hub v0.111.0: notice a deletion within a day (R-431); correct R-429; re-scope R-95
gates / gates (push) Successful in 17s
gates / gates (push) Successful in 17s
THE RECORD WAS TELLING A WORSE STORY THAN THE TRUTH FOR TWO MONTHS, and my own probe is why.
R-429 CORRECTED. Yesterday's spike searched for a directory called `.snapshots`. The vendor documents
the path as /.zfs/snapshot. The probe's CONTROLS were sound and its SUBJECT was wrong, so "not found"
was true and meant nothing. Re-probed at the documented path on both boxes, with controls:
- /.zfs lists (shares, snapshot) from inside the jail;
- a write into /.zfs/snapshot is REFUSED - dest open ...: Failure - while the identical write to
the account home SUCCEEDS and was cleaned up.
That is the append-only property PROVEN rather than cited, and it is the sentence the whole re-scope
rests on. Seven daily snapshots are confirmed in the panel. The mitigation works. What remains true,
and was always the actual finding: the row claiming it had no R-number, its "confirm tomorrow" went
36 days unanswered, and the DUE-CHECKS block built for that class was empty. THE FINDING WAS NEVER THE
SNAPSHOTS - IT WAS THAT NOBODY COULD TELL.
R-95 RE-SCOPED: the box can delete its LIVE repository but cannot write to the daily snapshots of it,
so a deletion costs at most one day plus a per-file recovery - not open-ended loss. The ranking is
Viktor's; it has been #1 since July on the old story.
R-432 FILED: a sub-account sees /.zfs/snapshot EMPTY while the same box holds seven snapshots, so
per-file recovery is operator-only today. One panel read settles whether a NAMED snapshot can still be
entered, which would make it product-reachable.
R-431 SHIPPED. Third signal in OffsiteChecker. On the hub deliberately: a detector on the box is one
the deletion can silence. Threshold REASONED, not invented - over 12 898 reports every decrease lands
on ZERO and predates stats_known, and in the stats_known window there are none, so observed churn gave
nothing to calibrate against. Retention cannot halve a total; a mass deletion goes to ~0. Hence: more
than half, and at least 5. Guarded by StatsKnown (R-331), the declared State (R-204) and run success
(R-100 - whose lesson lives in this very file).
ACCEPTANCE: 9 009 real report points replayed through the detector produced ZERO alarms.
Three red-proofs run. The escalation one only became real after the first version was found HOLLOW -
it re-swept the same report, so the baseline had already moved and the latch was never consulted.
07 row 10's status is NOT moved: the write-refusal is measured, but the recovery ROUTE has never been
walked, which is what PARTIAL means.
This commit is contained in:
@@ -210,3 +210,4 @@ likewise silent. Evidence: `audits/DRILL-r361-2026-08-22/evidence/06-part3-decis
|
||||
| **The hub reports LOSS OF VISIBILITY into either off-site store (not just how full it is)** | hub **v0.106.0** (R-339) | **IMPLEMENTED — deliberately NOT proven-live** | Both box checkers count consecutive failed fetch windows and emit `pbsdr_box_unreachable` / `offsite_box_unreachable` (severity `warning`) past a default 3 windows (≈30–45 min), each with a paired `*_recovered` all-clear routed via `recoveredPairedDownTypes` — required because the recoveries are severity `info`, which `severityNotifies` drops. Scopes stay customer-less (`pbsdr-box` / `pool-box`) → operator channel only. Fill logic untouched: a degraded read still drives no band transition. Evidence: `internal/monitor/box_reachability_test.go` + the cross-package wiring test in `internal/notify/`, which asserts an actual operator mail rather than a map entry. **Filed BECAUSE of a measured gap**, not a hypothesis: the 2026-08-18 ep0 outage ran 9 h 37 m with the hub silent | **The gap that remains is R-340**, and it is not small: the ep0 read is the `usage` op, which rides the LOCAL API daemon — the daemon that incident explicitly cleared — so this check would have shown GREEN for that entire outage. It closes "ep0 is unreachable as a host"; it does not close what actually happened. **No live or constructed outage has exercised the emit path**, and one cannot be manufactured against ep0 (Tier 2, protected) |
|
||||
| Secrets hygiene: bearer in k8s Secret, no secrets in git, single-quote credential store | hub v0.53, conventions | **IMPLEMENTED** | 07-13 closing bundle | |
|
||||
| Operator login password changeable from UI | hub v0.54 | **IMPLEMENTED** | 07-13 | |
|
||||
| **A deletion of a customer's off-site history is NOTICED within a day** | hub **v0.111.0** (R-431) | **IMPLEMENTED — not yet PROVEN-LIVE** | 09-01 | `hub/internal/monitor/offsite.go` — third signal beside FILL and STALENESS. **On the hub deliberately:** a detector on the box is one the deletion can silence. Alarms when the reported count falls by more than HALF and by at least 5, guarded by `StatsKnown` (R-331), the declared `State` (R-204) and run success (R-100). **Threshold reasoned, not invented:** over 12 898 reports every decrease lands on ZERO and predates `stats_known`; in the 380-report `stats_known` window there are none. **ACCEPTANCE: 9 009 real points replayed → ZERO alarms** (`offsite_r431_test.go`, fixture committed). **What PROVEN-LIVE would need and this does NOT have:** a real drop observed on a live box producing a real mail — the live firing done at ship time was driven through the hub's own path with synthetic counts, which is an end-to-end delivery proof, not a proof that a genuine deletion is caught. |
|
||||
@@ -885,7 +885,7 @@ crosses the line — **R-158**.
|
||||
| 7 | **Guest stopped and does not come back** | everything | guest-power watchdog (60 s, `onboot` as the deliberate-stop discriminator) | automatic | **120 s** | | **PROVEN** | agent v0.107.0 replay — 120 s unattended vs the incident's 587 s with a human |
|
||||
| 8 | **Host dies (hardware), drives intact** | the data drives; ep0's PBS namespace; the hub's Recipe + Escrow | **install a new host, then** `--selftest=bring-up -mode dr` per guest, **then** re-attach drives by `durable_id` | **operator** (SSH) | | 7 d | **IMPLEMENTED** | bring-up code exists and has **never been executed** (`CAMPAIGN-8…:522`); the drive half of the plan is empty on every box (**R-105**) |
|
||||
| 9 | **Whole box lost (fire/theft) — host and drives gone** | ep0 PBS namespace; the restic repo; the hub's Recipe + Escrow | new hardware → day-0 → escrow-consume with **R** → restore guests from PBS → app data from Tier-3 | **operator + customer** (R) | | 7 d (guest) · 24 h (app data) | **IMPLEMENTED / UNPROVEN** | every leg exists; the composed path has never been run. The destructive S5 drill is operator-gated and unrun (`06-offsite-connectivity.md:327`) |
|
||||
| 10 | **Ransomware / malicious deletion inside the guest** | PBS offsite (the box **cannot** delete its own snapshots); **the restic repo is NOT protected the same way** | whole-guest restore from PBS to a point before the event | **operator** (SSH) | | 7 d | **PARTIAL** | R-89 proved the box is refused when deleting its own PBS snapshot (`CAMPAIGN-8…:514-515`). **R-95 (open, ranked #1):** the restic credential **can delete** — `readonly=False`, `forget --prune` runs from the box, and SFTP cannot express append-only (`OPEN-ITEMS.md:13`) **2026-08-30 (R-359): the store is now VERIFIED on a cadence** — a daily `offsite-integrity` job runs `restic check` when the last successful one is over 7 days old. **This is a readability check, not a restore-test (R-87 stays open).** The depth that ships ON now re-reads **100%** of the pack data (R-399 CLOSED, controller v0.228.0) — the sentence here previously said the opposite and was stale. **2026-08-31 (SPIKE R-87):** a scratch restore of every app on demo-hp was measured at **25 s for 8 snapshots / 774 MB**, against 40.3 s for one weekly check — but restic 0.14.0's `--verify` checks size and mtime, **not content**, so nothing available today can vouch for the restored BYTES. Row 4's verdict is UNCHANGED by the spike. `audits/SPIKE-restic-restore-test-2026-08-31.md`. **2026-09-01 (SPIKE R-95, `audits/SPIKE-r95-offsite-delete-2026-09-01.md`) — THIS ROW'S STATUS IS UNCHANGED, but the seven-day bound beside it is now known to be unverified.** Measured on BOTH boxes over their own SFTP credential, with controls: **no `.snapshots` is visible to either sub-account**, and the account is jailed. Either none exist or a sub-account cannot see them — and a snapshot the box cannot see is one the box cannot restore from, so recovery would be an operator act at the Hetzner panel, not a product capability (R-429; the confirming field needs the provider API, fenced by §11-D). **The PBS shape does NOT transfer**: the sub-account API has one permission axis, `readonly`, and a backup target cannot be read-only — PBS is a server that can refuse, a Storage Box is a filesystem that runs nothing. **What the spike removed as a fear:** withdrawing delete does NOT wedge the store (measured — restic treats a dead owner's lock as stale and proceeds). **What it added:** `unlock --remove-all` reports success while deleting nothing (R-430). |
|
||||
| 10 | **Ransomware / malicious deletion inside the guest** | PBS offsite (the box **cannot** delete its own snapshots); **the restic repo is NOT protected the same way** | whole-guest restore from PBS to a point before the event | **operator** (SSH) | | 7 d | **PARTIAL** | R-89 proved the box is refused when deleting its own PBS snapshot (`CAMPAIGN-8…:514-515`). **R-95 (open, ranked #1):** the restic credential **can delete** — `readonly=False`, `forget --prune` runs from the box, and SFTP cannot express append-only (`OPEN-ITEMS.md:13`) **2026-08-30 (R-359): the store is now VERIFIED on a cadence** — a daily `offsite-integrity` job runs `restic check` when the last successful one is over 7 days old. **This is a readability check, not a restore-test (R-87 stays open).** The depth that ships ON now re-reads **100%** of the pack data (R-399 CLOSED, controller v0.228.0) — the sentence here previously said the opposite and was stale. **2026-08-31 (SPIKE R-87):** a scratch restore of every app on demo-hp was measured at **25 s for 8 snapshots / 774 MB**, against 40.3 s for one weekly check — but restic 0.14.0's `--verify` checks size and mtime, **not content**, so nothing available today can vouch for the restored BYTES. Row 4's verdict is UNCHANGED by the spike. `audits/SPIKE-restic-restore-test-2026-08-31.md`. **2026-09-01 (SPIKE R-95, `audits/SPIKE-r95-offsite-delete-2026-09-01.md`) — THIS ROW'S STATUS IS UNCHANGED, but the seven-day bound beside it is now known to be unverified.** Measured on BOTH boxes over their own SFTP credential, with controls: **no `.snapshots` is visible to either sub-account**, and the account is jailed. Either none exist or a sub-account cannot see them — and a snapshot the box cannot see is one the box cannot restore from, so recovery would be an operator act at the Hetzner panel, not a product capability (R-429; the confirming field needs the provider API, fenced by §11-D). **The PBS shape does NOT transfer**: the sub-account API has one permission axis, `readonly`, and a backup target cannot be read-only — PBS is a server that can refuse, a Storage Box is a filesystem that runs nothing. **What the spike removed as a fear:** withdrawing delete does NOT wedge the store (measured — restic treats a dead owner's lock as stale and proceeds). **What it added:** `unlock --remove-all` reports success while deleting nothing (R-430). **2026-09-01, LATER THE SAME DAY — THE PROBE ABOVE LOOKED FOR THE WRONG NAME AND THIS ROW'S SECOND CLAUSE IS NOW HALF WRONG.** It searched `.snapshots`; the vendor documents `/.zfs/snapshot`. Re-probed at the documented path on BOTH boxes, with controls, the two halves separate cleanly: **(a) the LIVE repository is deletable by the box — unchanged, R-95 stands;** **(b) the daily SNAPSHOTS of it are not writable by anything — PROVEN, not cited:** a write into `/.zfs/snapshot` is refused (`dest open …: Failure`) while the identical write to the account home succeeds. Seven daily snapshots confirmed in the panel (R-429). **So this row's "the restic repo is NOT protected the same way" is true of the repository and FALSE of its snapshots** — a deletion costs at most the day since the last snapshot, recoverable per-file (vendor). **Two limits kept honest:** a sub-account sees the snapshot directory EMPTY, so per-file recovery is operator-only today (R-432); and a panel-driven restore rolls back the WHOLE box. **The status is NOT moved** — evidence (b) is measured, but the recovery ROUTE has never been walked, which is what PARTIAL means. **Detection shipped hub v0.111.0 (R-431):** an unexplained fall in the snapshot count is noticed within a day. |
|
||||
| 11 | **Hub lost** | every box's data plane, every tier, every Lane-1 route | none needed for recovery **of a box**; the hub itself restores from its Longhorn volume backup | operator (`kubectl`) | | 24 h + weekly (Longhorn `RETAIN 1` each) | **UNPROVEN** | a hub restore has never been performed. The backup target is `nfs://192.168.0.180` — **DooPlex itself** — and exactly **2** restore points exist (LIVE, INV Part D2.2) |
|
||||
| 11b | *consequences while the hub is gone* | — | — | — | | | **[FACT]** | **day:** nothing customer-visible breaks; events queue (`settings.go:1466-1485`). **week:** the operator alarm plane is dark, no claim/reset codes, no config or floor convergence, no PBS-secret re-issue. **permanently:** escrow custody and break-glass credentials are gone (INV Part D2.4) |
|
||||
| 12 | **Offsite provider lost (Hetzner)** | everything on-premises: both drives, both whole-guest tiers | none needed — on-premises recovery is unaffected. Re-provision a new offsite target. | operator | | | **[FACT]** | **restic and PBS share the provider** (INV Part E.1). Whether they share an account and payment method is **UNKNOWN** → §11-D |
|
||||
@@ -1085,7 +1085,7 @@ does **not** hold as written. → **R-108**
|
||||
| **R-356** | The off-site restore resolved its destination with the raw `HDD_PATH` and read an empty answer as "not installed" — **CLOSED, controller v0.219.0, 2026-08-22** | 40 of 53 apps were refused permanently while running (§6.3 `[DESIGN]`) |
|
||||
| ~~**R-108**~~ | ~~Network storage can host an app's namespace~~ | **CLOSED 2026-07-30, controller v0.187.0 — D5 UNBLOCKED.** An app namespace may no longer be placed on network storage (5 surfaces guarded by one fail-closed predicate); the share-root bind is deliberately UNCHANGED because it is load-bearing and unscopable (§10.1). `audits/R108-network-app-namespace-2026-07-30.md` |
|
||||
| **R-126** | A `.fab` bundle — plaintext secrets, optional password — can be exported ONTO a NAS: `storageDriveList()` (`internal/web/handler_export.go`) does not filter network paths | split out of R-108, which closed without it. NOT a D5 precondition: an explicit customer-chosen export destination, not a browsing surface reaching a backup tree (§5, §7.3) |
|
||||
| R-95 (open) | The restic offsite credential **can delete** — the box can `forget --prune` its own repo, from **two** call sites (`offbox.go:1388` retention and `offbox.go:1759` over-quota) | the tier holding the customer's documents and photos is the one whose credential can destroy it (matrix row 10). **SPIKE 2026-09-01:** prevention needs a transport change (restic 0.14.0 does speak `rest:` — measured; append-only is a rest-server flag, not a restic one), because the sub-account API cannot express write-without-delete. **Detection is nearly free and is recommended first** — `snapshot_count` already reaches the hub and the hub appends reports, so the comparison needs no box change. `audits/SPIKE-r95-offsite-delete-2026-09-01.md` |
|
||||
| R-95 (open) | The restic offsite credential **can delete** — the box can `forget --prune` its own repo, from **two** call sites (`offbox.go:1388` retention and `offbox.go:1759` over-quota) | the tier holding the customer's documents and photos is the one whose credential can destroy it (matrix row 10). **SPIKE 2026-09-01:** prevention needs a transport change (restic 0.14.0 does speak `rest:` — measured; append-only is a rest-server flag, not a restic one), because the sub-account API cannot express write-without-delete. **Detection is nearly free and is recommended first** — `snapshot_count` already reaches the hub and the hub appends reports, so the comparison needs no box change. `audits/SPIKE-r95-offsite-delete-2026-09-01.md`. **RE-SCOPED 2026-09-01:** the box can delete the LIVE repository but **cannot write to the daily snapshots of it** (measured, both boxes) — so the exposure is at most one day's data plus an operator-driven per-file recovery, not open-ended loss. Detection shipped hub v0.111.0 (R-431). |
|
||||
| ~~R-86~~ | ~~Restore-tests are interval-scheduled, not backup-aligned~~ | **CLOSED 2026-08-03 — agent v0.121.0 + hub v0.91.0.** Restore-testing is now **per archive generation**: a tier is due when its newest archive that has settled ~24 h has not been proven, so a daily tier is proved daily on its own archive and a weekly tier weekly on its own. The ticker survives only as the evaluation interval (6 h, chosen from a measured cost). The hub's staleness window moved with it — per tier, from that tier's observed archive rhythm — because a weekly tier proved weekly sat EXACTLY on the old flat 7-day line (§3, Lane 2's per-archive rule) |
|
||||
| ~~R-353~~ | ~~A local unit restore reported a bare completion whether it returned an entire dataset or nothing~~ | **CLOSED 2026-08-30, controller v0.226.0.** The volume count already existed (`restoreDockerVolumesFrom`) and was discarded by a one-line wrapper, so the surface was structurally unable to say what came back. `RestoreFromRecoveryUnit` now returns `UnitRestoreResult` and the sentence has three cases, keyed on replayed-vs-**listed** — zero-replayed has two causes that are opposite news. Every sentence is a claim about the BACKUP, never the app: this path has no `SafetyDump` discriminator, and §6.3 is why that is not pedantry. Proven live on `demo-hp` |
|
||||
| ~~R-357~~ | ~~The destructive reconstitute had no free-space gate; all three that existed guarded non-destructive paths~~ | **CLOSED 2026-08-30, controller v0.226.0.** The gate sits before `writeSafetyDump` and `StopStack`, so a refusal costs no outage — the test asserts the StopStack call count, not the sentence. No headroom multiplier (a measured tree, not a predicted download); fail-closed when either probe reads ≤ 0, which was a real fail-open hole. **Not live-validated** — filling a filesystem is a drill step |
|
||||
|
||||
@@ -26,6 +26,8 @@
|
||||
|
||||
---
|
||||
|
||||
| **R-429** | **CORRECTED 2026-09-01 — the snapshots ARE being taken; what was broken was that nobody could tell, and my own probe looked for the wrong name.** Viktor read the control panel on 2026-09-01: **seven automatic daily snapshots** on `storage-box-pool-1` (plan BX11), six days old to ~9 h old, filesystem ~2.5 GB, per-snapshot 0–14 MB, *Display snapshot directory* ON. **The mitigation works.** **MY ERROR, NAMED:** the 2026-09-01 spike probed for a directory called `.snapshots`; the vendor documents the path as **`/.zfs/snapshot`**. The probe's controls were sound and its subject was wrong, so `not found` was true and meant nothing. **The task that set the spike asserted the mechanism without citing the vendor documentation, and I did not check it** — that is how a correct instrument produced a wrong headline. **WHAT REMAINS TRUE, and is the actual finding:** the row claiming it had no R-number so nothing could cite it; its "confirm tomorrow" went **36 days** unanswered; and the `DUE-CHECKS` block built for that exact class (R-341) was **empty**. **The finding was never the snapshots. It was that nobody could tell.** Re-probed at the documented path — see R-432 for what a sub-account can actually reach. Evidence: `audits/SPIKE-r95-offsite-delete-2026-09-01.md` §Q1 and this row. | **CLOSED 2026-09-01 — mitigation CONFIRMED WORKING; the visibility gap is R-432** | panel read 2026-09-01 (Viktor); re-probe at `/.zfs/snapshot` in `audits/SPIKE-r95-offsite-delete-2026-09-01.md` |
|
||||
| **R-431** | **An unexplained fall in a customer's off-site snapshot count is now noticed within a day — SHIPPED hub v0.111.0.** Third signal in `OffsiteChecker`, beside FILL and STALENESS. **It lives on the HUB deliberately:** the event being detected is a box deleting its own backups, so a detector on that box is one the same event can silence; the hub already receives the count and keeps the history. **Threshold: a fall of more than HALF the previous count and at least 5** — reasoned, not invented, because the measurement gave nothing to calibrate against: over 12 898 reports (2026-06-05 → 2026-09-01) every one of the nine decreases lands exactly on ZERO and every one predates `stats_known` (the R-331 shape), and in the 380-report window where `stats_known` is true there are **zero** decreases. Retention keeps 7 daily + 4 weekly + 6 monthly per group, so it **cannot halve a total**; a mass deletion goes to ~0. **Three pre-conditions, each with its scar:** `StatsKnown` (R-331 — a zero is not a zero when unmeasured), the declared `State` (R-204 — the box names its own situation), and run success (R-100 — presence is not success; `incomplete` excluded too). An untrustworthy report neither alarms nor moves the baseline. **ACCEPTANCE: 9 009 real report points replayed through the detector produced ZERO alarms**, fixture committed. Severity `error`, operator-only, no customer template. | **CLOSED 2026-09-01 — SHIPPED hub v0.111.0** | `hub/internal/monitor/offsite.go`; `offsite_r431_test.go` incl. 9 009-point real-history replay; hub v0.111.0 |
|
||||
| **R-419** | **`observations_gate.py` accepted an observation whose body merely CONTAINED the string `NOT-A-FINDING`, even in prose disclaiming it.** Found by accident on 2026-09-01 when a planted test observation reading *"it carries no `FILED:` and no `NOT-A-FINDING:` marker"* was reported `OK 1. NOT-A-FINDING` and a real push went green over an unfiled finding. The gate's whole job is to force an explicit choice, and a sentence disclaiming the choice counted as making it. **FIXED 2026-09-01:** a marker must now start a line or follow a sentence boundary, and inline code spans are stripped before matching — a marker inside backticks is being talked about, never used. | **CLOSED — FIXED + PINNED** (2026-09-01, R-421 sweep) | verified in BOTH directions: the decoy and a backticked mention are convicted; a real `**FILED: R-419**` and a real `**NOT-A-FINDING: ...**` still pass. Decoy kept in `felhom.eu/scripts/test_gate_decoys.py` |
|
||||
| **R-404** | **DECISION — should a documents-only push be subject to the golden-currency gate? RULED 2026-09-01: NEITHER option as framed. Block the push that can create the debt; notify the push that cannot.** The two options on the table were *narrow the gate* and *leave it and build a waiver*, and both were wrong for the same reason: they argued about the GATE, and the gate was never the problem. **The DIAGNOSIS, measured from live source, is that the check was aimed at the wrong repository.** `golden_currency_gate.py` never looks at the push at all — it compares the controller's newest CHANGELOG heading against this repo's bake evidence and returns the same verdict whatever you are pushing, which is correct for a standing invariant and wrong as a push gate. Meanwhile `controller_gates.py` had NO golden-currency entry, so **the repo where a release happens never checked, and the repo that cannot create the debt enforced it on every push.** 18 of the last 24 pushes here touched no code — MEASURED, and the classifier agrees exactly — most of them for a structural reason: the controller's code is in one repo and its register, architecture and status live in this one, so **every controller change produces a documents-only push here by construction.** Six of those 18 were bake records — **the push that PAYS the debt is itself documents-only, so the gate was blocking its own cure.** **WHY NOT THE WAIVER** the gate's own docstring prescribes: that clause was written for *a release nobody wants a golden for*. The case that actually occurred (R-417) was *a release we did want a golden for, on a night the runbook forbade baking*. A waiver would have recorded a lie. **SHIPPED:** `scripts/push_scope.py` (allow-list; every uncertainty answers `code`), a fifth `exemptible` field in `repo_gates.py` + `--scope`, a new **ADVISORY** verdict printed in its own block, the pre-push hook reading git's stdin, the same rule in CI from the push event payload, and `felhom-controller/controller/scripts/golden_notice.py` — a NON-BLOCKING notice at the moment a release is committed. **The gate's own logic, exit codes and wording are byte-identical**; only the consequence changed. The exemption is ONE gate wide and `TestR3`/Scenario C pins it, red-proved by widening it. Proven live on the real hook: docs+debt → ADVISORY, pushed; code+debt → refused; docs+debt+a second gate → refused for that gate alone | **CLOSED 2026-09-01** — ruled and shipped | full text and the ruling: `git show 1e6c387:documentation/backlog/CLOSED-ITEMS.md`; the diagnosis is in `CONTEXT.md` and `scripts/CHANGELOG.md`; `scripts/push_scope.py` + `test_repo_gates_scope.py` |
|
||||
| **R-417** | **A drill night that forbids baking a golden made `golden_currency_gate.py` red, so pushing the drill's own evidence needed `--no-verify` — the very signal CI e-mails about.** Measured 2026-09-01: five consecutive felhom.eu CI runs red (jobs 469/470/471/473/476), all mine, all on step 3 `Run the gate entry point`; job 478 green the moment the golden-0.232.0 evidence was committed. Cause confirmed by isolation — moving that directory aside reproduces exit=1, restoring it gives exit=0. **The gate was right every time**: 0.231.0 and 0.232.0 were released with no golden carrying them. **CAUSE REMOVED, not worked around** (R-404): a drill's pushes are documents-only, so the conviction now prints as a loud ADVISORY and the push proceeds — in the hook AND in CI, so a drill night no longer produces red runs indistinguishable from real ones. The expectation is now written where the next drill author reads it (`documentation/runbooks/target-selection.md`), which is the half I had left out. | **CLOSED 2026-09-01** — by R-404 | five red CI jobs 469/470/471/473/476, green at 478; reproduced by isolation (`documentation/audits/AUDIT-gate-decoys-2026-09-01.md` records the technique) |
|
||||
|
||||
@@ -224,7 +224,7 @@ unconditional promise, which CAMPAIGN-11 §7 step 7 measured the customer-facing
|
||||
| **E-2d** | **Prove E-2 on a fresh VM** — a real `felhom-host-install.sh` 1.22.0 run, Case B naturally, a claimable customer, then add a drive (the offer) and unplug it (`backup_target_absent` end-to-end) | **CLOSED — PARTIALLY PROVEN** (2026-07-29) | — | **C1, C2 proven** (`audits/E2D-fresh-vm-2026-07-29.md`); **C3, C4 proven live** (`audits/SESSION-C-2026-07-29.md`); **C5 FAILED → R-116** — the gate fires and an alarm reaches the hub, but it is the generic event, so the alarm and its recovery cannot be paired. **R-116 is the single named open leg**; per the Session-C runbook §9, decided in advance, a failed claim closes the item as partially proven rather than triggering a re-run. Both audits carry the full record — the `local-lvm` fence, the ISO/PAIRING derivation, the Phase 0 answers, the per-claim observables and the teardown evidence — and are the place to read it, not this cell. **The arc's actual definition of done is R-106 + R-109, R-108 and D5**, none of which this detour touched | CC |
|
||||
| **R-121** | **A BOX's installed agent can sit releases behind the vouched one and nothing notices — the R-120 gate does not cover it.** demo-hp ran agent **0.113.0** while the hub vouched **0.116.0**, through the whole R-116/R-117 arc, and no signal existed on any channel | **READY (S) — NEW 2026-07-30** | — | **Fourth instance of the drift family** (R-111 golden's agent 17 releases behind, R-115 built+deployed but never published, R-120 golden a controller behind — and now installed-vs-vouched on a live box). **Confirmed at source that R-120's gate cannot catch it:** `hub/internal/web/configs.go:1165-1169` compares `goldenVer` against `store.NewestReportedControllerVersion()` — it is a **golden-artifact vs fleet-CONTROLLER** check and says nothing about the agent installed on a box. **`MinAgent` does not cover it either:** it is used to HOLD the controller floor for a box whose agent is too old (`hub/internal/api/handler.go:530-538`, `store.go:1857`) — protective, not an alarm — and demo-hp's 0.113.0 **equalled** `min_agent` 0.113.0, so even a floor comparison was satisfied. **The cost, measured:** R-117's whole subject is the R-113 conjunction, which landed in **0.114.0** — so the designated drill host could not exercise the code under investigation at all, and the R-117 spike had to route every predicate result through an out-of-repo probe built from `main` instead of the installed agent (`audits/SPIKE-r117-bind-liveness-2026-07-30.md` §1, §2.3). Discovered because the R-117 task made bringing the box current an explicit prerequisite. **Fix shape (not implemented):** the hub already receives `AgentVersion` on every host report, and already has semver comparison in Go — the missing piece is a checker comparing reported agent vs the **vouched** agent and surfacing it, operator-tier. Note the honest tension: a box legitimately lags between publish and deploy, so this wants a staleness window rather than an instant alarm | CC |
|
||||
| **R-118** | **An absent drive's union row advertises the ROOT filesystem's capacity as its own.** In the absent-state payload the registry-union row reports `total_bytes: 49675956224 / used_bytes: 4584579072` — **byte-identical to the `local` row** (`durable_id: path:/var/lib/vz`, i.e. `pve-root`) in the same response. The real drive is **4 GB** | **READY (XS) — NEW 2026-07-30** | — | Cause: `statfsCapacity(d.MountPath)` (`disks.go:335-338`) statfs's `/mnt/cel`, which with the device gone is a **bare directory on the root filesystem**. `observe.go:176-183`'s comment warns about exactly this trap and guards the Observe path ("*an unmounted removable dir-storage's mountpoint reverts to a bare directory on root … catastrophic DR mis-id*"); **the union path has no equivalent guard.** **Not a DR mis-id** — `durable_id` on that row is still the correct `uuid:…`, so re-attach identity is safe. It is a **false capacity** reaching every consumer of `total_bytes`/`used_fraction` (fill monitors, storage cards): a detached 4 GB drive advertises 46 GiB at 9.2 % used. Same class as `role.go:180-181` — an absent drive's fields decaying to the root filesystem's. Evidence: `audits/DIAG-r116-disks-payload-2026-07-30.md` §12 | CC |
|
||||
| **R-95** | restic offsite credential **can delete** (`readonly=False`, `forget --prune` runs from the box); SFTP cannot express append-only | **READY** | — | Root exposure still open. Mitigation now ARMED — split prune off-box or move to REST `--append-only` | CC **SPIKE 2026-09-01 — `audits/SPIKE-r95-offsite-delete-2026-09-01.md`. THE WORD "ARMED" ABOVE IS NOT SUPPORTED AND IS WITHDRAWN PENDING R-429:** no `.snapshots` is visible to either box's sub-account (measured, both machines, with controls), so the seven-day bound is unverified and unverifiable from the product side. **Q3 (documented, `hub/internal/hetznerapi/hetznerapi.go:38-45`): the sub-account API has ONE permission axis, `readonly` — there is no append-only, so the PBS shape does NOT transfer** (PBS is a server that can refuse; a Storage Box is a filesystem that runs nothing). **Q5 (measured): withdrawing delete does NOT wedge the store** — restic treats a dead owner's lock as stale and proceeds — so the constraint everyone feared is not the blocker; but `unlock --remove-all` lies about success (R-430), and the crash-lock window is UNKNOWN. **Q6 (measured): restic 0.14.0 DOES speak `rest:`** (control: `banana:` → invalid backend), and append-only is a rest-server flag, not a restic one — reachable, but it needs a machine in the recovery path and ep0 is protected. **Q7 (measured): detection is nearly free** — `snapshot_count` already reaches the hub and the hub APPENDS reports, so the comparison needs no box change. **RECOMMENDATION: answer R-429 first (Viktor, ten minutes), then build detection, then move retention off the box; defer the transport change.** Two `forget --prune` sites must be disarmed together — `offbox.go:1388` AND `offbox.go:1759` — or R-191 repeats. |
|
||||
| **R-95** | restic offsite credential **can delete** (`readonly=False`, `forget --prune` runs from the box); SFTP cannot express append-only | **READY** | — | Root exposure still open. Mitigation now ARMED — split prune off-box or move to REST `--append-only` | CC **SPIKE 2026-09-01 — `audits/SPIKE-r95-offsite-delete-2026-09-01.md`. THE WORD "ARMED" ABOVE IS NOT SUPPORTED AND IS WITHDRAWN PENDING R-429:** no `.snapshots` is visible to either box's sub-account (measured, both machines, with controls), so the seven-day bound is unverified and unverifiable from the product side. **Q3 (documented, `hub/internal/hetznerapi/hetznerapi.go:38-45`): the sub-account API has ONE permission axis, `readonly` — there is no append-only, so the PBS shape does NOT transfer** (PBS is a server that can refuse; a Storage Box is a filesystem that runs nothing). **Q5 (measured): withdrawing delete does NOT wedge the store** — restic treats a dead owner's lock as stale and proceeds — so the constraint everyone feared is not the blocker; but `unlock --remove-all` lies about success (R-430), and the crash-lock window is UNKNOWN. **Q6 (measured): restic 0.14.0 DOES speak `rest:`** (control: `banana:` → invalid backend), and append-only is a rest-server flag, not a restic one — reachable, but it needs a machine in the recovery path and ep0 is protected. **Q7 (measured): detection is nearly free** — `snapshot_count` already reaches the hub and the hub APPENDS reports, so the comparison needs no box change. **RECOMMENDATION: answer R-429 first (Viktor, ten minutes), then build detection, then move retention off the box; defer the transport change.** Two `forget --prune` sites must be disarmed together — `offbox.go:1388` AND `offbox.go:1759` — or R-191 repeats. **RE-SCOPED 2026-09-01 — THE STORY WAS WORSE THAN THE TRUTH FOR TWO MONTHS.** The box can delete its own LIVE repository, **but it cannot write to the daily snapshots of it** — MEASURED, not cited: a write into `/.zfs/snapshot` is refused on both boxes while the same write to the account home succeeds (R-432). Seven daily snapshots are confirmed in the panel (R-429). **So a deletion costs at most the data written since the last daily snapshot, and the rest is recoverable — file by file, one customer at a time, with no effect on anyone else** (vendor: *"You can download individual files or entire directories as usual"*; *"It is not possible to write to the `/.zfs` directory or its subfolder"*). **NOT open-ended loss.** Two caveats kept honest: a panel-driven snapshot restore rolls back the WHOLE Storage Box and deletes newer snapshots, which is why the per-file route matters; and per-file recovery is **operator-only today** (R-432). **DETECTION SHIPPED hub v0.111.0 (R-431)** — an unexplained fall is noticed within a day. **THE RANKING IS VIKTOR'S:** this has been #1 since July on the old story. On the new facts I would rank it below the items that can still lose data outright, but I am not re-ranking it myself. `audits/SPIKE-r95-offsite-delete-2026-09-01.md` |
|
||||
| **R-191** | **Every weekly offsite backup UPLOADS successfully and then FAILS the job on a prune the box is deliberately not allowed to do — on both demo boxes.** Measured on demo-felhom 2026-08-04 06:49–06:53: the upload completed (223 s, 629 MiB of 1.874 GiB, **67.2 % reused incrementally**), then `ERROR: prune 'ct/9201': proxmox-backup-client failed: Error: permission check failed - missing Datastore.Modify\|Datastore.Prune on /datastore/felhom-offsite/demo-felhom` → `ERROR: Backup of VM 9201 failed - error pruning backups` → `TASK ERROR: job errors`. The hub raised `whole_guest_backup_failed` | **CLOSED — SHIPPED 2026-08-04** (installer **1.25.0**; both live boxes corrected) | — | **This is R-89's rule not reaching the config.** R-89 moved PBS pruning SERVER-SIDE — *"boxes set `keep_last: 0`, ep0 runs prune jobs; box tokens stay write-only, never widen the grant"*. The token behaves exactly as designed: it refuses. But **both** demo boxes still arm the offsite tier with `keep_last=2 prune_pbs_allowed=true` (`backup_targets: [{target_id: felhom-pbs, cadence_seconds: 604800, keep_last: 2}]`), so every run asks for a prune that must fail. **The data is SAFE and that is why this is not a P1:** the snapshot lands before the prune is attempted; what is wrong is the job's VERDICT and the weekly operator e-mail it produces. **But it is corrosive in the specific way this project keeps finding:** a backup that reports FAILED while succeeding trains the operator to discount `whole_guest_backup_failed`, which is the same alert that would carry a real one — and it is exactly the failure the R-100 corollary warns about, an alarm whose text is true and whose trigger is not the thing you would act on. **Fix is one config line per box** (`keep_last: 0` on the PBS tier) plus whatever writes it on a fresh install; **deliberately NOT applied in this session** — the session was a runbook with an explicit "change nothing, and if a change appears necessary, stop and report" rule, and a retention field on a live backup tier is not a change to slip into an observation run. **Check before fixing:** whether ep0's prune jobs actually cover these two namespaces, or the snapshots simply accumulate once the box stops asking **THE GATE WAS RUN FIRST, AND IT MATTERED.** Before disabling anything, ep0 was read (read-only, Tier 2): prune jobs `prune-demo-felhom` and `prune-demo-hp` exist on datastore `felhom-offsite`, one per namespace, `schedule 03:30`, `keep-last 2`, comment *"R-82 retention keep-last=2, server-side (box tokens are write-only)"* — and they have run **every day since 2026-07-27: 18 tasks, all `status=OK`**. The newest task log reads `retention options: --ns demo-felhom --max-depth 0 --keep-last 2` / `keep ct/9201/2026-07-27…` / `keep ct/9201/2026-07-28…` / `TASK OK`. Retention happens, and it happens there. **A METHODOLOGICAL WARNING WORTH MORE THAN THE FIX.** Three separate queries said the OPPOSITE — *no prune jobs have ever run* — and **all three were broken instruments**: `worker-type` where the field is `worker_type`; the value `prune` where the worker type is `prunejob`; and `journalctl -u proxmox-backup` where the unit is `proxmox-backup-proxy`. A fourth reading (3 snapshots under keep-last 2) was mis-framed by CC and self-corrected — the third snapshot had landed AFTER that day's 03:30 window. Acting on any of them would have disabled the only pruning ATTEMPT while reporting that nothing prunes: a weekly false alarm traded for unbounded growth on the protected endpoint, invisible for months. **The gate is what caught it, and only because it demanded evidence rather than a verdict.** **Shipped:** installer **1.25.0** writes `keep_last: 0` on the offsite tier (the agent's existing guard `allowPBSPrune = !primary && keep_last > 0` already reads that as *never prune from the box* — no agent change), the justifying paragraph is rewritten to say where retention lives and cite R-89, and `hostinstall_gates.py` asserts it (red-proved: pinning `keep_last: 2` back fails the gate). **Both live boxes corrected in their own config** — `backup tier armed target=felhom-pbs … keep_last=0 … prune_pbs_allowed=false` on demo-felhom and demo-hp, with the local tier untouched at `keep_last=3`. Served over HTTPS at `1.25.0` with `"keep_last":0` in the served bytes. **STILL TO OBSERVE:** the next weekly offsite run completing OK end-to-end. The change removes the failing step; the *schedule* proving it is next week's event, and this row should carry that line when it happens. | CC |
|
||||
| **R-194** | **PVE's permission cache delays every grant-state verdict by an unknown amount, so "the agent can read it" and "the ACL exists" are not the same measurement.** Observed twice while validating R-190's self-repair on demo-felhom 2026-08-04: both ACL rows for `/storage/felhom-backup` were deleted, and `GET /access/permissions` continued to report `Datastore.AllocateSpace` present — for **~40 s** in one run and **~16 minutes** in another. During that window the capability probe reads healthy and the self-repair does not fire | **OPEN** | — | **Why it matters beyond the delay:** it puts a floor under how fast a lost grant can be noticed, it makes any single permission read a lagging indicator, and — the interesting part — **it is a candidate contributor to R-190's own timeline**: a grant removed at an unknown moment could keep working until a cache expiry, which is exactly the shape of *worked at 04:44, refused at 09:24*. That does not explain what removed it, but it may explain when the refusal SURFACED, and the two have been treated as the same instant. **Not a defect in our code** — it is PVE behaviour, and the mitigation already tolerates it (the repair fires on the next probe after the cache clears). **What is worth deciding:** whether the store-grant probe should ALSO consult the storage content listing as a second signal, since that appeared to reflect the loss immediately (`{"data":[]}` while the permission read still said present) — two signals disagreeing is itself information, and today only one of them is read | CC |
|
||||
| **R-200** | **The DR password-injection seam has a handler, a route and tests — and no form.** `POST /backup/offbox/inject-password` is routed (`controller/internal/web/server.go:510`) to `offboxInjectPasswordHandler` (`offbox_handlers.go:174-196`) → `InjectOffboxPassword` (`backup/offbox.go:541`). **No template in the repository contains that path or any form posting to it** (grep over `internal/web/templates/`: one unrelated hit, an XSS comment) | **PLUMBING COMPLETE** (controller **v0.196.0**); **the FORM is not built — still open** | — | **The tenth instance of this project's built-but-never-wired class, and the exact shape `CLAUDE.md` and `felhom.eu/CLAUDE.md`'s seam-wiring rule were written for:** handler tests that POST directly (`offbox_escrow_test.go:167,180`) prove nothing about reachability. To use the only implemented recovery seam today, a person must hand-craft an authenticated POST with a session cookie and CSRF token. **Note the layering while fixing it:** this form takes a **64-hex repo password**, not a recovery code (`offboxRepoPwPattern`, `offbox.go:543`) — they are different secrets at different layers, and the operator's 2026-08-04 ruling asks for a form that takes **R**. Build the R form and treat this one as the operator/DR fallback it was written as, but **ship it with a render test per branch of whatever gate it sits behind**. Source: `audits/RECON-offsite-dr-chain-2026-08-04.md` §3 link 9 **THE DIAGNOSTIC HALF IS DONE AND IT ANSWERED THE QUESTION.** `--recover-offsite-check` is a `docker exec` escape hatch in the shape of `--print-reset-code`: R on **STDIN** (never argv, never `ps`, never shell history, never a transcript), fetch+unseal via the agent, and a verdict of **two sha256 hashes**. **It compares and never installs** — the recovered password is not written to `offbox/repo_password`; a test asserts the data dir is byte-unchanged and its red-proof (adding the install call) fails it. Confirmed live: `repo_password` mtime still `2026-08-03 07:18:02` after the successful check at `2026-08-04 11:49`. **Exit codes are load-bearing** — `0` match, `2` a clean MISMATCH, `1` a step failed; "it failed" and "it worked and disagreed" must never share a status because only one is a finding about the system. A box with no local password reports distinctly (the rebuilt-box shape, where the next step is to INSTALL rather than compare). **WHAT IS NOT BUILT, deliberately:** no card, no form, no preview, no customer-facing text — building an interface on top of a chain nobody had walked is how the preceding three weeks went wrong. **What remains for this row:** link 9 (the recovered password placed so `WriteOffboxSecrets` keeps it) and the customer-facing shape the operator ruled on 2026-08-04 (yell → R form → preview → proceed), which is now priced against a chain that exists rather than one that is assumed **PART 0 SHIPPED 2026-08-04 (v0.196.0):** `--recover-offsite-install` is the sibling of the check — same fetch/unseal path, same STDIN discipline for R — and it **places** the recovered password via `InjectOffboxPassword`. The confirmation is a SECOND invocation (`--confirm-install`): without it, both hashes print and nothing is written, so the operator sees the comparison before a write is possible. Three outcomes named distinctly: **installed** (no local password — the rebuilt-box shape), **unchanged** (identical key, nothing written), **refused** (a DIFFERENT key present — installing would clobber the key the current repository is encrypted under; exit 2, no force offered). It re-reads the file after writing rather than trusting the call. **Red-proof observed:** removing the confirmation gate makes the dry run write the password. The R-persistence test carries a **positive control** (a planted copy found, then removed and not found). **NOT YET EXERCISED AGAINST A LIVE RECOVERY** — the R-201 drill halted before step 9, so this is unit-proven only. **What remains for this row:** the customer-facing shape the operator ruled on 2026-08-04 (yell → recovery-code form → preview → proceed) | CC |
|
||||
@@ -602,8 +602,8 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server`
|
||||
| **R-426** | **The decoy-coverage exemption list — 20 registered gates that ship WITHOUT a decoy test, each named.** `scripts/decoy_coverage_gate.py`'s `EXEMPT` map is debt, and this row owns it so it lives in the register and not only in a Python literal. **Four kinds:** (a) genuinely covered in the 2026-09-01 sweep but not yet moved into a suite — `hub-copy`, `instructions`, `docker-v`, `image-pins`; (b) blocked by an open hole and therefore un-assertable as rejecting — `site` (R-423), `one-register` (R-424), `offbox-rename` (R-425); (c) shared scripts whose decoy lives in `felhom.eu` and is counted there — `reuse-refs`, `instructions`, `observations` in the controller and agent runners; (d) **no plausible decoy constructed yet** — `hostinstall`, `wire-contract`, `due-checks`, `published`, `image-resolvable`, `volume-persistence`. Group (d) is the honest unknown: six gates whose soundness is UNTESTED, not established. **The list is green today and shrinks; a NEW gate with no decoy fails immediately.** | **OPEN — 20 names; group (d) is six untested gates** |
|
||||
| **R-427** | **`closed_register_gate.py` checks ONE direction only: an open word in a CLOSED row. The mirror — a CLOSED verdict on a row still sitting in `OPEN-ITEMS.md` — is unchecked, and there are TWELVE.** MEASURED 2026-09-01 during the decoy sweep, by reading the leading verdict of every open row with the gate's own predicate: **R-385, R-387, R-341, R-378, R-405, R-88a, R-88b** read unambiguously closed; **R-123, R-190, R-352** read `PARTLY CLOSED` / `MITIGATION SHIPPED` and almost certainly belong where they are. **The rows were NOT moved by this session** — telling a finished row from a partly-finished one is a judgement, and R-378 is itself the record of what happens when a machine makes that judgement on a substring (six still-open rows moved out of the register). **This is R-405's finding mirrored:** that row exists because R-87 sat in the CLOSED file while its state read READY, and the gate written for it looks only the way it was bitten. Fix: the same leading-verdict predicate applied to `OPEN-ITEMS.md`, reporting rather than convicting until the twelve are adjudicated by a person — a gate registered while twelve rows fail it would refuse every push. | **OPEN — 12 rows named; the adjudication is Viktor's, the gate is mine** |
|
||||
| **R-428** | **The decoy-coverage gate — written to catch instruments that match a NAME instead of a fact — identified a repository by its DIRECTORY NAME.** MEASURED on its own first CI run (felhom.eu job 490, 2026-09-01): `os.path.basename(root)` looked up in a `RUNNERS` map, and Gitea's act-runner checks the repo out into a directory called `hostexecutor`, so the gate reported *"unknown repo 'hostexecutor'"* and went INCONCLUSIVE. **The gate that hunts label-matching was matching a label, in the first ten lines of its own main loop, and it shipped that way.** FIXED the same day: it now identifies a repo by which registered runner FILE exists under the root, which is a fact. **Recorded rather than quietly patched because it is the strongest evidence in the sweep that this class is not a matter of carelessness** — it was written by a session that had spent the morning reading 29 gates for exactly this, with the four shapes on screen. Verified under a renamed directory before and after. | **CLOSED 2026-09-01 — fixed, and kept as the class's best example** |
|
||||
| **R-429** | **The Storage Box snapshot mitigation that R-95 calls "ARMED" has never been confirmed, cannot be seen from either box, and its row has no id.** MEASURED 2026-09-01 over the SFTP credential each box already holds, read-only, with positive and negative controls on both machines: **no `.snapshots` directory is visible to either sub-account** — not in the account home, not inside the repo path — while the account is jailed (`/` returns `Permission denied`) and a bogus path errors correctly. Two readings the box cannot distinguish: no snapshots exist, or they exist and a sub-account cannot see them. **Both are bad, and the second is not a reprieve — a snapshot the box cannot see is one the box cannot restore from, so recovery would be an operator act through the Hetzner panel, not a product capability.** The claim lives at `OPEN-ITEMS.md:233` as a row whose ID cell is a bare em-dash instead of an R-number, so no gate and no grep-the-register rule can cite it; its "confirm tomorrow" was **2026-07-27** (last touched in `72692e1`), **36 days** unconfirmed; and the `DUE-CHECKS` block built for exactly this (R-341) is **EMPTY**. **The confirming field, `size_snapshots`, is a Hetzner API field and §11-D fences it — so this spike stopped and left it for Viktor.** Fix: Viktor reads the Storage Box panel once (ten minutes) and the answer goes in this row; give the WATCHING row an id or delete it; and R-95's word "ARMED" must not stand until then. Evidence: `audits/SPIKE-r95-offsite-delete-2026-09-01.md` §Q1. | **OPEN — Q1 is Viktor's, and it re-ranks R-95** |
|
||||
| **R-430** | **`restic unlock --remove-all` printed `successfully removed locks` while the lock was still there.** MEASURED 2026-09-01 in a throwaway local repo (no live store touched), under a faithful append-only model — a sticky locks directory owned by root holding a root-owned lock, restic run as `nobody`; both controls passed first (create allowed, delete refused). The command reported success, returned, and `ls` showed the lock present. **`resticStep`'s crash-lock self-heal is built directly on this call** (`felhom-controller/controller/internal/backup/offbox.go:~768`), and its licence to escalate rests on the escalation actually working. **A self-heal that cannot fail is a self-heal that cannot be trusted** — this is this project's *"exit codes that lie"* class, in the one path that runs unattended against the customer's off-site history. It is harmless TODAY because the credential can delete and the removal really happens; it becomes load-bearing the moment delete is withdrawn, which is what R-95 is about. **Not yet established:** whether restic reports success because it removed zero locks by design, or because it did not check. Settling it: read restic 0.14.0's unlock source, or re-run with `--verbose`. | **OPEN — precondition on any R-95 build** |
|
||||
| **R-432** | **A customer's own sub-account can REACH the snapshot door and is REFUSED writes to it — but sees it EMPTY, so per-file recovery is not product-reachable.** MEASURED 2026-09-01 on BOTH live boxes, over the credential each already holds, with a positive and a negative control in the same run. **What is now PROVEN rather than cited:** `/.zfs` lists (`shares`, `snapshot`) from inside the jail; a write into `/.zfs/snapshot` is **REFUSED** — `dest open …: Failure` — while the identical write to the account home **succeeds** and was cleaned up. **That is the append-only property, measured, and it is the sentence the whole R-95 re-scope rests on.** **What is NOT available:** `/.zfs/snapshot` lists **empty** (link count 2) on both boxes, while the same Storage Box demonstrably holds seven snapshots — `storage-box-pool-1` IS `u629488` (`RUNBOOK-ep0-datastore-volume-2026-07-27.md:386`), the box these sub-accounts live on. So the contents are filtered from a sub-account. **CONSEQUENCE: recovery from a snapshot is an OPERATOR act in a browser, not something the product can drive** — which decides whether R-95's remedy can ever be customer-facing. **Cheapest next step, and it is Viktor's:** read one snapshot's name from the panel; a single `ls /.zfs/snapshot/<name>` from a box then settles whether a named snapshot can be entered even though the directory does not list (ZFS allows exactly that). If it can, per-file recovery becomes product-reachable and this closes cheaply. | **OPEN — one panel read settles it** |
|
||||
|
||||
<!-- DUE-CHECKS-BEGIN — machine-readable. Parsed by scripts/due_checks_gate.py.
|
||||
One row per dated check. The R-number must have a row above. Dates are UTC.
|
||||
|
||||
@@ -1,3 +1,51 @@
|
||||
## v0.111.0 — notice a deletion within a day (2026-09-01, R-431; corrects R-429, re-scopes R-95)
|
||||
|
||||
**Third signal in `OffsiteChecker`, beside FILL and STALENESS. No controller change, no golden owed.**
|
||||
|
||||
**WHY THE HUB AND NOT THE BOX.** The event being detected is a box deleting its own off-site backups,
|
||||
so a detector living on that box is one the same event can silence. The hub already receives
|
||||
`snapshot_count` in every report and already keeps the history. The box still *reports* the number and
|
||||
a sophisticated attacker could lie — a far higher bar than deleting files, and the weekly integrity
|
||||
check would then disagree with the lie.
|
||||
|
||||
**THE THRESHOLD IS REASONED, NOT INVENTED, and the measurement is in the code beside it.** Over the
|
||||
hub's own `reports` table — 12 898 reports, 4 customers, 2026-06-05 → 2026-09-01 — **every one of the
|
||||
nine decreases lands exactly on ZERO** (36→0, 18→0 ×2, 15→0, 12→0, 8→0, 3→0), and every one predates
|
||||
`stats_known`, i.e. they are the R-331 shape: a zero meaning *unmeasured*. In the 380-report window
|
||||
where `stats_known` is true there are **zero** decreases (demo-felhom flat at 10; demo-hp 67→68→69).
|
||||
So observed churn gave nothing to calibrate against, and saying so beats inventing a number (R-401).
|
||||
Retention keeps 7 daily + 4 weekly + 6 monthly per `host,tags` group, so it **cannot halve a total**;
|
||||
a mass deletion goes to ~0. Hence: **a fall of more than HALF the previous count, and at least 5.**
|
||||
|
||||
**Three pre-conditions, each with a scar.** `StatsKnown` (R-331 — a zero is not a zero when nobody
|
||||
measured), the declared `State` (R-204 — a re-initialised or credential-less box reports a real,
|
||||
correct drop), and run success (R-100, whose lesson lives in this very file — presence is not success;
|
||||
`incomplete` (R-203) is excluded too). **An untrustworthy report neither alarms nor moves the
|
||||
baseline**, so a `needs_credential` zero cannot become the baseline and make the recovery look like a
|
||||
rise.
|
||||
|
||||
**Escalation-only with recovery re-arm**, like its two siblings: a continuing deletion pages once, and
|
||||
a clean sweep re-arms so a later deletion is caught again.
|
||||
|
||||
**The alarm.** New type `offsite_snapshots_dropped`, severity **`error`** — inside the hub's exact
|
||||
vocabulary (08 §6.1; a value outside it is coerced to `info` and mailed to nobody, which has shipped
|
||||
twice). **Following yesterday's `offsite_proof_empty` precedent exactly, and saying so:** allowlisted
|
||||
in `allowedEventTypes` and added to `operatorOnlyEvents` **in the same commit**, **no `customerMessages`
|
||||
entry** (the message carries live numbers a static template would discard), and **not** in
|
||||
`perAppCooldownEvents` (a fenced act, 08 §6.2). **The message must not claim data loss** — after
|
||||
today's measurement that is usually false — so it says the snapshots still hold the older copy and the
|
||||
route is per-file.
|
||||
|
||||
**Tests.** `offsite_r431_test.go`: fires once on a mass deletion with a valid severity; silent on all
|
||||
four untrustworthy shapes AND leaves the baseline untouched; silent on ordinary retention (69→60,
|
||||
10→7, 10→5) and loud on 69→34; escalation-only latch with re-arm. **ACCEPTANCE — 9 009 real report
|
||||
points from both live boxes replayed through the detector: ZERO alarms**, fixture committed so it
|
||||
never depends on a live database. **Red-proofs run:** loosening the fraction to 0.99 breaks the firing
|
||||
test; removing the `StatsKnown` guard breaks both the guard test and the real-history replay; removing
|
||||
the latch breaks the escalation test — **that third one only became a real proof after the first
|
||||
version of the escalation test was found hollow** (it re-swept the same report, so the baseline had
|
||||
already moved and the latch was never consulted).
|
||||
|
||||
## v0.110.0 — `offsite_proof_empty`: a backup that is intact and holds nothing is not a damaged store (2026-08-31, R-87)
|
||||
|
||||
**Two lines of register, and the reason they are not one line of reuse.** Controller v0.231.0 adds a
|
||||
|
||||
@@ -2019,6 +2019,8 @@ var allowedEventTypes = map[string]bool{
|
||||
// Operator routing is enforced by `notify.operatorOnlyEvents`, in this same commit; this entry
|
||||
// alone does NOT make it operator-only (the v0.78.0 defect that register records).
|
||||
"offsite_proof_empty": true,
|
||||
// R-431 — the hub raises this itself; allowlisted so a hub-origin event is never 400'd.
|
||||
"offsite_snapshots_dropped": true,
|
||||
"crossdrive_completed": true,
|
||||
"crossdrive_failed": true,
|
||||
// controller v0.134.1 — enlarged offsite push refused by the quota gate (warning; the controller's
|
||||
|
||||
@@ -17,6 +17,9 @@ import (
|
||||
//
|
||||
// - FILL: repo_size_bytes vs the shared-model soft quota (quota_gb>0) at warn 90% / crit 95% — the
|
||||
// operator's early warning before the controller's own 100% run-refusal bites the customer.
|
||||
// - SNAPSHOT-DROP (R-431): the reported snapshot count falls by more than retention can explain —
|
||||
// the "something deleted this customer's off-site history" detector. See snapshotDrop below for
|
||||
// why it lives HERE and not on the box, and for the measurement behind its threshold.
|
||||
// - STALENESS: enabled + escrowed but no run in >48h (or never) — the silently-STUCK detector. A
|
||||
// RECENTLY-failing offsite is NOT stale (backup_failed already alerts it); staleness is the
|
||||
// complement: nothing is even trying. Pending/disabled targets are normal onboarding, never stale.
|
||||
@@ -36,6 +39,10 @@ type OffsiteChecker struct {
|
||||
mu sync.Mutex
|
||||
fillStates map[string]string // customerID → fill band
|
||||
staleStates map[string]string // customerID → "ok" | "stale"
|
||||
// R-431. dropStates is the escalation latch ("ok" | "dropped"); lastCounts is the baseline the
|
||||
// next sweep compares against. Only a TRUSTWORTHY report updates lastCounts — see snapshotDrop.
|
||||
dropStates map[string]string
|
||||
lastCounts map[string]int
|
||||
}
|
||||
|
||||
const defaultOffsiteStaleAfter = 48 * time.Hour
|
||||
@@ -52,6 +59,15 @@ type offsiteReport struct {
|
||||
SnapshotCount int `json:"snapshot_count"`
|
||||
RepoSizeBytes int64 `json:"repo_size_bytes"`
|
||||
QuotaGB int `json:"quota_gb"`
|
||||
// StatsKnown (R-331/R-225) — the ONLY thing separating "this repository holds nothing" from
|
||||
// "nobody has ever measured this repository". Both are `snapshot_count: 0` on the wire and they
|
||||
// are opposite news. ABSENT on a pre-v0.225.0 controller, and absence means the box CANNOT
|
||||
// ANSWER, never that the answer is no. snapshotDrop refuses to judge without it.
|
||||
StatsKnown bool `json:"stats_known,omitempty"`
|
||||
// State (R-204/R-193) — a DECLARED condition, empty on every healthy box. A box that has
|
||||
// re-initialised, lost its credential or been abandoned reports a real, correct, large drop.
|
||||
// The box knows its own situation; believe it rather than alarming on it.
|
||||
State string `json:"state,omitempty"`
|
||||
}
|
||||
|
||||
// NewOffsiteChecker builds the checker. Same seeding philosophy as StorageFillChecker: already-breached
|
||||
@@ -64,6 +80,7 @@ func NewOffsiteChecker(s *store.Store, staleAfter time.Duration, onEvent EventNo
|
||||
oc := &OffsiteChecker{
|
||||
store: s, logger: logger, onEvent: onEvent, staleAfter: staleAfter, now: time.Now,
|
||||
fillStates: make(map[string]string), staleStates: make(map[string]string),
|
||||
dropStates: make(map[string]string), lastCounts: make(map[string]int),
|
||||
}
|
||||
customers, err := s.GetCustomers()
|
||||
if err != nil {
|
||||
@@ -83,6 +100,13 @@ func NewOffsiteChecker(s *store.Store, staleAfter time.Duration, onEvent EventNo
|
||||
if !oc.isStale(c.CustomerID, off) {
|
||||
oc.staleStates[c.CustomerID] = "ok"
|
||||
}
|
||||
// R-431: seed the baseline from the current report so a hub restart does not read the first
|
||||
// sweep as a drop from nothing. Only a trustworthy report seeds — an untrustworthy one leaves
|
||||
// no baseline, and no baseline means no verdict.
|
||||
if oc.countIsTrustworthy(off) {
|
||||
oc.lastCounts[c.CustomerID] = off.SnapshotCount
|
||||
oc.dropStates[c.CustomerID] = "ok"
|
||||
}
|
||||
}
|
||||
logger.Printf("[INFO] Offsite checker initialized: fill warn=90%% crit=95%%, stale after %s, %d ok-seeded", staleAfter, seeded)
|
||||
return oc
|
||||
@@ -175,6 +199,124 @@ func (oc *OffsiteChecker) isStale(customerID string, off *offsiteReport) bool {
|
||||
return oc.now().Sub(t) > oc.staleAfter
|
||||
}
|
||||
|
||||
// ── SNAPSHOT-DROP (R-431) ───────────────────────────────────────────────────────────────────────
|
||||
//
|
||||
// WHY THIS LIVES ON THE HUB AND NOT ON THE BOX. The thing being detected is a box deleting its own
|
||||
// off-site backups — so a detector living on that box is a detector the same event can silence. The
|
||||
// hub already receives the count in every report and already keeps the history, so it can notice
|
||||
// without trusting the box's judgement. The box still REPORTS the number, and a sophisticated
|
||||
// attacker could lie about it; that is a far higher bar than deleting files, and the weekly integrity
|
||||
// check (R-359) would then disagree with the lie.
|
||||
//
|
||||
// WHAT IT IS FOR. R-95: the restic credential can delete its own repository, and the sub-account API
|
||||
// has no append-only axis, so prevention needs a transport change. The Storage Box's daily ZFS
|
||||
// snapshots bound the loss — MEASURED 2026-09-01: a write into `/.zfs/snapshot` is REFUSED
|
||||
// (`dest open …: Failure`) while the same write to the account home succeeds. So the remedy exists;
|
||||
// what was missing was NOTICING, and this is that.
|
||||
//
|
||||
// THE THRESHOLD, AND THE MEASUREMENT IT CAME FROM. Measured over the hub's own `reports` table on
|
||||
// 2026-09-01: 12 898 reports, 4 customers, 2026-06-05 → 2026-09-01.
|
||||
//
|
||||
// * In the whole history there are NINE decreases, and EVERY ONE of them lands exactly on ZERO
|
||||
// (36→0, 18→0 ×2, 15→0, 12→0, 8→0, 3→0). There is not one gradual retention decrease anywhere.
|
||||
// * Every one of those nine predates `stats_known`, i.e. they are the R-331 shape — a zero that
|
||||
// means "could not measure", not "nothing is there". Several carry a declared State
|
||||
// (`needs_credential`, `awaiting_recovery_key`), which says so outright.
|
||||
// * In the window where `stats_known` is TRUE (380 reports across both live boxes) there are ZERO
|
||||
// decreases: demo-felhom sat flat at 10, demo-hp moved 67→68→69. Only rises.
|
||||
//
|
||||
// So observed retention churn gives NOTHING to calibrate against, and saying so is the honest answer
|
||||
// rather than inventing a number (R-401's lesson). The threshold is therefore reasoned from what
|
||||
// retention CAN do, not from what it was seen to do:
|
||||
//
|
||||
// the box runs `forget --keep-daily 7 --keep-weekly 4 --keep-monthly 6 --group-by host,tags`
|
||||
// over ~8 apps. On a boundary day several groups can expire at once, so a legitimate pass can
|
||||
// plausibly remove low double digits. **What it can NEVER do is halve the total**: keeping 7 daily
|
||||
// + 4 weekly + 6 monthly per group is a floor, and a mass deletion goes to ~0.
|
||||
//
|
||||
// Hence: **a fall of MORE THAN HALF the previous count, and at least 5 snapshots.** The 50% cannot be
|
||||
// reached by retention; the floor of 5 stops a tiny-count box alarming on ordinary ageing. It is
|
||||
// deliberately NOT sensitive — a detector that cries wolf is switched off within a fortnight, and
|
||||
// this project has proved that twice in a week.
|
||||
const (
|
||||
snapshotDropFraction = 0.5 // more than half the history gone in one step
|
||||
snapshotDropFloor = 5 // and at least this many, so small counts do not twitch
|
||||
)
|
||||
|
||||
// countIsTrustworthy — the three pre-conditions, each with a scar behind it. A report failing ANY of
|
||||
// them is not evidence of anything: it neither alarms NOR updates the baseline, because comparing
|
||||
// against a number nobody could measure is how a detector invents its own findings.
|
||||
func (oc *OffsiteChecker) countIsTrustworthy(off *offsiteReport) bool {
|
||||
// 1. R-331 — a zero is not a zero when stats are unknown. Absent `stats_known` means the box
|
||||
// cannot answer. Every historical decrease in this hub's data is this shape.
|
||||
if !off.StatsKnown {
|
||||
return false
|
||||
}
|
||||
// 2. R-204/R-193 — the box DECLARES its own situation. A re-initialised, credential-less or
|
||||
// abandoned box reports a real, correct drop; alarming on it would be blaming the box for
|
||||
// telling the truth.
|
||||
if off.State != "" {
|
||||
return false
|
||||
}
|
||||
// 3. R-100 — PRESENCE IS NOT SUCCESS, and this file is where that lesson was learned. A count
|
||||
// from a failed or half-finished run is not a measurement of the store. "incomplete" (R-203)
|
||||
// is deliberately excluded too: a partial run legitimately counts less.
|
||||
if off.LastStatus != "ok" || off.LastSuccess == "" {
|
||||
return false
|
||||
}
|
||||
return true
|
||||
}
|
||||
|
||||
// snapshotDropped reports whether `off` is a drop worth alarming about, against the remembered
|
||||
// baseline. Caller holds oc.mu.
|
||||
func (oc *OffsiteChecker) snapshotDropped(customerID string, off *offsiteReport) (bool, int, int) {
|
||||
prev, had := oc.lastCounts[customerID]
|
||||
if !had || prev <= 0 {
|
||||
return false, prev, off.SnapshotCount
|
||||
}
|
||||
drop := prev - off.SnapshotCount
|
||||
if drop < snapshotDropFloor {
|
||||
return false, prev, off.SnapshotCount
|
||||
}
|
||||
return float64(drop) > float64(prev)*snapshotDropFraction, prev, off.SnapshotCount
|
||||
}
|
||||
|
||||
// emitSnapshotDrop — one event, at a severity inside the hub's exact vocabulary (08 §6.1).
|
||||
//
|
||||
// THE MESSAGE MUST NOT SAY THE DATA IS LOST, because after the 2026-09-01 measurement that is usually
|
||||
// false: the daily Storage Box snapshots are read-only to every account (proven, not cited) and hold
|
||||
// the older copy. It says what happened, what it means, and where the data still is.
|
||||
func (oc *OffsiteChecker) emitSnapshotDrop(customerID string, off *offsiteReport, prev, cur int) {
|
||||
message := fmt.Sprintf(
|
||||
"Customer %s: off-site backup count fell from %d to %d snapshot(s) in one report — more than "+
|
||||
"retention can explain. The daily Storage Box snapshots are read-only and still hold the "+
|
||||
"older copy, so this is recoverable file-by-file; it is NOT confirmed data loss. Check "+
|
||||
"whether a deletion ran on the box before restoring anything.",
|
||||
customerID, prev, cur)
|
||||
details, _ := json.Marshal(map[string]any{
|
||||
"customer_id": customerID, "previous_count": prev, "current_count": cur,
|
||||
"drop": prev - cur, "last_success": off.LastSuccess, "last_status": off.LastStatus,
|
||||
})
|
||||
oc.logger.Printf("[INFO] Offsite snapshot drop: %s %d -> %d (offsite_snapshots_dropped)", customerID, prev, cur)
|
||||
if _, err := oc.store.SaveEvent(customerID, "offsite_snapshots_dropped", "error", message, string(details), "hub"); err != nil {
|
||||
oc.logger.Printf("[WARN] Failed to save offsite snapshot-drop event for %s: %v", customerID, err)
|
||||
return
|
||||
}
|
||||
if oc.onEvent != nil {
|
||||
oc.onEvent(customerID, "offsite_snapshots_dropped", "error", message, string(details), "hub")
|
||||
}
|
||||
}
|
||||
|
||||
// GetDropState exposes the latch for tests.
|
||||
func (oc *OffsiteChecker) GetDropState(customerID string) string {
|
||||
oc.mu.Lock()
|
||||
defer oc.mu.Unlock()
|
||||
if s := oc.dropStates[customerID]; s != "" {
|
||||
return s
|
||||
}
|
||||
return "unknown"
|
||||
}
|
||||
|
||||
// warnLegacyOnce logs the legacy degrade a single time per customer. Once, because this is a
|
||||
// steady-state condition until the box upgrades — a per-cycle line would be pure noise — but it must be
|
||||
// logged at all, so a fleet silently running on the old anchor is visible rather than assumed.
|
||||
@@ -228,11 +370,15 @@ func (oc *OffsiteChecker) Check() {
|
||||
if off == nil {
|
||||
delete(oc.fillStates, c.CustomerID) // vanished object (disabled / downgraded) → re-arm
|
||||
delete(oc.staleStates, c.CustomerID)
|
||||
delete(oc.dropStates, c.CustomerID)
|
||||
delete(oc.lastCounts, c.CustomerID) // no baseline survives a vanished object
|
||||
continue
|
||||
}
|
||||
if oc.store.IsCustomerBlocked(c.CustomerID) {
|
||||
delete(oc.fillStates, c.CustomerID)
|
||||
delete(oc.staleStates, c.CustomerID)
|
||||
delete(oc.dropStates, c.CustomerID)
|
||||
delete(oc.lastCounts, c.CustomerID)
|
||||
continue
|
||||
}
|
||||
|
||||
@@ -261,6 +407,24 @@ func (oc *OffsiteChecker) Check() {
|
||||
}
|
||||
}
|
||||
oc.staleStates[c.CustomerID] = newStale
|
||||
|
||||
// SNAPSHOT-DROP (R-431). Escalation-only, exactly like the two signals above: one deletion
|
||||
// produces ONE alarm, not one per report cycle. The baseline then moves to the new value, so a
|
||||
// SECOND deletion later is still caught — from the new floor.
|
||||
//
|
||||
// An UNTRUSTWORTHY report is a no-op in both directions: no alarm, and the baseline is left
|
||||
// alone rather than being overwritten with a number nobody could measure. That is what stops
|
||||
// a `needs_credential` zero from becoming the baseline and making the RECOVERY look like a rise.
|
||||
if oc.countIsTrustworthy(off) {
|
||||
dropped, prev, cur := oc.snapshotDropped(c.CustomerID, off)
|
||||
if dropped && oc.dropStates[c.CustomerID] != "dropped" {
|
||||
oc.emitSnapshotDrop(c.CustomerID, off, prev, cur)
|
||||
oc.dropStates[c.CustomerID] = "dropped"
|
||||
} else if !dropped {
|
||||
oc.dropStates[c.CustomerID] = "ok"
|
||||
}
|
||||
oc.lastCounts[c.CustomerID] = cur
|
||||
}
|
||||
}
|
||||
for k := range oc.fillStates {
|
||||
if !seen[k] {
|
||||
@@ -272,6 +436,12 @@ func (oc *OffsiteChecker) Check() {
|
||||
delete(oc.staleStates, k)
|
||||
}
|
||||
}
|
||||
for k := range oc.dropStates {
|
||||
if !seen[k] {
|
||||
delete(oc.dropStates, k)
|
||||
delete(oc.lastCounts, k)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// GetFillState / GetStaleState expose current states for tests.
|
||||
|
||||
@@ -0,0 +1,278 @@
|
||||
package monitor
|
||||
|
||||
import (
|
||||
"encoding/json"
|
||||
"fmt"
|
||||
"os"
|
||||
"path/filepath"
|
||||
"strings"
|
||||
"testing"
|
||||
"time"
|
||||
)
|
||||
|
||||
// R-431 — the snapshot-drop signal, proven in BOTH directions.
|
||||
//
|
||||
// The acceptance test is TestR431_RealHistoryProducesZeroAlarms: a detector that fires on healthy
|
||||
// boxes is the mistake this project caught twice in one week, and it is the one that would get this
|
||||
// signal switched off within a fortnight.
|
||||
|
||||
// dropJSON builds a trustworthy offsite object (stats_known, no declared state, last run ok).
|
||||
func dropJSON(count int, statsKnown bool, state, lastStatus string) string {
|
||||
ts := time.Now().UTC().Add(-1 * time.Hour).Format(time.RFC3339)
|
||||
success := ts
|
||||
if lastStatus != "ok" {
|
||||
success = ""
|
||||
}
|
||||
return fmt.Sprintf(
|
||||
`{"enabled":true,"escrow_state":"escrowed","last_run":%q,"last_status":%q,"last_success":%q,`+
|
||||
`"snapshot_count":%d,"repo_size_bytes":1073741824,"quota_gb":0,"stats_known":%v,"state":%q}`,
|
||||
ts, lastStatus, success, count, statsKnown, state)
|
||||
}
|
||||
|
||||
// TestR431_FiresOnAMassDeletion — direction 1. A drop past the threshold alarms EXACTLY once.
|
||||
//
|
||||
// RED-PROOF (run 2026-09-01, recorded in REPORT.md): setting snapshotDropFraction to 0.99 makes this
|
||||
// fail — 69 → 4 is a 94% fall and would no longer qualify, which is what an over-loose threshold
|
||||
// looks like in production.
|
||||
func TestR431_FiresOnAMassDeletion(t *testing.T) {
|
||||
st := newDiskStore(t)
|
||||
var got []struct{ et, sev, msg string }
|
||||
saveOffsiteReport(t, st, "victim", dropJSON(69, true, "", "ok"))
|
||||
oc := NewOffsiteChecker(st, 48*time.Hour, func(_, et, sev, msg, _, _ string) {
|
||||
got = append(got, struct{ et, sev, msg string }{et, sev, msg})
|
||||
}, quietLog())
|
||||
|
||||
// the constructor seeded the baseline at 69; now the store is emptied
|
||||
saveOffsiteReport(t, st, "victim", dropJSON(4, true, "", "ok"))
|
||||
oc.Check()
|
||||
|
||||
var drops []struct{ et, sev, msg string }
|
||||
for _, g := range got {
|
||||
if g.et == "offsite_snapshots_dropped" {
|
||||
drops = append(drops, g)
|
||||
}
|
||||
}
|
||||
if len(drops) != 1 {
|
||||
t.Fatalf("want exactly 1 offsite_snapshots_dropped, got %d (%v)", len(drops), got)
|
||||
}
|
||||
// The severity MUST be in the hub's exact vocabulary — anything else is coerced to info and
|
||||
// mailed to nobody (08 §6.1, shipped twice).
|
||||
switch drops[0].sev {
|
||||
case "info", "warning", "error", "critical":
|
||||
default:
|
||||
t.Fatalf("severity %q is outside the hub vocabulary — it would be coerced to info and reach nobody", drops[0].sev)
|
||||
}
|
||||
for _, frag := range []string{"69", "4", "read-only", "NOT confirmed data loss"} {
|
||||
if !strings.Contains(drops[0].msg, frag) {
|
||||
t.Fatalf("message must contain %q; got: %s", frag, drops[0].msg)
|
||||
}
|
||||
}
|
||||
if strings.Contains(drops[0].msg, "data is lost") || strings.Contains(drops[0].msg, "data lost") {
|
||||
t.Fatalf("the message must NOT claim data loss — the snapshots usually still hold it: %s", drops[0].msg)
|
||||
}
|
||||
|
||||
}
|
||||
|
||||
// TestR431_EscalationOnlyLatch — a CONTINUING deletion must not page on every report cycle.
|
||||
//
|
||||
// WRITTEN AFTER A HOLLOW FIRST ATTEMPT, and the failure is recorded because it is instructive: the
|
||||
// original assertion re-swept the SAME report and called that "escalation-only". It proved nothing —
|
||||
// the baseline had already moved to the new count, so the second sweep saw a drop of zero and the
|
||||
// latch was never consulted. Its red-proof (removing the latch) PASSED, which is how it was caught.
|
||||
//
|
||||
// This drives a count that keeps FALLING, which is the only shape where the latch is load-bearing.
|
||||
//
|
||||
// RED-PROOF (run 2026-09-01): replacing `dropped && oc.dropStates[...] != "dropped"` with `dropped`
|
||||
// makes this fail with 2 alarms — one per report cycle, which is the noise that trains an operator
|
||||
// to ignore the alarm.
|
||||
func TestR431_EscalationOnlyLatch(t *testing.T) {
|
||||
st := newDiskStore(t)
|
||||
var n int
|
||||
saveOffsiteReport(t, st, "sliding", dropJSON(69, true, "", "ok"))
|
||||
oc := NewOffsiteChecker(st, 48*time.Hour, func(_, et, _, _, _, _ string) {
|
||||
if et == "offsite_snapshots_dropped" {
|
||||
n++
|
||||
}
|
||||
}, quietLog())
|
||||
|
||||
saveOffsiteReport(t, st, "sliding", dropJSON(30, true, "", "ok")) // 69 -> 30: alarm
|
||||
oc.Check()
|
||||
if n != 1 {
|
||||
t.Fatalf("the first large drop must alarm exactly once; got %d", n)
|
||||
}
|
||||
saveOffsiteReport(t, st, "sliding", dropJSON(2, true, "", "ok")) // 30 -> 2: still falling
|
||||
oc.Check()
|
||||
if n != 1 {
|
||||
t.Fatalf("a CONTINUING deletion must not re-page while the latch is set; got %d alarms", n)
|
||||
}
|
||||
if oc.GetDropState("sliding") != "dropped" {
|
||||
t.Fatalf("the latch must be held, got %q", oc.GetDropState("sliding"))
|
||||
}
|
||||
|
||||
// RECOVERY RE-ARMS: a clean sweep clears the latch, so a LATER deletion is caught again.
|
||||
saveOffsiteReport(t, st, "sliding", dropJSON(40, true, "", "ok")) // rebuilt, no drop
|
||||
oc.Check()
|
||||
if oc.GetDropState("sliding") != "ok" {
|
||||
t.Fatalf("a clean sweep must re-arm the latch, got %q", oc.GetDropState("sliding"))
|
||||
}
|
||||
saveOffsiteReport(t, st, "sliding", dropJSON(1, true, "", "ok")) // deleted again
|
||||
oc.Check()
|
||||
if n != 2 {
|
||||
t.Fatalf("after re-arming, a NEW deletion must alarm again; got %d", n)
|
||||
}
|
||||
}
|
||||
|
||||
// TestR431_SilentWhenNotTrustworthy — direction 2, the three pre-conditions, each with its scar.
|
||||
func TestR431_SilentWhenNotTrustworthy(t *testing.T) {
|
||||
cases := []struct {
|
||||
name, first, second string
|
||||
}{
|
||||
{"stats_known absent (R-331: a zero that means UNMEASURED)",
|
||||
dropJSON(69, true, "", "ok"), dropJSON(0, false, "", "ok")},
|
||||
{"a DECLARED state (R-204: the box says what happened)",
|
||||
dropJSON(69, true, "", "ok"), dropJSON(0, true, "needs_credential", "ok")},
|
||||
{"the run FAILED (R-100: presence is not success)",
|
||||
dropJSON(69, true, "", "ok"), dropJSON(0, true, "", "error")},
|
||||
{"the run was INCOMPLETE (R-203: a partial run counts less)",
|
||||
dropJSON(69, true, "", "ok"), dropJSON(0, true, "", "incomplete")},
|
||||
}
|
||||
for i, c := range cases {
|
||||
t.Run(c.name, func(t *testing.T) {
|
||||
st := newDiskStore(t)
|
||||
cid := fmt.Sprintf("c%d", i)
|
||||
var n int
|
||||
saveOffsiteReport(t, st, cid, c.first)
|
||||
oc := NewOffsiteChecker(st, 48*time.Hour, func(_, et, _, _, _, _ string) {
|
||||
if et == "offsite_snapshots_dropped" {
|
||||
n++
|
||||
}
|
||||
}, quietLog())
|
||||
saveOffsiteReport(t, st, cid, c.second)
|
||||
oc.Check()
|
||||
if n != 0 {
|
||||
t.Fatalf("%s: must NOT alarm; got %d", c.name, n)
|
||||
}
|
||||
// AND the baseline must be untouched, so the RECOVERY does not read as a rise-then-drop.
|
||||
oc.mu.Lock()
|
||||
base := oc.lastCounts[cid]
|
||||
oc.mu.Unlock()
|
||||
if base != 69 {
|
||||
t.Fatalf("%s: an untrustworthy report must not overwrite the baseline; got %d", c.name, base)
|
||||
}
|
||||
})
|
||||
}
|
||||
}
|
||||
|
||||
// TestR431_OrdinaryRetentionIsSilent — a fall that retention CAN explain must not alarm.
|
||||
func TestR431_OrdinaryRetentionIsSilent(t *testing.T) {
|
||||
for _, c := range []struct {
|
||||
name string
|
||||
from, to int
|
||||
wantAlarms int
|
||||
}{
|
||||
{"69 -> 60 (9 gone, under half)", 69, 60, 0},
|
||||
{"10 -> 7 (3 gone, under the floor of 5)", 10, 7, 0},
|
||||
{"10 -> 5 (5 gone, exactly half — NOT more than half)", 10, 5, 0},
|
||||
{"69 -> 34 (35 gone, more than half)", 69, 34, 1},
|
||||
} {
|
||||
t.Run(c.name, func(t *testing.T) {
|
||||
st := newDiskStore(t)
|
||||
cid := fmt.Sprintf("r%d%d", c.from, c.to)
|
||||
var n int
|
||||
saveOffsiteReport(t, st, cid, dropJSON(c.from, true, "", "ok"))
|
||||
oc := NewOffsiteChecker(st, 48*time.Hour, func(_, et, _, _, _, _ string) {
|
||||
if et == "offsite_snapshots_dropped" {
|
||||
n++
|
||||
}
|
||||
}, quietLog())
|
||||
saveOffsiteReport(t, st, cid, dropJSON(c.to, true, "", "ok"))
|
||||
oc.Check()
|
||||
if n != c.wantAlarms {
|
||||
t.Fatalf("%s: want %d alarm(s), got %d", c.name, c.wantAlarms, n)
|
||||
}
|
||||
})
|
||||
}
|
||||
}
|
||||
|
||||
// TestR431_RealHistoryProducesZeroAlarms — THE ACCEPTANCE STEP.
|
||||
//
|
||||
// Replays the ACTUAL snapshot-count history of both live boxes, exported from the hub's own reports
|
||||
// table on 2026-09-01, through the real detector. It must produce ZERO alarms. A warning that fires
|
||||
// on healthy boxes is worse than no warning at all.
|
||||
//
|
||||
// The fixture is committed beside this test so the assertion does not depend on a live database.
|
||||
// If it is absent the test FAILS rather than skipping — a silent skip is how a green tick comes to
|
||||
// mean nothing.
|
||||
func TestR431_RealHistoryProducesZeroAlarms(t *testing.T) {
|
||||
path := filepath.Join("testdata", "r431_real_history.json")
|
||||
raw, err := os.ReadFile(path)
|
||||
if err != nil {
|
||||
t.Fatalf("the real-history fixture is missing (%v) — this is a FAILURE, never a skip: "+
|
||||
"without it the acceptance step asserts nothing", err)
|
||||
}
|
||||
var hist map[string][]struct {
|
||||
At string `json:"at"`
|
||||
Count int `json:"count"`
|
||||
StatsKnown bool `json:"stats_known"`
|
||||
State string `json:"state"`
|
||||
LastStatus string `json:"last_status"`
|
||||
}
|
||||
if err := json.Unmarshal(raw, &hist); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if len(hist) == 0 {
|
||||
t.Fatal("the fixture is empty — it would pass vacuously")
|
||||
}
|
||||
|
||||
total := 0
|
||||
for cust, points := range hist {
|
||||
if len(points) < 2 {
|
||||
t.Fatalf("%s: fewer than 2 points — nothing to compare", cust)
|
||||
}
|
||||
total += len(points)
|
||||
st := newDiskStore(t)
|
||||
var alarms int
|
||||
var first = points[0]
|
||||
saveOffsiteReport(t, st, cust, dropJSON(first.Count, first.StatsKnown, first.State, first.LastStatus))
|
||||
oc := NewOffsiteChecker(st, 48*time.Hour, func(_, et, _, msg, _, _ string) {
|
||||
if et == "offsite_snapshots_dropped" {
|
||||
alarms++
|
||||
t.Errorf("%s: FIRED ON REAL HISTORY: %s", cust, msg)
|
||||
}
|
||||
}, quietLog())
|
||||
// Feed the remaining points straight through the detector. Driving Check() per point would
|
||||
// need a 1.1s sleep each time (received_at is second-resolution) — hours for 5 000 points —
|
||||
// so the sweep's own decision path is exercised directly instead, with the same guards.
|
||||
for _, p := range points[1:] {
|
||||
off := &offsiteReport{
|
||||
SnapshotCount: p.Count, StatsKnown: p.StatsKnown, State: p.State,
|
||||
LastStatus: p.LastStatus, LastSuccess: "2026-09-01T00:00:00Z",
|
||||
}
|
||||
if p.LastStatus != "ok" {
|
||||
off.LastSuccess = ""
|
||||
}
|
||||
oc.mu.Lock()
|
||||
if oc.countIsTrustworthy(off) {
|
||||
dropped, prev, cur := oc.snapshotDropped(cust, off)
|
||||
if dropped && oc.dropStates[cust] != "dropped" {
|
||||
alarms++
|
||||
t.Errorf("%s at %s: FIRED ON REAL HISTORY %d -> %d", cust, p.At, prev, cur)
|
||||
oc.dropStates[cust] = "dropped"
|
||||
} else if !dropped {
|
||||
oc.dropStates[cust] = "ok"
|
||||
}
|
||||
oc.lastCounts[cust] = cur
|
||||
}
|
||||
oc.mu.Unlock()
|
||||
}
|
||||
if alarms != 0 {
|
||||
t.Fatalf("%s: %d alarm(s) on real history — the threshold is wrong", cust, alarms)
|
||||
}
|
||||
}
|
||||
// POSITIVE CONTROL: the replay must actually have looked at something. Without this the test
|
||||
// passes when the fixture is a list of empty lists.
|
||||
if total < 100 {
|
||||
t.Fatalf("only %d points replayed — too few for this to mean anything", total)
|
||||
}
|
||||
t.Logf("replayed %d real report points across %d customers: ZERO alarms", total, len(hist))
|
||||
}
|
||||
File diff suppressed because one or more lines are too long
@@ -527,6 +527,10 @@ var operatorOnlyEvents = map[string]bool{
|
||||
// a block — FormatCustomerEmail falls back to the raw message, which is the v0.78.0 defect this
|
||||
// register was built for.
|
||||
"offsite_proof_empty": true,
|
||||
// R-431 — OPERATOR ONLY. A count that fell is a question for the operator, not a customer:
|
||||
// the snapshots still hold the data and telling a customer "your backups were deleted"
|
||||
// would be wrong on the usual reading.
|
||||
"offsite_snapshots_dropped": true,
|
||||
// R-197 (v0.93.0). "The sealed offsite repository key changed" is a custody fact about escrow
|
||||
// blobs. A customer can take no action on it — the remedy is the operator's inspection of the
|
||||
// off-site tier — and the text is operator-grade English naming host ids and retained-blob
|
||||
|
||||
Reference in New Issue
Block a user