REPORT + R-437: the beta stopping line recorded, and the live trigger declined with its reason
gates / gates (push) Successful in 19s

REPORT.md carries the deployed sentence quoted from the RUNNING binary (kubectl cp + byte
grep, both controls), the stopping line as it reads in all three places, the enumerated
deferred set, the two provider questions, the register census, and the ArgoCD verification.

R-437 filed: the register compression sweep is OWED and was deliberately not run here.
Measured first — 12 of 181 rows / ~25 KB of 316 KB (about 7%) carry a closed leading
verdict — so it buys little and touches everything, and it is the exact operation that
misfiled seven rows in August (R-378; the seventh, R-87, sat wrong for nine days, R-405).
The row carries the scope so it can be picked up cold.

The live alarm trigger was NOT run and the report says so in its own section rather than
substituting quietly: this alarm only fires on a real fall in a real customer's snapshot
count, so firing it means either deleting real backups or POSTing a falsified report
claiming demo-hp lost its own. That would write a fabricated point into a customer's report
history, move its latch and baseline, and mail the operator a second alarm about a real box
hours after the first one already confused him. Covered instead by the deployed-bytes proof
plus three red-proofed tests driving saveReport -> Check -> notify. What remains unproven is
named: that the dispatcher delivers THIS wording to a mailbox.

Register 688 -> 700 lines; 182 rows; open-state 170.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LB8FmJaGd2cyjvy6dbEjpM
This commit is contained in:
2026-09-01 18:41:50 +02:00
parent 0f65f7a197
commit 1a1b32b3dd
2 changed files with 276 additions and 147 deletions
+1
View File
@@ -674,6 +674,7 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server`
| **R-434** | **The snapshot-drop alarm promises a recovery that cannot be performed.** `hub/internal/monitor/offsite.go` `emitSnapshotDrop` ships this text, live in hub **v0.111.0**: *"The daily Storage Box snapshots are read-only and still hold the older copy, **so this is recoverable file-by-file**; it is NOT confirmed data loss."* The first clause is true. The second is not reachable: not by the product (R-433), and not by the operator without a browser and a main-account credential that does not exist here. Its own comment states the intent — *"THE MESSAGE MUST NOT SAY THE DATA IS LOST, because after the 2026-09-01 measurement that is usually false"* — and the measurement it rests on was superseded the same day. **This is this project's own corollary landing on the alarm shipped that morning:** when a verdict changes which fact it counts from, the alarm text has to change with it, or the operator acts on a promise nobody can keep. **Fix is text-only and must not be made before R-433 settles what IS true** — an alarm rewritten twice in a week is worse than one rewritten once. **✅ CLOSED 2026-09-01, hub v0.111.1 — AND THE BLOCK ABOVE WAS WRONG, WHICH IS THE POINT WORTH KEEPING.** This row said the fix had to wait for R-433 to establish what IS true. **It did not, because the fix is a DELETION and not a REPLACEMENT.** The promise was withdrawn rather than swapped for a new one: *"The daily Storage Box snapshots are read-only and still hold the older copy. The route back out of them is not yet established, so treat this as neither confirmed data loss nor confirmed recovery. Get in touch before restoring anything, and check whether a deletion ran on the box."* **That sentence is true under EVERY possible answer to the provider questions, so it never needs a second rewrite** — which is the whole reason it was not blocked. A replacement would have been. **Three tests in `hub/internal/monitor/offsite_r434_test.go`**, all driving the production path so they assert the sentence an operator RECEIVES: the withdrawal is present, the promise is absent in three shapes, `confirmed data loss` may appear only inside its negation, and the stored row must not drift from the delivered mail. **RED-PROOF: restoring the v0.111.0 sentence failed all three**, on every fragment, with the offending sentence printed. **One existing test was edited and it had caught this fix correctly** — `TestR431_FiresOnAMassDeletion` asserted `"NOT confirmed data loss"`; the fragment was REMOVED rather than updated so the wording keeps ONE home. | **CLOSED 2026-09-01 — hub v0.111.1; the "blocked on R-433" verdict above was mine and it was wrong** |
| **R-435** | **The snapshot-drop detector is blind to a single-app deletion — the exact shape the box can produce.** `snapshotDropFraction = 0.5` and `snapshotDropFloor = 5` (`hub/internal/monitor/offsite.go`) require a fall of MORE than half the previous count. demo-hp's baseline is **69** across 9 apps, so ~35 snapshots must go before it speaks; **one app's tag is ~9 and is invisible.** `offbox.go:1388` runs `forget --prune` **grouped by host,tags** — a per-tag wipe is precisely the shape a faulty retention or a targeted deletion produces. **This is deliberate, not accidental:** the constant's own comment argues the insensitivity, and *"a detector that cries wolf is switched off within a fortnight"* is a lesson this project paid for. **So this row is NOT a demand to lower the threshold.** It is a demand that the blind spot be written where the operator reads it, because "an unexplained fall is noticed within a day" (R-431, and `STATUS.md`) is true only of falls above half. **Discovered by arithmetic while planning the drill's Phase 2/3 pairing, which could not have worked:** Phase 2 deletes one app and Phase 3 expects the alarm to fire. **LIMITATION NOW WRITTEN INTO THE ALARM'S OWN DOCUMENTATION, hub v0.111.1** — the comment above `snapshotDropFraction` in `hub/internal/monitor/offsite.go` now states what the detector does NOT see, with the demo-hp arithmetic and the `forget --prune` grouping that makes the blind spot sit on the most likely single-app failure. **It also says explicitly that the numbers must NOT be lowered to "fix" this** and that per-app detection needs a SECOND signal keyed on the per-tag count. **The row stays OPEN because documenting a blind spot is not covering it** — and because `STATUS.md` and R-431 both still say "noticed within a day", which is true only of falls above half. | **OPEN — documented in code v0.111.1; the coverage gap itself is unclosed** |
| **R-436** | **LEAD, NOT A DEFECT — append-only may be reachable without a new machine, which would make R-95's real prevention far cheaper than the spike concluded.** Hetzner's port-23 restricted shell advertises, in its own `help`, these server-side backends: `borg`, `rsync`, `scp`, `sftp`, **`rclone serve restic --stdio`**. And restic 0.14.0 **recognises the `rclone:` backend** — MEASURED 2026-09-01 with a control: `banana:` → `Fatal: parsing repository location failed: invalid backend`, while `rclone:` → `exec: "rclone": executable file not found in $PATH` (i.e. the backend parsed and it tried to run the helper). rclone is **not** in the controller image today. `rclone serve restic` carries an **`--append-only`** flag. **Why this matters:** the R-95 spike's option 3 was priced at a new always-on service in the recovery path plus either a mount in the hot path or migrating every customer's history — and it was deferred on that price. This route needs neither: the server side already runs at the provider. **THE CAVEAT, STATED FIRST because it may kill the idea:** the **client** supplies the server command line, so a compromised guest could simply omit `--append-only` unless the provider pins it. **NOT ESTABLISHED:** whether Hetzner pins the flag or accepts client-supplied arguments. **That is a vendor question and it is cheap** — it should be asked before any code is written, because if the answer is "client-supplied" this lead is worth nothing. | **OPEN — ask the vendor before building anything** |
| **R-437** | **The register compression sweep is OWED, and it was deliberately NOT run inside the 2026-09-01 beta-line session — this row is the record of that choice, not a note.** **The ask:** compress what has closed in `OPEN-ITEMS.md`. **The measurement, taken before deciding:** 181 rows, 316 KB of row text, of which **12 rows / ~25 KB (about 7 %) carry a CLOSED/DECIDED/ANSWERED leading verdict.** So the sweep buys little and touches everything. **Why it was refused as a side-task, and the citation matters:** a compression sweep is the exact operation that has already gone wrong here. The 2026-08-22 sweep (`ef6ac6f`, R-376..R-378) matched a status word ANYWHERE in the row, moved rows that were not closed, and **R-378 caught six in the same session and missed a seventh** — R-87 sat in the wrong register for nine days while the ranking paragraph pointed at nothing (R-405). **That is a session-scale hazard, and running it as the tail end of a session about something else is how it happened the first time.** **WHAT IS OWED, scoped so it can be picked up cold:** (1) classify by the **LEADING VERDICT** of the state cell only — the rule `closed_register_gate.py` already implements and red-proofs, never a whole-row match; (2) move, never rewrite — a compressed row that loses its evidence is worse than a long one; (3) run `closed_register_gate.py` before and after and quote both; (4) re-read the ranking paragraph afterwards, because that is the surface that silently went stale last time. **Not urgent:** the file is 688 lines and every gate reads it in well under a second. | **OPEN — owed; needs its own session, not a tail end** |
<!-- DUE-CHECKS-BEGIN — machine-readable. Parsed by scripts/due_checks_gate.py.
One row per dated check. The R-number must have a row above. Dates are UTC.