REPORT + R-437: the beta stopping line recorded, and the live trigger declined with its reason
gates / gates (push) Successful in 19s

REPORT.md carries the deployed sentence quoted from the RUNNING binary (kubectl cp + byte
grep, both controls), the stopping line as it reads in all three places, the enumerated
deferred set, the two provider questions, the register census, and the ArgoCD verification.

R-437 filed: the register compression sweep is OWED and was deliberately not run here.
Measured first — 12 of 181 rows / ~25 KB of 316 KB (about 7%) carry a closed leading
verdict — so it buys little and touches everything, and it is the exact operation that
misfiled seven rows in August (R-378; the seventh, R-87, sat wrong for nine days, R-405).
The row carries the scope so it can be picked up cold.

The live alarm trigger was NOT run and the report says so in its own section rather than
substituting quietly: this alarm only fires on a real fall in a real customer's snapshot
count, so firing it means either deleting real backups or POSTing a falsified report
claiming demo-hp lost its own. That would write a fabricated point into a customer's report
history, move its latch and baseline, and mail the operator a second alarm about a real box
hours after the first one already confused him. Covered instead by the deployed-bytes proof
plus three red-proofed tests driving saveReport -> Check -> notify. What remains unproven is
named: that the dispatcher delivers THIS wording to a mailbox.

Register 688 -> 700 lines; 182 rows; open-state 170.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LB8FmJaGd2cyjvy6dbEjpM
This commit is contained in:
2026-09-01 18:41:50 +02:00
parent 0f65f7a197
commit 1a1b32b3dd
2 changed files with 276 additions and 147 deletions
+275 -147
View File
@@ -1,166 +1,294 @@
# REPORT — the deletion we said is survivable: the recovery route does not exist (R-95 drill, 2026-09-01)
# REPORT — the line under the backup arc, and one alarm that was telling the operator something untrue
**RUNBOOK, destructive class, `demo-hp` only. STOPPED at the end of Phase 1 on the operator's ruling,
before any destructive step. No delete verb was issued against any live store; no byte on either
Storage Box sub-account was written, moved or removed.** No production code, no version bump, no
image, no golden. Evidence: `documentation/audits/evidence-drill-r95-recovery-2026-09-01/`.
**hub v0.111.0 → v0.111.1 · controller v0.232.0 UNTOUCHED · 2026-09-01**
| # | phase | verdict | one sentence |
|---|---|---|---|
| 1 | snapshot reachable, and its name | **NO — and it has no reachable name** | 777,600 exact names across nine days in the vendor's own format, plus 126 alternative shapes; zero resolve, with a control proving the sweep detects a path that exists. |
| 2 | the deletion | **NOT RUN — operator ruling** | With no recovery route, the deletion would have destroyed real history to buy nothing; put as a two-option decision, the ruling was stop. |
| 3 | the alarm fired | **NOT RUN — and it could not have fired at the specified size** | The shipped threshold needs a fall of more than half; one app's tag is ~9 of 69. → **R-435** |
| 4 | **the recovery** | **NOT RUN — no route exists that is not fenced** | Box-side: proven impossible. Panel: fenced and browserless. Hetzner API: fenced (§11-D), and the hub's client has no snapshot method at all. |
| — | **RTO from T₀** | **STILL BLANK** | Row 10's RTO cell is unchanged and remains a finding. |
| — | **data lost, quantified** | **NOT MEASURABLE THIS WAY** | The quantity only has meaning if the rest is recoverable, and the route that would recover it is not reachable. |
| 5 | re-arm | **NOT RUN** | Depended on Phase 4. |
| 6 | teardown | **PASS** | Store untouched at 69 snapshots; all three scratch layers removed; hub DB copies shredded; both boxes healthy. |
**No controller release. No agent release. No golden owed, no floor change.** The only behaviour
change in this session is one sentence in one alarm. Everything else is register and documentation.
---
## 1. Did the recovery work — and does yesterday's re-scope survive?
## 1. The alarm's new text, and proof the old promise is gone
**The recovery was never reachable, and the re-scope does not survive intact. Its first half stands;
its second half does not.**
Yesterday's re-scope has two clauses. They must now be separated:
* **(a) "The box can delete its live repository, but cannot write to the daily snapshots of it."**
**STANDS.** Re-confirmed here: `/.zfs/snapshot` is reachable and the write-refusal measurement is
unchanged. Nothing in this drill weakens it.
* **(b) "…so the rest is recoverable — file by file, one customer at a time."** **NOT SUPPORTED.**
A snapshot that cannot be opened cannot be copied out of. R-432 recorded the directory listing
empty and named the cheapest next step: *"a single `ls /.zfs/snapshot/<name>` from a box then
settles whether a named snapshot can be entered even though the directory does not list (ZFS
allows exactly that)."* **That step is now done, exhaustively, and the answer is no.**
**What was measured.** The port-23 restricted shell accepts a batched `stat`, which makes a cheap
existence oracle: 500–600 paths per round trip, stdout carrying only paths that exist.
| sweep | candidates | hits |
|---|---|---|
| `/.zfs/snapshot/YYYY-MM-DDTHH-MM-SS`, nine full days, second granularity | **777,600** | **0** |
| 126 alternative name shapes and snapshot paths (`daily`, `snapshot-1`, colon and compact time forms, `/home/.snapshot`, …) | 126 | 0 |
| **control — the identical 600-name batch shape with one real path appended** | 6 batches | **6/6 returned it** |
**And there is a structural reason, which is why I stopped sweeping.** The customer's data and the
snapshot door are on **different filesystems**:
**Quoted from the RUNNING binary** — `kubectl cp` out of pod `hub-857678f9b4-v95s6`, byte-grepped:
```
df → u629488-sub3 mounted on /home
stat /home → Device 0,82
stat /.zfs/snapshot → Device 0,276 ← a different device
stat /home/.zfs → cannot statx: No such file or directory
Customer %s: off-site backup count fell from %d to %d snapshot(s) in one report — more than
retention can explain. The daily Storage Box snapshots are read-only and still hold the older
copy. The route back out of them is not yet established, so treat this as neither confirmed
data loss nor confirmed recovery. Get in touch before restoring anything, and check whether a
deletion ran on the box.
```
A ZFS snapshot under `/.zfs/snapshot` belongs to the dataset that owns that `.zfs` — not to the
child mounted at `/home`. **So even a correctly named snapshot there could not contain
`felhom-repo`,** and the dataset that does hold it exposes no `.zfs` at all to this account. The
empty listing is not a display toggle hiding a reachable tree; from a sub-account there is no tree.
**Every shape of the withdrawn promise is absent from the deployed bytes:**
**Three tools agree, each with controls in the same run:** SFTP, the port-23 shell, and
`rsync --list-only`.
**What that does to R-95.** Its *exposure* is unchanged and its *remedy* is not. Yesterday the row
could say a deletion costs about a day because the rest comes back per-file. Today the only routes
to "the rest" are a whole-box panel rollback (which deletes newer snapshots and hits every customer
on the box) and the provider API (fenced, and unimplemented in the hub's client). **The re-scope's
comfort was resting on a route nobody had walked — which is precisely the standard this project
applies, and it is the reason this drill was called.**
**The ranking is Viktor's and I am not re-ranking it.** What I will say plainly: the argument that
moved R-95 down yesterday is the argument this drill removed. On these facts I would put it back
where it was.
## 2. The RTO
**Still blank, and it stays a finding.** `07` §8 row 10's RTO cell has been empty since July and this
drill did not fill it. Nothing was recovered, so nothing was timed. The summary line at
`07-backup-architecture.md:948` — *"no ransomware-shaped recovery has ever been run"* — is still
true, and is now true for a sharper reason: **not "nobody has run it" but "from the box, it cannot
be run."**
## 3. R-432's answer, and the naming scheme
**R-432 is ANSWERED, and negatively. It did not need the panel read it was waiting on.**
* **The naming scheme is `YYYY-MM-DDTHH-MM-SS`** — vendor-documented examples `2025-12-03T13-47-47`,
`2025-02-12T11-35-19`. Recorded so nobody hunts a console again.
* **Knowing it does not help.** Every name in that format for nine days is refused, and the st_dev
split above says why. **Per-file recovery is not operator-only — from the box it is nobody's,** and
for the operator it is a browser act against the main account that no credential in this project
can perform.
* **The panel cannot supply the missing piece either.** It offers Restore and Delete on a row and
does not show names; and the one name-shaped thing it could give would be tried against a door
that leads to the wrong dataset.
## 4. The alarm's first real firing
**It did not happen, and the drill as written could not have produced it.** The detector fires on a
fall of **more than half** the previous count **and at least 5** (`hub/internal/monitor/offsite.go`,
`snapshotDropFraction = 0.5`, `snapshotDropFloor = 5`). demo-hp's baseline is **69**. Phase 2 deletes
**one app's** history — about **9** snapshots. 9 is over the floor and nowhere near half, so the
alarm stays silent, **correctly and by design**. Firing it for real needs ~35+ snapshots destroyed,
i.e. most of demo-hp's off-site history. That trade is what the operator was asked to rule on. → **R-435**
**One thing the alarm says is now wrong.** Its message, live in hub 0.111.0, reads:
> "The daily Storage Box snapshots are read-only and still hold the older copy, **so this is
> recoverable file-by-file**; it is NOT confirmed data loss."
The first clause is true; **the second promises a recovery the product cannot perform and the
operator cannot perform without a browser and the main account.** This is this project's own
corollary — *when a verdict changes which field it counts from, the alarm text has to change with
it* — landing on the alarm shipped the same day. → **R-434**
## 5. Findings, as register rows
All four filed in `documentation/backlog/OPEN-ITEMS.md`.
| row | finding |
| fragment | in the deployed binary |
|---|---|
| **R-433** | A sub-account cannot reach any Storage Box snapshot **by any name**; `/home` and `/.zfs` are different filesystems and `/home/.zfs` does not exist. Answers R-432 negatively and removes clause (b) of the R-95 re-scope. |
| **R-434** | `emitSnapshotDrop`'s message promises file-by-file recovery that is not reachable. Live in hub 0.111.0. |
| **R-435** | The drop detector cannot see a single-app deletion (>50% of 69 ⇒ ~35 needed). `offbox.go:1388` forgets **by tag**, so a single-tag wipe is exactly the shape the detector is blind to. Deliberate insensitivity, but the blind spot should be stated where the operator reads it. |
| **R-436** | **LEAD, not a defect.** Hetzner's port-23 shell offers `rclone serve restic --stdio` as a server-side backend, and restic 0.14.0 recognises the `rclone:` backend (measured; control `banana:` → invalid backend; rclone is absent from the controller image). `rclone serve restic` carries `--append-only`. **This could make R-95's real prevention far cheaper than the spike concluded — no new always-on machine, no data migration.** Caveat stated up front: the **client** supplies the server command line, so a compromised box could omit the flag unless the provider pins it. Settling that is a vendor question, not a code change. |
| `recoverable file-by-file` | **absent** |
| `recoverable file by file` | **absent** |
| `so this is recoverable` | **absent** |
| `it is NOT confirmed data loss` | **absent** |
| *(negative control)* `zzz-no-such-string-r434` | absent — so the search can report absence |
| *(positive control)* `offsite_snapshots_dropped` | present, 2 occurrences |
**R-432 is marked ANSWERED**; its "one panel read settles it" next step is withdrawn as unnecessary.
All seven fragments of the new sentence were checked individually and are present.
## 6. Does `07` §8 row 10 move?
**THE FIX IS A DELETION, NOT A REPLACEMENT — and that is what unblocked it.** R-434's own row said
the fix was *"blocked on R-433"*, i.e. on first establishing what IS true. **That verdict was mine
and it was wrong.** A sentence that asserts **neither** loss **nor** recovery is true under every
possible answer to the provider questions, so it never needs a second rewrite. A *replacement* would
have been blocked; a *withdrawal* is not. The reasoning is recorded in R-434's closing cell and in
the function's doc comment, because the distinction is the transferable part.
**No. It stays `PARTIAL`, and its RTO stays blank.** The status was already correct for the right
reason — *"the recovery ROUTE has never been walked, which is what PARTIAL means"* — and this drill
found the route is not walkable from the box at all. **What the row needs is a text correction, not a
status change:** its clause *"recoverable per-file (vendor)"* and its limit *"per-file recovery is
operator-only today (R-432)"* both overstate what exists. Updated in place with the citation. Moving
it only as far as the evidence goes means not moving it.
**It does not swing the other way either.** `TestR434_AlarmStillDoesNotClaimDataLoss` pins that:
"your backups are gone" is still usually false — the snapshots exist and hold the older copy; what is
unproven is our route to them. Clause (a) of the 2026-09-01 measurement — the box **cannot write**
into the snapshot area — stands and is re-confirmed.
## 7. What could not be tested, and why
### Tests, and the red-proof
* **Whether the main account can see the snapshots.** No main-account credential exists in this
project — the hub holds only per-customer sub-accounts. This is the one question that would decide
whether per-file recovery exists *at all*, for anyone.
* **Whether the Hetzner API can list or read a snapshot.** Fenced by the runbook (§11-D). Separately,
`hub/internal/hetznerapi/hetznerapi.go` has **no snapshot method** — so this route needs new code
regardless of the fence.
* **The deletion, the alarm's first real firing, the recovery, the RTO, the re-arm.** Phases 2–5, not
run, on the operator's ruling.
* **Whether `rclone serve restic --stdio` is pinned server-side with `--append-only`** (R-436).
Three tests in `hub/internal/monitor/offsite_r434_test.go`, all driving the production path
(`saveOffsiteReport` → `oc.Check()` → the notify callback), so they assert the sentence an operator
**receives** rather than the function that formats it. ASCII-only fragments, with a positive control
(a phrase in every version of the alarm) and a negative control (a phrase that cannot exist).
## 8. My own mistakes
**RED-PROOF (run before the fix was restored):** the v0.111.0 sentence was put back and **all three
FAILED** — on `recoverable file-by-file`, on `so this is recoverable`, on all three required
fragments, on the un-negated `confirmed data loss`, and on the stored row — each with the offending
sentence printed in the failure message. Fix restored, `git diff` clean, full hub suite green.
* **I ran Phase 1 before finishing Phase 0's subject-app choice, and did not say so up front.** The
choice depends on the snapshot's timestamp, so the order was right, but the runbook's order is the
runbook's and a silent reordering is the thing this project keeps getting caught by. Stated at the
time in the session, recorded here.
* **My first sweep guessed the schedule instead of establishing it.** I probed 00:00 UTC and 22:00
UTC — 600 names — on the strength of a register line reading *"daily 00:00"*, got nothing, and only
then widened to whole days. The narrow sweep was worth nothing on its own: a zero over a guessed
window is not evidence, and I should have gone to full days first or not run it at all.
* **I nearly reported the empty listing as "the display toggle is hiding it".** The vendor documents
exactly such a toggle and it fitted. The st_dev comparison — which I only ran because `df` printed
a filesystem name I did not expect — says the tree is on another dataset entirely. **A plausible
cause that fits the symptom is not a measured one**, and I had the wrong one for about ten minutes.
* **`REPORT.md` held the only copy of the R-331 report** (hub v0.109.0, 2026-08-30) — durable content
living only in the overwritten file, which `CLAUDE.md:82-87` forbids. Preserved as
`REPORT-r331-backup-card.md` before this report replaced it.
**One existing test was edited, and it caught this fix correctly.**
`TestR431_FiresOnAMassDeletion` asserted the fragment `"NOT confirmed data loss"` and went red on the
new wording. **The fragment was REMOVED, not updated** — the wording now has one home
(`offsite_r434_test.go`), because duplicating it creates the second source that makes the next
correction land in one file and not the other. Its signal fragments (`69`, `4`, `read-only`) are
unchanged.
### R-435 written into the alarm's own documentation, no threshold changed
The comment above `snapshotDropFraction` now states what the detector does **not** see: **a mass
deletion, yes; one app being wiped, no.** demo-hp's baseline is 69 across 9 apps, so ~35 must go
before it speaks and one app's tag is ~9; `offbox.go:1388` runs `forget --prune` grouped by
`host,tags`, so the blind spot sits on the most likely single-app failure. The comment says
explicitly that the numbers must **not** be lowered to "fix" this, and that per-app detection needs a
**second** signal keyed on the per-tag count.
## 2. The stopping line, as it now reads in all three places
**Register — `documentation/backlog/OPEN-ITEMS.md`**, a new section in the voice this file uses for a
settled decision:
> **DECIDED — the backup and restore arc is CLOSED FOR BETA (2026-09-01)** … **CLOSED FOR BETA at
> controller v0.232.0 / hub v0.111.1.** … **What is finished, and proven live:** everything a
> customer does for themselves … **Rows 1, 2, 3, 3b, 3c, 6, 7 and 14 of `07` §8 — every one PROVEN.**
> … **THIS REOPENS IF:** a customer-facing recovery path is found broken; **or** Hetzner's answers
> change what the snapshots are worth …; **or** a real customer's data is at stake in one of the
> deferred rows.
**Architecture — `documentation/architecture/07-backup-architecture.md`, at the head of §8**, so a
reader of the matrix meets it before the blanks:
> **THE ARC IS CLOSED FOR BETA — read this before the blanks below (2026-09-01)** … **The blanks in
> the RTO column below are now blank ON PURPOSE, and that is the whole difference.** … **NO STATUS
> MOVED ON THE DAY THIS WAS WRITTEN, because nothing was proven that day. A stopping line that
> promotes a row is a stopping line that lies.**
**Operator page — `STATUS.md`**, under *Decided — and what would reopen each*:
> **THE BACKUP AND RESTORE WORK IS FINISHED FOR BETA. DECIDED 2026-09-01.** … **What is parked until
> after beta:** everything **only I do, with you** — rebuilding a machine as itself, losing a whole
> box, recovering from ransomware, restoring the hub, and losing Hetzner. **Six of these have never
> been timed, and the hub has never been restored.** They are written down, they are real, and **none
> of them stops a beta customer.**
## 3. The deferred set — row numbers, not a description
`07` §8 rows **4, 8, 9, 10, 11 (and 11b), 12**, each tagged **`[BETA-DEFERRED]`** in its status cell.
| row | failure | status today |
|---|---|---|
| **4** | primary drive dies — the drive-loss **journey** | `PARTIAL` |
| **8** | host dies, drives intact — a host rebuilt as itself | `IMPLEMENTED`, never executed |
| **9** | whole box lost (fire/theft) | `IMPLEMENTED / UNPROVEN` |
| **10** | ransomware / malicious deletion | `PARTIAL` |
| **11** (+**11b**) | hub lost — a hub restore | `UNPROVEN`, never performed |
| **12** | off-site provider lost (Hetzner) | `[FACT]` only |
`grep -n '\[BETA-DEFERRED\]' documentation/architecture/07-backup-architecture.md` returns **eight**
lines — the seven tagged rows plus the one line in §8's header that defines the marker. That is
stated in the register rather than left for the reader to trip over.
**A NUMBER IN THE BRIEF WAS WRONG AND IS CORRECTED IN PLACE.** The brief said *"six rows of §8 still
have no measured time"*. **Six rows are DEFERRED; ELEVEN carry a blank RTO** — counted, not
estimated: 4, 5, 8, 9, 10, 11, 11b, 12, 13, 14, 15. The other five are blank for reasons that are not
deferred work, and collapsing them into one number is how a blank stops meaning anything:
* **row 5** — `PROVEN` by construction; a derived copy, so there is no recovery to time.
* **row 13** — `NONE for host-loss` **by design**; R exists in zero system copies.
* **row 14** — `PROVEN`; the break-glass route works and has simply never been stopwatched.
* **row 11b** — a consequences note attached to row 11, not a recovery row.
* **row 15** — an **open DEFECT** (R-104, the stale-lock path). **It is NOT inside the stopping line**
and must not be read as parked by it. This one matters most: parking a live defect by accident is
the failure mode a stopping line invites.
**No status was moved.** Nothing was proven today.
## 4. The two questions
`documentation/runbooks/provider-questions-2026-09-01.md` — both drafted ready to paste, linked from
R-95 and R-433, **not sent**, and **no provider API was called** (§11-D stands).
* **Q1 — can the MAIN account retrieve individual files from a snapshot, without a whole-box restore?**
*Why:* if it cannot, the only route is a rollback that hits every customer on the box and destroys
newer snapshots — the snapshots would then protect almost nobody. **Decides how urgent R-95 is.**
* **Q2 — on `rclone serve restic --stdio`, is `--append-only` enforced server-side or chosen by the
client?** *Why:* if enforced, the box cannot delete its own backups, with no new machine and no
data move. **Could make R-95 disappear.**
Each carries how to read either answer, including the branch where the lead is worthless and should
be recorded as dead rather than left looking promising. **A dated DUE-CHECKS row (R-433, 2026-09-15)
tracks the reply** — the 2026-07-27 snapshot check that sat unconfirmed for 36 days is the scar that
block exists for, and it was never entered.
## 5. The register
| | before | after |
|---|---|---|
| file lines | **621** | **688** |
| total `R-` rows | **181** | **182** |
| open-state rows | **170** | **169** |
| closed / decided / answered rows | **11** | **13** |
*(The brief's "620 lines and 181 open" matches the file length and the row count; the count of rows
whose leading verdict is actually open was 170.)*
**Rows closed:** **R-434** — with the deletion-not-replacement reasoning, and with the fact that its
own "blocked on R-433" verdict was wrong recorded in the closing cell.
**Rows kept open and marked:**
* **R-95**, **R-433** — **`BLOCKED-ON-PROVIDER`**, both pointing at Part 3's file. R-95 records that
it spent **one day demoted on a clause that did not hold**, and that on today's evidence it belongs
back near the top. **That is a proposal. I have not re-ranked it; the order is Viktor's.**
* **R-430** — **`LATENT`**, with the trigger stated as a trigger: it becomes live **the moment delete
is withdrawn from the box**, which is exactly what R-95's remedy does by either route. So it is a
**precondition on the R-95 build, not a follow-up** — settle it in the same change or the
crash-lock self-heal ships already broken.
* **R-435** — open. The limitation is now in the code's own documentation, but **documenting a blind
spot is not covering it**, and `STATUS.md` and R-431 both still say "noticed within a day".
**Row filed:** **R-437** — the register compression sweep, owed and scoped.
**Compression was measured and deliberately not run, and the reason is one line as the standing rule
requires:** only **12 rows / ~25 KB of 316 KB (about 7 %)** carry a closed leading verdict, so the
sweep buys little and touches everything — and it is the exact operation that misfiled seven rows in
August (R-378, and the seventh, R-87, sat wrong for nine days — R-405). Running it as the tail end of
a session about something else is how that happened the first time. R-437 carries the scope so it can
be picked up cold.
**Also corrected in place:** the ranking paragraph's item 1 said the snapshot mitigation was *"armed
(daily 00:00, keep 7), but it has taken zero snapshots so far"*. **Both halves were wrong** — seven
snapshots exist (R-429) and "armed" was withdrawn the same day. **The ORDER of that list is
unchanged.**
## 6. Hub deploy and its verification
| step | result |
|---|---|
| image built + pushed | `gitea.dooplex.hu/admin/felhom-hub:0.111.1` (25 M) |
| clean-tree gate before build | `git status --porcelain` empty, `HEAD == origin/main` @ `db38f4c` |
| manifest bumped | `manifests/hub.yaml:128` → `:0.111.1`, committed `0f65f7a`, pushed |
| ArgoCD hard refresh + **deliberate** sync | `sync=Synced` `health=Healthy` `rev=0f65f7a…` — the revision equals HEAD |
| rollout | `deployment "hub" successfully rolled out` |
| deployed image | `gitea.dooplex.hu/admin/felhom-hub:0.111.1` |
| running binary | `felhom-hub 0.111.1 (built 2026-09-01T16:35:04Z)` |
No `kubectl set image`, no `kubectl apply`.
### The live trigger was NOT run, and the reason is not a shortcut
The brief asked me to *"trigger the alarm once through the hub's own path"*. **I did not, and I am
naming it rather than quietly substituting.**
**This alarm only fires on a real fall in a real customer's snapshot count.** Firing it live therefore
means one of two things: deleting a real customer's off-site backups (the destructive act this
session explicitly is not), or **POSTing a falsified report claiming `demo-hp` had lost its
backups** — which would write a fabricated data point permanently into that customer's report
history, move its drop latch and baseline, and send Viktor a **second** alarm mail about a real box,
five hours after the first one already confused him (`STATUS.md` item 5). Fabricating customer
telemetry to satisfy a verification step is the wrong trade in a project whose whole doctrine is that
a measurement must mean what it says.
**What was done instead covers both halves of what the live trigger was for:**
1. **Does the deployed artifact carry the sentence?** Proved on the bytes of the running pod's
binary — §1, with both controls.
2. **Does the production path emit it?** Proved by three tests that drive
`saveOffsiteReport → Check() → notify` and read the message the operator would receive, plus the
stored-row test that pins mail and audit row together — and all three red-proofed.
**What is therefore still unproven, stated plainly:** that the *dispatcher* delivers this particular
message to a mailbox. That leg was exercised for real on 2026-09-01 at 12:29 UTC by the previous
session with the old text, so the routing is known good; only the new wording has not travelled it.
## 7. Explicitly
**No controller release. No agent release. No image built for either. No golden baked, and none
owed** — the golden-currency gate reads `newest released controller 0.232.0 / newest golden baked
0.232.0`. **The floor is unchanged at 0.232.0.** `felhom-controller` and `felhom-agent` working trees
were not touched.
`python3 scripts/unproven.py --summary`: **NOT WALKED 35 of 55 — unchanged.** No claim moved, which
is correct: nothing was proven today. All **14** `felhom.eu` gates green before every push.
## 8. Observations
1. **`strings` stops at the em dash, so the alarm sentence appeared truncated in the deployed
binary and briefly looked like a bad build.** `strings` scans ASCII by default and the message's
`—` terminates the run; the fix is `LC_ALL=C grep -aoP` on the bytes.
**NOT-A-FINDING:** this is the project's already-recorded accented-grep trap appearing on a new
surface, so it needs no new row — it is a hazard of my verification method, not a defect in any
product code. Recorded here so the next person grepping a binary for Hungarian or em-dashed copy
does not read a truncation as a bad build, which is exactly how it read for a minute.
2. **The dispatcher's cooldowns are in-memory and are lost on every hub restart**, so any deploy
re-arms every alarm's 6-hour cooldown. **NOT-A-FINDING:** it is stated in
`hub/internal/notify/dispatcher.go:18` as a deliberate, accepted trade. Noted because it is why a
live trigger today would definitely have mailed, rather than being absorbed by yesterday's
cooldown — it changed the decision in §6.
3. **The brief's baseline for `felhom-agent` was two commits stale** — it names `058b945`
(2026-08-23) while `main` is `4586f0f` (2026-09-01). **NOT-A-FINDING:** the agent was untouched
either way, and `058b945` is a real commit, so nothing was ambiguous. Flagged only so the number
is not copied forward into the next brief.
4. **`OPEN-ITEMS.md` has no section called *"Decided — and what would reopen each"*; `STATUS.md`
does.** The brief sent the register text to that section by name.
**NOT-A-FINDING:** resolved in the session by writing a new register section in that section's
voice — the decision, then the condition that reopens it — which is plainly what the instruction
meant, so there is nothing left to file. Recorded only so the next session does not hunt
`OPEN-ITEMS.md` for a heading that has never existed there.
## 9. My own mistakes
* **I wrote "blocked on R-433" into R-434 yesterday, and it was wrong.** It cost nothing because the
block lasted one day, but the reasoning error is the interesting part: I treated *"we do not know
what is true"* as a reason not to touch a sentence that was **known to be false**. Removing a false
claim never needs the true one. The row and the code comment now say so, and it is the only part of
this session I would call a lesson rather than a task.
* **I asserted the brief's "six rows" instead of counting, for about ten minutes.** I wrote the
register block with "six" in it before running the count that produced eleven, and only caught it
because I decided to enumerate the rows rather than describe them — which the brief had insisted
on for a different reason. **The instruction that saved it was not the one aimed at this.**
* **My first `[BETA-DEFERRED]` claim over-promised.** I wrote in the register that the grep "returns
the set as a group and nothing else", then found the grep returns eight lines because §8's own
header defines the marker. Corrected in place with the real count. It is a small instance of
exactly the class the register spent 2026-09-01 documenting — a claim about an instrument that had
not been run.
* **I preserved `REPORT.md` twice in one day and should have noticed the pattern the first time.**
`REPORT.md` held the only copy of the R-331 report this morning and the only copy of the drill
report this afternoon; both are now siblings (`REPORT-r331-backup-card.md`,
`REPORT-drill-r95-recovery-2026-09-01.md`). The convention says durable content may not live only
in the overwritten file, and it has now been violated twice in a day by two different sessions —
which suggests the convention needs a gate, not more diligence. **Not filed:** I am not filing a
row for it in a session that already declined a register sweep; it is named here for whoever picks
up R-437.
+1
View File
@@ -674,6 +674,7 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server`
| **R-434** | **The snapshot-drop alarm promises a recovery that cannot be performed.** `hub/internal/monitor/offsite.go` `emitSnapshotDrop` ships this text, live in hub **v0.111.0**: *"The daily Storage Box snapshots are read-only and still hold the older copy, **so this is recoverable file-by-file**; it is NOT confirmed data loss."* The first clause is true. The second is not reachable: not by the product (R-433), and not by the operator without a browser and a main-account credential that does not exist here. Its own comment states the intent — *"THE MESSAGE MUST NOT SAY THE DATA IS LOST, because after the 2026-09-01 measurement that is usually false"* — and the measurement it rests on was superseded the same day. **This is this project's own corollary landing on the alarm shipped that morning:** when a verdict changes which fact it counts from, the alarm text has to change with it, or the operator acts on a promise nobody can keep. **Fix is text-only and must not be made before R-433 settles what IS true** — an alarm rewritten twice in a week is worse than one rewritten once. **✅ CLOSED 2026-09-01, hub v0.111.1 — AND THE BLOCK ABOVE WAS WRONG, WHICH IS THE POINT WORTH KEEPING.** This row said the fix had to wait for R-433 to establish what IS true. **It did not, because the fix is a DELETION and not a REPLACEMENT.** The promise was withdrawn rather than swapped for a new one: *"The daily Storage Box snapshots are read-only and still hold the older copy. The route back out of them is not yet established, so treat this as neither confirmed data loss nor confirmed recovery. Get in touch before restoring anything, and check whether a deletion ran on the box."* **That sentence is true under EVERY possible answer to the provider questions, so it never needs a second rewrite** — which is the whole reason it was not blocked. A replacement would have been. **Three tests in `hub/internal/monitor/offsite_r434_test.go`**, all driving the production path so they assert the sentence an operator RECEIVES: the withdrawal is present, the promise is absent in three shapes, `confirmed data loss` may appear only inside its negation, and the stored row must not drift from the delivered mail. **RED-PROOF: restoring the v0.111.0 sentence failed all three**, on every fragment, with the offending sentence printed. **One existing test was edited and it had caught this fix correctly** — `TestR431_FiresOnAMassDeletion` asserted `"NOT confirmed data loss"`; the fragment was REMOVED rather than updated so the wording keeps ONE home. | **CLOSED 2026-09-01 — hub v0.111.1; the "blocked on R-433" verdict above was mine and it was wrong** |
| **R-435** | **The snapshot-drop detector is blind to a single-app deletion — the exact shape the box can produce.** `snapshotDropFraction = 0.5` and `snapshotDropFloor = 5` (`hub/internal/monitor/offsite.go`) require a fall of MORE than half the previous count. demo-hp's baseline is **69** across 9 apps, so ~35 snapshots must go before it speaks; **one app's tag is ~9 and is invisible.** `offbox.go:1388` runs `forget --prune` **grouped by host,tags** — a per-tag wipe is precisely the shape a faulty retention or a targeted deletion produces. **This is deliberate, not accidental:** the constant's own comment argues the insensitivity, and *"a detector that cries wolf is switched off within a fortnight"* is a lesson this project paid for. **So this row is NOT a demand to lower the threshold.** It is a demand that the blind spot be written where the operator reads it, because "an unexplained fall is noticed within a day" (R-431, and `STATUS.md`) is true only of falls above half. **Discovered by arithmetic while planning the drill's Phase 2/3 pairing, which could not have worked:** Phase 2 deletes one app and Phase 3 expects the alarm to fire. **LIMITATION NOW WRITTEN INTO THE ALARM'S OWN DOCUMENTATION, hub v0.111.1** — the comment above `snapshotDropFraction` in `hub/internal/monitor/offsite.go` now states what the detector does NOT see, with the demo-hp arithmetic and the `forget --prune` grouping that makes the blind spot sit on the most likely single-app failure. **It also says explicitly that the numbers must NOT be lowered to "fix" this** and that per-app detection needs a SECOND signal keyed on the per-tag count. **The row stays OPEN because documenting a blind spot is not covering it** — and because `STATUS.md` and R-431 both still say "noticed within a day", which is true only of falls above half. | **OPEN — documented in code v0.111.1; the coverage gap itself is unclosed** |
| **R-436** | **LEAD, NOT A DEFECT — append-only may be reachable without a new machine, which would make R-95's real prevention far cheaper than the spike concluded.** Hetzner's port-23 restricted shell advertises, in its own `help`, these server-side backends: `borg`, `rsync`, `scp`, `sftp`, **`rclone serve restic --stdio`**. And restic 0.14.0 **recognises the `rclone:` backend** — MEASURED 2026-09-01 with a control: `banana:` → `Fatal: parsing repository location failed: invalid backend`, while `rclone:` → `exec: "rclone": executable file not found in $PATH` (i.e. the backend parsed and it tried to run the helper). rclone is **not** in the controller image today. `rclone serve restic` carries an **`--append-only`** flag. **Why this matters:** the R-95 spike's option 3 was priced at a new always-on service in the recovery path plus either a mount in the hot path or migrating every customer's history — and it was deferred on that price. This route needs neither: the server side already runs at the provider. **THE CAVEAT, STATED FIRST because it may kill the idea:** the **client** supplies the server command line, so a compromised guest could simply omit `--append-only` unless the provider pins it. **NOT ESTABLISHED:** whether Hetzner pins the flag or accepts client-supplied arguments. **That is a vendor question and it is cheap** — it should be asked before any code is written, because if the answer is "client-supplied" this lead is worth nothing. | **OPEN — ask the vendor before building anything** |
| **R-437** | **The register compression sweep is OWED, and it was deliberately NOT run inside the 2026-09-01 beta-line session — this row is the record of that choice, not a note.** **The ask:** compress what has closed in `OPEN-ITEMS.md`. **The measurement, taken before deciding:** 181 rows, 316 KB of row text, of which **12 rows / ~25 KB (about 7 %) carry a CLOSED/DECIDED/ANSWERED leading verdict.** So the sweep buys little and touches everything. **Why it was refused as a side-task, and the citation matters:** a compression sweep is the exact operation that has already gone wrong here. The 2026-08-22 sweep (`ef6ac6f`, R-376..R-378) matched a status word ANYWHERE in the row, moved rows that were not closed, and **R-378 caught six in the same session and missed a seventh** — R-87 sat in the wrong register for nine days while the ranking paragraph pointed at nothing (R-405). **That is a session-scale hazard, and running it as the tail end of a session about something else is how it happened the first time.** **WHAT IS OWED, scoped so it can be picked up cold:** (1) classify by the **LEADING VERDICT** of the state cell only — the rule `closed_register_gate.py` already implements and red-proofs, never a whole-row match; (2) move, never rewrite — a compressed row that loses its evidence is worse than a long one; (3) run `closed_register_gate.py` before and after and quote both; (4) re-read the ranking paragraph afterwards, because that is the surface that silently went stale last time. **Not urgent:** the file is 688 lines and every gate reads it in well under a second. | **OPEN — owed; needs its own session, not a tail end** |
<!-- DUE-CHECKS-BEGIN — machine-readable. Parsed by scripts/due_checks_gate.py.
One row per dated check. The R-number must have a row above. Dates are UTC.