REPORT + R-437: the beta stopping line recorded, and the live trigger declined with its reason
gates / gates (push) Successful in 19s
gates / gates (push) Successful in 19s
REPORT.md carries the deployed sentence quoted from the RUNNING binary (kubectl cp + byte grep, both controls), the stopping line as it reads in all three places, the enumerated deferred set, the two provider questions, the register census, and the ArgoCD verification. R-437 filed: the register compression sweep is OWED and was deliberately not run here. Measured first — 12 of 181 rows / ~25 KB of 316 KB (about 7%) carry a closed leading verdict — so it buys little and touches everything, and it is the exact operation that misfiled seven rows in August (R-378; the seventh, R-87, sat wrong for nine days, R-405). The row carries the scope so it can be picked up cold. The live alarm trigger was NOT run and the report says so in its own section rather than substituting quietly: this alarm only fires on a real fall in a real customer's snapshot count, so firing it means either deleting real backups or POSTing a falsified report claiming demo-hp lost its own. That would write a fabricated point into a customer's report history, move its latch and baseline, and mail the operator a second alarm about a real box hours after the first one already confused him. Covered instead by the deployed-bytes proof plus three red-proofed tests driving saveReport -> Check -> notify. What remains unproven is named: that the dispatcher delivers THIS wording to a mailbox. Register 688 -> 700 lines; 182 rows; open-state 170. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01LB8FmJaGd2cyjvy6dbEjpM
This commit is contained in:
@@ -1,166 +1,294 @@
|
||||
# REPORT — the deletion we said is survivable: the recovery route does not exist (R-95 drill, 2026-09-01)
|
||||
# REPORT — the line under the backup arc, and one alarm that was telling the operator something untrue
|
||||
|
||||
**RUNBOOK, destructive class, `demo-hp` only. STOPPED at the end of Phase 1 on the operator's ruling,
|
||||
before any destructive step. No delete verb was issued against any live store; no byte on either
|
||||
Storage Box sub-account was written, moved or removed.** No production code, no version bump, no
|
||||
image, no golden. Evidence: `documentation/audits/evidence-drill-r95-recovery-2026-09-01/`.
|
||||
**hub v0.111.0 → v0.111.1 · controller v0.232.0 UNTOUCHED · 2026-09-01**
|
||||
|
||||
| # | phase | verdict | one sentence |
|
||||
|---|---|---|---|
|
||||
| 1 | snapshot reachable, and its name | **NO — and it has no reachable name** | 777,600 exact names across nine days in the vendor's own format, plus 126 alternative shapes; zero resolve, with a control proving the sweep detects a path that exists. |
|
||||
| 2 | the deletion | **NOT RUN — operator ruling** | With no recovery route, the deletion would have destroyed real history to buy nothing; put as a two-option decision, the ruling was stop. |
|
||||
| 3 | the alarm fired | **NOT RUN — and it could not have fired at the specified size** | The shipped threshold needs a fall of more than half; one app's tag is ~9 of 69. → **R-435** |
|
||||
| 4 | **the recovery** | **NOT RUN — no route exists that is not fenced** | Box-side: proven impossible. Panel: fenced and browserless. Hetzner API: fenced (§11-D), and the hub's client has no snapshot method at all. |
|
||||
| — | **RTO from T₀** | **STILL BLANK** | Row 10's RTO cell is unchanged and remains a finding. |
|
||||
| — | **data lost, quantified** | **NOT MEASURABLE THIS WAY** | The quantity only has meaning if the rest is recoverable, and the route that would recover it is not reachable. |
|
||||
| 5 | re-arm | **NOT RUN** | Depended on Phase 4. |
|
||||
| 6 | teardown | **PASS** | Store untouched at 69 snapshots; all three scratch layers removed; hub DB copies shredded; both boxes healthy. |
|
||||
**No controller release. No agent release. No golden owed, no floor change.** The only behaviour
|
||||
change in this session is one sentence in one alarm. Everything else is register and documentation.
|
||||
|
||||
---
|
||||
|
||||
## 1. Did the recovery work — and does yesterday's re-scope survive?
|
||||
## 1. The alarm's new text, and proof the old promise is gone
|
||||
|
||||
**The recovery was never reachable, and the re-scope does not survive intact. Its first half stands;
|
||||
its second half does not.**
|
||||
|
||||
Yesterday's re-scope has two clauses. They must now be separated:
|
||||
|
||||
* **(a) "The box can delete its live repository, but cannot write to the daily snapshots of it."**
|
||||
**STANDS.** Re-confirmed here: `/.zfs/snapshot` is reachable and the write-refusal measurement is
|
||||
unchanged. Nothing in this drill weakens it.
|
||||
* **(b) "…so the rest is recoverable — file by file, one customer at a time."** **NOT SUPPORTED.**
|
||||
A snapshot that cannot be opened cannot be copied out of. R-432 recorded the directory listing
|
||||
empty and named the cheapest next step: *"a single `ls /.zfs/snapshot/<name>` from a box then
|
||||
settles whether a named snapshot can be entered even though the directory does not list (ZFS
|
||||
allows exactly that)."* **That step is now done, exhaustively, and the answer is no.**
|
||||
|
||||
**What was measured.** The port-23 restricted shell accepts a batched `stat`, which makes a cheap
|
||||
existence oracle: 500–600 paths per round trip, stdout carrying only paths that exist.
|
||||
|
||||
| sweep | candidates | hits |
|
||||
|---|---|---|
|
||||
| `/.zfs/snapshot/YYYY-MM-DDTHH-MM-SS`, nine full days, second granularity | **777,600** | **0** |
|
||||
| 126 alternative name shapes and snapshot paths (`daily`, `snapshot-1`, colon and compact time forms, `/home/.snapshot`, …) | 126 | 0 |
|
||||
| **control — the identical 600-name batch shape with one real path appended** | 6 batches | **6/6 returned it** |
|
||||
|
||||
**And there is a structural reason, which is why I stopped sweeping.** The customer's data and the
|
||||
snapshot door are on **different filesystems**:
|
||||
**Quoted from the RUNNING binary** — `kubectl cp` out of pod `hub-857678f9b4-v95s6`, byte-grepped:
|
||||
|
||||
```
|
||||
df → u629488-sub3 mounted on /home
|
||||
stat /home → Device 0,82
|
||||
stat /.zfs/snapshot → Device 0,276 ← a different device
|
||||
stat /home/.zfs → cannot statx: No such file or directory
|
||||
Customer %s: off-site backup count fell from %d to %d snapshot(s) in one report — more than
|
||||
retention can explain. The daily Storage Box snapshots are read-only and still hold the older
|
||||
copy. The route back out of them is not yet established, so treat this as neither confirmed
|
||||
data loss nor confirmed recovery. Get in touch before restoring anything, and check whether a
|
||||
deletion ran on the box.
|
||||
```
|
||||
|
||||
A ZFS snapshot under `/.zfs/snapshot` belongs to the dataset that owns that `.zfs` — not to the
|
||||
child mounted at `/home`. **So even a correctly named snapshot there could not contain
|
||||
`felhom-repo`,** and the dataset that does hold it exposes no `.zfs` at all to this account. The
|
||||
empty listing is not a display toggle hiding a reachable tree; from a sub-account there is no tree.
|
||||
**Every shape of the withdrawn promise is absent from the deployed bytes:**
|
||||
|
||||
**Three tools agree, each with controls in the same run:** SFTP, the port-23 shell, and
|
||||
`rsync --list-only`.
|
||||
|
||||
**What that does to R-95.** Its *exposure* is unchanged and its *remedy* is not. Yesterday the row
|
||||
could say a deletion costs about a day because the rest comes back per-file. Today the only routes
|
||||
to "the rest" are a whole-box panel rollback (which deletes newer snapshots and hits every customer
|
||||
on the box) and the provider API (fenced, and unimplemented in the hub's client). **The re-scope's
|
||||
comfort was resting on a route nobody had walked — which is precisely the standard this project
|
||||
applies, and it is the reason this drill was called.**
|
||||
|
||||
**The ranking is Viktor's and I am not re-ranking it.** What I will say plainly: the argument that
|
||||
moved R-95 down yesterday is the argument this drill removed. On these facts I would put it back
|
||||
where it was.
|
||||
|
||||
## 2. The RTO
|
||||
|
||||
**Still blank, and it stays a finding.** `07` §8 row 10's RTO cell has been empty since July and this
|
||||
drill did not fill it. Nothing was recovered, so nothing was timed. The summary line at
|
||||
`07-backup-architecture.md:948` — *"no ransomware-shaped recovery has ever been run"* — is still
|
||||
true, and is now true for a sharper reason: **not "nobody has run it" but "from the box, it cannot
|
||||
be run."**
|
||||
|
||||
## 3. R-432's answer, and the naming scheme
|
||||
|
||||
**R-432 is ANSWERED, and negatively. It did not need the panel read it was waiting on.**
|
||||
|
||||
* **The naming scheme is `YYYY-MM-DDTHH-MM-SS`** — vendor-documented examples `2025-12-03T13-47-47`,
|
||||
`2025-02-12T11-35-19`. Recorded so nobody hunts a console again.
|
||||
* **Knowing it does not help.** Every name in that format for nine days is refused, and the st_dev
|
||||
split above says why. **Per-file recovery is not operator-only — from the box it is nobody's,** and
|
||||
for the operator it is a browser act against the main account that no credential in this project
|
||||
can perform.
|
||||
* **The panel cannot supply the missing piece either.** It offers Restore and Delete on a row and
|
||||
does not show names; and the one name-shaped thing it could give would be tried against a door
|
||||
that leads to the wrong dataset.
|
||||
|
||||
## 4. The alarm's first real firing
|
||||
|
||||
**It did not happen, and the drill as written could not have produced it.** The detector fires on a
|
||||
fall of **more than half** the previous count **and at least 5** (`hub/internal/monitor/offsite.go`,
|
||||
`snapshotDropFraction = 0.5`, `snapshotDropFloor = 5`). demo-hp's baseline is **69**. Phase 2 deletes
|
||||
**one app's** history — about **9** snapshots. 9 is over the floor and nowhere near half, so the
|
||||
alarm stays silent, **correctly and by design**. Firing it for real needs ~35+ snapshots destroyed,
|
||||
i.e. most of demo-hp's off-site history. That trade is what the operator was asked to rule on. → **R-435**
|
||||
|
||||
**One thing the alarm says is now wrong.** Its message, live in hub 0.111.0, reads:
|
||||
|
||||
> "The daily Storage Box snapshots are read-only and still hold the older copy, **so this is
|
||||
> recoverable file-by-file**; it is NOT confirmed data loss."
|
||||
|
||||
The first clause is true; **the second promises a recovery the product cannot perform and the
|
||||
operator cannot perform without a browser and the main account.** This is this project's own
|
||||
corollary — *when a verdict changes which field it counts from, the alarm text has to change with
|
||||
it* — landing on the alarm shipped the same day. → **R-434**
|
||||
|
||||
## 5. Findings, as register rows
|
||||
|
||||
All four filed in `documentation/backlog/OPEN-ITEMS.md`.
|
||||
|
||||
| row | finding |
|
||||
| fragment | in the deployed binary |
|
||||
|---|---|
|
||||
| **R-433** | A sub-account cannot reach any Storage Box snapshot **by any name**; `/home` and `/.zfs` are different filesystems and `/home/.zfs` does not exist. Answers R-432 negatively and removes clause (b) of the R-95 re-scope. |
|
||||
| **R-434** | `emitSnapshotDrop`'s message promises file-by-file recovery that is not reachable. Live in hub 0.111.0. |
|
||||
| **R-435** | The drop detector cannot see a single-app deletion (>50% of 69 ⇒ ~35 needed). `offbox.go:1388` forgets **by tag**, so a single-tag wipe is exactly the shape the detector is blind to. Deliberate insensitivity, but the blind spot should be stated where the operator reads it. |
|
||||
| **R-436** | **LEAD, not a defect.** Hetzner's port-23 shell offers `rclone serve restic --stdio` as a server-side backend, and restic 0.14.0 recognises the `rclone:` backend (measured; control `banana:` → invalid backend; rclone is absent from the controller image). `rclone serve restic` carries `--append-only`. **This could make R-95's real prevention far cheaper than the spike concluded — no new always-on machine, no data migration.** Caveat stated up front: the **client** supplies the server command line, so a compromised box could omit the flag unless the provider pins it. Settling that is a vendor question, not a code change. |
|
||||
| `recoverable file-by-file` | **absent** |
|
||||
| `recoverable file by file` | **absent** |
|
||||
| `so this is recoverable` | **absent** |
|
||||
| `it is NOT confirmed data loss` | **absent** |
|
||||
| *(negative control)* `zzz-no-such-string-r434` | absent — so the search can report absence |
|
||||
| *(positive control)* `offsite_snapshots_dropped` | present, 2 occurrences |
|
||||
|
||||
**R-432 is marked ANSWERED**; its "one panel read settles it" next step is withdrawn as unnecessary.
|
||||
All seven fragments of the new sentence were checked individually and are present.
|
||||
|
||||
## 6. Does `07` §8 row 10 move?
|
||||
**THE FIX IS A DELETION, NOT A REPLACEMENT — and that is what unblocked it.** R-434's own row said
|
||||
the fix was *"blocked on R-433"*, i.e. on first establishing what IS true. **That verdict was mine
|
||||
and it was wrong.** A sentence that asserts **neither** loss **nor** recovery is true under every
|
||||
possible answer to the provider questions, so it never needs a second rewrite. A *replacement* would
|
||||
have been blocked; a *withdrawal* is not. The reasoning is recorded in R-434's closing cell and in
|
||||
the function's doc comment, because the distinction is the transferable part.
|
||||
|
||||
**No. It stays `PARTIAL`, and its RTO stays blank.** The status was already correct for the right
|
||||
reason — *"the recovery ROUTE has never been walked, which is what PARTIAL means"* — and this drill
|
||||
found the route is not walkable from the box at all. **What the row needs is a text correction, not a
|
||||
status change:** its clause *"recoverable per-file (vendor)"* and its limit *"per-file recovery is
|
||||
operator-only today (R-432)"* both overstate what exists. Updated in place with the citation. Moving
|
||||
it only as far as the evidence goes means not moving it.
|
||||
**It does not swing the other way either.** `TestR434_AlarmStillDoesNotClaimDataLoss` pins that:
|
||||
"your backups are gone" is still usually false — the snapshots exist and hold the older copy; what is
|
||||
unproven is our route to them. Clause (a) of the 2026-09-01 measurement — the box **cannot write**
|
||||
into the snapshot area — stands and is re-confirmed.
|
||||
|
||||
## 7. What could not be tested, and why
|
||||
### Tests, and the red-proof
|
||||
|
||||
* **Whether the main account can see the snapshots.** No main-account credential exists in this
|
||||
project — the hub holds only per-customer sub-accounts. This is the one question that would decide
|
||||
whether per-file recovery exists *at all*, for anyone.
|
||||
* **Whether the Hetzner API can list or read a snapshot.** Fenced by the runbook (§11-D). Separately,
|
||||
`hub/internal/hetznerapi/hetznerapi.go` has **no snapshot method** — so this route needs new code
|
||||
regardless of the fence.
|
||||
* **The deletion, the alarm's first real firing, the recovery, the RTO, the re-arm.** Phases 2–5, not
|
||||
run, on the operator's ruling.
|
||||
* **Whether `rclone serve restic --stdio` is pinned server-side with `--append-only`** (R-436).
|
||||
Three tests in `hub/internal/monitor/offsite_r434_test.go`, all driving the production path
|
||||
(`saveOffsiteReport` → `oc.Check()` → the notify callback), so they assert the sentence an operator
|
||||
**receives** rather than the function that formats it. ASCII-only fragments, with a positive control
|
||||
(a phrase in every version of the alarm) and a negative control (a phrase that cannot exist).
|
||||
|
||||
## 8. My own mistakes
|
||||
**RED-PROOF (run before the fix was restored):** the v0.111.0 sentence was put back and **all three
|
||||
FAILED** — on `recoverable file-by-file`, on `so this is recoverable`, on all three required
|
||||
fragments, on the un-negated `confirmed data loss`, and on the stored row — each with the offending
|
||||
sentence printed in the failure message. Fix restored, `git diff` clean, full hub suite green.
|
||||
|
||||
* **I ran Phase 1 before finishing Phase 0's subject-app choice, and did not say so up front.** The
|
||||
choice depends on the snapshot's timestamp, so the order was right, but the runbook's order is the
|
||||
runbook's and a silent reordering is the thing this project keeps getting caught by. Stated at the
|
||||
time in the session, recorded here.
|
||||
* **My first sweep guessed the schedule instead of establishing it.** I probed 00:00 UTC and 22:00
|
||||
UTC — 600 names — on the strength of a register line reading *"daily 00:00"*, got nothing, and only
|
||||
then widened to whole days. The narrow sweep was worth nothing on its own: a zero over a guessed
|
||||
window is not evidence, and I should have gone to full days first or not run it at all.
|
||||
* **I nearly reported the empty listing as "the display toggle is hiding it".** The vendor documents
|
||||
exactly such a toggle and it fitted. The st_dev comparison — which I only ran because `df` printed
|
||||
a filesystem name I did not expect — says the tree is on another dataset entirely. **A plausible
|
||||
cause that fits the symptom is not a measured one**, and I had the wrong one for about ten minutes.
|
||||
* **`REPORT.md` held the only copy of the R-331 report** (hub v0.109.0, 2026-08-30) — durable content
|
||||
living only in the overwritten file, which `CLAUDE.md:82-87` forbids. Preserved as
|
||||
`REPORT-r331-backup-card.md` before this report replaced it.
|
||||
**One existing test was edited, and it caught this fix correctly.**
|
||||
`TestR431_FiresOnAMassDeletion` asserted the fragment `"NOT confirmed data loss"` and went red on the
|
||||
new wording. **The fragment was REMOVED, not updated** — the wording now has one home
|
||||
(`offsite_r434_test.go`), because duplicating it creates the second source that makes the next
|
||||
correction land in one file and not the other. Its signal fragments (`69`, `4`, `read-only`) are
|
||||
unchanged.
|
||||
|
||||
### R-435 written into the alarm's own documentation, no threshold changed
|
||||
|
||||
The comment above `snapshotDropFraction` now states what the detector does **not** see: **a mass
|
||||
deletion, yes; one app being wiped, no.** demo-hp's baseline is 69 across 9 apps, so ~35 must go
|
||||
before it speaks and one app's tag is ~9; `offbox.go:1388` runs `forget --prune` grouped by
|
||||
`host,tags`, so the blind spot sits on the most likely single-app failure. The comment says
|
||||
explicitly that the numbers must **not** be lowered to "fix" this, and that per-app detection needs a
|
||||
**second** signal keyed on the per-tag count.
|
||||
|
||||
## 2. The stopping line, as it now reads in all three places
|
||||
|
||||
**Register — `documentation/backlog/OPEN-ITEMS.md`**, a new section in the voice this file uses for a
|
||||
settled decision:
|
||||
|
||||
> **DECIDED — the backup and restore arc is CLOSED FOR BETA (2026-09-01)** … **CLOSED FOR BETA at
|
||||
> controller v0.232.0 / hub v0.111.1.** … **What is finished, and proven live:** everything a
|
||||
> customer does for themselves … **Rows 1, 2, 3, 3b, 3c, 6, 7 and 14 of `07` §8 — every one PROVEN.**
|
||||
> … **THIS REOPENS IF:** a customer-facing recovery path is found broken; **or** Hetzner's answers
|
||||
> change what the snapshots are worth …; **or** a real customer's data is at stake in one of the
|
||||
> deferred rows.
|
||||
|
||||
**Architecture — `documentation/architecture/07-backup-architecture.md`, at the head of §8**, so a
|
||||
reader of the matrix meets it before the blanks:
|
||||
|
||||
> **THE ARC IS CLOSED FOR BETA — read this before the blanks below (2026-09-01)** … **The blanks in
|
||||
> the RTO column below are now blank ON PURPOSE, and that is the whole difference.** … **NO STATUS
|
||||
> MOVED ON THE DAY THIS WAS WRITTEN, because nothing was proven that day. A stopping line that
|
||||
> promotes a row is a stopping line that lies.**
|
||||
|
||||
**Operator page — `STATUS.md`**, under *Decided — and what would reopen each*:
|
||||
|
||||
> **THE BACKUP AND RESTORE WORK IS FINISHED FOR BETA. DECIDED 2026-09-01.** … **What is parked until
|
||||
> after beta:** everything **only I do, with you** — rebuilding a machine as itself, losing a whole
|
||||
> box, recovering from ransomware, restoring the hub, and losing Hetzner. **Six of these have never
|
||||
> been timed, and the hub has never been restored.** They are written down, they are real, and **none
|
||||
> of them stops a beta customer.**
|
||||
|
||||
## 3. The deferred set — row numbers, not a description
|
||||
|
||||
`07` §8 rows **4, 8, 9, 10, 11 (and 11b), 12**, each tagged **`[BETA-DEFERRED]`** in its status cell.
|
||||
|
||||
| row | failure | status today |
|
||||
|---|---|---|
|
||||
| **4** | primary drive dies — the drive-loss **journey** | `PARTIAL` |
|
||||
| **8** | host dies, drives intact — a host rebuilt as itself | `IMPLEMENTED`, never executed |
|
||||
| **9** | whole box lost (fire/theft) | `IMPLEMENTED / UNPROVEN` |
|
||||
| **10** | ransomware / malicious deletion | `PARTIAL` |
|
||||
| **11** (+**11b**) | hub lost — a hub restore | `UNPROVEN`, never performed |
|
||||
| **12** | off-site provider lost (Hetzner) | `[FACT]` only |
|
||||
|
||||
`grep -n '\[BETA-DEFERRED\]' documentation/architecture/07-backup-architecture.md` returns **eight**
|
||||
lines — the seven tagged rows plus the one line in §8's header that defines the marker. That is
|
||||
stated in the register rather than left for the reader to trip over.
|
||||
|
||||
**A NUMBER IN THE BRIEF WAS WRONG AND IS CORRECTED IN PLACE.** The brief said *"six rows of §8 still
|
||||
have no measured time"*. **Six rows are DEFERRED; ELEVEN carry a blank RTO** — counted, not
|
||||
estimated: 4, 5, 8, 9, 10, 11, 11b, 12, 13, 14, 15. The other five are blank for reasons that are not
|
||||
deferred work, and collapsing them into one number is how a blank stops meaning anything:
|
||||
|
||||
* **row 5** — `PROVEN` by construction; a derived copy, so there is no recovery to time.
|
||||
* **row 13** — `NONE for host-loss` **by design**; R exists in zero system copies.
|
||||
* **row 14** — `PROVEN`; the break-glass route works and has simply never been stopwatched.
|
||||
* **row 11b** — a consequences note attached to row 11, not a recovery row.
|
||||
* **row 15** — an **open DEFECT** (R-104, the stale-lock path). **It is NOT inside the stopping line**
|
||||
and must not be read as parked by it. This one matters most: parking a live defect by accident is
|
||||
the failure mode a stopping line invites.
|
||||
|
||||
**No status was moved.** Nothing was proven today.
|
||||
|
||||
## 4. The two questions
|
||||
|
||||
`documentation/runbooks/provider-questions-2026-09-01.md` — both drafted ready to paste, linked from
|
||||
R-95 and R-433, **not sent**, and **no provider API was called** (§11-D stands).
|
||||
|
||||
* **Q1 — can the MAIN account retrieve individual files from a snapshot, without a whole-box restore?**
|
||||
*Why:* if it cannot, the only route is a rollback that hits every customer on the box and destroys
|
||||
newer snapshots — the snapshots would then protect almost nobody. **Decides how urgent R-95 is.**
|
||||
* **Q2 — on `rclone serve restic --stdio`, is `--append-only` enforced server-side or chosen by the
|
||||
client?** *Why:* if enforced, the box cannot delete its own backups, with no new machine and no
|
||||
data move. **Could make R-95 disappear.**
|
||||
|
||||
Each carries how to read either answer, including the branch where the lead is worthless and should
|
||||
be recorded as dead rather than left looking promising. **A dated DUE-CHECKS row (R-433, 2026-09-15)
|
||||
tracks the reply** — the 2026-07-27 snapshot check that sat unconfirmed for 36 days is the scar that
|
||||
block exists for, and it was never entered.
|
||||
|
||||
## 5. The register
|
||||
|
||||
| | before | after |
|
||||
|---|---|---|
|
||||
| file lines | **621** | **688** |
|
||||
| total `R-` rows | **181** | **182** |
|
||||
| open-state rows | **170** | **169** |
|
||||
| closed / decided / answered rows | **11** | **13** |
|
||||
|
||||
*(The brief's "620 lines and 181 open" matches the file length and the row count; the count of rows
|
||||
whose leading verdict is actually open was 170.)*
|
||||
|
||||
**Rows closed:** **R-434** — with the deletion-not-replacement reasoning, and with the fact that its
|
||||
own "blocked on R-433" verdict was wrong recorded in the closing cell.
|
||||
|
||||
**Rows kept open and marked:**
|
||||
|
||||
* **R-95**, **R-433** — **`BLOCKED-ON-PROVIDER`**, both pointing at Part 3's file. R-95 records that
|
||||
it spent **one day demoted on a clause that did not hold**, and that on today's evidence it belongs
|
||||
back near the top. **That is a proposal. I have not re-ranked it; the order is Viktor's.**
|
||||
* **R-430** — **`LATENT`**, with the trigger stated as a trigger: it becomes live **the moment delete
|
||||
is withdrawn from the box**, which is exactly what R-95's remedy does by either route. So it is a
|
||||
**precondition on the R-95 build, not a follow-up** — settle it in the same change or the
|
||||
crash-lock self-heal ships already broken.
|
||||
* **R-435** — open. The limitation is now in the code's own documentation, but **documenting a blind
|
||||
spot is not covering it**, and `STATUS.md` and R-431 both still say "noticed within a day".
|
||||
|
||||
**Row filed:** **R-437** — the register compression sweep, owed and scoped.
|
||||
|
||||
**Compression was measured and deliberately not run, and the reason is one line as the standing rule
|
||||
requires:** only **12 rows / ~25 KB of 316 KB (about 7 %)** carry a closed leading verdict, so the
|
||||
sweep buys little and touches everything — and it is the exact operation that misfiled seven rows in
|
||||
August (R-378, and the seventh, R-87, sat wrong for nine days — R-405). Running it as the tail end of
|
||||
a session about something else is how that happened the first time. R-437 carries the scope so it can
|
||||
be picked up cold.
|
||||
|
||||
**Also corrected in place:** the ranking paragraph's item 1 said the snapshot mitigation was *"armed
|
||||
(daily 00:00, keep 7), but it has taken zero snapshots so far"*. **Both halves were wrong** — seven
|
||||
snapshots exist (R-429) and "armed" was withdrawn the same day. **The ORDER of that list is
|
||||
unchanged.**
|
||||
|
||||
## 6. Hub deploy and its verification
|
||||
|
||||
| step | result |
|
||||
|---|---|
|
||||
| image built + pushed | `gitea.dooplex.hu/admin/felhom-hub:0.111.1` (25 M) |
|
||||
| clean-tree gate before build | `git status --porcelain` empty, `HEAD == origin/main` @ `db38f4c` |
|
||||
| manifest bumped | `manifests/hub.yaml:128` → `:0.111.1`, committed `0f65f7a`, pushed |
|
||||
| ArgoCD hard refresh + **deliberate** sync | `sync=Synced` `health=Healthy` `rev=0f65f7a…` — the revision equals HEAD |
|
||||
| rollout | `deployment "hub" successfully rolled out` |
|
||||
| deployed image | `gitea.dooplex.hu/admin/felhom-hub:0.111.1` |
|
||||
| running binary | `felhom-hub 0.111.1 (built 2026-09-01T16:35:04Z)` |
|
||||
|
||||
No `kubectl set image`, no `kubectl apply`.
|
||||
|
||||
### The live trigger was NOT run, and the reason is not a shortcut
|
||||
|
||||
The brief asked me to *"trigger the alarm once through the hub's own path"*. **I did not, and I am
|
||||
naming it rather than quietly substituting.**
|
||||
|
||||
**This alarm only fires on a real fall in a real customer's snapshot count.** Firing it live therefore
|
||||
means one of two things: deleting a real customer's off-site backups (the destructive act this
|
||||
session explicitly is not), or **POSTing a falsified report claiming `demo-hp` had lost its
|
||||
backups** — which would write a fabricated data point permanently into that customer's report
|
||||
history, move its drop latch and baseline, and send Viktor a **second** alarm mail about a real box,
|
||||
five hours after the first one already confused him (`STATUS.md` item 5). Fabricating customer
|
||||
telemetry to satisfy a verification step is the wrong trade in a project whose whole doctrine is that
|
||||
a measurement must mean what it says.
|
||||
|
||||
**What was done instead covers both halves of what the live trigger was for:**
|
||||
|
||||
1. **Does the deployed artifact carry the sentence?** Proved on the bytes of the running pod's
|
||||
binary — §1, with both controls.
|
||||
2. **Does the production path emit it?** Proved by three tests that drive
|
||||
`saveOffsiteReport → Check() → notify` and read the message the operator would receive, plus the
|
||||
stored-row test that pins mail and audit row together — and all three red-proofed.
|
||||
|
||||
**What is therefore still unproven, stated plainly:** that the *dispatcher* delivers this particular
|
||||
message to a mailbox. That leg was exercised for real on 2026-09-01 at 12:29 UTC by the previous
|
||||
session with the old text, so the routing is known good; only the new wording has not travelled it.
|
||||
|
||||
## 7. Explicitly
|
||||
|
||||
**No controller release. No agent release. No image built for either. No golden baked, and none
|
||||
owed** — the golden-currency gate reads `newest released controller 0.232.0 / newest golden baked
|
||||
0.232.0`. **The floor is unchanged at 0.232.0.** `felhom-controller` and `felhom-agent` working trees
|
||||
were not touched.
|
||||
|
||||
`python3 scripts/unproven.py --summary`: **NOT WALKED 35 of 55 — unchanged.** No claim moved, which
|
||||
is correct: nothing was proven today. All **14** `felhom.eu` gates green before every push.
|
||||
|
||||
## 8. Observations
|
||||
|
||||
1. **`strings` stops at the em dash, so the alarm sentence appeared truncated in the deployed
|
||||
binary and briefly looked like a bad build.** `strings` scans ASCII by default and the message's
|
||||
`—` terminates the run; the fix is `LC_ALL=C grep -aoP` on the bytes.
|
||||
**NOT-A-FINDING:** this is the project's already-recorded accented-grep trap appearing on a new
|
||||
surface, so it needs no new row — it is a hazard of my verification method, not a defect in any
|
||||
product code. Recorded here so the next person grepping a binary for Hungarian or em-dashed copy
|
||||
does not read a truncation as a bad build, which is exactly how it read for a minute.
|
||||
|
||||
2. **The dispatcher's cooldowns are in-memory and are lost on every hub restart**, so any deploy
|
||||
re-arms every alarm's 6-hour cooldown. **NOT-A-FINDING:** it is stated in
|
||||
`hub/internal/notify/dispatcher.go:18` as a deliberate, accepted trade. Noted because it is why a
|
||||
live trigger today would definitely have mailed, rather than being absorbed by yesterday's
|
||||
cooldown — it changed the decision in §6.
|
||||
|
||||
3. **The brief's baseline for `felhom-agent` was two commits stale** — it names `058b945`
|
||||
(2026-08-23) while `main` is `4586f0f` (2026-09-01). **NOT-A-FINDING:** the agent was untouched
|
||||
either way, and `058b945` is a real commit, so nothing was ambiguous. Flagged only so the number
|
||||
is not copied forward into the next brief.
|
||||
|
||||
4. **`OPEN-ITEMS.md` has no section called *"Decided — and what would reopen each"*; `STATUS.md`
|
||||
does.** The brief sent the register text to that section by name.
|
||||
**NOT-A-FINDING:** resolved in the session by writing a new register section in that section's
|
||||
voice — the decision, then the condition that reopens it — which is plainly what the instruction
|
||||
meant, so there is nothing left to file. Recorded only so the next session does not hunt
|
||||
`OPEN-ITEMS.md` for a heading that has never existed there.
|
||||
|
||||
## 9. My own mistakes
|
||||
|
||||
* **I wrote "blocked on R-433" into R-434 yesterday, and it was wrong.** It cost nothing because the
|
||||
block lasted one day, but the reasoning error is the interesting part: I treated *"we do not know
|
||||
what is true"* as a reason not to touch a sentence that was **known to be false**. Removing a false
|
||||
claim never needs the true one. The row and the code comment now say so, and it is the only part of
|
||||
this session I would call a lesson rather than a task.
|
||||
* **I asserted the brief's "six rows" instead of counting, for about ten minutes.** I wrote the
|
||||
register block with "six" in it before running the count that produced eleven, and only caught it
|
||||
because I decided to enumerate the rows rather than describe them — which the brief had insisted
|
||||
on for a different reason. **The instruction that saved it was not the one aimed at this.**
|
||||
* **My first `[BETA-DEFERRED]` claim over-promised.** I wrote in the register that the grep "returns
|
||||
the set as a group and nothing else", then found the grep returns eight lines because §8's own
|
||||
header defines the marker. Corrected in place with the real count. It is a small instance of
|
||||
exactly the class the register spent 2026-09-01 documenting — a claim about an instrument that had
|
||||
not been run.
|
||||
* **I preserved `REPORT.md` twice in one day and should have noticed the pattern the first time.**
|
||||
`REPORT.md` held the only copy of the R-331 report this morning and the only copy of the drill
|
||||
report this afternoon; both are now siblings (`REPORT-r331-backup-card.md`,
|
||||
`REPORT-drill-r95-recovery-2026-09-01.md`). The convention says durable content may not live only
|
||||
in the overwritten file, and it has now been violated twice in a day by two different sessions —
|
||||
which suggests the convention needs a gate, not more diligence. **Not filed:** I am not filing a
|
||||
row for it in a session that already declined a register sweep; it is named here for whoever picks
|
||||
up R-437.
|
||||
|
||||
@@ -674,6 +674,7 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server`
|
||||
| **R-434** | **The snapshot-drop alarm promises a recovery that cannot be performed.** `hub/internal/monitor/offsite.go` `emitSnapshotDrop` ships this text, live in hub **v0.111.0**: *"The daily Storage Box snapshots are read-only and still hold the older copy, **so this is recoverable file-by-file**; it is NOT confirmed data loss."* The first clause is true. The second is not reachable: not by the product (R-433), and not by the operator without a browser and a main-account credential that does not exist here. Its own comment states the intent — *"THE MESSAGE MUST NOT SAY THE DATA IS LOST, because after the 2026-09-01 measurement that is usually false"* — and the measurement it rests on was superseded the same day. **This is this project's own corollary landing on the alarm shipped that morning:** when a verdict changes which fact it counts from, the alarm text has to change with it, or the operator acts on a promise nobody can keep. **Fix is text-only and must not be made before R-433 settles what IS true** — an alarm rewritten twice in a week is worse than one rewritten once. **✅ CLOSED 2026-09-01, hub v0.111.1 — AND THE BLOCK ABOVE WAS WRONG, WHICH IS THE POINT WORTH KEEPING.** This row said the fix had to wait for R-433 to establish what IS true. **It did not, because the fix is a DELETION and not a REPLACEMENT.** The promise was withdrawn rather than swapped for a new one: *"The daily Storage Box snapshots are read-only and still hold the older copy. The route back out of them is not yet established, so treat this as neither confirmed data loss nor confirmed recovery. Get in touch before restoring anything, and check whether a deletion ran on the box."* **That sentence is true under EVERY possible answer to the provider questions, so it never needs a second rewrite** — which is the whole reason it was not blocked. A replacement would have been. **Three tests in `hub/internal/monitor/offsite_r434_test.go`**, all driving the production path so they assert the sentence an operator RECEIVES: the withdrawal is present, the promise is absent in three shapes, `confirmed data loss` may appear only inside its negation, and the stored row must not drift from the delivered mail. **RED-PROOF: restoring the v0.111.0 sentence failed all three**, on every fragment, with the offending sentence printed. **One existing test was edited and it had caught this fix correctly** — `TestR431_FiresOnAMassDeletion` asserted `"NOT confirmed data loss"`; the fragment was REMOVED rather than updated so the wording keeps ONE home. | **CLOSED 2026-09-01 — hub v0.111.1; the "blocked on R-433" verdict above was mine and it was wrong** |
|
||||
| **R-435** | **The snapshot-drop detector is blind to a single-app deletion — the exact shape the box can produce.** `snapshotDropFraction = 0.5` and `snapshotDropFloor = 5` (`hub/internal/monitor/offsite.go`) require a fall of MORE than half the previous count. demo-hp's baseline is **69** across 9 apps, so ~35 snapshots must go before it speaks; **one app's tag is ~9 and is invisible.** `offbox.go:1388` runs `forget --prune` **grouped by host,tags** — a per-tag wipe is precisely the shape a faulty retention or a targeted deletion produces. **This is deliberate, not accidental:** the constant's own comment argues the insensitivity, and *"a detector that cries wolf is switched off within a fortnight"* is a lesson this project paid for. **So this row is NOT a demand to lower the threshold.** It is a demand that the blind spot be written where the operator reads it, because "an unexplained fall is noticed within a day" (R-431, and `STATUS.md`) is true only of falls above half. **Discovered by arithmetic while planning the drill's Phase 2/3 pairing, which could not have worked:** Phase 2 deletes one app and Phase 3 expects the alarm to fire. **LIMITATION NOW WRITTEN INTO THE ALARM'S OWN DOCUMENTATION, hub v0.111.1** — the comment above `snapshotDropFraction` in `hub/internal/monitor/offsite.go` now states what the detector does NOT see, with the demo-hp arithmetic and the `forget --prune` grouping that makes the blind spot sit on the most likely single-app failure. **It also says explicitly that the numbers must NOT be lowered to "fix" this** and that per-app detection needs a SECOND signal keyed on the per-tag count. **The row stays OPEN because documenting a blind spot is not covering it** — and because `STATUS.md` and R-431 both still say "noticed within a day", which is true only of falls above half. | **OPEN — documented in code v0.111.1; the coverage gap itself is unclosed** |
|
||||
| **R-436** | **LEAD, NOT A DEFECT — append-only may be reachable without a new machine, which would make R-95's real prevention far cheaper than the spike concluded.** Hetzner's port-23 restricted shell advertises, in its own `help`, these server-side backends: `borg`, `rsync`, `scp`, `sftp`, **`rclone serve restic --stdio`**. And restic 0.14.0 **recognises the `rclone:` backend** — MEASURED 2026-09-01 with a control: `banana:` → `Fatal: parsing repository location failed: invalid backend`, while `rclone:` → `exec: "rclone": executable file not found in $PATH` (i.e. the backend parsed and it tried to run the helper). rclone is **not** in the controller image today. `rclone serve restic` carries an **`--append-only`** flag. **Why this matters:** the R-95 spike's option 3 was priced at a new always-on service in the recovery path plus either a mount in the hot path or migrating every customer's history — and it was deferred on that price. This route needs neither: the server side already runs at the provider. **THE CAVEAT, STATED FIRST because it may kill the idea:** the **client** supplies the server command line, so a compromised guest could simply omit `--append-only` unless the provider pins it. **NOT ESTABLISHED:** whether Hetzner pins the flag or accepts client-supplied arguments. **That is a vendor question and it is cheap** — it should be asked before any code is written, because if the answer is "client-supplied" this lead is worth nothing. | **OPEN — ask the vendor before building anything** |
|
||||
| **R-437** | **The register compression sweep is OWED, and it was deliberately NOT run inside the 2026-09-01 beta-line session — this row is the record of that choice, not a note.** **The ask:** compress what has closed in `OPEN-ITEMS.md`. **The measurement, taken before deciding:** 181 rows, 316 KB of row text, of which **12 rows / ~25 KB (about 7 %) carry a CLOSED/DECIDED/ANSWERED leading verdict.** So the sweep buys little and touches everything. **Why it was refused as a side-task, and the citation matters:** a compression sweep is the exact operation that has already gone wrong here. The 2026-08-22 sweep (`ef6ac6f`, R-376..R-378) matched a status word ANYWHERE in the row, moved rows that were not closed, and **R-378 caught six in the same session and missed a seventh** — R-87 sat in the wrong register for nine days while the ranking paragraph pointed at nothing (R-405). **That is a session-scale hazard, and running it as the tail end of a session about something else is how it happened the first time.** **WHAT IS OWED, scoped so it can be picked up cold:** (1) classify by the **LEADING VERDICT** of the state cell only — the rule `closed_register_gate.py` already implements and red-proofs, never a whole-row match; (2) move, never rewrite — a compressed row that loses its evidence is worse than a long one; (3) run `closed_register_gate.py` before and after and quote both; (4) re-read the ranking paragraph afterwards, because that is the surface that silently went stale last time. **Not urgent:** the file is 688 lines and every gate reads it in well under a second. | **OPEN — owed; needs its own session, not a tail end** |
|
||||
|
||||
<!-- DUE-CHECKS-BEGIN — machine-readable. Parsed by scripts/due_checks_gate.py.
|
||||
One row per dated check. The R-number must have a row above. Dates are UTC.
|
||||
|
||||
Reference in New Issue
Block a user