Files
felhom.eu/REPORT.md
T
admin f380c6d43c
gates / gates (push) Successful in 18s
REPORT: correct the register line count (620 -> 688, one measure) and record the CI run ids
Both earlier commit messages carried a wrong line count - 621->688 and 688->700 - because I
mixed wc -l with a Python line split and then carried the error forward. Corrected in the
report rather than by rewriting history, and named there.

CI verified by run id against head_sha: 500, 501, 502 all success.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LB8FmJaGd2cyjvy6dbEjpM
2026-09-01 18:43:20 +02:00

312 lines
19 KiB
Markdown

# REPORT — the line under the backup arc, and one alarm that was telling the operator something untrue
**hub v0.111.0 → v0.111.1 · controller v0.232.0 UNTOUCHED · 2026-09-01**
**No controller release. No agent release. No golden owed, no floor change.** The only behaviour
change in this session is one sentence in one alarm. Everything else is register and documentation.
---
## 1. The alarm's new text, and proof the old promise is gone
**Quoted from the RUNNING binary** — `kubectl cp` out of pod `hub-857678f9b4-v95s6`, byte-grepped:
```
Customer %s: off-site backup count fell from %d to %d snapshot(s) in one report — more than
retention can explain. The daily Storage Box snapshots are read-only and still hold the older
copy. The route back out of them is not yet established, so treat this as neither confirmed
data loss nor confirmed recovery. Get in touch before restoring anything, and check whether a
deletion ran on the box.
```
**Every shape of the withdrawn promise is absent from the deployed bytes:**
| fragment | in the deployed binary |
|---|---|
| `recoverable file-by-file` | **absent** |
| `recoverable file by file` | **absent** |
| `so this is recoverable` | **absent** |
| `it is NOT confirmed data loss` | **absent** |
| *(negative control)* `zzz-no-such-string-r434` | absent — so the search can report absence |
| *(positive control)* `offsite_snapshots_dropped` | present, 2 occurrences |
All seven fragments of the new sentence were checked individually and are present.
**THE FIX IS A DELETION, NOT A REPLACEMENT — and that is what unblocked it.** R-434's own row said
the fix was *"blocked on R-433"*, i.e. on first establishing what IS true. **That verdict was mine
and it was wrong.** A sentence that asserts **neither** loss **nor** recovery is true under every
possible answer to the provider questions, so it never needs a second rewrite. A *replacement* would
have been blocked; a *withdrawal* is not. The reasoning is recorded in R-434's closing cell and in
the function's doc comment, because the distinction is the transferable part.
**It does not swing the other way either.** `TestR434_AlarmStillDoesNotClaimDataLoss` pins that:
"your backups are gone" is still usually false — the snapshots exist and hold the older copy; what is
unproven is our route to them. Clause (a) of the 2026-09-01 measurement — the box **cannot write**
into the snapshot area — stands and is re-confirmed.
### Tests, and the red-proof
Three tests in `hub/internal/monitor/offsite_r434_test.go`, all driving the production path
(`saveOffsiteReport` → `oc.Check()` → the notify callback), so they assert the sentence an operator
**receives** rather than the function that formats it. ASCII-only fragments, with a positive control
(a phrase in every version of the alarm) and a negative control (a phrase that cannot exist).
**RED-PROOF (run before the fix was restored):** the v0.111.0 sentence was put back and **all three
FAILED** — on `recoverable file-by-file`, on `so this is recoverable`, on all three required
fragments, on the un-negated `confirmed data loss`, and on the stored row — each with the offending
sentence printed in the failure message. Fix restored, `git diff` clean, full hub suite green.
**One existing test was edited, and it caught this fix correctly.**
`TestR431_FiresOnAMassDeletion` asserted the fragment `"NOT confirmed data loss"` and went red on the
new wording. **The fragment was REMOVED, not updated** — the wording now has one home
(`offsite_r434_test.go`), because duplicating it creates the second source that makes the next
correction land in one file and not the other. Its signal fragments (`69`, `4`, `read-only`) are
unchanged.
### R-435 written into the alarm's own documentation, no threshold changed
The comment above `snapshotDropFraction` now states what the detector does **not** see: **a mass
deletion, yes; one app being wiped, no.** demo-hp's baseline is 69 across 9 apps, so ~35 must go
before it speaks and one app's tag is ~9; `offbox.go:1388` runs `forget --prune` grouped by
`host,tags`, so the blind spot sits on the most likely single-app failure. The comment says
explicitly that the numbers must **not** be lowered to "fix" this, and that per-app detection needs a
**second** signal keyed on the per-tag count.
## 2. The stopping line, as it now reads in all three places
**Register — `documentation/backlog/OPEN-ITEMS.md`**, a new section in the voice this file uses for a
settled decision:
> **DECIDED — the backup and restore arc is CLOSED FOR BETA (2026-09-01)** … **CLOSED FOR BETA at
> controller v0.232.0 / hub v0.111.1.** … **What is finished, and proven live:** everything a
> customer does for themselves … **Rows 1, 2, 3, 3b, 3c, 6, 7 and 14 of `07` §8 — every one PROVEN.**
> … **THIS REOPENS IF:** a customer-facing recovery path is found broken; **or** Hetzner's answers
> change what the snapshots are worth …; **or** a real customer's data is at stake in one of the
> deferred rows.
**Architecture — `documentation/architecture/07-backup-architecture.md`, at the head of §8**, so a
reader of the matrix meets it before the blanks:
> **THE ARC IS CLOSED FOR BETA — read this before the blanks below (2026-09-01)** … **The blanks in
> the RTO column below are now blank ON PURPOSE, and that is the whole difference.** … **NO STATUS
> MOVED ON THE DAY THIS WAS WRITTEN, because nothing was proven that day. A stopping line that
> promotes a row is a stopping line that lies.**
**Operator page — `STATUS.md`**, under *Decided — and what would reopen each*:
> **THE BACKUP AND RESTORE WORK IS FINISHED FOR BETA. DECIDED 2026-09-01.** … **What is parked until
> after beta:** everything **only I do, with you** — rebuilding a machine as itself, losing a whole
> box, recovering from ransomware, restoring the hub, and losing Hetzner. **Six of these have never
> been timed, and the hub has never been restored.** They are written down, they are real, and **none
> of them stops a beta customer.**
## 3. The deferred set — row numbers, not a description
`07` §8 rows **4, 8, 9, 10, 11 (and 11b), 12**, each tagged **`[BETA-DEFERRED]`** in its status cell.
| row | failure | status today |
|---|---|---|
| **4** | primary drive dies — the drive-loss **journey** | `PARTIAL` |
| **8** | host dies, drives intact — a host rebuilt as itself | `IMPLEMENTED`, never executed |
| **9** | whole box lost (fire/theft) | `IMPLEMENTED / UNPROVEN` |
| **10** | ransomware / malicious deletion | `PARTIAL` |
| **11** (+**11b**) | hub lost — a hub restore | `UNPROVEN`, never performed |
| **12** | off-site provider lost (Hetzner) | `[FACT]` only |
`grep -n '\[BETA-DEFERRED\]' documentation/architecture/07-backup-architecture.md` returns **eight**
lines — the seven tagged rows plus the one line in §8's header that defines the marker. That is
stated in the register rather than left for the reader to trip over.
**A NUMBER IN THE BRIEF WAS WRONG AND IS CORRECTED IN PLACE.** The brief said *"six rows of §8 still
have no measured time"*. **Six rows are DEFERRED; ELEVEN carry a blank RTO** — counted, not
estimated: 4, 5, 8, 9, 10, 11, 11b, 12, 13, 14, 15. The other five are blank for reasons that are not
deferred work, and collapsing them into one number is how a blank stops meaning anything:
* **row 5** — `PROVEN` by construction; a derived copy, so there is no recovery to time.
* **row 13** — `NONE for host-loss` **by design**; R exists in zero system copies.
* **row 14** — `PROVEN`; the break-glass route works and has simply never been stopwatched.
* **row 11b** — a consequences note attached to row 11, not a recovery row.
* **row 15** — an **open DEFECT** (R-104, the stale-lock path). **It is NOT inside the stopping line**
and must not be read as parked by it. This one matters most: parking a live defect by accident is
the failure mode a stopping line invites.
**No status was moved.** Nothing was proven today.
## 4. The two questions
`documentation/runbooks/provider-questions-2026-09-01.md` — both drafted ready to paste, linked from
R-95 and R-433, **not sent**, and **no provider API was called** (§11-D stands).
* **Q1 — can the MAIN account retrieve individual files from a snapshot, without a whole-box restore?**
*Why:* if it cannot, the only route is a rollback that hits every customer on the box and destroys
newer snapshots — the snapshots would then protect almost nobody. **Decides how urgent R-95 is.**
* **Q2 — on `rclone serve restic --stdio`, is `--append-only` enforced server-side or chosen by the
client?** *Why:* if enforced, the box cannot delete its own backups, with no new machine and no
data move. **Could make R-95 disappear.**
Each carries how to read either answer, including the branch where the lead is worthless and should
be recorded as dead rather than left looking promising. **A dated DUE-CHECKS row (R-433, 2026-09-15)
tracks the reply** — the 2026-07-27 snapshot check that sat unconfirmed for 36 days is the scar that
block exists for, and it was never entered.
## 5. The register
| | before | after |
|---|---|---|
| file lines (`wc -l`) | **620** | **688** |
| total `R-` rows | **181** | **182** |
| open-state rows | **170** | **169** |
| closed / decided / answered rows | **11** | **13** |
*(The brief's "620 lines and 181 open" matches the file length and the row count exactly; the count
of rows whose LEADING verdict is actually open was 170, which is the number that matters and is not
the same thing.)*
**A correction to my own commit messages, made here rather than by rewriting history:** the
`db38f4c` message says `621 -> 688` and the `1a1b32b` message says `688 -> 700`. Both are wrong. I
mixed two measures — `wc -l` and a Python line split that counts the trailing newline as a line — and
then carried the error forward. **The table above is `wc -l` throughout and is the number to trust:
620 → 688.** It changes nothing about the work and it is exactly the class of sloppiness this project
files rows about, so it is named rather than quietly fixed.
**Rows closed:** **R-434** — with the deletion-not-replacement reasoning, and with the fact that its
own "blocked on R-433" verdict was wrong recorded in the closing cell.
**Rows kept open and marked:**
* **R-95**, **R-433** — **`BLOCKED-ON-PROVIDER`**, both pointing at Part 3's file. R-95 records that
it spent **one day demoted on a clause that did not hold**, and that on today's evidence it belongs
back near the top. **That is a proposal. I have not re-ranked it; the order is Viktor's.**
* **R-430** — **`LATENT`**, with the trigger stated as a trigger: it becomes live **the moment delete
is withdrawn from the box**, which is exactly what R-95's remedy does by either route. So it is a
**precondition on the R-95 build, not a follow-up** — settle it in the same change or the
crash-lock self-heal ships already broken.
* **R-435** — open. The limitation is now in the code's own documentation, but **documenting a blind
spot is not covering it**, and `STATUS.md` and R-431 both still say "noticed within a day".
**Row filed:** **R-437** — the register compression sweep, owed and scoped.
**Compression was measured and deliberately not run, and the reason is one line as the standing rule
requires:** only **12 rows / ~25 KB of 316 KB (about 7 %)** carry a closed leading verdict, so the
sweep buys little and touches everything — and it is the exact operation that misfiled seven rows in
August (R-378, and the seventh, R-87, sat wrong for nine days — R-405). Running it as the tail end of
a session about something else is how that happened the first time. R-437 carries the scope so it can
be picked up cold.
**Also corrected in place:** the ranking paragraph's item 1 said the snapshot mitigation was *"armed
(daily 00:00, keep 7), but it has taken zero snapshots so far"*. **Both halves were wrong** — seven
snapshots exist (R-429) and "armed" was withdrawn the same day. **The ORDER of that list is
unchanged.**
## 6. Hub deploy and its verification
| step | result |
|---|---|
| image built + pushed | `gitea.dooplex.hu/admin/felhom-hub:0.111.1` (25 M) |
| clean-tree gate before build | `git status --porcelain` empty, `HEAD == origin/main` @ `db38f4c` |
| manifest bumped | `manifests/hub.yaml:128` → `:0.111.1`, committed `0f65f7a`, pushed |
| ArgoCD hard refresh + **deliberate** sync | `sync=Synced` `health=Healthy` `rev=0f65f7a…` — the revision equals HEAD |
| rollout | `deployment "hub" successfully rolled out` |
| deployed image | `gitea.dooplex.hu/admin/felhom-hub:0.111.1` |
| running binary | `felhom-hub 0.111.1 (built 2026-09-01T16:35:04Z)` |
No `kubectl set image`, no `kubectl apply`.
### The live trigger was NOT run, and the reason is not a shortcut
The brief asked me to *"trigger the alarm once through the hub's own path"*. **I did not, and I am
naming it rather than quietly substituting.**
**This alarm only fires on a real fall in a real customer's snapshot count.** Firing it live therefore
means one of two things: deleting a real customer's off-site backups (the destructive act this
session explicitly is not), or **POSTing a falsified report claiming `demo-hp` had lost its
backups** — which would write a fabricated data point permanently into that customer's report
history, move its drop latch and baseline, and send Viktor a **second** alarm mail about a real box,
five hours after the first one already confused him (`STATUS.md` item 5). Fabricating customer
telemetry to satisfy a verification step is the wrong trade in a project whose whole doctrine is that
a measurement must mean what it says.
**What was done instead covers both halves of what the live trigger was for:**
1. **Does the deployed artifact carry the sentence?** Proved on the bytes of the running pod's
binary — §1, with both controls.
2. **Does the production path emit it?** Proved by three tests that drive
`saveOffsiteReport → Check() → notify` and read the message the operator would receive, plus the
stored-row test that pins mail and audit row together — and all three red-proofed.
**What is therefore still unproven, stated plainly:** that the *dispatcher* delivers this particular
message to a mailbox. That leg was exercised for real on 2026-09-01 at 12:29 UTC by the previous
session with the old text, so the routing is known good; only the new wording has not travelled it.
## 7. Explicitly
**No controller release. No agent release. No image built for either. No golden baked, and none
owed** — the golden-currency gate reads `newest released controller 0.232.0 / newest golden baked
0.232.0`. **The floor is unchanged at 0.232.0.** `felhom-controller` and `felhom-agent` working trees
were not touched.
`python3 scripts/unproven.py --summary`: **NOT WALKED 35 of 55 — unchanged.** No claim moved, which
is correct: nothing was proven today. All **14** `felhom.eu` gates green before every push.
**CI, checked by run id against `head_sha` rather than assumed** (the R-417 recipe — the
`actions/jobs` endpoint, paged to the end):
| run id | commit | conclusion |
|---|---|---|
| **500** | `db38f4c` — hub v0.111.1 + the stopping line | **success** |
| **501** | `0f65f7a` — manifest bump to 0.111.1 | **success** |
| **502** | `1a1b32b` — REPORT + R-437 | **success** |
## 8. Observations
1. **`strings` stops at the em dash, so the alarm sentence appeared truncated in the deployed
binary and briefly looked like a bad build.** `strings` scans ASCII by default and the message's
`—` terminates the run; the fix is `LC_ALL=C grep -aoP` on the bytes.
**NOT-A-FINDING:** this is the project's already-recorded accented-grep trap appearing on a new
surface, so it needs no new row — it is a hazard of my verification method, not a defect in any
product code. Recorded here so the next person grepping a binary for Hungarian or em-dashed copy
does not read a truncation as a bad build, which is exactly how it read for a minute.
2. **The dispatcher's cooldowns are in-memory and are lost on every hub restart**, so any deploy
re-arms every alarm's 6-hour cooldown. **NOT-A-FINDING:** it is stated in
`hub/internal/notify/dispatcher.go:18` as a deliberate, accepted trade. Noted because it is why a
live trigger today would definitely have mailed, rather than being absorbed by yesterday's
cooldown — it changed the decision in §6.
3. **The brief's baseline for `felhom-agent` was two commits stale** — it names `058b945`
(2026-08-23) while `main` is `4586f0f` (2026-09-01). **NOT-A-FINDING:** the agent was untouched
either way, and `058b945` is a real commit, so nothing was ambiguous. Flagged only so the number
is not copied forward into the next brief.
4. **`OPEN-ITEMS.md` has no section called *"Decided — and what would reopen each"*; `STATUS.md`
does.** The brief sent the register text to that section by name.
**NOT-A-FINDING:** resolved in the session by writing a new register section in that section's
voice — the decision, then the condition that reopens it — which is plainly what the instruction
meant, so there is nothing left to file. Recorded only so the next session does not hunt
`OPEN-ITEMS.md` for a heading that has never existed there.
## 9. My own mistakes
* **I wrote "blocked on R-433" into R-434 yesterday, and it was wrong.** It cost nothing because the
block lasted one day, but the reasoning error is the interesting part: I treated *"we do not know
what is true"* as a reason not to touch a sentence that was **known to be false**. Removing a false
claim never needs the true one. The row and the code comment now say so, and it is the only part of
this session I would call a lesson rather than a task.
* **I asserted the brief's "six rows" instead of counting, for about ten minutes.** I wrote the
register block with "six" in it before running the count that produced eleven, and only caught it
because I decided to enumerate the rows rather than describe them — which the brief had insisted
on for a different reason. **The instruction that saved it was not the one aimed at this.**
* **My first `[BETA-DEFERRED]` claim over-promised.** I wrote in the register that the grep "returns
the set as a group and nothing else", then found the grep returns eight lines because §8's own
header defines the marker. Corrected in place with the real count. It is a small instance of
exactly the class the register spent 2026-09-01 documenting — a claim about an instrument that had
not been run.
* **I preserved `REPORT.md` twice in one day and should have noticed the pattern the first time.**
`REPORT.md` held the only copy of the R-331 report this morning and the only copy of the drill
report this afternoon; both are now siblings (`REPORT-r331-backup-card.md`,
`REPORT-drill-r95-recovery-2026-09-01.md`). The convention says durable content may not live only
in the overwritten file, and it has now been violated twice in a day by two different sessions —
which suggests the convention needs a gate, not more diligence. **Not filed:** I am not filing a
row for it in a session that already declined a register sweep; it is named here for whoever picks
up R-437.