Files
felhom.eu/REPORT.md
T
admin f380c6d43c
gates / gates (push) Successful in 18s
REPORT: correct the register line count (620 -> 688, one measure) and record the CI run ids
Both earlier commit messages carried a wrong line count - 621->688 and 688->700 - because I
mixed wc -l with a Python line split and then carried the error forward. Corrected in the
report rather than by rewriting history, and named there.

CI verified by run id against head_sha: 500, 501, 502 all success.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LB8FmJaGd2cyjvy6dbEjpM
2026-09-01 18:43:20 +02:00

19 KiB

REPORT — the line under the backup arc, and one alarm that was telling the operator something untrue

hub v0.111.0 → v0.111.1 · controller v0.232.0 UNTOUCHED · 2026-09-01

No controller release. No agent release. No golden owed, no floor change. The only behaviour change in this session is one sentence in one alarm. Everything else is register and documentation.


1. The alarm's new text, and proof the old promise is gone

Quoted from the RUNNING binary — kubectl cp out of pod hub-857678f9b4-v95s6, byte-grepped:

Customer %s: off-site backup count fell from %d to %d snapshot(s) in one report — more than
retention can explain. The daily Storage Box snapshots are read-only and still hold the older
copy. The route back out of them is not yet established, so treat this as neither confirmed
data loss nor confirmed recovery. Get in touch before restoring anything, and check whether a
deletion ran on the box.

Every shape of the withdrawn promise is absent from the deployed bytes:

fragment in the deployed binary
recoverable file-by-file absent
recoverable file by file absent
so this is recoverable absent
it is NOT confirmed data loss absent
(negative control) zzz-no-such-string-r434 absent — so the search can report absence
(positive control) offsite_snapshots_dropped present, 2 occurrences

All seven fragments of the new sentence were checked individually and are present.

THE FIX IS A DELETION, NOT A REPLACEMENT — and that is what unblocked it. R-434's own row said the fix was "blocked on R-433", i.e. on first establishing what IS true. That verdict was mine and it was wrong. A sentence that asserts neither loss nor recovery is true under every possible answer to the provider questions, so it never needs a second rewrite. A replacement would have been blocked; a withdrawal is not. The reasoning is recorded in R-434's closing cell and in the function's doc comment, because the distinction is the transferable part.

It does not swing the other way either. TestR434_AlarmStillDoesNotClaimDataLoss pins that: "your backups are gone" is still usually false — the snapshots exist and hold the older copy; what is unproven is our route to them. Clause (a) of the 2026-09-01 measurement — the box cannot write into the snapshot area — stands and is re-confirmed.

Tests, and the red-proof

Three tests in hub/internal/monitor/offsite_r434_test.go, all driving the production path (saveOffsiteReport → oc.Check() → the notify callback), so they assert the sentence an operator receives rather than the function that formats it. ASCII-only fragments, with a positive control (a phrase in every version of the alarm) and a negative control (a phrase that cannot exist).

RED-PROOF (run before the fix was restored): the v0.111.0 sentence was put back and all three FAILED — on recoverable file-by-file, on so this is recoverable, on all three required fragments, on the un-negated confirmed data loss, and on the stored row — each with the offending sentence printed in the failure message. Fix restored, git diff clean, full hub suite green.

One existing test was edited, and it caught this fix correctly. TestR431_FiresOnAMassDeletion asserted the fragment "NOT confirmed data loss" and went red on the new wording. The fragment was REMOVED, not updated — the wording now has one home (offsite_r434_test.go), because duplicating it creates the second source that makes the next correction land in one file and not the other. Its signal fragments (69, 4, read-only) are unchanged.

R-435 written into the alarm's own documentation, no threshold changed

The comment above snapshotDropFraction now states what the detector does not see: a mass deletion, yes; one app being wiped, no. demo-hp's baseline is 69 across 9 apps, so ~35 must go before it speaks and one app's tag is ~9; offbox.go:1388 runs forget --prune grouped by host,tags, so the blind spot sits on the most likely single-app failure. The comment says explicitly that the numbers must not be lowered to "fix" this, and that per-app detection needs a second signal keyed on the per-tag count.

2. The stopping line, as it now reads in all three places

Register — documentation/backlog/OPEN-ITEMS.md, a new section in the voice this file uses for a settled decision:

DECIDED — the backup and restore arc is CLOSED FOR BETA (2026-09-01) … CLOSED FOR BETA at controller v0.232.0 / hub v0.111.1. … What is finished, and proven live: everything a customer does for themselves … Rows 1, 2, 3, 3b, 3c, 6, 7 and 14 of 07 §8 — every one PROVEN. … THIS REOPENS IF: a customer-facing recovery path is found broken; or Hetzner's answers change what the snapshots are worth …; or a real customer's data is at stake in one of the deferred rows.

Architecture — documentation/architecture/07-backup-architecture.md, at the head of §8, so a reader of the matrix meets it before the blanks:

THE ARC IS CLOSED FOR BETA — read this before the blanks below (2026-09-01) … The blanks in the RTO column below are now blank ON PURPOSE, and that is the whole difference. … NO STATUS MOVED ON THE DAY THIS WAS WRITTEN, because nothing was proven that day. A stopping line that promotes a row is a stopping line that lies.

Operator page — STATUS.md, under Decided — and what would reopen each:

THE BACKUP AND RESTORE WORK IS FINISHED FOR BETA. DECIDED 2026-09-01. … What is parked until after beta: everything only I do, with you — rebuilding a machine as itself, losing a whole box, recovering from ransomware, restoring the hub, and losing Hetzner. Six of these have never been timed, and the hub has never been restored. They are written down, they are real, and none of them stops a beta customer.

3. The deferred set — row numbers, not a description

07 §8 rows 4, 8, 9, 10, 11 (and 11b), 12, each tagged [BETA-DEFERRED] in its status cell.

row failure status today
4 primary drive dies — the drive-loss journey PARTIAL
8 host dies, drives intact — a host rebuilt as itself IMPLEMENTED, never executed
9 whole box lost (fire/theft) IMPLEMENTED / UNPROVEN
10 ransomware / malicious deletion PARTIAL
11 (+11b) hub lost — a hub restore UNPROVEN, never performed
12 off-site provider lost (Hetzner) [FACT] only

grep -n '\[BETA-DEFERRED\]' documentation/architecture/07-backup-architecture.md returns eight lines — the seven tagged rows plus the one line in §8's header that defines the marker. That is stated in the register rather than left for the reader to trip over.

A NUMBER IN THE BRIEF WAS WRONG AND IS CORRECTED IN PLACE. The brief said "six rows of §8 still have no measured time". Six rows are DEFERRED; ELEVEN carry a blank RTO — counted, not estimated: 4, 5, 8, 9, 10, 11, 11b, 12, 13, 14, 15. The other five are blank for reasons that are not deferred work, and collapsing them into one number is how a blank stops meaning anything:

  • row 5 — PROVEN by construction; a derived copy, so there is no recovery to time.
  • row 13 — NONE for host-loss by design; R exists in zero system copies.
  • row 14 — PROVEN; the break-glass route works and has simply never been stopwatched.
  • row 11b — a consequences note attached to row 11, not a recovery row.
  • row 15 — an open DEFECT (R-104, the stale-lock path). It is NOT inside the stopping line and must not be read as parked by it. This one matters most: parking a live defect by accident is the failure mode a stopping line invites.

No status was moved. Nothing was proven today.

4. The two questions

documentation/runbooks/provider-questions-2026-09-01.md — both drafted ready to paste, linked from R-95 and R-433, not sent, and no provider API was called (§11-D stands).

  • Q1 — can the MAIN account retrieve individual files from a snapshot, without a whole-box restore? Why: if it cannot, the only route is a rollback that hits every customer on the box and destroys newer snapshots — the snapshots would then protect almost nobody. Decides how urgent R-95 is.
  • Q2 — on rclone serve restic --stdio, is --append-only enforced server-side or chosen by the client? Why: if enforced, the box cannot delete its own backups, with no new machine and no data move. Could make R-95 disappear.

Each carries how to read either answer, including the branch where the lead is worthless and should be recorded as dead rather than left looking promising. A dated DUE-CHECKS row (R-433, 2026-09-15) tracks the reply — the 2026-07-27 snapshot check that sat unconfirmed for 36 days is the scar that block exists for, and it was never entered.

5. The register

before after
file lines (wc -l) 620 688
total R- rows 181 182
open-state rows 170 169
closed / decided / answered rows 11 13

(The brief's "620 lines and 181 open" matches the file length and the row count exactly; the count of rows whose LEADING verdict is actually open was 170, which is the number that matters and is not the same thing.)

A correction to my own commit messages, made here rather than by rewriting history: the db38f4c message says 621 -> 688 and the 1a1b32b message says 688 -> 700. Both are wrong. I mixed two measures — wc -l and a Python line split that counts the trailing newline as a line — and then carried the error forward. The table above is wc -l throughout and is the number to trust: 620 → 688. It changes nothing about the work and it is exactly the class of sloppiness this project files rows about, so it is named rather than quietly fixed.

Rows closed: R-434 — with the deletion-not-replacement reasoning, and with the fact that its own "blocked on R-433" verdict was wrong recorded in the closing cell.

Rows kept open and marked:

  • R-95, R-433 — BLOCKED-ON-PROVIDER, both pointing at Part 3's file. R-95 records that it spent one day demoted on a clause that did not hold, and that on today's evidence it belongs back near the top. That is a proposal. I have not re-ranked it; the order is Viktor's.
  • R-430 — LATENT, with the trigger stated as a trigger: it becomes live the moment delete is withdrawn from the box, which is exactly what R-95's remedy does by either route. So it is a precondition on the R-95 build, not a follow-up — settle it in the same change or the crash-lock self-heal ships already broken.
  • R-435 — open. The limitation is now in the code's own documentation, but documenting a blind spot is not covering it, and STATUS.md and R-431 both still say "noticed within a day".

Row filed: R-437 — the register compression sweep, owed and scoped.

Compression was measured and deliberately not run, and the reason is one line as the standing rule requires: only 12 rows / ~25 KB of 316 KB (about 7 %) carry a closed leading verdict, so the sweep buys little and touches everything — and it is the exact operation that misfiled seven rows in August (R-378, and the seventh, R-87, sat wrong for nine days — R-405). Running it as the tail end of a session about something else is how that happened the first time. R-437 carries the scope so it can be picked up cold.

Also corrected in place: the ranking paragraph's item 1 said the snapshot mitigation was "armed (daily 00:00, keep 7), but it has taken zero snapshots so far". Both halves were wrong — seven snapshots exist (R-429) and "armed" was withdrawn the same day. The ORDER of that list is unchanged.

6. Hub deploy and its verification

step result
image built + pushed gitea.dooplex.hu/admin/felhom-hub:0.111.1 (25 M)
clean-tree gate before build git status --porcelain empty, HEAD == origin/main @ db38f4c
manifest bumped manifests/hub.yaml:128 → :0.111.1, committed 0f65f7a, pushed
ArgoCD hard refresh + deliberate sync sync=Synced health=Healthy rev=0f65f7a… — the revision equals HEAD
rollout deployment "hub" successfully rolled out
deployed image gitea.dooplex.hu/admin/felhom-hub:0.111.1
running binary felhom-hub 0.111.1 (built 2026-09-01T16:35:04Z)

No kubectl set image, no kubectl apply.

The live trigger was NOT run, and the reason is not a shortcut

The brief asked me to "trigger the alarm once through the hub's own path". I did not, and I am naming it rather than quietly substituting.

This alarm only fires on a real fall in a real customer's snapshot count. Firing it live therefore means one of two things: deleting a real customer's off-site backups (the destructive act this session explicitly is not), or POSTing a falsified report claiming demo-hp had lost its backups — which would write a fabricated data point permanently into that customer's report history, move its drop latch and baseline, and send Viktor a second alarm mail about a real box, five hours after the first one already confused him (STATUS.md item 5). Fabricating customer telemetry to satisfy a verification step is the wrong trade in a project whose whole doctrine is that a measurement must mean what it says.

What was done instead covers both halves of what the live trigger was for:

  1. Does the deployed artifact carry the sentence? Proved on the bytes of the running pod's binary — §1, with both controls.
  2. Does the production path emit it? Proved by three tests that drive saveOffsiteReport → Check() → notify and read the message the operator would receive, plus the stored-row test that pins mail and audit row together — and all three red-proofed.

What is therefore still unproven, stated plainly: that the dispatcher delivers this particular message to a mailbox. That leg was exercised for real on 2026-09-01 at 12:29 UTC by the previous session with the old text, so the routing is known good; only the new wording has not travelled it.

7. Explicitly

No controller release. No agent release. No image built for either. No golden baked, and none owed — the golden-currency gate reads newest released controller 0.232.0 / newest golden baked 0.232.0. The floor is unchanged at 0.232.0. felhom-controller and felhom-agent working trees were not touched.

python3 scripts/unproven.py --summary: NOT WALKED 35 of 55 — unchanged. No claim moved, which is correct: nothing was proven today. All 14 felhom.eu gates green before every push.

CI, checked by run id against head_sha rather than assumed (the R-417 recipe — the actions/jobs endpoint, paged to the end):

run id commit conclusion
500 db38f4c — hub v0.111.1 + the stopping line success
501 0f65f7a — manifest bump to 0.111.1 success
502 1a1b32b — REPORT + R-437 success

8. Observations

  1. strings stops at the em dash, so the alarm sentence appeared truncated in the deployed binary and briefly looked like a bad build. strings scans ASCII by default and the message's — terminates the run; the fix is LC_ALL=C grep -aoP on the bytes. NOT-A-FINDING: this is the project's already-recorded accented-grep trap appearing on a new surface, so it needs no new row — it is a hazard of my verification method, not a defect in any product code. Recorded here so the next person grepping a binary for Hungarian or em-dashed copy does not read a truncation as a bad build, which is exactly how it read for a minute.

  2. The dispatcher's cooldowns are in-memory and are lost on every hub restart, so any deploy re-arms every alarm's 6-hour cooldown. NOT-A-FINDING: it is stated in hub/internal/notify/dispatcher.go:18 as a deliberate, accepted trade. Noted because it is why a live trigger today would definitely have mailed, rather than being absorbed by yesterday's cooldown — it changed the decision in §6.

  3. The brief's baseline for felhom-agent was two commits stale — it names 058b945 (2026-08-23) while main is 4586f0f (2026-09-01). NOT-A-FINDING: the agent was untouched either way, and 058b945 is a real commit, so nothing was ambiguous. Flagged only so the number is not copied forward into the next brief.

  4. OPEN-ITEMS.md has no section called "Decided — and what would reopen each"; STATUS.md does. The brief sent the register text to that section by name. NOT-A-FINDING: resolved in the session by writing a new register section in that section's voice — the decision, then the condition that reopens it — which is plainly what the instruction meant, so there is nothing left to file. Recorded only so the next session does not hunt OPEN-ITEMS.md for a heading that has never existed there.

9. My own mistakes

  • I wrote "blocked on R-433" into R-434 yesterday, and it was wrong. It cost nothing because the block lasted one day, but the reasoning error is the interesting part: I treated "we do not know what is true" as a reason not to touch a sentence that was known to be false. Removing a false claim never needs the true one. The row and the code comment now say so, and it is the only part of this session I would call a lesson rather than a task.
  • I asserted the brief's "six rows" instead of counting, for about ten minutes. I wrote the register block with "six" in it before running the count that produced eleven, and only caught it because I decided to enumerate the rows rather than describe them — which the brief had insisted on for a different reason. The instruction that saved it was not the one aimed at this.
  • My first [BETA-DEFERRED] claim over-promised. I wrote in the register that the grep "returns the set as a group and nothing else", then found the grep returns eight lines because §8's own header defines the marker. Corrected in place with the real count. It is a small instance of exactly the class the register spent 2026-09-01 documenting — a claim about an instrument that had not been run.
  • I preserved REPORT.md twice in one day and should have noticed the pattern the first time. REPORT.md held the only copy of the R-331 report this morning and the only copy of the drill report this afternoon; both are now siblings (REPORT-r331-backup-card.md, REPORT-drill-r95-recovery-2026-09-01.md). The convention says durable content may not live only in the overwritten file, and it has now been violated twice in a day by two different sessions — which suggests the convention needs a gate, not more diligence. Not filed: I am not filing a row for it in a session that already declined a register sweep; it is named here for whoever picks up R-437.