Files
felhom-controller/REPORT.md
T
admin 300d7e87d7
gates / gates (push) Failing after 12s
REPORT: R-359/R-397 shipped and validated; the measurement changed what it is worth
The headline is not the feature, it is what measuring it revealed: THE STRUCTURE
CHECK THAT SHIPS ON DOES NOT CATCH SILENT CORRUPTION. A pack corrupted without a
size change returned `no errors were found`, exit 0. Only --read-data caught it.
So R-399 is not merely a bandwidth question -- at the shipped default a class of
damage is not checked at all.

The three numbers R-399 needed are MEASURED, not estimated: store 134.3 MB / 67
snapshots; structure check 35.0 s; and the full curve 10% 35.9 s, 50% 37.3 s,
100% 39.2 s. At this size re-reading everything costs four seconds more than
reading none. Stated limit: they do not extrapolate.

Records four things that went wrong and were caught rather than shipped:
  - R-398 was MY OWN mistaken row. The seam already existed, Part 0 was not
    built, and building it would have HIDDEN the unlock --remove-all escalation
    from the assertions that must see it. Corrected, not closed.
  - the damage classifier matched restic's ORDINARY progress output; the
    negative control caught it (v0.227.1).
  - an exit code I misread through a pipe, corrected by re-measuring.
  - wire-contract convicted a field the hub cannot decode; allowlisted WITH A
    REASON rather than skipped, because building the hub display is a decision
    R-331 already ruled belongs to the operator.

And the sweep the task asked for: EIGHT debug buttons post to endpoints that do
not exist, not one. 24 referenced, 17 dispatched. Filed as R-400.

Not live-validated and each says why: the weekly firing (a week away), read-data
on a large store, and the join between "restic catches it" and "my code
classifies it" -- both proven, the join is not, and the seam is named.
2026-08-30 21:29:15 +02:00

228 lines
14 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# REPORT — R-359 / R-397: the off-site store gets checked
**Controller v0.227.0 → v0.227.1 · 2026-08-30 · implemented on DooPlex, validated live on `demo-hp`**
---
## ⚠ READ THIS FIRST — the measurement changed what the feature is worth
Part 5 corrupted one pack of a throwaway repository **without changing its size** (64 zero bytes at
offset 1024 — the subtlest form of rot). Both depths were run against it:
| depth | exit | verdict |
|---|---|---|
| `restic check` — **the depth that ships ON** | **0** | **`no errors were found`** |
| every `--read-data*` form — **ships OFF** | 1 | `Pack ID does not match, want 288afd3e…, got 4b6847bb…` |
**The structure check reported a corrupted store as healthy.** It verifies the index, the pack
inventory and the snapshot graph — it catches missing packs, broken indexes and unreadable snapshots,
which are real — but it does **not** re-hash pack contents. **R-399 is therefore not only a bandwidth
question: at the shipped default a class of damage is not checked at all**, and it is the class that
silently eats a customer's photos. The task could not have known this; nothing had measured it.
## 1. Confirmed baselines
| repo | spec baseline | actual at start | version |
|---|---|---|---|
| `felhom-controller` | `c476ae51` | **`e64c84ae`** (drift: one docs-only commit — my own golden-bake report) | v0.226.1 → **v0.227.1** |
| `felhom.eu` | `4f875174` | `4f875174` — **exact match** | docs only |
| `felhom-agent` | not touched | v0.130.0 | unchanged |
Both trees clean and equal to `origin/main`. **Every symbol in §5 was re-verified present** (20/20),
and all three claimed absences confirmed: 0 occurrences of a restic `check` verb, 0 callers of
`NotifyIntegrity*` in the controller, and the debug button present with no dispatch case.
`MinAgent` stays **0.129.0**.
## 2. Files created / modified
**`felhom-controller`:** `controller/internal/backup/offbox_integrity.go` (**new**),
`offbox.go` (report fields), `offbox_restore.go`+`r358_scratch_marker_test.go` (AST→execution test),
`internal/settings/settings.go`, `internal/config/config.go`,
`cmd/controller/main.go` (job + the one caller + the debug callback),
`internal/web/handler_debug.go`, `internal/web/handlers.go`,
tests: `r359_integrity_test.go`, `r359_lock_safety_test.go`, `r359_dueness_test.go` (**new**),
`cmd/controller/r359_wiring_test.go`, `internal/web/r397_debug_integrity_test.go` (**new**);
docs: `CHANGELOG.md`, `CONTEXT.md`, `REUSE.md`, `controller/README.md`, `REPORT.md`.
**`felhom.eu`:** `STATUS.md`, `backlog/OPEN-ITEMS.md`, `backlog/CLOSED-ITEMS.md`,
`architecture/07-backup-architecture.md`, `architecture/08-alarm-ladder.md`,
`architecture/00-capability-map.md`, `scripts/wire_contract_gate.py` (allowlist **with a reason**),
`documentation/tests/r359-integrity-2026-08-30/` (**new**, 10 files).
## 3. Commits pushed to `main`
| repo | commit | contents |
|---|---|---|
| `felhom-controller` | **`0d52a42c`** | v0.227.0 — the check, the job, the notifiers, the route, all tests |
| `felhom-controller` | **`45770f22`** | v0.227.1 — the classifier fix + CONTEXT/README/REUSE |
| `felhom.eu` | *(this report's commit)* | register, architecture, capability map, STATUS, evidence |
## 4. Tests and red-proofs
**All pass.** 22 new tests across 5 new files. Groups A (the check), B (the lock hazard), C (due-ness),
D (wiring), plus two fixtures built from restic's **real** bytes.
| red-proof | mutation | observed failure |
|---|---|---|
| **A2** | classifier treats any non-zero as OK | `a repository restic said contains errors was reported as OK` |
| **B1** | remove the `acquireRunning` guard | `RESTIC WAS INVOKED while another operation held the single-writer flag: [… cat config] [… check]` |
| **C2** | gate due-ness on `time.Sunday` | `a 21-day-old check was NOT due on Saturday` |
| **D2** | remove the debug dispatch case | `POST /api/debug/backup/integrity was not dispatched` |
**E1 — no existing test was modified.** `TestBackupReport_DeadFieldsStayZero` (R-331, D4) passes
unmodified, which is the check that I did not write to the retired report fields.
**Test count:** 28 packages green before and after; +22 tests. Green gate
`go build ./... && go vet ./... && go test ./...` → **rc 0**.
## 5. Deployed version
```
gitea.dooplex.hu/admin/felhom-controller:0.227.1 Up 13 seconds (healthy)
[INFO] [scheduler] Daily job offsite-integrity scheduled for 2026-08-31 06:00 CEST
```
**`demo-hp` only.** `demo-felhom` is on 0.226.1; the fleet floor and golden are 0.226.1.
**v0.227.1 is a patch, not a rebuilt 0.227.0** — 0.227.0 was already running on `demo-hp` when the
classifier defect was found, and re-pushing changed bytes under a live tag is the `:latest` hazard.
The live validation in §8 ran against **0.227.0**; 0.227.1 differs only by the damage signatures and
their test.
## 6. The schedule I chose, and what I read to choose it
Live on `demo-hp`: `db-dump 02:30` · `tier2-backup 03:30` · `fill-watch 03:30` · `metrics-prune 04:00`
· **`offbox-backup 04:15`** · `offsite-abandon-sweep 05:10` · whole-guest gate `[04:30, 08:30)`.
**Chose 06:00** — 1h45 clear of the off-site leg (which runs ~2m20s, measured) and 50 min clear of the
sweep. The backup window is customer-configurable, so **no fixed time is collision-proof for every
box**; a collision costs one skipped day, not a missed check, because due-ness makes tomorrow try
again. A slot that collided *every* night would be a real defect; this is not one.
## 7. NOT yet live-validated
1. **A real weekly firing.** It is a week away. The job is confirmed **registered** on the box, which
is a different and weaker claim, and the capability map says so.
2. **Every `--read-data-subset` claim about a LARGE store.** The curve in §9 is measured on 134.3 MB
and does not extrapolate.
3. **The failure branch end-to-end through my code.** restic's damage output was captured live and fed
to the classifier as a fixture (`TestR359_RealResticDamageOutputIsClassifiedAsDamage`), but
`CheckOffboxIntegrity` reads the *configured* target, so pointing it at the damaged scratch repo
would have meant repointing the live off-site target. **The seam is named rather than glossed:**
restic-catches-it and my-code-classifies-it are each proven; the join is not.
4. **The skip path via a real backup.** See §8 — the intended trigger does not exist (R-279).
## 8. Part 5 and the live validation, in full
Evidence: `felhom.eu/documentation/tests/r359-integrity-2026-08-30/` (10 files, copied off **before**
teardown).
- **Scratch repo:** `/mnt/felhom-drives/hdd_1/r359-scratch-repo`, 3 files (195.4 KiB), snapshot
`3611a338`, hashes recorded.
- **Negative control FIRST:** healthy repo → `no errors were found`, exit 0, **703 ms**.
- **Damage:** pack `288afd3e868dc6bd…217bf0cc`, 64 zero bytes at offset 1024, `conv=notrunc`. Size
**unchanged** at 200 333 B; sha256 → `4b6847bb5eec…dc496d2b`.
- **Positive control:** the table at the top of this report.
- **The debug button that had never done anything:**
`{"data":{"duration_ms":35550,"ok":true,…},"message":"Az ellenőrzés rendben lezajlott","ok":true}`
- **R-397's notifier, end to end:**
`Event pushed: backup_integrity_ok (info) — A távoli mentés ellenőrzése rendben lezajlott. (35s)`
— pushed to the hub, HTTP 200, severity `info`, so it mailed nobody.
- **THE HAZARD CONTROL, live.** The intended demonstration could not be run: `POST
/api/backup/offbox/run` returns **404** — there is no operator-triggerable off-site backup, which is
**R-279 and stays open**. So the same flag was exercised by its other holder — two checks 6 s apart:
```
B (while A held the flag): {"skipped":true,"skip_reason":"a backup or restore is already running","duration_ms":0}
A (completed): {"skipped":false,"ok":true,"duration_ms":34953}
```
**`duration_ms: 0` is the observable that matters: B never ran restic at all.**
## 9. The three numbers R-399 needs — all MEASURED
| | |
|---|---|
| live store size | **140 829 678 B (134.3 MB)**, 2 651 blobs, 67 snapshots (`restic stats --mode raw-data`) |
| structure check wall-clock | **35.0 s** |
| 10% read-data | **35.9 s** (+0.9 s, **+3%**) |
| 50% | 37.3 s (+6%) |
| **100%** | **39.2 s (+12%)** |
**Not an estimate — measured against the live store, read-only.** At this size, re-reading *all* the
data costs about four seconds more than reading none, because the wall clock is SFTP round-trips over
the WireGuard tunnel rather than transfer. **Stated assumption and its limit:** these do not
extrapolate — the structure check's cost tracks the INDEX, a read-data run's tracks the DATA, so a
50 GB store is ~370× the data and this curve says nothing about it.
## 10. Teardown — all three layers
- **Layer 1 (host):** nothing provisioned on `demo-hp`. No VM, no CT, no storage entry.
- **Layer 2 (guest):** the scratch restic repo was created and **removed** — **1 511 424 bytes**
returned to `/mnt/felhom-drives/hdd_1`; seven throwaway scripts in the guest and the container
removed; the `mode=unit` scratch from yesterday untouched.
- **Layer 3 (hub):** no customer, appliance or host record created, so none to discard.
`demo-felhom`, `ep0`, DooPlex and Peti's box were not touched.
## 11. Register
**Size: 163 before → 165 after filing → 163 after housekeeping.**
- **Closed:** R-359, R-397 — with version and evidence path, then compressed into `CLOSED-ITEMS.md`.
- **CORRECTED, not closed: R-398** — see §12 item 1.
- **Filed:** **R-399** (the depth decision, now with three measured numbers and a re-framed premise)
and **R-400** (the debug buttons — see §12 item 2).
- **Explicitly left open and said so:** **R-87** (the restic tier is never restore-**tested** — a check
is not a restore-test, and the two rows are adjacent, which is exactly how they would get
conflated), **R-279** (no operator-triggerable off-site backup), **R-104**.
## 12. Observations
1. **NOT-A-FINDING: acted on instead — R-398 was my own mistake and is CORRECTED, not closed.**
The row said `resticStep` is not a seam so no test can drive a restic-backed path. **The first half
is true and the conclusion was false:** `offboxRunner` / `SetOffboxRunner` / `m.runner()` has been
injectable since the off-site tier shipped, and `offbox_3a_test.go` has been using it five times
over. I read one function and generalised. **Part 0 was therefore not built, and building it would
have been actively harmful:** a `resticStepFn` seam replaces `resticStep`, hiding its
`unlock --remove-all` escalation from the very assertions that must observe it. What the row asked
for that *was* real is done — R-358's AST ordering test is now an execution test, which immediately
surfaced something the AST walk could not: `unlockStale` legitimately runs before the restore.
2. **FILED: R-400 — and the sweep found EIGHT dead buttons, not one.** Comparing every
`/api/debug/...` reference in `debug.html` against every `subpath ==` case: **24 referenced, 17
dispatched.** The seven others are `backup/crossdrive`, `backup/infra`, `dr/infra-status`,
`hub/infra-push`, `storage/simulate-disconnect`, `storage/simulate-reconnect`,
`storage/watchdog-status`. Single dispatcher, exact-match switch, `default: http.NotFound` — so they
404 rather than silently succeed. **`controller/README.md` documented four `backup/*` debug routes
when two existed**; corrected. A third of a debug page does nothing, on the surface an operator
reaches for when something is already wrong.
3. **NOT-A-FINDING: noted for R-359's own row, not a separate one.** `offbox_progress.go:328` carries a
near-identical lock-retry block to `resticStep`'s. Noted, **not merged** — §12 forbade it and the
duplication is load-bearing until someone proves otherwise.
4. **NOT-A-FINDING: caught and fixed in the same session.** The damage classifier matched restic's
ordinary progress output (`"pack "` vs `check all packs`), so a check failing for a non-damage
reason would have told the customer their backups were corrupt. **The negative control caught it** —
which is precisely why a control that has only seen the failing case is worth nothing. Shipped as
v0.227.1.
5. **NOT-A-FINDING: my own measurement error, corrected before it reached a conclusion.** An early run
showed `exit=0` on the read-data subset forms while they printed `Fatal: repository contains
errors`. That was not a restic defect — the commands were piped through `tail`, so `$?` was tail's.
Re-measured without pipes: every read-data form exits 1. This project's own "exit codes that lie"
trap, caught by re-measuring rather than reasoning.
## 13. The push bypassed one gate, deliberately
**`felhom.eu` — `golden-currency` CONVICTED.** v0.227.0/0.227.1 are released and the vouched golden
carries 0.226.1. **The gate is right.** A **BYPASS, not a waiver**, and the spec directs it: §13 says
golden and fleet delivery are Viktor's (R-242). **A golden carrying 0.227.1 is OWED**, and it is item 3
under "Waiting on you" in `STATUS.md`.
**`wire-contract` was NOT bypassed — it caught something real and was answered.** It convicted
`offsite.last_integrity_ok`: the controller emits it and no hub struct can decode it. That is the
"emitting into the void" shape. It is now **allowlisted with a written reason** rather than skipped:
building a hub display is a hub change, and R-331 ruled that class a decision for the operator — the
*previous* integrity display was removed precisely because it rendered a value nothing wrote. The entry
says to delete it when a surface is built.