REPORT.md — v0.228.0 (R-399 + R-400): proven live on demo-hp at both depths
gates / gates (push) Failing after 12s

This commit is contained in:
2026-08-31 10:38:49 +02:00
parent 3c49dc8ea4
commit b670da748d
+336 -170
View File
@@ -1,227 +1,393 @@
# REPORT — R-359 / R-397: the off-site store gets checked
# REPORT — controller v0.228.0 (R-399 + R-400), 2026-08-31
**Controller v0.227.0 → v0.227.1 · 2026-08-30 · implemented on DooPlex, validated live on `demo-hp`**
**The check now reads the data, and the debug page no longer lies.** Built, pushed, deployed to
`demo-hp`, and proven there at both depths with the restic argument list read off the running process.
---
## ⚠ READ THIS FIRST — the measurement changed what the feature is worth
## 1. Confirmed baselines — re-checked at the start, no drift
Part 5 corrupted one pack of a throwaway repository **without changing its size** (64 zero bytes at
offset 1024 — the subtlest form of rot). Both depths were run against it:
| depth | exit | verdict |
|---|---|---|
| `restic check` — **the depth that ships ON** | **0** | **`no errors were found`** |
| every `--read-data*` form — **ships OFF** | 1 | `Pack ID does not match, want 288afd3e…, got 4b6847bb…` |
**The structure check reported a corrupted store as healthy.** It verifies the index, the pack
inventory and the snapshot graph — it catches missing packs, broken indexes and unreadable snapshots,
which are real — but it does **not** re-hash pack contents. **R-399 is therefore not only a bandwidth
question: at the shipped default a class of damage is not checked at all**, and it is the class that
silently eats a customer's photos. The task could not have known this; nothing had measured it.
## 1. Confirmed baselines
| repo | spec baseline | actual at start | version |
| Repo | `main` @ start | matched `origin/main` | version |
|---|---|---|---|
| `felhom-controller` | `c476ae51` | **`e64c84ae`** (drift: one docs-only commit — my own golden-bake report) | v0.226.1 → **v0.227.1** |
| `felhom.eu` | `4f875174` | `4f875174` — **exact match** | docs only |
| `felhom-agent` | not touched | v0.130.0 | unchanged |
| `felhom-controller` | `300d7e87d7cfcbb6dd355594f5d8936c06184a27` | yes | v0.227.1 → **v0.228.0** |
| `felhom.eu` | `db0812b6f261dbb925ef89ed27bfc2d2d1d3b5b9` | yes | docs only |
Both trees clean and equal to `origin/main`. **Every symbol in §5 was re-verified present** (20/20),
and all three claimed absences confirmed: 0 occurrences of a restic `check` verb, 0 callers of
`NotifyIntegrity*` in the controller, and the debug button present with no dispatch case.
`MinAgent` stays **0.129.0**.
`git status --porcelain` was empty in both. `MinAgent` stays **0.129.0**. No drift to report.
## 2. Files created / modified
---
**`felhom-controller`:** `controller/internal/backup/offbox_integrity.go` (**new**),
`offbox.go` (report fields), `offbox_restore.go`+`r358_scratch_marker_test.go` (AST→execution test),
`internal/settings/settings.go`, `internal/config/config.go`,
`cmd/controller/main.go` (job + the one caller + the debug callback),
`internal/web/handler_debug.go`, `internal/web/handlers.go`,
tests: `r359_integrity_test.go`, `r359_lock_safety_test.go`, `r359_dueness_test.go` (**new**),
`cmd/controller/r359_wiring_test.go`, `internal/web/r397_debug_integrity_test.go` (**new**);
docs: `CHANGELOG.md`, `CONTEXT.md`, `REUSE.md`, `controller/README.md`, `REPORT.md`.
## 2. Files created / modified / deleted
**`felhom.eu`:** `STATUS.md`, `backlog/OPEN-ITEMS.md`, `backlog/CLOSED-ITEMS.md`,
`architecture/07-backup-architecture.md`, `architecture/08-alarm-ladder.md`,
`architecture/00-capability-map.md`, `scripts/wire_contract_gate.py` (allowlist **with a reason**),
`documentation/tests/r359-integrity-2026-08-30/` (**new**, 10 files).
**Created (6)**
- `controller/internal/backup/r399_depth_test.go` — Group A (8 tests)
- `controller/internal/backup/r399_slow_notice_test.go` — Group B1–B5
- `controller/cmd/controller/r399_no_event_test.go` — B6, the non-effect
- `controller/internal/web/r400_debug_routes_test.go` — Group D
- `controller/scripts/debug_route_gate.py` — the gate
- `controller/scripts/test_debug_route_gate.py` — Group C, incl. both red-proofs
**Modified (17)**
`controller/internal/backup/offbox_integrity.go` · `controller/internal/backup/offbox.go` ·
`controller/internal/settings/settings.go` · `controller/internal/config/config.go` ·
`controller/cmd/controller/main.go` · `controller/internal/web/handler_debug.go` ·
`controller/internal/web/templates/debug.html` · `controller/internal/report/types.go` ·
`controller/configs/controller.yaml.example` · `controller/scripts/controller_gates.py` ·
`controller/scripts/test_controller_gates.py` · `controller/internal/backup/r359_integrity_test.go` ·
`CHANGELOG.md` · `CONTEXT.md` · `REUSE.md` · `controller/README.md` · `.claude/rules/gates.md`
`felhom.eu`: `STATUS.md` · `documentation/architecture/00-capability-map.md` ·
`documentation/architecture/07-backup-architecture.md` · `documentation/backlog/OPEN-ITEMS.md` ·
`documentation/backlog/CLOSED-ITEMS.md` · `scripts/wire_contract_gate.py`
**DELETED — half this task, so listed explicitly**
*From `debug.html`:* six `/api/debug/...` references, their buttons and result spans, the entire
„Tárhely teszt" card (`section-storage`), the `dr-status` panel, and the JavaScript functions
`loadWatchdogStatus`, `renderWatchdogStatus`, `simulateDisconnect`, `simulateReconnect`, `loadDRStatus`,
plus their two `loadSectionData` cases. Template shrank 49 081 → 42 646 bytes.
*From the test tree:* `TestR359_StructureCheckPassesNoReadDataFlag` and
`TestR359_MalformedReadDataSubsetIsTreatedAsOff` — both asserted the ruling this release reverses.
They are **replaced, not weakened**, and a paragraph stands where each was naming its successor, so a
later reader does not re-derive the old ruling from an absence. See §4.
*From `controller/README.md`:* the „Tárhely teszt" debug-section row, and the `infra-push` /
`dr/infra-status` route mentions.
---
## 3. Commits pushed to `main`
| repo | commit | contents |
| Repo | Commit | What |
|---|---|---|
| `felhom-controller` | **`0d52a42c`** | v0.227.0 — the check, the job, the notifiers, the route, all tests |
| `felhom-controller` | **`45770f22`** | v0.227.1 — the classifier fix + CONTEXT/README/REUSE |
| `felhom.eu` | *(this report's commit)* | register, architecture, capability map, STATUS, evidence |
| `felhom-controller` | `3c49dc8ea42df6c84a4bc9495d0d8a7662163ae4` | the whole of Parts 1–3 + tests + gate + repo docs |
| `felhom.eu` | see §13 | STATUS, capability map, 07 §10.2, register, wire allowlist |
## 4. Tests and red-proofs
---
**All pass.** 22 new tests across 5 new files. Groups A (the check), B (the lock hazard), C (due-ness),
D (wiring), plus two fixtures built from restic's **real** bytes.
## 4. Tests — per group, and the red-proofs by name
| red-proof | mutation | observed failure |
Green gate: `go build ./... && go vet ./... && go test ./...` → **all three exit 0**, full suite.
`python3 scripts/controller_gates.py --fast` → **13/13 OK**.
| Test | Result |
|---|---|
| A1 `TestR399_AbsentConfigRunsFullDepth` | PASS |
| A2 `TestR399_EmptyStringIsNotOff` | PASS |
| A3 `TestR399_OffTokenRunsStructureOnly` | PASS |
| A4 `TestR399_OffTokenIsCaseInsensitive` | PASS |
| A5 `TestR399_ExplicitValueWins` | PASS |
| A6 `TestR399_MalformedFallsBackToTheDefault` | PASS |
| A7 `TestR399_DepthIsStatedInTheOutcome` | PASS |
| A8 `TestR399_DefaultResolvesThroughTheRealConfigPath` (the seam) | PASS |
| B1 `TestR399_SlowCheckWarns` | PASS |
| B2 `TestR399_FastCheckIsSilent` | PASS |
| B3 `TestR399_SkipNeverWarns` | PASS |
| B4 `TestR399_UnreachableNeverWarns` | PASS |
| B5 `TestR399_SlowAndFailedProducesBoth` | PASS |
| B6 `TestR399_SlownessRaisesNoHubEvent` (AST, with a positive control) | PASS |
| C1 gate on the shipped tree | PASS — `18 referenced address(es), all dispatched, none orphaned` |
| C2 `test_fails_on_an_unwired_reference` | PASS |
| C3 `test_fails_on_an_unreached_handler` | PASS |
| C4 `test_gate_is_registered_in_the_runner` (AST over `GATES`) | PASS |
| D `TestR400_DeletedControlsAreGoneFromTheTemplate` · `…LeftNoPanelOrScript` · `…CrossDriveRouteDispatches` · `…CrossDriveReportsWhichAppsItStarted` | PASS |
| D3 `TestBackupReport_DeadFieldsStayZero` | PASS, **unmodified** |
| E1 full existing suite | PASS — see §4a for the two superseded tests |
### Red-proofs — mutate, confirm failure, revert
| # | Mutation | Result |
|---|---|---|
| **A2** | classifier treats any non-zero as OK | `a repository restic said contains errors was reported as OK` |
| **B1** | remove the `acquireRunning` guard | `RESTIC WAS INVOKED while another operation held the single-writer flag: [… cat config] [… check]` |
| **C2** | gate due-ness on `time.Sunday` | `a 21-day-old check was NOT due on Saturday` |
| **D2** | remove the debug dispatch case | `POST /api/debug/backup/integrity was not dispatched` |
| **A1** | `defaultIntegrityReadDataSubset = ""` | `--- FAIL: TestR399_AbsentConfigRunsFullDepth` |
| **A6** | malformed falls back to `""` instead of the default | `--- FAIL: TestR399_MalformedFallsBackToTheDefault` |
| **B1** | `noticeIfSlow` returns immediately | `--- FAIL: TestR399_SlowCheckWarns` **and** `--- FAIL: TestR399_SlowAndFailedProducesBoth` |
| **C2** | a `/api/debug/storage/simulate-disconnect` reference added to a sandbox copy | gate exit 1, naming `storage/simulate-disconnect` |
| **C3** | a `ghost/handler` case added to a sandbox copy | gate exit 1, naming `ghost/handler` |
**E1 — no existing test was modified.** `TestBackupReport_DeadFieldsStayZero` (R-331, D4) passes
unmodified, which is the check that I did not write to the retired report fields.
Each was reverted from a pre-mutation copy and the suite re-run green. C2/C3 mutate a **throwaway copy
of the two files**, never the repo — a red-proof that leaves a broken tree behind when an assertion
fires mid-test is its own hazard.
**Test count:** 28 packages green before and after; +22 tests. Green gate
`go build ./... && go vet ./... && go test ./...` → **rc 0**.
### 4a. Two existing tests were superseded — stated rather than absorbed
§9 rule E1 says to stop and report if an existing test needs editing. Two did, and the reason is not
that they were inconvenient: **they asserted the ruling this release reverses.**
- `TestR359_StructureCheckPassesNoReadDataFlag` asserted an unconfigured box passes NO `--read-data`
flag. Correct on 2026-08-30; overturned on 2026-08-31 by measurement and by Viktor's ruling.
Replaced by **A1**, and the off token it left room for by **A3**.
- `TestR359_MalformedReadDataSubsetIsTreatedAsOff` — **its name was the defect.** Treating a typo as
"off" is the quiet downgrade. The half that still holds (nothing malformed reaches restic; it WARNs)
is asserted by **A6**, which also pins the new direction.
Nothing else in the suite was touched.
### Test count
| | Go test/fuzz/bench funcs | gate test files |
|---|---|---|
| before (`300d7e8`) | 1 590 | 1 |
| after | **1 606** (+18 new, −2 superseded) | **2** (+4 tests) |
---
## 5. Deployed version
```
gitea.dooplex.hu/admin/felhom-controller:0.227.1 Up 13 seconds (healthy)
[INFO] [scheduler] Daily job offsite-integrity scheduled for 2026-08-31 06:00 CEST
$ ssh hp "pct exec 9201 -- docker ps --filter name=felhom-controller --format '{{.Image}} {{.Status}}'"
gitea.dooplex.hu/admin/felhom-controller:0.228.0 Up 16 seconds (healthy)
```
**`demo-hp` only.** `demo-felhom` is on 0.226.1; the fleet floor and golden are 0.226.1.
Image digest `sha256:ed28f160fac742dc6d6b22ccee3b7774b49b4f31d67d18099c342f5ce10f0c8d`, 150 MB.
Deploy path: `docker pull` → `/etc/felhom-controller-image` → restart
`felhom-controller-bootstrap.service`. Before: 0.227.1 (up 13 h, healthy).
**v0.227.1 is a patch, not a rebuilt 0.227.0** — 0.227.0 was already running on `demo-hp` when the
classifier defect was found, and re-pushing changed bytes under a live tag is the `:latest` hazard.
The live validation in §8 ran against **0.227.0**; 0.227.1 differs only by the damage signatures and
their test.
**The fleet is on 0.227.1. A golden carrying 0.228.0 is OWED, and it is Viktor's (R-242).** The
`golden-currency` gate says so out loud — see §13.
## 6. The schedule I chose, and what I read to choose it
---
Live on `demo-hp`: `db-dump 02:30` · `tier2-backup 03:30` · `fill-watch 03:30` · `metrics-prune 04:00`
· **`offbox-backup 04:15`** · `offsite-abandon-sweep 05:10` · whole-guest gate `[04:30, 08:30)`.
## 6. The threshold I chose for the slow-check notice, and why
**Chose 06:00** — 1h45 clear of the off-site leg (which runs ~2m20s, measured) and 50 min clear of the
sweep. The backup window is customer-configurable, so **no fixed time is collision-proof for every
box**; a collision costs one skipped day, not a missed check, because due-ness makes tomorrow try
again. A slot that collided *every* night would be a real defect; this is not one.
**`integritySlowNoticeThreshold = 5 * time.Minute`.**
## 7. NOT yet live-validated
The only full-depth number that exists is **39.2 s**, on a 134 MB store. Five minutes is ≈7.6× that,
so it cannot fire on anything resembling today's fleet — and it is well under `integrityCheckTimeout`
(30 min), so the operator hears *"this is getting slow"* long before a check is killed for running too
long.
1. **A real weekly firing.** It is a week away. The job is confirmed **registered** on the box, which
is a different and weaker claim, and the capability map says so.
2. **Every `--read-data-subset` claim about a LARGE store.** The curve in §9 is measured on 134.3 MB
and does not extrapolate.
3. **The failure branch end-to-end through my code.** restic's damage output was captured live and fed
to the classifier as a fixture (`TestR359_RealResticDamageOutputIsClassifiedAsDamage`), but
`CheckOffboxIntegrity` reads the *configured* target, so pointing it at the damaged scratch repo
would have meant repointing the live off-site target. **The seam is named rather than glossed:**
restic-catches-it and my-code-classifies-it are each proven; the join is not.
4. **The skip path via a real backup.** See §8 — the intended trigger does not exist (R-279).
**Deliberately imprecise, and that is the argument.** A notice changes no behaviour, so an imprecise
number costs nothing; a precise-looking threshold derived from one measurement on one small store
would be the exact shape of the four production designs this project has already specced against
unvalidated mechanisms. The number that matters is not 5 minutes — it is that *something* says the
setting needs revisiting before a customer's upload does.
## 8. Part 5 and the live validation, in full
It is a WARN in the operator log and nothing else: no hub event, no customer alarm, and it never
changes the depth by itself. An event type would cost the severity contract, the grain table and three
registers to say "this took a while", against 08 §6.2's coarse-by-default rule.
Evidence: `felhom.eu/documentation/tests/r359-integrity-2026-08-30/` (10 files, copied off **before**
teardown).
---
- **Scratch repo:** `/mnt/felhom-drives/hdd_1/r359-scratch-repo`, 3 files (195.4 KiB), snapshot
`3611a338`, hashes recorded.
- **Negative control FIRST:** healthy repo → `no errors were found`, exit 0, **703 ms**.
- **Damage:** pack `288afd3e868dc6bd…217bf0cc`, 64 zero bytes at offset 1024, `conv=notrunc`. Size
**unchanged** at 200 333 B; sha256 → `4b6847bb5eec…dc496d2b`.
- **Positive control:** the table at the top of this report.
- **The debug button that had never done anything:**
`{"data":{"duration_ms":35550,"ok":true,…},"message":"Az ellenőrzés rendben lezajlott","ok":true}`
- **R-397's notifier, end to end:**
`Event pushed: backup_integrity_ok (info) — A távoli mentés ellenőrzése rendben lezajlott. (35s)`
— pushed to the hub, HTTP 200, severity `info`, so it mailed nobody.
- **THE HAZARD CONTROL, live.** The intended demonstration could not be run: `POST
/api/backup/offbox/run` returns **404** — there is no operator-triggerable off-site backup, which is
**R-279 and stays open**. So the same flag was exercised by its other holder — two checks 6 s apart:
```
B (while A held the flag): {"skipped":true,"skip_reason":"a backup or restore is already running","duration_ms":0}
A (completed): {"skipped":false,"ok":true,"duration_ms":34953}
```
**`duration_ms: 0` is the observable that matters: B never ran restic at all.**
## 7. NOT yet live-validated — stated, not implied
## 9. The three numbers R-399 needs — all MEASURED
1. **A real WEEKLY firing at the new depth.** The job is confirmed REGISTERED on the box
(`Daily job offsite-integrity scheduled for 2026-09-01 06:00 CEST`, `totalJobs=13`). That is not the
same claim as observing it fire. Both live runs below were hand-forced through the debug route,
which is the same code path with due-ness skipped and every other guard intact.
2. **Everything about a LARGE store.** There is one data point, on 134 MB. The slow notice has never
fired on real hardware — nothing on this fleet is slow enough. R-401 owns this.
3. **The slow-notice WARN text as rendered on a box.** Proven in tests through the production path;
not observed live, because no live check exceeds the threshold.
4. **The `crossdrive` route's zero-app branch on a real box.** demo-hp had three eligible apps, so the
non-empty branch was proven live and the empty one only in tests.
| | |
|---|---|
| live store size | **140 829 678 B (134.3 MB)**, 2 651 blobs, 67 snapshots (`restic stats --mode raw-data`) |
| structure check wall-clock | **35.0 s** |
| 10% read-data | **35.9 s** (+0.9 s, **+3%**) |
| 50% | 37.3 s (+6%) |
| **100%** | **39.2 s (+12%)** |
---
**Not an estimate — measured against the live store, read-only.** At this size, re-reading *all* the
data costs about four seconds more than reading none, because the wall clock is SFTP round-trips over
the WireGuard tunnel rather than transfer. **Stated assumption and its limit:** these do not
extrapolate — the structure check's cost tracks the INDEX, a read-data run's tracks the DATA, so a
50 GB store is ~370× the data and this curve says nothing about it.
## 8. The depth evidence — argv verbatim, and the new wall-clock
The box's `controller.yaml` was confirmed to have **no `integrity:` key at all** before run 1 — the
Scenario A condition, live, on a real box.
**Run 1 — the default, no config:**
```
restic -r sftp:u629488-sub3@u629488-sub3.your-storagebox.de:/home/felhom-repo -o sftp.command=ssh
u629488-sub3@u629488-sub3.your-storagebox.de -p 23 -oBatchMode=yes -oConnectTimeout=10
-oStrictHostKeyChecking=yes -oUserKnownHostsFile=/opt/docker/felhom-controller/data/offbox/known_hosts
-i /opt/docker/felhom-controller/data/offbox/ssh_key -s sftp check --read-data-subset=100%
```
```json
{"data":{"depth":"100%","duration_ms":38745,"ok":true,"read_data_subset":"100%","skipped":false,"unreachable":false},
"message":"Az ellenőrzés rendben lezajlott","ok":true}
```
```
[INFO] [offbox] integrity: check PASSED in 43s (structure, index, and 100% of the pack data re-read)
[INFO] Event pushed: backup_integrity_ok (info) — A távoli mentés ellenőrzése rendben lezajlott.
(43s, a mentett adatok 100%-át újraolvasva)
```
**Run 2 — `read_data_subset: "off"` added to the box's `controller.yaml`, container restarted:**
```
restic -r sftp:u629488-sub3@u629488-sub3.your-storagebox.de:/home/felhom-repo -o sftp.command=ssh
u629488-sub3@u629488-sub3.your-storagebox.de -p 23 -oBatchMode=yes -oConnectTimeout=10
-oStrictHostKeyChecking=yes -oUserKnownHostsFile=/opt/docker/felhom-controller/data/offbox/known_hosts
-i /opt/docker/felhom-controller/data/offbox/ssh_key -s sftp check
```
```json
{"data":{"depth":"structure","duration_ms":34742,"ok":true,"read_data_subset":"","skipped":false,"unreachable":false},
"message":"Az ellenőrzés rendben lezajlott","ok":true}
```
```
[INFO] [offbox] integrity: check PASSED in 35s (structure and index only — no pack data was downloaded)
```
**The config was then put back** (the `integrity:` block removed, container restarted, `grep -c
integrity` = 0), and the scratch backup file on the guest deleted.
| | 2026-08-30 (0.227.x) | 2026-08-31 (0.228.0) |
|---|---|---|
| structure depth | 35.0 s | **34.7 s** (via `off`) |
| 100% re-read | 39.2 s | **38.7 s** · 39.8 s · 43.2 s over three runs |
The numbers reproduce. The 43.2 s run was the first after a container restart, with a cold SFTP path.
**Method:** endpoint-level, via `POST /api/debug/backup/integrity` — the exact route the „Restic
integritás" button invokes. No browser is available on DooPlex. The argv was read from the **guest's**
process table while each check was in flight (`pct exec 9201 -- ps -eo args | grep '[r]estic'`); the
container has no `ps`. All evidence was captured **before** the config revert (R-320).
---
## 9. The seven controls — disposition, by name
| # | control | invoked by | disposition | why |
|---|---|---|---|---|
| 1 | `backup/crossdrive` | button „Csak cross-drive" | **IMPLEMENTED** | `Manager.RunTier2(stackName)` is live at `tier2.go:289` and already called from the app config page. Only the route was missing. Async (a Tier-2 sweep is bounded by disk, not by a timeout) and it answers with the app LIST, because zero apps and eight apps are different facts |
| 2 | `backup/infra` | button „Infra mentés" | **DELETED** | no backing function exists. `backup.go:22`: disk-tier backup (restic, cross-drive, drive-recovery, **infra-backup**) moved to the host agent in slice 8C |
| 3 | `hub/infra-push` | button „Infra backup küldése" | **DELETED** | `report/pusher.go:170`: *"PushInfraBackup removed 2026-06-16 — the infra-backup mechanism was retired hub-side. It was dead since slice 8C, had no callers, and pushed plaintext secrets to the hub."* |
| 4 | `dr/infra-status` | **fetch on page LOAD** | **DELETED** | it rendered per-drive infra backups and the hub infra push — the two mechanisms above, both retired. The panel had been permanently blank |
| 5 | `storage/watchdog-status` | **fetch on page LOAD, twice** (initial + 5 s poll) | **DELETED** | `web/server.go:254`: the slice-8C watchdog is retired and the drive-gate reconcile replaced it. Nothing publishes a per-path probe status; the panel had been permanently blank |
| 6 | `storage/simulate-disconnect` | button in the watchdog table | **DELETED** | no backing capability at all, **and it writes storage state**. A debug button that fakes a drive disconnect on a customer's machine is a foot-gun — that is where drives get unenrolled and data gets stranded. No live need was shown |
| 7 | `storage/simulate-reconnect` | button in the watchdog table | **DELETED** | same |
No control was left in the third state. Each deletion took its panel and its JavaScript; the
„Tárhely teszt" section had nothing left and went entirely.
**Counts, so the gate has a baseline:**
| | references in `debug.html` | cases in `handler_debug.go` |
|---|---|---|
| before | 24 | 17 |
| after | **18** | **18** |
**Proven live on the served page** (`GET /debug`, 200, 75 430 bytes, ASCII-only fragments per R-364):
`backup/infra`, `hub/infra-push`, `dr/infra-status`, `storage/watchdog-status`,
`storage/simulate-disconnect`, `storage/simulate-reconnect`, `watchdog-status`, `simulateDisconnect`,
`loadDRStatus`, `section-storage` → **0 occurrences each**. Positive controls on the same fetch:
`backup/crossdrive` 1, `backup/integrity` 1, `dr/trigger-setup` 1, `btn-dr-trigger` 5.
**The implemented control, invoked live:**
```json
{"data":{"apps":["calibre-web","paperless-ngx","romm"],"count":3},
"message":"Cross-drive mentés elindítva 3 alkalmazásra","ok":true}
```
and it did the work, which is the positive observable — not an absent error:
```
[INFO] [backup] Tier 2 copied calibre-web → …/backups/secondary/calibre-web (5.6 MB, 1 leg(s))
[INFO] [backup] Tier 2 copied paperless-ngx → …/backups/secondary/paperless-ngx (79.6 MB, 1 leg(s))
[INFO] [backup] Tier 2 copied romm → …/backups/secondary/romm (176.2 MB, 0 leg(s))
[INFO] [web] debug cross-drive run for {calibre-web,paperless-ngx,romm} completed
```
**The gate, and its red-proof:**
```
### gate on the shipped tree
debug route gate OK - 18 referenced address(es), all dispatched, none orphaned
rc=0
### red-proofs
Ran 4 tests in 0.106s — OK
```
---
## 10. Teardown — all three layers
- **Layer 1 (host):** nothing provisioned on `demo-hp`. No VM, no CT, no storage entry.
- **Layer 2 (guest):** the scratch restic repo was created and **removed** — **1 511 424 bytes**
returned to `/mnt/felhom-drives/hdd_1`; seven throwaway scripts in the guest and the container
removed; the `mode=unit` scratch from yesterday untouched.
- **Layer 3 (hub):** no customer, appliance or host record created, so none to discard.
- **DooPlex.** Nothing provisioned. One image built and pushed to the registry
(`felhom-controller:0.228.0`), which is the deliverable, not scratch. Evidence and logs live in the
session scratchpad and are reproduced verbatim in this report; nothing was left in `/tmp` beyond it.
- **`demo-hp` (host).** Nothing provisioned. No storage added, no VM, no guest.
- **Guest 9201.** One scratch file, `/root/controller.yaml.pre-r399` (the config backup taken before
the off-token run) — **deleted**, absence confirmed. `controller.yaml` restored byte-for-byte and the
restored state verified (`grep -c integrity` = 0). The off-site store was **read** at full depth
three times and never written, pruned, unlocked or forgotten.
- **Not touched at all:** `demo-felhom`, `ep0`, DooPlex's own k3s/Longhorn/PBS, Peti's box.
`demo-felhom`, `ep0`, DooPlex and Peti's box were not touched.
---
## 11. Register
**Size: 163 before → 165 after filing → 163 after housekeeping.**
| Row | Action |
|---|---|
| **R-399** | **CLOSED** — controller v0.228.0. Moved to `CLOSED-ITEMS.md` with its reasoning kept |
| **R-400** | **CLOSED** — controller v0.228.0. Moved to `CLOSED-ITEMS.md` with its reasoning kept |
| **R-401** | **FILED** — revisit the depth when a real store is large. **Trigger is the slow-check WARN firing, not a date.** Owner CC |
| **R-402** | **FILED** — the integrity verdict AND its depth are on the wire and no hub surface reads either. See §12 |
| **R-87** | **UNTOUCHED and still OPEN.** Reading the bytes back out of the store is not a restore. Restated in `07 §10.2` where the two rows sit adjacent |
- **Closed:** R-359, R-397 — with version and evidence path, then compressed into `CLOSED-ITEMS.md`.
- **CORRECTED, not closed: R-398** — see §12 item 1.
- **Filed:** **R-399** (the depth decision, now with three measured numbers and a re-framed premise)
and **R-400** (the debug buttons — see §12 item 2).
- **Explicitly left open and said so:** **R-87** (the restic tier is never restore-**tested** — a check
is not a restore-test, and the two rows are adjacent, which is exactly how they would get
conflated), **R-279** (no operator-triggerable off-site backup), **R-104**.
**Register size:** `OPEN-ITEMS.md` 166 → **165** rows (−2 closed, +2 filed). `CLOSED-ITEMS.md`
148 → **150** rows. Each closed entry names `300d7e8` as the commit whose
`git show 300d7e8:documentation/backlog/OPEN-ITEMS.md` returns the original text verbatim.
---
## 12. Observations
1. **NOT-A-FINDING: acted on instead — R-398 was my own mistake and is CORRECTED, not closed.**
The row said `resticStep` is not a seam so no test can drive a restic-backed path. **The first half
is true and the conclusion was false:** `offboxRunner` / `SetOffboxRunner` / `m.runner()` has been
injectable since the off-site tier shipped, and `offbox_3a_test.go` has been using it five times
over. I read one function and generalised. **Part 0 was therefore not built, and building it would
have been actively harmful:** a `resticStepFn` seam replaces `resticStep`, hiding its
`unlock --remove-all` escalation from the very assertions that must observe it. What the row asked
for that *was* real is done — R-358's AST ordering test is now an execution test, which immediately
surfaced something the AST walk could not: `unlockStale` legitimately runs before the restore.
1. **FILED: R-402 — a wire field with no receiver, caught by a gate, not by me.**
`scripts/wire_contract_gate.py` convicted the new `offsite.last_integrity_depth`: the literal
string occurs nowhere in the hub, so `encoding/json` discards it on arrival. Its sibling
`offsite.last_integrity_ok` has been in the same state since v0.227.0 and was already allowlisted
**with its reason**. I added the new field to that allowlist beside it and filed the row rather than
modelling it hub-side, because that is a hub release this task does not authorise and R-331 ruled
the *display* a decision for the operator. **This is the opposite order to the one that produced
R-331:** publish the value first, build the screen when someone decides what it should say.
2. **FILED: R-400 — and the sweep found EIGHT dead buttons, not one.** Comparing every
`/api/debug/...` reference in `debug.html` against every `subpath ==` case: **24 referenced, 17
dispatched.** The seven others are `backup/crossdrive`, `backup/infra`, `dr/infra-status`,
`hub/infra-push`, `storage/simulate-disconnect`, `storage/simulate-reconnect`,
`storage/watchdog-status`. Single dispatcher, exact-match switch, `default: http.NotFound` — so they
404 rather than silently succeed. **`controller/README.md` documented four `backup/*` debug routes
when two existed**; corrected. A third of a debug page does nothing, on the surface an operator
reaches for when something is already wrong.
2. **NOT-A-FINDING: `integrityCheckTimeout`'s comment was made false by this change, and I rewrote it
rather than leaving it.** It said read-data "ships OFF (R-399); whoever turns it on must revisit
this number, and this comment is the note that says so." Both halves stopped being true in this
release. It is now the number a large store will meet first, and it says so, pointing at R-401.
Strictly outside the listed scope; leaving it would have been the R-395 family defect this task
itself corrects three instances of.
3. **NOT-A-FINDING: noted for R-359's own row, not a separate one.** `offbox_progress.go:328` carries a
near-identical lock-retry block to `resticStep`'s. Noted, **not merged** — §12 forbade it and the
duplication is load-bearing until someone proves otherwise.
3. **NOT-A-FINDING: `.claude/rules/gates.md` said "all seven local gates" while nine were registered.**
Found while registering the tenth. Corrected to point at the runner's `GATES` table instead of
listing them again — the duplicate list is what drifted.
4. **NOT-A-FINDING: caught and fixed in the same session.** The damage classifier matched restic's
ordinary progress output (`"pack "` vs `check all packs`), so a check failing for a non-damage
reason would have told the customer their backups were corrupt. **The negative control caught it** —
which is precisely why a control that has only seen the failing case is worth nothing. Shipped as
v0.227.1.
4. **NOT-A-FINDING: `demo-hp`'s dashboard password is NOT stale, and my first read of it was wrong.**
`POST /login` returned 200-with-login-page and the box logged `Failed login`, which matches the
documented "the password drifted" symptom exactly. It had not: the value in
`~/.config/credentials` is **single-quoted**, and the extraction recipe recorded in project memory
strips only double quotes, so the leading and trailing `'` were being sent as part of the password.
With both quote characters stripped, `POST /login` → 302 + `felhom_session`. Nothing on the box was
changed to fix this. The memory
`credentials-file-values-are-quoted` says values are quoted but its example strips `"` only.
5. **NOT-A-FINDING: my own measurement error, corrected before it reached a conclusion.** An early run
showed `exit=0` on the read-data subset forms while they printed `Fatal: repository contains
errors`. That was not a restic defect — the commands were piped through `tail`, so `$?` was tail's.
Re-measured without pipes: every read-data form exits 1. This project's own "exit codes that lie"
trap, caught by re-measuring rather than reasoning.
5. **NOT-A-FINDING: the two `storage/simulate-*` controls are the only deletions that removed a
capability someone might want back.** They wrote state, so §2.1's rule deleted them absent a shown
live need. If drive-absent behaviour ever needs exercising by hand again, that is a new feature
with a gate in front of it, not a restored button — and the drive-gate reconcile it would be
testing did not exist when those buttons were written.
---
## 13. The push bypassed one gate, deliberately
**`felhom.eu` — `golden-currency` CONVICTED.** v0.227.0/0.227.1 are released and the vouched golden
carries 0.226.1. **The gate is right.** A **BYPASS, not a waiver**, and the spec directs it: §13 says
golden and fleet delivery are Viktor's (R-242). **A golden carrying 0.227.1 is OWED**, and it is item 3
under "Waiting on you" in `STATUS.md`.
`felhom.eu`'s `golden-currency` gate is **RED and correctly so**: controller v0.228.0 is released and
the newest golden bake is 0.227.1, so a machine installed right now would receive 0.227.1. That is
not a defect in this work — **it is the state the gate exists to make visible, and baking the golden
is Viktor's (R-242).** It is not a waiver either: a golden is genuinely owed, so recording one would
be false.
**`wire-contract` was NOT bypassed — it caught something real and was answered.** It convicted
`offsite.last_integrity_ok`: the controller emits it and no hub struct can decode it. That is the
"emitting into the void" shape. It is now **allowlisted with a written reason** rather than skipped:
building a hub display is a hub change, and R-331 ruled that class a decision for the operator — the
*previous* integrity display was removed precisely because it rendered a value nothing wrote. The entry
says to delete it when a surface is built.
The `felhom.eu` docs push therefore used `git push --no-verify`, which the hook explicitly sanctions
on condition that the session report says so. This is that sentence. Every other gate in that repo is
green, including `wire-contract`, `one-register`, `due-checks` and `observations`.
The `felhom-controller` push used the hook normally and all 13 gates passed.
---
## 14. What Viktor is owed
1. **Bake and vouch a golden carrying controller 0.228.0, and raise the fleet floor** (currently
0.227.1). Until then the release is written, tested, pushed — and not delivered.
2. Nothing else. No customer action, no data migration, no credential change. The debug page is
operator-only and no customer sees any part of it.