25 KiB
REPORT — controller v0.228.0 (R-399 + R-400), 2026-08-31
The check now reads the data, and the debug page no longer lies. Built, pushed, deployed to
demo-hp, and proven there at both depths with the restic argument list read off the running process.
1. Confirmed baselines — re-checked at the start, no drift
| Repo | main @ start |
matched origin/main |
version |
|---|---|---|---|
felhom-controller |
300d7e87d7cfcbb6dd355594f5d8936c06184a27 |
yes | v0.227.1 → v0.228.0 |
felhom.eu |
db0812b6f261dbb925ef89ed27bfc2d2d1d3b5b9 |
yes | docs only |
git status --porcelain was empty in both. MinAgent stays 0.129.0. No drift to report.
2. Files created / modified / deleted
Created (6)
controller/internal/backup/r399_depth_test.go— Group A (8 tests)controller/internal/backup/r399_slow_notice_test.go— Group B1–B5controller/cmd/controller/r399_no_event_test.go— B6, the non-effectcontroller/internal/web/r400_debug_routes_test.go— Group Dcontroller/scripts/debug_route_gate.py— the gatecontroller/scripts/test_debug_route_gate.py— Group C, incl. both red-proofs
Modified (17)
controller/internal/backup/offbox_integrity.go · controller/internal/backup/offbox.go ·
controller/internal/settings/settings.go · controller/internal/config/config.go ·
controller/cmd/controller/main.go · controller/internal/web/handler_debug.go ·
controller/internal/web/templates/debug.html · controller/internal/report/types.go ·
controller/configs/controller.yaml.example · controller/scripts/controller_gates.py ·
controller/scripts/test_controller_gates.py · controller/internal/backup/r359_integrity_test.go ·
CHANGELOG.md · CONTEXT.md · REUSE.md · controller/README.md · .claude/rules/gates.md
felhom.eu: STATUS.md · documentation/architecture/00-capability-map.md ·
documentation/architecture/07-backup-architecture.md · documentation/backlog/OPEN-ITEMS.md ·
documentation/backlog/CLOSED-ITEMS.md · scripts/wire_contract_gate.py
DELETED — half this task, so listed explicitly
From debug.html: six /api/debug/... references, their buttons and result spans, the entire
„Tárhely teszt" card (section-storage), the dr-status panel, and the JavaScript functions
loadWatchdogStatus, renderWatchdogStatus, simulateDisconnect, simulateReconnect, loadDRStatus,
plus their two loadSectionData cases. Template shrank 49 081 → 42 646 bytes.
From the test tree: TestR359_StructureCheckPassesNoReadDataFlag and
TestR359_MalformedReadDataSubsetIsTreatedAsOff — both asserted the ruling this release reverses.
They are replaced, not weakened, and a paragraph stands where each was naming its successor, so a
later reader does not re-derive the old ruling from an absence. See §4.
From controller/README.md: the „Tárhely teszt" debug-section row, and the infra-push /
dr/infra-status route mentions.
3. Commits pushed to main
| Repo | Commit | What |
|---|---|---|
felhom-controller |
3c49dc8ea42df6c84a4bc9495d0d8a7662163ae4 |
the whole of Parts 1–3 + tests + gate + repo docs |
felhom.eu |
77a5a115 |
STATUS, capability map, 07 §10.2, register, wire allowlist |
4. Tests — per group, and the red-proofs by name
Green gate: go build ./... && go vet ./... && go test ./... → all three exit 0, full suite.
python3 scripts/controller_gates.py --fast → 13/13 OK.
| Test | Result |
|---|---|
A1 TestR399_AbsentConfigRunsFullDepth |
PASS |
A2 TestR399_EmptyStringIsNotOff |
PASS |
A3 TestR399_OffTokenRunsStructureOnly |
PASS |
A4 TestR399_OffTokenIsCaseInsensitive |
PASS |
A5 TestR399_ExplicitValueWins |
PASS |
A6 TestR399_MalformedFallsBackToTheDefault |
PASS |
A7 TestR399_DepthIsStatedInTheOutcome |
PASS |
A8 TestR399_DefaultResolvesThroughTheRealConfigPath (the seam) |
PASS |
B1 TestR399_SlowCheckWarns |
PASS |
B2 TestR399_FastCheckIsSilent |
PASS |
B3 TestR399_SkipNeverWarns |
PASS |
B4 TestR399_UnreachableNeverWarns |
PASS |
B5 TestR399_SlowAndFailedProducesBoth |
PASS |
B6 TestR399_SlownessRaisesNoHubEvent (AST, with a positive control) |
PASS |
| C1 gate on the shipped tree | PASS — 18 referenced address(es), all dispatched, none orphaned |
C2 test_fails_on_an_unwired_reference |
PASS |
C3 test_fails_on_an_unreached_handler |
PASS |
C4 test_gate_is_registered_in_the_runner (AST over GATES) |
PASS |
D TestR400_DeletedControlsAreGoneFromTheTemplate · …LeftNoPanelOrScript · …CrossDriveRouteDispatches · …CrossDriveReportsWhichAppsItStarted |
PASS |
D3 TestBackupReport_DeadFieldsStayZero |
PASS, unmodified |
| E1 full existing suite | PASS — see §4a for the two superseded tests |
Red-proofs — mutate, confirm failure, revert
| # | Mutation | Result |
|---|---|---|
| A1 | defaultIntegrityReadDataSubset = "" |
--- FAIL: TestR399_AbsentConfigRunsFullDepth |
| A6 | malformed falls back to "" instead of the default |
--- FAIL: TestR399_MalformedFallsBackToTheDefault |
| B1 | noticeIfSlow returns immediately |
--- FAIL: TestR399_SlowCheckWarns and --- FAIL: TestR399_SlowAndFailedProducesBoth |
| C2 | a /api/debug/storage/simulate-disconnect reference added to a sandbox copy |
gate exit 1, naming storage/simulate-disconnect |
| C3 | a ghost/handler case added to a sandbox copy |
gate exit 1, naming ghost/handler |
Each was reverted from a pre-mutation copy and the suite re-run green. C2/C3 mutate a throwaway copy of the two files, never the repo — a red-proof that leaves a broken tree behind when an assertion fires mid-test is its own hazard.
4a. Two existing tests were superseded — stated rather than absorbed
§9 rule E1 says to stop and report if an existing test needs editing. Two did, and the reason is not that they were inconvenient: they asserted the ruling this release reverses.
TestR359_StructureCheckPassesNoReadDataFlagasserted an unconfigured box passes NO--read-dataflag. Correct on 2026-08-30; overturned on 2026-08-31 by measurement and by Viktor's ruling. Replaced by A1, and the off token it left room for by A3.TestR359_MalformedReadDataSubsetIsTreatedAsOff— its name was the defect. Treating a typo as "off" is the quiet downgrade. The half that still holds (nothing malformed reaches restic; it WARNs) is asserted by A6, which also pins the new direction.
Nothing else in the suite was touched.
Test count
| Go test/fuzz/bench funcs | gate test files | |
|---|---|---|
before (300d7e8) |
1 590 | 1 |
| after | 1 606 (+18 new, −2 superseded) | 2 (+4 tests) |
5. Deployed version
$ ssh hp "pct exec 9201 -- docker ps --filter name=felhom-controller --format '{{.Image}} {{.Status}}'"
gitea.dooplex.hu/admin/felhom-controller:0.228.0 Up 16 seconds (healthy)
Image digest sha256:ed28f160fac742dc6d6b22ccee3b7774b49b4f31d67d18099c342f5ce10f0c8d, 150 MB.
Deploy path: docker pull → /etc/felhom-controller-image → restart
felhom-controller-bootstrap.service. Before: 0.227.1 (up 13 h, healthy).
DELIVERED the same session, on the operator's instruction. Golden 0.228.0 baked, published,
round-trip verified and vouched; the fleet floor raised 0.227.1 → 0.228.0. demo-felhom then
self-updated in under a minute and re-registered offsite-integrity by itself — so both demo boxes
now re-read their whole off-site store weekly, and only demo-hp was ever touched by hand. Full
evidence: felhom.eu/documentation/tests/golden-0.228.0-2026-08-31/. See §15.
6. The threshold I chose for the slow-check notice, and why
integritySlowNoticeThreshold = 5 * time.Minute.
The only full-depth number that exists is 39.2 s, on a 134 MB store. Five minutes is ≈7.6× that,
so it cannot fire on anything resembling today's fleet — and it is well under integrityCheckTimeout
(30 min), so the operator hears "this is getting slow" long before a check is killed for running too
long.
Deliberately imprecise, and that is the argument. A notice changes no behaviour, so an imprecise number costs nothing; a precise-looking threshold derived from one measurement on one small store would be the exact shape of the four production designs this project has already specced against unvalidated mechanisms. The number that matters is not 5 minutes — it is that something says the setting needs revisiting before a customer's upload does.
It is a WARN in the operator log and nothing else: no hub event, no customer alarm, and it never changes the depth by itself. An event type would cost the severity contract, the grain table and three registers to say "this took a while", against 08 §6.2's coarse-by-default rule.
7. NOT yet live-validated — stated, not implied
- A real WEEKLY firing at the new depth. The job is confirmed REGISTERED on the box
(
Daily job offsite-integrity scheduled for 2026-09-01 06:00 CEST,totalJobs=13). That is not the same claim as observing it fire. Both live runs below were hand-forced through the debug route, which is the same code path with due-ness skipped and every other guard intact. - Everything about a LARGE store. There is one data point, on 134 MB. The slow notice has never fired on real hardware — nothing on this fleet is slow enough. R-401 owns this.
- The slow-notice WARN text as rendered on a box. Proven in tests through the production path; not observed live, because no live check exceeds the threshold.
- The
crossdriveroute's zero-app branch on a real box. demo-hp had three eligible apps, so the non-empty branch was proven live and the empty one only in tests.
8. The depth evidence — argv verbatim, and the new wall-clock
The box's controller.yaml was confirmed to have no integrity: key at all before run 1 — the
Scenario A condition, live, on a real box.
Run 1 — the default, no config:
restic -r sftp:u629488-sub3@u629488-sub3.your-storagebox.de:/home/felhom-repo -o sftp.command=ssh
u629488-sub3@u629488-sub3.your-storagebox.de -p 23 -oBatchMode=yes -oConnectTimeout=10
-oStrictHostKeyChecking=yes -oUserKnownHostsFile=/opt/docker/felhom-controller/data/offbox/known_hosts
-i /opt/docker/felhom-controller/data/offbox/ssh_key -s sftp check --read-data-subset=100%
{"data":{"depth":"100%","duration_ms":38745,"ok":true,"read_data_subset":"100%","skipped":false,"unreachable":false},
"message":"Az ellenőrzés rendben lezajlott","ok":true}
[INFO] [offbox] integrity: check PASSED in 43s (structure, index, and 100% of the pack data re-read)
[INFO] Event pushed: backup_integrity_ok (info) — A távoli mentés ellenőrzése rendben lezajlott.
(43s, a mentett adatok 100%-át újraolvasva)
Run 2 — read_data_subset: "off" added to the box's controller.yaml, container restarted:
restic -r sftp:u629488-sub3@u629488-sub3.your-storagebox.de:/home/felhom-repo -o sftp.command=ssh
u629488-sub3@u629488-sub3.your-storagebox.de -p 23 -oBatchMode=yes -oConnectTimeout=10
-oStrictHostKeyChecking=yes -oUserKnownHostsFile=/opt/docker/felhom-controller/data/offbox/known_hosts
-i /opt/docker/felhom-controller/data/offbox/ssh_key -s sftp check
{"data":{"depth":"structure","duration_ms":34742,"ok":true,"read_data_subset":"","skipped":false,"unreachable":false},
"message":"Az ellenőrzés rendben lezajlott","ok":true}
[INFO] [offbox] integrity: check PASSED in 35s (structure and index only — no pack data was downloaded)
The config was then put back (the integrity: block removed, container restarted, grep -c integrity = 0), and the scratch backup file on the guest deleted.
| 2026-08-30 (0.227.x) | 2026-08-31 (0.228.0) | |
|---|---|---|
| structure depth | 35.0 s | 34.7 s (via off) |
| 100% re-read | 39.2 s | 38.7 s · 39.8 s · 43.2 s over three runs |
The numbers reproduce. The 43.2 s run was the first after a container restart, with a cold SFTP path.
Method: endpoint-level, via POST /api/debug/backup/integrity — the exact route the „Restic
integritás" button invokes. No browser is available on DooPlex. The argv was read from the guest's
process table while each check was in flight (pct exec 9201 -- ps -eo args | grep '[r]estic'); the
container has no ps. All evidence was captured before the config revert (R-320).
9. The seven controls — disposition, by name
| # | control | invoked by | disposition | why |
|---|---|---|---|---|
| 1 | backup/crossdrive |
button „Csak cross-drive" | IMPLEMENTED | Manager.RunTier2(stackName) is live at tier2.go:289 and already called from the app config page. Only the route was missing. Async (a Tier-2 sweep is bounded by disk, not by a timeout) and it answers with the app LIST, because zero apps and eight apps are different facts |
| 2 | backup/infra |
button „Infra mentés" | DELETED | no backing function exists. backup.go:22: disk-tier backup (restic, cross-drive, drive-recovery, infra-backup) moved to the host agent in slice 8C |
| 3 | hub/infra-push |
button „Infra backup küldése" | DELETED | report/pusher.go:170: "PushInfraBackup removed 2026-06-16 — the infra-backup mechanism was retired hub-side. It was dead since slice 8C, had no callers, and pushed plaintext secrets to the hub." |
| 4 | dr/infra-status |
fetch on page LOAD | DELETED | it rendered per-drive infra backups and the hub infra push — the two mechanisms above, both retired. The panel had been permanently blank |
| 5 | storage/watchdog-status |
fetch on page LOAD, twice (initial + 5 s poll) | DELETED | web/server.go:254: the slice-8C watchdog is retired and the drive-gate reconcile replaced it. Nothing publishes a per-path probe status; the panel had been permanently blank |
| 6 | storage/simulate-disconnect |
button in the watchdog table | DELETED | no backing capability at all, and it writes storage state. A debug button that fakes a drive disconnect on a customer's machine is a foot-gun — that is where drives get unenrolled and data gets stranded. No live need was shown |
| 7 | storage/simulate-reconnect |
button in the watchdog table | DELETED | same |
No control was left in the third state. Each deletion took its panel and its JavaScript; the „Tárhely teszt" section had nothing left and went entirely.
Counts, so the gate has a baseline:
references in debug.html |
cases in handler_debug.go |
|
|---|---|---|
| before | 24 | 17 |
| after | 18 | 18 |
Proven live on the served page (GET /debug, 200, 75 430 bytes, ASCII-only fragments per R-364):
backup/infra, hub/infra-push, dr/infra-status, storage/watchdog-status,
storage/simulate-disconnect, storage/simulate-reconnect, watchdog-status, simulateDisconnect,
loadDRStatus, section-storage → 0 occurrences each. Positive controls on the same fetch:
backup/crossdrive 1, backup/integrity 1, dr/trigger-setup 1, btn-dr-trigger 5.
The implemented control, invoked live:
{"data":{"apps":["calibre-web","paperless-ngx","romm"],"count":3},
"message":"Cross-drive mentés elindítva 3 alkalmazásra","ok":true}
and it did the work, which is the positive observable — not an absent error:
[INFO] [backup] Tier 2 copied calibre-web → …/backups/secondary/calibre-web (5.6 MB, 1 leg(s))
[INFO] [backup] Tier 2 copied paperless-ngx → …/backups/secondary/paperless-ngx (79.6 MB, 1 leg(s))
[INFO] [backup] Tier 2 copied romm → …/backups/secondary/romm (176.2 MB, 0 leg(s))
[INFO] [web] debug cross-drive run for {calibre-web,paperless-ngx,romm} completed
The gate, and its red-proof:
### gate on the shipped tree
debug route gate OK - 18 referenced address(es), all dispatched, none orphaned
rc=0
### red-proofs
Ran 4 tests in 0.106s — OK
10. Teardown — all three layers
- DooPlex. Nothing provisioned. One image built and pushed to the registry
(
felhom-controller:0.228.0), which is the deliverable, not scratch. Evidence and logs live in the session scratchpad and are reproduced verbatim in this report; nothing was left in/tmpbeyond it. demo-hp(host). Nothing provisioned. No storage added, no VM, no guest.- Guest 9201. One scratch file,
/root/controller.yaml.pre-r399(the config backup taken before the off-token run) — deleted, absence confirmed.controller.yamlrestored byte-for-byte and the restored state verified (grep -c integrity= 0). The off-site store was read at full depth three times and never written, pruned, unlocked or forgotten. - Not touched at all:
demo-felhom,ep0, DooPlex's own k3s/Longhorn/PBS, Peti's box.
11. Register
| Row | Action |
|---|---|
| R-399 | CLOSED — controller v0.228.0. Moved to CLOSED-ITEMS.md with its reasoning kept |
| R-400 | CLOSED — controller v0.228.0. Moved to CLOSED-ITEMS.md with its reasoning kept |
| R-401 | FILED — revisit the depth when a real store is large. Trigger is the slow-check WARN firing, not a date. Owner CC |
| R-402 | FILED — the integrity verdict AND its depth are on the wire and no hub surface reads either. See §12 |
| R-87 | UNTOUCHED and still OPEN. Reading the bytes back out of the store is not a restore. Restated in 07 §10.2 where the two rows sit adjacent |
Register size: OPEN-ITEMS.md 166 → 165 rows (−2 closed, +2 filed). CLOSED-ITEMS.md
148 → 150 rows. Each closed entry names 300d7e8 as the commit whose
git show 300d7e8:documentation/backlog/OPEN-ITEMS.md returns the original text verbatim.
12. Observations
-
FILED: R-402 — a wire field with no receiver, caught by a gate, not by me.
scripts/wire_contract_gate.pyconvicted the newoffsite.last_integrity_depth: the literal string occurs nowhere in the hub, soencoding/jsondiscards it on arrival. Its siblingoffsite.last_integrity_okhas been in the same state since v0.227.0 and was already allowlisted with its reason. I added the new field to that allowlist beside it and filed the row rather than modelling it hub-side, because that is a hub release this task does not authorise and R-331 ruled the display a decision for the operator. This is the opposite order to the one that produced R-331: publish the value first, build the screen when someone decides what it should say. -
NOT-A-FINDING:
integrityCheckTimeout's comment was made false by this change, and I rewrote it rather than leaving it. It said read-data "ships OFF (R-399); whoever turns it on must revisit this number, and this comment is the note that says so." Both halves stopped being true in this release. It is now the number a large store will meet first, and it says so, pointing at R-401. Strictly outside the listed scope; leaving it would have been the R-395 family defect this task itself corrects three instances of. -
NOT-A-FINDING:
.claude/rules/gates.mdsaid "all seven local gates" while nine were registered. Found while registering the tenth. Corrected to point at the runner'sGATEStable instead of listing them again — the duplicate list is what drifted. -
NOT-A-FINDING:
demo-hp's dashboard password is NOT stale — I made the exact mistake the project already has a memory about.POST /loginreturned 200-with-login-page and the box loggedFailed login, which matches the documented "the password drifted" symptom exactly. It had not: values in~/.config/credentialsare single-quoted, my extraction stripped only", and the two'characters were being sent as part of the password. With both quote characters stripped,POST /login→ 302 +felhom_session. Nothing on the box was changed. The memorycredentials-file-values-are-quotedstates this correctly, gives the right recipe (tr -d "\"'"), and records the identical misdiagnosis from 2026-07-20 — where it was written up three times as "the stored password is stale" before being caught. The memory is right; I did not follow it. Its own lesson is the one that applies: an auth failure is evidence about the bytes you sent, not proof about the stored secret. -
NOT-A-FINDING: the two
storage/simulate-*controls are the only deletions that removed a capability someone might want back. They wrote state, so §2.1's rule deleted them absent a shown live need. If drive-absent behaviour ever needs exercising by hand again, that is a new feature with a gate in front of it, not a restored button — and the drive-gate reconcile it would be testing did not exist when those buttons were written.
13. One push bypassed a gate, deliberately — and the gate is now green
The first felhom.eu docs push (77a5a115) used git push --no-verify. At that moment
golden-currency was RED and correctly so: v0.228.0 was released and the newest golden bake was
0.227.1, so a machine installed right then would have received 0.227.1. That was not a defect in the
work — it is the state the gate exists to make visible. It was not a waiver case either: a golden was
genuinely owed, so recording one would have been false. The hook sanctions --no-verify on condition
that the session report says so; this is that sentence.
It is no longer red. The golden was baked and vouched later in the same session (§15), the gate
went red → green, and the follow-up push (1623a4d5) passed the hook normally with all 12 gates
OK. Both felhom-controller pushes used the hook normally, 13/13.
14. What Viktor is owed
Nothing. No customer action, no data migration, no credential change. The debug page is operator-only and no customer sees any part of it. The delivery that §13 recorded as owed was carried out on the operator's instruction and is verified in §15.
15. Delivery — golden 0.228.0 baked, vouched, and the floor raised
GOLDEN_VERSION |
0.228.0 |
GOLDEN_SHA256 |
76a3a98b9e7cc23bf8ae51b38a6272f576df285cb34cd22235ac3f06a31e53ec |
| size | 658 079 744 B |
| script | build-golden.sh v3.0.0, drill VM reverted to virgin and cold-booted |
| template | debian-13-standard_13.6-1_amd64.tar.zst, after pveam update |
Acceptance markers, counted: docker OK (overlay2 1 · including mount point rootfs 1 ·
including mount point mp0 1 · upload OK (HTTP 201) 1 · excluding 0 · FATAL 0 ·
mount point mp1 0. felhom-controller:0.228.0 appears 4× in the bake log.
The 404 pre-gate was proven before its 404 was believed: the target URL returned 404 while the existing 0.227.1 package returned 200 on the same command. A 404 from a check that cannot see anything is not a measurement.
The evidence is the round trip. Downloaded size and sha match the bake exactly, and
./etc/felhom-controller-image read out of the downloaded archive says
gitea.dooplex.hu/admin/felhom-controller:0.228.0 — the golden naming its controller from the bytes a
customer's box would actually fetch.
The vouch — three fields together: golden_version 0.228.0, agent_version 0.130.0,
min_agent 0.129.0 (read from this release's CHANGELOG header, not assumed), wrapper_sha256
carried through explicitly because the handler clears it when omitted. agent ≥ min_agent, so not
the R-216 shape. Verified by RE-READING the manifest, never by the flash. The R-120 gate passed
rather than being bypassed.
The floor is ACTING, not merely set. Impact preview {"below":3,"valid":true,"version":"0.228.0"};
re-read after the POST confirms DB override: v0.228.0. Then, with nobody touching it:
[INFO] [selfupdate] Post-update startup: update successful (0.227.1 → 0.228.0)
[INFO] [scheduler] Daily job offsite-integrity scheduled for 2026-09-01 06:00 CEST
[INFO] [offsite-apply] settle-gate: GO — at/above floor 0.228.0 (we are 0.228.0)
That is demo-felhom. The second line is the one worth keeping: a box nobody deployed to now runs
the deeper off-site check on its own schedule — R-399 reaching the fleet, observed rather than assumed.
Token hygiene and teardown are recorded in full at
felhom.eu/documentation/tests/golden-0.228.0-2026-08-31/README.md. In short: file → file scp, a
runner script inside the VM so the token never reached a command line (systemctl show … | grep -c
→ 0), the committed log's leak grep proven with a planted copy (1) before its 0 was believed,
build guest 9100 purged, secrets shredded after the log was copied out (standing rule 5), VM
powered off and the disk reverted to virgin.
Still NOT live-validated, unchanged from §7: a real weekly firing at the new depth (next is 2026-09-01 06:00 CEST, now on both boxes), and everything about a large store (R-401).