Files
felhom-controller/REPORT.md
T
2026-08-31 10:38:49 +02:00

21 KiB
Raw Blame History

REPORT — controller v0.228.0 (R-399 + R-400), 2026-08-31

The check now reads the data, and the debug page no longer lies. Built, pushed, deployed to demo-hp, and proven there at both depths with the restic argument list read off the running process.


1. Confirmed baselines — re-checked at the start, no drift

Repo main @ start matched origin/main version
felhom-controller 300d7e87d7cfcbb6dd355594f5d8936c06184a27 yes v0.227.1 → v0.228.0
felhom.eu db0812b6f261dbb925ef89ed27bfc2d2d1d3b5b9 yes docs only

git status --porcelain was empty in both. MinAgent stays 0.129.0. No drift to report.


2. Files created / modified / deleted

Created (6)

  • controller/internal/backup/r399_depth_test.go — Group A (8 tests)
  • controller/internal/backup/r399_slow_notice_test.go — Group B1–B5
  • controller/cmd/controller/r399_no_event_test.go — B6, the non-effect
  • controller/internal/web/r400_debug_routes_test.go — Group D
  • controller/scripts/debug_route_gate.py — the gate
  • controller/scripts/test_debug_route_gate.py — Group C, incl. both red-proofs

Modified (17)

controller/internal/backup/offbox_integrity.go · controller/internal/backup/offbox.go · controller/internal/settings/settings.go · controller/internal/config/config.go · controller/cmd/controller/main.go · controller/internal/web/handler_debug.go · controller/internal/web/templates/debug.html · controller/internal/report/types.go · controller/configs/controller.yaml.example · controller/scripts/controller_gates.py · controller/scripts/test_controller_gates.py · controller/internal/backup/r359_integrity_test.go · CHANGELOG.md · CONTEXT.md · REUSE.md · controller/README.md · .claude/rules/gates.md

felhom.eu: STATUS.md · documentation/architecture/00-capability-map.md · documentation/architecture/07-backup-architecture.md · documentation/backlog/OPEN-ITEMS.md · documentation/backlog/CLOSED-ITEMS.md · scripts/wire_contract_gate.py

DELETED — half this task, so listed explicitly

From debug.html: six /api/debug/... references, their buttons and result spans, the entire „Tárhely teszt" card (section-storage), the dr-status panel, and the JavaScript functions loadWatchdogStatus, renderWatchdogStatus, simulateDisconnect, simulateReconnect, loadDRStatus, plus their two loadSectionData cases. Template shrank 49 081 → 42 646 bytes.

From the test tree: TestR359_StructureCheckPassesNoReadDataFlag and TestR359_MalformedReadDataSubsetIsTreatedAsOff — both asserted the ruling this release reverses. They are replaced, not weakened, and a paragraph stands where each was naming its successor, so a later reader does not re-derive the old ruling from an absence. See §4.

From controller/README.md: the „Tárhely teszt" debug-section row, and the infra-push / dr/infra-status route mentions.


3. Commits pushed to main

Repo Commit What
felhom-controller 3c49dc8ea42df6c84a4bc9495d0d8a7662163ae4 the whole of Parts 1–3 + tests + gate + repo docs
felhom.eu see §13 STATUS, capability map, 07 §10.2, register, wire allowlist

4. Tests — per group, and the red-proofs by name

Green gate: go build ./... && go vet ./... && go test ./... → all three exit 0, full suite. python3 scripts/controller_gates.py --fast → 13/13 OK.

Test Result
A1 TestR399_AbsentConfigRunsFullDepth PASS
A2 TestR399_EmptyStringIsNotOff PASS
A3 TestR399_OffTokenRunsStructureOnly PASS
A4 TestR399_OffTokenIsCaseInsensitive PASS
A5 TestR399_ExplicitValueWins PASS
A6 TestR399_MalformedFallsBackToTheDefault PASS
A7 TestR399_DepthIsStatedInTheOutcome PASS
A8 TestR399_DefaultResolvesThroughTheRealConfigPath (the seam) PASS
B1 TestR399_SlowCheckWarns PASS
B2 TestR399_FastCheckIsSilent PASS
B3 TestR399_SkipNeverWarns PASS
B4 TestR399_UnreachableNeverWarns PASS
B5 TestR399_SlowAndFailedProducesBoth PASS
B6 TestR399_SlownessRaisesNoHubEvent (AST, with a positive control) PASS
C1 gate on the shipped tree PASS — 18 referenced address(es), all dispatched, none orphaned
C2 test_fails_on_an_unwired_reference PASS
C3 test_fails_on_an_unreached_handler PASS
C4 test_gate_is_registered_in_the_runner (AST over GATES) PASS
D TestR400_DeletedControlsAreGoneFromTheTemplate · …LeftNoPanelOrScript · …CrossDriveRouteDispatches · …CrossDriveReportsWhichAppsItStarted PASS
D3 TestBackupReport_DeadFieldsStayZero PASS, unmodified
E1 full existing suite PASS — see §4a for the two superseded tests

Red-proofs — mutate, confirm failure, revert

# Mutation Result
A1 defaultIntegrityReadDataSubset = "" --- FAIL: TestR399_AbsentConfigRunsFullDepth
A6 malformed falls back to "" instead of the default --- FAIL: TestR399_MalformedFallsBackToTheDefault
B1 noticeIfSlow returns immediately --- FAIL: TestR399_SlowCheckWarns and --- FAIL: TestR399_SlowAndFailedProducesBoth
C2 a /api/debug/storage/simulate-disconnect reference added to a sandbox copy gate exit 1, naming storage/simulate-disconnect
C3 a ghost/handler case added to a sandbox copy gate exit 1, naming ghost/handler

Each was reverted from a pre-mutation copy and the suite re-run green. C2/C3 mutate a throwaway copy of the two files, never the repo — a red-proof that leaves a broken tree behind when an assertion fires mid-test is its own hazard.

4a. Two existing tests were superseded — stated rather than absorbed

§9 rule E1 says to stop and report if an existing test needs editing. Two did, and the reason is not that they were inconvenient: they asserted the ruling this release reverses.

  • TestR359_StructureCheckPassesNoReadDataFlag asserted an unconfigured box passes NO --read-data flag. Correct on 2026-08-30; overturned on 2026-08-31 by measurement and by Viktor's ruling. Replaced by A1, and the off token it left room for by A3.
  • TestR359_MalformedReadDataSubsetIsTreatedAsOff — its name was the defect. Treating a typo as "off" is the quiet downgrade. The half that still holds (nothing malformed reaches restic; it WARNs) is asserted by A6, which also pins the new direction.

Nothing else in the suite was touched.

Test count

Go test/fuzz/bench funcs gate test files
before (300d7e8) 1 590 1
after 1 606 (+18 new, −2 superseded) 2 (+4 tests)

5. Deployed version

$ ssh hp "pct exec 9201 -- docker ps --filter name=felhom-controller --format '{{.Image}} {{.Status}}'"
gitea.dooplex.hu/admin/felhom-controller:0.228.0 Up 16 seconds (healthy)

Image digest sha256:ed28f160fac742dc6d6b22ccee3b7774b49b4f31d67d18099c342f5ce10f0c8d, 150 MB. Deploy path: docker pull → /etc/felhom-controller-image → restart felhom-controller-bootstrap.service. Before: 0.227.1 (up 13 h, healthy).

The fleet is on 0.227.1. A golden carrying 0.228.0 is OWED, and it is Viktor's (R-242). The golden-currency gate says so out loud — see §13.


6. The threshold I chose for the slow-check notice, and why

integritySlowNoticeThreshold = 5 * time.Minute.

The only full-depth number that exists is 39.2 s, on a 134 MB store. Five minutes is ≈7.6× that, so it cannot fire on anything resembling today's fleet — and it is well under integrityCheckTimeout (30 min), so the operator hears "this is getting slow" long before a check is killed for running too long.

Deliberately imprecise, and that is the argument. A notice changes no behaviour, so an imprecise number costs nothing; a precise-looking threshold derived from one measurement on one small store would be the exact shape of the four production designs this project has already specced against unvalidated mechanisms. The number that matters is not 5 minutes — it is that something says the setting needs revisiting before a customer's upload does.

It is a WARN in the operator log and nothing else: no hub event, no customer alarm, and it never changes the depth by itself. An event type would cost the severity contract, the grain table and three registers to say "this took a while", against 08 §6.2's coarse-by-default rule.


7. NOT yet live-validated — stated, not implied

  1. A real WEEKLY firing at the new depth. The job is confirmed REGISTERED on the box (Daily job offsite-integrity scheduled for 2026-09-01 06:00 CEST, totalJobs=13). That is not the same claim as observing it fire. Both live runs below were hand-forced through the debug route, which is the same code path with due-ness skipped and every other guard intact.
  2. Everything about a LARGE store. There is one data point, on 134 MB. The slow notice has never fired on real hardware — nothing on this fleet is slow enough. R-401 owns this.
  3. The slow-notice WARN text as rendered on a box. Proven in tests through the production path; not observed live, because no live check exceeds the threshold.
  4. The crossdrive route's zero-app branch on a real box. demo-hp had three eligible apps, so the non-empty branch was proven live and the empty one only in tests.

8. The depth evidence — argv verbatim, and the new wall-clock

The box's controller.yaml was confirmed to have no integrity: key at all before run 1 — the Scenario A condition, live, on a real box.

Run 1 — the default, no config:

restic -r sftp:u629488-sub3@u629488-sub3.your-storagebox.de:/home/felhom-repo -o sftp.command=ssh
u629488-sub3@u629488-sub3.your-storagebox.de -p 23 -oBatchMode=yes -oConnectTimeout=10
-oStrictHostKeyChecking=yes -oUserKnownHostsFile=/opt/docker/felhom-controller/data/offbox/known_hosts
-i /opt/docker/felhom-controller/data/offbox/ssh_key -s sftp check --read-data-subset=100%
{"data":{"depth":"100%","duration_ms":38745,"ok":true,"read_data_subset":"100%","skipped":false,"unreachable":false},
 "message":"Az ellenőrzés rendben lezajlott","ok":true}
[INFO] [offbox] integrity: check PASSED in 43s (structure, index, and 100% of the pack data re-read)
[INFO] Event pushed: backup_integrity_ok (info) — A távoli mentés ellenőrzése rendben lezajlott.
       (43s, a mentett adatok 100%-át újraolvasva)

Run 2 — read_data_subset: "off" added to the box's controller.yaml, container restarted:

restic -r sftp:u629488-sub3@u629488-sub3.your-storagebox.de:/home/felhom-repo -o sftp.command=ssh
u629488-sub3@u629488-sub3.your-storagebox.de -p 23 -oBatchMode=yes -oConnectTimeout=10
-oStrictHostKeyChecking=yes -oUserKnownHostsFile=/opt/docker/felhom-controller/data/offbox/known_hosts
-i /opt/docker/felhom-controller/data/offbox/ssh_key -s sftp check
{"data":{"depth":"structure","duration_ms":34742,"ok":true,"read_data_subset":"","skipped":false,"unreachable":false},
 "message":"Az ellenőrzés rendben lezajlott","ok":true}
[INFO] [offbox] integrity: check PASSED in 35s (structure and index only — no pack data was downloaded)

The config was then put back (the integrity: block removed, container restarted, grep -c integrity = 0), and the scratch backup file on the guest deleted.

2026-08-30 (0.227.x) 2026-08-31 (0.228.0)
structure depth 35.0 s 34.7 s (via off)
100% re-read 39.2 s 38.7 s · 39.8 s · 43.2 s over three runs

The numbers reproduce. The 43.2 s run was the first after a container restart, with a cold SFTP path.

Method: endpoint-level, via POST /api/debug/backup/integrity — the exact route the „Restic integritás" button invokes. No browser is available on DooPlex. The argv was read from the guest's process table while each check was in flight (pct exec 9201 -- ps -eo args | grep '[r]estic'); the container has no ps. All evidence was captured before the config revert (R-320).


9. The seven controls — disposition, by name

# control invoked by disposition why
1 backup/crossdrive button „Csak cross-drive" IMPLEMENTED Manager.RunTier2(stackName) is live at tier2.go:289 and already called from the app config page. Only the route was missing. Async (a Tier-2 sweep is bounded by disk, not by a timeout) and it answers with the app LIST, because zero apps and eight apps are different facts
2 backup/infra button „Infra mentés" DELETED no backing function exists. backup.go:22: disk-tier backup (restic, cross-drive, drive-recovery, infra-backup) moved to the host agent in slice 8C
3 hub/infra-push button „Infra backup küldése" DELETED report/pusher.go:170: "PushInfraBackup removed 2026-06-16 — the infra-backup mechanism was retired hub-side. It was dead since slice 8C, had no callers, and pushed plaintext secrets to the hub."
4 dr/infra-status fetch on page LOAD DELETED it rendered per-drive infra backups and the hub infra push — the two mechanisms above, both retired. The panel had been permanently blank
5 storage/watchdog-status fetch on page LOAD, twice (initial + 5 s poll) DELETED web/server.go:254: the slice-8C watchdog is retired and the drive-gate reconcile replaced it. Nothing publishes a per-path probe status; the panel had been permanently blank
6 storage/simulate-disconnect button in the watchdog table DELETED no backing capability at all, and it writes storage state. A debug button that fakes a drive disconnect on a customer's machine is a foot-gun — that is where drives get unenrolled and data gets stranded. No live need was shown
7 storage/simulate-reconnect button in the watchdog table DELETED same

No control was left in the third state. Each deletion took its panel and its JavaScript; the „Tárhely teszt" section had nothing left and went entirely.

Counts, so the gate has a baseline:

references in debug.html cases in handler_debug.go
before 24 17
after 18 18

Proven live on the served page (GET /debug, 200, 75 430 bytes, ASCII-only fragments per R-364): backup/infra, hub/infra-push, dr/infra-status, storage/watchdog-status, storage/simulate-disconnect, storage/simulate-reconnect, watchdog-status, simulateDisconnect, loadDRStatus, section-storage → 0 occurrences each. Positive controls on the same fetch: backup/crossdrive 1, backup/integrity 1, dr/trigger-setup 1, btn-dr-trigger 5.

The implemented control, invoked live:

{"data":{"apps":["calibre-web","paperless-ngx","romm"],"count":3},
 "message":"Cross-drive mentés elindítva 3 alkalmazásra","ok":true}

and it did the work, which is the positive observable — not an absent error:

[INFO] [backup] Tier 2 copied calibre-web   → …/backups/secondary/calibre-web   (5.6 MB, 1 leg(s))
[INFO] [backup] Tier 2 copied paperless-ngx → …/backups/secondary/paperless-ngx (79.6 MB, 1 leg(s))
[INFO] [backup] Tier 2 copied romm          → …/backups/secondary/romm          (176.2 MB, 0 leg(s))
[INFO] [web] debug cross-drive run for {calibre-web,paperless-ngx,romm} completed

The gate, and its red-proof:

### gate on the shipped tree
debug route gate OK - 18 referenced address(es), all dispatched, none orphaned
rc=0
### red-proofs
Ran 4 tests in 0.106s — OK

10. Teardown — all three layers

  • DooPlex. Nothing provisioned. One image built and pushed to the registry (felhom-controller:0.228.0), which is the deliverable, not scratch. Evidence and logs live in the session scratchpad and are reproduced verbatim in this report; nothing was left in /tmp beyond it.
  • demo-hp (host). Nothing provisioned. No storage added, no VM, no guest.
  • Guest 9201. One scratch file, /root/controller.yaml.pre-r399 (the config backup taken before the off-token run) — deleted, absence confirmed. controller.yaml restored byte-for-byte and the restored state verified (grep -c integrity = 0). The off-site store was read at full depth three times and never written, pruned, unlocked or forgotten.
  • Not touched at all: demo-felhom, ep0, DooPlex's own k3s/Longhorn/PBS, Peti's box.

11. Register

Row Action
R-399 CLOSED — controller v0.228.0. Moved to CLOSED-ITEMS.md with its reasoning kept
R-400 CLOSED — controller v0.228.0. Moved to CLOSED-ITEMS.md with its reasoning kept
R-401 FILED — revisit the depth when a real store is large. Trigger is the slow-check WARN firing, not a date. Owner CC
R-402 FILED — the integrity verdict AND its depth are on the wire and no hub surface reads either. See §12
R-87 UNTOUCHED and still OPEN. Reading the bytes back out of the store is not a restore. Restated in 07 §10.2 where the two rows sit adjacent

Register size: OPEN-ITEMS.md 166 → 165 rows (−2 closed, +2 filed). CLOSED-ITEMS.md 148 → 150 rows. Each closed entry names 300d7e8 as the commit whose git show 300d7e8:documentation/backlog/OPEN-ITEMS.md returns the original text verbatim.


12. Observations

  1. FILED: R-402 — a wire field with no receiver, caught by a gate, not by me. scripts/wire_contract_gate.py convicted the new offsite.last_integrity_depth: the literal string occurs nowhere in the hub, so encoding/json discards it on arrival. Its sibling offsite.last_integrity_ok has been in the same state since v0.227.0 and was already allowlisted with its reason. I added the new field to that allowlist beside it and filed the row rather than modelling it hub-side, because that is a hub release this task does not authorise and R-331 ruled the display a decision for the operator. This is the opposite order to the one that produced R-331: publish the value first, build the screen when someone decides what it should say.

  2. NOT-A-FINDING: integrityCheckTimeout's comment was made false by this change, and I rewrote it rather than leaving it. It said read-data "ships OFF (R-399); whoever turns it on must revisit this number, and this comment is the note that says so." Both halves stopped being true in this release. It is now the number a large store will meet first, and it says so, pointing at R-401. Strictly outside the listed scope; leaving it would have been the R-395 family defect this task itself corrects three instances of.

  3. NOT-A-FINDING: .claude/rules/gates.md said "all seven local gates" while nine were registered. Found while registering the tenth. Corrected to point at the runner's GATES table instead of listing them again — the duplicate list is what drifted.

  4. NOT-A-FINDING: demo-hp's dashboard password is NOT stale, and my first read of it was wrong. POST /login returned 200-with-login-page and the box logged Failed login, which matches the documented "the password drifted" symptom exactly. It had not: the value in ~/.config/credentials is single-quoted, and the extraction recipe recorded in project memory strips only double quotes, so the leading and trailing ' were being sent as part of the password. With both quote characters stripped, POST /login → 302 + felhom_session. Nothing on the box was changed to fix this. The memory credentials-file-values-are-quoted says values are quoted but its example strips " only.

  5. NOT-A-FINDING: the two storage/simulate-* controls are the only deletions that removed a capability someone might want back. They wrote state, so §2.1's rule deleted them absent a shown live need. If drive-absent behaviour ever needs exercising by hand again, that is a new feature with a gate in front of it, not a restored button — and the drive-gate reconcile it would be testing did not exist when those buttons were written.


13. The push bypassed one gate, deliberately

felhom.eu's golden-currency gate is RED and correctly so: controller v0.228.0 is released and the newest golden bake is 0.227.1, so a machine installed right now would receive 0.227.1. That is not a defect in this work — it is the state the gate exists to make visible, and baking the golden is Viktor's (R-242). It is not a waiver either: a golden is genuinely owed, so recording one would be false.

The felhom.eu docs push therefore used git push --no-verify, which the hook explicitly sanctions on condition that the session report says so. This is that sentence. Every other gate in that repo is green, including wire-contract, one-register, due-checks and observations. The felhom-controller push used the hook normally and all 13 gates passed.


14. What Viktor is owed

  1. Bake and vouch a golden carrying controller 0.228.0, and raise the fleet floor (currently 0.227.1). Until then the release is written, tested, pushed — and not delivered.
  2. Nothing else. No customer action, no data migration, no credential change. The debug page is operator-only and no customer sees any part of it.