# REPORT — R-77: endpoint-drift detection, samba protected-set gate, channel log honesty (2026-07-26) **Overwritten** per the standing rule. Controller **v0.173.0** live on demo-felhom 9201, demo-hp 9201 and drill-r50 9201; hub **v0.74.0** live. MinAgent unchanged; `felhom-agent` untouched. --- ## 1. Baselines used | Repo | start | end | version | |---|---|---|---| | felhom-controller | head after R-75 (`c7a3a90`) | `9056f01` | `v0.172.0` → **`v0.173.0`** | | felhom.eu | `9e94479` | `9cfa619` | hub `v0.73.1` live → **`v0.74.0`** | | felhom-agent | `dfd5d731` | untouched | `v0.96.0` | | app-catalog-felhom.eu | `4252121` | untouched | — | --- ## 2. Part 0 — operational repair (supervised, done first) Its own STOP condition — *"if `fingerprint` or `token` also disagree, STOP"* — **did not trigger**: only the address had moved on either box. | | demo-felhom 9201 | demo-hp 9201 | |---|---|---| | endpoint BEFORE | `192.168.0.162:8443` | `192.168.0.87:8443` | | endpoint AFTER | `169.254.253.1:8443` | `169.254.253.1:8443` | | **fingerprint agrees** | **true** | **true** | | **token agrees** | **true** | **true** | | backup | `/var/lib/docker/volumes/felhom-controller-data/_data/controller.yaml.pre-r77.bak` | same path | | `[channel]` lines in 90 s after restart | **0** | **0** | | hub `agent_channel_*` events since | **0** | **0** | Values were redacted at the point of collection — the probe compared `fingerprint`/`token` and emitted booleans plus a 12-char digest; neither value left the box. ### On the verbatim `agent channel recovered` line the brief asked for: it does not exist, by design `Check` returns early on a healthy probe — `return nil // healthy first-obs / steady-up → no notify` — and `NotifyRecovered` fires only on a `down→up` transition **within one process's lifetime**. The repair required a restart, so the new process's first observation is healthy and therefore silent. There is no recovery line to quote, and reporting one would mean fabricating it. The positive evidence instead: - **zero** `[channel]` lines 90 s after restart on both boxes, where the pre-repair container logged `transient down` within 60 s and confirmed `DOWN` within 120 s; - the dashboard banner cleared (fetched through the authed endpoint, ASCII-safe grep); - **zero** new `agent_channel_*` hub events since 06:45 UTC; - demo-felhom health `ok` continuously. This is now written into the runbook, because "wait for the recovered line" would leave an operator watching an empty log forever. --- ## 3. Files changed **felhom-controller** (`9056f01`, plus `REPORT.md`): | file | change | |---|---| | `internal/bootstrap/bootstrap.go` | `DetectEndpointDrift`, `EndpointDrift` + its two message builders | | `internal/bootstrap/drift_test.go` | NEW — Scenarios A/B/C/D + the secret-leak assertions | | `internal/web/alerts.go` | `endpointDriftAlert` field, `SetEndpointDriftAlert`, ordered ABOVE the channel banner in `GetAlerts` | | `internal/notify/notifier.go` | `NotifyEndpointDrift` + `EndpointDriftDetails` | | `cmd/controller/main.go` | startup wiring (once, after config settles) | | `internal/monitor/healthcheck.go` | `EffectiveProtected` gate + the corrected doc comment | | `internal/monitor/effective_protected_test.go` | Scenario E; existing toggle test updated | | `internal/channelhealth/checker.go` | `stateUnconfirmed`, `orUnseeded`, re-arm condition, field comment | | `internal/channelhealth/checker_test.go` | Scenario F ×2 | **felhom.eu** (`9cfa619`): `hub/internal/api/handler.go` (allowlist), `hub/CHANGELOG.md`, `manifests/hub.yaml`, `documentation/runbooks/RUNBOOK-local-api-endpoint-drift.md` (NEW), ROADMAP (R-77 shipped + R-78/79/80), capability-map note. ### Scope addition, flagged: the hub was NOT in the brief's scope, and had to be `handleEvent` validates `event_type` against `allowedEventTypes` and returns `400 Invalid event_type` otherwise. Shipping `local_api_endpoint_drift` controller-side alone would have produced an alert that never reaches the operator — the inert-seam class this project has hit four times, and precisely the failure mode R-77 exists to prevent. I added the one-line allowlist entry, bumped the hub to v0.74.0 and deployed it. Live leg 2 proves the end-to-end path. --- ## 4. Tests + the three mandatory red-proofs **951 → 959 test functions; 27 controller packages green; hub green.** `go build ./... && go vet ./... && go test ./...` clean in both repos. ### Red-proof A — run in BOTH failure directions *(i) the pre-fix shape (v0.172.0: no detection at all):* ``` --- FAIL: TestScenarioA_DriftDetectedAndNamed_NoWrite (0.00s) drift_test.go:78: drift must be DETECTED — this is the exact live shape from the 2026-07-25 outage ``` *(ii) the tempting wrong turn — a variant that detects **and helpfully corrects**, i.e. R-78 done prematurely:* ``` --- FAIL: TestScenarioA_DriftDetectedAndNamed_NoWrite (0.00s) drift_test.go:101: controller.yaml was MODIFIED — detection must never write (that is R-78) ``` The second is the one that matters: "an error was logged" is a hollow assertion; "and nothing was written" is the contract. ### Red-proof E ``` --- FAIL: TestScenarioE_SambaProtectedOnlyWhenActuallyDeployed (0.00s) effective_protected_test.go:99: (2) sharing ON, no password: samba protected = true, want false — THE BUG: reconcileSambaAt refuses to deploy without a password — a deliberate state, not a fault ``` ### Red-proof F ``` --- FAIL: TestScenarioF_BornDownLogsUnseededNotUp (0.00s) checker_test.go:308: an unseeded checker must not report state "up" after a suppressed first failure checker_test.go:318: born-down must log "unseeded->down:unreachable"; got: [WARN] [channel] agent channel DOWN (up->down:unreachable): dial tcp 192.168.0.87:8443: … ``` And the non-change half: all **nine** pre-existing `channelhealth` tests still pass, including `TestF2_BornDownNonTransient_AlertsOnce`, `TestF2_OldSeedSilentLogicWouldNotAlert` and `TestF2_BornDownTransient_Debounced`. Scenario F asserts sink call **count and arguments** (`reason`, `eventType`, `severity`, dashboard flag, probe count), not just the log string. ### One pre-existing test was deliberately changed `TestEffectiveProtectedTracksSharingToggle` used `SMBSettings{Enabled: true}` and asserted samba WAS watched — i.e. it encoded the bug. Updated to `Enabled: true, UserSet: true` with a comment saying why, and the three-state matrix added alongside. Flagged here because "changed a passing test" always deserves to be visible. ### Gates `template_id_gate`, `emoji_gate`, `mojibake_gate`, `native_confirm_gate`, `app_row_dedup_gate`, `offbox_rename_gate` — rc=0. `docker_run_volume_path_gate` **rc=1, PRE-EXISTING** (R-29(a), red since v0.129.0, in `appexport/estimate.go` which this change does not touch). --- ## 5. Deployed versions ``` demo-felhom 9201 : gitea.dooplex.hu/admin/felhom-controller:0.173.0 Up (healthy) demo-hp 9201 : gitea.dooplex.hu/admin/felhom-controller:0.173.0 Up (healthy) drill-r50 9201 : gitea.dooplex.hu/admin/felhom-controller:0.173.0 Up (healthy) hub : gitea.dooplex.hu/admin/felhom-hub:0.74.0 Synced / Healthy, rolled out ``` --- ## 6. Live legs **Leg 1 — no drift alert on the aligned production boxes (Scenario B live).** Zero `DRIFT` lines and zero `[channel]` lines on both after the v0.173.0 restart. Agreement really is silent; the check does not cry wolf on every healthy boot. **Leg 2 — `drill-r50`, the untouched real drift (Scenario A live).** It was left broken on purpose. Before (controller 0.161.0): ``` controller.yaml sha256 BEFORE : 54892f1bc691c29b70ff9c323aab9e0b6ad758c602c533130064006e9a2af737 controller.yaml endpoint : 192.168.0.176:8443 bootstrap.json endpoint : 169.254.253.1:8443 DRIFTED : True ``` After deploying v0.173.0: ``` [ERROR] bootstrap: local_api endpoint DRIFT — /opt/docker/felhom-controller/controller.yaml says "192.168.0.176:8443" but /etc/felhom-bootstrap/bootstrap.json says "169.254.253.1:8443"; the controller is dialling the FORMER. Pin agrees: true. Not auto-corrected (R-78 owns the authority ruling) — fix the intended file and restart the controller. [INFO] Event pushed: local_api_endpoint_drift (error) — … controller.yaml sha256 AFTER : 54892f1bc691c29b70ff9c323aab9e0b6ad758c602c533130064006e9a2af737 ``` **Checksums identical — zero write, proven on a live divergence.** And hub-side, proving the allowlist end-to-end: ``` events: 2026-07-26 07:21:31 drill-r50 local_api_endpoint_drift error notification_log: 2026-07-26 07:21:32 drill-r50 local_api_endpoint_drift operator sent ``` `drill-r50` was **left drifted** — repairing it is the operator's call, and it is now the only live fixture for this alert. **Leg 3 — demo-hp's false samba alarm is gone, sharing still on without a password.** ``` [DEBUG] [monitor] Checking 4 protected containers: [traefik cloudflared felhom-controller filebrowser] [DEBUG] [monitor] All protected containers running [DEBUG] [monitor] Final status: ok (issues=0, warnings=0, info=5) ``` Hub-side the flip is unambiguous: `fail` on 0.172.0 at 07:15:56 → **`ok` on 0.173.0 at 07:18:27**. **Leg 4 — the other half of Scenario E, live (operator-present state change).** Household password set through the real `POST /sharing/password` endpoint (303): ``` felhom-samba Up 45 seconds [DEBUG] [monitor] Checking 5 protected containers: [traefik cloudflared felhom-controller filebrowser felhom-samba] [DEBUG] [monitor] Final status: ok (issues=0, warnings=0, info=5) ``` Samba deployed and **is watched again** — the fix suppresses the deliberate state, not the real one. --- ## 7. NOT yet live-validated - **A drift where the FINGERPRINT also moved.** Unit-tested (`TestDrift_FingerprintDisagreementIsSurfacedNotLeaked`); no live fixture exists, and manufacturing one would mean deliberately corrupting a pin. - **The drift banner rendered in a browser.** The alert and event are proven; the Hungarian dashboard line is unit/`AlertManager`-level only. No browser on DooPlex. - **Drift on a box where `controller.yaml` is the CORRECT file.** The detector is symmetric by construction, but every live case so far has `bootstrap.json` correct. This is exactly the asymmetry R-78 must rule on. - **A recovery-line transition** (`down→up` within one process). The repair path restarts, so it is structurally unreachable; the `recovered` branch is unit-tested only. - **The samba gate through a full enable→password→disable cycle.** Enable-without-password and password-set were both exercised; disabling sharing again was not. - **Whether `drill-r50`'s channel actually recovers** once its endpoint is aligned — it was left drifted on purpose. --- ## 8. Observations — noticed, not acted on 1. **`manifests/hub.yaml` was pinned to `0.73.1` while `hub/CHANGELOG.md` topped out at `0.73.2`** — the documented "live image can lag the CHANGELOG" trap, found while bumping. v0.73.2's content is in the 0.74.0 image now, but **v0.73.2 as a released artifact was never deployed**. Worth a glance at what else it contained. 2. **`expected_backup_missed` still fires nightly on all three customers** — filed as **R-80**, with the arithmetic the DIAG hedged on: 7.3 days of stale backup materially exceeds the ~1.5-day channel outage, so backups were already stale ~6 days *before* the channel broke. R-77's repair will not resolve it, and it reaches a **customer** channel. It plausibly outranks this task. 3. **`report.Issues`/`Warnings` remain English on a Hungarian surface** — R-79, whole-surface, deliberately not swept here. 4. **The drift check is startup-only.** Correct today (both files are read at boot and neither changes under a running controller), but if anything ever rewrites `bootstrap.json` live, the check would not notice until the next restart. Named in the code comment so the assumption is visible. 5. **`mergeLocalAPI` replaces the whole block**, so endpoint/fingerprint/token cannot be reconciled independently today. That constrains R-78's design space and is recorded in its ROADMAP entry. 6. **Sharing on demo-hp now has a household password and samba is running** (leg 4). That is a real, deliberate state change from the brief, and it changes the box's LAN exposure — SMB is now actually served. Flagging so it is a decision, not a surprise.