# REPORT — remote-site remediation (F1/F2/Immich) + dashboard backup-card fix (F3) **Date:** 2026-07-20 · **Repo:** felhom-controller (v0.148.0 → **v0.149.0**) · Trunk, pushed to `main`. **Origin:** `felhom.eu/documentation/audits/AUDIT-vacation-remote-ops-2026-07-20.md` (findings F1–F7). **Run class:** implementation (Part 2) + supervised operational remediation (Part 1). --- ## 1. Baselines used | Repo | Baseline @ start | Gate | After | |------|------------------|------|-------| | felhom-controller | `9d00177` | `HEAD == origin/main`, tree clean ✔ | `c059fe4` (code+CHANGELOG+CONTEXT), then this docs commit | | felhom.eu | `648ee67` | `HEAD == origin/main` ✔ | docs commit (audit addendum, ROADMAP, capability map) | Deployed controller **before**: `felhom-controller:0.148.0` (Up, healthy) — matched the baseline. Deployed controller **after**: `felhom-controller:0.149.0` (Up, healthy). --- ## 2. Part 1 — remote-site remediation, evidence chain ### 2.1 Pre-flight — the situation had already moved The router reservation had taken effect before this run began: `vmbr0` already held **`192.168.0.162`, but via DHCP** (`dhclient.vmbr0.leases` stamped 08:13 CEST). Nothing had retried the agent since systemd gave up at 07:25. The spec's duplicate-address STOP needed interpretation rather than obedience: `ping -c1 192.168.0.162` **answered** — because the host answers *itself*. The genuine test is from another node: ARP for `.162` seen from guest 9201 resolved to `68:1d:ef:5d:a6:64`, which **is** this host's own vmbr0 MAC. No duplicate → the STOP did not apply. `ip neigh` on the host was empty (a host does not ARP its own address), consistent with the same conclusion. ### 2.2 Interfaces diff (`.bak-2026-07-20` → new) One stanza; `bridge-ports` / `bridge-stp` / `bridge-fd` and every other line preserved byte-for-byte: ```diff auto vmbr0 -iface vmbr0 inet dhcp +iface vmbr0 inet static + address 192.168.0.162/24 + gateway 192.168.0.1 bridge-ports enp1s0 ``` Gateway taken from the **live** default route, not from the spec. Tabs verified with `cat -A`. Applied detached: `setsid ifreload -a >/tmp/ifreload-2026-07-20.log` → log contained only `Killed old client process` (the vmbr0 dhclient exiting). ### 2.3 Post-flip verification ``` inet 192.168.0.162/24 scope global vmbr0 (no "dynamic" flag → static) default via 192.168.0.1 dev vmbr0 proto kernel onlink curl https://gitea.dooplex.hu → 200 tailscale: dooplex … active; direct 37.191.56.193:45127 ``` Because the address was unchanged (.162 → .162) connectivity never blipped; cloudflared needed no reconnect and the public dashboard kept serving. ### 2.4 Agent recovery — Scenario A satisfied ``` systemctl reset-failed felhom-agent; systemctl start felhom-agent → active, ActiveEnterTimestamp Mon 2026-07-20 08:40:11 CEST → LISTEN 192.168.0.162:8443 users:(("felhom-agent",pid=72859)) → level=INFO msg="local-api server listening" addr=192.168.0.162:8443 → wg-felhom latest handshake: 1 minute, 16 seconds ago (< 3 min) ``` Controller side: **zero** `agentapi` / `no route to host` / `channel not verified` lines in the next 5 minutes — explicitly **not** the "agent active but controller still erroring" wrong outcome the scenario names. On an authenticated fetch of the real dashboard the red **„A tárolókezelő ügynök nem elérhető"** banner is **gone**; the only `alert-warning` strings left on the page are two modal "this cannot be undone" confirmations. The one remaining agent WARN is the pre-existing F6 (`/etc/pve/priv/storage/felhom-pbs.pw: permission denied`) — not a new failure reason, so no STOP. ### 2.5 F2 — gap closed (and it closed itself first) Restoring the agent was sufficient: the quiesce loop found the overdue whole-guest backup and ran it **unprompted** — the strongest available evidence that the seam is healthy end to end. ``` 06:40:36Z [quiesce] backup due — quiescing 3 stack(s): [bookstack calibre-web immich] 06:40:45Z backup job backup-9201-1784529645204076480 started 06:40:55Z … snapshotted — resuming app early (8B.2) 06:42:52Z … done ``` A manual app-data run followed through the real UI endpoint (`POST /api/backup/run`, authenticated session + `X-CSRF-Token` read from the page's `` — the call the „Teljes mentés" button makes): ``` 06:45:27Z DB dump: bookstack-db → bookstack-mariadb.sql (55.9 KB, 228ms, 41 tables) 06:45:28Z DB dump: immich-postgres → immich-postgres.sql (49.9 MB, 1.28s, 66 tables) 06:45:44Z Volume dump: calibre-web/…_config → 360.0 KB 06:45:55Z / 06:45:58Z / 06:45:59Z immich ml_cache 785.5 MB, postgres_data 241.5 MB, redis_data 6.6 MB 06:46:10Z App-data backup completed: 2 databases (50.0 MB total), 3 volume dump(s) (43.348s) ``` **T = 2026-07-20 06:46:10Z** (completion; newest dump artifact 06:45:28Z). All 12 containers verified `Up (healthy)` afterwards — every stack the run stopped for a safe volume dump came back. ### 2.6 Immich — RESOLVED, plus a correction to the audit Started via `POST /api/stacks/immich/start` → all four containers `Up (healthy)`. **Correction:** the audit's §3.3 probe used the wrong hostname. The traefik router rule is ``Host(`photos.demo-felhom.eu`)``, not `immich.demo-felhom.eu`. The container-down diagnosis stands (and the 404 was genuinely traefik having no backend), but that specific probe URL was invalid. On the correct host: **`https://photos.demo-felhom.eu` → 200** in 0.28 s. --- ## 3. Part 2 — the F3 fix **File:** `controller/internal/web/handlers.go`, `dashboardHandler` backup block. `dashboard.html:116` branches on `{{if .BackupStatus}}` and reads `.Success` / `.LastRun`, but the handler never set that key — so the `{{if}}` arm was unreachable and the „Még nem futott" else-branch rendered unconditionally on **every** box. The adjacent „Adatbázisok: N mentve" row kept working because it reads `DBDumpStatus`, which *was* passed: exactly the contradiction seen live. Verified before editing: `DBDumpStatus` exposes precisely the `LastRun time.Time` / `Success bool` the template consumes, and `dashboard.html` is the **only** consumer of `BackupStatus` (the `debug.html` hits are unrelated JS function names). Added inside the existing `if s.backupMgr != nil`: ```go data["BackupStatus"] = fullStatus.LastDBDump ``` No template change, no new view-model, no change to what "utolsó mentés" means. `*DBDumpStatus` nil/non-nil maps exactly onto the template's branch, so a fresh box still reads „Még nem futott". ### Files changed | File | Change | |------|--------| | `controller/internal/web/handlers.go` | +1 assignment (+4 comment lines citing the audit) | | `controller/internal/web/dashboard_backup_card_test.go` | **new** — 3 tests + 2 helpers | | `CHANGELOG.md` | v0.149.0 entry (newest on top) | | `CONTEXT.md` | remediation + v0.149.0 block | | `REPORT.md` | this file (overwritten) | --- ## 4. Tests and the red-proof The tests drive the **real handler** through `ServeHTTP` (not the template alone), backed by a real `backup.Manager` whose per-drive dump scan finds a planted `.sql` file — so they bite on the handler wiring, which is what the red-proof requires. | Scenario | Test | Result | |---|---|---| | B — real dump ⇒ its timestamp | `TestDashboardBackupCard_ShowsLastRun` | **PASS** | | C — no dump ⇒ still honest, no `0001-01-01` | `TestDashboardBackupCard_FreshBoxStaysHonest` | **PASS** | | D — failed run ⇒ „Sikertelen" | `TestDashboardBackupCard_FailedRunShowsSikertelen` | **PASS** | **Companion red-proof (§10).** Deleted `data["BackupStatus"] = fullStatus.LastDBDump` from `dashboardHandler` and re-ran: ``` --- FAIL: TestDashboardBackupCard_ShowsLastRun (0.21s) dashboard_backup_card_test.go:118: card must show the real last run "2026-07-19 03:30", got: … dashboard_backup_card_test.go:121: a box WITH a dump must not claim it never ran, got: … ``` Scenario B failed on **both** assertions (C and D correctly still passed — they do not depend on the handler line for their outcome). Line restored; suite re-run green. Scenario D asserts the template branch directly rather than through the handler: the failure flag lives only in the manager's in-memory run state, unreachable from the `web` package without shelling out to docker. That limitation is stated in the test's own comment, not hidden. **Test count: 736 → 739 (+3).** ### Green gate — honest result `go build ./...` **OK** · `go vet ./...` **OK** · `go test ./...` → **22 packages ok, 1 FAIL**. The failing package is `internal/backup` with 7 tests (`TestTier2V2_*`, `TestSharesTier2*`). **These are pre-existing and unrelated to this change** — verified by checking the baseline commit `9d00177` out into a throwaway `git worktree` and running the same tests there: **identical failures with none of my changes present**. They look environment-dependent on DooPlex (the tier-2 tests reason about distinct physical devices, and every `t.TempDir()` here lands on the same filesystem) — the same class as the recorded "Windows green gate silently red" lesson. `internal/web`, the package this change touches, is **green**. I did not attempt to fix them: out of scope for this task, and quietly absorbing a red gate would be worse than reporting it. --- ## 5. Build / deploy / live verification ``` build.sh 0.149.0 --push → Build complete ✓ gitea.dooplex.hu/admin/felhom-controller:0.149.0 (145M) pct exec 9201: docker pull … && echo … > /etc/felhom-controller-image && systemctl restart felhom-controller-bootstrap.service docker ps → gitea.dooplex.hu/admin/felhom-controller:0.149.0 Up 5 seconds (healthy) ``` **Scenario B, live on the real dashboard** — authenticated endpoint fetch of the rendered page (no browser on DooPlex; `claude-in-chrome` unavailable per CLAUDE.md): ```html Utolsó mentés: 2026-07-20 06:45 Adatbázisok: 3 mentve ``` The identical page rendered „Még nem futott" before the deploy. The card shows 06:45 (newest dump artifact) rather than 06:46 (run completion) because the deploy restarted the controller and cleared the in-memory run state, so it renders from the dumps on disk — the honest value, and precisely the path that must never fabricate a timestamp when there are no dumps. --- ## 6. Not done / open | Item | Status | |---|---| | **Durable F1** — control plane pinned to a LAN literal | **OPEN → R-50** (island-bridge, **spike-first**). Today's static `.162` is a *window mitigation*: it still assumes the site's subnet is `192.168.0.0/24` and `.162` is free | | **F4** — dead-primary alerting in multi-container apps | **OPEN → R-51** (constraint recorded: must NOT simply fold `unhealthy` into down — that reverts fix-3) | | **F5** — boot desired-state reconciliation | **OPEN → R-52** (includes root-causing the `unless-stopped` non-resurrection; the evidence is gone, so it needs a fresh repro) | | **F6** — agent cannot read the PBS `.pw` | **OPEN, already tracked** as R-39's remaining item (b). No duplicate item minted | | **F7** — `app_export.html` CSRF-token-as-domain | **OPEN → R-53** | | `internal/backup` 7 red tests | **OPEN, pre-existing** — not investigated (out of scope) | Capability map gained a **PARTIAL** row: *"Box survives a site/network change (relocation, different subnet, DHCP re-lease) with the control plane intact"*, citing the audit and pointing at R-50/51/52. --- ## 7. Observations (noticed, not acted on) 1. **An orphaned `dhclient` runs on the host for a non-existent `eth0`** (`pgrep -a dhclient` → `/sbin/dhclient … eth0`, pid 1922) while no `eth0` exists in `ip addr`. Harmless today, but it is leftover state of exactly the kind that makes a future network change confusing to diagnose. 2. **Accented grep patterns are unreliable through the `ssh → pct exec → bash -c` chain.** A search for „nem elérhető" returned a false `0` while the string was present — which reads exactly like "the banner cleared". The page content is fine; the *pattern* gets mangled in transit. Use ASCII-only substrings for any grep that gates a conclusion. This nearly produced a wrong "banner gone" claim before an ASCII re-check caught it. 3. **The controller password contains `'` and `!`.** A helper script written via an unquoted heredoc had its `!!` mangled by history expansion and produced one failed login (visible in the controller log as `[WARN] [web] Failed login`). The proven-safe form is the inline `-d "password=$PW"` the audit run used. 4. The guest journal is **volatile** (current boot only) and the controller's logs are lost on container recreate — which is why F5's root cause could not be established after the fact. Worth knowing before any post-mortem that depends on in-guest history.