Files
felhom.eu/documentation/audits/AUDIT-vacation-remote-ops-2026-07-20.md

25 KiB
Raw Permalink Blame History

AUDIT — vacation remote-ops: temp access, connection/health diagnosis, app triage

Date: 2026-07-20 · Status: ACTIVE (remote-site window until ~2026-08-02) Operator: Claude Code on DooPlex · Target: demo host demo-felhom (N100) + guest 9201 Run class: runbook / diagnosis. Docs-only commits. No code changes, no version bumps, no builds.

Baselines: felhom-controller 29eda5d86e (v0.148.0, = deployed) · felhom-agent d81fddc7b1 (deployed binary 0.90.0) · felhom.eu (audit doc only).


1. Access change summary

felhom-pve was physically relocated to a remote site. Access is via Tailscale.

Item Result
~/.ssh/config on DooPlex Already correctHost felhom-pveHostName 100.70.170.35. No edit needed. A separate Host felhom-pve-lan192.168.0.162 entry preserves the home-LAN path.
Host-key identity check PASS — ed25519 key at 100.70.170.35 is byte-identical to the stored key for 192.168.0.162 (…OOz0BjPXLRCI3Z9C…). Same machine, not a reinstall.
Tailnet path direct, 37.191.56.193:45127 (not DERP-relayed). ICMP 4/4, rtt min/avg/max 34.7 / 36.8 / 38.3 ms.
Phase 0.4 alias verification ssh felhom-pve "pct exec 9201 -- docker ps …"gitea.dooplex.hu/admin/felhom-controller:0.148.0 | Up 31 minutes (healthy). Full CC hop works.

The tailscale health check on the host notes "Some peers are advertising routes but --accept-routes is false" — expected and correct per the travel-only decision recorded in tailscale-n100-location-independent-2026-07-19.


2. Ground-truth snapshot (Phase A)

Host demo-felhom

vmbr0        192.168.0.147/24  dynamic     <-- DHCP, NOT the former static .162
tailscale0   100.70.170.35/32
wg-felhom    10.77.0.2/32
default via 192.168.0.1 dev vmbr0
/etc/resolv.conf: nameserver 192.168.0.1          (no Pi-hole; remote router is the resolver)
timedatectl: 2026-07-20 07:56 CEST / 05:56 UTC — synchronized: yes, NTP active, RTC sane

The remote site uses the same 192.168.0.0/24 subnet as home, but the host now holds .147, not .162. This single fact is the root cause of §3.1.

Boot history — the box was powered off for ~17h45m in transport:

-1  Sat 2026-07-18 18:17:00 CEST → Sun 2026-07-19 13:40:40 CEST
 0  Mon 2026-07-20 07:25:05 CEST → (current)

WireGuard (PBS offsite leg) is healthy: peer 167.233.158.164:443, latest handshake 56 s, keepalive 25 s, 1.70 MiB rx / 1.00 MiB sx.

Guest 9201

eth0  192.168.0.104/24  dynamic
/etc/resolv.conf: nameserver 192.168.0.1
curl https://gitea.dooplex.hu → 200 (0.196 s)      outbound + DNS healthy

DooPlex

tailscale is not a host binary here — it runs as the k3s pod admin-system/tailscale-6d7bdc9c5c-hwtjq (uptime 6d22h). The peer view from the host shows dooplex 100.107.87.53 … active; direct.

Nothing in this section was changed.


3. Per-symptom evidence and root cause

3.1 "A tárolókezelő ügynök nem elérhető" — ROOT-CAUSED (critical)

The agent service is failed, not merely unreachable. Smoking gun from journalctl -u felhom-agent:

level=ERROR msg="daemon: exited with error"
  err="localapi: bind 192.168.0.162:8443: listen tcp 192.168.0.162:8443:
       bind: cannot assign requested address"
felhom-agent.service: Start request repeated too quickly.
Active: failed (Result: exit-code) since Mon 2026-07-20 07:25:51 CEST

The agent's localapi listener is pinned to the literal 192.168.0.162. That address no longer exists on any interface (the host DHCP'd .147), so bind() fails and the daemon exits ~1.1 s after start. systemd retried 4× and gave up. The agent has therefore been down since the first boot at the remote site — it has never run there.

Controller side, the matching error class is no route to host (not a TLS-pin or auth failure — those would be STOP conditions of a different kind):

[ERROR] [quiesce] cycle error: check due: agentapi: GET /backup/due:
  Get "https://192.168.0.162:8443/backup/due": dial tcp 192.168.0.162:8443: connect: no route to host
[WARN] local-api: GET /storage failed (… no route to host) — channel not verified

So both ends are independently pinned to 192.168.0.162: the agent binds it, the controller dials it. Nothing is listening (ss -tlnp | grep felhom → empty). This is precisely the OPEN item recorded on 2026-07-19 — "control plane pinned to 192.168.0.162 (agent listen_addr + guest bootstrap.json) → not portable off-LAN" — now realized in production.

Allowed action taken (§3.2 of the runbook): systemctl reset-failed && systemctl restart felhom-agent. It failed identically within 6 s:

Jul 20 08:00:02 … err="localapi: bind 192.168.0.162:8443: … cannot assign requested address"
Jul 20 08:00:02 … felhom-agent.service: Failed with result 'exit-code'.

STOP condition hit (runbook §3.4, "agent endpoint stale because the host IP changed"). The fix requires editing /etc/felhom-agent/agent.json (listen_addr) and the guest's bootstrap-managed endpoint — both forbidden in this run. Escalated as F1.

Note for the fix: re-pinning to the literal 192.168.0.147 would work today but is a DHCP lease away from breaking again, and the agent pins its served leaf cert, whose SAN set must cover whatever address the controller dials. Options are laid out in F1.

Blast radius while F1 stands: every agent-backed function is dead — storage management (Tárhely), PBS whole-guest backups, the quiesce cycle, restore-tests, and the DR path. Guest-local functions (dashboard, apps, app-data/DB backup cache, Samba, cloudflared) are unaffected.

3.2 Calibre-Web Automated stopped — RESOLVED

exit=128  finished=2026-07-19T11:40:15Z  policy=unless-stopped  restarts=0

Container log ends with a clean application shutdown (webserver stop (restart=False), Performing shutdown of Calibre-Web Automated) at 13:40:15 CEST — ~25 s before the host powered off for transport. Not a crash: an orderly stop caught by the shutdown, which then never restarted.

Allowed action taken: started via the controller's real endpoint (the exact path the UI uses — session login, <meta name="csrf-token"> read from the page, X-CSRF-Token header):

POST /api/stacks/calibre-web/start → {"ok":true,"message":"Stack calibre-web start completed"}
docker ps → calibre-web | Up 20 seconds (healthy)

Verified healthy. Stack state now reads running.

3.3 Immich "Nem egészséges" + "URL nem elérhető" — ROOT-CAUSED, NOT fixed (per runbook D.3)

The two flags have one cause: immich-server is not running at all.

immich-server            Exited (143) 18 hours ago   policy=unless-stopped  restarts=0
                         finished=2026-07-19T11:40:15Z   health=unhealthy
immich-postgres          Up 31 minutes (healthy)
immich-machine-learning  Up 31 minutes (healthy)
immich-redis             Up 31 minutes (healthy)

Exit 143 = SIGTERM at 11:40:15Z — the same second as Calibre-Web, i.e. the same pre-transport shutdown. Its final log lines are a fully healthy server (Immich Server is listening … [v3.0.3], ML server healthy, websocket clients connected) right up to the shutdown. This is not leftover state from the 2026-07-19 offsite-restore test (R-43/R-44) — hypothesis (a) is disproved by the logs: the server ran normally for ~42 min after that work, and there is no migration or DB error.

The URL probe is likewise not a remote-path/latency artifact — hypothesis (c) is disproved:

curl https://immich.demo-felhom.eu   → 404 in 0.239 s   (traefik has no live backend)
curl https://gitea.dooplex.hu        → 200 in 0.196 s   (same path, fine)

Fast 404, not a timeout. The public CF path is healthy; there is simply nothing behind the route. Left down deliberately per runbook D.3. One-line fix available under GO — see F5.

3.4 "Utolsó mentés: Még nem futott" — ROOT-CAUSED as a display bug (independent of F1)

This is hypothesis (c), and it is provable from code rather than inferred. dashboardHandler (controller/internal/web/handlers.go:134) populates the backup card with:

data["BackupEnabled"]    = s.cfg.Backup.Enabled        // :172
data["DBDumpStatus"]     = fullStatus.LastDBDump       // :175
data["BackupRunning"]    = fullStatus.Running          // :177
data["BackupMaxAgeHours"] =                           // :178

It never sets data["BackupStatus"] — and no other code path does either (repo-wide, the only occurrences of BackupStatus are the agent-proxy handler and this template). But dashboard.html:116 gates the whole "Utolsó mentés" value on it:

{{if .BackupStatus}}   … LastRun / "Sikertelen" …
{{else}}               <span class="backup-status-none">Még nem futott</span>
{{end}}

.BackupStatus is therefore always nil, the {{if}} branch is dead code, and the dashboard reports "Még nem futott" permanently on every box regardless of actual backup history. The adjacent "Adatbázisok: 2 mentve" renders from DBDumpStatus, which is set — which is exactly the contradiction observed.

Independent corroboration that backups had in fact been running: the hub's own host-reports consistently carry 1 backups, and hub events on 2026-07-19 include crossdrive_completed for bookstack / calibre-web / immich / Megosztások at 03:30 and db_dump_completed at 02:30.

Separately, and not a display artifact: no backup can run at all while F1 stands — the quiesce cycle fails every 5 min on GET /backup/due. The guest-local backup-cache job still runs fine (Found 2 DB dump files across drives, Discovered 2 databases, Discovered app data: 3 apps every 5 min). No manual backup was triggered: the runbook permits one only if the cause was (a) and the agent is healthy — neither holds.

3.5 Hub report freshness — HEALTHY, no gap

Reports land on the normal 15-min cadence from the remote site, and the staleness state machine recovered by itself on boot:

07:25:42  Event from demo-felhom: controller_started (info) — Controller elindult (0.148.0)
07:25:43  Staleness: demo-felhom down → ok (node_recovered)
07:27:07  Event: app_start_failed — Telepített alkalmazás nem fut: Calibre-Web Automated
07:27:40  Event: agent_channel_unknown (warning) — Controller→agent channel down: … no route to host
07:27:41  Operator email sent for demo-felhom/agent_channel_unknown
07:40:37 / 07:55:37  Received report from demo-felhom (7603 / 6232 bytes)

Hub telemetry is location-independent (it rides the public path), and the alerting chain worked end to end — the agent-channel failure produced an operator email within 33 s of the boot grace expiring. The transport gap appears exactly as it should: last report 2026-07-19 11:40:22Z, next 2026-07-20 05:25 UTC.


4. Ranked findings

# Sev Finding Evidence Proposed fix Viktor GO?
F1 CRITICAL Agent localapi binds the literal 192.168.0.162; host is .147 at the remote site → bind: cannot assign requested address → service failed since first remote boot. Controller dials the same literal. All agent-backed features (storage, PBS backup, quiesce, restore-test, DR) are dead for the whole 2-week window. §3.1 Bind 0.0.0.0:8443 (or the vmbr0 address resolved at start-up) instead of a literal, and give the controller a stable dial target; ensure the served leaf's SAN covers it. Short-term unblock: set listen_addr to 192.168.0.147 + reissue/verify the pin + update the guest endpoint — but this re-breaks on the next DHCP lease. YES — config and code; touches bootstrap-managed guest state
F2 HIGH No PBS/whole-guest backup can run while F1 stands → a ~2-week data-protection gap on the demo box. Quiesce errors every 5 min. §3.1, §3.4 Resolve F1; then verify a backup completes. Follows F1
F3 HIGH Dashboard permanently shows "Utolsó mentés: Még nem futott" on every box: dashboardHandler never sets BackupStatus, so the template's {{if .BackupStatus}} branch is unreachable dead code. Customer-facing misinformation about backup state. §3.4 (handlers.go:134-178, dashboard.html:116-131) Populate data["BackupStatus"] from fullStatus in dashboardHandler, or drop the dead branch and render from DBDumpStatus/GuestBackup. Needs a red-proof test (the current code has none covering this key). YES — code
F4 MEDIUM Multi-container apps under-alert. IsDownState (stacks/manager.go:46) counts only stopped/exited; it deliberately excludes unhealthy. Immich's stack aggregates to unhealthy because 3 of 4 containers are up — so with its primary container fully exited for 18 h and the app 100 % unreachable, it produced no dead-app banner and no app_start_failed hub event. Calibre-Web (single container) alerted correctly within 90 s. §3.3, live /api/stacksimmich unhealthy; hub log has an event for Calibre-Web only Treat "primary/named container down" as down for a multi-container stack, or add an unhealthy-for-N-cycles escalation. YES — code
F5 MEDIUM Two app containers were left Exited by the pre-transport shutdown and were not restarted on boot despite restart: unless-stopped; the controller reports deployed-but-stopped apps (30 s deadapp-check) but never starts them. Immich has now been down 18 h+ and stays down. §3.2, §3.3, cmd/controller/main.go:435 Start immich via the same real endpoint used for Calibre-Web (immediate, no code). Then consider a bounded boot-time reconciliation (start deployed-but-stopped apps once, N attempts, with the existing boot grace). Start = YES (runbook D.3 forbade it here); reconciler = code, YES
F6 LOW Agent logs pbs: cannot read token secret; skipping datastore storage=felhom-pbs err="open /etc/pve/priv/storage/felhom-pbs.pw: permission denied" on every start (pre-existing, R-39-adjacent — the non-root agent cannot read the PVE priv file). §3.1 journal Confirm whether the wrapper/sudoers path is meant to cover this read; fold into the R-39 follow-up. YES
F7 LOW Incidental defect found while tracing CSRF: app_export.html:93 builds var domain = '<subdomain>.{{$.CSRFToken}}' — the session CSRF token is substituted where the domain belongs. Produces a broken "open in browser" link and puts a CSRF token into a URL. controller/internal/web/templates/app_export.html:93 Replace {{$.CSRFToken}} with the domain value (cf. the neighbouring correct csrfH() helper). YES — code

Healthy / no action: tailnet path direct at 37 ms · host clock NTP-synced (no transport RTC drift) · WireGuard PBS leg handshaking at 56 s · guest outbound + DNS · hub reports on cadence and staleness auto-recovered · operator alert email fired correctly · 6 apps + Samba + traefik + cloudflared running · USB HDD mounted and visible.


5. Actions taken

Action Authority Outcome
systemctl reset-failed && systemctl restart felhom-agent runbook §3.2 (one allowed restart, after evidence) Failed identically in 6 s — same bind error. Confirms F1 is deterministic config, not a transient.
POST /api/stacks/calibre-web/start via the controller's real session+CSRF path runbook §3.2 (allowed, after evidence) SuccessUp (healthy), stack state running.

Nothing else was mutated. The ssh alias needed no edit (already on the tailnet IP).


6. Explicitly NOT fixed / not touched

  • F1 — no edit to /etc/felhom-agent/agent.json, the guest bootstrap.json, or any bootstrap/golden-managed state. STOP condition §3.4 hit and honoured.
  • Immich — left down per runbook D.3 (diagnose only), despite the fix being a one-call start.
  • F3, F4, F7 — code changes, out of scope for this run.
  • No network/nftables/tailscale/route changes on either machine. No manual backup triggered. No version bumps, image builds, deploys, or REPORT.md overwrites. No containers or volumes deleted. No secrets written to any committed file.

7. Return-day checklist (~2026-08-02)

  1. If F1 was fixed with a literal .147, revert it to the home-LAN value or (preferred) confirm the address-independent fix still binds correctly back on the home LAN.
  2. ~/.ssh/config needs no revert — Host felhom-pve on the tailnet IP works from both locations; felhom-pve-lan remains as the direct-LAN path. Re-verify the .162/LAN path once the host is home and holds its expected address.
  3. Delete the TEMPORARY blocks from felhom-controller/CLAUDE.md and felhom-agent/CLAUDE.md.
  4. Confirm the host reacquires its intended home address (it is now on DHCP — the former static .162 config did not survive/apply; decide whether to restore a static or keep the reservation).
  5. Re-verify: agent active (running), dashboard banner clear, Tárhely renders, a PBS backup completes, Immich healthy and reachable.
  6. Flip this document's status from ACTIVE to CLOSED.

Remediation — 2026-07-20

Executed the same day as the audit, on the relocated host. Mutations: one network file, one service start, one app start, two backup runs (one of them the product's own scheduler), one code line.

R.1 Static-IP flip (F1 window mitigation)

State found had already moved. By the time remediation started, the router's reservation had taken effect and vmbr0 held 192.168.0.162 again — but via DHCP (dhclient.vmbr0.leases stamped 08:13 CEST). So the address existed; nothing had retried the agent since systemd gave up.

The 1.1 duplicate-address check needed interpretation, not blind obedience: ping -c1 192.168.0.162 answered, because the host was answering itself. The genuine test is from another node — ARP for .162 seen from guest 9201 resolved to 68:1d:ef:5d:a6:64, which is this host's vmbr0 MAC. No duplicate; the STOP did not apply.

Interfaces diff (.bak-2026-07-20 → new) — one stanza, nothing else touched:

 auto vmbr0
-iface vmbr0 inet dhcp
+iface vmbr0 inet static
+	address 192.168.0.162/24
+	gateway 192.168.0.1
 	bridge-ports enp1s0
 	bridge-stp off
 	bridge-fd 0

The gateway was taken from the live default route (192.168.0.1), not from the spec. Applied detached: setsid ifreload -a → log contained only Killed old client process (the vmbr0 dhclient). Post-flip: inet 192.168.0.162/24 scope global vmbr0 (no dynamic flag), default via 192.168.0.1 … proto kernel onlink, curl https://gitea.dooplex.hu200, tailscale peer still active; direct. Because the address was unchanged, connectivity never blipped.

R.2 Agent recovery (Scenario A)

systemctl reset-failed && systemctl start felhom-agent — up at 08:40:11 CEST:

active
LISTEN 192.168.0.162:8443  users:(("felhom-agent",pid=72859))
level=INFO msg="local-api server listening" addr=192.168.0.162:8443
wg-felhom latest handshake: 1 minute, 16 seconds ago

Controller side, agentapi errors stopped entirely (no no route to host, no channel not verified in the following 5 min) — so this is not the "agent active but controller still erroring" wrong outcome the scenario warns about. On an authenticated fetch of the real dashboard the red „A tárolókezelő ügynök nem elérhető" banner is gone; the only alert-warning strings left on the page are two modal "this cannot be undone" confirmations.

Method note for anyone repeating this: grep the fetched page with ASCII-only patterns. Accented patterns sent through the ssh→pct exec→shell layers get mangled and return a false 0, which reads exactly like "the banner is gone" when it is not. The page content itself is fine — it is the pattern that gets corrupted.

R.3 F2 — the protection gap closed itself, then a manual run confirmed it

Restoring the agent was sufficient: the quiesce loop noticed the overdue whole-guest backup and ran it unprompted, which is the strongest possible evidence the seam is healthy end to end:

06:40:36Z [quiesce] backup due — quiescing 3 stack(s): [bookstack calibre-web immich]
06:40:45Z [quiesce] backup job backup-9201-1784529645204076480 started — polling to completion
06:40:55Z [quiesce] … snapshotted — resuming app early (8B.2)
06:42:52Z [quiesce] … done

A manual app-data run then followed via the real UI endpoint (POST /api/backup/run, authenticated session + X-CSRF-Token read from the page's <meta> — the same call the „Teljes mentés" button makes), started 06:45:26Z:

06:45:27Z DB dump: bookstack-db → bookstack-mariadb.sql (55.9 KB, 228ms, 41 tables)
06:45:28Z DB dump: immich-postgres → immich-postgres.sql (49.9 MB, 1.28s, 66 tables)
06:45:44Z Volume dump: calibre-web/…_config → 360.0 KB
06:45:55Z Volume dump: immich/…_ml_cache → 785.5 MB
06:45:58Z Volume dump: immich/…_postgres_data → 241.5 MB
06:45:59Z Volume dump: immich/…_redis_data → 6.6 MB
06:46:10Z App-data backup completed: 2 databases (50.0 MB total), 3 volume dump(s) (43.348s)

T = 2026-07-20 06:46:10Z (run completion; newest dump artifact 06:45:28Z). Every container the run stopped for a safe volume dump was restarted — all 12 containers verified Up (healthy) afterwards.

R.4 Immich (Scenario: app triage)

Started via the controller's real endpoint (POST /api/stacks/immich/start) → all four containers Up (healthy).

Correction to §3.3 of the audit. That section reported the URL probe returning 404 and treated it as a live symptom. The probe used the wrong hostname: the traefik router rule is Host(`photos.demo-felhom.eu`), not immich.demo-felhom.eu. The container-down diagnosis stands (and the 404 was genuinely traefik having no backend), but the specific probe URL in §3.3 was invalid. Post-start, on the correct host: https://photos.demo-felhom.eu200 in 0.28 s.

R.5 F3 — dashboard backup card (controller v0.149.0)

Handler-side one-liner, no template change: data["BackupStatus"] = fullStatus.LastDBDump in dashboardHandler. Deployed to guest 9201 and verified on the real, authenticated dashboard:

<span class="backup-label">Utolsó mentés:</span>
<span class="backup-value"><span>2026-07-20 06:45</span></span>
<span class="backup-label">Adatbázisok:</span>
<span class="backup-value">3 mentve</span>

Was „Még nem futott" on the identical page before the deploy. (The controller restart cleared the in-memory run state, so the card renders from the dumps on disk — newest artifact 06:45 — which is the honest value and exactly what the fresh-box branch must NOT fabricate.)

R.6 Per-finding status after remediation

# Was Now
F1 CRITICAL, open MITIGATED for the windowvmbr0 static 192.168.0.162/24; agent up. Durable fix still open → R-50 (island-bridge control plane, spike-first). Still depends on the site's subnet being 192.168.0.0/24 and .162 free
F2 HIGH, open CLOSED — whole-guest backup ran (06:42:52Z) + manual app-data run completed (T = 06:46:10Z)
F3 HIGH, open FIXED in controller v0.149.0, live-verified on the dashboard
Immich ROOT-CAUSED, down RESOLVED — healthy, photos.demo-felhom.eu → 200
F4 MEDIUM, open Roadmap R-51 (constraint recorded: must NOT simply fold unhealthy into down)
F5 MEDIUM, open Roadmap R-52 (includes root-causing the unless-stopped non-resurrection)
F6 LOW, open Already tracked — it is R-39's remaining item (b), the agent's PBS .pw read path. No duplicate item minted
F7 LOW, open Roadmap R-53

Capability map: added a PARTIAL row — "Box survives a site/network change (relocation, different subnet, DHCP re-lease) with the control plane intact" — citing this audit and pointing at R-50/R-51/R-52.

R.7 Not done / out of scope

No agent code, guest bootstrap config, golden state, nftables, tailscale, router, guest-9201 network, or resolv.conf was touched. No restores or drills. F4/F5/F6/F7 diagnosed and ranked only.

R.8 Return-day checklist — amendments

Supersedes items 1 and 4 above:

1a. Decide whether static .162 stays when the box comes home. Recommendation: keep it. It removes the home router's DHCP/reservation from the control plane's dependency chain, which is the same class of dependency that broke this window. Either way, R-50 is the real fix — the static pin is a crutch, and it will fail at any site whose subnet is not 192.168.0.0/24. 4a. /etc/network/interfaces.bak-2026-07-20 on the host is the pre-flip DHCP config — delete it once the decision in 1a is made and verified. 7. Note the deployed controller is v0.149.0 (was 0.148.0 at audit time).