25 KiB
AUDIT — vacation remote-ops: temp access, connection/health diagnosis, app triage
Date: 2026-07-20 · Status: ACTIVE (remote-site window until ~2026-08-02)
Operator: Claude Code on DooPlex · Target: demo host demo-felhom (N100) + guest 9201
Run class: runbook / diagnosis. Docs-only commits. No code changes, no version bumps, no builds.
Baselines: felhom-controller 29eda5d86e (v0.148.0, = deployed) · felhom-agent d81fddc7b1
(deployed binary 0.90.0) · felhom.eu (audit doc only).
1. Access change summary
felhom-pve was physically relocated to a remote site. Access is via Tailscale.
| Item | Result |
|---|---|
~/.ssh/config on DooPlex |
Already correct — Host felhom-pve → HostName 100.70.170.35. No edit needed. A separate Host felhom-pve-lan → 192.168.0.162 entry preserves the home-LAN path. |
| Host-key identity check | PASS — ed25519 key at 100.70.170.35 is byte-identical to the stored key for 192.168.0.162 (…OOz0BjPXLRCI3Z9C…). Same machine, not a reinstall. |
| Tailnet path | direct, 37.191.56.193:45127 (not DERP-relayed). ICMP 4/4, rtt min/avg/max 34.7 / 36.8 / 38.3 ms. |
| Phase 0.4 alias verification | ssh felhom-pve "pct exec 9201 -- docker ps …" → gitea.dooplex.hu/admin/felhom-controller:0.148.0 | Up 31 minutes (healthy). Full CC hop works. |
The tailscale health check on the host notes "Some peers are advertising routes but
--accept-routes is false" — expected and correct per the travel-only decision recorded in
tailscale-n100-location-independent-2026-07-19.
2. Ground-truth snapshot (Phase A)
Host demo-felhom
vmbr0 192.168.0.147/24 dynamic <-- DHCP, NOT the former static .162
tailscale0 100.70.170.35/32
wg-felhom 10.77.0.2/32
default via 192.168.0.1 dev vmbr0
/etc/resolv.conf: nameserver 192.168.0.1 (no Pi-hole; remote router is the resolver)
timedatectl: 2026-07-20 07:56 CEST / 05:56 UTC — synchronized: yes, NTP active, RTC sane
The remote site uses the same 192.168.0.0/24 subnet as home, but the host now holds .147,
not .162. This single fact is the root cause of §3.1.
Boot history — the box was powered off for ~17h45m in transport:
-1 Sat 2026-07-18 18:17:00 CEST → Sun 2026-07-19 13:40:40 CEST
0 Mon 2026-07-20 07:25:05 CEST → (current)
WireGuard (PBS offsite leg) is healthy: peer 167.233.158.164:443, latest handshake 56 s,
keepalive 25 s, 1.70 MiB rx / 1.00 MiB sx.
Guest 9201
eth0 192.168.0.104/24 dynamic
/etc/resolv.conf: nameserver 192.168.0.1
curl https://gitea.dooplex.hu → 200 (0.196 s) outbound + DNS healthy
DooPlex
tailscale is not a host binary here — it runs as the k3s pod admin-system/tailscale-6d7bdc9c5c-hwtjq
(uptime 6d22h). The peer view from the host shows dooplex 100.107.87.53 … active; direct.
Nothing in this section was changed.
3. Per-symptom evidence and root cause
3.1 "A tárolókezelő ügynök nem elérhető" — ROOT-CAUSED (critical)
The agent service is failed, not merely unreachable. Smoking gun from
journalctl -u felhom-agent:
level=ERROR msg="daemon: exited with error"
err="localapi: bind 192.168.0.162:8443: listen tcp 192.168.0.162:8443:
bind: cannot assign requested address"
felhom-agent.service: Start request repeated too quickly.
Active: failed (Result: exit-code) since Mon 2026-07-20 07:25:51 CEST
The agent's localapi listener is pinned to the literal 192.168.0.162. That address no longer
exists on any interface (the host DHCP'd .147), so bind() fails and the daemon exits ~1.1 s after
start. systemd retried 4× and gave up. The agent has therefore been down since the first boot at
the remote site — it has never run there.
Controller side, the matching error class is no route to host (not a TLS-pin or auth failure —
those would be STOP conditions of a different kind):
[ERROR] [quiesce] cycle error: check due: agentapi: GET /backup/due:
Get "https://192.168.0.162:8443/backup/due": dial tcp 192.168.0.162:8443: connect: no route to host
[WARN] local-api: GET /storage failed (… no route to host) — channel not verified
So both ends are independently pinned to 192.168.0.162: the agent binds it, the controller
dials it. Nothing is listening (ss -tlnp | grep felhom → empty). This is precisely the OPEN item
recorded on 2026-07-19 — "control plane pinned to 192.168.0.162 (agent listen_addr + guest
bootstrap.json) → not portable off-LAN" — now realized in production.
Allowed action taken (§3.2 of the runbook): systemctl reset-failed && systemctl restart felhom-agent.
It failed identically within 6 s:
Jul 20 08:00:02 … err="localapi: bind 192.168.0.162:8443: … cannot assign requested address"
Jul 20 08:00:02 … felhom-agent.service: Failed with result 'exit-code'.
STOP condition hit (runbook §3.4, "agent endpoint stale because the host IP changed"). The fix
requires editing /etc/felhom-agent/agent.json (listen_addr) and the guest's
bootstrap-managed endpoint — both forbidden in this run. Escalated as F1.
Note for the fix: re-pinning to the literal
192.168.0.147would work today but is a DHCP lease away from breaking again, and the agent pins its served leaf cert, whose SAN set must cover whatever address the controller dials. Options are laid out in F1.
Blast radius while F1 stands: every agent-backed function is dead — storage management (Tárhely), PBS whole-guest backups, the quiesce cycle, restore-tests, and the DR path. Guest-local functions (dashboard, apps, app-data/DB backup cache, Samba, cloudflared) are unaffected.
3.2 Calibre-Web Automated stopped — RESOLVED
exit=128 finished=2026-07-19T11:40:15Z policy=unless-stopped restarts=0
Container log ends with a clean application shutdown (webserver stop (restart=False),
Performing shutdown of Calibre-Web Automated) at 13:40:15 CEST — ~25 s before the host powered
off for transport. Not a crash: an orderly stop caught by the shutdown, which then never restarted.
Allowed action taken: started via the controller's real endpoint (the exact path the UI uses —
session login, <meta name="csrf-token"> read from the page, X-CSRF-Token header):
POST /api/stacks/calibre-web/start → {"ok":true,"message":"Stack calibre-web start completed"}
docker ps → calibre-web | Up 20 seconds (healthy)
Verified healthy. Stack state now reads running.
3.3 Immich "Nem egészséges" + "URL nem elérhető" — ROOT-CAUSED, NOT fixed (per runbook D.3)
The two flags have one cause: immich-server is not running at all.
immich-server Exited (143) 18 hours ago policy=unless-stopped restarts=0
finished=2026-07-19T11:40:15Z health=unhealthy
immich-postgres Up 31 minutes (healthy)
immich-machine-learning Up 31 minutes (healthy)
immich-redis Up 31 minutes (healthy)
Exit 143 = SIGTERM at 11:40:15Z — the same second as Calibre-Web, i.e. the same pre-transport
shutdown. Its final log lines are a fully healthy server (Immich Server is listening … [v3.0.3],
ML server healthy, websocket clients connected) right up to the shutdown. This is not leftover
state from the 2026-07-19 offsite-restore test (R-43/R-44) — hypothesis (a) is disproved by the
logs: the server ran normally for ~42 min after that work, and there is no migration or DB error.
The URL probe is likewise not a remote-path/latency artifact — hypothesis (c) is disproved:
curl https://immich.demo-felhom.eu → 404 in 0.239 s (traefik has no live backend)
curl https://gitea.dooplex.hu → 200 in 0.196 s (same path, fine)
Fast 404, not a timeout. The public CF path is healthy; there is simply nothing behind the route. Left down deliberately per runbook D.3. One-line fix available under GO — see F5.
3.4 "Utolsó mentés: Még nem futott" — ROOT-CAUSED as a display bug (independent of F1)
This is hypothesis (c), and it is provable from code rather than inferred.
dashboardHandler (controller/internal/web/handlers.go:134) populates the backup card with:
data["BackupEnabled"] = s.cfg.Backup.Enabled // :172
data["DBDumpStatus"] = fullStatus.LastDBDump // :175
data["BackupRunning"] = fullStatus.Running // :177
data["BackupMaxAgeHours"] = … // :178
It never sets data["BackupStatus"] — and no other code path does either (repo-wide, the only
occurrences of BackupStatus are the agent-proxy handler and this template). But
dashboard.html:116 gates the whole "Utolsó mentés" value on it:
{{if .BackupStatus}} … LastRun / "Sikertelen" …
{{else}} <span class="backup-status-none">Még nem futott</span>
{{end}}
.BackupStatus is therefore always nil, the {{if}} branch is dead code, and the dashboard
reports "Még nem futott" permanently on every box regardless of actual backup history. The
adjacent "Adatbázisok: 2 mentve" renders from DBDumpStatus, which is set — which is exactly the
contradiction observed.
Independent corroboration that backups had in fact been running: the hub's own host-reports
consistently carry 1 backups, and hub events on 2026-07-19 include
crossdrive_completed for bookstack / calibre-web / immich / Megosztások at 03:30 and
db_dump_completed at 02:30.
Separately, and not a display artifact: no backup can run at all while F1 stands — the quiesce
cycle fails every 5 min on GET /backup/due. The guest-local backup-cache job still runs fine
(Found 2 DB dump files across drives, Discovered 2 databases, Discovered app data: 3 apps
every 5 min). No manual backup was triggered: the runbook permits one only if the cause was (a) and
the agent is healthy — neither holds.
3.5 Hub report freshness — HEALTHY, no gap
Reports land on the normal 15-min cadence from the remote site, and the staleness state machine recovered by itself on boot:
07:25:42 Event from demo-felhom: controller_started (info) — Controller elindult (0.148.0)
07:25:43 Staleness: demo-felhom down → ok (node_recovered)
07:27:07 Event: app_start_failed — Telepített alkalmazás nem fut: Calibre-Web Automated
07:27:40 Event: agent_channel_unknown (warning) — Controller→agent channel down: … no route to host
07:27:41 Operator email sent for demo-felhom/agent_channel_unknown
07:40:37 / 07:55:37 Received report from demo-felhom (7603 / 6232 bytes)
Hub telemetry is location-independent (it rides the public path), and the alerting chain worked end to end — the agent-channel failure produced an operator email within 33 s of the boot grace expiring. The transport gap appears exactly as it should: last report 2026-07-19 11:40:22Z, next 2026-07-20 05:25 UTC.
4. Ranked findings
| # | Sev | Finding | Evidence | Proposed fix | Viktor GO? |
|---|---|---|---|---|---|
| F1 | CRITICAL | Agent localapi binds the literal 192.168.0.162; host is .147 at the remote site → bind: cannot assign requested address → service failed since first remote boot. Controller dials the same literal. All agent-backed features (storage, PBS backup, quiesce, restore-test, DR) are dead for the whole 2-week window. |
§3.1 | Bind 0.0.0.0:8443 (or the vmbr0 address resolved at start-up) instead of a literal, and give the controller a stable dial target; ensure the served leaf's SAN covers it. Short-term unblock: set listen_addr to 192.168.0.147 + reissue/verify the pin + update the guest endpoint — but this re-breaks on the next DHCP lease. |
YES — config and code; touches bootstrap-managed guest state |
| F2 | HIGH | No PBS/whole-guest backup can run while F1 stands → a ~2-week data-protection gap on the demo box. Quiesce errors every 5 min. | §3.1, §3.4 | Resolve F1; then verify a backup completes. | Follows F1 |
| F3 | HIGH | Dashboard permanently shows "Utolsó mentés: Még nem futott" on every box: dashboardHandler never sets BackupStatus, so the template's {{if .BackupStatus}} branch is unreachable dead code. Customer-facing misinformation about backup state. |
§3.4 (handlers.go:134-178, dashboard.html:116-131) |
Populate data["BackupStatus"] from fullStatus in dashboardHandler, or drop the dead branch and render from DBDumpStatus/GuestBackup. Needs a red-proof test (the current code has none covering this key). |
YES — code |
| F4 | MEDIUM | Multi-container apps under-alert. IsDownState (stacks/manager.go:46) counts only stopped/exited; it deliberately excludes unhealthy. Immich's stack aggregates to unhealthy because 3 of 4 containers are up — so with its primary container fully exited for 18 h and the app 100 % unreachable, it produced no dead-app banner and no app_start_failed hub event. Calibre-Web (single container) alerted correctly within 90 s. |
§3.3, live /api/stacks → immich unhealthy; hub log has an event for Calibre-Web only |
Treat "primary/named container down" as down for a multi-container stack, or add an unhealthy-for-N-cycles escalation. |
YES — code |
| F5 | MEDIUM | Two app containers were left Exited by the pre-transport shutdown and were not restarted on boot despite restart: unless-stopped; the controller reports deployed-but-stopped apps (30 s deadapp-check) but never starts them. Immich has now been down 18 h+ and stays down. |
§3.2, §3.3, cmd/controller/main.go:435 |
Start immich via the same real endpoint used for Calibre-Web (immediate, no code). Then consider a bounded boot-time reconciliation (start deployed-but-stopped apps once, N attempts, with the existing boot grace). |
Start = YES (runbook D.3 forbade it here); reconciler = code, YES |
| F6 | LOW | Agent logs pbs: cannot read token secret; skipping datastore storage=felhom-pbs err="open /etc/pve/priv/storage/felhom-pbs.pw: permission denied" on every start (pre-existing, R-39-adjacent — the non-root agent cannot read the PVE priv file). |
§3.1 journal | Confirm whether the wrapper/sudoers path is meant to cover this read; fold into the R-39 follow-up. | YES |
| F7 | LOW | Incidental defect found while tracing CSRF: app_export.html:93 builds var domain = '<subdomain>.{{$.CSRFToken}}' — the session CSRF token is substituted where the domain belongs. Produces a broken "open in browser" link and puts a CSRF token into a URL. |
controller/internal/web/templates/app_export.html:93 |
Replace {{$.CSRFToken}} with the domain value (cf. the neighbouring correct csrfH() helper). |
YES — code |
Healthy / no action: tailnet path direct at 37 ms · host clock NTP-synced (no transport RTC drift) · WireGuard PBS leg handshaking at 56 s · guest outbound + DNS · hub reports on cadence and staleness auto-recovered · operator alert email fired correctly · 6 apps + Samba + traefik + cloudflared running · USB HDD mounted and visible.
5. Actions taken
| Action | Authority | Outcome |
|---|---|---|
systemctl reset-failed && systemctl restart felhom-agent |
runbook §3.2 (one allowed restart, after evidence) | Failed identically in 6 s — same bind error. Confirms F1 is deterministic config, not a transient. |
POST /api/stacks/calibre-web/start via the controller's real session+CSRF path |
runbook §3.2 (allowed, after evidence) | Success — Up (healthy), stack state running. |
Nothing else was mutated. The ssh alias needed no edit (already on the tailnet IP).
6. Explicitly NOT fixed / not touched
- F1 — no edit to
/etc/felhom-agent/agent.json, the guestbootstrap.json, or any bootstrap/golden-managed state. STOP condition §3.4 hit and honoured. - Immich — left down per runbook D.3 (diagnose only), despite the fix being a one-call start.
- F3, F4, F7 — code changes, out of scope for this run.
- No network/nftables/tailscale/route changes on either machine. No manual backup triggered.
No version bumps, image builds, deploys, or
REPORT.mdoverwrites. No containers or volumes deleted. No secrets written to any committed file.
7. Return-day checklist (~2026-08-02)
- If F1 was fixed with a literal
.147, revert it to the home-LAN value or (preferred) confirm the address-independent fix still binds correctly back on the home LAN. ~/.ssh/configneeds no revert —Host felhom-pveon the tailnet IP works from both locations;felhom-pve-lanremains as the direct-LAN path. Re-verify the.162/LAN path once the host is home and holds its expected address.- Delete the TEMPORARY blocks from
felhom-controller/CLAUDE.mdandfelhom-agent/CLAUDE.md. - Confirm the host reacquires its intended home address (it is now on DHCP — the former static
.162config did not survive/apply; decide whether to restore a static or keep the reservation). - Re-verify: agent
active (running), dashboard banner clear, Tárhely renders, a PBS backup completes, Immich healthy and reachable. - Flip this document's status from ACTIVE to CLOSED.
Remediation — 2026-07-20
Executed the same day as the audit, on the relocated host. Mutations: one network file, one service start, one app start, two backup runs (one of them the product's own scheduler), one code line.
R.1 Static-IP flip (F1 window mitigation)
State found had already moved. By the time remediation started, the router's reservation had
taken effect and vmbr0 held 192.168.0.162 again — but via DHCP (dhclient.vmbr0.leases
stamped 08:13 CEST). So the address existed; nothing had retried the agent since systemd gave up.
The 1.1 duplicate-address check needed interpretation, not blind obedience: ping -c1 192.168.0.162
answered, because the host was answering itself. The genuine test is from another node —
ARP for .162 seen from guest 9201 resolved to 68:1d:ef:5d:a6:64, which is this host's vmbr0
MAC. No duplicate; the STOP did not apply.
Interfaces diff (.bak-2026-07-20 → new) — one stanza, nothing else touched:
auto vmbr0
-iface vmbr0 inet dhcp
+iface vmbr0 inet static
+ address 192.168.0.162/24
+ gateway 192.168.0.1
bridge-ports enp1s0
bridge-stp off
bridge-fd 0
The gateway was taken from the live default route (192.168.0.1), not from the spec. Applied
detached: setsid ifreload -a → log contained only Killed old client process (the vmbr0 dhclient).
Post-flip: inet 192.168.0.162/24 scope global vmbr0 (no dynamic flag), default via 192.168.0.1 … proto kernel onlink, curl https://gitea.dooplex.hu → 200, tailscale peer still
active; direct. Because the address was unchanged, connectivity never blipped.
R.2 Agent recovery (Scenario A)
systemctl reset-failed && systemctl start felhom-agent — up at 08:40:11 CEST:
active
LISTEN 192.168.0.162:8443 users:(("felhom-agent",pid=72859))
level=INFO msg="local-api server listening" addr=192.168.0.162:8443
wg-felhom latest handshake: 1 minute, 16 seconds ago
Controller side, agentapi errors stopped entirely (no no route to host, no
channel not verified in the following 5 min) — so this is not the "agent active but controller
still erroring" wrong outcome the scenario warns about. On an authenticated fetch of the real
dashboard the red „A tárolókezelő ügynök nem elérhető" banner is gone; the only
alert-warning strings left on the page are two modal "this cannot be undone" confirmations.
Method note for anyone repeating this: grep the fetched page with ASCII-only patterns. Accented patterns sent through the ssh→
pct exec→shell layers get mangled and return a false0, which reads exactly like "the banner is gone" when it is not. The page content itself is fine — it is the pattern that gets corrupted.
R.3 F2 — the protection gap closed itself, then a manual run confirmed it
Restoring the agent was sufficient: the quiesce loop noticed the overdue whole-guest backup and ran it unprompted, which is the strongest possible evidence the seam is healthy end to end:
06:40:36Z [quiesce] backup due — quiescing 3 stack(s): [bookstack calibre-web immich]
06:40:45Z [quiesce] backup job backup-9201-1784529645204076480 started — polling to completion
06:40:55Z [quiesce] … snapshotted — resuming app early (8B.2)
06:42:52Z [quiesce] … done
A manual app-data run then followed via the real UI endpoint (POST /api/backup/run, authenticated
session + X-CSRF-Token read from the page's <meta> — the same call the „Teljes mentés" button
makes), started 06:45:26Z:
06:45:27Z DB dump: bookstack-db → bookstack-mariadb.sql (55.9 KB, 228ms, 41 tables)
06:45:28Z DB dump: immich-postgres → immich-postgres.sql (49.9 MB, 1.28s, 66 tables)
06:45:44Z Volume dump: calibre-web/…_config → 360.0 KB
06:45:55Z Volume dump: immich/…_ml_cache → 785.5 MB
06:45:58Z Volume dump: immich/…_postgres_data → 241.5 MB
06:45:59Z Volume dump: immich/…_redis_data → 6.6 MB
06:46:10Z App-data backup completed: 2 databases (50.0 MB total), 3 volume dump(s) (43.348s)
T = 2026-07-20 06:46:10Z (run completion; newest dump artifact 06:45:28Z). Every container the
run stopped for a safe volume dump was restarted — all 12 containers verified Up (healthy)
afterwards.
R.4 Immich (Scenario: app triage)
Started via the controller's real endpoint (POST /api/stacks/immich/start) → all four containers
Up (healthy).
Correction to §3.3 of the audit. That section reported the URL probe returning 404 and treated
it as a live symptom. The probe used the wrong hostname: the traefik router rule is
Host(`photos.demo-felhom.eu`), not immich.demo-felhom.eu. The container-down diagnosis stands
(and the 404 was genuinely traefik having no backend), but the specific probe URL in §3.3 was
invalid. Post-start, on the correct host: https://photos.demo-felhom.eu → 200 in 0.28 s.
R.5 F3 — dashboard backup card (controller v0.149.0)
Handler-side one-liner, no template change: data["BackupStatus"] = fullStatus.LastDBDump in
dashboardHandler. Deployed to guest 9201 and verified on the real, authenticated dashboard:
<span class="backup-label">Utolsó mentés:</span>
<span class="backup-value"><span>2026-07-20 06:45</span></span>
<span class="backup-label">Adatbázisok:</span>
<span class="backup-value">3 mentve</span>
Was „Még nem futott" on the identical page before the deploy. (The controller restart cleared the in-memory run state, so the card renders from the dumps on disk — newest artifact 06:45 — which is the honest value and exactly what the fresh-box branch must NOT fabricate.)
R.6 Per-finding status after remediation
| # | Was | Now |
|---|---|---|
| F1 | CRITICAL, open | MITIGATED for the window — vmbr0 static 192.168.0.162/24; agent up. Durable fix still open → R-50 (island-bridge control plane, spike-first). Still depends on the site's subnet being 192.168.0.0/24 and .162 free |
| F2 | HIGH, open | CLOSED — whole-guest backup ran (06:42:52Z) + manual app-data run completed (T = 06:46:10Z) |
| F3 | HIGH, open | FIXED in controller v0.149.0, live-verified on the dashboard |
| Immich | ROOT-CAUSED, down | RESOLVED — healthy, photos.demo-felhom.eu → 200 |
| F4 | MEDIUM, open | Roadmap R-51 (constraint recorded: must NOT simply fold unhealthy into down) |
| F5 | MEDIUM, open | Roadmap R-52 (includes root-causing the unless-stopped non-resurrection) |
| F6 | LOW, open | Already tracked — it is R-39's remaining item (b), the agent's PBS .pw read path. No duplicate item minted |
| F7 | LOW, open | Roadmap R-53 |
Capability map: added a PARTIAL row — "Box survives a site/network change (relocation, different subnet, DHCP re-lease) with the control plane intact" — citing this audit and pointing at R-50/R-51/R-52.
R.7 Not done / out of scope
No agent code, guest bootstrap config, golden state, nftables, tailscale, router, guest-9201
network, or resolv.conf was touched. No restores or drills. F4/F5/F6/F7 diagnosed and ranked only.
R.8 Return-day checklist — amendments
Supersedes items 1 and 4 above:
1a. Decide whether static .162 stays when the box comes home. Recommendation: keep it. It
removes the home router's DHCP/reservation from the control plane's dependency chain, which is
the same class of dependency that broke this window. Either way, R-50 is the real fix — the
static pin is a crutch, and it will fail at any site whose subnet is not 192.168.0.0/24.
4a. /etc/network/interfaces.bak-2026-07-20 on the host is the pre-flip DHCP config — delete it
once the decision in 1a is made and verified.
7. Note the deployed controller is v0.149.0 (was 0.148.0 at audit time).