Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Nhk3eBHT8Mg5L8c2aj57aU
17 KiB
AUDIT — vacation remote-ops: temp access, connection/health diagnosis, app triage
Date: 2026-07-20 · Status: ACTIVE (remote-site window until ~2026-08-02)
Operator: Claude Code on DooPlex · Target: demo host demo-felhom (N100) + guest 9201
Run class: runbook / diagnosis. Docs-only commits. No code changes, no version bumps, no builds.
Baselines: felhom-controller 29eda5d86e (v0.148.0, = deployed) · felhom-agent d81fddc7b1
(deployed binary 0.90.0) · felhom.eu (audit doc only).
1. Access change summary
felhom-pve was physically relocated to a remote site. Access is via Tailscale.
| Item | Result |
|---|---|
~/.ssh/config on DooPlex |
Already correct — Host felhom-pve → HostName 100.70.170.35. No edit needed. A separate Host felhom-pve-lan → 192.168.0.162 entry preserves the home-LAN path. |
| Host-key identity check | PASS — ed25519 key at 100.70.170.35 is byte-identical to the stored key for 192.168.0.162 (…OOz0BjPXLRCI3Z9C…). Same machine, not a reinstall. |
| Tailnet path | direct, 37.191.56.193:45127 (not DERP-relayed). ICMP 4/4, rtt min/avg/max 34.7 / 36.8 / 38.3 ms. |
| Phase 0.4 alias verification | ssh felhom-pve "pct exec 9201 -- docker ps …" → gitea.dooplex.hu/admin/felhom-controller:0.148.0 | Up 31 minutes (healthy). Full CC hop works. |
The tailscale health check on the host notes "Some peers are advertising routes but
--accept-routes is false" — expected and correct per the travel-only decision recorded in
tailscale-n100-location-independent-2026-07-19.
2. Ground-truth snapshot (Phase A)
Host demo-felhom
vmbr0 192.168.0.147/24 dynamic <-- DHCP, NOT the former static .162
tailscale0 100.70.170.35/32
wg-felhom 10.77.0.2/32
default via 192.168.0.1 dev vmbr0
/etc/resolv.conf: nameserver 192.168.0.1 (no Pi-hole; remote router is the resolver)
timedatectl: 2026-07-20 07:56 CEST / 05:56 UTC — synchronized: yes, NTP active, RTC sane
The remote site uses the same 192.168.0.0/24 subnet as home, but the host now holds .147,
not .162. This single fact is the root cause of §3.1.
Boot history — the box was powered off for ~17h45m in transport:
-1 Sat 2026-07-18 18:17:00 CEST → Sun 2026-07-19 13:40:40 CEST
0 Mon 2026-07-20 07:25:05 CEST → (current)
WireGuard (PBS offsite leg) is healthy: peer 167.233.158.164:443, latest handshake 56 s,
keepalive 25 s, 1.70 MiB rx / 1.00 MiB sx.
Guest 9201
eth0 192.168.0.104/24 dynamic
/etc/resolv.conf: nameserver 192.168.0.1
curl https://gitea.dooplex.hu → 200 (0.196 s) outbound + DNS healthy
DooPlex
tailscale is not a host binary here — it runs as the k3s pod admin-system/tailscale-6d7bdc9c5c-hwtjq
(uptime 6d22h). The peer view from the host shows dooplex 100.107.87.53 … active; direct.
Nothing in this section was changed.
3. Per-symptom evidence and root cause
3.1 "A tárolókezelő ügynök nem elérhető" — ROOT-CAUSED (critical)
The agent service is failed, not merely unreachable. Smoking gun from
journalctl -u felhom-agent:
level=ERROR msg="daemon: exited with error"
err="localapi: bind 192.168.0.162:8443: listen tcp 192.168.0.162:8443:
bind: cannot assign requested address"
felhom-agent.service: Start request repeated too quickly.
Active: failed (Result: exit-code) since Mon 2026-07-20 07:25:51 CEST
The agent's localapi listener is pinned to the literal 192.168.0.162. That address no longer
exists on any interface (the host DHCP'd .147), so bind() fails and the daemon exits ~1.1 s after
start. systemd retried 4× and gave up. The agent has therefore been down since the first boot at
the remote site — it has never run there.
Controller side, the matching error class is no route to host (not a TLS-pin or auth failure —
those would be STOP conditions of a different kind):
[ERROR] [quiesce] cycle error: check due: agentapi: GET /backup/due:
Get "https://192.168.0.162:8443/backup/due": dial tcp 192.168.0.162:8443: connect: no route to host
[WARN] local-api: GET /storage failed (… no route to host) — channel not verified
So both ends are independently pinned to 192.168.0.162: the agent binds it, the controller
dials it. Nothing is listening (ss -tlnp | grep felhom → empty). This is precisely the OPEN item
recorded on 2026-07-19 — "control plane pinned to 192.168.0.162 (agent listen_addr + guest
bootstrap.json) → not portable off-LAN" — now realized in production.
Allowed action taken (§3.2 of the runbook): systemctl reset-failed && systemctl restart felhom-agent.
It failed identically within 6 s:
Jul 20 08:00:02 … err="localapi: bind 192.168.0.162:8443: … cannot assign requested address"
Jul 20 08:00:02 … felhom-agent.service: Failed with result 'exit-code'.
STOP condition hit (runbook §3.4, "agent endpoint stale because the host IP changed"). The fix
requires editing /etc/felhom-agent/agent.json (listen_addr) and the guest's
bootstrap-managed endpoint — both forbidden in this run. Escalated as F1.
Note for the fix: re-pinning to the literal
192.168.0.147would work today but is a DHCP lease away from breaking again, and the agent pins its served leaf cert, whose SAN set must cover whatever address the controller dials. Options are laid out in F1.
Blast radius while F1 stands: every agent-backed function is dead — storage management (Tárhely), PBS whole-guest backups, the quiesce cycle, restore-tests, and the DR path. Guest-local functions (dashboard, apps, app-data/DB backup cache, Samba, cloudflared) are unaffected.
3.2 Calibre-Web Automated stopped — RESOLVED
exit=128 finished=2026-07-19T11:40:15Z policy=unless-stopped restarts=0
Container log ends with a clean application shutdown (webserver stop (restart=False),
Performing shutdown of Calibre-Web Automated) at 13:40:15 CEST — ~25 s before the host powered
off for transport. Not a crash: an orderly stop caught by the shutdown, which then never restarted.
Allowed action taken: started via the controller's real endpoint (the exact path the UI uses —
session login, <meta name="csrf-token"> read from the page, X-CSRF-Token header):
POST /api/stacks/calibre-web/start → {"ok":true,"message":"Stack calibre-web start completed"}
docker ps → calibre-web | Up 20 seconds (healthy)
Verified healthy. Stack state now reads running.
3.3 Immich "Nem egészséges" + "URL nem elérhető" — ROOT-CAUSED, NOT fixed (per runbook D.3)
The two flags have one cause: immich-server is not running at all.
immich-server Exited (143) 18 hours ago policy=unless-stopped restarts=0
finished=2026-07-19T11:40:15Z health=unhealthy
immich-postgres Up 31 minutes (healthy)
immich-machine-learning Up 31 minutes (healthy)
immich-redis Up 31 minutes (healthy)
Exit 143 = SIGTERM at 11:40:15Z — the same second as Calibre-Web, i.e. the same pre-transport
shutdown. Its final log lines are a fully healthy server (Immich Server is listening … [v3.0.3],
ML server healthy, websocket clients connected) right up to the shutdown. This is not leftover
state from the 2026-07-19 offsite-restore test (R-43/R-44) — hypothesis (a) is disproved by the
logs: the server ran normally for ~42 min after that work, and there is no migration or DB error.
The URL probe is likewise not a remote-path/latency artifact — hypothesis (c) is disproved:
curl https://immich.demo-felhom.eu → 404 in 0.239 s (traefik has no live backend)
curl https://gitea.dooplex.hu → 200 in 0.196 s (same path, fine)
Fast 404, not a timeout. The public CF path is healthy; there is simply nothing behind the route. Left down deliberately per runbook D.3. One-line fix available under GO — see F5.
3.4 "Utolsó mentés: Még nem futott" — ROOT-CAUSED as a display bug (independent of F1)
This is hypothesis (c), and it is provable from code rather than inferred.
dashboardHandler (controller/internal/web/handlers.go:134) populates the backup card with:
data["BackupEnabled"] = s.cfg.Backup.Enabled // :172
data["DBDumpStatus"] = fullStatus.LastDBDump // :175
data["BackupRunning"] = fullStatus.Running // :177
data["BackupMaxAgeHours"] = … // :178
It never sets data["BackupStatus"] — and no other code path does either (repo-wide, the only
occurrences of BackupStatus are the agent-proxy handler and this template). But
dashboard.html:116 gates the whole "Utolsó mentés" value on it:
{{if .BackupStatus}} … LastRun / "Sikertelen" …
{{else}} <span class="backup-status-none">Még nem futott</span>
{{end}}
.BackupStatus is therefore always nil, the {{if}} branch is dead code, and the dashboard
reports "Még nem futott" permanently on every box regardless of actual backup history. The
adjacent "Adatbázisok: 2 mentve" renders from DBDumpStatus, which is set — which is exactly the
contradiction observed.
Independent corroboration that backups had in fact been running: the hub's own host-reports
consistently carry 1 backups, and hub events on 2026-07-19 include
crossdrive_completed for bookstack / calibre-web / immich / Megosztások at 03:30 and
db_dump_completed at 02:30.
Separately, and not a display artifact: no backup can run at all while F1 stands — the quiesce
cycle fails every 5 min on GET /backup/due. The guest-local backup-cache job still runs fine
(Found 2 DB dump files across drives, Discovered 2 databases, Discovered app data: 3 apps
every 5 min). No manual backup was triggered: the runbook permits one only if the cause was (a) and
the agent is healthy — neither holds.
3.5 Hub report freshness — HEALTHY, no gap
Reports land on the normal 15-min cadence from the remote site, and the staleness state machine recovered by itself on boot:
07:25:42 Event from demo-felhom: controller_started (info) — Controller elindult (0.148.0)
07:25:43 Staleness: demo-felhom down → ok (node_recovered)
07:27:07 Event: app_start_failed — Telepített alkalmazás nem fut: Calibre-Web Automated
07:27:40 Event: agent_channel_unknown (warning) — Controller→agent channel down: … no route to host
07:27:41 Operator email sent for demo-felhom/agent_channel_unknown
07:40:37 / 07:55:37 Received report from demo-felhom (7603 / 6232 bytes)
Hub telemetry is location-independent (it rides the public path), and the alerting chain worked end to end — the agent-channel failure produced an operator email within 33 s of the boot grace expiring. The transport gap appears exactly as it should: last report 2026-07-19 11:40:22Z, next 2026-07-20 05:25 UTC.
4. Ranked findings
| # | Sev | Finding | Evidence | Proposed fix | Viktor GO? |
|---|---|---|---|---|---|
| F1 | CRITICAL | Agent localapi binds the literal 192.168.0.162; host is .147 at the remote site → bind: cannot assign requested address → service failed since first remote boot. Controller dials the same literal. All agent-backed features (storage, PBS backup, quiesce, restore-test, DR) are dead for the whole 2-week window. |
§3.1 | Bind 0.0.0.0:8443 (or the vmbr0 address resolved at start-up) instead of a literal, and give the controller a stable dial target; ensure the served leaf's SAN covers it. Short-term unblock: set listen_addr to 192.168.0.147 + reissue/verify the pin + update the guest endpoint — but this re-breaks on the next DHCP lease. |
YES — config and code; touches bootstrap-managed guest state |
| F2 | HIGH | No PBS/whole-guest backup can run while F1 stands → a ~2-week data-protection gap on the demo box. Quiesce errors every 5 min. | §3.1, §3.4 | Resolve F1; then verify a backup completes. | Follows F1 |
| F3 | HIGH | Dashboard permanently shows "Utolsó mentés: Még nem futott" on every box: dashboardHandler never sets BackupStatus, so the template's {{if .BackupStatus}} branch is unreachable dead code. Customer-facing misinformation about backup state. |
§3.4 (handlers.go:134-178, dashboard.html:116-131) |
Populate data["BackupStatus"] from fullStatus in dashboardHandler, or drop the dead branch and render from DBDumpStatus/GuestBackup. Needs a red-proof test (the current code has none covering this key). |
YES — code |
| F4 | MEDIUM | Multi-container apps under-alert. IsDownState (stacks/manager.go:46) counts only stopped/exited; it deliberately excludes unhealthy. Immich's stack aggregates to unhealthy because 3 of 4 containers are up — so with its primary container fully exited for 18 h and the app 100 % unreachable, it produced no dead-app banner and no app_start_failed hub event. Calibre-Web (single container) alerted correctly within 90 s. |
§3.3, live /api/stacks → immich unhealthy; hub log has an event for Calibre-Web only |
Treat "primary/named container down" as down for a multi-container stack, or add an unhealthy-for-N-cycles escalation. |
YES — code |
| F5 | MEDIUM | Two app containers were left Exited by the pre-transport shutdown and were not restarted on boot despite restart: unless-stopped; the controller reports deployed-but-stopped apps (30 s deadapp-check) but never starts them. Immich has now been down 18 h+ and stays down. |
§3.2, §3.3, cmd/controller/main.go:435 |
Start immich via the same real endpoint used for Calibre-Web (immediate, no code). Then consider a bounded boot-time reconciliation (start deployed-but-stopped apps once, N attempts, with the existing boot grace). |
Start = YES (runbook D.3 forbade it here); reconciler = code, YES |
| F6 | LOW | Agent logs pbs: cannot read token secret; skipping datastore storage=felhom-pbs err="open /etc/pve/priv/storage/felhom-pbs.pw: permission denied" on every start (pre-existing, R-39-adjacent — the non-root agent cannot read the PVE priv file). |
§3.1 journal | Confirm whether the wrapper/sudoers path is meant to cover this read; fold into the R-39 follow-up. | YES |
| F7 | LOW | Incidental defect found while tracing CSRF: app_export.html:93 builds var domain = '<subdomain>.{{$.CSRFToken}}' — the session CSRF token is substituted where the domain belongs. Produces a broken "open in browser" link and puts a CSRF token into a URL. |
controller/internal/web/templates/app_export.html:93 |
Replace {{$.CSRFToken}} with the domain value (cf. the neighbouring correct csrfH() helper). |
YES — code |
Healthy / no action: tailnet path direct at 37 ms · host clock NTP-synced (no transport RTC drift) · WireGuard PBS leg handshaking at 56 s · guest outbound + DNS · hub reports on cadence and staleness auto-recovered · operator alert email fired correctly · 6 apps + Samba + traefik + cloudflared running · USB HDD mounted and visible.
5. Actions taken
| Action | Authority | Outcome |
|---|---|---|
systemctl reset-failed && systemctl restart felhom-agent |
runbook §3.2 (one allowed restart, after evidence) | Failed identically in 6 s — same bind error. Confirms F1 is deterministic config, not a transient. |
POST /api/stacks/calibre-web/start via the controller's real session+CSRF path |
runbook §3.2 (allowed, after evidence) | Success — Up (healthy), stack state running. |
Nothing else was mutated. The ssh alias needed no edit (already on the tailnet IP).
6. Explicitly NOT fixed / not touched
- F1 — no edit to
/etc/felhom-agent/agent.json, the guestbootstrap.json, or any bootstrap/golden-managed state. STOP condition §3.4 hit and honoured. - Immich — left down per runbook D.3 (diagnose only), despite the fix being a one-call start.
- F3, F4, F7 — code changes, out of scope for this run.
- No network/nftables/tailscale/route changes on either machine. No manual backup triggered.
No version bumps, image builds, deploys, or
REPORT.mdoverwrites. No containers or volumes deleted. No secrets written to any committed file.
7. Return-day checklist (~2026-08-02)
- If F1 was fixed with a literal
.147, revert it to the home-LAN value or (preferred) confirm the address-independent fix still binds correctly back on the home LAN. ~/.ssh/configneeds no revert —Host felhom-pveon the tailnet IP works from both locations;felhom-pve-lanremains as the direct-LAN path. Re-verify the.162/LAN path once the host is home and holds its expected address.- Delete the TEMPORARY blocks from
felhom-controller/CLAUDE.mdandfelhom-agent/CLAUDE.md. - Confirm the host reacquires its intended home address (it is now on DHCP — the former static
.162config did not survive/apply; decide whether to restore a static or keep the reservation). - Re-verify: agent
active (running), dashboard banner clear, Tárhely renders, a PBS backup completes, Immich healthy and reachable. - Flip this document's status from ACTIVE to CLOSED.