docs(audits): vacation remote-ops — access change + demo-node health diagnosis
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Nhk3eBHT8Mg5L8c2aj57aU
This commit is contained in:
@@ -0,0 +1,287 @@
|
||||
# AUDIT — vacation remote-ops: temp access, connection/health diagnosis, app triage
|
||||
|
||||
**Date:** 2026-07-20 · **Status:** ACTIVE (remote-site window until ~2026-08-02)
|
||||
**Operator:** Claude Code on DooPlex · **Target:** demo host `demo-felhom` (N100) + guest 9201
|
||||
**Run class:** runbook / diagnosis. Docs-only commits. No code changes, no version bumps, no builds.
|
||||
|
||||
Baselines: felhom-controller `29eda5d86e` (v0.148.0, = deployed) · felhom-agent `d81fddc7b1`
|
||||
(deployed binary **0.90.0**) · felhom.eu (audit doc only).
|
||||
|
||||
---
|
||||
|
||||
## 1. Access change summary
|
||||
|
||||
felhom-pve was physically relocated to a remote site. Access is via Tailscale.
|
||||
|
||||
| Item | Result |
|
||||
|---|---|
|
||||
| `~/.ssh/config` on DooPlex | **Already correct** — `Host felhom-pve` → `HostName 100.70.170.35`. No edit needed. A separate `Host felhom-pve-lan` → `192.168.0.162` entry preserves the home-LAN path. |
|
||||
| Host-key identity check | **PASS** — ed25519 key at `100.70.170.35` is byte-identical to the stored key for `192.168.0.162` (`…OOz0BjPXLRCI3Z9C…`). Same machine, not a reinstall. |
|
||||
| Tailnet path | **direct**, `37.191.56.193:45127` (not DERP-relayed). ICMP 4/4, rtt min/avg/max **34.7 / 36.8 / 38.3 ms**. |
|
||||
| Phase 0.4 alias verification | `ssh felhom-pve "pct exec 9201 -- docker ps …"` → `gitea.dooplex.hu/admin/felhom-controller:0.148.0 \| Up 31 minutes (healthy)`. Full CC hop works. |
|
||||
|
||||
The tailscale health check on the host notes *"Some peers are advertising routes but
|
||||
`--accept-routes` is false"* — expected and correct per the travel-only decision recorded in
|
||||
`tailscale-n100-location-independent-2026-07-19`.
|
||||
|
||||
---
|
||||
|
||||
## 2. Ground-truth snapshot (Phase A)
|
||||
|
||||
### Host `demo-felhom`
|
||||
|
||||
```
|
||||
vmbr0 192.168.0.147/24 dynamic <-- DHCP, NOT the former static .162
|
||||
tailscale0 100.70.170.35/32
|
||||
wg-felhom 10.77.0.2/32
|
||||
default via 192.168.0.1 dev vmbr0
|
||||
/etc/resolv.conf: nameserver 192.168.0.1 (no Pi-hole; remote router is the resolver)
|
||||
timedatectl: 2026-07-20 07:56 CEST / 05:56 UTC — synchronized: yes, NTP active, RTC sane
|
||||
```
|
||||
|
||||
**The remote site uses the same `192.168.0.0/24` subnet as home**, but the host now holds **`.147`**,
|
||||
not `.162`. This single fact is the root cause of §3.1.
|
||||
|
||||
Boot history — the box was **powered off for ~17h45m** in transport:
|
||||
|
||||
```
|
||||
-1 Sat 2026-07-18 18:17:00 CEST → Sun 2026-07-19 13:40:40 CEST
|
||||
0 Mon 2026-07-20 07:25:05 CEST → (current)
|
||||
```
|
||||
|
||||
WireGuard (PBS offsite leg) is **healthy**: peer `167.233.158.164:443`, latest handshake 56 s,
|
||||
keepalive 25 s, 1.70 MiB rx / 1.00 MiB sx.
|
||||
|
||||
### Guest 9201
|
||||
|
||||
```
|
||||
eth0 192.168.0.104/24 dynamic
|
||||
/etc/resolv.conf: nameserver 192.168.0.1
|
||||
curl https://gitea.dooplex.hu → 200 (0.196 s) outbound + DNS healthy
|
||||
```
|
||||
|
||||
### DooPlex
|
||||
|
||||
`tailscale` is not a host binary here — it runs as the k3s pod `admin-system/tailscale-6d7bdc9c5c-hwtjq`
|
||||
(uptime 6d22h). The peer view from the host shows `dooplex 100.107.87.53 … active; direct`.
|
||||
|
||||
**Nothing in this section was changed.**
|
||||
|
||||
---
|
||||
|
||||
## 3. Per-symptom evidence and root cause
|
||||
|
||||
### 3.1 "A tárolókezelő ügynök nem elérhető" — ROOT-CAUSED (critical)
|
||||
|
||||
The agent service is **failed**, not merely unreachable. Smoking gun from
|
||||
`journalctl -u felhom-agent`:
|
||||
|
||||
```
|
||||
level=ERROR msg="daemon: exited with error"
|
||||
err="localapi: bind 192.168.0.162:8443: listen tcp 192.168.0.162:8443:
|
||||
bind: cannot assign requested address"
|
||||
felhom-agent.service: Start request repeated too quickly.
|
||||
Active: failed (Result: exit-code) since Mon 2026-07-20 07:25:51 CEST
|
||||
```
|
||||
|
||||
The agent's `localapi` listener is pinned to the **literal `192.168.0.162`**. That address no longer
|
||||
exists on any interface (the host DHCP'd `.147`), so `bind()` fails and the daemon exits ~1.1 s after
|
||||
start. systemd retried 4× and gave up. **The agent has therefore been down since the first boot at
|
||||
the remote site — it has never run there.**
|
||||
|
||||
Controller side, the matching error class is **`no route to host`** (not a TLS-pin or auth failure —
|
||||
those would be STOP conditions of a different kind):
|
||||
|
||||
```
|
||||
[ERROR] [quiesce] cycle error: check due: agentapi: GET /backup/due:
|
||||
Get "https://192.168.0.162:8443/backup/due": dial tcp 192.168.0.162:8443: connect: no route to host
|
||||
[WARN] local-api: GET /storage failed (… no route to host) — channel not verified
|
||||
```
|
||||
|
||||
So **both ends are independently pinned to `192.168.0.162`**: the agent binds it, the controller
|
||||
dials it. Nothing is listening (`ss -tlnp | grep felhom` → empty). This is precisely the OPEN item
|
||||
recorded on 2026-07-19 — *"control plane pinned to 192.168.0.162 (agent `listen_addr` + guest
|
||||
`bootstrap.json`) → not portable off-LAN"* — now realized in production.
|
||||
|
||||
**Allowed action taken** (§3.2 of the runbook): `systemctl reset-failed && systemctl restart felhom-agent`.
|
||||
It failed identically within 6 s:
|
||||
|
||||
```
|
||||
Jul 20 08:00:02 … err="localapi: bind 192.168.0.162:8443: … cannot assign requested address"
|
||||
Jul 20 08:00:02 … felhom-agent.service: Failed with result 'exit-code'.
|
||||
```
|
||||
|
||||
**STOP condition hit** (runbook §3.4, "agent endpoint stale because the host IP changed"). The fix
|
||||
requires editing `/etc/felhom-agent/agent.json` (`listen_addr`) **and** the guest's
|
||||
bootstrap-managed endpoint — both forbidden in this run. Escalated as **F1**.
|
||||
|
||||
> Note for the fix: re-pinning to the literal `192.168.0.147` would work today but is a DHCP lease
|
||||
> away from breaking again, and the agent pins its **served leaf cert**, whose SAN set must cover
|
||||
> whatever address the controller dials. Options are laid out in F1.
|
||||
|
||||
**Blast radius while F1 stands:** every agent-backed function is dead — storage management
|
||||
(Tárhely), PBS whole-guest backups, the quiesce cycle, restore-tests, and the DR path. Guest-local
|
||||
functions (dashboard, apps, app-data/DB backup cache, Samba, cloudflared) are unaffected.
|
||||
|
||||
### 3.2 Calibre-Web Automated stopped — RESOLVED
|
||||
|
||||
```
|
||||
exit=128 finished=2026-07-19T11:40:15Z policy=unless-stopped restarts=0
|
||||
```
|
||||
|
||||
Container log ends with a clean application shutdown (`webserver stop (restart=False)`,
|
||||
`Performing shutdown of Calibre-Web Automated`) at 13:40:15 CEST — **~25 s before the host powered
|
||||
off for transport**. Not a crash: an orderly stop caught by the shutdown, which then never restarted.
|
||||
|
||||
**Allowed action taken:** started via the controller's real endpoint (the exact path the UI uses —
|
||||
session login, `<meta name="csrf-token">` read from the page, `X-CSRF-Token` header):
|
||||
|
||||
```
|
||||
POST /api/stacks/calibre-web/start → {"ok":true,"message":"Stack calibre-web start completed"}
|
||||
docker ps → calibre-web | Up 20 seconds (healthy)
|
||||
```
|
||||
|
||||
Verified healthy. Stack state now reads `running`.
|
||||
|
||||
### 3.3 Immich "Nem egészséges" + "URL nem elérhető" — ROOT-CAUSED, NOT fixed (per runbook D.3)
|
||||
|
||||
The two flags have **one** cause: `immich-server` is **not running at all**.
|
||||
|
||||
```
|
||||
immich-server Exited (143) 18 hours ago policy=unless-stopped restarts=0
|
||||
finished=2026-07-19T11:40:15Z health=unhealthy
|
||||
immich-postgres Up 31 minutes (healthy)
|
||||
immich-machine-learning Up 31 minutes (healthy)
|
||||
immich-redis Up 31 minutes (healthy)
|
||||
```
|
||||
|
||||
Exit **143 = SIGTERM** at 11:40:15Z — the *same second* as Calibre-Web, i.e. the same pre-transport
|
||||
shutdown. Its final log lines are a fully healthy server (`Immich Server is listening … [v3.0.3]`,
|
||||
ML server healthy, websocket clients connected) right up to the shutdown. **This is not leftover
|
||||
state from the 2026-07-19 offsite-restore test (R-43/R-44)** — hypothesis (a) is disproved by the
|
||||
logs: the server ran normally for ~42 min after that work, and there is no migration or DB error.
|
||||
|
||||
The URL probe is likewise **not** a remote-path/latency artifact — hypothesis (c) is disproved:
|
||||
|
||||
```
|
||||
curl https://immich.demo-felhom.eu → 404 in 0.239 s (traefik has no live backend)
|
||||
curl https://gitea.dooplex.hu → 200 in 0.196 s (same path, fine)
|
||||
```
|
||||
|
||||
Fast 404, not a timeout. The public CF path is healthy; there is simply nothing behind the route.
|
||||
**Left down deliberately** per runbook D.3. One-line fix available under GO — see **F5**.
|
||||
|
||||
### 3.4 "Utolsó mentés: Még nem futott" — ROOT-CAUSED as a **display bug** (independent of F1)
|
||||
|
||||
This is hypothesis (c), and it is provable from code rather than inferred.
|
||||
`dashboardHandler` (`controller/internal/web/handlers.go:134`) populates the backup card with:
|
||||
|
||||
```go
|
||||
data["BackupEnabled"] = s.cfg.Backup.Enabled // :172
|
||||
data["DBDumpStatus"] = fullStatus.LastDBDump // :175
|
||||
data["BackupRunning"] = fullStatus.Running // :177
|
||||
data["BackupMaxAgeHours"] = … // :178
|
||||
```
|
||||
|
||||
It **never sets `data["BackupStatus"]`** — and no other code path does either (repo-wide, the only
|
||||
occurrences of `BackupStatus` are the agent-proxy handler and this template). But
|
||||
`dashboard.html:116` gates the whole "Utolsó mentés" value on it:
|
||||
|
||||
```
|
||||
{{if .BackupStatus}} … LastRun / "Sikertelen" …
|
||||
{{else}} <span class="backup-status-none">Még nem futott</span>
|
||||
{{end}}
|
||||
```
|
||||
|
||||
`.BackupStatus` is therefore **always nil**, the `{{if}}` branch is dead code, and the dashboard
|
||||
reports **"Még nem futott" permanently on every box regardless of actual backup history**. The
|
||||
adjacent "Adatbázisok: 2 mentve" renders from `DBDumpStatus`, which *is* set — which is exactly the
|
||||
contradiction observed.
|
||||
|
||||
Independent corroboration that backups had in fact been running: the hub's own host-reports
|
||||
consistently carry `1 backups`, and hub events on 2026-07-19 include
|
||||
`crossdrive_completed` for bookstack / calibre-web / immich / Megosztások at 03:30 and
|
||||
`db_dump_completed` at 02:30.
|
||||
|
||||
Separately, and *not* a display artifact: **no backup can run at all while F1 stands** — the quiesce
|
||||
cycle fails every 5 min on `GET /backup/due`. The guest-local `backup-cache` job still runs fine
|
||||
(`Found 2 DB dump files across drives`, `Discovered 2 databases`, `Discovered app data: 3 apps`
|
||||
every 5 min). No manual backup was triggered: the runbook permits one only if the cause was (a) *and*
|
||||
the agent is healthy — neither holds.
|
||||
|
||||
### 3.5 Hub report freshness — HEALTHY, no gap
|
||||
|
||||
Reports land on the normal 15-min cadence from the remote site, and the staleness state machine
|
||||
recovered by itself on boot:
|
||||
|
||||
```
|
||||
07:25:42 Event from demo-felhom: controller_started (info) — Controller elindult (0.148.0)
|
||||
07:25:43 Staleness: demo-felhom down → ok (node_recovered)
|
||||
07:27:07 Event: app_start_failed — Telepített alkalmazás nem fut: Calibre-Web Automated
|
||||
07:27:40 Event: agent_channel_unknown (warning) — Controller→agent channel down: … no route to host
|
||||
07:27:41 Operator email sent for demo-felhom/agent_channel_unknown
|
||||
07:40:37 / 07:55:37 Received report from demo-felhom (7603 / 6232 bytes)
|
||||
```
|
||||
|
||||
Hub telemetry is location-independent (it rides the public path), and **the alerting chain worked
|
||||
end to end** — the agent-channel failure produced an operator email within 33 s of the boot grace
|
||||
expiring. The transport gap appears exactly as it should: last report 2026-07-19 11:40:22Z, next
|
||||
2026-07-20 05:25 UTC.
|
||||
|
||||
---
|
||||
|
||||
## 4. Ranked findings
|
||||
|
||||
| # | Sev | Finding | Evidence | Proposed fix | Viktor GO? |
|
||||
|---|-----|---------|----------|--------------|------------|
|
||||
| **F1** | **CRITICAL** | Agent `localapi` binds the literal `192.168.0.162`; host is `.147` at the remote site → `bind: cannot assign requested address` → service **failed since first remote boot**. Controller dials the same literal. All agent-backed features (storage, PBS backup, quiesce, restore-test, DR) are dead for the whole 2-week window. | §3.1 | Bind `0.0.0.0:8443` (or the vmbr0 address resolved at start-up) instead of a literal, and give the controller a stable dial target; ensure the served leaf's SAN covers it. Short-term unblock: set `listen_addr` to `192.168.0.147` + reissue/verify the pin + update the guest endpoint — but this re-breaks on the next DHCP lease. | **YES** — config **and** code; touches bootstrap-managed guest state |
|
||||
| **F2** | **HIGH** | No PBS/whole-guest backup can run while F1 stands → a ~2-week data-protection gap on the demo box. Quiesce errors every 5 min. | §3.1, §3.4 | Resolve F1; then verify a backup completes. | Follows F1 |
|
||||
| **F3** | **HIGH** | Dashboard permanently shows "Utolsó mentés: **Még nem futott**" on **every** box: `dashboardHandler` never sets `BackupStatus`, so the template's `{{if .BackupStatus}}` branch is unreachable dead code. Customer-facing misinformation about backup state. | §3.4 (`handlers.go:134-178`, `dashboard.html:116-131`) | Populate `data["BackupStatus"]` from `fullStatus` in `dashboardHandler`, or drop the dead branch and render from `DBDumpStatus`/`GuestBackup`. Needs a red-proof test (the current code has none covering this key). | **YES** — code |
|
||||
| **F4** | **MEDIUM** | **Multi-container apps under-alert.** `IsDownState` (`stacks/manager.go:46`) counts only `stopped`/`exited`; it deliberately excludes `unhealthy`. Immich's stack aggregates to `unhealthy` because 3 of 4 containers are up — so with its **primary container fully exited for 18 h and the app 100 % unreachable**, it produced **no** dead-app banner and **no** `app_start_failed` hub event. Calibre-Web (single container) alerted correctly within 90 s. | §3.3, live `/api/stacks` → `immich unhealthy`; hub log has an event for Calibre-Web only | Treat "primary/named container down" as down for a multi-container stack, or add an `unhealthy-for-N-cycles` escalation. | **YES** — code |
|
||||
| **F5** | **MEDIUM** | Two app containers were left `Exited` by the pre-transport shutdown and were **not** restarted on boot despite `restart: unless-stopped`; the controller **reports** deployed-but-stopped apps (30 s `deadapp-check`) but never **starts** them. Immich has now been down 18 h+ and stays down. | §3.2, §3.3, `cmd/controller/main.go:435` | Start `immich` via the same real endpoint used for Calibre-Web (immediate, no code). Then consider a bounded boot-time reconciliation (start deployed-but-stopped apps once, N attempts, with the existing boot grace). | Start = **YES** (runbook D.3 forbade it here); reconciler = code, YES |
|
||||
| **F6** | **LOW** | Agent logs `pbs: cannot read token secret; skipping datastore storage=felhom-pbs err="open /etc/pve/priv/storage/felhom-pbs.pw: permission denied"` on every start (pre-existing, R-39-adjacent — the non-root agent cannot read the PVE priv file). | §3.1 journal | Confirm whether the wrapper/sudoers path is meant to cover this read; fold into the R-39 follow-up. | YES |
|
||||
| **F7** | **LOW** | Incidental defect found while tracing CSRF: `app_export.html:93` builds `var domain = '<subdomain>.{{$.CSRFToken}}'` — the **session CSRF token is substituted where the domain belongs**. Produces a broken "open in browser" link and puts a CSRF token into a URL. | `controller/internal/web/templates/app_export.html:93` | Replace `{{$.CSRFToken}}` with the domain value (cf. the neighbouring correct `csrfH()` helper). | YES — code |
|
||||
|
||||
**Healthy / no action:** tailnet path direct at 37 ms · host clock NTP-synced (no transport RTC drift)
|
||||
· WireGuard PBS leg handshaking at 56 s · guest outbound + DNS · hub reports on cadence and
|
||||
staleness auto-recovered · operator alert email fired correctly · 6 apps + Samba + traefik +
|
||||
cloudflared running · USB HDD mounted and visible.
|
||||
|
||||
---
|
||||
|
||||
## 5. Actions taken
|
||||
|
||||
| Action | Authority | Outcome |
|
||||
|---|---|---|
|
||||
| `systemctl reset-failed && systemctl restart felhom-agent` | runbook §3.2 (one allowed restart, after evidence) | **Failed identically** in 6 s — same bind error. Confirms F1 is deterministic config, not a transient. |
|
||||
| `POST /api/stacks/calibre-web/start` via the controller's real session+CSRF path | runbook §3.2 (allowed, after evidence) | **Success** — `Up (healthy)`, stack state `running`. |
|
||||
|
||||
Nothing else was mutated. The ssh alias needed **no** edit (already on the tailnet IP).
|
||||
|
||||
---
|
||||
|
||||
## 6. Explicitly NOT fixed / not touched
|
||||
|
||||
- **F1** — no edit to `/etc/felhom-agent/agent.json`, the guest `bootstrap.json`, or any
|
||||
bootstrap/golden-managed state. STOP condition §3.4 hit and honoured.
|
||||
- **Immich** — left down per runbook D.3 (diagnose only), despite the fix being a one-call start.
|
||||
- **F3, F4, F7** — code changes, out of scope for this run.
|
||||
- **No** network/nftables/tailscale/route changes on either machine. **No** manual backup triggered.
|
||||
**No** version bumps, image builds, deploys, or `REPORT.md` overwrites. **No** containers or
|
||||
volumes deleted. **No** secrets written to any committed file.
|
||||
|
||||
---
|
||||
|
||||
## 7. Return-day checklist (~2026-08-02)
|
||||
|
||||
1. If F1 was fixed with a literal `.147`, revert it to the home-LAN value **or** (preferred) confirm
|
||||
the address-independent fix still binds correctly back on the home LAN.
|
||||
2. `~/.ssh/config` needs **no** revert — `Host felhom-pve` on the tailnet IP works from both
|
||||
locations; `felhom-pve-lan` remains as the direct-LAN path. Re-verify the `.162`/LAN path once the
|
||||
host is home and holds its expected address.
|
||||
3. Delete the **TEMPORARY** blocks from `felhom-controller/CLAUDE.md` and `felhom-agent/CLAUDE.md`.
|
||||
4. Confirm the host reacquires its intended home address (it is now on **DHCP** — the former static
|
||||
`.162` config did not survive/apply; decide whether to restore a static or keep the reservation).
|
||||
5. Re-verify: agent `active (running)`, dashboard banner clear, Tárhely renders, a PBS backup
|
||||
completes, Immich healthy and reachable.
|
||||
6. Flip this document's status from **ACTIVE** to **CLOSED**.
|
||||
Reference in New Issue
Block a user