docs(tests): unattended test campaign findings (2026-06-22, demo 9201)

Full-feature validation across the running demo: PBS backup, komga/gitea HC fixes,
deploy sweep, backup+restore, DR rebuild, storage, monitoring/alerts/hub, guest reboot.
Headline finding: controller->agent local-API (8443) unreachable from the controller
container, gating storage UI + host metrics + whole-guest backup.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
This commit is contained in:
2026-06-22 16:29:10 +02:00
parent 4a2d72e238
commit f70c011749
@@ -0,0 +1,262 @@
# Unattended test campaign — full-feature validation (demo guest 9201) — 2026-06-22
> **Status: empirical findings from an unattended live test campaign (2026-06-22).** Operator away,
> glancing remotely. Goal: maximize feature coverage across the running demo (Felhom on Proxmox:
> host agent `felhom-agent v0.39.0` on `felhom-pve`, in-guest controller `felhom-controller v0.73.0`
> in LXC 9201, hub `felhom-hub` on k3s/DooPlex) **without leaving the demo unrecoverably broken.**
> Only code change permitted: the komga catalog healthcheck fix (+ one analogous gitea HC fix the
> sweep surfaced, which the runbook sanctions). No controller/agent build. All destructive DR work
> targeted throwaway guests; the live demo was restored to baseline (and is healthier — komga fixed).
>
> Companion live report (per-check raw evidence, continuously updated during the run):
> `felhom-controller/TEST-REPORT.md`. This document is the consolidated findings writeup.
## Topology under test
- **Host / agent:** `felhom-pve` (`demo-felhom`, PVE 9.2.2, Intel N100, 16 GB), `felhom-agent v0.39.0`
(systemd active). SSH `root@felhom-pve`.
- **Guest / controller:** LXC **9201** (`demo-felhom`, 12 GB, 32 GB rootfs + 256 GB `/var/lib/docker`
vol), `felhom-controller:0.73.0` (docker container, no auth wall currently). Customer id `demo-felhom`.
- **Hub:** `felhom-hub` on k3s @ DooPlex `192.168.0.180`; `hub.felhom.eu`.
- **Backup target:** PBS `felhom-pbs` (datastore `felhom-spike` on DooPlex), 9.6 TB, 21 % used.
- **Drives (enrolled):** `felhom-flash` (USB, `uuid:81a26531…`, `/mnt/felhom-drives/felhom-flash`,
117 GB, holds app userdata for 8 apps) · `felhom-usb` (local-dir, `uuid:da9e7089…`,
`/mnt/felhom-drives/felhom-usb`, 916 GB, empty) · `local-lvm` · `local` · `felhom-pbs`.
- **Baseline apps:** 14 deployed app stacks (25 containers incl. 4 infra: controller, traefik,
cloudflared, filebrowser); 55 in catalog, 38 installable. komga `unhealthy` at baseline.
### Access note — how the UI was driven
On the LAN, split-horizon DNS (resolver `192.168.0.250`) maps `*.demo-felhom.eu` → guest 9201
(`192.168.0.141`), bypassing Cloudflare. The guest's traefik **`:443` (websecure) returns Go
`404 page not found` for every host** (controller and apps alike) while serving a valid LE
`*.demo-felhom.eu` cert; `:80` 301-redirects to https → so a browser dead-ends at the 404 on-LAN.
**Workaround:** a short-lived python TCP forwarder inside 9201 (`0.0.0.0:18080` → controller
`<container-ip>:8080`); the browser hit `http://felhom.demo-felhom.eu:18080/` (LAN DNS → guest:18080),
and the controller's `CatchAllMiddleware` strips the port so the Host matches `felhom.<domain>` and the
dashboard serves over plain HTTP. Forwarder removed at end of run. The controller dashboard is
reachable internally at traefik file-route `Host(felhom.demo-felhom.eu) → felhom-controller:8080`.
---
## Phase 0 — Baseline + PBS safety-net backup
| Check | Result | Evidence |
|---|---|---|
| Baseline capture | ✅ | agent v0.39.0; ctrl v0.73.0; 25 containers, komga unhealthy; drives + pvesm healthy |
| Fresh PBS whole-CT backup of 9201 | ✅ | `felhom-agent --selftest=backup -vmid=9201``felhom-pbs:backup/ct/9201/2026-06-22T12:45:37Z`, 19.62 GB, crash-consistent, 161 s. Uncovered (by design): `/etc/felhom-bootstrap`, `/mnt/felhom-drives` (drives backed up separately) |
| PBS verify | ✅ | `--selftest=pbs-verify``verify_state: ok` for the new snapshot |
The verified backup is the DR source for Phase 4. **Gate to destructive phases: opened.**
---
## Phase 1 — komga healthcheck fix + git-sync + health state machine
**Diagnosis (live):** the docker HC `curl -f http://localhost:25600/api/v1/actuator/health` exited 22
(HTTP **401** — the `/api/v1` prefix is auth-gated) → `unhealthy` for 18 h while komga served fine.
Endpoint probe matrix (wget inside the image): `/`, `/actuator/health`, `/api/v1/oauth2/providers`,
`/login` → all 200; only `/api/v1/actuator/health` → 401. **Komga's actuator is served unauthenticated
at `/actuator/health`, off the `/api/v1` API prefix.** A second probe matters too: the controller's own
`.felhom.yml` `healthcheck.checks[].path` (which drives the dashboard badge + route publishing) pointed
at the same 401 path. The `gotson/komga:1.20.0` image ships `curl` (verified).
**Fix (app-catalog `main`):** repoint **both** probes to `/actuator/health` — docker HC
(`3faa5ae`) and `.felhom.yml` (`9b066de`).
**Validation (end-to-end through the real UI + server pipeline):**
- GUI "Sablonok frissítése" → controller git-sync: `[sync] Updated komga/docker-compose.yml` +
`komga/.felhom.yml`; content-hash copy landed in `/opt/docker/stacks/komga/`.
- GUI "Frissítés" → `POST /api/stacks/komga/update` 200 → `docker compose up -d --remove-orphans`
recreated the container (4.2 s).
- Result: `docker ps``komga Up (healthy)`, HC now `/actuator/health`; controller log
`Health probes: N ok`; dashboard badge `Nem egészséges`**Fut**; the
"⚠ URL nem elérhető útvonal nincs publikálva" route warning cleared.
This one phase exercised: catalog fix, git-sync (content-hash copy), redeploy, the health-aware state
machine, and route publishing.
---
## Phase 2 — App deployment sweep
Scope note: catalog is **55 apps, 38 not deployed** (the runbook's "14 enabled" was stale). Deploying
all 38 would exhaust the 12 GB guest, so a representative sample was deployed via the GUI; the heavy
multi-container case (immich, 4 containers) is already deployed + healthy (cited).
| App | Result | Evidence |
|---|---|---|
| **gitea** (GUI deploy) | ✅ after HC fix | GUI deploy page → memory-gate bar (two-segment current+new) → 3-step progress panel → containers up. Initially `unhealthy`: HC `curl -f /api/v1/version`**404** (404 until install-lock) while `/api/healthz` → 200. **Fixed catalog** (`17e00b7`, start_period 30→90 s) → re-sync → redeploy → `healthy`; endpoint `:3000/` → 200 |
| **glance** (GUI deploy) | ❌ logged, not fixed | crash-loop `Restarting (1)`. Log: `open /app/config/glance.yml: no such file or directory`. Catalog template mounts an empty `glance_config` named volume but **never seeds the required `glance.yml`**. Not a trivial HC fix → logged. Removed cleanly (`stop`+`remove`) |
| **immich** (pre-deployed) | ✅ cited | 4 containers healthy = heavy-app coverage |
Bonus coverage: the stop + remove endpoints (glance + gitea teardown). Both test deployments were
removed at cleanup to restore the baseline app set.
---
## Phase 3 — Backup + restore (per-app)
| Check | Result | Evidence |
|---|---|---|
| Manual DB-dump backup | ✅ | `POST /api/backup/run` → 3 DBs: romm-mariadb (14 tbl, 38.5 KB), paperless-postgres (67 tbl, 360 KB), immich-postgres (60 tbl, 43.2 MB); 43.6 MB / 2.18 s. On-disk fresh (`…/primary/<app>/db-dumps/*.sql`) |
| Tier-2 cross-drive copy | ✅ | `POST /api/backup/tier2` → 8 HDD apps → `/mnt/sys_drive/felhom-data/backups/secondary/<app>` (immich 156 MB on disk; romm appdata + recovery-unit{db-dump,manifest,compose}); `crossdrive_completed` events pushed to hub |
| Per-app restore | ✅ | `POST /backup/restore stack_name=romm``RestoreFromRecoveryUnit`: compose down → redeploy (6 env, 3 encrypted recovered) → `Imported DB dump romm-mariadb.sql` → containers up → "Restore-from-unit completed" |
| Restore non-hollow | ✅ | planted a marker table pre-restore; post-restore romm DB has the real `platforms` + 17 tables, 3 containers healthy; marker dropped afterward to restore baseline |
**Findings:**
- **`GET /api/backup/snapshots` has no server handler** (the api-router returns "endpoint not found").
The restore UI's snapshot dropdown (`backups.html:682`) therefore cannot populate → a **pure-UI
restore is blocked at snapshot selection.** The restore itself works via the form action
`POST /backup/restore` (which calls `RestoreFromRecoveryUnit(stack)` and ignores `snapshot_id`'s value).
- **Restore DB import is additive** — `RestoreFromRecoveryUnit` replays the dump but does not drop
tables absent from it (the planted marker table survived). Snapshot data is restored; stale/extra
tables are not pruned. Minor behavioral note.
---
## Phase 4 — DR rebuild from backup (non-destructive to 9201)
Proven **twice**, both collision-safe (no tunnel/hub identity clash with the live 9201):
1. **Agent DR primitive**`felhom-agent --selftest=restore-test -archive=<Phase-0 volid>`: restored
into scratch **990000** → bind-mounts neutralized → **net link-down** → boot → `verified: boot+running`
→ torn down. `pass:true`, 2 m 56 s, self-cleaning.
2. **Deep manual restore into throwaway 9300**`pct restore 9300 <archive> --storage local-lvm`.
**Safety before start:** removed the shared binds `mp3:/mnt/felhom-drives` + `mp9:/etc/felhom-bootstrap`,
removed `net0`, set `onboot 0` (kept only the restored `mp0` docker-data + rootfs — independent
copies). Then started and inspected:
- `pct status: running`; `felhom-controller Up (healthy)` inside.
- `/opt/docker/stacks/` had all 57 stack dirs; romm `app.yaml` restored (`deployed:true`, env, locked_fields).
- **`docker ps` inside 9300: 25 containers up** from restored volumes — identical set to 9201.
- Restored data non-hollow: 9300 `romm-db` has 17 tables.
- No collision: 9300 had **no eth0**; live 9201 `cloudflared Up` unchanged throughout.
- Cleanup: `pct stop 9300 && pct destroy 9300` (both LVs removed); local-lvm freed.
**Core DR claim proven end-to-end:** PBS whole-CT restore → guest boots → controller + all app stacks +
DB data recovered from restored volumes. The bind-mount-neutralization recipe (remove host-path binds +
`net0` + `onboot` before start) is the safe way to inspect a restored customer image alongside a live
one.
---
## Phase 5 — Storage lifecycle — **BLOCKED by a controller→agent connectivity defect**
### 🔴 FINDING #1 (highest priority) — controller cannot reach the agent local API (8443)
The controller container's calls to the agent local API fail: `agentapi: GET /disks: dial tcp
192.168.0.162:8443: connect: cannot assign requested address` (EADDRNOTAVAIL, persistent — 6/6). The
same error surfaces in the Settings UI ("Meghajtók (ügynök nézet)") and in the backup page
("A host-ügynök jelenleg nem elérhető").
Reachability matrix **from the controller container**:
- `192.168.0.162:8006` (pveproxy, binds `0.0.0.0`) → **200** (reachable)
- `192.168.0.162:8443` (felhom-agent local API) → **refused / EADDRNOTAVAIL**
- `192.168.0.141:443` (guest's own traefik) → 404 (reachable)
From the guest **host namespace**, `192.168.0.162:8443` is OPEN. Confirmed: `ss` shows the agent
`LISTEN 192.168.0.162:8443`**bound to the specific host LAN IP, not `0.0.0.0`** like pveproxy;
**no iptables rule** references 8443; **pve-firewall disabled**. The agent itself is healthy
(`--selftest=storage` enumerates all 5 targets with SMART/FS/durable_id fine). So this is a
**container→host-IP routing / source-bind mismatch on port 8443**, not a firewall and not an agent
crash.
**Impact — the entire agent-backed feature set is down in this demo:** storage management UI
(`/api/disks`, `/api/storage`: scan/label/eject/enroll/format/migrate/decommission), the live
host-metrics API (`/api/host-metrics`), and whole-guest backup trigger. (Monitoring still shows host
CPU/mem/temp — those are read from `/proc`, since the LXC shares the host kernel, not via the agent.)
**Fix direction (supervised):** bind the agent local API to `0.0.0.0` (or the `vmbr0` bridge address the
guest routes through), or repoint the controller's `local_api.endpoint` (bootstrap.json) at a
container-reachable address; then re-run Phase 5's mutating checks. Not changed in this run (network
reconfig is out of scope for an unattended test).
| Check | Result | Evidence |
|---|---|---|
| Scan / health / FS / model / durable_id (agent tier) | ✅ | `--selftest=storage`: 5 targets; felhom-usb SMART **PASSED** temp=35 poh=2736 realloc=0; felhom-flash uuid:81a26531; thin-pool data=9.7 % |
| Label edit + revert (UI) | ⛔ blocked | controller→agent down (finding #1) |
| Eject → re-enroll cycle (UI) | ⛔ blocked | same |
| Destructive loopback (scan→format→mount→migrate→decommission) | ⛔ skipped | wizard drives the agent via the broken path; agent-direct needs the per-guest token + TLS-pin reconstruction → high-effort/low-confidence unattended. **No scratch loopback created** (nothing to clean up). Deferred to supervised. |
---
## Phase 6 — Monitoring / alerts / notifications / settings / hub
| Check | Result | Evidence |
|---|---|---|
| Monitoring — host metrics | ✅ | CPU/load/mem (7.6/15.4 GB — host's 16 GB via `/proc`/LXC kernel-share)/temp/uptime |
| Monitoring — container/app metrics | ✅ | per-container memory breakdown (26 containers), per-app CPU/MEM charts, time-range charts 1h30d |
| Alerts — state-based | ✅ | komga `0 ok, 1 unhealthy` before → `N ok` after the Phase-1 fix; badge red→green |
| Notification path (exercised ONCE) | ✅ | set email+events+cooldown → hub `Notification preferences updated for demo-felhom`. Test → hub **`Event from demo-felhom: test``Test email sent to nagyfenyvesi.viktor@gmail.com`** = controller→hub→Resend→email confirmed. One email only (§0.5) |
| Settings persistence | ✅ | notification prefs set → persisted to `settings.json` → reverted (both directions) |
| Settings — password change | n/a | password protection is **operator-only and not configured** ("A jelszavas védelem nincs beállítva"); dashboard currently serves with no auth wall |
| Hub — report received | ✅ | `Received report from demo-felhom (11983 bytes)` every 15 min; `host-report (3 guests, 5 storage targets, 1 backups, 1 restore-tests, 9 pbs-snapshots)`; UI "Hub: Kapcsolódva" |
| Hub — DR recipe carries the PBS coord | ✅ | `dr_recipe` (demo-felhom) host-half `pbs={"repo_id":"felhom-pbs","namespace":"root","latest_snapshot_id":"9201"}`; app-half 15 apps; **no secrets in either half** |
| Hub — update panel | ✅ | version 0.73.0, "naprakész", last check timestamp shown |
| Hub GUI tiles (analytics/healthchecks) | skipped | no operator GUI creds — verified hub-side state via logs + DB instead |
The DR-recipe app-half + host-half are both stored at the hub each report (`DR-recipe …-half stored for
customer demo-felhom (v1)`), and the assembled recipe carries the live PBS coordinate (the agent v0.39.0
live-reporter feature) — verified directly in the hub DB.
---
## Phase 7 — Resilience
| Check | Result | Evidence |
|---|---|---|
| Guest 9201 reboot | ✅ | `pct reboot 9201`; recovered to 26→**25** containers, 0 unhealthy in ~2 min |
| Drives re-bind | ✅ | felhom-flash (sdc1) + felhom-usb (sdb1) remounted ext4 at correct mountpoints |
| FileBrowser convergence | ✅ | binds reconverged to `<drive>/userdata → /srv/<drive>` |
| Tunnel reconnect | ✅ | cloudflared QUIC + HTTP/2 prechecks to region1/2 argotunnel pass |
| Controller + apps + Phase-1/2 fixes persist | ✅ | komga + gitea still healthy post-reboot |
| Host reboot | ⏭ skipped (§7) | unattended risk — deferred to supervised |
Incidental: the controller's docker container IP changed across the reboot
(`172.18.0.10``172.18.0.8`); harmless to the system (traefik routes by docker DNS name), only required
re-pointing the test forwarder.
---
## Consolidated findings (for supervised follow-up)
1. **🔴 controller→agent local-API (8443) unreachable** (Phase 5) — gates storage UI + host-metrics API
+ whole-guest backup. Agent binds `192.168.0.162` not `0.0.0.0`; no firewall; container reaches
`:8006` but not `:8443` on the same host. **Highest priority.** Fix: bind `0.0.0.0`/bridge or
repoint `local_api.endpoint`.
2. **`GET /api/backup/snapshots` has no handler** (Phase 3) — restore UI snapshot dropdown can't
populate (restore works via `POST /backup/restore`).
3. **glance catalog template** never seeds `glance.yml` → crash-loop on deploy (Phase 2). Needs a
default-config seed (init step / entrypoint), not a trivial HC fix.
4. **Restore DB import is additive** (Phase 3) — doesn't drop tables absent from the dump. Minor.
5. **On-LAN HTTPS to dashboard/apps 404s** at traefik `:443` for all hosts (cert valid) — supported
path is Cloudflare; on-LAN HTTPS appears non-functional. Worth confirming the Cloudflare path is the
intended sole ingress.
6. Minor: a few catalog logo SVGs 404 (adventurelog, cloudflared, crafty-controller); dashboard
"Adatbázisok mentve" count flickered 3→2 post-reboot (cosmetic).
## Code shipped (app-catalog `main`, trunk)
- komga HC `/api/v1/actuator/health`(401) → `/actuator/health` (docker HC + `.felhom.yml` probe) —
`3faa5ae`, `9b066de`.
- gitea HC `/api/v1/version`(404 pre-install) → `/api/healthz`, start_period 90 s — `17e00b7`.
- CHANGELOG + REPORT updated. **No felhom-controller / felhom-agent code changed** (catalog-only).
## Cleanup confirmation
- Throwaway guest **9300 destroyed**; restore-test scratch 990000 auto-torn-down; `pct list` =
9201 (running) + pre-existing 9001/9999 (stopped, not created by this run).
- **No scratch loopback devices created** (destructive storage sub-phase skipped).
- Test apps removed (glance, gitea) → baseline app set.
- Notification settings reverted; password never set; drives never ejected.
- Test forwarder (`/tmp/felhom-test-fwd.py`) removed from the guest; temp `hub.db` copy removed from
DooPlex. **DooPlex used in normal roles only** (PBS target, gitea/hub reads).
- **komga healthy** (improvement over the unhealthy baseline).
## After-state vs before
9201 running; **25 containers Up, 0 unhealthy** (before: komga unhealthy); controller v0.73.0; tunnel up;
5 storage targets reachable. Net result: demo restored to baseline **+ komga fixed**.
## Deferred — needs a supervised run
- **Agent 8443 reachability** root-cause + fix (finding #1) — gates the whole agent-backed subsystem.
- **Storage lifecycle UI** (label edit, eject→re-enroll, init-wizard format/migrate/decommission on a
loopback) — re-run once finding #1 is fixed.
- **Host reboot** resilience (full host AC-recovery + remount-by-UUID under sdX reshuffle).
- **glance** config-seeding fix (finding #3); **`/api/backup/snapshots`** handler + UI snapshot
selection (finding #2).