Files
felhom.eu/documentation/tests/unattended-test-campaign-2026-06-22-findings.md
T
admin f70c011749 docs(tests): unattended test campaign findings (2026-06-22, demo 9201)
Full-feature validation across the running demo: PBS backup, komga/gitea HC fixes,
deploy sweep, backup+restore, DR rebuild, storage, monitoring/alerts/hub, guest reboot.
Headline finding: controller->agent local-API (8443) unreachable from the controller
container, gating storage UI + host metrics + whole-guest backup.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-22 16:29:10 +02:00

263 lines
18 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Unattended test campaign — full-feature validation (demo guest 9201) — 2026-06-22
> **Status: empirical findings from an unattended live test campaign (2026-06-22).** Operator away,
> glancing remotely. Goal: maximize feature coverage across the running demo (Felhom on Proxmox:
> host agent `felhom-agent v0.39.0` on `felhom-pve`, in-guest controller `felhom-controller v0.73.0`
> in LXC 9201, hub `felhom-hub` on k3s/DooPlex) **without leaving the demo unrecoverably broken.**
> Only code change permitted: the komga catalog healthcheck fix (+ one analogous gitea HC fix the
> sweep surfaced, which the runbook sanctions). No controller/agent build. All destructive DR work
> targeted throwaway guests; the live demo was restored to baseline (and is healthier — komga fixed).
>
> Companion live report (per-check raw evidence, continuously updated during the run):
> `felhom-controller/TEST-REPORT.md`. This document is the consolidated findings writeup.
## Topology under test
- **Host / agent:** `felhom-pve` (`demo-felhom`, PVE 9.2.2, Intel N100, 16 GB), `felhom-agent v0.39.0`
(systemd active). SSH `root@felhom-pve`.
- **Guest / controller:** LXC **9201** (`demo-felhom`, 12 GB, 32 GB rootfs + 256 GB `/var/lib/docker`
vol), `felhom-controller:0.73.0` (docker container, no auth wall currently). Customer id `demo-felhom`.
- **Hub:** `felhom-hub` on k3s @ DooPlex `192.168.0.180`; `hub.felhom.eu`.
- **Backup target:** PBS `felhom-pbs` (datastore `felhom-spike` on DooPlex), 9.6 TB, 21 % used.
- **Drives (enrolled):** `felhom-flash` (USB, `uuid:81a26531…`, `/mnt/felhom-drives/felhom-flash`,
117 GB, holds app userdata for 8 apps) · `felhom-usb` (local-dir, `uuid:da9e7089…`,
`/mnt/felhom-drives/felhom-usb`, 916 GB, empty) · `local-lvm` · `local` · `felhom-pbs`.
- **Baseline apps:** 14 deployed app stacks (25 containers incl. 4 infra: controller, traefik,
cloudflared, filebrowser); 55 in catalog, 38 installable. komga `unhealthy` at baseline.
### Access note — how the UI was driven
On the LAN, split-horizon DNS (resolver `192.168.0.250`) maps `*.demo-felhom.eu` → guest 9201
(`192.168.0.141`), bypassing Cloudflare. The guest's traefik **`:443` (websecure) returns Go
`404 page not found` for every host** (controller and apps alike) while serving a valid LE
`*.demo-felhom.eu` cert; `:80` 301-redirects to https → so a browser dead-ends at the 404 on-LAN.
**Workaround:** a short-lived python TCP forwarder inside 9201 (`0.0.0.0:18080` → controller
`<container-ip>:8080`); the browser hit `http://felhom.demo-felhom.eu:18080/` (LAN DNS → guest:18080),
and the controller's `CatchAllMiddleware` strips the port so the Host matches `felhom.<domain>` and the
dashboard serves over plain HTTP. Forwarder removed at end of run. The controller dashboard is
reachable internally at traefik file-route `Host(felhom.demo-felhom.eu) → felhom-controller:8080`.
---
## Phase 0 — Baseline + PBS safety-net backup
| Check | Result | Evidence |
|---|---|---|
| Baseline capture | ✅ | agent v0.39.0; ctrl v0.73.0; 25 containers, komga unhealthy; drives + pvesm healthy |
| Fresh PBS whole-CT backup of 9201 | ✅ | `felhom-agent --selftest=backup -vmid=9201``felhom-pbs:backup/ct/9201/2026-06-22T12:45:37Z`, 19.62 GB, crash-consistent, 161 s. Uncovered (by design): `/etc/felhom-bootstrap`, `/mnt/felhom-drives` (drives backed up separately) |
| PBS verify | ✅ | `--selftest=pbs-verify``verify_state: ok` for the new snapshot |
The verified backup is the DR source for Phase 4. **Gate to destructive phases: opened.**
---
## Phase 1 — komga healthcheck fix + git-sync + health state machine
**Diagnosis (live):** the docker HC `curl -f http://localhost:25600/api/v1/actuator/health` exited 22
(HTTP **401** — the `/api/v1` prefix is auth-gated) → `unhealthy` for 18 h while komga served fine.
Endpoint probe matrix (wget inside the image): `/`, `/actuator/health`, `/api/v1/oauth2/providers`,
`/login` → all 200; only `/api/v1/actuator/health` → 401. **Komga's actuator is served unauthenticated
at `/actuator/health`, off the `/api/v1` API prefix.** A second probe matters too: the controller's own
`.felhom.yml` `healthcheck.checks[].path` (which drives the dashboard badge + route publishing) pointed
at the same 401 path. The `gotson/komga:1.20.0` image ships `curl` (verified).
**Fix (app-catalog `main`):** repoint **both** probes to `/actuator/health` — docker HC
(`3faa5ae`) and `.felhom.yml` (`9b066de`).
**Validation (end-to-end through the real UI + server pipeline):**
- GUI "Sablonok frissítése" → controller git-sync: `[sync] Updated komga/docker-compose.yml` +
`komga/.felhom.yml`; content-hash copy landed in `/opt/docker/stacks/komga/`.
- GUI "Frissítés" → `POST /api/stacks/komga/update` 200 → `docker compose up -d --remove-orphans`
recreated the container (4.2 s).
- Result: `docker ps``komga Up (healthy)`, HC now `/actuator/health`; controller log
`Health probes: N ok`; dashboard badge `Nem egészséges`**Fut**; the
"⚠ URL nem elérhető útvonal nincs publikálva" route warning cleared.
This one phase exercised: catalog fix, git-sync (content-hash copy), redeploy, the health-aware state
machine, and route publishing.
---
## Phase 2 — App deployment sweep
Scope note: catalog is **55 apps, 38 not deployed** (the runbook's "14 enabled" was stale). Deploying
all 38 would exhaust the 12 GB guest, so a representative sample was deployed via the GUI; the heavy
multi-container case (immich, 4 containers) is already deployed + healthy (cited).
| App | Result | Evidence |
|---|---|---|
| **gitea** (GUI deploy) | ✅ after HC fix | GUI deploy page → memory-gate bar (two-segment current+new) → 3-step progress panel → containers up. Initially `unhealthy`: HC `curl -f /api/v1/version`**404** (404 until install-lock) while `/api/healthz` → 200. **Fixed catalog** (`17e00b7`, start_period 30→90 s) → re-sync → redeploy → `healthy`; endpoint `:3000/` → 200 |
| **glance** (GUI deploy) | ❌ logged, not fixed | crash-loop `Restarting (1)`. Log: `open /app/config/glance.yml: no such file or directory`. Catalog template mounts an empty `glance_config` named volume but **never seeds the required `glance.yml`**. Not a trivial HC fix → logged. Removed cleanly (`stop`+`remove`) |
| **immich** (pre-deployed) | ✅ cited | 4 containers healthy = heavy-app coverage |
Bonus coverage: the stop + remove endpoints (glance + gitea teardown). Both test deployments were
removed at cleanup to restore the baseline app set.
---
## Phase 3 — Backup + restore (per-app)
| Check | Result | Evidence |
|---|---|---|
| Manual DB-dump backup | ✅ | `POST /api/backup/run` → 3 DBs: romm-mariadb (14 tbl, 38.5 KB), paperless-postgres (67 tbl, 360 KB), immich-postgres (60 tbl, 43.2 MB); 43.6 MB / 2.18 s. On-disk fresh (`…/primary/<app>/db-dumps/*.sql`) |
| Tier-2 cross-drive copy | ✅ | `POST /api/backup/tier2` → 8 HDD apps → `/mnt/sys_drive/felhom-data/backups/secondary/<app>` (immich 156 MB on disk; romm appdata + recovery-unit{db-dump,manifest,compose}); `crossdrive_completed` events pushed to hub |
| Per-app restore | ✅ | `POST /backup/restore stack_name=romm``RestoreFromRecoveryUnit`: compose down → redeploy (6 env, 3 encrypted recovered) → `Imported DB dump romm-mariadb.sql` → containers up → "Restore-from-unit completed" |
| Restore non-hollow | ✅ | planted a marker table pre-restore; post-restore romm DB has the real `platforms` + 17 tables, 3 containers healthy; marker dropped afterward to restore baseline |
**Findings:**
- **`GET /api/backup/snapshots` has no server handler** (the api-router returns "endpoint not found").
The restore UI's snapshot dropdown (`backups.html:682`) therefore cannot populate → a **pure-UI
restore is blocked at snapshot selection.** The restore itself works via the form action
`POST /backup/restore` (which calls `RestoreFromRecoveryUnit(stack)` and ignores `snapshot_id`'s value).
- **Restore DB import is additive** — `RestoreFromRecoveryUnit` replays the dump but does not drop
tables absent from it (the planted marker table survived). Snapshot data is restored; stale/extra
tables are not pruned. Minor behavioral note.
---
## Phase 4 — DR rebuild from backup (non-destructive to 9201)
Proven **twice**, both collision-safe (no tunnel/hub identity clash with the live 9201):
1. **Agent DR primitive**`felhom-agent --selftest=restore-test -archive=<Phase-0 volid>`: restored
into scratch **990000** → bind-mounts neutralized → **net link-down** → boot → `verified: boot+running`
→ torn down. `pass:true`, 2 m 56 s, self-cleaning.
2. **Deep manual restore into throwaway 9300**`pct restore 9300 <archive> --storage local-lvm`.
**Safety before start:** removed the shared binds `mp3:/mnt/felhom-drives` + `mp9:/etc/felhom-bootstrap`,
removed `net0`, set `onboot 0` (kept only the restored `mp0` docker-data + rootfs — independent
copies). Then started and inspected:
- `pct status: running`; `felhom-controller Up (healthy)` inside.
- `/opt/docker/stacks/` had all 57 stack dirs; romm `app.yaml` restored (`deployed:true`, env, locked_fields).
- **`docker ps` inside 9300: 25 containers up** from restored volumes — identical set to 9201.
- Restored data non-hollow: 9300 `romm-db` has 17 tables.
- No collision: 9300 had **no eth0**; live 9201 `cloudflared Up` unchanged throughout.
- Cleanup: `pct stop 9300 && pct destroy 9300` (both LVs removed); local-lvm freed.
**Core DR claim proven end-to-end:** PBS whole-CT restore → guest boots → controller + all app stacks +
DB data recovered from restored volumes. The bind-mount-neutralization recipe (remove host-path binds +
`net0` + `onboot` before start) is the safe way to inspect a restored customer image alongside a live
one.
---
## Phase 5 — Storage lifecycle — **BLOCKED by a controller→agent connectivity defect**
### 🔴 FINDING #1 (highest priority) — controller cannot reach the agent local API (8443)
The controller container's calls to the agent local API fail: `agentapi: GET /disks: dial tcp
192.168.0.162:8443: connect: cannot assign requested address` (EADDRNOTAVAIL, persistent — 6/6). The
same error surfaces in the Settings UI ("Meghajtók (ügynök nézet)") and in the backup page
("A host-ügynök jelenleg nem elérhető").
Reachability matrix **from the controller container**:
- `192.168.0.162:8006` (pveproxy, binds `0.0.0.0`) → **200** (reachable)
- `192.168.0.162:8443` (felhom-agent local API) → **refused / EADDRNOTAVAIL**
- `192.168.0.141:443` (guest's own traefik) → 404 (reachable)
From the guest **host namespace**, `192.168.0.162:8443` is OPEN. Confirmed: `ss` shows the agent
`LISTEN 192.168.0.162:8443`**bound to the specific host LAN IP, not `0.0.0.0`** like pveproxy;
**no iptables rule** references 8443; **pve-firewall disabled**. The agent itself is healthy
(`--selftest=storage` enumerates all 5 targets with SMART/FS/durable_id fine). So this is a
**container→host-IP routing / source-bind mismatch on port 8443**, not a firewall and not an agent
crash.
**Impact — the entire agent-backed feature set is down in this demo:** storage management UI
(`/api/disks`, `/api/storage`: scan/label/eject/enroll/format/migrate/decommission), the live
host-metrics API (`/api/host-metrics`), and whole-guest backup trigger. (Monitoring still shows host
CPU/mem/temp — those are read from `/proc`, since the LXC shares the host kernel, not via the agent.)
**Fix direction (supervised):** bind the agent local API to `0.0.0.0` (or the `vmbr0` bridge address the
guest routes through), or repoint the controller's `local_api.endpoint` (bootstrap.json) at a
container-reachable address; then re-run Phase 5's mutating checks. Not changed in this run (network
reconfig is out of scope for an unattended test).
| Check | Result | Evidence |
|---|---|---|
| Scan / health / FS / model / durable_id (agent tier) | ✅ | `--selftest=storage`: 5 targets; felhom-usb SMART **PASSED** temp=35 poh=2736 realloc=0; felhom-flash uuid:81a26531; thin-pool data=9.7 % |
| Label edit + revert (UI) | ⛔ blocked | controller→agent down (finding #1) |
| Eject → re-enroll cycle (UI) | ⛔ blocked | same |
| Destructive loopback (scan→format→mount→migrate→decommission) | ⛔ skipped | wizard drives the agent via the broken path; agent-direct needs the per-guest token + TLS-pin reconstruction → high-effort/low-confidence unattended. **No scratch loopback created** (nothing to clean up). Deferred to supervised. |
---
## Phase 6 — Monitoring / alerts / notifications / settings / hub
| Check | Result | Evidence |
|---|---|---|
| Monitoring — host metrics | ✅ | CPU/load/mem (7.6/15.4 GB — host's 16 GB via `/proc`/LXC kernel-share)/temp/uptime |
| Monitoring — container/app metrics | ✅ | per-container memory breakdown (26 containers), per-app CPU/MEM charts, time-range charts 1h30d |
| Alerts — state-based | ✅ | komga `0 ok, 1 unhealthy` before → `N ok` after the Phase-1 fix; badge red→green |
| Notification path (exercised ONCE) | ✅ | set email+events+cooldown → hub `Notification preferences updated for demo-felhom`. Test → hub **`Event from demo-felhom: test``Test email sent to nagyfenyvesi.viktor@gmail.com`** = controller→hub→Resend→email confirmed. One email only (§0.5) |
| Settings persistence | ✅ | notification prefs set → persisted to `settings.json` → reverted (both directions) |
| Settings — password change | n/a | password protection is **operator-only and not configured** ("A jelszavas védelem nincs beállítva"); dashboard currently serves with no auth wall |
| Hub — report received | ✅ | `Received report from demo-felhom (11983 bytes)` every 15 min; `host-report (3 guests, 5 storage targets, 1 backups, 1 restore-tests, 9 pbs-snapshots)`; UI "Hub: Kapcsolódva" |
| Hub — DR recipe carries the PBS coord | ✅ | `dr_recipe` (demo-felhom) host-half `pbs={"repo_id":"felhom-pbs","namespace":"root","latest_snapshot_id":"9201"}`; app-half 15 apps; **no secrets in either half** |
| Hub — update panel | ✅ | version 0.73.0, "naprakész", last check timestamp shown |
| Hub GUI tiles (analytics/healthchecks) | skipped | no operator GUI creds — verified hub-side state via logs + DB instead |
The DR-recipe app-half + host-half are both stored at the hub each report (`DR-recipe …-half stored for
customer demo-felhom (v1)`), and the assembled recipe carries the live PBS coordinate (the agent v0.39.0
live-reporter feature) — verified directly in the hub DB.
---
## Phase 7 — Resilience
| Check | Result | Evidence |
|---|---|---|
| Guest 9201 reboot | ✅ | `pct reboot 9201`; recovered to 26→**25** containers, 0 unhealthy in ~2 min |
| Drives re-bind | ✅ | felhom-flash (sdc1) + felhom-usb (sdb1) remounted ext4 at correct mountpoints |
| FileBrowser convergence | ✅ | binds reconverged to `<drive>/userdata → /srv/<drive>` |
| Tunnel reconnect | ✅ | cloudflared QUIC + HTTP/2 prechecks to region1/2 argotunnel pass |
| Controller + apps + Phase-1/2 fixes persist | ✅ | komga + gitea still healthy post-reboot |
| Host reboot | ⏭ skipped (§7) | unattended risk — deferred to supervised |
Incidental: the controller's docker container IP changed across the reboot
(`172.18.0.10``172.18.0.8`); harmless to the system (traefik routes by docker DNS name), only required
re-pointing the test forwarder.
---
## Consolidated findings (for supervised follow-up)
1. **🔴 controller→agent local-API (8443) unreachable** (Phase 5) — gates storage UI + host-metrics API
+ whole-guest backup. Agent binds `192.168.0.162` not `0.0.0.0`; no firewall; container reaches
`:8006` but not `:8443` on the same host. **Highest priority.** Fix: bind `0.0.0.0`/bridge or
repoint `local_api.endpoint`.
2. **`GET /api/backup/snapshots` has no handler** (Phase 3) — restore UI snapshot dropdown can't
populate (restore works via `POST /backup/restore`).
3. **glance catalog template** never seeds `glance.yml` → crash-loop on deploy (Phase 2). Needs a
default-config seed (init step / entrypoint), not a trivial HC fix.
4. **Restore DB import is additive** (Phase 3) — doesn't drop tables absent from the dump. Minor.
5. **On-LAN HTTPS to dashboard/apps 404s** at traefik `:443` for all hosts (cert valid) — supported
path is Cloudflare; on-LAN HTTPS appears non-functional. Worth confirming the Cloudflare path is the
intended sole ingress.
6. Minor: a few catalog logo SVGs 404 (adventurelog, cloudflared, crafty-controller); dashboard
"Adatbázisok mentve" count flickered 3→2 post-reboot (cosmetic).
## Code shipped (app-catalog `main`, trunk)
- komga HC `/api/v1/actuator/health`(401) → `/actuator/health` (docker HC + `.felhom.yml` probe) —
`3faa5ae`, `9b066de`.
- gitea HC `/api/v1/version`(404 pre-install) → `/api/healthz`, start_period 90 s — `17e00b7`.
- CHANGELOG + REPORT updated. **No felhom-controller / felhom-agent code changed** (catalog-only).
## Cleanup confirmation
- Throwaway guest **9300 destroyed**; restore-test scratch 990000 auto-torn-down; `pct list` =
9201 (running) + pre-existing 9001/9999 (stopped, not created by this run).
- **No scratch loopback devices created** (destructive storage sub-phase skipped).
- Test apps removed (glance, gitea) → baseline app set.
- Notification settings reverted; password never set; drives never ejected.
- Test forwarder (`/tmp/felhom-test-fwd.py`) removed from the guest; temp `hub.db` copy removed from
DooPlex. **DooPlex used in normal roles only** (PBS target, gitea/hub reads).
- **komga healthy** (improvement over the unhealthy baseline).
## After-state vs before
9201 running; **25 containers Up, 0 unhealthy** (before: komga unhealthy); controller v0.73.0; tunnel up;
5 storage targets reachable. Net result: demo restored to baseline **+ komga fixed**.
## Deferred — needs a supervised run
- **Agent 8443 reachability** root-cause + fix (finding #1) — gates the whole agent-backed subsystem.
- **Storage lifecycle UI** (label edit, eject→re-enroll, init-wizard format/migrate/decommission on a
loopback) — re-run once finding #1 is fixed.
- **Host reboot** resilience (full host AC-recovery + remount-by-UUID under sdX reshuffle).
- **glance** config-seeding fix (finding #3); **`/api/backup/snapshots`** handler + UI snapshot
selection (finding #2).