diff --git a/documentation/tests/unattended-test-campaign-2026-06-22-findings.md b/documentation/tests/unattended-test-campaign-2026-06-22-findings.md new file mode 100644 index 0000000..1e42129 --- /dev/null +++ b/documentation/tests/unattended-test-campaign-2026-06-22-findings.md @@ -0,0 +1,262 @@ +# Unattended test campaign — full-feature validation (demo guest 9201) — 2026-06-22 + +> **Status: empirical findings from an unattended live test campaign (2026-06-22).** Operator away, +> glancing remotely. Goal: maximize feature coverage across the running demo (Felhom on Proxmox: +> host agent `felhom-agent v0.39.0` on `felhom-pve`, in-guest controller `felhom-controller v0.73.0` +> in LXC 9201, hub `felhom-hub` on k3s/DooPlex) **without leaving the demo unrecoverably broken.** +> Only code change permitted: the komga catalog healthcheck fix (+ one analogous gitea HC fix the +> sweep surfaced, which the runbook sanctions). No controller/agent build. All destructive DR work +> targeted throwaway guests; the live demo was restored to baseline (and is healthier — komga fixed). +> +> Companion live report (per-check raw evidence, continuously updated during the run): +> `felhom-controller/TEST-REPORT.md`. This document is the consolidated findings writeup. + +## Topology under test + +- **Host / agent:** `felhom-pve` (`demo-felhom`, PVE 9.2.2, Intel N100, 16 GB), `felhom-agent v0.39.0` + (systemd active). SSH `root@felhom-pve`. +- **Guest / controller:** LXC **9201** (`demo-felhom`, 12 GB, 32 GB rootfs + 256 GB `/var/lib/docker` + vol), `felhom-controller:0.73.0` (docker container, no auth wall currently). Customer id `demo-felhom`. +- **Hub:** `felhom-hub` on k3s @ DooPlex `192.168.0.180`; `hub.felhom.eu`. +- **Backup target:** PBS `felhom-pbs` (datastore `felhom-spike` on DooPlex), 9.6 TB, 21 % used. +- **Drives (enrolled):** `felhom-flash` (USB, `uuid:81a26531…`, `/mnt/felhom-drives/felhom-flash`, + 117 GB, holds app userdata for 8 apps) · `felhom-usb` (local-dir, `uuid:da9e7089…`, + `/mnt/felhom-drives/felhom-usb`, 916 GB, empty) · `local-lvm` · `local` · `felhom-pbs`. +- **Baseline apps:** 14 deployed app stacks (25 containers incl. 4 infra: controller, traefik, + cloudflared, filebrowser); 55 in catalog, 38 installable. komga `unhealthy` at baseline. + +### Access note — how the UI was driven +On the LAN, split-horizon DNS (resolver `192.168.0.250`) maps `*.demo-felhom.eu` → guest 9201 +(`192.168.0.141`), bypassing Cloudflare. The guest's traefik **`:443` (websecure) returns Go +`404 page not found` for every host** (controller and apps alike) while serving a valid LE +`*.demo-felhom.eu` cert; `:80` 301-redirects to https → so a browser dead-ends at the 404 on-LAN. +**Workaround:** a short-lived python TCP forwarder inside 9201 (`0.0.0.0:18080` → controller +`:8080`); the browser hit `http://felhom.demo-felhom.eu:18080/` (LAN DNS → guest:18080), +and the controller's `CatchAllMiddleware` strips the port so the Host matches `felhom.` and the +dashboard serves over plain HTTP. Forwarder removed at end of run. The controller dashboard is +reachable internally at traefik file-route `Host(felhom.demo-felhom.eu) → felhom-controller:8080`. + +--- + +## Phase 0 — Baseline + PBS safety-net backup + +| Check | Result | Evidence | +|---|---|---| +| Baseline capture | ✅ | agent v0.39.0; ctrl v0.73.0; 25 containers, komga unhealthy; drives + pvesm healthy | +| Fresh PBS whole-CT backup of 9201 | ✅ | `felhom-agent --selftest=backup -vmid=9201` → `felhom-pbs:backup/ct/9201/2026-06-22T12:45:37Z`, 19.62 GB, crash-consistent, 161 s. Uncovered (by design): `/etc/felhom-bootstrap`, `/mnt/felhom-drives` (drives backed up separately) | +| PBS verify | ✅ | `--selftest=pbs-verify` → `verify_state: ok` for the new snapshot | + +The verified backup is the DR source for Phase 4. **Gate to destructive phases: opened.** + +--- + +## Phase 1 — komga healthcheck fix + git-sync + health state machine + +**Diagnosis (live):** the docker HC `curl -f http://localhost:25600/api/v1/actuator/health` exited 22 +(HTTP **401** — the `/api/v1` prefix is auth-gated) → `unhealthy` for 18 h while komga served fine. +Endpoint probe matrix (wget inside the image): `/`, `/actuator/health`, `/api/v1/oauth2/providers`, +`/login` → all 200; only `/api/v1/actuator/health` → 401. **Komga's actuator is served unauthenticated +at `/actuator/health`, off the `/api/v1` API prefix.** A second probe matters too: the controller's own +`.felhom.yml` `healthcheck.checks[].path` (which drives the dashboard badge + route publishing) pointed +at the same 401 path. The `gotson/komga:1.20.0` image ships `curl` (verified). + +**Fix (app-catalog `main`):** repoint **both** probes to `/actuator/health` — docker HC +(`3faa5ae`) and `.felhom.yml` (`9b066de`). + +**Validation (end-to-end through the real UI + server pipeline):** +- GUI "Sablonok frissítése" → controller git-sync: `[sync] Updated komga/docker-compose.yml` + + `komga/.felhom.yml`; content-hash copy landed in `/opt/docker/stacks/komga/`. +- GUI "Frissítés" → `POST /api/stacks/komga/update` 200 → `docker compose up -d --remove-orphans` + recreated the container (4.2 s). +- Result: `docker ps` → `komga Up (healthy)`, HC now `/actuator/health`; controller log + `Health probes: N ok`; dashboard badge `Nem egészséges` → **Fut**; the + "⚠ URL nem elérhető – útvonal nincs publikálva" route warning cleared. + +This one phase exercised: catalog fix, git-sync (content-hash copy), redeploy, the health-aware state +machine, and route publishing. + +--- + +## Phase 2 — App deployment sweep + +Scope note: catalog is **55 apps, 38 not deployed** (the runbook's "14 enabled" was stale). Deploying +all 38 would exhaust the 12 GB guest, so a representative sample was deployed via the GUI; the heavy +multi-container case (immich, 4 containers) is already deployed + healthy (cited). + +| App | Result | Evidence | +|---|---|---| +| **gitea** (GUI deploy) | ✅ after HC fix | GUI deploy page → memory-gate bar (two-segment current+new) → 3-step progress panel → containers up. Initially `unhealthy`: HC `curl -f /api/v1/version` → **404** (404 until install-lock) while `/api/healthz` → 200. **Fixed catalog** (`17e00b7`, start_period 30→90 s) → re-sync → redeploy → `healthy`; endpoint `:3000/` → 200 | +| **glance** (GUI deploy) | ❌ logged, not fixed | crash-loop `Restarting (1)`. Log: `open /app/config/glance.yml: no such file or directory`. Catalog template mounts an empty `glance_config` named volume but **never seeds the required `glance.yml`**. Not a trivial HC fix → logged. Removed cleanly (`stop`+`remove`) | +| **immich** (pre-deployed) | ✅ cited | 4 containers healthy = heavy-app coverage | + +Bonus coverage: the stop + remove endpoints (glance + gitea teardown). Both test deployments were +removed at cleanup to restore the baseline app set. + +--- + +## Phase 3 — Backup + restore (per-app) + +| Check | Result | Evidence | +|---|---|---| +| Manual DB-dump backup | ✅ | `POST /api/backup/run` → 3 DBs: romm-mariadb (14 tbl, 38.5 KB), paperless-postgres (67 tbl, 360 KB), immich-postgres (60 tbl, 43.2 MB); 43.6 MB / 2.18 s. On-disk fresh (`…/primary//db-dumps/*.sql`) | +| Tier-2 cross-drive copy | ✅ | `POST /api/backup/tier2` → 8 HDD apps → `/mnt/sys_drive/felhom-data/backups/secondary/` (immich 156 MB on disk; romm appdata + recovery-unit{db-dump,manifest,compose}); `crossdrive_completed` events pushed to hub | +| Per-app restore | ✅ | `POST /backup/restore stack_name=romm` → `RestoreFromRecoveryUnit`: compose down → redeploy (6 env, 3 encrypted recovered) → `Imported DB dump romm-mariadb.sql` → containers up → "Restore-from-unit completed" | +| Restore non-hollow | ✅ | planted a marker table pre-restore; post-restore romm DB has the real `platforms` + 17 tables, 3 containers healthy; marker dropped afterward to restore baseline | + +**Findings:** +- **`GET /api/backup/snapshots` has no server handler** (the api-router returns "endpoint not found"). + The restore UI's snapshot dropdown (`backups.html:682`) therefore cannot populate → a **pure-UI + restore is blocked at snapshot selection.** The restore itself works via the form action + `POST /backup/restore` (which calls `RestoreFromRecoveryUnit(stack)` and ignores `snapshot_id`'s value). +- **Restore DB import is additive** — `RestoreFromRecoveryUnit` replays the dump but does not drop + tables absent from it (the planted marker table survived). Snapshot data is restored; stale/extra + tables are not pruned. Minor behavioral note. + +--- + +## Phase 4 — DR rebuild from backup (non-destructive to 9201) + +Proven **twice**, both collision-safe (no tunnel/hub identity clash with the live 9201): + +1. **Agent DR primitive** — `felhom-agent --selftest=restore-test -archive=`: restored + into scratch **990000** → bind-mounts neutralized → **net link-down** → boot → `verified: boot+running` + → torn down. `pass:true`, 2 m 56 s, self-cleaning. +2. **Deep manual restore into throwaway 9300** — `pct restore 9300 --storage local-lvm`. + **Safety before start:** removed the shared binds `mp3:/mnt/felhom-drives` + `mp9:/etc/felhom-bootstrap`, + removed `net0`, set `onboot 0` (kept only the restored `mp0` docker-data + rootfs — independent + copies). Then started and inspected: + - `pct status: running`; `felhom-controller Up (healthy)` inside. + - `/opt/docker/stacks/` had all 57 stack dirs; romm `app.yaml` restored (`deployed:true`, env, locked_fields). + - **`docker ps` inside 9300: 25 containers up** from restored volumes — identical set to 9201. + - Restored data non-hollow: 9300 `romm-db` has 17 tables. + - No collision: 9300 had **no eth0**; live 9201 `cloudflared Up` unchanged throughout. + - Cleanup: `pct stop 9300 && pct destroy 9300` (both LVs removed); local-lvm freed. + +**Core DR claim proven end-to-end:** PBS whole-CT restore → guest boots → controller + all app stacks + +DB data recovered from restored volumes. The bind-mount-neutralization recipe (remove host-path binds + +`net0` + `onboot` before start) is the safe way to inspect a restored customer image alongside a live +one. + +--- + +## Phase 5 — Storage lifecycle — **BLOCKED by a controller→agent connectivity defect** + +### 🔴 FINDING #1 (highest priority) — controller cannot reach the agent local API (8443) +The controller container's calls to the agent local API fail: `agentapi: GET /disks: dial tcp +192.168.0.162:8443: connect: cannot assign requested address` (EADDRNOTAVAIL, persistent — 6/6). The +same error surfaces in the Settings UI ("Meghajtók (ügynök nézet)") and in the backup page +("A host-ügynök jelenleg nem elérhető"). + +Reachability matrix **from the controller container**: +- `192.168.0.162:8006` (pveproxy, binds `0.0.0.0`) → **200** (reachable) +- `192.168.0.162:8443` (felhom-agent local API) → **refused / EADDRNOTAVAIL** +- `192.168.0.141:443` (guest's own traefik) → 404 (reachable) + +From the guest **host namespace**, `192.168.0.162:8443` is OPEN. Confirmed: `ss` shows the agent +`LISTEN 192.168.0.162:8443` — **bound to the specific host LAN IP, not `0.0.0.0`** like pveproxy; +**no iptables rule** references 8443; **pve-firewall disabled**. The agent itself is healthy +(`--selftest=storage` enumerates all 5 targets with SMART/FS/durable_id fine). So this is a +**container→host-IP routing / source-bind mismatch on port 8443**, not a firewall and not an agent +crash. + +**Impact — the entire agent-backed feature set is down in this demo:** storage management UI +(`/api/disks`, `/api/storage`: scan/label/eject/enroll/format/migrate/decommission), the live +host-metrics API (`/api/host-metrics`), and whole-guest backup trigger. (Monitoring still shows host +CPU/mem/temp — those are read from `/proc`, since the LXC shares the host kernel, not via the agent.) + +**Fix direction (supervised):** bind the agent local API to `0.0.0.0` (or the `vmbr0` bridge address the +guest routes through), or repoint the controller's `local_api.endpoint` (bootstrap.json) at a +container-reachable address; then re-run Phase 5's mutating checks. Not changed in this run (network +reconfig is out of scope for an unattended test). + +| Check | Result | Evidence | +|---|---|---| +| Scan / health / FS / model / durable_id (agent tier) | ✅ | `--selftest=storage`: 5 targets; felhom-usb SMART **PASSED** temp=35 poh=2736 realloc=0; felhom-flash uuid:81a26531; thin-pool data=9.7 % | +| Label edit + revert (UI) | ⛔ blocked | controller→agent down (finding #1) | +| Eject → re-enroll cycle (UI) | ⛔ blocked | same | +| Destructive loopback (scan→format→mount→migrate→decommission) | ⛔ skipped | wizard drives the agent via the broken path; agent-direct needs the per-guest token + TLS-pin reconstruction → high-effort/low-confidence unattended. **No scratch loopback created** (nothing to clean up). Deferred to supervised. | + +--- + +## Phase 6 — Monitoring / alerts / notifications / settings / hub + +| Check | Result | Evidence | +|---|---|---| +| Monitoring — host metrics | ✅ | CPU/load/mem (7.6/15.4 GB — host's 16 GB via `/proc`/LXC kernel-share)/temp/uptime | +| Monitoring — container/app metrics | ✅ | per-container memory breakdown (26 containers), per-app CPU/MEM charts, time-range charts 1h–30d | +| Alerts — state-based | ✅ | komga `0 ok, 1 unhealthy` before → `N ok` after the Phase-1 fix; badge red→green | +| Notification path (exercised ONCE) | ✅ | set email+events+cooldown → hub `Notification preferences updated for demo-felhom`. Test → hub **`Event from demo-felhom: test` → `Test email sent to nagyfenyvesi.viktor@gmail.com`** = controller→hub→Resend→email confirmed. One email only (§0.5) | +| Settings persistence | ✅ | notification prefs set → persisted to `settings.json` → reverted (both directions) | +| Settings — password change | n/a | password protection is **operator-only and not configured** ("A jelszavas védelem nincs beállítva"); dashboard currently serves with no auth wall | +| Hub — report received | ✅ | `Received report from demo-felhom (11983 bytes)` every 15 min; `host-report (3 guests, 5 storage targets, 1 backups, 1 restore-tests, 9 pbs-snapshots)`; UI "Hub: Kapcsolódva" | +| Hub — DR recipe carries the PBS coord | ✅ | `dr_recipe` (demo-felhom) host-half `pbs={"repo_id":"felhom-pbs","namespace":"root","latest_snapshot_id":"9201"}`; app-half 15 apps; **no secrets in either half** | +| Hub — update panel | ✅ | version 0.73.0, "naprakész", last check timestamp shown | +| Hub GUI tiles (analytics/healthchecks) | skipped | no operator GUI creds — verified hub-side state via logs + DB instead | + +The DR-recipe app-half + host-half are both stored at the hub each report (`DR-recipe …-half stored for +customer demo-felhom (v1)`), and the assembled recipe carries the live PBS coordinate (the agent v0.39.0 +live-reporter feature) — verified directly in the hub DB. + +--- + +## Phase 7 — Resilience + +| Check | Result | Evidence | +|---|---|---| +| Guest 9201 reboot | ✅ | `pct reboot 9201`; recovered to 26→**25** containers, 0 unhealthy in ~2 min | +| Drives re-bind | ✅ | felhom-flash (sdc1) + felhom-usb (sdb1) remounted ext4 at correct mountpoints | +| FileBrowser convergence | ✅ | binds reconverged to `/userdata → /srv/` | +| Tunnel reconnect | ✅ | cloudflared QUIC + HTTP/2 prechecks to region1/2 argotunnel pass | +| Controller + apps + Phase-1/2 fixes persist | ✅ | komga + gitea still healthy post-reboot | +| Host reboot | ⏭ skipped (§7) | unattended risk — deferred to supervised | + +Incidental: the controller's docker container IP changed across the reboot +(`172.18.0.10`→`172.18.0.8`); harmless to the system (traefik routes by docker DNS name), only required +re-pointing the test forwarder. + +--- + +## Consolidated findings (for supervised follow-up) + +1. **🔴 controller→agent local-API (8443) unreachable** (Phase 5) — gates storage UI + host-metrics API + + whole-guest backup. Agent binds `192.168.0.162` not `0.0.0.0`; no firewall; container reaches + `:8006` but not `:8443` on the same host. **Highest priority.** Fix: bind `0.0.0.0`/bridge or + repoint `local_api.endpoint`. +2. **`GET /api/backup/snapshots` has no handler** (Phase 3) — restore UI snapshot dropdown can't + populate (restore works via `POST /backup/restore`). +3. **glance catalog template** never seeds `glance.yml` → crash-loop on deploy (Phase 2). Needs a + default-config seed (init step / entrypoint), not a trivial HC fix. +4. **Restore DB import is additive** (Phase 3) — doesn't drop tables absent from the dump. Minor. +5. **On-LAN HTTPS to dashboard/apps 404s** at traefik `:443` for all hosts (cert valid) — supported + path is Cloudflare; on-LAN HTTPS appears non-functional. Worth confirming the Cloudflare path is the + intended sole ingress. +6. Minor: a few catalog logo SVGs 404 (adventurelog, cloudflared, crafty-controller); dashboard + "Adatbázisok mentve" count flickered 3→2 post-reboot (cosmetic). + +## Code shipped (app-catalog `main`, trunk) +- komga HC `/api/v1/actuator/health`(401) → `/actuator/health` (docker HC + `.felhom.yml` probe) — + `3faa5ae`, `9b066de`. +- gitea HC `/api/v1/version`(404 pre-install) → `/api/healthz`, start_period 90 s — `17e00b7`. +- CHANGELOG + REPORT updated. **No felhom-controller / felhom-agent code changed** (catalog-only). + +## Cleanup confirmation +- Throwaway guest **9300 destroyed**; restore-test scratch 990000 auto-torn-down; `pct list` = + 9201 (running) + pre-existing 9001/9999 (stopped, not created by this run). +- **No scratch loopback devices created** (destructive storage sub-phase skipped). +- Test apps removed (glance, gitea) → baseline app set. +- Notification settings reverted; password never set; drives never ejected. +- Test forwarder (`/tmp/felhom-test-fwd.py`) removed from the guest; temp `hub.db` copy removed from + DooPlex. **DooPlex used in normal roles only** (PBS target, gitea/hub reads). +- **komga healthy** (improvement over the unhealthy baseline). + +## After-state vs before +9201 running; **25 containers Up, 0 unhealthy** (before: komga unhealthy); controller v0.73.0; tunnel up; +5 storage targets reachable. Net result: demo restored to baseline **+ komga fixed**. + +## Deferred — needs a supervised run +- **Agent 8443 reachability** root-cause + fix (finding #1) — gates the whole agent-backed subsystem. +- **Storage lifecycle UI** (label edit, eject→re-enroll, init-wizard format/migrate/decommission on a + loopback) — re-run once finding #1 is fixed. +- **Host reboot** resilience (full host AC-recovery + remount-by-UUID under sdX reshuffle). +- **glance** config-seeding fix (finding #3); **`/api/backup/snapshots`** handler + UI snapshot + selection (finding #2).