Files
felhom-controller/TEST-REPORT.md
T

122 lines
12 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# TEST-REPORT — Unattended test campaign (N100 / guest 9201)
**Run start:** 2026-06-22 (CC, unattended)
**Operator:** Viktor (away, glancing remotely)
**Scope:** RUNBOOK phased test matrix. Only code change permitted = komga catalog HC fix.
**Legend:** PASS / FAIL / SKIPPED(reason). Evidence is raw (command output / volid / screenshot path).
---
## Baseline state (Phase 0 "before")
- **Host:** felhom-pve (`demo-felhom`), agent **v0.39.0** (service active).
- **Controller:** image `gitea.dooplex.hu/admin/felhom-controller:0.73.0`, container `Up 5 days (healthy)`.
- **`pct list`:** 9201 running (`demo-felhom`); 9001 stopped (`spike-lxc`); 9999 stopped (`felhom-selftest-scratch`).
- **Infra containers (9201):** felhom-controller, cloudflared, traefik, filebrowser — all Up 5 days, healthy.
- **App containers (9201):** 24 containers Up 18h. All healthy **except `komga` = unhealthy** (Phase-1 target).
- **Disk (9201):** rootfs `/` 32G (4% used); `/var/lib/docker` 252G (8%, 19G); `/mnt/felhom-drives` (pve-root) 94G (31%).
- **Drives:** `felhom-flash` sdc1 118G (~159M used) `/mnt/felhom-drives/felhom-flash`; `felhom-usb` sdb1 916G (~2M used) `/mnt/felhom-drives/felhom-usb`.
- **PBS storage:** `felhom-pbs` active, 9.6 TB total, 21% used. Nightly 9201 backups present (latest `2026-06-21T18:55:12Z`, ~19.5 GB).
- **pvesm:** felhom-flash, felhom-usb, felhom-pbs, local, local-lvm — all active.
---
## Phase 0 — Baseline + safety net
State: 9201 healthy, controller up, tunnel up (cloudflared Up 5 days).
| Check | Result | Evidence |
|---|---|---|
| Baseline capture | PASS | see above |
| GUI screenshots (before) | _pending_ | Chrome bridge check next |
| Fresh PBS whole-CT backup of 9201 | **PASS** | volid `felhom-pbs:backup/ct/9201/2026-06-22T12:45:37Z`, 19.62 GB, crash-consistent=true, 161.5s, success=true. uncovered_volumes (by design): /etc/felhom-bootstrap, /mnt/felhom-drives (drives backed up separately) |
| PBS verify of new snapshot | **PASS** | `verify_state: ok` for backup_id 9201 backup_time 2026-06-22T12:45:37Z (datastore felhom-spike, verify UPID present) |
**DR source for Phase 4:** `felhom-pbs:backup/ct/9201/2026-06-22T12:45:37Z` (verified ok). Gate to destructive phases: **OPEN**.
**GUI render smoke-test (before):** dashboard PASS (17 running / 0 stopped / 55 total; mem 5.8/12 GB; CPU 5%; 58°C; v0.73.0; app-backup "Még nem futott"), `/monitoring` PASS, `/stacks` PASS.
### Access path note (important finding)
- On-LAN, split-horizon DNS (resolver 192.168.0.250) maps `*.demo-felhom.eu` → guest 9201 (192.168.0.141), bypassing Cloudflare (no `cf-ray` on responses).
- **The guest's traefik `:443` (websecure) returns Go `404 page not found` for ALL hosts** (controller AND apps, e.g. romm), while serving a valid LE `*.demo-felhom.eu` cert. `:80` 301-redirects to https → so the browser dead-ends at the 404. → on-LAN HTTPS to the dashboard/apps is non-functional. (The supported external path is Cloudflare; couldn't validate it from the LAN due to DNS hijack.) **FINDING — flag for supervised review.**
- **Workaround used for UI testing (reversible):** a python TCP forwarder inside guest 9201 (`/tmp/felhom-test-fwd.py`, listens `0.0.0.0:18080` → controller `172.18.0.10:8080`). Browser → `http://felhom.demo-felhom.eu:18080/` (LAN DNS → guest:18080 → controller). `CatchAllMiddleware` strips the port so Host matches `felhom.<domain>` → dashboard serves over plain HTTP. **To be killed at cleanup.**
- The dashboard currently serves with **no auth wall** (`/` → 200, no login redirect, no session cookie) — no password configured/enforced. (Bears on Phase 6 "password change" step.)
### Scope reality vs runbook
- Runbook assumed "14 enabled apps"; catalog is now **55 total, 38 not-deployed**. Deploying all 38 unattended would exhaust the 12 GiB guest. **Phase 2 will deploy a representative sample** (a few light apps) + cite the already-deployed heavy multi-container case (immich, 4 containers healthy). Documented, not silently truncated.
**Phase 0 verdict: PASS — gate OPEN.**
---
## Phase 1 — komga HC fix + git-sync + health detection — **PASS**
State: 9201 healthy, controller up, tunnel up.
**Diagnosis (live, guest 9201):** docker HC `curl -f .../api/v1/actuator/health` → exit 22 (HTTP 401, auth-gated) → `unhealthy` for 18h while serving fine on :25600. Endpoint probe matrix (wget in image): `/`, `/actuator/health`, `/api/v1/oauth2/providers`, `/login` all 200; only `/api/v1/actuator/health` → 401. Image ships `curl` (`/usr/bin/curl`). **Found a second probe:** the controller's own `.felhom.yml` `healthcheck.checks[].path` (drives the UI badge + route publishing) also pointed at the 401 path.
| Check | Result | Evidence |
|---|---|---|
| Catalog HC fix (docker-compose.yml) | PASS | `/api/v1/actuator/health``/actuator/health`; commit `3faa5ae` pushed to `main` |
| Catalog probe fix (.felhom.yml) | PASS | controller-side probe path fixed; commit `9b066de` pushed to `main` |
| CHANGELOG + REPORT (app-catalog) | PASS | created `CHANGELOG.md` (newest-on-top) + overwrote `REPORT.md` (`f51f794`) |
| GUI git-sync ("Sablonok frissítése") | PASS | logs: `[sync] Updated komga/docker-compose.yml` + `komga/.felhom.yml`; `Sablonok frissítve — frissítve: komga`. Live stack `/opt/docker/stacks/komga/docker-compose.yml` updated to `/actuator/health` (content-hash copy verified) |
| GUI redeploy ("Frissítés") | PASS | `POST /api/stacks/komga/update` → 200; `docker compose up -d --remove-orphans`; "Stack komga updated successfully (4.2s)"; container recreated (StartedAt 2026-06-22T13:09:50Z) |
| komga reaches healthy (docker) | PASS | `docker ps`: `komga Up (healthy)`; HC now `http://localhost:25600/actuator/health` |
| Controller health-state machine | PASS | controller log `Health probes: N ok (of N probed)` (no more komga WARN); card badge `Nem egészséges`**Fut** |
| Route publishing | PASS | "⚠ URL nem elérhető útvonal nincs publikálva" warning cleared; `komga.demo-felhom.eu ↗` link shown |
Minor cosmetic note (not fixed): a few catalog logo SVGs 404 (`adventurelog-logo.svg`, `cloudflared-logo.svg`, `crafty-controller-logo.svg`).
## Phase 2 — App deployment sweep — **PASS (with 1 logged catalog defect)**
State: 9201 healthy, controller up, tunnel up. Scope: representative sample (see Phase-0 scope note) — catalog has 38 not-deployed apps; deploying all on a 12 GiB guest is unsafe.
| App | Result | Evidence |
|---|---|---|
| **gitea** (deploy via GUI) | PASS (after HC fix) | GUI deploy page → memory-gate bar (`6043/11904 MB 50%`, two-segment current+new) → 3-step progress panel (✅ Konfiguráció mentve → ✅ Konténerek elindultak → ⏳ inicializálás) → `POST /api/stacks/gitea/update` 200. Initially **unhealthy**: HC `curl -f /api/v1/version` → 404 (404 until install-lock) while `/api/healthz` → 200. **Fixed catalog** (commit `17e00b7` + CHANGELOG `1a44c14`), re-synced, redeployed → `docker inspect`: `healthy`, HC now `/api/healthz`; endpoint `http://127.0.0.1:3000/``HTTP/1.1 200 OK`. |
| **glance** (deploy via GUI) | **FAIL — logged, not fixed** | GUI deploy → 3-step panel OK, but container `Restarting (1)` crash-loop. Logs: `parsing config: reading main YAML file: open /app/config/glance.yml: no such file or directory`. **Root cause:** catalog template mounts an empty `glance_config:/app/config` named volume but never seeds the required `glance.yml`; glance refuses to start without it. **Not a trivial HC/start_period fix** (needs a seeded default config asset + wiring) → logged per runbook, not fixed unattended. Cleaned up: `POST .../stop` + `POST .../remove` → container gone. |
| **immich** (heavy multi-container, pre-deployed) | PASS (cited) | 4 containers (server/ML/postgres/redis) all `Up (healthy)` — the heavy-app coverage case. |
**Bonus coverage:** stop + remove endpoints exercised (glance teardown). **Catalog defects found this phase:** gitea HC (fixed), glance missing-config (logged for supervised fix).
## Phase 3 — Backup + restore — **PASS (2 findings)**
State: 9201 healthy, controller up, tunnel up.
| Check | Result | Evidence |
|---|---|---|
| Manual DB-dump backup ("Mentés most") | PASS | `POST /api/backup/run` → 3 DBs dumped: romm-mariadb 38.5 KB/14 tbl, paperless-postgres 360 KB/67 tbl, immich-postgres 43.2 MB/60 tbl; 43.6 MB total in 2.18s; `status.success=true`, `last_run 2026-06-22T13:27:26Z`. Files on disk fresh **13:28**: `…/primary/immich/db-dumps/immich-postgres.sql` (45.3 MB), `…/primary/romm/db-dumps/romm-mariadb.sql` (39 KB) |
| Cross-drive copy (Tier-2) | PASS | `POST /api/backup/tier2` → 8 HDD apps copied to `/mnt/sys_drive/felhom-data/backups/secondary/<app>` (immich 155.3 MB verified 156 MB on disk; romm has appdata + recovery-unit{db-dump,manifest,compose}). `crossdrive_completed` events pushed |
| Backup page reflects status | PASS | page shows "Adatmentés aktív", 3 DBs, schedule 02:30, last DB backup timestamp; status API success=true |
| Restore flow (per-app) | PASS | `POST /backup/restore stack_name=romm``RestoreFromRecoveryUnit`: compose down → redeploy from recovery unit (6 env, 3 encrypted) → "Imported DB dump romm-mariadb.sql into romm-db" → containers up → "Restore-from-unit completed: romm" |
| Restore is non-hollow | PASS | post-restore romm DB has real `platforms` table + 17 tables; all 3 romm containers `Up (healthy)`. Marker-table test: planted `restore_test_marker`(id=42) pre-restore → **survived** restore (see finding) but real snapshot tables were reimported; marker dropped afterward (DB back to 17-table baseline) |
**Findings:**
1. **`GET /api/backup/snapshots` has no server handler** (api-router returns "endpoint not found"). The restore UI's snapshot dropdown (`backups.html:682`) can't populate → a pure-UI restore is blocked at snapshot selection. Restore itself works via `POST /backup/restore` (the form's action). **Flag for fix.**
2. **Restore DB import is additive**`RestoreFromRecoveryUnit` replays the dump but does not drop tables absent from it (my marker table persisted). Real snapshot data is restored; stale/extra tables are not pruned. Minor behavioral note.
3. Whole-server ("Rendszermentés") backup shows **"A host-ügynök jelenleg nem elérhető"** (agent local API unreachable from controller) — same root as the monitoring host-metrics gap; tracked in Phase 6.
## Phase 4 — DR rebuild from backup — **PASS** (non-destructive to 9201)
State: 9201 healthy throughout, controller up, tunnel up (verified before + after).
| Check | Result | Evidence |
|---|---|---|
| Agent DR primitive (`--selftest=restore-test`) | PASS | restored `felhom-pbs:backup/ct/9201/2026-06-22T12:45:37Z` into scratch **990000** (band 990000-990009, no collision) → bind-mounts neutralized → **net link-down** → boot → `verified: boot+running` → torn down. `pass:true`, 2m56s, self-cleaning |
| Deep DR restore into throwaway **9300** | PASS | `pct restore 9300 <archive> --storage local-lvm`. **Safety:** before start, removed shared binds `mp3:/mnt/felhom-drives` + `mp9:/etc/felhom-bootstrap`, removed `net0`, set `onboot 0` — verified config clean (only restored `mp0` docker-data + rootfs remain). |
| 9300 boots + controller process up | PASS | `pct status 9300: running`; `felhom-controller Up (healthy)` inside |
| Restored volumes + app.yaml + stacks intact | PASS | `/opt/docker/stacks/` has all 57 stack dirs; romm `app.yaml` restored (`deployed:true`, env, locked_fields) |
| `docker ps` inside 9300 | PASS | **25 containers up** from restored volumes — identical set to 9201 (filebrowser, vikunja, uptime-kuma, romm×3, radarr, privatebin, paperless×3, mealie, komga, jellyfin, immich×4, gokapi, calibre-web, audiobookshelf, actualbudget, felhom-controller healthy, cloudflared, traefik) |
| Restored data non-hollow | PASS | 9300 `romm-db`: 17 tables (the real restored DB) |
| No collision with live 9201 | PASS | 9300 has **no eth0** (net removed) → cloudflared cannot reach Cloudflare/hub; live 9201 `cloudflared Up 5 days` unchanged throughout |
| Cleanup | PASS | `pct stop 9300 && pct destroy 9300` → both LVs removed; `pct list` shows only 9001/9201/9999; local-lvm freed to 9.71% |
**Core DR claim proven end-to-end:** PBS whole-CT restore → guest boots → controller + all app stacks + DB data recovered from restored volumes — twice (agent self-cleaning primitive + manual deep inspection), both collision-safe per §0.6.
## Phase 5 — Storage lifecycle — _pending_
## Phase 6 — Monitoring/alerts/hub — _pending_
## Phase 7 — Resilience (optional) — _pending_
---
## Feature matrix (filled at end)
## Cleanup confirmation (filled at end)
## Deferred — needs supervised run (filled at end)