Files
felhom-controller/TEST-REPORT.md
T

79 lines
11 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# TEST-REPORT — Test campaign #2 (broader apps, edge cases, break-and-recover) — N100 / guest 9201
**Run start:** 2026-06-22 (CC, unattended). **Operator:** Viktor (away).
**Scope:** push past campaign #1 — more apps, edge cases, chaos break-and-recover. Storage lifecycle now
unblocked (ctrl v0.74.0 fixed the agent leak). Only code changes: trivial catalog HC/config-seed fixes.
**Legend:** PASS / FAIL / SKIPPED(reason). Evidence raw (docker ps / ss / curl / screenshot).
(Campaign #1 report preserved in git history + `felhom.eu/documentation/tests/unattended-test-campaign-2026-06-22-*.md`.)
---
## Phase 0 — Baseline + safety net + leak-fix gate — **PASS (gate OPEN)**
- **Baseline:** ctrl `v0.74.0` healthy, agent `v0.39.0`; `pct list` = 9201 running (+ 9001/9999 stopped, pre-existing); **25 containers, 0 unhealthy**; rootfs `/` 4%, `/var/lib/docker` 8%, drives felhom-flash 1% / felhom-usb 1%.
- **Master safety net + DR source:** PBS backup `felhom-pbs:backup/ct/9201/2026-06-22T16:13:44Z`, success=true, crash-consistent; PBS verify shows the snapshot present.
- **Leak-fix gate (the #1 unblock):** `/api/disks` lists felhom-flash/felhom-usb/local; `/api/host-metrics` returns `cpu_percent`; controller→agent idle sockets to `:8443` **flat at 2 before and after a 60-call burst**. Agent path fully working. **Gate to chaos/destructive phases: OPEN.**
## Phase 1 — Broader app sweep + edge cases — **PASS (with noted non-triggerables)**
| Check | Result | Evidence |
|---|---|---|
| Deploy navidrome (HDD-bound, Go) via GUI | PASS | GUI deploy → 3-step panel → `running/healthy` in 8s; `:4533` → HTTP 302 (serving); HDD binding `…/felhom-flash/userdata/media/music → /music` |
| Pattern coverage | PASS (cited) | JVM=komga (healthy), Python+DB=immich/paperless (healthy), Go=navidrome/gitea, HDD-bound=navidrome/romm/immich, optional-config=romm |
| Edge: rapid deploy→stop→redeploy race | PASS | stop → immediate update → `navidrome Up (healthy)`, `app.yaml deployed:true`, API `state=running deployed=true` — no stale button / no Deployed-flag race |
| Edge: slow-DB-migration start_period | PASS (cited) | mealie deployed healthy (start_period 60s covers migration — no premature unhealthy flap; lesson #7) |
| Edge: missing required config → graceful signal | PASS (cited #1) | campaign-1 glance: crash-loop with a clear logged error (`open /app/config/glance.yml: no such file`), controller marks it — not a silent hang |
| Edge: mem-gate hard-block | PARTIAL | the deploy-page memory projection renders correctly (e.g. navidrome 4065→+50 MB / 11904); a RAM hard-block isn't naturally triggerable (no single catalog app's request > ~7.6 GB free without mass-filling RAM). Disk-headroom hard-refuse (`GetDockerVolumeHeadroom`) is tested in Phase 3/C4. |
| Edge: Host()/port overlap | N/A (noted) | each catalog app has a unique subdomain by construction; an overlap isn't naturally producible without editing a template — not triggered |
navidrome left deployed (used as the drive-backed app for Phase-3 C6); removed at cleanup.
## Phase 2 — Storage lifecycle (the campaign-#1 blocker, NOW UNBLOCKED) — **PASS (destructive-format skipped, documented)**
Agent path works (v0.74.0), so these ran through the real controller→agent pipeline.
| Check | Result | Evidence |
|---|---|---|
| Scan / observe / health / FS / role / durable-id | PASS | `/api/disks`: 5 targets — felhom-flash (user-data, uuid:81a26531), felhom-usb (user-data, uuid:da9e7089), local/local-lvm (system), felhom-pbs (backup) with capacities + data_bearing |
| Label edit → revert (live) | PASS | felhom-usb label → "TEST-LABEL-c2" (log "Storage label updated") → reverted to exact UTF-8 original "Tárhely (felhom-usb)" (byte-verified). [my shell mangled the accent on first revert; corrected via `%C3%A1`] |
| **Disconnect → reconnect (reversible four-state)** | PASS | enrolled→**disconnected** (`/api/storage/disconnect`: unmounted, no apps) → **enrolled** (`/api/storage/reconnect`: remounted `/dev/sdb1`, disconnected mark cleared) |
| **Decommission → re-enroll (non-destructive)** | PASS | `/api/storage/decommission mode=anyway` (type-to-confirm `felhom-usb`): unmounted, `decommissioned=True` soft-marker, **fs intact — uuid da9e7089 preserved, NOT formatted**. `/api/storage/reconnect`: `reenrolled=true`, MOUNTED, marker cleared, label preserved |
| Destructive format/wipe on **loopback scratch** | **BLOCKED (documented)** | created `/dev/loop0` (truncate 2G + losetup); `/api/storage/init` → agent **`refusing to operate on non-whitelisted block device "/dev/loop0"`** — a deliberate device-whitelist safety control. So the clean scratch target the runbook prescribed is unavailable. Per §0.3's escape clause (no clean scratch target → skip), the destructive *format* was not run on the live felhom-usb either. Loopback torn down. |
| Migrate (data) | DEFERRED | no safe scratch target (loopback blocked); a live migrate would move a real app's data — not run unattended. `internal/stacks/migrate.go` has unit coverage. |
**Findings:**
1. **Agent device whitelist** refuses non-whitelisted block devices (loopback) — good security, but means a loopback can't serve as the destructive-disk scratch target the runbook assumed. Destructive format/migrate now need either a real spare disk or a supervised run.
2. **FileBrowser sync** logs `mkdir …/userdata: permission denied` when ensuring the skeleton on a just-unmounted/decommissioned path (recurs on disconnect + decommission). Cosmetic (the drive is intentionally unmounted), but noisy — could gate the skeleton-ensure on "is the path a live mountpoint."
## Phase 3 — Break-and-recover chaos (C1C8) — **PASS (7/8 run, C4 skipped; 2 findings)**
| # | Break → Detect → Recover → Verify | Result | Evidence |
|---|---|---|---|
| C1 | `docker kill navidrome` → controller detects → restart via UI → healthy | **PASS** | killed → `Exited(137)`; controller state `stopped`; `POST /api/stacks/navidrome/restart``running/healthy`, state `running` |
| C2 | `docker kill felhom-controller` → (expected bootstrap self-heal) → recover → verify | **PASS (w/ 🔴 finding)** | killed → **stayed Exited 48s, NO auto-recovery** (docker `unless-stopped` didn't fire; bootstrap.service is oneshot/exited). Data plane unaffected (apps+cloudflared+traefik up). Recovered via `systemctl restart felhom-controller-bootstrap.service` → running/healthy; UI 200; 26 containers/0 unhealthy. **Finding: a crashed controller has NO automatic recovery.** |
| C3 | stop `felhom-agent` → controller degrades gracefully → restart agent → immediate success | **PASS** | agent down → `/api/disks` **fast ~98ms clean error** ("connection refused"), no hang/crash; agent restarted → `/api/disks` OK on the **first** try (4/4), no stale-conn errors; idle sockets bounded at 1 — the v0.74.0 client's restarted-agent drain works |
| C4 | fill a drive near-full → disk warning → clear | **SKIPPED(impractical)** | drives are 118 GB / 916 GB and `/var/lib/docker` 252 GB; reaching the 90% warning / the deploy gate's `max(5GB,10%)`≈25 GB-free needs 100+ GB of fill — unsafe/slow unattended. Gate logic `GetDockerVolumeHeadroom` (hard-refuse deploy below the reserve) is present; monitoring reports per-drive usage. |
| C5 | rename `navidrome/app.yaml` away → rescan → restore | **PASS** | rescan after removal: controller stays `running/healthy` (no crash), container keeps running (data plane), card → `state=running deployed=false` (graceful orphan). Restored + rescan → `state=running deployed=true` |
| C6 | disconnect felhom-flash (9 apps) → recover | **PASS (recovered) + finding** | disconnect stopped all 9 apps + unmounted (`UNMOUNTED`, komga `stopped`); reconnect → felhom-flash remounted `/dev/sdc1`, all 9 restarted → **26 running, 0 unhealthy** after start_period. **Finding:** during the drive-absent window the controller's userdata-belt/recreate tried to bring drive-backed apps up while the mount was gone → `mkdir …/userdata: permission denied` + transient `Created` states (rather than a clean held "missing-storage" stop). Reconnect cleared it; but the disconnect→recreate interaction is noisy — worth a guard ("don't recreate/ensure-userdata when the drive is absent"). |
| C7 | kill controller MID-quiesce → Recover resumes | **PASS** | triggered whole-guest backup → quiesce marker written **upfront** (`active:true`, 15 stopped_stacks) → `docker kill felhom-controller` mid-quiesce → **marker persisted across the crash**; controller restart logged `[quiesce] crash recovery: a quiesce was in progress (15 stack(s) stopped) — restarting them`**26 running, 0 unhealthy, marker cleared**, no stack stranded. **Crash-safety contract proven.** Bonus: a concurrent trigger was rejected `mentés már folyamatban van` (single-flight works). |
| C8 | guest reboot during a backup op | **PASS** | triggered DB-dump + `pct reboot 9201` mid-op → recovered to 26 containers/0 unhealthy; **drives remounted by UUID** (felhom-flash /dev/sdc1, felhom-usb /dev/sdb1, ext4); cloudflared+controller healthy; **romm DB intact (17 tables — no corruption)**; UI 200 |
**🔴 Finding (C1/C2) — docker restart policy not honored:** containers with `restart: unless-stopped` do **not** auto-restart after `docker kill` (navidrome `RestartCount=0` over two trials; dockerd logged `stopping restart-manager`; Docker 29.5.3, Live Restore off). The controller is not stopping them (no reconcile-stop in logs/source). Impact: **no docker-level self-heal for crashed app or controller containers** — the controller detects the stopped state and a UI/operator restart recovers apps, but a crashed *controller* itself has no automatic recovery (oneshot bootstrap + dead restart policy). **Needs supervised investigation** (docker daemon restart-manager behaviour) + a controller-recovery watchdog (e.g. bootstrap as a `Restart=always` monitor, or a systemd unit watching the container).
## Phase 4 — Backup/restore edge cases — **PASS (gaps re-confirmed)**
| Check | Result | Evidence |
|---|---|---|
| Concurrent restore → running-mutex blocks the 2nd | **PASS** | restore #1 (romm) running → restore #2 → 302 flash_error "Visszaállítás sikertelen: **backup or restore already in progress**"; #1 completed; romm healthy. No double-run/corruption |
| Backup with a stack already stopped | **PASS** | stopped uptime-kuma → `/api/backup/run` → ok:true, completed, **no backup errors in logs**; uptime-kuma restarted → healthy. (Tier-2 then returned single-flight "Mentés már folyamatban" — the DB-dump was still running) |
| Gap re-confirm: `/api/backup/snapshots` | **STILL 404** | `{"ok":false,"error":"endpoint not found"}` — restore-UI snapshot dropdown still can't populate (restore works via `POST /backup/restore`) |
| Gap re-confirm: restore DB import is additive | **CONFIRMED** | planted `c2_restore_marker` (romm 17→18 tbl) → restore → marker **survived** (count=1, tables still 18); dropped afterward (back to 17) |
| Low-space backup behavior | SKIPPED | tied to C4 (filling 100+ GB impractical/unsafe unattended) |
Bonus across the phase: **single-flight** is robust — observed rejecting concurrent restore (mutex), concurrent whole-guest backup (C7), and Tier-2 vs DB-dump.
## Phase 5 — Host reboot (optional) — _pending_
---
## Feature matrix — _filled at end_
## Cleanup confirmation — _filled at end_
## Deferred / needs supervised — _filled at end_