diff --git a/TEST-REPORT.md b/TEST-REPORT.md index 494aa21..de19b22 100644 --- a/TEST-REPORT.md +++ b/TEST-REPORT.md @@ -1,139 +1,22 @@ -# TEST-REPORT — Test campaign #2 (broader apps, edge cases, break-and-recover) — N100 / guest 9201 +# TEST-REPORT — Test campaign #3: NO MERCY (N100 / guest 9201) -**Run start:** 2026-06-22 (CC, unattended). **Operator:** Viktor (away). -**Scope:** push past campaign #1 — more apps, edge cases, chaos break-and-recover. Storage lifecycle now -unblocked (ctrl v0.74.0 fixed the agent leak). Only code changes: trivial catalog HC/config-seed fixes. -**Legend:** PASS / FAIL / SKIPPED(reason). Evidence raw (docker ps / ss / curl / screenshot). -(Campaign #1 report preserved in git history + `felhom.eu/documentation/tests/unattended-test-campaign-2026-06-22-*.md`.) +**Run start:** 2026-06-22 (CC, unattended). Brutal chaos/edge — push to unintended states, verify +handles-or-fails-safe, then recovers. Every break carries a timed auto-revert; Phase-0 backup is the floor. +**Legend:** PASS / FAIL / SKIP + raw evidence. Each chaos line: break→detect→recover→verify. +(Campaigns #1/#2 + diagnoses in git history + `felhom.eu/documentation/tests/`.) + +## Phase 0 — Baseline + floor — **PASS (gate OPEN)** +- ctrl **v0.75.0** healthy; agent v0.39.0; **25 containers, 0 unhealthy**; rootfs `/` 4%, `/var/lib/docker` 8%, drives 1%; mem **8.8 Gi available**. +- Floor: PBS `felhom-pbs:backup/ct/9201/2026-06-22T19:01:56Z` (success, crash-consistent, verified). + +## Phase 1 — Resource starvation — _pending_ +## Phase 2 — State corruption [HIGH SEV] — _pending_ +## Phase 3 — Concurrency storms — _pending_ +## Phase 4 — Time chaos — _pending_ +## Phase 5 — Network partitions — _pending_ +## Phase 6 — Input/security fuzzing [HIGH SEV] — _pending_ +## Phase 7 — Brutal recovery — _pending_ --- - -## Phase 0 — Baseline + safety net + leak-fix gate — **PASS (gate OPEN)** - -- **Baseline:** ctrl `v0.74.0` healthy, agent `v0.39.0`; `pct list` = 9201 running (+ 9001/9999 stopped, pre-existing); **25 containers, 0 unhealthy**; rootfs `/` 4%, `/var/lib/docker` 8%, drives felhom-flash 1% / felhom-usb 1%. -- **Master safety net + DR source:** PBS backup `felhom-pbs:backup/ct/9201/2026-06-22T16:13:44Z`, success=true, crash-consistent; PBS verify shows the snapshot present. -- **Leak-fix gate (the #1 unblock):** `/api/disks` lists felhom-flash/felhom-usb/local; `/api/host-metrics` returns `cpu_percent`; controller→agent idle sockets to `:8443` **flat at 2 before and after a 60-call burst**. Agent path fully working. **Gate to chaos/destructive phases: OPEN.** - -## Phase 1 — Broader app sweep + edge cases — **PASS (with noted non-triggerables)** - -| Check | Result | Evidence | -|---|---|---| -| Deploy navidrome (HDD-bound, Go) via GUI | PASS | GUI deploy → 3-step panel → `running/healthy` in 8s; `:4533` → HTTP 302 (serving); HDD binding `…/felhom-flash/userdata/media/music → /music` | -| Pattern coverage | PASS (cited) | JVM=komga (healthy), Python+DB=immich/paperless (healthy), Go=navidrome/gitea, HDD-bound=navidrome/romm/immich, optional-config=romm | -| Edge: rapid deploy→stop→redeploy race | PASS | stop → immediate update → `navidrome Up (healthy)`, `app.yaml deployed:true`, API `state=running deployed=true` — no stale button / no Deployed-flag race | -| Edge: slow-DB-migration start_period | PASS (cited) | mealie deployed healthy (start_period 60s covers migration — no premature unhealthy flap; lesson #7) | -| Edge: missing required config → graceful signal | PASS (cited #1) | campaign-1 glance: crash-loop with a clear logged error (`open /app/config/glance.yml: no such file`), controller marks it — not a silent hang | -| Edge: mem-gate hard-block | PARTIAL | the deploy-page memory projection renders correctly (e.g. navidrome 4065→+50 MB / 11904); a RAM hard-block isn't naturally triggerable (no single catalog app's request > ~7.6 GB free without mass-filling RAM). Disk-headroom hard-refuse (`GetDockerVolumeHeadroom`) is tested in Phase 3/C4. | -| Edge: Host()/port overlap | N/A (noted) | each catalog app has a unique subdomain by construction; an overlap isn't naturally producible without editing a template — not triggered | - -navidrome left deployed (used as the drive-backed app for Phase-3 C6); removed at cleanup. -## Phase 2 — Storage lifecycle (the campaign-#1 blocker, NOW UNBLOCKED) — **PASS (destructive-format skipped, documented)** - -Agent path works (v0.74.0), so these ran through the real controller→agent pipeline. - -| Check | Result | Evidence | -|---|---|---| -| Scan / observe / health / FS / role / durable-id | PASS | `/api/disks`: 5 targets — felhom-flash (user-data, uuid:81a26531), felhom-usb (user-data, uuid:da9e7089), local/local-lvm (system), felhom-pbs (backup) with capacities + data_bearing | -| Label edit → revert (live) | PASS | felhom-usb label → "TEST-LABEL-c2" (log "Storage label updated") → reverted to exact UTF-8 original "Tárhely (felhom-usb)" (byte-verified). [my shell mangled the accent on first revert; corrected via `%C3%A1`] | -| **Disconnect → reconnect (reversible four-state)** | PASS | enrolled→**disconnected** (`/api/storage/disconnect`: unmounted, no apps) → **enrolled** (`/api/storage/reconnect`: remounted `/dev/sdb1`, disconnected mark cleared) | -| **Decommission → re-enroll (non-destructive)** | PASS | `/api/storage/decommission mode=anyway` (type-to-confirm `felhom-usb`): unmounted, `decommissioned=True` soft-marker, **fs intact — uuid da9e7089 preserved, NOT formatted**. `/api/storage/reconnect`: `reenrolled=true`, MOUNTED, marker cleared, label preserved | -| Destructive format/wipe on **loopback scratch** | **BLOCKED (documented)** | created `/dev/loop0` (truncate 2G + losetup); `/api/storage/init` → agent **`refusing to operate on non-whitelisted block device "/dev/loop0"`** — a deliberate device-whitelist safety control. So the clean scratch target the runbook prescribed is unavailable. Per §0.3's escape clause (no clean scratch target → skip), the destructive *format* was not run on the live felhom-usb either. Loopback torn down. | -| Migrate (data) | DEFERRED | no safe scratch target (loopback blocked); a live migrate would move a real app's data — not run unattended. `internal/stacks/migrate.go` has unit coverage. | - -**Findings:** -1. **Agent device whitelist** refuses non-whitelisted block devices (loopback) — good security, but means a loopback can't serve as the destructive-disk scratch target the runbook assumed. Destructive format/migrate now need either a real spare disk or a supervised run. -2. **FileBrowser sync** logs `mkdir …/userdata: permission denied` when ensuring the skeleton on a just-unmounted/decommissioned path (recurs on disconnect + decommission). Cosmetic (the drive is intentionally unmounted), but noisy — could gate the skeleton-ensure on "is the path a live mountpoint." -## Phase 3 — Break-and-recover chaos (C1–C8) — **PASS (7/8 run, C4 skipped; 2 findings)** - -| # | Break → Detect → Recover → Verify | Result | Evidence | -|---|---|---|---| -| C1 | `docker kill navidrome` → controller detects → restart via UI → healthy | **PASS** | killed → `Exited(137)`; controller state `stopped`; `POST /api/stacks/navidrome/restart` → `running/healthy`, state `running` | -| C2 | `docker kill felhom-controller` → (expected bootstrap self-heal) → recover → verify | **PASS (w/ 🔴 finding)** | killed → **stayed Exited 48s, NO auto-recovery** (docker `unless-stopped` didn't fire; bootstrap.service is oneshot/exited). Data plane unaffected (apps+cloudflared+traefik up). Recovered via `systemctl restart felhom-controller-bootstrap.service` → running/healthy; UI 200; 26 containers/0 unhealthy. **Finding: a crashed controller has NO automatic recovery.** | -| C3 | stop `felhom-agent` → controller degrades gracefully → restart agent → immediate success | **PASS** | agent down → `/api/disks` **fast ~98ms clean error** ("connection refused"), no hang/crash; agent restarted → `/api/disks` OK on the **first** try (4/4), no stale-conn errors; idle sockets bounded at 1 — the v0.74.0 client's restarted-agent drain works | - -| C4 | fill a drive near-full → disk warning → clear | **SKIPPED(impractical)** | drives are 118 GB / 916 GB and `/var/lib/docker` 252 GB; reaching the 90% warning / the deploy gate's `max(5GB,10%)`≈25 GB-free needs 100+ GB of fill — unsafe/slow unattended. Gate logic `GetDockerVolumeHeadroom` (hard-refuse deploy below the reserve) is present; monitoring reports per-drive usage. | -| C5 | rename `navidrome/app.yaml` away → rescan → restore | **PASS** | rescan after removal: controller stays `running/healthy` (no crash), container keeps running (data plane), card → `state=running deployed=false` (graceful orphan). Restored + rescan → `state=running deployed=true` | -| C6 | disconnect felhom-flash (9 apps) → recover | **PASS (recovered) + finding** | disconnect stopped all 9 apps + unmounted (`UNMOUNTED`, komga `stopped`); reconnect → felhom-flash remounted `/dev/sdc1`, all 9 restarted → **26 running, 0 unhealthy** after start_period. **Finding:** during the drive-absent window the controller's userdata-belt/recreate tried to bring drive-backed apps up while the mount was gone → `mkdir …/userdata: permission denied` + transient `Created` states (rather than a clean held "missing-storage" stop). Reconnect cleared it; but the disconnect→recreate interaction is noisy — worth a guard ("don't recreate/ensure-userdata when the drive is absent"). | - -| C7 | kill controller MID-quiesce → Recover resumes | **PASS** | triggered whole-guest backup → quiesce marker written **upfront** (`active:true`, 15 stopped_stacks) → `docker kill felhom-controller` mid-quiesce → **marker persisted across the crash**; controller restart logged `[quiesce] crash recovery: a quiesce was in progress (15 stack(s) stopped) — restarting them` → **26 running, 0 unhealthy, marker cleared**, no stack stranded. **Crash-safety contract proven.** Bonus: a concurrent trigger was rejected `mentés már folyamatban van` (single-flight works). | -| C8 | guest reboot during a backup op | **PASS** | triggered DB-dump + `pct reboot 9201` mid-op → recovered to 26 containers/0 unhealthy; **drives remounted by UUID** (felhom-flash /dev/sdc1, felhom-usb /dev/sdb1, ext4); cloudflared+controller healthy; **romm DB intact (17 tables — no corruption)**; UI 200 | - -**🔴 Finding (C1/C2) — docker restart policy not honored:** containers with `restart: unless-stopped` do **not** auto-restart after `docker kill` (navidrome `RestartCount=0` over two trials; dockerd logged `stopping restart-manager`; Docker 29.5.3, Live Restore off). The controller is not stopping them (no reconcile-stop in logs/source). Impact: **no docker-level self-heal for crashed app or controller containers** — the controller detects the stopped state and a UI/operator restart recovers apps, but a crashed *controller* itself has no automatic recovery (oneshot bootstrap + dead restart policy). **Needs supervised investigation** (docker daemon restart-manager behaviour) + a controller-recovery watchdog (e.g. bootstrap as a `Restart=always` monitor, or a systemd unit watching the container). -## Phase 4 — Backup/restore edge cases — **PASS (gaps re-confirmed)** - -| Check | Result | Evidence | -|---|---|---| -| Concurrent restore → running-mutex blocks the 2nd | **PASS** | restore #1 (romm) running → restore #2 → 302 flash_error "Visszaállítás sikertelen: **backup or restore already in progress**"; #1 completed; romm healthy. No double-run/corruption | -| Backup with a stack already stopped | **PASS** | stopped uptime-kuma → `/api/backup/run` → ok:true, completed, **no backup errors in logs**; uptime-kuma restarted → healthy. (Tier-2 then returned single-flight "Mentés már folyamatban" — the DB-dump was still running) | -| Gap re-confirm: `/api/backup/snapshots` | **STILL 404** | `{"ok":false,"error":"endpoint not found"}` — restore-UI snapshot dropdown still can't populate (restore works via `POST /backup/restore`) | -| Gap re-confirm: restore DB import is additive | **CONFIRMED** | planted `c2_restore_marker` (romm 17→18 tbl) → restore → marker **survived** (count=1, tables still 18); dropped afterward (back to 17) | -| Low-space backup behavior | SKIPPED | tied to C4 (filling 100+ GB impractical/unsafe unattended) | - -Bonus across the phase: **single-flight** is robust — observed rejecting concurrent restore (mutex), concurrent whole-guest backup (C7), and Tier-2 vs DB-dump. -## Phase 5 — Host reboot (optional) — **SKIPPED (unattended risk)** -A host reboot of the N100 is the highest-risk action unattended — if it doesn't POST/return (BIOS -AC-recovery etc.) the demo is down until physical access, which CC doesn't have. Deferred to supervised -(consistent with campaign #1). Guest-reboot recovery is proven (C8 + campaign #1). - ---- -## Feature matrix -| Area | Verdict | -|---|---| -| PBS backup + verify + leak-fix gate | PASS | -| App deploy (HDD-bound navidrome) + redeploy race | PASS | -| Storage: scan/health/durable-id | PASS | -| Storage: label edit/revert | PASS | -| Storage: disconnect↔reconnect (4-state) | PASS | -| Storage: decommission↔re-enroll (non-destructive) | PASS | -| Storage: destructive format on loopback | BLOCKED (agent device-whitelist) | -| Chaos C1 kill app → restart | PASS | -| Chaos C2 kill controller | PASS recovery / 🔴 no auto-heal | -| Chaos C3 stop agent (graceful + drain) | PASS | -| Chaos C5 remove app.yaml | PASS | -| Chaos C6 eject drive w/ apps | PASS recovery / finding (recreate noise) | -| Chaos C7 kill mid-quiesce → Recover | PASS (crash-safety proven) | -| Chaos C8 reboot during op | PASS | -| Backup: concurrent mutex | PASS | -| Backup: stopped-stack | PASS | -| Single-flight (backup/restore/migrate) | PASS | -| `/api/backup/snapshots` | 404 (gap, carried from #1) | -| restore DB import additive | confirmed (gap, carried from #1) | -| C4 disk-full / low-space | SKIPPED (impractical) | -| Phase 5 host reboot | SKIPPED (unattended risk) | - -## Findings (new this campaign) -1. **🔴 docker restart policy not honored** — `unless-stopped` containers don't auto-restart after a - crash (`docker kill` → `RestartCount=0`, dockerd `stopping restart-manager`; Docker 29.5.3). **No - docker-level self-heal for crashed apps OR the controller.** The controller detects + a UI/operator - restart recovers apps; but a crashed **controller** has **no automatic recovery** (oneshot bootstrap - + dead restart policy) — manual `systemctl restart felhom-controller-bootstrap.service` needed. - *Recommend:* a controller watchdog (e.g. bootstrap `Restart=always` monitor or a systemd unit - watching the container) + investigate why the daemon's restart-manager is cancelled. -2. **Disconnect-recreate race (C6)** — disconnecting a drive-with-apps triggers the controller's - userdata-belt/recreate to try bringing drive-backed apps up *while the drive is absent* → - `mkdir …/userdata: permission denied` + transient `Created` states, instead of a clean held - missing-storage stop. Reconnect recovers it. *Recommend:* gate ensure-userdata/recreate on - "drive is a live mountpoint." -3. **FileBrowser skeleton on unmounted path** — disconnect/decommission log - `mkdir …/userdata: permission denied` from the FileBrowser sync on the just-unmounted path. Cosmetic - but noisy; gate on live-mountpoint. -4. **Agent device whitelist** refuses loopback (`/dev/loop0`) — good security, but means a loopback - can't be the destructive-disk scratch target; destructive format/migrate need a real spare disk or - a supervised run. - -Gaps carried from campaign #1 (still open): `/api/backup/snapshots` 404; restore DB import additive. - -## Cleanup confirmation -- No throwaway guest created this campaign (DR was campaign #1); scratch loopback torn down (`losetup -d` + rm). -- navidrome (test app) removed → back to **25 containers** baseline; uptime-kuma restarted healthy. -- All chaos breaks recovered + verified (C1–C8); agent **active**; felhom-flash + felhom-usb **re-enrolled** + mounted. -- All timed guards killed (0 `sleep` guards); test forwarder + host temp files removed; **0 loopbacks**. -- Notifications untouched this campaign (none configured); no password changes. -- **9201 running, 25 containers, 0 unhealthy; controller + cloudflared healthy; agent active.** - -## Deferred / needs supervised -- **Controller-recovery watchdog** + docker restart-policy investigation (finding #1) — highest priority. -- Destructive storage **format/migrate** on a real spare disk (loopback blocked) — supervised. -- **Host reboot** resilience (Phase 5). -- Disconnect-recreate gate (finding #2); FileBrowser unmounted-skeleton gate (finding #3). -- `/api/backup/snapshots` handler + UI snapshot selection; restore `--clean` option (campaign-#1 gaps). +## Findings (ranked by severity) — _filled at end_ +## Cleanup confirmation — _filled at end_