# TEST-REPORT — Test campaign #3: NO MERCY (N100 / guest 9201) **Run start:** 2026-06-22 (CC, unattended). Brutal chaos/edge — push to unintended states, verify handles-or-fails-safe, then recovers. Every break carries a timed auto-revert; Phase-0 backup is the floor. **Legend:** PASS / FAIL / SKIP + raw evidence. Each chaos line: break→detect→recover→verify. (Campaigns #1/#2 + diagnoses in git history + `felhom.eu/documentation/tests/`.) ## Phase 0 — Baseline + floor — **PASS (gate OPEN)** - ctrl **v0.75.0** healthy; agent v0.39.0; **25 containers, 0 unhealthy**; rootfs `/` 4%, `/var/lib/docker` 8%, drives 1%; mem **8.8 Gi available**. - Floor: PBS `felhom-pbs:backup/ct/9201/2026-06-22T19:01:56Z` (success, crash-consistent, verified). ## Phase 1 — Resource starvation — **PASS (fail-safe held)** | # | Break (timed-revert) | Result | Evidence | |---|---|---|---| | R1 | python hog ~8 GB (~91% RAM), 75s | **PASS** | mid-stress: 10Gi used / 1.1Gi avail — controller **healthy**, 25/0 unhealthy; post: **dockerd NRestarts=0, no OOM** (alloc fit via cache eviction), recovered to 8.9Gi. Backstop pkill guard armed. | | R2 | settings save on full disk | **PASS (code-verified)** | `settings.save()` = atomic write-`.tmp`-then-`os.Rename` (`settings.go:236-261`) → a disk-full fails at WriteFile, original untouched. Live disk-full not reproducible: settings.json is on the 252 GB docker vol. | | R3 | fill rootfs (holds `/mnt/sys_drive` backups) to 94% (timed rm @150s) | **PASS** | controller stayed healthy (it's on the separate 252 GB vol); DB-dump + Tier-2 (8 apps, immich 155MB) **succeeded**; **settings.json stayed valid JSON** (no corruption). Reverted → rootfs 4%. (Literal-full refusal not pushed — near-0 rootfs risks wedging the guest OS unattended; the 2 GB margin had room so the gate didn't need to refuse.) | | R4 | memory gate hard-block | **PASS (code-verified)** | `deploy.go:186` is a **hard block** (returns an error on `committed+new > usable`), using **committed-memory** accounting (sum of deployed `mem_request`), reserved 384 MB, usable 11904 MB. Live committed ≈ 4.4 GB → ~7.5 GB headroom, so tripping needs mass-deploying ~7.5 GB of requests (impractical/risky unattended); the branch + math are verified. | ## Phase 2 — State corruption [HIGH SEV] — **mostly PASS; 1 medium finding (S1)** | # | Break (backed up first) | Result | Evidence | |---|---|---|---| | S1 | truncate settings.json → restart | **⚠ FINDING (medium)** | controller **crash-loops**: `[FATAL] Failed to load settings … unexpected end of JSON input`, RestartCount=7, restarting. **No safe-defaults fallback** — a corrupt settings.json takes the management plane down. *Not silent* (FATAL logged — better than the worst case). Restore → running/healthy, 3 storage_paths back. (Docker restart-manager cycling it re-confirms Finding #1.) | | S2 | garbage in uptime-kuma `app.yaml` → rescan | **PASS** | controller stays **healthy** (25/0), logs `[WARN] LoadAppConfig: yaml: … did not find expected key` (not silent, not fatal), other apps unaffected. Restore → `running/deployed`. **Contrast with S1:** per-stack app.yaml corruption is graceful; settings.json is fatal. | | S3 | corrupt quiesce marker (bad JSON) → restart | **PASS (minor finding)** | no panic/crash-loop, controller healthy, **stacks not stranded** (restarts=0). But the unparseable marker is **silently ignored** (no log) **and left in place** (not cleared/quarantined) — a real corrupted-mid-quiesce marker would skip recovery with no signal. | | S4 | symlink `→ /etc` in userdata/appdata → backup | **PASS (security holds)** 🔒 | Tier-2 rsync **preserved the symlink** (`EVIL_ETC -> /etc`), did **not** follow it; **/etc contents did NOT leak** into the backup (no passwd/shadow). No path-escape/exfil. | ## Phase 3 — Concurrency storms — **PASS** | # | Break | Result | Evidence | |---|---|---|---| | C1 | deploy + backup + restore + git-sync simultaneously | **PASS** | backup/run + tier2 both single-flighted ("Mentés már folyamatban"); restore mutex'd (302); sync ran independently ("nincs változás"); after settle **no stuck flag** (`running:false`), no deadlock, 25/0 healthy | | C2 | rapid felhom-usb flap ×5 in ~10s | **PASS** | converged → felhom-usb **MOUNTED**, disconnected mark `None` (not stuck); **0 `permission denied`** during the flap; the v0.75 mountpoint-gate fired **5 clean skips** ("not mounted") in the disconnected windows; 25/0 | | C3 | kill `felhom-agent` mid-quiesce-backup | **PASS** | quiesce stopped 14 stacks → agent killed → controller `[quiesce] unquiescing (backup start failed): restarting 14 stack(s)` (clean error + defer-unwind, **no stack stranded**); agent restarted → `/api/disks` OK; **idle sockets→8443 = 1** (v0.74 bound holds); recovered 25/0 | ## Phase 4 — Time chaos — **SKIPPED (not safely isolable)** T1 (+25h) / T2 (−2h): the guest is an **unprivileged LXC** — it **cannot set its own clock** (`date -s` → `Operation not permitted`, no CAP_SYS_TIME) and shares the host's `CLOCK_REALTIME` (time namespaces don't isolate the wall clock). The only way to jump the guest's wall-clock is to change the **host** (felhom-pve) clock, which would risk: cloudflared tunnel cert time-validation (demo down), the agent's leaf-cert pin to PVE/PBS, and **agent→DooPlex hub/PBS TLS auth** (§0.1 forbids degrading DooPlex). NTP is active on the host. **Not safely isolable to the guest unattended → SKIP.** Proper venue: a dedicated throwaway VM with its own clock, or an injectable-clock unit test. **Deferred to supervised / implementation.** ## Phase 5 — Network partitions — _pending_ ## Phase 6 — Input/security fuzzing [HIGH SEV] — _pending_ ## Phase 7 — Brutal recovery — _pending_ --- ## Findings (ranked by severity) — _filled at end_ ## Cleanup confirmation — _filled at end_