8845496c0e
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
5.0 KiB
5.0 KiB
TEST-REPORT — Test campaign #3: NO MERCY (N100 / guest 9201)
Run start: 2026-06-22 (CC, unattended). Brutal chaos/edge — push to unintended states, verify
handles-or-fails-safe, then recovers. Every break carries a timed auto-revert; Phase-0 backup is the floor.
Legend: PASS / FAIL / SKIP + raw evidence. Each chaos line: break→detect→recover→verify.
(Campaigns #1/#2 + diagnoses in git history + felhom.eu/documentation/tests/.)
Phase 0 — Baseline + floor — PASS (gate OPEN)
- ctrl v0.75.0 healthy; agent v0.39.0; 25 containers, 0 unhealthy; rootfs
/4%,/var/lib/docker8%, drives 1%; mem 8.8 Gi available. - Floor: PBS
felhom-pbs:backup/ct/9201/2026-06-22T19:01:56Z(success, crash-consistent, verified).
Phase 1 — Resource starvation — PASS (fail-safe held)
| # | Break (timed-revert) | Result | Evidence |
|---|---|---|---|
| R1 | python hog ~8 GB (~91% RAM), 75s | PASS | mid-stress: 10Gi used / 1.1Gi avail — controller healthy, 25/0 unhealthy; post: dockerd NRestarts=0, no OOM (alloc fit via cache eviction), recovered to 8.9Gi. Backstop pkill guard armed. |
| R2 | settings save on full disk | PASS (code-verified) | settings.save() = atomic write-.tmp-then-os.Rename (settings.go:236-261) → a disk-full fails at WriteFile, original untouched. Live disk-full not reproducible: settings.json is on the 252 GB docker vol. |
| R3 | fill rootfs (holds /mnt/sys_drive backups) to 94% (timed rm @150s) |
PASS | controller stayed healthy (it's on the separate 252 GB vol); DB-dump + Tier-2 (8 apps, immich 155MB) succeeded; settings.json stayed valid JSON (no corruption). Reverted → rootfs 4%. (Literal-full refusal not pushed — near-0 rootfs risks wedging the guest OS unattended; the 2 GB margin had room so the gate didn't need to refuse.) |
| R4 | memory gate hard-block | PASS (code-verified) | deploy.go:186 is a hard block (returns an error on committed+new > usable), using committed-memory accounting (sum of deployed mem_request), reserved 384 MB, usable 11904 MB. Live committed ≈ 4.4 GB → ~7.5 GB headroom, so tripping needs mass-deploying ~7.5 GB of requests (impractical/risky unattended); the branch + math are verified. |
Phase 2 — State corruption [HIGH SEV] — mostly PASS; 1 medium finding (S1)
| # | Break (backed up first) | Result | Evidence |
|---|---|---|---|
| S1 | truncate settings.json → restart | ⚠ FINDING (medium) | controller crash-loops: [FATAL] Failed to load settings … unexpected end of JSON input, RestartCount=7, restarting. No safe-defaults fallback — a corrupt settings.json takes the management plane down. Not silent (FATAL logged — better than the worst case). Restore → running/healthy, 3 storage_paths back. (Docker restart-manager cycling it re-confirms Finding #1.) |
| S2 | garbage in uptime-kuma app.yaml → rescan |
PASS | controller stays healthy (25/0), logs [WARN] LoadAppConfig: yaml: … did not find expected key (not silent, not fatal), other apps unaffected. Restore → running/deployed. Contrast with S1: per-stack app.yaml corruption is graceful; settings.json is fatal. |
| S3 | corrupt quiesce marker (bad JSON) → restart | PASS (minor finding) | no panic/crash-loop, controller healthy, stacks not stranded (restarts=0). But the unparseable marker is silently ignored (no log) and left in place (not cleared/quarantined) — a real corrupted-mid-quiesce marker would skip recovery with no signal. |
| S4 | symlink → /etc in userdata/appdata → backup |
PASS (security holds) 🔒 | Tier-2 rsync preserved the symlink (EVIL_ETC -> /etc), did not follow it; /etc contents did NOT leak into the backup (no passwd/shadow). No path-escape/exfil. |
Phase 3 — Concurrency storms — PASS
| # | Break | Result | Evidence |
|---|---|---|---|
| C1 | deploy + backup + restore + git-sync simultaneously | PASS | backup/run + tier2 both single-flighted ("Mentés már folyamatban"); restore mutex'd (302); sync ran independently ("nincs változás"); after settle no stuck flag (running:false), no deadlock, 25/0 healthy |
| C2 | rapid felhom-usb flap ×5 in ~10s | PASS | converged → felhom-usb MOUNTED, disconnected mark None (not stuck); 0 permission denied during the flap; the v0.75 mountpoint-gate fired 5 clean skips ("not mounted") in the disconnected windows; 25/0 |
| C3 | kill felhom-agent mid-quiesce-backup |
PASS | quiesce stopped 14 stacks → agent killed → controller [quiesce] unquiescing (backup start failed): restarting 14 stack(s) (clean error + defer-unwind, no stack stranded); agent restarted → /api/disks OK; idle sockets→8443 = 1 (v0.74 bound holds); recovered 25/0 |