Files
felhom.eu/documentation/tests/test-campaign-3-2026-06-22-findings.md
T
admin add5be5bd5 docs(tests): test campaign #3 'NO MERCY' findings (brutal chaos/edge, demo 9201)
7 phases. No silent-corruption, no path-escape/exfil found. Findings: S1 settings.json
crash-loop (medium, no safe-defaults), F2 restore stack_name traversal not validated
upfront (medium defense-in-depth, no escape observed), S3 corrupt quiesce marker silently
ignored (low). All else fail-safe held. Phase4 time-chaos SKIP (unprivileged LXC shares
host clock). No code changes. Demo restored to baseline.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-22 22:06:00 +02:00

91 lines
6.5 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Test campaign #3 — "NO MERCY" brutal chaos/edge — 2026-06-22 (demo 9201)
> **Status: empirical findings from an unattended brutal chaos campaign (2026-06-22).** Push the
> system into unintended states; verify it handles-or-fails-safe, then recovers. Every break carried a
> timed auto-revert; the Phase-0 verified PBS backup was the floor. ctrl **v0.75.0** + agent v0.39.0.
> Demo restored to baseline (25 containers, 0 unhealthy) and healthy at the end. Per-check raw evidence:
> `felhom-controller/TEST-REPORT.md`. **No silent-corruption and no path-escape/exfil found** (the two
> highest-severity classes). **No code changes shipped** — findings are behaviours logged for supervised fix.
## Result matrix
| Phase | Result |
|---|---|
| 0 baseline + verified PBS backup floor | PASS (gate open) |
| 1 resource starvation (R1 RAM 91%, R2 settings-atomicity, R3 rootfs 94%, R4 mem-gate) | PASS (fail-safe held) |
| 2 state corruption (S1 settings, S2 app.yaml, S3 quiesce marker, S4 symlink) | PASS + 1 medium (S1) |
| 3 concurrency (C1 4-way ops, C2 drive-flap×5, C3 kill-agent-mid-backup) | PASS |
| 4 time chaos (clock jumps) | SKIP (architectural — see below) |
| 5 network partitions (N1 tunnel block, N2 PBS block mid-backup) | PASS |
| 6 security fuzzing (F1 JSON, F2 traversal, F3 hostile compose, F4 weird filenames) | PASS (no escape) + 1 medium (F2) |
| 7 brutal recovery (B1 dockerd restart, B2 kill-9 ctrl, B3 reboot+flash-detached, B4 host reboot) | PASS (B1-B3); B4 SKIP |
## Findings, ranked by real-world severity
### 1. [MEDIUM] Corrupt `settings.json` → controller FATAL crash-loop (Phase 2 / S1)
A truncated/invalid `settings.json` makes the controller `[FATAL] Failed to load settings … unexpected
end of JSON input` and **crash-loop** (RestartCount climbs via docker's restart-manager) — **no
safe-defaults fallback**; the management plane is down until the file is fixed. It is *not silent*
(FATAL is logged). Apps + tunnel keep running (control/data separation). **Contrast:** per-stack
`app.yaml` corruption (S2) is handled gracefully — controller logs `[WARN] LoadAppConfig: yaml: …`,
marks that app, keeps serving others, no crash. **Recommend:** make the settings.json load path match
app.yaml's — log a WARN + load safe defaults (or boot read-only), not `[FATAL]`.
### 2. [MEDIUM, defense-in-depth] Restore `stack_name` not validated against path traversal (Phase 6 / F2)
`POST /backup/restore` proceeds with a traversal `stack_name` (`../../../etc`) into
`RestoreFromRecoveryUnit`/`RestoreApp` (`GetAppDrivePath` returns the default `/mnt/sys_drive`).
**No escape occurred**`/etc/passwd` intact, no `/opt/etc`/`/etc/passwd/` artifacts, no `/etc` writes
**only because** downstream `StopStack`/`StartStack` are map-based (`stack "../../../etc" not found`)
and there's no recovery-unit/volumes for a bogus name (no-op). The security boundary held, but by
luck of downstream behaviour, not by input validation. A future code path that built a filesystem path
from the raw name could escape. **Recommend:** explicit `stack_name` validation at the restore handler
(reject `/`, `..`, NUL). (Storage `where=` IS validated — `gateWhere`: `path.Clean`+`HasPrefix("/mnt/")`
rejected all 5 traversals.)
### 3. [LOW] Corrupt quiesce marker silently ignored (Phase 2 / S3)
A bad-JSON `quiesce-state.json` does **not** panic the controller at startup (good — stacks not
stranded), but the unparseable marker is **neither logged nor cleared/quarantined**. A real
corrupted-mid-quiesce marker would silently skip stack-recovery. **Recommend:** log + quarantine a
marker that fails to parse.
### 4. [INFO] Boot-ordering, re-confirmed (Phase 7 / B3)
On a guest reboot with a drive detached, docker's boot-restore attempts drive-backed bind-source
`mkdir` before mounts converge — but the **unprivileged-LXC mapping denies it** (2 denied attempts,
**no rootfs shadow dirs**), and the v0.75 mountpoint-gate covers the controller's belt/FileBrowser.
Benign but noisy. (Tracked from the v0.75 task.)
## What held (highlights)
- **No path-escape / no /etc exfil** anywhere: storage `where=` validation (F2), symlink-in-backup not
followed (S4 — rsync preserved `EVIL_ETC -> /etc`, no contents copied), restore-traversal no-op (F2).
- **Resource pressure** (R1 91% RAM, R3 94% rootfs): controller + 25 apps stayed healthy, no OOM, no
daemon restart, **settings.json stayed valid** (atomic tmp+rename, R2).
- **Concurrency**: single-flight serialized 4-way op storm (C1), drive flap×5 converged with the v0.75
gate firing 5 clean skips and **0 permission-denied** (C2), agent-kill-mid-backup unwound the quiesce
with no stack stranded (C3).
- **Partitions**: tunnel block → apps serve locally + reconnect (N1); PBS block mid-backup → backup
fails safe + quiesce unwinds + next backup succeeds (N2).
- **Recovery**: dockerd restart → all back <8s (B1); `kill -9` controller → `unless-stopped`
auto-recovers in 8s (B2, re-confirms Finding #1); reboot+flash-detached → apps held, **no shadow
dirs**, reconnect recovers (B3).
## Phase 4 SKIP — time chaos not safely isolable (architectural)
The guest is an **unprivileged LXC**: it cannot set its own clock (`date -s``Operation not
permitted`) and shares the host's `CLOCK_REALTIME` (time namespaces don't isolate the wall clock). The
only way to jump the guest wall-clock is to change the **host** clock, which would risk the cloudflared
tunnel cert, the agent leaf-cert pin, and **agent→DooPlex hub/PBS TLS auth** (§0.1 forbids degrading
DooPlex). Proper venue: a dedicated throwaway VM with its own clock, or an injectable-clock unit test.
## Deferred / supervised / implementation
- **Fix S1** (settings.json: safe-defaults + WARN instead of FATAL crash-loop) — highest priority.
- **Fix F2** (explicit restore `stack_name` traversal validation).
- **Fix S3** (log + quarantine a corrupt quiesce marker).
- **Time chaos** (Phase 4) on a dedicated VM / injectable clock.
- **Host reboot** (B4) — supervised (no physical recovery unattended).
- `.fab` import compose-fuzzing (F3) — the only customer compose vector; not exercised.
## Cleanup
9201 running, 25 containers, 0 unhealthy; controller + cloudflared healthy; agent active; both drives
mounted; rootfs 4%/29 G free. No leftover iptables rules (host 8007, guest DOCKER-USER both 0); all
campaign timed-guards killed; corrupted files restored (settings valid, app.yaml restored, quiesce
marker removed, symlinks removed); F4 test files removed; no 9300/scratch/loopback created. Demo =
Phase-0 baseline.