diff --git a/TEST-REPORT.md b/TEST-REPORT.md index df29426..192a7f1 100644 --- a/TEST-REPORT.md +++ b/TEST-REPORT.md @@ -56,8 +56,49 @@ supervised / implementation.** | F2 | path traversal (storage `where`, restore `stack_name`) | **PASS (no escape) 🔒 + FINDING** | **storage:** all 5 traversals (`/etc`, `../../etc`, `/mnt/../etc`, `/mnt/felhom-drives/../../../etc`, `/etc/shadow`) **rejected** by `gateWhere` (`path.Clean`+`HasPrefix("/mnt/")`). **restore:** the boundary HELD — `/etc/passwd` **intact**, no `/opt/etc`/`/etc/passwd/` artifacts, no `/etc` writes — but **the restore handler does NOT reject a traversal `stack_name` upfront**: `RestoreFromRecoveryUnit("../../../etc")` proceeded (`GetAppDrivePath` → default `/mnt/sys_drive`), saved only by downstream **map-based** `StopStack`/`StartStack` (`stack "../../../etc" not found`) + no recovery-unit/volumes (no-op). **Defense-in-depth gap — recommend explicit `stack_name` validation at the restore handler.** (My own batched test looped these restores ×90s each, briefly holding the mutex — a test artifact that cleared when stopped, NOT a stuck-flag bug.) | | F3 | hostile compose on a scratch app | **N/A (good property)** | the standard deploy renders only from the git-synced **catalog** — there is **no customer-facing arbitrary-compose injection vector**. (`.fab` import is the only path; not fuzzed — lower priority.) | | F4 | weird userdata filenames (unicode, spaces, `(N)`, 200-char) | **PASS** | files created + Tier-2 backup ran clean ("Tier 2 run complete: 8 apps"), no crash/panic, 25/0. (`(N)` merge/dedup is migration-specific — `migrate.go` unit-tested, not live-run here.) | -## Phase 7 — Brutal recovery — _pending_ +## Phase 7 — Brutal recovery — **PASS (B1-B3); B4 skipped** + +| # | Break | Result | Evidence | +|---|---|---|---| +| B1 | restart dockerd in 9201 | **PASS** | boot-restore brought all **25 containers + controller** back in <8s (drives stay mounted on a daemon-only restart → no boot-ordering issue); controller+cloudflared healthy | +| B2 | `kill -9` controller PID (real crash) | **PASS** | `unless-stopped` auto-recovered: RestartCount 0→1, **healthy in 8s** — re-confirms Finding #1 (docker kill ≠ crash) live | +| B3 | reboot 9201 with **felhom-flash detached** | **PASS** 🔑 | booted; flash **DETACHED** (disconnect intent persisted across reboot); 8 HDD apps **held** (absent, not crash-looping); **NO rootfs shadow dirs** (`…/felhom-flash/userdata` doesn't exist). Docker boot-restore made 2 bind-source mkdir attempts → **denied by the unprivileged-LXC mapping** → no escape (the v0.75 gate covers the controller's post-boot belt/FileBrowser; the unprivileged mapping covers docker's boot-restore). Reconnect → flash remounted, 8 apps recovered → 25/0. | +| B4 | host reboot of felhom-pve | **SKIPPED** | unattended risk — no physical recovery if the N100 doesn't POST/return. Deferred to supervised (consistent with #1/#2). | --- -## Findings (ranked by severity) — _filled at end_ -## Cleanup confirmation — _filled at end_ +## Findings (ranked by real-world severity) + +1. **[MEDIUM] Corrupt `settings.json` → controller FATAL crash-loop (S1).** No safe-defaults fallback — + a truncated/invalid settings.json takes the management plane down (RestartCount climbs via docker's + restart-manager). *Not silent* (FATAL is logged — strictly better than the worst case). Apps/tunnel + keep running (control/data separation). **Fix:** on parse failure, log a WARN + load safe defaults + (or boot read-only) instead of `[FATAL]`. (Contrast: per-stack `app.yaml` corruption (S2) is handled + gracefully with a WARN — settings.json should match that.) +2. **[MEDIUM, defense-in-depth] Restore `stack_name` not validated against path traversal (F2).** The + `/backup/restore` handler proceeds with a traversal name (`../../../etc`) into + `RestoreFromRecoveryUnit`/`RestoreApp`; **no escape occurred** (`/etc` intact, no artifacts) only + because downstream `StopStack`/`StartStack` are map-based ("stack not found") + there's no + recovery-unit/volumes for a bogus name. A future code path that built a filesystem path from the raw + name could escape. **Fix:** reject `stack_name` containing `/`, `..`, NUL, etc. at the handler. +3. **[LOW] Corrupt quiesce marker silently ignored (S3).** `Recover()` doesn't panic (good) but an + unparseable marker is **neither logged nor cleared/quarantined** — a real corrupted-mid-quiesce + marker would skip stack-recovery with no signal. **Fix:** log + quarantine a bad marker. +4. **[INFO] Boot-ordering (re-confirmed, B3).** Docker's boot-restore attempts drive-backed bind-source + `mkdir` before mounts converge; the **unprivileged-LXC mapping denies it** (no shadow dirs), and the + v0.75 gate covers the controller side — benign but noisy. (Tracked from the v0.75 task.) + +**Everything else fail-safe held:** R1-R4, S2, S4, C1-C3, N1-N2, F1, F3, F4, B1-B3 all PASS. No +silent-corruption and **no path-escape/exfil** found (the two highest-severity classes the runbook +prioritized). **No code changes shipped** (no broken apps surfaced; the findings are behaviours logged +for supervised fix). Phase 4 (time chaos) and B4 (host reboot) SKIPPED for architectural/unattended-risk +reasons (documented). + +## Cleanup confirmation +- **9201 running, 25 containers, 0 unhealthy**; controller + cloudflared healthy; agent active. +- Drives mounted (flash sdc1, usb sdb1); rootfs back to **4% / 29 G free** (R3 fill reverted). +- **No leftover rules:** host `OUTPUT` 8007-DROP = 0 (N2), guest `DOCKER-USER` cloudflared-DROP = 0 (N1). +- **No leftover guards** (the 2 residual `sleep` procs are host system monitors, parents 747/700843 — not mine). +- Corrupted files restored (settings.json valid; app.yaml restored; quiesce marker removed; symlinks removed); F4 test files removed. +- **No 9300 / scratch guest created** this campaign; no loopbacks (Phase 4 N/A, destructive disk used live drives only reversibly). +- Pre-existing (NOT this run): 9001 (spike-lxc) + 9999 (felhom-selftest-scratch) stopped. +- Net: demo restored to baseline (= Phase-0 state).