From aa37e6e56eb504d03575eb38b035135e07043ef4 Mon Sep 17 00:00:00 2001 From: kisfenyo Date: Sat, 13 Jun 2026 23:12:24 +0200 Subject: [PATCH] runbook: e2e live-drive (supervised) stressing CTRL-T2-1/CTRL-001/AGENT-001 + session done Co-Authored-By: Claude Opus 4.8 (1M context) --- RUNBOOK-e2e-live-drive.md | 132 ++++++++++++++++++++++++++++++++++ SESSION-2026-06-13-cleanup.md | 12 +++- 2 files changed, 142 insertions(+), 2 deletions(-) create mode 100644 RUNBOOK-e2e-live-drive.md diff --git a/RUNBOOK-e2e-live-drive.md b/RUNBOOK-e2e-live-drive.md new file mode 100644 index 0000000..07fa647 --- /dev/null +++ b/RUNBOOK-e2e-live-drive.md @@ -0,0 +1,132 @@ +# RUNBOOK — End-to-end live drive (SUPERVISED) — felhom-controller + +**Purpose:** a hands-on end-to-end exercise of the demo stack that deliberately stresses the fixes +shipped 2026-06-13 — **CTRL-T2-1** (crash-safe deploy state), **CTRL-001** (import path traversal), +**AGENT-001** (anti-retarget wipe), plus backup/restore — and surfaces UX/latency/error-handling +friction. **Execute SUPERVISED (operator present).** Do the non-destructive sections (1–6) first; the +destructive section (7) is last and gated. + +**Status:** NOT executed yet — written 2026-06-13 for the supervised session. + +## Environment & conventions + +- **Demo controller:** guest **9201** (`demo-felhom`) on Proxmox host `felhom-pve` (192.168.0.162), + bootstrap-managed. Public dashboard/API: **https://felhom.demo-felhom.eu** (no dashboard password → + the API is open; drive it via the PUBLIC URL, not the container IP). +- **Versions at writing:** controller **v0.60.0**, agent **v0.30.0**, hub v0.11.0. +- `SSH=/c/Windows/System32/OpenSSH/ssh.exe`; host root via SSH alias `felhom-pve`; `export MSYS_NO_PATHCONV=1` for `pct exec`. +- **Findings log:** record every observation (✓/✗ + notes on UX friction, latency, confusing labels, + error handling) in a new `REPORT-e2e-live-drive-.md`. Each step says what "good" looks like and + what to watch for. +- **Markers:** **[DESTRUCTIVE — operator confirm]** = needs operator eyes + explicit go-ahead. + **[HUMAN]** = cannot be done by the agent (physical / real-decision). + +--- + +## 1. Baseline (read-only) + +1. Controller healthy + version: + - `curl -s https://felhom.demo-felhom.eu/api/health` → `{"ok":true,...}`. + - `$SSH felhom-pve "pct exec 9201 -- docker ps --filter name=felhom-controller --format '{{.Image}} {{.Status}}'"` → `:0.60.0 Up ... (healthy)`. + - Dashboard loads (Hungarian UI), no error banners. +2. Agent healthy + version: `$SSH felhom-pve "systemctl is-active felhom-agent; /usr/local/bin/felhom-agent --version"` → `active`, `0.30.0`. +3. Headroom (deploys pull images — **bound to ≤3 small apps**): + - Docker-data volume free: dashboard storage bars, or `$SSH felhom-pve "pct exec 9201 -- df -h /var/lib/docker /"`. Need comfortably above the v0.58 reserve (`max(5GB,10%)`) or deploys will be gated **507**. + - RAM: `$SSH felhom-pve "pct exec 9201 -- free -h"`. + - Disk list sane: `curl -s https://felhom.demo-felhom.eu/api/disks` → felhom-usb (user-data, data_bearing), local/local-lvm (system), felhom-pbs (backup). +4. Record current deployed apps (so cleanup is unambiguous): dashboard "Alkalmazások", or `curl -s https://felhom.demo-felhom.eu/api/stacks/rescan` then the stacks list. (actualbudget is expected already deployed.) + - **Good:** all green/healthy; free space well above reserve. **Watch for:** any app stuck "Telepítés alatt" (deploying) from a prior run — note and resolve before starting. + +--- + +## 2. Deploy flow — happy path (stresses CTRL-T2-1, normal case) + +Pick **two small apps** not currently deployed (suggest: `vikunja`, `mealie` — small images; avoid immich/nextcloud/paperless which are large pulls). UI: **Alkalmazások → → Telepítés**, or API. + +1. Deploy app #1 via UI (fill any required fields, generate passwords where prompted), click **Telepítés**. + - Watch the **3-step progress panel** (config → containers → health) and the live polling + (`GET /api/stacks/` every 3s). API equivalent: `curl -s -X POST https://felhom.demo-felhom.eu/api/stacks//deploy -H 'Content-Type: application/json' -d '{"values":{}}'` then poll `GET /api/stacks/`. + - **Good:** progresses config→containers→health; ends `running`/healthy within ~120s; the card flips to deployed; no "Telepítés" button reappears mid-pull (in-memory Deployed=true during pull). + - **Watch for:** stuck at a step, health-probe never going green (check the app's healthcheck tool exists), confusing Hungarian labels, the deploy gate returning **507** (insufficient Docker-data headroom — expected if low on space; note the banner wording). +2. Deploy app #2; same checks. +3. Confirm on disk the durable record is correct (CTRL-T2-1, happy case): `$SSH felhom-pve "pct exec 9201 -- docker exec felhom-controller cat /opt/docker/stacks//app.yaml | grep deployed"` → `deployed: true` (only after success). + - **Good:** `deployed: true` on disk after a successful deploy. **Watch for:** secrets appearing in plaintext in app.yaml (they must be `enc:`-prefixed — H10/encryption check). + +--- + +## 3. Deploy crash-window probe (THE CTRL-T2-1 test) + +Goal: prove a crash during the image-pull window leaves the stack **NOT-deployed and redeployable**, not ghost-stuck. + +1. Pick a **third app with a non-trivial image pull** (so the pull window is a few seconds — e.g. `paperless-ngx` if space allows, else `mealie`). Start the deploy (UI Telepítés or API POST), and **immediately** — while it is still pulling (status `deploying`, before `running`) — kill the controller: + - `$SSH felhom-pve "pct exec 9201 -- docker kill felhom-controller"` **[operator: time this during the pull]** + - The bootstrap service (`felhom-controller-bootstrap.service`) restarts it within seconds. Confirm back up: `curl -s https://felhom.demo-felhom.eu/api/health`. +2. After restart, check the stack state: + - On disk: `$SSH felhom-pve "pct exec 9201 -- docker exec felhom-controller cat /opt/docker/stacks//app.yaml | grep deployed"` → **`deployed: false`** (transitional — the fix). + - UI/API: `GET /api/stacks/` → state `not_deployed` (the card shows **Telepítés**, not a ghost "deployed"). +3. **Redeploy** the same app — it must be **allowed** (no "already deployed; use update instead" refusal) and complete normally. + - **Good:** post-crash the app reads not-deployed and redeploys cleanly. **PRE-FIX behaviour (must NOT occur):** app.yaml `deployed: true` with no containers, and redeploy refused — that's the ghost-stuck regression the fix removes. + - **Watch for:** any orphaned containers from the killed pull (`docker ps -a`); the half-pulled image is fine (compose re-pulls). + +--- + +## 4. Import flow — round-trip + negative (THE CTRL-001 test) + +1. **Export** a deployed app to a `.fab` bundle. UI: app page → **Exportálás** (or API `POST /api/export/start` with `{"stack_name":"","drive":""}`, poll `GET /api/export/status`). The `.fab` lands under the chosen drive's `exports/`. + - **Good:** export completes; a `.fab` appears; manifest readable via `POST /api/export/manifest`. +2. **Re-import** the same `.fab`: UI import flow (or `POST /api/export/import` with the file path), poll `GET /api/export/import/status`. + - **Good:** imports, recreates the stack, data restored; fail-closed data-key gate honored if the app has a data-encrypting key. +3. **Negative — path traversal (CTRL-001):** craft a hostile `.fab` and confirm it is **rejected at parse**, not written. + - Build a minimal bundle whose `manifest.json` has `"app_name":"../evil"` (and/or an `hdd_subdirs` / `volume_names` entry with `../`). Place it under a registered `exports/` dir on the host: + `$SSH felhom-pve "pct exec 9201 -- docker exec felhom-controller sh -c 'ls /mnt/felhom-usb/exports/'"` to find the dir. + - Attempt import of the hostile bundle. + - **Good (the fix):** import **fails immediately** with a manifest/validation error; **no directory is created outside the stacks dir** (verify: `$SSH felhom-pve "pct exec 9201 -- docker exec felhom-controller ls -la /opt/docker/evil /etc/evil 2>/dev/null"` → nothing). **PRE-FIX (must NOT occur):** a dir/file written outside `/opt/docker/stacks/`. + - **Watch for:** the error message clarity (does the UI explain why it was rejected?). + +--- + +## 5. Backup — per-app + full (3-2-1 tiers) + +1. **Per-app / on-demand DB-dump + recovery-unit refresh:** trigger `POST /api/backup/run` (or the dashboard "Biztonsági mentés" action). Watch `GET /api/backup/status`. + - **Good:** completes; DB dumps written under `/backups/primary//db-dumps/`; recovery unit refreshed (`backups/primary//` has compose + manifest.json). +2. **Tier-2 off-drive copy:** `POST /api/backup/tier2`. (Single-drive demo → Tier-2 goes to the internal SSD for SMALL units only, DB/config-only, with the rootfs-headroom guard.) + - **Good:** the per-app "2. mentés" card in the UI populates (success → target drive, or an honest "needs 2nd HDD" no-target reason). **Watch for:** the rootfs-headroom guard correctly **refusing** rather than filling rootfs when the unit is too big for the SSD. +3. **Whole-guest (PBS) tier** is the agent's: confirm via the monitoring page that the PBS/offsite tier shows recent snapshots (the controller surfaces agent/PBS status; it does not run vzdump itself). + - **Good:** 3-2-1 visible — primary (on-drive recovery unit) + secondary (Tier-2 off-drive) + offsite (PBS). **Watch for:** any tier showing stale/never — note it. +4. Verify completion is reflected, not just started: re-check `GET /api/backup/status` returns `done` and the UI timestamps update. + +--- + +## 6. Wipe gate re-check (AGENT-001 — NON-destructive) + +Re-confirm the two refusals proven on 2026-06-13. **Do NOT send a matching confirmation for a real data drive.** Target the real data-bearing device `/dev/sdb1` (felhom-usb) — both calls REFUSE, nothing is formatted. + +1. **Refusal A — no confirmation:** `curl -s -w '\n%{http_code}\n' -X POST https://felhom.demo-felhom.eu/api/disks/format -H 'Content-Type: application/json' -d '{"device":"/dev/sdb1","fstype":"ext4"}'` + - **Good:** `formatted:false`, `needs_confirmation:true`, **HTTP 409**; no mkfs. +2. **Refusal B — wrong durable_id:** same call with `"confirmed":true,"durable_id":"byid:wwn-0xDEADBEEF-DOES-NOT-EXIST"`. + - **Good:** `formatted:false`, refused **409**; a non-matching confirmation does not authorize a wipe. +3. **Data-safety assertion:** `$SSH felhom-pve "findmnt /mnt/felhom-usb -o TARGET,SOURCE,FSTYPE; pct exec 9201 -- docker exec felhom-controller sh -c 'df -h /mnt/felhom-usb'"` → still mounted, used space unchanged. +4. **Happy-path destructive wipe** = **[HUMAN]** — never wipe a real/customer drive to test; covered by the agent unit test `retarget-mismatch-refused`. Only on a genuinely disposable blank device, supervised. **[DESTRUCTIVE — operator confirm]** + +--- + +## 7. Restore — DESTRUCTIVE **[DESTRUCTIVE — operator confirm]** + +Do this LAST, on a NON-critical app (e.g. one deployed in §2, not customer RomM data). Restore overwrites the app's current data. + +1. Pick an app with a backup/snapshot from §5. Note its current data state (log in, note an item) so you can confirm the restore round-trip. +2. **[operator confirm]** Restore: `POST /backup/restore` (web form: `stack_name=&snapshot_id=`), or the dashboard restore UI. + - **Good:** the app stops, restores volume tars / recovery unit, restarts; data returns to the snapshot state; the fail-closed data-key gate refuses (with a clear message) if a data-encrypting key can't be recovered — in which case STOP and do NOT force. + - **Watch for:** data loss vs the snapshot, secret-regeneration (must recover from the guest's own app.yaml, regenerate nothing), confusing progress/labels. +3. **[operator]** Any **delete/remove** of a stack to clean up the §2/§3 test apps is **[DESTRUCTIVE — operator confirm]** — protected stacks (traefik/cloudflared/filebrowser/controller) must remain unstoppable server-side; confirm a delete attempt on one of those is refused. + +--- + +## 8. Wrap-up + +- Tear down the test apps deployed in §2/§3 (operator-confirmed deletes), leaving the demo as found (actualbudget + RomM intact). +- Write `REPORT-e2e-live-drive-.md`: per-section ✓/✗, every UX/latency/error-handling finding, and any regression (especially: did §3 leave a ghost-deployed stack? did §4 write outside the stacks dir? did any wipe call format?). File new bugs against the audit record. +- Confirm final: controller `:0.60.0` healthy, agent `0.30.0` active, felhom-usb intact, no orphaned containers. + +### Order summary (non-destructive first) +1 Baseline → 2 Deploy happy → 3 Crash-window probe → 4 Import + traversal-negative → 5 Backup → 6 Wipe-gate refusals → **7 Restore [DESTRUCTIVE]** → 8 Wrap-up. diff --git a/SESSION-2026-06-13-cleanup.md b/SESSION-2026-06-13-cleanup.md index a459bb8..8c81395 100644 --- a/SESSION-2026-06-13-cleanup.md +++ b/SESSION-2026-06-13-cleanup.md @@ -23,5 +23,13 @@ Working/checkpoint log. Commit/push after each Part. Controller-only (agent v0.3 - M25 `-race` proven clean on the build server. M18/M19: notes + fix plan on `fix/m18-dump-validation-cache` / `fix/m19-stackname-crossref` (pushed, NOT deployed). - **Deployed v0.60.0** to demo 9201 (rollback ref v0.59.0). Verified: `0.60.0 Up (healthy)`, self-test 6/0, "Controller elindult (0.60.0)", public `/api/health` 200, no errors. No rollback. -## PART 2 — docs (in progress) -## PART 3 — runbook (pending) +## PART 2 — docs: DONE +- `felhom-agent/CLAUDE.md`: stale v0.13.0 banner → v0.30.0 current-state summary (agent `238ec02`). +- `felhom-controller/CLAUDE.md`: corrected the WRONG "bulk strip has NOT happened" status → de-priv DONE (slice 8C); `CONTEXT.md` → v0.60.0 banner (ctrl `c67c53b`). +- Doc-gaps closed (felhom.eu `4c0eb2f`): hub v0.6.3→v0.11.0 (05 doc); onlyoffice:nextcloud occ internals (auth-hub-sync doc); metricsDBPath **verified** on the live guest (coincides with volume-backed data_dir → persists; hardcoding is latent fragility only) + storage doc note. REORG-NOTES updated (gaps 2/3 closed, 5 partial). Remaining gaps (volume-dump-consistency #1, selfUpdateAuthMiddleware #4, design-drafts 01/03/04) left flagged. +- controller/README.md slim-down: DEFERRED again (1987-line rewrite; banner already points to central docs — safety/budget). + +## PART 3 — runbook: DONE +- `felhom-controller/RUNBOOK-e2e-live-drive.md` written for tomorrow's supervised drive. 8 sections, non-destructive first (baseline → deploy happy → **CTRL-T2-1 crash-window probe** → **CTRL-001 import + traversal-negative** → backup 3-2-1 → **AGENT-001 wipe-gate refusals**), then **[DESTRUCTIVE] restore** last. Each step: exact UI path/API call, what "good" looks like, what to watch for. All routes verified against current source. + +## ALL PARTS DONE.