Files
felhom.eu/documentation/audits/CAMPAIGN-6B-2026-07-14-PROMPT.md
T
admin eb6b3bba56 @
CAMPAIGN-6A (Phase 1: reboot-driven NAS re-arm matrix) + 6B unattended continuation prompt

Supervised run, operator authorized unattended reboots mid-run. Phase 1 COMPLETE:
1A idle-remediate 3/3, 1B active skip-active+heal 2/2 (F13 absent), 1C F10 guest-reboot
boot-safety, 1D host-reboot re-arm survival (0 cycles, USB retirement-proof), 1E unit
drift self-repair. F8 confirmed. Findings C6-1 (skip-active no-op on pct reboot),
C6-2 (NAS-outage-across-reboot strands share until agent restart), C6-3 (fresh
all_squash export blocks docker chown). Phases 2-5 -> CAMPAIGN-6B (unattended prompt).
No credential/R/blob committed.

Claude-Session: https://claude.ai/code/session_01LbMm4T7Ayzs1unB9pN6Uqd
@
2026-07-14 12:17:26 +02:00

116 lines
8.8 KiB
Markdown

# CAMPAIGN-6B — UNATTENDED close-out of the C6 remainder (.fab upload circle · browser planes · backup tiers · regression)
**Class:** Unattended "no mercy" continuation of CAMPAIGN-6. Findings only — NO code fixes, NO
spec-writing however obvious. Output: `felhom.eu/documentation/audits/CAMPAIGN-6B-<date>.md`
(C3/4/5/6A structure). Continue the ledger + evidence at `180:~/campaign6/`. Record a new launch
seed. No fixed time budget — the run ends when the checklist is complete.
## THE OPERATING MODE — read this first (it is the whole point)
**This run is UNATTENDED. Viktor is NOT at the keyboard. Do NOT BLOCK-and-wait.**
- **All reboots — guest AND host, both boxes — are PRE-AUTHORIZED. Drive them yourself**
(`pct reboot 9201`, host `systemctl reboot`). This is a dev/test environment; there is no
evidence-banking delay and no per-reboot approval. The CAMPAIGN-6A supervised run already proved
the reboot plane is safe; 6B does not re-ask.
- **`exportfs` toggles on the campaign share are pre-authorized** (same safety line: campaign
`felhom-campaign6` export ONLY; never non-felhom exports; never stop DooPlex services).
- **`teszt_enroll` drive destructive ops are pre-authorized.** Demo's ~20 existing apps: start/stop/
backup only — never delete/wipe/redeploy-over. Campaign apps + drill: free chaos. Escrow
ceremonies: drill only; every R scratch + uncommitted.
- **DEFERRAL IS STILL BANNED, but so is BLOCKING.** Do not stop to ask permission. **HALT ONLY IF a
genuine, unrecoverable decision is required** (e.g. a truly ambiguous destructive choice with no
safe default, or missing credentials) — and even then, prefer the safe default and ledger it.
Running low on context is NOT a halt reason: instead, wrap what is done into the doc, write a
`CAMPAIGN-6C` continuation prompt, and stop cleanly (the 6A pattern). Every checklist item ends as
PASS / FAIL / FINDING / (operator-authorized) split-to-6C — never blank, never "deferred".
- If a reboot/host action does not come back, use the CAMPAIGN-6A morning-recovery steps and
ledger it as a FAIL with the diagnostic, then continue (continue-on-failure).
## Pre-existing state (left ready by CAMPAIGN-6A — verify at P0, don't trust this)
- controller **0.129.0** both guests · agent **0.88.0** both hosts · hub **0.54.0**. caps 63/63 both.
- **campaign6 NFS share still enrolled** on demo (`192.168.0.180:/mnt/5_hdd/felhom-campaign6`
`/mnt/felhom-drives/campaign6`, idle); **sonarr deployed on it (stopped)**; DooPlex export active.
- Credentials **unchanged from CAMPAIGN-4/5/6A**: controller AND hub dashboard = the campaign
credential. **Never commit it; ledger says "campaign credential"; rotate only at the very end.**
- P7 samplers (`c5demo`/`c5drill`, `c4-sampler.sh`) still running on both PVE hosts — reuse them.
- Access: `SSH=/c/Windows/System32/OpenSSH/ssh.exe`. 180=DooPlex (passwordless sudo), felhom-pve=demo
host, root@192.168.0.152=drill host. Controllers driven via `docker exec <ctrl> curl 127.0.0.1:8080`
(real login→CSRF) or claude-in-chrome for the browser planes.
## Findings already banked in 6A (do NOT re-hunt; watch for related regressions)
- **C6-1 (LOW-MED):** `skip-active` is a no-op on `pct reboot` — the guest-hook heal re-arm carries
every active-share reboot; the "fresh namespaces inherit real mounts" log is misleading.
- **C6-2 (MED):** a NAS outage spanning a guest reboot can strand the share `failed` until an agent
restart (or periodic sweep — timing unmeasured; **6B: measure the periodic-sweep interval** if a
cheap opportunity arises).
- **C6-3 (LOW, setup):** the agent/wizard don't pre-create an app's userdata tree on a fresh
`all_squash` NFS export → docker chown fails → container stuck `Created` until pre-created.
- **F-A / F-B / F-C are already confirmed-fixed-live** (CAMPAIGN-5, same 0.129.0 fleet). 6B may
spot-re-confirm opportunistically but need not re-prove them.
## P0 (fast) — verify the pre-existing state, record a seed, then GO
Versions both; caps 63/63; campaign6 enrolled+health; sonarr present; escrow states; samplers alive.
Push to `180:~/campaign6/evidence/P0-6B/`.
## PHASE 2 — the `.fab` full circle, UPLOAD half (demo; twice-deferred flagship)
Fresh **campaign** app + **~4 GiB** `/dev/urandom` (varied sizes incl one >1 GiB); sha256 manifest.
Export → `.fab` (progress honesty + atomicity). Download twice — browser/LAN AND `curl` with the
session cookie **through the real Cloudflare edge** (`--resolve <host>:443:<CF-public-IP>` — Pi-hole
split-horizon bypasses the edge otherwise); both hashes == the docker-cp reference. Re-prove the edge
cap: single >100 MB POST → **413 from Cloudflare**. Delete the campaign app + data. Upload back —
browser/LAN AND `curl --resolve` through the edge in 64 MiB chunks. Import/restore; **byte-compare
every file vs the manifest → zero mismatches = PASS.** Edge volley (live over the unit tests): cancel
mid-upload → abort + `.part` gone; controller restart mid-upload → startup GC count logged; offset
replay → 409+resync; oversize-vs-free → Hungarian refusal (both numbers) **AND F-A real multi-GB
size**; collision → " (1)" then " (2)"; second concurrent init → 409; wrong extension refused; idle
15-min abort (park one, return). (Note the storage-bearing box + a proper export are required — see
C6-3; pre-create the userdata tree.)
## PHASE 3 — browser planes (claude-in-chrome; twice-deferred)
Session must be started AFTER the chrome bridge connected. Campaign credential for login.
- **3A escrow wizard, full browser pass (drill):** preflight all-green + Hungarian details (no raw
English leak), warnings, re-auth, run, reveal, **typed-back with the two highlighted words** (read
from screen), manual hide/show toggle, finish → auto-confirm flips + hub row hash matches
(server-side check via 180). Then re-claim → 410 UI; unclaimed → TTL → `unclaimed_void` screen;
**F-C live** (`phase:none` claim → clean 4xx UI, not 502); out-of-band CLI ceremony w/o staged
secret → stale card fires → wizard clears it.
- **3B DOM / native-alert / session sweep (both boxes):** grep served DOM for `alert(`/`confirm(`
(must be zero); session expiry mid-wizard and mid-upload → JSON 401 on `/api/`, redirect on pages;
CSRF stale-token on storage-wizard/escrow/upload POSTs → 403 JSON; zero-toggle honesty + inline
two-step confirms; **F-B live** quick re-check.
- **3C hub 8-tab customer-detail ring (hub credential = campaign credential):** hash-nav across all
8 tabs; auto-refresh scoped to live tabs; **dirty-form suppression** (start editing → refresh
holds); events tab under volume; no stale-host deletion against real hosts (synthetic row only).
## PHASE 4-rest — backup tiers depth + NAS integrity
- **F7 mid-backup NAS cut:** `POST /api/backup/run` on a NAS app; `exportfs -u` at T+~6s; the
last-good volume dump must survive (tmp+rename) — no 0-byte artifact; `success:false` is the only
signal. (Give the app enough data that the dump takes >6 s.)
- Four backup sub-pages truth vs live; offsite card three honest states; Tier-3 "Távoli mentés most"
additive + quota bar; **restic stale-lock self-heal** (kill controller mid-offsite-run → next run
self-heals the lock); tier-1 replace-semantics; offsite restore-to-verify to a scratch target
(campaign app) — byte-identical; volume-only app gets its tier-2 secondary (F6 fix); per-app
toggles round-trip; snapshot coherence across pages. (F8 already CONFIRMED in 6A — cite it.)
## PHASE 5 — regression spot-checks (fast)
- Agent-restart re-arm still logs per-share verdicts (6A re-showed it via 1E `netmigrate` — one
quick re-confirm). **F1/F2 residue** after the campaign share's final removal — zero mounts/units/
dirs, host+guest. **F4** `mapped_uid:101000` → friendly Hungarian 400 (not `agent_error`).
Dead-app alert + email on a killed campaign container (cooldown math vs the run's emails).
## Wrap — completeness gate (same as 6A)
Checklist table of EVERY 6B item = PASS / FAIL / FINDING / (operator-authorized) split-to-6C. Verdict;
ranked findings + repros; timings; deviations; box state (both boxes; "campaign credential active —
Viktor rotates"; drill fresh-R note); morning recovery; evidence index. Commit the doc + overwrite
`felhom.eu/REPORT.md`. **No credential, no R, no blob committed.**
## Final cleanup (unattended — do it, no BLOCK; show before/after in the ledger)
Remove campaign6 share + campaign apps (product flow); `sudo exportfs -u
192.168.0.162:/mnt/5_hdd/felhom-campaign6 && sudo rm -rf /mnt/5_hdd/felhom-campaign6` on 180 (confirm
`exportfs -v` back to the P0 baseline, felhom-data intact); remove scratch `.fab` + injected data;
stop the samplers (`pkill -f c4-sampler.sh` both PVE hosts, `pkill -f hub-sampler.sh` on 180). Leave
the campaign credential in place and note in the doc that **Viktor rotates it** — that is the one
human step that remains after 6B.