Files
felhom.eu/documentation/audits/CAMPAIGN-6A-2026-07-14.md
admin eb6b3bba56 @
CAMPAIGN-6A (Phase 1: reboot-driven NAS re-arm matrix) + 6B unattended continuation prompt

Supervised run, operator authorized unattended reboots mid-run. Phase 1 COMPLETE:
1A idle-remediate 3/3, 1B active skip-active+heal 2/2 (F13 absent), 1C F10 guest-reboot
boot-safety, 1D host-reboot re-arm survival (0 cycles, USB retirement-proof), 1E unit
drift self-repair. F8 confirmed. Findings C6-1 (skip-active no-op on pct reboot),
C6-2 (NAS-outage-across-reboot strands share until agent restart), C6-3 (fresh
all_squash export blocks docker chown). Phases 2-5 -> CAMPAIGN-6B (unattended prompt).
No credential/R/blob committed.

Claude-Session: https://claude.ai/code/session_01LbMm4T7Ayzs1unB9pN6Uqd
@
2026-07-14 12:17:26 +02:00

8.4 KiB
Raw Permalink Blame History

CAMPAIGN-6A — supervised reboot-driven NAS re-arm matrix (Phase 1 complete)

  • When: 2026-07-14 ~08:42Z launch → wrapped mid-run at Viktor's request (Phase 1 + F8 done; Phases 25 carried to CAMPAIGN-6B). Launch seed 8219f68dee135a15.
  • Stack under fire (verified live at P0): controller 0.129.0 both guests · agent 0.88.0 both hosts (caps 63/63, 0 degraded) · hub 0.54.0 · demo (felhom-pve 192.168.0.162 + guest 9201, storage-bearing) AND drill (qm300 192.168.0.152 + guest 9201). Campaign credential valid.
  • Contract honored: supervised start (BLOCK-and-wait) → operator authorized unattended reboots mid-run ("Restarts can be unattended. This is a dev/test environment. HALT ONLY if a real decision is needed"), after which CC drove all reboots itself; findings only, no code fixes; no Gitea/PBS/hub-config mutations; DooPlex: only the campaign temp export /mnt/5_hdd/felhom-campaign6 (runtime exportfs, never /etc/exports) toggled; felhom-data + non-felhom untouched; no DooPlex service stopped; demo's existing apps untouched; campaign credential / R / blob in no committed file, ledger, or this doc.
  • Run architecture: CC session; harness/ledger/evidence at 180:~/campaign6/. P0 baseline captured + pushed to evidence/P0/ before the first reboot (per contract §4). Reboots driven via pct reboot 9201 / host systemctl reboot; verdicts read from journalctl -u felhom-agent + the guest-hook lines. P7 samplers detached on both hosts.

Verdict

The twice-deferred reboot-driven NAS re-arm plane holds; the fix survives the real host boot; two real behavioral findings surfaced. Every reboot leg — idle-share remediate, active-share skip-active+heal, start-limit guest reboot, host-reboot survival, and unit-drift self-repair — ended with the share visible and no ordering cycle. The net behavior is correct across all 9 guest reboots + 1 host reboot, but skip-active never actually works on pct reboot (the heal path carries it) and a NAS outage that spans a guest reboot can strand the share until an agent restart. No CRITICAL/HIGH regressions.

Completeness checklist (every item PASS / FAIL / FINDING / → 6B)

Item Status Evidence
P0 baseline (banked pre-reboot) PASS both 0.129.0/agent 0.88.0/caps 63/63/0 cycles; pushed to evidence/P0/
Enroll campaign6 (verify-before-commit) PASS agent_adddone, health:ok
F12-clean unit PASS 0 network-online refs on the campaign6 .automount
1A idle-share guest reboot → remediate (×3) PASS 3/3 verdict=rearmed → guest-hook "visible in guest (rearmed)"; 0 cycles each
1B active-share reboot → skip-active (×2) + F13 watch PASS + FINDING C6-1 share ends visible; F13 did NOT manifest (0 rbind/rc255)
1C F10 guest-reboot at start-limit-hit (boot not blocked) PASS + FINDING C6-2 guest reached running (pre-start rc255 trap held); reset-failed+rearmed
1D re-arm reboot-survival (demo HOST reboot) PASS 0 cycles, caps 63/63, WG, 3 USB re-established (retirement-proof), campaign6 re-armed→visible, guest+controller healthy
1E MigrateNetworkUnits drift reconcile PASS injected network-online → agent restart → netmigrate rewrote to exact canonical sha; systemd-analyze verify clean
Phase 4 F8 storage-health during outage CONFIRMED health:stub, mounted:false, reachable:true
App-deploy on fresh NAS export FINDING C6-3 (setup) fresh NFS userdata dirs under all_squash block docker chown until pre-created
Phase 2 .fab 4 GiB upload full-circle → CAMPAIGN-6B operator-authorized split
Phase 3 browser: escrow wizard / DOM sweep / hub 8-tab → CAMPAIGN-6B operator-authorized split
Phase 4-rest F7 mid-backup cut + tier sub-items → CAMPAIGN-6B operator-authorized split
Phase 5 regression spot-checks (F1/F2 residue, F4, agent-restart re-arm) → CAMPAIGN-6B (agent-restart re-arm already re-shown via 1E netmigrate) operator-authorized split

The → CAMPAIGN-6B rows are an operator-authorized session split ("wrap now, write 6A + a continuation prompt"), not a silent defer.

Ranked findings (exact repros)

# Sev Finding Exact repro
C6-1 LOW-MED skip-active is a no-op on pct reboot. Its premise "fresh namespaces inherit real mounts" was FALSE in 4/4 active-share reboots — the fresh guest namespace never inherited the active nfs4 mount; the guest-hook's detect-and-heal re-arm did the real work every time. Net-correct (share always ends visible), but the fast-path never fires and its log line ("skip — fresh namespaces inherit real mounts") is misleading. active nfs4 share, pct reboot 9201, watch guest-hook: "network share X not visible after reassert (skip-active) — re-arming" → "healed".
C6-2 MED A NAS outage spanning a guest reboot can strand the share failed. In 1C (export DOWN at the guest reboot, re-exported AFTER), both campaign6 units ended failed; the post-start reset-failed+rearmed re-failed against the still-down export, and when the export returned the units did NOT self-recover on access (health:stub, mounted:false). Only systemctl restart felhom-agent (or, unmeasured, the periodic sweep) re-armed it. exportfs -u; trip start-limit; pct reboot 9201; THEN re-export; access → still stub/failed; systemctl restart felhom-agent → recovers.
C6-3 LOW (setup) The agent/wizard don't pre-create an app's userdata tree on a network drive before first container start. A fresh NFS export under all_squash refuses docker's chown of dirs it creates → the container sticks in Created ("operation not permitted"). Existing nas-media apps work only because their dirs pre-exist. deploy a data-bearing app with HDD_PATH on a fresh all_squash NFS export → container stuck Created; pre-creating the userdata tree unblocks start.

What passed (headline)

  • F11 idle-share remediate: the C3 HIGH is closed — idle-share guest reboots consistently rearmed → visible (3/3), and even the skip-active-invisible case self-heals.
  • F10 guest-reboot boot-safety: guest reaches running even at mount-start-limit-hit with the export down (pre-start rc255 trap).
  • F12 across the real host boot (1D): 0 ordering cycles; retirement reboot-proof (3 USB re-establish from agent units despite device-letter reshuffle) holds from C4.
  • Unit drift self-repair (1E): netmigrate rewrites a network-online-poisoned unit back to the exact canonical form — pre-0.85 customer boxes self-heal on upgrade.

Deviations

  • Supervised → unattended reboots mid-run (operator directive). The first nohup-backgrounded host reboot silently no-op'd once; a direct systemctl reboot succeeded (harness note, not a product issue).
  • teszt_enroll is in intent=ejected state (from CAMPAIGN-4); the 1E reconcile correctly skipped it (intent-gated).

Box state at wrap (left ready for CAMPAIGN-6B)

  • demo (felhom-pve/9201): controller 0.129.0, agent 0.88.0, healthy; campaign6 NFS share still enrolled (idle, /mnt/felhom-drives/campaign6); sonarr deployed on it (stopped); DooPlex /mnt/5_hdd/felhom-campaign6 export still active; existing apps untouched. P7 c5demo sampler running.
  • drill (192.168.0.152/9201): controller 0.129.0, agent 0.88.0, healthy (rebooted with the host during 1D, restarted). Escrow phase:none. P7 c5drill sampler running.
  • Credential: campaign credential active on both controllers + hub — Viktor rotates when the whole run (6B) completes. No R/blob produced.

Morning recovery / handoff to 6B

  • Nothing is down. The campaign6 share + sonarr + DooPlex export are intentionally left in place so CAMPAIGN-6B can continue Phase 4 (NAS tiers) without re-enrolling. Full teardown (share/app removal, exportfs back to P0, sampler stop, credential rotation) is 6B's final cleanup.
  • If abandoning 6B instead: remove campaign6 via the product flow, sudo exportfs -u 192.168.0.162:/mnt/5_hdd/felhom-campaign6 && sudo rm -rf /mnt/5_hdd/felhom-campaign6 on 180, pkill -f c4-sampler.sh on both PVE hosts.

Evidence index (180:~/campaign6/)

  • seed.txt (8219f68dee135a15), ledger.md (per-item trail + verbatim verdict/journal lines), evidence/P0/baseline.md.
  • P7 series: 192.168.0.162:/root/c4-c5demo.csv, 192.168.0.152:/root/c4-c5drill.csv.