@
CAMPAIGN-6A (Phase 1: reboot-driven NAS re-arm matrix) + 6B unattended continuation prompt Supervised run, operator authorized unattended reboots mid-run. Phase 1 COMPLETE: 1A idle-remediate 3/3, 1B active skip-active+heal 2/2 (F13 absent), 1C F10 guest-reboot boot-safety, 1D host-reboot re-arm survival (0 cycles, USB retirement-proof), 1E unit drift self-repair. F8 confirmed. Findings C6-1 (skip-active no-op on pct reboot), C6-2 (NAS-outage-across-reboot strands share until agent restart), C6-3 (fresh all_squash export blocks docker chown). Phases 2-5 -> CAMPAIGN-6B (unattended prompt). No credential/R/blob committed. Claude-Session: https://claude.ai/code/session_01LbMm4T7Ayzs1unB9pN6Uqd @
This commit is contained in:
@@ -2,27 +2,29 @@
|
||||
|
||||
> **Overwrite** this file with a summary of the most recent task only (uniform with the other repos; not cumulative). The cumulative hub history lives in [hub/CHANGELOG.md](hub/CHANGELOG.md); the scripts history lives in [scripts/CHANGELOG.md](scripts/CHANGELOG.md).
|
||||
|
||||
## CAMPAIGN-5 — NAS re-arm ring core + v0.129.0 fix live-proof — 2026-07-14
|
||||
## CAMPAIGN-6A — supervised reboot-driven NAS re-arm matrix (Phase 1) — 2026-07-14
|
||||
|
||||
Full report: [documentation/audits/CAMPAIGN-5-2026-07-14.md](documentation/audits/CAMPAIGN-5-2026-07-14.md). Launch seed `22cf2c0983034e61`. Findings-only.
|
||||
Full report: [documentation/audits/CAMPAIGN-6A-2026-07-14.md](documentation/audits/CAMPAIGN-6A-2026-07-14.md). Continuation prompt: [CAMPAIGN-6B-2026-07-14-PROMPT.md](documentation/audits/CAMPAIGN-6B-2026-07-14-PROMPT.md). Seed `8219f68dee135a15`. Findings-only. Wrapped mid-run at operator request (Phases 2–5 → CAMPAIGN-6B).
|
||||
|
||||
### Verdict
|
||||
The v0.129.0 fixes are **confirmed fixed live**, and the agent v0.85 NAS **re-arm plane holds** on the core matrix. No CRITICAL/HIGH regressions in the exercised scope.
|
||||
The twice-deferred reboot-driven NAS re-arm plane **holds**; the fix survives the real host boot; two real behavioral findings surfaced. 9 guest reboots + 1 host reboot, all ending share-visible with 0 ordering cycles.
|
||||
|
||||
### v0.129.0 fixes — all confirmed fixed live (the campaign's mandate)
|
||||
- **F-B:** drill 0.129.0 — 6 direct wrong logins (no XFF, distinct ports) → limiter engages on attempt 6; proxied path also limits (pre-fix: direct never limited).
|
||||
- **F-A:** demo download estimate → `data_size_bytes:74375`, `size_unknown:false` (real container-view du; pre-fix: 0).
|
||||
- **F-C:** drill escrow `phase:none` claim → HTTP 404 with clean Hungarian message (pre-fix: 502).
|
||||
### Completeness (Phase 1 = every reboot leg)
|
||||
| Item | Result |
|
||||
|---|---|
|
||||
| P0 (banked pre-reboot), enroll campaign6, F12-clean | PASS |
|
||||
| 1A idle-share guest reboot → remediate ×3 | **PASS 3/3** |
|
||||
| 1B active-share reboot ×2 + F13 watch | **PASS + FINDING C6-1**; F13 absent |
|
||||
| 1C F10 guest-reboot at start-limit (boot not blocked) | **PASS + FINDING C6-2** |
|
||||
| 1D re-arm reboot-survival (demo HOST reboot) | **PASS** (0 cycles, caps 63/63, USB retirement-proof, share re-armed) |
|
||||
| 1E MigrateNetworkUnits drift reconcile | **PASS** (exact canonical rewrite, verify clean) |
|
||||
| Phase 4 F8 storage-health | **CONFIRMED** (stub/false/true) |
|
||||
| Phases 2/3/4-rest/5 | **→ CAMPAIGN-6B** (operator-authorized split) |
|
||||
|
||||
### NAS re-arm ring core (fresh campaign NFS share, DooPlex→demo) — evidenced PASS
|
||||
- Enrollment (add wizard, verify-before-commit): PASS. Unit F12-clean (no network-online ordering).
|
||||
- **F10 core:** `exportfs -u` → force-unmount → 6 accesses → `mount-start-limit-hit`; `systemctl restart felhom-agent` → **reset-failed on both units** + `enable --now`, `verdict=reset-failed+rearmed`; access re-mounts cleanly, **guest stayed up**. PASS.
|
||||
- **F9 no-empty-sweep:** every share emits one verdict line (campaign5 reset-failed+rearmed, nas-media skip-active). PASS.
|
||||
- **F1/F2 residue:** clean removal — 0 mounts/units/failed/dirs on host+guest (contra C3). PASS.
|
||||
- **F8:** during the real outage → `health:"stub", mounted:false, reachable:true` — more honest than C3, but `reachable:true` still server-level. Observation.
|
||||
|
||||
### Deferred (ready procedures in the audit doc)
|
||||
The reboot half of the matrix (F10 guest-reboot, F11 idle/active on reboot, re-arm reboot-survival), F7 mid-backup cut, E drift-reconcile, the `.fab` upload full-circle, the browser escrow wizard, DOM sweep, backups tiers, hub 8-tab ring — deferred for runway (browser available; no fabrication).
|
||||
### Findings
|
||||
- **C6-1 (LOW-MED):** `skip-active` is a no-op on `pct reboot` — the guest-hook heal re-arm carries every active-share reboot (4/4 not inherited); log line misleading.
|
||||
- **C6-2 (MED):** a NAS outage spanning a guest reboot can strand the share `failed` until an agent restart; access alone doesn't recover it.
|
||||
- **C6-3 (LOW, setup):** agent/wizard don't pre-create an app's userdata tree on a fresh `all_squash` NFS export → docker chown fails → container stuck `Created`.
|
||||
|
||||
### Box state
|
||||
Nothing down. Campaign5 NFS export + dir fully removed from DooPlex (no `/etc/exports` change ever made); demo apps untouched; nas-media intact. Both boxes healthy on 0.129.0. Campaign credential active — **Viktor rotates.** No R/blob produced. Samplers running (stop: `pkill -f c4-sampler.sh` on both PVE hosts, `pkill -f hub-sampler.sh` on 180). Evidence at `180:~/campaign5/`.
|
||||
Nothing down. **campaign6 share + sonarr + the DooPlex export left in place** for CAMPAIGN-6B (Phase 4 needs the NAS); demo apps untouched; both boxes healthy on 0.129.0; samplers running. Campaign credential active on controllers + hub — **Viktor rotates at the end of 6B.** No R/blob produced. Evidence at `180:~/campaign6/`.
|
||||
|
||||
@@ -0,0 +1,62 @@
|
||||
# CAMPAIGN-6A — supervised reboot-driven NAS re-arm matrix (Phase 1 complete)
|
||||
|
||||
- **When:** 2026-07-14 ~08:42Z launch → wrapped mid-run at Viktor's request (Phase 1 + F8 done; Phases 2–5 carried to **CAMPAIGN-6B**). Launch seed `8219f68dee135a15`.
|
||||
- **Stack under fire (verified live at P0):** controller **0.129.0** both guests · agent **0.88.0** both hosts (caps **63/63**, 0 degraded) · hub **0.54.0** · demo (felhom-pve 192.168.0.162 + guest 9201, storage-bearing) AND drill (qm300 192.168.0.152 + guest 9201). Campaign credential valid.
|
||||
- **Contract honored:** supervised start (BLOCK-and-wait) → **operator authorized unattended reboots mid-run** ("Restarts can be unattended. This is a dev/test environment. HALT ONLY if a real decision is needed"), after which CC drove all reboots itself; findings only, no code fixes; no Gitea/PBS/hub-config mutations; **DooPlex: only the campaign temp export `/mnt/5_hdd/felhom-campaign6` (runtime `exportfs`, never `/etc/exports`) toggled; felhom-data + non-felhom untouched; no DooPlex service stopped**; demo's existing apps untouched; campaign credential / R / blob in **no** committed file, ledger, or this doc.
|
||||
- **Run architecture:** CC session; harness/ledger/evidence at `180:~/campaign6/`. P0 baseline captured + pushed to `evidence/P0/` **before the first reboot** (per contract §4). Reboots driven via `pct reboot 9201` / host `systemctl reboot`; verdicts read from `journalctl -u felhom-agent` + the guest-hook lines. P7 samplers detached on both hosts.
|
||||
|
||||
## Verdict
|
||||
|
||||
**The twice-deferred reboot-driven NAS re-arm plane holds; the fix survives the real host boot; two real behavioral findings surfaced.** Every reboot leg — idle-share remediate, active-share skip-active+heal, start-limit guest reboot, host-reboot survival, and unit-drift self-repair — ended with the share visible and no ordering cycle. The net behavior is correct across all 9 guest reboots + 1 host reboot, but **`skip-active` never actually works on `pct reboot` (the heal path carries it)** and **a NAS outage that spans a guest reboot can strand the share until an agent restart.** No CRITICAL/HIGH regressions.
|
||||
|
||||
## Completeness checklist (every item PASS / FAIL / FINDING / → 6B)
|
||||
|
||||
| Item | Status | Evidence |
|
||||
|---|---|---|
|
||||
| P0 baseline (banked pre-reboot) | **PASS** | both 0.129.0/agent 0.88.0/caps 63/63/0 cycles; pushed to `evidence/P0/` |
|
||||
| Enroll campaign6 (verify-before-commit) | **PASS** | `agent_add`→`done`, health:ok |
|
||||
| F12-clean unit | **PASS** | 0 network-online refs on the campaign6 `.automount` |
|
||||
| **1A** idle-share guest reboot → remediate (×3) | **PASS 3/3** | `verdict=rearmed` → guest-hook "visible in guest (rearmed)"; 0 cycles each |
|
||||
| **1B** active-share reboot → skip-active (×2) + F13 watch | **PASS + FINDING C6-1** | share ends visible; **F13 did NOT manifest** (0 rbind/rc255) |
|
||||
| **1C** F10 guest-reboot at start-limit-hit (boot not blocked) | **PASS + FINDING C6-2** | guest reached running (pre-start rc255 trap held); `reset-failed+rearmed` |
|
||||
| **1D** re-arm reboot-survival (demo HOST reboot) | **PASS** | 0 cycles, caps 63/63, WG, **3 USB re-established (retirement-proof)**, campaign6 re-armed→visible, guest+controller healthy |
|
||||
| **1E** MigrateNetworkUnits drift reconcile | **PASS** | injected `network-online` → agent restart → `netmigrate` rewrote to **exact canonical sha**; `systemd-analyze verify` clean |
|
||||
| **Phase 4 F8** storage-health during outage | **CONFIRMED** | `health:stub, mounted:false, reachable:true` |
|
||||
| App-deploy on fresh NAS export | **FINDING C6-3 (setup)** | fresh NFS userdata dirs under `all_squash` block docker chown until pre-created |
|
||||
| **Phase 2** `.fab` 4 GiB upload full-circle | **→ CAMPAIGN-6B** | operator-authorized split |
|
||||
| **Phase 3** browser: escrow wizard / DOM sweep / hub 8-tab | **→ CAMPAIGN-6B** | operator-authorized split |
|
||||
| **Phase 4-rest** F7 mid-backup cut + tier sub-items | **→ CAMPAIGN-6B** | operator-authorized split |
|
||||
| **Phase 5** regression spot-checks (F1/F2 residue, F4, agent-restart re-arm) | **→ CAMPAIGN-6B** (agent-restart re-arm already re-shown via 1E `netmigrate`) | operator-authorized split |
|
||||
|
||||
> The `→ CAMPAIGN-6B` rows are an **operator-authorized session split** ("wrap now, write 6A + a continuation prompt"), not a silent defer.
|
||||
|
||||
## Ranked findings (exact repros)
|
||||
|
||||
| # | Sev | Finding | Exact repro |
|
||||
|---|-----|---------|-------------|
|
||||
| **C6-1** | LOW-MED | **`skip-active` is a no-op on `pct reboot`.** Its premise "fresh namespaces inherit real mounts" was FALSE in **4/4** active-share reboots — the fresh guest namespace never inherited the active nfs4 mount; the guest-hook's detect-and-heal re-arm did the real work every time. Net-correct (share always ends visible), but the fast-path never fires and its log line ("skip — fresh namespaces inherit real mounts") is misleading. | active nfs4 share, `pct reboot 9201`, watch guest-hook: "network share X not visible after reassert (skip-active) — re-arming" → "healed". |
|
||||
| **C6-2** | MED | **A NAS outage spanning a guest reboot can strand the share `failed`.** In 1C (export DOWN at the guest reboot, re-exported AFTER), both campaign6 units ended `failed`; the post-start `reset-failed+rearmed` re-failed against the still-down export, and when the export returned the units did NOT self-recover on access (`health:stub, mounted:false`). Only `systemctl restart felhom-agent` (or, unmeasured, the periodic sweep) re-armed it. | `exportfs -u`; trip start-limit; `pct reboot 9201`; THEN re-export; access → still stub/failed; `systemctl restart felhom-agent` → recovers. |
|
||||
| **C6-3** | LOW (setup) | **The agent/wizard don't pre-create an app's userdata tree on a network drive before first container start.** A fresh NFS export under `all_squash` refuses docker's chown of dirs it creates → the container sticks in `Created` ("operation not permitted"). Existing nas-media apps work only because their dirs pre-exist. | deploy a data-bearing app with `HDD_PATH` on a fresh `all_squash` NFS export → container stuck `Created`; pre-creating the userdata tree unblocks start. |
|
||||
|
||||
## What passed (headline)
|
||||
- **F11 idle-share remediate:** the C3 HIGH is closed — idle-share guest reboots consistently `rearmed` → visible (3/3), and even the skip-active-invisible case self-heals.
|
||||
- **F10 guest-reboot boot-safety:** guest reaches running even at `mount-start-limit-hit` with the export down (pre-start rc255 trap).
|
||||
- **F12 across the real host boot (1D):** 0 ordering cycles; retirement reboot-proof (3 USB re-establish from agent units despite device-letter reshuffle) holds from C4.
|
||||
- **Unit drift self-repair (1E):** `netmigrate` rewrites a network-online-poisoned unit back to the exact canonical form — pre-0.85 customer boxes self-heal on upgrade.
|
||||
|
||||
## Deviations
|
||||
- Supervised → unattended reboots mid-run (operator directive). The first `nohup`-backgrounded host reboot silently no-op'd once; a direct `systemctl reboot` succeeded (harness note, not a product issue).
|
||||
- teszt_enroll is in `intent=ejected` state (from CAMPAIGN-4); the 1E reconcile correctly skipped it (intent-gated).
|
||||
|
||||
## Box state at wrap (left ready for CAMPAIGN-6B)
|
||||
- **demo (felhom-pve/9201):** controller 0.129.0, agent 0.88.0, healthy; **campaign6 NFS share still enrolled** (idle, `/mnt/felhom-drives/campaign6`); **sonarr deployed on it (stopped)**; DooPlex `/mnt/5_hdd/felhom-campaign6` export still active; existing apps untouched. P7 `c5demo` sampler running.
|
||||
- **drill (192.168.0.152/9201):** controller 0.129.0, agent 0.88.0, healthy (rebooted with the host during 1D, restarted). Escrow `phase:none`. P7 `c5drill` sampler running.
|
||||
- **Credential:** campaign credential active on both controllers + hub — **Viktor rotates when the whole run (6B) completes.** No R/blob produced.
|
||||
|
||||
## Morning recovery / handoff to 6B
|
||||
- Nothing is down. The campaign6 share + sonarr + DooPlex export are intentionally **left in place** so CAMPAIGN-6B can continue Phase 4 (NAS tiers) without re-enrolling. Full teardown (share/app removal, exportfs back to P0, sampler stop, credential rotation) is 6B's final cleanup.
|
||||
- If abandoning 6B instead: remove campaign6 via the product flow, `sudo exportfs -u 192.168.0.162:/mnt/5_hdd/felhom-campaign6 && sudo rm -rf /mnt/5_hdd/felhom-campaign6` on 180, `pkill -f c4-sampler.sh` on both PVE hosts.
|
||||
|
||||
## Evidence index (`180:~/campaign6/`)
|
||||
- `seed.txt` (`8219f68dee135a15`), `ledger.md` (per-item trail + verbatim verdict/journal lines), `evidence/P0/baseline.md`.
|
||||
- P7 series: `192.168.0.162:/root/c4-c5demo.csv`, `192.168.0.152:/root/c4-c5drill.csv`.
|
||||
@@ -0,0 +1,115 @@
|
||||
# CAMPAIGN-6B — UNATTENDED close-out of the C6 remainder (.fab upload circle · browser planes · backup tiers · regression)
|
||||
|
||||
**Class:** Unattended "no mercy" continuation of CAMPAIGN-6. Findings only — NO code fixes, NO
|
||||
spec-writing however obvious. Output: `felhom.eu/documentation/audits/CAMPAIGN-6B-<date>.md`
|
||||
(C3/4/5/6A structure). Continue the ledger + evidence at `180:~/campaign6/`. Record a new launch
|
||||
seed. No fixed time budget — the run ends when the checklist is complete.
|
||||
|
||||
## THE OPERATING MODE — read this first (it is the whole point)
|
||||
|
||||
**This run is UNATTENDED. Viktor is NOT at the keyboard. Do NOT BLOCK-and-wait.**
|
||||
|
||||
- **All reboots — guest AND host, both boxes — are PRE-AUTHORIZED. Drive them yourself**
|
||||
(`pct reboot 9201`, host `systemctl reboot`). This is a dev/test environment; there is no
|
||||
evidence-banking delay and no per-reboot approval. The CAMPAIGN-6A supervised run already proved
|
||||
the reboot plane is safe; 6B does not re-ask.
|
||||
- **`exportfs` toggles on the campaign share are pre-authorized** (same safety line: campaign
|
||||
`felhom-campaign6` export ONLY; never non-felhom exports; never stop DooPlex services).
|
||||
- **`teszt_enroll` drive destructive ops are pre-authorized.** Demo's ~20 existing apps: start/stop/
|
||||
backup only — never delete/wipe/redeploy-over. Campaign apps + drill: free chaos. Escrow
|
||||
ceremonies: drill only; every R scratch + uncommitted.
|
||||
- **DEFERRAL IS STILL BANNED, but so is BLOCKING.** Do not stop to ask permission. **HALT ONLY IF a
|
||||
genuine, unrecoverable decision is required** (e.g. a truly ambiguous destructive choice with no
|
||||
safe default, or missing credentials) — and even then, prefer the safe default and ledger it.
|
||||
Running low on context is NOT a halt reason: instead, wrap what is done into the doc, write a
|
||||
`CAMPAIGN-6C` continuation prompt, and stop cleanly (the 6A pattern). Every checklist item ends as
|
||||
PASS / FAIL / FINDING / (operator-authorized) split-to-6C — never blank, never "deferred".
|
||||
- If a reboot/host action does not come back, use the CAMPAIGN-6A morning-recovery steps and
|
||||
ledger it as a FAIL with the diagnostic, then continue (continue-on-failure).
|
||||
|
||||
## Pre-existing state (left ready by CAMPAIGN-6A — verify at P0, don't trust this)
|
||||
|
||||
- controller **0.129.0** both guests · agent **0.88.0** both hosts · hub **0.54.0**. caps 63/63 both.
|
||||
- **campaign6 NFS share still enrolled** on demo (`192.168.0.180:/mnt/5_hdd/felhom-campaign6` →
|
||||
`/mnt/felhom-drives/campaign6`, idle); **sonarr deployed on it (stopped)**; DooPlex export active.
|
||||
- Credentials **unchanged from CAMPAIGN-4/5/6A**: controller AND hub dashboard = the campaign
|
||||
credential. **Never commit it; ledger says "campaign credential"; rotate only at the very end.**
|
||||
- P7 samplers (`c5demo`/`c5drill`, `c4-sampler.sh`) still running on both PVE hosts — reuse them.
|
||||
- Access: `SSH=/c/Windows/System32/OpenSSH/ssh.exe`. 180=DooPlex (passwordless sudo), felhom-pve=demo
|
||||
host, root@192.168.0.152=drill host. Controllers driven via `docker exec <ctrl> curl 127.0.0.1:8080`
|
||||
(real login→CSRF) or claude-in-chrome for the browser planes.
|
||||
|
||||
## Findings already banked in 6A (do NOT re-hunt; watch for related regressions)
|
||||
- **C6-1 (LOW-MED):** `skip-active` is a no-op on `pct reboot` — the guest-hook heal re-arm carries
|
||||
every active-share reboot; the "fresh namespaces inherit real mounts" log is misleading.
|
||||
- **C6-2 (MED):** a NAS outage spanning a guest reboot can strand the share `failed` until an agent
|
||||
restart (or periodic sweep — timing unmeasured; **6B: measure the periodic-sweep interval** if a
|
||||
cheap opportunity arises).
|
||||
- **C6-3 (LOW, setup):** the agent/wizard don't pre-create an app's userdata tree on a fresh
|
||||
`all_squash` NFS export → docker chown fails → container stuck `Created` until pre-created.
|
||||
- **F-A / F-B / F-C are already confirmed-fixed-live** (CAMPAIGN-5, same 0.129.0 fleet). 6B may
|
||||
spot-re-confirm opportunistically but need not re-prove them.
|
||||
|
||||
## P0 (fast) — verify the pre-existing state, record a seed, then GO
|
||||
Versions both; caps 63/63; campaign6 enrolled+health; sonarr present; escrow states; samplers alive.
|
||||
Push to `180:~/campaign6/evidence/P0-6B/`.
|
||||
|
||||
## PHASE 2 — the `.fab` full circle, UPLOAD half (demo; twice-deferred flagship)
|
||||
Fresh **campaign** app + **~4 GiB** `/dev/urandom` (varied sizes incl one >1 GiB); sha256 manifest.
|
||||
Export → `.fab` (progress honesty + atomicity). Download twice — browser/LAN AND `curl` with the
|
||||
session cookie **through the real Cloudflare edge** (`--resolve <host>:443:<CF-public-IP>` — Pi-hole
|
||||
split-horizon bypasses the edge otherwise); both hashes == the docker-cp reference. Re-prove the edge
|
||||
cap: single >100 MB POST → **413 from Cloudflare**. Delete the campaign app + data. Upload back —
|
||||
browser/LAN AND `curl --resolve` through the edge in 64 MiB chunks. Import/restore; **byte-compare
|
||||
every file vs the manifest → zero mismatches = PASS.** Edge volley (live over the unit tests): cancel
|
||||
mid-upload → abort + `.part` gone; controller restart mid-upload → startup GC count logged; offset
|
||||
replay → 409+resync; oversize-vs-free → Hungarian refusal (both numbers) **AND F-A real multi-GB
|
||||
size**; collision → " (1)" then " (2)"; second concurrent init → 409; wrong extension refused; idle
|
||||
15-min abort (park one, return). (Note the storage-bearing box + a proper export are required — see
|
||||
C6-3; pre-create the userdata tree.)
|
||||
|
||||
## PHASE 3 — browser planes (claude-in-chrome; twice-deferred)
|
||||
Session must be started AFTER the chrome bridge connected. Campaign credential for login.
|
||||
- **3A escrow wizard, full browser pass (drill):** preflight all-green + Hungarian details (no raw
|
||||
English leak), warnings, re-auth, run, reveal, **typed-back with the two highlighted words** (read
|
||||
from screen), manual hide/show toggle, finish → auto-confirm flips + hub row hash matches
|
||||
(server-side check via 180). Then re-claim → 410 UI; unclaimed → TTL → `unclaimed_void` screen;
|
||||
**F-C live** (`phase:none` claim → clean 4xx UI, not 502); out-of-band CLI ceremony w/o staged
|
||||
secret → stale card fires → wizard clears it.
|
||||
- **3B DOM / native-alert / session sweep (both boxes):** grep served DOM for `alert(`/`confirm(`
|
||||
(must be zero); session expiry mid-wizard and mid-upload → JSON 401 on `/api/`, redirect on pages;
|
||||
CSRF stale-token on storage-wizard/escrow/upload POSTs → 403 JSON; zero-toggle honesty + inline
|
||||
two-step confirms; **F-B live** quick re-check.
|
||||
- **3C hub 8-tab customer-detail ring (hub credential = campaign credential):** hash-nav across all
|
||||
8 tabs; auto-refresh scoped to live tabs; **dirty-form suppression** (start editing → refresh
|
||||
holds); events tab under volume; no stale-host deletion against real hosts (synthetic row only).
|
||||
|
||||
## PHASE 4-rest — backup tiers depth + NAS integrity
|
||||
- **F7 mid-backup NAS cut:** `POST /api/backup/run` on a NAS app; `exportfs -u` at T+~6s; the
|
||||
last-good volume dump must survive (tmp+rename) — no 0-byte artifact; `success:false` is the only
|
||||
signal. (Give the app enough data that the dump takes >6 s.)
|
||||
- Four backup sub-pages truth vs live; offsite card three honest states; Tier-3 "Távoli mentés most"
|
||||
additive + quota bar; **restic stale-lock self-heal** (kill controller mid-offsite-run → next run
|
||||
self-heals the lock); tier-1 replace-semantics; offsite restore-to-verify to a scratch target
|
||||
(campaign app) — byte-identical; volume-only app gets its tier-2 secondary (F6 fix); per-app
|
||||
toggles round-trip; snapshot coherence across pages. (F8 already CONFIRMED in 6A — cite it.)
|
||||
|
||||
## PHASE 5 — regression spot-checks (fast)
|
||||
- Agent-restart re-arm still logs per-share verdicts (6A re-showed it via 1E `netmigrate` — one
|
||||
quick re-confirm). **F1/F2 residue** after the campaign share's final removal — zero mounts/units/
|
||||
dirs, host+guest. **F4** `mapped_uid:101000` → friendly Hungarian 400 (not `agent_error`).
|
||||
Dead-app alert + email on a killed campaign container (cooldown math vs the run's emails).
|
||||
|
||||
## Wrap — completeness gate (same as 6A)
|
||||
Checklist table of EVERY 6B item = PASS / FAIL / FINDING / (operator-authorized) split-to-6C. Verdict;
|
||||
ranked findings + repros; timings; deviations; box state (both boxes; "campaign credential active —
|
||||
Viktor rotates"; drill fresh-R note); morning recovery; evidence index. Commit the doc + overwrite
|
||||
`felhom.eu/REPORT.md`. **No credential, no R, no blob committed.**
|
||||
|
||||
## Final cleanup (unattended — do it, no BLOCK; show before/after in the ledger)
|
||||
Remove campaign6 share + campaign apps (product flow); `sudo exportfs -u
|
||||
192.168.0.162:/mnt/5_hdd/felhom-campaign6 && sudo rm -rf /mnt/5_hdd/felhom-campaign6` on 180 (confirm
|
||||
`exportfs -v` back to the P0 baseline, felhom-data intact); remove scratch `.fab` + injected data;
|
||||
stop the samplers (`pkill -f c4-sampler.sh` both PVE hosts, `pkill -f hub-sampler.sh` on 180). Leave
|
||||
the campaign credential in place and note in the doc that **Viktor rotates it** — that is the one
|
||||
human step that remains after 6B.
|
||||
Reference in New Issue
Block a user