Files
felhom.eu/documentation/audits/CAMPAIGN-6E-2026-07-15.md
T
admin c2b7145cbd docs(audit): CAMPAIGN-6E — P-TIER2 deep-4 + P4-DEEP timing (all 8 legs PASS)
Closes the two deferred deep campaigns from 6D on live infra: tier-2 v2 layout
engine (A1-A4) and F7 atomic-write + restic C2 self-heal (B1-B4). No production
code shipped. Two LOW findings (F-6E-1 demo drive drift; T-6E-1 DB-dump dir-fsync
parity). Drill repo torn down; fixtures reverted; box healthy.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HEPuEwyyGDJdcsXLFsTWJn
2026-07-15 21:11:39 +02:00

159 lines
12 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# CAMPAIGN-6E — P-TIER2 deep-4 + P4-DEEP timing (pre-broad-launch resilience) (2026-07-15)
**Operator:** Viktor (present). **Driver:** Claude Code (live box). **Validator:** project Claude.
Addendum to [CAMPAIGN-6D](CAMPAIGN-6D-2026-07-15.md) — closes the two deferred deep campaigns before
broad (Facebook) launch. NOT a friend-alpha gate (that is already GREEN in 6D). Operational,
supervised; **no new production code shipped**. A surfaced issue is RECORDED + ranked, never fixed
inline. Every mechanism under test is already SHIPPED + unit-tested; these legs confirm the shipped
behavior survives a real adverse event on live infra.
Baselines (E0): controller **0.137.0** (Up healthy), agent **0.88.0**, hub **0.56.0**.
Evidence sink: `180:~/campaign6/6E/evidence/`. Subject box: demo guest 9201 (`demo-felhom`).
(Sibling doc, not an append to 6D — 6E is a distinct addendum campaign.)
---
## Headline verdict
**Both deep campaigns GREEN.** The tier-2 v2 layout engine (Task 3b, v0.135.0) and the F7 atomic-write
+ restic single-writer self-heal (CAMPAIGN-3 / C2) all survive real adverse events on live infra.
- **Phase A (P-TIER2 deep-4): 4/4 PASS.** class-driven legs, the lifted N>1 refusal, deferred-prune
reconcile, and old-flat→v2 migration all behave exactly as designed; no removal ever left
`backups/secondary/`; markers written last.
- **Phase B (P4-DEEP timing): 4/4 PASS.** A mid-write SIGKILL never promotes a partial dump over a good
restore point (volume tar AND DB dump); the restic C2 crash-lock self-heal fires correctly against a
**real SFTP repo** with a genuine foreign-hostname exclusive lock; an offbox restore round-trips
byte-identically.
- **No HIGH/MED findings.** Two LOW observations (below), one with a candidate follow-up TASK.
---
## Pre-flight (E0E5)
| Gate | Result |
|------|--------|
| E0 baselines | ctrl 0.137.0 / agent 0.88.0 / hub 0.56.0 — all match. |
| E1 guest health | 9201 controller 0.137.0 Up (healthy). |
| E2 credential | `C4_PASSWORD` present in `180:~/.config/credentials` (single-quoted → stripped on use). |
| E3 drill-repo isolation | drill restic repo `sftp:…:/home/6e-timing-1784142113` — a dedicated path, never the real `/home/felhom-repo`. Torn down (proof below). |
| E4 evidence sink + headroom | `180:~/campaign6/6E/evidence/` created; 153 G free on 180. |
| E5 ≥2 local drives | three non-network drives registered: **felhom-usb** (default), **teszt_enroll**, **felhom-flash** (+ nas-media, network — correctly never a tier-2 target). |
---
## Phase A — P-TIER2 deep-4 (writes ONLY under `backups/secondary/`)
Driver: the real per-app endpoint `POST /stacks/{name}/backup` (`enabled=on`, `target=<pin>`, CSRF) →
`RunTier2(name)`, target pinned for determinism. Subject: **immich** (the one HDD app on the demo with a
drive registration consistent with where its data physically lives — see Finding F-6E-1). Source
felhom-usb → target teszt_enroll. Classification edits were made to the **deployed**
`/opt/docker/stacks/immich/.felhom.yml` / `docker-compose.yml` (the exact files `ClassifiedBinds` reads)
and reverted byte-identically; all RECORDED via `/root/*.orig` backups.
| Leg | Result | Evidence |
|-----|--------|----------|
| **A1** no-block legacy app | **PASS** | immich un-classified → legacy resolver path → single `hdd/appdata/immich` leg, **byte-identical to source (286/286 files)**, `.felhom-tier2-layout=2` present, log `1 leg(s)`, no `errTier2MultiDir`. The now-unclassified `userdata/media` leg was correctly pruned by reconcile. Confirms the legacy→v2 byte-identity promise (SQ5 footprint). |
| **A2** N>1 lifted refusal | **PASS** | Added a 2nd appdata bind (`appdata/immich-6e2`, classified mandatory) → run yields **two `hdd/appdata/*` legs** (`immich` + `immich-6e2`) coexisting, log `3 leg(s)` (both appdata + `userdata/media/photos`), **no `errTier2MultiDir`** — the v0.131.0 flat-appdata refusal is structurally gone. 2nd-appdata leg byte-identical. |
| **A3** re-class reconcile | **PASS** | Re-classed `immich-6e2``excluded` → run logs `[INFO] Tier 2 reconcile: removed stale dest dir hdd/appdata/immich-6e2`; `hdd/appdata/immich` + `userdata/media/photos` kept; marker rewritten. Removal strictly within `backups/secondary/` (`tier2SafeRemove` guard). |
| **A4** old-flat → v2 migration | **PASS** | Synthesized a genuine pre-v2 flat copy at the target (`appdata/` flat + `recovery-unit/`, **no marker**) → run: flat `appdata/` **removed**, v2 legs (`hdd/appdata/immich` + `userdata/media/photos`) rebuilt, **`recovery-unit/` content-identical (untouched by migration)**, `.felhom-tier2-layout=2` written **last**. |
**Phase A boundary check:** every reconcile/migration removal stayed within `backups/secondary/<stack>/`
(log paths + `tier2SafeRemove`); the surrounding tree (recovery-unit, sibling legs) intact each time.
Markers written last in every case.
Log evidence (`6E-phaseA-logs.txt`):
```
Tier 2 copied immich → …/teszt_enroll/backups/secondary/immich (466.4 MB, 3 leg(s), 0s) [A2]
Tier 2 reconcile: removed stale dest dir hdd/appdata/immich-6e2 [A3]
Tier 2 copied immich → … (466.4 MB, 2 leg(s), 0s) [A3]
Tier 2 reconcile: removed stale dest dir userdata/media [A1 legacy prune]
Tier 2 copied immich → … (466.4 MB, 1 leg(s), 0s) [A1]
Tier 2 copied immich → … (466.4 MB, 2 leg(s), 1s) [A4 migration]
```
---
## Phase B — P4-DEEP timing (process-kill only; no power-cut)
### B1 — F7 atomic tar-write (volume path) — **PASS**
Controlled real docker volume (2 GB) + the identical controller tar seam
(`docker run … alpine tar cf …/<vol>.tar.tmp`). tar **SIGKILLed mid-write** (exit 137) → left a
**partial** `.tar.tmp` (1.21 GB < 2.0 GB full); the good `.tar` was **byte-unchanged** (sha256 match).
The orphan `.tar.tmp` is invisible to restore-point scans (matches `*.tar` only). Clean re-run + the
`atomicPromoteTar` sequence (fsync + atomic rename + dir-fsync) → a new **valid, extractable** `.tar`.
→ A partial `.tar.tmp` never replaces a good restore point.
### B2 — F7 atomic write (DB-dump path) + the asymmetry check — **PASS**
Real `pg_dump` of the immich DB (read-only) into `<name>.sql.tmp`, **SIGKILLed mid-write** (exit 137) →
partial `.sql.tmp` (41 MB < 45.8 MB); the good `.sql` was **byte-unchanged**. `DumpOne`
(`appbackup/dbdump.go`) removes the tmp on error and only renames on success.
**Asymmetry observation (confirmed from code):** `DumpOne` fsyncs the **file** (`tmpFile.Sync()`,
dbdump.go:277) but **not the directory entry** after `os.Rename` — whereas `atomicPromoteTar`
(backup.go:599) additionally `dir.Sync()`s. Under a **process-kill** (this test) both are safe: the
rename is visible in the running kernel regardless of dir-fsync. The gap matters only for **power-loss**
durability parity. → candidate follow-up TASK **T-6E-1 (LOW)**, below.
### B3 — restic single-writer self-heal (drill repo only) — **PASS**
On the isolated drill repo (E3), reproduced the exact C2 condition: a **foreign-hostname exclusive
lock** created by a helper container (`--hostname felhom-old-ctnr`) whose `restic prune` was SIGKILLed
mid-run. From the controller (host `demo-felhom`):
1. plain `restic unlock` (stale-only) reports "successfully removed locks" but the foreign lock
**remains** — restic cannot prove `felhom-old-ctnr`'s PID dead, so it is not treated as stale. This is
exactly why the pre-run `unlockStale` is insufficient for a recreated-container crash lock.
2. `restic backup` → `unable to create lock in backend: repository is already locked exclusively by PID
1 on felhom-old-ctnr by root` — the error `offboxLockRe` matches.
3. `restic unlock --remove-all` → locks empty `[]`.
4. retry `restic backup` **once** → `snapshot 33568faf saved`.
This is precisely `resticStep`'s self-heal path (offbox.go:349): op → `repository is already locked` →
`unlock --remove-all` → retry once. The "never loops" bound (a 2nd lock failure surfaces the original
error, offbox.go:360/363) is a code guarantee, not force-tested live. Evidence: `6E-B3-exclusive.txt`.
Real repo `/home/felhom-repo` never touched (all writes to `/home/6e-timing-…`).
### B4 — restore-to-verify-compare — **PASS**
Read-only `restic restore` of immich's mandatory `appdata/immich` subtree (snapshot `d606269e`) from the
**real** repo to a **local** scratch (`/tmp/6e-b4` on the container rootfs — not network, per Part 0.5).
sha256 file-by-file compare vs the live source: **286/286 files byte-identical**. Restore is read-only,
so this ran safely against the real repo. → "a tier is only real once a restore round-trip is proven",
as a pure integrity compare.
---
## Findings (ranked)
| ID | Sev | Finding | Disposition |
|----|-----|---------|-------------|
| **F-6E-1** | LOW | **Demo-data drive drift:** `nextcloud` (and likely `romm`) have no `app.yaml`/`HDD_PATH`, so `GetAppDrivePath` defaults to `/mnt/sys_drive` while their real appdata lives on `felhom-usb`. Tier-2 therefore WARNs `mandatory data path missing on disk` and produces no copy. **The code is correct** (it warns loudly and does not silently succeed). This is demo-data churn from app-drive migration, not a product defect — a real deployment sets `HDD_PATH`. | No code TASK. Demo-hygiene cleanup only (re-point or redeploy nextcloud/romm on the demo if a working tier-2 copy is wanted there). Optional product idea: surface "tier-2 has never produced a copy (source missing)" more prominently in the operator UI — not filed. |
| **T-6E-1** | LOW | **DB-dump dir-fsync asymmetry** (B2): `DumpOne` fsyncs the file but not the directory after `os.Rename`; `atomicPromoteTar` fsyncs both. Safe under process-kill; only power-loss durability parity would benefit. | Candidate follow-up TASK: add a best-effort `dir.Sync()` after the rename in `DumpOne` (mirror `atomicPromoteTar`). LOW — not required for the shipped safety property; do NOT ship inside this campaign. |
No HIGH or MED findings.
---
## Teardown & box health
- **Drill repo torn down (proof):** `ssh …storagebox rm -rf /home/6e-timing-1784142113` then `ls /home`
→ `felhom-repo` only. Real repo intact.
- **All temporary fixtures reverted byte-identically:** immich `docker-compose.yml` + `.felhom.yml` ==
`/root/*.orig` (verified `diff -q`); scratch source `appdata/immich-6e2` removed; nextcloud
classification restored. immich DB never modified (a scratch-table attempt failed on a role error and
created nothing).
- **Tier-2 preferences:** immich + nextcloud were left at default participation (enabled, no pin,
`user_disabled=None`). immich now carries a healthy auto v2 copy on teszt_enroll; nextcloud reverts to
its pre-existing F-6E-1 state. teszt_enroll/secondary otherwise returned to pre-state.
- **Scratch/secrets removed:** B1/B2 scratch volume + dumps, B3/B4 scratch, and the session
cookie-jar/CSRF/password files all deleted; no `felhom-6e*` volumes or containers remain.
- **Box healthy:** controller 0.137.0 Up (healthy); immich-server Up (healthy).
## Evidence index — `180:~/campaign6/6E/evidence/`
`6E-prestate-secondary.txt`, `6E-A2-evidence.txt`, `6E-phaseA-logs.txt`, `6E-phaseB-B1B2.txt`,
`6E-B3-out.txt`, `6E-B3-exclusive.txt`, `6E-B4-result.txt`.
## Queue delta
- **Closed by 6E:** P-TIER2 deep-4 (A1A4) and P4-DEEP timing (B1B4) — the two deferred deep campaigns
from 6D. The v2 layout engine and the F7 / C2 self-heal mechanisms are live-confirmed.
- **Remains (broad-launch, non-blocking):** T-6E-1 (LOW, optional power-loss durability parity for the
DB-dump path); demo-hygiene for F-6E-1 (nextcloud/romm drive drift). Neither blocks broad launch.