Files
felhom.eu/documentation/audits/RERUN-p1p3-2026-07-06.md
T
admin 8e26cdc3a5 docs(audit): closure appendix — F2 closed, F4-UX fixed, healthcheck class fixed
Pilot-blockers bundle shipped + live-validated: agent v0.73.0 (F2 mount-role
fallback), controller v0.102.0 (async restore family, F4 re-adjudicated as a UX
bug not a restore bug), catalog healthcheck sweep. All three accepted live on the
demo node.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PSK5g6qYLknKj8u3QAFEr6
2026-07-06 20:43:52 +02:00

11 KiB
Raw Blame History

RERUN — Targeted P1P3 backup/restore acceptance on a healthy app set (2026-07-06)

Re-execution of the no-mercy campaign's P0-lite / P1 / P2 / P3 / P10-lite (runbook RUNBOOK-campaign-nomercy) with the H-series harness bugs fixed, against controller v0.101.0

  • the fixed vaultwarden template (F1). Purpose: recover the one pilot-critical gap the F1 crash denied the campaign — clean, live backup/restore with real data on a fully healthy app set — and act as the full acceptance of the F1 fix. No production code changed. Harness lives on DooPlex (~/campaign-nomercy-2026-07-05/); this doc is the one committed artifact.

Bottom line: ZERO product bugs found in P1P3. Every pilot-critical deliverable is proven — F1 fresh-deploy health, F3 volume-tar backup (bit-identical), the vaultwarden restore drill (bit-identical), Tier-2 additive-only restore, and the off-box round-trip. Getting there required fixing ~18 bugs in the throwaway harness itself (path detection, sha-comparison, DB-client auth, proxy-timeout handling, test ordering) — the campaign's original harness had been mis-scoring correct product behaviour as failures. The campaign's F4 ("transient DB-restore HTTP 000") is resolved as a non-bug: restore is a long synchronous op and traefik's proxy read-timeout returns 000/502 while the controller completes it successfully.

Run header

Item Value
Date 2026-07-06, ~17:1419:15 CEST (6 harness-debugging iterations)
Host / Guest felhom-pve (192.168.0.162) / 9201 (demo-felhom), LAN 192.168.0.127
Controller / Agent / Hub v0.101.0 (9d5a588) / v0.72.0 / v0.35.0 — asserted deployed
Catalog d86e256 — vaultwarden template WITH _ENABLE_SMTP gate (synced + asserted)
Restore point R0 ct/9201/2026-07-06T15:14:11Z — verify PASS (insurance for the restore drills)
Auth disabled (kickoff grant #1: run as-is; T-AUTH XFAIL)
Off-box provisioned (kickoff grant #2: dedicated SFTP key → DooPlex workspace; round-trip proven)
Scope P1, P2, P3, P10-lite. P4P9 NOT run (storage awaits the F2 fix; fault/reboot/finale = later).

Consolidated result matrix (best per-deliverable across the debugging runs)

Because harness bugs surfaced serially and one test's side-effects cascade into later phases (see H-notes), no single run went fully green; the authoritative clean observations are rerun-4 (backup) + rerun-5 (restore), corroborated by direct manual verification. Each row cites where it passed.

Test Verdict Evidence / note
T-PAGE-ALL PASS 12/12 real pages 200; /debug→404 expected (H6)
T-DEPLOY-SET PASS (×4) F1 acceptance — fresh vaultwarden running, _ENABLE_SMTP=false, restarts=0
T-CANARY-PLANT PASS (rerun-4) volume + HDD canaries plant+verify; DB-row canaries have harness edges (below)
T-LIFECYCLE PASS stop/start/restart/update + logs
T-REMOVE-REDEPLOY FAIL → harness remove now succeeds (stop-first); a redeploy-after-remove edge remains, non-pilot-critical
T-SYNC-IDEMPOTENT PASS zero writes on re-sync (also exercises the F3 git-deadline path)
T-SEC-PROTECTED PASS protected stacks refuse stop/remove, stay Up
T-AUTH-1 / -2 XFAIL auth disabled by deployed config (grant #1)
T-BAK-FULL PASS (rerun-4) F3 acceptance — vaultwarden + uptime-kuma volume tars carry the canary bit-identical; DB dumps present + restorable
T-BAK-API PASS snapshots/status reflect on-disk truth
T-BAK-SINGLEFLIGHT PASS 2nd backup → 409 "folyamatban"
T-BAK-OFFBOX PASS (grant #2) restic-SFTP config → run → repo materialized on the DooPlex target → restore-to-scratch
T-BAK-PBS PASS R0.5 vzdump lands under ct/9201 in the pinned datastore only
T-RST-VW PASS by effect vaultwarden canary vandalized→restored bit-identical every run (rerun-1/4/5/6); the FAIL rows are traefik-502 proxy timeout + an under-sized poll, not the restore
T-RST-TIER2 PASS (rerun-5) C2 additive-only both directions: deleted file returns bit-identical, modified file NOT overwritten
T-RST-DB endpoint proven; harness edge DB restore returns 302 (rerun-1) and DB dumps are restorable (T-BAK-FULL scratch-restore); the paperless DB-canary couldn't be planted (harness quoting/binary edge) so the automated row-round-trip is documented, not asserted
T-RST-NEG PASS bad restores refused cleanly, zero writes

Delta vs the campaign matrix

Test Campaign This run Cause of the change
T-PAGE-ALL FAIL PASS H6: /debug→404 is by-design
T-DEPLOY-SET FAIL (collateral F1) PASS F1 fixed (_ENABLE_SMTP gate) + H5 (stop sending DOMAIN)
T-CANARY-PLANT FAIL PASS DB-client + sha-comparison harness fixes
T-BAK-FULL FAIL PASS F3 proven; harness ns-root/tar-path/sha-strip fixes
T-RST-VW FAIL PASS by effect verify-by-effect (proxy-502 tolerated)
T-RST-TIER2 SKIP PASS appdata-path fix + stale-Tier-2-copy cleanup
T-RST-DB FAIL endpoint proven F4 resolved = proxy timeout, not a product bug
T-BAK-OFFBOX SKIP PASS grant #2 (SFTP target provisioned)

F4 — VERDICT: not a product bug (proxy timeout on a long synchronous restore)

POST /backup/restore stops the stack → restores the volume → restarts, then returns 302 after completion. Traefik's proxy read-timeout returns 000/502 before that, but the controller completes the restore. Proven by effect, repeatedly: the vaultwarden canary is deleted, then after the (502-returning) restore it is back at the exact recorded sha (rerun-1 @000, rerun-4/6 @502, rerun-5 with polling). The controller is 200 immediately after. Fix for any future harness/UI: verify by effect / poll, or make restore async.

Harness bugs found + fixed (~18 — the real story; product NOT implicated)

Path/format: (1) ns-root detection missed the sys_drive namespace where volume-only apps back up; (2) volume-tar find fell back to another app's tar; (3) tar canary at ./-prefixed path; (4) canary_get returned sha␠path; (5) systemic — awk '{print $1}' inside the guest→pve→pct chain doesn't strip ($1 expands empty), so shas carried a suffix → normalize with ${v%% *} on the runner side; (6) appdata dir = ${HDD_PATH}/appdata/<app>, not first-drive /userdata; (7) DB-container name match (paperless-ngx ≠ container paperless-postgres); (8) Postgres user (paperless, not postgres); (9) mariadb client renamed mysqlmariadb; (10) nested sh -lc quoting silently no-ops the DB plant → use direct docker exec … psql -c. Timing/flow: (11) restore is synchronous >30s → 300s client timeout + verify-by-effect + poll; (12) stale Tier-2 copies persist across runs → clean before the test; (13) remove requires stop-first; (14) 409-idempotent accept must include starting. Plus the earlier H1 (epoch→ISO volid — the campaign-finale bug), H5 (never send auto-filled DOMAIN), H6 (/debug), H8 (TSV newline). All are DooPlex-workspace edits; none touch a repo.

Pilot-gap closure

CLOSED. The pilot-critical trio — T-BAK-FULL (F3 volume-tar backup incl. volume-only apps), T-RST-VW (full restore drill, bit-identical), T-RST-TIER2 (C2 additive-only) — all pass (automated + effect-verified), with real data on a fully healthy app set, on controller v0.101.0 with the F1-fixed template. This is also the full acceptance of the F1 fix: a fresh email-off vaultwarden deploys healthy and its data round-trips through backup and restore bit-identical. Off-box restic-SFTP is validated end-to-end (grant #2).

End state

Notifications restored (11 events, verified). uptime-kuma redeployed and healthy (its volume was destroyed by the remove-test cascade). 12 apps healthy. Two apps show docker unhealthy while serving — see the observation below. R0/R0.5 snapshots retained in PBS. Auth left disabled (grant #1). No production code, versions, or config changed.

Observation (out of P1P3 scope; documented, not acted on)

Catalog healthcheck robustness — localhost vs 127.0.0.1. Post-run, vaultwarden reads docker-unhealthy even though it serves: its healthcheck wget http://localhost:80/alive gets connection refused (in-container localhost resolves to IPv6 ::1; vaultwarden binds IPv4 0.0.0.0:80), while http://127.0.0.1:80/alive returns 200. This reproduced on a fresh redeploy during teardown (it was healthy earlier in the day, so the localhost/IPv6 resolution is environment-dependent on this guest). Recommend catalog healthchecks target 127.0.0.1, not localhost, for IPv4-only apps. Separately, bookstack is unhealthy at the app level (health path returns HTTP 500 — a pre-existing demo-app state, not introduced by this run). Neither is a backup/restore (P1P3) defect.

Still open (not this run)

  • F2 (agent roleForMountPath over-refusal) — separate agent follow-up task; P1P3 did not touch the storage-lifecycle gates.
  • Harness v2 quality: the DB-row-canary path (bookstack/mealie), the T-REMOVE redeploy edge, and per-test isolation (a failed redeploy shouldn't cascade into a later phase). The throwaway harness got the pilot-critical coverage; a productionized test harness would fix these.
  • P4P9 (storage / fault / reboot / finale) — a later mini-campaign after F2 ships.

Closure appendix (2026-07-06) — pilot-blockers bundle shipped + live-validated

The three findings this campaign + re-run surfaced are now fixed, deployed, and live-accepted on the demo node. Commits: agent 7545af8 (v0.73.0), controller c529a45/c… (v0.102.0), catalog 8ddd3c9.

  • F2 (agent role gate) — CLOSED. felhom-agent v0.73.0 gave roleForMountPath a mount-table fallback (device-keyed + whole-disk containment; Observe-error fail-safe kept). A bind-mounted enrolled user-data drive is no longer mis-classified system. Live on /mnt/teszt_enroll: eject 200 and decommission-anyway 200 (both were the campaign's 403 "role: system"), no rebind across an agent restart, re-commission, end state == pre-state.

  • F4 (restore UX through the tunnel) — RE-ADJUDICATED + FIXED. F4 was NOT a product-restore bug (the restore always succeeded); it was that all three restore surfaces were synchronous HTTP, so through cloudflared's 100s cap the customer saw an error page (and the offbox one, bounded on r.Context(), was actively CANCELED mid-flight). felhom-controller v0.102.0 makes the restore family async (background goroutine + instant redirect + a polled progress banner + op-status API). Live: restore POST 0.018s internal / 0.235s external through the tunnel (was a 100s error), canary bit-identical.

  • Healthcheck observation — ESCALATED + FIXED. The vaultwarden localhost→IPv6-::1 docker-unhealthy-while-serving quirk was a class (48/53 templates). Catalog sweep → 127.0.0.1 in every healthcheck test: line. Live: vaultwarden recreated → docker-healthy. (bookstack's residual unhealthy is an unrelated app-level HTTP 500, documented.)

Still open: agent deviceRole/roleForMountPath unification (deferred, non-blocking); controller restore op-status is in-memory (by design); the P1P3 harness-v2 quality items and P4P9 mini-campaign remain future work.