Live: backup of 9001 (crash-consistent, 934MiB), restore-test (restore→boot→verify→ teardown, gate-benign, no leak), AND a real crash-recovery proof — a DestroyLXC DELETE bug leaked scratch 990000, the fix landed, and Recover reaped the leaked guest on the next run before a clean test. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
5.0 KiB
REPORT — Slice 6 Phase A: backup + the self-restore-test (v0.6.0-rc1) (2026-06-09)
Overwrite-latest report (most recent significant work only). Cumulative history lives in CHANGELOG.md.
Outcome
Slice 6 Phase A is implemented, live-validated on demo-felhom, and pushed as
v0.6.0-rc1 — the guest-level backup layer + the journaled self-restore-test that
closes "a backup you haven't restored isn't a backup" (doc 03 §8). Everything is benign
(backup, restore-to-NEW, scratch teardown): it reuses the slice-4 classifier/gate/journal —
no new destructive class, no new crypto. Local target only; PBS is Phase B. Restore is
to a NEW guest only (no overwrite). Backups are crash-consistent (marked so). This is
the hard checkpoint: awaiting validation before Phase B.
What landed
- proxmox:
DestroyLXC(the scratch-teardown primitive),VzdumpOptions.Notes→notes-template,LatestBackupVolID(resolve a produced archive — the task status carries no result volid). - The self-restore-test (
reconcile.Engine.RunRestoreTest): pick a free scratch VMID (configured band, excludes 9999; full band → skip) → journal aScratch-owned entry BEFORE any mutation → restore-to-new → benign net link-down SetConfig (so the clone can't conflict with a running source's MAC/IP — test-safety, NOT slice-7 identity reset) → boot → verify reachesrunning→ defer teardown that ALWAYS runs (benignClassGuestDestroy+ agent-tagged-scratch provenance, through the gate). - Crash-safe recovery (
Recover): aScratchentry is resolved by teardown, not by re-checking the restore sub-task's UPID — special-cased before the generic path (else a leaked guest would be marked "succeeded"). Idempotent: already-gone → clean; list-unreadable → left in-flight. internal/backup:BackupRunner(vzdump + archive/size resolve + the bulk-volume gap — a mountpoint is uncovered unless it carries an explicitbackup=1, so an unsetbackup=counts as uncovered, the safe DR direction),PickRestoreCandidate, an in-memoryStore(latest backup-per-target + latest restore-test) feeding the hub reporter seams, and a cadenceScheduler(default 24h; the fourth daemon goroutine; disabled cleanly when off/misconfigured).- hub report:
Backup+RestoreTestfilled (PBSSnapshotstays a Phase-B stub); collector seams; cross-repo golden byte-identical + bidirectional key-set tests in both repos; hubhandler.goparses + persists them and logs a FAILED restore-test prominently (the loudest DR signal). - config
BackupConfig(band 990000–990009 default);--selftest=backupand--selftest=restore-test(Recover-then restore→boot→verify→teardown).
Tests
go test ./... green; go test -race ./... green on the build server (the scheduler
goroutine). Notable unit tests: vzdump async UPID→WaitTask + bulk-gap (backup!=1 incl.
unset); restore-test pass + teardown-on-failed-verify + restore-failure-still-tears-down +
full-band-skip; the DELETE wire contract (no body + force=1); and the headline recover-
leaked-scratch trio (destroyed / already-gone-clean / list-unreadable-left-in-flight).
Live validation on demo-felhom
--selftest=backup -vmid 9001→ vzdump of the stopped spike-lxc tolocal: success, crash-consistent, archivelocal:backup/vzdump-lxc-9001-…tar.zst(~934 MiB) resolved, 26 s. Bulk gap empty (9001 has no extra mountpoints).--selftest=restore-test→ restored that backup into scratch 990000 on local-lvm, net link-down, booted, verifiedrunning, torn down (gate-benign) — pass, 15 s, no leak.- Crash-recovery proven for real: the first teardown hit a
DestroyLXCbug (DELETE with a form body → PVE HTTP 501; a booted guest also needsforce=1), which leaked scratch 990000. The fix moved the flags to the query string +force=1, no body (regression- tested). On the next run,Recoverfound the journaled in-flightScratchentry, confirmed 990000 existed, and destroyed the leaked guest (ScratchDestroyed:1) through the benign gate — then ran a clean test.pct listafter: no 99000x guests. The accidental leak + reap is the best possible live proof of the slice's load-bearing crash-safety path.
Notes / cleanup
- The demo config (
/root/.config/felhom-agent/agent.json) gained abackupblock with the cadence disabled (restore_test_cadence_seconds: -1); the validation used the selftest harness. The example config ships cadence default-ON (24h) with the 990000-band + local-lvm. - One test backup of 9001 remains on
local(a valid artifact; ~934 MiB of ~88 GiB free). - The binary on felhom-pve is v0.6.0-rc1; token/TLS pin per memory
demo-felhom-live-agent.
Deferred → Phase B
PBS: the datastore on the USB, zero-knowledge key custody, restore-from-PBS, and the PBS integrity-verify as the lighter frequent check. Then slice 7 (provisioning + identity-reset + golden base, §9).