R-672/R-673: 03 + 08 + register + CONTEXT (restore test off, 9201 repaired, agent v0.133.0, hub v0.124.0); R-684
gates / gates (push) Successful in 24s

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
2026-09-24 16:29:54 +02:00
parent 068e06537a
commit 54bff69f88
5 changed files with 51 additions and 2 deletions
@@ -308,6 +308,31 @@ per Part 1: **snapshot** (LVM-thin, transient, whole-guest rollback — not a ba
> Unchanged and load-bearing: `onboot=0` on the scratch at restore time, **every NIC link-down before
> boot**, journal-before-mutate, guaranteed teardown, and the per-tier restore-task timeout.
> **A restore-test can never fill a box's disk (R-672, agent v0.133.0 + hub v0.124.0).** Measured
> 2026-09-24 on demo-hp: the scheduled restore-test restored 9201's archive into `local-lvm` — the pool
> holding 9201 — with no space check; the pool reached 100 % and 9201's disks remounted read-only.
> Since v0.133.0, before anything is journaled or created:
> - **Space first.** Free data ≥ restored × 1.2 + 5 GiB (`backup.restore_test_space_factor`,
> `…_reserve_gib`) and room in a thin pool's metadata. `restored` is the **uncompressed** size — the
> vzdump log's "Total bytes written" or the PBS snapshot size. **Never the archive file:** 9201's file
> was 6.9 GB and its restore wrote 22.6 GB, so "file × 1.2 + 5 GiB" would have let that test run.
> - **Off the tested guest's pool** when another storage is eligible (active, `rootdir`, and the agent
> holds `Datastore.AllocateSpace` there) and fits. demo-hp has none: `nvme-scratch` carries no grant.
> - **Unknown refuses.** A refusal is the test's result — `pass=false`, `skipped`, "skipped: not enough
> space on …" — so the hub raises `restore_test_failed`; it is never a pass and never dropped.
> - **Leftovers on a timer.** A failed scratch teardown and the stale-lock sweep (R-673) run every 10
> minutes, not only at agent start; the sweep holds the one-heavy-operation gate; after 3 failed
> teardown tries the operator is told.
> - **A thin pool ≥ 90 % requests an immediate report**; the hub judges a thin pool on the worse of data
> and metadata, critical at 90 %, one alarm per pool per 6 hours (`08` §6.2).
>
> **Operator ruling 2026-09-24 (evening):** the scheduled restore-test is **OFF on both demo hosts**
> (`backup.restore_test_eval_interval_seconds: -1` — **0 does not disable, it means the 6-hour default**)
> until v0.133.0 is delivered there; then it is switched back on. Saved configs:
> `/etc/felhom-agent/agent.json.pre-r672`. On demo-hp, a full restore-test of 9201 does not fit today
> (21.1 GiB restored needs 30.3 GiB; 22.1 GiB free) — the preflight will refuse it, correctly, until the
> pool has room.
- **Quiescing (controller-driven for app-consistency) — implemented (slice 8B):** an LXC has no
fsfreeze (`proxmox-platform.md` §4.2), so app-consistency is the controller's job: it learns a
backup is due (`GET /backup/due`, §6) → **quiesces** (stops its app stacks) → `POST /backup` →