v0.174.0 — R-82 Slice B: one quiesce window, two backup tiers

MinAgent UNCHANGED — degrades gracefully against ANY older agent.

The agent gained per-target tiers in v0.97.0. The controller owns quiescing,
so the multi-tier schedule is reconciled here: every due tier is collected up
front and run inside ONE quiesce window (one stop, N sequential backups, one
resume). Two cycles on the weekly night would mean two app outages for one
night's work.

Dedup rule: local-only -> one quiesce; PBS-only -> one quiesce; BOTH due ->
ONE window with both backups inside; neither -> no quiesce.

- quiesce.TieredBackend + BackupTier + ErrTiersUnsupported (optional extension)
- agentapi: BackupTiers/BackupDueFor/StartBackupFor/BackupStatusFor;
  targetQuery("") yields an EMPTY suffix so untargeted hits the pre-R-82 route
  byte-for-byte
- Loop.resolveDueTiers = the dedup rule in one place, agent order preserved
- quiesceAndPollTiers + pollTier: app stays quiesced until the LAST tier
  snapshots (resuming earlier loses app-consistency on the DR tier). Consequence
  stated in the docs: both-due-night downtime = first tier's full backup + last
  tier's snapshot, which is why tiers run fast-first.
- Manual 'Mentes most' covers EVERY tier, due-ness ignored.
- Window-gate safety valve now uses the OLDEST due tier, so a stale DR tier
  cannot be starved by a fresher local one.

Capability detection: /backup/tiers 404 = pre-R-82 agent (the documented
route-probe mechanism). Not a featureProbes row on purpose — the loop needs the
tier LIST, not a yes/no. Degrade logged exactly once per process.

Tests +11, full suite green. Red-proofs #2 and #3 observed and restored.
This commit is contained in:
Claude Code
2026-07-26 14:40:44 +02:00
parent 47fda06ba1
commit de96efc0c5
8 changed files with 924 additions and 42 deletions
+14
View File
@@ -701,6 +701,20 @@ The nightly backup has two phases that run sequentially. All paths are **per-dri
> the three daily legs via `scheduler.UpdateDaily` and takes effect **without a restart**. Precedence:
> settings > controller.yaml `db_dump_schedule` > "02:30". See `internal/backupwindow`.
> **Multi-tier whole-guest backup (v0.174.0, R-82 Slice B).** The agent can serve SEVERAL whole-guest
> backup tiers with independent cadences — "local daily + PBS weekly" (agent >= v0.97.0,
> `GET /backup/tiers`). The controller owns quiescing, so it reconciles them: it collects EVERY due
> tier up front and runs them inside **ONE quiesce window** — one stop, N sequential backups (vzdump
> holds a guest lock), one resume. Two cycles on the weekly night would mean two app outages for one
> night's work. The app stays quiesced until the **LAST** tier snapshots, so every tier is
> app-consistent; the consequence is that both-due-night downtime is *(first tier's full backup)* +
> *(last tier's snapshot)*, which is why tiers run fast-first (the agent advertises primary/local
> first). A manual **"Mentés most"** covers every tier, due-ness ignored. The window gate's safety
> valve evaluates the OLDEST due tier, so a stale DR tier cannot be starved by a fresher local one.
> Against a **pre-R-82 agent** (`/backup/tiers` 404s) the loop degrades to the single untargeted
> tier, logs it once, and still takes the backup — MinAgent is unchanged. See `internal/quiesce`
> (`tiers.go`) and `internal/agentapi/backup_tiers.go`.
> **Atomic dump writes (v0.118.0, CAMPAIGN-3 F7).** BOTH dump paths are crash-safe: the DB dump
> (`dbdump.go` DumpOne) and the Docker-volume dump (`DumpAppVolumes`) write to a `.tmp` sibling, fsync,
> then `os.Rename` over the restore point ONLY on success. A mid-write failure (a NFS cut mid-tar, an