docs(pbs): move PBS prune server-side, close the write proof, schedule GC
Supervised runbook execution. No code, no version bump. The felhom-pbs tier had reported `job errors` on EVERY demo-hp backup since the tier was created on 07-26, while the data landed correctly every time: `DatastoreBackup` grants Datastore.Backup but not Datastore.Prune, so the box's keep_last=2 prune was denied. Operator ruling: retention is a COMMERCIAL attribute owned by the hub; ep0 executes. Box tokens therefore stay write-only - a compromised box must not be able to delete its own offsite backups. No grant was widened and felhom-tenantsync.sh is unchanged (the ruling makes it correct). Increment 1: - boxes stop attempting prune. allowPBSPrune is DERIVED (`!t.Primary && t.KeepLast > 0`), so keep_last: 0 on the PBS tier disables both the --prune-backups value and the gate in one config edit, and the tier stays armed. Verified prune_pbs_allowed=false on both boxes with no tier REJECTED line. - per-namespace prune jobs on ep0, keep-last 2, daily 03:30 UTC (05:30 CEST), dry-run gated. demo-hp 3->2, demo-felhom untouched, chunk count unchanged (prune removes indexes, not chunks). Write proof CLOSED: 08:25:47 job errors -> 09:37:29 TASK OK, snapshot 2026-07-27T09:37:29Z, chunks 9787->9813, prune step absent entirely. Driven through POST /api/guest-backup/trigger (the UI path), not --selftest and not raw vzdump. Hub gauge evidence explicitly NOT satisfied - the delta is below its 0.1 GB display granularity. GC scheduled sun 04:30 UTC and deliberately NOT run: every chunk still carries a fresh atime from the migration copy, so a run today would reclaim nothing. verify-new enabled per operator ruling, turning an inert hub alarm live. Legacy demo-felhom-01 namespace deleted with its two ACL entries and its token (operator ruling, confirmed twice) so nothing dangles. R-89 records the target architecture and carries the unanswered parallel question: does the restic key on storage-box-pool-1 have DELETE rights? If so the daily app-data tier has the identical exposure and append-only is the equivalent answer. ep0 is Etc/UTC, not CEST - corrected in the record.
This commit is contained in:
+11
@@ -57,6 +57,17 @@ or retiring it is → **R-83**.
|
||||
thresholds (host 26h / offsite 8d); fresh-install default; an unprovisioned tier DEFERS.
|
||||
**Operator rulings:** 2-week offsite retention, first backup runs as long as it needs, one backup
|
||||
at a time per guest, drill box dropped from the rollout.
|
||||
**RETENTION IS A COMMERCIAL ATTRIBUTE — the hub decides, ep0 executes (operator ruling 2026-07-27,
|
||||
R-89).** A paid tier may buy longer retention, so the policy belongs with customer config on the
|
||||
hub, never in ep0's PBS config and never in a box's config. Execution stays server-side: a
|
||||
reconciler writes a **PBS prune job** and PBS's own scheduler runs it, so hub downtime leaves the
|
||||
last-known policy running rather than silently stopping retention. **Box tokens stay write-only
|
||||
(`DatastoreBackup`) — never widen a grant to fix a prune error:** a compromised box must not be
|
||||
able to delete its own offsite backups, which is the scenario offsite DR exists to survive.
|
||||
Increment 1 shipped 2026-07-27 (boxes stop attempting prune via `keep_last: 0`; per-namespace prune
|
||||
jobs on ep0, daily 03:30 UTC) — `runbooks/RUNBOOK-pbs-prune-serverside-2026-07-27.md`. This closed a
|
||||
live false-negative: **every** demo-hp PBS backup since 07-26 reported `job errors` while the data
|
||||
landed correctly, because `DatastoreBackup` carries no `Datastore.Prune`.
|
||||
**Four defects found by RUNNING it, not reviewing it** — a 30-min wait bound against a 41-min
|
||||
backup (the agent recorded `success:false` while the backup was still going); the restore tier read
|
||||
from the configured target instead of the archive (**a silent regression of the S4.1 fix** — the
|
||||
|
||||
Reference in New Issue
Block a user