docs(pbs): move PBS prune server-side, close the write proof, schedule GC

Supervised runbook execution. No code, no version bump.

The felhom-pbs tier had reported `job errors` on EVERY demo-hp backup
since the tier was created on 07-26, while the data landed correctly
every time: `DatastoreBackup` grants Datastore.Backup but not
Datastore.Prune, so the box's keep_last=2 prune was denied.

Operator ruling: retention is a COMMERCIAL attribute owned by the hub;
ep0 executes. Box tokens therefore stay write-only - a compromised box
must not be able to delete its own offsite backups. No grant was widened
and felhom-tenantsync.sh is unchanged (the ruling makes it correct).

Increment 1:
- boxes stop attempting prune. allowPBSPrune is DERIVED
  (`!t.Primary && t.KeepLast > 0`), so keep_last: 0 on the PBS tier
  disables both the --prune-backups value and the gate in one config
  edit, and the tier stays armed. Verified prune_pbs_allowed=false on
  both boxes with no tier REJECTED line.
- per-namespace prune jobs on ep0, keep-last 2, daily 03:30 UTC
  (05:30 CEST), dry-run gated. demo-hp 3->2, demo-felhom untouched,
  chunk count unchanged (prune removes indexes, not chunks).

Write proof CLOSED: 08:25:47 job errors -> 09:37:29 TASK OK, snapshot
2026-07-27T09:37:29Z, chunks 9787->9813, prune step absent entirely.
Driven through POST /api/guest-backup/trigger (the UI path), not
--selftest and not raw vzdump. Hub gauge evidence explicitly NOT
satisfied - the delta is below its 0.1 GB display granularity.

GC scheduled sun 04:30 UTC and deliberately NOT run: every chunk still
carries a fresh atime from the migration copy, so a run today would
reclaim nothing. verify-new enabled per operator ruling, turning an
inert hub alarm live.

Legacy demo-felhom-01 namespace deleted with its two ACL entries and its
token (operator ruling, confirmed twice) so nothing dangles.

R-89 records the target architecture and carries the unanswered parallel
question: does the restic key on storage-box-pool-1 have DELETE rights?
If so the daily app-data tier has the identical exposure and append-only
is the equivalent answer.

ep0 is Etc/UTC, not CEST - corrected in the record.
This commit is contained in:
2026-07-27 15:16:11 +02:00
parent a16896af86
commit a31872ea24
4 changed files with 408 additions and 0 deletions
+11
View File
@@ -57,6 +57,17 @@ or retiring it is → **R-83**.
thresholds (host 26h / offsite 8d); fresh-install default; an unprovisioned tier DEFERS.
**Operator rulings:** 2-week offsite retention, first backup runs as long as it needs, one backup
at a time per guest, drill box dropped from the rollout.
**RETENTION IS A COMMERCIAL ATTRIBUTE — the hub decides, ep0 executes (operator ruling 2026-07-27,
R-89).** A paid tier may buy longer retention, so the policy belongs with customer config on the
hub, never in ep0's PBS config and never in a box's config. Execution stays server-side: a
reconciler writes a **PBS prune job** and PBS's own scheduler runs it, so hub downtime leaves the
last-known policy running rather than silently stopping retention. **Box tokens stay write-only
(`DatastoreBackup`) — never widen a grant to fix a prune error:** a compromised box must not be
able to delete its own offsite backups, which is the scenario offsite DR exists to survive.
Increment 1 shipped 2026-07-27 (boxes stop attempting prune via `keep_last: 0`; per-namespace prune
jobs on ep0, daily 03:30 UTC) — `runbooks/RUNBOOK-pbs-prune-serverside-2026-07-27.md`. This closed a
live false-negative: **every** demo-hp PBS backup since 07-26 reported `job errors` while the data
landed correctly, because `DatastoreBackup` carries no `Datastore.Prune`.
**Four defects found by RUNNING it, not reviewing it** — a 30-min wait bound against a 41-min
backup (the agent recorded `success:false` while the backup was still going); the restore tier read
from the configured target instead of the archive (**a silent regression of the S4.1 fix** — the