R-99 runbook (pbs-phantom-cleanup) + read-only ep0 listing: no phantom today, nothing deleted; R-444 trim measured on demo-hp; R-618 closed by ruling (150 -> 149)
gates / gates (push) Successful in 2m18s
gates / gates (push) Successful in 2m18s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
@@ -0,0 +1,16 @@
|
||||
== R-444 trim measurement, demo-hp, 2026-10-06T09:13:53Z
|
||||
-- before
|
||||
data 53.87g 65.53
|
||||
vm-9201-disk-0 32.00g 11.45
|
||||
vm-9201-disk-1 70.00g 45.20
|
||||
/dev/mapper/pve-vm--9201--disk--0 32716560 1161396 29861060 4% /
|
||||
/dev/mapper/pve-vm--9201--disk--1 72064432 15456272 52921772 23% /var/lib/felhom
|
||||
-- pct fstrim 9201 (start 09:13:59)
|
||||
/var/lib/lxc/9201/rootfs/: 30.1 GiB (32277680128 bytes) trimmed
|
||||
/var/lib/lxc/9201/rootfs/var/lib/felhom: 53.9 GiB (57865633792 bytes) trimmed
|
||||
rc=0 duration_s=24.367926462
|
||||
-- after
|
||||
data 53.87g 33.40
|
||||
vm-9201-disk-0 32.00g 6.01
|
||||
vm-9201-disk-1 70.00g 22.96
|
||||
-- probe (UTC time, http code, seconds): 18 samples; non-200: 0; max latency: 09:14:01 200 1.115079192
|
||||
@@ -0,0 +1,16 @@
|
||||
== R-99 ep0 listing, read only, 2026-10-06 ~11:25 CEST: datastore felhom-offsite (/mnt/pbs-datastore), every snapshot directory
|
||||
ns | group | snapshot | manifest (index.json.blob) | directory bytes
|
||||
Tester-2 | ct/9201 | 2026-10-04T16:31:13Z | True | 35841
|
||||
demo-felhom | ct/9201 | 2026-09-22T04:12:20Z | True | 86404
|
||||
demo-felhom | ct/9201 | 2026-09-29T04:16:43Z | True | 101070
|
||||
demo-felhom | ct/9201 | 2026-10-06T04:21:09Z | True | 43485
|
||||
demo-hp | ct/9201 | 2026-09-24T20:06:25Z | True | 279204
|
||||
demo-hp | ct/9201 | 2026-10-01T20:15:29Z | True | 187064
|
||||
operator | host/dooplex-hub | 2026-10-05T14:23:25Z | True | 10371
|
||||
operator | host/dooplex-hub | 2026-10-06T00:32:21Z | True | 10372
|
||||
tester-1 | ct/9201 | 2026-10-04T19:57:18Z | True | 72686
|
||||
leftover candidates (no manifest, or the PBS-reported size < 1 MiB): 0
|
||||
== the backup server's own view (proxmox-backup-debug api get …/snapshots, every namespace): 9 snapshots, sizes
|
||||
2096481498, 6956268441, 5840921439, 2749038679, 23236593219, 15223817295, 372921211, 369808250, 4913374016 bytes; every
|
||||
verification state "ok". Smallest 369808250 B, 352x above the 1 MiB phantom line. Real snapshots per namespace:
|
||||
Tester-2 1, demo-felhom 3, demo-hp 2, operator 2, tester-1 1. Nothing deleted (nothing to delete).
|
||||
@@ -26,6 +26,14 @@
|
||||
|
||||
---
|
||||
|
||||
## 2026-10-06 (midday) — the ten answers
|
||||
|
||||
The full text of every row below: `git show 8c65ff0c:documentation/backlog/OPEN-ITEMS.md`.
|
||||
|
||||
| Row | What | Closed | Evidence |
|
||||
|---|---|---|---|
|
||||
| **R-618** | **[P1-HIGH] THREE apps are presented to the household as UNHEALTHY while they are working perfectly — and because the guarded Update waits on that same probe, a SUCCESSFUL update ends by STOPPING the working app and sending the household to a restore they do not need.** (P4) | CLOSED 2026-10-06 — RULED (operator, `09` §3 decision 141): app updates keep our own health check only | The three wrong probes and their gate shipped 2026-09-22 (the row). The open idea — let `verifying` accept Docker's own `healthy` — is REJECTED by the ruling: Docker's healthy does not end a wait early (R-635: a probe can be green on a broken app). Nothing to build. |
|
||||
|
||||
## 2026-10-06 (morning) — R-889 delivered, the held catalog fixes proven and pushed
|
||||
|
||||
The full text of every row below: `git show e6cfe6d6:documentation/backlog/OPEN-ITEMS.md`.
|
||||
|
||||
File diff suppressed because one or more lines are too long
@@ -0,0 +1,62 @@
|
||||
# Runbook — delete a phantom backup snapshot (an aborted upload's leftover) on the off-site backup server
|
||||
|
||||
**Why it exists:** operator ruling 2026-10-06 (`architecture/09-update-architecture.md` §3 decision 140, R-99): *"Leftovers
|
||||
should be deleted."* A PBS daemon killed in the middle of an upload leaves a snapshot that holds nothing (measured 1 B,
|
||||
2026-07-28, F-CRIT-2). The box already refuses to count it as a backup (felhom-agent `internal/backup/runner.go`,
|
||||
`archivePlausiblyComplete`, the 1 MiB floor) and logs it once at WARN. Server-side prune never removes it (dry-run
|
||||
2026-07-28: keep-last 2 kept two real snapshots PLUS the phantom). So one accumulates per aborted upload, until a person
|
||||
deletes it — by this runbook, when one is seen. **ep0 is Tier 2, protected** (`runbooks/target-selection.md`): this runbook
|
||||
is the ONLY deletion the ruling allows there, and only of a snapshot this runbook proves to be a phantom.
|
||||
|
||||
## When you see one
|
||||
|
||||
The box's agent log says `… size N B is below the 1048576 B plausibility floor — an aborted/incomplete archive, not a
|
||||
successful backup` (once per volid). Or a session reads one in step 1.
|
||||
|
||||
## 1. List — read only
|
||||
|
||||
From DooPlex (`ssh root@167.233.158.164`; never via `felhom-pve → 10.77.0.1`):
|
||||
|
||||
```bash
|
||||
# the server's own view, every namespace (size = what the manifest references; verification state)
|
||||
for ns in $(ls /mnt/pbs-datastore/ns); do
|
||||
proxmox-backup-debug api get /admin/datastore/felhom-offsite/snapshots --ns "$ns" --output-format json-pretty
|
||||
done
|
||||
# the directories themselves (a snapshot with no manifest may not be listed by the API at all)
|
||||
python3 - < documentation/runbooks/pbs-phantom-list.py # run on ep0: ssh root@… python3 - < that file
|
||||
```
|
||||
|
||||
**A snapshot is a PHANTOM only when BOTH hold:**
|
||||
1. the server reports `size` below **1 MiB** (1,048,576 B) — or the directory has **no `index.json.blob`** (no manifest);
|
||||
2. it is **not** `protected`, and it is **not the newest** snapshot of its group while an upload for that group may still be
|
||||
running (check `proxmox-backup-manager task list --all --limit 20` — no running `backup` task for that group).
|
||||
|
||||
The smallest real backup ever measured is ~584 MiB; real ones on 2026-10-06 were ≥ 352 MiB. Anything between 1 MiB and the
|
||||
smallest real size is **NOT** a phantom by this runbook — leave it and ask the operator.
|
||||
|
||||
**Write down, per namespace, the count of real snapshots** (everything that is not a phantom). This count is the control.
|
||||
|
||||
## 2. Delete — one snapshot at a time, the command shown first
|
||||
|
||||
```bash
|
||||
# print it, read it, then run it
|
||||
echo proxmox-backup-debug api delete /admin/datastore/felhom-offsite/snapshots \
|
||||
--ns <namespace> --backup-type <ct|vm|host> --backup-id <id> --backup-time <unix time>
|
||||
proxmox-backup-debug api delete /admin/datastore/felhom-offsite/snapshots \
|
||||
--ns <namespace> --backup-type <ct|vm|host> --backup-id <id> --backup-time <unix time>
|
||||
```
|
||||
|
||||
Never `rm -rf` a snapshot directory, never `prune`, never touch a datastore, a namespace or a group. One `delete` per
|
||||
proven phantom.
|
||||
|
||||
## 3. Prove nothing real moved
|
||||
|
||||
Re-run step 1. **The count of real snapshots in every namespace must be exactly what it was before.** The deleted
|
||||
phantom is gone; nothing else changed. If a count moved, stop and tell the operator at once. Save both listings (before,
|
||||
after) under `documentation/audits/<session>/`.
|
||||
|
||||
## The record of each use
|
||||
|
||||
| Date | Namespace / group / time | Size | Why a phantom | Real counts before = after |
|
||||
|---|---|---|---|---|
|
||||
| 2026-10-06 | — | — | **No phantom existed** (9 snapshots, all ≥ 369,808,250 B, all verification `ok`; no directory without a manifest) — nothing deleted | Tester-2 1, demo-felhom 3, demo-hp 2, operator 2, tester-1 1 (`audits/ten-answers-2026-10-06/r99-ep0-listing.txt`) |
|
||||
@@ -0,0 +1,23 @@
|
||||
import os, re, json
|
||||
root = "/mnt/pbs-datastore"
|
||||
rows = []
|
||||
def walk_ns(base, ns):
|
||||
for typ in ("ct", "vm", "host"):
|
||||
tdir = os.path.join(base, typ)
|
||||
if not os.path.isdir(tdir): continue
|
||||
for gid in sorted(os.listdir(tdir)):
|
||||
gdir = os.path.join(tdir, gid)
|
||||
for snap in sorted(os.listdir(gdir)):
|
||||
sdir = os.path.join(gdir, snap)
|
||||
if not os.path.isdir(sdir) or not re.match(r"\d{4}-\d{2}-\d{2}T", snap): continue
|
||||
files = os.listdir(sdir)
|
||||
size = sum(os.path.getsize(os.path.join(sdir, f)) for f in files)
|
||||
man = "index.json.blob" in files
|
||||
manifest_files = []
|
||||
rows.append({"ns": ns or "(root)", "group": f"{typ}/{gid}", "snap": snap, "files": sorted(files), "dir_bytes": size, "manifest": man})
|
||||
nsdir = os.path.join(base, "ns")
|
||||
if os.path.isdir(nsdir):
|
||||
for n in sorted(os.listdir(nsdir)):
|
||||
walk_ns(os.path.join(nsdir, n), (ns + "/" if ns else "") + n)
|
||||
walk_ns(root, "")
|
||||
print(json.dumps(rows))
|
||||
Reference in New Issue
Block a user