Files
felhom.eu/documentation/audits/RETIRE-peti-2026-09-25.md
T
2026-09-25 11:33:27 +02:00

102 lines
7.8 KiB
Markdown

# RETIRE — Peti's box (`peti-felhom`), 2026-09-25
Operator ruling 2026-09-25: Peti's box is retired (the tester wiped his server; it will not return); its off-site
backup holds no user data and is to be deleted; it leaves the protected list. Architecture read first:
`runbooks/target-selection.md`, `05-hub-architecture.md` §14 (what a customer delete deprovisions),
`06-offsite-connectivity.md`, `04-control-plane-authorization.md`. Evidence: `audits/retire-peti-2026-09-25/`.
## Not done, or changed
- **Nothing of Peti's was on ep0.** No PBS namespace (`ns`: demo-felhom, demo-hp, tester-1), no WireGuard peer in
the live `wg show`, no config naming him. His only off-site item was a Storage Box **sub-account**
(`u629488-sub2`, id 269130, label `felhom-customer=peti-felhom`, home `felhom-peti-felhom`) on the pool box.
- **The "backup" was never a backup.** The sub-account's folder held ONE file: `.ssh/authorized_keys`, 81 bytes. No
restic repository was ever created — all 482 of his reports (2026-07-10 … 07-15) show 0 off-site snapshots and
0 bytes; his escrow stayed pending. The operator's "no user data" is confirmed by the listing, not only by word.
- **To list and empty the folder I reset the sub-account's password through the Hetzner API** (the product's own
`reset_subaccount_password` action, scoped to id 269130 after re-reading its label). The key file was removed from
inside the folder, then the hub deleted the sub-account. Whether Hetzner deletes a sub-account's folder with it
is NOT measured here — so the folder was emptied first, and nothing of his could remain either way.
- **Cloudflare:** his customer config carried a Cloudflare tunnel token and an API token (for `sajatfelhom.hu`).
They were purged with the customer record. The hub's delete dialog says "tunnel/zone removed", but no leg of the
cascade calls Cloudflare — the tunnel and DNS records on Cloudflare's side, if any remain, were NOT inventoried or
removed (a row: R-688).
- **A second, older Storage Box** (`PBS-storage-1`, u629193, box 611421) once held a folder `felhom-peti-spike/`
(spike leftovers, per `RUNBOOK-ep0-datastore-volume-2026-07-27.md`). Its mount on ep0 is gone, and the box is
not visible to either API token (404). The register row "delete the box" is the operator's. Not touched.
- **Code comments that name Peti were left** (hub `rollup.go`, `appliances.go`, `notify/templates.go`; agent pbsdr
tests). They explain why code is shaped as it is; no behaviour depends on Peti.
## A1 — inventory (before anything was removed)
| where | item | how found |
|---|---|---|
| hub | customer `peti-felhom` („Peti Proxmox", domain `sajatfelhom.hu`, dr_tier 0) with 482 reports, 123 events, 1,036 telemetry rows, DR recipe, one-time secret, claim, 1 log tail | copy of the hub DB (with its -wal), every table's `customer_id` / `host_id` |
| hub | host `peti-felhom-86d37d` — already DELETED 2026-07-15 08:56:22 (`host_deletions`, escrow_acked 0); no host, guest, escrow, recovery or PBS-secret rows | same |
| hub | WireGuard peer — none in `wg_peers` | same |
| Storage Box | sub-account id 269130 `u629488-sub2`, home `felhom-peti-felhom`, label `felhom-customer=peti-felhom`, created 2026-07-10 | Hetzner API with the hub's own token, matched on the LABEL |
| ep0 | nothing (namespaces, peers, `/etc`, `/root`, `/srv` grep — one false hit, "competing" in `lvm.conf`) | ssh read-only |
| documents | 48 non-history files + workspace rules + 21 memory notes (whole-word pattern; positive control target-selection.md 7 hits, negative control catalog templates 0) | `A1-documents.txt` |
**Control, before:** pool box `size_data 3,120,562,176`, sub-accounts sub1 (demo-felhom), sub3 (demo-hp), sub4
(tester-1); ep0 namespaces demo-felhom 2 snapshots / 220K, demo-hp 2 / 540K, tester-1 2 / 368K; demo boxes'
off-site `last_status ok` (11 and 90 snapshots, runs 02:15:46Z / 02:18:41Z).
## A3 — removal
1. `POST …/subaccounts/269130/actions/reset_subaccount_password` → success (password kept 0600 in the scratchpad,
deleted afterwards).
2. As `u629488-sub2`: `rm .ssh/authorized_keys`, `rmdir .ssh` → `ls -la` empty, `du -s .` 1.
3. Hub preview `GET /configs/peti-felhom/delete` — 0 hosts, off-site `u629488-sub2`, residue 1,519. Then
`POST /configs/peti-felhom/delete` (ack_hosts, ack_reset, ack_purge, confirm_id, expect_hosts=0) → 303. The hub's
log: sub-account 269130 deprovisioned; PBS tenancy `existed=false`; claim reset; residue purged (reports 482,
telemetry 1,036, log tails 1); cascade COMPLETE (journal #20).
## A4 — nothing else moved
| item | before | after |
|---|---|---|
| sub-accounts | sub1, sub2 (Peti), sub3, sub4 | sub1, sub3, sub4 — labels and homes unchanged |
| pool box `size_data` | 3,120,562,176 | 3,120,562,176 |
| login as `u629488-sub2` | worked | "Permission denied" |
| ep0 namespaces (snapshots / du) | demo-felhom 2/220K, demo-hp 2/540K, tester-1 2/368K | identical |
| ep0 WireGuard peers | .2 .3 .4 .250 | identical |
| hub hosts | demo-felhom, demo-hp, drill-r50 | identical |
| hub customers | demo-felhom, demo-hp, drill-r50, peti-felhom, tester-1 | without peti-felhom |
| hub rows naming peti | — | events 124, notification_log 80, host_deletions 1, customer_resets 1 (audit, by design), app_log_issues 8 (R-244) |
## A5 — documents changed
`runbooks/target-selection.md` (Tier 2 = DooPlex + ep0, ep0's reason rewritten, Peti's section retired); the
unprompted-work rule (4 identical copies); `03`, `06`, `00-capability-map`; runbooks `TASK-identity-only-escrow`
(obsolete note), `RUNBOOK-vzdump-target-move` (row 9), `RUNBOOK-island-migration`; retired banners on
`pilot/PETI-tester-agreement.md`, `RUNBOOK-peti-return`, `RUNBOOK-peti-pbsdr`; register (PETI closed as retired,
R-530, R-244, R-600 annotated); CONTEXT; STATUS; memory notes. Historic audits, tests and CHANGELOGs keep their text.
## Claims in the brief
1. *Peti's backup is on ep0* — **WRONG.** ep0 held nothing of his; the item was a Storage Box sub-account.
2. *The hub's host delete removes the WireGuard peer* — **not testable here:** the host was deleted in July and
ep0 carries no peer for it now; whether that delete removed one is not recorded (R-600 annotated).
3. *The backup holds no user data* — **TRUE, and stronger:** it held no backup at all (one 81-byte key file).
## Part B — controller v0.272.0 (`44ae4de`), floor 0.272.0
Red-proofs, each seen failing (`retire-peti-2026-09-25/B/redproof-*.txt`): R-670 (undo read with the validating
loader → "the undo logged a false backup-block ERROR"), R-671 (remover call dropped → "copies were left behind"; wiring
dropped → "main.go never calls SetUndoCopyRemover"), R-677 (always catalog_since → "got 3 days, want 0"), R-685 page
(line dropped → "a space skip must be said on the page"). Full suite + controller gates green; parity unchanged.
Live on 9202, drill catalog (`B/live/`): navidrome installed at 0.64.0 with a wrong probe; the step to 0.64.1 failed and
its undo failed → HOLD, 1 undo copy kept; **R-670:** 0 "backup block rejected" lines while the undo ran (my first count
of 3 "unreadable" was disk-watch lines — corrected in `r670-verdict.txt`); **R-671:** the backup page's restore →
"update hold … CLEARED", "removed 1 undo cop(y/ies)", 0 copies left, the seeded user read back; **R-677:** the head
re-tested at a new digest → „Frissítés elérhető — ma" / "Update available — today" (catalog_since two days old).
**R-685 page line: not live** — 9202 has no agent, and no demo box is short of space. Teardown: navidrome removed
through the product (0 volumes, 0 copies), 9202 back on the live catalog, drill reset to live `main`.
Floor saved 09:31:31Z; both demo boxes on 0.272.0 at 09:31:55Z; demo-hp's backup page renders in both languages.
## Part C — the night watch: NOT REACHED
The session ended in the daytime (≈11:40 CEST).