Files
felhom.eu/documentation/runbooks/ep0-datastore-copy.md
T
2026-10-04 08:56:56 +02:00

81 lines
6.3 KiB
Markdown

# Runbook — the ep0 datastore copy on DooPlex (decision 70, R-342)
ep0's PBS datastore `felhom-offsite` (`/mnt/pbs-datastore`, a separate Hetzner Volume that no server snapshot
covers and Hetzner cannot snapshot) is pulled to DooPlex every night. The copy holds **ciphertext only** — each
household's whole-box backups are encrypted with that household's own `encryption-key`; DooPlex cannot read them.
## What exists
| Where | Object | Purpose |
|---|---|---|
| ep0 | API token `root@pam!dooplex-sync`, ACL `DatastoreReader` on `/datastore/felhom-offsite` | read-only pull. **The only change on ep0.** |
| DooPlex | `felhom-ep0-pbs-tunnel.service` (systemd, runs as `kisfenyo`, `Restart=always`) | `ssh -N -L 127.0.0.1:18007:127.0.0.1:8007 root@ep0` — ep0's PBS listens on `wg0` only; DooPlex is not a WireGuard peer |
| DooPlex PBS | remote `ep0` (127.0.0.1:18007, ep0's cert fingerprint pinned) | the pull source |
| DooPlex PBS | datastore `ep0-copy` at `/mnt/5_hdd/backup/ep0-copy` | the copy |
| DooPlex PBS | sync job `ep0-felhom-offsite`, daily 05:00, `remove-vanished false` | the nightly pull (ep0's prune runs 03:30). **It never removes what ep0 removed** — a deletion on ep0 does not reach the copy |
| DooPlex PBS | verify job `verify-ep0-copy`, Saturdays 06:30 | reads the copy back |
| DooPlex PBS | prune job `prune-ep0-copy`, daily 07:30, **keep-weekly 8**, all namespaces (decision 71) | keeps the last 8 weekly copies per group; runs after the 05:00 pull, never during it |
| DooPlex PBS | garbage collection on `ep0-copy`, Sundays 08:30 | frees the chunks the prune released |
| DooPlex PBS | notification target `felhom-operator` (SMTP via Resend → admin@felhom.eu) + matcher `felhom-operator-errors` (every error) | a failed pull or verify reaches the operator. Proven 2026-10-03 with a test mail |
Secrets, all out of git: the ep0 token secret in `/etc/proxmox-backup/remote.cfg` (root:backup 0640, base64 —
PBS's own format), the Resend key in `/etc/proxmox-backup/notifications-priv.cfg` (root:root 0600).
## Checks
```bash
systemctl is-active felhom-ep0-pbs-tunnel.service
curl -sk -o /dev/null -w '%{http_code}\n' https://127.0.0.1:18007/ # 200 = ep0's PBS reachable
sudo proxmox-backup-manager task list --limit 10 | grep -E 'syncjob|verif' # last runs
sudo find /mnt/5_hdd/backup/ep0-copy/ns -mindepth 4 -maxdepth 4 -type d | sort # snapshots per customer
```
## If ep0 is lost — restore a household's whole box from the DooPlex copy
The copy is a normal PBS datastore. Two routes, both needing the household's PBS `encryption-key` (escrowed,
recovered with the household's recovery code — the same as restoring from ep0).
### Route 1 — the host reads DooPlex directly. **WALKED 2026-10-04 on demo-hp** (`audits/offsite-finish-2026-10-04/partD-restore-walk.txt`)
1. **On DooPlex:** a read-only token for the restore (`proxmox-backup-manager user generate-token root@pam <name>`, then
`acl update /datastore/ep0-copy DatastoreReader --auth-id 'root@pam!<name>'`). Keep the secret in a file only.
2. **On the host:** put the token in `/etc/pve/priv/storage/<id>.pw` (0600) and the household's PBS key in
`/etc/pve/priv/storage/<id>.enc`, then append the storage entry to `/etc/pve/storage.cfg`:
`pbs: <id>` / `datastore ep0-copy` / `server <DooPlex>` / `content backup` / `fingerprint <DooPlex PBS cert>` /
`namespace <customer>` / `username root@pam!<name>`.
**Do not use `pvesm add pbs` without `--password`:** it validates with the password from its command line, fails 401,
and on failure DELETES the `.pw`/`.enc` files you placed (measured). Passing `--password` puts the token on argv.
3. `pvesm list <id>` → the household's snapshots (measured: 2 s). Then restore with the safe script (R-834):
`felhom-restore-beside.sh <scratch VMID> <id>:backup/ct/<vmid>/<time> <dir storage>` (copy it from
`felhom.eu/scripts/`; measured 2026-10-03 by hand: **186 s for a 15 GB-logical / 14 GB-on-disk backup** over the LAN).
4. **⚠ Why the script, not a bare `pct restore`:** the archive's config is the PRODUCTION one — `onboot: 1`, `mp8` bound
to the host's REAL household drives (`/mnt/felhom-drives`) and `mp9` to the original guest's bootstrap. Started, or
after a host reboot, it is a second controller for the same household on the same drives. The script restores with
`--onboot 0`, removes every host-path bind, takes every NIC down, never starts it, and reads the config back
(proven live 2026-10-04, `audits/backup-close-2026-10-04/partA/`). On a true replacement host, where the original is
gone, the binds are what you want: use `RUNBOOK-manual-guest-restore.md` §3 or the agent's DR bring-up instead (the
DR bring-up refuses beside a live original since agent v0.139.0).
5. Read the data without starting it: `pct mount <vmid>` → `/var/lib/lxc/<vmid>/rootfs` (measured: 1 s; rootfs, the
controller data volume and `/var/lib/felhom` present, `settings.json` dated 10 min before the backup) → `pct unmount`.
6. Teardown: `pct destroy <vmid> --purge`; `pvesm remove <id>` (removes the entry and its priv files — never the
household's own `felhom-pbs.enc`); on DooPlex delete the token and its ACL.
**Reachability.** DooPlex's PBS (`:8007`) is reachable on DooPlex's LAN only. For a host on another network — a
household's home — the options, none built (the operator decides at the time, R-832 is the long-term answer):
(a) a temporary SSH forward from the host to DooPlex, as DooPlex already does to ep0 (needs an SSH key on DooPlex for
that host); (b) a WireGuard peer on DooPlex's existing tailscale/k3s network for the duration; (c) route 2 below —
rebuild an endpoint and pull back, so every host reconnects the usual way.
### Route 2 — rebuild the endpoint. Not walked.
Provision a new ep0 (06 §5), then on it add DooPlex as a remote and run
`proxmox-backup-manager pull <dooplex-remote> ep0-copy felhom-offsite`. Every box then reconnects as before.
Do **not** prune the copy tighter than decision 71 (8 weekly copies); never tighter than ep0's own retention.
## Remove
`sudo proxmox-backup-manager sync-job remove ep0-felhom-offsite; … verify-job remove verify-ep0-copy; … prune-job remove prune-ep0-copy;`
`sudo systemctl disable --now felhom-ep0-pbs-tunnel.service`; on ep0
`proxmox-backup-manager user delete-token root@pam dooplex-sync`. The datastore's bytes stay until removed by hand.