Files
felhom.eu/documentation/runbooks/ep0-datastore-copy.md
T

8.6 KiB

Runbook — the ep0 datastore copy on DooPlex (decision 70, R-342)

ep0's PBS datastore felhom-offsite (/mnt/pbs-datastore, a separate Hetzner Volume that no server snapshot covers and Hetzner cannot snapshot) is pulled to DooPlex every night. The copy holds ciphertext only — each household's whole-box backups are encrypted with that household's own encryption-key; DooPlex cannot read them.

What exists

Where Object Purpose
ep0 API token root@pam!dooplex-sync, ACL DatastoreReader on /datastore/felhom-offsite read-only pull. The only change on ep0.
DooPlex felhom-ep0-pbs-tunnel.service (systemd, runs as kisfenyo, Restart=always) ssh -N -L 127.0.0.1:18007:127.0.0.1:8007 root@ep0 — ep0's PBS listens on wg0 only; DooPlex is not a WireGuard peer
DooPlex PBS remote ep0 (127.0.0.1:18007, ep0's cert fingerprint pinned) the pull source
DooPlex PBS datastore ep0-copy at /mnt/5_hdd/backup/ep0-copy the copy
DooPlex PBS sync job ep0-felhom-offsite, daily 05:00, remove-vanished false the nightly pull (ep0's prune runs 03:30). It never removes what ep0 removed — a deletion on ep0 does not reach the copy
DooPlex PBS verify job verify-ep0-copy, Saturdays 06:30 reads the copy back
DooPlex PBS prune job prune-ep0-copy, daily 07:30, keep-weekly 8, all namespaces (decision 71) keeps the last 8 weekly copies per group; runs after the 05:00 pull, never during it
DooPlex PBS garbage collection on ep0-copy, Sundays 08:30 frees the chunks the prune released
DooPlex PBS notification target felhom-operator (SMTP via Resend → admin@felhom.eu) + matcher felhom-operator-errors (every error) a failed pull or verify reaches the operator. Proven 2026-10-03 with a test mail

Secrets, all out of git: the ep0 token secret in /etc/proxmox-backup/remote.cfg (root:backup 0640, base64 — PBS's own format), the Resend key in /etc/proxmox-backup/notifications-priv.cfg (root:root 0600).

Checks

systemctl is-active felhom-ep0-pbs-tunnel.service
curl -sk -o /dev/null -w '%{http_code}\n' https://127.0.0.1:18007/          # 200 = ep0's PBS reachable
sudo proxmox-backup-manager task list --limit 10 | grep -E 'syncjob|verif'   # last runs
sudo find /mnt/5_hdd/backup/ep0-copy/ns -mindepth 4 -maxdepth 4 -type d | sort # snapshots per customer

If ep0 is lost — restore a household's whole box from the DooPlex copy

The copy is a normal PBS datastore. Two routes, both needing the household's PBS encryption-key (escrowed, recovered with the household's recovery code — the same as restoring from ep0).

Route 1 — the host reads DooPlex directly. WALKED 2026-10-04 on demo-hp (audits/offsite-finish-2026-10-04/partD-restore-walk.txt)

  1. On DooPlex: a read-only token for the restore (proxmox-backup-manager user generate-token root@pam <name>, then acl update /datastore/ep0-copy DatastoreReader --auth-id 'root@pam!<name>'). Keep the secret in a file only.
  2. On the host: put the token in /etc/pve/priv/storage/<id>.pw (0600) and the household's PBS key in /etc/pve/priv/storage/<id>.enc, then append the storage entry to /etc/pve/storage.cfg: pbs: <id> / datastore ep0-copy / server <DooPlex> / content backup / fingerprint <DooPlex PBS cert> / namespace <customer> / username root@pam!<name>. The namespace comes from the hub's Recipe (pbs.namespace, read WITH pbs.namespace_state): resolved and a name → namespace <name>; resolved and an empty value → the ROOT namespace: write no namespace line (and pass no --ns to proxmox-backup-client). Agents before v0.147.0 wrote the word root there — no namespace has that name, so treat a recorded root as empty (R-124). unknown → read the namespace off the copy's ns/ folder above. Do not use pvesm add pbs without --password: it validates with the password from its command line, fails 401, and on failure DELETES the .pw/.enc files you placed (measured). Passing --password puts the token on argv.
  3. pvesm list <id> → the household's snapshots (measured: 2 s). Then restore with the safe script (R-834): felhom-restore-beside.sh <scratch VMID> <id>:backup/ct/<vmid>/<time> <dir storage> (copy it from felhom.eu/scripts/; measured 2026-10-03 by hand: 186 s for a 15 GB-logical / 14 GB-on-disk backup over the LAN).
  4. ⚠ Why the script, not a bare pct restore: the archive's config is the PRODUCTION one — onboot: 1, mp8 bound to the host's REAL household drives (/mnt/felhom-drives) and mp9 to the original guest's bootstrap. Started, or after a host reboot, it is a second controller for the same household on the same drives. The script restores with --onboot 0, removes every host-path bind, takes every NIC down, never starts it, and reads the config back (proven live 2026-10-04, audits/backup-close-2026-10-04/partA/). On a true replacement host, where the original is gone, the binds are what you want: use RUNBOOK-manual-guest-restore.md §3 or the agent's DR bring-up instead (the DR bring-up refuses beside a live original since agent v0.139.0).
  5. Read the data without starting it: pct mount <vmid> → /var/lib/lxc/<vmid>/rootfs (measured: 1 s; rootfs, the controller data volume and /var/lib/felhom present, settings.json dated 10 min before the backup) → pct unmount.
  6. Teardown: pct destroy <vmid> --purge; pvesm remove <id> (removes the entry and its priv files — never the household's own felhom-pbs.enc); on DooPlex delete the token and its ACL.

Reachability. DooPlex's PBS (:8007) is reachable on DooPlex's LAN only. For a host on another network — a household's home — the options, none built (the operator decides at the time, R-832 is the long-term answer): (a) a temporary SSH forward from the host to DooPlex, as DooPlex already does to ep0 (needs an SSH key on DooPlex for that host); (b) a WireGuard peer on DooPlex's existing tailscale/k3s network for the duration; (c) route 2 below — rebuild an endpoint and pull back, so every host reconnects the usual way.

Route 2 — rebuild the endpoint. Not walked.

Provision a new ep0 (06 §5), then on it add DooPlex as a remote and run proxmox-backup-manager pull <dooplex-remote> ep0-copy felhom-offsite. Every box then reconnects as before.

Do not prune the copy tighter than decision 71 (8 weekly copies); never tighter than ep0's own retention.

Removing a deleted customer's copy (R-901, 09 §3 decision 181) — WRITTEN 2026-10-08, NOT INSTALLED

The sync keeps what ep0 removed (remove-vanished false), so a deleted customer's namespace stayed here for ever. Ruling: removed within 30 days. The job scripts/ep0-copy-gc/felhom-ep0-copy-gc (tests beside it) runs daily at 08:00: a top-level namespace in the copy that ep0 no longer lists is remembered with the day it was first seen absent; after 7 days absent it is deleted from the copy (proxmox-backup-client namespace delete <ns> --delete-groups true); one that reappears is forgotten; operator is never deleted. If ep0's list cannot be read, or reads empty, it does nothing. Without --apply it only prints. Worst case: ≤ 1 day to notice + 7 days grace + 1 day = inside 30 days.

Not installed. To install (needs the operator's word — it is a change on DooPlex and it deletes data):

  1. Two tokens, secrets in root-only files under /etc/felhom/ep0-copy-gc/ (0600), never printed:
    • ep0-reader.secret — the existing ep0 read token root@pam!dooplex-sync (its secret is in remote.cfg, base64); ep0.fingerprint — ep0's pinned certificate fingerprint (also in remote.cfg).
    • local-gc.secret — a NEW local token: proxmox-backup-manager user generate-token root@pam ep0-copy-gc, then proxmox-backup-manager acl update /datastore/ep0-copy DatastoreAdmin --auth-id 'root@pam!ep0-copy-gc'.
  2. install -m 755 scripts/ep0-copy-gc/felhom-ep0-copy-gc /usr/local/sbin/; the .service and .timer into /etc/systemd/system/; systemctl daemon-reload.
  3. First run by hand, DRY: sudo /usr/local/sbin/felhom-ep0-copy-gc — read which namespaces it would remove; then systemctl enable --now felhom-ep0-copy-gc.timer.
  4. The chunks are freed by the Sunday 08:30 garbage collection.

Remove

sudo proxmox-backup-manager sync-job remove ep0-felhom-offsite; … verify-job remove verify-ep0-copy; … prune-job remove prune-ep0-copy; sudo systemctl disable --now felhom-ep0-pbs-tunnel.service; on ep0 proxmox-backup-manager user delete-token root@pam dooplex-sync. The datastore's bytes stay until removed by hand.