Files
felhom.eu/documentation/runbooks/ep0-datastore-copy.md
T

51 lines
3.4 KiB
Markdown

# Runbook — the ep0 datastore copy on DooPlex (decision 70, R-342)
ep0's PBS datastore `felhom-offsite` (`/mnt/pbs-datastore`, a separate Hetzner Volume that no server snapshot
covers and Hetzner cannot snapshot) is pulled to DooPlex every night. The copy holds **ciphertext only** — each
household's whole-box backups are encrypted with that household's own `encryption-key`; DooPlex cannot read them.
## What exists
| Where | Object | Purpose |
|---|---|---|
| ep0 | API token `root@pam!dooplex-sync`, ACL `DatastoreReader` on `/datastore/felhom-offsite` | read-only pull. **The only change on ep0.** |
| DooPlex | `felhom-ep0-pbs-tunnel.service` (systemd, runs as `kisfenyo`, `Restart=always`) | `ssh -N -L 127.0.0.1:18007:127.0.0.1:8007 root@ep0` — ep0's PBS listens on `wg0` only; DooPlex is not a WireGuard peer |
| DooPlex PBS | remote `ep0` (127.0.0.1:18007, ep0's cert fingerprint pinned) | the pull source |
| DooPlex PBS | datastore `ep0-copy` at `/mnt/5_hdd/backup/ep0-copy` | the copy |
| DooPlex PBS | sync job `ep0-felhom-offsite`, daily 05:00, `remove-vanished false` | the nightly pull (ep0's prune runs 03:30). **It never removes what ep0 removed** — a deletion on ep0 does not reach the copy |
| DooPlex PBS | verify job `verify-ep0-copy`, Saturdays 06:30 | reads the copy back |
| DooPlex PBS | notification target `felhom-operator` (SMTP via Resend → admin@felhom.eu) + matcher `felhom-operator-errors` (every error) | a failed pull or verify reaches the operator. Proven 2026-10-03 with a test mail |
Secrets, all out of git: the ep0 token secret in `/etc/proxmox-backup/remote.cfg` (root:backup 0640, base64 —
PBS's own format), the Resend key in `/etc/proxmox-backup/notifications-priv.cfg` (root:root 0600).
## Checks
```bash
systemctl is-active felhom-ep0-pbs-tunnel.service
curl -sk -o /dev/null -w '%{http_code}\n' https://127.0.0.1:18007/ # 200 = ep0's PBS reachable
sudo proxmox-backup-manager task list --limit 10 | grep -E 'syncjob|verif' # last runs
sudo find /mnt/5_hdd/backup/ep0-copy/ns -mindepth 4 -maxdepth 4 -type d | sort # snapshots per customer
```
## If ep0 is lost — restore a household's whole box from the DooPlex copy
The copy is a normal PBS datastore. Two routes, both needing the household's PBS `encryption-key` (escrowed,
recovered with the household's recovery code — the same as restoring from ep0):
1. **Point the box's host at DooPlex instead of ep0.** On the household's Proxmox host, add a PBS storage for
DooPlex's PBS (`ep0-copy`, namespace = the customer id) with the household's key, then restore the CT from it as
from ep0. DooPlex's PBS must be reachable from the host (it is not public today — the operator decides the route
at the time: a temporary tunnel, or a new endpoint).
2. **Rebuild the endpoint.** Provision a new ep0 (06 §5), then pull back: on the new ep0 add DooPlex as a remote
and run `proxmox-backup-manager pull <dooplex-remote> ep0-copy felhom-offsite`. Every box then reconnects as before.
Do **not** prune or garbage-collect the copy tighter than ep0's own retention. The copy has no prune job today and
grows with every nightly backup (R-828).
## Remove
`sudo proxmox-backup-manager sync-job remove ep0-felhom-offsite; … verify-job remove verify-ep0-copy;`
`sudo systemctl disable --now felhom-ep0-pbs-tunnel.service`; on ep0
`proxmox-backup-manager user delete-token root@pam dooplex-sync`. The datastore's bytes stay until removed by hand.