207ad19746
gates / gates (push) Successful in 41s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
51 lines
3.4 KiB
Markdown
51 lines
3.4 KiB
Markdown
# Runbook — the ep0 datastore copy on DooPlex (decision 70, R-342)
|
|
|
|
ep0's PBS datastore `felhom-offsite` (`/mnt/pbs-datastore`, a separate Hetzner Volume that no server snapshot
|
|
covers and Hetzner cannot snapshot) is pulled to DooPlex every night. The copy holds **ciphertext only** — each
|
|
household's whole-box backups are encrypted with that household's own `encryption-key`; DooPlex cannot read them.
|
|
|
|
## What exists
|
|
|
|
| Where | Object | Purpose |
|
|
|---|---|---|
|
|
| ep0 | API token `root@pam!dooplex-sync`, ACL `DatastoreReader` on `/datastore/felhom-offsite` | read-only pull. **The only change on ep0.** |
|
|
| DooPlex | `felhom-ep0-pbs-tunnel.service` (systemd, runs as `kisfenyo`, `Restart=always`) | `ssh -N -L 127.0.0.1:18007:127.0.0.1:8007 root@ep0` — ep0's PBS listens on `wg0` only; DooPlex is not a WireGuard peer |
|
|
| DooPlex PBS | remote `ep0` (127.0.0.1:18007, ep0's cert fingerprint pinned) | the pull source |
|
|
| DooPlex PBS | datastore `ep0-copy` at `/mnt/5_hdd/backup/ep0-copy` | the copy |
|
|
| DooPlex PBS | sync job `ep0-felhom-offsite`, daily 05:00, `remove-vanished false` | the nightly pull (ep0's prune runs 03:30). **It never removes what ep0 removed** — a deletion on ep0 does not reach the copy |
|
|
| DooPlex PBS | verify job `verify-ep0-copy`, Saturdays 06:30 | reads the copy back |
|
|
| DooPlex PBS | notification target `felhom-operator` (SMTP via Resend → admin@felhom.eu) + matcher `felhom-operator-errors` (every error) | a failed pull or verify reaches the operator. Proven 2026-10-03 with a test mail |
|
|
|
|
Secrets, all out of git: the ep0 token secret in `/etc/proxmox-backup/remote.cfg` (root:backup 0640, base64 —
|
|
PBS's own format), the Resend key in `/etc/proxmox-backup/notifications-priv.cfg` (root:root 0600).
|
|
|
|
## Checks
|
|
|
|
```bash
|
|
systemctl is-active felhom-ep0-pbs-tunnel.service
|
|
curl -sk -o /dev/null -w '%{http_code}\n' https://127.0.0.1:18007/ # 200 = ep0's PBS reachable
|
|
sudo proxmox-backup-manager task list --limit 10 | grep -E 'syncjob|verif' # last runs
|
|
sudo find /mnt/5_hdd/backup/ep0-copy/ns -mindepth 4 -maxdepth 4 -type d | sort # snapshots per customer
|
|
```
|
|
|
|
## If ep0 is lost — restore a household's whole box from the DooPlex copy
|
|
|
|
The copy is a normal PBS datastore. Two routes, both needing the household's PBS `encryption-key` (escrowed,
|
|
recovered with the household's recovery code — the same as restoring from ep0):
|
|
|
|
1. **Point the box's host at DooPlex instead of ep0.** On the household's Proxmox host, add a PBS storage for
|
|
DooPlex's PBS (`ep0-copy`, namespace = the customer id) with the household's key, then restore the CT from it as
|
|
from ep0. DooPlex's PBS must be reachable from the host (it is not public today — the operator decides the route
|
|
at the time: a temporary tunnel, or a new endpoint).
|
|
2. **Rebuild the endpoint.** Provision a new ep0 (06 §5), then pull back: on the new ep0 add DooPlex as a remote
|
|
and run `proxmox-backup-manager pull <dooplex-remote> ep0-copy felhom-offsite`. Every box then reconnects as before.
|
|
|
|
Do **not** prune or garbage-collect the copy tighter than ep0's own retention. The copy has no prune job today and
|
|
grows with every nightly backup (R-828).
|
|
|
|
## Remove
|
|
|
|
`sudo proxmox-backup-manager sync-job remove ep0-felhom-offsite; … verify-job remove verify-ep0-copy;`
|
|
`sudo systemctl disable --now felhom-ep0-pbs-tunnel.service`; on ep0
|
|
`proxmox-backup-manager user delete-token root@pam dooplex-sync`. The datastore's bytes stay until removed by hand.
|