Files
felhom.eu/documentation/runbooks/RUNBOOK-hub-db-offsite-backup.md
T

196 lines
15 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Runbook — the hub database in a backup that is NOT on DooPlex (R-173, R-232) — IN FORCE since 2026-10-05
> **Status: IN FORCE 2026-10-05** (option A, `09` decision 125). Steps 0–7 DONE, each marked below; §3 is a tested
> procedure. Evidence: `audits/hub-db-offsite-2026-10-05/` (part A–D). The readings the plan rested on:
> `audits/hub-safety-2026-10-05/partC/readings.txt`. **Corrections found while doing it are marked „Corrected".**
## 1. What is true today (measured 2026-10-05)
| Question | Answer |
|---|---|
| Where the hub database lives | `/data/hub.db` (+ `-wal`, `-shm`) in the hub pod, PVC `hub-data` (Longhorn, **2 Gi since 2026-10-05**, was 1 Gi; replicas on DooPlex's `sdb1`). 357 MiB; a snapshot is 353 MiB. |
| Is it in a backup? | **Yes, but only on DooPlex.** Longhorn's `backup-daily` (04:00) and `backup-weekly` (Sun 05:00), `retain=1`, write to `nfs://192.168.0.180:/mnt/5_hdd/backup/longhorn-pvc` — DooPlex's own `sda1`. Last: 2026-10-05 02:06 UTC, Completed. |
| Why R-173 said "excluded" | The PVC carries `recurring-job-group.longhorn.io/default: disabled` (git, `manifests/hub.yaml`, commit `868e8465` of 2026-02-16, no reason given). The live Longhorn **Volume** carries `enabled` — set by hand at some point, so the backups run. **This is drift:** the PVC label is the source of truth (Longhorn syncs PVC → Volume), so the backups can stop at any sync without anyone seeing it. **Corrected 2026-10-05:** the Volume kept `enabled` through every sync since February while the PVC said `disabled` — Longhorn did NOT copy the PVC label down. The PVC now says `enabled` too (Step 1), so the two agree; which one Longhorn reads was not measured. |
| What DooPlex's own backup covers | `dooplex-backup.timer` (03:19): k3s state, k8s Secrets (GPG files), Gitea mirrors, user data, PostgreSQL dumps — **all onto `sda1`, the same machine.** Nothing leaves DooPlex (`audits/RECON-dooplex-backup-2026-08-06.md`, R-232). |
| What tells anyone a backup failed | **Nothing.** `NOTIFY_WEBHOOK_URL` is commented out, so `notify_failure` is a no-op. No Prometheus rule watches a Longhorn backup's success or age, nor `dooplex-backup.service`. |
| What the database holds | Box→hub API keys, customer configs (incl. the owner passphrase), escrow custody blobs (opaque), the PBS-DR token values, the off-site sub-account passwords and — since hub v0.135.0 — the console passwords **sealed** under `OFFSITE_SECRET_KEY`. |
| What a copy is worth without the key | The sealed columns (console passwords, off-site passwords) are useless without `OFFSITE_SECRET_KEY`. **The key lives only in `Secret/offsite-secret-key` on DooPlex** (and in the GPG secrets export on the same disk). A backup off DooPlex without a key off DooPlex restores a hub that cannot open any console password. |
## 2. The plan (option A — my pick): a nightly, encrypted, consistent copy on ep0's PBS
ep0 is already off-site (Hetzner), already runs PBS, and DooPlex already reaches it through `felhom-ep0-pbs-tunnel`
(127.0.0.1:18007). ep0's datastore is also pulled back to DooPlex nightly (`ep0-copy`), so the copy exists in two places,
one of them off DooPlex. The copy is encrypted on DooPlex with a key ep0 never sees.
### Step 0 — the keys go off DooPlex first (operator, at the keyboard, 5 min) — DONE 2026-10-05 (operator: both saved)
Without these, every later step backs up something nobody can open after a DooPlex loss.
```bash
# 1. the hub's seal key → the operator's password manager (never a file, never a chat)
sudo kubectl -n felhom-system get secret offsite-secret-key -o jsonpath='{.data.OFFSITE_SECRET_KEY}' | base64 -d; echo
# 2. (after Step 2) the backup encryption key's paper copy → the password manager
sudo proxmox-backup-client key paperkey /etc/felhom-hub-backup/enc.key --output-format text
```
**The `data` field of that output is enough** (the key has no passphrase, `kdf: null`): a key file rebuilt from it alone —
`{"kdf": null, "created": "<any RFC 3339 time>", "modified": "<same>", "data": "<saved value>"}` — restored and
decrypted the copy on 2026-10-05 (`audits/hub-db-offsite-2026-10-05/partD/token-limits-and-paperkey.txt`). Note:
`proxmox-backup-client key show` prints a fingerprint only when the file stores one, so it cannot check a rebuilt key —
a restore can.
### Step 1 — end the label drift (CC, a felhom.eu commit + ArgoCD sync; reversible) — DONE 2026-10-05
`manifests/hub.yaml`: `recurring-job-group.longhorn.io/default: disabled` → `enabled`. Sync. Check:
`sudo kubectl -n felhom-system get pvc hub-data -o jsonpath='{.metadata.labels}'` and the Volume label both read `enabled`.
This keeps today's on-DooPlex copy alive; it is not the off-site copy.
Done with the volume growth to 2 Gi (two snapshots of 353 MiB do not fit in 1 Gi). **The online growth failed** — Longhorn's
`instance-manager` (116 days up) called a host PID that no longer existed (`nsenter: cannot open /host/proc/196610/ns/mnt`);
an offline growth was impossible too (the expansion holds its own attachment ticket). The operator approved restarting
the instance-manager: all 77 DooPlex volumes back `attached/healthy` in 110 s, the volume grew, and one app (zipline, on
`:latest`) came back on a newer release that refused its database — pinned to 4.7.0 (homelab-manifests). Evidence:
`audits/hub-db-offsite-2026-10-05/partA/step1-*.txt`.
### Step 2 — a write-only place on ep0 (on ep0, root; the operator's go for an ep0 change) — DONE 2026-10-05
What was run (PBS 4.2.8 on ep0; the proposal's commands were wrong in three places, corrected here):
```bash
proxmox-backup-manager user create dooplex-hub@pbs --comment "..." # no password: cannot log in
proxmox-backup-debug api create /admin/datastore/felhom-offsite/namespace --name operator
# (the CLI crashes AFTER creating it, printing the result — 'not implemented'; check it exists, don't re-run)
# two tokens, each secret file → file into a root 0600 file on DooPlex, never printed:
ssh root@<ep0> "proxmox-backup-debug api create /access/users/dooplex-hub@pbs/token/push --output-format json" \
| python3 -c '<print json["value"]>' | sudo sh -c 'umask 077; cat > /etc/felhom-hub-backup/token-push'
# (same for token/restore → token-restore; `user generate-token` has no --output-format)
P=/datastore/felhom-offsite/operator
proxmox-backup-manager acl update $P DatastoreBackup --auth-id dooplex-hub@pbs # a token's rights are cut
proxmox-backup-manager acl update $P DatastoreReader --auth-id dooplex-hub@pbs # down by its user's rights
proxmox-backup-manager acl update $P DatastoreBackup --auth-id 'dooplex-hub@pbs!push'
proxmox-backup-manager acl update $P DatastoreReader --auth-id 'dooplex-hub@pbs!restore'
proxmox-backup-manager prune-job create prune-operator-hubdb --store felhom-offsite --ns operator --max-depth 0 \
--schedule '03:45' --keep-daily 14 --keep-weekly 8 # 'daily 03:45' is not a PBS calendar event
```
Measured (`partB/`, `partD/token-limits-and-paperkey.txt`): neither token can forget a snapshot or list the datastore
root (the households); the restore token cannot write. **Corrected: the push token CAN restore its own copies** — PBS
lets a backup's owner read it back (`DatastoreBackup` = `Datastore.Backup`, owner-scoped). It reaches only `operator`,
and everything it can read is encrypted with a key ep0 never sees. The households' two prune jobs and the Sunday GC
are unchanged (before/after in `partB/`).
### Step 3 — a consistent snapshot of the live database (CC, a hub release) — DONE, hub v0.136.0 (`05` §16.3)
`hub.db` is in WAL mode and is written every few seconds; copying the three files is not one point in time. The hub
gets a nightly `VACUUM INTO '/data/snapshots/hub-<UTC date>.db'` (keeps 2, logs size and duration) — one SQLite
statement, consistent by construction, WAL-aware. No `sqlite3` exists in the hub image, so the copy is made by the hub
itself. **Corrected:** one snapshot took 44 s on the live volume (0.63 s on a local scratch copy), 353 MiB.
### Step 4 — the push (on DooPlex, root; `felhom-hub-db-backup.service` + `.timer` 02:30) — DONE 2026-10-05
**The real script is `scripts/hub-db-backup/felhom-hub-db-backup`** (versioned, R-231; installed by `install.sh`;
15 tests in `test_hub_db_backup.py`, run by hand — not in CI). It adds to the sketch below: it refuses a snapshot older
than 26 h (the hub stopped snapshotting), a copy whose size differs from the pod's file, and a copy with no hosts; it
encrypts with `--crypt-mode encrypt`. The sketch, as proposed:
```bash
#!/bin/sh -eu
# /usr/local/sbin/felhom-hub-db-backup — push the newest hub snapshot to ep0 (R-173). Root, 0755.
STAGE=/var/lib/felhom-hub-backup/stage; mkdir -p "$STAGE"; chmod 700 "$STAGE"
SNAP=$(kubectl -n felhom-system exec deploy/hub -- sh -c 'ls -1t /data/snapshots/hub-*.db | head -1')
kubectl -n felhom-system exec deploy/hub -- cat "$SNAP" > "$STAGE/hub.db"
sqlite3 -readonly "$STAGE/hub.db" 'PRAGMA integrity_check' | grep -qx ok # never push a broken copy
export PBS_PASSWORD_FILE=/etc/felhom-hub-backup/token PBS_FINGERPRINT=<ep0 cert fingerprint, as in ep0-datastore-copy.md>
proxmox-backup-client backup hubdb.pxar:"$STAGE" --ns operator --backup-id dooplex-hub \
--keyfile /etc/felhom-hub-backup/enc.key --repository 'dooplex-hub@pbs!push@127.0.0.1:18007:felhom-offsite'
shred -u "$STAGE/hub.db"
# the positive signal the alarm reads (written ONLY on success):
echo "felhom_hub_db_backup_last_success_timestamp_seconds $(date +%s)" > /var/lib/node_exporter/textfile_collector/felhom_hub_db.prom.$$ \
&& mv /var/lib/node_exporter/textfile_collector/felhom_hub_db.prom.$$ /var/lib/node_exporter/textfile_collector/felhom_hub_db.prom
```
### Step 5 — the restore test (weekly, Sun 04:30, same unit family) — DONE 2026-10-05
**The real script is `scripts/hub-db-backup/felhom-hub-db-restore-test`**, with the READ-ONLY token. It also refuses a
newest copy older than 50 h. The sketch, as proposed:
```bash
T=$(mktemp -d); chmod 700 "$T"
proxmox-backup-client restore "host/dooplex-hub/$(newest snapshot)" hubdb.pxar "$T" --ns operator \
--keyfile /etc/felhom-hub-backup/enc.key --repository 'dooplex-hub@pbs!push@127.0.0.1:18007:felhom-offsite'
sqlite3 -readonly "$T/hub.db" 'PRAGMA integrity_check' | grep -qx ok
test "$(sqlite3 -readonly "$T/hub.db" 'SELECT COUNT(*) FROM hosts')" -gt 0
test "$(sqlite3 -readonly "$T/hub.db" "SELECT COUNT(*) FROM host_recovery WHERE secret NOT LIKE 'enc:v1:%'")" -eq 0
shred -u "$T/hub.db"*; rmdir "$T"
echo "felhom_hub_db_restore_test_last_success_timestamp_seconds $(date +%s)" > …/felhom_hub_db_restore.prom # same tmp+mv
```
Done with a second, read-only token (`token-restore`).
### Step 6 — the alarm (homelab-manifests `prometheus-rules`, then `POST /-/reload` — the Prometheus there has no reloader) — DONE 2026-10-05
In the `backup-freshness` group; `promtool test rules` proves it (`partC/bf_test.yml`, two red-proofs). The Prometheus
Deployment is OutOfSync in ArgoCD for a reason unrelated to this; only the rules ConfigMap was synced.
```yaml
- alert: HubDBBackupStale
expr: time() - felhom_hub_db_backup_last_success_timestamp_seconds > 26*3600 or absent(felhom_hub_db_backup_last_success_timestamp_seconds)
for: 30m
labels: {severity: critical}
annotations: {summary: "The hub database has not reached ep0 for 26 h (R-173)"}
- alert: HubDBRestoreTestStale
expr: time() - felhom_hub_db_restore_test_last_success_timestamp_seconds > 8*24*3600 or absent(felhom_hub_db_restore_test_last_success_timestamp_seconds)
for: 1h
labels: {severity: warning}
```
Both reach the existing `email-notifications` receiver. `absent()` makes "the script never ran" an alarm too — an empty
log is not a success.
### Step 7 — prove it once (CC, with the operator's go) — DONE 2026-10-05 (`partD/`)
Run the unit by hand; read the snapshot on ep0 (`proxmox-backup-client snapshot list --ns operator`); run the restore
test by hand; stop the timer for a day on purpose and see `HubDBBackupStale` mail arrive (positive observable), then
start it again.
## 3. Bringing the hub back from this copy (the procedure the plan exists for) — TESTED 2026-10-05
Steps 1–3 were run on 2026-10-05 against the real copy on ep0 (`audits/hub-db-offsite-2026-10-05/partD/restore-procedure/drill.txt`):
4 hosts, **4 of 4 console passwords opened with the saved seal key, 0 of 4 with a random key**. Steps 4–5 (into a live
PVC) were NOT run — that needs the hub down; they are the ordinary scale-copy-scale.
What you need, all from the password manager: the seal key (`OFFSITE_SECRET_KEY`), the backup key's `data` field, and
the read-only token (or ep0 root to mint a new one: Step 2).
1. **The backup key file.** On the machine doing the restore, as root, `umask 077`, write
`{"kdf": null, "created": "2026-01-01T00:00:00+00:00", "modified": "2026-01-01T00:00:00+00:00", "data": "<data>"}`
to `enc.key` (Step 0 note). A copy of DooPlex's `/etc/felhom-hub-backup/enc.key` works as is.
2. **Restore the newest copy** (from any machine that reaches ep0's PBS on 8007 — DooPlex uses the tunnel 127.0.0.1:18007):
```bash
export PBS_PASSWORD_FILE=<token-restore file> PBS_FINGERPRINT=<ep0 cert fingerprint, /etc/felhom-hub-backup/env>
R='dooplex-hub@pbs!restore@<ep0>:8007:felhom-offsite'
proxmox-backup-client snapshot list host/dooplex-hub --ns operator --repository "$R" # pick the newest
proxmox-backup-client restore host/dooplex-hub/<time> hubdb.pxar ./out --ns operator --keyfile enc.key --repository "$R"
sqlite3 -readonly out/hub.db 'PRAGMA integrity_check' # must print: ok
```
If DooPlex's `ep0-copy` datastore survived, the same copy is there too (pulled nightly).
3. **Prove the seal key matches BEFORE putting the copy in place** — on a COPY of `out/hub.db` (the check migrates it):
```bash
cd felhom.eu/hub && go build -o hubdb-check ./cmd/hubdb-check
printf '%s' "<OFFSITE_SECRET_KEY>" > k; chmod 600 k # from the password manager — not on a command line in a shared shell
./hubdb-check copy-of-hub.db k # want: hosts=N console_passwords_opened=N failed=0; exit 0
```
`failed>0` means the wrong seal key: the hub would start but could open no console password (`05` §16.2).
4. **A k3s with the `felhom` ArgoCD app**, and `Secret/offsite-secret-key` recreated with the SAME value:
`kubectl -n felhom-system create secret generic offsite-secret-key --from-file=OFFSITE_SECRET_KEY=k`.
5. **Into the PVC:** scale `deploy/hub` to 0; put `out/hub.db` into the volume as `/data/hub.db` (a helper pod mounting
`hub-data`; delete any `hub.db-wal`/`-shm` there — the snapshot is a whole database); scale to 1. The start-up log
line `console passwords sealed at rest (0 legacy plaintext row(s) sealed now)` and one reveal on a host page confirm it.
6. Shred `k`, `enc.key` copies and `out/` when done.
## 4. Option B (not my pick): restic to a dedicated Hetzner Storage Box sub-account
Same Steps 0, 1, 3, 5, 6; the push is `restic backup` to a new sub-account with the R-820 append-only key pin. Costs a
new sub-account and its own key custody; the Storage Box sub-account shell can `rm` (memory: storagebox-subaccount-shell)
unless the pin is right. ep0 already has the server-side prune and the return copy, so A is less new machinery.