Files
felhom.eu/documentation/runbooks/RUNBOOK-hub-db-offsite-backup.md

15 KiB
Raw Permalink Blame History

Runbook — the hub database in a backup that is NOT on DooPlex (R-173, R-232) — IN FORCE since 2026-10-05

Status: IN FORCE 2026-10-05 (option A, 09 decision 125). Steps 0–7 DONE, each marked below; §3 is a tested procedure. Evidence: audits/hub-db-offsite-2026-10-05/ (part A–D). The readings the plan rested on: audits/hub-safety-2026-10-05/partC/readings.txt. Corrections found while doing it are marked „Corrected".

1. What is true today (measured 2026-10-05)

Question Answer
Where the hub database lives /data/hub.db (+ -wal, -shm) in the hub pod, PVC hub-data (Longhorn, 2 Gi since 2026-10-05, was 1 Gi; replicas on DooPlex's sdb1). 357 MiB; a snapshot is 353 MiB.
Is it in a backup? Yes, but only on DooPlex. Longhorn's backup-daily (04:00) and backup-weekly (Sun 05:00), retain=1, write to nfs://192.168.0.180:/mnt/5_hdd/backup/longhorn-pvc — DooPlex's own sda1. Last: 2026-10-05 02:06 UTC, Completed.
Why R-173 said "excluded" The PVC carries recurring-job-group.longhorn.io/default: disabled (git, manifests/hub.yaml, commit 868e8465 of 2026-02-16, no reason given). The live Longhorn Volume carries enabled — set by hand at some point, so the backups run. This is drift: the PVC label is the source of truth (Longhorn syncs PVC → Volume), so the backups can stop at any sync without anyone seeing it. Corrected 2026-10-05: the Volume kept enabled through every sync since February while the PVC said disabled — Longhorn did NOT copy the PVC label down. The PVC now says enabled too (Step 1), so the two agree; which one Longhorn reads was not measured.
What DooPlex's own backup covers dooplex-backup.timer (03:19): k3s state, k8s Secrets (GPG files), Gitea mirrors, user data, PostgreSQL dumps — all onto sda1, the same machine. Nothing leaves DooPlex (audits/RECON-dooplex-backup-2026-08-06.md, R-232).
What tells anyone a backup failed Nothing. NOTIFY_WEBHOOK_URL is commented out, so notify_failure is a no-op. No Prometheus rule watches a Longhorn backup's success or age, nor dooplex-backup.service.
What the database holds Box→hub API keys, customer configs (incl. the owner passphrase), escrow custody blobs (opaque), the PBS-DR token values, the off-site sub-account passwords and — since hub v0.135.0 — the console passwords sealed under OFFSITE_SECRET_KEY.
What a copy is worth without the key The sealed columns (console passwords, off-site passwords) are useless without OFFSITE_SECRET_KEY. The key lives only in Secret/offsite-secret-key on DooPlex (and in the GPG secrets export on the same disk). A backup off DooPlex without a key off DooPlex restores a hub that cannot open any console password.

2. The plan (option A — my pick): a nightly, encrypted, consistent copy on ep0's PBS

ep0 is already off-site (Hetzner), already runs PBS, and DooPlex already reaches it through felhom-ep0-pbs-tunnel (127.0.0.1:18007). ep0's datastore is also pulled back to DooPlex nightly (ep0-copy), so the copy exists in two places, one of them off DooPlex. The copy is encrypted on DooPlex with a key ep0 never sees.

Step 0 — the keys go off DooPlex first (operator, at the keyboard, 5 min) — DONE 2026-10-05 (operator: both saved)

Without these, every later step backs up something nobody can open after a DooPlex loss.

# 1. the hub's seal key → the operator's password manager (never a file, never a chat)
sudo kubectl -n felhom-system get secret offsite-secret-key -o jsonpath='{.data.OFFSITE_SECRET_KEY}' | base64 -d; echo
# 2. (after Step 2) the backup encryption key's paper copy → the password manager
sudo proxmox-backup-client key paperkey /etc/felhom-hub-backup/enc.key --output-format text

The data field of that output is enough (the key has no passphrase, kdf: null): a key file rebuilt from it alone — {"kdf": null, "created": "<any RFC 3339 time>", "modified": "<same>", "data": "<saved value>"} — restored and decrypted the copy on 2026-10-05 (audits/hub-db-offsite-2026-10-05/partD/token-limits-and-paperkey.txt). Note: proxmox-backup-client key show prints a fingerprint only when the file stores one, so it cannot check a rebuilt key — a restore can.

Step 1 — end the label drift (CC, a felhom.eu commit + ArgoCD sync; reversible) — DONE 2026-10-05

manifests/hub.yaml: recurring-job-group.longhorn.io/default: disabled → enabled. Sync. Check: sudo kubectl -n felhom-system get pvc hub-data -o jsonpath='{.metadata.labels}' and the Volume label both read enabled. This keeps today's on-DooPlex copy alive; it is not the off-site copy.

Done with the volume growth to 2 Gi (two snapshots of 353 MiB do not fit in 1 Gi). The online growth failed — Longhorn's instance-manager (116 days up) called a host PID that no longer existed (nsenter: cannot open /host/proc/196610/ns/mnt); an offline growth was impossible too (the expansion holds its own attachment ticket). The operator approved restarting the instance-manager: all 77 DooPlex volumes back attached/healthy in 110 s, the volume grew, and one app (zipline, on :latest) came back on a newer release that refused its database — pinned to 4.7.0 (homelab-manifests). Evidence: audits/hub-db-offsite-2026-10-05/partA/step1-*.txt.

Step 2 — a write-only place on ep0 (on ep0, root; the operator's go for an ep0 change) — DONE 2026-10-05

What was run (PBS 4.2.8 on ep0; the proposal's commands were wrong in three places, corrected here):

proxmox-backup-manager user create dooplex-hub@pbs --comment "..."          # no password: cannot log in
proxmox-backup-debug api create /admin/datastore/felhom-offsite/namespace --name operator
#   (the CLI crashes AFTER creating it, printing the result — 'not implemented'; check it exists, don't re-run)
# two tokens, each secret file → file into a root 0600 file on DooPlex, never printed:
ssh root@<ep0> "proxmox-backup-debug api create /access/users/dooplex-hub@pbs/token/push --output-format json" \
  | python3 -c '<print json["value"]>' | sudo sh -c 'umask 077; cat > /etc/felhom-hub-backup/token-push'
#   (same for token/restore → token-restore; `user generate-token` has no --output-format)
P=/datastore/felhom-offsite/operator
proxmox-backup-manager acl update $P DatastoreBackup --auth-id dooplex-hub@pbs        # a token's rights are cut
proxmox-backup-manager acl update $P DatastoreReader --auth-id dooplex-hub@pbs        # down by its user's rights
proxmox-backup-manager acl update $P DatastoreBackup --auth-id 'dooplex-hub@pbs!push'
proxmox-backup-manager acl update $P DatastoreReader --auth-id 'dooplex-hub@pbs!restore'
proxmox-backup-manager prune-job create prune-operator-hubdb --store felhom-offsite --ns operator --max-depth 0 \
  --schedule '03:45' --keep-daily 14 --keep-weekly 8      # 'daily 03:45' is not a PBS calendar event

Measured (partB/, partD/token-limits-and-paperkey.txt): neither token can forget a snapshot or list the datastore root (the households); the restore token cannot write. Corrected: the push token CAN restore its own copies — PBS lets a backup's owner read it back (DatastoreBackup = Datastore.Backup, owner-scoped). It reaches only operator, and everything it can read is encrypted with a key ep0 never sees. The households' two prune jobs and the Sunday GC are unchanged (before/after in partB/).

Step 3 — a consistent snapshot of the live database (CC, a hub release) — DONE, hub v0.136.0 (05 §16.3)

hub.db is in WAL mode and is written every few seconds; copying the three files is not one point in time. The hub gets a nightly VACUUM INTO '/data/snapshots/hub-<UTC date>.db' (keeps 2, logs size and duration) — one SQLite statement, consistent by construction, WAL-aware. No sqlite3 exists in the hub image, so the copy is made by the hub itself. Corrected: one snapshot took 44 s on the live volume (0.63 s on a local scratch copy), 353 MiB.

Step 4 — the push (on DooPlex, root; felhom-hub-db-backup.service + .timer 02:30) — DONE 2026-10-05

The real script is scripts/hub-db-backup/felhom-hub-db-backup (versioned, R-231; installed by install.sh; 15 tests in test_hub_db_backup.py, run by hand — not in CI). It adds to the sketch below: it refuses a snapshot older than 26 h (the hub stopped snapshotting), a copy whose size differs from the pod's file, and a copy with no hosts; it encrypts with --crypt-mode encrypt. The sketch, as proposed:

#!/bin/sh -eu
# /usr/local/sbin/felhom-hub-db-backup — push the newest hub snapshot to ep0 (R-173). Root, 0755.
STAGE=/var/lib/felhom-hub-backup/stage; mkdir -p "$STAGE"; chmod 700 "$STAGE"
SNAP=$(kubectl -n felhom-system exec deploy/hub -- sh -c 'ls -1t /data/snapshots/hub-*.db | head -1')
kubectl -n felhom-system exec deploy/hub -- cat "$SNAP" > "$STAGE/hub.db"
sqlite3 -readonly "$STAGE/hub.db" 'PRAGMA integrity_check' | grep -qx ok          # never push a broken copy
export PBS_PASSWORD_FILE=/etc/felhom-hub-backup/token PBS_FINGERPRINT=<ep0 cert fingerprint, as in ep0-datastore-copy.md>
proxmox-backup-client backup hubdb.pxar:"$STAGE" --ns operator --backup-id dooplex-hub \
  --keyfile /etc/felhom-hub-backup/enc.key --repository 'dooplex-hub@pbs!push@127.0.0.1:18007:felhom-offsite'
shred -u "$STAGE/hub.db"
# the positive signal the alarm reads (written ONLY on success):
echo "felhom_hub_db_backup_last_success_timestamp_seconds $(date +%s)" > /var/lib/node_exporter/textfile_collector/felhom_hub_db.prom.$$ \
  && mv /var/lib/node_exporter/textfile_collector/felhom_hub_db.prom.$$ /var/lib/node_exporter/textfile_collector/felhom_hub_db.prom

Step 5 — the restore test (weekly, Sun 04:30, same unit family) — DONE 2026-10-05

The real script is scripts/hub-db-backup/felhom-hub-db-restore-test, with the READ-ONLY token. It also refuses a newest copy older than 50 h. The sketch, as proposed:

T=$(mktemp -d); chmod 700 "$T"
proxmox-backup-client restore "host/dooplex-hub/$(newest snapshot)" hubdb.pxar "$T" --ns operator \
  --keyfile /etc/felhom-hub-backup/enc.key --repository 'dooplex-hub@pbs!push@127.0.0.1:18007:felhom-offsite'
sqlite3 -readonly "$T/hub.db" 'PRAGMA integrity_check' | grep -qx ok
test "$(sqlite3 -readonly "$T/hub.db" 'SELECT COUNT(*) FROM hosts')" -gt 0
test "$(sqlite3 -readonly "$T/hub.db" "SELECT COUNT(*) FROM host_recovery WHERE secret NOT LIKE 'enc:v1:%'")" -eq 0
shred -u "$T/hub.db"*; rmdir "$T"
echo "felhom_hub_db_restore_test_last_success_timestamp_seconds $(date +%s)" > …/felhom_hub_db_restore.prom   # same tmp+mv

Done with a second, read-only token (token-restore).

Step 6 — the alarm (homelab-manifests prometheus-rules, then POST /-/reload — the Prometheus there has no reloader) — DONE 2026-10-05

In the backup-freshness group; promtool test rules proves it (partC/bf_test.yml, two red-proofs). The Prometheus Deployment is OutOfSync in ArgoCD for a reason unrelated to this; only the rules ConfigMap was synced.

- alert: HubDBBackupStale
  expr: time() - felhom_hub_db_backup_last_success_timestamp_seconds > 26*3600 or absent(felhom_hub_db_backup_last_success_timestamp_seconds)
  for: 30m
  labels: {severity: critical}
  annotations: {summary: "The hub database has not reached ep0 for 26 h (R-173)"}
- alert: HubDBRestoreTestStale
  expr: time() - felhom_hub_db_restore_test_last_success_timestamp_seconds > 8*24*3600 or absent(felhom_hub_db_restore_test_last_success_timestamp_seconds)
  for: 1h
  labels: {severity: warning}

Both reach the existing email-notifications receiver. absent() makes "the script never ran" an alarm too — an empty log is not a success.

Step 7 — prove it once (CC, with the operator's go) — DONE 2026-10-05 (partD/)

Run the unit by hand; read the snapshot on ep0 (proxmox-backup-client snapshot list --ns operator); run the restore test by hand; stop the timer for a day on purpose and see HubDBBackupStale mail arrive (positive observable), then start it again.

3. Bringing the hub back from this copy (the procedure the plan exists for) — TESTED 2026-10-05

Steps 1–3 were run on 2026-10-05 against the real copy on ep0 (audits/hub-db-offsite-2026-10-05/partD/restore-procedure/drill.txt): 4 hosts, 4 of 4 console passwords opened with the saved seal key, 0 of 4 with a random key. Steps 4–5 (into a live PVC) were NOT run — that needs the hub down; they are the ordinary scale-copy-scale.

What you need, all from the password manager: the seal key (OFFSITE_SECRET_KEY), the backup key's data field, and the read-only token (or ep0 root to mint a new one: Step 2).

  1. The backup key file. On the machine doing the restore, as root, umask 077, write {"kdf": null, "created": "2026-01-01T00:00:00+00:00", "modified": "2026-01-01T00:00:00+00:00", "data": "<data>"} to enc.key (Step 0 note). A copy of DooPlex's /etc/felhom-hub-backup/enc.key works as is.
  2. Restore the newest copy (from any machine that reaches ep0's PBS on 8007 — DooPlex uses the tunnel 127.0.0.1:18007):
    export PBS_PASSWORD_FILE=<token-restore file> PBS_FINGERPRINT=<ep0 cert fingerprint, /etc/felhom-hub-backup/env>
    R='dooplex-hub@pbs!restore@<ep0>:8007:felhom-offsite'
    proxmox-backup-client snapshot list host/dooplex-hub --ns operator --repository "$R"      # pick the newest
    proxmox-backup-client restore host/dooplex-hub/<time> hubdb.pxar ./out --ns operator --keyfile enc.key --repository "$R"
    sqlite3 -readonly out/hub.db 'PRAGMA integrity_check'                                      # must print: ok
    
    If DooPlex's ep0-copy datastore survived, the same copy is there too (pulled nightly).
  3. Prove the seal key matches BEFORE putting the copy in place — on a COPY of out/hub.db (the check migrates it):
    cd felhom.eu/hub && go build -o hubdb-check ./cmd/hubdb-check
    printf '%s' "<OFFSITE_SECRET_KEY>" > k; chmod 600 k        # from the password manager — not on a command line in a shared shell
    ./hubdb-check copy-of-hub.db k     # want: hosts=N console_passwords_opened=N failed=0; exit 0
    
    failed>0 means the wrong seal key: the hub would start but could open no console password (05 §16.2).
  4. A k3s with the felhom ArgoCD app, and Secret/offsite-secret-key recreated with the SAME value: kubectl -n felhom-system create secret generic offsite-secret-key --from-file=OFFSITE_SECRET_KEY=k.
  5. Into the PVC: scale deploy/hub to 0; put out/hub.db into the volume as /data/hub.db (a helper pod mounting hub-data; delete any hub.db-wal/-shm there — the snapshot is a whole database); scale to 1. The start-up log line console passwords sealed at rest (0 legacy plaintext row(s) sealed now) and one reveal on a host page confirm it.
  6. Shred k, enc.key copies and out/ when done.

4. Option B (not my pick): restic to a dedicated Hetzner Storage Box sub-account

Same Steps 0, 1, 3, 5, 6; the push is restic backup to a new sub-account with the R-820 append-only key pin. Costs a new sub-account and its own key custody; the Storage Box sub-account shell can rm (memory: storagebox-subaccount-shell) unless the pin is right. ep0 already has the server-side prune and the return copy, so A is less new machinery.