R-173 option A in force (hub DB nightly to ep0, restore-tested, alarmed); R-519 proven live on 9202 and closed; R-173/R-232 narrowed; R-882..R-885 opened (332 -> 335); runbook §3 tested; hubdb-check
gates / gates (push) Successful in 59s
gates / gates (push) Successful in 59s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
@@ -1,16 +1,16 @@
|
||||
# Runbook — the hub database in a backup that is NOT on DooPlex (R-173, R-232) — PROPOSED, needs the operator's go
|
||||
# Runbook — the hub database in a backup that is NOT on DooPlex (R-173, R-232) — IN FORCE since 2026-10-05
|
||||
|
||||
> **Status: PROPOSED 2026-10-05. Nothing here has been done.** DooPlex and ep0 are protected; every step below changes
|
||||
> one of them, so each waits for the operator's go (the decision is in `STATUS.md`). The readings this plan rests on:
|
||||
> `audits/hub-safety-2026-10-05/partC/readings.txt` (read only).
|
||||
> **Status: IN FORCE 2026-10-05** (option A, `09` decision 125). Steps 0–7 DONE, each marked below; §3 is a tested
|
||||
> procedure. Evidence: `audits/hub-db-offsite-2026-10-05/` (part A–D). The readings the plan rested on:
|
||||
> `audits/hub-safety-2026-10-05/partC/readings.txt`. **Corrections found while doing it are marked „Corrected".**
|
||||
|
||||
## 1. What is true today (measured 2026-10-05)
|
||||
|
||||
| Question | Answer |
|
||||
|---|---|
|
||||
| Where the hub database lives | `/data/hub.db` (+ `-wal`, `-shm`) in the hub pod, PVC `hub-data` (Longhorn, 1 Gi, replicas on DooPlex's `sdb1`). 357 MiB. |
|
||||
| Where the hub database lives | `/data/hub.db` (+ `-wal`, `-shm`) in the hub pod, PVC `hub-data` (Longhorn, **2 Gi since 2026-10-05**, was 1 Gi; replicas on DooPlex's `sdb1`). 357 MiB; a snapshot is 353 MiB. |
|
||||
| Is it in a backup? | **Yes, but only on DooPlex.** Longhorn's `backup-daily` (04:00) and `backup-weekly` (Sun 05:00), `retain=1`, write to `nfs://192.168.0.180:/mnt/5_hdd/backup/longhorn-pvc` — DooPlex's own `sda1`. Last: 2026-10-05 02:06 UTC, Completed. |
|
||||
| Why R-173 said "excluded" | The PVC carries `recurring-job-group.longhorn.io/default: disabled` (git, `manifests/hub.yaml`, commit `868e8465` of 2026-02-16, no reason given). The live Longhorn **Volume** carries `enabled` — set by hand at some point, so the backups run. **This is drift:** the PVC label is the source of truth (Longhorn syncs PVC → Volume), so the backups can stop at any sync without anyone seeing it. |
|
||||
| Why R-173 said "excluded" | The PVC carries `recurring-job-group.longhorn.io/default: disabled` (git, `manifests/hub.yaml`, commit `868e8465` of 2026-02-16, no reason given). The live Longhorn **Volume** carries `enabled` — set by hand at some point, so the backups run. **This is drift:** the PVC label is the source of truth (Longhorn syncs PVC → Volume), so the backups can stop at any sync without anyone seeing it. **Corrected 2026-10-05:** the Volume kept `enabled` through every sync since February while the PVC said `disabled` — Longhorn did NOT copy the PVC label down. The PVC now says `enabled` too (Step 1), so the two agree; which one Longhorn reads was not measured. |
|
||||
| What DooPlex's own backup covers | `dooplex-backup.timer` (03:19): k3s state, k8s Secrets (GPG files), Gitea mirrors, user data, PostgreSQL dumps — **all onto `sda1`, the same machine.** Nothing leaves DooPlex (`audits/RECON-dooplex-backup-2026-08-06.md`, R-232). |
|
||||
| What tells anyone a backup failed | **Nothing.** `NOTIFY_WEBHOOK_URL` is commented out, so `notify_failure` is a no-op. No Prometheus rule watches a Longhorn backup's success or age, nor `dooplex-backup.service`. |
|
||||
| What the database holds | Box→hub API keys, customer configs (incl. the owner passphrase), escrow custody blobs (opaque), the PBS-DR token values, the off-site sub-account passwords and — since hub v0.135.0 — the console passwords **sealed** under `OFFSITE_SECRET_KEY`. |
|
||||
@@ -22,7 +22,7 @@ ep0 is already off-site (Hetzner), already runs PBS, and DooPlex already reaches
|
||||
(127.0.0.1:18007). ep0's datastore is also pulled back to DooPlex nightly (`ep0-copy`), so the copy exists in two places,
|
||||
one of them off DooPlex. The copy is encrypted on DooPlex with a key ep0 never sees.
|
||||
|
||||
### Step 0 — the keys go off DooPlex first (operator, at the keyboard, 5 min)
|
||||
### Step 0 — the keys go off DooPlex first (operator, at the keyboard, 5 min) — DONE 2026-10-05 (operator: both saved)
|
||||
|
||||
Without these, every later step backs up something nobody can open after a DooPlex loss.
|
||||
|
||||
@@ -33,33 +33,65 @@ sudo kubectl -n felhom-system get secret offsite-secret-key -o jsonpath='{.data.
|
||||
sudo proxmox-backup-client key paperkey /etc/felhom-hub-backup/enc.key --output-format text
|
||||
```
|
||||
|
||||
### Step 1 — end the label drift (CC, a felhom.eu commit + ArgoCD sync; reversible)
|
||||
**The `data` field of that output is enough** (the key has no passphrase, `kdf: null`): a key file rebuilt from it alone —
|
||||
`{"kdf": null, "created": "<any RFC 3339 time>", "modified": "<same>", "data": "<saved value>"}` — restored and
|
||||
decrypted the copy on 2026-10-05 (`audits/hub-db-offsite-2026-10-05/partD/token-limits-and-paperkey.txt`). Note:
|
||||
`proxmox-backup-client key show` prints a fingerprint only when the file stores one, so it cannot check a rebuilt key —
|
||||
a restore can.
|
||||
|
||||
### Step 1 — end the label drift (CC, a felhom.eu commit + ArgoCD sync; reversible) — DONE 2026-10-05
|
||||
|
||||
`manifests/hub.yaml`: `recurring-job-group.longhorn.io/default: disabled` → `enabled`. Sync. Check:
|
||||
`sudo kubectl -n felhom-system get pvc hub-data -o jsonpath='{.metadata.labels}'` and the Volume label both read `enabled`.
|
||||
This keeps today's on-DooPlex copy alive; it is not the off-site copy.
|
||||
|
||||
### Step 2 — a write-only place on ep0 (on ep0, root; the operator's go for an ep0 change)
|
||||
Done with the volume growth to 2 Gi (two snapshots of 353 MiB do not fit in 1 Gi). **The online growth failed** — Longhorn's
|
||||
`instance-manager` (116 days up) called a host PID that no longer existed (`nsenter: cannot open /host/proc/196610/ns/mnt`);
|
||||
an offline growth was impossible too (the expansion holds its own attachment ticket). The operator approved restarting
|
||||
the instance-manager: all 77 DooPlex volumes back `attached/healthy` in 110 s, the volume grew, and one app (zipline, on
|
||||
`:latest`) came back on a newer release that refused its database — pinned to 4.7.0 (homelab-manifests). Evidence:
|
||||
`audits/hub-db-offsite-2026-10-05/partA/step1-*.txt`.
|
||||
|
||||
### Step 2 — a write-only place on ep0 (on ep0, root; the operator's go for an ep0 change) — DONE 2026-10-05
|
||||
|
||||
What was run (PBS 4.2.8 on ep0; the proposal's commands were wrong in three places, corrected here):
|
||||
|
||||
```bash
|
||||
proxmox-backup-manager user create dooplex-hub@pbs --comment "DooPlex pushes the hub DB (R-173)"
|
||||
proxmox-backup-manager user generate-token dooplex-hub@pbs push # the secret → a 0600 file on DooPlex, file → file
|
||||
# namespace for operator data, apart from the households' namespaces
|
||||
proxmox-backup-client namespace create operator --repository 'root@pam@127.0.0.1:8007:felhom-offsite'
|
||||
proxmox-backup-manager acl update /datastore/felhom-offsite/operator DatastoreBackup --auth-id 'dooplex-hub@pbs!push'
|
||||
# retention on ep0 (the server prunes; the pushing token cannot delete — DatastoreBackup has no Prune)
|
||||
proxmox-backup-manager prune-job create prune-operator-hubdb --store felhom-offsite --ns operator \
|
||||
--schedule 'daily 03:45' --keep-daily 14 --keep-weekly 8
|
||||
proxmox-backup-manager user create dooplex-hub@pbs --comment "..." # no password: cannot log in
|
||||
proxmox-backup-debug api create /admin/datastore/felhom-offsite/namespace --name operator
|
||||
# (the CLI crashes AFTER creating it, printing the result — 'not implemented'; check it exists, don't re-run)
|
||||
# two tokens, each secret file → file into a root 0600 file on DooPlex, never printed:
|
||||
ssh root@<ep0> "proxmox-backup-debug api create /access/users/dooplex-hub@pbs/token/push --output-format json" \
|
||||
| python3 -c '<print json["value"]>' | sudo sh -c 'umask 077; cat > /etc/felhom-hub-backup/token-push'
|
||||
# (same for token/restore → token-restore; `user generate-token` has no --output-format)
|
||||
P=/datastore/felhom-offsite/operator
|
||||
proxmox-backup-manager acl update $P DatastoreBackup --auth-id dooplex-hub@pbs # a token's rights are cut
|
||||
proxmox-backup-manager acl update $P DatastoreReader --auth-id dooplex-hub@pbs # down by its user's rights
|
||||
proxmox-backup-manager acl update $P DatastoreBackup --auth-id 'dooplex-hub@pbs!push'
|
||||
proxmox-backup-manager acl update $P DatastoreReader --auth-id 'dooplex-hub@pbs!restore'
|
||||
proxmox-backup-manager prune-job create prune-operator-hubdb --store felhom-offsite --ns operator --max-depth 0 \
|
||||
--schedule '03:45' --keep-daily 14 --keep-weekly 8 # 'daily 03:45' is not a PBS calendar event
|
||||
```
|
||||
|
||||
### Step 3 — a consistent snapshot of the live database (CC, a hub release)
|
||||
Measured (`partB/`, `partD/token-limits-and-paperkey.txt`): neither token can forget a snapshot or list the datastore
|
||||
root (the households); the restore token cannot write. **Corrected: the push token CAN restore its own copies** — PBS
|
||||
lets a backup's owner read it back (`DatastoreBackup` = `Datastore.Backup`, owner-scoped). It reaches only `operator`,
|
||||
and everything it can read is encrypted with a key ep0 never sees. The households' two prune jobs and the Sunday GC
|
||||
are unchanged (before/after in `partB/`).
|
||||
|
||||
### Step 3 — a consistent snapshot of the live database (CC, a hub release) — DONE, hub v0.136.0 (`05` §16.3)
|
||||
|
||||
`hub.db` is in WAL mode and is written every few seconds; copying the three files is not one point in time. The hub
|
||||
gets a nightly `VACUUM INTO '/data/snapshots/hub-<UTC date>.db'` (keeps 2, logs size and duration) — one SQLite
|
||||
statement, consistent by construction, WAL-aware. **Needs a hub release** (filed under R-173). No `sqlite3` exists in the
|
||||
hub image, so the copy must be made by the hub itself.
|
||||
statement, consistent by construction, WAL-aware. No `sqlite3` exists in the hub image, so the copy is made by the hub
|
||||
itself. **Corrected:** one snapshot took 44 s on the live volume (0.63 s on a local scratch copy), 353 MiB.
|
||||
|
||||
### Step 4 — the push (on DooPlex, root; `felhom-hub-db-backup.service` + `.timer` 02:30, CC writes, operator approves)
|
||||
### Step 4 — the push (on DooPlex, root; `felhom-hub-db-backup.service` + `.timer` 02:30) — DONE 2026-10-05
|
||||
|
||||
**The real script is `scripts/hub-db-backup/felhom-hub-db-backup`** (versioned, R-231; installed by `install.sh`;
|
||||
15 tests in `test_hub_db_backup.py`, run by hand — not in CI). It adds to the sketch below: it refuses a snapshot older
|
||||
than 26 h (the hub stopped snapshotting), a copy whose size differs from the pod's file, and a copy with no hosts; it
|
||||
encrypts with `--crypt-mode encrypt`. The sketch, as proposed:
|
||||
|
||||
```bash
|
||||
#!/bin/sh -eu
|
||||
@@ -77,7 +109,10 @@ echo "felhom_hub_db_backup_last_success_timestamp_seconds $(date +%s)" > /var/li
|
||||
&& mv /var/lib/node_exporter/textfile_collector/felhom_hub_db.prom.$$ /var/lib/node_exporter/textfile_collector/felhom_hub_db.prom
|
||||
```
|
||||
|
||||
### Step 5 — the restore test (weekly, Sun 04:30, same unit family)
|
||||
### Step 5 — the restore test (weekly, Sun 04:30, same unit family) — DONE 2026-10-05
|
||||
|
||||
**The real script is `scripts/hub-db-backup/felhom-hub-db-restore-test`**, with the READ-ONLY token. It also refuses a
|
||||
newest copy older than 50 h. The sketch, as proposed:
|
||||
|
||||
```bash
|
||||
T=$(mktemp -d); chmod 700 "$T"
|
||||
@@ -90,9 +125,12 @@ shred -u "$T/hub.db"*; rmdir "$T"
|
||||
echo "felhom_hub_db_restore_test_last_success_timestamp_seconds $(date +%s)" > …/felhom_hub_db_restore.prom # same tmp+mv
|
||||
```
|
||||
|
||||
The push token needs `DatastoreReader` on `operator` too for the restore (or a second, read-only token — cleaner).
|
||||
Done with a second, read-only token (`token-restore`).
|
||||
|
||||
### Step 6 — the alarm (homelab-manifests `prometheus-rules`, then `POST /-/reload` — the Prometheus there has no reloader)
|
||||
### Step 6 — the alarm (homelab-manifests `prometheus-rules`, then `POST /-/reload` — the Prometheus there has no reloader) — DONE 2026-10-05
|
||||
|
||||
In the `backup-freshness` group; `promtool test rules` proves it (`partC/bf_test.yml`, two red-proofs). The Prometheus
|
||||
Deployment is OutOfSync in ArgoCD for a reason unrelated to this; only the rules ConfigMap was synced.
|
||||
|
||||
```yaml
|
||||
- alert: HubDBBackupStale
|
||||
@@ -109,18 +147,46 @@ The push token needs `DatastoreReader` on `operator` too for the restore (or a s
|
||||
Both reach the existing `email-notifications` receiver. `absent()` makes "the script never ran" an alarm too — an empty
|
||||
log is not a success.
|
||||
|
||||
### Step 7 — prove it once (CC, with the operator's go)
|
||||
### Step 7 — prove it once (CC, with the operator's go) — DONE 2026-10-05 (`partD/`)
|
||||
|
||||
Run the unit by hand; read the snapshot on ep0 (`proxmox-backup-client snapshot list --ns operator`); run the restore
|
||||
test by hand; stop the timer for a day on purpose and see `HubDBBackupStale` mail arrive (positive observable), then
|
||||
start it again.
|
||||
|
||||
## 3. Bringing the hub back from this copy (the procedure the plan exists for)
|
||||
## 3. Bringing the hub back from this copy (the procedure the plan exists for) — TESTED 2026-10-05
|
||||
|
||||
1. A k3s with the `felhom` ArgoCD app, and **`Secret/offsite-secret-key` recreated with the SAME value** (Step 0 copy).
|
||||
2. Restore the newest snapshot (Step 5's first command, with the paper key), scale `deploy/hub` to 0, copy `hub.db` into
|
||||
the PVC (no `-wal`/`-shm` — the snapshot is a whole database), scale to 1. The log line
|
||||
`console passwords sealed at rest (0 legacy plaintext row(s) sealed now)` and a working reveal prove the key matches.
|
||||
Steps 1–3 were run on 2026-10-05 against the real copy on ep0 (`audits/hub-db-offsite-2026-10-05/partD/restore-procedure/drill.txt`):
|
||||
4 hosts, **4 of 4 console passwords opened with the saved seal key, 0 of 4 with a random key**. Steps 4–5 (into a live
|
||||
PVC) were NOT run — that needs the hub down; they are the ordinary scale-copy-scale.
|
||||
|
||||
What you need, all from the password manager: the seal key (`OFFSITE_SECRET_KEY`), the backup key's `data` field, and
|
||||
the read-only token (or ep0 root to mint a new one: Step 2).
|
||||
|
||||
1. **The backup key file.** On the machine doing the restore, as root, `umask 077`, write
|
||||
`{"kdf": null, "created": "2026-01-01T00:00:00+00:00", "modified": "2026-01-01T00:00:00+00:00", "data": "<data>"}`
|
||||
to `enc.key` (Step 0 note). A copy of DooPlex's `/etc/felhom-hub-backup/enc.key` works as is.
|
||||
2. **Restore the newest copy** (from any machine that reaches ep0's PBS on 8007 — DooPlex uses the tunnel 127.0.0.1:18007):
|
||||
```bash
|
||||
export PBS_PASSWORD_FILE=<token-restore file> PBS_FINGERPRINT=<ep0 cert fingerprint, /etc/felhom-hub-backup/env>
|
||||
R='dooplex-hub@pbs!restore@<ep0>:8007:felhom-offsite'
|
||||
proxmox-backup-client snapshot list host/dooplex-hub --ns operator --repository "$R" # pick the newest
|
||||
proxmox-backup-client restore host/dooplex-hub/<time> hubdb.pxar ./out --ns operator --keyfile enc.key --repository "$R"
|
||||
sqlite3 -readonly out/hub.db 'PRAGMA integrity_check' # must print: ok
|
||||
```
|
||||
If DooPlex's `ep0-copy` datastore survived, the same copy is there too (pulled nightly).
|
||||
3. **Prove the seal key matches BEFORE putting the copy in place** — on a COPY of `out/hub.db` (the check migrates it):
|
||||
```bash
|
||||
cd felhom.eu/hub && go build -o hubdb-check ./cmd/hubdb-check
|
||||
printf '%s' "<OFFSITE_SECRET_KEY>" > k; chmod 600 k # from the password manager — not on a command line in a shared shell
|
||||
./hubdb-check copy-of-hub.db k # want: hosts=N console_passwords_opened=N failed=0; exit 0
|
||||
```
|
||||
`failed>0` means the wrong seal key: the hub would start but could open no console password (`05` §16.2).
|
||||
4. **A k3s with the `felhom` ArgoCD app**, and `Secret/offsite-secret-key` recreated with the SAME value:
|
||||
`kubectl -n felhom-system create secret generic offsite-secret-key --from-file=OFFSITE_SECRET_KEY=k`.
|
||||
5. **Into the PVC:** scale `deploy/hub` to 0; put `out/hub.db` into the volume as `/data/hub.db` (a helper pod mounting
|
||||
`hub-data`; delete any `hub.db-wal`/`-shm` there — the snapshot is a whole database); scale to 1. The start-up log
|
||||
line `console passwords sealed at rest (0 legacy plaintext row(s) sealed now)` and one reveal on a host page confirm it.
|
||||
6. Shred `k`, `enc.key` copies and `out/` when done.
|
||||
|
||||
## 4. Option B (not my pick): restic to a dedicated Hetzner Storage Box sub-account
|
||||
|
||||
|
||||
Reference in New Issue
Block a user