hub-safety session: R-135/R-133/R-604/R-530/R-508/R-509/R-880 closed, R-861/R-173/R-518/R-519 narrowed, R-879/R-881 opened (336 → 332); 03 §3.1, 05 §16, golden 0.296.0, the hub-DB off-site plan, STATUS
gates / gates (push) Successful in 32s
gates / gates (push) Successful in 32s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
@@ -0,0 +1,129 @@
|
||||
# Runbook — the hub database in a backup that is NOT on DooPlex (R-173, R-232) — PROPOSED, needs the operator's go
|
||||
|
||||
> **Status: PROPOSED 2026-10-05. Nothing here has been done.** DooPlex and ep0 are protected; every step below changes
|
||||
> one of them, so each waits for the operator's go (the decision is in `STATUS.md`). The readings this plan rests on:
|
||||
> `audits/hub-safety-2026-10-05/partC/readings.txt` (read only).
|
||||
|
||||
## 1. What is true today (measured 2026-10-05)
|
||||
|
||||
| Question | Answer |
|
||||
|---|---|
|
||||
| Where the hub database lives | `/data/hub.db` (+ `-wal`, `-shm`) in the hub pod, PVC `hub-data` (Longhorn, 1 Gi, replicas on DooPlex's `sdb1`). 357 MiB. |
|
||||
| Is it in a backup? | **Yes, but only on DooPlex.** Longhorn's `backup-daily` (04:00) and `backup-weekly` (Sun 05:00), `retain=1`, write to `nfs://192.168.0.180:/mnt/5_hdd/backup/longhorn-pvc` — DooPlex's own `sda1`. Last: 2026-10-05 02:06 UTC, Completed. |
|
||||
| Why R-173 said "excluded" | The PVC carries `recurring-job-group.longhorn.io/default: disabled` (git, `manifests/hub.yaml`, commit `868e8465` of 2026-02-16, no reason given). The live Longhorn **Volume** carries `enabled` — set by hand at some point, so the backups run. **This is drift:** the PVC label is the source of truth (Longhorn syncs PVC → Volume), so the backups can stop at any sync without anyone seeing it. |
|
||||
| What DooPlex's own backup covers | `dooplex-backup.timer` (03:19): k3s state, k8s Secrets (GPG files), Gitea mirrors, user data, PostgreSQL dumps — **all onto `sda1`, the same machine.** Nothing leaves DooPlex (`audits/RECON-dooplex-backup-2026-08-06.md`, R-232). |
|
||||
| What tells anyone a backup failed | **Nothing.** `NOTIFY_WEBHOOK_URL` is commented out, so `notify_failure` is a no-op. No Prometheus rule watches a Longhorn backup's success or age, nor `dooplex-backup.service`. |
|
||||
| What the database holds | Box→hub API keys, customer configs (incl. the owner passphrase), escrow custody blobs (opaque), the PBS-DR token values, the off-site sub-account passwords and — since hub v0.135.0 — the console passwords **sealed** under `OFFSITE_SECRET_KEY`. |
|
||||
| What a copy is worth without the key | The sealed columns (console passwords, off-site passwords) are useless without `OFFSITE_SECRET_KEY`. **The key lives only in `Secret/offsite-secret-key` on DooPlex** (and in the GPG secrets export on the same disk). A backup off DooPlex without a key off DooPlex restores a hub that cannot open any console password. |
|
||||
|
||||
## 2. The plan (option A — my pick): a nightly, encrypted, consistent copy on ep0's PBS
|
||||
|
||||
ep0 is already off-site (Hetzner), already runs PBS, and DooPlex already reaches it through `felhom-ep0-pbs-tunnel`
|
||||
(127.0.0.1:18007). ep0's datastore is also pulled back to DooPlex nightly (`ep0-copy`), so the copy exists in two places,
|
||||
one of them off DooPlex. The copy is encrypted on DooPlex with a key ep0 never sees.
|
||||
|
||||
### Step 0 — the keys go off DooPlex first (operator, at the keyboard, 5 min)
|
||||
|
||||
Without these, every later step backs up something nobody can open after a DooPlex loss.
|
||||
|
||||
```bash
|
||||
# 1. the hub's seal key → the operator's password manager (never a file, never a chat)
|
||||
sudo kubectl -n felhom-system get secret offsite-secret-key -o jsonpath='{.data.OFFSITE_SECRET_KEY}' | base64 -d; echo
|
||||
# 2. (after Step 2) the backup encryption key's paper copy → the password manager
|
||||
sudo proxmox-backup-client key paperkey /etc/felhom-hub-backup/enc.key --output-format text
|
||||
```
|
||||
|
||||
### Step 1 — end the label drift (CC, a felhom.eu commit + ArgoCD sync; reversible)
|
||||
|
||||
`manifests/hub.yaml`: `recurring-job-group.longhorn.io/default: disabled` → `enabled`. Sync. Check:
|
||||
`sudo kubectl -n felhom-system get pvc hub-data -o jsonpath='{.metadata.labels}'` and the Volume label both read `enabled`.
|
||||
This keeps today's on-DooPlex copy alive; it is not the off-site copy.
|
||||
|
||||
### Step 2 — a write-only place on ep0 (on ep0, root; the operator's go for an ep0 change)
|
||||
|
||||
```bash
|
||||
proxmox-backup-manager user create dooplex-hub@pbs --comment "DooPlex pushes the hub DB (R-173)"
|
||||
proxmox-backup-manager user generate-token dooplex-hub@pbs push # the secret → a 0600 file on DooPlex, file → file
|
||||
# namespace for operator data, apart from the households' namespaces
|
||||
proxmox-backup-client namespace create operator --repository 'root@pam@127.0.0.1:8007:felhom-offsite'
|
||||
proxmox-backup-manager acl update /datastore/felhom-offsite/operator DatastoreBackup --auth-id 'dooplex-hub@pbs!push'
|
||||
# retention on ep0 (the server prunes; the pushing token cannot delete — DatastoreBackup has no Prune)
|
||||
proxmox-backup-manager prune-job create prune-operator-hubdb --store felhom-offsite --ns operator \
|
||||
--schedule 'daily 03:45' --keep-daily 14 --keep-weekly 8
|
||||
```
|
||||
|
||||
### Step 3 — a consistent snapshot of the live database (CC, a hub release)
|
||||
|
||||
`hub.db` is in WAL mode and is written every few seconds; copying the three files is not one point in time. The hub
|
||||
gets a nightly `VACUUM INTO '/data/snapshots/hub-<UTC date>.db'` (keeps 2, logs size and duration) — one SQLite
|
||||
statement, consistent by construction, WAL-aware. **Needs a hub release** (filed under R-173). No `sqlite3` exists in the
|
||||
hub image, so the copy must be made by the hub itself.
|
||||
|
||||
### Step 4 — the push (on DooPlex, root; `felhom-hub-db-backup.service` + `.timer` 02:30, CC writes, operator approves)
|
||||
|
||||
```bash
|
||||
#!/bin/sh -eu
|
||||
# /usr/local/sbin/felhom-hub-db-backup — push the newest hub snapshot to ep0 (R-173). Root, 0755.
|
||||
STAGE=/var/lib/felhom-hub-backup/stage; mkdir -p "$STAGE"; chmod 700 "$STAGE"
|
||||
SNAP=$(kubectl -n felhom-system exec deploy/hub -- sh -c 'ls -1t /data/snapshots/hub-*.db | head -1')
|
||||
kubectl -n felhom-system exec deploy/hub -- cat "$SNAP" > "$STAGE/hub.db"
|
||||
sqlite3 -readonly "$STAGE/hub.db" 'PRAGMA integrity_check' | grep -qx ok # never push a broken copy
|
||||
export PBS_PASSWORD_FILE=/etc/felhom-hub-backup/token PBS_FINGERPRINT=<ep0 cert fingerprint, as in ep0-datastore-copy.md>
|
||||
proxmox-backup-client backup hubdb.pxar:"$STAGE" --ns operator --backup-id dooplex-hub \
|
||||
--keyfile /etc/felhom-hub-backup/enc.key --repository 'dooplex-hub@pbs!push@127.0.0.1:18007:felhom-offsite'
|
||||
shred -u "$STAGE/hub.db"
|
||||
# the positive signal the alarm reads (written ONLY on success):
|
||||
echo "felhom_hub_db_backup_last_success_timestamp_seconds $(date +%s)" > /var/lib/node_exporter/textfile_collector/felhom_hub_db.prom.$$ \
|
||||
&& mv /var/lib/node_exporter/textfile_collector/felhom_hub_db.prom.$$ /var/lib/node_exporter/textfile_collector/felhom_hub_db.prom
|
||||
```
|
||||
|
||||
### Step 5 — the restore test (weekly, Sun 04:30, same unit family)
|
||||
|
||||
```bash
|
||||
T=$(mktemp -d); chmod 700 "$T"
|
||||
proxmox-backup-client restore "host/dooplex-hub/$(newest snapshot)" hubdb.pxar "$T" --ns operator \
|
||||
--keyfile /etc/felhom-hub-backup/enc.key --repository 'dooplex-hub@pbs!push@127.0.0.1:18007:felhom-offsite'
|
||||
sqlite3 -readonly "$T/hub.db" 'PRAGMA integrity_check' | grep -qx ok
|
||||
test "$(sqlite3 -readonly "$T/hub.db" 'SELECT COUNT(*) FROM hosts')" -gt 0
|
||||
test "$(sqlite3 -readonly "$T/hub.db" "SELECT COUNT(*) FROM host_recovery WHERE secret NOT LIKE 'enc:v1:%'")" -eq 0
|
||||
shred -u "$T/hub.db"*; rmdir "$T"
|
||||
echo "felhom_hub_db_restore_test_last_success_timestamp_seconds $(date +%s)" > …/felhom_hub_db_restore.prom # same tmp+mv
|
||||
```
|
||||
|
||||
The push token needs `DatastoreReader` on `operator` too for the restore (or a second, read-only token — cleaner).
|
||||
|
||||
### Step 6 — the alarm (homelab-manifests `prometheus-rules`, then `POST /-/reload` — the Prometheus there has no reloader)
|
||||
|
||||
```yaml
|
||||
- alert: HubDBBackupStale
|
||||
expr: time() - felhom_hub_db_backup_last_success_timestamp_seconds > 26*3600 or absent(felhom_hub_db_backup_last_success_timestamp_seconds)
|
||||
for: 30m
|
||||
labels: {severity: critical}
|
||||
annotations: {summary: "The hub database has not reached ep0 for 26 h (R-173)"}
|
||||
- alert: HubDBRestoreTestStale
|
||||
expr: time() - felhom_hub_db_restore_test_last_success_timestamp_seconds > 8*24*3600 or absent(felhom_hub_db_restore_test_last_success_timestamp_seconds)
|
||||
for: 1h
|
||||
labels: {severity: warning}
|
||||
```
|
||||
|
||||
Both reach the existing `email-notifications` receiver. `absent()` makes "the script never ran" an alarm too — an empty
|
||||
log is not a success.
|
||||
|
||||
### Step 7 — prove it once (CC, with the operator's go)
|
||||
|
||||
Run the unit by hand; read the snapshot on ep0 (`proxmox-backup-client snapshot list --ns operator`); run the restore
|
||||
test by hand; stop the timer for a day on purpose and see `HubDBBackupStale` mail arrive (positive observable), then
|
||||
start it again.
|
||||
|
||||
## 3. Bringing the hub back from this copy (the procedure the plan exists for)
|
||||
|
||||
1. A k3s with the `felhom` ArgoCD app, and **`Secret/offsite-secret-key` recreated with the SAME value** (Step 0 copy).
|
||||
2. Restore the newest snapshot (Step 5's first command, with the paper key), scale `deploy/hub` to 0, copy `hub.db` into
|
||||
the PVC (no `-wal`/`-shm` — the snapshot is a whole database), scale to 1. The log line
|
||||
`console passwords sealed at rest (0 legacy plaintext row(s) sealed now)` and a working reveal prove the key matches.
|
||||
|
||||
## 4. Option B (not my pick): restic to a dedicated Hetzner Storage Box sub-account
|
||||
|
||||
Same Steps 0, 1, 3, 5, 6; the push is `restic backup` to a new sub-account with the R-820 append-only key pin. Costs a
|
||||
new sub-account and its own key custody; the Storage Box sub-account shell can `rm` (memory: storagebox-subaccount-shell)
|
||||
unless the pin is right. ep0 already has the server-side prune and the return copy, so A is less new machinery.
|
||||
@@ -31,6 +31,24 @@ sign; no bundle may add, remove or change them. A box that has no signers file g
|
||||
4. **Undo** = send the previous release's bundle the same way. The previous copies also stay on the box in
|
||||
`/var/lib/felhom-os-apply/bundle-prev/<time>-before-<version>/` (the last 3).
|
||||
|
||||
## A release whose bundle ADDS a path — the step bundle (R-880, decision 124)
|
||||
|
||||
The box's INSTALLED `felhom-os-apply` checks every path of an incoming bundle against its OWN table (R16). So when a
|
||||
release adds a path (agent v0.146.1 added four), every box on an older bundle refuses it. Send a step first:
|
||||
|
||||
```bash
|
||||
# the bundle the boxes run now — check its sha against the hub's Root files / config-bundle record
|
||||
curl -fsS -o base.json https://gitea.dooplex.hu/api/packages/admin/generic/felhom-agent/<old>/felhom-config-bundle.json
|
||||
python3 felhom-agent/scripts/build-step-bundle.py base.json <new>-step1 step.json # prints the step sha
|
||||
curl -u admin:<token from a file> -X PUT --upload-file step.json \
|
||||
https://gitea.dooplex.hu/api/packages/admin/generic/felhom-agent/<new>-step1/felhom-config-bundle.json # 201
|
||||
# per box: agent_update <new> → agent_config_update <new>-step1 (step sha) → agent_config_update <new> (release sha)
|
||||
```
|
||||
|
||||
The step is the old bundle with ONLY `felhom-os-apply` replaced, so the old wrapper accepts it (`written=1 same=20`);
|
||||
the new wrapper then accepts the release's bundle. Done this way on demo-hp, demo-felhom and Tester 1 on 2026-10-05
|
||||
(`audits/hub-safety-2026-10-05/partH/`). Keep the step package while any box may still be on the old bundle.
|
||||
|
||||
## A box from before agent v0.143.0 — the ONE by-hand step (bootstrap)
|
||||
|
||||
Such a box's `felhom-os-apply` has no bundle mode, and no signed job can write a root file there (that gap IS
|
||||
|
||||
Reference in New Issue
Block a user