R-173 option A in force (hub DB nightly to ep0, restore-tested, alarmed); R-519 proven live on 9202 and closed; R-173/R-232 narrowed; R-882..R-885 opened (332 -> 335); runbook §3 tested; hubdb-check
gates / gates (push) Successful in 59s
gates / gates (push) Successful in 59s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
@@ -235,6 +235,7 @@ likewise silent. Evidence: `audits/DRILL-r361-2026-08-22/evidence/06-part3-decis
|
||||
| Secrets hygiene: bearer in k8s Secret, no secrets in git, single-quote credential store | hub v0.53, conventions | **IMPLEMENTED** | 07-13 closing bundle | |
|
||||
| Operator login password changeable from UI | hub v0.54 | **IMPLEMENTED** | 07-13 | |
|
||||
| **The operator sees boxes left behind: per-customer controller floors with their age (and which a global raise cannot move), each box's agent against the vouched one, a 7-day "agent behind" alarm and a "floor raise skipped boxes" mail** | hub **v0.135.0** (R-604, R-530) | **IMPLEMENTED — the page PROVEN-LIVE, the alarm and the mail unit-proven** | `audits/hub-safety-2026-10-05/partD/` (live System page: three per-customer floors, Tester 2 `0.142.0 → 0.145.0`); `osupdates/r530_agent_alarm_test.go`, `web/r604_floor_held_back_test.go` | the mail fires only on a GLOBAL raise below an override — not exercised live |
|
||||
| **The hub database survives the loss of DooPlex: a nightly consistent copy, encrypted, on ep0; restore-tested weekly; an alarm when either stops** | hub **v0.136.0** (R-173), `scripts/hub-db-backup/`, homelab-manifests rules | **PROVEN-LIVE (2026-10-05)** — first push, ep0 listing, restore test (4 hosts, 4 sealed, 0 readable), token limits, a key rebuilt from the paper copy decrypts, the saved seal key opens 4/4 console passwords in the restored copy; the alarm by `promtool` rule test + red-proofs | `audits/hub-db-offsite-2026-10-05/`; `runbooks/RUNBOOK-hub-db-offsite-backup.md` | Runbook §3 steps 4–5 (into a live PVC) not exercised (R-173) |
|
||||
| **The operator surface refuses a cross-site form post on BOTH login paths (session token; Basic auth + `X-Felhom-Operator`)** | hub **v0.135.0** (R-135) | **PROVEN-LIVE (2026-10-05)** | `audits/hub-safety-2026-10-05/partA/live.txt` (Basic, no header → 403 even on an unknown path; with the header → passes; header without credentials → 401); `web/r135_csrf_test.go` (39 paths) | |
|
||||
| Box operating-system security updates (Proxmox host, guest Debian, Docker engine) | agent v0.143.0, hub v0.133.0 | **PARTIAL — the GUEST and HOST Debian fast lanes and the DOCKER engine slow lane are PROVEN-LIVE (2026-10-04), with the System page, the fleet view and the alarms; the KERNEL lane is MISSING** | Guest: `audits/os-guest-lane-2026-10-04/`. Host + fleet + alarms: `audits/os-host-lane-2026-10-04/`. Docker + System page + crash guard: `audits/os-docker-crash-2026-10-04/` — live-restore on with the same container ids on every box; Docker 29.8.2 on both demo boxes; operator-approved Docker release; a signed undo and a signed ring-1 step; a replay refused; the crash guard restarted demo-hp twice and kept it off the third time. Design `architecture/11-os-updates.md` §5.8, §5.9, §8 | **No automatic undo** (guest: last night's backup; host: by-hand runbook; Docker: a signed undo job); existing boxes get root-owned files by the signed config bundle since 2026-10-04 (R-840 CLOSED; a box from before agent 0.143.0 needs one by-hand bootstrap — Tester 2: R-862; `audits/r840-config-bundle-2026-10-04/`); test approvals now end with the test (R-859); the agent's sudoers is root-equivalent (R-861); the kernel lane (R-836); facts reach the hub late after a boot (R-853). **2026-10-05 (agent v0.144.1):** R8 measures the real download (R-865); a killed pass still reports (R-868, live); the debug pass runs with the hub away (R-866, live); **a power cut mid-update was proven by day on demo-hp — the box came back by itself in 37 s, but the next pass fails until `dpkg --configure -a` is run by hand (R-876, P2, open)** — `audits/night-fixes-2026-10-05/`. **2026-10-05 afternoon (agent v0.145.0): R-876 FIXED and proven live — after a second crash mid-unpack the next pass repaired dpkg by itself (`REPAIR … journal=1`) and finished** — `audits/catchup-2026-10-05/partD/` |
|
||||
| **An ENGLISH-SPEAKING household's first hour: download, install, pair, bind, claim, two apps** | controller **v0.259.0** + hub **v0.119.0** + ISO 1.29.0 + the whole catalog | **PROVEN-LIVE on 0.258.0 with one blocker; THE BLOCKER IS FIXED AND PROVEN, THE WALK IS NOT REPEATED** | `audits/DRILL-first-hour-en-0258-2026-09-20.md` — a fresh install 2026-09-20, one intervention (R-494), stop rule not reached. Then `audits/i18n-closing-2026-09-21/live/` — the three blockers fixed and each proven on a live box or in the operator's inbox: the claim page answers English through the real cookie path; the Backup page's tier names follow the language; and the setup mail carries **four plain-ASCII English words** where the drill's carried `képző-szkítia-ásatás`, one day apart in the same inbox. | **R-596, R-597 and R-598 are CLOSED.** What this row still does NOT claim: **the fixed journey has not been walked end to end by a stranger on a fresh install.** Three fixes proven at the endpoint are not an hour proven by a person, and this project's own rule is that fixes are not a journey (see the recovery-journey row). **Also not walked:** the recovery code (needs ep0), backup/restore/remove/power-cut (proven 2026-09-14), and the two Backup-page *warnings* themselves — guest 9201 is healthy and a healthy box renders none, so they are covered by handler render tests, not live. **Verdict: nothing known now stands between an English-speaking tester and their box — and that is a different sentence from "the walk passed".** |
|
||||
|
||||
@@ -140,8 +140,9 @@ mert a doboz vagy a vezérlő újraindult…"), kept until a run ends with every
|
||||
before this, both pages said nothing and the synthesised „Utolsó adatbázis mentés … OK" was read off the fresh `.sql`
|
||||
the cut run left beside last night's tars; that line now reads failed after a cut. Each restore point's time was
|
||||
already its OLDEST part (the data block, v0.275.0) — so a torn unit is dated by its stale tars, never by its new dump.
|
||||
*Live: unit-proven and red-proved; the live cut on 9202 was refused by the permission check and waits for the
|
||||
operator (R-519 narrowed).*
|
||||
*Live: PROVEN 2026-10-05 on 9202 (operator ruling 126): a run cut by a controller restart between bookstack's config
|
||||
dump and its database volume → both pages carry the notice, the restore point reads the older volume's time, the next
|
||||
complete run clears it (`audits/hub-db-offsite-2026-10-05/partD/r519/`; R-519 closed).*
|
||||
|
||||
### Lane 2 — the operator: guest and host recovery
|
||||
|
||||
|
||||
@@ -0,0 +1,26 @@
|
||||
## promtool, in pod/prometheus-55b675779d-t8c74, 2026-10-05T13:42:17Z
|
||||
### green
|
||||
SUCCESS
|
||||
|
||||
rc=0
|
||||
### red: threshold 26h -> 260h
|
||||
FAILED:
|
||||
alertname: HubDBBackupStale, time: 1d2h40m,
|
||||
exp:[
|
||||
0:
|
||||
Labels:{alertname="HubDBBackupStale", component="backup", instance="dooplex", severity="critical"}
|
||||
Annotations:{description="No successful push of the hub DB snapshot to ep0 for >26h (daily at 02:30). Check `journalctl -u felhom-hub-db-backup.service`, the tunnel `systemctl status felhom-ep0-pbs-tunnel`, and that the hub logs `db snapshot written` at 02:00.", summary="The Felhom hub database has not reached ep0 for 26 h (R-173)"}
|
||||
],
|
||||
got:[]
|
||||
|
||||
|
||||
command terminated with exit code 1
|
||||
rc=0
|
||||
### red: absent() removed
|
||||
SUCCESS
|
||||
|
||||
### red (re-run): absent() removed from HubDBBackupStale — first attempt above did NOT apply (indent mismatch)
|
||||
FAILED:
|
||||
alertname: HubDBBackupStale, time: 40m,
|
||||
Labels:{alertname="HubDBBackupStale", component="backup", severity="critical"}
|
||||
got:[]
|
||||
@@ -0,0 +1,52 @@
|
||||
rule_files: [bf.yml]
|
||||
evaluation_interval: 1m
|
||||
tests:
|
||||
# 1. the push stops: last success at t=0, never again → fires once 26 h + 30 m have passed
|
||||
- interval: 5m
|
||||
input_series:
|
||||
- series: 'felhom_hub_db_backup_last_success_timestamp_seconds{instance="dooplex"}'
|
||||
values: '0+0x400'
|
||||
- series: 'felhom_hub_db_restore_test_last_success_timestamp_seconds{instance="dooplex"}'
|
||||
values: '0+0x400'
|
||||
alert_rule_test:
|
||||
- eval_time: 26h
|
||||
alertname: HubDBBackupStale
|
||||
exp_alerts: []
|
||||
- eval_time: 26h40m
|
||||
alertname: HubDBBackupStale
|
||||
exp_alerts:
|
||||
- exp_labels: {severity: critical, component: backup, instance: dooplex}
|
||||
exp_annotations:
|
||||
summary: "The Felhom hub database has not reached ep0 for 26 h (R-173)"
|
||||
description: "No successful push of the hub DB snapshot to ep0 for >26h (daily at 02:30). Check `journalctl -u felhom-hub-db-backup.service`, the tunnel `systemctl status felhom-ep0-pbs-tunnel`, and that the hub logs `db snapshot written` at 02:00."
|
||||
# 2. healthy: the push succeeds every 24 h → never fires
|
||||
- interval: 1h
|
||||
input_series:
|
||||
- series: 'felhom_hub_db_backup_last_success_timestamp_seconds{instance="dooplex"}'
|
||||
values: '0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 86400 86400 86400 86400 86400 86400 86400 86400 86400 86400 86400 86400 86400 86400 86400 86400 86400 86400 86400 86400 86400 86400 86400 86400 172800 172800 172800'
|
||||
- series: 'felhom_hub_db_restore_test_last_success_timestamp_seconds{instance="dooplex"}'
|
||||
values: '0+0x50'
|
||||
alert_rule_test:
|
||||
- eval_time: 50h
|
||||
alertname: HubDBBackupStale
|
||||
exp_alerts: []
|
||||
# 3. the metric never existed (script never ran) → fires on absent()
|
||||
- interval: 5m
|
||||
input_series:
|
||||
- series: 'up{job="node"}'
|
||||
values: '1+0x30'
|
||||
alert_rule_test:
|
||||
- eval_time: 40m
|
||||
alertname: HubDBBackupStale
|
||||
exp_alerts:
|
||||
- exp_labels: {severity: critical, component: backup}
|
||||
exp_annotations:
|
||||
summary: "The Felhom hub database has not reached ep0 for 26 h (R-173)"
|
||||
description: "No successful push of the hub DB snapshot to ep0 for >26h (daily at 02:30). Check `journalctl -u felhom-hub-db-backup.service`, the tunnel `systemctl status felhom-ep0-pbs-tunnel`, and that the hub logs `db snapshot written` at 02:00."
|
||||
- eval_time: 2h
|
||||
alertname: HubDBRestoreTestStale
|
||||
exp_alerts:
|
||||
- exp_labels: {severity: warning, component: backup}
|
||||
exp_annotations:
|
||||
summary: "The hub database copy on ep0 has not passed a restore test for 8 days (R-173)"
|
||||
description: "The weekly restore test (Sun 04:30) has not succeeded for >8 days. Check `journalctl -u felhom-hub-db-restore-test.service`."
|
||||
@@ -0,0 +1,5 @@
|
||||
## alarm drill start 2026-10-05T13:51:46Z: success file moved aside (absent case)
|
||||
backup_freshness.prom
|
||||
fan_metrics.prom
|
||||
felhom_hub_db_restore.prom
|
||||
node_housekeeping.prom
|
||||
@@ -0,0 +1,27 @@
|
||||
## manual push via unit, 2026-10-05T13:50:30Z
|
||||
Result=success
|
||||
ExecMainStatus=0
|
||||
2026-10-05T15:50:08+02:00 dooplex systemd[1]: Starting felhom-hub-db-backup.service - Felhom: push the hub DB snapshot to ep0 (R-173)...
|
||||
2026-10-05T15:50:09+02:00 dooplex felhom-hub-db-backup[3036511]: felhom-hub-db-backup: snapshot hub-20261005T123037Z.db, 79 min old
|
||||
2026-10-05T15:50:18+02:00 dooplex felhom-hub-db-backup[3036511]: felhom-hub-db-backup: checked: 369807360 bytes, integrity ok, 4 host(s)
|
||||
2026-10-05T15:50:18+02:00 dooplex felhom-hub-db-backup[3037790]: Starting backup: [operator]:host/dooplex-hub/2026-10-05T13:50:18Z
|
||||
2026-10-05T15:50:18+02:00 dooplex felhom-hub-db-backup[3037790]: Client name: dooplex
|
||||
2026-10-05T15:50:18+02:00 dooplex felhom-hub-db-backup[3037790]: Starting backup protocol: Mon Oct 5 15:50:18 2026
|
||||
2026-10-05T15:50:18+02:00 dooplex felhom-hub-db-backup[3037790]: Using encryption key from '/etc/felhom-hub-backup/enc.key'..
|
||||
2026-10-05T15:50:18+02:00 dooplex felhom-hub-db-backup[3037790]: Encryption key fingerprint: b2:19:bf:36:3b:97:3d:6c
|
||||
2026-10-05T15:50:18+02:00 dooplex felhom-hub-db-backup[3037790]: No previous manifest available.
|
||||
2026-10-05T15:50:18+02:00 dooplex felhom-hub-db-backup[3037790]: Upload directory '/var/lib/felhom-hub-backup/stage' to 'dooplex-hub@pbs!push@127.0.0.1:18007:felhom-offsite' as hubdb.pxar.didx
|
||||
2026-10-05T15:50:25+02:00 dooplex felhom-hub-db-backup[3037790]: hubdb.pxar: had to backup 352.676 MiB of 352.676 MiB (compressed 17.242 MiB) in 6.58 s (average 53.604 MiB/s)
|
||||
2026-10-05T15:50:25+02:00 dooplex felhom-hub-db-backup[3037790]: Uploaded backup catalog (56 B)
|
||||
2026-10-05T15:50:25+02:00 dooplex felhom-hub-db-backup[3037790]: Duration: 7.19s
|
||||
2026-10-05T15:50:25+02:00 dooplex felhom-hub-db-backup[3037790]: End Time: Mon Oct 5 15:50:25 2026
|
||||
2026-10-05T15:50:25+02:00 dooplex felhom-hub-db-backup[3036511]: felhom-hub-db-backup: pushed hub-20261005T123037Z.db to ep0 (ns operator) in 7 s
|
||||
2026-10-05T15:50:25+02:00 dooplex felhom-hub-db-backup[3036511]: felhom-hub-db-backup: success signal written
|
||||
2026-10-05T15:50:30+02:00 dooplex systemd[1]: felhom-hub-db-backup.service: Deactivated successfully.
|
||||
2026-10-05T15:50:30+02:00 dooplex systemd[1]: Finished felhom-hub-db-backup.service - Felhom: push the hub DB snapshot to ep0 (R-173).
|
||||
2026-10-05T15:50:30+02:00 dooplex systemd[1]: felhom-hub-db-backup.service: Consumed 9.566s CPU time, 584.8M memory peak.
|
||||
## textfile
|
||||
# HELP felhom_hub_db_backup_last_success_timestamp_seconds Last successful push of the hub DB snapshot to ep0 (R-173).
|
||||
# TYPE felhom_hub_db_backup_last_success_timestamp_seconds gauge
|
||||
felhom_hub_db_backup_last_success_timestamp_seconds 1791208225
|
||||
felhom_hub_db_backup_last_success_bytes 369807360
|
||||
@@ -0,0 +1 @@
|
||||
gitea.dooplex.hu/admin/felhom-controller:0.296.0 Up 6 seconds (healthy)
|
||||
@@ -0,0 +1,3 @@
|
||||
## deploy bookstack 2026-10-05T13:53:26Z (fields: DOMAIN APP_KEY DB_PASSWORD ADMIN_PASSWORD; SUBDOMAIN = catalog default)
|
||||
|
||||
HTTP 202
|
||||
@@ -0,0 +1,36 @@
|
||||
## run 1 (baseline) start 2026-10-05T14:04:36Z
|
||||
|
||||
HTTP 200
|
||||
{"ok":true,"data":{"enabled":true,"running":true}}
|
||||
|
||||
HTTP 200
|
||||
|
||||
## run 1 end 2026-10-05T14:05:17Z: {"ok":true,"data":{"db_dump":{"count":2,"duration":"32.784678377s","last_run":"2026-10-05T14:05:12.109790785Z","success":true},"enabled":true,"running":false}}
|
||||
2026/10/05 14:04:44 data_versions.go:131: [DEBUG] [backup] bookstack: stamped volume-dumps/bookstack_bookstack_config.tar (70144 B) with pins [lscr.io/linuxserver/bookstack:26.09.1@sha256:99cd1f5707c1911afad213adec5c9739763b76f843d1477142231834ecdcb6f7 mariadb:12.3@sha256:805c8e104bd563d5bfa24fadd3f31cd419ea859cb5277f32b5dbf2db714f9ed1]
|
||||
2026/10/05 14:04:44 backup.go:848: [DEBUG] [backup] Dumping volume bookstack_bookstack_db_data for bookstack
|
||||
2026/10/05 14:04:45 backup.go:873: [INFO] [backup] Volume dump: bookstack/bookstack_bookstack_db_data → 153.4 MB
|
||||
2026/10/05 14:04:45 data_versions.go:131: [DEBUG] [backup] bookstack: stamped volume-dumps/bookstack_bookstack_db_data.tar (160868864 B) with pins [lscr.io/linuxserver/bookstack:26.09.1@sha256:99cd1f5707c1911afad213adec5c9739763b76f843d1477142231834ecdcb6f7 mariadb:12.3@sha256:805c8e104bd563d5bfa24fadd3f31cd419ea859cb5277f32b5dbf2db714f9ed1]
|
||||
2026/10/05 14:04:45 backup.go:990: [INFO] [backup] Restarting bookstack after volume dump
|
||||
2026/10/05 14:04:51 backup.go:974: [INFO] [backup] Stopping paperless-ngx for safe volume dump
|
||||
2026/10/05 14:04:58 backup.go:848: [DEBUG] [backup] Dumping volume paperless-ngx_paperless_redis_data for paperless-ngx
|
||||
2026/10/05 14:04:58 backup.go:873: [INFO] [backup] Volume dump: paperless-ngx/paperless-ngx_paperless_redis_data → 7.4 MB
|
||||
2026/10/05 14:04:58 data_versions.go:131: [DEBUG] [backup] paperless-ngx: stamped volume-dumps/paperless-ngx_paperless_redis_data.tar (7792128 B) with pins [ghcr.io/paperless-ngx/paperless-ngx:2.20.15 postgres:18-alpine redis:7-alpine]
|
||||
2026/10/05 14:04:58 backup.go:848: [DEBUG] [backup] Dumping volume paperless-ngx_paperless_data for paperless-ngx
|
||||
2026/10/05 14:04:59 backup.go:873: [INFO] [backup] Volume dump: paperless-ngx/paperless-ngx_paperless_data → 5.8 MB
|
||||
2026/10/05 14:04:59 data_versions.go:131: [DEBUG] [backup] paperless-ngx: stamped volume-dumps/paperless-ngx_paperless_data.tar (6052352 B) with pins [ghcr.io/paperless-ngx/paperless-ngx:2.20.15 postgres:18-alpine redis:7-alpine]
|
||||
2026/10/05 14:04:59 backup.go:848: [DEBUG] [backup] Dumping volume paperless-ngx_paperless_postgres_data for paperless-ngx
|
||||
2026/10/05 14:04:59 backup.go:873: [INFO] [backup] Volume dump: paperless-ngx/paperless-ngx_paperless_postgres_data → 68.8 MB
|
||||
2026/10/05 14:04:59 data_versions.go:131: [DEBUG] [backup] paperless-ngx: stamped volume-dumps/paperless-ngx_paperless_postgres_data.tar (72107008 B) with pins [ghcr.io/paperless-ngx/paperless-ngx:2.20.15 postgres:18-alpine redis:7-alpine]
|
||||
2026/10/05 14:04:59 backup.go:990: [INFO] [backup] Restarting paperless-ngx after volume dump
|
||||
2026/10/05 14:05:11 backup.go:974: [INFO] [backup] Stopping privatebin for safe volume dump
|
||||
2026/10/05 14:05:11 backup.go:848: [DEBUG] [backup] Dumping volume privatebin_privatebin_data for privatebin
|
||||
2026/10/05 14:05:12 backup.go:873: [INFO] [backup] Volume dump: privatebin/privatebin_privatebin_data → 1.5 KB
|
||||
2026/10/05 14:05:12 data_versions.go:131: [DEBUG] [backup] privatebin: stamped volume-dumps/privatebin_privatebin_data.tar (1536 B) with pins [privatebin/pdo:2.0.6@sha256:4c141b2326f8b353598ce9ce7507a9cfecf2dad5c60a39fea903d430e296d8f5]
|
||||
2026/10/05 14:05:12 backup.go:986: [INFO] [backup] privatebin NOT restarted after the volume dump: the household stopped it meanwhile
|
||||
2026/10/05 14:05:12 backup.go:684: [INFO] [backup] App-data backup completed: 2 databases (483.2 KB total), 3 volume dump(s) (32.785s)
|
||||
2026/10/05 14:05:12 recovery_unit.go:289: [INFO] [backup] Recovery unit captured for bookstack → /mnt/sys_drive/felhom-data/backups/primary/bookstack (images=2, secrets-referenced=3, data_keys=0, portable-carried=2/2, withheld=1)
|
||||
2026/10/05 14:05:12 recovery_unit.go:289: [INFO] [backup] Recovery unit captured for paperless-ngx → /mnt/felhom-drives/scratch_hdd/userdata/paperless-ngx/backups/primary/paperless-ngx (images=3, secrets-referenced=3, data_keys=0, portable-carried=2/2, withheld=1)
|
||||
2026/10/05 14:05:12 recovery_unit.go:289: [INFO] [backup] Recovery unit captured for privatebin → /mnt/sys_drive/felhom-data/backups/primary/privatebin (images=1, secrets-referenced=0, data_keys=0, portable-carried=0/0, withheld=0)
|
||||
|
||||
/var/lib/felhom/docker/volumes/felhom-controller-data/_data/data/appdata-run.json
|
||||
/var/lib/docker/volumes/felhom-controller-data/_data/data/appdata-run.json
|
||||
@@ -0,0 +1,12 @@
|
||||
## after run 1, 2026-10-05T14:05:34Z
|
||||
### run record
|
||||
{"running":false,"started_at":"0001-01-01T00:00:00Z","interrupted":"0001-01-01T00:00:00Z"}
|
||||
### restore points (bookstack)
|
||||
{"ok":true,"data":[{"time":"2026-10-05T14:04:39Z","short_id":"helyi","tier":1,"drive_label":"Belső SSD (rendszer)"}]}
|
||||
|
||||
HTTP 200
|
||||
|
||||
### volume dump file times
|
||||
total 157172
|
||||
-rw-r--r-- 1 root root 70144 2026-10-05T14:04:44 bookstack_bookstack_config.tar
|
||||
-rw-r--r-- 1 root root 160868864 2026-10-05T14:04:44 bookstack_bookstack_db_data.tar
|
||||
@@ -0,0 +1,34 @@
|
||||
## run 2 start 2026-10-05T14:05:50Z
|
||||
HTTP 200
|
||||
### watcher
|
||||
match 2026-10-05T14:05:58.856456447Z
|
||||
restarted 2026-10-05T14:05:59.669364637Z rc=0
|
||||
Up 10 seconds (healthy)
|
||||
### record at the moment of the cut
|
||||
{"running":true,"started_at":"2026-10-05T14:05:53.282495446Z","interrupted":"0001-01-01T00:00:00Z"}
|
||||
### record after restart
|
||||
{"running":false,"started_at":"2026-10-05T14:05:53.282495446Z","interrupted":"2026-10-05T14:05:53.282495446Z"}
|
||||
### controller log around the cut (previous + new process)
|
||||
2026/10/05 14:05:53 backup.go:557: [INFO] [backup] Starting database dump run
|
||||
2026/10/05 14:05:53 dbdump.go:208: [INFO] [backup] Discovered 2 databases
|
||||
2026/10/05 14:05:53 backup.go:596: [INFO] [backup] Discovered 2 database(s): paperless-postgres(postgres), bookstack-db(mariadb)
|
||||
2026/10/05 14:05:53 dbdump.go:410: [INFO] [backup] DB dump: paperless-postgres → paperless-ngx-postgres.sql (428.1 KB, 321ms, 72 tables)
|
||||
2026/10/05 14:05:54 dbdump.go:410: [INFO] [backup] DB dump: bookstack-db → bookstack-mariadb.sql (55.9 KB, 310ms, 41 tables)
|
||||
2026/10/05 14:05:54 backup.go:974: [INFO] [backup] Stopping bookstack for safe volume dump
|
||||
2026/10/05 14:05:58 backup.go:873: [INFO] [backup] Volume dump: bookstack/bookstack_bookstack_config → 49.0 KB
|
||||
2026/10/05 14:05:59 main.go:340: [INFO] felhom-controller 0.296.0 starting (customer: demo-hp, domain: enkisfelhom.hu)
|
||||
2026/10/05 14:05:59 appstop_marker.go:278: [WARN] [appstop] crash recovery: an app-data backup (volume dump) (op "volume-dump:bookstack") was interrupted and left 1 app(s) stopped — restarting them: [bookstack]
|
||||
2026/10/05 14:05:59 manager.go:1235: [INFO] [stacks] Starting stack: bookstack
|
||||
2026/10/05 14:06:06 appstop_marker.go:304: [INFO] [appstop] crash recovery: restarted bookstack after the interrupted an app-data backup (volume dump)
|
||||
2026/10/05 14:06:06 sync.go:117: [INFO] [sync] Starting catalog sync (repo: https://gitea.dooplex.hu/admin/app-catalog-felhom.eu.git, interval: 15m0s)
|
||||
2026/10/05 14:06:06 sync.go:208: [INFO] [sync] Starting catalog sync
|
||||
2026/10/05 14:06:06 restore_record_wiring.go:35: [INFO] [backup] restore record wired: /opt/docker/felhom-controller/data/restore-status.json (interrupted at startup: false)
|
||||
2026/10/05 14:06:06 run_record_wiring.go:27: [WARN] [backup] the app-data backup run started 2026-10-05T14:05:53Z was cut off by the stop — the backup pages say so until the next complete run (R-519)
|
||||
2026/10/05 14:06:06 scheduler.go:223: [INFO] [scheduler] Starting scheduler with 19 jobs
|
||||
2026/10/05 14:06:06 backup.go:1179: [INFO] [backup] Found 3 DB dump files across drives
|
||||
2026/10/05 14:06:06 dbdump.go:208: [INFO] [backup] Discovered 2 databases
|
||||
2026/10/05 14:06:06 [INFO] [backup] Discovered app data: 3 apps
|
||||
2026/10/05 14:06:06 recovery_unit.go:289: [INFO] [backup] Recovery unit captured for bookstack → /mnt/sys_drive/felhom-data/backups/primary/bookstack (images=2, secrets-referenced=3, data_keys=0, portable-carried=2/2, withheld=1)
|
||||
2026/10/05 14:06:06 recovery_unit.go:289: [INFO] [backup] Recovery unit captured for paperless-ngx → /mnt/felhom-drives/scratch_hdd/userdata/paperless-ngx/backups/primary/paperless-ngx (images=3, secrets-referenced=3, data_keys=0, portable-carried=2/2, withheld=1)
|
||||
2026/10/05 14:06:06 backup.go:1250: [INFO] [backup] Backup status cache refreshed
|
||||
2026/10/05 14:06:09 manager.go:1601: [INFO] [stacks] bookstack lscr.io/linuxserver/bookstack:26.09.1@sha256:99cd1f5707c1911afad213adec5c9739763b76f843d1477142231834ecdcb6f7 running Up 3 seconds (health: starting)
|
||||
@@ -0,0 +1,23 @@
|
||||
## after the cut, 2026-10-05T14:06:23Z
|
||||
### GET /backups
|
||||
data-interrupted-run present: 1
|
||||
<div class="alert alert-warning" data-interrupted-run="true">A legutóbbi mentés (2026-10-05 16:05) megszakadt, mert a doboz vagy a vezérlő újraindult. Amit nem fejezett be, annak a korábbi mentése maradt meg — minden visszaállítási pont annyira friss, amennyire a legrégebbi része. A következő teljes mentés után ez az üzenet eltűnik.
|
||||
HTTP 200
|
||||
### GET /backups/apps
|
||||
data-interrupted-run present: 1
|
||||
<div class="alert alert-warning" data-interrupted-run="true">A legutóbbi mentés (2026-10-05 16:05) megszakadt, mert a doboz vagy a vezérlő újraindult. Amit nem fejezett be, annak a korábbi mentése maradt meg — minden visszaállítási pont annyira friss, amennyire a legrégebbi része. A következő teljes mentés után ez az üzenet eltűnik.
|
||||
HTTP 200
|
||||
### negative control: a marker that must NOT be on the page
|
||||
0
|
||||
### restore points (bookstack)
|
||||
{"ok":true,"data":[{"time":"2026-10-05T14:04:44Z","short_id":"helyi","tier":1,"drive_label":"Belső SSD (rendszer)"}]}
|
||||
### dump file times
|
||||
total 314252
|
||||
-rw-r--r-- 1 root root 50176 2026-10-05T14:05:58 bookstack_bookstack_config.tar
|
||||
-rw-r--r-- 1 root root 160868864 2026-10-05T14:04:44 bookstack_bookstack_db_data.tar
|
||||
-rw-r--r-- 1 root root 160867328 2026-10-05T14:05:58 bookstack_bookstack_db_data.tar.tmp
|
||||
drwxr-xr-x 2 root root 4096 2026-10-05T14:05:54 db-dumps
|
||||
drwxr-xr-x 2 root root 4096 2026-10-05T14:05:58 volume-dumps
|
||||
### bookstack after
|
||||
bookstack Up 27 seconds (healthy)
|
||||
bookstack-db Up 32 seconds (healthy)
|
||||
@@ -0,0 +1,12 @@
|
||||
## run 3 (complete) start 2026-10-05T14:06:44Z
|
||||
HTTP 200
|
||||
## run 3 end 2026-10-05T14:07:28Z: {"ok":true,"data":{"db_dump":{"count":2,"duration":"33.315240638s","last_run":"2026-10-05T14:07:20.615183031Z","success":true},"enabled":true,"running":false}}
|
||||
2026/10/05 14:05:12 backup.go:684: [INFO] [backup] App-data backup completed: 2 databases (483.2 KB total), 3 volume dump(s) (32.785s)
|
||||
2026/10/05 14:07:20 backup.go:684: [INFO] [backup] App-data backup completed: 2 databases (483.9 KB total), 3 volume dump(s) (33.315s)
|
||||
{"running":false,"started_at":"0001-01-01T00:00:00Z","interrupted":"0001-01-01T00:00:00Z"}
|
||||
total 157152
|
||||
-rw-r--r-- 1 root root 50688 2026-10-05T14:06:52 bookstack_bookstack_config.tar
|
||||
-rw-r--r-- 1 root root 160867328 2026-10-05T14:06:53 bookstack_bookstack_db_data.tar
|
||||
/backups data-interrupted-run: 0
|
||||
/backups/apps data-interrupted-run: 0
|
||||
restore points: {"ok":true,"data":[{"time":"2026-10-05T14:06:47Z","short_id":"helyi","tier":1,"drive_label":"Belső SSD (rendszer)"}]}
|
||||
@@ -0,0 +1,8 @@
|
||||
## teardown 2026-10-05T14:07:48Z
|
||||
stop: HTTP 200
|
||||
remove: HTTP 200
|
||||
containers: 0
|
||||
volumes: 0
|
||||
stackdir: none
|
||||
backups: none
|
||||
/root/.dbody
|
||||
@@ -0,0 +1,9 @@
|
||||
## runbook §3 drill 2026-10-05T14:09:41Z (scratch /var/lib/felhom-hub-backup/sec3.65qk, root 0700)
|
||||
step 1 restore host/dooplex-hub/2026-10-05T13:50:18Z
|
||||
restore complete (352.676 MiB processed in 3.9s, average 90.961 MiB/s)
|
||||
-rw------- 369807360 hub.db
|
||||
step 2 the seal key from Secret/offsite-secret-key → /var/lib/felhom-hub-backup/sec3.65qk/k (0600, not printed)
|
||||
key file bytes: 64
|
||||
step 3 hubdb-check (same key): hosts=4 console_passwords_opened=4 failed=0 absent=0 rc=0
|
||||
control (a random key): hosts=4 console_passwords_opened=0 failed=4 absent=0 hubdb-check: FAILED: not every console password opened with this key rc=0
|
||||
scratch shredded: gone
|
||||
@@ -0,0 +1,7 @@
|
||||
### R7 hubdb-check counts a failed open as opened
|
||||
=== RUN TestCheck_WrongKeyOpensNothing
|
||||
main_test.go:63: got {hosts:2 opened:1 failed:0 absent:1}, want opened=0 failed=1
|
||||
--- FAIL: TestCheck_WrongKeyOpensNothing (0.03s)
|
||||
FAIL
|
||||
FAIL gitea.dooplex.hu/admin/felhom-hub/cmd/hubdb-check 0.040s
|
||||
FAIL
|
||||
@@ -0,0 +1,17 @@
|
||||
## snapshot list on ep0 with the READ-ONLY token, 2026-10-05T13:50:40Z
|
||||
+=======================================+=============+=====================================+
|
||||
| snapshot | size | files |
|
||||
+=======================================+=============+=====================================+
|
||||
| host/dooplex-hub/2026-10-05T13:50:18Z | 352.677 MiB | catalog.pcat1 hubdb.pxar index.json |
|
||||
+=======================================+=============+=====================================+
|
||||
unit rc=0
|
||||
Result=success
|
||||
2026-10-05T15:50:41+02:00 dooplex felhom-hub-db-restore-test[3040582]: felhom-hub-db-restore-test: restoring host/dooplex-hub/2026-10-05T13:50:18Z
|
||||
2026-10-05T15:50:46+02:00 dooplex felhom-hub-db-restore-test[3040582]: felhom-hub-db-restore-test: checked: integrity ok, 4 host(s), 4 sealed console password(s), 0 readable
|
||||
2026-10-05T15:50:46+02:00 dooplex felhom-hub-db-restore-test[3040582]: felhom-hub-db-restore-test: success signal written
|
||||
# HELP felhom_hub_db_restore_test_last_success_timestamp_seconds Last successful restore test of the hub DB copy on ep0 (R-173).
|
||||
# TYPE felhom_hub_db_restore_test_last_success_timestamp_seconds gauge
|
||||
felhom_hub_db_restore_test_last_success_timestamp_seconds 1791208246
|
||||
.cache
|
||||
.kube
|
||||
stage
|
||||
@@ -0,0 +1,12 @@
|
||||
## 2026-10-05T13:51:07Z token limits (each line = the tool output)
|
||||
push forget : Error: permission check failed - missing Datastore.Modify|Datastore.Prune on /datastore/felhom-offsite/operator
|
||||
restore forget : Error: permission check failed - missing Datastore.Modify|Datastore.Prune on /datastore/felhom-offsite/operator
|
||||
restore backup : Error: missing permissions 'Datastore.Backup' on '/datastore/felhom-offsite/operator'
|
||||
push list ns root (households): Error: permission check failed - missing Datastore.Audit|Datastore.Backup on /datastore/felhom-offsite
|
||||
restore list ns root (households): Error: permission check failed - missing Datastore.Audit|Datastore.Backup on /datastore/felhom-offsite
|
||||
push restore own copy: Error: missing key - manifest was created with key b2:19:bf:36:3b:97:3d:6c exit-file=absent
|
||||
## paper-key proof: restore with a key file rebuilt from the data field only
|
||||
restore rc=0
|
||||
integrity: ok hosts: 4
|
||||
## and with NO key: Error: missing key - manifest was created with key b2:19:bf:36:3b:97:3d:6c
|
||||
push restore own copy WITH key: restore complete (352.676 MiB processed in 4.6s, average 76.793 MiB/s) file=PRESENT
|
||||
@@ -26,6 +26,16 @@
|
||||
|
||||
---
|
||||
|
||||
## 2026-10-05 (evening) — the hub database off DooPlex (hub v0.136.0; `09` rulings 125–127)
|
||||
|
||||
The full text of every row below: `git show cea8502f:documentation/backlog/OPEN-ITEMS.md`.
|
||||
|
||||
| Row | What | Closed | Evidence |
|
||||
|---|---|---|---|
|
||||
| **R-519** | **After a backup torn by a power cut, a restore point carried the new database dump's time over the previous run's files, and no screen said the run was interrupted (P2).** Fixed in controller v0.296.0 (page notice + synthesised status; dating by the oldest part since v0.275.0). **Live on 9202 (operator ruling 126):** a complete run, then a run cut by `docker restart felhom-controller` 0.8 s after bookstack's config dump and before its database volume; afterwards both /backups and /backups/apps carry `data-interrupted-run` („A legutóbbi mentés (2026-10-05 16:05) megszakadt …", negative control 0), the restore point reads 14:04:44Z = the run-1 database volume (its oldest part; config 14:05:58, SQL 14:05:54), the controller restarted bookstack itself; the next complete run cleared the notice and replaced the torn `.tar.tmp`. The throwaway bookstack was removed through the product (0 containers, volumes, folders, backups). | CLOSED 2026-10-05 — FIXED controller v0.296.0, proven live | `audits/hub-db-offsite-2026-10-05/partD/r519/`; `internal/backup/run_record_test.go` |
|
||||
|
||||
---
|
||||
|
||||
## 2026-10-05 (late afternoon) — the hub's own safety, boxes left behind, the agent's root grants (hub v0.135.0, controller v0.296.0, agent v0.146.1, golden 0.296.0; CC decisions 119–124)
|
||||
|
||||
The full text of every row below: `git show 9bb45eaa:documentation/backlog/OPEN-ITEMS.md` (R-880 was opened and closed in this session).
|
||||
|
||||
File diff suppressed because one or more lines are too long
@@ -1,16 +1,16 @@
|
||||
# Runbook — the hub database in a backup that is NOT on DooPlex (R-173, R-232) — PROPOSED, needs the operator's go
|
||||
# Runbook — the hub database in a backup that is NOT on DooPlex (R-173, R-232) — IN FORCE since 2026-10-05
|
||||
|
||||
> **Status: PROPOSED 2026-10-05. Nothing here has been done.** DooPlex and ep0 are protected; every step below changes
|
||||
> one of them, so each waits for the operator's go (the decision is in `STATUS.md`). The readings this plan rests on:
|
||||
> `audits/hub-safety-2026-10-05/partC/readings.txt` (read only).
|
||||
> **Status: IN FORCE 2026-10-05** (option A, `09` decision 125). Steps 0–7 DONE, each marked below; §3 is a tested
|
||||
> procedure. Evidence: `audits/hub-db-offsite-2026-10-05/` (part A–D). The readings the plan rested on:
|
||||
> `audits/hub-safety-2026-10-05/partC/readings.txt`. **Corrections found while doing it are marked „Corrected".**
|
||||
|
||||
## 1. What is true today (measured 2026-10-05)
|
||||
|
||||
| Question | Answer |
|
||||
|---|---|
|
||||
| Where the hub database lives | `/data/hub.db` (+ `-wal`, `-shm`) in the hub pod, PVC `hub-data` (Longhorn, 1 Gi, replicas on DooPlex's `sdb1`). 357 MiB. |
|
||||
| Where the hub database lives | `/data/hub.db` (+ `-wal`, `-shm`) in the hub pod, PVC `hub-data` (Longhorn, **2 Gi since 2026-10-05**, was 1 Gi; replicas on DooPlex's `sdb1`). 357 MiB; a snapshot is 353 MiB. |
|
||||
| Is it in a backup? | **Yes, but only on DooPlex.** Longhorn's `backup-daily` (04:00) and `backup-weekly` (Sun 05:00), `retain=1`, write to `nfs://192.168.0.180:/mnt/5_hdd/backup/longhorn-pvc` — DooPlex's own `sda1`. Last: 2026-10-05 02:06 UTC, Completed. |
|
||||
| Why R-173 said "excluded" | The PVC carries `recurring-job-group.longhorn.io/default: disabled` (git, `manifests/hub.yaml`, commit `868e8465` of 2026-02-16, no reason given). The live Longhorn **Volume** carries `enabled` — set by hand at some point, so the backups run. **This is drift:** the PVC label is the source of truth (Longhorn syncs PVC → Volume), so the backups can stop at any sync without anyone seeing it. |
|
||||
| Why R-173 said "excluded" | The PVC carries `recurring-job-group.longhorn.io/default: disabled` (git, `manifests/hub.yaml`, commit `868e8465` of 2026-02-16, no reason given). The live Longhorn **Volume** carries `enabled` — set by hand at some point, so the backups run. **This is drift:** the PVC label is the source of truth (Longhorn syncs PVC → Volume), so the backups can stop at any sync without anyone seeing it. **Corrected 2026-10-05:** the Volume kept `enabled` through every sync since February while the PVC said `disabled` — Longhorn did NOT copy the PVC label down. The PVC now says `enabled` too (Step 1), so the two agree; which one Longhorn reads was not measured. |
|
||||
| What DooPlex's own backup covers | `dooplex-backup.timer` (03:19): k3s state, k8s Secrets (GPG files), Gitea mirrors, user data, PostgreSQL dumps — **all onto `sda1`, the same machine.** Nothing leaves DooPlex (`audits/RECON-dooplex-backup-2026-08-06.md`, R-232). |
|
||||
| What tells anyone a backup failed | **Nothing.** `NOTIFY_WEBHOOK_URL` is commented out, so `notify_failure` is a no-op. No Prometheus rule watches a Longhorn backup's success or age, nor `dooplex-backup.service`. |
|
||||
| What the database holds | Box→hub API keys, customer configs (incl. the owner passphrase), escrow custody blobs (opaque), the PBS-DR token values, the off-site sub-account passwords and — since hub v0.135.0 — the console passwords **sealed** under `OFFSITE_SECRET_KEY`. |
|
||||
@@ -22,7 +22,7 @@ ep0 is already off-site (Hetzner), already runs PBS, and DooPlex already reaches
|
||||
(127.0.0.1:18007). ep0's datastore is also pulled back to DooPlex nightly (`ep0-copy`), so the copy exists in two places,
|
||||
one of them off DooPlex. The copy is encrypted on DooPlex with a key ep0 never sees.
|
||||
|
||||
### Step 0 — the keys go off DooPlex first (operator, at the keyboard, 5 min)
|
||||
### Step 0 — the keys go off DooPlex first (operator, at the keyboard, 5 min) — DONE 2026-10-05 (operator: both saved)
|
||||
|
||||
Without these, every later step backs up something nobody can open after a DooPlex loss.
|
||||
|
||||
@@ -33,33 +33,65 @@ sudo kubectl -n felhom-system get secret offsite-secret-key -o jsonpath='{.data.
|
||||
sudo proxmox-backup-client key paperkey /etc/felhom-hub-backup/enc.key --output-format text
|
||||
```
|
||||
|
||||
### Step 1 — end the label drift (CC, a felhom.eu commit + ArgoCD sync; reversible)
|
||||
**The `data` field of that output is enough** (the key has no passphrase, `kdf: null`): a key file rebuilt from it alone —
|
||||
`{"kdf": null, "created": "<any RFC 3339 time>", "modified": "<same>", "data": "<saved value>"}` — restored and
|
||||
decrypted the copy on 2026-10-05 (`audits/hub-db-offsite-2026-10-05/partD/token-limits-and-paperkey.txt`). Note:
|
||||
`proxmox-backup-client key show` prints a fingerprint only when the file stores one, so it cannot check a rebuilt key —
|
||||
a restore can.
|
||||
|
||||
### Step 1 — end the label drift (CC, a felhom.eu commit + ArgoCD sync; reversible) — DONE 2026-10-05
|
||||
|
||||
`manifests/hub.yaml`: `recurring-job-group.longhorn.io/default: disabled` → `enabled`. Sync. Check:
|
||||
`sudo kubectl -n felhom-system get pvc hub-data -o jsonpath='{.metadata.labels}'` and the Volume label both read `enabled`.
|
||||
This keeps today's on-DooPlex copy alive; it is not the off-site copy.
|
||||
|
||||
### Step 2 — a write-only place on ep0 (on ep0, root; the operator's go for an ep0 change)
|
||||
Done with the volume growth to 2 Gi (two snapshots of 353 MiB do not fit in 1 Gi). **The online growth failed** — Longhorn's
|
||||
`instance-manager` (116 days up) called a host PID that no longer existed (`nsenter: cannot open /host/proc/196610/ns/mnt`);
|
||||
an offline growth was impossible too (the expansion holds its own attachment ticket). The operator approved restarting
|
||||
the instance-manager: all 77 DooPlex volumes back `attached/healthy` in 110 s, the volume grew, and one app (zipline, on
|
||||
`:latest`) came back on a newer release that refused its database — pinned to 4.7.0 (homelab-manifests). Evidence:
|
||||
`audits/hub-db-offsite-2026-10-05/partA/step1-*.txt`.
|
||||
|
||||
### Step 2 — a write-only place on ep0 (on ep0, root; the operator's go for an ep0 change) — DONE 2026-10-05
|
||||
|
||||
What was run (PBS 4.2.8 on ep0; the proposal's commands were wrong in three places, corrected here):
|
||||
|
||||
```bash
|
||||
proxmox-backup-manager user create dooplex-hub@pbs --comment "DooPlex pushes the hub DB (R-173)"
|
||||
proxmox-backup-manager user generate-token dooplex-hub@pbs push # the secret → a 0600 file on DooPlex, file → file
|
||||
# namespace for operator data, apart from the households' namespaces
|
||||
proxmox-backup-client namespace create operator --repository 'root@pam@127.0.0.1:8007:felhom-offsite'
|
||||
proxmox-backup-manager acl update /datastore/felhom-offsite/operator DatastoreBackup --auth-id 'dooplex-hub@pbs!push'
|
||||
# retention on ep0 (the server prunes; the pushing token cannot delete — DatastoreBackup has no Prune)
|
||||
proxmox-backup-manager prune-job create prune-operator-hubdb --store felhom-offsite --ns operator \
|
||||
--schedule 'daily 03:45' --keep-daily 14 --keep-weekly 8
|
||||
proxmox-backup-manager user create dooplex-hub@pbs --comment "..." # no password: cannot log in
|
||||
proxmox-backup-debug api create /admin/datastore/felhom-offsite/namespace --name operator
|
||||
# (the CLI crashes AFTER creating it, printing the result — 'not implemented'; check it exists, don't re-run)
|
||||
# two tokens, each secret file → file into a root 0600 file on DooPlex, never printed:
|
||||
ssh root@<ep0> "proxmox-backup-debug api create /access/users/dooplex-hub@pbs/token/push --output-format json" \
|
||||
| python3 -c '<print json["value"]>' | sudo sh -c 'umask 077; cat > /etc/felhom-hub-backup/token-push'
|
||||
# (same for token/restore → token-restore; `user generate-token` has no --output-format)
|
||||
P=/datastore/felhom-offsite/operator
|
||||
proxmox-backup-manager acl update $P DatastoreBackup --auth-id dooplex-hub@pbs # a token's rights are cut
|
||||
proxmox-backup-manager acl update $P DatastoreReader --auth-id dooplex-hub@pbs # down by its user's rights
|
||||
proxmox-backup-manager acl update $P DatastoreBackup --auth-id 'dooplex-hub@pbs!push'
|
||||
proxmox-backup-manager acl update $P DatastoreReader --auth-id 'dooplex-hub@pbs!restore'
|
||||
proxmox-backup-manager prune-job create prune-operator-hubdb --store felhom-offsite --ns operator --max-depth 0 \
|
||||
--schedule '03:45' --keep-daily 14 --keep-weekly 8 # 'daily 03:45' is not a PBS calendar event
|
||||
```
|
||||
|
||||
### Step 3 — a consistent snapshot of the live database (CC, a hub release)
|
||||
Measured (`partB/`, `partD/token-limits-and-paperkey.txt`): neither token can forget a snapshot or list the datastore
|
||||
root (the households); the restore token cannot write. **Corrected: the push token CAN restore its own copies** — PBS
|
||||
lets a backup's owner read it back (`DatastoreBackup` = `Datastore.Backup`, owner-scoped). It reaches only `operator`,
|
||||
and everything it can read is encrypted with a key ep0 never sees. The households' two prune jobs and the Sunday GC
|
||||
are unchanged (before/after in `partB/`).
|
||||
|
||||
### Step 3 — a consistent snapshot of the live database (CC, a hub release) — DONE, hub v0.136.0 (`05` §16.3)
|
||||
|
||||
`hub.db` is in WAL mode and is written every few seconds; copying the three files is not one point in time. The hub
|
||||
gets a nightly `VACUUM INTO '/data/snapshots/hub-<UTC date>.db'` (keeps 2, logs size and duration) — one SQLite
|
||||
statement, consistent by construction, WAL-aware. **Needs a hub release** (filed under R-173). No `sqlite3` exists in the
|
||||
hub image, so the copy must be made by the hub itself.
|
||||
statement, consistent by construction, WAL-aware. No `sqlite3` exists in the hub image, so the copy is made by the hub
|
||||
itself. **Corrected:** one snapshot took 44 s on the live volume (0.63 s on a local scratch copy), 353 MiB.
|
||||
|
||||
### Step 4 — the push (on DooPlex, root; `felhom-hub-db-backup.service` + `.timer` 02:30, CC writes, operator approves)
|
||||
### Step 4 — the push (on DooPlex, root; `felhom-hub-db-backup.service` + `.timer` 02:30) — DONE 2026-10-05
|
||||
|
||||
**The real script is `scripts/hub-db-backup/felhom-hub-db-backup`** (versioned, R-231; installed by `install.sh`;
|
||||
15 tests in `test_hub_db_backup.py`, run by hand — not in CI). It adds to the sketch below: it refuses a snapshot older
|
||||
than 26 h (the hub stopped snapshotting), a copy whose size differs from the pod's file, and a copy with no hosts; it
|
||||
encrypts with `--crypt-mode encrypt`. The sketch, as proposed:
|
||||
|
||||
```bash
|
||||
#!/bin/sh -eu
|
||||
@@ -77,7 +109,10 @@ echo "felhom_hub_db_backup_last_success_timestamp_seconds $(date +%s)" > /var/li
|
||||
&& mv /var/lib/node_exporter/textfile_collector/felhom_hub_db.prom.$$ /var/lib/node_exporter/textfile_collector/felhom_hub_db.prom
|
||||
```
|
||||
|
||||
### Step 5 — the restore test (weekly, Sun 04:30, same unit family)
|
||||
### Step 5 — the restore test (weekly, Sun 04:30, same unit family) — DONE 2026-10-05
|
||||
|
||||
**The real script is `scripts/hub-db-backup/felhom-hub-db-restore-test`**, with the READ-ONLY token. It also refuses a
|
||||
newest copy older than 50 h. The sketch, as proposed:
|
||||
|
||||
```bash
|
||||
T=$(mktemp -d); chmod 700 "$T"
|
||||
@@ -90,9 +125,12 @@ shred -u "$T/hub.db"*; rmdir "$T"
|
||||
echo "felhom_hub_db_restore_test_last_success_timestamp_seconds $(date +%s)" > …/felhom_hub_db_restore.prom # same tmp+mv
|
||||
```
|
||||
|
||||
The push token needs `DatastoreReader` on `operator` too for the restore (or a second, read-only token — cleaner).
|
||||
Done with a second, read-only token (`token-restore`).
|
||||
|
||||
### Step 6 — the alarm (homelab-manifests `prometheus-rules`, then `POST /-/reload` — the Prometheus there has no reloader)
|
||||
### Step 6 — the alarm (homelab-manifests `prometheus-rules`, then `POST /-/reload` — the Prometheus there has no reloader) — DONE 2026-10-05
|
||||
|
||||
In the `backup-freshness` group; `promtool test rules` proves it (`partC/bf_test.yml`, two red-proofs). The Prometheus
|
||||
Deployment is OutOfSync in ArgoCD for a reason unrelated to this; only the rules ConfigMap was synced.
|
||||
|
||||
```yaml
|
||||
- alert: HubDBBackupStale
|
||||
@@ -109,18 +147,46 @@ The push token needs `DatastoreReader` on `operator` too for the restore (or a s
|
||||
Both reach the existing `email-notifications` receiver. `absent()` makes "the script never ran" an alarm too — an empty
|
||||
log is not a success.
|
||||
|
||||
### Step 7 — prove it once (CC, with the operator's go)
|
||||
### Step 7 — prove it once (CC, with the operator's go) — DONE 2026-10-05 (`partD/`)
|
||||
|
||||
Run the unit by hand; read the snapshot on ep0 (`proxmox-backup-client snapshot list --ns operator`); run the restore
|
||||
test by hand; stop the timer for a day on purpose and see `HubDBBackupStale` mail arrive (positive observable), then
|
||||
start it again.
|
||||
|
||||
## 3. Bringing the hub back from this copy (the procedure the plan exists for)
|
||||
## 3. Bringing the hub back from this copy (the procedure the plan exists for) — TESTED 2026-10-05
|
||||
|
||||
1. A k3s with the `felhom` ArgoCD app, and **`Secret/offsite-secret-key` recreated with the SAME value** (Step 0 copy).
|
||||
2. Restore the newest snapshot (Step 5's first command, with the paper key), scale `deploy/hub` to 0, copy `hub.db` into
|
||||
the PVC (no `-wal`/`-shm` — the snapshot is a whole database), scale to 1. The log line
|
||||
`console passwords sealed at rest (0 legacy plaintext row(s) sealed now)` and a working reveal prove the key matches.
|
||||
Steps 1–3 were run on 2026-10-05 against the real copy on ep0 (`audits/hub-db-offsite-2026-10-05/partD/restore-procedure/drill.txt`):
|
||||
4 hosts, **4 of 4 console passwords opened with the saved seal key, 0 of 4 with a random key**. Steps 4–5 (into a live
|
||||
PVC) were NOT run — that needs the hub down; they are the ordinary scale-copy-scale.
|
||||
|
||||
What you need, all from the password manager: the seal key (`OFFSITE_SECRET_KEY`), the backup key's `data` field, and
|
||||
the read-only token (or ep0 root to mint a new one: Step 2).
|
||||
|
||||
1. **The backup key file.** On the machine doing the restore, as root, `umask 077`, write
|
||||
`{"kdf": null, "created": "2026-01-01T00:00:00+00:00", "modified": "2026-01-01T00:00:00+00:00", "data": "<data>"}`
|
||||
to `enc.key` (Step 0 note). A copy of DooPlex's `/etc/felhom-hub-backup/enc.key` works as is.
|
||||
2. **Restore the newest copy** (from any machine that reaches ep0's PBS on 8007 — DooPlex uses the tunnel 127.0.0.1:18007):
|
||||
```bash
|
||||
export PBS_PASSWORD_FILE=<token-restore file> PBS_FINGERPRINT=<ep0 cert fingerprint, /etc/felhom-hub-backup/env>
|
||||
R='dooplex-hub@pbs!restore@<ep0>:8007:felhom-offsite'
|
||||
proxmox-backup-client snapshot list host/dooplex-hub --ns operator --repository "$R" # pick the newest
|
||||
proxmox-backup-client restore host/dooplex-hub/<time> hubdb.pxar ./out --ns operator --keyfile enc.key --repository "$R"
|
||||
sqlite3 -readonly out/hub.db 'PRAGMA integrity_check' # must print: ok
|
||||
```
|
||||
If DooPlex's `ep0-copy` datastore survived, the same copy is there too (pulled nightly).
|
||||
3. **Prove the seal key matches BEFORE putting the copy in place** — on a COPY of `out/hub.db` (the check migrates it):
|
||||
```bash
|
||||
cd felhom.eu/hub && go build -o hubdb-check ./cmd/hubdb-check
|
||||
printf '%s' "<OFFSITE_SECRET_KEY>" > k; chmod 600 k # from the password manager — not on a command line in a shared shell
|
||||
./hubdb-check copy-of-hub.db k # want: hosts=N console_passwords_opened=N failed=0; exit 0
|
||||
```
|
||||
`failed>0` means the wrong seal key: the hub would start but could open no console password (`05` §16.2).
|
||||
4. **A k3s with the `felhom` ArgoCD app**, and `Secret/offsite-secret-key` recreated with the SAME value:
|
||||
`kubectl -n felhom-system create secret generic offsite-secret-key --from-file=OFFSITE_SECRET_KEY=k`.
|
||||
5. **Into the PVC:** scale `deploy/hub` to 0; put `out/hub.db` into the volume as `/data/hub.db` (a helper pod mounting
|
||||
`hub-data`; delete any `hub.db-wal`/`-shm` there — the snapshot is a whole database); scale to 1. The start-up log
|
||||
line `console passwords sealed at rest (0 legacy plaintext row(s) sealed now)` and one reveal on a host page confirm it.
|
||||
6. Shred `k`, `enc.key` copies and `out/` when done.
|
||||
|
||||
## 4. Option B (not my pick): restic to a dedicated Hetzner Storage Box sub-account
|
||||
|
||||
|
||||
Reference in New Issue
Block a user