R-173 option A in force (hub DB nightly to ep0, restore-tested, alarmed); R-519 proven live on 9202 and closed; R-173/R-232 narrowed; R-882..R-885 opened (332 -> 335); runbook §3 tested; hubdb-check
gates / gates (push) Successful in 59s

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
2026-10-05 16:14:13 +02:00
parent cea8502f0b
commit afba622fb6
27 changed files with 611 additions and 44 deletions
+12
View File
@@ -16,6 +16,18 @@
> and holds nothing of its own; this file does hold its own content, namely the standing rulings below.
> **2026-10-05 (evening) — the hub database off DooPlex (hub v0.136.0; operator rulings `09` 125–127).** R-173 option A
> IN FORCE: hub `internal/dbsnap` writes `VACUUM INTO /data/snapshots/hub-<UTC>.db` at 02:00 Budapest (keep 2, `ErrBusy`
> on overlap, start-up catch-up when >24 h; `05` §16.3); DooPlex `scripts/hub-db-backup/` (installed by `install.sh` to
> `/usr/local/sbin`, units + timers in `/etc/systemd/system`, config `/etc/felhom-hub-backup/{env,token-push,
> token-restore,enc.key}` all root 0600) pushes at 02:30 to ep0 `felhom-offsite` ns `operator` as `dooplex-hub@pbs!push`
> (`DatastoreBackup`), restore-tests Sun 04:30 as `!restore` (`DatastoreReader`); ep0 prune job `prune-operator-hubdb`
> (03:45, keep-daily 14, keep-weekly 8, max-depth 0). Success-only textfile metrics → `HubDBBackupStale` (26 h, critical)
> and `HubDBRestoreTestStale` (8 d) in homelab-manifests `backup-freshness`. `hub/cmd/hubdb-check` opens a restored copy
> with the seal key (runbook §3). Hub PVC 2 Gi + label `enabled`. Collateral: Longhorn instance-manager restart (R-882),
> zipline pinned 4.7.0 (R-883). R-519 proven live on 9202 (now controller 0.296.0) and CLOSED. Register 332 → 335.
> Report: `REPORT-hub-db-offsite-2026-10-05.md`.
> **2026-10-05 (late afternoon) — the hub's own safety, boxes left behind, the agent's root grants (hub v0.135.0, controller
> v0.296.0, agent v0.146.1 + bundle `42333e96…`, golden 0.296.0 vouched with agent 0.146.1, min_agent 0.131.0).** CC decisions
> 119–124, *operator may reverse*. Hub: R-135 a cookie-less state change needs Basic + `X-Felhom-Operator` (`05` §16.1); R-133
+40 -5
View File
@@ -3,9 +3,44 @@
**Ready for the first real tester (Tester-2): yes. Tester 2 (a laptop, off at night) was offline again; nothing was
sent to it.**
**Updated 2026-10-05 (late afternoon, the hub-safety session): every box of ours healthy. The hub refuses forged form
posts and keeps the console passwords locked; the System page shows boxes left behind; the agent can no longer make
itself root. Report: `REPORT-hub-safety-2026-10-05.md`.**
**Updated 2026-10-05 (evening, the hub-database session): every box of ours healthy. The hub database now leaves
DooPlex every night, locked, to ep0, is test-restored every Sunday, and an alarm mails you if either stops. Report:
`REPORT-hub-db-offsite-2026-10-05.md`.**
## Tonight (2026-10-05, evening): the hub database off DooPlex; the cut-backup check on the scratch box
**Decisions:** none of mine. Yours (`09` 125–127): ep0 for the copy; the scratch-box restart allowed; the agent's three
by-design abilities stay. Today you also chose: grow the hub's disk to 2 GiB, and restart Longhorn's disk manager.
**What works now (proven live):**
- **The hub makes a clean copy of its database every night at 02:00** (hub 0.136.0). The first one: 353 MB, 44 s.
- **DooPlex checks it, locks it with a key ep0 never sees, and sends it to ep0 at 02:30.** First send: 7 s.
- **Every Sunday at 04:30 DooPlex takes the copy back from ep0 and checks it**: it opens, it is whole, every console
password in it is still locked. Done once by hand today: 4 boxes, 4 locked passwords, 0 readable.
- **The two ep0 accounts can only do their one job:** the sending one cannot delete, the checking one cannot write,
neither can see the households' backups. (The sending one can read back its own locked copies — that is how ep0 works.)
- **Your two saved keys work:** a key rebuilt from the paper copy you saved opened the copy, and your saved lock key
opened all 4 console passwords in it (a wrong key opened none).
- **The alarm:** proven by Prometheus' own rule test; the real alarm mail is below.
- **A backup cut off by a restart is now said on the backup pages** (scratch box): the page said so, the restore point
kept the older time of the part that was not redone, and the next full backup cleared the message.
**What broke, and what I did:**
- **The hub's disk would not grow:** Longhorn's disk manager on DooPlex was stuck. My first try (stopping the hub so the
disk could grow offline) did not work and kept the hub **down about 9.5 minutes**. You approved restarting the disk
manager: all 77 disks were back in under 2 minutes, and the hub's disk is 2 GiB now.
- **Zipline did not come back after that restart:** it is set to "always the newest", so it pulled a new release that
refused its database. I pinned it to the previous release; it runs and its database is updated. 7 more apps on
DooPlex use "always the newest" (new row).
**Register:** 332 → 335 rows (1 closed: the cut-backup check; 4 opened: the Longhorn fault, the "always newest" apps, a
monitoring sync drift, script tests not in CI).
**Needs you:**
1. **Nothing urgent.** If you do nothing, the copy runs every night and you get a mail only if it stops.
2. **When convenient:** pin the 7 other DooPlex apps that use "always the newest" (or tell me to list them for you). If
you do nothing, any restart may upgrade one of them by surprise, as it did zipline.
3. **Tester 2's one-time step** is unchanged (below).
## Today (2026-10-05, late afternoon): the hub's own safety; boxes left behind; the agent's admin rights
@@ -38,7 +73,7 @@ itself root. Report: `REPORT-hub-safety-2026-10-05.md`.**
minutes. Both figures are on the page now.
**Needs you:**
1. **Where the hub database's off-site copy goes** (it holds every box's keys and your customers' settings):
1. **(DECIDED 2026-10-05 evening: A, done — see Tonight)** **Where the hub database's off-site copy goes** (it holds every box's keys and your customers' settings):
- **A — my pick: ep0's backup server**, encrypted on DooPlex before it leaves, with a weekly restore test and an
alarm mail. Costs one small change on ep0 (a write-only account) and keeping two keys in your password manager.
- **B: a separate Hetzner Storage Box account** with restic. More new parts to look after than A.
@@ -47,7 +82,7 @@ itself root. Report: `REPORT-hub-safety-2026-10-05.md`.**
`documentation/runbooks/RUNBOOK-hub-db-offsite-backup.md`.
- **Either way, first:** put the hub's lock key (`OFFSITE_SECRET_KEY`) in your password manager — without it a copy
of the database cannot open the console passwords.
2. **The power-cut-during-backup check (R-519) on the scratch box:** the permission check refused my restarting the
2. **(DONE 2026-10-05 evening — see Tonight)** **The power-cut-during-backup check (R-519) on the scratch box:** the permission check refused my restarting the
controller in the middle of a backup. Say "go" and the next session does it once on 9202; if not, the fix stays
proven by tests only.
3. **Three things the agent can still do, by design** (each written in `03` §3.1): pick which controller image its own
@@ -235,6 +235,7 @@ likewise silent. Evidence: `audits/DRILL-r361-2026-08-22/evidence/06-part3-decis
| Secrets hygiene: bearer in k8s Secret, no secrets in git, single-quote credential store | hub v0.53, conventions | **IMPLEMENTED** | 07-13 closing bundle | |
| Operator login password changeable from UI | hub v0.54 | **IMPLEMENTED** | 07-13 | |
| **The operator sees boxes left behind: per-customer controller floors with their age (and which a global raise cannot move), each box's agent against the vouched one, a 7-day "agent behind" alarm and a "floor raise skipped boxes" mail** | hub **v0.135.0** (R-604, R-530) | **IMPLEMENTED — the page PROVEN-LIVE, the alarm and the mail unit-proven** | `audits/hub-safety-2026-10-05/partD/` (live System page: three per-customer floors, Tester 2 `0.142.0 → 0.145.0`); `osupdates/r530_agent_alarm_test.go`, `web/r604_floor_held_back_test.go` | the mail fires only on a GLOBAL raise below an override — not exercised live |
| **The hub database survives the loss of DooPlex: a nightly consistent copy, encrypted, on ep0; restore-tested weekly; an alarm when either stops** | hub **v0.136.0** (R-173), `scripts/hub-db-backup/`, homelab-manifests rules | **PROVEN-LIVE (2026-10-05)** — first push, ep0 listing, restore test (4 hosts, 4 sealed, 0 readable), token limits, a key rebuilt from the paper copy decrypts, the saved seal key opens 4/4 console passwords in the restored copy; the alarm by `promtool` rule test + red-proofs | `audits/hub-db-offsite-2026-10-05/`; `runbooks/RUNBOOK-hub-db-offsite-backup.md` | Runbook §3 steps 4–5 (into a live PVC) not exercised (R-173) |
| **The operator surface refuses a cross-site form post on BOTH login paths (session token; Basic auth + `X-Felhom-Operator`)** | hub **v0.135.0** (R-135) | **PROVEN-LIVE (2026-10-05)** | `audits/hub-safety-2026-10-05/partA/live.txt` (Basic, no header → 403 even on an unknown path; with the header → passes; header without credentials → 401); `web/r135_csrf_test.go` (39 paths) | |
| Box operating-system security updates (Proxmox host, guest Debian, Docker engine) | agent v0.143.0, hub v0.133.0 | **PARTIAL — the GUEST and HOST Debian fast lanes and the DOCKER engine slow lane are PROVEN-LIVE (2026-10-04), with the System page, the fleet view and the alarms; the KERNEL lane is MISSING** | Guest: `audits/os-guest-lane-2026-10-04/`. Host + fleet + alarms: `audits/os-host-lane-2026-10-04/`. Docker + System page + crash guard: `audits/os-docker-crash-2026-10-04/` — live-restore on with the same container ids on every box; Docker 29.8.2 on both demo boxes; operator-approved Docker release; a signed undo and a signed ring-1 step; a replay refused; the crash guard restarted demo-hp twice and kept it off the third time. Design `architecture/11-os-updates.md` §5.8, §5.9, §8 | **No automatic undo** (guest: last night's backup; host: by-hand runbook; Docker: a signed undo job); existing boxes get root-owned files by the signed config bundle since 2026-10-04 (R-840 CLOSED; a box from before agent 0.143.0 needs one by-hand bootstrap — Tester 2: R-862; `audits/r840-config-bundle-2026-10-04/`); test approvals now end with the test (R-859); the agent's sudoers is root-equivalent (R-861); the kernel lane (R-836); facts reach the hub late after a boot (R-853). **2026-10-05 (agent v0.144.1):** R8 measures the real download (R-865); a killed pass still reports (R-868, live); the debug pass runs with the hub away (R-866, live); **a power cut mid-update was proven by day on demo-hp — the box came back by itself in 37 s, but the next pass fails until `dpkg --configure -a` is run by hand (R-876, P2, open)** — `audits/night-fixes-2026-10-05/`. **2026-10-05 afternoon (agent v0.145.0): R-876 FIXED and proven live — after a second crash mid-unpack the next pass repaired dpkg by itself (`REPAIR … journal=1`) and finished** — `audits/catchup-2026-10-05/partD/` |
| **An ENGLISH-SPEAKING household's first hour: download, install, pair, bind, claim, two apps** | controller **v0.259.0** + hub **v0.119.0** + ISO 1.29.0 + the whole catalog | **PROVEN-LIVE on 0.258.0 with one blocker; THE BLOCKER IS FIXED AND PROVEN, THE WALK IS NOT REPEATED** | `audits/DRILL-first-hour-en-0258-2026-09-20.md` — a fresh install 2026-09-20, one intervention (R-494), stop rule not reached. Then `audits/i18n-closing-2026-09-21/live/` — the three blockers fixed and each proven on a live box or in the operator's inbox: the claim page answers English through the real cookie path; the Backup page's tier names follow the language; and the setup mail carries **four plain-ASCII English words** where the drill's carried `képző-szkítia-ásatás`, one day apart in the same inbox. | **R-596, R-597 and R-598 are CLOSED.** What this row still does NOT claim: **the fixed journey has not been walked end to end by a stranger on a fresh install.** Three fixes proven at the endpoint are not an hour proven by a person, and this project's own rule is that fixes are not a journey (see the recovery-journey row). **Also not walked:** the recovery code (needs ep0), backup/restore/remove/power-cut (proven 2026-09-14), and the two Backup-page *warnings* themselves — guest 9201 is healthy and a healthy box renders none, so they are covered by handler render tests, not live. **Verdict: nothing known now stands between an English-speaking tester and their box — and that is a different sentence from "the walk passed".** |
@@ -140,8 +140,9 @@ mert a doboz vagy a vezérlő újraindult…"), kept until a run ends with every
before this, both pages said nothing and the synthesised „Utolsó adatbázis mentés … OK" was read off the fresh `.sql`
the cut run left beside last night's tars; that line now reads failed after a cut. Each restore point's time was
already its OLDEST part (the data block, v0.275.0) — so a torn unit is dated by its stale tars, never by its new dump.
*Live: unit-proven and red-proved; the live cut on 9202 was refused by the permission check and waits for the
operator (R-519 narrowed).*
*Live: PROVEN 2026-10-05 on 9202 (operator ruling 126): a run cut by a controller restart between bookstack's config
dump and its database volume → both pages carry the notice, the restore point reads the older volume's time, the next
complete run clears it (`audits/hub-db-offsite-2026-10-05/partD/r519/`; R-519 closed).*
### Lane 2 — the operator: guest and host recovery
@@ -0,0 +1,26 @@
## promtool, in pod/prometheus-55b675779d-t8c74, 2026-10-05T13:42:17Z
### green
SUCCESS
rc=0
### red: threshold 26h -> 260h
FAILED:
alertname: HubDBBackupStale, time: 1d2h40m,
exp:[
0:
Labels:{alertname="HubDBBackupStale", component="backup", instance="dooplex", severity="critical"}
Annotations:{description="No successful push of the hub DB snapshot to ep0 for >26h (daily at 02:30). Check `journalctl -u felhom-hub-db-backup.service`, the tunnel `systemctl status felhom-ep0-pbs-tunnel`, and that the hub logs `db snapshot written` at 02:00.", summary="The Felhom hub database has not reached ep0 for 26 h (R-173)"}
],
got:[]
command terminated with exit code 1
rc=0
### red: absent() removed
SUCCESS
### red (re-run): absent() removed from HubDBBackupStale — first attempt above did NOT apply (indent mismatch)
FAILED:
alertname: HubDBBackupStale, time: 40m,
Labels:{alertname="HubDBBackupStale", component="backup", severity="critical"}
got:[]
@@ -0,0 +1,52 @@
rule_files: [bf.yml]
evaluation_interval: 1m
tests:
# 1. the push stops: last success at t=0, never again → fires once 26 h + 30 m have passed
- interval: 5m
input_series:
- series: 'felhom_hub_db_backup_last_success_timestamp_seconds{instance="dooplex"}'
values: '0+0x400'
- series: 'felhom_hub_db_restore_test_last_success_timestamp_seconds{instance="dooplex"}'
values: '0+0x400'
alert_rule_test:
- eval_time: 26h
alertname: HubDBBackupStale
exp_alerts: []
- eval_time: 26h40m
alertname: HubDBBackupStale
exp_alerts:
- exp_labels: {severity: critical, component: backup, instance: dooplex}
exp_annotations:
summary: "The Felhom hub database has not reached ep0 for 26 h (R-173)"
description: "No successful push of the hub DB snapshot to ep0 for >26h (daily at 02:30). Check `journalctl -u felhom-hub-db-backup.service`, the tunnel `systemctl status felhom-ep0-pbs-tunnel`, and that the hub logs `db snapshot written` at 02:00."
# 2. healthy: the push succeeds every 24 h → never fires
- interval: 1h
input_series:
- series: 'felhom_hub_db_backup_last_success_timestamp_seconds{instance="dooplex"}'
values: '0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 86400 86400 86400 86400 86400 86400 86400 86400 86400 86400 86400 86400 86400 86400 86400 86400 86400 86400 86400 86400 86400 86400 86400 86400 172800 172800 172800'
- series: 'felhom_hub_db_restore_test_last_success_timestamp_seconds{instance="dooplex"}'
values: '0+0x50'
alert_rule_test:
- eval_time: 50h
alertname: HubDBBackupStale
exp_alerts: []
# 3. the metric never existed (script never ran) → fires on absent()
- interval: 5m
input_series:
- series: 'up{job="node"}'
values: '1+0x30'
alert_rule_test:
- eval_time: 40m
alertname: HubDBBackupStale
exp_alerts:
- exp_labels: {severity: critical, component: backup}
exp_annotations:
summary: "The Felhom hub database has not reached ep0 for 26 h (R-173)"
description: "No successful push of the hub DB snapshot to ep0 for >26h (daily at 02:30). Check `journalctl -u felhom-hub-db-backup.service`, the tunnel `systemctl status felhom-ep0-pbs-tunnel`, and that the hub logs `db snapshot written` at 02:00."
- eval_time: 2h
alertname: HubDBRestoreTestStale
exp_alerts:
- exp_labels: {severity: warning, component: backup}
exp_annotations:
summary: "The hub database copy on ep0 has not passed a restore test for 8 days (R-173)"
description: "The weekly restore test (Sun 04:30) has not succeeded for >8 days. Check `journalctl -u felhom-hub-db-restore-test.service`."
@@ -0,0 +1,5 @@
## alarm drill start 2026-10-05T13:51:46Z: success file moved aside (absent case)
backup_freshness.prom
fan_metrics.prom
felhom_hub_db_restore.prom
node_housekeeping.prom
@@ -0,0 +1,27 @@
## manual push via unit, 2026-10-05T13:50:30Z
Result=success
ExecMainStatus=0
2026-10-05T15:50:08+02:00 dooplex systemd[1]: Starting felhom-hub-db-backup.service - Felhom: push the hub DB snapshot to ep0 (R-173)...
2026-10-05T15:50:09+02:00 dooplex felhom-hub-db-backup[3036511]: felhom-hub-db-backup: snapshot hub-20261005T123037Z.db, 79 min old
2026-10-05T15:50:18+02:00 dooplex felhom-hub-db-backup[3036511]: felhom-hub-db-backup: checked: 369807360 bytes, integrity ok, 4 host(s)
2026-10-05T15:50:18+02:00 dooplex felhom-hub-db-backup[3037790]: Starting backup: [operator]:host/dooplex-hub/2026-10-05T13:50:18Z
2026-10-05T15:50:18+02:00 dooplex felhom-hub-db-backup[3037790]: Client name: dooplex
2026-10-05T15:50:18+02:00 dooplex felhom-hub-db-backup[3037790]: Starting backup protocol: Mon Oct 5 15:50:18 2026
2026-10-05T15:50:18+02:00 dooplex felhom-hub-db-backup[3037790]: Using encryption key from '/etc/felhom-hub-backup/enc.key'..
2026-10-05T15:50:18+02:00 dooplex felhom-hub-db-backup[3037790]: Encryption key fingerprint: b2:19:bf:36:3b:97:3d:6c
2026-10-05T15:50:18+02:00 dooplex felhom-hub-db-backup[3037790]: No previous manifest available.
2026-10-05T15:50:18+02:00 dooplex felhom-hub-db-backup[3037790]: Upload directory '/var/lib/felhom-hub-backup/stage' to 'dooplex-hub@pbs!push@127.0.0.1:18007:felhom-offsite' as hubdb.pxar.didx
2026-10-05T15:50:25+02:00 dooplex felhom-hub-db-backup[3037790]: hubdb.pxar: had to backup 352.676 MiB of 352.676 MiB (compressed 17.242 MiB) in 6.58 s (average 53.604 MiB/s)
2026-10-05T15:50:25+02:00 dooplex felhom-hub-db-backup[3037790]: Uploaded backup catalog (56 B)
2026-10-05T15:50:25+02:00 dooplex felhom-hub-db-backup[3037790]: Duration: 7.19s
2026-10-05T15:50:25+02:00 dooplex felhom-hub-db-backup[3037790]: End Time: Mon Oct 5 15:50:25 2026
2026-10-05T15:50:25+02:00 dooplex felhom-hub-db-backup[3036511]: felhom-hub-db-backup: pushed hub-20261005T123037Z.db to ep0 (ns operator) in 7 s
2026-10-05T15:50:25+02:00 dooplex felhom-hub-db-backup[3036511]: felhom-hub-db-backup: success signal written
2026-10-05T15:50:30+02:00 dooplex systemd[1]: felhom-hub-db-backup.service: Deactivated successfully.
2026-10-05T15:50:30+02:00 dooplex systemd[1]: Finished felhom-hub-db-backup.service - Felhom: push the hub DB snapshot to ep0 (R-173).
2026-10-05T15:50:30+02:00 dooplex systemd[1]: felhom-hub-db-backup.service: Consumed 9.566s CPU time, 584.8M memory peak.
## textfile
# HELP felhom_hub_db_backup_last_success_timestamp_seconds Last successful push of the hub DB snapshot to ep0 (R-173).
# TYPE felhom_hub_db_backup_last_success_timestamp_seconds gauge
felhom_hub_db_backup_last_success_timestamp_seconds 1791208225
felhom_hub_db_backup_last_success_bytes 369807360
@@ -0,0 +1 @@
gitea.dooplex.hu/admin/felhom-controller:0.296.0 Up 6 seconds (healthy)
@@ -0,0 +1,3 @@
## deploy bookstack 2026-10-05T13:53:26Z (fields: DOMAIN APP_KEY DB_PASSWORD ADMIN_PASSWORD; SUBDOMAIN = catalog default)
HTTP 202
@@ -0,0 +1,36 @@
## run 1 (baseline) start 2026-10-05T14:04:36Z
HTTP 200
{"ok":true,"data":{"enabled":true,"running":true}}
HTTP 200
## run 1 end 2026-10-05T14:05:17Z: {"ok":true,"data":{"db_dump":{"count":2,"duration":"32.784678377s","last_run":"2026-10-05T14:05:12.109790785Z","success":true},"enabled":true,"running":false}}
2026/10/05 14:04:44 data_versions.go:131: [DEBUG] [backup] bookstack: stamped volume-dumps/bookstack_bookstack_config.tar (70144 B) with pins [lscr.io/linuxserver/bookstack:26.09.1@sha256:99cd1f5707c1911afad213adec5c9739763b76f843d1477142231834ecdcb6f7 mariadb:12.3@sha256:805c8e104bd563d5bfa24fadd3f31cd419ea859cb5277f32b5dbf2db714f9ed1]
2026/10/05 14:04:44 backup.go:848: [DEBUG] [backup] Dumping volume bookstack_bookstack_db_data for bookstack
2026/10/05 14:04:45 backup.go:873: [INFO] [backup] Volume dump: bookstack/bookstack_bookstack_db_data → 153.4 MB
2026/10/05 14:04:45 data_versions.go:131: [DEBUG] [backup] bookstack: stamped volume-dumps/bookstack_bookstack_db_data.tar (160868864 B) with pins [lscr.io/linuxserver/bookstack:26.09.1@sha256:99cd1f5707c1911afad213adec5c9739763b76f843d1477142231834ecdcb6f7 mariadb:12.3@sha256:805c8e104bd563d5bfa24fadd3f31cd419ea859cb5277f32b5dbf2db714f9ed1]
2026/10/05 14:04:45 backup.go:990: [INFO] [backup] Restarting bookstack after volume dump
2026/10/05 14:04:51 backup.go:974: [INFO] [backup] Stopping paperless-ngx for safe volume dump
2026/10/05 14:04:58 backup.go:848: [DEBUG] [backup] Dumping volume paperless-ngx_paperless_redis_data for paperless-ngx
2026/10/05 14:04:58 backup.go:873: [INFO] [backup] Volume dump: paperless-ngx/paperless-ngx_paperless_redis_data → 7.4 MB
2026/10/05 14:04:58 data_versions.go:131: [DEBUG] [backup] paperless-ngx: stamped volume-dumps/paperless-ngx_paperless_redis_data.tar (7792128 B) with pins [ghcr.io/paperless-ngx/paperless-ngx:2.20.15 postgres:18-alpine redis:7-alpine]
2026/10/05 14:04:58 backup.go:848: [DEBUG] [backup] Dumping volume paperless-ngx_paperless_data for paperless-ngx
2026/10/05 14:04:59 backup.go:873: [INFO] [backup] Volume dump: paperless-ngx/paperless-ngx_paperless_data → 5.8 MB
2026/10/05 14:04:59 data_versions.go:131: [DEBUG] [backup] paperless-ngx: stamped volume-dumps/paperless-ngx_paperless_data.tar (6052352 B) with pins [ghcr.io/paperless-ngx/paperless-ngx:2.20.15 postgres:18-alpine redis:7-alpine]
2026/10/05 14:04:59 backup.go:848: [DEBUG] [backup] Dumping volume paperless-ngx_paperless_postgres_data for paperless-ngx
2026/10/05 14:04:59 backup.go:873: [INFO] [backup] Volume dump: paperless-ngx/paperless-ngx_paperless_postgres_data → 68.8 MB
2026/10/05 14:04:59 data_versions.go:131: [DEBUG] [backup] paperless-ngx: stamped volume-dumps/paperless-ngx_paperless_postgres_data.tar (72107008 B) with pins [ghcr.io/paperless-ngx/paperless-ngx:2.20.15 postgres:18-alpine redis:7-alpine]
2026/10/05 14:04:59 backup.go:990: [INFO] [backup] Restarting paperless-ngx after volume dump
2026/10/05 14:05:11 backup.go:974: [INFO] [backup] Stopping privatebin for safe volume dump
2026/10/05 14:05:11 backup.go:848: [DEBUG] [backup] Dumping volume privatebin_privatebin_data for privatebin
2026/10/05 14:05:12 backup.go:873: [INFO] [backup] Volume dump: privatebin/privatebin_privatebin_data → 1.5 KB
2026/10/05 14:05:12 data_versions.go:131: [DEBUG] [backup] privatebin: stamped volume-dumps/privatebin_privatebin_data.tar (1536 B) with pins [privatebin/pdo:2.0.6@sha256:4c141b2326f8b353598ce9ce7507a9cfecf2dad5c60a39fea903d430e296d8f5]
2026/10/05 14:05:12 backup.go:986: [INFO] [backup] privatebin NOT restarted after the volume dump: the household stopped it meanwhile
2026/10/05 14:05:12 backup.go:684: [INFO] [backup] App-data backup completed: 2 databases (483.2 KB total), 3 volume dump(s) (32.785s)
2026/10/05 14:05:12 recovery_unit.go:289: [INFO] [backup] Recovery unit captured for bookstack → /mnt/sys_drive/felhom-data/backups/primary/bookstack (images=2, secrets-referenced=3, data_keys=0, portable-carried=2/2, withheld=1)
2026/10/05 14:05:12 recovery_unit.go:289: [INFO] [backup] Recovery unit captured for paperless-ngx → /mnt/felhom-drives/scratch_hdd/userdata/paperless-ngx/backups/primary/paperless-ngx (images=3, secrets-referenced=3, data_keys=0, portable-carried=2/2, withheld=1)
2026/10/05 14:05:12 recovery_unit.go:289: [INFO] [backup] Recovery unit captured for privatebin → /mnt/sys_drive/felhom-data/backups/primary/privatebin (images=1, secrets-referenced=0, data_keys=0, portable-carried=0/0, withheld=0)
/var/lib/felhom/docker/volumes/felhom-controller-data/_data/data/appdata-run.json
/var/lib/docker/volumes/felhom-controller-data/_data/data/appdata-run.json
@@ -0,0 +1,12 @@
## after run 1, 2026-10-05T14:05:34Z
### run record
{"running":false,"started_at":"0001-01-01T00:00:00Z","interrupted":"0001-01-01T00:00:00Z"}
### restore points (bookstack)
{"ok":true,"data":[{"time":"2026-10-05T14:04:39Z","short_id":"helyi","tier":1,"drive_label":"Belső SSD (rendszer)"}]}
HTTP 200
### volume dump file times
total 157172
-rw-r--r-- 1 root root 70144 2026-10-05T14:04:44 bookstack_bookstack_config.tar
-rw-r--r-- 1 root root 160868864 2026-10-05T14:04:44 bookstack_bookstack_db_data.tar
@@ -0,0 +1,34 @@
## run 2 start 2026-10-05T14:05:50Z
HTTP 200
### watcher
match 2026-10-05T14:05:58.856456447Z
restarted 2026-10-05T14:05:59.669364637Z rc=0
Up 10 seconds (healthy)
### record at the moment of the cut
{"running":true,"started_at":"2026-10-05T14:05:53.282495446Z","interrupted":"0001-01-01T00:00:00Z"}
### record after restart
{"running":false,"started_at":"2026-10-05T14:05:53.282495446Z","interrupted":"2026-10-05T14:05:53.282495446Z"}
### controller log around the cut (previous + new process)
2026/10/05 14:05:53 backup.go:557: [INFO] [backup] Starting database dump run
2026/10/05 14:05:53 dbdump.go:208: [INFO] [backup] Discovered 2 databases
2026/10/05 14:05:53 backup.go:596: [INFO] [backup] Discovered 2 database(s): paperless-postgres(postgres), bookstack-db(mariadb)
2026/10/05 14:05:53 dbdump.go:410: [INFO] [backup] DB dump: paperless-postgres → paperless-ngx-postgres.sql (428.1 KB, 321ms, 72 tables)
2026/10/05 14:05:54 dbdump.go:410: [INFO] [backup] DB dump: bookstack-db → bookstack-mariadb.sql (55.9 KB, 310ms, 41 tables)
2026/10/05 14:05:54 backup.go:974: [INFO] [backup] Stopping bookstack for safe volume dump
2026/10/05 14:05:58 backup.go:873: [INFO] [backup] Volume dump: bookstack/bookstack_bookstack_config → 49.0 KB
2026/10/05 14:05:59 main.go:340: [INFO] felhom-controller 0.296.0 starting (customer: demo-hp, domain: enkisfelhom.hu)
2026/10/05 14:05:59 appstop_marker.go:278: [WARN] [appstop] crash recovery: an app-data backup (volume dump) (op "volume-dump:bookstack") was interrupted and left 1 app(s) stopped — restarting them: [bookstack]
2026/10/05 14:05:59 manager.go:1235: [INFO] [stacks] Starting stack: bookstack
2026/10/05 14:06:06 appstop_marker.go:304: [INFO] [appstop] crash recovery: restarted bookstack after the interrupted an app-data backup (volume dump)
2026/10/05 14:06:06 sync.go:117: [INFO] [sync] Starting catalog sync (repo: https://gitea.dooplex.hu/admin/app-catalog-felhom.eu.git, interval: 15m0s)
2026/10/05 14:06:06 sync.go:208: [INFO] [sync] Starting catalog sync
2026/10/05 14:06:06 restore_record_wiring.go:35: [INFO] [backup] restore record wired: /opt/docker/felhom-controller/data/restore-status.json (interrupted at startup: false)
2026/10/05 14:06:06 run_record_wiring.go:27: [WARN] [backup] the app-data backup run started 2026-10-05T14:05:53Z was cut off by the stop — the backup pages say so until the next complete run (R-519)
2026/10/05 14:06:06 scheduler.go:223: [INFO] [scheduler] Starting scheduler with 19 jobs
2026/10/05 14:06:06 backup.go:1179: [INFO] [backup] Found 3 DB dump files across drives
2026/10/05 14:06:06 dbdump.go:208: [INFO] [backup] Discovered 2 databases
2026/10/05 14:06:06 [INFO] [backup] Discovered app data: 3 apps
2026/10/05 14:06:06 recovery_unit.go:289: [INFO] [backup] Recovery unit captured for bookstack → /mnt/sys_drive/felhom-data/backups/primary/bookstack (images=2, secrets-referenced=3, data_keys=0, portable-carried=2/2, withheld=1)
2026/10/05 14:06:06 recovery_unit.go:289: [INFO] [backup] Recovery unit captured for paperless-ngx → /mnt/felhom-drives/scratch_hdd/userdata/paperless-ngx/backups/primary/paperless-ngx (images=3, secrets-referenced=3, data_keys=0, portable-carried=2/2, withheld=1)
2026/10/05 14:06:06 backup.go:1250: [INFO] [backup] Backup status cache refreshed
2026/10/05 14:06:09 manager.go:1601: [INFO] [stacks] bookstack lscr.io/linuxserver/bookstack:26.09.1@sha256:99cd1f5707c1911afad213adec5c9739763b76f843d1477142231834ecdcb6f7 running Up 3 seconds (health: starting)
@@ -0,0 +1,23 @@
## after the cut, 2026-10-05T14:06:23Z
### GET /backups
data-interrupted-run present: 1
<div class="alert alert-warning" data-interrupted-run="true">A legutóbbi mentés (2026-10-05 16:05) megszakadt, mert a doboz vagy a vezérlő újraindult. Amit nem fejezett be, annak a korábbi mentése maradt meg — minden visszaállítási pont annyira friss, amennyire a legrégebbi része. A következő teljes mentés után ez az üzenet eltűnik.
HTTP 200
### GET /backups/apps
data-interrupted-run present: 1
<div class="alert alert-warning" data-interrupted-run="true">A legutóbbi mentés (2026-10-05 16:05) megszakadt, mert a doboz vagy a vezérlő újraindult. Amit nem fejezett be, annak a korábbi mentése maradt meg — minden visszaállítási pont annyira friss, amennyire a legrégebbi része. A következő teljes mentés után ez az üzenet eltűnik.
HTTP 200
### negative control: a marker that must NOT be on the page
0
### restore points (bookstack)
{"ok":true,"data":[{"time":"2026-10-05T14:04:44Z","short_id":"helyi","tier":1,"drive_label":"Belső SSD (rendszer)"}]}
### dump file times
total 314252
-rw-r--r-- 1 root root 50176 2026-10-05T14:05:58 bookstack_bookstack_config.tar
-rw-r--r-- 1 root root 160868864 2026-10-05T14:04:44 bookstack_bookstack_db_data.tar
-rw-r--r-- 1 root root 160867328 2026-10-05T14:05:58 bookstack_bookstack_db_data.tar.tmp
drwxr-xr-x 2 root root 4096 2026-10-05T14:05:54 db-dumps
drwxr-xr-x 2 root root 4096 2026-10-05T14:05:58 volume-dumps
### bookstack after
bookstack Up 27 seconds (healthy)
bookstack-db Up 32 seconds (healthy)
@@ -0,0 +1,12 @@
## run 3 (complete) start 2026-10-05T14:06:44Z
HTTP 200
## run 3 end 2026-10-05T14:07:28Z: {"ok":true,"data":{"db_dump":{"count":2,"duration":"33.315240638s","last_run":"2026-10-05T14:07:20.615183031Z","success":true},"enabled":true,"running":false}}
2026/10/05 14:05:12 backup.go:684: [INFO] [backup] App-data backup completed: 2 databases (483.2 KB total), 3 volume dump(s) (32.785s)
2026/10/05 14:07:20 backup.go:684: [INFO] [backup] App-data backup completed: 2 databases (483.9 KB total), 3 volume dump(s) (33.315s)
{"running":false,"started_at":"0001-01-01T00:00:00Z","interrupted":"0001-01-01T00:00:00Z"}
total 157152
-rw-r--r-- 1 root root 50688 2026-10-05T14:06:52 bookstack_bookstack_config.tar
-rw-r--r-- 1 root root 160867328 2026-10-05T14:06:53 bookstack_bookstack_db_data.tar
/backups data-interrupted-run: 0
/backups/apps data-interrupted-run: 0
restore points: {"ok":true,"data":[{"time":"2026-10-05T14:06:47Z","short_id":"helyi","tier":1,"drive_label":"Belső SSD (rendszer)"}]}
@@ -0,0 +1,8 @@
## teardown 2026-10-05T14:07:48Z
stop: HTTP 200
remove: HTTP 200
containers: 0
volumes: 0
stackdir: none
backups: none
/root/.dbody
@@ -0,0 +1,9 @@
## runbook §3 drill 2026-10-05T14:09:41Z (scratch /var/lib/felhom-hub-backup/sec3.65qk, root 0700)
step 1 restore host/dooplex-hub/2026-10-05T13:50:18Z
restore complete (352.676 MiB processed in 3.9s, average 90.961 MiB/s)
-rw------- 369807360 hub.db
step 2 the seal key from Secret/offsite-secret-key → /var/lib/felhom-hub-backup/sec3.65qk/k (0600, not printed)
key file bytes: 64
step 3 hubdb-check (same key): hosts=4 console_passwords_opened=4 failed=0 absent=0 rc=0
control (a random key): hosts=4 console_passwords_opened=0 failed=4 absent=0 hubdb-check: FAILED: not every console password opened with this key rc=0
scratch shredded: gone
@@ -0,0 +1,7 @@
### R7 hubdb-check counts a failed open as opened
=== RUN TestCheck_WrongKeyOpensNothing
main_test.go:63: got {hosts:2 opened:1 failed:0 absent:1}, want opened=0 failed=1
--- FAIL: TestCheck_WrongKeyOpensNothing (0.03s)
FAIL
FAIL gitea.dooplex.hu/admin/felhom-hub/cmd/hubdb-check 0.040s
FAIL
@@ -0,0 +1,17 @@
## snapshot list on ep0 with the READ-ONLY token, 2026-10-05T13:50:40Z
+=======================================+=============+=====================================+
| snapshot | size | files |
+=======================================+=============+=====================================+
| host/dooplex-hub/2026-10-05T13:50:18Z | 352.677 MiB | catalog.pcat1 hubdb.pxar index.json |
+=======================================+=============+=====================================+
unit rc=0
Result=success
2026-10-05T15:50:41+02:00 dooplex felhom-hub-db-restore-test[3040582]: felhom-hub-db-restore-test: restoring host/dooplex-hub/2026-10-05T13:50:18Z
2026-10-05T15:50:46+02:00 dooplex felhom-hub-db-restore-test[3040582]: felhom-hub-db-restore-test: checked: integrity ok, 4 host(s), 4 sealed console password(s), 0 readable
2026-10-05T15:50:46+02:00 dooplex felhom-hub-db-restore-test[3040582]: felhom-hub-db-restore-test: success signal written
# HELP felhom_hub_db_restore_test_last_success_timestamp_seconds Last successful restore test of the hub DB copy on ep0 (R-173).
# TYPE felhom_hub_db_restore_test_last_success_timestamp_seconds gauge
felhom_hub_db_restore_test_last_success_timestamp_seconds 1791208246
.cache
.kube
stage
@@ -0,0 +1,12 @@
## 2026-10-05T13:51:07Z token limits (each line = the tool output)
push forget : Error: permission check failed - missing Datastore.Modify|Datastore.Prune on /datastore/felhom-offsite/operator
restore forget : Error: permission check failed - missing Datastore.Modify|Datastore.Prune on /datastore/felhom-offsite/operator
restore backup : Error: missing permissions 'Datastore.Backup' on '/datastore/felhom-offsite/operator'
push list ns root (households): Error: permission check failed - missing Datastore.Audit|Datastore.Backup on /datastore/felhom-offsite
restore list ns root (households): Error: permission check failed - missing Datastore.Audit|Datastore.Backup on /datastore/felhom-offsite
push restore own copy: Error: missing key - manifest was created with key b2:19:bf:36:3b:97:3d:6c exit-file=absent
## paper-key proof: restore with a key file rebuilt from the data field only
restore rc=0
integrity: ok hosts: 4
## and with NO key: Error: missing key - manifest was created with key b2:19:bf:36:3b:97:3d:6c
push restore own copy WITH key: restore complete (352.676 MiB processed in 4.6s, average 76.793 MiB/s) file=PRESENT
+10
View File
@@ -26,6 +26,16 @@
---
## 2026-10-05 (evening) — the hub database off DooPlex (hub v0.136.0; `09` rulings 125–127)
The full text of every row below: `git show cea8502f:documentation/backlog/OPEN-ITEMS.md`.
| Row | What | Closed | Evidence |
|---|---|---|---|
| **R-519** | **After a backup torn by a power cut, a restore point carried the new database dump's time over the previous run's files, and no screen said the run was interrupted (P2).** Fixed in controller v0.296.0 (page notice + synthesised status; dating by the oldest part since v0.275.0). **Live on 9202 (operator ruling 126):** a complete run, then a run cut by `docker restart felhom-controller` 0.8 s after bookstack's config dump and before its database volume; afterwards both /backups and /backups/apps carry `data-interrupted-run` („A legutóbbi mentés (2026-10-05 16:05) megszakadt …", negative control 0), the restore point reads 14:04:44Z = the run-1 database volume (its oldest part; config 14:05:58, SQL 14:05:54), the controller restarted bookstack itself; the next complete run cleared the notice and replaced the torn `.tar.tmp`. The throwaway bookstack was removed through the product (0 containers, volumes, folders, backups). | CLOSED 2026-10-05 — FIXED controller v0.296.0, proven live | `audits/hub-db-offsite-2026-10-05/partD/r519/`; `internal/backup/run_record_test.go` |
---
## 2026-10-05 (late afternoon) — the hub's own safety, boxes left behind, the agent's root grants (hub v0.135.0, controller v0.296.0, agent v0.146.1, golden 0.296.0; CC decisions 119–124)
The full text of every row below: `git show 9bb45eaa:documentation/backlog/OPEN-ITEMS.md` (R-880 was opened and closed in this session).
File diff suppressed because one or more lines are too long
@@ -1,16 +1,16 @@
# Runbook — the hub database in a backup that is NOT on DooPlex (R-173, R-232) — PROPOSED, needs the operator's go
# Runbook — the hub database in a backup that is NOT on DooPlex (R-173, R-232) — IN FORCE since 2026-10-05
> **Status: PROPOSED 2026-10-05. Nothing here has been done.** DooPlex and ep0 are protected; every step below changes
> one of them, so each waits for the operator's go (the decision is in `STATUS.md`). The readings this plan rests on:
> `audits/hub-safety-2026-10-05/partC/readings.txt` (read only).
> **Status: IN FORCE 2026-10-05** (option A, `09` decision 125). Steps 0–7 DONE, each marked below; §3 is a tested
> procedure. Evidence: `audits/hub-db-offsite-2026-10-05/` (part A–D). The readings the plan rested on:
> `audits/hub-safety-2026-10-05/partC/readings.txt`. **Corrections found while doing it are marked „Corrected".**
## 1. What is true today (measured 2026-10-05)
| Question | Answer |
|---|---|
| Where the hub database lives | `/data/hub.db` (+ `-wal`, `-shm`) in the hub pod, PVC `hub-data` (Longhorn, 1 Gi, replicas on DooPlex's `sdb1`). 357 MiB. |
| Where the hub database lives | `/data/hub.db` (+ `-wal`, `-shm`) in the hub pod, PVC `hub-data` (Longhorn, **2 Gi since 2026-10-05**, was 1 Gi; replicas on DooPlex's `sdb1`). 357 MiB; a snapshot is 353 MiB. |
| Is it in a backup? | **Yes, but only on DooPlex.** Longhorn's `backup-daily` (04:00) and `backup-weekly` (Sun 05:00), `retain=1`, write to `nfs://192.168.0.180:/mnt/5_hdd/backup/longhorn-pvc` — DooPlex's own `sda1`. Last: 2026-10-05 02:06 UTC, Completed. |
| Why R-173 said "excluded" | The PVC carries `recurring-job-group.longhorn.io/default: disabled` (git, `manifests/hub.yaml`, commit `868e8465` of 2026-02-16, no reason given). The live Longhorn **Volume** carries `enabled` — set by hand at some point, so the backups run. **This is drift:** the PVC label is the source of truth (Longhorn syncs PVC → Volume), so the backups can stop at any sync without anyone seeing it. |
| Why R-173 said "excluded" | The PVC carries `recurring-job-group.longhorn.io/default: disabled` (git, `manifests/hub.yaml`, commit `868e8465` of 2026-02-16, no reason given). The live Longhorn **Volume** carries `enabled` — set by hand at some point, so the backups run. **This is drift:** the PVC label is the source of truth (Longhorn syncs PVC → Volume), so the backups can stop at any sync without anyone seeing it. **Corrected 2026-10-05:** the Volume kept `enabled` through every sync since February while the PVC said `disabled` — Longhorn did NOT copy the PVC label down. The PVC now says `enabled` too (Step 1), so the two agree; which one Longhorn reads was not measured. |
| What DooPlex's own backup covers | `dooplex-backup.timer` (03:19): k3s state, k8s Secrets (GPG files), Gitea mirrors, user data, PostgreSQL dumps — **all onto `sda1`, the same machine.** Nothing leaves DooPlex (`audits/RECON-dooplex-backup-2026-08-06.md`, R-232). |
| What tells anyone a backup failed | **Nothing.** `NOTIFY_WEBHOOK_URL` is commented out, so `notify_failure` is a no-op. No Prometheus rule watches a Longhorn backup's success or age, nor `dooplex-backup.service`. |
| What the database holds | Box→hub API keys, customer configs (incl. the owner passphrase), escrow custody blobs (opaque), the PBS-DR token values, the off-site sub-account passwords and — since hub v0.135.0 — the console passwords **sealed** under `OFFSITE_SECRET_KEY`. |
@@ -22,7 +22,7 @@ ep0 is already off-site (Hetzner), already runs PBS, and DooPlex already reaches
(127.0.0.1:18007). ep0's datastore is also pulled back to DooPlex nightly (`ep0-copy`), so the copy exists in two places,
one of them off DooPlex. The copy is encrypted on DooPlex with a key ep0 never sees.
### Step 0 — the keys go off DooPlex first (operator, at the keyboard, 5 min)
### Step 0 — the keys go off DooPlex first (operator, at the keyboard, 5 min) — DONE 2026-10-05 (operator: both saved)
Without these, every later step backs up something nobody can open after a DooPlex loss.
@@ -33,33 +33,65 @@ sudo kubectl -n felhom-system get secret offsite-secret-key -o jsonpath='{.data.
sudo proxmox-backup-client key paperkey /etc/felhom-hub-backup/enc.key --output-format text
```
### Step 1 — end the label drift (CC, a felhom.eu commit + ArgoCD sync; reversible)
**The `data` field of that output is enough** (the key has no passphrase, `kdf: null`): a key file rebuilt from it alone —
`{"kdf": null, "created": "<any RFC 3339 time>", "modified": "<same>", "data": "<saved value>"}` — restored and
decrypted the copy on 2026-10-05 (`audits/hub-db-offsite-2026-10-05/partD/token-limits-and-paperkey.txt`). Note:
`proxmox-backup-client key show` prints a fingerprint only when the file stores one, so it cannot check a rebuilt key —
a restore can.
### Step 1 — end the label drift (CC, a felhom.eu commit + ArgoCD sync; reversible) — DONE 2026-10-05
`manifests/hub.yaml`: `recurring-job-group.longhorn.io/default: disabled` → `enabled`. Sync. Check:
`sudo kubectl -n felhom-system get pvc hub-data -o jsonpath='{.metadata.labels}'` and the Volume label both read `enabled`.
This keeps today's on-DooPlex copy alive; it is not the off-site copy.
### Step 2 — a write-only place on ep0 (on ep0, root; the operator's go for an ep0 change)
Done with the volume growth to 2 Gi (two snapshots of 353 MiB do not fit in 1 Gi). **The online growth failed** — Longhorn's
`instance-manager` (116 days up) called a host PID that no longer existed (`nsenter: cannot open /host/proc/196610/ns/mnt`);
an offline growth was impossible too (the expansion holds its own attachment ticket). The operator approved restarting
the instance-manager: all 77 DooPlex volumes back `attached/healthy` in 110 s, the volume grew, and one app (zipline, on
`:latest`) came back on a newer release that refused its database — pinned to 4.7.0 (homelab-manifests). Evidence:
`audits/hub-db-offsite-2026-10-05/partA/step1-*.txt`.
### Step 2 — a write-only place on ep0 (on ep0, root; the operator's go for an ep0 change) — DONE 2026-10-05
What was run (PBS 4.2.8 on ep0; the proposal's commands were wrong in three places, corrected here):
```bash
proxmox-backup-manager user create dooplex-hub@pbs --comment "DooPlex pushes the hub DB (R-173)"
proxmox-backup-manager user generate-token dooplex-hub@pbs push # the secret → a 0600 file on DooPlex, file → file
# namespace for operator data, apart from the households' namespaces
proxmox-backup-client namespace create operator --repository 'root@pam@127.0.0.1:8007:felhom-offsite'
proxmox-backup-manager acl update /datastore/felhom-offsite/operator DatastoreBackup --auth-id 'dooplex-hub@pbs!push'
# retention on ep0 (the server prunes; the pushing token cannot delete — DatastoreBackup has no Prune)
proxmox-backup-manager prune-job create prune-operator-hubdb --store felhom-offsite --ns operator \
--schedule 'daily 03:45' --keep-daily 14 --keep-weekly 8
proxmox-backup-manager user create dooplex-hub@pbs --comment "..." # no password: cannot log in
proxmox-backup-debug api create /admin/datastore/felhom-offsite/namespace --name operator
# (the CLI crashes AFTER creating it, printing the result — 'not implemented'; check it exists, don't re-run)
# two tokens, each secret file → file into a root 0600 file on DooPlex, never printed:
ssh root@<ep0> "proxmox-backup-debug api create /access/users/dooplex-hub@pbs/token/push --output-format json" \
| python3 -c '<print json["value"]>' | sudo sh -c 'umask 077; cat > /etc/felhom-hub-backup/token-push'
# (same for token/restore → token-restore; `user generate-token` has no --output-format)
P=/datastore/felhom-offsite/operator
proxmox-backup-manager acl update $P DatastoreBackup --auth-id dooplex-hub@pbs # a token's rights are cut
proxmox-backup-manager acl update $P DatastoreReader --auth-id dooplex-hub@pbs # down by its user's rights
proxmox-backup-manager acl update $P DatastoreBackup --auth-id 'dooplex-hub@pbs!push'
proxmox-backup-manager acl update $P DatastoreReader --auth-id 'dooplex-hub@pbs!restore'
proxmox-backup-manager prune-job create prune-operator-hubdb --store felhom-offsite --ns operator --max-depth 0 \
--schedule '03:45' --keep-daily 14 --keep-weekly 8 # 'daily 03:45' is not a PBS calendar event
```
### Step 3 — a consistent snapshot of the live database (CC, a hub release)
Measured (`partB/`, `partD/token-limits-and-paperkey.txt`): neither token can forget a snapshot or list the datastore
root (the households); the restore token cannot write. **Corrected: the push token CAN restore its own copies** — PBS
lets a backup's owner read it back (`DatastoreBackup` = `Datastore.Backup`, owner-scoped). It reaches only `operator`,
and everything it can read is encrypted with a key ep0 never sees. The households' two prune jobs and the Sunday GC
are unchanged (before/after in `partB/`).
### Step 3 — a consistent snapshot of the live database (CC, a hub release) — DONE, hub v0.136.0 (`05` §16.3)
`hub.db` is in WAL mode and is written every few seconds; copying the three files is not one point in time. The hub
gets a nightly `VACUUM INTO '/data/snapshots/hub-<UTC date>.db'` (keeps 2, logs size and duration) — one SQLite
statement, consistent by construction, WAL-aware. **Needs a hub release** (filed under R-173). No `sqlite3` exists in the
hub image, so the copy must be made by the hub itself.
statement, consistent by construction, WAL-aware. No `sqlite3` exists in the hub image, so the copy is made by the hub
itself. **Corrected:** one snapshot took 44 s on the live volume (0.63 s on a local scratch copy), 353 MiB.
### Step 4 — the push (on DooPlex, root; `felhom-hub-db-backup.service` + `.timer` 02:30, CC writes, operator approves)
### Step 4 — the push (on DooPlex, root; `felhom-hub-db-backup.service` + `.timer` 02:30) — DONE 2026-10-05
**The real script is `scripts/hub-db-backup/felhom-hub-db-backup`** (versioned, R-231; installed by `install.sh`;
15 tests in `test_hub_db_backup.py`, run by hand — not in CI). It adds to the sketch below: it refuses a snapshot older
than 26 h (the hub stopped snapshotting), a copy whose size differs from the pod's file, and a copy with no hosts; it
encrypts with `--crypt-mode encrypt`. The sketch, as proposed:
```bash
#!/bin/sh -eu
@@ -77,7 +109,10 @@ echo "felhom_hub_db_backup_last_success_timestamp_seconds $(date +%s)" > /var/li
&& mv /var/lib/node_exporter/textfile_collector/felhom_hub_db.prom.$$ /var/lib/node_exporter/textfile_collector/felhom_hub_db.prom
```
### Step 5 — the restore test (weekly, Sun 04:30, same unit family)
### Step 5 — the restore test (weekly, Sun 04:30, same unit family) — DONE 2026-10-05
**The real script is `scripts/hub-db-backup/felhom-hub-db-restore-test`**, with the READ-ONLY token. It also refuses a
newest copy older than 50 h. The sketch, as proposed:
```bash
T=$(mktemp -d); chmod 700 "$T"
@@ -90,9 +125,12 @@ shred -u "$T/hub.db"*; rmdir "$T"
echo "felhom_hub_db_restore_test_last_success_timestamp_seconds $(date +%s)" > …/felhom_hub_db_restore.prom # same tmp+mv
```
The push token needs `DatastoreReader` on `operator` too for the restore (or a second, read-only token — cleaner).
Done with a second, read-only token (`token-restore`).
### Step 6 — the alarm (homelab-manifests `prometheus-rules`, then `POST /-/reload` — the Prometheus there has no reloader)
### Step 6 — the alarm (homelab-manifests `prometheus-rules`, then `POST /-/reload` — the Prometheus there has no reloader) — DONE 2026-10-05
In the `backup-freshness` group; `promtool test rules` proves it (`partC/bf_test.yml`, two red-proofs). The Prometheus
Deployment is OutOfSync in ArgoCD for a reason unrelated to this; only the rules ConfigMap was synced.
```yaml
- alert: HubDBBackupStale
@@ -109,18 +147,46 @@ The push token needs `DatastoreReader` on `operator` too for the restore (or a s
Both reach the existing `email-notifications` receiver. `absent()` makes "the script never ran" an alarm too — an empty
log is not a success.
### Step 7 — prove it once (CC, with the operator's go)
### Step 7 — prove it once (CC, with the operator's go) — DONE 2026-10-05 (`partD/`)
Run the unit by hand; read the snapshot on ep0 (`proxmox-backup-client snapshot list --ns operator`); run the restore
test by hand; stop the timer for a day on purpose and see `HubDBBackupStale` mail arrive (positive observable), then
start it again.
## 3. Bringing the hub back from this copy (the procedure the plan exists for)
## 3. Bringing the hub back from this copy (the procedure the plan exists for) — TESTED 2026-10-05
1. A k3s with the `felhom` ArgoCD app, and **`Secret/offsite-secret-key` recreated with the SAME value** (Step 0 copy).
2. Restore the newest snapshot (Step 5's first command, with the paper key), scale `deploy/hub` to 0, copy `hub.db` into
the PVC (no `-wal`/`-shm` — the snapshot is a whole database), scale to 1. The log line
`console passwords sealed at rest (0 legacy plaintext row(s) sealed now)` and a working reveal prove the key matches.
Steps 1–3 were run on 2026-10-05 against the real copy on ep0 (`audits/hub-db-offsite-2026-10-05/partD/restore-procedure/drill.txt`):
4 hosts, **4 of 4 console passwords opened with the saved seal key, 0 of 4 with a random key**. Steps 4–5 (into a live
PVC) were NOT run — that needs the hub down; they are the ordinary scale-copy-scale.
What you need, all from the password manager: the seal key (`OFFSITE_SECRET_KEY`), the backup key's `data` field, and
the read-only token (or ep0 root to mint a new one: Step 2).
1. **The backup key file.** On the machine doing the restore, as root, `umask 077`, write
`{"kdf": null, "created": "2026-01-01T00:00:00+00:00", "modified": "2026-01-01T00:00:00+00:00", "data": "<data>"}`
to `enc.key` (Step 0 note). A copy of DooPlex's `/etc/felhom-hub-backup/enc.key` works as is.
2. **Restore the newest copy** (from any machine that reaches ep0's PBS on 8007 — DooPlex uses the tunnel 127.0.0.1:18007):
```bash
export PBS_PASSWORD_FILE=<token-restore file> PBS_FINGERPRINT=<ep0 cert fingerprint, /etc/felhom-hub-backup/env>
R='dooplex-hub@pbs!restore@<ep0>:8007:felhom-offsite'
proxmox-backup-client snapshot list host/dooplex-hub --ns operator --repository "$R" # pick the newest
proxmox-backup-client restore host/dooplex-hub/<time> hubdb.pxar ./out --ns operator --keyfile enc.key --repository "$R"
sqlite3 -readonly out/hub.db 'PRAGMA integrity_check' # must print: ok
```
If DooPlex's `ep0-copy` datastore survived, the same copy is there too (pulled nightly).
3. **Prove the seal key matches BEFORE putting the copy in place** — on a COPY of `out/hub.db` (the check migrates it):
```bash
cd felhom.eu/hub && go build -o hubdb-check ./cmd/hubdb-check
printf '%s' "<OFFSITE_SECRET_KEY>" > k; chmod 600 k # from the password manager — not on a command line in a shared shell
./hubdb-check copy-of-hub.db k # want: hosts=N console_passwords_opened=N failed=0; exit 0
```
`failed>0` means the wrong seal key: the hub would start but could open no console password (`05` §16.2).
4. **A k3s with the `felhom` ArgoCD app**, and `Secret/offsite-secret-key` recreated with the SAME value:
`kubectl -n felhom-system create secret generic offsite-secret-key --from-file=OFFSITE_SECRET_KEY=k`.
5. **Into the PVC:** scale `deploy/hub` to 0; put `out/hub.db` into the volume as `/data/hub.db` (a helper pod mounting
`hub-data`; delete any `hub.db-wal`/`-shm` there — the snapshot is a whole database); scale to 1. The start-up log
line `console passwords sealed at rest (0 legacy plaintext row(s) sealed now)` and one reveal on a host page confirm it.
6. Shred `k`, `enc.key` copies and `out/` when done.
## 4. Option B (not my pick): restic to a dedicated Hetzner Storage Box sub-account
+3
View File
@@ -12,6 +12,9 @@ group (`manifests/hub.yaml`: `recurring-job-group.longhorn.io/default: enabled`)
the WAL (precondition asserted: a plain file copy loses them), keep-2, never-two-at-once, a failed write leaves no
file, the catch-up age; `cmd/hub/r173_wiring_test.go` pins the 02:00 schedule and the catch-up in `main()`.
Red-proofs R1–R6, all convict: `documentation/audits/hub-db-offsite-2026-10-05/partA/red-proof.txt`.
- **`cmd/hubdb-check`** (added after the release; a tool, not in the image): opens a RESTORED copy of hub.db with the
seal key and prints only counts — hosts, console passwords opened / failed / absent; exit 0 only when every vaulted one
opens. Runbook §3 step 3. Tests: same key opens, a wrong key opens nothing (red-proof R7, `partD/restore-procedure/`).
## v0.135.0 — the hub's own safety: form protection for the password path, console passwords sealed at rest; boxes left behind are listed and alarmed; a waiting customer with no e-mail is flagged (R-135, R-133, R-604, R-530, R-508) (2026-10-05)
+75
View File
@@ -0,0 +1,75 @@
// hubdb-check — open a RESTORED copy of hub.db with the seal key and report, as counts only, whether every vaulted
// console password opens (R-173, runbooks/RUNBOOK-hub-db-offsite-backup.md §3, step 2).
//
// Usage: hubdb-check <restored hub.db> <file holding OFFSITE_SECRET_KEY>
//
// Run it on a scratch COPY, never on the live database: opening runs the store's migrations. It prints no secret —
// only the number of hosts and of console passwords that opened, failed or are absent. Exit 0 only when every
// vaulted console password opened, at least one host exists and at least one password was vaulted; a wrong key
// fails every row (the seal is AES-GCM, authenticated). Pinned by main_test.go.
package main
import (
"fmt"
"io"
"log"
"os"
"gitea.dooplex.hu/admin/felhom-hub/internal/store"
)
type result struct{ hosts, opened, failed, absent int }
func check(dbPath, keyFile string) (result, error) {
var r result
raw, err := os.ReadFile(keyFile)
if err != nil {
return r, fmt.Errorf("reading the key file: %w", err)
}
key, err := store.ParseOffsiteSecretKey(string(raw))
if err != nil {
return r, fmt.Errorf("the key file does not hold a valid OFFSITE_SECRET_KEY: %w", err)
}
s, err := store.New(dbPath, log.New(io.Discard, "", 0))
if err != nil {
return r, err
}
defer s.Close()
if err := s.SetOffsiteSecretKey(key); err != nil {
return r, err
}
hosts, err := s.ListHosts()
if err != nil {
return r, err
}
r.hosts = len(hosts)
for _, h := range hosts {
c, err := s.GetHostRecoveryCredential(h.HostID)
switch {
case err != nil:
r.failed++
case c == nil:
r.absent++
default:
r.opened++
}
}
return r, nil
}
func main() {
if len(os.Args) != 3 {
fmt.Fprintln(os.Stderr, "usage: hubdb-check <restored hub.db> <OFFSITE_SECRET_KEY file>")
os.Exit(2)
}
r, err := check(os.Args[1], os.Args[2])
if err != nil {
fmt.Fprintln(os.Stderr, "hubdb-check: FAILED:", err)
os.Exit(1)
}
fmt.Printf("hosts=%d console_passwords_opened=%d failed=%d absent=%d\n", r.hosts, r.opened, r.failed, r.absent)
if r.hosts == 0 || r.opened == 0 || r.failed > 0 {
fmt.Fprintln(os.Stderr, "hubdb-check: FAILED: not every console password opened with this key")
os.Exit(1)
}
}
+65
View File
@@ -0,0 +1,65 @@
package main
import (
"encoding/hex"
"io"
"log"
"os"
"path/filepath"
"testing"
"gitea.dooplex.hu/admin/felhom-hub/internal/store"
)
// Under `go test` every store seals with this fixed key (store.New, testing.Testing()).
const testKey = "felhom-hub-test-only-seal-key-32"
func seededCopy(t *testing.T) string {
t.Helper()
p := filepath.Join(t.TempDir(), "hub.db")
s, err := store.New(p, log.New(io.Discard, "", 0))
if err != nil {
t.Fatal(err)
}
for _, h := range []string{"h1", "h2"} {
if err := s.UpsertHost(&store.Host{HostID: h, CustomerID: "c", APIKey: "k-" + h}); err != nil {
t.Fatal(err)
}
}
if err := s.SaveHostRecoveryCredential("h1", "root@pam", "console-pw"); err != nil {
t.Fatal(err)
}
s.Close()
return p
}
func keyFile(t *testing.T, key string) string {
t.Helper()
p := filepath.Join(t.TempDir(), "k")
if err := os.WriteFile(p, []byte(hex.EncodeToString([]byte(key))+"\n"), 0o600); err != nil {
t.Fatal(err)
}
return p
}
// The consequence the runbook relies on: the SAME key opens every vaulted console password of the restored copy.
func TestCheck_SameKeyOpensEveryConsolePassword(t *testing.T) {
r, err := check(seededCopy(t), keyFile(t, testKey))
if err != nil {
t.Fatal(err)
}
if r.hosts != 2 || r.opened != 1 || r.failed != 0 || r.absent != 1 {
t.Fatalf("got %+v, want hosts=2 opened=1 failed=0 absent=1", r)
}
}
// A different key opens nothing — so "opened" is evidence that the key matches, not that the column is readable.
func TestCheck_WrongKeyOpensNothing(t *testing.T) {
r, err := check(seededCopy(t), keyFile(t, "a-completely-different-key-32byt"))
if err != nil {
t.Fatal(err)
}
if r.opened != 0 || r.failed != 1 {
t.Fatalf("got %+v, want opened=0 failed=1", r)
}
}
+12
View File
@@ -1,3 +1,15 @@
## hub-db-backup 1.0 — the hub database leaves DooPlex every night, encrypted, and is restore-tested weekly (R-173) (2026-10-05)
New, in `scripts/hub-db-backup/` (versioned from day one, R-231): `felhom-hub-db-backup` (02:30 — copy the hub's newest
nightly snapshot out of the pod, refuse a snapshot >26 h old, a size mismatch, a failed `integrity_check` or no hosts;
push it to ep0's PBS ns `operator`, `--crypt-mode encrypt`, write-only token; write the success timestamp only after the
push), `felhom-hub-db-restore-test` (Sun 04:30 — restore the newest copy with the read-only token, refuse a copy >50 h
old, a failed integrity check, no hosts, or any console password stored readable), four systemd units, `install.sh`
(does not enable the timers). `test_hub_db_backup.py`: 15 tests with fake `kubectl` / `proxmox-backup-client` (nothing
reaches the hub, PBS or ep0); red-proofs P1–P9 (`documentation/audits/hub-db-offsite-2026-10-05/partC/red-proof.txt` —
P4 and P5 needed a stricter test first). Installed on DooPlex 2026-10-05; timers enabled after the first manual run.
Runbook: `documentation/runbooks/RUNBOOK-hub-db-offsite-backup.md`.
## felhom-host-install.sh 1.31.0 — the root-owned files come from the agent's config bundle (R-840) (2026-10-04)
Needs a vouched agent ≥ 0.143.0 for the bundle (hub ≥ 0.133.0 serves its sha); an older vouched agent: the per-file