R-173 option A in force (hub DB nightly to ep0, restore-tested, alarmed); R-519 proven live on 9202 and closed; R-173/R-232 narrowed; R-882..R-885 opened (332 -> 335); runbook §3 tested; hubdb-check
gates / gates (push) Successful in 59s
gates / gates (push) Successful in 59s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
+12
@@ -16,6 +16,18 @@
|
||||
> and holds nothing of its own; this file does hold its own content, namely the standing rulings below.
|
||||
|
||||
|
||||
> **2026-10-05 (evening) — the hub database off DooPlex (hub v0.136.0; operator rulings `09` 125–127).** R-173 option A
|
||||
> IN FORCE: hub `internal/dbsnap` writes `VACUUM INTO /data/snapshots/hub-<UTC>.db` at 02:00 Budapest (keep 2, `ErrBusy`
|
||||
> on overlap, start-up catch-up when >24 h; `05` §16.3); DooPlex `scripts/hub-db-backup/` (installed by `install.sh` to
|
||||
> `/usr/local/sbin`, units + timers in `/etc/systemd/system`, config `/etc/felhom-hub-backup/{env,token-push,
|
||||
> token-restore,enc.key}` all root 0600) pushes at 02:30 to ep0 `felhom-offsite` ns `operator` as `dooplex-hub@pbs!push`
|
||||
> (`DatastoreBackup`), restore-tests Sun 04:30 as `!restore` (`DatastoreReader`); ep0 prune job `prune-operator-hubdb`
|
||||
> (03:45, keep-daily 14, keep-weekly 8, max-depth 0). Success-only textfile metrics → `HubDBBackupStale` (26 h, critical)
|
||||
> and `HubDBRestoreTestStale` (8 d) in homelab-manifests `backup-freshness`. `hub/cmd/hubdb-check` opens a restored copy
|
||||
> with the seal key (runbook §3). Hub PVC 2 Gi + label `enabled`. Collateral: Longhorn instance-manager restart (R-882),
|
||||
> zipline pinned 4.7.0 (R-883). R-519 proven live on 9202 (now controller 0.296.0) and CLOSED. Register 332 → 335.
|
||||
> Report: `REPORT-hub-db-offsite-2026-10-05.md`.
|
||||
|
||||
> **2026-10-05 (late afternoon) — the hub's own safety, boxes left behind, the agent's root grants (hub v0.135.0, controller
|
||||
> v0.296.0, agent v0.146.1 + bundle `42333e96…`, golden 0.296.0 vouched with agent 0.146.1, min_agent 0.131.0).** CC decisions
|
||||
> 119–124, *operator may reverse*. Hub: R-135 a cookie-less state change needs Basic + `X-Felhom-Operator` (`05` §16.1); R-133
|
||||
|
||||
@@ -3,9 +3,44 @@
|
||||
**Ready for the first real tester (Tester-2): yes. Tester 2 (a laptop, off at night) was offline again; nothing was
|
||||
sent to it.**
|
||||
|
||||
**Updated 2026-10-05 (late afternoon, the hub-safety session): every box of ours healthy. The hub refuses forged form
|
||||
posts and keeps the console passwords locked; the System page shows boxes left behind; the agent can no longer make
|
||||
itself root. Report: `REPORT-hub-safety-2026-10-05.md`.**
|
||||
**Updated 2026-10-05 (evening, the hub-database session): every box of ours healthy. The hub database now leaves
|
||||
DooPlex every night, locked, to ep0, is test-restored every Sunday, and an alarm mails you if either stops. Report:
|
||||
`REPORT-hub-db-offsite-2026-10-05.md`.**
|
||||
|
||||
## Tonight (2026-10-05, evening): the hub database off DooPlex; the cut-backup check on the scratch box
|
||||
|
||||
**Decisions:** none of mine. Yours (`09` 125–127): ep0 for the copy; the scratch-box restart allowed; the agent's three
|
||||
by-design abilities stay. Today you also chose: grow the hub's disk to 2 GiB, and restart Longhorn's disk manager.
|
||||
|
||||
**What works now (proven live):**
|
||||
- **The hub makes a clean copy of its database every night at 02:00** (hub 0.136.0). The first one: 353 MB, 44 s.
|
||||
- **DooPlex checks it, locks it with a key ep0 never sees, and sends it to ep0 at 02:30.** First send: 7 s.
|
||||
- **Every Sunday at 04:30 DooPlex takes the copy back from ep0 and checks it**: it opens, it is whole, every console
|
||||
password in it is still locked. Done once by hand today: 4 boxes, 4 locked passwords, 0 readable.
|
||||
- **The two ep0 accounts can only do their one job:** the sending one cannot delete, the checking one cannot write,
|
||||
neither can see the households' backups. (The sending one can read back its own locked copies — that is how ep0 works.)
|
||||
- **Your two saved keys work:** a key rebuilt from the paper copy you saved opened the copy, and your saved lock key
|
||||
opened all 4 console passwords in it (a wrong key opened none).
|
||||
- **The alarm:** proven by Prometheus' own rule test; the real alarm mail is below.
|
||||
- **A backup cut off by a restart is now said on the backup pages** (scratch box): the page said so, the restore point
|
||||
kept the older time of the part that was not redone, and the next full backup cleared the message.
|
||||
|
||||
**What broke, and what I did:**
|
||||
- **The hub's disk would not grow:** Longhorn's disk manager on DooPlex was stuck. My first try (stopping the hub so the
|
||||
disk could grow offline) did not work and kept the hub **down about 9.5 minutes**. You approved restarting the disk
|
||||
manager: all 77 disks were back in under 2 minutes, and the hub's disk is 2 GiB now.
|
||||
- **Zipline did not come back after that restart:** it is set to "always the newest", so it pulled a new release that
|
||||
refused its database. I pinned it to the previous release; it runs and its database is updated. 7 more apps on
|
||||
DooPlex use "always the newest" (new row).
|
||||
|
||||
**Register:** 332 → 335 rows (1 closed: the cut-backup check; 4 opened: the Longhorn fault, the "always newest" apps, a
|
||||
monitoring sync drift, script tests not in CI).
|
||||
|
||||
**Needs you:**
|
||||
1. **Nothing urgent.** If you do nothing, the copy runs every night and you get a mail only if it stops.
|
||||
2. **When convenient:** pin the 7 other DooPlex apps that use "always the newest" (or tell me to list them for you). If
|
||||
you do nothing, any restart may upgrade one of them by surprise, as it did zipline.
|
||||
3. **Tester 2's one-time step** is unchanged (below).
|
||||
|
||||
## Today (2026-10-05, late afternoon): the hub's own safety; boxes left behind; the agent's admin rights
|
||||
|
||||
@@ -38,7 +73,7 @@ itself root. Report: `REPORT-hub-safety-2026-10-05.md`.**
|
||||
minutes. Both figures are on the page now.
|
||||
|
||||
**Needs you:**
|
||||
1. **Where the hub database's off-site copy goes** (it holds every box's keys and your customers' settings):
|
||||
1. **(DECIDED 2026-10-05 evening: A, done — see Tonight)** **Where the hub database's off-site copy goes** (it holds every box's keys and your customers' settings):
|
||||
- **A — my pick: ep0's backup server**, encrypted on DooPlex before it leaves, with a weekly restore test and an
|
||||
alarm mail. Costs one small change on ep0 (a write-only account) and keeping two keys in your password manager.
|
||||
- **B: a separate Hetzner Storage Box account** with restic. More new parts to look after than A.
|
||||
@@ -47,7 +82,7 @@ itself root. Report: `REPORT-hub-safety-2026-10-05.md`.**
|
||||
`documentation/runbooks/RUNBOOK-hub-db-offsite-backup.md`.
|
||||
- **Either way, first:** put the hub's lock key (`OFFSITE_SECRET_KEY`) in your password manager — without it a copy
|
||||
of the database cannot open the console passwords.
|
||||
2. **The power-cut-during-backup check (R-519) on the scratch box:** the permission check refused my restarting the
|
||||
2. **(DONE 2026-10-05 evening — see Tonight)** **The power-cut-during-backup check (R-519) on the scratch box:** the permission check refused my restarting the
|
||||
controller in the middle of a backup. Say "go" and the next session does it once on 9202; if not, the fix stays
|
||||
proven by tests only.
|
||||
3. **Three things the agent can still do, by design** (each written in `03` §3.1): pick which controller image its own
|
||||
|
||||
@@ -235,6 +235,7 @@ likewise silent. Evidence: `audits/DRILL-r361-2026-08-22/evidence/06-part3-decis
|
||||
| Secrets hygiene: bearer in k8s Secret, no secrets in git, single-quote credential store | hub v0.53, conventions | **IMPLEMENTED** | 07-13 closing bundle | |
|
||||
| Operator login password changeable from UI | hub v0.54 | **IMPLEMENTED** | 07-13 | |
|
||||
| **The operator sees boxes left behind: per-customer controller floors with their age (and which a global raise cannot move), each box's agent against the vouched one, a 7-day "agent behind" alarm and a "floor raise skipped boxes" mail** | hub **v0.135.0** (R-604, R-530) | **IMPLEMENTED — the page PROVEN-LIVE, the alarm and the mail unit-proven** | `audits/hub-safety-2026-10-05/partD/` (live System page: three per-customer floors, Tester 2 `0.142.0 → 0.145.0`); `osupdates/r530_agent_alarm_test.go`, `web/r604_floor_held_back_test.go` | the mail fires only on a GLOBAL raise below an override — not exercised live |
|
||||
| **The hub database survives the loss of DooPlex: a nightly consistent copy, encrypted, on ep0; restore-tested weekly; an alarm when either stops** | hub **v0.136.0** (R-173), `scripts/hub-db-backup/`, homelab-manifests rules | **PROVEN-LIVE (2026-10-05)** — first push, ep0 listing, restore test (4 hosts, 4 sealed, 0 readable), token limits, a key rebuilt from the paper copy decrypts, the saved seal key opens 4/4 console passwords in the restored copy; the alarm by `promtool` rule test + red-proofs | `audits/hub-db-offsite-2026-10-05/`; `runbooks/RUNBOOK-hub-db-offsite-backup.md` | Runbook §3 steps 4–5 (into a live PVC) not exercised (R-173) |
|
||||
| **The operator surface refuses a cross-site form post on BOTH login paths (session token; Basic auth + `X-Felhom-Operator`)** | hub **v0.135.0** (R-135) | **PROVEN-LIVE (2026-10-05)** | `audits/hub-safety-2026-10-05/partA/live.txt` (Basic, no header → 403 even on an unknown path; with the header → passes; header without credentials → 401); `web/r135_csrf_test.go` (39 paths) | |
|
||||
| Box operating-system security updates (Proxmox host, guest Debian, Docker engine) | agent v0.143.0, hub v0.133.0 | **PARTIAL — the GUEST and HOST Debian fast lanes and the DOCKER engine slow lane are PROVEN-LIVE (2026-10-04), with the System page, the fleet view and the alarms; the KERNEL lane is MISSING** | Guest: `audits/os-guest-lane-2026-10-04/`. Host + fleet + alarms: `audits/os-host-lane-2026-10-04/`. Docker + System page + crash guard: `audits/os-docker-crash-2026-10-04/` — live-restore on with the same container ids on every box; Docker 29.8.2 on both demo boxes; operator-approved Docker release; a signed undo and a signed ring-1 step; a replay refused; the crash guard restarted demo-hp twice and kept it off the third time. Design `architecture/11-os-updates.md` §5.8, §5.9, §8 | **No automatic undo** (guest: last night's backup; host: by-hand runbook; Docker: a signed undo job); existing boxes get root-owned files by the signed config bundle since 2026-10-04 (R-840 CLOSED; a box from before agent 0.143.0 needs one by-hand bootstrap — Tester 2: R-862; `audits/r840-config-bundle-2026-10-04/`); test approvals now end with the test (R-859); the agent's sudoers is root-equivalent (R-861); the kernel lane (R-836); facts reach the hub late after a boot (R-853). **2026-10-05 (agent v0.144.1):** R8 measures the real download (R-865); a killed pass still reports (R-868, live); the debug pass runs with the hub away (R-866, live); **a power cut mid-update was proven by day on demo-hp — the box came back by itself in 37 s, but the next pass fails until `dpkg --configure -a` is run by hand (R-876, P2, open)** — `audits/night-fixes-2026-10-05/`. **2026-10-05 afternoon (agent v0.145.0): R-876 FIXED and proven live — after a second crash mid-unpack the next pass repaired dpkg by itself (`REPAIR … journal=1`) and finished** — `audits/catchup-2026-10-05/partD/` |
|
||||
| **An ENGLISH-SPEAKING household's first hour: download, install, pair, bind, claim, two apps** | controller **v0.259.0** + hub **v0.119.0** + ISO 1.29.0 + the whole catalog | **PROVEN-LIVE on 0.258.0 with one blocker; THE BLOCKER IS FIXED AND PROVEN, THE WALK IS NOT REPEATED** | `audits/DRILL-first-hour-en-0258-2026-09-20.md` — a fresh install 2026-09-20, one intervention (R-494), stop rule not reached. Then `audits/i18n-closing-2026-09-21/live/` — the three blockers fixed and each proven on a live box or in the operator's inbox: the claim page answers English through the real cookie path; the Backup page's tier names follow the language; and the setup mail carries **four plain-ASCII English words** where the drill's carried `képző-szkítia-ásatás`, one day apart in the same inbox. | **R-596, R-597 and R-598 are CLOSED.** What this row still does NOT claim: **the fixed journey has not been walked end to end by a stranger on a fresh install.** Three fixes proven at the endpoint are not an hour proven by a person, and this project's own rule is that fixes are not a journey (see the recovery-journey row). **Also not walked:** the recovery code (needs ep0), backup/restore/remove/power-cut (proven 2026-09-14), and the two Backup-page *warnings* themselves — guest 9201 is healthy and a healthy box renders none, so they are covered by handler render tests, not live. **Verdict: nothing known now stands between an English-speaking tester and their box — and that is a different sentence from "the walk passed".** |
|
||||
|
||||
@@ -140,8 +140,9 @@ mert a doboz vagy a vezérlő újraindult…"), kept until a run ends with every
|
||||
before this, both pages said nothing and the synthesised „Utolsó adatbázis mentés … OK" was read off the fresh `.sql`
|
||||
the cut run left beside last night's tars; that line now reads failed after a cut. Each restore point's time was
|
||||
already its OLDEST part (the data block, v0.275.0) — so a torn unit is dated by its stale tars, never by its new dump.
|
||||
*Live: unit-proven and red-proved; the live cut on 9202 was refused by the permission check and waits for the
|
||||
operator (R-519 narrowed).*
|
||||
*Live: PROVEN 2026-10-05 on 9202 (operator ruling 126): a run cut by a controller restart between bookstack's config
|
||||
dump and its database volume → both pages carry the notice, the restore point reads the older volume's time, the next
|
||||
complete run clears it (`audits/hub-db-offsite-2026-10-05/partD/r519/`; R-519 closed).*
|
||||
|
||||
### Lane 2 — the operator: guest and host recovery
|
||||
|
||||
|
||||
@@ -0,0 +1,26 @@
|
||||
## promtool, in pod/prometheus-55b675779d-t8c74, 2026-10-05T13:42:17Z
|
||||
### green
|
||||
SUCCESS
|
||||
|
||||
rc=0
|
||||
### red: threshold 26h -> 260h
|
||||
FAILED:
|
||||
alertname: HubDBBackupStale, time: 1d2h40m,
|
||||
exp:[
|
||||
0:
|
||||
Labels:{alertname="HubDBBackupStale", component="backup", instance="dooplex", severity="critical"}
|
||||
Annotations:{description="No successful push of the hub DB snapshot to ep0 for >26h (daily at 02:30). Check `journalctl -u felhom-hub-db-backup.service`, the tunnel `systemctl status felhom-ep0-pbs-tunnel`, and that the hub logs `db snapshot written` at 02:00.", summary="The Felhom hub database has not reached ep0 for 26 h (R-173)"}
|
||||
],
|
||||
got:[]
|
||||
|
||||
|
||||
command terminated with exit code 1
|
||||
rc=0
|
||||
### red: absent() removed
|
||||
SUCCESS
|
||||
|
||||
### red (re-run): absent() removed from HubDBBackupStale — first attempt above did NOT apply (indent mismatch)
|
||||
FAILED:
|
||||
alertname: HubDBBackupStale, time: 40m,
|
||||
Labels:{alertname="HubDBBackupStale", component="backup", severity="critical"}
|
||||
got:[]
|
||||
@@ -0,0 +1,52 @@
|
||||
rule_files: [bf.yml]
|
||||
evaluation_interval: 1m
|
||||
tests:
|
||||
# 1. the push stops: last success at t=0, never again → fires once 26 h + 30 m have passed
|
||||
- interval: 5m
|
||||
input_series:
|
||||
- series: 'felhom_hub_db_backup_last_success_timestamp_seconds{instance="dooplex"}'
|
||||
values: '0+0x400'
|
||||
- series: 'felhom_hub_db_restore_test_last_success_timestamp_seconds{instance="dooplex"}'
|
||||
values: '0+0x400'
|
||||
alert_rule_test:
|
||||
- eval_time: 26h
|
||||
alertname: HubDBBackupStale
|
||||
exp_alerts: []
|
||||
- eval_time: 26h40m
|
||||
alertname: HubDBBackupStale
|
||||
exp_alerts:
|
||||
- exp_labels: {severity: critical, component: backup, instance: dooplex}
|
||||
exp_annotations:
|
||||
summary: "The Felhom hub database has not reached ep0 for 26 h (R-173)"
|
||||
description: "No successful push of the hub DB snapshot to ep0 for >26h (daily at 02:30). Check `journalctl -u felhom-hub-db-backup.service`, the tunnel `systemctl status felhom-ep0-pbs-tunnel`, and that the hub logs `db snapshot written` at 02:00."
|
||||
# 2. healthy: the push succeeds every 24 h → never fires
|
||||
- interval: 1h
|
||||
input_series:
|
||||
- series: 'felhom_hub_db_backup_last_success_timestamp_seconds{instance="dooplex"}'
|
||||
values: '0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 86400 86400 86400 86400 86400 86400 86400 86400 86400 86400 86400 86400 86400 86400 86400 86400 86400 86400 86400 86400 86400 86400 86400 86400 172800 172800 172800'
|
||||
- series: 'felhom_hub_db_restore_test_last_success_timestamp_seconds{instance="dooplex"}'
|
||||
values: '0+0x50'
|
||||
alert_rule_test:
|
||||
- eval_time: 50h
|
||||
alertname: HubDBBackupStale
|
||||
exp_alerts: []
|
||||
# 3. the metric never existed (script never ran) → fires on absent()
|
||||
- interval: 5m
|
||||
input_series:
|
||||
- series: 'up{job="node"}'
|
||||
values: '1+0x30'
|
||||
alert_rule_test:
|
||||
- eval_time: 40m
|
||||
alertname: HubDBBackupStale
|
||||
exp_alerts:
|
||||
- exp_labels: {severity: critical, component: backup}
|
||||
exp_annotations:
|
||||
summary: "The Felhom hub database has not reached ep0 for 26 h (R-173)"
|
||||
description: "No successful push of the hub DB snapshot to ep0 for >26h (daily at 02:30). Check `journalctl -u felhom-hub-db-backup.service`, the tunnel `systemctl status felhom-ep0-pbs-tunnel`, and that the hub logs `db snapshot written` at 02:00."
|
||||
- eval_time: 2h
|
||||
alertname: HubDBRestoreTestStale
|
||||
exp_alerts:
|
||||
- exp_labels: {severity: warning, component: backup}
|
||||
exp_annotations:
|
||||
summary: "The hub database copy on ep0 has not passed a restore test for 8 days (R-173)"
|
||||
description: "The weekly restore test (Sun 04:30) has not succeeded for >8 days. Check `journalctl -u felhom-hub-db-restore-test.service`."
|
||||
@@ -0,0 +1,5 @@
|
||||
## alarm drill start 2026-10-05T13:51:46Z: success file moved aside (absent case)
|
||||
backup_freshness.prom
|
||||
fan_metrics.prom
|
||||
felhom_hub_db_restore.prom
|
||||
node_housekeeping.prom
|
||||
@@ -0,0 +1,27 @@
|
||||
## manual push via unit, 2026-10-05T13:50:30Z
|
||||
Result=success
|
||||
ExecMainStatus=0
|
||||
2026-10-05T15:50:08+02:00 dooplex systemd[1]: Starting felhom-hub-db-backup.service - Felhom: push the hub DB snapshot to ep0 (R-173)...
|
||||
2026-10-05T15:50:09+02:00 dooplex felhom-hub-db-backup[3036511]: felhom-hub-db-backup: snapshot hub-20261005T123037Z.db, 79 min old
|
||||
2026-10-05T15:50:18+02:00 dooplex felhom-hub-db-backup[3036511]: felhom-hub-db-backup: checked: 369807360 bytes, integrity ok, 4 host(s)
|
||||
2026-10-05T15:50:18+02:00 dooplex felhom-hub-db-backup[3037790]: Starting backup: [operator]:host/dooplex-hub/2026-10-05T13:50:18Z
|
||||
2026-10-05T15:50:18+02:00 dooplex felhom-hub-db-backup[3037790]: Client name: dooplex
|
||||
2026-10-05T15:50:18+02:00 dooplex felhom-hub-db-backup[3037790]: Starting backup protocol: Mon Oct 5 15:50:18 2026
|
||||
2026-10-05T15:50:18+02:00 dooplex felhom-hub-db-backup[3037790]: Using encryption key from '/etc/felhom-hub-backup/enc.key'..
|
||||
2026-10-05T15:50:18+02:00 dooplex felhom-hub-db-backup[3037790]: Encryption key fingerprint: b2:19:bf:36:3b:97:3d:6c
|
||||
2026-10-05T15:50:18+02:00 dooplex felhom-hub-db-backup[3037790]: No previous manifest available.
|
||||
2026-10-05T15:50:18+02:00 dooplex felhom-hub-db-backup[3037790]: Upload directory '/var/lib/felhom-hub-backup/stage' to 'dooplex-hub@pbs!push@127.0.0.1:18007:felhom-offsite' as hubdb.pxar.didx
|
||||
2026-10-05T15:50:25+02:00 dooplex felhom-hub-db-backup[3037790]: hubdb.pxar: had to backup 352.676 MiB of 352.676 MiB (compressed 17.242 MiB) in 6.58 s (average 53.604 MiB/s)
|
||||
2026-10-05T15:50:25+02:00 dooplex felhom-hub-db-backup[3037790]: Uploaded backup catalog (56 B)
|
||||
2026-10-05T15:50:25+02:00 dooplex felhom-hub-db-backup[3037790]: Duration: 7.19s
|
||||
2026-10-05T15:50:25+02:00 dooplex felhom-hub-db-backup[3037790]: End Time: Mon Oct 5 15:50:25 2026
|
||||
2026-10-05T15:50:25+02:00 dooplex felhom-hub-db-backup[3036511]: felhom-hub-db-backup: pushed hub-20261005T123037Z.db to ep0 (ns operator) in 7 s
|
||||
2026-10-05T15:50:25+02:00 dooplex felhom-hub-db-backup[3036511]: felhom-hub-db-backup: success signal written
|
||||
2026-10-05T15:50:30+02:00 dooplex systemd[1]: felhom-hub-db-backup.service: Deactivated successfully.
|
||||
2026-10-05T15:50:30+02:00 dooplex systemd[1]: Finished felhom-hub-db-backup.service - Felhom: push the hub DB snapshot to ep0 (R-173).
|
||||
2026-10-05T15:50:30+02:00 dooplex systemd[1]: felhom-hub-db-backup.service: Consumed 9.566s CPU time, 584.8M memory peak.
|
||||
## textfile
|
||||
# HELP felhom_hub_db_backup_last_success_timestamp_seconds Last successful push of the hub DB snapshot to ep0 (R-173).
|
||||
# TYPE felhom_hub_db_backup_last_success_timestamp_seconds gauge
|
||||
felhom_hub_db_backup_last_success_timestamp_seconds 1791208225
|
||||
felhom_hub_db_backup_last_success_bytes 369807360
|
||||
@@ -0,0 +1 @@
|
||||
gitea.dooplex.hu/admin/felhom-controller:0.296.0 Up 6 seconds (healthy)
|
||||
@@ -0,0 +1,3 @@
|
||||
## deploy bookstack 2026-10-05T13:53:26Z (fields: DOMAIN APP_KEY DB_PASSWORD ADMIN_PASSWORD; SUBDOMAIN = catalog default)
|
||||
|
||||
HTTP 202
|
||||
@@ -0,0 +1,36 @@
|
||||
## run 1 (baseline) start 2026-10-05T14:04:36Z
|
||||
|
||||
HTTP 200
|
||||
{"ok":true,"data":{"enabled":true,"running":true}}
|
||||
|
||||
HTTP 200
|
||||
|
||||
## run 1 end 2026-10-05T14:05:17Z: {"ok":true,"data":{"db_dump":{"count":2,"duration":"32.784678377s","last_run":"2026-10-05T14:05:12.109790785Z","success":true},"enabled":true,"running":false}}
|
||||
2026/10/05 14:04:44 data_versions.go:131: [DEBUG] [backup] bookstack: stamped volume-dumps/bookstack_bookstack_config.tar (70144 B) with pins [lscr.io/linuxserver/bookstack:26.09.1@sha256:99cd1f5707c1911afad213adec5c9739763b76f843d1477142231834ecdcb6f7 mariadb:12.3@sha256:805c8e104bd563d5bfa24fadd3f31cd419ea859cb5277f32b5dbf2db714f9ed1]
|
||||
2026/10/05 14:04:44 backup.go:848: [DEBUG] [backup] Dumping volume bookstack_bookstack_db_data for bookstack
|
||||
2026/10/05 14:04:45 backup.go:873: [INFO] [backup] Volume dump: bookstack/bookstack_bookstack_db_data → 153.4 MB
|
||||
2026/10/05 14:04:45 data_versions.go:131: [DEBUG] [backup] bookstack: stamped volume-dumps/bookstack_bookstack_db_data.tar (160868864 B) with pins [lscr.io/linuxserver/bookstack:26.09.1@sha256:99cd1f5707c1911afad213adec5c9739763b76f843d1477142231834ecdcb6f7 mariadb:12.3@sha256:805c8e104bd563d5bfa24fadd3f31cd419ea859cb5277f32b5dbf2db714f9ed1]
|
||||
2026/10/05 14:04:45 backup.go:990: [INFO] [backup] Restarting bookstack after volume dump
|
||||
2026/10/05 14:04:51 backup.go:974: [INFO] [backup] Stopping paperless-ngx for safe volume dump
|
||||
2026/10/05 14:04:58 backup.go:848: [DEBUG] [backup] Dumping volume paperless-ngx_paperless_redis_data for paperless-ngx
|
||||
2026/10/05 14:04:58 backup.go:873: [INFO] [backup] Volume dump: paperless-ngx/paperless-ngx_paperless_redis_data → 7.4 MB
|
||||
2026/10/05 14:04:58 data_versions.go:131: [DEBUG] [backup] paperless-ngx: stamped volume-dumps/paperless-ngx_paperless_redis_data.tar (7792128 B) with pins [ghcr.io/paperless-ngx/paperless-ngx:2.20.15 postgres:18-alpine redis:7-alpine]
|
||||
2026/10/05 14:04:58 backup.go:848: [DEBUG] [backup] Dumping volume paperless-ngx_paperless_data for paperless-ngx
|
||||
2026/10/05 14:04:59 backup.go:873: [INFO] [backup] Volume dump: paperless-ngx/paperless-ngx_paperless_data → 5.8 MB
|
||||
2026/10/05 14:04:59 data_versions.go:131: [DEBUG] [backup] paperless-ngx: stamped volume-dumps/paperless-ngx_paperless_data.tar (6052352 B) with pins [ghcr.io/paperless-ngx/paperless-ngx:2.20.15 postgres:18-alpine redis:7-alpine]
|
||||
2026/10/05 14:04:59 backup.go:848: [DEBUG] [backup] Dumping volume paperless-ngx_paperless_postgres_data for paperless-ngx
|
||||
2026/10/05 14:04:59 backup.go:873: [INFO] [backup] Volume dump: paperless-ngx/paperless-ngx_paperless_postgres_data → 68.8 MB
|
||||
2026/10/05 14:04:59 data_versions.go:131: [DEBUG] [backup] paperless-ngx: stamped volume-dumps/paperless-ngx_paperless_postgres_data.tar (72107008 B) with pins [ghcr.io/paperless-ngx/paperless-ngx:2.20.15 postgres:18-alpine redis:7-alpine]
|
||||
2026/10/05 14:04:59 backup.go:990: [INFO] [backup] Restarting paperless-ngx after volume dump
|
||||
2026/10/05 14:05:11 backup.go:974: [INFO] [backup] Stopping privatebin for safe volume dump
|
||||
2026/10/05 14:05:11 backup.go:848: [DEBUG] [backup] Dumping volume privatebin_privatebin_data for privatebin
|
||||
2026/10/05 14:05:12 backup.go:873: [INFO] [backup] Volume dump: privatebin/privatebin_privatebin_data → 1.5 KB
|
||||
2026/10/05 14:05:12 data_versions.go:131: [DEBUG] [backup] privatebin: stamped volume-dumps/privatebin_privatebin_data.tar (1536 B) with pins [privatebin/pdo:2.0.6@sha256:4c141b2326f8b353598ce9ce7507a9cfecf2dad5c60a39fea903d430e296d8f5]
|
||||
2026/10/05 14:05:12 backup.go:986: [INFO] [backup] privatebin NOT restarted after the volume dump: the household stopped it meanwhile
|
||||
2026/10/05 14:05:12 backup.go:684: [INFO] [backup] App-data backup completed: 2 databases (483.2 KB total), 3 volume dump(s) (32.785s)
|
||||
2026/10/05 14:05:12 recovery_unit.go:289: [INFO] [backup] Recovery unit captured for bookstack → /mnt/sys_drive/felhom-data/backups/primary/bookstack (images=2, secrets-referenced=3, data_keys=0, portable-carried=2/2, withheld=1)
|
||||
2026/10/05 14:05:12 recovery_unit.go:289: [INFO] [backup] Recovery unit captured for paperless-ngx → /mnt/felhom-drives/scratch_hdd/userdata/paperless-ngx/backups/primary/paperless-ngx (images=3, secrets-referenced=3, data_keys=0, portable-carried=2/2, withheld=1)
|
||||
2026/10/05 14:05:12 recovery_unit.go:289: [INFO] [backup] Recovery unit captured for privatebin → /mnt/sys_drive/felhom-data/backups/primary/privatebin (images=1, secrets-referenced=0, data_keys=0, portable-carried=0/0, withheld=0)
|
||||
|
||||
/var/lib/felhom/docker/volumes/felhom-controller-data/_data/data/appdata-run.json
|
||||
/var/lib/docker/volumes/felhom-controller-data/_data/data/appdata-run.json
|
||||
@@ -0,0 +1,12 @@
|
||||
## after run 1, 2026-10-05T14:05:34Z
|
||||
### run record
|
||||
{"running":false,"started_at":"0001-01-01T00:00:00Z","interrupted":"0001-01-01T00:00:00Z"}
|
||||
### restore points (bookstack)
|
||||
{"ok":true,"data":[{"time":"2026-10-05T14:04:39Z","short_id":"helyi","tier":1,"drive_label":"Belső SSD (rendszer)"}]}
|
||||
|
||||
HTTP 200
|
||||
|
||||
### volume dump file times
|
||||
total 157172
|
||||
-rw-r--r-- 1 root root 70144 2026-10-05T14:04:44 bookstack_bookstack_config.tar
|
||||
-rw-r--r-- 1 root root 160868864 2026-10-05T14:04:44 bookstack_bookstack_db_data.tar
|
||||
@@ -0,0 +1,34 @@
|
||||
## run 2 start 2026-10-05T14:05:50Z
|
||||
HTTP 200
|
||||
### watcher
|
||||
match 2026-10-05T14:05:58.856456447Z
|
||||
restarted 2026-10-05T14:05:59.669364637Z rc=0
|
||||
Up 10 seconds (healthy)
|
||||
### record at the moment of the cut
|
||||
{"running":true,"started_at":"2026-10-05T14:05:53.282495446Z","interrupted":"0001-01-01T00:00:00Z"}
|
||||
### record after restart
|
||||
{"running":false,"started_at":"2026-10-05T14:05:53.282495446Z","interrupted":"2026-10-05T14:05:53.282495446Z"}
|
||||
### controller log around the cut (previous + new process)
|
||||
2026/10/05 14:05:53 backup.go:557: [INFO] [backup] Starting database dump run
|
||||
2026/10/05 14:05:53 dbdump.go:208: [INFO] [backup] Discovered 2 databases
|
||||
2026/10/05 14:05:53 backup.go:596: [INFO] [backup] Discovered 2 database(s): paperless-postgres(postgres), bookstack-db(mariadb)
|
||||
2026/10/05 14:05:53 dbdump.go:410: [INFO] [backup] DB dump: paperless-postgres → paperless-ngx-postgres.sql (428.1 KB, 321ms, 72 tables)
|
||||
2026/10/05 14:05:54 dbdump.go:410: [INFO] [backup] DB dump: bookstack-db → bookstack-mariadb.sql (55.9 KB, 310ms, 41 tables)
|
||||
2026/10/05 14:05:54 backup.go:974: [INFO] [backup] Stopping bookstack for safe volume dump
|
||||
2026/10/05 14:05:58 backup.go:873: [INFO] [backup] Volume dump: bookstack/bookstack_bookstack_config → 49.0 KB
|
||||
2026/10/05 14:05:59 main.go:340: [INFO] felhom-controller 0.296.0 starting (customer: demo-hp, domain: enkisfelhom.hu)
|
||||
2026/10/05 14:05:59 appstop_marker.go:278: [WARN] [appstop] crash recovery: an app-data backup (volume dump) (op "volume-dump:bookstack") was interrupted and left 1 app(s) stopped — restarting them: [bookstack]
|
||||
2026/10/05 14:05:59 manager.go:1235: [INFO] [stacks] Starting stack: bookstack
|
||||
2026/10/05 14:06:06 appstop_marker.go:304: [INFO] [appstop] crash recovery: restarted bookstack after the interrupted an app-data backup (volume dump)
|
||||
2026/10/05 14:06:06 sync.go:117: [INFO] [sync] Starting catalog sync (repo: https://gitea.dooplex.hu/admin/app-catalog-felhom.eu.git, interval: 15m0s)
|
||||
2026/10/05 14:06:06 sync.go:208: [INFO] [sync] Starting catalog sync
|
||||
2026/10/05 14:06:06 restore_record_wiring.go:35: [INFO] [backup] restore record wired: /opt/docker/felhom-controller/data/restore-status.json (interrupted at startup: false)
|
||||
2026/10/05 14:06:06 run_record_wiring.go:27: [WARN] [backup] the app-data backup run started 2026-10-05T14:05:53Z was cut off by the stop — the backup pages say so until the next complete run (R-519)
|
||||
2026/10/05 14:06:06 scheduler.go:223: [INFO] [scheduler] Starting scheduler with 19 jobs
|
||||
2026/10/05 14:06:06 backup.go:1179: [INFO] [backup] Found 3 DB dump files across drives
|
||||
2026/10/05 14:06:06 dbdump.go:208: [INFO] [backup] Discovered 2 databases
|
||||
2026/10/05 14:06:06 [INFO] [backup] Discovered app data: 3 apps
|
||||
2026/10/05 14:06:06 recovery_unit.go:289: [INFO] [backup] Recovery unit captured for bookstack → /mnt/sys_drive/felhom-data/backups/primary/bookstack (images=2, secrets-referenced=3, data_keys=0, portable-carried=2/2, withheld=1)
|
||||
2026/10/05 14:06:06 recovery_unit.go:289: [INFO] [backup] Recovery unit captured for paperless-ngx → /mnt/felhom-drives/scratch_hdd/userdata/paperless-ngx/backups/primary/paperless-ngx (images=3, secrets-referenced=3, data_keys=0, portable-carried=2/2, withheld=1)
|
||||
2026/10/05 14:06:06 backup.go:1250: [INFO] [backup] Backup status cache refreshed
|
||||
2026/10/05 14:06:09 manager.go:1601: [INFO] [stacks] bookstack lscr.io/linuxserver/bookstack:26.09.1@sha256:99cd1f5707c1911afad213adec5c9739763b76f843d1477142231834ecdcb6f7 running Up 3 seconds (health: starting)
|
||||
@@ -0,0 +1,23 @@
|
||||
## after the cut, 2026-10-05T14:06:23Z
|
||||
### GET /backups
|
||||
data-interrupted-run present: 1
|
||||
<div class="alert alert-warning" data-interrupted-run="true">A legutóbbi mentés (2026-10-05 16:05) megszakadt, mert a doboz vagy a vezérlő újraindult. Amit nem fejezett be, annak a korábbi mentése maradt meg — minden visszaállítási pont annyira friss, amennyire a legrégebbi része. A következő teljes mentés után ez az üzenet eltűnik.
|
||||
HTTP 200
|
||||
### GET /backups/apps
|
||||
data-interrupted-run present: 1
|
||||
<div class="alert alert-warning" data-interrupted-run="true">A legutóbbi mentés (2026-10-05 16:05) megszakadt, mert a doboz vagy a vezérlő újraindult. Amit nem fejezett be, annak a korábbi mentése maradt meg — minden visszaállítási pont annyira friss, amennyire a legrégebbi része. A következő teljes mentés után ez az üzenet eltűnik.
|
||||
HTTP 200
|
||||
### negative control: a marker that must NOT be on the page
|
||||
0
|
||||
### restore points (bookstack)
|
||||
{"ok":true,"data":[{"time":"2026-10-05T14:04:44Z","short_id":"helyi","tier":1,"drive_label":"Belső SSD (rendszer)"}]}
|
||||
### dump file times
|
||||
total 314252
|
||||
-rw-r--r-- 1 root root 50176 2026-10-05T14:05:58 bookstack_bookstack_config.tar
|
||||
-rw-r--r-- 1 root root 160868864 2026-10-05T14:04:44 bookstack_bookstack_db_data.tar
|
||||
-rw-r--r-- 1 root root 160867328 2026-10-05T14:05:58 bookstack_bookstack_db_data.tar.tmp
|
||||
drwxr-xr-x 2 root root 4096 2026-10-05T14:05:54 db-dumps
|
||||
drwxr-xr-x 2 root root 4096 2026-10-05T14:05:58 volume-dumps
|
||||
### bookstack after
|
||||
bookstack Up 27 seconds (healthy)
|
||||
bookstack-db Up 32 seconds (healthy)
|
||||
@@ -0,0 +1,12 @@
|
||||
## run 3 (complete) start 2026-10-05T14:06:44Z
|
||||
HTTP 200
|
||||
## run 3 end 2026-10-05T14:07:28Z: {"ok":true,"data":{"db_dump":{"count":2,"duration":"33.315240638s","last_run":"2026-10-05T14:07:20.615183031Z","success":true},"enabled":true,"running":false}}
|
||||
2026/10/05 14:05:12 backup.go:684: [INFO] [backup] App-data backup completed: 2 databases (483.2 KB total), 3 volume dump(s) (32.785s)
|
||||
2026/10/05 14:07:20 backup.go:684: [INFO] [backup] App-data backup completed: 2 databases (483.9 KB total), 3 volume dump(s) (33.315s)
|
||||
{"running":false,"started_at":"0001-01-01T00:00:00Z","interrupted":"0001-01-01T00:00:00Z"}
|
||||
total 157152
|
||||
-rw-r--r-- 1 root root 50688 2026-10-05T14:06:52 bookstack_bookstack_config.tar
|
||||
-rw-r--r-- 1 root root 160867328 2026-10-05T14:06:53 bookstack_bookstack_db_data.tar
|
||||
/backups data-interrupted-run: 0
|
||||
/backups/apps data-interrupted-run: 0
|
||||
restore points: {"ok":true,"data":[{"time":"2026-10-05T14:06:47Z","short_id":"helyi","tier":1,"drive_label":"Belső SSD (rendszer)"}]}
|
||||
@@ -0,0 +1,8 @@
|
||||
## teardown 2026-10-05T14:07:48Z
|
||||
stop: HTTP 200
|
||||
remove: HTTP 200
|
||||
containers: 0
|
||||
volumes: 0
|
||||
stackdir: none
|
||||
backups: none
|
||||
/root/.dbody
|
||||
@@ -0,0 +1,9 @@
|
||||
## runbook §3 drill 2026-10-05T14:09:41Z (scratch /var/lib/felhom-hub-backup/sec3.65qk, root 0700)
|
||||
step 1 restore host/dooplex-hub/2026-10-05T13:50:18Z
|
||||
restore complete (352.676 MiB processed in 3.9s, average 90.961 MiB/s)
|
||||
-rw------- 369807360 hub.db
|
||||
step 2 the seal key from Secret/offsite-secret-key → /var/lib/felhom-hub-backup/sec3.65qk/k (0600, not printed)
|
||||
key file bytes: 64
|
||||
step 3 hubdb-check (same key): hosts=4 console_passwords_opened=4 failed=0 absent=0 rc=0
|
||||
control (a random key): hosts=4 console_passwords_opened=0 failed=4 absent=0 hubdb-check: FAILED: not every console password opened with this key rc=0
|
||||
scratch shredded: gone
|
||||
@@ -0,0 +1,7 @@
|
||||
### R7 hubdb-check counts a failed open as opened
|
||||
=== RUN TestCheck_WrongKeyOpensNothing
|
||||
main_test.go:63: got {hosts:2 opened:1 failed:0 absent:1}, want opened=0 failed=1
|
||||
--- FAIL: TestCheck_WrongKeyOpensNothing (0.03s)
|
||||
FAIL
|
||||
FAIL gitea.dooplex.hu/admin/felhom-hub/cmd/hubdb-check 0.040s
|
||||
FAIL
|
||||
@@ -0,0 +1,17 @@
|
||||
## snapshot list on ep0 with the READ-ONLY token, 2026-10-05T13:50:40Z
|
||||
+=======================================+=============+=====================================+
|
||||
| snapshot | size | files |
|
||||
+=======================================+=============+=====================================+
|
||||
| host/dooplex-hub/2026-10-05T13:50:18Z | 352.677 MiB | catalog.pcat1 hubdb.pxar index.json |
|
||||
+=======================================+=============+=====================================+
|
||||
unit rc=0
|
||||
Result=success
|
||||
2026-10-05T15:50:41+02:00 dooplex felhom-hub-db-restore-test[3040582]: felhom-hub-db-restore-test: restoring host/dooplex-hub/2026-10-05T13:50:18Z
|
||||
2026-10-05T15:50:46+02:00 dooplex felhom-hub-db-restore-test[3040582]: felhom-hub-db-restore-test: checked: integrity ok, 4 host(s), 4 sealed console password(s), 0 readable
|
||||
2026-10-05T15:50:46+02:00 dooplex felhom-hub-db-restore-test[3040582]: felhom-hub-db-restore-test: success signal written
|
||||
# HELP felhom_hub_db_restore_test_last_success_timestamp_seconds Last successful restore test of the hub DB copy on ep0 (R-173).
|
||||
# TYPE felhom_hub_db_restore_test_last_success_timestamp_seconds gauge
|
||||
felhom_hub_db_restore_test_last_success_timestamp_seconds 1791208246
|
||||
.cache
|
||||
.kube
|
||||
stage
|
||||
@@ -0,0 +1,12 @@
|
||||
## 2026-10-05T13:51:07Z token limits (each line = the tool output)
|
||||
push forget : Error: permission check failed - missing Datastore.Modify|Datastore.Prune on /datastore/felhom-offsite/operator
|
||||
restore forget : Error: permission check failed - missing Datastore.Modify|Datastore.Prune on /datastore/felhom-offsite/operator
|
||||
restore backup : Error: missing permissions 'Datastore.Backup' on '/datastore/felhom-offsite/operator'
|
||||
push list ns root (households): Error: permission check failed - missing Datastore.Audit|Datastore.Backup on /datastore/felhom-offsite
|
||||
restore list ns root (households): Error: permission check failed - missing Datastore.Audit|Datastore.Backup on /datastore/felhom-offsite
|
||||
push restore own copy: Error: missing key - manifest was created with key b2:19:bf:36:3b:97:3d:6c exit-file=absent
|
||||
## paper-key proof: restore with a key file rebuilt from the data field only
|
||||
restore rc=0
|
||||
integrity: ok hosts: 4
|
||||
## and with NO key: Error: missing key - manifest was created with key b2:19:bf:36:3b:97:3d:6c
|
||||
push restore own copy WITH key: restore complete (352.676 MiB processed in 4.6s, average 76.793 MiB/s) file=PRESENT
|
||||
@@ -26,6 +26,16 @@
|
||||
|
||||
---
|
||||
|
||||
## 2026-10-05 (evening) — the hub database off DooPlex (hub v0.136.0; `09` rulings 125–127)
|
||||
|
||||
The full text of every row below: `git show cea8502f:documentation/backlog/OPEN-ITEMS.md`.
|
||||
|
||||
| Row | What | Closed | Evidence |
|
||||
|---|---|---|---|
|
||||
| **R-519** | **After a backup torn by a power cut, a restore point carried the new database dump's time over the previous run's files, and no screen said the run was interrupted (P2).** Fixed in controller v0.296.0 (page notice + synthesised status; dating by the oldest part since v0.275.0). **Live on 9202 (operator ruling 126):** a complete run, then a run cut by `docker restart felhom-controller` 0.8 s after bookstack's config dump and before its database volume; afterwards both /backups and /backups/apps carry `data-interrupted-run` („A legutóbbi mentés (2026-10-05 16:05) megszakadt …", negative control 0), the restore point reads 14:04:44Z = the run-1 database volume (its oldest part; config 14:05:58, SQL 14:05:54), the controller restarted bookstack itself; the next complete run cleared the notice and replaced the torn `.tar.tmp`. The throwaway bookstack was removed through the product (0 containers, volumes, folders, backups). | CLOSED 2026-10-05 — FIXED controller v0.296.0, proven live | `audits/hub-db-offsite-2026-10-05/partD/r519/`; `internal/backup/run_record_test.go` |
|
||||
|
||||
---
|
||||
|
||||
## 2026-10-05 (late afternoon) — the hub's own safety, boxes left behind, the agent's root grants (hub v0.135.0, controller v0.296.0, agent v0.146.1, golden 0.296.0; CC decisions 119–124)
|
||||
|
||||
The full text of every row below: `git show 9bb45eaa:documentation/backlog/OPEN-ITEMS.md` (R-880 was opened and closed in this session).
|
||||
|
||||
File diff suppressed because one or more lines are too long
@@ -1,16 +1,16 @@
|
||||
# Runbook — the hub database in a backup that is NOT on DooPlex (R-173, R-232) — PROPOSED, needs the operator's go
|
||||
# Runbook — the hub database in a backup that is NOT on DooPlex (R-173, R-232) — IN FORCE since 2026-10-05
|
||||
|
||||
> **Status: PROPOSED 2026-10-05. Nothing here has been done.** DooPlex and ep0 are protected; every step below changes
|
||||
> one of them, so each waits for the operator's go (the decision is in `STATUS.md`). The readings this plan rests on:
|
||||
> `audits/hub-safety-2026-10-05/partC/readings.txt` (read only).
|
||||
> **Status: IN FORCE 2026-10-05** (option A, `09` decision 125). Steps 0–7 DONE, each marked below; §3 is a tested
|
||||
> procedure. Evidence: `audits/hub-db-offsite-2026-10-05/` (part A–D). The readings the plan rested on:
|
||||
> `audits/hub-safety-2026-10-05/partC/readings.txt`. **Corrections found while doing it are marked „Corrected".**
|
||||
|
||||
## 1. What is true today (measured 2026-10-05)
|
||||
|
||||
| Question | Answer |
|
||||
|---|---|
|
||||
| Where the hub database lives | `/data/hub.db` (+ `-wal`, `-shm`) in the hub pod, PVC `hub-data` (Longhorn, 1 Gi, replicas on DooPlex's `sdb1`). 357 MiB. |
|
||||
| Where the hub database lives | `/data/hub.db` (+ `-wal`, `-shm`) in the hub pod, PVC `hub-data` (Longhorn, **2 Gi since 2026-10-05**, was 1 Gi; replicas on DooPlex's `sdb1`). 357 MiB; a snapshot is 353 MiB. |
|
||||
| Is it in a backup? | **Yes, but only on DooPlex.** Longhorn's `backup-daily` (04:00) and `backup-weekly` (Sun 05:00), `retain=1`, write to `nfs://192.168.0.180:/mnt/5_hdd/backup/longhorn-pvc` — DooPlex's own `sda1`. Last: 2026-10-05 02:06 UTC, Completed. |
|
||||
| Why R-173 said "excluded" | The PVC carries `recurring-job-group.longhorn.io/default: disabled` (git, `manifests/hub.yaml`, commit `868e8465` of 2026-02-16, no reason given). The live Longhorn **Volume** carries `enabled` — set by hand at some point, so the backups run. **This is drift:** the PVC label is the source of truth (Longhorn syncs PVC → Volume), so the backups can stop at any sync without anyone seeing it. |
|
||||
| Why R-173 said "excluded" | The PVC carries `recurring-job-group.longhorn.io/default: disabled` (git, `manifests/hub.yaml`, commit `868e8465` of 2026-02-16, no reason given). The live Longhorn **Volume** carries `enabled` — set by hand at some point, so the backups run. **This is drift:** the PVC label is the source of truth (Longhorn syncs PVC → Volume), so the backups can stop at any sync without anyone seeing it. **Corrected 2026-10-05:** the Volume kept `enabled` through every sync since February while the PVC said `disabled` — Longhorn did NOT copy the PVC label down. The PVC now says `enabled` too (Step 1), so the two agree; which one Longhorn reads was not measured. |
|
||||
| What DooPlex's own backup covers | `dooplex-backup.timer` (03:19): k3s state, k8s Secrets (GPG files), Gitea mirrors, user data, PostgreSQL dumps — **all onto `sda1`, the same machine.** Nothing leaves DooPlex (`audits/RECON-dooplex-backup-2026-08-06.md`, R-232). |
|
||||
| What tells anyone a backup failed | **Nothing.** `NOTIFY_WEBHOOK_URL` is commented out, so `notify_failure` is a no-op. No Prometheus rule watches a Longhorn backup's success or age, nor `dooplex-backup.service`. |
|
||||
| What the database holds | Box→hub API keys, customer configs (incl. the owner passphrase), escrow custody blobs (opaque), the PBS-DR token values, the off-site sub-account passwords and — since hub v0.135.0 — the console passwords **sealed** under `OFFSITE_SECRET_KEY`. |
|
||||
@@ -22,7 +22,7 @@ ep0 is already off-site (Hetzner), already runs PBS, and DooPlex already reaches
|
||||
(127.0.0.1:18007). ep0's datastore is also pulled back to DooPlex nightly (`ep0-copy`), so the copy exists in two places,
|
||||
one of them off DooPlex. The copy is encrypted on DooPlex with a key ep0 never sees.
|
||||
|
||||
### Step 0 — the keys go off DooPlex first (operator, at the keyboard, 5 min)
|
||||
### Step 0 — the keys go off DooPlex first (operator, at the keyboard, 5 min) — DONE 2026-10-05 (operator: both saved)
|
||||
|
||||
Without these, every later step backs up something nobody can open after a DooPlex loss.
|
||||
|
||||
@@ -33,33 +33,65 @@ sudo kubectl -n felhom-system get secret offsite-secret-key -o jsonpath='{.data.
|
||||
sudo proxmox-backup-client key paperkey /etc/felhom-hub-backup/enc.key --output-format text
|
||||
```
|
||||
|
||||
### Step 1 — end the label drift (CC, a felhom.eu commit + ArgoCD sync; reversible)
|
||||
**The `data` field of that output is enough** (the key has no passphrase, `kdf: null`): a key file rebuilt from it alone —
|
||||
`{"kdf": null, "created": "<any RFC 3339 time>", "modified": "<same>", "data": "<saved value>"}` — restored and
|
||||
decrypted the copy on 2026-10-05 (`audits/hub-db-offsite-2026-10-05/partD/token-limits-and-paperkey.txt`). Note:
|
||||
`proxmox-backup-client key show` prints a fingerprint only when the file stores one, so it cannot check a rebuilt key —
|
||||
a restore can.
|
||||
|
||||
### Step 1 — end the label drift (CC, a felhom.eu commit + ArgoCD sync; reversible) — DONE 2026-10-05
|
||||
|
||||
`manifests/hub.yaml`: `recurring-job-group.longhorn.io/default: disabled` → `enabled`. Sync. Check:
|
||||
`sudo kubectl -n felhom-system get pvc hub-data -o jsonpath='{.metadata.labels}'` and the Volume label both read `enabled`.
|
||||
This keeps today's on-DooPlex copy alive; it is not the off-site copy.
|
||||
|
||||
### Step 2 — a write-only place on ep0 (on ep0, root; the operator's go for an ep0 change)
|
||||
Done with the volume growth to 2 Gi (two snapshots of 353 MiB do not fit in 1 Gi). **The online growth failed** — Longhorn's
|
||||
`instance-manager` (116 days up) called a host PID that no longer existed (`nsenter: cannot open /host/proc/196610/ns/mnt`);
|
||||
an offline growth was impossible too (the expansion holds its own attachment ticket). The operator approved restarting
|
||||
the instance-manager: all 77 DooPlex volumes back `attached/healthy` in 110 s, the volume grew, and one app (zipline, on
|
||||
`:latest`) came back on a newer release that refused its database — pinned to 4.7.0 (homelab-manifests). Evidence:
|
||||
`audits/hub-db-offsite-2026-10-05/partA/step1-*.txt`.
|
||||
|
||||
### Step 2 — a write-only place on ep0 (on ep0, root; the operator's go for an ep0 change) — DONE 2026-10-05
|
||||
|
||||
What was run (PBS 4.2.8 on ep0; the proposal's commands were wrong in three places, corrected here):
|
||||
|
||||
```bash
|
||||
proxmox-backup-manager user create dooplex-hub@pbs --comment "DooPlex pushes the hub DB (R-173)"
|
||||
proxmox-backup-manager user generate-token dooplex-hub@pbs push # the secret → a 0600 file on DooPlex, file → file
|
||||
# namespace for operator data, apart from the households' namespaces
|
||||
proxmox-backup-client namespace create operator --repository 'root@pam@127.0.0.1:8007:felhom-offsite'
|
||||
proxmox-backup-manager acl update /datastore/felhom-offsite/operator DatastoreBackup --auth-id 'dooplex-hub@pbs!push'
|
||||
# retention on ep0 (the server prunes; the pushing token cannot delete — DatastoreBackup has no Prune)
|
||||
proxmox-backup-manager prune-job create prune-operator-hubdb --store felhom-offsite --ns operator \
|
||||
--schedule 'daily 03:45' --keep-daily 14 --keep-weekly 8
|
||||
proxmox-backup-manager user create dooplex-hub@pbs --comment "..." # no password: cannot log in
|
||||
proxmox-backup-debug api create /admin/datastore/felhom-offsite/namespace --name operator
|
||||
# (the CLI crashes AFTER creating it, printing the result — 'not implemented'; check it exists, don't re-run)
|
||||
# two tokens, each secret file → file into a root 0600 file on DooPlex, never printed:
|
||||
ssh root@<ep0> "proxmox-backup-debug api create /access/users/dooplex-hub@pbs/token/push --output-format json" \
|
||||
| python3 -c '<print json["value"]>' | sudo sh -c 'umask 077; cat > /etc/felhom-hub-backup/token-push'
|
||||
# (same for token/restore → token-restore; `user generate-token` has no --output-format)
|
||||
P=/datastore/felhom-offsite/operator
|
||||
proxmox-backup-manager acl update $P DatastoreBackup --auth-id dooplex-hub@pbs # a token's rights are cut
|
||||
proxmox-backup-manager acl update $P DatastoreReader --auth-id dooplex-hub@pbs # down by its user's rights
|
||||
proxmox-backup-manager acl update $P DatastoreBackup --auth-id 'dooplex-hub@pbs!push'
|
||||
proxmox-backup-manager acl update $P DatastoreReader --auth-id 'dooplex-hub@pbs!restore'
|
||||
proxmox-backup-manager prune-job create prune-operator-hubdb --store felhom-offsite --ns operator --max-depth 0 \
|
||||
--schedule '03:45' --keep-daily 14 --keep-weekly 8 # 'daily 03:45' is not a PBS calendar event
|
||||
```
|
||||
|
||||
### Step 3 — a consistent snapshot of the live database (CC, a hub release)
|
||||
Measured (`partB/`, `partD/token-limits-and-paperkey.txt`): neither token can forget a snapshot or list the datastore
|
||||
root (the households); the restore token cannot write. **Corrected: the push token CAN restore its own copies** — PBS
|
||||
lets a backup's owner read it back (`DatastoreBackup` = `Datastore.Backup`, owner-scoped). It reaches only `operator`,
|
||||
and everything it can read is encrypted with a key ep0 never sees. The households' two prune jobs and the Sunday GC
|
||||
are unchanged (before/after in `partB/`).
|
||||
|
||||
### Step 3 — a consistent snapshot of the live database (CC, a hub release) — DONE, hub v0.136.0 (`05` §16.3)
|
||||
|
||||
`hub.db` is in WAL mode and is written every few seconds; copying the three files is not one point in time. The hub
|
||||
gets a nightly `VACUUM INTO '/data/snapshots/hub-<UTC date>.db'` (keeps 2, logs size and duration) — one SQLite
|
||||
statement, consistent by construction, WAL-aware. **Needs a hub release** (filed under R-173). No `sqlite3` exists in the
|
||||
hub image, so the copy must be made by the hub itself.
|
||||
statement, consistent by construction, WAL-aware. No `sqlite3` exists in the hub image, so the copy is made by the hub
|
||||
itself. **Corrected:** one snapshot took 44 s on the live volume (0.63 s on a local scratch copy), 353 MiB.
|
||||
|
||||
### Step 4 — the push (on DooPlex, root; `felhom-hub-db-backup.service` + `.timer` 02:30, CC writes, operator approves)
|
||||
### Step 4 — the push (on DooPlex, root; `felhom-hub-db-backup.service` + `.timer` 02:30) — DONE 2026-10-05
|
||||
|
||||
**The real script is `scripts/hub-db-backup/felhom-hub-db-backup`** (versioned, R-231; installed by `install.sh`;
|
||||
15 tests in `test_hub_db_backup.py`, run by hand — not in CI). It adds to the sketch below: it refuses a snapshot older
|
||||
than 26 h (the hub stopped snapshotting), a copy whose size differs from the pod's file, and a copy with no hosts; it
|
||||
encrypts with `--crypt-mode encrypt`. The sketch, as proposed:
|
||||
|
||||
```bash
|
||||
#!/bin/sh -eu
|
||||
@@ -77,7 +109,10 @@ echo "felhom_hub_db_backup_last_success_timestamp_seconds $(date +%s)" > /var/li
|
||||
&& mv /var/lib/node_exporter/textfile_collector/felhom_hub_db.prom.$$ /var/lib/node_exporter/textfile_collector/felhom_hub_db.prom
|
||||
```
|
||||
|
||||
### Step 5 — the restore test (weekly, Sun 04:30, same unit family)
|
||||
### Step 5 — the restore test (weekly, Sun 04:30, same unit family) — DONE 2026-10-05
|
||||
|
||||
**The real script is `scripts/hub-db-backup/felhom-hub-db-restore-test`**, with the READ-ONLY token. It also refuses a
|
||||
newest copy older than 50 h. The sketch, as proposed:
|
||||
|
||||
```bash
|
||||
T=$(mktemp -d); chmod 700 "$T"
|
||||
@@ -90,9 +125,12 @@ shred -u "$T/hub.db"*; rmdir "$T"
|
||||
echo "felhom_hub_db_restore_test_last_success_timestamp_seconds $(date +%s)" > …/felhom_hub_db_restore.prom # same tmp+mv
|
||||
```
|
||||
|
||||
The push token needs `DatastoreReader` on `operator` too for the restore (or a second, read-only token — cleaner).
|
||||
Done with a second, read-only token (`token-restore`).
|
||||
|
||||
### Step 6 — the alarm (homelab-manifests `prometheus-rules`, then `POST /-/reload` — the Prometheus there has no reloader)
|
||||
### Step 6 — the alarm (homelab-manifests `prometheus-rules`, then `POST /-/reload` — the Prometheus there has no reloader) — DONE 2026-10-05
|
||||
|
||||
In the `backup-freshness` group; `promtool test rules` proves it (`partC/bf_test.yml`, two red-proofs). The Prometheus
|
||||
Deployment is OutOfSync in ArgoCD for a reason unrelated to this; only the rules ConfigMap was synced.
|
||||
|
||||
```yaml
|
||||
- alert: HubDBBackupStale
|
||||
@@ -109,18 +147,46 @@ The push token needs `DatastoreReader` on `operator` too for the restore (or a s
|
||||
Both reach the existing `email-notifications` receiver. `absent()` makes "the script never ran" an alarm too — an empty
|
||||
log is not a success.
|
||||
|
||||
### Step 7 — prove it once (CC, with the operator's go)
|
||||
### Step 7 — prove it once (CC, with the operator's go) — DONE 2026-10-05 (`partD/`)
|
||||
|
||||
Run the unit by hand; read the snapshot on ep0 (`proxmox-backup-client snapshot list --ns operator`); run the restore
|
||||
test by hand; stop the timer for a day on purpose and see `HubDBBackupStale` mail arrive (positive observable), then
|
||||
start it again.
|
||||
|
||||
## 3. Bringing the hub back from this copy (the procedure the plan exists for)
|
||||
## 3. Bringing the hub back from this copy (the procedure the plan exists for) — TESTED 2026-10-05
|
||||
|
||||
1. A k3s with the `felhom` ArgoCD app, and **`Secret/offsite-secret-key` recreated with the SAME value** (Step 0 copy).
|
||||
2. Restore the newest snapshot (Step 5's first command, with the paper key), scale `deploy/hub` to 0, copy `hub.db` into
|
||||
the PVC (no `-wal`/`-shm` — the snapshot is a whole database), scale to 1. The log line
|
||||
`console passwords sealed at rest (0 legacy plaintext row(s) sealed now)` and a working reveal prove the key matches.
|
||||
Steps 1–3 were run on 2026-10-05 against the real copy on ep0 (`audits/hub-db-offsite-2026-10-05/partD/restore-procedure/drill.txt`):
|
||||
4 hosts, **4 of 4 console passwords opened with the saved seal key, 0 of 4 with a random key**. Steps 4–5 (into a live
|
||||
PVC) were NOT run — that needs the hub down; they are the ordinary scale-copy-scale.
|
||||
|
||||
What you need, all from the password manager: the seal key (`OFFSITE_SECRET_KEY`), the backup key's `data` field, and
|
||||
the read-only token (or ep0 root to mint a new one: Step 2).
|
||||
|
||||
1. **The backup key file.** On the machine doing the restore, as root, `umask 077`, write
|
||||
`{"kdf": null, "created": "2026-01-01T00:00:00+00:00", "modified": "2026-01-01T00:00:00+00:00", "data": "<data>"}`
|
||||
to `enc.key` (Step 0 note). A copy of DooPlex's `/etc/felhom-hub-backup/enc.key` works as is.
|
||||
2. **Restore the newest copy** (from any machine that reaches ep0's PBS on 8007 — DooPlex uses the tunnel 127.0.0.1:18007):
|
||||
```bash
|
||||
export PBS_PASSWORD_FILE=<token-restore file> PBS_FINGERPRINT=<ep0 cert fingerprint, /etc/felhom-hub-backup/env>
|
||||
R='dooplex-hub@pbs!restore@<ep0>:8007:felhom-offsite'
|
||||
proxmox-backup-client snapshot list host/dooplex-hub --ns operator --repository "$R" # pick the newest
|
||||
proxmox-backup-client restore host/dooplex-hub/<time> hubdb.pxar ./out --ns operator --keyfile enc.key --repository "$R"
|
||||
sqlite3 -readonly out/hub.db 'PRAGMA integrity_check' # must print: ok
|
||||
```
|
||||
If DooPlex's `ep0-copy` datastore survived, the same copy is there too (pulled nightly).
|
||||
3. **Prove the seal key matches BEFORE putting the copy in place** — on a COPY of `out/hub.db` (the check migrates it):
|
||||
```bash
|
||||
cd felhom.eu/hub && go build -o hubdb-check ./cmd/hubdb-check
|
||||
printf '%s' "<OFFSITE_SECRET_KEY>" > k; chmod 600 k # from the password manager — not on a command line in a shared shell
|
||||
./hubdb-check copy-of-hub.db k # want: hosts=N console_passwords_opened=N failed=0; exit 0
|
||||
```
|
||||
`failed>0` means the wrong seal key: the hub would start but could open no console password (`05` §16.2).
|
||||
4. **A k3s with the `felhom` ArgoCD app**, and `Secret/offsite-secret-key` recreated with the SAME value:
|
||||
`kubectl -n felhom-system create secret generic offsite-secret-key --from-file=OFFSITE_SECRET_KEY=k`.
|
||||
5. **Into the PVC:** scale `deploy/hub` to 0; put `out/hub.db` into the volume as `/data/hub.db` (a helper pod mounting
|
||||
`hub-data`; delete any `hub.db-wal`/`-shm` there — the snapshot is a whole database); scale to 1. The start-up log
|
||||
line `console passwords sealed at rest (0 legacy plaintext row(s) sealed now)` and one reveal on a host page confirm it.
|
||||
6. Shred `k`, `enc.key` copies and `out/` when done.
|
||||
|
||||
## 4. Option B (not my pick): restic to a dedicated Hetzner Storage Box sub-account
|
||||
|
||||
|
||||
@@ -12,6 +12,9 @@ group (`manifests/hub.yaml`: `recurring-job-group.longhorn.io/default: enabled`)
|
||||
the WAL (precondition asserted: a plain file copy loses them), keep-2, never-two-at-once, a failed write leaves no
|
||||
file, the catch-up age; `cmd/hub/r173_wiring_test.go` pins the 02:00 schedule and the catch-up in `main()`.
|
||||
Red-proofs R1–R6, all convict: `documentation/audits/hub-db-offsite-2026-10-05/partA/red-proof.txt`.
|
||||
- **`cmd/hubdb-check`** (added after the release; a tool, not in the image): opens a RESTORED copy of hub.db with the
|
||||
seal key and prints only counts — hosts, console passwords opened / failed / absent; exit 0 only when every vaulted one
|
||||
opens. Runbook §3 step 3. Tests: same key opens, a wrong key opens nothing (red-proof R7, `partD/restore-procedure/`).
|
||||
|
||||
## v0.135.0 — the hub's own safety: form protection for the password path, console passwords sealed at rest; boxes left behind are listed and alarmed; a waiting customer with no e-mail is flagged (R-135, R-133, R-604, R-530, R-508) (2026-10-05)
|
||||
|
||||
|
||||
@@ -0,0 +1,75 @@
|
||||
// hubdb-check — open a RESTORED copy of hub.db with the seal key and report, as counts only, whether every vaulted
|
||||
// console password opens (R-173, runbooks/RUNBOOK-hub-db-offsite-backup.md §3, step 2).
|
||||
//
|
||||
// Usage: hubdb-check <restored hub.db> <file holding OFFSITE_SECRET_KEY>
|
||||
//
|
||||
// Run it on a scratch COPY, never on the live database: opening runs the store's migrations. It prints no secret —
|
||||
// only the number of hosts and of console passwords that opened, failed or are absent. Exit 0 only when every
|
||||
// vaulted console password opened, at least one host exists and at least one password was vaulted; a wrong key
|
||||
// fails every row (the seal is AES-GCM, authenticated). Pinned by main_test.go.
|
||||
package main
|
||||
|
||||
import (
|
||||
"fmt"
|
||||
"io"
|
||||
"log"
|
||||
"os"
|
||||
|
||||
"gitea.dooplex.hu/admin/felhom-hub/internal/store"
|
||||
)
|
||||
|
||||
type result struct{ hosts, opened, failed, absent int }
|
||||
|
||||
func check(dbPath, keyFile string) (result, error) {
|
||||
var r result
|
||||
raw, err := os.ReadFile(keyFile)
|
||||
if err != nil {
|
||||
return r, fmt.Errorf("reading the key file: %w", err)
|
||||
}
|
||||
key, err := store.ParseOffsiteSecretKey(string(raw))
|
||||
if err != nil {
|
||||
return r, fmt.Errorf("the key file does not hold a valid OFFSITE_SECRET_KEY: %w", err)
|
||||
}
|
||||
s, err := store.New(dbPath, log.New(io.Discard, "", 0))
|
||||
if err != nil {
|
||||
return r, err
|
||||
}
|
||||
defer s.Close()
|
||||
if err := s.SetOffsiteSecretKey(key); err != nil {
|
||||
return r, err
|
||||
}
|
||||
hosts, err := s.ListHosts()
|
||||
if err != nil {
|
||||
return r, err
|
||||
}
|
||||
r.hosts = len(hosts)
|
||||
for _, h := range hosts {
|
||||
c, err := s.GetHostRecoveryCredential(h.HostID)
|
||||
switch {
|
||||
case err != nil:
|
||||
r.failed++
|
||||
case c == nil:
|
||||
r.absent++
|
||||
default:
|
||||
r.opened++
|
||||
}
|
||||
}
|
||||
return r, nil
|
||||
}
|
||||
|
||||
func main() {
|
||||
if len(os.Args) != 3 {
|
||||
fmt.Fprintln(os.Stderr, "usage: hubdb-check <restored hub.db> <OFFSITE_SECRET_KEY file>")
|
||||
os.Exit(2)
|
||||
}
|
||||
r, err := check(os.Args[1], os.Args[2])
|
||||
if err != nil {
|
||||
fmt.Fprintln(os.Stderr, "hubdb-check: FAILED:", err)
|
||||
os.Exit(1)
|
||||
}
|
||||
fmt.Printf("hosts=%d console_passwords_opened=%d failed=%d absent=%d\n", r.hosts, r.opened, r.failed, r.absent)
|
||||
if r.hosts == 0 || r.opened == 0 || r.failed > 0 {
|
||||
fmt.Fprintln(os.Stderr, "hubdb-check: FAILED: not every console password opened with this key")
|
||||
os.Exit(1)
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,65 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"encoding/hex"
|
||||
"io"
|
||||
"log"
|
||||
"os"
|
||||
"path/filepath"
|
||||
"testing"
|
||||
|
||||
"gitea.dooplex.hu/admin/felhom-hub/internal/store"
|
||||
)
|
||||
|
||||
// Under `go test` every store seals with this fixed key (store.New, testing.Testing()).
|
||||
const testKey = "felhom-hub-test-only-seal-key-32"
|
||||
|
||||
func seededCopy(t *testing.T) string {
|
||||
t.Helper()
|
||||
p := filepath.Join(t.TempDir(), "hub.db")
|
||||
s, err := store.New(p, log.New(io.Discard, "", 0))
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
for _, h := range []string{"h1", "h2"} {
|
||||
if err := s.UpsertHost(&store.Host{HostID: h, CustomerID: "c", APIKey: "k-" + h}); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
}
|
||||
if err := s.SaveHostRecoveryCredential("h1", "root@pam", "console-pw"); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
s.Close()
|
||||
return p
|
||||
}
|
||||
|
||||
func keyFile(t *testing.T, key string) string {
|
||||
t.Helper()
|
||||
p := filepath.Join(t.TempDir(), "k")
|
||||
if err := os.WriteFile(p, []byte(hex.EncodeToString([]byte(key))+"\n"), 0o600); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
return p
|
||||
}
|
||||
|
||||
// The consequence the runbook relies on: the SAME key opens every vaulted console password of the restored copy.
|
||||
func TestCheck_SameKeyOpensEveryConsolePassword(t *testing.T) {
|
||||
r, err := check(seededCopy(t), keyFile(t, testKey))
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if r.hosts != 2 || r.opened != 1 || r.failed != 0 || r.absent != 1 {
|
||||
t.Fatalf("got %+v, want hosts=2 opened=1 failed=0 absent=1", r)
|
||||
}
|
||||
}
|
||||
|
||||
// A different key opens nothing — so "opened" is evidence that the key matches, not that the column is readable.
|
||||
func TestCheck_WrongKeyOpensNothing(t *testing.T) {
|
||||
r, err := check(seededCopy(t), keyFile(t, "a-completely-different-key-32byt"))
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if r.opened != 0 || r.failed != 1 {
|
||||
t.Fatalf("got %+v, want opened=0 failed=1", r)
|
||||
}
|
||||
}
|
||||
@@ -1,3 +1,15 @@
|
||||
## hub-db-backup 1.0 — the hub database leaves DooPlex every night, encrypted, and is restore-tested weekly (R-173) (2026-10-05)
|
||||
|
||||
New, in `scripts/hub-db-backup/` (versioned from day one, R-231): `felhom-hub-db-backup` (02:30 — copy the hub's newest
|
||||
nightly snapshot out of the pod, refuse a snapshot >26 h old, a size mismatch, a failed `integrity_check` or no hosts;
|
||||
push it to ep0's PBS ns `operator`, `--crypt-mode encrypt`, write-only token; write the success timestamp only after the
|
||||
push), `felhom-hub-db-restore-test` (Sun 04:30 — restore the newest copy with the read-only token, refuse a copy >50 h
|
||||
old, a failed integrity check, no hosts, or any console password stored readable), four systemd units, `install.sh`
|
||||
(does not enable the timers). `test_hub_db_backup.py`: 15 tests with fake `kubectl` / `proxmox-backup-client` (nothing
|
||||
reaches the hub, PBS or ep0); red-proofs P1–P9 (`documentation/audits/hub-db-offsite-2026-10-05/partC/red-proof.txt` —
|
||||
P4 and P5 needed a stricter test first). Installed on DooPlex 2026-10-05; timers enabled after the first manual run.
|
||||
Runbook: `documentation/runbooks/RUNBOOK-hub-db-offsite-backup.md`.
|
||||
|
||||
## felhom-host-install.sh 1.31.0 — the root-owned files come from the agent's config bundle (R-840) (2026-10-04)
|
||||
|
||||
Needs a vouched agent ≥ 0.143.0 for the bundle (hub ≥ 0.133.0 serves its sha); an older vouched agent: the per-file
|
||||
|
||||
Reference in New Issue
Block a user