From afba622fb63d22320160fa4462fc5aba013d8ff0 Mon Sep 17 00:00:00 2001 From: kisfenyo Date: Mon, 5 Oct 2026 16:14:13 +0200 Subject: [PATCH] =?UTF-8?q?R-173=20option=20A=20in=20force=20(hub=20DB=20n?= =?UTF-8?q?ightly=20to=20ep0,=20restore-tested,=20alarmed);=20R-519=20prov?= =?UTF-8?q?en=20live=20on=209202=20and=20closed;=20R-173/R-232=20narrowed;?= =?UTF-8?q?=20R-882..R-885=20opened=20(332=20->=20335);=20runbook=20=C2=A7?= =?UTF-8?q?3=20tested;=20hubdb-check?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS --- CONTEXT.md | 12 ++ STATUS.md | 45 ++++++- .../architecture/00-capability-map.md | 1 + .../architecture/07-backup-architecture.md | 5 +- .../partC/alarm-rule-test.txt | 26 ++++ .../partC/bf_test.yml | 52 ++++++++ .../partD/alarm-drill.txt | 5 + .../partD/push-1.txt | 27 ++++ .../partD/r519/00-upgrade-9202.txt | 1 + .../partD/r519/01-deploy.txt | 3 + .../partD/r519/02-run1.txt | 36 +++++ .../partD/r519/03-before-cut.txt | 12 ++ .../partD/r519/04-run2-cut.txt | 34 +++++ .../partD/r519/05-after-cut.txt | 23 ++++ .../partD/r519/06-run3.txt | 12 ++ .../partD/r519/07-teardown.txt | 8 ++ .../partD/restore-procedure/drill.txt | 9 ++ .../partD/restore-procedure/red-proof.txt | 7 + .../partD/restore-test-1.txt | 17 +++ .../partD/token-limits-and-paperkey.txt | 12 ++ documentation/backlog/CLOSED-ITEMS.md | 10 ++ documentation/backlog/OPEN-ITEMS.md | 17 ++- .../runbooks/RUNBOOK-hub-db-offsite-backup.md | 126 +++++++++++++----- hub/CHANGELOG.md | 3 + hub/cmd/hubdb-check/main.go | 75 +++++++++++ hub/cmd/hubdb-check/main_test.go | 65 +++++++++ scripts/CHANGELOG.md | 12 ++ 27 files changed, 611 insertions(+), 44 deletions(-) create mode 100644 documentation/audits/hub-db-offsite-2026-10-05/partC/alarm-rule-test.txt create mode 100644 documentation/audits/hub-db-offsite-2026-10-05/partC/bf_test.yml create mode 100644 documentation/audits/hub-db-offsite-2026-10-05/partD/alarm-drill.txt create mode 100644 documentation/audits/hub-db-offsite-2026-10-05/partD/push-1.txt create mode 100644 documentation/audits/hub-db-offsite-2026-10-05/partD/r519/00-upgrade-9202.txt create mode 100644 documentation/audits/hub-db-offsite-2026-10-05/partD/r519/01-deploy.txt create mode 100644 documentation/audits/hub-db-offsite-2026-10-05/partD/r519/02-run1.txt create mode 100644 documentation/audits/hub-db-offsite-2026-10-05/partD/r519/03-before-cut.txt create mode 100644 documentation/audits/hub-db-offsite-2026-10-05/partD/r519/04-run2-cut.txt create mode 100644 documentation/audits/hub-db-offsite-2026-10-05/partD/r519/05-after-cut.txt create mode 100644 documentation/audits/hub-db-offsite-2026-10-05/partD/r519/06-run3.txt create mode 100644 documentation/audits/hub-db-offsite-2026-10-05/partD/r519/07-teardown.txt create mode 100644 documentation/audits/hub-db-offsite-2026-10-05/partD/restore-procedure/drill.txt create mode 100644 documentation/audits/hub-db-offsite-2026-10-05/partD/restore-procedure/red-proof.txt create mode 100644 documentation/audits/hub-db-offsite-2026-10-05/partD/restore-test-1.txt create mode 100644 documentation/audits/hub-db-offsite-2026-10-05/partD/token-limits-and-paperkey.txt create mode 100644 hub/cmd/hubdb-check/main.go create mode 100644 hub/cmd/hubdb-check/main_test.go diff --git a/CONTEXT.md b/CONTEXT.md index 6b1eba3e..ba76795d 100644 --- a/CONTEXT.md +++ b/CONTEXT.md @@ -16,6 +16,18 @@ > and holds nothing of its own; this file does hold its own content, namely the standing rulings below. +> **2026-10-05 (evening) — the hub database off DooPlex (hub v0.136.0; operator rulings `09` 125–127).** R-173 option A +> IN FORCE: hub `internal/dbsnap` writes `VACUUM INTO /data/snapshots/hub-.db` at 02:00 Budapest (keep 2, `ErrBusy` +> on overlap, start-up catch-up when >24 h; `05` §16.3); DooPlex `scripts/hub-db-backup/` (installed by `install.sh` to +> `/usr/local/sbin`, units + timers in `/etc/systemd/system`, config `/etc/felhom-hub-backup/{env,token-push, +> token-restore,enc.key}` all root 0600) pushes at 02:30 to ep0 `felhom-offsite` ns `operator` as `dooplex-hub@pbs!push` +> (`DatastoreBackup`), restore-tests Sun 04:30 as `!restore` (`DatastoreReader`); ep0 prune job `prune-operator-hubdb` +> (03:45, keep-daily 14, keep-weekly 8, max-depth 0). Success-only textfile metrics → `HubDBBackupStale` (26 h, critical) +> and `HubDBRestoreTestStale` (8 d) in homelab-manifests `backup-freshness`. `hub/cmd/hubdb-check` opens a restored copy +> with the seal key (runbook §3). Hub PVC 2 Gi + label `enabled`. Collateral: Longhorn instance-manager restart (R-882), +> zipline pinned 4.7.0 (R-883). R-519 proven live on 9202 (now controller 0.296.0) and CLOSED. Register 332 → 335. +> Report: `REPORT-hub-db-offsite-2026-10-05.md`. + > **2026-10-05 (late afternoon) — the hub's own safety, boxes left behind, the agent's root grants (hub v0.135.0, controller > v0.296.0, agent v0.146.1 + bundle `42333e96…`, golden 0.296.0 vouched with agent 0.146.1, min_agent 0.131.0).** CC decisions > 119–124, *operator may reverse*. Hub: R-135 a cookie-less state change needs Basic + `X-Felhom-Operator` (`05` §16.1); R-133 diff --git a/STATUS.md b/STATUS.md index f6ad432e..d3384314 100644 --- a/STATUS.md +++ b/STATUS.md @@ -3,9 +3,44 @@ **Ready for the first real tester (Tester-2): yes. Tester 2 (a laptop, off at night) was offline again; nothing was sent to it.** -**Updated 2026-10-05 (late afternoon, the hub-safety session): every box of ours healthy. The hub refuses forged form -posts and keeps the console passwords locked; the System page shows boxes left behind; the agent can no longer make -itself root. Report: `REPORT-hub-safety-2026-10-05.md`.** +**Updated 2026-10-05 (evening, the hub-database session): every box of ours healthy. The hub database now leaves +DooPlex every night, locked, to ep0, is test-restored every Sunday, and an alarm mails you if either stops. Report: +`REPORT-hub-db-offsite-2026-10-05.md`.** + +## Tonight (2026-10-05, evening): the hub database off DooPlex; the cut-backup check on the scratch box + +**Decisions:** none of mine. Yours (`09` 125–127): ep0 for the copy; the scratch-box restart allowed; the agent's three +by-design abilities stay. Today you also chose: grow the hub's disk to 2 GiB, and restart Longhorn's disk manager. + +**What works now (proven live):** +- **The hub makes a clean copy of its database every night at 02:00** (hub 0.136.0). The first one: 353 MB, 44 s. +- **DooPlex checks it, locks it with a key ep0 never sees, and sends it to ep0 at 02:30.** First send: 7 s. +- **Every Sunday at 04:30 DooPlex takes the copy back from ep0 and checks it**: it opens, it is whole, every console + password in it is still locked. Done once by hand today: 4 boxes, 4 locked passwords, 0 readable. +- **The two ep0 accounts can only do their one job:** the sending one cannot delete, the checking one cannot write, + neither can see the households' backups. (The sending one can read back its own locked copies — that is how ep0 works.) +- **Your two saved keys work:** a key rebuilt from the paper copy you saved opened the copy, and your saved lock key + opened all 4 console passwords in it (a wrong key opened none). +- **The alarm:** proven by Prometheus' own rule test; the real alarm mail is below. +- **A backup cut off by a restart is now said on the backup pages** (scratch box): the page said so, the restore point + kept the older time of the part that was not redone, and the next full backup cleared the message. + +**What broke, and what I did:** +- **The hub's disk would not grow:** Longhorn's disk manager on DooPlex was stuck. My first try (stopping the hub so the + disk could grow offline) did not work and kept the hub **down about 9.5 minutes**. You approved restarting the disk + manager: all 77 disks were back in under 2 minutes, and the hub's disk is 2 GiB now. +- **Zipline did not come back after that restart:** it is set to "always the newest", so it pulled a new release that + refused its database. I pinned it to the previous release; it runs and its database is updated. 7 more apps on + DooPlex use "always the newest" (new row). + +**Register:** 332 → 335 rows (1 closed: the cut-backup check; 4 opened: the Longhorn fault, the "always newest" apps, a +monitoring sync drift, script tests not in CI). + +**Needs you:** +1. **Nothing urgent.** If you do nothing, the copy runs every night and you get a mail only if it stops. +2. **When convenient:** pin the 7 other DooPlex apps that use "always the newest" (or tell me to list them for you). If + you do nothing, any restart may upgrade one of them by surprise, as it did zipline. +3. **Tester 2's one-time step** is unchanged (below). ## Today (2026-10-05, late afternoon): the hub's own safety; boxes left behind; the agent's admin rights @@ -38,7 +73,7 @@ itself root. Report: `REPORT-hub-safety-2026-10-05.md`.** minutes. Both figures are on the page now. **Needs you:** -1. **Where the hub database's off-site copy goes** (it holds every box's keys and your customers' settings): +1. **(DECIDED 2026-10-05 evening: A, done — see Tonight)** **Where the hub database's off-site copy goes** (it holds every box's keys and your customers' settings): - **A — my pick: ep0's backup server**, encrypted on DooPlex before it leaves, with a weekly restore test and an alarm mail. Costs one small change on ep0 (a write-only account) and keeping two keys in your password manager. - **B: a separate Hetzner Storage Box account** with restic. More new parts to look after than A. @@ -47,7 +82,7 @@ itself root. Report: `REPORT-hub-safety-2026-10-05.md`.** `documentation/runbooks/RUNBOOK-hub-db-offsite-backup.md`. - **Either way, first:** put the hub's lock key (`OFFSITE_SECRET_KEY`) in your password manager — without it a copy of the database cannot open the console passwords. -2. **The power-cut-during-backup check (R-519) on the scratch box:** the permission check refused my restarting the +2. **(DONE 2026-10-05 evening — see Tonight)** **The power-cut-during-backup check (R-519) on the scratch box:** the permission check refused my restarting the controller in the middle of a backup. Say "go" and the next session does it once on 9202; if not, the fix stays proven by tests only. 3. **Three things the agent can still do, by design** (each written in `03` §3.1): pick which controller image its own diff --git a/documentation/architecture/00-capability-map.md b/documentation/architecture/00-capability-map.md index 3a287d45..37697580 100644 --- a/documentation/architecture/00-capability-map.md +++ b/documentation/architecture/00-capability-map.md @@ -235,6 +235,7 @@ likewise silent. Evidence: `audits/DRILL-r361-2026-08-22/evidence/06-part3-decis | Secrets hygiene: bearer in k8s Secret, no secrets in git, single-quote credential store | hub v0.53, conventions | **IMPLEMENTED** | 07-13 closing bundle | | | Operator login password changeable from UI | hub v0.54 | **IMPLEMENTED** | 07-13 | | | **The operator sees boxes left behind: per-customer controller floors with their age (and which a global raise cannot move), each box's agent against the vouched one, a 7-day "agent behind" alarm and a "floor raise skipped boxes" mail** | hub **v0.135.0** (R-604, R-530) | **IMPLEMENTED — the page PROVEN-LIVE, the alarm and the mail unit-proven** | `audits/hub-safety-2026-10-05/partD/` (live System page: three per-customer floors, Tester 2 `0.142.0 → 0.145.0`); `osupdates/r530_agent_alarm_test.go`, `web/r604_floor_held_back_test.go` | the mail fires only on a GLOBAL raise below an override — not exercised live | +| **The hub database survives the loss of DooPlex: a nightly consistent copy, encrypted, on ep0; restore-tested weekly; an alarm when either stops** | hub **v0.136.0** (R-173), `scripts/hub-db-backup/`, homelab-manifests rules | **PROVEN-LIVE (2026-10-05)** — first push, ep0 listing, restore test (4 hosts, 4 sealed, 0 readable), token limits, a key rebuilt from the paper copy decrypts, the saved seal key opens 4/4 console passwords in the restored copy; the alarm by `promtool` rule test + red-proofs | `audits/hub-db-offsite-2026-10-05/`; `runbooks/RUNBOOK-hub-db-offsite-backup.md` | Runbook §3 steps 4–5 (into a live PVC) not exercised (R-173) | | **The operator surface refuses a cross-site form post on BOTH login paths (session token; Basic auth + `X-Felhom-Operator`)** | hub **v0.135.0** (R-135) | **PROVEN-LIVE (2026-10-05)** | `audits/hub-safety-2026-10-05/partA/live.txt` (Basic, no header → 403 even on an unknown path; with the header → passes; header without credentials → 401); `web/r135_csrf_test.go` (39 paths) | | | Box operating-system security updates (Proxmox host, guest Debian, Docker engine) | agent v0.143.0, hub v0.133.0 | **PARTIAL — the GUEST and HOST Debian fast lanes and the DOCKER engine slow lane are PROVEN-LIVE (2026-10-04), with the System page, the fleet view and the alarms; the KERNEL lane is MISSING** | Guest: `audits/os-guest-lane-2026-10-04/`. Host + fleet + alarms: `audits/os-host-lane-2026-10-04/`. Docker + System page + crash guard: `audits/os-docker-crash-2026-10-04/` — live-restore on with the same container ids on every box; Docker 29.8.2 on both demo boxes; operator-approved Docker release; a signed undo and a signed ring-1 step; a replay refused; the crash guard restarted demo-hp twice and kept it off the third time. Design `architecture/11-os-updates.md` §5.8, §5.9, §8 | **No automatic undo** (guest: last night's backup; host: by-hand runbook; Docker: a signed undo job); existing boxes get root-owned files by the signed config bundle since 2026-10-04 (R-840 CLOSED; a box from before agent 0.143.0 needs one by-hand bootstrap — Tester 2: R-862; `audits/r840-config-bundle-2026-10-04/`); test approvals now end with the test (R-859); the agent's sudoers is root-equivalent (R-861); the kernel lane (R-836); facts reach the hub late after a boot (R-853). **2026-10-05 (agent v0.144.1):** R8 measures the real download (R-865); a killed pass still reports (R-868, live); the debug pass runs with the hub away (R-866, live); **a power cut mid-update was proven by day on demo-hp — the box came back by itself in 37 s, but the next pass fails until `dpkg --configure -a` is run by hand (R-876, P2, open)** — `audits/night-fixes-2026-10-05/`. **2026-10-05 afternoon (agent v0.145.0): R-876 FIXED and proven live — after a second crash mid-unpack the next pass repaired dpkg by itself (`REPAIR … journal=1`) and finished** — `audits/catchup-2026-10-05/partD/` | | **An ENGLISH-SPEAKING household's first hour: download, install, pair, bind, claim, two apps** | controller **v0.259.0** + hub **v0.119.0** + ISO 1.29.0 + the whole catalog | **PROVEN-LIVE on 0.258.0 with one blocker; THE BLOCKER IS FIXED AND PROVEN, THE WALK IS NOT REPEATED** | `audits/DRILL-first-hour-en-0258-2026-09-20.md` — a fresh install 2026-09-20, one intervention (R-494), stop rule not reached. Then `audits/i18n-closing-2026-09-21/live/` — the three blockers fixed and each proven on a live box or in the operator's inbox: the claim page answers English through the real cookie path; the Backup page's tier names follow the language; and the setup mail carries **four plain-ASCII English words** where the drill's carried `képző-szkítia-ásatás`, one day apart in the same inbox. | **R-596, R-597 and R-598 are CLOSED.** What this row still does NOT claim: **the fixed journey has not been walked end to end by a stranger on a fresh install.** Three fixes proven at the endpoint are not an hour proven by a person, and this project's own rule is that fixes are not a journey (see the recovery-journey row). **Also not walked:** the recovery code (needs ep0), backup/restore/remove/power-cut (proven 2026-09-14), and the two Backup-page *warnings* themselves — guest 9201 is healthy and a healthy box renders none, so they are covered by handler render tests, not live. **Verdict: nothing known now stands between an English-speaking tester and their box — and that is a different sentence from "the walk passed".** | diff --git a/documentation/architecture/07-backup-architecture.md b/documentation/architecture/07-backup-architecture.md index cbadf1cd..552d9a3f 100644 --- a/documentation/architecture/07-backup-architecture.md +++ b/documentation/architecture/07-backup-architecture.md @@ -140,8 +140,9 @@ mert a doboz vagy a vezérlő újraindult…"), kept until a run ends with every before this, both pages said nothing and the synthesised „Utolsó adatbázis mentés … OK" was read off the fresh `.sql` the cut run left beside last night's tars; that line now reads failed after a cut. Each restore point's time was already its OLDEST part (the data block, v0.275.0) — so a torn unit is dated by its stale tars, never by its new dump. -*Live: unit-proven and red-proved; the live cut on 9202 was refused by the permission check and waits for the -operator (R-519 narrowed).* +*Live: PROVEN 2026-10-05 on 9202 (operator ruling 126): a run cut by a controller restart between bookstack's config +dump and its database volume → both pages carry the notice, the restore point reads the older volume's time, the next +complete run clears it (`audits/hub-db-offsite-2026-10-05/partD/r519/`; R-519 closed).* ### Lane 2 — the operator: guest and host recovery diff --git a/documentation/audits/hub-db-offsite-2026-10-05/partC/alarm-rule-test.txt b/documentation/audits/hub-db-offsite-2026-10-05/partC/alarm-rule-test.txt new file mode 100644 index 00000000..5f930637 --- /dev/null +++ b/documentation/audits/hub-db-offsite-2026-10-05/partC/alarm-rule-test.txt @@ -0,0 +1,26 @@ +## promtool, in pod/prometheus-55b675779d-t8c74, 2026-10-05T13:42:17Z +### green + SUCCESS + +rc=0 +### red: threshold 26h -> 260h + FAILED: + alertname: HubDBBackupStale, time: 1d2h40m, + exp:[ + 0: + Labels:{alertname="HubDBBackupStale", component="backup", instance="dooplex", severity="critical"} + Annotations:{description="No successful push of the hub DB snapshot to ep0 for >26h (daily at 02:30). Check `journalctl -u felhom-hub-db-backup.service`, the tunnel `systemctl status felhom-ep0-pbs-tunnel`, and that the hub logs `db snapshot written` at 02:00.", summary="The Felhom hub database has not reached ep0 for 26 h (R-173)"} + ], + got:[] + + +command terminated with exit code 1 +rc=0 +### red: absent() removed + SUCCESS + +### red (re-run): absent() removed from HubDBBackupStale — first attempt above did NOT apply (indent mismatch) + FAILED: + alertname: HubDBBackupStale, time: 40m, + Labels:{alertname="HubDBBackupStale", component="backup", severity="critical"} + got:[] diff --git a/documentation/audits/hub-db-offsite-2026-10-05/partC/bf_test.yml b/documentation/audits/hub-db-offsite-2026-10-05/partC/bf_test.yml new file mode 100644 index 00000000..ee9959a0 --- /dev/null +++ b/documentation/audits/hub-db-offsite-2026-10-05/partC/bf_test.yml @@ -0,0 +1,52 @@ +rule_files: [bf.yml] +evaluation_interval: 1m +tests: + # 1. the push stops: last success at t=0, never again → fires once 26 h + 30 m have passed + - interval: 5m + input_series: + - series: 'felhom_hub_db_backup_last_success_timestamp_seconds{instance="dooplex"}' + values: '0+0x400' + - series: 'felhom_hub_db_restore_test_last_success_timestamp_seconds{instance="dooplex"}' + values: '0+0x400' + alert_rule_test: + - eval_time: 26h + alertname: HubDBBackupStale + exp_alerts: [] + - eval_time: 26h40m + alertname: HubDBBackupStale + exp_alerts: + - exp_labels: {severity: critical, component: backup, instance: dooplex} + exp_annotations: + summary: "The Felhom hub database has not reached ep0 for 26 h (R-173)" + description: "No successful push of the hub DB snapshot to ep0 for >26h (daily at 02:30). Check `journalctl -u felhom-hub-db-backup.service`, the tunnel `systemctl status felhom-ep0-pbs-tunnel`, and that the hub logs `db snapshot written` at 02:00." + # 2. healthy: the push succeeds every 24 h → never fires + - interval: 1h + input_series: + - series: 'felhom_hub_db_backup_last_success_timestamp_seconds{instance="dooplex"}' + values: '0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 86400 86400 86400 86400 86400 86400 86400 86400 86400 86400 86400 86400 86400 86400 86400 86400 86400 86400 86400 86400 86400 86400 86400 86400 172800 172800 172800' + - series: 'felhom_hub_db_restore_test_last_success_timestamp_seconds{instance="dooplex"}' + values: '0+0x50' + alert_rule_test: + - eval_time: 50h + alertname: HubDBBackupStale + exp_alerts: [] + # 3. the metric never existed (script never ran) → fires on absent() + - interval: 5m + input_series: + - series: 'up{job="node"}' + values: '1+0x30' + alert_rule_test: + - eval_time: 40m + alertname: HubDBBackupStale + exp_alerts: + - exp_labels: {severity: critical, component: backup} + exp_annotations: + summary: "The Felhom hub database has not reached ep0 for 26 h (R-173)" + description: "No successful push of the hub DB snapshot to ep0 for >26h (daily at 02:30). Check `journalctl -u felhom-hub-db-backup.service`, the tunnel `systemctl status felhom-ep0-pbs-tunnel`, and that the hub logs `db snapshot written` at 02:00." + - eval_time: 2h + alertname: HubDBRestoreTestStale + exp_alerts: + - exp_labels: {severity: warning, component: backup} + exp_annotations: + summary: "The hub database copy on ep0 has not passed a restore test for 8 days (R-173)" + description: "The weekly restore test (Sun 04:30) has not succeeded for >8 days. Check `journalctl -u felhom-hub-db-restore-test.service`." diff --git a/documentation/audits/hub-db-offsite-2026-10-05/partD/alarm-drill.txt b/documentation/audits/hub-db-offsite-2026-10-05/partD/alarm-drill.txt new file mode 100644 index 00000000..d9874a2b --- /dev/null +++ b/documentation/audits/hub-db-offsite-2026-10-05/partD/alarm-drill.txt @@ -0,0 +1,5 @@ +## alarm drill start 2026-10-05T13:51:46Z: success file moved aside (absent case) +backup_freshness.prom +fan_metrics.prom +felhom_hub_db_restore.prom +node_housekeeping.prom diff --git a/documentation/audits/hub-db-offsite-2026-10-05/partD/push-1.txt b/documentation/audits/hub-db-offsite-2026-10-05/partD/push-1.txt new file mode 100644 index 00000000..640e2c2e --- /dev/null +++ b/documentation/audits/hub-db-offsite-2026-10-05/partD/push-1.txt @@ -0,0 +1,27 @@ +## manual push via unit, 2026-10-05T13:50:30Z +Result=success +ExecMainStatus=0 +2026-10-05T15:50:08+02:00 dooplex systemd[1]: Starting felhom-hub-db-backup.service - Felhom: push the hub DB snapshot to ep0 (R-173)... +2026-10-05T15:50:09+02:00 dooplex felhom-hub-db-backup[3036511]: felhom-hub-db-backup: snapshot hub-20261005T123037Z.db, 79 min old +2026-10-05T15:50:18+02:00 dooplex felhom-hub-db-backup[3036511]: felhom-hub-db-backup: checked: 369807360 bytes, integrity ok, 4 host(s) +2026-10-05T15:50:18+02:00 dooplex felhom-hub-db-backup[3037790]: Starting backup: [operator]:host/dooplex-hub/2026-10-05T13:50:18Z +2026-10-05T15:50:18+02:00 dooplex felhom-hub-db-backup[3037790]: Client name: dooplex +2026-10-05T15:50:18+02:00 dooplex felhom-hub-db-backup[3037790]: Starting backup protocol: Mon Oct 5 15:50:18 2026 +2026-10-05T15:50:18+02:00 dooplex felhom-hub-db-backup[3037790]: Using encryption key from '/etc/felhom-hub-backup/enc.key'.. +2026-10-05T15:50:18+02:00 dooplex felhom-hub-db-backup[3037790]: Encryption key fingerprint: b2:19:bf:36:3b:97:3d:6c +2026-10-05T15:50:18+02:00 dooplex felhom-hub-db-backup[3037790]: No previous manifest available. +2026-10-05T15:50:18+02:00 dooplex felhom-hub-db-backup[3037790]: Upload directory '/var/lib/felhom-hub-backup/stage' to 'dooplex-hub@pbs!push@127.0.0.1:18007:felhom-offsite' as hubdb.pxar.didx +2026-10-05T15:50:25+02:00 dooplex felhom-hub-db-backup[3037790]: hubdb.pxar: had to backup 352.676 MiB of 352.676 MiB (compressed 17.242 MiB) in 6.58 s (average 53.604 MiB/s) +2026-10-05T15:50:25+02:00 dooplex felhom-hub-db-backup[3037790]: Uploaded backup catalog (56 B) +2026-10-05T15:50:25+02:00 dooplex felhom-hub-db-backup[3037790]: Duration: 7.19s +2026-10-05T15:50:25+02:00 dooplex felhom-hub-db-backup[3037790]: End Time: Mon Oct 5 15:50:25 2026 +2026-10-05T15:50:25+02:00 dooplex felhom-hub-db-backup[3036511]: felhom-hub-db-backup: pushed hub-20261005T123037Z.db to ep0 (ns operator) in 7 s +2026-10-05T15:50:25+02:00 dooplex felhom-hub-db-backup[3036511]: felhom-hub-db-backup: success signal written +2026-10-05T15:50:30+02:00 dooplex systemd[1]: felhom-hub-db-backup.service: Deactivated successfully. +2026-10-05T15:50:30+02:00 dooplex systemd[1]: Finished felhom-hub-db-backup.service - Felhom: push the hub DB snapshot to ep0 (R-173). +2026-10-05T15:50:30+02:00 dooplex systemd[1]: felhom-hub-db-backup.service: Consumed 9.566s CPU time, 584.8M memory peak. +## textfile +# HELP felhom_hub_db_backup_last_success_timestamp_seconds Last successful push of the hub DB snapshot to ep0 (R-173). +# TYPE felhom_hub_db_backup_last_success_timestamp_seconds gauge +felhom_hub_db_backup_last_success_timestamp_seconds 1791208225 +felhom_hub_db_backup_last_success_bytes 369807360 diff --git a/documentation/audits/hub-db-offsite-2026-10-05/partD/r519/00-upgrade-9202.txt b/documentation/audits/hub-db-offsite-2026-10-05/partD/r519/00-upgrade-9202.txt new file mode 100644 index 00000000..45c972b9 --- /dev/null +++ b/documentation/audits/hub-db-offsite-2026-10-05/partD/r519/00-upgrade-9202.txt @@ -0,0 +1 @@ +gitea.dooplex.hu/admin/felhom-controller:0.296.0 Up 6 seconds (healthy) diff --git a/documentation/audits/hub-db-offsite-2026-10-05/partD/r519/01-deploy.txt b/documentation/audits/hub-db-offsite-2026-10-05/partD/r519/01-deploy.txt new file mode 100644 index 00000000..c8621e68 --- /dev/null +++ b/documentation/audits/hub-db-offsite-2026-10-05/partD/r519/01-deploy.txt @@ -0,0 +1,3 @@ +## deploy bookstack 2026-10-05T13:53:26Z (fields: DOMAIN APP_KEY DB_PASSWORD ADMIN_PASSWORD; SUBDOMAIN = catalog default) + +HTTP 202 diff --git a/documentation/audits/hub-db-offsite-2026-10-05/partD/r519/02-run1.txt b/documentation/audits/hub-db-offsite-2026-10-05/partD/r519/02-run1.txt new file mode 100644 index 00000000..966106f6 --- /dev/null +++ b/documentation/audits/hub-db-offsite-2026-10-05/partD/r519/02-run1.txt @@ -0,0 +1,36 @@ +## run 1 (baseline) start 2026-10-05T14:04:36Z + +HTTP 200 +{"ok":true,"data":{"enabled":true,"running":true}} + +HTTP 200 + +## run 1 end 2026-10-05T14:05:17Z: {"ok":true,"data":{"db_dump":{"count":2,"duration":"32.784678377s","last_run":"2026-10-05T14:05:12.109790785Z","success":true},"enabled":true,"running":false}} +2026/10/05 14:04:44 data_versions.go:131: [DEBUG] [backup] bookstack: stamped volume-dumps/bookstack_bookstack_config.tar (70144 B) with pins [lscr.io/linuxserver/bookstack:26.09.1@sha256:99cd1f5707c1911afad213adec5c9739763b76f843d1477142231834ecdcb6f7 mariadb:12.3@sha256:805c8e104bd563d5bfa24fadd3f31cd419ea859cb5277f32b5dbf2db714f9ed1] +2026/10/05 14:04:44 backup.go:848: [DEBUG] [backup] Dumping volume bookstack_bookstack_db_data for bookstack +2026/10/05 14:04:45 backup.go:873: [INFO] [backup] Volume dump: bookstack/bookstack_bookstack_db_data → 153.4 MB +2026/10/05 14:04:45 data_versions.go:131: [DEBUG] [backup] bookstack: stamped volume-dumps/bookstack_bookstack_db_data.tar (160868864 B) with pins [lscr.io/linuxserver/bookstack:26.09.1@sha256:99cd1f5707c1911afad213adec5c9739763b76f843d1477142231834ecdcb6f7 mariadb:12.3@sha256:805c8e104bd563d5bfa24fadd3f31cd419ea859cb5277f32b5dbf2db714f9ed1] +2026/10/05 14:04:45 backup.go:990: [INFO] [backup] Restarting bookstack after volume dump +2026/10/05 14:04:51 backup.go:974: [INFO] [backup] Stopping paperless-ngx for safe volume dump +2026/10/05 14:04:58 backup.go:848: [DEBUG] [backup] Dumping volume paperless-ngx_paperless_redis_data for paperless-ngx +2026/10/05 14:04:58 backup.go:873: [INFO] [backup] Volume dump: paperless-ngx/paperless-ngx_paperless_redis_data → 7.4 MB +2026/10/05 14:04:58 data_versions.go:131: [DEBUG] [backup] paperless-ngx: stamped volume-dumps/paperless-ngx_paperless_redis_data.tar (7792128 B) with pins [ghcr.io/paperless-ngx/paperless-ngx:2.20.15 postgres:18-alpine redis:7-alpine] +2026/10/05 14:04:58 backup.go:848: [DEBUG] [backup] Dumping volume paperless-ngx_paperless_data for paperless-ngx +2026/10/05 14:04:59 backup.go:873: [INFO] [backup] Volume dump: paperless-ngx/paperless-ngx_paperless_data → 5.8 MB +2026/10/05 14:04:59 data_versions.go:131: [DEBUG] [backup] paperless-ngx: stamped volume-dumps/paperless-ngx_paperless_data.tar (6052352 B) with pins [ghcr.io/paperless-ngx/paperless-ngx:2.20.15 postgres:18-alpine redis:7-alpine] +2026/10/05 14:04:59 backup.go:848: [DEBUG] [backup] Dumping volume paperless-ngx_paperless_postgres_data for paperless-ngx +2026/10/05 14:04:59 backup.go:873: [INFO] [backup] Volume dump: paperless-ngx/paperless-ngx_paperless_postgres_data → 68.8 MB +2026/10/05 14:04:59 data_versions.go:131: [DEBUG] [backup] paperless-ngx: stamped volume-dumps/paperless-ngx_paperless_postgres_data.tar (72107008 B) with pins [ghcr.io/paperless-ngx/paperless-ngx:2.20.15 postgres:18-alpine redis:7-alpine] +2026/10/05 14:04:59 backup.go:990: [INFO] [backup] Restarting paperless-ngx after volume dump +2026/10/05 14:05:11 backup.go:974: [INFO] [backup] Stopping privatebin for safe volume dump +2026/10/05 14:05:11 backup.go:848: [DEBUG] [backup] Dumping volume privatebin_privatebin_data for privatebin +2026/10/05 14:05:12 backup.go:873: [INFO] [backup] Volume dump: privatebin/privatebin_privatebin_data → 1.5 KB +2026/10/05 14:05:12 data_versions.go:131: [DEBUG] [backup] privatebin: stamped volume-dumps/privatebin_privatebin_data.tar (1536 B) with pins [privatebin/pdo:2.0.6@sha256:4c141b2326f8b353598ce9ce7507a9cfecf2dad5c60a39fea903d430e296d8f5] +2026/10/05 14:05:12 backup.go:986: [INFO] [backup] privatebin NOT restarted after the volume dump: the household stopped it meanwhile +2026/10/05 14:05:12 backup.go:684: [INFO] [backup] App-data backup completed: 2 databases (483.2 KB total), 3 volume dump(s) (32.785s) +2026/10/05 14:05:12 recovery_unit.go:289: [INFO] [backup] Recovery unit captured for bookstack → /mnt/sys_drive/felhom-data/backups/primary/bookstack (images=2, secrets-referenced=3, data_keys=0, portable-carried=2/2, withheld=1) +2026/10/05 14:05:12 recovery_unit.go:289: [INFO] [backup] Recovery unit captured for paperless-ngx → /mnt/felhom-drives/scratch_hdd/userdata/paperless-ngx/backups/primary/paperless-ngx (images=3, secrets-referenced=3, data_keys=0, portable-carried=2/2, withheld=1) +2026/10/05 14:05:12 recovery_unit.go:289: [INFO] [backup] Recovery unit captured for privatebin → /mnt/sys_drive/felhom-data/backups/primary/privatebin (images=1, secrets-referenced=0, data_keys=0, portable-carried=0/0, withheld=0) + +/var/lib/felhom/docker/volumes/felhom-controller-data/_data/data/appdata-run.json +/var/lib/docker/volumes/felhom-controller-data/_data/data/appdata-run.json diff --git a/documentation/audits/hub-db-offsite-2026-10-05/partD/r519/03-before-cut.txt b/documentation/audits/hub-db-offsite-2026-10-05/partD/r519/03-before-cut.txt new file mode 100644 index 00000000..1b61aaf0 --- /dev/null +++ b/documentation/audits/hub-db-offsite-2026-10-05/partD/r519/03-before-cut.txt @@ -0,0 +1,12 @@ +## after run 1, 2026-10-05T14:05:34Z +### run record +{"running":false,"started_at":"0001-01-01T00:00:00Z","interrupted":"0001-01-01T00:00:00Z"} +### restore points (bookstack) +{"ok":true,"data":[{"time":"2026-10-05T14:04:39Z","short_id":"helyi","tier":1,"drive_label":"Belső SSD (rendszer)"}]} + +HTTP 200 + +### volume dump file times +total 157172 +-rw-r--r-- 1 root root 70144 2026-10-05T14:04:44 bookstack_bookstack_config.tar +-rw-r--r-- 1 root root 160868864 2026-10-05T14:04:44 bookstack_bookstack_db_data.tar diff --git a/documentation/audits/hub-db-offsite-2026-10-05/partD/r519/04-run2-cut.txt b/documentation/audits/hub-db-offsite-2026-10-05/partD/r519/04-run2-cut.txt new file mode 100644 index 00000000..3deaf319 --- /dev/null +++ b/documentation/audits/hub-db-offsite-2026-10-05/partD/r519/04-run2-cut.txt @@ -0,0 +1,34 @@ +## run 2 start 2026-10-05T14:05:50Z +HTTP 200 +### watcher +match 2026-10-05T14:05:58.856456447Z +restarted 2026-10-05T14:05:59.669364637Z rc=0 +Up 10 seconds (healthy) +### record at the moment of the cut +{"running":true,"started_at":"2026-10-05T14:05:53.282495446Z","interrupted":"0001-01-01T00:00:00Z"} +### record after restart +{"running":false,"started_at":"2026-10-05T14:05:53.282495446Z","interrupted":"2026-10-05T14:05:53.282495446Z"} +### controller log around the cut (previous + new process) +2026/10/05 14:05:53 backup.go:557: [INFO] [backup] Starting database dump run +2026/10/05 14:05:53 dbdump.go:208: [INFO] [backup] Discovered 2 databases +2026/10/05 14:05:53 backup.go:596: [INFO] [backup] Discovered 2 database(s): paperless-postgres(postgres), bookstack-db(mariadb) +2026/10/05 14:05:53 dbdump.go:410: [INFO] [backup] DB dump: paperless-postgres → paperless-ngx-postgres.sql (428.1 KB, 321ms, 72 tables) +2026/10/05 14:05:54 dbdump.go:410: [INFO] [backup] DB dump: bookstack-db → bookstack-mariadb.sql (55.9 KB, 310ms, 41 tables) +2026/10/05 14:05:54 backup.go:974: [INFO] [backup] Stopping bookstack for safe volume dump +2026/10/05 14:05:58 backup.go:873: [INFO] [backup] Volume dump: bookstack/bookstack_bookstack_config → 49.0 KB +2026/10/05 14:05:59 main.go:340: [INFO] felhom-controller 0.296.0 starting (customer: demo-hp, domain: enkisfelhom.hu) +2026/10/05 14:05:59 appstop_marker.go:278: [WARN] [appstop] crash recovery: an app-data backup (volume dump) (op "volume-dump:bookstack") was interrupted and left 1 app(s) stopped — restarting them: [bookstack] +2026/10/05 14:05:59 manager.go:1235: [INFO] [stacks] Starting stack: bookstack +2026/10/05 14:06:06 appstop_marker.go:304: [INFO] [appstop] crash recovery: restarted bookstack after the interrupted an app-data backup (volume dump) +2026/10/05 14:06:06 sync.go:117: [INFO] [sync] Starting catalog sync (repo: https://gitea.dooplex.hu/admin/app-catalog-felhom.eu.git, interval: 15m0s) +2026/10/05 14:06:06 sync.go:208: [INFO] [sync] Starting catalog sync +2026/10/05 14:06:06 restore_record_wiring.go:35: [INFO] [backup] restore record wired: /opt/docker/felhom-controller/data/restore-status.json (interrupted at startup: false) +2026/10/05 14:06:06 run_record_wiring.go:27: [WARN] [backup] the app-data backup run started 2026-10-05T14:05:53Z was cut off by the stop — the backup pages say so until the next complete run (R-519) +2026/10/05 14:06:06 scheduler.go:223: [INFO] [scheduler] Starting scheduler with 19 jobs +2026/10/05 14:06:06 backup.go:1179: [INFO] [backup] Found 3 DB dump files across drives +2026/10/05 14:06:06 dbdump.go:208: [INFO] [backup] Discovered 2 databases +2026/10/05 14:06:06 [INFO] [backup] Discovered app data: 3 apps +2026/10/05 14:06:06 recovery_unit.go:289: [INFO] [backup] Recovery unit captured for bookstack → /mnt/sys_drive/felhom-data/backups/primary/bookstack (images=2, secrets-referenced=3, data_keys=0, portable-carried=2/2, withheld=1) +2026/10/05 14:06:06 recovery_unit.go:289: [INFO] [backup] Recovery unit captured for paperless-ngx → /mnt/felhom-drives/scratch_hdd/userdata/paperless-ngx/backups/primary/paperless-ngx (images=3, secrets-referenced=3, data_keys=0, portable-carried=2/2, withheld=1) +2026/10/05 14:06:06 backup.go:1250: [INFO] [backup] Backup status cache refreshed +2026/10/05 14:06:09 manager.go:1601: [INFO] [stacks] bookstack lscr.io/linuxserver/bookstack:26.09.1@sha256:99cd1f5707c1911afad213adec5c9739763b76f843d1477142231834ecdcb6f7 running Up 3 seconds (health: starting) diff --git a/documentation/audits/hub-db-offsite-2026-10-05/partD/r519/05-after-cut.txt b/documentation/audits/hub-db-offsite-2026-10-05/partD/r519/05-after-cut.txt new file mode 100644 index 00000000..61039764 --- /dev/null +++ b/documentation/audits/hub-db-offsite-2026-10-05/partD/r519/05-after-cut.txt @@ -0,0 +1,23 @@ +## after the cut, 2026-10-05T14:06:23Z +### GET /backups +data-interrupted-run present: 1 +
A legutóbbi mentés (2026-10-05 16:05) megszakadt, mert a doboz vagy a vezérlő újraindult. Amit nem fejezett be, annak a korábbi mentése maradt meg — minden visszaállítási pont annyira friss, amennyire a legrégebbi része. A következő teljes mentés után ez az üzenet eltűnik. +HTTP 200 +### GET /backups/apps +data-interrupted-run present: 1 +
A legutóbbi mentés (2026-10-05 16:05) megszakadt, mert a doboz vagy a vezérlő újraindult. Amit nem fejezett be, annak a korábbi mentése maradt meg — minden visszaállítási pont annyira friss, amennyire a legrégebbi része. A következő teljes mentés után ez az üzenet eltűnik. +HTTP 200 +### negative control: a marker that must NOT be on the page +0 +### restore points (bookstack) +{"ok":true,"data":[{"time":"2026-10-05T14:04:44Z","short_id":"helyi","tier":1,"drive_label":"Belső SSD (rendszer)"}]} +### dump file times +total 314252 +-rw-r--r-- 1 root root 50176 2026-10-05T14:05:58 bookstack_bookstack_config.tar +-rw-r--r-- 1 root root 160868864 2026-10-05T14:04:44 bookstack_bookstack_db_data.tar +-rw-r--r-- 1 root root 160867328 2026-10-05T14:05:58 bookstack_bookstack_db_data.tar.tmp +drwxr-xr-x 2 root root 4096 2026-10-05T14:05:54 db-dumps +drwxr-xr-x 2 root root 4096 2026-10-05T14:05:58 volume-dumps +### bookstack after +bookstack Up 27 seconds (healthy) +bookstack-db Up 32 seconds (healthy) diff --git a/documentation/audits/hub-db-offsite-2026-10-05/partD/r519/06-run3.txt b/documentation/audits/hub-db-offsite-2026-10-05/partD/r519/06-run3.txt new file mode 100644 index 00000000..b0ac4d4b --- /dev/null +++ b/documentation/audits/hub-db-offsite-2026-10-05/partD/r519/06-run3.txt @@ -0,0 +1,12 @@ +## run 3 (complete) start 2026-10-05T14:06:44Z +HTTP 200 +## run 3 end 2026-10-05T14:07:28Z: {"ok":true,"data":{"db_dump":{"count":2,"duration":"33.315240638s","last_run":"2026-10-05T14:07:20.615183031Z","success":true},"enabled":true,"running":false}} +2026/10/05 14:05:12 backup.go:684: [INFO] [backup] App-data backup completed: 2 databases (483.2 KB total), 3 volume dump(s) (32.785s) +2026/10/05 14:07:20 backup.go:684: [INFO] [backup] App-data backup completed: 2 databases (483.9 KB total), 3 volume dump(s) (33.315s) +{"running":false,"started_at":"0001-01-01T00:00:00Z","interrupted":"0001-01-01T00:00:00Z"} +total 157152 +-rw-r--r-- 1 root root 50688 2026-10-05T14:06:52 bookstack_bookstack_config.tar +-rw-r--r-- 1 root root 160867328 2026-10-05T14:06:53 bookstack_bookstack_db_data.tar +/backups data-interrupted-run: 0 +/backups/apps data-interrupted-run: 0 +restore points: {"ok":true,"data":[{"time":"2026-10-05T14:06:47Z","short_id":"helyi","tier":1,"drive_label":"Belső SSD (rendszer)"}]} diff --git a/documentation/audits/hub-db-offsite-2026-10-05/partD/r519/07-teardown.txt b/documentation/audits/hub-db-offsite-2026-10-05/partD/r519/07-teardown.txt new file mode 100644 index 00000000..9831a352 --- /dev/null +++ b/documentation/audits/hub-db-offsite-2026-10-05/partD/r519/07-teardown.txt @@ -0,0 +1,8 @@ +## teardown 2026-10-05T14:07:48Z +stop: HTTP 200 +remove: HTTP 200 +containers: 0 +volumes: 0 +stackdir: none +backups: none +/root/.dbody diff --git a/documentation/audits/hub-db-offsite-2026-10-05/partD/restore-procedure/drill.txt b/documentation/audits/hub-db-offsite-2026-10-05/partD/restore-procedure/drill.txt new file mode 100644 index 00000000..16388e33 --- /dev/null +++ b/documentation/audits/hub-db-offsite-2026-10-05/partD/restore-procedure/drill.txt @@ -0,0 +1,9 @@ +## runbook §3 drill 2026-10-05T14:09:41Z (scratch /var/lib/felhom-hub-backup/sec3.65qk, root 0700) +step 1 restore host/dooplex-hub/2026-10-05T13:50:18Z +restore complete (352.676 MiB processed in 3.9s, average 90.961 MiB/s) +-rw------- 369807360 hub.db +step 2 the seal key from Secret/offsite-secret-key → /var/lib/felhom-hub-backup/sec3.65qk/k (0600, not printed) +key file bytes: 64 +step 3 hubdb-check (same key): hosts=4 console_passwords_opened=4 failed=0 absent=0 rc=0 +control (a random key): hosts=4 console_passwords_opened=0 failed=4 absent=0 hubdb-check: FAILED: not every console password opened with this key rc=0 +scratch shredded: gone diff --git a/documentation/audits/hub-db-offsite-2026-10-05/partD/restore-procedure/red-proof.txt b/documentation/audits/hub-db-offsite-2026-10-05/partD/restore-procedure/red-proof.txt new file mode 100644 index 00000000..3d46b962 --- /dev/null +++ b/documentation/audits/hub-db-offsite-2026-10-05/partD/restore-procedure/red-proof.txt @@ -0,0 +1,7 @@ +### R7 hubdb-check counts a failed open as opened +=== RUN TestCheck_WrongKeyOpensNothing + main_test.go:63: got {hosts:2 opened:1 failed:0 absent:1}, want opened=0 failed=1 +--- FAIL: TestCheck_WrongKeyOpensNothing (0.03s) +FAIL +FAIL gitea.dooplex.hu/admin/felhom-hub/cmd/hubdb-check 0.040s +FAIL diff --git a/documentation/audits/hub-db-offsite-2026-10-05/partD/restore-test-1.txt b/documentation/audits/hub-db-offsite-2026-10-05/partD/restore-test-1.txt new file mode 100644 index 00000000..38fa4d98 --- /dev/null +++ b/documentation/audits/hub-db-offsite-2026-10-05/partD/restore-test-1.txt @@ -0,0 +1,17 @@ +## snapshot list on ep0 with the READ-ONLY token, 2026-10-05T13:50:40Z ++=======================================+=============+=====================================+ +| snapshot | size | files | ++=======================================+=============+=====================================+ +| host/dooplex-hub/2026-10-05T13:50:18Z | 352.677 MiB | catalog.pcat1 hubdb.pxar index.json | ++=======================================+=============+=====================================+ +unit rc=0 +Result=success +2026-10-05T15:50:41+02:00 dooplex felhom-hub-db-restore-test[3040582]: felhom-hub-db-restore-test: restoring host/dooplex-hub/2026-10-05T13:50:18Z +2026-10-05T15:50:46+02:00 dooplex felhom-hub-db-restore-test[3040582]: felhom-hub-db-restore-test: checked: integrity ok, 4 host(s), 4 sealed console password(s), 0 readable +2026-10-05T15:50:46+02:00 dooplex felhom-hub-db-restore-test[3040582]: felhom-hub-db-restore-test: success signal written +# HELP felhom_hub_db_restore_test_last_success_timestamp_seconds Last successful restore test of the hub DB copy on ep0 (R-173). +# TYPE felhom_hub_db_restore_test_last_success_timestamp_seconds gauge +felhom_hub_db_restore_test_last_success_timestamp_seconds 1791208246 +.cache +.kube +stage diff --git a/documentation/audits/hub-db-offsite-2026-10-05/partD/token-limits-and-paperkey.txt b/documentation/audits/hub-db-offsite-2026-10-05/partD/token-limits-and-paperkey.txt new file mode 100644 index 00000000..40a5a0e7 --- /dev/null +++ b/documentation/audits/hub-db-offsite-2026-10-05/partD/token-limits-and-paperkey.txt @@ -0,0 +1,12 @@ +## 2026-10-05T13:51:07Z token limits (each line = the tool output) +push forget : Error: permission check failed - missing Datastore.Modify|Datastore.Prune on /datastore/felhom-offsite/operator +restore forget : Error: permission check failed - missing Datastore.Modify|Datastore.Prune on /datastore/felhom-offsite/operator +restore backup : Error: missing permissions 'Datastore.Backup' on '/datastore/felhom-offsite/operator' +push list ns root (households): Error: permission check failed - missing Datastore.Audit|Datastore.Backup on /datastore/felhom-offsite +restore list ns root (households): Error: permission check failed - missing Datastore.Audit|Datastore.Backup on /datastore/felhom-offsite +push restore own copy: Error: missing key - manifest was created with key b2:19:bf:36:3b:97:3d:6c exit-file=absent +## paper-key proof: restore with a key file rebuilt from the data field only +restore rc=0 +integrity: ok hosts: 4 +## and with NO key: Error: missing key - manifest was created with key b2:19:bf:36:3b:97:3d:6c +push restore own copy WITH key: restore complete (352.676 MiB processed in 4.6s, average 76.793 MiB/s) file=PRESENT diff --git a/documentation/backlog/CLOSED-ITEMS.md b/documentation/backlog/CLOSED-ITEMS.md index 19fe9dd1..8314db46 100644 --- a/documentation/backlog/CLOSED-ITEMS.md +++ b/documentation/backlog/CLOSED-ITEMS.md @@ -26,6 +26,16 @@ --- +## 2026-10-05 (evening) — the hub database off DooPlex (hub v0.136.0; `09` rulings 125–127) + +The full text of every row below: `git show cea8502f:documentation/backlog/OPEN-ITEMS.md`. + +| Row | What | Closed | Evidence | +|---|---|---|---| +| **R-519** | **After a backup torn by a power cut, a restore point carried the new database dump's time over the previous run's files, and no screen said the run was interrupted (P2).** Fixed in controller v0.296.0 (page notice + synthesised status; dating by the oldest part since v0.275.0). **Live on 9202 (operator ruling 126):** a complete run, then a run cut by `docker restart felhom-controller` 0.8 s after bookstack's config dump and before its database volume; afterwards both /backups and /backups/apps carry `data-interrupted-run` („A legutóbbi mentés (2026-10-05 16:05) megszakadt …", negative control 0), the restore point reads 14:04:44Z = the run-1 database volume (its oldest part; config 14:05:58, SQL 14:05:54), the controller restarted bookstack itself; the next complete run cleared the notice and replaced the torn `.tar.tmp`. The throwaway bookstack was removed through the product (0 containers, volumes, folders, backups). | CLOSED 2026-10-05 — FIXED controller v0.296.0, proven live | `audits/hub-db-offsite-2026-10-05/partD/r519/`; `internal/backup/run_record_test.go` | + +--- + ## 2026-10-05 (late afternoon) — the hub's own safety, boxes left behind, the agent's root grants (hub v0.135.0, controller v0.296.0, agent v0.146.1, golden 0.296.0; CC decisions 119–124) The full text of every row below: `git show 9bb45eaa:documentation/backlog/OPEN-ITEMS.md` (R-880 was opened and closed in this session). diff --git a/documentation/backlog/OPEN-ITEMS.md b/documentation/backlog/OPEN-ITEMS.md index 36935088..a71e3e4e 100644 --- a/documentation/backlog/OPEN-ITEMS.md +++ b/documentation/backlog/OPEN-ITEMS.md @@ -190,17 +190,16 @@ stopping line that lies. | **R-687** | App updates | P4 | **[P3-LOW] Part 7's live proof has four gaps a scratch box cannot close, and one observability gap.** (1) **W+5h reached with steps left** is proven by unit test only (`TestLeg_NoStepAtOrAfterW5h`) — the leg starts at W+105m and would need a 3-hour leg live; (2) **the off-site leg FAILING** before the update leg: 9202 has no off-site tier, so only the "no target" path ran live — failure and panic are `TestChainUpdateLeg_EveryPath`; (3) **a `files_may_change` step WITHOUT a whole copy**: both drill apps given the mark (wishlist, navidrome, romm) turned out whole on 9202 by the backup side's truth table (why, per app, is not logged — see the observability gap), so only "with a copy" ran live; (4) **the full-system gate waiting** cannot run on 9202 (no agent), and **did not occur on the demo boxes' real night either** (2026-09-25: both legs ended by 04:19, before the gate opened at 04:30, and no whole-box backup was due on either box) — unit + red-proof only (`TestD20_GateWaitsForTheLeg`). **Also cosmetic:** a leg with no steps reports `"steps": null` to the hub, not `[]`. **Observability:** when the leg TAKES a `files_may_change` step it does not log which whole copy allowed it (only the skip says why). `audits/night-2026-09-25/C/` **-- NARROWED 2026-09-25 (controller v0.273.0):** the cosmetic `"steps": null` → `[]` and the taken `files_may_change` step's missing log line are FIXED (red-proofed, `audits/night-2026-09-26/F/`). Items (1)–(4) stay; (4) did not occur on 2026-09-25 either (demo-felhom's whole-box backup ran at 07:29, three hours after its leg; demo-hp had none due). **-- 2026-09-28 (night 27/28):** (4) did not occur again — on demo-hp the leg ended 04:23:54 and the whole-guest backup began 04:37:06, after the gate opened at 04:30; demo-felhom's backup ran at 07:36 (`audits/evidence-golden-0276-2026-09-28/phaseD2-night-read.txt`). **-- 2026-09-30 (by day, demo-hp 9201): item (4) PROVEN LIVE.** The night chain pressed by hand, the window moved to W = now − 2h05m the moment the leg started, `quiesce.poll_interval` 1m: `[quiesce] full-system backup due and inside its window, but the automatic update leg is running … deferring` at 11:35:11 and 11:36:11 UTC while bookstack (55.1 s) and kimai (75.1 s) stepped; the leg's end line at 11:36:29; the backup quiesced at 11:37:11 (the first poll after), job done 11:47:19, the agent's `backup: completed` 9.98 GB. Config and window put back and read back (`audits/pg-last-six-2026-09-30/C/`). **Found, cosmetic, manual chain only:** the deferral names the moved window's W+5h (16:29) while the manual leg's own deadline was its start + the leg length (16:49). | **OPEN — P3, gaps (1)–(3) + the manual-chain deferral text; item (4) proven live 2026-09-30; owner: CC** **Re-ranked 2026-10-03: P3→P4: the gaps are covered by unit tests; left is live-proof completeness and one log text.** | — | — | CC | | **R-734** | App updates | P4 | **[P3-LOW] The harness marks immich `files_may_change` because immich rewrites six 13-byte `.immich` folder markers at every start.** MEASURED 2026-09-30 on the bench (v3.2.2 → v3.2.4): the bind-tree hash of `appdata/immich` changed; the only changed files were `{encoded-video,library,backups,profile,thumbs,upload}/.immich`, rewritten at each start — no household file. The mark is honest by the harness's rule and the ladder writer copies it (never edited by hand), so immich's v3.2.4 night step needs a fresh WHOLE copy (decision 13); on a box without one the night leg skips it and a person presses. The 2026-09-23 immich entry did not carry it (`files_changed []`). **Needs:** a decision whether app-owned marker files are excluded from the file hash (a per-template ignore list, or a size/name rule), or the mark stays. **-- 2026-09-30 (evening):** the harness now NAMES the files behind the mark (`files_changed_detail`, catalog `5b1972b`); on immich's step `0b82…` re-proof it named exactly the six `.immich` markers again. calibre-web's step (v4.0.6 → v4.0.8) carries the mark too, and the named files are its LIBRARY DATABASE: `media/books/metadata.db`, `metadata.db-shm`, `metadata.db-wal` changed; the book file did not (bench names re-run, `A/calibre-names/`). That is household data (the template's backup class `mandatory` holds DB and books as one unit), so the mark is right there and the night leg takes that step only with a fresh whole copy. On 9202 the same step changed no file in that folder (per-file hashes before/after) — not explained. `audits/more-night-apps-2026-09-30/` | **READY — rank P3-LOW; owner: CC (harness); the rule change needs a word** **Re-ranked 2026-10-03: P3→P4: the effect is an update that waits for a press; no data risk.** | — | — | CC + operator | -## Backup & restore — 54 rows (P2 9, P3 23, P4 22) +## Backup & restore — 53 rows (P2 8, P3 23, P4 22) | ID | Category | Sev | What | State | Blocked on | Next action | Owner | |---|---|---|---|---|---|---|---| | **R-32** | Backup & restore | P2 | **[P2-HIGH] RESET must purge the customer base dir; the orphan card must stay honest; unattributed bytes must be visible.** The rehearsal's S7 said in advance that an orphan card would BE a finding — and one appeared (16:58:14). Cause: RESET's `"hetzner":"ok"` leg destroys the sub-account, but **a Hetzner sub-account is an access-control object, not a data object** — its directory survives, so re-enabling offsite recreated an account over the previous lifecycle's ciphertext, encrypted under a key that same RESET had destroyed. **MIGRATED FROM `ROADMAP.md` 2026-08-22 (R-369) — originally filed 2026-07-21, size M, roadmap state `idea`.** Moved verbatim; nothing added or reinterpreted. The roadmap keeps its copy as history, marked moved. | **OPEN — migrated from ROADMAP 2026-08-22, rank unchanged** | — | **Ruling from the run (three parts, deliberately separate):** (1) because RESET destroys custody, the ciphertext it leaves behind is unrecoverable **BY DESIGN** → RESET gains a **main-account purge of the customer base dir** (the existing operator ack already covers it); (2) the **move-aside guard STAYS** for reinstall-*without*-RESET — there custody survives and the card's "history recoverable" promise is true (R-26 depends on exactly that); (3) the operator **Restic tab shows per-customer directory bytes vs attributed snapshot bytes**, so dead data cannot hide. Measured on the pool box that night: **49 M attributed** (2 snapshots, 48.717 MiB) against **1.4 G + 3.0 M unattributed** across TWO `.orphaned-*` dirs. Evidence `restic-and-pool.txt` | CC | | **R-105** | Backup & restore | P2 | **Three hub-held DR records are empty on the entire live fleet.** `hosts.dr_record_json` = `{}` on all 3 hosts; `host_escrow.directive_json` = `{}` on both escrowed hosts; `dr_recipe.host_half.drives` = `[]` on every customer **including two with enrolled data drives** (916 GB USB on demo-felhom, 938 GB NVMe on demo-hp) **MIGRATED FROM `ROADMAP.md` 2026-08-22 (R-369) — originally filed 2026-07-28, size M, roadmap state `READY — 2026-07-28`.** Moved verbatim; nothing added or reinterpreted. The roadmap keeps its copy as history, marked moved. ****PARTLY FIXED BY ITS OWN UPDATE.** The `drives` third was traced and populated on both demo boxes on 2026-07-28 (the enrolled drives were never PVE storages, so `isUserDataDrive` never saw them). The other two thirds — `hosts.dr_record_json` and `host_escrow.directive_json` — were NOT re-verified this session and are carried as written.** | **OPEN — migrated from ROADMAP 2026-08-22, rank unchanged** | — | These are exactly the fields a host-loss recovery reads: `05-hub-architecture.md:175-176,186` names the slim DR record as one of four durable sources; `06-offsite-connectivity.md:148-150` says the escrow upload carried the DR directive; `felhom-agent/internal/dr/plan.go:34-35` makes `PlannedDrive` the re-attach-by-`durable_id` wrong-disk guard. **The three may have different causes** — `isUserDataDrive` (`internal/hub/dr_recipe.go:129-136`) requires type `usb`/`local-dir` **and** a non-empty `DurableID` **and** `MountPath`, and which of the three fails was not traced. Evidence: `architecture/_recovery-inventory-2026-07-28.md` Part D2.3. **UPDATE 2026-07-28 (vzdump-target move): the `drives` third is TRACED and now POPULATED on both demo boxes.** Cause: the enrolled data drives were never PVE storages at all — only agent-generated systemd mounts — so they never entered `report.StorageTargets` and `isUserDataDrive` never saw them. Giving each drive a `dir` storage at its own mountpoint supplied all three required fields at once (type `local-dir`, fs-UUID durable id, mount path), and the recipe now emits `uuid:91d2dc2d-…`/`/mnt/nvme-1tb` on demo-hp and `uuid:47a3361a-…`/`/mnt/hdd_1` o | CC | -| **R-232** | Backup & restore | P2 | **DooPlex's backup makes every copy inside the same box — and nothing tells anyone when it fails.** Surveyed read-only 2026-08-06 (`audits/RECON-dooplex-backup-2026-08-06.md`). **What works:** five sets, 14/14 successful runs in 14 days; a file was restored from the `data` repo and matched the live original **byte for byte**; every set except two is cross-disk; k3s is integrity-checked on every run. **What the matrix exposes, ranked:** (a) **`notify_failure` is a no-op** — `NOTIFY_ON_FAILURE=true` but `NOTIFY_WEBHOOK_URL` is commented out, so a failed backup notifies **nobody**; the project already has a working Resend path that CI uses. Cheapest item, and it makes every other failure visible. (b) **Nothing leaves the box** — no rclone, no remote repo, no off-site target anywhere; Longhorn's target is `nfs://192.168.0.180:` pointing at DooPlex itself, and the only outbound-looking cron pulls *inbound* from Hetzner for a different project. The machine that runs the hub managing the customers' off-site chain has no off-site copy of its own. (c) **The backup tree is a single writable path** and the restic repos are not append-only — one bad script or ransomware destroys every copy at once. (d) **Two same-disk sets**: `.claude-memory` and the PostgreSQL dumps, whose source directory sits *inside* the backup tree. (e) **Longhorn `retain=1`** — one generation per volume, so a corruption noticed a day late has no earlier copy. (f) **`/opt/backup/docs/BACKUP-RESTORE.md` does not exist** though the systemd unit advertises it. (g) **`secrets/restic-repo` has never held a snapshot** — `backup-secrets.sh` contains no `restic` call; the secrets are GPG files on `sda1` only. (h) **No restore has ever been run** beyond today's single-file probe — the matrix's "ever demonstrated?" column is otherwise entirely empty. **Not a finding:** the restic passphrase. The on-box copy is on `sdb1`, a different disk from the backups, and the **operator holds an offline copy out of band** — so a disk loss is recoverable. The narrow residual is that it is operator-held rather than system-held, unlike the customer case's hub-vaulted escrow, so it should be confirmed current and findable by someone else. **Nothing was changed by the recon.** | **READY** — owner Viktor | — | — | operator | +| **R-232** | Backup & restore | P2 | **DooPlex's backup makes every copy inside the same box — and nothing tells anyone when it fails.** Surveyed read-only 2026-08-06 (`audits/RECON-dooplex-backup-2026-08-06.md`). **What works:** five sets, 14/14 successful runs in 14 days; a file was restored from the `data` repo and matched the live original **byte for byte**; every set except two is cross-disk; k3s is integrity-checked on every run. **What the matrix exposes, ranked:** (a) **`notify_failure` is a no-op** — `NOTIFY_ON_FAILURE=true` but `NOTIFY_WEBHOOK_URL` is commented out, so a failed backup notifies **nobody**; the project already has a working Resend path that CI uses. Cheapest item, and it makes every other failure visible. (b) **Nothing leaves the box** — no rclone, no remote repo, no off-site target anywhere; Longhorn's target is `nfs://192.168.0.180:` pointing at DooPlex itself, and the only outbound-looking cron pulls *inbound* from Hetzner for a different project. The machine that runs the hub managing the customers' off-site chain has no off-site copy of its own. (c) **The backup tree is a single writable path** and the restic repos are not append-only — one bad script or ransomware destroys every copy at once. (d) **Two same-disk sets**: `.claude-memory` and the PostgreSQL dumps, whose source directory sits *inside* the backup tree. (e) **Longhorn `retain=1`** — one generation per volume, so a corruption noticed a day late has no earlier copy. (f) **`/opt/backup/docs/BACKUP-RESTORE.md` does not exist** though the systemd unit advertises it. (g) **`secrets/restic-repo` has never held a snapshot** — `backup-secrets.sh` contains no `restic` call; the secrets are GPG files on `sda1` only. (h) **No restore has ever been run** beyond today's single-file probe — the matrix's "ever demonstrated?" column is otherwise entirely empty. **Not a finding:** the restic passphrase. The on-box copy is on `sdb1`, a different disk from the backups, and the **operator holds an offline copy out of band** — so a disk loss is recoverable. The narrow residual is that it is operator-held rather than system-held, unlike the customer case's hub-vaulted escrow, so it should be confirmed current and findable by someone else. **Nothing was changed by the recon.** | **NARROWED 2026-10-05 — owner Viktor.** (b) partly: the hub database now leaves DooPlex nightly, encrypted, to ep0 (R-173); everything else in DooPlex's backup still stays on the box. (a) partly: the hub copy alarms through Prometheus (`HubDBBackupStale`); `notify_failure` is still a no-op for the rest. (c)–(h) unchanged. **READY** for the rest | — | — | operator | | **R-304** | Backup & restore | P2 | **The retained escrow key works, and the customer is told their correct code is wrong.** DRILL 2026-08-12 answered the three questions separately, on `demo-felhom`, with planted data. **(a) retention: WORKS** — the first retained row in fleet history to carry material (`host_escrow_superseded` id 11, `identity_blob` 572 B), byte-identical (`sha256 a10032341c8584ed…`) to the pre-supersession `host_escrow` row. **(b) the material opens the old store: YES** — unsealed with the OLD recovery code it yielded a password byte-identical to the pre-change one (`sha c60c8bc737a6b7c6…`), and restored three planted files **byte-identical** from a store the box itself could no longer open (negative control first: `Fatal: wrong password or no key found`), **including a Hungarian accented filename verified as raw bytes**. **(c) the customer's route: DOES NOT EXIST, and misinforms.** `ListSupersededEscrow` (`store.go:2841`) is the only reader of a retained `identity_blob` and has **zero production callers** — five call sites, all `_test.go`; the product path (`POST /escrow/recover-offsite-password` → `FetchIdentityEscrow` → `GetHostDRBundle`, `store.go:3152`) selects `FROM host_escrow` — the CURRENT row only. Asked for the old password with the code that demonstrably opens the retained row, the product answered **"the recovery code did not open the sealed bundle — nothing was written"**. **This is the R-224 class again**: there an unreachable hub was reported as a bad code; here a VALID code for retained history is reported as a bad code, and the customer's attempt ends there. **Consequence:** the census answer stands (it was about retention); the countdown banner's promise is true in substance and false in practice; **any capability-map claim that the customer can recover the old history with their recovery code is false today and must move** | **READY (L) — NEW 2026-08-12, RANK 1** | R-198, R-199, R-224, R-241 | Decide the shape: serve retained rows on the recovery path (needs a "which package?" choice — a customer may have several), or stop promising retrieval anywhere the customer cannot perform it. **Until one of those, the honest position is that retention is an operator-only capability.** At minimum, the refusal must stop asserting the code is wrong when the hub simply never looked | operator + CC | | **R-366** | Backup & restore | P2 | **The 21 August reinstall orphaned `demo-hp`'s PBS whole-guest archives as well as its off-site repo — the box can no longer read its own pre-reinstall backups, and this surfaces only as a restore-test failure.** Hub event 3016, 2026-08-21 21:59:28Z, unprompted: `Restore-test FAILED on the pbs tier: archive felhom-pbs:backup/ct/9201/2026-08-18T03:58:43Z could not be restored+booted … proxmox-backup-client failed: Error: wrong key - unable to verify signature since manifest's key 3f:4f:65:c0:d8:f3:9f:3c does not match provided key dd:d1:d8:53:44:62:5e:0b`. The archive predates the reinstall by three days. **This is the PBS-tier analogue of R-193** (a guest rebuild mints a fresh secret and orphans the history), and the two together mean a rebuilt box loses BOTH off-premises tiers at once: the restic repo needed a self-heal + re-toggle (see the drill report), and the PBS archives are simply unreadable to it. **Credit: the restore-test caught it and said so precisely** — the mechanism works. **The gap is what it is called:** it is reported as *a restore test that failed*, which reads as a flaky verification, not as *every whole-guest backup you took before the reinstall is unreadable on this machine*. **Found incidentally by the 2026-08-21 backup-truth drill; nobody was looking for it.** | **OPEN — HIGH** | related: R-193 | Establish whether the pre-reinstall PBS archives are recoverable at all (the old key's whereabouts), and separate the two verdicts: a tier whose ARCHIVES ARE ORPHANED is a different alarm from a tier whose restore test failed. Do not close on the strength of the restore-test wording alone. | CC | | **R-518** | Backup & restore | P2 | **[P2-MEDIUM] „Mentés most" on the whole-system backup stops every app for about eight minutes while the page promises „csak néhány másodpercre".** MEASURED 2026-09-14 (BIGNIGHT, VM 333, 12 apps): the button's call quiesced all 12 stacks at 19:03:23Z (first stopped 19:03:27Z); the local vzdump ran 19:03:49 → 19:09:59Z; the controller then kept the apps stopped for the second (PBS) tier and restarted them at 19:10:09Z after it failed, the last started 19:11:12Z (`phase4/guest-backup-quiesce-log.txt`) — **≈ 7 m 45 s** with every app answering 404. The page under the button: „Pillanatkép-mód: az alkalmazások csak néhány másodpercre állnak le." A household pressing it at dinner loses every app for the length of the dump, and longer on a bigger box. **Fix shape:** state the real expected downtime (it scales with data), or quiesce per tier and not across a second tier's attempt; do not start a tier whose storage is absent (see R-517). **NARROWED 2026-09-15 (controller v0.243.0 + agent v0.131.0):** a tier whose storage the agent reports absent is skipped before anything stops (`backup_tier_skipped`, once per absence; unknown never skipped), and the button copy now says „általában néhány perc, nagyobb adatnál több". Unit-proven with red-proofs. **Still open:** quiesce per tier, so a slow second tier does not keep every app down. **— NIGHT 2026-09-23 (controller v0.267.0):** the copy half is DONE: the page and the confirm now state the measured stop (≈ 8 minutes on a 12-app box), both languages, red-proofed (`audits/night-2026-09-23/A5-*`). The brief's „csak néhány másodpercre" had already gone in v0.243.0. **Still open:** quiesce per tier, so a slow second tier does not keep every app down. | **READY — P2, narrowed to per-tier quiesce; owner: CC (controller). 2026-10-05: the copy now states today's measurement too (demo-hp, 9 apps, local tier only: 5 min 47 s) — controller v0.296.0, `audits/hub-safety-2026-10-05/partE/`.** | — | — | CC | -| **R-519** | Backup & restore | P2 | **[P2-MEDIUM] After a backup torn by a power cut, an app's restore point carries the new database dump's time while its files are from the previous run — and no customer screen says the run was interrupted.** MEASURED 2026-09-14 (BIGNIGHT F2, VM 333): „Mentés most" 19:40:03Z; the power was cut 19:40:08Z while adventurelog was stopped for its volume dump. On disk afterwards, `backups/primary/adventurelog`: `db-dumps/adventurelog-postgres.sql` **19:40:07**, `volume-dumps/*.tar` **19:00:35**, `manifest.json created_at 19:02:34Z`; bookstack the same shape (sql 19:40:07, tars 19:00:48). `GET /api/backup/snapshots` for both → `time 2026-09-14T19:40:07Z, helyi` (`phase5/F2/units-on-disk.txt`, `backup-honesty.txt`). `/backups/apps` shows „Utolsó adatbázis mentés 2026-09-14 21:40 · … OK" and every app „Utolsó: 5 perce"; `/backups` and `/dashboard` contain no word of an interruption (fragments `megszakad|sikertelen|nem sikerült` = 0; control „hiba" appears in the standing warning text). The controller itself knew: `[appstop] crash recovery: an app-data backup (volume dump) … was interrupted … restarting them: [adventurelog]` and pushed `backup_failed (error)` to the hub. A household restoring „the 21:40 backup" gets 21:00 files for BookStack's uploads. **Fix shape:** date a point by the oldest part it contains (or mark it partial) and show the interrupted run on the backups page until the next complete one. **F6 (drive unplugged 1 s into a backup, 20:37:59Z) adds three facts:** the run skipped four apps' volume dumps („Skipping volume dump for immich — drive disconnected", also jellyfin, nextcloud, paperless-ngx) and still reported `db_dump {"count":6, "success":true}`; nextcloud's point is dated 20:37:59Z (its SQL finished before the unplug) beside 19:01 volume tars; and a torn `immich-postgres.sql.tmp` (20:38:00) plus F3's `pre-restore-…-nextcloud-mariadb.sql.tmp` are left in the units on the drive. Immich's point correctly stayed at 19:02:34Z (the `.tmp` was not promoted). `/backups/apps` fragments `kihagy|sikertelen|részleges` = 0 (`phase5/F6/`). | **NARROWED 2026-10-05 — FIXED controller v0.296.0, unit-proven (3 red-proofs): a run cut off by a stop is said on /backups and /backups/apps until a run ends with every step OK; the synthesised "last database backup" no longer reads OK over a cut run; the restore point's time was already its oldest part (v0.275.0). LEFT: the live cut on 9202 — the permission check refused restarting the controller mid-backup; the operator is asked (STATUS).** | — | the operator's go for one controller restart mid-backup on 9202 | CC | | **R-638** | Backup & restore | P2 | **[P2-MEDIUM] The product's database loader cannot replay a copy over a NEWER schema: on PostgreSQL it FAILS, on MariaDB it leaves the newer version's tables behind.** MEASURED 2026-09-23 on 9202. `ImportDump` (`appbackup/dbdump.go:719`, `psql -v ON_ERROR_STOP=1 --single-transaction`) replays a `pg_dump --clean --if-exists` file over the live database. After docmost 0.95.0 → 0.96.0 migrated, the replay of the pre-update copy was refused in 0.40 s, rc 3: *cannot drop constraint workspaces_pkey on table public.workspaces because other objects depend on it / DETAIL: constraint oauth_clients_workspace_id_fkey …* — the new version created six tables whose foreign keys point at old ones, and `--clean` only drops what the dump knows. Database unchanged (the transaction rolled back). On MariaDB (`mariadb-dump`, `FOREIGN_KEY_CHECKS=0`) the same replay after romm 5.0.0 → 5.3.0 returned rc 0 in 1.25 s and left **12 base tables** of the new version behind; RomM 5.0.0 happened to ignore them. **What worked:** `DROP SCHEMA public CASCADE; CREATE SCHEMA public;` + the dump in ONE transaction — rc 0 in 1.38 s, every table, index and extension back. **Why this is a row of its own and not only part of R-637:** the SAME loader backs shipped paths — `rollbackSafetyDump` (off-site restore's undo) and the dump replay of the restores — so **any restore of a copy taken BEFORE an update that migrated, replayed over the migrated database, may fail the same way. NOT MEASURED:** whether the unit restore the hold sentence names does this (it also carries the data VOLUME tar, which may make the replay moot). That is the measurement owed, on 9202, before anyone relies on it. Evidence: `audits/update-rulings-2026-09-23/README.md` Part 1, `docmost-45`, `romm-44`. **-- NARROWED 2026-09-23:** the undo no longer touches this loader — it copies folders (decision 19, controller v0.263.0). **What stays open is the part about SHIPPED paths:** `rollbackSafetyDump` and the restores' dump replay still replay over whatever schema is live, and whether the unit restore the hold sentence names works after a real schema migration is STILL UNMEASURED. | **OPEN — P2, narrowed to the restore paths; owner: CC; measure the named restore after a real schema migration first** | — | — | CC | | **R-822** | Backup & restore | P2 | **An add-only key does not make retention safe: an attacker who can only ADD snapshots can make the honest pruner erase every real one.** MEASURED 2026-10-03 (lab, restic 0.14.0 from the controller image, rclone `--append-only`): 13 empty snapshots dated in the future with the same host and tag, added through the add-only key (all allowed), make the box's exact policy `forget --group-by host,tags --keep-daily 7 --keep-weekly 4 --keep-monthly 6` keep only the fakes and select **all 3 real snapshots** for removal (`--dry-run`). **Whoever prunes an append-only repo — the box in a window, or a Felhom-side worker — inherits this.** Not a defect today (today the box can simply delete, R-95); a PRECONDITION on the R-95 build, like R-430 was. `audits/offsite-append-only-2026-10-03/lab/C3-retention-poisoning.txt` | **NARROWED 2026-10-03 — the guard ships in controller v0.289.0 (future-dated / newer-than-hub / recent-removal refusals, oldest-first cap, the lab's 13-fake shape refused in a test). RESIDUAL, not closable by a guard: an add-only attacker can plant PAST-dated snapshots interleaved with real ones and so steer weekly/monthly keeps; bounded per window by `MaxRemove` and the hub's count check, not prevented.** | — | Before any `forget`: refuse when a snapshot is dated in the future or newer than the newest the hub has seen reported; dry-run first and abort above the count the policy can remove in a week; hub compares the count before/after (DESIGN.md §3) | CC | | **R-49** | Backup & restore | P3 | **[P2] The offsite capture set is ~90% cache and duplication — 1.1 GB of a 1.2 GB immich "photo backup".** Measured 2026-07-19: `immich_ml_cache.tar` **823 660 032 B (~60%)** — re-downloadable ML model weights; `immich_postgres_data.tar` **308 251 136 B (~23%)** — a raw tar of the postgres data dir that DUPLICATES the logical `.sql` dump captured beside it; `upload/backups/` **18 MB** — immich's own nightly dump, a backup inside the backup, growing daily; plus the stranded pre-v3 `dccc13fe…` tree (~36 MB) no DB has ever referenced. Actual irreplaceable content: **72 MB of originals**. **MIGRATED FROM `ROADMAP.md` 2026-08-22 (R-369) — originally filed 2026-07-19, size S–M, roadmap state `idea`.** Moved verbatim; nothing added or reinterpreted. The roadmap keeps its copy as history, marked moved. | **OPEN — migrated from ROADMAP 2026-08-22, rank unchanged** **Re-ranked 2026-10-03: P2→P3: wasted space and transfer, no data risk; needs a capture-set ruling, not a sale blocker.** | — | **Evidence: `audits/DIAG-immich-restore-round2-2026-07-19.md` §4 (full byte breakdown).** This is the customer's offsite quota and transfer cost, and it lands on the Hetzner sub-account they are billed for. **Recorded, deliberately not changed** — a capture-set exclusion is a data-loss-shaped decision and gets its own ruling, not a drive-by edit. Candidates in priority order: (a) `immich_ml_cache` — pure cache, strongest case; (b) the `postgres_data` volume tar where a logical dump of the same DB is already captured (the dump is what the restore path actually replays); (c) `upload/backups/`. Likely generalises past immich into a template-classification rule about cache volumes and self-backup directories, so it should be specified against the catalog, not one app | CC | @@ -326,7 +325,7 @@ stopping line that lies. | **R-853** | Box system & updates | P3 | **After a boot the box's versions and crash facts reach the hub up to ~15 minutes late.** MEASURED 2026-10-04 on demo-hp (crash-guard test): the agent's first report after a boot has no `system.facts` — the facts read needs a RUNNING customer guest (`firstGuest`), the guest starts ~1–2 min after the agent, and the failed read is cached for 10 minutes; so the HOST half (the crash guard, the kernel) is lost too. The crash events arrived 15 min after the boot (17:17 → 17:32 CEST); nothing was lost (the guard keeps 7 days). Fix direction: the facts mode reads the host without a guest (guest fields `unknown`), and a failed read is not cached. `audits/os-docker-crash-2026-10-04/partC/c4-hub-events.txt` | **READY — owner: CC** | — | — | CC | | **R-839** | Box system & updates | P3 | **After a Docker restart that stopped every container, the boot sweep HELD an app whose `app.yaml` `HDD_PATH` names its user folder instead of the drive.** MEASURED 2026-10-04 on scratch 9202 (paperless-ngx): `bootrecon` logged `drive /mnt/felhom-drives/scratch_hdd/userdata/paperless-ngx is not a live mountpoint — NOT starting it`, although the drive `/mnt/felhom-drives/scratch_hdd` IS a mountpoint; the app stayed down until started by hand. The customer boxes' 9201s showed no hold. **Not diagnosed:** which writer put a per-app path in `HDD_PATH` on 9202, and whether a household box can get it. `audits/os-updates-spike-2026-10-04/partG/SUMMARY.md` | **READY — diagnose; owner: CC** | — | — | CC | -## Monitoring & notifications — 26 rows (P2 3, P3 16, P4 7) +## Monitoring & notifications — 27 rows (P2 3, P3 16, P4 8) | ID | Category | Sev | What | State | Blocked on | Next action | Owner | |---|---|---|---|---|---|---|---| @@ -356,12 +355,13 @@ stopping line that lies. | **R-371** | Monitoring & notifications | P4 | **The off-site tier is the only backup tier that announces nothing on success.** Written down 2026-08-05 in `audits/CAMPAIGN-11-recovery-journey-2026-08-05.md:508-513` and explicitly *"recorded, not filed"*: the off-site run emits **no hub event at all**, while both lesser tiers do (`db_dump_completed`, `crossdrive_completed`). Failures are covered by `backup_run_failures` and staleness by the hub's 8-day tier deadline, which is why it was judged a wrinkle. **Still true 2026-08-22** — the 2026-08-21 drill's own event dump shows `db_dump_completed` and six `crossdrive_completed` rows and no off-site success event. **Age when filed: 17 days.** | **OPEN — LOW** | — | Either emit one, or record deliberately that the highest-value tier is silent on success and say why. | CC | | **R-856** | Monitoring & notifications | P4 | **A crash restart reaches the household twice: the hub's "restarted after an unexpected stop" line AND the controller's app mails.** 2026-10-04 crash-guard test on demo-hp: after the third crash and the power-on, the controller sent `app_start_failed` (operator) and `app_stopped_unhealthy` (operator AND the household's address) for apps that were still coming up. Each is true on its own; the app ladder has no "the host just crashed" suppression like its boot grace for an ordinary restart (`08` §5). A design question for the operator, not a defect yet. `audits/os-docker-crash-2026-10-04/partC/c4-hub-events.txt` | **READY — operator decision** | — | — | operator | | **R-872** | Monitoring & notifications | P2 | **A box that is off every night never raises a missed-backup alarm: the 05:00 deadline check skips every customer whose node is `down`, so missing database dumps, second copies and off-site copies stay silent indefinitely; the only nightly signal is `node_down`.** MEASURED 2026-10-05 05:00 Budapest, hub log: `Deadline check: … 0 backup missed … 1 skipped (down)` — the skipped one is Tester 2, off since 18:06 UTC (`hub/internal/monitor/deadline.go` ~360: `if st == "down" \|\| st == StateDisabled { skipped++; continue }`). R-195 / R-321's shape again — a skip keyed off the wrong fact: "down now" was meant to avoid a double alarm, but a box down at every deadline is never checked at all. Fix direction (after R-871): count the days since the last success regardless of the node state, and alarm on N missed nights. `audits/night-fixes-2026-10-05/partF/FINDINGS.md` | **NARROWED 2026-10-05 — FIXED hub v0.134.0, proven by tests (3 red-proofs, `audits/catchup-2026-10-05/partC/`): a down box is judged on 48 h (dump) / 72 h (whole-guest) lines (`08` §6.4, decision 115). LEFT: the first live 05:00 run — DATED CHECK 2026-10-06 (DUE-CHECKS): the hub log line `Deadline check: Tester-2 is DOWN — judged on the longer lines … dump missed=1 backup missed=1` (if Tester 2 is still off at 05:00), and the two events in `events`. Holds → close; does not → a new row.** | R-871 | — | CC | +| **R-884** | Monitoring & notifications | P4 | **ArgoCD app `monitoring` shows `Deployment/prometheus` OutOfSync** (seen 2026-10-05 while syncing the R-173 alarm rules; only the rules ConfigMap was synced, so the Deployment drift is untouched and its cause unknown). A full sync would change the running Prometheus in an unknown way. | **OPEN** | — | `argocd app diff monitoring` (or the CR's resource diff) to see what differs, then decide git or live | operator | -## Hub & operator — 23 rows (P2 1, P3 7, P4 15) +## Hub & operator — 25 rows (P2 1, P3 9, P4 15) | ID | Category | Sev | What | State | Blocked on | Next action | Owner | |---|---|---|---|---|---|---|---| -| **R-173** | Hub & operator | P2 | **The hub's SQLite PVC is excluded from every Longhorn backup job.** `pvc/hub-data` carries `recurring-job-group.longhorn.io/default: disabled`, and `backup-daily` + `backup-weekly` (04:00 / Sun 05:00) are the ONLY recurring jobs and both target the `default` group — so the 128 MB `/data/hub.db` has **no volume-level backup**. That database holds `host_recovery` (every managed box's break-glass root password), `host_escrow` + `host_escrow_superseded` (escrow custody), `host_pbs_secrets`, `customer_configs`, `dr_recipe` and the wg endpoints/peers — i.e. the material several documented recovery routes depend on | **NARROWED 2026-10-05 — WAITING-ON-OPERATOR.** Measured (read only): the database IS backed up nightly — only on DooPlex: Longhorn `backup-daily`/`backup-weekly` (last 2026-10-05 02:06 UTC, Completed, 713 MB) to DooPlex's own `sda1`, because the live Volume carries `default: enabled` although the PVC (git, 2026-02-16, no reason) says `disabled` — a hand-set drift any sync may undo. Nothing leaves DooPlex; nothing alarms on failure (R-232). Since hub v0.135.0 the console passwords are sealed, so a copy is worth little without `OFFSITE_SECRET_KEY`, which also exists only on DooPlex. The off-site plan (keys off the box, the label fixed, a consistent `VACUUM INTO` snapshot pushed encrypted to ep0's PBS, a weekly restore test, Prometheus alarms): `runbooks/RUNBOOK-hub-db-offsite-backup.md`; the decision is in STATUS. `audits/hub-safety-2026-10-05/partC/` | — | **Noticed while checking the blast radius of the R-172 WAL change, not by a failure** — the WAL work needed to know who copies this file, and the answer turned out to be nobody on a schedule. **Establish before designing:** (a) whether the exclusion is deliberate (a 1 Gi RWO Longhorn volume snapshotting a 128 MB SQLite file is cheap, so the label looks like a leftover rather than a decision) and by whom; (b) whether anything else backs it up out-of-band that this census missed — the `_recovery-inventory-2026-07-28.md` records a MANUAL hot copy, which is not a backup. **When it is designed, it must be WAL-aware** (R-172): a volume snapshot of a live WAL database is crash-consistent and replays on open, which is fine, but any file-level copy must take `hub.db-wal` too or it silently loses the newest writes. **Grep establishing the ID was free:** `grep -ro "R-173\b" documentation/ *.md` → 0 hits | CC | +| **R-173** | Hub & operator | P2 | **The hub's SQLite PVC is excluded from every Longhorn backup job.** `pvc/hub-data` carries `recurring-job-group.longhorn.io/default: disabled`, and `backup-daily` + `backup-weekly` (04:00 / Sun 05:00) are the ONLY recurring jobs and both target the `default` group — so the 128 MB `/data/hub.db` has **no volume-level backup**. That database holds `host_recovery` (every managed box's break-glass root password), `host_escrow` + `host_escrow_superseded` (escrow custody), `host_pbs_secrets`, `customer_configs`, `dr_recipe` and the wg endpoints/peers — i.e. the material several documented recovery routes depend on | **NARROWED 2026-10-05 (evening) — option A IN FORCE** (`09` decision 125). The hub writes a nightly `VACUUM INTO` snapshot at 02:00 (hub v0.136.0, `05` §16.3, keep 2; the volume grew to 2 Gi); DooPlex checks it (`integrity_check`, size, ≥1 host, ≤26 h old), encrypts it with a key ep0 never sees and pushes it at 02:30 to ep0's `operator` namespace with a write-only token; a read-only token restore-tests it every Sunday 04:30 (and refuses a readable console password); `HubDBBackupStale`/`HubDBRestoreTestStale` alarm on success-only timestamps, `absent()` included. Both keys are off DooPlex (operator, 2026-10-05). PVC label fixed (`enabled`). Proven live: first push 7 s, restore test, token limits, a key rebuilt from the paper `data` field decrypts, runbook §3 steps 1–3 (4/4 console passwords open with the saved seal key, 0/4 with a random one). `audits/hub-db-offsite-2026-10-05/` | — | **LEFT:** runbook §3 steps 4–5 (the copy into a live PVC) are not exercised — they need the hub down; do them at the next planned hub maintenance or a DR drill on a scratch k3s. Close then. | CC | | **R-30** | Hub & operator | P3 | **[P2-HIGH] Liveness presence should come from the wait channel, not the report clock.** The box was powered off at the start of the rehearsal, yet the hub carried it as healthy until the staleness threshold expired ~30 min later (`host_stale` 16:05:24 "no report for 30m"; cleared 16:33:24 "was stale for 27m"). The host-delete guard compounds it: RESET refuses while any host row exists, so a stale-but-"Online" host stalls a forced teardown. **MIGRATED FROM `ROADMAP.md` 2026-08-22 (R-369) — originally filed 2026-07-21, size M, roadmap state `idea`.** Moved verbatim; nothing added or reinterpreted. The roadmap keeps its copy as history, marked moved. | **OPEN — migrated from ROADMAP 2026-08-22, rank unchanged** **Re-ranked 2026-10-03: P2→P3: operator-side presence delay; alarms still fire after 30 minutes and no household data is at risk.** | — | Direction: derive presence from **Dir-2 long-poll connectedness (~90 s grace)**, decoupled from notification hysteresis (the hysteresis is right for *alerting*, wrong for *presence*); an agent/ep0 analog can follow. Pairs with R-13/R-23 — the transport already exists, this is about believing it. *(Discussed in-session as "R-29"; that number was already taken by the gate-rot item earlier the same day, so it is R-30.)* | CC | | **R-31** | Hub & operator | P3 | **[P2-HIGH] Offsite provisioning is synchronous with no status affordance.** Save runs the Hetzner sync in-request, so the request can hit the nginx 504 **while succeeding server-side**: the operator cannot tell failed from slow, and a retry races the first attempt. **MIGRATED FROM `ROADMAP.md` 2026-08-22 (R-369) — originally filed 2026-07-21, size M, roadmap state `idea`.** Moved verbatim; nothing added or reinterpreted. The roadmap keeps its copy as history, marked moved. | **OPEN — migrated from ROADMAP 2026-08-22, rank unchanged** **Re-ranked 2026-10-03: P2→P3: operator-only; a known workaround (click once, wait, verify) exists.** | — | Direction: make it async + a status card, reusing the proven **awaiting-card/poll idiom** (v0.138.0 escrow card). **Interim mitigation belongs in R-3 as an operator note: click once, wait, verify — do not re-click.** | CC | | **R-244** | Hub & operator | P3 | **The customer DELETE cascade leaves `app_log_issues` behind, and it is systematic across every venue ever torn down.** Found **2026-08-07** while verifying the `finalwalk` teardown with a **full census** (every table, every column) rather than a per-table query. After a cascade that logged `COMPLETE … full teardown`, **61 rows still matched `finalwalk`**. Four of the five sources are **deliberate and correct** — the cascade's own header states *"Provenance/events are NEVER wiped — audit outlives every tier"*: `events` 16, `notification_log` 14, `host_deletions` 1, `customer_resets` 1. **The fifth is a gap:** `app_log_issues` 29 rows, which the residue purge does not touch (its logged leg covers `reports`/`app_telemetry`/`app_log_tails`/`log_tail_requests`/`notif_prefs`/`selfbind_tokens`/`appliance_registrations` — not this table). **It is not a `finalwalk` quirk:** rows still reference **`c11` 40, `rewalk` 20, `part4` 24** — all three torn down 2026-08-06, whose ledger recorded *"0 occurrences"*. **That prior claim was measured with a narrower query and does not survive a full census; the correction is recorded rather than the measurement quietly redone.** **Why it was probably never written, established rather than assumed:** the table is a **fleet-wide aggregate** keyed on `app_name`+`fingerprint` with an `affected_customers` JSON list — of the 29 `finalwalk` rows, **12 reference only `finalwalk`** (orphans, safely deletable) and **17 are shared with LIVE customers** (`demo-felhom`, `peti-felhom`, …) and **must not be deleted, only de-referenced.** A naive `DELETE … WHERE customer LIKE` would destroy a live customer's issue history — which is very likely why the leg does not exist, and is the reason this is not a one-line fix. **Severity is LOW and stated plainly: no secret material is involved** — app name, fingerprint, message text, counts, timestamps. What survives is a deleted customer's *identifier* inside an aggregate row. **Proposed shape:** a residue leg that (a) removes the customer id from `affected_customers`/`context_customer`, and (b) deletes rows whose `affected_customers` becomes empty; plus a one-off sweep for the four already-torn-down venues. **The general lesson is the reusable part:** *a per-table absence query is not a census.* The teardown verification is now a full-schema sweep, and that is what found this. **Not fixed** — a cascade change needs its own red-proof and this session was scoped as a spike plus two operations. Evidence: `tests/teardown-finalwalk-2026-08-07.md`. **⚠ STILL OWED, AND NOW MEASURED RATHER THAN ESTIMATED (2026-08-08 census, read-only, no truncation).** `app_log_issues` holds **1309 rows**; **71 reference a torn-down venue** (`finalwalk`, `c11`, `rewalk`, `part4`); of those **44 are ORPHANS** — they name only torn-down customers and are safely deletable — and **27 are SHARED with a live customer** (`demo-felhom`, `peti-felhom`, …) and **must be de-referenced, never deleted**. 1238 rows are untouched. **The 27 are exactly why the leg was never written**, and why a `DELETE … WHERE customer LIKE` would destroy a live customer's issue history. **What it needs, precisely:** a cascade leg that (a) removes the customer id from `affected_customers` / `context_customer`, and (b) deletes only rows whose `affected_customers` becomes empty; plus a one-off sweep for the four venues already gone. **Why it was NOT done on 2026-08-08:** the fix is hub code, and that session's scope forbade a hub version bump; a hand-run SQL mutation over 71 rows — 27 of them needing surgical de-referencing — with no tested code path and no red-proof is precisely the shape that goes wrong on a live database. **It accumulates one venue at a time, so the next walk adds to it**; the numbers above mean the next session starts from data rather than a guess. **⚠ IT GREW AGAIN, AS PREDICTED — walk5 teardown, 2026-08-08.** The fifth walk's venue was torn down with a full-schema census taken **before and after**: **168 rows → 67**. Of the 67, **37 are by design** (`events` 21, `notification_log` 14, `host_deletions` 1, `customer_resets` 1) and **30 are `app_log_issues`** — this row's gap, and the count was **predicted in the pre-run enumeration rather than discovered afterwards**, which is the difference from the ledger that once recorded *"0 occurrences"* from a narrower query. **The running total across torn-down venues therefore rises from 71 to ~101 rows** (`finalwalk`, `c11`, `rewalk`, `part4`, now `walk5`) — the shared-with-a-live-customer subset must still be **de-referenced, never deleted**. **It accumulates one venue at a time and it did so again.** Evidence: `tests/walk5-r201-2026-08-07/teardown-walk5-2026-08-08.md`. **2026-09-25:** `peti-felhom` is no longer a live customer (deleted through the cascade, journal #20); 8 `app_log_issues` rows still name it — the same gap. | **READY** — owner Viktor | — | — | operator | @@ -369,6 +369,8 @@ stopping line that lies. | **R-581** | Hub & operator | P3 | **[P2-MED] `ORDER BY received_at` cannot answer "the newest report" — the column has SECOND granularity.** FOUND 2026-09-18 building hub v0.118.0 (R-558): `Store.CustomerLanguage` read the household's language from the newest report ordered by `received_at`, and `TestNewestReportedLanguageWins` failed — four reports written in the same test tick all carry the same `datetime('now')` string, so the "newest" was whichever row SQLite felt like returning. On a real box the same shape appears whenever two reports land in one second (a settle burst, a restart race), and the symptom would have been a household switching language on their dashboard and getting the old language back at random. FIXED in the same release by ordering on the autoincrement `id`, which is the real insertion order. **What is still open, and it is the reason this is a row rather than a note: `GetCustomers()` has the same shape** — `INNER JOIN (SELECT customer_id, MAX(received_at) …)` with no tie-break — and it is what the whole operator dashboard and `countBoxesBelowFloor` read. A same-second tie there picks an arbitrary report's health, version and vitals. Not observed in the wild; not looked for either. **Fix shape:** tie-break every newest-report query on `id DESC`, or give `reports` a monotonic ordering column and use it everywhere; then a test that writes two reports in one tick and asserts which one wins. | **READY - rank P2-MED; owner: CC** **Re-ranked 2026-10-03: P2->P3: operator dashboard only; the household-facing language case was fixed; never observed.** | — | — | CC | | **R-600** | Hub & operator | P3 | **[P2-MEDIUM] "Full teardown" is logged while the deleted box's WireGuard peer is still configured on ep0.** FOUND 2026-09-20 by the slice-6 drill's teardown, **measured on ep0 rather than inferred from the hub**. The customer delete cascade finished at 19:41:54 with `customer DELETE cascade COMPLETE for drill-en-0920 (journal #18) — full teardown`, and every hub-side row was gone (0 configs, 0 hosts, 17 residue rows purged, PBS tenancy deprovisioned, escrow demoted). **Three minutes later `wg show wg0 allowed-ips` on ep0 still listed `10.77.0.5/32`** — the drill box's peer — because `wgsync` pushes on its own cycle. **Watched to the end rather than assumed: the peer was gone by 17:47:46Z — it outlived the *full teardown* line by about 6 minutes.** (My first estimate said ~35, read off the gap between two log lines; the sync runs oftener and only LOGS when something changes. That is the second time in this session that a period inferred from two log lines was wrong — the other was the delete's own staleness window. **A period read off two log lines is not a measurement.**) **The 2026-09-14 drill's findings say "the teardown removes it through the host delete"; measured, the host delete removes the hub's RECORD and the peer goes on the next push.** The mechanism is not broken — it is asynchronous, and the log line claims a completeness it does not yet have. **Why it is P2 rather than P3:** a session that tears down, reads *full teardown*, and leaves is the normal case; the peer outlives it by minutes, and the third teardown layer is the one the workspace rules single out as the one that gets forgotten. Six minutes is short — but the session that reads *full teardown* and leaves has no way to know it is six and not six hours. **Fix shape (smallest first):** the cascade triggers a wgsync push before it logs COMPLETE, or the log line says what is still pending and when ("wg peer removal queued; next push in N min"). A session should not have to read ep0 to know whether a teardown finished. **2026-09-25 (Peti's retirement):** nothing to remove for `peti-felhom-86d37d` — its host was deleted 2026-07-15 and ep0's live `wg show` carries no peer beyond the demo boxes', drill-r50 and the operator OOB (`audits/retire-peti-2026-09-25/A1-ep0-before.txt`); whether that July delete removed a peer, or none existed, is not recorded. **-- 2026-09-28: the Day-0 test install's peer (`drill-g0276`, key `Ly0yjK…`, 10.77.0.5) was checked on ep0 the day after its customer DELETE: gone from `wg show`, absent from `/etc/wireguard/*.conf` (control: a live key found), no namespace or file names the customer — the hub's sync removed it; nothing was removed by hand.** `audits/logins-nvme-2026-09-28/E/`. **MEASURED AGAIN 2026-09-30 (a HOST delete, not a customer delete; new-household drill):** the host was deleted 07:23:03Z and `wgsync: pushed 4 peers` at 07:23:51Z — `10.77.0.5/32` gone from ep0 48 s later, the other four peers unchanged (`audits/evidence-drill-new-household-2026-09-30/teardown/layer4-ep0.txt`). | **READY - rank P2-MEDIUM; owner: CC (hub)** **Re-ranked 2026-10-03: P2->P3: measured as seconds to minutes and asynchronous by design; operator-only log wording.** | — | — | CC | | **R-728** | Hub & operator | P3 | **[P3-LOW] A customer created with one press was created TWICE, and the first of its two connect mails holds a dead link.** MEASURED 2026-09-30 on `Tester-2`: the hub logged `Customer config created: Tester-2` twice in the same second and two self-bind mints (hashes `c40df008…`, `6a1cbef4…`); a mint replaces the previous link (single-active), so one of the two identical mails the tester received answers „expired". Cause not established (a double form submit, or the handler run twice). **Fix direction:** make the create idempotent within a few seconds (or disable the button on submit), and pin it. The workaround for the tester is in STATUS. | **READY — rank P3-LOW; owner: CC (hub)** | — | — | CC | +| **R-882** | Hub & operator | P3 | **Longhorn on DooPlex could not grow a volume online: its `instance-manager` (116 days up) called a host process that no longer existed** — `nsenter: cannot open /host/proc/196610/ns/mnt` on every expansion retry, and an offline growth was blocked by the expansion's own attachment ticket (found 2026-10-05 growing `hub-data` to 2 Gi). A restart of the instance-manager (operator-approved) fixed it: 77/77 volumes back `attached/healthy` in 110 s. **Why the cached PID went stale was not established** (likely a containerd/k3s or iscsid restart after the instance-manager started), so it will recur after the next such restart and stay invisible until a volume needs to grow. `audits/hub-db-offsite-2026-10-05/partA/step1-*.txt` | **OPEN** | — | Find which host process the PID was and whether Longhorn 1.10.x re-resolves it; until then, before growing any volume, check the instance-manager's age against the last k3s/containerd restart | operator | +| **R-883** | Hub & operator | P3 | **8 DooPlex workloads run an image by a moving tag (`:latest` or none), so any pod restart is a silent upgrade.** Measured 2026-10-05: the Longhorn restart restarted zipline on `ghcr.io/diced/zipline:latest` (pull Always), which pulled 4.8.0; 4.8.0 refused its database (`cannot safely migrate from prisma to drizzle: expected migration 20260508022000 … was not applied`) and crash-looped. Fixed for zipline by pinning `4.7.0` (homelab-manifests `90f60e4`, `4c8ec7a`; 4.7.0 applied the four missing migrations; a dump from before is kept out-of-band). The other 7 were counted, not named or changed. | **OPEN** | — | List the 7 (`kubectl get deploy,sts -A` images without a fixed tag), pin each to the running version, and let Renovate move them | operator | | **R-92** | Hub & operator | P4 | Hub PBS-DR gauge is 0.1 GB-granular — small deltas unverifiable | READY (XS) | — | Widen precision when retention becomes customer-visible | CC | | **R-261** | Hub & operator | P4 | **C6 — `CountSelfBindTokens` exists so that callers can assert an invariant, and no production caller asserts it.** `hub/internal/store/selfbind.go:106-111`. Its doc comment: *"it exists so callers can assert the 'after this runs, the only live link is one we just issued — or none' invariant that the auto-mint at customer-create / RESET-completion depends on."* **Census: only its own declaration in production; the two callers are `selfbind_automint_test.go:29` and `customer_delete_test.go:510`.** Tests are not callers (the campaign's rule), so the invariant the auto-mint *depends on* is checked in the test suite and never at the moment it matters. **This is the smallest of the eight rows and is filed at its true size, because the rest of the C6 sweep found INERT dead accessors rather than defects:** `OffboxOrphanedRenamedTo` and `OffboxEscrowState` have no caller but their data reaches the card another way (the template reads the settings field directly, `backups_remote.html:80`) — **R-228 is genuinely closed, and the sweep's first reading that it had regressed was wrong.** **The more consequential C6 result is a method result and is in the report, not here:** `golang.org/x/tools/cmd/deadcode` re-finds **neither** known instance, and a planted probe measured why — it reports an unreachable exported FUNCTION and not an unreachable exported METHOD on a widely-used type, and both known instances are methods | **READY** — owner Viktor | — | — | operator | | **R-264** | Hub & operator | P4 | **Twenty-one facts the boxes report that the hub can now decode nowhere, each allowlisted with a reason rather than silently skipped — and for these the reason is "no consumer today, and one is arguably owed".** Split out of R-260 on 2026-08-08 so that closing the CLASS (gated) and fixing its sharpest instance (`operator_key_configured`) could not be mistaken for having decided what the hub should do with the rest. **The list, grouped by what a consumer would be for.** **(a) Guest-network health — `guest_net` and its seven children** (`checked_at`, `has_route`, `dhclient_alive`, `heal_succeeded`, `heals_last_hour`, `last_heal_at`, `damped`). The R-54 watchdog reports per-guest network state and self-heal counts every cycle and the hub — the component that emails the operator — models none of it. There is a live incident in this project's own record where a killed `dhclient` took a tunnel down for 1 h 15 m (`audits/INCIDENT-guest-dhclient-killed-2026-07-20.md`); a recurring-heal signal is exactly what would have surfaced it. **This is the strongest candidate of the twenty-one.** **(b) `selfupdate_pending` + `selfupdate_pending_version`** — an agent that has flipped its binary and never committed reports pending on every heartbeat so that "the operator sees WHY the version isn't advancing", and no operator can see it. **(c) `mgmt_plane.healed_recently`** — bounded: the hub DOES alarm on the `privsep_healed_at` timestamp beside it, so the recurring-clobber signal is not lost, only this flag. **(d) `restore_tests.mount_parity` + `mount_inventory`** — R-262's subject; the verdict is not lost (a mismatch fails before `Pass` is set) but the hub cannot tell a full-fidelity pass from a boot-only one. **(e) `pbs_dr.applied_at`.** **(f) Controller-side: `config_hash`, `reporting_disabled`, `stacks`, `storage.migrated_to`, `backup.last_db_dump`, `backup.last_integrity_check`** — the last two are backup-integrity timestamps, which is the "presence is not success" neighbourhood. **For each the question is the same and is NOT answered here: is it wanted? If the hub should act on it, model it and name what consults it. If it should not, the honest end is that the emitter stops sending it** — a fact emitted forever and consumed nowhere is a future false green waiting for someone to write a check against it. **Deliberately not decided in the G-1 session**, whose scope was the gate plus the operator-access instance; unilaterally removing emitters would also break the byte-identical cross-repo host-report golden and is a coordinated two-repo change. **⚠ DISPOSITIONS RECORDED 2026-08-13 — and the first thing to say is that they were NOT in this register before today.** The rulings were made on 2026-08-12; this row still read **READY — owner Viktor** and the gate's twenty entries still all said *"arguably owed"*, so a session told to *"re-read the dispositions from the register"* would have found none. They are written down now, which is the point of writing them down. **THE COUNT WAS ALSO WRONG:** this row says *twenty-one*; the gate's allowlist held **twenty**, measured. Twenty is the number the dispositions below account for, exactly. **(1) BUILD A READER — four groups, fourteen facts.** (a) guest-network health, (b) the staged-update-pending pair, (d) the restore-test depth pair, (f-part) the two backup-integrity timestamps. **(2) NO READER WANTED — five facts, now recorded as `not consumed, DELIBERATELY` with the ruling and its date** in `scripts/wire_contract_gate.py`, each with its own reason rather than a bare refusal: `mgmt_plane.healed_recently` (the hub already alarms on the timestamp beside it), `pbs_dr.applied_at` (`pbs_dr.state` is the verdict; the timestamp alone is the attempt-read-as-result trap), `config_hash` (the hub authors the config and knows its own generation), `stacks` (the app view is built from the purpose-built `app_telemetry` wire), `storage.migrated_to` (box-local bookkeeping with no hub-side intent to reconcile against). **The emitters are deliberately left alone** — removing one is a coordinated two-repo change and breaks the host-report golden; the honest end here is a recorded decision, not a deletion. The gate grew a THIRD entry kind to carry them, because an undecided fact and a decided one must not read alike. **(3) `reporting_disabled`, decided on its own merits: RECLASSIFIED `redundant`** — `health.status = "disabled"` travels in the same minimal report, is decoded into `reports.health_status`, and IS rendered. **The decision surfaced a real defect the flag would not have fixed: the staleness checker is age-only, so a deliberately-silent box still alarms → R-321.** **PROGRESS, 2026-08-13: the first reader is BUILT — guest-network health, R-319.** Its eight allowlist entries are **removed** (an allowlisted tag is skipped, so leaving them would have meant the new reader's fields were never checked); the gate's checked-tag count rose **182 → 190** and skipped fell **88 → 80**, which is the positive control that the wiring is real. **WHERE THE TWENTY NOW STAND: 8 read · 5 deliberately unread · 1 redundant · 6 still owed a reader** (`selfupdate_pending`, `selfupdate_pending_version`, `restore_tests.mount_parity`, `restore_tests.mount_inventory`, `backup.last_db_dump`, `backup.last_integrity_check`) — counts measured from the allowlist, not estimated. **Only ONE reader was built on purpose:** four at once is a design session pretending to be an implementation, and this one now tells us what the other three cost | **OPEN — 6 of 20 still owed a reader; dispositions recorded 2026-08-13, first reader shipped (R-319)** — owner Viktor | — | — | operator | @@ -397,7 +399,7 @@ stopping line that lies. | **R-793** | Business & legal | P4 | **[P3-LOW] Enterprise / BUSL code ships inside four open images — Cal.com and Docmost (EE folders, off without a key), Outline (BUSL-1.1: no commercial "Document Service"), meilisearch v1.36 in Wanderer (EE modules).** READ 2026-10-02 (`audits/licences-2026-10-02/TABLE.md`). Each is fine as the catalog runs them: no EE key, the household's own Outline is not a Document Service, Wanderer uses plain search. **Watch:** never turn on an EE feature, never switch Karakeep's/Wanderer's meilisearch to the `-enterprise` image, and re-read on each major. | **WATCHING — rank P3-LOW; owner: CC** **Re-ranked 2026-10-03: P3→P4: a watch item; nothing is wrong as the catalog runs them.** | — | — | CC | | **R-794** | Business & legal | P4 | **[P3-LOW] redis 7.4 (RSALv2 / SSPL, not OSI) runs as a private cache in seven apps: dawarich, docmost, immich, nextcloud, outline, paperless-ngx, romm.** READ 2026-10-02 (`audits/licences-2026-10-02/TABLE.md`). Read as permitted (a private cache only its app uses is not Redis offered as a service — inferred). Valkey (BSD-3) or redis 8 (AGPL option) removes the question. **Needs:** a ladder step per app to valkey or redis 8, through the harness — no hurry. | **READY — rank P3-LOW; owner: CC** **Re-ranked 2026-10-03: P3→P4: the row itself says no hurry; usage read as permitted.** | — | — | CC | -## Process & tooling — 88 rows (P3 4, P4 84) +## Process & tooling — 89 rows (P3 4, P4 85) | ID | Category | Sev | What | State | Blocked on | Next action | Owner | |---|---|---|---|---|---|---|---| @@ -489,6 +491,7 @@ stopping line that lies. | **R-818** | Process & tooling | P4 | **Two changelogs cite register ids for other findings.** `hub/CHANGELOG.md:526-538` (hub v0.109.0, the Backup card) and its note at `:533` use **R-331** and **R-330** for a hub display fault and a nightly false app alarm; in the register R-330 and R-331 are the disk-health Phase 2 and Phase 3 rows. A reader following the id lands on the wrong finding. Found 2026-10-03 by the triage. | **READY — filed 2026-10-03 (triage); owner: CC.** Add a dated correction line under each changelog entry naming the right rows (the real ids are in `CLOSED-ITEMS.md`); do not renumber anything. | — | — | CC | | **R-819** | Process & tooling | P4 | **`scripts/check_stands.py` is red and runs in no runner.** Measured 2026-10-03 on `9e2786c` (before the triage): it convicts `where-felhom-stands.yaml` for citing R-273 and R-356, which are in neither `OPEN-ITEMS.md` nor anywhere it reads. After the triage it also convicts R-281, R-198 and R-201, because its rule 3 reads only `OPEN-ITEMS.md` and those rows are closed. It is in neither `repo_gates.py` nor CI, so nobody saw it — the R-29 shape. | **READY — filed 2026-10-03 (triage); owner: CC.** Let rule 3 accept an id in `CLOSED-ITEMS.md` (and check the stand's status agrees), fix the two dangling ids, then register it in `repo_gates.py` with a decoy. | — | — | CC | | **R-857** | Process & tooling | P4 | **Baking a golden twice under the SAME version leaves two stale facts.** 2026-10-04 (golden 0.292.0 re-baked with live-restore): (1) `golden_currency_gate.py` keeps reporting the FIRST bake directory's sha (`golden-0.292.0-2026-10-04`, d6cf8b33…) — it picks one of two directories with the same version, not the newest; (2) the publish replaces the package in place (pre-delete + upload), so until the operator re-vouches, the hub vouches a sha the registry no longer holds — a fresh install in that window fails its sha check (fail-closed; none ran; re-vouched 17:48). Fix direction: the gate prefers the newest bake directory (or refuses two for one version); the runbook says "re-vouch at once after a same-version re-bake". `documentation/tests/golden-0.292.0-2026-10-04-rebake/` | **READY — owner: CC** | — | — | CC | +| **R-885** | Process & tooling | P4 | **The Python tests under `felhom.eu/scripts/` do not run in CI.** `.gitea/workflows/gates.yml` runs `repo_gates.py`, which runs gates, not the `test_*.py` suites (`scripts/test_*.py`, `scripts/hub-db-backup/test_hub_db_backup.py` — 15 tests of the DooPlex push/restore scripts, added 2026-10-05 and run by hand). A change that breaks one is caught only if someone runs it. | **OPEN** | — | Add a CI step (or a gate) that runs every `scripts/**/test_*.py`, with a decoy (R-421) | CC |