From 81c835827ec5d3de4a8fdf137438178a9f48377e Mon Sep 17 00:00:00 2001 From: kisfenyo Date: Fri, 9 Oct 2026 10:54:59 +0200 Subject: [PATCH] =?UTF-8?q?R-232/R-173:=20hub=20restored=20into=20a=20thro?= =?UTF-8?q?waway=20k3s=20(runbook=20=C2=A73=20steps=204=E2=80=935=20proven?= =?UTF-8?q?,=20corrected);=20R-173,=20R-861,=20R-518=20closed;=20R-921,=20?= =?UTF-8?q?R-922=20filed;=20STATUS,=20capability=20map,=20report?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS --- REPORT-dooplex-survival-2026-10-09.md | 92 +++++++++ STATUS.md | 18 +- .../architecture/00-capability-map.md | 3 +- .../partD-01-k3s.txt | 27 +++ .../partD-02-hubdb-restore.txt | 19 ++ .../partD-03-transfer.txt | 5 + .../partD-04-deploy.txt | 49 +++++ .../partD-05-checks.txt | 26 +++ .../partD-06-teardown.txt | 175 ++++++++++++++++++ documentation/backlog/CLOSED-ITEMS.md | 10 + documentation/backlog/OPEN-ITEMS.md | 13 +- .../runbooks/RUNBOOK-hub-db-offsite-backup.md | 24 ++- 12 files changed, 448 insertions(+), 13 deletions(-) create mode 100644 REPORT-dooplex-survival-2026-10-09.md create mode 100644 documentation/audits/dooplex-survival-2026-10-09/partD-01-k3s.txt create mode 100644 documentation/audits/dooplex-survival-2026-10-09/partD-02-hubdb-restore.txt create mode 100644 documentation/audits/dooplex-survival-2026-10-09/partD-03-transfer.txt create mode 100644 documentation/audits/dooplex-survival-2026-10-09/partD-04-deploy.txt create mode 100644 documentation/audits/dooplex-survival-2026-10-09/partD-05-checks.txt create mode 100644 documentation/audits/dooplex-survival-2026-10-09/partD-06-teardown.txt diff --git a/REPORT-dooplex-survival-2026-10-09.md b/REPORT-dooplex-survival-2026-10-09.md new file mode 100644 index 00000000..4e327b8e --- /dev/null +++ b/REPORT-dooplex-survival-2026-10-09.md @@ -0,0 +1,92 @@ +# REPORT — can the business survive losing DooPlex? (2026-10-09) + +| Part | What | Result | +|---|---|---| +| A | Measure; plan | DONE — `audits/dooplex-survival-2026-10-09/PLAN.md`; operator said **yes** in chat | +| B | Nightly encrypted copy of Gitea + secrets to ep0 | DONE, LIVE — first push 57 s; read back; alarms in force; failure mail proven | +| C | Gitea restored into a throwaway | DONE — 10/10 repos, `main` = live, a file byte for byte, a login; deleted | +| D | Hub restored into a throwaway k3s (R-173 steps 4–5) | DONE — customers 4/4, hosts 4/4, console passwords 4/4; deleted; R-173 closed | +| E | R-861 (a), R-518, the mail-address row | R-861 closed, R-518 closed (+ R-921), R-922 filed | + +**Register: before 135 · after 135 · opened 2 (R-921, R-922) · closed 3 (R-173, R-861, R-518).** A parallel session +added one row meanwhile (the file read 136 before this session's edit). + +Architecture read before any claim: `07-backup-architecture.md` (trust model, encryption), `06-offsite-connectivity.md`, +`01-topology-and-trust.md`, `runbooks/target-selection.md`, R-232 + its recon, the R-173 runbook. + +## Part A — measured (read only) + +- Gitea 1.26.2: repositories **614 MB** (10 repos), database `gitea` 64 MB live = **6.2 MB** as a dump (~150 KB/day), + `app.ini` 2 KB (holds Gitea's secrets). Registry 27.7 GB — **left out on purpose** (images rebuild from code). +- Secrets set: three GPG files per night, **4.4 MB**, flat. +- **Correction to the recon (§7):** the API mirror holds ONE repo (`homelab-manifests`). The product repos had one + backup only: Longhorn, retain=1, on DooPlex. +- ep0 `operator`: 83 GB free of 98 GB; the hub-DB's write-only token already reaches it. **Nothing to create on ep0.** + +## Part B — what was installed on DooPlex (operator yes) + +- `/usr/local/sbin/felhom-dooplex-offsite` (+ restore test, + `felhom-backup-failmail`), from `scripts/dooplex-offsite/` + at felhom.eu `1707c928`/`02a54e26`. Timers: daily 00:20 and Sun 05:30. +- New key `/etc/felhom-dooplex-offsite/enc.key` (root 0600, `--kdf none`, never printed). Tokens: the hub-DB's, read in place. +- `OnFailure=felhom-backup-failmail@%n.service` on the two new units **and the two hub-DB units**. +- homelab-manifests `691db39`: `DooplexGiteaOffsiteStale` (26 h) + `DooplexGiteaRestoreTestStale` (8 days), `absent()` + included; only the rules ConfigMap synced; `/-/reload` 200; `/api/v1/rules` reads both `ok`. +- **Consistency choice:** Gitea's docs say stop it for a consistent backup. Not chosen — a nightly stop costs CI and the + registry. Instead: the newest complete database dump (taken first), then the files; the restore test runs `git fsck` + on every repo. +- Proof: first push 544 MB in 57 s; restore test with the read-only token: 27 805 files match, 10 repos pass `git fsck`; + the push token asked to forget a snapshot → `permission check failed`; ep0's disk shows the new group. +- **Found and fixed live:** the failure-mail script died under `set -u` on the shared config's unset variable (the first + dry run sent no mail). Red test, fix, second dry run → Resend id, mail in the inbox 08:17:00Z. +- Security review findings, both fixed and tested: links in the pod's archive are refused; root `git fsck` never reads a + repo's own `config` (red-proved with a config git refuses). +- Tests: 20, green with GNU and BusyBox tools; red-proofs for 7 checks (`partB/red-proof.txt`); promtool green + 2 reds. + +## Part C — Gitea from ep0 into a throwaway (the bench, LXC 9401 on demo-hp) + +Restored on DooPlex with the read-only token (22 s), streamed to the bench **without the secrets files**, Postgres 17.2 +and Gitea 1.26.2 on an `--internal` Docker network (no route out: `wget gitea.com` → bad address). 10/10 repos; three +product `main` equal live; `felhom.eu` one commit behind — that commit (`02a54e26`) was pushed 3 min after the copy and +the copy's `1707c928` is its parent; `CLAUDE.md` sha256 equal; throwaway admin logged in. Runbook: `runbooks/gitea-restore.md`. + +## Part D — the hub into a throwaway k3s (same bench) + +k3s v1.33.6 in a Docker container on an internal network (two fixes needed: the cgroup v2 move, `--flannel-iface eth0`). +Hub image of the live digest, by file. Runbook §3 steps 4–5 as written. Start-up: `console passwords sealed at rest +(0 legacy …)`. Customers 4/4 and hosts 4/4 equal live (`GET /configs`, `/hosts`, live read GET-only). 4/4 reveals → +32-char passwords (length only); second channel `hubdb-check`: 4/4 with the saved key, 0/4 random. **Found:** the +restored hub tried to mail two households a pending kernel notice within a minute — only the cut network stopped it; +operator-mail-off does not. Runbook §3 corrected (and: three more Secrets are not optional). + +## Teardown (three layers) + +- **Machine (bench 9401):** containers 0, volumes 0, networks back to the original four, images removed, `/root` as before, + stopped again (it was stopped). **Found:** Postgres and k3s leave anonymous volumes holding the data; 1 + 8 found by + counting, each checked by creation time, removed by name (no prune). +- **Host (DooPlex):** both scratch dirs shredded and removed; the seal-key and password files shredded. +- **Hub:** provisioned nothing. The live hub was only read (GET). + +## Part E + +- **R-861 (a) CLOSED:** all three 0.304.0 swaps went through `felhom-priv-apply controller-image` (sudo log + wrapper log + + agent journal) and the guests run 0.304.0 (Docker). No `tee` that day. +- **R-518 CLOSED:** 2026-10-08 on demo-hp both tiers ran, each with its own stop under ~2 min. A third stop of ~1 min + with no copy happened when the off-site tier answered BUSY → **R-921** (P3, needs a release). +- **R-922 filed** (P2): a household's cleared mail address stays on the hub (the no-clobber guard). Rule: operator. + +## Other + +- CI: `1707c928`, `02a54e26` first failed with every step red and no runner log (R-887, dropped jobs); re-run → success. + `994826ec` success. I pushed the second commit before reading the first run — the rule says check first. +- `unproven.py --summary`: not walked 35 of 55 (unchanged). +- Instruction files: none edited. + +## For the operator + +1. **Save the new key's paper copy** (at your own terminal, not through `!` here): + `sudo proxmox-backup-client key paperkey /etc/felhom-dooplex-offsite/enc.key --output-format text` → the `data` field + into the password manager as „DooPlex off-site (Gitea) key". **If you do nothing:** after a DooPlex loss the copy on + ep0 cannot be opened. Pick: do it today. +2. **R-922 — a household clears its mail address:** (A) the hub drops the address when the household clears it (the + controller says so explicitly); (B) keep it, and write the keep time into the privacy notice. **Pick A.** **If you do + nothing:** the address stays on the hub with no end date. diff --git a/STATUS.md b/STATUS.md index 2d9eee45..d0f2b1ec 100644 --- a/STATUS.md +++ b/STATUS.md @@ -2,8 +2,22 @@ **Ready for the first real tester (Tester-2): yes. Tester 2 (a laptop) is off; nothing was sent to it.** -**Updated 2026-10-09 (morning): hub 0.144.0; demo-hp, demo-felhom and Tester 1 run agent 0.154.0 and controller -0.304.0. The open-items list is at 135. Report: `REPORT-day4-2026-10-09.md`.** +**Updated 2026-10-09 (midday): hub 0.144.0; demo-hp, demo-felhom and Tester 1 run agent 0.154.0 and controller +0.304.0. The open-items list is at 135. Reports: `REPORT-dooplex-survival-2026-10-09.md`, `REPORT-day4-2026-10-09.md`.** + +## Midday (2026-10-09): the code and the hub can now survive losing DooPlex + +- **Every night at 00:20, all the code (Gitea) and DooPlex's secrets go to ep0, encrypted.** You said yes in chat. + The key that writes the copy cannot delete or read old copies. Nothing changed on ep0. +- **Gitea was brought back from that copy on a throwaway machine** — the first real restore. All 10 repositories are + there, the newest commits match, a file matched byte for byte, a login worked. +- **The hub was brought back from its copy into a throwaway k3s.** Same 4 customers, same 4 boxes, all 4 console + passwords open. Found: a restored hub mails households at once. The test had no network, so nothing went out. + The runbook now says so. +- **Every backup job on DooPlex now mails you when it fails.** A test mail reached the inbox. +- **You need to do one thing:** save the new key's paper copy in your password manager (the report has the command). + If you do nothing, the copy on ep0 cannot be opened after DooPlex is lost. +- **Not copied on purpose:** the container registry (27.7 GB). The images rebuild from the code. ## Morning (2026-10-09): released, delivered, and four of your answers proven live diff --git a/documentation/architecture/00-capability-map.md b/documentation/architecture/00-capability-map.md index 37697580..c990e12c 100644 --- a/documentation/architecture/00-capability-map.md +++ b/documentation/architecture/00-capability-map.md @@ -235,7 +235,8 @@ likewise silent. Evidence: `audits/DRILL-r361-2026-08-22/evidence/06-part3-decis | Secrets hygiene: bearer in k8s Secret, no secrets in git, single-quote credential store | hub v0.53, conventions | **IMPLEMENTED** | 07-13 closing bundle | | | Operator login password changeable from UI | hub v0.54 | **IMPLEMENTED** | 07-13 | | | **The operator sees boxes left behind: per-customer controller floors with their age (and which a global raise cannot move), each box's agent against the vouched one, a 7-day "agent behind" alarm and a "floor raise skipped boxes" mail** | hub **v0.135.0** (R-604, R-530) | **IMPLEMENTED — the page PROVEN-LIVE, the alarm and the mail unit-proven** | `audits/hub-safety-2026-10-05/partD/` (live System page: three per-customer floors, Tester 2 `0.142.0 → 0.145.0`); `osupdates/r530_agent_alarm_test.go`, `web/r604_floor_held_back_test.go` | the mail fires only on a GLOBAL raise below an override — not exercised live | -| **The hub database survives the loss of DooPlex: a nightly consistent copy, encrypted, on ep0; restore-tested weekly; an alarm when either stops** | hub **v0.136.0** (R-173), `scripts/hub-db-backup/`, homelab-manifests rules | **PROVEN-LIVE (2026-10-05)** — first push, ep0 listing, restore test (4 hosts, 4 sealed, 0 readable), token limits, a key rebuilt from the paper copy decrypts, the saved seal key opens 4/4 console passwords in the restored copy; the alarm by `promtool` rule test + red-proofs | `audits/hub-db-offsite-2026-10-05/`; `runbooks/RUNBOOK-hub-db-offsite-backup.md` | Runbook §3 steps 4–5 (into a live PVC) not exercised (R-173) | +| **The hub database survives the loss of DooPlex: a nightly consistent copy, encrypted, on ep0; restore-tested weekly; an alarm when either stops** | hub **v0.136.0** (R-173), `scripts/hub-db-backup/`, homelab-manifests rules | **PROVEN-LIVE (2026-10-05)** — first push, ep0 listing, restore test (4 hosts, 4 sealed, 0 readable), token limits, a key rebuilt from the paper copy decrypts, the saved seal key opens 4/4 console passwords in the restored copy; the alarm by `promtool` rule test + red-proofs ; **2026-10-09: the whole recovery (§3 steps 1–5) PROVEN on a throwaway k3s** — the live hub image started on the restored copy, customers 4/4 and hosts 4/4 equal live, 4/4 console passwords revealed | `audits/hub-db-offsite-2026-10-05/`; `audits/dooplex-survival-2026-10-09/partD-*.txt`; `runbooks/RUNBOOK-hub-db-offsite-backup.md` | A restored hub mails households pending notices at once — a test restore has no network (runbook §3) | +| **Gitea (all code) and DooPlex's secrets survive the loss of DooPlex: a nightly encrypted copy on ep0; restore-tested weekly; an alarm when either stops; a failure mail** | `scripts/dooplex-offsite/` (R-232), homelab-manifests rules | **PROVEN-LIVE (2026-10-09)** — first push 57 s, read back with the read-only token (27 805 files, 10 repos pass `git fsck`); Gitea restored into a throwaway and started: 10/10 repos, product `main` = live, a file byte for byte, a login; alarm by `promtool` + red-proofs; failure mail reached the inbox | `audits/dooplex-survival-2026-10-09/`; `runbooks/gitea-restore.md` | The container registry is NOT copied (images rebuild from the code); the on-box backup tree is still one writable path (R-232 c) | | **The operator surface refuses a cross-site form post on BOTH login paths (session token; Basic auth + `X-Felhom-Operator`)** | hub **v0.135.0** (R-135) | **PROVEN-LIVE (2026-10-05)** | `audits/hub-safety-2026-10-05/partA/live.txt` (Basic, no header → 403 even on an unknown path; with the header → passes; header without credentials → 401); `web/r135_csrf_test.go` (39 paths) | | | Box operating-system security updates (Proxmox host, guest Debian, Docker engine) | agent v0.143.0, hub v0.133.0 | **PARTIAL — the GUEST and HOST Debian fast lanes and the DOCKER engine slow lane are PROVEN-LIVE (2026-10-04), with the System page, the fleet view and the alarms; the KERNEL lane is MISSING** | Guest: `audits/os-guest-lane-2026-10-04/`. Host + fleet + alarms: `audits/os-host-lane-2026-10-04/`. Docker + System page + crash guard: `audits/os-docker-crash-2026-10-04/` — live-restore on with the same container ids on every box; Docker 29.8.2 on both demo boxes; operator-approved Docker release; a signed undo and a signed ring-1 step; a replay refused; the crash guard restarted demo-hp twice and kept it off the third time. Design `architecture/11-os-updates.md` §5.8, §5.9, §8 | **No automatic undo** (guest: last night's backup; host: by-hand runbook; Docker: a signed undo job); existing boxes get root-owned files by the signed config bundle since 2026-10-04 (R-840 CLOSED; a box from before agent 0.143.0 needs one by-hand bootstrap — Tester 2: R-862; `audits/r840-config-bundle-2026-10-04/`); test approvals now end with the test (R-859); the agent's sudoers is root-equivalent (R-861); the kernel lane (R-836); facts reach the hub late after a boot (R-853). **2026-10-05 (agent v0.144.1):** R8 measures the real download (R-865); a killed pass still reports (R-868, live); the debug pass runs with the hub away (R-866, live); **a power cut mid-update was proven by day on demo-hp — the box came back by itself in 37 s, but the next pass fails until `dpkg --configure -a` is run by hand (R-876, P2, open)** — `audits/night-fixes-2026-10-05/`. **2026-10-05 afternoon (agent v0.145.0): R-876 FIXED and proven live — after a second crash mid-unpack the next pass repaired dpkg by itself (`REPAIR … journal=1`) and finished** — `audits/catchup-2026-10-05/partD/` | | **An ENGLISH-SPEAKING household's first hour: download, install, pair, bind, claim, two apps** | controller **v0.259.0** + hub **v0.119.0** + ISO 1.29.0 + the whole catalog | **PROVEN-LIVE on 0.258.0 with one blocker; THE BLOCKER IS FIXED AND PROVEN, THE WALK IS NOT REPEATED** | `audits/DRILL-first-hour-en-0258-2026-09-20.md` — a fresh install 2026-09-20, one intervention (R-494), stop rule not reached. Then `audits/i18n-closing-2026-09-21/live/` — the three blockers fixed and each proven on a live box or in the operator's inbox: the claim page answers English through the real cookie path; the Backup page's tier names follow the language; and the setup mail carries **four plain-ASCII English words** where the drill's carried `képző-szkítia-ásatás`, one day apart in the same inbox. | **R-596, R-597 and R-598 are CLOSED.** What this row still does NOT claim: **the fixed journey has not been walked end to end by a stranger on a fresh install.** Three fixes proven at the endpoint are not an hour proven by a person, and this project's own rule is that fixes are not a journey (see the recovery-journey row). **Also not walked:** the recovery code (needs ep0), backup/restore/remove/power-cut (proven 2026-09-14), and the two Backup-page *warnings* themselves — guest 9201 is healthy and a healthy box renders none, so they are covered by handler render tests, not live. **Verdict: nothing known now stands between an English-speaking tester and their box — and that is a different sentence from "the walk passed".** | diff --git a/documentation/audits/dooplex-survival-2026-10-09/partD-01-k3s.txt b/documentation/audits/dooplex-survival-2026-10-09/partD-01-k3s.txt new file mode 100644 index 00000000..01efcf7a --- /dev/null +++ b/documentation/audits/dooplex-survival-2026-10-09/partD-01-k3s.txt @@ -0,0 +1,27 @@ +## 2026-10-09T08:42:44Z fetch the k3s image + airgap set (bench network, before isolation) +docker.io/rancher/k3s:v1.33.6-k3s1 +total 140296 +drwxr-xr-x 2 root root 4096 Oct 9 08:36 . +drwxr-xr-x 3 root root 4096 Oct 9 08:36 .. +-rw-r--r-- 1 root root 143649404 Oct 9 08:36 k3s-airgap-images-amd64.tar.zst +pd-net internal=true +## after ~15 s +NAME STATUS ROLES AGE VERSION INTERNAL-IP EXTERNAL-IP OS-IMAGE KERNEL-VERSION CONTAINER-RUNTIME +pd-k3s Ready control-plane,master 4s v1.33.6+k3s1 172.30.0.10 K3s v1.33.6+k3s1 7.0.14-20-pve containerd://2.1.5-k3s1.33 +No resources found +time="2026-10-09T08:42:48Z" level=error msg="Sending HTTP/1.1 503 response to 127.0.0.1:48848: runtime core not ready" +## 2026-10-09T08:44:46Z fetch the k3s image + airgap set (bench network, before isolation) +docker.io/rancher/k3s:v1.33.6-k3s1 +total 140296 +drwxr-xr-x 2 root root 4096 Oct 9 08:36 . +drwxr-xr-x 3 root root 4096 Oct 9 08:44 .. +-rw-r--r-- 1 root root 143649404 Oct 9 08:36 k3s-airgap-images-amd64.tar.zst +pd-net internal=true +pd-k3s Up 55 seconds +## after ~15 s +NAME STATUS ROLES AGE VERSION INTERNAL-IP EXTERNAL-IP OS-IMAGE KERNEL-VERSION CONTAINER-RUNTIME +pd-k3s Ready control-plane,master 49s v1.33.6+k3s1 172.30.0.10 K3s v1.33.6+k3s1 7.0.14-20-pve containerd://2.1.5-k3s1.33 +NAMESPACE NAME READY STATUS RESTARTS AGE +kube-system coredns-6d668d687-bvh4d 1/1 Running 0 43s +kube-system local-path-provisioner-869c44bfbd-rf27x 1/1 Running 0 43s +time="2026-10-09T08:44:50Z" level=error msg="Sending HTTP/1.1 503 response to 127.0.0.1:57484: runtime core not ready" diff --git a/documentation/audits/dooplex-survival-2026-10-09/partD-02-hubdb-restore.txt b/documentation/audits/dooplex-survival-2026-10-09/partD-02-hubdb-restore.txt new file mode 100644 index 00000000..5cb073de --- /dev/null +++ b/documentation/audits/dooplex-survival-2026-10-09/partD-02-hubdb-restore.txt @@ -0,0 +1,19 @@ +## 2026-10-09T08:43:12Z hub DB restore on DooPlex, read-only token (as felhom-hub-db-restore-test) +newest: host/dooplex-hub/2026-10-09T00:31:46Z +restore complete (380.047 MiB processed in 4.5s, average 84.169 MiB/s) +-rw------- 1 root root 398508032 Oct 9 02:31 /var/lib/felhom-hub-backup/partd/out/hub.db +integrity: ok +Error: in prepare, no such table: customers +hosts: 4 customers: +Error: in prepare, no such table: customers + +Error: in prepare, no such column: id + select id from hosts order by id + ^--- error here + +sealed console passwords: 4 +seal key file bytes: 64 (from Secret/offsite-secret-key, the value the operator saved 2026-10-05; not printed) +## (corrected queries: no customers table; customers live in customer_configs) +customer_configs: 4 +customer ids: Tester-2 demo-felhom demo-hp tester-1 +hosts: Tester-2-be8404→Tester-2 demo-felhom-8363b5→demo-felhom demo-hp-bb76ea→demo-hp tester-1-d70be4→tester-1 diff --git a/documentation/audits/dooplex-survival-2026-10-09/partD-03-transfer.txt b/documentation/audits/dooplex-survival-2026-10-09/partD-03-transfer.txt new file mode 100644 index 00000000..38abb319 --- /dev/null +++ b/documentation/audits/dooplex-survival-2026-10-09/partD-03-transfer.txt @@ -0,0 +1,5 @@ +## 2026-10-09T08:44:01Z transfer to bench 9401 +-rw------- 1 root root 26974208 Oct 9 08:44 /root/pd/hub-image.tar +95b67d69afdeddaf +DooPlex side: 95b67d69afdeddaf +64 diff --git a/documentation/audits/dooplex-survival-2026-10-09/partD-04-deploy.txt b/documentation/audits/dooplex-survival-2026-10-09/partD-04-deploy.txt new file mode 100644 index 00000000..f68ad785 --- /dev/null +++ b/documentation/audits/dooplex-survival-2026-10-09/partD-04-deploy.txt @@ -0,0 +1,49 @@ +## 2026-10-09T08:45:51Z system pods +coredns-6d668d687-bvh4d Running +local-path-provisioner-869c44bfbd-rf27x Running +## import the hub image (by file; k3s cannot pull) +Importing elapsed: 1.7 s total: 0.0 B (0.0 B/s) +helper image: docker.io/rancher/mirrored-library-busybox:1.36.1 +## secrets: the REAL seal key (step 4: same value); throwaway dummies for mail, report API, registry +secret/offsite-secret-key created +secret/resend-api created +secret/report-api created +secret/gitea-creds created +persistentvolumeclaim/hub-data created +configmap/hub-config created +deployment.apps/hub created +service/hub created +## first start on an EMPTY volume (after ~20 s) +hub-7f775576fd-fmddj 1/1 Running 0 16s +## step 5: scale to 0, copy the restored hub.db in with a helper pod, delete -wal/-shm, scale to 1 +deployment.apps/hub scaled +pod/hub-7f775576fd-fmddj condition met +pod/hubdb-copy created +pod/hubdb-copy condition met +total 412 +drwxrwxrwx 4 root root 4096 Oct 9 08:46 . +drwxr-xr-x 1 root root 4096 Oct 9 08:46 .. +drwxr-xr-x 2 root root 20480 Oct 9 08:46 assets +-rw-r--r-- 1 root root 389120 Oct 9 08:46 hub.db +drwx------ 2 root root 4096 Oct 9 08:46 snapshots +total 389200 +drwxrwxrwx 4 root root 4096 Oct 9 08:46 . +drwxr-xr-x 1 root root 4096 Oct 9 08:46 .. +drwxr-xr-x 2 root root 20480 Oct 9 08:46 assets +-rw------- 1 root root 398508032 Oct 9 08:46 hub.db +drwx------ 2 root root 4096 Oct 9 08:46 snapshots +95b67d69afdeddaf +bench copy: 95b67d69afdeddaf +pod "hubdb-copy" deleted +deployment.apps/hub scaled +## started on the restored DB (after ~20 s) +hub-7f775576fd-cp9nj 1/1 Running 0 16s +## start-up log (sealed / version / errors) +2026/10/09 10:46:52 [INFO] off-site secrets sealed at rest (0 legacy plaintext row(s) sealed now) +2026/10/09 10:46:52 [INFO] console passwords sealed at rest (0 legacy plaintext row(s) sealed now) +2026/10/09 10:46:52 [INFO] box secrets sealed at rest (0 legacy plaintext value(s) sealed now) +2026/10/09 10:46:52 [INFO] Default controller-version floor: 0.120.0 +2026/10/09 10:46:52 [INFO] Gitea artifact browser enabled (Day-0 version dropdowns) via http://gitea.gitea-system.svc.cluster.local:3000 +2026/10/09 10:46:52 [INFO] Registry version checker started (every 6h) +2026/10/09 10:46:52 [INFO] Listening on :8080 +2026/10/09 10:46:56 [WARN] Registry version check failed: HTTP request failed: Get "https://gitea.dooplex.hu/v2/admin/felhom-controller/tags/list": dial tcp: lookup gitea.dooplex.hu on 10.43.0.10:53: server misbehaving diff --git a/documentation/audits/dooplex-survival-2026-10-09/partD-05-checks.txt b/documentation/audits/dooplex-survival-2026-10-09/partD-05-checks.txt new file mode 100644 index 00000000..d4d4cadc --- /dev/null +++ b/documentation/audits/dooplex-survival-2026-10-09/partD-05-checks.txt @@ -0,0 +1,26 @@ +## LIVE hub (GET only; /configs + /hosts) 2026-10-09T08:48:00Z +customers: 4 Tester-2 demo-felhom demo-hp tester-1 +hosts: 4 Tester-2-be8404 demo-felhom-8363b5 demo-hp-bb76ea tester-1-d70be4 +## THROWAWAY hub on the bench k3s 2026-10-09T08:48:12Z (port-forward inside pd-k3s; same queries + a reveal per host) +customers: 4 Tester-2 demo-felhom demo-hp tester-1 +hosts: 4 Tester-2-be8404 demo-felhom-8363b5 demo-hp-bb76ea tester-1-d70be4 +reveal Tester-2-be8404: http 200, password 32 chars, sealed-form=False (value not printed) +reveal demo-felhom-8363b5: http 200, password 32 chars, sealed-form=False (value not printed) +reveal demo-hp-bb76ea: http 200, password 32 chars, sealed-form=False (value not printed) +reveal tester-1-d70be4: http 200, password 32 chars, sealed-form=False (value not printed) + +## the throwaway hub log for the reveals +2026/10/09 10:46:56 [WARN] Registry version check failed: HTTP request failed: Get "https://gitea.dooplex.hu/v2/admin/felhom-controller/tags/list": dial tcp: lookup gitea.dooplex.hu on 10.43.0.10:53: +2026/10/09 10:47:54 [ERROR] kernel notice (7.0.14-22-pve) to customer demo-felhom failed: sending request: Post "https://api.resend.com/emails": dial tcp: lookup api.resend.com on 10.43.0.10:53: serve +2026/10/09 10:47:58 [ERROR] kernel notice (7.0.14-22-pve) to customer demo-hp failed: sending request: Post "https://api.resend.com/emails": dial tcp: lookup api.resend.com on 10.43.0.10:53: server mi +2026/10/09 10:48:17 [INFO] operator revealed break-glass console credential for host Tester-2-be8404 (user=root@pam, secret 32 chars) +2026/10/09 10:48:17 [INFO] operator revealed break-glass console credential for host demo-felhom-8363b5 (user=root@pam, secret 32 chars) +2026/10/09 10:48:18 [INFO] operator revealed break-glass console credential for host demo-hp-bb76ea (user=root@pam, secret 32 chars) +2026/10/09 10:48:18 [INFO] operator revealed break-glass console credential for host tester-1-d70be4 (user=root@pam, secret 32 chars) +## hubdb-check on DooPlex (second channel), on a COPY of the restored DB 2026-10-09T08:48:30Z +hosts=4 console_passwords_opened=4 failed=0 absent=0 +rc=0 +control (a random key): +hosts=4 console_passwords_opened=0 failed=4 absent=0 +hubdb-check: FAILED: not every console password opened with this key +rc=1 diff --git a/documentation/audits/dooplex-survival-2026-10-09/partD-06-teardown.txt b/documentation/audits/dooplex-survival-2026-10-09/partD-06-teardown.txt new file mode 100644 index 00000000..3f3ece3e --- /dev/null +++ b/documentation/audits/dooplex-survival-2026-10-09/partD-06-teardown.txt @@ -0,0 +1,175 @@ +## 2026-10-09T08:48:46Z teardown — bench 9401 +pd-k3s volumes: c57fb08341e5b6fa44a389ddec08f866774ad7acde602870e2e574d044ab33d4 403bfbca4a80250b0e49a8131844b17e59552746513dac0d01f409f942a6e95d d8c1d37af50fe0473724a093c518ad86ec878c2ed96018c8a2ba908346f97daf 585175f25b48d2f86c1a782ee532cac86f099c94c83fe7e84cfd9da986251698 +pd-k3s +c57fb08341e5b6fa44a389ddec08f866774ad7acde602870e2e574d044ab33d4 +403bfbca4a80250b0e49a8131844b17e59552746513dac0d01f409f942a6e95d +d8c1d37af50fe0473724a093c518ad86ec878c2ed96018c8a2ba908346f97daf +585175f25b48d2f86c1a782ee532cac86f099c94c83fe7e84cfd9da986251698 +pd-net +image-removed +containers=0 volumes=8 networks=bridge host none traefik-public +.bashrc +.profile +.ssh +/dev/loop2 59G 3.8G 53G 7% / + are supported and installed on your system. + are supported and installed on your system. +status: stopped +## teardown — DooPlex +.cache +.kube +stage +-home-kisfenyo-git-homelab-manifests +-home-kisfenyo-git-jarr +-mnt-5-hdd-felhom-eu-build-wt-kept +-mnt-5-hdd-felhom-eu-drill-app-catalog-drill +-mnt-5-hdd-felhom-eu-git +-mnt-5-hdd-felhom-eu-git-app-catalog-felhom-eu +-mnt-5-hdd-felhom-eu-git-felhom-agent +-mnt-5-hdd-felhom-eu-git-felhom-controller +-mnt-5-hdd-felhom-eu-git-felhom-eu +-mnt-5-hdd-felhom-eu-git-wt-catalog +-mnt-5-hdd-felhom-eu-worktrees-ctrl-c-msgref +-mnt-5-hdd-felhom-eu-worktrees-ctrl-d4-sessions +-mnt-5-hdd-felhom-eu-worktrees-ctrl-integ +-mnt-5-hdd-felhom-eu-worktrees-hub-cf-mail-drop +ag.txt +au_tests.py +backups_degraded.old +backups_empty.old +backups_full.old +backups_interrupted_run.old +backups_nobackup_yet.old +backups_tier_due.old +bash-edit-diff +bb +bh.bak +bh.new +bookstackfix.py +cache-break-state-35d4e820-990d-465a-b5ee-c7b72649d29a.json +cache-break-state-55a70710-3ae0-4d71-b946-d9db725be672.json +cache-break-state-9a980473-7466-41a7-9c4c-d5f6c1785531.json +cache-break-state-c6c2aaef-766c-4184-8225-3610d20724ca.json +cache-break-state-e6fbe918-56ae-4fc3-af26-4a8a10574e30.json +cache-break-state-f6c29d80-39ec-4c54-960a-646ec49d5e71.json +cat_gates.txt +cfg.bak +cg.txt +cg2.txt +cg3.txt +cgi.txt +cgi2.txt +ci-1594.log +ci-ctl.log +ci.py +ci1357.log +ci1358.log +ci1361.log +ciwait.sh +claperfix.py +cmd_controller_main.go.bak +configs.go.bak +crg.bak +ctl-build.txt +ctlbuild.txt +ctrlgates.txt +deb.I8Wk +deb2.y8f3 +deb3.XEFL +docs45.py +dr.bak +drafts +dump_test.go +dumpgo +en.old +felhom-opsign +g.txt +g1.txt +g2.txt +g3.txt +gates-b.txt +gates.txt +gg.bak +gr.txt +h.bak +hb.new +hc.bak +hr.bak +hu.old +hubbak +hubbuild.txt +hubgate.txt +hubrepogates.txt +internal_backup_backup.go.bak +internal_i18n_locales_en.json.bak +internal_i18n_locales_hu.json.bak +isogate.5cI4 +m.go +main.bak +main.go.bak +mk.cur +mk.new +mk.old +mnt-hdd_1.mount +newrules.txt +o.bak +off.go +ok.bak +op.bak +opa.bak +osa.bak +pa.bak +partC-ids.json +partC.md +partd.py +partd2.py +partd3.py +promtest +r.bak +r262.bak +r705a.py +r706.py +rec.bak +reg3.py +rel147.txt +rg.bak +rg.txt +rgi.txt +rgi2.txt +rp +rp.bak +rp2 +rp2.bak +rr.txt +server.go.bak +service.go.bak +sg.bak +snap.bak +srv.bak +ss.bak +st.bak +states.py +stg.bak +sudoers.v0145 +sudotest +t.bak +t.go +t.txt +test_bb.py +uf.bak +unchecked-table.md +v279docs.py +vik.bak +w.bak +x.go +## 2026-10-09T08:49:18Z leftover volumes from the two failed pd-k3s starts (08:36, 08:42) +9bf479a381ed84a255645772d9424cc60a73d024bfc537485536b4e897831ac1 2026-10-09T08:36:57Z map[com.docker.volume.anonymous:] +88e8070b44ca866db658800c9c013b7d6095880b86033d2200d99c91ea97ce0e 2026-10-09T08:42:45Z map[com.docker.volume.anonymous:] +7119dcb4fa0e1ada468bd477c3bf7f00bda498d3cf8d984f8ff903d674bf5fb5 2026-10-09T08:42:45Z map[com.docker.volume.anonymous:] +54334eaba740d0cd68c33610c767f361b5882b2a1711ab6abbd6de22a246f416 2026-10-09T08:42:45Z map[com.docker.volume.anonymous:] +71831f601ff91a65952659dc60d05c5b7c6712158505d20c9665b3e0561bb5df 2026-10-09T08:36:58Z map[com.docker.volume.anonymous:] +a7d0645cf17a95b33018daeb558756bb01f54ab3a59d8553e09b4b8933649d78 2026-10-09T08:36:58Z map[com.docker.volume.anonymous:] +b922c6e0612a135afcdc7d9eb49a964682e4940b45b4fefe311b2258967e275d 2026-10-09T08:36:58Z map[com.docker.volume.anonymous:] +cfabb1005614015b82c12090ad7520c6ba877a4dcfafba317803dae13c0079ef 2026-10-09T08:42:45Z map[com.docker.volume.anonymous:] +volumes=0 containers=0 +status: stopped diff --git a/documentation/backlog/CLOSED-ITEMS.md b/documentation/backlog/CLOSED-ITEMS.md index 953a3174..d3e05520 100644 --- a/documentation/backlog/CLOSED-ITEMS.md +++ b/documentation/backlog/CLOSED-ITEMS.md @@ -26,6 +26,16 @@ --- +## 2026-10-09 — can the business survive losing DooPlex: Gitea + secrets off-site, Gitea and the hub restored into throwaways + +The full text of every row below: `git show 59f1ad2b86:documentation/backlog/OPEN-ITEMS.md`. + +| Row | What | Closed | Evidence | +|---|---|---|---| +| **R-173** | **The hub's SQLite PVC is excluded from every Longhorn backup job.** (P2) | CLOSED 2026-10-09 — the last steps of the hub's recovery runbook (§3 steps 4–5) proven on a throwaway single-node k3s (the bench, an INTERNAL Docker network, no route out): the newest ep0 copy restored with the read-only token, the live hub image (same digest) deployed, scaled to 0, `hub.db` copied into the PVC by a helper pod (-wal/-shm removed), scaled to 1 → `console passwords sealed at rest (0 legacy …)`; customers 4/4 and hosts 4/4 equal live; 4/4 console passwords revealed (length only) and, second channel, `hubdb-check` 4/4 with the saved key, 0/4 with a random one. Found: a restored hub at once mails customers their pending notices (operator mail off does NOT stop them) — runbook §3 corrected. Deleted: the k3s container, its volumes, every copy (shredded). | `audits/dooplex-survival-2026-10-09/partD-*.txt`, `runbooks/RUNBOOK-hub-db-offsite-backup.md` §3 | +| **R-861** | **The agent's sudoers lets the agent user reach root without the operator key.** (P2) | CLOSED 2026-10-09 — (a) proven: all three controller swaps of 0.304.0 (demo-hp, demo-felhom, Tester 1, ~05:05Z) ran `felhom-priv-apply controller-image 9201` (sudo log + the wrapper's `WROTE … 0.304.0` + the agent's "new controller healthy"), and the guests run 0.304.0 (Docker); the `tee` route 0 times that day. (b) B2 delivered + B3 accepted, (c) C2 accepted — `09` §3 decision 165. | `audits/dooplex-survival-2026-10-09/partE/R-861.txt` | +| **R-518** | **„Mentés most" stopped every app for ~8 min.** (P2) | CLOSED 2026-10-09 — the two-tier night under one-stop-per-tier read back on demo-hp (2026-10-08): local tier 20:20–20:25Z, off-site tier 20:34–20:37Z, each with its own stop under ~2 min (agent journal + Proxmox task log; controller metrics.db; ep0 listing `2026-10-08T20:34:27Z`, 7-day cadence held). A third, useless stop when the off-site tier answered BUSY → new row **R-921**. | `audits/dooplex-survival-2026-10-09/partE/R-518.txt` | + ## 2026-10-09 — the release day: D1–D4 proven live, the ep0-copy job installed The full text of every row below: `git show b9073e8fb6:documentation/backlog/OPEN-ITEMS.md`. diff --git a/documentation/backlog/OPEN-ITEMS.md b/documentation/backlog/OPEN-ITEMS.md index b66d409f..cd34c1d7 100644 --- a/documentation/backlog/OPEN-ITEMS.md +++ b/documentation/backlog/OPEN-ITEMS.md @@ -147,13 +147,12 @@ stopping line that lies. | **R-683** | App updates | P3 | **[P3-LOW] Watch: after a power cut during an update's health check, the hold named an HOUR-OLD second-drive copy, not the one the update's own backup should have just made.** 2026-09-24 chaos round 3 (nextcloud, `backup_max_age: 1m`): no `backing-up` phase was seen and the hold named Tier 2 at 13:04 for an update pressed at 14:04; the pre-cut controller log was lost with the container (the runner now saves it at arm time — R-320). Round 11, the same action without a power cut, named a fresh 14:34 copy and logged the Tier-2 copy. The sentence was TRUE (it named the copy it offered); the question is why the update did not back up first. Not reproduced; watch the next power-cut drill. `audits/night-2026-09-24/E/round-03*.json`, `E/round-11-controller-pre.log` | **OPEN — P3; owner: CC (watch)** | — | — | CC | | **R-785** | App updates | P3 | **[P3-LOW] SparkyFitness is pinned 11 releases and a major behind upstream (v0.17.3; upstream v1.7.3, v1.6.0 dated 2026-07-24).** READ 2026-10-01 (`audits/visitors-2026-10-01/C/bench/C1-previous-tag.txt`). **Needs:** an update walk 0.17 → 1.x through the ladder (bench + box), after R-784 is decided. | **OPEN — rank P3-LOW; owner: CC (after R-784)** | — | — | CC | -## Backup & restore — 30 rows (P2 5, P3 11, P4 14) +## Backup & restore — 29 rows (P2 3, P3 12, P4 14) | ID | Category | Sev | What | State | Blocked on | Next action | Owner | |---|---|---|---|---|---|---|---| -| **R-232** | Backup & restore | P2 | **DooPlex's backup makes every copy inside the same box — and nothing tells anyone when it fails.** Surveyed read-only 2026-08-06 (`audits/RECON-dooplex-backup-2026-08-06.md`). **What works:** five sets, 14/14 successful runs in 14 days; a file was restored from the `data` repo and matched the live original **byte for byte**; every set except two is cross-disk; k3s is integrity-checked on every run. **What the matrix exposes, ranked:** (a) **`notify_failure` is a no-op** — `NOTIFY_ON_FAILURE=true` but `NOTIFY_WEBHOOK_URL` is commented out, so a failed backup notifies **nobody**; the project already has a working Resend path that CI uses. Cheapest item, and it makes every other failure visible. (b) **Nothing leaves the box** — no rclone, no remote repo, no off-site target anywhere; Longhorn's target is `nfs://192.168.0.180:` pointing at DooPlex itself, and the only outbound-looking cron pulls *inbound* from Hetzner for a different project. The machine that runs the hub managing the customers' off-site chain has no off-site copy of its own. (c) **The backup tree is a single writable path** and the restic repos are not append-only — one bad script or ransomware destroys every copy at once. (d) **Two same-disk sets**: `.claude-memory` and the PostgreSQL dumps, whose source directory sits *inside* the backup tree. (e) **Longhorn `retain=1`** — one generation per volume, so a corruption noticed a day late has no earlier copy. (f) **`/opt/backup/docs/BACKUP-RESTORE.md` does not exist** though the systemd unit advertises it. (g) **`secrets/restic-repo` has never held a snapshot** — `backup-secrets.sh` contains no `restic` call; the secrets are GPG files on `sda1` only. (h) **No restore has ever been run** beyond today's single-file probe — the matrix's "ever demonstrated?" column is otherwise entirely empty. **Not a finding:** the restic passphrase. The on-box copy is on `sdb1`, a different disk from the backups, and the **operator holds an offline copy out of band** — so a disk loss is recoverable. The narrow residual is that it is operator-held rather than system-held, unlike the customer case's hub-vaulted escrow, so it should be confirmed current and findable by someone else. **Nothing was changed by the recon.** | **NARROWED 2026-10-08 — owner Viktor.** (a) DONE with the operator's yes in chat: `notify_failure` now mails admin@felhom.eu through Resend; proven by one test mail that reached the inbox (`audits/day-2026-10-08/r232/`; no backup was started). (b) partly: the hub database leaves DooPlex nightly to ep0 (R-173); everything else stays on the box. (c)–(h) unchanged. **READY** for the rest | — | — | operator | +| **R-232** | Backup & restore | P2 | **DooPlex's backup makes every copy inside the same box — and nothing tells anyone when it fails.** Surveyed read-only 2026-08-06 (`audits/RECON-dooplex-backup-2026-08-06.md`). **What works:** five sets, 14/14 successful runs in 14 days; a file was restored from the `data` repo and matched the live original **byte for byte**; every set except two is cross-disk; k3s is integrity-checked on every run. **What the matrix exposes, ranked:** (a) **`notify_failure` is a no-op** — `NOTIFY_ON_FAILURE=true` but `NOTIFY_WEBHOOK_URL` is commented out, so a failed backup notifies **nobody**; the project already has a working Resend path that CI uses. Cheapest item, and it makes every other failure visible. (b) **Nothing leaves the box** — no rclone, no remote repo, no off-site target anywhere; Longhorn's target is `nfs://192.168.0.180:` pointing at DooPlex itself, and the only outbound-looking cron pulls *inbound* from Hetzner for a different project. The machine that runs the hub managing the customers' off-site chain has no off-site copy of its own. (c) **The backup tree is a single writable path** and the restic repos are not append-only — one bad script or ransomware destroys every copy at once. (d) **Two same-disk sets**: `.claude-memory` and the PostgreSQL dumps, whose source directory sits *inside* the backup tree. (e) **Longhorn `retain=1`** — one generation per volume, so a corruption noticed a day late has no earlier copy. (f) **`/opt/backup/docs/BACKUP-RESTORE.md` does not exist** though the systemd unit advertises it. (g) **`secrets/restic-repo` has never held a snapshot** — `backup-secrets.sh` contains no `restic` call; the secrets are GPG files on `sda1` only. (h) **No restore has ever been run** beyond today's single-file probe — the matrix's "ever demonstrated?" column is otherwise entirely empty. **Not a finding:** the restic passphrase. The on-box copy is on `sdb1`, a different disk from the backups, and the **operator holds an offline copy out of band** — so a disk loss is recoverable. The narrow residual is that it is operator-held rather than system-held, unlike the customer case's hub-vaulted escrow, so it should be confirmed current and findable by someone else. **Nothing was changed by the recon.** | **NARROWED 2026-10-08 — owner Viktor.** (a) DONE with the operator's yes in chat: `notify_failure` now mails admin@felhom.eu through Resend; proven by one test mail that reached the inbox (`audits/day-2026-10-08/r232/`; no backup was started). (b) partly: the hub database leaves DooPlex nightly to ep0 (R-173); everything else stays on the box. (c)–(h) unchanged. **2026-10-09 (operator yes in chat): (b) and (h) DONE for Gitea and the secrets.** Nightly 00:20 `felhom-dooplex-offsite` pushes Gitea (614 MB of repositories, the `gitea` database dump taken first, `app.ini`) and the nightly GPG secrets export, encrypted with a NEW key, to ep0 `operator` (`host/dooplex-gitea`) on the hub-DB write-only token — no change on ep0; first push 57 s; restore test weekly (manifest, `git fsck` every repo, `pg_restore --list`), run once: 27 805 files, 10 repos OK; `DooplexGiteaOffsiteStale`/`DooplexGiteaRestoreTestStale` in force (`absent()` included, promtool + 2 red-proofs); every backup unit now mails on failure (`OnFailure=`, dry run reached the inbox). **First real Gitea restore:** into a throwaway on the bench, 10/10 repos, product `main` = live, a file byte for byte, a login (`runbooks/gitea-restore.md`). The hub came back too (R-173, closed). **The registry is left out on purpose:** its images rebuild from the code. **Correction to the recon §7:** the API mirror holds ONE repo (`homelab-manifests`), not all — Gitea had a single copy, on DooPlex. `audits/dooplex-survival-2026-10-09/` **LEFT:** (c) append-only for the on-box tree, (d)–(g) as written; the paper copy of the new key (operator). **READY** for the rest | — | — | operator | | **R-304** | Backup & restore | P2 | **The retained escrow key works, and the customer is told their correct code is wrong.** DRILL 2026-08-12 answered the three questions separately, on `demo-felhom`, with planted data. **(a) retention: WORKS** — the first retained row in fleet history to carry material (`host_escrow_superseded` id 11, `identity_blob` 572 B), byte-identical (`sha256 a10032341c8584ed…`) to the pre-supersession `host_escrow` row. **(b) the material opens the old store: YES** — unsealed with the OLD recovery code it yielded a password byte-identical to the pre-change one (`sha c60c8bc737a6b7c6…`), and restored three planted files **byte-identical** from a store the box itself could no longer open (negative control first: `Fatal: wrong password or no key found`), **including a Hungarian accented filename verified as raw bytes**. **(c) the customer's route: DOES NOT EXIST, and misinforms.** `ListSupersededEscrow` (`store.go:2841`) is the only reader of a retained `identity_blob` and has **zero production callers** — five call sites, all `_test.go`; the product path (`POST /escrow/recover-offsite-password` → `FetchIdentityEscrow` → `GetHostDRBundle`, `store.go:3152`) selects `FROM host_escrow` — the CURRENT row only. Asked for the old password with the code that demonstrably opens the retained row, the product answered **"the recovery code did not open the sealed bundle — nothing was written"**. **This is the R-224 class again**: there an unreachable hub was reported as a bad code; here a VALID code for retained history is reported as a bad code, and the customer's attempt ends there. **Consequence:** the census answer stands (it was about retention); the countdown banner's promise is true in substance and false in practice; **any capability-map claim that the customer can recover the old history with their recovery code is false today and must move** | **READY — the operator mail is BUILT on main 2026-10-08 (decision 183; controller `a40729a` + hub, ships with the next releases).** The honesty fix is on main too (agent `91b9405`, controller `75b3b39`). Left: the runbook „open a retained package for a household" (design option C, second slice). Design `audits/day-2026-10-08/design-R-304.md`. | R-198, R-199, R-224, R-241 | Release controller + hub; write the retained-package runbook from the 2026-08-12 drill §4; then close | CC | -| **R-518** | Backup & restore | P2 | **[P2-MEDIUM] „Mentés most" on the whole-system backup stops every app for about eight minutes while the page promises „csak néhány másodpercre".** MEASURED 2026-09-14 (BIGNIGHT, VM 333, 12 apps): the button's call quiesced all 12 stacks at 19:03:23Z (first stopped 19:03:27Z); the local vzdump ran 19:03:49 → 19:09:59Z; the controller then kept the apps stopped for the second (PBS) tier and restarted them at 19:10:09Z after it failed, the last started 19:11:12Z (`phase4/guest-backup-quiesce-log.txt`) — **≈ 7 m 45 s** with every app answering 404. The page under the button: „Pillanatkép-mód: az alkalmazások csak néhány másodpercre állnak le." A household pressing it at dinner loses every app for the length of the dump, and longer on a bigger box. **Fix shape:** state the real expected downtime (it scales with data), or quiesce per tier and not across a second tier's attempt; do not start a tier whose storage is absent (see R-517). **NARROWED 2026-09-15 (controller v0.243.0 + agent v0.131.0):** a tier whose storage the agent reports absent is skipped before anything stops (`backup_tier_skipped`, once per absence; unknown never skipped), and the button copy now says „általában néhány perc, nagyobb adatnál több". Unit-proven with red-proofs. **Still open:** quiesce per tier, so a slow second tier does not keep every app down. **— NIGHT 2026-09-23 (controller v0.267.0):** the copy half is DONE: the page and the confirm now state the measured stop (≈ 8 minutes on a 12-app box), both languages, red-proofed (`audits/night-2026-09-23/A5-*`). The brief's „csak néhány másodpercre" had already gone in v0.243.0. **Still open:** quiesce per tier, so a slow second tier does not keep every app down. | **READY — P2, narrowed to per-tier quiesce; owner: CC (controller). 2026-10-05: the copy now states today's measurement too (demo-hp, 9 apps, local tier only: 5 min 47 s) — controller v0.296.0, `audits/hub-safety-2026-10-05/partE/`.** **2026-10-05 (burn-down night): a one-page design proposal (no code) is in `audits/night-burndown-2026-10-05/design-R-518.md`** — for the operator. **2026-10-06: BUILT — controller v0.301.0, `09` §3 decision 156 (reverses R-82's one window).** One stop per tier; the button makes the local copy only. Measured first, read-only: demo-felhom's night off-site job reached `snapshotted` 2 s after it started, the app back 8 s later (the off-site part of a stop is seconds). **Not shown live:** a press under the new rule — scratch 9202 has no agent connection and the demo boxes take deliveries only. Red tests and the build: `audits/design-build-2026-10-06/`D/. **Risk noted, unmeasured:** after a local copy the agent runs its OS step, and the off-site tier then answered BUSY (2026-10-05) — under the new rule that costs one short stop with no copy before the 15-min backoff. **2026-10-06 (night), from Part C:** demo-hp's off-site tier was NOT overdue — its last copy is 2026-10-01 20:15Z (ep0's listing, verify ok), so with the 7-day cadence it is due ~2026-10-08; the night of 2026-10-06→07 is most likely local-only on both demo boxes (demo-felhom's off-site landed 2026-10-06 04:21Z). The two-tier night under the new rule is then ~2026-10-08 on demo-hp. **2026-10-07 (morning): the local-tier night and one press READ BACK** (`audits/readback-2026-10-07/RESULT-B-D.md`): the night stop on demo-hp (9 apps) was ~91 s (was 5 min 47 s), demo-felhom (1 app) ~11 s; one press on demo-hp: 80 s from press to the last app (per app 39–79 s), the copy finished 4 min later with the apps running, only the local tier ran; the page's „kb. 1–1,5 perc" holds. Two channels each (controller log + agent journal / container StartedAt + a 5-s HTTP sampler). **Left:** the first night with both tiers due on demo-hp (~2026-10-08). | — | Read back the ~2026-10-08 night (both tiers on demo-hp); then close | CC | | **R-893** | Backup & restore | P3 | **After a failed OFF-SITE replay, the rollback pours the NEWER pre-restore copy over the OLDER volume just put back.** Read in source 2026-10-06 (R-638 option A, not measured): `internal/backup/offbox_reconstitute.go` writes the undo copy from the live (newer) database, replaces the volumes with the snapshot's older tars, then — when the replay fails — `rollbackSafetyDump` loads that newer dump over the older database volume. The loader only drops what the dump knows, so tables the newer migration removed stay; and when the snapshot's older definition was written, the rollback branch does not put the newer definition back, so the older app starts on rolled-back data; non-database volumes stay at the snapshot's state. An order change cannot fix it (the only undo is a logical dump, and its volume was replaced). Known limit in `07` §6.3. **-- 2026-10-08 design:** `audits/day-2026-10-08/design-R-893.md` — re-verified; and the screen's „your data is back as it was" (`err.backup.db_restore_failed_rolled_back`) is false for files, other volumes and the app version. Question D8 on STATUS's decision sheet. **-- 2026-10-08 14:16 operator ruling D8 (`09` §3 decision 192):** yes, both, in that order — first the app stays stopped for support (option C), next „put back exactly as it was" (option A). | **NARROWED — 2026-10-09: the first half DELIVERED in controller 0.304.0. Slice 0 (the live measurement) NOT run: 9202 has no off-site, and a dump cannot be cut by hand inside an encrypted off-site copy on Tester 1; it needs a scratch off-site repo — a plan, not a finding.** **NARROWED 2026-10-08 — the first half (D8 option C, the app held stopped) is built on controller main and ships tomorrow; what stays open is the second half, „put back exactly as it was” (option A: a pre-restore copy of the app's volumes and placed files, a fit check, disk for one copy; the slice-0 measurement on 9202 first).** **OPEN — filed 2026-10-06** **2026-10-06 night: verified in source, no code** — `offbox_reconstitute.go:758` (undo dump from the live DB), `:778-860` (files and volumes from the snapshot), `:804` (the snapshot's definition is written when its version differs), `:898` (the rollback loads the newer dump over the older volume; nothing writes the live definition back). Not a reorder fix: it needs R-638 option B (a rebuilding loader) or a pre-restore volume copy (disk cost; R-685's class). Which state a household gets after a failed off-site replay is the operator's call. Next: the 9202 measurement, then the design. | a design: R-638 option B (a loader that rebuilds instead of overlays) or a pre-restore volume copy | Measure it once on 9202 (a forced replay failure after an off-site restore over a migrated app); then a design for the operator | CC | | **R-895** | Backup & restore | P2 | **The hub's clean-up-window check trusts the snapshot counts the box sends, so a broken-into box (or past-dated fakes added through the add-only key) can shrink the real off-site history without an alarm.** READ 2026-10-06 night in source (R-822's design): the before/after comparison uses counts the box itself reports (`hub/internal/offsitekeys/service.go:284`, `:343`); new fakes keep the count level. Decision 68 already accepts a box-trusted count. | **OPEN — filed 2026-10-06 night** **2026-10-07 07:58: kept open for later (`09` §3 decision 166).** | a design + one read-only measurement (does the Storage Box shell on port 23 show snapshot file upload times?) | Option B of `audits/night-burndown-2026-10-06/design-R-822.md`: the hub lists the repo's `snapshots/` files over its own login before and after a window and alarms on snapshots no box run explains | CC | | **R-49** | Backup & restore | P3 | **[P2] The offsite capture set is ~90% cache and duplication — 1.1 GB of a 1.2 GB immich "photo backup".** Measured 2026-07-19: `immich_ml_cache.tar` **823 660 032 B (~60%)** — re-downloadable ML model weights; `immich_postgres_data.tar` **308 251 136 B (~23%)** — a raw tar of the postgres data dir that DUPLICATES the logical `.sql` dump captured beside it; `upload/backups/` **18 MB** — immich's own nightly dump, a backup inside the backup, growing daily; plus the stranded pre-v3 `dccc13fe…` tree (~36 MB) no DB has ever referenced. Actual irreplaceable content: **72 MB of originals**. **MIGRATED FROM `ROADMAP.md` 2026-08-22 (R-369) — originally filed 2026-07-19, size S–M, roadmap state `idea`.** Moved verbatim; nothing added or reinterpreted. The roadmap keeps its copy as history, marked moved. | **OPEN — migrated from ROADMAP 2026-08-22, rank unchanged** **Re-ranked 2026-10-03: P2→P3: wasted space and transfer, no data risk; needs a capture-set ruling, not a sale blocker.** | — | **Evidence: `audits/DIAG-immich-restore-round2-2026-07-19.md` §4 (full byte breakdown).** This is the customer's offsite quota and transfer cost, and it lands on the Hetzner sub-account they are billed for. **Recorded, deliberately not changed** — a capture-set exclusion is a data-loss-shaped decision and gets its own ruling, not a drive-by edit. Candidates in priority order: (a) `immich_ml_cache` — pure cache, strongest case; (b) the `postgres_data` volume tar where a logical dump of the same DB is already captured (the dump is what the restore path actually replays); (c) `upload/backups/`. Likely generalises past immich into a template-classification rule about cache volumes and self-backup directories, so it should be specified against the catalog, not one app | CC | @@ -166,6 +165,7 @@ stopping line that lies. | **R-540** | Backup & restore | P3 | **[P3-LOW] The hub knows exactly ONE off-site pool box, so there is no rule for what happens when it fills.** Read from source 2026-09-16 while making off-site the default: `HETZNER_POOL_BOX_ID` is a single value, and every shared customer becomes a sub-account on that box. With off-site now ON for every new customer (hub v0.116.0) the box fills faster, and the fill warning (80%/90% of the box, `monitor/offsite.go`) tells the operator it is filling but nothing says which box a new customer should land on. **Needs a selection rule** (least-full, or explicit per-customer), not a bigger box. No customer is at risk today: the pool box read 0.3% full (2.7 GB of 1 TB), Σ shared quota 150 GB, oversub 0.15x. | **READY — rank P3-LOW; owner: CC (hub)** **2026-10-05 (burn-down night): NEEDS A DESIGN (and money).** A selection rule needs a multi-box configuration and a second box. Next: an operator pick of the rule and its trigger (e.g. add box 2 at 70 %). | — | — | CC | | **R-548** | Backup & restore | P3 | **[P3-LOW] The whole-guest backup’s LOCAL tier cannot fit on a small-system-disk box, and will retry on that tier for ever.** MEASURED 2026-09-17 (chaos night) on `tester-1-022354`: a whole-guest backup wrote a **~29 GB source** (`mp0` = `local-lvm:vm-9201-disk-1`, 70 G provisioned, 40.58 % used, `backup=1`) into `pve-root`, which on a 32 GB system disk is **14 GB total with ~4.9 GB free**. Two samples thirty seconds apart showed the archive growing ~497 MB while free space fell ~475 MB — **~16 MB/s, i.e. under four minutes to a full `/`** on the nested PVE. **The product’s behaviour is correct and legible throughout:** it failed the tier and said which one — `whole_guest_backup_failed` (error, operator-only): „Whole-guest backup FAILED on the local tier — retrying with backoff (next attempt in 15m0s)” — its status surface agreed (`target_id:"local"`, `success:false`, `size_bytes:0`), and the **off-site tier then ran from the same snapshot and succeeded in ~8½ minutes, encrypted to ep0, consuming no local disk at all**. So the data still left the house. What is filed is the loop: on a box shaped like this the local tier can **never** succeed, and it keeps retrying on a backoff for ever, burning I/O and risking `/` each time. **Honest caveat:** the 32 GB system disk is this drill’s fixture choice, so the row is conditional on disk size — but nothing in the product checks whether the local target could ever hold the source before trying. **Fix shape:** compare the source size against the target’s free space before starting the local tier and skip it with a clear reason, rather than discovering it at ~16 MB/s. Evidence: `audits/evidence-chaos-night-2026-09-17/round-6.txt`. **-- 2026-09-30, the class on a real box: demo-hp's whole-guest LOCAL tier was refused by the space preflight (R-685) 10 times since 2026-09-27** („local has 12.1 GiB free; the last archive of guest 9201 was 8.9 GiB, so a new one needs about 12.1 GiB") — its newest local archive was 2026-09-29 04:42, 30 h old at the day's read (the weekly PBS tier, last 2026-09-24, is its whole copy meanwhile). The refusal is correct and named; what is missing is that nothing gives the box the room back (the host's `local` holds the 8.9 GiB archive of the only guest plus templates on a 39 GB root). `audits/pg-last-six-2026-09-30/C/`. | **READY — rank P3-LOW; owner: CC** | — | — | CC | | **R-698** | Backup & restore | P3 | **[P3-LOW] A backup stores the image's NAME, not the image — a restore of a version its maker has deleted cannot start.** `RecoveryManifest.image_pins` ("image NOT stored — re-pulled on restore"); since controller v0.275.0 each data file also records its running `ref@digest`, and a restore brings the data back AT ITS OWN VERSION (`07` §6.6) — so a restore asks for exactly the old image. **Measured 2026-09-26** (`audits/version-travel-2026-09-26/A7/`, registry HEADs, no pulls): the catalog's 42 ladder `ref@digest` pairs all resolve (200); an invented digest answers 404 on Docker Hub and ghcr.io (negative control). Not measured: the digests recorded on boxes (older than any ladder entry), how often makers delete versions, the catalog's 66 digest-less compose lines. **Options (decide nothing yet):** (a) keep — a restore of a deleted version fails at the pull and the household uses the next copy or a newer version; (b) mirror every INSTALLED image into the DooPlex registry, restore falls back to it — storage + bandwidth on DooPlex, a new part on the recovery path; (c) mirror only ladder-named versions — bounded, misses pre-ladder boxes; (d) `docker save` into the unit — hundreds of MB per app per copy on every tier. **-- 2026-09-30 late (decision 53):** a box now keeps only an app's running and previous image; a restore to an older version re-pulls it — as every restore already did. The limit above is unchanged. | **OPEN — P3; owner: operator (a decision), CC measures** **Operator ruling 2026-10-05 18:23: kept OPEN as a known risk to a household's restore; owner the operator; not worked on in the burn-down.** | — | — | operator | +| **R-921** | Backup & restore | P3 | **When the off-site tier answers BUSY, the household still gets a short stop of every app — with no copy.** MEASURED 2026-10-08 night on demo-hp (read-back 2026-10-09, `audits/dooplex-survival-2026-10-09/partE/R-518.txt`): local tier 20:20–20:25Z (own stop, < 2 min), then at 20:26:14Z the agent refused the off-site request (`backup refused — a heavy operation is already in flight`, `busy=backup:local` — the night OS step ran 20:27–20:29Z), yet the controller's metrics show all 15 app containers down at 20:26:19Z and back at 20:27:19Z; the retry ran the off-site tier at 20:34Z with its own short stop. So a two-tier night costs THREE stops, one of them for nothing. This was R-518's noted "unmeasured" risk. Fix shape: ask the agent whether it is free (or reserve the slot) BEFORE quiescing, and resume at once on BUSY. | **READY** — needs a controller (and maybe agent) release; this session had no release budget. | — | Build the pre-quiesce check; read back a two-tier night | CC | | **R-91** | Backup & restore | P4 | Old 13 GB datastore copy at `/srv/pbs-felhom` on ep0's root disk **Checked from source 2026-10-05 (burn-down round 2):** Gate is long past (row waits on demo-felhom's first post-migration PBS backup, migration 2026-07-27). Last positive record of the copy: audits/CAMPAIGN-9-restore-proof-2026-07-28.md:759 'ep0 : /srv/pbs-felhom rollback copy intact (13G)'; CONTEXT.md:3666 still says it is 13 G of dead weight awaiting R-91. No later record of deletion found (grep srv/pbs-felhom across felhom.eu). Deleting is on ep0 (protected) and needs an operator word. | WATCHING | demo-felhom's first **post-migration** PBS backup | Delete once it lands; fix `CONTEXT.md:1018` same commit | CC | | **R-164** | Backup & restore | P4 | **C2's chain: the DB volume tar cannot be dropped until a SOUND dump predicate exists.** The unit carries both a volume tar and a SQL dump; the restore uses **both** — the dump is authoritative and replayed *after* the tar so it WINS (F17), with only the DB service up (R-47) — `internal/backup/restore_unit.go:262-266`. Dropping the DB container's tar would halve DB-app units **and** close the R-127(b) initdb-skip password trap (restored PGDATA ⇒ `POSTGRES_PASSWORD` ignored). | **BLOCKED** — on the predicate | a dump-validity predicate that is not `accounts has rows` | **The obvious gate is DEAD, measured:** `ValidateDump` warns when the `accounts` table is empty, and that warning was **correct** — the live DB genuinely had 0 accounts, and seeding one stopped the warning and put the row in the dump. But **a fresh appliance legitimately has zero accounts**, so promoting that predicate to a gate would **block every new customer's first backup**. Order: (1) a sound predicate — dump vs **live** per-table counts, not an absolute expectation; (2) warn→gate; (3) tar-drop. **Until (1), the tar is load-bearing** — not because dumps are bad, but because nothing can yet prove one is good. Pairs with **R-127** | CC | | **R-213** | Backup & restore | P4 | **Putting files back in place — the half the recovery screen deliberately does not do.** The screen (R-193, controller v0.200.0) unlocks the repository and LISTS what is in it; restoring is per-app and lives in the backups area, and the operator ruled the two separate on 2026-08-05: *a screen that unlocks and then offers to overwrite is two decisions wearing one button*. What is missing is the step after the listing — a customer who can now SEE their files still has to work out, per app, which restore to choose. **The operator named its requirement: a live-versus-backup comparison** — the customer must be able to see what would change before anything is overwritten | **OPEN — not started, deliberately** | the comparison design (nothing exists for it yet) | Design the live-vs-backup comparison, then the put-back flow on top of it. Do NOT fold it into the recovery screen | Operator + CC | @@ -191,11 +191,10 @@ stopping line that lies. | **R-331** | Storage & devices | P4 | **Disk health Phase 3 — growth-rate detection, and retiring the static 64.** The v0.215.0 count backstop (64 unreadable sectors → Hiba) is **a judgement from ONE drive**: the observed benign excursion peaked at 16 and cleared inside an hour, and the terminal run passed 64 at 13 Aug 11:28 and never came back. It is deliberately a backstop BEHIND the sustain rule, not the primary signal, but it is still a magic number tuned on a single sample and it will be wrong for some drive. With Phase 2's history the box can ask the question that actually matters — *is this count climbing, and how fast* — which distinguishes a drive with eight stable aging sectors from one adding forty a day, something no static threshold can do. Revisit 64 when that exists | **READY (M) — NEW 2026-08-14** | R-330 | Growth-rate rule over persisted samples; re-derive or delete the static 64 | CC | | **R-352** | Storage & devices | P4 | **Four screens state something untrue about where an app's data goes, and the configured default is consulted by nothing that places data.** Measured on `demo-hp` 2026-08-21. **(1)** **40 of 53** catalogue templates declare no data path (`grep -rl 'env_var: HDD_PATH' --include='.felhom.yml'` → 13; total 53); those apps get **no storage field and no default** — their data lands in a named Docker volume on the system drive. **(2)** `GetDefaultStoragePath()` has exactly **three** non-test callers — the metrics collector (`cmd/controller/main.go:410`), the dashboard SystemInfo panel (`web/server.go:733`) and `.fab` import landing (`handler_export_upload.go:154`). **The deploy route never reads it.** Its field comment `// new apps use this by default` (`internal/settings/settings.go:453`) has never been true — an invariant with no test pinning it. **(3)** The first-tier backup follows the data onto the same disk (`backup/backup.go:324-334` → `systemDataPath`), so data and nearest copy share one device for a customer doing nothing wrong — the posture Tier 2 refuses outright at `tier2.go:329`. **(4)** „1 alkalmazás használja" on the Drives page counts only `Env["HDD_PATH"] == path` (`web/handlers.go:2118`), so it can never include the 40-class; it truthfully means *„1 of the apps that CAN use a drive does"*. | **NARROWED** — **PARTLY CLOSED 2026-08-21** — visibility shipped; **placement OPEN** | — | **Shipped tonight (visibility only, no placement change, nothing migrated):** the deploy page now states where the app's data will live before the button is pressed, naming the system drive for the 40-class and the selected drive for the 13. **⚠ RE-FRAMED 2026-08-22 — THE FOUR MEASUREMENTS STAND; TWO OF THE CONCLUSIONS DRAWN FROM THEM DO NOT.** (1) is a measurement and is correct, but "those apps get no storage field and no default" is not a deprivation: the architecture places **hot** data (DB/config/cache) on fast storage inside the guest and states that placement is **ENFORCED** (`documentation/architecture/01-topology-and-trust.md:150-152`). The 40 are all-hot apps; the 13 are the ones with **bulk** content, which belongs on an attached drive. There is no choice being denied. (3) **overstated one risk and understated a distinction.** Since R-165 the guest carries a small OS rootfs plus **ONE** data volume at `/var/lib/felhom`; `/var/lib/docker` and `/mnt/sys_drive` are two **binds of that same volume** (`felhom-agent/configs/build-golden.sh:29-40, 99`) — the `mp0`/`mp1` split assumed here was retired 2026-08-03. **Real risk:** a physical-disk failure loses the data and its first-tier copy together — which is what the off-site and whole-machine tiers exist for, and which is equally true of a drive-resident app whose unit sits beside its data by design. **Overstated risk:** a full data volume stopping the operating system — the OS rootfs is a separate volume and the capture floor refuses per app before exhaustion (`00-capability-map.md:94`), watched working 2026-08-21 with the volume at 99% and all 15 containers healthy. **The comparison to Tier 2's same-disk refusal (`tier2.go:329`) is withdrawn:** Tier 2 refuses a SECOND copy on the same disk; Tier 1's unit is meant to sit beside the data. **(2) and (4) are untouched and remain correct** — (2) is now filed on its own as **R-368** with its scope measured, and (4) needs no ruling: the count is honest and only easy to misread. **The specification for the rest is filed at `documentation/backlog/SPEC-app-data-placement-2026-08-21.md`** (corrected 2026-08-22, framing marked inline, measurements kept) and lists the five points a ruling must settle (compose-template vs controller, existing deployments, when the SSD is legitimately right, `IsDefault` must become true or go away *with a test*, and the Drives-page count). **An earlier recommendation to refuse deployment until a drive is registered was WITHDRAWN** — it assumed the customer had failed to choose; they had no choice to make. | **Viktor rules**, CC executes | -## Security & access — 15 rows (P2 1, P3 11, P4 3) +## Security & access — 14 rows (P3 10, P4 4) | ID | Category | Sev | What | State | Blocked on | Next action | Owner | |---|---|---|---|---|---|---|---| -| **R-861** | Security & access | P2 | **The agent's sudoers lets the agent user reach root without the operator key, so "root-minimized" (`03` §3) overstates it and the root-owned trust files (decision 93, the bundle's R17) are defence in depth, not a boundary.** READ 2026-10-04 from `felhom-agent/configs/felhom-agent.sudoers` (not exploited): `FELHOM_GUESTHOOK` installs `/tmp/felhom-guest-hook-*.sh` as a hookscript Proxmox runs as root at guest start, and `pct reboot` is granted; `FELHOM_INTERMEDIARY` installs a script + a systemd unit that run as root at boot; `FELHOM_ESCROW` runs `/usr/local/bin/felhom-agent` as root, and `FELHOM_SELFUPDATE apply` accepts a sha the agent itself passes. A compromised agent PROCESS is therefore root on its host. Fix direction: each of the four becomes a root-owned wrapper that checks its own input (fixed content or a signature), like `felhom-os-apply`; delivered by the config bundle. `11` §5.4.2, `03` §11. | **NARROWED 2026-10-05 — FIXED agent v0.146.1 for every root path found (nine, not four), delivered to demo-hp, demo-felhom and Tester 1 by a step bundle (R-880); live on both demo boxes: `sudo -l` 93/93 (64 commands allowed, 29 attacks refused — 23 of them allowed before), capability probe 67/67, a staged unit over /etc/sudoers.d refused. Design `03` §3.1, decision 122. LEFT, each named there: (a) the controller-swap image ref is guest-scoped (a compromised agent can run a chosen pinned-registry image in the guest); (b) the felhom-op SSH key is hub-delivered, not signed (felhom-op's sudo is scoped, not root); (c) the escrow ceremony hands the agent R by design (the box's PBS key). Tester 2: not delivered (offline).** **2026-10-06 night: design written** (`audits/night-burndown-2026-10-06/design-R-861.md`). Correction: (a) is not "pinned-registry" — the `tee` content is unchecked by sudo, so any image from any registry runs in the guest with the docker socket (`03` §3.1 corrected). Pick: (a) close before the first paying customer (a `felhom-priv-apply controller-image` verb; ~1–2 h, rides the bundle); (b) and (c) accept for the first customers. Waits for the operator. **2026-10-07 07:58: `09` §3 decision 165 — (a) A1 yes before the first paying customer; (b) B3 accept + B2 hygiene in the same bundle; (c) C2 accept.** **2026-10-07: (a) A1 and (b) B2 DELIVERED (agent 0.151.0 + bundle on demo-hp, demo-felhom, Tester 1; probe 68/68).** Live on demo-hp: no `tee` grant left in `sudo -l -U felhom-agent`; the verb `felhom-priv-apply ^controller-image [0-9]+$` is the route; a hand-fed `docker.io/library/alpine:latest` → `REFUSED [I1]` rc 3, the guest's image file unchanged; the old `pct exec … tee` asks for a password; felhom-op's pct lines anchored (`audits/day-2026-10-07/C/C-live-demo-hp.txt`). **LEFT:** one managed controller swap seen through the verb — no newer controller existed today; the next controller release shows it. (c) accepted (decision 165). | — | the operator decides whether (a)–(c) are accepted or need work before the first paying customer | CC | | **R-132** | Security & access | P3 | **`curl -w '%{redirect_url}'` reconstructs the request URL WITH its basic-auth credential** — so a `-u ":$HUB_PW"` call that never put the password in a URL still printed it **Merged 2026-10-05 from R-350 (duplicate):** (1) 2026-08-20 occurrence: POST /configuration/artifacts answers 303; leak lives only in the CC transcript under ~/.claude/projects/ on DooPlex, not in git/evidence (checked then). (2) `-v` and `--libcurl` also re-render the credential, not only %{redirect_url}; confirm redirects with %{http_code} + follow-up GET. (3) Rotation path: hub /configuration form (current_password/new_password/confirm_password); DB override wins over ConfigMap (break-glass); CC can rotate file-to-file without printing (operator-present-one-time-secrets) if asked. | **WAITING-ON-OPERATOR** — **ACTION: rotate `HUB_PW`** **Folded R-580 2026-10-03** (the same `curl -w %{redirect_url}` credential echo, seen again 2026-09-18). | — | Happened on 2026-07-31 while red-proofing the R-120 gate: the hub operator password was written to the session transcript by the write-out format, not by the request. `-u` is safe; the *reporting* was not. Rule: read the redirect from `-D -` and grep `^Location:`, never `%{redirect_url}`, on any authenticated call. Rotate the hub password (`/configuration` → Login password; ConfigMap `auth.password_hash` is the reset path) and update `~/.config/credentials` | Viktor | | **R-137** | Security & access | P3 | **Cloudflare geo-WAF rules are zone-scoped and non-namespaced — four cross-tenant faults.** `globalRuleDesc = "[felhom-geo] Global"` (`waf.go:18`) is one literal description per ZONE; `appRuleDescPrefix` keys by app name with no customer (`waf.go:21`); `BuildGlobalExpression` has no positive hostname scoping (`waf.go:241`); `applyDiff` deletes every `[felhom-geo]` rule not in THIS box's desired set (`geosync.go:320`) | READY (M) — **blocks shared-zone onboarding** | — | With two customers in one zone: they overwrite each other's Global rule forever; one customer's country policy applies zone-wide; per-app rules collide by name; and disabling the feature for one (or the hub's `RemoveGeoRules`) wipes them all. Interim mitigation, no code: keep geo-restriction OFF for every shared-zone customer. Fix = namespace descriptions by `customer_id` + add `http.host ends_with ""` to both expressions — a TWO-REPO change (controller + hub `RemoveGeoRules`). Same audit §5.1 | CC | | **R-255** | Security & access | P3 | **The check that would catch a fourth secret-in-the-body covers 4 of 27 pages, and the cheap gate that covers all 36 templates is blind to the shape that actually shipped.** Filed 2026-08-08 while closing R-254, **because a partial guard reported as complete is worse than no guard — it stops the next person looking.** **Two nets, both measured.** **(1) `scripts/secret_in_markup_gate.py`** reads all 36 templates and convicts any `{{ … }}` naming a secret unless allowlisted with a reason. It catches `{{.RetrievalPassword}}` and `{{.InitialCreds.Password}}`, **and it catches a launder through a local variable** because the assignment itself names the secret (`{{$v := .InitialCreds.Password}}` is convicted — verified). **It is blind to a secret arriving under a NEUTRAL PAGE-DATA KEY** — `data["Tagline"] = creds.Password` then `{{.AppInfo.Tagline}}` passes it cleanly, also verified. **That is exactly the shape of R-254 site two** (`value="{{$val}}"` inside an `{{if eq .Type "secret"}}` branch), so the gate **would not have caught one of the three instances it was written for.** **(2) The runtime body assertion** — render the page and grep the response for a sentinel — catches every shape, including that one (demonstrated on the same planted leak the gate missed). But it needs each page's data to be constructible in a test, and **only 4 of 27 page templates have that today**: `settings_security`, `app_info`, `deploy`, `backups_restore` — the four that were touched by R-249/R-252/R-253/R-254 and therefore got their own tests. **The other 23 pages have no runtime coverage at all.** **What closing this needs, so the cost is not re-estimated:** a per-page data fixture for the remaining 23 (most need a wired `Server` — `stackMgr`, `backupMgr`, agent seams), then one table-driven test that renders each with a sentinel substituted for every string in its data and asserts the sentinel is absent. **That is real scaffolding, which is why it was NOT built inside R-254's session** rather than half-built and declared done. | **READY** — owner Viktor | — | — | operator | @@ -242,11 +241,11 @@ stopping line that lies. | **R-906** | Monitoring & notifications | P4 | **A page can show „+ 5 további figyelmeztetés" (+5 more warnings) with no warning above it.** SEEN 2026-10-08 on scratch 9202 (controller 0.303.0): `GetAlerts` caps the list at 5 and counts the rest into the overflow line BEFORE the layout drops the alerts that belong on other pages (`disk-not-separate` is `Inline` + `PageOnly` dashboard/monitoring, `web/alerts.go` ~L281); on the launcher, apps, backups and system pages all five visible ones were such alerts, so only the overflow line rendered. Not fixed here (controller out of scope). | **VERIFY — fixed on main 2026-10-08, ships with tomorrow's controller release** (controller `d5f2e47`). `AlertManager.GetBannerAlerts(page, lang)` filters with the layout's own rule BEFORE the cap; `baseData` and /monitoring use it. `TestR906_*` (the 9202 case: nine inline disk warnings + hub off, on /launcher, /stacks, /backups, /monitoring, hu + en — red against 05e12921). Live on 9202 with a test build: the overflow line is gone and the three real banner warnings show (`audits/dashboard-layout-2026-10-08/`) | — | Release tomorrow; read one page on a box with inline warnings; close | CC | | **R-911** | Monitoring & notifications | P4 | **On an English page the „data storage not reachable" banner stays Hungarian: „Adattároló nem elérhető: ".** SEEN 2026-10-08 on 9202 (test build, `?lang=en`), once R-906 let the real banner warnings through. Cause: the health check appends `warnFmtStorageUnavailable` as plain text (`internal/monitor/healthcheck.go:322`) with no `MsgRef`, so the banner has no key to render in English (the R-516 item 10 pattern). | **OPEN — owner: CC** | — | Give the warning a `MsgRef` (new key `alert`/`health` hu + en, the wire text unchanged); English page test for the banner | CC | -## Hub & operator — 12 rows (P2 1, P3 5, P4 6) +## Hub & operator — 11 rows (P2 1, P3 4, P4 6) | ID | Category | Sev | What | State | Blocked on | Next action | Owner | |---|---|---|---|---|---|---|---| -| **R-173** | Hub & operator | P2 | **The hub's SQLite PVC is excluded from every Longhorn backup job.** `pvc/hub-data` carries `recurring-job-group.longhorn.io/default: disabled`, and `backup-daily` + `backup-weekly` (04:00 / Sun 05:00) are the ONLY recurring jobs and both target the `default` group — so the 128 MB `/data/hub.db` has **no volume-level backup**. That database holds `host_recovery` (every managed box's break-glass root password), `host_escrow` + `host_escrow_superseded` (escrow custody), `host_pbs_secrets`, `customer_configs`, `dr_recipe` and the wg endpoints/peers — i.e. the material several documented recovery routes depend on | **NARROWED 2026-10-05 (evening) — option A IN FORCE** (`09` decision 125). The hub writes a nightly `VACUUM INTO` snapshot at 02:00 (hub v0.136.0, `05` §16.3, keep 2; the volume grew to 2 Gi); DooPlex checks it (`integrity_check`, size, ≥1 host, ≤26 h old), encrypts it with a key ep0 never sees and pushes it at 02:30 to ep0's `operator` namespace with a write-only token; a read-only token restore-tests it every Sunday 04:30 (and refuses a readable console password); `HubDBBackupStale`/`HubDBRestoreTestStale` alarm on success-only timestamps, `absent()` included. Both keys are off DooPlex (operator, 2026-10-05). PVC label fixed (`enabled`). Proven live: first push 7 s, restore test, token limits, a key rebuilt from the paper `data` field decrypts, runbook §3 steps 1–3 (4/4 console passwords open with the saved seal key, 0/4 with a random one). `audits/hub-db-offsite-2026-10-05/` | — | **LEFT:** runbook §3 steps 4–5 (the copy into a live PVC) are not exercised — they need the hub down; do them at the next planned hub maintenance or a DR drill on a scratch k3s. Close then. | CC | +| **R-922** | Hub & operator | P2 | **A household that clears its mail address on the dashboard does not get it cleared on the hub — the hub keeps the old address.** SEEN 2026-10-09 on Tester 1 (`audits/release-2026-10-09/proofs/D3/d3.txt`): the push with an empty address answered 200 and the hub row kept the old address (with no events, so no mail is sent). Cause: the empty-email no-clobber guard (`hub/internal/api/handler.go`, v0.71.0, audit F12) protects against an unconfigured box wiping a seeded address, and cannot tell that from a household's deliberate clear. Personal data the household removed stays on our side with no stated end — the same family as R-901's deletion rules. | **WAITING-ON-OPERATOR** — the rule: is a deliberate clear a deletion (the hub drops the address), and how is it told apart from an unconfigured box (e.g. an explicit `cleared: true` from the controller)? Then CC builds it (controller + hub). | operator decision | Operator: the rule; CC: the fix in both repos | operator (rule), CC (fix) | | **R-31** | Hub & operator | P3 | **[P2-HIGH] Offsite provisioning is synchronous with no status affordance.** Save runs the Hetzner sync in-request, so the request can hit the nginx 504 **while succeeding server-side**: the operator cannot tell failed from slow, and a retry races the first attempt. **MIGRATED FROM `ROADMAP.md` 2026-08-22 (R-369) — originally filed 2026-07-21, size M, roadmap state `idea`.** Moved verbatim; nothing added or reinterpreted. The roadmap keeps its copy as history, marked moved. | **OPEN — migrated from ROADMAP 2026-08-22, rank unchanged** **Re-ranked 2026-10-03: P2→P3: operator-only; a known workaround (click once, wait, verify) exists.** **2026-10-06 night: the race half fixed on felhom.eu main (hub, unreleased):** a second Save while the first still provisions is refused with 409 („already running — wait about a minute, then reload; do not save again"); nothing saved, nothing created; per customer, in memory. `TestProvision_R31_*`, red-proof `audits/night-burndown-2026-10-06/hub/R-31-red.txt`. LEFT: the async save with a status card (the escrow-card idiom). | — | Direction: make it async + a status card, reusing the proven **awaiting-card/poll idiom** (v0.138.0 escrow card). **Interim mitigation belongs in R-3 as an operator note: click once, wait, verify — do not re-click.** | CC | | **R-244** | Hub & operator | P3 | **The customer DELETE cascade leaves `app_log_issues` behind, and it is systematic across every venue ever torn down.** Found **2026-08-07** while verifying the `finalwalk` teardown with a **full census** (every table, every column) rather than a per-table query. After a cascade that logged `COMPLETE … full teardown`, **61 rows still matched `finalwalk`**. Four of the five sources are **deliberate and correct** — the cascade's own header states *"Provenance/events are NEVER wiped — audit outlives every tier"*: `events` 16, `notification_log` 14, `host_deletions` 1, `customer_resets` 1. **The fifth is a gap:** `app_log_issues` 29 rows, which the residue purge does not touch (its logged leg covers `reports`/`app_telemetry`/`app_log_tails`/`log_tail_requests`/`notif_prefs`/`selfbind_tokens`/`appliance_registrations` — not this table). **It is not a `finalwalk` quirk:** rows still reference **`c11` 40, `rewalk` 20, `part4` 24** — all three torn down 2026-08-06, whose ledger recorded *"0 occurrences"*. **That prior claim was measured with a narrower query and does not survive a full census; the correction is recorded rather than the measurement quietly redone.** **Why it was probably never written, established rather than assumed:** the table is a **fleet-wide aggregate** keyed on `app_name`+`fingerprint` with an `affected_customers` JSON list — of the 29 `finalwalk` rows, **12 reference only `finalwalk`** (orphans, safely deletable) and **17 are shared with LIVE customers** (`demo-felhom`, `peti-felhom`, …) and **must not be deleted, only de-referenced.** A naive `DELETE … WHERE customer LIKE` would destroy a live customer's issue history — which is very likely why the leg does not exist, and is the reason this is not a one-line fix. **Severity is LOW and stated plainly: no secret material is involved** — app name, fingerprint, message text, counts, timestamps. What survives is a deleted customer's *identifier* inside an aggregate row. **Proposed shape:** a residue leg that (a) removes the customer id from `affected_customers`/`context_customer`, and (b) deletes rows whose `affected_customers` becomes empty; plus a one-off sweep for the four already-torn-down venues. **The general lesson is the reusable part:** *a per-table absence query is not a census.* The teardown verification is now a full-schema sweep, and that is what found this. **Not fixed** — a cascade change needs its own red-proof and this session was scoped as a spike plus two operations. Evidence: `tests/teardown-finalwalk-2026-08-07.md`. **⚠ STILL OWED, AND NOW MEASURED RATHER THAN ESTIMATED (2026-08-08 census, read-only, no truncation).** `app_log_issues` holds **1309 rows**; **71 reference a torn-down venue** (`finalwalk`, `c11`, `rewalk`, `part4`); of those **44 are ORPHANS** — they name only torn-down customers and are safely deletable — and **27 are SHARED with a live customer** (`demo-felhom`, `peti-felhom`, …) and **must be de-referenced, never deleted**. 1238 rows are untouched. **The 27 are exactly why the leg was never written**, and why a `DELETE … WHERE customer LIKE` would destroy a live customer's issue history. **What it needs, precisely:** a cascade leg that (a) removes the customer id from `affected_customers` / `context_customer`, and (b) deletes only rows whose `affected_customers` becomes empty; plus a one-off sweep for the four venues already gone. **Why it was NOT done on 2026-08-08:** the fix is hub code, and that session's scope forbade a hub version bump; a hand-run SQL mutation over 71 rows — 27 of them needing surgical de-referencing — with no tested code path and no red-proof is precisely the shape that goes wrong on a live database. **It accumulates one venue at a time, so the next walk adds to it**; the numbers above mean the next session starts from data rather than a guess. **⚠ IT GREW AGAIN, AS PREDICTED — walk5 teardown, 2026-08-08.** The fifth walk's venue was torn down with a full-schema census taken **before and after**: **168 rows → 67**. Of the 67, **37 are by design** (`events` 21, `notification_log` 14, `host_deletions` 1, `customer_resets` 1) and **30 are `app_log_issues`** — this row's gap, and the count was **predicted in the pre-run enumeration rather than discovered afterwards**, which is the difference from the ledger that once recorded *"0 occurrences"* from a narrower query. **The running total across torn-down venues therefore rises from 71 to ~101 rows** (`finalwalk`, `c11`, `rewalk`, `part4`, now `walk5`) — the shared-with-a-live-customer subset must still be **de-referenced, never deleted**. **It accumulates one venue at a time and it did so again.** Evidence: `tests/walk5-r201-2026-08-07/teardown-walk5-2026-08-08.md`. **2026-09-25:** `peti-felhom` is no longer a live customer (deleted through the cascade, journal #20); 8 `app_log_issues` rows still name it — the same gap. | **READY** — owner Viktor | — | — | operator | | **R-882** | Hub & operator | P3 | **Longhorn on DooPlex could not grow a volume online: its `instance-manager` (116 days up) called a host process that no longer existed** — `nsenter: cannot open /host/proc/196610/ns/mnt` on every expansion retry, and an offline growth was blocked by the expansion's own attachment ticket (found 2026-10-05 growing `hub-data` to 2 Gi). A restart of the instance-manager (operator-approved) fixed it: 77/77 volumes back `attached/healthy` in 110 s. **Why the cached PID went stale was not established** (likely a containerd/k3s or iscsid restart after the instance-manager started), so it will recur after the next such restart and stay invisible until a volume needs to grow. `audits/hub-db-offsite-2026-10-05/partA/step1-*.txt` | **OPEN** | — | Find which host process the PID was and whether Longhorn 1.10.x re-resolves it; until then, before growing any volume, check the instance-manager's age against the last k3s/containerd restart | operator | diff --git a/documentation/runbooks/RUNBOOK-hub-db-offsite-backup.md b/documentation/runbooks/RUNBOOK-hub-db-offsite-backup.md index 671a365d..bf199841 100644 --- a/documentation/runbooks/RUNBOOK-hub-db-offsite-backup.md +++ b/documentation/runbooks/RUNBOOK-hub-db-offsite-backup.md @@ -153,11 +153,19 @@ Run the unit by hand; read the snapshot on ep0 (`proxmox-backup-client snapshot test by hand; stop the timer for a day on purpose and see `HubDBBackupStale` mail arrive (positive observable), then start it again. -## 3. Bringing the hub back from this copy (the procedure the plan exists for) — TESTED 2026-10-05 +## 3. Bringing the hub back from this copy (the procedure the plan exists for) — TESTED 2026-10-05 (steps 1–3) and 2026-10-09 (steps 4–5) Steps 1–3 were run on 2026-10-05 against the real copy on ep0 (`audits/hub-db-offsite-2026-10-05/partD/restore-procedure/drill.txt`): -4 hosts, **4 of 4 console passwords opened with the saved seal key, 0 of 4 with a random key**. Steps 4–5 (into a live -PVC) were NOT run — that needs the hub down; they are the ordinary scale-copy-scale. +4 hosts, **4 of 4 console passwords opened with the saved seal key, 0 of 4 with a random key**. **Steps 4–5 were run on +2026-10-09 into a throwaway single-node k3s** (the bench, LXC 9401, k3s v1.33.6 in a Docker container on an INTERNAL +network — `audits/dooplex-survival-2026-10-09/partD-*.txt`): the hub image of the live version started on the restored +copy, the customer list and host list equal live (4/4, 4/4), all 4 console passwords revealed (R-173 closed). + +> **⚠ Corrected 2026-10-09 — a restored hub mails households AT ONCE.** Within a minute of starting on the copy it +> tried to send two households a pending „kernel notice" (blocked only because the test had no network). Turning +> operator mail off (`operator_enabled: false`) does NOT stop household mail. So: **a test restore has no network at +> all** (no route to Resend, to ep0, to the boxes). **A real recovery** should expect the notices the copy still holds +> to go out once — check the hub's pending notices before giving it a network if that matters. What you need, all from the password manager: the seal key (`OFFSITE_SECRET_KEY`), the backup key's `data` field, and the read-only token (or ep0 root to mint a new one: Step 2). @@ -183,9 +191,19 @@ the read-only token (or ep0 root to mint a new one: Step 2). `failed>0` means the wrong seal key: the hub would start but could open no console password (`05` §16.2). 4. **A k3s with the `felhom` ArgoCD app**, and `Secret/offsite-secret-key` recreated with the SAME value: `kubectl -n felhom-system create secret generic offsite-secret-key --from-file=OFFSITE_SECRET_KEY=k`. + **Corrected 2026-10-09:** the Deployment also needs, NOT optional, `Secret/resend-api` (`RESEND_API_KEY`), + `Secret/report-api` (`REPORT_API_KEY`) and `Secret/gitea-creds` (`username`, `password`); without them the pod does + not start. On a real rebuild their values come from DooPlex's nightly secrets export, which since 2026-10-09 is also + off-site (`runbooks/gitea-restore.md`, `secrets/*.gpg`, opened with DooPlex's restic passphrase). A test uses dummies. + The hub's image is pulled from Gitea's registry — on a rebuild with no registry, `docker save` it from any machine + that has it, or build it from the restored code. 5. **Into the PVC:** scale `deploy/hub` to 0; put `out/hub.db` into the volume as `/data/hub.db` (a helper pod mounting `hub-data`; delete any `hub.db-wal`/`-shm` there — the snapshot is a whole database); scale to 1. The start-up log line `console passwords sealed at rest (0 legacy plaintext row(s) sealed now)` and one reveal on a host page confirm it. + As run 2026-10-09: `kubectl scale deploy/hub --replicas=0`; a `busybox` pod mounting `hub-data` at `/data`; + `kubectl cp hub.db hubdb-copy:/data/hub.db.new`, then in the pod `rm -f /data/hub.db-wal /data/hub.db-shm && mv + /data/hub.db.new /data/hub.db`; delete the pod; scale to 1 (20 s to Ready). The customer list is `GET /configs` + (rows link to `/customers/`), the hosts `GET /hosts`. 6. Shred `k`, `enc.key` copies and `out/` when done. ## 4. Option B (not my pick): restic to a dedicated Hetzner Storage Box sub-account