From 07773bf58faf20c49596f28dfa97c9dc5fd692ca Mon Sep 17 00:00:00 2001 From: kisfenyo Date: Fri, 9 Oct 2026 13:01:29 +0200 Subject: [PATCH] R-923: STATUS, capability map, session report (break the circle) Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS --- REPORT-break-the-circle-2026-10-09.md | 61 +++++++++++++++++++ STATUS.md | 16 ++++- .../architecture/00-capability-map.md | 2 +- 3 files changed, 76 insertions(+), 3 deletions(-) create mode 100644 REPORT-break-the-circle-2026-10-09.md diff --git a/REPORT-break-the-circle-2026-10-09.md b/REPORT-break-the-circle-2026-10-09.md new file mode 100644 index 00000000..65e3dfc4 --- /dev/null +++ b/REPORT-break-the-circle-2026-10-09.md @@ -0,0 +1,61 @@ +# REPORT — break the circle: the password manager off-site, every recovery key listed, one sheet to print (2026-10-09) + +| Part | What | Result | +|---|---|---| +| A | Vaultwarden's data in the nightly off-site copy | DONE, LIVE (operator yes in chat) — run by hand, restore test, throwaway start | +| B | Recovery-key inventory + printable sheet + paper walk | DONE — `runbooks/total-loss-of-dooplex.md`, `runbooks/break-glass-sheet.md` | +| C | R-922 A, a quiet restored hub, R-921 | BUILT, committed, NOT released (hub + controller); R-921 partly | + +**Register: before 135 · after 138 · opened 2 (R-923, R-924) · closed 0.** A parallel session added R-925 meanwhile. + +**Rulings recorded first:** (1) the operator's finding → **R-923** (P2); "off DooPlex" corrected in the hub-DB runbook +(Step 0 and §3), `gitea-restore.md`, `secrets.md`, STATUS; (2) R-922 → option A, recorded in the row. + +## Part A — Vaultwarden (`vaultwarden-system`, 1.37.3, 7.4 MB) + +- **Method:** Vaultwarden's own `vaultwarden backup` (SQLite `VACUUM INTO`, consistent while it runs; its backup + guidance), plus `rsa_key.pem` and — when present — `attachments/`, `sends/`, `config.json`. The one backup file it + writes into `/data` is copied out and removed (also on refusal). Same archive, key and write-only token as Gitea. +- **Live:** first run 2026-10-09T10:02Z: `vaultwarden: integrity ok, 1 user(s)`, pushed in 13 s; `/data` clean after. +- **Restore test (Sunday job, run by hand):** 27 969 files match, 10 repos pass `git fsck`, `Vaultwarden 1 user(s) / 797 item(s)`. +- **Throwaway (bench 9401):** row counts only — users 1, ciphers 797, folders 5, **twofactor 0**, sso_users 1; a + Vaultwarden 1.37.3 started on the copy, `/alive` 200, no route out. Deleted (0 containers, 0 volumes, bench stopped); + DooPlex scratch shredded. No login, no item opened, the master password never asked for. +- Tests: 26 (6 new), green with GNU, BusyBox and the fake `sqlite3`; 4 red-proofs (`audits/dooplex-survival-2026-10-09/vaultwarden/`). + The tests found a defect before the live run (the optional-file listing failed when the last file was absent). +- Seen, not changed: Vaultwarden warns its `ADMIN_TOKEN` is plain text (in R-923). + +## Part B — what a total loss needs + +18 secrets/logins listed with where each lives today. **Exists only on DooPlex (lost with it):** S1 the DooPlex +off-site key, S2 the hub-DB key (otherwise only in Vaultwarden, which is behind S1), **S6 the signing keys (in no +backup at all, no passphrase — R-924)**, S7 the restore token (re-mintable). **Circles:** S8 Hetzner and S11 Cloudflare +logins if kept only in Vaultwarden. Checked: the nightly Secrets export opens with the DooPlex restic passphrase (S5) and +holds all eight Secrets the hub uses (names only read). **Paper walk:** with a filled sheet and the master password, +no Felhom recovery step needs DooPlex; without it, S1, S6 and S8 block. + +## Part C — code (waits for the next release) + +- **Hub** (`d55c590c`, `e5b9151e`, CI 1605 and 1606 success): R-922 A (`email_cleared` deletes the notification address; + F12 guard kept); **MAIL-HOLD** (`/MAIL-HOLD` → no mail on any path, dropped not queued, banner, release + button) — now step 1 of the hub restore runbook; two log lines no longer print the address; **SECURITY:** + `/preferences` and `/notify` accepted any box's key for any household — now 403 (red-proved). +- **Controller** (`ec9b997`, CI 1607 success): R-922 sends `email_cleared` after a household's clear; R-921 pre-check — + no stop while another tier's job is in flight. **Not covered:** the agent's host-wide busy lock (the case measured on + demo-hp) — no agent endpoint serves it; needs an agent field (R-921 LEFT). +- **Release order:** hub first, then the controller (the controller sends a field an older hub ignores). +- Built by two helpers under this brief's fences; reviewed, tested again and committed here. One stash of mine briefly + held a helper's uncommitted work (about a minute; nothing lost). + +## Teardown + +Machine: bench 9401 — 0 containers, 0 volumes, image removed, stopped as it was. Host: DooPlex scratch dirs shredded. +Hub: provisioned nothing. Vaultwarden: its one test backup file removed by name. + +## Decisions for the operator + +1. **R-924 — the signing keys.** (A) Print both on the sheet AND add them to the nightly encrypted copy (CC changes the + DooPlex job with your yes). (B) Print only the recovery key and keep it away from home. **Pick A.** **If you do + nothing:** after a loss of DooPlex no box accepts a signed update, bundle or OS step until it is re-enrolled on site. +2. **Print the sheet** (`runbooks/break-glass-sheet.md`, commands at its top, run from your workstation). **If you do + nothing:** the off-site copies exist but cannot be opened after a loss of DooPlex. diff --git a/STATUS.md b/STATUS.md index 5e7e0afa..eba88a08 100644 --- a/STATUS.md +++ b/STATUS.md @@ -2,8 +2,20 @@ **Ready for the first real tester (Tester-2): yes. Tester 2 (a laptop) is off; nothing was sent to it.** -**Updated 2026-10-09 (midday): hub 0.144.0; demo-hp, demo-felhom and Tester 1 run agent 0.154.0 and controller -0.304.0. The open-items list is at 135. Reports: `REPORT-dooplex-survival-2026-10-09.md`, `REPORT-day4-2026-10-09.md`.** +**Updated 2026-10-09 (afternoon): hub 0.144.0; demo-hp, demo-felhom and Tester 1 run agent 0.154.0 and controller +0.304.0. The open-items list is at 138. Reports: `REPORT-break-the-circle-2026-10-09.md`, `REPORT-dooplex-survival-2026-10-09.md`, `REPORT-day4-2026-10-09.md`.** + +## Afternoon (2026-10-09): the password manager is copied off-site; one page to print + +- **Your password manager (Vaultwarden) now goes to ep0 every night**, with the code. You said yes. A test brought it + back on a throwaway machine: it starts, with your 1 account and 797 items. Nobody opened the items. +- **The circle is broken only when you print the sheet.** The copy on ep0 opens with one key, and that key was only + in the password manager on DooPlex. The sheet lists the keys to print, with the commands. It holds no value itself. +- **New finding: the signing keys live only on DooPlex**, with no password and in no backup. If DooPlex is lost, no + box can take a signed update again without a visit. Your decision (see the report). +- **Built, not released:** a household that clears its mail address gets it deleted on the hub (your answer A); a + restored hub sends no mail until you release it; the backup no longer stops the apps when another backup is still + running. Found and fixed on the way: one box's key could change another household's mail settings. ## Midday (2026-10-09): the code and the hub can now survive losing DooPlex diff --git a/documentation/architecture/00-capability-map.md b/documentation/architecture/00-capability-map.md index c990e12c..0337709c 100644 --- a/documentation/architecture/00-capability-map.md +++ b/documentation/architecture/00-capability-map.md @@ -236,7 +236,7 @@ likewise silent. Evidence: `audits/DRILL-r361-2026-08-22/evidence/06-part3-decis | Operator login password changeable from UI | hub v0.54 | **IMPLEMENTED** | 07-13 | | | **The operator sees boxes left behind: per-customer controller floors with their age (and which a global raise cannot move), each box's agent against the vouched one, a 7-day "agent behind" alarm and a "floor raise skipped boxes" mail** | hub **v0.135.0** (R-604, R-530) | **IMPLEMENTED — the page PROVEN-LIVE, the alarm and the mail unit-proven** | `audits/hub-safety-2026-10-05/partD/` (live System page: three per-customer floors, Tester 2 `0.142.0 → 0.145.0`); `osupdates/r530_agent_alarm_test.go`, `web/r604_floor_held_back_test.go` | the mail fires only on a GLOBAL raise below an override — not exercised live | | **The hub database survives the loss of DooPlex: a nightly consistent copy, encrypted, on ep0; restore-tested weekly; an alarm when either stops** | hub **v0.136.0** (R-173), `scripts/hub-db-backup/`, homelab-manifests rules | **PROVEN-LIVE (2026-10-05)** — first push, ep0 listing, restore test (4 hosts, 4 sealed, 0 readable), token limits, a key rebuilt from the paper copy decrypts, the saved seal key opens 4/4 console passwords in the restored copy; the alarm by `promtool` rule test + red-proofs ; **2026-10-09: the whole recovery (§3 steps 1–5) PROVEN on a throwaway k3s** — the live hub image started on the restored copy, customers 4/4 and hosts 4/4 equal live, 4/4 console passwords revealed | `audits/hub-db-offsite-2026-10-05/`; `audits/dooplex-survival-2026-10-09/partD-*.txt`; `runbooks/RUNBOOK-hub-db-offsite-backup.md` | A restored hub mails households pending notices at once — a test restore has no network (runbook §3) | -| **Gitea (all code) and DooPlex's secrets survive the loss of DooPlex: a nightly encrypted copy on ep0; restore-tested weekly; an alarm when either stops; a failure mail** | `scripts/dooplex-offsite/` (R-232), homelab-manifests rules | **PROVEN-LIVE (2026-10-09)** — first push 57 s, read back with the read-only token (27 805 files, 10 repos pass `git fsck`); Gitea restored into a throwaway and started: 10/10 repos, product `main` = live, a file byte for byte, a login; alarm by `promtool` + red-proofs; failure mail reached the inbox | `audits/dooplex-survival-2026-10-09/`; `runbooks/gitea-restore.md` | The container registry is NOT copied (images rebuild from the code); the on-box backup tree is still one writable path (R-232 c) | +| **Gitea (all code), the operator's password manager (Vaultwarden, since 2026-10-09 afternoon, R-923) and DooPlex's secrets survive the loss of DooPlex: a nightly encrypted copy on ep0; restore-tested weekly; an alarm when either stops; a failure mail** | `scripts/dooplex-offsite/` (R-232), homelab-manifests rules | **PROVEN-LIVE (2026-10-09)** — first push 57 s, read back with the read-only token (27 805 files, 10 repos pass `git fsck`); Gitea restored into a throwaway and started: 10/10 repos, product `main` = live, a file byte for byte, a login; alarm by `promtool` + red-proofs; failure mail reached the inbox | `audits/dooplex-survival-2026-10-09/`; `runbooks/gitea-restore.md` | The container registry is NOT copied (images rebuild from the code); the on-box backup tree is still one writable path (R-232 c); the copy opens only with keys that must be on paper (`runbooks/break-glass-sheet.md`; the signing keys are in no copy, R-924) | | **The operator surface refuses a cross-site form post on BOTH login paths (session token; Basic auth + `X-Felhom-Operator`)** | hub **v0.135.0** (R-135) | **PROVEN-LIVE (2026-10-05)** | `audits/hub-safety-2026-10-05/partA/live.txt` (Basic, no header → 403 even on an unknown path; with the header → passes; header without credentials → 401); `web/r135_csrf_test.go` (39 paths) | | | Box operating-system security updates (Proxmox host, guest Debian, Docker engine) | agent v0.143.0, hub v0.133.0 | **PARTIAL — the GUEST and HOST Debian fast lanes and the DOCKER engine slow lane are PROVEN-LIVE (2026-10-04), with the System page, the fleet view and the alarms; the KERNEL lane is MISSING** | Guest: `audits/os-guest-lane-2026-10-04/`. Host + fleet + alarms: `audits/os-host-lane-2026-10-04/`. Docker + System page + crash guard: `audits/os-docker-crash-2026-10-04/` — live-restore on with the same container ids on every box; Docker 29.8.2 on both demo boxes; operator-approved Docker release; a signed undo and a signed ring-1 step; a replay refused; the crash guard restarted demo-hp twice and kept it off the third time. Design `architecture/11-os-updates.md` §5.8, §5.9, §8 | **No automatic undo** (guest: last night's backup; host: by-hand runbook; Docker: a signed undo job); existing boxes get root-owned files by the signed config bundle since 2026-10-04 (R-840 CLOSED; a box from before agent 0.143.0 needs one by-hand bootstrap — Tester 2: R-862; `audits/r840-config-bundle-2026-10-04/`); test approvals now end with the test (R-859); the agent's sudoers is root-equivalent (R-861); the kernel lane (R-836); facts reach the hub late after a boot (R-853). **2026-10-05 (agent v0.144.1):** R8 measures the real download (R-865); a killed pass still reports (R-868, live); the debug pass runs with the hub away (R-866, live); **a power cut mid-update was proven by day on demo-hp — the box came back by itself in 37 s, but the next pass fails until `dpkg --configure -a` is run by hand (R-876, P2, open)** — `audits/night-fixes-2026-10-05/`. **2026-10-05 afternoon (agent v0.145.0): R-876 FIXED and proven live — after a second crash mid-unpack the next pass repaired dpkg by itself (`REPAIR … journal=1`) and finished** — `audits/catchup-2026-10-05/partD/` | | **An ENGLISH-SPEAKING household's first hour: download, install, pair, bind, claim, two apps** | controller **v0.259.0** + hub **v0.119.0** + ISO 1.29.0 + the whole catalog | **PROVEN-LIVE on 0.258.0 with one blocker; THE BLOCKER IS FIXED AND PROVEN, THE WALK IS NOT REPEATED** | `audits/DRILL-first-hour-en-0258-2026-09-20.md` — a fresh install 2026-09-20, one intervention (R-494), stop rule not reached. Then `audits/i18n-closing-2026-09-21/live/` — the three blockers fixed and each proven on a live box or in the operator's inbox: the claim page answers English through the real cookie path; the Backup page's tier names follow the language; and the setup mail carries **four plain-ASCII English words** where the drill's carried `képző-szkítia-ásatás`, one day apart in the same inbox. | **R-596, R-597 and R-598 are CLOSED.** What this row still does NOT claim: **the fixed journey has not been walked end to end by a stranger on a fresh install.** Three fixes proven at the endpoint are not an hour proven by a person, and this project's own rule is that fixes are not a journey (see the recovery-journey row). **Also not walked:** the recovery code (needs ep0), backup/restore/remove/power-cut (proven 2026-09-14), and the two Backup-page *warnings* themselves — guest 9201 is healthy and a healthy box renders none, so they are covered by handler render tests, not live. **Verdict: nothing known now stands between an English-speaking tester and their box — and that is a different sentence from "the walk passed".** |