R-923: STATUS, capability map, session report (break the circle)
gates / gates (push) Successful in 5m38s

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
2026-10-09 13:01:29 +02:00
parent e5b9151e05
commit 07773bf58f
3 changed files with 76 additions and 3 deletions
+61
View File
@@ -0,0 +1,61 @@
# REPORT — break the circle: the password manager off-site, every recovery key listed, one sheet to print (2026-10-09)
| Part | What | Result |
|---|---|---|
| A | Vaultwarden's data in the nightly off-site copy | DONE, LIVE (operator yes in chat) — run by hand, restore test, throwaway start |
| B | Recovery-key inventory + printable sheet + paper walk | DONE — `runbooks/total-loss-of-dooplex.md`, `runbooks/break-glass-sheet.md` |
| C | R-922 A, a quiet restored hub, R-921 | BUILT, committed, NOT released (hub + controller); R-921 partly |
**Register: before 135 · after 138 · opened 2 (R-923, R-924) · closed 0.** A parallel session added R-925 meanwhile.
**Rulings recorded first:** (1) the operator's finding → **R-923** (P2); "off DooPlex" corrected in the hub-DB runbook
(Step 0 and §3), `gitea-restore.md`, `secrets.md`, STATUS; (2) R-922 → option A, recorded in the row.
## Part A — Vaultwarden (`vaultwarden-system`, 1.37.3, 7.4 MB)
- **Method:** Vaultwarden's own `vaultwarden backup` (SQLite `VACUUM INTO`, consistent while it runs; its backup
guidance), plus `rsa_key.pem` and — when present — `attachments/`, `sends/`, `config.json`. The one backup file it
writes into `/data` is copied out and removed (also on refusal). Same archive, key and write-only token as Gitea.
- **Live:** first run 2026-10-09T10:02Z: `vaultwarden: integrity ok, 1 user(s)`, pushed in 13 s; `/data` clean after.
- **Restore test (Sunday job, run by hand):** 27 969 files match, 10 repos pass `git fsck`, `Vaultwarden 1 user(s) / 797 item(s)`.
- **Throwaway (bench 9401):** row counts only — users 1, ciphers 797, folders 5, **twofactor 0**, sso_users 1; a
Vaultwarden 1.37.3 started on the copy, `/alive` 200, no route out. Deleted (0 containers, 0 volumes, bench stopped);
DooPlex scratch shredded. No login, no item opened, the master password never asked for.
- Tests: 26 (6 new), green with GNU, BusyBox and the fake `sqlite3`; 4 red-proofs (`audits/dooplex-survival-2026-10-09/vaultwarden/`).
The tests found a defect before the live run (the optional-file listing failed when the last file was absent).
- Seen, not changed: Vaultwarden warns its `ADMIN_TOKEN` is plain text (in R-923).
## Part B — what a total loss needs
18 secrets/logins listed with where each lives today. **Exists only on DooPlex (lost with it):** S1 the DooPlex
off-site key, S2 the hub-DB key (otherwise only in Vaultwarden, which is behind S1), **S6 the signing keys (in no
backup at all, no passphrase — R-924)**, S7 the restore token (re-mintable). **Circles:** S8 Hetzner and S11 Cloudflare
logins if kept only in Vaultwarden. Checked: the nightly Secrets export opens with the DooPlex restic passphrase (S5) and
holds all eight Secrets the hub uses (names only read). **Paper walk:** with a filled sheet and the master password,
no Felhom recovery step needs DooPlex; without it, S1, S6 and S8 block.
## Part C — code (waits for the next release)
- **Hub** (`d55c590c`, `e5b9151e`, CI 1605 and 1606 success): R-922 A (`email_cleared` deletes the notification address;
F12 guard kept); **MAIL-HOLD** (`<data_dir>/MAIL-HOLD` → no mail on any path, dropped not queued, banner, release
button) — now step 1 of the hub restore runbook; two log lines no longer print the address; **SECURITY:**
`/preferences` and `/notify` accepted any box's key for any household — now 403 (red-proved).
- **Controller** (`ec9b997`, CI 1607 success): R-922 sends `email_cleared` after a household's clear; R-921 pre-check —
no stop while another tier's job is in flight. **Not covered:** the agent's host-wide busy lock (the case measured on
demo-hp) — no agent endpoint serves it; needs an agent field (R-921 LEFT).
- **Release order:** hub first, then the controller (the controller sends a field an older hub ignores).
- Built by two helpers under this brief's fences; reviewed, tested again and committed here. One stash of mine briefly
held a helper's uncommitted work (about a minute; nothing lost).
## Teardown
Machine: bench 9401 — 0 containers, 0 volumes, image removed, stopped as it was. Host: DooPlex scratch dirs shredded.
Hub: provisioned nothing. Vaultwarden: its one test backup file removed by name.
## Decisions for the operator
1. **R-924 — the signing keys.** (A) Print both on the sheet AND add them to the nightly encrypted copy (CC changes the
DooPlex job with your yes). (B) Print only the recovery key and keep it away from home. **Pick A.** **If you do
nothing:** after a loss of DooPlex no box accepts a signed update, bundle or OS step until it is re-enrolled on site.
2. **Print the sheet** (`runbooks/break-glass-sheet.md`, commands at its top, run from your workstation). **If you do
nothing:** the off-site copies exist but cannot be opened after a loss of DooPlex.
+14 -2
View File
@@ -2,8 +2,20 @@
**Ready for the first real tester (Tester-2): yes. Tester 2 (a laptop) is off; nothing was sent to it.** **Ready for the first real tester (Tester-2): yes. Tester 2 (a laptop) is off; nothing was sent to it.**
**Updated 2026-10-09 (midday): hub 0.144.0; demo-hp, demo-felhom and Tester 1 run agent 0.154.0 and controller **Updated 2026-10-09 (afternoon): hub 0.144.0; demo-hp, demo-felhom and Tester 1 run agent 0.154.0 and controller
0.304.0. The open-items list is at 135. Reports: `REPORT-dooplex-survival-2026-10-09.md`, `REPORT-day4-2026-10-09.md`.** 0.304.0. The open-items list is at 138. Reports: `REPORT-break-the-circle-2026-10-09.md`, `REPORT-dooplex-survival-2026-10-09.md`, `REPORT-day4-2026-10-09.md`.**
## Afternoon (2026-10-09): the password manager is copied off-site; one page to print
- **Your password manager (Vaultwarden) now goes to ep0 every night**, with the code. You said yes. A test brought it
back on a throwaway machine: it starts, with your 1 account and 797 items. Nobody opened the items.
- **The circle is broken only when you print the sheet.** The copy on ep0 opens with one key, and that key was only
in the password manager on DooPlex. The sheet lists the keys to print, with the commands. It holds no value itself.
- **New finding: the signing keys live only on DooPlex**, with no password and in no backup. If DooPlex is lost, no
box can take a signed update again without a visit. Your decision (see the report).
- **Built, not released:** a household that clears its mail address gets it deleted on the hub (your answer A); a
restored hub sends no mail until you release it; the backup no longer stops the apps when another backup is still
running. Found and fixed on the way: one box's key could change another household's mail settings.
## Midday (2026-10-09): the code and the hub can now survive losing DooPlex ## Midday (2026-10-09): the code and the hub can now survive losing DooPlex
@@ -236,7 +236,7 @@ likewise silent. Evidence: `audits/DRILL-r361-2026-08-22/evidence/06-part3-decis
| Operator login password changeable from UI | hub v0.54 | **IMPLEMENTED** | 07-13 | | | Operator login password changeable from UI | hub v0.54 | **IMPLEMENTED** | 07-13 | |
| **The operator sees boxes left behind: per-customer controller floors with their age (and which a global raise cannot move), each box's agent against the vouched one, a 7-day "agent behind" alarm and a "floor raise skipped boxes" mail** | hub **v0.135.0** (R-604, R-530) | **IMPLEMENTED — the page PROVEN-LIVE, the alarm and the mail unit-proven** | `audits/hub-safety-2026-10-05/partD/` (live System page: three per-customer floors, Tester 2 `0.142.0 → 0.145.0`); `osupdates/r530_agent_alarm_test.go`, `web/r604_floor_held_back_test.go` | the mail fires only on a GLOBAL raise below an override — not exercised live | | **The operator sees boxes left behind: per-customer controller floors with their age (and which a global raise cannot move), each box's agent against the vouched one, a 7-day "agent behind" alarm and a "floor raise skipped boxes" mail** | hub **v0.135.0** (R-604, R-530) | **IMPLEMENTED — the page PROVEN-LIVE, the alarm and the mail unit-proven** | `audits/hub-safety-2026-10-05/partD/` (live System page: three per-customer floors, Tester 2 `0.142.0 → 0.145.0`); `osupdates/r530_agent_alarm_test.go`, `web/r604_floor_held_back_test.go` | the mail fires only on a GLOBAL raise below an override — not exercised live |
| **The hub database survives the loss of DooPlex: a nightly consistent copy, encrypted, on ep0; restore-tested weekly; an alarm when either stops** | hub **v0.136.0** (R-173), `scripts/hub-db-backup/`, homelab-manifests rules | **PROVEN-LIVE (2026-10-05)** — first push, ep0 listing, restore test (4 hosts, 4 sealed, 0 readable), token limits, a key rebuilt from the paper copy decrypts, the saved seal key opens 4/4 console passwords in the restored copy; the alarm by `promtool` rule test + red-proofs ; **2026-10-09: the whole recovery (§3 steps 1–5) PROVEN on a throwaway k3s** — the live hub image started on the restored copy, customers 4/4 and hosts 4/4 equal live, 4/4 console passwords revealed | `audits/hub-db-offsite-2026-10-05/`; `audits/dooplex-survival-2026-10-09/partD-*.txt`; `runbooks/RUNBOOK-hub-db-offsite-backup.md` | A restored hub mails households pending notices at once — a test restore has no network (runbook §3) | | **The hub database survives the loss of DooPlex: a nightly consistent copy, encrypted, on ep0; restore-tested weekly; an alarm when either stops** | hub **v0.136.0** (R-173), `scripts/hub-db-backup/`, homelab-manifests rules | **PROVEN-LIVE (2026-10-05)** — first push, ep0 listing, restore test (4 hosts, 4 sealed, 0 readable), token limits, a key rebuilt from the paper copy decrypts, the saved seal key opens 4/4 console passwords in the restored copy; the alarm by `promtool` rule test + red-proofs ; **2026-10-09: the whole recovery (§3 steps 1–5) PROVEN on a throwaway k3s** — the live hub image started on the restored copy, customers 4/4 and hosts 4/4 equal live, 4/4 console passwords revealed | `audits/hub-db-offsite-2026-10-05/`; `audits/dooplex-survival-2026-10-09/partD-*.txt`; `runbooks/RUNBOOK-hub-db-offsite-backup.md` | A restored hub mails households pending notices at once — a test restore has no network (runbook §3) |
| **Gitea (all code) and DooPlex's secrets survive the loss of DooPlex: a nightly encrypted copy on ep0; restore-tested weekly; an alarm when either stops; a failure mail** | `scripts/dooplex-offsite/` (R-232), homelab-manifests rules | **PROVEN-LIVE (2026-10-09)** — first push 57 s, read back with the read-only token (27 805 files, 10 repos pass `git fsck`); Gitea restored into a throwaway and started: 10/10 repos, product `main` = live, a file byte for byte, a login; alarm by `promtool` + red-proofs; failure mail reached the inbox | `audits/dooplex-survival-2026-10-09/`; `runbooks/gitea-restore.md` | The container registry is NOT copied (images rebuild from the code); the on-box backup tree is still one writable path (R-232 c) | | **Gitea (all code), the operator's password manager (Vaultwarden, since 2026-10-09 afternoon, R-923) and DooPlex's secrets survive the loss of DooPlex: a nightly encrypted copy on ep0; restore-tested weekly; an alarm when either stops; a failure mail** | `scripts/dooplex-offsite/` (R-232), homelab-manifests rules | **PROVEN-LIVE (2026-10-09)** — first push 57 s, read back with the read-only token (27 805 files, 10 repos pass `git fsck`); Gitea restored into a throwaway and started: 10/10 repos, product `main` = live, a file byte for byte, a login; alarm by `promtool` + red-proofs; failure mail reached the inbox | `audits/dooplex-survival-2026-10-09/`; `runbooks/gitea-restore.md` | The container registry is NOT copied (images rebuild from the code); the on-box backup tree is still one writable path (R-232 c); the copy opens only with keys that must be on paper (`runbooks/break-glass-sheet.md`; the signing keys are in no copy, R-924) |
| **The operator surface refuses a cross-site form post on BOTH login paths (session token; Basic auth + `X-Felhom-Operator`)** | hub **v0.135.0** (R-135) | **PROVEN-LIVE (2026-10-05)** | `audits/hub-safety-2026-10-05/partA/live.txt` (Basic, no header → 403 even on an unknown path; with the header → passes; header without credentials → 401); `web/r135_csrf_test.go` (39 paths) | | | **The operator surface refuses a cross-site form post on BOTH login paths (session token; Basic auth + `X-Felhom-Operator`)** | hub **v0.135.0** (R-135) | **PROVEN-LIVE (2026-10-05)** | `audits/hub-safety-2026-10-05/partA/live.txt` (Basic, no header → 403 even on an unknown path; with the header → passes; header without credentials → 401); `web/r135_csrf_test.go` (39 paths) | |
| Box operating-system security updates (Proxmox host, guest Debian, Docker engine) | agent v0.143.0, hub v0.133.0 | **PARTIAL — the GUEST and HOST Debian fast lanes and the DOCKER engine slow lane are PROVEN-LIVE (2026-10-04), with the System page, the fleet view and the alarms; the KERNEL lane is MISSING** | Guest: `audits/os-guest-lane-2026-10-04/`. Host + fleet + alarms: `audits/os-host-lane-2026-10-04/`. Docker + System page + crash guard: `audits/os-docker-crash-2026-10-04/` — live-restore on with the same container ids on every box; Docker 29.8.2 on both demo boxes; operator-approved Docker release; a signed undo and a signed ring-1 step; a replay refused; the crash guard restarted demo-hp twice and kept it off the third time. Design `architecture/11-os-updates.md` §5.8, §5.9, §8 | **No automatic undo** (guest: last night's backup; host: by-hand runbook; Docker: a signed undo job); existing boxes get root-owned files by the signed config bundle since 2026-10-04 (R-840 CLOSED; a box from before agent 0.143.0 needs one by-hand bootstrap — Tester 2: R-862; `audits/r840-config-bundle-2026-10-04/`); test approvals now end with the test (R-859); the agent's sudoers is root-equivalent (R-861); the kernel lane (R-836); facts reach the hub late after a boot (R-853). **2026-10-05 (agent v0.144.1):** R8 measures the real download (R-865); a killed pass still reports (R-868, live); the debug pass runs with the hub away (R-866, live); **a power cut mid-update was proven by day on demo-hp — the box came back by itself in 37 s, but the next pass fails until `dpkg --configure -a` is run by hand (R-876, P2, open)** — `audits/night-fixes-2026-10-05/`. **2026-10-05 afternoon (agent v0.145.0): R-876 FIXED and proven live — after a second crash mid-unpack the next pass repaired dpkg by itself (`REPAIR … journal=1`) and finished** — `audits/catchup-2026-10-05/partD/` | | Box operating-system security updates (Proxmox host, guest Debian, Docker engine) | agent v0.143.0, hub v0.133.0 | **PARTIAL — the GUEST and HOST Debian fast lanes and the DOCKER engine slow lane are PROVEN-LIVE (2026-10-04), with the System page, the fleet view and the alarms; the KERNEL lane is MISSING** | Guest: `audits/os-guest-lane-2026-10-04/`. Host + fleet + alarms: `audits/os-host-lane-2026-10-04/`. Docker + System page + crash guard: `audits/os-docker-crash-2026-10-04/` — live-restore on with the same container ids on every box; Docker 29.8.2 on both demo boxes; operator-approved Docker release; a signed undo and a signed ring-1 step; a replay refused; the crash guard restarted demo-hp twice and kept it off the third time. Design `architecture/11-os-updates.md` §5.8, §5.9, §8 | **No automatic undo** (guest: last night's backup; host: by-hand runbook; Docker: a signed undo job); existing boxes get root-owned files by the signed config bundle since 2026-10-04 (R-840 CLOSED; a box from before agent 0.143.0 needs one by-hand bootstrap — Tester 2: R-862; `audits/r840-config-bundle-2026-10-04/`); test approvals now end with the test (R-859); the agent's sudoers is root-equivalent (R-861); the kernel lane (R-836); facts reach the hub late after a boot (R-853). **2026-10-05 (agent v0.144.1):** R8 measures the real download (R-865); a killed pass still reports (R-868, live); the debug pass runs with the hub away (R-866, live); **a power cut mid-update was proven by day on demo-hp — the box came back by itself in 37 s, but the next pass fails until `dpkg --configure -a` is run by hand (R-876, P2, open)** — `audits/night-fixes-2026-10-05/`. **2026-10-05 afternoon (agent v0.145.0): R-876 FIXED and proven live — after a second crash mid-unpack the next pass repaired dpkg by itself (`REPAIR … journal=1`) and finished** — `audits/catchup-2026-10-05/partD/` |
| **An ENGLISH-SPEAKING household's first hour: download, install, pair, bind, claim, two apps** | controller **v0.259.0** + hub **v0.119.0** + ISO 1.29.0 + the whole catalog | **PROVEN-LIVE on 0.258.0 with one blocker; THE BLOCKER IS FIXED AND PROVEN, THE WALK IS NOT REPEATED** | `audits/DRILL-first-hour-en-0258-2026-09-20.md` — a fresh install 2026-09-20, one intervention (R-494), stop rule not reached. Then `audits/i18n-closing-2026-09-21/live/` — the three blockers fixed and each proven on a live box or in the operator's inbox: the claim page answers English through the real cookie path; the Backup page's tier names follow the language; and the setup mail carries **four plain-ASCII English words** where the drill's carried `képző-szkítia-ásatás`, one day apart in the same inbox. | **R-596, R-597 and R-598 are CLOSED.** What this row still does NOT claim: **the fixed journey has not been walked end to end by a stranger on a fresh install.** Three fixes proven at the endpoint are not an hour proven by a person, and this project's own rule is that fixes are not a journey (see the recovery-journey row). **Also not walked:** the recovery code (needs ep0), backup/restore/remove/power-cut (proven 2026-09-14), and the two Backup-page *warnings* themselves — guest 9201 is healthy and a healthy box renders none, so they are covered by handler render tests, not live. **Verdict: nothing known now stands between an English-speaking tester and their box — and that is a different sentence from "the walk passed".** | | **An ENGLISH-SPEAKING household's first hour: download, install, pair, bind, claim, two apps** | controller **v0.259.0** + hub **v0.119.0** + ISO 1.29.0 + the whole catalog | **PROVEN-LIVE on 0.258.0 with one blocker; THE BLOCKER IS FIXED AND PROVEN, THE WALK IS NOT REPEATED** | `audits/DRILL-first-hour-en-0258-2026-09-20.md` — a fresh install 2026-09-20, one intervention (R-494), stop rule not reached. Then `audits/i18n-closing-2026-09-21/live/` — the three blockers fixed and each proven on a live box or in the operator's inbox: the claim page answers English through the real cookie path; the Backup page's tier names follow the language; and the setup mail carries **four plain-ASCII English words** where the drill's carried `képző-szkítia-ásatás`, one day apart in the same inbox. | **R-596, R-597 and R-598 are CLOSED.** What this row still does NOT claim: **the fixed journey has not been walked end to end by a stranger on a fresh install.** Three fixes proven at the endpoint are not an hour proven by a person, and this project's own rule is that fixes are not a journey (see the recovery-journey row). **Also not walked:** the recovery code (needs ep0), backup/restore/remove/power-cut (proven 2026-09-14), and the two Backup-page *warnings* themselves — guest 9201 is healthy and a healthy box renders none, so they are covered by handler render tests, not live. **Verdict: nothing known now stands between an English-speaking tester and their box — and that is a different sentence from "the walk passed".** |