diff --git a/documentation/audits/dooplex-survival-2026-10-09/vaultwarden/red-proof.txt b/documentation/audits/dooplex-survival-2026-10-09/vaultwarden/red-proof.txt new file mode 100644 index 00000000..b305fbff --- /dev/null +++ b/documentation/audits/dooplex-survival-2026-10-09/vaultwarden/red-proof.txt @@ -0,0 +1,25 @@ +### RED: no-user check removed +FAIL: test_vaultwarden_with_no_user_refuses (__main__.Push.test_vaultwarden_with_no_user_refuses) +Ran 1 test in 0.527s +FAILED (failures=1) +### RED: backup-file removal removed +FAIL: test_vaultwarden_is_copied_and_its_backup_file_removed (__main__.Push.test_vaultwarden_is_copied_and_its_backup_file_removed) +Ran 1 test in 0.492s +FAILED (failures=1) +### RED: integrity check removed +Ran 1 test in 0.468s +OK +### RED: restore: Vaultwarden presence check removed +FAIL: test_restore_without_vaultwarden_fails (__main__.RestoreTest.test_restore_without_vaultwarden_fails) +Ran 1 test in 0.726s +FAILED (failures=1) +### GREEN +Ran 26 tests in 40.521s +OK +### RED (re-run, test now names the check): integrity check removed +FAIL: test_a_corrupt_vaultwarden_copy_refuses (__main__.Push.test_a_corrupt_vaultwarden_copy_refuses) +Ran 1 test in 0.447s +FAILED (failures=1) +### GREEN +Ran 26 tests in 40.608s +OK diff --git a/documentation/backlog/OPEN-ITEMS.md b/documentation/backlog/OPEN-ITEMS.md index cd34c1d7..8061a085 100644 --- a/documentation/backlog/OPEN-ITEMS.md +++ b/documentation/backlog/OPEN-ITEMS.md @@ -147,7 +147,7 @@ stopping line that lies. | **R-683** | App updates | P3 | **[P3-LOW] Watch: after a power cut during an update's health check, the hold named an HOUR-OLD second-drive copy, not the one the update's own backup should have just made.** 2026-09-24 chaos round 3 (nextcloud, `backup_max_age: 1m`): no `backing-up` phase was seen and the hold named Tier 2 at 13:04 for an update pressed at 14:04; the pre-cut controller log was lost with the container (the runner now saves it at arm time — R-320). Round 11, the same action without a power cut, named a fresh 14:34 copy and logged the Tier-2 copy. The sentence was TRUE (it named the copy it offered); the question is why the update did not back up first. Not reproduced; watch the next power-cut drill. `audits/night-2026-09-24/E/round-03*.json`, `E/round-11-controller-pre.log` | **OPEN — P3; owner: CC (watch)** | — | — | CC | | **R-785** | App updates | P3 | **[P3-LOW] SparkyFitness is pinned 11 releases and a major behind upstream (v0.17.3; upstream v1.7.3, v1.6.0 dated 2026-07-24).** READ 2026-10-01 (`audits/visitors-2026-10-01/C/bench/C1-previous-tag.txt`). **Needs:** an update walk 0.17 → 1.x through the ladder (bench + box), after R-784 is decided. | **OPEN — rank P3-LOW; owner: CC (after R-784)** | — | — | CC | -## Backup & restore — 29 rows (P2 3, P3 12, P4 14) +## Backup & restore — 30 rows (P2 4, P3 12, P4 14) | ID | Category | Sev | What | State | Blocked on | Next action | Owner | |---|---|---|---|---|---|---|---| @@ -155,6 +155,7 @@ stopping line that lies. | **R-304** | Backup & restore | P2 | **The retained escrow key works, and the customer is told their correct code is wrong.** DRILL 2026-08-12 answered the three questions separately, on `demo-felhom`, with planted data. **(a) retention: WORKS** — the first retained row in fleet history to carry material (`host_escrow_superseded` id 11, `identity_blob` 572 B), byte-identical (`sha256 a10032341c8584ed…`) to the pre-supersession `host_escrow` row. **(b) the material opens the old store: YES** — unsealed with the OLD recovery code it yielded a password byte-identical to the pre-change one (`sha c60c8bc737a6b7c6…`), and restored three planted files **byte-identical** from a store the box itself could no longer open (negative control first: `Fatal: wrong password or no key found`), **including a Hungarian accented filename verified as raw bytes**. **(c) the customer's route: DOES NOT EXIST, and misinforms.** `ListSupersededEscrow` (`store.go:2841`) is the only reader of a retained `identity_blob` and has **zero production callers** — five call sites, all `_test.go`; the product path (`POST /escrow/recover-offsite-password` → `FetchIdentityEscrow` → `GetHostDRBundle`, `store.go:3152`) selects `FROM host_escrow` — the CURRENT row only. Asked for the old password with the code that demonstrably opens the retained row, the product answered **"the recovery code did not open the sealed bundle — nothing was written"**. **This is the R-224 class again**: there an unreachable hub was reported as a bad code; here a VALID code for retained history is reported as a bad code, and the customer's attempt ends there. **Consequence:** the census answer stands (it was about retention); the countdown banner's promise is true in substance and false in practice; **any capability-map claim that the customer can recover the old history with their recovery code is false today and must move** | **READY — the operator mail is BUILT on main 2026-10-08 (decision 183; controller `a40729a` + hub, ships with the next releases).** The honesty fix is on main too (agent `91b9405`, controller `75b3b39`). Left: the runbook „open a retained package for a household" (design option C, second slice). Design `audits/day-2026-10-08/design-R-304.md`. | R-198, R-199, R-224, R-241 | Release controller + hub; write the retained-package runbook from the 2026-08-12 drill §4; then close | CC | | **R-893** | Backup & restore | P3 | **After a failed OFF-SITE replay, the rollback pours the NEWER pre-restore copy over the OLDER volume just put back.** Read in source 2026-10-06 (R-638 option A, not measured): `internal/backup/offbox_reconstitute.go` writes the undo copy from the live (newer) database, replaces the volumes with the snapshot's older tars, then — when the replay fails — `rollbackSafetyDump` loads that newer dump over the older database volume. The loader only drops what the dump knows, so tables the newer migration removed stay; and when the snapshot's older definition was written, the rollback branch does not put the newer definition back, so the older app starts on rolled-back data; non-database volumes stay at the snapshot's state. An order change cannot fix it (the only undo is a logical dump, and its volume was replaced). Known limit in `07` §6.3. **-- 2026-10-08 design:** `audits/day-2026-10-08/design-R-893.md` — re-verified; and the screen's „your data is back as it was" (`err.backup.db_restore_failed_rolled_back`) is false for files, other volumes and the app version. Question D8 on STATUS's decision sheet. **-- 2026-10-08 14:16 operator ruling D8 (`09` §3 decision 192):** yes, both, in that order — first the app stays stopped for support (option C), next „put back exactly as it was" (option A). | **NARROWED — 2026-10-09: the first half DELIVERED in controller 0.304.0. Slice 0 (the live measurement) NOT run: 9202 has no off-site, and a dump cannot be cut by hand inside an encrypted off-site copy on Tester 1; it needs a scratch off-site repo — a plan, not a finding.** **NARROWED 2026-10-08 — the first half (D8 option C, the app held stopped) is built on controller main and ships tomorrow; what stays open is the second half, „put back exactly as it was” (option A: a pre-restore copy of the app's volumes and placed files, a fit check, disk for one copy; the slice-0 measurement on 9202 first).** **OPEN — filed 2026-10-06** **2026-10-06 night: verified in source, no code** — `offbox_reconstitute.go:758` (undo dump from the live DB), `:778-860` (files and volumes from the snapshot), `:804` (the snapshot's definition is written when its version differs), `:898` (the rollback loads the newer dump over the older volume; nothing writes the live definition back). Not a reorder fix: it needs R-638 option B (a rebuilding loader) or a pre-restore volume copy (disk cost; R-685's class). Which state a household gets after a failed off-site replay is the operator's call. Next: the 9202 measurement, then the design. | a design: R-638 option B (a loader that rebuilds instead of overlays) or a pre-restore volume copy | Measure it once on 9202 (a forced replay failure after an off-site restore over a migrated app); then a design for the operator | CC | | **R-895** | Backup & restore | P2 | **The hub's clean-up-window check trusts the snapshot counts the box sends, so a broken-into box (or past-dated fakes added through the add-only key) can shrink the real off-site history without an alarm.** READ 2026-10-06 night in source (R-822's design): the before/after comparison uses counts the box itself reports (`hub/internal/offsitekeys/service.go:284`, `:343`); new fakes keep the count level. Decision 68 already accepts a box-trusted count. | **OPEN — filed 2026-10-06 night** **2026-10-07 07:58: kept open for later (`09` §3 decision 166).** | a design + one read-only measurement (does the Storage Box shell on port 23 show snapshot file upload times?) | Option B of `audits/night-burndown-2026-10-06/design-R-822.md`: the hub lists the repo's `snapshots/` files over its own login before and after a window and alarms on snapshots no box run explains | CC | +| **R-923** | Backup & restore | P2 | **The password manager (Vaultwarden) runs on DooPlex and holds the keys that open DooPlex's own off-site copies — a loss of DooPlex loses the keys with the copies.** Operator finding, 2026-10-09 11:19: "the password manager also runs on DooPlex (Vaultwarden), so I probably need to back up its DB somewhere too." Measured: `vaultwarden-system/vaultwarden` (1.37.3), PVC `vaultwarden-data` 7.4 MB (`db.sqlite3` + WAL, `rsa_key.pem`, no attachments); its only backups were Longhorn on DooPlex. **A key in a password manager that runs on DooPlex is not off DooPlex** — R-173's "both keys are off DooPlex" (2026-10-05) and the hub-DB runbook Step 0 were wrong in that sense. Also seen: Vaultwarden warns its `ADMIN_TOKEN` is plain text (an Argon2 hash is its guidance). | **NARROWED 2026-10-09 (operator yes in chat):** Vaultwarden's own `backup` copy + `rsa_key.pem` ride the nightly encrypted DooPlex→ep0 job (`scripts/dooplex-offsite/`), checked on every push and every Sunday (integrity, users, items — rows only). That copy opens only with the DooPlex off-site key, and the vault only with the master password — so the circle is broken only once both are on paper outside DooPlex: `runbooks/break-glass-sheet.md`, `runbooks/total-loss-of-dooplex.md`. **LEFT:** the operator prints the sheet and stores it away from home; the `ADMIN_TOKEN` hash (a DooPlex change, operator's word). | operator: print the sheet | Operator: print + store the sheet; then close | operator | | **R-49** | Backup & restore | P3 | **[P2] The offsite capture set is ~90% cache and duplication — 1.1 GB of a 1.2 GB immich "photo backup".** Measured 2026-07-19: `immich_ml_cache.tar` **823 660 032 B (~60%)** — re-downloadable ML model weights; `immich_postgres_data.tar` **308 251 136 B (~23%)** — a raw tar of the postgres data dir that DUPLICATES the logical `.sql` dump captured beside it; `upload/backups/` **18 MB** — immich's own nightly dump, a backup inside the backup, growing daily; plus the stranded pre-v3 `dccc13fe…` tree (~36 MB) no DB has ever referenced. Actual irreplaceable content: **72 MB of originals**. **MIGRATED FROM `ROADMAP.md` 2026-08-22 (R-369) — originally filed 2026-07-19, size S–M, roadmap state `idea`.** Moved verbatim; nothing added or reinterpreted. The roadmap keeps its copy as history, marked moved. | **OPEN — migrated from ROADMAP 2026-08-22, rank unchanged** **Re-ranked 2026-10-03: P2→P3: wasted space and transfer, no data risk; needs a capture-set ruling, not a sale blocker.** | — | **Evidence: `audits/DIAG-immich-restore-round2-2026-07-19.md` §4 (full byte breakdown).** This is the customer's offsite quota and transfer cost, and it lands on the Hetzner sub-account they are billed for. **Recorded, deliberately not changed** — a capture-set exclusion is a data-loss-shaped decision and gets its own ruling, not a drive-by edit. Candidates in priority order: (a) `immich_ml_cache` — pure cache, strongest case; (b) the `postgres_data` volume tar where a logical dump of the same DB is already captured (the dump is what the restore path actually replays); (c) `upload/backups/`. Likely generalises past immich into a template-classification rule about cache volumes and self-backup directories, so it should be specified against the catalog, not one app | CC | | **R-127** | Backup & restore | P3 | **The catalog's `data_key: true` flag is UNRELIABLE — at least four data-encrypting keys the catalog itself labels as encryption keys are unflagged; and the O4 restore path can regenerate a DB password that then does not match the restored data directory** | **READY (S/M)** **2026-10-06: leg (a) PUSHED** to the live catalog (`c265b37`). **2026-10-06 (burn-down night, later): leg (b) NEEDS A DESIGN.** `07` §7.4 sets no direction; refusing the restore without the DB password, or `ALTER USER` after it, each change restore behaviour on customer data. | — | **Found by D5's Part 0, and it is why D5's boundary is `type: secret` rather than `data_key`.** Two separable legs. **(a) The misclassification.** Only 5 fields across 4 apps set `data_key: true` (`adventurelog/SECRET_KEY`, `homebox/HBOX_AUTH_API_KEY_PEPPER`, `papra/AUTH_SECRET`, `sparkyfitness/{API_ENCRYPTION_KEY,BETTER_AUTH_SECRET}`), yet `n8n/N8N_ENCRYPTION_KEY` („Titkosítási kulcs"), `wanderer/POCKETBASE_ENCRYPTION_KEY` („Adatbázis titkosítási kulcs"), `calcom/CALENDSO_ENCRYPTION_KEY` and `bookstack/APP_KEY` are unflagged — the catalog's own Hungarian labels contradict the flag. **D5 makes this non-urgent but not harmless:** everything `type: secret` now travels, so the keys DO reach the drive; what stays wrong is the **fail-closed gate**, which only refuses for `data_key` names — so if one of these is missing from both sources the restore proceeds onto data it cannot decrypt instead of refusing. Fix = flag them (`app-catalog-felhom.eu`, a catalog-only change) + a gate/test that the flag set and the label set agree. **(b) The regenerated-DB-password trap.** `internal/backup/restore_unit.go` O4 generates a replacement for any missing non-data-key secret. Proven on `postgres:16-alpine`: with PGDATA restored from the volume tar, `POSTGRES_PASSWORD` is ignored (initdb skipped), so the app fails over the compose network while the dump replay still succeeds through the container's local **trust** socket — success reported, data unreachable. v0.188.0 corrected the WARN's false claim that *"stored data is unaffected"* and scoped it, but did **not** add a guard: D5 shrinks this to the rare case (the secret was empty at capture AND absent from the guest). Real fix = either treat a DB password as fail-closed like a data key, or `ALTER USER` to the regenerated value after the volume restore. 18 DB/root-password fields are in scope; MariaDB fails loudly instead (Access denied), which is the safer half | CC | | **R-231** | Backup & restore | P3 | **`/opt/backup/scripts/` on DooPlex is unversioned host state** — found 2026-08-06 while adding the auto-memory store to the backup set. No repository tracks the scripts that protect the recovery chain, so the edit made that day (`CLAUDE_MEMORY_DIR` in `backup-config.sh`, multi-path restic call in `backup-data.sh`) exists only on the box. This is the same class the part-2 session was closing, found inside the fix for it; the change is transcribed in `felhom.eu/workspace/README.md` so it is at least *recorded*. **Two related facts, both understating current safety:** the backup destination (`/mnt/5_hdd/backup`) is on the **same physical disk** as the workspace it protects, and the DooPlex backup set has **no off-site leg** (`sync-hetzner-backups.sh` is jarrs.eu and pulls *from* Hetzner *to* DooPlex). Bringing a root-owned production backup script under version control, and deciding what installs it, is its own scoped change. | **READY** — owner Viktor | — | — | operator | @@ -245,7 +246,7 @@ stopping line that lies. | ID | Category | Sev | What | State | Blocked on | Next action | Owner | |---|---|---|---|---|---|---|---| -| **R-922** | Hub & operator | P2 | **A household that clears its mail address on the dashboard does not get it cleared on the hub — the hub keeps the old address.** SEEN 2026-10-09 on Tester 1 (`audits/release-2026-10-09/proofs/D3/d3.txt`): the push with an empty address answered 200 and the hub row kept the old address (with no events, so no mail is sent). Cause: the empty-email no-clobber guard (`hub/internal/api/handler.go`, v0.71.0, audit F12) protects against an unconfigured box wiping a seeded address, and cannot tell that from a household's deliberate clear. Personal data the household removed stays on our side with no stated end — the same family as R-901's deletion rules. | **WAITING-ON-OPERATOR** — the rule: is a deliberate clear a deletion (the hub drops the address), and how is it told apart from an unconfigured box (e.g. an explicit `cleared: true` from the controller)? Then CC builds it (controller + hub). | operator decision | Operator: the rule; CC: the fix in both repos | operator (rule), CC (fix) | +| **R-922** | Hub & operator | P2 | **A household that clears its mail address on the dashboard does not get it cleared on the hub — the hub keeps the old address.** SEEN 2026-10-09 on Tester 1 (`audits/release-2026-10-09/proofs/D3/d3.txt`): the push with an empty address answered 200 and the hub row kept the old address (with no events, so no mail is sent). Cause: the empty-email no-clobber guard (`hub/internal/api/handler.go`, v0.71.0, audit F12) protects against an unconfigured box wiping a seeded address, and cannot tell that from a household's deliberate clear. Personal data the household removed stays on our side with no stated end — the same family as R-901's deletion rules. | **READY — operator ruling 2026-10-09 11:19: option A** (the hub deletes the address when the household clears it). CC builds it; it ships with the next release. | — | Build A (controller + hub), red tests, privacy-notice line | CC | | **R-31** | Hub & operator | P3 | **[P2-HIGH] Offsite provisioning is synchronous with no status affordance.** Save runs the Hetzner sync in-request, so the request can hit the nginx 504 **while succeeding server-side**: the operator cannot tell failed from slow, and a retry races the first attempt. **MIGRATED FROM `ROADMAP.md` 2026-08-22 (R-369) — originally filed 2026-07-21, size M, roadmap state `idea`.** Moved verbatim; nothing added or reinterpreted. The roadmap keeps its copy as history, marked moved. | **OPEN — migrated from ROADMAP 2026-08-22, rank unchanged** **Re-ranked 2026-10-03: P2→P3: operator-only; a known workaround (click once, wait, verify) exists.** **2026-10-06 night: the race half fixed on felhom.eu main (hub, unreleased):** a second Save while the first still provisions is refused with 409 („already running — wait about a minute, then reload; do not save again"); nothing saved, nothing created; per customer, in memory. `TestProvision_R31_*`, red-proof `audits/night-burndown-2026-10-06/hub/R-31-red.txt`. LEFT: the async save with a status card (the escrow-card idiom). | — | Direction: make it async + a status card, reusing the proven **awaiting-card/poll idiom** (v0.138.0 escrow card). **Interim mitigation belongs in R-3 as an operator note: click once, wait, verify — do not re-click.** | CC | | **R-244** | Hub & operator | P3 | **The customer DELETE cascade leaves `app_log_issues` behind, and it is systematic across every venue ever torn down.** Found **2026-08-07** while verifying the `finalwalk` teardown with a **full census** (every table, every column) rather than a per-table query. After a cascade that logged `COMPLETE … full teardown`, **61 rows still matched `finalwalk`**. Four of the five sources are **deliberate and correct** — the cascade's own header states *"Provenance/events are NEVER wiped — audit outlives every tier"*: `events` 16, `notification_log` 14, `host_deletions` 1, `customer_resets` 1. **The fifth is a gap:** `app_log_issues` 29 rows, which the residue purge does not touch (its logged leg covers `reports`/`app_telemetry`/`app_log_tails`/`log_tail_requests`/`notif_prefs`/`selfbind_tokens`/`appliance_registrations` — not this table). **It is not a `finalwalk` quirk:** rows still reference **`c11` 40, `rewalk` 20, `part4` 24** — all three torn down 2026-08-06, whose ledger recorded *"0 occurrences"*. **That prior claim was measured with a narrower query and does not survive a full census; the correction is recorded rather than the measurement quietly redone.** **Why it was probably never written, established rather than assumed:** the table is a **fleet-wide aggregate** keyed on `app_name`+`fingerprint` with an `affected_customers` JSON list — of the 29 `finalwalk` rows, **12 reference only `finalwalk`** (orphans, safely deletable) and **17 are shared with LIVE customers** (`demo-felhom`, `peti-felhom`, …) and **must not be deleted, only de-referenced.** A naive `DELETE … WHERE customer LIKE` would destroy a live customer's issue history — which is very likely why the leg does not exist, and is the reason this is not a one-line fix. **Severity is LOW and stated plainly: no secret material is involved** — app name, fingerprint, message text, counts, timestamps. What survives is a deleted customer's *identifier* inside an aggregate row. **Proposed shape:** a residue leg that (a) removes the customer id from `affected_customers`/`context_customer`, and (b) deletes rows whose `affected_customers` becomes empty; plus a one-off sweep for the four already-torn-down venues. **The general lesson is the reusable part:** *a per-table absence query is not a census.* The teardown verification is now a full-schema sweep, and that is what found this. **Not fixed** — a cascade change needs its own red-proof and this session was scoped as a spike plus two operations. Evidence: `tests/teardown-finalwalk-2026-08-07.md`. **⚠ STILL OWED, AND NOW MEASURED RATHER THAN ESTIMATED (2026-08-08 census, read-only, no truncation).** `app_log_issues` holds **1309 rows**; **71 reference a torn-down venue** (`finalwalk`, `c11`, `rewalk`, `part4`); of those **44 are ORPHANS** — they name only torn-down customers and are safely deletable — and **27 are SHARED with a live customer** (`demo-felhom`, `peti-felhom`, …) and **must be de-referenced, never deleted**. 1238 rows are untouched. **The 27 are exactly why the leg was never written**, and why a `DELETE … WHERE customer LIKE` would destroy a live customer's issue history. **What it needs, precisely:** a cascade leg that (a) removes the customer id from `affected_customers` / `context_customer`, and (b) deletes only rows whose `affected_customers` becomes empty; plus a one-off sweep for the four venues already gone. **Why it was NOT done on 2026-08-08:** the fix is hub code, and that session's scope forbade a hub version bump; a hand-run SQL mutation over 71 rows — 27 of them needing surgical de-referencing — with no tested code path and no red-proof is precisely the shape that goes wrong on a live database. **It accumulates one venue at a time, so the next walk adds to it**; the numbers above mean the next session starts from data rather than a guess. **⚠ IT GREW AGAIN, AS PREDICTED — walk5 teardown, 2026-08-08.** The fifth walk's venue was torn down with a full-schema census taken **before and after**: **168 rows → 67**. Of the 67, **37 are by design** (`events` 21, `notification_log` 14, `host_deletions` 1, `customer_resets` 1) and **30 are `app_log_issues`** — this row's gap, and the count was **predicted in the pre-run enumeration rather than discovered afterwards**, which is the difference from the ledger that once recorded *"0 occurrences"* from a narrower query. **The running total across torn-down venues therefore rises from 71 to ~101 rows** (`finalwalk`, `c11`, `rewalk`, `part4`, now `walk5`) — the shared-with-a-live-customer subset must still be **de-referenced, never deleted**. **It accumulates one venue at a time and it did so again.** Evidence: `tests/walk5-r201-2026-08-07/teardown-walk5-2026-08-08.md`. **2026-09-25:** `peti-felhom` is no longer a live customer (deleted through the cascade, journal #20); 8 `app_log_issues` rows still name it — the same gap. | **READY** — owner Viktor | — | — | operator | | **R-882** | Hub & operator | P3 | **Longhorn on DooPlex could not grow a volume online: its `instance-manager` (116 days up) called a host process that no longer existed** — `nsenter: cannot open /host/proc/196610/ns/mnt` on every expansion retry, and an offline growth was blocked by the expansion's own attachment ticket (found 2026-10-05 growing `hub-data` to 2 Gi). A restart of the instance-manager (operator-approved) fixed it: 77/77 volumes back `attached/healthy` in 110 s. **Why the cached PID went stale was not established** (likely a containerd/k3s or iscsid restart after the instance-manager started), so it will recur after the next such restart and stay invisible until a volume needs to grow. `audits/hub-db-offsite-2026-10-05/partA/step1-*.txt` | **OPEN** | — | Find which host process the PID was and whether Longhorn 1.10.x re-resolves it; until then, before growing any volume, check the instance-manager's age against the last k3s/containerd restart | operator | diff --git a/scripts/CHANGELOG.md b/scripts/CHANGELOG.md index 1012f853..c6a5a847 100644 --- a/scripts/CHANGELOG.md +++ b/scripts/CHANGELOG.md @@ -1,3 +1,16 @@ +## 2026-10-09 (afternoon) — dooplex-offsite carries the password manager too (R-923) + +- `felhom-dooplex-offsite`: a new step runs Vaultwarden's own `vaultwarden backup` (SQLite `VACUUM INTO`, consistent + while it runs) in its pod, copies that one file out and removes it, plus `rsa_key.pem` and (when they exist) + `attachments/`, `sends/`, `config.json` — Vaultwarden's backup guidance. Refuses on a failed backup command, an + unexpected file name, a link in the archive, a failed `integrity_check`, or no user; the backup file is removed even + on refusal. New metric `felhom_dooplex_offsite_last_success_vaultwarden_users`. +- `felhom-dooplex-offsite-restore-test`: checks the Vaultwarden copy — integrity, users = `USERS`, items > 0 (row + counts only, never contents). +- 6 new tests (26 total), green with GNU tools, BusyBox tools and the fake `sqlite3`; red-proofs in + `audits/dooplex-survival-2026-10-09/vaultwarden/red-proof.txt`. The tests found a real defect before any live run: + the optional-file listing returned non-zero when the last optional file was absent. + ## 2026-10-09 — dooplex-offsite: Gitea + DooPlex's secrets leave DooPlex nightly, encrypted (R-232 (b), (h)) - New `scripts/dooplex-offsite/`: `felhom-dooplex-offsite` (daily 00:20) copies the newest complete `gitea.dump` diff --git a/scripts/dooplex-offsite/felhom-dooplex-offsite b/scripts/dooplex-offsite/felhom-dooplex-offsite index 58e6bc35..a1d61386 100755 --- a/scripts/dooplex-offsite/felhom-dooplex-offsite +++ b/scripts/dooplex-offsite/felhom-dooplex-offsite @@ -1,5 +1,5 @@ #!/bin/sh -# felhom-dooplex-offsite — push Gitea (repositories, database dump, config) and DooPlex's nightly secrets export to +# felhom-dooplex-offsite — push Gitea (repositories, database dump, config), Vaultwarden (R-923) and DooPlex's nightly secrets export to # ep0's PBS, encrypted on DooPlex (R-232 (b)). Runs on DooPlex as root from felhom-dooplex-offsite.timer (00:20). # Plan: documentation/audits/dooplex-survival-2026-10-09/PLAN.md. Restore: documentation/runbooks/gitea-restore.md. # Pinned by test_dooplex_offsite.py. @@ -27,13 +27,15 @@ RETRY_SLEEP=${FELHOM_DXOFF_RETRY_SLEEP:-30} log() { echo "felhom-dooplex-offsite: $*"; } die() { echo "felhom-dooplex-offsite: FAILED: $*" >&2; exit 1; } K() { kubectl -n gitea-system exec deploy/gitea -c gitea -- "$@"; } +V() { kubectl -n vaultwarden-system exec deploy/vaultwarden -- "$@"; } umask 077 STAGE="$STATE/stage" mkdir -p "$STATE"; chmod 700 "$STATE" -rm -rf "$STAGE"; mkdir -p "$STAGE/root/gitea" "$STAGE/root/db" "$STAGE/root/secrets" +rm -rf "$STAGE"; mkdir -p "$STAGE/root/gitea" "$STAGE/root/db" "$STAGE/root/secrets" "$STAGE/root/vaultwarden" cleanup() { - for f in "$STAGE/root/gitea/gitea/conf/app.ini" "$STAGE/root/db/gitea.dump" "$STAGE/root/db/globals.sql"; do + for f in "$STAGE/root/gitea/gitea/conf/app.ini" "$STAGE/root/db/gitea.dump" "$STAGE/root/db/globals.sql" \ + "$STAGE/root/vaultwarden/db.sqlite3" "$STAGE/root/vaultwarden/rsa_key.pem"; do [ -f "$f" ] && [ ! -L "$f" ] && { shred -u "$f" 2>/dev/null || rm -f "$f"; } done rm -rf "$STAGE" @@ -83,6 +85,34 @@ SAGE=$((NOW - $(stat -c %Y "$NEWEST"))) cp "$SECRETS"/*-"$STAMP".*gpg "$STAGE/root/secrets/" log "secrets: $(ls "$STAGE/root/secrets" | wc -l | tr -d ' ') file(s) of $STAMP" +# 3b. the password manager (Vaultwarden, R-923) — its own `vaultwarden backup` (SQLite VACUUM INTO, consistent while it +# runs) writes ONE file into /data; it is copied out and that one file removed. Plus rsa_key.pem and, when they exist, +# attachments/, sends/, config.json (Vaultwarden's backup guidance). The vault items are encrypted under each user's +# master password; this job never opens them. Refuses on: the backup command failing, an unexpected file name, a +# failed integrity_check, no user. +VOUT=$(V /vaultwarden backup 2>&1) || die "vaultwarden backup: $(printf '%s' "$VOUT" | tail -n 1)" +VFILE=$(printf '%s\n' "$VOUT" | sed -n "s#^Backup to '\(.*\)' was successful.*#\1#p" | tail -n 1) +case "$VFILE" in data/db_[0-9]*_[0-9]*.sqlite3) ;; *) die "vaultwarden backup: unexpected output (no backup file named)" ;; esac +VPATH="/$VFILE" +VRC=0; V cat "$VPATH" > "$STAGE/root/vaultwarden/db.sqlite3" || VRC=$? +V rm -f "$VPATH" || log "WARNING: could not remove $VPATH from Vaultwarden's /data" +[ "$VRC" -eq 0 ] || die "copying $VPATH out of the Vaultwarden pod" +VLIST=$(V sh -c 'cd /data && for p in rsa_key.pem rsa_key.pub.pem config.json attachments sends; do if [ -e "$p" ]; then echo "$p"; fi; done') \ + || die "listing Vaultwarden's files" +[ -n "$VLIST" ] && { V tar -cf - -C /data $VLIST > "$STAGE/vw.tar" || die "copying Vaultwarden's files"; } +if [ -s "$STAGE/vw.tar" ]; then + LINKS=$(tar -tvf "$STAGE/vw.tar" | grep -c '^[lh]') || LINKS=0 + [ "$LINKS" -eq 0 ] || die "Vaultwarden's archive holds $LINKS link(s) — refused" + tar -xof "$STAGE/vw.tar" -C "$STAGE/root/vaultwarden" || die "unpacking Vaultwarden's files" + rm -f "$STAGE/vw.tar" +fi +VIC=$(sqlite3 -readonly "$STAGE/root/vaultwarden/db.sqlite3" 'PRAGMA integrity_check;' 2>&1 | head -n 5) || true +[ "$VIC" = "ok" ] || die "Vaultwarden integrity_check: $VIC" +VUSERS=$(sqlite3 -readonly "$STAGE/root/vaultwarden/db.sqlite3" 'SELECT COUNT(*) FROM users;' 2>/dev/null) || die "cannot count Vaultwarden users" +[ "${VUSERS:-0}" -gt 0 ] || die "the Vaultwarden copy holds no user" +echo "$VUSERS" > "$STAGE/root/vaultwarden/USERS" +log "vaultwarden: integrity ok, $VUSERS user(s), files: db.sqlite3 $(echo $VLIST)" + # 4. the manifest the restore test checks, then the push echo "$GOT" > "$STAGE/root/REPOS" (cd "$STAGE/root" && find . -type f ! -name MANIFEST.sha256 -print0 | sort -z | xargs -0 sha256sum > MANIFEST.sha256) \ @@ -103,6 +133,7 @@ TMP="$TEXTFILE_DIR/felhom_dooplex_offsite.prom.$$" echo "felhom_dooplex_offsite_last_success_timestamp_seconds $(date +%s)" echo "felhom_dooplex_offsite_last_success_bytes $BYTES" echo "felhom_dooplex_offsite_last_success_repositories $GOT" + echo "felhom_dooplex_offsite_last_success_vaultwarden_users $VUSERS" } > "$TMP" chmod 644 "$TMP" mv "$TMP" "$TEXTFILE_DIR/felhom_dooplex_offsite.prom" diff --git a/scripts/dooplex-offsite/felhom-dooplex-offsite-restore-test b/scripts/dooplex-offsite/felhom-dooplex-offsite-restore-test index 4c38e874..a9fad36f 100755 --- a/scripts/dooplex-offsite/felhom-dooplex-offsite-restore-test +++ b/scripts/dooplex-offsite/felhom-dooplex-offsite-restore-test @@ -21,7 +21,7 @@ die() { echo "felhom-dooplex-offsite-restore-test: FAILED: $*" >&2; exit 1; } umask 077 mkdir -p "$STATE"; chmod 700 "$STATE" T=$(mktemp -d "$STATE/restore.XXXXXX") -trap 'for f in "$T"/out/gitea/gitea/conf/app.ini "$T"/out/db/gitea.dump "$T"/out/db/globals.sql; do [ -f "$f" ] && shred -u "$f" 2>/dev/null; done; rm -rf "$T"' EXIT +trap 'for f in "$T"/out/gitea/gitea/conf/app.ini "$T"/out/db/gitea.dump "$T"/out/db/globals.sql "$T"/out/vaultwarden/db.sqlite3 "$T"/out/vaultwarden/rsa_key.pem; do [ -f "$f" ] && shred -u "$f" 2>/dev/null; done; rm -rf "$T"' EXIT export PBS_PASSWORD_FILE="$TOKENS/token-restore" PBS_FINGERPRINT LIST=$(proxmox-backup-client snapshot list host/dooplex-gitea --ns operator --output-format json --repository "$PBS_REPOSITORY_RESTORE") \ @@ -66,7 +66,17 @@ rm -rf "$T/fsck.git" [ -s "$O/gitea/gitea/conf/app.ini" ] || die "app.ini missing" pg_restore --list "$O/db/gitea.dump" >/dev/null || die "pg_restore cannot read gitea.dump" ls "$O"/secrets/*.gpg >/dev/null 2>&1 || die "no secrets file in the copy" -log "checked: $FILES files match the manifest, $GOT repositories pass git fsck, gitea.dump readable, $(ls "$O"/secrets | wc -l | tr -d ' ') secrets file(s)" +# Vaultwarden (R-923): rows, never contents. Copies made before 2026-10-09 midday have no vaultwarden/ — refused, since +# the newest copy is the one tested and every copy since then carries it. +VDB="$O/vaultwarden/db.sqlite3" +[ -s "$VDB" ] || die "no Vaultwarden database in the copy" +VIC=$(sqlite3 -readonly "$VDB" 'PRAGMA integrity_check;' 2>&1 | head -n 5) || true +[ "$VIC" = "ok" ] || die "Vaultwarden integrity_check: $VIC" +VUSERS=$(sqlite3 -readonly "$VDB" 'SELECT COUNT(*) FROM users;' 2>/dev/null) || die "cannot count Vaultwarden users" +VITEMS=$(sqlite3 -readonly "$VDB" 'SELECT COUNT(*) FROM ciphers;' 2>/dev/null) || die "cannot count Vaultwarden items" +[ "${VUSERS:-0}" -gt 0 ] && [ "$VUSERS" = "$(cat "$O/vaultwarden/USERS" 2>/dev/null)" ] || die "Vaultwarden users: $VUSERS, USERS says $(cat "$O/vaultwarden/USERS" 2>/dev/null)" +[ "${VITEMS:-0}" -gt 0 ] || die "the Vaultwarden copy holds no item" +log "checked: $FILES files match the manifest, $GOT repositories pass git fsck, gitea.dump readable, $(ls "$O"/secrets | wc -l | tr -d ' ') secrets file(s), Vaultwarden $VUSERS user(s) / $VITEMS item(s)" TMP="$TEXTFILE_DIR/felhom_dooplex_offsite_restore.prom.$$" { diff --git a/scripts/dooplex-offsite/test_dooplex_offsite.py b/scripts/dooplex-offsite/test_dooplex_offsite.py index 76ddcb91..9f397ae4 100644 --- a/scripts/dooplex-offsite/test_dooplex_offsite.py +++ b/scripts/dooplex-offsite/test_dooplex_offsite.py @@ -25,6 +25,27 @@ import os, subprocess, sys a = sys.argv[1:] pod = os.environ["FAKE_POD_DATA"] cmd = a[a.index("--") + 1:] +if "vaultwarden-system" in a: + vw = os.environ["FAKE_VW_DATA"] + if cmd == ["/vaultwarden", "backup"]: + if os.environ.get("FAKE_VW_BACKUP_FAIL"): + print("[ERROR] backup failed: database is locked"); sys.exit(1) + name = "db_20261009_094401.sqlite3" + import shutil as sh; sh.copy(os.path.join(vw, "db.sqlite3"), os.path.join(vw, name)) + print("[NOTICE] You are using a plain text `ADMIN_TOKEN` which is insecure.") + print("Backup to 'data/%s' was successful" % name); sys.exit(0) + if cmd[0] == "cat": + sys.exit(subprocess.call(["cat", vw + cmd[1][len("/data"):]])) + if cmd[:2] == ["rm", "-f"]: + p = vw + cmd[2][len("/data"):] + open(os.path.join(os.environ["FAKE_STATE"], "vw-removed"), "a").write(cmd[2] + "\n") + if os.path.exists(p): os.remove(p) + sys.exit(0) + if cmd[:2] == ["sh", "-c"]: + sys.exit(subprocess.call(["sh", "-c", cmd[2].replace("/data", vw)])) + if cmd[0] == "tar": + sys.exit(subprocess.call([vw if x == "/data" else x for x in cmd])) + sys.exit(96) if cmd[0] == "tar": cnt = os.path.join(os.environ["FAKE_STATE"], "tar-calls") n = int(open(cnt).read()) if os.path.exists(cnt) else 0 @@ -73,6 +94,31 @@ head -c 5 "$2" | grep -q '^PGDMP' || { echo "pg_restore: error: input file does ''' +# The CI runner has no sqlite3 CLI: a stand-in on Python's own sqlite3 module (the same SQLite library) answers the calls +# the scripts make — `sqlite3 -readonly ''`, one row per line, exit 1 on an error. +FAKE_SQLITE3 = r'''#!/usr/bin/env python3 +import sqlite3, sys +a = [x for x in sys.argv[1:] if x != "-readonly"] +try: + db = sqlite3.connect("file:" + a[0] + "?mode=ro", uri=True) + for row in db.execute(a[1]): + print("|".join("" if v is None else str(v) for v in row)) +except Exception as e: + print("Error: " + str(e), file=sys.stderr); sys.exit(1) +''' +REAL_SQLITE3 = None if os.environ.get("FORCE_FAKE_SQLITE3") else shutil.which("sqlite3") + + +def make_vw_db(path, users=1, items=3): + import sqlite3 + db = sqlite3.connect(path) + db.execute("CREATE TABLE users (uuid TEXT PRIMARY KEY, email TEXT)") + db.execute("CREATE TABLE ciphers (uuid TEXT PRIMARY KEY, data TEXT)") + for i in range(users): db.execute("INSERT INTO users VALUES (?, ?)", ("u%d" % i, "u%d@example.invalid" % i)) + for i in range(items): db.execute("INSERT INTO ciphers VALUES (?, ?)", ("c%d" % i, "2.encrypted")) + db.commit(); db.close() + + def git(*a, cwd=None): env = dict(os.environ, GIT_AUTHOR_NAME="t", GIT_AUTHOR_EMAIL="t@t", GIT_COMMITTER_NAME="t", GIT_COMMITTER_EMAIL="t@t") subprocess.run(["git", *a], cwd=cwd, check=True, capture_output=True, env=env) @@ -109,7 +155,14 @@ class Base(unittest.TestCase): open(os.path.join(self.conf, "enc.key"), "w").write("{}") open(os.path.join(self.conf, "env"), "w").write( "PBS_REPOSITORY_PUSH='u!push@h:1:s'\nPBS_REPOSITORY_RESTORE='u!restore@h:1:s'\nPBS_FINGERPRINT='aa'\n") - for name, body in (("kubectl", FAKE_KUBECTL), ("proxmox-backup-client", FAKE_PBS), ("pg_restore", FAKE_PG_RESTORE)): + self.vw = j("vw") + os.makedirs(self.vw) + make_vw_db(os.path.join(self.vw, "db.sqlite3")) + open(os.path.join(self.vw, "rsa_key.pem"), "w").write("-----TEST KEY-----\n") + fakes = [("kubectl", FAKE_KUBECTL), ("proxmox-backup-client", FAKE_PBS), ("pg_restore", FAKE_PG_RESTORE)] + if not REAL_SQLITE3: + fakes.append(("sqlite3", FAKE_SQLITE3)) + for name, body in fakes: p = os.path.join(self.bin, name) open(p, "w").write(body); os.chmod(p, 0o755) @@ -130,7 +183,7 @@ class Base(unittest.TestCase): def env(self, **kw): e = dict(os.environ, PATH=self.bin + ":" + os.environ["PATH"], FAKE_POD_DATA=self.pod, FAKE_PBS_DIR=self.pbs, - FAKE_STATE=self.t, FELHOM_DXOFF_CONF=self.conf, FELHOM_DXOFF_TOKENS=self.tokens, + FAKE_STATE=self.t, FAKE_VW_DATA=self.vw, FELHOM_DXOFF_CONF=self.conf, FELHOM_DXOFF_TOKENS=self.tokens, FELHOM_DXOFF_STATE=self.state, FELHOM_DXOFF_TEXTFILE_DIR=self.text, FELHOM_DXOFF_DUMPS=self.dumps, FELHOM_DXOFF_SECRETS=self.secrets) e.update({k: str(v) for k, v in kw.items()}) @@ -228,6 +281,39 @@ class Push(Base): self.assertIn("link(s)", r.stderr) self.assertEqual(self.pushed(), []); self.assertFalse(self.signal()) + def test_vaultwarden_is_copied_and_its_backup_file_removed(self): + r = self.push() + self.assertEqual(r.returncode, 0, r.stderr) + snap = os.path.join(self.pbs, "snaps", self.pushed()[0]) + self.assertTrue(os.path.isfile(os.path.join(snap, "vaultwarden/db.sqlite3"))) + self.assertTrue(os.path.isfile(os.path.join(snap, "vaultwarden/rsa_key.pem"))) + self.assertEqual(open(os.path.join(snap, "vaultwarden/USERS")).read().strip(), "1") + self.assertFalse(os.path.exists(os.path.join(self.vw, "db_20261009_094401.sqlite3")), "the backup file must be removed") + self.assertEqual(open(os.path.join(self.t, "vw-removed")).read().strip(), "/data/db_20261009_094401.sqlite3") + self.assertIn("vaultwarden_users 1", open(os.path.join(self.text, "felhom_dooplex_offsite.prom")).read()) + + def test_vaultwarden_backup_failing_refuses(self): + r = self.push(FAKE_VW_BACKUP_FAIL=1) + self.assertNotEqual(r.returncode, 0) + self.assertIn("vaultwarden backup", r.stderr) + self.assertEqual(self.pushed(), []); self.assertFalse(self.signal()) + + def test_vaultwarden_with_no_user_refuses(self): + os.remove(os.path.join(self.vw, "db.sqlite3")) + make_vw_db(os.path.join(self.vw, "db.sqlite3"), users=0) + r = self.push() + self.assertNotEqual(r.returncode, 0) + self.assertIn("no user", r.stderr) + self.assertEqual(self.pushed(), []); self.assertFalse(self.signal()) + self.assertFalse(os.path.exists(os.path.join(self.vw, "db_20261009_094401.sqlite3")), "removed even on refusal") + + def test_a_corrupt_vaultwarden_copy_refuses(self): + open(os.path.join(self.vw, "db.sqlite3"), "wb").write(b"SQLite format 3\x00" + b"\xff" * 4096) + r = self.push() + self.assertNotEqual(r.returncode, 0) + self.assertIn("Vaultwarden integrity_check", r.stderr) + self.assertEqual(self.pushed(), []); self.assertFalse(self.signal()) + def test_failed_push_writes_no_signal(self): r = self.push(FAKE_PBS_FAIL=1) self.assertNotEqual(r.returncode, 0) @@ -285,6 +371,24 @@ class RestoreTest(Base): r = self.restore() self.assertEqual(r.returncode, 0, r.stderr) + def test_restore_reports_vaultwarden_rows_never_contents(self): + self.pushed_copy() + r = self.restore() + self.assertEqual(r.returncode, 0, r.stderr) + self.assertIn("Vaultwarden 1 user(s) / 3 item(s)", r.stdout) + self.assertNotIn("encrypted", r.stdout + r.stderr) + + def test_restore_without_vaultwarden_fails(self): + snap = self.pushed_copy() + os.remove(os.path.join(snap, "vaultwarden/db.sqlite3")) + man = os.path.join(snap, "MANIFEST.sha256") + kept = [l for l in open(man).readlines() if "vaultwarden/db.sqlite3" not in l] + open(man, "w").writelines(kept) + r = self.restore() + self.assertNotEqual(r.returncode, 0) + self.assertIn("no Vaultwarden database", r.stderr) + self.assertFalse(self.signal("felhom_dooplex_offsite_restore.prom")) + def test_an_unreadable_dump_fails(self): shutil.rmtree(self.dumps); os.makedirs(self.dumps) self.add_dump("20261008-220001", self.now - 1200, magic=b"XXXXX")