diff --git a/documentation/architecture/06-offsite-connectivity.md b/documentation/architecture/06-offsite-connectivity.md index c53fd566..e4a2b1ff 100644 --- a/documentation/architecture/06-offsite-connectivity.md +++ b/documentation/architecture/06-offsite-connectivity.md @@ -182,7 +182,7 @@ production endpoint exists. ### 3.6 The ep0 datastore has a second copy — decision 70 (2026-10-03) -**[DESIGN]** A nightly PBS pull-sync copies `felhom-offsite` from ep0 to DooPlex's PBS over a read-only token. ep0's server snapshot never covered this volume and Hetzner has no volume snapshots (R-342). The copy is ciphertext per customer. Built per the 2026-10-03 lock brief Part F; the restore route lives in `runbooks/`. +**[DESIGN]** A nightly PBS pull-sync copies `felhom-offsite` from ep0 to DooPlex's PBS over a read-only token. ep0's server snapshot never covered this volume and Hetzner has no volume snapshots (R-342). The copy is ciphertext per customer. **`[FACT]` BUILT 2026-10-03:** ep0 token `root@pam!dooplex-sync` (`DatastoreReader` only — the one change on ep0); ep0's PBS listens on `wg0` only, so DooPlex reaches it through an SSH forward (`felhom-ep0-pbs-tunnel.service`, operator ruling the same day); DooPlex datastore `ep0-copy`, sync daily 05:00 with `remove-vanished false`, verify Saturdays, failures mailed via Resend. First pull 201 s / 12 GB / 4 of 4 snapshots. Restore route: `runbooks/ep0-datastore-copy.md`. Evidence `audits/offsite-lock-build-2026-10-03/partF/`. ## 4. Robustness (production details beyond the spike) diff --git a/documentation/architecture/07-backup-architecture.md b/documentation/architecture/07-backup-architecture.md index e05fb908..eb82ed37 100644 --- a/documentation/architecture/07-backup-architecture.md +++ b/documentation/architecture/07-backup-architecture.md @@ -222,7 +222,7 @@ It carries guest sizing, `pve_storage[]`, and an app half with per-app `storage_ the identity bundle's shape is `{tunnel_token, pbs_token, wg_private_key, restic_repo_password}` (`felhom-agent/internal/escrow/identity.go:26-39`). -**[DESIGN] Off-site deletion custody (decisions 68–69, 2026-10-03).** The restic repository password stays on the box only. The box's off-site key is append-only (pinned in `authorized_keys`, written by the hub — the key registrar); the box never receives the sub-account password, which the hub stores encrypted at rest. Old snapshots are pruned by the box itself, only in a weekly window the hub opens, behind a fake-snapshot guard (R-822); until that ships nothing prunes. A Felhom-side pruner holding repository passwords is rejected. +**[DESIGN] Off-site deletion custody (decisions 68–69, 2026-10-03) — BUILT hub v0.127.0 / controller v0.289.1, live on both demo boxes.** The restic repository password stays on the box only. The box's off-site key is append-only (pinned in `authorized_keys`, written by the hub — the key registrar); the box never receives the sub-account password, which the hub stores encrypted at rest. Old snapshots are pruned by the box itself, only in a weekly window the hub opens, behind a fake-snapshot guard (R-822); until that ships nothing prunes. A Felhom-side pruner holding repository passwords is rejected. **[FACT] Three parts of the Recipe are empty or wrong on the live fleet**, and they are exactly the parts a host-loss recovery would read (INV Part D2.3): `hosts.dr_record_json` is `{}` on all three @@ -1115,11 +1115,11 @@ crosses the line — **R-158**. | 6 | **Guest lost or corrupted** | the host, both whole-guest tiers, the data drives (they are host binds) | `pct restore` from `local:` or `felhom-pbs:` | **operator** (SSH) | **84–112 s** local · **1101 s** PBS (both = restore-test into a scratch guest, boot + verify + teardown) | 24 h local · 7 d offsite | **PROVEN** | CAMPAIGN-2 T-P9; CAMPAIGN-8 Phase C (exact mount parity, `unprivileged: 1` preserved); LIVE restore-tests on both boxes this session | | 7 | **Guest stopped and does not come back** | everything | guest-power watchdog (60 s, `onboot` as the deliberate-stop discriminator) | automatic | **120 s** | | **PROVEN** | agent v0.107.0 replay — 120 s unattended vs the incident's 587 s with a human | | 8 | **Host dies (hardware), drives intact** | the data drives; ep0's PBS namespace; the hub's Recipe + Escrow | **install a new host, then** `--selftest=bring-up -mode dr` per guest, **then** re-attach drives by `durable_id` | **operator** (SSH) | | 7 d | **IMPLEMENTED** **`[BETA-DEFERRED]`** | bring-up code exists and has **never been executed** (`CAMPAIGN-8…:522`); the drive half of the plan is empty on every box (**R-105**) | -| 9 | **Whole box lost (fire/theft) — host and drives gone** | ep0 PBS namespace; the restic repo; the hub's Recipe + Escrow | new hardware → day-0 → escrow-consume with **R** → restore guests from PBS → app data from Tier-3 | **operator + customer** (R) | | 7 d (guest) · 24 h (app data) | **IMPLEMENTED / UNPROVEN** **`[BETA-DEFERRED]`** | every leg exists; the composed path has never been run. The destructive S5 drill is operator-gated and unrun (`06-offsite-connectivity.md:327`) | -| 10 | **Ransomware / malicious deletion inside the guest** | PBS offsite (the box **cannot** delete its own snapshots); **the restic repo is NOT protected the same way** | whole-guest restore from PBS to a point before the event | **operator** (SSH) | | 7 d | **PARTIAL** **`[BETA-DEFERRED]`** | R-89 proved the box is refused when deleting its own PBS snapshot (`CAMPAIGN-8…:514-515`). **R-95 (open, ranked #1):** the restic credential **can delete** — `readonly=False`, `forget --prune` runs from the box, and SFTP cannot express append-only (`OPEN-ITEMS.md:13`) **2026-08-30 (R-359): the store is now VERIFIED on a cadence** — a daily `offsite-integrity` job runs `restic check` when the last successful one is over 7 days old. **This is a readability check, not a restore-test (R-87 stays open).** The depth that ships ON now re-reads **100%** of the pack data (R-399 CLOSED, controller v0.228.0) — the sentence here previously said the opposite and was stale. **2026-08-31 (SPIKE R-87):** a scratch restore of every app on demo-hp was measured at **25 s for 8 snapshots / 774 MB**, against 40.3 s for one weekly check — but restic 0.14.0's `--verify` checks size and mtime, **not content**, so nothing available today can vouch for the restored BYTES. Row 4's verdict is UNCHANGED by the spike. `audits/SPIKE-restic-restore-test-2026-08-31.md`. **2026-09-01 (SPIKE R-95, `audits/SPIKE-r95-offsite-delete-2026-09-01.md`) — THIS ROW'S STATUS IS UNCHANGED, but the seven-day bound beside it is now known to be unverified.** Measured on BOTH boxes over their own SFTP credential, with controls: **no `.snapshots` is visible to either sub-account**, and the account is jailed. Either none exist or a sub-account cannot see them — and a snapshot the box cannot see is one the box cannot restore from, so recovery would be an operator act at the Hetzner panel, not a product capability (R-429; the confirming field needs the provider API, fenced by §11-D). **The PBS shape does NOT transfer**: the sub-account API has one permission axis, `readonly`, and a backup target cannot be read-only — PBS is a server that can refuse, a Storage Box is a filesystem that runs nothing. **What the spike removed as a fear:** withdrawing delete does NOT wedge the store (measured — restic treats a dead owner's lock as stale and proceeds). **What it added:** `unlock --remove-all` reports success while deleting nothing (R-430). **2026-09-01, LATER THE SAME DAY — THE PROBE ABOVE LOOKED FOR THE WRONG NAME AND THIS ROW'S SECOND CLAUSE IS NOW HALF WRONG.** It searched `.snapshots`; the vendor documents `/.zfs/snapshot`. Re-probed at the documented path on BOTH boxes, with controls, the two halves separate cleanly: **(a) the LIVE repository is deletable by the box — unchanged, R-95 stands;** **(b) the daily SNAPSHOTS of it are not writable by anything — PROVEN, not cited:** a write into `/.zfs/snapshot` is refused (`dest open …: Failure`) while the identical write to the account home succeeds. Seven daily snapshots confirmed in the panel (R-429). **So this row's "the restic repo is NOT protected the same way" is true of the repository and FALSE of its snapshots** — a deletion costs at most the day since the last snapshot, recoverable per-file (vendor). **Two limits kept honest:** a sub-account sees the snapshot directory EMPTY, so per-file recovery is operator-only today (R-432); and a panel-driven restore rolls back the WHOLE box. **The status is NOT moved** — evidence (b) is measured, but the recovery ROUTE has never been walked, which is what PARTIAL means. **Detection shipped hub v0.111.0 (R-431):** an unexplained fall in the snapshot count is noticed within a day — **but only a fall of MORE than half (R-435).** **2026-09-01, THE DRILL — CLAUSE (b) OF THE RE-SCOPE ABOVE IS WITHDRAWN; THE STATUS IS STILL NOT MOVED.** The re-scope said a deletion is *recoverable per-file*. **It is not, from the box: no snapshot is reachable by ANY name.** MEASURED on demo-hp, read-only, no delete verb issued: **777,600** exact names in the vendor form `YYYY-MM-DDTHH-MM-SS` (nine full days, second granularity) plus 126 alternative shapes — **zero hits**, against a control where the identical 600-name batch returns a path that exists (6/6). **Structural cause:** `/home` (`u629488-sub3`) is **st_dev 0,82**, `/.zfs/snapshot` is **st_dev 0,276**, and `/home/.zfs` does not exist — a snapshot under `/.zfs/snapshot` belongs to a different dataset than the one holding the repository. **Clause (a) — the box cannot WRITE into the snapshot area — is unchanged and re-confirmed.** So this row's *"recoverable per-file (vendor)"* and its limit *"per-file recovery is operator-only today (R-432)"* both overstate what exists: the remaining routes are a panel rollback of the WHOLE Storage Box and the provider API (fenced, and the hub's client has no snapshot method at all). **R-432 ANSWERED negatively; R-433 opened. RTO still blank — nothing was recovered, so nothing was timed.** `audits/evidence-drill-r95-recovery-2026-09-01/`. | +| 9 | **Whole box lost (fire/theft) — host and drives gone** | ep0 PBS namespace; the restic repo; the hub's Recipe + Escrow | new hardware → day-0 → escrow-consume with **R** → restore guests from PBS → app data from Tier-3 | **operator + customer** (R) | | 7 d (guest) · 24 h (app data) | **IMPLEMENTED / UNPROVEN** **`[BETA-DEFERRED]`** | every leg exists; the composed path has never been run. The destructive S5 drill is operator-gated and unrun (`06-offsite-connectivity.md:327`) **`[FACT]` 2026-10-03 (decision 70): ep0's datastore now has a nightly copy on DooPlex** (PBS pull-sync `ep0-felhom-offsite`, `remove-vanished false`, weekly verify, failures mailed to the operator); first pull 201 s, 12 GB, 4 of 4 snapshots, matching ep0 namespace for namespace. Restore route: `runbooks/ep0-datastore-copy.md` — never walked. `audits/offsite-lock-build-2026-10-03/partF/`. | +| 10 | **Ransomware / malicious deletion inside the guest** | PBS offsite (the box **cannot** delete its own snapshots); **the restic repo is NOT protected the same way** | whole-guest restore from PBS to a point before the event | **operator** (SSH) | | 7 d | **PARTIAL** **`[BETA-DEFERRED]`** | R-89 proved the box is refused when deleting its own PBS snapshot (`CAMPAIGN-8…:514-515`). **R-95 (open, ranked #1):** the restic credential **can delete** — `readonly=False`, `forget --prune` runs from the box, and SFTP cannot express append-only (`OPEN-ITEMS.md:13`) **2026-08-30 (R-359): the store is now VERIFIED on a cadence** — a daily `offsite-integrity` job runs `restic check` when the last successful one is over 7 days old. **This is a readability check, not a restore-test (R-87 stays open).** The depth that ships ON now re-reads **100%** of the pack data (R-399 CLOSED, controller v0.228.0) — the sentence here previously said the opposite and was stale. **2026-08-31 (SPIKE R-87):** a scratch restore of every app on demo-hp was measured at **25 s for 8 snapshots / 774 MB**, against 40.3 s for one weekly check — but restic 0.14.0's `--verify` checks size and mtime, **not content**, so nothing available today can vouch for the restored BYTES. Row 4's verdict is UNCHANGED by the spike. `audits/SPIKE-restic-restore-test-2026-08-31.md`. **2026-09-01 (SPIKE R-95, `audits/SPIKE-r95-offsite-delete-2026-09-01.md`) — THIS ROW'S STATUS IS UNCHANGED, but the seven-day bound beside it is now known to be unverified.** Measured on BOTH boxes over their own SFTP credential, with controls: **no `.snapshots` is visible to either sub-account**, and the account is jailed. Either none exist or a sub-account cannot see them — and a snapshot the box cannot see is one the box cannot restore from, so recovery would be an operator act at the Hetzner panel, not a product capability (R-429; the confirming field needs the provider API, fenced by §11-D). **The PBS shape does NOT transfer**: the sub-account API has one permission axis, `readonly`, and a backup target cannot be read-only — PBS is a server that can refuse, a Storage Box is a filesystem that runs nothing. **What the spike removed as a fear:** withdrawing delete does NOT wedge the store (measured — restic treats a dead owner's lock as stale and proceeds). **What it added:** `unlock --remove-all` reports success while deleting nothing (R-430). **2026-09-01, LATER THE SAME DAY — THE PROBE ABOVE LOOKED FOR THE WRONG NAME AND THIS ROW'S SECOND CLAUSE IS NOW HALF WRONG.** It searched `.snapshots`; the vendor documents `/.zfs/snapshot`. Re-probed at the documented path on BOTH boxes, with controls, the two halves separate cleanly: **(a) the LIVE repository is deletable by the box — unchanged, R-95 stands;** **(b) the daily SNAPSHOTS of it are not writable by anything — PROVEN, not cited:** a write into `/.zfs/snapshot` is refused (`dest open …: Failure`) while the identical write to the account home succeeds. Seven daily snapshots confirmed in the panel (R-429). **So this row's "the restic repo is NOT protected the same way" is true of the repository and FALSE of its snapshots** — a deletion costs at most the day since the last snapshot, recoverable per-file (vendor). **Two limits kept honest:** a sub-account sees the snapshot directory EMPTY, so per-file recovery is operator-only today (R-432); and a panel-driven restore rolls back the WHOLE box. **The status is NOT moved** — evidence (b) is measured, but the recovery ROUTE has never been walked, which is what PARTIAL means. **Detection shipped hub v0.111.0 (R-431):** an unexplained fall in the snapshot count is noticed within a day — **but only a fall of MORE than half (R-435).** **2026-09-01, THE DRILL — CLAUSE (b) OF THE RE-SCOPE ABOVE IS WITHDRAWN; THE STATUS IS STILL NOT MOVED.** The re-scope said a deletion is *recoverable per-file*. **It is not, from the box: no snapshot is reachable by ANY name.** MEASURED on demo-hp, read-only, no delete verb issued: **777,600** exact names in the vendor form `YYYY-MM-DDTHH-MM-SS` (nine full days, second granularity) plus 126 alternative shapes — **zero hits**, against a control where the identical 600-name batch returns a path that exists (6/6). **Structural cause:** `/home` (`u629488-sub3`) is **st_dev 0,82**, `/.zfs/snapshot` is **st_dev 0,276**, and `/home/.zfs` does not exist — a snapshot under `/.zfs/snapshot` belongs to a different dataset than the one holding the repository. **Clause (a) — the box cannot WRITE into the snapshot area — is unchanged and re-confirmed.** So this row's *"recoverable per-file (vendor)"* and its limit *"per-file recovery is operator-only today (R-432)"* both overstate what exists: the remaining routes are a panel rollback of the WHOLE Storage Box and the provider API (fenced, and the hub's client has no snapshot method at all). **R-432 ANSWERED negatively; R-433 opened. RTO still blank — nothing was recovered, so nothing was timed.** `audits/evidence-drill-r95-recovery-2026-09-01/`. **`[FACT]` 2026-10-03 (decisions 68–69, hub v0.127.0, controller v0.289.1): THE BOX CAN NO LONGER DELETE ITS RESTIC HISTORY.** Both demo boxes run on an append-only key the hub pinned; a delete from the box is refused (`403 Forbidden`, measured on each, counts 13→13 and 100→100); the box never receives the sub-account password (the old endpoint answers 410, measured with demo-hp's own key); the hub reads every key file daily. Retention runs only in a hub-opened window behind a fake-snapshot guard — mechanics proven live (window 1 on demo-hp: opened, guard refused, closed in 3 s, operator mailed); **a real prune inside a window is NOT yet proven** (R-824). Residual: an add-only attacker can still fill the quota and plant past-dated snapshots (R-822). `audits/offsite-lock-build-2026-10-03/`. | | 11 | **Hub lost** | every box's data plane, every tier, every Lane-1 route | none needed for recovery **of a box**; the hub itself restores from its Longhorn volume backup | operator (`kubectl`) | | 24 h + weekly (Longhorn `RETAIN 1` each) | **UNPROVEN** **`[BETA-DEFERRED]`** | a hub restore has never been performed. The backup target is `nfs://192.168.0.180` — **DooPlex itself** — and exactly **2** restore points exist (LIVE, INV Part D2.2) | | 11b | *consequences while the hub is gone* | — | — | — | | | **[FACT]** **`[BETA-DEFERRED]`** | **day:** nothing customer-visible breaks; events queue (`settings.go:1466-1485`). **week:** the operator alarm plane is dark, no claim/reset codes, no config or floor convergence, no PBS-secret re-issue. **permanently:** escrow custody and break-glass credentials are gone (INV Part D2.4) | -| 12 | **Offsite provider lost (Hetzner)** | everything on-premises: both drives, both whole-guest tiers | none needed — on-premises recovery is unaffected. Re-provision a new offsite target. | operator | | | **[FACT]** **`[BETA-DEFERRED]`** | **restic and PBS share the provider** (INV Part E.1). Whether they share an account and payment method is **UNKNOWN** → §11-D | +| 12 | **Offsite provider lost (Hetzner)** | everything on-premises: both drives, both whole-guest tiers | none needed — on-premises recovery is unaffected. Re-provision a new offsite target. | operator | | | **[FACT]** **`[BETA-DEFERRED]`** | **restic and PBS share the provider** (INV Part E.1). Whether they share an account and payment method is **UNKNOWN** → §11-D **`[FACT]` 2026-10-03:** the PBS (whole-guest) off-site copy now also exists OFF Hetzner — DooPlex pulls ep0's datastore nightly (decision 70). The restic tier is still Hetzner-only. | | 13 | **Customer loses R** | every byte, every tier, the Recipe | **on-premises recovery is unaffected** — Tier-1/2/3 and the whole-guest tiers all work while the guest and host live. What is lost: the ability to re-establish **host identity** and to open the escrowed offsite key material after a host loss | — | | | **NONE for host-loss** | R exists in **zero** system copies by design (INV Part C, row 24). Whether the operator should be able to recover is **open** → §11-A, §11-B. **D5 narrows this further:** Tier-1/2 app recovery needs only the drive, so losing R now costs the offsite route and host identity, never local app recovery (§7.4) | | 14 | **Host SSH management plane dies** | everything | 3-layer break-glass: tmpfiles → 60 s watchdog → hub-vaulted `root@pam` on the PVE console | operator | | | **PROVEN** | `runbooks/break-glass.md`; the original incident and its fix | | 15 | **Interrupted offsite run leaves a stale lock** | everything | manual `restic unlock --remove-all` | **operator** (SSH) | | | **DEFECT** | the self-heal exists (`offbox.go:634-648`) but the probe fails first and `classifyResticProbe` has no lock case → fail-fast, operator told *„ismeretlen okból"* → **R-104** | diff --git a/documentation/audits/offsite-lock-build-2026-10-03/partB/hub-rollout.txt b/documentation/audits/offsite-lock-build-2026-10-03/partB/hub-rollout.txt new file mode 100644 index 00000000..ebc2e75e --- /dev/null +++ b/documentation/audits/offsite-lock-build-2026-10-03/partB/hub-rollout.txt @@ -0,0 +1,8 @@ +## hub v0.127.0 rollout 2026-10-03T14:59:40Z +startup: off-site secrets sealed at rest (4 legacy plaintext row(s) sealed now) +startup: Off-site key registrar enabled; daily key check at 07:10 Budapest +raw DB after (prefix, length only): + demo-felhom enc:v1: 83 + demo-hp enc:v1: 83 + tester-1 enc:v1: 83 + Tester-2 enc:v1: 83 diff --git a/documentation/audits/offsite-lock-build-2026-10-03/partC/red-proofs-controller.txt b/documentation/audits/offsite-lock-build-2026-10-03/partC/red-proofs-controller.txt index 3160e5e5..bfed55b8 100644 --- a/documentation/audits/offsite-lock-build-2026-10-03/partC/red-proofs-controller.txt +++ b/documentation/audits/offsite-lock-build-2026-10-03/partC/red-proofs-controller.txt @@ -19,3 +19,9 @@ ## restored: ok gitea.dooplex.hu/admin/felhom-controller/internal/backup 433.945s ok gitea.dooplex.hu/admin/felhom-controller/internal/offsiteapply (cached) +FAIL + +## RPC4: an unreadable count recorded as a measured zero (the v0.289.0 shape, live on demo-felhom) +=== RUN TestRunOffbox_UnreadableCountIsNotZero + offbox_window_test.go:266: an unreadable count was recorded as a measured zero: 0 +--- FAIL: TestRunOffbox_UnreadableCountIsNotZero (0.00s) diff --git a/documentation/audits/offsite-lock-build-2026-10-03/partD/demo-felhom-before.txt b/documentation/audits/offsite-lock-build-2026-10-03/partD/demo-felhom-before.txt new file mode 100644 index 00000000..58a39657 --- /dev/null +++ b/documentation/audits/offsite-lock-build-2026-10-03/partD/demo-felhom-before.txt @@ -0,0 +1,2 @@ +demo-felhom before migration (controller 0.288.0, sftp transport, u629488-sub1): 11 snapshots +after two chain runs on 0.289.0/0.289.1 (pinned): 13 snapshots diff --git a/documentation/audits/offsite-lock-build-2026-10-03/partD/demo-felhom-proofs.txt b/documentation/audits/offsite-lock-build-2026-10-03/partD/demo-felhom-proofs.txt new file mode 100644 index 00000000..c575b78e --- /dev/null +++ b/documentation/audits/offsite-lock-build-2026-10-03/partD/demo-felhom-proofs.txt @@ -0,0 +1,17 @@ +transport=rclone-pinned target=u629488-sub1@u629488-sub1.your-storagebox.de:/home/felhom-repo +## count before +13 snapshots +## restore: one file from the newest snapshot 5dd1f0b9 +restoring to /tmp/rt +restored files: 2 +sample: /mnt/sys_drive/felhom-data/backups/primary/opengist/data-stamps.json (398 bytes) +## check (exclusive lock — the weekly integrity job's op) + +no errors were found +## delete attempt through the box's key: forget 6ea85413 (the OLDEST) +unable to remove from the repository +[0:48] 0.00% 0 / 1 files deleted +blob not removed, server response: 403 Forbidden (403) +rc-line done +## count after +13 snapshots diff --git a/documentation/audits/offsite-lock-build-2026-10-03/partD/demo-hp-before.txt b/documentation/audits/offsite-lock-build-2026-10-03/partD/demo-hp-before.txt new file mode 100644 index 00000000..43208a79 --- /dev/null +++ b/documentation/audits/offsite-lock-build-2026-10-03/partD/demo-hp-before.txt @@ -0,0 +1,4 @@ +gitea.dooplex.hu/admin/felhom-controller:0.288.0 +offbox files: applied_marker known_hosts repo_password ssh_key +target: u629488-sub3@u629488-sub3.your-storagebox.de port=23 repo=/home/felhom-repo transport=sftp settings.snapshot_count= +91 snapshots diff --git a/documentation/audits/offsite-lock-build-2026-10-03/partD/demo-hp-migration.txt b/documentation/audits/offsite-lock-build-2026-10-03/partD/demo-hp-migration.txt new file mode 100644 index 00000000..7361e9d4 --- /dev/null +++ b/documentation/audits/offsite-lock-build-2026-10-03/partD/demo-hp-migration.txt @@ -0,0 +1,8 @@ +gitea.dooplex.hu/admin/felhom-controller:0.289.1 +gitea.dooplex.hu/admin/felhom-controller:0.289.1 Up About a minute (healthy) +2026/10/03 15:19:24 offsiteapply.go:148: [INFO] [offsite-apply] settle-gate: awaiting floor knowledge (first report ACK) before offsite apply +2026/10/03 15:19:34 offsiteapply.go:148: [INFO] [offsite-apply] settle-gate: GO — at/above floor 0.288.0 (we are 0.289.1), no managed update running +2026/10/03 15:19:37 offsiteapply.go:148: [INFO] [offsite-apply] the hub installed key SHA256:16GpHoTF0WHLbj15Pog0mpmDMit4/BLFLRfIqmHJQc8 append-only on u629488-sub3@u629488-sub3.your-storagebox.de (fresh=false) +2026/10/03 15:19:37 offsiteapply.go:148: [INFO] [offsite-apply] offsite configured append-only for u629488-sub3@u629488-sub3.your-storagebox.de:/home/felhom-repo (key SHA256:16GpHoTF0WHLbj15Pog0mpmDMit4/BLFLRfIqmHJQc8) +2026/10/03 17:19:36 [INFO] offsitekeys: installed box key SHA256:16GpHoTF0WHLbj15Pog0mpmDMit4/BLFLRfIqmHJQc8 for demo-hp pinned append-only (u629488-sub3@u629488-sub3.your-storagebox.de, dropped 5 unpinned line(s)) in 938ms +2026/10/03 17:19:37 [INFO] offsitekeys: box confirmed key SHA256:16GpHoTF0WHLbj15Pog0mpmDMit4/BLFLRfIqmHJQc8 for demo-hp; 0 other line(s) removed diff --git a/documentation/audits/offsite-lock-build-2026-10-03/partD/demo-hp-proofs.txt b/documentation/audits/offsite-lock-build-2026-10-03/partD/demo-hp-proofs.txt new file mode 100644 index 00000000..872f7bed --- /dev/null +++ b/documentation/audits/offsite-lock-build-2026-10-03/partD/demo-hp-proofs.txt @@ -0,0 +1,24 @@ +2026/10/03 15:22:53 night_chain.go:77: [INFO] [night-chain] offsite: started +2026/10/03 15:22:53 offbox.go:991: [INFO] [offbox] backup run started (9 app(s) toggled) +2026/10/03 15:24:49 offbox.go:1466: [WARN] [offbox] backed up bentopdf (/mnt/sys_drive/felhom-data/backups/primary/bentopdf, 0 mandatory path(s)) — but the recovery unit carried NO database dump and NO volume tar, so this snapshot holds none of the app's data; the next run with a dump leg will replace it +2026/10/03 15:25:12 offbox_window.go:178: [INFO] [offbox] retention skipped (after-run): no clean-up window now (weekly windows are off) — nothing deleted (decision 68) +2026/10/03 15:25:15 offbox.go:1234: [INFO] [offbox] backup OK: 9 app(s) backed up, 100 snapshot(s), 2m19s +2026/10/03 15:25:15 night_chain.go:82: [INFO] [night-chain] offsite: done in 2m23s +2026/10/03 15:25:15 night_chain.go:93: [INFO] [night-chain] finished in 4m14s +transport=rclone-pinned target=u629488-sub3@u629488-sub3.your-storagebox.de:/home/felhom-repo +## count before +100 snapshots +## restore: one file from the newest snapshot c43f2d0e +restoring to /tmp/rt +restored files: 2 +sample: /mnt/sys_drive/felhom-data/backups/primary/opengist/manifest.json (1767 bytes) +## check (exclusive lock — the weekly integrity job's op) + +no errors were found +## delete attempt through the box's key: forget 05c3346a (the OLDEST) +unable to remove from the repository +[0:48] 0.00% 0 / 1 files deleted +blob not removed, server response: 403 Forbidden (403) +rc-line done +## count after +100 snapshots diff --git a/documentation/audits/offsite-lock-build-2026-10-03/partD/key-audit-1.txt b/documentation/audits/offsite-lock-build-2026-10-03/partD/key-audit-1.txt new file mode 100644 index 00000000..b1a37cae --- /dev/null +++ b/documentation/audits/offsite-lock-build-2026-10-03/partD/key-audit-1.txt @@ -0,0 +1,40 @@ +## POST /offsite/key-audit (operator, Basic auth) 2026-10-03T15:26:40Z +[ + { + "customer": "Tester-2", + "lines": 0, + "pinned": 0, + "findings": null + }, + { + "customer": "demo-felhom", + "lines": 1, + "pinned": 1, + "findings": null + }, + { + "customer": "demo-hp", + "lines": 1, + "pinned": 1, + "findings": null + }, + { + "customer": "tester-1", + "lines": 3, + "pinned": 0, + "findings": [ + { + "Fingerprint": "SHA256:3UoXpMIvo9gat9AL1380UK39UGAtO2T2CBBllo8n3Mg", + "Kind": "unpinned" + }, + { + "Fingerprint": "SHA256:Jri1gf2AGHCTj8cJrFSpRj8+BxdY64OQ5Rnms/tbOZI", + "Kind": "unpinned" + }, + { + "Fingerprint": "SHA256:gAhxqeDkxAPTOOeKQDHBDNe/AYpARkxI8h7Rrc8pLD8", + "Kind": "unpinned" + } + ] + } +] diff --git a/documentation/audits/offsite-lock-build-2026-10-03/partD/no-password-for-box.txt b/documentation/audits/offsite-lock-build-2026-10-03/partD/no-password-for-box.txt new file mode 100644 index 00000000..19abd29a --- /dev/null +++ b/documentation/audits/offsite-lock-build-2026-10-03/partD/no-password-for-box.txt @@ -0,0 +1,5 @@ +hub=https://hub.felhom.eu keylen=64 +consume-password HTTP 410 +body: gone: the hub no longer serves the storage password; register the box's public key at /api/v1/offsite/register-key/ +lines mentioning password field: 1 +2026/10/03 17:26:59 [WARN] offsite consume-password called by demo-hp — retired (decision 69); the box must register its public key (controller >= 0.289.0) diff --git a/documentation/audits/offsite-lock-build-2026-10-03/partE/live-window-demo-hp.txt b/documentation/audits/offsite-lock-build-2026-10-03/partE/live-window-demo-hp.txt new file mode 100644 index 00000000..8b7ae2f6 --- /dev/null +++ b/documentation/audits/offsite-lock-build-2026-10-03/partE/live-window-demo-hp.txt @@ -0,0 +1,22 @@ +grant: {"ok":true} +HTTP 202 +2026/10/03 15:22:24 controller_image_retention.go:184: [INFO] [stacks] controller image retention: deleted gitea.dooplex.hu/admin/felhom-controller:0.287.0 (244595237106, 409MB) — older than the previous controller and no container uses it (decision 56) +2026/10/03 15:22:24 controller_image_retention.go:115: [INFO] [stacks] controller image retention: pass over 3 controller image(s) — running 0.289.1, previous "0.288.0" (by version order (no swap record names one present)), 1 candidate(s), 1 deleted, the rest kept +2026/10/03 15:25:12 offbox_window.go:178: [INFO] [offbox] retention skipped (after-run): no clean-up window now (weekly windows are off) — nothing deleted (decision 68) +2026/10/03 15:25:15 offbox.go:1234: [INFO] [offbox] backup OK: 9 app(s) backed up, 100 snapshot(s), 2m19s +2026/10/03 15:22:24 controller_image_retention.go:184: [INFO] [stacks] controller image retention: deleted gitea.dooplex.hu/admin/felhom-controller:0.287.0 (244595237106, 409MB) — older than the previous controller and no container uses it (decision 56) +2026/10/03 15:22:24 controller_image_retention.go:115: [INFO] [stacks] controller image retention: pass over 3 controller image(s) — running 0.289.1, previous "0.288.0" (by version order (no swap record names one present)), 1 candidate(s), 1 deleted, the rest kept +2026/10/03 15:25:12 offbox_window.go:178: [INFO] [offbox] retention skipped (after-run): no clean-up window now (weekly windows are off) — nothing deleted (decision 68) +2026/10/03 15:25:15 offbox.go:1234: [INFO] [offbox] backup OK: 9 app(s) backed up, 100 snapshot(s), 2m19s +2026/10/03 15:25:15 night_chain.go:93: [INFO] [night-chain] finished in 4m14s +2026/10/03 15:32:05 offbox_window.go:199: [ERROR] [offbox] clean-up window 1: the fake-snapshot guard REFUSED — nothing deleted: the policy would remove snapshot c6b5c67b from 2026-10-03T15:24:51Z — younger than 8 days, which honest retention never does (R-822) +2026/10/03 15:32:10 offbox.go:1234: [INFO] [offbox] backup OK: 9 app(s) backed up, 109 snapshot(s), 2m22s +2026/10/03 15:32:10 night_chain.go:93: [INFO] [night-chain] finished in 4m18s +2026/10/03 17:32:03 [WARN] offsitekeys: clean-up window 1 OPENED for demo-hp (key SHA256:16GpHoTF0WHLbj15Pog0mpmDMit4/BLFLRfIqmHJQc8, 109 snapshot(s), max 43 removed, closes by 2026-10-03T15:52:03Z, one-shot=true) +2026/10/03 17:32:06 [INFO] offsitekeys: clean-up window 1 CLOSED for demo-hp: outcome=guard-refused, 109 -> 109 (drop 0, allowed 43) +2026/10/03 17:32:06 [INFO] Operator email sent for demo-hp/offsite_prune_guard_refused +## key check after window 1 2026-10-03T15:32:32Z +Tester-2 lines 0 pinned 0 findings 0 +demo-felhom lines 1 pinned 1 findings 0 +demo-hp lines 1 pinned 1 findings 0 +tester-1 lines 3 pinned 0 findings 3 diff --git a/documentation/audits/offsite-lock-build-2026-10-03/partF/copy-listing.txt b/documentation/audits/offsite-lock-build-2026-10-03/partF/copy-listing.txt new file mode 100644 index 00000000..596d30d8 --- /dev/null +++ b/documentation/audits/offsite-lock-build-2026-10-03/partF/copy-listing.txt @@ -0,0 +1,21 @@ +## copy contents per namespace (filesystem listing, read-only) +ns/demo-felhom/ct/9201/2026-09-22T04:12:20Z +ns/demo-felhom/ct/9201/2026-09-29T04:16:43Z +ns/demo-hp/ct/9201/2026-09-24T20:06:25Z +ns/demo-hp/ct/9201/2026-10-01T20:15:29Z + +## one snapshot's archives (no decryption — the copy is ciphertext; the catalog needs the customer's key) + +4096 . +4096 .. +5576 catalog.pcat1.didx +1220 client.log.blob +690 index.json.blob +401 pct.conf.blob +179176 root.pxar.didx + +## same view on ep0 (source) +ns/demo-felhom/ct/9201/2026-09-22T04:12:20Z +ns/demo-felhom/ct/9201/2026-09-29T04:16:43Z +ns/demo-hp/ct/9201/2026-09-24T20:06:25Z +ns/demo-hp/ct/9201/2026-10-01T20:15:29Z diff --git a/documentation/audits/offsite-lock-build-2026-10-03/partF/first-sync.txt b/documentation/audits/offsite-lock-build-2026-10-03/partF/first-sync.txt new file mode 100644 index 00000000..87c9e59c --- /dev/null +++ b/documentation/audits/offsite-lock-build-2026-10-03/partF/first-sync.txt @@ -0,0 +1,29 @@ +## first pull 2026-10-03T19:13:17Z +---- +Syncing datastore 'felhom-offsite', namespace 'demo-felhom' into datastore 'ep0-copy', namespace 'demo-felhom' +Created namespace demo-felhom +Found 1 groups to sync (out of 1 total) +[ct/9201]: 2026-09-22T04:12:20Z: start sync +[ct/9201]: 2026-09-22T04:12:20Z/pct.conf.blob: sync archive +[ct/9201]: 2026-09-22T04:12:20Z/root.pxar.didx: sync archive +[ct/9201]: 2026-09-22T04:12:20Z/root.pxar.didx: downloaded 2.057 GiB (62.788 MiB/s) +[ct/9201]: 2026-09-22T04:12:20Z/catalog.pcat1.didx: sync archive +[ct/9201]: 2026-09-22T04:12:20Z/catalog.pcat1.didx: downloaded 652.916 KiB (10.225 MiB/s) +[ct/9201]: Snapshot ct/9201/2026-09-22T04:12:20Z: got backup log file client.log.blob +[ct/9201]: 2026-09-22T04:12:20Z: sync done +[ct/9201]: percentage done: 50.00% (1/2 snapshots) +[ct/9201]: 2026-09-29T04:16:43Z: start sync +[ct/9201]: 2026-09-29T04:16:43Z/pct.conf.blob: sync archive +[ct/9201]: 2026-09-29T04:16:43Z/root.pxar.didx: sync archive +[ct/9201]: 2026-09-29T04:16:43Z/root.pxar.didx: downloaded 641.072 MiB (61.537 MiB/s) +[ct/9201]: 2026-09-29T04:16:43Z/catalog.pcat1.didx: sync archive +[ct/9201]: 2026-09-29T04:16:43Z/catalog.pcat1.didx: downloaded 768.894 KiB (16.864 MiB/s) +[ct/9201]: Snapshot ct/9201/2026-09-29T04:16:43Z: got backup log file client.log.blob +[ct/9201]: 2026-09-29T04:16:43Z: sync done +[ct/9201]: percentage done: 100.00% (2/2 snapshots) +Finished syncing namespace demo-felhom, current progress: 2 groups, 0 snapshots +pull datastore 'ep0-copy' end +TASK OK +elapsed 201 s +## bytes on DooPlex +12G /mnt/5_hdd/backup/ep0-copy diff --git a/documentation/audits/offsite-lock-build-2026-10-03/partF/jobs.txt b/documentation/audits/offsite-lock-build-2026-10-03/partF/jobs.txt new file mode 100644 index 00000000..16daa4f1 --- /dev/null +++ b/documentation/audits/offsite-lock-build-2026-10-03/partF/jobs.txt @@ -0,0 +1,10 @@ ++====================+================+==========+========+================+==========+==============+=========+=================================================================================+ +| id | sync-direction | store | remote | remote-store | schedule | group-filter | rate-in | comment | ++====================+================+==========+========+================+==========+==============+=========+=================================================================================+ +| ep0-felhom-offsite | | ep0-copy | ep0 | felhom-offsite | 05:00 | all | | decision 70: nightly copy of ep0 felhom-offsite; never removes what ep0 removed | ++====================+================+==========+========+================+==========+==============+=========+=================================================================================+ ++=================+==========+===========+=================+================+============================================+ +| id | store | schedule | ignore-verified | outdated-after | comment | ++=================+==========+===========+=================+================+============================================+ +| verify-ep0-copy | ep0-copy | sat 06:30 | 1 | 30 | decision 70: weekly verify of the ep0 copy | ++=================+==========+===========+=================+================+============================================+ diff --git a/documentation/audits/offsite-lock-build-2026-10-03/partF/setup.txt b/documentation/audits/offsite-lock-build-2026-10-03/partF/setup.txt new file mode 100644 index 00000000..cd28c849 --- /dev/null +++ b/documentation/audits/offsite-lock-build-2026-10-03/partF/setup.txt @@ -0,0 +1,37 @@ +## tunnel unit +[Unit] +Description=Felhom: SSH tunnel DooPlex 127.0.0.1:18007 -> ep0 PBS 127.0.0.1:8007 (decision 70, nightly pull-sync of felhom-offsite) +Documentation=file:///mnt/5_hdd/felhom.eu/git/felhom.eu/documentation/runbooks/ep0-datastore-copy.md +After=network-online.target +Wants=network-online.target + +[Service] +User=kisfenyo +ExecStart=/usr/bin/ssh -N -o BatchMode=yes -o ExitOnForwardFailure=yes -o ServerAliveInterval=30 -o ServerAliveCountMax=3 -L 127.0.0.1:18007:127.0.0.1:8007 root@167.233.158.164 +Restart=always +RestartSec=30 + +[Install] +WantedBy=multi-user.target +active + +## notification target + matcher (password lives in notifications-priv.cfg, root:root 0600, not shown) +smtp: felhom-operator + author Felhom DooPlex PBS + comment decision 70: ep0 copy job failures to the operator (Resend) + from-address monitoring@felhom.eu + mailto admin@felhom.eu + mode tls + port 465 + server smtp.resend.com + username resend + +matcher: felhom-operator-errors + comment decision 70: any error (the ep0 pull-sync and verify jobs) reaches the operator + match-severity error + mode all + target felhom-operator + +## test mail: received in the operator mailbox 2026-10-03T19:17:41Z, subject 'Test notification', from monitoring@felhom.eu to admin@felhom.eu (Gmail connector search) + +## ep0 side: token root@pam!dooplex-sync, ACL DatastoreReader on /datastore/felhom-offsite (propagate) — the only change on ep0 diff --git a/documentation/backlog/CLOSED-ITEMS.md b/documentation/backlog/CLOSED-ITEMS.md index b53d9fd2..9332e787 100644 --- a/documentation/backlog/CLOSED-ITEMS.md +++ b/documentation/backlog/CLOSED-ITEMS.md @@ -26,6 +26,19 @@ --- +## 2026-10-03 (evening) — the off-site lock built (hub v0.127.0, controller v0.289.0/0.289.1, decisions 68–70) + +> Evidence: `audits/offsite-lock-build-2026-10-03/`. + +| Row | What | Closed | Evidence | +|---|---|---|---| +| **R-820** | **A box could obtain its sub-account password at will (the hub's self-heal re-armed it; the box consumed it) — and that password removes any append-only pin.** Fixed: the box sends only its PUBLIC key; the hub (key registrar) writes it pinned; `consume-password` answers 410. Measured live: demo-hp's own API key gets `410 gone`, no password. **Reasoning kept: a pinned key protects nothing while any route can rewrite `authorized_keys` — the hub is now that file's only writer, and it reads every file daily.** | CLOSED 2026-10-03 — FIXED hub v0.127.0 + controller v0.289.0/0.289.1, proven live on both demo boxes | `partD/no-password-for-box.txt`, `partB/red-proofs-hub.txt`; hub `internal/offsitekeys` | +| **R-821** | **The hub DB held every sub-account password in the clear.** Fixed: AES-256-GCM at rest under `OFFSITE_SECRET_KEY` (Secret/offsite-secret-key, out of git); the 4 legacy rows sealed at start-up (raw rows read back `enc:v1:`); no key → the hub refuses to store or use one. **Reasoning kept: the running hub still holds the key and can open the passwords — a hub compromise remains an off-site compromise; this closes the database-copy route only.** | CLOSED 2026-10-03 — FIXED hub v0.127.0, verified on the live DB | `partB/hub-rollout.txt`; `TestOffsiteSecret_*` | +| **R-342** | **ep0's server snapshot never covered `/mnt/pbs-datastore`.** Built (decision 70): nightly PBS pull-sync to DooPlex (`ep0-copy`), `remove-vanished false`, weekly verify, failures mailed (test mail received). First pull 201 s / 12 GB / 4 of 4 snapshots, matching ep0. ep0 changed by one read-only token only; reached through an SSH forward from DooPlex (operator ruling). **Reasoning kept: Hetzner cannot snapshot a Volume — the copy is the only safeguard; never prune it tighter than ep0.** | CLOSED 2026-10-03 — BUILT; restore route unwalked (R-830), growth unbounded (R-828) | `partF/`; `runbooks/ep0-datastore-copy.md` | +| **R-825** | **Controller v0.289.0 reported 0 off-site snapshots as MEASURED over a store holding 12 (demo-felhom), and the hub mailed `offsite_snapshots_dropped` 11→0 — a false alarm.** Cause: the provider's rclone prints a NOTICE line that restic forwards into the combined output; every `--json` parse failed. Fixed in v0.289.1 within 15 minutes: the notice is stripped, and an unreadable count is never a measured zero (keeps the last value, `stats_known=false`). **Reasoning kept: a failed measurement must never be written as a measurement (R-331) — the detector that caught it is the one it would have blinded.** | CLOSED 2026-10-03 — FIXED controller v0.289.1 (red-proved), found and fixed in-session | operator mail 2026-10-03 17:17 CEST; `partC/red-proofs-controller.txt` RPC4 | + +--- + ## 2026-10-03 — off-site append-only, measured on the provider (R-436, R-430) > Spike, no product change. Evidence and design: `audits/offsite-append-only-2026-10-03/`. diff --git a/documentation/backlog/OPEN-ITEMS.md b/documentation/backlog/OPEN-ITEMS.md index e471b57e..c172ff6e 100644 --- a/documentation/backlog/OPEN-ITEMS.md +++ b/documentation/backlog/OPEN-ITEMS.md @@ -191,22 +191,22 @@ stopping line that lies. | **R-687** | App updates | P4 | **[P3-LOW] Part 7's live proof has four gaps a scratch box cannot close, and one observability gap.** (1) **W+5h reached with steps left** is proven by unit test only (`TestLeg_NoStepAtOrAfterW5h`) — the leg starts at W+105m and would need a 3-hour leg live; (2) **the off-site leg FAILING** before the update leg: 9202 has no off-site tier, so only the "no target" path ran live — failure and panic are `TestChainUpdateLeg_EveryPath`; (3) **a `files_may_change` step WITHOUT a whole copy**: both drill apps given the mark (wishlist, navidrome, romm) turned out whole on 9202 by the backup side's truth table (why, per app, is not logged — see the observability gap), so only "with a copy" ran live; (4) **the full-system gate waiting** cannot run on 9202 (no agent), and **did not occur on the demo boxes' real night either** (2026-09-25: both legs ended by 04:19, before the gate opened at 04:30, and no whole-box backup was due on either box) — unit + red-proof only (`TestD20_GateWaitsForTheLeg`). **Also cosmetic:** a leg with no steps reports `"steps": null` to the hub, not `[]`. **Observability:** when the leg TAKES a `files_may_change` step it does not log which whole copy allowed it (only the skip says why). `audits/night-2026-09-25/C/` **-- NARROWED 2026-09-25 (controller v0.273.0):** the cosmetic `"steps": null` → `[]` and the taken `files_may_change` step's missing log line are FIXED (red-proofed, `audits/night-2026-09-26/F/`). Items (1)–(4) stay; (4) did not occur on 2026-09-25 either (demo-felhom's whole-box backup ran at 07:29, three hours after its leg; demo-hp had none due). **-- 2026-09-28 (night 27/28):** (4) did not occur again — on demo-hp the leg ended 04:23:54 and the whole-guest backup began 04:37:06, after the gate opened at 04:30; demo-felhom's backup ran at 07:36 (`audits/evidence-golden-0276-2026-09-28/phaseD2-night-read.txt`). **-- 2026-09-30 (by day, demo-hp 9201): item (4) PROVEN LIVE.** The night chain pressed by hand, the window moved to W = now − 2h05m the moment the leg started, `quiesce.poll_interval` 1m: `[quiesce] full-system backup due and inside its window, but the automatic update leg is running … deferring` at 11:35:11 and 11:36:11 UTC while bookstack (55.1 s) and kimai (75.1 s) stepped; the leg's end line at 11:36:29; the backup quiesced at 11:37:11 (the first poll after), job done 11:47:19, the agent's `backup: completed` 9.98 GB. Config and window put back and read back (`audits/pg-last-six-2026-09-30/C/`). **Found, cosmetic, manual chain only:** the deferral names the moved window's W+5h (16:29) while the manual leg's own deadline was its start + the leg length (16:49). | **OPEN — P3, gaps (1)–(3) + the manual-chain deferral text; item (4) proven live 2026-09-30; owner: CC** **Re-ranked 2026-10-03: P3→P4: the gaps are covered by unit tests; left is live-proof completeness and one log text.** | — | — | CC | | **R-734** | App updates | P4 | **[P3-LOW] The harness marks immich `files_may_change` because immich rewrites six 13-byte `.immich` folder markers at every start.** MEASURED 2026-09-30 on the bench (v3.2.2 → v3.2.4): the bind-tree hash of `appdata/immich` changed; the only changed files were `{encoded-video,library,backups,profile,thumbs,upload}/.immich`, rewritten at each start — no household file. The mark is honest by the harness's rule and the ladder writer copies it (never edited by hand), so immich's v3.2.4 night step needs a fresh WHOLE copy (decision 13); on a box without one the night leg skips it and a person presses. The 2026-09-23 immich entry did not carry it (`files_changed []`). **Needs:** a decision whether app-owned marker files are excluded from the file hash (a per-template ignore list, or a size/name rule), or the mark stays. **-- 2026-09-30 (evening):** the harness now NAMES the files behind the mark (`files_changed_detail`, catalog `5b1972b`); on immich's step `0b82…` re-proof it named exactly the six `.immich` markers again. calibre-web's step (v4.0.6 → v4.0.8) carries the mark too, and the named files are its LIBRARY DATABASE: `media/books/metadata.db`, `metadata.db-shm`, `metadata.db-wal` changed; the book file did not (bench names re-run, `A/calibre-names/`). That is household data (the template's backup class `mandatory` holds DB and books as one unit), so the mark is right there and the night leg takes that step only with a fresh whole copy. On 9202 the same step changed no file in that folder (per-file hashes before/after) — not explained. `audits/more-night-apps-2026-09-30/` | **READY — rank P3-LOW; owner: CC (harness); the rule change needs a word** **Re-ranked 2026-10-03: P3→P4: the effect is an update that waits for a press; no data risk.** | — | — | CC + operator | -## Backup & restore — 55 rows (P2 12, P3 23, P4 20) +## Backup & restore — 58 rows (P2 12, P3 26, P4 20) | ID | Category | Sev | What | State | Blocked on | Next action | Owner | |---|---|---|---|---|---|---|---| | **R-32** | Backup & restore | P2 | **[P2-HIGH] RESET must purge the customer base dir; the orphan card must stay honest; unattributed bytes must be visible.** The rehearsal's S7 said in advance that an orphan card would BE a finding — and one appeared (16:58:14). Cause: RESET's `"hetzner":"ok"` leg destroys the sub-account, but **a Hetzner sub-account is an access-control object, not a data object** — its directory survives, so re-enabling offsite recreated an account over the previous lifecycle's ciphertext, encrypted under a key that same RESET had destroyed. **MIGRATED FROM `ROADMAP.md` 2026-08-22 (R-369) — originally filed 2026-07-21, size M, roadmap state `idea`.** Moved verbatim; nothing added or reinterpreted. The roadmap keeps its copy as history, marked moved. | **OPEN — migrated from ROADMAP 2026-08-22, rank unchanged** | — | **Ruling from the run (three parts, deliberately separate):** (1) because RESET destroys custody, the ciphertext it leaves behind is unrecoverable **BY DESIGN** → RESET gains a **main-account purge of the customer base dir** (the existing operator ack already covers it); (2) the **move-aside guard STAYS** for reinstall-*without*-RESET — there custody survives and the card's "history recoverable" promise is true (R-26 depends on exactly that); (3) the operator **Restic tab shows per-customer directory bytes vs attributed snapshot bytes**, so dead data cannot hide. Measured on the pool box that night: **49 M attributed** (2 snapshots, 48.717 MiB) against **1.4 G + 3.0 M unattributed** across TWO `.orphaned-*` dirs. Evidence `restic-and-pool.txt` | CC | -| **R-95** | Backup & restore | P2 | restic offsite credential **can delete** (`readonly=False`, `forget --prune` runs from the box); SFTP cannot express append-only | **READY** **MEASURED 2026-10-03 (R-436 CLOSED) — append-only HOLDS for a pinned key; it is NOT yet protection.** On the provider (`u629488-sub4`, scratch repo, removed): a key pinned to `command="rclone serve restic --stdio --append-only ",restrict` backs up, restores (bytes identical) and checks; `forget --prune`, `forget --keep-last 1` and a real `prune` are refused with `blob not removed, server response: 403 Forbidden (403)`; the unpinned control deletes. **What still defeats it:** the sub-account PASSWORD logs in on ports 22 and 23 and rewrote `authorized_keys` in this session, and a box can obtain that password from the hub at will (R-820). **Also found:** an add-only key can plant future-dated snapshots that make the box's own retention policy select every real snapshot (R-822); the hub holds every sub-account password in the clear (R-821); FOUR box features need a deleting login, not two — both `forget` sites, the orphan move-aside (`mv`) and the abandonment (`rm -rf`). **PROPOSAL (operator decides — STATUS):** the hub becomes the key registrar so the box never sees the password; switch boxes to the pinned key with no box-side retention first (quota headroom is months); then a weekly hub-opened clean-up window with a poisoning guard. `audits/offsite-append-only-2026-10-03/DESIGN.md`. The rank stays the operator's. | — | Root exposure still open. Mitigation now ARMED — split prune off-box or move to REST `--append-only` **-- UNBLOCKED 2026-09-22, not closed.** The premise in this row's own title - *SFTP cannot express append-only* - is answered by Hetzner (ticket #2026090103040671, recorded in R-433 and R-436): on a Storage Box the `rclone serve restic --stdio --append-only` backend can be pinned to an SSH key as a FORCED COMMAND, so append-only IS expressible without a new machine. **The row stays open because nothing has been measured**: no forced-command key exists on `u629488`, no `forget --prune` has been refused through one, and the protection only holds if EVERY key that can reach the repository carries the prefix. The next step is R-436's, and it is small. | CC **SPIKE 2026-09-01 — `audits/SPIKE-r95-offsite-delete-2026-09-01.md`. THE WORD "ARMED" ABOVE IS NOT SUPPORTED AND IS WITHDRAWN PENDING R-429:** no `.snapshots` is visible to either box's sub-account (measured, both machines, with controls), so the seven-day bound is unverified and unverifiable from the product side. **Q3 (documented, `hub/internal/hetznerapi/hetznerapi.go:38-45`): the sub-account API has ONE permission axis, `readonly` — there is no append-only, so the PBS shape does NOT transfer** (PBS is a server that can refuse; a Storage Box is a filesystem that runs nothing). **Q5 (measured): withdrawing delete does NOT wedge the store** — restic treats a dead owner's lock as stale and proceeds — so the constraint everyone feared is not the blocker; but `unlock --remove-all` lies about success (R-430), and the crash-lock window is UNKNOWN. **Q6 (measured): restic 0.14.0 DOES speak `rest:`** (control: `banana:` → invalid backend), and append-only is a rest-server flag, not a restic one — reachable, but it needs a machine in the recovery path and ep0 is protected. **Q7 (measured): detection is nearly free** — `snapshot_count` already reaches the hub and the hub APPENDS reports, so the comparison needs no box change. **RECOMMENDATION: answer R-429 first (Viktor, ten minutes), then build detection, then move retention off the box; defer the transport change.** Two `forget --prune` sites must be disarmed together — `offbox.go:1388` AND `offbox.go:1759` — or R-191 repeats. **RE-SCOPED 2026-09-01 — THE STORY WAS WORSE THAN THE TRUTH FOR TWO MONTHS.** The box can delete its own LIVE repository, **but it cannot write to the daily snapshots of it** — MEASURED, not cited: a write into `/.zfs/snapshot` is refused on both boxes while the same write to the account home succeeds (R-432). Seven daily snapshots are confirmed in the panel (R-429). **So a deletion costs at most the data written since the last daily snapshot, and the rest is recoverable — file by file, one customer at a time, with no effect on anyone else** (vendor: *"You can download individual files or entire directories as usual"*; *"It is not possible to write to the `/.zfs` directory or its subfolder"*). **NOT open-ended loss.** Two caveats kept honest: a panel-driven snapshot restore rolls back the WHOLE Storage Box and deletes newer snapshots, which is why the per-file route matters; and per-file recovery is **operator-only today** (R-432). **DETECTION SHIPPED hub v0.111.0 (R-431)** — an unexplained fall is noticed within a day. **THE RANKING IS VIKTOR'S:** this has been #1 since July on the old story. On the new facts I would rank it below the items that can still lose data outright, but I am not re-ranking it myself. `audits/SPIKE-r95-offsite-delete-2026-09-01.md` **DRILL 2026-09-01, LATER THE SAME DAY — THE RE-SCOPE'S SECOND HALF IS WITHDRAWN.** The drill that was to walk the recovery found there is no route to walk: **no snapshot is reachable from a sub-account by ANY name** (R-433) — 777,600 exact names in the vendor format over nine days, zero hits, with a passing control, plus the structural reason (`/home` st_dev 0,82 vs `/.zfs/snapshot` st_dev 0,276, and `/home/.zfs` absent). **Clause (a) stands: the box can delete the live repo and cannot write into the snapshot area.** Clause (b) — *"the rest is recoverable file by file"* — is NOT SUPPORTED. **The drill was STOPPED before its destructive phase on the operator's ruling**, because with no recovery route the deletion would have destroyed real history to buy only an alarm test that could not fire at the specified size (R-435). Nothing was deleted; the store is verified untouched at 69 snapshots. **So the comfort that lowered this row rested on an unwalked route, and the route does not exist.** Two new leads decide what happens next: **R-433** (can the MAIN account see them? nobody here holds that credential) and **R-436** (`rclone serve restic --stdio` is offered server-side and restic speaks `rclone:` — measured — which could make real prevention cheap, IF the provider pins `--append-only`). **The rank stays Viktor's. Plainly: the argument that moved this row down is the argument the drill removed.** `audits/evidence-drill-r95-recovery-2026-09-01/` **BLOCKED-ON-PROVIDER 2026-09-01, and the record of the demotion is kept deliberately: this row spent ONE DAY demoted on a clause that did not hold.** It was re-scoped down on the morning of 2026-09-01 on the strength of *"recoverable file by file"*, and that clause was measured false the same afternoon (R-433). **On today's evidence it belongs back near the top — that is a proposal, not an action; CC has not re-ranked it and will not.** Both questions that can settle it are drafted in `documentation/runbooks/provider-questions-2026-09-01.md`: **Q1** decides how urgent this is, **Q2** (R-436) could remove the root cause cheaply. **Precondition on any build here: R-430**, which is harmless only while the box can still delete. | +| **R-95** | Backup & restore | P2 | restic offsite credential **can delete** (`readonly=False`, `forget --prune` runs from the box); SFTP cannot express append-only | **NARROWED — BUILT 2026-10-03 (hub v0.127.0, controller v0.289.1): the box can no longer delete; weekly windows OFF; a real prune inside a window not yet proven (R-824)** **MEASURED 2026-10-03 (R-436 CLOSED) — append-only HOLDS for a pinned key; it is NOT yet protection.** On the provider (`u629488-sub4`, scratch repo, removed): a key pinned to `command="rclone serve restic --stdio --append-only ",restrict` backs up, restores (bytes identical) and checks; `forget --prune`, `forget --keep-last 1` and a real `prune` are refused with `blob not removed, server response: 403 Forbidden (403)`; the unpinned control deletes. **What still defeats it:** the sub-account PASSWORD logs in on ports 22 and 23 and rewrote `authorized_keys` in this session, and a box can obtain that password from the hub at will (R-820). **Also found:** an add-only key can plant future-dated snapshots that make the box's own retention policy select every real snapshot (R-822); the hub holds every sub-account password in the clear (R-821); FOUR box features need a deleting login, not two — both `forget` sites, the orphan move-aside (`mv`) and the abandonment (`rm -rf`). **PROPOSAL (operator decides — STATUS):** the hub becomes the key registrar so the box never sees the password; switch boxes to the pinned key with no box-side retention first (quota headroom is months); then a weekly hub-opened clean-up window with a poisoning guard. `audits/offsite-append-only-2026-10-03/DESIGN.md`. The rank stays the operator's. **BUILT 2026-10-03 (decisions 68–69).** Both demo boxes migrated to a hub-pinned append-only key with history kept (11→13, 91→100); a delete from each box is refused (403); the old password endpoint answers 410 to demo-hp's own key; the daily key check is live; the clean-up window opened, refused by the guard and closed live on demo-hp. Left: R-824 (the guard refuses after any manual run, so no real prune has run), R-822 residual, R-823. `audits/offsite-lock-build-2026-10-03/` | — | Root exposure still open. Mitigation now ARMED — split prune off-box or move to REST `--append-only` **-- UNBLOCKED 2026-09-22, not closed.** The premise in this row's own title - *SFTP cannot express append-only* - is answered by Hetzner (ticket #2026090103040671, recorded in R-433 and R-436): on a Storage Box the `rclone serve restic --stdio --append-only` backend can be pinned to an SSH key as a FORCED COMMAND, so append-only IS expressible without a new machine. **The row stays open because nothing has been measured**: no forced-command key exists on `u629488`, no `forget --prune` has been refused through one, and the protection only holds if EVERY key that can reach the repository carries the prefix. The next step is R-436's, and it is small. | CC **SPIKE 2026-09-01 — `audits/SPIKE-r95-offsite-delete-2026-09-01.md`. THE WORD "ARMED" ABOVE IS NOT SUPPORTED AND IS WITHDRAWN PENDING R-429:** no `.snapshots` is visible to either box's sub-account (measured, both machines, with controls), so the seven-day bound is unverified and unverifiable from the product side. **Q3 (documented, `hub/internal/hetznerapi/hetznerapi.go:38-45`): the sub-account API has ONE permission axis, `readonly` — there is no append-only, so the PBS shape does NOT transfer** (PBS is a server that can refuse; a Storage Box is a filesystem that runs nothing). **Q5 (measured): withdrawing delete does NOT wedge the store** — restic treats a dead owner's lock as stale and proceeds — so the constraint everyone feared is not the blocker; but `unlock --remove-all` lies about success (R-430), and the crash-lock window is UNKNOWN. **Q6 (measured): restic 0.14.0 DOES speak `rest:`** (control: `banana:` → invalid backend), and append-only is a rest-server flag, not a restic one — reachable, but it needs a machine in the recovery path and ep0 is protected. **Q7 (measured): detection is nearly free** — `snapshot_count` already reaches the hub and the hub APPENDS reports, so the comparison needs no box change. **RECOMMENDATION: answer R-429 first (Viktor, ten minutes), then build detection, then move retention off the box; defer the transport change.** Two `forget --prune` sites must be disarmed together — `offbox.go:1388` AND `offbox.go:1759` — or R-191 repeats. **RE-SCOPED 2026-09-01 — THE STORY WAS WORSE THAN THE TRUTH FOR TWO MONTHS.** The box can delete its own LIVE repository, **but it cannot write to the daily snapshots of it** — MEASURED, not cited: a write into `/.zfs/snapshot` is refused on both boxes while the same write to the account home succeeds (R-432). Seven daily snapshots are confirmed in the panel (R-429). **So a deletion costs at most the data written since the last daily snapshot, and the rest is recoverable — file by file, one customer at a time, with no effect on anyone else** (vendor: *"You can download individual files or entire directories as usual"*; *"It is not possible to write to the `/.zfs` directory or its subfolder"*). **NOT open-ended loss.** Two caveats kept honest: a panel-driven snapshot restore rolls back the WHOLE Storage Box and deletes newer snapshots, which is why the per-file route matters; and per-file recovery is **operator-only today** (R-432). **DETECTION SHIPPED hub v0.111.0 (R-431)** — an unexplained fall is noticed within a day. **THE RANKING IS VIKTOR'S:** this has been #1 since July on the old story. On the new facts I would rank it below the items that can still lose data outright, but I am not re-ranking it myself. `audits/SPIKE-r95-offsite-delete-2026-09-01.md` **DRILL 2026-09-01, LATER THE SAME DAY — THE RE-SCOPE'S SECOND HALF IS WITHDRAWN.** The drill that was to walk the recovery found there is no route to walk: **no snapshot is reachable from a sub-account by ANY name** (R-433) — 777,600 exact names in the vendor format over nine days, zero hits, with a passing control, plus the structural reason (`/home` st_dev 0,82 vs `/.zfs/snapshot` st_dev 0,276, and `/home/.zfs` absent). **Clause (a) stands: the box can delete the live repo and cannot write into the snapshot area.** Clause (b) — *"the rest is recoverable file by file"* — is NOT SUPPORTED. **The drill was STOPPED before its destructive phase on the operator's ruling**, because with no recovery route the deletion would have destroyed real history to buy only an alarm test that could not fire at the specified size (R-435). Nothing was deleted; the store is verified untouched at 69 snapshots. **So the comfort that lowered this row rested on an unwalked route, and the route does not exist.** Two new leads decide what happens next: **R-433** (can the MAIN account see them? nobody here holds that credential) and **R-436** (`rclone serve restic --stdio` is offered server-side and restic speaks `rclone:` — measured — which could make real prevention cheap, IF the provider pins `--append-only`). **The rank stays Viktor's. Plainly: the argument that moved this row down is the argument the drill removed.** `audits/evidence-drill-r95-recovery-2026-09-01/` **BLOCKED-ON-PROVIDER 2026-09-01, and the record of the demotion is kept deliberately: this row spent ONE DAY demoted on a clause that did not hold.** It was re-scoped down on the morning of 2026-09-01 on the strength of *"recoverable file by file"*, and that clause was measured false the same afternoon (R-433). **On today's evidence it belongs back near the top — that is a proposal, not an action; CC has not re-ranked it and will not.** Both questions that can settle it are drafted in `documentation/runbooks/provider-questions-2026-09-01.md`: **Q1** decides how urgent this is, **Q2** (R-436) could remove the root cause cheaply. **Precondition on any build here: R-430**, which is harmless only while the box can still delete. | | **R-105** | Backup & restore | P2 | **Three hub-held DR records are empty on the entire live fleet.** `hosts.dr_record_json` = `{}` on all 3 hosts; `host_escrow.directive_json` = `{}` on both escrowed hosts; `dr_recipe.host_half.drives` = `[]` on every customer **including two with enrolled data drives** (916 GB USB on demo-felhom, 938 GB NVMe on demo-hp) **MIGRATED FROM `ROADMAP.md` 2026-08-22 (R-369) — originally filed 2026-07-28, size M, roadmap state `READY — 2026-07-28`.** Moved verbatim; nothing added or reinterpreted. The roadmap keeps its copy as history, marked moved. ****PARTLY FIXED BY ITS OWN UPDATE.** The `drives` third was traced and populated on both demo boxes on 2026-07-28 (the enrolled drives were never PVE storages, so `isUserDataDrive` never saw them). The other two thirds — `hosts.dr_record_json` and `host_escrow.directive_json` — were NOT re-verified this session and are carried as written.** | **OPEN — migrated from ROADMAP 2026-08-22, rank unchanged** | — | These are exactly the fields a host-loss recovery reads: `05-hub-architecture.md:175-176,186` names the slim DR record as one of four durable sources; `06-offsite-connectivity.md:148-150` says the escrow upload carried the DR directive; `felhom-agent/internal/dr/plan.go:34-35` makes `PlannedDrive` the re-attach-by-`durable_id` wrong-disk guard. **The three may have different causes** — `isUserDataDrive` (`internal/hub/dr_recipe.go:129-136`) requires type `usb`/`local-dir` **and** a non-empty `DurableID` **and** `MountPath`, and which of the three fails was not traced. Evidence: `architecture/_recovery-inventory-2026-07-28.md` Part D2.3. **UPDATE 2026-07-28 (vzdump-target move): the `drives` third is TRACED and now POPULATED on both demo boxes.** Cause: the enrolled data drives were never PVE storages at all — only agent-generated systemd mounts — so they never entered `report.StorageTargets` and `isUserDataDrive` never saw them. Giving each drive a `dir` storage at its own mountpoint supplied all three required fields at once (type `local-dir`, fs-UUID durable id, mount path), and the recipe now emits `uuid:91d2dc2d-…`/`/mnt/nvme-1tb` on demo-hp and `uuid:47a3361a-…`/`/mnt/hdd_1` o | CC | | **R-232** | Backup & restore | P2 | **DooPlex's backup makes every copy inside the same box — and nothing tells anyone when it fails.** Surveyed read-only 2026-08-06 (`audits/RECON-dooplex-backup-2026-08-06.md`). **What works:** five sets, 14/14 successful runs in 14 days; a file was restored from the `data` repo and matched the live original **byte for byte**; every set except two is cross-disk; k3s is integrity-checked on every run. **What the matrix exposes, ranked:** (a) **`notify_failure` is a no-op** — `NOTIFY_ON_FAILURE=true` but `NOTIFY_WEBHOOK_URL` is commented out, so a failed backup notifies **nobody**; the project already has a working Resend path that CI uses. Cheapest item, and it makes every other failure visible. (b) **Nothing leaves the box** — no rclone, no remote repo, no off-site target anywhere; Longhorn's target is `nfs://192.168.0.180:` pointing at DooPlex itself, and the only outbound-looking cron pulls *inbound* from Hetzner for a different project. The machine that runs the hub managing the customers' off-site chain has no off-site copy of its own. (c) **The backup tree is a single writable path** and the restic repos are not append-only — one bad script or ransomware destroys every copy at once. (d) **Two same-disk sets**: `.claude-memory` and the PostgreSQL dumps, whose source directory sits *inside* the backup tree. (e) **Longhorn `retain=1`** — one generation per volume, so a corruption noticed a day late has no earlier copy. (f) **`/opt/backup/docs/BACKUP-RESTORE.md` does not exist** though the systemd unit advertises it. (g) **`secrets/restic-repo` has never held a snapshot** — `backup-secrets.sh` contains no `restic` call; the secrets are GPG files on `sda1` only. (h) **No restore has ever been run** beyond today's single-file probe — the matrix's "ever demonstrated?" column is otherwise entirely empty. **Not a finding:** the restic passphrase. The on-box copy is on `sdb1`, a different disk from the backups, and the **operator holds an offline copy out of band** — so a disk loss is recoverable. The narrow residual is that it is operator-held rather than system-held, unlike the customer case's hub-vaulted escrow, so it should be confirmed current and findable by someone else. **Nothing was changed by the recon.** | **READY** — owner Viktor | — | — | operator | | **R-304** | Backup & restore | P2 | **The retained escrow key works, and the customer is told their correct code is wrong.** DRILL 2026-08-12 answered the three questions separately, on `demo-felhom`, with planted data. **(a) retention: WORKS** — the first retained row in fleet history to carry material (`host_escrow_superseded` id 11, `identity_blob` 572 B), byte-identical (`sha256 a10032341c8584ed…`) to the pre-supersession `host_escrow` row. **(b) the material opens the old store: YES** — unsealed with the OLD recovery code it yielded a password byte-identical to the pre-change one (`sha c60c8bc737a6b7c6…`), and restored three planted files **byte-identical** from a store the box itself could no longer open (negative control first: `Fatal: wrong password or no key found`), **including a Hungarian accented filename verified as raw bytes**. **(c) the customer's route: DOES NOT EXIST, and misinforms.** `ListSupersededEscrow` (`store.go:2841`) is the only reader of a retained `identity_blob` and has **zero production callers** — five call sites, all `_test.go`; the product path (`POST /escrow/recover-offsite-password` → `FetchIdentityEscrow` → `GetHostDRBundle`, `store.go:3152`) selects `FROM host_escrow` — the CURRENT row only. Asked for the old password with the code that demonstrably opens the retained row, the product answered **"the recovery code did not open the sealed bundle — nothing was written"**. **This is the R-224 class again**: there an unreachable hub was reported as a bad code; here a VALID code for retained history is reported as a bad code, and the customer's attempt ends there. **Consequence:** the census answer stands (it was about retention); the countdown banner's promise is true in substance and false in practice; **any capability-map claim that the customer can recover the old history with their recovery code is false today and must move** | **READY (L) — NEW 2026-08-12, RANK 1** | R-198, R-199, R-224, R-241 | Decide the shape: serve retained rows on the recovery path (needs a "which package?" choice — a customer may have several), or stop promising retrieval anywhere the customer cannot perform it. **Until one of those, the honest position is that retention is an operator-only capability.** At minimum, the refusal must stop asserting the code is wrong when the hub simply never looked | operator + CC | -| **R-342** | Backup & restore | P2 | **The ep0 snapshot covers less than it looks like it covers, and the next person will assume otherwise.** Quoting `audits/evidence-ep0-pbs-upgrade-2026-08-18/stop2-snapshot.txt` verbatim: *"covers — the 38 GB system disk /dev/sda (root), i.e. the PBS packages, unit files, /etc/systemd drop-ins, nftables and wg config. DOES NOT — /mnt/pbs-datastore. That is /dev/sdb, a separate 100 GB VOLUME, and Hetzner server snapshots do not include attached volumes. The backup data is therefore NOT protected by this snapshot."* Snapshot **421440873** (`felhom-hetzner-20260818`, 15.06 GB, Available) was taken as the rollback for the 4.2.2→4.2.5 PBS upgrade. **Rolling it back restores software state, not the datastore.** That was *acceptable for that change* — a package install writes no datastore content — and the file says so. **The problem is what happens next:** this fact lives in an evidence file nobody will open again, and a snapshot named as "the rollback" reads as protecting everything on the box. ep0 holds the only off-premises copy of a real customer's data | **READY (S) — NEW 2026-08-18** | — | **Decide the safeguard for any future ep0 procedure that could touch `/mnt/pbs-datastore` — it does not exist and has not been designed.** Candidates: a Hetzner **Volume** snapshot (a different object from the server snapshot), a PBS-level sync to a second location, or an explicit written acceptance that the datastore is unprotected for the duration. **Nothing may be added to a runbook implying a safeguard exists until one does** **2026-10-03 — options costed, one candidate removed:** Hetzner Cloud has NO volume snapshots (server snapshots exclude volumes; open vendor request since 2019), so the first candidate does not exist. Options left: (A) a PBS pull-sync of `felhom-offsite` to DooPlex's PBS — €0/month, 1–2 h, protects against datastore loss AND losing Hetzner, ciphertext only (per-customer `encryption-key`), but a new job on DooPlex; (B) a written acceptance plus a copy before each risky procedure. CC recommends A. Datastore today 19.5 GB of 97.9 GB. `audits/offsite-append-only-2026-10-03/PART-D-ep0-safeguard.md`; decision in STATUS. | **Viktor decides; CC executes** — a risk-to-customer-data question | | **R-366** | Backup & restore | P2 | **The 21 August reinstall orphaned `demo-hp`'s PBS whole-guest archives as well as its off-site repo — the box can no longer read its own pre-reinstall backups, and this surfaces only as a restore-test failure.** Hub event 3016, 2026-08-21 21:59:28Z, unprompted: `Restore-test FAILED on the pbs tier: archive felhom-pbs:backup/ct/9201/2026-08-18T03:58:43Z could not be restored+booted … proxmox-backup-client failed: Error: wrong key - unable to verify signature since manifest's key 3f:4f:65:c0:d8:f3:9f:3c does not match provided key dd:d1:d8:53:44:62:5e:0b`. The archive predates the reinstall by three days. **This is the PBS-tier analogue of R-193** (a guest rebuild mints a fresh secret and orphans the history), and the two together mean a rebuilt box loses BOTH off-premises tiers at once: the restic repo needed a self-heal + re-toggle (see the drill report), and the PBS archives are simply unreadable to it. **Credit: the restore-test caught it and said so precisely** — the mechanism works. **The gap is what it is called:** it is reported as *a restore test that failed*, which reads as a flaky verification, not as *every whole-guest backup you took before the reinstall is unreadable on this machine*. **Found incidentally by the 2026-08-21 backup-truth drill; nobody was looking for it.** | **OPEN — HIGH** | related: R-193 | Establish whether the pre-reinstall PBS archives are recoverable at all (the old key's whereabouts), and separate the two verdicts: a tier whose ARCHIVES ARE ORPHANED is a different alarm from a tier whose restore test failed. Do not close on the strength of the restore-test wording alone. | CC | | **R-518** | Backup & restore | P2 | **[P2-MEDIUM] „Mentés most" on the whole-system backup stops every app for about eight minutes while the page promises „csak néhány másodpercre".** MEASURED 2026-09-14 (BIGNIGHT, VM 333, 12 apps): the button's call quiesced all 12 stacks at 19:03:23Z (first stopped 19:03:27Z); the local vzdump ran 19:03:49 → 19:09:59Z; the controller then kept the apps stopped for the second (PBS) tier and restarted them at 19:10:09Z after it failed, the last started 19:11:12Z (`phase4/guest-backup-quiesce-log.txt`) — **≈ 7 m 45 s** with every app answering 404. The page under the button: „Pillanatkép-mód: az alkalmazások csak néhány másodpercre állnak le." A household pressing it at dinner loses every app for the length of the dump, and longer on a bigger box. **Fix shape:** state the real expected downtime (it scales with data), or quiesce per tier and not across a second tier's attempt; do not start a tier whose storage is absent (see R-517). **NARROWED 2026-09-15 (controller v0.243.0 + agent v0.131.0):** a tier whose storage the agent reports absent is skipped before anything stops (`backup_tier_skipped`, once per absence; unknown never skipped), and the button copy now says „általában néhány perc, nagyobb adatnál több". Unit-proven with red-proofs. **Still open:** quiesce per tier, so a slow second tier does not keep every app down. **— NIGHT 2026-09-23 (controller v0.267.0):** the copy half is DONE: the page and the confirm now state the measured stop (≈ 8 minutes on a 12-app box), both languages, red-proofed (`audits/night-2026-09-23/A5-*`). The brief's „csak néhány másodpercre" had already gone in v0.243.0. **Still open:** quiesce per tier, so a slow second tier does not keep every app down. | **READY — P2, narrowed to per-tier quiesce; owner: CC (controller)** | — | — | CC | | **R-519** | Backup & restore | P2 | **[P2-MEDIUM] After a backup torn by a power cut, an app's restore point carries the new database dump's time while its files are from the previous run — and no customer screen says the run was interrupted.** MEASURED 2026-09-14 (BIGNIGHT F2, VM 333): „Mentés most" 19:40:03Z; the power was cut 19:40:08Z while adventurelog was stopped for its volume dump. On disk afterwards, `backups/primary/adventurelog`: `db-dumps/adventurelog-postgres.sql` **19:40:07**, `volume-dumps/*.tar` **19:00:35**, `manifest.json created_at 19:02:34Z`; bookstack the same shape (sql 19:40:07, tars 19:00:48). `GET /api/backup/snapshots` for both → `time 2026-09-14T19:40:07Z, helyi` (`phase5/F2/units-on-disk.txt`, `backup-honesty.txt`). `/backups/apps` shows „Utolsó adatbázis mentés 2026-09-14 21:40 · … OK" and every app „Utolsó: 5 perce"; `/backups` and `/dashboard` contain no word of an interruption (fragments `megszakad|sikertelen|nem sikerült` = 0; control „hiba" appears in the standing warning text). The controller itself knew: `[appstop] crash recovery: an app-data backup (volume dump) … was interrupted … restarting them: [adventurelog]` and pushed `backup_failed (error)` to the hub. A household restoring „the 21:40 backup" gets 21:00 files for BookStack's uploads. **Fix shape:** date a point by the oldest part it contains (or mark it partial) and show the interrupted run on the backups page until the next complete one. **F6 (drive unplugged 1 s into a backup, 20:37:59Z) adds three facts:** the run skipped four apps' volume dumps („Skipping volume dump for immich — drive disconnected", also jellyfin, nextcloud, paperless-ngx) and still reported `db_dump {"count":6, "success":true}`; nextcloud's point is dated 20:37:59Z (its SQL finished before the unplug) beside 19:01 volume tars; and a torn `immich-postgres.sql.tmp` (20:38:00) plus F3's `pre-restore-…-nextcloud-mariadb.sql.tmp` are left in the units on the drive. Immich's point correctly stayed at 19:02:34Z (the `.tmp` was not promoted). `/backups/apps` fragments `kihagy|sikertelen|részleges` = 0 (`phase5/F6/`). | **READY — rank P2-MEDIUM; owner: CC (controller)** | — | — | CC | | **R-638** | Backup & restore | P2 | **[P2-MEDIUM] The product's database loader cannot replay a copy over a NEWER schema: on PostgreSQL it FAILS, on MariaDB it leaves the newer version's tables behind.** MEASURED 2026-09-23 on 9202. `ImportDump` (`appbackup/dbdump.go:719`, `psql -v ON_ERROR_STOP=1 --single-transaction`) replays a `pg_dump --clean --if-exists` file over the live database. After docmost 0.95.0 → 0.96.0 migrated, the replay of the pre-update copy was refused in 0.40 s, rc 3: *cannot drop constraint workspaces_pkey on table public.workspaces because other objects depend on it / DETAIL: constraint oauth_clients_workspace_id_fkey …* — the new version created six tables whose foreign keys point at old ones, and `--clean` only drops what the dump knows. Database unchanged (the transaction rolled back). On MariaDB (`mariadb-dump`, `FOREIGN_KEY_CHECKS=0`) the same replay after romm 5.0.0 → 5.3.0 returned rc 0 in 1.25 s and left **12 base tables** of the new version behind; RomM 5.0.0 happened to ignore them. **What worked:** `DROP SCHEMA public CASCADE; CREATE SCHEMA public;` + the dump in ONE transaction — rc 0 in 1.38 s, every table, index and extension back. **Why this is a row of its own and not only part of R-637:** the SAME loader backs shipped paths — `rollbackSafetyDump` (off-site restore's undo) and the dump replay of the restores — so **any restore of a copy taken BEFORE an update that migrated, replayed over the migrated database, may fail the same way. NOT MEASURED:** whether the unit restore the hold sentence names does this (it also carries the data VOLUME tar, which may make the replay moot). That is the measurement owed, on 9202, before anyone relies on it. Evidence: `audits/update-rulings-2026-09-23/README.md` Part 1, `docmost-45`, `romm-44`. **-- NARROWED 2026-09-23:** the undo no longer touches this loader — it copies folders (decision 19, controller v0.263.0). **What stays open is the part about SHIPPED paths:** `rollbackSafetyDump` and the restores' dump replay still replay over whatever schema is live, and whether the unit restore the hold sentence names works after a real schema migration is STILL UNMEASURED. | **OPEN — P2, narrowed to the restore paths; owner: CC; measure the named restore after a real schema migration first** | — | — | CC | | **R-726** | Backup & restore | P2 | **[P2-MEDIUM] A new box for a customer who had one before makes NO off-site copy on night one: the old repository is found orphaned, and the fix is a button nobody pointed the household to.** MEASURED 2026-09-30 00:15 UTC on the new-household drill box (`tester-1`, whose previous box was deleted 2026-09-17): `[offbox] offsite repo ORPHANED — remote holds backups written under a previous, no-longer-available key; runs will skip until reset` → `offbox_repo_orphaned` (warning) to the household's timeline and an operator mail; `offsite-integrity` then checked nothing. The household's page is honest („A távoli tároló másik kulccsal készült mentéseket tartalmaz … Új távoli mentés indítása…", old history set aside, never deleted), but the evening before, the recovery-code ceremony and the off-site page raised nothing, and the guide does not mention it. The hub re-issued the off-site credentials on re-enroll by itself; it could have known the repository would orphan. Customer data was never at risk (the old repository is untouched); the household simply has no off-site copy until someone presses the button. **Fix direction:** offer the reset at the recovery-code ceremony when the repository already holds another key's snapshots, or the hub's re-enroll re-issue sets the old history aside the same way (it is the same move-aside), and the guide says so. | **READY — rank P2-MEDIUM; owner: CC (design first) / operator (which)** | — | — | CC + operator | -| **R-822** | Backup & restore | P2 | **An add-only key does not make retention safe: an attacker who can only ADD snapshots can make the honest pruner erase every real one.** MEASURED 2026-10-03 (lab, restic 0.14.0 from the controller image, rclone `--append-only`): 13 empty snapshots dated in the future with the same host and tag, added through the add-only key (all allowed), make the box's exact policy `forget --group-by host,tags --keep-daily 7 --keep-weekly 4 --keep-monthly 6` keep only the fakes and select **all 3 real snapshots** for removal (`--dry-run`). **Whoever prunes an append-only repo — the box in a window, or a Felhom-side worker — inherits this.** Not a defect today (today the box can simply delete, R-95); a PRECONDITION on the R-95 build, like R-430 was. `audits/offsite-append-only-2026-10-03/lab/C3-retention-poisoning.txt` | **OPEN — precondition on any R-95 build** | — | Before any `forget`: refuse when a snapshot is dated in the future or newer than the newest the hub has seen reported; dry-run first and abort above the count the policy can remove in a week; hub compares the count before/after (DESIGN.md §3) | CC | +| **R-822** | Backup & restore | P2 | **An add-only key does not make retention safe: an attacker who can only ADD snapshots can make the honest pruner erase every real one.** MEASURED 2026-10-03 (lab, restic 0.14.0 from the controller image, rclone `--append-only`): 13 empty snapshots dated in the future with the same host and tag, added through the add-only key (all allowed), make the box's exact policy `forget --group-by host,tags --keep-daily 7 --keep-weekly 4 --keep-monthly 6` keep only the fakes and select **all 3 real snapshots** for removal (`--dry-run`). **Whoever prunes an append-only repo — the box in a window, or a Felhom-side worker — inherits this.** Not a defect today (today the box can simply delete, R-95); a PRECONDITION on the R-95 build, like R-430 was. `audits/offsite-append-only-2026-10-03/lab/C3-retention-poisoning.txt` | **NARROWED 2026-10-03 — the guard ships in controller v0.289.0 (future-dated / newer-than-hub / recent-removal refusals, oldest-first cap, the lab's 13-fake shape refused in a test). RESIDUAL, not closable by a guard: an add-only attacker can plant PAST-dated snapshots interleaved with real ones and so steer weekly/monthly keeps; bounded per window by `MaxRemove` and the hub's count check, not prevented.** | — | Before any `forget`: refuse when a snapshot is dated in the future or newer than the newest the hub has seen reported; dry-run first and abort above the count the policy can remove in a week; hub compares the count before/after (DESIGN.md §3) | CC | +| **R-824** | Backup & restore | P2 | **The fake-snapshot guard refuses whenever a manual run happened in the last 8 days — so on an active household no prune ever runs.** MEASURED live 2026-10-03, window 1 on demo-hp: the only snapshots the honest policy removes are same-day older copies superseded by a manual run, and those are younger than 8 days, so the guard refused (`the policy would remove snapshot c6b5c67b … younger than 8 days`). The window mechanics all worked (opened, refused, closed in 3 s, operator mailed). **No real prune inside a window has run yet.** Fix direction: EXCLUDE young same-day superseded snapshots from the plan (they are removed in a later window once older) instead of refusing; keep refusing on future-dated snapshots. Then one live window with a real removal. `audits/offsite-lock-build-2026-10-03/partE/` | **READY — next session; weekly windows stay OFF until then** | — | controller fix + one live window on demo-hp | CC | | **R-49** | Backup & restore | P3 | **[P2] The offsite capture set is ~90% cache and duplication — 1.1 GB of a 1.2 GB immich "photo backup".** Measured 2026-07-19: `immich_ml_cache.tar` **823 660 032 B (~60%)** — re-downloadable ML model weights; `immich_postgres_data.tar` **308 251 136 B (~23%)** — a raw tar of the postgres data dir that DUPLICATES the logical `.sql` dump captured beside it; `upload/backups/` **18 MB** — immich's own nightly dump, a backup inside the backup, growing daily; plus the stranded pre-v3 `dccc13fe…` tree (~36 MB) no DB has ever referenced. Actual irreplaceable content: **72 MB of originals**. **MIGRATED FROM `ROADMAP.md` 2026-08-22 (R-369) — originally filed 2026-07-19, size S–M, roadmap state `idea`.** Moved verbatim; nothing added or reinterpreted. The roadmap keeps its copy as history, marked moved. | **OPEN — migrated from ROADMAP 2026-08-22, rank unchanged** **Re-ranked 2026-10-03: P2→P3: wasted space and transfer, no data risk; needs a capture-set ruling, not a sale blocker.** | — | **Evidence: `audits/DIAG-immich-restore-round2-2026-07-19.md` §4 (full byte breakdown).** This is the customer's offsite quota and transfer cost, and it lands on the Hetzner sub-account they are billed for. **Recorded, deliberately not changed** — a capture-set exclusion is a data-loss-shaped decision and gets its own ruling, not a drive-by edit. Candidates in priority order: (a) `immich_ml_cache` — pure cache, strongest case; (b) the `postgres_data` volume tar where a logical dump of the same DB is already captured (the dump is what the restore path actually replays); (c) `upload/backups/`. Likely generalises past immich into a template-classification rule about cache volumes and self-backup directories, so it should be specified against the catalog, not one app | CC | | **R-127** | Backup & restore | P3 | **The catalog's `data_key: true` flag is UNRELIABLE — at least four data-encrypting keys the catalog itself labels as encryption keys are unflagged; and the O4 restore path can regenerate a DB password that then does not match the restored data directory** | **READY (S/M)** | — | **Found by D5's Part 0, and it is why D5's boundary is `type: secret` rather than `data_key`.** Two separable legs. **(a) The misclassification.** Only 5 fields across 4 apps set `data_key: true` (`adventurelog/SECRET_KEY`, `homebox/HBOX_AUTH_API_KEY_PEPPER`, `papra/AUTH_SECRET`, `sparkyfitness/{API_ENCRYPTION_KEY,BETTER_AUTH_SECRET}`), yet `n8n/N8N_ENCRYPTION_KEY` („Titkosítási kulcs"), `wanderer/POCKETBASE_ENCRYPTION_KEY` („Adatbázis titkosítási kulcs"), `calcom/CALENDSO_ENCRYPTION_KEY` and `bookstack/APP_KEY` are unflagged — the catalog's own Hungarian labels contradict the flag. **D5 makes this non-urgent but not harmless:** everything `type: secret` now travels, so the keys DO reach the drive; what stays wrong is the **fail-closed gate**, which only refuses for `data_key` names — so if one of these is missing from both sources the restore proceeds onto data it cannot decrypt instead of refusing. Fix = flag them (`app-catalog-felhom.eu`, a catalog-only change) + a gate/test that the flag set and the label set agree. **(b) The regenerated-DB-password trap.** `internal/backup/restore_unit.go` O4 generates a replacement for any missing non-data-key secret. Proven on `postgres:16-alpine`: with PGDATA restored from the volume tar, `POSTGRES_PASSWORD` is ignored (initdb skipped), so the app fails over the compose network while the dump replay still succeeds through the container's local **trust** socket — success reported, data unreachable. v0.188.0 corrected the WARN's false claim that *"stored data is unaffected"* and scoped it, but did **not** add a guard: D5 shrinks this to the rare case (the secret was empty at capture AND absent from the guest). Real fix = either treat a DB password as fail-closed like a data key, or `ALTER USER` to the regenerated value after the volume restore. 18 DB/root-password fields are in scope; MariaDB fails loudly instead (Access denied), which is the safer half | CC | | **R-200** | Backup & restore | P3 | **The DR password-injection seam has a handler, a route and tests — and no form.** `POST /backup/offbox/inject-password` is routed (`controller/internal/web/server.go:510`) to `offboxInjectPasswordHandler` (`offbox_handlers.go:174-196`) → `InjectOffboxPassword` (`backup/offbox.go:541`). **No template in the repository contains that path or any form posting to it** (grep over `internal/web/templates/`: one unrelated hit, an XSS comment) | **NARROWED** — **PLUMBING COMPLETE** (controller **v0.196.0**); **the FORM is not built — still open** | — | **The tenth instance of this project's built-but-never-wired class, and the exact shape `CLAUDE.md` and `felhom.eu/CLAUDE.md`'s seam-wiring rule were written for:** handler tests that POST directly (`offbox_escrow_test.go:167,180`) prove nothing about reachability. To use the only implemented recovery seam today, a person must hand-craft an authenticated POST with a session cookie and CSRF token. **Note the layering while fixing it:** this form takes a **64-hex repo password**, not a recovery code (`offboxRepoPwPattern`, `offbox.go:543`) — they are different secrets at different layers, and the operator's 2026-08-04 ruling asks for a form that takes **R**. Build the R form and treat this one as the operator/DR fallback it was written as, but **ship it with a render test per branch of whatever gate it sits behind**. Source: `audits/RECON-offsite-dr-chain-2026-08-04.md` §3 link 9 **THE DIAGNOSTIC HALF IS DONE AND IT ANSWERED THE QUESTION.** `--recover-offsite-check` is a `docker exec` escape hatch in the shape of `--print-reset-code`: R on **STDIN** (never argv, never `ps`, never shell history, never a transcript), fetch+unseal via the agent, and a verdict of **two sha256 hashes**. **It compares and never installs** — the recovered password is not written to `offbox/repo_password`; a test asserts the data dir is byte-unchanged and its red-proof (adding the install call) fails it. Confirmed live: `repo_password` mtime still `2026-08-03 07:18:02` after the successful check at `2026-08-04 11:49`. **Exit codes are load-bearing** — `0` match, `2` a clean MISMATCH, `1` a step failed; "it failed" and "it worked and disagreed" must never share a status because only one is a finding about the system. A box with no local password reports distinctly (the rebuilt-box shape, where the next step is to INSTALL rather than compare). **WHAT IS NOT BUILT, deliberately:** no card, no form, no preview, no customer-facing text — building an interface on top of a chain nobody had walked is how the preceding three weeks went wrong. **What remains for this row:** link 9 (the recovered password placed so `WriteOffboxSecrets` keeps it) and the customer-facing shape the operator ruled on 2026-08-04 (yell → R form → preview → proceed), which is now priced against a chain that exists rather than one that is assumed **PART 0 SHIPPED 2026-08-04 (v0.196.0):** `--recover-offsite-install` is the sibling of the check — same fetch/unseal path, same STDIN discipline for R — and it **places** the recovered password via `InjectOffboxPassword`. The confirmation is a SECOND invocation (`--confirm-install`): without it, both hashes print and nothing is written, so the operator sees the comparison before a write is possible. Three outcomes named distinctly: **installed** (no local password — the rebuilt-box shape), **unchanged** (identical key, nothing written), **refused** (a DIFFERENT key present — installing would clobber the key the current repository is encrypted under; exit 2, no force offered). It re-reads the file after writing rather than trusting the call. **Red-proof observed:** removing the confirmation gate makes the dry run write the password. The R-persistence test carries a **positive control** (a planted copy found, then removed and not found). **NOT YET EXERCISED AGAINST A LIVE RECOVERY** — the R-201 drill halted before step 9, so this is unit-proven only. **What remains for this row:** the customer-facing shape the operator ruled on 2026-08-04 (yell → recovery-code form → preview → proceed) | CC | @@ -230,6 +230,9 @@ stopping line that lies. | **R-698** | Backup & restore | P3 | **[P3-LOW] A backup stores the image's NAME, not the image — a restore of a version its maker has deleted cannot start.** `RecoveryManifest.image_pins` ("image NOT stored — re-pulled on restore"); since controller v0.275.0 each data file also records its running `ref@digest`, and a restore brings the data back AT ITS OWN VERSION (`07` §6.6) — so a restore asks for exactly the old image. **Measured 2026-09-26** (`audits/version-travel-2026-09-26/A7/`, registry HEADs, no pulls): the catalog's 42 ladder `ref@digest` pairs all resolve (200); an invented digest answers 404 on Docker Hub and ghcr.io (negative control). Not measured: the digests recorded on boxes (older than any ladder entry), how often makers delete versions, the catalog's 66 digest-less compose lines. **Options (decide nothing yet):** (a) keep — a restore of a deleted version fails at the pull and the household uses the next copy or a newer version; (b) mirror every INSTALLED image into the DooPlex registry, restore falls back to it — storage + bandwidth on DooPlex, a new part on the recovery path; (c) mirror only ladder-named versions — bounded, misses pre-ladder boxes; (d) `docker save` into the unit — hundreds of MB per app per copy on every tier. **-- 2026-09-30 late (decision 53):** a box now keeps only an app's running and previous image; a restore to an older version re-pulls it — as every restore already did. The limit above is unchanged. | **OPEN — P3; owner: operator (a decision), CC measures** | — | — | CC + operator | | **R-706** | Backup & restore | P3 | **[P3-LOW] Removing an app "with its backups" leaves its off-site verification copy on the drive.** Measured 2026-09-28 on demo-hp: after a full off-site restore of nextcloud (which leaves the downloaded copy in `backups/offsite-restore/nextcloud`, ~1 GB, by design, for the household to inspect), `POST /api/stacks/nextcloud/remove` with `remove_backups: true` removed the unit and listed `backup_paths_removed` WITHOUT the verification copy; it stayed until the restore page's own delete (`POST /backup/offbox/verify-copy/delete`, 302 `scratch_deleted`). A household that removes an app to free space keeps 1 GB it cannot see on the app list. **Fix direction:** the removal with backups also deletes the app's verification copy (the same `DeleteOffsiteRestoreCopy`). Evidence: `audits/kept-offsite-2026-09-28/E/E9-teardown.txt`. **-- 2026-09-28 evening: FIXED in controller v0.279.0** — a removal with its backups also deletes the verification copy and lists it among the removed paths (`TestR706_…`, red-proofed RP7). Not yet seen live (needs a full off-site restore then a removal). | **WATCHING — P3; owner: CC (live)** | — | — | CC | | **R-729** | Backup & restore | P3 | **[P3-LOW] An off-site target, once saved on the page, cannot be removed through the product.** MEASURED 2026-09-30 on 9202: `/backup/offbox/config` refuses an empty address and no route clears the target; the session removed its throwaway target from `settings.json` by hand, with the controller stopped (harness teardown on a scratch guest). A household that tries its own NAS and gives up keeps a disabled target forever. **Fix direction:** a „Távoli mentési cél törlése" press that clears the target (never the repository). | **READY — rank P3-LOW; owner: CC (controller)** | — | — | CC | +| **R-823** | Backup & restore | P3 | **The household-chosen deletion of set-aside off-site history no longer happens on the hub-provisioned tier.** Since controller v0.289.0 the box's key is append-only, so a due abandonment deletes NOTHING; the box raises `offbox_abandon_deferred` (operator-only) and closes the countdown. The set-aside copy stays and costs quota until the operator removes it. Not yet checked: whether the household's page still names a deletion date. Fix direction: the hub deletes the set-aside directory on the box's request, with a hub-enforced delay so a broken-into box cannot use it to erase history. | **READY** | — | read the page copy; hub-side deletion with a delay | CC + operator | +| **R-828** | Backup & restore | P3 | **The ep0 copy on DooPlex grows without bound.** The nightly pull keeps everything (`remove-vanished false`, deliberately — a deletion on ep0 must not reach the copy) and there is no prune or garbage collection on `ep0-copy`. 12 GB on day one. Fix direction: a prune job on the copy that is LOOSER than ep0's own (ep0 keeps 2 per namespace) plus a weekly GC — an operator choice of how much history the copy keeps. `runbooks/ep0-datastore-copy.md` | **READY — operator picks the copy's retention** | — | — | operator + CC | +| **R-830** | Backup & restore | P3 | **The restore route from the DooPlex ep0 copy has never been walked, and DooPlex's PBS is not reachable from a household's host.** `runbooks/ep0-datastore-copy.md` describes two routes (point the host at DooPlex; rebuild an endpoint and pull back); neither has run. | **READY** | — | walk route 2 on a scratch endpoint | CC | | **R-10** | Backup & restore | P4 | T-6E-1: DB-dump dir-fsync asymmetry (LOW, confirmed in 6E) **MIGRATED FROM `ROADMAP.md` 2026-08-22 (R-369) — originally filed 2026-07-15, size XS, roadmap state `idea`.** Moved verbatim; nothing added or reinterpreted. The roadmap keeps its copy as history, marked moved. **T-6E-1, confirmed in CAMPAIGN-6E.** | **OPEN — migrated from ROADMAP 2026-08-22, rank unchanged** | — | One-line hardening; batch with the next controller task | CC | | **R-91** | Backup & restore | P4 | Old 13 GB datastore copy at `/srv/pbs-felhom` on ep0's root disk | WATCHING | demo-felhom's first **post-migration** PBS backup | Delete once it lands; fix `CONTEXT.md:1018` same commit | CC | | **R-99** | Backup & restore | P4 | Server-side prune **never removes** a phantom snapshot. Confirmed it does NOT count them toward `keep-last` (dry-run kept 2 real + the phantom) so there is **no retention/data-loss bug** — but one accumulates per aborted upload, forever | READY (S) | — | Decide a cleanup path. Deletion on a **customer** datastore is a separate ruling — detection shipped, removal deliberately not automated | CC | @@ -268,15 +271,13 @@ stopping line that lies. | **R-368** | Storage & devices | P4 | **The storage default DOES apply at deploy time — the earlier claim that it never does was wrong, and the residual defect is smaller and different.** R-352 and `SPEC-app-data-placement-2026-08-21.md` §2.2 stated *"the deploy route never reads it"*, from `grep -nE 'GetDefaultStoragePath|primaryHDDPath|IsDefault' internal/stacks/deploy.go internal/stacks/manager.go` → nothing. **That grep searched Go files only and never the templates.** `internal/web/templates/deploy.html:612` reads `.IsDefault` directly off each `DeployStoragePath` (which embeds `settings.StoragePath`, `web/handlers.go:89-99`) and **pre-selects the default drive for a new deploy**: `{{else if and .IsDefault (not .NotAllowed)}}selected{{end}}`. So `// new apps use this by default` (`settings.go:453`) is **IMPRECISE ABOUT THE MECHANISM, NOT FALSE** — nobody calls `GetDefaultStoragePath()` on that route, but the value is honoured. The customer-facing label promises exactly this and no more: **„Legyen alapértelmezett új telepítéseknél"** (`storage.html:469`). **THE RESIDUAL, and it is the whole finding:** the default lives in the TEMPLATE, not in the server. `POST /api/stacks//deploy` accepts `values` verbatim; omit `HDD_PATH` and `withPathVars` (`stacks/deploy.go:584`) receives `""` and no default is applied. **That is why the invariant has no test — there is nothing server-side to test.** | **OPEN — LOW** | corrects R-352(2); supersedes SPEC §2.2 | Either move the default into the server so the API and the form agree and a test can pin it, or reword the comment to say the template owns it. Do not "fix" the behaviour: it is correct on the path customers use. | CC | | **R-568** | Storage & devices | P4 | **[P3-LOW] The dashboard's drive-health rows swap order between visits — the same two disks, listed in a different order a minute apart.** MEASURED 2026-09-17 on demo-hp 9201 during slice 1 release C's live proof: `/dashboard` fetched on 0.249.0 listed „KXG50PNV1T02 NVMe TOSHIBA 1024GB” then „SanDisk X600 M.2 2280 SATA 128GB”; fetched on 0.250.0 a minute later, the reverse (`audits/i18n-slice1-2026-09-17/C/live/hu-before-vs-after.txt`). `diskHealthRows` (`disk_health.go` L135–141) keeps the agent's response order and does not sort; the agent's order is therefore not stable. Cosmetic, but a household that reads „the second disk” finds a different one. **Fix shape:** sort the rows controller-side by a durable key (device path or serial), with a test that feeds two orders and expects one. | **READY - rank P3-LOW; owner: CC** **Re-ranked 2026-10-03: P3->P4: cosmetic.** | — | — | CC | -## Security & access — 31 rows (P2 5, P3 23, P4 3) +## Security & access — 30 rows (P2 3, P3 24, P4 3) | ID | Category | Sev | What | State | Blocked on | Next action | Owner | |---|---|---|---|---|---|---|---| | **R-133** | Security & access | P2 | **The vaulted break-glass console credential is PLAINTEXT AT REST — every hub DB backup is a fleet-wide console-credential dump.** `host_recovery.secret` holds each managed box's `root@pam` password verbatim, so any copy of the SQLite DB (Longhorn snapshot, PBS backup of the hub PVC, a hand-taken copy during a diagnosis) carries root console access to every Felhom host in one file | **READY (M) — NEW 2026-07-31** | — | **The deferred leg of hub v0.84.0** (Console access card), filed separately because v0.84.0 changed only WHO can retrieve the secret, never how it is stored. v0.84.0 makes it more worth doing, not more broken: retrieval now rides the hub SESSION, so the DB and the login password are jointly the whole protection (ruling **S-4**, `CONTEXT.md`). Fix shape: **envelope-encrypt the `host_recovery.secret` column under a KEK held outside the DB** — the hub already proves it can hold something it cannot itself read (escrow blobs), and that contrast is the argument. Two constraints the design must respect: the credential must stay retrievable **when the box is unreachable** (that is the whole point of break-glass), so the KEK cannot live on the box or depend on the agent; and the global-key API path must keep working with the hub UI down. Would flip the capability-map row **"Break-glass management-plane recovery"**, which today reads IMPLEMENTED with this as its caveat | CC | | **R-135** | Security & access | P2 | **`validateCSRF` returns TRUE when there is no session cookie** (`hub/internal/web/server.go:678-683`) — measured live: `POST` with Basic auth and no cookie goes straight past the CSRF gate (404, not 403), while the same POST with a cookie and no token is 403 | READY (S) — **security** | — | Browsers cache HTTP Basic credentials per origin and resend them automatically on cross-origin requests, and `SameSite` does not govern the `Authorization` header. So if the operator has ever Basic-authed to the hub in a browser, any attacker page can POST to every mutating route. Latent on the condition, not guaranteed absent. Fix = require the token whenever the request is not provably programmatic, or drop browser-usable Basic auth. Same audit §4.3 | CC | | **R-777** | Security & access | P2 | **[P2-MEDIUM] Emby and Jellyfin treat every internet visitor as being on the LAN — users with "remote access" off can sign in from the internet, IP filters and remote limits are skipped.** READ in source (`audits/visitors-2026-10-01/A/sweep/sweep-1.md`, `sweep-2.md`), not measured live: Jellyfin with `KnownProxies` empty uses the TCP peer (traefik, private) → "LAN"; Emby reads the leftmost XFF (its chain is now removed by the R-753 reset, so it sees cloudflared's private address → "LAN", as before). True before R-753 too; R-753 neither caused nor fixed it. **Needs:** measure on 9202 (a user with remote access off, through the simulated tunnel); Jellyfin: `KnownProxies` `172.16.0.0/12` in `network.xml` (no env — an `after_install` or a seed file); Emby: no setting fixes it (its `LocalNetworkSubnets` still counts private ranges) — a page sentence or a decision. | **READY — rank P2-MEDIUM; owner: CC (measure), operator (Emby route)** | — | — | CC + operator | -| **R-820** | Security & access | P2 | **A box can obtain its own off-site sub-account password at will, and that password removes any append-only lock.** MEASURED 2026-10-03 on `u629488-sub4`: the password logs in on port 22 AND port 23, and through it `.ssh/authorized_keys` was read and rewritten (the test keys were installed that way). READ from source: the box declares `needs_credential` in two reports → the hub's off-site self-heal re-arms the STORED value (`RestageOneTimeSecret`) or re-issues it at the provider → the box consumes it with its own API key (`hub/internal/offsiteheal/reconciler.go`, `controller/internal/offsiteapply/seams.go`). **Not P1:** today the box's own key can already delete (R-95, P2), so this adds no harm TODAY; it is the bypass the moment a pinned key ships, so it is a precondition on R-95. `audits/offsite-append-only-2026-10-03/live/B1-B2-ports-password-forward.txt` | **OPEN — precondition on any R-95 build** | — | The hub becomes the key registrar (box sends its PUBLIC key, the hub writes the pinned line); the box never receives the password; a daily hub check alarms on any `authorized_keys` line without the pin (DESIGN.md §2) | CC + operator | -| **R-821** | Security & access | P2 | **The hub database holds every customer's off-site sub-account password in the clear, indefinitely — whoever reads it can delete every household's off-site history.** `one_time_secrets.value` is never cleared after a consume (deliberate, so the self-heal can re-arm it — `hub/internal/store/store.go` `RestageOneTimeSecret`). Confirmed 2026-10-03: the stored value for `tester-1` (re-issued 2026-09-29) still logged in to its sub-account. The same holds for any copy of the hub DB. **Not P1:** it needs the hub or a copy of its DB, which is operator-tier; it is the central version of R-95. | **OPEN — needs an operator ruling on who holds the password** | — | With R-820's registrar the hub still needs the password; options: keep it but encrypted at rest, or rotate it after each use and keep only the current one; decide with the R-95 ruling | CC + operator | | **R-126** | Security & access | P3 | **A `.fab` bundle — plaintext secrets, OPTIONAL password — can be exported ONTO a NAS.** `storageDriveList()` (`internal/web/handler_export.go`) does not filter network paths | READY (S) | — | Split out of R-108, which closed without it: this is an explicit customer-chosen **export destination**, not a browsing surface reaching a backup tree, so it was never part of D5's precondition (`07` §7.3 records that reasoning). Was recorded inside R-108's row as its "second effect, independent of D5"; promoted to its own row so it does not vanish with R-108's closure. Fix = filter network paths out of the export destination list, or require the bundle password when the destination is a share | CC | | **R-132** | Security & access | P3 | **`curl -w '%{redirect_url}'` reconstructs the request URL WITH its basic-auth credential** — so a `-u ":$HUB_PW"` call that never put the password in a URL still printed it | **WAITING-ON-OPERATOR** — **ACTION: rotate `HUB_PW`** **Folded R-580 2026-10-03** (the same `curl -w %{redirect_url}` credential echo, seen again 2026-09-18). | — | Happened on 2026-07-31 while red-proofing the R-120 gate: the hub operator password was written to the session transcript by the write-out format, not by the request. `-u` is safe; the *reporting* was not. Rule: read the redirect from `-D -` and grep `^Location:`, never `%{redirect_url}`, on any authenticated call. Rotate the hub password (`/configuration` → Login password; ConfigMap `auth.password_hash` is the reset path) and update `~/.config/credentials` | Viktor | | **R-136** | Security & access | P3 | **Rename `hub_session` → `__Host-hub_session`** — makes cookie tossing structurally impossible | READY (XS, one line) | — | Verified on the live production response that all three prefix preconditions already hold: `Path=/`, `Secure`, no `Domain`. **Caveat for the ticket:** browsers reject a `__Host-` cookie without `Secure`, and `isSecure` is conditional on `r.TLS`/`X-Forwarded-Proto`, so plain-HTTP *browser* access to the hub would stop working (non-browser access uses Basic auth, unaffected). Tested consequence: `r.Cookie` returns the FIRST match and never tries the others, so a tossed cookie wins outright. Same audit §4.1-4.2 | CC | @@ -300,11 +301,12 @@ stopping line that lies. | **R-778** | Security & access | P3 | **[P3-LOW] A box that rolls back to a controller ≤ 0.285 after v0.286 keeps the new traefik (an old controller never rewrites a running traefik) — and the old `clientIP` believes the LEFTMOST X-Forwarded-For, which a stranger then writes: the dashboard's login counter becomes dodgeable until the box moves forward again.** Reasoned from the code (old `claim.go` clientIP + v0.286 traefik trust), not measured. The floor never moves back; the window is the self-update's crash roll-back. **Needs:** decide whether that window matters (it closes at the next floor); if it does, a 0.285.x patch that reads the rightmost hop, or the self-update refusing to roll back across v0.286. | **OPEN — rank P3-LOW; owner: CC** | — | — | CC | | **R-782** | Security & access | P3 | **[P3-LOW] Two side observations of the R-753 sweep, inferred, not measured:** glance's seeded `glance.yml` has no `auth:` block (the dashboard is public to anyone with the address), and homepage's `/api/*` refuses a Host not in `HOMEPAGE_ALLOWED_HOSTS`, which the template does not set (widgets may 400). **Needs:** measure both on 9202; glance: decide whether a public link dashboard is intended (the setup gate does not cover it after setup). | **READY — rank P3-LOW; owner: CC (catalog)** | — | — | CC | | **R-783** | Security & access | P3 | **[P3-LOW] SparkyFitness: three wrong sign-ins by anyone shut EVERY visitor out of sign-in for ~10 s — a stranger retrying every 10 s keeps the household out.** MEASURED 2026-10-01 on 9202 through the simulated tunnel (`audits/visitors-2026-10-01/C/box/sparky-box.txt`): better-auth's sign-in limit (3 per 10 s) is keyed on one address — it reads the LEFTMOST X-Forwarded-For, whose chain the R-753 router reset removes, so its frontend nginx hands it traefik's address; the household from another address got 429 at 3 s and 10 s, in at 16 s. Not forgeable (a rotating forged address did not escape). better-auth's header setting is not exposed as an env by SparkyFitness. **Needs:** accept (10 s), or an upstream setting for better-auth's `ipAddressHeaders` + a right-walking reader. | **OPEN — rank P3-LOW; owner: CC** | — | — | CC | +| **R-826** | Security & access | P3 | **Storage Box sub-accounts kept every earlier box's key, unpinned — each a route to delete history.** The registrar's first install dropped 4 stale unpinned lines on demo-felhom's sub-account and 5 on demo-hp's (every reinstall had added a key; nothing removed one). tester-1's sub-account still holds 3 (its box is gone), so the daily key check raises `offsite_key_unlocked` for it every day until they go. Fix: the next tester-1 box's registration removes them, or the operator lets CC rewrite that file through the registrar. | **READY — tester-1 alarm stands until then** | — | — | CC | | **R-134** | Security & access | P4 | **Two zone-resolvers disagree on depth.** The controller strips labels progressively (`controller/internal/cloudflare/zone.go:18`); the hub's `resolveZone` tries the exact name then `parentDomain`, which strips exactly ONE label (`hub/internal/cloudflare/unblock.go:115,136`) | READY (XS) | — | For a one-label Felhom-issued subdomain both work; for anything deeper the hub silently fails to find the zone while the controller succeeds — the geo-unblock would then no-op with a "no active zone found" error. One concept, two implementations. Same audit §2.6 | CC | | **R-525** | Security & access | P4 | **[P3-LOW] FileBrowser has its own login; putting it behind the dashboard session (traefik forwardAuth or Quantum proxy auth) is a new mechanism nobody has measured.** Filed 2026-09-15 by the P1-fixes task (B.5). R-513 closed the default-password hole with a generated password; a household still has two logins. **What it needs:** a spike on a scratch guest — forwardAuth to the controller session, and what FileBrowser Quantum does with a trusted header. | **READY — rank P3-LOW; owner: CC (spike)** **Re-ranked 2026-10-03: P3->P4: comfort feature needing a new unmeasured mechanism; the default password hole is closed.** | — | — | CC | | **R-779** | Security & access | P4 | **[P3-LOW] Part A's "two outside addresses seen as two" is proven through the simulated tunnel only; on the REAL tunnel the second outside address (ep0, one request allowed) was refused by Cloudflare's edge with 403 and never reached the box.** Measured 2026-10-01 19:51 UTC (`audits/visitors-2026-10-01/A/L2-demo-hp-real-tunnel.txt`): no log line on demo-hp; demo-hp's box has no geo restriction in its settings, so a Cloudflare ZONE rule (country or bot, not read) refused a German datacenter address. DooPlex's own address on the real tunnel was seen as itself. **Needs:** one sign-in from a second Hungarian address (the operator's phone off wifi) while DooPlex is locked out — 2 minutes; and say which Cloudflare rule refused ep0. | **WAITING-ON-OPERATOR — rank P3-LOW; owner: operator (a phone), CC reads the logs** **Re-ranked 2026-10-03: P3→P4: a proof gap on the real tunnel; operator-only follow-up.** | — | — | CC + operator | -## Box system & updates — 17 rows (P2 3, P3 11, P4 3) +## Box system & updates — 16 rows (P2 3, P3 10, P4 3) | ID | Category | Sev | What | State | Blocked on | Next action | Owner | |---|---|---|---|---|---|---|---| @@ -355,7 +357,7 @@ stopping line that lies. | **R-348** | Monitoring & notifications | P4 | **Every agent restart blanks the reported backup list for up to ~18 hours, and the comment that covers it says "unaffected".** Observed 2026-08-20 while deploying R-344: the first host reports after `demo-hp`'s agent restart carry **`0 backups`** (11:15:50 and 11:30:52 CEST, two consecutive), while the box's own `pvesm list` shows archives present on **both** tiers. `internal/backup/store.go`'s `Store` is in-memory and `byTarget` is repopulated only when a backup **runs** — daily for the local tier, weekly for offsite — so the field reads 0 until the next run. `restore_tests` did **not** blank, because that half has a durable on-disk companion (`RestoreTestState`, R-189). **It blinds no alarm, and that was CHECKED rather than assumed.** `hub/internal/monitor/deadline.go` scans back over stored reports with a 7-day `backupEvidenceLookback` whose own comment names this exact case — *"when the LATEST report carries none... and against an agent that stayed restarted for days"* — and `pbs_snapshots` stayed populated at 2 regardless. So this is an observability wart, **not** a safety hole, and it is filed at that severity deliberately. **What is actually wrong is the comment.** The `Store` doc says *"Backups are unaffected — their freshness has a ground truth on the storage (R-84)"*. That is true of the **consequence** and false of the **field**, and it sits three lines below a paragraph explaining that the very same sentence about restore-tests *"used to be here and it is now FALSE"* — so the file already carries one correction of this shape and invites the next reader to trust the surviving half. | **READY (XS) — NEW 2026-08-20** | — | Say what is measured: the field IS lost on restart and repopulates only when a backup runs; the freshness VERDICT is unaffected because the hub looks back 7 days. **Name `backupEvidenceLookback` in the comment** so the cross-repo dependency is visible from the agent side — today the agent's claim of safety rests on a hub constant it does not mention. Per the workspace rule, a comment asserting an invariant needs a test pinning it: the pin belongs on the HUB side, asserting the verdict survives a report carrying `backups: []`. | CC | | **R-371** | Monitoring & notifications | P4 | **The off-site tier is the only backup tier that announces nothing on success.** Written down 2026-08-05 in `audits/CAMPAIGN-11-recovery-journey-2026-08-05.md:508-513` and explicitly *"recorded, not filed"*: the off-site run emits **no hub event at all**, while both lesser tiers do (`db_dump_completed`, `crossdrive_completed`). Failures are covered by `backup_run_failures` and staleness by the hub's 8-day tier deadline, which is why it was judged a wrinkle. **Still true 2026-08-22** — the 2026-08-21 drill's own event dump shows `db_dump_completed` and six `crossdrive_completed` rows and no off-site success event. **Age when filed: 17 days.** | **OPEN — LOW** | — | Either emit one, or record deliberately that the highest-value tier is silent on success and say why. | CC | -## Hub & operator — 21 rows (P2 1, P3 7, P4 13) +## Hub & operator — 22 rows (P2 1, P3 7, P4 14) | ID | Category | Sev | What | State | Blocked on | Next action | Owner | |---|---|---|---|---|---|---|---| @@ -380,6 +382,7 @@ stopping line that lies. | **R-688** | Hub & operator | P4 | **[P3-LOW] The customer delete says it removes the tunnel and zone, but no leg of it calls Cloudflare.** The dialog's acknowledgement reads "the customer will be RESET — offsite repo DESTROYED, PBS revoked, tunnel/zone removed" (`hub/internal/web/customer_delete.go` `deleteCascadeAcks`), while `commitCustomerReset` has legs for Hetzner, PBS, claim, descriptor and DB only. Seen 2026-09-25 retiring `peti-felhom`, whose config carried a Cloudflare tunnel token and API token (`sajatfelhom.hu`): the tokens went with the record; any tunnel or DNS record on Cloudflare's side was neither listed nor removed. **Fix direction:** either a Cloudflare leg (tunnel + DNS by the customer's ids), or the dialog stops promising it and lists what to remove by hand. `audits/RETIRE-peti-2026-09-25.md` **-- HALF DONE 2026-09-25 (hub v0.125.0):** the dialog no longer promises a Cloudflare removal; the preview lists what the operator removes by hand, by domain (the tunnel, the DNS records), never the token — proven live on the hub (`audits/night-2026-09-26/F/`). The Cloudflare leg itself is NOT built. | **NARROWED — the Cloudflare leg only; owner: operator (decide if it is wanted) / CC (build)** **Re-ranked 2026-10-03: P3→P4: the dialog no longer promises it; what is left is operator comfort.** | — | — | CC + operator | | **R-719** | Hub & operator | P4 | **[P2-MEDIUM] A customer who already exists never gets a fresh self-bind link when their new box registers: the last link expires in 7 days and nothing re-sends it.** MEASURED 2026-09-29 (new-household drill, `tester-1`): the previous link went out 2026-09-17 07:25 UTC at a host delete and expired 2026-09-24; the box registered at 19:11:30 UTC and its console told the volunteer to open the link from their e-mail — there was none that worked. Hub source: the link is sent at customer creation, RESET, e-mail set on a box-less customer and host delete (`selfbind_mint.go` callers `hosts.go:908`, `configs.go:850`, `customer_reset.go:162`) — never on appliance registration. The volunteer guide says the operator needs to press nothing. The operator pressed „Send self-bind link" (the mail arrived in 1 s) — an operator step the volunteer depends on, recorded, not an intervention. **Fix direction:** send the link when an unclaimed appliance registers while a box-less customer waits with no live link (R-509's first fix shape), or the guide's operator part says: press it the day the volunteer installs. Evidence: `audits/evidence-drill-new-household-2026-09-30/` `phase0/operator-steps.txt`. **CHANGED AND BUILT 2026-09-30 (hub v0.126.0) — the brief's shape was not buildable:** a box registers UNCLAIMED (uuid, MACs, host keys, hardware — nothing of a customer), so "send the link when their box registers" would mail every waiting customer. Built instead: the expired AND used link pages offer „Új linket kérek" → a fresh link to the address registered for that link's customer, only when it has no box, ≤1/h per customer, identical answer for any token (no oracle). Live: the button on the real hub, the same page for a made-up token, no mail; the mint+send path unit-proven (RP42). Evidence: `audits/evidence-fixes-first-tester-2026-09-30/``partD/`. | **WAITING-ON-OPERATOR** (2026-10-03 triage: the row's verdict was finished, but it names open work no other row carries — the operator has not reviewed the changed page shape ("operator may prefer another"), and mint+send is unit-proven only) — **CLOSED 2026-09-30 — hub v0.126.0 (changed shape; operator may prefer another)** | — | — | operator | | **R-814** | Hub & operator | P4 | `PBS-storage-1` (u629193, box 611421) still `status=active`, 19.9 MB | **VERIFY** (2026-10-03 triage: a July watch row with no id; given R-814. WAITING-ON-OPERATOR — no record found that the box was deleted.) — WAITING-ON-OPERATOR | operator console | Delete the box | operator | +| **R-827** | Hub & operator | P4 | **The daily off-site key check is not strictly read-only:** its read step creates `.ssh` when the directory is absent (`offsitekeys.read` → `mkdir .ssh`). Tester-2's sub-account read back 0 lines on 2026-10-03 — it may have been given an empty `.ssh`. Harmless (an empty directory) but the code comment and REUSE row say read-only. Fix: create `.ssh` only on the write paths. | **READY** | — | — | CC | ## Business & legal — 7 rows (P2 4, P4 3) @@ -393,7 +396,7 @@ stopping line that lies. | **R-793** | Business & legal | P4 | **[P3-LOW] Enterprise / BUSL code ships inside four open images — Cal.com and Docmost (EE folders, off without a key), Outline (BUSL-1.1: no commercial "Document Service"), meilisearch v1.36 in Wanderer (EE modules).** READ 2026-10-02 (`audits/licences-2026-10-02/TABLE.md`). Each is fine as the catalog runs them: no EE key, the household's own Outline is not a Document Service, Wanderer uses plain search. **Watch:** never turn on an EE feature, never switch Karakeep's/Wanderer's meilisearch to the `-enterprise` image, and re-read on each major. | **WATCHING — rank P3-LOW; owner: CC** **Re-ranked 2026-10-03: P3→P4: a watch item; nothing is wrong as the catalog runs them.** | — | — | CC | | **R-794** | Business & legal | P4 | **[P3-LOW] redis 7.4 (RSALv2 / SSPL, not OSI) runs as a private cache in seven apps: dawarich, docmost, immich, nextcloud, outline, paperless-ngx, romm.** READ 2026-10-02 (`audits/licences-2026-10-02/TABLE.md`). Read as permitted (a private cache only its app uses is not Redis offered as a service — inferred). Valkey (BSD-3) or redis 8 (AGPL option) removes the question. **Needs:** a ladder step per app to valkey or redis 8, through the harness — no hurry. | **READY — rank P3-LOW; owner: CC** **Re-ranked 2026-10-03: P3→P4: the row itself says no hurry; usage read as permitted.** | — | — | CC | -## Process & tooling — 87 rows (P3 4, P4 83) +## Process & tooling — 86 rows (P3 4, P4 82) | ID | Category | Sev | What | State | Blocked on | Next action | Owner | |---|---|---|---|---|---|---|---| diff --git a/documentation/runbooks/ep0-datastore-copy.md b/documentation/runbooks/ep0-datastore-copy.md new file mode 100644 index 00000000..049f88a2 --- /dev/null +++ b/documentation/runbooks/ep0-datastore-copy.md @@ -0,0 +1,50 @@ +# Runbook — the ep0 datastore copy on DooPlex (decision 70, R-342) + +ep0's PBS datastore `felhom-offsite` (`/mnt/pbs-datastore`, a separate Hetzner Volume that no server snapshot +covers and Hetzner cannot snapshot) is pulled to DooPlex every night. The copy holds **ciphertext only** — each +household's whole-box backups are encrypted with that household's own `encryption-key`; DooPlex cannot read them. + +## What exists + +| Where | Object | Purpose | +|---|---|---| +| ep0 | API token `root@pam!dooplex-sync`, ACL `DatastoreReader` on `/datastore/felhom-offsite` | read-only pull. **The only change on ep0.** | +| DooPlex | `felhom-ep0-pbs-tunnel.service` (systemd, runs as `kisfenyo`, `Restart=always`) | `ssh -N -L 127.0.0.1:18007:127.0.0.1:8007 root@ep0` — ep0's PBS listens on `wg0` only; DooPlex is not a WireGuard peer | +| DooPlex PBS | remote `ep0` (127.0.0.1:18007, ep0's cert fingerprint pinned) | the pull source | +| DooPlex PBS | datastore `ep0-copy` at `/mnt/5_hdd/backup/ep0-copy` | the copy | +| DooPlex PBS | sync job `ep0-felhom-offsite`, daily 05:00, `remove-vanished false` | the nightly pull (ep0's prune runs 03:30). **It never removes what ep0 removed** — a deletion on ep0 does not reach the copy | +| DooPlex PBS | verify job `verify-ep0-copy`, Saturdays 06:30 | reads the copy back | +| DooPlex PBS | notification target `felhom-operator` (SMTP via Resend → admin@felhom.eu) + matcher `felhom-operator-errors` (every error) | a failed pull or verify reaches the operator. Proven 2026-10-03 with a test mail | + +Secrets, all out of git: the ep0 token secret in `/etc/proxmox-backup/remote.cfg` (root:backup 0640, base64 — +PBS's own format), the Resend key in `/etc/proxmox-backup/notifications-priv.cfg` (root:root 0600). + +## Checks + +```bash +systemctl is-active felhom-ep0-pbs-tunnel.service +curl -sk -o /dev/null -w '%{http_code}\n' https://127.0.0.1:18007/ # 200 = ep0's PBS reachable +sudo proxmox-backup-manager task list --limit 10 | grep -E 'syncjob|verif' # last runs +sudo find /mnt/5_hdd/backup/ep0-copy/ns -mindepth 4 -maxdepth 4 -type d | sort # snapshots per customer +``` + +## If ep0 is lost — restore a household's whole box from the DooPlex copy + +The copy is a normal PBS datastore. Two routes, both needing the household's PBS `encryption-key` (escrowed, +recovered with the household's recovery code — the same as restoring from ep0): + +1. **Point the box's host at DooPlex instead of ep0.** On the household's Proxmox host, add a PBS storage for + DooPlex's PBS (`ep0-copy`, namespace = the customer id) with the household's key, then restore the CT from it as + from ep0. DooPlex's PBS must be reachable from the host (it is not public today — the operator decides the route + at the time: a temporary tunnel, or a new endpoint). +2. **Rebuild the endpoint.** Provision a new ep0 (06 §5), then pull back: on the new ep0 add DooPlex as a remote + and run `proxmox-backup-manager pull ep0-copy felhom-offsite`. Every box then reconnects as before. + +Do **not** prune or garbage-collect the copy tighter than ep0's own retention. The copy has no prune job today and +grows with every nightly backup (R-828). + +## Remove + +`sudo proxmox-backup-manager sync-job remove ep0-felhom-offsite; … verify-job remove verify-ep0-copy;` +`sudo systemctl disable --now felhom-ep0-pbs-tunnel.service`; on ep0 +`proxmox-backup-manager user delete-token root@pam dooplex-sync`. The datastore's bytes stay until removed by hand. diff --git a/documentation/runbooks/secrets.md b/documentation/runbooks/secrets.md index 49fddcc2..360d1e1a 100644 --- a/documentation/runbooks/secrets.md +++ b/documentation/runbooks/secrets.md @@ -140,6 +140,21 @@ for every off-site customer. --- +## DooPlex PBS — the ep0 copy (decision 70, 2026-10-03) + +Not k8s Secrets — PBS's own private files on DooPlex, never in git: + +| File | Holds | Created | +|---|---|---| +| `/etc/proxmox-backup/remote.cfg` (root:backup 0640) | ep0 API token `root@pam!dooplex-sync` (base64, PBS format) | from the token's one-time output, file → file | +| `/etc/proxmox-backup/notifications-priv.cfg` (root:root 0600) | the Resend API key, as the SMTP password of target `felhom-operator` | from `$RESEND_API`, file → file | + +**Rotating the Resend key (§ above) must also rewrite `notifications-priv.cfg`**, or the ep0-copy failure mails stop +silently. Rotate the ep0 token with `proxmox-backup-manager user generate-token` on ep0 (delete the old) and rewrite +`remote.cfg`. See `runbooks/ep0-datastore-copy.md`. + +--- + ## Other committed secrets (tracked, NOT yet de-gitted — backlog) `manifests/felhom.secret.yaml` still commits other plaintext secrets (`healthchecks-config` `SECRET_KEY` diff --git a/hub/internal/offsitekeys/service_test.go b/hub/internal/offsitekeys/service_test.go index d8f17cf1..14e2de03 100644 --- a/hub/internal/offsitekeys/service_test.go +++ b/hub/internal/offsitekeys/service_test.go @@ -92,3 +92,38 @@ func TestWindow_NoConfirmedKeyRefused(t *testing.T) { t.Fatal("granted without a confirmed key") } } + +// A window the box never closes is closed by the hub at its bound: the deleting line goes, the ledger +// row closes with reason "timeout", and the operator hears offsite_window_failed. +func TestWindow_LeftOpenIsClosedByTheSweep(t *testing.T) { + s, fs, events := svcFixture(t) + ctx := context.Background() + pub, fp := newKey(t) + if _, err := s.RegisterKey(ctx, "c1", pub); err != nil { + t.Fatal(err) + } + if _, err := s.ConfirmKey(ctx, "c1", fp); err != nil { + t.Fatal(err) + } + _ = s.Store.GrantOffsiteWindowOnce("c1") + g, err := s.OpenWindowFor(ctx, "c1", 10) + if err != nil || !g.Granted { + t.Fatalf("%+v %v", g, err) + } + // Make it overdue: the box crashed and never reported. + if err := s.Store.ForceOffsiteWindowDueForTest(g.WindowID); err != nil { + t.Fatal(err) + } + s.SweepExpiredWindows(ctx) + if a := audit(fs.files[".ssh/authorized_keys"], "/home/felhom-repo", false); len(a.Findings) != 0 { + t.Fatalf("the sweep left the deleting line: %+v", a) + } + w, _ := s.Store.GetOffsiteWindow(g.WindowID) + if w == nil || w.ClosedAt.IsZero() || w.CloseReason != "timeout" { + t.Fatalf("ledger = %+v", w) + } + last := (*events)[len(*events)-1] + if last != EventWindowFailed { + t.Fatalf("last event = %s", last) + } +} diff --git a/hub/internal/store/offsite_keys.go b/hub/internal/store/offsite_keys.go index 2de3126d..2b1ed386 100644 --- a/hub/internal/store/offsite_keys.go +++ b/hub/internal/store/offsite_keys.go @@ -175,3 +175,9 @@ func (s *Store) TakeOffsiteWindowGrant(customerID string) bool { _ = s.setSetting(k, "") return true } + +// ForceOffsiteWindowDueForTest back-dates a window's closes_by. TEST-ONLY. +func (s *Store) ForceOffsiteWindowDueForTest(id int64) error { + _, err := s.db.Exec(`UPDATE offsite_windows SET closes_by = datetime('now', '-1 minute') WHERE id = ?`, id) + return err +}