Off-site lock live: Parts D/E/F evidence, ep0 copy runbook, 06/07 facts, register (R-820/R-821/R-342 closed, R-825 opened+closed, R-95/R-822 narrowed, R-823/R-824/R-826/R-827/R-828/R-830 opened; 327 -> 330); hub window-sweep test (test-only)
gates / gates (push) Successful in 41s

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
2026-10-03 21:21:04 +02:00
parent cdfcc47b15
commit 207ad19746
22 changed files with 370 additions and 15 deletions
@@ -182,7 +182,7 @@ production endpoint exists.
### 3.6 The ep0 datastore has a second copy — decision 70 (2026-10-03)
**[DESIGN]** A nightly PBS pull-sync copies `felhom-offsite` from ep0 to DooPlex's PBS over a read-only token. ep0's server snapshot never covered this volume and Hetzner has no volume snapshots (R-342). The copy is ciphertext per customer. Built per the 2026-10-03 lock brief Part F; the restore route lives in `runbooks/`.
**[DESIGN]** A nightly PBS pull-sync copies `felhom-offsite` from ep0 to DooPlex's PBS over a read-only token. ep0's server snapshot never covered this volume and Hetzner has no volume snapshots (R-342). The copy is ciphertext per customer. **`[FACT]` BUILT 2026-10-03:** ep0 token `root@pam!dooplex-sync` (`DatastoreReader` only — the one change on ep0); ep0's PBS listens on `wg0` only, so DooPlex reaches it through an SSH forward (`felhom-ep0-pbs-tunnel.service`, operator ruling the same day); DooPlex datastore `ep0-copy`, sync daily 05:00 with `remove-vanished false`, verify Saturdays, failures mailed via Resend. First pull 201 s / 12 GB / 4 of 4 snapshots. Restore route: `runbooks/ep0-datastore-copy.md`. Evidence `audits/offsite-lock-build-2026-10-03/partF/`.
## 4. Robustness (production details beyond the spike)
File diff suppressed because one or more lines are too long
@@ -0,0 +1,8 @@
## hub v0.127.0 rollout 2026-10-03T14:59:40Z
startup: off-site secrets sealed at rest (4 legacy plaintext row(s) sealed now)
startup: Off-site key registrar enabled; daily key check at 07:10 Budapest
raw DB after (prefix, length only):
demo-felhom enc:v1: 83
demo-hp enc:v1: 83
tester-1 enc:v1: 83
Tester-2 enc:v1: 83
@@ -19,3 +19,9 @@
## restored:
ok gitea.dooplex.hu/admin/felhom-controller/internal/backup 433.945s
ok gitea.dooplex.hu/admin/felhom-controller/internal/offsiteapply (cached)
FAIL
## RPC4: an unreadable count recorded as a measured zero (the v0.289.0 shape, live on demo-felhom)
=== RUN TestRunOffbox_UnreadableCountIsNotZero
offbox_window_test.go:266: an unreadable count was recorded as a measured zero: 0
--- FAIL: TestRunOffbox_UnreadableCountIsNotZero (0.00s)
@@ -0,0 +1,2 @@
demo-felhom before migration (controller 0.288.0, sftp transport, u629488-sub1): 11 snapshots
after two chain runs on 0.289.0/0.289.1 (pinned): 13 snapshots
@@ -0,0 +1,17 @@
transport=rclone-pinned target=u629488-sub1@u629488-sub1.your-storagebox.de:/home/felhom-repo
## count before
13 snapshots
## restore: one file from the newest snapshot 5dd1f0b9
restoring <Snapshot 5dd1f0b9 of [/mnt/sys_drive/felhom-data/backups/primary/opengist] at 2026-10-03 15:17:27.992813997 +0000 UTC by root@demo-felhom> to /tmp/rt
restored files: 2
sample: /mnt/sys_drive/felhom-data/backups/primary/opengist/data-stamps.json (398 bytes)
## check (exclusive lock — the weekly integrity job's op)
no errors were found
## delete attempt through the box's key: forget 6ea85413 (the OLDEST)
unable to remove <snapshot/6ea854132e> from the repository
[0:48] 0.00% 0 / 1 files deleted
blob not removed, server response: 403 Forbidden (403)
rc-line done
## count after
13 snapshots
@@ -0,0 +1,4 @@
gitea.dooplex.hu/admin/felhom-controller:0.288.0
offbox files: applied_marker known_hosts repo_password ssh_key
target: u629488-sub3@u629488-sub3.your-storagebox.de port=23 repo=/home/felhom-repo transport=sftp settings.snapshot_count=
91 snapshots
@@ -0,0 +1,8 @@
gitea.dooplex.hu/admin/felhom-controller:0.289.1
gitea.dooplex.hu/admin/felhom-controller:0.289.1 Up About a minute (healthy)
2026/10/03 15:19:24 offsiteapply.go:148: [INFO] [offsite-apply] settle-gate: awaiting floor knowledge (first report ACK) before offsite apply
2026/10/03 15:19:34 offsiteapply.go:148: [INFO] [offsite-apply] settle-gate: GO — at/above floor 0.288.0 (we are 0.289.1), no managed update running
2026/10/03 15:19:37 offsiteapply.go:148: [INFO] [offsite-apply] the hub installed key SHA256:16GpHoTF0WHLbj15Pog0mpmDMit4/BLFLRfIqmHJQc8 append-only on u629488-sub3@u629488-sub3.your-storagebox.de (fresh=false)
2026/10/03 15:19:37 offsiteapply.go:148: [INFO] [offsite-apply] offsite configured append-only for u629488-sub3@u629488-sub3.your-storagebox.de:/home/felhom-repo (key SHA256:16GpHoTF0WHLbj15Pog0mpmDMit4/BLFLRfIqmHJQc8)
2026/10/03 17:19:36 [INFO] offsitekeys: installed box key SHA256:16GpHoTF0WHLbj15Pog0mpmDMit4/BLFLRfIqmHJQc8 for demo-hp pinned append-only (u629488-sub3@u629488-sub3.your-storagebox.de, dropped 5 unpinned line(s)) in 938ms
2026/10/03 17:19:37 [INFO] offsitekeys: box confirmed key SHA256:16GpHoTF0WHLbj15Pog0mpmDMit4/BLFLRfIqmHJQc8 for demo-hp; 0 other line(s) removed
@@ -0,0 +1,24 @@
2026/10/03 15:22:53 night_chain.go:77: [INFO] [night-chain] offsite: started
2026/10/03 15:22:53 offbox.go:991: [INFO] [offbox] backup run started (9 app(s) toggled)
2026/10/03 15:24:49 offbox.go:1466: [WARN] [offbox] backed up bentopdf (/mnt/sys_drive/felhom-data/backups/primary/bentopdf, 0 mandatory path(s)) — but the recovery unit carried NO database dump and NO volume tar, so this snapshot holds none of the app's data; the next run with a dump leg will replace it
2026/10/03 15:25:12 offbox_window.go:178: [INFO] [offbox] retention skipped (after-run): no clean-up window now (weekly windows are off) — nothing deleted (decision 68)
2026/10/03 15:25:15 offbox.go:1234: [INFO] [offbox] backup OK: 9 app(s) backed up, 100 snapshot(s), 2m19s
2026/10/03 15:25:15 night_chain.go:82: [INFO] [night-chain] offsite: done in 2m23s
2026/10/03 15:25:15 night_chain.go:93: [INFO] [night-chain] finished in 4m14s
transport=rclone-pinned target=u629488-sub3@u629488-sub3.your-storagebox.de:/home/felhom-repo
## count before
100 snapshots
## restore: one file from the newest snapshot c43f2d0e
restoring <Snapshot c43f2d0e of [/mnt/sys_drive/felhom-data/backups/primary/opengist] at 2026-10-03 15:25:07.97694046 +0000 UTC by root@demo-hp> to /tmp/rt
restored files: 2
sample: /mnt/sys_drive/felhom-data/backups/primary/opengist/manifest.json (1767 bytes)
## check (exclusive lock — the weekly integrity job's op)
no errors were found
## delete attempt through the box's key: forget 05c3346a (the OLDEST)
unable to remove <snapshot/05c3346abc> from the repository
[0:48] 0.00% 0 / 1 files deleted
blob not removed, server response: 403 Forbidden (403)
rc-line done
## count after
100 snapshots
@@ -0,0 +1,40 @@
## POST /offsite/key-audit (operator, Basic auth) 2026-10-03T15:26:40Z
[
{
"customer": "Tester-2",
"lines": 0,
"pinned": 0,
"findings": null
},
{
"customer": "demo-felhom",
"lines": 1,
"pinned": 1,
"findings": null
},
{
"customer": "demo-hp",
"lines": 1,
"pinned": 1,
"findings": null
},
{
"customer": "tester-1",
"lines": 3,
"pinned": 0,
"findings": [
{
"Fingerprint": "SHA256:3UoXpMIvo9gat9AL1380UK39UGAtO2T2CBBllo8n3Mg",
"Kind": "unpinned"
},
{
"Fingerprint": "SHA256:Jri1gf2AGHCTj8cJrFSpRj8+BxdY64OQ5Rnms/tbOZI",
"Kind": "unpinned"
},
{
"Fingerprint": "SHA256:gAhxqeDkxAPTOOeKQDHBDNe/AYpARkxI8h7Rrc8pLD8",
"Kind": "unpinned"
}
]
}
]
@@ -0,0 +1,5 @@
hub=https://hub.felhom.eu keylen=64
consume-password HTTP 410
body: gone: the hub no longer serves the storage password; register the box's public key at /api/v1/offsite/register-key/
lines mentioning password field: 1
2026/10/03 17:26:59 [WARN] offsite consume-password called by demo-hp — retired (decision 69); the box must register its public key (controller >= 0.289.0)
@@ -0,0 +1,22 @@
grant: {"ok":true}
HTTP 202
2026/10/03 15:22:24 controller_image_retention.go:184: [INFO] [stacks] controller image retention: deleted gitea.dooplex.hu/admin/felhom-controller:0.287.0 (244595237106, 409MB) — older than the previous controller and no container uses it (decision 56)
2026/10/03 15:22:24 controller_image_retention.go:115: [INFO] [stacks] controller image retention: pass over 3 controller image(s) — running 0.289.1, previous "0.288.0" (by version order (no swap record names one present)), 1 candidate(s), 1 deleted, the rest kept
2026/10/03 15:25:12 offbox_window.go:178: [INFO] [offbox] retention skipped (after-run): no clean-up window now (weekly windows are off) — nothing deleted (decision 68)
2026/10/03 15:25:15 offbox.go:1234: [INFO] [offbox] backup OK: 9 app(s) backed up, 100 snapshot(s), 2m19s
2026/10/03 15:22:24 controller_image_retention.go:184: [INFO] [stacks] controller image retention: deleted gitea.dooplex.hu/admin/felhom-controller:0.287.0 (244595237106, 409MB) — older than the previous controller and no container uses it (decision 56)
2026/10/03 15:22:24 controller_image_retention.go:115: [INFO] [stacks] controller image retention: pass over 3 controller image(s) — running 0.289.1, previous "0.288.0" (by version order (no swap record names one present)), 1 candidate(s), 1 deleted, the rest kept
2026/10/03 15:25:12 offbox_window.go:178: [INFO] [offbox] retention skipped (after-run): no clean-up window now (weekly windows are off) — nothing deleted (decision 68)
2026/10/03 15:25:15 offbox.go:1234: [INFO] [offbox] backup OK: 9 app(s) backed up, 100 snapshot(s), 2m19s
2026/10/03 15:25:15 night_chain.go:93: [INFO] [night-chain] finished in 4m14s
2026/10/03 15:32:05 offbox_window.go:199: [ERROR] [offbox] clean-up window 1: the fake-snapshot guard REFUSED — nothing deleted: the policy would remove snapshot c6b5c67b from 2026-10-03T15:24:51Z — younger than 8 days, which honest retention never does (R-822)
2026/10/03 15:32:10 offbox.go:1234: [INFO] [offbox] backup OK: 9 app(s) backed up, 109 snapshot(s), 2m22s
2026/10/03 15:32:10 night_chain.go:93: [INFO] [night-chain] finished in 4m18s
2026/10/03 17:32:03 [WARN] offsitekeys: clean-up window 1 OPENED for demo-hp (key SHA256:16GpHoTF0WHLbj15Pog0mpmDMit4/BLFLRfIqmHJQc8, 109 snapshot(s), max 43 removed, closes by 2026-10-03T15:52:03Z, one-shot=true)
2026/10/03 17:32:06 [INFO] offsitekeys: clean-up window 1 CLOSED for demo-hp: outcome=guard-refused, 109 -> 109 (drop 0, allowed 43)
2026/10/03 17:32:06 [INFO] Operator email sent for demo-hp/offsite_prune_guard_refused
## key check after window 1 2026-10-03T15:32:32Z
Tester-2 lines 0 pinned 0 findings 0
demo-felhom lines 1 pinned 1 findings 0
demo-hp lines 1 pinned 1 findings 0
tester-1 lines 3 pinned 0 findings 3
@@ -0,0 +1,21 @@
## copy contents per namespace (filesystem listing, read-only)
ns/demo-felhom/ct/9201/2026-09-22T04:12:20Z
ns/demo-felhom/ct/9201/2026-09-29T04:16:43Z
ns/demo-hp/ct/9201/2026-09-24T20:06:25Z
ns/demo-hp/ct/9201/2026-10-01T20:15:29Z
## one snapshot's archives (no decryption — the copy is ciphertext; the catalog needs the customer's key)
4096 .
4096 ..
5576 catalog.pcat1.didx
1220 client.log.blob
690 index.json.blob
401 pct.conf.blob
179176 root.pxar.didx
## same view on ep0 (source)
ns/demo-felhom/ct/9201/2026-09-22T04:12:20Z
ns/demo-felhom/ct/9201/2026-09-29T04:16:43Z
ns/demo-hp/ct/9201/2026-09-24T20:06:25Z
ns/demo-hp/ct/9201/2026-10-01T20:15:29Z
@@ -0,0 +1,29 @@
## first pull 2026-10-03T19:13:17Z
----
Syncing datastore 'felhom-offsite', namespace 'demo-felhom' into datastore 'ep0-copy', namespace 'demo-felhom'
Created namespace demo-felhom
Found 1 groups to sync (out of 1 total)
[ct/9201]: 2026-09-22T04:12:20Z: start sync
[ct/9201]: 2026-09-22T04:12:20Z/pct.conf.blob: sync archive
[ct/9201]: 2026-09-22T04:12:20Z/root.pxar.didx: sync archive
[ct/9201]: 2026-09-22T04:12:20Z/root.pxar.didx: downloaded 2.057 GiB (62.788 MiB/s)
[ct/9201]: 2026-09-22T04:12:20Z/catalog.pcat1.didx: sync archive
[ct/9201]: 2026-09-22T04:12:20Z/catalog.pcat1.didx: downloaded 652.916 KiB (10.225 MiB/s)
[ct/9201]: Snapshot ct/9201/2026-09-22T04:12:20Z: got backup log file client.log.blob
[ct/9201]: 2026-09-22T04:12:20Z: sync done
[ct/9201]: percentage done: 50.00% (1/2 snapshots)
[ct/9201]: 2026-09-29T04:16:43Z: start sync
[ct/9201]: 2026-09-29T04:16:43Z/pct.conf.blob: sync archive
[ct/9201]: 2026-09-29T04:16:43Z/root.pxar.didx: sync archive
[ct/9201]: 2026-09-29T04:16:43Z/root.pxar.didx: downloaded 641.072 MiB (61.537 MiB/s)
[ct/9201]: 2026-09-29T04:16:43Z/catalog.pcat1.didx: sync archive
[ct/9201]: 2026-09-29T04:16:43Z/catalog.pcat1.didx: downloaded 768.894 KiB (16.864 MiB/s)
[ct/9201]: Snapshot ct/9201/2026-09-29T04:16:43Z: got backup log file client.log.blob
[ct/9201]: 2026-09-29T04:16:43Z: sync done
[ct/9201]: percentage done: 100.00% (2/2 snapshots)
Finished syncing namespace demo-felhom, current progress: 2 groups, 0 snapshots
pull datastore 'ep0-copy' end
TASK OK
elapsed 201 s
## bytes on DooPlex
12G /mnt/5_hdd/backup/ep0-copy
@@ -0,0 +1,10 @@
+====================+================+==========+========+================+==========+==============+=========+=================================================================================+
| id | sync-direction | store | remote | remote-store | schedule | group-filter | rate-in | comment |
+====================+================+==========+========+================+==========+==============+=========+=================================================================================+
| ep0-felhom-offsite | | ep0-copy | ep0 | felhom-offsite | 05:00 | all | | decision 70: nightly copy of ep0 felhom-offsite; never removes what ep0 removed |
+====================+================+==========+========+================+==========+==============+=========+=================================================================================+
+=================+==========+===========+=================+================+============================================+
| id | store | schedule | ignore-verified | outdated-after | comment |
+=================+==========+===========+=================+================+============================================+
| verify-ep0-copy | ep0-copy | sat 06:30 | 1 | 30 | decision 70: weekly verify of the ep0 copy |
+=================+==========+===========+=================+================+============================================+
@@ -0,0 +1,37 @@
## tunnel unit
[Unit]
Description=Felhom: SSH tunnel DooPlex 127.0.0.1:18007 -> ep0 PBS 127.0.0.1:8007 (decision 70, nightly pull-sync of felhom-offsite)
Documentation=file:///mnt/5_hdd/felhom.eu/git/felhom.eu/documentation/runbooks/ep0-datastore-copy.md
After=network-online.target
Wants=network-online.target
[Service]
User=kisfenyo
ExecStart=/usr/bin/ssh -N -o BatchMode=yes -o ExitOnForwardFailure=yes -o ServerAliveInterval=30 -o ServerAliveCountMax=3 -L 127.0.0.1:18007:127.0.0.1:8007 root@167.233.158.164
Restart=always
RestartSec=30
[Install]
WantedBy=multi-user.target
active
## notification target + matcher (password lives in notifications-priv.cfg, root:root 0600, not shown)
smtp: felhom-operator
author Felhom DooPlex PBS
comment decision 70: ep0 copy job failures to the operator (Resend)
from-address monitoring@felhom.eu
mailto admin@felhom.eu
mode tls
port 465
server smtp.resend.com
username resend
matcher: felhom-operator-errors
comment decision 70: any error (the ep0 pull-sync and verify jobs) reaches the operator
match-severity error
mode all
target felhom-operator
## test mail: received in the operator mailbox 2026-10-03T19:17:41Z, subject 'Test notification', from monitoring@felhom.eu to admin@felhom.eu (Gmail connector search)
## ep0 side: token root@pam!dooplex-sync, ACL DatastoreReader on /datastore/felhom-offsite (propagate) — the only change on ep0
+13
View File
@@ -26,6 +26,19 @@
---
## 2026-10-03 (evening) — the off-site lock built (hub v0.127.0, controller v0.289.0/0.289.1, decisions 68–70)
> Evidence: `audits/offsite-lock-build-2026-10-03/`.
| Row | What | Closed | Evidence |
|---|---|---|---|
| **R-820** | **A box could obtain its sub-account password at will (the hub's self-heal re-armed it; the box consumed it) — and that password removes any append-only pin.** Fixed: the box sends only its PUBLIC key; the hub (key registrar) writes it pinned; `consume-password` answers 410. Measured live: demo-hp's own API key gets `410 gone`, no password. **Reasoning kept: a pinned key protects nothing while any route can rewrite `authorized_keys` — the hub is now that file's only writer, and it reads every file daily.** | CLOSED 2026-10-03 — FIXED hub v0.127.0 + controller v0.289.0/0.289.1, proven live on both demo boxes | `partD/no-password-for-box.txt`, `partB/red-proofs-hub.txt`; hub `internal/offsitekeys` |
| **R-821** | **The hub DB held every sub-account password in the clear.** Fixed: AES-256-GCM at rest under `OFFSITE_SECRET_KEY` (Secret/offsite-secret-key, out of git); the 4 legacy rows sealed at start-up (raw rows read back `enc:v1:`); no key → the hub refuses to store or use one. **Reasoning kept: the running hub still holds the key and can open the passwords — a hub compromise remains an off-site compromise; this closes the database-copy route only.** | CLOSED 2026-10-03 — FIXED hub v0.127.0, verified on the live DB | `partB/hub-rollout.txt`; `TestOffsiteSecret_*` |
| **R-342** | **ep0's server snapshot never covered `/mnt/pbs-datastore`.** Built (decision 70): nightly PBS pull-sync to DooPlex (`ep0-copy`), `remove-vanished false`, weekly verify, failures mailed (test mail received). First pull 201 s / 12 GB / 4 of 4 snapshots, matching ep0. ep0 changed by one read-only token only; reached through an SSH forward from DooPlex (operator ruling). **Reasoning kept: Hetzner cannot snapshot a Volume — the copy is the only safeguard; never prune it tighter than ep0.** | CLOSED 2026-10-03 — BUILT; restore route unwalked (R-830), growth unbounded (R-828) | `partF/`; `runbooks/ep0-datastore-copy.md` |
| **R-825** | **Controller v0.289.0 reported 0 off-site snapshots as MEASURED over a store holding 12 (demo-felhom), and the hub mailed `offsite_snapshots_dropped` 11→0 — a false alarm.** Cause: the provider's rclone prints a NOTICE line that restic forwards into the combined output; every `--json` parse failed. Fixed in v0.289.1 within 15 minutes: the notice is stripped, and an unreadable count is never a measured zero (keeps the last value, `stats_known=false`). **Reasoning kept: a failed measurement must never be written as a measurement (R-331) — the detector that caught it is the one it would have blinded.** | CLOSED 2026-10-03 — FIXED controller v0.289.1 (red-proved), found and fixed in-session | operator mail 2026-10-03 17:17 CEST; `partC/red-proofs-controller.txt` RPC4 |
---
## 2026-10-03 — off-site append-only, measured on the provider (R-436, R-430)
> Spike, no product change. Evidence and design: `audits/offsite-append-only-2026-10-03/`.
File diff suppressed because one or more lines are too long
@@ -0,0 +1,50 @@
# Runbook — the ep0 datastore copy on DooPlex (decision 70, R-342)
ep0's PBS datastore `felhom-offsite` (`/mnt/pbs-datastore`, a separate Hetzner Volume that no server snapshot
covers and Hetzner cannot snapshot) is pulled to DooPlex every night. The copy holds **ciphertext only** — each
household's whole-box backups are encrypted with that household's own `encryption-key`; DooPlex cannot read them.
## What exists
| Where | Object | Purpose |
|---|---|---|
| ep0 | API token `root@pam!dooplex-sync`, ACL `DatastoreReader` on `/datastore/felhom-offsite` | read-only pull. **The only change on ep0.** |
| DooPlex | `felhom-ep0-pbs-tunnel.service` (systemd, runs as `kisfenyo`, `Restart=always`) | `ssh -N -L 127.0.0.1:18007:127.0.0.1:8007 root@ep0` — ep0's PBS listens on `wg0` only; DooPlex is not a WireGuard peer |
| DooPlex PBS | remote `ep0` (127.0.0.1:18007, ep0's cert fingerprint pinned) | the pull source |
| DooPlex PBS | datastore `ep0-copy` at `/mnt/5_hdd/backup/ep0-copy` | the copy |
| DooPlex PBS | sync job `ep0-felhom-offsite`, daily 05:00, `remove-vanished false` | the nightly pull (ep0's prune runs 03:30). **It never removes what ep0 removed** — a deletion on ep0 does not reach the copy |
| DooPlex PBS | verify job `verify-ep0-copy`, Saturdays 06:30 | reads the copy back |
| DooPlex PBS | notification target `felhom-operator` (SMTP via Resend → admin@felhom.eu) + matcher `felhom-operator-errors` (every error) | a failed pull or verify reaches the operator. Proven 2026-10-03 with a test mail |
Secrets, all out of git: the ep0 token secret in `/etc/proxmox-backup/remote.cfg` (root:backup 0640, base64 —
PBS's own format), the Resend key in `/etc/proxmox-backup/notifications-priv.cfg` (root:root 0600).
## Checks
```bash
systemctl is-active felhom-ep0-pbs-tunnel.service
curl -sk -o /dev/null -w '%{http_code}\n' https://127.0.0.1:18007/ # 200 = ep0's PBS reachable
sudo proxmox-backup-manager task list --limit 10 | grep -E 'syncjob|verif' # last runs
sudo find /mnt/5_hdd/backup/ep0-copy/ns -mindepth 4 -maxdepth 4 -type d | sort # snapshots per customer
```
## If ep0 is lost — restore a household's whole box from the DooPlex copy
The copy is a normal PBS datastore. Two routes, both needing the household's PBS `encryption-key` (escrowed,
recovered with the household's recovery code — the same as restoring from ep0):
1. **Point the box's host at DooPlex instead of ep0.** On the household's Proxmox host, add a PBS storage for
DooPlex's PBS (`ep0-copy`, namespace = the customer id) with the household's key, then restore the CT from it as
from ep0. DooPlex's PBS must be reachable from the host (it is not public today — the operator decides the route
at the time: a temporary tunnel, or a new endpoint).
2. **Rebuild the endpoint.** Provision a new ep0 (06 §5), then pull back: on the new ep0 add DooPlex as a remote
and run `proxmox-backup-manager pull <dooplex-remote> ep0-copy felhom-offsite`. Every box then reconnects as before.
Do **not** prune or garbage-collect the copy tighter than ep0's own retention. The copy has no prune job today and
grows with every nightly backup (R-828).
## Remove
`sudo proxmox-backup-manager sync-job remove ep0-felhom-offsite; … verify-job remove verify-ep0-copy;`
`sudo systemctl disable --now felhom-ep0-pbs-tunnel.service`; on ep0
`proxmox-backup-manager user delete-token root@pam dooplex-sync`. The datastore's bytes stay until removed by hand.
+15
View File
@@ -140,6 +140,21 @@ for every off-site customer.
---
## DooPlex PBS — the ep0 copy (decision 70, 2026-10-03)
Not k8s Secrets — PBS's own private files on DooPlex, never in git:
| File | Holds | Created |
|---|---|---|
| `/etc/proxmox-backup/remote.cfg` (root:backup 0640) | ep0 API token `root@pam!dooplex-sync` (base64, PBS format) | from the token's one-time output, file → file |
| `/etc/proxmox-backup/notifications-priv.cfg` (root:root 0600) | the Resend API key, as the SMTP password of target `felhom-operator` | from `$RESEND_API`, file → file |
**Rotating the Resend key (§ above) must also rewrite `notifications-priv.cfg`**, or the ep0-copy failure mails stop
silently. Rotate the ep0 token with `proxmox-backup-manager user generate-token` on ep0 (delete the old) and rewrite
`remote.cfg`. See `runbooks/ep0-datastore-copy.md`.
---
## Other committed secrets (tracked, NOT yet de-gitted — backlog)
`manifests/felhom.secret.yaml` still commits other plaintext secrets (`healthchecks-config` `SECRET_KEY`
+35
View File
@@ -92,3 +92,38 @@ func TestWindow_NoConfirmedKeyRefused(t *testing.T) {
t.Fatal("granted without a confirmed key")
}
}
// A window the box never closes is closed by the hub at its bound: the deleting line goes, the ledger
// row closes with reason "timeout", and the operator hears offsite_window_failed.
func TestWindow_LeftOpenIsClosedByTheSweep(t *testing.T) {
s, fs, events := svcFixture(t)
ctx := context.Background()
pub, fp := newKey(t)
if _, err := s.RegisterKey(ctx, "c1", pub); err != nil {
t.Fatal(err)
}
if _, err := s.ConfirmKey(ctx, "c1", fp); err != nil {
t.Fatal(err)
}
_ = s.Store.GrantOffsiteWindowOnce("c1")
g, err := s.OpenWindowFor(ctx, "c1", 10)
if err != nil || !g.Granted {
t.Fatalf("%+v %v", g, err)
}
// Make it overdue: the box crashed and never reported.
if err := s.Store.ForceOffsiteWindowDueForTest(g.WindowID); err != nil {
t.Fatal(err)
}
s.SweepExpiredWindows(ctx)
if a := audit(fs.files[".ssh/authorized_keys"], "/home/felhom-repo", false); len(a.Findings) != 0 {
t.Fatalf("the sweep left the deleting line: %+v", a)
}
w, _ := s.Store.GetOffsiteWindow(g.WindowID)
if w == nil || w.ClosedAt.IsZero() || w.CloseReason != "timeout" {
t.Fatalf("ledger = %+v", w)
}
last := (*events)[len(*events)-1]
if last != EventWindowFailed {
t.Fatalf("last event = %s", last)
}
}
+6
View File
@@ -175,3 +175,9 @@ func (s *Store) TakeOffsiteWindowGrant(customerID string) bool {
_ = s.setSetting(k, "")
return true
}
// ForceOffsiteWindowDueForTest back-dates a window's closes_by. TEST-ONLY.
func (s *Store) ForceOffsiteWindowDueForTest(id int64) error {
_, err := s.db.Exec(`UPDATE offsite_windows SET closes_by = datetime('now', '-1 minute') WHERE id = ?`, id)
return err
}