diff --git a/REPORT-offsite-append-only-2026-10-03.md b/REPORT-offsite-append-only-2026-10-03.md new file mode 100644 index 00000000..49f9cc40 --- /dev/null +++ b/REPORT-offsite-append-only-2026-10-03.md @@ -0,0 +1,85 @@ +# REPORT — off-site backup safety, step 1: the append-only lock measured on the provider (2026-10-03) + +A spike. No product code changed, no release. Evidence, exit test and design: +`documentation/audits/offsite-append-only-2026-10-03/`. Architecture read: `07-backup-architecture.md` +§8a, threat row 10, §D; `06-offsite-connectivity.md` (PBS/tunnel only — it does not describe the restic +tier, so the facts went to `07` §D). Baselines (re-verified): felhom.eu `f4c5466`, controller `0945332` +(v0.288.0), register 326 rows, highest id R-819. + +## The Part table + +| Part | done / not done / changed | why | +|---|---|---| +| 0 — venue | **changed** — `u629488-sub4` (tester-1) instead of a new scratch customer | operator ruled "use Tester1" in-session. tester-1's box was deleted 2026-09-30; nothing writes there. Credential: the hub's stored tester-1 value, read from a copy of the hub DB on a second operator ruling (copy deleted, value never printed or written to a committed file). A new repo dir `spike-r436` only; `felhom-repo` never read or written | +| A — the lock | **done**, exit test written first (`EXIT-TEST.md`) | E1–E8 and C1–C2 as stated; locks measured | +| B — the attacker | **done**, one item lab-only | raw-HTTP path escape through the pinned server measured in the lab only — a live HTTP/2 bridge over the forced ssh could not be made to work in the time box | +| C — design | **done** — `DESIGN.md`, STATUS decision 0 | | +| D — ep0 | **done** — `PART-D-ep0-safeguard.md`, STATUS decision 0b | read only; ep0 not touched | +| E — records | **done** | below | + +## Claims in the brief (and the register) that turned out wrong + +1. **"The box holds no sub-account password"** — it does not STORE one, but it can **obtain it at will**: + declare `needs_credential` twice → the hub re-arms the stored value → the box consumes it (R-820). +2. **"A forced command cannot be bypassed by the sub-account itself"** — the pinned key cannot; the + **password can** (logs in on ports 22 and 23, rewrote `authorized_keys` this session). +3. **"The hub cannot prune because of custody"** — true for *pruning*; but the hub can **delete**: it + holds every sub-account password in the clear (R-821). +4. **"Both `forget` sites must change together"** — there are **four** deleting features on the box: + both `forget` sites, the orphan move-aside (`mv`) and the abandonment (`rm -rf`). +5. **rclone in the image** (R-436 row: "rclone is not in the controller image today", implying it is + needed) — **not needed**; restic 0.14.0 with `-o rclone.program="ssh … rclone"` is enough. +6. **R-342's first candidate, a Hetzner Volume snapshot** — does not exist. +7. **R-430's model** (a locks dir where deletion is refused) — does not describe this transport; the + append-only server allows lock deletion and `unlock --remove-all` works. +8. **The vendor's cited blog** (`fluix.one`) shows the line WITHOUT `--append-only`; only Hetzner's + ticket reply adds it. Copying the blog would give a deleting key. +9. Held: restic is **0.14.0** (`0.14.0-1+b5`); **restore works through the add-only key**. + +## Part A — results (verbatim refusal) + +`blob not removed, server response: 403 Forbidden (403)` for `forget d807418c --prune`, +`forget --keep-last 1` and a real `prune` (each ~45–48 s of retries, rc=1); snapshot count unchanged; +control key: `1 / 1 files deleted`. Crash lock: blocks `check`, not `backup`; plain `unlock` prints +success and removes nothing; `--remove-all` removes it. Files: `live/E1-E3…`, `live/E4-E6…`, `live/C2-A5…`. + +## Part B — the attacker table + +| Route | Tried how | Result | What closes it | +|---|---|---|---| +| Password, port 23 | `sshpass ssh -p 23` | **logs in**; `authorized_keys` read and **rewritten** | box never receives it (hub = key registrar) | +| Password, port 22 | `sshpass sftp -P 22` | **logs in** (SFTP), `.ssh` listed | same | +| Box obtains the password | source read | **yes, at will** (self-heal re-arm + consume) | same — R-820 | +| Pinned key: shell / `rm -rf` | `ssh … 'ls'`, `'rm -rf spike-r436'` | runs the forced rclone; repo intact | — (holds) | +| Pinned key: sftp / scp / rsync | each | refused / protocol error | — (holds) | +| Pinned key: port forward | `-L`, then connect | `administratively prohibited` | — (holds) | +| Pinned key: other path, no flag | `rclone serve restic --stdio felhom-repo` | pinned dir served, append-only | — (holds) | +| Pinned key: `../` escapes, overwrite | raw HTTP (lab) | 400 / 403 | — (holds; lab rclone) | +| Pinned key: add junk / new `keys/` | raw HTTP (lab) | allowed | quota fills — R-431/quota alarms | +| Pinned key: future-dated snapshots | restic (lab) | allowed → retention erases real history | poisoning guard — R-822 | +| Any key on port 22 | both test keys | refused (port 22 takes no OpenSSH key) | — | +| Hetzner API / panel | box code read | nothing on the box reaches either | — | +| Hub DB | operator-tier | every sub-account password in clear | R-821 | + +**A route defeats the lock: the password (R-820).** The lock alone is not protection until it is closed. + +## Records + +- **Closed:** R-436 (measured; the 2026-10-06 due-check is cleared — the block is now empty), R-430. +- **Opened:** R-820 (P2, Security), R-821 (P2, Security), R-822 (P2, Backup). None is P1 by the + scale: today the box's own key can already delete (R-95), so none adds harm *today*. +- **Updated:** R-95 (the measurement, the four sites, the proposal; rank untouched), R-342 (options costed). +- **Register: 326 → 327** (`register_shape_gate`). All felhom.eu gates green. +- `07` §D: one `[FACT]` block. STATUS: two decisions in the operator's format. +- `unproven.py --summary`: NOT WALKED 35 of 55 — unchanged. + +## Teardown + +- **Provider:** `authorized_keys` restored — sha256 `795e7153…` before and after, identical; `spike-r436` + removed; `~/.config/rclone/` (created by the provider's rclone during the test) removed; home is back to + `.ssh`, `felhom-repo`. Both test keys refused afterwards. (`live/TEARDOWN.txt`) +- **DooPlex:** lab container, network and image removed; test keys, the password file, the hub DB copy + and hub page copies deleted from the scratchpad. +- **Hub:** nothing changed (two reads). +- **Left as is, on purpose:** the tester-1 sub-account password was NOT rotated — the next tester-1 install + needs the stored value. R-821 covers why that is itself a risk. diff --git a/STATUS.md b/STATUS.md index 907ada0c..d3ee22b4 100644 --- a/STATUS.md +++ b/STATUS.md @@ -2,9 +2,22 @@ **Ready for the first real tester (Tester-2): yes. You confirmed the tunnel route and the connect mails (2026-09-30).** -**Updated 2026-10-03 (the to-do list put in order). Versions unchanged since 2026-10-02 (afternoon). Both demo boxes run controller 0.288.0 and host agent 0.138.0. Hub 0.126.0. New +**Updated 2026-10-03 (off-site backup safety, step 1: measured on Hetzner). Versions unchanged since 2026-10-02 (afternoon). Both demo boxes run controller 0.288.0 and host agent 0.138.0. Hub 0.126.0. New installs get golden 0.288.0 with agent 0.138.0.** +## Today (2026-10-03, later): off-site backup safety, step 1 — measured, nothing built + +- **The "add only" lock works.** I tested it on tester-1's storage account (your choice; no box uses it now). + New backups go in. Restore works. Every delete is refused. I put the account back exactly as it was. +- **But the lock alone does not protect us yet.** A broken-into box can ask the hub for the storage password. + With that password it can log in and remove the lock. First the box must stop getting the password. +- **A second trap:** with "add only", an attacker can add fake backups dated in the future. The normal + clean-up rule then deletes all the real backups. Any clean-up must check for this. +- **The hub keeps every storage password in plain form.** Anyone who reads the hub database can delete + every household's off-site backups. New row. +- **ep0:** Hetzner cannot snapshot the extra disk at all. One of the three old ideas does not exist. +- **Rows:** 2 closed, 3 opened. The list went from 326 to 327 rows. The dated check for 6 October is done. + ## Today (2026-10-03): the to-do list is in order — paperwork only, no machine touched - **Finished items left the open list.** It went from 442 rows to 326. Nothing was deleted; each moved row names @@ -55,10 +68,23 @@ Your licence decisions are recorded: Emby, Plex and n8n stay. recipe-importer ne ## What needs you -0. **Pick the next theme.** **A — off-site backup safety** (recommended): a dated check on it turns red on - 6 October and then refuses every push until it is done; it guards the household's own files. **B — box system - security updates**: nothing patches a box today; a test on a throwaway box comes first. **If you say nothing:** - the next session starts A. 18 rows wait on you; the list is in the triage recommendation. +0. **Who may delete old off-site backups, once boxes can only add?** (Details: the design in the + 2026-10-03 off-site audit folder.) + - **A — the box, in a short weekly window the hub opens** (recommended). The backup password stays only on + the box, as we promise today. Cost: during the window a broken-into box could delete; the hub checks the + count before and after. + - **B — a Felhom machine does it for every box.** Cost: that machine must hold every household's backup + password, so it could read every household's backups. That changes a promise to the customer, and adds a + new always-on machine. + - **If you say nothing:** nothing is built. Boxes keep the key that can delete (today's risk stays). + - Either way, first: the hub installs the box's key, so the box never gets the storage password. +0b. **ep0's backup disk has no copy of its own. Which safeguard?** + - **A — DooPlex copies it every night** (recommended). €0 a month, about 1–2 hours to set up. Protects against + losing the disk and losing Hetzner. The copy is encrypted per household, so DooPlex cannot read it. Cost: a + new job on DooPlex. + - **B — accept the risk in writing,** and copy the disk off by hand before any risky work on ep0. + - **If you say nothing:** the disk stays unprotected; a bad day on ep0 loses every household's whole-box + off-site copy. 1. **plant-it:** keep the hidden template as it is, or remove it entirely (its image no longer exists). **If you say nothing:** it stays hidden; nothing runs it. 2. **Send the SparkyFitness request, and ask the Tandoor authors** (the "Before the first paying customer" list). diff --git a/documentation/architecture/07-backup-architecture.md b/documentation/architecture/07-backup-architecture.md index ad4315f7..d7d82d64 100644 --- a/documentation/architecture/07-backup-architecture.md +++ b/documentation/architecture/07-backup-architecture.md @@ -1367,6 +1367,19 @@ answered by calling the provider API with the production token. Accept explicitl > spike stopped here rather than answering it by acting**, and left it as R-429 — ten minutes in > the Storage Box panel, and it re-ranks R-95. `audits/SPIKE-r95-offsite-delete-2026-09-01.md` §Q1. +> **`[FACT]` 2026-10-03 — append-only on the Storage Box, MEASURED (R-436).** On the provider +> (`u629488-sub4`, scratch repo, since removed), a key pinned in `authorized_keys` to +> `command="rclone serve restic --stdio --append-only ",restrict` backs up, lists, restores and +> checks, and every delete is refused: `blob not removed, server response: 403 Forbidden (403)`; the +> same command through an unpinned key deletes. The pinned key gets no shell, sftp, scp, rsync or port +> forward. **But the PASSWORD logs in on ports 22 and 23 and can rewrite `authorized_keys`, and a box can +> obtain that password from the hub's off-site self-heal at will** — so the pin protects nothing until +> the box stops receiving the password (R-820). Port 22 accepts no OpenSSH-format key; 23 does. +> Locks: a crash lock blocks `check`, not `backup`; `unlock --remove-all` clears it through the pinned +> key. An add-only key can still plant future-dated snapshots that make the box's retention policy +> select every real snapshot (R-822). Who prunes is a PROPOSAL awaiting the operator, not design: +> `audits/offsite-append-only-2026-10-03/DESIGN.md`; evidence `…/live/`, `…/lab/`. + **E. `local` vzdump shares a physical device with the guest it backs up.** LIVE on both hosts: `/var/lib/vz` (the archive target) and `local-lvm` (the guest's rootfs and both `backup=1` mountpoints) are both on `/dev/sda3` → VG `pve`. The tier therefore protects against **corruption diff --git a/documentation/audits/offsite-append-only-2026-10-03/DESIGN.md b/documentation/audits/offsite-append-only-2026-10-03/DESIGN.md new file mode 100644 index 00000000..8a259a23 --- /dev/null +++ b/documentation/audits/offsite-append-only-2026-10-03/DESIGN.md @@ -0,0 +1,118 @@ +# Off-site append-only — who deletes old backups when the box cannot? (PROPOSAL, 2026-10-03) + +**Status: a PROPOSAL for the operator to rule on. Nothing here is built. Nothing here is `[DESIGN]`.** +Evidence: `live/` (the provider, `u629488-sub4`) and `lab/` (OpenSSH + rclone v1.75.1 on DooPlex). +Baselines: felhom.eu `f4c5466`, controller `0945332` (v0.288.0, restic **0.14.0** Debian `0.14.0-1+b5`). + +## 1. What was measured (the facts this design stands on) + +| # | Fact | Where | +|---|---|---| +| F1 | A key whose `authorized_keys` line is `command="rclone serve restic --stdio --append-only ",restrict` can `init`, `backup`, `snapshots`, `restore` (bytes identical) and `check`. | `live/E1-E3`, `live/E4-E6`, `live/C2-A5` | +| F2 | Through that key `forget --prune`, `forget --keep-last 1` and a real `prune` are REFUSED: `blob not removed, server response: 403 Forbidden (403)`, rc=1, snapshot count unchanged. Each refusal costs **~45–48 s** of restic retries. A refused `prune` has already **written a new index** first (`check`: duplicate index, non-critical). | `live/E4-E6` | +| F3 | Control: the same `forget` through a key WITHOUT the forced command deletes (1/1 files deleted). | `live/E4-E6` (C1) | +| F4 | The client's requested path and flags are ignored: the pinned directory is served even when the client names another path, and the client never asked for `--append-only` (restic sends `serve restic --stdio --b2-hard-delete`). | `live/C2-A5`, `live/B2` | +| F5 | The forced key cannot get a shell, `sftp`, `scp`, `rsync` or a port forward (`administratively prohibited`); a command such as `rm -rf ` just starts the forced rclone. | `live/B2`, `live/B1-B2` | +| F6 | Port **22 accepts no OpenSSH-format key at all** (both test keys refused there); keys work on port **23** only — the port the product uses (`hub/internal/offsite/offsite.go` `sftpPort = 23`). **The PASSWORD logs in on BOTH ports**, and through it `.ssh/authorized_keys` was read and REWRITTEN (that is how the test keys were installed). | `live/B1-B2`, this session | +| F7 | Crash locks: a killed backup leaves a lock. The next **backup is not blocked**; an **exclusive** operation (`check` — the weekly integrity job) **is**. Plain `unlock` prints `successfully removed locks` and removes nothing (the lock is not yet stale — by design, R-430 reproduced); `unlock --remove-all` **does remove it** — the append-only server allows lock deletion. | `live/C2-A5`, `lab/A5` | +| F8 | A crashed upload leaves unreferenced packs (lab: 17) that only a deleting key can clean. | `lab/A5` | +| F9 | The append-only server refuses delete/overwrite of existing files (403) and path escapes (`../`, `%2e%2e`, `..%2f` → 400), but ACCEPTS new files of any name inside the repo, including a new `keys/` entry (quota can be filled). | `lab/B-rclone-server-surface` (lab rclone; the provider's rclone version cannot be read — `rclone version` is not a shell command there) | +| F10 | **Retention poisoning.** Through the add-only key, 13 future-dated empty snapshots with the same host+tag make the box's exact policy (`--group-by host,tags --keep-daily 7 --keep-weekly 4 --keep-monthly 6`) select **all 3 real snapshots for removal**. | `lab/C3-retention-poisoning` | +| F11 | The provider's rclone creates `~/.config/rclone/` (empty) in the sub-account home on first use. | `live/TEARDOWN` | + +From the code (read, not run): + +| # | Fact | Where | +|---|---|---| +| K1 | The box can obtain its sub-account password **at will**: it declares `needs_credential` in two reports, the hub's off-site self-heal re-arms the **stored** value (`RestageOneTimeSecret`; the value is never cleared after a consume) or escalates to a provider re-issue, and the box consumes it with its own API key. | `hub/internal/offsiteheal/reconciler.go`, `hub/internal/store/store.go` `ConsumeOneTimeSecret`/`RestageOneTimeSecret`, `controller/internal/offsiteapply/seams.go` | +| K2 | The hub DB holds every sub-account password in the clear, indefinitely (`one_time_secrets.value`). | `hub/internal/store/store.go` | +| K3 | Four box features need a normal (deleting) login, not just the two `forget` sites: `forget --prune` after a run (`offbox.go` ~1422), over quota (`offbox.go` ~1793), the orphan **move-aside** (`mv`, `offbox.go` `resetOrphanedRepo`), and the customer-chosen **abandonment** (`rm -rf`, `offbox_abandon.go` `AbandonSweep`). | controller `0945332` | +| K4 | restic has no prune-only credential: every key unlocks the same master key, so whoever prunes can read every backup in that repo. | restic design (documented, not measured) | + +## 2. The prerequisite every option shares — the lock is worthless until this is closed + +**F6 + K1: a box that is broken into can get the password, log in on port 23 with it, and rewrite +`authorized_keys` — removing the forced command from its own key.** So before any lock ships: + +1. **The box never receives the sub-account password.** The hub becomes the key registrar: the box + generates its key pair and sends the PUBLIC key to the hub; the hub (which already holds the + password, K2) writes the forced line into `authorized_keys` over SFTP port 23. The off-site + self-heal re-arms the HUB's install job, not a password for the box. +2. **The hub checks the file.** Each day the hub reads `authorized_keys` and alarms if any line lacks + the forced prefix (outside an open clean-up window). This is also the answer to "an operator adds a + key later": a key without the prefix re-opens deletion, so a check — not a rule — catches it. +3. Optional hardening, operator's call: rotate the sub-account password after each hub use and keep it + only in the hub (K2 stays true either way: **the hub can already delete every household's off-site + history today** — new row). + +## 3. The options + +### Option 1 — a clean-up window opened by the hub (the box keeps pruning its own repo) + +The box normally holds only its add-only key. Once a week the hub adds a second, deleting key line for +the same box key pair (or a second key the box holds) for ~15 minutes; the box runs its retention; the +hub removes the line and checks the file. + +- **Custody: unchanged.** The repository password never leaves the box (07 §8a stays true). +- **What a compromised box can do:** delete during a window. Detection backstop: R-431 (a fall of more + than half), plus a new per-window check — the hub compares the snapshot count before and after and + alarms on a fall larger than the retention could cause. +- **Poisoning (F10) must be guarded even here**, because an attacker who was on the box and left can + plant future snapshots that the next honest window then obeys: before `forget`, refuse when any + snapshot is dated in the future or newer than the newest the hub has seen reported; run `--dry-run` + first and abort if it would remove more snapshots than the policy can remove in a week. +- **Build cost:** hub: key registrar + window open/close + `authorized_keys` check (SFTP with the + password it holds — no provider API, no main account). Controller: switch the transport from `sftp:` + to `rclone:` with `-o rclone.program="ssh -p 23 … -i … rclone"` (restic 0.14.0 is enough, F1; + **rclone is NOT needed in the image**); move `forget --prune` (both sites, together — R-191) into + the window; move the move-aside and the abandonment to the hub (K3). +- **Money:** none. + +### Option 2 — clean-up runs off the box + +A Felhom-side worker prunes each repo with a deleting key. + +- **Custody: CHANGES.** restic has no prune-only key (K4): the worker must hold every repository + password in a form it can use unattended, and could read every household's backups. That changes a + promise to the customer (07 §8a) — not CC's to change. +- **Poisoning (F10):** same guard needed. +- **Build cost:** a new always-on worker in the recovery path, its own credentials store, its own + failure alarms — plus everything in §2. + +### Option 3 — never delete; let the quota grow + +- **Money:** small. The pool box (BX11, 1 TB) was €4.06/month when recorded (`SPIKE-ep0-storagebox- + 2026-07-09`; Hetzner raised prices in April 2026 — re-check) ≈ **€0.004 per GB per month**. A + household adding 10 GB a year unpruned costs ≈ **€0.50 a year**. The demo boxes cannot calibrate + this (0.5 GB and 0.0 GB used today), so 10 GB/year is an assumption, stated as one. +- **The real cost is the quota (50–100 GB):** at the quota `offbox_fit.go` refuses new pushes and the + over-quota `forget` (K3) would itself be refused, so the household's off-site copy STOPS. Crash + leftovers (F8) and refused-prune index files (F2) also accumulate. +- Good as an **interim**, not an end state. + +## 4. Recommendation + +**Option 1, with Option 3 as the interim while it is built.** Order: + +1. §2 first (key registrar + the `authorized_keys` check) — without it nothing below protects anything. +2. Switch every box to the forced key with **no** box-side retention (Option 3 interim). Quotas today + are 50–100 GB against ≤0.5 GB used, so this costs nothing for months. +3. Then the weekly window (Option 1) with the poisoning guard. + +Why not Option 2: it trades a box-level risk for a custody change that lets one Felhom machine read +every household's backups, and it adds an always-on service to the recovery path. + +## 5. Migration, rotation, restore + +- **Existing repos (both demo boxes, any household):** only the `authorized_keys` line and the + box's transport change; the repository bytes are not touched. **Not measured:** reading a repo + written over `sftp:` through `rclone:` — restic's on-disk layout is backend-independent, so it is + expected to work; measure on the scratch account before any household. History is kept. +- **Key rotation:** the hub writes the new forced line, the box switches, the hub removes the old line — + the same registrar path. +- **Restore** works through the add-only key (F1) — no second key is needed for the household. +- **Locks:** the self-heal (`resticStep` → `unlock --remove-all`) keeps working under this transport + (F7). R-430's fear does not apply to it. +- **Failure speed:** with today's code a forced key makes each night's `forget --prune` fail after + ~45 s per refused file and grow the index (F2) — the two `forget` sites must leave the box in the + same change that switches the key. diff --git a/documentation/audits/offsite-append-only-2026-10-03/EXIT-TEST.md b/documentation/audits/offsite-append-only-2026-10-03/EXIT-TEST.md new file mode 100644 index 00000000..e29e17ca --- /dev/null +++ b/documentation/audits/offsite-append-only-2026-10-03/EXIT-TEST.md @@ -0,0 +1,33 @@ +# R-436 exit test — written BEFORE any command ran (2026-10-03 ~11:10 CEST) + +The lock HOLDS only if ALL of these are true on the LIVE Storage Box (`u629488`, scratch venue +`u629488-sub4`, a dedicated repo path that is NOT `felhom-repo`): + +| # | Through the FORCED key (`command="rclone serve restic --stdio --append-only "`) | Must be | +|---|---|---| +| E1 | `restic init` (scratch repo) | succeeds | +| E2 | `restic backup` of a small tree | succeeds, snapshot count +1 | +| E3 | `restic snapshots`, `restic restore` of one file (bytes compared) | succeeds | +| E4 | `restic forget --prune` | REFUSED — verbatim error quoted; count unchanged | +| E5 | `restic prune` | REFUSED — verbatim; `check` clean | +| E6 | `restic forget --keep-last 1` | REFUSED — verbatim; count unchanged | +| E7 | the same key asking for a shell / `sftp` / `scp` / `rsync` / a port forward | REFUSED | +| E8 | the same key, client asks for a plain `rclone serve restic --stdio` (no flag) | still append-only | + +Controls (without them a refusal means nothing): + +| # | Control | Must be | +|---|---|---| +| C1 | the same `forget --prune` through a key WITHOUT the forced command | SUCCEEDS (the test can see a delete) | +| C2 | the forced key, client names a DIFFERENT repo path | recorded as observed (expected: the pinned path is served regardless) | + +Locks (R-430): kill a backup mid-run through the forced key; record whether the next backup +wedges, and what `unlock` / `unlock --remove-all` report AND what remains on disk. + +**Verdict rule:** E1–E8 as stated and C1 succeeding ⇒ the lock holds *for that key*. Part B then asks +whether any OTHER credential reachable from a box defeats it; if one does, the lock alone is not +protection, and that is reported first. + +A LOCAL LAB (`lab/`) runs the same matrix against OpenSSH + rclone on DooPlex. It measures rclone's +and restic's behaviour; it CANNOT stand in for the provider's sshd honouring `command=`. Only `live/` +can close R-436. diff --git a/documentation/audits/offsite-append-only-2026-10-03/PART-D-ep0-safeguard.md b/documentation/audits/offsite-append-only-2026-10-03/PART-D-ep0-safeguard.md new file mode 100644 index 00000000..c0c79226 --- /dev/null +++ b/documentation/audits/offsite-append-only-2026-10-03/PART-D-ep0-safeguard.md @@ -0,0 +1,29 @@ +# R-342 — the ep0 datastore safeguard: options (2026-10-03, read only, nothing changed on ep0) + +**The gap:** ep0's server snapshot covers the 38 GB system disk, not `/mnt/pbs-datastore` (a separate +100 GB Hetzner Cloud Volume). The hub's Offsite page read today: datastore `felhom-offsite`, +**19.5 GB used of 97.9 GB (20 %)**. ep0 holds the only off-premises copy of the households' +whole-box backups. Those backups are **client-side encrypted per customer** (`encryption-key` in each +box's `storage.cfg`, 07 §8a) — a copy elsewhere holds ciphertext only. + +## A correction first + +**R-342's first candidate does not exist.** Hetzner Cloud has **no Volume snapshots**: server +snapshots and backups exclude attached volumes, and volume snapshots have been an open feature request +since 2019 (hetznercloud/csi-driver issue #79; simplebackups.com Hetzner note). Not measured on the +panel; documented by third parties and consistent with R-342's own evidence file. + +## The options + +| | What | Money / month | Setup | What it protects | +|---|---|---|---|---| +| **A. PBS pull-sync to DooPlex's PBS** | A sync job on DooPlex's existing PBS pulls the `felhom-offsite` datastore from ep0 (read-only token on ep0, `DatastoreReader`), e.g. nightly | **€0** (DooPlex `/mnt/5_hdd` has 5.5 TB free; ~20 GB now) | ~1–2 h: a PBS remote + sync job + a read-only token; one route from DooPlex to ep0 :8007 must be chosen (the tunnel or the public address) | A lost/corrupted ep0 datastore, a bad ep0 procedure, and losing Hetzner entirely (a different provider and site). Ciphertext only — no custody change. **Cost:** a new job on DooPlex (Tier 2 — the operator's call, not CC's) | +| **B. Written acceptance + a copy before each risky procedure** | Accept that the datastore is unprotected in normal running; any runbook step that can touch `/mnt/pbs-datastore` first copies it off (e.g. to DooPlex, `rsync -a` **without** `-H` — `-H` was measured running out of memory on a PBS chunk store here, and PBS uses no hard links) | €0 | ~10 min to write the rule; ~20–40 min per procedure | Only the procedure itself. Not disk loss, not the provider, not a mistake between procedures | +| ~~Hetzner Volume snapshot~~ | does not exist | — | — | — | +| A second Hetzner volume / server | another volume ≈ €0.057/GB/month (April 2026 price, third-party source) → 100 GB ≈ **€5.70** + a server to attach it to | €5–10 | hours | Disk loss only; same provider, same failure domain (07 §D) | + +## Recommendation + +**A.** It costs no money, protects against everything B protects against and more, and keeps the +custody model because the data is already encrypted per customer. Its one real cost is a new job on +DooPlex, which is why it is the operator's decision. diff --git a/documentation/audits/offsite-append-only-2026-10-03/lab/A1-A2-init-backup.txt b/documentation/audits/offsite-append-only-2026-10-03/lab/A1-A2-init-backup.txt new file mode 100644 index 00000000..f1aed3c2 --- /dev/null +++ b/documentation/audits/offsite-append-only-2026-10-03/lab/A1-A2-init-backup.txt @@ -0,0 +1,43 @@ +$ c.sh forced spike-repo init +rclone: 2026/10/03 09:07:47 NOTICE: Config file "/home/sub/.config/rclone/rclone.conf" not found - using defaults +created restic repository 54423dc079 at rclone:lab:spike-repo + +Please note that knowledge of your password is required to access +the repository. Losing your password means that your data is +irrecoverably lost. +[rc=0] + +$ c.sh forced spike-repo backup /d/tree --host labbox --tag app1 +rclone: 2026/10/03 09:07:51 NOTICE: Config file "/home/sub/.config/rclone/rclone.conf" not found - using defaults +no parent snapshot found, will read all files + +Files: 4 new, 0 changed, 0 unmodified +Dirs: 3 new, 0 changed, 0 unmodified +Added to the repository: 588.177 KiB (587.567 KiB stored) + +processed 4 files, 585.948 KiB in 0:00 +snapshot 63a075dc saved +[rc=0] + +$ c.sh forced spike-repo backup /d/tree --host labbox --tag app1 +rclone: 2026/10/03 09:07:53 NOTICE: Config file "/home/sub/.config/rclone/rclone.conf" not found - using defaults +using parent snapshot 63a075dc + +Files: 0 new, 0 changed, 4 unmodified +Dirs: 0 new, 0 changed, 3 unmodified +Added to the repository: 0 B (0 B stored) + +processed 4 files, 585.948 KiB in 0:00 +snapshot 066a039d saved +[rc=0] + +$ c.sh forced spike-repo snapshots +rclone: 2026/10/03 09:07:55 NOTICE: Config file "/home/sub/.config/rclone/rclone.conf" not found - using defaults +ID Time Host Tags Paths +-------------------------------------------------------------- +63a075dc 2026-10-03 09:07:50 labbox app1 /d/tree +066a039d 2026-10-03 09:07:52 labbox app1 /d/tree +-------------------------------------------------------------- +2 snapshots +[rc=0] + diff --git a/documentation/audits/offsite-append-only-2026-10-03/lab/A3-deletes-forced.txt b/documentation/audits/offsite-append-only-2026-10-03/lab/A3-deletes-forced.txt new file mode 100644 index 00000000..d70d63eb --- /dev/null +++ b/documentation/audits/offsite-append-only-2026-10-03/lab/A3-deletes-forced.txt @@ -0,0 +1,138 @@ +server files before: snapshots=2 keys=1 index=1 data=2 locks=0 +$ c.sh forced spike-repo restore 63a075dc --target /d/restored --include /d/tree/sub/note.txt +restoring to /d/restored +[rc=0] + +restored bytes: +$ c.sh forced spike-repo forget 63a075dc --prune +Remove() returned error, retrying after 720.254544ms: blob not removed, server response: 403 Forbidden (403) +Remove() returned error, retrying after 873.42004ms: blob not removed, server response: 403 Forbidden (403) +Remove() returned error, retrying after 1.054928461s: blob not removed, server response: 403 Forbidden (403) +Remove() returned error, retrying after 1.560325776s: blob not removed, server response: 403 Forbidden (403) +Remove() returned error, retrying after 3.004145903s: blob not removed, server response: 403 Forbidden (403) +Remove() returned error, retrying after 2.147653057s: blob not removed, server response: 403 Forbidden (403) +Remove() returned error, retrying after 3.739082318s: blob not removed, server response: 403 Forbidden (403) +Remove() returned error, retrying after 5.099891944s: blob not removed, server response: 403 Forbidden (403) +Remove() returned error, retrying after 10.263247495s: blob not removed, server response: 403 Forbidden (403) +Remove() returned error, retrying after 19.514091959s: blob not removed, server response: 403 Forbidden (403) +unable to remove from the repository +[0:48] 0.00% 0 / 1 files deleted + +blob not removed, server response: 403 Forbidden (403) +github.com/restic/restic/internal/backend/rest.(*Backend).Remove + github.com/restic/restic/internal/backend/rest/rest.go:399 +github.com/restic/restic/internal/backend.(*RetryBackend).Remove.func1 + github.com/restic/restic/internal/backend/backend_retry.go:108 +github.com/cenkalti/backoff.RetryNotifyWithTimer + github.com/cenkalti/backoff/retry.go:55 +github.com/cenkalti/backoff.RetryNotify + github.com/cenkalti/backoff/retry.go:34 +github.com/restic/restic/internal/backend.(*RetryBackend).retry + github.com/restic/restic/internal/backend/backend_retry.go:46 +github.com/restic/restic/internal/backend.(*RetryBackend).Remove + github.com/restic/restic/internal/backend/backend_retry.go:107 +github.com/restic/restic/internal/cache.(*Backend).Remove + github.com/restic/restic/internal/cache/backend.go:38 +main.deleteFiles.func2 + github.com/restic/restic/cmd/restic/delete.go:47 +golang.org/x/sync/errgroup.(*Group).Go.func1 + golang.org/x/sync/errgroup/errgroup.go:75 +runtime.goexit + runtime/asm_amd64.s:1594 +[rc=1] + +$ c.sh forced spike-repo prune +loading indexes... +loading all snapshots... +finding data that is still in use for 2 snapshots +[0:00] 100.00% 2 / 2 snapshots + +searching used packs... +collecting packs for deletion and repacking +[0:00] 100.00% 2 / 2 packs processed + + +to repack: 0 blobs / 0 B +this removes: 0 blobs / 0 B +to delete: 0 blobs / 0 B +total prune: 0 blobs / 0 B +remaining: 8 blobs / 587.247 KiB +unused size after prune: 0 B (0.00% of remaining size) + +done +[rc=0] + +$ c.sh forced spike-repo forget --keep-last 1 +Applying Policy: keep 1 latest snapshots +keep 1 snapshots: +ID Time Host Tags Reasons Paths +----------------------------------------------------------------------------- +066a039d 2026-10-03 09:07:52 labbox app1 last snapshot /d/tree +----------------------------------------------------------------------------- +1 snapshots + +remove 1 snapshots: +ID Time Host Tags Paths +-------------------------------------------------------------- +63a075dc 2026-10-03 09:07:50 labbox app1 /d/tree +-------------------------------------------------------------- +1 snapshots + +Remove() returned error, retrying after 720.254544ms: blob not removed, server response: 403 Forbidden (403) +Remove() returned error, retrying after 873.42004ms: blob not removed, server response: 403 Forbidden (403) +Remove() returned error, retrying after 1.054928461s: blob not removed, server response: 403 Forbidden (403) +Remove() returned error, retrying after 1.560325776s: blob not removed, server response: 403 Forbidden (403) +Remove() returned error, retrying after 3.004145903s: blob not removed, server response: 403 Forbidden (403) +Remove() returned error, retrying after 2.147653057s: blob not removed, server response: 403 Forbidden (403) +Remove() returned error, retrying after 3.739082318s: blob not removed, server response: 403 Forbidden (403) +Remove() returned error, retrying after 5.099891944s: blob not removed, server response: 403 Forbidden (403) +Remove() returned error, retrying after 10.263247495s: blob not removed, server response: 403 Forbidden (403) +Remove() returned error, retrying after 19.514091959s: blob not removed, server response: 403 Forbidden (403) +unable to remove from the repository +[0:48] 0.00% 0 / 1 files deleted + +blob not removed, server response: 403 Forbidden (403) +github.com/restic/restic/internal/backend/rest.(*Backend).Remove + github.com/restic/restic/internal/backend/rest/rest.go:399 +github.com/restic/restic/internal/backend.(*RetryBackend).Remove.func1 + github.com/restic/restic/internal/backend/backend_retry.go:108 +github.com/cenkalti/backoff.RetryNotifyWithTimer + github.com/cenkalti/backoff/retry.go:55 +github.com/cenkalti/backoff.RetryNotify + github.com/cenkalti/backoff/retry.go:34 +github.com/restic/restic/internal/backend.(*RetryBackend).retry + github.com/restic/restic/internal/backend/backend_retry.go:46 +github.com/restic/restic/internal/backend.(*RetryBackend).Remove + github.com/restic/restic/internal/backend/backend_retry.go:107 +github.com/restic/restic/internal/cache.(*Backend).Remove + github.com/restic/restic/internal/cache/backend.go:38 +main.deleteFiles.func2 + github.com/restic/restic/cmd/restic/delete.go:47 +golang.org/x/sync/errgroup.(*Group).Go.func1 + golang.org/x/sync/errgroup/errgroup.go:75 +runtime.goexit + runtime/asm_amd64.s:1594 +[rc=1] + +$ c.sh forced spike-repo snapshots +ID Time Host Tags Paths +-------------------------------------------------------------- +63a075dc 2026-10-03 09:07:50 labbox app1 /d/tree +066a039d 2026-10-03 09:07:52 labbox app1 /d/tree +-------------------------------------------------------------- +2 snapshots +[rc=0] + +$ c.sh forced spike-repo check +using temporary cache in /tmp/restic-check-cache-3313406781 +create exclusive lock for repository +load indexes +check all packs +check snapshots, trees and blobs +[0:00] 100.00% 2 / 2 snapshots + +no errors were found +[rc=0] + +server files after: snapshots=2 keys=1 index=1 data=2 locks=0 +cmp restored vs source: IDENTICAL (hello-r436) diff --git a/documentation/audits/offsite-append-only-2026-10-03/lab/A4-controls-and-real-prune.txt b/documentation/audits/offsite-append-only-2026-10-03/lab/A4-controls-and-real-prune.txt new file mode 100644 index 00000000..0a4c2eaf --- /dev/null +++ b/documentation/audits/offsite-append-only-2026-10-03/lab/A4-controls-and-real-prune.txt @@ -0,0 +1,67 @@ +## E5 made real: unreferenced data exists, then prune through the FORCED key +$ c.sh forced spike-repo backup /d/tree2 --host labbox --tag app2 +no parent snapshot found, will read all files + +Files: 1 new, 0 changed, 0 unmodified +Dirs: 2 new, 0 changed, 0 unmodified +Added to the repository: 293.934 KiB (293.869 KiB stored) + +processed 1 files, 292.969 KiB in 0:00 +snapshot 4137acee saved +[rc=0] + +app2 snapshot = 4137acee +## C1 control: the PLAIN key (no forced command) removes a snapshot +$ c.sh plain spike-repo forget 4137acee +rclone: 2026/10/03 09:10:36 CRITICAL: Failed to create file system for "lab:spike-repo": didn't find section in config file ("lab") +Fatal: unable to open repository at rclone:lab:spike-repo: error talking HTTP to rclone: Get "http://localhost/file-5577006791947779410": unexpected EOF +[rc=1] + +server files now: snapshots=3 keys=1 index=2 data=4 locks=0 +## E5: prune through the FORCED key must now try to delete a pack +$ c.sh forced spike-repo prune +loading indexes... +loading all snapshots... +finding data that is still in use for 3 snapshots +[0:00] 100.00% 3 / 3 snapshots + +searching used packs... +collecting packs for deletion and repacking +[0:00] 100.00% 4 / 4 packs processed + + +to repack: 0 blobs / 0 B +this removes: 0 blobs / 0 B +to delete: 0 blobs / 0 B +total prune: 0 blobs / 0 B +remaining: 12 blobs / 880.956 KiB +unused size after prune: 0 B (0.00% of remaining size) + +done +[rc=0] + +server files after forced prune: snapshots=3 keys=1 index=2 data=4 locks=0 +## C1 control: prune through the PLAIN key +$ c.sh plain spike-repo prune +rclone: 2026/10/03 09:10:39 CRITICAL: Failed to create file system for "lab:spike-repo": didn't find section in config file ("lab") +Fatal: unable to open repository at rclone:lab:spike-repo: error talking HTTP to rclone: Get "http://localhost/file-5577006791947779410": unexpected EOF +[rc=1] + +server files after plain prune: snapshots=3 keys=1 index=2 data=4 locks=0 +## C2: forced key, client names a DIFFERENT repo path (other-repo) +$ c.sh forced other-repo snapshots +ID Time Host Tags Paths +--------------------------------------------------------------- +63a075dc 2026-10-03 09:07:50 labbox app1 /d/tree +066a039d 2026-10-03 09:07:52 labbox app1 /d/tree +4137acee 2026-10-03 09:10:32 labbox app2 /d/tree2 +--------------------------------------------------------------- +3 snapshots +[rc=0] + +$ c.sh forced other-repo init +Fatal: create repository at rclone:lab:other-repo failed: Fatal: config file already exists + +[rc=1] + +spike-repo diff --git a/documentation/audits/offsite-append-only-2026-10-03/lab/A4b-controls-and-real-prune.txt b/documentation/audits/offsite-append-only-2026-10-03/lab/A4b-controls-and-real-prune.txt new file mode 100644 index 00000000..538ea81d --- /dev/null +++ b/documentation/audits/offsite-append-only-2026-10-03/lab/A4b-controls-and-real-prune.txt @@ -0,0 +1,105 @@ +## (re-run; the first attempt named an rclone remote 'lab:' the PLAIN server has no config for — the plain key runs rclone with the CLIENT's path, which is the point) +## C1 control: the PLAIN key (no forced command) removes snapshot 4137acee +$ c.sh plain spike-repo forget 4137acee +[0:00] 100.00% 1 / 1 files deleted + +[rc=0] + +server files now: snapshots=2 keys=1 index=2 data=4 locks=0 +## E5: prune through the FORCED key — unreferenced pack now exists +$ c.sh forced spike-repo prune +loading indexes... +loading all snapshots... +finding data that is still in use for 2 snapshots +[0:00] 100.00% 2 / 2 snapshots + +searching used packs... +collecting packs for deletion and repacking +[0:00] 100.00% 4 / 4 packs processed + + +to repack: 0 blobs / 0 B +this removes: 0 blobs / 0 B +to delete: 4 blobs / 293.709 KiB +total prune: 4 blobs / 293.709 KiB +remaining: 8 blobs / 587.247 KiB +unused size after prune: 0 B (0.00% of remaining size) + +rebuilding index +[0:00] 100.00% 2 / 2 packs processed + +deleting obsolete index files +Remove() returned error, retrying after 720.254544ms: blob not removed, server response: 403 Forbidden (403) +Remove() returned error, retrying after 582.280027ms: blob not removed, server response: 403 Forbidden (403) +Remove() returned error, retrying after 703.28564ms: blob not removed, server response: 403 Forbidden (403) +Remove() returned error, retrying after 693.478123ms: blob not removed, server response: 403 Forbidden (403) +Remove() returned error, retrying after 1.335175957s: blob not removed, server response: 403 Forbidden (403) +Remove() returned error, retrying after 636.341646ms: blob not removed, server response: 403 Forbidden (403) +Remove() returned error, retrying after 1.107876242s: blob not removed, server response: 403 Forbidden (403) +Remove() returned error, retrying after 1.007386063s: blob not removed, server response: 403 Forbidden (403) +Remove() returned error, retrying after 2.027308147s: blob not removed, server response: 403 Forbidden (403) +Remove() returned error, retrying after 2.569756966s: blob not removed, server response: 403 Forbidden (403) +Remove() returned error, retrying after 4.987726727s: blob not removed, server response: 403 Forbidden (403) +Remove() returned error, retrying after 2.711970641s: blob not removed, server response: 403 Forbidden (403) +Remove() returned error, retrying after 5.015617898s: blob not removed, server response: 403 Forbidden (403) +Remove() returned error, retrying after 4.659096946s: blob not removed, server response: 403 Forbidden (403) +Remove() returned error, retrying after 8.277195667s: blob not removed, server response: 403 Forbidden (403) +Remove() returned error, retrying after 6.689436284s: blob not removed, server response: 403 Forbidden (403) +Remove() returned error, retrying after 10.163166574s: blob not removed, server response: 403 Forbidden (403) +Remove() returned error, retrying after 15.10932531s: blob not removed, server response: 403 Forbidden (403) +Remove() returned error, retrying after 13.811796615s: blob not removed, server response: 403 Forbidden (403) +Remove() returned error, retrying after 13.516432903s: blob not removed, server response: 403 Forbidden (403) +unable to remove from the repository +unable to remove from the repository +[0:45] 0.00% 0 / 2 files deleted + +Fatal: blob not removed, server response: 403 Forbidden (403) +[rc=1] + +server files after forced prune: snapshots=2 keys=1 index=3 data=4 locks=0 +$ c.sh forced spike-repo check +using temporary cache in /tmp/restic-check-cache-3704335306 +create exclusive lock for repository +load indexes +pack 1d5237362eb08ed6ae49c8d57dc28123e665d6efadb34bb1771764f874ae2ba9 contained in several indexes: {2ac5484c c4e8a9de} +pack f3a63740213e99cde370d4f8f7f63d808c0e921237ed522678c730e3e81391e5 contained in several indexes: {2ac5484c c4e8a9de} +This is non-critical, you can run `restic rebuild-index' to correct this +check all packs +check snapshots, trees and blobs +[0:00] 100.00% 2 / 2 snapshots + +no errors were found +[rc=0] + +## C1 control: prune through the PLAIN key +$ c.sh plain spike-repo prune +loading indexes... +loading all snapshots... +finding data that is still in use for 2 snapshots +[0:00] 100.00% 2 / 2 snapshots + +searching used packs... +collecting packs for deletion and repacking +[0:00] 100.00% 4 / 4 packs processed + + +to repack: 0 blobs / 0 B +this removes: 0 blobs / 0 B +to delete: 4 blobs / 293.709 KiB +total prune: 4 blobs / 293.709 KiB +remaining: 8 blobs / 587.247 KiB +unused size after prune: 0 B (0.00% of remaining size) + +rebuilding index +[0:00] 100.00% 2 / 2 packs processed + +deleting obsolete index files +[0:00] 100.00% 3 / 3 files deleted + +removing 2 old packs +[0:00] 100.00% 2 / 2 files deleted + +done +[rc=0] + +server files after plain prune: snapshots=2 keys=1 index=1 data=2 locks=0 diff --git a/documentation/audits/offsite-append-only-2026-10-03/lab/A5-locks.txt b/documentation/audits/offsite-append-only-2026-10-03/lab/A5-locks.txt new file mode 100644 index 00000000..15257dd8 --- /dev/null +++ b/documentation/audits/offsite-append-only-2026-10-03/lab/A5-locks.txt @@ -0,0 +1,70 @@ +## A5: kill a backup mid-run through the FORCED key (docker kill -s KILL on the client = a crash) +killing r436-cli-2801178 at 09:12:15 +locks on server after the crash: +-rw-r--r-- 1 sub sub 141 09:12:12 b70bc18cbb4e907bc81308a463dd0a86839ef7200eb74c8ba8baf20d51f74fca +## next BACKUP (shared lock) from a NEW container (new hostname — as after a controller recreate) +$ c.sh forced spike-repo backup /d/tree --host labbox --tag app1 +using parent snapshot 066a039d + +Files: 0 new, 0 changed, 4 unmodified +Dirs: 0 new, 1 changed, 2 unmodified +Added to the repository: 325 B (274 B stored) + +processed 4 files, 585.948 KiB in 0:00 +snapshot 16f27a03 saved +[rc=0] + +## next EXCLUSIVE op (check) — does the stale lock wedge it? +$ c.sh forced spike-repo check +using temporary cache in /tmp/restic-check-cache-108793622 +create exclusive lock for repository +unable to create lock in backend: repository is already locked by PID 1 on 6999b25ca937 by root (UID 0, GID 0) +lock was created at 2026-10-03 09:12:12 (9.203992203s ago) +storage ID b70bc18c +the `unlock` command can be used to remove stale locks +[rc=1] + +## plain unlock (stale-only) +$ c.sh forced spike-repo unlock +successfully removed locks +[rc=0] + +locks: +-rw-r--r-- 1 sub sub 141 09:12:12 b70bc18cbb4e907bc81308a463dd0a86839ef7200eb74c8ba8baf20d51f74fca +## unlock --remove-all +$ c.sh forced spike-repo unlock --remove-all +successfully removed locks +[rc=0] + +locks: +## check again +$ c.sh forced spike-repo check +using temporary cache in /tmp/restic-check-cache-3364593304 +create exclusive lock for repository +load indexes +check all packs +pack 1a9dd015f3b76fd43ed4f4019bb15d024a36763d2e52f1a9216284a53f9be971: not referenced in any index +pack 8184618eaa974e3094796144f9b466468113f9e2bcd82194f9ca58f959d0f822: not referenced in any index +pack 63e40e54bedfc638a55e9fe5225231edebc84bb7263b4e4ca25b5933188ebb1c: not referenced in any index +pack 97b9da6b48256ec083071902fea61ef56ea5a18f3af869f7ec6eb2ead7772674: not referenced in any index +pack b213484fe31fd640b6b519a5ddbe486bfeed32688fea8a0185ff98dcf7165eb3: not referenced in any index +pack 5ccfd475d12254508662e5940c923ae132dc0a63b48ba65ff8e3970925f264fb: not referenced in any index +pack 0b27c44a9453ac56fb8c171447927d6a45b30ad9233aaa71e4dd9db9a3102b43: not referenced in any index +pack af691cecf756d65918ffd1430d8bebdfb720b657ba1ab24c37fd42e1506185ca: not referenced in any index +pack 14883d87a878ea91e66a471e892eb1a2f17cc916dcf3be3f6f32f9eb353b755c: not referenced in any index +pack 70c49c3190eba6522d6b493b2c35009bd8a562edd5fb706a0ef29a865c56d52a: not referenced in any index +pack fa2445b724ba511629fb1dadc75a86f6ae413a61147f7152e50d58effba4ea88: not referenced in any index +pack 3902d8453121f05f590775ff617782e0da7027553edcc81305a5629abf6bcde9: not referenced in any index +pack f86e1c71608c4ddcde8ad66ada62c42a0459b798f4ce4bc1f8a76a4728a85b1d: not referenced in any index +pack 4a080216c5d17e9f2b5eb91cc797901d7f858b86f97cbedb20192e6a90160599: not referenced in any index +pack 10a1dfa78d932c22e68f4bf59fde5e5fdb997d79711d67b8029285ab4c9ae934: not referenced in any index +pack 70fb6c036df351cc2f9f46ad69ecc8d234f635b4deee44fbc711e7f81174ea62: not referenced in any index +pack c7d8a4a140ac753cb9e9232a014e521b5d419c5c29cb29f4713e7b8b71cd0e47: not referenced in any index +17 additional files were found in the repo, which likely contain duplicate data. +This is non-critical, you can run `restic prune` to correct this. +check snapshots, trees and blobs +[0:00] 100.00% 3 / 3 snapshots + +no errors were found +[rc=0] + diff --git a/documentation/audits/offsite-append-only-2026-10-03/lab/B-rclone-server-surface.txt b/documentation/audits/offsite-append-only-2026-10-03/lab/B-rclone-server-surface.txt new file mode 100644 index 00000000..28526dda --- /dev/null +++ b/documentation/audits/offsite-append-only-2026-10-03/lab/B-rclone-server-surface.txt @@ -0,0 +1,32 @@ +rclone: rclone v1.75.1 (lab; the provider runs its own version) +T1 DELETE snapshots/ -> 403 Forbidden +T2 POST over existing snapshots/ -> 403 Forbidden +T3 POST over existing config -> 403 Forbidden +T4 DELETE data/ -> 403 Forbidden +T5 POST ../.ssh/authorized_keys -> 400 Bad Request +T6 POST %2e%2e/.ssh/authorized_keys -> 400 Bad Request +T7 POST data/..%2f..%2f.ssh/x -> 400 Bad Request +T8 POST a NEW keys/aaaa -> 200 +T9 POST arbitrary top-level foo -> 200 +T10 DELETE the new keys/aaaa -> 403 Forbidden +T11 POST a NEW locks/bbbb then DELETE -> 200 / 200 +after: snapshot present: yes; config unchanged: yes; authorized_keys unchanged: yes; stray files: +/home/sub: +spike-repo + +/home/sub/.ssh: +authorized_keys +/home/sub/spike-repo: +config +data +foo +index +keys +locks +snapshots + +/home/sub/spike-repo/keys: +aaaa +f651e7eec6b06d2594a748cb05cfaca39f7488092af5c4da5c332a7c3a7a8e5e + +/home/sub/spike-repo/locks: diff --git a/documentation/audits/offsite-append-only-2026-10-03/lab/C3-retention-poisoning.txt b/documentation/audits/offsite-append-only-2026-10-03/lab/C3-retention-poisoning.txt new file mode 100644 index 00000000..c9f138de --- /dev/null +++ b/documentation/audits/offsite-append-only-2026-10-03/lab/C3-retention-poisoning.txt @@ -0,0 +1,73 @@ +## C3 — retention poisoning. An attacker holding ONLY the forced key + the repo password ADDS snapshots (allowed) dated in the future, same host+tag. +real snapshots before: +ID Time Host Tags Paths +-------------------------------------------------------------- +63a075dc 2026-10-03 09:07:50 labbox app1 /d/tree +066a039d 2026-10-03 09:07:52 labbox app1 /d/tree +16f27a03 2026-10-03 09:12:18 labbox app1 /d/tree +-------------------------------------------------------------- +3 snapshots +after the attacker's adds (forced key, all succeeded): +ID Time Host Tags +--------------------------------------------- +63a075dc 2026-10-03 09:07:50 labbox app1 +066a039d 2026-10-03 09:07:52 labbox app1 +16f27a03 2026-10-03 09:12:18 labbox app1 +cf04458a 2027-01-01 03:00:00 labbox app1 +19394898 2027-01-02 03:00:00 labbox app1 +bfaf4619 2027-01-03 03:00:00 labbox app1 +3ba72ec1 2027-01-04 03:00:00 labbox app1 +8b3b6385 2027-01-05 03:00:00 labbox app1 +7034c9da 2027-01-06 03:00:00 labbox app1 +2c18c893 2027-01-07 03:00:00 labbox app1 +f0caba56 2027-02-15 03:00:00 labbox app1 +cb7c4ec4 2027-03-15 03:00:00 labbox app1 +064066cf 2027-04-15 03:00:00 labbox app1 +bc83265e 2027-05-15 03:00:00 labbox app1 +af39799e 2027-06-15 03:00:00 labbox app1 +e887befe 2027-07-15 03:00:00 labbox app1 +--------------------------------------------- +16 snapshots +## the HONEST retention (the box's exact policy) run later by whoever holds a deleting key — DRY RUN: +Applying Policy: keep 7 daily, 4 weekly, 6 monthly snapshots +keep 7 snapshots: +ID Time Host Tags Reasons Paths +--------------------------------------------------------------------------------- +2c18c893 2027-01-07 03:00:00 labbox app1 daily snapshot /d/empty +f0caba56 2027-02-15 03:00:00 labbox app1 daily snapshot /d/empty + monthly snapshot +cb7c4ec4 2027-03-15 03:00:00 labbox app1 daily snapshot /d/empty + monthly snapshot +064066cf 2027-04-15 03:00:00 labbox app1 daily snapshot /d/empty + weekly snapshot + monthly snapshot +bc83265e 2027-05-15 03:00:00 labbox app1 daily snapshot /d/empty + weekly snapshot + monthly snapshot +af39799e 2027-06-15 03:00:00 labbox app1 daily snapshot /d/empty + weekly snapshot + monthly snapshot +e887befe 2027-07-15 03:00:00 labbox app1 daily snapshot /d/empty + weekly snapshot + monthly snapshot +--------------------------------------------------------------------------------- +7 snapshots + +remove 9 snapshots: +ID Time Host Tags Paths +--------------------------------------------------------------- +63a075dc 2026-10-03 09:07:50 labbox app1 /d/tree +066a039d 2026-10-03 09:07:52 labbox app1 /d/tree +16f27a03 2026-10-03 09:12:18 labbox app1 /d/tree +cf04458a 2027-01-01 03:00:00 labbox app1 /d/empty +19394898 2027-01-02 03:00:00 labbox app1 /d/empty +bfaf4619 2027-01-03 03:00:00 labbox app1 /d/empty +3ba72ec1 2027-01-04 03:00:00 labbox app1 /d/empty +8b3b6385 2027-01-05 03:00:00 labbox app1 /d/empty +7034c9da 2027-01-06 03:00:00 labbox app1 /d/empty +--------------------------------------------------------------- +9 snapshots + +Would have removed the following snapshots: +{066a039d 16f27a03 19394898 3ba72ec1 63a075dc 7034c9da 8b3b6385 bfaf4619 cf04458a} + diff --git a/documentation/audits/offsite-append-only-2026-10-03/live/B1-B2-ports-password-forward.txt b/documentation/audits/offsite-append-only-2026-10-03/live/B1-B2-ports-password-forward.txt new file mode 100644 index 00000000..b8cac17f --- /dev/null +++ b/documentation/audits/offsite-append-only-2026-10-03/live/B1-B2-ports-password-forward.txt @@ -0,0 +1,27 @@ +## after B2: spike-r436 still there (the 'rm -rf' through the forced key ran rclone, not rm) +$ password p23: ls +felhom-repo +spike-r436 +[rc=0] + +## port 22 controls +$ plain p22: sftp ls .ssh +u629488-sub4@u629488-sub4.your-storagebox.de: Permission denied (publickey,password). +Connection closed +[rc=255] + +$ password p22: sftp ls .ssh (stdin batch) +Connected to u629488-sub4.your-storagebox.de. +sftp> ls -la .ssh +drwx------ 2 u629488-sub4 u629488 3 Sep 16 16:02 . +drwxr-xr-x 6 u629488-sub4 u629488 6 Oct 3 11:20 .. +-rw------- 1 u629488-sub4 u629488 510 Oct 3 11:19 authorized_keys +sftp> bye +[rc=0] + +## B2 port forward, positive test: open -L, then try to use it +$ connect through the forward +[rc=0] + +ssh -L log: +channel 1: open failed: administratively prohibited: open failed diff --git a/documentation/audits/offsite-append-only-2026-10-03/live/B2-forced-key-misuse.txt b/documentation/audits/offsite-append-only-2026-10-03/live/B2-forced-key-misuse.txt new file mode 100644 index 00000000..f19f86a4 --- /dev/null +++ b/documentation/audits/offsite-append-only-2026-10-03/live/B2-forced-key-misuse.txt @@ -0,0 +1,78 @@ +## B2 — the FORCED key asked for anything else (port 23 and 22) +$ forced p23: shell command 'ls -la .ssh' +2026/10/03 11:24:27 NOTICE: Config file "/home/.config/rclone/rclone.conf" not found - using defaults +[rc=0] + +$ forced p23: 'rm -rf spike-r436' +2026/10/03 11:24:28 NOTICE: Config file "/home/.config/rclone/rclone.conf" not found - using defaults +[rc=0] + +$ forced p23: sftp subsystem +Connection closed +[rc=255] + +$ forced p23: scp download +scp: Connection closed +[rc=255] + +$ forced p23: rsync --server +2026/10/03 11:24:50 NOTICE: Config file "/home/.config/rclone/rclone.conf" not found - using defaults +protocol version mismatch -- is your shell clean? +(see the rsync manpage for an explanation) +rsync error: protocol incompatibility (code 2) at compat.c(622) [Receiver=3.4.1] +[rc=2] + +$ forced p23: local port forward +[rc=124] + +$ forced p23: rclone serve restic WITHOUT the flag, other path +2026/10/03 11:25:16 NOTICE: Config file "/home/.config/rclone/rclone.conf" not found - using defaults +[rc=0] + +$ forced p22: shell command 'ls -la .ssh' +u629488-sub4@u629488-sub4.your-storagebox.de: Permission denied (publickey,password). +[rc=255] + +$ forced p22: 'rm -rf spike-r436' +u629488-sub4@u629488-sub4.your-storagebox.de: Permission denied (publickey,password). +[rc=255] + +$ forced p22: sftp subsystem +u629488-sub4@u629488-sub4.your-storagebox.de: Permission denied (publickey,password). +Connection closed +[rc=255] + +$ forced p22: scp download +u629488-sub4@u629488-sub4.your-storagebox.de: Permission denied (publickey,password). +scp: Connection closed +[rc=255] + +$ forced p22: rsync --server +u629488-sub4@u629488-sub4.your-storagebox.de: Permission denied (publickey,password). +rsync: connection unexpectedly closed (0 bytes received so far) [Receiver] +rsync error: unexplained error (code 255) at io.c(232) [Receiver=3.4.1] +[rc=255] + +$ forced p22: local port forward +u629488-sub4@u629488-sub4.your-storagebox.de: Permission denied (publickey,password). +[rc=255] + +$ forced p22: rclone serve restic WITHOUT the flag, other path +u629488-sub4@u629488-sub4.your-storagebox.de: Permission denied (publickey,password). +[rc=255] + +## B-control — the PLAIN key (what every box holds today) on port 23 +$ plain p23: ls -la .ssh +total 3 +drwx------ 2 u629488-sub4 1061 3 Sep 16 16:02 . +drwxr-xr-x 6 u629488-sub4 1061 6 Oct 3 11:20 .. +-rw------- 1 u629488-sub4 1061 510 Oct 3 11:19 authorized_keys +[rc=0] + +$ plain p23: sftp ls .ssh +sftp> ls -la .ssh +drwx------ ? u629488-sub4 1061 3 Sep 16 18:02 .ssh/. +drwxr-xr-x ? u629488-sub4 1061 6 Oct 3 13:20 .ssh/.. +-rw------- ? u629488-sub4 1061 510 Oct 3 13:19 .ssh/authorized_keys +[rc=0] + diff --git a/documentation/audits/offsite-append-only-2026-10-03/live/C2-A5-locks.txt b/documentation/audits/offsite-append-only-2026-10-03/live/C2-A5-locks.txt new file mode 100644 index 00000000..2796c09a --- /dev/null +++ b/documentation/audits/offsite-append-only-2026-10-03/live/C2-A5-locks.txt @@ -0,0 +1,68 @@ +## C2: forced key, client names another path +$ lc.sh forced other-path snapshots +ID Time Host Tags Paths +-------------------------------------------------------------- +d807418c 2026-10-03 11:20:19 livebox app1 /d/tree +68885c41 2026-10-03 11:20:22 livebox app1 /d/tree +-------------------------------------------------------------- +2 snapshots +[rc=0] + +## A5: crash a backup mid-run (docker kill -s KILL) +killing r436-live-3883038 at 11:31:14 +locks after crash: +-rw-rw-r-- 1 u629488-sub4 1061 141 Oct 3 11:31 96a4f54c6477793c31050da87ad5e835bbbed2ce76549634a1ccf6cf5a065524 +$ lc.sh forced spike-r436 backup /d/tree --host livebox --tag app1 +using parent snapshot 68885c41 + +Files: 0 new, 0 changed, 4 unmodified +Dirs: 0 new, 1 changed, 2 unmodified +Added to the repository: 325 B (274 B stored) + +processed 4 files, 585.948 KiB in 0:00 +snapshot 8f720f1d saved +[rc=0] + +$ lc.sh forced spike-r436 check +using temporary cache in /tmp/restic-check-cache-2564422886 +create exclusive lock for repository +unable to create lock in backend: repository is already locked by PID 1 on 91671d11c1a8 by root (UID 0, GID 0) +lock was created at 2026-10-03 11:31:08 (14.102493775s ago) +storage ID 96a4f54c +the `unlock` command can be used to remove stale locks +[rc=1] + +## plain unlock +$ lc.sh forced spike-r436 unlock +successfully removed locks +[rc=0] + +locks: +-rw-rw-r-- 1 u629488-sub4 1061 141 Oct 3 11:31 96a4f54c6477793c31050da87ad5e835bbbed2ce76549634a1ccf6cf5a065524 +## unlock --remove-all +$ lc.sh forced spike-r436 unlock --remove-all +successfully removed locks +[rc=0] + +locks: +$ lc.sh forced spike-r436 check +using temporary cache in /tmp/restic-check-cache-3405178813 +create exclusive lock for repository +load indexes +pack c67a24b49d4e41a91d6de9448cb7c72623b6eb12ee8a50b7f4db2e2dfb465480 contained in several indexes: {85027deb a5b0181c} +pack 92475c24d7b0495584f168c7a73e6f0e35c9c27881c2ecbb0e18695b889cfb7f contained in several indexes: {85027deb a5b0181c} +This is non-critical, you can run `restic rebuild-index' to correct this +check all packs +pack 9783961d590d0675569230cf7662548d03914af0eebe1d4b1eb6b5573dafb9a8: not referenced in any index +pack 5aa289972c6799f5a18c991f3f2c36a2ac8fcdeb41ed5b2708807a774bc730ae: not referenced in any index +pack f16746415ed296cc92e78df06859b3fd2c0ba3deb73a31c97e8ed301ee7c214a: not referenced in any index +pack 51690a831d4a96fbeed192e0ff0ce879d309994ccb044ead6a1f2096baeea610: not referenced in any index +pack 38026bf50bc4028ef383ea19fbdeff362a925d6ecd4be66a3c635b8796fe623e: not referenced in any index +5 additional files were found in the repo, which likely contain duplicate data. +This is non-critical, you can run `restic prune` to correct this. +check snapshots, trees and blobs +[0:00] 100.00% 3 / 3 snapshots + +no errors were found +[rc=0] + diff --git a/documentation/audits/offsite-append-only-2026-10-03/live/E1-E3-init-backup.txt b/documentation/audits/offsite-append-only-2026-10-03/live/E1-E3-init-backup.txt new file mode 100644 index 00000000..11825e61 --- /dev/null +++ b/documentation/audits/offsite-append-only-2026-10-03/live/E1-E3-init-backup.txt @@ -0,0 +1,44 @@ +Sat Oct 3 11:20:12 UTC 2026 +$ lc.sh forced spike-r436 init +rclone: 2026/10/03 11:20:16 NOTICE: Config file "/home/.config/rclone/rclone.conf" not found - using defaults +created restic repository 3b129d9736 at rclone:spike-r436 + +Please note that knowledge of your password is required to access +the repository. Losing your password means that your data is +irrecoverably lost. +[rc=0] + +$ lc.sh forced spike-r436 backup /d/tree --host livebox --tag app1 +rclone: 2026/10/03 11:20:20 NOTICE: Config file "/home/.config/rclone/rclone.conf" not found - using defaults +no parent snapshot found, will read all files + +Files: 4 new, 0 changed, 0 unmodified +Dirs: 3 new, 0 changed, 0 unmodified +Added to the repository: 588.177 KiB (587.535 KiB stored) + +processed 4 files, 585.948 KiB in 0:00 +snapshot d807418c saved +[rc=0] + +$ lc.sh forced spike-r436 backup /d/tree --host livebox --tag app1 +rclone: 2026/10/03 11:20:23 NOTICE: Config file "/home/.config/rclone/rclone.conf" not found - using defaults +using parent snapshot d807418c + +Files: 0 new, 0 changed, 4 unmodified +Dirs: 0 new, 0 changed, 3 unmodified +Added to the repository: 0 B (0 B stored) + +processed 4 files, 585.948 KiB in 0:00 +snapshot 68885c41 saved +[rc=0] + +$ lc.sh forced spike-r436 snapshots +rclone: 2026/10/03 11:20:26 NOTICE: Config file "/home/.config/rclone/rclone.conf" not found - using defaults +ID Time Host Tags Paths +-------------------------------------------------------------- +d807418c 2026-10-03 11:20:19 livebox app1 /d/tree +68885c41 2026-10-03 11:20:22 livebox app1 /d/tree +-------------------------------------------------------------- +2 snapshots +[rc=0] + diff --git a/documentation/audits/offsite-append-only-2026-10-03/live/E4-E6-deletes-and-C1.txt b/documentation/audits/offsite-append-only-2026-10-03/live/E4-E6-deletes-and-C1.txt new file mode 100644 index 00000000..7cfc0d7e --- /dev/null +++ b/documentation/audits/offsite-append-only-2026-10-03/live/E4-E6-deletes-and-C1.txt @@ -0,0 +1,112 @@ +server files (via password shell, find): snapshots=0 index=0 data=0 keys=0 locks=0 +$ lc.sh forced spike-r436 restore d807418c --target /d/restored --include /d/tree/sub/note.txt +restoring to /d/restored +[rc=0] + +restored vs source: IDENTICAL +## E4 +$ lc.sh forced spike-r436 forget d807418c --prune +[0:48] 0.00% 0 / 1 files deleted + +unable to remove from the repository +blob not removed, server response: 403 Forbidden (403) +[rc=1] + +## E6 +$ lc.sh forced spike-r436 forget --keep-last 1 +Applying Policy: keep 1 latest snapshots +keep 1 snapshots: +ID Time Host Tags Reasons Paths +----------------------------------------------------------------------------- +68885c41 2026-10-03 11:20:22 livebox app1 last snapshot /d/tree +----------------------------------------------------------------------------- +1 snapshots + +remove 1 snapshots: +ID Time Host Tags Paths +-------------------------------------------------------------- +d807418c 2026-10-03 11:20:19 livebox app1 /d/tree +-------------------------------------------------------------- +1 snapshots + +unable to remove from the repository +[0:48] 0.00% 0 / 1 files deleted + +blob not removed, server response: 403 Forbidden (403) +[rc=1] + +server files: snapshots=0 index=0 data=0 keys=0 locks=0 +## E5 made real: app2 snapshot, removed by the PLAIN key (C1), then prune through the FORCED key +$ lc.sh forced spike-r436 backup /d/tree2 --host livebox --tag app2 +no parent snapshot found, will read all files + +Files: 1 new, 0 changed, 0 unmodified +Dirs: 2 new, 0 changed, 0 unmodified +Added to the repository: 293.934 KiB (293.869 KiB stored) + +processed 1 files, 292.969 KiB in 0:00 +snapshot 973e0dac saved +[rc=0] + +app2 = 973e0dac +## C1 control: PLAIN key forget +$ lc.sh plain spike-r436 forget 973e0dac +[0:00] 100.00% 1 / 1 files deleted + +[rc=0] + +server files: snapshots=0 index=0 data=0 keys=0 locks=0 +## E5 +$ lc.sh forced spike-r436 prune +loading indexes... +loading all snapshots... +finding data that is still in use for 2 snapshots +[0:00] 100.00% 2 / 2 snapshots + +searching used packs... +collecting packs for deletion and repacking +[0:00] 100.00% 4 / 4 packs processed + + +to repack: 0 blobs / 0 B +this removes: 0 blobs / 0 B +to delete: 4 blobs / 293.709 KiB +total prune: 4 blobs / 293.709 KiB +remaining: 8 blobs / 587.215 KiB +unused size after prune: 0 B (0.00% of remaining size) + +rebuilding index +[0:00] 100.00% 2 / 2 packs processed + +deleting obsolete index files +unable to remove from the repository +unable to remove from the repository +[0:45] 0.00% 0 / 2 files deleted + +Fatal: blob not removed, server response: 403 Forbidden (403) +[rc=1] + +server files: snapshots=0 index=0 data=0 keys=0 locks=0 +$ lc.sh forced spike-r436 snapshots +ID Time Host Tags Paths +-------------------------------------------------------------- +d807418c 2026-10-03 11:20:19 livebox app1 /d/tree +68885c41 2026-10-03 11:20:22 livebox app1 /d/tree +-------------------------------------------------------------- +2 snapshots +[rc=0] + +$ lc.sh forced spike-r436 check +using temporary cache in /tmp/restic-check-cache-3559416072 +create exclusive lock for repository +load indexes +pack c67a24b49d4e41a91d6de9448cb7c72623b6eb12ee8a50b7f4db2e2dfb465480 contained in several indexes: {85027deb a5b0181c} +pack 92475c24d7b0495584f168c7a73e6f0e35c9c27881c2ecbb0e18695b889cfb7f contained in several indexes: {85027deb a5b0181c} +This is non-critical, you can run `restic rebuild-index' to correct this +check all packs +check snapshots, trees and blobs +[0:00] 100.00% 2 / 2 snapshots + +no errors were found +[rc=0] + diff --git a/documentation/audits/offsite-append-only-2026-10-03/live/TEARDOWN.txt b/documentation/audits/offsite-append-only-2026-10-03/live/TEARDOWN.txt new file mode 100644 index 00000000..66441345 --- /dev/null +++ b/documentation/audits/offsite-append-only-2026-10-03/live/TEARDOWN.txt @@ -0,0 +1,16 @@ +## teardown 2026-10-03T11:31:45Z +Connected to u629488-sub4.your-storagebox.de. +sftp> put ak.orig .ssh/authorized_keys +Uploading ak.orig to /home/.ssh/authorized_keys +authorized_keys sha256 now : 795e715315973740bed25d99df39aebb159b79867423f72be9c92f7ed1072c94 +authorized_keys sha256 orig: 795e715315973740bed25d99df39aebb159b79867423f72be9c92f7ed1072c94 +rm spike-r436: rc=0 +home now: . .. .config .ssh felhom-repo +forced key after teardown: u629488-sub4@u629488-sub4.your-storagebox.de: Permission denied (publickey,password). +plain key after teardown: u629488-sub4@u629488-sub4.your-storagebox.de: Permission denied (publickey,password). +## .config was NOT present before the test (initial listing: .ssh, felhom-repo) — created by the provider's rclone: +.config +└── rclone + +2 directories, 0 files +home after: . .. .ssh felhom-repo diff --git a/documentation/backlog/CLOSED-ITEMS.md b/documentation/backlog/CLOSED-ITEMS.md index c8531400..b53d9fd2 100644 --- a/documentation/backlog/CLOSED-ITEMS.md +++ b/documentation/backlog/CLOSED-ITEMS.md @@ -26,6 +26,17 @@ --- +## 2026-10-03 — off-site append-only, measured on the provider (R-436, R-430) + +> Spike, no product change. Evidence and design: `audits/offsite-append-only-2026-10-03/`. + +| Row | What | Closed | Evidence | +|---|---|---|---| +| **R-436** | **Hetzner's `--append-only` forced command holds for the key it is pinned to.** On `u629488-sub4` (tester-1's, operator-ruled venue; scratch repo `spike-r436`, removed): pinned key `command="rclone serve restic --stdio --append-only spike-r436",restrict` — `init`, two `backup`s, `snapshots`, `restore` (bytes identical), `check` OK; `forget d807418c --prune`, `forget --keep-last 1`, `prune` → `blob not removed, server response: 403 Forbidden (403)`, rc=1, count unchanged; control with an unpinned key: `1 / 1 files deleted`. The client's path and flags are ignored; no shell, sftp, scp, rsync or port forward (`administratively prohibited`). **Reasoning kept: the pin protects a repository only if no other route can rewrite `authorized_keys` — and the password can (R-820).** | CLOSED 2026-10-03 — MEASURED; the due-check (2026-10-06) is cleared by this measurement | `live/E1-E3-init-backup.txt`, `live/E4-E6-deletes-and-C1.txt`, `live/B2-forced-key-misuse.txt`, `live/TEARDOWN.txt` (authorized_keys restored, sha256 identical) | +| **R-430** | **`restic unlock` prints `successfully removed locks` after removing nothing — by design; and `unlock --remove-all` DOES work through the append-only key.** Measured live and in the lab: a crash lock (not yet stale: under 30 min, new hostname) survives plain `unlock`, which still prints success; `--remove-all` removes it because the rclone append-only server allows lock deletion. A crash lock blocks `check`, not `backup`. So `resticStep`'s self-heal stays valid under R-436's transport; the earlier sticky-directory model does not describe it. | CLOSED 2026-10-03 — ANSWERED; not a precondition for the rclone transport | `live/C2-A5-locks.txt`, `lab/A5-locks.txt` | + +--- + ## 2026-10-03 — the triage: finished rows moved out of the open register > Every row below sat in `OPEN-ITEMS.md` with a finished LEADING verdict (or was verified finished against diff --git a/documentation/backlog/OPEN-ITEMS.md b/documentation/backlog/OPEN-ITEMS.md index 8ea25443..e471b57e 100644 --- a/documentation/backlog/OPEN-ITEMS.md +++ b/documentation/backlog/OPEN-ITEMS.md @@ -191,22 +191,22 @@ stopping line that lies. | **R-687** | App updates | P4 | **[P3-LOW] Part 7's live proof has four gaps a scratch box cannot close, and one observability gap.** (1) **W+5h reached with steps left** is proven by unit test only (`TestLeg_NoStepAtOrAfterW5h`) — the leg starts at W+105m and would need a 3-hour leg live; (2) **the off-site leg FAILING** before the update leg: 9202 has no off-site tier, so only the "no target" path ran live — failure and panic are `TestChainUpdateLeg_EveryPath`; (3) **a `files_may_change` step WITHOUT a whole copy**: both drill apps given the mark (wishlist, navidrome, romm) turned out whole on 9202 by the backup side's truth table (why, per app, is not logged — see the observability gap), so only "with a copy" ran live; (4) **the full-system gate waiting** cannot run on 9202 (no agent), and **did not occur on the demo boxes' real night either** (2026-09-25: both legs ended by 04:19, before the gate opened at 04:30, and no whole-box backup was due on either box) — unit + red-proof only (`TestD20_GateWaitsForTheLeg`). **Also cosmetic:** a leg with no steps reports `"steps": null` to the hub, not `[]`. **Observability:** when the leg TAKES a `files_may_change` step it does not log which whole copy allowed it (only the skip says why). `audits/night-2026-09-25/C/` **-- NARROWED 2026-09-25 (controller v0.273.0):** the cosmetic `"steps": null` → `[]` and the taken `files_may_change` step's missing log line are FIXED (red-proofed, `audits/night-2026-09-26/F/`). Items (1)–(4) stay; (4) did not occur on 2026-09-25 either (demo-felhom's whole-box backup ran at 07:29, three hours after its leg; demo-hp had none due). **-- 2026-09-28 (night 27/28):** (4) did not occur again — on demo-hp the leg ended 04:23:54 and the whole-guest backup began 04:37:06, after the gate opened at 04:30; demo-felhom's backup ran at 07:36 (`audits/evidence-golden-0276-2026-09-28/phaseD2-night-read.txt`). **-- 2026-09-30 (by day, demo-hp 9201): item (4) PROVEN LIVE.** The night chain pressed by hand, the window moved to W = now − 2h05m the moment the leg started, `quiesce.poll_interval` 1m: `[quiesce] full-system backup due and inside its window, but the automatic update leg is running … deferring` at 11:35:11 and 11:36:11 UTC while bookstack (55.1 s) and kimai (75.1 s) stepped; the leg's end line at 11:36:29; the backup quiesced at 11:37:11 (the first poll after), job done 11:47:19, the agent's `backup: completed` 9.98 GB. Config and window put back and read back (`audits/pg-last-six-2026-09-30/C/`). **Found, cosmetic, manual chain only:** the deferral names the moved window's W+5h (16:29) while the manual leg's own deadline was its start + the leg length (16:49). | **OPEN — P3, gaps (1)–(3) + the manual-chain deferral text; item (4) proven live 2026-09-30; owner: CC** **Re-ranked 2026-10-03: P3→P4: the gaps are covered by unit tests; left is live-proof completeness and one log text.** | — | — | CC | | **R-734** | App updates | P4 | **[P3-LOW] The harness marks immich `files_may_change` because immich rewrites six 13-byte `.immich` folder markers at every start.** MEASURED 2026-09-30 on the bench (v3.2.2 → v3.2.4): the bind-tree hash of `appdata/immich` changed; the only changed files were `{encoded-video,library,backups,profile,thumbs,upload}/.immich`, rewritten at each start — no household file. The mark is honest by the harness's rule and the ladder writer copies it (never edited by hand), so immich's v3.2.4 night step needs a fresh WHOLE copy (decision 13); on a box without one the night leg skips it and a person presses. The 2026-09-23 immich entry did not carry it (`files_changed []`). **Needs:** a decision whether app-owned marker files are excluded from the file hash (a per-template ignore list, or a size/name rule), or the mark stays. **-- 2026-09-30 (evening):** the harness now NAMES the files behind the mark (`files_changed_detail`, catalog `5b1972b`); on immich's step `0b82…` re-proof it named exactly the six `.immich` markers again. calibre-web's step (v4.0.6 → v4.0.8) carries the mark too, and the named files are its LIBRARY DATABASE: `media/books/metadata.db`, `metadata.db-shm`, `metadata.db-wal` changed; the book file did not (bench names re-run, `A/calibre-names/`). That is household data (the template's backup class `mandatory` holds DB and books as one unit), so the mark is right there and the night leg takes that step only with a fresh whole copy. On 9202 the same step changed no file in that folder (per-file hashes before/after) — not explained. `audits/more-night-apps-2026-09-30/` | **READY — rank P3-LOW; owner: CC (harness); the rule change needs a word** **Re-ranked 2026-10-03: P3→P4: the effect is an update that waits for a press; no data risk.** | — | — | CC + operator | -## Backup & restore — 56 rows (P2 12, P3 24, P4 20) +## Backup & restore — 55 rows (P2 12, P3 23, P4 20) | ID | Category | Sev | What | State | Blocked on | Next action | Owner | |---|---|---|---|---|---|---|---| | **R-32** | Backup & restore | P2 | **[P2-HIGH] RESET must purge the customer base dir; the orphan card must stay honest; unattributed bytes must be visible.** The rehearsal's S7 said in advance that an orphan card would BE a finding — and one appeared (16:58:14). Cause: RESET's `"hetzner":"ok"` leg destroys the sub-account, but **a Hetzner sub-account is an access-control object, not a data object** — its directory survives, so re-enabling offsite recreated an account over the previous lifecycle's ciphertext, encrypted under a key that same RESET had destroyed. **MIGRATED FROM `ROADMAP.md` 2026-08-22 (R-369) — originally filed 2026-07-21, size M, roadmap state `idea`.** Moved verbatim; nothing added or reinterpreted. The roadmap keeps its copy as history, marked moved. | **OPEN — migrated from ROADMAP 2026-08-22, rank unchanged** | — | **Ruling from the run (three parts, deliberately separate):** (1) because RESET destroys custody, the ciphertext it leaves behind is unrecoverable **BY DESIGN** → RESET gains a **main-account purge of the customer base dir** (the existing operator ack already covers it); (2) the **move-aside guard STAYS** for reinstall-*without*-RESET — there custody survives and the card's "history recoverable" promise is true (R-26 depends on exactly that); (3) the operator **Restic tab shows per-customer directory bytes vs attributed snapshot bytes**, so dead data cannot hide. Measured on the pool box that night: **49 M attributed** (2 snapshots, 48.717 MiB) against **1.4 G + 3.0 M unattributed** across TWO `.orphaned-*` dirs. Evidence `restic-and-pool.txt` | CC | -| **R-95** | Backup & restore | P2 | restic offsite credential **can delete** (`readonly=False`, `forget --prune` runs from the box); SFTP cannot express append-only | **READY** | — | Root exposure still open. Mitigation now ARMED — split prune off-box or move to REST `--append-only` **-- UNBLOCKED 2026-09-22, not closed.** The premise in this row's own title - *SFTP cannot express append-only* - is answered by Hetzner (ticket #2026090103040671, recorded in R-433 and R-436): on a Storage Box the `rclone serve restic --stdio --append-only` backend can be pinned to an SSH key as a FORCED COMMAND, so append-only IS expressible without a new machine. **The row stays open because nothing has been measured**: no forced-command key exists on `u629488`, no `forget --prune` has been refused through one, and the protection only holds if EVERY key that can reach the repository carries the prefix. The next step is R-436's, and it is small. | CC **SPIKE 2026-09-01 — `audits/SPIKE-r95-offsite-delete-2026-09-01.md`. THE WORD "ARMED" ABOVE IS NOT SUPPORTED AND IS WITHDRAWN PENDING R-429:** no `.snapshots` is visible to either box's sub-account (measured, both machines, with controls), so the seven-day bound is unverified and unverifiable from the product side. **Q3 (documented, `hub/internal/hetznerapi/hetznerapi.go:38-45`): the sub-account API has ONE permission axis, `readonly` — there is no append-only, so the PBS shape does NOT transfer** (PBS is a server that can refuse; a Storage Box is a filesystem that runs nothing). **Q5 (measured): withdrawing delete does NOT wedge the store** — restic treats a dead owner's lock as stale and proceeds — so the constraint everyone feared is not the blocker; but `unlock --remove-all` lies about success (R-430), and the crash-lock window is UNKNOWN. **Q6 (measured): restic 0.14.0 DOES speak `rest:`** (control: `banana:` → invalid backend), and append-only is a rest-server flag, not a restic one — reachable, but it needs a machine in the recovery path and ep0 is protected. **Q7 (measured): detection is nearly free** — `snapshot_count` already reaches the hub and the hub APPENDS reports, so the comparison needs no box change. **RECOMMENDATION: answer R-429 first (Viktor, ten minutes), then build detection, then move retention off the box; defer the transport change.** Two `forget --prune` sites must be disarmed together — `offbox.go:1388` AND `offbox.go:1759` — or R-191 repeats. **RE-SCOPED 2026-09-01 — THE STORY WAS WORSE THAN THE TRUTH FOR TWO MONTHS.** The box can delete its own LIVE repository, **but it cannot write to the daily snapshots of it** — MEASURED, not cited: a write into `/.zfs/snapshot` is refused on both boxes while the same write to the account home succeeds (R-432). Seven daily snapshots are confirmed in the panel (R-429). **So a deletion costs at most the data written since the last daily snapshot, and the rest is recoverable — file by file, one customer at a time, with no effect on anyone else** (vendor: *"You can download individual files or entire directories as usual"*; *"It is not possible to write to the `/.zfs` directory or its subfolder"*). **NOT open-ended loss.** Two caveats kept honest: a panel-driven snapshot restore rolls back the WHOLE Storage Box and deletes newer snapshots, which is why the per-file route matters; and per-file recovery is **operator-only today** (R-432). **DETECTION SHIPPED hub v0.111.0 (R-431)** — an unexplained fall is noticed within a day. **THE RANKING IS VIKTOR'S:** this has been #1 since July on the old story. On the new facts I would rank it below the items that can still lose data outright, but I am not re-ranking it myself. `audits/SPIKE-r95-offsite-delete-2026-09-01.md` **DRILL 2026-09-01, LATER THE SAME DAY — THE RE-SCOPE'S SECOND HALF IS WITHDRAWN.** The drill that was to walk the recovery found there is no route to walk: **no snapshot is reachable from a sub-account by ANY name** (R-433) — 777,600 exact names in the vendor format over nine days, zero hits, with a passing control, plus the structural reason (`/home` st_dev 0,82 vs `/.zfs/snapshot` st_dev 0,276, and `/home/.zfs` absent). **Clause (a) stands: the box can delete the live repo and cannot write into the snapshot area.** Clause (b) — *"the rest is recoverable file by file"* — is NOT SUPPORTED. **The drill was STOPPED before its destructive phase on the operator's ruling**, because with no recovery route the deletion would have destroyed real history to buy only an alarm test that could not fire at the specified size (R-435). Nothing was deleted; the store is verified untouched at 69 snapshots. **So the comfort that lowered this row rested on an unwalked route, and the route does not exist.** Two new leads decide what happens next: **R-433** (can the MAIN account see them? nobody here holds that credential) and **R-436** (`rclone serve restic --stdio` is offered server-side and restic speaks `rclone:` — measured — which could make real prevention cheap, IF the provider pins `--append-only`). **The rank stays Viktor's. Plainly: the argument that moved this row down is the argument the drill removed.** `audits/evidence-drill-r95-recovery-2026-09-01/` **BLOCKED-ON-PROVIDER 2026-09-01, and the record of the demotion is kept deliberately: this row spent ONE DAY demoted on a clause that did not hold.** It was re-scoped down on the morning of 2026-09-01 on the strength of *"recoverable file by file"*, and that clause was measured false the same afternoon (R-433). **On today's evidence it belongs back near the top — that is a proposal, not an action; CC has not re-ranked it and will not.** Both questions that can settle it are drafted in `documentation/runbooks/provider-questions-2026-09-01.md`: **Q1** decides how urgent this is, **Q2** (R-436) could remove the root cause cheaply. **Precondition on any build here: R-430**, which is harmless only while the box can still delete. | +| **R-95** | Backup & restore | P2 | restic offsite credential **can delete** (`readonly=False`, `forget --prune` runs from the box); SFTP cannot express append-only | **READY** **MEASURED 2026-10-03 (R-436 CLOSED) — append-only HOLDS for a pinned key; it is NOT yet protection.** On the provider (`u629488-sub4`, scratch repo, removed): a key pinned to `command="rclone serve restic --stdio --append-only ",restrict` backs up, restores (bytes identical) and checks; `forget --prune`, `forget --keep-last 1` and a real `prune` are refused with `blob not removed, server response: 403 Forbidden (403)`; the unpinned control deletes. **What still defeats it:** the sub-account PASSWORD logs in on ports 22 and 23 and rewrote `authorized_keys` in this session, and a box can obtain that password from the hub at will (R-820). **Also found:** an add-only key can plant future-dated snapshots that make the box's own retention policy select every real snapshot (R-822); the hub holds every sub-account password in the clear (R-821); FOUR box features need a deleting login, not two — both `forget` sites, the orphan move-aside (`mv`) and the abandonment (`rm -rf`). **PROPOSAL (operator decides — STATUS):** the hub becomes the key registrar so the box never sees the password; switch boxes to the pinned key with no box-side retention first (quota headroom is months); then a weekly hub-opened clean-up window with a poisoning guard. `audits/offsite-append-only-2026-10-03/DESIGN.md`. The rank stays the operator's. | — | Root exposure still open. Mitigation now ARMED — split prune off-box or move to REST `--append-only` **-- UNBLOCKED 2026-09-22, not closed.** The premise in this row's own title - *SFTP cannot express append-only* - is answered by Hetzner (ticket #2026090103040671, recorded in R-433 and R-436): on a Storage Box the `rclone serve restic --stdio --append-only` backend can be pinned to an SSH key as a FORCED COMMAND, so append-only IS expressible without a new machine. **The row stays open because nothing has been measured**: no forced-command key exists on `u629488`, no `forget --prune` has been refused through one, and the protection only holds if EVERY key that can reach the repository carries the prefix. The next step is R-436's, and it is small. | CC **SPIKE 2026-09-01 — `audits/SPIKE-r95-offsite-delete-2026-09-01.md`. THE WORD "ARMED" ABOVE IS NOT SUPPORTED AND IS WITHDRAWN PENDING R-429:** no `.snapshots` is visible to either box's sub-account (measured, both machines, with controls), so the seven-day bound is unverified and unverifiable from the product side. **Q3 (documented, `hub/internal/hetznerapi/hetznerapi.go:38-45`): the sub-account API has ONE permission axis, `readonly` — there is no append-only, so the PBS shape does NOT transfer** (PBS is a server that can refuse; a Storage Box is a filesystem that runs nothing). **Q5 (measured): withdrawing delete does NOT wedge the store** — restic treats a dead owner's lock as stale and proceeds — so the constraint everyone feared is not the blocker; but `unlock --remove-all` lies about success (R-430), and the crash-lock window is UNKNOWN. **Q6 (measured): restic 0.14.0 DOES speak `rest:`** (control: `banana:` → invalid backend), and append-only is a rest-server flag, not a restic one — reachable, but it needs a machine in the recovery path and ep0 is protected. **Q7 (measured): detection is nearly free** — `snapshot_count` already reaches the hub and the hub APPENDS reports, so the comparison needs no box change. **RECOMMENDATION: answer R-429 first (Viktor, ten minutes), then build detection, then move retention off the box; defer the transport change.** Two `forget --prune` sites must be disarmed together — `offbox.go:1388` AND `offbox.go:1759` — or R-191 repeats. **RE-SCOPED 2026-09-01 — THE STORY WAS WORSE THAN THE TRUTH FOR TWO MONTHS.** The box can delete its own LIVE repository, **but it cannot write to the daily snapshots of it** — MEASURED, not cited: a write into `/.zfs/snapshot` is refused on both boxes while the same write to the account home succeeds (R-432). Seven daily snapshots are confirmed in the panel (R-429). **So a deletion costs at most the data written since the last daily snapshot, and the rest is recoverable — file by file, one customer at a time, with no effect on anyone else** (vendor: *"You can download individual files or entire directories as usual"*; *"It is not possible to write to the `/.zfs` directory or its subfolder"*). **NOT open-ended loss.** Two caveats kept honest: a panel-driven snapshot restore rolls back the WHOLE Storage Box and deletes newer snapshots, which is why the per-file route matters; and per-file recovery is **operator-only today** (R-432). **DETECTION SHIPPED hub v0.111.0 (R-431)** — an unexplained fall is noticed within a day. **THE RANKING IS VIKTOR'S:** this has been #1 since July on the old story. On the new facts I would rank it below the items that can still lose data outright, but I am not re-ranking it myself. `audits/SPIKE-r95-offsite-delete-2026-09-01.md` **DRILL 2026-09-01, LATER THE SAME DAY — THE RE-SCOPE'S SECOND HALF IS WITHDRAWN.** The drill that was to walk the recovery found there is no route to walk: **no snapshot is reachable from a sub-account by ANY name** (R-433) — 777,600 exact names in the vendor format over nine days, zero hits, with a passing control, plus the structural reason (`/home` st_dev 0,82 vs `/.zfs/snapshot` st_dev 0,276, and `/home/.zfs` absent). **Clause (a) stands: the box can delete the live repo and cannot write into the snapshot area.** Clause (b) — *"the rest is recoverable file by file"* — is NOT SUPPORTED. **The drill was STOPPED before its destructive phase on the operator's ruling**, because with no recovery route the deletion would have destroyed real history to buy only an alarm test that could not fire at the specified size (R-435). Nothing was deleted; the store is verified untouched at 69 snapshots. **So the comfort that lowered this row rested on an unwalked route, and the route does not exist.** Two new leads decide what happens next: **R-433** (can the MAIN account see them? nobody here holds that credential) and **R-436** (`rclone serve restic --stdio` is offered server-side and restic speaks `rclone:` — measured — which could make real prevention cheap, IF the provider pins `--append-only`). **The rank stays Viktor's. Plainly: the argument that moved this row down is the argument the drill removed.** `audits/evidence-drill-r95-recovery-2026-09-01/` **BLOCKED-ON-PROVIDER 2026-09-01, and the record of the demotion is kept deliberately: this row spent ONE DAY demoted on a clause that did not hold.** It was re-scoped down on the morning of 2026-09-01 on the strength of *"recoverable file by file"*, and that clause was measured false the same afternoon (R-433). **On today's evidence it belongs back near the top — that is a proposal, not an action; CC has not re-ranked it and will not.** Both questions that can settle it are drafted in `documentation/runbooks/provider-questions-2026-09-01.md`: **Q1** decides how urgent this is, **Q2** (R-436) could remove the root cause cheaply. **Precondition on any build here: R-430**, which is harmless only while the box can still delete. | | **R-105** | Backup & restore | P2 | **Three hub-held DR records are empty on the entire live fleet.** `hosts.dr_record_json` = `{}` on all 3 hosts; `host_escrow.directive_json` = `{}` on both escrowed hosts; `dr_recipe.host_half.drives` = `[]` on every customer **including two with enrolled data drives** (916 GB USB on demo-felhom, 938 GB NVMe on demo-hp) **MIGRATED FROM `ROADMAP.md` 2026-08-22 (R-369) — originally filed 2026-07-28, size M, roadmap state `READY — 2026-07-28`.** Moved verbatim; nothing added or reinterpreted. The roadmap keeps its copy as history, marked moved. ****PARTLY FIXED BY ITS OWN UPDATE.** The `drives` third was traced and populated on both demo boxes on 2026-07-28 (the enrolled drives were never PVE storages, so `isUserDataDrive` never saw them). The other two thirds — `hosts.dr_record_json` and `host_escrow.directive_json` — were NOT re-verified this session and are carried as written.** | **OPEN — migrated from ROADMAP 2026-08-22, rank unchanged** | — | These are exactly the fields a host-loss recovery reads: `05-hub-architecture.md:175-176,186` names the slim DR record as one of four durable sources; `06-offsite-connectivity.md:148-150` says the escrow upload carried the DR directive; `felhom-agent/internal/dr/plan.go:34-35` makes `PlannedDrive` the re-attach-by-`durable_id` wrong-disk guard. **The three may have different causes** — `isUserDataDrive` (`internal/hub/dr_recipe.go:129-136`) requires type `usb`/`local-dir` **and** a non-empty `DurableID` **and** `MountPath`, and which of the three fails was not traced. Evidence: `architecture/_recovery-inventory-2026-07-28.md` Part D2.3. **UPDATE 2026-07-28 (vzdump-target move): the `drives` third is TRACED and now POPULATED on both demo boxes.** Cause: the enrolled data drives were never PVE storages at all — only agent-generated systemd mounts — so they never entered `report.StorageTargets` and `isUserDataDrive` never saw them. Giving each drive a `dir` storage at its own mountpoint supplied all three required fields at once (type `local-dir`, fs-UUID durable id, mount path), and the recipe now emits `uuid:91d2dc2d-…`/`/mnt/nvme-1tb` on demo-hp and `uuid:47a3361a-…`/`/mnt/hdd_1` o | CC | | **R-232** | Backup & restore | P2 | **DooPlex's backup makes every copy inside the same box — and nothing tells anyone when it fails.** Surveyed read-only 2026-08-06 (`audits/RECON-dooplex-backup-2026-08-06.md`). **What works:** five sets, 14/14 successful runs in 14 days; a file was restored from the `data` repo and matched the live original **byte for byte**; every set except two is cross-disk; k3s is integrity-checked on every run. **What the matrix exposes, ranked:** (a) **`notify_failure` is a no-op** — `NOTIFY_ON_FAILURE=true` but `NOTIFY_WEBHOOK_URL` is commented out, so a failed backup notifies **nobody**; the project already has a working Resend path that CI uses. Cheapest item, and it makes every other failure visible. (b) **Nothing leaves the box** — no rclone, no remote repo, no off-site target anywhere; Longhorn's target is `nfs://192.168.0.180:` pointing at DooPlex itself, and the only outbound-looking cron pulls *inbound* from Hetzner for a different project. The machine that runs the hub managing the customers' off-site chain has no off-site copy of its own. (c) **The backup tree is a single writable path** and the restic repos are not append-only — one bad script or ransomware destroys every copy at once. (d) **Two same-disk sets**: `.claude-memory` and the PostgreSQL dumps, whose source directory sits *inside* the backup tree. (e) **Longhorn `retain=1`** — one generation per volume, so a corruption noticed a day late has no earlier copy. (f) **`/opt/backup/docs/BACKUP-RESTORE.md` does not exist** though the systemd unit advertises it. (g) **`secrets/restic-repo` has never held a snapshot** — `backup-secrets.sh` contains no `restic` call; the secrets are GPG files on `sda1` only. (h) **No restore has ever been run** beyond today's single-file probe — the matrix's "ever demonstrated?" column is otherwise entirely empty. **Not a finding:** the restic passphrase. The on-box copy is on `sdb1`, a different disk from the backups, and the **operator holds an offline copy out of band** — so a disk loss is recoverable. The narrow residual is that it is operator-held rather than system-held, unlike the customer case's hub-vaulted escrow, so it should be confirmed current and findable by someone else. **Nothing was changed by the recon.** | **READY** — owner Viktor | — | — | operator | | **R-304** | Backup & restore | P2 | **The retained escrow key works, and the customer is told their correct code is wrong.** DRILL 2026-08-12 answered the three questions separately, on `demo-felhom`, with planted data. **(a) retention: WORKS** — the first retained row in fleet history to carry material (`host_escrow_superseded` id 11, `identity_blob` 572 B), byte-identical (`sha256 a10032341c8584ed…`) to the pre-supersession `host_escrow` row. **(b) the material opens the old store: YES** — unsealed with the OLD recovery code it yielded a password byte-identical to the pre-change one (`sha c60c8bc737a6b7c6…`), and restored three planted files **byte-identical** from a store the box itself could no longer open (negative control first: `Fatal: wrong password or no key found`), **including a Hungarian accented filename verified as raw bytes**. **(c) the customer's route: DOES NOT EXIST, and misinforms.** `ListSupersededEscrow` (`store.go:2841`) is the only reader of a retained `identity_blob` and has **zero production callers** — five call sites, all `_test.go`; the product path (`POST /escrow/recover-offsite-password` → `FetchIdentityEscrow` → `GetHostDRBundle`, `store.go:3152`) selects `FROM host_escrow` — the CURRENT row only. Asked for the old password with the code that demonstrably opens the retained row, the product answered **"the recovery code did not open the sealed bundle — nothing was written"**. **This is the R-224 class again**: there an unreachable hub was reported as a bad code; here a VALID code for retained history is reported as a bad code, and the customer's attempt ends there. **Consequence:** the census answer stands (it was about retention); the countdown banner's promise is true in substance and false in practice; **any capability-map claim that the customer can recover the old history with their recovery code is false today and must move** | **READY (L) — NEW 2026-08-12, RANK 1** | R-198, R-199, R-224, R-241 | Decide the shape: serve retained rows on the recovery path (needs a "which package?" choice — a customer may have several), or stop promising retrieval anywhere the customer cannot perform it. **Until one of those, the honest position is that retention is an operator-only capability.** At minimum, the refusal must stop asserting the code is wrong when the hub simply never looked | operator + CC | -| **R-342** | Backup & restore | P2 | **The ep0 snapshot covers less than it looks like it covers, and the next person will assume otherwise.** Quoting `audits/evidence-ep0-pbs-upgrade-2026-08-18/stop2-snapshot.txt` verbatim: *"covers — the 38 GB system disk /dev/sda (root), i.e. the PBS packages, unit files, /etc/systemd drop-ins, nftables and wg config. DOES NOT — /mnt/pbs-datastore. That is /dev/sdb, a separate 100 GB VOLUME, and Hetzner server snapshots do not include attached volumes. The backup data is therefore NOT protected by this snapshot."* Snapshot **421440873** (`felhom-hetzner-20260818`, 15.06 GB, Available) was taken as the rollback for the 4.2.2→4.2.5 PBS upgrade. **Rolling it back restores software state, not the datastore.** That was *acceptable for that change* — a package install writes no datastore content — and the file says so. **The problem is what happens next:** this fact lives in an evidence file nobody will open again, and a snapshot named as "the rollback" reads as protecting everything on the box. ep0 holds the only off-premises copy of a real customer's data | **READY (S) — NEW 2026-08-18** | — | **Decide the safeguard for any future ep0 procedure that could touch `/mnt/pbs-datastore` — it does not exist and has not been designed.** Candidates: a Hetzner **Volume** snapshot (a different object from the server snapshot), a PBS-level sync to a second location, or an explicit written acceptance that the datastore is unprotected for the duration. **Nothing may be added to a runbook implying a safeguard exists until one does** | **Viktor decides; CC executes** — a risk-to-customer-data question | +| **R-342** | Backup & restore | P2 | **The ep0 snapshot covers less than it looks like it covers, and the next person will assume otherwise.** Quoting `audits/evidence-ep0-pbs-upgrade-2026-08-18/stop2-snapshot.txt` verbatim: *"covers — the 38 GB system disk /dev/sda (root), i.e. the PBS packages, unit files, /etc/systemd drop-ins, nftables and wg config. DOES NOT — /mnt/pbs-datastore. That is /dev/sdb, a separate 100 GB VOLUME, and Hetzner server snapshots do not include attached volumes. The backup data is therefore NOT protected by this snapshot."* Snapshot **421440873** (`felhom-hetzner-20260818`, 15.06 GB, Available) was taken as the rollback for the 4.2.2→4.2.5 PBS upgrade. **Rolling it back restores software state, not the datastore.** That was *acceptable for that change* — a package install writes no datastore content — and the file says so. **The problem is what happens next:** this fact lives in an evidence file nobody will open again, and a snapshot named as "the rollback" reads as protecting everything on the box. ep0 holds the only off-premises copy of a real customer's data | **READY (S) — NEW 2026-08-18** | — | **Decide the safeguard for any future ep0 procedure that could touch `/mnt/pbs-datastore` — it does not exist and has not been designed.** Candidates: a Hetzner **Volume** snapshot (a different object from the server snapshot), a PBS-level sync to a second location, or an explicit written acceptance that the datastore is unprotected for the duration. **Nothing may be added to a runbook implying a safeguard exists until one does** **2026-10-03 — options costed, one candidate removed:** Hetzner Cloud has NO volume snapshots (server snapshots exclude volumes; open vendor request since 2019), so the first candidate does not exist. Options left: (A) a PBS pull-sync of `felhom-offsite` to DooPlex's PBS — €0/month, 1–2 h, protects against datastore loss AND losing Hetzner, ciphertext only (per-customer `encryption-key`), but a new job on DooPlex; (B) a written acceptance plus a copy before each risky procedure. CC recommends A. Datastore today 19.5 GB of 97.9 GB. `audits/offsite-append-only-2026-10-03/PART-D-ep0-safeguard.md`; decision in STATUS. | **Viktor decides; CC executes** — a risk-to-customer-data question | | **R-366** | Backup & restore | P2 | **The 21 August reinstall orphaned `demo-hp`'s PBS whole-guest archives as well as its off-site repo — the box can no longer read its own pre-reinstall backups, and this surfaces only as a restore-test failure.** Hub event 3016, 2026-08-21 21:59:28Z, unprompted: `Restore-test FAILED on the pbs tier: archive felhom-pbs:backup/ct/9201/2026-08-18T03:58:43Z could not be restored+booted … proxmox-backup-client failed: Error: wrong key - unable to verify signature since manifest's key 3f:4f:65:c0:d8:f3:9f:3c does not match provided key dd:d1:d8:53:44:62:5e:0b`. The archive predates the reinstall by three days. **This is the PBS-tier analogue of R-193** (a guest rebuild mints a fresh secret and orphans the history), and the two together mean a rebuilt box loses BOTH off-premises tiers at once: the restic repo needed a self-heal + re-toggle (see the drill report), and the PBS archives are simply unreadable to it. **Credit: the restore-test caught it and said so precisely** — the mechanism works. **The gap is what it is called:** it is reported as *a restore test that failed*, which reads as a flaky verification, not as *every whole-guest backup you took before the reinstall is unreadable on this machine*. **Found incidentally by the 2026-08-21 backup-truth drill; nobody was looking for it.** | **OPEN — HIGH** | related: R-193 | Establish whether the pre-reinstall PBS archives are recoverable at all (the old key's whereabouts), and separate the two verdicts: a tier whose ARCHIVES ARE ORPHANED is a different alarm from a tier whose restore test failed. Do not close on the strength of the restore-test wording alone. | CC | -| **R-436** | Backup & restore | P2 | **LEAD, NOT A DEFECT — append-only may be reachable without a new machine, which would make R-95's real prevention far cheaper than the spike concluded.** Hetzner's port-23 restricted shell advertises, in its own `help`, these server-side backends: `borg`, `rsync`, `scp`, `sftp`, **`rclone serve restic --stdio`**. And restic 0.14.0 **recognises the `rclone:` backend** — MEASURED 2026-09-01 with a control: `banana:` → `Fatal: parsing repository location failed: invalid backend`, while `rclone:` → `exec: "rclone": executable file not found in $PATH` (i.e. the backend parsed and it tried to run the helper). rclone is **not** in the controller image today. `rclone serve restic` carries an **`--append-only`** flag. **Why this matters:** the R-95 spike's option 3 was priced at a new always-on service in the recovery path plus either a mount in the hot path or migrating every customer's history — and it was deferred on that price. This route needs neither: the server side already runs at the provider. **THE CAVEAT, STATED FIRST because it may kill the idea:** the **client** supplies the server command line, so a compromised guest could simply omit `--append-only` unless the provider pins it. **NOT ESTABLISHED:** whether Hetzner pins the flag or accepts client-supplied arguments. **That is a vendor question and it is cheap** — it should be asked before any code is written, because if the answer is "client-supplied" this lead is worth nothing. **STRENGTHENED 2026-09-01 (operator supplied the page): the backend is OFFICIALLY DOCUMENTED, not merely advertised in a shell banner.** `docs.hetzner.com/storage/storage-box/access/access-ssh-rsync-borg/#restic` reads: *"Restic is natively supported with the SFTP backend. As another option, we support the restic backend, which is provided by Rclone over SSH."* So the transport exists as a supported product feature and the client half is already proven (restic 0.14.0 parses `rclone:`, measured with a control). **AND THE SAME PAGE SETTLES THAT THE DOCS CANNOT ANSWER THE CAVEAT: neither its Rclone nor its Restic section mentions append-only at all.** That is worth stating because it closes the cheapest alternative to asking — nobody need re-read the documentation hoping for it. **Corroboration, unlooked for:** that page's table of port-23 commands matches, item for item, the `help` output measured live on our own sub-account — independent confirmation that the live measurement was reading the right product's surface. **-- ANSWERED BY HETZNER 2026-09-22 (ticket #2026090103040671), and the lead is GOOD.** Asked directly whether `--append-only` is enforced by Hetzner or taken from what the client sends, support replied: *"You can use --append-only and it would look like this: `command="rclone serve restic --stdio --append-only path/to/repo" `"*, pointing at `fluix.one/blog/hetzner-restic-append-only/`. **The mechanism is an SSH FORCED COMMAND pinned to the key in the Storage Box's `authorized_keys`** - so the flag is enforced on the server side, by a line WE write, and a client asking for a plain `rclone serve restic --stdio` on that key gets the forced one instead. **That is real prevention without a new machine, which is what this lead claimed and could not confirm.** **What it does NOT say, and must not be read as saying:** a key WITHOUT that prefix still gets a deleting server, so this protects a repository only if every key that can reach it carries the forced command - including any key an operator adds later. **NOTHING IS MEASURED YET.** Owed, and cheap: a forced-command key on `u629488`, a restic backup through it that SUCCEEDS, and a `forget --prune` through it that is REFUSED - with the refusal quoted. Until that runs this is a support answer, not a property of our backups. | **OPEN — ask the vendor before building anything** | — | — | CC | | **R-518** | Backup & restore | P2 | **[P2-MEDIUM] „Mentés most" on the whole-system backup stops every app for about eight minutes while the page promises „csak néhány másodpercre".** MEASURED 2026-09-14 (BIGNIGHT, VM 333, 12 apps): the button's call quiesced all 12 stacks at 19:03:23Z (first stopped 19:03:27Z); the local vzdump ran 19:03:49 → 19:09:59Z; the controller then kept the apps stopped for the second (PBS) tier and restarted them at 19:10:09Z after it failed, the last started 19:11:12Z (`phase4/guest-backup-quiesce-log.txt`) — **≈ 7 m 45 s** with every app answering 404. The page under the button: „Pillanatkép-mód: az alkalmazások csak néhány másodpercre állnak le." A household pressing it at dinner loses every app for the length of the dump, and longer on a bigger box. **Fix shape:** state the real expected downtime (it scales with data), or quiesce per tier and not across a second tier's attempt; do not start a tier whose storage is absent (see R-517). **NARROWED 2026-09-15 (controller v0.243.0 + agent v0.131.0):** a tier whose storage the agent reports absent is skipped before anything stops (`backup_tier_skipped`, once per absence; unknown never skipped), and the button copy now says „általában néhány perc, nagyobb adatnál több". Unit-proven with red-proofs. **Still open:** quiesce per tier, so a slow second tier does not keep every app down. **— NIGHT 2026-09-23 (controller v0.267.0):** the copy half is DONE: the page and the confirm now state the measured stop (≈ 8 minutes on a 12-app box), both languages, red-proofed (`audits/night-2026-09-23/A5-*`). The brief's „csak néhány másodpercre" had already gone in v0.243.0. **Still open:** quiesce per tier, so a slow second tier does not keep every app down. | **READY — P2, narrowed to per-tier quiesce; owner: CC (controller)** | — | — | CC | | **R-519** | Backup & restore | P2 | **[P2-MEDIUM] After a backup torn by a power cut, an app's restore point carries the new database dump's time while its files are from the previous run — and no customer screen says the run was interrupted.** MEASURED 2026-09-14 (BIGNIGHT F2, VM 333): „Mentés most" 19:40:03Z; the power was cut 19:40:08Z while adventurelog was stopped for its volume dump. On disk afterwards, `backups/primary/adventurelog`: `db-dumps/adventurelog-postgres.sql` **19:40:07**, `volume-dumps/*.tar` **19:00:35**, `manifest.json created_at 19:02:34Z`; bookstack the same shape (sql 19:40:07, tars 19:00:48). `GET /api/backup/snapshots` for both → `time 2026-09-14T19:40:07Z, helyi` (`phase5/F2/units-on-disk.txt`, `backup-honesty.txt`). `/backups/apps` shows „Utolsó adatbázis mentés 2026-09-14 21:40 · … OK" and every app „Utolsó: 5 perce"; `/backups` and `/dashboard` contain no word of an interruption (fragments `megszakad|sikertelen|nem sikerült` = 0; control „hiba" appears in the standing warning text). The controller itself knew: `[appstop] crash recovery: an app-data backup (volume dump) … was interrupted … restarting them: [adventurelog]` and pushed `backup_failed (error)` to the hub. A household restoring „the 21:40 backup" gets 21:00 files for BookStack's uploads. **Fix shape:** date a point by the oldest part it contains (or mark it partial) and show the interrupted run on the backups page until the next complete one. **F6 (drive unplugged 1 s into a backup, 20:37:59Z) adds three facts:** the run skipped four apps' volume dumps („Skipping volume dump for immich — drive disconnected", also jellyfin, nextcloud, paperless-ngx) and still reported `db_dump {"count":6, "success":true}`; nextcloud's point is dated 20:37:59Z (its SQL finished before the unplug) beside 19:01 volume tars; and a torn `immich-postgres.sql.tmp` (20:38:00) plus F3's `pre-restore-…-nextcloud-mariadb.sql.tmp` are left in the units on the drive. Immich's point correctly stayed at 19:02:34Z (the `.tmp` was not promoted). `/backups/apps` fragments `kihagy|sikertelen|részleges` = 0 (`phase5/F6/`). | **READY — rank P2-MEDIUM; owner: CC (controller)** | — | — | CC | | **R-638** | Backup & restore | P2 | **[P2-MEDIUM] The product's database loader cannot replay a copy over a NEWER schema: on PostgreSQL it FAILS, on MariaDB it leaves the newer version's tables behind.** MEASURED 2026-09-23 on 9202. `ImportDump` (`appbackup/dbdump.go:719`, `psql -v ON_ERROR_STOP=1 --single-transaction`) replays a `pg_dump --clean --if-exists` file over the live database. After docmost 0.95.0 → 0.96.0 migrated, the replay of the pre-update copy was refused in 0.40 s, rc 3: *cannot drop constraint workspaces_pkey on table public.workspaces because other objects depend on it / DETAIL: constraint oauth_clients_workspace_id_fkey …* — the new version created six tables whose foreign keys point at old ones, and `--clean` only drops what the dump knows. Database unchanged (the transaction rolled back). On MariaDB (`mariadb-dump`, `FOREIGN_KEY_CHECKS=0`) the same replay after romm 5.0.0 → 5.3.0 returned rc 0 in 1.25 s and left **12 base tables** of the new version behind; RomM 5.0.0 happened to ignore them. **What worked:** `DROP SCHEMA public CASCADE; CREATE SCHEMA public;` + the dump in ONE transaction — rc 0 in 1.38 s, every table, index and extension back. **Why this is a row of its own and not only part of R-637:** the SAME loader backs shipped paths — `rollbackSafetyDump` (off-site restore's undo) and the dump replay of the restores — so **any restore of a copy taken BEFORE an update that migrated, replayed over the migrated database, may fail the same way. NOT MEASURED:** whether the unit restore the hold sentence names does this (it also carries the data VOLUME tar, which may make the replay moot). That is the measurement owed, on 9202, before anyone relies on it. Evidence: `audits/update-rulings-2026-09-23/README.md` Part 1, `docmost-45`, `romm-44`. **-- NARROWED 2026-09-23:** the undo no longer touches this loader — it copies folders (decision 19, controller v0.263.0). **What stays open is the part about SHIPPED paths:** `rollbackSafetyDump` and the restores' dump replay still replay over whatever schema is live, and whether the unit restore the hold sentence names works after a real schema migration is STILL UNMEASURED. | **OPEN — P2, narrowed to the restore paths; owner: CC; measure the named restore after a real schema migration first** | — | — | CC | | **R-726** | Backup & restore | P2 | **[P2-MEDIUM] A new box for a customer who had one before makes NO off-site copy on night one: the old repository is found orphaned, and the fix is a button nobody pointed the household to.** MEASURED 2026-09-30 00:15 UTC on the new-household drill box (`tester-1`, whose previous box was deleted 2026-09-17): `[offbox] offsite repo ORPHANED — remote holds backups written under a previous, no-longer-available key; runs will skip until reset` → `offbox_repo_orphaned` (warning) to the household's timeline and an operator mail; `offsite-integrity` then checked nothing. The household's page is honest („A távoli tároló másik kulccsal készült mentéseket tartalmaz … Új távoli mentés indítása…", old history set aside, never deleted), but the evening before, the recovery-code ceremony and the off-site page raised nothing, and the guide does not mention it. The hub re-issued the off-site credentials on re-enroll by itself; it could have known the repository would orphan. Customer data was never at risk (the old repository is untouched); the household simply has no off-site copy until someone presses the button. **Fix direction:** offer the reset at the recovery-code ceremony when the repository already holds another key's snapshots, or the hub's re-enroll re-issue sets the old history aside the same way (it is the same move-aside), and the guide says so. | **READY — rank P2-MEDIUM; owner: CC (design first) / operator (which)** | — | — | CC + operator | +| **R-822** | Backup & restore | P2 | **An add-only key does not make retention safe: an attacker who can only ADD snapshots can make the honest pruner erase every real one.** MEASURED 2026-10-03 (lab, restic 0.14.0 from the controller image, rclone `--append-only`): 13 empty snapshots dated in the future with the same host and tag, added through the add-only key (all allowed), make the box's exact policy `forget --group-by host,tags --keep-daily 7 --keep-weekly 4 --keep-monthly 6` keep only the fakes and select **all 3 real snapshots** for removal (`--dry-run`). **Whoever prunes an append-only repo — the box in a window, or a Felhom-side worker — inherits this.** Not a defect today (today the box can simply delete, R-95); a PRECONDITION on the R-95 build, like R-430 was. `audits/offsite-append-only-2026-10-03/lab/C3-retention-poisoning.txt` | **OPEN — precondition on any R-95 build** | — | Before any `forget`: refuse when a snapshot is dated in the future or newer than the newest the hub has seen reported; dry-run first and abort above the count the policy can remove in a week; hub compares the count before/after (DESIGN.md §3) | CC | | **R-49** | Backup & restore | P3 | **[P2] The offsite capture set is ~90% cache and duplication — 1.1 GB of a 1.2 GB immich "photo backup".** Measured 2026-07-19: `immich_ml_cache.tar` **823 660 032 B (~60%)** — re-downloadable ML model weights; `immich_postgres_data.tar` **308 251 136 B (~23%)** — a raw tar of the postgres data dir that DUPLICATES the logical `.sql` dump captured beside it; `upload/backups/` **18 MB** — immich's own nightly dump, a backup inside the backup, growing daily; plus the stranded pre-v3 `dccc13fe…` tree (~36 MB) no DB has ever referenced. Actual irreplaceable content: **72 MB of originals**. **MIGRATED FROM `ROADMAP.md` 2026-08-22 (R-369) — originally filed 2026-07-19, size S–M, roadmap state `idea`.** Moved verbatim; nothing added or reinterpreted. The roadmap keeps its copy as history, marked moved. | **OPEN — migrated from ROADMAP 2026-08-22, rank unchanged** **Re-ranked 2026-10-03: P2→P3: wasted space and transfer, no data risk; needs a capture-set ruling, not a sale blocker.** | — | **Evidence: `audits/DIAG-immich-restore-round2-2026-07-19.md` §4 (full byte breakdown).** This is the customer's offsite quota and transfer cost, and it lands on the Hetzner sub-account they are billed for. **Recorded, deliberately not changed** — a capture-set exclusion is a data-loss-shaped decision and gets its own ruling, not a drive-by edit. Candidates in priority order: (a) `immich_ml_cache` — pure cache, strongest case; (b) the `postgres_data` volume tar where a logical dump of the same DB is already captured (the dump is what the restore path actually replays); (c) `upload/backups/`. Likely generalises past immich into a template-classification rule about cache volumes and self-backup directories, so it should be specified against the catalog, not one app | CC | | **R-127** | Backup & restore | P3 | **The catalog's `data_key: true` flag is UNRELIABLE — at least four data-encrypting keys the catalog itself labels as encryption keys are unflagged; and the O4 restore path can regenerate a DB password that then does not match the restored data directory** | **READY (S/M)** | — | **Found by D5's Part 0, and it is why D5's boundary is `type: secret` rather than `data_key`.** Two separable legs. **(a) The misclassification.** Only 5 fields across 4 apps set `data_key: true` (`adventurelog/SECRET_KEY`, `homebox/HBOX_AUTH_API_KEY_PEPPER`, `papra/AUTH_SECRET`, `sparkyfitness/{API_ENCRYPTION_KEY,BETTER_AUTH_SECRET}`), yet `n8n/N8N_ENCRYPTION_KEY` („Titkosítási kulcs"), `wanderer/POCKETBASE_ENCRYPTION_KEY` („Adatbázis titkosítási kulcs"), `calcom/CALENDSO_ENCRYPTION_KEY` and `bookstack/APP_KEY` are unflagged — the catalog's own Hungarian labels contradict the flag. **D5 makes this non-urgent but not harmless:** everything `type: secret` now travels, so the keys DO reach the drive; what stays wrong is the **fail-closed gate**, which only refuses for `data_key` names — so if one of these is missing from both sources the restore proceeds onto data it cannot decrypt instead of refusing. Fix = flag them (`app-catalog-felhom.eu`, a catalog-only change) + a gate/test that the flag set and the label set agree. **(b) The regenerated-DB-password trap.** `internal/backup/restore_unit.go` O4 generates a replacement for any missing non-data-key secret. Proven on `postgres:16-alpine`: with PGDATA restored from the volume tar, `POSTGRES_PASSWORD` is ignored (initdb skipped), so the app fails over the compose network while the dump replay still succeeds through the container's local **trust** socket — success reported, data unreachable. v0.188.0 corrected the WARN's false claim that *"stored data is unaffected"* and scoped it, but did **not** add a guard: D5 shrinks this to the rare case (the secret was empty at capture AND absent from the guest). Real fix = either treat a DB password as fail-closed like a data key, or `ALTER USER` to the regenerated value after the volume restore. 18 DB/root-password fields are in scope; MariaDB fails loudly instead (Access denied), which is the safer half | CC | | **R-200** | Backup & restore | P3 | **The DR password-injection seam has a handler, a route and tests — and no form.** `POST /backup/offbox/inject-password` is routed (`controller/internal/web/server.go:510`) to `offboxInjectPasswordHandler` (`offbox_handlers.go:174-196`) → `InjectOffboxPassword` (`backup/offbox.go:541`). **No template in the repository contains that path or any form posting to it** (grep over `internal/web/templates/`: one unrelated hit, an XSS comment) | **NARROWED** — **PLUMBING COMPLETE** (controller **v0.196.0**); **the FORM is not built — still open** | — | **The tenth instance of this project's built-but-never-wired class, and the exact shape `CLAUDE.md` and `felhom.eu/CLAUDE.md`'s seam-wiring rule were written for:** handler tests that POST directly (`offbox_escrow_test.go:167,180`) prove nothing about reachability. To use the only implemented recovery seam today, a person must hand-craft an authenticated POST with a session cookie and CSRF token. **Note the layering while fixing it:** this form takes a **64-hex repo password**, not a recovery code (`offboxRepoPwPattern`, `offbox.go:543`) — they are different secrets at different layers, and the operator's 2026-08-04 ruling asks for a form that takes **R**. Build the R form and treat this one as the operator/DR fallback it was written as, but **ship it with a render test per branch of whatever gate it sits behind**. Source: `audits/RECON-offsite-dr-chain-2026-08-04.md` §3 link 9 **THE DIAGNOSTIC HALF IS DONE AND IT ANSWERED THE QUESTION.** `--recover-offsite-check` is a `docker exec` escape hatch in the shape of `--print-reset-code`: R on **STDIN** (never argv, never `ps`, never shell history, never a transcript), fetch+unseal via the agent, and a verdict of **two sha256 hashes**. **It compares and never installs** — the recovered password is not written to `offbox/repo_password`; a test asserts the data dir is byte-unchanged and its red-proof (adding the install call) fails it. Confirmed live: `repo_password` mtime still `2026-08-03 07:18:02` after the successful check at `2026-08-04 11:49`. **Exit codes are load-bearing** — `0` match, `2` a clean MISMATCH, `1` a step failed; "it failed" and "it worked and disagreed" must never share a status because only one is a finding about the system. A box with no local password reports distinctly (the rebuilt-box shape, where the next step is to INSTALL rather than compare). **WHAT IS NOT BUILT, deliberately:** no card, no form, no preview, no customer-facing text — building an interface on top of a chain nobody had walked is how the preceding three weeks went wrong. **What remains for this row:** link 9 (the recovered password placed so `WriteOffboxSecrets` keeps it) and the customer-facing shape the operator ruled on 2026-08-04 (yell → R form → preview → proceed), which is now priced against a chain that exists rather than one that is assumed **PART 0 SHIPPED 2026-08-04 (v0.196.0):** `--recover-offsite-install` is the sibling of the check — same fetch/unseal path, same STDIN discipline for R — and it **places** the recovered password via `InjectOffboxPassword`. The confirmation is a SECOND invocation (`--confirm-install`): without it, both hashes print and nothing is written, so the operator sees the comparison before a write is possible. Three outcomes named distinctly: **installed** (no local password — the rebuilt-box shape), **unchanged** (identical key, nothing written), **refused** (a DIFFERENT key present — installing would clobber the key the current repository is encrypted under; exit 2, no force offered). It re-reads the file after writing rather than trusting the call. **Red-proof observed:** removing the confirmation gate makes the dry run write the password. The R-persistence test carries a **positive control** (a planted copy found, then removed and not found). **NOT YET EXERCISED AGAINST A LIVE RECOVERY** — the R-201 drill halted before step 9, so this is unit-proven only. **What remains for this row:** the customer-facing shape the operator ruled on 2026-08-04 (yell → recovery-code form → preview → proceed) | CC | @@ -220,7 +220,6 @@ stopping line that lies. | **R-401** | Backup & restore | P3 | **Revisit the off-site integrity depth when a real store is LARGE — the default rests on ONE measurement, on ONE 134 MB store.** Controller v0.228.0 (R-399) made `--read-data-subset=100%` the default for every box. The whole justification is a single data point: `demo-hp`, 2026-08-30, 140 829 678 B / 2 651 blobs / 67 snapshots, structure 35.0 s vs 100% **39.2 s** — four seconds. Re-proven live 2026-08-31 at 38.7 s. **It does not extrapolate, and the reason is structural: the structure check's cost tracks the INDEX, a read-data run's tracks the DATA.** A 50 GB store is ~370x the data and this curve says nothing about it. **Nothing was invented from that one point** — no rotation schedule, no size threshold, no bandwidth budget — because four production designs in this project were specced against unvalidated mechanisms and all four were wrong. **`readDataSubsetRe` already accepts `n/m`**, so a rotating schedule (`1/7` on a different seventh each week) needs no parser work when the time comes; the missing input is a measurement on a large store, not code. **THE TRIGGER IS AN EVENT, NOT A DATE:** the slow-check WARN from v0.228.0 firing on any box (`integritySlowNoticeThreshold`, 5 min) — that line names the duration, the depth and this row. **WHAT HAPPENS IF NOBODY ACTS:** every box re-reads its entire store every week, however large it grows, and the first person to notice is a customer whose upload is saturated. **Whoever acts must also revisit `integrityCheckTimeout` (30 min)**, which is now the number a large store meets first. | **OPEN — WATCHING** | — | When the WARN fires: measure the curve on that store, then choose between a rotation (`n/m`), a size-conditional default, or leaving it. Do NOT choose from this row's numbers — they are the small-store case. | CC | | **R-409** | Backup & restore | P3 | **Nothing in the product can vouch for the bytes of a restored recovery unit — the only hash record covers 0.002 % of it.** MEASURED on demo-hp 2026-08-31 against kimai's restored unit: `manifest.json`'s `checksums` object carries sha256 for `.felhom.yml` (2 235 B), `app.yaml` (488 B) and `docker-compose.yml` (2 195 B) — **4 918 bytes of a 213 231 242-byte unit**. The database dump (48 217 B) and the two named-volume tars (160 331 776 B + 52 845 056 B) — the recoverable data, 99.998 % of the bytes — have no recorded hash anywhere. **And nothing else supplies one:** restic 0.14.0's `restore --verify` is a size-and-mtime reconciliation (a size-and-mtime-preserving one-byte corruption of a 160 MB tar passed clean, red-proofed), and `restic ls --json` file nodes in 0.14.0 carry name, size, mode, uid/gid and three timestamps and **no content hash**. **So "the restore produced correct files" is currently unanswerable by any automated means.** **What is NOT claimed here:** `restic check --read-data-subset=100%` already proves the STORE's packs, and the config files that ARE hashed are the ones a wrong-content failure would be hardest to spot in. | **OPEN — MEDIUM** | R-87, R-361 | Cheapest fix, and it is already half-built: extend the capture's `checksums` to cover `db_dumps` and `volume_dumps` — R-361 already computes a canonical dump sha256 to prove itself, so the value exists at capture time. Then a restore-test has a real reference and R-87's narrow version becomes a content check rather than a completeness one. Evidence: `audits/SPIKE-restic-restore-test-2026-08-31.md` §Q2, §Q3. | CC | | **R-412** | Backup & restore | P3 | **A recovery unit lost DURING an off-site run — after its own dump leg, before its push — is shipped hollow and the run reports success.** **CORRECTED 2026-09-01 04:22, and the first wording of this row OVERSTATED it.** As first filed it claimed the hollow unit sat in the store for a whole cycle because "the volume-dump leg runs on the backup schedule, not on capture". **That is wrong, and measuring it overnight is what showed it:** the off-site run has its OWN pre-push dump leg — *"Stopping calibre-web for safe volume dump"*, *"Volume dump: calibre-web/calibre-web_calibre_web_config -> 877.5 KB"* — so a unit that is hollow when a run starts is **REPAIRED before it is pushed**. Proven twice: `opengist` (2026-08-31 21:0x) and `calibre-web` (2026-09-01 04:15) both went in hollow and came out complete, and the snapshot pulled back from the store (`6fee3b5a`) holds the volume tar and all 17 userdata files. **WHAT REMAINS REAL, and it is narrower:** the one hollow snapshot that DID reach the store (`35ba9fe7`, opengist) was created when the unit was destroyed **inside** a run that had already completed opengist's dump leg — so the push shipped what the capture had just rebuilt empty, and logged *"backed up opengist (… 0 mandatory path(s))"*, **a success line over a backup holding none of the app's data**. That race is real, it was observed, and the success wording is wrong either way. **The R-403 mirror guard holds throughout** — proven live: *"unit leg SKIPPED … The copy was PRESERVED rather than replaced with an empty one"*, secondary byte-identical. | **NARROWED** — **LEG 1 CLOSED 2026-09-01 (controller v0.232.0) — LEG 2 STILL OPEN, LOW** | R-403, R-87, R-413 | **LEG 1 IS DONE:** a per-app push whose unit carried no database dump and no volume tar now logs at WARN and says what it did not carry, using the existing `unitIsHollow` predicate. Wording only — no guard, and the capture is untouched (08 §8.2). Pinned by `TestR412a_EmptyPushDoesNotReadAsAPlainSuccess`, which asserts the hollow line carries the words, the SOUND line does not, and neither is at INFO; red-proofed by restoring the single unconditional line. **LEG 2 IS STILL OPEN and is the remaining work on this row:** whether the push should RE-READ the unit it is about to send, or whether the race window is small enough to accept. Two separable things. (1) The success line: a per-app push that carried no dumps and no tars should not read as a plain success — that is a wording fix in the run's own reporting, not a new guard. (2) The race: decide whether the push should re-read the unit it is about to send, or whether the window is small enough to accept. **Do NOT guard the capture** (08 §8.2). Evidence: `audits/DRILL-soak-2026-08-31/phase2-guard-interactions/` and `phase5-mutated-cycle/09-what-reached-the-store.txt`. | CC | -| **R-430** | Backup & restore | P3 | **`restic unlock --remove-all` printed `successfully removed locks` while the lock was still there.** MEASURED 2026-09-01 in a throwaway local repo (no live store touched), under a faithful append-only model — a sticky locks directory owned by root holding a root-owned lock, restic run as `nobody`; both controls passed first (create allowed, delete refused). The command reported success, returned, and `ls` showed the lock present. **`resticStep`'s crash-lock self-heal is built directly on this call** (`felhom-controller/controller/internal/backup/offbox.go:~768`), and its licence to escalate rests on the escalation actually working. **A self-heal that cannot fail is a self-heal that cannot be trusted** — this is this project's *"exit codes that lie"* class, in the one path that runs unattended against the customer's off-site history. It is harmless TODAY because the credential can delete and the removal really happens; it becomes load-bearing the moment delete is withdrawn, which is what R-95 is about. **Not yet established:** whether restic reports success because it removed zero locks by design, or because it did not check. Settling it: read restic 0.14.0's unlock source, or re-run with `--verbose`. **MARKED LATENT 2026-09-01.** It is **harmless today** and the reason is precise: the credential CAN delete, so the removal really happens and the success it reports is accidentally true. **THE CONDITION THAT MAKES IT LIVE — the only one, so it is stated as a trigger and not as prose: the moment delete is withdrawn from the box.** That is exactly what R-95's remedy does, by either route (retention moved off-box, or an append-only transport via R-436). From that moment `resticStep`'s crash-lock self-heal is escalating with a call that cannot fail, in the one path that runs unattended against the customer's off-site history. **So this is a PRECONDITION on the R-95 build, not a follow-up to it** — settle it in the same change or the self-heal ships already broken. | **OPEN — LATENT; becomes live the moment delete is withdrawn (precondition on any R-95 build)** | — | — | CC | | **R-433** | Backup & restore | P3 | **A sub-account cannot reach ANY Storage Box snapshot, by any name — so clause (b) of the 2026-09-01 R-95 re-scope ("the rest is recoverable file by file") is NOT SUPPORTED.** MEASURED 2026-09-01 on `demo-hp` over the credential the box already holds, read-only, no delete verb issued. **The sweep:** a batched `stat -c %n` over Hetzner's port-23 restricted shell (500–600 paths per round trip, stdout carrying only paths that exist) tried **777,600** names of the vendor form `YYYY-MM-DDTHH-MM-SS` across nine full days at second granularity, and 126 alternative shapes — **zero resolved.** **The control is what makes the zero mean anything:** the identical 600-name batch with one real path appended returned it in 6 of 6 batches. **The structural cause:** `/home` (the customer data, `u629488-sub3`) is **st_dev 0,82**; `/.zfs/snapshot` is **st_dev 0,276**; `/home/.zfs` does not exist. A snapshot under `/.zfs/snapshot` belongs to a different dataset than the one holding `felhom-repo`. **What still stands:** clause (a) — the box can delete its live repository but cannot WRITE into `/.zfs/snapshot` — is unchanged and re-confirmed. **What is now open again:** the only routes to the older copy are a panel rollback of the WHOLE Storage Box (deletes newer snapshots, hits every customer on it) and the provider API (fenced by §11-D, and `hub/internal/hetznerapi/hetznerapi.go` has **no snapshot method at all**, so it needs new code regardless). **NOT ESTABLISHED, and it is the question that decides whether per-file recovery exists for anyone:** whether the MAIN account can see the snapshots. No main-account credential exists in this project. **The ranking is Viktor's; I am not re-ranking R-95 — but the argument that moved it down is the argument this drill removed.** `audits/evidence-drill-r95-recovery-2026-09-01/` **BLOCKED-ON-PROVIDER 2026-09-01.** The one question that can move this is drafted and ready to send: **Question 1** of `documentation/runbooks/provider-questions-2026-09-01.md` — *can the MAIN account retrieve individual files from a snapshot, without a whole-box restore?* **Neither answer leaves this row where it is:** "yes" makes per-file recovery real but permanently operator-only, and closes this; "no, full restore only" means the snapshots do not bound a single customer's exposure at all, because using them costs every other customer on the box their newer snapshots — **and R-95 becomes urgent.** Nothing here can progress without it, and it is not CC's to send (§11-D). Tracked by a dated check in the DUE-CHECKS block. **⚠ A TRAP ON THE WAY TO THIS ANSWER, found by the operator 2026-09-01 and armed against in the questions file: HETZNER HAS TWO SIMILARLY-NAMED PRODUCTS WHOSE DOCS SAY OPPOSITE THINGS.** **Storage SHARE** (a managed Nextcloud, NOT us) documents *"Currently, we only support restores for the full backup ZFS snapshot to a specific point in time"* (`docs.hetzner.com/storage/storage-share/faq/backup-snapshot/`). **Storage BOX** (ours) documents the opposite — *"You can download individual files or entire directories as usual"* (`docs.hetzner.com/storage/storage-box/snapshots/`). **A web search for the obvious phrasing surfaces the SHARE page, and it reads like a definitive NO.** Tell them apart by the giveaways: the Share page talks about *Nextcloud's data cache*, a *database dump* and the *konsoleH* interface, and never mentions Storage Box. **Why this is filed and not left as trivia: a support agent could answer Question 1 from the wrong page, and that answer would push this row to the top of the register for no reason.** Question 1 now names the product, quotes the Storage Box line, and says explicitly that we know what the Share FAQ says. **If an answer cites full-snapshot-only, check WHICH PRODUCT it is about before acting on it.** **Nothing in this repository ever leaned on the Share claim** — verified by grep at the time; the only vendor line we cite is the Storage Box one. **RE-DATED 2026-09-15 → 2026-09-22 (due-checks gate):** the operator mailbox read through the Gmail connector (`(from:hetzner OR subject:hetzner OR "Storage Box") after:2026/09/01`) holds no Hetzner reply — one match, our own `offsite_snapshots_dropped` alarm. Whether the two tickets were ever sent is not visible to CC; if they were not, sending them is the operator's act. **-- ANSWERED. Hetzner replied on ticket #2026090103040671; the operator supplied the thread on 2026-09-22 after this session searched for it and WRONGLY reported no answer (see the instrument note below and R-628).** **Q1 - can the MAIN account read individual files and directories out of a specific snapshot, without restoring the whole box?** Hetzner: *"With the main account it should be possible to download files and directories from the snapshot as it is stated in the documentation"*, citing `docs.hetzner.com/storage/storage-box/snapshots#access-to-snapshots`; and separately *"A restore of a snapshot will revert the Storage Box completely to the state it was in when the snapshot was taken."* **So the Storage BOX documentation governs, not the Storage Share FAQ** - which is exactly the trap `provider-questions-2026-09-01.md` warned the reader about, and the answer came back on the right side of it. **File-level snapshot access is a MAIN-ACCOUNT capability; the sub-account cannot do it**, which matches R-433's own measured finding (777,600 names swept from a sub-account, zero hits). **HEDGE, kept because it is in the reply: "should be possible" is not "is", and nobody has yet read a file out of a snapshot from the main account on `u629488`. That is now the measurement this row needs, and it is cheap.** **Q2 - is `--append-only` enforced by Hetzner, or taken from what the client sends?** Hetzner: *"You can use --append-only and it would look like this: `command="rclone serve restic --stdio --append-only path/to/repo" `"*, with `fluix.one/blog/hetzner-restic-append-only/`. **So it is enforced by US, by pinning a FORCED COMMAND to the SSH key in the Storage Box's `authorized_keys`** - not by Hetzner globally, and not by anything the client asks for. A key without that prefix still gets a deleting server. **This is the answer R-95 has been blocked on** and it says append-only IS expressible on a Storage Box, which R-95 records as impossible over SFTP - see R-95 and R-436. **NOT YET MEASURED, and that distinction is the whole of what this row is worth: this is a support answer, not a proof.** What is owed is a real forced-command key on `u629488`, a restic `forget --prune` through it that is REFUSED, and a backup through it that still succeeds. | **BLOCKED** — **BLOCKED-ON-PROVIDER — Question 1 of `runbooks/provider-questions-2026-09-01.md`** | — | — | CC + operator | | **R-540** | Backup & restore | P3 | **[P3-LOW] The hub knows exactly ONE off-site pool box, so there is no rule for what happens when it fills.** Read from source 2026-09-16 while making off-site the default: `HETZNER_POOL_BOX_ID` is a single value, and every shared customer becomes a sub-account on that box. With off-site now ON for every new customer (hub v0.116.0) the box fills faster, and the fill warning (80%/90% of the box, `monitor/offsite.go`) tells the operator it is filling but nothing says which box a new customer should land on. **Needs a selection rule** (least-full, or explicit per-customer), not a bigger box. No customer is at risk today: the pool box read 0.3% full (2.7 GB of 1 TB), Σ shared quota 150 GB, oversub 0.15x. | **READY — rank P3-LOW; owner: CC (hub)** | — | — | CC | | **R-545** | Backup & restore | P3 | **[P3-LOW] There is no product action that un-configures an off-site target — only one that re-starts an ORPHANED repository.** FOUND 2026-09-16 while exercising R-543's paused state on the scratch guest: `POST /backup/offbox/config` configures a target and can disable it (`enabled` unchecked), but nothing removes it. `POST /backup/offbox/reset` refuses unless `OffboxOrphaned()` is true („Az offsite tároló nincs elárvult állapotban."), and it means „start a new remote backup, set the old history aside" — not „forget this destination". So a household that sets up the wrong NAS, or a box being handed to someone else, keeps the host, user, path, ssh key and minted repo password on disk with no route to clear them; a disabled target still holds its secrets in `data/offbox/`. **Why P3 and not higher:** a disabled target runs nothing and the secrets are 0600 on the box's own disk, so nothing leaks and no copy is lost. **Fix shape:** a „Távoli cél törlése" action beside the config form that clears the target and shreds `data/offbox/`, REFUSING while the hub holds a sealed package for this box (the R-241 rule — dropping the key would orphan the history that package protects). Teardown for this session's proof had to clear it out-of-band for exactly this reason, which is the measurement. | **READY — rank P3-LOW; owner: CC** | — | — | CC | @@ -269,13 +268,15 @@ stopping line that lies. | **R-368** | Storage & devices | P4 | **The storage default DOES apply at deploy time — the earlier claim that it never does was wrong, and the residual defect is smaller and different.** R-352 and `SPEC-app-data-placement-2026-08-21.md` §2.2 stated *"the deploy route never reads it"*, from `grep -nE 'GetDefaultStoragePath|primaryHDDPath|IsDefault' internal/stacks/deploy.go internal/stacks/manager.go` → nothing. **That grep searched Go files only and never the templates.** `internal/web/templates/deploy.html:612` reads `.IsDefault` directly off each `DeployStoragePath` (which embeds `settings.StoragePath`, `web/handlers.go:89-99`) and **pre-selects the default drive for a new deploy**: `{{else if and .IsDefault (not .NotAllowed)}}selected{{end}}`. So `// new apps use this by default` (`settings.go:453`) is **IMPRECISE ABOUT THE MECHANISM, NOT FALSE** — nobody calls `GetDefaultStoragePath()` on that route, but the value is honoured. The customer-facing label promises exactly this and no more: **„Legyen alapértelmezett új telepítéseknél"** (`storage.html:469`). **THE RESIDUAL, and it is the whole finding:** the default lives in the TEMPLATE, not in the server. `POST /api/stacks//deploy` accepts `values` verbatim; omit `HDD_PATH` and `withPathVars` (`stacks/deploy.go:584`) receives `""` and no default is applied. **That is why the invariant has no test — there is nothing server-side to test.** | **OPEN — LOW** | corrects R-352(2); supersedes SPEC §2.2 | Either move the default into the server so the API and the form agree and a test can pin it, or reword the comment to say the template owns it. Do not "fix" the behaviour: it is correct on the path customers use. | CC | | **R-568** | Storage & devices | P4 | **[P3-LOW] The dashboard's drive-health rows swap order between visits — the same two disks, listed in a different order a minute apart.** MEASURED 2026-09-17 on demo-hp 9201 during slice 1 release C's live proof: `/dashboard` fetched on 0.249.0 listed „KXG50PNV1T02 NVMe TOSHIBA 1024GB” then „SanDisk X600 M.2 2280 SATA 128GB”; fetched on 0.250.0 a minute later, the reverse (`audits/i18n-slice1-2026-09-17/C/live/hu-before-vs-after.txt`). `diskHealthRows` (`disk_health.go` L135–141) keeps the agent's response order and does not sort; the agent's order is therefore not stable. Cosmetic, but a household that reads „the second disk” finds a different one. **Fix shape:** sort the rows controller-side by a durable key (device path or serial), with a test that feeds two orders and expects one. | **READY - rank P3-LOW; owner: CC** **Re-ranked 2026-10-03: P3->P4: cosmetic.** | — | — | CC | -## Security & access — 29 rows (P2 3, P3 23, P4 3) +## Security & access — 31 rows (P2 5, P3 23, P4 3) | ID | Category | Sev | What | State | Blocked on | Next action | Owner | |---|---|---|---|---|---|---|---| | **R-133** | Security & access | P2 | **The vaulted break-glass console credential is PLAINTEXT AT REST — every hub DB backup is a fleet-wide console-credential dump.** `host_recovery.secret` holds each managed box's `root@pam` password verbatim, so any copy of the SQLite DB (Longhorn snapshot, PBS backup of the hub PVC, a hand-taken copy during a diagnosis) carries root console access to every Felhom host in one file | **READY (M) — NEW 2026-07-31** | — | **The deferred leg of hub v0.84.0** (Console access card), filed separately because v0.84.0 changed only WHO can retrieve the secret, never how it is stored. v0.84.0 makes it more worth doing, not more broken: retrieval now rides the hub SESSION, so the DB and the login password are jointly the whole protection (ruling **S-4**, `CONTEXT.md`). Fix shape: **envelope-encrypt the `host_recovery.secret` column under a KEK held outside the DB** — the hub already proves it can hold something it cannot itself read (escrow blobs), and that contrast is the argument. Two constraints the design must respect: the credential must stay retrievable **when the box is unreachable** (that is the whole point of break-glass), so the KEK cannot live on the box or depend on the agent; and the global-key API path must keep working with the hub UI down. Would flip the capability-map row **"Break-glass management-plane recovery"**, which today reads IMPLEMENTED with this as its caveat | CC | | **R-135** | Security & access | P2 | **`validateCSRF` returns TRUE when there is no session cookie** (`hub/internal/web/server.go:678-683`) — measured live: `POST` with Basic auth and no cookie goes straight past the CSRF gate (404, not 403), while the same POST with a cookie and no token is 403 | READY (S) — **security** | — | Browsers cache HTTP Basic credentials per origin and resend them automatically on cross-origin requests, and `SameSite` does not govern the `Authorization` header. So if the operator has ever Basic-authed to the hub in a browser, any attacker page can POST to every mutating route. Latent on the condition, not guaranteed absent. Fix = require the token whenever the request is not provably programmatic, or drop browser-usable Basic auth. Same audit §4.3 | CC | | **R-777** | Security & access | P2 | **[P2-MEDIUM] Emby and Jellyfin treat every internet visitor as being on the LAN — users with "remote access" off can sign in from the internet, IP filters and remote limits are skipped.** READ in source (`audits/visitors-2026-10-01/A/sweep/sweep-1.md`, `sweep-2.md`), not measured live: Jellyfin with `KnownProxies` empty uses the TCP peer (traefik, private) → "LAN"; Emby reads the leftmost XFF (its chain is now removed by the R-753 reset, so it sees cloudflared's private address → "LAN", as before). True before R-753 too; R-753 neither caused nor fixed it. **Needs:** measure on 9202 (a user with remote access off, through the simulated tunnel); Jellyfin: `KnownProxies` `172.16.0.0/12` in `network.xml` (no env — an `after_install` or a seed file); Emby: no setting fixes it (its `LocalNetworkSubnets` still counts private ranges) — a page sentence or a decision. | **READY — rank P2-MEDIUM; owner: CC (measure), operator (Emby route)** | — | — | CC + operator | +| **R-820** | Security & access | P2 | **A box can obtain its own off-site sub-account password at will, and that password removes any append-only lock.** MEASURED 2026-10-03 on `u629488-sub4`: the password logs in on port 22 AND port 23, and through it `.ssh/authorized_keys` was read and rewritten (the test keys were installed that way). READ from source: the box declares `needs_credential` in two reports → the hub's off-site self-heal re-arms the STORED value (`RestageOneTimeSecret`) or re-issues it at the provider → the box consumes it with its own API key (`hub/internal/offsiteheal/reconciler.go`, `controller/internal/offsiteapply/seams.go`). **Not P1:** today the box's own key can already delete (R-95, P2), so this adds no harm TODAY; it is the bypass the moment a pinned key ships, so it is a precondition on R-95. `audits/offsite-append-only-2026-10-03/live/B1-B2-ports-password-forward.txt` | **OPEN — precondition on any R-95 build** | — | The hub becomes the key registrar (box sends its PUBLIC key, the hub writes the pinned line); the box never receives the password; a daily hub check alarms on any `authorized_keys` line without the pin (DESIGN.md §2) | CC + operator | +| **R-821** | Security & access | P2 | **The hub database holds every customer's off-site sub-account password in the clear, indefinitely — whoever reads it can delete every household's off-site history.** `one_time_secrets.value` is never cleared after a consume (deliberate, so the self-heal can re-arm it — `hub/internal/store/store.go` `RestageOneTimeSecret`). Confirmed 2026-10-03: the stored value for `tester-1` (re-issued 2026-09-29) still logged in to its sub-account. The same holds for any copy of the hub DB. **Not P1:** it needs the hub or a copy of its DB, which is operator-tier; it is the central version of R-95. | **OPEN — needs an operator ruling on who holds the password** | — | With R-820's registrar the hub still needs the password; options: keep it but encrypted at rest, or rotate it after each use and keep only the current one; decide with the R-95 ruling | CC + operator | | **R-126** | Security & access | P3 | **A `.fab` bundle — plaintext secrets, OPTIONAL password — can be exported ONTO a NAS.** `storageDriveList()` (`internal/web/handler_export.go`) does not filter network paths | READY (S) | — | Split out of R-108, which closed without it: this is an explicit customer-chosen **export destination**, not a browsing surface reaching a backup tree, so it was never part of D5's precondition (`07` §7.3 records that reasoning). Was recorded inside R-108's row as its "second effect, independent of D5"; promoted to its own row so it does not vanish with R-108's closure. Fix = filter network paths out of the export destination list, or require the bundle password when the destination is a share | CC | | **R-132** | Security & access | P3 | **`curl -w '%{redirect_url}'` reconstructs the request URL WITH its basic-auth credential** — so a `-u ":$HUB_PW"` call that never put the password in a URL still printed it | **WAITING-ON-OPERATOR** — **ACTION: rotate `HUB_PW`** **Folded R-580 2026-10-03** (the same `curl -w %{redirect_url}` credential echo, seen again 2026-09-18). | — | Happened on 2026-07-31 while red-proofing the R-120 gate: the hub operator password was written to the session transcript by the write-out format, not by the request. `-u` is safe; the *reporting* was not. Rule: read the redirect from `-D -` and grep `^Location:`, never `%{redirect_url}`, on any authenticated call. Rotate the hub password (`/configuration` → Login password; ConfigMap `auth.password_hash` is the reset path) and update `~/.config/credentials` | Viktor | | **R-136** | Security & access | P3 | **Rename `hub_session` → `__Host-hub_session`** — makes cookie tossing structurally impossible | READY (XS, one line) | — | Verified on the live production response that all three prefix preconditions already hold: `Path=/`, `Secure`, no `Domain`. **Caveat for the ticket:** browsers reject a `__Host-` cookie without `Secure`, and `isSecure` is conditional on `r.TLS`/`X-Forwarded-Proto`, so plain-HTTP *browser* access to the hub would stop working (non-browser access uses Basic auth, unaffected). Tested consequence: `r.Cookie` returns the FIRST match and never tries the others, so a tossed cookie wins outright. Same audit §4.1-4.2 | CC | @@ -492,5 +493,4 @@ stopping line that lies. the R-row. Duplicating them here would create the second source this design avoids. --> | item | due (UTC) | what to measure | |---|---|---| -| R-436 | 2026-10-06 | **The question is ANSWERED; what is owed is the MEASUREMENT.** Hetzner (ticket #2026090103040671, recorded in R-433) says `--append-only` is enforced by pinning `command="rclone serve restic --stdio --append-only path/to/repo"` to the key in the Storage Box's `authorized_keys`. Measure it on `u629488`: a forced-command key, a restic backup through it that SUCCEEDS, and a `forget --prune` through it that is REFUSED — quote the refusal. Record in R-436 and R-95. A support answer is not a property of our backups. |