# Off-site append-only — who deletes old backups when the box cannot? (PROPOSAL, 2026-10-03) **Status: a PROPOSAL for the operator to rule on. Nothing here is built. Nothing here is `[DESIGN]`.** Evidence: `live/` (the provider, `u629488-sub4`) and `lab/` (OpenSSH + rclone v1.75.1 on DooPlex). Baselines: felhom.eu `f4c5466`, controller `0945332` (v0.288.0, restic **0.14.0** Debian `0.14.0-1+b5`). ## 1. What was measured (the facts this design stands on) | # | Fact | Where | |---|---|---| | F1 | A key whose `authorized_keys` line is `command="rclone serve restic --stdio --append-only ",restrict` can `init`, `backup`, `snapshots`, `restore` (bytes identical) and `check`. | `live/E1-E3`, `live/E4-E6`, `live/C2-A5` | | F2 | Through that key `forget --prune`, `forget --keep-last 1` and a real `prune` are REFUSED: `blob not removed, server response: 403 Forbidden (403)`, rc=1, snapshot count unchanged. Each refusal costs **~45–48 s** of restic retries. A refused `prune` has already **written a new index** first (`check`: duplicate index, non-critical). | `live/E4-E6` | | F3 | Control: the same `forget` through a key WITHOUT the forced command deletes (1/1 files deleted). | `live/E4-E6` (C1) | | F4 | The client's requested path and flags are ignored: the pinned directory is served even when the client names another path, and the client never asked for `--append-only` (restic sends `serve restic --stdio --b2-hard-delete`). | `live/C2-A5`, `live/B2` | | F5 | The forced key cannot get a shell, `sftp`, `scp`, `rsync` or a port forward (`administratively prohibited`); a command such as `rm -rf ` just starts the forced rclone. | `live/B2`, `live/B1-B2` | | F6 | Port **22 accepts no OpenSSH-format key at all** (both test keys refused there); keys work on port **23** only — the port the product uses (`hub/internal/offsite/offsite.go` `sftpPort = 23`). **The PASSWORD logs in on BOTH ports**, and through it `.ssh/authorized_keys` was read and REWRITTEN (that is how the test keys were installed). | `live/B1-B2`, this session | | F7 | Crash locks: a killed backup leaves a lock. The next **backup is not blocked**; an **exclusive** operation (`check` — the weekly integrity job) **is**. Plain `unlock` prints `successfully removed locks` and removes nothing (the lock is not yet stale — by design, R-430 reproduced); `unlock --remove-all` **does remove it** — the append-only server allows lock deletion. | `live/C2-A5`, `lab/A5` | | F8 | A crashed upload leaves unreferenced packs (lab: 17) that only a deleting key can clean. | `lab/A5` | | F9 | The append-only server refuses delete/overwrite of existing files (403) and path escapes (`../`, `%2e%2e`, `..%2f` → 400), but ACCEPTS new files of any name inside the repo, including a new `keys/` entry (quota can be filled). | `lab/B-rclone-server-surface` (lab rclone; the provider's rclone version cannot be read — `rclone version` is not a shell command there) | | F10 | **Retention poisoning.** Through the add-only key, 13 future-dated empty snapshots with the same host+tag make the box's exact policy (`--group-by host,tags --keep-daily 7 --keep-weekly 4 --keep-monthly 6`) select **all 3 real snapshots for removal**. | `lab/C3-retention-poisoning` | | F11 | The provider's rclone creates `~/.config/rclone/` (empty) in the sub-account home on first use. | `live/TEARDOWN` | From the code (read, not run): | # | Fact | Where | |---|---|---| | K1 | The box can obtain its sub-account password **at will**: it declares `needs_credential` in two reports, the hub's off-site self-heal re-arms the **stored** value (`RestageOneTimeSecret`; the value is never cleared after a consume) or escalates to a provider re-issue, and the box consumes it with its own API key. | `hub/internal/offsiteheal/reconciler.go`, `hub/internal/store/store.go` `ConsumeOneTimeSecret`/`RestageOneTimeSecret`, `controller/internal/offsiteapply/seams.go` | | K2 | The hub DB holds every sub-account password in the clear, indefinitely (`one_time_secrets.value`). | `hub/internal/store/store.go` | | K3 | Four box features need a normal (deleting) login, not just the two `forget` sites: `forget --prune` after a run (`offbox.go` ~1422), over quota (`offbox.go` ~1793), the orphan **move-aside** (`mv`, `offbox.go` `resetOrphanedRepo`), and the customer-chosen **abandonment** (`rm -rf`, `offbox_abandon.go` `AbandonSweep`). | controller `0945332` | | K4 | restic has no prune-only credential: every key unlocks the same master key, so whoever prunes can read every backup in that repo. | restic design (documented, not measured) | ## 2. The prerequisite every option shares — the lock is worthless until this is closed **F6 + K1: a box that is broken into can get the password, log in on port 23 with it, and rewrite `authorized_keys` — removing the forced command from its own key.** So before any lock ships: 1. **The box never receives the sub-account password.** The hub becomes the key registrar: the box generates its key pair and sends the PUBLIC key to the hub; the hub (which already holds the password, K2) writes the forced line into `authorized_keys` over SFTP port 23. The off-site self-heal re-arms the HUB's install job, not a password for the box. 2. **The hub checks the file.** Each day the hub reads `authorized_keys` and alarms if any line lacks the forced prefix (outside an open clean-up window). This is also the answer to "an operator adds a key later": a key without the prefix re-opens deletion, so a check — not a rule — catches it. 3. Optional hardening, operator's call: rotate the sub-account password after each hub use and keep it only in the hub (K2 stays true either way: **the hub can already delete every household's off-site history today** — new row). ## 3. The options ### Option 1 — a clean-up window opened by the hub (the box keeps pruning its own repo) The box normally holds only its add-only key. Once a week the hub adds a second, deleting key line for the same box key pair (or a second key the box holds) for ~15 minutes; the box runs its retention; the hub removes the line and checks the file. - **Custody: unchanged.** The repository password never leaves the box (07 §8a stays true). - **What a compromised box can do:** delete during a window. Detection backstop: R-431 (a fall of more than half), plus a new per-window check — the hub compares the snapshot count before and after and alarms on a fall larger than the retention could cause. - **Poisoning (F10) must be guarded even here**, because an attacker who was on the box and left can plant future snapshots that the next honest window then obeys: before `forget`, refuse when any snapshot is dated in the future or newer than the newest the hub has seen reported; run `--dry-run` first and abort if it would remove more snapshots than the policy can remove in a week. - **Build cost:** hub: key registrar + window open/close + `authorized_keys` check (SFTP with the password it holds — no provider API, no main account). Controller: switch the transport from `sftp:` to `rclone:` with `-o rclone.program="ssh -p 23 … -i … rclone"` (restic 0.14.0 is enough, F1; **rclone is NOT needed in the image**); move `forget --prune` (both sites, together — R-191) into the window; move the move-aside and the abandonment to the hub (K3). - **Money:** none. ### Option 2 — clean-up runs off the box A Felhom-side worker prunes each repo with a deleting key. - **Custody: CHANGES.** restic has no prune-only key (K4): the worker must hold every repository password in a form it can use unattended, and could read every household's backups. That changes a promise to the customer (07 §8a) — not CC's to change. - **Poisoning (F10):** same guard needed. - **Build cost:** a new always-on worker in the recovery path, its own credentials store, its own failure alarms — plus everything in §2. ### Option 3 — never delete; let the quota grow - **Money:** small. The pool box (BX11, 1 TB) was €4.06/month when recorded (`SPIKE-ep0-storagebox- 2026-07-09`; Hetzner raised prices in April 2026 — re-check) ≈ **€0.004 per GB per month**. A household adding 10 GB a year unpruned costs ≈ **€0.50 a year**. The demo boxes cannot calibrate this (0.5 GB and 0.0 GB used today), so 10 GB/year is an assumption, stated as one. - **The real cost is the quota (50–100 GB):** at the quota `offbox_fit.go` refuses new pushes and the over-quota `forget` (K3) would itself be refused, so the household's off-site copy STOPS. Crash leftovers (F8) and refused-prune index files (F2) also accumulate. - Good as an **interim**, not an end state. ## 4. Recommendation **Option 1, with Option 3 as the interim while it is built.** Order: 1. §2 first (key registrar + the `authorized_keys` check) — without it nothing below protects anything. 2. Switch every box to the forced key with **no** box-side retention (Option 3 interim). Quotas today are 50–100 GB against ≤0.5 GB used, so this costs nothing for months. 3. Then the weekly window (Option 1) with the poisoning guard. Why not Option 2: it trades a box-level risk for a custody change that lets one Felhom machine read every household's backups, and it adds an always-on service to the recovery path. ## 5. Migration, rotation, restore - **Existing repos (both demo boxes, any household):** only the `authorized_keys` line and the box's transport change; the repository bytes are not touched. **Not measured:** reading a repo written over `sftp:` through `rclone:` — restic's on-disk layout is backend-independent, so it is expected to work; measure on the scratch account before any household. History is kept. - **Key rotation:** the hub writes the new forced line, the box switches, the hub removes the old line — the same registrar path. - **Restore** works through the add-only key (F1) — no second key is needed for the household. - **Locks:** the self-heal (`resticStep` → `unlock --remove-all`) keeps working under this transport (F7). R-430's fear does not apply to it. - **Failure speed:** with today's code a forced key makes each night's `forget --prune` fail after ~45 s per refused file and grow the index (F2) — the two `forget` sites must leave the box in the same change that switches the key.