# Off-site append-only — who deletes old backups when the box cannot? (PROPOSAL, 2026-10-03)
**Status: a PROPOSAL for the operator to rule on. Nothing here is built. Nothing here is `[DESIGN]`.**
Evidence: `live/` (the provider, `u629488-sub4`) and `lab/` (OpenSSH + rclone v1.75.1 on DooPlex).
Baselines: felhom.eu `f4c5466`, controller `0945332` (v0.288.0, restic **0.14.0** Debian `0.14.0-1+b5`).
## 1. What was measured (the facts this design stands on)
| # | Fact | Where |
|---|---|---|
| F1 | A key whose `authorized_keys` line is `command="rclone serve restic --stdio --append-only
",restrict` can `init`, `backup`, `snapshots`, `restore` (bytes identical) and `check`. | `live/E1-E3`, `live/E4-E6`, `live/C2-A5` |
| F2 | Through that key `forget --prune`, `forget --keep-last 1` and a real `prune` are REFUSED: `blob not removed, server response: 403 Forbidden (403)`, rc=1, snapshot count unchanged. Each refusal costs **~45–48 s** of restic retries. A refused `prune` has already **written a new index** first (`check`: duplicate index, non-critical). | `live/E4-E6` |
| F3 | Control: the same `forget` through a key WITHOUT the forced command deletes (1/1 files deleted). | `live/E4-E6` (C1) |
| F4 | The client's requested path and flags are ignored: the pinned directory is served even when the client names another path, and the client never asked for `--append-only` (restic sends `serve restic --stdio --b2-hard-delete`). | `live/C2-A5`, `live/B2` |
| F5 | The forced key cannot get a shell, `sftp`, `scp`, `rsync` or a port forward (`administratively prohibited`); a command such as `rm -rf ` just starts the forced rclone. | `live/B2`, `live/B1-B2` |
| F6 | Port **22 accepts no OpenSSH-format key at all** (both test keys refused there); keys work on port **23** only — the port the product uses (`hub/internal/offsite/offsite.go` `sftpPort = 23`). **The PASSWORD logs in on BOTH ports**, and through it `.ssh/authorized_keys` was read and REWRITTEN (that is how the test keys were installed). | `live/B1-B2`, this session |
| F7 | Crash locks: a killed backup leaves a lock. The next **backup is not blocked**; an **exclusive** operation (`check` — the weekly integrity job) **is**. Plain `unlock` prints `successfully removed locks` and removes nothing (the lock is not yet stale — by design, R-430 reproduced); `unlock --remove-all` **does remove it** — the append-only server allows lock deletion. | `live/C2-A5`, `lab/A5` |
| F8 | A crashed upload leaves unreferenced packs (lab: 17) that only a deleting key can clean. | `lab/A5` |
| F9 | The append-only server refuses delete/overwrite of existing files (403) and path escapes (`../`, `%2e%2e`, `..%2f` → 400), but ACCEPTS new files of any name inside the repo, including a new `keys/` entry (quota can be filled). | `lab/B-rclone-server-surface` (lab rclone; the provider's rclone version cannot be read — `rclone version` is not a shell command there) |
| F10 | **Retention poisoning.** Through the add-only key, 13 future-dated empty snapshots with the same host+tag make the box's exact policy (`--group-by host,tags --keep-daily 7 --keep-weekly 4 --keep-monthly 6`) select **all 3 real snapshots for removal**. | `lab/C3-retention-poisoning` |
| F11 | The provider's rclone creates `~/.config/rclone/` (empty) in the sub-account home on first use. | `live/TEARDOWN` |
From the code (read, not run):
| # | Fact | Where |
|---|---|---|
| K1 | The box can obtain its sub-account password **at will**: it declares `needs_credential` in two reports, the hub's off-site self-heal re-arms the **stored** value (`RestageOneTimeSecret`; the value is never cleared after a consume) or escalates to a provider re-issue, and the box consumes it with its own API key. | `hub/internal/offsiteheal/reconciler.go`, `hub/internal/store/store.go` `ConsumeOneTimeSecret`/`RestageOneTimeSecret`, `controller/internal/offsiteapply/seams.go` |
| K2 | The hub DB holds every sub-account password in the clear, indefinitely (`one_time_secrets.value`). | `hub/internal/store/store.go` |
| K3 | Four box features need a normal (deleting) login, not just the two `forget` sites: `forget --prune` after a run (`offbox.go` ~1422), over quota (`offbox.go` ~1793), the orphan **move-aside** (`mv`, `offbox.go` `resetOrphanedRepo`), and the customer-chosen **abandonment** (`rm -rf`, `offbox_abandon.go` `AbandonSweep`). | controller `0945332` |
| K4 | restic has no prune-only credential: every key unlocks the same master key, so whoever prunes can read every backup in that repo. | restic design (documented, not measured) |
## 2. The prerequisite every option shares — the lock is worthless until this is closed
**F6 + K1: a box that is broken into can get the password, log in on port 23 with it, and rewrite
`authorized_keys` — removing the forced command from its own key.** So before any lock ships:
1. **The box never receives the sub-account password.** The hub becomes the key registrar: the box
generates its key pair and sends the PUBLIC key to the hub; the hub (which already holds the
password, K2) writes the forced line into `authorized_keys` over SFTP port 23. The off-site
self-heal re-arms the HUB's install job, not a password for the box.
2. **The hub checks the file.** Each day the hub reads `authorized_keys` and alarms if any line lacks
the forced prefix (outside an open clean-up window). This is also the answer to "an operator adds a
key later": a key without the prefix re-opens deletion, so a check — not a rule — catches it.
3. Optional hardening, operator's call: rotate the sub-account password after each hub use and keep it
only in the hub (K2 stays true either way: **the hub can already delete every household's off-site
history today** — new row).
## 3. The options
### Option 1 — a clean-up window opened by the hub (the box keeps pruning its own repo)
The box normally holds only its add-only key. Once a week the hub adds a second, deleting key line for
the same box key pair (or a second key the box holds) for ~15 minutes; the box runs its retention; the
hub removes the line and checks the file.
- **Custody: unchanged.** The repository password never leaves the box (07 §8a stays true).
- **What a compromised box can do:** delete during a window. Detection backstop: R-431 (a fall of more
than half), plus a new per-window check — the hub compares the snapshot count before and after and
alarms on a fall larger than the retention could cause.
- **Poisoning (F10) must be guarded even here**, because an attacker who was on the box and left can
plant future snapshots that the next honest window then obeys: before `forget`, refuse when any
snapshot is dated in the future or newer than the newest the hub has seen reported; run `--dry-run`
first and abort if it would remove more snapshots than the policy can remove in a week.
- **Build cost:** hub: key registrar + window open/close + `authorized_keys` check (SFTP with the
password it holds — no provider API, no main account). Controller: switch the transport from `sftp:`
to `rclone:` with `-o rclone.program="ssh -p 23 … -i … rclone"` (restic 0.14.0 is enough, F1;
**rclone is NOT needed in the image**); move `forget --prune` (both sites, together — R-191) into
the window; move the move-aside and the abandonment to the hub (K3).
- **Money:** none.
### Option 2 — clean-up runs off the box
A Felhom-side worker prunes each repo with a deleting key.
- **Custody: CHANGES.** restic has no prune-only key (K4): the worker must hold every repository
password in a form it can use unattended, and could read every household's backups. That changes a
promise to the customer (07 §8a) — not CC's to change.
- **Poisoning (F10):** same guard needed.
- **Build cost:** a new always-on worker in the recovery path, its own credentials store, its own
failure alarms — plus everything in §2.
### Option 3 — never delete; let the quota grow
- **Money:** small. The pool box (BX11, 1 TB) was €4.06/month when recorded (`SPIKE-ep0-storagebox-
2026-07-09`; Hetzner raised prices in April 2026 — re-check) ≈ **€0.004 per GB per month**. A
household adding 10 GB a year unpruned costs ≈ **€0.50 a year**. The demo boxes cannot calibrate
this (0.5 GB and 0.0 GB used today), so 10 GB/year is an assumption, stated as one.
- **The real cost is the quota (50–100 GB):** at the quota `offbox_fit.go` refuses new pushes and the
over-quota `forget` (K3) would itself be refused, so the household's off-site copy STOPS. Crash
leftovers (F8) and refused-prune index files (F2) also accumulate.
- Good as an **interim**, not an end state.
## 4. Recommendation
**Option 1, with Option 3 as the interim while it is built.** Order:
1. §2 first (key registrar + the `authorized_keys` check) — without it nothing below protects anything.
2. Switch every box to the forced key with **no** box-side retention (Option 3 interim). Quotas today
are 50–100 GB against ≤0.5 GB used, so this costs nothing for months.
3. Then the weekly window (Option 1) with the poisoning guard.
Why not Option 2: it trades a box-level risk for a custody change that lets one Felhom machine read
every household's backups, and it adds an always-on service to the recovery path.
## 5. Migration, rotation, restore
- **Existing repos (both demo boxes, any household):** only the `authorized_keys` line and the
box's transport change; the repository bytes are not touched. **Not measured:** reading a repo
written over `sftp:` through `rclone:` — restic's on-disk layout is backend-independent, so it is
expected to work; measure on the scratch account before any household. History is kept.
- **Key rotation:** the hub writes the new forced line, the box switches, the hub removes the old line —
the same registrar path.
- **Restore** works through the add-only key (F1) — no second key is needed for the household.
- **Locks:** the self-heal (`resticStep` → `unlock --remove-all`) keeps working under this transport
(F7). R-430's fear does not apply to it.
- **Failure speed:** with today's code a forced key makes each night's `forget --prune` fail after
~45 s per refused file and grow the index (F2) — the two `forget` sites must leave the box in the
same change that switches the key.