Spike, no product change. Venue u629488-sub4 (tester-1, operator ruling); scratch repo removed, authorized_keys restored byte-identical. Closed R-436 (due-check cleared), R-430. Opened R-820, R-821, R-822. R-95 and R-342 updated. 07 §D [FACT] block. STATUS: two operator decisions. Register 326 -> 327. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
10 KiB
Off-site append-only — who deletes old backups when the box cannot? (PROPOSAL, 2026-10-03)
Status: a PROPOSAL for the operator to rule on. Nothing here is built. Nothing here is [DESIGN].
Evidence: live/ (the provider, u629488-sub4) and lab/ (OpenSSH + rclone v1.75.1 on DooPlex).
Baselines: felhom.eu f4c5466, controller 0945332 (v0.288.0, restic 0.14.0 Debian 0.14.0-1+b5).
1. What was measured (the facts this design stands on)
| # | Fact | Where |
|---|---|---|
| F1 | A key whose authorized_keys line is command="rclone serve restic --stdio --append-only <dir>",restrict can init, backup, snapshots, restore (bytes identical) and check. |
live/E1-E3, live/E4-E6, live/C2-A5 |
| F2 | Through that key forget <id> --prune, forget --keep-last 1 and a real prune are REFUSED: blob not removed, server response: 403 Forbidden (403), rc=1, snapshot count unchanged. Each refusal costs ~45–48 s of restic retries. A refused prune has already written a new index first (check: duplicate index, non-critical). |
live/E4-E6 |
| F3 | Control: the same forget through a key WITHOUT the forced command deletes (1/1 files deleted). |
live/E4-E6 (C1) |
| F4 | The client's requested path and flags are ignored: the pinned directory is served even when the client names another path, and the client never asked for --append-only (restic sends serve restic --stdio --b2-hard-delete). |
live/C2-A5, live/B2 |
| F5 | The forced key cannot get a shell, sftp, scp, rsync or a port forward (administratively prohibited); a command such as rm -rf <dir> just starts the forced rclone. |
live/B2, live/B1-B2 |
| F6 | Port 22 accepts no OpenSSH-format key at all (both test keys refused there); keys work on port 23 only — the port the product uses (hub/internal/offsite/offsite.go sftpPort = 23). The PASSWORD logs in on BOTH ports, and through it .ssh/authorized_keys was read and REWRITTEN (that is how the test keys were installed). |
live/B1-B2, this session |
| F7 | Crash locks: a killed backup leaves a lock. The next backup is not blocked; an exclusive operation (check — the weekly integrity job) is. Plain unlock prints successfully removed locks and removes nothing (the lock is not yet stale — by design, R-430 reproduced); unlock --remove-all does remove it — the append-only server allows lock deletion. |
live/C2-A5, lab/A5 |
| F8 | A crashed upload leaves unreferenced packs (lab: 17) that only a deleting key can clean. | lab/A5 |
| F9 | The append-only server refuses delete/overwrite of existing files (403) and path escapes (../, %2e%2e, ..%2f → 400), but ACCEPTS new files of any name inside the repo, including a new keys/ entry (quota can be filled). |
lab/B-rclone-server-surface (lab rclone; the provider's rclone version cannot be read — rclone version is not a shell command there) |
| F10 | Retention poisoning. Through the add-only key, 13 future-dated empty snapshots with the same host+tag make the box's exact policy (--group-by host,tags --keep-daily 7 --keep-weekly 4 --keep-monthly 6) select all 3 real snapshots for removal. |
lab/C3-retention-poisoning |
| F11 | The provider's rclone creates ~/.config/rclone/ (empty) in the sub-account home on first use. |
live/TEARDOWN |
From the code (read, not run):
| # | Fact | Where |
|---|---|---|
| K1 | The box can obtain its sub-account password at will: it declares needs_credential in two reports, the hub's off-site self-heal re-arms the stored value (RestageOneTimeSecret; the value is never cleared after a consume) or escalates to a provider re-issue, and the box consumes it with its own API key. |
hub/internal/offsiteheal/reconciler.go, hub/internal/store/store.go ConsumeOneTimeSecret/RestageOneTimeSecret, controller/internal/offsiteapply/seams.go |
| K2 | The hub DB holds every sub-account password in the clear, indefinitely (one_time_secrets.value). |
hub/internal/store/store.go |
| K3 | Four box features need a normal (deleting) login, not just the two forget sites: forget --prune after a run (offbox.go ~1422), over quota (offbox.go ~1793), the orphan move-aside (mv, offbox.go resetOrphanedRepo), and the customer-chosen abandonment (rm -rf, offbox_abandon.go AbandonSweep). |
controller 0945332 |
| K4 | restic has no prune-only credential: every key unlocks the same master key, so whoever prunes can read every backup in that repo. | restic design (documented, not measured) |
2. The prerequisite every option shares — the lock is worthless until this is closed
F6 + K1: a box that is broken into can get the password, log in on port 23 with it, and rewrite
authorized_keys — removing the forced command from its own key. So before any lock ships:
- The box never receives the sub-account password. The hub becomes the key registrar: the box
generates its key pair and sends the PUBLIC key to the hub; the hub (which already holds the
password, K2) writes the forced line into
authorized_keysover SFTP port 23. The off-site self-heal re-arms the HUB's install job, not a password for the box. - The hub checks the file. Each day the hub reads
authorized_keysand alarms if any line lacks the forced prefix (outside an open clean-up window). This is also the answer to "an operator adds a key later": a key without the prefix re-opens deletion, so a check — not a rule — catches it. - Optional hardening, operator's call: rotate the sub-account password after each hub use and keep it only in the hub (K2 stays true either way: the hub can already delete every household's off-site history today — new row).
3. The options
Option 1 — a clean-up window opened by the hub (the box keeps pruning its own repo)
The box normally holds only its add-only key. Once a week the hub adds a second, deleting key line for the same box key pair (or a second key the box holds) for ~15 minutes; the box runs its retention; the hub removes the line and checks the file.
- Custody: unchanged. The repository password never leaves the box (07 §8a stays true).
- What a compromised box can do: delete during a window. Detection backstop: R-431 (a fall of more than half), plus a new per-window check — the hub compares the snapshot count before and after and alarms on a fall larger than the retention could cause.
- Poisoning (F10) must be guarded even here, because an attacker who was on the box and left can
plant future snapshots that the next honest window then obeys: before
forget, refuse when any snapshot is dated in the future or newer than the newest the hub has seen reported; run--dry-runfirst and abort if it would remove more snapshots than the policy can remove in a week. - Build cost: hub: key registrar + window open/close +
authorized_keyscheck (SFTP with the password it holds — no provider API, no main account). Controller: switch the transport fromsftp:torclone:with-o rclone.program="ssh -p 23 … -i <key> … rclone"(restic 0.14.0 is enough, F1; rclone is NOT needed in the image); moveforget --prune(both sites, together — R-191) into the window; move the move-aside and the abandonment to the hub (K3). - Money: none.
Option 2 — clean-up runs off the box
A Felhom-side worker prunes each repo with a deleting key.
- Custody: CHANGES. restic has no prune-only key (K4): the worker must hold every repository password in a form it can use unattended, and could read every household's backups. That changes a promise to the customer (07 §8a) — not CC's to change.
- Poisoning (F10): same guard needed.
- Build cost: a new always-on worker in the recovery path, its own credentials store, its own failure alarms — plus everything in §2.
Option 3 — never delete; let the quota grow
- Money: small. The pool box (BX11, 1 TB) was €4.06/month when recorded (
SPIKE-ep0-storagebox- 2026-07-09; Hetzner raised prices in April 2026 — re-check) ≈ €0.004 per GB per month. A household adding 10 GB a year unpruned costs ≈ €0.50 a year. The demo boxes cannot calibrate this (0.5 GB and 0.0 GB used today), so 10 GB/year is an assumption, stated as one. - The real cost is the quota (50–100 GB): at the quota
offbox_fit.gorefuses new pushes and the over-quotaforget(K3) would itself be refused, so the household's off-site copy STOPS. Crash leftovers (F8) and refused-prune index files (F2) also accumulate. - Good as an interim, not an end state.
4. Recommendation
Option 1, with Option 3 as the interim while it is built. Order:
- §2 first (key registrar + the
authorized_keyscheck) — without it nothing below protects anything. - Switch every box to the forced key with no box-side retention (Option 3 interim). Quotas today are 50–100 GB against ≤0.5 GB used, so this costs nothing for months.
- Then the weekly window (Option 1) with the poisoning guard.
Why not Option 2: it trades a box-level risk for a custody change that lets one Felhom machine read every household's backups, and it adds an always-on service to the recovery path.
5. Migration, rotation, restore
- Existing repos (both demo boxes, any household): only the
authorized_keysline and the box's transport change; the repository bytes are not touched. Not measured: reading a repo written oversftp:throughrclone:— restic's on-disk layout is backend-independent, so it is expected to work; measure on the scratch account before any household. History is kept. - Key rotation: the hub writes the new forced line, the box switches, the hub removes the old line — the same registrar path.
- Restore works through the add-only key (F1) — no second key is needed for the household.
- Locks: the self-heal (
resticStep→unlock --remove-all) keeps working under this transport (F7). R-430's fear does not apply to it. - Failure speed: with today's code a forced key makes each night's
forget --prunefail after ~45 s per refused file and grow the index (F2) — the twoforgetsites must leave the box in the same change that switches the key.