Files
felhom.eu/documentation/audits/offsite-append-only-2026-10-03/DESIGN.md
T
admin 9268d9933b
gates / gates (push) Successful in 29s
R-436 measured on the provider: append-only forced key HOLDS (403 on every delete), but the sub-account password defeats it (R-820); design proposal + ep0 options
Spike, no product change. Venue u629488-sub4 (tester-1, operator ruling); scratch repo removed,
authorized_keys restored byte-identical. Closed R-436 (due-check cleared), R-430. Opened R-820,
R-821, R-822. R-95 and R-342 updated. 07 §D [FACT] block. STATUS: two operator decisions.
Register 326 -> 327.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-10-03 13:40:35 +02:00

10 KiB
Raw Blame History

Off-site append-only — who deletes old backups when the box cannot? (PROPOSAL, 2026-10-03)

Status: a PROPOSAL for the operator to rule on. Nothing here is built. Nothing here is [DESIGN]. Evidence: live/ (the provider, u629488-sub4) and lab/ (OpenSSH + rclone v1.75.1 on DooPlex). Baselines: felhom.eu f4c5466, controller 0945332 (v0.288.0, restic 0.14.0 Debian 0.14.0-1+b5).

1. What was measured (the facts this design stands on)

# Fact Where
F1 A key whose authorized_keys line is command="rclone serve restic --stdio --append-only <dir>",restrict can init, backup, snapshots, restore (bytes identical) and check. live/E1-E3, live/E4-E6, live/C2-A5
F2 Through that key forget <id> --prune, forget --keep-last 1 and a real prune are REFUSED: blob not removed, server response: 403 Forbidden (403), rc=1, snapshot count unchanged. Each refusal costs ~45–48 s of restic retries. A refused prune has already written a new index first (check: duplicate index, non-critical). live/E4-E6
F3 Control: the same forget through a key WITHOUT the forced command deletes (1/1 files deleted). live/E4-E6 (C1)
F4 The client's requested path and flags are ignored: the pinned directory is served even when the client names another path, and the client never asked for --append-only (restic sends serve restic --stdio --b2-hard-delete). live/C2-A5, live/B2
F5 The forced key cannot get a shell, sftp, scp, rsync or a port forward (administratively prohibited); a command such as rm -rf <dir> just starts the forced rclone. live/B2, live/B1-B2
F6 Port 22 accepts no OpenSSH-format key at all (both test keys refused there); keys work on port 23 only — the port the product uses (hub/internal/offsite/offsite.go sftpPort = 23). The PASSWORD logs in on BOTH ports, and through it .ssh/authorized_keys was read and REWRITTEN (that is how the test keys were installed). live/B1-B2, this session
F7 Crash locks: a killed backup leaves a lock. The next backup is not blocked; an exclusive operation (check — the weekly integrity job) is. Plain unlock prints successfully removed locks and removes nothing (the lock is not yet stale — by design, R-430 reproduced); unlock --remove-all does remove it — the append-only server allows lock deletion. live/C2-A5, lab/A5
F8 A crashed upload leaves unreferenced packs (lab: 17) that only a deleting key can clean. lab/A5
F9 The append-only server refuses delete/overwrite of existing files (403) and path escapes (../, %2e%2e, ..%2f → 400), but ACCEPTS new files of any name inside the repo, including a new keys/ entry (quota can be filled). lab/B-rclone-server-surface (lab rclone; the provider's rclone version cannot be read — rclone version is not a shell command there)
F10 Retention poisoning. Through the add-only key, 13 future-dated empty snapshots with the same host+tag make the box's exact policy (--group-by host,tags --keep-daily 7 --keep-weekly 4 --keep-monthly 6) select all 3 real snapshots for removal. lab/C3-retention-poisoning
F11 The provider's rclone creates ~/.config/rclone/ (empty) in the sub-account home on first use. live/TEARDOWN

From the code (read, not run):

# Fact Where
K1 The box can obtain its sub-account password at will: it declares needs_credential in two reports, the hub's off-site self-heal re-arms the stored value (RestageOneTimeSecret; the value is never cleared after a consume) or escalates to a provider re-issue, and the box consumes it with its own API key. hub/internal/offsiteheal/reconciler.go, hub/internal/store/store.go ConsumeOneTimeSecret/RestageOneTimeSecret, controller/internal/offsiteapply/seams.go
K2 The hub DB holds every sub-account password in the clear, indefinitely (one_time_secrets.value). hub/internal/store/store.go
K3 Four box features need a normal (deleting) login, not just the two forget sites: forget --prune after a run (offbox.go ~1422), over quota (offbox.go ~1793), the orphan move-aside (mv, offbox.go resetOrphanedRepo), and the customer-chosen abandonment (rm -rf, offbox_abandon.go AbandonSweep). controller 0945332
K4 restic has no prune-only credential: every key unlocks the same master key, so whoever prunes can read every backup in that repo. restic design (documented, not measured)

2. The prerequisite every option shares — the lock is worthless until this is closed

F6 + K1: a box that is broken into can get the password, log in on port 23 with it, and rewrite authorized_keys — removing the forced command from its own key. So before any lock ships:

  1. The box never receives the sub-account password. The hub becomes the key registrar: the box generates its key pair and sends the PUBLIC key to the hub; the hub (which already holds the password, K2) writes the forced line into authorized_keys over SFTP port 23. The off-site self-heal re-arms the HUB's install job, not a password for the box.
  2. The hub checks the file. Each day the hub reads authorized_keys and alarms if any line lacks the forced prefix (outside an open clean-up window). This is also the answer to "an operator adds a key later": a key without the prefix re-opens deletion, so a check — not a rule — catches it.
  3. Optional hardening, operator's call: rotate the sub-account password after each hub use and keep it only in the hub (K2 stays true either way: the hub can already delete every household's off-site history today — new row).

3. The options

Option 1 — a clean-up window opened by the hub (the box keeps pruning its own repo)

The box normally holds only its add-only key. Once a week the hub adds a second, deleting key line for the same box key pair (or a second key the box holds) for ~15 minutes; the box runs its retention; the hub removes the line and checks the file.

  • Custody: unchanged. The repository password never leaves the box (07 §8a stays true).
  • What a compromised box can do: delete during a window. Detection backstop: R-431 (a fall of more than half), plus a new per-window check — the hub compares the snapshot count before and after and alarms on a fall larger than the retention could cause.
  • Poisoning (F10) must be guarded even here, because an attacker who was on the box and left can plant future snapshots that the next honest window then obeys: before forget, refuse when any snapshot is dated in the future or newer than the newest the hub has seen reported; run --dry-run first and abort if it would remove more snapshots than the policy can remove in a week.
  • Build cost: hub: key registrar + window open/close + authorized_keys check (SFTP with the password it holds — no provider API, no main account). Controller: switch the transport from sftp: to rclone: with -o rclone.program="ssh -p 23 … -i <key> … rclone" (restic 0.14.0 is enough, F1; rclone is NOT needed in the image); move forget --prune (both sites, together — R-191) into the window; move the move-aside and the abandonment to the hub (K3).
  • Money: none.

Option 2 — clean-up runs off the box

A Felhom-side worker prunes each repo with a deleting key.

  • Custody: CHANGES. restic has no prune-only key (K4): the worker must hold every repository password in a form it can use unattended, and could read every household's backups. That changes a promise to the customer (07 §8a) — not CC's to change.
  • Poisoning (F10): same guard needed.
  • Build cost: a new always-on worker in the recovery path, its own credentials store, its own failure alarms — plus everything in §2.

Option 3 — never delete; let the quota grow

  • Money: small. The pool box (BX11, 1 TB) was €4.06/month when recorded (SPIKE-ep0-storagebox- 2026-07-09; Hetzner raised prices in April 2026 — re-check) ≈ €0.004 per GB per month. A household adding 10 GB a year unpruned costs ≈ €0.50 a year. The demo boxes cannot calibrate this (0.5 GB and 0.0 GB used today), so 10 GB/year is an assumption, stated as one.
  • The real cost is the quota (50–100 GB): at the quota offbox_fit.go refuses new pushes and the over-quota forget (K3) would itself be refused, so the household's off-site copy STOPS. Crash leftovers (F8) and refused-prune index files (F2) also accumulate.
  • Good as an interim, not an end state.

4. Recommendation

Option 1, with Option 3 as the interim while it is built. Order:

  1. §2 first (key registrar + the authorized_keys check) — without it nothing below protects anything.
  2. Switch every box to the forced key with no box-side retention (Option 3 interim). Quotas today are 50–100 GB against ≤0.5 GB used, so this costs nothing for months.
  3. Then the weekly window (Option 1) with the poisoning guard.

Why not Option 2: it trades a box-level risk for a custody change that lets one Felhom machine read every household's backups, and it adds an always-on service to the recovery path.

5. Migration, rotation, restore

  • Existing repos (both demo boxes, any household): only the authorized_keys line and the box's transport change; the repository bytes are not touched. Not measured: reading a repo written over sftp: through rclone: — restic's on-disk layout is backend-independent, so it is expected to work; measure on the scratch account before any household. History is kept.
  • Key rotation: the hub writes the new forced line, the box switches, the hub removes the old line — the same registrar path.
  • Restore works through the add-only key (F1) — no second key is needed for the household.
  • Locks: the self-heal (resticStep → unlock --remove-all) keeps working under this transport (F7). R-430's fear does not apply to it.
  • Failure speed: with today's code a forced key makes each night's forget --prune fail after ~45 s per refused file and grow the index (F2) — the two forget sites must leave the box in the same change that switches the key.