Files
felhom.eu/documentation/audits/SPIKE-r95-offsite-delete-2026-09-01.md
T
admin 0476a8d8e6
gates / gates (push) Successful in 17s
SPIKE R-95: the safety net cannot be seen from the box - and that RAISES the urgency
READ-ONLY STUDY. No code, no version, no image, no golden. No delete verb was issued against any
live store. ep0, DooPlex and Peti's box were not touched at all.

Q1 FIRST, AND IT DID NOT GO THE EXPECTED WAY. The brief supposed a working seven-day snapshot net
might bound the worst case. Measured on BOTH boxes over their own SFTP credential, with a positive
and a negative control on each: NO .snapshots is visible to either sub-account - not in the account
home, not inside the repo - and the account is jailed at /. Either none exist or a sub-account
cannot see them, and the second is not a reprieve: a snapshot the box cannot see is one the box
cannot restore from, so recovery would be an operator act at the Hetzner panel.

The register's claim rests on nothing that was checked. The row has NO R-number, so nothing can cite
it; its "confirm tomorrow" was 2026-07-27, 36 days ago; and the DUE-CHECKS block built for exactly
this (R-341) is EMPTY. R-95's word "ARMED" is withdrawn pending R-429. The confirming field is a
Hetzner API field, so this spike STOPPED at the section 11-D fence and left it for Viktor - ten
minutes in the panel, and it re-ranks everything.

Q2: TEN verbs, not nine. `check` was missing from the brief's list; `dump` is not a verb (it is a
progress phase constant) and was withdrawn. There are TWO `forget --prune` sites - offbox.go:1388
AND offbox.go:1759 - and disarming one without the other reproduces R-191 exactly.

Q3 (documented, from our own API mirror): AccessSettings has five booleans and `readonly` is the
only permission axis. No append-only. So the PBS shape does NOT transfer - PBS is a server that can
refuse; a Storage Box is a filesystem that runs nothing.

Q5 (measured, faithful append-only model, both controls passed first): withdrawing delete does NOT
wedge the store - restic treats a dead owner's lock as stale and proceeds. The constraint everyone
feared is not the blocker. But `unlock --remove-all` printed "successfully removed locks" while the
lock was still there, and resticStep's crash-lock self-heal is built on that call - R-430.

Q6 (measured, with a control): restic 0.14.0 DOES speak rest:. Append-only is a rest-server flag,
not a restic one.

Q7 (measured): detection is nearly free. snapshot_count already reaches the hub and the hub APPENDS
reports, so the history to compare against is already on disk.

RECOMMENDATION: answer Q1 today (Viktor, ten minutes), then build detection, then move retention off
the box. Defer the transport change until Q1 is answered.

Register: OPEN 179 -> 181. Filed R-429, R-430; R-95 updated and kept OPEN. 07 row 10's status is
deliberately UNCHANGED.
2026-09-01 13:55:35 +02:00

17 KiB

SPIKE — can the box be stopped from deleting its own off-site history? (R-95, 2026-09-01)

Read-only study. No production code, no version bump, no image, no golden. No delete verb was issued against any live store. ep0, DooPlex and Peti's box were not touched at all.


Q1 — Is the safety net real? NO — and worse than "no": it cannot be seen from the box at all, so the exposure must be treated as UNBOUNDED until Viktor looks at the control panel.

Measured on both live boxes, over the SFTP credential each box already holds, read-only.

probe demo-hp (sub-account A, gid 1058) demo-felhom (gid 1019)
positive control — the account's own home lists .ssh, <repo> lists .ssh, <repo>, <repo>.orphaned-20260810
negative control — ./zzz-no-such-r95 not found not found
./.snapshots not found not found
repo path itself config data index keys locks snapshots data index keys locks snapshots
<repo>/.snapshots not found not found
/ (box root) Permission denied — the account is jailed to /home same

Two readings, and the box cannot distinguish them: either no snapshots exist, or they exist and are invisible to a sub-account. Both are bad, and the second is not a reprieve — a snapshot the box cannot see is a snapshot the box cannot restore from either. Recovery would be an operator act through the Hetzner panel, not something the product can do.

What the register actually claims, OPEN-ITEMS.md:233, verbatim:

| — | Storage Box **snapshots** on storage-box-pool-1 — plan SET (daily 00:00, keep 7) but **0 taken yet** | WATCHING | first run tonight 00:00 | Confirm size_snapshots > 0 tomorrow; ... until then the mitigation is armed, not proven |

Three facts about that row:

  1. It has no R-number (| — |), so no gate and no grep-the-register rule can ever cite it.
  2. Its "confirm tomorrow" was 2026-07-27 (git log -S, last touched in 72692e1). That is 36 days ago. Nobody confirmed.
  3. The DUE-CHECKS block — the mechanism this project built (R-341) for exactly this — is EMPTY. The one dated commitment that mattered was never entered into it.

And R-95 leans on it: its own text reads "Mitigation now ARMED". That word is doing work it has not earned. → R-429.

Not answerable here: whether snapshots exist. The register's own confirming field, size_snapshots, is a Hetzner API field, and §11-D records the deliberate decision not to call that API with the production token. This spike stopped at that fence. → a question for Viktor.

Explicitly NOT done: no restore from a snapshot was attempted. Whether one can be restored is a separate question and is named as such.

What this does to the urgency: it raises it. The task supposed Q1 might bound the worst case to seven days. It does not. The bound is unverified, and unverifiable from the product side.


Q2 — What must the box write, and what must it delete? Ten verbs, not nine — and there are TWO forget --prune sites, not one.

Census by grep over internal/backup/*.go at 960d29b, each confirmed by reading its context.

verb file:line class
init offbox.go:321, offbox.go:811 write
backup offbox.go:927, offbox.go:1332, offbox_shares.go:107 write
forget … --prune offbox.go:1388 (retention) and offbox.go:1759 (over-quota) DELETE
unlock offbox.go:746 (stale-only), offbox.go:768 (--remove-all) DELETE
restore offbox.go:1827, offbox_restore.go:393, offbox_proof.go:328 (--no-lock) read
snapshots offbox.go:1771, offbox_inventory.go:73, offbox_restore.go:126 read
stats offbox.go:1786, offbox_restore.go:160 read (takes a lock)
cat offbox.go:782 read
check offbox_integrity.go:316 read (takes a lock)

Two corrections to the task's list, both mine to report:

  • check is missing from it. It is a real off-site verb (offbox_integrity.go:316) and it is lock-taking, which makes it central to Q5.
  • dump is NOT a verb. offbox_progress.go:185 is OffboxPhaseDump = "dump", a progress phase name. I withdrew it after reading the context rather than counting it.

Does the box need the delete verbs? forget/prune — no, that is retention and it is exactly what R-89 moved off the box for PBS. unlock — this is the whole of Q5.


Q3 — Can the sub-account express "write but never delete"? No. The API offers exactly one axis, and it is all-or-nothing.

DOCUMENTED, not measured — from this project's own mirror of the API, hub/internal/hetznerapi/hetznerapi.go:38-45, which is the authority we have without calling the provider:

type AccessSettings struct {
    SSHEnabled          bool `json:"ssh_enabled"`
    ReachableExternally bool `json:"reachable_externally"`
    SambaEnabled        bool `json:"samba_enabled"`
    WebDAVEnabled       bool `json:"webdav_enabled"`
    Readonly            bool `json:"readonly"`
}

Five booleans. readonly is the only permission axis, and a backup target cannot be read-only. There is no append-only, no per-directory grant, no write-without-unlink. This is the finding that forces Q6, and the register already suspected it.

Measured corroboration, no new action taken: demo-felhom's home contains a <repo>.orphaned-20260810 directory beside <repo>. That directory was renamed by the controller (the orphan guard). A rename is a delete-class operation. So the credential's reach is not a theory — the existing state on the live store is evidence of it.

Not established: whether the live API would report anything the struct omits. Settling it needs the provider token — fenced (§11-D).


Q4 — What did the PBS fix do, and what would copying it cost? The shape does not transfer, because the far end is a disk, not a server.

What moved (R-89, PROVEN-LIVE): retention became a hub-owned commercial attribute; boxes set keep_last: 0; ep0 runs the prune jobs; box tokens stay write-only. 07 §8 row 10 records the box being refused when deleting its own PBS snapshot.

Why it does not transfer as-is: PBS is a server that can refuse. The restic tier writes to a Storage Box over SFTP — a filesystem. A filesystem runs nothing. So "move retention to the far end" has no far end to move it to. Candidates, with their costs:

who runs retention instead cost new risk it creates
the hub a scheduler + a per-customer credential the hub gains a credential that can delete every customer's history — one compromise instead of N
DooPlex a cron + the same credentials same concentration, on a Tier-2 box that is itself the recovery chain
ep0 a service on a protected machine an architecture change, not a config change; ep0 is protected
nobody — never prune £0 today the store grows without bound; the over-quota path at offbox.go:1759 exists precisely because quota is already a live concern

What the migration cost last time, and the trap is identical here. R-191: R-89 changed the contract, one client-side retention setting did not follow, and every successful weekly backup then reported as a failure — upload complete, 67.2 % reused, then TASK ERROR: job errors and a whole_guest_backup_failed alarm to the operator. "A doc that states the contract does not enforce it — the gate does."

The same trap, doubled: withdrawing delete without disarming retention would make every off-site run log [WARN] forget --prune failed — and there are two call sites, offbox.go:1388 and offbox.go:1759. The task's brief names only the first. Disarming one and not the other reproduces R-191 exactly.


Q5 — The lock problem. Measured, and it is NOT the blocker I expected — but it exposed something worse: unlock reports success on a deletion that did not happen.

Method. A throwaway local restic repo in the controller container's /tmp on demo-hp (60 MB, removed at the end, no live store touched). Append-only was modelled faithfully: a sticky locks directory (1777) owned by root, holding a root-owned lock, with restic run as nobody. That is exactly append-only semantics — create allowed, delete refused.

Both controls passed before anything was believed:

  • nobody can create in the locks directory → the model is not simply "read-only".
  • nobody cannot delete root's lock → the model really does refuse deletes.

A genuine restic lock was captured (copied out while a real check --read-data held it), not hand-forged.

verb, running as nobody against an undeletable stale lock result
restic unlock --remove-all prints successfully removed locks — and the lock is STILL THERE
restic backup succeeds — snapshot a7928829 saved
restic check succeeds — no errors were found
restic snapshots --no-lock succeeds
restic restore --no-lock (the R-87 proof pattern) succeeds

So the answer to "does the store wedge?" is: not in this case. restic 0.14.0 treats a lock whose owner is provably dead as stale and proceeds without needing to remove it. Withdrawing delete does not, by itself, wedge the store. That removes the constraint the task suspected would be deciding.

But the measurement found a different defect. unlock --remove-all reported success while deleting nothing. This project already has a named class for that — "Exit codes that lie" — and resticStep's crash-lock self-heal is built directly on top of this call. A self-heal that cannot fail is a self-heal that cannot be trusted. → R-430.

UNKNOWN, and I am not claiming otherwise: the crash-lock case — a lock left by a container that no longer exists, whose hostname restic cannot match, so it will not treat it as stale for ~30 minutes. That is documented in resticStep's own comment (offbox.go:~760) and is the reason --remove-all exists at all. My model could not reproduce it: the captured lock carried this container's own hostname. What would settle it: a lock captured from a container with a different hostname, replayed against an append-only endpoint. In that window, the only remedy is the very call that Q5 just showed reports success while doing nothing.

How far --no-lock reaches: restic's own help says it "allows some operations on read-only repositories". Measured: snapshots and restore work under it. backup and check still take a lock — but creating a lock is a write, which append-only permits. So the read paths are safe and the write paths are unaffected; only removal is denied, and only the crash-lock window depends on removal.


Q6 — Is an append-only transport reachable? Yes in principle — restic 0.14.0 does speak REST — but not without either moving the data or putting a machine in front of it.

Measured, in the controller container, with a control:

repo string result
banana:http://127.0.0.1:1/x (control) Fatal: parsing repository location failed: invalid backend
rest:http://127.0.0.1:1/ Fatal: unable to open config file: … dial tcp 127.0.0.1:1: connect: connection refused

A connection error, not a parse error — the REST backend IS recognised by restic 0.14.0.

Append-only is not a restic feature. Measured: restic help and restic backup --help contain zero occurrences of "append". It is a rest-server flag (--append-only). So the client is ready and the server does not exist yet.

In front, or move the data? rest-server serves a local directory. The Storage Box is a remote share. So either (a) a machine mounts the Storage Box and runs rest-server over that mount — keeping the bytes where they are, at the cost of a mount in the hot path — or (b) the data moves to storage attached to whatever runs rest-server. These are very different prices and the choice is not obvious; (a) keeps the current bill, (b) is a migration of every customer's history.

Where would it run? ep0 is protected. Adding a service to it is an architecture change, not a configuration change — and it also makes ep0 a single point of failure for both tiers, which is precisely the concern §11-D already has open.

Cost: one new always-on service in the recovery path, plus whatever machine hosts it. The moving part matters more than the money: if rest-server is down, backups stop — and this project's own history (R-191) is a warning about what happens when a change in the off-site contract is not carried everywhere at once.


Q7 — What does DETECTION cost? Almost nothing. The number already arrives at the hub, and the hub already keeps the history to compare it against.

Measured from source:

  • The box already reports snapshot_count — controller/internal/report/types.go:131, and it is already rendered on the hub's Backup card (hub/internal/web/backup_card.go:118).
  • The hub retains report history: store.go:968 is an INSERT INTO reports, and the read is SELECT … ORDER BY id DESC LIMIT 1 (store.go:1016). It appends; it does not replace. So every previous snapshot_count for every customer is already on disk.

So an unexplained drop is a comparison between two rows the hub already has. No new box code, no new credential, no new moving part, no new bill. The alarm vocabulary and the operator-only routing already exist.

The honest caveats: a legitimate forget also drops the count, so the rule needs to know retention (and today the box prunes itself, which is what makes the number noisy — the same change that disarms box-side retention is what makes this signal clean). And detection is not prevention: it tells you within a day that history was destroyed; it does not stop it. snapshot_count: 0 must mean UNKNOWN, not EMPTY — backup_card.go:29 records that exact defect being fixed once already.


Ranked options for Viktor

All four were considered. My recommendation is 4 + 2, in that order, and explicitly not 3 yet.

1. Do nothing — NOT acceptable as it stands, and Q1 is why

Before this spike, "do nothing" meant "exposed, but bounded to seven days by snapshots." That sentence is not supported by any evidence. The box cannot see a single snapshot, on either machine, and the register's confirmation step was never done. Doing nothing now means an unbounded exposure on the tier holding the customer's documents and photos. Cost: £0, no evenings. What happens if you choose it: the risk stays exactly where it is, and the register keeps saying "ARMED", which is the part I would not accept. One cheap act rescues most of this option: look at the Storage Box panel and answer Q1. If snapshots are real, "do nothing" becomes defensible again. That is ten minutes and it is yours to do — the fence in §11-D stopped me.

2. Copy the PBS shape — worth doing, but it is smaller than it sounds and it is not the safety it appears to be

Withdraw box-side retention; someone else prunes. Q5 says this no longer looks dangerous — a stale lock does not wedge the store. But Q3 says the credential still cannot be narrowed: the API has one axis, readonly, and a backup target cannot be read-only. So the box would keep the ability to delete, and simply stop using it. That is a discipline, not a guarantee — a compromised guest is unaffected by it. Cost: one evening, plus wherever retention moves (each candidate concentrates the credential). Trap: both forget sites (offbox.go:1388 and :1759) must be disarmed in the same change, or you get R-191 again — successful backups reported as failures.

3. Change the transport (rest-server --append-only) — the only real prevention, and I would not start it yet

Q6 says it is reachable: restic 0.14.0 speaks REST. This is the option that actually makes the box unable to delete. Cost: a new always-on service in the recovery path; a machine to host it (ep0 is protected — this is an architecture change); and either a mount in the hot path or moving every customer's history. Money is the small part. Why not yet: it should be chosen against a measured picture, and one measurement is still missing — Q1. Building the expensive prevention while nobody knows whether a seven-day net already exists is the wrong order.

4. Detect instead of prevent — cheapest by a wide margin, and I would do this first

Q7 measured that the material is already there: the count is on the wire, the hub keeps the history, the alarm path exists. An unexplained drop in snapshot_count becomes noticeable within a day. Cost: hub-side only. No box change, no credential change, no new service, no new bill. What it does not do: it does not stop the deletion. It converts "we would never know" into "we know tomorrow" — and against ransomware inside the guest, knowing tomorrow is the difference between losing a day and losing everything silently. This week's pattern held twice already (R-87, R-404): proving beats preventing when preventing is expensive. This is the third instance.

My pick: answer Q1 today (yours, ten minutes), then build 4, then 2. Revisit 3 once Q1 is answered — and if Q1 comes back "no snapshots", 3 moves up sharply.