# SPIKE — can the box be stopped from deleting its own off-site history? (R-95, 2026-09-01) **Read-only study. No production code, no version bump, no image, no golden. No delete verb was issued against any live store. `ep0`, DooPlex and Peti's box were not touched at all.** --- ## Q1 — Is the safety net real? **NO — and worse than "no": it cannot be seen from the box at all, so the exposure must be treated as UNBOUNDED until Viktor looks at the control panel.** **Measured** on both live boxes, over the SFTP credential each box already holds, read-only. | probe | demo-hp (sub-account A, gid 1058) | demo-felhom (gid 1019) | |---|---|---| | **positive control** — the account's own home | lists `.ssh`, `` | lists `.ssh`, ``, `.orphaned-20260810` | | **negative control** — `./zzz-no-such-r95` | `not found` | `not found` | | `./.snapshots` | **`not found`** | **`not found`** | | repo path itself | `config data index keys locks snapshots` | `data index keys locks snapshots` | | `/.snapshots` | `not found` | `not found` | | `/` (box root) | **`Permission denied`** — the account is jailed to `/home` | same | **Two readings, and the box cannot distinguish them:** either no snapshots exist, or they exist and are invisible to a sub-account. **Both are bad, and the second is not a reprieve** — a snapshot the box cannot see is a snapshot the box cannot restore from either. Recovery would be an operator act through the Hetzner panel, not something the product can do. **What the register actually claims**, `OPEN-ITEMS.md:233`, verbatim: > `| — | Storage Box **snapshots** on storage-box-pool-1 — plan SET (daily 00:00, keep 7) but **0 taken yet** | WATCHING | first run tonight 00:00 | Confirm size_snapshots > 0 tomorrow; ... until then the mitigation is armed, not proven |` **Three facts about that row:** 1. **It has no R-number** (`| — |`), so no gate and no grep-the-register rule can ever cite it. 2. **Its "confirm tomorrow" was 2026-07-27** (`git log -S`, last touched in `72692e1`). That is **36 days** ago. Nobody confirmed. 3. **The `DUE-CHECKS` block — the mechanism this project built (R-341) for exactly this — is EMPTY.** The one dated commitment that mattered was never entered into it. **And R-95 leans on it**: its own text reads *"Mitigation now ARMED"*. **That word is doing work it has not earned.** → **R-429**. **Not answerable here:** whether snapshots exist. The register's own confirming field, `size_snapshots`, is a **Hetzner API field**, and §11-D records the deliberate decision not to call that API with the production token. **This spike stopped at that fence.** → a question for Viktor. **Explicitly NOT done:** no restore from a snapshot was attempted. Whether one *can* be restored is a separate question and is named as such. **What this does to the urgency: it raises it.** The task supposed Q1 might bound the worst case to seven days. It does not. **The bound is unverified, and unverifiable from the product side.** --- ## Q2 — What must the box write, and what must it delete? **Ten verbs, not nine — and there are TWO `forget --prune` sites, not one.** Census by grep over `internal/backup/*.go` at `960d29b`, each confirmed by reading its context. | verb | `file:line` | class | |---|---|---| | `init` | `offbox.go:321`, `offbox.go:811` | **write** | | `backup` | `offbox.go:927`, `offbox.go:1332`, `offbox_shares.go:107` | **write** | | `forget … --prune` | **`offbox.go:1388`** (retention) and **`offbox.go:1759`** (over-quota) | **DELETE** | | `unlock` | `offbox.go:746` (stale-only), `offbox.go:768` (`--remove-all`) | **DELETE** | | `restore` | `offbox.go:1827`, `offbox_restore.go:393`, `offbox_proof.go:328` (`--no-lock`) | read | | `snapshots` | `offbox.go:1771`, `offbox_inventory.go:73`, `offbox_restore.go:126` | read | | `stats` | `offbox.go:1786`, `offbox_restore.go:160` | read (**takes a lock**) | | `cat` | `offbox.go:782` | read | | `check` | `offbox_integrity.go:316` | read (**takes a lock**) | **Two corrections to the task's list, both mine to report:** * **`check` is missing from it.** It is a real off-site verb (`offbox_integrity.go:316`) and it is lock-taking, which makes it central to Q5. * **`dump` is NOT a verb.** `offbox_progress.go:185` is `OffboxPhaseDump = "dump"`, a progress phase name. I withdrew it after reading the context rather than counting it. **Does the box need the delete verbs?** `forget`/`prune` — **no**, that is retention and it is exactly what R-89 moved off the box for PBS. `unlock` — **this is the whole of Q5**. --- ## Q3 — Can the sub-account express "write but never delete"? **No. The API offers exactly one axis, and it is all-or-nothing.** **DOCUMENTED, not measured** — from this project's own mirror of the API, `hub/internal/hetznerapi/hetznerapi.go:38-45`, which is the authority we have without calling the provider: ```go type AccessSettings struct { SSHEnabled bool `json:"ssh_enabled"` ReachableExternally bool `json:"reachable_externally"` SambaEnabled bool `json:"samba_enabled"` WebDAVEnabled bool `json:"webdav_enabled"` Readonly bool `json:"readonly"` } ``` **Five booleans. `readonly` is the only permission axis, and a backup target cannot be read-only.** There is no append-only, no per-directory grant, no write-without-unlink. **This is the finding that forces Q6**, and the register already suspected it. **Measured corroboration, no new action taken:** demo-felhom's home contains a `.orphaned-20260810` directory beside ``. That directory was **renamed by the controller** (the orphan guard). A rename is a delete-class operation. **So the credential's reach is not a theory — the existing state on the live store is evidence of it.** **Not established:** whether the live API would report anything the struct omits. Settling it needs the provider token — **fenced** (§11-D). --- ## Q4 — What did the PBS fix do, and what would copying it cost? **The shape does not transfer, because the far end is a disk, not a server.** **What moved (R-89, PROVEN-LIVE):** retention became a hub-owned commercial attribute; boxes set `keep_last: 0`; **ep0 runs the prune jobs**; box tokens stay write-only. `07` §8 row 10 records the box being *refused* when deleting its own PBS snapshot. **Why it does not transfer as-is:** PBS is a **server** that can refuse. The restic tier writes to a **Storage Box over SFTP** — a filesystem. **A filesystem runs nothing.** So "move retention to the far end" has no far end to move it to. Candidates, with their costs: | who runs retention instead | cost | new risk it creates | |---|---|---| | **the hub** | a scheduler + a per-customer credential | **the hub gains a credential that can delete every customer's history** — one compromise instead of N | | **DooPlex** | a cron + the same credentials | same concentration, on a **Tier-2 box that is itself the recovery chain** | | **ep0** | a service on a protected machine | **an architecture change, not a config change**; ep0 is protected | | **nobody — never prune** | £0 today | the store grows without bound; the over-quota path at `offbox.go:1759` exists precisely because quota is already a live concern | **What the migration cost last time, and the trap is identical here.** R-191: R-89 changed the contract, one client-side retention setting did not follow, and **every successful weekly backup then reported as a failure** — upload complete, 67.2 % reused, then `TASK ERROR: job errors` and a `whole_guest_backup_failed` alarm to the operator. *"A doc that states the contract does not enforce it — the gate does."* **The same trap, doubled:** withdrawing delete without disarming retention would make every off-site run log `[WARN] forget --prune failed` — and there are **two** call sites, `offbox.go:1388` **and `offbox.go:1759`**. The task's brief names only the first. **Disarming one and not the other reproduces R-191 exactly.** --- ## Q5 — The lock problem. **Measured, and it is NOT the blocker I expected — but it exposed something worse: `unlock` reports success on a deletion that did not happen.** **Method.** A throwaway **local** restic repo in the controller container's `/tmp` on demo-hp (60 MB, removed at the end, **no live store touched**). Append-only was modelled faithfully: a sticky locks directory (`1777`) owned by root, holding a root-owned lock, with restic run as `nobody`. That is exactly append-only semantics — **create allowed, delete refused**. **Both controls passed before anything was believed:** * `nobody` **can** create in the locks directory → the model is not simply "read-only". * `nobody` **cannot** delete root's lock → the model really does refuse deletes. A **genuine** restic lock was captured (copied out while a real `check --read-data` held it), not hand-forged. | verb, running as `nobody` against an undeletable stale lock | result | |---|---| | `restic unlock --remove-all` | **prints `successfully removed locks` — and the lock is STILL THERE** | | `restic backup` | **succeeds** — `snapshot a7928829 saved` | | `restic check` | **succeeds** — `no errors were found` | | `restic snapshots --no-lock` | succeeds | | `restic restore --no-lock` (the R-87 proof pattern) | succeeds | **So the answer to "does the store wedge?" is: not in this case.** restic 0.14.0 treats a lock whose owner is provably dead as stale and proceeds without needing to remove it. **Withdrawing delete does not, by itself, wedge the store.** That removes the constraint the task suspected would be deciding. **But the measurement found a different defect.** `unlock --remove-all` **reported success while deleting nothing.** This project already has a named class for that — *"Exit codes that lie"* — and `resticStep`'s crash-lock self-heal is built directly on top of this call. **A self-heal that cannot fail is a self-heal that cannot be trusted.** → **R-430**. **UNKNOWN, and I am not claiming otherwise:** the **crash-lock** case — a lock left by a container that no longer exists, whose hostname restic cannot match, so it will **not** treat it as stale for ~30 minutes. That is documented in `resticStep`'s own comment (`offbox.go:~760`) and is the reason `--remove-all` exists at all. My model could not reproduce it: the captured lock carried this container's own hostname. **What would settle it:** a lock captured from a container with a different hostname, replayed against an append-only endpoint. **In that window, the only remedy is the very call that Q5 just showed reports success while doing nothing.** **How far `--no-lock` reaches:** restic's own help says it *"allows some operations on read-only repositories"*. Measured: `snapshots` and `restore` work under it. `backup` and `check` still take a lock — but **creating** a lock is a write, which append-only permits. So the read paths are safe and the write paths are unaffected; **only removal is denied, and only the crash-lock window depends on removal.** --- ## Q6 — Is an append-only transport reachable? **Yes in principle — restic 0.14.0 does speak REST — but not without either moving the data or putting a machine in front of it.** **Measured**, in the controller container, with a control: | repo string | result | |---|---| | `banana:http://127.0.0.1:1/x` (control) | `Fatal: parsing repository location failed: invalid backend` | | `rest:http://127.0.0.1:1/` | `Fatal: unable to open config file: … dial tcp 127.0.0.1:1: connect: connection refused` | **A connection error, not a parse error — the REST backend IS recognised by restic 0.14.0.** **Append-only is not a restic feature.** Measured: `restic help` and `restic backup --help` contain **zero** occurrences of "append". It is a **rest-server** flag (`--append-only`). So the client is ready and the server does not exist yet. **In front, or move the data?** rest-server serves a **local directory**. The Storage Box is a remote share. So either (a) a machine mounts the Storage Box and runs rest-server over that mount — keeping the bytes where they are, at the cost of a mount in the hot path — or (b) the data moves to storage attached to whatever runs rest-server. **These are very different prices and the choice is not obvious**; (a) keeps the current bill, (b) is a migration of every customer's history. **Where would it run?** **`ep0` is protected. Adding a service to it is an architecture change, not a configuration change** — and it also makes ep0 a single point of failure for both tiers, which is precisely the concern §11-D already has open. **Cost:** one new always-on service in the recovery path, plus whatever machine hosts it. **The moving part matters more than the money:** if rest-server is down, backups stop — and this project's own history (R-191) is a warning about what happens when a change in the off-site contract is not carried everywhere at once. --- ## Q7 — What does DETECTION cost? **Almost nothing. The number already arrives at the hub, and the hub already keeps the history to compare it against.** **Measured from source:** * The box already reports `snapshot_count` — `controller/internal/report/types.go:131`, and it is already rendered on the hub's Backup card (`hub/internal/web/backup_card.go:118`). * **The hub retains report history**: `store.go:968` is an `INSERT INTO reports`, and the read is `SELECT … ORDER BY id DESC LIMIT 1` (`store.go:1016`). **It appends; it does not replace.** So every previous `snapshot_count` for every customer is already on disk. **So an unexplained drop is a comparison between two rows the hub already has.** No new box code, no new credential, no new moving part, no new bill. The alarm vocabulary and the operator-only routing already exist. **The honest caveats:** a legitimate `forget` also drops the count, so the rule needs to know retention (and today the box prunes itself, which is what makes the number noisy — **the same change that disarms box-side retention is what makes this signal clean**). And detection is not prevention: it tells you within a day that history was destroyed; it does not stop it. **`snapshot_count: 0` must mean UNKNOWN, not EMPTY** — `backup_card.go:29` records that exact defect being fixed once already. --- # Ranked options for Viktor **All four were considered. My recommendation is 4 + 2, in that order, and explicitly not 3 yet.** ### 1. Do nothing — **NOT acceptable as it stands, and Q1 is why** Before this spike, "do nothing" meant *"exposed, but bounded to seven days by snapshots."* **That sentence is not supported by any evidence.** The box cannot see a single snapshot, on either machine, and the register's confirmation step was never done. **Doing nothing now means an unbounded exposure on the tier holding the customer's documents and photos.** **Cost:** £0, no evenings. **What happens if you choose it:** the risk stays exactly where it is, and the register keeps saying "ARMED", which is the part I would not accept. **One cheap act rescues most of this option: look at the Storage Box panel and answer Q1.** If snapshots are real, "do nothing" becomes defensible again. **That is ten minutes and it is yours to do — the fence in §11-D stopped me.** ### 2. Copy the PBS shape — **worth doing, but it is smaller than it sounds and it is not the safety it appears to be** Withdraw box-side retention; someone else prunes. **Q5 says this no longer looks dangerous** — a stale lock does not wedge the store. **But Q3 says the credential still cannot be narrowed**: the API has one axis, `readonly`, and a backup target cannot be read-only. **So the box would keep the ability to delete, and simply stop using it.** That is a discipline, not a guarantee — a compromised guest is unaffected by it. **Cost:** one evening, plus wherever retention moves (each candidate concentrates the credential). **Trap:** both `forget` sites (`offbox.go:1388` **and** `:1759`) must be disarmed in the same change, or you get R-191 again — successful backups reported as failures. ### 3. Change the transport (rest-server `--append-only`) — **the only real prevention, and I would not start it yet** **Q6 says it is reachable**: restic 0.14.0 speaks REST. This is the option that actually makes the box *unable* to delete. **Cost:** a new always-on service in the recovery path; a machine to host it (**ep0 is protected — this is an architecture change**); and either a mount in the hot path or moving every customer's history. **Money is the small part.** **Why not yet:** it should be chosen against a measured picture, and one measurement is still missing — Q1. Building the expensive prevention while nobody knows whether a seven-day net already exists is the wrong order. ### 4. Detect instead of prevent — **cheapest by a wide margin, and I would do this first** **Q7 measured that the material is already there**: the count is on the wire, the hub keeps the history, the alarm path exists. An unexplained drop in `snapshot_count` becomes noticeable within a day. **Cost:** hub-side only. No box change, no credential change, no new service, no new bill. **What it does not do:** it does not stop the deletion. It converts "we would never know" into "we know tomorrow" — and against ransomware inside the guest, knowing tomorrow is the difference between losing a day and losing everything silently. **This week's pattern held twice already** (R-87, R-404): **proving beats preventing when preventing is expensive.** This is the third instance. **My pick: answer Q1 today (yours, ten minutes), then build 4, then 2. Revisit 3 once Q1 is answered** — and if Q1 comes back "no snapshots", 3 moves up sharply.