READ-ONLY STUDY. No code, no version, no image, no golden. No delete verb was issued against any live store. ep0, DooPlex and Peti's box were not touched at all. Q1 FIRST, AND IT DID NOT GO THE EXPECTED WAY. The brief supposed a working seven-day snapshot net might bound the worst case. Measured on BOTH boxes over their own SFTP credential, with a positive and a negative control on each: NO .snapshots is visible to either sub-account - not in the account home, not inside the repo - and the account is jailed at /. Either none exist or a sub-account cannot see them, and the second is not a reprieve: a snapshot the box cannot see is one the box cannot restore from, so recovery would be an operator act at the Hetzner panel. The register's claim rests on nothing that was checked. The row has NO R-number, so nothing can cite it; its "confirm tomorrow" was 2026-07-27, 36 days ago; and the DUE-CHECKS block built for exactly this (R-341) is EMPTY. R-95's word "ARMED" is withdrawn pending R-429. The confirming field is a Hetzner API field, so this spike STOPPED at the section 11-D fence and left it for Viktor - ten minutes in the panel, and it re-ranks everything. Q2: TEN verbs, not nine. `check` was missing from the brief's list; `dump` is not a verb (it is a progress phase constant) and was withdrawn. There are TWO `forget --prune` sites - offbox.go:1388 AND offbox.go:1759 - and disarming one without the other reproduces R-191 exactly. Q3 (documented, from our own API mirror): AccessSettings has five booleans and `readonly` is the only permission axis. No append-only. So the PBS shape does NOT transfer - PBS is a server that can refuse; a Storage Box is a filesystem that runs nothing. Q5 (measured, faithful append-only model, both controls passed first): withdrawing delete does NOT wedge the store - restic treats a dead owner's lock as stale and proceeds. The constraint everyone feared is not the blocker. But `unlock --remove-all` printed "successfully removed locks" while the lock was still there, and resticStep's crash-lock self-heal is built on that call - R-430. Q6 (measured, with a control): restic 0.14.0 DOES speak rest:. Append-only is a rest-server flag, not a restic one. Q7 (measured): detection is nearly free. snapshot_count already reaches the hub and the hub APPENDS reports, so the history to compare against is already on disk. RECOMMENDATION: answer Q1 today (Viktor, ten minutes), then build detection, then move retention off the box. Defer the transport change until Q1 is answered. Register: OPEN 179 -> 181. Filed R-429, R-430; R-95 updated and kept OPEN. 07 row 10's status is deliberately UNCHANGED.
17 KiB
SPIKE — can the box be stopped from deleting its own off-site history? (R-95, 2026-09-01)
Read-only study. No production code, no version bump, no image, no golden. No delete verb was
issued against any live store. ep0, DooPlex and Peti's box were not touched at all.
Q1 — Is the safety net real? NO — and worse than "no": it cannot be seen from the box at all, so the exposure must be treated as UNBOUNDED until Viktor looks at the control panel.
Measured on both live boxes, over the SFTP credential each box already holds, read-only.
| probe | demo-hp (sub-account A, gid 1058) | demo-felhom (gid 1019) |
|---|---|---|
| positive control — the account's own home | lists .ssh, <repo> |
lists .ssh, <repo>, <repo>.orphaned-20260810 |
negative control — ./zzz-no-such-r95 |
not found |
not found |
./.snapshots |
not found |
not found |
| repo path itself | config data index keys locks snapshots |
data index keys locks snapshots |
<repo>/.snapshots |
not found |
not found |
/ (box root) |
Permission denied — the account is jailed to /home |
same |
Two readings, and the box cannot distinguish them: either no snapshots exist, or they exist and are invisible to a sub-account. Both are bad, and the second is not a reprieve — a snapshot the box cannot see is a snapshot the box cannot restore from either. Recovery would be an operator act through the Hetzner panel, not something the product can do.
What the register actually claims, OPEN-ITEMS.md:233, verbatim:
| — | Storage Box **snapshots** on storage-box-pool-1 — plan SET (daily 00:00, keep 7) but **0 taken yet** | WATCHING | first run tonight 00:00 | Confirm size_snapshots > 0 tomorrow; ... until then the mitigation is armed, not proven |
Three facts about that row:
- It has no R-number (
| — |), so no gate and no grep-the-register rule can ever cite it. - Its "confirm tomorrow" was 2026-07-27 (
git log -S, last touched in72692e1). That is 36 days ago. Nobody confirmed. - The
DUE-CHECKSblock — the mechanism this project built (R-341) for exactly this — is EMPTY. The one dated commitment that mattered was never entered into it.
And R-95 leans on it: its own text reads "Mitigation now ARMED". That word is doing work it has not earned. → R-429.
Not answerable here: whether snapshots exist. The register's own confirming field,
size_snapshots, is a Hetzner API field, and §11-D records the deliberate decision not to call
that API with the production token. This spike stopped at that fence. → a question for Viktor.
Explicitly NOT done: no restore from a snapshot was attempted. Whether one can be restored is a separate question and is named as such.
What this does to the urgency: it raises it. The task supposed Q1 might bound the worst case to seven days. It does not. The bound is unverified, and unverifiable from the product side.
Q2 — What must the box write, and what must it delete? Ten verbs, not nine — and there are TWO forget --prune sites, not one.
Census by grep over internal/backup/*.go at 960d29b, each confirmed by reading its context.
| verb | file:line |
class |
|---|---|---|
init |
offbox.go:321, offbox.go:811 |
write |
backup |
offbox.go:927, offbox.go:1332, offbox_shares.go:107 |
write |
forget … --prune |
offbox.go:1388 (retention) and offbox.go:1759 (over-quota) |
DELETE |
unlock |
offbox.go:746 (stale-only), offbox.go:768 (--remove-all) |
DELETE |
restore |
offbox.go:1827, offbox_restore.go:393, offbox_proof.go:328 (--no-lock) |
read |
snapshots |
offbox.go:1771, offbox_inventory.go:73, offbox_restore.go:126 |
read |
stats |
offbox.go:1786, offbox_restore.go:160 |
read (takes a lock) |
cat |
offbox.go:782 |
read |
check |
offbox_integrity.go:316 |
read (takes a lock) |
Two corrections to the task's list, both mine to report:
checkis missing from it. It is a real off-site verb (offbox_integrity.go:316) and it is lock-taking, which makes it central to Q5.dumpis NOT a verb.offbox_progress.go:185isOffboxPhaseDump = "dump", a progress phase name. I withdrew it after reading the context rather than counting it.
Does the box need the delete verbs? forget/prune — no, that is retention and it is exactly
what R-89 moved off the box for PBS. unlock — this is the whole of Q5.
Q3 — Can the sub-account express "write but never delete"? No. The API offers exactly one axis, and it is all-or-nothing.
DOCUMENTED, not measured — from this project's own mirror of the API,
hub/internal/hetznerapi/hetznerapi.go:38-45, which is the authority we have without calling the
provider:
type AccessSettings struct {
SSHEnabled bool `json:"ssh_enabled"`
ReachableExternally bool `json:"reachable_externally"`
SambaEnabled bool `json:"samba_enabled"`
WebDAVEnabled bool `json:"webdav_enabled"`
Readonly bool `json:"readonly"`
}
Five booleans. readonly is the only permission axis, and a backup target cannot be read-only.
There is no append-only, no per-directory grant, no write-without-unlink. This is the finding that
forces Q6, and the register already suspected it.
Measured corroboration, no new action taken: demo-felhom's home contains
a <repo>.orphaned-20260810 directory beside <repo>. That directory was renamed by the controller
(the orphan guard). A rename is a delete-class operation. So the credential's reach is not a
theory — the existing state on the live store is evidence of it.
Not established: whether the live API would report anything the struct omits. Settling it needs the provider token — fenced (§11-D).
Q4 — What did the PBS fix do, and what would copying it cost? The shape does not transfer, because the far end is a disk, not a server.
What moved (R-89, PROVEN-LIVE): retention became a hub-owned commercial attribute; boxes set
keep_last: 0; ep0 runs the prune jobs; box tokens stay write-only. 07 §8 row 10 records the
box being refused when deleting its own PBS snapshot.
Why it does not transfer as-is: PBS is a server that can refuse. The restic tier writes to a Storage Box over SFTP — a filesystem. A filesystem runs nothing. So "move retention to the far end" has no far end to move it to. Candidates, with their costs:
| who runs retention instead | cost | new risk it creates |
|---|---|---|
| the hub | a scheduler + a per-customer credential | the hub gains a credential that can delete every customer's history — one compromise instead of N |
| DooPlex | a cron + the same credentials | same concentration, on a Tier-2 box that is itself the recovery chain |
| ep0 | a service on a protected machine | an architecture change, not a config change; ep0 is protected |
| nobody — never prune | £0 today | the store grows without bound; the over-quota path at offbox.go:1759 exists precisely because quota is already a live concern |
What the migration cost last time, and the trap is identical here. R-191: R-89 changed the
contract, one client-side retention setting did not follow, and every successful weekly backup then
reported as a failure — upload complete, 67.2 % reused, then TASK ERROR: job errors and a
whole_guest_backup_failed alarm to the operator. "A doc that states the contract does not enforce
it — the gate does."
The same trap, doubled: withdrawing delete without disarming retention would make every off-site
run log [WARN] forget --prune failed — and there are two call sites, offbox.go:1388 and
offbox.go:1759. The task's brief names only the first. Disarming one and not the other
reproduces R-191 exactly.
Q5 — The lock problem. Measured, and it is NOT the blocker I expected — but it exposed something worse: unlock reports success on a deletion that did not happen.
Method. A throwaway local restic repo in the controller container's /tmp on demo-hp (60 MB,
removed at the end, no live store touched). Append-only was modelled faithfully: a sticky locks
directory (1777) owned by root, holding a root-owned lock, with restic run as nobody. That is
exactly append-only semantics — create allowed, delete refused.
Both controls passed before anything was believed:
nobodycan create in the locks directory → the model is not simply "read-only".nobodycannot delete root's lock → the model really does refuse deletes.
A genuine restic lock was captured (copied out while a real check --read-data held it), not
hand-forged.
verb, running as nobody against an undeletable stale lock |
result |
|---|---|
restic unlock --remove-all |
prints successfully removed locks — and the lock is STILL THERE |
restic backup |
succeeds — snapshot a7928829 saved |
restic check |
succeeds — no errors were found |
restic snapshots --no-lock |
succeeds |
restic restore --no-lock (the R-87 proof pattern) |
succeeds |
So the answer to "does the store wedge?" is: not in this case. restic 0.14.0 treats a lock whose owner is provably dead as stale and proceeds without needing to remove it. Withdrawing delete does not, by itself, wedge the store. That removes the constraint the task suspected would be deciding.
But the measurement found a different defect. unlock --remove-all reported success while
deleting nothing. This project already has a named class for that — "Exit codes that lie" — and
resticStep's crash-lock self-heal is built directly on top of this call. A self-heal that cannot
fail is a self-heal that cannot be trusted. → R-430.
UNKNOWN, and I am not claiming otherwise: the crash-lock case — a lock left by a container
that no longer exists, whose hostname restic cannot match, so it will not treat it as stale for
~30 minutes. That is documented in resticStep's own comment (offbox.go:~760) and is the reason
--remove-all exists at all. My model could not reproduce it: the captured lock carried this
container's own hostname. What would settle it: a lock captured from a container with a different
hostname, replayed against an append-only endpoint. In that window, the only remedy is the very
call that Q5 just showed reports success while doing nothing.
How far --no-lock reaches: restic's own help says it "allows some operations on read-only
repositories". Measured: snapshots and restore work under it. backup and check still take a
lock — but creating a lock is a write, which append-only permits. So the read paths are safe and
the write paths are unaffected; only removal is denied, and only the crash-lock window depends on
removal.
Q6 — Is an append-only transport reachable? Yes in principle — restic 0.14.0 does speak REST — but not without either moving the data or putting a machine in front of it.
Measured, in the controller container, with a control:
| repo string | result |
|---|---|
banana:http://127.0.0.1:1/x (control) |
Fatal: parsing repository location failed: invalid backend |
rest:http://127.0.0.1:1/ |
Fatal: unable to open config file: … dial tcp 127.0.0.1:1: connect: connection refused |
A connection error, not a parse error — the REST backend IS recognised by restic 0.14.0.
Append-only is not a restic feature. Measured: restic help and restic backup --help contain
zero occurrences of "append". It is a rest-server flag (--append-only). So the client is
ready and the server does not exist yet.
In front, or move the data? rest-server serves a local directory. The Storage Box is a remote share. So either (a) a machine mounts the Storage Box and runs rest-server over that mount — keeping the bytes where they are, at the cost of a mount in the hot path — or (b) the data moves to storage attached to whatever runs rest-server. These are very different prices and the choice is not obvious; (a) keeps the current bill, (b) is a migration of every customer's history.
Where would it run? ep0 is protected. Adding a service to it is an architecture change, not a
configuration change — and it also makes ep0 a single point of failure for both tiers, which is
precisely the concern §11-D already has open.
Cost: one new always-on service in the recovery path, plus whatever machine hosts it. The moving part matters more than the money: if rest-server is down, backups stop — and this project's own history (R-191) is a warning about what happens when a change in the off-site contract is not carried everywhere at once.
Q7 — What does DETECTION cost? Almost nothing. The number already arrives at the hub, and the hub already keeps the history to compare it against.
Measured from source:
- The box already reports
snapshot_count—controller/internal/report/types.go:131, and it is already rendered on the hub's Backup card (hub/internal/web/backup_card.go:118). - The hub retains report history:
store.go:968is anINSERT INTO reports, and the read isSELECT … ORDER BY id DESC LIMIT 1(store.go:1016). It appends; it does not replace. So every previoussnapshot_countfor every customer is already on disk.
So an unexplained drop is a comparison between two rows the hub already has. No new box code, no new credential, no new moving part, no new bill. The alarm vocabulary and the operator-only routing already exist.
The honest caveats: a legitimate forget also drops the count, so the rule needs to know
retention (and today the box prunes itself, which is what makes the number noisy — the same change
that disarms box-side retention is what makes this signal clean). And detection is not prevention:
it tells you within a day that history was destroyed; it does not stop it. snapshot_count: 0 must
mean UNKNOWN, not EMPTY — backup_card.go:29 records that exact defect being fixed once already.
Ranked options for Viktor
All four were considered. My recommendation is 4 + 2, in that order, and explicitly not 3 yet.
1. Do nothing — NOT acceptable as it stands, and Q1 is why
Before this spike, "do nothing" meant "exposed, but bounded to seven days by snapshots." That sentence is not supported by any evidence. The box cannot see a single snapshot, on either machine, and the register's confirmation step was never done. Doing nothing now means an unbounded exposure on the tier holding the customer's documents and photos. Cost: £0, no evenings. What happens if you choose it: the risk stays exactly where it is, and the register keeps saying "ARMED", which is the part I would not accept. One cheap act rescues most of this option: look at the Storage Box panel and answer Q1. If snapshots are real, "do nothing" becomes defensible again. That is ten minutes and it is yours to do — the fence in §11-D stopped me.
2. Copy the PBS shape — worth doing, but it is smaller than it sounds and it is not the safety it appears to be
Withdraw box-side retention; someone else prunes. Q5 says this no longer looks dangerous — a stale
lock does not wedge the store.
But Q3 says the credential still cannot be narrowed: the API has one axis, readonly, and a
backup target cannot be read-only. So the box would keep the ability to delete, and simply stop
using it. That is a discipline, not a guarantee — a compromised guest is unaffected by it.
Cost: one evening, plus wherever retention moves (each candidate concentrates the credential).
Trap: both forget sites (offbox.go:1388 and :1759) must be disarmed in the same change,
or you get R-191 again — successful backups reported as failures.
3. Change the transport (rest-server --append-only) — the only real prevention, and I would not start it yet
Q6 says it is reachable: restic 0.14.0 speaks REST. This is the option that actually makes the box unable to delete. Cost: a new always-on service in the recovery path; a machine to host it (ep0 is protected — this is an architecture change); and either a mount in the hot path or moving every customer's history. Money is the small part. Why not yet: it should be chosen against a measured picture, and one measurement is still missing — Q1. Building the expensive prevention while nobody knows whether a seven-day net already exists is the wrong order.
4. Detect instead of prevent — cheapest by a wide margin, and I would do this first
Q7 measured that the material is already there: the count is on the wire, the hub keeps the
history, the alarm path exists. An unexplained drop in snapshot_count becomes noticeable within a
day.
Cost: hub-side only. No box change, no credential change, no new service, no new bill.
What it does not do: it does not stop the deletion. It converts "we would never know" into "we
know tomorrow" — and against ransomware inside the guest, knowing tomorrow is the difference between
losing a day and losing everything silently.
This week's pattern held twice already (R-87, R-404): proving beats preventing when preventing
is expensive. This is the third instance.
My pick: answer Q1 today (yours, ten minutes), then build 4, then 2. Revisit 3 once Q1 is answered — and if Q1 comes back "no snapshots", 3 moves up sharply.