Files
felhom.eu/REPORT-r95-spike.md
admin 0476a8d8e6
gates / gates (push) Successful in 17s
SPIKE R-95: the safety net cannot be seen from the box - and that RAISES the urgency
READ-ONLY STUDY. No code, no version, no image, no golden. No delete verb was issued against any
live store. ep0, DooPlex and Peti's box were not touched at all.

Q1 FIRST, AND IT DID NOT GO THE EXPECTED WAY. The brief supposed a working seven-day snapshot net
might bound the worst case. Measured on BOTH boxes over their own SFTP credential, with a positive
and a negative control on each: NO .snapshots is visible to either sub-account - not in the account
home, not inside the repo - and the account is jailed at /. Either none exist or a sub-account
cannot see them, and the second is not a reprieve: a snapshot the box cannot see is one the box
cannot restore from, so recovery would be an operator act at the Hetzner panel.

The register's claim rests on nothing that was checked. The row has NO R-number, so nothing can cite
it; its "confirm tomorrow" was 2026-07-27, 36 days ago; and the DUE-CHECKS block built for exactly
this (R-341) is EMPTY. R-95's word "ARMED" is withdrawn pending R-429. The confirming field is a
Hetzner API field, so this spike STOPPED at the section 11-D fence and left it for Viktor - ten
minutes in the panel, and it re-ranks everything.

Q2: TEN verbs, not nine. `check` was missing from the brief's list; `dump` is not a verb (it is a
progress phase constant) and was withdrawn. There are TWO `forget --prune` sites - offbox.go:1388
AND offbox.go:1759 - and disarming one without the other reproduces R-191 exactly.

Q3 (documented, from our own API mirror): AccessSettings has five booleans and `readonly` is the
only permission axis. No append-only. So the PBS shape does NOT transfer - PBS is a server that can
refuse; a Storage Box is a filesystem that runs nothing.

Q5 (measured, faithful append-only model, both controls passed first): withdrawing delete does NOT
wedge the store - restic treats a dead owner's lock as stale and proceeds. The constraint everyone
feared is not the blocker. But `unlock --remove-all` printed "successfully removed locks" while the
lock was still there, and resticStep's crash-lock self-heal is built on that call - R-430.

Q6 (measured, with a control): restic 0.14.0 DOES speak rest:. Append-only is a rest-server flag,
not a restic one.

Q7 (measured): detection is nearly free. snapshot_count already reaches the hub and the hub APPENDS
reports, so the history to compare against is already on disk.

RECOMMENDATION: answer Q1 today (Viktor, ten minutes), then build detection, then move retention off
the box. Defer the transport change until Q1 is answered.

Register: OPEN 179 -> 181. Filed R-429, R-430; R-95 updated and kept OPEN. 07 row 10's status is
deliberately UNCHANGED.
2026-09-01 13:55:35 +02:00

10 KiB
Raw Permalink Blame History

REPORT — SPIKE R-95: can the box be stopped from deleting its own off-site history? (2026-09-01)

Q1 first, and it raises the urgency rather than lowering it

The safety net cannot be seen from the box, on either machine, so the seven-day bound is unverified — and unverifiable from the product side.

Measured over each box's own SFTP credential, read-only, with controls that passed first:

probe demo-hp (sub-account A) demo-felhom
positive control — account home lists .ssh, <repo> lists .ssh, <repo>, <repo>.orphaned-20260810
negative control — bogus name not found not found
./.snapshots not found not found
<repo>/.snapshots not found not found
/ Permission denied (jailed) same

Either no snapshots exist, or a sub-account cannot see them. The box cannot tell which, and neither is a reprieve — a snapshot the box cannot see is one it cannot restore from, so recovery is an operator act at the Hetzner panel, not a product capability.

The register's claim rests on nothing that was ever checked. OPEN-ITEMS.md:233 is a row with no R-number, its "confirm tomorrow" was 2026-07-27 (36 days), and the DUE-CHECKS block built for exactly this (R-341) is empty. R-95's own text says "Mitigation now ARMED"; that word is withdrawn pending R-429.

The task expected Q1 might bound the exposure to seven days. It does not. I stopped at the §11-D fence rather than answering it with the provider token.

Q1–Q7, one answer each

Q answer
Q1 NO / unknowable from the box. Measured on both machines with controls. → R-429
Q2 Ten verbs, not nine, and TWO forget --prune sites. check (offbox_integrity.go:316) is missing from the task's list; dump is not a verb (offbox_progress.go:185 is a phase constant) — withdrawn. Delete-capable: forget/prune (offbox.go:1388 and :1759), unlock (:746, :768).
Q3 No. DOCUMENTED from hub/internal/hetznerapi/hetznerapi.go:38-45: AccessSettings has five booleans and readonly is the only permission axis. No append-only. Corroborated by existing state, no new action: demo-felhom's home holds a <repo>.orphaned-20260810 directory produced by a controller rename, which is delete-class.
Q4 The PBS shape does not transfer. PBS is a server that can refuse; a Storage Box is a filesystem that runs nothing, so there is no far end to move retention to. Four candidates costed; each concentrates a delete-capable credential. R-191's trap is doubled here — both forget sites must be disarmed in the same change or every successful backup reports failure.
Q5 NOT the blocker — measured. Under a faithful append-only model (both controls passed), a stale lock does not wedge the store: backup, check, snapshots --no-lock and restore --no-lock all succeeded. But unlock --remove-all printed successfully removed locks while the lock survived → R-430. The crash-lock (foreign hostname) window is UNKNOWN.
Q6 Reachable. MEASURED: rest: gives a connection error where the control banana: gives invalid backend — restic 0.14.0 speaks REST. Append-only is a rest-server flag; restic help contains zero occurrences of "append". Needs a machine in the recovery path (ep0 is protected — architecture change) and either a mount in the hot path or moving every customer's history.
Q7 Nearly free. MEASURED: snapshot_count already reaches the hub (report/types.go:131 → backup_card.go:118) and the hub APPENDS reports (store.go:968 INSERT; read is ORDER BY id DESC LIMIT 1), so the history to compare against is already on disk. No box change, no credential, no new service.

The four options, ranked

My pick: answer Q1 today (Viktor, ten minutes) → build 4 → then 2. Defer 3.

  1. Do nothing — not acceptable as it stands. It used to mean "bounded to seven days". Q1 shows that sentence is unsupported. £0, no evenings, and an unbounded exposure on the tier holding the customer's documents.
  2. Copy the PBS shape — worth doing, smaller than it sounds. Q5 removed the fear that it would wedge the store. But Q3 means the credential still can delete; the box would merely stop using it. That is discipline, not a guarantee, and a compromised guest is unaffected. One evening + a new home for retention.
  3. Change the transport — the only real prevention, and not yet. Q6 says it is reachable. It costs an always-on service in the recovery path and a protected machine's architecture. Choosing it before Q1 is answered is the wrong order.
  4. Detect instead of prevent — cheapest by a wide margin, do this first. Q7 measured that the material already exists. Converts "we would never know" into "we know tomorrow". Third instance this week of proving beats preventing when preventing is expensive.

What I could not measure, and what would settle it

unknown what would settle it
Do snapshots exist on the Storage Box? the Hetzner panel or size_snapshots via the API — fenced by §11-D; stopped and left for Viktor (R-429)
Whether the live API exposes any permission the Go struct omits the provider token — same fence
The crash-lock case (a lock whose hostname restic cannot match, non-stale for ~30 min) a lock captured from a container with a different hostname, replayed against an append-only endpoint. My model could not build it — the captured lock carried this container's own hostname
Whether unlock reports success because it removed zero locks by design, or because it never checked read restic 0.14.0's unlock source, or re-run with --verbose
Whether a Storage Box snapshot can be restored deliberately not attempted — named as the next question, per the brief

Register

Before: OPEN 179 · CLOSED 161. After: OPEN 181 · CLOSED 161. Filed R-429 (the unconfirmed snapshot mitigation, and the id-less row), R-430 (unlock lies about success). R-95 updated with the verdict and kept OPEN; its word "ARMED" withdrawn. Docs: 07 §8 row 10 and §10.2 gained the verdict (row 10's status deliberately NOT moved); §11-D records that the fence was reached again and held; STATUS.md carries one plain-language item.

Compliance

  • No code changed. No version bumped. No image built. No golden owed. golden_currency_gate.py exits 0; golden and floor remain 0.232.0.
  • No delete verb was issued against any live store — no forget, prune, unlock or init. Every live-store interaction was an SFTP ls.
  • ep0, DooPlex and Peti's box were not touched at all, not even read — the two demo boxes were the only machines used.
  • The Hetzner API and control panel were not called.
  • Scratch resources: a throwaway local restic repo under /tmp/r95s (and /tmp/r95scratch in the first attempt) inside the controller container on demo-hp, 60 MB, plus three probe scripts. All removed, verified by the scripts' own teardown output (scratch: gone). Nothing was created on any Storage Box.

Observations, and my own mistakes by name

  1. The register calls a mitigation ARMED that has never been confirmed, in a row that cannot be cited because it has no id, with the dated-check mechanism sitting empty beside it. FILED: R-429.
  2. restic unlock --remove-all reports success on a deletion that did not happen, and the crash-lock self-heal is built on it. FILED: R-430.
  3. The task's own verb list was missing check and included dump, which is not a verb, and it names one forget --prune site where there are two. NOT-A-FINDING: the brief invited me to confirm the list myself, which is what this is; both corrections are in the spike document and the second one is carried into R-95's row, because disarming one site and not the other reproduces R-191 exactly.
  4. My mistake — my first Q5 model proved nothing. I used chmod a-w and ran restic as root, which ignores permission bits, so every verb succeeded and I nearly recorded "no wedge" on a test that tested nothing. Rebuilt as a sticky directory with a root-owned lock and restic run as nobody, with two controls. NOT-A-FINDING: caught inside the same session by the result being too clean; the corrected model is the one reported, and the first is described so nobody repeats it.
  5. My mistake — my first sftp probe used -p for the port, which sftp reads as "preserve", so the port became the destination and all three probes returned identical usage errors. The controls are what exposed it — a positive and a negative control failing the same way is an instrument fault, not a result. NOT-A-FINDING: a flag error of mine, corrected in one command; it is recorded because the failure mode it demonstrates — three identical errors reading as three findings — is the one this project keeps paying for.
  6. My mistake — I wrote the §8 verdict onto row 4 instead of row 10. The anchor text I matched appears in both rows and I replaced the first occurrence. Caught by checking the line number, reverted from row 4 and applied to row 10, both verified by grep. NOT-A-FINDING: an editing error of mine, corrected within the session and verified in both directions, so no wrong claim ever reached a push.
  7. My mistake — I put two escaped pipes inside a register row, which makes it a five-column row in a three-column table and would have made it unreadable to closed_register_gate.py — the exact defect I fixed in that gate yesterday. Caught by counting pipes before committing. NOT-A-FINDING: corrected before the push; recorded because I introduced the same shape twice in two days.
  8. The task's baseline table lists felhom-agent at 058b945; it is at 4586f0f. NOT-A-FINDING: that is my own push from the previous session, so the table was stale rather than wrong about anything that matters here; the agent repo was not touched by this spike at all.