From 0476a8d8e6c8e9a80f3c5e05b4ddef599ed10515 Mon Sep 17 00:00:00 2001 From: kisfenyo Date: Tue, 1 Sep 2026 13:55:35 +0200 Subject: [PATCH] SPIKE R-95: the safety net cannot be seen from the box - and that RAISES the urgency READ-ONLY STUDY. No code, no version, no image, no golden. No delete verb was issued against any live store. ep0, DooPlex and Peti's box were not touched at all. Q1 FIRST, AND IT DID NOT GO THE EXPECTED WAY. The brief supposed a working seven-day snapshot net might bound the worst case. Measured on BOTH boxes over their own SFTP credential, with a positive and a negative control on each: NO .snapshots is visible to either sub-account - not in the account home, not inside the repo - and the account is jailed at /. Either none exist or a sub-account cannot see them, and the second is not a reprieve: a snapshot the box cannot see is one the box cannot restore from, so recovery would be an operator act at the Hetzner panel. The register's claim rests on nothing that was checked. The row has NO R-number, so nothing can cite it; its "confirm tomorrow" was 2026-07-27, 36 days ago; and the DUE-CHECKS block built for exactly this (R-341) is EMPTY. R-95's word "ARMED" is withdrawn pending R-429. The confirming field is a Hetzner API field, so this spike STOPPED at the section 11-D fence and left it for Viktor - ten minutes in the panel, and it re-ranks everything. Q2: TEN verbs, not nine. `check` was missing from the brief's list; `dump` is not a verb (it is a progress phase constant) and was withdrawn. There are TWO `forget --prune` sites - offbox.go:1388 AND offbox.go:1759 - and disarming one without the other reproduces R-191 exactly. Q3 (documented, from our own API mirror): AccessSettings has five booleans and `readonly` is the only permission axis. No append-only. So the PBS shape does NOT transfer - PBS is a server that can refuse; a Storage Box is a filesystem that runs nothing. Q5 (measured, faithful append-only model, both controls passed first): withdrawing delete does NOT wedge the store - restic treats a dead owner's lock as stale and proceeds. The constraint everyone feared is not the blocker. But `unlock --remove-all` printed "successfully removed locks" while the lock was still there, and resticStep's crash-lock self-heal is built on that call - R-430. Q6 (measured, with a control): restic 0.14.0 DOES speak rest:. Append-only is a rest-server flag, not a restic one. Q7 (measured): detection is nearly free. snapshot_count already reaches the hub and the hub APPENDS reports, so the history to compare against is already on disk. RECOMMENDATION: answer Q1 today (Viktor, ten minutes), then build detection, then move retention off the box. Defer the transport change until Q1 is answered. Register: OPEN 179 -> 181. Filed R-429, R-430; R-95 updated and kept OPEN. 07 row 10's status is deliberately UNCHANGED. --- REPORT-r95-spike.md | 127 ++++++++ STATUS.md | 19 +- .../architecture/07-backup-architecture.md | 15 +- .../SPIKE-r95-offsite-delete-2026-09-01.md | 294 ++++++++++++++++++ documentation/backlog/OPEN-ITEMS.md | 4 +- 5 files changed, 452 insertions(+), 7 deletions(-) create mode 100644 REPORT-r95-spike.md create mode 100644 documentation/audits/SPIKE-r95-offsite-delete-2026-09-01.md diff --git a/REPORT-r95-spike.md b/REPORT-r95-spike.md new file mode 100644 index 00000000..4bb271f6 --- /dev/null +++ b/REPORT-r95-spike.md @@ -0,0 +1,127 @@ +# REPORT — SPIKE R-95: can the box be stopped from deleting its own off-site history? (2026-09-01) + +## Q1 first, and it raises the urgency rather than lowering it + +**The safety net cannot be seen from the box, on either machine, so the seven-day bound is +unverified — and unverifiable from the product side.** + +Measured over each box's own SFTP credential, read-only, with controls that passed first: + +| probe | demo-hp (sub-account A) | demo-felhom | +|---|---|---| +| positive control — account home | lists `.ssh`, `` | lists `.ssh`, ``, `.orphaned-20260810` | +| negative control — bogus name | `not found` | `not found` | +| `./.snapshots` | **`not found`** | **`not found`** | +| `/.snapshots` | `not found` | `not found` | +| `/` | `Permission denied` (jailed) | same | + +**Either no snapshots exist, or a sub-account cannot see them. The box cannot tell which, and neither +is a reprieve** — a snapshot the box cannot see is one it cannot restore from, so recovery is an +operator act at the Hetzner panel, not a product capability. + +**The register's claim rests on nothing that was ever checked.** `OPEN-ITEMS.md:233` is a row with +**no R-number**, its "confirm tomorrow" was **2026-07-27** (36 days), and the `DUE-CHECKS` block built +for exactly this (R-341) is **empty**. R-95's own text says "Mitigation now ARMED"; that word is +**withdrawn** pending R-429. + +**The task expected Q1 might bound the exposure to seven days. It does not.** I stopped at the §11-D +fence rather than answering it with the provider token. + +## Q1–Q7, one answer each + +| Q | answer | +|---|---| +| **Q1** | **NO / unknowable from the box.** Measured on both machines with controls. → R-429 | +| **Q2** | **Ten verbs, not nine, and TWO `forget --prune` sites.** `check` (`offbox_integrity.go:316`) is missing from the task's list; `dump` is not a verb (`offbox_progress.go:185` is a phase constant) — withdrawn. Delete-capable: `forget`/`prune` (`offbox.go:1388` **and `:1759`**), `unlock` (`:746`, `:768`). | +| **Q3** | **No. DOCUMENTED** from `hub/internal/hetznerapi/hetznerapi.go:38-45`: `AccessSettings` has five booleans and **`readonly` is the only permission axis**. No append-only. Corroborated by existing state, no new action: demo-felhom's home holds a `.orphaned-20260810` directory produced by a controller **rename**, which is delete-class. | +| **Q4** | **The PBS shape does not transfer.** PBS is a server that can refuse; a Storage Box is a filesystem that **runs nothing**, so there is no far end to move retention to. Four candidates costed; each concentrates a delete-capable credential. **R-191's trap is doubled here** — both `forget` sites must be disarmed in the same change or every successful backup reports failure. | +| **Q5** | **NOT the blocker — measured.** Under a faithful append-only model (both controls passed), a stale lock does **not** wedge the store: `backup`, `check`, `snapshots --no-lock` and `restore --no-lock` all succeeded. **But `unlock --remove-all` printed `successfully removed locks` while the lock survived** → R-430. The crash-lock (foreign hostname) window is **UNKNOWN**. | +| **Q6** | **Reachable. MEASURED:** `rest:` gives a connection error where the control `banana:` gives `invalid backend` — **restic 0.14.0 speaks REST**. Append-only is a **rest-server** flag; `restic help` contains zero occurrences of "append". Needs a machine in the recovery path (**ep0 is protected — architecture change**) and either a mount in the hot path or moving every customer's history. | +| **Q7** | **Nearly free. MEASURED:** `snapshot_count` already reaches the hub (`report/types.go:131` → `backup_card.go:118`) and **the hub APPENDS reports** (`store.go:968` INSERT; read is `ORDER BY id DESC LIMIT 1`), so the history to compare against is already on disk. No box change, no credential, no new service. | + +## The four options, ranked + +**My pick: answer Q1 today (Viktor, ten minutes) → build 4 → then 2. Defer 3.** + +1. **Do nothing — not acceptable as it stands.** It used to mean "bounded to seven days". Q1 shows + that sentence is unsupported. £0, no evenings, and an unbounded exposure on the tier holding the + customer's documents. +2. **Copy the PBS shape — worth doing, smaller than it sounds.** Q5 removed the fear that it would + wedge the store. But Q3 means the credential still *can* delete; the box would merely stop using + it. That is discipline, not a guarantee, and a compromised guest is unaffected. One evening + a new + home for retention. +3. **Change the transport — the only real prevention, and not yet.** Q6 says it is reachable. It costs + an always-on service in the recovery path and a protected machine's architecture. Choosing it + before Q1 is answered is the wrong order. +4. **Detect instead of prevent — cheapest by a wide margin, do this first.** Q7 measured that the + material already exists. Converts "we would never know" into "we know tomorrow". Third instance + this week of *proving beats preventing when preventing is expensive*. + +## What I could not measure, and what would settle it + +| unknown | what would settle it | +|---|---| +| **Do snapshots exist on the Storage Box?** | the Hetzner panel or `size_snapshots` via the API — **fenced by §11-D; stopped and left for Viktor (R-429)** | +| Whether the live API exposes any permission the Go struct omits | the provider token — **same fence** | +| The **crash-lock** case (a lock whose hostname restic cannot match, non-stale for ~30 min) | a lock captured from a container with a **different hostname**, replayed against an append-only endpoint. My model could not build it — the captured lock carried this container's own hostname | +| Whether `unlock` reports success because it removed zero locks by design, or because it never checked | read restic 0.14.0's unlock source, or re-run with `--verbose` | +| Whether a Storage Box snapshot can be **restored** | deliberately not attempted — named as the next question, per the brief | + +## Register + +**Before:** OPEN 179 · CLOSED 161. **After:** OPEN 181 · CLOSED 161. +**Filed R-429** (the unconfirmed snapshot mitigation, and the id-less row), **R-430** (`unlock` lies +about success). **R-95 updated** with the verdict and kept OPEN; its word "ARMED" withdrawn. +**Docs:** `07` §8 row 10 and §10.2 gained the verdict (**row 10's status deliberately NOT moved**); +§11-D records that the fence was reached again and held; `STATUS.md` carries one plain-language item. + +## Compliance + +- **No code changed. No version bumped. No image built. No golden owed.** `golden_currency_gate.py` + exits 0; golden and floor remain **0.232.0**. +- **No delete verb was issued against any live store** — no `forget`, `prune`, `unlock` or `init`. + Every live-store interaction was an SFTP `ls`. +- **`ep0`, DooPlex and Peti's box were not touched at all**, not even read — the two demo boxes were + the only machines used. +- **The Hetzner API and control panel were not called.** +- **Scratch resources:** a throwaway local restic repo under `/tmp/r95s` (and `/tmp/r95scratch` in the + first attempt) inside the controller container on demo-hp, 60 MB, plus three probe scripts. **All + removed**, verified by the scripts' own teardown output (`scratch: gone`). Nothing was created on + any Storage Box. + +## Observations, and my own mistakes by name + +1. **The register calls a mitigation ARMED that has never been confirmed, in a row that cannot be + cited because it has no id, with the dated-check mechanism sitting empty beside it.** + **FILED: R-429.** +2. **`restic unlock --remove-all` reports success on a deletion that did not happen**, and the + crash-lock self-heal is built on it. **FILED: R-430.** +3. **The task's own verb list was missing `check` and included `dump`, which is not a verb**, and it + names one `forget --prune` site where there are two. **NOT-A-FINDING: the brief invited me to + confirm the list myself, which is what this is; both corrections are in the spike document and the + second one is carried into R-95's row, because disarming one site and not the other reproduces + R-191 exactly.** +4. **My mistake — my first Q5 model proved nothing.** I used `chmod a-w` and ran restic as **root**, + which ignores permission bits, so every verb succeeded and I nearly recorded "no wedge" on a test + that tested nothing. Rebuilt as a sticky directory with a root-owned lock and restic run as + `nobody`, with two controls. **NOT-A-FINDING: caught inside the same session by the result being + too clean; the corrected model is the one reported, and the first is described so nobody repeats + it.** +5. **My mistake — my first `sftp` probe used `-p` for the port**, which `sftp` reads as "preserve", so + the port became the destination and all three probes returned identical usage errors. **The + controls are what exposed it** — a positive and a negative control failing the same way is an + instrument fault, not a result. **NOT-A-FINDING: a flag error of mine, corrected in one command; + it is recorded because the failure mode it demonstrates — three identical errors reading as three + findings — is the one this project keeps paying for.** +6. **My mistake — I wrote the §8 verdict onto row 4 instead of row 10.** The anchor text I matched + appears in both rows and I replaced the first occurrence. Caught by checking the line number, + reverted from row 4 and applied to row 10, both verified by grep. **NOT-A-FINDING: an editing error + of mine, corrected within the session and verified in both directions, so no wrong claim ever + reached a push.** +7. **My mistake — I put two escaped pipes inside a register row**, which makes it a five-column row in + a three-column table and would have made it unreadable to `closed_register_gate.py` — **the exact + defect I fixed in that gate yesterday.** Caught by counting pipes before committing. **NOT-A-FINDING: + corrected before the push; recorded because I introduced the same shape twice in two days.** +8. **The task's baseline table lists `felhom-agent` at `058b945`; it is at `4586f0f`.** + **NOT-A-FINDING: that is my own push from the previous session, so the table was stale rather than + wrong about anything that matters here; the agent repo was not touched by this spike at all.** diff --git a/STATUS.md b/STATUS.md index 94a961cb..7b0ddcb4 100644 --- a/STATUS.md +++ b/STATUS.md @@ -18,7 +18,7 @@ not an evening's work.** *This section is allowed to be longer than one screen, and each item says what happens if you do nothing.* -1. **Nothing is waiting on you.** Both problems the overnight test found are fixed and proven on +1. **One thing is waiting on you — item 4, and it takes ten minutes.** Otherwise: Both problems the overnight test found are fixed and proven on the real machines: - the background job that could delete a live restore's lock now waits its turn — and the check that finds the next one like it is a test, not a comment, so it cannot come back quietly; @@ -37,11 +37,24 @@ nothing.* two register lines in the hub (already live). No customer action, no data migration, no credential change. -4. **Whether to change the hub password** (R-350). I printed it into my own session log on 20 August. +4. **The copy that holds the customers' documents and photos can still be deleted by the box that + made it** (R-95 — first on the list since July, and this is the first time it has reached this + page). I studied it today and did not change anything. **One thing is yours and it takes ten + minutes.** Our notes say a safety net is switched on: the storage provider takes a snapshot every + night and keeps seven. **I could not find a single one.** I looked from both machines, using the + credentials they already have, and checked that my method could see other things and could + correctly fail to see a made-up name. Either the snapshots are not being taken, or they are + invisible to the machines — and if they are invisible, they are also useless to them: getting one + back would be you, in the provider's control panel. **I did not log in to check, because your own + notes say that question is yours.** **If you do nothing:** the register keeps saying the net is + armed, and nobody knows whether it is. **Please open the Storage Box panel and tell me whether + snapshots exist.** The answer changes which fix is worth building. + +5. **Whether to change the hub password** (R-350). I printed it into my own session log on 20 August. Not in git, not in any saved file — in the log on this machine. **If you do nothing:** it stays as it is, at the risk you accept by leaving it. I can change it without ever showing you the new one. -5. **`demo-hp`'s network setup does not match our own notes** (R-338) — the machine works, the page is +6. **`demo-hp`'s network setup does not match our own notes** (R-338) — the machine works, the page is wrong, or the other way round. **If you do nothing:** the page keeps misleading the next session, as it misled one by an hour. diff --git a/documentation/architecture/07-backup-architecture.md b/documentation/architecture/07-backup-architecture.md index cb6be913..c7aec04e 100644 --- a/documentation/architecture/07-backup-architecture.md +++ b/documentation/architecture/07-backup-architecture.md @@ -98,7 +98,7 @@ with nothing but their dashboard password. No operator, no ticket, no scheduling | 1 | „Visszaállítás indítása" | `POST /backup/restore` | destructive — rebuilds the app from its Tier-1 unit | | 2 | „Fájlok visszaállítása" | `POST /backup/tier2/restore` | additive, missing-only | | 3 | „Visszaállítás a távoli tárolóból" | `POST /backup/offbox/restore` | non-destructive, to a verification copy | -| 4 | „Helyreállítás az élő adatok közé" | `POST /backup/offbox/place` | additive, missing-only **2026-08-30 (R-359): the store is now VERIFIED on a cadence** — a daily `offsite-integrity` job runs `restic check` when the last successful one is over 7 days old. **This is a readability check, not a restore-test (R-87 stays open).** The depth that ships ON now re-reads **100%** of the pack data (R-399 CLOSED, controller v0.228.0) — the sentence here previously said the opposite and was stale. **2026-08-31 (SPIKE R-87):** a scratch restore of every app on demo-hp was measured at **25 s for 8 snapshots / 774 MB**, against 40.3 s for one weekly check — but restic 0.14.0's `--verify` checks size and mtime, **not content**, so nothing available today can vouch for the restored BYTES. Row 4's verdict is UNCHANGED by the spike. `audits/SPIKE-restic-restore-test-2026-08-31.md`. | +| 4 | „Helyreállítás az élő adatok közé" | `POST /backup/offbox/place` | additive, missing-only **2026-08-30 (R-359): the store is now VERIFIED on a cadence** — a daily `offsite-integrity` job runs `restic check` when the last successful one is over 7 days old. **This is a readability check, not a restore-test (R-87 stays open).** The depth that ships ON now re-reads **100%** of the pack data (R-399 CLOSED, controller v0.228.0) — the sentence here previously said the opposite and was stale. **2026-08-31 (SPIKE R-87):** a scratch restore of every app on demo-hp was measured at **25 s for 8 snapshots / 774 MB**, against 40.3 s for one weekly check — but restic 0.14.0's `--verify` checks size and mtime, **not content**, so nothing available today can vouch for the restored BYTES. Row 4's verdict is UNCHANGED by the spike. `audits/SPIKE-restic-restore-test-2026-08-31.md`. `audits/SPIKE-restic-restore-test-2026-08-31.md`. | | 5 | „Teljes visszaállítás (fájlok + adatbázis)" | `POST /backup/offbox/reconstitute` | destructive to files + DB, never deleting | | 6 | „Megosztások visszaállítása" + place | `POST /backup/shares/{restore,place}` | additive | | 7 | `.fab` import | `POST` → `apiImportStart` | destructive re-import of one app | @@ -885,7 +885,7 @@ crosses the line — **R-158**. | 7 | **Guest stopped and does not come back** | everything | guest-power watchdog (60 s, `onboot` as the deliberate-stop discriminator) | automatic | **120 s** | | **PROVEN** | agent v0.107.0 replay — 120 s unattended vs the incident's 587 s with a human | | 8 | **Host dies (hardware), drives intact** | the data drives; ep0's PBS namespace; the hub's Recipe + Escrow | **install a new host, then** `--selftest=bring-up -mode dr` per guest, **then** re-attach drives by `durable_id` | **operator** (SSH) | | 7 d | **IMPLEMENTED** | bring-up code exists and has **never been executed** (`CAMPAIGN-8…:522`); the drive half of the plan is empty on every box (**R-105**) | | 9 | **Whole box lost (fire/theft) — host and drives gone** | ep0 PBS namespace; the restic repo; the hub's Recipe + Escrow | new hardware → day-0 → escrow-consume with **R** → restore guests from PBS → app data from Tier-3 | **operator + customer** (R) | | 7 d (guest) · 24 h (app data) | **IMPLEMENTED / UNPROVEN** | every leg exists; the composed path has never been run. The destructive S5 drill is operator-gated and unrun (`06-offsite-connectivity.md:327`) | -| 10 | **Ransomware / malicious deletion inside the guest** | PBS offsite (the box **cannot** delete its own snapshots); **the restic repo is NOT protected the same way** | whole-guest restore from PBS to a point before the event | **operator** (SSH) | | 7 d | **PARTIAL** | R-89 proved the box is refused when deleting its own PBS snapshot (`CAMPAIGN-8…:514-515`). **R-95 (open, ranked #1):** the restic credential **can delete** — `readonly=False`, `forget --prune` runs from the box, and SFTP cannot express append-only (`OPEN-ITEMS.md:13`) **2026-08-30 (R-359): the store is now VERIFIED on a cadence** — a daily `offsite-integrity` job runs `restic check` when the last successful one is over 7 days old. **This is a readability check, not a restore-test (R-87 stays open).** The depth that ships ON now re-reads **100%** of the pack data (R-399 CLOSED, controller v0.228.0) — the sentence here previously said the opposite and was stale. **2026-08-31 (SPIKE R-87):** a scratch restore of every app on demo-hp was measured at **25 s for 8 snapshots / 774 MB**, against 40.3 s for one weekly check — but restic 0.14.0's `--verify` checks size and mtime, **not content**, so nothing available today can vouch for the restored BYTES. Row 4's verdict is UNCHANGED by the spike. `audits/SPIKE-restic-restore-test-2026-08-31.md`. | +| 10 | **Ransomware / malicious deletion inside the guest** | PBS offsite (the box **cannot** delete its own snapshots); **the restic repo is NOT protected the same way** | whole-guest restore from PBS to a point before the event | **operator** (SSH) | | 7 d | **PARTIAL** | R-89 proved the box is refused when deleting its own PBS snapshot (`CAMPAIGN-8…:514-515`). **R-95 (open, ranked #1):** the restic credential **can delete** — `readonly=False`, `forget --prune` runs from the box, and SFTP cannot express append-only (`OPEN-ITEMS.md:13`) **2026-08-30 (R-359): the store is now VERIFIED on a cadence** — a daily `offsite-integrity` job runs `restic check` when the last successful one is over 7 days old. **This is a readability check, not a restore-test (R-87 stays open).** The depth that ships ON now re-reads **100%** of the pack data (R-399 CLOSED, controller v0.228.0) — the sentence here previously said the opposite and was stale. **2026-08-31 (SPIKE R-87):** a scratch restore of every app on demo-hp was measured at **25 s for 8 snapshots / 774 MB**, against 40.3 s for one weekly check — but restic 0.14.0's `--verify` checks size and mtime, **not content**, so nothing available today can vouch for the restored BYTES. Row 4's verdict is UNCHANGED by the spike. `audits/SPIKE-restic-restore-test-2026-08-31.md`. **2026-09-01 (SPIKE R-95, `audits/SPIKE-r95-offsite-delete-2026-09-01.md`) — THIS ROW'S STATUS IS UNCHANGED, but the seven-day bound beside it is now known to be unverified.** Measured on BOTH boxes over their own SFTP credential, with controls: **no `.snapshots` is visible to either sub-account**, and the account is jailed. Either none exist or a sub-account cannot see them — and a snapshot the box cannot see is one the box cannot restore from, so recovery would be an operator act at the Hetzner panel, not a product capability (R-429; the confirming field needs the provider API, fenced by §11-D). **The PBS shape does NOT transfer**: the sub-account API has one permission axis, `readonly`, and a backup target cannot be read-only — PBS is a server that can refuse, a Storage Box is a filesystem that runs nothing. **What the spike removed as a fear:** withdrawing delete does NOT wedge the store (measured — restic treats a dead owner's lock as stale and proceeds). **What it added:** `unlock --remove-all` reports success while deleting nothing (R-430). | | 11 | **Hub lost** | every box's data plane, every tier, every Lane-1 route | none needed for recovery **of a box**; the hub itself restores from its Longhorn volume backup | operator (`kubectl`) | | 24 h + weekly (Longhorn `RETAIN 1` each) | **UNPROVEN** | a hub restore has never been performed. The backup target is `nfs://192.168.0.180` — **DooPlex itself** — and exactly **2** restore points exist (LIVE, INV Part D2.2) | | 11b | *consequences while the hub is gone* | — | — | — | | | **[FACT]** | **day:** nothing customer-visible breaks; events queue (`settings.go:1466-1485`). **week:** the operator alarm plane is dark, no claim/reset codes, no config or floor convergence, no PBS-secret re-issue. **permanently:** escrow custody and break-glass credentials are gone (INV Part D2.4) | | 12 | **Offsite provider lost (Hetzner)** | everything on-premises: both drives, both whole-guest tiers | none needed — on-premises recovery is unaffected. Re-provision a new offsite target. | operator | | | **[FACT]** | **restic and PBS share the provider** (INV Part E.1). Whether they share an account and payment method is **UNKNOWN** → §11-D | @@ -1085,7 +1085,7 @@ does **not** hold as written. → **R-108** | **R-356** | The off-site restore resolved its destination with the raw `HDD_PATH` and read an empty answer as "not installed" — **CLOSED, controller v0.219.0, 2026-08-22** | 40 of 53 apps were refused permanently while running (§6.3 `[DESIGN]`) | | ~~**R-108**~~ | ~~Network storage can host an app's namespace~~ | **CLOSED 2026-07-30, controller v0.187.0 — D5 UNBLOCKED.** An app namespace may no longer be placed on network storage (5 surfaces guarded by one fail-closed predicate); the share-root bind is deliberately UNCHANGED because it is load-bearing and unscopable (§10.1). `audits/R108-network-app-namespace-2026-07-30.md` | | **R-126** | A `.fab` bundle — plaintext secrets, optional password — can be exported ONTO a NAS: `storageDriveList()` (`internal/web/handler_export.go`) does not filter network paths | split out of R-108, which closed without it. NOT a D5 precondition: an explicit customer-chosen export destination, not a browsing surface reaching a backup tree (§5, §7.3) | -| R-95 (open) | The restic offsite credential **can delete** — the box can `forget --prune` its own repo | the tier holding the customer's documents and photos is the one whose credential can destroy it (matrix row 10) | +| R-95 (open) | The restic offsite credential **can delete** — the box can `forget --prune` its own repo, from **two** call sites (`offbox.go:1388` retention and `offbox.go:1759` over-quota) | the tier holding the customer's documents and photos is the one whose credential can destroy it (matrix row 10). **SPIKE 2026-09-01:** prevention needs a transport change (restic 0.14.0 does speak `rest:` — measured; append-only is a rest-server flag, not a restic one), because the sub-account API cannot express write-without-delete. **Detection is nearly free and is recommended first** — `snapshot_count` already reaches the hub and the hub appends reports, so the comparison needs no box change. `audits/SPIKE-r95-offsite-delete-2026-09-01.md` | | ~~R-86~~ | ~~Restore-tests are interval-scheduled, not backup-aligned~~ | **CLOSED 2026-08-03 — agent v0.121.0 + hub v0.91.0.** Restore-testing is now **per archive generation**: a tier is due when its newest archive that has settled ~24 h has not been proven, so a daily tier is proved daily on its own archive and a weekly tier weekly on its own. The ticker survives only as the evaluation interval (6 h, chosen from a measured cost). The hub's staleness window moved with it — per tier, from that tier's observed archive rhythm — because a weekly tier proved weekly sat EXACTLY on the old flat 7-day line (§3, Lane 2's per-archive rule) | | ~~R-353~~ | ~~A local unit restore reported a bare completion whether it returned an entire dataset or nothing~~ | **CLOSED 2026-08-30, controller v0.226.0.** The volume count already existed (`restoreDockerVolumesFrom`) and was discarded by a one-line wrapper, so the surface was structurally unable to say what came back. `RestoreFromRecoveryUnit` now returns `UnitRestoreResult` and the sentence has three cases, keyed on replayed-vs-**listed** — zero-replayed has two causes that are opposite news. Every sentence is a claim about the BACKUP, never the app: this path has no `SafetyDump` discriminator, and §6.3 is why that is not pedantry. Proven live on `demo-hp` | | ~~R-357~~ | ~~The destructive reconstitute had no free-space gate; all three that existed guarded non-destructive paths~~ | **CLOSED 2026-08-30, controller v0.226.0.** The gate sits before `writeSafetyDump` and `StopStack`, so a refusal costs no outage — the test asserts the StopStack call count, not the sentence. No headroom multiplier (a measured tree, not a predicted download); fail-closed when either probe reads ≤ 0, which was a real fail-open hole. **Not live-validated** — filling a filesystem is a drill step | @@ -1129,6 +1129,15 @@ and PBS (Cloud server ep0 + a Cloud Volume) are both Hetzner. Whether they share login and payment method** is unverified (INV Unknown U-1) — the question was deliberately not answered by calling the provider API with the production token. Accept explicitly, or mitigate. +> **The fence was reached again, and held, 2026-09-01 (SPIKE R-95).** The R-95 spike needed one +> provider-side fact — **do daily snapshots exist on the Storage Box, and how many** — because the +> register calls that mitigation ARMED and the whole ranking of R-95's options turns on it. It is +> readable from the box: measured on BOTH machines over their own SFTP credential, with positive +> and negative controls, **no `.snapshots` is visible to either sub-account**, and the account is +> jailed at `/`. The register's own confirming field, `size_snapshots`, is an API field. **So the +> spike stopped here rather than answering it by acting**, and left it as R-429 — ten minutes in +> the Storage Box panel, and it re-ranks R-95. `audits/SPIKE-r95-offsite-delete-2026-09-01.md` §Q1. + **E. `local` vzdump shares a physical device with the guest it backs up.** LIVE on both hosts: `/var/lib/vz` (the archive target) and `local-lvm` (the guest's rootfs and both `backup=1` mountpoints) are both on `/dev/sda3` → VG `pve`. The tier therefore protects against **corruption diff --git a/documentation/audits/SPIKE-r95-offsite-delete-2026-09-01.md b/documentation/audits/SPIKE-r95-offsite-delete-2026-09-01.md new file mode 100644 index 00000000..46e1ab99 --- /dev/null +++ b/documentation/audits/SPIKE-r95-offsite-delete-2026-09-01.md @@ -0,0 +1,294 @@ +# SPIKE — can the box be stopped from deleting its own off-site history? (R-95, 2026-09-01) + +**Read-only study. No production code, no version bump, no image, no golden. No delete verb was +issued against any live store. `ep0`, DooPlex and Peti's box were not touched at all.** + +--- + +## Q1 — Is the safety net real? **NO — and worse than "no": it cannot be seen from the box at all, so the exposure must be treated as UNBOUNDED until Viktor looks at the control panel.** + +**Measured** on both live boxes, over the SFTP credential each box already holds, read-only. + +| probe | demo-hp (sub-account A, gid 1058) | demo-felhom (gid 1019) | +|---|---|---| +| **positive control** — the account's own home | lists `.ssh`, `` | lists `.ssh`, ``, `.orphaned-20260810` | +| **negative control** — `./zzz-no-such-r95` | `not found` | `not found` | +| `./.snapshots` | **`not found`** | **`not found`** | +| repo path itself | `config data index keys locks snapshots` | `data index keys locks snapshots` | +| `/.snapshots` | `not found` | `not found` | +| `/` (box root) | **`Permission denied`** — the account is jailed to `/home` | same | + +**Two readings, and the box cannot distinguish them:** either no snapshots exist, or they exist and +are invisible to a sub-account. **Both are bad, and the second is not a reprieve** — a snapshot the +box cannot see is a snapshot the box cannot restore from either. Recovery would be an operator act +through the Hetzner panel, not something the product can do. + +**What the register actually claims**, `OPEN-ITEMS.md:233`, verbatim: + +> `| — | Storage Box **snapshots** on storage-box-pool-1 — plan SET (daily 00:00, keep 7) but **0 taken yet** | WATCHING | first run tonight 00:00 | Confirm size_snapshots > 0 tomorrow; ... until then the mitigation is armed, not proven |` + +**Three facts about that row:** + +1. **It has no R-number** (`| — |`), so no gate and no grep-the-register rule can ever cite it. +2. **Its "confirm tomorrow" was 2026-07-27** (`git log -S`, last touched in `72692e1`). That is + **36 days** ago. Nobody confirmed. +3. **The `DUE-CHECKS` block — the mechanism this project built (R-341) for exactly this — is + EMPTY.** The one dated commitment that mattered was never entered into it. + +**And R-95 leans on it**: its own text reads *"Mitigation now ARMED"*. **That word is doing work it +has not earned.** → **R-429**. + +**Not answerable here:** whether snapshots exist. The register's own confirming field, +`size_snapshots`, is a **Hetzner API field**, and §11-D records the deliberate decision not to call +that API with the production token. **This spike stopped at that fence.** → a question for Viktor. + +**Explicitly NOT done:** no restore from a snapshot was attempted. Whether one *can* be restored is a +separate question and is named as such. + +**What this does to the urgency: it raises it.** The task supposed Q1 might bound the worst case to +seven days. It does not. **The bound is unverified, and unverifiable from the product side.** + +--- + +## Q2 — What must the box write, and what must it delete? **Ten verbs, not nine — and there are TWO `forget --prune` sites, not one.** + +Census by grep over `internal/backup/*.go` at `960d29b`, each confirmed by reading its context. + +| verb | `file:line` | class | +|---|---|---| +| `init` | `offbox.go:321`, `offbox.go:811` | **write** | +| `backup` | `offbox.go:927`, `offbox.go:1332`, `offbox_shares.go:107` | **write** | +| `forget … --prune` | **`offbox.go:1388`** (retention) and **`offbox.go:1759`** (over-quota) | **DELETE** | +| `unlock` | `offbox.go:746` (stale-only), `offbox.go:768` (`--remove-all`) | **DELETE** | +| `restore` | `offbox.go:1827`, `offbox_restore.go:393`, `offbox_proof.go:328` (`--no-lock`) | read | +| `snapshots` | `offbox.go:1771`, `offbox_inventory.go:73`, `offbox_restore.go:126` | read | +| `stats` | `offbox.go:1786`, `offbox_restore.go:160` | read (**takes a lock**) | +| `cat` | `offbox.go:782` | read | +| `check` | `offbox_integrity.go:316` | read (**takes a lock**) | + +**Two corrections to the task's list, both mine to report:** + +* **`check` is missing from it.** It is a real off-site verb (`offbox_integrity.go:316`) and it is + lock-taking, which makes it central to Q5. +* **`dump` is NOT a verb.** `offbox_progress.go:185` is `OffboxPhaseDump = "dump"`, a progress phase + name. I withdrew it after reading the context rather than counting it. + +**Does the box need the delete verbs?** `forget`/`prune` — **no**, that is retention and it is exactly +what R-89 moved off the box for PBS. `unlock` — **this is the whole of Q5**. + +--- + +## Q3 — Can the sub-account express "write but never delete"? **No. The API offers exactly one axis, and it is all-or-nothing.** + +**DOCUMENTED, not measured** — from this project's own mirror of the API, +`hub/internal/hetznerapi/hetznerapi.go:38-45`, which is the authority we have without calling the +provider: + +```go +type AccessSettings struct { + SSHEnabled bool `json:"ssh_enabled"` + ReachableExternally bool `json:"reachable_externally"` + SambaEnabled bool `json:"samba_enabled"` + WebDAVEnabled bool `json:"webdav_enabled"` + Readonly bool `json:"readonly"` +} +``` + +**Five booleans. `readonly` is the only permission axis, and a backup target cannot be read-only.** +There is no append-only, no per-directory grant, no write-without-unlink. **This is the finding that +forces Q6**, and the register already suspected it. + +**Measured corroboration, no new action taken:** demo-felhom's home contains +a `.orphaned-20260810` directory beside ``. That directory was **renamed by the controller** +(the orphan guard). A rename is a delete-class operation. **So the credential's reach is not a +theory — the existing state on the live store is evidence of it.** + +**Not established:** whether the live API would report anything the struct omits. Settling it needs +the provider token — **fenced** (§11-D). + +--- + +## Q4 — What did the PBS fix do, and what would copying it cost? **The shape does not transfer, because the far end is a disk, not a server.** + +**What moved (R-89, PROVEN-LIVE):** retention became a hub-owned commercial attribute; boxes set +`keep_last: 0`; **ep0 runs the prune jobs**; box tokens stay write-only. `07` §8 row 10 records the +box being *refused* when deleting its own PBS snapshot. + +**Why it does not transfer as-is:** PBS is a **server** that can refuse. The restic tier writes to a +**Storage Box over SFTP** — a filesystem. **A filesystem runs nothing.** So "move retention to the far +end" has no far end to move it to. Candidates, with their costs: + +| who runs retention instead | cost | new risk it creates | +|---|---|---| +| **the hub** | a scheduler + a per-customer credential | **the hub gains a credential that can delete every customer's history** — one compromise instead of N | +| **DooPlex** | a cron + the same credentials | same concentration, on a **Tier-2 box that is itself the recovery chain** | +| **ep0** | a service on a protected machine | **an architecture change, not a config change**; ep0 is protected | +| **nobody — never prune** | £0 today | the store grows without bound; the over-quota path at `offbox.go:1759` exists precisely because quota is already a live concern | + +**What the migration cost last time, and the trap is identical here.** R-191: R-89 changed the +contract, one client-side retention setting did not follow, and **every successful weekly backup then +reported as a failure** — upload complete, 67.2 % reused, then `TASK ERROR: job errors` and a +`whole_guest_backup_failed` alarm to the operator. *"A doc that states the contract does not enforce +it — the gate does."* + +**The same trap, doubled:** withdrawing delete without disarming retention would make every off-site +run log `[WARN] forget --prune failed` — and there are **two** call sites, `offbox.go:1388` **and +`offbox.go:1759`**. The task's brief names only the first. **Disarming one and not the other +reproduces R-191 exactly.** + +--- + +## Q5 — The lock problem. **Measured, and it is NOT the blocker I expected — but it exposed something worse: `unlock` reports success on a deletion that did not happen.** + +**Method.** A throwaway **local** restic repo in the controller container's `/tmp` on demo-hp (60 MB, +removed at the end, **no live store touched**). Append-only was modelled faithfully: a sticky locks +directory (`1777`) owned by root, holding a root-owned lock, with restic run as `nobody`. That is +exactly append-only semantics — **create allowed, delete refused**. + +**Both controls passed before anything was believed:** + +* `nobody` **can** create in the locks directory → the model is not simply "read-only". +* `nobody` **cannot** delete root's lock → the model really does refuse deletes. + +A **genuine** restic lock was captured (copied out while a real `check --read-data` held it), not +hand-forged. + +| verb, running as `nobody` against an undeletable stale lock | result | +|---|---| +| `restic unlock --remove-all` | **prints `successfully removed locks` — and the lock is STILL THERE** | +| `restic backup` | **succeeds** — `snapshot a7928829 saved` | +| `restic check` | **succeeds** — `no errors were found` | +| `restic snapshots --no-lock` | succeeds | +| `restic restore --no-lock` (the R-87 proof pattern) | succeeds | + +**So the answer to "does the store wedge?" is: not in this case.** restic 0.14.0 treats a lock whose +owner is provably dead as stale and proceeds without needing to remove it. **Withdrawing delete does +not, by itself, wedge the store.** That removes the constraint the task suspected would be deciding. + +**But the measurement found a different defect.** `unlock --remove-all` **reported success while +deleting nothing.** This project already has a named class for that — *"Exit codes that lie"* — and +`resticStep`'s crash-lock self-heal is built directly on top of this call. **A self-heal that cannot +fail is a self-heal that cannot be trusted.** → **R-430**. + +**UNKNOWN, and I am not claiming otherwise:** the **crash-lock** case — a lock left by a container +that no longer exists, whose hostname restic cannot match, so it will **not** treat it as stale for +~30 minutes. That is documented in `resticStep`'s own comment (`offbox.go:~760`) and is the reason +`--remove-all` exists at all. My model could not reproduce it: the captured lock carried this +container's own hostname. **What would settle it:** a lock captured from a container with a different +hostname, replayed against an append-only endpoint. **In that window, the only remedy is the very +call that Q5 just showed reports success while doing nothing.** + +**How far `--no-lock` reaches:** restic's own help says it *"allows some operations on read-only +repositories"*. Measured: `snapshots` and `restore` work under it. `backup` and `check` still take a +lock — but **creating** a lock is a write, which append-only permits. So the read paths are safe and +the write paths are unaffected; **only removal is denied, and only the crash-lock window depends on +removal.** + +--- + +## Q6 — Is an append-only transport reachable? **Yes in principle — restic 0.14.0 does speak REST — but not without either moving the data or putting a machine in front of it.** + +**Measured**, in the controller container, with a control: + +| repo string | result | +|---|---| +| `banana:http://127.0.0.1:1/x` (control) | `Fatal: parsing repository location failed: invalid backend` | +| `rest:http://127.0.0.1:1/` | `Fatal: unable to open config file: … dial tcp 127.0.0.1:1: connect: connection refused` | + +**A connection error, not a parse error — the REST backend IS recognised by restic 0.14.0.** + +**Append-only is not a restic feature.** Measured: `restic help` and `restic backup --help` contain +**zero** occurrences of "append". It is a **rest-server** flag (`--append-only`). So the client is +ready and the server does not exist yet. + +**In front, or move the data?** rest-server serves a **local directory**. The Storage Box is a remote +share. So either (a) a machine mounts the Storage Box and runs rest-server over that mount — keeping +the bytes where they are, at the cost of a mount in the hot path — or (b) the data moves to storage +attached to whatever runs rest-server. **These are very different prices and the choice is not +obvious**; (a) keeps the current bill, (b) is a migration of every customer's history. + +**Where would it run?** **`ep0` is protected. Adding a service to it is an architecture change, not a +configuration change** — and it also makes ep0 a single point of failure for both tiers, which is +precisely the concern §11-D already has open. + +**Cost:** one new always-on service in the recovery path, plus whatever machine hosts it. **The moving +part matters more than the money:** if rest-server is down, backups stop — and this project's own +history (R-191) is a warning about what happens when a change in the off-site contract is not carried +everywhere at once. + +--- + +## Q7 — What does DETECTION cost? **Almost nothing. The number already arrives at the hub, and the hub already keeps the history to compare it against.** + +**Measured from source:** + +* The box already reports `snapshot_count` — `controller/internal/report/types.go:131`, and it is + already rendered on the hub's Backup card (`hub/internal/web/backup_card.go:118`). +* **The hub retains report history**: `store.go:968` is an `INSERT INTO reports`, and the read is + `SELECT … ORDER BY id DESC LIMIT 1` (`store.go:1016`). **It appends; it does not replace.** So every + previous `snapshot_count` for every customer is already on disk. + +**So an unexplained drop is a comparison between two rows the hub already has.** No new box code, no +new credential, no new moving part, no new bill. The alarm vocabulary and the operator-only routing +already exist. + +**The honest caveats:** a legitimate `forget` also drops the count, so the rule needs to know +retention (and today the box prunes itself, which is what makes the number noisy — **the same change +that disarms box-side retention is what makes this signal clean**). And detection is not prevention: +it tells you within a day that history was destroyed; it does not stop it. **`snapshot_count: 0` must +mean UNKNOWN, not EMPTY** — `backup_card.go:29` records that exact defect being fixed once already. + +--- + +# Ranked options for Viktor + +**All four were considered. My recommendation is 4 + 2, in that order, and explicitly not 3 yet.** + +### 1. Do nothing — **NOT acceptable as it stands, and Q1 is why** + +Before this spike, "do nothing" meant *"exposed, but bounded to seven days by snapshots."* **That +sentence is not supported by any evidence.** The box cannot see a single snapshot, on either machine, +and the register's confirmation step was never done. **Doing nothing now means an unbounded exposure +on the tier holding the customer's documents and photos.** +**Cost:** £0, no evenings. **What happens if you choose it:** the risk stays exactly where it is, and +the register keeps saying "ARMED", which is the part I would not accept. +**One cheap act rescues most of this option: look at the Storage Box panel and answer Q1.** If +snapshots are real, "do nothing" becomes defensible again. **That is ten minutes and it is yours to +do — the fence in §11-D stopped me.** + +### 2. Copy the PBS shape — **worth doing, but it is smaller than it sounds and it is not the safety it appears to be** + +Withdraw box-side retention; someone else prunes. **Q5 says this no longer looks dangerous** — a stale +lock does not wedge the store. +**But Q3 says the credential still cannot be narrowed**: the API has one axis, `readonly`, and a +backup target cannot be read-only. **So the box would keep the ability to delete, and simply stop +using it.** That is a discipline, not a guarantee — a compromised guest is unaffected by it. +**Cost:** one evening, plus wherever retention moves (each candidate concentrates the credential). +**Trap:** both `forget` sites (`offbox.go:1388` **and** `:1759`) must be disarmed in the same change, +or you get R-191 again — successful backups reported as failures. + +### 3. Change the transport (rest-server `--append-only`) — **the only real prevention, and I would not start it yet** + +**Q6 says it is reachable**: restic 0.14.0 speaks REST. This is the option that actually makes the +box *unable* to delete. +**Cost:** a new always-on service in the recovery path; a machine to host it (**ep0 is protected — +this is an architecture change**); and either a mount in the hot path or moving every customer's +history. **Money is the small part.** +**Why not yet:** it should be chosen against a measured picture, and one measurement is still missing +— Q1. Building the expensive prevention while nobody knows whether a seven-day net already exists is +the wrong order. + +### 4. Detect instead of prevent — **cheapest by a wide margin, and I would do this first** + +**Q7 measured that the material is already there**: the count is on the wire, the hub keeps the +history, the alarm path exists. An unexplained drop in `snapshot_count` becomes noticeable within a +day. +**Cost:** hub-side only. No box change, no credential change, no new service, no new bill. +**What it does not do:** it does not stop the deletion. It converts "we would never know" into "we +know tomorrow" — and against ransomware inside the guest, knowing tomorrow is the difference between +losing a day and losing everything silently. +**This week's pattern held twice already** (R-87, R-404): **proving beats preventing when preventing +is expensive.** This is the third instance. + +**My pick: answer Q1 today (yours, ten minutes), then build 4, then 2. Revisit 3 once Q1 is +answered** — and if Q1 comes back "no snapshots", 3 moves up sharply. diff --git a/documentation/backlog/OPEN-ITEMS.md b/documentation/backlog/OPEN-ITEMS.md index 4e6340ed..11e85a33 100644 --- a/documentation/backlog/OPEN-ITEMS.md +++ b/documentation/backlog/OPEN-ITEMS.md @@ -224,7 +224,7 @@ unconditional promise, which CAMPAIGN-11 §7 step 7 measured the customer-facing | **E-2d** | **Prove E-2 on a fresh VM** — a real `felhom-host-install.sh` 1.22.0 run, Case B naturally, a claimable customer, then add a drive (the offer) and unplug it (`backup_target_absent` end-to-end) | **CLOSED — PARTIALLY PROVEN** (2026-07-29) | — | **C1, C2 proven** (`audits/E2D-fresh-vm-2026-07-29.md`); **C3, C4 proven live** (`audits/SESSION-C-2026-07-29.md`); **C5 FAILED → R-116** — the gate fires and an alarm reaches the hub, but it is the generic event, so the alarm and its recovery cannot be paired. **R-116 is the single named open leg**; per the Session-C runbook §9, decided in advance, a failed claim closes the item as partially proven rather than triggering a re-run. Both audits carry the full record — the `local-lvm` fence, the ISO/PAIRING derivation, the Phase 0 answers, the per-claim observables and the teardown evidence — and are the place to read it, not this cell. **The arc's actual definition of done is R-106 + R-109, R-108 and D5**, none of which this detour touched | CC | | **R-121** | **A BOX's installed agent can sit releases behind the vouched one and nothing notices — the R-120 gate does not cover it.** demo-hp ran agent **0.113.0** while the hub vouched **0.116.0**, through the whole R-116/R-117 arc, and no signal existed on any channel | **READY (S) — NEW 2026-07-30** | — | **Fourth instance of the drift family** (R-111 golden's agent 17 releases behind, R-115 built+deployed but never published, R-120 golden a controller behind — and now installed-vs-vouched on a live box). **Confirmed at source that R-120's gate cannot catch it:** `hub/internal/web/configs.go:1165-1169` compares `goldenVer` against `store.NewestReportedControllerVersion()` — it is a **golden-artifact vs fleet-CONTROLLER** check and says nothing about the agent installed on a box. **`MinAgent` does not cover it either:** it is used to HOLD the controller floor for a box whose agent is too old (`hub/internal/api/handler.go:530-538`, `store.go:1857`) — protective, not an alarm — and demo-hp's 0.113.0 **equalled** `min_agent` 0.113.0, so even a floor comparison was satisfied. **The cost, measured:** R-117's whole subject is the R-113 conjunction, which landed in **0.114.0** — so the designated drill host could not exercise the code under investigation at all, and the R-117 spike had to route every predicate result through an out-of-repo probe built from `main` instead of the installed agent (`audits/SPIKE-r117-bind-liveness-2026-07-30.md` §1, §2.3). Discovered because the R-117 task made bringing the box current an explicit prerequisite. **Fix shape (not implemented):** the hub already receives `AgentVersion` on every host report, and already has semver comparison in Go — the missing piece is a checker comparing reported agent vs the **vouched** agent and surfacing it, operator-tier. Note the honest tension: a box legitimately lags between publish and deploy, so this wants a staleness window rather than an instant alarm | CC | | **R-118** | **An absent drive's union row advertises the ROOT filesystem's capacity as its own.** In the absent-state payload the registry-union row reports `total_bytes: 49675956224 / used_bytes: 4584579072` — **byte-identical to the `local` row** (`durable_id: path:/var/lib/vz`, i.e. `pve-root`) in the same response. The real drive is **4 GB** | **READY (XS) — NEW 2026-07-30** | — | Cause: `statfsCapacity(d.MountPath)` (`disks.go:335-338`) statfs's `/mnt/cel`, which with the device gone is a **bare directory on the root filesystem**. `observe.go:176-183`'s comment warns about exactly this trap and guards the Observe path ("*an unmounted removable dir-storage's mountpoint reverts to a bare directory on root … catastrophic DR mis-id*"); **the union path has no equivalent guard.** **Not a DR mis-id** — `durable_id` on that row is still the correct `uuid:…`, so re-attach identity is safe. It is a **false capacity** reaching every consumer of `total_bytes`/`used_fraction` (fill monitors, storage cards): a detached 4 GB drive advertises 46 GiB at 9.2 % used. Same class as `role.go:180-181` — an absent drive's fields decaying to the root filesystem's. Evidence: `audits/DIAG-r116-disks-payload-2026-07-30.md` §12 | CC | -| **R-95** | restic offsite credential **can delete** (`readonly=False`, `forget --prune` runs from the box); SFTP cannot express append-only | **READY** | — | Root exposure still open. Mitigation now ARMED — split prune off-box or move to REST `--append-only` | CC | +| **R-95** | restic offsite credential **can delete** (`readonly=False`, `forget --prune` runs from the box); SFTP cannot express append-only | **READY** | — | Root exposure still open. Mitigation now ARMED — split prune off-box or move to REST `--append-only` | CC **SPIKE 2026-09-01 — `audits/SPIKE-r95-offsite-delete-2026-09-01.md`. THE WORD "ARMED" ABOVE IS NOT SUPPORTED AND IS WITHDRAWN PENDING R-429:** no `.snapshots` is visible to either box's sub-account (measured, both machines, with controls), so the seven-day bound is unverified and unverifiable from the product side. **Q3 (documented, `hub/internal/hetznerapi/hetznerapi.go:38-45`): the sub-account API has ONE permission axis, `readonly` — there is no append-only, so the PBS shape does NOT transfer** (PBS is a server that can refuse; a Storage Box is a filesystem that runs nothing). **Q5 (measured): withdrawing delete does NOT wedge the store** — restic treats a dead owner's lock as stale and proceeds — so the constraint everyone feared is not the blocker; but `unlock --remove-all` lies about success (R-430), and the crash-lock window is UNKNOWN. **Q6 (measured): restic 0.14.0 DOES speak `rest:`** (control: `banana:` → invalid backend), and append-only is a rest-server flag, not a restic one — reachable, but it needs a machine in the recovery path and ep0 is protected. **Q7 (measured): detection is nearly free** — `snapshot_count` already reaches the hub and the hub APPENDS reports, so the comparison needs no box change. **RECOMMENDATION: answer R-429 first (Viktor, ten minutes), then build detection, then move retention off the box; defer the transport change.** Two `forget --prune` sites must be disarmed together — `offbox.go:1388` AND `offbox.go:1759` — or R-191 repeats. | | **R-191** | **Every weekly offsite backup UPLOADS successfully and then FAILS the job on a prune the box is deliberately not allowed to do — on both demo boxes.** Measured on demo-felhom 2026-08-04 06:49–06:53: the upload completed (223 s, 629 MiB of 1.874 GiB, **67.2 % reused incrementally**), then `ERROR: prune 'ct/9201': proxmox-backup-client failed: Error: permission check failed - missing Datastore.Modify\|Datastore.Prune on /datastore/felhom-offsite/demo-felhom` → `ERROR: Backup of VM 9201 failed - error pruning backups` → `TASK ERROR: job errors`. The hub raised `whole_guest_backup_failed` | **CLOSED — SHIPPED 2026-08-04** (installer **1.25.0**; both live boxes corrected) | — | **This is R-89's rule not reaching the config.** R-89 moved PBS pruning SERVER-SIDE — *"boxes set `keep_last: 0`, ep0 runs prune jobs; box tokens stay write-only, never widen the grant"*. The token behaves exactly as designed: it refuses. But **both** demo boxes still arm the offsite tier with `keep_last=2 prune_pbs_allowed=true` (`backup_targets: [{target_id: felhom-pbs, cadence_seconds: 604800, keep_last: 2}]`), so every run asks for a prune that must fail. **The data is SAFE and that is why this is not a P1:** the snapshot lands before the prune is attempted; what is wrong is the job's VERDICT and the weekly operator e-mail it produces. **But it is corrosive in the specific way this project keeps finding:** a backup that reports FAILED while succeeding trains the operator to discount `whole_guest_backup_failed`, which is the same alert that would carry a real one — and it is exactly the failure the R-100 corollary warns about, an alarm whose text is true and whose trigger is not the thing you would act on. **Fix is one config line per box** (`keep_last: 0` on the PBS tier) plus whatever writes it on a fresh install; **deliberately NOT applied in this session** — the session was a runbook with an explicit "change nothing, and if a change appears necessary, stop and report" rule, and a retention field on a live backup tier is not a change to slip into an observation run. **Check before fixing:** whether ep0's prune jobs actually cover these two namespaces, or the snapshots simply accumulate once the box stops asking **THE GATE WAS RUN FIRST, AND IT MATTERED.** Before disabling anything, ep0 was read (read-only, Tier 2): prune jobs `prune-demo-felhom` and `prune-demo-hp` exist on datastore `felhom-offsite`, one per namespace, `schedule 03:30`, `keep-last 2`, comment *"R-82 retention keep-last=2, server-side (box tokens are write-only)"* — and they have run **every day since 2026-07-27: 18 tasks, all `status=OK`**. The newest task log reads `retention options: --ns demo-felhom --max-depth 0 --keep-last 2` / `keep ct/9201/2026-07-27…` / `keep ct/9201/2026-07-28…` / `TASK OK`. Retention happens, and it happens there. **A METHODOLOGICAL WARNING WORTH MORE THAN THE FIX.** Three separate queries said the OPPOSITE — *no prune jobs have ever run* — and **all three were broken instruments**: `worker-type` where the field is `worker_type`; the value `prune` where the worker type is `prunejob`; and `journalctl -u proxmox-backup` where the unit is `proxmox-backup-proxy`. A fourth reading (3 snapshots under keep-last 2) was mis-framed by CC and self-corrected — the third snapshot had landed AFTER that day's 03:30 window. Acting on any of them would have disabled the only pruning ATTEMPT while reporting that nothing prunes: a weekly false alarm traded for unbounded growth on the protected endpoint, invisible for months. **The gate is what caught it, and only because it demanded evidence rather than a verdict.** **Shipped:** installer **1.25.0** writes `keep_last: 0` on the offsite tier (the agent's existing guard `allowPBSPrune = !primary && keep_last > 0` already reads that as *never prune from the box* — no agent change), the justifying paragraph is rewritten to say where retention lives and cite R-89, and `hostinstall_gates.py` asserts it (red-proved: pinning `keep_last: 2` back fails the gate). **Both live boxes corrected in their own config** — `backup tier armed target=felhom-pbs … keep_last=0 … prune_pbs_allowed=false` on demo-felhom and demo-hp, with the local tier untouched at `keep_last=3`. Served over HTTPS at `1.25.0` with `"keep_last":0` in the served bytes. **STILL TO OBSERVE:** the next weekly offsite run completing OK end-to-end. The change removes the failing step; the *schedule* proving it is next week's event, and this row should carry that line when it happens. | CC | | **R-194** | **PVE's permission cache delays every grant-state verdict by an unknown amount, so "the agent can read it" and "the ACL exists" are not the same measurement.** Observed twice while validating R-190's self-repair on demo-felhom 2026-08-04: both ACL rows for `/storage/felhom-backup` were deleted, and `GET /access/permissions` continued to report `Datastore.AllocateSpace` present — for **~40 s** in one run and **~16 minutes** in another. During that window the capability probe reads healthy and the self-repair does not fire | **OPEN** | — | **Why it matters beyond the delay:** it puts a floor under how fast a lost grant can be noticed, it makes any single permission read a lagging indicator, and — the interesting part — **it is a candidate contributor to R-190's own timeline**: a grant removed at an unknown moment could keep working until a cache expiry, which is exactly the shape of *worked at 04:44, refused at 09:24*. That does not explain what removed it, but it may explain when the refusal SURFACED, and the two have been treated as the same instant. **Not a defect in our code** — it is PVE behaviour, and the mitigation already tolerates it (the repair fires on the next probe after the cache clears). **What is worth deciding:** whether the store-grant probe should ALSO consult the storage content listing as a second signal, since that appeared to reflect the loss immediately (`{"data":[]}` while the permission read still said present) — two signals disagreeing is itself information, and today only one of them is read | CC | | **R-200** | **The DR password-injection seam has a handler, a route and tests — and no form.** `POST /backup/offbox/inject-password` is routed (`controller/internal/web/server.go:510`) to `offboxInjectPasswordHandler` (`offbox_handlers.go:174-196`) → `InjectOffboxPassword` (`backup/offbox.go:541`). **No template in the repository contains that path or any form posting to it** (grep over `internal/web/templates/`: one unrelated hit, an XSS comment) | **PLUMBING COMPLETE** (controller **v0.196.0**); **the FORM is not built — still open** | — | **The tenth instance of this project's built-but-never-wired class, and the exact shape `CLAUDE.md` and `felhom.eu/CLAUDE.md`'s seam-wiring rule were written for:** handler tests that POST directly (`offbox_escrow_test.go:167,180`) prove nothing about reachability. To use the only implemented recovery seam today, a person must hand-craft an authenticated POST with a session cookie and CSRF token. **Note the layering while fixing it:** this form takes a **64-hex repo password**, not a recovery code (`offboxRepoPwPattern`, `offbox.go:543`) — they are different secrets at different layers, and the operator's 2026-08-04 ruling asks for a form that takes **R**. Build the R form and treat this one as the operator/DR fallback it was written as, but **ship it with a render test per branch of whatever gate it sits behind**. Source: `audits/RECON-offsite-dr-chain-2026-08-04.md` §3 link 9 **THE DIAGNOSTIC HALF IS DONE AND IT ANSWERED THE QUESTION.** `--recover-offsite-check` is a `docker exec` escape hatch in the shape of `--print-reset-code`: R on **STDIN** (never argv, never `ps`, never shell history, never a transcript), fetch+unseal via the agent, and a verdict of **two sha256 hashes**. **It compares and never installs** — the recovered password is not written to `offbox/repo_password`; a test asserts the data dir is byte-unchanged and its red-proof (adding the install call) fails it. Confirmed live: `repo_password` mtime still `2026-08-03 07:18:02` after the successful check at `2026-08-04 11:49`. **Exit codes are load-bearing** — `0` match, `2` a clean MISMATCH, `1` a step failed; "it failed" and "it worked and disagreed" must never share a status because only one is a finding about the system. A box with no local password reports distinctly (the rebuilt-box shape, where the next step is to INSTALL rather than compare). **WHAT IS NOT BUILT, deliberately:** no card, no form, no preview, no customer-facing text — building an interface on top of a chain nobody had walked is how the preceding three weeks went wrong. **What remains for this row:** link 9 (the recovered password placed so `WriteOffboxSecrets` keeps it) and the customer-facing shape the operator ruled on 2026-08-04 (yell → R form → preview → proceed), which is now priced against a chain that exists rather than one that is assumed **PART 0 SHIPPED 2026-08-04 (v0.196.0):** `--recover-offsite-install` is the sibling of the check — same fetch/unseal path, same STDIN discipline for R — and it **places** the recovered password via `InjectOffboxPassword`. The confirmation is a SECOND invocation (`--confirm-install`): without it, both hashes print and nothing is written, so the operator sees the comparison before a write is possible. Three outcomes named distinctly: **installed** (no local password — the rebuilt-box shape), **unchanged** (identical key, nothing written), **refused** (a DIFFERENT key present — installing would clobber the key the current repository is encrypted under; exit 2, no force offered). It re-reads the file after writing rather than trusting the call. **Red-proof observed:** removing the confirmation gate makes the dry run write the password. The R-persistence test carries a **positive control** (a planted copy found, then removed and not found). **NOT YET EXERCISED AGAINST A LIVE RECOVERY** — the R-201 drill halted before step 9, so this is unit-proven only. **What remains for this row:** the customer-facing shape the operator ruled on 2026-08-04 (yell → recovery-code form → preview → proceed) | CC | @@ -602,6 +602,8 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server` | **R-426** | **The decoy-coverage exemption list — 20 registered gates that ship WITHOUT a decoy test, each named.** `scripts/decoy_coverage_gate.py`'s `EXEMPT` map is debt, and this row owns it so it lives in the register and not only in a Python literal. **Four kinds:** (a) genuinely covered in the 2026-09-01 sweep but not yet moved into a suite — `hub-copy`, `instructions`, `docker-v`, `image-pins`; (b) blocked by an open hole and therefore un-assertable as rejecting — `site` (R-423), `one-register` (R-424), `offbox-rename` (R-425); (c) shared scripts whose decoy lives in `felhom.eu` and is counted there — `reuse-refs`, `instructions`, `observations` in the controller and agent runners; (d) **no plausible decoy constructed yet** — `hostinstall`, `wire-contract`, `due-checks`, `published`, `image-resolvable`, `volume-persistence`. Group (d) is the honest unknown: six gates whose soundness is UNTESTED, not established. **The list is green today and shrinks; a NEW gate with no decoy fails immediately.** | **OPEN — 20 names; group (d) is six untested gates** | | **R-427** | **`closed_register_gate.py` checks ONE direction only: an open word in a CLOSED row. The mirror — a CLOSED verdict on a row still sitting in `OPEN-ITEMS.md` — is unchecked, and there are TWELVE.** MEASURED 2026-09-01 during the decoy sweep, by reading the leading verdict of every open row with the gate's own predicate: **R-385, R-387, R-341, R-378, R-405, R-88a, R-88b** read unambiguously closed; **R-123, R-190, R-352** read `PARTLY CLOSED` / `MITIGATION SHIPPED` and almost certainly belong where they are. **The rows were NOT moved by this session** — telling a finished row from a partly-finished one is a judgement, and R-378 is itself the record of what happens when a machine makes that judgement on a substring (six still-open rows moved out of the register). **This is R-405's finding mirrored:** that row exists because R-87 sat in the CLOSED file while its state read READY, and the gate written for it looks only the way it was bitten. Fix: the same leading-verdict predicate applied to `OPEN-ITEMS.md`, reporting rather than convicting until the twelve are adjudicated by a person — a gate registered while twelve rows fail it would refuse every push. | **OPEN — 12 rows named; the adjudication is Viktor's, the gate is mine** | | **R-428** | **The decoy-coverage gate — written to catch instruments that match a NAME instead of a fact — identified a repository by its DIRECTORY NAME.** MEASURED on its own first CI run (felhom.eu job 490, 2026-09-01): `os.path.basename(root)` looked up in a `RUNNERS` map, and Gitea's act-runner checks the repo out into a directory called `hostexecutor`, so the gate reported *"unknown repo 'hostexecutor'"* and went INCONCLUSIVE. **The gate that hunts label-matching was matching a label, in the first ten lines of its own main loop, and it shipped that way.** FIXED the same day: it now identifies a repo by which registered runner FILE exists under the root, which is a fact. **Recorded rather than quietly patched because it is the strongest evidence in the sweep that this class is not a matter of carelessness** — it was written by a session that had spent the morning reading 29 gates for exactly this, with the four shapes on screen. Verified under a renamed directory before and after. | **CLOSED 2026-09-01 — fixed, and kept as the class's best example** | +| **R-429** | **The Storage Box snapshot mitigation that R-95 calls "ARMED" has never been confirmed, cannot be seen from either box, and its row has no id.** MEASURED 2026-09-01 over the SFTP credential each box already holds, read-only, with positive and negative controls on both machines: **no `.snapshots` directory is visible to either sub-account** — not in the account home, not inside the repo path — while the account is jailed (`/` returns `Permission denied`) and a bogus path errors correctly. Two readings the box cannot distinguish: no snapshots exist, or they exist and a sub-account cannot see them. **Both are bad, and the second is not a reprieve — a snapshot the box cannot see is one the box cannot restore from, so recovery would be an operator act through the Hetzner panel, not a product capability.** The claim lives at `OPEN-ITEMS.md:233` as a row whose ID cell is a bare em-dash instead of an R-number, so no gate and no grep-the-register rule can cite it; its "confirm tomorrow" was **2026-07-27** (last touched in `72692e1`), **36 days** unconfirmed; and the `DUE-CHECKS` block built for exactly this (R-341) is **EMPTY**. **The confirming field, `size_snapshots`, is a Hetzner API field and §11-D fences it — so this spike stopped and left it for Viktor.** Fix: Viktor reads the Storage Box panel once (ten minutes) and the answer goes in this row; give the WATCHING row an id or delete it; and R-95's word "ARMED" must not stand until then. Evidence: `audits/SPIKE-r95-offsite-delete-2026-09-01.md` §Q1. | **OPEN — Q1 is Viktor's, and it re-ranks R-95** | +| **R-430** | **`restic unlock --remove-all` printed `successfully removed locks` while the lock was still there.** MEASURED 2026-09-01 in a throwaway local repo (no live store touched), under a faithful append-only model — a sticky locks directory owned by root holding a root-owned lock, restic run as `nobody`; both controls passed first (create allowed, delete refused). The command reported success, returned, and `ls` showed the lock present. **`resticStep`'s crash-lock self-heal is built directly on this call** (`felhom-controller/controller/internal/backup/offbox.go:~768`), and its licence to escalate rests on the escalation actually working. **A self-heal that cannot fail is a self-heal that cannot be trusted** — this is this project's *"exit codes that lie"* class, in the one path that runs unattended against the customer's off-site history. It is harmless TODAY because the credential can delete and the removal really happens; it becomes load-bearing the moment delete is withdrawn, which is what R-95 is about. **Not yet established:** whether restic reports success because it removed zero locks by design, or because it did not check. Settling it: read restic 0.14.0's unlock source, or re-run with `--verbose`. | **OPEN — precondition on any R-95 build** |