diff --git a/REPORT-r95-spike.md b/REPORT-r95-spike.md new file mode 100644 index 00000000..4bb271f6 --- /dev/null +++ b/REPORT-r95-spike.md @@ -0,0 +1,127 @@ +# REPORT — SPIKE R-95: can the box be stopped from deleting its own off-site history? (2026-09-01) + +## Q1 first, and it raises the urgency rather than lowering it + +**The safety net cannot be seen from the box, on either machine, so the seven-day bound is +unverified — and unverifiable from the product side.** + +Measured over each box's own SFTP credential, read-only, with controls that passed first: + +| probe | demo-hp (sub-account A) | demo-felhom | +|---|---|---| +| positive control — account home | lists `.ssh`, `` | lists `.ssh`, ``, `.orphaned-20260810` | +| negative control — bogus name | `not found` | `not found` | +| `./.snapshots` | **`not found`** | **`not found`** | +| `/.snapshots` | `not found` | `not found` | +| `/` | `Permission denied` (jailed) | same | + +**Either no snapshots exist, or a sub-account cannot see them. The box cannot tell which, and neither +is a reprieve** — a snapshot the box cannot see is one it cannot restore from, so recovery is an +operator act at the Hetzner panel, not a product capability. + +**The register's claim rests on nothing that was ever checked.** `OPEN-ITEMS.md:233` is a row with +**no R-number**, its "confirm tomorrow" was **2026-07-27** (36 days), and the `DUE-CHECKS` block built +for exactly this (R-341) is **empty**. R-95's own text says "Mitigation now ARMED"; that word is +**withdrawn** pending R-429. + +**The task expected Q1 might bound the exposure to seven days. It does not.** I stopped at the §11-D +fence rather than answering it with the provider token. + +## Q1–Q7, one answer each + +| Q | answer | +|---|---| +| **Q1** | **NO / unknowable from the box.** Measured on both machines with controls. → R-429 | +| **Q2** | **Ten verbs, not nine, and TWO `forget --prune` sites.** `check` (`offbox_integrity.go:316`) is missing from the task's list; `dump` is not a verb (`offbox_progress.go:185` is a phase constant) — withdrawn. Delete-capable: `forget`/`prune` (`offbox.go:1388` **and `:1759`**), `unlock` (`:746`, `:768`). | +| **Q3** | **No. DOCUMENTED** from `hub/internal/hetznerapi/hetznerapi.go:38-45`: `AccessSettings` has five booleans and **`readonly` is the only permission axis**. No append-only. Corroborated by existing state, no new action: demo-felhom's home holds a `.orphaned-20260810` directory produced by a controller **rename**, which is delete-class. | +| **Q4** | **The PBS shape does not transfer.** PBS is a server that can refuse; a Storage Box is a filesystem that **runs nothing**, so there is no far end to move retention to. Four candidates costed; each concentrates a delete-capable credential. **R-191's trap is doubled here** — both `forget` sites must be disarmed in the same change or every successful backup reports failure. | +| **Q5** | **NOT the blocker — measured.** Under a faithful append-only model (both controls passed), a stale lock does **not** wedge the store: `backup`, `check`, `snapshots --no-lock` and `restore --no-lock` all succeeded. **But `unlock --remove-all` printed `successfully removed locks` while the lock survived** → R-430. The crash-lock (foreign hostname) window is **UNKNOWN**. | +| **Q6** | **Reachable. MEASURED:** `rest:` gives a connection error where the control `banana:` gives `invalid backend` — **restic 0.14.0 speaks REST**. Append-only is a **rest-server** flag; `restic help` contains zero occurrences of "append". Needs a machine in the recovery path (**ep0 is protected — architecture change**) and either a mount in the hot path or moving every customer's history. | +| **Q7** | **Nearly free. MEASURED:** `snapshot_count` already reaches the hub (`report/types.go:131` → `backup_card.go:118`) and **the hub APPENDS reports** (`store.go:968` INSERT; read is `ORDER BY id DESC LIMIT 1`), so the history to compare against is already on disk. No box change, no credential, no new service. | + +## The four options, ranked + +**My pick: answer Q1 today (Viktor, ten minutes) → build 4 → then 2. Defer 3.** + +1. **Do nothing — not acceptable as it stands.** It used to mean "bounded to seven days". Q1 shows + that sentence is unsupported. £0, no evenings, and an unbounded exposure on the tier holding the + customer's documents. +2. **Copy the PBS shape — worth doing, smaller than it sounds.** Q5 removed the fear that it would + wedge the store. But Q3 means the credential still *can* delete; the box would merely stop using + it. That is discipline, not a guarantee, and a compromised guest is unaffected. One evening + a new + home for retention. +3. **Change the transport — the only real prevention, and not yet.** Q6 says it is reachable. It costs + an always-on service in the recovery path and a protected machine's architecture. Choosing it + before Q1 is answered is the wrong order. +4. **Detect instead of prevent — cheapest by a wide margin, do this first.** Q7 measured that the + material already exists. Converts "we would never know" into "we know tomorrow". Third instance + this week of *proving beats preventing when preventing is expensive*. + +## What I could not measure, and what would settle it + +| unknown | what would settle it | +|---|---| +| **Do snapshots exist on the Storage Box?** | the Hetzner panel or `size_snapshots` via the API — **fenced by §11-D; stopped and left for Viktor (R-429)** | +| Whether the live API exposes any permission the Go struct omits | the provider token — **same fence** | +| The **crash-lock** case (a lock whose hostname restic cannot match, non-stale for ~30 min) | a lock captured from a container with a **different hostname**, replayed against an append-only endpoint. My model could not build it — the captured lock carried this container's own hostname | +| Whether `unlock` reports success because it removed zero locks by design, or because it never checked | read restic 0.14.0's unlock source, or re-run with `--verbose` | +| Whether a Storage Box snapshot can be **restored** | deliberately not attempted — named as the next question, per the brief | + +## Register + +**Before:** OPEN 179 · CLOSED 161. **After:** OPEN 181 · CLOSED 161. +**Filed R-429** (the unconfirmed snapshot mitigation, and the id-less row), **R-430** (`unlock` lies +about success). **R-95 updated** with the verdict and kept OPEN; its word "ARMED" withdrawn. +**Docs:** `07` §8 row 10 and §10.2 gained the verdict (**row 10's status deliberately NOT moved**); +§11-D records that the fence was reached again and held; `STATUS.md` carries one plain-language item. + +## Compliance + +- **No code changed. No version bumped. No image built. No golden owed.** `golden_currency_gate.py` + exits 0; golden and floor remain **0.232.0**. +- **No delete verb was issued against any live store** — no `forget`, `prune`, `unlock` or `init`. + Every live-store interaction was an SFTP `ls`. +- **`ep0`, DooPlex and Peti's box were not touched at all**, not even read — the two demo boxes were + the only machines used. +- **The Hetzner API and control panel were not called.** +- **Scratch resources:** a throwaway local restic repo under `/tmp/r95s` (and `/tmp/r95scratch` in the + first attempt) inside the controller container on demo-hp, 60 MB, plus three probe scripts. **All + removed**, verified by the scripts' own teardown output (`scratch: gone`). Nothing was created on + any Storage Box. + +## Observations, and my own mistakes by name + +1. **The register calls a mitigation ARMED that has never been confirmed, in a row that cannot be + cited because it has no id, with the dated-check mechanism sitting empty beside it.** + **FILED: R-429.** +2. **`restic unlock --remove-all` reports success on a deletion that did not happen**, and the + crash-lock self-heal is built on it. **FILED: R-430.** +3. **The task's own verb list was missing `check` and included `dump`, which is not a verb**, and it + names one `forget --prune` site where there are two. **NOT-A-FINDING: the brief invited me to + confirm the list myself, which is what this is; both corrections are in the spike document and the + second one is carried into R-95's row, because disarming one site and not the other reproduces + R-191 exactly.** +4. **My mistake — my first Q5 model proved nothing.** I used `chmod a-w` and ran restic as **root**, + which ignores permission bits, so every verb succeeded and I nearly recorded "no wedge" on a test + that tested nothing. Rebuilt as a sticky directory with a root-owned lock and restic run as + `nobody`, with two controls. **NOT-A-FINDING: caught inside the same session by the result being + too clean; the corrected model is the one reported, and the first is described so nobody repeats + it.** +5. **My mistake — my first `sftp` probe used `-p` for the port**, which `sftp` reads as "preserve", so + the port became the destination and all three probes returned identical usage errors. **The + controls are what exposed it** — a positive and a negative control failing the same way is an + instrument fault, not a result. **NOT-A-FINDING: a flag error of mine, corrected in one command; + it is recorded because the failure mode it demonstrates — three identical errors reading as three + findings — is the one this project keeps paying for.** +6. **My mistake — I wrote the §8 verdict onto row 4 instead of row 10.** The anchor text I matched + appears in both rows and I replaced the first occurrence. Caught by checking the line number, + reverted from row 4 and applied to row 10, both verified by grep. **NOT-A-FINDING: an editing error + of mine, corrected within the session and verified in both directions, so no wrong claim ever + reached a push.** +7. **My mistake — I put two escaped pipes inside a register row**, which makes it a five-column row in + a three-column table and would have made it unreadable to `closed_register_gate.py` — **the exact + defect I fixed in that gate yesterday.** Caught by counting pipes before committing. **NOT-A-FINDING: + corrected before the push; recorded because I introduced the same shape twice in two days.** +8. **The task's baseline table lists `felhom-agent` at `058b945`; it is at `4586f0f`.** + **NOT-A-FINDING: that is my own push from the previous session, so the table was stale rather than + wrong about anything that matters here; the agent repo was not touched by this spike at all.** diff --git a/STATUS.md b/STATUS.md index 94a961cb..7b0ddcb4 100644 --- a/STATUS.md +++ b/STATUS.md @@ -18,7 +18,7 @@ not an evening's work.** *This section is allowed to be longer than one screen, and each item says what happens if you do nothing.* -1. **Nothing is waiting on you.** Both problems the overnight test found are fixed and proven on +1. **One thing is waiting on you — item 4, and it takes ten minutes.** Otherwise: Both problems the overnight test found are fixed and proven on the real machines: - the background job that could delete a live restore's lock now waits its turn — and the check that finds the next one like it is a test, not a comment, so it cannot come back quietly; @@ -37,11 +37,24 @@ nothing.* two register lines in the hub (already live). No customer action, no data migration, no credential change. -4. **Whether to change the hub password** (R-350). I printed it into my own session log on 20 August. +4. **The copy that holds the customers' documents and photos can still be deleted by the box that + made it** (R-95 — first on the list since July, and this is the first time it has reached this + page). I studied it today and did not change anything. **One thing is yours and it takes ten + minutes.** Our notes say a safety net is switched on: the storage provider takes a snapshot every + night and keeps seven. **I could not find a single one.** I looked from both machines, using the + credentials they already have, and checked that my method could see other things and could + correctly fail to see a made-up name. Either the snapshots are not being taken, or they are + invisible to the machines — and if they are invisible, they are also useless to them: getting one + back would be you, in the provider's control panel. **I did not log in to check, because your own + notes say that question is yours.** **If you do nothing:** the register keeps saying the net is + armed, and nobody knows whether it is. **Please open the Storage Box panel and tell me whether + snapshots exist.** The answer changes which fix is worth building. + +5. **Whether to change the hub password** (R-350). I printed it into my own session log on 20 August. Not in git, not in any saved file — in the log on this machine. **If you do nothing:** it stays as it is, at the risk you accept by leaving it. I can change it without ever showing you the new one. -5. **`demo-hp`'s network setup does not match our own notes** (R-338) — the machine works, the page is +6. **`demo-hp`'s network setup does not match our own notes** (R-338) — the machine works, the page is wrong, or the other way round. **If you do nothing:** the page keeps misleading the next session, as it misled one by an hour. diff --git a/documentation/architecture/07-backup-architecture.md b/documentation/architecture/07-backup-architecture.md index cb6be913..c7aec04e 100644 --- a/documentation/architecture/07-backup-architecture.md +++ b/documentation/architecture/07-backup-architecture.md @@ -98,7 +98,7 @@ with nothing but their dashboard password. No operator, no ticket, no scheduling | 1 | „Visszaállítás indítása" | `POST /backup/restore` | destructive — rebuilds the app from its Tier-1 unit | | 2 | „Fájlok visszaállítása" | `POST /backup/tier2/restore` | additive, missing-only | | 3 | „Visszaállítás a távoli tárolóból" | `POST /backup/offbox/restore` | non-destructive, to a verification copy | -| 4 | „Helyreállítás az élő adatok közé" | `POST /backup/offbox/place` | additive, missing-only **2026-08-30 (R-359): the store is now VERIFIED on a cadence** — a daily `offsite-integrity` job runs `restic check` when the last successful one is over 7 days old. **This is a readability check, not a restore-test (R-87 stays open).** The depth that ships ON now re-reads **100%** of the pack data (R-399 CLOSED, controller v0.228.0) — the sentence here previously said the opposite and was stale. **2026-08-31 (SPIKE R-87):** a scratch restore of every app on demo-hp was measured at **25 s for 8 snapshots / 774 MB**, against 40.3 s for one weekly check — but restic 0.14.0's `--verify` checks size and mtime, **not content**, so nothing available today can vouch for the restored BYTES. Row 4's verdict is UNCHANGED by the spike. `audits/SPIKE-restic-restore-test-2026-08-31.md`. | +| 4 | „Helyreállítás az élő adatok közé" | `POST /backup/offbox/place` | additive, missing-only **2026-08-30 (R-359): the store is now VERIFIED on a cadence** — a daily `offsite-integrity` job runs `restic check` when the last successful one is over 7 days old. **This is a readability check, not a restore-test (R-87 stays open).** The depth that ships ON now re-reads **100%** of the pack data (R-399 CLOSED, controller v0.228.0) — the sentence here previously said the opposite and was stale. **2026-08-31 (SPIKE R-87):** a scratch restore of every app on demo-hp was measured at **25 s for 8 snapshots / 774 MB**, against 40.3 s for one weekly check — but restic 0.14.0's `--verify` checks size and mtime, **not content**, so nothing available today can vouch for the restored BYTES. Row 4's verdict is UNCHANGED by the spike. `audits/SPIKE-restic-restore-test-2026-08-31.md`. `audits/SPIKE-restic-restore-test-2026-08-31.md`. | | 5 | „Teljes visszaállítás (fájlok + adatbázis)" | `POST /backup/offbox/reconstitute` | destructive to files + DB, never deleting | | 6 | „Megosztások visszaállítása" + place | `POST /backup/shares/{restore,place}` | additive | | 7 | `.fab` import | `POST` → `apiImportStart` | destructive re-import of one app | @@ -885,7 +885,7 @@ crosses the line — **R-158**. | 7 | **Guest stopped and does not come back** | everything | guest-power watchdog (60 s, `onboot` as the deliberate-stop discriminator) | automatic | **120 s** | | **PROVEN** | agent v0.107.0 replay — 120 s unattended vs the incident's 587 s with a human | | 8 | **Host dies (hardware), drives intact** | the data drives; ep0's PBS namespace; the hub's Recipe + Escrow | **install a new host, then** `--selftest=bring-up -mode dr` per guest, **then** re-attach drives by `durable_id` | **operator** (SSH) | | 7 d | **IMPLEMENTED** | bring-up code exists and has **never been executed** (`CAMPAIGN-8…:522`); the drive half of the plan is empty on every box (**R-105**) | | 9 | **Whole box lost (fire/theft) — host and drives gone** | ep0 PBS namespace; the restic repo; the hub's Recipe + Escrow | new hardware → day-0 → escrow-consume with **R** → restore guests from PBS → app data from Tier-3 | **operator + customer** (R) | | 7 d (guest) · 24 h (app data) | **IMPLEMENTED / UNPROVEN** | every leg exists; the composed path has never been run. The destructive S5 drill is operator-gated and unrun (`06-offsite-connectivity.md:327`) | -| 10 | **Ransomware / malicious deletion inside the guest** | PBS offsite (the box **cannot** delete its own snapshots); **the restic repo is NOT protected the same way** | whole-guest restore from PBS to a point before the event | **operator** (SSH) | | 7 d | **PARTIAL** | R-89 proved the box is refused when deleting its own PBS snapshot (`CAMPAIGN-8…:514-515`). **R-95 (open, ranked #1):** the restic credential **can delete** — `readonly=False`, `forget --prune` runs from the box, and SFTP cannot express append-only (`OPEN-ITEMS.md:13`) **2026-08-30 (R-359): the store is now VERIFIED on a cadence** — a daily `offsite-integrity` job runs `restic check` when the last successful one is over 7 days old. **This is a readability check, not a restore-test (R-87 stays open).** The depth that ships ON now re-reads **100%** of the pack data (R-399 CLOSED, controller v0.228.0) — the sentence here previously said the opposite and was stale. **2026-08-31 (SPIKE R-87):** a scratch restore of every app on demo-hp was measured at **25 s for 8 snapshots / 774 MB**, against 40.3 s for one weekly check — but restic 0.14.0's `--verify` checks size and mtime, **not content**, so nothing available today can vouch for the restored BYTES. Row 4's verdict is UNCHANGED by the spike. `audits/SPIKE-restic-restore-test-2026-08-31.md`. | +| 10 | **Ransomware / malicious deletion inside the guest** | PBS offsite (the box **cannot** delete its own snapshots); **the restic repo is NOT protected the same way** | whole-guest restore from PBS to a point before the event | **operator** (SSH) | | 7 d | **PARTIAL** | R-89 proved the box is refused when deleting its own PBS snapshot (`CAMPAIGN-8…:514-515`). **R-95 (open, ranked #1):** the restic credential **can delete** — `readonly=False`, `forget --prune` runs from the box, and SFTP cannot express append-only (`OPEN-ITEMS.md:13`) **2026-08-30 (R-359): the store is now VERIFIED on a cadence** — a daily `offsite-integrity` job runs `restic check` when the last successful one is over 7 days old. **This is a readability check, not a restore-test (R-87 stays open).** The depth that ships ON now re-reads **100%** of the pack data (R-399 CLOSED, controller v0.228.0) — the sentence here previously said the opposite and was stale. **2026-08-31 (SPIKE R-87):** a scratch restore of every app on demo-hp was measured at **25 s for 8 snapshots / 774 MB**, against 40.3 s for one weekly check — but restic 0.14.0's `--verify` checks size and mtime, **not content**, so nothing available today can vouch for the restored BYTES. Row 4's verdict is UNCHANGED by the spike. `audits/SPIKE-restic-restore-test-2026-08-31.md`. **2026-09-01 (SPIKE R-95, `audits/SPIKE-r95-offsite-delete-2026-09-01.md`) — THIS ROW'S STATUS IS UNCHANGED, but the seven-day bound beside it is now known to be unverified.** Measured on BOTH boxes over their own SFTP credential, with controls: **no `.snapshots` is visible to either sub-account**, and the account is jailed. Either none exist or a sub-account cannot see them — and a snapshot the box cannot see is one the box cannot restore from, so recovery would be an operator act at the Hetzner panel, not a product capability (R-429; the confirming field needs the provider API, fenced by §11-D). **The PBS shape does NOT transfer**: the sub-account API has one permission axis, `readonly`, and a backup target cannot be read-only — PBS is a server that can refuse, a Storage Box is a filesystem that runs nothing. **What the spike removed as a fear:** withdrawing delete does NOT wedge the store (measured — restic treats a dead owner's lock as stale and proceeds). **What it added:** `unlock --remove-all` reports success while deleting nothing (R-430). | | 11 | **Hub lost** | every box's data plane, every tier, every Lane-1 route | none needed for recovery **of a box**; the hub itself restores from its Longhorn volume backup | operator (`kubectl`) | | 24 h + weekly (Longhorn `RETAIN 1` each) | **UNPROVEN** | a hub restore has never been performed. The backup target is `nfs://192.168.0.180` — **DooPlex itself** — and exactly **2** restore points exist (LIVE, INV Part D2.2) | | 11b | *consequences while the hub is gone* | — | — | — | | | **[FACT]** | **day:** nothing customer-visible breaks; events queue (`settings.go:1466-1485`). **week:** the operator alarm plane is dark, no claim/reset codes, no config or floor convergence, no PBS-secret re-issue. **permanently:** escrow custody and break-glass credentials are gone (INV Part D2.4) | | 12 | **Offsite provider lost (Hetzner)** | everything on-premises: both drives, both whole-guest tiers | none needed — on-premises recovery is unaffected. Re-provision a new offsite target. | operator | | | **[FACT]** | **restic and PBS share the provider** (INV Part E.1). Whether they share an account and payment method is **UNKNOWN** → §11-D | @@ -1085,7 +1085,7 @@ does **not** hold as written. → **R-108** | **R-356** | The off-site restore resolved its destination with the raw `HDD_PATH` and read an empty answer as "not installed" — **CLOSED, controller v0.219.0, 2026-08-22** | 40 of 53 apps were refused permanently while running (§6.3 `[DESIGN]`) | | ~~**R-108**~~ | ~~Network storage can host an app's namespace~~ | **CLOSED 2026-07-30, controller v0.187.0 — D5 UNBLOCKED.** An app namespace may no longer be placed on network storage (5 surfaces guarded by one fail-closed predicate); the share-root bind is deliberately UNCHANGED because it is load-bearing and unscopable (§10.1). `audits/R108-network-app-namespace-2026-07-30.md` | | **R-126** | A `.fab` bundle — plaintext secrets, optional password — can be exported ONTO a NAS: `storageDriveList()` (`internal/web/handler_export.go`) does not filter network paths | split out of R-108, which closed without it. NOT a D5 precondition: an explicit customer-chosen export destination, not a browsing surface reaching a backup tree (§5, §7.3) | -| R-95 (open) | The restic offsite credential **can delete** — the box can `forget --prune` its own repo | the tier holding the customer's documents and photos is the one whose credential can destroy it (matrix row 10) | +| R-95 (open) | The restic offsite credential **can delete** — the box can `forget --prune` its own repo, from **two** call sites (`offbox.go:1388` retention and `offbox.go:1759` over-quota) | the tier holding the customer's documents and photos is the one whose credential can destroy it (matrix row 10). **SPIKE 2026-09-01:** prevention needs a transport change (restic 0.14.0 does speak `rest:` — measured; append-only is a rest-server flag, not a restic one), because the sub-account API cannot express write-without-delete. **Detection is nearly free and is recommended first** — `snapshot_count` already reaches the hub and the hub appends reports, so the comparison needs no box change. `audits/SPIKE-r95-offsite-delete-2026-09-01.md` | | ~~R-86~~ | ~~Restore-tests are interval-scheduled, not backup-aligned~~ | **CLOSED 2026-08-03 — agent v0.121.0 + hub v0.91.0.** Restore-testing is now **per archive generation**: a tier is due when its newest archive that has settled ~24 h has not been proven, so a daily tier is proved daily on its own archive and a weekly tier weekly on its own. The ticker survives only as the evaluation interval (6 h, chosen from a measured cost). The hub's staleness window moved with it — per tier, from that tier's observed archive rhythm — because a weekly tier proved weekly sat EXACTLY on the old flat 7-day line (§3, Lane 2's per-archive rule) | | ~~R-353~~ | ~~A local unit restore reported a bare completion whether it returned an entire dataset or nothing~~ | **CLOSED 2026-08-30, controller v0.226.0.** The volume count already existed (`restoreDockerVolumesFrom`) and was discarded by a one-line wrapper, so the surface was structurally unable to say what came back. `RestoreFromRecoveryUnit` now returns `UnitRestoreResult` and the sentence has three cases, keyed on replayed-vs-**listed** — zero-replayed has two causes that are opposite news. Every sentence is a claim about the BACKUP, never the app: this path has no `SafetyDump` discriminator, and §6.3 is why that is not pedantry. Proven live on `demo-hp` | | ~~R-357~~ | ~~The destructive reconstitute had no free-space gate; all three that existed guarded non-destructive paths~~ | **CLOSED 2026-08-30, controller v0.226.0.** The gate sits before `writeSafetyDump` and `StopStack`, so a refusal costs no outage — the test asserts the StopStack call count, not the sentence. No headroom multiplier (a measured tree, not a predicted download); fail-closed when either probe reads ≤ 0, which was a real fail-open hole. **Not live-validated** — filling a filesystem is a drill step | @@ -1129,6 +1129,15 @@ and PBS (Cloud server ep0 + a Cloud Volume) are both Hetzner. Whether they share login and payment method** is unverified (INV Unknown U-1) — the question was deliberately not answered by calling the provider API with the production token. Accept explicitly, or mitigate. +> **The fence was reached again, and held, 2026-09-01 (SPIKE R-95).** The R-95 spike needed one +> provider-side fact — **do daily snapshots exist on the Storage Box, and how many** — because the +> register calls that mitigation ARMED and the whole ranking of R-95's options turns on it. It is +> readable from the box: measured on BOTH machines over their own SFTP credential, with positive +> and negative controls, **no `.snapshots` is visible to either sub-account**, and the account is +> jailed at `/`. The register's own confirming field, `size_snapshots`, is an API field. **So the +> spike stopped here rather than answering it by acting**, and left it as R-429 — ten minutes in +> the Storage Box panel, and it re-ranks R-95. `audits/SPIKE-r95-offsite-delete-2026-09-01.md` §Q1. + **E. `local` vzdump shares a physical device with the guest it backs up.** LIVE on both hosts: `/var/lib/vz` (the archive target) and `local-lvm` (the guest's rootfs and both `backup=1` mountpoints) are both on `/dev/sda3` → VG `pve`. The tier therefore protects against **corruption diff --git a/documentation/audits/SPIKE-r95-offsite-delete-2026-09-01.md b/documentation/audits/SPIKE-r95-offsite-delete-2026-09-01.md new file mode 100644 index 00000000..46e1ab99 --- /dev/null +++ b/documentation/audits/SPIKE-r95-offsite-delete-2026-09-01.md @@ -0,0 +1,294 @@ +# SPIKE — can the box be stopped from deleting its own off-site history? (R-95, 2026-09-01) + +**Read-only study. No production code, no version bump, no image, no golden. No delete verb was +issued against any live store. `ep0`, DooPlex and Peti's box were not touched at all.** + +--- + +## Q1 — Is the safety net real? **NO — and worse than "no": it cannot be seen from the box at all, so the exposure must be treated as UNBOUNDED until Viktor looks at the control panel.** + +**Measured** on both live boxes, over the SFTP credential each box already holds, read-only. + +| probe | demo-hp (sub-account A, gid 1058) | demo-felhom (gid 1019) | +|---|---|---| +| **positive control** — the account's own home | lists `.ssh`, `` | lists `.ssh`, ``, `.orphaned-20260810` | +| **negative control** — `./zzz-no-such-r95` | `not found` | `not found` | +| `./.snapshots` | **`not found`** | **`not found`** | +| repo path itself | `config data index keys locks snapshots` | `data index keys locks snapshots` | +| `/.snapshots` | `not found` | `not found` | +| `/` (box root) | **`Permission denied`** — the account is jailed to `/home` | same | + +**Two readings, and the box cannot distinguish them:** either no snapshots exist, or they exist and +are invisible to a sub-account. **Both are bad, and the second is not a reprieve** — a snapshot the +box cannot see is a snapshot the box cannot restore from either. Recovery would be an operator act +through the Hetzner panel, not something the product can do. + +**What the register actually claims**, `OPEN-ITEMS.md:233`, verbatim: + +> `| — | Storage Box **snapshots** on storage-box-pool-1 — plan SET (daily 00:00, keep 7) but **0 taken yet** | WATCHING | first run tonight 00:00 | Confirm size_snapshots > 0 tomorrow; ... until then the mitigation is armed, not proven |` + +**Three facts about that row:** + +1. **It has no R-number** (`| — |`), so no gate and no grep-the-register rule can ever cite it. +2. **Its "confirm tomorrow" was 2026-07-27** (`git log -S`, last touched in `72692e1`). That is + **36 days** ago. Nobody confirmed. +3. **The `DUE-CHECKS` block — the mechanism this project built (R-341) for exactly this — is + EMPTY.** The one dated commitment that mattered was never entered into it. + +**And R-95 leans on it**: its own text reads *"Mitigation now ARMED"*. **That word is doing work it +has not earned.** → **R-429**. + +**Not answerable here:** whether snapshots exist. The register's own confirming field, +`size_snapshots`, is a **Hetzner API field**, and §11-D records the deliberate decision not to call +that API with the production token. **This spike stopped at that fence.** → a question for Viktor. + +**Explicitly NOT done:** no restore from a snapshot was attempted. Whether one *can* be restored is a +separate question and is named as such. + +**What this does to the urgency: it raises it.** The task supposed Q1 might bound the worst case to +seven days. It does not. **The bound is unverified, and unverifiable from the product side.** + +--- + +## Q2 — What must the box write, and what must it delete? **Ten verbs, not nine — and there are TWO `forget --prune` sites, not one.** + +Census by grep over `internal/backup/*.go` at `960d29b`, each confirmed by reading its context. + +| verb | `file:line` | class | +|---|---|---| +| `init` | `offbox.go:321`, `offbox.go:811` | **write** | +| `backup` | `offbox.go:927`, `offbox.go:1332`, `offbox_shares.go:107` | **write** | +| `forget … --prune` | **`offbox.go:1388`** (retention) and **`offbox.go:1759`** (over-quota) | **DELETE** | +| `unlock` | `offbox.go:746` (stale-only), `offbox.go:768` (`--remove-all`) | **DELETE** | +| `restore` | `offbox.go:1827`, `offbox_restore.go:393`, `offbox_proof.go:328` (`--no-lock`) | read | +| `snapshots` | `offbox.go:1771`, `offbox_inventory.go:73`, `offbox_restore.go:126` | read | +| `stats` | `offbox.go:1786`, `offbox_restore.go:160` | read (**takes a lock**) | +| `cat` | `offbox.go:782` | read | +| `check` | `offbox_integrity.go:316` | read (**takes a lock**) | + +**Two corrections to the task's list, both mine to report:** + +* **`check` is missing from it.** It is a real off-site verb (`offbox_integrity.go:316`) and it is + lock-taking, which makes it central to Q5. +* **`dump` is NOT a verb.** `offbox_progress.go:185` is `OffboxPhaseDump = "dump"`, a progress phase + name. I withdrew it after reading the context rather than counting it. + +**Does the box need the delete verbs?** `forget`/`prune` — **no**, that is retention and it is exactly +what R-89 moved off the box for PBS. `unlock` — **this is the whole of Q5**. + +--- + +## Q3 — Can the sub-account express "write but never delete"? **No. The API offers exactly one axis, and it is all-or-nothing.** + +**DOCUMENTED, not measured** — from this project's own mirror of the API, +`hub/internal/hetznerapi/hetznerapi.go:38-45`, which is the authority we have without calling the +provider: + +```go +type AccessSettings struct { + SSHEnabled bool `json:"ssh_enabled"` + ReachableExternally bool `json:"reachable_externally"` + SambaEnabled bool `json:"samba_enabled"` + WebDAVEnabled bool `json:"webdav_enabled"` + Readonly bool `json:"readonly"` +} +``` + +**Five booleans. `readonly` is the only permission axis, and a backup target cannot be read-only.** +There is no append-only, no per-directory grant, no write-without-unlink. **This is the finding that +forces Q6**, and the register already suspected it. + +**Measured corroboration, no new action taken:** demo-felhom's home contains +a `.orphaned-20260810` directory beside ``. That directory was **renamed by the controller** +(the orphan guard). A rename is a delete-class operation. **So the credential's reach is not a +theory — the existing state on the live store is evidence of it.** + +**Not established:** whether the live API would report anything the struct omits. Settling it needs +the provider token — **fenced** (§11-D). + +--- + +## Q4 — What did the PBS fix do, and what would copying it cost? **The shape does not transfer, because the far end is a disk, not a server.** + +**What moved (R-89, PROVEN-LIVE):** retention became a hub-owned commercial attribute; boxes set +`keep_last: 0`; **ep0 runs the prune jobs**; box tokens stay write-only. `07` §8 row 10 records the +box being *refused* when deleting its own PBS snapshot. + +**Why it does not transfer as-is:** PBS is a **server** that can refuse. The restic tier writes to a +**Storage Box over SFTP** — a filesystem. **A filesystem runs nothing.** So "move retention to the far +end" has no far end to move it to. Candidates, with their costs: + +| who runs retention instead | cost | new risk it creates | +|---|---|---| +| **the hub** | a scheduler + a per-customer credential | **the hub gains a credential that can delete every customer's history** — one compromise instead of N | +| **DooPlex** | a cron + the same credentials | same concentration, on a **Tier-2 box that is itself the recovery chain** | +| **ep0** | a service on a protected machine | **an architecture change, not a config change**; ep0 is protected | +| **nobody — never prune** | £0 today | the store grows without bound; the over-quota path at `offbox.go:1759` exists precisely because quota is already a live concern | + +**What the migration cost last time, and the trap is identical here.** R-191: R-89 changed the +contract, one client-side retention setting did not follow, and **every successful weekly backup then +reported as a failure** — upload complete, 67.2 % reused, then `TASK ERROR: job errors` and a +`whole_guest_backup_failed` alarm to the operator. *"A doc that states the contract does not enforce +it — the gate does."* + +**The same trap, doubled:** withdrawing delete without disarming retention would make every off-site +run log `[WARN] forget --prune failed` — and there are **two** call sites, `offbox.go:1388` **and +`offbox.go:1759`**. The task's brief names only the first. **Disarming one and not the other +reproduces R-191 exactly.** + +--- + +## Q5 — The lock problem. **Measured, and it is NOT the blocker I expected — but it exposed something worse: `unlock` reports success on a deletion that did not happen.** + +**Method.** A throwaway **local** restic repo in the controller container's `/tmp` on demo-hp (60 MB, +removed at the end, **no live store touched**). Append-only was modelled faithfully: a sticky locks +directory (`1777`) owned by root, holding a root-owned lock, with restic run as `nobody`. That is +exactly append-only semantics — **create allowed, delete refused**. + +**Both controls passed before anything was believed:** + +* `nobody` **can** create in the locks directory → the model is not simply "read-only". +* `nobody` **cannot** delete root's lock → the model really does refuse deletes. + +A **genuine** restic lock was captured (copied out while a real `check --read-data` held it), not +hand-forged. + +| verb, running as `nobody` against an undeletable stale lock | result | +|---|---| +| `restic unlock --remove-all` | **prints `successfully removed locks` — and the lock is STILL THERE** | +| `restic backup` | **succeeds** — `snapshot a7928829 saved` | +| `restic check` | **succeeds** — `no errors were found` | +| `restic snapshots --no-lock` | succeeds | +| `restic restore --no-lock` (the R-87 proof pattern) | succeeds | + +**So the answer to "does the store wedge?" is: not in this case.** restic 0.14.0 treats a lock whose +owner is provably dead as stale and proceeds without needing to remove it. **Withdrawing delete does +not, by itself, wedge the store.** That removes the constraint the task suspected would be deciding. + +**But the measurement found a different defect.** `unlock --remove-all` **reported success while +deleting nothing.** This project already has a named class for that — *"Exit codes that lie"* — and +`resticStep`'s crash-lock self-heal is built directly on top of this call. **A self-heal that cannot +fail is a self-heal that cannot be trusted.** → **R-430**. + +**UNKNOWN, and I am not claiming otherwise:** the **crash-lock** case — a lock left by a container +that no longer exists, whose hostname restic cannot match, so it will **not** treat it as stale for +~30 minutes. That is documented in `resticStep`'s own comment (`offbox.go:~760`) and is the reason +`--remove-all` exists at all. My model could not reproduce it: the captured lock carried this +container's own hostname. **What would settle it:** a lock captured from a container with a different +hostname, replayed against an append-only endpoint. **In that window, the only remedy is the very +call that Q5 just showed reports success while doing nothing.** + +**How far `--no-lock` reaches:** restic's own help says it *"allows some operations on read-only +repositories"*. Measured: `snapshots` and `restore` work under it. `backup` and `check` still take a +lock — but **creating** a lock is a write, which append-only permits. So the read paths are safe and +the write paths are unaffected; **only removal is denied, and only the crash-lock window depends on +removal.** + +--- + +## Q6 — Is an append-only transport reachable? **Yes in principle — restic 0.14.0 does speak REST — but not without either moving the data or putting a machine in front of it.** + +**Measured**, in the controller container, with a control: + +| repo string | result | +|---|---| +| `banana:http://127.0.0.1:1/x` (control) | `Fatal: parsing repository location failed: invalid backend` | +| `rest:http://127.0.0.1:1/` | `Fatal: unable to open config file: … dial tcp 127.0.0.1:1: connect: connection refused` | + +**A connection error, not a parse error — the REST backend IS recognised by restic 0.14.0.** + +**Append-only is not a restic feature.** Measured: `restic help` and `restic backup --help` contain +**zero** occurrences of "append". It is a **rest-server** flag (`--append-only`). So the client is +ready and the server does not exist yet. + +**In front, or move the data?** rest-server serves a **local directory**. The Storage Box is a remote +share. So either (a) a machine mounts the Storage Box and runs rest-server over that mount — keeping +the bytes where they are, at the cost of a mount in the hot path — or (b) the data moves to storage +attached to whatever runs rest-server. **These are very different prices and the choice is not +obvious**; (a) keeps the current bill, (b) is a migration of every customer's history. + +**Where would it run?** **`ep0` is protected. Adding a service to it is an architecture change, not a +configuration change** — and it also makes ep0 a single point of failure for both tiers, which is +precisely the concern §11-D already has open. + +**Cost:** one new always-on service in the recovery path, plus whatever machine hosts it. **The moving +part matters more than the money:** if rest-server is down, backups stop — and this project's own +history (R-191) is a warning about what happens when a change in the off-site contract is not carried +everywhere at once. + +--- + +## Q7 — What does DETECTION cost? **Almost nothing. The number already arrives at the hub, and the hub already keeps the history to compare it against.** + +**Measured from source:** + +* The box already reports `snapshot_count` — `controller/internal/report/types.go:131`, and it is + already rendered on the hub's Backup card (`hub/internal/web/backup_card.go:118`). +* **The hub retains report history**: `store.go:968` is an `INSERT INTO reports`, and the read is + `SELECT … ORDER BY id DESC LIMIT 1` (`store.go:1016`). **It appends; it does not replace.** So every + previous `snapshot_count` for every customer is already on disk. + +**So an unexplained drop is a comparison between two rows the hub already has.** No new box code, no +new credential, no new moving part, no new bill. The alarm vocabulary and the operator-only routing +already exist. + +**The honest caveats:** a legitimate `forget` also drops the count, so the rule needs to know +retention (and today the box prunes itself, which is what makes the number noisy — **the same change +that disarms box-side retention is what makes this signal clean**). And detection is not prevention: +it tells you within a day that history was destroyed; it does not stop it. **`snapshot_count: 0` must +mean UNKNOWN, not EMPTY** — `backup_card.go:29` records that exact defect being fixed once already. + +--- + +# Ranked options for Viktor + +**All four were considered. My recommendation is 4 + 2, in that order, and explicitly not 3 yet.** + +### 1. Do nothing — **NOT acceptable as it stands, and Q1 is why** + +Before this spike, "do nothing" meant *"exposed, but bounded to seven days by snapshots."* **That +sentence is not supported by any evidence.** The box cannot see a single snapshot, on either machine, +and the register's confirmation step was never done. **Doing nothing now means an unbounded exposure +on the tier holding the customer's documents and photos.** +**Cost:** £0, no evenings. **What happens if you choose it:** the risk stays exactly where it is, and +the register keeps saying "ARMED", which is the part I would not accept. +**One cheap act rescues most of this option: look at the Storage Box panel and answer Q1.** If +snapshots are real, "do nothing" becomes defensible again. **That is ten minutes and it is yours to +do — the fence in §11-D stopped me.** + +### 2. Copy the PBS shape — **worth doing, but it is smaller than it sounds and it is not the safety it appears to be** + +Withdraw box-side retention; someone else prunes. **Q5 says this no longer looks dangerous** — a stale +lock does not wedge the store. +**But Q3 says the credential still cannot be narrowed**: the API has one axis, `readonly`, and a +backup target cannot be read-only. **So the box would keep the ability to delete, and simply stop +using it.** That is a discipline, not a guarantee — a compromised guest is unaffected by it. +**Cost:** one evening, plus wherever retention moves (each candidate concentrates the credential). +**Trap:** both `forget` sites (`offbox.go:1388` **and** `:1759`) must be disarmed in the same change, +or you get R-191 again — successful backups reported as failures. + +### 3. Change the transport (rest-server `--append-only`) — **the only real prevention, and I would not start it yet** + +**Q6 says it is reachable**: restic 0.14.0 speaks REST. This is the option that actually makes the +box *unable* to delete. +**Cost:** a new always-on service in the recovery path; a machine to host it (**ep0 is protected — +this is an architecture change**); and either a mount in the hot path or moving every customer's +history. **Money is the small part.** +**Why not yet:** it should be chosen against a measured picture, and one measurement is still missing +— Q1. Building the expensive prevention while nobody knows whether a seven-day net already exists is +the wrong order. + +### 4. Detect instead of prevent — **cheapest by a wide margin, and I would do this first** + +**Q7 measured that the material is already there**: the count is on the wire, the hub keeps the +history, the alarm path exists. An unexplained drop in `snapshot_count` becomes noticeable within a +day. +**Cost:** hub-side only. No box change, no credential change, no new service, no new bill. +**What it does not do:** it does not stop the deletion. It converts "we would never know" into "we +know tomorrow" — and against ransomware inside the guest, knowing tomorrow is the difference between +losing a day and losing everything silently. +**This week's pattern held twice already** (R-87, R-404): **proving beats preventing when preventing +is expensive.** This is the third instance. + +**My pick: answer Q1 today (yours, ten minutes), then build 4, then 2. Revisit 3 once Q1 is +answered** — and if Q1 comes back "no snapshots", 3 moves up sharply. diff --git a/documentation/backlog/OPEN-ITEMS.md b/documentation/backlog/OPEN-ITEMS.md index 4e6340ed..11e85a33 100644 --- a/documentation/backlog/OPEN-ITEMS.md +++ b/documentation/backlog/OPEN-ITEMS.md @@ -224,7 +224,7 @@ unconditional promise, which CAMPAIGN-11 §7 step 7 measured the customer-facing | **E-2d** | **Prove E-2 on a fresh VM** — a real `felhom-host-install.sh` 1.22.0 run, Case B naturally, a claimable customer, then add a drive (the offer) and unplug it (`backup_target_absent` end-to-end) | **CLOSED — PARTIALLY PROVEN** (2026-07-29) | — | **C1, C2 proven** (`audits/E2D-fresh-vm-2026-07-29.md`); **C3, C4 proven live** (`audits/SESSION-C-2026-07-29.md`); **C5 FAILED → R-116** — the gate fires and an alarm reaches the hub, but it is the generic event, so the alarm and its recovery cannot be paired. **R-116 is the single named open leg**; per the Session-C runbook §9, decided in advance, a failed claim closes the item as partially proven rather than triggering a re-run. Both audits carry the full record — the `local-lvm` fence, the ISO/PAIRING derivation, the Phase 0 answers, the per-claim observables and the teardown evidence — and are the place to read it, not this cell. **The arc's actual definition of done is R-106 + R-109, R-108 and D5**, none of which this detour touched | CC | | **R-121** | **A BOX's installed agent can sit releases behind the vouched one and nothing notices — the R-120 gate does not cover it.** demo-hp ran agent **0.113.0** while the hub vouched **0.116.0**, through the whole R-116/R-117 arc, and no signal existed on any channel | **READY (S) — NEW 2026-07-30** | — | **Fourth instance of the drift family** (R-111 golden's agent 17 releases behind, R-115 built+deployed but never published, R-120 golden a controller behind — and now installed-vs-vouched on a live box). **Confirmed at source that R-120's gate cannot catch it:** `hub/internal/web/configs.go:1165-1169` compares `goldenVer` against `store.NewestReportedControllerVersion()` — it is a **golden-artifact vs fleet-CONTROLLER** check and says nothing about the agent installed on a box. **`MinAgent` does not cover it either:** it is used to HOLD the controller floor for a box whose agent is too old (`hub/internal/api/handler.go:530-538`, `store.go:1857`) — protective, not an alarm — and demo-hp's 0.113.0 **equalled** `min_agent` 0.113.0, so even a floor comparison was satisfied. **The cost, measured:** R-117's whole subject is the R-113 conjunction, which landed in **0.114.0** — so the designated drill host could not exercise the code under investigation at all, and the R-117 spike had to route every predicate result through an out-of-repo probe built from `main` instead of the installed agent (`audits/SPIKE-r117-bind-liveness-2026-07-30.md` §1, §2.3). Discovered because the R-117 task made bringing the box current an explicit prerequisite. **Fix shape (not implemented):** the hub already receives `AgentVersion` on every host report, and already has semver comparison in Go — the missing piece is a checker comparing reported agent vs the **vouched** agent and surfacing it, operator-tier. Note the honest tension: a box legitimately lags between publish and deploy, so this wants a staleness window rather than an instant alarm | CC | | **R-118** | **An absent drive's union row advertises the ROOT filesystem's capacity as its own.** In the absent-state payload the registry-union row reports `total_bytes: 49675956224 / used_bytes: 4584579072` — **byte-identical to the `local` row** (`durable_id: path:/var/lib/vz`, i.e. `pve-root`) in the same response. The real drive is **4 GB** | **READY (XS) — NEW 2026-07-30** | — | Cause: `statfsCapacity(d.MountPath)` (`disks.go:335-338`) statfs's `/mnt/cel`, which with the device gone is a **bare directory on the root filesystem**. `observe.go:176-183`'s comment warns about exactly this trap and guards the Observe path ("*an unmounted removable dir-storage's mountpoint reverts to a bare directory on root … catastrophic DR mis-id*"); **the union path has no equivalent guard.** **Not a DR mis-id** — `durable_id` on that row is still the correct `uuid:…`, so re-attach identity is safe. It is a **false capacity** reaching every consumer of `total_bytes`/`used_fraction` (fill monitors, storage cards): a detached 4 GB drive advertises 46 GiB at 9.2 % used. Same class as `role.go:180-181` — an absent drive's fields decaying to the root filesystem's. Evidence: `audits/DIAG-r116-disks-payload-2026-07-30.md` §12 | CC | -| **R-95** | restic offsite credential **can delete** (`readonly=False`, `forget --prune` runs from the box); SFTP cannot express append-only | **READY** | — | Root exposure still open. Mitigation now ARMED — split prune off-box or move to REST `--append-only` | CC | +| **R-95** | restic offsite credential **can delete** (`readonly=False`, `forget --prune` runs from the box); SFTP cannot express append-only | **READY** | — | Root exposure still open. Mitigation now ARMED — split prune off-box or move to REST `--append-only` | CC **SPIKE 2026-09-01 — `audits/SPIKE-r95-offsite-delete-2026-09-01.md`. THE WORD "ARMED" ABOVE IS NOT SUPPORTED AND IS WITHDRAWN PENDING R-429:** no `.snapshots` is visible to either box's sub-account (measured, both machines, with controls), so the seven-day bound is unverified and unverifiable from the product side. **Q3 (documented, `hub/internal/hetznerapi/hetznerapi.go:38-45`): the sub-account API has ONE permission axis, `readonly` — there is no append-only, so the PBS shape does NOT transfer** (PBS is a server that can refuse; a Storage Box is a filesystem that runs nothing). **Q5 (measured): withdrawing delete does NOT wedge the store** — restic treats a dead owner's lock as stale and proceeds — so the constraint everyone feared is not the blocker; but `unlock --remove-all` lies about success (R-430), and the crash-lock window is UNKNOWN. **Q6 (measured): restic 0.14.0 DOES speak `rest:`** (control: `banana:` → invalid backend), and append-only is a rest-server flag, not a restic one — reachable, but it needs a machine in the recovery path and ep0 is protected. **Q7 (measured): detection is nearly free** — `snapshot_count` already reaches the hub and the hub APPENDS reports, so the comparison needs no box change. **RECOMMENDATION: answer R-429 first (Viktor, ten minutes), then build detection, then move retention off the box; defer the transport change.** Two `forget --prune` sites must be disarmed together — `offbox.go:1388` AND `offbox.go:1759` — or R-191 repeats. | | **R-191** | **Every weekly offsite backup UPLOADS successfully and then FAILS the job on a prune the box is deliberately not allowed to do — on both demo boxes.** Measured on demo-felhom 2026-08-04 06:49–06:53: the upload completed (223 s, 629 MiB of 1.874 GiB, **67.2 % reused incrementally**), then `ERROR: prune 'ct/9201': proxmox-backup-client failed: Error: permission check failed - missing Datastore.Modify\|Datastore.Prune on /datastore/felhom-offsite/demo-felhom` → `ERROR: Backup of VM 9201 failed - error pruning backups` → `TASK ERROR: job errors`. The hub raised `whole_guest_backup_failed` | **CLOSED — SHIPPED 2026-08-04** (installer **1.25.0**; both live boxes corrected) | — | **This is R-89's rule not reaching the config.** R-89 moved PBS pruning SERVER-SIDE — *"boxes set `keep_last: 0`, ep0 runs prune jobs; box tokens stay write-only, never widen the grant"*. The token behaves exactly as designed: it refuses. But **both** demo boxes still arm the offsite tier with `keep_last=2 prune_pbs_allowed=true` (`backup_targets: [{target_id: felhom-pbs, cadence_seconds: 604800, keep_last: 2}]`), so every run asks for a prune that must fail. **The data is SAFE and that is why this is not a P1:** the snapshot lands before the prune is attempted; what is wrong is the job's VERDICT and the weekly operator e-mail it produces. **But it is corrosive in the specific way this project keeps finding:** a backup that reports FAILED while succeeding trains the operator to discount `whole_guest_backup_failed`, which is the same alert that would carry a real one — and it is exactly the failure the R-100 corollary warns about, an alarm whose text is true and whose trigger is not the thing you would act on. **Fix is one config line per box** (`keep_last: 0` on the PBS tier) plus whatever writes it on a fresh install; **deliberately NOT applied in this session** — the session was a runbook with an explicit "change nothing, and if a change appears necessary, stop and report" rule, and a retention field on a live backup tier is not a change to slip into an observation run. **Check before fixing:** whether ep0's prune jobs actually cover these two namespaces, or the snapshots simply accumulate once the box stops asking **THE GATE WAS RUN FIRST, AND IT MATTERED.** Before disabling anything, ep0 was read (read-only, Tier 2): prune jobs `prune-demo-felhom` and `prune-demo-hp` exist on datastore `felhom-offsite`, one per namespace, `schedule 03:30`, `keep-last 2`, comment *"R-82 retention keep-last=2, server-side (box tokens are write-only)"* — and they have run **every day since 2026-07-27: 18 tasks, all `status=OK`**. The newest task log reads `retention options: --ns demo-felhom --max-depth 0 --keep-last 2` / `keep ct/9201/2026-07-27…` / `keep ct/9201/2026-07-28…` / `TASK OK`. Retention happens, and it happens there. **A METHODOLOGICAL WARNING WORTH MORE THAN THE FIX.** Three separate queries said the OPPOSITE — *no prune jobs have ever run* — and **all three were broken instruments**: `worker-type` where the field is `worker_type`; the value `prune` where the worker type is `prunejob`; and `journalctl -u proxmox-backup` where the unit is `proxmox-backup-proxy`. A fourth reading (3 snapshots under keep-last 2) was mis-framed by CC and self-corrected — the third snapshot had landed AFTER that day's 03:30 window. Acting on any of them would have disabled the only pruning ATTEMPT while reporting that nothing prunes: a weekly false alarm traded for unbounded growth on the protected endpoint, invisible for months. **The gate is what caught it, and only because it demanded evidence rather than a verdict.** **Shipped:** installer **1.25.0** writes `keep_last: 0` on the offsite tier (the agent's existing guard `allowPBSPrune = !primary && keep_last > 0` already reads that as *never prune from the box* — no agent change), the justifying paragraph is rewritten to say where retention lives and cite R-89, and `hostinstall_gates.py` asserts it (red-proved: pinning `keep_last: 2` back fails the gate). **Both live boxes corrected in their own config** — `backup tier armed target=felhom-pbs … keep_last=0 … prune_pbs_allowed=false` on demo-felhom and demo-hp, with the local tier untouched at `keep_last=3`. Served over HTTPS at `1.25.0` with `"keep_last":0` in the served bytes. **STILL TO OBSERVE:** the next weekly offsite run completing OK end-to-end. The change removes the failing step; the *schedule* proving it is next week's event, and this row should carry that line when it happens. | CC | | **R-194** | **PVE's permission cache delays every grant-state verdict by an unknown amount, so "the agent can read it" and "the ACL exists" are not the same measurement.** Observed twice while validating R-190's self-repair on demo-felhom 2026-08-04: both ACL rows for `/storage/felhom-backup` were deleted, and `GET /access/permissions` continued to report `Datastore.AllocateSpace` present — for **~40 s** in one run and **~16 minutes** in another. During that window the capability probe reads healthy and the self-repair does not fire | **OPEN** | — | **Why it matters beyond the delay:** it puts a floor under how fast a lost grant can be noticed, it makes any single permission read a lagging indicator, and — the interesting part — **it is a candidate contributor to R-190's own timeline**: a grant removed at an unknown moment could keep working until a cache expiry, which is exactly the shape of *worked at 04:44, refused at 09:24*. That does not explain what removed it, but it may explain when the refusal SURFACED, and the two have been treated as the same instant. **Not a defect in our code** — it is PVE behaviour, and the mitigation already tolerates it (the repair fires on the next probe after the cache clears). **What is worth deciding:** whether the store-grant probe should ALSO consult the storage content listing as a second signal, since that appeared to reflect the loss immediately (`{"data":[]}` while the permission read still said present) — two signals disagreeing is itself information, and today only one of them is read | CC | | **R-200** | **The DR password-injection seam has a handler, a route and tests — and no form.** `POST /backup/offbox/inject-password` is routed (`controller/internal/web/server.go:510`) to `offboxInjectPasswordHandler` (`offbox_handlers.go:174-196`) → `InjectOffboxPassword` (`backup/offbox.go:541`). **No template in the repository contains that path or any form posting to it** (grep over `internal/web/templates/`: one unrelated hit, an XSS comment) | **PLUMBING COMPLETE** (controller **v0.196.0**); **the FORM is not built — still open** | — | **The tenth instance of this project's built-but-never-wired class, and the exact shape `CLAUDE.md` and `felhom.eu/CLAUDE.md`'s seam-wiring rule were written for:** handler tests that POST directly (`offbox_escrow_test.go:167,180`) prove nothing about reachability. To use the only implemented recovery seam today, a person must hand-craft an authenticated POST with a session cookie and CSRF token. **Note the layering while fixing it:** this form takes a **64-hex repo password**, not a recovery code (`offboxRepoPwPattern`, `offbox.go:543`) — they are different secrets at different layers, and the operator's 2026-08-04 ruling asks for a form that takes **R**. Build the R form and treat this one as the operator/DR fallback it was written as, but **ship it with a render test per branch of whatever gate it sits behind**. Source: `audits/RECON-offsite-dr-chain-2026-08-04.md` §3 link 9 **THE DIAGNOSTIC HALF IS DONE AND IT ANSWERED THE QUESTION.** `--recover-offsite-check` is a `docker exec` escape hatch in the shape of `--print-reset-code`: R on **STDIN** (never argv, never `ps`, never shell history, never a transcript), fetch+unseal via the agent, and a verdict of **two sha256 hashes**. **It compares and never installs** — the recovered password is not written to `offbox/repo_password`; a test asserts the data dir is byte-unchanged and its red-proof (adding the install call) fails it. Confirmed live: `repo_password` mtime still `2026-08-03 07:18:02` after the successful check at `2026-08-04 11:49`. **Exit codes are load-bearing** — `0` match, `2` a clean MISMATCH, `1` a step failed; "it failed" and "it worked and disagreed" must never share a status because only one is a finding about the system. A box with no local password reports distinctly (the rebuilt-box shape, where the next step is to INSTALL rather than compare). **WHAT IS NOT BUILT, deliberately:** no card, no form, no preview, no customer-facing text — building an interface on top of a chain nobody had walked is how the preceding three weeks went wrong. **What remains for this row:** link 9 (the recovered password placed so `WriteOffboxSecrets` keeps it) and the customer-facing shape the operator ruled on 2026-08-04 (yell → R form → preview → proceed), which is now priced against a chain that exists rather than one that is assumed **PART 0 SHIPPED 2026-08-04 (v0.196.0):** `--recover-offsite-install` is the sibling of the check — same fetch/unseal path, same STDIN discipline for R — and it **places** the recovered password via `InjectOffboxPassword`. The confirmation is a SECOND invocation (`--confirm-install`): without it, both hashes print and nothing is written, so the operator sees the comparison before a write is possible. Three outcomes named distinctly: **installed** (no local password — the rebuilt-box shape), **unchanged** (identical key, nothing written), **refused** (a DIFFERENT key present — installing would clobber the key the current repository is encrypted under; exit 2, no force offered). It re-reads the file after writing rather than trusting the call. **Red-proof observed:** removing the confirmation gate makes the dry run write the password. The R-persistence test carries a **positive control** (a planted copy found, then removed and not found). **NOT YET EXERCISED AGAINST A LIVE RECOVERY** — the R-201 drill halted before step 9, so this is unit-proven only. **What remains for this row:** the customer-facing shape the operator ruled on 2026-08-04 (yell → recovery-code form → preview → proceed) | CC | @@ -602,6 +602,8 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server` | **R-426** | **The decoy-coverage exemption list — 20 registered gates that ship WITHOUT a decoy test, each named.** `scripts/decoy_coverage_gate.py`'s `EXEMPT` map is debt, and this row owns it so it lives in the register and not only in a Python literal. **Four kinds:** (a) genuinely covered in the 2026-09-01 sweep but not yet moved into a suite — `hub-copy`, `instructions`, `docker-v`, `image-pins`; (b) blocked by an open hole and therefore un-assertable as rejecting — `site` (R-423), `one-register` (R-424), `offbox-rename` (R-425); (c) shared scripts whose decoy lives in `felhom.eu` and is counted there — `reuse-refs`, `instructions`, `observations` in the controller and agent runners; (d) **no plausible decoy constructed yet** — `hostinstall`, `wire-contract`, `due-checks`, `published`, `image-resolvable`, `volume-persistence`. Group (d) is the honest unknown: six gates whose soundness is UNTESTED, not established. **The list is green today and shrinks; a NEW gate with no decoy fails immediately.** | **OPEN — 20 names; group (d) is six untested gates** | | **R-427** | **`closed_register_gate.py` checks ONE direction only: an open word in a CLOSED row. The mirror — a CLOSED verdict on a row still sitting in `OPEN-ITEMS.md` — is unchecked, and there are TWELVE.** MEASURED 2026-09-01 during the decoy sweep, by reading the leading verdict of every open row with the gate's own predicate: **R-385, R-387, R-341, R-378, R-405, R-88a, R-88b** read unambiguously closed; **R-123, R-190, R-352** read `PARTLY CLOSED` / `MITIGATION SHIPPED` and almost certainly belong where they are. **The rows were NOT moved by this session** — telling a finished row from a partly-finished one is a judgement, and R-378 is itself the record of what happens when a machine makes that judgement on a substring (six still-open rows moved out of the register). **This is R-405's finding mirrored:** that row exists because R-87 sat in the CLOSED file while its state read READY, and the gate written for it looks only the way it was bitten. Fix: the same leading-verdict predicate applied to `OPEN-ITEMS.md`, reporting rather than convicting until the twelve are adjudicated by a person — a gate registered while twelve rows fail it would refuse every push. | **OPEN — 12 rows named; the adjudication is Viktor's, the gate is mine** | | **R-428** | **The decoy-coverage gate — written to catch instruments that match a NAME instead of a fact — identified a repository by its DIRECTORY NAME.** MEASURED on its own first CI run (felhom.eu job 490, 2026-09-01): `os.path.basename(root)` looked up in a `RUNNERS` map, and Gitea's act-runner checks the repo out into a directory called `hostexecutor`, so the gate reported *"unknown repo 'hostexecutor'"* and went INCONCLUSIVE. **The gate that hunts label-matching was matching a label, in the first ten lines of its own main loop, and it shipped that way.** FIXED the same day: it now identifies a repo by which registered runner FILE exists under the root, which is a fact. **Recorded rather than quietly patched because it is the strongest evidence in the sweep that this class is not a matter of carelessness** — it was written by a session that had spent the morning reading 29 gates for exactly this, with the four shapes on screen. Verified under a renamed directory before and after. | **CLOSED 2026-09-01 — fixed, and kept as the class's best example** | +| **R-429** | **The Storage Box snapshot mitigation that R-95 calls "ARMED" has never been confirmed, cannot be seen from either box, and its row has no id.** MEASURED 2026-09-01 over the SFTP credential each box already holds, read-only, with positive and negative controls on both machines: **no `.snapshots` directory is visible to either sub-account** — not in the account home, not inside the repo path — while the account is jailed (`/` returns `Permission denied`) and a bogus path errors correctly. Two readings the box cannot distinguish: no snapshots exist, or they exist and a sub-account cannot see them. **Both are bad, and the second is not a reprieve — a snapshot the box cannot see is one the box cannot restore from, so recovery would be an operator act through the Hetzner panel, not a product capability.** The claim lives at `OPEN-ITEMS.md:233` as a row whose ID cell is a bare em-dash instead of an R-number, so no gate and no grep-the-register rule can cite it; its "confirm tomorrow" was **2026-07-27** (last touched in `72692e1`), **36 days** unconfirmed; and the `DUE-CHECKS` block built for exactly this (R-341) is **EMPTY**. **The confirming field, `size_snapshots`, is a Hetzner API field and §11-D fences it — so this spike stopped and left it for Viktor.** Fix: Viktor reads the Storage Box panel once (ten minutes) and the answer goes in this row; give the WATCHING row an id or delete it; and R-95's word "ARMED" must not stand until then. Evidence: `audits/SPIKE-r95-offsite-delete-2026-09-01.md` §Q1. | **OPEN — Q1 is Viktor's, and it re-ranks R-95** | +| **R-430** | **`restic unlock --remove-all` printed `successfully removed locks` while the lock was still there.** MEASURED 2026-09-01 in a throwaway local repo (no live store touched), under a faithful append-only model — a sticky locks directory owned by root holding a root-owned lock, restic run as `nobody`; both controls passed first (create allowed, delete refused). The command reported success, returned, and `ls` showed the lock present. **`resticStep`'s crash-lock self-heal is built directly on this call** (`felhom-controller/controller/internal/backup/offbox.go:~768`), and its licence to escalate rests on the escalation actually working. **A self-heal that cannot fail is a self-heal that cannot be trusted** — this is this project's *"exit codes that lie"* class, in the one path that runs unattended against the customer's off-site history. It is harmless TODAY because the credential can delete and the removal really happens; it becomes load-bearing the moment delete is withdrawn, which is what R-95 is about. **Not yet established:** whether restic reports success because it removed zero locks by design, or because it did not check. Settling it: read restic 0.14.0's unlock source, or re-run with `--verbose`. | **OPEN — precondition on any R-95 build** |