SPIKE R-95: the safety net cannot be seen from the box - and that RAISES the urgency
gates / gates (push) Successful in 17s
gates / gates (push) Successful in 17s
READ-ONLY STUDY. No code, no version, no image, no golden. No delete verb was issued against any live store. ep0, DooPlex and Peti's box were not touched at all. Q1 FIRST, AND IT DID NOT GO THE EXPECTED WAY. The brief supposed a working seven-day snapshot net might bound the worst case. Measured on BOTH boxes over their own SFTP credential, with a positive and a negative control on each: NO .snapshots is visible to either sub-account - not in the account home, not inside the repo - and the account is jailed at /. Either none exist or a sub-account cannot see them, and the second is not a reprieve: a snapshot the box cannot see is one the box cannot restore from, so recovery would be an operator act at the Hetzner panel. The register's claim rests on nothing that was checked. The row has NO R-number, so nothing can cite it; its "confirm tomorrow" was 2026-07-27, 36 days ago; and the DUE-CHECKS block built for exactly this (R-341) is EMPTY. R-95's word "ARMED" is withdrawn pending R-429. The confirming field is a Hetzner API field, so this spike STOPPED at the section 11-D fence and left it for Viktor - ten minutes in the panel, and it re-ranks everything. Q2: TEN verbs, not nine. `check` was missing from the brief's list; `dump` is not a verb (it is a progress phase constant) and was withdrawn. There are TWO `forget --prune` sites - offbox.go:1388 AND offbox.go:1759 - and disarming one without the other reproduces R-191 exactly. Q3 (documented, from our own API mirror): AccessSettings has five booleans and `readonly` is the only permission axis. No append-only. So the PBS shape does NOT transfer - PBS is a server that can refuse; a Storage Box is a filesystem that runs nothing. Q5 (measured, faithful append-only model, both controls passed first): withdrawing delete does NOT wedge the store - restic treats a dead owner's lock as stale and proceeds. The constraint everyone feared is not the blocker. But `unlock --remove-all` printed "successfully removed locks" while the lock was still there, and resticStep's crash-lock self-heal is built on that call - R-430. Q6 (measured, with a control): restic 0.14.0 DOES speak rest:. Append-only is a rest-server flag, not a restic one. Q7 (measured): detection is nearly free. snapshot_count already reaches the hub and the hub APPENDS reports, so the history to compare against is already on disk. RECOMMENDATION: answer Q1 today (Viktor, ten minutes), then build detection, then move retention off the box. Defer the transport change until Q1 is answered. Register: OPEN 179 -> 181. Filed R-429, R-430; R-95 updated and kept OPEN. 07 row 10's status is deliberately UNCHANGED.
This commit is contained in:
@@ -98,7 +98,7 @@ with nothing but their dashboard password. No operator, no ticket, no scheduling
|
||||
| 1 | „Visszaállítás indítása" | `POST /backup/restore` | destructive — rebuilds the app from its Tier-1 unit |
|
||||
| 2 | „Fájlok visszaállítása" | `POST /backup/tier2/restore` | additive, missing-only |
|
||||
| 3 | „Visszaállítás a távoli tárolóból" | `POST /backup/offbox/restore` | non-destructive, to a verification copy |
|
||||
| 4 | „Helyreállítás az élő adatok közé" | `POST /backup/offbox/place` | additive, missing-only **2026-08-30 (R-359): the store is now VERIFIED on a cadence** — a daily `offsite-integrity` job runs `restic check` when the last successful one is over 7 days old. **This is a readability check, not a restore-test (R-87 stays open).** The depth that ships ON now re-reads **100%** of the pack data (R-399 CLOSED, controller v0.228.0) — the sentence here previously said the opposite and was stale. **2026-08-31 (SPIKE R-87):** a scratch restore of every app on demo-hp was measured at **25 s for 8 snapshots / 774 MB**, against 40.3 s for one weekly check — but restic 0.14.0's `--verify` checks size and mtime, **not content**, so nothing available today can vouch for the restored BYTES. Row 4's verdict is UNCHANGED by the spike. `audits/SPIKE-restic-restore-test-2026-08-31.md`. |
|
||||
| 4 | „Helyreállítás az élő adatok közé" | `POST /backup/offbox/place` | additive, missing-only **2026-08-30 (R-359): the store is now VERIFIED on a cadence** — a daily `offsite-integrity` job runs `restic check` when the last successful one is over 7 days old. **This is a readability check, not a restore-test (R-87 stays open).** The depth that ships ON now re-reads **100%** of the pack data (R-399 CLOSED, controller v0.228.0) — the sentence here previously said the opposite and was stale. **2026-08-31 (SPIKE R-87):** a scratch restore of every app on demo-hp was measured at **25 s for 8 snapshots / 774 MB**, against 40.3 s for one weekly check — but restic 0.14.0's `--verify` checks size and mtime, **not content**, so nothing available today can vouch for the restored BYTES. Row 4's verdict is UNCHANGED by the spike. `audits/SPIKE-restic-restore-test-2026-08-31.md`. `audits/SPIKE-restic-restore-test-2026-08-31.md`. |
|
||||
| 5 | „Teljes visszaállítás (fájlok + adatbázis)" | `POST /backup/offbox/reconstitute` | destructive to files + DB, never deleting |
|
||||
| 6 | „Megosztások visszaállítása" + place | `POST /backup/shares/{restore,place}` | additive |
|
||||
| 7 | `.fab` import | `POST` → `apiImportStart` | destructive re-import of one app |
|
||||
@@ -885,7 +885,7 @@ crosses the line — **R-158**.
|
||||
| 7 | **Guest stopped and does not come back** | everything | guest-power watchdog (60 s, `onboot` as the deliberate-stop discriminator) | automatic | **120 s** | | **PROVEN** | agent v0.107.0 replay — 120 s unattended vs the incident's 587 s with a human |
|
||||
| 8 | **Host dies (hardware), drives intact** | the data drives; ep0's PBS namespace; the hub's Recipe + Escrow | **install a new host, then** `--selftest=bring-up -mode dr` per guest, **then** re-attach drives by `durable_id` | **operator** (SSH) | | 7 d | **IMPLEMENTED** | bring-up code exists and has **never been executed** (`CAMPAIGN-8…:522`); the drive half of the plan is empty on every box (**R-105**) |
|
||||
| 9 | **Whole box lost (fire/theft) — host and drives gone** | ep0 PBS namespace; the restic repo; the hub's Recipe + Escrow | new hardware → day-0 → escrow-consume with **R** → restore guests from PBS → app data from Tier-3 | **operator + customer** (R) | | 7 d (guest) · 24 h (app data) | **IMPLEMENTED / UNPROVEN** | every leg exists; the composed path has never been run. The destructive S5 drill is operator-gated and unrun (`06-offsite-connectivity.md:327`) |
|
||||
| 10 | **Ransomware / malicious deletion inside the guest** | PBS offsite (the box **cannot** delete its own snapshots); **the restic repo is NOT protected the same way** | whole-guest restore from PBS to a point before the event | **operator** (SSH) | | 7 d | **PARTIAL** | R-89 proved the box is refused when deleting its own PBS snapshot (`CAMPAIGN-8…:514-515`). **R-95 (open, ranked #1):** the restic credential **can delete** — `readonly=False`, `forget --prune` runs from the box, and SFTP cannot express append-only (`OPEN-ITEMS.md:13`) **2026-08-30 (R-359): the store is now VERIFIED on a cadence** — a daily `offsite-integrity` job runs `restic check` when the last successful one is over 7 days old. **This is a readability check, not a restore-test (R-87 stays open).** The depth that ships ON now re-reads **100%** of the pack data (R-399 CLOSED, controller v0.228.0) — the sentence here previously said the opposite and was stale. **2026-08-31 (SPIKE R-87):** a scratch restore of every app on demo-hp was measured at **25 s for 8 snapshots / 774 MB**, against 40.3 s for one weekly check — but restic 0.14.0's `--verify` checks size and mtime, **not content**, so nothing available today can vouch for the restored BYTES. Row 4's verdict is UNCHANGED by the spike. `audits/SPIKE-restic-restore-test-2026-08-31.md`. |
|
||||
| 10 | **Ransomware / malicious deletion inside the guest** | PBS offsite (the box **cannot** delete its own snapshots); **the restic repo is NOT protected the same way** | whole-guest restore from PBS to a point before the event | **operator** (SSH) | | 7 d | **PARTIAL** | R-89 proved the box is refused when deleting its own PBS snapshot (`CAMPAIGN-8…:514-515`). **R-95 (open, ranked #1):** the restic credential **can delete** — `readonly=False`, `forget --prune` runs from the box, and SFTP cannot express append-only (`OPEN-ITEMS.md:13`) **2026-08-30 (R-359): the store is now VERIFIED on a cadence** — a daily `offsite-integrity` job runs `restic check` when the last successful one is over 7 days old. **This is a readability check, not a restore-test (R-87 stays open).** The depth that ships ON now re-reads **100%** of the pack data (R-399 CLOSED, controller v0.228.0) — the sentence here previously said the opposite and was stale. **2026-08-31 (SPIKE R-87):** a scratch restore of every app on demo-hp was measured at **25 s for 8 snapshots / 774 MB**, against 40.3 s for one weekly check — but restic 0.14.0's `--verify` checks size and mtime, **not content**, so nothing available today can vouch for the restored BYTES. Row 4's verdict is UNCHANGED by the spike. `audits/SPIKE-restic-restore-test-2026-08-31.md`. **2026-09-01 (SPIKE R-95, `audits/SPIKE-r95-offsite-delete-2026-09-01.md`) — THIS ROW'S STATUS IS UNCHANGED, but the seven-day bound beside it is now known to be unverified.** Measured on BOTH boxes over their own SFTP credential, with controls: **no `.snapshots` is visible to either sub-account**, and the account is jailed. Either none exist or a sub-account cannot see them — and a snapshot the box cannot see is one the box cannot restore from, so recovery would be an operator act at the Hetzner panel, not a product capability (R-429; the confirming field needs the provider API, fenced by §11-D). **The PBS shape does NOT transfer**: the sub-account API has one permission axis, `readonly`, and a backup target cannot be read-only — PBS is a server that can refuse, a Storage Box is a filesystem that runs nothing. **What the spike removed as a fear:** withdrawing delete does NOT wedge the store (measured — restic treats a dead owner's lock as stale and proceeds). **What it added:** `unlock --remove-all` reports success while deleting nothing (R-430). |
|
||||
| 11 | **Hub lost** | every box's data plane, every tier, every Lane-1 route | none needed for recovery **of a box**; the hub itself restores from its Longhorn volume backup | operator (`kubectl`) | | 24 h + weekly (Longhorn `RETAIN 1` each) | **UNPROVEN** | a hub restore has never been performed. The backup target is `nfs://192.168.0.180` — **DooPlex itself** — and exactly **2** restore points exist (LIVE, INV Part D2.2) |
|
||||
| 11b | *consequences while the hub is gone* | — | — | — | | | **[FACT]** | **day:** nothing customer-visible breaks; events queue (`settings.go:1466-1485`). **week:** the operator alarm plane is dark, no claim/reset codes, no config or floor convergence, no PBS-secret re-issue. **permanently:** escrow custody and break-glass credentials are gone (INV Part D2.4) |
|
||||
| 12 | **Offsite provider lost (Hetzner)** | everything on-premises: both drives, both whole-guest tiers | none needed — on-premises recovery is unaffected. Re-provision a new offsite target. | operator | | | **[FACT]** | **restic and PBS share the provider** (INV Part E.1). Whether they share an account and payment method is **UNKNOWN** → §11-D |
|
||||
@@ -1085,7 +1085,7 @@ does **not** hold as written. → **R-108**
|
||||
| **R-356** | The off-site restore resolved its destination with the raw `HDD_PATH` and read an empty answer as "not installed" — **CLOSED, controller v0.219.0, 2026-08-22** | 40 of 53 apps were refused permanently while running (§6.3 `[DESIGN]`) |
|
||||
| ~~**R-108**~~ | ~~Network storage can host an app's namespace~~ | **CLOSED 2026-07-30, controller v0.187.0 — D5 UNBLOCKED.** An app namespace may no longer be placed on network storage (5 surfaces guarded by one fail-closed predicate); the share-root bind is deliberately UNCHANGED because it is load-bearing and unscopable (§10.1). `audits/R108-network-app-namespace-2026-07-30.md` |
|
||||
| **R-126** | A `.fab` bundle — plaintext secrets, optional password — can be exported ONTO a NAS: `storageDriveList()` (`internal/web/handler_export.go`) does not filter network paths | split out of R-108, which closed without it. NOT a D5 precondition: an explicit customer-chosen export destination, not a browsing surface reaching a backup tree (§5, §7.3) |
|
||||
| R-95 (open) | The restic offsite credential **can delete** — the box can `forget --prune` its own repo | the tier holding the customer's documents and photos is the one whose credential can destroy it (matrix row 10) |
|
||||
| R-95 (open) | The restic offsite credential **can delete** — the box can `forget --prune` its own repo, from **two** call sites (`offbox.go:1388` retention and `offbox.go:1759` over-quota) | the tier holding the customer's documents and photos is the one whose credential can destroy it (matrix row 10). **SPIKE 2026-09-01:** prevention needs a transport change (restic 0.14.0 does speak `rest:` — measured; append-only is a rest-server flag, not a restic one), because the sub-account API cannot express write-without-delete. **Detection is nearly free and is recommended first** — `snapshot_count` already reaches the hub and the hub appends reports, so the comparison needs no box change. `audits/SPIKE-r95-offsite-delete-2026-09-01.md` |
|
||||
| ~~R-86~~ | ~~Restore-tests are interval-scheduled, not backup-aligned~~ | **CLOSED 2026-08-03 — agent v0.121.0 + hub v0.91.0.** Restore-testing is now **per archive generation**: a tier is due when its newest archive that has settled ~24 h has not been proven, so a daily tier is proved daily on its own archive and a weekly tier weekly on its own. The ticker survives only as the evaluation interval (6 h, chosen from a measured cost). The hub's staleness window moved with it — per tier, from that tier's observed archive rhythm — because a weekly tier proved weekly sat EXACTLY on the old flat 7-day line (§3, Lane 2's per-archive rule) |
|
||||
| ~~R-353~~ | ~~A local unit restore reported a bare completion whether it returned an entire dataset or nothing~~ | **CLOSED 2026-08-30, controller v0.226.0.** The volume count already existed (`restoreDockerVolumesFrom`) and was discarded by a one-line wrapper, so the surface was structurally unable to say what came back. `RestoreFromRecoveryUnit` now returns `UnitRestoreResult` and the sentence has three cases, keyed on replayed-vs-**listed** — zero-replayed has two causes that are opposite news. Every sentence is a claim about the BACKUP, never the app: this path has no `SafetyDump` discriminator, and §6.3 is why that is not pedantry. Proven live on `demo-hp` |
|
||||
| ~~R-357~~ | ~~The destructive reconstitute had no free-space gate; all three that existed guarded non-destructive paths~~ | **CLOSED 2026-08-30, controller v0.226.0.** The gate sits before `writeSafetyDump` and `StopStack`, so a refusal costs no outage — the test asserts the StopStack call count, not the sentence. No headroom multiplier (a measured tree, not a predicted download); fail-closed when either probe reads ≤ 0, which was a real fail-open hole. **Not live-validated** — filling a filesystem is a drill step |
|
||||
@@ -1129,6 +1129,15 @@ and PBS (Cloud server ep0 + a Cloud Volume) are both Hetzner. Whether they share
|
||||
login and payment method** is unverified (INV Unknown U-1) — the question was deliberately not
|
||||
answered by calling the provider API with the production token. Accept explicitly, or mitigate.
|
||||
|
||||
> **The fence was reached again, and held, 2026-09-01 (SPIKE R-95).** The R-95 spike needed one
|
||||
> provider-side fact — **do daily snapshots exist on the Storage Box, and how many** — because the
|
||||
> register calls that mitigation ARMED and the whole ranking of R-95's options turns on it. It is
|
||||
> readable from the box: measured on BOTH machines over their own SFTP credential, with positive
|
||||
> and negative controls, **no `.snapshots` is visible to either sub-account**, and the account is
|
||||
> jailed at `/`. The register's own confirming field, `size_snapshots`, is an API field. **So the
|
||||
> spike stopped here rather than answering it by acting**, and left it as R-429 — ten minutes in
|
||||
> the Storage Box panel, and it re-ranks R-95. `audits/SPIKE-r95-offsite-delete-2026-09-01.md` §Q1.
|
||||
|
||||
**E. `local` vzdump shares a physical device with the guest it backs up.** LIVE on both hosts:
|
||||
`/var/lib/vz` (the archive target) and `local-lvm` (the guest's rootfs and both `backup=1`
|
||||
mountpoints) are both on `/dev/sda3` → VG `pve`. The tier therefore protects against **corruption
|
||||
|
||||
Reference in New Issue
Block a user