From 10c223bdfe2a36528859e1471f79792039bfa805 Mon Sep 17 00:00:00 2001 From: kisfenyo Date: Tue, 1 Sep 2026 16:55:50 +0200 Subject: [PATCH] =?UTF-8?q?DRILL=20R-95:=20the=20recovery=20route=20does?= =?UTF-8?q?=20not=20exist=20=E2=80=94=20stopped=20before=20the=20destructi?= =?UTF-8?q?ve=20phase?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The drill was to delete demo-hp's off-site history and get it back out of a Storage Box snapshot, filling row 10's blank RTO. Phase 1 found there is nothing to get it back from: no snapshot is reachable from a sub-account BY ANY NAME. Measured, read-only, no delete verb issued against any live store: - 777,600 exact names in the vendor form YYYY-MM-DDTHH-MM-SS, nine full days at second granularity, plus 126 alternative shapes -> ZERO hits. - The control is what makes that mean anything: the identical 600-name batch shape with one real path appended returned it, 6 of 6. - Structural cause: /home (u629488-sub3) is st_dev 0,82; /.zfs/snapshot is st_dev 0,276; /home/.zfs does not exist. A snapshot under /.zfs/snapshot belongs to a different dataset than the one holding felhom-repo. - Three tools agree with controls in the same run: SFTP, the port-23 shell, rsync --list-only. So yesterday's re-scope splits: clause (a) "the box cannot write into the snapshot area" STANDS and is re-confirmed; clause (b) "the rest is recoverable file by file" is NOT SUPPORTED. STOPPED before Phase 2 on the operator's ruling — with no recovery leg the deletion would have destroyed real history to buy only an alarm test that could not fire at the specified size. Store verified untouched at 69 snapshots. R-432 ANSWERED (negatively; its panel-read next step withdrawn as unnecessary). R-433 no snapshot reachable by any name — decides R-95's remedy and its rank. R-434 the drop alarm's text promises a file-by-file recovery that cannot be performed. R-435 the drop detector is blind to a single-app deletion (>50% of 69 needed, ~9 given). R-436 LEAD: the provider offers `rclone serve restic --stdio` and restic 0.14.0 speaks `rclone:` (measured, controlled) — real prevention may need no new machine, IF the vendor pins --append-only. Ask before building. 07 §8 row 10: text corrected, status NOT moved, RTO still blank. No code, no version bump, no image, no golden. REPORT.md's only copy of the R-331 report preserved as REPORT-r331-backup-card.md before overwrite. Co-Authored-By: Claude Opus 5 (1M context) Claude-Session: https://claude.ai/code/session_01LB8FmJaGd2cyjvy6dbEjpM --- REPORT-r331-backup-card.md | 187 +++++++++++ REPORT.md | 295 ++++++++---------- .../architecture/07-backup-architecture.md | 4 +- .../README.md | 52 +++ .../01-offsite-inventory.txt | 162 ++++++++++ .../phase0-ground-truth/02-hub-view.txt | 21 ++ .../01-probe-tree.txt | 31 ++ .../02-shell-and-rsync-doors.txt | 21 ++ .../03-shell-capabilities.txt | 54 ++++ .../04-enumeration-attempts.txt | 41 +++ .../05-home-zfs-door.txt | 16 + .../06-name-sweep.txt | 165 ++++++++++ .../07-sweep-control.txt | 3 + .../08-alt-formats.txt | 4 + .../09-rclone-restic-lead.txt | 7 + .../phase2-the-attack/NOT-RUN.md | 13 + .../phase3-the-alarm/NOT-RUN.md | 13 + .../phase4-the-recovery/NOT-RUN.md | 13 + .../phase5-rearm/NOT-RUN.md | 13 + .../phase6-teardown/01-verify-then-clean.txt | 23 ++ .../02-guest-and-host-clean.txt | 22 ++ .../03-dooplex-clean-and-final-state.txt | 12 + documentation/backlog/OPEN-ITEMS.md | 8 +- 23 files changed, 1018 insertions(+), 162 deletions(-) create mode 100644 REPORT-r331-backup-card.md create mode 100644 documentation/audits/evidence-drill-r95-recovery-2026-09-01/README.md create mode 100644 documentation/audits/evidence-drill-r95-recovery-2026-09-01/phase0-ground-truth/01-offsite-inventory.txt create mode 100644 documentation/audits/evidence-drill-r95-recovery-2026-09-01/phase0-ground-truth/02-hub-view.txt create mode 100644 documentation/audits/evidence-drill-r95-recovery-2026-09-01/phase1-find-the-snapshot/01-probe-tree.txt create mode 100644 documentation/audits/evidence-drill-r95-recovery-2026-09-01/phase1-find-the-snapshot/02-shell-and-rsync-doors.txt create mode 100644 documentation/audits/evidence-drill-r95-recovery-2026-09-01/phase1-find-the-snapshot/03-shell-capabilities.txt create mode 100644 documentation/audits/evidence-drill-r95-recovery-2026-09-01/phase1-find-the-snapshot/04-enumeration-attempts.txt create mode 100644 documentation/audits/evidence-drill-r95-recovery-2026-09-01/phase1-find-the-snapshot/05-home-zfs-door.txt create mode 100644 documentation/audits/evidence-drill-r95-recovery-2026-09-01/phase1-find-the-snapshot/06-name-sweep.txt create mode 100644 documentation/audits/evidence-drill-r95-recovery-2026-09-01/phase1-find-the-snapshot/07-sweep-control.txt create mode 100644 documentation/audits/evidence-drill-r95-recovery-2026-09-01/phase1-find-the-snapshot/08-alt-formats.txt create mode 100644 documentation/audits/evidence-drill-r95-recovery-2026-09-01/phase1-find-the-snapshot/09-rclone-restic-lead.txt create mode 100644 documentation/audits/evidence-drill-r95-recovery-2026-09-01/phase2-the-attack/NOT-RUN.md create mode 100644 documentation/audits/evidence-drill-r95-recovery-2026-09-01/phase3-the-alarm/NOT-RUN.md create mode 100644 documentation/audits/evidence-drill-r95-recovery-2026-09-01/phase4-the-recovery/NOT-RUN.md create mode 100644 documentation/audits/evidence-drill-r95-recovery-2026-09-01/phase5-rearm/NOT-RUN.md create mode 100644 documentation/audits/evidence-drill-r95-recovery-2026-09-01/phase6-teardown/01-verify-then-clean.txt create mode 100644 documentation/audits/evidence-drill-r95-recovery-2026-09-01/phase6-teardown/02-guest-and-host-clean.txt create mode 100644 documentation/audits/evidence-drill-r95-recovery-2026-09-01/phase6-teardown/03-dooplex-clean-and-final-state.txt diff --git a/REPORT-r331-backup-card.md b/REPORT-r331-backup-card.md new file mode 100644 index 00000000..ad0d01bb --- /dev/null +++ b/REPORT-r331-backup-card.md @@ -0,0 +1,187 @@ +# REPORT — R-331: the operator Backup card said every customer had no backups + +**Hub v0.109.0 (with controller v0.225.0) · 2026-08-30** + +--- + +## 1. What was wrong + +The hub customer page's **Backup** card read, for **every customer, indefinitely**: + +``` +Enabled Yes Snapshots 0 +Repo Size 0 MB Integrity Unknown +``` + +Measured on `demo-hp` 2026-08-30, at which moment the truth was: + +| source | value | +|---|---| +| the box's own `settings.json` | `snapshot_count: 67, repo_size_bytes: 140829678, stats_known: true` | +| that night's controller log | `[offbox] backup OK: 8 app(s) backed up, 67 snapshot(s), 2m14s` | +| **this hub's own Offsite page** | `0.1 GB` used of a `50 GB` quota — read from the same stored report | + +**A card that reads "no backups" over a working backup is worse than no card.** It is the R-88 +direction of failure — degrading to *no backup* rather than to *unknown* — on the one screen an +operator consults to answer "is this customer protected?". + +## 2. Root cause + +The card rendered the report's **`backup`** object. Its `snapshot_count`, `repo_size_mb` and +`integrity_ok` fields have had **no producer** since disk-tier restic moved to the host agent (slice +8C) — the controller's `buildBackupReport` leaves them zero *deliberately* and says so in a comment. +The zeros were correct values for dead fields, rendered as if live. + +**The data was never missing.** The live numbers ride in the report's **`offsite`** object, which this +package **already** reads for the Offsite page (`offsiteUsageBytes`) and which `monitor.OffsiteChecker` +**already** drives fill and staleness alarms from. That the Offsite page rendered demo-hp's real usage +from the same stored report, at the same moment the Backup card said `0 MB`, is the proof the bytes +were arriving. This is a **render fix over an existing feed**, not a new pipeline. + +## 3. Why it was not a one-line template swap + +`snapshot_count: 0` means two opposite things — *this repository holds nothing* and *nobody has ever +measured this repository*. **R-225 measured that confusion one layer down**: a rebuilt box rendered +„Tarolo meret · 0 pillanatkep" over a store that really held snapshot `f3d9cd67`, and the controller's +`StatsKnown` fixed it there. It was never on the wire, so rendering the count without it would have +**moved R-225 up to the hub instead of fixing anything**. Controller v0.225.0 now forwards +`stats_known`. + +## 4. What changed + +`hub/internal/web/backup_card.go` builds a typed `backupCardView` — resolved in Go, because the card's +whole subject is a distinction a template `{{if}}` chain over `map[string]interface{}` float64s cannot +keep: + +| report state | card shows | +|---|---| +| no `offsite` object at all | "No off-site data reported" — **and says explicitly this is not the same as "no backups"** | +| `enabled:false` + declared `state` | the blocker by name (`needs_credential`) — a different operator action from "not enabled" | +| enabled, `stats_known:false` | **—**, plus "never been measured". Never `0` | +| enabled, `stats_known:true` | the real count and size, **including a real `0`** — measured empty is knowledge | + +**A pre-v0.225.0 controller sends no `stats_known`, which unmarshals to false → "unknown".** That is +the fail-safe direction: upgrading the hub ahead of the fleet must not tell the operator that every +un-upgraded customer has zero backups. Pinned by a test. + +**The Integrity row is deleted, not re-sourced.** Nothing produces it: the controller runs no integrity +check, and `NotifyIntegrityOK` / `NotifyIntegrityFailed` exist and are **called from nowhere**. A row +that can only ever read "Unknown" is not information, and one that could read "OK" from an unwritten +field would be a lie. + +`fmtBytesAuto` is new rather than reusing `fmtBytesGB`: that one is fixed at GB because it renders +against GB quotas, and it turns demo-hp's real 140 829 678 bytes into `0.1 GB` — which on a card whose +entire defect was under-reporting a real backup reads as "nearly nothing". + +## 5. Tests and the red-proof + +`r331_backup_card_test.go` asserts the **rendered page**, using demo-hp's real reported values, so a +regression fails against the same numbers the defect was measured against. **The defect lived in the +template's choice of source object, so a test one layer below it would have been green against the +shipped bug** — which is why these drive `handleCustomerUnified` and grep the HTML. + +**RED-PROOF (run 2026-08-30):** restoring the pre-fix card markup fails all four tests — +`the rendered Backup card does not contain demo-hp's real snapshot count (67)`, +`the card does not carry the real repository size (134.3 MB ...)`, +`the card still shows an Integrity row`, plus every branch of the three-way ruling. Restored +immediately; `git diff` clean. + +**Green gate:** `go build ./... && go vet ./... && go test ./...` in `hub/` — 18 packages, rc 0. + +## 6. Deployment and live verification + +Hub **0.109.0** built, pushed, manifest bumped, ArgoCD hard-refreshed and synced. Controller +**0.225.0** deployed to both demo boxes. Verified from the live objects, not from a rollout message +(an ArgoCD "rolled out" can name the old image): + +``` +argocd: sync=Synced rev=36f86300209e9ce914f2a97dac16ee9eeea249e2 (== HEAD) +deploy: gitea.dooplex.hu/admin/felhom-hub:0.109.0 +pod: hub-6795879c4b-pf4lz ...felhom-hub:0.109.0 Running +boxes: ...felhom-controller:0.225.0 Up (healthy) on demo-felhom AND demo-hp +``` + +**The card, fetched from the live hub** (endpoint-level: the exact URL the operator's browser +requests; the residual is client-side rendering only — there is no browser on DooPlex): + +| | demo-hp | demo-felhom | +|---|---|---| +| Off-site snapshots | **67** | **10** | +| Repo size | **134.3 MB** | **132.5 KB** | +| Last successful run | 15h ago | 15h ago | +| Soft quota | 50 GB | 50 GB | +| Integrity row | **absent** (grep count 0) | **absent** (grep count 0) | + +Both read `Snapshots 0 · Repo Size 0 MB · Integrity Unknown` before this change. + +**Cross-checked against the source, not just against itself** — the numbers on the card are the +numbers on the boxes: + +``` +demo-hp settings.json offbox: snapshot_count 67, repo_size_bytes 140829678, stats_known true +demo-felhom settings.json offbox: snapshot_count 10, repo_size_bytes 135635, stats_known true + 135635 / 1024 = 132.5 KB → matches the rendered value +``` + +**One honest detail worth keeping:** demo-felhom's "Last DB dump" reads `—`. That is correct, not a +regression — its only app (`opengist`) has no database, so the box has never taken a DB dump. + +**Not verified live: the "never measured" branch.** Both boxes report `stats_known: true`, so the +degradation path could not be exercised on real hardware without falsifying a box's state. It is +covered by `TestBackupCard_ThreeWayRuling` and `TestBackupCard_OldControllerDegradesToUnknownNotEmpty` +at render level, and this is stated rather than implied. + +## 7. The push bypassed a gate, deliberately, and here is the declaration + +**`git push --no-verify` was used for this change.** `repo_gates.py`'s `golden-currency` gate was +CONVICTED and it was RIGHT: controller **v0.224.0** and **v0.225.0** are released and the newest golden +bake carries **0.223.0**, so a machine installed right now receives neither fix. + +**This is a BYPASS, not a waiver.** The gate offers a waiver only for a release that *deliberately needs +no golden*; these need one. **The operator was asked and ruled bypass-now-bake-later**, on the stated +ground that neither fix bites a day-0 box — R-330 is a nightly false alarm about apps a new box has not +installed yet, R-331 is a hub-side display over backups a new box has not taken yet — and both arrive by +self-update afterwards. That ground is recorded on R-242 precisely because it is the thing to re-check: +**it does not extend to a release that changes first-boot behaviour.** + +**A golden carrying 0.225.0 is OWED** (`RUNBOOK-manual-build.md` §4.1; the vouch is a three-field change, +`MinAgent 0.129.0`). This is the **fourth** bypass of this gate, and the gap it names is now two releases +wide rather than one. + +**The other failing gate was fixed, not bypassed.** `due-checks` was red on R-341's `+7 d` measurement, +five days overdue. It was **taken** during this session — see §8. + +## 8. R-341's overdue check was taken, and its premise did not survive + +Unrelated to R-331; it blocked the same push, so it was done rather than deferred. Evidence: +`documentation/audits/evidence-r341-plus7d-2026-08-30/step1-fd-and-sockets.txt`. + +**Precondition passed**, which is what makes the reading interpretable: `proxmox-backup-proxy` still +`MainPID 551655`, `ps -o lstart=` still `2026-08-18 09:51:04`, `NRestarts=0` — the same proxy generation +as t0, so nothing restarted and re-based the count. (The anchor is `ps`, not `ActiveEnterTimestamp`, +which reads 03:54:54Z here — R-346's trap, avoided.) + +**Result: fd = 17.** Not 17 more — seventeen total, exactly the documented baseline, against **405** at +the first check on 2026-08-20. The socket histogram holds **one LISTEN and nothing else**: ESTAB 0, +CLOSE-WAIT 0. + +**The verdict is "unanswerable", not "the upgrade fixed it".** R-341 asks whether the PBS 4.2.5-1 +upgrade changed the fd slope. Inside this interval **we removed the leak ourselves** (R-344, agent +0.130.0, now live on both boxes). A slope of ~0 measures our fix, not the upgrade — and reading it the +other way would credit a changelog that was read in advance and found to contain no such mechanism. The +perturbation pre-registered for this window was Phase C at ~3%; the actual perturbation was the removal +of the entire phenomenon. The question is now **moot**, and the row is closed as such. + +**What it does establish, which is worth more than the original question:** twelve days after the R-344 +fix, on the same proxy generation with no restart to hide behind, ep0 sits at baseline with zero +established connections. The 388-descriptor accumulation has not returned, and R-336's ~323-day runway +concern retires with it. + +## 9. Not done, and why + +- **No staleness verdict on the card.** `monitor.OffsiteChecker` already owns that and alarms on it. A + second verdict over the same data is two things that can disagree — a shape this codebase has already + paid for (`LastRun` vs `LastSuccess`, R-100). +- **The dead `backup` fields were not removed from the controller's wire format.** Removing them would + stop historical reports already in this hub's store from parsing, for no gain — nothing renders them + now, and a controller-side test fails if anything starts producing them. diff --git a/REPORT.md b/REPORT.md index ad0d01bb..a9755c33 100644 --- a/REPORT.md +++ b/REPORT.md @@ -1,187 +1,166 @@ -# REPORT — R-331: the operator Backup card said every customer had no backups +# REPORT — the deletion we said is survivable: the recovery route does not exist (R-95 drill, 2026-09-01) -**Hub v0.109.0 (with controller v0.225.0) · 2026-08-30** +**RUNBOOK, destructive class, `demo-hp` only. STOPPED at the end of Phase 1 on the operator's ruling, +before any destructive step. No delete verb was issued against any live store; no byte on either +Storage Box sub-account was written, moved or removed.** No production code, no version bump, no +image, no golden. Evidence: `documentation/audits/evidence-drill-r95-recovery-2026-09-01/`. + +| # | phase | verdict | one sentence | +|---|---|---|---| +| 1 | snapshot reachable, and its name | **NO — and it has no reachable name** | 777,600 exact names across nine days in the vendor's own format, plus 126 alternative shapes; zero resolve, with a control proving the sweep detects a path that exists. | +| 2 | the deletion | **NOT RUN — operator ruling** | With no recovery route, the deletion would have destroyed real history to buy nothing; put as a two-option decision, the ruling was stop. | +| 3 | the alarm fired | **NOT RUN — and it could not have fired at the specified size** | The shipped threshold needs a fall of more than half; one app's tag is ~9 of 69. → **R-435** | +| 4 | **the recovery** | **NOT RUN — no route exists that is not fenced** | Box-side: proven impossible. Panel: fenced and browserless. Hetzner API: fenced (§11-D), and the hub's client has no snapshot method at all. | +| — | **RTO from T₀** | **STILL BLANK** | Row 10's RTO cell is unchanged and remains a finding. | +| — | **data lost, quantified** | **NOT MEASURABLE THIS WAY** | The quantity only has meaning if the rest is recoverable, and the route that would recover it is not reachable. | +| 5 | re-arm | **NOT RUN** | Depended on Phase 4. | +| 6 | teardown | **PASS** | Store untouched at 69 snapshots; all three scratch layers removed; hub DB copies shredded; both boxes healthy. | --- -## 1. What was wrong +## 1. Did the recovery work — and does yesterday's re-scope survive? -The hub customer page's **Backup** card read, for **every customer, indefinitely**: +**The recovery was never reachable, and the re-scope does not survive intact. Its first half stands; +its second half does not.** -``` -Enabled Yes Snapshots 0 -Repo Size 0 MB Integrity Unknown -``` +Yesterday's re-scope has two clauses. They must now be separated: -Measured on `demo-hp` 2026-08-30, at which moment the truth was: +* **(a) "The box can delete its live repository, but cannot write to the daily snapshots of it."** + **STANDS.** Re-confirmed here: `/.zfs/snapshot` is reachable and the write-refusal measurement is + unchanged. Nothing in this drill weakens it. +* **(b) "…so the rest is recoverable — file by file, one customer at a time."** **NOT SUPPORTED.** + A snapshot that cannot be opened cannot be copied out of. R-432 recorded the directory listing + empty and named the cheapest next step: *"a single `ls /.zfs/snapshot/` from a box then + settles whether a named snapshot can be entered even though the directory does not list (ZFS + allows exactly that)."* **That step is now done, exhaustively, and the answer is no.** -| source | value | -|---|---| -| the box's own `settings.json` | `snapshot_count: 67, repo_size_bytes: 140829678, stats_known: true` | -| that night's controller log | `[offbox] backup OK: 8 app(s) backed up, 67 snapshot(s), 2m14s` | -| **this hub's own Offsite page** | `0.1 GB` used of a `50 GB` quota — read from the same stored report | +**What was measured.** The port-23 restricted shell accepts a batched `stat`, which makes a cheap +existence oracle: 500–600 paths per round trip, stdout carrying only paths that exist. -**A card that reads "no backups" over a working backup is worse than no card.** It is the R-88 -direction of failure — degrading to *no backup* rather than to *unknown* — on the one screen an -operator consults to answer "is this customer protected?". - -## 2. Root cause - -The card rendered the report's **`backup`** object. Its `snapshot_count`, `repo_size_mb` and -`integrity_ok` fields have had **no producer** since disk-tier restic moved to the host agent (slice -8C) — the controller's `buildBackupReport` leaves them zero *deliberately* and says so in a comment. -The zeros were correct values for dead fields, rendered as if live. - -**The data was never missing.** The live numbers ride in the report's **`offsite`** object, which this -package **already** reads for the Offsite page (`offsiteUsageBytes`) and which `monitor.OffsiteChecker` -**already** drives fill and staleness alarms from. That the Offsite page rendered demo-hp's real usage -from the same stored report, at the same moment the Backup card said `0 MB`, is the proof the bytes -were arriving. This is a **render fix over an existing feed**, not a new pipeline. - -## 3. Why it was not a one-line template swap - -`snapshot_count: 0` means two opposite things — *this repository holds nothing* and *nobody has ever -measured this repository*. **R-225 measured that confusion one layer down**: a rebuilt box rendered -„Tarolo meret · 0 pillanatkep" over a store that really held snapshot `f3d9cd67`, and the controller's -`StatsKnown` fixed it there. It was never on the wire, so rendering the count without it would have -**moved R-225 up to the hub instead of fixing anything**. Controller v0.225.0 now forwards -`stats_known`. - -## 4. What changed - -`hub/internal/web/backup_card.go` builds a typed `backupCardView` — resolved in Go, because the card's -whole subject is a distinction a template `{{if}}` chain over `map[string]interface{}` float64s cannot -keep: - -| report state | card shows | -|---|---| -| no `offsite` object at all | "No off-site data reported" — **and says explicitly this is not the same as "no backups"** | -| `enabled:false` + declared `state` | the blocker by name (`needs_credential`) — a different operator action from "not enabled" | -| enabled, `stats_known:false` | **—**, plus "never been measured". Never `0` | -| enabled, `stats_known:true` | the real count and size, **including a real `0`** — measured empty is knowledge | - -**A pre-v0.225.0 controller sends no `stats_known`, which unmarshals to false → "unknown".** That is -the fail-safe direction: upgrading the hub ahead of the fleet must not tell the operator that every -un-upgraded customer has zero backups. Pinned by a test. - -**The Integrity row is deleted, not re-sourced.** Nothing produces it: the controller runs no integrity -check, and `NotifyIntegrityOK` / `NotifyIntegrityFailed` exist and are **called from nowhere**. A row -that can only ever read "Unknown" is not information, and one that could read "OK" from an unwritten -field would be a lie. - -`fmtBytesAuto` is new rather than reusing `fmtBytesGB`: that one is fixed at GB because it renders -against GB quotas, and it turns demo-hp's real 140 829 678 bytes into `0.1 GB` — which on a card whose -entire defect was under-reporting a real backup reads as "nearly nothing". - -## 5. Tests and the red-proof - -`r331_backup_card_test.go` asserts the **rendered page**, using demo-hp's real reported values, so a -regression fails against the same numbers the defect was measured against. **The defect lived in the -template's choice of source object, so a test one layer below it would have been green against the -shipped bug** — which is why these drive `handleCustomerUnified` and grep the HTML. - -**RED-PROOF (run 2026-08-30):** restoring the pre-fix card markup fails all four tests — -`the rendered Backup card does not contain demo-hp's real snapshot count (67)`, -`the card does not carry the real repository size (134.3 MB ...)`, -`the card still shows an Integrity row`, plus every branch of the three-way ruling. Restored -immediately; `git diff` clean. - -**Green gate:** `go build ./... && go vet ./... && go test ./...` in `hub/` — 18 packages, rc 0. - -## 6. Deployment and live verification - -Hub **0.109.0** built, pushed, manifest bumped, ArgoCD hard-refreshed and synced. Controller -**0.225.0** deployed to both demo boxes. Verified from the live objects, not from a rollout message -(an ArgoCD "rolled out" can name the old image): - -``` -argocd: sync=Synced rev=36f86300209e9ce914f2a97dac16ee9eeea249e2 (== HEAD) -deploy: gitea.dooplex.hu/admin/felhom-hub:0.109.0 -pod: hub-6795879c4b-pf4lz ...felhom-hub:0.109.0 Running -boxes: ...felhom-controller:0.225.0 Up (healthy) on demo-felhom AND demo-hp -``` - -**The card, fetched from the live hub** (endpoint-level: the exact URL the operator's browser -requests; the residual is client-side rendering only — there is no browser on DooPlex): - -| | demo-hp | demo-felhom | +| sweep | candidates | hits | |---|---|---| -| Off-site snapshots | **67** | **10** | -| Repo size | **134.3 MB** | **132.5 KB** | -| Last successful run | 15h ago | 15h ago | -| Soft quota | 50 GB | 50 GB | -| Integrity row | **absent** (grep count 0) | **absent** (grep count 0) | +| `/.zfs/snapshot/YYYY-MM-DDTHH-MM-SS`, nine full days, second granularity | **777,600** | **0** | +| 126 alternative name shapes and snapshot paths (`daily`, `snapshot-1`, colon and compact time forms, `/home/.snapshot`, …) | 126 | 0 | +| **control — the identical 600-name batch shape with one real path appended** | 6 batches | **6/6 returned it** | -Both read `Snapshots 0 · Repo Size 0 MB · Integrity Unknown` before this change. - -**Cross-checked against the source, not just against itself** — the numbers on the card are the -numbers on the boxes: +**And there is a structural reason, which is why I stopped sweeping.** The customer's data and the +snapshot door are on **different filesystems**: ``` -demo-hp settings.json offbox: snapshot_count 67, repo_size_bytes 140829678, stats_known true -demo-felhom settings.json offbox: snapshot_count 10, repo_size_bytes 135635, stats_known true - 135635 / 1024 = 132.5 KB → matches the rendered value +df → u629488-sub3 mounted on /home +stat /home → Device 0,82 +stat /.zfs/snapshot → Device 0,276 ← a different device +stat /home/.zfs → cannot statx: No such file or directory ``` -**One honest detail worth keeping:** demo-felhom's "Last DB dump" reads `—`. That is correct, not a -regression — its only app (`opengist`) has no database, so the box has never taken a DB dump. +A ZFS snapshot under `/.zfs/snapshot` belongs to the dataset that owns that `.zfs` — not to the +child mounted at `/home`. **So even a correctly named snapshot there could not contain +`felhom-repo`,** and the dataset that does hold it exposes no `.zfs` at all to this account. The +empty listing is not a display toggle hiding a reachable tree; from a sub-account there is no tree. -**Not verified live: the "never measured" branch.** Both boxes report `stats_known: true`, so the -degradation path could not be exercised on real hardware without falsifying a box's state. It is -covered by `TestBackupCard_ThreeWayRuling` and `TestBackupCard_OldControllerDegradesToUnknownNotEmpty` -at render level, and this is stated rather than implied. +**Three tools agree, each with controls in the same run:** SFTP, the port-23 shell, and +`rsync --list-only`. -## 7. The push bypassed a gate, deliberately, and here is the declaration +**What that does to R-95.** Its *exposure* is unchanged and its *remedy* is not. Yesterday the row +could say a deletion costs about a day because the rest comes back per-file. Today the only routes +to "the rest" are a whole-box panel rollback (which deletes newer snapshots and hits every customer +on the box) and the provider API (fenced, and unimplemented in the hub's client). **The re-scope's +comfort was resting on a route nobody had walked — which is precisely the standard this project +applies, and it is the reason this drill was called.** -**`git push --no-verify` was used for this change.** `repo_gates.py`'s `golden-currency` gate was -CONVICTED and it was RIGHT: controller **v0.224.0** and **v0.225.0** are released and the newest golden -bake carries **0.223.0**, so a machine installed right now receives neither fix. +**The ranking is Viktor's and I am not re-ranking it.** What I will say plainly: the argument that +moved R-95 down yesterday is the argument this drill removed. On these facts I would put it back +where it was. -**This is a BYPASS, not a waiver.** The gate offers a waiver only for a release that *deliberately needs -no golden*; these need one. **The operator was asked and ruled bypass-now-bake-later**, on the stated -ground that neither fix bites a day-0 box — R-330 is a nightly false alarm about apps a new box has not -installed yet, R-331 is a hub-side display over backups a new box has not taken yet — and both arrive by -self-update afterwards. That ground is recorded on R-242 precisely because it is the thing to re-check: -**it does not extend to a release that changes first-boot behaviour.** +## 2. The RTO -**A golden carrying 0.225.0 is OWED** (`RUNBOOK-manual-build.md` §4.1; the vouch is a three-field change, -`MinAgent 0.129.0`). This is the **fourth** bypass of this gate, and the gap it names is now two releases -wide rather than one. +**Still blank, and it stays a finding.** `07` §8 row 10's RTO cell has been empty since July and this +drill did not fill it. Nothing was recovered, so nothing was timed. The summary line at +`07-backup-architecture.md:948` — *"no ransomware-shaped recovery has ever been run"* — is still +true, and is now true for a sharper reason: **not "nobody has run it" but "from the box, it cannot +be run."** -**The other failing gate was fixed, not bypassed.** `due-checks` was red on R-341's `+7 d` measurement, -five days overdue. It was **taken** during this session — see §8. +## 3. R-432's answer, and the naming scheme -## 8. R-341's overdue check was taken, and its premise did not survive +**R-432 is ANSWERED, and negatively. It did not need the panel read it was waiting on.** -Unrelated to R-331; it blocked the same push, so it was done rather than deferred. Evidence: -`documentation/audits/evidence-r341-plus7d-2026-08-30/step1-fd-and-sockets.txt`. +* **The naming scheme is `YYYY-MM-DDTHH-MM-SS`** — vendor-documented examples `2025-12-03T13-47-47`, + `2025-02-12T11-35-19`. Recorded so nobody hunts a console again. +* **Knowing it does not help.** Every name in that format for nine days is refused, and the st_dev + split above says why. **Per-file recovery is not operator-only — from the box it is nobody's,** and + for the operator it is a browser act against the main account that no credential in this project + can perform. +* **The panel cannot supply the missing piece either.** It offers Restore and Delete on a row and + does not show names; and the one name-shaped thing it could give would be tried against a door + that leads to the wrong dataset. -**Precondition passed**, which is what makes the reading interpretable: `proxmox-backup-proxy` still -`MainPID 551655`, `ps -o lstart=` still `2026-08-18 09:51:04`, `NRestarts=0` — the same proxy generation -as t0, so nothing restarted and re-based the count. (The anchor is `ps`, not `ActiveEnterTimestamp`, -which reads 03:54:54Z here — R-346's trap, avoided.) +## 4. The alarm's first real firing -**Result: fd = 17.** Not 17 more — seventeen total, exactly the documented baseline, against **405** at -the first check on 2026-08-20. The socket histogram holds **one LISTEN and nothing else**: ESTAB 0, -CLOSE-WAIT 0. +**It did not happen, and the drill as written could not have produced it.** The detector fires on a +fall of **more than half** the previous count **and at least 5** (`hub/internal/monitor/offsite.go`, +`snapshotDropFraction = 0.5`, `snapshotDropFloor = 5`). demo-hp's baseline is **69**. Phase 2 deletes +**one app's** history — about **9** snapshots. 9 is over the floor and nowhere near half, so the +alarm stays silent, **correctly and by design**. Firing it for real needs ~35+ snapshots destroyed, +i.e. most of demo-hp's off-site history. That trade is what the operator was asked to rule on. → **R-435** -**The verdict is "unanswerable", not "the upgrade fixed it".** R-341 asks whether the PBS 4.2.5-1 -upgrade changed the fd slope. Inside this interval **we removed the leak ourselves** (R-344, agent -0.130.0, now live on both boxes). A slope of ~0 measures our fix, not the upgrade — and reading it the -other way would credit a changelog that was read in advance and found to contain no such mechanism. The -perturbation pre-registered for this window was Phase C at ~3%; the actual perturbation was the removal -of the entire phenomenon. The question is now **moot**, and the row is closed as such. +**One thing the alarm says is now wrong.** Its message, live in hub 0.111.0, reads: -**What it does establish, which is worth more than the original question:** twelve days after the R-344 -fix, on the same proxy generation with no restart to hide behind, ep0 sits at baseline with zero -established connections. The 388-descriptor accumulation has not returned, and R-336's ~323-day runway -concern retires with it. +> "The daily Storage Box snapshots are read-only and still hold the older copy, **so this is +> recoverable file-by-file**; it is NOT confirmed data loss." -## 9. Not done, and why +The first clause is true; **the second promises a recovery the product cannot perform and the +operator cannot perform without a browser and the main account.** This is this project's own +corollary — *when a verdict changes which field it counts from, the alarm text has to change with +it* — landing on the alarm shipped the same day. → **R-434** -- **No staleness verdict on the card.** `monitor.OffsiteChecker` already owns that and alarms on it. A - second verdict over the same data is two things that can disagree — a shape this codebase has already - paid for (`LastRun` vs `LastSuccess`, R-100). -- **The dead `backup` fields were not removed from the controller's wire format.** Removing them would - stop historical reports already in this hub's store from parsing, for no gain — nothing renders them - now, and a controller-side test fails if anything starts producing them. +## 5. Findings, as register rows + +All four filed in `documentation/backlog/OPEN-ITEMS.md`. + +| row | finding | +|---|---| +| **R-433** | A sub-account cannot reach any Storage Box snapshot **by any name**; `/home` and `/.zfs` are different filesystems and `/home/.zfs` does not exist. Answers R-432 negatively and removes clause (b) of the R-95 re-scope. | +| **R-434** | `emitSnapshotDrop`'s message promises file-by-file recovery that is not reachable. Live in hub 0.111.0. | +| **R-435** | The drop detector cannot see a single-app deletion (>50% of 69 ⇒ ~35 needed). `offbox.go:1388` forgets **by tag**, so a single-tag wipe is exactly the shape the detector is blind to. Deliberate insensitivity, but the blind spot should be stated where the operator reads it. | +| **R-436** | **LEAD, not a defect.** Hetzner's port-23 shell offers `rclone serve restic --stdio` as a server-side backend, and restic 0.14.0 recognises the `rclone:` backend (measured; control `banana:` → invalid backend; rclone is absent from the controller image). `rclone serve restic` carries `--append-only`. **This could make R-95's real prevention far cheaper than the spike concluded — no new always-on machine, no data migration.** Caveat stated up front: the **client** supplies the server command line, so a compromised box could omit the flag unless the provider pins it. Settling that is a vendor question, not a code change. | + +**R-432 is marked ANSWERED**; its "one panel read settles it" next step is withdrawn as unnecessary. + +## 6. Does `07` §8 row 10 move? + +**No. It stays `PARTIAL`, and its RTO stays blank.** The status was already correct for the right +reason — *"the recovery ROUTE has never been walked, which is what PARTIAL means"* — and this drill +found the route is not walkable from the box at all. **What the row needs is a text correction, not a +status change:** its clause *"recoverable per-file (vendor)"* and its limit *"per-file recovery is +operator-only today (R-432)"* both overstate what exists. Updated in place with the citation. Moving +it only as far as the evidence goes means not moving it. + +## 7. What could not be tested, and why + +* **Whether the main account can see the snapshots.** No main-account credential exists in this + project — the hub holds only per-customer sub-accounts. This is the one question that would decide + whether per-file recovery exists *at all*, for anyone. +* **Whether the Hetzner API can list or read a snapshot.** Fenced by the runbook (§11-D). Separately, + `hub/internal/hetznerapi/hetznerapi.go` has **no snapshot method** — so this route needs new code + regardless of the fence. +* **The deletion, the alarm's first real firing, the recovery, the RTO, the re-arm.** Phases 2–5, not + run, on the operator's ruling. +* **Whether `rclone serve restic --stdio` is pinned server-side with `--append-only`** (R-436). + +## 8. My own mistakes + +* **I ran Phase 1 before finishing Phase 0's subject-app choice, and did not say so up front.** The + choice depends on the snapshot's timestamp, so the order was right, but the runbook's order is the + runbook's and a silent reordering is the thing this project keeps getting caught by. Stated at the + time in the session, recorded here. +* **My first sweep guessed the schedule instead of establishing it.** I probed 00:00 UTC and 22:00 + UTC — 600 names — on the strength of a register line reading *"daily 00:00"*, got nothing, and only + then widened to whole days. The narrow sweep was worth nothing on its own: a zero over a guessed + window is not evidence, and I should have gone to full days first or not run it at all. +* **I nearly reported the empty listing as "the display toggle is hiding it".** The vendor documents + exactly such a toggle and it fitted. The st_dev comparison — which I only ran because `df` printed + a filesystem name I did not expect — says the tree is on another dataset entirely. **A plausible + cause that fits the symptom is not a measured one**, and I had the wrong one for about ten minutes. +* **`REPORT.md` held the only copy of the R-331 report** (hub v0.109.0, 2026-08-30) — durable content + living only in the overwritten file, which `CLAUDE.md:82-87` forbids. Preserved as + `REPORT-r331-backup-card.md` before this report replaced it. diff --git a/documentation/architecture/07-backup-architecture.md b/documentation/architecture/07-backup-architecture.md index 7a85d3e8..727b003a 100644 --- a/documentation/architecture/07-backup-architecture.md +++ b/documentation/architecture/07-backup-architecture.md @@ -885,7 +885,7 @@ crosses the line — **R-158**. | 7 | **Guest stopped and does not come back** | everything | guest-power watchdog (60 s, `onboot` as the deliberate-stop discriminator) | automatic | **120 s** | | **PROVEN** | agent v0.107.0 replay — 120 s unattended vs the incident's 587 s with a human | | 8 | **Host dies (hardware), drives intact** | the data drives; ep0's PBS namespace; the hub's Recipe + Escrow | **install a new host, then** `--selftest=bring-up -mode dr` per guest, **then** re-attach drives by `durable_id` | **operator** (SSH) | | 7 d | **IMPLEMENTED** | bring-up code exists and has **never been executed** (`CAMPAIGN-8…:522`); the drive half of the plan is empty on every box (**R-105**) | | 9 | **Whole box lost (fire/theft) — host and drives gone** | ep0 PBS namespace; the restic repo; the hub's Recipe + Escrow | new hardware → day-0 → escrow-consume with **R** → restore guests from PBS → app data from Tier-3 | **operator + customer** (R) | | 7 d (guest) · 24 h (app data) | **IMPLEMENTED / UNPROVEN** | every leg exists; the composed path has never been run. The destructive S5 drill is operator-gated and unrun (`06-offsite-connectivity.md:327`) | -| 10 | **Ransomware / malicious deletion inside the guest** | PBS offsite (the box **cannot** delete its own snapshots); **the restic repo is NOT protected the same way** | whole-guest restore from PBS to a point before the event | **operator** (SSH) | | 7 d | **PARTIAL** | R-89 proved the box is refused when deleting its own PBS snapshot (`CAMPAIGN-8…:514-515`). **R-95 (open, ranked #1):** the restic credential **can delete** — `readonly=False`, `forget --prune` runs from the box, and SFTP cannot express append-only (`OPEN-ITEMS.md:13`) **2026-08-30 (R-359): the store is now VERIFIED on a cadence** — a daily `offsite-integrity` job runs `restic check` when the last successful one is over 7 days old. **This is a readability check, not a restore-test (R-87 stays open).** The depth that ships ON now re-reads **100%** of the pack data (R-399 CLOSED, controller v0.228.0) — the sentence here previously said the opposite and was stale. **2026-08-31 (SPIKE R-87):** a scratch restore of every app on demo-hp was measured at **25 s for 8 snapshots / 774 MB**, against 40.3 s for one weekly check — but restic 0.14.0's `--verify` checks size and mtime, **not content**, so nothing available today can vouch for the restored BYTES. Row 4's verdict is UNCHANGED by the spike. `audits/SPIKE-restic-restore-test-2026-08-31.md`. **2026-09-01 (SPIKE R-95, `audits/SPIKE-r95-offsite-delete-2026-09-01.md`) — THIS ROW'S STATUS IS UNCHANGED, but the seven-day bound beside it is now known to be unverified.** Measured on BOTH boxes over their own SFTP credential, with controls: **no `.snapshots` is visible to either sub-account**, and the account is jailed. Either none exist or a sub-account cannot see them — and a snapshot the box cannot see is one the box cannot restore from, so recovery would be an operator act at the Hetzner panel, not a product capability (R-429; the confirming field needs the provider API, fenced by §11-D). **The PBS shape does NOT transfer**: the sub-account API has one permission axis, `readonly`, and a backup target cannot be read-only — PBS is a server that can refuse, a Storage Box is a filesystem that runs nothing. **What the spike removed as a fear:** withdrawing delete does NOT wedge the store (measured — restic treats a dead owner's lock as stale and proceeds). **What it added:** `unlock --remove-all` reports success while deleting nothing (R-430). **2026-09-01, LATER THE SAME DAY — THE PROBE ABOVE LOOKED FOR THE WRONG NAME AND THIS ROW'S SECOND CLAUSE IS NOW HALF WRONG.** It searched `.snapshots`; the vendor documents `/.zfs/snapshot`. Re-probed at the documented path on BOTH boxes, with controls, the two halves separate cleanly: **(a) the LIVE repository is deletable by the box — unchanged, R-95 stands;** **(b) the daily SNAPSHOTS of it are not writable by anything — PROVEN, not cited:** a write into `/.zfs/snapshot` is refused (`dest open …: Failure`) while the identical write to the account home succeeds. Seven daily snapshots confirmed in the panel (R-429). **So this row's "the restic repo is NOT protected the same way" is true of the repository and FALSE of its snapshots** — a deletion costs at most the day since the last snapshot, recoverable per-file (vendor). **Two limits kept honest:** a sub-account sees the snapshot directory EMPTY, so per-file recovery is operator-only today (R-432); and a panel-driven restore rolls back the WHOLE box. **The status is NOT moved** — evidence (b) is measured, but the recovery ROUTE has never been walked, which is what PARTIAL means. **Detection shipped hub v0.111.0 (R-431):** an unexplained fall in the snapshot count is noticed within a day. | +| 10 | **Ransomware / malicious deletion inside the guest** | PBS offsite (the box **cannot** delete its own snapshots); **the restic repo is NOT protected the same way** | whole-guest restore from PBS to a point before the event | **operator** (SSH) | | 7 d | **PARTIAL** | R-89 proved the box is refused when deleting its own PBS snapshot (`CAMPAIGN-8…:514-515`). **R-95 (open, ranked #1):** the restic credential **can delete** — `readonly=False`, `forget --prune` runs from the box, and SFTP cannot express append-only (`OPEN-ITEMS.md:13`) **2026-08-30 (R-359): the store is now VERIFIED on a cadence** — a daily `offsite-integrity` job runs `restic check` when the last successful one is over 7 days old. **This is a readability check, not a restore-test (R-87 stays open).** The depth that ships ON now re-reads **100%** of the pack data (R-399 CLOSED, controller v0.228.0) — the sentence here previously said the opposite and was stale. **2026-08-31 (SPIKE R-87):** a scratch restore of every app on demo-hp was measured at **25 s for 8 snapshots / 774 MB**, against 40.3 s for one weekly check — but restic 0.14.0's `--verify` checks size and mtime, **not content**, so nothing available today can vouch for the restored BYTES. Row 4's verdict is UNCHANGED by the spike. `audits/SPIKE-restic-restore-test-2026-08-31.md`. **2026-09-01 (SPIKE R-95, `audits/SPIKE-r95-offsite-delete-2026-09-01.md`) — THIS ROW'S STATUS IS UNCHANGED, but the seven-day bound beside it is now known to be unverified.** Measured on BOTH boxes over their own SFTP credential, with controls: **no `.snapshots` is visible to either sub-account**, and the account is jailed. Either none exist or a sub-account cannot see them — and a snapshot the box cannot see is one the box cannot restore from, so recovery would be an operator act at the Hetzner panel, not a product capability (R-429; the confirming field needs the provider API, fenced by §11-D). **The PBS shape does NOT transfer**: the sub-account API has one permission axis, `readonly`, and a backup target cannot be read-only — PBS is a server that can refuse, a Storage Box is a filesystem that runs nothing. **What the spike removed as a fear:** withdrawing delete does NOT wedge the store (measured — restic treats a dead owner's lock as stale and proceeds). **What it added:** `unlock --remove-all` reports success while deleting nothing (R-430). **2026-09-01, LATER THE SAME DAY — THE PROBE ABOVE LOOKED FOR THE WRONG NAME AND THIS ROW'S SECOND CLAUSE IS NOW HALF WRONG.** It searched `.snapshots`; the vendor documents `/.zfs/snapshot`. Re-probed at the documented path on BOTH boxes, with controls, the two halves separate cleanly: **(a) the LIVE repository is deletable by the box — unchanged, R-95 stands;** **(b) the daily SNAPSHOTS of it are not writable by anything — PROVEN, not cited:** a write into `/.zfs/snapshot` is refused (`dest open …: Failure`) while the identical write to the account home succeeds. Seven daily snapshots confirmed in the panel (R-429). **So this row's "the restic repo is NOT protected the same way" is true of the repository and FALSE of its snapshots** — a deletion costs at most the day since the last snapshot, recoverable per-file (vendor). **Two limits kept honest:** a sub-account sees the snapshot directory EMPTY, so per-file recovery is operator-only today (R-432); and a panel-driven restore rolls back the WHOLE box. **The status is NOT moved** — evidence (b) is measured, but the recovery ROUTE has never been walked, which is what PARTIAL means. **Detection shipped hub v0.111.0 (R-431):** an unexplained fall in the snapshot count is noticed within a day — **but only a fall of MORE than half (R-435).** **2026-09-01, THE DRILL — CLAUSE (b) OF THE RE-SCOPE ABOVE IS WITHDRAWN; THE STATUS IS STILL NOT MOVED.** The re-scope said a deletion is *recoverable per-file*. **It is not, from the box: no snapshot is reachable by ANY name.** MEASURED on demo-hp, read-only, no delete verb issued: **777,600** exact names in the vendor form `YYYY-MM-DDTHH-MM-SS` (nine full days, second granularity) plus 126 alternative shapes — **zero hits**, against a control where the identical 600-name batch returns a path that exists (6/6). **Structural cause:** `/home` (`u629488-sub3`) is **st_dev 0,82**, `/.zfs/snapshot` is **st_dev 0,276**, and `/home/.zfs` does not exist — a snapshot under `/.zfs/snapshot` belongs to a different dataset than the one holding the repository. **Clause (a) — the box cannot WRITE into the snapshot area — is unchanged and re-confirmed.** So this row's *"recoverable per-file (vendor)"* and its limit *"per-file recovery is operator-only today (R-432)"* both overstate what exists: the remaining routes are a panel rollback of the WHOLE Storage Box and the provider API (fenced, and the hub's client has no snapshot method at all). **R-432 ANSWERED negatively; R-433 opened. RTO still blank — nothing was recovered, so nothing was timed.** `audits/evidence-drill-r95-recovery-2026-09-01/`. | | 11 | **Hub lost** | every box's data plane, every tier, every Lane-1 route | none needed for recovery **of a box**; the hub itself restores from its Longhorn volume backup | operator (`kubectl`) | | 24 h + weekly (Longhorn `RETAIN 1` each) | **UNPROVEN** | a hub restore has never been performed. The backup target is `nfs://192.168.0.180` — **DooPlex itself** — and exactly **2** restore points exist (LIVE, INV Part D2.2) | | 11b | *consequences while the hub is gone* | — | — | — | | | **[FACT]** | **day:** nothing customer-visible breaks; events queue (`settings.go:1466-1485`). **week:** the operator alarm plane is dark, no claim/reset codes, no config or floor convergence, no PBS-secret re-issue. **permanently:** escrow custody and break-glass credentials are gone (INV Part D2.4) | | 12 | **Offsite provider lost (Hetzner)** | everything on-premises: both drives, both whole-guest tiers | none needed — on-premises recovery is unaffected. Re-provision a new offsite target. | operator | | | **[FACT]** | **restic and PBS share the provider** (INV Part E.1). Whether they share an account and payment method is **UNKNOWN** → §11-D | @@ -1085,7 +1085,7 @@ does **not** hold as written. → **R-108** | **R-356** | The off-site restore resolved its destination with the raw `HDD_PATH` and read an empty answer as "not installed" — **CLOSED, controller v0.219.0, 2026-08-22** | 40 of 53 apps were refused permanently while running (§6.3 `[DESIGN]`) | | ~~**R-108**~~ | ~~Network storage can host an app's namespace~~ | **CLOSED 2026-07-30, controller v0.187.0 — D5 UNBLOCKED.** An app namespace may no longer be placed on network storage (5 surfaces guarded by one fail-closed predicate); the share-root bind is deliberately UNCHANGED because it is load-bearing and unscopable (§10.1). `audits/R108-network-app-namespace-2026-07-30.md` | | **R-126** | A `.fab` bundle — plaintext secrets, optional password — can be exported ONTO a NAS: `storageDriveList()` (`internal/web/handler_export.go`) does not filter network paths | split out of R-108, which closed without it. NOT a D5 precondition: an explicit customer-chosen export destination, not a browsing surface reaching a backup tree (§5, §7.3) | -| R-95 (open) | The restic offsite credential **can delete** — the box can `forget --prune` its own repo, from **two** call sites (`offbox.go:1388` retention and `offbox.go:1759` over-quota) | the tier holding the customer's documents and photos is the one whose credential can destroy it (matrix row 10). **SPIKE 2026-09-01:** prevention needs a transport change (restic 0.14.0 does speak `rest:` — measured; append-only is a rest-server flag, not a restic one), because the sub-account API cannot express write-without-delete. **Detection is nearly free and is recommended first** — `snapshot_count` already reaches the hub and the hub appends reports, so the comparison needs no box change. `audits/SPIKE-r95-offsite-delete-2026-09-01.md`. **RE-SCOPED 2026-09-01:** the box can delete the LIVE repository but **cannot write to the daily snapshots of it** (measured, both boxes) — so the exposure is at most one day's data plus an operator-driven per-file recovery, not open-ended loss. Detection shipped hub v0.111.0 (R-431). | +| R-95 (open) | The restic offsite credential **can delete** — the box can `forget --prune` its own repo, from **two** call sites (`offbox.go:1388` retention and `offbox.go:1759` over-quota) | the tier holding the customer's documents and photos is the one whose credential can destroy it (matrix row 10). **SPIKE 2026-09-01:** prevention needs a transport change (restic 0.14.0 does speak `rest:` — measured; append-only is a rest-server flag, not a restic one), because the sub-account API cannot express write-without-delete. **Detection is nearly free and is recommended first** — `snapshot_count` already reaches the hub and the hub appends reports, so the comparison needs no box change. `audits/SPIKE-r95-offsite-delete-2026-09-01.md`. **RE-SCOPED 2026-09-01:** the box can delete the LIVE repository but **cannot write to the daily snapshots of it** (measured, both boxes) — so the exposure is at most one day's data plus an operator-driven per-file recovery, not open-ended loss. Detection shipped hub v0.111.0 (R-431). **DRILL 2026-09-01: the re-scope's recovery clause is WITHDRAWN — no snapshot is reachable from a sub-account by any name (R-433, 777,600 names, controlled), so the exposure is NOT bounded by an operator-driven per-file recovery. R-436 is the new cheap lead: the provider already offers `rclone serve restic --stdio` server-side and restic 0.14.0 speaks `rclone:` (measured) — but the client supplies the server command line, so ask the vendor whether `--append-only` is pinned BEFORE building anything.** | | ~~R-86~~ | ~~Restore-tests are interval-scheduled, not backup-aligned~~ | **CLOSED 2026-08-03 — agent v0.121.0 + hub v0.91.0.** Restore-testing is now **per archive generation**: a tier is due when its newest archive that has settled ~24 h has not been proven, so a daily tier is proved daily on its own archive and a weekly tier weekly on its own. The ticker survives only as the evaluation interval (6 h, chosen from a measured cost). The hub's staleness window moved with it — per tier, from that tier's observed archive rhythm — because a weekly tier proved weekly sat EXACTLY on the old flat 7-day line (§3, Lane 2's per-archive rule) | | ~~R-353~~ | ~~A local unit restore reported a bare completion whether it returned an entire dataset or nothing~~ | **CLOSED 2026-08-30, controller v0.226.0.** The volume count already existed (`restoreDockerVolumesFrom`) and was discarded by a one-line wrapper, so the surface was structurally unable to say what came back. `RestoreFromRecoveryUnit` now returns `UnitRestoreResult` and the sentence has three cases, keyed on replayed-vs-**listed** — zero-replayed has two causes that are opposite news. Every sentence is a claim about the BACKUP, never the app: this path has no `SafetyDump` discriminator, and §6.3 is why that is not pedantry. Proven live on `demo-hp` | | ~~R-357~~ | ~~The destructive reconstitute had no free-space gate; all three that existed guarded non-destructive paths~~ | **CLOSED 2026-08-30, controller v0.226.0.** The gate sits before `writeSafetyDump` and `StopStack`, so a refusal costs no outage — the test asserts the StopStack call count, not the sentence. No headroom multiplier (a measured tree, not a predicted download); fail-closed when either probe reads ≤ 0, which was a real fail-open hole. **Not live-validated** — filling a filesystem is a drill step | diff --git a/documentation/audits/evidence-drill-r95-recovery-2026-09-01/README.md b/documentation/audits/evidence-drill-r95-recovery-2026-09-01/README.md new file mode 100644 index 00000000..2ed2657d --- /dev/null +++ b/documentation/audits/evidence-drill-r95-recovery-2026-09-01/README.md @@ -0,0 +1,52 @@ +# Evidence — DRILL: the deletion we said is survivable (R-95 recovery drill, 2026-09-01) + +**The drill STOPPED at the end of Phase 1, on the operator's ruling, before any destructive step.** +**No delete verb was issued against any live store. No byte on either Storage Box sub-account was +written, moved or removed.** Every probe in this directory is a read: `ls`, `stat`, `tree`, `du`, +`df`, `rsync --list-only`, `restic snapshots`. + +## Why it stopped + +Phase 1 was supposed to *find* a Storage Box snapshot so Phase 4 could recover from it. It found +that **no snapshot is reachable from the box by any name**, and that **no unfenced route to one +exists for CC at all**. The recovery leg therefore could not run, and deleting a real app's off-site +history would have bought only an alarm test that — measured against the shipped threshold — could +not have fired at the size the runbook specifies. Put to the operator as a two-option decision; +the ruling was **stop and report**. + +## Files + +| file | what it is | +|---|---| +| `phase0-ground-truth/01-offsite-inventory.txt` | demo-hp's full off-site inventory — 69 restic snapshots, 9 apps, ids + tags + paths. The baseline. | +| `phase0-ground-truth/02-hub-view.txt` | the hub's latest `offsite` object for demo-hp (`snapshot_count: 69`, `stats_known: true`, `last_status: ok`). Cross-checks the restic count exactly. | +| `phase1-find-the-snapshot/01-probe-tree.txt` | SFTP probes of `/.zfs`, `/.zfs/snapshot`, `/home/.zfs`, with a positive and a negative control. Reproduces R-432. | +| `phase1-find-the-snapshot/02-shell-and-rsync-doors.txt` | two independent tools (a real shell on port 23, and `rsync --list-only`) agree the tree lists empty. Both controlled. | +| `phase1-find-the-snapshot/03-shell-capabilities.txt` | the port-23 restricted shell's full `help` — the command set the credential actually reaches. | +| `phase1-find-the-snapshot/04-enumeration-attempts.txt` | `df`/`stat`/`tree`/`du` on the snapshot door. **Where the st_dev split was found.** | +| `phase1-find-the-snapshot/05-home-zfs-door.txt` | `/home/.zfs` does not exist — three tools, negative control in the same run. | +| `phase1-find-the-snapshot/06-name-sweep.txt` | the name sweep: **nine full days, second granularity, 777,600 candidate names, zero hits.** | +| `phase1-find-the-snapshot/07-sweep-control.txt` | **the control for that sweep** — the identical 600-name batch shape with one real path appended; 6/6 batches returned it. Without this the zero result would prove nothing. | +| `phase1-find-the-snapshot/08-alt-formats.txt` | 126 non-timestamp name shapes and alternative snapshot paths. Only the controls resolved. | +| `phase1-find-the-snapshot/09-rclone-restic-lead.txt` | the `rclone:` backend measurement behind R-436, with the `banana:` invalid-backend control. | +| `phase6-teardown/01-verify-then-clean.txt` | the store is untouched: still 69 snapshots, home is exactly `.ssh` + `felhom-repo`, repo top level intact. | +| `phase6-teardown/02-guest-and-host-clean.txt` | scratch removed from the guest and the PVE host; controller 0.232.0 healthy, 19 healthy containers. | +| `phase6-teardown/03-dooplex-clean-and-final-state.txt` | hub DB copies shredded (both, `-wal` included); both customers' tiers still `ok`; no alarm event raised by this session. | + +## The method that makes the negative result trustworthy + +A sweep that finds nothing is worthless unless it is shown it *could* have found something. The +oracle is a batched `stat -c %n` over the port-23 shell: 500–600 paths per round trip, and stdout +carries **only the paths that exist**. `07-sweep-control.txt` runs the identical batch shape with +`/home` appended and gets `/home` back, 6 times out of 6. So the zero in `06-name-sweep.txt` is a +measurement, not a silence. + +## Not done, and why + +* **No Hetzner API call.** Fenced by the runbook (§11-D). The hub's own API client has no snapshot + method at all, so even unfenced it would have needed new code. +* **No panel action.** No browser on DooPlex, and "Restore snapshot" is fenced in any case — it rolls + back the whole Storage Box and deletes newer snapshots. +* **`demo-felhom` never addressed.** Every storage-box connection in this session authenticated as + `u629488-sub3`, demo-hp's own sub-account. Its tier is verified still reporting `ok` in + `phase6-teardown/03-*`. diff --git a/documentation/audits/evidence-drill-r95-recovery-2026-09-01/phase0-ground-truth/01-offsite-inventory.txt b/documentation/audits/evidence-drill-r95-recovery-2026-09-01/phase0-ground-truth/01-offsite-inventory.txt new file mode 100644 index 00000000..0a6439ba --- /dev/null +++ b/documentation/audits/evidence-drill-r95-recovery-2026-09-01/phase0-ground-truth/01-offsite-inventory.txt @@ -0,0 +1,162 @@ +### date (guest UTC): 2026-09-01T14:30:31Z +### restic snapshots --json is too noisy; table form: +ID Time Host Tags Paths +------------------------------------------------------------------------------------------------------------------------------ +41c830db 2026-08-09 08:30:38 demo-hp felhom-offbox,calibre-web /mnt/sys_drive/felhom-data/backups/primary/calibre-web + /mnt/sys_drive/felhom-data/userdata/media/books + +9e38b84c 2026-08-09 08:30:49 demo-hp felhom-offbox,opengist /mnt/sys_drive/felhom-data/backups/primary/opengist + +78b93f04 2026-08-09 08:30:53 demo-hp felhom-offbox,privatebin /mnt/sys_drive/felhom-data/backups/primary/privatebin + +f06ad130 2026-08-23 02:16:25 demo-hp felhom-offbox,bookstack /mnt/sys_drive/felhom-data/backups/primary/bookstack + +e68f6444 2026-08-23 02:16:32 demo-hp felhom-offbox,calibre-web /mnt/felhom-drives/hdd_1/backups/primary/calibre-web + /mnt/felhom-drives/hdd_1/userdata/media/books + +d1a614b5 2026-08-23 02:16:36 demo-hp felhom-offbox,privatebin /mnt/sys_drive/felhom-data/backups/primary/privatebin + +e23e34a5 2026-08-23 02:16:39 demo-hp felhom-offbox,docmost /mnt/sys_drive/felhom-data/backups/primary/docmost + +5c3492ad 2026-08-23 02:16:45 demo-hp felhom-offbox,kimai /mnt/sys_drive/felhom-data/backups/primary/kimai + +82ef8692 2026-08-23 02:16:55 demo-hp felhom-offbox,opengist /mnt/sys_drive/felhom-data/backups/primary/opengist + +778e8feb 2026-08-23 02:16:58 demo-hp felhom-offbox,paperless-ngx /mnt/felhom-drives/hdd_1/appdata/paperless/media + /mnt/felhom-drives/hdd_1/backups/primary/paperless-ngx + +8cc0e4b9 2026-08-23 02:17:04 demo-hp felhom-offbox,romm /mnt/felhom-drives/hdd_1/backups/primary/romm + +936a30c3 2026-08-26 02:16:29 demo-hp felhom-offbox,kimai /mnt/sys_drive/felhom-data/backups/primary/kimai + +13c3cf9d 2026-08-26 02:16:39 demo-hp felhom-offbox,opengist /mnt/sys_drive/felhom-data/backups/primary/opengist + +9e3474ec 2026-08-26 02:16:43 demo-hp felhom-offbox,romm /mnt/felhom-drives/hdd_1/backups/primary/romm + +a2aec191 2026-08-26 02:16:49 demo-hp felhom-offbox,paperless-ngx /mnt/felhom-drives/hdd_1/appdata/paperless/media + /mnt/felhom-drives/hdd_1/backups/primary/paperless-ngx + +389a3a6d 2026-08-26 02:16:55 demo-hp felhom-offbox,privatebin /mnt/sys_drive/felhom-data/backups/primary/privatebin + +c7408e7d 2026-08-26 02:16:59 demo-hp felhom-offbox,bookstack /mnt/sys_drive/felhom-data/backups/primary/bookstack + +f40c8fd9 2026-08-26 02:17:03 demo-hp felhom-offbox,calibre-web /mnt/felhom-drives/hdd_1/backups/primary/calibre-web + /mnt/felhom-drives/hdd_1/userdata/media/books + +ea5e9a25 2026-08-26 02:17:07 demo-hp felhom-offbox,docmost /mnt/sys_drive/felhom-data/backups/primary/docmost + +e6ee92ac 2026-08-27 02:16:26 demo-hp felhom-offbox,paperless-ngx /mnt/felhom-drives/hdd_1/appdata/paperless/media + /mnt/felhom-drives/hdd_1/backups/primary/paperless-ngx + +73266947 2026-08-27 02:16:34 demo-hp felhom-offbox,privatebin /mnt/sys_drive/felhom-data/backups/primary/privatebin + +794d8219 2026-08-27 02:16:37 demo-hp felhom-offbox,bookstack /mnt/sys_drive/felhom-data/backups/primary/bookstack + +b1e4d6a5 2026-08-27 02:16:41 demo-hp felhom-offbox,calibre-web /mnt/felhom-drives/hdd_1/backups/primary/calibre-web + /mnt/felhom-drives/hdd_1/userdata/media/books + +7eea66d2 2026-08-27 02:16:45 demo-hp felhom-offbox,docmost /mnt/sys_drive/felhom-data/backups/primary/docmost + +b9db5519 2026-08-27 02:16:50 demo-hp felhom-offbox,kimai /mnt/sys_drive/felhom-data/backups/primary/kimai + +0751dad2 2026-08-27 02:17:00 demo-hp felhom-offbox,opengist /mnt/sys_drive/felhom-data/backups/primary/opengist + +3714731c 2026-08-27 02:17:03 demo-hp felhom-offbox,romm /mnt/felhom-drives/hdd_1/backups/primary/romm + +8dcbd53a 2026-08-28 02:16:30 demo-hp felhom-offbox,bookstack /mnt/sys_drive/felhom-data/backups/primary/bookstack + +5a635785 2026-08-28 02:16:34 demo-hp felhom-offbox,calibre-web /mnt/felhom-drives/hdd_1/backups/primary/calibre-web + /mnt/felhom-drives/hdd_1/userdata/media/books + +035e923f 2026-08-28 02:16:38 demo-hp felhom-offbox,docmost /mnt/sys_drive/felhom-data/backups/primary/docmost + +6f613bbd 2026-08-28 02:16:43 demo-hp felhom-offbox,kimai /mnt/sys_drive/felhom-data/backups/primary/kimai + +f0c6ce70 2026-08-28 02:16:53 demo-hp felhom-offbox,opengist /mnt/sys_drive/felhom-data/backups/primary/opengist + +02f1de9d 2026-08-28 02:16:56 demo-hp felhom-offbox,romm /mnt/felhom-drives/hdd_1/backups/primary/romm + +726c1ce8 2026-08-28 02:17:02 demo-hp felhom-offbox,paperless-ngx /mnt/felhom-drives/hdd_1/appdata/paperless/media + /mnt/felhom-drives/hdd_1/backups/primary/paperless-ngx + +c0e9dc25 2026-08-28 02:17:08 demo-hp felhom-offbox,privatebin /mnt/sys_drive/felhom-data/backups/primary/privatebin + +0b90da52 2026-08-29 02:16:26 demo-hp felhom-offbox,bookstack /mnt/sys_drive/felhom-data/backups/primary/bookstack + +6aee1818 2026-08-29 02:16:30 demo-hp felhom-offbox,calibre-web /mnt/felhom-drives/hdd_1/backups/primary/calibre-web + /mnt/felhom-drives/hdd_1/userdata/media/books + +7aea3e6d 2026-08-29 02:16:34 demo-hp felhom-offbox,docmost /mnt/sys_drive/felhom-data/backups/primary/docmost + +2f043b35 2026-08-29 02:16:39 demo-hp felhom-offbox,kimai /mnt/sys_drive/felhom-data/backups/primary/kimai + +fa862b02 2026-08-29 02:16:49 demo-hp felhom-offbox,opengist /mnt/sys_drive/felhom-data/backups/primary/opengist + +3f93ec62 2026-08-29 02:16:52 demo-hp felhom-offbox,romm /mnt/felhom-drives/hdd_1/backups/primary/romm + +ebade956 2026-08-29 02:16:58 demo-hp felhom-offbox,paperless-ngx /mnt/felhom-drives/hdd_1/appdata/paperless/media + /mnt/felhom-drives/hdd_1/backups/primary/paperless-ngx + +983a0365 2026-08-29 02:17:06 demo-hp felhom-offbox,privatebin /mnt/sys_drive/felhom-data/backups/primary/privatebin + +9ae0aa88 2026-08-30 02:16:29 demo-hp felhom-offbox,romm /mnt/felhom-drives/hdd_1/backups/primary/romm + +2404511d 2026-08-30 02:16:35 demo-hp felhom-offbox,paperless-ngx /mnt/felhom-drives/hdd_1/appdata/paperless/media + /mnt/felhom-drives/hdd_1/backups/primary/paperless-ngx + +4a5db84d 2026-08-30 02:16:42 demo-hp felhom-offbox,privatebin /mnt/sys_drive/felhom-data/backups/primary/privatebin + +ea40c93a 2026-08-30 02:16:45 demo-hp felhom-offbox,bookstack /mnt/sys_drive/felhom-data/backups/primary/bookstack + +37b7b51a 2026-08-30 02:16:49 demo-hp felhom-offbox,calibre-web /mnt/felhom-drives/hdd_1/backups/primary/calibre-web + /mnt/felhom-drives/hdd_1/userdata/media/books + +ee67d85d 2026-08-30 02:16:53 demo-hp felhom-offbox,docmost /mnt/sys_drive/felhom-data/backups/primary/docmost + +84542ec8 2026-08-30 02:16:58 demo-hp felhom-offbox,kimai /mnt/sys_drive/felhom-data/backups/primary/kimai + +a5c5c973 2026-08-30 02:17:08 demo-hp felhom-offbox,opengist /mnt/sys_drive/felhom-data/backups/primary/opengist + +05c3346a 2026-08-31 21:29:03 demo-hp felhom-offbox,bookstack /mnt/sys_drive/felhom-data/backups/primary/bookstack + +b6b7c205 2026-08-31 21:29:08 demo-hp felhom-offbox,calibre-web /mnt/felhom-drives/hdd_1/backups/primary/calibre-web + /mnt/felhom-drives/hdd_1/userdata/media/books + +0b5781f5 2026-08-31 21:29:12 demo-hp felhom-offbox,docmost /mnt/sys_drive/felhom-data/backups/primary/docmost + +6a2672b0 2026-08-31 21:29:16 demo-hp felhom-offbox,paperless-ngx /mnt/felhom-drives/hdd_1/appdata/paperless/media + /mnt/felhom-drives/hdd_1/backups/primary/paperless-ngx + +d3970ee3 2026-08-31 21:29:21 demo-hp felhom-offbox,kimai /mnt/sys_drive/felhom-data/backups/primary/kimai + +5d5bee7f 2026-08-31 21:29:31 demo-hp felhom-offbox,opengist /mnt/sys_drive/felhom-data/backups/primary/opengist + +7b0a7161 2026-08-31 21:29:34 demo-hp felhom-offbox,privatebin /mnt/sys_drive/felhom-data/backups/primary/privatebin + +1b8c7361 2026-08-31 21:29:38 demo-hp felhom-offbox,romm /mnt/felhom-drives/hdd_1/backups/primary/romm + +5dcee6e3 2026-08-31 21:29:44 demo-hp felhom-offbox,bentopdf /mnt/sys_drive/felhom-data/backups/primary/bentopdf + +6fee3b5a 2026-09-01 02:16:30 demo-hp felhom-offbox,calibre-web /mnt/felhom-drives/hdd_1/backups/primary/calibre-web + /mnt/felhom-drives/hdd_1/userdata/media/books + +b2e059ee 2026-09-01 02:16:34 demo-hp felhom-offbox,docmost /mnt/sys_drive/felhom-data/backups/primary/docmost + +1e179cfe 2026-09-01 02:16:39 demo-hp felhom-offbox,paperless-ngx /mnt/felhom-drives/hdd_1/appdata/paperless/media + /mnt/felhom-drives/hdd_1/backups/primary/paperless-ngx + +f69e0510 2026-09-01 02:16:45 demo-hp felhom-offbox,kimai /mnt/sys_drive/felhom-data/backups/primary/kimai + +8b9d9a94 2026-09-01 02:16:55 demo-hp felhom-offbox,opengist /mnt/sys_drive/felhom-data/backups/primary/opengist + +1e548d3b 2026-09-01 02:16:58 demo-hp felhom-offbox,privatebin /mnt/sys_drive/felhom-data/backups/primary/privatebin + +ae903b66 2026-09-01 02:17:01 demo-hp felhom-offbox,romm /mnt/felhom-drives/hdd_1/backups/primary/romm + +9d002b38 2026-09-01 02:17:08 demo-hp felhom-offbox,bentopdf /mnt/sys_drive/felhom-data/backups/primary/bentopdf + +7a855a65 2026-09-01 02:17:11 demo-hp felhom-offbox,bookstack /mnt/sys_drive/felhom-data/backups/primary/bookstack +------------------------------------------------------------------------------------------------------------------------------ +69 snapshots + +### rc=0 diff --git a/documentation/audits/evidence-drill-r95-recovery-2026-09-01/phase0-ground-truth/02-hub-view.txt b/documentation/audits/evidence-drill-r95-recovery-2026-09-01/phase0-ground-truth/02-hub-view.txt new file mode 100644 index 00000000..92492486 --- /dev/null +++ b/documentation/audits/evidence-drill-r95-recovery-2026-09-01/phase0-ground-truth/02-hub-view.txt @@ -0,0 +1,21 @@ +report id 22221 received_at 2026-09-01 14:45:18 +top-level keys: ['app_telemetry', 'backup', 'claimed', 'config_hash', 'containers', 'controller_url', 'controller_version', 'customer_id', 'customer_name', 'dr_recipe', 'geo_restriction', 'health', 'offsite', 'stacks', 'storage', 'system', 'timestamp', 'version'] +--- offsite --- +{ + "enabled": true, + "escrow_state": "escrowed", + "last_run": "2026-09-01T02:17:57Z", + "last_status": "ok", + "last_success": "2026-09-01T02:17:57Z", + "snapshot_count": 69, + "repo_size_bytes": 147274432, + "quota_gb": 50, + "stats_known": true, + "last_integrity_check": "2026-09-01T08:24:24Z", + "last_integrity_ok": true, + "last_integrity_depth": "100%", + "last_proof_run": "2026-09-01T04:04:48Z", + "last_proof_stack": "bookstack", + "last_proof_snapshot": "7a855a65", + "last_proof_result": "pass" +} diff --git a/documentation/audits/evidence-drill-r95-recovery-2026-09-01/phase1-find-the-snapshot/01-probe-tree.txt b/documentation/audits/evidence-drill-r95-recovery-2026-09-01/phase1-find-the-snapshot/01-probe-tree.txt new file mode 100644 index 00000000..f23ca103 --- /dev/null +++ b/documentation/audits/evidence-drill-r95-recovery-2026-09-01/phase1-find-the-snapshot/01-probe-tree.txt @@ -0,0 +1,31 @@ +--- ls -la . +Connected to u629488-sub3.your-storagebox.de. +sftp> ls -la . +drwxr-xr-x ? u629488-sub3 1058 4 Sep 1 12:12 ./. +dr-x--x--x ? root root 11 Jul 21 16:01 ./.. +drwx------ ? u629488-sub3 1058 3 Jul 23 09:53 ./.ssh +drwxrwxr-x ? u629488-sub3 1058 8 Aug 4 12:38 ./felhom-repo +--- ls -la ./zzz-no-such-r95-negative-control +Connected to u629488-sub3.your-storagebox.de. +sftp> ls -la ./zzz-no-such-r95-negative-control +Can't ls: "/home/./zzz-no-such-r95-negative-control" not found +--- ls -la /.zfs +Connected to u629488-sub3.your-storagebox.de. +sftp> ls -la /.zfs +drwxrwxrwx ? root root 0 Jul 21 16:01 /.zfs/. +dr-x--x--x ? root root 11 Jul 21 16:01 /.zfs/.. +drwxrwxrwx ? root root 2 Sep 1 12:12 /.zfs/shares +drwxrwxrwx ? root root 2 Jan 1 1970 /.zfs/snapshot +--- ls -la /.zfs/snapshot +Connected to u629488-sub3.your-storagebox.de. +sftp> ls -la /.zfs/snapshot +drwxrwxrwx ? root root 2 Jan 1 1970 /.zfs/snapshot/. +drwxrwxrwx ? root root 0 Jul 21 16:01 /.zfs/snapshot/.. +--- ls -la /home/.zfs +Connected to u629488-sub3.your-storagebox.de. +sftp> ls -la /home/.zfs +Can't ls: "/home/.zfs" not found +--- ls -la /home/.zfs/snapshot +Connected to u629488-sub3.your-storagebox.de. +sftp> ls -la /home/.zfs/snapshot +Can't ls: "/home/.zfs/snapshot" not found diff --git a/documentation/audits/evidence-drill-r95-recovery-2026-09-01/phase1-find-the-snapshot/02-shell-and-rsync-doors.txt b/documentation/audits/evidence-drill-r95-recovery-2026-09-01/phase1-find-the-snapshot/02-shell-and-rsync-doors.txt new file mode 100644 index 00000000..6cac9db4 --- /dev/null +++ b/documentation/audits/evidence-drill-r95-recovery-2026-09-01/phase1-find-the-snapshot/02-shell-and-rsync-doors.txt @@ -0,0 +1,21 @@ +=== A. ssh with a command (limited shell?) === +total 0 +drwxrwxrwx 2 root root 2 Jan 1 1970 . +drwxrwxrwx 1 root root 0 Jul 21 16:01 .. +rc=0 + +=== B. rsync --list-only on the account home (POSITIVE CONTROL) === +drwxr-xr-x 4 2026/09/01 12:12:48 . +drwx------ 3 2026/07/23 09:53:43 .ssh +drwxrwxr-x 8 2026/08/04 12:38:16 felhom-repo +rc=0 + +=== C. rsync --list-only on /.zfs/snapshot/ (THE PROBE) === +drwxrwxrwx 2 1970/01/01 00:00:00 . +rc=0 + +=== D. rsync --list-only on a path that must NOT exist (NEGATIVE CONTROL) === +rsync: [sender] change_dir "/zzz-no-such-r95" failed: No such file or directory (2) +rsync error: some files/attrs were not transferred (see previous errors) (code 23) at main.c(1874) [Receiver=3.2.7] +rsync: [Receiver] write error: Broken pipe (32) +rc=0 diff --git a/documentation/audits/evidence-drill-r95-recovery-2026-09-01/phase1-find-the-snapshot/03-shell-capabilities.txt b/documentation/audits/evidence-drill-r95-recovery-2026-09-01/phase1-find-the-snapshot/03-shell-capabilities.txt new file mode 100644 index 00000000..239c9eca --- /dev/null +++ b/documentation/audits/evidence-drill-r95-recovery-2026-09-01/phase1-find-the-snapshot/03-shell-capabilities.txt @@ -0,0 +1,54 @@ +=== id / pwd / shell === +Command not found. Use 'help' to get a list of available commands. +=== ls / === +/usr/bin/ls: cannot open directory '/': Permission denied +=== mount === +/usr/bin/cat: /proc/self/mountinfo: No such file or directory +=== which tools === +Command not found. Use 'help' to get a list of available commands. +=== help === ++------------------------------------------------------------------------------+ +| The following commands are available: | +| ls|ll list directory content | +| tree list directory content | +| cd change current working directory | +| pwd show current working directory | +| mkdir create new directory | +| rmdir delete directory | +| du disk usage of files/directories | +| df show disk usage | +| dd read and write files | +| cat output file content | +| touch create new file | +| cp copy files/directories | +| rm delete files/directories | +| unlink delete file/directory | +| mv move files/directories | +| chmod change file/directory permissions | +| md5|sha1|sha256|sha512 create hash sum of file | +| md5sum|sha1sum|sha256sum|sha512sum create hash sum of file | +| head show first lines of file | +| tail show last lines of file | +| grep search for specific string in files | +| stat stat files/directory | +| version show version of this SSH environment | +| | +| Available as server side backend: | +| borg | +| rsync | +| scp | +| sftp | +| rclone serve restic --stdio | +| | +| Please note that this is only a restricted shell which do not | +| support shell features like redirects or pipes. | +| | +| You can find more information in our Docs: | +| https://docs.hetzner.com/storage/storage-box/ | ++------------------------------------------------------------------------------+ +=== snapshot-ish commands === +Command not found. Use 'help' to get a list of available commands. +-- +Command not found. Use 'help' to get a list of available commands. +-- +Command not found. Use 'help' to get a list of available commands. diff --git a/documentation/audits/evidence-drill-r95-recovery-2026-09-01/phase1-find-the-snapshot/04-enumeration-attempts.txt b/documentation/audits/evidence-drill-r95-recovery-2026-09-01/phase1-find-the-snapshot/04-enumeration-attempts.txt new file mode 100644 index 00000000..10900974 --- /dev/null +++ b/documentation/audits/evidence-drill-r95-recovery-2026-09-01/phase1-find-the-snapshot/04-enumeration-attempts.txt @@ -0,0 +1,41 @@ +=== df === +Filesystem 1K-blocks Used Available Use% Mounted on +u629488-sub3 1073645568 2685568 1070960000 1% /home +=== stat /.zfs/snapshot === + File: /.zfs/snapshot + Size: 2 Blocks: 0 IO Block: 512 directory +Device: 0,276 Inode: 281474976710653 Links: 2 +Access: (0777/drwxrwxrwx) Uid: ( 0/ root) Gid: ( 0/ root) +Access: 2026-09-01 14:36:07.861516179 +0000 +Modify: 1970-01-01 00:00:00.000000000 +0000 +Change: 1970-01-01 00:00:00.000000000 +0000 + Birth: - +=== ls -a /.zfs/snapshot === +. +.. +=== tree /.zfs === +/.zfs +├── shares +└── snapshot + +3 directories, 0 files +=== du /.zfs/snapshot === +0 /.zfs/snapshot +=== multi-arg stat test (2 real, 1 fake) === + File: /home + Size: 4 Blocks: 1 IO Block: 131072 directory +Device: 0,82 Inode: 2273 Links: 4 +Access: (0755/drwxr-xr-x) Uid: ( 1057/u629488-sub3) Gid: ( 1058/ UNKNOWN) +Access: 2026-07-21 16:01:41.810597438 +0000 +Modify: 2026-09-01 12:12:48.637707029 +0000 +Change: 2026-09-01 12:12:48.637707029 +0000 + Birth: 2026-07-21 16:01:41.810597438 +0000 + File: /.zfs/snapshot + Size: 2 Blocks: 0 IO Block: 512 directory +Device: 0,276 Inode: 281474976710653 Links: 2 +Access: (0777/drwxrwxrwx) Uid: ( 0/ root) Gid: ( 0/ root) +Access: 2026-09-01 14:36:10.597470894 +0000 +Modify: 1970-01-01 00:00:00.000000000 +0000 +Change: 1970-01-01 00:00:00.000000000 +0000 + Birth: - +/usr/bin/stat: cannot statx '/zzz-no-such-r95': No such file or directory diff --git a/documentation/audits/evidence-drill-r95-recovery-2026-09-01/phase1-find-the-snapshot/05-home-zfs-door.txt b/documentation/audits/evidence-drill-r95-recovery-2026-09-01/phase1-find-the-snapshot/05-home-zfs-door.txt new file mode 100644 index 00000000..d7deb7d5 --- /dev/null +++ b/documentation/audits/evidence-drill-r95-recovery-2026-09-01/phase1-find-the-snapshot/05-home-zfs-door.txt @@ -0,0 +1,16 @@ +=== stat /home/.zfs and /home/.zfs/snapshot (fake path = negative control) === +/usr/bin/stat: cannot statx '/home/.zfs': No such file or directory +/usr/bin/stat: cannot statx '/home/.zfs/snapshot': No such file or directory +/usr/bin/stat: cannot statx '/home/zzz-no-such-r95': No such file or directory + +=== ls -a /home/.zfs/snapshot === +/usr/bin/ls: cannot access '/home/.zfs/snapshot': No such file or directory +=== ls -a /home === +. +.. +.ssh +felhom-repo +=== tree /home/.zfs === +/home/.zfs [error opening dir] + +0 directories, 0 files diff --git a/documentation/audits/evidence-drill-r95-recovery-2026-09-01/phase1-find-the-snapshot/06-name-sweep.txt b/documentation/audits/evidence-drill-r95-recovery-2026-09-01/phase1-find-the-snapshot/06-name-sweep.txt new file mode 100644 index 00000000..87d77e0c --- /dev/null +++ b/documentation/audits/evidence-drill-r95-recovery-2026-09-01/phase1-find-the-snapshot/06-name-sweep.txt @@ -0,0 +1,165 @@ +warm-up / positive control: /home +........................ day 2026-09-01 done, batches-with-hits=0 +........................ day 2026-08-31 done, batches-with-hits=0 +### CONTROL RUN — same code path, /home injected into every batch; MUST hit 144/144 +HIT(2026-09-01 00:0x): /home +HIT(2026-09-01 00:1x): /home +HIT(2026-09-01 00:2x): /home +HIT(2026-09-01 00:3x): /home +HIT(2026-09-01 00:4x): /home +HIT(2026-09-01 00:5x): /home +.HIT(2026-09-01 01:0x): /home +HIT(2026-09-01 01:1x): /home +HIT(2026-09-01 01:2x): /home +HIT(2026-09-01 01:3x): /home +HIT(2026-09-01 01:4x): /home +HIT(2026-09-01 01:5x): /home +.HIT(2026-09-01 02:0x): /home +HIT(2026-09-01 02:1x): /home +HIT(2026-09-01 02:2x): /home +HIT(2026-09-01 02:3x): /home +HIT(2026-09-01 02:4x): /home +HIT(2026-09-01 02:5x): /home +.HIT(2026-09-01 03:0x): /home +HIT(2026-09-01 03:1x): /home +HIT(2026-09-01 03:2x): /home +HIT(2026-09-01 03:3x): /home +HIT(2026-09-01 03:4x): /home +HIT(2026-09-01 03:5x): /home +.HIT(2026-09-01 04:0x): /home +HIT(2026-09-01 04:1x): /home +HIT(2026-09-01 04:2x): /home +HIT(2026-09-01 04:3x): /home +HIT(2026-09-01 04:4x): /home +HIT(2026-09-01 04:5x): /home +.HIT(2026-09-01 05:0x): /home +HIT(2026-09-01 05:1x): /home +HIT(2026-09-01 05:2x): /home +HIT(2026-09-01 05:3x): /home +HIT(2026-09-01 05:4x): /home +HIT(2026-09-01 05:5x): /home +.HIT(2026-09-01 06:0x): /home +HIT(2026-09-01 06:1x): /home +HIT(2026-09-01 06:2x): /home +HIT(2026-09-01 06:3x): /home +HIT(2026-09-01 06:4x): /home +HIT(2026-09-01 06:5x): /home +.HIT(2026-09-01 07:0x): /home +HIT(2026-09-01 07:1x): /home +HIT(2026-09-01 07:2x): /home +HIT(2026-09-01 07:3x): /home +HIT(2026-09-01 07:4x): /home +HIT(2026-09-01 07:5x): /home +.HIT(2026-09-01 08:0x): /home +HIT(2026-09-01 08:1x): /home +HIT(2026-09-01 08:2x): /home +HIT(2026-09-01 08:3x): /home +HIT(2026-09-01 08:4x): /home +HIT(2026-09-01 08:5x): /home +.HIT(2026-09-01 09:0x): /home +HIT(2026-09-01 09:1x): /home +HIT(2026-09-01 09:2x): /home +HIT(2026-09-01 09:3x): /home +HIT(2026-09-01 09:4x): /home +HIT(2026-09-01 09:5x): /home +.HIT(2026-09-01 10:0x): /home +HIT(2026-09-01 10:1x): /home +HIT(2026-09-01 10:2x): /home +HIT(2026-09-01 10:3x): /home +HIT(2026-09-01 10:4x): /home +HIT(2026-09-01 10:5x): /home +.HIT(2026-09-01 11:0x): /home +HIT(2026-09-01 11:1x): /home +HIT(2026-09-01 11:2x): /home +HIT(2026-09-01 11:3x): /home +HIT(2026-09-01 11:4x): /home +HIT(2026-09-01 11:5x): /home +.HIT(2026-09-01 12:0x): /home +HIT(2026-09-01 12:1x): /home +HIT(2026-09-01 12:2x): /home +HIT(2026-09-01 12:3x): /home +HIT(2026-09-01 12:4x): /home +HIT(2026-09-01 12:5x): /home +.HIT(2026-09-01 13:0x): /home +HIT(2026-09-01 13:1x): /home +HIT(2026-09-01 13:2x): /home +HIT(2026-09-01 13:3x): /home +HIT(2026-09-01 13:4x): /home +HIT(2026-09-01 13:5x): /home +.HIT(2026-09-01 14:0x): /home +HIT(2026-09-01 14:1x): /home +HIT(2026-09-01 14:2x): /home +HIT(2026-09-01 14:3x): /home +HIT(2026-09-01 14:4x): /home +HIT(2026-09-01 14:5x): /home +.HIT(2026-09-01 15:0x): /home +HIT(2026-09-01 15:1x): /home +HIT(2026-09-01 15:2x): /home +HIT(2026-09-01 15:3x): /home +HIT(2026-09-01 15:4x): /home +HIT(2026-09-01 15:5x): /home +.HIT(2026-09-01 16:0x): /home +HIT(2026-09-01 16:1x): /home +HIT(2026-09-01 16:2x): /home +HIT(2026-09-01 16:3x): /home +HIT(2026-09-01 16:4x): /home +HIT(2026-09-01 16:5x): /home +.HIT(2026-09-01 17:0x): /home +HIT(2026-09-01 17:1x): /home +HIT(2026-09-01 17:2x): /home +HIT(2026-09-01 17:3x): /home +HIT(2026-09-01 17:4x): /home +HIT(2026-09-01 17:5x): /home +.HIT(2026-09-01 18:0x): /home +HIT(2026-09-01 18:1x): /home +HIT(2026-09-01 18:2x): /home +HIT(2026-09-01 18:3x): /home +HIT(2026-09-01 18:4x): /home +HIT(2026-09-01 18:5x): /home +.HIT(2026-09-01 19:0x): /home +HIT(2026-09-01 19:1x): /home +HIT(2026-09-01 19:2x): /home +HIT(2026-09-01 19:3x): /home +HIT(2026-09-01 19:4x): /home +HIT(2026-09-01 19:5x): /home +.HIT(2026-09-01 20:0x): /home +HIT(2026-09-01 20:1x): /home +HIT(2026-09-01 20:2x): /home +HIT(2026-09-01 20:3x): /home +HIT(2026-09-01 20:4x): /home +HIT(2026-09-01 20:5x): /home +.HIT(2026-09-01 21:0x): /home +HIT(2026-09-01 21:1x): /home +HIT(2026-09-01 21:2x): /home +HIT(2026-09-01 21:3x): /home +HIT(2026-09-01 21:4x): /home +HIT(2026-09-01 21:5x): /home +.HIT(2026-09-01 22:0x): /home +HIT(2026-09-01 22:1x): /home +HIT(2026-09-01 22:2x): /home +HIT(2026-09-01 22:3x): /home +HIT(2026-09-01 22:4x): /home +HIT(2026-09-01 22:5x): /home +.HIT(2026-09-01 23:0x): /home +HIT(2026-09-01 23:1x): /home +HIT(2026-09-01 23:2x): /home +HIT(2026-09-01 23:3x): /home +HIT(2026-09-01 23:4x): /home +HIT(2026-09-01 23:5x): /home +. day 2026-09-01 done, batches-with-hits=144 + +### REMAINING FIVE DAYS — no control injected +........................ day 2026-08-30 done, batches-with-hits=0 +........................ day 2026-08-29 done, batches-with-hits=0 +........................ day 2026-08-28 done, batches-with-hits=0 +........................ day 2026-08-27 done, batches-with-hits=0 +........................ day 2026-08-26 done, batches-with-hits=0 +.......................### REMAINING DAYS, full 24h each, second granularity +day 2026-08-30: hits=0 +day 2026-08-29: hits=0 +day 2026-08-28: hits=0 +day 2026-08-27: hits=0 +day 2026-08-26: hits=0 +day 2026-08-25: hits=0 +day 2026-08-24: hits=0 +### done diff --git a/documentation/audits/evidence-drill-r95-recovery-2026-09-01/phase1-find-the-snapshot/07-sweep-control.txt b/documentation/audits/evidence-drill-r95-recovery-2026-09-01/phase1-find-the-snapshot/07-sweep-control.txt new file mode 100644 index 00000000..3b82400a --- /dev/null +++ b/documentation/audits/evidence-drill-r95-recovery-2026-09-01/phase1-find-the-snapshot/07-sweep-control.txt @@ -0,0 +1,3 @@ +### CONTROL: exactly the sweep's batch shape (600 fake names) + 1 real path appended. +### If the sweep can see a directory that exists, this MUST print /home. 6 batches tried. +control batches: 6, batches that correctly returned exactly /home: 6 diff --git a/documentation/audits/evidence-drill-r95-recovery-2026-09-01/phase1-find-the-snapshot/08-alt-formats.txt b/documentation/audits/evidence-drill-r95-recovery-2026-09-01/phase1-find-the-snapshot/08-alt-formats.txt new file mode 100644 index 00000000..98bfbde5 --- /dev/null +++ b/documentation/audits/evidence-drill-r95-recovery-2026-09-01/phase1-find-the-snapshot/08-alt-formats.txt @@ -0,0 +1,4 @@ +candidates: 126 +=== HITS === +/.zfs/shares/. +/home diff --git a/documentation/audits/evidence-drill-r95-recovery-2026-09-01/phase1-find-the-snapshot/09-rclone-restic-lead.txt b/documentation/audits/evidence-drill-r95-recovery-2026-09-01/phase1-find-the-snapshot/09-rclone-restic-lead.txt new file mode 100644 index 00000000..218a43db --- /dev/null +++ b/documentation/audits/evidence-drill-r95-recovery-2026-09-01/phase1-find-the-snapshot/09-rclone-restic-lead.txt @@ -0,0 +1,7 @@ +=== is rclone in the controller container? === +NO rclone in container +=== does restic 0.14.0 know the rclone backend? (control: banana:) === +Fatal: parsing repository location failed: invalid backend +If the repository is in a local directory, you need to add a `local:` prefix +--- rclone: backend --- +Fatal: unable to open repository at rclone:nosuchremote:/x: exec: "rclone": executable file not found in $PATH diff --git a/documentation/audits/evidence-drill-r95-recovery-2026-09-01/phase2-the-attack/NOT-RUN.md b/documentation/audits/evidence-drill-r95-recovery-2026-09-01/phase2-the-attack/NOT-RUN.md new file mode 100644 index 00000000..59aab83d --- /dev/null +++ b/documentation/audits/evidence-drill-r95-recovery-2026-09-01/phase2-the-attack/NOT-RUN.md @@ -0,0 +1,13 @@ +# phase2-the-attack — NOT RUN + +The drill stopped at the end of Phase 1 on the operator's ruling. **This directory is empty on +purpose, not by omission.** + +Phase 1 established that no Storage Box snapshot is reachable from the box by any name (R-433), and +that every other route to one is fenced (the Hetzner API, §11-D) or unavailable (a browser and a +main-account credential nobody here holds). With no recovery leg, the deletion in Phase 2 would have +destroyed real off-site history to buy only an alarm test that could not have fired at the size the +runbook specifies — the detector needs a fall of more than half of 69, and one app's tag is ~9 +(R-435). Put to the operator as a two-option decision; the ruling was stop and report. + +See `../README.md` and `felhom.eu/REPORT.md`. diff --git a/documentation/audits/evidence-drill-r95-recovery-2026-09-01/phase3-the-alarm/NOT-RUN.md b/documentation/audits/evidence-drill-r95-recovery-2026-09-01/phase3-the-alarm/NOT-RUN.md new file mode 100644 index 00000000..e0fbe189 --- /dev/null +++ b/documentation/audits/evidence-drill-r95-recovery-2026-09-01/phase3-the-alarm/NOT-RUN.md @@ -0,0 +1,13 @@ +# phase3-the-alarm — NOT RUN + +The drill stopped at the end of Phase 1 on the operator's ruling. **This directory is empty on +purpose, not by omission.** + +Phase 1 established that no Storage Box snapshot is reachable from the box by any name (R-433), and +that every other route to one is fenced (the Hetzner API, §11-D) or unavailable (a browser and a +main-account credential nobody here holds). With no recovery leg, the deletion in Phase 2 would have +destroyed real off-site history to buy only an alarm test that could not have fired at the size the +runbook specifies — the detector needs a fall of more than half of 69, and one app's tag is ~9 +(R-435). Put to the operator as a two-option decision; the ruling was stop and report. + +See `../README.md` and `felhom.eu/REPORT.md`. diff --git a/documentation/audits/evidence-drill-r95-recovery-2026-09-01/phase4-the-recovery/NOT-RUN.md b/documentation/audits/evidence-drill-r95-recovery-2026-09-01/phase4-the-recovery/NOT-RUN.md new file mode 100644 index 00000000..34800998 --- /dev/null +++ b/documentation/audits/evidence-drill-r95-recovery-2026-09-01/phase4-the-recovery/NOT-RUN.md @@ -0,0 +1,13 @@ +# phase4-the-recovery — NOT RUN + +The drill stopped at the end of Phase 1 on the operator's ruling. **This directory is empty on +purpose, not by omission.** + +Phase 1 established that no Storage Box snapshot is reachable from the box by any name (R-433), and +that every other route to one is fenced (the Hetzner API, §11-D) or unavailable (a browser and a +main-account credential nobody here holds). With no recovery leg, the deletion in Phase 2 would have +destroyed real off-site history to buy only an alarm test that could not have fired at the size the +runbook specifies — the detector needs a fall of more than half of 69, and one app's tag is ~9 +(R-435). Put to the operator as a two-option decision; the ruling was stop and report. + +See `../README.md` and `felhom.eu/REPORT.md`. diff --git a/documentation/audits/evidence-drill-r95-recovery-2026-09-01/phase5-rearm/NOT-RUN.md b/documentation/audits/evidence-drill-r95-recovery-2026-09-01/phase5-rearm/NOT-RUN.md new file mode 100644 index 00000000..846bfda0 --- /dev/null +++ b/documentation/audits/evidence-drill-r95-recovery-2026-09-01/phase5-rearm/NOT-RUN.md @@ -0,0 +1,13 @@ +# phase5-rearm — NOT RUN + +The drill stopped at the end of Phase 1 on the operator's ruling. **This directory is empty on +purpose, not by omission.** + +Phase 1 established that no Storage Box snapshot is reachable from the box by any name (R-433), and +that every other route to one is fenced (the Hetzner API, §11-D) or unavailable (a browser and a +main-account credential nobody here holds). With no recovery leg, the deletion in Phase 2 would have +destroyed real off-site history to buy only an alarm test that could not have fired at the size the +runbook specifies — the detector needs a fall of more than half of 69, and one app's tag is ~9 +(R-435). Put to the operator as a two-option decision; the ruling was stop and report. + +See `../README.md` and `felhom.eu/REPORT.md`. diff --git a/documentation/audits/evidence-drill-r95-recovery-2026-09-01/phase6-teardown/01-verify-then-clean.txt b/documentation/audits/evidence-drill-r95-recovery-2026-09-01/phase6-teardown/01-verify-then-clean.txt new file mode 100644 index 00000000..8f686931 --- /dev/null +++ b/documentation/audits/evidence-drill-r95-recovery-2026-09-01/phase6-teardown/01-verify-then-clean.txt @@ -0,0 +1,23 @@ +=== A. off-site inventory UNCHANGED? (must still be 69) === + bookstack +----------------------------------------------------- +69 snapshots + +=== B. account home contents (must be exactly .ssh + felhom-repo, nothing of mine) === +. +.. +.ssh +felhom-repo + +=== C. repo top level (untouched) === +config +data +index +keys +locks +snapshots + +=== D. LAYER 1 CLEAN: ssh control sockets inside the container === +/tmp/r95-cm4-u629488-sub3 +removed +verified: no control sockets left in container diff --git a/documentation/audits/evidence-drill-r95-recovery-2026-09-01/phase6-teardown/02-guest-and-host-clean.txt b/documentation/audits/evidence-drill-r95-recovery-2026-09-01/phase6-teardown/02-guest-and-host-clean.txt new file mode 100644 index 00000000..87e1d08e --- /dev/null +++ b/documentation/audits/evidence-drill-r95-recovery-2026-09-01/phase6-teardown/02-guest-and-host-clean.txt @@ -0,0 +1,22 @@ +=== LAYER 2 (guest 9201) + LAYER 3 (demo-hp PVE host) cleanup === +-- guest scratch before: +. +.. +cands.txt +cmd.sh +env.sh +probe.txt +sweep1.txt +-- guest scratch after: +ls: cannot access '/mnt/r95drill': No such file or directory +-- guest /mnt now: +. +.. +felhom-drives +sys_drive +-- host /tmp/r95cmd.sh after: +ls: cannot access '/tmp/r95cmd.sh': No such file or directory +-- controller container healthy: +felhom-controller gitea.dooplex.hu/admin/felhom-controller:0.232.0 Up 6 hours (healthy) +-- app containers healthy count: +19 diff --git a/documentation/audits/evidence-drill-r95-recovery-2026-09-01/phase6-teardown/03-dooplex-clean-and-final-state.txt b/documentation/audits/evidence-drill-r95-recovery-2026-09-01/phase6-teardown/03-dooplex-clean-and-final-state.txt new file mode 100644 index 00000000..087442f5 --- /dev/null +++ b/documentation/audits/evidence-drill-r95-recovery-2026-09-01/phase6-teardown/03-dooplex-clean-and-final-state.txt @@ -0,0 +1,12 @@ +=== LAYER 4 (DooPlex) — shred the hub DB copy (it holds every host's secret) === +-rw-rw-r-- 1 kisfenyo kisfenyo 236482560 Sep 1 16:45 /tmp/claude-1000/hubdb.db +ls: cannot access '/tmp/claude-1000/hubdb.db*': No such file or directory +verified: hub DB copies gone + +=== demo-felhom NOT TOUCHED — positive check, its own tier still reporting ok === +demo-hp: report 22221 @2026-09-01 14:45:18 snapshot_count=69 stats_known=True last_status=ok last_success=2026-09-01T02:17:57Z +demo-felhom: report 22222 @2026-09-01 14:47:19 snapshot_count=10 stats_known=True last_status=ok last_success=2026-09-01T08:32:13Z + +--- any offsite_snapshots_dropped events raised today? (must be only the constructed one) --- +(3368, 'demo-hp', '2026-09-01 12:29:06', 'offsite_snapshots_dropped') +second copy shredded diff --git a/documentation/backlog/OPEN-ITEMS.md b/documentation/backlog/OPEN-ITEMS.md index f8bffe41..43348c12 100644 --- a/documentation/backlog/OPEN-ITEMS.md +++ b/documentation/backlog/OPEN-ITEMS.md @@ -224,7 +224,7 @@ unconditional promise, which CAMPAIGN-11 §7 step 7 measured the customer-facing | **E-2d** | **Prove E-2 on a fresh VM** — a real `felhom-host-install.sh` 1.22.0 run, Case B naturally, a claimable customer, then add a drive (the offer) and unplug it (`backup_target_absent` end-to-end) | **CLOSED — PARTIALLY PROVEN** (2026-07-29) | — | **C1, C2 proven** (`audits/E2D-fresh-vm-2026-07-29.md`); **C3, C4 proven live** (`audits/SESSION-C-2026-07-29.md`); **C5 FAILED → R-116** — the gate fires and an alarm reaches the hub, but it is the generic event, so the alarm and its recovery cannot be paired. **R-116 is the single named open leg**; per the Session-C runbook §9, decided in advance, a failed claim closes the item as partially proven rather than triggering a re-run. Both audits carry the full record — the `local-lvm` fence, the ISO/PAIRING derivation, the Phase 0 answers, the per-claim observables and the teardown evidence — and are the place to read it, not this cell. **The arc's actual definition of done is R-106 + R-109, R-108 and D5**, none of which this detour touched | CC | | **R-121** | **A BOX's installed agent can sit releases behind the vouched one and nothing notices — the R-120 gate does not cover it.** demo-hp ran agent **0.113.0** while the hub vouched **0.116.0**, through the whole R-116/R-117 arc, and no signal existed on any channel | **READY (S) — NEW 2026-07-30** | — | **Fourth instance of the drift family** (R-111 golden's agent 17 releases behind, R-115 built+deployed but never published, R-120 golden a controller behind — and now installed-vs-vouched on a live box). **Confirmed at source that R-120's gate cannot catch it:** `hub/internal/web/configs.go:1165-1169` compares `goldenVer` against `store.NewestReportedControllerVersion()` — it is a **golden-artifact vs fleet-CONTROLLER** check and says nothing about the agent installed on a box. **`MinAgent` does not cover it either:** it is used to HOLD the controller floor for a box whose agent is too old (`hub/internal/api/handler.go:530-538`, `store.go:1857`) — protective, not an alarm — and demo-hp's 0.113.0 **equalled** `min_agent` 0.113.0, so even a floor comparison was satisfied. **The cost, measured:** R-117's whole subject is the R-113 conjunction, which landed in **0.114.0** — so the designated drill host could not exercise the code under investigation at all, and the R-117 spike had to route every predicate result through an out-of-repo probe built from `main` instead of the installed agent (`audits/SPIKE-r117-bind-liveness-2026-07-30.md` §1, §2.3). Discovered because the R-117 task made bringing the box current an explicit prerequisite. **Fix shape (not implemented):** the hub already receives `AgentVersion` on every host report, and already has semver comparison in Go — the missing piece is a checker comparing reported agent vs the **vouched** agent and surfacing it, operator-tier. Note the honest tension: a box legitimately lags between publish and deploy, so this wants a staleness window rather than an instant alarm | CC | | **R-118** | **An absent drive's union row advertises the ROOT filesystem's capacity as its own.** In the absent-state payload the registry-union row reports `total_bytes: 49675956224 / used_bytes: 4584579072` — **byte-identical to the `local` row** (`durable_id: path:/var/lib/vz`, i.e. `pve-root`) in the same response. The real drive is **4 GB** | **READY (XS) — NEW 2026-07-30** | — | Cause: `statfsCapacity(d.MountPath)` (`disks.go:335-338`) statfs's `/mnt/cel`, which with the device gone is a **bare directory on the root filesystem**. `observe.go:176-183`'s comment warns about exactly this trap and guards the Observe path ("*an unmounted removable dir-storage's mountpoint reverts to a bare directory on root … catastrophic DR mis-id*"); **the union path has no equivalent guard.** **Not a DR mis-id** — `durable_id` on that row is still the correct `uuid:…`, so re-attach identity is safe. It is a **false capacity** reaching every consumer of `total_bytes`/`used_fraction` (fill monitors, storage cards): a detached 4 GB drive advertises 46 GiB at 9.2 % used. Same class as `role.go:180-181` — an absent drive's fields decaying to the root filesystem's. Evidence: `audits/DIAG-r116-disks-payload-2026-07-30.md` §12 | CC | -| **R-95** | restic offsite credential **can delete** (`readonly=False`, `forget --prune` runs from the box); SFTP cannot express append-only | **READY** | — | Root exposure still open. Mitigation now ARMED — split prune off-box or move to REST `--append-only` | CC **SPIKE 2026-09-01 — `audits/SPIKE-r95-offsite-delete-2026-09-01.md`. THE WORD "ARMED" ABOVE IS NOT SUPPORTED AND IS WITHDRAWN PENDING R-429:** no `.snapshots` is visible to either box's sub-account (measured, both machines, with controls), so the seven-day bound is unverified and unverifiable from the product side. **Q3 (documented, `hub/internal/hetznerapi/hetznerapi.go:38-45`): the sub-account API has ONE permission axis, `readonly` — there is no append-only, so the PBS shape does NOT transfer** (PBS is a server that can refuse; a Storage Box is a filesystem that runs nothing). **Q5 (measured): withdrawing delete does NOT wedge the store** — restic treats a dead owner's lock as stale and proceeds — so the constraint everyone feared is not the blocker; but `unlock --remove-all` lies about success (R-430), and the crash-lock window is UNKNOWN. **Q6 (measured): restic 0.14.0 DOES speak `rest:`** (control: `banana:` → invalid backend), and append-only is a rest-server flag, not a restic one — reachable, but it needs a machine in the recovery path and ep0 is protected. **Q7 (measured): detection is nearly free** — `snapshot_count` already reaches the hub and the hub APPENDS reports, so the comparison needs no box change. **RECOMMENDATION: answer R-429 first (Viktor, ten minutes), then build detection, then move retention off the box; defer the transport change.** Two `forget --prune` sites must be disarmed together — `offbox.go:1388` AND `offbox.go:1759` — or R-191 repeats. **RE-SCOPED 2026-09-01 — THE STORY WAS WORSE THAN THE TRUTH FOR TWO MONTHS.** The box can delete its own LIVE repository, **but it cannot write to the daily snapshots of it** — MEASURED, not cited: a write into `/.zfs/snapshot` is refused on both boxes while the same write to the account home succeeds (R-432). Seven daily snapshots are confirmed in the panel (R-429). **So a deletion costs at most the data written since the last daily snapshot, and the rest is recoverable — file by file, one customer at a time, with no effect on anyone else** (vendor: *"You can download individual files or entire directories as usual"*; *"It is not possible to write to the `/.zfs` directory or its subfolder"*). **NOT open-ended loss.** Two caveats kept honest: a panel-driven snapshot restore rolls back the WHOLE Storage Box and deletes newer snapshots, which is why the per-file route matters; and per-file recovery is **operator-only today** (R-432). **DETECTION SHIPPED hub v0.111.0 (R-431)** — an unexplained fall is noticed within a day. **THE RANKING IS VIKTOR'S:** this has been #1 since July on the old story. On the new facts I would rank it below the items that can still lose data outright, but I am not re-ranking it myself. `audits/SPIKE-r95-offsite-delete-2026-09-01.md` | +| **R-95** | restic offsite credential **can delete** (`readonly=False`, `forget --prune` runs from the box); SFTP cannot express append-only | **READY** | — | Root exposure still open. Mitigation now ARMED — split prune off-box or move to REST `--append-only` | CC **SPIKE 2026-09-01 — `audits/SPIKE-r95-offsite-delete-2026-09-01.md`. THE WORD "ARMED" ABOVE IS NOT SUPPORTED AND IS WITHDRAWN PENDING R-429:** no `.snapshots` is visible to either box's sub-account (measured, both machines, with controls), so the seven-day bound is unverified and unverifiable from the product side. **Q3 (documented, `hub/internal/hetznerapi/hetznerapi.go:38-45`): the sub-account API has ONE permission axis, `readonly` — there is no append-only, so the PBS shape does NOT transfer** (PBS is a server that can refuse; a Storage Box is a filesystem that runs nothing). **Q5 (measured): withdrawing delete does NOT wedge the store** — restic treats a dead owner's lock as stale and proceeds — so the constraint everyone feared is not the blocker; but `unlock --remove-all` lies about success (R-430), and the crash-lock window is UNKNOWN. **Q6 (measured): restic 0.14.0 DOES speak `rest:`** (control: `banana:` → invalid backend), and append-only is a rest-server flag, not a restic one — reachable, but it needs a machine in the recovery path and ep0 is protected. **Q7 (measured): detection is nearly free** — `snapshot_count` already reaches the hub and the hub APPENDS reports, so the comparison needs no box change. **RECOMMENDATION: answer R-429 first (Viktor, ten minutes), then build detection, then move retention off the box; defer the transport change.** Two `forget --prune` sites must be disarmed together — `offbox.go:1388` AND `offbox.go:1759` — or R-191 repeats. **RE-SCOPED 2026-09-01 — THE STORY WAS WORSE THAN THE TRUTH FOR TWO MONTHS.** The box can delete its own LIVE repository, **but it cannot write to the daily snapshots of it** — MEASURED, not cited: a write into `/.zfs/snapshot` is refused on both boxes while the same write to the account home succeeds (R-432). Seven daily snapshots are confirmed in the panel (R-429). **So a deletion costs at most the data written since the last daily snapshot, and the rest is recoverable — file by file, one customer at a time, with no effect on anyone else** (vendor: *"You can download individual files or entire directories as usual"*; *"It is not possible to write to the `/.zfs` directory or its subfolder"*). **NOT open-ended loss.** Two caveats kept honest: a panel-driven snapshot restore rolls back the WHOLE Storage Box and deletes newer snapshots, which is why the per-file route matters; and per-file recovery is **operator-only today** (R-432). **DETECTION SHIPPED hub v0.111.0 (R-431)** — an unexplained fall is noticed within a day. **THE RANKING IS VIKTOR'S:** this has been #1 since July on the old story. On the new facts I would rank it below the items that can still lose data outright, but I am not re-ranking it myself. `audits/SPIKE-r95-offsite-delete-2026-09-01.md` **DRILL 2026-09-01, LATER THE SAME DAY — THE RE-SCOPE'S SECOND HALF IS WITHDRAWN.** The drill that was to walk the recovery found there is no route to walk: **no snapshot is reachable from a sub-account by ANY name** (R-433) — 777,600 exact names in the vendor format over nine days, zero hits, with a passing control, plus the structural reason (`/home` st_dev 0,82 vs `/.zfs/snapshot` st_dev 0,276, and `/home/.zfs` absent). **Clause (a) stands: the box can delete the live repo and cannot write into the snapshot area.** Clause (b) — *"the rest is recoverable file by file"* — is NOT SUPPORTED. **The drill was STOPPED before its destructive phase on the operator's ruling**, because with no recovery route the deletion would have destroyed real history to buy only an alarm test that could not fire at the specified size (R-435). Nothing was deleted; the store is verified untouched at 69 snapshots. **So the comfort that lowered this row rested on an unwalked route, and the route does not exist.** Two new leads decide what happens next: **R-433** (can the MAIN account see them? nobody here holds that credential) and **R-436** (`rclone serve restic --stdio` is offered server-side and restic speaks `rclone:` — measured — which could make real prevention cheap, IF the provider pins `--append-only`). **The rank stays Viktor's. Plainly: the argument that moved this row down is the argument the drill removed.** `audits/evidence-drill-r95-recovery-2026-09-01/` | | **R-191** | **Every weekly offsite backup UPLOADS successfully and then FAILS the job on a prune the box is deliberately not allowed to do — on both demo boxes.** Measured on demo-felhom 2026-08-04 06:49–06:53: the upload completed (223 s, 629 MiB of 1.874 GiB, **67.2 % reused incrementally**), then `ERROR: prune 'ct/9201': proxmox-backup-client failed: Error: permission check failed - missing Datastore.Modify\|Datastore.Prune on /datastore/felhom-offsite/demo-felhom` → `ERROR: Backup of VM 9201 failed - error pruning backups` → `TASK ERROR: job errors`. The hub raised `whole_guest_backup_failed` | **CLOSED — SHIPPED 2026-08-04** (installer **1.25.0**; both live boxes corrected) | — | **This is R-89's rule not reaching the config.** R-89 moved PBS pruning SERVER-SIDE — *"boxes set `keep_last: 0`, ep0 runs prune jobs; box tokens stay write-only, never widen the grant"*. The token behaves exactly as designed: it refuses. But **both** demo boxes still arm the offsite tier with `keep_last=2 prune_pbs_allowed=true` (`backup_targets: [{target_id: felhom-pbs, cadence_seconds: 604800, keep_last: 2}]`), so every run asks for a prune that must fail. **The data is SAFE and that is why this is not a P1:** the snapshot lands before the prune is attempted; what is wrong is the job's VERDICT and the weekly operator e-mail it produces. **But it is corrosive in the specific way this project keeps finding:** a backup that reports FAILED while succeeding trains the operator to discount `whole_guest_backup_failed`, which is the same alert that would carry a real one — and it is exactly the failure the R-100 corollary warns about, an alarm whose text is true and whose trigger is not the thing you would act on. **Fix is one config line per box** (`keep_last: 0` on the PBS tier) plus whatever writes it on a fresh install; **deliberately NOT applied in this session** — the session was a runbook with an explicit "change nothing, and if a change appears necessary, stop and report" rule, and a retention field on a live backup tier is not a change to slip into an observation run. **Check before fixing:** whether ep0's prune jobs actually cover these two namespaces, or the snapshots simply accumulate once the box stops asking **THE GATE WAS RUN FIRST, AND IT MATTERED.** Before disabling anything, ep0 was read (read-only, Tier 2): prune jobs `prune-demo-felhom` and `prune-demo-hp` exist on datastore `felhom-offsite`, one per namespace, `schedule 03:30`, `keep-last 2`, comment *"R-82 retention keep-last=2, server-side (box tokens are write-only)"* — and they have run **every day since 2026-07-27: 18 tasks, all `status=OK`**. The newest task log reads `retention options: --ns demo-felhom --max-depth 0 --keep-last 2` / `keep ct/9201/2026-07-27…` / `keep ct/9201/2026-07-28…` / `TASK OK`. Retention happens, and it happens there. **A METHODOLOGICAL WARNING WORTH MORE THAN THE FIX.** Three separate queries said the OPPOSITE — *no prune jobs have ever run* — and **all three were broken instruments**: `worker-type` where the field is `worker_type`; the value `prune` where the worker type is `prunejob`; and `journalctl -u proxmox-backup` where the unit is `proxmox-backup-proxy`. A fourth reading (3 snapshots under keep-last 2) was mis-framed by CC and self-corrected — the third snapshot had landed AFTER that day's 03:30 window. Acting on any of them would have disabled the only pruning ATTEMPT while reporting that nothing prunes: a weekly false alarm traded for unbounded growth on the protected endpoint, invisible for months. **The gate is what caught it, and only because it demanded evidence rather than a verdict.** **Shipped:** installer **1.25.0** writes `keep_last: 0` on the offsite tier (the agent's existing guard `allowPBSPrune = !primary && keep_last > 0` already reads that as *never prune from the box* — no agent change), the justifying paragraph is rewritten to say where retention lives and cite R-89, and `hostinstall_gates.py` asserts it (red-proved: pinning `keep_last: 2` back fails the gate). **Both live boxes corrected in their own config** — `backup tier armed target=felhom-pbs … keep_last=0 … prune_pbs_allowed=false` on demo-felhom and demo-hp, with the local tier untouched at `keep_last=3`. Served over HTTPS at `1.25.0` with `"keep_last":0` in the served bytes. **STILL TO OBSERVE:** the next weekly offsite run completing OK end-to-end. The change removes the failing step; the *schedule* proving it is next week's event, and this row should carry that line when it happens. | CC | | **R-194** | **PVE's permission cache delays every grant-state verdict by an unknown amount, so "the agent can read it" and "the ACL exists" are not the same measurement.** Observed twice while validating R-190's self-repair on demo-felhom 2026-08-04: both ACL rows for `/storage/felhom-backup` were deleted, and `GET /access/permissions` continued to report `Datastore.AllocateSpace` present — for **~40 s** in one run and **~16 minutes** in another. During that window the capability probe reads healthy and the self-repair does not fire | **OPEN** | — | **Why it matters beyond the delay:** it puts a floor under how fast a lost grant can be noticed, it makes any single permission read a lagging indicator, and — the interesting part — **it is a candidate contributor to R-190's own timeline**: a grant removed at an unknown moment could keep working until a cache expiry, which is exactly the shape of *worked at 04:44, refused at 09:24*. That does not explain what removed it, but it may explain when the refusal SURFACED, and the two have been treated as the same instant. **Not a defect in our code** — it is PVE behaviour, and the mitigation already tolerates it (the repair fires on the next probe after the cache clears). **What is worth deciding:** whether the store-grant probe should ALSO consult the storage content listing as a second signal, since that appeared to reflect the loss immediately (`{"data":[]}` while the permission read still said present) — two signals disagreeing is itself information, and today only one of them is read | CC | | **R-200** | **The DR password-injection seam has a handler, a route and tests — and no form.** `POST /backup/offbox/inject-password` is routed (`controller/internal/web/server.go:510`) to `offboxInjectPasswordHandler` (`offbox_handlers.go:174-196`) → `InjectOffboxPassword` (`backup/offbox.go:541`). **No template in the repository contains that path or any form posting to it** (grep over `internal/web/templates/`: one unrelated hit, an XSS comment) | **PLUMBING COMPLETE** (controller **v0.196.0**); **the FORM is not built — still open** | — | **The tenth instance of this project's built-but-never-wired class, and the exact shape `CLAUDE.md` and `felhom.eu/CLAUDE.md`'s seam-wiring rule were written for:** handler tests that POST directly (`offbox_escrow_test.go:167,180`) prove nothing about reachability. To use the only implemented recovery seam today, a person must hand-craft an authenticated POST with a session cookie and CSRF token. **Note the layering while fixing it:** this form takes a **64-hex repo password**, not a recovery code (`offboxRepoPwPattern`, `offbox.go:543`) — they are different secrets at different layers, and the operator's 2026-08-04 ruling asks for a form that takes **R**. Build the R form and treat this one as the operator/DR fallback it was written as, but **ship it with a render test per branch of whatever gate it sits behind**. Source: `audits/RECON-offsite-dr-chain-2026-08-04.md` §3 link 9 **THE DIAGNOSTIC HALF IS DONE AND IT ANSWERED THE QUESTION.** `--recover-offsite-check` is a `docker exec` escape hatch in the shape of `--print-reset-code`: R on **STDIN** (never argv, never `ps`, never shell history, never a transcript), fetch+unseal via the agent, and a verdict of **two sha256 hashes**. **It compares and never installs** — the recovered password is not written to `offbox/repo_password`; a test asserts the data dir is byte-unchanged and its red-proof (adding the install call) fails it. Confirmed live: `repo_password` mtime still `2026-08-03 07:18:02` after the successful check at `2026-08-04 11:49`. **Exit codes are load-bearing** — `0` match, `2` a clean MISMATCH, `1` a step failed; "it failed" and "it worked and disagreed" must never share a status because only one is a finding about the system. A box with no local password reports distinctly (the rebuilt-box shape, where the next step is to INSTALL rather than compare). **WHAT IS NOT BUILT, deliberately:** no card, no form, no preview, no customer-facing text — building an interface on top of a chain nobody had walked is how the preceding three weeks went wrong. **What remains for this row:** link 9 (the recovered password placed so `WriteOffboxSecrets` keeps it) and the customer-facing shape the operator ruled on 2026-08-04 (yell → R form → preview → proceed), which is now priced against a chain that exists rather than one that is assumed **PART 0 SHIPPED 2026-08-04 (v0.196.0):** `--recover-offsite-install` is the sibling of the check — same fetch/unseal path, same STDIN discipline for R — and it **places** the recovered password via `InjectOffboxPassword`. The confirmation is a SECOND invocation (`--confirm-install`): without it, both hashes print and nothing is written, so the operator sees the comparison before a write is possible. Three outcomes named distinctly: **installed** (no local password — the rebuilt-box shape), **unchanged** (identical key, nothing written), **refused** (a DIFFERENT key present — installing would clobber the key the current repository is encrypted under; exit 2, no force offered). It re-reads the file after writing rather than trusting the call. **Red-proof observed:** removing the confirmation gate makes the dry run write the password. The R-persistence test carries a **positive control** (a planted copy found, then removed and not found). **NOT YET EXERCISED AGAINST A LIVE RECOVERY** — the R-201 drill halted before step 9, so this is unit-proven only. **What remains for this row:** the customer-facing shape the operator ruled on 2026-08-04 (yell → recovery-code form → preview → proceed) | CC | @@ -603,7 +603,11 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server` | **R-427** | **`closed_register_gate.py` checks ONE direction only: an open word in a CLOSED row. The mirror — a CLOSED verdict on a row still sitting in `OPEN-ITEMS.md` — is unchecked, and there are TWELVE.** MEASURED 2026-09-01 during the decoy sweep, by reading the leading verdict of every open row with the gate's own predicate: **R-385, R-387, R-341, R-378, R-405, R-88a, R-88b** read unambiguously closed; **R-123, R-190, R-352** read `PARTLY CLOSED` / `MITIGATION SHIPPED` and almost certainly belong where they are. **The rows were NOT moved by this session** — telling a finished row from a partly-finished one is a judgement, and R-378 is itself the record of what happens when a machine makes that judgement on a substring (six still-open rows moved out of the register). **This is R-405's finding mirrored:** that row exists because R-87 sat in the CLOSED file while its state read READY, and the gate written for it looks only the way it was bitten. Fix: the same leading-verdict predicate applied to `OPEN-ITEMS.md`, reporting rather than convicting until the twelve are adjudicated by a person — a gate registered while twelve rows fail it would refuse every push. | **OPEN — 12 rows named; the adjudication is Viktor's, the gate is mine** | | **R-428** | **The decoy-coverage gate — written to catch instruments that match a NAME instead of a fact — identified a repository by its DIRECTORY NAME.** MEASURED on its own first CI run (felhom.eu job 490, 2026-09-01): `os.path.basename(root)` looked up in a `RUNNERS` map, and Gitea's act-runner checks the repo out into a directory called `hostexecutor`, so the gate reported *"unknown repo 'hostexecutor'"* and went INCONCLUSIVE. **The gate that hunts label-matching was matching a label, in the first ten lines of its own main loop, and it shipped that way.** FIXED the same day: it now identifies a repo by which registered runner FILE exists under the root, which is a fact. **Recorded rather than quietly patched because it is the strongest evidence in the sweep that this class is not a matter of carelessness** — it was written by a session that had spent the morning reading 29 gates for exactly this, with the four shapes on screen. Verified under a renamed directory before and after. | **CLOSED 2026-09-01 — fixed, and kept as the class's best example** | | **R-430** | **`restic unlock --remove-all` printed `successfully removed locks` while the lock was still there.** MEASURED 2026-09-01 in a throwaway local repo (no live store touched), under a faithful append-only model — a sticky locks directory owned by root holding a root-owned lock, restic run as `nobody`; both controls passed first (create allowed, delete refused). The command reported success, returned, and `ls` showed the lock present. **`resticStep`'s crash-lock self-heal is built directly on this call** (`felhom-controller/controller/internal/backup/offbox.go:~768`), and its licence to escalate rests on the escalation actually working. **A self-heal that cannot fail is a self-heal that cannot be trusted** — this is this project's *"exit codes that lie"* class, in the one path that runs unattended against the customer's off-site history. It is harmless TODAY because the credential can delete and the removal really happens; it becomes load-bearing the moment delete is withdrawn, which is what R-95 is about. **Not yet established:** whether restic reports success because it removed zero locks by design, or because it did not check. Settling it: read restic 0.14.0's unlock source, or re-run with `--verbose`. | **OPEN — precondition on any R-95 build** | -| **R-432** | **A customer's own sub-account can REACH the snapshot door and is REFUSED writes to it — but sees it EMPTY, so per-file recovery is not product-reachable.** MEASURED 2026-09-01 on BOTH live boxes, over the credential each already holds, with a positive and a negative control in the same run. **What is now PROVEN rather than cited:** `/.zfs` lists (`shares`, `snapshot`) from inside the jail; a write into `/.zfs/snapshot` is **REFUSED** — `dest open …: Failure` — while the identical write to the account home **succeeds** and was cleaned up. **That is the append-only property, measured, and it is the sentence the whole R-95 re-scope rests on.** **What is NOT available:** `/.zfs/snapshot` lists **empty** (link count 2) on both boxes, while the same Storage Box demonstrably holds seven snapshots — `storage-box-pool-1` IS `u629488` (`RUNBOOK-ep0-datastore-volume-2026-07-27.md:386`), the box these sub-accounts live on. So the contents are filtered from a sub-account. **CONSEQUENCE: recovery from a snapshot is an OPERATOR act in a browser, not something the product can drive** — which decides whether R-95's remedy can ever be customer-facing. **Cheapest next step, and it is Viktor's:** read one snapshot's name from the panel; a single `ls /.zfs/snapshot/` from a box then settles whether a named snapshot can be entered even though the directory does not list (ZFS allows exactly that). If it can, per-file recovery becomes product-reachable and this closes cheaply. | **OPEN — one panel read settles it** | +| **R-432** | **A customer's own sub-account can REACH the snapshot door and is REFUSED writes to it — but sees it EMPTY, so per-file recovery is not product-reachable.** MEASURED 2026-09-01 on BOTH live boxes, over the credential each already holds, with a positive and a negative control in the same run. **What is now PROVEN rather than cited:** `/.zfs` lists (`shares`, `snapshot`) from inside the jail; a write into `/.zfs/snapshot` is **REFUSED** — `dest open …: Failure` — while the identical write to the account home **succeeds** and was cleaned up. **That is the append-only property, measured, and it is the sentence the whole R-95 re-scope rests on.** **What is NOT available:** `/.zfs/snapshot` lists **empty** (link count 2) on both boxes, while the same Storage Box demonstrably holds seven snapshots — `storage-box-pool-1` IS `u629488` (`RUNBOOK-ep0-datastore-volume-2026-07-27.md:386`), the box these sub-accounts live on. So the contents are filtered from a sub-account. **CONSEQUENCE: recovery from a snapshot is an OPERATOR act in a browser, not something the product can drive** — which decides whether R-95's remedy can ever be customer-facing. **Cheapest next step, and it is Viktor's:** read one snapshot's name from the panel; a single `ls /.zfs/snapshot/` from a box then settles whether a named snapshot can be entered even though the directory does not list (ZFS allows exactly that). If it can, per-file recovery becomes product-reachable and this closes cheaply. **ANSWERED 2026-09-01 (DRILL, `audits/evidence-drill-r95-recovery-2026-09-01/`) — NO, AND THE PANEL READ IS NOT NEEDED.** The named-entry hypothesis was tested exhaustively and fails: **777,600 exact names** in the vendor-documented format `YYYY-MM-DDTHH-MM-SS` (nine full days, second granularity) plus 126 alternative shapes, **zero hits**, with a control proving the identical batch shape returns a path that does exist (6/6). **And there is a structural reason:** `df` reports `u629488-sub3` mounted on `/home` at **st_dev 0,82** while `/.zfs/snapshot` is **st_dev 0,276** — a different filesystem — and `/home/.zfs` does not exist. A ZFS snapshot under `/.zfs/snapshot` belongs to the dataset owning that `.zfs`, not to the child at `/home`, so a correctly-named snapshot there **could not contain `felhom-repo`**. Three tools agree with controls in the same run (SFTP, the port-23 shell, `rsync --list-only`). The empty listing is not a display toggle hiding a reachable tree — from a sub-account there is no tree. **CONSEQUENCE: per-file recovery is not "operator-only", it is unreachable from the box entirely** → R-433. | **ANSWERED 2026-09-01 — negatively; the panel-read next step is WITHDRAWN as unnecessary** | +| **R-433** | **A sub-account cannot reach ANY Storage Box snapshot, by any name — so clause (b) of the 2026-09-01 R-95 re-scope ("the rest is recoverable file by file") is NOT SUPPORTED.** MEASURED 2026-09-01 on `demo-hp` over the credential the box already holds, read-only, no delete verb issued. **The sweep:** a batched `stat -c %n` over Hetzner's port-23 restricted shell (500–600 paths per round trip, stdout carrying only paths that exist) tried **777,600** names of the vendor form `YYYY-MM-DDTHH-MM-SS` across nine full days at second granularity, and 126 alternative shapes — **zero resolved.** **The control is what makes the zero mean anything:** the identical 600-name batch with one real path appended returned it in 6 of 6 batches. **The structural cause:** `/home` (the customer data, `u629488-sub3`) is **st_dev 0,82**; `/.zfs/snapshot` is **st_dev 0,276**; `/home/.zfs` does not exist. A snapshot under `/.zfs/snapshot` belongs to a different dataset than the one holding `felhom-repo`. **What still stands:** clause (a) — the box can delete its live repository but cannot WRITE into `/.zfs/snapshot` — is unchanged and re-confirmed. **What is now open again:** the only routes to the older copy are a panel rollback of the WHOLE Storage Box (deletes newer snapshots, hits every customer on it) and the provider API (fenced by §11-D, and `hub/internal/hetznerapi/hetznerapi.go` has **no snapshot method at all**, so it needs new code regardless). **NOT ESTABLISHED, and it is the question that decides whether per-file recovery exists for anyone:** whether the MAIN account can see the snapshots. No main-account credential exists in this project. **The ranking is Viktor's; I am not re-ranking R-95 — but the argument that moved it down is the argument this drill removed.** `audits/evidence-drill-r95-recovery-2026-09-01/` | **OPEN — decides R-95's remedy and its rank** | +| **R-434** | **The snapshot-drop alarm promises a recovery that cannot be performed.** `hub/internal/monitor/offsite.go` `emitSnapshotDrop` ships this text, live in hub **v0.111.0**: *"The daily Storage Box snapshots are read-only and still hold the older copy, **so this is recoverable file-by-file**; it is NOT confirmed data loss."* The first clause is true. The second is not reachable: not by the product (R-433), and not by the operator without a browser and a main-account credential that does not exist here. Its own comment states the intent — *"THE MESSAGE MUST NOT SAY THE DATA IS LOST, because after the 2026-09-01 measurement that is usually false"* — and the measurement it rests on was superseded the same day. **This is this project's own corollary landing on the alarm shipped that morning:** when a verdict changes which fact it counts from, the alarm text has to change with it, or the operator acts on a promise nobody can keep. **Fix is text-only and must not be made before R-433 settles what IS true** — an alarm rewritten twice in a week is worse than one rewritten once. | **OPEN — text-only, blocked on R-433** | +| **R-435** | **The snapshot-drop detector is blind to a single-app deletion — the exact shape the box can produce.** `snapshotDropFraction = 0.5` and `snapshotDropFloor = 5` (`hub/internal/monitor/offsite.go`) require a fall of MORE than half the previous count. demo-hp's baseline is **69** across 9 apps, so ~35 snapshots must go before it speaks; **one app's tag is ~9 and is invisible.** `offbox.go:1388` runs `forget --prune` **grouped by host,tags** — a per-tag wipe is precisely the shape a faulty retention or a targeted deletion produces. **This is deliberate, not accidental:** the constant's own comment argues the insensitivity, and *"a detector that cries wolf is switched off within a fortnight"* is a lesson this project paid for. **So this row is NOT a demand to lower the threshold.** It is a demand that the blind spot be written where the operator reads it, because "an unexplained fall is noticed within a day" (R-431, and `STATUS.md`) is true only of falls above half. **Discovered by arithmetic while planning the drill's Phase 2/3 pairing, which could not have worked:** Phase 2 deletes one app and Phase 3 expects the alarm to fire. | **OPEN — documentation, not a threshold change** | +| **R-436** | **LEAD, NOT A DEFECT — append-only may be reachable without a new machine, which would make R-95's real prevention far cheaper than the spike concluded.** Hetzner's port-23 restricted shell advertises, in its own `help`, these server-side backends: `borg`, `rsync`, `scp`, `sftp`, **`rclone serve restic --stdio`**. And restic 0.14.0 **recognises the `rclone:` backend** — MEASURED 2026-09-01 with a control: `banana:` → `Fatal: parsing repository location failed: invalid backend`, while `rclone:` → `exec: "rclone": executable file not found in $PATH` (i.e. the backend parsed and it tried to run the helper). rclone is **not** in the controller image today. `rclone serve restic` carries an **`--append-only`** flag. **Why this matters:** the R-95 spike's option 3 was priced at a new always-on service in the recovery path plus either a mount in the hot path or migrating every customer's history — and it was deferred on that price. This route needs neither: the server side already runs at the provider. **THE CAVEAT, STATED FIRST because it may kill the idea:** the **client** supplies the server command line, so a compromised guest could simply omit `--append-only` unless the provider pins it. **NOT ESTABLISHED:** whether Hetzner pins the flag or accepts client-supplied arguments. **That is a vendor question and it is cheap** — it should be asked before any code is written, because if the answer is "client-supplied" this lead is worth nothing. | **OPEN — ask the vendor before building anything** |