diff --git a/REPORT-r331-backup-card.md b/REPORT-r331-backup-card.md new file mode 100644 index 00000000..ad0d01bb --- /dev/null +++ b/REPORT-r331-backup-card.md @@ -0,0 +1,187 @@ +# REPORT — R-331: the operator Backup card said every customer had no backups + +**Hub v0.109.0 (with controller v0.225.0) · 2026-08-30** + +--- + +## 1. What was wrong + +The hub customer page's **Backup** card read, for **every customer, indefinitely**: + +``` +Enabled Yes Snapshots 0 +Repo Size 0 MB Integrity Unknown +``` + +Measured on `demo-hp` 2026-08-30, at which moment the truth was: + +| source | value | +|---|---| +| the box's own `settings.json` | `snapshot_count: 67, repo_size_bytes: 140829678, stats_known: true` | +| that night's controller log | `[offbox] backup OK: 8 app(s) backed up, 67 snapshot(s), 2m14s` | +| **this hub's own Offsite page** | `0.1 GB` used of a `50 GB` quota — read from the same stored report | + +**A card that reads "no backups" over a working backup is worse than no card.** It is the R-88 +direction of failure — degrading to *no backup* rather than to *unknown* — on the one screen an +operator consults to answer "is this customer protected?". + +## 2. Root cause + +The card rendered the report's **`backup`** object. Its `snapshot_count`, `repo_size_mb` and +`integrity_ok` fields have had **no producer** since disk-tier restic moved to the host agent (slice +8C) — the controller's `buildBackupReport` leaves them zero *deliberately* and says so in a comment. +The zeros were correct values for dead fields, rendered as if live. + +**The data was never missing.** The live numbers ride in the report's **`offsite`** object, which this +package **already** reads for the Offsite page (`offsiteUsageBytes`) and which `monitor.OffsiteChecker` +**already** drives fill and staleness alarms from. That the Offsite page rendered demo-hp's real usage +from the same stored report, at the same moment the Backup card said `0 MB`, is the proof the bytes +were arriving. This is a **render fix over an existing feed**, not a new pipeline. + +## 3. Why it was not a one-line template swap + +`snapshot_count: 0` means two opposite things — *this repository holds nothing* and *nobody has ever +measured this repository*. **R-225 measured that confusion one layer down**: a rebuilt box rendered +„Tarolo meret · 0 pillanatkep" over a store that really held snapshot `f3d9cd67`, and the controller's +`StatsKnown` fixed it there. It was never on the wire, so rendering the count without it would have +**moved R-225 up to the hub instead of fixing anything**. Controller v0.225.0 now forwards +`stats_known`. + +## 4. What changed + +`hub/internal/web/backup_card.go` builds a typed `backupCardView` — resolved in Go, because the card's +whole subject is a distinction a template `{{if}}` chain over `map[string]interface{}` float64s cannot +keep: + +| report state | card shows | +|---|---| +| no `offsite` object at all | "No off-site data reported" — **and says explicitly this is not the same as "no backups"** | +| `enabled:false` + declared `state` | the blocker by name (`needs_credential`) — a different operator action from "not enabled" | +| enabled, `stats_known:false` | **—**, plus "never been measured". Never `0` | +| enabled, `stats_known:true` | the real count and size, **including a real `0`** — measured empty is knowledge | + +**A pre-v0.225.0 controller sends no `stats_known`, which unmarshals to false → "unknown".** That is +the fail-safe direction: upgrading the hub ahead of the fleet must not tell the operator that every +un-upgraded customer has zero backups. Pinned by a test. + +**The Integrity row is deleted, not re-sourced.** Nothing produces it: the controller runs no integrity +check, and `NotifyIntegrityOK` / `NotifyIntegrityFailed` exist and are **called from nowhere**. A row +that can only ever read "Unknown" is not information, and one that could read "OK" from an unwritten +field would be a lie. + +`fmtBytesAuto` is new rather than reusing `fmtBytesGB`: that one is fixed at GB because it renders +against GB quotas, and it turns demo-hp's real 140 829 678 bytes into `0.1 GB` — which on a card whose +entire defect was under-reporting a real backup reads as "nearly nothing". + +## 5. Tests and the red-proof + +`r331_backup_card_test.go` asserts the **rendered page**, using demo-hp's real reported values, so a +regression fails against the same numbers the defect was measured against. **The defect lived in the +template's choice of source object, so a test one layer below it would have been green against the +shipped bug** — which is why these drive `handleCustomerUnified` and grep the HTML. + +**RED-PROOF (run 2026-08-30):** restoring the pre-fix card markup fails all four tests — +`the rendered Backup card does not contain demo-hp's real snapshot count (67)`, +`the card does not carry the real repository size (134.3 MB ...)`, +`the card still shows an Integrity row`, plus every branch of the three-way ruling. Restored +immediately; `git diff` clean. + +**Green gate:** `go build ./... && go vet ./... && go test ./...` in `hub/` — 18 packages, rc 0. + +## 6. Deployment and live verification + +Hub **0.109.0** built, pushed, manifest bumped, ArgoCD hard-refreshed and synced. Controller +**0.225.0** deployed to both demo boxes. Verified from the live objects, not from a rollout message +(an ArgoCD "rolled out" can name the old image): + +``` +argocd: sync=Synced rev=36f86300209e9ce914f2a97dac16ee9eeea249e2 (== HEAD) +deploy: gitea.dooplex.hu/admin/felhom-hub:0.109.0 +pod: hub-6795879c4b-pf4lz ...felhom-hub:0.109.0 Running +boxes: ...felhom-controller:0.225.0 Up (healthy) on demo-felhom AND demo-hp +``` + +**The card, fetched from the live hub** (endpoint-level: the exact URL the operator's browser +requests; the residual is client-side rendering only — there is no browser on DooPlex): + +| | demo-hp | demo-felhom | +|---|---|---| +| Off-site snapshots | **67** | **10** | +| Repo size | **134.3 MB** | **132.5 KB** | +| Last successful run | 15h ago | 15h ago | +| Soft quota | 50 GB | 50 GB | +| Integrity row | **absent** (grep count 0) | **absent** (grep count 0) | + +Both read `Snapshots 0 · Repo Size 0 MB · Integrity Unknown` before this change. + +**Cross-checked against the source, not just against itself** — the numbers on the card are the +numbers on the boxes: + +``` +demo-hp settings.json offbox: snapshot_count 67, repo_size_bytes 140829678, stats_known true +demo-felhom settings.json offbox: snapshot_count 10, repo_size_bytes 135635, stats_known true + 135635 / 1024 = 132.5 KB → matches the rendered value +``` + +**One honest detail worth keeping:** demo-felhom's "Last DB dump" reads `—`. That is correct, not a +regression — its only app (`opengist`) has no database, so the box has never taken a DB dump. + +**Not verified live: the "never measured" branch.** Both boxes report `stats_known: true`, so the +degradation path could not be exercised on real hardware without falsifying a box's state. It is +covered by `TestBackupCard_ThreeWayRuling` and `TestBackupCard_OldControllerDegradesToUnknownNotEmpty` +at render level, and this is stated rather than implied. + +## 7. The push bypassed a gate, deliberately, and here is the declaration + +**`git push --no-verify` was used for this change.** `repo_gates.py`'s `golden-currency` gate was +CONVICTED and it was RIGHT: controller **v0.224.0** and **v0.225.0** are released and the newest golden +bake carries **0.223.0**, so a machine installed right now receives neither fix. + +**This is a BYPASS, not a waiver.** The gate offers a waiver only for a release that *deliberately needs +no golden*; these need one. **The operator was asked and ruled bypass-now-bake-later**, on the stated +ground that neither fix bites a day-0 box — R-330 is a nightly false alarm about apps a new box has not +installed yet, R-331 is a hub-side display over backups a new box has not taken yet — and both arrive by +self-update afterwards. That ground is recorded on R-242 precisely because it is the thing to re-check: +**it does not extend to a release that changes first-boot behaviour.** + +**A golden carrying 0.225.0 is OWED** (`RUNBOOK-manual-build.md` §4.1; the vouch is a three-field change, +`MinAgent 0.129.0`). This is the **fourth** bypass of this gate, and the gap it names is now two releases +wide rather than one. + +**The other failing gate was fixed, not bypassed.** `due-checks` was red on R-341's `+7 d` measurement, +five days overdue. It was **taken** during this session — see §8. + +## 8. R-341's overdue check was taken, and its premise did not survive + +Unrelated to R-331; it blocked the same push, so it was done rather than deferred. Evidence: +`documentation/audits/evidence-r341-plus7d-2026-08-30/step1-fd-and-sockets.txt`. + +**Precondition passed**, which is what makes the reading interpretable: `proxmox-backup-proxy` still +`MainPID 551655`, `ps -o lstart=` still `2026-08-18 09:51:04`, `NRestarts=0` — the same proxy generation +as t0, so nothing restarted and re-based the count. (The anchor is `ps`, not `ActiveEnterTimestamp`, +which reads 03:54:54Z here — R-346's trap, avoided.) + +**Result: fd = 17.** Not 17 more — seventeen total, exactly the documented baseline, against **405** at +the first check on 2026-08-20. The socket histogram holds **one LISTEN and nothing else**: ESTAB 0, +CLOSE-WAIT 0. + +**The verdict is "unanswerable", not "the upgrade fixed it".** R-341 asks whether the PBS 4.2.5-1 +upgrade changed the fd slope. Inside this interval **we removed the leak ourselves** (R-344, agent +0.130.0, now live on both boxes). A slope of ~0 measures our fix, not the upgrade — and reading it the +other way would credit a changelog that was read in advance and found to contain no such mechanism. The +perturbation pre-registered for this window was Phase C at ~3%; the actual perturbation was the removal +of the entire phenomenon. The question is now **moot**, and the row is closed as such. + +**What it does establish, which is worth more than the original question:** twelve days after the R-344 +fix, on the same proxy generation with no restart to hide behind, ep0 sits at baseline with zero +established connections. The 388-descriptor accumulation has not returned, and R-336's ~323-day runway +concern retires with it. + +## 9. Not done, and why + +- **No staleness verdict on the card.** `monitor.OffsiteChecker` already owns that and alarms on it. A + second verdict over the same data is two things that can disagree — a shape this codebase has already + paid for (`LastRun` vs `LastSuccess`, R-100). +- **The dead `backup` fields were not removed from the controller's wire format.** Removing them would + stop historical reports already in this hub's store from parsing, for no gain — nothing renders them + now, and a controller-side test fails if anything starts producing them. diff --git a/REPORT.md b/REPORT.md index ad0d01bb..a9755c33 100644 --- a/REPORT.md +++ b/REPORT.md @@ -1,187 +1,166 @@ -# REPORT — R-331: the operator Backup card said every customer had no backups +# REPORT — the deletion we said is survivable: the recovery route does not exist (R-95 drill, 2026-09-01) -**Hub v0.109.0 (with controller v0.225.0) · 2026-08-30** +**RUNBOOK, destructive class, `demo-hp` only. STOPPED at the end of Phase 1 on the operator's ruling, +before any destructive step. No delete verb was issued against any live store; no byte on either +Storage Box sub-account was written, moved or removed.** No production code, no version bump, no +image, no golden. Evidence: `documentation/audits/evidence-drill-r95-recovery-2026-09-01/`. + +| # | phase | verdict | one sentence | +|---|---|---|---| +| 1 | snapshot reachable, and its name | **NO — and it has no reachable name** | 777,600 exact names across nine days in the vendor's own format, plus 126 alternative shapes; zero resolve, with a control proving the sweep detects a path that exists. | +| 2 | the deletion | **NOT RUN — operator ruling** | With no recovery route, the deletion would have destroyed real history to buy nothing; put as a two-option decision, the ruling was stop. | +| 3 | the alarm fired | **NOT RUN — and it could not have fired at the specified size** | The shipped threshold needs a fall of more than half; one app's tag is ~9 of 69. → **R-435** | +| 4 | **the recovery** | **NOT RUN — no route exists that is not fenced** | Box-side: proven impossible. Panel: fenced and browserless. Hetzner API: fenced (§11-D), and the hub's client has no snapshot method at all. | +| — | **RTO from T₀** | **STILL BLANK** | Row 10's RTO cell is unchanged and remains a finding. | +| — | **data lost, quantified** | **NOT MEASURABLE THIS WAY** | The quantity only has meaning if the rest is recoverable, and the route that would recover it is not reachable. | +| 5 | re-arm | **NOT RUN** | Depended on Phase 4. | +| 6 | teardown | **PASS** | Store untouched at 69 snapshots; all three scratch layers removed; hub DB copies shredded; both boxes healthy. | --- -## 1. What was wrong +## 1. Did the recovery work — and does yesterday's re-scope survive? -The hub customer page's **Backup** card read, for **every customer, indefinitely**: +**The recovery was never reachable, and the re-scope does not survive intact. Its first half stands; +its second half does not.** -``` -Enabled Yes Snapshots 0 -Repo Size 0 MB Integrity Unknown -``` +Yesterday's re-scope has two clauses. They must now be separated: -Measured on `demo-hp` 2026-08-30, at which moment the truth was: +* **(a) "The box can delete its live repository, but cannot write to the daily snapshots of it."** + **STANDS.** Re-confirmed here: `/.zfs/snapshot` is reachable and the write-refusal measurement is + unchanged. Nothing in this drill weakens it. +* **(b) "…so the rest is recoverable — file by file, one customer at a time."** **NOT SUPPORTED.** + A snapshot that cannot be opened cannot be copied out of. R-432 recorded the directory listing + empty and named the cheapest next step: *"a single `ls /.zfs/snapshot/` from a box then + settles whether a named snapshot can be entered even though the directory does not list (ZFS + allows exactly that)."* **That step is now done, exhaustively, and the answer is no.** -| source | value | -|---|---| -| the box's own `settings.json` | `snapshot_count: 67, repo_size_bytes: 140829678, stats_known: true` | -| that night's controller log | `[offbox] backup OK: 8 app(s) backed up, 67 snapshot(s), 2m14s` | -| **this hub's own Offsite page** | `0.1 GB` used of a `50 GB` quota — read from the same stored report | +**What was measured.** The port-23 restricted shell accepts a batched `stat`, which makes a cheap +existence oracle: 500–600 paths per round trip, stdout carrying only paths that exist. -**A card that reads "no backups" over a working backup is worse than no card.** It is the R-88 -direction of failure — degrading to *no backup* rather than to *unknown* — on the one screen an -operator consults to answer "is this customer protected?". - -## 2. Root cause - -The card rendered the report's **`backup`** object. Its `snapshot_count`, `repo_size_mb` and -`integrity_ok` fields have had **no producer** since disk-tier restic moved to the host agent (slice -8C) — the controller's `buildBackupReport` leaves them zero *deliberately* and says so in a comment. -The zeros were correct values for dead fields, rendered as if live. - -**The data was never missing.** The live numbers ride in the report's **`offsite`** object, which this -package **already** reads for the Offsite page (`offsiteUsageBytes`) and which `monitor.OffsiteChecker` -**already** drives fill and staleness alarms from. That the Offsite page rendered demo-hp's real usage -from the same stored report, at the same moment the Backup card said `0 MB`, is the proof the bytes -were arriving. This is a **render fix over an existing feed**, not a new pipeline. - -## 3. Why it was not a one-line template swap - -`snapshot_count: 0` means two opposite things — *this repository holds nothing* and *nobody has ever -measured this repository*. **R-225 measured that confusion one layer down**: a rebuilt box rendered -„Tarolo meret · 0 pillanatkep" over a store that really held snapshot `f3d9cd67`, and the controller's -`StatsKnown` fixed it there. It was never on the wire, so rendering the count without it would have -**moved R-225 up to the hub instead of fixing anything**. Controller v0.225.0 now forwards -`stats_known`. - -## 4. What changed - -`hub/internal/web/backup_card.go` builds a typed `backupCardView` — resolved in Go, because the card's -whole subject is a distinction a template `{{if}}` chain over `map[string]interface{}` float64s cannot -keep: - -| report state | card shows | -|---|---| -| no `offsite` object at all | "No off-site data reported" — **and says explicitly this is not the same as "no backups"** | -| `enabled:false` + declared `state` | the blocker by name (`needs_credential`) — a different operator action from "not enabled" | -| enabled, `stats_known:false` | **—**, plus "never been measured". Never `0` | -| enabled, `stats_known:true` | the real count and size, **including a real `0`** — measured empty is knowledge | - -**A pre-v0.225.0 controller sends no `stats_known`, which unmarshals to false → "unknown".** That is -the fail-safe direction: upgrading the hub ahead of the fleet must not tell the operator that every -un-upgraded customer has zero backups. Pinned by a test. - -**The Integrity row is deleted, not re-sourced.** Nothing produces it: the controller runs no integrity -check, and `NotifyIntegrityOK` / `NotifyIntegrityFailed` exist and are **called from nowhere**. A row -that can only ever read "Unknown" is not information, and one that could read "OK" from an unwritten -field would be a lie. - -`fmtBytesAuto` is new rather than reusing `fmtBytesGB`: that one is fixed at GB because it renders -against GB quotas, and it turns demo-hp's real 140 829 678 bytes into `0.1 GB` — which on a card whose -entire defect was under-reporting a real backup reads as "nearly nothing". - -## 5. Tests and the red-proof - -`r331_backup_card_test.go` asserts the **rendered page**, using demo-hp's real reported values, so a -regression fails against the same numbers the defect was measured against. **The defect lived in the -template's choice of source object, so a test one layer below it would have been green against the -shipped bug** — which is why these drive `handleCustomerUnified` and grep the HTML. - -**RED-PROOF (run 2026-08-30):** restoring the pre-fix card markup fails all four tests — -`the rendered Backup card does not contain demo-hp's real snapshot count (67)`, -`the card does not carry the real repository size (134.3 MB ...)`, -`the card still shows an Integrity row`, plus every branch of the three-way ruling. Restored -immediately; `git diff` clean. - -**Green gate:** `go build ./... && go vet ./... && go test ./...` in `hub/` — 18 packages, rc 0. - -## 6. Deployment and live verification - -Hub **0.109.0** built, pushed, manifest bumped, ArgoCD hard-refreshed and synced. Controller -**0.225.0** deployed to both demo boxes. Verified from the live objects, not from a rollout message -(an ArgoCD "rolled out" can name the old image): - -``` -argocd: sync=Synced rev=36f86300209e9ce914f2a97dac16ee9eeea249e2 (== HEAD) -deploy: gitea.dooplex.hu/admin/felhom-hub:0.109.0 -pod: hub-6795879c4b-pf4lz ...felhom-hub:0.109.0 Running -boxes: ...felhom-controller:0.225.0 Up (healthy) on demo-felhom AND demo-hp -``` - -**The card, fetched from the live hub** (endpoint-level: the exact URL the operator's browser -requests; the residual is client-side rendering only — there is no browser on DooPlex): - -| | demo-hp | demo-felhom | +| sweep | candidates | hits | |---|---|---| -| Off-site snapshots | **67** | **10** | -| Repo size | **134.3 MB** | **132.5 KB** | -| Last successful run | 15h ago | 15h ago | -| Soft quota | 50 GB | 50 GB | -| Integrity row | **absent** (grep count 0) | **absent** (grep count 0) | +| `/.zfs/snapshot/YYYY-MM-DDTHH-MM-SS`, nine full days, second granularity | **777,600** | **0** | +| 126 alternative name shapes and snapshot paths (`daily`, `snapshot-1`, colon and compact time forms, `/home/.snapshot`, …) | 126 | 0 | +| **control — the identical 600-name batch shape with one real path appended** | 6 batches | **6/6 returned it** | -Both read `Snapshots 0 · Repo Size 0 MB · Integrity Unknown` before this change. - -**Cross-checked against the source, not just against itself** — the numbers on the card are the -numbers on the boxes: +**And there is a structural reason, which is why I stopped sweeping.** The customer's data and the +snapshot door are on **different filesystems**: ``` -demo-hp settings.json offbox: snapshot_count 67, repo_size_bytes 140829678, stats_known true -demo-felhom settings.json offbox: snapshot_count 10, repo_size_bytes 135635, stats_known true - 135635 / 1024 = 132.5 KB → matches the rendered value +df → u629488-sub3 mounted on /home +stat /home → Device 0,82 +stat /.zfs/snapshot → Device 0,276 ← a different device +stat /home/.zfs → cannot statx: No such file or directory ``` -**One honest detail worth keeping:** demo-felhom's "Last DB dump" reads `—`. That is correct, not a -regression — its only app (`opengist`) has no database, so the box has never taken a DB dump. +A ZFS snapshot under `/.zfs/snapshot` belongs to the dataset that owns that `.zfs` — not to the +child mounted at `/home`. **So even a correctly named snapshot there could not contain +`felhom-repo`,** and the dataset that does hold it exposes no `.zfs` at all to this account. The +empty listing is not a display toggle hiding a reachable tree; from a sub-account there is no tree. -**Not verified live: the "never measured" branch.** Both boxes report `stats_known: true`, so the -degradation path could not be exercised on real hardware without falsifying a box's state. It is -covered by `TestBackupCard_ThreeWayRuling` and `TestBackupCard_OldControllerDegradesToUnknownNotEmpty` -at render level, and this is stated rather than implied. +**Three tools agree, each with controls in the same run:** SFTP, the port-23 shell, and +`rsync --list-only`. -## 7. The push bypassed a gate, deliberately, and here is the declaration +**What that does to R-95.** Its *exposure* is unchanged and its *remedy* is not. Yesterday the row +could say a deletion costs about a day because the rest comes back per-file. Today the only routes +to "the rest" are a whole-box panel rollback (which deletes newer snapshots and hits every customer +on the box) and the provider API (fenced, and unimplemented in the hub's client). **The re-scope's +comfort was resting on a route nobody had walked — which is precisely the standard this project +applies, and it is the reason this drill was called.** -**`git push --no-verify` was used for this change.** `repo_gates.py`'s `golden-currency` gate was -CONVICTED and it was RIGHT: controller **v0.224.0** and **v0.225.0** are released and the newest golden -bake carries **0.223.0**, so a machine installed right now receives neither fix. +**The ranking is Viktor's and I am not re-ranking it.** What I will say plainly: the argument that +moved R-95 down yesterday is the argument this drill removed. On these facts I would put it back +where it was. -**This is a BYPASS, not a waiver.** The gate offers a waiver only for a release that *deliberately needs -no golden*; these need one. **The operator was asked and ruled bypass-now-bake-later**, on the stated -ground that neither fix bites a day-0 box — R-330 is a nightly false alarm about apps a new box has not -installed yet, R-331 is a hub-side display over backups a new box has not taken yet — and both arrive by -self-update afterwards. That ground is recorded on R-242 precisely because it is the thing to re-check: -**it does not extend to a release that changes first-boot behaviour.** +## 2. The RTO -**A golden carrying 0.225.0 is OWED** (`RUNBOOK-manual-build.md` §4.1; the vouch is a three-field change, -`MinAgent 0.129.0`). This is the **fourth** bypass of this gate, and the gap it names is now two releases -wide rather than one. +**Still blank, and it stays a finding.** `07` §8 row 10's RTO cell has been empty since July and this +drill did not fill it. Nothing was recovered, so nothing was timed. The summary line at +`07-backup-architecture.md:948` — *"no ransomware-shaped recovery has ever been run"* — is still +true, and is now true for a sharper reason: **not "nobody has run it" but "from the box, it cannot +be run."** -**The other failing gate was fixed, not bypassed.** `due-checks` was red on R-341's `+7 d` measurement, -five days overdue. It was **taken** during this session — see §8. +## 3. R-432's answer, and the naming scheme -## 8. R-341's overdue check was taken, and its premise did not survive +**R-432 is ANSWERED, and negatively. It did not need the panel read it was waiting on.** -Unrelated to R-331; it blocked the same push, so it was done rather than deferred. Evidence: -`documentation/audits/evidence-r341-plus7d-2026-08-30/step1-fd-and-sockets.txt`. +* **The naming scheme is `YYYY-MM-DDTHH-MM-SS`** — vendor-documented examples `2025-12-03T13-47-47`, + `2025-02-12T11-35-19`. Recorded so nobody hunts a console again. +* **Knowing it does not help.** Every name in that format for nine days is refused, and the st_dev + split above says why. **Per-file recovery is not operator-only — from the box it is nobody's,** and + for the operator it is a browser act against the main account that no credential in this project + can perform. +* **The panel cannot supply the missing piece either.** It offers Restore and Delete on a row and + does not show names; and the one name-shaped thing it could give would be tried against a door + that leads to the wrong dataset. -**Precondition passed**, which is what makes the reading interpretable: `proxmox-backup-proxy` still -`MainPID 551655`, `ps -o lstart=` still `2026-08-18 09:51:04`, `NRestarts=0` — the same proxy generation -as t0, so nothing restarted and re-based the count. (The anchor is `ps`, not `ActiveEnterTimestamp`, -which reads 03:54:54Z here — R-346's trap, avoided.) +## 4. The alarm's first real firing -**Result: fd = 17.** Not 17 more — seventeen total, exactly the documented baseline, against **405** at -the first check on 2026-08-20. The socket histogram holds **one LISTEN and nothing else**: ESTAB 0, -CLOSE-WAIT 0. +**It did not happen, and the drill as written could not have produced it.** The detector fires on a +fall of **more than half** the previous count **and at least 5** (`hub/internal/monitor/offsite.go`, +`snapshotDropFraction = 0.5`, `snapshotDropFloor = 5`). demo-hp's baseline is **69**. Phase 2 deletes +**one app's** history — about **9** snapshots. 9 is over the floor and nowhere near half, so the +alarm stays silent, **correctly and by design**. Firing it for real needs ~35+ snapshots destroyed, +i.e. most of demo-hp's off-site history. That trade is what the operator was asked to rule on. → **R-435** -**The verdict is "unanswerable", not "the upgrade fixed it".** R-341 asks whether the PBS 4.2.5-1 -upgrade changed the fd slope. Inside this interval **we removed the leak ourselves** (R-344, agent -0.130.0, now live on both boxes). A slope of ~0 measures our fix, not the upgrade — and reading it the -other way would credit a changelog that was read in advance and found to contain no such mechanism. The -perturbation pre-registered for this window was Phase C at ~3%; the actual perturbation was the removal -of the entire phenomenon. The question is now **moot**, and the row is closed as such. +**One thing the alarm says is now wrong.** Its message, live in hub 0.111.0, reads: -**What it does establish, which is worth more than the original question:** twelve days after the R-344 -fix, on the same proxy generation with no restart to hide behind, ep0 sits at baseline with zero -established connections. The 388-descriptor accumulation has not returned, and R-336's ~323-day runway -concern retires with it. +> "The daily Storage Box snapshots are read-only and still hold the older copy, **so this is +> recoverable file-by-file**; it is NOT confirmed data loss." -## 9. Not done, and why +The first clause is true; **the second promises a recovery the product cannot perform and the +operator cannot perform without a browser and the main account.** This is this project's own +corollary — *when a verdict changes which field it counts from, the alarm text has to change with +it* — landing on the alarm shipped the same day. → **R-434** -- **No staleness verdict on the card.** `monitor.OffsiteChecker` already owns that and alarms on it. A - second verdict over the same data is two things that can disagree — a shape this codebase has already - paid for (`LastRun` vs `LastSuccess`, R-100). -- **The dead `backup` fields were not removed from the controller's wire format.** Removing them would - stop historical reports already in this hub's store from parsing, for no gain — nothing renders them - now, and a controller-side test fails if anything starts producing them. +## 5. Findings, as register rows + +All four filed in `documentation/backlog/OPEN-ITEMS.md`. + +| row | finding | +|---|---| +| **R-433** | A sub-account cannot reach any Storage Box snapshot **by any name**; `/home` and `/.zfs` are different filesystems and `/home/.zfs` does not exist. Answers R-432 negatively and removes clause (b) of the R-95 re-scope. | +| **R-434** | `emitSnapshotDrop`'s message promises file-by-file recovery that is not reachable. Live in hub 0.111.0. | +| **R-435** | The drop detector cannot see a single-app deletion (>50% of 69 ⇒ ~35 needed). `offbox.go:1388` forgets **by tag**, so a single-tag wipe is exactly the shape the detector is blind to. Deliberate insensitivity, but the blind spot should be stated where the operator reads it. | +| **R-436** | **LEAD, not a defect.** Hetzner's port-23 shell offers `rclone serve restic --stdio` as a server-side backend, and restic 0.14.0 recognises the `rclone:` backend (measured; control `banana:` → invalid backend; rclone is absent from the controller image). `rclone serve restic` carries `--append-only`. **This could make R-95's real prevention far cheaper than the spike concluded — no new always-on machine, no data migration.** Caveat stated up front: the **client** supplies the server command line, so a compromised box could omit the flag unless the provider pins it. Settling that is a vendor question, not a code change. | + +**R-432 is marked ANSWERED**; its "one panel read settles it" next step is withdrawn as unnecessary. + +## 6. Does `07` §8 row 10 move? + +**No. It stays `PARTIAL`, and its RTO stays blank.** The status was already correct for the right +reason — *"the recovery ROUTE has never been walked, which is what PARTIAL means"* — and this drill +found the route is not walkable from the box at all. **What the row needs is a text correction, not a +status change:** its clause *"recoverable per-file (vendor)"* and its limit *"per-file recovery is +operator-only today (R-432)"* both overstate what exists. Updated in place with the citation. Moving +it only as far as the evidence goes means not moving it. + +## 7. What could not be tested, and why + +* **Whether the main account can see the snapshots.** No main-account credential exists in this + project — the hub holds only per-customer sub-accounts. This is the one question that would decide + whether per-file recovery exists *at all*, for anyone. +* **Whether the Hetzner API can list or read a snapshot.** Fenced by the runbook (§11-D). Separately, + `hub/internal/hetznerapi/hetznerapi.go` has **no snapshot method** — so this route needs new code + regardless of the fence. +* **The deletion, the alarm's first real firing, the recovery, the RTO, the re-arm.** Phases 2–5, not + run, on the operator's ruling. +* **Whether `rclone serve restic --stdio` is pinned server-side with `--append-only`** (R-436). + +## 8. My own mistakes + +* **I ran Phase 1 before finishing Phase 0's subject-app choice, and did not say so up front.** The + choice depends on the snapshot's timestamp, so the order was right, but the runbook's order is the + runbook's and a silent reordering is the thing this project keeps getting caught by. Stated at the + time in the session, recorded here. +* **My first sweep guessed the schedule instead of establishing it.** I probed 00:00 UTC and 22:00 + UTC — 600 names — on the strength of a register line reading *"daily 00:00"*, got nothing, and only + then widened to whole days. The narrow sweep was worth nothing on its own: a zero over a guessed + window is not evidence, and I should have gone to full days first or not run it at all. +* **I nearly reported the empty listing as "the display toggle is hiding it".** The vendor documents + exactly such a toggle and it fitted. The st_dev comparison — which I only ran because `df` printed + a filesystem name I did not expect — says the tree is on another dataset entirely. **A plausible + cause that fits the symptom is not a measured one**, and I had the wrong one for about ten minutes. +* **`REPORT.md` held the only copy of the R-331 report** (hub v0.109.0, 2026-08-30) — durable content + living only in the overwritten file, which `CLAUDE.md:82-87` forbids. Preserved as + `REPORT-r331-backup-card.md` before this report replaced it. diff --git a/documentation/architecture/07-backup-architecture.md b/documentation/architecture/07-backup-architecture.md index 7a85d3e8..727b003a 100644 --- a/documentation/architecture/07-backup-architecture.md +++ b/documentation/architecture/07-backup-architecture.md @@ -885,7 +885,7 @@ crosses the line — **R-158**. | 7 | **Guest stopped and does not come back** | everything | guest-power watchdog (60 s, `onboot` as the deliberate-stop discriminator) | automatic | **120 s** | | **PROVEN** | agent v0.107.0 replay — 120 s unattended vs the incident's 587 s with a human | | 8 | **Host dies (hardware), drives intact** | the data drives; ep0's PBS namespace; the hub's Recipe + Escrow | **install a new host, then** `--selftest=bring-up -mode dr` per guest, **then** re-attach drives by `durable_id` | **operator** (SSH) | | 7 d | **IMPLEMENTED** | bring-up code exists and has **never been executed** (`CAMPAIGN-8…:522`); the drive half of the plan is empty on every box (**R-105**) | | 9 | **Whole box lost (fire/theft) — host and drives gone** | ep0 PBS namespace; the restic repo; the hub's Recipe + Escrow | new hardware → day-0 → escrow-consume with **R** → restore guests from PBS → app data from Tier-3 | **operator + customer** (R) | | 7 d (guest) · 24 h (app data) | **IMPLEMENTED / UNPROVEN** | every leg exists; the composed path has never been run. The destructive S5 drill is operator-gated and unrun (`06-offsite-connectivity.md:327`) | -| 10 | **Ransomware / malicious deletion inside the guest** | PBS offsite (the box **cannot** delete its own snapshots); **the restic repo is NOT protected the same way** | whole-guest restore from PBS to a point before the event | **operator** (SSH) | | 7 d | **PARTIAL** | R-89 proved the box is refused when deleting its own PBS snapshot (`CAMPAIGN-8…:514-515`). **R-95 (open, ranked #1):** the restic credential **can delete** — `readonly=False`, `forget --prune` runs from the box, and SFTP cannot express append-only (`OPEN-ITEMS.md:13`) **2026-08-30 (R-359): the store is now VERIFIED on a cadence** — a daily `offsite-integrity` job runs `restic check` when the last successful one is over 7 days old. **This is a readability check, not a restore-test (R-87 stays open).** The depth that ships ON now re-reads **100%** of the pack data (R-399 CLOSED, controller v0.228.0) — the sentence here previously said the opposite and was stale. **2026-08-31 (SPIKE R-87):** a scratch restore of every app on demo-hp was measured at **25 s for 8 snapshots / 774 MB**, against 40.3 s for one weekly check — but restic 0.14.0's `--verify` checks size and mtime, **not content**, so nothing available today can vouch for the restored BYTES. Row 4's verdict is UNCHANGED by the spike. `audits/SPIKE-restic-restore-test-2026-08-31.md`. **2026-09-01 (SPIKE R-95, `audits/SPIKE-r95-offsite-delete-2026-09-01.md`) — THIS ROW'S STATUS IS UNCHANGED, but the seven-day bound beside it is now known to be unverified.** Measured on BOTH boxes over their own SFTP credential, with controls: **no `.snapshots` is visible to either sub-account**, and the account is jailed. Either none exist or a sub-account cannot see them — and a snapshot the box cannot see is one the box cannot restore from, so recovery would be an operator act at the Hetzner panel, not a product capability (R-429; the confirming field needs the provider API, fenced by §11-D). **The PBS shape does NOT transfer**: the sub-account API has one permission axis, `readonly`, and a backup target cannot be read-only — PBS is a server that can refuse, a Storage Box is a filesystem that runs nothing. **What the spike removed as a fear:** withdrawing delete does NOT wedge the store (measured — restic treats a dead owner's lock as stale and proceeds). **What it added:** `unlock --remove-all` reports success while deleting nothing (R-430). **2026-09-01, LATER THE SAME DAY — THE PROBE ABOVE LOOKED FOR THE WRONG NAME AND THIS ROW'S SECOND CLAUSE IS NOW HALF WRONG.** It searched `.snapshots`; the vendor documents `/.zfs/snapshot`. Re-probed at the documented path on BOTH boxes, with controls, the two halves separate cleanly: **(a) the LIVE repository is deletable by the box — unchanged, R-95 stands;** **(b) the daily SNAPSHOTS of it are not writable by anything — PROVEN, not cited:** a write into `/.zfs/snapshot` is refused (`dest open …: Failure`) while the identical write to the account home succeeds. Seven daily snapshots confirmed in the panel (R-429). **So this row's "the restic repo is NOT protected the same way" is true of the repository and FALSE of its snapshots** — a deletion costs at most the day since the last snapshot, recoverable per-file (vendor). **Two limits kept honest:** a sub-account sees the snapshot directory EMPTY, so per-file recovery is operator-only today (R-432); and a panel-driven restore rolls back the WHOLE box. **The status is NOT moved** — evidence (b) is measured, but the recovery ROUTE has never been walked, which is what PARTIAL means. **Detection shipped hub v0.111.0 (R-431):** an unexplained fall in the snapshot count is noticed within a day. | +| 10 | **Ransomware / malicious deletion inside the guest** | PBS offsite (the box **cannot** delete its own snapshots); **the restic repo is NOT protected the same way** | whole-guest restore from PBS to a point before the event | **operator** (SSH) | | 7 d | **PARTIAL** | R-89 proved the box is refused when deleting its own PBS snapshot (`CAMPAIGN-8…:514-515`). **R-95 (open, ranked #1):** the restic credential **can delete** — `readonly=False`, `forget --prune` runs from the box, and SFTP cannot express append-only (`OPEN-ITEMS.md:13`) **2026-08-30 (R-359): the store is now VERIFIED on a cadence** — a daily `offsite-integrity` job runs `restic check` when the last successful one is over 7 days old. **This is a readability check, not a restore-test (R-87 stays open).** The depth that ships ON now re-reads **100%** of the pack data (R-399 CLOSED, controller v0.228.0) — the sentence here previously said the opposite and was stale. **2026-08-31 (SPIKE R-87):** a scratch restore of every app on demo-hp was measured at **25 s for 8 snapshots / 774 MB**, against 40.3 s for one weekly check — but restic 0.14.0's `--verify` checks size and mtime, **not content**, so nothing available today can vouch for the restored BYTES. Row 4's verdict is UNCHANGED by the spike. `audits/SPIKE-restic-restore-test-2026-08-31.md`. **2026-09-01 (SPIKE R-95, `audits/SPIKE-r95-offsite-delete-2026-09-01.md`) — THIS ROW'S STATUS IS UNCHANGED, but the seven-day bound beside it is now known to be unverified.** Measured on BOTH boxes over their own SFTP credential, with controls: **no `.snapshots` is visible to either sub-account**, and the account is jailed. Either none exist or a sub-account cannot see them — and a snapshot the box cannot see is one the box cannot restore from, so recovery would be an operator act at the Hetzner panel, not a product capability (R-429; the confirming field needs the provider API, fenced by §11-D). **The PBS shape does NOT transfer**: the sub-account API has one permission axis, `readonly`, and a backup target cannot be read-only — PBS is a server that can refuse, a Storage Box is a filesystem that runs nothing. **What the spike removed as a fear:** withdrawing delete does NOT wedge the store (measured — restic treats a dead owner's lock as stale and proceeds). **What it added:** `unlock --remove-all` reports success while deleting nothing (R-430). **2026-09-01, LATER THE SAME DAY — THE PROBE ABOVE LOOKED FOR THE WRONG NAME AND THIS ROW'S SECOND CLAUSE IS NOW HALF WRONG.** It searched `.snapshots`; the vendor documents `/.zfs/snapshot`. Re-probed at the documented path on BOTH boxes, with controls, the two halves separate cleanly: **(a) the LIVE repository is deletable by the box — unchanged, R-95 stands;** **(b) the daily SNAPSHOTS of it are not writable by anything — PROVEN, not cited:** a write into `/.zfs/snapshot` is refused (`dest open …: Failure`) while the identical write to the account home succeeds. Seven daily snapshots confirmed in the panel (R-429). **So this row's "the restic repo is NOT protected the same way" is true of the repository and FALSE of its snapshots** — a deletion costs at most the day since the last snapshot, recoverable per-file (vendor). **Two limits kept honest:** a sub-account sees the snapshot directory EMPTY, so per-file recovery is operator-only today (R-432); and a panel-driven restore rolls back the WHOLE box. **The status is NOT moved** — evidence (b) is measured, but the recovery ROUTE has never been walked, which is what PARTIAL means. **Detection shipped hub v0.111.0 (R-431):** an unexplained fall in the snapshot count is noticed within a day — **but only a fall of MORE than half (R-435).** **2026-09-01, THE DRILL — CLAUSE (b) OF THE RE-SCOPE ABOVE IS WITHDRAWN; THE STATUS IS STILL NOT MOVED.** The re-scope said a deletion is *recoverable per-file*. **It is not, from the box: no snapshot is reachable by ANY name.** MEASURED on demo-hp, read-only, no delete verb issued: **777,600** exact names in the vendor form `YYYY-MM-DDTHH-MM-SS` (nine full days, second granularity) plus 126 alternative shapes — **zero hits**, against a control where the identical 600-name batch returns a path that exists (6/6). **Structural cause:** `/home` (`u629488-sub3`) is **st_dev 0,82**, `/.zfs/snapshot` is **st_dev 0,276**, and `/home/.zfs` does not exist — a snapshot under `/.zfs/snapshot` belongs to a different dataset than the one holding the repository. **Clause (a) — the box cannot WRITE into the snapshot area — is unchanged and re-confirmed.** So this row's *"recoverable per-file (vendor)"* and its limit *"per-file recovery is operator-only today (R-432)"* both overstate what exists: the remaining routes are a panel rollback of the WHOLE Storage Box and the provider API (fenced, and the hub's client has no snapshot method at all). **R-432 ANSWERED negatively; R-433 opened. RTO still blank — nothing was recovered, so nothing was timed.** `audits/evidence-drill-r95-recovery-2026-09-01/`. | | 11 | **Hub lost** | every box's data plane, every tier, every Lane-1 route | none needed for recovery **of a box**; the hub itself restores from its Longhorn volume backup | operator (`kubectl`) | | 24 h + weekly (Longhorn `RETAIN 1` each) | **UNPROVEN** | a hub restore has never been performed. The backup target is `nfs://192.168.0.180` — **DooPlex itself** — and exactly **2** restore points exist (LIVE, INV Part D2.2) | | 11b | *consequences while the hub is gone* | — | — | — | | | **[FACT]** | **day:** nothing customer-visible breaks; events queue (`settings.go:1466-1485`). **week:** the operator alarm plane is dark, no claim/reset codes, no config or floor convergence, no PBS-secret re-issue. **permanently:** escrow custody and break-glass credentials are gone (INV Part D2.4) | | 12 | **Offsite provider lost (Hetzner)** | everything on-premises: both drives, both whole-guest tiers | none needed — on-premises recovery is unaffected. Re-provision a new offsite target. | operator | | | **[FACT]** | **restic and PBS share the provider** (INV Part E.1). Whether they share an account and payment method is **UNKNOWN** → §11-D | @@ -1085,7 +1085,7 @@ does **not** hold as written. → **R-108** | **R-356** | The off-site restore resolved its destination with the raw `HDD_PATH` and read an empty answer as "not installed" — **CLOSED, controller v0.219.0, 2026-08-22** | 40 of 53 apps were refused permanently while running (§6.3 `[DESIGN]`) | | ~~**R-108**~~ | ~~Network storage can host an app's namespace~~ | **CLOSED 2026-07-30, controller v0.187.0 — D5 UNBLOCKED.** An app namespace may no longer be placed on network storage (5 surfaces guarded by one fail-closed predicate); the share-root bind is deliberately UNCHANGED because it is load-bearing and unscopable (§10.1). `audits/R108-network-app-namespace-2026-07-30.md` | | **R-126** | A `.fab` bundle — plaintext secrets, optional password — can be exported ONTO a NAS: `storageDriveList()` (`internal/web/handler_export.go`) does not filter network paths | split out of R-108, which closed without it. NOT a D5 precondition: an explicit customer-chosen export destination, not a browsing surface reaching a backup tree (§5, §7.3) | -| R-95 (open) | The restic offsite credential **can delete** — the box can `forget --prune` its own repo, from **two** call sites (`offbox.go:1388` retention and `offbox.go:1759` over-quota) | the tier holding the customer's documents and photos is the one whose credential can destroy it (matrix row 10). **SPIKE 2026-09-01:** prevention needs a transport change (restic 0.14.0 does speak `rest:` — measured; append-only is a rest-server flag, not a restic one), because the sub-account API cannot express write-without-delete. **Detection is nearly free and is recommended first** — `snapshot_count` already reaches the hub and the hub appends reports, so the comparison needs no box change. `audits/SPIKE-r95-offsite-delete-2026-09-01.md`. **RE-SCOPED 2026-09-01:** the box can delete the LIVE repository but **cannot write to the daily snapshots of it** (measured, both boxes) — so the exposure is at most one day's data plus an operator-driven per-file recovery, not open-ended loss. Detection shipped hub v0.111.0 (R-431). | +| R-95 (open) | The restic offsite credential **can delete** — the box can `forget --prune` its own repo, from **two** call sites (`offbox.go:1388` retention and `offbox.go:1759` over-quota) | the tier holding the customer's documents and photos is the one whose credential can destroy it (matrix row 10). **SPIKE 2026-09-01:** prevention needs a transport change (restic 0.14.0 does speak `rest:` — measured; append-only is a rest-server flag, not a restic one), because the sub-account API cannot express write-without-delete. **Detection is nearly free and is recommended first** — `snapshot_count` already reaches the hub and the hub appends reports, so the comparison needs no box change. `audits/SPIKE-r95-offsite-delete-2026-09-01.md`. **RE-SCOPED 2026-09-01:** the box can delete the LIVE repository but **cannot write to the daily snapshots of it** (measured, both boxes) — so the exposure is at most one day's data plus an operator-driven per-file recovery, not open-ended loss. Detection shipped hub v0.111.0 (R-431). **DRILL 2026-09-01: the re-scope's recovery clause is WITHDRAWN — no snapshot is reachable from a sub-account by any name (R-433, 777,600 names, controlled), so the exposure is NOT bounded by an operator-driven per-file recovery. R-436 is the new cheap lead: the provider already offers `rclone serve restic --stdio` server-side and restic 0.14.0 speaks `rclone:` (measured) — but the client supplies the server command line, so ask the vendor whether `--append-only` is pinned BEFORE building anything.** | | ~~R-86~~ | ~~Restore-tests are interval-scheduled, not backup-aligned~~ | **CLOSED 2026-08-03 — agent v0.121.0 + hub v0.91.0.** Restore-testing is now **per archive generation**: a tier is due when its newest archive that has settled ~24 h has not been proven, so a daily tier is proved daily on its own archive and a weekly tier weekly on its own. The ticker survives only as the evaluation interval (6 h, chosen from a measured cost). The hub's staleness window moved with it — per tier, from that tier's observed archive rhythm — because a weekly tier proved weekly sat EXACTLY on the old flat 7-day line (§3, Lane 2's per-archive rule) | | ~~R-353~~ | ~~A local unit restore reported a bare completion whether it returned an entire dataset or nothing~~ | **CLOSED 2026-08-30, controller v0.226.0.** The volume count already existed (`restoreDockerVolumesFrom`) and was discarded by a one-line wrapper, so the surface was structurally unable to say what came back. `RestoreFromRecoveryUnit` now returns `UnitRestoreResult` and the sentence has three cases, keyed on replayed-vs-**listed** — zero-replayed has two causes that are opposite news. Every sentence is a claim about the BACKUP, never the app: this path has no `SafetyDump` discriminator, and §6.3 is why that is not pedantry. Proven live on `demo-hp` | | ~~R-357~~ | ~~The destructive reconstitute had no free-space gate; all three that existed guarded non-destructive paths~~ | **CLOSED 2026-08-30, controller v0.226.0.** The gate sits before `writeSafetyDump` and `StopStack`, so a refusal costs no outage — the test asserts the StopStack call count, not the sentence. No headroom multiplier (a measured tree, not a predicted download); fail-closed when either probe reads ≤ 0, which was a real fail-open hole. **Not live-validated** — filling a filesystem is a drill step | diff --git a/documentation/audits/evidence-drill-r95-recovery-2026-09-01/README.md b/documentation/audits/evidence-drill-r95-recovery-2026-09-01/README.md new file mode 100644 index 00000000..2ed2657d --- /dev/null +++ b/documentation/audits/evidence-drill-r95-recovery-2026-09-01/README.md @@ -0,0 +1,52 @@ +# Evidence — DRILL: the deletion we said is survivable (R-95 recovery drill, 2026-09-01) + +**The drill STOPPED at the end of Phase 1, on the operator's ruling, before any destructive step.** +**No delete verb was issued against any live store. No byte on either Storage Box sub-account was +written, moved or removed.** Every probe in this directory is a read: `ls`, `stat`, `tree`, `du`, +`df`, `rsync --list-only`, `restic snapshots`. + +## Why it stopped + +Phase 1 was supposed to *find* a Storage Box snapshot so Phase 4 could recover from it. It found +that **no snapshot is reachable from the box by any name**, and that **no unfenced route to one +exists for CC at all**. The recovery leg therefore could not run, and deleting a real app's off-site +history would have bought only an alarm test that — measured against the shipped threshold — could +not have fired at the size the runbook specifies. Put to the operator as a two-option decision; +the ruling was **stop and report**. + +## Files + +| file | what it is | +|---|---| +| `phase0-ground-truth/01-offsite-inventory.txt` | demo-hp's full off-site inventory — 69 restic snapshots, 9 apps, ids + tags + paths. The baseline. | +| `phase0-ground-truth/02-hub-view.txt` | the hub's latest `offsite` object for demo-hp (`snapshot_count: 69`, `stats_known: true`, `last_status: ok`). Cross-checks the restic count exactly. | +| `phase1-find-the-snapshot/01-probe-tree.txt` | SFTP probes of `/.zfs`, `/.zfs/snapshot`, `/home/.zfs`, with a positive and a negative control. Reproduces R-432. | +| `phase1-find-the-snapshot/02-shell-and-rsync-doors.txt` | two independent tools (a real shell on port 23, and `rsync --list-only`) agree the tree lists empty. Both controlled. | +| `phase1-find-the-snapshot/03-shell-capabilities.txt` | the port-23 restricted shell's full `help` — the command set the credential actually reaches. | +| `phase1-find-the-snapshot/04-enumeration-attempts.txt` | `df`/`stat`/`tree`/`du` on the snapshot door. **Where the st_dev split was found.** | +| `phase1-find-the-snapshot/05-home-zfs-door.txt` | `/home/.zfs` does not exist — three tools, negative control in the same run. | +| `phase1-find-the-snapshot/06-name-sweep.txt` | the name sweep: **nine full days, second granularity, 777,600 candidate names, zero hits.** | +| `phase1-find-the-snapshot/07-sweep-control.txt` | **the control for that sweep** — the identical 600-name batch shape with one real path appended; 6/6 batches returned it. Without this the zero result would prove nothing. | +| `phase1-find-the-snapshot/08-alt-formats.txt` | 126 non-timestamp name shapes and alternative snapshot paths. Only the controls resolved. | +| `phase1-find-the-snapshot/09-rclone-restic-lead.txt` | the `rclone:` backend measurement behind R-436, with the `banana:` invalid-backend control. | +| `phase6-teardown/01-verify-then-clean.txt` | the store is untouched: still 69 snapshots, home is exactly `.ssh` + `felhom-repo`, repo top level intact. | +| `phase6-teardown/02-guest-and-host-clean.txt` | scratch removed from the guest and the PVE host; controller 0.232.0 healthy, 19 healthy containers. | +| `phase6-teardown/03-dooplex-clean-and-final-state.txt` | hub DB copies shredded (both, `-wal` included); both customers' tiers still `ok`; no alarm event raised by this session. | + +## The method that makes the negative result trustworthy + +A sweep that finds nothing is worthless unless it is shown it *could* have found something. The +oracle is a batched `stat -c %n` over the port-23 shell: 500–600 paths per round trip, and stdout +carries **only the paths that exist**. `07-sweep-control.txt` runs the identical batch shape with +`/home` appended and gets `/home` back, 6 times out of 6. So the zero in `06-name-sweep.txt` is a +measurement, not a silence. + +## Not done, and why + +* **No Hetzner API call.** Fenced by the runbook (§11-D). The hub's own API client has no snapshot + method at all, so even unfenced it would have needed new code. +* **No panel action.** No browser on DooPlex, and "Restore snapshot" is fenced in any case — it rolls + back the whole Storage Box and deletes newer snapshots. +* **`demo-felhom` never addressed.** Every storage-box connection in this session authenticated as + `u629488-sub3`, demo-hp's own sub-account. Its tier is verified still reporting `ok` in + `phase6-teardown/03-*`. diff --git a/documentation/audits/evidence-drill-r95-recovery-2026-09-01/phase0-ground-truth/01-offsite-inventory.txt b/documentation/audits/evidence-drill-r95-recovery-2026-09-01/phase0-ground-truth/01-offsite-inventory.txt new file mode 100644 index 00000000..0a6439ba --- /dev/null +++ b/documentation/audits/evidence-drill-r95-recovery-2026-09-01/phase0-ground-truth/01-offsite-inventory.txt @@ -0,0 +1,162 @@ +### date (guest UTC): 2026-09-01T14:30:31Z +### restic snapshots --json is too noisy; table form: +ID Time Host Tags Paths +------------------------------------------------------------------------------------------------------------------------------ +41c830db 2026-08-09 08:30:38 demo-hp felhom-offbox,calibre-web /mnt/sys_drive/felhom-data/backups/primary/calibre-web + /mnt/sys_drive/felhom-data/userdata/media/books + +9e38b84c 2026-08-09 08:30:49 demo-hp felhom-offbox,opengist /mnt/sys_drive/felhom-data/backups/primary/opengist + +78b93f04 2026-08-09 08:30:53 demo-hp felhom-offbox,privatebin /mnt/sys_drive/felhom-data/backups/primary/privatebin + +f06ad130 2026-08-23 02:16:25 demo-hp felhom-offbox,bookstack /mnt/sys_drive/felhom-data/backups/primary/bookstack + +e68f6444 2026-08-23 02:16:32 demo-hp felhom-offbox,calibre-web /mnt/felhom-drives/hdd_1/backups/primary/calibre-web + /mnt/felhom-drives/hdd_1/userdata/media/books + +d1a614b5 2026-08-23 02:16:36 demo-hp felhom-offbox,privatebin /mnt/sys_drive/felhom-data/backups/primary/privatebin + +e23e34a5 2026-08-23 02:16:39 demo-hp felhom-offbox,docmost /mnt/sys_drive/felhom-data/backups/primary/docmost + +5c3492ad 2026-08-23 02:16:45 demo-hp felhom-offbox,kimai /mnt/sys_drive/felhom-data/backups/primary/kimai + +82ef8692 2026-08-23 02:16:55 demo-hp felhom-offbox,opengist /mnt/sys_drive/felhom-data/backups/primary/opengist + +778e8feb 2026-08-23 02:16:58 demo-hp felhom-offbox,paperless-ngx /mnt/felhom-drives/hdd_1/appdata/paperless/media + /mnt/felhom-drives/hdd_1/backups/primary/paperless-ngx + +8cc0e4b9 2026-08-23 02:17:04 demo-hp felhom-offbox,romm /mnt/felhom-drives/hdd_1/backups/primary/romm + +936a30c3 2026-08-26 02:16:29 demo-hp felhom-offbox,kimai /mnt/sys_drive/felhom-data/backups/primary/kimai + +13c3cf9d 2026-08-26 02:16:39 demo-hp felhom-offbox,opengist /mnt/sys_drive/felhom-data/backups/primary/opengist + +9e3474ec 2026-08-26 02:16:43 demo-hp felhom-offbox,romm /mnt/felhom-drives/hdd_1/backups/primary/romm + +a2aec191 2026-08-26 02:16:49 demo-hp felhom-offbox,paperless-ngx /mnt/felhom-drives/hdd_1/appdata/paperless/media + /mnt/felhom-drives/hdd_1/backups/primary/paperless-ngx + +389a3a6d 2026-08-26 02:16:55 demo-hp felhom-offbox,privatebin /mnt/sys_drive/felhom-data/backups/primary/privatebin + +c7408e7d 2026-08-26 02:16:59 demo-hp felhom-offbox,bookstack /mnt/sys_drive/felhom-data/backups/primary/bookstack + +f40c8fd9 2026-08-26 02:17:03 demo-hp felhom-offbox,calibre-web /mnt/felhom-drives/hdd_1/backups/primary/calibre-web + /mnt/felhom-drives/hdd_1/userdata/media/books + +ea5e9a25 2026-08-26 02:17:07 demo-hp felhom-offbox,docmost /mnt/sys_drive/felhom-data/backups/primary/docmost + +e6ee92ac 2026-08-27 02:16:26 demo-hp felhom-offbox,paperless-ngx /mnt/felhom-drives/hdd_1/appdata/paperless/media + /mnt/felhom-drives/hdd_1/backups/primary/paperless-ngx + +73266947 2026-08-27 02:16:34 demo-hp felhom-offbox,privatebin /mnt/sys_drive/felhom-data/backups/primary/privatebin + +794d8219 2026-08-27 02:16:37 demo-hp felhom-offbox,bookstack /mnt/sys_drive/felhom-data/backups/primary/bookstack + +b1e4d6a5 2026-08-27 02:16:41 demo-hp felhom-offbox,calibre-web /mnt/felhom-drives/hdd_1/backups/primary/calibre-web + /mnt/felhom-drives/hdd_1/userdata/media/books + +7eea66d2 2026-08-27 02:16:45 demo-hp felhom-offbox,docmost /mnt/sys_drive/felhom-data/backups/primary/docmost + +b9db5519 2026-08-27 02:16:50 demo-hp felhom-offbox,kimai /mnt/sys_drive/felhom-data/backups/primary/kimai + +0751dad2 2026-08-27 02:17:00 demo-hp felhom-offbox,opengist /mnt/sys_drive/felhom-data/backups/primary/opengist + +3714731c 2026-08-27 02:17:03 demo-hp felhom-offbox,romm /mnt/felhom-drives/hdd_1/backups/primary/romm + +8dcbd53a 2026-08-28 02:16:30 demo-hp felhom-offbox,bookstack /mnt/sys_drive/felhom-data/backups/primary/bookstack + +5a635785 2026-08-28 02:16:34 demo-hp felhom-offbox,calibre-web /mnt/felhom-drives/hdd_1/backups/primary/calibre-web + /mnt/felhom-drives/hdd_1/userdata/media/books + +035e923f 2026-08-28 02:16:38 demo-hp felhom-offbox,docmost /mnt/sys_drive/felhom-data/backups/primary/docmost + +6f613bbd 2026-08-28 02:16:43 demo-hp felhom-offbox,kimai /mnt/sys_drive/felhom-data/backups/primary/kimai + +f0c6ce70 2026-08-28 02:16:53 demo-hp felhom-offbox,opengist /mnt/sys_drive/felhom-data/backups/primary/opengist + +02f1de9d 2026-08-28 02:16:56 demo-hp felhom-offbox,romm /mnt/felhom-drives/hdd_1/backups/primary/romm + +726c1ce8 2026-08-28 02:17:02 demo-hp felhom-offbox,paperless-ngx /mnt/felhom-drives/hdd_1/appdata/paperless/media + /mnt/felhom-drives/hdd_1/backups/primary/paperless-ngx + +c0e9dc25 2026-08-28 02:17:08 demo-hp felhom-offbox,privatebin /mnt/sys_drive/felhom-data/backups/primary/privatebin + +0b90da52 2026-08-29 02:16:26 demo-hp felhom-offbox,bookstack /mnt/sys_drive/felhom-data/backups/primary/bookstack + +6aee1818 2026-08-29 02:16:30 demo-hp felhom-offbox,calibre-web /mnt/felhom-drives/hdd_1/backups/primary/calibre-web + /mnt/felhom-drives/hdd_1/userdata/media/books + +7aea3e6d 2026-08-29 02:16:34 demo-hp felhom-offbox,docmost /mnt/sys_drive/felhom-data/backups/primary/docmost + +2f043b35 2026-08-29 02:16:39 demo-hp felhom-offbox,kimai /mnt/sys_drive/felhom-data/backups/primary/kimai + +fa862b02 2026-08-29 02:16:49 demo-hp felhom-offbox,opengist /mnt/sys_drive/felhom-data/backups/primary/opengist + +3f93ec62 2026-08-29 02:16:52 demo-hp felhom-offbox,romm /mnt/felhom-drives/hdd_1/backups/primary/romm + +ebade956 2026-08-29 02:16:58 demo-hp felhom-offbox,paperless-ngx /mnt/felhom-drives/hdd_1/appdata/paperless/media + /mnt/felhom-drives/hdd_1/backups/primary/paperless-ngx + +983a0365 2026-08-29 02:17:06 demo-hp felhom-offbox,privatebin /mnt/sys_drive/felhom-data/backups/primary/privatebin + +9ae0aa88 2026-08-30 02:16:29 demo-hp felhom-offbox,romm /mnt/felhom-drives/hdd_1/backups/primary/romm + +2404511d 2026-08-30 02:16:35 demo-hp felhom-offbox,paperless-ngx /mnt/felhom-drives/hdd_1/appdata/paperless/media + /mnt/felhom-drives/hdd_1/backups/primary/paperless-ngx + +4a5db84d 2026-08-30 02:16:42 demo-hp felhom-offbox,privatebin /mnt/sys_drive/felhom-data/backups/primary/privatebin + +ea40c93a 2026-08-30 02:16:45 demo-hp felhom-offbox,bookstack /mnt/sys_drive/felhom-data/backups/primary/bookstack + +37b7b51a 2026-08-30 02:16:49 demo-hp felhom-offbox,calibre-web /mnt/felhom-drives/hdd_1/backups/primary/calibre-web + /mnt/felhom-drives/hdd_1/userdata/media/books + +ee67d85d 2026-08-30 02:16:53 demo-hp felhom-offbox,docmost /mnt/sys_drive/felhom-data/backups/primary/docmost + +84542ec8 2026-08-30 02:16:58 demo-hp felhom-offbox,kimai /mnt/sys_drive/felhom-data/backups/primary/kimai + +a5c5c973 2026-08-30 02:17:08 demo-hp felhom-offbox,opengist /mnt/sys_drive/felhom-data/backups/primary/opengist + +05c3346a 2026-08-31 21:29:03 demo-hp felhom-offbox,bookstack /mnt/sys_drive/felhom-data/backups/primary/bookstack + +b6b7c205 2026-08-31 21:29:08 demo-hp felhom-offbox,calibre-web /mnt/felhom-drives/hdd_1/backups/primary/calibre-web + /mnt/felhom-drives/hdd_1/userdata/media/books + +0b5781f5 2026-08-31 21:29:12 demo-hp felhom-offbox,docmost /mnt/sys_drive/felhom-data/backups/primary/docmost + +6a2672b0 2026-08-31 21:29:16 demo-hp felhom-offbox,paperless-ngx /mnt/felhom-drives/hdd_1/appdata/paperless/media + /mnt/felhom-drives/hdd_1/backups/primary/paperless-ngx + +d3970ee3 2026-08-31 21:29:21 demo-hp felhom-offbox,kimai /mnt/sys_drive/felhom-data/backups/primary/kimai + +5d5bee7f 2026-08-31 21:29:31 demo-hp felhom-offbox,opengist /mnt/sys_drive/felhom-data/backups/primary/opengist + +7b0a7161 2026-08-31 21:29:34 demo-hp felhom-offbox,privatebin /mnt/sys_drive/felhom-data/backups/primary/privatebin + +1b8c7361 2026-08-31 21:29:38 demo-hp felhom-offbox,romm /mnt/felhom-drives/hdd_1/backups/primary/romm + +5dcee6e3 2026-08-31 21:29:44 demo-hp felhom-offbox,bentopdf /mnt/sys_drive/felhom-data/backups/primary/bentopdf + +6fee3b5a 2026-09-01 02:16:30 demo-hp felhom-offbox,calibre-web /mnt/felhom-drives/hdd_1/backups/primary/calibre-web + /mnt/felhom-drives/hdd_1/userdata/media/books + +b2e059ee 2026-09-01 02:16:34 demo-hp felhom-offbox,docmost /mnt/sys_drive/felhom-data/backups/primary/docmost + +1e179cfe 2026-09-01 02:16:39 demo-hp felhom-offbox,paperless-ngx /mnt/felhom-drives/hdd_1/appdata/paperless/media + /mnt/felhom-drives/hdd_1/backups/primary/paperless-ngx + +f69e0510 2026-09-01 02:16:45 demo-hp felhom-offbox,kimai /mnt/sys_drive/felhom-data/backups/primary/kimai + +8b9d9a94 2026-09-01 02:16:55 demo-hp felhom-offbox,opengist /mnt/sys_drive/felhom-data/backups/primary/opengist + +1e548d3b 2026-09-01 02:16:58 demo-hp felhom-offbox,privatebin /mnt/sys_drive/felhom-data/backups/primary/privatebin + +ae903b66 2026-09-01 02:17:01 demo-hp felhom-offbox,romm /mnt/felhom-drives/hdd_1/backups/primary/romm + +9d002b38 2026-09-01 02:17:08 demo-hp felhom-offbox,bentopdf /mnt/sys_drive/felhom-data/backups/primary/bentopdf + +7a855a65 2026-09-01 02:17:11 demo-hp felhom-offbox,bookstack /mnt/sys_drive/felhom-data/backups/primary/bookstack +------------------------------------------------------------------------------------------------------------------------------ +69 snapshots + +### rc=0 diff --git a/documentation/audits/evidence-drill-r95-recovery-2026-09-01/phase0-ground-truth/02-hub-view.txt b/documentation/audits/evidence-drill-r95-recovery-2026-09-01/phase0-ground-truth/02-hub-view.txt new file mode 100644 index 00000000..92492486 --- /dev/null +++ b/documentation/audits/evidence-drill-r95-recovery-2026-09-01/phase0-ground-truth/02-hub-view.txt @@ -0,0 +1,21 @@ +report id 22221 received_at 2026-09-01 14:45:18 +top-level keys: ['app_telemetry', 'backup', 'claimed', 'config_hash', 'containers', 'controller_url', 'controller_version', 'customer_id', 'customer_name', 'dr_recipe', 'geo_restriction', 'health', 'offsite', 'stacks', 'storage', 'system', 'timestamp', 'version'] +--- offsite --- +{ + "enabled": true, + "escrow_state": "escrowed", + "last_run": "2026-09-01T02:17:57Z", + "last_status": "ok", + "last_success": "2026-09-01T02:17:57Z", + "snapshot_count": 69, + "repo_size_bytes": 147274432, + "quota_gb": 50, + "stats_known": true, + "last_integrity_check": "2026-09-01T08:24:24Z", + "last_integrity_ok": true, + "last_integrity_depth": "100%", + "last_proof_run": "2026-09-01T04:04:48Z", + "last_proof_stack": "bookstack", + "last_proof_snapshot": "7a855a65", + "last_proof_result": "pass" +} diff --git a/documentation/audits/evidence-drill-r95-recovery-2026-09-01/phase1-find-the-snapshot/01-probe-tree.txt b/documentation/audits/evidence-drill-r95-recovery-2026-09-01/phase1-find-the-snapshot/01-probe-tree.txt new file mode 100644 index 00000000..f23ca103 --- /dev/null +++ b/documentation/audits/evidence-drill-r95-recovery-2026-09-01/phase1-find-the-snapshot/01-probe-tree.txt @@ -0,0 +1,31 @@ +--- ls -la . +Connected to u629488-sub3.your-storagebox.de. +sftp> ls -la . +drwxr-xr-x ? u629488-sub3 1058 4 Sep 1 12:12 ./. +dr-x--x--x ? root root 11 Jul 21 16:01 ./.. +drwx------ ? u629488-sub3 1058 3 Jul 23 09:53 ./.ssh +drwxrwxr-x ? u629488-sub3 1058 8 Aug 4 12:38 ./felhom-repo +--- ls -la ./zzz-no-such-r95-negative-control +Connected to u629488-sub3.your-storagebox.de. +sftp> ls -la ./zzz-no-such-r95-negative-control +Can't ls: "/home/./zzz-no-such-r95-negative-control" not found +--- ls -la /.zfs +Connected to u629488-sub3.your-storagebox.de. +sftp> ls -la /.zfs +drwxrwxrwx ? root root 0 Jul 21 16:01 /.zfs/. +dr-x--x--x ? root root 11 Jul 21 16:01 /.zfs/.. +drwxrwxrwx ? root root 2 Sep 1 12:12 /.zfs/shares +drwxrwxrwx ? root root 2 Jan 1 1970 /.zfs/snapshot +--- ls -la /.zfs/snapshot +Connected to u629488-sub3.your-storagebox.de. +sftp> ls -la /.zfs/snapshot +drwxrwxrwx ? root root 2 Jan 1 1970 /.zfs/snapshot/. +drwxrwxrwx ? root root 0 Jul 21 16:01 /.zfs/snapshot/.. +--- ls -la /home/.zfs +Connected to u629488-sub3.your-storagebox.de. +sftp> ls -la /home/.zfs +Can't ls: "/home/.zfs" not found +--- ls -la /home/.zfs/snapshot +Connected to u629488-sub3.your-storagebox.de. +sftp> ls -la /home/.zfs/snapshot +Can't ls: "/home/.zfs/snapshot" not found diff --git a/documentation/audits/evidence-drill-r95-recovery-2026-09-01/phase1-find-the-snapshot/02-shell-and-rsync-doors.txt b/documentation/audits/evidence-drill-r95-recovery-2026-09-01/phase1-find-the-snapshot/02-shell-and-rsync-doors.txt new file mode 100644 index 00000000..6cac9db4 --- /dev/null +++ b/documentation/audits/evidence-drill-r95-recovery-2026-09-01/phase1-find-the-snapshot/02-shell-and-rsync-doors.txt @@ -0,0 +1,21 @@ +=== A. ssh with a command (limited shell?) === +total 0 +drwxrwxrwx 2 root root 2 Jan 1 1970 . +drwxrwxrwx 1 root root 0 Jul 21 16:01 .. +rc=0 + +=== B. rsync --list-only on the account home (POSITIVE CONTROL) === +drwxr-xr-x 4 2026/09/01 12:12:48 . +drwx------ 3 2026/07/23 09:53:43 .ssh +drwxrwxr-x 8 2026/08/04 12:38:16 felhom-repo +rc=0 + +=== C. rsync --list-only on /.zfs/snapshot/ (THE PROBE) === +drwxrwxrwx 2 1970/01/01 00:00:00 . +rc=0 + +=== D. rsync --list-only on a path that must NOT exist (NEGATIVE CONTROL) === +rsync: [sender] change_dir "/zzz-no-such-r95" failed: No such file or directory (2) +rsync error: some files/attrs were not transferred (see previous errors) (code 23) at main.c(1874) [Receiver=3.2.7] +rsync: [Receiver] write error: Broken pipe (32) +rc=0 diff --git a/documentation/audits/evidence-drill-r95-recovery-2026-09-01/phase1-find-the-snapshot/03-shell-capabilities.txt b/documentation/audits/evidence-drill-r95-recovery-2026-09-01/phase1-find-the-snapshot/03-shell-capabilities.txt new file mode 100644 index 00000000..239c9eca --- /dev/null +++ b/documentation/audits/evidence-drill-r95-recovery-2026-09-01/phase1-find-the-snapshot/03-shell-capabilities.txt @@ -0,0 +1,54 @@ +=== id / pwd / shell === +Command not found. Use 'help' to get a list of available commands. +=== ls / === +/usr/bin/ls: cannot open directory '/': Permission denied +=== mount === +/usr/bin/cat: /proc/self/mountinfo: No such file or directory +=== which tools === +Command not found. Use 'help' to get a list of available commands. +=== help === ++------------------------------------------------------------------------------+ +| The following commands are available: | +| ls|ll list directory content | +| tree list directory content | +| cd change current working directory | +| pwd show current working directory | +| mkdir create new directory | +| rmdir delete directory | +| du disk usage of files/directories | +| df show disk usage | +| dd read and write files | +| cat output file content | +| touch create new file | +| cp copy files/directories | +| rm delete files/directories | +| unlink delete file/directory | +| mv move files/directories | +| chmod change file/directory permissions | +| md5|sha1|sha256|sha512 create hash sum of file | +| md5sum|sha1sum|sha256sum|sha512sum create hash sum of file | +| head show first lines of file | +| tail show last lines of file | +| grep search for specific string in files | +| stat stat files/directory | +| version show version of this SSH environment | +| | +| Available as server side backend: | +| borg | +| rsync | +| scp | +| sftp | +| rclone serve restic --stdio | +| | +| Please note that this is only a restricted shell which do not | +| support shell features like redirects or pipes. | +| | +| You can find more information in our Docs: | +| https://docs.hetzner.com/storage/storage-box/ | ++------------------------------------------------------------------------------+ +=== snapshot-ish commands === +Command not found. Use 'help' to get a list of available commands. +-- +Command not found. Use 'help' to get a list of available commands. +-- +Command not found. Use 'help' to get a list of available commands. diff --git a/documentation/audits/evidence-drill-r95-recovery-2026-09-01/phase1-find-the-snapshot/04-enumeration-attempts.txt b/documentation/audits/evidence-drill-r95-recovery-2026-09-01/phase1-find-the-snapshot/04-enumeration-attempts.txt new file mode 100644 index 00000000..10900974 --- /dev/null +++ b/documentation/audits/evidence-drill-r95-recovery-2026-09-01/phase1-find-the-snapshot/04-enumeration-attempts.txt @@ -0,0 +1,41 @@ +=== df === +Filesystem 1K-blocks Used Available Use% Mounted on +u629488-sub3 1073645568 2685568 1070960000 1% /home +=== stat /.zfs/snapshot === + File: /.zfs/snapshot + Size: 2 Blocks: 0 IO Block: 512 directory +Device: 0,276 Inode: 281474976710653 Links: 2 +Access: (0777/drwxrwxrwx) Uid: ( 0/ root) Gid: ( 0/ root) +Access: 2026-09-01 14:36:07.861516179 +0000 +Modify: 1970-01-01 00:00:00.000000000 +0000 +Change: 1970-01-01 00:00:00.000000000 +0000 + Birth: - +=== ls -a /.zfs/snapshot === +. +.. +=== tree /.zfs === +/.zfs +├── shares +└── snapshot + +3 directories, 0 files +=== du /.zfs/snapshot === +0 /.zfs/snapshot +=== multi-arg stat test (2 real, 1 fake) === + File: /home + Size: 4 Blocks: 1 IO Block: 131072 directory +Device: 0,82 Inode: 2273 Links: 4 +Access: (0755/drwxr-xr-x) Uid: ( 1057/u629488-sub3) Gid: ( 1058/ UNKNOWN) +Access: 2026-07-21 16:01:41.810597438 +0000 +Modify: 2026-09-01 12:12:48.637707029 +0000 +Change: 2026-09-01 12:12:48.637707029 +0000 + Birth: 2026-07-21 16:01:41.810597438 +0000 + File: /.zfs/snapshot + Size: 2 Blocks: 0 IO Block: 512 directory +Device: 0,276 Inode: 281474976710653 Links: 2 +Access: (0777/drwxrwxrwx) Uid: ( 0/ root) Gid: ( 0/ root) +Access: 2026-09-01 14:36:10.597470894 +0000 +Modify: 1970-01-01 00:00:00.000000000 +0000 +Change: 1970-01-01 00:00:00.000000000 +0000 + Birth: - +/usr/bin/stat: cannot statx '/zzz-no-such-r95': No such file or directory diff --git a/documentation/audits/evidence-drill-r95-recovery-2026-09-01/phase1-find-the-snapshot/05-home-zfs-door.txt b/documentation/audits/evidence-drill-r95-recovery-2026-09-01/phase1-find-the-snapshot/05-home-zfs-door.txt new file mode 100644 index 00000000..d7deb7d5 --- /dev/null +++ b/documentation/audits/evidence-drill-r95-recovery-2026-09-01/phase1-find-the-snapshot/05-home-zfs-door.txt @@ -0,0 +1,16 @@ +=== stat /home/.zfs and /home/.zfs/snapshot (fake path = negative control) === +/usr/bin/stat: cannot statx '/home/.zfs': No such file or directory +/usr/bin/stat: cannot statx '/home/.zfs/snapshot': No such file or directory +/usr/bin/stat: cannot statx '/home/zzz-no-such-r95': No such file or directory + +=== ls -a /home/.zfs/snapshot === +/usr/bin/ls: cannot access '/home/.zfs/snapshot': No such file or directory +=== ls -a /home === +. +.. +.ssh +felhom-repo +=== tree /home/.zfs === +/home/.zfs [error opening dir] + +0 directories, 0 files diff --git a/documentation/audits/evidence-drill-r95-recovery-2026-09-01/phase1-find-the-snapshot/06-name-sweep.txt b/documentation/audits/evidence-drill-r95-recovery-2026-09-01/phase1-find-the-snapshot/06-name-sweep.txt new file mode 100644 index 00000000..87d77e0c --- /dev/null +++ b/documentation/audits/evidence-drill-r95-recovery-2026-09-01/phase1-find-the-snapshot/06-name-sweep.txt @@ -0,0 +1,165 @@ +warm-up / positive control: /home +........................ day 2026-09-01 done, batches-with-hits=0 +........................ day 2026-08-31 done, batches-with-hits=0 +### CONTROL RUN — same code path, /home injected into every batch; MUST hit 144/144 +HIT(2026-09-01 00:0x): /home +HIT(2026-09-01 00:1x): /home +HIT(2026-09-01 00:2x): /home +HIT(2026-09-01 00:3x): /home +HIT(2026-09-01 00:4x): /home +HIT(2026-09-01 00:5x): /home +.HIT(2026-09-01 01:0x): /home +HIT(2026-09-01 01:1x): /home +HIT(2026-09-01 01:2x): /home +HIT(2026-09-01 01:3x): /home +HIT(2026-09-01 01:4x): /home +HIT(2026-09-01 01:5x): /home +.HIT(2026-09-01 02:0x): /home +HIT(2026-09-01 02:1x): /home +HIT(2026-09-01 02:2x): /home +HIT(2026-09-01 02:3x): /home +HIT(2026-09-01 02:4x): /home +HIT(2026-09-01 02:5x): /home +.HIT(2026-09-01 03:0x): /home +HIT(2026-09-01 03:1x): /home +HIT(2026-09-01 03:2x): /home +HIT(2026-09-01 03:3x): /home +HIT(2026-09-01 03:4x): /home +HIT(2026-09-01 03:5x): /home +.HIT(2026-09-01 04:0x): /home +HIT(2026-09-01 04:1x): /home +HIT(2026-09-01 04:2x): /home +HIT(2026-09-01 04:3x): /home +HIT(2026-09-01 04:4x): /home +HIT(2026-09-01 04:5x): /home +.HIT(2026-09-01 05:0x): /home +HIT(2026-09-01 05:1x): /home +HIT(2026-09-01 05:2x): /home +HIT(2026-09-01 05:3x): /home +HIT(2026-09-01 05:4x): /home +HIT(2026-09-01 05:5x): /home +.HIT(2026-09-01 06:0x): /home +HIT(2026-09-01 06:1x): /home +HIT(2026-09-01 06:2x): /home +HIT(2026-09-01 06:3x): /home +HIT(2026-09-01 06:4x): /home +HIT(2026-09-01 06:5x): /home +.HIT(2026-09-01 07:0x): /home +HIT(2026-09-01 07:1x): /home +HIT(2026-09-01 07:2x): /home +HIT(2026-09-01 07:3x): /home +HIT(2026-09-01 07:4x): /home +HIT(2026-09-01 07:5x): /home +.HIT(2026-09-01 08:0x): /home +HIT(2026-09-01 08:1x): /home +HIT(2026-09-01 08:2x): /home +HIT(2026-09-01 08:3x): /home +HIT(2026-09-01 08:4x): /home +HIT(2026-09-01 08:5x): /home +.HIT(2026-09-01 09:0x): /home +HIT(2026-09-01 09:1x): /home +HIT(2026-09-01 09:2x): /home +HIT(2026-09-01 09:3x): /home +HIT(2026-09-01 09:4x): /home +HIT(2026-09-01 09:5x): /home +.HIT(2026-09-01 10:0x): /home +HIT(2026-09-01 10:1x): /home +HIT(2026-09-01 10:2x): /home +HIT(2026-09-01 10:3x): /home +HIT(2026-09-01 10:4x): /home +HIT(2026-09-01 10:5x): /home +.HIT(2026-09-01 11:0x): /home +HIT(2026-09-01 11:1x): /home +HIT(2026-09-01 11:2x): /home +HIT(2026-09-01 11:3x): /home +HIT(2026-09-01 11:4x): /home +HIT(2026-09-01 11:5x): /home +.HIT(2026-09-01 12:0x): /home +HIT(2026-09-01 12:1x): /home +HIT(2026-09-01 12:2x): /home +HIT(2026-09-01 12:3x): /home +HIT(2026-09-01 12:4x): /home +HIT(2026-09-01 12:5x): /home +.HIT(2026-09-01 13:0x): /home +HIT(2026-09-01 13:1x): /home +HIT(2026-09-01 13:2x): /home +HIT(2026-09-01 13:3x): /home +HIT(2026-09-01 13:4x): /home +HIT(2026-09-01 13:5x): /home +.HIT(2026-09-01 14:0x): /home +HIT(2026-09-01 14:1x): /home +HIT(2026-09-01 14:2x): /home +HIT(2026-09-01 14:3x): /home +HIT(2026-09-01 14:4x): /home +HIT(2026-09-01 14:5x): /home +.HIT(2026-09-01 15:0x): /home +HIT(2026-09-01 15:1x): /home +HIT(2026-09-01 15:2x): /home +HIT(2026-09-01 15:3x): /home +HIT(2026-09-01 15:4x): /home +HIT(2026-09-01 15:5x): /home +.HIT(2026-09-01 16:0x): /home +HIT(2026-09-01 16:1x): /home +HIT(2026-09-01 16:2x): /home +HIT(2026-09-01 16:3x): /home +HIT(2026-09-01 16:4x): /home +HIT(2026-09-01 16:5x): /home +.HIT(2026-09-01 17:0x): /home +HIT(2026-09-01 17:1x): /home +HIT(2026-09-01 17:2x): /home +HIT(2026-09-01 17:3x): /home +HIT(2026-09-01 17:4x): /home +HIT(2026-09-01 17:5x): /home +.HIT(2026-09-01 18:0x): /home +HIT(2026-09-01 18:1x): /home +HIT(2026-09-01 18:2x): /home +HIT(2026-09-01 18:3x): /home +HIT(2026-09-01 18:4x): /home +HIT(2026-09-01 18:5x): /home +.HIT(2026-09-01 19:0x): /home +HIT(2026-09-01 19:1x): /home +HIT(2026-09-01 19:2x): /home +HIT(2026-09-01 19:3x): /home +HIT(2026-09-01 19:4x): /home +HIT(2026-09-01 19:5x): /home +.HIT(2026-09-01 20:0x): /home +HIT(2026-09-01 20:1x): /home +HIT(2026-09-01 20:2x): /home +HIT(2026-09-01 20:3x): /home +HIT(2026-09-01 20:4x): /home +HIT(2026-09-01 20:5x): /home +.HIT(2026-09-01 21:0x): /home +HIT(2026-09-01 21:1x): /home +HIT(2026-09-01 21:2x): /home +HIT(2026-09-01 21:3x): /home +HIT(2026-09-01 21:4x): /home +HIT(2026-09-01 21:5x): /home +.HIT(2026-09-01 22:0x): /home +HIT(2026-09-01 22:1x): /home +HIT(2026-09-01 22:2x): /home +HIT(2026-09-01 22:3x): /home +HIT(2026-09-01 22:4x): /home +HIT(2026-09-01 22:5x): /home +.HIT(2026-09-01 23:0x): /home +HIT(2026-09-01 23:1x): /home +HIT(2026-09-01 23:2x): /home +HIT(2026-09-01 23:3x): /home +HIT(2026-09-01 23:4x): /home +HIT(2026-09-01 23:5x): /home +. day 2026-09-01 done, batches-with-hits=144 + +### REMAINING FIVE DAYS — no control injected +........................ day 2026-08-30 done, batches-with-hits=0 +........................ day 2026-08-29 done, batches-with-hits=0 +........................ day 2026-08-28 done, batches-with-hits=0 +........................ day 2026-08-27 done, batches-with-hits=0 +........................ day 2026-08-26 done, batches-with-hits=0 +.......................### REMAINING DAYS, full 24h each, second granularity +day 2026-08-30: hits=0 +day 2026-08-29: hits=0 +day 2026-08-28: hits=0 +day 2026-08-27: hits=0 +day 2026-08-26: hits=0 +day 2026-08-25: hits=0 +day 2026-08-24: hits=0 +### done diff --git a/documentation/audits/evidence-drill-r95-recovery-2026-09-01/phase1-find-the-snapshot/07-sweep-control.txt b/documentation/audits/evidence-drill-r95-recovery-2026-09-01/phase1-find-the-snapshot/07-sweep-control.txt new file mode 100644 index 00000000..3b82400a --- /dev/null +++ b/documentation/audits/evidence-drill-r95-recovery-2026-09-01/phase1-find-the-snapshot/07-sweep-control.txt @@ -0,0 +1,3 @@ +### CONTROL: exactly the sweep's batch shape (600 fake names) + 1 real path appended. +### If the sweep can see a directory that exists, this MUST print /home. 6 batches tried. +control batches: 6, batches that correctly returned exactly /home: 6 diff --git a/documentation/audits/evidence-drill-r95-recovery-2026-09-01/phase1-find-the-snapshot/08-alt-formats.txt b/documentation/audits/evidence-drill-r95-recovery-2026-09-01/phase1-find-the-snapshot/08-alt-formats.txt new file mode 100644 index 00000000..98bfbde5 --- /dev/null +++ b/documentation/audits/evidence-drill-r95-recovery-2026-09-01/phase1-find-the-snapshot/08-alt-formats.txt @@ -0,0 +1,4 @@ +candidates: 126 +=== HITS === +/.zfs/shares/. +/home diff --git a/documentation/audits/evidence-drill-r95-recovery-2026-09-01/phase1-find-the-snapshot/09-rclone-restic-lead.txt b/documentation/audits/evidence-drill-r95-recovery-2026-09-01/phase1-find-the-snapshot/09-rclone-restic-lead.txt new file mode 100644 index 00000000..218a43db --- /dev/null +++ b/documentation/audits/evidence-drill-r95-recovery-2026-09-01/phase1-find-the-snapshot/09-rclone-restic-lead.txt @@ -0,0 +1,7 @@ +=== is rclone in the controller container? === +NO rclone in container +=== does restic 0.14.0 know the rclone backend? (control: banana:) === +Fatal: parsing repository location failed: invalid backend +If the repository is in a local directory, you need to add a `local:` prefix +--- rclone: backend --- +Fatal: unable to open repository at rclone:nosuchremote:/x: exec: "rclone": executable file not found in $PATH diff --git a/documentation/audits/evidence-drill-r95-recovery-2026-09-01/phase2-the-attack/NOT-RUN.md b/documentation/audits/evidence-drill-r95-recovery-2026-09-01/phase2-the-attack/NOT-RUN.md new file mode 100644 index 00000000..59aab83d --- /dev/null +++ b/documentation/audits/evidence-drill-r95-recovery-2026-09-01/phase2-the-attack/NOT-RUN.md @@ -0,0 +1,13 @@ +# phase2-the-attack — NOT RUN + +The drill stopped at the end of Phase 1 on the operator's ruling. **This directory is empty on +purpose, not by omission.** + +Phase 1 established that no Storage Box snapshot is reachable from the box by any name (R-433), and +that every other route to one is fenced (the Hetzner API, §11-D) or unavailable (a browser and a +main-account credential nobody here holds). With no recovery leg, the deletion in Phase 2 would have +destroyed real off-site history to buy only an alarm test that could not have fired at the size the +runbook specifies — the detector needs a fall of more than half of 69, and one app's tag is ~9 +(R-435). Put to the operator as a two-option decision; the ruling was stop and report. + +See `../README.md` and `felhom.eu/REPORT.md`. diff --git a/documentation/audits/evidence-drill-r95-recovery-2026-09-01/phase3-the-alarm/NOT-RUN.md b/documentation/audits/evidence-drill-r95-recovery-2026-09-01/phase3-the-alarm/NOT-RUN.md new file mode 100644 index 00000000..e0fbe189 --- /dev/null +++ b/documentation/audits/evidence-drill-r95-recovery-2026-09-01/phase3-the-alarm/NOT-RUN.md @@ -0,0 +1,13 @@ +# phase3-the-alarm — NOT RUN + +The drill stopped at the end of Phase 1 on the operator's ruling. **This directory is empty on +purpose, not by omission.** + +Phase 1 established that no Storage Box snapshot is reachable from the box by any name (R-433), and +that every other route to one is fenced (the Hetzner API, §11-D) or unavailable (a browser and a +main-account credential nobody here holds). With no recovery leg, the deletion in Phase 2 would have +destroyed real off-site history to buy only an alarm test that could not have fired at the size the +runbook specifies — the detector needs a fall of more than half of 69, and one app's tag is ~9 +(R-435). Put to the operator as a two-option decision; the ruling was stop and report. + +See `../README.md` and `felhom.eu/REPORT.md`. diff --git a/documentation/audits/evidence-drill-r95-recovery-2026-09-01/phase4-the-recovery/NOT-RUN.md b/documentation/audits/evidence-drill-r95-recovery-2026-09-01/phase4-the-recovery/NOT-RUN.md new file mode 100644 index 00000000..34800998 --- /dev/null +++ b/documentation/audits/evidence-drill-r95-recovery-2026-09-01/phase4-the-recovery/NOT-RUN.md @@ -0,0 +1,13 @@ +# phase4-the-recovery — NOT RUN + +The drill stopped at the end of Phase 1 on the operator's ruling. **This directory is empty on +purpose, not by omission.** + +Phase 1 established that no Storage Box snapshot is reachable from the box by any name (R-433), and +that every other route to one is fenced (the Hetzner API, §11-D) or unavailable (a browser and a +main-account credential nobody here holds). With no recovery leg, the deletion in Phase 2 would have +destroyed real off-site history to buy only an alarm test that could not have fired at the size the +runbook specifies — the detector needs a fall of more than half of 69, and one app's tag is ~9 +(R-435). Put to the operator as a two-option decision; the ruling was stop and report. + +See `../README.md` and `felhom.eu/REPORT.md`. diff --git a/documentation/audits/evidence-drill-r95-recovery-2026-09-01/phase5-rearm/NOT-RUN.md b/documentation/audits/evidence-drill-r95-recovery-2026-09-01/phase5-rearm/NOT-RUN.md new file mode 100644 index 00000000..846bfda0 --- /dev/null +++ b/documentation/audits/evidence-drill-r95-recovery-2026-09-01/phase5-rearm/NOT-RUN.md @@ -0,0 +1,13 @@ +# phase5-rearm — NOT RUN + +The drill stopped at the end of Phase 1 on the operator's ruling. **This directory is empty on +purpose, not by omission.** + +Phase 1 established that no Storage Box snapshot is reachable from the box by any name (R-433), and +that every other route to one is fenced (the Hetzner API, §11-D) or unavailable (a browser and a +main-account credential nobody here holds). With no recovery leg, the deletion in Phase 2 would have +destroyed real off-site history to buy only an alarm test that could not have fired at the size the +runbook specifies — the detector needs a fall of more than half of 69, and one app's tag is ~9 +(R-435). Put to the operator as a two-option decision; the ruling was stop and report. + +See `../README.md` and `felhom.eu/REPORT.md`. diff --git a/documentation/audits/evidence-drill-r95-recovery-2026-09-01/phase6-teardown/01-verify-then-clean.txt b/documentation/audits/evidence-drill-r95-recovery-2026-09-01/phase6-teardown/01-verify-then-clean.txt new file mode 100644 index 00000000..8f686931 --- /dev/null +++ b/documentation/audits/evidence-drill-r95-recovery-2026-09-01/phase6-teardown/01-verify-then-clean.txt @@ -0,0 +1,23 @@ +=== A. off-site inventory UNCHANGED? (must still be 69) === + bookstack +----------------------------------------------------- +69 snapshots + +=== B. account home contents (must be exactly .ssh + felhom-repo, nothing of mine) === +. +.. +.ssh +felhom-repo + +=== C. repo top level (untouched) === +config +data +index +keys +locks +snapshots + +=== D. LAYER 1 CLEAN: ssh control sockets inside the container === +/tmp/r95-cm4-u629488-sub3 +removed +verified: no control sockets left in container diff --git a/documentation/audits/evidence-drill-r95-recovery-2026-09-01/phase6-teardown/02-guest-and-host-clean.txt b/documentation/audits/evidence-drill-r95-recovery-2026-09-01/phase6-teardown/02-guest-and-host-clean.txt new file mode 100644 index 00000000..87e1d08e --- /dev/null +++ b/documentation/audits/evidence-drill-r95-recovery-2026-09-01/phase6-teardown/02-guest-and-host-clean.txt @@ -0,0 +1,22 @@ +=== LAYER 2 (guest 9201) + LAYER 3 (demo-hp PVE host) cleanup === +-- guest scratch before: +. +.. +cands.txt +cmd.sh +env.sh +probe.txt +sweep1.txt +-- guest scratch after: +ls: cannot access '/mnt/r95drill': No such file or directory +-- guest /mnt now: +. +.. +felhom-drives +sys_drive +-- host /tmp/r95cmd.sh after: +ls: cannot access '/tmp/r95cmd.sh': No such file or directory +-- controller container healthy: +felhom-controller gitea.dooplex.hu/admin/felhom-controller:0.232.0 Up 6 hours (healthy) +-- app containers healthy count: +19 diff --git a/documentation/audits/evidence-drill-r95-recovery-2026-09-01/phase6-teardown/03-dooplex-clean-and-final-state.txt b/documentation/audits/evidence-drill-r95-recovery-2026-09-01/phase6-teardown/03-dooplex-clean-and-final-state.txt new file mode 100644 index 00000000..087442f5 --- /dev/null +++ b/documentation/audits/evidence-drill-r95-recovery-2026-09-01/phase6-teardown/03-dooplex-clean-and-final-state.txt @@ -0,0 +1,12 @@ +=== LAYER 4 (DooPlex) — shred the hub DB copy (it holds every host's secret) === +-rw-rw-r-- 1 kisfenyo kisfenyo 236482560 Sep 1 16:45 /tmp/claude-1000/hubdb.db +ls: cannot access '/tmp/claude-1000/hubdb.db*': No such file or directory +verified: hub DB copies gone + +=== demo-felhom NOT TOUCHED — positive check, its own tier still reporting ok === +demo-hp: report 22221 @2026-09-01 14:45:18 snapshot_count=69 stats_known=True last_status=ok last_success=2026-09-01T02:17:57Z +demo-felhom: report 22222 @2026-09-01 14:47:19 snapshot_count=10 stats_known=True last_status=ok last_success=2026-09-01T08:32:13Z + +--- any offsite_snapshots_dropped events raised today? (must be only the constructed one) --- +(3368, 'demo-hp', '2026-09-01 12:29:06', 'offsite_snapshots_dropped') +second copy shredded diff --git a/documentation/backlog/OPEN-ITEMS.md b/documentation/backlog/OPEN-ITEMS.md index f8bffe41..43348c12 100644 --- a/documentation/backlog/OPEN-ITEMS.md +++ b/documentation/backlog/OPEN-ITEMS.md @@ -224,7 +224,7 @@ unconditional promise, which CAMPAIGN-11 §7 step 7 measured the customer-facing | **E-2d** | **Prove E-2 on a fresh VM** — a real `felhom-host-install.sh` 1.22.0 run, Case B naturally, a claimable customer, then add a drive (the offer) and unplug it (`backup_target_absent` end-to-end) | **CLOSED — PARTIALLY PROVEN** (2026-07-29) | — | **C1, C2 proven** (`audits/E2D-fresh-vm-2026-07-29.md`); **C3, C4 proven live** (`audits/SESSION-C-2026-07-29.md`); **C5 FAILED → R-116** — the gate fires and an alarm reaches the hub, but it is the generic event, so the alarm and its recovery cannot be paired. **R-116 is the single named open leg**; per the Session-C runbook §9, decided in advance, a failed claim closes the item as partially proven rather than triggering a re-run. Both audits carry the full record — the `local-lvm` fence, the ISO/PAIRING derivation, the Phase 0 answers, the per-claim observables and the teardown evidence — and are the place to read it, not this cell. **The arc's actual definition of done is R-106 + R-109, R-108 and D5**, none of which this detour touched | CC | | **R-121** | **A BOX's installed agent can sit releases behind the vouched one and nothing notices — the R-120 gate does not cover it.** demo-hp ran agent **0.113.0** while the hub vouched **0.116.0**, through the whole R-116/R-117 arc, and no signal existed on any channel | **READY (S) — NEW 2026-07-30** | — | **Fourth instance of the drift family** (R-111 golden's agent 17 releases behind, R-115 built+deployed but never published, R-120 golden a controller behind — and now installed-vs-vouched on a live box). **Confirmed at source that R-120's gate cannot catch it:** `hub/internal/web/configs.go:1165-1169` compares `goldenVer` against `store.NewestReportedControllerVersion()` — it is a **golden-artifact vs fleet-CONTROLLER** check and says nothing about the agent installed on a box. **`MinAgent` does not cover it either:** it is used to HOLD the controller floor for a box whose agent is too old (`hub/internal/api/handler.go:530-538`, `store.go:1857`) — protective, not an alarm — and demo-hp's 0.113.0 **equalled** `min_agent` 0.113.0, so even a floor comparison was satisfied. **The cost, measured:** R-117's whole subject is the R-113 conjunction, which landed in **0.114.0** — so the designated drill host could not exercise the code under investigation at all, and the R-117 spike had to route every predicate result through an out-of-repo probe built from `main` instead of the installed agent (`audits/SPIKE-r117-bind-liveness-2026-07-30.md` §1, §2.3). Discovered because the R-117 task made bringing the box current an explicit prerequisite. **Fix shape (not implemented):** the hub already receives `AgentVersion` on every host report, and already has semver comparison in Go — the missing piece is a checker comparing reported agent vs the **vouched** agent and surfacing it, operator-tier. Note the honest tension: a box legitimately lags between publish and deploy, so this wants a staleness window rather than an instant alarm | CC | | **R-118** | **An absent drive's union row advertises the ROOT filesystem's capacity as its own.** In the absent-state payload the registry-union row reports `total_bytes: 49675956224 / used_bytes: 4584579072` — **byte-identical to the `local` row** (`durable_id: path:/var/lib/vz`, i.e. `pve-root`) in the same response. The real drive is **4 GB** | **READY (XS) — NEW 2026-07-30** | — | Cause: `statfsCapacity(d.MountPath)` (`disks.go:335-338`) statfs's `/mnt/cel`, which with the device gone is a **bare directory on the root filesystem**. `observe.go:176-183`'s comment warns about exactly this trap and guards the Observe path ("*an unmounted removable dir-storage's mountpoint reverts to a bare directory on root … catastrophic DR mis-id*"); **the union path has no equivalent guard.** **Not a DR mis-id** — `durable_id` on that row is still the correct `uuid:…`, so re-attach identity is safe. It is a **false capacity** reaching every consumer of `total_bytes`/`used_fraction` (fill monitors, storage cards): a detached 4 GB drive advertises 46 GiB at 9.2 % used. Same class as `role.go:180-181` — an absent drive's fields decaying to the root filesystem's. Evidence: `audits/DIAG-r116-disks-payload-2026-07-30.md` §12 | CC | -| **R-95** | restic offsite credential **can delete** (`readonly=False`, `forget --prune` runs from the box); SFTP cannot express append-only | **READY** | — | Root exposure still open. Mitigation now ARMED — split prune off-box or move to REST `--append-only` | CC **SPIKE 2026-09-01 — `audits/SPIKE-r95-offsite-delete-2026-09-01.md`. THE WORD "ARMED" ABOVE IS NOT SUPPORTED AND IS WITHDRAWN PENDING R-429:** no `.snapshots` is visible to either box's sub-account (measured, both machines, with controls), so the seven-day bound is unverified and unverifiable from the product side. **Q3 (documented, `hub/internal/hetznerapi/hetznerapi.go:38-45`): the sub-account API has ONE permission axis, `readonly` — there is no append-only, so the PBS shape does NOT transfer** (PBS is a server that can refuse; a Storage Box is a filesystem that runs nothing). **Q5 (measured): withdrawing delete does NOT wedge the store** — restic treats a dead owner's lock as stale and proceeds — so the constraint everyone feared is not the blocker; but `unlock --remove-all` lies about success (R-430), and the crash-lock window is UNKNOWN. **Q6 (measured): restic 0.14.0 DOES speak `rest:`** (control: `banana:` → invalid backend), and append-only is a rest-server flag, not a restic one — reachable, but it needs a machine in the recovery path and ep0 is protected. **Q7 (measured): detection is nearly free** — `snapshot_count` already reaches the hub and the hub APPENDS reports, so the comparison needs no box change. **RECOMMENDATION: answer R-429 first (Viktor, ten minutes), then build detection, then move retention off the box; defer the transport change.** Two `forget --prune` sites must be disarmed together — `offbox.go:1388` AND `offbox.go:1759` — or R-191 repeats. **RE-SCOPED 2026-09-01 — THE STORY WAS WORSE THAN THE TRUTH FOR TWO MONTHS.** The box can delete its own LIVE repository, **but it cannot write to the daily snapshots of it** — MEASURED, not cited: a write into `/.zfs/snapshot` is refused on both boxes while the same write to the account home succeeds (R-432). Seven daily snapshots are confirmed in the panel (R-429). **So a deletion costs at most the data written since the last daily snapshot, and the rest is recoverable — file by file, one customer at a time, with no effect on anyone else** (vendor: *"You can download individual files or entire directories as usual"*; *"It is not possible to write to the `/.zfs` directory or its subfolder"*). **NOT open-ended loss.** Two caveats kept honest: a panel-driven snapshot restore rolls back the WHOLE Storage Box and deletes newer snapshots, which is why the per-file route matters; and per-file recovery is **operator-only today** (R-432). **DETECTION SHIPPED hub v0.111.0 (R-431)** — an unexplained fall is noticed within a day. **THE RANKING IS VIKTOR'S:** this has been #1 since July on the old story. On the new facts I would rank it below the items that can still lose data outright, but I am not re-ranking it myself. `audits/SPIKE-r95-offsite-delete-2026-09-01.md` | +| **R-95** | restic offsite credential **can delete** (`readonly=False`, `forget --prune` runs from the box); SFTP cannot express append-only | **READY** | — | Root exposure still open. Mitigation now ARMED — split prune off-box or move to REST `--append-only` | CC **SPIKE 2026-09-01 — `audits/SPIKE-r95-offsite-delete-2026-09-01.md`. THE WORD "ARMED" ABOVE IS NOT SUPPORTED AND IS WITHDRAWN PENDING R-429:** no `.snapshots` is visible to either box's sub-account (measured, both machines, with controls), so the seven-day bound is unverified and unverifiable from the product side. **Q3 (documented, `hub/internal/hetznerapi/hetznerapi.go:38-45`): the sub-account API has ONE permission axis, `readonly` — there is no append-only, so the PBS shape does NOT transfer** (PBS is a server that can refuse; a Storage Box is a filesystem that runs nothing). **Q5 (measured): withdrawing delete does NOT wedge the store** — restic treats a dead owner's lock as stale and proceeds — so the constraint everyone feared is not the blocker; but `unlock --remove-all` lies about success (R-430), and the crash-lock window is UNKNOWN. **Q6 (measured): restic 0.14.0 DOES speak `rest:`** (control: `banana:` → invalid backend), and append-only is a rest-server flag, not a restic one — reachable, but it needs a machine in the recovery path and ep0 is protected. **Q7 (measured): detection is nearly free** — `snapshot_count` already reaches the hub and the hub APPENDS reports, so the comparison needs no box change. **RECOMMENDATION: answer R-429 first (Viktor, ten minutes), then build detection, then move retention off the box; defer the transport change.** Two `forget --prune` sites must be disarmed together — `offbox.go:1388` AND `offbox.go:1759` — or R-191 repeats. **RE-SCOPED 2026-09-01 — THE STORY WAS WORSE THAN THE TRUTH FOR TWO MONTHS.** The box can delete its own LIVE repository, **but it cannot write to the daily snapshots of it** — MEASURED, not cited: a write into `/.zfs/snapshot` is refused on both boxes while the same write to the account home succeeds (R-432). Seven daily snapshots are confirmed in the panel (R-429). **So a deletion costs at most the data written since the last daily snapshot, and the rest is recoverable — file by file, one customer at a time, with no effect on anyone else** (vendor: *"You can download individual files or entire directories as usual"*; *"It is not possible to write to the `/.zfs` directory or its subfolder"*). **NOT open-ended loss.** Two caveats kept honest: a panel-driven snapshot restore rolls back the WHOLE Storage Box and deletes newer snapshots, which is why the per-file route matters; and per-file recovery is **operator-only today** (R-432). **DETECTION SHIPPED hub v0.111.0 (R-431)** — an unexplained fall is noticed within a day. **THE RANKING IS VIKTOR'S:** this has been #1 since July on the old story. On the new facts I would rank it below the items that can still lose data outright, but I am not re-ranking it myself. `audits/SPIKE-r95-offsite-delete-2026-09-01.md` **DRILL 2026-09-01, LATER THE SAME DAY — THE RE-SCOPE'S SECOND HALF IS WITHDRAWN.** The drill that was to walk the recovery found there is no route to walk: **no snapshot is reachable from a sub-account by ANY name** (R-433) — 777,600 exact names in the vendor format over nine days, zero hits, with a passing control, plus the structural reason (`/home` st_dev 0,82 vs `/.zfs/snapshot` st_dev 0,276, and `/home/.zfs` absent). **Clause (a) stands: the box can delete the live repo and cannot write into the snapshot area.** Clause (b) — *"the rest is recoverable file by file"* — is NOT SUPPORTED. **The drill was STOPPED before its destructive phase on the operator's ruling**, because with no recovery route the deletion would have destroyed real history to buy only an alarm test that could not fire at the specified size (R-435). Nothing was deleted; the store is verified untouched at 69 snapshots. **So the comfort that lowered this row rested on an unwalked route, and the route does not exist.** Two new leads decide what happens next: **R-433** (can the MAIN account see them? nobody here holds that credential) and **R-436** (`rclone serve restic --stdio` is offered server-side and restic speaks `rclone:` — measured — which could make real prevention cheap, IF the provider pins `--append-only`). **The rank stays Viktor's. Plainly: the argument that moved this row down is the argument the drill removed.** `audits/evidence-drill-r95-recovery-2026-09-01/` | | **R-191** | **Every weekly offsite backup UPLOADS successfully and then FAILS the job on a prune the box is deliberately not allowed to do — on both demo boxes.** Measured on demo-felhom 2026-08-04 06:49–06:53: the upload completed (223 s, 629 MiB of 1.874 GiB, **67.2 % reused incrementally**), then `ERROR: prune 'ct/9201': proxmox-backup-client failed: Error: permission check failed - missing Datastore.Modify\|Datastore.Prune on /datastore/felhom-offsite/demo-felhom` → `ERROR: Backup of VM 9201 failed - error pruning backups` → `TASK ERROR: job errors`. The hub raised `whole_guest_backup_failed` | **CLOSED — SHIPPED 2026-08-04** (installer **1.25.0**; both live boxes corrected) | — | **This is R-89's rule not reaching the config.** R-89 moved PBS pruning SERVER-SIDE — *"boxes set `keep_last: 0`, ep0 runs prune jobs; box tokens stay write-only, never widen the grant"*. The token behaves exactly as designed: it refuses. But **both** demo boxes still arm the offsite tier with `keep_last=2 prune_pbs_allowed=true` (`backup_targets: [{target_id: felhom-pbs, cadence_seconds: 604800, keep_last: 2}]`), so every run asks for a prune that must fail. **The data is SAFE and that is why this is not a P1:** the snapshot lands before the prune is attempted; what is wrong is the job's VERDICT and the weekly operator e-mail it produces. **But it is corrosive in the specific way this project keeps finding:** a backup that reports FAILED while succeeding trains the operator to discount `whole_guest_backup_failed`, which is the same alert that would carry a real one — and it is exactly the failure the R-100 corollary warns about, an alarm whose text is true and whose trigger is not the thing you would act on. **Fix is one config line per box** (`keep_last: 0` on the PBS tier) plus whatever writes it on a fresh install; **deliberately NOT applied in this session** — the session was a runbook with an explicit "change nothing, and if a change appears necessary, stop and report" rule, and a retention field on a live backup tier is not a change to slip into an observation run. **Check before fixing:** whether ep0's prune jobs actually cover these two namespaces, or the snapshots simply accumulate once the box stops asking **THE GATE WAS RUN FIRST, AND IT MATTERED.** Before disabling anything, ep0 was read (read-only, Tier 2): prune jobs `prune-demo-felhom` and `prune-demo-hp` exist on datastore `felhom-offsite`, one per namespace, `schedule 03:30`, `keep-last 2`, comment *"R-82 retention keep-last=2, server-side (box tokens are write-only)"* — and they have run **every day since 2026-07-27: 18 tasks, all `status=OK`**. The newest task log reads `retention options: --ns demo-felhom --max-depth 0 --keep-last 2` / `keep ct/9201/2026-07-27…` / `keep ct/9201/2026-07-28…` / `TASK OK`. Retention happens, and it happens there. **A METHODOLOGICAL WARNING WORTH MORE THAN THE FIX.** Three separate queries said the OPPOSITE — *no prune jobs have ever run* — and **all three were broken instruments**: `worker-type` where the field is `worker_type`; the value `prune` where the worker type is `prunejob`; and `journalctl -u proxmox-backup` where the unit is `proxmox-backup-proxy`. A fourth reading (3 snapshots under keep-last 2) was mis-framed by CC and self-corrected — the third snapshot had landed AFTER that day's 03:30 window. Acting on any of them would have disabled the only pruning ATTEMPT while reporting that nothing prunes: a weekly false alarm traded for unbounded growth on the protected endpoint, invisible for months. **The gate is what caught it, and only because it demanded evidence rather than a verdict.** **Shipped:** installer **1.25.0** writes `keep_last: 0` on the offsite tier (the agent's existing guard `allowPBSPrune = !primary && keep_last > 0` already reads that as *never prune from the box* — no agent change), the justifying paragraph is rewritten to say where retention lives and cite R-89, and `hostinstall_gates.py` asserts it (red-proved: pinning `keep_last: 2` back fails the gate). **Both live boxes corrected in their own config** — `backup tier armed target=felhom-pbs … keep_last=0 … prune_pbs_allowed=false` on demo-felhom and demo-hp, with the local tier untouched at `keep_last=3`. Served over HTTPS at `1.25.0` with `"keep_last":0` in the served bytes. **STILL TO OBSERVE:** the next weekly offsite run completing OK end-to-end. The change removes the failing step; the *schedule* proving it is next week's event, and this row should carry that line when it happens. | CC | | **R-194** | **PVE's permission cache delays every grant-state verdict by an unknown amount, so "the agent can read it" and "the ACL exists" are not the same measurement.** Observed twice while validating R-190's self-repair on demo-felhom 2026-08-04: both ACL rows for `/storage/felhom-backup` were deleted, and `GET /access/permissions` continued to report `Datastore.AllocateSpace` present — for **~40 s** in one run and **~16 minutes** in another. During that window the capability probe reads healthy and the self-repair does not fire | **OPEN** | — | **Why it matters beyond the delay:** it puts a floor under how fast a lost grant can be noticed, it makes any single permission read a lagging indicator, and — the interesting part — **it is a candidate contributor to R-190's own timeline**: a grant removed at an unknown moment could keep working until a cache expiry, which is exactly the shape of *worked at 04:44, refused at 09:24*. That does not explain what removed it, but it may explain when the refusal SURFACED, and the two have been treated as the same instant. **Not a defect in our code** — it is PVE behaviour, and the mitigation already tolerates it (the repair fires on the next probe after the cache clears). **What is worth deciding:** whether the store-grant probe should ALSO consult the storage content listing as a second signal, since that appeared to reflect the loss immediately (`{"data":[]}` while the permission read still said present) — two signals disagreeing is itself information, and today only one of them is read | CC | | **R-200** | **The DR password-injection seam has a handler, a route and tests — and no form.** `POST /backup/offbox/inject-password` is routed (`controller/internal/web/server.go:510`) to `offboxInjectPasswordHandler` (`offbox_handlers.go:174-196`) → `InjectOffboxPassword` (`backup/offbox.go:541`). **No template in the repository contains that path or any form posting to it** (grep over `internal/web/templates/`: one unrelated hit, an XSS comment) | **PLUMBING COMPLETE** (controller **v0.196.0**); **the FORM is not built — still open** | — | **The tenth instance of this project's built-but-never-wired class, and the exact shape `CLAUDE.md` and `felhom.eu/CLAUDE.md`'s seam-wiring rule were written for:** handler tests that POST directly (`offbox_escrow_test.go:167,180`) prove nothing about reachability. To use the only implemented recovery seam today, a person must hand-craft an authenticated POST with a session cookie and CSRF token. **Note the layering while fixing it:** this form takes a **64-hex repo password**, not a recovery code (`offboxRepoPwPattern`, `offbox.go:543`) — they are different secrets at different layers, and the operator's 2026-08-04 ruling asks for a form that takes **R**. Build the R form and treat this one as the operator/DR fallback it was written as, but **ship it with a render test per branch of whatever gate it sits behind**. Source: `audits/RECON-offsite-dr-chain-2026-08-04.md` §3 link 9 **THE DIAGNOSTIC HALF IS DONE AND IT ANSWERED THE QUESTION.** `--recover-offsite-check` is a `docker exec` escape hatch in the shape of `--print-reset-code`: R on **STDIN** (never argv, never `ps`, never shell history, never a transcript), fetch+unseal via the agent, and a verdict of **two sha256 hashes**. **It compares and never installs** — the recovered password is not written to `offbox/repo_password`; a test asserts the data dir is byte-unchanged and its red-proof (adding the install call) fails it. Confirmed live: `repo_password` mtime still `2026-08-03 07:18:02` after the successful check at `2026-08-04 11:49`. **Exit codes are load-bearing** — `0` match, `2` a clean MISMATCH, `1` a step failed; "it failed" and "it worked and disagreed" must never share a status because only one is a finding about the system. A box with no local password reports distinctly (the rebuilt-box shape, where the next step is to INSTALL rather than compare). **WHAT IS NOT BUILT, deliberately:** no card, no form, no preview, no customer-facing text — building an interface on top of a chain nobody had walked is how the preceding three weeks went wrong. **What remains for this row:** link 9 (the recovered password placed so `WriteOffboxSecrets` keeps it) and the customer-facing shape the operator ruled on 2026-08-04 (yell → R form → preview → proceed), which is now priced against a chain that exists rather than one that is assumed **PART 0 SHIPPED 2026-08-04 (v0.196.0):** `--recover-offsite-install` is the sibling of the check — same fetch/unseal path, same STDIN discipline for R — and it **places** the recovered password via `InjectOffboxPassword`. The confirmation is a SECOND invocation (`--confirm-install`): without it, both hashes print and nothing is written, so the operator sees the comparison before a write is possible. Three outcomes named distinctly: **installed** (no local password — the rebuilt-box shape), **unchanged** (identical key, nothing written), **refused** (a DIFFERENT key present — installing would clobber the key the current repository is encrypted under; exit 2, no force offered). It re-reads the file after writing rather than trusting the call. **Red-proof observed:** removing the confirmation gate makes the dry run write the password. The R-persistence test carries a **positive control** (a planted copy found, then removed and not found). **NOT YET EXERCISED AGAINST A LIVE RECOVERY** — the R-201 drill halted before step 9, so this is unit-proven only. **What remains for this row:** the customer-facing shape the operator ruled on 2026-08-04 (yell → recovery-code form → preview → proceed) | CC | @@ -603,7 +603,11 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server` | **R-427** | **`closed_register_gate.py` checks ONE direction only: an open word in a CLOSED row. The mirror — a CLOSED verdict on a row still sitting in `OPEN-ITEMS.md` — is unchecked, and there are TWELVE.** MEASURED 2026-09-01 during the decoy sweep, by reading the leading verdict of every open row with the gate's own predicate: **R-385, R-387, R-341, R-378, R-405, R-88a, R-88b** read unambiguously closed; **R-123, R-190, R-352** read `PARTLY CLOSED` / `MITIGATION SHIPPED` and almost certainly belong where they are. **The rows were NOT moved by this session** — telling a finished row from a partly-finished one is a judgement, and R-378 is itself the record of what happens when a machine makes that judgement on a substring (six still-open rows moved out of the register). **This is R-405's finding mirrored:** that row exists because R-87 sat in the CLOSED file while its state read READY, and the gate written for it looks only the way it was bitten. Fix: the same leading-verdict predicate applied to `OPEN-ITEMS.md`, reporting rather than convicting until the twelve are adjudicated by a person — a gate registered while twelve rows fail it would refuse every push. | **OPEN — 12 rows named; the adjudication is Viktor's, the gate is mine** | | **R-428** | **The decoy-coverage gate — written to catch instruments that match a NAME instead of a fact — identified a repository by its DIRECTORY NAME.** MEASURED on its own first CI run (felhom.eu job 490, 2026-09-01): `os.path.basename(root)` looked up in a `RUNNERS` map, and Gitea's act-runner checks the repo out into a directory called `hostexecutor`, so the gate reported *"unknown repo 'hostexecutor'"* and went INCONCLUSIVE. **The gate that hunts label-matching was matching a label, in the first ten lines of its own main loop, and it shipped that way.** FIXED the same day: it now identifies a repo by which registered runner FILE exists under the root, which is a fact. **Recorded rather than quietly patched because it is the strongest evidence in the sweep that this class is not a matter of carelessness** — it was written by a session that had spent the morning reading 29 gates for exactly this, with the four shapes on screen. Verified under a renamed directory before and after. | **CLOSED 2026-09-01 — fixed, and kept as the class's best example** | | **R-430** | **`restic unlock --remove-all` printed `successfully removed locks` while the lock was still there.** MEASURED 2026-09-01 in a throwaway local repo (no live store touched), under a faithful append-only model — a sticky locks directory owned by root holding a root-owned lock, restic run as `nobody`; both controls passed first (create allowed, delete refused). The command reported success, returned, and `ls` showed the lock present. **`resticStep`'s crash-lock self-heal is built directly on this call** (`felhom-controller/controller/internal/backup/offbox.go:~768`), and its licence to escalate rests on the escalation actually working. **A self-heal that cannot fail is a self-heal that cannot be trusted** — this is this project's *"exit codes that lie"* class, in the one path that runs unattended against the customer's off-site history. It is harmless TODAY because the credential can delete and the removal really happens; it becomes load-bearing the moment delete is withdrawn, which is what R-95 is about. **Not yet established:** whether restic reports success because it removed zero locks by design, or because it did not check. Settling it: read restic 0.14.0's unlock source, or re-run with `--verbose`. | **OPEN — precondition on any R-95 build** | -| **R-432** | **A customer's own sub-account can REACH the snapshot door and is REFUSED writes to it — but sees it EMPTY, so per-file recovery is not product-reachable.** MEASURED 2026-09-01 on BOTH live boxes, over the credential each already holds, with a positive and a negative control in the same run. **What is now PROVEN rather than cited:** `/.zfs` lists (`shares`, `snapshot`) from inside the jail; a write into `/.zfs/snapshot` is **REFUSED** — `dest open …: Failure` — while the identical write to the account home **succeeds** and was cleaned up. **That is the append-only property, measured, and it is the sentence the whole R-95 re-scope rests on.** **What is NOT available:** `/.zfs/snapshot` lists **empty** (link count 2) on both boxes, while the same Storage Box demonstrably holds seven snapshots — `storage-box-pool-1` IS `u629488` (`RUNBOOK-ep0-datastore-volume-2026-07-27.md:386`), the box these sub-accounts live on. So the contents are filtered from a sub-account. **CONSEQUENCE: recovery from a snapshot is an OPERATOR act in a browser, not something the product can drive** — which decides whether R-95's remedy can ever be customer-facing. **Cheapest next step, and it is Viktor's:** read one snapshot's name from the panel; a single `ls /.zfs/snapshot/` from a box then settles whether a named snapshot can be entered even though the directory does not list (ZFS allows exactly that). If it can, per-file recovery becomes product-reachable and this closes cheaply. | **OPEN — one panel read settles it** | +| **R-432** | **A customer's own sub-account can REACH the snapshot door and is REFUSED writes to it — but sees it EMPTY, so per-file recovery is not product-reachable.** MEASURED 2026-09-01 on BOTH live boxes, over the credential each already holds, with a positive and a negative control in the same run. **What is now PROVEN rather than cited:** `/.zfs` lists (`shares`, `snapshot`) from inside the jail; a write into `/.zfs/snapshot` is **REFUSED** — `dest open …: Failure` — while the identical write to the account home **succeeds** and was cleaned up. **That is the append-only property, measured, and it is the sentence the whole R-95 re-scope rests on.** **What is NOT available:** `/.zfs/snapshot` lists **empty** (link count 2) on both boxes, while the same Storage Box demonstrably holds seven snapshots — `storage-box-pool-1` IS `u629488` (`RUNBOOK-ep0-datastore-volume-2026-07-27.md:386`), the box these sub-accounts live on. So the contents are filtered from a sub-account. **CONSEQUENCE: recovery from a snapshot is an OPERATOR act in a browser, not something the product can drive** — which decides whether R-95's remedy can ever be customer-facing. **Cheapest next step, and it is Viktor's:** read one snapshot's name from the panel; a single `ls /.zfs/snapshot/` from a box then settles whether a named snapshot can be entered even though the directory does not list (ZFS allows exactly that). If it can, per-file recovery becomes product-reachable and this closes cheaply. **ANSWERED 2026-09-01 (DRILL, `audits/evidence-drill-r95-recovery-2026-09-01/`) — NO, AND THE PANEL READ IS NOT NEEDED.** The named-entry hypothesis was tested exhaustively and fails: **777,600 exact names** in the vendor-documented format `YYYY-MM-DDTHH-MM-SS` (nine full days, second granularity) plus 126 alternative shapes, **zero hits**, with a control proving the identical batch shape returns a path that does exist (6/6). **And there is a structural reason:** `df` reports `u629488-sub3` mounted on `/home` at **st_dev 0,82** while `/.zfs/snapshot` is **st_dev 0,276** — a different filesystem — and `/home/.zfs` does not exist. A ZFS snapshot under `/.zfs/snapshot` belongs to the dataset owning that `.zfs`, not to the child at `/home`, so a correctly-named snapshot there **could not contain `felhom-repo`**. Three tools agree with controls in the same run (SFTP, the port-23 shell, `rsync --list-only`). The empty listing is not a display toggle hiding a reachable tree — from a sub-account there is no tree. **CONSEQUENCE: per-file recovery is not "operator-only", it is unreachable from the box entirely** → R-433. | **ANSWERED 2026-09-01 — negatively; the panel-read next step is WITHDRAWN as unnecessary** | +| **R-433** | **A sub-account cannot reach ANY Storage Box snapshot, by any name — so clause (b) of the 2026-09-01 R-95 re-scope ("the rest is recoverable file by file") is NOT SUPPORTED.** MEASURED 2026-09-01 on `demo-hp` over the credential the box already holds, read-only, no delete verb issued. **The sweep:** a batched `stat -c %n` over Hetzner's port-23 restricted shell (500–600 paths per round trip, stdout carrying only paths that exist) tried **777,600** names of the vendor form `YYYY-MM-DDTHH-MM-SS` across nine full days at second granularity, and 126 alternative shapes — **zero resolved.** **The control is what makes the zero mean anything:** the identical 600-name batch with one real path appended returned it in 6 of 6 batches. **The structural cause:** `/home` (the customer data, `u629488-sub3`) is **st_dev 0,82**; `/.zfs/snapshot` is **st_dev 0,276**; `/home/.zfs` does not exist. A snapshot under `/.zfs/snapshot` belongs to a different dataset than the one holding `felhom-repo`. **What still stands:** clause (a) — the box can delete its live repository but cannot WRITE into `/.zfs/snapshot` — is unchanged and re-confirmed. **What is now open again:** the only routes to the older copy are a panel rollback of the WHOLE Storage Box (deletes newer snapshots, hits every customer on it) and the provider API (fenced by §11-D, and `hub/internal/hetznerapi/hetznerapi.go` has **no snapshot method at all**, so it needs new code regardless). **NOT ESTABLISHED, and it is the question that decides whether per-file recovery exists for anyone:** whether the MAIN account can see the snapshots. No main-account credential exists in this project. **The ranking is Viktor's; I am not re-ranking R-95 — but the argument that moved it down is the argument this drill removed.** `audits/evidence-drill-r95-recovery-2026-09-01/` | **OPEN — decides R-95's remedy and its rank** | +| **R-434** | **The snapshot-drop alarm promises a recovery that cannot be performed.** `hub/internal/monitor/offsite.go` `emitSnapshotDrop` ships this text, live in hub **v0.111.0**: *"The daily Storage Box snapshots are read-only and still hold the older copy, **so this is recoverable file-by-file**; it is NOT confirmed data loss."* The first clause is true. The second is not reachable: not by the product (R-433), and not by the operator without a browser and a main-account credential that does not exist here. Its own comment states the intent — *"THE MESSAGE MUST NOT SAY THE DATA IS LOST, because after the 2026-09-01 measurement that is usually false"* — and the measurement it rests on was superseded the same day. **This is this project's own corollary landing on the alarm shipped that morning:** when a verdict changes which fact it counts from, the alarm text has to change with it, or the operator acts on a promise nobody can keep. **Fix is text-only and must not be made before R-433 settles what IS true** — an alarm rewritten twice in a week is worse than one rewritten once. | **OPEN — text-only, blocked on R-433** | +| **R-435** | **The snapshot-drop detector is blind to a single-app deletion — the exact shape the box can produce.** `snapshotDropFraction = 0.5` and `snapshotDropFloor = 5` (`hub/internal/monitor/offsite.go`) require a fall of MORE than half the previous count. demo-hp's baseline is **69** across 9 apps, so ~35 snapshots must go before it speaks; **one app's tag is ~9 and is invisible.** `offbox.go:1388` runs `forget --prune` **grouped by host,tags** — a per-tag wipe is precisely the shape a faulty retention or a targeted deletion produces. **This is deliberate, not accidental:** the constant's own comment argues the insensitivity, and *"a detector that cries wolf is switched off within a fortnight"* is a lesson this project paid for. **So this row is NOT a demand to lower the threshold.** It is a demand that the blind spot be written where the operator reads it, because "an unexplained fall is noticed within a day" (R-431, and `STATUS.md`) is true only of falls above half. **Discovered by arithmetic while planning the drill's Phase 2/3 pairing, which could not have worked:** Phase 2 deletes one app and Phase 3 expects the alarm to fire. | **OPEN — documentation, not a threshold change** | +| **R-436** | **LEAD, NOT A DEFECT — append-only may be reachable without a new machine, which would make R-95's real prevention far cheaper than the spike concluded.** Hetzner's port-23 restricted shell advertises, in its own `help`, these server-side backends: `borg`, `rsync`, `scp`, `sftp`, **`rclone serve restic --stdio`**. And restic 0.14.0 **recognises the `rclone:` backend** — MEASURED 2026-09-01 with a control: `banana:` → `Fatal: parsing repository location failed: invalid backend`, while `rclone:` → `exec: "rclone": executable file not found in $PATH` (i.e. the backend parsed and it tried to run the helper). rclone is **not** in the controller image today. `rclone serve restic` carries an **`--append-only`** flag. **Why this matters:** the R-95 spike's option 3 was priced at a new always-on service in the recovery path plus either a mount in the hot path or migrating every customer's history — and it was deferred on that price. This route needs neither: the server side already runs at the provider. **THE CAVEAT, STATED FIRST because it may kill the idea:** the **client** supplies the server command line, so a compromised guest could simply omit `--append-only` unless the provider pins it. **NOT ESTABLISHED:** whether Hetzner pins the flag or accepts client-supplied arguments. **That is a vendor question and it is cheap** — it should be asked before any code is written, because if the answer is "client-supplied" this lead is worth nothing. | **OPEN — ask the vendor before building anything** |