diff --git a/STATUS.md b/STATUS.md index 490e45b..a4238f8 100644 --- a/STATUS.md +++ b/STATUS.md @@ -58,8 +58,11 @@ So I took your stated fallback: the old store was **moved aside, not deleted** ( **One read of the hub's own records, no machine touched.** The hub holds backup keys for exactly **three** machines. Both demo boxes lost their old key in the same four-hour window on 4 August, before the retention fix was in force — that is the whole population of the problem, and it is -entirely ours. **The tester's machine has no record at all**, so it cannot be affected; and anything -enrolled from now on is covered, because the fix has been in force since 4 August. +entirely ours. **The tester's machine has no record at all** — so it cannot be affected by *this* +defect, and that is the only reassuring thing about it: the reason it has no record is that its host +row was deleted on 15 July, and it has **no off-site copy, no key and no local backup either**. See +the `PETI` row in the register. Anything enrolled from now on is covered, because the fix has been in +force since 4 August. I ran a control before trusting the query: it had to say *material present* for a machine known to have it and *absent* for one known not to. It did both. diff --git a/documentation/backlog/OPEN-ITEMS.md b/documentation/backlog/OPEN-ITEMS.md index b5fc8fd..66086ca 100644 --- a/documentation/backlog/OPEN-ITEMS.md +++ b/documentation/backlog/OPEN-ITEMS.md @@ -307,7 +307,7 @@ unconditional promise, which CAMPAIGN-11 §7 step 7 measured the customer-facing | **E-2a** | ~~The target move needs a root-fenced wrapper — the agent cannot do it~~ | **SHIPPED + PROVEN-LIVE** (agent v0.113.0 + host-install v1.22.0, 2026-07-29) | — | `felhom-backup-target-apply` behind a literal `FELHOM_BACKUPTARGET` sudoers alias; the agent's PVE role was NOT widened. Enforces F-1 (`mountpoint -q`) and F-2 (`is_mountpoint 1` hardcoded), refuses a root-device target, has NO storage-removal path (grep-assertable), is idempotent and refuses to repoint. All five laws proven live as root on demo-hp with 0 stray storages | — | | **E-2b** | ~~`NotifyStorageDisconnected`/`Reconnected` defined and called NOWHERE — a drive going absent emitted no event on any channel~~ | **SHIPPED + PROVEN-LIVE** (controller v0.184.1 + agent v0.112.0 + hub v0.81.0, 2026-07-29) | — | Seam wired in `ReconcileDriveGates`; a target drive raises the specific `backup_target_absent` instead. **A keying bug was caught before deploy:** `a.Path` is the registered GUEST path, not the agent's host `MountPath`, so the target branch was unreachable — every absent drive, target included, fell through to the generic event (v0.184.1). Tests observe the WIRE (httptest hub), not a mock | — | | **E-2c** | ~~E-1 put the whole-guest backups on a drive `POST /disks/eject` would eject~~ | **SHIPPED + PROVEN-LIVE** (agent v0.112.0, 2026-07-29) | — | Eject + decommission refuse 409 on the backup-target mount, naming the storage and the remedy. **Live on BOTH boxes:** demo-hp `/mnt/nvme-1tb` and demo-felhom `/mnt/hdd_1` both refused, drives unmoved. NOT a role reclassification — `RoleForStorage` untouched, because on both boxes that drive is ALSO the enrolled user-data drive; `TestEjectStillAllowedOnANonTargetDrive` pins the non-over-correction and `/var/lib/vz` is still refused by the PRE-EXISTING role gate, not this one | — | -| **PETI** | **`peti-felhom` deliberately NOT migrated.** Its whole-guest backup still shares a device with its guest, so a drive failure there is **offsite-only recovery** | **ACCEPTED RISK — parked** | operator's next visit (tester reinstalling from scratch) | **Accepted until the reinstall; re-evaluate if that slips past ~2026-09-01.** Do not migrate, do not touch | operator | +| **PETI** | **`peti-felhom` deliberately NOT migrated — and the mitigation this row used to name DOES NOT EXIST.** This row said a drive failure there is *"offsite-only recovery"*. Re-read from the hub's own store on 2026-08-10 and again on 2026-08-12, with a control run first (the escrow query returns 1+1 rows for each demo box and 0+0 for `drill-r50`, so it distinguishes the states): **there is no off-site copy, no key, and no local backup either.** Three independent reasons, each a fact rather than an inference: **(1)** the host row was DELETED — `host_deletions` id=1, `peti-felhom-86d37d`, **2026-07-15 08:56:22**, `escrow_acked = 0` — long before host-delete-demotes-escrow-to-retained-custody existed, so nothing was carried over; **(2)** there is **no escrow row of any kind**, current or superseded (only 4 exist hub-wide, all belonging to the two demo boxes), and the off-site restic REPOSITORY password lives in `identity_blob` on that row (`hub/internal/store/store.go:381`) — its only other copy is `/offbox/repo_password` (`controller/internal/backup/offbox.go:395`) **on the very disk whose failure is the scenario**; **(3)** off-site backup **never ran once** — its last report carried `offsite: {escrow_state: "pending", snapshot_count: 0, repo_size_bytes: 0}`, and that is the fork-4 guard working exactly as designed (`controller/internal/settings/settings.go:317-320`: *"no offsite run proceeds until an operator confirms the escrow ceremony"*), not a fault. The local app-data restic repo was also empty (`snapshot_count: 0, integrity_ok: false`), and the whole-guest vzdump shares the failing device. **SO: if that drive fails today, everything on it is lost.** **Size, so this is not read as larger than it is:** one lightly-used test box — a single `/` mount, **3.6 GB used of 48.9 GB**, one catalogued app (`rallly`); the dashboard was never claimed (`customer_claims.claimed_at` NULL). **BOUNDARY:** every fact is as of the last report, **2026-07-15 08:39:00 UTC** (controller 0.115.0); the hub has heard nothing since and confirming today's state would mean contacting the machine, which is fenced. **Contact since deletion:** no inbound row of any kind after 2026-07-15 08:39 — the only later rows are the hub's OWN staleness alarms (`source = hub`: `node_stale` 09:09:32, `node_down` 09:39:32) — and no contact attempt, accepted or rejected, in the current hub pod's logs (since 2026-08-09 17:26 UTC; grep proven by 851 `demo-hp` hits against 0 for peti, 0 unauthorized). **The window 2026-07-15 → 2026-08-09 cannot be answered from records**: a report from a deleted host 401s and is not persisted, and those logs are gone. **THE RULING IS LEFT OPEN DELIBERATELY** — whether the machine stays parked is the operator's call and does not need restating here; this row records the FACT, which does not need his opinion to be true | **PARKED — the recorded mitigation is void; ruling OPEN for the operator** | — | **First act of the visit: copy that ~3.6 GB off before anything is reinstalled — it is currently the only copy in existence.** Do not migrate, do not contact | operator | | **R-109** | ~~The DR recipe records no backup target~~ | **SHIPPED + PROVEN-LIVE** (agent v0.118.1 + hub v0.83.0, 2026-07-30) | — | `backup_target` resolves from the PRIMARY tier of `cfg.Backup.BackupTiers()` — the function the scheduler consults, not a re-derivation — plus the mountpoint, which is what actually separates `/mnt/hdd_1` from `/var/lib/vz`. Three states, and unresolvable is recorded as unresolvable (`agent_backup_config_unavailable` / `not_a_known_storage`), never a default. **The resolver reads the daemon-start config on purpose:** a target move rewrites `agent.json` and deliberately does NOT restart, so a disk re-read would name a storage no archive had reached. **Needed a HUB half nobody had scoped** — `AssembleDRRecipe` allow-lists top-level keys, so the field would have been stored intact and dropped before any operator saw it (→ **R-122**). Evidence: `audits/R106-R109-recipe-completeness-2026-07-30.md` | — | | **R-106** | ~~The DR recipe records the PBS namespace as `"root"` on every box~~ | **SHIPPED + PROVEN-LIVE** (agent v0.118.1, 2026-07-30) | — | **Was open-but-UNREGISTERED on this page until 2026-07-30 (→ R-123)** — `ROADMAP.md:109` had it READY and the only mention here was inside R-109's prose. Namespace now resolves from the pbs STORAGE (storage.cfg's `namespace`), the same field `vzdump --storage ` makes PVE read, so the recipe cannot disagree with the backup that produced the snapshot. An unconfigured namespace still reads `"root"` — that is an ANSWER, and `namespace_state` separates it from not knowing. Live: `demo-felhom` and `demo-hp` now report their own namespaces. Evidence: same audit | — | | **R-122** | ~~`AssembleDRRecipe` silently DROPPED `offsite_restic` — the offsite recovery location never reached any recipe~~ | **SHIPPED** (hub v0.83.0, 2026-07-30) | — | **Found 2026-07-30 while scoping R-109's hub half; it had already shipped and nobody knew.** The controller has emitted `offsite_restic` since fork-4 (*"so DR knows WHERE to recover from"*), the hub stored it for **all three real customers**, and `appHalfShape` never listed the key — so no delivered recipe has ever contained it. No error, no log, green suite, because the fixture `drAppHalf` is hand-written and omits the field. `hostHalfShape`/`appHalfShape` are **ALLOW-LISTS dressed as forward-compat**; `TestAssembleDRRecipe_CarriesEveryEmittedSection` is now the guard, built on halves read verbatim out of the live `dr_recipe` table. `REUSE.md` (both repos) records that a recipe section is a TWO-REPO change | — |