From f761f69d469bf828033d52fb0795c06dca61fae1 Mon Sep 17 00:00:00 2001 From: kisfenyo Date: Tue, 21 Jul 2026 10:52:52 +0200 Subject: [PATCH] =?UTF-8?q?docs:=20R-39=20CLOSED=20=E2=80=94=20STOP-2=20pr?= =?UTF-8?q?oven=20live;=20DR-tier=20row=20->=20PROVEN-LIVE?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The operator pressed Re-issue PBS credentials and the chain closed in 13 seconds. The identical click on 2026-07-18 did nothing at all. hub 08:39:31Z fresh mint, generation 0 -> 1; descriptor gains secret_generation: 1 (token_id + fingerprint BYTE-IDENTICAL — the invisible re-key shape) agent 10:39:34 felhom-pbs-apply read felhom-pbs (leg b: the impossible read) agent 10:39:38 ERROR REJECTED ... applied and DEAD, previous_state=applied (leg c: the R-39 state, loud) hub 08:39:45Z consumed_at stamped agent 10:39:45 one-time token secret consumed (leg a: NO short-circuit) agent 10:39:45 reconcile (set-only, no --server) agent 10:39:47 pbsdr: converged state=applied Corroboration: marker hash moved to afbb3b41… (it was byte-identical to the pre-reissue marker in the failure); secret mtime 2026-07-18 -> 2026-07-21 10:39:45; new credential probes 200; three consecutive reports trace applied -> auth_failed -> applied; ZERO self-heal escalations, one mint, one consume, no consumed-failed.json — the box healed through the descriptor path before the damper was ever needed. Recorded for future runbooks: the operator first pressed the OFFSITE re-issue (two distinct Re-issue actions exist). Harmless to PBS-DR, but it rotated the restic password and correctly marked the escrow STALE, so the ceremony had to be re-run. Name the surface explicitly next time. --- documentation/architecture/00-capability-map.md | 2 +- documentation/backlog/ROADMAP.md | 2 +- 2 files changed, 2 insertions(+), 2 deletions(-) diff --git a/documentation/architecture/00-capability-map.md b/documentation/architecture/00-capability-map.md index c432a51..7744af5 100644 --- a/documentation/architecture/00-capability-map.md +++ b/documentation/architecture/00-capability-map.md @@ -35,7 +35,7 @@ | Customer claim: one-time emailed code → customer sets own password (bcrypt, operator never sees it) | controller v0.122, hub v0.50 | **PROVEN-LIVE** (drill VM) | `DRILL-day0-vm-2026-07-12` §10/F-4 (gate ON via real edge; claimed, code consumed) | Never executed by a non-Viktor human → R-3. **Deliverability (R-4), gmail half DONE 2026-07-18:** the rehearsal's claim email was the first sent under the tightened DMARC `p=quarantine` and **landed in the gmail Inbox, not spam** (`tests/VALIDATION-n100-rehearsal-2026-07-18.md`). **freemail.hu remains Viktor's open half.** (Dropped mis-cited `CAMPAIGN-4` F-C — that is the escrow-claim 502, not password claim) | | Customer binds their own appliance (self-service): operator-sent 7-day tokenized capability link → public two-factor `/bind/` (console pairing code + retrieval passphrase) → hub stages the bind, no operator | hub v0.66.0 + ISO scripts v1.20.0 | **PROVEN-LIVE** (real customer-zero bind on metal, 2026-07-18) | **`tests/VALIDATION-n100-rehearsal-2026-07-18.md`:** operator minted + emailed the link 16:28:55 (7-day TTL, expiry 2026-07-25 recorded); **the customer bound their own box at 16:29:55 with `attempts=0`, `locked=0`** — `appliance_bound` carries source **`customer_selfbind`**, and the credential was delivered **26 s later** with no operator action. Hub-side lifecycle in `hub-state.txt` (`selfbind_tokens` mint→email→consume). Prior unit evidence: hub v0.66.0 (`web/selfbind.go`, `store/selfbind.go`; Scenarios A–F + F1/F2; 4 red-proofs verified red — THE TRAP `/bind/` exemption, no-oracle, lockout, single-active); GC verdict §3 (no appliance GC → TTL stands alone) | R-27 **slice 1**. No appliance list ever rendered; wrong code == wrong passphrase (one generic failure); 5-attempt lockout → call support; expiry falls back to operator-bind. **Live first-run DONE 2026-07-18** (rehearsal; the console banner rendered on the real ISO). **R-27b** (controller second-box dismissable prompt) deferred; **multi-box-per-link** = repeated operator sends | | Escrow ceremony: customer-facing wizard, one-shot R claim, operator zero-knowledge | controller v0.127, agent v0.88/0.89 | **PROVEN-LIVE** (drill VM, endpoint-exact) | agent v0.88.0 REPORT (ceremony ~4s, one-shot claim 200→410, R absent from every payload); `SPIKE-controller-escrow-2026-07-13` | **Customer-facing browser wizard FIRST LIVE FIRING 2026-07-18** (`tests/VALIDATION-n100-rehearsal-2026-07-18.md`, S6): customer zero drove the wizard on the reborn box — ceremony started 16:56:29, recovery code claimed one-shot 16:56:39 (absent from logs by design), hub-verified and `EscrowState` auto-confirmed 16:56:41, **offsite runs enabled 12 s after the ceremony began**; the v0.138.0 „megerősítésre vár, legfeljebb 15 perc" awaiting card rendered and flipped on the ACK (operator screenshots: Viktor's set). Honest caveat: at a 12-second confirm the awaiting window is so short that catching *both* states on screen is luck, not procedure. Prior: endpoints driven on the drill VM. **agent v0.89.0:** `/escrow/preflight` `pbs_storage_id` row now live-reloads (reads current agent.json) — a pbsdr convergence that seeds the id flips it green with NO service restart. **hub v0.60.0 (data-first retention):** a re-escrow with a DIFFERENT sealed passphrase no longer destroys the old blob — the hub RETAINS it (`host_escrow_superseded`), so a previous passphrase stays recoverable with its recovery code (turns the reinstall-orphan incident from "history destroyed" into "history recoverable"). Guided-recovery flow = R-26. Red-proof `TestSaveHostEscrow_RetainsSuperseded`. **hub v0.60.1 — custody survives the host lifecycle:** host deletion (with the escrow ack) DEMOTES the current blob to retained custody (moved into `host_escrow_superseded`, never destroyed; existing superseded rows spared); the customer Danger-zone Delete is the one true purge point (cascades both escrow tables incl. already-deleted hosts). No operator path through host lifecycle can lose a blob. Red-proofs `TestDeleteHost_DemotesEscrowNeverDestroys` + `TestDeleteCustomer_PurgesEscrowCustody` | -| DR tier by default: PBS + WireGuard base infra on every install, hub-controlled activation | installer v1.15, agent v0.86, hub v0.51 | **IMPLEMENTED** | `DRILL-day0-take2-2026-07-12` §2 (WG enabled both modes, PBS-DR descriptor auto-provisioned ~1s after WG registration, zero operator steps); ships installer v1.15/agent v0.86/hub v0.51 | Live only on demo/drill fleet. (Cited spike was slice-0 mechanics — shipped nothing; corrected.) **⚠ The candidate upgrade to PROVEN-LIVE is WITHDRAWN — the 2026-07-18 rehearsal produced a live counter-example (R-39).** On the reborn N100 the descriptor auto-provisioned and the agent reported `converged state=applied` (16:45:53), yet **the storage is dead**: `pvesm status` → `felhom-pbs: error fetching datastores - 401 Unauthorized` / `inactive`, and a direct probe with the stored credential returns **401 on every endpoint including `/version`** while the WG transport is healthy (handshake 9 s, 27.9 ms RTT) — i.e. authentication failure, not ACL scope. Root cause in the evidence: **the hub minted a SECOND token secret at 16:47:52, two minutes after the agent had applied the first, and `consumed_at` is still NULL**; the converged state machine will not re-apply, and the agent's 15-minute verify loop **cannot even read the credential to notice** (`open /etc/pve/priv/storage/felhom-pbs.pw: permission denied` — non-root agent reading a file it writes through a root wrapper). A tier that reports `applied` while silently unable to authenticate is exactly the shape that must not carry a PROVEN-LIVE badge. See `tests/VALIDATION-n100-rehearsal-2026-07-18.md` F2 and `pbs-dr-state.txt`. **agent v0.89.0 closes the F4 non-default-storage-id gap (R-22) — PROVEN-LIVE 2026-07-17:** the reconcile self-grants the ACL through the root wrapper on a pre-check 403 instead of dead-locking. Reproduced F4 on the demo (marker moved aside = reinstall fresh-state + felhom-offsite ACLs revoked) → next reconcile tick `pbsdr: pre-check 403 … self-granting … (R-22)` → `converged state=adopted` in ~3 s, ACLs self-restored, `pvesm status felhom-offsite`=active, zero operator action. No more one-shot `pveum` grant **2026-07-21 — the R-39 fleet fix SHIPPED (hub v0.68.0 + agent v0.91.2), closing the self-heal chain end to end.** The three defects that let a box be `applied` and dead simultaneously are each addressed: the hub stamps a monotonic `secret_generation` into the descriptor so a credential re-key finally MOVES the content hash the agent re-applies on; the wrapper gains a narrow `read` verb so the non-root agent can read the credential it writes (it never could — `/etc/pve/priv` is 0700 root:www-data, which made the verify loop blind by construction); and `pbs.ProbeAuth` turns a 401 into a loud `auth_failed` that the existing `pbsdrheal` damper escalates to a fresh mint. Plus a consumed_at honesty gauge for the disagreement no single tier can see (box says `applied`, hub's staged secret never consumed). Proven live on felhom-pve: the agent read its credential through the wrapper (`rc=0`) and probed successfully (`credential probe OK storage=felhom-pbs`). **This row is NOT upgraded to PROVEN-LIVE yet** — the decisive evidence is STOP-2, the operator pressing Re-issue and the box converging where the identical click did nothing on 2026-07-18. Until that runs, the capability is SHIPPED but the destroy-then-recover proof this row asks for is not in hand. Evidence: `felhom-agent/REPORT.md` (2026-07-21). | +| DR tier by default: PBS + WireGuard base infra on every install, hub-controlled activation | installer v1.15, agent v0.86, hub v0.51 | **PROVEN-LIVE** (2026-07-21) | `DRILL-day0-take2-2026-07-12` §2 (WG enabled both modes, PBS-DR descriptor auto-provisioned ~1s after WG registration, zero operator steps); ships installer v1.15/agent v0.86/hub v0.51 | Live only on demo/drill fleet. (Cited spike was slice-0 mechanics — shipped nothing; corrected.) **⚠ The candidate upgrade to PROVEN-LIVE is WITHDRAWN — the 2026-07-18 rehearsal produced a live counter-example (R-39).** On the reborn N100 the descriptor auto-provisioned and the agent reported `converged state=applied` (16:45:53), yet **the storage is dead**: `pvesm status` → `felhom-pbs: error fetching datastores - 401 Unauthorized` / `inactive`, and a direct probe with the stored credential returns **401 on every endpoint including `/version`** while the WG transport is healthy (handshake 9 s, 27.9 ms RTT) — i.e. authentication failure, not ACL scope. Root cause in the evidence: **the hub minted a SECOND token secret at 16:47:52, two minutes after the agent had applied the first, and `consumed_at` is still NULL**; the converged state machine will not re-apply, and the agent's 15-minute verify loop **cannot even read the credential to notice** (`open /etc/pve/priv/storage/felhom-pbs.pw: permission denied` — non-root agent reading a file it writes through a root wrapper). A tier that reports `applied` while silently unable to authenticate is exactly the shape that must not carry a PROVEN-LIVE badge. See `tests/VALIDATION-n100-rehearsal-2026-07-18.md` F2 and `pbs-dr-state.txt`. **agent v0.89.0 closes the F4 non-default-storage-id gap (R-22) — PROVEN-LIVE 2026-07-17:** the reconcile self-grants the ACL through the root wrapper on a pre-check 403 instead of dead-locking. Reproduced F4 on the demo (marker moved aside = reinstall fresh-state + felhom-offsite ACLs revoked) → next reconcile tick `pbsdr: pre-check 403 … self-granting … (R-22)` → `converged state=adopted` in ~3 s, ACLs self-restored, `pvesm status felhom-offsite`=active, zero operator action. No more one-shot `pveum` grant **2026-07-21 — the R-39 fleet fix SHIPPED (hub v0.68.0 + agent v0.91.2), closing the self-heal chain end to end.** The three defects that let a box be `applied` and dead simultaneously are each addressed: the hub stamps a monotonic `secret_generation` into the descriptor so a credential re-key finally MOVES the content hash the agent re-applies on; the wrapper gains a narrow `read` verb so the non-root agent can read the credential it writes (it never could — `/etc/pve/priv` is 0700 root:www-data, which made the verify loop blind by construction); and `pbs.ProbeAuth` turns a 401 into a loud `auth_failed` that the existing `pbsdrheal` damper escalates to a fresh mint. Plus a consumed_at honesty gauge for the disagreement no single tier can see (box says `applied`, hub's staged secret never consumed). Proven live on felhom-pve: the agent read its credential through the wrapper (`rc=0`) and probed successfully (`credential probe OK storage=felhom-pbs`). **STOP-2 RAN 2026-07-21 AND THE CHAIN CLOSED — 13 SECONDS, operator click to converged.** The operator pressed **Re-issue PBS credentials**; the identical click on 2026-07-18 did nothing at all. Full chain (hub UTC / host CEST = UTC+2): `08:39:31Z` hub mints a fresh secret, **generation 0 → 1**, and the descriptor gains `"secret_generation": 1` — with `token_id` and `fingerprint` **byte-identical**, i.e. exactly the re-key shape that used to be invisible → `10:39:34` the agent READS its credential through the wrapper (leg b — the read that was impossible until v0.91.0) → `10:39:38` **`ERROR pbsdr: the DR endpoint REJECTED this box's credential — the tier is applied and DEAD` `previous_state=applied`** (leg c: the exact R-39 failure state, detected out loud for the first time ever) → `10:39:45` **`one-time token secret consumed`** `secret_len=36` (leg a: **NO short-circuit** — this is the line that never appeared on 2026-07-18) → `10:39:45` `felhom-pbs-apply reconcile` (the set-only wrapper, no `--server`) → `10:39:47` **`pbsdr: converged state=applied`**. Corroboration: the agent marker hash moved to `afbb3b41…` (it was byte-identical to the pre-reissue marker in the failure); `consumed_at` stamped `08:39:45Z`; the on-disk secret's mtime moved `2026-07-18 20:28:52` → `2026-07-21 10:39:45`; a live probe with the NEW credential returns **200**; three consecutive hub reports trace the whole state machine `applied → auth_failed → applied`; and **zero** `pbsdr_selfheal` escalations fired — the box healed through the descriptor path before the damper was ever needed, with exactly ONE mint and ONE consume and no `consumed-failed.json`. **Row upgraded to PROVEN-LIVE (2026-07-21).** Evidence: `felhom-agent/REPORT.md` (2026-07-21). | | Customer RESET (middle lifecycle tier: host delete < RESET < customer Delete): one operator action → pre-first-install; all operational state destroyed, identity + basic config survive | hub v0.61.0, felhom-tenantsync v1.1.0 | **PROVEN-LIVE (external teardown, incl. two real firings)** | **`tests/VALIDATION-n100-rehearsal-2026-07-18.md` — two live firings, both host-delete-first, on two different customers** (`demo-vm-felhom` 15:49:57, `demo-felhom` 16:08:51): every leg `ok` (`claim`, `db_purge`, `descriptor`, `hetzner`, `pbs`), escrow acked separately, each completing in 8–9 s (`hub-state.txt` `customer_resets`). The **Hetzner sub-account destruction is now verified against the live pool box** — and produced the run's sharpest lesson: **a sub-account is an access-control object, not a data object.** Deleting it left its `/home` intact, so re-enabling offsite recreated an account over the previous lifecycle's ciphertext under a key this same RESET had destroyed — which is why the orphan guard fired at 16:58:14 (**a finding by S7's own criterion**) and why RESET now needs a base-dir purge → **R-32**. Prior: hub v0.61.0 REPORT; **ep0 live drill 2026-07-17** (throwaway `drill-reset-01` with a real backup: deprovision `deleted:true` destroyed the namespace + backup group + token, idempotent re-run `deleted:false`, all 3 real tenants + shared user survived); red-proofs (ack-gate, partial-failure resumability) + orchestration/store/offsite/render tests | External teardown FIRST, DB purge LAST, every leg idempotent; refuses while any host row exists; separate escrow-custody ack; clears claim (fresh code next onboarding); keeps the offsite tier CHOICE, drops provisioned fields. **Live-clicked 2026-07-18** (twice, by Viktor) — this supersedes the earlier "not live-clicked / Hetzner delete unit-tested only" note. **Consistency gap → R-25b:** the Danger-zone DELETE leaves host rows and doesn't run this teardown | | Uninstall: KEPT-vs-WIPED statement, secret purge, enrolled-drive handling | installer | **PARTIAL** | `DRILL-GL6-2026-07-08` Phase 1/5 (KEPT-vs-WIPED printed verbatim; drive data intact ×3); GL-4 code | Secret purge (GL6-F1 `.bak` residue) fixed v1.12.0; enrolled-drive `mnt-*.mount` units survive (GL6-F2, open); cluster-aware `felhom_guests` guard + saferemove cost warning missing → R-9 | diff --git a/documentation/backlog/ROADMAP.md b/documentation/backlog/ROADMAP.md index 119192e..07f0a41 100644 --- a/documentation/backlog/ROADMAP.md +++ b/documentation/backlog/ROADMAP.md @@ -35,7 +35,7 @@ | ID | Item | Size | Status | Notes | |----|------|------|--------|-------| -| R-39 | **[P2-HIGH] The PBS DR tier can be `applied` and dead at the same time — and nothing notices.** On the reborn N100 the descriptor auto-provisioned and the agent converged `state=applied`, yet `pvesm status` reports `felhom-pbs: error fetching datastores - 401 Unauthorized` / `inactive` and a direct probe with the stored credential 401s on **every** endpoint including `/version` (WG transport healthy: handshake 9 s, 27.9 ms RTT — so authentication, not ACL scope). Three compounding defects: **(a)** a **mint/consume race** — the hub minted a SECOND token secret at 16:47:52, two minutes *after* the agent applied the first, and `consumed_at` is still NULL; **(b)** the converged state machine will not re-apply, so the box is pinned to a stale secret; **(c)** the agent's 15-minute PBS verify loop **cannot read the credential to detect any of it** (`open /etc/pve/priv/storage/felhom-pbs.pw: permission denied` — the non-root agent writes that file through a root sudo wrapper, then reads it directly). | M | **SHIPPED 2026-07-21 (hub 0.68.0 + agent 0.91.2); awaiting operator STOP-2/STOP-3** | **DIAGNOSIS (2026-07-18, live on the N100 — supersedes the initial hypothesis).** The brief guessed "the re-mint fails to bump the generation". **That is FALSE and no hub fix was shipped:** `store.SetHostDesired` bumps `desired_generation` unconditionally (it went 2→3 on the re-issue), and `web/configs.go`'s `applyPBSDR` is likewise exonerated — its "no re-key, no second secret, no spurious generation bump" comment is accurate, guarded by the `cur != nil && cur.Namespace != ""` early return, and the hub log shows mint #2 came from the **re-issue** path, not from an Edit-tab Save. **The real mechanism is a signal mismatch between the two tiers.** The hub's re-consume signal is *a generation bump + a poke*; the agent's re-apply trigger is *a change in the DESCRIPTOR CONTENT HASH* (`felhom-agent internal/pbsdr/manager.go` ~L235: `if mk := m.loadMarker(); mk != nil && mk.Hash == h && (cf == nil || cf.Hash != h) { return }`). An ep0 credential re-issue re-keys the **secret of an existing token**, so `token_id` and `fingerprint` are unchanged and the descriptor is **byte-identical** — only the side-table `host_pbs_secrets` row rotates. Same hash → the converged agent short-circuits → the fresh secret is never consumed → the box keeps presenting a revoked credential → **401 forever**. Proof in one line: the agent's `consumed-failed.json` carries hash `a4e5424…`, **identical** to the `marker.json` written 2 min before the re-issue. The comment at `hub/internal/web/pbsdr.go:320` asserts the reissue refreshes the descriptor "with the NEW token_id/fingerprint" — that assumption is simply false for this op. **A SECOND, independent defect was found while healing** and is fixed: `configs/felhom-pbs-apply`'s `reconcile` passed `--server` to `pvesm set`, which PVE rejects wholesale as a create-only parameter, so *every* re-apply exited 255 — and because the agent consumes the one-time secret BEFORE calling the wrapper, each re-issue **burned a credential**. Fixed in **agent v0.90.1** (one argv line + red-proof `TestReconcileNeverPassesServerToPvesmSet`); proven live (`pvesm set --server ` rejected, without it rc 0). **BOX HEALED 2026-07-18:** after the wrapper hotfix, Viktor's Re-issue click converged in 9 s and the tier went `401/inactive` → **`active`** (token probe 401 → 200) with a real backup landing PBS-side — `felhom-pbs:backup/ct/9201/2026-07-18T18:31:06Z`, 9 744 319 312 B, encrypted under the escrowed key. **2026-07-21 — the "cheap half" is ALREADY CLOSED in the field, and the planned artifact publish was CANCELLED as a false signal (operator ruling, same day).** v0.90.1 was to be built, published and deployed as "the PBS wrapper argv fix". Inspection of `9596d5a` shows it changes **zero non-test Go files** — `CHANGELOG.md`, `REPORT.md`, `configs/felhom-pbs-apply` (the fix) and `internal/pbsdr/manager_test.go` (the red-proof); its own commit message states *"the Go binary is unchanged"*. So the 0.90.1 binary is functionally identical to the 0.90.0 in the field. The fix is live anyway by two independent paths: felhom-pve carries the hotfixed wrapper since 2026-07-18 (`args=(--fingerprint "$fp")` at L107, `.bak-20260718-preR39` retained), and **every new install fetches the wrapper unversioned** — `felhom-host-install.sh:1914` `fetch_raw` pulls `configs/felhom-pbs-apply` from `raw/branch/main`, and `9596d5a` is an ancestor of `main`. Publishing 0.90.1 would therefore have delivered no behaviour change, restarted the agent on a production host at a remote site for nothing, and — once the Day-0 manifest was saved to 0.90.1 — advertised that the fix shipped as a versioned artifact when the artifact channel never carried it. **Ruled: leave 0.90.0 published; record the closure here instead.** Evidence: `felhom-controller/REPORT.md` §5 (2026-07-21). **This surfaced a NEW item — see R-50b: a root-owned privileged host artifact is delivered from `main` with no version, so "which wrapper is on this host" is not answerable from any manifest.** **FLEET FIX SHIPPED 2026-07-21 — hub v0.68.0 + agent v0.91.2.** All three legs closed. **(a) the re-key is finally VISIBLE:** `host_pbs_secrets` gains a monotonic per-host `generation`, advanced by every fresh MINT and by nothing else, stamped into the descriptor as `secret_generation` — the only field a re-key moves, and because `descriptorHash` marshals the parsed struct, the thing that finally re-arms a converged agent. A re-STAGE deliberately does not advance it (same secret, same descriptor content). `omitempty` is load-bearing: emitting a zero would move every pre-existing descriptor's hash at once. *Spec deviation, deliberate:* the brief said to reuse "the new row's id, no schema change" — there is no row id (the table is `host_id PRIMARY KEY`, UPSERTed last-write-wins) and `created_at` collides within a second, so an additive counter column is the only monotonic source. **(b) the agent can finally READ its own credential:** a narrow wrapper `read` verb + exactly one sudoers line + a `pbsdr-read` capability row. Proven live on felhom-pve: `sudo -u felhom-agent sudo -n felhom-pbs-apply read felhom-pbs /etc/pve/priv/storage` → `rc=0`, 37 bytes, empty stderr. **(c) authentication is PROBED:** `pbs.ProbeAuth` (`GET /version` + an `ErrUnauthorized` sentinel) runs on the 15-minute collect path and becomes a loud `auth_failed` that `pbsdrheal` escalates to a fresh mint through the EXISTING damper — closing the loop end to end. Proven live: `level=DEBUG msg="pbs: credential probe OK" storage=felhom-pbs datastore=felhom-offsite`. 403 is deliberately NOT unauthorized (a narrow ACL must not be re-keyed forever); a transport error is UNKNOWN, never a rejection (a blip must not burn a credential). Plus a **consumed_at honesty gauge**: an unconsumed secret past a 15-min grace under a box reporting `applied` — the exact July-18 fingerprint, a disagreement no single tier can see — is surfaced with its own event, deliberately as a SURFACE not a heal (minting on top of an unconsumed secret is the R-39(a) race). **A LOAD-BEARING FACT the spec did not flag, checked rather than trusted:** `Apply` bails out if the storage status probe ERRORS, and `adopt` converges WITHOUT consuming when the storage reads active — so the whole fix depended on PVE's 401 behaviour. PVE's `storage_info` wraps `activate_storage`/`$plugin->status` in `eval{}` and leaves the pre-initialised `active => 0`, so a 401 returns **HTTP 200 with `active: 0`, never an API error** — `Apply` correctly falls through to verify → consume → reconcile. **A defect I shipped and caught:** v0.91.0 built the probe seam and `main.go` never called `SetAuthSink`, so the whole leg was INERT and every test still passed (the seam was injected directly) — same class as controller v0.154.0 the day before; fixed in v0.91.1, artifact superseded not overwritten, and v0.91.2 made a healthy probe observable so "no auth_failed" can never again be confused with "never probed". Four red-proofs, all at the assertion level. **REMAINING = the operator: STOP-2 (press Re-issue — the same click that did nothing on 2026-07-18 must now converge) and STOP-3 (manifest save: agent 0.91.2 / sha `34d309be…` / wrapper sha `104db0a4…` / MinAgent 0.91.2 — safe, the fleet is ONE host and it already runs 0.91.2).** Evidence: `felhom-agent/REPORT.md`, `felhom.eu/REPORT.md`. *(Superseded — the spec was written and shipped:)* **REMAINING (fleet, needs its own spec — deliberately NOT improvised):** (a) make a fresh unconsumed secret actually un-converge the agent — either `consumed_at` becomes authoritative or the descriptor carries a secret generation/nonce so the hash moves; (b) fix the verify loop's read path — it reads `/etc/pve/priv/storage/.pw` **directly as non-root**, a file it can only ever *write* through the root wrapper (`/etc/pve/priv` is `0700 root:www-data`; sudoers exposes `create|reconcile|grant` and **no read verb**), so the one loop that could catch this is permanently blind; (c) an auth probe in the reconciler/gauge/ceremony-precheck so `applied` can never mean `401`. — **Original finding note:** discovered by CC while collecting Phase-A evidence; not on the brief's finding list, so its P2-HIGH rank is provisional pending Viktor. The severity case: this is the DR tier, the failure is silent, and it would surface first at a real restore. Suggested shape: make `consumed_at` authoritative (a fresh unconsumed secret must un-converge the reconciler), fix the verify loop's read path (read via the same root wrapper that writes it), and make a failing `pvesm status` a LOUD state rather than a skipped datastore. **Blocks the DR-tier map row's candidate upgrade to PROVEN-LIVE — that upgrade is now explicitly WITHDRAWN.** Evidence `pbs-dr-state.txt`, `hub-state.txt` | +| R-39 | **[P2-HIGH] The PBS DR tier can be `applied` and dead at the same time — and nothing notices.** On the reborn N100 the descriptor auto-provisioned and the agent converged `state=applied`, yet `pvesm status` reports `felhom-pbs: error fetching datastores - 401 Unauthorized` / `inactive` and a direct probe with the stored credential 401s on **every** endpoint including `/version` (WG transport healthy: handshake 9 s, 27.9 ms RTT — so authentication, not ACL scope). Three compounding defects: **(a)** a **mint/consume race** — the hub minted a SECOND token secret at 16:47:52, two minutes *after* the agent applied the first, and `consumed_at` is still NULL; **(b)** the converged state machine will not re-apply, so the box is pinned to a stale secret; **(c)** the agent's 15-minute PBS verify loop **cannot read the credential to detect any of it** (`open /etc/pve/priv/storage/felhom-pbs.pw: permission denied` — the non-root agent writes that file through a root sudo wrapper, then reads it directly). | M | **CLOSED 2026-07-21 — PROVEN LIVE (hub 0.68.1 + agent 0.91.2)** | **DIAGNOSIS (2026-07-18, live on the N100 — supersedes the initial hypothesis).** The brief guessed "the re-mint fails to bump the generation". **That is FALSE and no hub fix was shipped:** `store.SetHostDesired` bumps `desired_generation` unconditionally (it went 2→3 on the re-issue), and `web/configs.go`'s `applyPBSDR` is likewise exonerated — its "no re-key, no second secret, no spurious generation bump" comment is accurate, guarded by the `cur != nil && cur.Namespace != ""` early return, and the hub log shows mint #2 came from the **re-issue** path, not from an Edit-tab Save. **The real mechanism is a signal mismatch between the two tiers.** The hub's re-consume signal is *a generation bump + a poke*; the agent's re-apply trigger is *a change in the DESCRIPTOR CONTENT HASH* (`felhom-agent internal/pbsdr/manager.go` ~L235: `if mk := m.loadMarker(); mk != nil && mk.Hash == h && (cf == nil || cf.Hash != h) { return }`). An ep0 credential re-issue re-keys the **secret of an existing token**, so `token_id` and `fingerprint` are unchanged and the descriptor is **byte-identical** — only the side-table `host_pbs_secrets` row rotates. Same hash → the converged agent short-circuits → the fresh secret is never consumed → the box keeps presenting a revoked credential → **401 forever**. Proof in one line: the agent's `consumed-failed.json` carries hash `a4e5424…`, **identical** to the `marker.json` written 2 min before the re-issue. The comment at `hub/internal/web/pbsdr.go:320` asserts the reissue refreshes the descriptor "with the NEW token_id/fingerprint" — that assumption is simply false for this op. **A SECOND, independent defect was found while healing** and is fixed: `configs/felhom-pbs-apply`'s `reconcile` passed `--server` to `pvesm set`, which PVE rejects wholesale as a create-only parameter, so *every* re-apply exited 255 — and because the agent consumes the one-time secret BEFORE calling the wrapper, each re-issue **burned a credential**. Fixed in **agent v0.90.1** (one argv line + red-proof `TestReconcileNeverPassesServerToPvesmSet`); proven live (`pvesm set --server ` rejected, without it rc 0). **BOX HEALED 2026-07-18:** after the wrapper hotfix, Viktor's Re-issue click converged in 9 s and the tier went `401/inactive` → **`active`** (token probe 401 → 200) with a real backup landing PBS-side — `felhom-pbs:backup/ct/9201/2026-07-18T18:31:06Z`, 9 744 319 312 B, encrypted under the escrowed key. **2026-07-21 — the "cheap half" is ALREADY CLOSED in the field, and the planned artifact publish was CANCELLED as a false signal (operator ruling, same day).** v0.90.1 was to be built, published and deployed as "the PBS wrapper argv fix". Inspection of `9596d5a` shows it changes **zero non-test Go files** — `CHANGELOG.md`, `REPORT.md`, `configs/felhom-pbs-apply` (the fix) and `internal/pbsdr/manager_test.go` (the red-proof); its own commit message states *"the Go binary is unchanged"*. So the 0.90.1 binary is functionally identical to the 0.90.0 in the field. The fix is live anyway by two independent paths: felhom-pve carries the hotfixed wrapper since 2026-07-18 (`args=(--fingerprint "$fp")` at L107, `.bak-20260718-preR39` retained), and **every new install fetches the wrapper unversioned** — `felhom-host-install.sh:1914` `fetch_raw` pulls `configs/felhom-pbs-apply` from `raw/branch/main`, and `9596d5a` is an ancestor of `main`. Publishing 0.90.1 would therefore have delivered no behaviour change, restarted the agent on a production host at a remote site for nothing, and — once the Day-0 manifest was saved to 0.90.1 — advertised that the fix shipped as a versioned artifact when the artifact channel never carried it. **Ruled: leave 0.90.0 published; record the closure here instead.** Evidence: `felhom-controller/REPORT.md` §5 (2026-07-21). **This surfaced a NEW item — see R-50b: a root-owned privileged host artifact is delivered from `main` with no version, so "which wrapper is on this host" is not answerable from any manifest.** **FLEET FIX SHIPPED 2026-07-21 — hub v0.68.0 + agent v0.91.2.** All three legs closed. **(a) the re-key is finally VISIBLE:** `host_pbs_secrets` gains a monotonic per-host `generation`, advanced by every fresh MINT and by nothing else, stamped into the descriptor as `secret_generation` — the only field a re-key moves, and because `descriptorHash` marshals the parsed struct, the thing that finally re-arms a converged agent. A re-STAGE deliberately does not advance it (same secret, same descriptor content). `omitempty` is load-bearing: emitting a zero would move every pre-existing descriptor's hash at once. *Spec deviation, deliberate:* the brief said to reuse "the new row's id, no schema change" — there is no row id (the table is `host_id PRIMARY KEY`, UPSERTed last-write-wins) and `created_at` collides within a second, so an additive counter column is the only monotonic source. **(b) the agent can finally READ its own credential:** a narrow wrapper `read` verb + exactly one sudoers line + a `pbsdr-read` capability row. Proven live on felhom-pve: `sudo -u felhom-agent sudo -n felhom-pbs-apply read felhom-pbs /etc/pve/priv/storage` → `rc=0`, 37 bytes, empty stderr. **(c) authentication is PROBED:** `pbs.ProbeAuth` (`GET /version` + an `ErrUnauthorized` sentinel) runs on the 15-minute collect path and becomes a loud `auth_failed` that `pbsdrheal` escalates to a fresh mint through the EXISTING damper — closing the loop end to end. Proven live: `level=DEBUG msg="pbs: credential probe OK" storage=felhom-pbs datastore=felhom-offsite`. 403 is deliberately NOT unauthorized (a narrow ACL must not be re-keyed forever); a transport error is UNKNOWN, never a rejection (a blip must not burn a credential). Plus a **consumed_at honesty gauge**: an unconsumed secret past a 15-min grace under a box reporting `applied` — the exact July-18 fingerprint, a disagreement no single tier can see — is surfaced with its own event, deliberately as a SURFACE not a heal (minting on top of an unconsumed secret is the R-39(a) race). **A LOAD-BEARING FACT the spec did not flag, checked rather than trusted:** `Apply` bails out if the storage status probe ERRORS, and `adopt` converges WITHOUT consuming when the storage reads active — so the whole fix depended on PVE's 401 behaviour. PVE's `storage_info` wraps `activate_storage`/`$plugin->status` in `eval{}` and leaves the pre-initialised `active => 0`, so a 401 returns **HTTP 200 with `active: 0`, never an API error** — `Apply` correctly falls through to verify → consume → reconcile. **A defect I shipped and caught:** v0.91.0 built the probe seam and `main.go` never called `SetAuthSink`, so the whole leg was INERT and every test still passed (the seam was injected directly) — same class as controller v0.154.0 the day before; fixed in v0.91.1, artifact superseded not overwritten, and v0.91.2 made a healthy probe observable so "no auth_failed" can never again be confused with "never probed". Four red-proofs, all at the assertion level. **CLOSED 2026-07-21 — STOP-2 and STOP-3 both done, chain proven end to end in 13 seconds.** STOP-3 saved (agent 0.91.2 / sha `34d309be…` / **wrapper sha `104db0a4…`** / MinAgent 0.91.2; the agent reports the same wrapper hash so the R-50b drift gauge reads ok). STOP-2: hub `08:39:31Z` fresh mint **generation 0 → 1** + descriptor `"secret_generation": 1` (token_id/fingerprint byte-identical — the re-key shape that used to be invisible) → agent `10:39:34` reads the credential via the wrapper → `10:39:38` **`REJECTED … applied and DEAD` `previous_state=applied`** (the R-39 state, loud for the first time) → `10:39:45` **`one-time token secret consumed`** (NO short-circuit — the line that never appeared on 2026-07-18) → `10:39:45` set-only `reconcile` → `10:39:47` **`converged state=applied`**. Marker hash moved to `afbb3b41…`; `consumed_at` stamped; secret mtime `2026-07-18` → `2026-07-21 10:39:45`; new credential probes **200**; reports trace `applied → auth_failed → applied`; **zero** self-heal escalations, ONE mint, ONE consume, no `consumed-failed.json`. **A NOTE ON THE FIRST ATTEMPT:** the operator initially pressed the **offsite** re-issue (two different Re-issue actions exist) — harmless to PBS-DR, but it rotated the restic password and correctly marked the escrow STALE, so the recovery-code ceremony had to be re-run. Worth naming the surface explicitly in any future runbook step. Evidence: `felhom-agent/REPORT.md`, `felhom.eu/REPORT.md`. *(Superseded — the spec was written and shipped:)* **REMAINING (fleet, needs its own spec — deliberately NOT improvised):** (a) make a fresh unconsumed secret actually un-converge the agent — either `consumed_at` becomes authoritative or the descriptor carries a secret generation/nonce so the hash moves; (b) fix the verify loop's read path — it reads `/etc/pve/priv/storage/.pw` **directly as non-root**, a file it can only ever *write* through the root wrapper (`/etc/pve/priv` is `0700 root:www-data`; sudoers exposes `create|reconcile|grant` and **no read verb**), so the one loop that could catch this is permanently blind; (c) an auth probe in the reconciler/gauge/ceremony-precheck so `applied` can never mean `401`. — **Original finding note:** discovered by CC while collecting Phase-A evidence; not on the brief's finding list, so its P2-HIGH rank is provisional pending Viktor. The severity case: this is the DR tier, the failure is silent, and it would surface first at a real restore. Suggested shape: make `consumed_at` authoritative (a fresh unconsumed secret must un-converge the reconciler), fix the verify loop's read path (read via the same root wrapper that writes it), and make a failing `pvesm status` a LOUD state rather than a skipped datastore. **Blocks the DR-tier map row's candidate upgrade to PROVEN-LIVE — that upgrade is now explicitly WITHDRAWN.** Evidence `pbs-dr-state.txt`, `hub-state.txt` | | R-30 | **[P2-HIGH] Liveness presence should come from the wait channel, not the report clock.** The box was powered off at the start of the rehearsal, yet the hub carried it as healthy until the staleness threshold expired ~30 min later (`host_stale` 16:05:24 "no report for 30m"; cleared 16:33:24 "was stale for 27m"). The host-delete guard compounds it: RESET refuses while any host row exists, so a stale-but-"Online" host stalls a forced teardown. | M | idea | Direction: derive presence from **Dir-2 long-poll connectedness (~90 s grace)**, decoupled from notification hysteresis (the hysteresis is right for *alerting*, wrong for *presence*); an agent/ep0 analog can follow. Pairs with R-13/R-23 — the transport already exists, this is about believing it. *(Discussed in-session as "R-29"; that number was already taken by the gate-rot item earlier the same day, so it is R-30.)* | | R-31 | **[P2-HIGH] Offsite provisioning is synchronous with no status affordance.** Save runs the Hetzner sync in-request, so the request can hit the nginx 504 **while succeeding server-side**: the operator cannot tell failed from slow, and a retry races the first attempt. | M | idea | Direction: make it async + a status card, reusing the proven **awaiting-card/poll idiom** (v0.138.0 escrow card). **Interim mitigation belongs in R-3 as an operator note: click once, wait, verify — do not re-click.** | | R-32 | **[P2-HIGH] RESET must purge the customer base dir; the orphan card must stay honest; unattributed bytes must be visible.** The rehearsal's S7 said in advance that an orphan card would BE a finding — and one appeared (16:58:14). Cause: RESET's `"hetzner":"ok"` leg destroys the sub-account, but **a Hetzner sub-account is an access-control object, not a data object** — its directory survives, so re-enabling offsite recreated an account over the previous lifecycle's ciphertext, encrypted under a key that same RESET had destroyed. | M | idea | **Ruling from the run (three parts, deliberately separate):** (1) because RESET destroys custody, the ciphertext it leaves behind is unrecoverable **BY DESIGN** → RESET gains a **main-account purge of the customer base dir** (the existing operator ack already covers it); (2) the **move-aside guard STAYS** for reinstall-*without*-RESET — there custody survives and the card's "history recoverable" promise is true (R-26 depends on exactly that); (3) the operator **Restic tab shows per-customer directory bytes vs attributed snapshot bytes**, so dead data cannot hide. Measured on the pool box that night: **49 M attributed** (2 snapshots, 48.717 MiB) against **1.4 G + 3.0 M unattributed** across TWO `.orphaned-*` dirs. Evidence `restic-and-pool.txt` |