# SPIKE — PBS-DR day-0 self-heal: root cause + re-trigger validation — 2026-07-15 **Class:** spike (findings only — no production code, config, or schema edits; the only non-scratch write is this doc + its commit to `felhom.eu` main). Gates the **PBS-DR self-heal TASK** spec, which is BLOCKED until these verdicts land. All hub-DB mutations were throwaway probes scoped to the drill host; every one is logged + reverted/self-reverted below (§Mutations). Evidence sink: `180:~/spikes/pbsdr-selfheal-2026-07-15/`. ## Verdict up front 1. **ROOT CAUSE (SQ-1) — the prompt's surviving theory is REFUTED by live evidence.** The descriptor is **NOT** absent. On the drill host the `pbs_dr` descriptor is **present and enabled** in `desired_json`, the WG peer **exists**, and the fingerprint **verifies over the tunnel**. The block is that the **one-time PBS secret was consumed on 2026-07-12 and no code path ever re-mints or re-stages it** for a host that has since lost its agent-side converged marker (fresh day-0 / re-install / snapshot rollback). The agent sits in `pbs_dr.state="waiting_secret"` **forever**, re-applying the descriptor every 60 s tick but unable to consume a burned secret. The reused-peer angle is real but its consequence is *not* "descriptor never written" — it is "the WG-registration cascade cannot re-fire (same pubkey → `changed==false`), so nothing re-mints a consumable secret." 2. **THE MISSING PIECE IS A CONSUMABLE SECRET, NOT THE DESCRIPTOR (SQ-2, live-proven).** Re-asserting the descriptor alone (generation bump) converges **nothing** (negative control). Staging **any** consumable secret converges the box in **one tick**: re-exposing the *existing, still-valid* ep0 token secret (`consumed_at → NULL`, **zero ep0 churn**) made the agent verify → consume → create the `felhom-pbs` storage entry → mark converged within ~30 s. The operator **"Re-issue PBS credentials"** action IS the manual recovery; its DB effect is a **superset** of that proven re-stage (fresh secret + descriptor bump). 3. **A PERIODIC RE-TRIGGER IS SAFE — IF it re-serves a secret and never blind-timer-reissues (SQ-3).** Re-asserting a *converged* descriptor is a clean idempotent no-op (live-proven: marker untouched, no re-consume, no generation thrash — the `descriptorHash` marker gates it). Re-running the **provision atom** does NOT churn ep0 (`Provision` short-circuits with `ErrTokenExists`, mints nothing) — but it also does NOT fix a stuck host (it refuses with "use Re-issue"). **`Reissue` rotates the ep0 token every call** — safe against orphans (delete+recreate) but a blind-timer `Reissue` = descriptor-hash thrash + constant re-consume. Safe design: **gate on the agent's reported `waiting_secret`/`consumed_failed` state, re-stage the stored secret first (no churn, no gen bump), escalate to `Reissue` only if that fails.** 4. **⚠️ ARCHITECTURE IMPACT.** There is **no automatic recovery** for the single most common production event — a customer box re-installed / restored / rolled back onto its stable `host_id`. The hub keeps the durable descriptor + a durable *consumed* secret; the WG hook can't re-fire; `applyPBSDR`'s "already provisioned → no-op" branch means even a config re-save won't re-mint. The PBS-DR tier silently stays unconverged (and therefore escrow + offsite never arm) until an operator notices and clicks Re-issue. This is the P-DAY0-DEEP block, generalized. --- ## 0. Baselines (live-verified at session start) | Item | Value | How verified | |---|---|---| | `felhom.eu` main head | **`aff272c4be40da33855c0c45f7e66f407f33c99c`** (CAMPAIGN-6D P-DAY0 doc) | `git log -1` local == pulled on 180 | | Hub live image / version | **`felhom-hub:0.55.0`** (`hub-6fddb45766-k9f7f`, Running 9 h) | `kubectl get deploy hub -o jsonpath`; startup log `felhom-hub 0.55.0 starting` | | Agent live version (drill) | **`felhom-agent 0.88.0`** | `felhom-agent --version` on drill VM `192.168.0.152` | | Controller golden | **0.136.0** (not exercised here) | baseline per prompt | | Drill host (hub `host_id`) | **`demo-vm-felhom-2f4b00`** / customer `demo-vm-felhom`, gen 3 at start | `hosts` table | | Drill guest (Proxmox) | **VM 300 `drill-day0`** on demo-pve (`192.168.0.152`, MAC `bc:24:11:85:98:7b`) | `qm status 300`; ARP on felhom-pve | | Rollback point | snapshot **`post_day0_golden136`** (2026-07-15 15:32:02) — *"golden 0.136.0 + agent 0.88.0 … escrow pending (no DR chain)"* | `qm listsnapshot 300` | **Read seam (both directions):** a throwaway pure-Go tool `hubdbpoke` (built on 180 inside the hub module so `modernc.org/sqlite` resolves; `CGO_ENABLED=0`), `kubectl cp`'d into the hub pod, run against `/data/hub.db` — `q` = read query, `e` = exec. Removed from the pod at session end; source never committed. Secret **values** were never selected (only `length(value)` + timestamps). --- ## The two theories the reviewer already refuted (restated with file:line, both CONFIRMED here) - **(a) "the agent lacks self-heal" — FALSE.** `pbsdr/loop.go:62-77` re-runs `Manager.Apply` on every 60 s tick *and* on each desired-state nudge. `pbsdr/manager.go:265-272` (verify-pin-before-consume): a probe failure leaves the one-time secret **untouched** and returns, so a slow/flapping tunnel self-heals **as long as a consumable secret is in hand**. Live-proven: the drill agent oscillated through `verify_failed` (tunnel i/o timeout) and recovered without operator help — see §SQ-1 note. - **(b) "the hub provision atom probes the customer tunnel" — FALSE.** `web/pbsdr.go:184-198`: the atom's preconditions are all **hub-DB-local** (`tenantsync!=nil`, `GetWGPeerForHost`, `GetWGEndpoint`). It never dials the box. The `i/o timeout` in the field is therefore the **agent's** PBS fingerprint probe (`dial 10.77.0.1:8007`), not a hub failure — confirmed verbatim in the drill agent journal (§SQ-1). --- ## SQ-1 — delivery-path root cause **Question:** is the descriptor's only write path the first-WG cascade — and on the drill host, is the descriptor absent (cascade never fired) or present-but-something-else? **Probes (hubdbpoke `q` against live `/data/hub.db`, scoped `host_id='demo-vm-felhom-2f4b00'`):** | # | Query | Verbatim result | |---|---|---| | a | `wg_peers WHERE host_id=…` | 1 row: pubkey `zxy+y029eYri+g8HhP4qGwep/Nmv8VR7l3iyf39ovmo=`, ip `10.77.0.3`, created `2026-07-12 19:50:52`, updated `2026-07-15 13:14:43` | | b | `desired_json` | `{"pbs_dr":{"enabled":true,"storage_id":"felhom-pbs","pbs_tunnel_ip":"10.77.0.1","datastore":"felhom-offsite","namespace":"demo-vm-felhom","token_id":"felhom@pbs!demo-vm-felhom","fingerprint":"c6:07:…:fd"}}` | | c | `host_pbs_secrets` (len only) | 1 row: `val_len=36`, created `2026-07-12 19:50:53`, **`consumed_at=2026-07-12 20:05:53`** | | d | agent pbsdr status (latest host-report) | `"pbs_dr":{"state":"waiting_secret","storage_id":"felhom-pbs","namespace":"demo-vm-felhom","message":"verified; no unconsumed token secret staged on the hub"}` | | e | `customer_configs.dr_tier` | `1` (DR tier ON) | **Agent journal (drill VM, local CEST = UTC+2), verbatim:** ``` 17:07:26 WARN pbsdr: PBS fingerprint verify failed BEFORE consume (nothing consumed; retrying) err="pbs: fingerprint probe dial 10.77.0.1:8007: i/o timeout" 15:14:42 INFO pbsdr: bridge enabled (hub-driven; no-op until a pbs_dr descriptor arrives) ``` At session time the tunnel was **up** (`wg show`: handshake 10 s ago; `10.77.0.1:8007` OPEN 3/3). No `marker.json` and no `felhom-pbs` entry in `/etc/pve/storage.cfg` existed → the agent had **never** gotten past consume on this box's current life. **Forensic tie-off (reused peer):** the agent's own WG pubkey on the drill VM is `zxy+y029eYri+g8HhP4qGwep/Nmv8VR7l3iyf39ovmo=` — **identical** to the stored `wg_peers` row. The key survives the snapshot/re-install, so on re-registration `RegisterWGPeerForHost` returns `changed==false` → the `if changed { … wgRegisteredHook(…) }` block (`api/wg.go:288-305`) never fires → `PBSDRAutoProvision` (`web/pbsdr.go:274-300`) never runs → no new secret. **Verdict:** the descriptor is **PRESENT and enabled** (delivered correctly at first enrollment on 07-12, and durable in `desired_json` ever since). The block is that the **one-time secret was consumed on 07-12 and is never re-minted/re-staged** after the box lost its converged marker (the `post_day0_golden136` snapshot is a *fresh* day-0 with "no DR chain", branched off `pre-day0-clean`). The agent is permanently `waiting_secret`. The transient `verify_failed`/i/o-timeout is a separate, self-healing environmental flap of the nested drill VM's WG tunnel — **not** the root cause. --- ## SQ-2 — prove the gap directly + validate the re-trigger The literal SQ-2a premise ("`host_pbs_secrets` holds an UNCONSUMED secret") is **false** — it is consumed. Two probes isolate the exact missing piece. **SQ-2a′ (negative control — descriptor re-assert only).** MUT-1: `desired_generation` 3→4 (re-serve the identical descriptor; secret untouched). Result: the agent kept reporting `waiting_secret`; the descriptor had *already* been re-applied on every 60 s tick since rollback with no convergence. **A generation bump / descriptor re-delivery changes nothing.** → descriptor delivery is **not** the gap. **SQ-2b′ (stage a consumable secret — zero ep0 churn).** Tunnel confirmed up first. MUT-2: `host_pbs_secrets.consumed_at → NULL` (re-expose the *existing* still-valid ep0 token secret; no `tenantsync` call, no new token). Within ~30 s the drill agent journal showed, verbatim: ``` 17:44:51 INFO pbsdr: one-time token secret consumed (single-use; value withheld from logs) storage_id=felhom-pbs secret_len=36 17:44:53 INFO pbsdr: converged state=applied storage_id=felhom-pbs ``` `marker.json` = `{"hash":"f1059b54…","state":"applied","applied_at":"2026-07-15T15:44:53Z"}`; `felhom-pbs` now PRESENT in `/etc/pve/storage.cfg`; hub `consumed_at` self-re-set to `15:44:51` (single-use tx). **The agent self-heals fully the instant a consumable secret exists** — the descriptor was never the missing thing. **Re-issue (SQ-2b, the literal operator action):** validated from code rather than fired live — the web handler `POST /configs/{id}/pbsdr-reissue` (`server.go:456-463` → `handlePBSDRReissue`, `web/pbsdr.go:305-355`) sits behind the password-gated operator session + CSRF (not drivable headless), and a live run would rotate an ep0 token against the STOP condition. Its DB effect is a **superset** of the proven re-stage: `tenantsync.Reissue` (fresh token) → `SaveHostPBSSecret` (new secret, `consumed_at` reset) → descriptor rewrite with the **new** `token_id`/`fingerprint` + bump (`web/pbsdr.go:335-346`). Since a fresh consumable secret is exactly what SQ-2b′ proved sufficient, Re-issue converges a reused-peer guest. Use it as the escalation when the stored secret is stale. **Verdict:** descriptor-only delivery converges **NO**; staging a consumable secret converges **YES** (one tick, existing token, no churn); Re-issue converges **YES** (superset, code-proven). **The fix is to re-serve a consumable secret** — re-stage the stored one (least-churn) or Re-issue (escalation) — **not** to re-assert the descriptor. --- ## SQ-3 — is a periodic re-trigger safe? **SQ-3a — re-serve for an ALREADY-CONVERGED host (live-proven).** MUT-3: `desired_generation` 4→5 on the now-converged host; waited ~75 s (≥1 tick). Result: `marker.json` **unchanged** (`applied_at` still `15:44:53Z`), **no** new `consumed`/`converged` journal line, hub `consumed_at` **unchanged** (`15:44:51`). The `descriptorHash` idempotency short-circuit (`manager.go:233-238`: `mk.Hash==h && cf==nil` → set status + return) makes re-serving the same descriptor a **clean no-op** — no re-consume, no re-adopt, **no generation thrash on the agent side**. **SQ-3b — the one-shot-secret interaction + ep0 churn (code-derived; STOP-condition forbids spraying tokens).** - `store/pbsdr.go:10-17` `SaveHostPBSSecret` = last-write-wins, resets `consumed_at`. It is only reached **after** a successful `Provision`/`Reissue`. - `tenantsync/client.go:93-95,151-155` `Provision` returns typed **`ErrTokenExists`** for an existing token — it **mints nothing** and never reaches `SaveHostPBSSecret`. So **re-running the provision atom does NOT churn ep0 and does NOT stomp** — but (per `web/pbsdr.go:206-222`) for a provisioned-and-not-acked-deleted host it **returns a loud "use Re-issue" error and does not fix the stuck host**. Re-running the atom is therefore the *wrong* re-trigger for this root cause. - `tenantsync/client.go:97-100` `Reissue` = delete+recreate the token (one rotation per call; no orphan accumulation). A **blind-timer** `Reissue` would rotate the token every tick → the descriptor hash (`token_id`/`fingerprint`) changes every tick → the agent re-consumes + re-reconciles every tick = **generation + apply thrash**. This is the real hazard to avoid. - **Least-churn primary:** re-stage the stored secret (`consumed_at → NULL`). Zero ep0 interaction, no hash change; proven in SQ-2b′. Valid whenever the stored token is still good on ep0 (the common re-install/rollback case — the token was never deleted). **SQ-3c — the `consumed-failed` dead-end (code-derived from `manager.go`).** - Same-hash re-serve **stays put**: with `consumed-failed.json` present and `cf.Hash==h`, the idempotent short-circuit is skipped (`manager.go:235`), and when `ConsumePBSToken` returns `ErrNoPBSSecret` the `cf.Hash==h` branch (`:275-281`) re-asserts the **LOUD** `consumed_failed` state — never a silent burned-secret retry. Matches the package law comment (`:17-19`). - **Re-issue breaks out** because it does two things at once: stages a **fresh** secret **and** changes the descriptor (new `token_id`/`fingerprint` → new hash `h'`). With `h'≠cf.Hash`, the dead-end gate no longer matches → the agent re-attempts, consume now succeeds, `finishConverged` runs and **`os.Remove(consumedFailedPath())`** clears the dead-end (`manager.go:348-362`). - Note (design input): **re-staging the stored secret also recovers `consumed_failed`** *when the stored token is still valid* — because `ConsumePBSToken` then **succeeds** (not `ErrNoPBSSecret`), so the dead-end branch is never entered and the apply proceeds. Only a genuinely stale/rotated ep0 token forces the `Reissue` escalation. **Verdict:** a safe re-trigger = **re-assert-secret, gated on the agent's reported stuck state, never blind-timer-reissue.** Re-staging the stored secret needs **no** generation bump (SQ-2b′ converged on the agent's own ticker); a generation bump must be reserved for an **actual descriptor content change** (Reissue / storage-id edit) — reuse the existing `readPBSDR`+compare no-op discipline (`web/pbsdr.go:148-164`). --- ## SQ-4 — where the re-trigger lives + the condition **Recommended design (spike-proven, one paragraph):** add a **hub periodic reconciler** (mirroring `wgsync/reconciler.go`) that, for each host with `dr_tier=1` and an **enabled `pbs_dr` descriptor**, reads the latest host-report `pbs_dr.state`; when that state is **`waiting_secret` or `consumed_failed` sustained across ≥1 report cycle** (debounce, so a transient `verify_failed` tunnel flap is *not* acted on), it **re-stages the stored one-time secret** (`consumed_at → NULL`, **no generation bump**) and records an event. If the host is still stuck after N recon_cycles (stored token stale → `consumed_failed` persists), it **escalates once to `Reissue`** (fresh token + descriptor bump). The capability/host-report path is the only trigger that can observe "DR-ON but agent reports pbsdr inactive" — the WG hook is structurally unable to help (same pubkey → `changed==false`), and config-save is operator-manual. **The one spike-proven reason this is correct:** convergence requires only a *consumable secret*, and staging the existing stored secret is a zero-ep0-churn, no-hash-change, no-gen-bump operation that the agent picks up on its own 60 s tick (SQ-2b′) and that no-ops cleanly once converged (SQ-3a) — so the reconciler self-limits to exactly the stuck hosts and stops the moment they converge. **Generation-bump discipline (exact condition):** bump **only** on a real descriptor **content** change (Reissue's new `token_id`/`fingerprint`, or a `storage_id` edit) — never for a secret re-stage. This is the existing `applyPBSDR` precedent: `readPBSDR` + field compare, no-op-if-equal (`web/pbsdr.go:148-164`); the agent re-fetches on generation change, so a spurious bump every tick is an agent-refetch loop. ### Design inputs for the PBS-DR self-heal TASK | Dimension | Spike-proven answer | |---|---| | **Root cause** | Durable hub descriptor + durable *consumed* one-time secret + stable `host_id`, after the agent loses its converged marker (re-install / rollback / disk loss). No path re-mints/re-stages the secret; WG cascade can't re-fire (`changed==false`); `applyPBSDR` "already provisioned" no-ops. Agent stuck in `waiting_secret`. | | **Safe re-trigger shape** | Re-serve a **consumable secret**, gated on the agent's reported `waiting_secret`/`consumed_failed` (debounced ≥1 cycle). Primary: re-stage stored secret (`consumed_at→NULL`, no ep0 call, no gen bump). Escalation (stored token stale): `Reissue` once. **Never** re-run the provision atom (refuses, doesn't fix) and **never** blind-timer `Reissue` (hash/gen thrash). | | **Trigger location** | Hub **periodic reconciler** reading host-report `pbs_dr.state` (like `wgsync/reconciler.go`). Not the WG hook (structurally can't re-fire); not config-save (manual). | | **Gen-bump condition** | Only on a real descriptor **content** change (Reissue/storage-id edit) via `readPBSDR`+compare no-op discipline (`web/pbsdr.go:148-164`). A secret re-stage bumps **nothing**. | | **Orphan-token cleanup** | **None needed** for the primary path (re-stage = no ep0 interaction) or repeat `Provision` (`ErrTokenExists`, mints nothing). `Reissue` is delete+recreate (no accumulation). The only orphan risk is a *blind-timer* Reissue design — explicitly excluded above. | --- ## Mutations made + reverts (all scoped `host_id='demo-vm-felhom-2f4b00'`; SELECT-verified 1 row before each) | ID | Mutation | Revert | |---|---|---| | MUT-1 | `UPDATE hosts SET desired_generation+1` (3→4) — SQ-2a′ | none needed (monotonic counter; agent compares, absolute value irrelevant) | | MUT-2 | `UPDATE host_pbs_secrets SET consumed_at=NULL WHERE … AND consumed_at IS NOT NULL` — SQ-2b′ re-stage of the **existing** token secret (no ep0 churn) | **self-reverted**: agent re-consumed at `15:44:51Z`, `consumed_at` re-set by `ConsumeHostPBSSecret` | | MUT-3 | `UPDATE hosts SET desired_generation+1` (4→5) — SQ-3a | none needed (monotonic) | **Net DB delta vs spike start:** `host_pbs_secrets.consumed_at` `2026-07-12T20:05:53Z` → `2026-07-15T15:44:51Z` (secret re-consumed *by real convergence*); `hosts.desired_generation` 3 → 5. **No orphaned/unconsumed secrets, no ep0 token churn** (existing token reused throughout). No non-drill host touched (every statement `WHERE host_id='demo-vm-felhom-2f4b00'`; the demo host `demo-felhom-01` was never in a WHERE clause). Throwaway `hubdbpoke` binary removed from the hub pod. ## Guest end-state **LEFT CONVERGED** (not rolled back). PBS-DR on the drill guest is now genuinely `state=applied` (`felhom-pbs` storage entry created, marker `f1059b54…`) — a real side-benefit: the P-DAY0-DEEP PBS-DR leg is now GREEN on the drill (escrow ceremony + offsite arming remain separate follow-ups). Rollback to `post_day0_golden136` is available if a pristine "no DR chain" fixture is wanted again. ## STOP conditions honored Every WHERE clause proven to hit only the drill host (SELECT-before-UPDATE, `count(*)=1`). No code fix attempted (that is the spec). No ep0 token churn — the primary proof deliberately reused the existing token; `Reissue` was code-validated, not fired. No non-drill/production host touched.