Refutes the surviving theory: on the drill host the pbs_dr descriptor is present and enabled, the WG peer exists, verify passes — the block is a one-time secret consumed 07-12 that no path re-mints after the agent loses its converged marker (re-install/rollback; same WG pubkey -> changed==false -> cascade can't re-fire). Live-proven: staging any consumable secret converges in one tick (existing token, zero ep0 churn); re-asserting a converged descriptor is a clean idempotent no-op. Safe re-trigger = re-serve a secret gated on agent waiting_secret/consumed_failed, never blind-timer Reissue. Design-inputs table handed to the self-heal TASK spec. Drill left CONVERGED (P-DAY0-DEEP PBS-DR leg now GREEN). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01HEPuEwyyGDJdcsXLFsTWJn
20 KiB
SPIKE — PBS-DR day-0 self-heal: root cause + re-trigger validation — 2026-07-15
Class: spike (findings only — no production code, config, or schema edits; the only non-scratch
write is this doc + its commit to felhom.eu main). Gates the PBS-DR self-heal TASK spec, which
is BLOCKED until these verdicts land. All hub-DB mutations were throwaway probes scoped to the drill
host; every one is logged + reverted/self-reverted below (§Mutations). Evidence sink:
180:~/spikes/pbsdr-selfheal-2026-07-15/.
Verdict up front
-
ROOT CAUSE (SQ-1) — the prompt's surviving theory is REFUTED by live evidence. The descriptor is NOT absent. On the drill host the
pbs_drdescriptor is present and enabled indesired_json, the WG peer exists, and the fingerprint verifies over the tunnel. The block is that the one-time PBS secret was consumed on 2026-07-12 and no code path ever re-mints or re-stages it for a host that has since lost its agent-side converged marker (fresh day-0 / re-install / snapshot rollback). The agent sits inpbs_dr.state="waiting_secret"forever, re-applying the descriptor every 60 s tick but unable to consume a burned secret. The reused-peer angle is real but its consequence is not "descriptor never written" — it is "the WG-registration cascade cannot re-fire (same pubkey →changed==false), so nothing re-mints a consumable secret." -
THE MISSING PIECE IS A CONSUMABLE SECRET, NOT THE DESCRIPTOR (SQ-2, live-proven). Re-asserting the descriptor alone (generation bump) converges nothing (negative control). Staging any consumable secret converges the box in one tick: re-exposing the existing, still-valid ep0 token secret (
consumed_at → NULL, zero ep0 churn) made the agent verify → consume → create thefelhom-pbsstorage entry → mark converged within ~30 s. The operator "Re-issue PBS credentials" action IS the manual recovery; its DB effect is a superset of that proven re-stage (fresh secret + descriptor bump). -
A PERIODIC RE-TRIGGER IS SAFE — IF it re-serves a secret and never blind-timer-reissues (SQ-3). Re-asserting a converged descriptor is a clean idempotent no-op (live-proven: marker untouched, no re-consume, no generation thrash — the
descriptorHashmarker gates it). Re-running the provision atom does NOT churn ep0 (Provisionshort-circuits withErrTokenExists, mints nothing) — but it also does NOT fix a stuck host (it refuses with "use Re-issue").Reissuerotates the ep0 token every call — safe against orphans (delete+recreate) but a blind-timerReissue= descriptor-hash thrash + constant re-consume. Safe design: gate on the agent's reportedwaiting_secret/consumed_failedstate, re-stage the stored secret first (no churn, no gen bump), escalate toReissueonly if that fails. -
⚠️ ARCHITECTURE IMPACT. There is no automatic recovery for the single most common production event — a customer box re-installed / restored / rolled back onto its stable
host_id. The hub keeps the durable descriptor + a durable consumed secret; the WG hook can't re-fire;applyPBSDR's "already provisioned → no-op" branch means even a config re-save won't re-mint. The PBS-DR tier silently stays unconverged (and therefore escrow + offsite never arm) until an operator notices and clicks Re-issue. This is the P-DAY0-DEEP block, generalized.
0. Baselines (live-verified at session start)
| Item | Value | How verified |
|---|---|---|
felhom.eu main head |
aff272c4be40da33855c0c45f7e66f407f33c99c (CAMPAIGN-6D P-DAY0 doc) |
git log -1 local == pulled on 180 |
| Hub live image / version | felhom-hub:0.55.0 (hub-6fddb45766-k9f7f, Running 9 h) |
kubectl get deploy hub -o jsonpath; startup log felhom-hub 0.55.0 starting |
| Agent live version (drill) | felhom-agent 0.88.0 |
felhom-agent --version on drill VM 192.168.0.152 |
| Controller golden | 0.136.0 (not exercised here) | baseline per prompt |
Drill host (hub host_id) |
demo-vm-felhom-2f4b00 / customer demo-vm-felhom, gen 3 at start |
hosts table |
| Drill guest (Proxmox) | VM 300 drill-day0 on demo-pve (192.168.0.152, MAC bc:24:11:85:98:7b) |
qm status 300; ARP on felhom-pve |
| Rollback point | snapshot post_day0_golden136 (2026-07-15 15:32:02) — "golden 0.136.0 + agent 0.88.0 … escrow pending (no DR chain)" |
qm listsnapshot 300 |
Read seam (both directions): a throwaway pure-Go tool hubdbpoke (built on 180 inside the hub
module so modernc.org/sqlite resolves; CGO_ENABLED=0), kubectl cp'd into the hub pod, run
against /data/hub.db — q = read query, e = exec. Removed from the pod at session end; source
never committed. Secret values were never selected (only length(value) + timestamps).
The two theories the reviewer already refuted (restated with file:line, both CONFIRMED here)
-
(a) "the agent lacks self-heal" — FALSE.
pbsdr/loop.go:62-77re-runsManager.Applyon every 60 s tick and on each desired-state nudge.pbsdr/manager.go:265-272(verify-pin-before-consume): a probe failure leaves the one-time secret untouched and returns, so a slow/flapping tunnel self-heals as long as a consumable secret is in hand. Live-proven: the drill agent oscillated throughverify_failed(tunnel i/o timeout) and recovered without operator help — see §SQ-1 note. -
(b) "the hub provision atom probes the customer tunnel" — FALSE.
web/pbsdr.go:184-198: the atom's preconditions are all hub-DB-local (tenantsync!=nil,GetWGPeerForHost,GetWGEndpoint). It never dials the box. Thei/o timeoutin the field is therefore the agent's PBS fingerprint probe (dial 10.77.0.1:8007), not a hub failure — confirmed verbatim in the drill agent journal (§SQ-1).
SQ-1 — delivery-path root cause
Question: is the descriptor's only write path the first-WG cascade — and on the drill host, is the descriptor absent (cascade never fired) or present-but-something-else?
Probes (hubdbpoke q against live /data/hub.db, scoped host_id='demo-vm-felhom-2f4b00'):
| # | Query | Verbatim result |
|---|---|---|
| a | wg_peers WHERE host_id=… |
1 row: pubkey zxy+y029eYri+g8HhP4qGwep/Nmv8VR7l3iyf39ovmo=, ip 10.77.0.3, created 2026-07-12 19:50:52, updated 2026-07-15 13:14:43 |
| b | desired_json |
{"pbs_dr":{"enabled":true,"storage_id":"felhom-pbs","pbs_tunnel_ip":"10.77.0.1","datastore":"felhom-offsite","namespace":"demo-vm-felhom","token_id":"felhom@pbs!demo-vm-felhom","fingerprint":"c6:07:…:fd"}} |
| c | host_pbs_secrets (len only) |
1 row: val_len=36, created 2026-07-12 19:50:53, consumed_at=2026-07-12 20:05:53 |
| d | agent pbsdr status (latest host-report) | "pbs_dr":{"state":"waiting_secret","storage_id":"felhom-pbs","namespace":"demo-vm-felhom","message":"verified; no unconsumed token secret staged on the hub"} |
| e | customer_configs.dr_tier |
1 (DR tier ON) |
Agent journal (drill VM, local CEST = UTC+2), verbatim:
17:07:26 WARN pbsdr: PBS fingerprint verify failed BEFORE consume (nothing consumed; retrying)
err="pbs: fingerprint probe dial 10.77.0.1:8007: i/o timeout"
15:14:42 INFO pbsdr: bridge enabled (hub-driven; no-op until a pbs_dr descriptor arrives)
At session time the tunnel was up (wg show: handshake 10 s ago; 10.77.0.1:8007 OPEN 3/3). No
marker.json and no felhom-pbs entry in /etc/pve/storage.cfg existed → the agent had never
gotten past consume on this box's current life.
Forensic tie-off (reused peer): the agent's own WG pubkey on the drill VM is
zxy+y029eYri+g8HhP4qGwep/Nmv8VR7l3iyf39ovmo= — identical to the stored wg_peers row. The key
survives the snapshot/re-install, so on re-registration RegisterWGPeerForHost returns
changed==false → the if changed { … wgRegisteredHook(…) } block (api/wg.go:288-305) never fires
→ PBSDRAutoProvision (web/pbsdr.go:274-300) never runs → no new secret.
Verdict: the descriptor is PRESENT and enabled (delivered correctly at first enrollment on
07-12, and durable in desired_json ever since). The block is that the one-time secret was consumed
on 07-12 and is never re-minted/re-staged after the box lost its converged marker (the
post_day0_golden136 snapshot is a fresh day-0 with "no DR chain", branched off pre-day0-clean).
The agent is permanently waiting_secret. The transient verify_failed/i/o-timeout is a separate,
self-healing environmental flap of the nested drill VM's WG tunnel — not the root cause.
SQ-2 — prove the gap directly + validate the re-trigger
The literal SQ-2a premise ("host_pbs_secrets holds an UNCONSUMED secret") is false — it is
consumed. Two probes isolate the exact missing piece.
SQ-2a′ (negative control — descriptor re-assert only). MUT-1: desired_generation 3→4 (re-serve
the identical descriptor; secret untouched). Result: the agent kept reporting waiting_secret; the
descriptor had already been re-applied on every 60 s tick since rollback with no convergence. A
generation bump / descriptor re-delivery changes nothing. → descriptor delivery is not the gap.
SQ-2b′ (stage a consumable secret — zero ep0 churn). Tunnel confirmed up first. MUT-2:
host_pbs_secrets.consumed_at → NULL (re-expose the existing still-valid ep0 token secret; no
tenantsync call, no new token). Within ~30 s the drill agent journal showed, verbatim:
17:44:51 INFO pbsdr: one-time token secret consumed (single-use; value withheld from logs) storage_id=felhom-pbs secret_len=36
17:44:53 INFO pbsdr: converged state=applied storage_id=felhom-pbs
marker.json = {"hash":"f1059b54…","state":"applied","applied_at":"2026-07-15T15:44:53Z"};
felhom-pbs now PRESENT in /etc/pve/storage.cfg; hub consumed_at self-re-set to 15:44:51
(single-use tx). The agent self-heals fully the instant a consumable secret exists — the descriptor
was never the missing thing.
Re-issue (SQ-2b, the literal operator action): validated from code rather than fired live — the
web handler POST /configs/{id}/pbsdr-reissue (server.go:456-463 → handlePBSDRReissue,
web/pbsdr.go:305-355) sits behind the password-gated operator session + CSRF (not drivable
headless), and a live run would rotate an ep0 token against the STOP condition. Its DB effect is a
superset of the proven re-stage: tenantsync.Reissue (fresh token) → SaveHostPBSSecret (new
secret, consumed_at reset) → descriptor rewrite with the new token_id/fingerprint + bump
(web/pbsdr.go:335-346). Since a fresh consumable secret is exactly what SQ-2b′ proved sufficient,
Re-issue converges a reused-peer guest. Use it as the escalation when the stored secret is stale.
Verdict: descriptor-only delivery converges NO; staging a consumable secret converges YES (one tick, existing token, no churn); Re-issue converges YES (superset, code-proven). The fix is to re-serve a consumable secret — re-stage the stored one (least-churn) or Re-issue (escalation) — not to re-assert the descriptor.
SQ-3 — is a periodic re-trigger safe?
SQ-3a — re-serve for an ALREADY-CONVERGED host (live-proven). MUT-3: desired_generation 4→5 on
the now-converged host; waited ~75 s (≥1 tick). Result: marker.json unchanged
(applied_at still 15:44:53Z), no new consumed/converged journal line, hub consumed_at
unchanged (15:44:51). The descriptorHash idempotency short-circuit (manager.go:233-238:
mk.Hash==h && cf==nil → set status + return) makes re-serving the same descriptor a clean no-op
— no re-consume, no re-adopt, no generation thrash on the agent side.
SQ-3b — the one-shot-secret interaction + ep0 churn (code-derived; STOP-condition forbids spraying tokens).
store/pbsdr.go:10-17SaveHostPBSSecret= last-write-wins, resetsconsumed_at. It is only reached after a successfulProvision/Reissue.tenantsync/client.go:93-95,151-155Provisionreturns typedErrTokenExistsfor an existing token — it mints nothing and never reachesSaveHostPBSSecret. So re-running the provision atom does NOT churn ep0 and does NOT stomp — but (perweb/pbsdr.go:206-222) for a provisioned-and-not-acked-deleted host it returns a loud "use Re-issue" error and does not fix the stuck host. Re-running the atom is therefore the wrong re-trigger for this root cause.tenantsync/client.go:97-100Reissue= delete+recreate the token (one rotation per call; no orphan accumulation). A blind-timerReissuewould rotate the token every tick → the descriptor hash (token_id/fingerprint) changes every tick → the agent re-consumes + re-reconciles every tick = generation + apply thrash. This is the real hazard to avoid.- Least-churn primary: re-stage the stored secret (
consumed_at → NULL). Zero ep0 interaction, no hash change; proven in SQ-2b′. Valid whenever the stored token is still good on ep0 (the common re-install/rollback case — the token was never deleted).
SQ-3c — the consumed-failed dead-end (code-derived from manager.go).
- Same-hash re-serve stays put: with
consumed-failed.jsonpresent andcf.Hash==h, the idempotent short-circuit is skipped (manager.go:235), and whenConsumePBSTokenreturnsErrNoPBSSecretthecf.Hash==hbranch (:275-281) re-asserts the LOUDconsumed_failedstate — never a silent burned-secret retry. Matches the package law comment (:17-19). - Re-issue breaks out because it does two things at once: stages a fresh secret and changes
the descriptor (new
token_id/fingerprint→ new hashh'). Withh'≠cf.Hash, the dead-end gate no longer matches → the agent re-attempts, consume now succeeds,finishConvergedruns andos.Remove(consumedFailedPath())clears the dead-end (manager.go:348-362). - Note (design input): re-staging the stored secret also recovers
consumed_failedwhen the stored token is still valid — becauseConsumePBSTokenthen succeeds (notErrNoPBSSecret), so the dead-end branch is never entered and the apply proceeds. Only a genuinely stale/rotated ep0 token forces theReissueescalation.
Verdict: a safe re-trigger = re-assert-secret, gated on the agent's reported stuck state, never
blind-timer-reissue. Re-staging the stored secret needs no generation bump (SQ-2b′ converged on
the agent's own ticker); a generation bump must be reserved for an actual descriptor content
change (Reissue / storage-id edit) — reuse the existing readPBSDR+compare no-op discipline
(web/pbsdr.go:148-164).
SQ-4 — where the re-trigger lives + the condition
Recommended design (spike-proven, one paragraph): add a hub periodic reconciler (mirroring
wgsync/reconciler.go) that, for each host with dr_tier=1 and an enabled pbs_dr descriptor,
reads the latest host-report pbs_dr.state; when that state is waiting_secret or consumed_failed
sustained across ≥1 report cycle (debounce, so a transient verify_failed tunnel flap is not
acted on), it re-stages the stored one-time secret (consumed_at → NULL, no generation bump)
and records an event. If the host is still stuck after N recon_cycles (stored token stale →
consumed_failed persists), it escalates once to Reissue (fresh token + descriptor bump). The
capability/host-report path is the only trigger that can observe "DR-ON but agent reports pbsdr
inactive" — the WG hook is structurally unable to help (same pubkey → changed==false), and
config-save is operator-manual. The one spike-proven reason this is correct: convergence requires
only a consumable secret, and staging the existing stored secret is a zero-ep0-churn, no-hash-change,
no-gen-bump operation that the agent picks up on its own 60 s tick (SQ-2b′) and that no-ops cleanly
once converged (SQ-3a) — so the reconciler self-limits to exactly the stuck hosts and stops the moment
they converge.
Generation-bump discipline (exact condition): bump only on a real descriptor content
change (Reissue's new token_id/fingerprint, or a storage_id edit) — never for a secret re-stage.
This is the existing applyPBSDR precedent: readPBSDR + field compare, no-op-if-equal
(web/pbsdr.go:148-164); the agent re-fetches on generation change, so a spurious bump every tick is
an agent-refetch loop.
Design inputs for the PBS-DR self-heal TASK
| Dimension | Spike-proven answer |
|---|---|
| Root cause | Durable hub descriptor + durable consumed one-time secret + stable host_id, after the agent loses its converged marker (re-install / rollback / disk loss). No path re-mints/re-stages the secret; WG cascade can't re-fire (changed==false); applyPBSDR "already provisioned" no-ops. Agent stuck in waiting_secret. |
| Safe re-trigger shape | Re-serve a consumable secret, gated on the agent's reported waiting_secret/consumed_failed (debounced ≥1 cycle). Primary: re-stage stored secret (consumed_at→NULL, no ep0 call, no gen bump). Escalation (stored token stale): Reissue once. Never re-run the provision atom (refuses, doesn't fix) and never blind-timer Reissue (hash/gen thrash). |
| Trigger location | Hub periodic reconciler reading host-report pbs_dr.state (like wgsync/reconciler.go). Not the WG hook (structurally can't re-fire); not config-save (manual). |
| Gen-bump condition | Only on a real descriptor content change (Reissue/storage-id edit) via readPBSDR+compare no-op discipline (web/pbsdr.go:148-164). A secret re-stage bumps nothing. |
| Orphan-token cleanup | None needed for the primary path (re-stage = no ep0 interaction) or repeat Provision (ErrTokenExists, mints nothing). Reissue is delete+recreate (no accumulation). The only orphan risk is a blind-timer Reissue design — explicitly excluded above. |
Mutations made + reverts (all scoped host_id='demo-vm-felhom-2f4b00'; SELECT-verified 1 row before each)
| ID | Mutation | Revert |
|---|---|---|
| MUT-1 | UPDATE hosts SET desired_generation+1 (3→4) — SQ-2a′ |
none needed (monotonic counter; agent compares, absolute value irrelevant) |
| MUT-2 | UPDATE host_pbs_secrets SET consumed_at=NULL WHERE … AND consumed_at IS NOT NULL — SQ-2b′ re-stage of the existing token secret (no ep0 churn) |
self-reverted: agent re-consumed at 15:44:51Z, consumed_at re-set by ConsumeHostPBSSecret |
| MUT-3 | UPDATE hosts SET desired_generation+1 (4→5) — SQ-3a |
none needed (monotonic) |
Net DB delta vs spike start: host_pbs_secrets.consumed_at 2026-07-12T20:05:53Z →
2026-07-15T15:44:51Z (secret re-consumed by real convergence); hosts.desired_generation 3 → 5.
No orphaned/unconsumed secrets, no ep0 token churn (existing token reused throughout). No non-drill
host touched (every statement WHERE host_id='demo-vm-felhom-2f4b00'; the demo host
demo-felhom-01 was never in a WHERE clause). Throwaway hubdbpoke binary removed from the hub pod.
Guest end-state
LEFT CONVERGED (not rolled back). PBS-DR on the drill guest is now genuinely state=applied
(felhom-pbs storage entry created, marker f1059b54…) — a real side-benefit: the P-DAY0-DEEP
PBS-DR leg is now GREEN on the drill (escrow ceremony + offsite arming remain separate follow-ups).
Rollback to post_day0_golden136 is available if a pristine "no DR chain" fixture is wanted again.
STOP conditions honored
Every WHERE clause proven to hit only the drill host (SELECT-before-UPDATE, count(*)=1). No code
fix attempted (that is the spec). No ep0 token churn — the primary proof deliberately reused the
existing token; Reissue was code-validated, not fired. No non-drill/production host touched.