Files
felhom.eu/documentation/audits/SPIKE-offsite-credential-recovery-2026-08-04.md
T
admin 3f2b7bc023
gates / gates (push) Successful in 7s
RECON: trace the offsite DR chain link by link — it does not join up (R-198..R-201)
Read-only recon of the escrow -> recovery chain, from a dead node to an open
repository. No production code, no build, no version bump.

Headline: the hub's superseded-escrow retention does NOT retain the offsite
repository password. host_escrow_superseded has no identity_blob column and
demoteCurrentEscrowTx copies only the K-escrow blob, so what survives a
supersession is the PBS datastore key, not the restic repo password. The next
escrow ceremony -- which the system tells a rebuilt box's customer to run --
destroys the last copy. Both demo boxes crossed that line on 2026-08-04.

Also established:
- the hub's blob-serving endpoints (re-enroll / restore-directive) have zero
  callers anywhere: agent, hub UI, scripts, runbooks (R-199)
- POST /backup/offbox/inject-password is routed and handled but no template
  contains the form (R-200)
- nothing in the recovery path has ever been exercised; the one live
  round-trip proof (2026-06-10) predates the ResticRepoPassword field (R-201)
- a fail-closed mint refusal IS implementable: the report ACK already carries
  escrow{identity_blob_present, restic_pw_sha256} and the controller discards
  it whenever no offbox target exists

Corrections: yesterday's spike annotated (candidate (b) overturned in part --
unattended recovery is impossible, customer-present is not); capability-map
retention claim struck through and replaced with what the code does.

Deliverable: documentation/audits/RECON-offsite-dr-chain-2026-08-04.md
Register: new R-198..R-201; R-193 and R-192 updated; STATUS.md refreshed.
2026-08-04 12:16:04 +02:00

34 KiB

SPIKE — what is actually one-shot, and what a guest rebuild really costs

Date: 2026-08-04 · Items: R-193 (rebuild drops offsite), R-192 (the alert says the opposite of what it measured) · Class: spike — no production code shipped for R-193 or R-192 Baselines read: felhom.eu @ f456835bbcc4 (hub v0.91.1) · felhom-controller @ 0887fd67 · felhom-agent @ 856a127

This document ends in ranked options and a STOP. It establishes facts; it does not choose. The operator's question is stated, unanswered, at the end.


ANNOTATION 2026-08-04 (later the same day) — read this before acting on §Q8

A follow-up recon traced the escrow→recovery chain link by link (audits/RECON-offsite-dr-chain-2026-08-04.md). Two corrections to this document, and they pull in opposite directions.

(1) Q8 candidate (b) — "not implementable as specified" — is TOO STRONG and is overturned in part. Zero-knowledge means the hub cannot open the blob; the customer can, with R, which is what the ceremony exists for. The mechanism to hand a rebuilt box its old repo password back — customer-present, not unattended — is a real design, and the operator has since ruled on its shape. What (b) got right is narrower than what it said: an UNATTENDED rebuild cannot recover from a zero-knowledge escrow. A customer-present one can. The recon prices that shape row by row.

(2) A LOAD-BEARING FACT THIS SPIKE DID NOT CHECK, and it makes (d)-alone weaker than ranked here. host_escrow_superseded has no identity_blob column, and demoteCurrentEscrowTx (hub/internal/store/store.go:2547-2556) copies only the K-escrow blob. So the retained "superseded" blob holds the PBS datastore key, not the restic repository password. The retention feature does not retain the secret it was built to retain, and the next escrow ceremony — which the system tells the customer to run whenever a rebuild leaves the escrow pending — is what destroys the last copy. Both demo boxes crossed that line on 2026-08-04 at 07:15:36 and 07:20:08. → R-198.

Consequence for this document's recommendation. "Ship (d), then decide (c)" still holds, but (d)'s implicit promise — the old history stays recoverable with the recovery code — is false as the code stands, and is repeated verbatim to the customer in Hungarian on the orphan card. (d) is not complete without R-198.

Unchanged and still correct: Q1, Q2, Q4 (R-196), Q5, Q6, Q7 (R-192) and the Q8 (a)/(c)/(d) analyses. Q3 remains UNMEASURED — neither box has run since; the 2026-08-05 ~02:15 UTC run is still the decisive one.


0. Why this exists

Two candidate fixes were named when R-193 was filed, and neither could be weighed, because two prior session reports described the same mechanism in contradictory terms:

  • R-39's record: Re-issue "rotates the restic password and makes the escrow stale".
  • R-193's own resolution note: the one-time password is "only the transport credential used once to install the box's own SSH key".

Those two worlds imply very different fates for the fifteen existing snapshots. This spike settles it from source first, live state second, so a measurement can contradict the code rather than be interpreted by it.

The headline, before the detail: both prior claims are wrong in the way that matters. Re-issue does not touch the restic repository password — but a guest rebuild does, silently, on every box, with nothing marking the escrow stale and nothing telling anyone. Both demo boxes lost repository continuity on 2026-08-03/04. The one that "recovered by luck" lost it too.


1. The instruments, and what each one can and cannot answer

Stated up front so a wrong instrument is visible here rather than inside a conclusion (§8 rule 2).

# Instrument Exact query / read Known limit
I1 hub SQLite snapshot kubectl exec deploy/hub -- cat /data/{hub.db,hub.db-wal,hub.db-shm} → local sqlite3 The -wal is load-bearing: hub.db alone was 2 h 15 m stale at copy time (mtime 08:34 vs wal 10:45 CEST). Copying only the main file would have produced confident answers from a stale file — the exact R-3 shape.
I1-check freshness proof PRAGMA integrity_checkok; MAX(host_reports.received_at) = 2026-08-04 08:45:38, datetime('now') = 08:49:46 Positive observable: newest row 4 min old. Not "the query returned no error".
I2 host_escrow.restic_pw_sha256 + host_escrow_superseded.restic_pw_sha256 see §2 sha256 of a 256-bit random secret — non-reversible, safe to compare. Answers "did the repo password change" without touching a box or a credential.
I3 reports.report_json → $.offsite grouped distinct-object query with first/last seen The report echoes the box's own target state; absence of the key = no target configured.
I4 one_time_secrets SELECT customer_id, created_at, consumed_at customer_id is the PRIMARY KEY — one row per customer, last-write-wins. It cannot answer "how many credentials were delivered"; a Re-issue destroys the prior row. demo-hp's 2026-07-23 consume is no longer in this table.
I5 source hub/internal/offsite/*, hub/internal/api/*, controller/internal/{backup,offsiteapply}, agent/internal/escrow

Not used, deliberately: no ReissueCredentials, no ceremony, no config change, no deletion, no write to any box. Nothing on either box or on the storage endpoint was altered by this session.


Q2 — Does a rebuilt controller reuse the escrowed repository password, or mint a new one?

RULING: it MINTS A NEW ONE. The escrowed password is never consulted on any automatic path.

Method — source. The apply-bridge's terminal step is Enabler.ConfigureOffbox(...) (controller/internal/offsiteapply/offsiteapply.go:221,256), which reaches Manager.ApplyOffsiteTargetWriteOffboxSecrets (controller/internal/backup/offbox.go:370). Quoted verbatim, offbox.go:392-401:

// Auto-generate the repo password once (0600), never log it.
if _, err := os.Stat(m.offboxPwPath()); os.IsNotExist(err) {
    pw, gerr := generateOffboxPassword()
    ...
    if werr := os.WriteFile(m.offboxPwPath(), []byte(pw), 0o600); werr != nil {

offboxPwPath() is <DataDir>/offbox/repo_password (offbox.go:363). A rebuilt guest has a fresh data dir, so the file is absent, so a fresh 256-bit password is minted.

The only path that recovers the old one is InjectOffboxPassword (offbox.go:541), whose sole caller in the entire repo is offboxInjectPasswordHandler (controller/internal/web/offbox_handlers.go:189) — a web form where a human pastes the recovered password. Verified by census: grep -rn "InjectOffboxPassword" --include=*.go . → three hits, one definition, one caller, one comment. The apply-bridge never calls it.

Method — live, and it is decisive without touching a box (I2). The hub already stores the sha256 of the escrowed repo password, so the question is answerable as a hash comparison:

host superseded (pre-rebuild) current (post-rebuild) verdict
demo-hp-bb76ea 8e03eddf9ff7… created 2026-07-23T10:01:17Z 8a9e33aa4da6… created 2026-08-04T07:15:36Z CHANGED
demo-felhom-8363b5 48741892f0ef… created 2026-07-21T08:38:33Z c60c8bc737a6… created 2026-08-04T07:20:07Z CHANGED

Both boxes minted a new repository password. demo-hp's pre-rebuild escrow (8e03…, sealed eight minutes after its 2026-07-23 09:53:41 apply) is the key to its 15 snapshots. That key now exists only inside a superseded, R-wrapped escrow blob that the hub cannot open.

This is bigger than R-193 as filed, and it hits the box R-193 called lucky.

R-193 records demo-felhom as having "survived the SAME rebuild by luck" — 76 seconds of downtime. Measured (I3), demo-felhom's pre-rebuild offsite object was:

last_run 2026-08-03T02:17:44Z · last_status ok · snapshot_count 36 · repo_size_bytes 1 136 685 919

36 snapshots, 1.14 GB. Since 2026-08-03 07:19:10 UTC it has reported snapshot_count: 0, repo_size_bytes: 0 in every one of 109 reports, and its repo password hash changed.

The luck recovered DELIVERY. It did not recover the REPOSITORY. The staged secret restored the transport in 76 seconds and the box then minted a brand-new repo password anyway. The contrast R-193 draws between the two boxes is real for the credential, and false for the data: both boxes lost repository continuity, one loudly and one silently. The silent one is worse.


Q3 — What did the 04:15 run report?

RULING: UNMEASURED — the decisive run has not happened yet. And the binary the question offers is the wrong one; source says the answer will be NEITHER 15 nor 1.

Why it has not happened. The scheduled off-box run fires at ~02:15 UTC (04:15 CEST). Measured from I3 — last_run values 2026-08-02T02:15:52Z, 2026-08-03T02:16:39Z on demo-hp; 2026-08-01T02:17:54Z, 2026-08-02T02:17:01Z, 2026-08-03T02:17:44Z on demo-felhom. Daily, both boxes. Then:

box why 2026-08-04 02:15 UTC produced nothing
demo-hp no offsite target existed at all — the offsite key was ABSENT from every report between 2026-08-03 06:12:19 and 2026-08-04 07:12:01 UTC (111 reports). The re-issue landed at 07:15:47, five hours after the window.
demo-felhom the target existed but escrow_state: pending from 2026-08-03 07:19:10 to 2026-08-04 07:20:13. OffboxRunnable requires configured AND escrowed (offbox.go:564-569), so the fork-4 gate blocked the run. The ceremony landed at 07:20:28.

Both boxes became runnable only this morning. The decisive run is 2026-08-05 ~02:15 UTC on both. Reported as unmeasured rather than inferred.

What source says will happen — and it is a third outcome

Same sub-account, same repo path, new password. Measured coordinates (I3, $.dr_recipe): u629488-sub3.your-storagebox.de:/home/felhom-repounchanged across the whole incident, and repoPath is a compile-time constant (hub/internal/offsite/offsite.go:85). So the next run opens an existing repo with the wrong key, which the codebase already has a name for (controller/internal/backup/offbox.go:67-93):

// ErrOffboxOrphaned is the sentinel returned when the offsite repo exists but is keyed under a
// passphrase this controller no longer has (the reinstall shape) ...
case strings.Contains(s, "wrong password or no key found"):
    return "orphaned"

and ensureOffboxRepo (offbox.go:663-679) branches on claim state:

if !m.settings.GetClaimed() {  // UNCLAIMED → auto move-aside + re-init
...
m.markOrphaned()
return ErrOffboxOrphaned      // CLAIMED → skip the run, show the orphan card

Both boxes report claimed: 1 (I3, $.claimed, latest report each). So the predicted outcome is: the run REFUSES with ErrOffboxOrphaned, produces no snapshot at all, and waits for the customer to confirm a reset. Not 15 (reattached) and not 1 (silently fresh) — a third outcome the question's framing did not contain. That is the good news half: the system is not going to quietly start a new history over the old one. It will stop and say so.

Consequence if the customer confirms the reset: resetOrphanedRepo (offbox.go:270-314) moves the old repo aside and re-inits — "move-aside, not deleted". The 40.9 MB + 1.14 GB of old ciphertext survives, unreadable without the superseded escrow, and keeps consuming the 50 GB soft quota indefinitely. Nothing prunes a moved-aside restic repo (S-24 covers ep0's PBS namespaces, a different tier).

To measure it tomorrow: re-read $.offsite.last_status / last_error on both boxes after 02:20 UTC, and $.offsite.repo_state for the orphan flag. Predicted ErrOffboxOrphaned, not a snapshot count. Whoever picks this up should record which of the three actually occurred — a prediction from source is not a measurement.


Q1 — What is one-shot, and what is not?

Three secrets, and they have almost nothing in common.

① storage-provider password ② the box's SFTP key ③ restic repository password
What it is the Hetzner sub-account / box password ed25519 keypair for sftp transport the data key — restic's repo encryption
Generated by hubgenPassword() (offsite.go:481) controllerKeyGen.Generate() (offsiteapply.go:234) controllergenerateOffboxPassword(), 32 random bytes hex (offbox.go:406)
Stored where one_time_secrets (hub, plaintext, one row per customer) <DataDir>/offbox/ssh_key 0600 (box only) <DataDir>/offbox/repo_password 0600 (box) + inside the R-wrapped escrow blob (hub, opaque)
Delivered how GET …/offsite/consume-password, served exactly once then marked consumed (api/offsite.go:9-31) never delivered — installed onto the provider by ssh-copy-id using ① never delivered anywhere; it never leaves the box except into the escrow ceremony
Re-issuable? YESReissueCredentials resets it at the provider, any time, operator-initiated YES — regenerated on every full apply; dr_recipe.go:45 states it plainly: "the SFTP access key is regenerated at DR" NO automatic path. Recoverable only by unsealing the escrow with the customer's recovery code and pasting it into offboxInjectPasswordHandler
What a guest rebuild does to it nothing (it lives on the hub / at the provider) destroys it; a fresh one is minted and installed — harmless destroys it; a fresh one is minted — CATASTROPHIC for the existing repo
Escrowed? no no yesIdentityBundle.ResticRepoPassword

The agent's own source names the asymmetry, and it is the sentence the whole spike turns on (felhom-agent/internal/escrow/identity.go:35-39):

ResticRepoPasswordIt is the DATA key for the offsite tier — irreplaceable (unlike the SFTP access key, which is regenerable at DR).

The one-shot thing (①) is the recoverable one. The irreplaceable thing (③) is the one nothing re-stages. R-193 was filed against ①. The damage is in ③.

Instrument caveat, worth carrying (I4): one_time_secrets has customer_id as PRIMARY KEY, and SaveOneTimeSecret is last-write-wins by design (the R-39(a) guard comment in monitor/offsite_delivery.go:26-30 depends on it). So the delivery ledger holds only the newest credential per customer: demo-hp's 2026-07-23 consume, on which R-192's whole diagnosis rests, was overwritten by the 2026-08-04 Re-issue and is no longer readable from that table.


Q4 — Which claim about Re-issue is true?

RULING: ReissueCredentials touches ONLY the provider credential (①). It does NOT touch the restic repository password. R-39's record is wrong — and the hub's own comments repeat the same wrong claim in three places.

Method — source, quoted rather than summarised. ReissueCredentials (hub/internal/offsite/offsite.go:150-228) does exactly four things:

  1. genPassword();
  2. ResetSubaccountPassword / ResetBoxPassword + WaitActionthe provider account password;
  3. Store.SaveOneTimeSecret(customerID, pw) — stage ① for delivery;
  4. bookkeeping: MarkEscrowStale + offsite_reissued + escrow_stale events.

There is no reference to a restic password anywhere in the function, and none is possible — the repo password is generated on the box and never leaves it except into the escrow blob. The hub does not hold it in any openable form (api/handler.go: "The hub stores the bytes and NEVER decrypts them (it has no recovery code)").

But step 4 is justified by a claim that is false, offsite.go:198-201:

// v0.57.0 (2.3, the escrow-honesty fix): the restic repo password just changed, so any existing // key-escrow blob — which sealed the OLD password — is now STALE.

The escrow is marked stale; the stated reason for marking it is not something this function does. The same false premise appears twice more: api/handler.go:1067-1069 ("The re-issuer resets the restic repo password, which makes the OLD escrow blob stale") and OPEN-ITEMS.md's R-193 row, which inherited it from R-39.

This is the EIGHTH entry in CLAUDE.md's table of comments asserting an invariant the code does not provide — and the first where the comment is not merely unenforced but factually describes a different function. It survived because it reads as settled and because its effect (a stale escrow) is real, so nobody checked the cause.

Consequence, and it is a live defect, not a documentation nit. On the ordinary Re-issue shape — a consumed-but-failed install on a box that still has its repo_password file — the box re-applies, WriteOffboxSecrets finds the file present and keeps it, and the repo password is unchanged. The hub has nonetheless told the customer, in Hungarian, that their recovery escrow is stale and asked them to re-run the ceremony. A false staleness alarm and an unnecessary ceremony.R-196.

Corroboration against what actually happened on demo-hp yesterday. The escrow went pending after the Re-issue and the task asks: if the repo password was untouched, what explains that? It is fully explained, and not by the Re-issue. Two independent mechanisms, both hub-side-invisible:

  1. ApplyOffsiteTarget (offbox.go:482-494) carries EscrowState over from the existing target — but the rebuilt guest had no existing target, so cur was nil, so the fallthrough if tgt.EscrowState != "escrowed" { tgt.EscrowState = "pending" } fired.
  2. The box had minted a new repo password moments earlier (Q2), so pending was correct — the new key genuinely was not escrowed.

The Re-issue's escrow_stale event fired at 07:11:51 for the wrong reason and happened to be true for a different one. Coincidence, not correctness. The proof that these are independent is demo-felhom: it received no Re-issue at all (measured — its only escrow_stale / offsite_reissued events in the entire history are dated 2026-07-21 08:29:29), its repo password changed anyway, and nothing marked its escrow stale for thirteen hours.

The mechanism is exactly inverted. escrow_stale is wired to the one path that does not change the repo password, and absent from the path that does.


Q5 — Why did one box recover itself and the other not?

RULING: confirmed from the ledger — an unconsumed staged secret, thirteen days old. And the recovery was partial in a way the narrative missed.

Method — I4, the one_time_secrets ledger, no values read:

customer created_at consumed_at
demo-felhom 2026-07-21 08:29:29 2026-08-03 07:17:58
demo-hp 2026-08-04 07:11:51 2026-08-04 07:12:06

demo-felhom's row was created 2026-07-21 and sat UNCONSUMED for 13 days — traceable to a Re-issue at that exact timestamp (its offsite_reissued + escrow_stale event pair, 2026-07-21 08:29:29). When the rebuild dropped its target on 2026-08-03, the fresh controller consumed the spare at 07:17:58 and was reporting offsite again by 07:19:10. 76 seconds, on a credential nobody had planned to leave there. Confirmed from the ledger, not the narrative.

What would have happened without one: exactly demo-hp — consume404 (api/offsite.go:20-22, no unconsumed offsite password), the apply-bridge returns offsite-apply: consume one-time password: … and retries forever, and the tier stays absent until an operator acts. demo-hp sat there 25 hours (2026-08-03 06:12:192026-08-04 07:12:01).

The correction to the contrast (see Q2): demo-felhom's spare restored delivery only. Its 36 snapshots / 1.14 GB were orphaned by the same rebuild. Neither box's data survived; one box's plumbing did.

Not established: demo-felhom's escrow generation before 2026-07-21. host_escrow_superseded holds only two rows in the whole database (ids 3 and 4, one per box, both superseded 2026-08-04), so whether the 2026-07-21 Re-issue's ceremony re-sealed an unchanged password — the predicted false-staleness shape — cannot be shown from this data. Stated as unknown rather than asserted.


Q6 — Why is delivery one-shot at all?

RULING: the design's reason is that ① is a LIVE PROVIDER PASSWORD, not a bootstrap token; and the honest answer to the security question is that an automatic restage weakens EXPOSURE WINDOW, not authentication.

The design's own reasons, from source:

  • api/offsite.go:9-12"serves the one-time transient offsite password to the controller EXACTLY ONCE … The value is returned once then marked consumed — a second call 404s. NEVER logged."
  • offsite.go:145-149"the EXPLICIT operator recovery for a consumed-password dead-end … It is NOT implicit rotation: ProvisionOffsite never calls this."
  • offsiteapply.go:1-6 — the apply-bridge is "idempotent (a descriptor hash marker prevents re-consuming a spent password)", i.e. the design treats a wasted consume as a real hazard, which R-71a's whole settle-gate exists to prevent.

The security half, stated rather than skipped. The question is: could a machine that is not the real one obtain a credential via an automatic restage?

What stands there today is authentication, and a restage does not remove it. The consume endpoint is gated by checkAuthCustomer (api/offsite.go:14-18) — the customer's API key, the same credential as the config pull. Anything that could consume a restaged password can already pull the customer's whole config. So an automatic restage grants no new access to an unauthenticated party.

What it genuinely weakens is the exposure window. Today a live Hetzner sub-account password is fetchable only in the seconds after an operator deliberately staged one. Under an automatic restage it becomes fetchable whenever the box reports no offsite target — a condition an attacker holding the API key can arguably induce. The credential's blast radius at the provider is the customer's own sub-account (label-scoped; ReissueCredentials refuses unless the label lookup finds exactly one, offsite.go:164,180), so this is a widening of when, not of what.

And one further guard must survive any change — R-39(a), monitor/offsite_delivery.go:26-30: SaveOneTimeSecret is last-write-wins, so restaging on top of an unconsumed secret clobbers a password a box may be about to consume. Any automatic restage must re-read the row immediately before acting and refuse unless it is still a consumed row. That guard is correct and is not what R-192 is about.


Q7 — R-192's guard: exactly what it reads, and why it read wrong

RULING: both halves confirmed, with numbers. The guard reads the OLDEST 500 reports since the consume, and for demo-hp all 500 predated the rebuild by six days.

The query, quoted (hub/internal/store/store.go:987):

SELECT report_json FROM reports WHERE customer_id = ? AND received_at > ? ORDER BY id LIMIT 500

ORDER BY id ascending with LIMIT 500 = the oldest 500, not the newest.

Reproduced against the live data (I1), with demo-hp's real consume anchor 2026-07-23 09:53:41 (recovered from R-192's record — it is no longer in one_time_secrets, see I4):

measurement value
total (what the guard sees) 500
withOffsite (what the guard sees) 500
oldest report in that window 2026-07-23 09:53:47
newest report in that window 2026-07-28 11:17:40
true total over the same window, no LIMIT 1174
true withOffsite over the same window, no LIMIT 1063 (⇒ 111 without — matching the 111 ABSENT reports in I3)

The whole 500-report evidence set ends 2026-07-28 — six days before the 2026-08-03 rebuild.

Defect (a) — the message is exactly inverted. maybeEmitStuck (monitor/offsite_delivery.go:~110) interpolates status.ReportsSinceConsume into a string that hardcodes the phrase "carry no offbox target", and never reads OffsiteReportsSinceConsume. The live e-mail, quoted in full from the events table:

"Offsite delivery stuck: one-time password consumed 284h19m0s ago and 500 report(s) since carry no offbox target — the credential is likely burned (apply died between consume and persist). Re-issue delivers a fresh one."

OffsiteReportsSinceConsume for that same call was 500. Every one of them carried a target. The message states the precise negation of its own measurement, and prescribes a remedy for a failure mode that did not occur.

Defect (b) — the heal refuses silently, on stale evidence. maybeHeal returns bare on status.OffsiteReportsSinceConsume != 0 (it was 500), with no log line. offsite_credential_restaged has never fired for any customer — confirmed: zero rows of that event type in the database.

What the guard SHOULD discriminate on — the shape, not the code. The guard's intent is right: "there is offbox evidence, so this is a regressed apply, not a burned credential — the operator's call." Its error is that it asks the question of the wrong time period. Offbox evidence from before a rebuild is not evidence that the credential still works. The discriminator must be recency-bounded and rebuild-aware: judge on evidence after the newest controller_started / config-hash change, or on the latest N reports — never on "everything since the consume". And a refusal must name its reason in the log, because "we chose not to heal" and "the heal never ran" must not look identical (the S-16 rule, one layer down).

Do not fix it in isolation. Under Q2's finding, a successful auto-restage on demo-hp would have restored the transport and the box would still have minted a new repo password. The heal cannot protect the data; it can only protect the plumbing. Whatever shape it takes must say so.


Q8 — What each candidate would cost

A fourth option exists and is the important one, because (a) and (b) both address ① and Q2 shows the damage is in ③.

(a) The hub restages automatically when a re-enrolled box reports no offsite

A version of this ALREADY EXISTS and is wired to the wrong event. reissueOnReenroll (hub/internal/api/handler.go:1051-1084), leg F3:

"F3 — offsite continuity: re-stage the one-time offsite password to the fresh controller (the one-time password only ever reached the OLD controller)."

It is called from handleHostEnroll — but after a mint-once-reuse short-circuit (handler.go:1004-1016): if existing != nil { …return… }. A guest rebuild leaves the hosts row intact (measured: demo-hp-bb76ea is the same host_id throughout the incident; only the agent leaf re-keyed, host_leaf_changed 2026-08-03 06:09:40), so re-enroll returns the existing credential and F3 is never reached.

  • Changes: hub only. Either lower the trigger from host-enrollment to a controller-level rebuild signal (the config_hash change + controller_started pair the hub already receives and already logs), or fix R-192's guard so the existing R-71c self-heal can act.
  • Weakens: the exposure window in Q6. Needs the R-39(a) unconsumed-secret guard kept intact.
  • Does NOT solve: ③. The box still mints a new repo password and still orphans the repo. This option restores the plumbing to a box whose data key is already gone.
  • Cost: small. Value against the actual harm: near zero.

(b) The credential becomes recoverable from escrow at re-bootstrap

ANNOTATED 2026-08-04 (recon) — this ruling is overturned in part. See the annotation at the top. The constraint is real but narrower than written: an unattended rebuild cannot recover from a zero-knowledge escrow; a customer-present one can, and the operator has ruled on that shape. The recon also found that the chain this option would use is clientless at three links and lossy at a fourth (RECON-offsite-dr-chain-2026-08-04.md §3–§4).

  • BLOCKED AS STATED, and this is a hard constraint, not an estimate. The escrow blob is R-wrapped and zero-knowledge — the hub holds opaque bytes and has no recovery code (api/handler.go, and D6 in CONTEXT.md S-3). A rebuilt box cannot unseal it without the customer's recovery code. Any design that says "recover it at re-bootstrap" is describing a customer-present ceremony, i.e. what the existing manual offboxInjectPasswordHandler already is.
  • Does NOT solve: an unattended rebuild, which is the whole scenario.
  • Cost: not implementable as specified.

(c) — NOT PREVIOUSLY NAMED — the agent retains the repo password across a guest rebuild

The agent lives on the Proxmox host and survives a guest rebuild. It already receives the repo password today: the controller pushes it over the pinned local API (POST /escrow/stage-secretagent/internal/localapi/escrow_stage.go), and the agent writes it to the fixed 0600 path /var/lib/felhom-agent/escrow-stage/restic_repo_password (agent/internal/escrow/identity.go:44-46) transiently — wiped by the ceremony (WipeStagedResticPassword). The seam, the transport, the pinning and the file are all already built.

  • Changes: agent (retain rather than wipe; serve back over the same pinned local API) + controller (on a fresh data dir, ask the agent before minting — the InjectOffboxPassword seam already exists and is exactly the right shape).
  • Weakens: the escrow's zero-knowledge posture in one specific way — the irreplaceable data key would now sit at rest on the Proxmox host, which the operator has root on. Measured against D6, that is not a new exposure in principle (D6 already states plainly that "the operator cannot read customer data" was never the security property, and the operator holds root on every box) — but it IS a new copy, and copies are the thing R-133 is already open about for the hub DB. This is a real trade and it is the operator's to make.
  • Solves: exactly the failure that occurred, unattended, on every rebuild, without any hub involvement and without any credential leaving the premises.
  • Cost: medium. Two repos, one existing seam each.

(d) Do neither — keep the manual Re-issue, and fix what LIES

  • Changes: R-192's message + refusal log; correct the three false comments about what Re-issue rotates; and — the part that actually matters — make a repo-password change VISIBLE. The hub already receives restic_pw_sha256 on every escrow upload and already stores it. Comparing the new hash against the superseded one is a two-line verdict, and it is the signal that was missing on demo-felhom for thirteen hours → R-197. (Note the shape: this spike answered its own hardest question with a comparison the hub could be making automatically and is not. Both values were already in the database; nothing read them.)
  • Weakens: nothing.
  • Does NOT solve: the data loss. Every rebuild still orphans the repo; the operator simply finds out the same day instead of never.
  • Cost: small.

Ranked, with a recommendation

rank option why
1 (d) now — the honesty pass It is cheap, it weakens nothing, and it is the only option that helps the incident already in flight. The alarm that lies is worse than no alarm; the repo-password-changed detector is the one signal whose absence let demo-felhom lose 1.14 GB of history silently. Ship this regardless of what is decided about the rest.
2 (c) — agent-retained repo password The only candidate that addresses ③, which is where the harm is. Every seam it needs exists. It costs one deliberate trade-off the operator must actually make.
3 (a) — auto-restage Worth doing after (c), never instead of it. On its own it restores the transport to a box whose data key is already gone, and would have made both boxes look healthy on 2026-08-04 while the snapshots were orphaned — a strictly worse outcome than the current loud failure.
4 (b) — recover from escrow at re-bootstrap Not implementable as stated; the escrow is zero-knowledge by design.

Recommendation: ship (d), then decide (c). And do not ship (a) first — it would have hidden this.


What could not be established

Named rather than guessed (§8 rule 4).

  1. The 2026-08-05 02:15 UTC run's actual outcome (Q3). Predicted ErrOffboxOrphaned from source + claim state; not measured.
  2. Whether the orphaned ciphertext still exists at the provider. 40.9 MB (demo-hp) + 1.14 GB (demo-felhom) are presumed present at /home/felhom-repo — reading the endpoint needs the SFTP credential, which this session did not touch. The Hetzner API has no directory-listing surface.
  3. Whether demo-felhom's 2026-07-21 Re-issue re-sealed an unchanged password (the predicted false-staleness shape, Q4/Q5). Only two host_escrow_superseded rows exist in the whole database; the generation before 2026-07-21 is not retained.
  4. What removed demo-hp's offsite target at the rebuild is fully explained (fresh data dir), but why the guest was rebuilt on both boxes on consecutive mornings is outside this spike.

The operator's question — stated, and NOT answered here

The irreplaceable secret in the offsite tier is the restic repository password. It is generated on the box, it is destroyed by a guest rebuild, and the only copy that survives is sealed under the customer's recovery code — which nothing but a human can open. Every machine is going to be reinstalled.

Do you want that key to survive a rebuild unattended — by keeping a copy on the Proxmox host, where you already have root — or do you want it to stay sealed under the customer's recovery code and accept that every rebuild starts a new repository and orphans the old one, provided the system says so loudly on the day it happens?

Option (c) is the first. Option (d) alone is the second. This spike does not choose.