# SPIKE — what is actually one-shot, and what a guest rebuild really costs **Date:** 2026-08-04 · **Items:** R-193 (rebuild drops offsite), R-192 (the alert says the opposite of what it measured) · **Class:** spike — **no production code shipped for R-193 or R-192** **Baselines read:** `felhom.eu` @ `f456835bbcc4` (hub v0.91.1) · `felhom-controller` @ `0887fd67` · `felhom-agent` @ `856a127` > **This document ends in ranked options and a STOP.** It establishes facts; it does not choose. > The operator's question is stated, unanswered, at the end. --- > ### ANNOTATION 2026-08-04 (later the same day) — read this before acting on §Q8 > > A follow-up recon traced the escrow→recovery chain link by link > (`audits/RECON-offsite-dr-chain-2026-08-04.md`). **Two corrections to this document, and they pull > in opposite directions.** > > **(1) Q8 candidate (b) — "not implementable as specified" — is TOO STRONG and is overturned in > part.** Zero-knowledge means the *hub* cannot open the blob; the *customer* can, with R, which is > what the ceremony exists for. The mechanism to hand a rebuilt box its old repo password back — > customer-present, not unattended — is a real design, and the operator has since ruled on its shape. > What (b) got right is narrower than what it said: **an UNATTENDED rebuild cannot recover from a > zero-knowledge escrow.** A customer-present one can. The recon prices that shape row by row. > > **(2) A LOAD-BEARING FACT THIS SPIKE DID NOT CHECK, and it makes (d)-alone weaker than ranked > here.** `host_escrow_superseded` **has no `identity_blob` column**, and `demoteCurrentEscrowTx` > (`hub/internal/store/store.go:2547-2556`) copies only the K-escrow `blob`. So the retained > "superseded" blob holds the **PBS datastore key**, not the **restic repository password**. The > retention feature does not retain the secret it was built to retain, and the **next escrow ceremony > — which the system tells the customer to run whenever a rebuild leaves the escrow `pending` — is > what destroys the last copy.** Both demo boxes crossed that line on 2026-08-04 at 07:15:36 and > 07:20:08. → **R-198**. > > **Consequence for this document's recommendation.** "Ship (d), then decide (c)" still holds, but > (d)'s implicit promise — *the old history stays recoverable with the recovery code* — is **false as > the code stands**, and is repeated verbatim to the customer in Hungarian on the orphan card. (d) is > not complete without R-198. > > **Unchanged and still correct:** Q1, Q2, Q4 (R-196), Q5, Q6, Q7 (R-192) and the Q8 (a)/(c)/(d) > analyses. **Q3 remains UNMEASURED** — neither box has run since; the 2026-08-05 ~02:15 UTC run is > still the decisive one. --- ## 0. Why this exists Two candidate fixes were named when R-193 was filed, and neither could be weighed, because two prior session reports described the same mechanism in contradictory terms: - R-39's record: Re-issue *"rotates the restic password and makes the escrow stale"*. - R-193's own resolution note: the one-time password is *"only the transport credential used once to install the box's own SSH key"*. Those two worlds imply very different fates for the fifteen existing snapshots. This spike settles it **from source first, live state second**, so a measurement can contradict the code rather than be interpreted by it. **The headline, before the detail:** *both prior claims are wrong in the way that matters.* Re-issue does **not** touch the restic repository password — but a **guest rebuild does**, silently, on every box, with nothing marking the escrow stale and nothing telling anyone. Both demo boxes lost repository continuity on 2026-08-03/04. The one that "recovered by luck" lost it too. --- ## 1. The instruments, and what each one can and cannot answer Stated up front so a wrong instrument is visible here rather than inside a conclusion (§8 rule 2). | # | Instrument | Exact query / read | Known limit | |---|---|---|---| | I1 | hub SQLite snapshot | `kubectl exec deploy/hub -- cat /data/{hub.db,hub.db-wal,hub.db-shm}` → local `sqlite3` | **The `-wal` is load-bearing**: `hub.db` alone was 2 h 15 m stale at copy time (mtime 08:34 vs wal 10:45 CEST). Copying only the main file would have produced confident answers from a stale file — the exact R-3 shape. | | I1-check | freshness proof | `PRAGMA integrity_check` → `ok`; `MAX(host_reports.received_at)` = `2026-08-04 08:45:38`, `datetime('now')` = `08:49:46` | **Positive observable**: newest row 4 min old. Not "the query returned no error". | | I2 | `host_escrow.restic_pw_sha256` + `host_escrow_superseded.restic_pw_sha256` | see §2 | sha256 of a 256-bit random secret — non-reversible, safe to compare. Answers *"did the repo password change"* **without touching a box or a credential**. | | I3 | `reports.report_json → $.offsite` | grouped distinct-object query with first/last seen | The report echoes the box's own target state; absence of the key = no target configured. | | I4 | `one_time_secrets` | `SELECT customer_id, created_at, consumed_at` | **`customer_id` is the PRIMARY KEY** — one row per customer, last-write-wins. It **cannot** answer "how many credentials were delivered"; a Re-issue destroys the prior row. demo-hp's 2026-07-23 consume is no longer in this table. | | I5 | source | `hub/internal/offsite/*`, `hub/internal/api/*`, `controller/internal/{backup,offsiteapply}`, `agent/internal/escrow` | — | **Not used, deliberately:** no `ReissueCredentials`, no ceremony, no config change, no deletion, no write to any box. Nothing on either box or on the storage endpoint was altered by this session. --- ## Q2 — Does a rebuilt controller reuse the escrowed repository password, or mint a new one? ### **RULING: it MINTS A NEW ONE. The escrowed password is never consulted on any automatic path.** **Method — source.** The apply-bridge's terminal step is `Enabler.ConfigureOffbox(...)` (`controller/internal/offsiteapply/offsiteapply.go:221,256`), which reaches `Manager.ApplyOffsiteTarget` → `WriteOffboxSecrets` (`controller/internal/backup/offbox.go:370`). Quoted verbatim, `offbox.go:392-401`: ```go // Auto-generate the repo password once (0600), never log it. if _, err := os.Stat(m.offboxPwPath()); os.IsNotExist(err) { pw, gerr := generateOffboxPassword() ... if werr := os.WriteFile(m.offboxPwPath(), []byte(pw), 0o600); werr != nil { ``` `offboxPwPath()` is `/offbox/repo_password` (`offbox.go:363`). A rebuilt guest has a fresh data dir, so the file is absent, so **a fresh 256-bit password is minted**. **The only path that recovers the old one is `InjectOffboxPassword`** (`offbox.go:541`), whose sole caller in the entire repo is `offboxInjectPasswordHandler` (`controller/internal/web/offbox_handlers.go:189`) — a **web form where a human pastes the recovered password**. Verified by census: `grep -rn "InjectOffboxPassword" --include=*.go .` → three hits, one definition, one caller, one comment. **The apply-bridge never calls it.** **Method — live, and it is decisive without touching a box (I2).** The hub already stores the sha256 of the escrowed repo password, so the question is answerable as a hash comparison: | host | superseded (pre-rebuild) | current (post-rebuild) | verdict | |---|---|---|---| | `demo-hp-bb76ea` | `8e03eddf9ff7…` created `2026-07-23T10:01:17Z` | `8a9e33aa4da6…` created `2026-08-04T07:15:36Z` | **CHANGED** | | `demo-felhom-8363b5` | `48741892f0ef…` created `2026-07-21T08:38:33Z` | `c60c8bc737a6…` created `2026-08-04T07:20:07Z` | **CHANGED** | **Both boxes minted a new repository password.** demo-hp's pre-rebuild escrow (`8e03…`, sealed eight minutes after its 2026-07-23 09:53:41 apply) is the key to its 15 snapshots. That key now exists **only inside a superseded, R-wrapped escrow blob** that the hub cannot open. ### **This is bigger than R-193 as filed, and it hits the box R-193 called lucky.** R-193 records demo-felhom as having *"survived the SAME rebuild by luck"* — 76 seconds of downtime. Measured (I3), demo-felhom's pre-rebuild offsite object was: ``` last_run 2026-08-03T02:17:44Z · last_status ok · snapshot_count 36 · repo_size_bytes 1 136 685 919 ``` **36 snapshots, 1.14 GB.** Since 2026-08-03 07:19:10 UTC it has reported `snapshot_count: 0, repo_size_bytes: 0` in every one of 109 reports, and its repo password hash changed. > **The luck recovered DELIVERY. It did not recover the REPOSITORY.** The staged secret restored the > transport in 76 seconds and the box then minted a brand-new repo password anyway. The contrast > R-193 draws between the two boxes is real for the *credential*, and **false for the data**: both > boxes lost repository continuity, one loudly and one silently. The silent one is worse. --- ## Q3 — What did the 04:15 run report? ### **RULING: UNMEASURED — the decisive run has not happened yet. And the binary the question offers is the wrong one; source says the answer will be NEITHER 15 nor 1.** **Why it has not happened.** The scheduled off-box run fires at ~02:15 UTC (04:15 CEST). Measured from I3 — `last_run` values `2026-08-02T02:15:52Z`, `2026-08-03T02:16:39Z` on demo-hp; `2026-08-01T02:17:54Z`, `2026-08-02T02:17:01Z`, `2026-08-03T02:17:44Z` on demo-felhom. **Daily, both boxes.** Then: | box | why 2026-08-04 02:15 UTC produced nothing | |---|---| | demo-hp | no offsite target existed at all — the `offsite` key was **ABSENT** from every report between `2026-08-03 06:12:19` and `2026-08-04 07:12:01` UTC (111 reports). The re-issue landed at 07:15:47, five hours after the window. | | demo-felhom | the target existed but `escrow_state: pending` from `2026-08-03 07:19:10` to `2026-08-04 07:20:13`. `OffboxRunnable` requires configured **AND** escrowed (`offbox.go:564-569`), so the fork-4 gate blocked the run. The ceremony landed at 07:20:28. | Both boxes became runnable only this morning. **The decisive run is 2026-08-05 ~02:15 UTC on both.** Reported as unmeasured rather than inferred. ### What source says will happen — and it is a third outcome Same sub-account, same repo path, new password. Measured coordinates (I3, `$.dr_recipe`): `u629488-sub3.your-storagebox.de:/home/felhom-repo` — **unchanged** across the whole incident, and `repoPath` is a compile-time constant (`hub/internal/offsite/offsite.go:85`). So the next run opens an **existing repo with the wrong key**, which the codebase already has a name for (`controller/internal/backup/offbox.go:67-93`): ```go // ErrOffboxOrphaned is the sentinel returned when the offsite repo exists but is keyed under a // passphrase this controller no longer has (the reinstall shape) ... case strings.Contains(s, "wrong password or no key found"): return "orphaned" ``` and `ensureOffboxRepo` (`offbox.go:663-679`) branches on claim state: ```go if !m.settings.GetClaimed() { // UNCLAIMED → auto move-aside + re-init ... m.markOrphaned() return ErrOffboxOrphaned // CLAIMED → skip the run, show the orphan card ``` Both boxes report `claimed: 1` (I3, `$.claimed`, latest report each). **So the predicted outcome is: the run REFUSES with `ErrOffboxOrphaned`, produces no snapshot at all, and waits for the customer to confirm a reset.** Not 15 (reattached) and not 1 (silently fresh) — a **third** outcome the question's framing did not contain. That is the good news half: the system is not going to quietly start a new history over the old one. It will stop and say so. **Consequence if the customer confirms the reset:** `resetOrphanedRepo` (`offbox.go:270-314`) moves the old repo aside and re-inits — *"move-aside, not deleted"*. The 40.9 MB + 1.14 GB of old ciphertext survives, unreadable without the superseded escrow, **and keeps consuming the 50 GB soft quota indefinitely.** Nothing prunes a moved-aside restic repo (S-24 covers ep0's PBS namespaces, a different tier). **To measure it tomorrow:** re-read `$.offsite.last_status` / `last_error` on both boxes after 02:20 UTC, and `$.offsite.repo_state` for the orphan flag. **Predicted `ErrOffboxOrphaned`, not a snapshot count.** Whoever picks this up should record which of the three actually occurred — a prediction from source is not a measurement. --- ## Q1 — What is one-shot, and what is not? Three secrets, and they have almost nothing in common. | | **① storage-provider password** | **② the box's SFTP key** | **③ restic repository password** | |---|---|---|---| | **What it is** | the Hetzner sub-account / box password | ed25519 keypair for `sftp` transport | the **data** key — restic's repo encryption | | **Generated by** | **hub** — `genPassword()` (`offsite.go:481`) | **controller** — `KeyGen.Generate()` (`offsiteapply.go:234`) | **controller** — `generateOffboxPassword()`, 32 random bytes hex (`offbox.go:406`) | | **Stored where** | `one_time_secrets` (hub, plaintext, **one row per customer**) | `/offbox/ssh_key` 0600 (box only) | `/offbox/repo_password` 0600 (box) + inside the R-wrapped escrow blob (hub, opaque) | | **Delivered how** | `GET …/offsite/consume-password`, **served exactly once** then marked consumed (`api/offsite.go:9-31`) | never delivered — installed onto the provider by ssh-copy-id using ① | never delivered anywhere; it never leaves the box except into the escrow ceremony | | **Re-issuable?** | **YES** — `ReissueCredentials` resets it at the provider, any time, operator-initiated | **YES** — regenerated on every full apply; `dr_recipe.go:45` states it plainly: *"the SFTP access key is regenerated at DR"* | **NO automatic path.** Recoverable only by unsealing the escrow with the customer's recovery code and pasting it into `offboxInjectPasswordHandler` | | **What a guest rebuild does to it** | nothing (it lives on the hub / at the provider) | destroys it; a fresh one is minted and installed — **harmless** | **destroys it; a fresh one is minted — CATASTROPHIC for the existing repo** | | **Escrowed?** | no | no | **yes** — `IdentityBundle.ResticRepoPassword` | The agent's own source names the asymmetry, and it is the sentence the whole spike turns on (`felhom-agent/internal/escrow/identity.go:35-39`): > `ResticRepoPassword` … **It is the DATA key for the offsite tier — irreplaceable (unlike the SFTP > access key, which is regenerable at DR).** **The one-shot thing (①) is the recoverable one. The irreplaceable thing (③) is the one nothing re-stages.** R-193 was filed against ①. The damage is in ③. **Instrument caveat, worth carrying (I4):** `one_time_secrets` has `customer_id` as PRIMARY KEY, and `SaveOneTimeSecret` is last-write-wins **by design** (the R-39(a) guard comment in `monitor/offsite_delivery.go:26-30` depends on it). So the delivery ledger holds only the newest credential per customer: demo-hp's 2026-07-23 consume, on which R-192's whole diagnosis rests, was **overwritten** by the 2026-08-04 Re-issue and is no longer readable from that table. --- ## Q4 — Which claim about Re-issue is true? ### **RULING: `ReissueCredentials` touches ONLY the provider credential (①). It does NOT touch the restic repository password. R-39's record is wrong — and the hub's own comments repeat the same wrong claim in three places.** **Method — source, quoted rather than summarised.** `ReissueCredentials` (`hub/internal/offsite/offsite.go:150-228`) does exactly four things: 1. `genPassword()`; 2. `ResetSubaccountPassword` / `ResetBoxPassword` + `WaitAction` — **the provider account password**; 3. `Store.SaveOneTimeSecret(customerID, pw)` — stage ① for delivery; 4. bookkeeping: `MarkEscrowStale` + `offsite_reissued` + `escrow_stale` events. There is **no reference to a restic password anywhere in the function**, and none is possible — the repo password is generated on the box and never leaves it except into the escrow blob. The hub does not hold it in any openable form (`api/handler.go`: *"The hub stores the bytes and NEVER decrypts them (it has no recovery code)"*). **But step 4 is justified by a claim that is false**, `offsite.go:198-201`: > `// v0.57.0 (2.3, the escrow-honesty fix): the restic repo password just changed, so any existing` > `// key-escrow blob — which sealed the OLD password — is now STALE.` The escrow *is* marked stale; the stated reason for marking it is not something this function does. The same false premise appears twice more: `api/handler.go:1067-1069` (*"The re-issuer resets the restic repo password, which makes the OLD escrow blob stale"*) and `OPEN-ITEMS.md`'s R-193 row, which inherited it from R-39. > **This is the EIGHTH entry in `CLAUDE.md`'s table of comments asserting an invariant the code does > not provide** — and the first where the comment is not merely unenforced but factually describes a > different function. It survived because it reads as settled and because its *effect* (a stale escrow) > is real, so nobody checked the *cause*. **Consequence, and it is a live defect, not a documentation nit.** On the ordinary Re-issue shape — a consumed-but-failed install on a box that still has its `repo_password` file — the box re-applies, `WriteOffboxSecrets` finds the file present and **keeps it**, and the repo password is unchanged. The hub has nonetheless told the customer, in Hungarian, that their recovery escrow is stale and asked them to re-run the ceremony. **A false staleness alarm and an unnecessary ceremony.** → **R-196**. **Corroboration against what actually happened on demo-hp yesterday.** The escrow went `pending` after the Re-issue and the task asks: if the repo password was untouched, what explains that? **It is fully explained, and not by the Re-issue.** Two independent mechanisms, both hub-side-invisible: 1. `ApplyOffsiteTarget` (`offbox.go:482-494`) carries `EscrowState` over from the *existing* target — but the rebuilt guest had **no existing target**, so `cur` was nil, so the fallthrough `if tgt.EscrowState != "escrowed" { tgt.EscrowState = "pending" }` fired. 2. The box had minted a **new** repo password moments earlier (Q2), so `pending` was *correct* — the new key genuinely was not escrowed. **The Re-issue's `escrow_stale` event fired at 07:11:51 for the wrong reason and happened to be true for a different one.** Coincidence, not correctness. The proof that these are independent is **demo-felhom**: it received **no Re-issue at all** (measured — its only `escrow_stale` / `offsite_reissued` events in the entire history are dated `2026-07-21 08:29:29`), its repo password changed anyway, and **nothing marked its escrow stale for thirteen hours.** > **The mechanism is exactly inverted.** `escrow_stale` is wired to the one path that does **not** > change the repo password, and absent from the path that **does**. --- ## Q5 — Why did one box recover itself and the other not? ### **RULING: confirmed from the ledger — an unconsumed staged secret, thirteen days old. And the recovery was partial in a way the narrative missed.** **Method — I4, the `one_time_secrets` ledger, no values read:** | customer | `created_at` | `consumed_at` | |---|---|---| | `demo-felhom` | `2026-07-21 08:29:29` | `2026-08-03 07:17:58` | | `demo-hp` | `2026-08-04 07:11:51` | `2026-08-04 07:12:06` | demo-felhom's row was **created 2026-07-21 and sat UNCONSUMED for 13 days** — traceable to a Re-issue at that exact timestamp (its `offsite_reissued` + `escrow_stale` event pair, `2026-07-21 08:29:29`). When the rebuild dropped its target on 2026-08-03, the fresh controller consumed the spare at 07:17:58 and was reporting `offsite` again by 07:19:10. **76 seconds, on a credential nobody had planned to leave there.** Confirmed from the ledger, not the narrative. **What would have happened without one:** exactly demo-hp — `consume` → **404** (`api/offsite.go:20-22`, `no unconsumed offsite password`), the apply-bridge returns `offsite-apply: consume one-time password: …` and retries forever, and the tier stays absent until an operator acts. demo-hp sat there **25 hours** (`2026-08-03 06:12:19` → `2026-08-04 07:12:01`). **The correction to the contrast (see Q2):** demo-felhom's spare restored *delivery* only. Its 36 snapshots / 1.14 GB were orphaned by the same rebuild. **Neither box's data survived; one box's plumbing did.** *Not established:* demo-felhom's escrow generation before 2026-07-21. `host_escrow_superseded` holds only two rows in the whole database (ids 3 and 4, one per box, both superseded 2026-08-04), so whether the 2026-07-21 Re-issue's ceremony re-sealed an *unchanged* password — the predicted false-staleness shape — **cannot be shown from this data**. Stated as unknown rather than asserted. --- ## Q6 — Why is delivery one-shot at all? ### **RULING: the design's reason is that ① is a LIVE PROVIDER PASSWORD, not a bootstrap token; and the honest answer to the security question is that an automatic restage weakens EXPOSURE WINDOW, not authentication.** **The design's own reasons, from source:** - `api/offsite.go:9-12` — *"serves the one-time transient offsite password to the controller **EXACTLY ONCE** … The value is returned once then marked consumed — a second call 404s. NEVER logged."* - `offsite.go:145-149` — *"the **EXPLICIT** operator recovery for a consumed-password dead-end … It is **NOT** implicit rotation: ProvisionOffsite never calls this."* - `offsiteapply.go:1-6` — the apply-bridge is *"idempotent (a descriptor hash marker prevents **re-consuming a spent password**)"*, i.e. the design treats a wasted consume as a real hazard, which R-71a's whole settle-gate exists to prevent. **The security half, stated rather than skipped.** The question is: *could a machine that is not the real one obtain a credential via an automatic restage?* **What stands there today is authentication, and a restage does not remove it.** The consume endpoint is gated by `checkAuthCustomer` (`api/offsite.go:14-18`) — **the customer's API key, the same credential as the config pull**. Anything that could consume a restaged password can already pull the customer's whole config. So an automatic restage **grants no new access to an unauthenticated party**. **What it genuinely weakens is the exposure window.** Today a live Hetzner sub-account password is fetchable only in the seconds after an operator deliberately staged one. Under an automatic restage it becomes fetchable **whenever the box reports no offsite target** — a condition an attacker holding the API key can arguably induce. The credential's blast radius at the provider is the customer's own sub-account (label-scoped; `ReissueCredentials` refuses unless the label lookup finds **exactly one**, `offsite.go:164,180`), so this is a widening of *when*, not of *what*. **And one further guard must survive any change** — R-39(a), `monitor/offsite_delivery.go:26-30`: `SaveOneTimeSecret` is last-write-wins, so restaging on top of an **unconsumed** secret clobbers a password a box may be about to consume. Any automatic restage must re-read the row immediately before acting and refuse unless it is still a **consumed** row. That guard is correct and is not what R-192 is about. --- ## Q7 — R-192's guard: exactly what it reads, and why it read wrong ### **RULING: both halves confirmed, with numbers. The guard reads the OLDEST 500 reports since the consume, and for demo-hp all 500 predated the rebuild by six days.** **The query, quoted** (`hub/internal/store/store.go:987`): ```sql SELECT report_json FROM reports WHERE customer_id = ? AND received_at > ? ORDER BY id LIMIT 500 ``` `ORDER BY id` ascending with `LIMIT 500` = **the oldest 500**, not the newest. **Reproduced against the live data (I1)**, with demo-hp's real consume anchor `2026-07-23 09:53:41` (recovered from R-192's record — it is no longer in `one_time_secrets`, see I4): | measurement | value | |---|---| | `total` (what the guard sees) | **500** | | `withOffsite` (what the guard sees) | **500** | | oldest report in that window | `2026-07-23 09:53:47` | | newest report in that window | **`2026-07-28 11:17:40`** | | true `total` over the same window, no LIMIT | 1174 | | true `withOffsite` over the same window, no LIMIT | 1063 (⇒ 111 without — matching the 111 ABSENT reports in I3) | **The whole 500-report evidence set ends 2026-07-28 — six days before the 2026-08-03 rebuild.** **Defect (a) — the message is exactly inverted.** `maybeEmitStuck` (`monitor/offsite_delivery.go:~110`) interpolates `status.ReportsSinceConsume` into a string that hardcodes the phrase *"carry no offbox target"*, and never reads `OffsiteReportsSinceConsume`. The live e-mail, quoted in full from the events table: > *"Offsite delivery stuck: one-time password consumed 284h19m0s ago and **500 report(s) since carry > no offbox target** — the credential is likely burned (apply died between consume and persist). > Re-issue delivers a fresh one."* `OffsiteReportsSinceConsume` for that same call was **500**. Every one of them carried a target. The message states the precise negation of its own measurement, and prescribes a remedy for a failure mode that did not occur. **Defect (b) — the heal refuses silently, on stale evidence.** `maybeHeal` returns bare on `status.OffsiteReportsSinceConsume != 0` (it was 500), with **no log line**. `offsite_credential_restaged` has never fired for any customer — confirmed: zero rows of that event type in the database. **What the guard SHOULD discriminate on — the shape, not the code.** The guard's intent is right: *"there is offbox evidence, so this is a regressed apply, not a burned credential — the operator's call."* Its error is that **it asks the question of the wrong time period**. Offbox evidence from before a rebuild is not evidence that the credential still works. The discriminator must be **recency-bounded and rebuild-aware**: judge on evidence *after the newest `controller_started` / config-hash change*, or on the latest N reports — never on "everything since the consume". And a refusal must **name its reason in the log**, because "we chose not to heal" and "the heal never ran" must not look identical (the S-16 rule, one layer down). **Do not fix it in isolation.** Under Q2's finding, a successful auto-restage on demo-hp would have restored the transport and the box would *still* have minted a new repo password. **The heal cannot protect the data; it can only protect the plumbing.** Whatever shape it takes must say so. --- ## Q8 — What each candidate would cost **A fourth option exists and is the important one**, because (a) and (b) both address ① and Q2 shows the damage is in ③. ### (a) The hub restages automatically when a re-enrolled box reports no offsite **A version of this ALREADY EXISTS and is wired to the wrong event.** `reissueOnReenroll` (`hub/internal/api/handler.go:1051-1084`), leg **F3**: > *"F3 — offsite continuity: re-stage the one-time offsite password to the fresh controller (the > one-time password only ever reached the OLD controller)."* It is called from `handleHostEnroll` — but **after** a mint-once-reuse short-circuit (`handler.go:1004-1016`): `if existing != nil { …return… }`. **A guest rebuild leaves the `hosts` row intact** (measured: `demo-hp-bb76ea` is the same host_id throughout the incident; only the agent leaf re-keyed, `host_leaf_changed` 2026-08-03 06:09:40), so re-enroll returns the existing credential and **F3 is never reached.** - **Changes:** hub only. Either lower the trigger from host-enrollment to a controller-level rebuild signal (the `config_hash` change + `controller_started` pair the hub already receives and already logs), or fix R-192's guard so the existing R-71c self-heal can act. - **Weakens:** the exposure window in Q6. Needs the R-39(a) unconsumed-secret guard kept intact. - **Does NOT solve:** ③. The box still mints a new repo password and still orphans the repo. **This option restores the plumbing to a box whose data key is already gone.** - **Cost:** small. **Value against the actual harm: near zero.** ### (b) The credential becomes recoverable from escrow at re-bootstrap > **ANNOTATED 2026-08-04 (recon) — this ruling is overturned in part. See the annotation at the top.** > The constraint is real but narrower than written: an **unattended** rebuild cannot recover from a > zero-knowledge escrow; a **customer-present** one can, and the operator has ruled on that shape. > The recon also found that the chain this option would use is clientless at three links and lossy at > a fourth (`RECON-offsite-dr-chain-2026-08-04.md` §3–§4). - **BLOCKED AS STATED, and this is a hard constraint, not an estimate.** The escrow blob is **R-wrapped and zero-knowledge** — the hub holds opaque bytes and has no recovery code (`api/handler.go`, and D6 in `CONTEXT.md` S-3). **A rebuilt box cannot unseal it without the customer's recovery code.** Any design that says "recover it at re-bootstrap" is describing a customer-present ceremony, i.e. what the existing manual `offboxInjectPasswordHandler` already is. - **Does NOT solve:** an unattended rebuild, which is the whole scenario. - **Cost: not implementable as specified.** ### (c) — NOT PREVIOUSLY NAMED — the agent retains the repo password across a guest rebuild The agent lives **on the Proxmox host** and survives a guest rebuild. It already receives the repo password today: the controller pushes it over the pinned local API (`POST /escrow/stage-secret` → `agent/internal/localapi/escrow_stage.go`), and the agent writes it to the fixed 0600 path `/var/lib/felhom-agent/escrow-stage/restic_repo_password` (`agent/internal/escrow/identity.go:44-46`) **transiently — wiped by the ceremony** (`WipeStagedResticPassword`). The seam, the transport, the pinning and the file are all already built. - **Changes:** agent (retain rather than wipe; serve back over the same pinned local API) + controller (on a fresh data dir, ask the agent before minting — the `InjectOffboxPassword` seam already exists and is exactly the right shape). - **Weakens:** the escrow's *zero-knowledge* posture in one specific way — the irreplaceable data key would now sit at rest on the Proxmox host, which the operator has root on. **Measured against D6, that is not a new exposure in principle** (D6 already states plainly that *"the operator cannot read customer data"* was never the security property, and the operator holds root on every box) — but it IS a new copy, and copies are the thing R-133 is already open about for the hub DB. **This is a real trade and it is the operator's to make.** - **Solves:** exactly the failure that occurred, unattended, on every rebuild, without any hub involvement and without any credential leaving the premises. - **Cost:** medium. Two repos, one existing seam each. ### (d) Do neither — keep the manual Re-issue, and fix what LIES - **Changes:** R-192's message + refusal log; correct the three false comments about what Re-issue rotates; and — the part that actually matters — **make a repo-password change VISIBLE**. The hub already receives `restic_pw_sha256` on every escrow upload and already stores it. Comparing the new hash against the superseded one is a two-line verdict, and it is the signal that was missing on demo-felhom for thirteen hours → **R-197**. (Note the shape: **this spike answered its own hardest question with a comparison the hub could be making automatically and is not.** Both values were already in the database; nothing read them.) - **Weakens:** nothing. - **Does NOT solve:** the data loss. Every rebuild still orphans the repo; the operator simply finds out the same day instead of never. - **Cost:** small. ### Ranked, with a recommendation | rank | option | why | |---|---|---| | **1** | **(d) now — the honesty pass** | It is cheap, it weakens nothing, and **it is the only option that helps the incident already in flight.** The alarm that lies is worse than no alarm; the repo-password-changed detector is the one signal whose absence let demo-felhom lose 1.14 GB of history silently. Ship this regardless of what is decided about the rest. | | **2** | **(c) — agent-retained repo password** | The only candidate that addresses ③, which is where the harm is. Every seam it needs exists. It costs one deliberate trade-off the operator must actually make. | | **3** | **(a) — auto-restage** | Worth doing **after** (c), never instead of it. On its own it restores the transport to a box whose data key is already gone, and would have made both boxes look healthy on 2026-08-04 while the snapshots were orphaned — a *strictly worse* outcome than the current loud failure. | | **4** | **(b) — recover from escrow at re-bootstrap** | Not implementable as stated; the escrow is zero-knowledge by design. | **Recommendation: ship (d), then decide (c).** And do not ship (a) first — it would have hidden this. --- ## What could not be established Named rather than guessed (§8 rule 4). 1. **The 2026-08-05 02:15 UTC run's actual outcome** (Q3). Predicted `ErrOffboxOrphaned` from source + claim state; **not measured**. 2. **Whether the orphaned ciphertext still exists at the provider.** 40.9 MB (demo-hp) + 1.14 GB (demo-felhom) are *presumed* present at `/home/felhom-repo` — reading the endpoint needs the SFTP credential, which this session did not touch. The Hetzner API has no directory-listing surface. 3. **Whether demo-felhom's 2026-07-21 Re-issue re-sealed an unchanged password** (the predicted false-staleness shape, Q4/Q5). Only two `host_escrow_superseded` rows exist in the whole database; the generation before 2026-07-21 is not retained. 4. **What removed demo-hp's offsite target at the rebuild** is fully explained (fresh data dir), but **why the guest was rebuilt on both boxes on consecutive mornings** is outside this spike. --- ## The operator's question — stated, and NOT answered here > **The irreplaceable secret in the offsite tier is the restic repository password. It is generated on > the box, it is destroyed by a guest rebuild, and the only copy that survives is sealed under the > customer's recovery code — which nothing but a human can open. Every machine is going to be > reinstalled.** > > **Do you want that key to survive a rebuild unattended — by keeping a copy on the Proxmox host, where > you already have root — or do you want it to stay sealed under the customer's recovery code and > accept that every rebuild starts a new repository and orphans the old one, provided the system says > so loudly on the day it happens?** Option (c) is the first. Option (d) alone is the second. **This spike does not choose.**