Files
felhom.eu/documentation/audits/SPIKE-r241-recovery-offer-2026-08-07.md
T
admin 71c43f87c2
gates / gates (push) Successful in 6s
SPIKE R-241 — ruled: a MINTING defect, not a screen-predicate defect
Read-only spike on the standing venue (demo-hp VM 324 finalwalk-appliance),
read before anything else touched the fleet. No product code written.

THE RULING. The recovery screen was telling the truth. There genuinely was
nothing recoverable under the key the box held, because the box minted that
key itself over the top of a sealed package it already knew the hub held.
Fixing the predicate would paper over a box quietly making its own history
unopenable.

Three measurements carry it:
  1. WriteOffboxSecrets (offbox.go:411) mints on ONE input - does the file
     exist. Its two neighbours in the same file both consult
     GetHubEscrowIdentityPresent(). Same fact on three paths, used on two.
  2. The flag was the PRECONDITION of the chain that reached the minting:
     the retry job only logs when the declaration is live, and the venue
     logged it at 02:48:03Z - 30 minutes and six ticks before the mint.
  3. The box KNEW and discarded it: at 03:28:03Z, 35 minutes before the
     customer looked, escrow_confirm.go:154 computed the exact
     discriminator and logged the mismatch. Never persisted, never surfaced.

And the hub explicitly disclaims doing this - offsiteheal's package doc:
'credential automatic, key customer-present ... must not quietly widen'.

Also answered: Q2 (shape (b) is structurally unreachable - the escrow gate
sits upstream of the only orphan producer, with a positive control that the
scheduler was alive); Q7 (the 'create a new code' button does NOT destroy
the data - R-198's retention holds - but it converts a self-service recovery
into one needing an unbuilt read path, and it re-enables the screen while
invalidating the code that screen accepts).

New rows:
  R-242 - a release that changes customer-visible behaviour is not delivered
          until a golden carries it, and nothing enforces that. Recorded
          with three proposed shapes; deliberately NOT built.
  R-243 - a box in this state silently stops backing up and NO alarm fires:
          three individually-correct exclusions leave one state unobserved.

Highest register ID moves R-241 -> R-243.
2026-08-07 08:53:16 +02:00

29 KiB
Raw Blame History

SPIKE R-241 — why the recovery screen retired itself

2026-08-07 · READ-ONLY · no product code written · venue read before anything else touched the fleet

Venue: demo-hp VM 324 finalwalk-appliance, guest LXC 9201, hub customer finalwalk, host finalwalk-ed05d6. Read in the state the overnight walk left it. Nothing on it was changed.

Evidence for the events themselves: documentation/tests/finalwalk-r201-2026-08-07/journal.md. Everything below that is labelled measured was measured in this session, from the venue, not copied from that journal. Where the journal is the only witness (the customer-side observations, which happened before this session) it is cited as such and marked [walk].


THE RULING, first

R-241 is a MINTING defect, not a screen-predicate defect.

The screen was telling the truth. At the moment the customer looked there genuinely was nothing recoverable under the key the box was holding, because the box had minted that key itself, thirty minutes earlier, over the top of a sealed package it already knew the hub was holding for it.

Fixing the screen's predicate would make the screen appear — and would leave in place a box that quietly makes its own backup history unopenable. The predicate is not where this is fixed.

The reviewer's hypothesis in the task was correct, and it is now measured rather than read off two call sites. Two independent findings carry it:

  1. The minting does not consult the flag — and the flag was not merely available, it was the precondition of the chain that reached the minting. §Q3.
  2. The hub's self-heal explicitly disclaims doing this, in its own package doc: "credential automatic, key customer-present — the ruling this session implements and must not quietly widen." The minting widens exactly that ruling, from the box's side, where neither side is looking. §Q6.

And the box knew. §Q5: ten minutes after minting, it computed the exact discriminator and wrote the answer to its own log.


The measured timeline

All times UTC, from the venue's own controller log (docker logs felhom-controller, guest 9201) and from file mtimes in the controller's data volume. Guest is UTC; the appliance is CEST.

time event source
2026-08-06 22:17:02 escrow ceremony seals the repository key 30ef574f…; hub stores it, fingerprint 81:dc:91:…:68:d5 hub host_escrow, measured
02:40:31 guest 9201 destroyed --purge [walk]
02:43:28 Day-0 provision SUCCESS; rebuilt guest, controller 0.203.0 [walk]
02:43:02 controller starts; settle-gate: awaiting floor knowledge (first report ACK) log measured
02:43:12 settle-gate: GO — at/above floor 0.156.0 (we are 0.203.0); reconcile finds no unconsumed password log measured
02:48:03 credential retry: … (the box still declares a need; retrying) — the first tick that proves the declaration is live log measured
02:53 · 02:58 · 03:03 · 03:08 · 03:13 same line, five more times log measured
02:59 onward hub: offsite-delivery: finalwalk: self-heal REFUSED — the box DECLARES offsite.state=needs_credential [walk]
03:15 hub offsiteheal re-stages the stored one-time secret, unaided [walk]
03:18:06 offsite configured for u629488-sub4@…:/home/felhom-repo (pending key escrow) + credential retry: the staged credential was collected and the tier appliedand WriteOffboxSecrets mints 9b4a9a9d… log measured; key hash measured
03:28:03 [WARN] [escrow-confirm] the hub's escrow blob does not cover the CURRENT repo password (hub hash 30ef574fe492… != local 9b4a9a9dcec7…) … staying pending log measured
~04:00 the customer logs in. / → „Indítópult". No recovery pointer. /recovery 302s away. [walk]
04:41:15 operator moves the self-heal key aside, places the recovered key 30ef574f… mtime measured
04:43:03 [INFO] [escrow-confirm] hub-verified: the escrow covers the current repo password (hash 30ef574fe492…) — EscrowState auto-confirmed escrowed log measured

The two keys, measured on the venue this session

offbox/repo_password                  sha256 = 30ef574fe492a43f89bf1a5071c44e89f51c44c6b08ebcb184b320a2634fad75   (mtime 04:41:15)
offbox/repo_password.selfheal-aside   sha256 = 9b4a9a9dcec7898e7544f35b18470aac77c3d9064e5d3a302897617fa62edd65   (mtime 03:18:06)

Both reproduce the journal's figures exactly, computed independently here. The values were never read, printed or moved — only hashed in place inside the guest.

The hub's side, measured this session

host_escrow           finalwalk-ed05d6
  key_fingerprint     81:dc:91:ce:a1:d0:50:3a:…:09:68:d5
  created_at          2026-08-06T22:17:02Z
  restic_pw_sha256    30ef574fe492a43f89bf1a5071c44e89f51c44c6b08ebcb184b320a2634fad75
  stale_at            NULL
  identity_blob       572 bytes  (present — R-198 material)
host_escrow_superseded finalwalk-ed05d6 → NO ROWS

So the hub held — and still holds — a sealed package covering 30ef574f…, the key that opens the history at ep0. The hash the box logged at 03:28:03 as "hub hash 30ef574fe492…" is this row. Both sides of the mismatch are now confirmed from independent sources.


Q1 — what did the screen evaluate, and to what?

Measured (state), [walk] (the customer-side observation).

recovery_handlers.go:34backup.OffsiteRecoveryOffer() (offbox.go:1411):

if !m.settings.GetHubEscrowIdentityPresent() { return false }   // fact 1
if _, ok := m.OffboxRepoPasswordHash(); !ok  { return true  }   // shape (a)
return m.OffboxOrphaned()                                       // shape (b)

At the moment the customer looked, the three inputs were:

input value how established
hub_escrow_identity_present true measured in settings.json; and see Q3 — it must have been true from 02:48:03 at the latest
repository password on disk present (9b4a9a9d…) measured: repo_password.selfheal-aside, mtime 03:18:06 — the file that was repo_password until the operator set it aside
RepoState == "orphaned" false (field absent) measured: settings.json carries no repo_state; and Q2 — no run ever probed

So OffsiteRecoveryOffer() = true && !(no password) && false = false.

recoveryPageHandler (recovery_handlers.go:79) then does exactly what it is written to do:

if !s.recoveryOffer() { http.Redirect(w, r, "/backups/remote", http.StatusFound); return }

That is the 302 the walk observed. It is the R-215 gate, shipped in v0.201.0 to stop the page asserting a recovery story on a box that has none — working correctly, on a box that had one.

The state today is different and must not be mistaken for the failed state. The venue now reads escrow_state: "escrowed" with repo_password = 30ef574f…, because the operator's 04:41 intervention made the local key match the hub's package and the auto-confirm flipped it at 04:43. OffsiteRecoveryOffer() is still false today, and correctly so — the box now holds the right key. The failed state survives only in the aside file, the mtimes and the log.


Q2 — can shape (b) ever fire on a box in this state?

Answer: no. It is structurally unreachable, and the block is upstream of the detection.

RepoState = "orphaned" has exactly one producer: markOrphaned() (offbox.go:270), called only from ensureOffboxRepo() (offbox.go:692, :697). ensureOffboxRepo is called only from runOffboxInternal (:1162) and offboxPruneOnly (:1494), both inside a run.

And every run passes this gate first (offbox.go:743, before any repo contact):

if !m.offboxEscrowed() {
    m.logger.Printf("[INFO] [offbox] skipped — pending key escrow (…)")
    return nil
}

offboxEscrowed() is EscrowState == "escrowed". The box was pending — and could not stop being pending, because the auto-confirm flips pending→escrowed only when the hub's hash equals the local hash, and those were 30ef574f… vs 9b4a9a9d…. The box said so itself at 03:28:03: "staying pending".

So the loop is closed in both directions, and it is self-locking exactly as the walk described:

no escrow → no run → no repo probe → no orphan state → shape (b) false. And: fresh key ≠ hub's key → auto-confirm never flips → no escrow. Forever.

Positive control, because an absent log line is not evidence. The controller log for the whole post-rebuild window (02:43→06:43) contains zero [offbox] lines — but that alone would be equally consistent with "the scheduler died". It did not: agent-channel-health ran 241 times, stack-scan 120, offsite-credential-retry 48, hub-report 16. The scheduler was alive throughout. offbox-backup is a sched.Daily leg (main.go:937), not a periodic job — its slot is ~02:15 controller time, which fell before the 02:40 destruction, so the next was ~02:15 tomorrow. No scheduled run was due, none was requested, and had one been, line 743 would have returned before the probe. The absence is explained, not merely observed.

Is the escrow block itself right — should a run be prevented while an escrow is pending?

Yes, and it should not be touched. It is the fork-4 atomicity guarantee: no offsite ciphertext may exist that nobody can decrypt. Relaxing it to let orphan detection run would create exactly the un-recoverable ciphertext the gate exists to prevent, and would be a much worse defect than the one it fixed. Shape (b) is not reachable by loosening this gate, and no option below proposes it.


Q3 — does WriteOffboxSecrets consult the hub's sealed-package flag?

Answer: NO. And the flag was not merely available — it was the precondition of the chain that reached the minting.

The minting, in full (offbox.go:410-419):

// Auto-generate the repo password once (0600), never log it.
if _, err := os.Stat(m.offboxPwPath()); os.IsNotExist(err) {
    pw, gerr := generateOffboxPassword()
    
    os.WriteFile(m.offboxPwPath(), []byte(pw), 0o600)
}

One input: does the file exist. No settings read, no GetHubEscrowIdentityPresent(), no hub hash, no escrow state. The Manager has m.settings in hand three lines away and does not ask it.

Compare the two neighbours that do ask, both in the same file:

call site consults the flag?
OffsiteRecoveryOffer()offbox.go:1412 yes, GetHubEscrowIdentityPresent() is fact 1
needsOffsiteCredential()offbox.go:1377 yes, same call
WriteOffboxSecrets()offbox.go:411 no

The same fact is available on all three paths and used on two. That is the shape the task predicted, and it is confirmed.

Was the flag available at 03:18:06? Yes — provably, and this is the strongest single measurement.

Not by timestamp comparison, but by the product's own control flow. The 5-minute retry job (main.go:787-795) only does work when the box declares:

attempted, err := offsiteBridge.RetryIfDeclared(ctx, func() bool {
    st := backupMgr.OffboxReportStatus()
    return st != nil && st.State == backup.OffsiteStateNeedsCredential
})
if !attempted { return nil }   // no log line at all

OffboxReportStatus() returns that state only via needsOffsiteCredential() — which returns false unless GetHubEscrowIdentityPresent() is true (offbox.go:1377-1380).

The venue logged credential retry: … (the box still declares a need; retrying) at 02:48:03, and five more times before the mint. That line is unreachable unless the flag was already true.

The flag was true at 02:48:03 — thirty minutes and six ticks before the minting at 03:18:06. The minting is the terminal step of a loop whose entry condition is "the hub is holding a sealed package for us". The box walked that whole loop and then minted a key as though it had never asked.


Q4 — the state table

P = hub holds a sealed package (hub_escrow_identity_present) · K = repository password on disk · O = RepoState=="orphaned" · H = hub's ACK restic_pw_sha256 vs local hash · S = hub_escrow_superseded_present.

Today's predicate: Offer = P && (!K || O).

# state P K O H today correct? should be
1 never had off-site backups n/a no offer no offer — nothing to recover
2 pristine rebuild, credential not yet arrived n/a offer (shape a) offer — but this window is ~1535 min and closes by itself
3 self-healed with a fresh key, hub holds the older packageTHE VENUE mismatch no offer offer — this is R-241
4 healthy, key matches the package match no offer no offer
5 re-escrowed — customer made a new code after a self-heal; old package retained (S=✓) match no offer no offer from the current package — but the retained older history is unreachable and unmentioned (R-199 read path unbuilt, R-222 message only)
6 orphaned — a run proved the repo will not open mismatch offer (shape b) offer
7 customer set the old data aside (ResetOrphanedRepo) ✗ (cleared) mismatch, until a new ceremony runs no offer no offer — and see the hazard below

Two things this table makes visible that the task did not ask for

(i) Row 7 is the trap in the obvious fix. resetOrphanedRepo (offbox.go:291-332) moves the old repo aside, re-inits under the current passphrase and clears RepoState — it does not run a new escrow ceremony. So the hub's package still seals the old key and H stays mismatch. A discriminator that offers the screen on H-mismatch alone would re-offer the recovery screen forever to a customer who has explicitly said they do not want the old data. Any fix needs a decided latch, or the reset flow must carry the ceremony with it. (Reasoned from code; not measured — reaching row 7 requires a run, which requires escrow, which the venue never had.)

(ii) Row 3's box cannot self-heal out of the state either, and NOTHING ALARMS. It is not just that the screen is missing: escrow_state is stuck pending forever, so off-site backups never run again. The box is not only unrecoverable-by-the-customer, it is silently not backing up.

All three signals that could catch it are excluded, each for its own reason — verified in the hub:

signal why it does not fire
offsite_stale isStale (monitor/offsite.go:135) returns false unless EscrowState == "escrowed". Its own comment: "Pending/disabled = normal onboarding, never stale." A row-3 box is classified as a box still being set up — forever.
offsite_delivery_stuck monitor/offsite_delivery.go:91 skips the applied shape. The credential was consumed and the target is present in reports, so delivery is applied and correct.
backup_failed never fires — the run returns nil at the escrow gate (offbox.go:743-746) before touching anything. Nothing fails, so nothing reports a failure.

Each exclusion is individually right. Together they leave the state unobserved. The walk found the recovery journey broken; the same state also silently stops the product's core promise, and that has not been filed. See R-243 below.


Q5 — is there a discriminator that answers the real question directly?

Answer: YES. It already exists, it is already computed on every report cycle, and the result is thrown away.

The question — does the hub hold a sealed package whose key differs from the one this box is using? — is precisely:

es.IdentityBlobPresent && es.ResticPwSHA256 != "" && localHash != es.ResticPwSHA256

Both halves are present and wired today:

  • the hub sends restic_pw_sha256 on every report ACK (report/escrow_confirm.go:21) — the hash of the password the sealed blob covers;
  • the box has LocalHash() (offbox.go:539, OffboxRepoPasswordHash).

And EscrowAutoConfirmer.Reconcile already performs exactly this comparison, at escrow_confirm.go:154. On the venue it produced, at 03:28:03:

[WARN] [escrow-confirm] the hub's escrow blob does not cover the CURRENT repo password (hub hash 30ef574fe492… != local 9b4a9a9dcec7…) — run the escrow ceremony …; staying pending

The box computed the right answer to the right question, thirty-five minutes before the customer looked, and wrote it to a log. It did not persist it, did not surface it, and the recovery screen — evaluating in the same process — went on asking a question that could not see it.

Would it give the right answer in every row of Q4?

No — not on its own. Honest assessment, row by row:

row H-mismatch discriminator alone
1 correct — P false, no ACK escrow object, no offer
2 needs shape (a) kept — there is no local key, so there is no hash to mismatch. LocalHash() returns ok=false and reconcileEscrowed returns early. Shape (a) is still required.
3 correct, and it is the fix
4 correct — hashes match, no offer
5 correct for the current package; says nothing about the retained one (unchanged, R-199)
6 correct — mismatch is exactly why the repo would not open
7 wrong — see hazard (i). Re-offers forever after the customer declined.

So the discriminator is necessary and nearly sufficient: it fixes rows 3 and 6, must keep shape (a) for row 2, and needs one additional latch for row 7.

There is a second, stronger check — CheckOffsiteKeyRecoverable (offbox_recovery_check.go, R-200) — which recovers the key through the agent and compares hashes. It is not usable here: it requires R, the customer's recovery code, which is the very thing the screen exists to ask for. It is the post-code diagnostic, not the pre-code discriminator.


Q6 — what did the self-heal intend?

Answer: the transport credential, and explicitly not key material. The minting is a side-effect of the path, not an intent of either side.

The hub's own package doc (hub/internal/offsiteheal/reconciler.go, header) is unusually explicit:

WHAT IT DELIBERATELY DOES NOT DO: it never runs, or asks for, an escrow ceremony. A credential is replaceable; the recovery code is not, because only the customer holds it. Credential automatic, key customer-present — the ruling this session implements and must not quietly widen.

And, on why an already-keyed box cannot reach it at all:

It cannot reach this reconciler at all, because the controller's declaration predicate requires the repository password to be ABSENT (backup.needsOffsiteCredential).

The hub therefore acts on a stated request whose stated meaning is "I have no repository password", re-stages a transport secret, makes no provider call and mints no new password, and hands back. The box then applies it through ApplyOffsiteTargetWriteOffboxSecrets, and that — a function whose job is writing an SSH key and a known_hosts line — mints a 256-bit repository key on the way past because a file was missing.

Neither side intended it and neither side can see it. The hub ruled "key customer-present" and honoured it. The box honoured the transport contract. The repository key was created in the seam between them, by a helper doing the thing its doc comment says it does. That is why this survived review: every individual comment is accurate.

This is R-218's shape one level up, exactly as the walk said — and the escalation is real: R-218 was "succeeding at recovery stopped the box asking for what it still needed"; this is "succeeding at the credential self-heal stopped the box being able to recover at all".


Q7 — the „Helyreállítási kód létrehozása" button

This was flagged as possibly the most important thing in the spike. The finding is a genuine near-miss, and the thing that stops it being a disaster is R-198 — shipped for exactly this, and it holds.

What the customer saw [walk]: /backups/remote offering „Helyreállítási kód szükséges … → Helyreállítási kód létrehozása". Measured: that is backups_remote.html:132-141, the branch {{else if and .OffboxConfigured (ne .Offbox.EscrowState "escrowed")}} — which is precisely the venue's state (configured, pending). It links to /backup/escrow.

It was the only affirmative action on the page. The recovery screen was gone, /recovery 302'd away, and this was what remained.

What it actually does, traced

(Reasoned from code and the hub schema. I did not press it — it is a state change and it would have destroyed the evidence this spike exists to read. Marked as such deliberately.)

A new ceremony seals the current key — 9b4a9a9d…, the self-heal's — under a new recovery code, and uploads it. On the hub, SaveHostEscrow (store.go:2779) finds a current row sealing a different hash and calls demoteCurrentEscrowTx before overwriting:

INSERT INTO host_escrow_superseded (host_id, blob, , identity_blob)
SELECT host_id, blob, , identity_blob FROM host_escrow WHERE host_id = ?

Copy-before-delete, and since R-198 the identity_blob is copied too — that column is "where the offsite restic REPOSITORY password lives", and its omission for two months is the defect R-198 fixed.

So: how close did last night come to a customer destroying their own history?

Not as close as it looks — and closer than anyone should be comfortable with.

the ciphertext at ep0 untouched. Nothing in this path writes to the repository.
the sealed package protecting it retained, not destroyed — demoted to host_escrow_superseded with its identity_blob. R-198 holds.
the customer's held R still opens that retained package. The material survives.
but the read path for a retained package UNBUILT — R-199's inventory. The product has no way to use it.
and the recovery engine fails closed against the current package. R would now produce the R-222 message: „a régebbi csomagot… nem töröltük — megőrizzük, megnyitni viszont innen egyelőre nem lehet."

The honest answer: the button does not destroy the data. It converts a situation that was one screen away from self-service recovery into one that needs an operator, a database and a read path that does not exist yet. Nothing is lost; everything becomes unreachable by any shipped path. R-198 is the whole margin, and R-198 shipped nine days ago.

One further consequence, which cuts the other way and is worth stating (reasoned from code): after a new ceremony escrow_state flips to escrowed, so runs unblock; the first run's ensureOffboxRepo probes a repo keyed under 30ef574f… with a local key of 9b4a9a9d…, classifies orphaned, and — the box being claimed — calls markOrphaned(). OffboxOrphaned() becomes true and shape (b) fires: the recovery screen comes back. But by then the code that screen accepts is the new one, and the customer's held R opens only the retained package the screen cannot read.

So the button re-enables the screen and invalidates the code it asks for, in the same act. That is a worse trap than a plain dead end, because it looks like progress.


What I could not establish

Named, per rule 9.

  1. Row 7 of Q4 and the post-button behaviour in Q7 are reasoned from code, not measured. Both require a state change on the venue (a run, or pressing the button). Reaching either would have destroyed the evidence for everything above. If the operator wants them measured, that is a separate scoped session with its own fixture — and it should be, because hazard (i) shapes the fix.
  2. The exact wall-clock at which hub_escrow_identity_present first became true is not measured — SetHubEscrowIdentityPresent writes only on change and logs nothing. What is established is the bound that matters: true at or before 02:48:03, thirty minutes before the mint, by the control- flow argument in Q3. A tighter figure would add nothing.
  3. Whether the customer would in fact have pressed the button. Unknowable. What is measured is that it was the only affirmative action offered.
  4. applied_marker (offbox/applied_marker, 64 bytes, mtime 03:18:06) — noted, not investigated; it plays no part in any predicate above.

Options, ranked

No code was written. No option below has been started.

Have EscrowAutoConfirmer.Reconcile persist the comparison it already makes (a new settings mirror beside hub_escrow_identity_present, same last-write-wins discipline), and add a third shape to OffsiteRecoveryOffer: the hub's package covers a different key than the one we hold.

  • Fixes rows 3 and 6; keeps shape (a) for row 2; needs a decided latch for row 7.
  • Uses only facts already on the wire — no hub change, no new ACK field, no protocol version.
  • The fact is already computed on every cycle; this stops discarding it.
  • Shape (b) can then be retired or kept as a corroborator — it becomes redundant for row 6, and it is the shape that could never fire on the state that needed it.
  • Cost: one settings field, one predicate branch, one latch, and the tests that pin them.

Option B — do not mint when the hub holds a package for us

Make WriteOffboxSecrets refuse to auto-generate when GetHubEscrowIdentityPresent() is true, and leave the box in shape (a) — the state the recovery screen was designed for.

  • Attacks the root cause directly, and is the smallest change that makes the walk's headline false.
  • But it changes a helper that six call sites and eleven tests depend on, and it makes ApplyOffsiteTarget return a target with no usable key — a state OffboxConfigured() currently treats as not-configured, which would need tracing before it could be trusted.
  • Strictly better as a companion to A than as a substitute. On its own it leaves row 6 to shape (b), which is the shape that failed.

Option C — A and B together ★ the honest end state

B stops the box creating the problem; A makes the box able to describe the problem when it exists anyway (a key minted before B ships, a hand-injected key, a restored volume). This is where the fix should land eventually. It is listed separately because doing both at once is a larger change than one session should make against a defect this subtle, and A alone already restores the journey.

Option D — surface it to the operator only

Turn the 03:28 WARN into a hub event and an operator alert; change nothing customer-facing.

  • Cheapest, and insufficient. It makes the failure visible on a dashboard while leaving the customer with no route to their data. It is worth doing as well — the state in Q4 hazard (ii) silently stops off-site backups and nothing alarms — but it is not a fix for R-241.

Rejected outright

Loosening the escrow-pending run gate so orphan detection can run. It would create un-recoverable offsite ciphertext, which is the exact hazard fork-4 exists to prevent. Named here so it is not rediscovered as clever. See Q2.


THE OPERATOR'S DECISION — owed, and deliberately not taken here

1. Which option, and in what order? The recommendation is A now, B next, converging on C. Taking A alone leaves boxes able to mint a redundant key that the screen will then correctly offer to fix — safe, but noisier than necessary.

2. Row 7's latch — what counts as "the customer has decided"? A persisted decision flag, or making ResetOrphanedRepo carry the new ceremony with it. This shapes A's design and A should not be started before it is answered, because it is the one way A can regress a shipped flow.

3. Is Q4 hazard (ii) — a box in row 3 silently stops backing up, with nothing alarming — worth its own fix, or does it fall out of A? It falls out of A for boxes that recover. It does not alarm for boxes that do not. Filed as R-243.

4. Should Q7's post-button behaviour and Q4 row 7 be measured before the fix is written? They are reasoned from code here. Measuring them needs a fresh fixture and its own session.

These are stated and left unanswered, per §3.4.


Register

ID
R-241 RULED — minting defect. Body updated in OPEN-ITEMS.md with this spike's evidence. Not fixed; no code written.
R-242 NEW — a controller release that changes customer-visible behaviour is not delivered until a golden carries it, and nothing enforces that (§4.1 of the task; recorded, not built).
R-243 NEW — a box in Q4 row 3 is stuck escrow_state: pending forever, so off-site backups silently never run again, and no alarm of any kind fires: isStale requires escrowed, offsite_delivery_stuck skips the applied shape, and backup_failed needs a run that never happens. Three correct exclusions leaving one state unobserved. Found by this spike; not part of the walk's finding.

Highest register ID moves R-241 → R-243.