A hub the agent could not reach was reported to the customer as a bad recovery code. Measured live 2026-08-05 (CAMPAIGN-11 F3): hub firewalled off, a CORRECT current code, and the customer told it did not open their package — in 0.0556s against ~1.0s for a real unseal. No unseal was attempted. The discriminator existed here and this boundary threw it away: recover.go fails at four distinguishable points and the local-api handler had cases for two, with a default answering 'the recovery code did not open the sealed bundle, OR the bundle could not be fetched'. escrow.ErrBundleFetch now joins the fetch leg and the handler routes it to 502 with its own words — the code was NOT used. 502 not 4xx: the request was not bad, an upstream dependency failed. Four situations, four statuses: 502 fetch / 400 fetched-and-refused / 404 no bundle / 409 predates the field. The controller classifies on the STATUS and never parses the sentence. A GREEN TEST NAMED THIS DEFECT AND DID NOT PREVENT IT. TestRecoverOffsiteRepoPassword_FetchErrorIsDistinct has said since v0.125.0 that the operator must not be sent to re-read their code because the hub was unreachable — and passed throughout, because it asserted this package's error STRING one layer below the merge, and a string is not something a caller can branch on. Re-pointed at the sentinel, with a consequence-level twin asserting the status. Red-proofs: removing the %w join fails the sentinel test; deleting the handler case makes fetch and wrong-code both answer 400 with the wrong-code sentence. 29 packages ok, vet clean, agent gates OK.
3.2 KiB
REPORT — felhom-agent v0.126.0: a fetch failure is not a wrong recovery code (R-224)
Scope: this repo's half of R-224. The controller half ships as felhom-controller v0.202.0.
Why the agent changed at all
The task that commissioned this work scoped felhom-agent as untouched. It could not be. Its
Scenario A (a hub outage must not blame the customer's code) and Scenario C (a genuine mistype must be
told to re-check the ten words) are mutually unsatisfiable while this agent answers both with one
HTTP 400 and one sentence. No value available to the controller separates them. The task's own §5
anticipates this — "if the step is not recoverable from the value, make it so, and say what that
cost" — and §4.3 says the source outranks the register's recorded shape. The cost is this version,
a publish, and a MinAgent coupling on the controller side.
What changed
| File | Change |
|---|---|
internal/escrow/recover.go |
new ErrBundleFetch sentinel; the fetch leg joins it with %w: %w so the cause survives for the operator log |
internal/localapi/escrow_recover.go |
new case errors.Is(err, escrow.ErrBundleFetch) → 502 with its own words; the default now carries only the wrong-code case and drops the "or" |
internal/escrow/recover_test.go |
three new tests; the pre-existing FetchErrorIsDistinct re-pointed from a string to the sentinel, with the reason it failed to protect |
internal/localapi/escrow_recover_class_test.go |
new — the consequence-level test: four situations, four statuses |
Four statuses: 502 fetch failed (the code was not used) · 400 fetched and refused ·
404 no bundle · 409 bundle predates the field.
The finding this turned up
A green test named the defect and did not prevent it. TestRecoverOffsiteRepoPassword_FetchErrorIsDistinct
has asserted since v0.125.0 that "the operator must not be sent to re-read their recovery code because
the hub was unreachable". It passed throughout, because it checked this package's error string one
layer below the local-api default that did the merging — and a string is not something a caller can
branch on. Mechanism asserted, consequence unpinned; the project's own rule names this exact case.
It is also a comment-vs-code entry: recover.go's header said the errors were "DISTINCT on purpose"
and named three situations while a fourth was silently folded into one of them.
Green gate
go build ./... clean · go vet ./... clean · go test ./... → 29 packages ok ·
python3 scripts/agent_gates.py --fast → all gates OK.
Red-proofs, each demonstrated failing then restored:
| Mutation | Result |
|---|---|
remove the %w: %w join (pre-R-224 wrap) |
FetchFailureIsClassifiedAsFetch FAILS |
delete the ErrBundleFetch handler case |
fetch answers 400 "the recovery code did not open the sealed bundle" — the exact defect, and both status tests FAIL |
Not changed
No Proxmox surface, no privileged path, no report/hub contract, no config schema. The route's
success path, its scoping and its R-handling discipline (R = "" on both paths, never logged, never
persisted) are untouched.