Files
felhom.eu/REPORT.md
T

16 KiB
Raw Blame History

REPORT — hub v0.93.0: the retention keeps the key it was built to keep (R-198) + the honesty pass (R-197, R-192, R-196)

Date: 2026-08-04 · Class: implementation. One additive schema change, no destructive operation. Deployed: gitea.dooplex.hu/admin/felhom-hub:0.93.0, ArgoCD Synced, rollout complete, schema change confirmed applied on the live database.

Supersedes the earlier 2026-08-04 report (the RECON session), per the overwrite convention. Nothing durable from it lived only here — it is in audits/RECON-offsite-dr-chain-2026-08-04.md, OPEN-ITEMS.md R-198…R-202 and the capability map.


1. Confirmed baselines, and the three landmarks

Repo on arrival end state
felhom.eu 3f2b7bc02386, clean, == origin/main, hub v0.92.0 40687b0, hub v0.93.0 deployed
felhom-controller 0887fd676d60, clean untouched — Part 5 hit its gate (§7)

§3's three landmarks, re-confirmed in the real source before editing — none had drifted:

  1. host_escrow_superseded (store.go:377) had blob and no identity_blob.
  2. identity_blob was added to host_escrow by a later ALTER TABLE (store.go:414).
  3. demoteCurrentEscrowTx (store.go:2547) selected six columns and not that one.

Highest register ID on arrival: R-201. Grepped OPEN-ITEMS.md, ROADMAP.md, STATUS.md, CONTEXT.md before minting. One new item: R-202.

2. §8.1's two facts — both LOOKED AT on the live database, not deduced

Read from a hub.db + -wal + -shm snapshot taken after the deploy (PRAGMA integrity_checkok; freshness proven positively: MAX(host_reports.received_at) = 2026-08-04 10:45:39 vs datetime('now') = 10:58:59).

(a) Nothing is backfillable, and here is the looking:

id host_id identity_blob IS NOT NULL superseded_at
3 demo-hp-bb76ea 0 2026-08-04 07:15:36
4 demo-felhom-8363b5 0 2026-08-04 07:20:08

Both retained rows are empty of it, and their source rows were overwritten at those same timestamps — host_escrow now holds the new generation. No path preserved them. (The prompt asked me to check whether some path had, because that would have been a better answer than expected. It had not.)

(b) Who the fix protects, i.e. whose NEXT ceremony now retains a recoverable off-site key:

host_id identity blob present bytes
demo-felhom-8363b5 yes 572
demo-hp-bb76ea yes 572

2 of 2 hosts with a current escrow. Both were one ceremony away from losing their repository password again.

Schema change confirmed applied livePRAGMA table_info(host_escrow_superseded) on the running pod's database returns identity_blob BLOB as column 8. It runs at startup, so this is the check that the startup actually did it rather than the code merely containing it.

3. Files modified, and the commits

Commit 91cabdd — code + docs:

File Change
hub/internal/store/store.go identity_blob on host_escrow_superseded (CREATE + additive ALTER); carried in demoteCurrentEscrowTx; ListSupersededEscrow reads it; HostEscrow.IdentityBlob; new CountCurrentEscrowWithIdentity; SaveHostEscrow returns the replaced hash; the table comment rewritten to describe what the code now does and to record that it once did not; three R-196 comment corrections
hub/internal/api/handler.go maybeEmitRepoKeyChanged + eventRepoKeyChanged; wired into the superseded branch of the escrow PUT; one R-196 comment correction
hub/internal/monitor/offsite_delivery.go deliveryShape (burned/regressed); two-shape message; offsite_reports_since_consume in the details; recordHealRefusal; refusal branches split so each names its reason
hub/internal/notify/dispatcher.go offsite_repo_key_changed, offsite_delivery_stuck, offsite_credential_restagedoperatorOnlyEvents
hub/internal/offsite/offsite.go one R-196 comment correction
new: store/escrow_identity_retention_test.go, api/escrow_repo_key_changed_test.go, monitor/offsite_delivery_honesty_test.go Scenarios AG
11 existing _test.go files mechanical: SaveHostEscrow now returns three values
hub/CHANGELOG.md, CONTEXT.md (S-28), REUSE.md, STATUS.md, 00-capability-map.md, 03-host-agent.md, OPEN-ITEMS.md, ROADMAP.md docs

Commit 40687b0manifests/hub.yaml 0.92.00.93.0.

4. Tests, and every red-proof outcome

Green gate after each phase: go build ./... && go vet ./... && go test ./... in hub/rc=0, full suite, run separately from every commit (never combined).

Group Scenario Result Red-proof — mutation → outcome
A re-escrow retains both keys PASS drop identity_blob from the shared copy routine (= production ≤ v0.92.0) → FAIL: "the retained row carries NO identity blob — the off-site repository password was destroyed by the ceremony that was supposed to preserve it"
B host delete demotes both PASS same mutation → FAIL. Plus the §8.2 mutation: fix only the re-escrow caller (patch the retained row in SaveHostEscrow instead of the shared routine) → A passes, B FAILS — the "fixed one caller" mistake, demonstrated
C legacy host with no identity blob supersedes cleanly PASS — (a nullable-column check; no guard to remove)
D changed key raises one signal, host named, no hash values PASS remove the comparison from the escrow PUT → FAIL: "the repository key demonstrably changed and NO signal was raised", while E still passes
E unchanged key is silent PASS
F message states the measured state PASS restore the single hardcoded sentence → both message tests FAIL
G heal refusal leaves a durable record PASS replace the regressed branch with a bare return (= production ≤ v0.92.0) → FAIL, and its cadence test with it
IsOperatorOnly for all three event types PASS
CountCurrentEscrowWithIdentity census PASS

A mutation that did not remove every guard, recorded because rule §9.12 requires it. My first Scenario-F mutation kept the switch's default branch, so the regressed test passed for the wrong reason and only the burned one failed. That is not a red-proof. The mutation that does fail both is deleting the shape switch entirely and emitting the one v0.92.0 sentence — done, observed, recorded.

Both callers proven separately (§8.2): TestSaveHostEscrow_RetainsIdentityBlob (re-escrow) and TestDeleteHost_DemotesIdentityBlob (host delete). The pre-existing TestSaveHostEscrow_RetainsSuperseded stayed green throughout the two months the bug existed, because it asserts the mechanism (a retained row exists, carrying the old K-blob). The new tests assert the consequence (the retained row can still yield a repository password) — CLAUDE.md's rule, applied to the case that earned it.

Every -run filter was verified to have matched (-v, === RUN lines counted) before any result was read as a proof.

5. The exact operator message text, quoted for review

R-197 — offsite_repo_key_changed (warning, operator-only, edge-triggered once per supersession):

Offsite repository key CHANGED for host <host_id>: the new escrow seals a different repository password than the one it replaced. The previous off-site history is no longer opened by this box's current key. The superseding blob was retained (<N> held), so that history stays recoverable with the recovery code that sealed it — verify the box's off-site tier reports a repository rather than an orphan card, and expect the next backup to start a fresh history.

Details JSON: {"host_id": "...", "retained_count": N, "repo_key": "changed"}no hash values.

R-192 — offsite_delivery_stuck, burned shape (unchanged in intent, corrected in wording):

Offsite delivery stuck (BURNED-credential shape): the one-time password was consumed <age> ago; of the first <N> report(s) after that consume, NONE carried an offbox target, and the latest report carries none either. The credential never reached a persisted apply. Re-issue delivers a fresh one.

R-192 — regressed shape (the demo-hp case, which previously got the burned text):

Offsite delivery stuck (REGRESSED-apply shape): the one-time password was consumed <age> ago; of the first <N> report(s) after that consume, <M> DID carry an offbox target — and the latest report carries none. The credential was applied and worked; the target was lost afterwards. Re-issue is NOT the indicated action: find what removed the offbox target (a guest rebuild does, R-193). Automatic restage is deliberately withheld for this shape. NOTE: the counts cover at most the first 500 reports after the consume, so on a long-lived box they describe the start of the window, not now (R-192, open).

R-192 — the refusal record (notification_log, channel operator, status refused), message "Automatic offsite credential restage was NOT performed." with the reason in error_message, e.g.:

regressed-apply shape: <M> of the first <N> report(s) after the consume DID carry an offbox target, so a burned credential is ruled out — a restage would treat a symptom whose cause is elsewhere. Operator's call (R-193).

6. The corrected comments — FIVE, not three, with their locations as found

The spec expected three and named two; a census of every restic mention in the hub found five.

# Location What it claimed
1 hub/internal/offsite/offsite.go:198 "the restic repo password just changed" as the justification for MarkEscrowStale
2 hub/internal/api/handler.go:1067 "The re-issuer resets the restic repo password, which makes the OLD escrow blob stale"
3 hub/internal/store/store.go:423 (the stale_at ALTER) "stale_at is set when the offsite repo password is re-issued"
4 hub/internal/store/store.go MarkEscrowStale doc comment "called when the offsite repo password is re-issued"
5 hub/internal/store/store.go EscrowStatus.Stale field comment "true when the offsite password was re-issued"

All five now state what the code does, name the correction with its date and item, and cite the recon. The staleness mark is documented as precautionary — the box's re-apply may mint a fresh repository password — with R-197's measured signal named as the evidential one. The behaviour is unchanged and R-196 stays OPEN, because the false-staleness alarm is a behaviour change that must not ride a comment-correction release. It also matters more now than when filed: under R-198 an unnecessary ceremony is no longer harmless bookkeeping, it supersedes a blob.

7. Part 5 — the gate was hit; the card is untouched; R-202 filed

The gate: ship it iff the hub can tell a box what it needs with one additional boolean on the escrow ACK it already sends.

The hub can compute such a boolean cheaply"≥1 retained blob for this host carries an identity blob", one correlated predicate in GetEscrowStatusForCustomer — and the controller even has the right seam already (SetEscrowStale/StaleBlob is exactly this shape). But that boolean does not answer the card's question. The card renders on RepoState == "orphaned", and its promise is about the key THIS orphaned repository was written under. A box does not know which escrow generation the orphaned remote belongs to. A box with a pre-v0.93.0 orphan and a post-v0.93.0 supersession would read the boolean TRUE and the promise would still be false — a conditional falsehood that looks verified, which on a customer-facing card is worse than today's hedged one. Making it truthful needs the orphan's generation, which is the same knowledge R-199/R-201's unassembled chain needs.

So: stopped, filed R-202, left the card alone, per §8.6. Stated plainly rather than buried: the false sentence is still live, in Hungarian, on both demo boxes. The cheapest honest interim — drop the recoverability clause and say only that the old history is set aside and not deleted, which is true unconditionally — is recorded in R-202 and not taken here, because it is a customer-copy change and the gate said leave it alone.

8. Register

  • R-198 → SHIPPED (hub v0.93.0), with the §8.1 facts and both red-proof outcomes recorded.
  • R-197 → SHIPPED (hub v0.93.0).
  • R-192 → HALF SHIPPED: both honesty halves done; the guard's 500-oldest-reports scoping stays OPEN, deliberately, because its correct shape depends on the recovery chain that is not yet assembled. The window is named inside the alert text so the limitation travels with the number.
  • R-196 → comments corrected; the BEHAVIOUR stays OPEN.
  • R-193 → still open, updated with the three operator rulings.
  • R-202 → NEW (Part 5's gate).
  • R-199, R-200 and R-201 remain OPEN and untouched — the hub's blob-serving endpoints still have no client, the injection seam still has no form, and nothing in the recovery path has ever been exercised. v0.93.0 makes the key survive; it does not assemble the chain that hands it back.
  • The three operator rulings of 2026-08-04 are recorded in a dedicated block at the top of OPEN-ITEMS.md (a ruling that lives only in a conversation binds nobody — the R-96 rule): run the drill after R-198 (R-198 has shipped; the drill is next); delete the orphaned ciphertext — STILL OWED, deliberately not done here because a destructive act on a protected endpoint does not ride a schema-change release; and accept the risk on R-193(c), which is what makes this retention load-bearing rather than tidy.
  • Capability map, 03-host-agent.md §8a (a new Custody across generations section — the contract changed, so S-1 applies), CONTEXT.md S-28, REUSE.md and ROADMAP.md all updated in the same commit.

9. Deploy & validation

  • Clean-tree gate held per repo before the build; the image was built from a pushed tree.
  • ./build.sh 0.93.0 --pushfelhom-hub:0.93.0, 25 MB. Manifest bumped in git, committed, pushed.
  • ArgoCD hard-refresh + deliberate sync (never kubectl set image). Synced; deploy/hub rolled out; live image gitea.dooplex.hu/admin/felhom-hub:0.93.0; startup log clean (all checkers initialized, Listening on :8080, no errors).
  • Live validation was structural, not manufactured (§9.3): the schema column read from the running pod's database, and the two censuses in §2. No ceremony was run on any box — today, a ceremony is the act that destroys the key, and it is also the one thing that would have "proved" the retention by consuming a real customer's generation.
  • Disk headroom checked before the build (/mnt/5_hdd 24%, / 86%).
  • Nothing provisioned; nothing to tear down.
  • Gates: python3 scripts/repo_gates.py --fast → all OK, after every docs change.

10. CI

  • Code commit 91cabdd and manifest commit 40687b0 — run numbers, task ids and conclusions quoted in the session summary; --no-verify was NOT used (the pre-push hook ran repo_gates.py --fast and passed on both pushes).

11. Observations — noticed, not acted on

  1. ListSupersededEscrow still has no production caller. v0.93.0 makes it return the identity blob, so a retained key is reachable from Go for the first time — but nothing reads it. That is R-199's territory and was deliberately not widened here.
  2. handleReEnroll returns a freshly minted API key in its response body on an endpoint nothing currently calls. Worth a look before anything starts calling it (R-199).
  3. CountReportsOffsiteSince's LIMIT 500 is the only instrument in the delivery checker, and its window is now named in operator-facing text. If R-192's remaining half is ever fixed, that sentence must change with it — the R-100 corollary (an alarm whose text stopped matching what its verdict counts) applies directly.
  4. offsite_delivery_stuck had been reaching the customer channel by construction for its whole life; the only reason no customer row exists is that the demo customers have no configured recipient. An absent notification row is not evidence a leg is blocked.