diff --git a/CONTEXT.md b/CONTEXT.md index 040a9f5..c07f37e 100644 --- a/CONTEXT.md +++ b/CONTEXT.md @@ -17,6 +17,41 @@ ## Standing rulings +**S-29 — a box may fetch its OWN sealed recovery blob with its OWN credential; the operator-driven DR +path is a separate thing and stays gated (2026-08-04, R-199; hub v0.94.0 + agent v0.125.0 + controller +v0.195.0).** + +**The decision, so it is never re-litigated from scratch.** Self-service recovery needs the box to +obtain its own sealed package. The only retrieval that existed was operator-gated behind recovery mode +and the global key (`dr.go`), which cannot serve a customer-present flow. `GET /hosts/{id}/escrow` now +serves that ONE object to its authenticated owner. + +- **Why it is safe:** the blob is opaque; the hub has never held R and has no decrypt path; a wrong + code fails closed at age's scrypt KDF (measured live — `exit status 1`, no plaintext, nothing + written); the caller already authenticates as this host for its report, desired state, WG + registration and PBS token, so this adds no new identity, only a new object. +- **What it changes, which is the whole of the trade:** obtaining the blob used to require the + OPERATOR to arm recovery mode. It now needs only that box's credential. **Fewer parties are + required.** The mitigation is that the capability is AUDITED, not silent: every retrieval raises + `escrow_blob_served` (warning, operator-only), written before the bytes leave. +- **The override is one constant.** `api.escrowSelfServiceRetrieval` — false re-imposes the + recovery-mode requirement on this endpoint and changes nothing else. Do not scatter the condition. +- **`handleReEnroll` / `handleGetRestoreDirective` are UNTOUCHED** and must stay so: they rotate the + host API key and serve the K-escrow and the directive as well. Different object, different gate. + Pinned by `TestEscrowGet_OperatorDRPathUnchanged`. + +**Facts a future session needs before touching this area:** +- **The unseal lives in the AGENT and must stay there.** `age` is an agent runtime dependency + (`/usr/bin/age`, present on both demo hosts) and is deliberately absent from the controller image; + the blob is host-scoped and its only writer is the agent under the per-host key. +- **The agent returns ONE field.** Not the tunnel token, not the PBS token, not the WG key — the + controller is a trust tier down. Narrowing cost nothing and is not recoverable later. +- **R's rules:** in memory for one call, cleared on every path, never on disk, never in argv, never + logged, never echoed. The controller CLI reads it from STDIN for exactly that reason. Verified live: + 0 log lines, 0 files, 0 leftover staging dirs, with a positive control proving the search worked. +- **The chain today: links 1–8 walked, 9–11 not.** The KEY comes back. Nothing installs it, reopens a + repository with it, or restores a file — R-200's remaining half and R-201. + **S-28 — the escrow retention now covers the OFF-SITE data key, and customer-present recovery is the accepted design, which makes that retention load-bearing (2026-08-04, R-198/R-197; hub v0.93.0).** diff --git a/REPORT.md b/REPORT.md index 485ad70..a71b167 100644 --- a/REPORT.md +++ b/REPORT.md @@ -1,241 +1,221 @@ -# REPORT — hub v0.93.0: the retention keeps the key it was built to keep (R-198) + the honesty pass (R-197, R-192, R-196) +# REPORT — the offsite backup key comes back (2026-08-04, R-199 closed, R-200 half) -**Date:** 2026-08-04 · **Class:** implementation. One additive schema change, no destructive operation. -**Deployed:** `gitea.dooplex.hu/admin/felhom-hub:0.93.0`, ArgoCD **Synced**, rollout complete, schema -change **confirmed applied on the live database**. - -> Supersedes the earlier 2026-08-04 report (the RECON session), per the overwrite convention. Nothing -> durable from it lived only here — it is in `audits/RECON-offsite-dr-chain-2026-08-04.md`, -> `OPEN-ITEMS.md` R-198…R-202 and the capability map. +**Class:** implementation across three repos + one decisive live proof. Nothing wiped, deleted, +rotated or reconfigured; no customer-facing change in any language. --- -## 1. Confirmed baselines, and the three landmarks +## 1. SCENARIO A — THE RESULT, FIRST, BECAUSE IT IS THE SESSION -| Repo | on arrival | end state | -|---|---|---| -| `felhom.eu` | `3f2b7bc02386`, clean, == `origin/main`, hub v0.92.0 | `40687b0`, hub **v0.93.0** deployed | -| `felhom-controller` | `0887fd676d60`, clean | **untouched** — Part 5 hit its gate (§7) | +**On `demo-felhom`, 2026-08-04 13:49 CEST, through the real endpoints end to end:** -**§3's three landmarks, re-confirmed in the real source before editing — none had drifted:** +``` +=== offsite key recovery check (R-200) — compares, never installs === + on-disk sha256: c60c8bc737a6b7c6647c7849283f52087f650a885babedb4ef5fdf9a5c9543cb + recovered sha256: c60c8bc737a6b7c6647c7849283f52087f650a885babedb4ef5fdf9a5c9543cb + [MATCH] the offsite repository password IS recoverable from the sealed escrow. + Nothing was written: this check compares and never installs. +EXIT=0 +``` -1. `host_escrow_superseded` (`store.go:377`) had `blob` and **no** `identity_blob`. -2. `identity_blob` was added to `host_escrow` by a **later** `ALTER TABLE` (`store.go:414`). -3. `demoteCurrentEscrowTx` (`store.go:2547`) selected six columns and not that one. +**The hashes matched.** And they match a **third, independent** source: `host_escrow.restic_pw_sha256` +on the hub reads `c60c8bc737a6…` for `demo-felhom-8363b5`, recorded at the ceremony and never involved +in this comparison. The box's file, the sealed bundle and the hub's record all agree. -**Highest register ID on arrival: R-201.** Grepped `OPEN-ITEMS.md`, `ROADMAP.md`, `STATUS.md`, -`CONTEXT.md` before minting. One new item: **R-202**. +**The offsite repository password is recoverable from the sealed escrow with the customer's recovery +code.** That had never been demonstrated — the one live round-trip on record (2026-06-10) predates the +field by a month. -## 2. §8.1's two facts — both LOOKED AT on the live database, not deduced +## 2. Part 0 — the recovery code -Read from a `hub.db` + `-wal` + `-shm` snapshot taken after the deploy (`PRAGMA integrity_check` → -`ok`; freshness proven positively: `MAX(host_reports.received_at)` = `2026-08-04 10:45:39` vs -`datetime('now')` = `10:58:59`). +**Answered by the operator before the session began:** the 2026-08-04 07:16/07:20 codes had not been +kept, so **the operator ran a fresh ceremony on both nodes and saved the codes** (supplied to this +session as `R_DEMO-FELHOM` in the DooPlex credentials file). No ceremony was run by CC. -**(a) Nothing is backfillable, and here is the looking:** +Measured hub-side, and worth recording because it exercised two of yesterday's behaviours live: +the new escrow rows are stamped `2026-08-04T11:11:37Z` (demo-hp) and `11:13:06Z` (demo-felhom), and +`restic_pw_sha256` is **unchanged** on both — so `SaveHostEscrow` correctly treated them as +same-password re-ceremonies: **no superseded row was created** and **no `offsite_repo_key_changed` +fired**. That is v0.93.0's Scenario E, live. It also means yesterday's `identity_blob` retention is +still unit-proven and structurally verified, **not yet exercised live** — no supersession has happened. -| id | host_id | `identity_blob IS NOT NULL` | superseded_at | +Subject box: **demo-felhom** (Tier 0, reachable with a baked key, current escrow). Not demo-hp +(break-glass only), not DooPlex, not ep0. + +## 3. The §8.2 decision as built, and where the gate is + +**Built as §8.2:** `GET /api/v1/hosts/{host_id}/escrow`, authenticated with the box's own per-host key, +self-scoped (global may read any — the same asymmetry the PUT has). + +**The single gate point is `api.escrowSelfServiceRetrieval`** (`hub/internal/api/handler.go`), a named +constant. `false` re-imposes the recovery-mode requirement on this endpoint and changes nothing else — +the same unseal, the same extraction, the same proof, minus the self-service property. The trade is +written above the handler in `dr.go`'s style: what it gives out, why it is safe, and **what it changes +about who is required** — obtaining the blob used to need an operator to arm recovery mode and now +needs only that box's credential. They still cannot open it. The mitigation is the audit row. + +## 4. Files, commits, deployed versions + +| Repo | Commits | Version | Deployed | |---|---|---|---| -| 3 | `demo-hp-bb76ea` | **0** | `2026-08-04 07:15:36` | -| 4 | `demo-felhom-8363b5` | **0** | `2026-08-04 07:20:08` | +| `felhom.eu` | `435f4a5` (code) · `dd08926` (manifest) · docs commit below | hub **v0.94.0** | ArgoCD Synced/Healthy, `felhom-hub:0.94.0` | +| `felhom-agent` | `6d79047` | **v0.125.0** | published sha `f7d8339b53d9…`, **verified by independent download**, installed on demo-felhom (`felhom-agent --version` → 0.125.0, `active`). **NOT vouched** | +| `felhom-controller` | `9640e51` | **v0.195.0** | `felhom-controller:0.195.0` on guest 9201, `Up (healthy)` | -Both retained rows are empty of it, and their source rows were overwritten at those same timestamps — -`host_escrow` now holds the *new* generation. **No path preserved them.** (The prompt asked me to check -whether some path had, because that would have been a better answer than expected. It had not.) +**Hub:** `internal/api/handler.go` (route + `handleHostEscrowGet` + `escrowSelfServiceRetrieval`), +`internal/notify/dispatcher.go` (`escrow_blob_served` → `operatorOnlyEvents`), +`internal/api/escrow_get_test.go` (new). +**Agent:** `internal/escrow/recover.go` (new — `OffsiteKeyRecoverer`, distinct errors), +`internal/hub/client.go` (`FetchIdentityEscrow`), `internal/localapi/escrow_recover.go` (new route), +`internal/localapi/server.go` (seam + registration), `cmd/felhom-agent/main.go` (production wiring + +the §8.6 selftest message fix), plus `recover_test.go` and `escrow_recover_wiring_test.go` (new). +**Controller:** `internal/backup/offbox_recovery_check.go` + `offbox_recovery_cli.go` (new), +`internal/agentapi/escrow.go` (`RecoverOffsiteRepoPassword`), `cmd/controller/main.go` +(`--recover-offsite-check`), plus `offbox_recovery_check_test.go` (new). -**(b) Who the fix protects, i.e. whose NEXT ceremony now retains a recoverable off-site key:** +**Green gate in each repo after each phase** (`go build && go vet && go test ./...`, rc=0), plus +`repo_gates.py` / `agent_gates.py` / `controller_gates.py` — all OK. **No test run was combined with a +commit.** -| host_id | identity blob present | bytes | +## 5. Tests and every red-proof + +| Scenario | Result | Red-proof — mutation → outcome | |---|---|---| -| `demo-felhom-8363b5` | yes | 572 | -| `demo-hp-bb76ea` | yes | 572 | +| **A** recovered key == on-disk key | PASS (unit, real `age`) + **PROVEN LIVE** | return `bundle.PBSToken` instead of `ResticRepoPassword` → **FAIL**: *"the recovered key is not the sealed repository password (len 24 vs 64)"*. This is the mutation that mattered: every field of that bundle is a non-empty string that looks like a secret | +| **B** wrong R fails closed, nothing written | PASS + **PROVEN LIVE** | — (the property is age's KDF, asserted rather than guarded, so there is no guard to remove) | +| **C** a box fetches only its own blob | PASS | delete the `!isGlobal && authHostID != pathHostID` check → **FAIL**: host A was served host B's blob, body and all | +| **D** no blob → a clean "none" | PASS | — | +| **E** every retrieval leaves a record | PASS + **PROVEN LIVE** | `if false &&` on the audit block → **FAIL**: *"a sealed recovery blob was served and NOTHING recorded it"* | +| **F** R persists nowhere | PASS + **PROVEN LIVE** | write R to a temp file inside the unseal → **FAIL** (see the correction below) | +| **G** the operator DR path is untouched | PASS | — | +| **H** the seam is wired (AST) | PASS | `_ = escrowRecoverer` with the `Options` field removed → **FAIL**: *"localapi.Options … carries no EscrowRecovery field"* | +| §8.5 compare-never-install | PASS + **PROVEN LIVE** | add `InjectOffboxPassword` to the check → **FAIL**: *"the check INSTALLED the recovered password"* | -**2 of 2 hosts with a current escrow.** Both were one ceremony away from losing their repository -password again. +**TWO MUTATIONS DID NOT REMOVE EVERY GUARD ON THE FIRST ATTEMPT, AND BOTH ARE RECORDED RATHER THAN +QUIETLY REDONE** (§9.13 — two of the last four sessions found a red-proof that passed): -**Schema change confirmed applied live** — `PRAGMA table_info(host_escrow_superseded)` on the running -pod's database returns `identity_blob BLOB` as column 8. It runs at startup, so this is the check that -the startup actually did it rather than the code merely containing it. - -## 3. Files modified, and the commits - -**Commit `91cabdd`** — code + docs: - -| File | Change | -|---|---| -| `hub/internal/store/store.go` | `identity_blob` on `host_escrow_superseded` (CREATE + additive `ALTER`); carried in `demoteCurrentEscrowTx`; `ListSupersededEscrow` reads it; `HostEscrow.IdentityBlob`; new `CountCurrentEscrowWithIdentity`; `SaveHostEscrow` returns the replaced hash; the table comment rewritten to describe what the code now does **and** to record that it once did not; three R-196 comment corrections | -| `hub/internal/api/handler.go` | `maybeEmitRepoKeyChanged` + `eventRepoKeyChanged`; wired into the superseded branch of the escrow PUT; one R-196 comment correction | -| `hub/internal/monitor/offsite_delivery.go` | `deliveryShape` (burned/regressed); two-shape message; `offsite_reports_since_consume` in the details; `recordHealRefusal`; refusal branches split so each names its reason | -| `hub/internal/notify/dispatcher.go` | `offsite_repo_key_changed`, `offsite_delivery_stuck`, `offsite_credential_restaged` → `operatorOnlyEvents` | -| `hub/internal/offsite/offsite.go` | one R-196 comment correction | -| new: `store/escrow_identity_retention_test.go`, `api/escrow_repo_key_changed_test.go`, `monitor/offsite_delivery_honesty_test.go` | Scenarios A–G | -| 11 existing `_test.go` files | mechanical: `SaveHostEscrow` now returns three values | -| `hub/CHANGELOG.md`, `CONTEXT.md` (S-28), `REUSE.md`, `STATUS.md`, `00-capability-map.md`, `03-host-agent.md`, `OPEN-ITEMS.md`, `ROADMAP.md` | docs | - -**Commit `40687b0`** — `manifests/hub.yaml` `0.92.0` → `0.93.0`. - -## 4. Tests, and every red-proof outcome - -Green gate after each phase: `go build ./... && go vet ./... && go test ./...` in `hub/` — **rc=0**, -full suite, run separately from every commit (never combined). - -| Group | Scenario | Result | Red-proof — mutation → outcome | -|---|---|---|---| -| A | re-escrow retains both keys | PASS | drop `identity_blob` from the shared copy routine (**= production ≤ v0.92.0**) → **FAIL**: *"the retained row carries NO identity blob — the off-site repository password was destroyed by the ceremony that was supposed to preserve it"* | -| B | host delete demotes both | PASS | same mutation → **FAIL**. **Plus the §8.2 mutation**: fix only the re-escrow caller (patch the retained row in `SaveHostEscrow` instead of the shared routine) → **A passes, B FAILS** — the "fixed one caller" mistake, demonstrated | -| C | legacy host with no identity blob supersedes cleanly | PASS | — (a nullable-column check; no guard to remove) | -| D | changed key raises one signal, host named, no hash values | PASS | remove the comparison from the escrow PUT → **FAIL**: *"the repository key demonstrably changed and NO signal was raised"*, while E still passes | -| E | unchanged key is silent | PASS | — | -| F | message states the measured state | PASS | restore the single hardcoded sentence → **both message tests FAIL** | -| G | heal refusal leaves a durable record | PASS | replace the regressed branch with a bare `return` (**= production ≤ v0.92.0**) → **FAIL**, and its cadence test with it | -| — | `IsOperatorOnly` for all three event types | PASS | — | -| — | `CountCurrentEscrowWithIdentity` census | PASS | — | - -**A mutation that did not remove every guard, recorded because rule §9.12 requires it.** My first -Scenario-F mutation kept the `switch`'s `default` branch, so the *regressed* test passed for the wrong -reason and only the *burned* one failed. That is not a red-proof. The mutation that does fail both is -deleting the shape switch entirely and emitting the one v0.92.0 sentence — done, observed, recorded. - -**Both callers proven separately** (§8.2): `TestSaveHostEscrow_RetainsIdentityBlob` (re-escrow) and -`TestDeleteHost_DemotesIdentityBlob` (host delete). The pre-existing -`TestSaveHostEscrow_RetainsSuperseded` stayed **green throughout the two months the bug existed**, -because it asserts the mechanism (a retained row exists, carrying the old K-blob). The new tests assert -the consequence (the retained row can still yield a repository password) — `CLAUDE.md`'s rule, applied -to the case that earned it. +1. **Scenario F.** The planted leak wrote R into the walked directory and **the test still passed** — + because a later failing call OVERWROTE the leak file with the wrong code, and the test only scanned + for the first one. Instrumenting the mutation (printing where it wrote) is what found it. **The test + was wrong, not the mutation.** It now asserts **emptiness** of the TMPDIR tree as its primary check — + nothing under it is created by the test, so any survivor is a leak regardless of content — with the + content scan kept as defence in depth over all three secrets. Re-run: **FAIL**, *"the unseal left + 1 file(s) behind under TMPDIR"*. +2. **Scenario H.** Commenting the wiring line out failed to **compile** (`declared and not used`), which + is not a red-proof. The mutation that compiles keeps the variable alive (`_ = escrowRecoverer`) and + drops the struct field. Re-run: **FAIL** as quoted above. **Every `-run` filter was verified to have matched** (`-v`, `=== RUN` lines counted) before any result was read as a proof. -## 5. The exact operator message text, quoted for review +## 6. The live validation trail, in order -**R-197 — `offsite_repo_key_changed`** (warning, operator-only, edge-triggered once per supersession): +1. **Wrong code first, deliberately** (13:44:44) — hub: `escrow blob SERVED to host demo-felhom-8363b5 + (572 opaque bytes, self_scope=true) — recovery path in use`; agent: `offsite key recovery FAILED … + err="escrow: the recovery code did not unwrap the identity escrow (wrong recovery code, or a corrupt + blob): exit status 1"`; CLI: `[FAIL] … nothing was written.`, exit 1. + **This is Scenario B live AND it proves links 6 and 7 ran independently of the success** — the blob + was fetched and the unseal was attempted and refused. +2. **The real code** (13:49:47) — §1's output, exit 0. +3. **Nothing was written** — `offbox/repo_password` mtime on the live box is still + `2026-08-03 07:18:02`, unchanged by a check that ran at `11:49` UTC. The whole `offbox/` directory is + byte-for-byte the pre-check listing. -> Offsite repository key CHANGED for host ``: the new escrow seals a different repository -> password than the one it replaced. The previous off-site history is no longer opened by this box's -> current key. The superseding blob was retained (`` held), so that history stays recoverable with -> the recovery code that sealed it — verify the box's off-site tier reports a repository rather than an -> orphan card, and expect the next backup to start a fresh history. +## 7. The retrieval record as it appears at the hub -Details JSON: `{"host_id": "...", "retained_count": N, "repo_key": "changed"}` — **no hash values**. +``` +customer_id severity created_at details_json +demo-felhom warning 2026-08-04 11:44:44 {"blob_bytes":572,"host_id":"demo-felhom-8363b5","self_scope":true} +demo-felhom warning 2026-08-04 11:49:47 {"blob_bytes":572,"host_id":"demo-felhom-8363b5","self_scope":true} +``` -**R-192 — `offsite_delivery_stuck`, burned shape** (unchanged in intent, corrected in wording): +And the routing, which is the half that shows the register works: -> Offsite delivery stuck (BURNED-credential shape): the one-time password was consumed `` ago; of -> the first `` report(s) after that consume, NONE carried an offbox target, and the latest report -> carries none either. The credential never reached a persisted apply. Re-issue delivers a fresh one. +``` +customer_id channel status created_at +demo-felhom operator sent 2026-08-04 11:44:44 +demo-felhom customer skipped 2026-08-04 11:44:44 ← operator_only +demo-felhom operator suppressed 2026-08-04 11:49:47 ← 1h operator cooldown, itself recorded (R-182) +demo-felhom customer skipped 2026-08-04 11:49:47 +``` -**R-192 — regressed shape** (the demo-hp case, which previously got the burned text): +**Both retrievals raised an event; the operator was mailed for the first; the second's mail was +cooldown-suppressed and that suppression is written down.** The audit is per-retrieval and complete +even where the mail is not. -> Offsite delivery stuck (REGRESSED-apply shape): the one-time password was consumed `` ago; of -> the first `` report(s) after that consume, `` DID carry an offbox target — and the latest report -> carries none. The credential was applied and worked; the target was lost afterwards. Re-issue is NOT -> the indicated action: find what removed the offbox target (a guest rebuild does, R-193). Automatic -> restage is deliberately withheld for this shape. NOTE: the counts cover at most the first 500 reports -> after the consume, so on a long-lived box they describe the start of the window, not now (R-192, open). +## 8. Evidence that R persisted nowhere — the search, not a claim -**R-192 — the refusal record** (`notification_log`, channel `operator`, status `refused`), message -*"Automatic offsite credential restage was NOT performed."* with the reason in `error_message`, e.g.: +Searched on the subject host with the pattern read from a file (never in argv), across the agent +journal, the controller's container log, and `/tmp`, `/var/tmp`, `/var/lib/felhom-agent`, `/root`: -> regressed-apply shape: `` of the first `` report(s) after the consume DID carry an offbox -> target, so a burned credential is ruled out — a restage would treat a symptom whose cause is -> elsewhere. Operator's call (R-193). +``` +agent journal lines containing R: 0 +controller log lines containing R: 0 +host files containing R: 0 +leftover felhom-idesc-* staging dirs: 0 +staged escrow secret file: 0 entries +``` -## 6. The corrected comments — FIVE, not three, with their locations as found +**With a positive control, because an absent match is not a measurement unless the instrument is shown +to work:** a copy of R was planted under `/tmp`, the same sweep found **1**, the copy was shredded, and +the sweep returned **0**. The instrument detects R when R is there. -The spec expected three and named two; a census of every `restic` mention in the hub found **five**. +R reached the box only as: the credentials-file value on DooPlex → a pipe → the container's stdin → +the request body → the agent's memory. It was never an argument, never a file on either machine, and +appears in no log line at any level. -| # | Location | What it claimed | -|---|---|---| -| 1 | `hub/internal/offsite/offsite.go:198` | *"the restic repo password just changed"* as the justification for `MarkEscrowStale` | -| 2 | `hub/internal/api/handler.go:1067` | *"The re-issuer resets the restic repo password, which makes the OLD escrow blob stale"* | -| 3 | `hub/internal/store/store.go:423` (the `stale_at` ALTER) | *"stale_at is set when the offsite repo password is re-issued"* | -| 4 | `hub/internal/store/store.go` `MarkEscrowStale` doc comment | *"called when the offsite repo password is re-issued"* | -| 5 | `hub/internal/store/store.go` `EscrowStatus.Stale` field comment | *"true when the offsite password was re-issued"* | +## 9. Register -All five now state what the code does, name the correction with its date and item, and cite the recon. -The staleness mark is documented as **precautionary** — the box's re-apply may mint a fresh repository -password — with R-197's measured signal named as the evidential one. **The behaviour is unchanged and -R-196 stays OPEN**, because the false-staleness alarm is a behaviour change that must not ride a -comment-correction release. It also matters *more* now than when filed: under R-198 an unnecessary -ceremony is no longer harmless bookkeeping, it supersedes a blob. +- **R-199 → SHIPPED + PROVEN-LIVE.** Links 6, 7 and 8 assembled and walked. +- **R-200 → plumbing shipped; the FORM is not built** and the row stays open for it and for link 9. +- **R-201 → OPEN, and the distinction is deliberate:** the KEY is proven recoverable; a recovered + password has never been **installed**, no repository has been **reopened** under one, and **no file + has ever been restored**. The drill's pass condition is unchanged — a byte-identical sentinel file, + not "the store opened". +- **R-202 → OPEN and untouched.** The orphan card still promises recoverability unconditionally. +- **The orphaned-ciphertext deletion (~1.2 GB) is STILL OWED** — operator ruling 2 of 2026-08-04, + deliberately not ridden along with a code release. +- Capability map, `03-host-agent.md` §8a, `CONTEXT.md` **S-29**, `REUSE.md` and `ROADMAP.md` updated; + `STATUS.md` rewritten for the operator. -## 7. Part 5 — the gate was hit; the card is untouched; R-202 filed +## 10. The capability-map row, and what it deliberately does not claim -**The gate:** ship it iff the hub can tell a box what it needs with **one** additional boolean on the -escrow ACK it already sends. +Added: *"The offsite repository password can be RECOVERED from the sealed escrow with the customer's +recovery code"* — **PROVEN-LIVE (2026-08-04)**, with the evidence above. -**The hub can compute such a boolean cheaply** — *"≥1 retained blob for this host carries an identity -blob"*, one correlated predicate in `GetEscrowStatusForCustomer` — and the controller even has the right -seam already (`SetEscrowStale`/`StaleBlob` is exactly this shape). **But that boolean does not answer -the card's question.** The card renders on `RepoState == "orphaned"`, and its promise is about *the key -THIS orphaned repository was written under*. A box does not know which escrow generation the orphaned -remote belongs to. A box with a pre-v0.93.0 orphan and a post-v0.93.0 supersession would read the -boolean TRUE and the promise would still be false — a **conditional** falsehood that looks verified, -which on a customer-facing card is worse than today's hedged one. Making it truthful needs the orphan's -generation, which is the same knowledge R-199/R-201's unassembled chain needs. +**What the row explicitly refuses to claim**, stated in the row itself because the previous over-claim +on this subject was struck out four hours earlier: it covers the **key**, not the **data**. No +recovered password has been installed, no repository reopened, no file restored. And the proof used a +box whose local key still exists — the rebuilt-box case, where there is nothing to compare against, is +exactly what the drill covers and it has not run. -**So: stopped, filed R-202, left the card alone**, per §8.6. **Stated plainly rather than buried: the -false sentence is still live**, in Hungarian, on both demo boxes. The cheapest honest interim — drop the -recoverability clause and say only that the old history is set aside and not deleted, which is true -unconditionally — is recorded in R-202 and **not taken here**, because it is a customer-copy change and -the gate said leave it alone. +## 11. CI -## 8. Register +Run numbers, task ids and conclusions for all commits are quoted in the session summary. +**`--no-verify` was NOT used** — every push ran its repo's pre-push gate and passed. -- **R-198 → SHIPPED** (hub v0.93.0), with the §8.1 facts and both red-proof outcomes recorded. -- **R-197 → SHIPPED** (hub v0.93.0). -- **R-192 → HALF SHIPPED**: both honesty halves done; **the guard's 500-oldest-reports scoping stays - OPEN**, deliberately, because its correct shape depends on the recovery chain that is not yet - assembled. The window is named inside the alert text so the limitation travels with the number. -- **R-196 → comments corrected; the BEHAVIOUR stays OPEN.** -- **R-193 → still open**, updated with the three operator rulings. -- **R-202 → NEW** (Part 5's gate). -- **R-199, R-200 and R-201 remain OPEN and untouched** — the hub's blob-serving endpoints still have no - client, the injection seam still has no form, and nothing in the recovery path has ever been - exercised. v0.93.0 makes the key survive; it does not assemble the chain that hands it back. -- **The three operator rulings of 2026-08-04 are recorded** in a dedicated block at the top of - `OPEN-ITEMS.md` (a ruling that lives only in a conversation binds nobody — the R-96 rule): - run the drill after R-198 (**R-198 has shipped; the drill is next**); **delete the orphaned - ciphertext — STILL OWED**, deliberately not done here because a destructive act on a protected - endpoint does not ride a schema-change release; and **accept the risk on R-193(c)**, which is what - makes this retention load-bearing rather than tidy. -- Capability map, `03-host-agent.md` §8a (a new *Custody across generations* section — the contract - changed, so S-1 applies), `CONTEXT.md` S-28, `REUSE.md` and `ROADMAP.md` all updated in the same - commit. +## 12. Teardown -## 9. Deploy & validation +**Nothing was provisioned.** The temporary pattern file used by the R sweep and its positive control +were shredded on the host by the same command that created them (verified: 0 files remain). No +diagnostic artifact persists on either machine. The agent's previous binary is retained as +`/usr/local/bin/felhom-agent.bak-0.124.1`, per the documented deploy. -- Clean-tree gate held per repo before the build; the image was built from a pushed tree. -- `./build.sh 0.93.0 --push` → `felhom-hub:0.93.0`, 25 MB. Manifest bumped in git, committed, pushed. -- ArgoCD hard-refresh + **deliberate sync** (never `kubectl set image`). **Synced**; `deploy/hub` - rolled out; live image `gitea.dooplex.hu/admin/felhom-hub:0.93.0`; startup log clean (all checkers - initialized, `Listening on :8080`, no errors). -- **Live validation was structural, not manufactured** (§9.3): the schema column read from the running - pod's database, and the two censuses in §2. **No ceremony was run on any box** — today, a ceremony is - the act that destroys the key, and it is also the one thing that would have "proved" the retention by - consuming a real customer's generation. -- Disk headroom checked before the build (`/mnt/5_hdd` 24%, `/` 86%). -- **Nothing provisioned; nothing to tear down.** -- Gates: `python3 scripts/repo_gates.py --fast` → all OK, after every docs change. +## 13. Observations — noticed, NOT acted on -## 10. CI - -- Code commit **`91cabdd`** and manifest commit **`40687b0`** — run numbers, task ids and conclusions - quoted in the session summary; **`--no-verify` was NOT used** (the pre-push hook ran `repo_gates.py - --fast` and passed on both pushes). - -## 11. Observations — noticed, not acted on - -1. **`ListSupersededEscrow` still has no production caller.** v0.93.0 makes it *return* the identity - blob, so a retained key is reachable from Go for the first time — but nothing reads it. That is - R-199's territory and was deliberately not widened here. -2. **`handleReEnroll` returns a freshly minted API key in its response body** on an endpoint nothing - currently calls. Worth a look *before* anything starts calling it (R-199). -3. **`CountReportsOffsiteSince`'s `LIMIT 500` is the only instrument in the delivery checker**, and its - window is now named in operator-facing text. If R-192's remaining half is ever fixed, that sentence - must change with it — the R-100 corollary (an alarm whose text stopped matching what its verdict - counts) applies directly. -4. **`offsite_delivery_stuck` had been reaching the customer channel by construction** for its whole - life; the only reason no customer row exists is that the demo customers have no configured - recipient. An absent notification row is not evidence a leg is blocked. - +1. **The `escrow_blob_served` mail is subject to the 1-hour operator cooldown** — visible above, where + the second retrieval's mail was suppressed. The EVENT rows are per-retrieval so the audit is + complete, but if retrieval ever becomes routine, a burst would produce one mail. Worth revisiting + **when** the customer-facing flow lands, not before. +2. **`ListSupersededEscrow` still has no production caller.** Recovery reads the CURRENT blob; a + customer needing a PREVIOUS generation's key (which is what R-198's retention exists for) has no + path at all. That is the next gap in this area after the drill. +3. **The controller binary is at `/usr/local/bin/felhom-controller`, not `/app/…`** — the first + diagnostic invocation failed on a guessed path. Worth a line in the runbook when one is written. +4. **`age` is at the hardcoded `/usr/bin/age` on both demo hosts** (1.2.1 on demo-felhom) — the recon + listed this as unestablished; it is now established for demo-felhom. `ageBinary` is a package + variable with no config override, so a host that installs `age` elsewhere would fail the unseal at + the last step. +5. **Yesterday's `identity_blob` retention has still not been exercised live** — the operator's fresh + ceremonies sealed the *same* password, so no supersession occurred. It remains unit-proven with a + verified schema; the first real supersession will be its live proof. diff --git a/REUSE.md b/REUSE.md index da5f4e9..28622bf 100644 --- a/REUSE.md +++ b/REUSE.md @@ -78,6 +78,7 @@ | `(*Store).HasEverBoundHost` (v0.92.0, R-195) | hub/internal/store/store.go | `(customerID) (bool, error)` | Any verdict that must not fire for a customer with **no machine ever bound** — "was anything ever expected of this customer" | `hosts` row **OR** `host_deletions` tombstone. **NOT a liveness check and never a substitute for one:** a box that was bound and went silent returns `true` and must keep alarming — that is the case any change here breaks first (pinned by `TestCheckBackupDeadlines_BoundButNeverReported_StillAlarms`). Callers **fail OPEN** on its error: an unreadable binding must never SUPPRESS an alarm. Do **not** re-derive this from report presence — `store.GetCustomers()` (and therefore the staleness checker's `down` state) is a query over `reports`, so a never-reported customer has no state at all, which is exactly how the daily false alarm reached `david`. | | `(*Server).configFormData` (v0.49.0) | hub/internal/web/configs.go (~L430) | `(r, isNew, cfg, overrides, errMsg) configFormView` | The ONE view-model builder for the customer config form (standalone chrome + the customer page Edit tab) | `overrides=nil` → parses the STORED cfg.ConfigJSON; pass the SUBMITTED map on the update validation-error re-render or typed values reset (red-proofed). | | `config_form_body` sub-template (v0.49.0) | hub/internal/web/templates/config_form_body.html | `{{template "config_form_body" }}` | Rendering the config form on ANY surface (config_form.html chrome + customer Edit tab) | The floor/geo/danger cards on the Edit tab are SIBLINGS after `` — never nest a form inside it (breaks the offsite/PBS formaction sub-buttons). Includes the F5 in-flight `