- OPEN-ITEMS: R-204 all four CLOSED with both 2026-08-05 rulings recorded (the declared-state trigger and its four-meanings-of-absence reasoning; the recovery preview's dashboard-password exposure accepted as metadata, not content). R-193's credential half CLOSED, screen + deletion still open. R-192 CLOSED by REPLACEMENT. R-202 untouched. - R-212 MINTED (R-211 was the highest, grepped): the orphaned-ciphertext deletion HALTED at its STOP because the measured paths do not match the register — three set-aside stores totalling ~1.45 GB, and the thing that is exactly 1.2 GB is demo-felhom's LIVE repo. Nothing was deleted. - capability map: all four interventions closed; the row KEEPS a qualifier for a new reason — no step needs an operator, but there is no customer-facing recovery screen, and the journey has not been re-walked end to end. - 07-backup-architecture 7.0: the four-step table updated; the declaration-vs- inference reasoning and the credential-automatic/key-customer-present split. - CONTEXT: standing ruling S-33. - STATUS: the headline change and the deletion STOP. - REPORT-r204-item4.md rather than REPORT.md: a parallel session is active in this shared clone.
14 KiB
REPORT — R-204 item 4 / R-193 credential half (hub v0.96.0), 2026-08-05
A rebuilt box asks for its credential back, and the hub answers. The last of the four manual
interventions the 2026-08-04 drill needed. Controller half: felhom-controller v0.199.0.
Written as
REPORT-r204-item4.md, notREPORT.md. A PARALLEL SESSION is active in this shared clone — it committedc917251(the R-205…R-211 disk containment) between my baseline read and my first commit. PerCLAUDE.md's parallel-session rule the second session never touches the sharedREPORT.md. Explicit per-file staging was used throughout; verified after the fact that none of my three commits carries a foreign file.
1. Baselines, re-read on arrival — with a drift
| Repo | Expected | Found |
|---|---|---|
felhom.eu |
0dbd954fec90 / hub v0.95.0 |
DRIFTED to ee9d9bf — one docs-only commit ahead (R-205…R-211 spike output). Hub code, manifests/hub.yaml and the deployed image were all still v0.95.0, so the drift did not affect the work. A second foreign docs commit (c917251) landed mid-session. |
felhom-controller |
68f195676b91 / v0.198.0 |
exact match, tree clean |
felhom-agent |
v0.125.0 | 3f5f61b, clean (rider only, no version bump) |
app-catalog-felhom.eu |
n/a | 122bbee, clean (rider only, no version bump) |
The task's "highest register ID in use: R-204" was STALE. Grepped before minting, as instructed: the highest is R-211. This session mints R-212 (§8 below).
2. Part 2.0 FIRST — does the stored value survive a consume? YES.
Established from the schema and the code, deliberately not from the PBS analogy:
one_time_secretsdeclaresvalue TEXT NOT NULL(store.go, the CREATE TABLE).ConsumeOneTimeSecretrunsUPDATE one_time_secrets SET consumed_at = datetime('now')— it touches nothing else. The value column is never cleared or overwritten.
So restage-before-mint is possible, and RestageOneTimeSecret mirrors RestageHostPBSSecret:
clear the consumed flag, never touch the value, never bump a generation, return false when no row
exists so the caller escalates.
This is asserted rather than assumed by TestRestageOneTimeSecret_ReArmsTheSameValue, which checks
the same value comes back — so a future hardening that cleared the column fails loudly instead of
silently turning every rebuild into an external mint. The two secrets are different objects with
different lifecycles, and assuming a shared shape is how two earlier sessions confused the credentials.
3. The declared state — name, shape, and how a pre-upgrade hub treats it
Name: needs_credential, on the report's existing offsite object (not a new top-level field).
Constants: backup.OffsiteStateNeedsCredential / offsiteheal.StateNeedsCredential, pinned to each
other by TestDeclaredStateStringMatchesTheReconciler.
Shape, as received by the hub live (report id 16743):
{"enabled": false, "escrow_state": "", "state": "needs_credential",
"snapshot_count": 0, "repo_size_bytes": 0, "quota_gb": 0}
A pre-upgrade hub treats it as INERT, established from the readers' code rather than assumed:
| Reader | Behaviour on the declaration | Why |
|---|---|---|
OffsiteChecker.isStale |
no staleness alarm | returns early on !off.Enabled |
OffsiteChecker.fillBand |
bandOK |
returns OK on a zero quota/size |
encoding/json |
ignores state |
unknown field, not an error |
reportHasOffsite |
would have MISREAD it | see below — the one reader that needed changing |
A configured box's report JSON is byte-identical to v0.198.0's — state is omitempty and is
never set on a configured box (TestOffsiteDeclare_ConfiguredBoxJSONIsUnchanged).
4. The debounce: TWO distinct reports, derived not chosen
The controller reports every ~15 minutes. Two distinct declarations mean the state survived a full report cycle, and a restart, a slow first report or a transient config read all resolve well inside one. One would act on a blip; three would leave a genuinely stranded customer waiting ~45 minutes for the one thing they cannot obtain any other way. The reconciler's own sweep is 5 minutes — deliberately faster than the report cadence so it adds no latency of its own — and the debounce counts fresh evidence, not ticks, so a fast sweep cannot shorten it.
5. §8.4 — is there a deliberately-unhealed state here? Yes, and it is excluded upstream.
pbsdrheal refuses to heal verify_failed because re-staging would not help it and it must stay
loud. The off-site analogue is the regressed shape: a box that HAD a working tier and lost its
target while still holding its repository password. Re-arming a credential would not help it either.
It cannot reach this reconciler at all — the controller's declaration predicate requires the repository password to be absent, so a box that still holds one never declares. The unhealable case is excluded by construction, upstream, rather than filtered out in a switch. Stated explicitly because "there is nothing like that here" is usually wrong, and this was checked rather than assumed.
6. Files modified
| File | Change |
|---|---|
hub/internal/store/store.go |
new RestageOneTimeSecret, new LatestReportOffsiteDeclaration; reportHasOffsite tightened to require enabled:true (+ its comment corrected) |
hub/internal/offsiteheal/reconciler.go |
new package — the reconciler |
hub/internal/offsiteheal/{reconciler,wiring}_test.go |
new — Scenarios A–F + the wiring AST test |
hub/internal/store/offsite_restage_test.go |
new — the value-survives proof + the reportHasOffsite equivalence table |
hub/internal/monitor/offsite_delivery.go |
shapeDeclared outranks both inferred shapes; maybeHeal stands down with a record |
hub/internal/monitor/offsite_declared_test.go |
new — the demo-hp shape + the cross-package string pin |
hub/cmd/hub/main.go |
wires + runs the reconciler (OFFSITEHEAL_ONLY_CUSTOMER for a supervised rollout) |
.githooks/pre-push |
the workspace-root assertion (rider) |
manifests/hub.yaml |
image 0.95.0 → 0.96.0 |
Commits on main: f62a115 (hub half) · fe1e816 (rider) · b2462b8 (CHANGELOG) ·
4114c5f (manifest). Deployed: felhom-hub:0.96.0, ArgoCD Synced/Healthy, deploy/hub
rolled out, startup log carries offsite credential self-heal reconciler started (interval 5m, debounce 2 reports).
7. Tests and red-proofs
Green gate: go build ./... && go vet ./... && go test ./... in hub/ — rc=0.
python3 scripts/repo_gates.py --fast — all five gates OK.
| Scenario | Test | Result | Red-proof — what was mutated | Outcome |
|---|---|---|---|---|
| A+D | TestScenarioAD_DeclaringBoxIsRestagedNotMinted |
PASS | inverted heal() so Reissue runs before Restage |
FAILED — "a sustained declaration was not served: restages=0" |
| B | TestScenarioB_SilentBoxIsNeverActedOn |
PASS | — | — |
| C | TestScenarioC_HealthyBoxIsAPureNoOp |
PASS | — | — |
| E | TestScenarioE_NothingToRestageEscalatesOnce |
PASS | — | — |
| F | TestScenarioF_BlipIsAbsorbedByTheDebounce |
PASS | debounceReportsDefault 2 → 1 |
FAILED — "a one-report blip triggered a credential action" (and Scenario A+D also failed, "acted on a SINGLE declaration") |
| H | TestMainWiresTheOffsiteHealReconciler |
PASS | (AST; comments dropped) | — |
| — | TestRestageOneTimeSecret_ReArmsTheSameValue |
PASS | — | — |
| — | TestReportHasOffsite_EnabledOnly |
PASS | — | — |
| — | TestShapeOf_DeclarationOutranksBothInferredShapes |
PASS | — | — |
| — | TestDisabledDescriptorIsNeverHealed, TestBlockedCustomerIsNeverHealed, TestRestageErrorDoesNotEscalate |
PASS | — | — |
A test that first passed for the wrong reason, caught and fixed. TestBlockedCustomerIsNeverHealed
initially seeded Status: "blocked" through SaveCustomerConfig, whose INSERT does not carry the
column — so the customer was never actually blocked and the assertion would have been vacuous. It now
goes through SetCustomerConfigStatus and asserts IsCustomerBlocked before proceeding.
8. Part 4 — HALTED AT ITS STOP. Nothing was deleted.
§8.7: "If the paths do not match R-193's record exactly, STOP. A near-match on a protected endpoint is not a match." They do not match.
Measured read-only over SFTP, using each box's own credential, from inside its guest:
| Customer | Path | Size | What it is |
|---|---|---|---|
demo-felhom (u629488-sub1) |
/home/felhom-repo |
1.2 G | LIVE — the configured repo_path. Unopenable by the box (R-193), but NOT a set-aside store |
| demo-felhom | /home/felhom-repo.orphaned-20260717 |
1.4 G | set aside |
| demo-felhom | /home/felhom-repo.orphaned-20260718 |
3.0 M | set aside |
demo-hp (u629488-sub3) |
/home/felhom-repo |
582 K | LIVE |
| demo-hp | /home/felhom-repo.orphaned-20260804 |
43 M | set aside |
The ruling says "~1.2 GB across the two demo boxes, in set-aside stores". Reality: three set-aside
stores totalling ~1.45 GB — and the figure that matches ~1.2 GB is demo-felhom's LIVE
felhom-repo. Had the size been used to identify the target, the live repository would have been
deleted. Filed as R-212, WAITING-ON-OPERATOR.
9. Live validation
| # | What | Observable |
|---|---|---|
| 1 | A healthy box: no action, no events | demo-hp reported healthy throughout (enabled:true, no state key — report id 16742). Zero offsite_selfheal_* rows in the hub DB. Honest limit: the reconciler is silent by design on a healthy sweep, so there is no per-tick positive observable; what I have is the startup line proving Run was entered, the DB showing no events, and Scenario C. |
| 2 | The declared state, produced live without wiping a box | demo-felhom 9201 arranged reversibly into the stranded shape (settings + offbox/ backed up first; the offbox key removed and repo_password moved aside). Report id=16743 reached the hub carrying {enabled:false, state:"needs_credential", quota_gb:0, repo_size_bytes:0}. Restored the same minute; report id=16744 is healthy again. The single declaration was absorbed by the debounce — no self-heal event fired — which is Scenario F demonstrated on live infrastructure rather than in a fake. |
| 2b | The ACK field is no longer discarded | Both demo boxes' settings.json now carry hub_escrow_identity_present = true — the recorder working on a HEALTHY box, which is the case that used to return early. |
| 3 | A push from outside the workspace is refused | A scratch clone at /tmp/.../outside-clone: pre-push: PUSH REFUSED - this clone is OUTSIDE the felhom workspace, naming /mnt/5_hdd/felhom.eu, before the gates run. Red-proof: with the assertion removed the same push succeeded (rc=0, new branch on a throwaway bare remote). In-workspace pushes ran normally all session. |
| 4 | Part 4 | Not performed — halted at the STOP. Listings above. |
Not fired live: the hub actually re-staging a credential. Doing so would have re-applied demo-felhom's off-site target mid-session and changed the very state Part 4's listing describes. It is proven by Scenarios A/D/E with the store primitive proven separately against a real SQLite database.
Teardown: the scratch clone and throwaway remote are removed; demo-felhom's settings.json and
offbox/ restored from the backup taken first (verified: enabled=True, escrow_state=escrowed,
repo_password present=True); the backup copy remains at /root/r204-backup on the guest for
traceability. Nothing else was provisioned.
10. Registers
- R-204 — ALL FOUR ITEMS CLOSED. Both 2026-08-05 rulings recorded on the row: the declared-state trigger with its four-meanings-of-absence reasoning, and that the recovery preview's dashboard-password exposure is metadata, not content, and accepted.
- R-193 — credential half CLOSED. Still open under this ID: the customer-facing recovery screen, and the ciphertext deletion (now R-212).
- R-192 — CLOSED, guard half by REPLACEMENT. The counting inference is outranked by the declaration.
- R-202 — untouched, still open.
- R-212 — NEW (R-211 was the highest; grepped): the halted deletion, with the full measured listing.
11. CI
felhom-agent 0404f60, app-catalog ee2c810, felhom-controller 992803c, felhom.eu
4114c5f — run IDs and conclusions confirmed in the controller's REPORT.md §11 and re-checked at
session end. --no-verify was NOT used; every push ran the pre-push gate, including the new
workspace-root assertion.
12. Observations — noticed, NOT acted on
- A second, accidental barrier exists outside the workspace and should not be relied on:
felhom-agent'sreuse-refsgate FAILS in a clone outside the workspace because the sharedreuse_refs_check.pylives in thefelhom.eusibling. That is why the first Scenario-G red-proof had to be redone withapp-catalog-felhom.eu, whose fast gate is self-contained. It is incidental, repo-specific and not a substitute for the assertion. offsite_credential_restaged(the R-71c event) has still never fired for any customer. The new reconciler emits its ownoffsite_selfheal_*events, so the old one may now be permanently dead — worth a deliberate look rather than leaving two event families for one concern.- demo-felhom's live
felhom-repois 1.2 GB of ciphertext its own box cannot open (0 snapshots,last_status: error). It is not a set-aside store, so it is out of R-212's scope as filed — but it is the largest single block of unrecoverable data in the fleet and nothing currently plans its disposal.