docs: v0.72.0 ship report + live validation, CONTEXT rulings, DIAG-f10 R-70/R-71 status annotations

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01NKSN3gSg4TKVBBqkwW2djR
This commit is contained in:
2026-07-23 13:05:03 +02:00
parent 92681bda6c
commit 527d81cf70
3 changed files with 105 additions and 50 deletions
+80 -50
View File
@@ -1,63 +1,93 @@
# REPORT — F10 diagnostic: demo-hp offsite "enabled at the hub, absent on the box" (2026-07-23)
# REPORT — hub v0.72.0: R-70 + R-71(c) — the offsite last mile becomes visible, burned credentials heal themselves (2026-07-23)
**Task:** F10 diagnostic spec (project Claude, 2026-07-23) — diagnose first, repair only via the
designed path, prove the tier. **No code changed in any repo.** Full evidence record:
`documentation/audits/DIAG-f10-demo-hp-offsite-2026-07-23.md`.
**Spec:** R-70 + R-71(c) prompt (project Claude, 2026-07-23). Companion controller leg: v0.161.0
(see `felhom-controller/REPORT.md`). Origin: `documentation/audits/DIAG-f10-demo-hp-offsite-2026-07-23.md`.
R-71(a) (day-0 ordering) untouched — separate upcoming spec.
## Phase-0 verdict
## Baselines
Neither of the spec's two candidate shapes. The evidence (hub DB + box state + live logs, every
claim cited in the DIAG) proves a third: **the day-0 managed floor-update (0.153.0→0.156.0,
07-21 16:28:17Z) killed the offsite apply-bridge ~35 s after it consumed the one-time password**
(16:27:42Z), before key-install/persist. Consume-then-persist + retry-only-on-restart
(`offsiteapply.go:106187`) ⇒ the credential was burned, no key was ever installed (so the
key-auth-first recovery path could never engage), and every later start logged the consume-404 WARN
and gave up. 153 reports over 2 days never carried an offbox object; the hub's "Provisioned…"
line is static copy that reads neither `consumed_at` nor the reports.
| Repo | start | shipped |
|---|---|---|
| felhom.eu | `c801cee6`, hub 0.71.0 | hub **0.72.0** (code + separate manifest chore commit), ArgoCD Synced/Healthy |
| felhom-controller | `0eba37d5`, v0.160.0 | **v0.161.0** (`ce85314`), deployed BOTH boxes, healthy |
- Shape B ruled out from source: managed offsite is fully automatic; the box's „Távoli mentési cél
beállítása" button is the BYO NAS/SFTP form only (`offbox_handlers.go:44126`).
- Strictly this was the spec's "consumed but persist failed → STOP" class; since the mechanism
provably held its fail-safe and the source itself designates the recovery ("the password is
spent; reset it on the hub to retry" = the offsite Re-issue), the operator ruled in-session:
proceed on the Re-issue path.
## What shipped (one detector, four consumers)
## Repair (designed path only)
1. **Detector**`internal/offsite/delivery.go` `DeliveryStateFor`: `applied` /
`consumed_awaiting_apply` / `staged_awaiting_consume` / `no_secret` from the secret-row
timestamps × report offsite-presence. Applied wins (the box's own report is the strongest
evidence); applied+unconsumed-staged (demo-felhom) = `applied` + `StaleStagedSince` flag.
New store reads (`GetOneTimeSecretInfo` — timestamps only, value never selected;
`LatestReportOffsitePresence`; `CountReportsOffsiteSince`; `LastEventAt`) + the PBSDR-style
test back-dater.
2. **Customer card**`deliveryViewFor` + `config_form_body.html`: the static "delivered to the
controller once" claim is deleted; the card renders state + age (badge `n-ok/n-warn/n-neutral`,
consumed goes amber past 30 min, stale-staged info line). Render test per branch.
3. **Loud event**`offsite_delivery_stuck` (WARNING) at ≥ 1 h of consumed_awaiting_apply;
24 h/customer cooldown, durable via the events table (restart-proof).
4. **Self-heal (R-71c)**`monitor.OffsiteDeliveryChecker` on the shared 60 s ticker invokes
`web.Server.ReissueOffsiteForCustomer` behind the narrow `monitor.OffsiteReissuer` interface
(pbsdrheal precedent; armed only when the provisioner exists — else a restaged event would
lie about a silent no-op). Trigger: consumed ≥ 1 h + ≥ 4 consecutive offbox-less reports +
zero offbox evidence since consume. One restage/customer/24 h (durable); every firing emits
`offsite_credential_restaged` (WARNING). **R-39(a) guard**: the heal re-reads the secret row
immediately before acting and refuses over an unconsumed row — the store's
`SaveOneTimeSecret` clobber semantics are untouched (Re-issue depends on supersede; the guard
lives in the caller, exactly as specced).
- Operator clicked **Re-issue offsite credentials** ONCE (R-31 click-once discipline; pre-verified
side-effect-free: no escrow blob existed, the box never held the old password, sub3 was empty).
Click 09:53:37Z → box consumed 09:53:41 → `offsite configured … (pending key escrow)` 09:53:45.
**8 seconds click-to-converged.**
- Escrow ceremony run by the operator through the real `/backup/escrow` wizard (one-shot R on the
operator's screen only): blob stored 10:01:17 (zero_knowledge, pw-hash recorded), hub-verified
auto-confirm 10:01:24 → `escrowed; offsite runs enabled`.
## Red-proofs (run, observed, restored — verbatim failures)
## Tier proof (F10 closure bar)
1. **THE CLOBBER RED-PROOF** — R-39(a) guard block removed from `maybeHeal`; the TOCTOU test
(operator Re-issue staged mid-tick via the onEvent hook) failed with:
`reissue calls = 1, want 0 — the R-39(a) guard must refuse over an unconsumed secret`
— i.e. the operator's fresh unconsumed secret would have been clobbered
(the fake reissuer mimics the production `SaveOneTimeSecret` side effect, so the clobber is
observed on the row, not inferred). Guard restored → green.
2. **Heal rate-limit**`LastEventAt`/`healCooldown` check removed; failed with:
`reissue calls after recurrence = 2, want STILL 1 (one restage per customer per 24h)`. Restored.
3. **Stuck-event cooldown** — cooldown check removed; failed with:
`stuck events = 2, want exactly 1 (24h per-customer cooldown)`. Restored.
paperless-ngx toggled into offsite scope via the real endpoint (per-app default is OFF). Then, all
via real endpoints from inside guest 9201 (endpoint-level method; no browser on DooPlex):
probe (md5 `9120e65d6a9f071072d827fc404dc840`) in the mandatory `appdata/paperless/media`
**first offsite run**: repo initialized fresh on sub3, 79.8 MB / 49 files, 1m19s, ok →
probe deleted → **`mode=full` restore** (size gate 79.8 MB → confirm): snapshot **`2bf7f2e1`** to
staging, staging md5-identical → **place**: `1 file(s) merged (missing-only)`, live md5-identical.
Cleanup: probe removed, second run (2m17s ok) leaves the latest snapshot probe-free (retention
pruned the probe-bearing one); zero residue on box/repo; break-glass + DB copies shredded.
Hub now reports demo-hp `offsite: enabled/escrowed/quota 50`; nightly run scheduled (04:15 UTC).
Plus: the full-Check demo-felhom-shape test (applied + stale staged → zero events, zero calls,
row byte-untouched), evidence gates (<4 reports → no heal but stuck event still fires; mixed
offbox history → no heal), young-consumed silence, disabled/blocked skip, dispatcher severity
tests (warning routes operator-only; an info variant would be silent — pinned beside the v0.71.0
guard, `severityNotifies` untouched), 5 card render tests + `deliveryViewFor` amber derivation.
Green gates: hub 17 packages ok + `hub_confirm_gate.py`; controller 25 packages ok + all template
gates (the pre-existing R-29 `docker_run_volume_path_gate` red noted, untouched).
## Product findings + docs
## Live validation (read-only, both fixtures intact)
- **R-70 (P2-HIGH)** minted: the offsite last mile is invisible on both surfaces (hub can't tell
staged/consumed/applied; box shows the generic empty state). Coupled to R-31's status-card idiom
and the R-39 consumed_at honesty-gauge precedent.
- **R-71 (P1)** minted: the race itself — recurs structurally on every fresh onboarding whose ISO
floor lags the managed floor. Spec-first directions listed in the row (ordering / two-phase
consume / hub-side auto-restage with the R-39(a) mint-race guard).
- Audit F10 row annotated: **offsite leg resolved**; PBS-DR half explicitly stays open (F13 +
DR ceremony R-moment). CONTEXT.md updated.
- **Checker silence:** 10 min of 60 s ticks on the live fleet → **0** `offsite-delivery` log
lines, **0** detector events in the DB — both fixtures are healthy and the detector agrees.
- **Fixture states from live data** (fresh DB copy, shredded after):
demo-hp `latest_report_offsite=True, secret consumed 09:53:41`**applied**;
demo-felhom `latest_report_offsite=True, secret_row=(2026-07-21 08:29:29, None)`
**applied + stale-staged since 07-21** — the live specimen SURVIVED the train untouched
(`consumed_at` still NULL, created_at unchanged); peti-felhom → applied.
- **demo-hp controller page** (authed endpoint fetch inside the guest, ASCII-safe greps):
banner count 0, `felhom-offsite-card` id count 0 (configured + 1 toggled app → no card at all),
configured markers present. v0.161.0 healthy on both boxes.
- **Method:** endpoint-level + DB-input derivation (no browser on DooPlex). The rendered card is
pinned by render tests; the operator's 10-second residual: demo-hp Edit page shows
`applied`, demo-felhom shows `applied` + the stale-staged note (since 2026-07-21).
- **The self-heal is NOT live-fired** — no broken box exists and none was broken for it (F9
rule). It ships unit-proven + red-proofed, PARTIAL/IMPLEMENTED on the ROADMAP with the
explicit "fires on next natural occurrence or a staged drill" note — never PROVEN-LIVE.
The controller banner leg is likewise unit-proven/live-pending (no box occupies the
enabled+no-offbox window; the next fresh onboarding is its natural live leg).
## Rulings recorded (CONTEXT.md)
State precedence (applied wins; stale-staged is a flag, never a downgrade); durable cooldowns via
the events table (restart-proof by design); both detector events operator-only (no
customerMessages entry, not in allowedEventTypes — the pbsdr_* precedent) until the mechanism has
history; heal disabled without a provisioner.
## Observed, not acted on
- demo-felhom's 07-21 staged offsite secret is still unconsumed (residue of the mistaken R-39-day
offsite Re-issue; box recovered via key-auth-first, which never consumes). Harmless; supports R-70.
- The 3 dead unclaimed-appliance records from the ISO train remain for operator discard.
- Hub pod log only reaches back to 07-22 20:58Z (restart); the 07-21 correlation came from the DB.
- R-29: `docker_run_volume_path_gate` red on `appexport/estimate.go:179` (pre-existing, has its
own row — the 3-line allowlist fix remains undone by design of this train's scope).
- The `szolg` single-hit on the demo-hp page grep is an accent-truncated unrelated word (the
ASCII-only-grep trap documented in felhom-controller/CLAUDE.md — verified benign via the
card-id count of 0).