docs: v0.72.0 ship report + live validation, CONTEXT rulings, DIAG-f10 R-70/R-71 status annotations
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01NKSN3gSg4TKVBBqkwW2djR
This commit is contained in:
@@ -1,63 +1,93 @@
|
||||
# REPORT — F10 diagnostic: demo-hp offsite "enabled at the hub, absent on the box" (2026-07-23)
|
||||
# REPORT — hub v0.72.0: R-70 + R-71(c) — the offsite last mile becomes visible, burned credentials heal themselves (2026-07-23)
|
||||
|
||||
**Task:** F10 diagnostic spec (project Claude, 2026-07-23) — diagnose first, repair only via the
|
||||
designed path, prove the tier. **No code changed in any repo.** Full evidence record:
|
||||
`documentation/audits/DIAG-f10-demo-hp-offsite-2026-07-23.md`.
|
||||
**Spec:** R-70 + R-71(c) prompt (project Claude, 2026-07-23). Companion controller leg: v0.161.0
|
||||
(see `felhom-controller/REPORT.md`). Origin: `documentation/audits/DIAG-f10-demo-hp-offsite-2026-07-23.md`.
|
||||
R-71(a) (day-0 ordering) untouched — separate upcoming spec.
|
||||
|
||||
## Phase-0 verdict
|
||||
## Baselines
|
||||
|
||||
Neither of the spec's two candidate shapes. The evidence (hub DB + box state + live logs, every
|
||||
claim cited in the DIAG) proves a third: **the day-0 managed floor-update (0.153.0→0.156.0,
|
||||
07-21 16:28:17Z) killed the offsite apply-bridge ~35 s after it consumed the one-time password**
|
||||
(16:27:42Z), before key-install/persist. Consume-then-persist + retry-only-on-restart
|
||||
(`offsiteapply.go:106–187`) ⇒ the credential was burned, no key was ever installed (so the
|
||||
key-auth-first recovery path could never engage), and every later start logged the consume-404 WARN
|
||||
and gave up. 153 reports over 2 days never carried an offbox object; the hub's "Provisioned…"
|
||||
line is static copy that reads neither `consumed_at` nor the reports.
|
||||
| Repo | start | shipped |
|
||||
|---|---|---|
|
||||
| felhom.eu | `c801cee6`, hub 0.71.0 | hub **0.72.0** (code + separate manifest chore commit), ArgoCD Synced/Healthy |
|
||||
| felhom-controller | `0eba37d5`, v0.160.0 | **v0.161.0** (`ce85314`), deployed BOTH boxes, healthy |
|
||||
|
||||
- Shape B ruled out from source: managed offsite is fully automatic; the box's „Távoli mentési cél
|
||||
beállítása" button is the BYO NAS/SFTP form only (`offbox_handlers.go:44–126`).
|
||||
- Strictly this was the spec's "consumed but persist failed → STOP" class; since the mechanism
|
||||
provably held its fail-safe and the source itself designates the recovery ("the password is
|
||||
spent; reset it on the hub to retry" = the offsite Re-issue), the operator ruled in-session:
|
||||
proceed on the Re-issue path.
|
||||
## What shipped (one detector, four consumers)
|
||||
|
||||
## Repair (designed path only)
|
||||
1. **Detector** — `internal/offsite/delivery.go` `DeliveryStateFor`: `applied` /
|
||||
`consumed_awaiting_apply` / `staged_awaiting_consume` / `no_secret` from the secret-row
|
||||
timestamps × report offsite-presence. Applied wins (the box's own report is the strongest
|
||||
evidence); applied+unconsumed-staged (demo-felhom) = `applied` + `StaleStagedSince` flag.
|
||||
New store reads (`GetOneTimeSecretInfo` — timestamps only, value never selected;
|
||||
`LatestReportOffsitePresence`; `CountReportsOffsiteSince`; `LastEventAt`) + the PBSDR-style
|
||||
test back-dater.
|
||||
2. **Customer card** — `deliveryViewFor` + `config_form_body.html`: the static "delivered to the
|
||||
controller once" claim is deleted; the card renders state + age (badge `n-ok/n-warn/n-neutral`,
|
||||
consumed goes amber past 30 min, stale-staged info line). Render test per branch.
|
||||
3. **Loud event** — `offsite_delivery_stuck` (WARNING) at ≥ 1 h of consumed_awaiting_apply;
|
||||
24 h/customer cooldown, durable via the events table (restart-proof).
|
||||
4. **Self-heal (R-71c)** — `monitor.OffsiteDeliveryChecker` on the shared 60 s ticker invokes
|
||||
`web.Server.ReissueOffsiteForCustomer` behind the narrow `monitor.OffsiteReissuer` interface
|
||||
(pbsdrheal precedent; armed only when the provisioner exists — else a restaged event would
|
||||
lie about a silent no-op). Trigger: consumed ≥ 1 h + ≥ 4 consecutive offbox-less reports +
|
||||
zero offbox evidence since consume. One restage/customer/24 h (durable); every firing emits
|
||||
`offsite_credential_restaged` (WARNING). **R-39(a) guard**: the heal re-reads the secret row
|
||||
immediately before acting and refuses over an unconsumed row — the store's
|
||||
`SaveOneTimeSecret` clobber semantics are untouched (Re-issue depends on supersede; the guard
|
||||
lives in the caller, exactly as specced).
|
||||
|
||||
- Operator clicked **Re-issue offsite credentials** ONCE (R-31 click-once discipline; pre-verified
|
||||
side-effect-free: no escrow blob existed, the box never held the old password, sub3 was empty).
|
||||
Click 09:53:37Z → box consumed 09:53:41 → `offsite configured … (pending key escrow)` 09:53:45.
|
||||
**8 seconds click-to-converged.**
|
||||
- Escrow ceremony run by the operator through the real `/backup/escrow` wizard (one-shot R on the
|
||||
operator's screen only): blob stored 10:01:17 (zero_knowledge, pw-hash recorded), hub-verified
|
||||
auto-confirm 10:01:24 → `escrowed; offsite runs enabled`.
|
||||
## Red-proofs (run, observed, restored — verbatim failures)
|
||||
|
||||
## Tier proof (F10 closure bar)
|
||||
1. **THE CLOBBER RED-PROOF** — R-39(a) guard block removed from `maybeHeal`; the TOCTOU test
|
||||
(operator Re-issue staged mid-tick via the onEvent hook) failed with:
|
||||
`reissue calls = 1, want 0 — the R-39(a) guard must refuse over an unconsumed secret`
|
||||
— i.e. the operator's fresh unconsumed secret would have been clobbered
|
||||
(the fake reissuer mimics the production `SaveOneTimeSecret` side effect, so the clobber is
|
||||
observed on the row, not inferred). Guard restored → green.
|
||||
2. **Heal rate-limit** — `LastEventAt`/`healCooldown` check removed; failed with:
|
||||
`reissue calls after recurrence = 2, want STILL 1 (one restage per customer per 24h)`. Restored.
|
||||
3. **Stuck-event cooldown** — cooldown check removed; failed with:
|
||||
`stuck events = 2, want exactly 1 (24h per-customer cooldown)`. Restored.
|
||||
|
||||
paperless-ngx toggled into offsite scope via the real endpoint (per-app default is OFF). Then, all
|
||||
via real endpoints from inside guest 9201 (endpoint-level method; no browser on DooPlex):
|
||||
probe (md5 `9120e65d6a9f071072d827fc404dc840`) in the mandatory `appdata/paperless/media` →
|
||||
**first offsite run**: repo initialized fresh on sub3, 79.8 MB / 49 files, 1m19s, ok →
|
||||
probe deleted → **`mode=full` restore** (size gate 79.8 MB → confirm): snapshot **`2bf7f2e1`** to
|
||||
staging, staging md5-identical → **place**: `1 file(s) merged (missing-only)`, live md5-identical.
|
||||
Cleanup: probe removed, second run (2m17s ok) leaves the latest snapshot probe-free (retention
|
||||
pruned the probe-bearing one); zero residue on box/repo; break-glass + DB copies shredded.
|
||||
Hub now reports demo-hp `offsite: enabled/escrowed/quota 50`; nightly run scheduled (04:15 UTC).
|
||||
Plus: the full-Check demo-felhom-shape test (applied + stale staged → zero events, zero calls,
|
||||
row byte-untouched), evidence gates (<4 reports → no heal but stuck event still fires; mixed
|
||||
offbox history → no heal), young-consumed silence, disabled/blocked skip, dispatcher severity
|
||||
tests (warning routes operator-only; an info variant would be silent — pinned beside the v0.71.0
|
||||
guard, `severityNotifies` untouched), 5 card render tests + `deliveryViewFor` amber derivation.
|
||||
Green gates: hub 17 packages ok + `hub_confirm_gate.py`; controller 25 packages ok + all template
|
||||
gates (the pre-existing R-29 `docker_run_volume_path_gate` red noted, untouched).
|
||||
|
||||
## Product findings + docs
|
||||
## Live validation (read-only, both fixtures intact)
|
||||
|
||||
- **R-70 (P2-HIGH)** minted: the offsite last mile is invisible on both surfaces (hub can't tell
|
||||
staged/consumed/applied; box shows the generic empty state). Coupled to R-31's status-card idiom
|
||||
and the R-39 consumed_at honesty-gauge precedent.
|
||||
- **R-71 (P1)** minted: the race itself — recurs structurally on every fresh onboarding whose ISO
|
||||
floor lags the managed floor. Spec-first directions listed in the row (ordering / two-phase
|
||||
consume / hub-side auto-restage with the R-39(a) mint-race guard).
|
||||
- Audit F10 row annotated: **offsite leg resolved**; PBS-DR half explicitly stays open (F13 +
|
||||
DR ceremony R-moment). CONTEXT.md updated.
|
||||
- **Checker silence:** 10 min of 60 s ticks on the live fleet → **0** `offsite-delivery` log
|
||||
lines, **0** detector events in the DB — both fixtures are healthy and the detector agrees.
|
||||
- **Fixture states from live data** (fresh DB copy, shredded after):
|
||||
demo-hp `latest_report_offsite=True, secret consumed 09:53:41` → **applied**;
|
||||
demo-felhom `latest_report_offsite=True, secret_row=(2026-07-21 08:29:29, None)` →
|
||||
**applied + stale-staged since 07-21** — the live specimen SURVIVED the train untouched
|
||||
(`consumed_at` still NULL, created_at unchanged); peti-felhom → applied.
|
||||
- **demo-hp controller page** (authed endpoint fetch inside the guest, ASCII-safe greps):
|
||||
banner count 0, `felhom-offsite-card` id count 0 (configured + 1 toggled app → no card at all),
|
||||
configured markers present. v0.161.0 healthy on both boxes.
|
||||
- **Method:** endpoint-level + DB-input derivation (no browser on DooPlex). The rendered card is
|
||||
pinned by render tests; the operator's 10-second residual: demo-hp Edit page shows
|
||||
`applied`, demo-felhom shows `applied` + the stale-staged note (since 2026-07-21).
|
||||
- **The self-heal is NOT live-fired** — no broken box exists and none was broken for it (F9
|
||||
rule). It ships unit-proven + red-proofed, PARTIAL/IMPLEMENTED on the ROADMAP with the
|
||||
explicit "fires on next natural occurrence or a staged drill" note — never PROVEN-LIVE.
|
||||
The controller banner leg is likewise unit-proven/live-pending (no box occupies the
|
||||
enabled+no-offbox window; the next fresh onboarding is its natural live leg).
|
||||
|
||||
## Rulings recorded (CONTEXT.md)
|
||||
|
||||
State precedence (applied wins; stale-staged is a flag, never a downgrade); durable cooldowns via
|
||||
the events table (restart-proof by design); both detector events operator-only (no
|
||||
customerMessages entry, not in allowedEventTypes — the pbsdr_* precedent) until the mechanism has
|
||||
history; heal disabled without a provisioner.
|
||||
|
||||
## Observed, not acted on
|
||||
|
||||
- demo-felhom's 07-21 staged offsite secret is still unconsumed (residue of the mistaken R-39-day
|
||||
offsite Re-issue; box recovered via key-auth-first, which never consumes). Harmless; supports R-70.
|
||||
- The 3 dead unclaimed-appliance records from the ISO train remain for operator discard.
|
||||
- Hub pod log only reaches back to 07-22 20:58Z (restart); the 07-21 correlation came from the DB.
|
||||
- R-29: `docker_run_volume_path_gate` red on `appexport/estimate.go:179` (pre-existing, has its
|
||||
own row — the 3-line allowlist fix remains undone by design of this train's scope).
|
||||
- The `szolg` single-hit on the demo-hp page grep is an accent-truncated unrelated word (the
|
||||
ASCII-only-grep trap documented in felhom-controller/CLAUDE.md — verified benign via the
|
||||
card-id count of 0).
|
||||
|
||||
Reference in New Issue
Block a user