5.2 KiB
felhom.eu — task reports
Overwrite this file with a summary of the most recent task only (uniform with the other repos; not cumulative). The cumulative hub history lives in hub/CHANGELOG.md; the scripts history lives in scripts/CHANGELOG.md.
Direction-2 immediate-sync — hub→box wait channel (hub v0.58.0) — 2026-07-16
Implements option (b) of documentation/audits/SPIKE-immediate-sync-transport-2026-07-16.md. An
operator action on the hub now reaches the box in seconds. Pairs with controller v0.140.0
(the long-poll client). Baselines: felhom.eu 10e07f57 → hub v0.58.0 (deployed); controller
f665bbed → v0.140.0.
What shipped (hub)
internal/intent— in-memory operator-intent notifier. Per-customer generation counter + waiter registry:Bumpadvances + wakes (coalesce to LATEST — counter, not queue);Waitreturns on gen-change / ctx / maxHold / Close, with a pre-register gen-check closing the bump-before-connect race;Closewakes all. In-memory BY DESIGN (a restart resets generations → the box's!=compare costs exactly one harmless full-state report, never a storm). No persistence, no schema.GET /api/v1/wait(api/wait.go).checkAuthCustomer; per-customer only (global key → 400; customer resolved from the key, nocustomer_idparam → A can't observe B). 240 s hold with a 25 s heartbeat newline — nginx'sproxy_read_timeoutis measured between upstream reads, so the heartbeat keeps the default 60 s from firing → no ingress annotation / manifest change. Contentless{"gen":N}completion; per-connection write-deadline viahttp.NewResponseController(globalhttp.Servertimeouts untouched).- Intent bumps (web), fire-after-commit. config create/update/delete, claim resend, offsite re-issue (UI + re-enroll seam), offsite freeze/unfreeze, retrieval-password regen, block/unblock, per-customer floor, global floor (bumps every config-managed customer), controller log-tail request, controller log-bundle request (the AGENT ring rides the heartbeat envelope — deliberately not bumped).
- Wiring + shutdown. One
intent.New()shared by web + API;intentHub.Close()beforeserver.Shutdown. nil-safe (unset hub → wait 503, bumps no-op).
Tests + red-proofs (all run-fail-reverted)
- intent: coalesce-to-latest (return the as-of-register snapshot →
TestWait_CoalescesBurstToLatestGenfails) and race-closer (drop the pre-register check →TestWait_RaceCloser_BumpBeforeWaitNotLostfails). Plus timeout/ctx/close/isolation/stress. api/wait: 401/400/503, completes-on-bump, timeout-with-heartbeats, race-closer, customer isolation. Full hub suite green (go build/vet/test ./...); intent race-tested on the build server.
Deploy (GitOps)
Image felhom-hub:0.58.0 built + pushed on 180; manifests/hub.yaml bumped 0.57.0→0.58.0 (commit
59eb1f8); ArgoCD hard-refresh + patch-sync → Synced/Healthy, rollout complete, live image
:0.58.0, listening. Route smoke: /api/v1/wait → 401 (no/bad auth), 404 (wrong method).
Live validation (demo arc, customer demo-felhom)
- Completion-on-difference path (the code path a bump triggers): a box-side curl with a
mismatched
?gen=999999returned{"gen":0}in 0.047 s through the real ingress — sub-100 ms wake delivery, DNS-only nginx path. - Hold proof / no-annotation claim: a box-side curl at
?gen=0(matching the current generation) held cleanly for 240.04 s (HTTP 200 →{"gen":0}, 10 heartbeat newlines observed) — surviving well past nginx's 60 s default with the 25 s heartbeat, no ingress annotation applied. The controller's own Waiter logged zero wait-channel errors over the observation window (it holds + reconnects silently), a second confirmation. No-annotation claim: HELD; the ingress fallback was NOT needed. - Scenario C — hub restart resilience (live):
kubectl rollout restart deploy/hubmid-hold → the box's Waiter logged exactly one WARN (wait channel error: unexpected EOF — backing off), reconnected cleanly, and fired 0 out-of-cycle reports (the generation was 0 before and after, so there was correctly nothing to deliver — no restart storm). The immediate-probe endpoint returned 200 the moment the pod was ready. (Cosmetic: the "recovered" INFO is deferred until the reconnected 240 s hold completes, sincepollOnceblocks — the reconnect itself is immediate.) The "exactly one fire" restart case needs a prior non-zero generation, i.e. an operator bump. - Round-trip save→applied via the operator UI requires an operator login (CC lacks the operator
password by standing constraint) — the mechanism is proven by the unit tests + the live
completion-path proof, and the delivery chain (report ACK → config_version →
ConfigRefresherre-pull + self-restart) is the already-live v0.139.0 / v0.26.0 path (unchanged).
Not changed / parked
Report ACK schema, the 15-min cycle, the 900 s agent poll, and the global http.Server timeouts are
byte-identical. The agent-plane ep0-relayed poke stays PARKED in the OOB / mutual-repair arc
(Viktor's ruling + spike P4).