From 9dc26459eaccaf58c09b4494039907c3b10b30fe Mon Sep 17 00:00:00 2001 From: kisfenyo Date: Thu, 6 Aug 2026 12:56:12 +0200 Subject: [PATCH] v0.203.0: the box collects what the hub staged for it (R-218 consume half) + R-220's message MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit R-218's declaration half shipped in v0.201.0 and works. Its consume half never existed. Reconcile ran exactly twice per process — at start-up and when the recovery screen drives it — and BOTH fire before the hub has anything staged, because the hub stages in RESPONSE to the declaration those runs precede. Measured on the R-201 re-walk: unlock reconcile 11:43:07, hub staged 11:44:57 saying 'next cycle', a full report cycle ran 11:55:46, still unconsumed at 12:06. A guest command line applied it in 18 seconds — everything correct except the trigger. Bridge.RetryIfDeclared re-runs the SAME reconcile on a 5-minute tick, driven from the box's own published declaration (OffboxReportStatus().State) — the very statement the hub acts on, so the two cannot disagree. Poll, not an ACK flag, decided on the promise: the no-target message says 'amint megvannak' (no deadline) and the card says 'within a day'. Five minutes is inside both by a wide margin and needs no hub change. It stops by construction — a healthy box does no work and logs nothing — and the settle gate is deliberately kept via ReconcileWhenSettled. The marker was investigated and left alone: applied_marker lives in the guest's DataDir, which a rebuild destroys, so it cannot suppress a legitimate re-run. R-220's customer half: the refusal no longer tells the customer to choose from a list that may be empty. It names the rebuild, points at the Meghajtók page, and promises no outcome. Red-proofs: remove the retry -> credential uncollected (the dead end reproduced); drop the stop condition -> a healthy box hammers the hub; call Reconcile instead of ReconcileWhenSettled -> settle gate bypassed; restore the old sentence -> the impossible action returns. 28 packages ok, vet clean, all controller gates OK. --- CHANGELOG.md | 51 +++++++++ controller/cmd/controller/main.go | 49 +++++++++ .../internal/offsiteapply/offsiteapply.go | 33 ++++++ .../internal/offsiteapply/retry_test.go | 101 ++++++++++++++++++ .../internal/settings/refuse_message_test.go | 40 +++++++ controller/internal/settings/settings.go | 17 ++- 6 files changed, 290 insertions(+), 1 deletion(-) create mode 100644 controller/internal/offsiteapply/retry_test.go create mode 100644 controller/internal/settings/refuse_message_test.go diff --git a/CHANGELOG.md b/CHANGELOG.md index fe35fcc..feabc14 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -1,3 +1,54 @@ +## v0.203.0 — the box collects what the hub staged for it (2026-08-06, R-218 consume half / R-220 message) — MinAgent 0.127.0 + +**R-218's declaration half shipped in v0.201.0 and works. Its consume half never existed.** + +A stranded box says `offsite.state=needs_credential`; the hub's `offsiteheal` re-stages the one-time +secret and logs *"the box re-consumes on its next cycle"*. **There was no next cycle.** `Reconcile` +ran exactly twice in a process's life — once at start-up (whose own comment said *"retries on next +config refresh/restart"*) and once when the recovery screen drives it (R-219) — and **both fire before +the hub has anything staged, because the hub stages in RESPONSE to the declaration those runs +precede.** So the hub held a credential the box would never fetch. + +**Measured on the R-201 re-walk, 2026-08-06:** unlock reconcile **11:43:07** · hub staged **11:44:57** +saying "next cycle" · a full report cycle ran **11:55:46** · **still unconsumed at 12:06**. A guest +command line applied it in **18 seconds** — proving the credential, the target and the key were all +correct and only the trigger was missing. That was the first of the two dead ends that kept the +recovery journey failing. + +**The fix: re-run the SAME reconcile, on a tick, for exactly as long as the box says it needs a +credential.** `Bridge.RetryIfDeclared` is driven from the box's own published declaration +(`OffboxReportStatus().State`) — **the very statement the hub acts on**, so the two can never disagree +about whether a retry is wanted. + +**Why a poll and not an ACK flag.** The deciding criterion was the promise the customer is given: the +no-target message says *"amint megvannak"* (no deadline) and the backups card says *"ha egy napon +belül nem áll be"* — **within a day**. A 5-minute tick is inside both by a wide margin and needs **no +hub change**. If either promise ever tightens to minutes, revisit. + +**It stops by construction.** The instant a target exists the declaration goes false: a healthy box +does no work and **logs nothing** (asserted). **The settle gate is deliberately kept** — the retry goes +through `ReconcileWhenSettled`, so the day-0 floor race it guards is unchanged. + +**The marker was investigated and left alone.** `applied_marker` lives at `/offbox/` — inside +the guest's data dir, which a rebuild destroys — so it cannot suppress a legitimate post-rebuild +re-run. It is not part of this defect. + +### R-220's customer-facing half — a refusal that named an impossible action + +The deploy refusal said *"Válasszon a listából csatlakoztatott meghajtót"* — choose an attached drive +from the list — **while the list was empty**, on a rebuilt box, for a reason the customer had no part +in. That is the I3 breach the campaign recorded. It now says what is true (the drive is not registered +**on this machine**, which is what a rebuild causes), points at the page where re-attaching happens +rather than at a possibly-empty list, and **promises no outcome**, because whether the drive can be +re-attached is not knowable from there. The NAS refusal is a different situation and is untouched. + +Tests: `internal/offsiteapply/retry_test.go` (a credential staged after start-up is collected; a +healthy box does nothing and logs nothing; a nil bridge is silent; the settle gate holds) and +`internal/settings/refuse_message_test.go`. **Red-proofs:** removing the retry leaves the credential +uncollected — the re-walk's dead end reproduced; dropping the stop condition makes a healthy box +hammer the hub; calling `Reconcile` instead of `ReconcileWhenSettled` bypasses the settle gate; +restoring the old sentence brings the impossible action back. + ## docs — the "CI is still owed" claim was stale; corrected (2026-08-06, R-229 part 2) — no version bump **One sentence, no code.** This file asserted that continuous integration was still owed diff --git a/controller/cmd/controller/main.go b/controller/cmd/controller/main.go index b9f2507..cc297f7 100644 --- a/controller/cmd/controller/main.go +++ b/controller/cmd/controller/main.go @@ -756,6 +756,55 @@ func main() { return err }) + // ── R-218, THE CONSUME HALF (v0.203.0) ──────────────────────────────────────────────── + // + // The declaration half shipped in v0.201.0 and works: a stranded box says + // `offsite.state=needs_credential`, and the hub's `offsiteheal` re-stages the one-time secret + // and logs *"the box re-consumes on its next cycle"*. + // + // **THERE WAS NO NEXT CYCLE.** `Reconcile` ran exactly twice in a process's life: once at + // start-up (the goroutine above, whose own comment says "retries on next config refresh/ + // restart") and once when the recovery screen drives it (R-219). Both fire BEFORE the hub has + // anything staged, because the hub only stages in response to the declaration those runs + // precede. + // + // Measured on the R-201 re-walk, 2026-08-06: unlock reconcile 11:43:07 · hub staged 11:44:57 + // saying "next cycle" · a full report cycle ran 11:55:46 · still unconsumed at 12:06. A guest + // command line moved it in **18 seconds**, which proves the credential, the target and the key + // were all fine and only the trigger was missing. + // + // So: re-run the SAME reconcile, on a tick, for exactly as long as the box itself says it needs + // a credential. The signal is the box's own `OffboxReportStatus().State` — the very statement + // the hub acts on, so the two can never disagree about whether a retry is wanted. + // + // WHY A POLL AND NOT AN ACK FLAG: the customer-facing text promises "amint megvannak" (no + // deadline) and the backups card promises "within a day"; a 5-minute tick is inside both by a + // wide margin, and it needs no hub change. If either promise ever tightens to minutes, revisit. + // + // IT STOPS BY CONSTRUCTION: the instant a target exists `OffboxReportStatus` stops returning + // the declared state, so a healthy box does no work and makes no noise. The settle gate is + // deliberately kept — `ReconcileWhenSettled` waits for floor knowledge exactly as at start-up. + sched.Every("offsite-credential-retry", 5*time.Minute, func(ctx context.Context) error { + if offsiteBridge == nil { + return nil // off-site not configured for this customer + } + attempted, err := offsiteBridge.RetryIfDeclared(ctx, func() bool { + st := backupMgr.OffboxReportStatus() + return st != nil && st.State == backup.OffsiteStateNeedsCredential + }) + if !attempted { + return nil // a target exists (or we never needed one) — no work, and no log line + } + if err != nil { + // NOT a job failure: the box still declares, so the next tick tries again. Logged at + // WARN because a credential that never arrives is exactly what this exists to surface. + logger.Printf("[WARN] [offsite-apply] credential retry: %v (the box still declares a need; retrying)", err) + return nil + } + logger.Printf("[INFO] [offsite-apply] credential retry: the staged credential was collected and the tier applied") + return nil + }) + // Cache refresh: every 5 minutes. Recompute the effective window each pass so the cached // "next DB dump" follows a runtime window change (the UI save also refreshes immediately). sched.Every("backup-cache", 5*time.Minute, func(ctx context.Context) error { diff --git a/controller/internal/offsiteapply/offsiteapply.go b/controller/internal/offsiteapply/offsiteapply.go index 1fb2724..e6d9ce0 100644 --- a/controller/internal/offsiteapply/offsiteapply.go +++ b/controller/internal/offsiteapply/offsiteapply.go @@ -341,3 +341,36 @@ func (b *Bridge) ReconcileWhenSettled(gateCtx context.Context) error { defer cancel() return b.Reconcile(ctx) } + +// ── R-218, THE CONSUME HALF ────────────────────────────────────────────────────────────────────── +// +// `Reconcile` was correct from the day it shipped and was simply never run again. It fires at +// start-up and once more when the recovery screen drives it (R-219) — and BOTH precede the moment the +// hub has anything staged, because the hub stages in RESPONSE to the declaration those runs come +// before. So the hub held a credential the box would never fetch. +// +// Measured on the R-201 re-walk, 2026-08-06: unlock reconcile 11:43:07 · hub staged 11:44:57 saying +// "the box re-consumes on its next cycle" · a full report cycle ran 11:55:46 · still unconsumed at +// 12:06. A guest command line moved it in 18 seconds — everything was fine except the trigger. + +// NeedsCredentialFunc reports whether the box STILL declares it needs a transport credential. It is +// deliberately the box's own published declaration (`backup.OffboxReportStatus().State`) rather than a +// second predicate: the hub acts on that statement, so driving the retry from anything else would let +// the two disagree about whether a retry is wanted. +type NeedsCredentialFunc func() bool + +// RetryIfDeclared is ONE tick of the consume half. +// +// It reconciles **only while the box declares a need**, which is what makes it stop: the instant a +// target exists the declaration goes false, this returns immediately, and a healthy box does no work +// and logs nothing. The settle gate is deliberately preserved — `ReconcileWhenSettled` waits for floor +// knowledge exactly as the start-up path does, because the day-0 race it guards is unchanged. +// +// Returns whether a reconcile was ATTEMPTED, so a caller (and a test) can tell "declined to run" from +// "ran and failed" without reading the log. +func (b *Bridge) RetryIfDeclared(ctx context.Context, declared NeedsCredentialFunc) (attempted bool, err error) { + if b == nil || declared == nil || !declared() { + return false, nil + } + return true, b.ReconcileWhenSettled(ctx) +} diff --git a/controller/internal/offsiteapply/retry_test.go b/controller/internal/offsiteapply/retry_test.go new file mode 100644 index 0000000..f229e95 --- /dev/null +++ b/controller/internal/offsiteapply/retry_test.go @@ -0,0 +1,101 @@ +package offsiteapply + +import ( + "context" + "strings" + "testing" +) + +// ── R-218's CONSUME HALF ───────────────────────────────────────────────────────────────────────── +// +// Measured on the R-201 re-walk, 2026-08-06: the hub staged a credential at 11:44:57 and logged +// "the box re-consumes on its next cycle"; a full report cycle ran at 11:55:46; at 12:06 it was still +// unconsumed, and a guest command line applied it in 18 seconds. Everything was correct except that +// nothing ever re-ran the reconcile. +// +// These assert the EFFECT — was a reconcile attempted, and did the tier get applied — not that a +// helper returned a bool. + +// ── SCENARIO A — a credential staged AFTER start-up is collected, unaided ──────────────────────── +// +// RED-PROOF: make RetryIfDeclared return (false, nil) unconditionally — i.e. remove the retry, which +// is the pre-v0.203.0 world — and this FAILS with the credential still sitting unconsumed. That is +// the re-walk's first dead end, reproduced. +func TestRetryIfDeclared_CollectsACredentialStagedAfterStartup(t *testing.T) { + b, cons, _, en, _ := newBridge(t, goodOffsite()) + + // The box declares: it was rebuilt, has no target, and the hub holds a package for it. + attempted, err := b.RetryIfDeclared(context.Background(), func() bool { return true }) + if err != nil { + t.Fatalf("retry: %v", err) + } + if !attempted { + t.Fatal("R-218 RETURNED: the box declared a need and no reconcile was attempted") + } + // EFFECT: the one-time password was actually consumed and the tier configured. + if cons.calls == 0 { + t.Fatal("the staged credential was never collected") + } + if en.calls == 0 { + t.Fatal("the off-site tier was never configured after collecting the credential") + } +} + +// ── SCENARIO C — a healthy box does not retry, and makes no noise ──────────────────────────────── +// +// RED-PROOF: drop the `!declared()` guard so the tick always reconciles → this FAILS, and a box whose +// tier already works hammers the hub forever. +func TestRetryIfDeclared_HealthyBoxDoesNothing(t *testing.T) { + b, cons, inst, en, logbuf := newBridge(t, goodOffsite()) + + attempted, err := b.RetryIfDeclared(context.Background(), func() bool { return false }) + if err != nil { + t.Fatalf("retry: %v", err) + } + if attempted { + t.Fatal("a box that declares NO need must not reconcile") + } + if cons.calls != 0 || inst.calls != 0 || en.calls != 0 { + t.Fatalf("a healthy box touched the hub: consume=%d install=%d enable=%d", cons.calls, inst.calls, en.calls) + } + if strings.TrimSpace(logbuf.String()) != "" { + t.Fatalf("a healthy box logged noise every tick: %q", logbuf.String()) + } +} + +// A nil bridge (off-site not configured for this customer) is a silent no-op, not a panic — main.go +// wires nil in exactly that case. +func TestRetryIfDeclared_NilBridgeIsSilent(t *testing.T) { + var b *Bridge + attempted, err := b.RetryIfDeclared(context.Background(), func() bool { return true }) + if attempted || err != nil { + t.Fatalf("a nil bridge must be a silent no-op, got attempted=%v err=%v", attempted, err) + } +} + +// ── SCENARIO D — the settle gate still holds on the retry path ─────────────────────────────────── +// +// The retry must not become a back door around the day-0 floor race the gate exists for. It goes +// through ReconcileWhenSettled, so an unsettled box WAITS rather than reconciling immediately. +// +// RED-PROOF: change RetryIfDeclared to call Reconcile directly instead of ReconcileWhenSettled → +// this FAILS, because the reconcile happens while the floor is still unknown. +func TestRetryIfDeclared_HonoursTheSettleGate(t *testing.T) { + b, cons, _, _, _ := newBridge(t, goodOffsite()) + // Floor never becomes known → the gate must hold the reconcile off until its own bound expires. + b.Settle = SettleFunc(func() (string, string, bool, bool) { return "0.203.0", "", false, false }) + + ctx, cancel := context.WithCancel(context.Background()) + cancel() // the gate observes a cancelled context and must not proceed to reconcile + + attempted, err := b.RetryIfDeclared(ctx, func() bool { return true }) + if !attempted { + t.Fatal("the box declared, so a retry attempt must be reported even when the gate stops it") + } + if err == nil { + t.Fatal("a cancelled gate must surface an error, not silently reconcile") + } + if cons.calls != 0 { + t.Fatal("SETTLE GATE BYPASSED: the retry consumed a password while the floor was unknown") + } +} diff --git a/controller/internal/settings/refuse_message_test.go b/controller/internal/settings/refuse_message_test.go new file mode 100644 index 0000000..89e1f0a --- /dev/null +++ b/controller/internal/settings/refuse_message_test.go @@ -0,0 +1,40 @@ +package settings + +import ( + "strings" + "testing" +) + +// ── SCENARIO G (R-220) — AN EMPTY LIST MUST EXPLAIN ITSELF ────────────────────────────────────── +// +// The old refusal said "choose an attached drive from the list" while the list was empty — on a +// rebuilt box, for a reason the customer had no part in and could not see. Measured three times live. +// A refusal that names an action the customer cannot perform is the I3 breach the campaign recorded. +// +// RED-PROOF: restore the old sentence and this FAILS on the impossible-action assertion. +func TestRefuseAppNamespace_DoesNotNameAnImpossibleAction(t *testing.T) { + msg := refuseAppNamespaceUndeterminable + + if strings.Contains(msg, "Válasszon a listából") { + t.Fatal("R-220's I3 breach RETURNED: the refusal tells the customer to choose from a list that may be empty") + } + // It must say WHY the list can be empty — the rebuild — so the state is explicable. + if !strings.Contains(msg, "újratelepítettük") && !strings.Contains(msg, "újra") { + t.Fatalf("the refusal must explain why the drive is unregistered; got %q", msg) + } + // And point somewhere a customer can actually go. + if !strings.Contains(msg, "Meghajtók") { + t.Fatalf("the refusal must name where re-attaching happens; got %q", msg) + } + // It must not promise an outcome it cannot know. + for _, forbidden := range []string{"biztosan", "garantál", "mindig sikerül"} { + if strings.Contains(msg, forbidden) { + t.Errorf("the refusal promises an outcome it cannot know (%q)", forbidden) + } + } + // The NAS refusal is a different situation and keeps its own wording — it points at a list that + // genuinely does have entries, so it is not the same defect. + if !strings.Contains(refuseAppNamespaceNetwork, "Válasszon csatlakoztatott meghajtót") { + t.Fatal("the NAS refusal was changed; it is a different situation and was not part of R-220") + } +} diff --git a/controller/internal/settings/settings.go b/controller/internal/settings/settings.go index 846a5fe..22b8a2c 100644 --- a/controller/internal/settings/settings.go +++ b/controller/internal/settings/settings.go @@ -1580,8 +1580,23 @@ func (s *Settings) RefuseAsAppNamespace(path string) (bool, string) { const ( refuseAppNamespaceNetwork = "Hálózati tárhelyen (NAS) nem futtatható alkalmazás adatkönyvtára — " + "a NAS megosztás tallózásra és médiatárolásra használható. Válasszon csatlakoztatott meghajtót." + // ⚠ R-220 / SCENARIO G — THIS SENTENCE USED TO NAME AN IMPOSSIBLE ACTION. + // + // It said *"Válasszon a listából csatlakoztatott meghajtót"* — choose an attached drive from the + // list — and on a rebuilt box that list is EMPTY, for a reason the customer had no part in and no + // way to see. Measured three times live (CAMPAIGN-11 Phase 1, and twice on the R-201 re-walk). + // Telling someone to pick from an empty list is the I3 breach the campaign recorded: a refusal must + // name a reason a person can act on. + // + // It now says what is true — the drive is not registered ON THIS MACHINE, which is what a rebuild + // causes — and points at the page where re-attaching happens, rather than at a list that may hold + // nothing. It promises no outcome, because whether the drive can be re-attached is not knowable + // from here. refuseAppNamespaceUndeterminable = "A megadott tárhely nem azonosítható regisztrált meghajtóként, " + - "ezért alkalmazás adatkönyvtáraként nem használható. Válasszon a listából csatlakoztatott meghajtót." + "ezért alkalmazás adatkönyvtáraként nem használható. Ha a gépet nemrég telepítettük újra, a " + + "meghajtóid megvannak, de még nincsenek újra csatlakoztatva ehhez a géphez — a Tárhely → " + + "Meghajtók oldalon csatlakoztathatod őket, és utána indítsd újra a telepítést. Ha ott sem " + + "látszanak, keresd a Felhom ügyfélszolgálatát." ) // IsStoragePathSchedulable returns whether a path belongs to a registered,