PART 1+2 — dated checks stop being wishes. R-341 booked two measurements as prose in a register row; nothing read those dates and nothing would have objected when they passed. The dates now live in a DUE-CHECKS block INSIDE OPEN-ITEMS.md (inside, so no sidecar can drift from it) and a new gate reads them. Registered as #10 in repo_gates.py, --fast, so it runs in BOTH the pre-push hook and CI. exit 0 nothing due (prints pending count + nearest date; empty block too) exit 1 a row is due/overdue (due <= today, UTC -- due TODAY counts), or a row names an item with no R-row exit 2 block absent/duplicated/unparseable -- INCONCLUSIVE, never 0 It REFUSES rather than warns, and its docstring states the limitation: it is NOT a scheduler, it fires on the next push, not on the date. 37 tests. BOTH red-proofs run and reverted -- and the first one earned its keep by catching a hollow assertion of MINE rather than confirming the gate: flipping <= to < left a due-today row in neither bucket, min() raised on an empty list, and the TRACEBACK exited 1, so "rc == 1" passed while the boundary was wrong. An exit code cannot tell a verdict from a crash. The test now asserts the conviction banner and the absence of a traceback, and the gate returns 2 rather than crashing if that partition breaks again. PART 3 — the floor raise, and the premise was WRONG. Read back from the store (not the form): min_controller_version = 0.216.0 @ 12:36:58Z, zero per-customer overrides, no "managed floor HELD" line. But read 5 shows the raise was NOT a no-op: demo-felhom had been on 0.214.0 since 12 Aug and auto-updated 0.214.0 -> 0.216.0 at 12:37:07Z -- NINE SECONDS after the save, exactly the immediate action publish-train rule 2 documents. No error events followed; it restarted clean. R-343 is therefore filed OPEN, not CLOSED: the closing condition was all five reads clean and no directive served. It went well, but a record calling it inert when it moved a customer box is what misleads the next reader. The row also states why the floor was behind -- rule 2 policy, not drift, earned by the 2026-07-11 skew onto Peti's box -- and cites ResolveManagedFloor (store.go:2068) plus the two build-felhom-iso.sh facts (build-time at :267, fails open at :78-82) rather than asserting them. Two boxes are below the floor and neither reports: drill-r50 (blocked, powered off) and peti-felhom (host row deleted). peti-felhom was NOT contacted -- its row records that a report from a deleted host 401s and is not persisted, so the raise cannot reach it. PART 4 — R-342 filed READY, quoting stop2-snapshot.txt verbatim: Hetzner server snapshot 421440873 covers /dev/sda only; /mnt/pbs-datastore is a separate Volume that snapshots exclude, so a rollback restores software state and NOT the datastore. Fine for that upgrade; the safeguard for any future procedure that could touch the datastore does not exist and is Viktor's call. Also: CLAUDE.md's gate list named 6 of 10 registered gates -- completed rather than appending a 7th to a wrong list (124 -> 128 effective, ceiling 200). Capability map deliberately unchanged; no row cites a floor or golden version. repo_gates.py fully green, 10/10.
6.2 KiB
Publish-train rules — coupled controller/agent releases
Standing gates for every artifact publish train (agent binary + controller golden + Day-0 manifest + floor). Each rule cites the incident that created it. The box-level backstop (rule 4) exists because rules 1–3 are operator discipline and discipline fails; the rules stay the primary control.
1. Manifest before floor (the GL-1 rule, restated)
Publish and vouch the artifacts in the Day-0 manifest before any floor movement. A floor that points at an unvouched (or unpublished) version bricks self-updates: boxes are told to move to a version they cannot verify. (GL-1 supply-chain arc; see the go-live package records.)
2. The manifest screen carries the LIVE DB floor — save the floor field LAST
The hub's Day-0 manifest UI also persists the global floor as a DB row
(hub_settings.min_controller_version — hub/internal/store/store.go,
GetGlobalMinControllerVersion). The DB row overrides the env floor and acts immediately on the
next report cycle — it is not staged by the manifest vouch.
Incident: publish train 0.81/0.113 (2026-07-11,
documentation/pilot/RUNBOOK-publish-0.81-0.113-2026-07-11.md): the operator saved the manifest
screen with the floor field filled; the floor acted at once and controller 0.113.0 reached Peti's
box ~9 minutes before agent 0.81.0 — the exact forbidden skew, benign only because the box had
zero NAS shares.
Rule: fill the floor field last, in a separate save, only after rule 3's fleet check passes.
3. MinAgent fleet gate — HUB-ENFORCED PER-BOX since hub v0.45.0
A controller release that depends on coupled agent behavior declares MinAgent: X.Y.Z in its
CHANGELOG entry header line (felhom-controller convention, since v0.114.0; retroactively, 0.113.0's
effective MinAgent was 0.81.0 for the NAS add). The floor may not effectively push a box past a
controller whose MinAgent that box's agent does not yet meet.
Hub-enforced per-box since hub v0.45.0 (the manual fleet check is retired): the operator sets
the golden's MinAgent in the Day-0 artifact manifest (Configuration → Day-0 artifacts →
"Min agent"; blank = uncoupled release, no gating). At report-ACK time the hub compares each box's
reported hosts.agent_version against that MinAgent
(store.ResolveManagedFloor, hub/internal/store/store.go): agent ≥ MinAgent → the controller
floor is served; agent below MinAgent or unknown → the floor is HELD (the ACK omits the
directive) and the box is flagged on the Hosts dashboard (floor held: agent <v> < MinAgent <w>) —
a held box is visible, never silently stale. The operator still SETS MinAgent at manifest time; the
hub does the per-box gating. This mechanises the "agent BEFORE controller floor" ordering that rule
2's incident violated by hand.
Effective-floor visibility (same v0.45.0): the floor card shows the resolved value + its source
(DB hub_settings vs env DEFAULT_MIN_CONTROLLER_VERSION, both raw values when they differ), and
the floor save is behind a type-to-confirm dialog stating the live below-floor blast radius — so
rule 2's "the floor acts immediately" is impossible to miss.
4. Box-level backstop: the controller's capability gate
Since controller v0.114.0, a coupled feature entry point probes the agent's capability
(controller/internal/agentapi/features.go — route probe: 2xx ⇒ supported, 404 ⇒ older agent,
transport/5xx ⇒ indeterminate, never "too old") and refuses up front with an honest Hungarian
message instead of failing mid-pipeline. This turns a violated ordering into a graceful refusal —
it does not license sloppy trains: rules 1–3 remain the primary control.
Convention for new coupled features: add a row to the featureProbes table AND the featureMinAgent
table + a Supports gate call at the feature's entry point, and declare MinAgent per rule 3. Since
controller v0.115.0 + agent v0.82.0 the agent reports its version in the X-Felhom-Agent-Version
response header, so Supports decides by version comparison when the version is known and only
falls back to the route probe for header-less (≤0.81) agents — capability detection is now explicit,
not probe-inferred.
5. Every ISO build asserts golden ≥ the managed floor (the R-71 gate)
Incident: DIAG-f10-demo-hp-offsite-2026-07-23.md / R-71. A box installed from an ISO whose
golden controller is BELOW the hub's managed floor boots below the floor, so the day-0 managed
update fires within minutes of first boot — racing the offsite apply-bridge in exactly the window
that burned demo-hp's one-time offsite credential (2 days unprotected). The gap is invisible at
build time unless something checks it.
Rule: a build where the managed floor exceeds golden must fail loudly, at build time, not
ship. build-felhom-iso.sh carries the assertion itself (assert_golden_ge_floor, gated on
FELHOM_ASSERT_GOLDEN = the hub's artifact_golden_version and FELHOM_ASSERT_FLOOR = its
min_controller_version); it prints both versions and dies on golden < floor. Resolve the two
values operator-side before the build and pass them in — e.g. from the hub DB
(hub_settings.artifact_golden_version / min_controller_version) or the operator artifacts UI —
so the gate is enforced, not skipped (an unset input warns LOUDLY and does not silently pass). The
fix for a tripped gate is never to lower the floor: republish golden ≥ floor and vouch it
(rule 1), then rebuild. This gate composes with rule 1 — the manifest still leads the floor; this
one stops an ISO from carrying a golden the floor has already outrun.
Dated note, 2026-08-18 — the current pair is golden
0.216.0/ floor0.216.0. Equal, so the R-71 gate passes. Recorded so the next build has numbers to hand, and deliberately not as a replacement for the instruction above: resolve both LIVE from the hub before a build. A snapshot that looks authoritative is exactly howbuild-golden.sh's hand-bumpedCONTROLLER_IMAGEdefault rotted twice. Treat these as a sanity check on what you read, never as a substitute for reading it. (Floor raised to 0.216.0 by the operator on 2026-08-18 in a separate save after the vouch, per rule 2 — see R-343, which records that the raise moveddemo-felhomnine seconds later.)