Files
felhom.eu/documentation/runbooks/publish-train-rules.md
T
admin 0a5e9b14dc
gates / gates (push) Successful in 14s
due-checks gate (R-341), floor raise recorded (R-343), snapshot coverage (R-342)
PART 1+2 — dated checks stop being wishes. R-341 booked two measurements as
prose in a register row; nothing read those dates and nothing would have
objected when they passed. The dates now live in a DUE-CHECKS block INSIDE
OPEN-ITEMS.md (inside, so no sidecar can drift from it) and a new gate reads
them. Registered as #10 in repo_gates.py, --fast, so it runs in BOTH the
pre-push hook and CI.

  exit 0  nothing due (prints pending count + nearest date; empty block too)
  exit 1  a row is due/overdue (due <= today, UTC -- due TODAY counts), or a
          row names an item with no R-row
  exit 2  block absent/duplicated/unparseable -- INCONCLUSIVE, never 0

It REFUSES rather than warns, and its docstring states the limitation: it is
NOT a scheduler, it fires on the next push, not on the date.

37 tests. BOTH red-proofs run and reverted -- and the first one earned its
keep by catching a hollow assertion of MINE rather than confirming the gate:
flipping <= to < left a due-today row in neither bucket, min() raised on an
empty list, and the TRACEBACK exited 1, so "rc == 1" passed while the
boundary was wrong. An exit code cannot tell a verdict from a crash. The test
now asserts the conviction banner and the absence of a traceback, and the gate
returns 2 rather than crashing if that partition breaks again.

PART 3 — the floor raise, and the premise was WRONG. Read back from the store
(not the form): min_controller_version = 0.216.0 @ 12:36:58Z, zero
per-customer overrides, no "managed floor HELD" line. But read 5 shows the
raise was NOT a no-op: demo-felhom had been on 0.214.0 since 12 Aug and
auto-updated 0.214.0 -> 0.216.0 at 12:37:07Z -- NINE SECONDS after the save,
exactly the immediate action publish-train rule 2 documents. No error events
followed; it restarted clean.

R-343 is therefore filed OPEN, not CLOSED: the closing condition was all five
reads clean and no directive served. It went well, but a record calling it
inert when it moved a customer box is what misleads the next reader. The row
also states why the floor was behind -- rule 2 policy, not drift, earned by
the 2026-07-11 skew onto Peti's box -- and cites ResolveManagedFloor
(store.go:2068) plus the two build-felhom-iso.sh facts (build-time at :267,
fails open at :78-82) rather than asserting them.

Two boxes are below the floor and neither reports: drill-r50 (blocked,
powered off) and peti-felhom (host row deleted). peti-felhom was NOT
contacted -- its row records that a report from a deleted host 401s and is
not persisted, so the raise cannot reach it.

PART 4 — R-342 filed READY, quoting stop2-snapshot.txt verbatim: Hetzner
server snapshot 421440873 covers /dev/sda only; /mnt/pbs-datastore is a
separate Volume that snapshots exclude, so a rollback restores software state
and NOT the datastore. Fine for that upgrade; the safeguard for any future
procedure that could touch the datastore does not exist and is Viktor's call.

Also: CLAUDE.md's gate list named 6 of 10 registered gates -- completed
rather than appending a 7th to a wrong list (124 -> 128 effective, ceiling
200). Capability map deliberately unchanged; no row cites a floor or golden
version. repo_gates.py fully green, 10/10.
2026-08-18 15:16:51 +02:00

92 lines
6.2 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Publish-train rules — coupled controller/agent releases
Standing gates for every artifact publish train (agent binary + controller golden + Day-0 manifest +
floor). Each rule cites the incident that created it. The box-level backstop (rule 4) exists because
rules 1–3 are operator discipline and discipline fails; the rules stay the primary control.
## 1. Manifest before floor (the GL-1 rule, restated)
Publish and vouch the artifacts in the Day-0 manifest **before** any floor movement. A floor that
points at an unvouched (or unpublished) version bricks self-updates: boxes are told to move to a
version they cannot verify. (GL-1 supply-chain arc; see the go-live package records.)
## 2. The manifest screen carries the LIVE DB floor — save the floor field LAST
The hub's Day-0 manifest UI also persists the global floor as a DB row
(`hub_settings.min_controller_version` — `hub/internal/store/store.go`,
`GetGlobalMinControllerVersion`). **The DB row overrides the env floor and acts immediately on the
next report cycle** — it is not staged by the manifest vouch.
**Incident:** publish train 0.81/0.113 (2026-07-11,
`documentation/pilot/RUNBOOK-publish-0.81-0.113-2026-07-11.md`): the operator saved the manifest
screen with the floor field filled; the floor acted at once and controller 0.113.0 reached Peti's
box ~9 minutes **before** agent 0.81.0 — the exact forbidden skew, benign only because the box had
zero NAS shares.
**Rule:** fill the floor field **last**, in a separate save, only after rule 3's fleet check passes.
## 3. MinAgent fleet gate — HUB-ENFORCED PER-BOX since hub v0.45.0
A controller release that depends on coupled agent behavior declares `MinAgent: X.Y.Z` in its
CHANGELOG entry header line (felhom-controller convention, since v0.114.0; retroactively, 0.113.0's
effective MinAgent was 0.81.0 for the NAS add). The floor may not effectively push a box past a
controller whose MinAgent that box's agent does not yet meet.
**Hub-enforced per-box since hub v0.45.0 (the manual fleet check is retired):** the operator sets
the golden's **`MinAgent`** in the Day-0 artifact manifest (Configuration → Day-0 artifacts →
"Min agent"; blank = uncoupled release, no gating). At report-ACK time the hub compares each box's
reported `hosts.agent_version` against that MinAgent
(`store.ResolveManagedFloor`, `hub/internal/store/store.go`): agent ≥ MinAgent → the controller
floor is served; agent below MinAgent **or unknown** → the floor is **HELD** (the ACK omits the
directive) and the box is flagged on the Hosts dashboard (`floor held: agent <v> < MinAgent <w>`) —
a held box is visible, never silently stale. The operator still SETS MinAgent at manifest time; the
hub does the per-box gating. This mechanises the "agent BEFORE controller floor" ordering that rule
2's incident violated by hand.
Effective-floor visibility (same v0.45.0): the floor card shows the resolved value + its source
(DB `hub_settings` vs env `DEFAULT_MIN_CONTROLLER_VERSION`, both raw values when they differ), and
the floor save is behind a type-to-confirm dialog stating the live below-floor blast radius — so
rule 2's "the floor acts immediately" is impossible to miss.
## 4. Box-level backstop: the controller's capability gate
Since controller v0.114.0, a coupled feature entry point probes the agent's capability
(`controller/internal/agentapi/features.go` — route probe: 2xx ⇒ supported, 404 ⇒ older agent,
transport/5xx ⇒ indeterminate, never "too old") and refuses up front with an honest Hungarian
message instead of failing mid-pipeline. This turns a violated ordering into a graceful refusal —
it does **not** license sloppy trains: rules 1–3 remain the primary control.
Convention for new coupled features: add a row to the `featureProbes` table AND the `featureMinAgent`
table + a `Supports` gate call at the feature's entry point, and declare MinAgent per rule 3. Since
controller v0.115.0 + agent v0.82.0 the agent reports its version in the `X-Felhom-Agent-Version`
response header, so `Supports` decides by version comparison when the version is known and only
falls back to the route probe for header-less (≤0.81) agents — capability detection is now explicit,
not probe-inferred.
## 5. Every ISO build asserts golden ≥ the managed floor (the R-71 gate)
**Incident:** `DIAG-f10-demo-hp-offsite-2026-07-23.md` / R-71. A box installed from an ISO whose
golden controller is BELOW the hub's managed floor boots below the floor, so the day-0 managed
update fires within minutes of first boot — racing the offsite apply-bridge in exactly the window
that burned demo-hp's one-time offsite credential (2 days unprotected). The gap is invisible at
build time unless something checks it.
**Rule:** a build where the managed floor exceeds golden must **fail loudly, at build time**, not
ship. `build-felhom-iso.sh` carries the assertion itself (`assert_golden_ge_floor`, gated on
`FELHOM_ASSERT_GOLDEN` = the hub's `artifact_golden_version` and `FELHOM_ASSERT_FLOOR` = its
`min_controller_version`); it prints both versions and `die`s on `golden < floor`. Resolve the two
values operator-side before the build and pass them in — e.g. from the hub DB
(`hub_settings.artifact_golden_version` / `min_controller_version`) or the operator artifacts UI —
so the gate is enforced, not skipped (an unset input warns LOUDLY and does not silently pass). The
fix for a tripped gate is never to lower the floor: **republish golden ≥ floor and vouch it**
(rule 1), then rebuild. This gate composes with rule 1 — the manifest still leads the floor; this
one stops an ISO from carrying a golden the floor has already outrun.
> **Dated note, 2026-08-18 — the current pair is golden `0.216.0` / floor `0.216.0`.** Equal, so the
> R-71 gate passes. Recorded so the next build has numbers to hand, and deliberately **not** as a
> replacement for the instruction above: **resolve both LIVE from the hub before a build.** A
> snapshot that looks authoritative is exactly how `build-golden.sh`'s hand-bumped `CONTROLLER_IMAGE`
> default rotted twice. Treat these as a sanity check on what you read, never as a substitute for
> reading it. (Floor raised to 0.216.0 by the operator on 2026-08-18 in a separate save after the
> vouch, per rule 2 — see R-343, which records that the raise moved `demo-felhom` nine seconds later.)