due-checks gate (R-341), floor raise recorded (R-343), snapshot coverage (R-342)
gates / gates (push) Successful in 14s

PART 1+2 — dated checks stop being wishes. R-341 booked two measurements as
prose in a register row; nothing read those dates and nothing would have
objected when they passed. The dates now live in a DUE-CHECKS block INSIDE
OPEN-ITEMS.md (inside, so no sidecar can drift from it) and a new gate reads
them. Registered as #10 in repo_gates.py, --fast, so it runs in BOTH the
pre-push hook and CI.

  exit 0  nothing due (prints pending count + nearest date; empty block too)
  exit 1  a row is due/overdue (due <= today, UTC -- due TODAY counts), or a
          row names an item with no R-row
  exit 2  block absent/duplicated/unparseable -- INCONCLUSIVE, never 0

It REFUSES rather than warns, and its docstring states the limitation: it is
NOT a scheduler, it fires on the next push, not on the date.

37 tests. BOTH red-proofs run and reverted -- and the first one earned its
keep by catching a hollow assertion of MINE rather than confirming the gate:
flipping <= to < left a due-today row in neither bucket, min() raised on an
empty list, and the TRACEBACK exited 1, so "rc == 1" passed while the
boundary was wrong. An exit code cannot tell a verdict from a crash. The test
now asserts the conviction banner and the absence of a traceback, and the gate
returns 2 rather than crashing if that partition breaks again.

PART 3 — the floor raise, and the premise was WRONG. Read back from the store
(not the form): min_controller_version = 0.216.0 @ 12:36:58Z, zero
per-customer overrides, no "managed floor HELD" line. But read 5 shows the
raise was NOT a no-op: demo-felhom had been on 0.214.0 since 12 Aug and
auto-updated 0.214.0 -> 0.216.0 at 12:37:07Z -- NINE SECONDS after the save,
exactly the immediate action publish-train rule 2 documents. No error events
followed; it restarted clean.

R-343 is therefore filed OPEN, not CLOSED: the closing condition was all five
reads clean and no directive served. It went well, but a record calling it
inert when it moved a customer box is what misleads the next reader. The row
also states why the floor was behind -- rule 2 policy, not drift, earned by
the 2026-07-11 skew onto Peti's box -- and cites ResolveManagedFloor
(store.go:2068) plus the two build-felhom-iso.sh facts (build-time at :267,
fails open at :78-82) rather than asserting them.

Two boxes are below the floor and neither reports: drill-r50 (blocked,
powered off) and peti-felhom (host row deleted). peti-felhom was NOT
contacted -- its row records that a report from a deleted host 401s and is
not persisted, so the raise cannot reach it.

PART 4 — R-342 filed READY, quoting stop2-snapshot.txt verbatim: Hetzner
server snapshot 421440873 covers /dev/sda only; /mnt/pbs-datastore is a
separate Volume that snapshots exclude, so a rollback restores software state
and NOT the datastore. Fine for that upgrade; the safeguard for any future
procedure that could touch the datastore does not exist and is Viktor's call.

Also: CLAUDE.md's gate list named 6 of 10 registered gates -- completed
rather than appending a 7th to a wrong list (124 -> 128 effective, ceiling
200). Capability map deliberately unchanged; no row cites a floor or golden
version. repo_gates.py fully green, 10/10.
This commit is contained in:
2026-08-18 15:16:51 +02:00
parent f267bc047f
commit 0a5e9b14dc
10 changed files with 905 additions and 1 deletions
+20
View File
@@ -15,6 +15,26 @@
> would make one of the two audiences stop reading. `STATUS.md` is also a **view of `OPEN-ITEMS.md`**
> and holds nothing of its own; this file does hold its own content, namely the standing rulings below.
## The managed floor now tracks the vouched golden — and that is rule 2 ARRIVING, not an exception (2026-08-18, R-343)
Since 2026-08-18 12:36:58Z the global managed controller floor
(`hub_settings.min_controller_version`) reads **0.216.0**, equal to the vouched golden. Read it from
the store (`GetGlobalMinControllerVersion`, `hub/internal/store/store.go:1751`), never from the form.
**Do not read the preceding gap as drift.** `publish-train-rules.md` rule 1 is *manifest before
floor* and rule 2 requires the floor field saved **LAST, in a separate save**, because the DB row
overrides the env floor and acts on the next report cycle. The floor therefore sits behind the newest
controller **by design** for as long as it takes to bake and vouch a golden — and it is raised
afterwards. Floor-equals-golden is the state that ordering is meant to arrive at; a session finding
the floor behind should check whether a vouch is outstanding before calling it a defect.
**The raise is not free, and this is the part worth remembering.** Rule 2's "acts immediately" is
literal: nine seconds after the save, `demo-felhom` auto-updated 0.214.0 → 0.216.0
(`controller_updated` 12:37:07Z, `controller_started` 12:37:12Z, no error events). So a floor raise is
a fleet action, not bookkeeping — plan it as one. `ResolveManagedFloor`
(`hub/internal/store/store.go:2068`) is what makes it safe: it holds the floor entirely when it
exceeds the golden (R-216), and per-box when that box's agent is below the manifest's `MinAgent`.
## A name must differ from its neighbours on a stem, not on a noun (2026-08-13, R-323)
The third near-homograph is renamed: the five-word phrase that proves an account owns the box being