Files
felhom.eu/REPORT-register-and-floor.md
T
admin 0a5e9b14dc
gates / gates (push) Successful in 14s
due-checks gate (R-341), floor raise recorded (R-343), snapshot coverage (R-342)
PART 1+2 — dated checks stop being wishes. R-341 booked two measurements as
prose in a register row; nothing read those dates and nothing would have
objected when they passed. The dates now live in a DUE-CHECKS block INSIDE
OPEN-ITEMS.md (inside, so no sidecar can drift from it) and a new gate reads
them. Registered as #10 in repo_gates.py, --fast, so it runs in BOTH the
pre-push hook and CI.

  exit 0  nothing due (prints pending count + nearest date; empty block too)
  exit 1  a row is due/overdue (due <= today, UTC -- due TODAY counts), or a
          row names an item with no R-row
  exit 2  block absent/duplicated/unparseable -- INCONCLUSIVE, never 0

It REFUSES rather than warns, and its docstring states the limitation: it is
NOT a scheduler, it fires on the next push, not on the date.

37 tests. BOTH red-proofs run and reverted -- and the first one earned its
keep by catching a hollow assertion of MINE rather than confirming the gate:
flipping <= to < left a due-today row in neither bucket, min() raised on an
empty list, and the TRACEBACK exited 1, so "rc == 1" passed while the
boundary was wrong. An exit code cannot tell a verdict from a crash. The test
now asserts the conviction banner and the absence of a traceback, and the gate
returns 2 rather than crashing if that partition breaks again.

PART 3 — the floor raise, and the premise was WRONG. Read back from the store
(not the form): min_controller_version = 0.216.0 @ 12:36:58Z, zero
per-customer overrides, no "managed floor HELD" line. But read 5 shows the
raise was NOT a no-op: demo-felhom had been on 0.214.0 since 12 Aug and
auto-updated 0.214.0 -> 0.216.0 at 12:37:07Z -- NINE SECONDS after the save,
exactly the immediate action publish-train rule 2 documents. No error events
followed; it restarted clean.

R-343 is therefore filed OPEN, not CLOSED: the closing condition was all five
reads clean and no directive served. It went well, but a record calling it
inert when it moved a customer box is what misleads the next reader. The row
also states why the floor was behind -- rule 2 policy, not drift, earned by
the 2026-07-11 skew onto Peti's box -- and cites ResolveManagedFloor
(store.go:2068) plus the two build-felhom-iso.sh facts (build-time at :267,
fails open at :78-82) rather than asserting them.

Two boxes are below the floor and neither reports: drill-r50 (blocked,
powered off) and peti-felhom (host row deleted). peti-felhom was NOT
contacted -- its row records that a report from a deleted host 401s and is
not persisted, so the raise cannot reach it.

PART 4 — R-342 filed READY, quoting stop2-snapshot.txt verbatim: Hetzner
server snapshot 421440873 covers /dev/sda only; /mnt/pbs-datastore is a
separate Volume that snapshots exclude, so a rollback restores software state
and NOT the datastore. Fine for that upgrade; the safeguard for any future
procedure that could touch the datastore does not exist and is Viktor's call.

Also: CLAUDE.md's gate list named 6 of 10 registered gates -- completed
rather than appending a 7th to a wrong list (124 -> 128 effective, ceiling
200). Capability map deliberately unchanged; no row cites a floor or golden
version. repo_gates.py fully green, 10/10.
2026-08-18 15:16:51 +02:00

14 KiB
Raw Blame History

REPORT — dated checks that bite, the floor raise on the record, the snapshot that covers less (2026-08-18)

Three pieces of bookkeeping, no machine put at risk. The hub was READ ONLY throughout.

The headline is that Part 3's premise was wrong. The floor raise was not a no-op: it moved a live customer box nine seconds after the save. That is the whole reason the task said to read it back rather than assume it. R-343 is therefore filed OPEN, not CLOSED, per the task's own condition.


1. Confirmed baselines

item value
felhom.eu main @ start f267bc047f198de4cb600068fdd8bcef557cff20 — matches the sheet
clean tree at start yes; HEAD == origin/main in felhom.eu, felhom-controller, felhom-agent
scripts/ version IN felhom-host-install.sh v1.28.0 (CHANGELOG head)
scripts/ version OUT due_checks_gate.py v1.0.0 (new head entry)

2. Files created / modified

Created: scripts/due_checks_gate.py, scripts/test_due_checks_gate.py, REPORT-register-and-floor.md. Modified: scripts/repo_gates.py (registration + docstring), scripts/CHANGELOG.md, CLAUDE.md, CONTEXT.md, STATUS.md, documentation/backlog/OPEN-ITEMS.md (block + R-342 + R-343), documentation/runbooks/publish-train-rules.md. Commit hashes are in §9.

3. Tests — 37 assertions, and a red-proof that caught my own test

python3 scripts/test_due_checks_gate.py → passed: 37, failed: 0, groups A–G.

Red-proof 1 — the boundary. It failed usefully: it exposed a HOLLOW assertion of mine.

Mutation: due_now = [... if r[1] <= today] → < today. First run, before the fix: Group C reported

  PASS  C: due TODAY exits 1  rc=1          <-- passed, and should NOT have
  FAIL  C: says DUE TODAY rather than overdue

The rc == 1 assertion passed for the wrong reason. With <, a row dated exactly today falls into neither due_now (<) nor pending (>), so min(pending, …) raised ValueError: min() iterable argument is empty and the traceback exited 1. Confirmed directly:

  File ".../due_checks_gate.py", line 232, in main
    nearest = min(pending, key=lambda r: r[1])
ValueError: min() iterable argument is empty
RC=1

An exit code alone cannot distinguish a verdict from a crash. Two fixes, both kept:

  1. the test now asserts DUE-CHECKS GATE FAILED is in the output and Traceback is not, plus a new test_c_gate_never_ends_in_a_traceback across overdue/future/empty inputs;
  2. the gate returns 2 (INCONCLUSIVE) with a message if the partition is ever broken again, because a crash is never a verdict.

Re-run after the fix — the mutation now bites properly:

  FAIL  C: due TODAY exits 1 (boundary is <=)  rc=2
  FAIL  C: exits 1 as a VERDICT, not a traceback
  PASS  C: did not crash
  FAIL  C: says DUE TODAY rather than overdue
passed: 34   failed: 3

Mutation reverted, verified by grep -n "MUTATED" returning nothing and the <= line restored.

Red-proof 2 — the missing-block path

Mutation: the missing-block branch sys.exit(2) → sys.exit(0). Seen failing:

  FAIL  E: missing block exits 2 (NOT 0)  rc=0
passed: 36   failed: 1

Reverted, sys.exit(2) restored on that branch.

4. The gate's real output in all three states

Overdue (fixture, today=2026-08-20):

DUE-CHECKS GATE FAILED: 1 dated check(s) are due or overdue as of 2026-08-20 (UTC).

  R-341    due 2026-08-19   1 day(s) OVERDUE
           measure: ep0 proxy fd count + ESTAB/CLOSE-WAIT split; PID must still be 551655
           the command and its preconditions are in the R-341 row of documentation/backlog/OPEN-ITEMS.md

Take the measurement, record the result in that R-row, then remove the row from the
DUE-CHECKS block. Moving the date instead is allowed — state the reason in the R-row.
NOTE: this gate fires on a PUSH, not on the date; it may be later than the date.

Pending / the LIVE run against the real register today (these are the same run):

due-checks gate OK — 2 dated check(s) pending, none due yet.
  today (UTC): 2026-08-18
  nearest: R-341 due 2026-08-19 (in 1 day(s)) — ep0 proxy fd count + ESTAB/CLOSE-WAIT split; PID must still be 551655
  (fires on the next PUSH after a date passes, not on the date itself — by design)

5. The runner's output

  site               OK            (exit 0)
  hostinstall        OK            (exit 0)
  hub-confirm        OK            (exit 0)
  manifest-bearer    OK            (exit 0)
  reuse-refs         OK            (exit 0)
  instructions       OK            (exit 0)
  golden-currency    OK            (exit 0)
  wire-contract      OK            (exit 0)
  hub-copy           OK            (exit 0)
  due-checks         OK            (exit 0)
all felhom.eu gates OK

Group G asserts registration by running the runner and matching due-checks in its output, never by grepping repo_gates.py's source — a commented-out entry still contains the string.

6. Part 3's five reads — evidence, not summary

READ 1 — the live floor, from the store

  artifact_agent_version     = '0.129.0'    (updated 2026-08-18 11:00:59)
  artifact_golden_version    = '0.216.0'    (updated 2026-08-18 11:00:59)
  artifact_min_agent         = '0.129.0'    (updated 2026-08-18 11:01:00)
  min_controller_version     = '0.216.0'    (updated 2026-08-18 12:36:58)

The raise landed, so Part 3 proceeded. Read from hub_settings, not the form.

READ 2 — per-customer overrides

  demo-felhom      status=active     override=''         config_version=12
  demo-hp          status=active     override=''         config_version=5
  drill-r50        status=blocked    override=''         config_version=1
  peti-felhom      status=active     override=''         config_version=6
  tester-1         status=active     override=''         config_version=1
  -> 0 customer(s) carry a non-empty override

Zero overrides, so the global applies to everyone and no box hides behind a lower one.

READ 3 — every box's controller version. Two are below the floor.

  demo-felhom      controller='0.216.0'    last_report=2026-08-18 13:07:07
  demo-hp          controller='0.216.0'    last_report=2026-08-18 13:01:34
  drill-r50        controller='0.213.0'    last_report=2026-08-12 15:33:25
  peti-felhom      controller='0.115.0'    last_report=2026-07-15 08:39:00

The two REPORTING boxes are both at 0.216.0, at the floor. The other two are below it and neither is a reporting box: drill-r50 is status=blocked, last heard from six days ago, powered off and reverted; peti-felhom's host row was deleted on 2026-07-15. Reported here rather than as a footnote, per the task's edge-case rule.

(Note: guests.controller_version is empty for every guest — the hub carries the controller version on reports.controller_version, not on the guest row. The first query I wrote read the guest field and would have reported "unknown" for every box.)

READ 4 — directives and holds

2026/08/18 14:36:58 [INFO] Global controller-version floor set to "0.216.0"

No managed floor HELD line exists — searched over 24 h of pod logs. (Hub log lines are CEST; the DB stores UTC, hence 14:36:58 here and 12:36:58 above — the same instant.)

READ 5 — THE FINDING: a controller DID auto-update after the raise

2026-08-18 12:37:07 | demo-felhom | controller_updated | Controller frissítve: 0.214.0 → 0.216.0
2026-08-18 12:37:12 | demo-felhom | controller_started | Controller elindult (0.216.0)

and the version trail confirms it:

   2026-08-18 11:14:55  controller=0.214.0
   2026-08-18 12:37:12  controller=0.216.0

demo-felhom had been on 0.214.0 since 2026-08-12 16:44 and the floor raise pulled it to 0.216.0 nine seconds after the save — exactly the "acts immediately on the next report cycle" that publish-train-rules.md rule 2 documents and that the 2026-07-11 incident was filed for. demo-hp was already on 0.216.0 (hand-deployed 2026-08-14 08:31) and did not move.

No error, warning or critical event followed — the update completed and the controller restarted. So: harmless in outcome, but not a no-op. "Every reporting box is at or above the floor" is true because of the raise, not independently of it.

R-343 is filed OPEN. The task's closing condition was all five reads clean and no directive served; read 5 shows a live box moved. It went well, and a record that called it inert would mislead the next reader.

7. peti-felhom — not contacted

The machine was not contacted in any way. Sourced from the PETI register row, quoted:

"a report from a deleted host 401s and is not persisted"

with its host row deleted 2026-07-15 08:56:22 (host_deletions id=1). It therefore cannot receive a floor directive and the raise cannot reach it. Its reports row still shows controller 0.115.0 from its last report on 2026-07-15 08:39:00 — a stale record, not a live box.

8. The two build-felhom-iso.sh facts, confirmed in the script

(a) It is a BUILD-TIME gate. assert_golden_ge_floor() is defined at :77 and called at :267, in the build flow.

(b) It FAILS OPEN with a warning when its inputs are absent — :78-82:

    local golden="${FELHOM_ASSERT_GOLDEN:-}" floor="${FELHOM_ASSERT_FLOOR:-}"
    if [[ -z "$golden" || -z "$floor" ]]; then
        log_warn "R-71 golden>=floor gate UNENFORCED — pass FELHOM_ASSERT_GOLDEN + FELHOM_ASSERT_FLOOR to enforce (golden='${golden:-unset}' floor='${floor:-unset}')"
        return 0
    fi

Both read as the task described. No ISO rebuild is required: the golden is fetched at first boot from the hub's manifest (0.216.0 — at the floor, not below it), and this gate governs future builds.

10. NOT yet validated

The gate has never fired on a real overdue date in the live register. Every conviction shown here is from a temp-file fixture or a FELHOM_GATE_TODAY override. Its first genuine firing will be the next push on or after 2026-08-19, when R-341's first check comes due. Until that happens, "it refuses the push" is proven in fixtures and inferred in production — the registration test proves it is wired into the runner, which is the part that could silently not be true.

Also unvalidated: the block's own upkeep. Nothing checks that a row removed from the block was removed because the measurement was taken rather than because it was inconvenient.

11. Teardown

This task provisioned nothing. No VM, no container, no machine touched. The hub was read-only — snapshots of hub.db + -wal were taken into the session scratchpad for querying and are not committed.

12. Register rows

The block, verbatim as committed:

<!-- DUE-CHECKS-BEGIN — machine-readable. Parsed by scripts/due_checks_gate.py.
     One row per dated check. The R-number must have a row above. Dates are UTC.
     Clearing a row means the check was DONE and its result recorded in that R-row —
     or the date was deliberately moved, with the reason stated in the R-row.
     This block is an INDEX, not the detail: the command and the preconditions live in
     the R-row. Duplicating them here would create the second source this design avoids. -->
| item | due (UTC) | what to measure |
|---|---|---|
| R-341 | 2026-08-19 | ep0 proxy fd count + ESTAB/CLOSE-WAIT split; PID must still be 551655 |
| R-341 | 2026-08-25 | same, +7 d |
<!-- DUE-CHECKS-END -->

R-341's dates in the register matched the sheet exactly — no disagreement to report.

R-343's verdict cell, verbatim:

OPEN — NEW 2026-08-18. Deliberately NOT closed: the task's closing condition was all five reads clean, no directive served, and read 5 shows a live box updated. It went cleanly and is the floor working as designed — but a change recorded as a no-op when it moved a customer box is exactly the kind of record that misleads later

R-342 filed READY (S), owner Viktor decides; CC executes, quoting stop2-snapshot.txt verbatim on what the snapshot covers and does not.

13. unproven.py --summary

where felhom stands — 55 claims, verified_on 2026-08-09
  walked   23
  partial  14   (6 cite evidence, 8 prose only)
  built    14   (0 cite evidence, 14 prose only)
  missing  4   (0 cite evidence, 4 prose only)
  NOT WALKED: 32 of 55

No number moved, correctly: this task added a gate and three register facts, and walked no claim in the standing picture.

14. Observations — noticed, deliberately not acted on

  • CLAUDE.md's gate list named only 6 of the 10 registered gates. It was missing golden_currency_gate.py, wire_contract_gate.py and hub_copy_gate.py — all registered weeks ago. I completed the list rather than appending a 7th name to a list that was already wrong, since the section's stated job is to name each gate. Effective line count 124 → 128 against a ceiling of 200, so no trim was needed.
  • documentation/architecture/00-capability-map.md — no change, and this is the explicit statement the task asked for. No row's evidence citation names the floor or golden version: line 153's publish-train row cites a runbook path, and line 44's golden literal is a dated historical citation on the recovery-journey row.
  • drill-r50 will be dragged 0.213.0 → 0.216.0 by this floor if it is ever booted and reports. Its agent (0.129.0) meets MinAgent, so the floor would be served, not held. That is the floor doing its job; noted so it is not read as a surprise later.
  • The hub's log timestamps are CEST while its DB stores UTC. Not a defect, but it makes a log line and an events row for the same instant look two hours apart, which is worth knowing before correlating them under pressure.
  • min_controller_version and artifact_* live in the same hub_settings table but are saved by different actions, two hours apart today (11:00:59 vouch, 12:36:58 floor). That separation is rule 2 working, and is why reading only one of them would give a misleading picture.