Files
felhom.eu/REPORT-register-and-floor.md
T
2026-08-18 15:17:48 +02:00

15 KiB
Raw Blame History

REPORT — dated checks that bite, the floor raise on the record, the snapshot that covers less (2026-08-18)

Three pieces of bookkeeping, no machine put at risk. The hub was READ ONLY throughout.

The headline is that Part 3's premise was wrong. The floor raise was not a no-op: it moved a live customer box nine seconds after the save. That is the whole reason the task said to read it back rather than assume it. R-343 is therefore filed OPEN, not CLOSED, per the task's own condition.


1. Confirmed baselines

item value
felhom.eu main @ start f267bc047f198de4cb600068fdd8bcef557cff20 — matches the sheet
clean tree at start yes; HEAD == origin/main in felhom.eu, felhom-controller, felhom-agent
scripts/ version IN felhom-host-install.sh v1.28.0 (CHANGELOG head)
scripts/ version OUT due_checks_gate.py v1.0.0 (new head entry)

2. Files created / modified

Created: scripts/due_checks_gate.py, scripts/test_due_checks_gate.py, REPORT-register-and-floor.md. Modified: scripts/repo_gates.py (registration + docstring), scripts/CHANGELOG.md, CLAUDE.md, CONTEXT.md, STATUS.md, documentation/backlog/OPEN-ITEMS.md (block + R-342 + R-343), documentation/runbooks/publish-train-rules.md. Commit hashes are in §9.

3. Tests — 37 assertions, and a red-proof that caught my own test

python3 scripts/test_due_checks_gate.py → passed: 37, failed: 0, groups A–G.

Red-proof 1 — the boundary. It failed usefully: it exposed a HOLLOW assertion of mine.

Mutation: due_now = [... if r[1] <= today] → < today. First run, before the fix: Group C reported

  PASS  C: due TODAY exits 1  rc=1          <-- passed, and should NOT have
  FAIL  C: says DUE TODAY rather than overdue

The rc == 1 assertion passed for the wrong reason. With <, a row dated exactly today falls into neither due_now (<) nor pending (>), so min(pending, …) raised ValueError: min() iterable argument is empty and the traceback exited 1. Confirmed directly:

  File ".../due_checks_gate.py", line 232, in main
    nearest = min(pending, key=lambda r: r[1])
ValueError: min() iterable argument is empty
RC=1

An exit code alone cannot distinguish a verdict from a crash. Two fixes, both kept:

  1. the test now asserts DUE-CHECKS GATE FAILED is in the output and Traceback is not, plus a new test_c_gate_never_ends_in_a_traceback across overdue/future/empty inputs;
  2. the gate returns 2 (INCONCLUSIVE) with a message if the partition is ever broken again, because a crash is never a verdict.

Re-run after the fix — the mutation now bites properly:

  FAIL  C: due TODAY exits 1 (boundary is <=)  rc=2
  FAIL  C: exits 1 as a VERDICT, not a traceback
  PASS  C: did not crash
  FAIL  C: says DUE TODAY rather than overdue
passed: 34   failed: 3

Mutation reverted, verified by grep -n "MUTATED" returning nothing and the <= line restored.

Red-proof 2 — the missing-block path

Mutation: the missing-block branch sys.exit(2) → sys.exit(0). Seen failing:

  FAIL  E: missing block exits 2 (NOT 0)  rc=0
passed: 36   failed: 1

Reverted, sys.exit(2) restored on that branch.

4. The gate's real output in all three states

Overdue (fixture, today=2026-08-20):

DUE-CHECKS GATE FAILED: 1 dated check(s) are due or overdue as of 2026-08-20 (UTC).

  R-341    due 2026-08-19   1 day(s) OVERDUE
           measure: ep0 proxy fd count + ESTAB/CLOSE-WAIT split; PID must still be 551655
           the command and its preconditions are in the R-341 row of documentation/backlog/OPEN-ITEMS.md

Take the measurement, record the result in that R-row, then remove the row from the
DUE-CHECKS block. Moving the date instead is allowed — state the reason in the R-row.
NOTE: this gate fires on a PUSH, not on the date; it may be later than the date.

Pending / the LIVE run against the real register today (these are the same run):

due-checks gate OK — 2 dated check(s) pending, none due yet.
  today (UTC): 2026-08-18
  nearest: R-341 due 2026-08-19 (in 1 day(s)) — ep0 proxy fd count + ESTAB/CLOSE-WAIT split; PID must still be 551655
  (fires on the next PUSH after a date passes, not on the date itself — by design)

5. The runner's output

  site               OK            (exit 0)
  hostinstall        OK            (exit 0)
  hub-confirm        OK            (exit 0)
  manifest-bearer    OK            (exit 0)
  reuse-refs         OK            (exit 0)
  instructions       OK            (exit 0)
  golden-currency    OK            (exit 0)
  wire-contract      OK            (exit 0)
  hub-copy           OK            (exit 0)
  due-checks         OK            (exit 0)
all felhom.eu gates OK

Group G asserts registration by running the runner and matching due-checks in its output, never by grepping repo_gates.py's source — a commented-out entry still contains the string.

6. Part 3's five reads — evidence, not summary

READ 1 — the live floor, from the store

  artifact_agent_version     = '0.129.0'    (updated 2026-08-18 11:00:59)
  artifact_golden_version    = '0.216.0'    (updated 2026-08-18 11:00:59)
  artifact_min_agent         = '0.129.0'    (updated 2026-08-18 11:01:00)
  min_controller_version     = '0.216.0'    (updated 2026-08-18 12:36:58)

The raise landed, so Part 3 proceeded. Read from hub_settings, not the form.

READ 2 — per-customer overrides

  demo-felhom      status=active     override=''         config_version=12
  demo-hp          status=active     override=''         config_version=5
  drill-r50        status=blocked    override=''         config_version=1
  peti-felhom      status=active     override=''         config_version=6
  tester-1         status=active     override=''         config_version=1
  -> 0 customer(s) carry a non-empty override

Zero overrides, so the global applies to everyone and no box hides behind a lower one.

READ 3 — every box's controller version. Two are below the floor.

  demo-felhom      controller='0.216.0'    last_report=2026-08-18 13:07:07
  demo-hp          controller='0.216.0'    last_report=2026-08-18 13:01:34
  drill-r50        controller='0.213.0'    last_report=2026-08-12 15:33:25
  peti-felhom      controller='0.115.0'    last_report=2026-07-15 08:39:00

The two REPORTING boxes are both at 0.216.0, at the floor. The other two are below it and neither is a reporting box: drill-r50 is status=blocked, last heard from six days ago, powered off and reverted; peti-felhom's host row was deleted on 2026-07-15. Reported here rather than as a footnote, per the task's edge-case rule.

(Note: guests.controller_version is empty for every guest — the hub carries the controller version on reports.controller_version, not on the guest row. The first query I wrote read the guest field and would have reported "unknown" for every box.)

READ 4 — directives and holds

2026/08/18 14:36:58 [INFO] Global controller-version floor set to "0.216.0"

No managed floor HELD line exists — searched over 24 h of pod logs. (Hub log lines are CEST; the DB stores UTC, hence 14:36:58 here and 12:36:58 above — the same instant.)

READ 5 — THE FINDING: a controller DID auto-update after the raise

2026-08-18 12:37:07 | demo-felhom | controller_updated | Controller frissítve: 0.214.0 → 0.216.0
2026-08-18 12:37:12 | demo-felhom | controller_started | Controller elindult (0.216.0)

and the version trail confirms it:

   2026-08-18 11:14:55  controller=0.214.0
   2026-08-18 12:37:12  controller=0.216.0

demo-felhom had been on 0.214.0 since 2026-08-12 16:44 and the floor raise pulled it to 0.216.0 nine seconds after the save — exactly the "acts immediately on the next report cycle" that publish-train-rules.md rule 2 documents and that the 2026-07-11 incident was filed for. demo-hp was already on 0.216.0 (hand-deployed 2026-08-14 08:31) and did not move.

No error, warning or critical event followed — the update completed and the controller restarted. So: harmless in outcome, but not a no-op. "Every reporting box is at or above the floor" is true because of the raise, not independently of it.

R-343 is filed OPEN. The task's closing condition was all five reads clean and no directive served; read 5 shows a live box moved. It went well, and a record that called it inert would mislead the next reader.

7. peti-felhom — not contacted

The machine was not contacted in any way. Sourced from the PETI register row, quoted:

"a report from a deleted host 401s and is not persisted"

with its host row deleted 2026-07-15 08:56:22 (host_deletions id=1). It therefore cannot receive a floor directive and the raise cannot reach it. Its reports row still shows controller 0.115.0 from its last report on 2026-07-15 08:39:00 — a stale record, not a live box.

8. The two build-felhom-iso.sh facts, confirmed in the script

(a) It is a BUILD-TIME gate. assert_golden_ge_floor() is defined at :77 and called at :267, in the build flow.

(b) It FAILS OPEN with a warning when its inputs are absent — :78-82:

    local golden="${FELHOM_ASSERT_GOLDEN:-}" floor="${FELHOM_ASSERT_FLOOR:-}"
    if [[ -z "$golden" || -z "$floor" ]]; then
        log_warn "R-71 golden>=floor gate UNENFORCED — pass FELHOM_ASSERT_GOLDEN + FELHOM_ASSERT_FLOOR to enforce (golden='${golden:-unset}' floor='${floor:-unset}')"
        return 0
    fi

Both read as the task described. No ISO rebuild is required: the golden is fetched at first boot from the hub's manifest (0.216.0 — at the floor, not below it), and this gate governs future builds.

9. Commits and CI

commit contents
0a5e9b14dc84ecfb179b2654d079c2f6d3f15fe2 the gate, its tests, registration, both register rows, the block, and all §5 documentation

CI run 355, head_sha 0a5e9b14d, conclusion success (started 2026-08-18T13:17:04Z). The previous run 354 on f267bc047 was also green, so this run's green is attributable to this change rather than inherited from a red baseline — and per §13 of the task, a red run here would have been mine to own.

The push needed no --no-verify. The pre-push hook ran repo_gates.py --fast, including the new due-checks gate, and passed — so the gate has now run in its real place, in both homes, not only in its own test suite.

10. NOT yet validated

The gate has never fired on a real overdue date in the live register. Every conviction shown here is from a temp-file fixture or a FELHOM_GATE_TODAY override. Its first genuine firing will be the next push on or after 2026-08-19, when R-341's first check comes due. Until that happens, "it refuses the push" is proven in fixtures and inferred in production — the registration test proves it is wired into the runner, which is the part that could silently not be true.

Also unvalidated: the block's own upkeep. Nothing checks that a row removed from the block was removed because the measurement was taken rather than because it was inconvenient.

11. Teardown

This task provisioned nothing. No VM, no container, no machine touched. The hub was read-only — snapshots of hub.db + -wal were taken into the session scratchpad for querying and are not committed.

12. Register rows

The block, verbatim as committed:

<!-- DUE-CHECKS-BEGIN — machine-readable. Parsed by scripts/due_checks_gate.py.
     One row per dated check. The R-number must have a row above. Dates are UTC.
     Clearing a row means the check was DONE and its result recorded in that R-row —
     or the date was deliberately moved, with the reason stated in the R-row.
     This block is an INDEX, not the detail: the command and the preconditions live in
     the R-row. Duplicating them here would create the second source this design avoids. -->
| item | due (UTC) | what to measure |
|---|---|---|
| R-341 | 2026-08-19 | ep0 proxy fd count + ESTAB/CLOSE-WAIT split; PID must still be 551655 |
| R-341 | 2026-08-25 | same, +7 d |
<!-- DUE-CHECKS-END -->

R-341's dates in the register matched the sheet exactly — no disagreement to report.

R-343's verdict cell, verbatim:

OPEN — NEW 2026-08-18. Deliberately NOT closed: the task's closing condition was all five reads clean, no directive served, and read 5 shows a live box updated. It went cleanly and is the floor working as designed — but a change recorded as a no-op when it moved a customer box is exactly the kind of record that misleads later

R-342 filed READY (S), owner Viktor decides; CC executes, quoting stop2-snapshot.txt verbatim on what the snapshot covers and does not.

13. unproven.py --summary

where felhom stands — 55 claims, verified_on 2026-08-09
  walked   23
  partial  14   (6 cite evidence, 8 prose only)
  built    14   (0 cite evidence, 14 prose only)
  missing  4   (0 cite evidence, 4 prose only)
  NOT WALKED: 32 of 55

No number moved, correctly: this task added a gate and three register facts, and walked no claim in the standing picture.

14. Observations — noticed, deliberately not acted on

  • CLAUDE.md's gate list named only 6 of the 10 registered gates. It was missing golden_currency_gate.py, wire_contract_gate.py and hub_copy_gate.py — all registered weeks ago. I completed the list rather than appending a 7th name to a list that was already wrong, since the section's stated job is to name each gate. Effective line count 124 → 128 against a ceiling of 200, so no trim was needed.
  • documentation/architecture/00-capability-map.md — no change, and this is the explicit statement the task asked for. No row's evidence citation names the floor or golden version: line 153's publish-train row cites a runbook path, and line 44's golden literal is a dated historical citation on the recovery-journey row.
  • drill-r50 will be dragged 0.213.0 → 0.216.0 by this floor if it is ever booted and reports. Its agent (0.129.0) meets MinAgent, so the floor would be served, not held. That is the floor doing its job; noted so it is not read as a surprise later.
  • The hub's log timestamps are CEST while its DB stores UTC. Not a defect, but it makes a log line and an events row for the same instant look two hours apart, which is worth knowing before correlating them under pressure.
  • min_controller_version and artifact_* live in the same hub_settings table but are saved by different actions, two hours apart today (11:00:59 vouch, 12:36:58 floor). That separation is rule 2 working, and is why reading only one of them would give a misleading picture.