Files
felhom.eu/REPORT-register-and-floor.md
T
admin 0a5e9b14dc
gates / gates (push) Successful in 14s
due-checks gate (R-341), floor raise recorded (R-343), snapshot coverage (R-342)
PART 1+2 — dated checks stop being wishes. R-341 booked two measurements as
prose in a register row; nothing read those dates and nothing would have
objected when they passed. The dates now live in a DUE-CHECKS block INSIDE
OPEN-ITEMS.md (inside, so no sidecar can drift from it) and a new gate reads
them. Registered as #10 in repo_gates.py, --fast, so it runs in BOTH the
pre-push hook and CI.

  exit 0  nothing due (prints pending count + nearest date; empty block too)
  exit 1  a row is due/overdue (due <= today, UTC -- due TODAY counts), or a
          row names an item with no R-row
  exit 2  block absent/duplicated/unparseable -- INCONCLUSIVE, never 0

It REFUSES rather than warns, and its docstring states the limitation: it is
NOT a scheduler, it fires on the next push, not on the date.

37 tests. BOTH red-proofs run and reverted -- and the first one earned its
keep by catching a hollow assertion of MINE rather than confirming the gate:
flipping <= to < left a due-today row in neither bucket, min() raised on an
empty list, and the TRACEBACK exited 1, so "rc == 1" passed while the
boundary was wrong. An exit code cannot tell a verdict from a crash. The test
now asserts the conviction banner and the absence of a traceback, and the gate
returns 2 rather than crashing if that partition breaks again.

PART 3 — the floor raise, and the premise was WRONG. Read back from the store
(not the form): min_controller_version = 0.216.0 @ 12:36:58Z, zero
per-customer overrides, no "managed floor HELD" line. But read 5 shows the
raise was NOT a no-op: demo-felhom had been on 0.214.0 since 12 Aug and
auto-updated 0.214.0 -> 0.216.0 at 12:37:07Z -- NINE SECONDS after the save,
exactly the immediate action publish-train rule 2 documents. No error events
followed; it restarted clean.

R-343 is therefore filed OPEN, not CLOSED: the closing condition was all five
reads clean and no directive served. It went well, but a record calling it
inert when it moved a customer box is what misleads the next reader. The row
also states why the floor was behind -- rule 2 policy, not drift, earned by
the 2026-07-11 skew onto Peti's box -- and cites ResolveManagedFloor
(store.go:2068) plus the two build-felhom-iso.sh facts (build-time at :267,
fails open at :78-82) rather than asserting them.

Two boxes are below the floor and neither reports: drill-r50 (blocked,
powered off) and peti-felhom (host row deleted). peti-felhom was NOT
contacted -- its row records that a report from a deleted host 401s and is
not persisted, so the raise cannot reach it.

PART 4 — R-342 filed READY, quoting stop2-snapshot.txt verbatim: Hetzner
server snapshot 421440873 covers /dev/sda only; /mnt/pbs-datastore is a
separate Volume that snapshots exclude, so a rollback restores software state
and NOT the datastore. Fine for that upgrade; the safeguard for any future
procedure that could touch the datastore does not exist and is Viktor's call.

Also: CLAUDE.md's gate list named 6 of 10 registered gates -- completed
rather than appending a 7th to a wrong list (124 -> 128 effective, ceiling
200). Capability map deliberately unchanged; no row cites a floor or golden
version. repo_gates.py fully green, 10/10.
2026-08-18 15:16:51 +02:00

318 lines
14 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# REPORT — dated checks that bite, the floor raise on the record, the snapshot that covers less (2026-08-18)
**Three pieces of bookkeeping, no machine put at risk. The hub was READ ONLY throughout.**
**The headline is that Part 3's premise was wrong.** The floor raise was *not* a no-op: it moved a
live customer box nine seconds after the save. That is the whole reason the task said to read it back
rather than assume it. **R-343 is therefore filed OPEN, not CLOSED**, per the task's own condition.
---
## 1. Confirmed baselines
| item | value |
|---|---|
| felhom.eu `main` @ start | `f267bc047f198de4cb600068fdd8bcef557cff20` — matches the sheet |
| clean tree at start | yes; `HEAD == origin/main` in felhom.eu, felhom-controller, felhom-agent |
| `scripts/` version IN | `felhom-host-install.sh v1.28.0` (CHANGELOG head) |
| `scripts/` version OUT | `due_checks_gate.py v1.0.0` (new head entry) |
## 2. Files created / modified
**Created:** `scripts/due_checks_gate.py`, `scripts/test_due_checks_gate.py`,
`REPORT-register-and-floor.md`.
**Modified:** `scripts/repo_gates.py` (registration + docstring), `scripts/CHANGELOG.md`, `CLAUDE.md`,
`CONTEXT.md`, `STATUS.md`, `documentation/backlog/OPEN-ITEMS.md` (block + R-342 + R-343),
`documentation/runbooks/publish-train-rules.md`.
Commit hashes are in §9.
## 3. Tests — 37 assertions, and a red-proof that caught my own test
`python3 scripts/test_due_checks_gate.py` → **passed: 37, failed: 0**, groups A–G.
### Red-proof 1 — the boundary. **It failed usefully: it exposed a HOLLOW assertion of mine.**
**Mutation:** `due_now = [... if r[1] <= today]` → `< today`.
**First run, before the fix:** Group C reported
```
PASS C: due TODAY exits 1 rc=1 <-- passed, and should NOT have
FAIL C: says DUE TODAY rather than overdue
```
**The `rc == 1` assertion passed for the wrong reason.** With `<`, a row dated exactly today falls
into neither `due_now` (`<`) nor `pending` (`>`), so `min(pending, …)` raised
`ValueError: min() iterable argument is empty` and the **traceback** exited 1. Confirmed directly:
```
File ".../due_checks_gate.py", line 232, in main
nearest = min(pending, key=lambda r: r[1])
ValueError: min() iterable argument is empty
RC=1
```
**An exit code alone cannot distinguish a verdict from a crash.** Two fixes, both kept:
1. the test now asserts `DUE-CHECKS GATE FAILED` is in the output **and** `Traceback` is not, plus a
new `test_c_gate_never_ends_in_a_traceback` across overdue/future/empty inputs;
2. the gate returns **2 (INCONCLUSIVE)** with a message if the partition is ever broken again,
because a crash is never a verdict.
**Re-run after the fix — the mutation now bites properly:**
```
FAIL C: due TODAY exits 1 (boundary is <=) rc=2
FAIL C: exits 1 as a VERDICT, not a traceback
PASS C: did not crash
FAIL C: says DUE TODAY rather than overdue
passed: 34 failed: 3
```
**Mutation reverted**, verified by `grep -n "MUTATED"` returning nothing and the `<=` line restored.
### Red-proof 2 — the missing-block path
**Mutation:** the missing-block branch `sys.exit(2)` → `sys.exit(0)`.
**Seen failing:**
```
FAIL E: missing block exits 2 (NOT 0) rc=0
passed: 36 failed: 1
```
**Reverted**, `sys.exit(2)` restored on that branch.
## 4. The gate's real output in all three states
**Overdue (fixture, today=2026-08-20):**
```
DUE-CHECKS GATE FAILED: 1 dated check(s) are due or overdue as of 2026-08-20 (UTC).
R-341 due 2026-08-19 1 day(s) OVERDUE
measure: ep0 proxy fd count + ESTAB/CLOSE-WAIT split; PID must still be 551655
the command and its preconditions are in the R-341 row of documentation/backlog/OPEN-ITEMS.md
Take the measurement, record the result in that R-row, then remove the row from the
DUE-CHECKS block. Moving the date instead is allowed — state the reason in the R-row.
NOTE: this gate fires on a PUSH, not on the date; it may be later than the date.
```
**Pending / the LIVE run against the real register today (these are the same run):**
```
due-checks gate OK — 2 dated check(s) pending, none due yet.
today (UTC): 2026-08-18
nearest: R-341 due 2026-08-19 (in 1 day(s)) — ep0 proxy fd count + ESTAB/CLOSE-WAIT split; PID must still be 551655
(fires on the next PUSH after a date passes, not on the date itself — by design)
```
## 5. The runner's output
```
site OK (exit 0)
hostinstall OK (exit 0)
hub-confirm OK (exit 0)
manifest-bearer OK (exit 0)
reuse-refs OK (exit 0)
instructions OK (exit 0)
golden-currency OK (exit 0)
wire-contract OK (exit 0)
hub-copy OK (exit 0)
due-checks OK (exit 0)
all felhom.eu gates OK
```
Group G asserts registration by **running the runner** and matching `due-checks` in its output, never
by grepping `repo_gates.py`'s source — a commented-out entry still contains the string.
## 6. Part 3's five reads — evidence, not summary
### READ 1 — the live floor, from the store
```
artifact_agent_version = '0.129.0' (updated 2026-08-18 11:00:59)
artifact_golden_version = '0.216.0' (updated 2026-08-18 11:00:59)
artifact_min_agent = '0.129.0' (updated 2026-08-18 11:01:00)
min_controller_version = '0.216.0' (updated 2026-08-18 12:36:58)
```
**The raise landed**, so Part 3 proceeded. Read from `hub_settings`, not the form.
### READ 2 — per-customer overrides
```
demo-felhom status=active override='' config_version=12
demo-hp status=active override='' config_version=5
drill-r50 status=blocked override='' config_version=1
peti-felhom status=active override='' config_version=6
tester-1 status=active override='' config_version=1
-> 0 customer(s) carry a non-empty override
```
**Zero overrides**, so the global applies to everyone and no box hides behind a lower one.
### READ 3 — every box's controller version. **Two are below the floor.**
```
demo-felhom controller='0.216.0' last_report=2026-08-18 13:07:07
demo-hp controller='0.216.0' last_report=2026-08-18 13:01:34
drill-r50 controller='0.213.0' last_report=2026-08-12 15:33:25
peti-felhom controller='0.115.0' last_report=2026-07-15 08:39:00
```
**The two REPORTING boxes are both at 0.216.0, at the floor.** The other two are below it and neither
is a reporting box: `drill-r50` is `status=blocked`, last heard from six days ago, powered off and
reverted; `peti-felhom`'s host row was deleted on 2026-07-15. Reported here rather than as a
footnote, per the task's edge-case rule.
*(Note: `guests.controller_version` is empty for every guest — the hub carries the controller version
on `reports.controller_version`, not on the guest row. The first query I wrote read the guest field
and would have reported "unknown" for every box.)*
### READ 4 — directives and holds
```
2026/08/18 14:36:58 [INFO] Global controller-version floor set to "0.216.0"
```
**No `managed floor HELD` line exists** — searched over 24 h of pod logs. (Hub log lines are CEST;
the DB stores UTC, hence 14:36:58 here and 12:36:58 above — the same instant.)
### READ 5 — **THE FINDING: a controller DID auto-update after the raise**
```
2026-08-18 12:37:07 | demo-felhom | controller_updated | Controller frissítve: 0.214.0 → 0.216.0
2026-08-18 12:37:12 | demo-felhom | controller_started | Controller elindult (0.216.0)
```
and the version trail confirms it:
```
2026-08-18 11:14:55 controller=0.214.0
2026-08-18 12:37:12 controller=0.216.0
```
**`demo-felhom` had been on 0.214.0 since 2026-08-12 16:44 and the floor raise pulled it to 0.216.0
nine seconds after the save** — exactly the "acts immediately on the next report cycle" that
`publish-train-rules.md` rule 2 documents and that the 2026-07-11 incident was filed for.
`demo-hp` was already on 0.216.0 (hand-deployed 2026-08-14 08:31) and did not move.
**No error, warning or critical event followed** — the update completed and the controller restarted.
So: harmless in outcome, but **not a no-op**. "Every reporting box is at or above the floor" is true
**because of** the raise, not independently of it.
**R-343 is filed OPEN.** The task's closing condition was *all five reads clean and no directive
served*; read 5 shows a live box moved. It went well, and a record that called it inert would mislead
the next reader.
## 7. `peti-felhom` — not contacted
**The machine was not contacted in any way.** Sourced from the PETI register row, quoted:
> *"a report from a deleted host 401s and is not persisted"*
with its host row deleted `2026-07-15 08:56:22` (`host_deletions` id=1). It therefore cannot receive a
floor directive and the raise cannot reach it. Its `reports` row still shows controller 0.115.0 from
its last report on 2026-07-15 08:39:00 — a stale record, not a live box.
## 8. The two `build-felhom-iso.sh` facts, confirmed in the script
**(a) It is a BUILD-TIME gate.** `assert_golden_ge_floor()` is defined at **`:77`** and called at
**`:267`**, in the build flow.
**(b) It FAILS OPEN with a warning when its inputs are absent** — `:78-82`:
```bash
local golden="${FELHOM_ASSERT_GOLDEN:-}" floor="${FELHOM_ASSERT_FLOOR:-}"
if [[ -z "$golden" || -z "$floor" ]]; then
log_warn "R-71 golden>=floor gate UNENFORCED — pass FELHOM_ASSERT_GOLDEN + FELHOM_ASSERT_FLOOR to enforce (golden='${golden:-unset}' floor='${floor:-unset}')"
return 0
fi
```
Both read as the task described. **No ISO rebuild is required:** the golden is fetched at first boot
from the hub's manifest (0.216.0 — at the floor, not below it), and this gate governs *future* builds.
## 10. NOT yet validated
**The gate has never fired on a real overdue date in the live register.** Every conviction shown here
is from a temp-file fixture or a `FELHOM_GATE_TODAY` override. Its first genuine firing will be the
next push on or after **2026-08-19**, when R-341's first check comes due. Until that happens, "it
refuses the push" is proven in fixtures and *inferred* in production — the registration test proves it
is wired into the runner, which is the part that could silently not be true.
Also unvalidated: the block's own upkeep. Nothing checks that a row removed from the block was removed
because the measurement was *taken* rather than because it was inconvenient.
## 11. Teardown
**This task provisioned nothing.** No VM, no container, no machine touched. The hub was read-only —
snapshots of `hub.db` + `-wal` were taken into the session scratchpad for querying and are not
committed.
## 12. Register rows
**The block, verbatim as committed:**
```markdown
<!-- DUE-CHECKS-BEGIN — machine-readable. Parsed by scripts/due_checks_gate.py.
One row per dated check. The R-number must have a row above. Dates are UTC.
Clearing a row means the check was DONE and its result recorded in that R-row —
or the date was deliberately moved, with the reason stated in the R-row.
This block is an INDEX, not the detail: the command and the preconditions live in
the R-row. Duplicating them here would create the second source this design avoids. -->
| item | due (UTC) | what to measure |
|---|---|---|
| R-341 | 2026-08-19 | ep0 proxy fd count + ESTAB/CLOSE-WAIT split; PID must still be 551655 |
| R-341 | 2026-08-25 | same, +7 d |
<!-- DUE-CHECKS-END -->
```
**R-341's dates in the register matched the sheet exactly** — no disagreement to report.
**R-343's verdict cell, verbatim:**
> **OPEN — NEW 2026-08-18.** Deliberately NOT closed: the task's closing condition was *all five reads
> clean, no directive served*, and read 5 shows a live box updated. It went cleanly and is the floor
> working as designed — but a change recorded as a no-op when it moved a customer box is exactly the
> kind of record that misleads later
**R-342** filed **READY (S)**, owner *Viktor decides; CC executes*, quoting `stop2-snapshot.txt`
verbatim on what the snapshot covers and does not.
## 13. `unproven.py --summary`
```
where felhom stands — 55 claims, verified_on 2026-08-09
walked 23
partial 14 (6 cite evidence, 8 prose only)
built 14 (0 cite evidence, 14 prose only)
missing 4 (0 cite evidence, 4 prose only)
NOT WALKED: 32 of 55
```
**No number moved**, correctly: this task added a gate and three register facts, and walked no claim
in the standing picture.
## 14. Observations — noticed, deliberately not acted on
- **`CLAUDE.md`'s gate list named only 6 of the 10 registered gates.** It was missing
`golden_currency_gate.py`, `wire_contract_gate.py` and `hub_copy_gate.py` — all registered weeks
ago. I completed the list rather than appending a 7th name to a list that was already wrong, since
the section's stated job is to name each gate. Effective line count 124 → 128 against a ceiling of
200, so no trim was needed.
- **`documentation/architecture/00-capability-map.md` — no change, and this is the explicit
statement the task asked for.** No row's evidence citation names the floor or golden *version*:
line 153's publish-train row cites a runbook path, and line 44's golden literal is a dated
historical citation on the recovery-journey row.
- **`drill-r50` will be dragged 0.213.0 → 0.216.0 by this floor if it is ever booted and reports.**
Its agent (0.129.0) meets `MinAgent`, so the floor would be served, not held. That is the floor
doing its job; noted so it is not read as a surprise later.
- **The hub's log timestamps are CEST while its DB stores UTC.** Not a defect, but it makes a log
line and an events row for the same instant look two hours apart, which is worth knowing before
correlating them under pressure.
- **`min_controller_version` and `artifact_*` live in the same `hub_settings` table but are saved by
different actions**, two hours apart today (11:00:59 vouch, 12:36:58 floor). That separation is
rule 2 working, and is why reading only one of them would give a misleading picture.