Five felhom.eu CI runs went red tonight (jobs 469/470/471/473/476) and all five were mine, every
one on step 3 `Run the gate entry point`. Job 478 is green. CAUSE CONFIRMED BY ISOLATION: the only
functional diff between the last red and the green is the golden-0.232.0 evidence directory;
moving it aside reproduces exit=1, restoring it gives exit=0, tree byte-identical after. The first
reproduction attempt used a detached worktree, where three gates go INCONCLUSIVE for want of the
sibling clones - that is the worktree, not the commit, so it is discarded rather than quoted.
THE GATE WAS RIGHT EVERY TIME. 0.231.0 and 0.232.0 were released with no golden carrying them, so
a machine installed in those hours would have received 0.230.0.
R-417 is the SHAPE, not the gate: the soak runbook forbade baking a golden that night, so red was
unavoidable and pushing the drill's own evidence needed --no-verify. The gate's failure text names
the remedy for exactly that case - record a waiver here, never a bypass - and I did not write one.
A red CI run on a drill night is now indistinguishable from a real one, which is the whole value
of the signal.
TWO INSTRUCTION DEFECTS, found by following the end-of-session checklist and being unable to:
- The CI-verification recipe cannot produce what it asks for. `actions/tasks` returns
"conclusion": null for every run, so a session following it quotes a conclusion it never read.
Its id is also offset from the `jobs` id for the same run (479 vs 478 for 63eff21a), and
`actions/runs/<n>` takes a JOB id - `runs/294` returned an unrelated job from 2026-08-10 and
looked like a valid answer. Now: the jobs endpoint, matched on head_sha, oldest-first paging.
- "Not-walked is 32 of 55" was stale; the tool says 35, and has for some time. The discrepancy is
written into the line so the next reader trusts the tool over the prose.
This session moved no claim status - unproven.py at ab8b8847~1 and at HEAD are identical.
9.5 KiB
CLAUDE.md — felhom.eu
Stable orientation only — current state lives in
CONTEXT.mdand the tops ofhub/CHANGELOG.md/scripts/CHANGELOG.md/website/CHANGELOG.md, never here. Cross-repo conventions (the three-component model, artifact taxonomy, access, clean-tree gate, secrets): workspace-root/mnt/5_hdd/felhom.eu/git/CLAUDE.md, whose versioned copy isdocumentation/runbooks/workspace-CLAUDE.md. Path-scoped detail:.claude/rules/.
What this repo is
Four surfaces in one repo, plus the design home for the whole system:
hub/— felhom-hub, the operator backend (Go, k3s,hub.felhom.eu).website/— static HTML at felhom.eu, served by k3s nginx + git-sync.manifests/— k3s manifests for felhom-system, GitOps via one ArgoCD app.scripts/— the public installer (felhom-host-install.sh) and this repo's gates.documentation/— the authoritative design home for all of Felhom, not just this repo.skills/— versioned source of the Claude Code skills; install withpython3 scripts/install_skills.py(symlink — repo edits are live immediately).
Doing X → read Y
| Doing | Read |
|---|---|
| writing any new code | REUSE.md — helpers, seams, extension points, traps |
| needing current state / roadmap | CONTEXT.md |
| hub work (architecture, deploy, patterns) | loads itself: .claude/rules/hub.md |
| website or installer work | loads itself: .claude/rules/website.md |
| manifests / ArgoCD / secrets | loads itself: .claude/rules/manifests.md |
| writing or routing a document | loads itself: .claude/rules/docs.md |
| build, deploy, publish, verify a version | the felhom-build-deploy skill |
| writing or reviewing a test, fixing a bug | the felhom-testing skill |
| UI, tokens, badges, Hungarian copy | the felhom-ui-design skill |
| host addresses, break-glass, node facts | documentation/operations/nodes.md — never restate them |
| which box may I break | documentation/runbooks/target-selection.md |
| what version is live anywhere | ask the hub (/hosts, /configs) or the box — never a doc |
| the authoritative design | documentation/architecture/01..05-*.md |
Code quality
- If you need more input or troubleshooting output, ask first — don't guess.
- A
go test -runpattern that matches no test printsokand exits 0. A red-proof using-runmust first prove the filter matched something (-v, look for=== RUN). Generally: an instrument that can drop results silently is not a measurement.
The installer publishes by TAG, not by push (R-110)
This fence is in the core deliberately: its trigger is editing scripts/felhom-host-install.sh, and
no path-scoped rule covers that file. It governs the one artifact that runs as root on a virgin
box.
- Pushing
scripts/felhom-host-install.shtomainpublishes NOTHING.manifests/webpage.yamlruns two git-syncs: the website frommain, and/scripts/from the taginstaller-v<SCRIPT_VERSION>. - To publish: cut
installer-v<new SCRIPT_VERSION>, bump the--refinwebpage.yaml(both the sidecar and the init container), commit, sync. - To roll back: move the tag back and wait ~30 s. No ArgoCD sync, no deploy — that is the emergency lever; fix forward afterwards.
- Do NOT pin the website to the tag, and the URL never carries a ref —
felhom-bootstrap.shand the hub's day-0 command follow the tag with no edit. hostinstall_gates.pygate 6 fails if the manifest stops naming aninstaller-v…tag or if the website stops trackingmain.
Workflow — what is specific to this repo
- Never
git add -Ahere — parallel sessions share the clone and it sweeps foreign WIP. Stage explicit paths only,git pull --rebasebefore every push. REPORT.mdis overwritten, so two sessions in this repo clobber each other. The second session writesREPORT-<topic>.mdand never touches the sharedREPORT.md.CHANGELOG.mdhere is per-area:hub/,scripts/,website/.
Gates — ONE entry point
Run python3 scripts/repo_gates.py after ANY change in this repo. It runs every gate —
site_gates.py, hostinstall_gates.py, hub_confirm_gate.py, manifest_bearer_gate.py,
reuse_refs_check.py, instructions_gate.py, golden_currency_gate.py, wire_contract_gate.py,
hub_copy_gate.py and due_checks_gate.py — streaming each gate's own output and exiting
non-zero if any fails. --fast selects the gates that touch no network and no container runtime;
today that is all of them. A missing gate script is a FAILURE, never a skip.
due_checks_gate.py refuses the push when a dated check in OPEN-ITEMS.md's DUE-CHECKS block has
come due (R-341). It is not a scheduler — it fires on the next push, not on the date.
site_gates.py is a gate, not a runner — do not model new work on it;
app-catalog-felhom.eu/scripts/catalog_gates.py is the canonical runner (R-161).
The pre-push hook (.githooks/pre-push) runs it with --fast and refuses a failing push. It is
per-clone — switch it on once with git config core.hooksPath .githooks, and a manual run WARNS
when this clone is unarmed. git push --no-verify bypasses it deliberately; say so in the session
report when you use it — CI re-runs the same entry point on every push and emails the operator
on failure, so a bypass is noticed even though it is not blocked (R-168, CLOSED 2026-08-02; CI
reports rather than refuses because there is no PR to gate — R-169).
End-of-session checklist
Registers first — a finding goes in documentation/backlog/OPEN-ITEMS.md first, never only in a
report, an audit or STATUS.md. Four items in this project were minted in a spike doc and lost
(R-153/154/155, R-156/157). This applies to every session that ships, breaks or decides
something, not only sessions that touch documentation/ — which is why it is here and not in
docs.md.
-
CHANGELOG.md+REPORT.mdin every repo touched (see the workspace root for the rule, and the parallel-session caveat above). -
REUSE.md, if a shared helper or pattern moved (same commit). -
OPEN-ITEMS.md— every finding, with a number. -
Root
STATUS.md— at the end of every session in which something shipped, broke or was decided. It is a view ofOPEN-ITEMS.md; nothing may exist only there. One screen, written for the operator in plain language, and deliberately notCONTEXT.md. -
The capability map (
documentation/architecture/00-capability-map.md), if a capability's status changed — with its new evidence citation. -
python3 scripts/unproven.py --summary— one line per status, and the not-walked total. Run it at the end of any session that shipped, broke or proved something, and say in the report if a number moved. It exists because "which claims are unproven?" was answerable only by a person reading a page: a session asked for "the nine grey claims" could not determine which nine and rightly refused to guess (R-326). Nine was real and answered a different question — it is the count of claims the 2026-08-09 pass DOWNGRADED. Not-walked is 35 of 55 as of 2026-09-01 -- the figure read 32 here for weeks while the tool said 35, so re-read the tool rather than this line. A status that moves without anyone noticing is how the picture stops being true. -
Confirm your own last push's CI run went green, by run ID. CI emails on failure, which is a PUSH signal; this is the PULL check that catches a lost, filtered or unread mail. Quote the run id and its conclusion. Use the
jobsendpoint and match onhead_sha, never on an id (R-417, measured 2026-09-01):actions/tasksreturns"conclusion": nullfor every run, so a session following the old recipe here quotes a conclusion it never read; itsidis also offset from thejobsid for the same run (479 vs 478), andactions/runs/<n>takes a JOB id, soruns/294cheerfully returns an unrelated job from three weeks earlier. The list is oldest-first — page to the end.T=$(curl -s -u "$U:$P" ".../actions/jobs?limit=1" | python3 -c 'import json,sys;print(json.load(sys.stdin)["total_count"])') curl -s -u "$U:$P" ".../actions/jobs?limit=50&page=$(( T/50 + 1 ))" # then grep your own head_shaAn unchecked green is an assumption, not an observation — and so is a green read off a field the API never populates.