Files
felhom.eu/CLAUDE.md
T

182 lines
12 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# CLAUDE.md — `felhom.eu`
> Stable orientation only — **current state lives in `CONTEXT.md` and the tops of
> `hub/CHANGELOG.md` / `scripts/CHANGELOG.md` / `website/CHANGELOG.md`**, never here. Cross-repo
> conventions (the three-component model, artifact taxonomy, access, clean-tree gate, secrets):
> workspace-root `/mnt/5_hdd/felhom.eu/git/CLAUDE.md`, whose versioned copy is
> `documentation/runbooks/workspace-CLAUDE.md`. Path-scoped detail: `.claude/rules/`.
## What this repo is
Four surfaces in one repo, plus the design home for the whole system:
- `hub/` — **felhom-hub**, the operator backend (Go, k3s, `hub.felhom.eu`).
- `website/` — static HTML at felhom.eu, served by k3s nginx + git-sync.
- `manifests/` — k3s manifests for felhom-system, GitOps via one ArgoCD app.
- `scripts/` — the **public installer** (`felhom-host-install.sh`) and this repo's gates.
- `documentation/` — the **authoritative design home for all of Felhom**, not just this repo.
- `skills/` — versioned source of the Claude Code skills; install with
`python3 scripts/install_skills.py` (symlink — repo edits are live immediately).
## Doing X → read Y
| Doing | Read |
|---|---|
| writing any new code | `REUSE.md` — helpers, seams, extension points, traps |
| needing current state / roadmap | `CONTEXT.md` |
| hub work (architecture, deploy, patterns) | loads itself: `.claude/rules/hub.md` |
| website or installer work | loads itself: `.claude/rules/website.md` |
| manifests / ArgoCD / secrets | loads itself: `.claude/rules/manifests.md` |
| writing or routing a document | loads itself: `.claude/rules/docs.md` |
| build, deploy, publish, verify a version | the **`felhom-build-deploy`** skill |
| writing or reviewing a test, fixing a bug | the **`felhom-testing`** skill |
| UI, tokens, badges, Hungarian copy | the **`felhom-ui-design`** skill |
| host addresses, break-glass, node facts | `documentation/operations/nodes.md` — never restate them |
| which box may I break | `documentation/runbooks/target-selection.md` |
| what version is live anywhere | ask the hub (`/hosts`, `/configs`) or the box — **never a doc** |
| the authoritative design | `documentation/architecture/01..05-*.md` |
## Code quality
- If you need more input or troubleshooting output, **ask first — don't guess**.
- **A `go test -run` pattern that matches no test prints `ok` and exits 0.** A red-proof using `-run`
must first prove the filter matched something (`-v`, look for `=== RUN`). Generally: **an
instrument that can drop results silently is not a measurement.**
<!--
go test -run: found 2026-08-02 while red-proofing — `-run TestCustomerUnified` matched nothing in the
target file and reported `ok ... 0.062s`, which was read as a passing red-proof. This is the "an
absent line is not evidence" rule aimed at the one place a false green costs most: the proof itself.
The same class bit twice that day — a `| tail -5` inside a census query silently dropped rows and
looked exactly like a real finding.
-->
## The installer publishes by TAG, not by push (R-110)
This fence is in the core deliberately: its trigger is editing `scripts/felhom-host-install.sh`, and
no path-scoped rule covers that file. It governs the one artifact that runs **as root on a virgin
box**.
- **Pushing `scripts/felhom-host-install.sh` to `main` publishes NOTHING.** `manifests/webpage.yaml`
runs two git-syncs: the website from `main`, and `/scripts/` from the tag
`installer-v<SCRIPT_VERSION>`.
- **To publish:** cut `installer-v<new SCRIPT_VERSION>`, bump the `--ref` in `webpage.yaml` (**both**
the sidecar and the init container), commit, sync.
- **To roll back:** move the tag back and wait ~30 s. No ArgoCD sync, no deploy — that is the
emergency lever; fix forward afterwards.
- **Do NOT pin the website to the tag**, and the URL never carries a ref — `felhom-bootstrap.sh` and
the hub's day-0 command follow the tag with no edit.
- `hostinstall_gates.py` gate 6 fails if the manifest stops naming an `installer-v…` tag or if the
website stops tracking `main`.
## Workflow — what is specific to this repo
- **Never `git add -A` here** — parallel sessions share the clone and it sweeps foreign WIP. Stage
explicit paths only, `git pull --rebase` before every push.
- **`REPORT.md` is overwritten, so two sessions in this repo clobber each other.** The second session
writes **`REPORT-<topic>.md`** and never touches the shared `REPORT.md`.
- `CHANGELOG.md` here is **per-area**: `hub/`, `scripts/`, `website/`.
<!-- The sweep incident: v0.47.0, commit 146d165 — a red-proof-mutated guard got swept to main. -->
## Gates — ONE entry point
**Run `python3 scripts/repo_gates.py` after ANY change in this repo.** It runs every gate —
`site_gates.py`, `hostinstall_gates.py`, `hub_confirm_gate.py`, `manifest_bearer_gate.py`,
`reuse_refs_check.py`, `instructions_gate.py`, `golden_currency_gate.py`, `wire_contract_gate.py`,
`hub_copy_gate.py` and `due_checks_gate.py` — streaming each gate's own output and exiting
non-zero if any fails. `--fast` selects the gates that touch no network and no container runtime;
today that is all of them. **A missing gate script is a FAILURE, never a skip.**
`due_checks_gate.py` refuses the push when a dated check in `OPEN-ITEMS.md`'s `DUE-CHECKS` block has
come due (R-341). **It is not a scheduler** — it fires on the next push, not on the date.
**A gate ships with a decoy test that has been seen to fail (R-421).** A decoy is the LABEL without
the FACT — a directory with the right name and no bake log, a note whose prose mentions the marker it
lacks. `scripts/decoy_coverage_gate.py` refuses a new gate that has neither a decoy nor a named
exemption carrying its row. The four shapes, the 2026-09-01 sweep that fooled 16 of 29 gates, and the
decoys withdrawn as illegitimate: `documentation/audits/AUDIT-gate-decoys-2026-09-01.md` and
`felhom-controller/.claude/rules/gates.md`. **Scope is a fact too** — prefer `os.walk` over
`os.listdir`, and a glob over a hand-maintained list.
`site_gates.py` is a *gate*, not a runner — do not model new work on it;
`app-catalog-felhom.eu/scripts/catalog_gates.py` is the canonical runner (R-161).
**The pre-push hook** (`.githooks/pre-push`) runs it with `--fast` and refuses a failing push. It is
**per-clone** — switch it on once with `git config core.hooksPath .githooks`, and a manual run WARNS
when this clone is unarmed. `git push --no-verify` bypasses it deliberately; **say so in the session
report when you use it** — CI re-runs the same entry point on every push and **emails the operator
on failure**, so a bypass is noticed even though it is not blocked (R-168, CLOSED 2026-08-02; CI
reports rather than refuses because there is no PR to gate — R-169).
<!--
WHY A RUNNER AND NOT FIVE INVOCATIONS (2026-08-02, R-29): a census of all thirteen gates across the
four repos found that every check a CLAUDE.md names was passing, and two of the four nobody is told
to run were failing — one since 14 July. The single-entry-point shape is the only one that
demonstrably gets run here.
instructions_gate joined the list on 2026-08-06 (R-229 deferred leg). It is registered in the
controller and agent runners too, and it is the SHARED copy in felhom.eu/scripts/ — never copied
into a sibling repo, because a copy would recreate the drift it detects.
-->
## End-of-session checklist
Registers first — **a finding goes in `documentation/backlog/OPEN-ITEMS.md` first**, never only in a
report, an audit or `STATUS.md`. Four items in this project were minted in a spike doc and lost
(R-153/154/155, R-156/157). This applies to **every** session that ships, breaks or decides
something, not only sessions that touch `documentation/` — which is why it is here and not in
`docs.md`.
- **`CHANGELOG.md` + `REPORT.md`** in every repo touched (see the workspace root for the rule, and
the parallel-session caveat above).
- **`REUSE.md`**, if a shared helper or pattern moved (same commit).
- **`OPEN-ITEMS.md`** — every finding, with a number. **A row you close moves to `CLOSED-ITEMS.md` in the
same commit** — `closed_register_gate.py` RULE 3 refuses a finished row left in the open register. A
**new row** goes into its category's section with one **Category** and one **Sev** (P1–P4); the scale
and the eleven names head `OPEN-ITEMS.md`, and `register_shape_gate.py` refuses anything else.
- **Root `STATUS.md`** — at the end of every session in which something shipped, broke or was
decided. It is a **view** of `OPEN-ITEMS.md`; nothing may exist only there. One screen, written for
the operator in plain language, and deliberately **not** `CONTEXT.md`.
- **The golden, on its cadence** (operator ruling 2026-09-13): **weekly, and before ANY drill or
fresh install**, bake + vouch + raise the floor per `documentation/runbooks/RUNBOOK-manual-build.md`
§4.1. Not per release. Between bakes the dated waiver (§4.2, `documentation/tests/golden-waiver.yml`,
≤ 14 days) keeps `golden_currency_gate.py` advisory; **when it expires the gate is red and stays
red until someone bakes or renews — that is the mechanism, so do not `--no-verify` past it.** A
nightly or drill session that starts on a fresh install checks the golden FIRST.
- **The capability map** (`documentation/architecture/00-capability-map.md`), if a capability's
status changed — with its new evidence citation.
- **`python3 scripts/unproven.py --summary`** — one line per status, and the not-walked total. Run it
at the end of any session that shipped, broke or proved something, and **say in the report if a
number moved**. It exists because "which claims are unproven?" was answerable only by a person
reading a page: a session asked for "the nine grey claims" could not determine which nine and
rightly refused to guess (R-326). *Nine was real and answered a different question — it is the
count of claims the 2026-08-09 pass DOWNGRADED. Not-walked is 35 of 55 as of 2026-09-01 -- the figure read 32 here for weeks while the tool said 35, so re-read the tool rather than this line.* A status that moves
without anyone noticing is how the picture stops being true.
- **Confirm your own last push's CI run went green, by run ID.** CI emails on failure, which is a
PUSH signal; this is the PULL check that catches a lost, filtered or unread mail. Quote the run id
and its conclusion. **Use the `jobs` endpoint and match on `head_sha`, never on an id** (R-417,
measured 2026-09-01): `actions/tasks` returns `"conclusion": null` for every run, so a session
following the old recipe here quotes a conclusion it never read; its `id` is also offset from the
`jobs` id for the same run (479 vs 478), and `actions/runs/<n>` takes a JOB id, so `runs/294`
cheerfully returns an unrelated job from three weeks earlier. The list is oldest-first — page to
the end.
```bash
T=$(curl -s -u "$U:$P" ".../actions/jobs?limit=1" | python3 -c 'import json,sys;print(json.load(sys.stdin)["total_count"])')
for pg in $(seq 1 $(( (T+49)/50 ))); do curl -s -u "$U:$P" ".../actions/jobs?limit=50&page=$pg"; done
# then match your own head_sha across ALL of them
```
**Two things this recipe got wrong until 2026-09-20, both measured the hard way:**
1. **The response key is `jobs`, not `workflow_runs`.** A parser reading `workflow_runs` gets an
empty list and prints nothing — which reads exactly like "no CI run for this commit" and is not.
That is R-417's own shape (a field the API never populates) committed while checking R-417.
2. **Job ids are NOT ordered within a page**, so `page = total/50 + 1` does not hold the newest
rows — a run can sit several pages earlier. **Scan every page** and match on `head_sha`; with a
few hundred jobs that is a handful of requests. Guessing the page produced three consecutive
false "no CI job for this commit" readings in one session.
An unchecked green is an assumption, not an observation — and so is a green read off a field the
API never populates.