Files
felhom.eu/CLAUDE.md
T
admin ae59c31a84 R-459 CLOSED (MariaDB converts itself, proven by harness + live), golden 0.236.0 (R-467), the golden waiver (R-468)
Operator rulings 2026-09-13, both shipped the same day:
- MariaDB finishes its own conversion (catalog eec1228/bd32830/3525e35). Harness E3/E3b `proven`
  with engine_state_after "already upgraded to 12.3.3-MariaDB [exit=1]", the skip line gone, C3
  still `failed`; landed on demo-hp through the real 15-min cycle, nothing recreated, one deliberate
  restart logged "MariaDB upgrade not required" with the app serving. Evidence:
  documentation/audits/r459-close-2026-09-13/. The engine-major rule + gate keep every engine
  inside its major until Slice 4 (R-448) — removal tracked as R-469.
- Goldens on a cadence, not per release. golden_currency_gate.py reads a dated waiver
  (documentation/tests/golden-waiver.yml, <= 14 days, row-bound): valid + BEHIND -> loud advisory,
  exit 0; expired -> red again naming the date; UNRECORDED (R-385) never covered; malformed -> 2,
  never 0. Tests cases 5-15 incl. the R-421 decoy; red-proof old-vs-new on the real behind tree.
  R-242's vouch half stays open. Cadence in RUNBOOK-manual-build.md §4.2 + the checklist.
- Golden 0.236.0 baked, round-tripped, vouched, floor raised 0.232.0 -> 0.236.0
  (documentation/tests/golden-0.236.0-2026-09-13/) — the last per-release bake; the waiver was
  issued AFTER it landed. No --no-verify anywhere in this session.

Rows: R-459 CLOSED, R-467 CLOSED, R-242 narrowed; R-468/R-469/R-470/R-471 opened. 09 §3 gains
decisions 5 and 6; STATUS items 11 and 12 closed; CONTEXT records the cadence ruling.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-13 10:14:37 +02:00

11 KiB

CLAUDE.md — felhom.eu

Stable orientation only — current state lives in CONTEXT.md and the tops of hub/CHANGELOG.md / scripts/CHANGELOG.md / website/CHANGELOG.md, never here. Cross-repo conventions (the three-component model, artifact taxonomy, access, clean-tree gate, secrets): workspace-root /mnt/5_hdd/felhom.eu/git/CLAUDE.md, whose versioned copy is documentation/runbooks/workspace-CLAUDE.md. Path-scoped detail: .claude/rules/.

What this repo is

Four surfaces in one repo, plus the design home for the whole system:

  • hub/ — felhom-hub, the operator backend (Go, k3s, hub.felhom.eu).
  • website/ — static HTML at felhom.eu, served by k3s nginx + git-sync.
  • manifests/ — k3s manifests for felhom-system, GitOps via one ArgoCD app.
  • scripts/ — the public installer (felhom-host-install.sh) and this repo's gates.
  • documentation/ — the authoritative design home for all of Felhom, not just this repo.
  • skills/ — versioned source of the Claude Code skills; install with python3 scripts/install_skills.py (symlink — repo edits are live immediately).

Doing X → read Y

Doing Read
writing any new code REUSE.md — helpers, seams, extension points, traps
needing current state / roadmap CONTEXT.md
hub work (architecture, deploy, patterns) loads itself: .claude/rules/hub.md
website or installer work loads itself: .claude/rules/website.md
manifests / ArgoCD / secrets loads itself: .claude/rules/manifests.md
writing or routing a document loads itself: .claude/rules/docs.md
build, deploy, publish, verify a version the felhom-build-deploy skill
writing or reviewing a test, fixing a bug the felhom-testing skill
UI, tokens, badges, Hungarian copy the felhom-ui-design skill
host addresses, break-glass, node facts documentation/operations/nodes.md — never restate them
which box may I break documentation/runbooks/target-selection.md
what version is live anywhere ask the hub (/hosts, /configs) or the box — never a doc
the authoritative design documentation/architecture/01..05-*.md

Code quality

  • If you need more input or troubleshooting output, ask first — don't guess.
  • A go test -run pattern that matches no test prints ok and exits 0. A red-proof using -run must first prove the filter matched something (-v, look for === RUN). Generally: an instrument that can drop results silently is not a measurement.

The installer publishes by TAG, not by push (R-110)

This fence is in the core deliberately: its trigger is editing scripts/felhom-host-install.sh, and no path-scoped rule covers that file. It governs the one artifact that runs as root on a virgin box.

  • Pushing scripts/felhom-host-install.sh to main publishes NOTHING. manifests/webpage.yaml runs two git-syncs: the website from main, and /scripts/ from the tag installer-v<SCRIPT_VERSION>.
  • To publish: cut installer-v<new SCRIPT_VERSION>, bump the --ref in webpage.yaml (both the sidecar and the init container), commit, sync.
  • To roll back: move the tag back and wait ~30 s. No ArgoCD sync, no deploy — that is the emergency lever; fix forward afterwards.
  • Do NOT pin the website to the tag, and the URL never carries a ref — felhom-bootstrap.sh and the hub's day-0 command follow the tag with no edit.
  • hostinstall_gates.py gate 6 fails if the manifest stops naming an installer-v… tag or if the website stops tracking main.

Workflow — what is specific to this repo

  • Never git add -A here — parallel sessions share the clone and it sweeps foreign WIP. Stage explicit paths only, git pull --rebase before every push.
  • REPORT.md is overwritten, so two sessions in this repo clobber each other. The second session writes REPORT-<topic>.md and never touches the shared REPORT.md.
  • CHANGELOG.md here is per-area: hub/, scripts/, website/.

Gates — ONE entry point

Run python3 scripts/repo_gates.py after ANY change in this repo. It runs every gate — site_gates.py, hostinstall_gates.py, hub_confirm_gate.py, manifest_bearer_gate.py, reuse_refs_check.py, instructions_gate.py, golden_currency_gate.py, wire_contract_gate.py, hub_copy_gate.py and due_checks_gate.py — streaming each gate's own output and exiting non-zero if any fails. --fast selects the gates that touch no network and no container runtime; today that is all of them. A missing gate script is a FAILURE, never a skip.

due_checks_gate.py refuses the push when a dated check in OPEN-ITEMS.md's DUE-CHECKS block has come due (R-341). It is not a scheduler — it fires on the next push, not on the date.

A gate ships with a decoy test that has been seen to fail (R-421). A decoy is the LABEL without the FACT — a directory with the right name and no bake log, a note whose prose mentions the marker it lacks. scripts/decoy_coverage_gate.py refuses a new gate that has neither a decoy nor a named exemption carrying its row. The four shapes, the 2026-09-01 sweep that fooled 16 of 29 gates, and the decoys withdrawn as illegitimate: documentation/audits/AUDIT-gate-decoys-2026-09-01.md and felhom-controller/.claude/rules/gates.md. Scope is a fact too — prefer os.walk over os.listdir, and a glob over a hand-maintained list.

site_gates.py is a gate, not a runner — do not model new work on it; app-catalog-felhom.eu/scripts/catalog_gates.py is the canonical runner (R-161).

The pre-push hook (.githooks/pre-push) runs it with --fast and refuses a failing push. It is per-clone — switch it on once with git config core.hooksPath .githooks, and a manual run WARNS when this clone is unarmed. git push --no-verify bypasses it deliberately; say so in the session report when you use it — CI re-runs the same entry point on every push and emails the operator on failure, so a bypass is noticed even though it is not blocked (R-168, CLOSED 2026-08-02; CI reports rather than refuses because there is no PR to gate — R-169).

End-of-session checklist

Registers first — a finding goes in documentation/backlog/OPEN-ITEMS.md first, never only in a report, an audit or STATUS.md. Four items in this project were minted in a spike doc and lost (R-153/154/155, R-156/157). This applies to every session that ships, breaks or decides something, not only sessions that touch documentation/ — which is why it is here and not in docs.md.

  • CHANGELOG.md + REPORT.md in every repo touched (see the workspace root for the rule, and the parallel-session caveat above).

  • REUSE.md, if a shared helper or pattern moved (same commit).

  • OPEN-ITEMS.md — every finding, with a number.

  • Root STATUS.md — at the end of every session in which something shipped, broke or was decided. It is a view of OPEN-ITEMS.md; nothing may exist only there. One screen, written for the operator in plain language, and deliberately not CONTEXT.md.

  • The golden, on its cadence (operator ruling 2026-09-13): weekly, and before ANY drill or fresh install, bake + vouch + raise the floor per documentation/runbooks/RUNBOOK-manual-build.md §4.1. Not per release. Between bakes the dated waiver (§4.2, documentation/tests/golden-waiver.yml, ≤ 14 days) keeps golden_currency_gate.py advisory; when it expires the gate is red and stays red until someone bakes or renews — that is the mechanism, so do not --no-verify past it. A nightly or drill session that starts on a fresh install checks the golden FIRST.

  • The capability map (documentation/architecture/00-capability-map.md), if a capability's status changed — with its new evidence citation.

  • python3 scripts/unproven.py --summary — one line per status, and the not-walked total. Run it at the end of any session that shipped, broke or proved something, and say in the report if a number moved. It exists because "which claims are unproven?" was answerable only by a person reading a page: a session asked for "the nine grey claims" could not determine which nine and rightly refused to guess (R-326). Nine was real and answered a different question — it is the count of claims the 2026-08-09 pass DOWNGRADED. Not-walked is 35 of 55 as of 2026-09-01 -- the figure read 32 here for weeks while the tool said 35, so re-read the tool rather than this line. A status that moves without anyone noticing is how the picture stops being true.

  • Confirm your own last push's CI run went green, by run ID. CI emails on failure, which is a PUSH signal; this is the PULL check that catches a lost, filtered or unread mail. Quote the run id and its conclusion. Use the jobs endpoint and match on head_sha, never on an id (R-417, measured 2026-09-01): actions/tasks returns "conclusion": null for every run, so a session following the old recipe here quotes a conclusion it never read; its id is also offset from the jobs id for the same run (479 vs 478), and actions/runs/<n> takes a JOB id, so runs/294 cheerfully returns an unrelated job from three weeks earlier. The list is oldest-first — page to the end.

    T=$(curl -s -u "$U:$P" ".../actions/jobs?limit=1" | python3 -c 'import json,sys;print(json.load(sys.stdin)["total_count"])')
    curl -s -u "$U:$P" ".../actions/jobs?limit=50&page=$(( T/50 + 1 ))"   # then grep your own head_sha
    

    An unchecked green is an assumption, not an observation — and so is a green read off a field the API never populates.