Files
felhom.eu/CLAUDE.md
T
admin 574f5df107
gates / gates (push) Failing after 17s
the decoy sweep: 29 gates read, 16 fooled, 10 fixed - and a gate that refuses the next one (R-421)
THE CLASS, now a row: an instrument that matches a LABEL rather than the fact it names. Five
instances - R-410, R-400, R-378, R-419, R-94 - and EVERY ONE was found by accident, by someone
looking at something else. The gates enforce every other rule in this project, including the rule
that findings must be written down rather than left in prose. Nothing had ever checked the gates.

METHOD, and it is the transferable part: for each gate, construct the label WITHOUT the fact - a
directory with the right name and no bake log, a handler case that exists only in a comment, a note
whose prose mentions the marker it lacks - run the gate, record what it says. No verdict was reached
by reading. Reading is how all five hid.

RESULT: 29 distinct scripts (35 registrations; three are shared across three runners). 19 sound, 4
holes left OPEN with rows, 6 that no plausible decoy could be built for and are named UNTESTED rather
than called sound. A gate nobody tried to fool is UNKNOWN.

SCOPE IS A FACT TOO - the largest single cause, and mundane. Eight gates decided what to look at with
os.listdir, one level. Every one was green AND CORRECT today, and every one would have gone blind the
moment anyone added a subdirectory. mojibake and docker-v already used os.walk, caught the identical
planted file, and are the control that proves the cause was the listing and not the decoy.

IN THIS REPO: hub-confirm and manifest-bearer now walk. observations_gate (R-419, CLOSED) requires a
marker at a line start or after a sentence boundary and strips inline code spans - a note SAYING it
carries no marker no longer satisfies the marker test. closed-register now CONVICTS on a row it
cannot parse instead of warning: FOUR rows were in that state, TWO of them written by the session
that closed them the day before, and every one was exempt from the only check that reads that file.
The rows were repaired first and the conviction added second - registering a failing gate refuses
every push.

THE META-GATE: decoy_coverage_gate.py refuses a gate registered without a decoy or a named exemption.
It convicted ITSELF the moment it was registered, which is how it came to have one. Coverage is a
DECLARATION the gate AST-parses, never a grep - searching a test file for a gate's name would be the
very shape this sweep exists to find. The 20 uncovered gates are listed by name (R-426).

NOT FIXED, each with a row and a decoy asserting TODAY's behaviour so the fix must be deliberate:
R-422 reuse-refs (only 7 extensions; a rotted .md citation is invisible), R-423 site (PAGES is a
hardcoded list of 7), R-424 one-register (a defect parked as `idea`), R-425 offbox-rename (fixed
FILES list). R-427: closed_register_gate checks ONE direction - twelve open rows carry a closed
verdict and were NOT moved, because telling finished from partly-finished is a judgement and R-378
is the record of a machine getting it wrong.

FIVE DECOYS WITHDRAWN AS ILLEGITIMATE, mine, named in the audit. A decoy nobody would write proves
nothing, and manufacturing a finding to fill a row is worse than an honest NO.

No product code. No version bump. No image. No golden owed. All four runners green.
Register: OPEN 172 -> 178, CLOSED 160 -> 161.
2026-09-01 12:39:45 +02:00

10 KiB

CLAUDE.md — felhom.eu

Stable orientation only — current state lives in CONTEXT.md and the tops of hub/CHANGELOG.md / scripts/CHANGELOG.md / website/CHANGELOG.md, never here. Cross-repo conventions (the three-component model, artifact taxonomy, access, clean-tree gate, secrets): workspace-root /mnt/5_hdd/felhom.eu/git/CLAUDE.md, whose versioned copy is documentation/runbooks/workspace-CLAUDE.md. Path-scoped detail: .claude/rules/.

What this repo is

Four surfaces in one repo, plus the design home for the whole system:

  • hub/ — felhom-hub, the operator backend (Go, k3s, hub.felhom.eu).
  • website/ — static HTML at felhom.eu, served by k3s nginx + git-sync.
  • manifests/ — k3s manifests for felhom-system, GitOps via one ArgoCD app.
  • scripts/ — the public installer (felhom-host-install.sh) and this repo's gates.
  • documentation/ — the authoritative design home for all of Felhom, not just this repo.
  • skills/ — versioned source of the Claude Code skills; install with python3 scripts/install_skills.py (symlink — repo edits are live immediately).

Doing X → read Y

Doing Read
writing any new code REUSE.md — helpers, seams, extension points, traps
needing current state / roadmap CONTEXT.md
hub work (architecture, deploy, patterns) loads itself: .claude/rules/hub.md
website or installer work loads itself: .claude/rules/website.md
manifests / ArgoCD / secrets loads itself: .claude/rules/manifests.md
writing or routing a document loads itself: .claude/rules/docs.md
build, deploy, publish, verify a version the felhom-build-deploy skill
writing or reviewing a test, fixing a bug the felhom-testing skill
UI, tokens, badges, Hungarian copy the felhom-ui-design skill
host addresses, break-glass, node facts documentation/operations/nodes.md — never restate them
which box may I break documentation/runbooks/target-selection.md
what version is live anywhere ask the hub (/hosts, /configs) or the box — never a doc
the authoritative design documentation/architecture/01..05-*.md

Code quality

  • If you need more input or troubleshooting output, ask first — don't guess.
  • A go test -run pattern that matches no test prints ok and exits 0. A red-proof using -run must first prove the filter matched something (-v, look for === RUN). Generally: an instrument that can drop results silently is not a measurement.

The installer publishes by TAG, not by push (R-110)

This fence is in the core deliberately: its trigger is editing scripts/felhom-host-install.sh, and no path-scoped rule covers that file. It governs the one artifact that runs as root on a virgin box.

  • Pushing scripts/felhom-host-install.sh to main publishes NOTHING. manifests/webpage.yaml runs two git-syncs: the website from main, and /scripts/ from the tag installer-v<SCRIPT_VERSION>.
  • To publish: cut installer-v<new SCRIPT_VERSION>, bump the --ref in webpage.yaml (both the sidecar and the init container), commit, sync.
  • To roll back: move the tag back and wait ~30 s. No ArgoCD sync, no deploy — that is the emergency lever; fix forward afterwards.
  • Do NOT pin the website to the tag, and the URL never carries a ref — felhom-bootstrap.sh and the hub's day-0 command follow the tag with no edit.
  • hostinstall_gates.py gate 6 fails if the manifest stops naming an installer-v… tag or if the website stops tracking main.

Workflow — what is specific to this repo

  • Never git add -A here — parallel sessions share the clone and it sweeps foreign WIP. Stage explicit paths only, git pull --rebase before every push.
  • REPORT.md is overwritten, so two sessions in this repo clobber each other. The second session writes REPORT-<topic>.md and never touches the shared REPORT.md.
  • CHANGELOG.md here is per-area: hub/, scripts/, website/.

Gates — ONE entry point

Run python3 scripts/repo_gates.py after ANY change in this repo. It runs every gate — site_gates.py, hostinstall_gates.py, hub_confirm_gate.py, manifest_bearer_gate.py, reuse_refs_check.py, instructions_gate.py, golden_currency_gate.py, wire_contract_gate.py, hub_copy_gate.py and due_checks_gate.py — streaming each gate's own output and exiting non-zero if any fails. --fast selects the gates that touch no network and no container runtime; today that is all of them. A missing gate script is a FAILURE, never a skip.

due_checks_gate.py refuses the push when a dated check in OPEN-ITEMS.md's DUE-CHECKS block has come due (R-341). It is not a scheduler — it fires on the next push, not on the date.

A gate ships with a decoy test that has been seen to fail (R-421). A decoy is the LABEL without the FACT — a directory with the right name and no bake log, a note whose prose mentions the marker it lacks. scripts/decoy_coverage_gate.py refuses a new gate that has neither a decoy nor a named exemption carrying its row. The four shapes, the 2026-09-01 sweep that fooled 16 of 29 gates, and the decoys withdrawn as illegitimate: documentation/audits/AUDIT-gate-decoys-2026-09-01.md and felhom-controller/.claude/rules/gates.md. Scope is a fact too — prefer os.walk over os.listdir, and a glob over a hand-maintained list.

site_gates.py is a gate, not a runner — do not model new work on it; app-catalog-felhom.eu/scripts/catalog_gates.py is the canonical runner (R-161).

The pre-push hook (.githooks/pre-push) runs it with --fast and refuses a failing push. It is per-clone — switch it on once with git config core.hooksPath .githooks, and a manual run WARNS when this clone is unarmed. git push --no-verify bypasses it deliberately; say so in the session report when you use it — CI re-runs the same entry point on every push and emails the operator on failure, so a bypass is noticed even though it is not blocked (R-168, CLOSED 2026-08-02; CI reports rather than refuses because there is no PR to gate — R-169).

End-of-session checklist

Registers first — a finding goes in documentation/backlog/OPEN-ITEMS.md first, never only in a report, an audit or STATUS.md. Four items in this project were minted in a spike doc and lost (R-153/154/155, R-156/157). This applies to every session that ships, breaks or decides something, not only sessions that touch documentation/ — which is why it is here and not in docs.md.

  • CHANGELOG.md + REPORT.md in every repo touched (see the workspace root for the rule, and the parallel-session caveat above).

  • REUSE.md, if a shared helper or pattern moved (same commit).

  • OPEN-ITEMS.md — every finding, with a number.

  • Root STATUS.md — at the end of every session in which something shipped, broke or was decided. It is a view of OPEN-ITEMS.md; nothing may exist only there. One screen, written for the operator in plain language, and deliberately not CONTEXT.md.

  • The capability map (documentation/architecture/00-capability-map.md), if a capability's status changed — with its new evidence citation.

  • python3 scripts/unproven.py --summary — one line per status, and the not-walked total. Run it at the end of any session that shipped, broke or proved something, and say in the report if a number moved. It exists because "which claims are unproven?" was answerable only by a person reading a page: a session asked for "the nine grey claims" could not determine which nine and rightly refused to guess (R-326). Nine was real and answered a different question — it is the count of claims the 2026-08-09 pass DOWNGRADED. Not-walked is 35 of 55 as of 2026-09-01 -- the figure read 32 here for weeks while the tool said 35, so re-read the tool rather than this line. A status that moves without anyone noticing is how the picture stops being true.

  • Confirm your own last push's CI run went green, by run ID. CI emails on failure, which is a PUSH signal; this is the PULL check that catches a lost, filtered or unread mail. Quote the run id and its conclusion. Use the jobs endpoint and match on head_sha, never on an id (R-417, measured 2026-09-01): actions/tasks returns "conclusion": null for every run, so a session following the old recipe here quotes a conclusion it never read; its id is also offset from the jobs id for the same run (479 vs 478), and actions/runs/<n> takes a JOB id, so runs/294 cheerfully returns an unrelated job from three weeks earlier. The list is oldest-first — page to the end.

    T=$(curl -s -u "$U:$P" ".../actions/jobs?limit=1" | python3 -c 'import json,sys;print(json.load(sys.stdin)["total_count"])')
    curl -s -u "$U:$P" ".../actions/jobs?limit=50&page=$(( T/50 + 1 ))"   # then grep your own head_sha
    

    An unchecked green is an assumption, not an observation — and so is a green read off a field the API never populates.