Files
felhom.eu/CLAUDE.md
T
admin c21bcf84f7
gates / gates (push) Successful in 7s
docs+gate: instruction files cannot silently regrow (R-229)
New shared scripts/instructions_gate.py, registered in controller_gates.py and
agent_gates.py, never copied into a sibling repo (the reuse_refs_check.py
precedent). 20 fixture tests, all asserting the effect: exit code AND that the
message names the file and the reason.

It is a consistency gate, not a budget gate, and the failure message says so. A
/context reading measured the instruction files at 15k tokens against 869k free in
a 1M window -- space is not the constraint, and a future reader must not re-derive
the wrong reason. The 200-line ceiling is adherence guidance; a file nobody can
hold in their head is where contradictions hide, and five were found here.

Checks run against effective text (HTML comments stripped, because they are
stripped before injection): the line ceiling; every .claude/rules/*.md declares
paths: or an explicit unconditional: true; no component version literal; no
TEMPORARY block carrying a past date; and the workspace-root CLAUDE.md is
byte-identical to its versioned copy -- the live file sits outside any git repo,
so that copy is its only version-controlled record.

Two traps recorded so they are not reintroduced: a bare \d+\.\d+\.\d+ matches the
first three octets of every IPv4 (the gate excludes dotted quads, or it fails on
192.168.0.180 in the agent's own file); and unconditional: true is NOT a Claude
Code feature but this project's own marker.

Workspace-root CLAUDE.md 208 -> 182 lines (142 effective), copy kept identical.
The nine-instance invariant table moved into the felhom-testing skill, which
triggers when writing or reviewing a test; all three directive bullets stayed in
the core. felhom.eu/CLAUDE.md got surgical corrections only and is knowingly still
over the ceiling at 227 effective lines -- closing it needs the restructure R-229
defers, said plainly rather than quietly absorbed.

CONTEXT.md gains standing ruling S-35. OPEN-ITEMS.md gains R-229.

Docs only -- no Go, no version bump, nothing built or deployed.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JJc8sAGRWmavP3rMtdpkr2
2026-08-06 09:38:52 +02:00

16 KiB
Raw Blame History

CLAUDE.md — Project Instructions for Claude Code (felhom.eu)

Read automatically when Claude Code works in this repo. Stable orientation only — current state lives in CONTEXT.md and the tops of hub/CHANGELOG.md / scripts/CHANGELOG.md / website/CHANGELOG.md, never here. Cross-repo orientation (the felhom system, artifact taxonomy, access): workspace-root /mnt/5_hdd/felhom.eu/git/CLAUDE.md; this file is felhom.eu-specific. A versioned copy of that workspace file lives at documentation/runbooks/workspace-CLAUDE.md.

Project overview

This repo contains:

  • Website (website/) — static HTML at felhom.eu, served via k3s nginx + git-sync sidecar.
  • Hub (hub/) — Go application (felhom-hub), the operator backend, on k3s at hub.felhom.eu.
  • K8s manifests (manifests/) — k3s deployment manifests for felhom-system services.
  • Architecture docs (documentation/) — the authoritative design home for the whole Felhom system: architecture/01..05-*.md, proxmox-platform.md, tests/phase*-findings.md, runbooks, audits. Read these before designing.
  • Skills (skills/) — the versioned source of the Claude Code skills; install/update with python3 scripts/install_skills.py (symlink — repo edits are live immediately).

See README.md for full architecture/DNS/email/SEO docs. See TASK.md for the current task (if any). See REUSE.md before writing new code.

The Felhom system (so the hub's role is in context)

Felhom is Proxmox-based, with a locked three-component model:

  • Hub (this repo, hub/) — operator backend. Authors operator intent; mirrors box reality; holds no data-plane role and never connects inbound to a box.
  • Host agent (repo felhom-agent/) — one per Proxmox host; owns all Proxmox interaction.
  • In-guest controller (repo felhom-controller/) — one per customer LXC; Docker-only.

Hub — architecture (version-free; current version = manifests/hub.yaml image tag)

The hub ingests two report streams — the agent's host-domain report (POST /api/v1/host-report, the heartbeat/dead-man's-switch) and the legacy controller report (POST /api/v1/report, frozen until the slice-10 cutover — do not modify) — plus structured controller events (POST /api/v1/event, gated by allowedEventTypes). Around them: staleness/disk/storage-fill/leaf/capability monitor checkers, the two-tier notification dispatcher (operator English / customer Hungarian, Resend, cooldowns), the app-mail relay, customer-config + Day-0 artifact-manifest management (the checksum trust root the host bootstrap verifies against), assets serving, and the password-gated operator web UI. Package map, helpers, seams, extension points: REUSE.md (e.g. new event types must enter allowedEventTypes + customerMessages together).

Code quality rules

  • If you need more input or troubleshooting output, ask first — don't guess.
  • Testing doctrine (non-hollow tests, red-proofs, seams): use the felhom-testing skill.
  • Seam-wiring rule — it covers TEMPLATE GATES: a feature is not shipped until its entry point is reachable. For UI, any conditional affordance ({{if .Flag}} around a button/form/script) ships with a render test per branch of the gate — handler tests that POST directly prove nothing about reachability.
  • A go test -run pattern that matches no test prints ok and exits 0. A red-proof that uses -run must first prove the filter matched something (-v, look for === RUN). Generally: an instrument that can drop results silently is not a measurement.
  • A health check issues no block I/O. A probe that touches a wedged device enters uninterruptible sleep, survives SIGKILL, and cannot be recovered until the device returns or the host reboots — so systemctl restart hangs too. A timeout protects the caller's control flow and nothing else: the blocked thread remains. Liveness is decided from /proc and the kernel's own state, never by reading or writing the filesystem.
  • UI/design work (tokens, gates, copy rules): use the felhom-ui-design skill.
  • Logging: levels/English/no-secrets rules per documentation/runbooks/logging-conventions.md (DEBUG = flow detail, INFO = state change + duration; logs are operator-tier English; keys never values — the hub's bundle secret-gate blocks violating pulls fail-closed).

Workflow & artifacts

The planning/architecture assistant ("project Claude", in claude.ai) writes specs and validates pushes; you (Claude Code) implement. A file being open in the editor is NOT an instruction.

  • TASK.md / TASK-*.md — a spec for you to implement. Then push and update hub/CHANGELOG.md and root REPORT.md per the convention below.
  • RUNBOOK-*.md — an operational procedure. CC executes the steps it has access and capability for, including live validation on the demo nodes and the demo Proxmox host (CC has root@felhom-pve SSH + the felhom-agent token). Mark a step HUMAN only when it genuinely needs physical presence, a real-world decision, or credentials CC truly lacks.
  • Validation of a push against a spec's criteria is project Claude's job, not yours, unless asked.
  • Browser automation is NOT available in the DooPlex environment (claude-in-chrome was a Windows-workstation capability). Validate at the endpoint level — invoke the exact endpoint the UI invokes — and via render tests; say which method was used. The hub UI is operator-password-gated anyway, so render tests were already the method for UI changes. Strict end-to-end UI coverage is a manual click-through by the operator.

In every repository where you make a change, update both files in that repo:

  • CHANGELOG.md — cumulative log, newest on top (here: per-area hub/, scripts/, website/).
  • REPORT.mdoverwrite with the most recent implementation/validation summary only. Parallel sessions: REPORT.md is overwritten, so two sessions working in this repo at once will clobber each other. The second session writes REPORT-<topic>.md instead and never touches the shared REPORT.md.

Never write secrets into any committed file — reference them as "stored out-of-band".

  • Update REUSE.md if you added/changed/deprecated a shared helper or pattern (same commit).
  • Never git add -A in this repo — parallel sessions share the clone and it sweeps foreign WIP. Stage explicit paths only, git pull --rebase before every push, and do not run two writing sessions on one clone (use git worktree if truly needed).

End-of-session checklist

  • CHANGELOG.md + REPORT.md per the rule above, in every repo touched.
  • REUSE.md, if a shared helper or pattern moved (same commit).
  • The capability map (documentation/architecture/00-capability-map.md), if a capability's status changed — with its new evidence citation.
  • The architecture doc that owns any changed contract (S-1, CONTEXT.md).
  • Root STATUS.mdupdate it at the end of every session in which something shipped, broke, or was decided. It is a view of documentation/backlog/OPEN-ITEMS.mdnothing may exist only there. One screen; cut items rather than extending it. It is written for the operator in plain language, and is deliberately not CONTEXT.md — do not consolidate the two.
  • A finding goes in OPEN-ITEMS.md first, never only in a report, an audit or STATUS.md. Four items in this project were minted in a spike doc and lost (R-153/154/155, R-156/157).
  • Confirm your own last push's CI run went green, by run ID. CI emails on failure, which is a PUSH signal — this is the PULL check that catches a lost, filtered or unread mail. Quote the run id and its conclusion in the session report, e.g. curl -s "https://gitea.dooplex.hu/api/v1/repos/admin/<repo>/actions/tasks?limit=3" → match the head_sha to your commit. An unchecked green is an assumption, not an observation.

Hub stack — the two constraints

The dependency list is hub/go.mod's business and the deploy shape is manifests/. Only the two rules that the code cannot tell you belong here:

  • No web frameworks. Go stdlib net/http + html/template, and it stays that way.
  • Secrets via out-of-band secretKeyRef — never inline stringData (REUSE.md §3).

Environment & access

Claude Code runs on DooPlex (192.168.0.180, Debian 13, user kisfenyo) — the k3s node itself. Repos in /mnt/5_hdd/felhom.eu/git/, build dirs in /mnt/5_hdd/felhom.eu/build/. kubectl and the image build/push are local commands; felhom-pve is one SSH hop.

Host addresses, routes, node names, break-glass and what is provisioned on each: documentation/operations/nodes.md — the single home. Do not restate them here; re-check an address rather than trusting one written down. Tailscale topology, the accept-dns rule, the accept-routes spike result and rollback: documentation/operations/tailscale.md.

Which box do I break?documentation/runbooks/target-selection.md — the tiers, and per machine what is freely permitted / needs care / forbidden, each with its reason. Read it before picking a machine for a drill, a destructive test or a throwaway VM. DooPlex is Tier 2 — precious: it is the recovery chain (hub, Gitea, registry, PBS, k3s+Longhorn), and never a drill target. Both demo Proxmox hosts are Tier 0 — disposable; drill and build VMs belong on the t740.

Build & deploy — Hub (GitOps via ArgoCD)

Full runbook: use the felhom-build-deploy skill. The load-bearing rules:

The whole cluster is GitOps via a single ArgoCD app felhom syncing this repo's manifests/ to felhom-system. Auto-sync is OFF — deploys are a deliberate manual sync. ArgoCD's source of truth is the manifest:

  • A code change + CHANGELOG bump deploys NOTHING. The running image changes only when manifests/hub.yaml's image: tag changes in git and the app is synced.
  • Pin explicit versions, never :latest. Never bare kubectl set image/kubectl apply (reverted on next sync).
  • The live image can lag the CHANGELOG when a bump was committed but the manifest/sync step never happened — reconcile via the manifest, not the changelog.
  • Green gate before any hub commit: go build ./... && go vet ./... && go test ./... in hub/.

Clean-tree gate before any build: git status --porcelain must be empty and git rev-parse HEAD must equal git rev-parse origin/main in the repo being built. An unpushed change does not exist — never build a dirty or unpushed tree. The git pull in the build step stays (it is a no-op when you work in this tree, and load-bearing if anything was pushed from elsewhere).

Steps: commit+push code → cd /mnt/5_hdd/felhom.eu/build/felhom-hub && ./build.sh <VER> --push (local) → bump manifests/hub.yaml tag + push → ArgoCD hard-refresh + sync (kubectl-patch method in the skill, now local sudo kubectl) → verify Synced/Healthy + rollout + image + startup log.

Gates — ONE entry point

Run python3 scripts/repo_gates.py after ANY change in this repo. It is the one entry point and runs every gate — site_gates.py, hostinstall_gates.py, hub_confirm_gate.py, manifest_bearer_gate.py and reuse_refs_check.py on this root — streaming each gate's own output and exiting non-zero if any fails. --fast selects only the gates that touch no network and no container runtime; today that is all of them. A missing gate script is a FAILURE, never a skip.

site_gates.py is a gate, not a runner — do not model new work on it; app-catalog-felhom.eu/scripts/catalog_gates.py is the canonical runner (R-161).

The pre-push hook. .githooks/pre-push runs repo_gates.py --fast and refuses the push if it fails. It is per-clone and switched on once with git config core.hooksPath .githooks — a clone does not carry it, and any manual repo_gates.py run WARNS when this clone is unarmed. git push --no-verify bypasses it deliberately; say so in the session report when you use it. Both facts are why continuous integration is still owed (OPEN-ITEMS.md R-168) — this hook is local and skippable, and only CI is neither.

Build & deploy — Website / Manifests

  • Website auto-deploys via git-sync; just push to main (live in 12 min). Website changes go through repo_gates.py above (it runs site_gates.py); new pages go into that gate's PAGES list. Emergency edits: https://files.felhom.eu. All website/ HTML is UTF-8 with BOM — preserve it.
  • THE INSTALLER DOES NOT (R-110, 2026-08-03). manifests/webpage.yaml runs two git-syncs: the website from main as above, and /scripts/ from the tag installer-v<SCRIPT_VERSION>. Pushing scripts/felhom-host-install.sh therefore changes nothing that any machine downloads — which it used to, within thirty seconds, for the one artifact that runs as root on a virgin box.
    • To publish: cut installer-v<new SCRIPT_VERSION>, bump the --ref in webpage.yaml (both the sidecar and the init container), commit, and sync. hostinstall_gates.py gate 6 fails if the manifest stops naming an installer-v… tag or if the website stops tracking main.
    • To roll back: move the tag back to the previous commit and wait ~30 s. No ArgoCD sync and no deploy — git-sync picks up a moved tag on its next period, measured live on 2026-08-03 in both directions. That is the emergency lever; fix forward with a new version afterwards.
    • Do NOT pin the website to the tag. The sparse-checkout used to cover /website/ and /scripts/ in one sync, and pinning that would turn every copy edit into a release.
    • The URL never carries a ref (https://felhom.eu/scripts/felhom-host-install.sh), so felhom-bootstrap.sh and the hub's day-0 command follow the tag with no edit — do not add one.
    • The installer's own sixteen run-time fetches are pinned separately, to raw/tag/v$ART_AGENT_VER in the agent repo (R-183) — they are the agent's configs, not this repo's.
  • Manifests are GitOps via the felhom app — commit to main, then deliberate sync.

Key patterns

  • Hub status logic: OK (report < 30m), WARN (30m1h or health=warn), DOWN (> 1h or health=fail); host liveness thresholds shared between UI and checker (never invent a second definition).
  • SQLite timestamps vary in format — always parseSQLiteTime().
  • Dashboard/detail auto-refresh every 60s via meta refresh. Geo-restricted to Hungary via nginx ingress annotation.
  • Helpers, seams, extension points, traps: REUSE.md — the map is maintained same-commit.