Files
felhom.eu/CLAUDE.md
T
admin 0a5e9b14dc
gates / gates (push) Successful in 14s
due-checks gate (R-341), floor raise recorded (R-343), snapshot coverage (R-342)
PART 1+2 — dated checks stop being wishes. R-341 booked two measurements as
prose in a register row; nothing read those dates and nothing would have
objected when they passed. The dates now live in a DUE-CHECKS block INSIDE
OPEN-ITEMS.md (inside, so no sidecar can drift from it) and a new gate reads
them. Registered as #10 in repo_gates.py, --fast, so it runs in BOTH the
pre-push hook and CI.

  exit 0  nothing due (prints pending count + nearest date; empty block too)
  exit 1  a row is due/overdue (due <= today, UTC -- due TODAY counts), or a
          row names an item with no R-row
  exit 2  block absent/duplicated/unparseable -- INCONCLUSIVE, never 0

It REFUSES rather than warns, and its docstring states the limitation: it is
NOT a scheduler, it fires on the next push, not on the date.

37 tests. BOTH red-proofs run and reverted -- and the first one earned its
keep by catching a hollow assertion of MINE rather than confirming the gate:
flipping <= to < left a due-today row in neither bucket, min() raised on an
empty list, and the TRACEBACK exited 1, so "rc == 1" passed while the
boundary was wrong. An exit code cannot tell a verdict from a crash. The test
now asserts the conviction banner and the absence of a traceback, and the gate
returns 2 rather than crashing if that partition breaks again.

PART 3 — the floor raise, and the premise was WRONG. Read back from the store
(not the form): min_controller_version = 0.216.0 @ 12:36:58Z, zero
per-customer overrides, no "managed floor HELD" line. But read 5 shows the
raise was NOT a no-op: demo-felhom had been on 0.214.0 since 12 Aug and
auto-updated 0.214.0 -> 0.216.0 at 12:37:07Z -- NINE SECONDS after the save,
exactly the immediate action publish-train rule 2 documents. No error events
followed; it restarted clean.

R-343 is therefore filed OPEN, not CLOSED: the closing condition was all five
reads clean and no directive served. It went well, but a record calling it
inert when it moved a customer box is what misleads the next reader. The row
also states why the floor was behind -- rule 2 policy, not drift, earned by
the 2026-07-11 skew onto Peti's box -- and cites ResolveManagedFloor
(store.go:2068) plus the two build-felhom-iso.sh facts (build-time at :267,
fails open at :78-82) rather than asserting them.

Two boxes are below the floor and neither reports: drill-r50 (blocked,
powered off) and peti-felhom (host row deleted). peti-felhom was NOT
contacted -- its row records that a report from a deleted host 401s and is
not persisted, so the raise cannot reach it.

PART 4 — R-342 filed READY, quoting stop2-snapshot.txt verbatim: Hetzner
server snapshot 421440873 covers /dev/sda only; /mnt/pbs-datastore is a
separate Volume that snapshots exclude, so a rollback restores software state
and NOT the datastore. Fine for that upgrade; the safeguard for any future
procedure that could touch the datastore does not exist and is Viktor's call.

Also: CLAUDE.md's gate list named 6 of 10 registered gates -- completed
rather than appending a 7th to a wrong list (124 -> 128 effective, ceiling
200). Capability map deliberately unchanged; no row cites a floor or golden
version. repo_gates.py fully green, 10/10.
2026-08-18 15:16:51 +02:00

8.7 KiB

CLAUDE.md — felhom.eu

Stable orientation only — current state lives in CONTEXT.md and the tops of hub/CHANGELOG.md / scripts/CHANGELOG.md / website/CHANGELOG.md, never here. Cross-repo conventions (the three-component model, artifact taxonomy, access, clean-tree gate, secrets): workspace-root /mnt/5_hdd/felhom.eu/git/CLAUDE.md, whose versioned copy is documentation/runbooks/workspace-CLAUDE.md. Path-scoped detail: .claude/rules/.

What this repo is

Four surfaces in one repo, plus the design home for the whole system:

  • hub/ — felhom-hub, the operator backend (Go, k3s, hub.felhom.eu).
  • website/ — static HTML at felhom.eu, served by k3s nginx + git-sync.
  • manifests/ — k3s manifests for felhom-system, GitOps via one ArgoCD app.
  • scripts/ — the public installer (felhom-host-install.sh) and this repo's gates.
  • documentation/ — the authoritative design home for all of Felhom, not just this repo.
  • skills/ — versioned source of the Claude Code skills; install with python3 scripts/install_skills.py (symlink — repo edits are live immediately).

Doing X → read Y

Doing Read
writing any new code REUSE.md — helpers, seams, extension points, traps
needing current state / roadmap CONTEXT.md
hub work (architecture, deploy, patterns) loads itself: .claude/rules/hub.md
website or installer work loads itself: .claude/rules/website.md
manifests / ArgoCD / secrets loads itself: .claude/rules/manifests.md
writing or routing a document loads itself: .claude/rules/docs.md
build, deploy, publish, verify a version the felhom-build-deploy skill
writing or reviewing a test, fixing a bug the felhom-testing skill
UI, tokens, badges, Hungarian copy the felhom-ui-design skill
host addresses, break-glass, node facts documentation/operations/nodes.md — never restate them
which box may I break documentation/runbooks/target-selection.md
what version is live anywhere ask the hub (/hosts, /configs) or the box — never a doc
the authoritative design documentation/architecture/01..05-*.md

Code quality

  • If you need more input or troubleshooting output, ask first — don't guess.
  • A go test -run pattern that matches no test prints ok and exits 0. A red-proof using -run must first prove the filter matched something (-v, look for === RUN). Generally: an instrument that can drop results silently is not a measurement.

The installer publishes by TAG, not by push (R-110)

This fence is in the core deliberately: its trigger is editing scripts/felhom-host-install.sh, and no path-scoped rule covers that file. It governs the one artifact that runs as root on a virgin box.

  • Pushing scripts/felhom-host-install.sh to main publishes NOTHING. manifests/webpage.yaml runs two git-syncs: the website from main, and /scripts/ from the tag installer-v<SCRIPT_VERSION>.
  • To publish: cut installer-v<new SCRIPT_VERSION>, bump the --ref in webpage.yaml (both the sidecar and the init container), commit, sync.
  • To roll back: move the tag back and wait ~30 s. No ArgoCD sync, no deploy — that is the emergency lever; fix forward afterwards.
  • Do NOT pin the website to the tag, and the URL never carries a ref — felhom-bootstrap.sh and the hub's day-0 command follow the tag with no edit.
  • hostinstall_gates.py gate 6 fails if the manifest stops naming an installer-v… tag or if the website stops tracking main.

Workflow — what is specific to this repo

  • Never git add -A here — parallel sessions share the clone and it sweeps foreign WIP. Stage explicit paths only, git pull --rebase before every push.
  • REPORT.md is overwritten, so two sessions in this repo clobber each other. The second session writes REPORT-<topic>.md and never touches the shared REPORT.md.
  • CHANGELOG.md here is per-area: hub/, scripts/, website/.

Gates — ONE entry point

Run python3 scripts/repo_gates.py after ANY change in this repo. It runs every gate — site_gates.py, hostinstall_gates.py, hub_confirm_gate.py, manifest_bearer_gate.py, reuse_refs_check.py, instructions_gate.py, golden_currency_gate.py, wire_contract_gate.py, hub_copy_gate.py and due_checks_gate.py — streaming each gate's own output and exiting non-zero if any fails. --fast selects the gates that touch no network and no container runtime; today that is all of them. A missing gate script is a FAILURE, never a skip.

due_checks_gate.py refuses the push when a dated check in OPEN-ITEMS.md's DUE-CHECKS block has come due (R-341). It is not a scheduler — it fires on the next push, not on the date.

site_gates.py is a gate, not a runner — do not model new work on it; app-catalog-felhom.eu/scripts/catalog_gates.py is the canonical runner (R-161).

The pre-push hook (.githooks/pre-push) runs it with --fast and refuses a failing push. It is per-clone — switch it on once with git config core.hooksPath .githooks, and a manual run WARNS when this clone is unarmed. git push --no-verify bypasses it deliberately; say so in the session report when you use it — CI re-runs the same entry point on every push and emails the operator on failure, so a bypass is noticed even though it is not blocked (R-168, CLOSED 2026-08-02; CI reports rather than refuses because there is no PR to gate — R-169).

End-of-session checklist

Registers first — a finding goes in documentation/backlog/OPEN-ITEMS.md first, never only in a report, an audit or STATUS.md. Four items in this project were minted in a spike doc and lost (R-153/154/155, R-156/157). This applies to every session that ships, breaks or decides something, not only sessions that touch documentation/ — which is why it is here and not in docs.md.

  • CHANGELOG.md + REPORT.md in every repo touched (see the workspace root for the rule, and the parallel-session caveat above).
  • REUSE.md, if a shared helper or pattern moved (same commit).
  • OPEN-ITEMS.md — every finding, with a number.
  • Root STATUS.md — at the end of every session in which something shipped, broke or was decided. It is a view of OPEN-ITEMS.md; nothing may exist only there. One screen, written for the operator in plain language, and deliberately not CONTEXT.md.
  • The capability map (documentation/architecture/00-capability-map.md), if a capability's status changed — with its new evidence citation.
  • python3 scripts/unproven.py --summary — one line per status, and the not-walked total. Run it at the end of any session that shipped, broke or proved something, and say in the report if a number moved. It exists because "which claims are unproven?" was answerable only by a person reading a page: a session asked for "the nine grey claims" could not determine which nine and rightly refused to guess (R-326). Nine was real and answered a different question — it is the count of claims the 2026-08-09 pass DOWNGRADED. Not-walked is 32 of 55. A status that moves without anyone noticing is how the picture stops being true.
  • Confirm your own last push's CI run went green, by run ID. CI emails on failure, which is a PUSH signal; this is the PULL check that catches a lost, filtered or unread mail. Quote the run id and its conclusion, e.g. curl -s "https://gitea.dooplex.hu/api/v1/repos/admin/<repo>/actions/tasks?limit=3" → match the head_sha to your commit. An unchecked green is an assumption, not an observation.