Files
felhom.eu/CLAUDE.md
T
admin 699790b12d docs: write down which boxes are disposable (target selection by blast radius)
Nothing in the repo said which machines are safe to break. The host table gave
access and role and stopped there, so a session needing a victim had to guess --
and the guessing inverted: the two boxes that exist to be broken were treated as
sacred, and DooPlex (the recovery chain) got used because it was the only box no
spec had fenced.

New documentation/runbooks/target-selection.md -- one page, three tiers, and per
machine what is freely permitted / needs care / forbidden, each carrying its
REASON so a rule can be correctly narrowed later instead of ossifying. States the
selection rule positively (start at Tier 0; a Tier 2 box only when a task says so
explicitly; an absent fence is not permission) and that fences name ACTS, not
machines -- demo-hp's over-subscribed local-lvm is one dangerous storage, not a
dangerous box.

CLAUDE.md: host table gains a Blast radius column, gains the missing demo-hp row
(it was where the drill VMs ran and it was not in the table at all), and a pointer
line to the new runbook.

CORRECTION to the spec's problem statement: the designation was not missing. The
2026-07-25 operator ruling naming the t740 as drill+build VM host -- explicitly
"moved off DooPlex" -- already existed in operations/nodes.md. It sat where no
session reads at start, while the prohibitions were repeated in every task spec.
The defect is reachability of the ruling, not its absence, and the R-116 drill on
DooPlex contradicted a written ruling rather than filling a vacuum.

CORRECTION to the R-116 record, same commit: the baseline claimed controller
0.186.0 on both demo boxes. Only felhom-pve was sampled and generalised; demo-hp
re-checked directly runs 0.185.1, so the fleet is split and R-114's TargetAbsent
branch is absent from demo-hp. Fixed in the audit table and REPORT-r116-diag.

Docs only -- no code, no build, no deploy, no host reconfigured, no host renamed.
2026-07-30 08:13:58 +02:00

11 KiB
Raw Blame History

CLAUDE.md — Project Instructions for Claude Code (felhom.eu)

Read automatically when Claude Code works in this repo. Stable orientation only — current state lives in CONTEXT.md and the tops of hub/CHANGELOG.md / scripts/CHANGELOG.md / website/CHANGELOG.md, never here. Cross-repo orientation (the felhom system, artifact taxonomy, access): workspace-root /mnt/5_hdd/felhom.eu/git/CLAUDE.md; this file is felhom.eu-specific. A versioned copy of that workspace file lives at documentation/runbooks/workspace-CLAUDE.md.

Project overview

This repo contains:

  • Website (website/) — static HTML at felhom.eu, served via k3s nginx + git-sync sidecar.
  • Hub (hub/) — Go application (felhom-hub), the operator backend, on k3s at hub.felhom.eu.
  • K8s manifests (manifests/) — k3s deployment manifests for felhom-system services.
  • Architecture docs (documentation/) — the authoritative design home for the whole Felhom system: architecture/01..05-*.md, proxmox-platform.md, tests/phase*-findings.md, runbooks, audits. Read these before designing.
  • Skills (skills/) — the versioned source of the Claude Code skills (felhom-build-deploy, felhom-ui-design, felhom-testing, felhom-app-catalog); install/update with python3 scripts/install_skills.py (symlink into ~/.claude/skills/ on POSIX, junction on Windows — either way repo edits are live immediately).

See README.md for full architecture/DNS/email/SEO docs. See TASK.md for the current task (if any). See REUSE.md before writing new code.

The Felhom system (so the hub's role is in context)

Felhom is Proxmox-based, with a locked three-component model:

  • Hub (this repo, hub/) — operator backend. Authors operator intent; mirrors box reality; holds no data-plane role and never connects inbound to a box.
  • Host agent (repo felhom-agent/) — one per Proxmox host; owns all Proxmox interaction.
  • In-guest controller (repo felhom-controller/) — one per customer LXC; Docker-only.

Hub — architecture (version-free; current version = manifests/hub.yaml image tag)

The hub ingests two report streams — the agent's host-domain report (POST /api/v1/host-report, the heartbeat/dead-man's-switch) and the legacy controller report (POST /api/v1/report, frozen until the slice-10 cutover — do not modify) — plus structured controller events (POST /api/v1/event, gated by allowedEventTypes). Around them: staleness/disk/storage-fill/leaf/capability monitor checkers, the two-tier notification dispatcher (operator English / customer Hungarian, Resend, cooldowns), the app-mail relay, customer-config + Day-0 artifact-manifest management (the checksum trust root the host bootstrap verifies against), assets serving, and the password-gated operator web UI. Package map, helpers, seams, extension points: REUSE.md (e.g. new event types must enter allowedEventTypes + customerMessages together).

Code quality rules

  • Always double-check generated code for bugs, logic issues, syntax errors.
  • Handle edge cases without overcomplicating.
  • Add debug capabilities (logging, verbose output).
  • If you need more input or troubleshooting output, ask first — don't guess.
  • Testing doctrine (non-hollow tests, red-proofs, seams): use the felhom-testing skill.
  • Seam-wiring rule — and it covers TEMPLATE GATES (fourth inert seam, hub v0.70.1): a feature is not shipped until its entry point is reachable. For UI, any conditional affordance ({{if .Flag}} around a button/form/script) ships with a render test per branch of the gate — handler tests that POST directly prove nothing about reachability. The v0.70.0 ghost-delete was fully implemented server-side and fully dead UI because the button sat inside the wrong gate.
  • UI/design work (tokens, gates, copy rules): use the felhom-ui-design skill.
  • Logging: levels/English/no-secrets rules per documentation/runbooks/logging-conventions.md (DEBUG = flow detail, INFO = state change + duration; logs are operator-tier English; keys never values — the hub's bundle secret-gate blocks violating pulls fail-closed).

Workflow & artifacts

The planning/architecture assistant ("project Claude", in claude.ai) writes specs and validates pushes; you (Claude Code) implement. A file being open in the editor is NOT an instruction.

  • TASK.md / TASK-*.md — a spec for you to implement. Then push and update hub/CHANGELOG.md and root REPORT.md per the convention below.
  • RUNBOOK-*.md — an operational procedure. CC executes the steps it has access and capability for, including live validation on the demo nodes and the demo Proxmox host (CC has root@felhom-pve SSH + the felhom-agent token). Mark a step HUMAN only when it genuinely needs physical presence, a real-world decision, or credentials CC truly lacks.
  • Validation of a push against a spec's criteria is project Claude's job, not yours, unless asked.
  • Browser automation is NOT available in the DooPlex environment (claude-in-chrome was a Windows-workstation capability). Validate at the endpoint level — invoke the exact endpoint the UI invokes — and via render tests; say which method was used. The hub UI is operator-password-gated anyway, so render tests were already the method for UI changes. Strict end-to-end UI coverage is a manual click-through by the operator.

In every repository where you make a change, update both files in that repo:

  • CHANGELOG.md — cumulative log, newest on top (here: per-area hub/, scripts/, website/).
  • REPORT.mdoverwrite with the most recent implementation/validation summary only. Parallel sessions: REPORT.md is overwritten, so two sessions working in this repo at once will clobber each other. The second session writes REPORT-<topic>.md instead and never touches the shared REPORT.md.

Never write secrets into any committed file — reference them as "stored out-of-band".

  • Update REUSE.md if you added/changed/deprecated a shared helper or pattern (same commit).
  • Never git add -A in this repo — parallel sessions share the clone and it sweeps foreign WIP (the v0.47.0 146d165 incident: a red-proof-mutated guard got swept to main). Stage explicit paths only, git pull --rebase before every push, and do not run two writing sessions on one clone (use git worktree if truly needed).

Tech stack (Hub)

  • Language: Go (stdlib net/http + html/template, no frameworks). DB: SQLite via modernc.org/sqlite (pure Go). Auth: bcrypt + Bearer tokens + session cookies + CSRF.
  • Deploy: Docker on k3s (felhom-system ns). Storage: Longhorn PVC at /data/ (SQLite DB).
  • Config: YAML via ConfigMap at /etc/felhom-hub/hub.yaml. Secrets via out-of-band secretKeyRef (never inline stringData — REUSE.md §3).

Environment & access

Claude Code runs on DooPlex (192.168.0.180, Debian 13, user kisfenyo) — the k3s node itself. Repos in /mnt/5_hdd/felhom.eu/git/, build dirs in /mnt/5_hdd/felhom.eu/build/. kubectl and the image build/push are local commands; felhom-pve is one SSH hop.

Host Access Role Blast radius
DooPlex (this host) local — /mnt/5_hdd/felhom.eu/{git,build}/ Build + push images, sudo kubectl Tier 2 — precious. It is the recovery chain (hub, Gitea, registry, PBS, k3s+Longhorn). Never a drill target
Demo Proxmox host (N100) ssh felhom-pve — via Tailscale 100.70.170.35 (location-independent); felhom-pve-lan = LAN 192.168.0.162 fallback pveum/pct + live Proxmox validation Tier 0 — disposable
Demo Proxmox host (HP t740) ssh demo-hp — via Tailscale 100.76.96.79; demo-hp-lan = LAN 192.168.0.87 (ProxyJump felhom-pve). No baked SSH key — G1 break-glass password vaulted in the hub The designated drill + build VM host (operator ruling 2026-07-25) Tier 0 — disposable. Reach here first

Which box do I break?documentation/runbooks/target-selection.md — the tiers, and per machine what is freely permitted / needs care / forbidden, each with its reason. Read it before picking a machine for a drill, a destructive test or a throwaway VM.

The felhom-pve transport is Tailscale (the N100 is travel-portable) — topology, the accept-dns rule, the accept-routes spike result, rollback, and the vacation-day checklist live in documentation/operations/tailscale.md.

Legacy: Windows workstation. Until 2026-07-19 CC ran on Windows 11 with repos in E:\git\, and every remote command needed SSH=/c/Windows/System32/OpenSSH/ssh.exe (Git Bash's ssh fails silently). Retained in case that environment is revived.

Build & deploy — Hub (GitOps via ArgoCD)

Full runbook: use the felhom-build-deploy skill. The load-bearing rules:

The whole cluster is GitOps via a single ArgoCD app felhom syncing this repo's manifests/ to felhom-system. Auto-sync is OFF — deploys are a deliberate manual sync. ArgoCD's source of truth is the manifest:

  • A code change + CHANGELOG bump deploys NOTHING. The running image changes only when manifests/hub.yaml's image: tag changes in git and the app is synced.
  • Pin explicit versions, never :latest. Never bare kubectl set image/kubectl apply (reverted on next sync).
  • The live image can lag the CHANGELOG when a bump was committed but the manifest/sync step never happened — reconcile via the manifest, not the changelog.
  • Green gate before any hub commit: go build ./... && go vet ./... && go test ./... in hub/.

Clean-tree gate before any build: git status --porcelain must be empty and git rev-parse HEAD must equal git rev-parse origin/main in the repo being built. An unpushed change does not exist — never build a dirty or unpushed tree. The git pull in the build step stays (it is a no-op when you work in this tree, and load-bearing if anything was pushed from elsewhere).

Steps: commit+push code → cd /mnt/5_hdd/felhom.eu/build/felhom-hub && ./build.sh <VER> --push (local) → bump manifests/hub.yaml tag + push → ArgoCD hard-refresh + sync (kubectl-patch method in the skill, now local sudo kubectl) → verify Synced/Healthy + rollout + image + startup log.

Build & deploy — Website / Manifests

  • Website auto-deploys via git-sync; just push to main (live in 12 min). Run python3 scripts/site_gates.py after ANY website change; new pages go into its PAGES list. Emergency edits: https://files.felhom.eu. All website/ HTML is UTF-8 with BOM — preserve it.
  • Manifests are GitOps via the felhom app — commit to main, then deliberate sync.

Key patterns

  • Hub status logic: OK (report < 30m), WARN (30m1h or health=warn), DOWN (> 1h or health=fail); host liveness thresholds shared between UI and checker (never invent a second definition).
  • SQLite timestamps vary in format — always parseSQLiteTime().
  • Dashboard/detail auto-refresh every 60s via meta refresh. Geo-restricted to Hungary via nginx ingress annotation.
  • Helpers, seams, extension points, traps: REUSE.md — the map is maintained same-commit.