# CLAUDE.md — `/mnt/5_hdd/felhom.eu/git` workspace root (DooPlex) ## What this workspace is `/mnt/5_hdd/felhom.eu/git` is a parent folder holding the felhom sibling repos. Most are one logical product — **Felhom**, a managed home-server service for Hungarian households — spread across several repos. (Any non-felhom repo is unrelated; ignore unless asked.) **Claude Code runs HERE, on DooPlex (192.168.0.180), as `kisfenyo`.** Builds are local commands; the Proxmox host is one SSH hop (`ssh felhom-pve`). The Windows workstation is no longer the orchestration point and its trees are stale — see "Legacy: Windows workstation" at the bottom. Run CC inside tmux so sessions survive SSH drops: **`tmux new -A -s cc`**. ## This host is production infrastructure DooPlex runs Gitea, the container registry, k3s + Longhorn, PBS, and the hub. Treat it accordingly: - NEVER run `docker system prune`, `docker image prune -a`, or any global Docker cleanup here. - NEVER touch k3s data dirs, Longhorn mounts, PBS datastores, or Gitea storage. Workspace is `/mnt/5_hdd/felhom.eu/` — stay inside it plus `~/build` symlinks/dirs. - Destructive disk/guest operations belong to felhom-pve via the agent — never on this host. - Do not run Claude Code with permission prompts disabled on this host. - Watch disk headroom before large builds: `df -h /mnt/5_hdd /` — abort if either is >90%. ## The Felhom system (three-component model, Proxmox-based) - **Hub** — operator backend on k3s (`hub.felhom.eu`). Lives in `felhom.eu/hub/`. - **Host agent** — one per Proxmox host, operator-tier, owns all Proxmox interaction. Repo `felhom-agent/`. - **In-guest controller** — one per customer LXC, Docker-only. Repo `felhom-controller/`. Other felhom repos: `app-catalog-felhom.eu/` (app templates), `homelab-manifests/` (DooPlex k3s). **Authoritative design docs (read these before designing anything):** `felhom.eu/documentation/architecture/01..05-*.md`, `felhom.eu/documentation/proxmox-platform.md`, `felhom.eu/documentation/tests/phase{0,1-2,3,4}-findings.md`. ## Per-repo guidance When you work in a repo, read its `CLAUDE.md` (it loads on-demand the moment you touch a file there): - `felhom-agent/CLAUDE.md` — the Go host agent. - `felhom.eu/CLAUDE.md` — hub + website + manifests + the architecture docs. - `felhom-controller/CLAUDE.md` — the in-guest controller. ## Skills Four Felhom skills exist (personal scope, `~/.claude/skills/`): **`felhom-build-deploy`** (all build/deploy/publish runbooks), **`felhom-ui-design`** (design-system v2 tokens/rules/gates), **`felhom-testing`** (non-hollow tests + red-proofs + seams), **`felhom-app-catalog`** (catalog authoring workflow). Source of truth: `felhom.eu/skills/`; install/update with `python3 felhom.eu/scripts/install_skills.py` (symlink — repo edits are live immediately). ## Memory The accumulated project memory (119 files) migrated from the Windows workstation lives at `/mnt/5_hdd/felhom.eu/git/.claude-memory/`, surfaced to Claude Code via `~/.claude/projects/-mnt-5-hdd-felhom-eu-git/memory` (symlink). `MEMORY.md` there is the index. Memories reflect what was true when written — verify a named file/flag still exists before acting on it. ## Artifact taxonomy (READ THIS — it prevents the "what do I do?" stall) The planning/architecture assistant (in claude.ai, "project Claude") produces files with distinct roles. **A file being open in the editor is NOT an instruction. If no task is stated, ask.** - **`TASK.md` / `TASK-*.md`** — a spec for **you (Claude Code) to implement**. Implement it when it is placed as `TASK.md` at a repo root, or when explicitly told "implement ". Then push, update `CHANGELOG.md`, and write the repo's `REPORT.md`. - **`RUNBOOK-*.md`** — an operational procedure. CC executes the steps it has access and capability for, including live validation on the demo nodes and the demo Proxmox host (CC has root@felhom-pve SSH + the felhom-agent token). A step is human-only only when it genuinely needs physical presence, a real-world decision, or credentials CC truly lacks — mark those steps HUMAN. Do not decline a whole procedure because it touches a live host or a privileged token. (Judgment still applies: confirm before irreversible ops on real customer data — but demo scratch guests are fair game.) - **Validation/review** — checking a push against a spec's criteria is **project Claude's** job, not yours, unless asked. ## Shared conventions ### Standing rules — each one earned by a real failure (R-96, committed 2026-07-27) These were agreed in conversation and lived nowhere, so they bound nobody. They do now. 1. **Never combine a test run and a commit in one command.** A combined command has ONE exit code and the interesting one gets swallowed. Three recorded occurrences; the worst pushed a red suite because `packages ok: 28` was read while `rc=1` was not. Run the suite, read `rc`, *then* commit. 2. **A "no access" claim must list what was tried.** "No access" is unfalsifiable unless it names its attempts. Two wrong verdicts on 2026-07-27 alone: ep0 (declared unreachable after trying exactly one route — `felhom-pve → 10.77.0.1`; `DooPlex → 167.233.158.164` worked and the project memory said so), and the storage-box API (`api.hetzner.cloud` 404s for every storage-box endpoint; `api.hetzner.com/v1` is the real one, and the hub's own `hetznerapi.go:3` records it). 3. **An absent log line is not evidence of correct behaviour.** Verify with a POSITIVE observable — something that MUST appear when the system is healthy. An empty log is equally consistent with "working" and "stopped entirely". Earned twice on 2026-07-27: the R-88 watcher (an empty quiesce log could not distinguish a healthy loop from a dead one — retired in favour of the per-tier `/backup/due` polls in `pveproxy/access.log`), and a hub DB copy whose write had silently failed, returning a confident "0 events in window" from a file a day stale until its mtime was checked. 4. **A recommendation that is not followed gets one line saying why.** Silence reads as agreement and the disagreement is lost. Twice in the R-88/R-97 arc a review point was absorbed rather than argued: R-84 was folded into R-82 without a word, and R-97a's operator-only guard was dropped while the claim it was meant to enforce got committed as a comment — which is how a false guarantee shipped and survived a release. Disagreeing is fine; disagreeing silently is not. - **Push to `main` directly** — no feature branches. > **Clean-tree gate before any build:** `git status --porcelain` must be empty and > `git rev-parse HEAD` must equal `git rev-parse origin/main` in the repo being built. An unpushed > change does not exist — never build a dirty or unpushed tree. The `git pull` in the build step > stays (it is a no-op when you work in this tree, and load-bearing if anything was pushed from > elsewhere). > **In every repository where you make a change, update both files in that repo:** > - **`CHANGELOG.md`** — a cumulative log of **all** changes; newest entry on top. > - **`REPORT.md`** — **overwrite** with a summary of the **most recent** implementation (or significant validation/operational run) only; not cumulative. > > **Never write secrets** — tokens, passwords, private keys, API keys — into `CHANGELOG.md`, `REPORT.md`, or any committed file. Reference them as "stored out-of-band" instead. - **Versioning** is via build-time ldflags (`-X main.version`/`-X main.Version`); bump on meaningful changes + add a CHANGELOG entry. - Code quality: double-check for bugs/edge cases; add debug logging; **ask rather than guess** when you'd otherwise need to invent input or output. ## Live validation — no browser here **`claude-in-chrome` is NOT available on DooPlex.** The standard method is endpoint-level: invoke the exact endpoint the UI invokes (no server logic is skipped, only rendering) and say which method was used. Strict end-to-end UI coverage is a manual click-through by the operator. ## Access Local (this host): repos `/mnt/5_hdd/felhom.eu/git/`, build dirs `/mnt/5_hdd/felhom.eu/build/felhom-{controller,hub,agent}`, `sudo kubectl`, Go toolchain, Docker build+push to `gitea.dooplex.hu/admin/`. | Host | Access | Use | Blast radius | |---|---|---|---| | **DooPlex (this host)** | local — Debian 13, `kisfenyo`, `/mnt/5_hdd/felhom.eu/` | build/push images, `sudo kubectl`, build+run the agent for tests | **Tier 2 — precious.** It *is* the recovery chain (hub, Gitea, registry, PBS, k3s+Longhorn). **Never a drill target** | | Demo Proxmox host `demo-hp` (HP t740) | `ssh demo-hp` (tailnet `100.76.96.79`; **no baked key** — G1 break-glass password vaulted in the hub) | **the designated drill + build VM host** (operator ruling 2026-07-25) | **Tier 0 — disposable. Reach here first** | | Demo Proxmox host `demo-felhom` (N100) | `ssh felhom-pve` (root, no sudo; tailnet `100.70.170.35`) | pveum/pct + live Proxmox validation | **Tier 0 — disposable** | | Demo guest 9201 | `ssh felhom-pve "pct exec 9201 -- ..."` | the live demo controller | Tier 0 (rides its host) | | felhotest (legacy) | `ssh -p 33022 kisfenyo@router.abonet.hu` — **`Connection refused` 2026-07-30** | OLD /opt/docker compose mechanism | untiered — assume nothing | **Which box do I break?** → **`felhom.eu/documentation/runbooks/target-selection.md`** — the tiers, and per machine what is freely permitted / needs care / forbidden, each with its reason. Read it before picking a machine for a drill, a destructive test or a throwaway VM. **A task that needs a victim names one; an absent fence is not permission.** **Component versions are not recorded in any inventory doc** — agent/controller/hub versions change several times a day and the fleet is not uniform. Ask the hub's `/hosts` + `/configs`, or `felhom-agent --version` / `pct exec -- docker ps` on the box. The demo Proxmox host key changes on reprovision (N100) → refresh with `ssh-keygen -R 192.168.0.162` then connect with `-o StrictHostKeyChecking=accept-new` (`ssh-keyscan` hangs — avoid it). ## Legacy: Windows workstation Kept so the old environment can be revived; **not the current setup**. - Repos were in `E:\git\` (`/e/git/` in Git Bash); this file lived at `E:\git\CLAUDE.md`. - **SSH binary had to be** `SSH=/c/Windows/System32/OpenSSH/ssh.exe` — Git Bash's `/usr/bin/ssh` lacks access to the Windows SSH Agent and fails silently. Every remote command was `$SSH kisfenyo@192.168.0.180 "..."`; details in `felhom-controller/docs/vscode-ssh-fix.md`. - `pct exec` over SSH needed `export MSYS_NO_PATHCONV=1` (MSYS mangled `/`-paths). - Agent deploy was a two-hop copy: build on 180 → `scp` to the Windows box (local path needed `cygpath -w`) → `scp` on to felhom-pve. Beware CRLF when scp-ing config files through Windows. - Skills were installed as Windows junctions (`mklink /J`) rather than POSIX symlinks. - `claude-in-chrome` browser automation WAS available there (attaching only to sessions started after the bridge connected). ### Presence is not success A timestamp recording an **attempt** must never be read as evidence of a **result**. Where a status field travels alongside a timestamp, the verdict consults both — or the timestamp records only successes. | # | instance | what happened | |---|---|---| | 1 | **F-CRIT-2** | a phantom snapshot's ctime set tier freshness — an aborted 1-byte upload made the tier look backed up | | 2 | **R-100** | `LastRun` is written on failure, so a nightly-failing offsite tier kept the staleness clock fresh forever | Both were found by asking of a timestamp: *what exactly must have happened for this to be set?* If the answer is "we tried", it cannot answer "did it work". Corollary, from R-100's fix: when a verdict changes which field it counts from, **the alarm text has to change with it**. Leaving the message reading `last run 8h ago` while alarming on a six-day-old success turns a true alarm into one the operator dismisses. ### A comment asserting an invariant needs a test pinning it, or it is a wish **Six instances in this project have shipped guarantees the code did not provide** — each survived review because the comment read as settled: | # | Comment | What it claimed | What the code did | |---|---|---|---| | 1 | `EffectiveProtected` | a stack was protected | it was not — the samba false alarm | | 2 | `newestArchiveOn` | *"errors degrade to unknown, never to no-backup"* | the `(time,bool)` signature made that impossible (R-88 Part 2) | | 3 | R-97a operator-only | the event *"cannot be routed to a customer"* | only configuration stopped it; fixed by a real `operatorOnlyEvents` register | | 4 | `classifyRunStates` I1 | *"StateStopped means deliberately stopped by the user"* | quiesce stops stacks the same way — a failed restart was silent (F-CRIT-1) | | 5 | `inflight.go` | *"a caller that cannot acquire DEFERS"* | the backup caller recorded a failure and paged the operator (F-A1) | | 6 | `quiesce.go` | the agent's 409 *prevents* "a spurious failure" | on the start path it produced one (F-A1) | Two of these (4 and 5/6) were found by Campaign 8 **on live hardware**, not by review or unit tests — #4 had a green, red-proofed test suite over a production path that was broken two independent ways. So: - If a comment states an invariant, **name the test that pins it**, or write one. - If an invariant has a stated dependency (*"if either invariant changes, revisit this"*), that is not a safeguard — nobody revisits. Pin it with a test that fails when the dependency moves. - Prefer a test that asserts the **consequence** (does the alarm fire?) over one that asserts the **mechanism** (does suppression expire?). R-97b's Scenario F proved the mechanism and the consequence was still broken.