Files
felhom.eu/documentation/runbooks/workspace-CLAUDE.md
T
admin e168600148 docs: F-CRIT-1 + F-A1 shipped (controller v0.179.0); invariant rule
Both marked SHIPPED + PROVEN-LIVE in OPEN-ITEMS and the campaign doc. All three
of Campaign 8's alarm findings are now closed (F-CRIT-1, F-CRIT-2, F-A1).

Adds the standing rule earned by this arc to the versioned workspace CLAUDE.md:
a comment asserting an invariant needs a test pinning it, or it is a wish — with
all six shipped-false-guarantee instances catalogued, and the corollary that a
test should assert the CONSEQUENCE (does the alarm fire?) not the MECHANISM
(does suppression expire?).
2026-07-28 09:48:51 +02:00

12 KiB

CLAUDE.md — /mnt/5_hdd/felhom.eu/git workspace root (DooPlex)

What this workspace is

/mnt/5_hdd/felhom.eu/git is a parent folder holding the felhom sibling repos. Most are one logical product — Felhom, a managed home-server service for Hungarian households — spread across several repos. (Any non-felhom repo is unrelated; ignore unless asked.)

Claude Code runs HERE, on DooPlex (192.168.0.180), as kisfenyo. Builds are local commands; the Proxmox host is one SSH hop (ssh felhom-pve). The Windows workstation is no longer the orchestration point and its trees are stale — see "Legacy: Windows workstation" at the bottom.

Run CC inside tmux so sessions survive SSH drops: tmux new -A -s cc.

This host is production infrastructure

DooPlex runs Gitea, the container registry, k3s + Longhorn, PBS, and the hub. Treat it accordingly:

  • NEVER run docker system prune, docker image prune -a, or any global Docker cleanup here.
  • NEVER touch k3s data dirs, Longhorn mounts, PBS datastores, or Gitea storage. Workspace is /mnt/5_hdd/felhom.eu/ — stay inside it plus ~/build symlinks/dirs.
  • Destructive disk/guest operations belong to felhom-pve via the agent — never on this host.
  • Do not run Claude Code with permission prompts disabled on this host.
  • Watch disk headroom before large builds: df -h /mnt/5_hdd / — abort if either is >90%.

The Felhom system (three-component model, Proxmox-based)

  • Hub — operator backend on k3s (hub.felhom.eu). Lives in felhom.eu/hub/.
  • Host agent — one per Proxmox host, operator-tier, owns all Proxmox interaction. Repo felhom-agent/.
  • In-guest controller — one per customer LXC, Docker-only. Repo felhom-controller/.

Other felhom repos: app-catalog-felhom.eu/ (app templates), homelab-manifests/ (DooPlex k3s).

Authoritative design docs (read these before designing anything): felhom.eu/documentation/architecture/01..05-*.md, felhom.eu/documentation/proxmox-platform.md, felhom.eu/documentation/tests/phase{0,1-2,3,4}-findings.md.

Per-repo guidance

When you work in a repo, read its CLAUDE.md (it loads on-demand the moment you touch a file there):

  • felhom-agent/CLAUDE.md — the Go host agent.
  • felhom.eu/CLAUDE.md — hub + website + manifests + the architecture docs.
  • felhom-controller/CLAUDE.md — the in-guest controller.

Skills

Four Felhom skills exist (personal scope, ~/.claude/skills/): felhom-build-deploy (all build/deploy/publish runbooks), felhom-ui-design (design-system v2 tokens/rules/gates), felhom-testing (non-hollow tests + red-proofs + seams), felhom-app-catalog (catalog authoring workflow). Source of truth: felhom.eu/skills/; install/update with python3 felhom.eu/scripts/install_skills.py (symlink — repo edits are live immediately).

Memory

The accumulated project memory (119 files) migrated from the Windows workstation lives at /mnt/5_hdd/felhom.eu/git/.claude-memory/, surfaced to Claude Code via ~/.claude/projects/-mnt-5-hdd-felhom-eu-git/memory (symlink). MEMORY.md there is the index. Memories reflect what was true when written — verify a named file/flag still exists before acting on it.

Artifact taxonomy (READ THIS — it prevents the "what do I do?" stall)

The planning/architecture assistant (in claude.ai, "project Claude") produces files with distinct roles. A file being open in the editor is NOT an instruction. If no task is stated, ask.

  • TASK.md / TASK-*.md — a spec for you (Claude Code) to implement. Implement it when it is placed as TASK.md at a repo root, or when explicitly told "implement ". Then push, update CHANGELOG.md, and write the repo's REPORT.md.
  • RUNBOOK-*.md — an operational procedure. CC executes the steps it has access and capability for, including live validation on the demo nodes and the demo Proxmox host (CC has root@felhom-pve SSH + the felhom-agent token). A step is human-only only when it genuinely needs physical presence, a real-world decision, or credentials CC truly lacks — mark those steps HUMAN. Do not decline a whole procedure because it touches a live host or a privileged token. (Judgment still applies: confirm before irreversible ops on real customer data — but demo scratch guests are fair game.)
  • Validation/review — checking a push against a spec's criteria is project Claude's job, not yours, unless asked.

Shared conventions

Standing rules — each one earned by a real failure (R-96, committed 2026-07-27)

These were agreed in conversation and lived nowhere, so they bound nobody. They do now.

  1. Never combine a test run and a commit in one command. A combined command has ONE exit code and the interesting one gets swallowed. Three recorded occurrences; the worst pushed a red suite because packages ok: 28 was read while rc=1 was not. Run the suite, read rc, then commit.

  2. A "no access" claim must list what was tried. "No access" is unfalsifiable unless it names its attempts. Two wrong verdicts on 2026-07-27 alone: ep0 (declared unreachable after trying exactly one route — felhom-pve → 10.77.0.1; DooPlex → 167.233.158.164 worked and the project memory said so), and the storage-box API (api.hetzner.cloud 404s for every storage-box endpoint; api.hetzner.com/v1 is the real one, and the hub's own hetznerapi.go:3 records it).

  3. An absent log line is not evidence of correct behaviour. Verify with a POSITIVE observable — something that MUST appear when the system is healthy. An empty log is equally consistent with "working" and "stopped entirely". Earned twice on 2026-07-27: the R-88 watcher (an empty quiesce log could not distinguish a healthy loop from a dead one — retired in favour of the per-tier /backup/due polls in pveproxy/access.log), and a hub DB copy whose write had silently failed, returning a confident "0 events in window" from a file a day stale until its mtime was checked.

  4. A recommendation that is not followed gets one line saying why. Silence reads as agreement and the disagreement is lost. Twice in the R-88/R-97 arc a review point was absorbed rather than argued: R-84 was folded into R-82 without a word, and R-97a's operator-only guard was dropped while the claim it was meant to enforce got committed as a comment — which is how a false guarantee shipped and survived a release. Disagreeing is fine; disagreeing silently is not.

  • Push to main directly — no feature branches.

Clean-tree gate before any build: git status --porcelain must be empty and git rev-parse HEAD must equal git rev-parse origin/main in the repo being built. An unpushed change does not exist — never build a dirty or unpushed tree. The git pull in the build step stays (it is a no-op when you work in this tree, and load-bearing if anything was pushed from elsewhere).

In every repository where you make a change, update both files in that repo:

  • CHANGELOG.md — a cumulative log of all changes; newest entry on top.
  • REPORT.mdoverwrite with a summary of the most recent implementation (or significant validation/operational run) only; not cumulative.

Never write secrets — tokens, passwords, private keys, API keys — into CHANGELOG.md, REPORT.md, or any committed file. Reference them as "stored out-of-band" instead.

  • Versioning is via build-time ldflags (-X main.version/-X main.Version); bump on meaningful changes + add a CHANGELOG entry.
  • Code quality: double-check for bugs/edge cases; add debug logging; ask rather than guess when you'd otherwise need to invent input or output.

Live validation — no browser here

claude-in-chrome is NOT available on DooPlex. The standard method is endpoint-level: invoke the exact endpoint the UI invokes (no server logic is skipped, only rendering) and say which method was used. Strict end-to-end UI coverage is a manual click-through by the operator.

Access

Local (this host): repos /mnt/5_hdd/felhom.eu/git/<repo>, build dirs /mnt/5_hdd/felhom.eu/build/felhom-{controller,hub,agent}, sudo kubectl, Go toolchain, Docker build+push to gitea.dooplex.hu/admin/.

Host Access Use
DooPlex (this host) local — Debian 13, kisfenyo, /mnt/5_hdd/felhom.eu/ build/push images, sudo kubectl, build+run the agent for tests
Demo Proxmox host demo-felhom ssh felhom-pve (root@192.168.0.162, no sudo) pveum/pct + live Proxmox validation
Demo guest 9201 ssh felhom-pve "pct exec 9201 -- ..." the live demo controller
felhotest (legacy) ssh -p 33022 kisfenyo@router.abonet.hu OLD /opt/docker compose mechanism

The demo Proxmox host key changes on reprovision (N100) → refresh with ssh-keygen -R 192.168.0.162 then connect with -o StrictHostKeyChecking=accept-new (ssh-keyscan hangs — avoid it).

Legacy: Windows workstation

Kept so the old environment can be revived; not the current setup.

  • Repos were in E:\git\ (/e/git/ in Git Bash); this file lived at E:\git\CLAUDE.md.
  • SSH binary had to be SSH=/c/Windows/System32/OpenSSH/ssh.exe — Git Bash's /usr/bin/ssh lacks access to the Windows SSH Agent and fails silently. Every remote command was $SSH kisfenyo@192.168.0.180 "..."; details in felhom-controller/docs/vscode-ssh-fix.md.
  • pct exec over SSH needed export MSYS_NO_PATHCONV=1 (MSYS mangled /-paths).
  • Agent deploy was a two-hop copy: build on 180 → scp to the Windows box (local path needed cygpath -w) → scp on to felhom-pve. Beware CRLF when scp-ing config files through Windows.
  • Skills were installed as Windows junctions (mklink /J) rather than POSIX symlinks.
  • claude-in-chrome browser automation WAS available there (attaching only to sessions started after the bridge connected).

A comment asserting an invariant needs a test pinning it, or it is a wish

Six instances in this project have shipped guarantees the code did not provide — each survived review because the comment read as settled:

# Comment What it claimed What the code did
1 EffectiveProtected a stack was protected it was not — the samba false alarm
2 newestArchiveOn "errors degrade to unknown, never to no-backup" the (time,bool) signature made that impossible (R-88 Part 2)
3 R-97a operator-only the event "cannot be routed to a customer" only configuration stopped it; fixed by a real operatorOnlyEvents register
4 classifyRunStates I1 "StateStopped means deliberately stopped by the user" quiesce stops stacks the same way — a failed restart was silent (F-CRIT-1)
5 inflight.go "a caller that cannot acquire DEFERS" the backup caller recorded a failure and paged the operator (F-A1)
6 quiesce.go the agent's 409 prevents "a spurious failure" on the start path it produced one (F-A1)

Two of these (4 and 5/6) were found by Campaign 8 on live hardware, not by review or unit tests — #4 had a green, red-proofed test suite over a production path that was broken two independent ways. So:

  • If a comment states an invariant, name the test that pins it, or write one.
  • If an invariant has a stated dependency ("if either invariant changes, revisit this"), that is not a safeguard — nobody revisits. Pin it with a test that fails when the dependency moves.
  • Prefer a test that asserts the consequence (does the alarm fire?) over one that asserts the mechanism (does suppression expire?). R-97b's Scenario F proved the mechanism and the consequence was still broken.