e6b5fa1e63
Follow-up acting on the observations filed with runbooks/target-selection.md. operations/nodes.md - The demo-hp NVMe was documented "PRESENT AND UNENROLLED -- do not touch" and listed under "What is NOT enrolled here (deliberately)". Both are FALSE and had been for eight days: it was enrolled 2026-07-22 through the normal Tarhely flow and is now /mnt/nvme-1tb -- the enrolled user-data drive AND the felhom-backup target (verified live 2026-07-30: nvme0n1 -> /mnt/nvme-1tb, and dir: felhom-backup / path /mnt/nvme-1tb / is_mountpoint 1). The fence's own condition (join via Tarhely, not the installer, not by hand) was SATISFIED, so the prohibition expired with it -- while still contradicting the task specs that correctly sent drill-VM disks there. Retracted with its reason recorded, and the caution that IS still live kept (dir storage at the mountpoint ROOT, else exactMount fails and the storage reads disconnected forever). - Component versions REMOVED and a note explains why: agent/controller/hub versions change several times a day, so a number written in an inventory is wrong within hours and then read as fact -- and the fleet is not uniform (on 2026-07-30 the two boxes ran different agent AND different controller versions). Points at the authorities instead: hub /hosts + /configs, felhom-agent --version, docker ps. - Site addresses now say re-check rather than asserting one (the N100 read .162, not the recorded .147); records that LAN literals are unreachable from DooPlex while the boxes are away. Adds the target-selection pointer: this page is what the hardware IS, that page is what may be done to it. PROMPT-TEMPLATE.md -- the upstream generator of the defect - Section 12's "Do NOT touch [the untouchable]" asked the spec author to name a THING. Now asks for the forbidden ACT plus its REASON, with the demo-hp case as the worked example of how a bare object-fence over-reads. - Section 13 gains the positive counterpart, which was the actual gap: if a task needs a machine to break, NAME IT. Listing only what is off-limits leaves the most valuable unfenced machine as the residual choice. runbooks/workspace-CLAUDE.md (+ the untracked root copy re-synced, verified identical) - Host table gains a Blast radius column and the missing demo-hp row, notes felhotest as Connection refused, and points at target-selection.md. This is the file that loads FIRST every session, so leaving it with the old table would have undercut the whole fix. No code, no build, no deploy, no host reconfigured or renamed.
203 lines
14 KiB
Markdown
203 lines
14 KiB
Markdown
# CLAUDE.md — `/mnt/5_hdd/felhom.eu/git` workspace root (DooPlex)
|
|
|
|
## What this workspace is
|
|
|
|
`/mnt/5_hdd/felhom.eu/git` is a parent folder holding the felhom sibling repos. Most are one logical
|
|
product — **Felhom**, a managed home-server service for Hungarian households — spread across several
|
|
repos. (Any non-felhom repo is unrelated; ignore unless asked.)
|
|
|
|
**Claude Code runs HERE, on DooPlex (192.168.0.180), as `kisfenyo`.** Builds are local commands; the
|
|
Proxmox host is one SSH hop (`ssh felhom-pve`). The Windows workstation is no longer the
|
|
orchestration point and its trees are stale — see "Legacy: Windows workstation" at the bottom.
|
|
|
|
Run CC inside tmux so sessions survive SSH drops: **`tmux new -A -s cc`**.
|
|
|
|
## This host is production infrastructure
|
|
|
|
DooPlex runs Gitea, the container registry, k3s + Longhorn, PBS, and the hub. Treat it accordingly:
|
|
|
|
- NEVER run `docker system prune`, `docker image prune -a`, or any global Docker cleanup here.
|
|
- NEVER touch k3s data dirs, Longhorn mounts, PBS datastores, or Gitea storage. Workspace is
|
|
`/mnt/5_hdd/felhom.eu/` — stay inside it plus `~/build` symlinks/dirs.
|
|
- Destructive disk/guest operations belong to felhom-pve via the agent — never on this host.
|
|
- Do not run Claude Code with permission prompts disabled on this host.
|
|
- Watch disk headroom before large builds: `df -h /mnt/5_hdd /` — abort if either is >90%.
|
|
|
|
## The Felhom system (three-component model, Proxmox-based)
|
|
|
|
- **Hub** — operator backend on k3s (`hub.felhom.eu`). Lives in `felhom.eu/hub/`.
|
|
- **Host agent** — one per Proxmox host, operator-tier, owns all Proxmox interaction. Repo `felhom-agent/`.
|
|
- **In-guest controller** — one per customer LXC, Docker-only. Repo `felhom-controller/`.
|
|
|
|
Other felhom repos: `app-catalog-felhom.eu/` (app templates), `homelab-manifests/` (DooPlex k3s).
|
|
|
|
**Authoritative design docs (read these before designing anything):** `felhom.eu/documentation/architecture/01..05-*.md`, `felhom.eu/documentation/proxmox-platform.md`, `felhom.eu/documentation/tests/phase{0,1-2,3,4}-findings.md`.
|
|
|
|
## Per-repo guidance
|
|
|
|
When you work in a repo, read its `CLAUDE.md` (it loads on-demand the moment you touch a file there):
|
|
- `felhom-agent/CLAUDE.md` — the Go host agent.
|
|
- `felhom.eu/CLAUDE.md` — hub + website + manifests + the architecture docs.
|
|
- `felhom-controller/CLAUDE.md` — the in-guest controller.
|
|
|
|
## Skills
|
|
|
|
Four Felhom skills exist (personal scope, `~/.claude/skills/`): **`felhom-build-deploy`** (all
|
|
build/deploy/publish runbooks), **`felhom-ui-design`** (design-system v2 tokens/rules/gates),
|
|
**`felhom-testing`** (non-hollow tests + red-proofs + seams), **`felhom-app-catalog`** (catalog
|
|
authoring workflow). Source of truth: `felhom.eu/skills/`; install/update with
|
|
`python3 felhom.eu/scripts/install_skills.py` (symlink — repo edits are live immediately).
|
|
|
|
## Memory
|
|
|
|
The accumulated project memory (119 files) migrated from the Windows workstation lives at
|
|
`/mnt/5_hdd/felhom.eu/git/.claude-memory/`, surfaced to Claude Code via
|
|
`~/.claude/projects/-mnt-5-hdd-felhom-eu-git/memory` (symlink). `MEMORY.md` there is the index.
|
|
Memories reflect what was true when written — verify a named file/flag still exists before acting
|
|
on it.
|
|
|
|
## Artifact taxonomy (READ THIS — it prevents the "what do I do?" stall)
|
|
|
|
The planning/architecture assistant (in claude.ai, "project Claude") produces files with distinct roles. **A file being open in the editor is NOT an instruction. If no task is stated, ask.**
|
|
|
|
- **`TASK.md` / `TASK-*.md`** — a spec for **you (Claude Code) to implement**. Implement it when it is placed as `TASK.md` at a repo root, or when explicitly told "implement <file>". Then push, update `CHANGELOG.md`, and write the repo's `REPORT.md`.
|
|
- **`RUNBOOK-*.md`** — an operational procedure. CC executes the steps it has access and capability for, including live validation on the demo nodes and the demo Proxmox host (CC has root@felhom-pve SSH + the felhom-agent token). A step is human-only only when it genuinely needs physical presence, a real-world decision, or credentials CC truly lacks — mark those steps HUMAN. Do not decline a whole procedure because it touches a live host or a privileged token. (Judgment still applies: confirm before irreversible ops on real customer data — but demo scratch guests are fair game.)
|
|
- **Validation/review** — checking a push against a spec's criteria is **project Claude's** job, not yours, unless asked.
|
|
|
|
## Shared conventions
|
|
|
|
### Standing rules — each one earned by a real failure (R-96, committed 2026-07-27)
|
|
|
|
These were agreed in conversation and lived nowhere, so they bound nobody. They do now.
|
|
|
|
1. **Never combine a test run and a commit in one command.** A combined command has ONE exit code and
|
|
the interesting one gets swallowed. Three recorded occurrences; the worst pushed a red suite
|
|
because `packages ok: 28` was read while `rc=1` was not. Run the suite, read `rc`, *then* commit.
|
|
|
|
2. **A "no access" claim must list what was tried.** "No access" is unfalsifiable unless it names its
|
|
attempts. Two wrong verdicts on 2026-07-27 alone: ep0 (declared unreachable after trying exactly
|
|
one route — `felhom-pve → 10.77.0.1`; `DooPlex → 167.233.158.164` worked and the project memory
|
|
said so), and the storage-box API (`api.hetzner.cloud` 404s for every storage-box endpoint;
|
|
`api.hetzner.com/v1` is the real one, and the hub's own `hetznerapi.go:3` records it).
|
|
|
|
3. **An absent log line is not evidence of correct behaviour.** Verify with a POSITIVE observable —
|
|
something that MUST appear when the system is healthy. An empty log is equally consistent with
|
|
"working" and "stopped entirely". Earned twice on 2026-07-27: the R-88 watcher (an empty quiesce
|
|
log could not distinguish a healthy loop from a dead one — retired in favour of the per-tier
|
|
`/backup/due` polls in `pveproxy/access.log`), and a hub DB copy whose write had silently failed,
|
|
returning a confident "0 events in window" from a file a day stale until its mtime was checked.
|
|
|
|
4. **A recommendation that is not followed gets one line saying why.** Silence reads as agreement and
|
|
the disagreement is lost. Twice in the R-88/R-97 arc a review point was absorbed rather than
|
|
argued: R-84 was folded into R-82 without a word, and R-97a's operator-only guard was dropped
|
|
while the claim it was meant to enforce got committed as a comment — which is how a false
|
|
guarantee shipped and survived a release. Disagreeing is fine; disagreeing silently is not.
|
|
|
|
- **Push to `main` directly** — no feature branches.
|
|
|
|
> **Clean-tree gate before any build:** `git status --porcelain` must be empty and
|
|
> `git rev-parse HEAD` must equal `git rev-parse origin/main` in the repo being built. An unpushed
|
|
> change does not exist — never build a dirty or unpushed tree. The `git pull` in the build step
|
|
> stays (it is a no-op when you work in this tree, and load-bearing if anything was pushed from
|
|
> elsewhere).
|
|
|
|
> **In every repository where you make a change, update both files in that repo:**
|
|
> - **`CHANGELOG.md`** — a cumulative log of **all** changes; newest entry on top.
|
|
> - **`REPORT.md`** — **overwrite** with a summary of the **most recent** implementation (or significant validation/operational run) only; not cumulative.
|
|
>
|
|
> **Never write secrets** — tokens, passwords, private keys, API keys — into `CHANGELOG.md`, `REPORT.md`, or any committed file. Reference them as "stored out-of-band" instead.
|
|
|
|
- **Versioning** is via build-time ldflags (`-X main.version`/`-X main.Version`); bump on meaningful changes + add a CHANGELOG entry.
|
|
- Code quality: double-check for bugs/edge cases; add debug logging; **ask rather than guess** when you'd otherwise need to invent input or output.
|
|
|
|
## Live validation — no browser here
|
|
|
|
**`claude-in-chrome` is NOT available on DooPlex.** The standard method is endpoint-level: invoke the
|
|
exact endpoint the UI invokes (no server logic is skipped, only rendering) and say which method was
|
|
used. Strict end-to-end UI coverage is a manual click-through by the operator.
|
|
|
|
## Access
|
|
|
|
Local (this host): repos `/mnt/5_hdd/felhom.eu/git/<repo>`, build dirs
|
|
`/mnt/5_hdd/felhom.eu/build/felhom-{controller,hub,agent}`, `sudo kubectl`, Go toolchain, Docker
|
|
build+push to `gitea.dooplex.hu/admin/`.
|
|
|
|
| Host | Access | Use | Blast radius |
|
|
|---|---|---|---|
|
|
| **DooPlex (this host)** | local — Debian 13, `kisfenyo`, `/mnt/5_hdd/felhom.eu/` | build/push images, `sudo kubectl`, build+run the agent for tests | **Tier 2 — precious.** It *is* the recovery chain (hub, Gitea, registry, PBS, k3s+Longhorn). **Never a drill target** |
|
|
| Demo Proxmox host `demo-hp` (HP t740) | `ssh demo-hp` (tailnet `100.76.96.79`; **no baked key** — G1 break-glass password vaulted in the hub) | **the designated drill + build VM host** (operator ruling 2026-07-25) | **Tier 0 — disposable. Reach here first** |
|
|
| Demo Proxmox host `demo-felhom` (N100) | `ssh felhom-pve` (root, no sudo; tailnet `100.70.170.35`) | pveum/pct + live Proxmox validation | **Tier 0 — disposable** |
|
|
| Demo guest 9201 | `ssh felhom-pve "pct exec 9201 -- ..."` | the live demo controller | Tier 0 (rides its host) |
|
|
| felhotest (legacy) | `ssh -p 33022 kisfenyo@router.abonet.hu` — **`Connection refused` 2026-07-30** | OLD /opt/docker compose mechanism | untiered — assume nothing |
|
|
|
|
**Which box do I break?** → **`felhom.eu/documentation/runbooks/target-selection.md`** — the tiers, and
|
|
per machine what is freely permitted / needs care / forbidden, each with its reason. Read it before
|
|
picking a machine for a drill, a destructive test or a throwaway VM. **A task that needs a victim names
|
|
one; an absent fence is not permission.**
|
|
|
|
**Component versions are not recorded in any inventory doc** — agent/controller/hub versions change
|
|
several times a day and the fleet is not uniform. Ask the hub's `/hosts` + `/configs`, or
|
|
`felhom-agent --version` / `pct exec <vmid> -- docker ps` on the box.
|
|
|
|
The demo Proxmox host key changes on reprovision (N100) → refresh with
|
|
`ssh-keygen -R 192.168.0.162` then connect with `-o StrictHostKeyChecking=accept-new`
|
|
(`ssh-keyscan` hangs — avoid it).
|
|
|
|
## Legacy: Windows workstation
|
|
|
|
Kept so the old environment can be revived; **not the current setup**.
|
|
|
|
- Repos were in `E:\git\` (`/e/git/` in Git Bash); this file lived at `E:\git\CLAUDE.md`.
|
|
- **SSH binary had to be** `SSH=/c/Windows/System32/OpenSSH/ssh.exe` — Git Bash's `/usr/bin/ssh`
|
|
lacks access to the Windows SSH Agent and fails silently. Every remote command was
|
|
`$SSH kisfenyo@192.168.0.180 "..."`; details in `felhom-controller/docs/vscode-ssh-fix.md`.
|
|
- `pct exec` over SSH needed `export MSYS_NO_PATHCONV=1` (MSYS mangled `/`-paths).
|
|
- Agent deploy was a two-hop copy: build on 180 → `scp` to the Windows box (local path needed
|
|
`cygpath -w`) → `scp` on to felhom-pve. Beware CRLF when scp-ing config files through Windows.
|
|
- Skills were installed as Windows junctions (`mklink /J`) rather than POSIX symlinks.
|
|
- `claude-in-chrome` browser automation WAS available there (attaching only to sessions started
|
|
after the bridge connected).
|
|
|
|
### Presence is not success
|
|
|
|
A timestamp recording an **attempt** must never be read as evidence of a **result**. Where a status
|
|
field travels alongside a timestamp, the verdict consults both — or the timestamp records only
|
|
successes.
|
|
|
|
| # | instance | what happened |
|
|
|---|---|---|
|
|
| 1 | **F-CRIT-2** | a phantom snapshot's ctime set tier freshness — an aborted 1-byte upload made the tier look backed up |
|
|
| 2 | **R-100** | `LastRun` is written on failure, so a nightly-failing offsite tier kept the staleness clock fresh forever |
|
|
|
|
Both were found by asking of a timestamp: *what exactly must have happened for this to be set?* If the
|
|
answer is "we tried", it cannot answer "did it work".
|
|
|
|
Corollary, from R-100's fix: when a verdict changes which field it counts from, **the alarm text has to
|
|
change with it**. Leaving the message reading `last run 8h ago` while alarming on a six-day-old success
|
|
turns a true alarm into one the operator dismisses.
|
|
|
|
### A comment asserting an invariant needs a test pinning it, or it is a wish
|
|
|
|
**Six instances in this project have shipped guarantees the code did not provide** — each survived
|
|
review because the comment read as settled:
|
|
|
|
| # | Comment | What it claimed | What the code did |
|
|
|---|---|---|---|
|
|
| 1 | `EffectiveProtected` | a stack was protected | it was not — the samba false alarm |
|
|
| 2 | `newestArchiveOn` | *"errors degrade to unknown, never to no-backup"* | the `(time,bool)` signature made that impossible (R-88 Part 2) |
|
|
| 3 | R-97a operator-only | the event *"cannot be routed to a customer"* | only configuration stopped it; fixed by a real `operatorOnlyEvents` register |
|
|
| 4 | `classifyRunStates` I1 | *"StateStopped means deliberately stopped by the user"* | quiesce stops stacks the same way — a failed restart was silent (F-CRIT-1) |
|
|
| 5 | `inflight.go` | *"a caller that cannot acquire DEFERS"* | the backup caller recorded a failure and paged the operator (F-A1) |
|
|
| 6 | `quiesce.go` | the agent's 409 *prevents* "a spurious failure" | on the start path it produced one (F-A1) |
|
|
|
|
Two of these (4 and 5/6) were found by Campaign 8 **on live hardware**, not by review or unit tests
|
|
— #4 had a green, red-proofed test suite over a production path that was broken two independent
|
|
ways. So:
|
|
|
|
- If a comment states an invariant, **name the test that pins it**, or write one.
|
|
- If an invariant has a stated dependency (*"if either invariant changes, revisit this"*), that is
|
|
not a safeguard — nobody revisits. Pin it with a test that fails when the dependency moves.
|
|
- Prefer a test that asserts the **consequence** (does the alarm fire?) over one that asserts the
|
|
**mechanism** (does suppression expire?). R-97b's Scenario F proved the mechanism and the
|
|
consequence was still broken.
|