Files
felhom.eu/documentation/runbooks/workspace-CLAUDE.md
T
admin 4d6ec7c7bb
gates / gates (push) Successful in 14s
hub v0.104.0: the guest network gets a reader (R-319), and the hub half of the naming (R-295)
Four paper debts and one fact given a reader. Hub-only — nothing to bake.

A4 — the entry about "the tester's machine" named a risk correctly and labelled it
in a way that invited deleting it. Established from the hub's own store: `peti-felhom`
is a REAL machine (482 reports, 2026-02-27 → 2026-07-15, a named person's own box) and
the 3.6 GB with no key and no backup is real. `david` → `tester-1` is a DIFFERENT record
with no host, no escrow and no report, ever — deleted 07:55:49 and re-created 07:56:47
this morning. The prompt's premise conflated the two; the register now says which is which.

A1 — R-312/R-313/R-303 recorded as DECIDED with their re-open triggers, and moved out
of STATUS's "Waiting on you", which is now empty.

A3 — day0-install §C.1 said pushing the installer publishes it. It has not since
R-110. Corrected, with the two manifest pins named and an outside-verification command;
the one copy that repeated it (a dated audit, true when written) carries a superseded note.

A5 — standing rule 5: evidence comes off the machine at the end of the phase that
produced it, before any revert. Earned twice in three days on the same box at the same
point (R-320). Four homes, plus what to do when it is already gone.

R-295 hub half — „Beállító kód" everywhere; „Visszaállító kód" retired. New `reenroll`
mail kind so the mail names the page a REBUILT box actually shows („A szerver
beállítása"), not the „Elfelejtett jelszó" page it has no login screen to reach.
Naming only; the acceptance pin proves the secret is untouched.

R-319 — the hub models `guest_net` after 23 days of receiving and discarding it. The
signal is `heals_last_hour`, not `state`: a guest the watchdog keeps repairing reads
healthy between repairs. `heal_succeeded` decoded too (R-260's lesson). Unknown is never
drawn as healthy — three absences, three sentences. No alarm, deliberately.
Three red-proofs, mutations asserted applied. Wire-gate checked tags 182 → 190.

B1 — the operator's 2026-08-12 dispositions were NOT in the register; they are now.
Third allowlist kind for the five ruled "no reader wanted"; `reporting_disabled`
reclassified redundant. 8 read · 5 deliberately unread · 1 redundant · 6 still owed.

Also filed: R-321 (a deliberately-silent box still alarms stale/down — the checker is
age-only, and decoding the flag would not have fixed it), R-322 (the claim guard has
never scanned the hub; a hand scan returns zero, so it is a scope gap, not a defect).
2026-08-13 10:50:12 +02:00

13 KiB

CLAUDE.md — /mnt/5_hdd/felhom.eu/git workspace root (DooPlex)

What this workspace is

A parent folder holding the felhom sibling repos. Most are one logical product — Felhom, a managed home-server service for Hungarian households. (Any non-felhom repo is unrelated; ignore unless asked.)

Claude Code runs HERE, on DooPlex (192.168.0.180), as kisfenyo. Builds are local; the Proxmox host is one SSH hop. Run CC inside tmux so sessions survive SSH drops: tmux new -A -s cc.

  • Hub — operator backend on k3s (hub.felhom.eu), in felhom.eu/hub/.
  • Host agent — one per Proxmox host, operator-tier, owns all Proxmox interaction: felhom-agent/.
  • In-guest controller — one per customer LXC, Docker-only: felhom-controller/.
  • Also: app-catalog-felhom.eu/ (app templates), homelab-manifests/ (DooPlex k3s).

Each repo's own CLAUDE.md and .claude/rules/ load when you touch files there. The four Felhom skills are installed from felhom.eu/skills/ with python3 felhom.eu/scripts/install_skills.py (symlink — repo edits are live immediately).

This host is production infrastructure

DooPlex runs Gitea, the container registry, k3s + Longhorn, PBS, and the hub. Treat it accordingly:

  • NEVER run docker system prune, docker image prune -a, or any global Docker cleanup here.
  • NEVER touch k3s data dirs, Longhorn mounts, PBS datastores, or Gitea storage. Workspace is /mnt/5_hdd/felhom.eu/ — stay inside it plus ~/build symlinks/dirs.
  • Destructive disk/guest operations belong to felhom-pve via the agent — never on this host.
  • Do not run Claude Code with permission prompts disabled on this host.
  • Watch disk headroom before large builds: df -h /mnt/5_hdd / — abort if either is >90%.

Artifact taxonomy (it prevents the "what do I do?" stall)

The planning/architecture assistant (in claude.ai, "project Claude") produces files with distinct roles. A file being open in the editor is NOT an instruction. If no task is stated, ask.

  • TASK.md / TASK-*.md — a spec for you (Claude Code) to implement. Implement it when it is placed as TASK.md at a repo root, or when explicitly told "implement ". Then push, update CHANGELOG.md, and write the repo's REPORT.md.
  • RUNBOOK-*.md — an operational procedure. CC executes every step it has access and capability for, live hosts included (CC has root@felhom-pve SSH + the felhom-agent token). Mark a step HUMAN only when it genuinely needs physical presence, a real-world decision, or credentials CC lacks. Do not decline a whole procedure because it touches a live host or a privileged token. Confirm before irreversible ops on real customer data; demo scratch guests are fair game.
  • Validation/review — checking a push against a spec's criteria is project Claude's job, not yours, unless asked.

Standing rules — each earned by a real failure (R-96)

  1. Never combine a test run and a commit in one command. A combined command has ONE exit code and the interesting one gets swallowed. Run the suite, read rc, then commit.
  2. A "no access" claim must list what was tried. "No access" is unfalsifiable unless it names its attempts.
  3. An absent log line is not evidence of correct behaviour. Verify with a POSITIVE observable — something that MUST appear when the system is healthy. An empty log is equally consistent with "working" and "stopped entirely".
  4. A recommendation that is not followed gets one line saying why. Silence reads as agreement and the disagreement is lost.
  5. Evidence is copied off the machine at the end of the phase that produced it — before any revert, snapshot restore or teardown. Not at the end of the session. The intermediate revert is the one that gets forgotten; both losses were the middle teardown, never the final one. The mechanism, because a rule without one is a wish: the last act of a phase that ran on a machine is scp/pct pull its logs to the evidence directory on DooPlex — the same act that ends the phase, not a separate step to remember later. And when a session notices the evidence is already gone: say so plainly in the report and REPRODUCE it independently. That is the documented expectation, not an improvisation to be invented under pressure — it is what both sessions did, and it is the only reason two sets of conclusions survived.

Shared conventions

  • Push to main directly — no feature branches.
  • Versioning via build-time ldflags (-X main.version/-X main.Version); bump on meaningful changes + a CHANGELOG entry.
  • Code quality: double-check for bugs/edge cases; add debug logging; ask rather than guess when you'd otherwise need to invent input or output.

Clean-tree gate before any build: git status --porcelain must be empty and git rev-parse HEAD must equal git rev-parse origin/main in the repo being built. An unpushed change does not exist — never build a dirty or unpushed tree. The git pull in the build step stays (it is a no-op when you work in this tree, and load-bearing if anything was pushed from elsewhere).

In every repository where you make a change, update both files in that repo:

  • CHANGELOG.md — a cumulative log of all changes; newest entry on top.
  • REPORT.mdoverwrite with a summary of the most recent implementation (or significant validation/operational run) only; not cumulative.

Never write secrets — tokens, passwords, private keys, API keys — into CHANGELOG.md, REPORT.md, or any committed file. Reference them as "stored out-of-band" instead.

Live validation — no browser here

claude-in-chrome is NOT available on DooPlex. The standard method is endpoint-level: invoke the exact endpoint the UI invokes (no server logic is skipped, only rendering) and say which method was used. Strict end-to-end UI coverage is a manual click-through by the operator.

Access

Local (this host): repos /mnt/5_hdd/felhom.eu/git/<repo>, build dirs /mnt/5_hdd/felhom.eu/build/felhom-{controller,hub,agent}, sudo kubectl, Go toolchain, Docker build+push to gitea.dooplex.hu/admin/.

Host addresses, routes, break-glass and per-node facts: felhom.eu/documentation/operations/nodes.md — the single home. Do not restate them elsewhere.

Which box do I break?felhom.eu/documentation/runbooks/target-selection.md — the tiers, and per machine what is freely permitted / needs care / forbidden, each with its reason. Read it before picking a machine for a drill, a destructive test or a throwaway VM. A task that needs a victim names one; an absent fence is not permission. DooPlex is Tier 2 — precious: it is the recovery chain, and never a drill target.

Component versions are not recorded in any inventory doc — agent/controller/hub versions change several times a day and the fleet is not uniform. Ask the hub's /hosts + /configs, or felhom-agent --version / pct exec <vmid> -- docker ps on the box.

Memory

Project memory lives at /mnt/5_hdd/felhom.eu/git/.claude-memory/, surfaced via ~/.claude/projects/-mnt-5-hdd-felhom-eu-git/memory (symlink); MEMORY.md is the index. Memories reflect what was true when written — verify a named file/flag still exists before acting on it.

Presence is not success

A timestamp recording an attempt must never be read as evidence of a result. Where a status field travels alongside a timestamp, the verdict consults both — or the timestamp records only successes. Ask of any timestamp: what exactly must have happened for this to be set? If the answer is "we tried", it cannot answer "did it work".

Corollary: when a verdict changes which field it counts from, the alarm text has to change with it. Leaving the message reading last run 8h ago while alarming on a six-day-old success turns a true alarm into one the operator dismisses.

A comment asserting an invariant needs a test pinning it, or it is a wish

Nine instances in this project have shipped guarantees the code did not provide — each survived review because the comment read as settled, and three were caught only on live hardware. The case table is in the felhom-testing skill, which loads when you write or review a test, harden a guard, or fix a bug.

  • If a comment states an invariant, name the test that pins it, or write one.
  • If an invariant has a stated dependency ("if either invariant changes, revisit this"), that is not a safeguard — nobody revisits. Pin it with a test that fails when the dependency moves.
  • Prefer a test that asserts the consequence (does the alarm fire?) over one that asserts the mechanism (does suppression expire?). R-97b's Scenario F proved the mechanism and the consequence was still broken.