Files
felhom.eu/documentation/runbooks/workspace-CLAUDE.md
T
admin 4d6ec7c7bb
gates / gates (push) Successful in 14s
hub v0.104.0: the guest network gets a reader (R-319), and the hub half of the naming (R-295)
Four paper debts and one fact given a reader. Hub-only — nothing to bake.

A4 — the entry about "the tester's machine" named a risk correctly and labelled it
in a way that invited deleting it. Established from the hub's own store: `peti-felhom`
is a REAL machine (482 reports, 2026-02-27 → 2026-07-15, a named person's own box) and
the 3.6 GB with no key and no backup is real. `david` → `tester-1` is a DIFFERENT record
with no host, no escrow and no report, ever — deleted 07:55:49 and re-created 07:56:47
this morning. The prompt's premise conflated the two; the register now says which is which.

A1 — R-312/R-313/R-303 recorded as DECIDED with their re-open triggers, and moved out
of STATUS's "Waiting on you", which is now empty.

A3 — day0-install §C.1 said pushing the installer publishes it. It has not since
R-110. Corrected, with the two manifest pins named and an outside-verification command;
the one copy that repeated it (a dated audit, true when written) carries a superseded note.

A5 — standing rule 5: evidence comes off the machine at the end of the phase that
produced it, before any revert. Earned twice in three days on the same box at the same
point (R-320). Four homes, plus what to do when it is already gone.

R-295 hub half — „Beállító kód" everywhere; „Visszaállító kód" retired. New `reenroll`
mail kind so the mail names the page a REBUILT box actually shows („A szerver
beállítása"), not the „Elfelejtett jelszó" page it has no login screen to reach.
Naming only; the acceptance pin proves the secret is untouched.

R-319 — the hub models `guest_net` after 23 days of receiving and discarding it. The
signal is `heals_last_hour`, not `state`: a guest the watchdog keeps repairing reads
healthy between repairs. `heal_succeeded` decoded too (R-260's lesson). Unknown is never
drawn as healthy — three absences, three sentences. No alarm, deliberately.
Three red-proofs, mutations asserted applied. Wire-gate checked tags 182 → 190.

B1 — the operator's 2026-08-12 dispositions were NOT in the register; they are now.
Third allowlist kind for the five ruled "no reader wanted"; `reporting_disabled`
reclassified redundant. 8 read · 5 deliberately unread · 1 redundant · 6 still owed.

Also filed: R-321 (a deliberately-silent box still alarms stale/down — the checker is
age-only, and decoding the flag would not have fixed it), R-322 (the claim guard has
never scanned the hub; a hand scan returns zero, so it is a scope gap, not a defect).
2026-08-13 10:50:12 +02:00

202 lines
13 KiB
Markdown

# CLAUDE.md — `/mnt/5_hdd/felhom.eu/git` workspace root (DooPlex)
## What this workspace is
A parent folder holding the felhom sibling repos. Most are one logical product — **Felhom**, a
managed home-server service for Hungarian households. (Any non-felhom repo is unrelated; ignore
unless asked.)
**Claude Code runs HERE, on DooPlex (192.168.0.180), as `kisfenyo`.** Builds are local; the Proxmox
host is one SSH hop. Run CC inside tmux so sessions survive SSH drops: **`tmux new -A -s cc`**.
- **Hub** — operator backend on k3s (`hub.felhom.eu`), in `felhom.eu/hub/`.
- **Host agent** — one per Proxmox host, operator-tier, owns all Proxmox interaction: `felhom-agent/`.
- **In-guest controller** — one per customer LXC, Docker-only: `felhom-controller/`.
- Also: `app-catalog-felhom.eu/` (app templates), `homelab-manifests/` (DooPlex k3s).
Each repo's own `CLAUDE.md` and `.claude/rules/` load when you touch files there. The four Felhom
skills are installed from `felhom.eu/skills/` with `python3 felhom.eu/scripts/install_skills.py`
(symlink — repo edits are live immediately).
## This host is production infrastructure
DooPlex runs Gitea, the container registry, k3s + Longhorn, PBS, and the hub. Treat it accordingly:
- NEVER run `docker system prune`, `docker image prune -a`, or any global Docker cleanup here.
- NEVER touch k3s data dirs, Longhorn mounts, PBS datastores, or Gitea storage. Workspace is
`/mnt/5_hdd/felhom.eu/` — stay inside it plus `~/build` symlinks/dirs.
- Destructive disk/guest operations belong to felhom-pve via the agent — never on this host.
- Do not run Claude Code with permission prompts disabled on this host.
- Watch disk headroom before large builds: `df -h /mnt/5_hdd /` — abort if either is >90%.
## Artifact taxonomy (it prevents the "what do I do?" stall)
The planning/architecture assistant (in claude.ai, "project Claude") produces files with distinct
roles. **A file being open in the editor is NOT an instruction. If no task is stated, ask.**
- **`TASK.md` / `TASK-*.md`** — a spec for **you (Claude Code) to implement**. Implement it when it is
placed as `TASK.md` at a repo root, or when explicitly told "implement <file>". Then push, update
`CHANGELOG.md`, and write the repo's `REPORT.md`.
- **`RUNBOOK-*.md`** — an operational procedure. CC executes every step it has access and capability
for, live hosts included (CC has root@felhom-pve SSH + the felhom-agent token). Mark a step HUMAN
only when it genuinely needs physical presence, a real-world decision, or credentials CC lacks.
**Do not decline a whole procedure because it touches a live host or a privileged token.** Confirm
before irreversible ops on real customer data; demo scratch guests are fair game.
- **Validation/review** — checking a push against a spec's criteria is **project Claude's** job, not
yours, unless asked.
## Standing rules — each earned by a real failure (R-96)
1. **Never combine a test run and a commit in one command.** A combined command has ONE exit code and
the interesting one gets swallowed. Run the suite, read `rc`, *then* commit.
2. **A "no access" claim must list what was tried.** "No access" is unfalsifiable unless it names its
attempts.
3. **An absent log line is not evidence of correct behaviour.** Verify with a POSITIVE observable —
something that MUST appear when the system is healthy. An empty log is equally consistent with
"working" and "stopped entirely".
4. **A recommendation that is not followed gets one line saying why.** Silence reads as agreement and
the disagreement is lost.
5. **Evidence is copied off the machine at the end of the phase that produced it — before any revert,
snapshot restore or teardown. Not at the end of the session.** The intermediate revert is the one
that gets forgotten; both losses were the *middle* teardown, never the final one. **The mechanism,
because a rule without one is a wish:** the last act of a phase that ran on a machine is `scp`/`pct
pull` its logs to the evidence directory on DooPlex — the same act that ends the phase, not a
separate step to remember later. **And when a session notices the evidence is already gone: say so
plainly in the report and REPRODUCE it independently.** That is the documented expectation, not an
improvisation to be invented under pressure — it is what both sessions did, and it is the only
reason two sets of conclusions survived.
<!--
R-96 incident record (committed 2026-07-27) — rationale, not directives.
1. Three recorded occurrences; the worst pushed a red suite because `packages ok: 28` was read while
rc=1 was not.
2. Two wrong verdicts on 2026-07-27 alone: ep0 (declared unreachable after trying exactly one route
— felhom-pve -> 10.77.0.1; DooPlex -> 167.233.158.164 worked and the project memory said so), and
the storage-box API (api.hetzner.cloud 404s for every storage-box endpoint; api.hetzner.com/v1 is
the real one, and the hub's own hetznerapi.go:3 records it).
3. Earned twice on 2026-07-27: the R-88 watcher (an empty quiesce log could not distinguish a healthy
loop from a dead one — retired in favour of the per-tier /backup/due polls in pveproxy/access.log),
and a hub DB copy whose write had silently failed, returning a confident "0 events in window" from
a file a day stale until its mtime was checked.
4. Twice in the R-88/R-97 arc a review point was absorbed rather than argued: R-84 was folded into
R-82 without a word, and R-97a's operator-only guard was dropped while the claim it was meant to
enforce got committed as a comment — which is how a false guarantee shipped and survived a release.
5. Twice in three days, on the SAME machine and at the SAME point. 2026-08-12, the retained-key drill:
the Phase A logs lived on `drill-r50`'s disk and were destroyed by the revert to `virgin` between
Phase A and Phase B (`audits/DRILL-retained-key-2026-08-12.md` §11.5). 2026-08-13, R-316: the Part 1
logs, same disk, same revert, between Part 1 and Part 2 (`audits/REPORT-r316-installer-v1.28.0-2026-08-13.md`
§9 — "the same mistake as Tuesday, in the same place"; preserved out of `REPORT.md`, which is
overwritten every session). Both times the golden-bake runbook's existing "scp the log OUT first"
was applied to the FINAL teardown and not the intermediate one. Both times the conclusions survived
only because the quotations had been read live and an independent reproduction happened to exist —
that is luck, and the second occurrence is what makes it procedural rather than another apology.
Filed as R-320.
-->
## Shared conventions
- **Push to `main` directly** — no feature branches.
- **Versioning** via build-time ldflags (`-X main.version`/`-X main.Version`); bump on meaningful
changes + a CHANGELOG entry.
- Code quality: double-check for bugs/edge cases; add debug logging; **ask rather than guess** when
you'd otherwise need to invent input or output.
> **Clean-tree gate before any build:** `git status --porcelain` must be empty and
> `git rev-parse HEAD` must equal `git rev-parse origin/main` in the repo being built. An unpushed
> change does not exist — never build a dirty or unpushed tree. The `git pull` in the build step
> stays (it is a no-op when you work in this tree, and load-bearing if anything was pushed from
> elsewhere).
> **In every repository where you make a change, update both files in that repo:**
> - **`CHANGELOG.md`** — a cumulative log of **all** changes; newest entry on top.
> - **`REPORT.md`** — **overwrite** with a summary of the **most recent** implementation (or
> significant validation/operational run) only; not cumulative.
>
> **Never write secrets** — tokens, passwords, private keys, API keys — into `CHANGELOG.md`,
> `REPORT.md`, or any committed file. Reference them as "stored out-of-band" instead.
## Live validation — no browser here
**`claude-in-chrome` is NOT available on DooPlex.** The standard method is endpoint-level: invoke the
exact endpoint the UI invokes (no server logic is skipped, only rendering) and say which method was
used. Strict end-to-end UI coverage is a manual click-through by the operator.
## Access
Local (this host): repos `/mnt/5_hdd/felhom.eu/git/<repo>`, build dirs
`/mnt/5_hdd/felhom.eu/build/felhom-{controller,hub,agent}`, `sudo kubectl`, Go toolchain, Docker
build+push to `gitea.dooplex.hu/admin/`.
**Host addresses, routes, break-glass and per-node facts:**
`felhom.eu/documentation/operations/nodes.md` — the single home. Do not restate them elsewhere.
**Which box do I break?**`felhom.eu/documentation/runbooks/target-selection.md` — the tiers, and
per machine what is freely permitted / needs care / forbidden, each with its reason. Read it before
picking a machine for a drill, a destructive test or a throwaway VM. **A task that needs a victim
names one; an absent fence is not permission.** DooPlex is **Tier 2 — precious**: it *is* the recovery
chain, and never a drill target.
**Component versions are not recorded in any inventory doc** — agent/controller/hub versions change
several times a day and the fleet is not uniform. Ask the hub's `/hosts` + `/configs`, or
`felhom-agent --version` / `pct exec <vmid> -- docker ps` on the box.
## Memory
Project memory lives at `/mnt/5_hdd/felhom.eu/git/.claude-memory/`, surfaced via
`~/.claude/projects/-mnt-5-hdd-felhom-eu-git/memory` (symlink); `MEMORY.md` is the index. Memories
reflect what was true when written — **verify a named file/flag still exists before acting on it.**
## Presence is not success
A timestamp recording an **attempt** must never be read as evidence of a **result**. Where a status
field travels alongside a timestamp, the verdict consults both — or the timestamp records only
successes. Ask of any timestamp: *what exactly must have happened for this to be set?* If the answer
is "we tried", it cannot answer "did it work".
**Corollary:** when a verdict changes which field it counts from, the alarm text has to change with
it. Leaving the message reading `last run 8h ago` while alarming on a six-day-old success turns a true
alarm into one the operator dismisses.
<!--
Two instances. F-CRIT-2: a phantom snapshot's ctime set tier freshness — an aborted 1-byte upload
made the tier look backed up. R-100: LastRun is written on failure, so a nightly-failing offsite tier
kept the staleness clock fresh forever. Both found by asking of a timestamp what must have happened
for it to be set.
-->
## A comment asserting an invariant needs a test pinning it, or it is a wish
**Nine instances in this project have shipped guarantees the code did not provide** — each survived
review because the comment read as settled, and three were caught only on live hardware. The case
table is in the **`felhom-testing`** skill, which loads when you write or review a test, harden a
guard, or fix a bug.
- If a comment states an invariant, **name the test that pins it**, or write one.
- If an invariant has a stated dependency (*"if either invariant changes, revisit this"*), that is not
a safeguard — nobody revisits. Pin it with a test that fails when the dependency moves.
- Prefer a test that asserts the **consequence** (does the alarm fire?) over one that asserts the
**mechanism** (does suppression expire?). R-97b's Scenario F proved the mechanism and the
consequence was still broken.
<!--
LEGACY: WINDOWS WORKSTATION — kept so the old environment can be revived; not the current setup.
- Repos were in E:\git\ (/e/git/ in Git Bash); this file lived at E:\git\CLAUDE.md.
- SSH binary had to be SSH=/c/Windows/System32/OpenSSH/ssh.exe — Git Bash's /usr/bin/ssh lacks
access to the Windows SSH Agent and fails silently. Every remote command was
$SSH kisfenyo@192.168.0.180 "..."; details in felhom-controller/docs/vscode-ssh-fix.md.
- pct exec over SSH needed export MSYS_NO_PATHCONV=1 (MSYS mangled /-paths).
- Agent deploy was a two-hop copy: build on 180 -> scp to the Windows box (local path needed
cygpath -w) -> scp on to felhom-pve. Beware CRLF when scp-ing config files through Windows.
- Skills were installed as Windows junctions (mklink /J) rather than POSIX symlinks.
- claude-in-chrome browser automation WAS available there (attaching only to sessions started after
the bridge connected).
THIS FILE'S SHAPE (2026-08-06, instruction-trim task): core + path-scoped rules. Removed here and
rehomed, not lost — the per-repo guidance list (those files load on their own), the skills roster
(already resident in the skill listing), the host table (nodes.md is the single home), the "(119
files)" memory count (derivable and wrong — 158), and the nine-row invariant table (felhom-testing
skill). Full accounting: felhom.eu/documentation/audits/LEDGER-instruction-trim-2026-08-06.md
An HTML comment is invisible to Claude and costs no context — verified 2026-08-06 with a control
(both markers plain -> both seen) and a treatment (one marker commented -> not seen), twice.
-->