Registers and evidence for the campaign and its fix pass. OPEN-ITEMS: R-214..R-223. Six SHIPPED (R-215/216/217/218/219/222); three deliberately still open and each blocks a real flow (R-214 console banner, R-220 drives unenrollable after a rebuild, R-221 a rebuilt box cannot run the escrow ceremony); R-223 minted and WAITING-ON-OPERATOR (vouch agent 0.125.0). R-213 and R-202 untouched. Capability map: a new row for the customer's UNAIDED journey, recorded FAILED and staying failed until a re-walk passes — fixes are not a journey. The existing rebuild row is corrected where it said R-198's retention was unit-proven only: it was proven in production on the first supersession since the fix, identity_blob retained at 572 B byte-length exact. CLAUDE.md comment-vs-code table: eighth entry — ResolveManagedFloor, the first where the false invariant was a GUARD rather than a comment alone. STATUS: the headline is now "the backup promise is proved, the recovery journey is not", and the one thing waiting on the operator.
15 KiB
CLAUDE.md — /mnt/5_hdd/felhom.eu/git workspace root (DooPlex)
What this workspace is
/mnt/5_hdd/felhom.eu/git is a parent folder holding the felhom sibling repos. Most are one logical
product — Felhom, a managed home-server service for Hungarian households — spread across several
repos. (Any non-felhom repo is unrelated; ignore unless asked.)
Claude Code runs HERE, on DooPlex (192.168.0.180), as kisfenyo. Builds are local commands; the
Proxmox host is one SSH hop (ssh felhom-pve). The Windows workstation is no longer the
orchestration point and its trees are stale — see "Legacy: Windows workstation" at the bottom.
Run CC inside tmux so sessions survive SSH drops: tmux new -A -s cc.
This host is production infrastructure
DooPlex runs Gitea, the container registry, k3s + Longhorn, PBS, and the hub. Treat it accordingly:
- NEVER run
docker system prune,docker image prune -a, or any global Docker cleanup here. - NEVER touch k3s data dirs, Longhorn mounts, PBS datastores, or Gitea storage. Workspace is
/mnt/5_hdd/felhom.eu/— stay inside it plus~/buildsymlinks/dirs. - Destructive disk/guest operations belong to felhom-pve via the agent — never on this host.
- Do not run Claude Code with permission prompts disabled on this host.
- Watch disk headroom before large builds:
df -h /mnt/5_hdd /— abort if either is >90%.
The Felhom system (three-component model, Proxmox-based)
- Hub — operator backend on k3s (
hub.felhom.eu). Lives infelhom.eu/hub/. - Host agent — one per Proxmox host, operator-tier, owns all Proxmox interaction. Repo
felhom-agent/. - In-guest controller — one per customer LXC, Docker-only. Repo
felhom-controller/.
Other felhom repos: app-catalog-felhom.eu/ (app templates), homelab-manifests/ (DooPlex k3s).
Authoritative design docs (read these before designing anything): felhom.eu/documentation/architecture/01..05-*.md, felhom.eu/documentation/proxmox-platform.md, felhom.eu/documentation/tests/phase{0,1-2,3,4}-findings.md.
Per-repo guidance
When you work in a repo, read its CLAUDE.md (it loads on-demand the moment you touch a file there):
felhom-agent/CLAUDE.md— the Go host agent.felhom.eu/CLAUDE.md— hub + website + manifests + the architecture docs.felhom-controller/CLAUDE.md— the in-guest controller.
Skills
Four Felhom skills exist (personal scope, ~/.claude/skills/): felhom-build-deploy (all
build/deploy/publish runbooks), felhom-ui-design (design-system v2 tokens/rules/gates),
felhom-testing (non-hollow tests + red-proofs + seams), felhom-app-catalog (catalog
authoring workflow). Source of truth: felhom.eu/skills/; install/update with
python3 felhom.eu/scripts/install_skills.py (symlink — repo edits are live immediately).
Memory
The accumulated project memory (119 files) migrated from the Windows workstation lives at
/mnt/5_hdd/felhom.eu/git/.claude-memory/, surfaced to Claude Code via
~/.claude/projects/-mnt-5-hdd-felhom-eu-git/memory (symlink). MEMORY.md there is the index.
Memories reflect what was true when written — verify a named file/flag still exists before acting
on it.
Artifact taxonomy (READ THIS — it prevents the "what do I do?" stall)
The planning/architecture assistant (in claude.ai, "project Claude") produces files with distinct roles. A file being open in the editor is NOT an instruction. If no task is stated, ask.
TASK.md/TASK-*.md— a spec for you (Claude Code) to implement. Implement it when it is placed asTASK.mdat a repo root, or when explicitly told "implement ". Then push, updateCHANGELOG.md, and write the repo'sREPORT.md.RUNBOOK-*.md— an operational procedure. CC executes the steps it has access and capability for, including live validation on the demo nodes and the demo Proxmox host (CC has root@felhom-pve SSH + the felhom-agent token). A step is human-only only when it genuinely needs physical presence, a real-world decision, or credentials CC truly lacks — mark those steps HUMAN. Do not decline a whole procedure because it touches a live host or a privileged token. (Judgment still applies: confirm before irreversible ops on real customer data — but demo scratch guests are fair game.)- Validation/review — checking a push against a spec's criteria is project Claude's job, not yours, unless asked.
Shared conventions
Standing rules — each one earned by a real failure (R-96, committed 2026-07-27)
These were agreed in conversation and lived nowhere, so they bound nobody. They do now.
-
Never combine a test run and a commit in one command. A combined command has ONE exit code and the interesting one gets swallowed. Three recorded occurrences; the worst pushed a red suite because
packages ok: 28was read whilerc=1was not. Run the suite, readrc, then commit. -
A "no access" claim must list what was tried. "No access" is unfalsifiable unless it names its attempts. Two wrong verdicts on 2026-07-27 alone: ep0 (declared unreachable after trying exactly one route —
felhom-pve → 10.77.0.1;DooPlex → 167.233.158.164worked and the project memory said so), and the storage-box API (api.hetzner.cloud404s for every storage-box endpoint;api.hetzner.com/v1is the real one, and the hub's ownhetznerapi.go:3records it). -
An absent log line is not evidence of correct behaviour. Verify with a POSITIVE observable — something that MUST appear when the system is healthy. An empty log is equally consistent with "working" and "stopped entirely". Earned twice on 2026-07-27: the R-88 watcher (an empty quiesce log could not distinguish a healthy loop from a dead one — retired in favour of the per-tier
/backup/duepolls inpveproxy/access.log), and a hub DB copy whose write had silently failed, returning a confident "0 events in window" from a file a day stale until its mtime was checked. -
A recommendation that is not followed gets one line saying why. Silence reads as agreement and the disagreement is lost. Twice in the R-88/R-97 arc a review point was absorbed rather than argued: R-84 was folded into R-82 without a word, and R-97a's operator-only guard was dropped while the claim it was meant to enforce got committed as a comment — which is how a false guarantee shipped and survived a release. Disagreeing is fine; disagreeing silently is not.
- Push to
maindirectly — no feature branches.
Clean-tree gate before any build:
git status --porcelainmust be empty andgit rev-parse HEADmust equalgit rev-parse origin/mainin the repo being built. An unpushed change does not exist — never build a dirty or unpushed tree. Thegit pullin the build step stays (it is a no-op when you work in this tree, and load-bearing if anything was pushed from elsewhere).
In every repository where you make a change, update both files in that repo:
CHANGELOG.md— a cumulative log of all changes; newest entry on top.REPORT.md— overwrite with a summary of the most recent implementation (or significant validation/operational run) only; not cumulative.Never write secrets — tokens, passwords, private keys, API keys — into
CHANGELOG.md,REPORT.md, or any committed file. Reference them as "stored out-of-band" instead.
- Versioning is via build-time ldflags (
-X main.version/-X main.Version); bump on meaningful changes + add a CHANGELOG entry. - Code quality: double-check for bugs/edge cases; add debug logging; ask rather than guess when you'd otherwise need to invent input or output.
Live validation — no browser here
claude-in-chrome is NOT available on DooPlex. The standard method is endpoint-level: invoke the
exact endpoint the UI invokes (no server logic is skipped, only rendering) and say which method was
used. Strict end-to-end UI coverage is a manual click-through by the operator.
Access
Local (this host): repos /mnt/5_hdd/felhom.eu/git/<repo>, build dirs
/mnt/5_hdd/felhom.eu/build/felhom-{controller,hub,agent}, sudo kubectl, Go toolchain, Docker
build+push to gitea.dooplex.hu/admin/.
| Host | Access | Use | Blast radius |
|---|---|---|---|
| DooPlex (this host) | local — Debian 13, kisfenyo, /mnt/5_hdd/felhom.eu/ |
build/push images, sudo kubectl, build+run the agent for tests |
Tier 2 — precious. It is the recovery chain (hub, Gitea, registry, PBS, k3s+Longhorn). Never a drill target |
Demo Proxmox host demo-hp (HP t740) |
ssh demo-hp (tailnet 100.76.96.79; no baked key — G1 break-glass password vaulted in the hub) |
the designated drill + build VM host (operator ruling 2026-07-25) | Tier 0 — disposable. Reach here first |
Demo Proxmox host demo-felhom (N100) |
ssh felhom-pve (root, no sudo; tailnet 100.70.170.35) |
pveum/pct + live Proxmox validation | Tier 0 — disposable |
| Demo guest 9201 | ssh felhom-pve "pct exec 9201 -- ..." |
the live demo controller | Tier 0 (rides its host) |
| felhotest (legacy) | ssh -p 33022 kisfenyo@router.abonet.hu — Connection refused 2026-07-30 |
OLD /opt/docker compose mechanism | untiered — assume nothing |
Which box do I break? → felhom.eu/documentation/runbooks/target-selection.md — the tiers, and
per machine what is freely permitted / needs care / forbidden, each with its reason. Read it before
picking a machine for a drill, a destructive test or a throwaway VM. A task that needs a victim names
one; an absent fence is not permission.
Component versions are not recorded in any inventory doc — agent/controller/hub versions change
several times a day and the fleet is not uniform. Ask the hub's /hosts + /configs, or
felhom-agent --version / pct exec <vmid> -- docker ps on the box.
The demo Proxmox host key changes on reprovision (N100) → refresh with
ssh-keygen -R 192.168.0.162 then connect with -o StrictHostKeyChecking=accept-new
(ssh-keyscan hangs — avoid it).
Legacy: Windows workstation
Kept so the old environment can be revived; not the current setup.
- Repos were in
E:\git\(/e/git/in Git Bash); this file lived atE:\git\CLAUDE.md. - SSH binary had to be
SSH=/c/Windows/System32/OpenSSH/ssh.exe— Git Bash's/usr/bin/sshlacks access to the Windows SSH Agent and fails silently. Every remote command was$SSH kisfenyo@192.168.0.180 "..."; details infelhom-controller/docs/vscode-ssh-fix.md. pct execover SSH neededexport MSYS_NO_PATHCONV=1(MSYS mangled/-paths).- Agent deploy was a two-hop copy: build on 180 →
scpto the Windows box (local path neededcygpath -w) →scpon to felhom-pve. Beware CRLF when scp-ing config files through Windows. - Skills were installed as Windows junctions (
mklink /J) rather than POSIX symlinks. claude-in-chromebrowser automation WAS available there (attaching only to sessions started after the bridge connected).
Presence is not success
A timestamp recording an attempt must never be read as evidence of a result. Where a status field travels alongside a timestamp, the verdict consults both — or the timestamp records only successes.
| # | instance | what happened |
|---|---|---|
| 1 | F-CRIT-2 | a phantom snapshot's ctime set tier freshness — an aborted 1-byte upload made the tier look backed up |
| 2 | R-100 | LastRun is written on failure, so a nightly-failing offsite tier kept the staleness clock fresh forever |
Both were found by asking of a timestamp: what exactly must have happened for this to be set? If the answer is "we tried", it cannot answer "did it work".
Corollary, from R-100's fix: when a verdict changes which field it counts from, the alarm text has to
change with it. Leaving the message reading last run 8h ago while alarming on a six-day-old success
turns a true alarm into one the operator dismisses.
A comment asserting an invariant needs a test pinning it, or it is a wish
Eight instances in this project have shipped guarantees the code did not provide — each survived review because the comment read as settled:
| # | Comment | What it claimed | What the code did |
|---|---|---|---|
| 1 | EffectiveProtected |
a stack was protected | it was not — the samba false alarm |
| 2 | newestArchiveOn |
"errors degrade to unknown, never to no-backup" | the (time,bool) signature made that impossible (R-88 Part 2) |
| 3 | R-97a operator-only | the event "cannot be routed to a customer" | only configuration stopped it; fixed by a real operatorOnlyEvents register |
| 4 | classifyRunStates I1 |
"StateStopped means deliberately stopped by the user" | quiesce stops stacks the same way — a failed restart was silent (F-CRIT-1) |
| 5 | inflight.go |
"a caller that cannot acquire DEFERS" | the backup caller recorded a failure and paged the operator (F-A1) |
| 6 | quiesce.go |
the agent's 409 prevents "a spurious failure" | on the start path it produced one (F-A1) |
| 7 | recovery_unit.go B2 refusal (R-181) |
"the previous unit is untouched and NOTHING was deleted" | nothing deleted held; untouched was measured false — the floor was checked ONLY in captureAllRecoveryUnits, while the two dump legs wrote the bulk into the same tree first and unguarded, so a 182,272 B tar became 2,147,666,432 B under a manifest that had not moved |
| 8 | ResolveManagedFloor (R-216) |
"never push a controller past the agent it depends on" | it compared the box's agent against the golden's MinAgent while serving a floor that could point elsewhere. Raise a floor above the vouched golden — which the day-0 runbook recommends and a per-customer override makes trivial — and the guard checks a version it is not serving. Measured live 2026-08-05: golden 0.192.0/MinAgent 0.113.0, floor 0.200.0, agent 0.120.0 → served, and the box landed on a controller needing 0.125.0. Its customer was then told their correct recovery code was wrong. The first entry in this table where the false invariant was a GUARD, not a comment alone. Fixed hub v0.97.0: a floor above the vouched golden is HELD, with its own reason |
Three of these (4, 5/6 and 7) were found on live hardware, not by review or unit tests — #4 had a
green, red-proofed test suite over a production path that was broken two independent ways, and #7
survived a full green suite plus three of its own red-proofs, because every one of them asserted the
mechanism inside captureAllRecoveryUnits and none asserted the consequence across the whole
backup run. The test that would have caught it is the one #7's fix ships: fingerprint the tree before
and after, and compare. So:
- If a comment states an invariant, name the test that pins it, or write one.
- If an invariant has a stated dependency ("if either invariant changes, revisit this"), that is not a safeguard — nobody revisits. Pin it with a test that fails when the dependency moves.
- Prefer a test that asserts the consequence (does the alarm fire?) over one that asserts the mechanism (does suppression expire?). R-97b's Scenario F proved the mechanism and the consequence was still broken.