Files
felhom.eu/documentation/PROMPT-TEMPLATE.md
T

35 KiB
Raw Blame History

Claude Code Prompt: [Feature / Fix Name]


For the operator — what this fixes, in one page (plain language)


0. Task class & scope

  • Implementation — code change with unit-test coverage; build + deploy + verify. (All sections.)
  • Spike — empirically validate an unproven mechanism BEFORE any production spec is written. Output is a findings doc in felhom.eu/documentation/audits/SPIKE-<topic>-<date>.md, not shipped production code. (§1, §2, §3, §4, §13-verify, §15. Skip §5-§12 unless the spike writes throwaway probes.)
  • Risky / supervised — touches agent / golden image / provisioning / destructive disk ops, or anything that mutates real customer data. Spec it; implement during the supervised session itself, directly on main. Mandatory STOP point in §13. (All sections + §13 STOP is load-bearing.)
  • Runbook-style validation — see note above; prefer a RUNBOOK-*.md.

Repos touched (check all — each gets its own baseline, green gate, build/deploy, and wrap-up):

  • felhom-controller (in-guest; Docker-only; no Proxmox creds) — app domain
  • felhom-agent (host; operator-tier; owns ALL Proxmox interaction)
  • felhom.eu (hub/, website/, documentation/)
  • app-catalog-felhom.eu (compose templates + .felhom.yml)

Don't confuse the two ex-"controllers": felhom-agent (host, was proxmox-controller) vs this felhom-controller (in-guest, was deploy-felhom-compose). Host/disk/Proxmox/Cloudflare logic lives in the agent; the controller keeps stack/deploy/UI/app-data-backup/metrics/git-sync/notifications.


1. Confirmed baselines (anti-drift anchor — MANDATORY)

Repo main @ commit Current version → Target version
felhom-controller <hash> v0.XX.0 v0.XX+1.0
felhom-agent <hash> v0.YY.0 v0.YY+1.0
felhom.eu (hub) <hash> v0.ZZ.0 v0.ZZ+1.0

An unpushed change does not exist for validation or deployment. CC commits directly to main (trunk-based — no feature/fix branches). If a change can't be cleanly shipped, revert + report; never park it on a branch.


2. Overview


3. Spike-first gate

  • Does this rely on any unvalidated mechanism? yes / no
  • If yes: is it already validated in a spike doc? Cite it: felhom.eu/documentation/audits/SPIKE-....
  • If yes and not validated: STOP — switch task class to Spike. Do not spec production code on assumptions.

4. Read these files first

  1. /mnt/5_hdd/felhom.eu/git/CLAUDE.md — workspace orientation (the felhom system, shared conventions, access). Versioned copy: felhom.eu/documentation/runbooks/workspace-CLAUDE.md.

  2. <repo>/CLAUDE.md — repo build/deploy, code-quality rules, trunk-based + live-validation rules.

  3. <repo>/CONTEXT.md — current project state / decisions / roadmap.

  4. THE ARCHITECTURE DOCUMENT FOR THE AREA THIS TASK TOUCHES — NAME IT AND SAY WHAT IT SAYS. Not "read the architecture folder": name the file, and state in one line what it rules about this area. A prompt that cannot name one says so explicitly, and that absence is itself recorded — an undocumented architectural decision is how a deliberate design gets "fixed" by someone who did not know it was one.

    area file
    what the platform does today, per scenario 00-capability-map.md
    topology, trust, app-data placement (hot vs bulk), backup scoping 01-topology-and-trust.md
    controller packages: KEEP/PORT/DELETE/MODIFY 02-controller-module-map.md
    the host agent 03-host-agent.md
    control-plane authorization 04-control-plane-authorization.md
    the hub 05-hub-architecture.md
    off-site connectivity 06-offsite-connectivity.md
    tiers, capture sets, restore paths, recovery model 07-backup-architecture.md

    Three sources, in this order, before any claim: the architecture folder holds the REASONING, the register holds the WORK, source holds the TRUTH. Skipping the first is how a decision gets reported as a bug.

    And the test that catches it: is what I am about to call a defect something we chose? If it was chosen and the choice is wrong, that is a proposal to change a decision — it goes to the operator as a decision, not filed as a bug. Cost of learning this (R-370): between 19 and 22 August a documented placement decision was called a defect in four places, because the register and live source were read and documentation/architecture/ was not.

    02-controller-module-map.md remains required reading before touching backup/, storage/, system/, config/.

  5. felhom.eu/documentation/controller/<feature>.md — code-verified feature docs (authoritative; match code, not summaries, if they drift).

  6. controller/README.md (or felhom-agent/README.md) — module map, feature reference, REST API.

  7. [exact source file(s) to modify] — [why it's relevant].

  8. [exact source file(s) to depend on] — [why it's relevant].


5. Existing functions / helpers to reuse (MUST use — do NOT reinvent)

Symbol File (landmark) Signature / use Modify?
Manager.StopStack internal/stacks/manager.go (~L682) StopStack(name string) error — idempotent No
Manager.RedeployFromEnv internal/stacks/deploy.go (~L478) writes app.yaml then compose up -d + post-start status No
appbackup.AppDataDir internal/.../paths.go namespace+app → data dir No
writeDiskJSON internal/.../agent_disk_handlers.go JSON envelope for disk endpoints No
[helper that needs a change] path/file.go Func(args) Yes — add intent param; gate on IntentEnrolled

Do NOT reuse: [name the dangerous-lookalike + why], e.g. rsyncMirror (has --delete).


6. Pattern references (copy structure from these canonical files)

Pattern Canonical file Key traits to copy
REST handler internal/api/router.go + a sibling handler stdlib net/http, JSON envelope, error logging
Local-API disk handler internal/localapi/disks.go (handleDiskEject) withGuest, scopedFromBody, role gate, fail-safe-to-protected
Scheduler job internal/.../scheduler register + skip-if-running, graceful shutdown
Crash-safe journal internal/stacks/backup.go journal-before-mutate, atomic tmp+fsync+rename, Recover on startup
Unit test internal/.../*_test.go (a recent non-hollow one) t.TempDir(), fakes/seams, effect assertions
HTML template internal/web/templates/*.html Hungarian UI, Go template funcs, no inline secrets

7. Integration scenarios (must pass)

Scenario A — [happy path]

Given: [concrete state]
When:  [exact endpoint / CC action]
Then:
  - [effect 1 — exact value/state]
  - [effect 2]
  - [the WRONG outcome this must NOT produce]

Scenario B — [alternative / role / mode]

Given: [what differs from A]
When:  [action]
Then:  [how outcomes differ and WHY]

Scenario C — [error / boundary / security gate]

Given: [bad input / unauthorised actor / data-bearing device]
When:  [action that must be refused]
Then:  [exact refusal — HTTP status, error, and the proven non-effect, e.g. "mkfs NOT called"]

8. Edge cases & deduplication rules

Condition A Condition B Result

9. Strict rules

  1. Green gate (Go): after every commit/phase, in each touched repo: go build ./... && go vet ./... && go test ./... — all green before proceeding. The build IS the typecheck; do not accumulate compile errors.
  2. Minimal changes: build only what's listed. No "while I'm here" refactors. Note anything worth fixing under "Observations" (§15) — don't act on it, but do file it: §15.9's marker rule means "not acted on" never means "not recorded".
  3. No silent failures: never swallow a parse/exec error — log it. Check a subprocess's own exit code; never pipe in a way that hides a 127. (The silent .felhom.yml quoting bug + the spike's exit-swallow lesson.)
  4. Secrets safety: log env var keys, never values. Never write secrets (tokens, passwords, keys) into CHANGELOG.md, REPORT.md, or any committed file — reference as "stored out-of-band".
  5. Line numbers are landmarks: reference code by symbol/landmark; reconfirm the line in the real source before editing (bundled files ≠ source files).
  6. Trunk-based: commit directly to main. No branches. Can't ship cleanly → revert + report.
  7. Version discipline: never :latest; bump the version on every task; one CHANGELOG entry per shippable change; overwrite REPORT.md.
  8. Coding standards: follow <repo>/CLAUDE.md (cross-cutting) and the feature doc / README for domain patterns.

10. Non-hollow test discipline (MANDATORY for any correctness/security fix)

  • Assert the effect happened, not just absence of error. Security example: prove mkfs was NOT called on a data-bearing device by inspecting the device/fake state — not by trusting the caller.
  • Companion red-proof: for every correctness fix, write the test so it FAILS on the pre-fix / trivial implementation. Demonstrate this: run the mutation (or guard-removed shape), confirm the test fails, revert. State the result in §15.
  • FS work uses t.TempDir(); orchestration uses fakes/seams (introduce a small stackOps / copier interface so tests don't shell out to docker/rsync).
  • Generation/idempotency: assert the negative (e.g. "fetch count does NOT increment on an unchanged heartbeat"; "a re-run finds its own prior (N) and adds no (1)(1)").
  • Seam discipline — every seam added gets ONE test through the PRODUCTION wiring path. An injected-seam test proves the component, never the caller. Three shipped defects in three days make this non-negotiable: controller v0.154.0 (the wizard read a flag the handler never sets), agent v0.91.0 (main.go never called SetAuthSink, so the whole auth-honesty leg was inert), and agent v0.92.0 (the watchdog had no sudoers grant for three of its four probes). Every one of them was fully green. Where the caller is func main() and cannot be invoked from a test, walk its AST for the call — and note that a strings.Contains on the source is NOT sufficient: a commented-out call still contains the string, which is how the controller's first version of that test passed its own red-proof (2026-07-21).

Part 1: [Foundation / engine]

1.1 [Task]

File(s): path/file.go Problem: [what's wrong / what to build, and why] DO: [explicit positive instructions; tricky code snippets only — CC fills the obvious parts] DO NOT: [the specific anti-pattern for THIS task] Crash-safety (if stateful): journal-before-mutate; atomic tmp+fsync+rename; Recover on startup; guaranteed cleanup via defer for graceful exits only — a defer does NOT run on SIGKILL, so a crash-safe cleanup needs an on-disk marker plus a startup Recover() (Campaign 8 fault 10 proved this on live hardware); single-flight mutex; ListLXC-style ground truth in recovery.

Green gate: go build ./... && go vet ./... && go test ./internal/<pkg>/


Part 2: [Core logic / external surface]

2.1 [Task]

Phase order (risky surfaces): Phase A structural → validate → Phase B security/external surface.

Green gate: full go test ./...


Part 3: [Web / UI — if touched]

3.1 [Task]

Hungarian strings: "[exact copy]" — e.g. "Adatok másolása: (X% — Y/Z GB)…".

Green gate: go build ./... + go test ./...


Part N: Documentation & versioning

N.1 CHANGELOG.md (each touched repo)

Cumulative, newest entry on top. New version number. Cite the files/commits for each item.

N.2 REPORT.md (each touched repo) — OVERWRITE

Most-recent implementation only (not cumulative). MUST contain: confirmed baselines, per-commit hashes, per-test results, deployed versions, and an explicit "NOT yet live-validated — awaiting supervised [Bx]" list. No secrets.

N.3 CONTEXT.md

Decisions made, architectural state, what's next.

N.4 README / architecture docs

<repo>/README.md if architecture/features changed; authoritative architecture/feature docs in felhom.eu/documentation/{architecture,controller}/. Spikes/audits → felhom.eu/documentation/audits/.

N.5 The coupling rule — FOUR artifacts, same session (end-of-session checklist)

If the task changed what the platform can do (new/removed capability, a capability moved status, or what is open changed), update all four in the SAME session:

  • architecture/00-capability-map.md — add/adjust the affected scenario row with the correct status (PROVEN-LIVE requires an audits/|tests/ citation; else IMPLEMENTED/PARTIAL) and an evidence citation. A leg not exercised live → PARTIAL/IMPLEMENTED with a note saying which leg and a → R-n pointer, never PROVEN-LIVE.

  • backlog/ROADMAP.md — add any new open work / deferred legs as an item (next free R-n), collapse a now-shipped item to its one-liner + version, and honor the coupling rule (every roadmap item names the map row it flips; every map gap row points back by ID). Parked sub-features (e.g. an alternative transport, an agent-plane leg) go as a note under the owning arc's item.

  • the owning architecture/*.md — ruled as S-1 in CONTEXT.md (2026-07-26, R-81): any task that changes an architectural contract (tiers, targets, cadences, trust boundaries) updates the owning design doc in the same session. It was ruled but never written here, in the template CC actually reads — so it bound nobody. Now it does.

  • backlog/OPEN-ITEMS.md — the register, and the single source of truth for open work. A new row is filed into its category's section with one Category and one Sev (P1–P4) — the scale and the eleven categories head the file, and register_shape_gate.py refuses a row without them (2026-10-03). A task that changes what is open without touching it re-creates exactly the thread-loss the register was built to solve: REPORT.md is overwritten every session, so nothing durable may live only there.

    AN ENUMERATED GAP BECOMES A ROW, IN THE SAME SESSION. PROSE IS NOT A RECORD.

    This binds surveys, inventories, spikes, reviews and diagnoses, not only implementation sessions — those are the documents that enumerate gaps, and they are the ones that have lost them. If a document says a thing is missing, unhandled, unreachable or "not currently filed", it does not leave the session as prose. It leaves as a row here, with a rank and an owner. Writing "not filed" is not a disposition; it is a note that the work was seen and dropped.

    A row in ROADMAP.md alone does not satisfy this. Both files hold open work and only this one calls itself the source of truth, so a finding recorded solely there is invisible to every standing rule that says "grep the register before minting" (R-369).

    The cost, recorded so the rule can be narrowed later rather than becoming permanent by accident: R-107 — "no offsite action unpacks the named-volume tars Tier-3 captures" — was enumerated on 2026-07-28, given a number, written into ROADMAP.md and 07-backup-architecture.md, and never entered here. It was rediscovered from scratch 25 days later by an overnight drill that planted files and watched them not come back, and shipped as R-354. The work was right the first time; only the filing was missing.

Report which OPEN-ITEMS.md rows the task opened, closed (and moved to CLOSED-ITEMS.md), narrowed or re-ranked (§15). Every row carries an owner — a row nobody owns is how items got lost in the first place.

N.6 Website version bump (if controller/hub version is shown on the site).

N.7 Housekeeping — before the report, not after (2026-08-22 ruling)

Nothing was ever pruned and the numbers got bad: the register reached 672 KB across 286 entries, over half of it finished work, one entry at 16 KB. A file that cannot be read is a file that cannot be checked, and this project has paid for that twice — a record nobody could find because it sat inside an entry about something else, and a finding rediscovered because nobody could see it.

  1. Move what this session closed — IN THE SAME COMMIT THAT CLOSES IT. A closed row keeps its title, the version it shipped in, its evidence paths, and any sentence stating a rule. Everything else goes, and it moves to backlog/CLOSED-ITEMS.md. Nothing is deleted: the compressed entry names the commit whose git show returns the full original text. Open rows are not touched — their detail is doing a job. This is a gate, not a habit (2026-10-03): scripts/closed_register_gate.py RULE 3 refuses a push whose OPEN-ITEMS.md holds a row LEADING with a finished word (CLOSED, SHIPPED, FIXED, DECIDED, …). The rule existed from 2026-08-22 without a gate and was followed only sometimes: on 2026-10-03 the open register held 113 finished rows, a quarter of the file.
  2. Rehome live reasoning before compressing it away. If a closed entry carries the reason a rule exists or a fence sits where it does, that reasoning moves — to CONTEXT.md if it is a decision, to the owning architecture/*.md if it is a shape. Where it is a decision, mark the resulting shape [DESIGN] in the architecture document and point it at the log entry (§4's map). Losing a reason is how a deliberate design becomes a bug in someone's eyes — that cost four mis-filed defect reports in August 2026 (R-370, R-376).
  3. State the register's size in the report, before and after. A number every session is what makes growth visible; prose about tidiness is not a mechanism.

11. Tests

Test Group A — [Scenario A]

- setup (t.TempDir / fakes)
- action (the exact call)
- assert (effects, exact values/counts)
- companion: [mutation that makes it fail on pre-fix code]

Test Group B — [Scenario B]

Test Group C — [Scenario C — error/security gate]


12. What NOT to do

  • Do NOT [specific out-of-scope thing CC might build].
  • Do NOT [the forbidden ACT, and its reason: e.g. "modify the operator-signed DecommissionExecutor — it is the signature boundary"]. Fence the act, not the object. A bare "do not touch X" is read as covering every act on X, including ones nobody meant to restrict — that is how "do not re-target demo-hp's backup target" became "do not use demo-hp at all", which pushed a destructive drill onto the one machine holding the recovery chain. State the reason too: a rule whose reason is recorded can be correctly narrowed later, and one without it becomes permanent by default.
  • Do NOT reuse [the dangerous lookalike from §5].
  • Do NOT refactor nearby code, change passing tests, or create a branch.
  • Do NOT add "Co-Authored by..." anywhere

13. Build / deploy / STOP

Clean-tree gate first: git status --porcelain empty AND HEAD == origin/main in the repo being built. An unpushed change does not exist.

If the task needs a machine to break — NAME IT. A drill, destructive test or throwaway VM gets an explicit target: "run this on demo-hp". Do not leave it to be inferred from the prohibition list; listing only what is off-limits leaves the most valuable unfenced machine as the residual choice, which is exactly how a drill landed on DooPlex. Tiers and per-machine permitted/care/forbidden: runbooks/target-selection.md.

Controller (local build on DooPlex → guest via golden/bootstrap):

FELHOM_ROOT=/mnt/5_hdd/felhom.eu
# build + push
cd $FELHOM_ROOT/build/felhom-controller && git -C $FELHOM_ROOT/git/felhom-controller pull && ./build.sh <VER> --push
# deploy to guest (golden/bootstrap): docker pull → write /etc/felhom-controller-image → restart bootstrap svc
# verify:
ssh felhom-pve "pct exec 9201 -- docker ps --filter name=felhom-controller --format '{{.Image}} {{.Status}}'"

Agent (build locally, scp /tmp/felhom-agent-<VER> felhom-pve:/tmp/ — one hop; deploy to felhom-pve: backup prior binary, replace, restart service; verify clean restart — e.g. ReassertGuestBinds does not rebind a non-enrolled drive).

Hub (felhom.eu/hub/): build + push locally, then bump manifests/hub.yaml's image: tag in git and do a deliberate ArgoCD sync — never kubectl set image (auto-sync is OFF and any imperative change is reverted on the next sync). Verify Synced/Healthy + rollout + image + logs.

Live validation rule: if validating a user-facing feature, exercise the real server pipeline (connect → enroll → deploy), or invoke the exact endpoint the UI invokes (acceptable proxy — no server logic skipped). Do NOT hand-set state via raw agent attach (the F9 bypass). Browser automation is NOT available on DooPlex — endpoint-level is the standard method; say which was used.

STOP point (risky/supervised tasks): build + unit-test + build/deploy images, but do NOT run a live destructive op or migrate real customer data — that is the supervised [Bx] step. Throwaway marker dirs under a scratch /mnt path are OK only if they touch no enrolled customer data.

Teardown. A run that provisions anything owns its removal. Three layers, and the third is the one that gets missed:

  1. The machine — the VM or guest and its volumes, deleted.
  2. The host — pvesm status before and after, and the space actually returned.
  3. The hub — the customer or appliance record the run created. State its disposition explicitly: deleted; or retained as a fixture with the reason; or blocked by the ONLINE→DOWN delete gate with the command recorded for later. Silence is how drill-r50 became both a blocked customer and the only drift fixture.

Precedent: three drills, three orphaned customers — drill-r50, sess-c, sess-d. sess-c was not recorded by its own report, so the record said the teardown was clean when it was not. Layers 1 and 2 are the ones that get remembered because they are visible on the box; layer 3 is invisible from there and has been missed every time.

Which machine to provision on in the first place: documentation/runbooks/target-selection.md.


14. Implementation order

Phase 1: [Part 1]            → go build ./... && go vet ./... && go test ./internal/<pkg>/
Phase 2: [Part 2]            → full go test ./...
Phase 3: [Part 3 — UI]       → go build ./... && go test ./...
Phase N-1: [Tests]           → full go test ./... (all green)
Phase N:   [Build/deploy/verify + docs + push to main]

15. Final verification & deliverables

Run, per touched repo: go build ./... && go vet ./... && go test ./...

Report MUST include:

  1. Confirmed baselines used (repo @ hash, version).

  2. Files created / modified (paths).

  3. Per-commit hashes pushed to main.

  4. Per-test results (name → pass/fail) + the §10 companion red-proof outcomes.

  5. Test count before / after; all green (or list failures).

  6. Deployed versions + docker ps / agent-service / hub-pod verification output.

  7. NOT yet live-validated — awaiting supervised [Bx]: explicit list (the real put-data → operate → integrity flow).

  8. Evidence copied off BEFORE each revert — for every phase that ran on a machine, the logs were pulled to the evidence directory at the end of that phase, before any revert, snapshot restore or teardown, including the intermediate ones. The intermediate revert is the one that gets forgotten: two sessions lost a phase's logs to a mid-run revert to virgin on 2026-08-12 and 2026-08-13 — same machine, same point, three days apart (R-320). If a phase's evidence is already gone, the report says so plainly and the finding is REPRODUCED independently — that is the expectation, not an improvisation. A quotation read live and no longer re-readable is named as such.

  9. Teardown evidence — all three §13 layers, if the run provisioned anything: the machine deleted, pvesm status before/after with the space returned, and the hub-side record's disposition named (deleted / retained-with-reason / gate-blocked-with-the-command). A run that provisioned nothing says so. "Teardown clean" without layer 3 is not a report — it is the sess-c failure.

  10. Observations: out-of-scope items noticed. Every item carries FILED: R-NNN naming the register row opened for it in THIS session, or NOT-A-FINDING: <reason> declaring plainly that it does not warrant one. Opening the row is the default; declaring is the exception and its reason is the whole of the marker.

    This wording replaces "documented, NOT acted on" (2026-08-24, R-389), and the old wording was the defect. "Documented" was satisfied by a paragraph — and REPORT.md is overwritten every session, so a paragraph has a lifetime of one session. On 2026-08-23 a live, reproducible finding (only the first broken app per hour reaches the operator) was written under Observations and nowhere else; it had no register row and had to be re-derived the next day. That is the same shape as R-341, and this project's own standard says a rule without a mechanism is a wish.

    The mechanism is gate 11 (scripts/observations_gate.py, registered in the repo runners), which refuses a push whose REPORT.md carries an observation with neither marker. Note what it deliberately does NOT accept: a passing mention of some other R-NNN. The lost item cited R-182 as an analogy, so "cites a register row" would have passed the very item the gate exists to catch.