The four existing skills cover the product; nothing covered how work is reported. Two rules this project has paid for — check the artifact rather than the report, and do not state a claim more firmly than the evidence allows — lived only in the operator's head and in chat, where Claude Code never read them. - felhom-evidence five confidence tiers, artifact-over-report - felhom-diagnosis no hypothesis until a command has been seen red - felhom-plain-language ASD-STE100, two options, the re-pitch - felhom-handoff the note goes to a FILE, not the conversation - felhom-doc-authoring the pointer decides whether material is reached scripts/check_skills.py asserts what decides whether a skill is EVER reached: frontmatter parses, name == directory, description and body non-empty, under 150 lines, installed copy still samefile()s into the repo. install_skills.py globs and never reads the file, so a missing description installs perfectly and then silently never loads. It convicted on its first run: felhom-build-deploy is 179 lines. NOT trimmed here (pre-existing skills are out of scope, and trimming a deploy skill without exercising its commands is how a wrong command reaches a live host) — a named single-entry GRANDFATHERED exception, WARNed every run, R-394. A new skill over the limit is convicted. Red-proof run and seen failing: description removed from felhom-evidence -> exit 1, "frontmatter field 'description' is missing or empty". Restored, tree clean. skills/SOURCES.md records both MIT upstreams, that these are adaptations not copies, and the six pieces deliberately EXCLUDED with reasons. Register: R-392 (no architecture doc covers the two-AI workflow), R-393 (decision-log skill deferred, with the reason), R-394.
4.8 KiB
name, description
| name | description |
|---|---|
| felhom-diagnosis | The order of operations for a hard Felhom bug — build a failing feedback loop BEFORE forming any theory. Triggers - a reported defect; something broken, throwing, failing, hanging or slow; a regression; a fault that only appears after a restart; and the phrases "diagnose", "debug", "why is this happening", "it stopped working". Contains the loop-first gate, the ways to build a loop on Felhom's surfaces, minimisation, plural hypotheses, and the stale-state rule. |
Diagnosis
The whole job of this skill is to stop a theory arriving before a repro.
1. Build the feedback loop first — this is the skill
Everything after this step is mechanical. With a tight command that goes red on this bug, the cause will be found. Without one, reading code will not save you; it will produce a confident story that fits the code you happened to read.
Ways to build one, roughly in order of preference on Felhom's surfaces:
- A failing Go test at whatever seam reaches the fault.
REUSE.md§4 in each repo lists the seams and the existing fakes. - A
curlagainst the running endpoint. Endpoint-level is this project's standard method — there is no browser on DooPlex. Invoke the exact endpoint the UI invokes; the residual is client-side rendering only. - A CLI invocation against a fixture, diffing the output against a known-good file.
- Replaying a captured payload or log line through the code path in isolation.
- A throwaway harness that exercises the path with one call and nothing else around it.
- A differential run — two versions, or two configs, one changed thing between them.
- A bisection harness, when the fault appeared between two known-good states.
2. Tighten it
Faster, sharper, more deterministic. A flaky thirty-second loop is barely a loop; a two-second deterministic one is a superpower, because it can be run fifty times while you think.
For an intermittent fault the goal is not a clean repro — it is a higher reproduction rate. Loop the trigger, add stress, narrow the timing window until the fault is frequent enough to be debuggable. A fault that fires 1 in 50 is a research project; the same fault at 1 in 3 is an afternoon.
3. The gate
You may not proceed to a hypothesis until you can name one command that you have already run at least once, and show its invocation and its output.
If you catch yourself reading code to build a theory before that command exists, stop. That is the exact failure this skill prevents, and this project has lost whole sessions to it.
If you genuinely cannot build a loop, say so explicitly, list what you tried (per
felhom-evidence standing rule 2), and ask the operator for the one thing that would settle it:
a log dump, a payload capture, or access to the environment that reproduces it. Do not proceed to
hypothesise without one and do not present the result as a diagnosis.
4. Reproduce, then minimise
First confirm the loop produces the symptom the operator described, not a nearby one. A repro of a different fault leads to a correct fix for the wrong bug, and it looks like success the whole way.
Then shrink. Cut one element at a time and re-run, keeping only what stays red. Done when removing any remaining element turns it green — that residue is the fault's actual surface, and it is usually much smaller than the first repro.
5. Hypothesise in the plural
Generate three to five ranked candidates before testing any of them. A single hypothesis anchors on the first plausible idea, and every subsequent observation gets read as support for it.
Each pass, take the split that eliminates the most remaining space, and get runtime evidence rather than reasoning. A print, a log line, a breakpoint value beats an argument about what the code must do.
6. State before code, when a restart is involved
When something fails only after a restart, suspect stale persistent state before code. Code does not change between two runs of the same binary. State does.
Check: config files, caches, lock files, serialised state, journals, registry files, anything on disk the process reads at start. If clearing a state file restores the behaviour, the fix is state validation on read, not a guard at the point that crashed.
7. Fix the cause
Resist adding a check that silences the crash. The crash is the messenger.
If a workaround needs a paragraph of comment to justify it, the code underneath is wrong and the paragraph is the tell. And when the cause is found, grep for the pattern, not only the instance — a bug that shipped once usually shipped in the places it was copied to.
8. Then hand off
The failing repro becomes the regression test, and the companion red-proof is mandatory for
every correctness or security fix. That procedure lives in felhom-testing — read it there.