Files
felhom.eu/skills/felhom-diagnosis/SKILL.md
T
admin c30430c530
gates / gates (push) Failing after 15s
skills: five process-domain skills + check_skills.py
The four existing skills cover the product; nothing covered how work is
reported. Two rules this project has paid for — check the artifact rather
than the report, and do not state a claim more firmly than the evidence
allows — lived only in the operator's head and in chat, where Claude Code
never read them.

- felhom-evidence      five confidence tiers, artifact-over-report
- felhom-diagnosis     no hypothesis until a command has been seen red
- felhom-plain-language ASD-STE100, two options, the re-pitch
- felhom-handoff       the note goes to a FILE, not the conversation
- felhom-doc-authoring the pointer decides whether material is reached

scripts/check_skills.py asserts what decides whether a skill is EVER
reached: frontmatter parses, name == directory, description and body
non-empty, under 150 lines, installed copy still samefile()s into the
repo. install_skills.py globs and never reads the file, so a missing
description installs perfectly and then silently never loads.

It convicted on its first run: felhom-build-deploy is 179 lines. NOT
trimmed here (pre-existing skills are out of scope, and trimming a
deploy skill without exercising its commands is how a wrong command
reaches a live host) — a named single-entry GRANDFATHERED exception,
WARNed every run, R-394. A new skill over the limit is convicted.

Red-proof run and seen failing: description removed from
felhom-evidence -> exit 1, "frontmatter field 'description' is missing
or empty". Restored, tree clean.

skills/SOURCES.md records both MIT upstreams, that these are adaptations
not copies, and the six pieces deliberately EXCLUDED with reasons.

Register: R-392 (no architecture doc covers the two-AI workflow),
R-393 (decision-log skill deferred, with the reason), R-394.
2026-08-25 09:36:20 +02:00

91 lines
4.8 KiB
Markdown

---
name: felhom-diagnosis
description: The order of operations for a hard Felhom bug — build a failing feedback loop BEFORE forming any theory. Triggers - a reported defect; something broken, throwing, failing, hanging or slow; a regression; a fault that only appears after a restart; and the phrases "diagnose", "debug", "why is this happening", "it stopped working". Contains the loop-first gate, the ways to build a loop on Felhom's surfaces, minimisation, plural hypotheses, and the stale-state rule.
---
# Diagnosis
The whole job of this skill is to stop a theory arriving before a repro.
## 1. Build the feedback loop first — this is the skill
Everything after this step is mechanical. With a tight command that goes red on **this** bug, the
cause will be found. Without one, reading code will not save you; it will produce a confident story
that fits the code you happened to read.
Ways to build one, roughly in order of preference on Felhom's surfaces:
- **A failing Go test** at whatever seam reaches the fault. `REUSE.md` §4 in each repo lists the
seams and the existing fakes.
- **A `curl` against the running endpoint.** Endpoint-level is this project's standard method —
there is no browser on DooPlex. Invoke the exact endpoint the UI invokes; the residual is
client-side rendering only.
- **A CLI invocation against a fixture**, diffing the output against a known-good file.
- **Replaying a captured payload or log line** through the code path in isolation.
- **A throwaway harness** that exercises the path with one call and nothing else around it.
- **A differential run** — two versions, or two configs, one changed thing between them.
- **A bisection harness**, when the fault appeared between two known-good states.
## 2. Tighten it
Faster, sharper, more deterministic. A flaky thirty-second loop is barely a loop; a two-second
deterministic one is a superpower, because it can be run fifty times while you think.
For an intermittent fault the goal is **not** a clean repro — it is a **higher reproduction rate**.
Loop the trigger, add stress, narrow the timing window until the fault is frequent enough to be
debuggable. A fault that fires 1 in 50 is a research project; the same fault at 1 in 3 is an
afternoon.
## 3. The gate
**You may not proceed to a hypothesis until you can name one command that you have already run at
least once, and show its invocation and its output.**
If you catch yourself reading code to build a theory before that command exists, stop. That is the
exact failure this skill prevents, and this project has lost whole sessions to it.
If you genuinely cannot build a loop, say so explicitly, list what you tried (per
**`felhom-evidence`** standing rule 2), and ask the operator for the one thing that would settle it:
a log dump, a payload capture, or access to the environment that reproduces it. Do not proceed to
hypothesise without one and do not present the result as a diagnosis.
## 4. Reproduce, then minimise
First confirm the loop produces the symptom **the operator described**, not a nearby one. A repro of
a different fault leads to a correct fix for the wrong bug, and it looks like success the whole way.
Then shrink. Cut one element at a time and re-run, keeping only what stays red. Done when removing
any remaining element turns it green — that residue is the fault's actual surface, and it is usually
much smaller than the first repro.
## 5. Hypothesise in the plural
Generate **three to five ranked candidates before testing any of them.** A single hypothesis anchors
on the first plausible idea, and every subsequent observation gets read as support for it.
Each pass, take the split that eliminates the most remaining space, and get **runtime evidence**
rather than reasoning. A print, a log line, a breakpoint value beats an argument about what the code
must do.
## 6. State before code, when a restart is involved
When something fails only after a restart, suspect **stale persistent state before code**. Code does
not change between two runs of the same binary. State does.
Check: config files, caches, lock files, serialised state, journals, registry files, anything on
disk the process reads at start. If clearing a state file restores the behaviour, the fix is **state
validation on read**, not a guard at the point that crashed.
## 7. Fix the cause
Resist adding a check that silences the crash. The crash is the messenger.
If a workaround needs a paragraph of comment to justify it, the code underneath is wrong and the
paragraph is the tell. And when the cause is found, **grep for the pattern, not only the
instance** — a bug that shipped once usually shipped in the places it was copied to.
## 8. Then hand off
The failing repro becomes the regression test, and the **companion red-proof is mandatory** for
every correctness or security fix. That procedure lives in **`felhom-testing`** — read it there.