--- name: felhom-diagnosis description: The order of operations for a hard Felhom bug — build a failing feedback loop BEFORE forming any theory. Triggers - a reported defect; something broken, throwing, failing, hanging or slow; a regression; a fault that only appears after a restart; and the phrases "diagnose", "debug", "why is this happening", "it stopped working". Contains the loop-first gate, the ways to build a loop on Felhom's surfaces, minimisation, plural hypotheses, and the stale-state rule. --- # Diagnosis The whole job of this skill is to stop a theory arriving before a repro. ## 1. Build the feedback loop first — this is the skill Everything after this step is mechanical. With a tight command that goes red on **this** bug, the cause will be found. Without one, reading code will not save you; it will produce a confident story that fits the code you happened to read. Ways to build one, roughly in order of preference on Felhom's surfaces: - **A failing Go test** at whatever seam reaches the fault. `REUSE.md` §4 in each repo lists the seams and the existing fakes. - **A `curl` against the running endpoint.** Endpoint-level is this project's standard method — there is no browser on DooPlex. Invoke the exact endpoint the UI invokes; the residual is client-side rendering only. - **A CLI invocation against a fixture**, diffing the output against a known-good file. - **Replaying a captured payload or log line** through the code path in isolation. - **A throwaway harness** that exercises the path with one call and nothing else around it. - **A differential run** — two versions, or two configs, one changed thing between them. - **A bisection harness**, when the fault appeared between two known-good states. ## 2. Tighten it Faster, sharper, more deterministic. A flaky thirty-second loop is barely a loop; a two-second deterministic one is a superpower, because it can be run fifty times while you think. For an intermittent fault the goal is **not** a clean repro — it is a **higher reproduction rate**. Loop the trigger, add stress, narrow the timing window until the fault is frequent enough to be debuggable. A fault that fires 1 in 50 is a research project; the same fault at 1 in 3 is an afternoon. ## 3. The gate **You may not proceed to a hypothesis until you can name one command that you have already run at least once, and show its invocation and its output.** If you catch yourself reading code to build a theory before that command exists, stop. That is the exact failure this skill prevents, and this project has lost whole sessions to it. If you genuinely cannot build a loop, say so explicitly, list what you tried (per **`felhom-evidence`** standing rule 2), and ask the operator for the one thing that would settle it: a log dump, a payload capture, or access to the environment that reproduces it. Do not proceed to hypothesise without one and do not present the result as a diagnosis. ## 4. Reproduce, then minimise First confirm the loop produces the symptom **the operator described**, not a nearby one. A repro of a different fault leads to a correct fix for the wrong bug, and it looks like success the whole way. Then shrink. Cut one element at a time and re-run, keeping only what stays red. Done when removing any remaining element turns it green — that residue is the fault's actual surface, and it is usually much smaller than the first repro. ## 5. Hypothesise in the plural Generate **three to five ranked candidates before testing any of them.** A single hypothesis anchors on the first plausible idea, and every subsequent observation gets read as support for it. Each pass, take the split that eliminates the most remaining space, and get **runtime evidence** rather than reasoning. A print, a log line, a breakpoint value beats an argument about what the code must do. ## 6. State before code, when a restart is involved When something fails only after a restart, suspect **stale persistent state before code**. Code does not change between two runs of the same binary. State does. Check: config files, caches, lock files, serialised state, journals, registry files, anything on disk the process reads at start. If clearing a state file restores the behaviour, the fix is **state validation on read**, not a guard at the point that crashed. ## 7. Fix the cause Resist adding a check that silences the crash. The crash is the messenger. If a workaround needs a paragraph of comment to justify it, the code underneath is wrong and the paragraph is the tell. And when the cause is found, **grep for the pattern, not only the instance** — a bug that shipped once usually shipped in the places it was copied to. ## 8. Then hand off The failing repro becomes the regression test, and the **companion red-proof is mandatory** for every correctness or security fix. That procedure lives in **`felhom-testing`** — read it there.