c30430c530
gates / gates (push) Failing after 15s
The four existing skills cover the product; nothing covered how work is reported. Two rules this project has paid for — check the artifact rather than the report, and do not state a claim more firmly than the evidence allows — lived only in the operator's head and in chat, where Claude Code never read them. - felhom-evidence five confidence tiers, artifact-over-report - felhom-diagnosis no hypothesis until a command has been seen red - felhom-plain-language ASD-STE100, two options, the re-pitch - felhom-handoff the note goes to a FILE, not the conversation - felhom-doc-authoring the pointer decides whether material is reached scripts/check_skills.py asserts what decides whether a skill is EVER reached: frontmatter parses, name == directory, description and body non-empty, under 150 lines, installed copy still samefile()s into the repo. install_skills.py globs and never reads the file, so a missing description installs perfectly and then silently never loads. It convicted on its first run: felhom-build-deploy is 179 lines. NOT trimmed here (pre-existing skills are out of scope, and trimming a deploy skill without exercising its commands is how a wrong command reaches a live host) — a named single-entry GRANDFATHERED exception, WARNed every run, R-394. A new skill over the limit is convicted. Red-proof run and seen failing: description removed from felhom-evidence -> exit 1, "frontmatter field 'description' is missing or empty". Restored, tree clean. skills/SOURCES.md records both MIT upstreams, that these are adaptations not copies, and the six pieces deliberately EXCLUDED with reasons. Register: R-392 (no architecture doc covers the two-AI workflow), R-393 (decision-log skill deferred, with the reason), R-394.
91 lines
4.8 KiB
Markdown
91 lines
4.8 KiB
Markdown
---
|
|
name: felhom-diagnosis
|
|
description: The order of operations for a hard Felhom bug — build a failing feedback loop BEFORE forming any theory. Triggers - a reported defect; something broken, throwing, failing, hanging or slow; a regression; a fault that only appears after a restart; and the phrases "diagnose", "debug", "why is this happening", "it stopped working". Contains the loop-first gate, the ways to build a loop on Felhom's surfaces, minimisation, plural hypotheses, and the stale-state rule.
|
|
---
|
|
|
|
# Diagnosis
|
|
|
|
The whole job of this skill is to stop a theory arriving before a repro.
|
|
|
|
## 1. Build the feedback loop first — this is the skill
|
|
|
|
Everything after this step is mechanical. With a tight command that goes red on **this** bug, the
|
|
cause will be found. Without one, reading code will not save you; it will produce a confident story
|
|
that fits the code you happened to read.
|
|
|
|
Ways to build one, roughly in order of preference on Felhom's surfaces:
|
|
|
|
- **A failing Go test** at whatever seam reaches the fault. `REUSE.md` §4 in each repo lists the
|
|
seams and the existing fakes.
|
|
- **A `curl` against the running endpoint.** Endpoint-level is this project's standard method —
|
|
there is no browser on DooPlex. Invoke the exact endpoint the UI invokes; the residual is
|
|
client-side rendering only.
|
|
- **A CLI invocation against a fixture**, diffing the output against a known-good file.
|
|
- **Replaying a captured payload or log line** through the code path in isolation.
|
|
- **A throwaway harness** that exercises the path with one call and nothing else around it.
|
|
- **A differential run** — two versions, or two configs, one changed thing between them.
|
|
- **A bisection harness**, when the fault appeared between two known-good states.
|
|
|
|
## 2. Tighten it
|
|
|
|
Faster, sharper, more deterministic. A flaky thirty-second loop is barely a loop; a two-second
|
|
deterministic one is a superpower, because it can be run fifty times while you think.
|
|
|
|
For an intermittent fault the goal is **not** a clean repro — it is a **higher reproduction rate**.
|
|
Loop the trigger, add stress, narrow the timing window until the fault is frequent enough to be
|
|
debuggable. A fault that fires 1 in 50 is a research project; the same fault at 1 in 3 is an
|
|
afternoon.
|
|
|
|
## 3. The gate
|
|
|
|
**You may not proceed to a hypothesis until you can name one command that you have already run at
|
|
least once, and show its invocation and its output.**
|
|
|
|
If you catch yourself reading code to build a theory before that command exists, stop. That is the
|
|
exact failure this skill prevents, and this project has lost whole sessions to it.
|
|
|
|
If you genuinely cannot build a loop, say so explicitly, list what you tried (per
|
|
**`felhom-evidence`** standing rule 2), and ask the operator for the one thing that would settle it:
|
|
a log dump, a payload capture, or access to the environment that reproduces it. Do not proceed to
|
|
hypothesise without one and do not present the result as a diagnosis.
|
|
|
|
## 4. Reproduce, then minimise
|
|
|
|
First confirm the loop produces the symptom **the operator described**, not a nearby one. A repro of
|
|
a different fault leads to a correct fix for the wrong bug, and it looks like success the whole way.
|
|
|
|
Then shrink. Cut one element at a time and re-run, keeping only what stays red. Done when removing
|
|
any remaining element turns it green — that residue is the fault's actual surface, and it is usually
|
|
much smaller than the first repro.
|
|
|
|
## 5. Hypothesise in the plural
|
|
|
|
Generate **three to five ranked candidates before testing any of them.** A single hypothesis anchors
|
|
on the first plausible idea, and every subsequent observation gets read as support for it.
|
|
|
|
Each pass, take the split that eliminates the most remaining space, and get **runtime evidence**
|
|
rather than reasoning. A print, a log line, a breakpoint value beats an argument about what the code
|
|
must do.
|
|
|
|
## 6. State before code, when a restart is involved
|
|
|
|
When something fails only after a restart, suspect **stale persistent state before code**. Code does
|
|
not change between two runs of the same binary. State does.
|
|
|
|
Check: config files, caches, lock files, serialised state, journals, registry files, anything on
|
|
disk the process reads at start. If clearing a state file restores the behaviour, the fix is **state
|
|
validation on read**, not a guard at the point that crashed.
|
|
|
|
## 7. Fix the cause
|
|
|
|
Resist adding a check that silences the crash. The crash is the messenger.
|
|
|
|
If a workaround needs a paragraph of comment to justify it, the code underneath is wrong and the
|
|
paragraph is the tell. And when the cause is found, **grep for the pattern, not only the
|
|
instance** — a bug that shipped once usually shipped in the places it was copied to.
|
|
|
|
## 8. Then hand off
|
|
|
|
The failing repro becomes the regression test, and the **companion red-proof is mandatory** for
|
|
every correctness or security fix. That procedure lives in **`felhom-testing`** — read it there.
|