The four existing skills cover the product; nothing covered how work is reported. Two rules this project has paid for — check the artifact rather than the report, and do not state a claim more firmly than the evidence allows — lived only in the operator's head and in chat, where Claude Code never read them. - felhom-evidence five confidence tiers, artifact-over-report - felhom-diagnosis no hypothesis until a command has been seen red - felhom-plain-language ASD-STE100, two options, the re-pitch - felhom-handoff the note goes to a FILE, not the conversation - felhom-doc-authoring the pointer decides whether material is reached scripts/check_skills.py asserts what decides whether a skill is EVER reached: frontmatter parses, name == directory, description and body non-empty, under 150 lines, installed copy still samefile()s into the repo. install_skills.py globs and never reads the file, so a missing description installs perfectly and then silently never loads. It convicted on its first run: felhom-build-deploy is 179 lines. NOT trimmed here (pre-existing skills are out of scope, and trimming a deploy skill without exercising its commands is how a wrong command reaches a live host) — a named single-entry GRANDFATHERED exception, WARNed every run, R-394. A new skill over the limit is convicted. Red-proof run and seen failing: description removed from felhom-evidence -> exit 1, "frontmatter field 'description' is missing or empty". Restored, tree clean. skills/SOURCES.md records both MIT upstreams, that these are adaptations not copies, and the six pieces deliberately EXCLUDED with reasons. Register: R-392 (no architecture doc covers the two-AI workflow), R-393 (decision-log skill deferred, with the reason), R-394.
4.8 KiB
name, description
| name | description |
|---|---|
| felhom-evidence | How to grade a claim about the Felhom system and how to verify work before reporting it done. Triggers - stating whether a behaviour is intentional or broken; reporting a result; validating another session's, another agent's or a REPORT.md's work; summarising a survey or an audit; and the questions "are we sure", "is X the case", "did that work". Contains the five confidence tiers, the words that carry confidence, the words to drop, and the artifact-over-report rule. |
Evidence and confidence
Applies to every claim about this system, in chat, in REPORT.md, in a register row, in a commit
message. Two of this project's five standing rules (documentation/runbooks/workspace-CLAUDE.md)
are the seed of this skill: rule 2 (a "no access" claim must list what was tried) and rule 3 (an
absent log line is not evidence).
1. Confidence tiers
Every claim sits in exactly one tier. The tier fixes the phrasing.
| Tier | Means | Phrasing |
|---|---|---|
| Direct | someone wrote it down, and the text answers the question | confident, present tense, citation immediately adjacent |
| Supported | several independent pieces converge; no single one states it | confident but visibly derived — name each piece |
| Inferred | a reasonable reading of context; nothing states it | hedged; make the inference chain explicit |
| Speculative | plausible, but other explanations fit equally well | explicitly marked a guess |
| Unknown | you looked and did not find it | name what you searched, and what you searched for |
Unknown is a result, not a failure. It is reported as one — with its search list, per standing
rule 2. "No access" or "not found" without a list of attempts is unfalsifiable.
2. Words that carry confidence
because · the reason is · was designed to · fixes · the decision was
Writing one of these asserts Direct or Supported. So a citation sits beside it — a
file:line, a commit, a register row, a measured observable — or the word changes.
3. Words to drop
clearly · obviously · of course · just · simply
Each one either restates a citation you already have, in which case it is noise, or hides that you do not have one, in which case it is a false claim wearing a confident coat.
4. Do not rationalise
Three named traps.
- Backwards justification. Assuming the author did the right thing, then reasoning back to a reason it must be right. The reason arrives before the evidence and is fitted to the conclusion.
- Repetition read as intent. A pattern that appears five times may be one decision copied four times. Frequency is not endorsement.
- Absence read as absence. Standing rule 3 covers the Felhom instance: an absent log line is equally consistent with "healthy" and "stopped entirely". The general form: prove a negative with a positive control. Plant the thing, find it, remove it, fail to find it. Only then does the silence mean something. Until then the silence may be the search being broken.
5. The embedded-hypothesis trap
A question often carries its own answer: "why is this slow, I assume it is the disk?"
The guess is a prompt to investigate, never a conclusion to confirm. Check it against the evidence like any other candidate, rank it with the others, and say plainly when it does not hold. Agreeing with an embedded guess costs nothing at the time and costs a whole session when it is wrong.
6. Verify the artifact, not the report
When checking work you did not do yourself — a prior session, a subagent, a handed-over
REPORT.md — read the thing itself:
- the pushed source at
file:line, on the remote, not the working tree; - the actual file contents, not a summary of them;
- the running state, not what a status line says about it.
An agent reports what it intended. That is not always what happened, and the gap is invisible in the report by construction.
The strongest proof is a script that re-runs the comparison. A one-time look proves one moment;
a checked-in script is an artifact someone else can re-run, and it is what turns "I looked" into
something a second person can confirm. This project's own instance: REPORT.md is never trusted on
its own — every claim it makes about code is confirmed against live Gitea before it is acted on.
7. Presence is not success
A timestamp that records an attempt is not evidence of a result. Ask of any timestamp: what exactly must have happened for this to be set? If the answer is "we tried", it cannot answer "did it work". Where a status field travels beside a timestamp, the verdict consults both.
DO NOT
Do not restate the companion red-proof procedure here. It belongs to felhom-testing, which
owns it together with the non-hollow rule and the nine shipped-false-invariant cases. Point at that
skill by name.