Files
admin c30430c530
gates / gates (push) Failing after 15s
skills: five process-domain skills + check_skills.py
The four existing skills cover the product; nothing covered how work is
reported. Two rules this project has paid for — check the artifact rather
than the report, and do not state a claim more firmly than the evidence
allows — lived only in the operator's head and in chat, where Claude Code
never read them.

- felhom-evidence      five confidence tiers, artifact-over-report
- felhom-diagnosis     no hypothesis until a command has been seen red
- felhom-plain-language ASD-STE100, two options, the re-pitch
- felhom-handoff       the note goes to a FILE, not the conversation
- felhom-doc-authoring the pointer decides whether material is reached

scripts/check_skills.py asserts what decides whether a skill is EVER
reached: frontmatter parses, name == directory, description and body
non-empty, under 150 lines, installed copy still samefile()s into the
repo. install_skills.py globs and never reads the file, so a missing
description installs perfectly and then silently never loads.

It convicted on its first run: felhom-build-deploy is 179 lines. NOT
trimmed here (pre-existing skills are out of scope, and trimming a
deploy skill without exercising its commands is how a wrong command
reaches a live host) — a named single-entry GRANDFATHERED exception,
WARNed every run, R-394. A new skill over the limit is convicted.

Red-proof run and seen failing: description removed from
felhom-evidence -> exit 1, "frontmatter field 'description' is missing
or empty". Restored, tree clean.

skills/SOURCES.md records both MIT upstreams, that these are adaptations
not copies, and the six pieces deliberately EXCLUDED with reasons.

Register: R-392 (no architecture doc covers the two-AI workflow),
R-393 (decision-log skill deferred, with the reason), R-394.
2026-08-25 09:36:20 +02:00

4.8 KiB

name, description
name description
felhom-evidence How to grade a claim about the Felhom system and how to verify work before reporting it done. Triggers - stating whether a behaviour is intentional or broken; reporting a result; validating another session's, another agent's or a REPORT.md's work; summarising a survey or an audit; and the questions "are we sure", "is X the case", "did that work". Contains the five confidence tiers, the words that carry confidence, the words to drop, and the artifact-over-report rule.

Evidence and confidence

Applies to every claim about this system, in chat, in REPORT.md, in a register row, in a commit message. Two of this project's five standing rules (documentation/runbooks/workspace-CLAUDE.md) are the seed of this skill: rule 2 (a "no access" claim must list what was tried) and rule 3 (an absent log line is not evidence).

1. Confidence tiers

Every claim sits in exactly one tier. The tier fixes the phrasing.

Tier Means Phrasing
Direct someone wrote it down, and the text answers the question confident, present tense, citation immediately adjacent
Supported several independent pieces converge; no single one states it confident but visibly derived — name each piece
Inferred a reasonable reading of context; nothing states it hedged; make the inference chain explicit
Speculative plausible, but other explanations fit equally well explicitly marked a guess
Unknown you looked and did not find it name what you searched, and what you searched for

Unknown is a result, not a failure. It is reported as one — with its search list, per standing rule 2. "No access" or "not found" without a list of attempts is unfalsifiable.

2. Words that carry confidence

because · the reason is · was designed to · fixes · the decision was

Writing one of these asserts Direct or Supported. So a citation sits beside it — a file:line, a commit, a register row, a measured observable — or the word changes.

3. Words to drop

clearly · obviously · of course · just · simply

Each one either restates a citation you already have, in which case it is noise, or hides that you do not have one, in which case it is a false claim wearing a confident coat.

4. Do not rationalise

Three named traps.

  • Backwards justification. Assuming the author did the right thing, then reasoning back to a reason it must be right. The reason arrives before the evidence and is fitted to the conclusion.
  • Repetition read as intent. A pattern that appears five times may be one decision copied four times. Frequency is not endorsement.
  • Absence read as absence. Standing rule 3 covers the Felhom instance: an absent log line is equally consistent with "healthy" and "stopped entirely". The general form: prove a negative with a positive control. Plant the thing, find it, remove it, fail to find it. Only then does the silence mean something. Until then the silence may be the search being broken.

5. The embedded-hypothesis trap

A question often carries its own answer: "why is this slow, I assume it is the disk?"

The guess is a prompt to investigate, never a conclusion to confirm. Check it against the evidence like any other candidate, rank it with the others, and say plainly when it does not hold. Agreeing with an embedded guess costs nothing at the time and costs a whole session when it is wrong.

6. Verify the artifact, not the report

When checking work you did not do yourself — a prior session, a subagent, a handed-over REPORT.md — read the thing itself:

  • the pushed source at file:line, on the remote, not the working tree;
  • the actual file contents, not a summary of them;
  • the running state, not what a status line says about it.

An agent reports what it intended. That is not always what happened, and the gap is invisible in the report by construction.

The strongest proof is a script that re-runs the comparison. A one-time look proves one moment; a checked-in script is an artifact someone else can re-run, and it is what turns "I looked" into something a second person can confirm. This project's own instance: REPORT.md is never trusted on its own — every claim it makes about code is confirmed against live Gitea before it is acted on.

7. Presence is not success

A timestamp that records an attempt is not evidence of a result. Ask of any timestamp: what exactly must have happened for this to be set? If the answer is "we tried", it cannot answer "did it work". Where a status field travels beside a timestamp, the verdict consults both.

DO NOT

Do not restate the companion red-proof procedure here. It belongs to felhom-testing, which owns it together with the non-hollow rule and the nine shipped-false-invariant cases. Point at that skill by name.