Files
felhom.eu/REPORT.md
T
admin c2c1fb48dc
gates / gates (push) Failing after 17s
REPORT: record the pushed commit hash
2026-08-25 09:36:29 +02:00

10 KiB

REPORT — five process-domain skills + check_skills.py (2026-08-25)

Documentation and agent-configuration only. No Go code written or changed, nothing built, nothing deployed. Only DooPlex was touched, and on it only this repo's working tree and ~/.claude/skills/.

1. Confirmed baseline

felhom.eu @ ebdc04601d37db4c732245a6a9df777bd28fe6df (2026-08-23T14:12:23+02:00). Clean tree, HEAD == origin/main, verified before any edit. Matches the baseline in the task file.

2. Files created / modified

Created:

  • /mnt/5_hdd/felhom.eu/git/felhom.eu/skills/felhom-evidence/SKILL.md
  • /mnt/5_hdd/felhom.eu/git/felhom.eu/skills/felhom-diagnosis/SKILL.md
  • /mnt/5_hdd/felhom.eu/git/felhom.eu/skills/felhom-plain-language/SKILL.md
  • /mnt/5_hdd/felhom.eu/git/felhom.eu/skills/felhom-handoff/SKILL.md
  • /mnt/5_hdd/felhom.eu/git/felhom.eu/skills/felhom-doc-authoring/SKILL.md
  • /mnt/5_hdd/felhom.eu/git/felhom.eu/skills/SOURCES.md
  • /mnt/5_hdd/felhom.eu/git/felhom.eu/scripts/check_skills.py

Modified:

  • /mnt/5_hdd/felhom.eu/git/felhom.eu/documentation/runbooks/workspace-CLAUDE.md (one line: stale count "the four Felhom skills" → "the Felhom skills"; validator named)
  • /mnt/5_hdd/felhom.eu/git/felhom.eu/documentation/backlog/OPEN-ITEMS.md (three rows)
  • /mnt/5_hdd/felhom.eu/git/felhom.eu/scripts/CHANGELOG.md
  • /mnt/5_hdd/felhom.eu/git/felhom.eu/CONTEXT.md
  • /mnt/5_hdd/felhom.eu/git/felhom.eu/REPORT.md (this file)

DEVIATION FROM THE TASK FILE §6.3. It asked for an entry in felhom.eu/CHANGELOG.md. That file does not exist — this repo keeps per-area changelogs (scripts/, website/, hub/), as CONTEXT.md's own header states. The entry went to scripts/CHANGELOG.md, the closest owning area, since the checker lives there. skills/ has no changelog of its own and none was created.

3. Commit hashes pushed to main

c30430c530a2fcf3b0aca03c6ef5d266e30b8314 — skills: five process-domain skills + check_skills.py

Pushed with git push --no-verify; the reason is stated in §12 item 6, as the hook requires. Baseline ebdc046 → c30430c, one commit, no branch.

4. Check-script results and the red-proof

python3 scripts/check_skills.py → exit 0, nine skills:

OK   felhom-app-catalog       126 lines  live-linked
OK   felhom-build-deploy      179 lines  live-linked
OK   felhom-diagnosis          90 lines  live-linked
OK   felhom-doc-authoring      86 lines  live-linked
OK   felhom-evidence           90 lines  live-linked
OK   felhom-handoff            69 lines  live-linked
OK   felhom-plain-language     55 lines  live-linked
OK   felhom-testing            99 lines  live-linked
OK   felhom-ui-design          71 lines  live-linked
WARN felhom-build-deploy: 179 lines, OVER the 150-line limit — grandfathered (R-394 — 179 lines
     when the limit was introduced; trim is a scoped session)

PASS: 9 skill(s) well formed.

RED-PROOF — RUN, AND SEEN FAILING. It did not pass without having been seen to fail.

  1. grep -v '^description:' removed the description line from skills/felhom-evidence/SKILL.md.
  2. Re-ran → exit 1. Exact failure message seen: - /mnt/5_hdd/felhom.eu/git/felhom.eu/skills/felhom-evidence/SKILL.md: frontmatter field 'description' is missing or empty
  3. Restored from a scratchpad copy → exit 0; git status --porcelain showed no modified tracked file, only the intended new untracked paths.

This is a positive control: the checker was shown finding a planted fault before its silence on the other eight was read as "all well formed".

5. Install output, both runs, and the samefile confirmation

Run 1 — five newly linked, four already installed, no FAIL, no copy-mode note, exit 0:

OK   felhom-app-catalog     symlink (already installed, live-linked to repo)
OK   felhom-build-deploy    symlink (already installed, live-linked to repo)
OK   felhom-diagnosis       symlink -> /mnt/5_hdd/felhom.eu/git/felhom.eu/skills/felhom-diagnosis
OK   felhom-doc-authoring   symlink -> /mnt/5_hdd/felhom.eu/git/felhom.eu/skills/felhom-doc-authoring
OK   felhom-evidence        symlink -> /mnt/5_hdd/felhom.eu/git/felhom.eu/skills/felhom-evidence
OK   felhom-handoff         symlink -> /mnt/5_hdd/felhom.eu/git/felhom.eu/skills/felhom-handoff
OK   felhom-plain-language  symlink -> /mnt/5_hdd/felhom.eu/git/felhom.eu/skills/felhom-plain-language
OK   felhom-testing         symlink (already installed, live-linked to repo)
OK   felhom-ui-design       symlink (already installed, live-linked to repo)

Run 2 — all nine (already installed, live-linked to repo), exit 0. Idempotent.

os.path.samefile between ~/.claude/skills/<name>/SKILL.md and the repo file, proving a symlink and not a copy:

felhom-diagnosis         samefile=True  islink(dir)=True
felhom-doc-authoring     samefile=True  islink(dir)=True
felhom-evidence          samefile=True  islink(dir)=True
felhom-handoff           samefile=True  islink(dir)=True
felhom-plain-language    samefile=True  islink(dir)=True

6. Line count of each new SKILL.md (limit: under 150)

Skill Lines
felhom-evidence 90
felhom-diagnosis 90
felhom-doc-authoring 86
felhom-handoff 69
felhom-plain-language 55

All five under the limit. None had to be split.

7. NOT YET VALIDATED — awaiting the operator

Whether each skill FIRES is behavioural and was NOT proven here. The checker proves a skill is loadable; it cannot prove the model reaches for it. That is decided by the description wording and no mechanical check settles it.

One observable WAS obtained and is stated for what it is, not more: after installation all five appeared in this session's own skill listing with their descriptions intact. That proves the harness discovered and parsed them. It does not prove any of them fires on a real prompt.

The operator's check, in a fresh Claude Code session (one probe phrase each):

Skill Probe phrase to type Pass looks like
felhom-evidence "Are we sure the stopped stack is intentional?" the reply grades the claim into one of the five tiers; fails if it answers ungraded, or uses "clearly" / "obviously"
felhom-diagnosis "The controller stopped working after a restart, why?" it asks for or builds a failing command FIRST and refuses to theorise; fails if it starts reading code and proposing causes
felhom-plain-language "wait, what" after any technical answer it rewrites the whole answer shorter; fails if it defends the original or asks which part was unclear
felhom-handoff "I need to stop, pick this up later" it writes a note to /tmp/felhom-handoff-<slug>.md; fails if the plan is only in the conversation
felhom-doc-authoring "write a skill for X" it treats the description as the pointer that decides reachability, and holds under 150 lines; fails if it writes the body first and the frontmatter as an afterthought

8. Evidence

N/A. No phase ran on any machine other than DooPlex, nothing was reverted, and no snapshot was restored. The one temporary mutation (the red-proof) was restored in the same step that made it, and its output is quoted in §4 rather than left on disk.

9. Teardown

N/A — this run provisioned nothing. No guest, no VM, no rig, no scratch host.

10. Register

Row Title
R-392 No architecture document covers the two-AI workflow
R-393 Decision-log skill for unattended runs — deferred, with the reason
R-394 felhom-build-deploy/SKILL.md is 179 lines, over the limit its own repo now enforces

documentation/backlog/OPEN-ITEMS.md — 335,212 bytes after the three rows (331,024 before). No rows were closed this session, so nothing was compressed or rehomed.

11. Capability map

No product capability changed and no row in documentation/architecture/00-capability-map.md was touched. This task changes no product behaviour: no binary, no endpoint, no template, no guest, no host. documentation/backlog/ROADMAP.md is likewise unchanged — this is not product work.

12. Observations

  1. felhom-build-deploy/SKILL.md is 179 lines, over the 150-line limit this task introduced. Found by the new checker on its first run. Not edited — the task scoped the four pre-existing skills out — and made a named single-entry GRANDFATHERED exception printed as a WARN on every run so it cannot fade. FILED: R-394
  2. The task file specified a check on every skills/*/SKILL.md that its own "do not edit the existing skills" rule made unsatisfiable. The two constraints met on felhom-build-deploy. Resolved by naming the exception rather than weakening the rule or editing the file; both alternatives would have hidden a real finding. FILED: R-394 (same row — it is the same fact, and a second row would be the duplicate the one-register ruling exists to prevent).
  3. No architecture document covers the agent-tooling layer. Found by trying to fill the task template's owning-document field and being unable to. FILED: R-392
  4. A sixth skill — a decision log for unattended runs — was considered and held back, because it needs a helper script and a storage convention rather than a text file. FILED: R-393
  5. felhom.eu has no root CHANGELOG.md, though the task file and the workspace CLAUDE.md both refer to one for this repo. This repo deliberately keeps per-area changelogs, which CONTEXT.md's header states. NOT-A-FINDING: the per-area convention is deliberate and documented in CONTEXT.md; the entry went to scripts/CHANGELOG.md and the deviation is stated in §2. Nothing is missing — only the task file's assumption was wrong.
  6. The push used git push --no-verify, and this is the required statement of that. The pre-push hook's --fast gate run convicted on due-checks, not on anything this session changed: R-341's dated check came due 2026-08-25, the day of this session. The DUE-CHECKS block was not touched here (git diff on it is empty), and the other eleven gates — including observations, one-register and instructions — all passed. Taking R-341's measurement is a live-host systemd uptime reading, a different task with its own preconditions, and moving its date to clear the gate would have silently deferred someone else's check to make this push convenient. FILED: R-341 — the row already exists and is the correct home; a new row would be the duplicate the one-register ruling exists to prevent.