skills: five process-domain skills + check_skills.py
gates / gates (push) Failing after 15s

The four existing skills cover the product; nothing covered how work is
reported. Two rules this project has paid for — check the artifact rather
than the report, and do not state a claim more firmly than the evidence
allows — lived only in the operator's head and in chat, where Claude Code
never read them.

- felhom-evidence      five confidence tiers, artifact-over-report
- felhom-diagnosis     no hypothesis until a command has been seen red
- felhom-plain-language ASD-STE100, two options, the re-pitch
- felhom-handoff       the note goes to a FILE, not the conversation
- felhom-doc-authoring the pointer decides whether material is reached

scripts/check_skills.py asserts what decides whether a skill is EVER
reached: frontmatter parses, name == directory, description and body
non-empty, under 150 lines, installed copy still samefile()s into the
repo. install_skills.py globs and never reads the file, so a missing
description installs perfectly and then silently never loads.

It convicted on its first run: felhom-build-deploy is 179 lines. NOT
trimmed here (pre-existing skills are out of scope, and trimming a
deploy skill without exercising its commands is how a wrong command
reaches a live host) — a named single-entry GRANDFATHERED exception,
WARNed every run, R-394. A new skill over the limit is convicted.

Red-proof run and seen failing: description removed from
felhom-evidence -> exit 1, "frontmatter field 'description' is missing
or empty". Restored, tree clean.

skills/SOURCES.md records both MIT upstreams, that these are adaptations
not copies, and the six pieces deliberately EXCLUDED with reasons.

Register: R-392 (no architecture doc covers the two-AI workflow),
R-393 (decision-log skill deferred, with the reason), R-394.
This commit is contained in:
2026-08-25 09:36:20 +02:00
parent ebdc04601d
commit c30430c530
12 changed files with 826 additions and 243 deletions
+3
View File
@@ -140,6 +140,9 @@ the fault was real. Full observables: `tests/campaign11-evidence-2026-08-05/jour
| **R-387** | **The hub REWRITES an unknown severity and says nothing, and the guard built to catch that sits downstream of the rewrite.** One handler, two fields, opposite discipline: an unknown `event_type` is rejected with a loud `400`, while an unknown `severity` was silently coerced to `info` — after which `severityNotifies` drops it and NEITHER delivery leg runs. **Two shipped features went out that way**: `DiskAlertKind.Severity` emitted `"warn"` until controller v0.215.0, `app_start_failed` until v0.223.0. **Measured on the live hub DB 2026-08-23: 91 `app_start_failed` events stored all-time and ZERO `notification_log` rows before that day** — not one, on any channel, while every POST returned 200. **The dispatcher's `unrecognized severity` line could never execute** for an API event, because the coercion one line upstream guarantees the value it looks for cannot arrive. | **CLOSED — hub v0.107.0, 2026-08-23** | — | **The coercion STAYS; only the silence is fixed** — a rejected event is a LOST event, and losing an alarm is worse than mis-routing one. A `WARN` now names the customer, the event type, the rejected value and the consequence. **The dispatcher branch was KEPT, on evidence not caution:** `cmd/hub/main.go` wires `dispatcher.ProcessEvent` DIRECTLY as the `monitor.EventNotifyFunc` for the staleness, host-staleness and offsite-box checkers, which never pass through the handler — for them it is the only severity guard there is; deleting it as "dead" would have removed the live half while the dead half supplied the justification. All 90 severity literals in `internal/monitor` verified already valid. Proven live: `[WARN] [api] Event from demo-hp: severity "warn" is not in {info,warning,error,critical}…`, with an `error` control silent. Evidence: `audits/DRILL-r329-r386-2026-08-23/evidence/live-19-scenarioH-after.txt`. | CC |
| **R-391** | **Gate 11 (observations) is registered in three of the four runners; `app-catalog-felhom.eu` is the exception.** The controller and agent runners already carried a shared-gate mechanism (`SHARED_REUSE`, `SHARED_INSTRUCTIONS` pointing into `felhom.eu/scripts/`), so registering there was one constant and one `GATES` line each. **`catalog_gates.py` has no such mechanism:** its `run_gate` joins every entry against its OWN `scripts/` directory, so it cannot invoke a sibling repo's script at all; and its loop appends `--all` to every gate unconditionally, which the observations gate would read as a path. Registering there therefore needs `run_gate`'s contract widened AND the argument handling changed — a refactor of a runner whose shape is deliberately different (per-app scoping, network/runtime gates excluded from `--fast`), in a repo this task marked out of scope. **The exposure today is nil** — `app-catalog-felhom.eu/REPORT.md` has no observations section, and the gate passes quietly on that — but a future catalog session could write one and nothing would read it. **Filed rather than left as a sentence in a report, which is the exact failure R-389 records.** | **OPEN — LOW** | — | Either give `catalog_gates.py` the `SHARED_*` absolute-path mechanism the other two runners already have and stop appending `--all` to gates that do not take it, or state in that repo's CLAUDE.md that its REPORT.md carries no observations section by convention. **Do not copy the gate script** — the shared checker lives in ONE place (`felhom.eu/scripts/`) and copying it is the drift the shared pattern exists to prevent. | CC |
| **R-390** | **The golden-bake runbook omits `pveam update`, and the failure it produces names the wrong cause.** `documentation/runbooks/RUNBOOK-manual-build.md` §4.1 step 2 says to list the current Debian template because "the exact point release rots" — but on the drill VM's `virgin` snapshot **the `pveam` INDEX is stale too**, so `pveam available` offers an old point release and `pveam download local <that>` fails with **`400 Parameter verification failed. template: no such template`**. That reads as a typo or a bad argument, not as an old index, and it costs a diagnosis every time. **Hit on two consecutive bakes** (golden 0.222.0 and 0.223.0, both 2026-08-23). The runbook is otherwise correct verbatim — the qemu launch line, the token-read-inside-the-VM pattern and the acceptance markers all worked unchanged. | **OPEN — LOW** | — | Add `pveam update` as its own numbered step before the listing, and say WHY: a snapshot that never changes carries an index that never updates, so the rot warning already in the step applies to the index as well as to the release. Recorded meanwhile in the workspace memory `golden-bake-needs-pveam-update` and in `documentation/tests/golden-0.223.0-2026-08-23/README.md`. | CC |
| **R-392** | **No architecture document covers the two-AI workflow.** `documentation/architecture/` holds eight documents and **all eight cover the product** — topology, host agent, control-plane authorization, hub, off-site connectivity, backup, controller modules, capability map. Nothing records how the Claude.ai / Claude Code split works, what each side owns, how skills and `.claude/rules/` are scoped, or why. **The absence was found by trying to fill the template field, not by a survey:** the task that added the five process skills (2026-08-25) had to name an owning architecture document and could not, and the template requires that be recorded rather than passed over. The exposure today is low — the split is stable and both sides work — but it lives entirely in the operator's head and in chat, which is precisely the shape of a commitment nothing enforces. | **OPEN — LOW** | — | Write one architecture document for the agent-tooling layer: which AI owns which artifact class (`TASK-*.md`, `RUNBOOK-*.md`, validation), how skills are scoped and installed, what belongs in a `CLAUDE.md` versus a skill versus a rules file, and the reasoning for each boundary. **The rules themselves already exist** in `skills/felhom-doc-authoring/SKILL.md`; what is missing is the map of who owns what. Do not restate the doc-authoring rules there — point at that skill. | CC |
| **R-393** | **A decision-log skill for unattended runs was considered and deliberately deferred.** Filed 2026-08-25 by the session that added the five process skills, so the deferral is a decision on the record rather than a thing that was dropped. **The gap it would close:** an overnight or unattended run makes dozens of decisions and the operator can only reconstruct them by reading the whole transcript, which is exactly what nobody does. The proposal is an appended row per decision — what was chosen, why, the evidence pointer, and the result — so a long run is reconstructable in a page. **Why it was NOT built with the other five:** the other five are text files that need nothing but the existing installer glob. This one needs a helper script to append rows and a storage convention for where the log lives and when it is rotated, which makes it an implementation task with its own acceptance criteria, not a skill file. | **OPEN — LOW** | — | Decide the storage convention FIRST — most likely a per-session file beside the session's evidence directory, never `REPORT.md`, which is overwritten every session (the R-341 shape). Then the skill, then the helper. **Check it does not duplicate `felhom-handoff`**, which already owns the end-of-session note; a decision log is the during-the-run half and the two must point at each other rather than overlap. | CC |
| **R-394** | **`felhom-build-deploy/SKILL.md` is 179 lines, over the 150-line limit its own repo now enforces.** Found 2026-08-25 by `scripts/check_skills.py` on its first run — the over-length was discovered BY the new checker, on the day the limit was written down, which is the checker working as intended. **It is not edited and not trimmed here:** the task that introduced the limit explicitly scoped the four pre-existing skills out, and trimming a build-and-deploy skill without exercising its commands is how a wrong command ships to a live host. **It is a named single-entry exception in `GRANDFATHERED` in `scripts/check_skills.py`, printed as a WARN on every run**, so it cannot fade; a NEW skill over the limit is convicted normally, and growing the set requires editing that file in a commit with a row to name. **The rationale for the limit** — attention thins across the excess, so the lines that matter are not the ones that survive — is in `skills/felhom-doc-authoring/SKILL.md` §5. | **OPEN — LOW** | — | Trim `felhom-build-deploy/SKILL.md` under 150 lines in a session that can VERIFY the commands it keeps, then delete its `GRANDFATHERED` entry in the same commit. The likely trim is the per-artifact command blocks moving behind a pointer to the runbooks, keeping the gotchas inline — but that is a judgement for a session with a build to run, not a line-count exercise. | CC |
| **R-388** | **PRODUCT DECISION (not a defect): the customer notification model is the wrong shape, and the settings page grows by one toggle per detector.** The operator's framing, recorded verbatim 2026-08-23: *"A customer should be notified only about things they can act on or are responsible for — the drive they unplugged, the storage they filled. **A failed backup is our incident, not theirs.** The intended shape is that we detect it, we tell them we noticed and are dealing with it, and they are not handed an error they cannot solve. The subscription should feel like being looked after, not like being on call."* Today's page is the opposite shape — one switch per detector, and it **grew from 12 to 15 in a single session** (one new alarm plus two compound toggles split into four). That growth is the argument, not an aside: a page that grows per detector keeps asking a household to make engineering decisions. | **OPEN — DIRECTION, operator's call** | a decision on scope; nothing here is a bug | Recorded as a dated **[DESIGN — DIRECTION]** entry at `documentation/architecture/08-alarm-ladder.md` §8, marked plainly as *not current behaviour*. **Deliberately NOT implemented in the session that recorded it.** `app_start_failed` defaulting OFF is consistent with the direction and reversible either way, but was ruled on its own merits and does not pre-judge the redesign. | Viktor |
| **R-229** | **The instruction-file rightsizing landed for `felhom-controller` and the workspace root; three pieces were deliberately deferred.** Done 2026-08-06: controller split into a 92-effective-line core plus four `paths:`-scoped `.claude/rules/*.md`; workspace root 208→142 effective lines with its versioned copy kept byte-identical; surgical corrections to `felhom-agent` and `felhom.eu` (expired TEMPORARY block, every version literal, the Legacy-Windows copies, the duplicated health-check rule); five contradictions resolved — including a drill-VM claim **measured live** (`qm list` on demo-hp shows VM 300 `drill-r50`; `felhom-agent` was right, `felhom-controller` was wrong); new shared `felhom.eu/scripts/instructions_gate.py` registered in `controller_gates.py` and `agent_gates.py`, 20 fixture tests + red-proof. **Leg (a) CLOSED 2026-08-06 (part 2):** `felhom.eu/CLAUDE.md` **227 → 115 effective lines**, split into a core plus `.claude/rules/{hub,website,manifests,docs}.md`; `instructions_gate` **registered in `scripts/repo_gates.py`** (six gates, all OK) in the required order — trim first, register second, because a registered-but-failing gate refuses every push. Scoping proven from the `InstructionsLoaded` hook log in two fresh sessions, not from frontmatter. **Still deferred:** (b) **CLOSED 2026-08-06 (close-out)** — `felhom-agent/CLAUDE.md` **175 → 99 effective lines** (measured 175, not 173: the CI correction added two), split into a core plus `.claude/rules/{proxmox,localapi,backup,storage}.md` beside the existing `health-checks.md`. The release section now points at the `felhom-build-deploy` skill instead of restating a table that drifts from the script. **Every `CLAUDE.md` in the workspace is now ≤120 effective lines except the workspace root at 142, which is deliberate — it is the only file re-injected after `/compact`.** (c) **CLOSED 2026-08-06 (part 2)** — all 44 orphans resolved with **zero deletions** (file count 158 before and after): 4 durable `reference`-type files indexed, 40 dated episode records moved to `.claude-memory/archive/`. `MEMORY.md` 145 → **150 lines / 17,977 bytes**, and `instructions_gate` check 6 now watches it (over-limit FAILS, orphan WARNS, absent store PASSES *printing its reason*). (d) **The spec-as-failing-test pilot** — moved to R-230. Full accounting: `audits/LEDGER-instruction-trim-2026-08-06.md` + `audits/LEDGER-instruction-trim-part2-2026-08-06.md` | **READY** — owner Viktor |
| **R-230** | **Three instruction/memory follow-ups deliberately left by the part-2 session (2026-08-06), each needing a decision rather than an implementation.** (a) **A ruling is owed on auto-written staleness.** The hand-written `CLAUDE.md` files are now clean of version literals and expired blocks — the gate enforces it — but `MEMORY.md`, which Claude writes and which is the LARGER half of what loads (8.4k tokens vs the root file's 6.6k), carries **21 lines with component version literals**, **5 with bare host addresses**, and an entry still reading *"demo boxes REMOTE till ~08-02"* — the same expired-TEMPORARY class the gate was built to kill, now surviving in the one file the gate's content rules do not cover. **Partly actioned 2026-08-06 (close-out), and the ruling is STILL OWED:** the **three statements that were actively false** were corrected — `R-193 decision open` (closed 2026-08-05), `demo boxes REMOTE till ~08-02` (the box answers on the home LAN), `OPEN R-25b` (shipped 2026-07-21) — and gate check 6 now **WARNs** on version literals, host addresses, expired statements and stale-open citations in the index. WARN, never FAIL: Claude writes that file between sessions, so a hard failure would refuse a human's push over a line no human typed, and the warning is read by the model that will next edit it. **The remaining 32 version literals and 4 host addresses were deliberately left** for that loop. What is still owed is the bulk-correction ruling. **Correcting the premise:** the earlier report's "three expired statements" were all FALSE POSITIVES — each matched an ISO date inside a markdown link target, i.e. a filename — while the one real expired claim carried no ISO date at all. (b) **CLOSED 2026-08-06 (close-out)** — the workspace-root `CLAUDE.md` **is now a relative symlink** to the versioned copy, so the divergence class is gone rather than policed. Check 5 learned two shapes: for a link it asserts the target resolves to a real file (**a dangling link is worse than a diverged copy — the instructions load NOTHING and there is no content left to notice is wrong**), for two files byte-identity as before, so a clone elsewhere is unaffected. **Proven, not assumed:** three fresh sessions logged `session_start` for the link path, and a fourth **with no tools at all** quoted standing rule 1 verbatim — the content reaches the model, not just the path. (c) **The spec-as-failing-test pilot**, approved in principle and not started (was R-229(d)). | **READY** — owner Viktor |
+3 -2
View File
@@ -14,9 +14,10 @@ host is one SSH hop. Run CC inside tmux so sessions survive SSH drops: **`tmux n
- **In-guest controller** — one per customer LXC, Docker-only: `felhom-controller/`.
- Also: `app-catalog-felhom.eu/` (app templates), `homelab-manifests/` (DooPlex k3s).
Each repo's own `CLAUDE.md` and `.claude/rules/` load when you touch files there. The four Felhom
Each repo's own `CLAUDE.md` and `.claude/rules/` load when you touch files there. The Felhom
skills are installed from `felhom.eu/skills/` with `python3 felhom.eu/scripts/install_skills.py`
(symlink — repo edits are live immediately).
(symlink — repo edits are live immediately), and validated with
`python3 felhom.eu/scripts/check_skills.py`.
## This host is production infrastructure