R-389: key the operator cooldown per app for app_start_failed; gate 11 makes an unfiled observation refuse the push
gates / gates (push) Successful in 16s
gates / gates (push) Successful in 16s
The cooldown key was customerID:eventType plus the tier and run suffixes, and none of them names an app, so every app going down inside the same hour collapsed onto one key and only the first was mailed. Measured on demo-hp: bookstack sent 09:27:51, privatebin suppressed 09:31:51 under key=demo-hp:app_start_failed. cooldownStackSuffix is the third sibling of cooldownTierSuffix and cooldownRunSuffix, and separate for the reason the second one's docstring already gives: the existing two keep byte-identical semantics for every type that uses them. It is ALLOW-LISTED to app_start_failed and takes the event type as well as the details, unlike its siblings, and that asymmetry is the safety property. The backup family's cooldown is coarse ON PURPOSE (R-97a, R-182) so one full disk sends one digest rather than one mail per app - and crossdrive_failed is severity error, reaches the operator leg, and carries stack_name through a DIFFERENT struct, so a payload-shape rule would have split it silently. The hour itself does not change. Gate 11 refuses a push whose REPORT.md carries an observation with neither `FILED: R-NNN` nor `NOT-A-FINDING: <reason>`. It deliberately does NOT accept a passing mention of some other R-number: the lost item cited R-182 as an analogy, so "cites a register row" would have passed the very item the gate exists to catch. That discrepancy with the spec is recorded in the gate's docstring. Registered here and in the controller and agent runners. NOT in the catalog runner - it has no shared-gate mechanism and appends --all to every gate; filed as R-391 rather than left as a sentence, which is this session's lesson. PROMPT-TEMPLATE.md §15.9 corrected: "documented, NOT acted on" was the wording that invited the gap, and it now names the markers and points at the gate. R-390 filed for the golden-bake runbook's missing `pveam update`. Hub tests 709 -> 716.
This commit is contained in:
@@ -256,7 +256,8 @@ Then: [exact refusal — HTTP status, error, and the proven non-effect, e.g. "m
|
||||
`go build ./... && go vet ./... && go test ./...` — all green before proceeding. The build IS the
|
||||
typecheck; do not accumulate compile errors.
|
||||
2. **Minimal changes:** build only what's listed. No "while I'm here" refactors. Note anything worth
|
||||
fixing under "Observations" (§15) — don't act on it.
|
||||
fixing under "Observations" (§15) — don't act on it, but **do file it**: §15.9's marker rule means
|
||||
"not acted on" never means "not recorded".
|
||||
3. **No silent failures:** never swallow a parse/exec error — log it. Check a subprocess's **own** exit
|
||||
code; never pipe in a way that hides a 127. (The silent `.felhom.yml` quoting bug + the spike's
|
||||
exit-swallow lesson.)
|
||||
@@ -553,7 +554,22 @@ Report MUST include:
|
||||
`pvesm status` before/after with the space returned, and the **hub-side record's disposition named**
|
||||
(deleted / retained-with-reason / gate-blocked-with-the-command). A run that provisioned nothing says
|
||||
so. "Teardown clean" without layer 3 is not a report — it is the `sess-c` failure.
|
||||
9. **Observations:** out-of-scope items noticed — documented, NOT acted on.
|
||||
9. **Observations:** out-of-scope items noticed. **Every item carries `FILED: R-NNN` naming the
|
||||
register row opened for it in THIS session, or `NOT-A-FINDING: <reason>` declaring plainly that it
|
||||
does not warrant one.** Opening the row is the default; declaring is the exception and its reason
|
||||
is the whole of the marker.
|
||||
|
||||
**This wording replaces "documented, NOT acted on" (2026-08-24, R-389), and the old wording was
|
||||
the defect.** "Documented" was satisfied by a paragraph — and `REPORT.md` is overwritten every
|
||||
session, so a paragraph has a lifetime of one session. On 2026-08-23 a live, reproducible finding
|
||||
(only the first broken app per hour reaches the operator) was written under Observations and
|
||||
nowhere else; it had no register row and had to be re-derived the next day. That is the same shape
|
||||
as R-341, and this project's own standard says **a rule without a mechanism is a wish**.
|
||||
|
||||
**The mechanism is gate 11** (`scripts/observations_gate.py`, registered in the repo runners),
|
||||
which refuses a push whose `REPORT.md` carries an observation with neither marker. Note what it
|
||||
deliberately does NOT accept: a passing mention of some other `R-NNN`. The lost item cited `R-182`
|
||||
as an analogy, so "cites a register row" would have passed the very item the gate exists to catch.
|
||||
|
||||
---
|
||||
|
||||
|
||||
@@ -138,6 +138,8 @@ the fault was real. Full observables: `tests/campaign11-evidence-2026-08-05/jour
|
||||
|---|---|---|
|
||||
| **R-385** | **A controller was built, baked AND vouched with no CHANGELOG entry of its own, and every gate stayed green.** Controller **0.221.1** shipped on 2026-08-23 while the newest heading in `felhom-controller/CHANGELOG.md` still read `v0.221.0` — the prune-ordering fix (commit `810b18a`) had been written INSIDE the v0.221.0 entry instead of getting its own. The image was never in question; the RECORD was, and the fleet ran a version the record did not name. **`scripts/golden_currency_gate.py` could not catch it by construction:** it failed only on `released > baked`, so a golden AHEAD of the record passed silently. Measured on the real history: `newest released 0.221.0 / newest golden baked 0.221.1 → OK, exit 0`. | **CLOSED — 2026-08-23** | — | **Both halves fixed, both directions red-proofed.** The record: `v0.221.1` has its own heading carrying the MOVED (not duplicated, not deleted) reasoning — commit `da75603`, pushed alone before anything else. The gate now asks *"is the baked version WRITTEN DOWN?"* — the baked version must have its own `## vX.Y.Z` heading **anywhere** in the CHANGELOG. **Membership, not `baked > released`, deliberately:** a comparison against the newest heading alone goes green the moment any later entry is written, leaving the unrecorded version permanently unrecorded and the gate permanently silent about it. INCONCLUSIVE (exit 2) preserved. Evidence: `audits/DRILL-r384-dead-db-alarm-2026-08-23/evidence/gate-0*.txt` — old gate/old record `exit 0`, new gate/old record `exit 1`, new gate/fixed record `exit 0`. | CC |
|
||||
| **R-387** | **The hub REWRITES an unknown severity and says nothing, and the guard built to catch that sits downstream of the rewrite.** One handler, two fields, opposite discipline: an unknown `event_type` is rejected with a loud `400`, while an unknown `severity` was silently coerced to `info` — after which `severityNotifies` drops it and NEITHER delivery leg runs. **Two shipped features went out that way**: `DiskAlertKind.Severity` emitted `"warn"` until controller v0.215.0, `app_start_failed` until v0.223.0. **Measured on the live hub DB 2026-08-23: 91 `app_start_failed` events stored all-time and ZERO `notification_log` rows before that day** — not one, on any channel, while every POST returned 200. **The dispatcher's `unrecognized severity` line could never execute** for an API event, because the coercion one line upstream guarantees the value it looks for cannot arrive. | **CLOSED — hub v0.107.0, 2026-08-23** | — | **The coercion STAYS; only the silence is fixed** — a rejected event is a LOST event, and losing an alarm is worse than mis-routing one. A `WARN` now names the customer, the event type, the rejected value and the consequence. **The dispatcher branch was KEPT, on evidence not caution:** `cmd/hub/main.go` wires `dispatcher.ProcessEvent` DIRECTLY as the `monitor.EventNotifyFunc` for the staleness, host-staleness and offsite-box checkers, which never pass through the handler — for them it is the only severity guard there is; deleting it as "dead" would have removed the live half while the dead half supplied the justification. All 90 severity literals in `internal/monitor` verified already valid. Proven live: `[WARN] [api] Event from demo-hp: severity "warn" is not in {info,warning,error,critical}…`, with an `error` control silent. Evidence: `audits/DRILL-r329-r386-2026-08-23/evidence/live-19-scenarioH-after.txt`. | CC |
|
||||
| **R-391** | **Gate 11 (observations) is registered in three of the four runners; `app-catalog-felhom.eu` is the exception.** The controller and agent runners already carried a shared-gate mechanism (`SHARED_REUSE`, `SHARED_INSTRUCTIONS` pointing into `felhom.eu/scripts/`), so registering there was one constant and one `GATES` line each. **`catalog_gates.py` has no such mechanism:** its `run_gate` joins every entry against its OWN `scripts/` directory, so it cannot invoke a sibling repo's script at all; and its loop appends `--all` to every gate unconditionally, which the observations gate would read as a path. Registering there therefore needs `run_gate`'s contract widened AND the argument handling changed — a refactor of a runner whose shape is deliberately different (per-app scoping, network/runtime gates excluded from `--fast`), in a repo this task marked out of scope. **The exposure today is nil** — `app-catalog-felhom.eu/REPORT.md` has no observations section, and the gate passes quietly on that — but a future catalog session could write one and nothing would read it. **Filed rather than left as a sentence in a report, which is the exact failure R-389 records.** | **OPEN — LOW** | — | Either give `catalog_gates.py` the `SHARED_*` absolute-path mechanism the other two runners already have and stop appending `--all` to gates that do not take it, or state in that repo's CLAUDE.md that its REPORT.md carries no observations section by convention. **Do not copy the gate script** — the shared checker lives in ONE place (`felhom.eu/scripts/`) and copying it is the drift the shared pattern exists to prevent. | CC |
|
||||
| **R-390** | **The golden-bake runbook omits `pveam update`, and the failure it produces names the wrong cause.** `documentation/runbooks/RUNBOOK-manual-build.md` §4.1 step 2 says to list the current Debian template because "the exact point release rots" — but on the drill VM's `virgin` snapshot **the `pveam` INDEX is stale too**, so `pveam available` offers an old point release and `pveam download local <that>` fails with **`400 Parameter verification failed. template: no such template`**. That reads as a typo or a bad argument, not as an old index, and it costs a diagnosis every time. **Hit on two consecutive bakes** (golden 0.222.0 and 0.223.0, both 2026-08-23). The runbook is otherwise correct verbatim — the qemu launch line, the token-read-inside-the-VM pattern and the acceptance markers all worked unchanged. | **OPEN — LOW** | — | Add `pveam update` as its own numbered step before the listing, and say WHY: a snapshot that never changes carries an index that never updates, so the rot warning already in the step applies to the index as well as to the release. Recorded meanwhile in the workspace memory `golden-bake-needs-pveam-update` and in `documentation/tests/golden-0.223.0-2026-08-23/README.md`. | CC |
|
||||
| **R-389** | **Only the FIRST broken app per hour reaches the operator — the cooldown key names the event type, not the app.** `dispatcher.go:337` builds the operator key as `customerID + ":" + eventType + cooldownTierSuffix(details) + cooldownRunSuffix(details)`, and **neither suffix reads an app name**. So every app that goes down inside the same hour collapses onto one key and only the first is mailed. **Measured live on `demo-hp` 2026-08-23:** `bookstack` alarmed at 09:27:51 and was `sent`; `privatebin` alarmed at 09:31:51, four minutes later, and was logged `suppressed — operator cooldown 1h, key=demo-hp:app_start_failed`. Three apps down together tonight would produce one mail. **The app's identity is already on the wire** — `AppDetails{StackName, DisplayName}` serialises as `stack_name` (`felhom-controller/internal/notify/notifier.go:146-149`), and the hub already makes exactly this kind of distinction twice, with `cooldownTierSuffix` and `cooldownRunSuffix`. **This was latent for as long as the cooldown has existed and only became reachable when R-329 made `app_start_failed` deliverable at all** — the same "a known-broken thing moves from unreachable to load-bearing" shape as R-329 itself. **AND IT WAS NEVER FILED:** it was written in a REPORT.md observations paragraph on 2026-08-23 and nowhere else — the register had no row for it until now, which is R-341's shape one surface over and is why gate 11 exists. | **OPEN — MEDIUM** | — | A third sibling suffix, `cooldownStackSuffix`, **allow-listed to `app_start_failed` and nothing else**. **Do NOT apply it globally:** the backup family's cooldown is coarse ON PURPOSE (R-97a, R-182) so that one full disk sends one digest rather than one mail per app — and `crossdrive_failed` (severity `error`) carries `stack_name` through a *different* struct (`CrossDriveDetails`), so a global suffix would silently split it per-app. The hour itself does not change; the grain is the complaint, not the length. Evidence: `audits/DRILL-cooldown-grain-2026-08-23/`. | CC |
|
||||
| **R-388** | **PRODUCT DECISION (not a defect): the customer notification model is the wrong shape, and the settings page grows by one toggle per detector.** The operator's framing, recorded verbatim 2026-08-23: *"A customer should be notified only about things they can act on or are responsible for — the drive they unplugged, the storage they filled. **A failed backup is our incident, not theirs.** The intended shape is that we detect it, we tell them we noticed and are dealing with it, and they are not handed an error they cannot solve. The subscription should feel like being looked after, not like being on call."* Today's page is the opposite shape — one switch per detector, and it **grew from 12 to 15 in a single session** (one new alarm plus two compound toggles split into four). That growth is the argument, not an aside: a page that grows per detector keeps asking a household to make engineering decisions. | **OPEN — DIRECTION, operator's call** | a decision on scope; nothing here is a bug | Recorded as a dated **[DESIGN — DIRECTION]** entry at `documentation/architecture/08-alarm-ladder.md` §8, marked plainly as *not current behaviour*. **Deliberately NOT implemented in the session that recorded it.** `app_start_failed` defaulting OFF is consistent with the direction and reversible either way, but was ruled on its own merits and does not pre-judge the redesign. | Viktor |
|
||||
| **R-229** | **The instruction-file rightsizing landed for `felhom-controller` and the workspace root; three pieces were deliberately deferred.** Done 2026-08-06: controller split into a 92-effective-line core plus four `paths:`-scoped `.claude/rules/*.md`; workspace root 208→142 effective lines with its versioned copy kept byte-identical; surgical corrections to `felhom-agent` and `felhom.eu` (expired TEMPORARY block, every version literal, the Legacy-Windows copies, the duplicated health-check rule); five contradictions resolved — including a drill-VM claim **measured live** (`qm list` on demo-hp shows VM 300 `drill-r50`; `felhom-agent` was right, `felhom-controller` was wrong); new shared `felhom.eu/scripts/instructions_gate.py` registered in `controller_gates.py` and `agent_gates.py`, 20 fixture tests + red-proof. **Leg (a) CLOSED 2026-08-06 (part 2):** `felhom.eu/CLAUDE.md` **227 → 115 effective lines**, split into a core plus `.claude/rules/{hub,website,manifests,docs}.md`; `instructions_gate` **registered in `scripts/repo_gates.py`** (six gates, all OK) in the required order — trim first, register second, because a registered-but-failing gate refuses every push. Scoping proven from the `InstructionsLoaded` hook log in two fresh sessions, not from frontmatter. **Still deferred:** (b) **CLOSED 2026-08-06 (close-out)** — `felhom-agent/CLAUDE.md` **175 → 99 effective lines** (measured 175, not 173: the CI correction added two), split into a core plus `.claude/rules/{proxmox,localapi,backup,storage}.md` beside the existing `health-checks.md`. The release section now points at the `felhom-build-deploy` skill instead of restating a table that drifts from the script. **Every `CLAUDE.md` in the workspace is now ≤120 effective lines except the workspace root at 142, which is deliberate — it is the only file re-injected after `/compact`.** (c) **CLOSED 2026-08-06 (part 2)** — all 44 orphans resolved with **zero deletions** (file count 158 before and after): 4 durable `reference`-type files indexed, 40 dated episode records moved to `.claude-memory/archive/`. `MEMORY.md` 145 → **150 lines / 17,977 bytes**, and `instructions_gate` check 6 now watches it (over-limit FAILS, orphan WARNS, absent store PASSES *printing its reason*). (d) **The spec-as-failing-test pilot** — moved to R-230. Full accounting: `audits/LEDGER-instruction-trim-2026-08-06.md` + `audits/LEDGER-instruction-trim-part2-2026-08-06.md` | **READY** — owner Viktor |
|
||||
|
||||
Reference in New Issue
Block a user