docs(hub v0.108.0): the delivery grain, the cooldown ruling, and gate 11's first subject
gates / gates (push) Successful in 15s
gates / gates (push) Successful in 15s
The alarm ladder gains §6.2 - which events are per-app, per-run, per-tier or coarse, and why the default is coarse. CONTEXT records two rulings: the grain is allow-listed rather than inferred from the payload, with crossdrive_failed as the proof that a payload rule would have been wrong; and a finding recorded only in REPORT.md has a lifetime of one session. R-389 closed and compressed, keeping its rules and naming the commit whose git show returns the full text. R-390 and R-391 left open. REPORT.md is gate 11's first real subject and passes: six observations, two FILED, four NOT-A-FINDING with their reasons. Three of those declarations are things a tidier report would have omitted - the gate's own spec would have passed the item it was built to catch, the burst has no ceiling, and ArgoCD said "successfully rolled out" while still running the old image. STATUS carries forward the one thing outstanding: the controller floor still reads 0.222.0 while the golden reads 0.223.0.
This commit is contained in:
@@ -164,3 +164,4 @@
|
||||
| **R-384** | **An app whose DATABASE had died raised no dead-app alarm — the wrong question answered first.** Shipped in controller v0.222.0. Evidence: `audits/DRILL-r384-dead-db-alarm-2026-08-23/`. **Reasoning kept:** *The defect was the ORDER of two questions, not the `unhealthy` exclusion.* "Is a SUPERVISED member dead?" and "is a RUNNING member failing its healthcheck?" are different questions, and the second was answering the first — a dying database drags its own front end `unhealthy`, so the symptom the fault causes was what suppressed the alarm for it. **`IsDownState` is byte-identical and `unhealthy` stays excluded** — an unhealthy container is RUNNING, and folding it in reintroduces the flapping that exclusion exists to stop; **no new state was minted**, `StateDegraded` already means this. **Two things had to move and either alone leaves the defect standing:** the hoist, AND widening "some members are up" from `running > 0` to *any member not in the down bucket* — the old guard made the R-51 block unreachable in precisely the case it was written for. **The register's own suggested fix was WRONG and is recorded as such:** it proposed a sustained-`unhealthy` threshold on the `crashLoopAfter` model; the actual defect needed no threshold at all. **PROVEN LIVE the only way it can be** — the same fixture that printed `0 currently down` on 2026-08-22 printed **`1 currently down`** on 2026-08-23, with `app_start_failed` 7 s after the stop and the banner reading *„…nem fut: BookStack (degraded)"*. Scenario D measured **0 alarms across 9 scans** through a full stop→start cycle. | **CLOSED — SHIPPED + PROVEN-LIVE** (controller v0.222.0, 2026-08-23) | full text: `git show 1eb64bec5183:documentation/backlog/OPEN-ITEMS.md` |
|
||||
| **R-329** | **`app_start_failed` was emitted with severity `"warn"`, so every one of them was delivered to nobody.** Shipped in controller v0.223.0 (+ hub v0.107.0). Evidence: `audits/DRILL-r329-r386-2026-08-23/`. **Reasoning kept:** *The vocabulary is EXACT and it is the HUB's, not ours* — `{info, warning, error, critical}`; anything else is coerced to `info` at ingest and dropped by `severityNotifies` before BOTH legs. **This was the SECOND occurrence** (`DiskAlertKind.Severity` until v0.215.0), and its comment had recorded the lesson — **a comment is not a guard**, so the guard is now an AST walk over the whole controller, with the six variable-passing call sites registered by name because a walk cannot follow a variable and *an unlisted limit is not a limit, it is a hole*. **The register's own framing was that the DECISION was the work** — should a stopped app mail the customer at all? Answered: **operator always, customer OFF by default**, because `processOperator` never consults customer preferences, so one word fixed the operator leg and left the customer leg exactly where the ruling wanted it. **Deliberately NOT added to `operatorOnlyEvents`** — that would make the new toggle visible, flickable and structurally incapable of delivering. **Measured on the live hub DB: 91 events stored all-time, ZERO notification rows before the fix; one operator row, `warning`/`sent`, after it.** | **CLOSED — SHIPPED + PROVEN-LIVE** (controller v0.223.0 + hub v0.107.0, 2026-08-23) | full text: `git show 68a9f5475cd2:documentation/backlog/OPEN-ITEMS.md` |
|
||||
| **R-386** | **A single-container app stopped out of band raised no alarm, and a comment stated the opposite as settled fact.** Shipped in controller v0.223.0. Evidence: `audits/DRILL-r329-r386-2026-08-23/`. **Reasoning kept:** *the state test was guessing at something the product already knows.* `DesiredState` records the customer's intent, has **exactly one writer**, and is tri-state; `StateExited` never survives aggregation, so no state test can separate an out-of-band stop from a customer stop. The ruling: `Stopped` → no alarm, `Running` → **alarm**, **absent → UNKNOWN, keep today's behaviour AND announce it**. *Reading unknown as "nobody asked" would, on the first cycle after upgrade, e-mail about every app any owner ever deliberately stopped — fleet-wide, from a field that predates the intent it is being asked about.* **A rule without a mechanism is a wish:** every such suppression sets `IntentUnknown` and the names are logged at INFO, so an operator can answer *"how many apps am I blind to?"*. **`failedRestart` must still lift a `Stopped` intent or F-CRIT-1 re-opens.** **Fenced act: adding a `DesiredState` WRITER** — twelve of `StopStack`'s fourteen callers are machines. Proven live: alarm 24 s after an out-of-band `docker compose stop`, heartbeat `1 currently down` against the previous day's `0`; and with intent removed, suppressed *plus* the log line naming the app. **0 of 8 deployed apps on `demo-hp` carry an absent intent.** | **CLOSED — SHIPPED + PROVEN-LIVE** (controller v0.223.0, 2026-08-23) | full text: `git show 68a9f5475cd2:documentation/backlog/OPEN-ITEMS.md` |
|
||||
| **R-389** | **Only the FIRST broken app per hour reached the operator — the cooldown key named the event type, not the app.** Shipped in hub v0.108.0. Evidence: `audits/DRILL-cooldown-grain-2026-08-23/`. **Reasoning kept:** the fix is a THIRD SIBLING of `cooldownTierSuffix`/`cooldownRunSuffix`, separate for the reason the second one's docstring already gives — *the existing two keep byte-identical semantics for every type that uses them.* **`cooldownStackSuffix` takes the EVENT TYPE as well as the details, unlike its siblings, and that asymmetry is the whole safety property:** `tier` and `run_id` appear only on types that want that grain, `stack_name` does not. **`perAppCooldownEvents` is a named allow-list with `app_start_failed` and nothing else** — *the backup family's cooldown is coarse ON PURPOSE (R-97a, R-182) so one full disk sends one digest rather than one mail per app*, and this is not hypothetical: **`crossdrive_failed` is severity `error`, reaches the operator leg, and carries `stack_name` through a different struct**, so a payload-shape rule would have split it silently. **The fenced act is adding an entry for a type whose family has a digest or a coarse-by-design cooldown.** `app_start_failed` qualifies precisely because it has NO digest — there is no `apps_down_run` the way `backup_run_failures` summarises a run. **The hour is unchanged; the grain was the complaint.** Fail-soft: absent or malformed details degrade to the old key and the mail still goes. **PROVEN LIVE 2026-08-23:** two apps four minutes apart gave **2 sent / 0 suppressed** where the same shape gave 1 and 1 the day before, each repeat suppressed under its OWN key (`…:opengist`, `…:calibre-web`) against the previous day's shared `key=demo-hp:app_start_failed`; and `crossdrive_failed` for two different apps stayed **coarse** under `key=demo-hp:crossdrive_failed`, byte-identical to the derived v0.107.0 value. **AND IT WAS NEVER FILED UNTIL THE DAY IT WAS FIXED** — it lived in a REPORT.md observations paragraph, which is why gate 11 now exists. | **CLOSED — SHIPPED + PROVEN-LIVE** (hub v0.108.0, 2026-08-23) | full text: `git show 45659bdc5a2f:documentation/backlog/OPEN-ITEMS.md` |
|
||||
|
||||
@@ -140,7 +140,6 @@ the fault was real. Full observables: `tests/campaign11-evidence-2026-08-05/jour
|
||||
| **R-387** | **The hub REWRITES an unknown severity and says nothing, and the guard built to catch that sits downstream of the rewrite.** One handler, two fields, opposite discipline: an unknown `event_type` is rejected with a loud `400`, while an unknown `severity` was silently coerced to `info` — after which `severityNotifies` drops it and NEITHER delivery leg runs. **Two shipped features went out that way**: `DiskAlertKind.Severity` emitted `"warn"` until controller v0.215.0, `app_start_failed` until v0.223.0. **Measured on the live hub DB 2026-08-23: 91 `app_start_failed` events stored all-time and ZERO `notification_log` rows before that day** — not one, on any channel, while every POST returned 200. **The dispatcher's `unrecognized severity` line could never execute** for an API event, because the coercion one line upstream guarantees the value it looks for cannot arrive. | **CLOSED — hub v0.107.0, 2026-08-23** | — | **The coercion STAYS; only the silence is fixed** — a rejected event is a LOST event, and losing an alarm is worse than mis-routing one. A `WARN` now names the customer, the event type, the rejected value and the consequence. **The dispatcher branch was KEPT, on evidence not caution:** `cmd/hub/main.go` wires `dispatcher.ProcessEvent` DIRECTLY as the `monitor.EventNotifyFunc` for the staleness, host-staleness and offsite-box checkers, which never pass through the handler — for them it is the only severity guard there is; deleting it as "dead" would have removed the live half while the dead half supplied the justification. All 90 severity literals in `internal/monitor` verified already valid. Proven live: `[WARN] [api] Event from demo-hp: severity "warn" is not in {info,warning,error,critical}…`, with an `error` control silent. Evidence: `audits/DRILL-r329-r386-2026-08-23/evidence/live-19-scenarioH-after.txt`. | CC |
|
||||
| **R-391** | **Gate 11 (observations) is registered in three of the four runners; `app-catalog-felhom.eu` is the exception.** The controller and agent runners already carried a shared-gate mechanism (`SHARED_REUSE`, `SHARED_INSTRUCTIONS` pointing into `felhom.eu/scripts/`), so registering there was one constant and one `GATES` line each. **`catalog_gates.py` has no such mechanism:** its `run_gate` joins every entry against its OWN `scripts/` directory, so it cannot invoke a sibling repo's script at all; and its loop appends `--all` to every gate unconditionally, which the observations gate would read as a path. Registering there therefore needs `run_gate`'s contract widened AND the argument handling changed — a refactor of a runner whose shape is deliberately different (per-app scoping, network/runtime gates excluded from `--fast`), in a repo this task marked out of scope. **The exposure today is nil** — `app-catalog-felhom.eu/REPORT.md` has no observations section, and the gate passes quietly on that — but a future catalog session could write one and nothing would read it. **Filed rather than left as a sentence in a report, which is the exact failure R-389 records.** | **OPEN — LOW** | — | Either give `catalog_gates.py` the `SHARED_*` absolute-path mechanism the other two runners already have and stop appending `--all` to gates that do not take it, or state in that repo's CLAUDE.md that its REPORT.md carries no observations section by convention. **Do not copy the gate script** — the shared checker lives in ONE place (`felhom.eu/scripts/`) and copying it is the drift the shared pattern exists to prevent. | CC |
|
||||
| **R-390** | **The golden-bake runbook omits `pveam update`, and the failure it produces names the wrong cause.** `documentation/runbooks/RUNBOOK-manual-build.md` §4.1 step 2 says to list the current Debian template because "the exact point release rots" — but on the drill VM's `virgin` snapshot **the `pveam` INDEX is stale too**, so `pveam available` offers an old point release and `pveam download local <that>` fails with **`400 Parameter verification failed. template: no such template`**. That reads as a typo or a bad argument, not as an old index, and it costs a diagnosis every time. **Hit on two consecutive bakes** (golden 0.222.0 and 0.223.0, both 2026-08-23). The runbook is otherwise correct verbatim — the qemu launch line, the token-read-inside-the-VM pattern and the acceptance markers all worked unchanged. | **OPEN — LOW** | — | Add `pveam update` as its own numbered step before the listing, and say WHY: a snapshot that never changes carries an index that never updates, so the rot warning already in the step applies to the index as well as to the release. Recorded meanwhile in the workspace memory `golden-bake-needs-pveam-update` and in `documentation/tests/golden-0.223.0-2026-08-23/README.md`. | CC |
|
||||
| **R-389** | **Only the FIRST broken app per hour reaches the operator — the cooldown key names the event type, not the app.** `dispatcher.go:337` builds the operator key as `customerID + ":" + eventType + cooldownTierSuffix(details) + cooldownRunSuffix(details)`, and **neither suffix reads an app name**. So every app that goes down inside the same hour collapses onto one key and only the first is mailed. **Measured live on `demo-hp` 2026-08-23:** `bookstack` alarmed at 09:27:51 and was `sent`; `privatebin` alarmed at 09:31:51, four minutes later, and was logged `suppressed — operator cooldown 1h, key=demo-hp:app_start_failed`. Three apps down together tonight would produce one mail. **The app's identity is already on the wire** — `AppDetails{StackName, DisplayName}` serialises as `stack_name` (`felhom-controller/internal/notify/notifier.go:146-149`), and the hub already makes exactly this kind of distinction twice, with `cooldownTierSuffix` and `cooldownRunSuffix`. **This was latent for as long as the cooldown has existed and only became reachable when R-329 made `app_start_failed` deliverable at all** — the same "a known-broken thing moves from unreachable to load-bearing" shape as R-329 itself. **AND IT WAS NEVER FILED:** it was written in a REPORT.md observations paragraph on 2026-08-23 and nowhere else — the register had no row for it until now, which is R-341's shape one surface over and is why gate 11 exists. | **OPEN — MEDIUM** | — | A third sibling suffix, `cooldownStackSuffix`, **allow-listed to `app_start_failed` and nothing else**. **Do NOT apply it globally:** the backup family's cooldown is coarse ON PURPOSE (R-97a, R-182) so that one full disk sends one digest rather than one mail per app — and `crossdrive_failed` (severity `error`) carries `stack_name` through a *different* struct (`CrossDriveDetails`), so a global suffix would silently split it per-app. The hour itself does not change; the grain is the complaint, not the length. Evidence: `audits/DRILL-cooldown-grain-2026-08-23/`. | CC |
|
||||
| **R-388** | **PRODUCT DECISION (not a defect): the customer notification model is the wrong shape, and the settings page grows by one toggle per detector.** The operator's framing, recorded verbatim 2026-08-23: *"A customer should be notified only about things they can act on or are responsible for — the drive they unplugged, the storage they filled. **A failed backup is our incident, not theirs.** The intended shape is that we detect it, we tell them we noticed and are dealing with it, and they are not handed an error they cannot solve. The subscription should feel like being looked after, not like being on call."* Today's page is the opposite shape — one switch per detector, and it **grew from 12 to 15 in a single session** (one new alarm plus two compound toggles split into four). That growth is the argument, not an aside: a page that grows per detector keeps asking a household to make engineering decisions. | **OPEN — DIRECTION, operator's call** | a decision on scope; nothing here is a bug | Recorded as a dated **[DESIGN — DIRECTION]** entry at `documentation/architecture/08-alarm-ladder.md` §8, marked plainly as *not current behaviour*. **Deliberately NOT implemented in the session that recorded it.** `app_start_failed` defaulting OFF is consistent with the direction and reversible either way, but was ruled on its own merits and does not pre-judge the redesign. | Viktor |
|
||||
| **R-229** | **The instruction-file rightsizing landed for `felhom-controller` and the workspace root; three pieces were deliberately deferred.** Done 2026-08-06: controller split into a 92-effective-line core plus four `paths:`-scoped `.claude/rules/*.md`; workspace root 208→142 effective lines with its versioned copy kept byte-identical; surgical corrections to `felhom-agent` and `felhom.eu` (expired TEMPORARY block, every version literal, the Legacy-Windows copies, the duplicated health-check rule); five contradictions resolved — including a drill-VM claim **measured live** (`qm list` on demo-hp shows VM 300 `drill-r50`; `felhom-agent` was right, `felhom-controller` was wrong); new shared `felhom.eu/scripts/instructions_gate.py` registered in `controller_gates.py` and `agent_gates.py`, 20 fixture tests + red-proof. **Leg (a) CLOSED 2026-08-06 (part 2):** `felhom.eu/CLAUDE.md` **227 → 115 effective lines**, split into a core plus `.claude/rules/{hub,website,manifests,docs}.md`; `instructions_gate` **registered in `scripts/repo_gates.py`** (six gates, all OK) in the required order — trim first, register second, because a registered-but-failing gate refuses every push. Scoping proven from the `InstructionsLoaded` hook log in two fresh sessions, not from frontmatter. **Still deferred:** (b) **CLOSED 2026-08-06 (close-out)** — `felhom-agent/CLAUDE.md` **175 → 99 effective lines** (measured 175, not 173: the CI correction added two), split into a core plus `.claude/rules/{proxmox,localapi,backup,storage}.md` beside the existing `health-checks.md`. The release section now points at the `felhom-build-deploy` skill instead of restating a table that drifts from the script. **Every `CLAUDE.md` in the workspace is now ≤120 effective lines except the workspace root at 142, which is deliberate — it is the only file re-injected after `/compact`.** (c) **CLOSED 2026-08-06 (part 2)** — all 44 orphans resolved with **zero deletions** (file count 158 before and after): 4 durable `reference`-type files indexed, 40 dated episode records moved to `.claude-memory/archive/`. `MEMORY.md` 145 → **150 lines / 17,977 bytes**, and `instructions_gate` check 6 now watches it (over-limit FAILS, orphan WARNS, absent store PASSES *printing its reason*). (d) **The spec-as-failing-test pilot** — moved to R-230. Full accounting: `audits/LEDGER-instruction-trim-2026-08-06.md` + `audits/LEDGER-instruction-trim-part2-2026-08-06.md` | **READY** — owner Viktor |
|
||||
| **R-230** | **Three instruction/memory follow-ups deliberately left by the part-2 session (2026-08-06), each needing a decision rather than an implementation.** (a) **A ruling is owed on auto-written staleness.** The hand-written `CLAUDE.md` files are now clean of version literals and expired blocks — the gate enforces it — but `MEMORY.md`, which Claude writes and which is the LARGER half of what loads (8.4k tokens vs the root file's 6.6k), carries **21 lines with component version literals**, **5 with bare host addresses**, and an entry still reading *"demo boxes REMOTE till ~08-02"* — the same expired-TEMPORARY class the gate was built to kill, now surviving in the one file the gate's content rules do not cover. **Partly actioned 2026-08-06 (close-out), and the ruling is STILL OWED:** the **three statements that were actively false** were corrected — `R-193 decision open` (closed 2026-08-05), `demo boxes REMOTE till ~08-02` (the box answers on the home LAN), `OPEN R-25b` (shipped 2026-07-21) — and gate check 6 now **WARNs** on version literals, host addresses, expired statements and stale-open citations in the index. WARN, never FAIL: Claude writes that file between sessions, so a hard failure would refuse a human's push over a line no human typed, and the warning is read by the model that will next edit it. **The remaining 32 version literals and 4 host addresses were deliberately left** for that loop. What is still owed is the bulk-correction ruling. **Correcting the premise:** the earlier report's "three expired statements" were all FALSE POSITIVES — each matched an ISO date inside a markdown link target, i.e. a filename — while the one real expired claim carried no ISO date at all. (b) **CLOSED 2026-08-06 (close-out)** — the workspace-root `CLAUDE.md` **is now a relative symlink** to the versioned copy, so the divergence class is gone rather than policed. Check 5 learned two shapes: for a link it asserts the target resolves to a real file (**a dangling link is worse than a diverged copy — the instructions load NOTHING and there is no content left to notice is wrong**), for two files byte-identity as before, so a clone elsewhere is unaffected. **Proven, not assumed:** three fresh sessions logged `session_start` for the link path, and a fourth **with no tools at all** quoted standing rule 1 verbatim — the content reaches the model, not just the path. (c) **The spec-as-failing-test pilot**, approved in principle and not started (was R-229(d)). | **READY** — owner Viktor |
|
||||
|
||||
Reference in New Issue
Block a user