docs(v0.223.0): REPORT, CONTEXT rulings, README severity contract
gates / gates (push) Successful in 11s
gates / gates (push) Successful in 11s
REPORT overwritten: the 1.1 sweep in full (one bad severity, nine legitimate "warn" strings that are healthcheck statuses), the hub manifest's real location since the task's premise was wrong, all five red-proofs with the layer each guard sits at, the live walk in six steps with the hub's own records quoted, and the absent-intent count (0 of 8). Three things are reported that a tidier account would omit: red-proof 5 passed first time because the mutation was INERT; Scenario G was silently refused twice behind an HTTP 200; and the live Scenario A does NOT prove the customer gate, because demo-hp has no prefs row at all. CONTEXT records the severity vocabulary as a ruling with its mechanism, the intent ruling with its three-way handling of unknown, both fences, and two traps worth more than the fixes: a 200 can be a refusal, and a passing red-proof can mean an inert mutation. README: the event table said `app_start_failed | warn` - the defect, written down as if correct. Now `warning`, with the vocabulary contract and who receives what. `disk_critical` also corrected from `error` to `critical`, which is what fillwatch has always sent.
This commit is contained in:
+18
-2
@@ -1971,6 +1971,22 @@ The controller pushes structured events to the Hub's `/api/v1/event` endpoint. T
|
||||
|
||||
**Core method:** `PushEvent(eventType, severity, message, details)` — non-blocking goroutine, 2 retries with 3s backoff, never blocks the caller.
|
||||
|
||||
> **⚠ THE SEVERITY VOCABULARY IS THE HUB'S, AND IT IS EXACT: `{info, warning, error, critical}`.**
|
||||
> The hub **coerces anything else to `info` at ingest**, and `info` is dropped by `severityNotifies`
|
||||
> **before both** delivery legs. So a severity outside that set means the event is stored, the POST
|
||||
> returns `200`, the dashboard shows it — and **it is e-mailed to nobody**.
|
||||
>
|
||||
> This shipped twice: `DiskAlertKind.Severity` sent `warn` until v0.215.0, and `app_start_failed` sent
|
||||
> it until **v0.223.0** — **91 of those events were stored and not one was ever delivered.** It is now
|
||||
> pinned by an AST walk over the whole controller
|
||||
> (`TestR329_EveryEmittedSeverityIsInTheHubVocabulary`); the six call sites that pass a *variable* are
|
||||
> registered by name, so a new one fails the test. Since v0.107.0 the hub also logs a `WARN` naming
|
||||
> any severity it had to rewrite. Full contract: `felhom.eu/documentation/architecture/08-alarm-ladder.md` §6.1.
|
||||
>
|
||||
> **Who receives what.** `processOperator` consults only the operator switch, the address and a
|
||||
> one-hour cooldown — **never customer preferences** — so a valid severity always reaches the operator.
|
||||
> The customer leg additionally consults `operatorOnlyEvents` and the customer's own enabled events.
|
||||
|
||||
#### Event Types
|
||||
|
||||
| Event Type | Severity | Trigger |
|
||||
@@ -1986,14 +2002,14 @@ The controller pushes structured events to the Hub's `/api/v1/event` endpoint. T
|
||||
| `health_critical` | error | Health status critical (any→fail) |
|
||||
| `health_recovered` | info | Health status recovers (fail/warn→ok) |
|
||||
| `disk_warning` | warning | Disk usage crosses 90% |
|
||||
| `disk_critical` | error | Disk usage crosses 95% |
|
||||
| `disk_critical` | **critical** | Disk usage crosses 95% (this row read `error` until v0.223.0; the emitter is `fillwatch.Band.Severity()` and it has always sent `critical`) |
|
||||
| `storage_disconnected` | error | Storage drive physically removed |
|
||||
| `storage_reconnected` | info | Storage drive reconnected |
|
||||
| `controller_started` | info | Controller process starts |
|
||||
| `controller_updated` | info/error | Self-update success or failure |
|
||||
| `app_deployed` | info | New app deployed via API |
|
||||
| `app_removed` | info | App removed via API |
|
||||
| `app_start_failed` | warn | A DEPLOYED app is not running (fix-3) — fired ONCE per running→down transition |
|
||||
| `app_start_failed` | **warning** | A DEPLOYED app is not running (fix-3) — fired ONCE per running→down transition. **Customer-switchable („Alkalmazás nem fut"), OFF by default; the OPERATOR is e-mailed regardless.** Was `warn` until v0.223.0 — see the severity note below |
|
||||
| `disaster_recovery_started` | warning | DR restore begins |
|
||||
| `disaster_recovery_completed` | info/error | DR restore finishes (success/partial) |
|
||||
|
||||
|
||||
Reference in New Issue
Block a user