docs(R-329/R-386/R-387): the severity contract, the intent ruling, and Part 5 recorded
gates / gates (push) Successful in 17s

The alarm ladder gains the severity contract (the hub's vocabulary is exact, it
coerces silently, and three things now hold it) and the intent test with its
three-way ruling on unknown. Both marked [DESIGN] with the live measurements.

Part 5 is RECORDED AND NOT IMPLEMENTED: the operator's notification philosophy,
verbatim, marked plainly as direction rather than current behaviour, with the
12 -> 15 toggle growth as the argument. Filed as R-388, a product decision.

R-329 and R-386 compressed into CLOSED-ITEMS with their rules kept and the
full-text commit named. R-387 filed closed - including WHY the dispatcher branch
was kept rather than deleted, which is evidence (three monitor checkers call
ProcessEvent directly) and not caution.

The drill record names three things that had to be re-run: an inert red-proof
mutation, Scenario G refused twice behind an HTTP 200, and the live Scenario A
NOT proving the customer gate because demo-hp has no prefs row at all.

Register: OPEN 328325 -> 328132 B, CLOSED 71441 -> 74642 B.
This commit is contained in:
2026-08-23 12:03:49 +02:00
parent 68a9f5475c
commit 2f7c9a6ce5
33 changed files with 5774 additions and 104 deletions
+81
View File
@@ -14,6 +14,87 @@
> language, one screen, no identifiers in the prose. Same subjects, different readers; merging them
> would make one of the two audiences stop reading. `STATUS.md` is also a **view of `OPEN-ITEMS.md`**
> and holds nothing of its own; this file does hold its own content, namely the standing rulings below.
## The severity a controller sends is the HUB's vocabulary, and getting it wrong deletes the alert (2026-08-23, R-329 / R-387)
**[RULING] The set is exactly `{info, warning, error, critical}`.** The hub **coerces anything else to
`info` at ingest**, and `severityNotifies` drops `info` **before both** delivery legs. So a severity
outside the set means: stored, `200` returned, dashboard shows it, **e-mailed to nobody**.
**It shipped twice** — `DiskAlertKind.Severity` until controller v0.215.0, `app_start_failed` until
v0.223.0. Measured on the live hub DB: **91 `app_start_failed` events stored all-time, ZERO
`notification_log` rows** before 2026-08-23.
**[MECHANISM, because a comment already recorded this lesson and it happened again]** An AST walk over
the whole controller pins every emitted severity. **grep cannot do this job** — `"warn"` is a
legitimate *healthcheck status* in `internal/monitor` and `internal/selftest`; the 2026-08-23 sweep hit
nine such strings and exactly one defect. The walk cannot follow a variable, so the six call sites that
pass one are **registered by name with the values each can take**; a new dynamic site fails the test.
**An unlisted limit is not a limit, it is a hole.**
**[RULING] The hub's coercion STAYS and now SPEAKS (hub v0.107.0).** A rejected event is a *lost*
event, and losing an alarm is worse than mis-routing one — that is why it was a coercion. The fix is
that it logs `WARN` naming customer, type and value. **The guard that should have caught this class sat
downstream of the rewrite and was structurally blind to it** — note the asymmetry it lived beside: an
unknown event *type* 400s loudly, an unknown *severity* was silent.
**[RULING] `dispatcher.go`'s `unrecognized severity` branch is KEPT, on evidence.** `cmd/hub/main.go`
wires `ProcessEvent` **directly** as the `monitor.EventNotifyFunc` for three checkers that never pass
the ingest handler. For those it is the only severity guard there is.
**[RULING] `app_start_failed`: operator always, customer OFF by default.** `processOperator` consults
only `operatorOn`, the address and a 1-hour cooldown — **never** customer preferences — so one word
fixed the operator leg and left the customer leg where the ruling wanted it. It is deliberately **not**
in `operatorOnlyEvents`: that would make the new toggle visible, flickable and structurally incapable
of delivering.
**[GOTCHA, now load-bearing] The operator cooldown key carries no app identifier**
(`customerID:eventType`), so **only the first app-down per hour e-mails the operator**; the rest are
logged `suppressed`. Harmless while the event was undeliverable. Not any more. R-182's shape, filed.
## "The customer stopped this" is a RECORD, never an inference (2026-08-23, R-386)
**[RULING] Ask `DesiredState`, do not read the container state.** `aggregateState` folds `StateExited`
into the stopped counter, so an all-down stack returns `StateStopped` whatever killed it — an
out-of-band stop and a customer's Stop are byte-identical on the Docker side. Measured: `privatebin`
stopped out of band, nine scans, **zero events, zero banner**, while a comment claimed it "still
alerts".
**[RULING] The tri-state, and the third value is the whole safety property.** `Stopped` → no alarm.
`Running` → **alarm**. **Absent → UNKNOWN, keep the old behaviour, AND announce it.** Reading absent as
"nobody asked" would, on the first cycle after upgrade, e-mail about every app any owner ever
deliberately stopped — fleet-wide, from a field that predates the intent being asked of it. **The
backfill cannot help: it seeds `Running` only from an observed-UP reading**, so anything stopped at
upgrade time stays unknown — exactly the ambiguous population.
**[MECHANISM] `AppRunState.IntentUnknown` + an INFO line naming the apps**, on the heartbeat cadence.
An operator must be able to ask *"how many apps am I blind to?"* and get a number. **A rule without a
mechanism is a wish.** Measured on `demo-hp`: **0 of 8** deployed apps had an absent intent.
**[FENCE] Adding a `DesiredState` WRITER is the fenced act; reading it anywhere is fine.** Twelve of
`StopStack`'s fourteen callers are machines, so recording intent in the primitive would make a nightly
backup indistinguishable from the customer pressing Stop.
**[FENCE] `failedRestart` must still lift a `Stopped` intent** — the quiesce loop stops stacks by the
same path a customer does, and one it stopped and could not restart must alarm whatever the intent
says. Removing that term re-opens F-CRIT-1.
## Two settings toggles each governed two alarms, and the labels named one (2026-08-23)
„Lemez figyelmeztetés (90%+)" also wrote `disk_critical` — the drive-is-**failing** alarm. Split into
four honest toggles; **12 → 15**. **[RULING] The migration, not the split, was the risk:** a save whose
event **set** is unchanged now stores the **existing slice verbatim**, so byte-identity is by
construction. Without that guard the defaults case reorders — the red-proof caught it. Legacy compound
form names are still read, so a stale browser tab cannot drop a key.
## Customer notification model — DIRECTION, recorded and deliberately NOT implemented (2026-08-23, R-388)
The operator's framing, verbatim: *"A customer should be notified only about things they can act on or
are responsible for… **A failed backup is our incident, not theirs.**… The subscription should feel
like being looked after, not like being on call."* Full entry at
`documentation/architecture/08-alarm-ladder.md` §8, marked **not current behaviour**. The settings page
grew 12 → 15 toggles in one session and grows by one per detector — that growth is the argument.
**Nothing in this session implements it.**
## App data placement is a DECISION, and it had never been written down as one (2026-08-22, R-376)
**Recorded here on 2026-08-22 to give an existing decision a home. It is NOT new, and this entry is
+91 -68
View File
@@ -1,97 +1,120 @@
# REPORT — felhom.eu: the golden-currency gate could not see an unrecorded golden (R-385)
# REPORT — felhom.eu: hub v0.107.0 (R-387), the alarm ladder, and golden 0.223.0
**Session 2026-08-23.** Companion to `felhom-controller` v0.222.0 (R-384, R-383) — see that repo's
**Session 2026-08-23.** Companion to `felhom-controller` v0.223.0 (R-329, R-386) — see that repo's
`REPORT.md` for the controller work and the full live walk.
## What was wrong here
## 1. Baselines, and the hub's four numbers as read
`scripts/golden_currency_gate.py` asked ONE question — *is the golden BEHIND the record?* — and
therefore could only ever catch a forgotten bake. **It said nothing when the golden was AHEAD of the
record**, and that direction is not harmless: a golden ahead of every CHANGELOG heading was built
from something never written down.
| Repo | at start | at end |
|---|---|---|
| felhom.eu | `55274d5e` | hub **v0.107.0** deployed |
| felhom-controller | `14137efa` (v0.222.0) | **v0.223.0** deployed |
| felhom-agent | `40d857b5` | untouched |
That is not hypothetical. Controller **0.221.1** was built, baked **and vouched** on 2026-08-23 while
the newest heading in the controller CHANGELOG still read `v0.221.0`. Measured on the real history,
with the old gate:
**Hub's four numbers, read live from `GET /configuration` before starting:**
`golden_version` **0.222.0** · `agent_version` **0.130.0** · `min_agent` **0.129.0** ·
controller floor **0.222.0**. All four as the task predicted; the operator's 0.222.0 vouch had landed.
```
newest released controller : 0.221.0 (## v0.221.0 — taking the undo copy destroyed …)
newest golden baked : 0.221.1 (documentation/tests/golden-0.221.1-2026-08-23)
golden currency gate OK …
EXIT=0
```
## 2. The hub's deployment path — §6's premise was wrong, and here it is
Every gate was green while the fleet ran a version the record did not name.
**`felhom.eu/manifests/hub.yaml`, line 128.** ArgoCD `Application/felhom` tracks
`https://gitea.dooplex.hu/admin/felhom.eu.git`, path `manifests`, with `syncPolicy.automated.enabled
= false`. Bumped `0.106.0 → 0.107.0` in commit **`68a9f54`**. **There is no out-of-git deployment
path** — the finding §6 braced for does not exist. The image was built and pushed to the registry
**before** the manifest landed, so a sync could never have pointed at a missing tag, and the sync was
then requested deliberately (`refresh=hard`, then a patched `operation`). Never `kubectl set image`.
## The fix, and why it is membership and not a comparison
## 3. R-387 — what was wrong
The gate now asks **"is the version we are shipping WRITTEN DOWN?"** — the baked version must have its
own `## vX.Y.Z` heading **anywhere** in the controller CHANGELOG, not merely at the top (an entry may
legitimately be overtaken by later ones; what may never happen is that it is absent).
One handler, two fields, opposite discipline: an unknown `event_type` is rejected with a loud `400`;
an unknown `severity` was rewritten to `info` **without a word**, and `severityNotifies` drops `info`
before *both* legs. **The guard built to catch exactly this sat downstream of the rewrite** — the
dispatcher's `unrecognized severity` line can never execute for an API event, because the coercion one
line earlier guarantees the value it looks for cannot arrive.
**Membership, not `baked > released`, deliberately:** a comparison against the newest heading alone
goes green the moment ANY later entry is written — which would have left 0.221.1 permanently
unrecorded and the gate permanently silent about it.
**Measured on the live hub DB:** `91` `app_start_failed` events stored all-time, **`0`
`notification_log` rows before this session** — not one, on any channel, while every POST returned 200.
Preserved unchanged: **INCONCLUSIVE (exit 2)** for an absent clone or an unparseable CHANGELOG — *not
knowing is never a pass, and never a conviction*. Every refusal names a reason **and** a route,
including what to do if a bake was a throwaway that must never be delivered.
**The coercion stays.** A rejected event is a *lost* event, and losing an alarm is worse than
mis-routing one. Only the silence is fixed.
## Red-proofs — both directions, against the real history
## 4. The dead-branch decision, and the reason
| Run | Gate | CHANGELOG | Golden baked | Exit | |
|---|---|---|---|---|---|
| `gate-01` | **old** | v0.221.0 | 0.221.1 | **0** | the blindness, reproduced |
| `gate-02` | **new** | v0.221.0 | 0.221.1 | **1** | convicted |
| `gate-03` | new | v0.221.1 | 0.221.1 | **0** | Part 0's heading makes it pass |
| `gate-04` | new | absent clone | — | **2** | INCONCLUSIVE preserved |
| `gate-05` | new | v0.222.0 | 0.222.0 | **0** | post-bake |
**KEPT.** Not caution — evidence. `cmd/hub/main.go` wires `dispatcher.ProcessEvent` **directly** as
the `monitor.EventNotifyFunc` for the staleness, host-staleness and offsite-box checkers, and those
hub-generated events never pass through the ingest handler at all. For every one of them that line is
the **only** severity guard there is. Deleting it as "dead" would have removed the live half while the
dead half supplied the justification.
Transcripts: `documentation/audits/DRILL-r384-dead-db-alarm-2026-08-23/evidence/gate-0*.txt`.
Verified while deciding: **all 90 severity literals in `internal/monitor` are already valid**, so the
guard is silent because the producers are correct. (`"warn"` in `internal/web` is UI badge vocabulary,
not a severity.)
## Files changed
## 5. Files changed, commits, CI
| File | Change |
| Commit | Contents |
|---|---|
| `scripts/golden_currency_gate.py` | the unrecorded-golden conviction; `newest_released` → `released_versions` returning the whole set; docstring records the second blindness |
| `documentation/architecture/08-alarm-ladder.md` | **NEW.** The alarm ladder as a dated [DESIGN] |
| `documentation/architecture/00-capability-map.md` | R-384 marked closed with its live evidence; points at the new doc |
| `documentation/backlog/OPEN-ITEMS.md` | R-383/R-384 removed (closed); **R-385** (closed) and **R-386** (open) filed |
| `documentation/backlog/CLOSED-ITEMS.md` | R-383 + R-384 compressed, each keeping its rules and naming `git show 1eb64bec5183:…` for the full text |
| `documentation/tests/golden-0.222.0-2026-08-23/` | **NEW.** Bake evidence + log + the vouching instructions |
| `STATUS.md` | the 0.222.0 vouch replaces the (now completed) 0.221.1 one; R-386 added in plain words |
| `documentation/audits/DRILL-r384-dead-db-alarm-2026-08-23/` | **NEW.** The drill record and 29 evidence files |
| **`68a9f54`** | hub v0.107.0 (ingest WARN + kept-branch note + tests), manifest bump, golden 0.223.0 evidence |
| **`<docs>`** | alarm ladder §6.1/§7/§8, register, `STATUS.md`, `REPORT.md`, drill record |
## The alarm ladder had no owning document — that absence is a finding
Files: `hub/internal/api/handler.go`, `hub/internal/notify/dispatcher.go`, `hub/CHANGELOG.md`,
`manifests/hub.yaml`, `documentation/architecture/08-alarm-ladder.md`,
`documentation/backlog/{OPEN,CLOSED}-ITEMS.md`, `documentation/tests/golden-0.223.0-2026-08-23/`,
`documentation/audits/DRILL-r329-r386-2026-08-23/`, plus two new test files.
Nothing in `documentation/architecture/` owned the question *"when does a customer's app being broken
raise an alarm?"* The rules lived as comments across four packages, each locally correct, with the
ordering between them legible only by reading `aggregateState` top to bottom. **That is precisely how
R-384 survived review**, and three separate defects in this ladder (R-51, C9-F2, R-384) were each
found on live hardware rather than by reading. `08-alarm-ladder.md` now owns it.
**CI runs confirmed BY ID** (`id` and `run_number` diverge — both printed): see §5 of the controller
REPORT for its runs; felhom.eu's are listed at the end of this file.
## Golden
## 6. Tests and red-proofs
**Baked and PUBLISHED: 0.222.0.** `GOLDEN_SHA256 = 19f5904f5379…`, `upload OK (HTTP 201)`, round-trip
`HTTP 206` from the package URL, all acceptance markers counted.
**VOUCHING IS THE OPERATOR'S ACT AND WAS NOT DONE HERE.**
`internal/api/r387_severity_visibility_test.go` — the event is **not lost**, the stored severity is
still `info`, and the WARN names customer + type + value; plus a guard that a **valid** severity stays
silent, because an alarm on the normal path is one people learn to ignore.
`internal/notify/r329_app_start_failed_test.go` — the routing consequence: operator emailed, customer
not, unless opted in, in which case both legs deliver and the customer's copy carries the Hungarian
template. Scenario B is also the **positive control** for Scenario A's absence claim.
**Deviation recorded:** the bake runbook's §4.1 is missing a `pveam update`. On the `virgin` snapshot
the template index is stale, so the listed template cannot be downloaded and the failure presents as
`400 Parameter verification failed. template: no such template` rather than as a stale index.
Test count **702 → 709**.
## Register size
**Red-proof (seen failing):** delete the ingest `WARN` → `the hub rewrote a severity and said
nothing`, with the log showing only the ordinary `[INFO] Event from c1: backup_failed (info)`. The
guard sits at **ingest**, because that is the last point at which the offending value still exists.
## 7. Golden
**Baked and PUBLISHED: 0.223.0.** `GOLDEN_SHA256 =
9eaf39ac39219b42ec9e6cbf890275febcdcc6f53325fe0c0f591d3431044f17`; `upload OK (HTTP 201)`; round-trip
**HTTP 206**; all five acceptance markers counted (`docker OK (overlay2` 1, `including mount point` 2,
`upload OK` 1, `excluding` 0, `FATAL` 0). **VOUCHING IS THE OPERATOR'S ACT AND WAS NOT DONE HERE.**
**Runbook deviation, second session running:** §4.1 omits `pveam update`, so the `virgin` snapshot's
stale template index fails as `400 … no such template`.
## 8. Scenario H, live against v0.107.0
```
[WARN] [api] Event from demo-hp: severity "warn" is not in {info,warning,error,critical} — coercing
to "info", which severityNotifies DROPS, so this backup_failed alert will reach NOBODY. Fix the
emitting controller; this event is stored but not routed.
```
Both the bad-severity POST and the `error` control returned **200** (nothing lost), and the control
produced **no** warning — the guard does not fire on the normal path.
## 9. Register size
| File | Before | After |
|---|---|---|
| `OPEN-ITEMS.md` | 327,266 B | **328,325 B** |
| `CLOSED-ITEMS.md` | 68,464 B | **71,441 B** |
| `OPEN-ITEMS.md` | 328,325 B | **328,132 B** |
| `CLOSED-ITEMS.md` | 71,441 B | **74,642 B** |
OPEN grew ~1 KB despite two closures, because R-386 is a substantial new finding. Recorded rather
than smoothed over.
R-329 and R-386 compressed into CLOSED with their rules kept; **R-387** (closed) and **R-388** (the
notification-model product decision, open, operator's call) filed.
## Hub numbers as read at session start (live, `GET /configuration`)
## 10. Observations — recorded, not acted on
`golden_version` **0.221.1** · `agent_version` **0.130.0** · `min_agent` **0.129.0** ·
controller floor **0.221.1**. The task expected 0.220.2/0.220.2; the operator had already vouched.
**The hub was READ ONLY this session** — nothing was written to it.
1. **The operator cooldown key carries no app identifier.** PrivateBin's alarm four minutes after
BookStack's was logged `suppressed — operator cooldown 1h, key=demo-hp:app_start_failed`, so **only
the first app-down per hour reaches the operator by e-mail**. R-182's known shape; harmless while
the event was undeliverable, and no longer. Not fixed here.
2. Two probe events remain as rows for `demo-hp` from Scenario H — inert, and named rather than left.
+15 -20
View File
@@ -1,7 +1,7 @@
# STATUS — what works, what's broken, what's next
**Updated 2026-08-23 — an app whose database dies now raises an alarm. It did not before, and the
watcher said "nothing is down" the whole time. Released and NOT yet delivered: step 1 is yours.**
**Updated 2026-08-23 — the app-down alarm now actually reaches you by e-mail. It never has: 91 of
them were filed and not one was ever sent. Released and NOT yet delivered: step 1 is yours.**
> **A view, not a source.** `documentation/backlog/OPEN-ITEMS.md` is the authority; this page restates
> part of it in plain words, and **nothing may exist only here**. **Items, not paragraphs. One screen.**
@@ -12,26 +12,21 @@ watcher said "nothing is down" the whole time. Released and NOT yet delivered: s
*This section is allowed to be longer than one screen, and each item says what happens if you do
nothing.*
1. **Vouch the golden carrying controller 0.222.0** — Hub → Configuration → Day-0 artifacts.
1. **Vouch the golden carrying controller 0.223.0** — Hub → Configuration → Day-0 artifacts.
**It is already baked, published and round-trip verified**
(`documentation/tests/golden-0.222.0-2026-08-23/`); only the vouch is left, and only you can do it.
**It is a THREE-field save:** `golden_version` → **0.222.0**, `agent_version` → **0.130.0**,
`min_agent` → **0.129.0**. **Then** raise the floor to **0.222.0**, last, in its own save.
**If you do nothing:** the fleet stays on 0.221.1, where an app whose database has died reports
nothing at all — no banner, no event — and the watcher keeps printing "0 currently down". New
machines still receive 0.221.1.
*(Thank you — the 0.221.1 vouch from earlier today has landed; the hub reads golden 0.221.1 and
floor 0.221.1. Nothing is owed on that one.)*
(`documentation/tests/golden-0.223.0-2026-08-23/`); only the vouch is left, and only you can do it.
**It is a THREE-field save:** `golden_version` → **0.223.0**, `agent_version` → **0.130.0**,
`min_agent` → **0.129.0**. **Then** raise the floor to **0.223.0**, last, in its own save.
**If you do nothing:** the fleet stays on 0.222.0, where the app-down alarm shows on the dashboard
and e-mails nobody, and where an app stopped from outside is still reported as if you had stopped
it yourself. New machines still receive 0.222.0.
*(The hub half is already live — v0.107.0 deployed itself through the manifest. Nothing owed there.)*
2. **A new switch has appeared for your customers, and it is OFF** — „Alkalmazás nem fut". **You get
the e-mail either way**; the switch only decides whether the customer also does. This is what you
asked for and it needs nothing from you. Mentioned so it is not a surprise the first time you see
the settings page.
2. **An app that is stopped from outside still reports nothing** (R-386) — and this one I found today
and deliberately did **not** fix. If a single-container app is stopped by hand on the machine
rather than through the product, nothing is said: no banner, no e-mail, no operator event. I
measured it: nine checks ran over four minutes and every one stayed silent. A comment in our own
code claims the opposite, which is why nobody noticed. **The reason I stopped rather than fixed it:**
from the outside this looks exactly like a customer pressing Stop, and the obvious fix would start
alarming every time somebody legitimately stops their own app. That trade is a decision, not a
patch. **If you do nothing:** it stays as it is — this is not a new fault, it has always been so;
it is newly *known*.
3. **Whether to change the hub password** (R-350). I printed it into my own session log on 20 August.
Not in git, not in any saved file — in the log on this machine. **If you do nothing:** it stays as
it is, at the risk you accept by leaving it. I can change it without ever showing you the new one.
+99 -14
View File
@@ -121,21 +121,106 @@ live 2026-08-23 — one event across 22 scans.
`[deadapp] check alive: N scans since boot, M deployed app(s) evaluated, K currently down`.
This line exists because an absent alarm and a stopped detector look identical in a log.
> ⚠ **R-329, OPEN and it bites here.** `app_start_failed` is pushed with severity **`warn`**, which is
> **not** in the hub's vocabulary (`{info, warning, error, critical}`) and is silently coerced to
> `info` — which e-mails nobody, while the POST still returns 200. Observed again on 2026-08-23:
> `PushEvent: type=app_start_failed severity=warn`. R-384 makes this event actually fire, so the
> severity bug now matters more than it did while the event was unreachable.
### 6.1 The severity contract [DESIGN, R-329 — CLOSED controller v0.223.0 / hub v0.107.0]
**The vocabulary is the HUB's and it is exact: `{info, warning, error, critical}`.** Anything else is
**coerced to `info` at ingest**, and `info` is dropped by `severityNotifies` before *both* delivery
legs. So a severity outside the set means the event is stored, answers `200`, shows on the dashboard —
and is e-mailed to **nobody**.
**This shipped twice.** `DiskAlertKind.Severity` emitted `"warn"` until controller v0.215.0;
`app_start_failed` emitted it until v0.223.0. Measured on the live hub DB 2026-08-23: **91
`app_start_failed` events stored all-time, ZERO `notification_log` rows before that day** — not one,
on any channel.
Three things now hold it:
1. **The emitter is pinned by an AST walk** over the whole controller
(`TestR329_EveryEmittedSeverityIsInTheHubVocabulary`). Not grep — "warn" is a legitimate
*healthcheck status* in `internal/monitor` and `internal/selftest`. The six call sites that pass a
variable are registered by name with the values each can take, so a new dynamic path fails.
2. **The hub SAYS SO** when it coerces (hub v0.107.0, R-387): a `WARN` naming the customer, the event
type and the rejected value. **The coercion stays** — a rejected event is a *lost* event, and
losing an alarm is worse than mis-routing one.
3. **The dispatcher's `unrecognized severity` branch is kept**, because the hub's own monitor checkers
call `ProcessEvent` directly and never pass the ingest handler. For them it is the only guard.
**Who gets it.** `processOperator` consults only `operatorOn`, the address and a 1-hour cooldown —
**never customer preferences** — so a valid severity always reaches the operator. `processCustomer`
consults `operatorOnlyEvents` and then the customer's `enabled_events`.
**`app_start_failed` is customer-switchable but OFF by default** [DESIGN, operator ruling 2026-08-23]:
it is deliberately absent from `DefaultEnabledEvents`, and deliberately **not** in `operatorOnlyEvents`
— being in that register would make the toggle visible, flickable and structurally unable to deliver.
---
## 7. Known gap, filed not fixed
## 7. The intent test [DESIGN, R-386 — CLOSED controller v0.223.0]
> **R-386 (filed 2026-08-23, OPEN).** An all-down stack aggregates to `stopped` — `StateExited` is
> folded into the same counter and never survives aggregation. `classifyRunStates` then whitelists
> `stopped` as a deliberate user stop. So a **single-container app stopped out of band raises no
> alarm at all**, which directly contradicts the comment at `cmd/controller/main.go`: *"An out-of-band
> `docker compose stop` leaves the containers present → StateExited → still alerts."*
> **Measured on `demo-hp` 2026-08-23:** `privatebin` stopped out of band, 9 dead-app scans over 4+
> minutes, `state=stopped`, **zero events and zero banner lines** — against a positive control from
> the same box 17 minutes earlier. Not fixed in v0.222.0 deliberately; it is a separate decision.
**"The customer stopped this" is asked of the FIELD THAT RECORDS IT, never inferred from the state.**
Until v0.223.0 `classifyRunStates` read `st.State == StateStopped` and assumed every stopped stack was
deliberate. It is not inferable: `aggregateState` folds `StateExited` into the stopped counter, so an
all-down stack returns `StateStopped` whatever killed it. Measured on `demo-hp` 2026-08-23:
`privatebin` stopped out of band, nine dead-app scans over four minutes, **zero events, zero banner
lines** — while a comment beside the code claimed an out-of-band stop *"still alerts"*.
`DesiredState` records the answer, has **exactly one writer** (the customer's own action), and is
tri-state:
| Intent | Verdict | Why |
|---|---|---|
| `Stopped` | **no alarm** | the customer asked |
| `Running` | **ALARM** | nobody asked — the R-386 case |
| absent (`""`) | **no alarm, and SAY SO** | UNKNOWN never means running |
**The absent case keeps the old behaviour deliberately.** Reading it as "nobody asked" would, on the
first cycle after upgrade, e-mail about every app any owner ever stopped — fleet-wide, from a field
that predates the intent being asked of it. The backfill cannot help: it seeds `Running` only from an
observed-**up** reading, so anything stopped at upgrade time stays unknown, which is precisely the
ambiguous population.
**The gap is BOUNDED, not silent.** Every such suppression sets `AppRunState.IntentUnknown`, and the
scheduler logs the names at `INFO` on the heartbeat cadence:
```
[deadapp] N stopped app(s) have NO recorded customer intent, so their dead-app alarm is
suppressed by the unknown-intent fallback (R-386): <names>. This closes itself as each app is
started or stopped through the interface.
```
**A rule without a mechanism is a wish.** Measured on `demo-hp` 2026-08-23: **0 of 8 deployed apps had
an absent intent** — the population is already empty on an exercised box; it will be larger on one
upgraded and left alone.
`failedRestart` still lifts a `Stopped` intent, and that ordering is load-bearing: the quiesce loop
stops stacks by the same path a customer does, so one it stopped and could not restart must alarm
whatever the intent says. Removing that term re-opens F-CRIT-1.
**Fenced act:** adding a `DesiredState` **writer**. Reading it anywhere is fine. Twelve of
`StopStack`'s fourteen callers are machines, so recording intent in the primitive would make a nightly
backup indistinguishable from the customer pressing Stop.
---
## 8. Direction — who a customer should be notified about at all
**[DESIGN — DIRECTION, NOT CURRENT BEHAVIOUR. Dated 2026-08-23, the operator's own framing.
Nothing in controller v0.223.0 / hub v0.107.0 implements this.]**
> **A customer should be notified only about things they can act on or are responsible for** — the
> drive they unplugged, the storage they filled. **A failed backup is our incident, not theirs.** The
> intended shape is that we detect it, we tell them we noticed and are dealing with it, and they are
> not handed an error they cannot solve. The subscription should feel like being looked after, not
> like being on call.
Today's settings page is the opposite shape: it exposes one toggle per detector and **grew from 12 to
15 in this session alone** (one new alarm, plus two compound toggles split into four). That growth is
the argument, not an aside — a page that grows by one per detector is a page that will keep asking a
household to make engineering decisions.
`app_start_failed` defaulting **off** is consistent with this direction and reversible either way; it
was ruled that way on its own merits and does not pre-judge the redesign.
**Filed as a PRODUCT DECISION, not a defect** — see the register. It is the operator's call to take
separately, and no part of it was implemented here.
@@ -0,0 +1,110 @@
# DRILL — R-329 + R-386: the alarm that fired but reached nobody, and the stop nobody heard (2026-08-23)
**Controller v0.222.0 → v0.223.0. Hub v0.106.0 → v0.107.0. Live leg on `demo-hp` (Tier 0), guest
9201. UNATTENDED.** Method: endpoint-level plus the hub's own SQLite records — no browser exists on
DooPlex. Guest and hub clocks are UTC; the hub pod logs CEST.
## Verdict
| Part | Outcome |
|---|---|
| 1.1 the severity word + sweep | ✅ **exactly one** bad severity in the whole controller |
| 1.2 the AST contract guard | ✅ and it found two dynamic sites the hand sweep missed |
| 1.3 customer toggle, default OFF | ✅ |
| 2 the hub says what it rewrote | ✅ hub v0.107.0, proven live |
| 3 ask the field that knows | ✅ proven live, both directions |
| 4 split the compound toggles | ✅ round trip byte-identical |
| 5 record the notification philosophy | ✅ recorded, **not implemented** |
**No halt condition fired.**
## The one number that says it all
Read from the live hub DB:
```
91 app_start_failed events stored, all-time
0 notification_log rows before 2026-08-23 09:00 <- not one, ever, on any channel
```
After the fix, at 09:27:51: **one row — `warning` / `sent` / `operator`.**
## The two live pairs
**Scenario A** — a database dies, customer has not opted in:
```
events : demo-hp app_start_failed warning 2026-08-23 09:27:51 <- v0.223.0
demo-hp app_start_failed info 2026-08-23 05:30:14 <- v0.222.0, coerced
notification_log: demo-hp app_start_failed warning sent operator 09:27:51
(no customer row)
```
**Scenario D** — an app stopped out of band, intent `running`:
```
2026/08/23 05:47:44 [deadapp] check alive: 40 scans since boot, 8 deployed app(s) evaluated, 0 currently down <- v0.222.0
2026/08/23 09:34:51 [deadapp] check alive: 20 scans since boot, 8 deployed app(s) evaluated, 1 currently down <- v0.223.0
```
Alarm fired **24 seconds** after the `docker compose stop`.
## Three things that had to be re-run, and why that matters
1. **Red-proof 5 passed first time — the mutation was INERT.** Changing `if next <= prev` to
`if next < prev` in fillwatch does nothing, because an earlier `if next == prev { continue }` had
already removed the equal case. The test was right to pass. Removing the guard outright convicts
it. **Check the mutation applied before believing either verdict.**
2. **Scenario G was silently refused TWICE behind an HTTP 200.** The empty-email wipe guard declines
the save and renders an error page — still `200`. The first run had no prefs at all; the second
read the email with a single-line grep from a `<input>` that spans **three lines**, got `""`, and
was refused again. Both times the before/after hashes matched — *because nothing was saved*, not
because nothing changed. Fixed by asserting the refusal banner is **absent**. **A warning beside a
success is read as a success.**
3. **The live Scenario A does NOT prove the customer gate**, and is not claimed to. `demo-hp` has no
`customer_notifications` row at all, so the customer leg could not have delivered regardless. The
toggle gate is proven by the unit tests, which configure prefs both ways. Stated rather than
implied.
## Evidence index (`evidence/`)
| File | What it shows |
|---|---|
| `redproof-1-R329-emitter.txt` | severity back to `"warn"` → AST guard names file, line and value |
| `redproof-2-R386-intent.txt` | intent test reverted → `dead-app banner = []`, the live symptom |
| `redproof-3-part4-migration.txt` | no-op-save guard removed → the `defaults` case reorders |
| `redproof-4-R387-ingest.txt` | hub WARN removed → "the hub rewrote a severity and said nothing" |
| `redproof-5-fillwatch-consequence.txt` | the inert first attempt, and the effective one |
| `live-01`…`live-04` | Scenario A: stop, controller send, **hub records**, customer-leg control |
| `live-05-step4-absent-intent-count.txt` | **0 of 8** deployed apps carry an absent intent |
| `live-07`…`live-09` | Scenario D: out-of-band stop, alarm in 24 s, the heartbeat pair |
| `live-10`, `live-11` | Scenario C: UI Stop → intent `running`→`stopped`, 0 alarms / 9 scans |
| `live-12`, `live-13` | Scenario E: intent removed → suppressed **and** the log line names the app |
| `live-16-scenarioG-roundtrip.txt` | all three attempts, ending byte-identical (`10840f3a…`) |
| `live-18-golden-bake.txt` | golden 0.223.0 markers, each counted |
| `live-19-scenarioH-after.txt` | the new hub WARN line, live, with a silent `error` control |
| `live-06`, `live-14` | full controller-log windows (1041 and 4044 lines), pulled before each revert |
## Observations — noticed, recorded, NOT acted on
1. **The operator cooldown has no app identifier, and it now bites.** PrivateBin's event 4 minutes
after BookStack's was logged `suppressed — operator cooldown 1h, key=demo-hp:app_start_failed`. So
**only the first app-down per hour e-mails the operator.** This is R-182's known cooldown-key
shape; it was harmless while the event was undeliverable and is not any more. Recorded, not fixed.
2. **The prompt's §6 premise was wrong and is corrected:** the hub's manifest **is** in version
control, at `felhom.eu/manifests/hub.yaml:128`, and ArgoCD app `felhom` tracks
`admin/felhom.eu.git` path `manifests`. No out-of-git deployment path exists.
3. **The golden-bake runbook still lacks `pveam update`** — second bake in a row to hit the stale
index on the `virgin` snapshot, presenting as `400 … no such template`.
4. `internal/notify/notifier.go` carries **pre-existing** gofmt drift in a const block, confirmed by
stashing this session's work. Not touched (§12).
## Teardown
Nothing provisioned. Every app restarted and confirmed healthy (17 containers). `privatebin`'s
`app.yaml` restored from its backup and the backup removed; its intent reads `running` again.
`demo-hp`'s notification settings restored to `enabled_events: null`, no e-mail — the state they were
in before the drill. The drill VM is powered off, its Gitea token shredded, and `drill.qcow2` reverted
to `virgin`. **Hub-side: two probe events (`backup_failed`, "R-387 scenario H probe"/"control") were
POSTed to the live hub for Scenario H and remain as event rows for customer `demo-hp`.** They are
inert records; named here rather than left for someone to find.
@@ -0,0 +1,10 @@
=== LIVE WALK PRE-STATE (controller v0.223.0) ===
2026-08-23T09:25:42Z
gitea.dooplex.hu/admin/felhom-controller:0.223.0 Up About a minute (healthy)
docmost Up 4 hours (healthy)
docmost-postgres Up 4 hours (healthy)
docmost-redis Up 4 hours (healthy)
privatebin Up 4 hours (healthy)
bookstack Up 7 hours (healthy)
bookstack-db Up 4 hours (healthy)
@@ -0,0 +1,6 @@
=== SCENARIO A — an app's database dies; customer has NOT opted in ===
waiting out the 90s dead-app boot grace first...
T0=2026-08-23T09:27:27Z
2026-08-23T09:27:32Z
bookstack Up 7 hours (healthy)
bookstack-db Exited (0) 5 seconds ago
@@ -0,0 +1,4 @@
-- the CONTROLLER only proves it SENT (quoted for the severity word) --
2026/08/23 09:27:51 notifier.go:206: [DEBUG] PushEvent: type=app_start_failed severity=warning url=https://hub.felhom.eu/api/v1/event
2026/08/23 09:27:51 notifier.go:232: [DEBUG] PushEvent: app_start_failed pushed OK (HTTP 200)
2026/08/23 09:27:51 notifier.go:234: [INFO] Event pushed: app_start_failed (warning) — Telepített alkalmazás nem fut: BookStack
@@ -0,0 +1,14 @@
=== SCENARIO A — THE HUB'S OWN RECORDS (not the controller's) ===
-- 1. the STORED event: severity as the hub filed it --
customer_id event_type severity message created_at
----------- ---------------- -------- ---------------------------------------- -------------------
demo-hp app_start_failed warning Telepített alkalmazás nem fut: BookStack 2026-08-23 09:27:51
demo-hp app_start_failed info Telepített alkalmazás nem fut: BookStack 2026-08-23 05:30:14
demo-hp app_start_failed info Telepített alkalmazás nem fut: BookStack 2026-08-22 21:27:18
demo-hp app_start_failed info Telepített alkalmazás nem fut: Kimai 2026-08-22 14:19:12
-- 2. the NOTIFICATION LOG: which channel actually delivered --
customer_id event_type severity status channel created_at
----------- ---------------- -------- ------ -------- -------------------
demo-hp app_start_failed warning sent operator 2026-08-23 09:27:51
@@ -0,0 +1,26 @@
-- 3. IS THE CUSTOMER LEG EVEN CAPABLE? (or is 'no customer row' true for the wrong reason) --
Error: in prepare, no such table: notification_prefs
-- the customer leg DOES deliver other event types for this same customer: --
event_type severity status channel created_at
------------------- -------- ------- -------- -------------------
backup_run_failures error skipped customer 2026-08-22 22:20:32
backup_run_failures error skipped customer 2026-08-22 22:17:44
backup_run_failures error skipped customer 2026-08-22 21:11:19
backup_run_failures error skipped customer 2026-08-22 16:06:21
backup_run_failures error skipped customer 2026-08-21 21:20:36
-- 4. HOW MANY notification rows did app_start_failed EVER produce before today? --
0 rows before 2026-08-23 09:00
91 app_start_failed EVENTS stored, all-time
-- 3b. the customer's own preferences (does app_start_failed appear?) --
-- 3b. the customer's own preferences (does app_start_failed appear?) --
-- HONESTY CHECK: demo-hp has NO customer_notifications row at all --
0 prefs rows for demo-hp
-- why the customer channel logs 'skipped' for other types: --
event_type status reason
------------------- ------- -------------
backup_run_failures skipped operator_only
backup_run_failures skipped operator_only
backup_run_failures skipped operator_only
@@ -0,0 +1,11 @@
=== STEP 4 — THE ABSENT-INTENT COUNT ON demo-hp (deployed apps only) ===
bookstack running
calibre-web running
docmost running
kimai running
opengist running
paperless-ngx running
privatebin running
romm running
DEPLOYED=8 running=8 stopped=0 ABSENT=0
@@ -0,0 +1,8 @@
=== SCENARIO D — an app stopped OUT OF BAND, intent recorded as Running ===
subject: privatebin (single container, desired_state: running) — R-386's measured-silent case
desired_state: running
T0=2026-08-23T09:31:27Z
2026-08-23T09:31:33Z
privatebin Exited (0) 5 seconds ago
bookstack Up 7 hours (healthy)
bookstack-db Up 10 seconds (healthy)
@@ -0,0 +1,6 @@
-- SCENARIO D verdict (T0=09:31:27Z) --
2026-08-23T09:32:41Z
privatebin Exited (0) About a minute ago
app_start_failed events since T0: 1
2026/08/23 09:31:51 notifier.go:234: [INFO] Event pushed: app_start_failed (warning) — Telepített alkalmazás nem fut: PrivateBin
@@ -0,0 +1,17 @@
-- SCENARIO D: the heartbeat, new beside last session's --
LAST SESSION (v0.222.0, R-386 open, privatebin stopped out of band):
2026/08/23 05:5x [deadapp] check alive: ... 8 deployed app(s) evaluated, 0 currently down
(9 scans, 0 events, 0 banner lines — evidence: DRILL-r384-.../live-18-sec4-verdict.txt)
NOW (v0.223.0):
2026/08/23 09:34:51 main.go:1731: [INFO] [deadapp] check alive: 20 scans since boot, 8 deployed app(s) evaluated, 1 currently down
THE PAIR, exactly as logged — same box, same fixture (privatebin stopped out of band), same 8 apps:
2026/08/23 05:47:44 [deadapp] check alive: 40 scans since boot, 8 deployed app(s) evaluated, 0 currently down <- v0.222.0
2026/08/23 09:34:51 [deadapp] check alive: 20 scans since boot, 8 deployed app(s) evaluated, 1 currently down <- v0.223.0
-- and the hub's own record for it --
event_type severity status channel message created_at
---------------- -------- ---------- -------- ----------------------------------------- -------------------
app_start_failed warning suppressed operator Telepített alkalmazás nem fut: PrivateBin 2026-08-23 09:31:51
app_start_failed warning sent operator Telepített alkalmazás nem fut: BookStack 2026-08-23 09:27:51
@@ -0,0 +1,11 @@
=== SCENARIO C — the customer presses Stop in the interface ===
-- intent BEFORE --
desired_state: running
-- restart it first, so the Stop is a real transition --
{"ok":true,"message":"Stack privatebin start completed"}
start=200
T_STOP=2026-08-23T09:36:13Z
{"ok":true,"message":"Stack privatebin stop completed"}
stop=200
-- intent AFTER --
desired_state: stopped
@@ -0,0 +1,7 @@
-- SCENARIO C verdict: 4 minutes past the Stop (T=09:36:13Z), every grace elapsed --
2026-08-23T09:40:37Z
app_start_failed since the Stop : 0
deadapp scans since the Stop : 9
unknown-intent log lines : 0
@@ -0,0 +1,7 @@
=== SCENARIO E — an app with NO recorded intent (the legacy/upgrade population) ===
demo-hp has ZERO such apps (all 8 read 'running'), so one is CREATED for the measurement:
the desired_state key is removed from privatebin's app.yaml, exactly as a box upgraded from
before R-166 would look. Backed up and restored afterwards. No code writer was added.
desired_state line now: [0 found]
T0=2026-08-23T09:40:55Z
Up 30 seconds (healthy)
@@ -0,0 +1,11 @@
-- SCENARIO E verdict --
2026-08-23T09:51:48Z
app_start_failed events : 0
deadapp scans : 21
-- is privatebin actually being EVALUATED? (a suppression nobody reaches proves nothing) --
privatebin deployed=True state=stopped
-- waiting for the heartbeat scan (deadAppScans lags the log count by the boot grace) --
2026/08/23 09:51:56 main.go:1731: [INFO] [deadapp] check alive: 20 scans since boot, 8 deployed app(s) evaluated, 0 currently down
2026/08/23 09:51:56 main.go:1760: [INFO] [deadapp] 1 stopped app(s) have NO recorded customer intent, so their dead-app alarm is suppressed by the unknown-intent fallback (R-386): privatebin. This closes itself as each app is started or stopped through the interface.
@@ -0,0 +1,4 @@
=== SCENARIO G — open the settings page and Save, changing NOTHING ===
-- stored BEFORE --
/var/lib/felhom/docker/volumes/felhom-controller-data/_data/data/settings.json
/var/lib/docker/volumes/felhom-controller-data/_data/data/settings.json
@@ -0,0 +1,46 @@
=== SCENARIO G — open the settings page and Save, changing NOTHING ===
-- stored BEFORE (sha256 + value) --
enabled_events: null
sha256 : 74234e98afe7498fb5daf1f36ac2d78a
email set : False
-- RENDER the page, take exactly what it ticked, POST it back unchanged --
render=200
ticked boxes : 10
save=200
-- stored AFTER --
enabled_events: null
sha256 : 74234e98afe7498fb5daf1f36ac2d78a
=== SCENARIO G — REDONE ===
First attempt was INCONCLUSIVE and is reported: demo-hp has NO notification prefs, so the
page rendered the DEFAULTS ticked and the save was refused by the empty-email wipe guard.
Stored bytes were unchanged — but for the wrong reason. Seeding prefs makes it a real round trip.
seed=200
BEFORE: ["backup_failed", "db_dump_failed", "node_down", "disk_warning", "disk_critical", "expected_backup_missed", "expected_dbdump_missed"]
sha : 10840f3a95bac168f0d7c79760998138
-- now RENDER and SAVE, changing nothing --
render=200
ticked: 7 email=
save=200
AFTER : ["backup_failed", "db_dump_failed", "node_down", "disk_warning", "disk_critical", "expected_backup_missed", "expected_dbdump_missed"]
sha : 10840f3a95bac168f0d7c79760998138
=== SCENARIO G — THIRD attempt, and the first CONCLUSIVE one ===
Attempt 2 also refused: the email <input> spans THREE LINES, so a single-line grep read it as
empty and the wipe guard declined the save — while still answering 200. A warning beside a
success reads as a success; the sha matching meant NOTHING was saved, not that nothing changed.
BEFORE: ["backup_failed", "db_dump_failed", "node_down", "disk_warning", "disk_critical", "expected_backup_missed", "expected_dbdump_missed"]
sha : 10840f3a95bac168f0d7c79760998138
render=200
email read : [drill@felhom.eu] ticked: 7
HTTP=200
refusal banner present? : 0 (0 = a REAL save)
AFTER : ["backup_failed", "db_dump_failed", "node_down", "disk_warning", "disk_critical", "expected_backup_missed", "expected_dbdump_missed"]
sha : 10840f3a95bac168f0d7c79760998138
-- restoring demo-hp's notification settings to their pre-drill state (no email, no events) --
restore=200
enabled_events: null
email set : False
@@ -0,0 +1,6 @@
=== SCENARIO H — the hub receives an unknown severity ===
NOTE: the LIVE hub is still v0.106.0 (the manifest bump is not synced yet), so this run is the
BEFORE picture. It is repeated after the hub deploy.
-- hub version now --
gitea.dooplex.hu/admin/felhom-hub:0.106.0
@@ -0,0 +1,11 @@
=== GOLDEN 0.223.0 — acceptance markers, each counted ===
docker OK (overlay2 : 1
including mount point : 2
upload OK (HTTP 201) : 1
excluding (must be 0) : 0
FATAL (must be 0) : 0
docker OK (overlay2; data-root /var/lib/docker)
GOLDEN_VERSION=0.223.0
GOLDEN_SHA256=9eaf39ac39219b42ec9e6cbf890275febcdcc6f53325fe0c0f591d3431044f17
-- round trip from the published URL --
ranged GET http=206
@@ -0,0 +1,9 @@
=== SCENARIO H — LIVE hub v0.107.0, bad severity ===
api key length: 64
POST /event (severity=warn) http=200
POST /event (severity=error, control) http=200
-- THE NEW LINE in the hub's own log --
2026/08/23 11:59:07 [WARN] [api] Event from demo-hp: severity "warn" is not in {info,warning,error,critical} — coercing to "info", which severityNotifies DROPS, so this backup_failed alert will reach NOBODY. Fix the emitting controller; this event is stored but not routed.
2026/08/23 11:59:07 [INFO] Event from demo-hp: backup_failed (info) — R-387 scenario H probe
2026/08/23 11:59:07 [INFO] Event from demo-hp: backup_failed (error) — R-387 scenario H control
@@ -0,0 +1,30 @@
=== END STATE — every app healthy, planted data untouched ===
2026-08-23T09:59:32Z
controller: gitea.dooplex.hu/admin/felhom-controller:0.223.0 Up 18 minutes (healthy)
bookstack Up 7 hours (healthy)
bookstack-db Up 28 minutes (healthy)
calibre-web Up 7 hours (healthy)
docmost Up 4 hours (healthy)
docmost-postgres Up 4 hours (healthy)
docmost-redis Up 4 hours (healthy)
filebrowser Up 42 hours (healthy)
kimai Up 7 hours (healthy)
kimai-db Up 7 hours (healthy)
opengist Up 7 hours (healthy)
paperless-postgres Up 7 hours (healthy)
paperless-redis Up 7 hours (healthy)
paperless-webserver Up 7 hours (healthy)
privatebin Up 7 minutes (healthy)
romm Up 7 hours (healthy)
romm-db Up 7 hours (healthy)
romm-redis Up 7 hours (healthy)
traefik Up 42 hours
-- intents restored --
privatebin desired_state: running
bookstack desired_state: running
docmost desired_state: running
/root/privatebin-app.yaml.bak
(backup file still present — removing)
notification prefs: null email_set= False
@@ -0,0 +1,8 @@
### RED-PROOF 1 (R-329) — mutation: app_start_failed severity back to "warn" ###
### layer: the EMITTER — the last point at which the bad value still exists ###
--- FAIL: TestR329_EveryEmittedSeverityIsInTheHubVocabulary (0.16s)
r329_severity_contract_test.go:203: /mnt/5_hdd/felhom.eu/git/felhom-controller/controller/internal/notify/notifier.go:561:30: emit(...) emits severity "warn", which is NOT in the hub's vocabulary {info, warning, error, critical}.
r329_severity_contract_test.go:138: checked 30 severity literals across the controller
FAIL
FAIL gitea.dooplex.hu/admin/felhom-controller/internal/notify 0.163s
FAIL
@@ -0,0 +1,11 @@
### RED-PROOF 2 (R-386) — mutation: userStopped reverted to the state guess ###
### layer: classifyRunStates — the single derivation point where the guess was made ###
--- FAIL: TestR386_OutOfBandStopWithRunningIntentAlarms (0.00s)
r386_intent_test.go:51: dead-app banner = [], want exactly privatebin — nobody asked for this app to be stopped, so its being stopped is a fault (measured silent on demo-hp 2026-08-23)
--- FAIL: TestR386_AbsentIntentSuppressesButIsAnnounced (0.00s)
r386_intent_test.go:94: IntentUnknown = false — the suppression happened but nothing records it, so an operator cannot answer 'how many apps am I blind to?'. A rule without a mechanism is not a rule
--- FAIL: TestR386_TheSchedulerWiresBothHalves (0.01s)
r386_intent_test.go:243: stacks.DesiredStateOf is never called from main.go — the classifier is back to guessing from the state (R-386)
FAIL
FAIL gitea.dooplex.hu/admin/felhom-controller/cmd/controller 0.023s
FAIL
@@ -0,0 +1,8 @@
### RED-PROOF 3 (Part 4) — mutation: sameEventSet no-op guard removed ###
### layer: the SAVE handler — where a render-then-save would rewrite stored bytes ###
--- FAIL: TestR329Part4_RoundTripIsByteIdentical (0.04s)
--- FAIL: TestR329Part4_RoundTripIsByteIdentical/defaults (0.01s)
r329_toggle_split_test.go:115: a no-op save CHANGED the stored settings.
FAIL
FAIL gitea.dooplex.hu/admin/felhom-controller/internal/web 0.048s
FAIL
@@ -0,0 +1,7 @@
### RED-PROOF 4 (R-387) — mutation: the ingest WARN line removed ###
### layer: INGEST — the last point at which the offending value still exists ###
--- FAIL: TestR387_UnknownSeverityIsCoercedAndAnnounced (0.02s)
r387_severity_visibility_test.go:89: the hub rewrote a severity and said nothing — this is R-387, and it is how two features shipped undeliverable for months. log="[INFO] Event from c1: backup_failed (info) — test message\n"
FAIL
FAIL gitea.dooplex.hu/admin/felhom-hub/internal/api 0.143s
FAIL
@@ -0,0 +1,11 @@
### RED-PROOF 5 (R-329, fillwatch) — REDONE ###
### first attempt was INERT: an earlier 'if next == prev { continue }' makes 'next < prev'
### and 'next <= prev' behave identically. Reported as a finding. ###
### mutation: the de-escalation guard REMOVED, so BandOK reaches the notify seam ###
### layer: Check() — the invariant lives THERE, not in Severity() ###
--- FAIL: TestR329_FillwatchNeverEmitsTheEmptySeverity (0.00s)
r329_severity_test.go:62: notified with band ok → severity "" (event type ""): the hub coerces an unknown severity to "info" and then drops it, so this alert would reach NOBODY
r329_severity_test.go:67: notified with severity "", outside the hub vocabulary
r329_severity_test.go:70: 4 crossings notified, every one with a routable severity
FAIL
FAIL gitea.dooplex.hu/admin/felhom-controller/internal/fillwatch 0.004s
+2
View File
@@ -162,3 +162,5 @@
| **R-107** | **No offsite action unpacks the named-volume tars Tier-3 captures on every run.** Shipped in v0.218.0. | **CLOSED — migrated from ROADMAP 2026-08-22** | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` |
| **R-383** | **The double-failure message told the customer their previous state was saved, and named a file that was not there.** Shipped in controller v0.222.0. Evidence: `audits/DRILL-r384-dead-db-alarm-2026-08-23/`. **Reasoning kept:** *One of the two ways a rollback fails is that the undo copy is missing — so the sentence was most likely to be false in exactly the case it was printed.* **Do NOT simply drop the filename:** an operator needs it, and R-351's lesson is that a refusal naming nothing forces someone to remember what the product already knows — so the absent case still names WHERE the file should have been. **A zero-length dump counts as MISSING**, because a 0-byte file restores nothing and calling it present is the same false reassurance one step smaller. The check is `os.Stat` and deliberately not an integrity test: this runs at the end of a failed restore on a machine that may be unwell, and presence is the honest claim available there. | **CLOSED — SHIPPED** (controller v0.222.0, 2026-08-23; `undoCopyPhrase`, four cases, plus an AST seam test that the message is still wired to the builder) | full text: `git show 1eb64bec5183:documentation/backlog/OPEN-ITEMS.md` |
| **R-384** | **An app whose DATABASE had died raised no dead-app alarm — the wrong question answered first.** Shipped in controller v0.222.0. Evidence: `audits/DRILL-r384-dead-db-alarm-2026-08-23/`. **Reasoning kept:** *The defect was the ORDER of two questions, not the `unhealthy` exclusion.* "Is a SUPERVISED member dead?" and "is a RUNNING member failing its healthcheck?" are different questions, and the second was answering the first — a dying database drags its own front end `unhealthy`, so the symptom the fault causes was what suppressed the alarm for it. **`IsDownState` is byte-identical and `unhealthy` stays excluded** — an unhealthy container is RUNNING, and folding it in reintroduces the flapping that exclusion exists to stop; **no new state was minted**, `StateDegraded` already means this. **Two things had to move and either alone leaves the defect standing:** the hoist, AND widening "some members are up" from `running > 0` to *any member not in the down bucket* — the old guard made the R-51 block unreachable in precisely the case it was written for. **The register's own suggested fix was WRONG and is recorded as such:** it proposed a sustained-`unhealthy` threshold on the `crashLoopAfter` model; the actual defect needed no threshold at all. **PROVEN LIVE the only way it can be** — the same fixture that printed `0 currently down` on 2026-08-22 printed **`1 currently down`** on 2026-08-23, with `app_start_failed` 7 s after the stop and the banner reading *„…nem fut: BookStack (degraded)"*. Scenario D measured **0 alarms across 9 scans** through a full stop→start cycle. | **CLOSED — SHIPPED + PROVEN-LIVE** (controller v0.222.0, 2026-08-23) | full text: `git show 1eb64bec5183:documentation/backlog/OPEN-ITEMS.md` |
| **R-329** | **`app_start_failed` was emitted with severity `"warn"`, so every one of them was delivered to nobody.** Shipped in controller v0.223.0 (+ hub v0.107.0). Evidence: `audits/DRILL-r329-r386-2026-08-23/`. **Reasoning kept:** *The vocabulary is EXACT and it is the HUB's, not ours* — `{info, warning, error, critical}`; anything else is coerced to `info` at ingest and dropped by `severityNotifies` before BOTH legs. **This was the SECOND occurrence** (`DiskAlertKind.Severity` until v0.215.0), and its comment had recorded the lesson — **a comment is not a guard**, so the guard is now an AST walk over the whole controller, with the six variable-passing call sites registered by name because a walk cannot follow a variable and *an unlisted limit is not a limit, it is a hole*. **The register's own framing was that the DECISION was the work** — should a stopped app mail the customer at all? Answered: **operator always, customer OFF by default**, because `processOperator` never consults customer preferences, so one word fixed the operator leg and left the customer leg exactly where the ruling wanted it. **Deliberately NOT added to `operatorOnlyEvents`** — that would make the new toggle visible, flickable and structurally incapable of delivering. **Measured on the live hub DB: 91 events stored all-time, ZERO notification rows before the fix; one operator row, `warning`/`sent`, after it.** | **CLOSED — SHIPPED + PROVEN-LIVE** (controller v0.223.0 + hub v0.107.0, 2026-08-23) | full text: `git show 68a9f5475cd2:documentation/backlog/OPEN-ITEMS.md` |
| **R-386** | **A single-container app stopped out of band raised no alarm, and a comment stated the opposite as settled fact.** Shipped in controller v0.223.0. Evidence: `audits/DRILL-r329-r386-2026-08-23/`. **Reasoning kept:** *the state test was guessing at something the product already knows.* `DesiredState` records the customer's intent, has **exactly one writer**, and is tri-state; `StateExited` never survives aggregation, so no state test can separate an out-of-band stop from a customer stop. The ruling: `Stopped` → no alarm, `Running` → **alarm**, **absent → UNKNOWN, keep today's behaviour AND announce it**. *Reading unknown as "nobody asked" would, on the first cycle after upgrade, e-mail about every app any owner ever deliberately stopped — fleet-wide, from a field that predates the intent it is being asked about.* **A rule without a mechanism is a wish:** every such suppression sets `IntentUnknown` and the names are logged at INFO, so an operator can answer *"how many apps am I blind to?"*. **`failedRestart` must still lift a `Stopped` intent or F-CRIT-1 re-opens.** **Fenced act: adding a `DesiredState` WRITER** — twelve of `StopStack`'s fourteen callers are machines. Proven live: alarm 24 s after an out-of-band `docker compose stop`, heartbeat `1 currently down` against the previous day's `0`; and with intent removed, suppressed *plus* the log line naming the app. **0 of 8 deployed apps on `demo-hp` carry an absent intent.** | **CLOSED — SHIPPED + PROVEN-LIVE** (controller v0.223.0, 2026-08-23) | full text: `git show 68a9f5475cd2:documentation/backlog/OPEN-ITEMS.md` |
+2 -2
View File
@@ -137,7 +137,8 @@ the fault was real. Full observables: `tests/campaign11-evidence-2026-08-05/jour
| ID | What | State |
|---|---|---|
| **R-385** | **A controller was built, baked AND vouched with no CHANGELOG entry of its own, and every gate stayed green.** Controller **0.221.1** shipped on 2026-08-23 while the newest heading in `felhom-controller/CHANGELOG.md` still read `v0.221.0` — the prune-ordering fix (commit `810b18a`) had been written INSIDE the v0.221.0 entry instead of getting its own. The image was never in question; the RECORD was, and the fleet ran a version the record did not name. **`scripts/golden_currency_gate.py` could not catch it by construction:** it failed only on `released > baked`, so a golden AHEAD of the record passed silently. Measured on the real history: `newest released 0.221.0 / newest golden baked 0.221.1 → OK, exit 0`. | **CLOSED — 2026-08-23** | — | **Both halves fixed, both directions red-proofed.** The record: `v0.221.1` has its own heading carrying the MOVED (not duplicated, not deleted) reasoning — commit `da75603`, pushed alone before anything else. The gate now asks *"is the baked version WRITTEN DOWN?"* — the baked version must have its own `## vX.Y.Z` heading **anywhere** in the CHANGELOG. **Membership, not `baked > released`, deliberately:** a comparison against the newest heading alone goes green the moment any later entry is written, leaving the unrecorded version permanently unrecorded and the gate permanently silent about it. INCONCLUSIVE (exit 2) preserved. Evidence: `audits/DRILL-r384-dead-db-alarm-2026-08-23/evidence/gate-0*.txt` — old gate/old record `exit 0`, new gate/old record `exit 1`, new gate/fixed record `exit 0`. | CC |
| **R-386** | **A single-container app stopped OUT OF BAND raises no alarm at all — and a comment states the opposite as settled fact.** `aggregateState` folds `StateExited` into the `stopped` counter, so an all-down stack returns `StateStopped` and **`StateExited` never survives aggregation**. `classifyRunStates` then whitelists `StateStopped` as a deliberate user stop unless the quiesce loop reports a failed restart. The comment at `cmd/controller/main.go` says: *"(An out-of-band `docker compose stop` leaves the containers present → StateExited → still alerts, which is correct: out-of-band tampering IS reportable.)"* — **measured FALSE.** The neighbouring I2 claim (*"a CRASHING app never comes to rest at `stopped` — faults surface as StateExited"*) is false in the same way. **Measured live on `demo-hp` 2026-08-23 (controller v0.222.0):** `privatebin` (1 container, `unless-stopped`) stopped out of band at 05:47:35Z; at 05:51:53Z it read `state=stopped`, **9 dead-app scans had run, and there were ZERO `app_start_failed` events and ZERO banner lines.** **The absence is trustworthy — positive control from the same box 17 minutes earlier:** `app_start_failed` fired for BookStack at 05:30:14Z, so the detector demonstrably works there. **SCOPE, stated so it is not overclaimed:** a genuine crash under `unless-stopped` is RESTARTED by Docker and surfaces as `restarting` → the 5-minute crash-loop path, which does alarm. The silent case is an explicit out-of-band stop of a stack with no surviving member. **This is case #10 of "a comment asserting an invariant the code does not provide".** Found by §4 of the R-384 task, which asked for a measurement and explicitly forbade a fix in that session. | **OPEN — MEDIUM** | — | Decide whether an out-of-band stop is distinguishable from a customer stop at all — they are byte-identical on the Docker side, exactly as invariant I1 says, so the answer is probably NOT a state test but a recorded intent (`DesiredStateOf` already exists and `bootrecon` already consumes it). **Do NOT simply un-whitelist `StateStopped`** — that re-alarms every genuine customer stop, which is the over-correction F-CRIT-1's fix was careful to avoid. **And fix the comment either way:** it is load-bearing and it is false. Evidence: `audits/DRILL-r384-dead-db-alarm-2026-08-23/evidence/live-17-sec4-stop.txt`, `live-18-sec4-verdict.txt`, `live-19-scenarioD-sec4-full-log.txt`. | CC |
| **R-387** | **The hub REWRITES an unknown severity and says nothing, and the guard built to catch that sits downstream of the rewrite.** One handler, two fields, opposite discipline: an unknown `event_type` is rejected with a loud `400`, while an unknown `severity` was silently coerced to `info` — after which `severityNotifies` drops it and NEITHER delivery leg runs. **Two shipped features went out that way**: `DiskAlertKind.Severity` emitted `"warn"` until controller v0.215.0, `app_start_failed` until v0.223.0. **Measured on the live hub DB 2026-08-23: 91 `app_start_failed` events stored all-time and ZERO `notification_log` rows before that day** — not one, on any channel, while every POST returned 200. **The dispatcher's `unrecognized severity` line could never execute** for an API event, because the coercion one line upstream guarantees the value it looks for cannot arrive. | **CLOSED — hub v0.107.0, 2026-08-23** | — | **The coercion STAYS; only the silence is fixed** — a rejected event is a LOST event, and losing an alarm is worse than mis-routing one. A `WARN` now names the customer, the event type, the rejected value and the consequence. **The dispatcher branch was KEPT, on evidence not caution:** `cmd/hub/main.go` wires `dispatcher.ProcessEvent` DIRECTLY as the `monitor.EventNotifyFunc` for the staleness, host-staleness and offsite-box checkers, which never pass through the handler — for them it is the only severity guard there is; deleting it as "dead" would have removed the live half while the dead half supplied the justification. All 90 severity literals in `internal/monitor` verified already valid. Proven live: `[WARN] [api] Event from demo-hp: severity "warn" is not in {info,warning,error,critical}…`, with an `error` control silent. Evidence: `audits/DRILL-r329-r386-2026-08-23/evidence/live-19-scenarioH-after.txt`. | CC |
| **R-388** | **PRODUCT DECISION (not a defect): the customer notification model is the wrong shape, and the settings page grows by one toggle per detector.** The operator's framing, recorded verbatim 2026-08-23: *"A customer should be notified only about things they can act on or are responsible for — the drive they unplugged, the storage they filled. **A failed backup is our incident, not theirs.** The intended shape is that we detect it, we tell them we noticed and are dealing with it, and they are not handed an error they cannot solve. The subscription should feel like being looked after, not like being on call."* Today's page is the opposite shape — one switch per detector, and it **grew from 12 to 15 in a single session** (one new alarm plus two compound toggles split into four). That growth is the argument, not an aside: a page that grows per detector keeps asking a household to make engineering decisions. | **OPEN — DIRECTION, operator's call** | a decision on scope; nothing here is a bug | Recorded as a dated **[DESIGN — DIRECTION]** entry at `documentation/architecture/08-alarm-ladder.md` §8, marked plainly as *not current behaviour*. **Deliberately NOT implemented in the session that recorded it.** `app_start_failed` defaulting OFF is consistent with the direction and reversible either way, but was ruled on its own merits and does not pre-judge the redesign. | Viktor |
| **R-229** | **The instruction-file rightsizing landed for `felhom-controller` and the workspace root; three pieces were deliberately deferred.** Done 2026-08-06: controller split into a 92-effective-line core plus four `paths:`-scoped `.claude/rules/*.md`; workspace root 208→142 effective lines with its versioned copy kept byte-identical; surgical corrections to `felhom-agent` and `felhom.eu` (expired TEMPORARY block, every version literal, the Legacy-Windows copies, the duplicated health-check rule); five contradictions resolved — including a drill-VM claim **measured live** (`qm list` on demo-hp shows VM 300 `drill-r50`; `felhom-agent` was right, `felhom-controller` was wrong); new shared `felhom.eu/scripts/instructions_gate.py` registered in `controller_gates.py` and `agent_gates.py`, 20 fixture tests + red-proof. **Leg (a) CLOSED 2026-08-06 (part 2):** `felhom.eu/CLAUDE.md` **227 → 115 effective lines**, split into a core plus `.claude/rules/{hub,website,manifests,docs}.md`; `instructions_gate` **registered in `scripts/repo_gates.py`** (six gates, all OK) in the required order — trim first, register second, because a registered-but-failing gate refuses every push. Scoping proven from the `InstructionsLoaded` hook log in two fresh sessions, not from frontmatter. **Still deferred:** (b) **CLOSED 2026-08-06 (close-out)** — `felhom-agent/CLAUDE.md` **175 → 99 effective lines** (measured 175, not 173: the CI correction added two), split into a core plus `.claude/rules/{proxmox,localapi,backup,storage}.md` beside the existing `health-checks.md`. The release section now points at the `felhom-build-deploy` skill instead of restating a table that drifts from the script. **Every `CLAUDE.md` in the workspace is now ≤120 effective lines except the workspace root at 142, which is deliberate — it is the only file re-injected after `/compact`.** (c) **CLOSED 2026-08-06 (part 2)** — all 44 orphans resolved with **zero deletions** (file count 158 before and after): 4 durable `reference`-type files indexed, 40 dated episode records moved to `.claude-memory/archive/`. `MEMORY.md` 145 → **150 lines / 17,977 bytes**, and `instructions_gate` check 6 now watches it (over-limit FAILS, orphan WARNS, absent store PASSES *printing its reason*). (d) **The spec-as-failing-test pilot** — moved to R-230. Full accounting: `audits/LEDGER-instruction-trim-2026-08-06.md` + `audits/LEDGER-instruction-trim-part2-2026-08-06.md` | **READY** — owner Viktor |
| **R-230** | **Three instruction/memory follow-ups deliberately left by the part-2 session (2026-08-06), each needing a decision rather than an implementation.** (a) **A ruling is owed on auto-written staleness.** The hand-written `CLAUDE.md` files are now clean of version literals and expired blocks — the gate enforces it — but `MEMORY.md`, which Claude writes and which is the LARGER half of what loads (8.4k tokens vs the root file's 6.6k), carries **21 lines with component version literals**, **5 with bare host addresses**, and an entry still reading *"demo boxes REMOTE till ~08-02"* — the same expired-TEMPORARY class the gate was built to kill, now surviving in the one file the gate's content rules do not cover. **Partly actioned 2026-08-06 (close-out), and the ruling is STILL OWED:** the **three statements that were actively false** were corrected — `R-193 decision open` (closed 2026-08-05), `demo boxes REMOTE till ~08-02` (the box answers on the home LAN), `OPEN R-25b` (shipped 2026-07-21) — and gate check 6 now **WARNs** on version literals, host addresses, expired statements and stale-open citations in the index. WARN, never FAIL: Claude writes that file between sessions, so a hard failure would refuse a human's push over a line no human typed, and the warning is read by the model that will next edit it. **The remaining 32 version literals and 4 host addresses were deliberately left** for that loop. What is still owed is the bulk-correction ruling. **Correcting the premise:** the earlier report's "three expired statements" were all FALSE POSITIVES — each matched an ISO date inside a markdown link target, i.e. a filename — while the one real expired claim carried no ISO date at all. (b) **CLOSED 2026-08-06 (close-out)** — the workspace-root `CLAUDE.md` **is now a relative symlink** to the versioned copy, so the divergence class is gone rather than policed. Check 5 learned two shapes: for a link it asserts the target resolves to a real file (**a dangling link is worse than a diverged copy — the instructions load NOTHING and there is no content left to notice is wrong**), for two files byte-identity as before, so a clone elsewhere is unaffected. **Proven, not assumed:** three fresh sessions logged `session_start` for the link path, and a fourth **with no tools at all** quoted standing rule 1 verbatim — the content reaches the model, not just the path. (c) **The spec-as-failing-test pilot**, approved in principle and not started (was R-229(d)). | **READY** — owner Viktor |
| **R-232** | **DooPlex's backup makes every copy inside the same box — and nothing tells anyone when it fails.** Surveyed read-only 2026-08-06 (`audits/RECON-dooplex-backup-2026-08-06.md`). **What works:** five sets, 14/14 successful runs in 14 days; a file was restored from the `data` repo and matched the live original **byte for byte**; every set except two is cross-disk; k3s is integrity-checked on every run. **What the matrix exposes, ranked:** (a) **`notify_failure` is a no-op** — `NOTIFY_ON_FAILURE=true` but `NOTIFY_WEBHOOK_URL` is commented out, so a failed backup notifies **nobody**; the project already has a working Resend path that CI uses. Cheapest item, and it makes every other failure visible. (b) **Nothing leaves the box** — no rclone, no remote repo, no off-site target anywhere; Longhorn's target is `nfs://192.168.0.180:` pointing at DooPlex itself, and the only outbound-looking cron pulls *inbound* from Hetzner for a different project. The machine that runs the hub managing the customers' off-site chain has no off-site copy of its own. (c) **The backup tree is a single writable path** and the restic repos are not append-only — one bad script or ransomware destroys every copy at once. (d) **Two same-disk sets**: `.claude-memory` and the PostgreSQL dumps, whose source directory sits *inside* the backup tree. (e) **Longhorn `retain=1`** — one generation per volume, so a corruption noticed a day late has no earlier copy. (f) **`/opt/backup/docs/BACKUP-RESTORE.md` does not exist** though the systemd unit advertises it. (g) **`secrets/restic-repo` has never held a snapshot** — `backup-secrets.sh` contains no `restic` call; the secrets are GPG files on `sda1` only. (h) **No restore has ever been run** beyond today's single-file probe — the matrix's "ever demonstrated?" column is otherwise entirely empty. **Not a finding:** the restic passphrase. The on-box copy is on `sdb1`, a different disk from the backups, and the **operator holds an offline copy out of band** — so a disk loss is recoverable. The narrow residual is that it is operator-held rather than system-held, unlike the customer case's hub-vaulted escrow, so it should be confirmed current and findable by someone else. **Nothing was changed by the recon.** | **READY** — owner Viktor |
@@ -509,7 +510,6 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server`
| **R-317** | **The agent decides whether to install dnsmasq by stat-ing a file the OTHER package owns.** `EnsureDnsmasq` (`felhom-agent/internal/lanresolver/lanresolver.go:105`) does `os.Stat("/usr/sbin/dnsmasq")` and skips the apt install when it exists — but that path is shipped by **`dnsmasq-base`**, while the systemd unit comes from **`dnsmasq`** (confirmed on the box: `dpkg -S /usr/sbin/dnsmasq` → `dnsmasq-base`; `dpkg -S /usr/lib/systemd/system/dnsmasq.service` → `dnsmasq`). So on any host carrying `dnsmasq-base` without `dnsmasq`, the agent skips the install and then runs `systemctl enable --now dnsmasq` against a unit that is not there: the resolver never comes up and the failure is a retried WARN in the journal rather than anything a customer or the install sees. **Pre-existing, NOT introduced by R-316** — but R-316 makes the shape reachable, because a host whose `dnsmasq-base` pre-dated Felhom now keeps it while `dnsmasq` is removed. R-316's uninstall says so explicitly instead of leaving it to be found from a silent resolver. **Ranked 2 (costs time), not 1:** the box installs fine, only LAN name resolution is missing | **READY (S) — NEW 2026-08-13** | R-316 | Probe what is actually needed — the unit or the `dnsmasq` package — rather than a path a sibling package owns. One-line change in the agent; deliberately NOT made here to keep this session to one repo | CC |
| **R-325** | **The shared copy vocabulary is imported by ONE of its two consumers, and drift-checked into the other.** `customer_copy_vocab.py` is the single list; `hub_copy_gate.py` imports it. **`felhom-controller/controller/scripts/retrieval_promise_gate.py` still carries its own `STEMS` literal**, because the session that created the shared module was under a hard end-state requirement to leave `felhom-controller` untouched — its target box was being re-deployed the same evening. **Two copies of a word list is not a theoretical risk in this project: it is the R-299 defect exactly**, where a guard asserted one inflection of a Hungarian verb and the plural walked past it. **So the gap is instrumented rather than left open: `hub_copy_gate.py` READS the controller gate's `STEMS` and FAILS if the two disagree** — single-source semantics tonight without a cross-repo edit. **Watched failing:** removing one stem from the shared list produced *"the shared vocabulary is no longer shared"* with both lists printed, and restoring it returned the gate to green. An ABSENT sibling clone is **INCONCLUSIVE (exit 2), never a pass** — the G-1 lesson. **This is a scaffold, not the destination** | **READY (S) — NEW 2026-08-13, RANK 3** | R-299, R-324 | Make `retrieval_promise_gate.py` import `felhom.eu/scripts/customer_copy_vocab.py` and delete its own literal — a felhom-controller change of a few lines, needing no bake (a gate is not shipped code). Then the drift check becomes redundant and should be removed with it, rather than left as a second mechanism nobody re-reads | CC |
| **R-327** | **The standing picture still describes a defect that has been fixed twice over.** Found by the first run of `unproven.py` (R-326), which is the argument for having built it. `where-felhom-stands.yaml`'s `claim.code-naming` is `status: partial` and its title reads *"The same word is used for two different secrets across three surfaces; the email points at a page a rebuilt machine does not show"* — **both halves of which are now false.** The box side shipped 2026-08-10 (R-295), the hub half and the page-naming fix on 2026-08-13 (R-295 hub, new `reenroll` mail kind), and the third near-homograph on 2026-08-13 (R-323). **NOT MOVED BY THIS SESSION, deliberately and by the dataset's own rule:** *"A status may not be RAISED here — if the evidence supports a stronger status than the capability map records, the MAP changes first and this file follows it."* Raising it here would be the exact inversion the file's header forbids, and the map edit is a separate judgement about what "walked" means for a naming change that no customer has yet met | **READY (S) — NEW 2026-08-13, RANK 4** | R-295, R-323, R-326 | Decide the capability-map status for the naming arc, then let the dataset follow it. **Note the honest difficulty: no customer has typed „Tulajdonosi jelmondat” yet**, so `walked` would be an over-claim; `built` is probably right, and the title needs rewriting either way because it describes a defect rather than a capability | operator + CC |
| **R-329** | **`app_start_failed` carries the IDENTICAL defect and was deliberately left alone.** `notifier.go` ~L546 emits severity `"warn"` for *„Telepített alkalmazás nem fut: %s"* — the same string that made R-328 undeliverable, so this event is also stored as `info` and emailed to nobody. It was NOT changed while fixing R-328 because it needs a decision first: **should a stopped app email the customer at all?** Flipping the string without answering that turns a silent event into a mail flood on a box where an app crash-loops. **Whoever changes it must check the hub side too** — a `customerMessages` entry and `DefaultEnabledEvents` membership decide who hears it, and the R-158/R-167 lesson is that routing a can't-act-on-it failure to a customer-enabled type is its own defect | **READY (XS code, the DECISION is the work) — NEW 2026-08-14** | a decision on whether `app_start_failed` should notify, and on which leg | Decide operator-only vs customer; then set the severity and the hub routing to match | Viktor |
| **R-330** | **Disk health Phase 2 — the three SMART attributes the wire does not carry.** The failing drive's most telling counter was **187 `Reported_Uncorrect`**, sitting at normalized **1** against threshold **0** with a raw count of **1001** — one point from failing and structurally unable to get there. Also wanted: **199 `UDMA_CRC_Error_Count`** (cabling) and **188 `Command_Timeout`**. None are on the agent→controller wire today, so v0.215.0's ladder could not use them. Phase 2 also persists periodic SMART **samples** (the right home is `metrics.MetricsStore`, NOT the Phase-1 state file, which is one record per disk and must stay that way). **This is a declared WIRE change, so under the G-1 gate the hub must model the new fields in the SAME session** — that is precisely why it was kept out of Phase 1, where it would have turned a one-word severity fix into a three-repo change | **READY (M) — NEW 2026-08-14** | R-328 (closed) | Add 187/199/188 to the agent's `SmartSummary` + hub model in one session; then persist samples | CC |
| **R-331** | **Disk health Phase 3 — growth-rate detection, and retiring the static 64.** The v0.215.0 count backstop (64 unreadable sectors → Hiba) is **a judgement from ONE drive**: the observed benign excursion peaked at 16 and cleared inside an hour, and the terminal run passed 64 at 13 Aug 11:28 and never came back. It is deliberately a backstop BEHIND the sustain rule, not the primary signal, but it is still a magic number tuned on a single sample and it will be wrong for some drive. With Phase 2's history the box can ask the question that actually matters — *is this count climbing, and how fast* — which distinguishes a drive with eight stable aging sectors from one adding forty a day, something no static threshold can do. Revisit 64 when that exists | **READY (M) — NEW 2026-08-14** | R-330 | Growth-rate rule over persisted samples; re-derive or delete the static 64 | CC |
| **R-332** | **The new Hiba-from-counters path has never fired on real hardware.** v0.215.0's whole point is a verdict the product could not previously reach, and it is proven only against the committed fixture's values in unit tests (12 scenario groups, 11 of 12 red-proofs failing as required). The live validation on demo-hp proved the **negative** — three healthy disks still read Rendben across the deploy, no false alert — and the **severity wire** end to end, but no live disk has actually reached Hiba. **This is the honest gap and it must not be closed by pointing at the fixture tests**: the drive that produced the fixture is in DooPlex, which is Tier 2 and never a drill target, and the demo boxes are all-flash and healthy | **WATCHING — NEW 2026-08-14, NARROWED same day.** One item originally in this gap is now PROVEN LIVE: the **persisted state surviving a controller restart**. The v0.215.0→v0.216.0 redeploy destroyed and rebuilt the container, and the new one read back a `changed_at` written by the PREVIOUS version (`2026-08-14T07:23:14.640216851Z`, still intact at 09:31:35Z) instead of re-baselining — Scenario L on real hardware, not just the production-path unit test. **What remains unproven is the verdict itself, plus the stronger restart half: an already-ALERTED disk not re-alerting** | a real degrading disk, or an injection harness | **Closing condition:** a live disk reaching Hiba from counters, OR a deliberate injection through the REAL pipeline (agent `/disks` → controller check → hub event), not a hand-set verdict | CC |