docs(hub v0.108.0): the delivery grain, the cooldown ruling, and gate 11's first subject
gates / gates (push) Successful in 15s
gates / gates (push) Successful in 15s
The alarm ladder gains §6.2 - which events are per-app, per-run, per-tier or coarse, and why the default is coarse. CONTEXT records two rulings: the grain is allow-listed rather than inferred from the payload, with crossdrive_failed as the proof that a payload rule would have been wrong; and a finding recorded only in REPORT.md has a lifetime of one session. R-389 closed and compressed, keeping its rules and naming the commit whose git show returns the full text. R-390 and R-391 left open. REPORT.md is gate 11's first real subject and passes: six observations, two FILED, four NOT-A-FINDING with their reasons. Three of those declarations are things a tidier report would have omitted - the gate's own spec would have passed the item it was built to catch, the burst has no ceiling, and ArgoCD said "successfully rolled out" while still running the old image. STATUS carries forward the one thing outstanding: the controller floor still reads 0.222.0 while the golden reads 0.223.0.
This commit is contained in:
+53
@@ -14,6 +14,59 @@
|
||||
> language, one screen, no identifiers in the prose. Same subjects, different readers; merging them
|
||||
> would make one of the two audiences stop reading. `STATUS.md` is also a **view of `OPEN-ITEMS.md`**
|
||||
> and holds nothing of its own; this file does hold its own content, namely the standing rulings below.
|
||||
## Cooldown GRAIN is allow-listed, never inferred from the payload (2026-08-23, R-389)
|
||||
|
||||
**[RULING] Per-app cooldown is a NAMED REGISTER (`perAppCooldownEvents`), not a rule of the form "if
|
||||
the details carry a `stack_name`, split per app".** The backup family's cooldown is coarse **on
|
||||
purpose** — R-97a and R-182 exist so one full disk sends one digest rather than one mail per app.
|
||||
|
||||
**[FACT] The payload-shape rule would have been wrong, and provably so.** `crossdrive_failed` is
|
||||
severity `error`, reaches the operator leg, and carries `stack_name` through **`CrossDriveDetails`**,
|
||||
a different struct from `AppDetails`. Anything keying off the field would have split it silently.
|
||||
**That is why `cooldownStackSuffix` takes the EVENT TYPE as well as the details** — an asymmetry with
|
||||
`cooldownTierSuffix`/`cooldownRunSuffix`, and the asymmetry IS the safety property: `tier` and
|
||||
`run_id` appear only on types that want that grain; `stack_name` does not have that property.
|
||||
|
||||
**[RULING] `app_start_failed` gets per-app because it has NO DIGEST.** There is no `apps_down_run`
|
||||
summarising a scan the way `backup_run_failures` summarises a run, so per-app is the only grain that
|
||||
does not lose alarms. **The hour did not change** — the grain was the complaint, not the length.
|
||||
|
||||
**[FENCE] The fenced act is adding an entry to `perAppCooldownEvents` for a type whose family has a
|
||||
digest or a deliberately coarse cooldown.** Reading it is fine.
|
||||
|
||||
**[MEASURED] Burst volume, so it is a number and not an impression:** three apps in one scan →
|
||||
**3 attempted, 3 sent, 0 suppressed**. Reference box has 8 deployed apps, so a total outage is 8
|
||||
mails. It scales linearly and has no ceiling; the reopening condition is a box big enough that a
|
||||
total outage is unreadable, and the answer there is a digest, not a wider cooldown.
|
||||
|
||||
## A finding recorded only in REPORT.md has a lifetime of ONE SESSION (2026-08-23, R-389 / gate 11)
|
||||
|
||||
**[FACT] This cost a day.** The cooldown-grain defect was measured live on 2026-08-23, written under
|
||||
`## Observations` in `REPORT.md`, and written **nowhere else**. `REPORT.md` is overwritten every
|
||||
session by this project's own convention. It had no register row and had to be re-derived the next
|
||||
day. Same shape as R-341.
|
||||
|
||||
**[FACT] The instruction invited it.** `PROMPT-TEMPLATE.md` §15 asked for observations *"documented,
|
||||
NOT acted on"*, and "documented" was satisfied by the paragraph. **The template was the defect**, and
|
||||
it was corrected in the same session that built the mechanism.
|
||||
|
||||
**[MECHANISM] Gate 11 (`scripts/observations_gate.py`).** Every numbered item in a `REPORT.md`
|
||||
observations section must carry `FILED: R-NNN` (which must resolve) or `NOT-A-FINDING: <reason>`. It
|
||||
REFUSES; it does not warn. It PASSES quietly when there is no observations section, deliberately — a
|
||||
gate that taxes every push gets disabled within a week.
|
||||
|
||||
**[RULING, and it contradicts the task that commissioned it] A bare mention of an R-number does NOT
|
||||
satisfy the gate.** The specification said an item may "cite an R-NNN that resolves" — **that rule
|
||||
would have passed the very item the gate was built to catch**, because the lost item cited `R-182` as
|
||||
an *analogy*. No parser can tell citation-as-precedent from citation-as-filing by reading prose. The
|
||||
marker is explicit for that reason, and the discrepancy is recorded in the gate's docstring rather
|
||||
than quietly resolved.
|
||||
|
||||
**[GOTCHA] Registered in three runners, not four.** `catalog_gates.py` has no shared-gate mechanism —
|
||||
`run_gate` joins against its own `scripts/` and the loop appends `--all` to every gate — so
|
||||
registering there needs its contract widened. Filed as **R-391** rather than left as a sentence,
|
||||
which is this session's whole lesson.
|
||||
|
||||
## The severity a controller sends is the HUB's vocabulary, and getting it wrong deletes the alert (2026-08-23, R-329 / R-387)
|
||||
|
||||
**[RULING] The set is exactly `{info, warning, error, critical}`.** The hub **coerces anything else to
|
||||
|
||||
@@ -1,126 +1,266 @@
|
||||
# REPORT — felhom.eu: hub v0.107.0 (R-387), the alarm ladder, and golden 0.223.0
|
||||
# REPORT — hub v0.108.0 (R-389), gate 11, and the instruction that invited the gap
|
||||
|
||||
**Session 2026-08-23.** Companion to `felhom-controller` v0.223.0 (R-329, R-386) — see that repo's
|
||||
`REPORT.md` for the controller work and the full live walk.
|
||||
**Session 2026-08-23.** `felhom.eu` is the subject; the controller and agent were touched only to
|
||||
register the shared gate. **No controller release — no golden bake, no vouch, no floor.**
|
||||
**No halt condition fired.** Nothing was dropped.
|
||||
|
||||
## 1. Baselines, and the hub's four numbers as read
|
||||
|
||||
| Repo | at start | at end |
|
||||
|---|---|---|
|
||||
| felhom.eu | `55274d5e` | hub **v0.107.0** deployed |
|
||||
| felhom-controller | `14137efa` (v0.222.0) | **v0.223.0** deployed |
|
||||
| felhom-agent | `40d857b5` | untouched |
|
||||
| felhom.eu | `2f7c9a6` (hub v0.107.0) | **hub v0.108.0** deployed |
|
||||
| felhom-controller | `1da2c9c` (v0.223.0) | **unchanged** — runner registration only |
|
||||
| felhom-agent | `40d857b` (v0.130.0) | **unchanged** — runner registration only |
|
||||
|
||||
**Hub's four numbers, read live from `GET /configuration` before starting:**
|
||||
`golden_version` **0.222.0** · `agent_version` **0.130.0** · `min_agent` **0.129.0** ·
|
||||
controller floor **0.222.0**. All four as the task predicted; the operator's 0.222.0 vouch had landed.
|
||||
**Hub's four numbers, live from `GET /configuration` before starting:**
|
||||
|
||||
## 2. The hub's deployment path — §6's premise was wrong, and here it is
|
||||
|
||||
**`felhom.eu/manifests/hub.yaml`, line 128.** ArgoCD `Application/felhom` tracks
|
||||
`https://gitea.dooplex.hu/admin/felhom.eu.git`, path `manifests`, with `syncPolicy.automated.enabled
|
||||
= false`. Bumped `0.106.0 → 0.107.0` in commit **`68a9f54`**. **There is no out-of-git deployment
|
||||
path** — the finding §6 braced for does not exist. The image was built and pushed to the registry
|
||||
**before** the manifest landed, so a sync could never have pointed at a missing tag, and the sync was
|
||||
then requested deliberately (`refresh=hard`, then a patched `operation`). Never `kubectl set image`.
|
||||
|
||||
## 3. R-387 — what was wrong
|
||||
|
||||
One handler, two fields, opposite discipline: an unknown `event_type` is rejected with a loud `400`;
|
||||
an unknown `severity` was rewritten to `info` **without a word**, and `severityNotifies` drops `info`
|
||||
before *both* legs. **The guard built to catch exactly this sat downstream of the rewrite** — the
|
||||
dispatcher's `unrecognized severity` line can never execute for an API event, because the coercion one
|
||||
line earlier guarantees the value it looks for cannot arrive.
|
||||
|
||||
**Measured on the live hub DB:** `91` `app_start_failed` events stored all-time, **`0`
|
||||
`notification_log` rows before this session** — not one, on any channel, while every POST returned 200.
|
||||
|
||||
**The coercion stays.** A rejected event is a *lost* event, and losing an alarm is worse than
|
||||
mis-routing one. Only the silence is fixed.
|
||||
|
||||
## 4. The dead-branch decision, and the reason
|
||||
|
||||
**KEPT.** Not caution — evidence. `cmd/hub/main.go` wires `dispatcher.ProcessEvent` **directly** as
|
||||
the `monitor.EventNotifyFunc` for the staleness, host-staleness and offsite-box checkers, and those
|
||||
hub-generated events never pass through the ingest handler at all. For every one of them that line is
|
||||
the **only** severity guard there is. Deleting it as "dead" would have removed the live half while the
|
||||
dead half supplied the justification.
|
||||
|
||||
Verified while deciding: **all 90 severity literals in `internal/monitor` are already valid**, so the
|
||||
guard is silent because the producers are correct. (`"warn"` in `internal/web` is UI badge vocabulary,
|
||||
not a severity.)
|
||||
|
||||
## 5. Files changed, commits, CI
|
||||
|
||||
| Commit | Contents |
|
||||
| Field | Value |
|
||||
|---|---|
|
||||
| **`68a9f54`** | hub v0.107.0 (ingest WARN + kept-branch note + tests), manifest bump, golden 0.223.0 evidence |
|
||||
| **`<docs>`** | alarm ladder §6.1/§7/§8, register, `STATUS.md`, `REPORT.md`, drill record |
|
||||
| `golden_version` | **0.223.0** |
|
||||
| `agent_version` | **0.130.0** |
|
||||
| `min_agent` | **0.129.0** |
|
||||
| controller floor (`min_controller_version`) | **0.222.0** |
|
||||
|
||||
Files: `hub/internal/api/handler.go`, `hub/internal/notify/dispatcher.go`, `hub/CHANGELOG.md`,
|
||||
`manifests/hub.yaml`, `documentation/architecture/08-alarm-ladder.md`,
|
||||
`documentation/backlog/{OPEN,CLOSED}-ITEMS.md`, `documentation/tests/golden-0.223.0-2026-08-23/`,
|
||||
`documentation/audits/DRILL-r329-r386-2026-08-23/`, plus two new test files.
|
||||
**The task predicted the floor at 0.223.0 and it reads 0.222.0.** The operator vouched the golden but
|
||||
has not yet raised the floor — the last step of the previous release, which `STATUS.md` says to do
|
||||
"last, in its own save". Carried forward as item 1 there. Not a halt; the correct reading is simply
|
||||
different from the prediction.
|
||||
|
||||
**CI runs confirmed BY ID** (`id` and `run_number` diverge — both printed): see §5 of the controller
|
||||
REPORT for its runs; felhom.eu's are listed at the end of this file.
|
||||
**The golden-currency gate stayed green throughout** and no bake was needed, exactly as §1 said it
|
||||
should be: the controller CHANGELOG's newest entry and the newest baked golden are both 0.223.0 and
|
||||
this session moved neither.
|
||||
|
||||
## 6. Tests and red-proofs
|
||||
## 2. Documents read
|
||||
|
||||
`internal/api/r387_severity_visibility_test.go` — the event is **not lost**, the stored severity is
|
||||
still `info`, and the WARN names customer + type + value; plus a guard that a **valid** severity stays
|
||||
silent, because an alarm on the normal path is one people learn to ignore.
|
||||
`internal/notify/r329_app_start_failed_test.go` — the routing consequence: operator emailed, customer
|
||||
not, unless opted in, in which case both legs deliver and the customer's copy carries the Hungarian
|
||||
template. Scenario B is also the **positive control** for Scenario A's absence claim.
|
||||
`hub/internal/notify/dispatcher.go` (`cooldownTierSuffix` + `cooldownRunSuffix` docstrings in full,
|
||||
`processOperator` and its R-182 suppression-logging block, `operatorOnlyEvents`),
|
||||
`scripts/repo_gates.py` (whole docstring, including "WHY 10 IS HERE" and "WHY 7 IS HERE"),
|
||||
`scripts/due_checks_gate.py`, `documentation/PROMPT-TEMPLATE.md` §15, and the alarm ladder at
|
||||
**`documentation/architecture/08-alarm-ladder.md`** — extended here with §6.2, the delivery grain.
|
||||
|
||||
Test count **702 → 709**.
|
||||
## 3. The `AppDetails` emitter count, measured
|
||||
|
||||
**Red-proof (seen failing):** delete the ingest `WARN` → `the hub rewrote a severity and said
|
||||
nothing`, with the log showing only the ordinary `[INFO] Event from c1: backup_failed (info)`. The
|
||||
guard sits at **ingest**, because that is the last point at which the offending value still exists.
|
||||
**Three, exactly as §3 said.** `grep -rn "AppDetails{" --include=*.go` over the controller, excluding
|
||||
tests:
|
||||
|
||||
## 7. Golden
|
||||
| Emitter | Event | Severity | Reaches the operator leg? |
|
||||
|---|---|---|---|
|
||||
| `notifier.go:502` | `app_deployed` | `info` | no — `info` is dropped by `severityNotifies` |
|
||||
| `notifier.go:563` | `app_start_failed` | `warning` | **yes** |
|
||||
| `notifier.go:661` | `app_removed` | `info` | no |
|
||||
|
||||
**Baked and PUBLISHED: 0.223.0.** `GOLDEN_SHA256 =
|
||||
9eaf39ac39219b42ec9e6cbf890275febcdcc6f53325fe0c0f591d3431044f17`; `upload OK (HTTP 201)`; round-trip
|
||||
**HTTP 206**; all five acceptance markers counted (`docker OK (overlay2` 1, `including mount point` 2,
|
||||
`upload OK` 1, `excluding` 0, `FATAL` 0). **VOUCHING IS THE OPERATOR'S ACT AND WAS NOT DONE HERE.**
|
||||
**No fourth emitter. No halt.**
|
||||
|
||||
**Runbook deviation, second session running:** §4.1 omits `pveam update`, so the `virgin` snapshot's
|
||||
stale template index fails as `400 … no such template`.
|
||||
**But the sweep found something the `AppDetails` question could not:** `stack_name` is also carried by
|
||||
a **different struct**, `CrossDriveDetails` (`notifier.go:151-157`), used by `crossdrive_failed`
|
||||
(severity **`error`**, so it *does* reach the operator leg) and `crossdrive_completed`. This is what
|
||||
makes the allow-list load-bearing in fact rather than in principle — a payload-shape rule would have
|
||||
split a backup-family event per app and silently undone R-182. It is Scenario C's live subject.
|
||||
|
||||
## 8. Scenario H, live against v0.107.0
|
||||
## 4. Files, commits, CI
|
||||
|
||||
| Commit | Repo | Contents |
|
||||
|---|---|---|
|
||||
| **`f751aea`** | felhom.eu | R-389 filed — **alone, before any code** (Phase 1) |
|
||||
| **`2fc4a15`** | felhom.eu | the suffix + allow-list + tests, gate 11, template fix, R-390/R-391 |
|
||||
| **`45659bd`** | felhom.eu | hub v0.108.0 CHANGELOG + manifest bump |
|
||||
| **`f8c9390`** | felhom-controller | gate 11 registered; its REPORT's observations marked up |
|
||||
| **`058b945`** | felhom-agent | gate 11 registered |
|
||||
|
||||
**CI runs confirmed BY ID** — and the listing was checked for truncation rather than trusted, since a
|
||||
silently truncated listing has already cost this arc a false claim:
|
||||
|
||||
| Commit | Repo | CI `id` | `run_number` | Result |
|
||||
|---|---|---|---|---|
|
||||
| `f751aea` | felhom.eu | **412** | 263 | success |
|
||||
| `2fc4a15` | felhom.eu | **413** | 264 | success |
|
||||
| `45659bd` | felhom.eu | **416** | 265 | success |
|
||||
| `f8c9390` | felhom-controller | **414** | 90 | success |
|
||||
| `058b945` | felhom-agent | **415** | 55 | success |
|
||||
|
||||
The API reported `total_count` 265 / 91 / 55 against 3 rows shown in each case — i.e. the pages were
|
||||
known-partial and the newest rows are the ones quoted, not the whole set.
|
||||
|
||||
## 5. Red-proofs — two planted, both seen failing
|
||||
|
||||
| # | Mutation | Layer, and why that layer | Observed |
|
||||
|---|---|---|---|
|
||||
| 1 | `cooldownStackSuffix` dropped from the key expression | **`processOperator`'s KEY** — where the collapse physically happens | `2 apps down inside the hour produced 1 operator mail(s), want 2 (suppressed=1)`, and the suppression row reads `key=c1:app_start_failed` — the live shape reproduced in a unit test |
|
||||
| 2 | the allow-list check removed from the suffix | **the REGISTER** — the fence that keeps the backup family coarse | `crossdrive_failed` … `= ":bookstack", want ""`, plus `app_deployed`, `app_removed`, `backup_failed` — the fence convicting exactly the types it was written for |
|
||||
|
||||
Both mutations asserted their pre-fix text was present before rewriting and printed `MUTATION
|
||||
APPLIED`. **Neither passed first time**, and the check for that was explicit after yesterday's inert
|
||||
mutation.
|
||||
|
||||
Gate 11 carries its own controls rather than a mutation, because the gate *is* the guard: ten cases
|
||||
in §Part 3 of the drill record, including the historical red-proof against yesterday's real file.
|
||||
|
||||
## 6. Test counts
|
||||
|
||||
| Repo | Before | After |
|
||||
|---|---|---|
|
||||
| felhom.eu hub | 709 | **716** |
|
||||
|
||||
`go build ./... && go vet ./... && go test ./...` in the hub → **exit 0, zero failures**.
|
||||
`python3 scripts/repo_gates.py --fast` → **12/12 OK** in felhom.eu; controller and agent runners both
|
||||
OK with gate 11 registered.
|
||||
|
||||
## 7. Deployed hub version, and the manifest commit
|
||||
|
||||
**`gitea.dooplex.hu/admin/felhom-hub:0.108.0`**, ArgoCD `Synced` / `Healthy`, rolled out.
|
||||
Deployed by **`45659bd`**, which bumped `manifests/hub.yaml:128`. The image was pushed to the registry
|
||||
**before** that commit landed, so a sync could never have pointed at a missing tag. Never
|
||||
`kubectl set image`.
|
||||
|
||||
**One thing worth stating because it looked like success and was not:** the first `refresh=hard` +
|
||||
sync reported `successfully rolled out` while the deployment still read **0.107.0** and the app read
|
||||
`OutOfSync` — ArgoCD had synced a pre-push revision. A second hard refresh took it to
|
||||
`Synced rev=45659bd` and the image then read 0.108.0. **The rollout message alone would have been a
|
||||
false confirmation**; the image tag is the observable that settles it.
|
||||
|
||||
## 8. The live walk
|
||||
|
||||
All counts filtered `created_at >= T0` (`2026-08-23 11:56:06Z`) so yesterday's two inert Scenario H
|
||||
probe rows cannot contaminate them.
|
||||
|
||||
### Step 1 — Scenario A: two different apps, four minutes apart ✅
|
||||
|
||||
```
|
||||
[WARN] [api] Event from demo-hp: severity "warn" is not in {info,warning,error,critical} — coercing
|
||||
to "info", which severityNotifies DROPS, so this backup_failed alert will reach NOBODY. Fix the
|
||||
emitting controller; this event is stored but not routed.
|
||||
sent operator Telepített alkalmazás nem fut: OpenGist 2026-08-23 11:56:57
|
||||
sent operator Telepített alkalmazás nem fut: Calibre-Web A 2026-08-23 12:00:57
|
||||
sent: 2 suppressed: 0
|
||||
```
|
||||
|
||||
Both the bad-severity POST and the `error` control returned **200** (nothing lost), and the control
|
||||
produced **no** warning — the guard does not fire on the normal path.
|
||||
Yesterday, the identical shape: `bookstack` **sent** 09:27:51, `privatebin` **suppressed** 09:31:51
|
||||
under `key=demo-hp:app_start_failed`.
|
||||
|
||||
## 9. Register size
|
||||
### Step 2 — Scenario B: each app again inside the hour ✅
|
||||
|
||||
```
|
||||
suppressed OpenGist operator cooldown 1h, key=demo-hp:app_start_failed:opengist 12:02:48
|
||||
suppressed Calibre-Web operator cooldown 1h, key=demo-hp:app_start_failed:calibre-web 12:02:48
|
||||
```
|
||||
|
||||
One `sent` and one `suppressed` per app — **the hour is unchanged** — and the two keys **differ by the
|
||||
app**, against v0.107.0's single shared key. The hub records the key only on a suppression, which is
|
||||
why this step is what exposes it.
|
||||
|
||||
*Method note:* the repeats were posted through `/api/v1/event`, the exact endpoint the controller
|
||||
invokes, with the controller's own `AppDetails` payload. The controller's own event is edge-triggered
|
||||
per app, so a down→down cycle is silent **by design** and cannot re-fire from the box.
|
||||
|
||||
### Step 3 — Scenario C: a backup-family event carrying `stack_name` ✅
|
||||
|
||||
```
|
||||
sent opengist
|
||||
suppressed calibre-web operator cooldown 1h, key=demo-hp:crossdrive_failed 12:03:44
|
||||
```
|
||||
|
||||
**Byte-identical to v0.107.0's key, with no app suffix.** The "before" value was obtained two
|
||||
independent ways, both stated: derivation from the v0.107.0 expression (which has no app term), and
|
||||
`TestR389_NoOtherEventTypeKeyChanges`, which models that expression inline for 12 event types **and
|
||||
carries a positive control proving it can see a key change before reporting that none occurred**.
|
||||
|
||||
### Step 4 — Part 2's burst ✅ (see §9)
|
||||
|
||||
### Step 5 — Part 3's gate ✅ all controls plus the historical red-proof (see §10)
|
||||
|
||||
## 9. Part 2's three counts, and the judgement
|
||||
|
||||
Three apps stopped in one scan, after checking none carried a live cooldown — a stale one would have
|
||||
halved the count and made the answer look better than it is:
|
||||
|
||||
| | |
|
||||
|---|---|
|
||||
| attempted | **3** |
|
||||
| sent | **3** |
|
||||
| suppressed | **0** |
|
||||
|
||||
**Plain judgement: per-app is the right grain and this volume is acceptable.** The reference box has 8
|
||||
deployed apps, so a total outage is 8 mails; the boot grace (90 s), the quiesce grace (180 s) and the
|
||||
per-app edge trigger absorb reboots, backup cycles and persistently-dead apps. **No burst-digest row
|
||||
was filed**, and the condition that would reopen it is recorded rather than left implicit: the volume
|
||||
scales linearly with app count and has no ceiling, so a box large enough that a total outage is
|
||||
unreadable is the point at which the answer becomes a digest with a customer message — not a wider
|
||||
cooldown.
|
||||
|
||||
## 10. Which repos gate 11 is registered in
|
||||
|
||||
| Repo | Registered | Note |
|
||||
|---|---|---|
|
||||
| `felhom.eu` | **yes** — gate 11 | its own `REPORT.md` is the gate's first real subject |
|
||||
| `felhom-controller` | **yes** | already had `SHARED_*` constants; one constant + one `GATES` line |
|
||||
| `felhom-agent` | **yes** | same; no observations section today, so it passes quietly |
|
||||
| `app-catalog-felhom.eu` | **NO** | filed as **R-391** |
|
||||
|
||||
**Why not the catalog.** `catalog_gates.py` has no shared-gate mechanism at all: `run_gate` joins
|
||||
every entry against its **own** `scripts/` directory, so it cannot invoke a sibling repo's script; and
|
||||
the loop appends `--all` to every gate unconditionally, which the observations gate would read as a
|
||||
path. Registering there needs `run_gate`'s contract widened **and** its argument handling changed — a
|
||||
refactor of a runner whose shape is deliberately different, in a repo this task marked out of scope.
|
||||
Exposure today is nil (that repo's `REPORT.md` has no observations section, and the gate passes
|
||||
quietly on that), but a future session could write one. **Filed rather than left as a sentence in a
|
||||
report, which is the exact failure this session exists to fix.**
|
||||
|
||||
## 11. Evidence
|
||||
|
||||
`documentation/audits/DRILL-cooldown-grain-2026-08-23/evidence/` — 17 files: 2 red-proof transcripts,
|
||||
3 gate-11 control files covering 10 cases, 11 live-walk files, and a 1803-line controller-log window
|
||||
**pulled off before the apps were restored**.
|
||||
|
||||
## 12. Teardown, three layers
|
||||
|
||||
1. **Guest 9201 / apps** — nothing provisioned. Five apps stopped across the walk (`opengist`,
|
||||
`calibre-web`, `kimai`, `romm`, `paperless-ngx`); **all restarted and confirmed healthy**, 17
|
||||
containers up. The three retained subjects (`docmost`, `bookstack`, `privatebin`) were not touched.
|
||||
No app rebuilt, redeployed or restored.
|
||||
2. **No VM, no bake** — this session built no golden and started no drill VM.
|
||||
3. **Hub-side, stated explicitly.** The hub *was* written: deployed to v0.108.0 via the manifest, and
|
||||
**six probe events POSTed for Scenarios B and C** (two `app_start_failed`, two `crossdrive_failed`,
|
||||
plus yesterday's two, left in place deliberately). They are inert event rows for `demo-hp`, named
|
||||
here rather than left to be found, and **every count in this report is `created_at`-filtered so
|
||||
they cannot contaminate it**. Nothing else: no appliance registered, no customer created, no
|
||||
artifact manifest changed, floor untouched.
|
||||
|
||||
## 13. Register size
|
||||
|
||||
| File | Before | After |
|
||||
|---|---|---|
|
||||
| `OPEN-ITEMS.md` | 328,325 B | **328,132 B** |
|
||||
| `CLOSED-ITEMS.md` | 71,441 B | **74,642 B** |
|
||||
| `OPEN-ITEMS.md` | 328,132 B | **331,024 B** |
|
||||
| `CLOSED-ITEMS.md` | 74,642 B | **76,855 B** |
|
||||
|
||||
R-329 and R-386 compressed into CLOSED with their rules kept; **R-387** (closed) and **R-388** (the
|
||||
notification-model product decision, open, operator's call) filed.
|
||||
R-389 filed, then closed and compressed into `CLOSED-ITEMS.md` in the same session. **R-390** (the
|
||||
golden-bake runbook's missing `pveam update`) and **R-391** (the catalog runner) filed open.
|
||||
|
||||
## 10. Observations
|
||||
## 14. Observations
|
||||
|
||||
> **Markers added 2026-08-24 (gate 11, R-389).** Item 1 is the finding that had no row; adding its
|
||||
> marker is the first thing the gate ever asked for. The observations' text is unchanged.
|
||||
> **Gate 11's first real subject is this section.** Each item carries `FILED: R-NNN` naming a row
|
||||
> opened this session, or `NOT-A-FINDING:` with its reason.
|
||||
|
||||
1. **The operator cooldown key carries no app identifier.** PrivateBin's alarm four minutes after
|
||||
BookStack's was logged `suppressed — operator cooldown 1h, key=demo-hp:app_start_failed`, so **only
|
||||
the first app-down per hour reaches the operator by e-mail**. R-182's known shape; harmless while
|
||||
the event was undeliverable, and no longer. Not fixed here.
|
||||
FILED: R-389
|
||||
2. Two probe events remain as rows for `demo-hp` from Scenario H — inert, and named rather than left.
|
||||
NOT-A-FINDING: two inert event rows on a Tier 0 demo box, created deliberately as a live
|
||||
control and named in that session's teardown; they carry no state and nothing reads them.
|
||||
1. **The gate's specification would have passed the item the gate exists to catch.** It said an
|
||||
observation may "cite an `R-NNN` that resolves" — but yesterday's lost item cites `R-182`, which
|
||||
resolves, as an **analogy** rather than as its own row. No parser can tell citation-as-precedent
|
||||
from citation-as-filing by reading prose, so the marker is explicit instead. The discrepancy is
|
||||
recorded in the gate's docstring and proven by `EDGE 7`.
|
||||
NOT-A-FINDING: this is a design decision taken and documented inside the deliverable itself, not a
|
||||
defect left behind — the gate ships with the stricter rule and its reasoning, so there is nothing
|
||||
outstanding for a row to track.
|
||||
2. **The burst has no ceiling.** Three apps in one scan produce three mails; the reference box's worst
|
||||
case is 8, and it scales linearly with app count.
|
||||
NOT-A-FINDING: measured and judged acceptable at today's scale in §9, with the reopening condition
|
||||
stated there; filing a row for a digest would queue work the operator has not asked for and that
|
||||
needs a customer message and their call on volume.
|
||||
3. **ArgoCD reported "successfully rolled out" while still running the old image.** The first sync ran
|
||||
against a pre-push revision; only a second hard refresh moved it. The rollout message alone was a
|
||||
false confirmation and the image tag was the observable that settled it.
|
||||
NOT-A-FINDING: the existing runbook already says to verify the image tag after a sync, and this run
|
||||
followed it and caught the discrepancy — the procedure worked; recording the near-miss here is the
|
||||
appropriate weight.
|
||||
4. **The golden-bake runbook still omits `pveam update`**, and its failure names the wrong cause.
|
||||
FILED: R-390
|
||||
5. **Gate 11 is registered in three runners, not four** — `catalog_gates.py` cannot invoke a sibling
|
||||
script and appends `--all` to every gate.
|
||||
FILED: R-391
|
||||
6. **Deliberately left open, untouched:** R-102, R-359, R-385, R-387, R-388's redesign.
|
||||
NOT-A-FINDING: a pointer to rows that already exist, carried so their absence from this session
|
||||
reads as deliberate rather than forgotten.
|
||||
|
||||
@@ -1,7 +1,7 @@
|
||||
# STATUS — what works, what's broken, what's next
|
||||
|
||||
**Updated 2026-08-23 — the app-down alarm now actually reaches you by e-mail. It never has: 91 of
|
||||
them were filed and not one was ever sent. Released and NOT yet delivered: step 1 is yours.**
|
||||
**Updated 2026-08-23 — you now hear about EVERY broken app, not just the first one each hour. The
|
||||
hub deployed itself; nothing is waiting on you except the floor from the last release.**
|
||||
|
||||
> **A view, not a source.** `documentation/backlog/OPEN-ITEMS.md` is the authority; this page restates
|
||||
> part of it in plain words, and **nothing may exist only here**. **Items, not paragraphs. One screen.**
|
||||
@@ -12,20 +12,14 @@ them were filed and not one was ever sent. Released and NOT yet delivered: step
|
||||
*This section is allowed to be longer than one screen, and each item says what happens if you do
|
||||
nothing.*
|
||||
|
||||
1. **Vouch the golden carrying controller 0.223.0** — Hub → Configuration → Day-0 artifacts.
|
||||
**It is already baked, published and round-trip verified**
|
||||
(`documentation/tests/golden-0.223.0-2026-08-23/`); only the vouch is left, and only you can do it.
|
||||
**It is a THREE-field save:** `golden_version` → **0.223.0**, `agent_version` → **0.130.0**,
|
||||
`min_agent` → **0.129.0**. **Then** raise the floor to **0.223.0**, last, in its own save.
|
||||
**If you do nothing:** the fleet stays on 0.222.0, where the app-down alarm shows on the dashboard
|
||||
and e-mails nobody, and where an app stopped from outside is still reported as if you had stopped
|
||||
it yourself. New machines still receive 0.222.0.
|
||||
*(The hub half is already live — v0.107.0 deployed itself through the manifest. Nothing owed there.)*
|
||||
1. **Raise the controller floor to 0.223.0** — Hub → Configuration, the global floor box, on its own.
|
||||
You already vouched the golden (the hub reads `golden_version` 0.223.0), but the floor still reads
|
||||
**0.222.0**, so the last step of that release is outstanding.
|
||||
**If you do nothing:** boxes are offered 0.223.0 but nothing requires it, so a machine that misses
|
||||
the offer stays on 0.222.0 — where an app stopped from outside is reported as if you stopped it.
|
||||
|
||||
2. **A new switch has appeared for your customers, and it is OFF** — „Alkalmazás nem fut". **You get
|
||||
the e-mail either way**; the switch only decides whether the customer also does. This is what you
|
||||
asked for and it needs nothing from you. Mentioned so it is not a surprise the first time you see
|
||||
the settings page.
|
||||
2. **Nothing else.** This release changed only the hub, and the hub deploys itself through its
|
||||
manifest. No golden, no vouch, no controller.
|
||||
|
||||
3. **Whether to change the hub password** (R-350). I printed it into my own session log on 20 August.
|
||||
Not in git, not in any saved file — in the log on this machine. **If you do nothing:** it stays as
|
||||
|
||||
@@ -155,6 +155,48 @@ it is deliberately absent from `DefaultEnabledEvents`, and deliberately **not**
|
||||
|
||||
---
|
||||
|
||||
## 6.2 The delivery grain — how often, and per what [DESIGN, R-389 — hub v0.108.0]
|
||||
|
||||
**"How loud" is a separate decision from "does it alarm", and it is made in one place**: the operator
|
||||
cooldown key at `processOperator`. Everything sharing a key is collapsed for **one hour**.
|
||||
|
||||
| Family | Grain | Key carries | Why |
|
||||
|---|---|---|---|
|
||||
| app down (`app_start_failed`) | **per APP** | `…:<stack_name>` | no digest exists for it — see below |
|
||||
| backup run (`backup_run_failures`) | per RUN | `…:<run_id>` | a digest already lists every failing app; one per run |
|
||||
| tiered backup (`whole_guest_backup_failed`, …) | per TIER | `…:<tier>` | the tiers fail independently and mean different things |
|
||||
| everything else, incl. `crossdrive_failed` | per TYPE, per hour | — | coarse **on purpose** |
|
||||
|
||||
**The default is COARSE and that is deliberate.** R-97a and R-182 exist so that one full disk produces
|
||||
**one** mail listing every affected app rather than one per app. Widening the grain is what makes an
|
||||
operator stop reading their alerts, which is the same failure as not sending them.
|
||||
|
||||
**`app_start_failed` is the exception because it has no digest.** There is no `apps_down_run`
|
||||
summarising a scan the way `backup_run_failures` summarises a run, so per-app is the only grain
|
||||
available that does not lose alarms. Until hub v0.108.0 it was keyed per type, and **only the first
|
||||
broken app per hour reached the operator** — measured 2026-08-23: `bookstack` sent at 09:27:51,
|
||||
`privatebin` suppressed at 09:31:51 under `key=demo-hp:app_start_failed`.
|
||||
|
||||
**The mechanism is a named ALLOW-LIST (`perAppCooldownEvents`), not a payload rule**, and the
|
||||
distinction is load-bearing rather than stylistic: **`crossdrive_failed` is severity `error`, reaches
|
||||
the operator leg, and carries `stack_name`** through `CrossDriveDetails`. A rule of the form "if the
|
||||
details carry a stack_name, split per app" would have split it, silently, and undone R-182.
|
||||
`cooldownStackSuffix` therefore takes the **event type** as well as the details — an asymmetry with
|
||||
its two siblings, and the reason for it is exactly this.
|
||||
|
||||
**Fenced act:** adding an entry to `perAppCooldownEvents` for a type whose family has a digest, or
|
||||
whose coarse cooldown is deliberate. Reading the register anywhere is fine.
|
||||
|
||||
**Measured burst, so the volume is a number and not an impression:** three apps stopped in one scan
|
||||
produced **three attempted, three sent, zero suppressed** (2026-08-23). The reference box has 8
|
||||
deployed apps, so a total outage is 8 mails. The boot grace (90 s), the quiesce grace (180 s) and the
|
||||
per-app edge trigger absorb reboots, backup cycles and persistently-dead apps. **This scales linearly
|
||||
with app count and has no ceiling** — the condition that would reopen the question is a box large
|
||||
enough that a total outage is unreadable, at which point the answer is a digest with a customer
|
||||
message, not a wider cooldown.
|
||||
|
||||
---
|
||||
|
||||
## 7. The intent test [DESIGN, R-386 — CLOSED controller v0.223.0]
|
||||
|
||||
**"The customer stopped this" is asked of the FIELD THAT RECORDS IT, never inferred from the state.**
|
||||
|
||||
@@ -0,0 +1,118 @@
|
||||
# DRILL — R-389: only the first broken app per hour reached the operator (2026-08-23)
|
||||
|
||||
**Hub v0.107.0 → v0.108.0. No controller release, so no golden and no floor.** Live leg on `demo-hp`
|
||||
(Tier 0) plus the hub. UNATTENDED. Method: the hub's own SQLite records plus the controller log.
|
||||
Guest and hub DB are UTC; the hub pod logs CEST.
|
||||
|
||||
**No halt condition fired.**
|
||||
|
||||
## The pair that says it
|
||||
|
||||
Same box, same shape — two apps down four minutes apart, inside one hour:
|
||||
|
||||
```
|
||||
2026-08-23 09:27:51 sent BookStack <- v0.107.0
|
||||
2026-08-23 09:31:51 suppressed PrivateBin operator cooldown 1h, key=demo-hp:app_start_failed
|
||||
|
||||
2026-08-23 11:56:57 sent OpenGist <- v0.108.0
|
||||
2026-08-23 12:00:57 sent Calibre-Web
|
||||
```
|
||||
|
||||
**2 sent / 0 suppressed**, where the day before the identical shape gave one of each.
|
||||
|
||||
And the keys, from the hub's own suppression rows — it records the key only when it declines:
|
||||
|
||||
```
|
||||
operator cooldown 1h, key=demo-hp:app_start_failed:opengist
|
||||
operator cooldown 1h, key=demo-hp:app_start_failed:calibre-web
|
||||
```
|
||||
|
||||
against v0.107.0's shared `key=demo-hp:app_start_failed`.
|
||||
|
||||
## The fence, and why it is not decoration
|
||||
|
||||
`crossdrive_failed` is severity `error`, reaches the operator leg, and carries `stack_name` — through
|
||||
`CrossDriveDetails`, **a different struct from `AppDetails`**. A rule of the form *"if the details
|
||||
carry a stack_name, split per app"* would have split it and silently undone R-182.
|
||||
|
||||
Proven live: two different apps' `crossdrive_failed`, one minute apart →
|
||||
|
||||
```
|
||||
sent opengist
|
||||
suppressed calibre-web operator cooldown 1h, key=demo-hp:crossdrive_failed
|
||||
```
|
||||
|
||||
**Byte-identical to the derived v0.107.0 key. No app suffix.** That is why `cooldownStackSuffix`
|
||||
takes the event type as well as the details, unlike its two siblings.
|
||||
|
||||
## Part 2 — the burst, measured
|
||||
|
||||
Three apps stopped in one scan (`kimai`, `romm`, `paperless-ngx`, none with a live cooldown — checked
|
||||
first, because a stale one would have halved the count and made the answer look better than it is):
|
||||
|
||||
| | |
|
||||
|---|---|
|
||||
| attempted | **3** |
|
||||
| sent | **3** |
|
||||
| suppressed | **0** |
|
||||
|
||||
**Judgement: per-app is the right grain, and this volume is acceptable.** The reference box has 8
|
||||
deployed apps, so a total outage is 8 mails; the boot grace (90 s), the quiesce grace (180 s) and the
|
||||
per-app edge trigger absorb reboots, backup cycles and persistently-dead apps. **No digest row was
|
||||
filed.** The reopening condition is stated rather than left implicit: it scales linearly with app
|
||||
count and has no ceiling, so a box large enough that a total outage is unreadable is the point at
|
||||
which the answer becomes a digest with a customer message — not a wider cooldown.
|
||||
|
||||
## Part 3 — gate 11, and the spec discrepancy it forced
|
||||
|
||||
The gate refuses a push whose `REPORT.md` carries an observation with neither `FILED: R-NNN` nor
|
||||
`NOT-A-FINDING: <reason>`.
|
||||
|
||||
**The specification said an item may "cite an R-NNN that resolves". That rule would have passed the
|
||||
very item the gate was built to catch.** Yesterday's lost observation reads *"This is R-182's known
|
||||
cooldown-key shape…"* — `R-182` resolves, and it is cited as an **analogy**, not as the row that files
|
||||
it. No parser can tell citation-as-precedent from citation-as-filing by reading prose. The marker is
|
||||
therefore explicit, and the discrepancy is recorded in the gate's docstring rather than quietly
|
||||
resolved. `EDGE 7` in the evidence is that exact case, convicted.
|
||||
|
||||
| Control | Expected | Observed |
|
||||
|---|---|---|
|
||||
| historical: yesterday's real section, verbatim | refuse | **exit 1**, naming both items |
|
||||
| plant an observation with no row | refuse | **exit 1** |
|
||||
| add the row | pass | **exit 0** |
|
||||
| remove the observation | pass quietly | **exit 0** |
|
||||
| `FILED:` a row that does not resolve | refuse | **exit 1** |
|
||||
| `NOT-A-FINDING:` with no reason | refuse | **exit 1** |
|
||||
| both markers on one item | refuse | **exit 1** |
|
||||
| section present, no numbered items | inconclusive | **exit 2**, saying what it could not read |
|
||||
| no `REPORT.md` at all | inconclusive | **exit 2** |
|
||||
| a bare `R-182` mention (the trap) | refuse | **exit 1** |
|
||||
|
||||
## Evidence index (`evidence/`)
|
||||
|
||||
| File | What it shows |
|
||||
|---|---|
|
||||
| `redproof-1-key.txt` | suffix dropped from the key → `1 operator mail(s), want 2`, and the live key shape reproduced |
|
||||
| `redproof-2-allowlist.txt` | allow-list removed → `crossdrive_failed`, `app_deployed`, `app_removed`, `backup_failed` all split per app |
|
||||
| `gate11-01-historical-redproof.txt` | yesterday's actual observations section, refused |
|
||||
| `gate11-02-three-controls.txt` | plant → refuse, file → pass, remove → pass |
|
||||
| `gate11-03-edges.txt` | seven boundaries incl. the R-182 trap |
|
||||
| `live-01`…`live-05` | Scenarios A and B: both sent, then each suppressed under its own key |
|
||||
| `live-06`, `live-07` | Scenario C: crossdrive stays coarse |
|
||||
| `live-08`, `live-09` | the burst, with its pre-check |
|
||||
| `live-10-full-controller-log.txt` | 1803 lines, pulled before the restore |
|
||||
|
||||
## Teardown
|
||||
|
||||
Nothing provisioned. Five apps were stopped across the walk (`opengist`, `calibre-web`, `kimai`,
|
||||
`romm`, `paperless-ngx`) and **all were restarted and confirmed healthy** — 17 containers up. No app
|
||||
was rebuilt, redeployed or restored; the three retained subjects (`docmost`, `bookstack`,
|
||||
`privatebin`) were not touched at all.
|
||||
|
||||
**Hub-side, stated explicitly.** The hub was written this session: deployment to v0.108.0 via the
|
||||
manifest, and **six probe events were POSTed to the live hub** for Scenarios B and C (two
|
||||
`app_start_failed`, two `crossdrive_failed`, plus the two from yesterday's Scenario H that were
|
||||
deliberately left in place). They are inert event rows for `demo-hp` and are named here rather than
|
||||
left to be found. **Every count in this drill is filtered by `created_at >= T0`** precisely so those
|
||||
rows cannot contaminate it. Nothing else: no appliance registered, no customer created, no artifact
|
||||
manifest changed, floor untouched.
|
||||
+16
@@ -0,0 +1,16 @@
|
||||
### HISTORICAL RED-PROOF — yesterday's ACTUAL observations section, verbatim from 2f7c9a6 ###
|
||||
report : ../../../../../tmp/claude-1000/-mnt-5-hdd-felhom-eu-git/3b17bdbd-9113-4908-9b30-15678f2978bc/scratchpad/gate11/REPORT.md
|
||||
section : ## 10. Observations — recorded, not acted on
|
||||
observation items : 2
|
||||
|
||||
OBSERVATIONS GATE FAILED: 2 observation(s) with nothing behind them.
|
||||
|
||||
item 1: **The operator cooldown key carries no app identifier.** PrivateBin's alarm fo…
|
||||
neither `FILED: R-NNN` nor `NOT-A-FINDING: <reason>`
|
||||
|
||||
item 2: Two probe events remain as rows for `demo-hp` from Scenario H — inert, and nam…
|
||||
neither `FILED: R-NNN` nor `NOT-A-FINDING: <reason>`
|
||||
|
||||
An observation that lives only in REPORT.md has a lifetime of ONE SESSION — this file is overwritten every time. That is how the cooldown-grain finding was lost on 2026-08-23 and had to be re-derived the next day.
|
||||
Fix: add `FILED: R-NNN` naming the row you opened for it, or `NOT-A-FINDING: <why this is not worth a row>`. Opening the row is the default; declaring is the exception and needs its reason stated.
|
||||
EXIT=1 (want 1)
|
||||
+25
@@ -0,0 +1,25 @@
|
||||
### CONTROL 1 — an observation with NO row: must REFUSE (1) ###
|
||||
report : ../../../../../tmp/claude-1000/-mnt-5-hdd-felhom-eu-git/3b17bdbd-9113-4908-9b30-15678f2978bc/scratchpad/gate11/c1.md
|
||||
section : ## 9. Observations
|
||||
observation items : 1
|
||||
|
||||
OBSERVATIONS GATE FAILED: 1 observation(s) with nothing behind them.
|
||||
|
||||
item 1: The operator cooldown key carries no app identifier, so only the first app per…
|
||||
neither `FILED: R-NNN` nor `NOT-A-FINDING: <reason>`
|
||||
|
||||
An observation that lives only in REPORT.md has a lifetime of ONE SESSION — this file is overwritten every time. That is how the cooldown-grain finding was lost on 2026-08-23 and had to be re-derived the next day.
|
||||
Fix: add `FILED: R-NNN` naming the row you opened for it, or `NOT-A-FINDING: <why this is not worth a row>`. Opening the row is the default; declaring is the exception and needs its reason stated.
|
||||
EXIT=1
|
||||
|
||||
### CONTROL 2 — add the row: must PASS (0) ###
|
||||
report : ../../../../../tmp/claude-1000/-mnt-5-hdd-felhom-eu-git/3b17bdbd-9113-4908-9b30-15678f2978bc/scratchpad/gate11/c2.md
|
||||
section : ## 9. Observations
|
||||
observation items : 1
|
||||
OK 1. FILED R-389
|
||||
observations gate OK — every observation is either filed or explicitly declared
|
||||
EXIT=0
|
||||
|
||||
### CONTROL 3 — remove the observation entirely: must PASS quietly (0) — Scenario F ###
|
||||
observations gate OK — c3.md has no observations section (nothing to check)
|
||||
EXIT=0
|
||||
@@ -0,0 +1,44 @@
|
||||
### EDGE 1 — NOT-A-FINDING with a reason: PASS (0) ###
|
||||
section : ## Observations
|
||||
observation items : 1
|
||||
OK 1. NOT-A-FINDING
|
||||
observations gate OK — every observation is either filed or explicitly declared
|
||||
EXIT=0
|
||||
|
||||
### EDGE 2 — FILED with a row that does NOT resolve: REFUSE (1) ###
|
||||
`FILED: R-99999` does not resolve — no row for R-99999 in OPEN-ITEMS.md or CLOSED-ITEMS.md
|
||||
|
||||
An observation that lives only in REPORT.md has a lifetime of ONE SESSION — this file is overwritten every time. That is how the cooldown-grain finding was lost on 2026-08-23 and had to be re-derived the next day.
|
||||
Fix: add `FILED: R-NNN` naming the row you opened for it, or `NOT-A-FINDING: <why this is not worth a row>`. Opening the row is the default; declaring is the exception and needs its reason stated.
|
||||
EXIT=1
|
||||
|
||||
### EDGE 3 — NOT-A-FINDING with no reason: REFUSE (1) ###
|
||||
`NOT-A-FINDING:` carries no reason — the reason is the whole point of the marker
|
||||
|
||||
An observation that lives only in REPORT.md has a lifetime of ONE SESSION — this file is overwritten every time. That is how the cooldown-grain finding was lost on 2026-08-23 and had to be re-derived the next day.
|
||||
Fix: add `FILED: R-NNN` naming the row you opened for it, or `NOT-A-FINDING: <why this is not worth a row>`. Opening the row is the default; declaring is the exception and needs its reason stated.
|
||||
EXIT=1
|
||||
|
||||
### EDGE 4 — BOTH markers on one item: REFUSE (1) ###
|
||||
carries BOTH `FILED:` and `NOT-A-FINDING:` — decide which it is
|
||||
|
||||
An observation that lives only in REPORT.md has a lifetime of ONE SESSION — this file is overwritten every time. That is how the cooldown-grain finding was lost on 2026-08-23 and had to be re-derived the next day.
|
||||
Fix: add `FILED: R-NNN` naming the row you opened for it, or `NOT-A-FINDING: <why this is not worth a row>`. Opening the row is the default; declaring is the exception and needs its reason stated.
|
||||
EXIT=1
|
||||
|
||||
### EDGE 5 — section present, NO numbered items: INCONCLUSIVE (2) ###
|
||||
OBSERVATIONS GATE INCONCLUSIVE: found the section '## Observations' but no numbered items under it.
|
||||
Items must be a numbered list (`1.`, `2.` …). If the section is deliberately empty, remove the heading — an empty section reads as coverage while providing none.
|
||||
EXIT=2
|
||||
|
||||
### EDGE 6 — REPORT.md absent: INCONCLUSIVE (2) ###
|
||||
OBSERVATIONS GATE INCONCLUSIVE: no REPORT.md at /tmp/claude-1000/-mnt-5-hdd-felhom-eu-git/3b17bdbd-9113-4908-9b30-15678f2978bc/scratchpad/gate11/does-not-exist.md
|
||||
This gate is registered here because this repo keeps one. If it was removed deliberately, unregister the gate in the runner rather than leaving it unreadable.
|
||||
EXIT=2
|
||||
|
||||
### EDGE 7 — the R-182 TRAP: a bare mention must NOT satisfy the gate (1) ###
|
||||
neither `FILED: R-NNN` nor `NOT-A-FINDING: <reason>`
|
||||
|
||||
An observation that lives only in REPORT.md has a lifetime of ONE SESSION — this file is overwritten every time. That is how the cooldown-grain finding was lost on 2026-08-23 and had to be re-derived the next day.
|
||||
Fix: add `FILED: R-NNN` naming the row you opened for it, or `NOT-A-FINDING: <why this is not worth a row>`. Opening the row is the default; declaring is the exception and needs its reason stated.
|
||||
EXIT=1
|
||||
@@ -0,0 +1,16 @@
|
||||
=== LIVE WALK — pre-state ===
|
||||
hub image: gitea.dooplex.hu/admin/felhom-hub:0.108.0
|
||||
2026-08-23T11:56:06Z
|
||||
controller: gitea.dooplex.hu/admin/felhom-controller:0.223.0 Up 2 hours (healthy)
|
||||
|
||||
bookstack Up 9 hours (healthy)
|
||||
bookstack-db Up 2 hours (healthy)
|
||||
calibre-web Up 9 hours (healthy)
|
||||
docmost Up 6 hours (healthy)
|
||||
docmost-postgres Up 6 hours (healthy)
|
||||
docmost-redis Up 6 hours (healthy)
|
||||
opengist Up 9 hours (healthy)
|
||||
privatebin Up 2 hours (healthy)
|
||||
|
||||
-- T0 for every count below: any notification_log row created at or after this instant --
|
||||
T0=2026-08-23T11:56:06Z
|
||||
+8
@@ -0,0 +1,8 @@
|
||||
=== SCENARIO A — two DIFFERENT apps down inside the hour ===
|
||||
subjects: opengist and calibre-web (single-container, intent running, neither on the
|
||||
do-not-rebuild list). Stopped OUT OF BAND so the controller's own Deploying skip cannot mask it.
|
||||
opengist desired_state: running
|
||||
calibre-web desired_state: running
|
||||
|
||||
APP 1 (opengist) down at: 2026-08-23T11:56:26Z
|
||||
opengist Exited (0) 5 seconds ago
|
||||
+8
@@ -0,0 +1,8 @@
|
||||
-- waiting 4 minutes, then APP 2 — the same gap that produced one mail yesterday --
|
||||
APP 2 (calibre-web) down at: 2026-08-23T12:00:48Z
|
||||
opengist Exited (0) 4 minutes ago
|
||||
calibre-web Exited (0) 5 seconds ago
|
||||
|
||||
-- letting the second alarm land --
|
||||
2026/08/23 11:56:56 notifier.go:234: [INFO] Event pushed: app_start_failed (warning) — Telepített alkalmazás nem fut: OpenGist
|
||||
2026/08/23 12:00:57 notifier.go:234: [INFO] Event pushed: app_start_failed (warning) — Telepített alkalmazás nem fut: Calibre-Web Automated
|
||||
+11
@@ -0,0 +1,11 @@
|
||||
=== SCENARIO A — the hub's notification log, filtered to created_at >= T0 (11:56:06Z) ===
|
||||
(the T0 filter is what keeps yesterday's two inert Scenario H probe rows out of every count)
|
||||
|
||||
status channel message key_or_reason created_at
|
||||
------ -------- -------------------------------------------- ------------- -------------------
|
||||
sent operator Telepített alkalmazás nem fut: OpenGist 2026-08-23 11:56:57
|
||||
sent operator Telepített alkalmazás nem fut: Calibre-Web A 2026-08-23 12:00:57
|
||||
|
||||
-- counts --
|
||||
sent : 2
|
||||
suppressed: 0
|
||||
+8
@@ -0,0 +1,8 @@
|
||||
=== SCENARIO B — the SAME app again inside the hour: one sent, one suppressed ===
|
||||
Posted through the exact endpoint the controller invokes (/api/v1/event), with the controller's
|
||||
own AppDetails payload. The controller's own event is edge-triggered per app, so a down->down
|
||||
cycle is silent by design and cannot re-fire it from the box.
|
||||
|
||||
T_B=2026-08-23T12:02:48Z
|
||||
repeat OpenGist http=200
|
||||
repeat Calibre-Web http=200
|
||||
+15
@@ -0,0 +1,15 @@
|
||||
=== SCENARIOS A + B — every operator row since T0, with the KEY where the hub records it ===
|
||||
status app key_or_reason created_at
|
||||
---------- ---------------------- -------------------------------------------------------------- -------------------
|
||||
sent ut: OpenGist 2026-08-23 11:56:57
|
||||
sent ut: Calibre-Web Automa 2026-08-23 12:00:57
|
||||
suppressed ut: OpenGist operator cooldown 1h, key=demo-hp:app_start_failed:opengist 2026-08-23 12:02:48
|
||||
suppressed ut: Calibre-Web operator cooldown 1h, key=demo-hp:app_start_failed:calibre-web 2026-08-23 12:02:48
|
||||
|
||||
-- THE KEYS, side by side. Yesterday BOTH apps shared key=demo-hp:app_start_failed. --
|
||||
operator cooldown 1h, key=demo-hp:app_start_failed:opengist
|
||||
operator cooldown 1h, key=demo-hp:app_start_failed:calibre-web
|
||||
|
||||
-- counts since T0 --
|
||||
sent : 2
|
||||
suppressed: 2
|
||||
+15
@@ -0,0 +1,15 @@
|
||||
=== SCENARIO C — a backup-family event carrying stack_name must NOT split per app ===
|
||||
subject: crossdrive_failed — severity error, reaches the operator leg, carries stack_name
|
||||
through CrossDriveDetails. Two DIFFERENT apps: if the key had split, both would send.
|
||||
|
||||
HOW THE 'BEFORE' VALUE WAS OBTAINED, two independent ways:
|
||||
(a) derivation from v0.107.0's expression, which is customerID:eventType + tier + run and
|
||||
carries no app term, so for this payload it is exactly 'demo-hp:crossdrive_failed';
|
||||
(b) TestR389_NoOtherEventTypeKeyChanges models that v0.107.0 expression INLINE and compares
|
||||
it against the live one for 12 event types — with a positive control proving the test
|
||||
can see a key change before it reports that none occurred.
|
||||
The live key printed below must equal that.
|
||||
|
||||
T_C=2026-08-23T12:03:44Z
|
||||
crossdrive_failed opengist http=200
|
||||
crossdrive_failed calibre-web http=200
|
||||
+8
@@ -0,0 +1,8 @@
|
||||
-- SCENARIO C result: the operator rows for crossdrive_failed since T_C --
|
||||
status app key_or_reason created_at
|
||||
---------- ---------- --------------------------------------------------- -------------------
|
||||
suppressed alibre-web operator cooldown 1h, key=demo-hp:crossdrive_failed 2026-08-23 12:03:44
|
||||
sent pengist 2026-08-23 12:03:45
|
||||
|
||||
-- the live key, to compare against the derived v0.107.0 value 'demo-hp:crossdrive_failed' --
|
||||
operator cooldown 1h, key=demo-hp:crossdrive_failed
|
||||
@@ -0,0 +1,21 @@
|
||||
-- restoring the two Scenario A/B subjects before the burst --
|
||||
opengist Up 20 seconds (healthy)
|
||||
calibre-web Up 20 seconds (healthy)
|
||||
|
||||
=== PART 2 — THE BURST: three apps down in ONE scan ===
|
||||
subjects: kimai, romm, paperless-ngx — none has alarmed inside the last hour, so no
|
||||
pre-existing cooldown can mask the count. Verified below before stopping them.
|
||||
-- CONTROL: no app_start_failed row for these three in the last hour (a stale cooldown would
|
||||
silently halve the count and the burst would look better than it is) --
|
||||
rows for kimai/romm/paperless since 11:04: 0
|
||||
|
||||
T_BURST=2026-08-23T12:05:03Z
|
||||
2026-08-23T12:05:13Z
|
||||
kimai Exited (0) 9 seconds ago
|
||||
kimai-db Exited (0) 8 seconds ago
|
||||
paperless-postgres Exited (0) 4 seconds ago
|
||||
paperless-redis Exited (0) 4 seconds ago
|
||||
paperless-webserver Exited (0) 4 seconds ago
|
||||
romm Exited (0) 9 seconds ago
|
||||
romm-db Exited (0) 9 seconds ago
|
||||
romm-redis Exited (0) 9 seconds ago
|
||||
+18
@@ -0,0 +1,18 @@
|
||||
-- PART 2 — THE BURST RESULT (T_BURST = 12:05:03Z) --
|
||||
|
||||
controller side, what it SENT:
|
||||
2026/08/23 12:05:26 notifier.go:234: [INFO] Event pushed: app_start_failed (warning) — Telepített alkalmazás nem fut: Kimai
|
||||
2026/08/23 12:05:26 notifier.go:234: [INFO] Event pushed: app_start_failed (warning) — Telepített alkalmazás nem fut: RomM
|
||||
2026/08/23 12:05:26 notifier.go:234: [INFO] Event pushed: app_start_failed (warning) — Telepített alkalmazás nem fut: Paperless-ngx
|
||||
|
||||
hub side, the operator rows:
|
||||
status app key_or_reason created_at
|
||||
------ ----------------- ------------- -------------------
|
||||
sent ut: RomM 2026-08-23 12:05:27
|
||||
sent ut: Kimai 2026-08-23 12:05:27
|
||||
sent ut: Paperless-ngx 2026-08-23 12:05:27
|
||||
|
||||
THE THREE COUNTS:
|
||||
attempted (rows written): 3
|
||||
sent : 3
|
||||
suppressed : 0
|
||||
+1803
File diff suppressed because it is too large
Load Diff
@@ -0,0 +1,19 @@
|
||||
-- restoring the three burst subjects --
|
||||
2026-08-23T12:08:39Z
|
||||
bookstack Up 10 hours (healthy)
|
||||
bookstack-db Up 3 hours (healthy)
|
||||
calibre-web Up 4 minutes (healthy)
|
||||
docmost Up 6 hours (healthy)
|
||||
docmost-postgres Up 6 hours (healthy)
|
||||
docmost-redis Up 6 hours (healthy)
|
||||
filebrowser Up 44 hours (healthy)
|
||||
kimai Up 52 seconds (healthy)
|
||||
kimai-db Up About a minute (healthy)
|
||||
opengist Up 4 minutes (healthy)
|
||||
paperless-postgres Up 40 seconds (healthy)
|
||||
paperless-redis Up 40 seconds (healthy)
|
||||
paperless-webserver Up 30 seconds (health: starting)
|
||||
privatebin Up 2 hours (healthy)
|
||||
romm Up 41 seconds (healthy)
|
||||
romm-db Up 51 seconds (healthy)
|
||||
romm-redis Up 51 seconds (healthy)
|
||||
@@ -0,0 +1,9 @@
|
||||
### RED-PROOF 1 (R-389) — mutation: cooldownStackSuffix removed from the key ###
|
||||
### layer: processOperator's KEY — where the collapse happens ###
|
||||
--- FAIL: TestR389_TwoAppsInsideTheHourBothReachTheOperator (0.02s)
|
||||
r389_cooldown_grain_test.go:180: 2 apps down inside the hour produced 1 operator mail(s), want 2 (suppressed=1). This is R-389: the second app's alarm took the first app's cooldown slot.
|
||||
--- FAIL: TestR389_SameAppTwiceInsideTheHourIsStillSuppressed (0.02s)
|
||||
r389_cooldown_grain_test.go:220: the suppression row does not name the app in its key — that visibility is R-182's contribution and is how this defect was found: [sent | Telepített alkalmazás nem fut: BookStack | suppressed | Telepített alkalmazás nem fut: BookStack | operator cooldown 1h, key=c1:app_start_failed]
|
||||
FAIL
|
||||
FAIL gitea.dooplex.hu/admin/felhom-hub/internal/notify 0.166s
|
||||
FAIL
|
||||
+12
@@ -0,0 +1,12 @@
|
||||
### RED-PROOF 2 (R-389) — mutation: allow-list removed, suffix applied globally ###
|
||||
### layer: cooldownStackSuffix's REGISTER — the fence that keeps the backup family coarse ###
|
||||
--- FAIL: TestR389_StackSuffixIsAllowListedAndFailSoft (0.00s)
|
||||
--- FAIL: TestR389_StackSuffixIsAllowListedAndFailSoft/crossdrive_failed_is_NOT_split_per_app (0.00s)
|
||||
r389_cooldown_grain_test.go:58: cooldownStackSuffix("crossdrive_failed", "{\"stack_name\":\"bookstack\",\"method\":\"rsync\"}") = ":bookstack", want ""
|
||||
--- FAIL: TestR389_StackSuffixIsAllowListedAndFailSoft/app_deployed_is_not_in_the_register (0.00s)
|
||||
r389_cooldown_grain_test.go:58: cooldownStackSuffix("app_deployed", "{\"stack_name\":\"bookstack\",\"display_name\":\"BookStack\"}") = ":bookstack", want ""
|
||||
--- FAIL: TestR389_StackSuffixIsAllowListedAndFailSoft/app_removed_is_not_in_the_register (0.00s)
|
||||
r389_cooldown_grain_test.go:58: cooldownStackSuffix("app_removed", "{\"stack_name\":\"bookstack\",\"display_name\":\"BookStack\"}") = ":bookstack", want ""
|
||||
--- FAIL: TestR389_StackSuffixIsAllowListedAndFailSoft/backup_failed_is_not_in_the_register (0.00s)
|
||||
r389_cooldown_grain_test.go:58: cooldownStackSuffix("backup_failed", "{\"stack_name\":\"bookstack\",\"display_name\":\"BookStack\"}") = ":bookstack", want ""
|
||||
--- FAIL: TestR389_NoOtherEventTypeKeyChanges (0.00s)
|
||||
@@ -164,3 +164,4 @@
|
||||
| **R-384** | **An app whose DATABASE had died raised no dead-app alarm — the wrong question answered first.** Shipped in controller v0.222.0. Evidence: `audits/DRILL-r384-dead-db-alarm-2026-08-23/`. **Reasoning kept:** *The defect was the ORDER of two questions, not the `unhealthy` exclusion.* "Is a SUPERVISED member dead?" and "is a RUNNING member failing its healthcheck?" are different questions, and the second was answering the first — a dying database drags its own front end `unhealthy`, so the symptom the fault causes was what suppressed the alarm for it. **`IsDownState` is byte-identical and `unhealthy` stays excluded** — an unhealthy container is RUNNING, and folding it in reintroduces the flapping that exclusion exists to stop; **no new state was minted**, `StateDegraded` already means this. **Two things had to move and either alone leaves the defect standing:** the hoist, AND widening "some members are up" from `running > 0` to *any member not in the down bucket* — the old guard made the R-51 block unreachable in precisely the case it was written for. **The register's own suggested fix was WRONG and is recorded as such:** it proposed a sustained-`unhealthy` threshold on the `crashLoopAfter` model; the actual defect needed no threshold at all. **PROVEN LIVE the only way it can be** — the same fixture that printed `0 currently down` on 2026-08-22 printed **`1 currently down`** on 2026-08-23, with `app_start_failed` 7 s after the stop and the banner reading *„…nem fut: BookStack (degraded)"*. Scenario D measured **0 alarms across 9 scans** through a full stop→start cycle. | **CLOSED — SHIPPED + PROVEN-LIVE** (controller v0.222.0, 2026-08-23) | full text: `git show 1eb64bec5183:documentation/backlog/OPEN-ITEMS.md` |
|
||||
| **R-329** | **`app_start_failed` was emitted with severity `"warn"`, so every one of them was delivered to nobody.** Shipped in controller v0.223.0 (+ hub v0.107.0). Evidence: `audits/DRILL-r329-r386-2026-08-23/`. **Reasoning kept:** *The vocabulary is EXACT and it is the HUB's, not ours* — `{info, warning, error, critical}`; anything else is coerced to `info` at ingest and dropped by `severityNotifies` before BOTH legs. **This was the SECOND occurrence** (`DiskAlertKind.Severity` until v0.215.0), and its comment had recorded the lesson — **a comment is not a guard**, so the guard is now an AST walk over the whole controller, with the six variable-passing call sites registered by name because a walk cannot follow a variable and *an unlisted limit is not a limit, it is a hole*. **The register's own framing was that the DECISION was the work** — should a stopped app mail the customer at all? Answered: **operator always, customer OFF by default**, because `processOperator` never consults customer preferences, so one word fixed the operator leg and left the customer leg exactly where the ruling wanted it. **Deliberately NOT added to `operatorOnlyEvents`** — that would make the new toggle visible, flickable and structurally incapable of delivering. **Measured on the live hub DB: 91 events stored all-time, ZERO notification rows before the fix; one operator row, `warning`/`sent`, after it.** | **CLOSED — SHIPPED + PROVEN-LIVE** (controller v0.223.0 + hub v0.107.0, 2026-08-23) | full text: `git show 68a9f5475cd2:documentation/backlog/OPEN-ITEMS.md` |
|
||||
| **R-386** | **A single-container app stopped out of band raised no alarm, and a comment stated the opposite as settled fact.** Shipped in controller v0.223.0. Evidence: `audits/DRILL-r329-r386-2026-08-23/`. **Reasoning kept:** *the state test was guessing at something the product already knows.* `DesiredState` records the customer's intent, has **exactly one writer**, and is tri-state; `StateExited` never survives aggregation, so no state test can separate an out-of-band stop from a customer stop. The ruling: `Stopped` → no alarm, `Running` → **alarm**, **absent → UNKNOWN, keep today's behaviour AND announce it**. *Reading unknown as "nobody asked" would, on the first cycle after upgrade, e-mail about every app any owner ever deliberately stopped — fleet-wide, from a field that predates the intent it is being asked about.* **A rule without a mechanism is a wish:** every such suppression sets `IntentUnknown` and the names are logged at INFO, so an operator can answer *"how many apps am I blind to?"*. **`failedRestart` must still lift a `Stopped` intent or F-CRIT-1 re-opens.** **Fenced act: adding a `DesiredState` WRITER** — twelve of `StopStack`'s fourteen callers are machines. Proven live: alarm 24 s after an out-of-band `docker compose stop`, heartbeat `1 currently down` against the previous day's `0`; and with intent removed, suppressed *plus* the log line naming the app. **0 of 8 deployed apps on `demo-hp` carry an absent intent.** | **CLOSED — SHIPPED + PROVEN-LIVE** (controller v0.223.0, 2026-08-23) | full text: `git show 68a9f5475cd2:documentation/backlog/OPEN-ITEMS.md` |
|
||||
| **R-389** | **Only the FIRST broken app per hour reached the operator — the cooldown key named the event type, not the app.** Shipped in hub v0.108.0. Evidence: `audits/DRILL-cooldown-grain-2026-08-23/`. **Reasoning kept:** the fix is a THIRD SIBLING of `cooldownTierSuffix`/`cooldownRunSuffix`, separate for the reason the second one's docstring already gives — *the existing two keep byte-identical semantics for every type that uses them.* **`cooldownStackSuffix` takes the EVENT TYPE as well as the details, unlike its siblings, and that asymmetry is the whole safety property:** `tier` and `run_id` appear only on types that want that grain, `stack_name` does not. **`perAppCooldownEvents` is a named allow-list with `app_start_failed` and nothing else** — *the backup family's cooldown is coarse ON PURPOSE (R-97a, R-182) so one full disk sends one digest rather than one mail per app*, and this is not hypothetical: **`crossdrive_failed` is severity `error`, reaches the operator leg, and carries `stack_name` through a different struct**, so a payload-shape rule would have split it silently. **The fenced act is adding an entry for a type whose family has a digest or a coarse-by-design cooldown.** `app_start_failed` qualifies precisely because it has NO digest — there is no `apps_down_run` the way `backup_run_failures` summarises a run. **The hour is unchanged; the grain was the complaint.** Fail-soft: absent or malformed details degrade to the old key and the mail still goes. **PROVEN LIVE 2026-08-23:** two apps four minutes apart gave **2 sent / 0 suppressed** where the same shape gave 1 and 1 the day before, each repeat suppressed under its OWN key (`…:opengist`, `…:calibre-web`) against the previous day's shared `key=demo-hp:app_start_failed`; and `crossdrive_failed` for two different apps stayed **coarse** under `key=demo-hp:crossdrive_failed`, byte-identical to the derived v0.107.0 value. **AND IT WAS NEVER FILED UNTIL THE DAY IT WAS FIXED** — it lived in a REPORT.md observations paragraph, which is why gate 11 now exists. | **CLOSED — SHIPPED + PROVEN-LIVE** (hub v0.108.0, 2026-08-23) | full text: `git show 45659bdc5a2f:documentation/backlog/OPEN-ITEMS.md` |
|
||||
|
||||
@@ -140,7 +140,6 @@ the fault was real. Full observables: `tests/campaign11-evidence-2026-08-05/jour
|
||||
| **R-387** | **The hub REWRITES an unknown severity and says nothing, and the guard built to catch that sits downstream of the rewrite.** One handler, two fields, opposite discipline: an unknown `event_type` is rejected with a loud `400`, while an unknown `severity` was silently coerced to `info` — after which `severityNotifies` drops it and NEITHER delivery leg runs. **Two shipped features went out that way**: `DiskAlertKind.Severity` emitted `"warn"` until controller v0.215.0, `app_start_failed` until v0.223.0. **Measured on the live hub DB 2026-08-23: 91 `app_start_failed` events stored all-time and ZERO `notification_log` rows before that day** — not one, on any channel, while every POST returned 200. **The dispatcher's `unrecognized severity` line could never execute** for an API event, because the coercion one line upstream guarantees the value it looks for cannot arrive. | **CLOSED — hub v0.107.0, 2026-08-23** | — | **The coercion STAYS; only the silence is fixed** — a rejected event is a LOST event, and losing an alarm is worse than mis-routing one. A `WARN` now names the customer, the event type, the rejected value and the consequence. **The dispatcher branch was KEPT, on evidence not caution:** `cmd/hub/main.go` wires `dispatcher.ProcessEvent` DIRECTLY as the `monitor.EventNotifyFunc` for the staleness, host-staleness and offsite-box checkers, which never pass through the handler — for them it is the only severity guard there is; deleting it as "dead" would have removed the live half while the dead half supplied the justification. All 90 severity literals in `internal/monitor` verified already valid. Proven live: `[WARN] [api] Event from demo-hp: severity "warn" is not in {info,warning,error,critical}…`, with an `error` control silent. Evidence: `audits/DRILL-r329-r386-2026-08-23/evidence/live-19-scenarioH-after.txt`. | CC |
|
||||
| **R-391** | **Gate 11 (observations) is registered in three of the four runners; `app-catalog-felhom.eu` is the exception.** The controller and agent runners already carried a shared-gate mechanism (`SHARED_REUSE`, `SHARED_INSTRUCTIONS` pointing into `felhom.eu/scripts/`), so registering there was one constant and one `GATES` line each. **`catalog_gates.py` has no such mechanism:** its `run_gate` joins every entry against its OWN `scripts/` directory, so it cannot invoke a sibling repo's script at all; and its loop appends `--all` to every gate unconditionally, which the observations gate would read as a path. Registering there therefore needs `run_gate`'s contract widened AND the argument handling changed — a refactor of a runner whose shape is deliberately different (per-app scoping, network/runtime gates excluded from `--fast`), in a repo this task marked out of scope. **The exposure today is nil** — `app-catalog-felhom.eu/REPORT.md` has no observations section, and the gate passes quietly on that — but a future catalog session could write one and nothing would read it. **Filed rather than left as a sentence in a report, which is the exact failure R-389 records.** | **OPEN — LOW** | — | Either give `catalog_gates.py` the `SHARED_*` absolute-path mechanism the other two runners already have and stop appending `--all` to gates that do not take it, or state in that repo's CLAUDE.md that its REPORT.md carries no observations section by convention. **Do not copy the gate script** — the shared checker lives in ONE place (`felhom.eu/scripts/`) and copying it is the drift the shared pattern exists to prevent. | CC |
|
||||
| **R-390** | **The golden-bake runbook omits `pveam update`, and the failure it produces names the wrong cause.** `documentation/runbooks/RUNBOOK-manual-build.md` §4.1 step 2 says to list the current Debian template because "the exact point release rots" — but on the drill VM's `virgin` snapshot **the `pveam` INDEX is stale too**, so `pveam available` offers an old point release and `pveam download local <that>` fails with **`400 Parameter verification failed. template: no such template`**. That reads as a typo or a bad argument, not as an old index, and it costs a diagnosis every time. **Hit on two consecutive bakes** (golden 0.222.0 and 0.223.0, both 2026-08-23). The runbook is otherwise correct verbatim — the qemu launch line, the token-read-inside-the-VM pattern and the acceptance markers all worked unchanged. | **OPEN — LOW** | — | Add `pveam update` as its own numbered step before the listing, and say WHY: a snapshot that never changes carries an index that never updates, so the rot warning already in the step applies to the index as well as to the release. Recorded meanwhile in the workspace memory `golden-bake-needs-pveam-update` and in `documentation/tests/golden-0.223.0-2026-08-23/README.md`. | CC |
|
||||
| **R-389** | **Only the FIRST broken app per hour reaches the operator — the cooldown key names the event type, not the app.** `dispatcher.go:337` builds the operator key as `customerID + ":" + eventType + cooldownTierSuffix(details) + cooldownRunSuffix(details)`, and **neither suffix reads an app name**. So every app that goes down inside the same hour collapses onto one key and only the first is mailed. **Measured live on `demo-hp` 2026-08-23:** `bookstack` alarmed at 09:27:51 and was `sent`; `privatebin` alarmed at 09:31:51, four minutes later, and was logged `suppressed — operator cooldown 1h, key=demo-hp:app_start_failed`. Three apps down together tonight would produce one mail. **The app's identity is already on the wire** — `AppDetails{StackName, DisplayName}` serialises as `stack_name` (`felhom-controller/internal/notify/notifier.go:146-149`), and the hub already makes exactly this kind of distinction twice, with `cooldownTierSuffix` and `cooldownRunSuffix`. **This was latent for as long as the cooldown has existed and only became reachable when R-329 made `app_start_failed` deliverable at all** — the same "a known-broken thing moves from unreachable to load-bearing" shape as R-329 itself. **AND IT WAS NEVER FILED:** it was written in a REPORT.md observations paragraph on 2026-08-23 and nowhere else — the register had no row for it until now, which is R-341's shape one surface over and is why gate 11 exists. | **OPEN — MEDIUM** | — | A third sibling suffix, `cooldownStackSuffix`, **allow-listed to `app_start_failed` and nothing else**. **Do NOT apply it globally:** the backup family's cooldown is coarse ON PURPOSE (R-97a, R-182) so that one full disk sends one digest rather than one mail per app — and `crossdrive_failed` (severity `error`) carries `stack_name` through a *different* struct (`CrossDriveDetails`), so a global suffix would silently split it per-app. The hour itself does not change; the grain is the complaint, not the length. Evidence: `audits/DRILL-cooldown-grain-2026-08-23/`. | CC |
|
||||
| **R-388** | **PRODUCT DECISION (not a defect): the customer notification model is the wrong shape, and the settings page grows by one toggle per detector.** The operator's framing, recorded verbatim 2026-08-23: *"A customer should be notified only about things they can act on or are responsible for — the drive they unplugged, the storage they filled. **A failed backup is our incident, not theirs.** The intended shape is that we detect it, we tell them we noticed and are dealing with it, and they are not handed an error they cannot solve. The subscription should feel like being looked after, not like being on call."* Today's page is the opposite shape — one switch per detector, and it **grew from 12 to 15 in a single session** (one new alarm plus two compound toggles split into four). That growth is the argument, not an aside: a page that grows per detector keeps asking a household to make engineering decisions. | **OPEN — DIRECTION, operator's call** | a decision on scope; nothing here is a bug | Recorded as a dated **[DESIGN — DIRECTION]** entry at `documentation/architecture/08-alarm-ladder.md` §8, marked plainly as *not current behaviour*. **Deliberately NOT implemented in the session that recorded it.** `app_start_failed` defaulting OFF is consistent with the direction and reversible either way, but was ruled on its own merits and does not pre-judge the redesign. | Viktor |
|
||||
| **R-229** | **The instruction-file rightsizing landed for `felhom-controller` and the workspace root; three pieces were deliberately deferred.** Done 2026-08-06: controller split into a 92-effective-line core plus four `paths:`-scoped `.claude/rules/*.md`; workspace root 208→142 effective lines with its versioned copy kept byte-identical; surgical corrections to `felhom-agent` and `felhom.eu` (expired TEMPORARY block, every version literal, the Legacy-Windows copies, the duplicated health-check rule); five contradictions resolved — including a drill-VM claim **measured live** (`qm list` on demo-hp shows VM 300 `drill-r50`; `felhom-agent` was right, `felhom-controller` was wrong); new shared `felhom.eu/scripts/instructions_gate.py` registered in `controller_gates.py` and `agent_gates.py`, 20 fixture tests + red-proof. **Leg (a) CLOSED 2026-08-06 (part 2):** `felhom.eu/CLAUDE.md` **227 → 115 effective lines**, split into a core plus `.claude/rules/{hub,website,manifests,docs}.md`; `instructions_gate` **registered in `scripts/repo_gates.py`** (six gates, all OK) in the required order — trim first, register second, because a registered-but-failing gate refuses every push. Scoping proven from the `InstructionsLoaded` hook log in two fresh sessions, not from frontmatter. **Still deferred:** (b) **CLOSED 2026-08-06 (close-out)** — `felhom-agent/CLAUDE.md` **175 → 99 effective lines** (measured 175, not 173: the CI correction added two), split into a core plus `.claude/rules/{proxmox,localapi,backup,storage}.md` beside the existing `health-checks.md`. The release section now points at the `felhom-build-deploy` skill instead of restating a table that drifts from the script. **Every `CLAUDE.md` in the workspace is now ≤120 effective lines except the workspace root at 142, which is deliberate — it is the only file re-injected after `/compact`.** (c) **CLOSED 2026-08-06 (part 2)** — all 44 orphans resolved with **zero deletions** (file count 158 before and after): 4 durable `reference`-type files indexed, 40 dated episode records moved to `.claude-memory/archive/`. `MEMORY.md` 145 → **150 lines / 17,977 bytes**, and `instructions_gate` check 6 now watches it (over-limit FAILS, orphan WARNS, absent store PASSES *printing its reason*). (d) **The spec-as-failing-test pilot** — moved to R-230. Full accounting: `audits/LEDGER-instruction-trim-2026-08-06.md` + `audits/LEDGER-instruction-trim-part2-2026-08-06.md` | **READY** — owner Viktor |
|
||||
| **R-230** | **Three instruction/memory follow-ups deliberately left by the part-2 session (2026-08-06), each needing a decision rather than an implementation.** (a) **A ruling is owed on auto-written staleness.** The hand-written `CLAUDE.md` files are now clean of version literals and expired blocks — the gate enforces it — but `MEMORY.md`, which Claude writes and which is the LARGER half of what loads (8.4k tokens vs the root file's 6.6k), carries **21 lines with component version literals**, **5 with bare host addresses**, and an entry still reading *"demo boxes REMOTE till ~08-02"* — the same expired-TEMPORARY class the gate was built to kill, now surviving in the one file the gate's content rules do not cover. **Partly actioned 2026-08-06 (close-out), and the ruling is STILL OWED:** the **three statements that were actively false** were corrected — `R-193 decision open` (closed 2026-08-05), `demo boxes REMOTE till ~08-02` (the box answers on the home LAN), `OPEN R-25b` (shipped 2026-07-21) — and gate check 6 now **WARNs** on version literals, host addresses, expired statements and stale-open citations in the index. WARN, never FAIL: Claude writes that file between sessions, so a hard failure would refuse a human's push over a line no human typed, and the warning is read by the model that will next edit it. **The remaining 32 version literals and 4 host addresses were deliberately left** for that loop. What is still owed is the bulk-correction ruling. **Correcting the premise:** the earlier report's "three expired statements" were all FALSE POSITIVES — each matched an ISO date inside a markdown link target, i.e. a filename — while the one real expired claim carried no ISO date at all. (b) **CLOSED 2026-08-06 (close-out)** — the workspace-root `CLAUDE.md` **is now a relative symlink** to the versioned copy, so the divergence class is gone rather than policed. Check 5 learned two shapes: for a link it asserts the target resolves to a real file (**a dangling link is worse than a diverged copy — the instructions load NOTHING and there is no content left to notice is wrong**), for two files byte-identity as before, so a clone elsewhere is unaffected. **Proven, not assumed:** three fresh sessions logged `session_start` for the link path, and a fourth **with no tools at all** quoted standing rule 1 verbatim — the content reaches the model, not just the path. (c) **The spec-as-failing-test pilot**, approved in principle and not started (was R-229(d)). | **READY** — owner Viktor |
|
||||
|
||||
Reference in New Issue
Block a user