2f7c9a6ce5
gates / gates (push) Successful in 17s
The alarm ladder gains the severity contract (the hub's vocabulary is exact, it coerces silently, and three things now hold it) and the intent test with its three-way ruling on unknown. Both marked [DESIGN] with the live measurements. Part 5 is RECORDED AND NOT IMPLEMENTED: the operator's notification philosophy, verbatim, marked plainly as direction rather than current behaviour, with the 12 -> 15 toggle growth as the argument. Filed as R-388, a product decision. R-329 and R-386 compressed into CLOSED-ITEMS with their rules kept and the full-text commit named. R-387 filed closed - including WHY the dispatcher branch was kept rather than deleted, which is evidence (three monitor checkers call ProcessEvent directly) and not caution. The drill record names three things that had to be re-run: an inert red-proof mutation, Scenario G refused twice behind an HTTP 200, and the live Scenario A NOT proving the customer gate because demo-hp has no prefs row at all. Register: OPEN 328325 -> 328132 B, CLOSED 71441 -> 74642 B.
111 lines
6.3 KiB
Markdown
111 lines
6.3 KiB
Markdown
# DRILL — R-329 + R-386: the alarm that fired but reached nobody, and the stop nobody heard (2026-08-23)
|
|
|
|
**Controller v0.222.0 → v0.223.0. Hub v0.106.0 → v0.107.0. Live leg on `demo-hp` (Tier 0), guest
|
|
9201. UNATTENDED.** Method: endpoint-level plus the hub's own SQLite records — no browser exists on
|
|
DooPlex. Guest and hub clocks are UTC; the hub pod logs CEST.
|
|
|
|
## Verdict
|
|
|
|
| Part | Outcome |
|
|
|---|---|
|
|
| 1.1 the severity word + sweep | ✅ **exactly one** bad severity in the whole controller |
|
|
| 1.2 the AST contract guard | ✅ and it found two dynamic sites the hand sweep missed |
|
|
| 1.3 customer toggle, default OFF | ✅ |
|
|
| 2 the hub says what it rewrote | ✅ hub v0.107.0, proven live |
|
|
| 3 ask the field that knows | ✅ proven live, both directions |
|
|
| 4 split the compound toggles | ✅ round trip byte-identical |
|
|
| 5 record the notification philosophy | ✅ recorded, **not implemented** |
|
|
|
|
**No halt condition fired.**
|
|
|
|
## The one number that says it all
|
|
|
|
Read from the live hub DB:
|
|
|
|
```
|
|
91 app_start_failed events stored, all-time
|
|
0 notification_log rows before 2026-08-23 09:00 <- not one, ever, on any channel
|
|
```
|
|
|
|
After the fix, at 09:27:51: **one row — `warning` / `sent` / `operator`.**
|
|
|
|
## The two live pairs
|
|
|
|
**Scenario A** — a database dies, customer has not opted in:
|
|
|
|
```
|
|
events : demo-hp app_start_failed warning 2026-08-23 09:27:51 <- v0.223.0
|
|
demo-hp app_start_failed info 2026-08-23 05:30:14 <- v0.222.0, coerced
|
|
notification_log: demo-hp app_start_failed warning sent operator 09:27:51
|
|
(no customer row)
|
|
```
|
|
|
|
**Scenario D** — an app stopped out of band, intent `running`:
|
|
|
|
```
|
|
2026/08/23 05:47:44 [deadapp] check alive: 40 scans since boot, 8 deployed app(s) evaluated, 0 currently down <- v0.222.0
|
|
2026/08/23 09:34:51 [deadapp] check alive: 20 scans since boot, 8 deployed app(s) evaluated, 1 currently down <- v0.223.0
|
|
```
|
|
|
|
Alarm fired **24 seconds** after the `docker compose stop`.
|
|
|
|
## Three things that had to be re-run, and why that matters
|
|
|
|
1. **Red-proof 5 passed first time — the mutation was INERT.** Changing `if next <= prev` to
|
|
`if next < prev` in fillwatch does nothing, because an earlier `if next == prev { continue }` had
|
|
already removed the equal case. The test was right to pass. Removing the guard outright convicts
|
|
it. **Check the mutation applied before believing either verdict.**
|
|
2. **Scenario G was silently refused TWICE behind an HTTP 200.** The empty-email wipe guard declines
|
|
the save and renders an error page — still `200`. The first run had no prefs at all; the second
|
|
read the email with a single-line grep from a `<input>` that spans **three lines**, got `""`, and
|
|
was refused again. Both times the before/after hashes matched — *because nothing was saved*, not
|
|
because nothing changed. Fixed by asserting the refusal banner is **absent**. **A warning beside a
|
|
success is read as a success.**
|
|
3. **The live Scenario A does NOT prove the customer gate**, and is not claimed to. `demo-hp` has no
|
|
`customer_notifications` row at all, so the customer leg could not have delivered regardless. The
|
|
toggle gate is proven by the unit tests, which configure prefs both ways. Stated rather than
|
|
implied.
|
|
|
|
## Evidence index (`evidence/`)
|
|
|
|
| File | What it shows |
|
|
|---|---|
|
|
| `redproof-1-R329-emitter.txt` | severity back to `"warn"` → AST guard names file, line and value |
|
|
| `redproof-2-R386-intent.txt` | intent test reverted → `dead-app banner = []`, the live symptom |
|
|
| `redproof-3-part4-migration.txt` | no-op-save guard removed → the `defaults` case reorders |
|
|
| `redproof-4-R387-ingest.txt` | hub WARN removed → "the hub rewrote a severity and said nothing" |
|
|
| `redproof-5-fillwatch-consequence.txt` | the inert first attempt, and the effective one |
|
|
| `live-01`…`live-04` | Scenario A: stop, controller send, **hub records**, customer-leg control |
|
|
| `live-05-step4-absent-intent-count.txt` | **0 of 8** deployed apps carry an absent intent |
|
|
| `live-07`…`live-09` | Scenario D: out-of-band stop, alarm in 24 s, the heartbeat pair |
|
|
| `live-10`, `live-11` | Scenario C: UI Stop → intent `running`→`stopped`, 0 alarms / 9 scans |
|
|
| `live-12`, `live-13` | Scenario E: intent removed → suppressed **and** the log line names the app |
|
|
| `live-16-scenarioG-roundtrip.txt` | all three attempts, ending byte-identical (`10840f3a…`) |
|
|
| `live-18-golden-bake.txt` | golden 0.223.0 markers, each counted |
|
|
| `live-19-scenarioH-after.txt` | the new hub WARN line, live, with a silent `error` control |
|
|
| `live-06`, `live-14` | full controller-log windows (1041 and 4044 lines), pulled before each revert |
|
|
|
|
## Observations — noticed, recorded, NOT acted on
|
|
|
|
1. **The operator cooldown has no app identifier, and it now bites.** PrivateBin's event 4 minutes
|
|
after BookStack's was logged `suppressed — operator cooldown 1h, key=demo-hp:app_start_failed`. So
|
|
**only the first app-down per hour e-mails the operator.** This is R-182's known cooldown-key
|
|
shape; it was harmless while the event was undeliverable and is not any more. Recorded, not fixed.
|
|
2. **The prompt's §6 premise was wrong and is corrected:** the hub's manifest **is** in version
|
|
control, at `felhom.eu/manifests/hub.yaml:128`, and ArgoCD app `felhom` tracks
|
|
`admin/felhom.eu.git` path `manifests`. No out-of-git deployment path exists.
|
|
3. **The golden-bake runbook still lacks `pveam update`** — second bake in a row to hit the stale
|
|
index on the `virgin` snapshot, presenting as `400 … no such template`.
|
|
4. `internal/notify/notifier.go` carries **pre-existing** gofmt drift in a const block, confirmed by
|
|
stashing this session's work. Not touched (§12).
|
|
|
|
## Teardown
|
|
|
|
Nothing provisioned. Every app restarted and confirmed healthy (17 containers). `privatebin`'s
|
|
`app.yaml` restored from its backup and the backup removed; its intent reads `running` again.
|
|
`demo-hp`'s notification settings restored to `enabled_events: null`, no e-mail — the state they were
|
|
in before the drill. The drill VM is powered off, its Gitea token shredded, and `drill.qcow2` reverted
|
|
to `virgin`. **Hub-side: two probe events (`backup_failed`, "R-387 scenario H probe"/"control") were
|
|
POSTed to the live hub for Scenario H and remain as event rows for customer `demo-hp`.** They are
|
|
inert records; named here rather than left for someone to find.
|