R-330: stop the backup alarming about the apps it is holding down (v0.224.0)
gates / gates (push) Successful in 11s

Measured live on demo-hp 2026-08-30 (controller 0.223.0): the nightly db-dump
and offbox-backup legs stop each stack ~13s to tar its volumes while the
deadapp-check job scans every 30s, so the scan caught whichever stack was
mid-cycle and pushed app_start_failed to the customer. 61 e-mails about apps
that were never broken.

The defect is not a missing mechanism. quiesce/suppress.go solved exactly this
in v0.179.0 and works -- but classifyRunStates read only the quiesce loop's set,
and that loop covers the WHOLE-GUEST backup. The per-app legs stop stacks
through Manager.DumpAppVolumesSafe, which registered with nothing. Two
mechanisms stop apps on purpose; only one told the alarm. Fifth instance of the
"seam built but never wired" class, and the first where the unwired half was a
consumer.

The suppression now rides AppStopGuard, which already brackets every deliberate
stop in the product (Begin before the stop, End after a successful restart) at
all three call sites, and which main.go hands as ONE object to the backup
manager and the exporter. scanDeployedAppRunStates takes the union of both sets.
All three per-app stop paths are covered, not only the reported nightly one.

It cannot latch -- End() runs only on a restart that SUCCEEDED, so unlike the
quiesce loop an open-ended hold is a real hazard here:
  1. ReleaseFailed drops the entry IMMEDIATELY on a restart that broke, wired at
     every failure path, so the app alarms on the next scan;
  2. Begin REPLACES the set (one marker file = one operation);
  3. appStopMaxHold (6h) caps a hold nothing released, logged at WARN.
Grace is 180s, deliberately quiesce's own constant and derivation. Suppression
is NOT persisted: after a crash the guard holds nothing and a down app must
alarm. ReleaseFailed keeps the durable crash marker; a test pins that.

Three companion red-proofs, each printing the pre-fix value (REPORT.md section 5):
  - drop markStopped from Begin      -> "suppressed at stop = map[]"
  - drop ReleaseFailed from the dump -> "map[bookstack:true] after a restart that FAILED"
  - pass nil instead of appStopGuard -> the AST wiring test fails
The third is load-bearing: the component was never the broken part, so a suite
that only injected it would have been green against the shipped defect.

Green gate clean: go build + go vet + go test ./... -- 28 packages, rc 0.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LB8FmJaGd2cyjvy6dbEjpM
This commit is contained in:
2026-08-30 17:58:15 +02:00
parent f8c9390946
commit 92cebb8c95
12 changed files with 811 additions and 246 deletions
+139 -238
View File
@@ -1,278 +1,179 @@
# REPORT — controller v0.223.0 (R-329, R-386, and the compound toggles)
# REPORT — R-330: the nightly backup alarmed about the apps it was holding down
**Session 2026-08-23, UNATTENDED.** Live leg on `demo-hp` (Tier 0), guest 9201.
**No halt condition fired.** Nothing was dropped.
**Controller v0.224.0 · 2026-08-30 · implemented on DooPlex, diagnosed live on `demo-hp`**
## 1. Baselines, and the hub's four numbers as read
This report covers Option A of a two-part request. Option B (the hub Backup card that always reads
`Snapshots 0`) is a separate, still-open defect and is described in §7.
| Repo | at start | at end |
---
## 1. What was reported
61 e-mails, arriving in two bursts every night from both demo boxes:
```
[Felhom] demo-hp: app_start_failed
Severity: warning
Time: 2026-08-30 02:30 CEST
Message: Telepített alkalmazás nem fut: Docmost
```
## 2. What was actually happening
**Nothing was broken.** Both boxes were healthy at every check:
| | `demo-felhom` (N100) | `demo-hp` (HP t740) |
|---|---|---|
| felhom-controller | `14137efa` (v0.222.0) | **v0.223.0** deployed |
| felhom.eu | `55274d5e` (hub v0.106.0) | **hub v0.107.0** deployed |
| felhom-agent | `40d857b5` (v0.130.0) | untouched |
| host uptime | 20 d | 8 d 23 h |
| agent | 0.130.0, `active` | 0.130.0, `active` |
| controller | 0.223.0 | 0.223.0 |
| containers | 1/1 | **16/16, all `healthy`** |
| SMART, all disks | PASSED | PASSED |
| `journalctl -u felhom-agent -p warning`, 3 days | no entries | no entries |
| hub health | `ok` | `ok` |
**Hub's four numbers, live from `GET /configuration` before starting:** `golden_version` **0.222.0**,
`agent_version` **0.130.0**, `min_agent` **0.129.0**, controller floor **0.222.0** — all four as the
task predicted.
The alarms are the box's own backup. Both bursts line up exactly with the nightly legs
(`backupwindow`: DB dump at W, tier-2 at W+60m, off-box at W+105m; default W = 02:30 CEST):
## 2. Documents read
`internal/notify/notifier.go:583-602` (`Severity()`'s doc comment — **it already stated the entire
contract and named both hub locations**), `hub/internal/api/handler.go` (the type rejection and the
severity coercion side by side), `hub/internal/notify/dispatcher.go` (`severityNotifies`,
`ProcessEvent`, `operatorOnlyEvents`, `processOperator`), `internal/stacks/deploy.go` (the
`DesiredState` comment and its one-owner rule), `cmd/controller/main.go` (`classifyRunStates`),
`internal/web/handlers.go` (the compound toggles).
**§2's conditional is answered: the alarm ladder DOES exist** —
`felhom.eu/documentation/architecture/08-alarm-ladder.md`, written last session. It has been extended
here with §6.1 (the severity contract), §7 (the intent test) and §8 (Part 5's direction).
## 3. The 1.1 sweep — the result in full
**Exactly ONE bad severity in the whole controller: `notifier.go:546`, `"warn"`.** Nothing else.
Verified across Go **and** templates **and** queued-event construction, because the task warned that
reading zero from Go files while the answer sat in a template has produced three wrong conclusions
here:
- every `emit(...)` / `PushEvent(...)` literal — one offender, the rest valid;
- `internal/channelhealth`'s classifier (the source of `NotifyAgentChannelDown`'s variable) — all
`"warning"`/`"error"`;
- `debug.html`'s operator-triggerable severity `<select>` — offers only `error`/`warning`/`info`;
- no `PendingEvent{}` literal construction exists anywhere.
Nine other `"warn"` strings exist and are **not** defects: `internal/monitor` and `internal/selftest`
use it as a *healthcheck status* vocabulary, and `statusRank`'s `case "warn"` maps it to the correct
`"warning"` severity. **This is why the guard is an AST walk and not grep.**
**One latent hazard found while sweeping, pinned rather than left:** `fillwatch.Band.Severity()`
returns `""` for `BandOK`. Unreachable, because `Check()` notifies only on an escalation — but that
safety lives in a *different function* from the one that looks unsafe, so the test asserts the
consequence.
**And the guard found two dynamic call sites the hand sweep missed** (`NotifyDRCompleted`, and
fillwatch's via `main.go`). All six are now registered by name with the values each can take.
## 4. The hub's manifest — §6's premise was wrong
**`felhom.eu/manifests/hub.yaml`, line 128.** ArgoCD `Application/felhom` tracks
`admin/felhom.eu.git` path `manifests`, `automated.enabled=false`. Bumped in **`68a9f54`**. **No
out-of-git deployment path exists.** The image was pushed to the registry *before* the manifest landed,
so a sync could never point at a missing tag; the sync was then requested deliberately. Never
`kubectl set image`.
## 5. Files, commits, CI
| Commit | Repo | Contents |
|---|---|---|
| **`9832760`** | controller | v0.223.0 — the severity word, the AST guard, the toggle, the intent test, the toggle split |
| **`<docs>`** | controller | REPORT / CONTEXT / README |
| **`68a9f54`** | felhom.eu | hub v0.107.0 + manifest bump + golden evidence |
| **`2f7c9a6`** | felhom.eu | alarm ladder, register, STATUS, drill record |
Modified: `internal/notify/notifier.go`, `cmd/controller/main.go`, `internal/web/handlers.go`,
`internal/web/templates/settings_notifications.html`, `CHANGELOG.md`.
Added: `internal/notify/r329_severity_contract_test.go`, `internal/fillwatch/r329_severity_test.go`,
`cmd/controller/r386_intent_test.go`, `internal/web/r329_toggle_split_test.go`.
**CI runs confirmed BY ID** (`id` and `run_number` diverge — both printed) — see §15 below.
## 6. Red-proofs — five planted, and ONE PASSED FIRST TIME
| # | Mutation | Layer, and why that layer | Observed |
| leg | fired (UTC) | = CEST | events pushed |
|---|---|---|---|
| 1 | severity back to `"warn"` | **the EMITTER** — the last point at which the bad value still exists; the hub deliberately destroys it one line later | `notifier.go:561:30: emit(...) emits severity "warn", which is NOT in the hub's vocabulary` |
| 2 | `userStopped` back to the state guess | **`classifyRunStates`** — the single derivation point where the guess was made | `dead-app banner = [], want exactly privatebin` — the exact live symptom — plus `IntentUnknown = false` and the wiring check |
| 3 | no-op-save guard removed | **the SAVE handler** — where a render-then-save rewrites stored bytes | `a no-op save CHANGED the stored settings` on the `defaults` shape |
| 4 | hub ingest `WARN` removed | **INGEST** — the last point the offending value exists | `the hub rewrote a severity and said nothing`, log showing only `[INFO] … (info)` |
| 5 | fillwatch de-escalation guard removed | **`Check()`** — the invariant lives there, not in `Severity()` | `notified with band ok → severity "" (event type "")` |
| `db-dump` | 00:30 | 02:30 | Docmost, Paperless-ngx, RomM |
| `tier2-backup` | 01:30 | 03:30 | — |
| `offbox-backup` | 02:15 | 04:15 | Docmost, Paperless-ngx |
**Red-proof 5 passed on the first attempt and that is reported, not omitted.** The first mutation —
`if next <= prev` → `if next < prev` — is **inert**: an earlier `if next == prev { continue }` had
already removed the equal case, so the code's behaviour did not change and the test was *right* to
pass. Removing the guard outright convicts it. **A red-proof that passes needs the mutation checked
before either verdict is believed.** Every other mutation asserted its pre-fix text was present before
rewriting and printed `MUTATION APPLIED`.
## 7. Test counts
| Repo | Before | After |
|---|---|---|
| felhom-controller | 1504 | **1522** |
| felhom.eu hub | 702 | **709** |
Full green gate `go build && go vet && go test ./...` → **exit 0, zero failures, both repos**. All 11
controller design gates OK; all 11 felhom.eu repo gates OK.
## 8. Deployed versions and the golden
Evidence, from the guest's own controller log (all copied off the box before any change):
```
gitea.dooplex.hu/admin/felhom-controller:0.223.0 Up 18 minutes (healthy)
gitea.dooplex.hu/admin/felhom-hub:0.107.0 Synced, rollout complete
00:30:01 backup.go:868 [INFO] [backup] Stopping bookstack for safe volume dump
00:30:06 manager.go:1129 [INFO] [stacks] Stack bookstack stopped successfully (took 4.4s)
00:30:14 manager.go:1054 [INFO] [stacks] Stack bookstack started successfully (took 6.2s)
00:30:26 notifier.go:234 [INFO] Event pushed: app_start_failed (warning) — Telepített alkalmazás nem fut: Docmost
```
**Golden BAKED and PUBLISHED: YES — 0.223.0.**
`sha256 9eaf39ac39219b42ec9e6cbf890275febcdcc6f53325fe0c0f591d3431044f17`, `upload OK (HTTP 201)`,
round-trip **HTTP 206**, all five acceptance markers counted.
**VOUCHING IS THE OPERATOR'S ACT AND WAS NOT DONE HERE.**
`DumpAppVolumesSafe` stops a stack (`docker compose down`), tars its volumes and starts it again —
**~13 s per stack, measured** — while the `deadapp-check` scheduler job runs every **30 s**. The scan
caught whichever stack was mid-cycle.
## 9. The live walk, all six steps
**The positive observable that proves the boxes were fine** (standing rule 3 — an absent alarm is not
evidence): every dead-app scan across the other 23 hours logged `8 deployed app(s) evaluated,
0 currently down`, and the same night's off-site run logged
`[offbox] backup OK: 8 app(s) backed up, 67 snapshot(s), 2m14s`.
### Step 1 — Scenario A: a database dies, customer has not opted in ✅
## 3. Root cause
**Quoted from the hub's own records, not the controller's:**
**A mechanism that exists, works, and was never consulted.** `quiesce/suppress.go` solved exactly this
problem in v0.179.0 (R-97b). But `classifyRunStates` read only `quiesce.Loop.SuppressedStacks()`, and
the quiesce loop covers the **whole-guest** (vzdump/PBS) backup. The **per-app** legs stop stacks
through `Manager.DumpAppVolumesSafe`, which registered with no suppressor at all.
```
events demo-hp app_start_failed warning 2026-08-23 09:27:51 <- v0.223.0
demo-hp app_start_failed info 2026-08-23 05:30:14 <- v0.222.0, coerced
notification_log demo-hp app_start_failed warning sent operator 2026-08-23 09:27:51
(no customer row)
```
Two mechanisms in this product stop a customer's app on purpose. Only one told the alarm.
**The number that says it all: 91 `app_start_failed` events stored all-time, ZERO `notification_log`
rows before 09:00 today.** Not one, ever, on any channel.
This is the **"seam built but never wired"** class, fifth instance — and the first where the unwired
half was a *consumer* rather than a producer. That distinction is why the tests below include an AST
wiring check: a test that injected the suppressor directly would have passed against the shipped bug.
**An honest limit, stated rather than implied:** this run does **not** prove the customer *gate*.
`demo-hp` has no `customer_notifications` row at all, so the customer leg could not have delivered
regardless. The gate is proven by the unit tests, which configure prefs both ways — and Scenario B
there is the positive control showing the customer leg *can* deliver for this event type.
## 4. The fix
### Step 2 — Scenario D: stopped out of band, intent `running` ✅
`internal/backup/appstop_suppress.go` (new) puts the suppression on **`AppStopGuard`**, which already
brackets every deliberate stop in the product (`Begin` before the stop, `End` after a successful
restart) at all three call sites — volume dump, off-site reconstitute, `.fab` export — and which
`main.go` hands as ONE object to both the backup manager and the exporter. The fact the alarm needs
already lived there with exactly one writer; a fourth registry beside it would have been drift.
`docker compose stop privatebin` at 09:31:27Z → **`app_start_failed (warning)` at 09:31:51Z, 24
seconds later.** Heartbeat, new beside last session's:
`scanDeployedAppRunStates` now passes `unionSuppressed(q.SuppressedStacks(), g.SuppressedStacks())`.
**All three per-app stop paths are fixed by the one change**, not only the nightly leg that was
reported — a `.fab` export and an off-site restore stop an app the same way and would alarm the same
way.
```
2026/08/23 05:47:44 [deadapp] check alive: 40 scans since boot, 8 deployed app(s) evaluated, 0 currently down <- v0.222.0
2026/08/23 09:34:51 [deadapp] check alive: 20 scans since boot, 8 deployed app(s) evaluated, 1 currently down <- v0.223.0
```
### It must never latch — the harder half
Hub-side the second alarm was logged `suppressed — operator cooldown 1h` (see §15.1).
Permanent suppression trades a loud false alarm for a silent real one (F-CRIT-1, R-88 Scenario D).
`End()` runs **only on a restart that succeeded**, so an open-ended hold is a genuine hazard here in a
way it is not for the quiesce loop, which always releases. Three independent guards:
### Step 3 — Scenario C: the customer presses Stop ✅
1. **`ReleaseFailed`** — a restart attempted and broken drops the entry **immediately**; the app
alarms on the very next scan, with no delay at all. Wired at every failure path (volume dump,
`restartStack` in the off-site reconstitution, the exporter's restart defer — the exporter's seam
interface grew the method rather than the exporter keeping separate bookkeeping).
2. **`Begin` replaces the set** — one marker file is one operation, so a set stranded by an operation
that died mid-window cannot survive into a later one.
3. **`appStopMaxHold` = 6 h** — a backstop for a hold nothing released, logged at WARN when it fires.
Through `POST /api/stacks/privatebin/stop`, the exact call the button makes. Intent moved
`running → stopped`. Four minutes and **9 dead-app scans later: 0 alarms, 0 unknown-intent lines.**
The post-restart grace is **180 s, the same constant and derivation as `quiesce.quiesceAlarmGrace`**.
Two windows over one alarm that disagreed on how long a restart takes would be a bug waiting to be
found on whichever path used the shorter one.
### Step 4 — Scenario E: intent absent ✅ (and the count)
The suppression is **deliberately not persisted**: after a crash the guard holds nothing, `Recover()`
either brings the apps back or leaves them genuinely down, and a down app must alarm. `ReleaseFailed`
drops the suppression and **keeps** the durable crash marker — the two are independent, and a test
pins that.
demo-hp has **zero** such apps, so one was created: `desired_state` removed from `privatebin`'s
`app.yaml`, as a pre-R-166 box would look (backed up, restored, no code writer added). Result —
evaluated (`deployed=True state=stopped`, 8 apps evaluated), **suppressed (0 alarms)**, and:
## 5. Tests, and the red-proofs
```
[deadapp] 1 stopped app(s) have NO recorded customer intent, so their dead-app alarm is suppressed
by the unknown-intent fallback (R-386): privatebin. This closes itself as each app is started or
stopped through the interface.
```
`internal/backup/appstop_suppress_test.go` drives the **real** `DumpAppVolumesSafe` (not `Begin`
directly) and asserts the suppression set the dead-app scanner actually reads — the consequence, not
a log line.
### Step 5 — Scenario G: save changing nothing ✅ byte-identical, **on the third attempt**
| test | pins |
|---|---|
| `TestVolumeDump_SuppressesTheAlarmForTheAppItIsHolding` | the fix, through the production path |
| `TestVolumeDump_SuppressionExpiresSoARealOutageStillAlarms` | the window is a bounded delay, never a lost alarm |
| `TestVolumeDump_FailedRestartAlarmsImmediately` | anti-latch #1, and that the crash marker survives |
| `TestSuppression_CannotOutliveTheBackstop` | anti-latch #3 |
| `TestBeginReplacesThePreviousOperationsSet` | anti-latch #2 |
| `TestSuppression_FailedBeginSuppressesNothing` | a refused stop suppresses nothing |
| `TestHeldAppIsNotReportedDown` | the consequence: held app silent, genuinely exited app still alarms |
| `TestScanDeployedAppRunStatesIsGivenTheAppStopGuard` | the wiring, by AST walk |
`sha256(enabled_events)` = `10840f3a95bac168f0d7c79760998138` **before and after** a real save
(7 boxes ticked, refusal banner absent).
**Three companion red-proofs were run, and each printed the pre-fix value:**
**The first two attempts were silently REFUSED behind an HTTP 200** — the empty-email wipe guard
declines and renders an error page, still `200`. Run 1: no prefs existed. Run 2: the email `<input>`
spans **three lines**, so a single-line grep read it as `""`. **Both times the hashes matched —
because nothing was saved, not because nothing changed.** Fixed by asserting the refusal banner is
absent. *A warning beside a success is read as a success.*
1. Delete `g.markStopped(stackNames)` from `Begin` →
`suppressed at stop = map[], want bookstack` — the exact shape that produced the e-mails.
(`TestVolumeDump_SuppressionExpires…` and `TestSuppression_CannotOutlive…` also went red.)
2. Delete `m.appStop.ReleaseFailed(stackName)` from the volume dump →
`suppressed = map[bookstack:true] after a restart that FAILED`.
3. Pass `nil` instead of `appStopGuard` in `main.go` → the AST wiring test failed.
### Step 6 — Scenario H: a bad severity to the hub ✅
Each was restored immediately and `git diff` verified clean afterwards. Red-proof 3 is the load-bearing
one: the component was never the broken part, so a suite that only injected it would have been green
against the shipped defect.
```
[WARN] [api] Event from demo-hp: severity "warn" is not in {info,warning,error,critical} — coercing
to "info", which severityNotifies DROPS, so this backup_failed alert will reach NOBODY. Fix the
emitting controller; this event is stored but not routed.
```
**Green gate:** `go build ./... && go vet ./... && go test ./...` in `felhom-controller/controller/` —
clean, no failures.
Both the bad POST and an `error` control returned **200** (nothing lost); the control produced **no**
warning.
## 6. Not done, and why
## 10. The absent-intent count on `demo-hp`, in plain words
- **No live deploy yet.** The build/deploy of 0.224.0 to the two boxes is the next step and is
reported separately; this report covers the change and its unit-land proof only.
- **The `restore-hold` path** (`offbox_reconstitute.go`, an app deliberately held down after a failed
replay) calls `End()`, so it gets the 180 s grace and then alarms. That is **today's behaviour plus
180 s** and is deliberate: the app really is down, the customer should learn that, and the hold has
its own operator notification (`restoreHoldNotify`) besides.
- **`HeldStacks()` was left alone.** It reads the marker from disk for the boot reconciler and covers
"held right now" but not the post-restart grace — which is precisely the window R-97b proved is
needed. The new in-memory set is a superset for alarm purposes; the durable one stays the recovery
record.
**Zero.** All **8** deployed apps carry `desired_state: running`; none is `stopped` and none is absent.
The unknown-intent fallback therefore suppresses nothing on this box today — the population is already
empty on a machine that has been exercised through the interface. It will be larger on a box upgraded
and left alone, which is why the log line exists rather than a one-off count.
## 7. Found while diagnosing — a second, still-open defect (Option B)
## 11. The dead-branch decision, and the reason
**The hub's customer Backup card is inert for every customer.** It reads
`Snapshots 0 · Repo Size 0 MB · Integrity Unknown` while the same box's log says
`[offbox] backup OK: 8 app(s) backed up, 67 snapshot(s)`.
**KEPT.** `cmd/hub/main.go` wires `dispatcher.ProcessEvent` **directly** as the
`monitor.EventNotifyFunc` for the staleness, host-staleness and offsite-box checkers — those events
never pass the ingest handler, so for them that line is the only severity guard there is. Deleting it
as "dead" would have removed the live half while the dead half supplied the justification. All 90
severity literals in `internal/monitor` were verified already valid, so the guard is silent because
the producers are correct.
`hub/internal/web/templates/customer_unified.html:284` renders `snapshot_count` / `repo_size_mb` /
`integrity_ok` from the report. Those fields are **declared** in
`controller/internal/report/types.go:105-108` and **assigned nowhere** — a repo-wide grep finds only
the declaration. The controller tracks the real numbers in `internal/backup/offbox.go` (`SnapshotCount`,
line 1043) and serves them on its own API (`internal/web/offbox_handlers.go:295`); they are simply
never copied into the hub report.
## 12. Evidence
**Until it is fixed, that card must not be read as evidence of a missing backup.** It is a two-repo
change (controller report builder + hub) and is Option B of this request.
`felhom.eu/documentation/audits/DRILL-r329-r386-2026-08-23/evidence/` — 26 files: 5 red-proof
transcripts, 20 live-walk files, two full controller-log windows (1041 and 4044 lines) **pulled off
before each revert**, and the hub's own DB queries.
## 8. Also observed (not changed)
## 13. Teardown, three layers, and the end state
1. **Guest 9201 / apps** — nothing provisioned. **All 17 app containers healthy.** `privatebin`'s
`app.yaml` restored from backup and the backup deleted; intent reads `running`. `demo-hp`'s
notification settings restored to `enabled_events: null`, no e-mail — their pre-drill state. No app
rebuilt, redeployed or restored; planted data untouched.
2. **Bake VM** — powered off, Gitea token and runner script **shredded**, `drill.qcow2` reverted to
`virgin`. Build guest 9100 exists only inside that reverted snapshot. No storage added anywhere, so
`pvesm status` has nothing to compare.
3. **Hub-side, stated explicitly.** The hub was **written** this session, unlike last: the deployment
is v0.107.0 via the manifest, and **two probe events remain as rows for `demo-hp`** from Scenario H
(`backup_failed`, "R-387 scenario H probe" and "…control"). They are inert records; named here
rather than left for someone to find. Nothing else: no appliance registered, no customer created,
no artifact manifest changed, floor untouched.
**End state:** controller **0.223.0** and hub **0.107.0** deployed; golden **0.223.0** baked and
published but **NOT vouched**; floor still **0.222.0**; all apps running; planted data present.
## 14. Register size
| File | Before | After |
|---|---|---|
| `OPEN-ITEMS.md` | 328,325 B | **328,132 B** |
| `CLOSED-ITEMS.md` | 71,441 B | **74,642 B** |
R-329 and R-386 closed and compressed; **R-387** (closed) and **R-388** (the notification-model
product decision — open, operator's call) filed.
## 15. Observations
> **Each item carries `FILED: R-NNN` or `NOT-A-FINDING: <reason>` (gate 11, added 2026-08-23).**
> The markers were added retrospectively on 2026-08-24 when the gate was built: item 1 was the very
> finding that had no row, and adding its marker is the first thing the gate ever asked for. The
> observations' text is unchanged.
1. **The operator cooldown key has no app identifier, and it now bites.** PrivateBin's alarm four
minutes after BookStack's was logged `suppressed — operator cooldown 1h,
key=demo-hp:app_start_failed`, so **only the first app-down per hour e-mails the operator**. This is
R-182's known cooldown-key shape; it was harmless while `app_start_failed` was undeliverable and is
not any more. **Same pattern as R-329 itself: a known-broken thing moved from unreachable to
load-bearing.** Not fixed here.
FILED: R-389
2. **The settings page grew 12 → 15 toggles in one session** — one new alarm, plus two compound
toggles split into four. Recorded as the argument inside R-388.
FILED: R-388
3. `internal/notify/notifier.go` carries **pre-existing** gofmt drift in an unrelated const block,
confirmed by stashing this session's work and re-running `gofmt -l`. Not touched (§12).
NOT-A-FINDING: cosmetic drift in an unrelated const block, predating this work and touching no
behaviour; §12 forbids nearby refactors, so filing a row would only queue a whitespace commit.
4. **The golden-bake runbook still lacks `pveam update`** — second consecutive bake to hit the stale
index on the `virgin` snapshot, presenting as `400 … no such template`.
FILED: R-390
5. **Deliberately left open, untouched:** R-102, R-359, R-385, and R-388's redesign.
NOT-A-FINDING: this item is a pointer to rows that already exist, not a new observation; it is
here so their absence from this session reads as deliberate rather than forgotten.
### CI runs, confirmed by ID
| Commit | Repo | CI `id` | `run_number` | Result |
|---|---|---|---|---|
| `9832760` | felhom-controller | **408** | 88 | success |
| `68a9f54` | felhom.eu | **409** | 261 | success |
| `2f7c9a6` | felhom.eu | **410** | 262 | success |
(The controller docs commit's run is confirmed after its push and is the next `id` in that repo.)
- **`ssh demo-hp` no longer works** — the tailnet peer `100.76.96.79` has been offline 8 days
(`tailscale status`: `offline, last seen 8d ago`). The box is reachable on the home LAN as
`ssh hp` → `192.168.0.104`, which is what every command in this report used.
- **`drill-r50-0a4f9a` still appears in the hub host list as `DOWN`** — a leftover record from the
R-50 drill whose rig was torn down; not a live box.