docs(v0.223.0): REPORT, CONTEXT rulings, README severity contract
gates / gates (push) Successful in 11s

REPORT overwritten: the 1.1 sweep in full (one bad severity, nine legitimate
"warn" strings that are healthcheck statuses), the hub manifest's real location
since the task's premise was wrong, all five red-proofs with the layer each
guard sits at, the live walk in six steps with the hub's own records quoted, and
the absent-intent count (0 of 8).

Three things are reported that a tidier account would omit: red-proof 5 passed
first time because the mutation was INERT; Scenario G was silently refused twice
behind an HTTP 200; and the live Scenario A does NOT prove the customer gate,
because demo-hp has no prefs row at all.

CONTEXT records the severity vocabulary as a ruling with its mechanism, the
intent ruling with its three-way handling of unknown, both fences, and two traps
worth more than the fixes: a 200 can be a refusal, and a passing red-proof can
mean an inert mutation.

README: the event table said `app_start_failed | warn` - the defect, written
down as if correct. Now `warning`, with the vocabulary contract and who receives
what. `disk_critical` also corrected from `error` to `critical`, which is what
fillwatch has always sent.
This commit is contained in:
2026-08-23 12:06:45 +02:00
parent 9832760027
commit 1da2c9c6c6
3 changed files with 273 additions and 204 deletions
+200 -201
View File
@@ -1,267 +1,266 @@
# REPORT — controller v0.221.1 (record) + v0.222.0 (R-384, R-383)
# REPORT — controller v0.223.0 (R-329, R-386, and the compound toggles)
**Session 2026-08-23, UNATTENDED. Live leg on `demo-hp` (Tier 0, disposable).**
**Session 2026-08-23, UNATTENDED.** Live leg on `demo-hp` (Tier 0), guest 9201.
**No halt condition fired.** Nothing was dropped.
> ## ⚠ HALT DECLARED — §4's measurement found a real defect, filed as R-386, NOT fixed
>
> The task's §4 asked for a measurement and named it a halt condition. It reproduced.
> **A single-container app stopped out of band raises no alarm at all.** Details in §10 below.
> Per §13 the fix was NOT attempted here. Live-walk step 4 (Scenario E) was dropped as a
> consequence — it is item (3) on the task's own drop list. Everything else completed.
## 1. Baselines, and the hub's four numbers as read
---
## 1. Baselines used, and the hub's four numbers as read
| Repo | `main` at start | Verified |
| Repo | at start | at end |
|---|---|---|
| felhom-controller | `f7881787f434` | matches the task |
| felhom.eu | `1eb64bec5183` | task said `4e488321bfd1+`; it had moved on |
| felhom-agent | untouched | — |
| felhom-controller | `14137efa` (v0.222.0) | **v0.223.0** deployed |
| felhom.eu | `55274d5e` (hub v0.106.0) | **hub v0.107.0** deployed |
| felhom-agent | `40d857b5` (v0.130.0) | untouched |
**Hub's own numbers, read live from `/configuration` (ClusterIP + Basic auth) 2026-08-23:**
**Hub's four numbers, live from `GET /configuration` before starting:** `golden_version` **0.222.0**,
`agent_version` **0.130.0**, `min_agent` **0.129.0**, controller floor **0.222.0** — all four as the
task predicted.
| Field | Value |
|---|---|
| `golden_version` | **0.221.1** |
| `agent_version` | **0.130.0** |
| `min_agent` | **0.129.0** |
| controller floor (`min_controller_version`) | **0.221.1** |
## 2. Documents read
The task expected golden/floor **0.220.2**; the operator had already vouched **0.221.1** and raised
the floor. Live controller on `demo-hp` at session start: **0.221.1** — so the running version, the
golden and the floor all agreed, and only the RECORD disagreed. That is exactly R-385's shape.
`internal/notify/notifier.go:583-602` (`Severity()`'s doc comment — **it already stated the entire
contract and named both hub locations**), `hub/internal/api/handler.go` (the type rejection and the
severity coercion side by side), `hub/internal/notify/dispatcher.go` (`severityNotifies`,
`ProcessEvent`, `operatorOnlyEvents`, `processOperator`), `internal/stacks/deploy.go` (the
`DesiredState` comment and its one-owner rule), `cmd/controller/main.go` (`classifyRunStates`),
`internal/web/handlers.go` (the compound toggles).
## 2. Architecture documents read
**§2's conditional is answered: the alarm ladder DOES exist** —
`felhom.eu/documentation/architecture/08-alarm-ladder.md`, written last session. It has been extended
here with §6.1 (the severity contract), §7 (the intent test) and §8 (Part 5's direction).
- `documentation/architecture/00-capability-map.md` — its 2026-08-22 paragraph already NAMED R-384 as
an open finding, from the held-app measurement.
- `felhom-controller/internal/stacks/manager.go` `IsDownState` + `aggregateState` + `supervisedPolicy`
- `cmd/controller/main.go` `classifyRunStates` and its three suppressions
- `internal/quiesce/suppress.go`, `internal/bootrecon/bootrecon.go`, `internal/stacks/desiredstate.go`
- `felhom.eu/scripts/golden_currency_gate.py` (all 133 lines)
## 3. The 1.1 sweep — the result in full
**§2's conditional applies and is answered: NO document owned the alarm ladder.** That absence is
reported as a finding, and `documentation/architecture/08-alarm-ladder.md` now owns it (Part N.4).
It is why the ordering defect was legible only by reading one function top to bottom.
**Exactly ONE bad severity in the whole controller: `notifier.go:546`, `"warn"`.** Nothing else.
## 3. Part 0 — the record, pushed ALONE
Verified across Go **and** templates **and** queued-event construction, because the task warned that
reading zero from Go files while the answer sat in a template has produced three wrong conclusions
here:
Exact heading written:
- every `emit(...)` / `PushEvent(...)` literal — one offender, the rest valid;
- `internal/channelhealth`'s classifier (the source of `NotifyAgentChannelDown`'s variable) — all
`"warning"`/`"error"`;
- `debug.html`'s operator-triggerable severity `<select>` — offers only `error`/`warning`/`info`;
- no `PendingEvent{}` literal construction exists anywhere.
```
## v0.221.1 — the undo-copy prune stopped running because another fix made its guard reachable (2026-08-23, R-361 follow-on)
```
Nine other `"warn"` strings exist and are **not** defects: `internal/monitor` and `internal/selftest`
use it as a *healthcheck status* vocabulary, and `statusRank`'s `case "warn"` maps it to the correct
`"warning"` severity. **This is why the guard is an AST walk and not grep.**
Commit **`da75603`**, pushed alone before anything else. The reasoning was **moved verbatim** from the
v0.221.0 entry (which no longer claims it), not rewritten and not duplicated.
**One latent hazard found while sweeping, pinned rather than left:** `fillwatch.Band.Severity()`
returns `""` for `BandOK`. Unreachable, because `Check()` notifies only on an escalation — but that
safety lives in a *different function* from the one that looks unsafe, so the test asserts the
consequence.
## 4. Files changed, commits, CI runs
**And the guard found two dynamic call sites the hand sweep missed** (`NotifyDRCompleted`, and
fillwatch's via `main.go`). All six are now registered by name with the values each can take.
| Commit | Contents |
|---|---|
| **`da75603`** | Part 0 — the `v0.221.1` heading, alone |
| **`5da11c4`** | v0.222.0 — R-384 + R-383, tests, CHANGELOG |
## 4. The hub's manifest — §6's premise was wrong
Modified: `CHANGELOG.md`, `internal/stacks/manager.go`, `internal/stacks/degraded_test.go`,
`internal/backup/offbox_reconstitute.go`, `cmd/controller/r361_classifier_control_test.go`.
Added: `cmd/controller/r384_dead_db_alarm_test.go`, `internal/backup/r383_undo_phrase_test.go`.
**`felhom.eu/manifests/hub.yaml`, line 128.** ArgoCD `Application/felhom` tracks
`admin/felhom.eu.git` path `manifests`, `automated.enabled=false`. Bumped in **`68a9f54`**. **No
out-of-git deployment path exists.** The image was pushed to the registry *before* the manifest landed,
so a sync could never point at a missing tag; the sync was then requested deliberately. Never
`kubectl set image`.
**CI runs confirmed BY ID** (`id` and `run_number` diverge, both printed):
## 5. Files, commits, CI
| Commit | CI `id` | `run_number` | Result |
| Commit | Repo | Contents |
|---|---|---|
| **`9832760`** | controller | v0.223.0 — the severity word, the AST guard, the toggle, the intent test, the toggle split |
| **`<docs>`** | controller | REPORT / CONTEXT / README |
| **`68a9f54`** | felhom.eu | hub v0.107.0 + manifest bump + golden evidence |
| **`2f7c9a6`** | felhom.eu | alarm ladder, register, STATUS, drill record |
Modified: `internal/notify/notifier.go`, `cmd/controller/main.go`, `internal/web/handlers.go`,
`internal/web/templates/settings_notifications.html`, `CHANGELOG.md`.
Added: `internal/notify/r329_severity_contract_test.go`, `internal/fillwatch/r329_severity_test.go`,
`cmd/controller/r386_intent_test.go`, `internal/web/r329_toggle_split_test.go`.
**CI runs confirmed BY ID** (`id` and `run_number` diverge — both printed) — see §15 below.
## 6. Red-proofs — five planted, and ONE PASSED FIRST TIME
| # | Mutation | Layer, and why that layer | Observed |
|---|---|---|---|
| `da75603` | **404** | 85 | success |
| `5da11c4` | **405** | 86 | success |
| 1 | severity back to `"warn"` | **the EMITTER** — the last point at which the bad value still exists; the hub deliberately destroys it one line later | `notifier.go:561:30: emit(...) emits severity "warn", which is NOT in the hub's vocabulary` |
| 2 | `userStopped` back to the state guess | **`classifyRunStates`** — the single derivation point where the guess was made | `dead-app banner = [], want exactly privatebin` — the exact live symptom — plus `IntentUnknown = false` and the wiring check |
| 3 | no-op-save guard removed | **the SAVE handler** — where a render-then-save rewrites stored bytes | `a no-op save CHANGED the stored settings` on the `defaults` shape |
| 4 | hub ingest `WARN` removed | **INGEST** — the last point the offending value exists | `the hub rewrote a severity and said nothing`, log showing only `[INFO] … (info)` |
| 5 | fillwatch de-escalation guard removed | **`Check()`** — the invariant lives there, not in `Severity()` | `notified with band ok → severity "" (event type "")` |
## 5. Red-proofs — four planted, FOUR SEEN FAILING
**Red-proof 5 passed on the first attempt and that is reported, not omitted.** The first mutation —
`if next <= prev` → `if next < prev` — is **inert**: an earlier `if next == prev { continue }` had
already removed the equal case, so the code's behaviour did not change and the test was *right* to
pass. Removing the guard outright convicts it. **A red-proof that passes needs the mutation checked
before either verdict is believed.** Every other mutation asserted its pre-fix text was present before
rewriting and printed `MUTATION APPLIED`.
| # | Mutation | Layer the guard sits at | Observed failure |
|---|---|---|---|
| 1 | hoisted block moved back **below** `unhealthy > 0` | `aggregateState` — the ORDERING | `aggregateState = "unhealthy", want "degraded"` **and** `bookstack state = "unhealthy", want "degraded"` (production-path wiring) |
| 2 | `up` narrowed back to `running` alone | `aggregateState` — the GUARD | same subtest, plus `"starting"` and `"restarting"` — all three survivor shapes convict |
| 3 | classifier drops `StateDegraded` from `down` | `classifyRunStates` — the CONSEQUENCE | `dead-app banner = [], want exactly one entry for bookstack` |
| 4 | `undoCopyPhrase` reverted to the unconditional claim | the phrase builder — where the CLAIM is made | `phrase "…mentése megvan: …mariadb.sql" contains "mentése megvan" — it asserts a file that is not on disk`; the empty set printed `megvan: .`, naming a file that never existed |
## 7. Test counts
**Every mutation was asserted to have applied** (the scripts `assert` the pre-fix text is present
before rewriting and print `MUTATION APPLIED`). **None passed first time.** Mutations 1 and 2 convict
independently, which is what proves the fix genuinely has two halves.
| Repo | Before | After |
|---|---|---|
| felhom-controller | 1504 | **1522** |
| felhom.eu hub | 702 | **709** |
## 6. Test count
Full green gate `go build && go vet && go test ./...` → **exit 0, zero failures, both repos**. All 11
controller design gates OK; all 11 felhom.eu repo gates OK.
**1494 → 1504** top-level test functions (measured by `go test ./... -list '.*'` on the stashed and
unstashed tree, not estimated). Full green gate `go build && go vet && go test ./...` → **exit 0, zero
failures**, run after Part 2 and again after Part 3.
## 7. Deployed version, and the golden
## 8. Deployed versions and the golden
```
gitea.dooplex.hu/admin/felhom-controller:0.222.0 Up 20 seconds (healthy)
gitea.dooplex.hu/admin/felhom-controller:0.223.0 Up 18 minutes (healthy)
gitea.dooplex.hu/admin/felhom-hub:0.107.0 Synced, rollout complete
```
**Golden BAKED and PUBLISHED: YES — version 0.222.0.**
`GOLDEN_SHA256 = 19f5904f53792684f046ec0bc25426645cb87ad73d5cfc6c03639d9f82706037`, `upload OK (HTTP 201)`,
round-trip `HTTP 206` from the package URL. All acceptance markers counted and recorded.
**Golden BAKED and PUBLISHED: YES — 0.223.0.**
`sha256 9eaf39ac39219b42ec9e6cbf890275febcdcc6f53325fe0c0f591d3431044f17`, `upload OK (HTTP 201)`,
round-trip **HTTP 206**, all five acceptance markers counted.
**VOUCHING IS THE OPERATOR'S ACT AND WAS NOT DONE HERE.**
## 8. The five `IsDownState` consumers, walked and named
## 9. The live walk, all six steps
| Consumer | What changes |
|---|---|
| `cmd/controller/main.go:2173` `classifyRunStates` | **THE INTENDED CHANGE.** A stack that read `unhealthy` now reads `degraded` → `down=true` → banner + `app_start_failed`. `userStopped` tests `StateStopped` specifically, so the whitelist cannot swallow `degraded`. |
| `internal/bootrecon/bootrecon.go:213` (DesiredStateRunning) | **CHANGES, and toward repair.** A half-started stack at boot now reads `degraded` → an orphan → `compose up -d`. Previously it read `unhealthy` → not an orphan → left half-dead. Aligned with the file's own stated intent. |
| `internal/bootrecon/bootrecon.go:216` (legacy DesiredStateUnknown) | Same shape, same direction. |
| `internal/bootrecon/bootrecon.go:288` (recovery check) | **CHANGES, and toward truth.** A stack that came back with a dead supervised member is no longer counted `Recovered`; it stays pending and is retried, bounded by `r.attempts`. It used to be declared recovered while half-dead. |
| `internal/stacks/desiredstate.go:150` `isObservedUp` | **UNAFFECTED — verified, not assumed.** It is an allow-list of `{running, starting}`; neither `unhealthy` nor `degraded` was ever in it, so a stack moving between them does not cross the boundary. |
| `internal/quiesce/suppress.go` | **UNAFFECTED.** It does not call `IsDownState` at all — the suppression is cycle-keyed and state-blind, which is precisely why R-97b's guarantee cannot be weakened by a state change. Pinned by `TestR384_QuiesceSuppressionStillHoldsForDegraded`. |
| dashboard state badge | **Already handled.** R-51 wired `degraded` through `handlers.go:158/169` (counts with stopped) and `funcmap.go:255` (filters with stopped). Verified live — the badge rendered `(degraded)`. |
### Step 1 — Scenario A: a database dies, customer has not opted in ✅
## 9. The live walk
Method: endpoint-level. No browser exists on DooPlex; every read below is either the exact endpoint
the UI calls (`POST /api/stacks/<name>/<action>`, `GET /api/stacks`, `GET /dashboard`) or the
controller's own log. Guest clock is UTC.
### Step 1 — Scenario A: a database dies behind a healthy-looking app ✅
| Observable | Result |
|---|---|
| `bookstack-db` stopped out of band | 05:30:07Z |
| front end went `unhealthy` | 05:31:21Z — **the state that used to swallow the alarm** |
| aggregate state read | **`degraded`** while the front end was `unhealthy` (05:32:03Z) |
| `app_start_failed` | **fired at 05:30:14Z**, 7 s after the stop |
| new code path visible | `manager.go:703: restart-policy of down member "bookstack-db" = "unless-stopped"` |
| banner | *„Telepített alkalmazás nem fut: BookStack (degraded)"* on **both** `/launcher` and `/dashboard` |
| edge-triggered, not per-scan | **1 event across 22 scans** |
**The heartbeat, old beside new — same 8 apps evaluated:**
**Quoted from the hub's own records, not the controller's:**
```
2026/08/22 21:13:48 [deadapp] check alive: 580 scans since boot, 8 deployed app(s) evaluated, 0 currently down ← v0.220.2/0.221.1
2026/08/23 05:37:44 [deadapp] check alive: 20 scans since boot, 8 deployed app(s) evaluated, 1 currently down ← v0.222.0
events demo-hp app_start_failed warning 2026-08-23 09:27:51 <- v0.223.0
demo-hp app_start_failed info 2026-08-23 05:30:14 <- v0.222.0, coerced
notification_log demo-hp app_start_failed warning sent operator 2026-08-23 09:27:51
(no customer row)
```
### Step 2 — Scenario B: unhealthy with nothing dead ✅
**The number that says it all: 91 `app_start_failed` events stored all-time, ZERO `notification_log`
rows before 09:00 today.** Not one, ever, on any channel.
The database was restarted; the front end stayed `unhealthy` with nothing down. Aggregate read
**`unhealthy`**, not `degraded`. **0 new `app_start_failed`**, and — the positive observable —
**no `restart-policy of down member` line at all**, meaning the supervised path was not entered. That
absence is trustworthy because the same line HAD appeared on this box 10 minutes earlier.
Banner cleared; `bookstack state=running`.
**An honest limit, stated rather than implied:** this run does **not** prove the customer *gate*.
`demo-hp` has no `customer_notifications` row at all, so the customer leg could not have delivered
regardless. The gate is proven by the unit tests, which configure prefs both ways — and Scenario B
there is the positive control showing the customer leg *can* deliver for this event type.
*Honest limit:* the live window in which docker reported the front end unhealthy with the database up
was ~11 s wide, and the controller's cached read was taken at its edge. The three-shape unit test
`TestR384_UnhealthyWithNothingDeadDoesNotAlarm` carries the rest of this case.
### Step 2 — Scenario D: stopped out of band, intent `running` ✅
### Step 3 — Scenario D: a full deploy cycle ✅ **0 alarms**
`docker compose stop privatebin` at 09:31:27Z → **`app_start_failed (warning)` at 09:31:51Z, 24
seconds later.** Heartbeat, new beside last session's:
`POST /api/stacks/docmost/stop` then `/start` — the exact calls the launcher's buttons make — on a
3-container stack, watched for 5 minutes to settled healthy.
```
2026/08/23 05:47:44 [deadapp] check alive: 40 scans since boot, 8 deployed app(s) evaluated, 0 currently down <- v0.222.0
2026/08/23 09:34:51 [deadapp] check alive: 20 scans since boot, 8 deployed app(s) evaluated, 1 currently down <- v0.223.0
```
| Observable | Result |
|---|---|
| `app_start_failed` across the cycle | **0** |
| dead-app scans in the window | **9** |
| supervised-down path entered | 1, at 05:42:34Z (docmost itself momentarily down beside two live members) |
Hub-side the second alarm was logged `suppressed — operator cooldown 1h` (see §15.1).
**Alarms that v0.221.1 would NOT have produced: ZERO.** The single `degraded` reading at 05:42:34Z
has `running > 0`, so v0.221.1's mixed-case branch reaches the identical verdict. No moment in the
cycle had a down supervised member with only non-`running` survivors, which is the only shape where
the two versions differ.
### Step 3 — Scenario C: the customer presses Stop ✅
### Step 4 — Scenario E: the quiesce cycle — **DROPPED, and why**
Through `POST /api/stacks/privatebin/stop`, the exact call the button makes. Intent moved
`running → stopped`. Four minutes and **9 dead-app scans later: 0 alarms, 0 unknown-intent lines.**
Dropped as a direct consequence of the §4 halt (see the banner at the top), and it is item **(3)** on
the task's own drop list. Running it would have meant triggering a real backup cycle on the box
*after* a halt condition had already fired. **Covered at unit level instead** by
`TestR384_QuiesceSuppressionStillHoldsForDegraded`, which asserts a `degraded` stack inside the
quiesce set produces no banner and `Down=false`. **Not proven live in this session — stated plainly
rather than implied.**
### Step 4 — Scenario E: intent absent ✅ (and the count)
### Step 5 — §4's measurement ✅ (it reproduced — see §10)
demo-hp has **zero** such apps, so one was created: `desired_state` removed from `privatebin`'s
`app.yaml`, as a pre-R-166 box would look (backed up, restored, no code writer added). Result —
evaluated (`deployed=True state=stopped`, 8 apps evaluated), **suppressed (0 alarms)**, and:
### Step 6 — Part 1's gate, both directions ✅
```
[deadapp] 1 stopped app(s) have NO recorded customer intent, so their dead-app alarm is suppressed
by the unknown-intent fallback (R-386): privatebin. This closes itself as each app is started or
stopped through the interface.
```
| Run | Gate | CHANGELOG | Golden | Exit |
|---|---|---|---|---|
| `gate-01` | **old** | v0.221.0 | 0.221.1 | **0** — the blindness, on the real history |
| `gate-02` | **new** | v0.221.0 | 0.221.1 | **1** — convicted, naming the missing heading and the route |
| `gate-03` | new | v0.221.1 | 0.221.1 | 0 |
| `gate-04` | new | absent clone | — | **2** — INCONCLUSIVE preserved |
| `gate-05` | new | v0.222.0 | 0.222.0 | 0 — post-bake |
### Step 5 — Scenario G: save changing nothing ✅ byte-identical, **on the third attempt**
## 10. §4's answer, in plain words
`sha256(enabled_events)` = `10840f3a95bac168f0d7c79760998138` **before and after** a real save
(7 boxes ticked, refusal banner absent).
**A single-container app that is stopped out of band raises no alarm at all, and a comment in the
code says the opposite.**
**The first two attempts were silently REFUSED behind an HTTP 200** — the empty-email wipe guard
declines and renders an error page, still `200`. Run 1: no prefs existed. Run 2: the email `<input>`
spans **three lines**, so a single-line grep read it as `""`. **Both times the hashes matched —
because nothing was saved, not because nothing changed.** Fixed by asserting the refusal banner is
absent. *A warning beside a success is read as a success.*
`aggregateState` folds `StateExited` into the `stopped` counter, so when every member is down it
returns `StateStopped` — **`StateExited` never survives aggregation**, which is the path the task
suspected and could not find. `classifyRunStates` then whitelists `StateStopped` as a deliberate user
stop. So the comment at `cmd/controller/main.go` — *"An out-of-band `docker compose stop` leaves the
containers present → StateExited → still alerts, which is correct: out-of-band tampering IS
reportable"* — is **false**, and so is the neighbouring I2 claim that a crashing app never comes to
rest at `stopped`.
### Step 6 — Scenario H: a bad severity to the hub ✅
**Measured:** `privatebin` (1 container, `unless-stopped`) stopped 05:47:35Z. At 05:51:53Z:
`state=stopped`, **9 dead-app scans had run, 0 events, 0 banner lines.**
**Positive control first, per standing rule 3:** `app_start_failed` fired for BookStack at 05:30:14Z
on the same box 17 minutes earlier, so the detector was demonstrably alive.
```
[WARN] [api] Event from demo-hp: severity "warn" is not in {info,warning,error,critical} — coercing
to "info", which severityNotifies DROPS, so this backup_failed alert will reach NOBODY. Fix the
emitting controller; this event is stored but not routed.
```
**Scoped honestly:** a genuine *crash* under `unless-stopped` is restarted by Docker and surfaces as
`restarting` → the 5-minute crash-loop path, which does alarm. The silent case is an explicit
out-of-band stop of a stack with no surviving member.
Both the bad POST and an `error` control returned **200** (nothing lost); the control produced **no**
warning.
**Filed as R-386 (OPEN — MEDIUM). Not fixed here, per §12 and §13.**
## 10. The absent-intent count on `demo-hp`, in plain words
## 11. Evidence
**Zero.** All **8** deployed apps carry `desired_state: running`; none is `stopped` and none is absent.
The unknown-intent fallback therefore suppresses nothing on this box today — the population is already
empty on a machine that has been exercised through the interface. It will be larger on a box upgraded
and left alone, which is why the log line exists rather than a one-off count.
`felhom.eu/documentation/audits/DRILL-r384-dead-db-alarm-2026-08-23/evidence/` — 5 gate runs, 4
red-proof transcripts, 20 live-walk files including two full controller-log windows (1808 and 2139
lines) pulled off the guest. **Both log windows were copied off before the app was restarted**, per
standing rule 5.
## 11. The dead-branch decision, and the reason
## 12. Teardown, three layers, and the box's end state
**KEPT.** `cmd/hub/main.go` wires `dispatcher.ProcessEvent` **directly** as the
`monitor.EventNotifyFunc` for the staleness, host-staleness and offsite-box checkers — those events
never pass the ingest handler, so for them that line is the only severity guard there is. Deleting it
as "dead" would have removed the live half while the dead half supplied the justification. All 90
severity literals in `internal/monitor` were verified already valid, so the guard is silent because
the producers are correct.
1. **Guest 9201 / apps** — nothing provisioned. `bookstack-db` restarted and **`bookstack` confirmed
healthy**; `privatebin` restarted and healthy; `docmost` (all 3) healthy. Planted data untouched
throughout — no app was rebuilt, redeployed or restored.
2. **Bake VM** — the drill VM on DooPlex ran the bake and its build guest 9100 is stopped inside it;
the qcow2 reverts to the `virgin` snapshot. No storage was added anywhere, so `pvesm status` has
nothing to compare.
3. **Hub-side record — stated explicitly even though there is none.** No appliance was registered, no
customer created, no config written, no artifact manifest changed. **The hub was READ ONLY**
(`GET /configuration`, `GET /events`). Nothing to discard.
## 12. Evidence
**End state:** `demo-hp` guest 9201 runs controller **0.222.0**, all 8 deployed apps healthy, golden
0.222.0 baked and published but **NOT vouched** — floor still **0.221.1**.
`felhom.eu/documentation/audits/DRILL-r329-r386-2026-08-23/evidence/` — 26 files: 5 red-proof
transcripts, 20 live-walk files, two full controller-log windows (1041 and 4044 lines) **pulled off
before each revert**, and the hub's own DB queries.
## 13. Register size
## 13. Teardown, three layers, and the end state
1. **Guest 9201 / apps** — nothing provisioned. **All 17 app containers healthy.** `privatebin`'s
`app.yaml` restored from backup and the backup deleted; intent reads `running`. `demo-hp`'s
notification settings restored to `enabled_events: null`, no e-mail — their pre-drill state. No app
rebuilt, redeployed or restored; planted data untouched.
2. **Bake VM** — powered off, Gitea token and runner script **shredded**, `drill.qcow2` reverted to
`virgin`. Build guest 9100 exists only inside that reverted snapshot. No storage added anywhere, so
`pvesm status` has nothing to compare.
3. **Hub-side, stated explicitly.** The hub was **written** this session, unlike last: the deployment
is v0.107.0 via the manifest, and **two probe events remain as rows for `demo-hp`** from Scenario H
(`backup_failed`, "R-387 scenario H probe" and "…control"). They are inert records; named here
rather than left for someone to find. Nothing else: no appliance registered, no customer created,
no artifact manifest changed, floor untouched.
**End state:** controller **0.223.0** and hub **0.107.0** deployed; golden **0.223.0** baked and
published but **NOT vouched**; floor still **0.222.0**; all apps running; planted data present.
## 14. Register size
| File | Before | After |
|---|---|---|
| `OPEN-ITEMS.md` | 327,266 B | **328,325 B** |
| `CLOSED-ITEMS.md` | 68,464 B | **71,441 B** |
| `OPEN-ITEMS.md` | 328,325 B | **328,132 B** |
| `CLOSED-ITEMS.md` | 71,441 B | **74,642 B** |
R-383 and R-384 moved to CLOSED compressed; R-385 (closed) and R-386 (open) filed. OPEN grew by
~1 KB despite two closures because R-386 is a substantial new finding — recorded rather than
smoothed over.
R-329 and R-386 closed and compressed; **R-387** (closed) and **R-388** (the notification-model
product decision — open, operator's call) filed.
## 14. Observations — noticed, documented, NOT acted on
## 15. Observations — noticed, documented, NOT acted on
1. **R-329 is live and now matters much more.** `app_start_failed` is pushed with severity **`warn`**,
which is not in the hub's vocabulary (`{info, warning, error, critical}`) and coerces silently to
`info`, e-mailing nobody, while the POST still returns 200. Observed again today:
`PushEvent: type=app_start_failed severity=warn`. **R-384 makes this event actually fire, so a
known-broken severity moved from unreachable to load-bearing.** Not in scope; not touched.
2. **Two files carry pre-existing `gofmt` drift** — `internal/backup/offbox.go` and
`internal/backup/offbox_recovery_cli.go`. Confirmed pre-existing by stashing this session's work
and re-running `gofmt -l`. Not touched (§12 forbids nearby refactors).
3. **The runbook's golden-bake step is missing `pveam update`.** On the `virgin` snapshot the template
index is stale, so `pveam available` offers `13.1-2` and downloading it fails with
`400 Parameter verification failed. template: no such template`. Recorded in the bake evidence
README; the runbook itself was not edited.
4. **The register's own suggested fix for R-384 was wrong** — it proposed a sustained-`unhealthy`
threshold on the `crashLoopAfter` model. The defect needed no threshold at all, only an ordering.
Recorded in the CLOSED entry so the next reader sees that a register remedy is a hypothesis.
5. **Deliberately left open, untouched:** R-102, R-359, R-361's sibling surfaces.
1. **The operator cooldown key has no app identifier, and it now bites.** PrivateBin's alarm four
minutes after BookStack's was logged `suppressed — operator cooldown 1h,
key=demo-hp:app_start_failed`, so **only the first app-down per hour e-mails the operator**. This is
R-182's known cooldown-key shape; it was harmless while `app_start_failed` was undeliverable and is
not any more. **Same pattern as R-329 itself: a known-broken thing moved from unreachable to
load-bearing.** Not fixed here.
2. **The settings page grew 12 → 15 toggles in one session** — one new alarm, plus two compound
toggles split into four. Recorded as the argument inside R-388.
3. `internal/notify/notifier.go` carries **pre-existing** gofmt drift in an unrelated const block,
confirmed by stashing this session's work and re-running `gofmt -l`. Not touched (§12).
4. **The golden-bake runbook still lacks `pveam update`** — second consecutive bake to hit the stale
index on the `virgin` snapshot, presenting as `400 … no such template`.
5. **Deliberately left open, untouched:** R-102, R-359, R-385, and R-388's redesign.
### CI runs, confirmed by ID
| Commit | Repo | CI `id` | `run_number` | Result |
|---|---|---|---|---|
| `9832760` | felhom-controller | **408** | 88 | success |
| `68a9f54` | felhom.eu | **409** | 261 | success |
| `2f7c9a6` | felhom.eu | **410** | 262 | success |
(The controller docs commit's run is confirmed after its push and is the next `id` in that repo.)