docs(v0.223.0): REPORT, CONTEXT rulings, README severity contract
gates / gates (push) Successful in 11s
gates / gates (push) Successful in 11s
REPORT overwritten: the 1.1 sweep in full (one bad severity, nine legitimate "warn" strings that are healthcheck statuses), the hub manifest's real location since the task's premise was wrong, all five red-proofs with the layer each guard sits at, the live walk in six steps with the hub's own records quoted, and the absent-intent count (0 of 8). Three things are reported that a tidier account would omit: red-proof 5 passed first time because the mutation was INERT; Scenario G was silently refused twice behind an HTTP 200; and the live Scenario A does NOT prove the customer gate, because demo-hp has no prefs row at all. CONTEXT records the severity vocabulary as a ruling with its mechanism, the intent ruling with its three-way handling of unknown, both fences, and two traps worth more than the fixes: a 200 can be a refusal, and a passing red-proof can mean an inert mutation. README: the event table said `app_start_failed | warn` - the defect, written down as if correct. Now `warning`, with the vocabulary contract and who receives what. `disk_critical` also corrected from `error` to `critical`, which is what fillwatch has always sent.
This commit is contained in:
@@ -1,267 +1,266 @@
|
||||
# REPORT — controller v0.221.1 (record) + v0.222.0 (R-384, R-383)
|
||||
# REPORT — controller v0.223.0 (R-329, R-386, and the compound toggles)
|
||||
|
||||
**Session 2026-08-23, UNATTENDED. Live leg on `demo-hp` (Tier 0, disposable).**
|
||||
**Session 2026-08-23, UNATTENDED.** Live leg on `demo-hp` (Tier 0), guest 9201.
|
||||
**No halt condition fired.** Nothing was dropped.
|
||||
|
||||
> ## ⚠ HALT DECLARED — §4's measurement found a real defect, filed as R-386, NOT fixed
|
||||
>
|
||||
> The task's §4 asked for a measurement and named it a halt condition. It reproduced.
|
||||
> **A single-container app stopped out of band raises no alarm at all.** Details in §10 below.
|
||||
> Per §13 the fix was NOT attempted here. Live-walk step 4 (Scenario E) was dropped as a
|
||||
> consequence — it is item (3) on the task's own drop list. Everything else completed.
|
||||
## 1. Baselines, and the hub's four numbers as read
|
||||
|
||||
---
|
||||
|
||||
## 1. Baselines used, and the hub's four numbers as read
|
||||
|
||||
| Repo | `main` at start | Verified |
|
||||
| Repo | at start | at end |
|
||||
|---|---|---|
|
||||
| felhom-controller | `f7881787f434` | matches the task |
|
||||
| felhom.eu | `1eb64bec5183` | task said `4e488321bfd1+`; it had moved on |
|
||||
| felhom-agent | untouched | — |
|
||||
| felhom-controller | `14137efa` (v0.222.0) | **v0.223.0** deployed |
|
||||
| felhom.eu | `55274d5e` (hub v0.106.0) | **hub v0.107.0** deployed |
|
||||
| felhom-agent | `40d857b5` (v0.130.0) | untouched |
|
||||
|
||||
**Hub's own numbers, read live from `/configuration` (ClusterIP + Basic auth) 2026-08-23:**
|
||||
**Hub's four numbers, live from `GET /configuration` before starting:** `golden_version` **0.222.0**,
|
||||
`agent_version` **0.130.0**, `min_agent` **0.129.0**, controller floor **0.222.0** — all four as the
|
||||
task predicted.
|
||||
|
||||
| Field | Value |
|
||||
|---|---|
|
||||
| `golden_version` | **0.221.1** |
|
||||
| `agent_version` | **0.130.0** |
|
||||
| `min_agent` | **0.129.0** |
|
||||
| controller floor (`min_controller_version`) | **0.221.1** |
|
||||
## 2. Documents read
|
||||
|
||||
The task expected golden/floor **0.220.2**; the operator had already vouched **0.221.1** and raised
|
||||
the floor. Live controller on `demo-hp` at session start: **0.221.1** — so the running version, the
|
||||
golden and the floor all agreed, and only the RECORD disagreed. That is exactly R-385's shape.
|
||||
`internal/notify/notifier.go:583-602` (`Severity()`'s doc comment — **it already stated the entire
|
||||
contract and named both hub locations**), `hub/internal/api/handler.go` (the type rejection and the
|
||||
severity coercion side by side), `hub/internal/notify/dispatcher.go` (`severityNotifies`,
|
||||
`ProcessEvent`, `operatorOnlyEvents`, `processOperator`), `internal/stacks/deploy.go` (the
|
||||
`DesiredState` comment and its one-owner rule), `cmd/controller/main.go` (`classifyRunStates`),
|
||||
`internal/web/handlers.go` (the compound toggles).
|
||||
|
||||
## 2. Architecture documents read
|
||||
**§2's conditional is answered: the alarm ladder DOES exist** —
|
||||
`felhom.eu/documentation/architecture/08-alarm-ladder.md`, written last session. It has been extended
|
||||
here with §6.1 (the severity contract), §7 (the intent test) and §8 (Part 5's direction).
|
||||
|
||||
- `documentation/architecture/00-capability-map.md` — its 2026-08-22 paragraph already NAMED R-384 as
|
||||
an open finding, from the held-app measurement.
|
||||
- `felhom-controller/internal/stacks/manager.go` `IsDownState` + `aggregateState` + `supervisedPolicy`
|
||||
- `cmd/controller/main.go` `classifyRunStates` and its three suppressions
|
||||
- `internal/quiesce/suppress.go`, `internal/bootrecon/bootrecon.go`, `internal/stacks/desiredstate.go`
|
||||
- `felhom.eu/scripts/golden_currency_gate.py` (all 133 lines)
|
||||
## 3. The 1.1 sweep — the result in full
|
||||
|
||||
**§2's conditional applies and is answered: NO document owned the alarm ladder.** That absence is
|
||||
reported as a finding, and `documentation/architecture/08-alarm-ladder.md` now owns it (Part N.4).
|
||||
It is why the ordering defect was legible only by reading one function top to bottom.
|
||||
**Exactly ONE bad severity in the whole controller: `notifier.go:546`, `"warn"`.** Nothing else.
|
||||
|
||||
## 3. Part 0 — the record, pushed ALONE
|
||||
Verified across Go **and** templates **and** queued-event construction, because the task warned that
|
||||
reading zero from Go files while the answer sat in a template has produced three wrong conclusions
|
||||
here:
|
||||
|
||||
Exact heading written:
|
||||
- every `emit(...)` / `PushEvent(...)` literal — one offender, the rest valid;
|
||||
- `internal/channelhealth`'s classifier (the source of `NotifyAgentChannelDown`'s variable) — all
|
||||
`"warning"`/`"error"`;
|
||||
- `debug.html`'s operator-triggerable severity `<select>` — offers only `error`/`warning`/`info`;
|
||||
- no `PendingEvent{}` literal construction exists anywhere.
|
||||
|
||||
```
|
||||
## v0.221.1 — the undo-copy prune stopped running because another fix made its guard reachable (2026-08-23, R-361 follow-on)
|
||||
```
|
||||
Nine other `"warn"` strings exist and are **not** defects: `internal/monitor` and `internal/selftest`
|
||||
use it as a *healthcheck status* vocabulary, and `statusRank`'s `case "warn"` maps it to the correct
|
||||
`"warning"` severity. **This is why the guard is an AST walk and not grep.**
|
||||
|
||||
Commit **`da75603`**, pushed alone before anything else. The reasoning was **moved verbatim** from the
|
||||
v0.221.0 entry (which no longer claims it), not rewritten and not duplicated.
|
||||
**One latent hazard found while sweeping, pinned rather than left:** `fillwatch.Band.Severity()`
|
||||
returns `""` for `BandOK`. Unreachable, because `Check()` notifies only on an escalation — but that
|
||||
safety lives in a *different function* from the one that looks unsafe, so the test asserts the
|
||||
consequence.
|
||||
|
||||
## 4. Files changed, commits, CI runs
|
||||
**And the guard found two dynamic call sites the hand sweep missed** (`NotifyDRCompleted`, and
|
||||
fillwatch's via `main.go`). All six are now registered by name with the values each can take.
|
||||
|
||||
| Commit | Contents |
|
||||
|---|---|
|
||||
| **`da75603`** | Part 0 — the `v0.221.1` heading, alone |
|
||||
| **`5da11c4`** | v0.222.0 — R-384 + R-383, tests, CHANGELOG |
|
||||
## 4. The hub's manifest — §6's premise was wrong
|
||||
|
||||
Modified: `CHANGELOG.md`, `internal/stacks/manager.go`, `internal/stacks/degraded_test.go`,
|
||||
`internal/backup/offbox_reconstitute.go`, `cmd/controller/r361_classifier_control_test.go`.
|
||||
Added: `cmd/controller/r384_dead_db_alarm_test.go`, `internal/backup/r383_undo_phrase_test.go`.
|
||||
**`felhom.eu/manifests/hub.yaml`, line 128.** ArgoCD `Application/felhom` tracks
|
||||
`admin/felhom.eu.git` path `manifests`, `automated.enabled=false`. Bumped in **`68a9f54`**. **No
|
||||
out-of-git deployment path exists.** The image was pushed to the registry *before* the manifest landed,
|
||||
so a sync could never point at a missing tag; the sync was then requested deliberately. Never
|
||||
`kubectl set image`.
|
||||
|
||||
**CI runs confirmed BY ID** (`id` and `run_number` diverge, both printed):
|
||||
## 5. Files, commits, CI
|
||||
|
||||
| Commit | CI `id` | `run_number` | Result |
|
||||
| Commit | Repo | Contents |
|
||||
|---|---|---|
|
||||
| **`9832760`** | controller | v0.223.0 — the severity word, the AST guard, the toggle, the intent test, the toggle split |
|
||||
| **`<docs>`** | controller | REPORT / CONTEXT / README |
|
||||
| **`68a9f54`** | felhom.eu | hub v0.107.0 + manifest bump + golden evidence |
|
||||
| **`2f7c9a6`** | felhom.eu | alarm ladder, register, STATUS, drill record |
|
||||
|
||||
Modified: `internal/notify/notifier.go`, `cmd/controller/main.go`, `internal/web/handlers.go`,
|
||||
`internal/web/templates/settings_notifications.html`, `CHANGELOG.md`.
|
||||
Added: `internal/notify/r329_severity_contract_test.go`, `internal/fillwatch/r329_severity_test.go`,
|
||||
`cmd/controller/r386_intent_test.go`, `internal/web/r329_toggle_split_test.go`.
|
||||
|
||||
**CI runs confirmed BY ID** (`id` and `run_number` diverge — both printed) — see §15 below.
|
||||
|
||||
## 6. Red-proofs — five planted, and ONE PASSED FIRST TIME
|
||||
|
||||
| # | Mutation | Layer, and why that layer | Observed |
|
||||
|---|---|---|---|
|
||||
| `da75603` | **404** | 85 | success |
|
||||
| `5da11c4` | **405** | 86 | success |
|
||||
| 1 | severity back to `"warn"` | **the EMITTER** — the last point at which the bad value still exists; the hub deliberately destroys it one line later | `notifier.go:561:30: emit(...) emits severity "warn", which is NOT in the hub's vocabulary` |
|
||||
| 2 | `userStopped` back to the state guess | **`classifyRunStates`** — the single derivation point where the guess was made | `dead-app banner = [], want exactly privatebin` — the exact live symptom — plus `IntentUnknown = false` and the wiring check |
|
||||
| 3 | no-op-save guard removed | **the SAVE handler** — where a render-then-save rewrites stored bytes | `a no-op save CHANGED the stored settings` on the `defaults` shape |
|
||||
| 4 | hub ingest `WARN` removed | **INGEST** — the last point the offending value exists | `the hub rewrote a severity and said nothing`, log showing only `[INFO] … (info)` |
|
||||
| 5 | fillwatch de-escalation guard removed | **`Check()`** — the invariant lives there, not in `Severity()` | `notified with band ok → severity "" (event type "")` |
|
||||
|
||||
## 5. Red-proofs — four planted, FOUR SEEN FAILING
|
||||
**Red-proof 5 passed on the first attempt and that is reported, not omitted.** The first mutation —
|
||||
`if next <= prev` → `if next < prev` — is **inert**: an earlier `if next == prev { continue }` had
|
||||
already removed the equal case, so the code's behaviour did not change and the test was *right* to
|
||||
pass. Removing the guard outright convicts it. **A red-proof that passes needs the mutation checked
|
||||
before either verdict is believed.** Every other mutation asserted its pre-fix text was present before
|
||||
rewriting and printed `MUTATION APPLIED`.
|
||||
|
||||
| # | Mutation | Layer the guard sits at | Observed failure |
|
||||
|---|---|---|---|
|
||||
| 1 | hoisted block moved back **below** `unhealthy > 0` | `aggregateState` — the ORDERING | `aggregateState = "unhealthy", want "degraded"` **and** `bookstack state = "unhealthy", want "degraded"` (production-path wiring) |
|
||||
| 2 | `up` narrowed back to `running` alone | `aggregateState` — the GUARD | same subtest, plus `"starting"` and `"restarting"` — all three survivor shapes convict |
|
||||
| 3 | classifier drops `StateDegraded` from `down` | `classifyRunStates` — the CONSEQUENCE | `dead-app banner = [], want exactly one entry for bookstack` |
|
||||
| 4 | `undoCopyPhrase` reverted to the unconditional claim | the phrase builder — where the CLAIM is made | `phrase "…mentése megvan: …mariadb.sql" contains "mentése megvan" — it asserts a file that is not on disk`; the empty set printed `megvan: .`, naming a file that never existed |
|
||||
## 7. Test counts
|
||||
|
||||
**Every mutation was asserted to have applied** (the scripts `assert` the pre-fix text is present
|
||||
before rewriting and print `MUTATION APPLIED`). **None passed first time.** Mutations 1 and 2 convict
|
||||
independently, which is what proves the fix genuinely has two halves.
|
||||
| Repo | Before | After |
|
||||
|---|---|---|
|
||||
| felhom-controller | 1504 | **1522** |
|
||||
| felhom.eu hub | 702 | **709** |
|
||||
|
||||
## 6. Test count
|
||||
Full green gate `go build && go vet && go test ./...` → **exit 0, zero failures, both repos**. All 11
|
||||
controller design gates OK; all 11 felhom.eu repo gates OK.
|
||||
|
||||
**1494 → 1504** top-level test functions (measured by `go test ./... -list '.*'` on the stashed and
|
||||
unstashed tree, not estimated). Full green gate `go build && go vet && go test ./...` → **exit 0, zero
|
||||
failures**, run after Part 2 and again after Part 3.
|
||||
|
||||
## 7. Deployed version, and the golden
|
||||
## 8. Deployed versions and the golden
|
||||
|
||||
```
|
||||
gitea.dooplex.hu/admin/felhom-controller:0.222.0 Up 20 seconds (healthy)
|
||||
gitea.dooplex.hu/admin/felhom-controller:0.223.0 Up 18 minutes (healthy)
|
||||
gitea.dooplex.hu/admin/felhom-hub:0.107.0 Synced, rollout complete
|
||||
```
|
||||
|
||||
**Golden BAKED and PUBLISHED: YES — version 0.222.0.**
|
||||
`GOLDEN_SHA256 = 19f5904f53792684f046ec0bc25426645cb87ad73d5cfc6c03639d9f82706037`, `upload OK (HTTP 201)`,
|
||||
round-trip `HTTP 206` from the package URL. All acceptance markers counted and recorded.
|
||||
**Golden BAKED and PUBLISHED: YES — 0.223.0.**
|
||||
`sha256 9eaf39ac39219b42ec9e6cbf890275febcdcc6f53325fe0c0f591d3431044f17`, `upload OK (HTTP 201)`,
|
||||
round-trip **HTTP 206**, all five acceptance markers counted.
|
||||
**VOUCHING IS THE OPERATOR'S ACT AND WAS NOT DONE HERE.**
|
||||
|
||||
## 8. The five `IsDownState` consumers, walked and named
|
||||
## 9. The live walk, all six steps
|
||||
|
||||
| Consumer | What changes |
|
||||
|---|---|
|
||||
| `cmd/controller/main.go:2173` `classifyRunStates` | **THE INTENDED CHANGE.** A stack that read `unhealthy` now reads `degraded` → `down=true` → banner + `app_start_failed`. `userStopped` tests `StateStopped` specifically, so the whitelist cannot swallow `degraded`. |
|
||||
| `internal/bootrecon/bootrecon.go:213` (DesiredStateRunning) | **CHANGES, and toward repair.** A half-started stack at boot now reads `degraded` → an orphan → `compose up -d`. Previously it read `unhealthy` → not an orphan → left half-dead. Aligned with the file's own stated intent. |
|
||||
| `internal/bootrecon/bootrecon.go:216` (legacy DesiredStateUnknown) | Same shape, same direction. |
|
||||
| `internal/bootrecon/bootrecon.go:288` (recovery check) | **CHANGES, and toward truth.** A stack that came back with a dead supervised member is no longer counted `Recovered`; it stays pending and is retried, bounded by `r.attempts`. It used to be declared recovered while half-dead. |
|
||||
| `internal/stacks/desiredstate.go:150` `isObservedUp` | **UNAFFECTED — verified, not assumed.** It is an allow-list of `{running, starting}`; neither `unhealthy` nor `degraded` was ever in it, so a stack moving between them does not cross the boundary. |
|
||||
| `internal/quiesce/suppress.go` | **UNAFFECTED.** It does not call `IsDownState` at all — the suppression is cycle-keyed and state-blind, which is precisely why R-97b's guarantee cannot be weakened by a state change. Pinned by `TestR384_QuiesceSuppressionStillHoldsForDegraded`. |
|
||||
| dashboard state badge | **Already handled.** R-51 wired `degraded` through `handlers.go:158/169` (counts with stopped) and `funcmap.go:255` (filters with stopped). Verified live — the badge rendered `(degraded)`. |
|
||||
### Step 1 — Scenario A: a database dies, customer has not opted in ✅
|
||||
|
||||
## 9. The live walk
|
||||
|
||||
Method: endpoint-level. No browser exists on DooPlex; every read below is either the exact endpoint
|
||||
the UI calls (`POST /api/stacks/<name>/<action>`, `GET /api/stacks`, `GET /dashboard`) or the
|
||||
controller's own log. Guest clock is UTC.
|
||||
|
||||
### Step 1 — Scenario A: a database dies behind a healthy-looking app ✅
|
||||
|
||||
| Observable | Result |
|
||||
|---|---|
|
||||
| `bookstack-db` stopped out of band | 05:30:07Z |
|
||||
| front end went `unhealthy` | 05:31:21Z — **the state that used to swallow the alarm** |
|
||||
| aggregate state read | **`degraded`** while the front end was `unhealthy` (05:32:03Z) |
|
||||
| `app_start_failed` | **fired at 05:30:14Z**, 7 s after the stop |
|
||||
| new code path visible | `manager.go:703: restart-policy of down member "bookstack-db" = "unless-stopped"` |
|
||||
| banner | *„Telepített alkalmazás nem fut: BookStack (degraded)"* on **both** `/launcher` and `/dashboard` |
|
||||
| edge-triggered, not per-scan | **1 event across 22 scans** |
|
||||
|
||||
**The heartbeat, old beside new — same 8 apps evaluated:**
|
||||
**Quoted from the hub's own records, not the controller's:**
|
||||
|
||||
```
|
||||
2026/08/22 21:13:48 [deadapp] check alive: 580 scans since boot, 8 deployed app(s) evaluated, 0 currently down ← v0.220.2/0.221.1
|
||||
2026/08/23 05:37:44 [deadapp] check alive: 20 scans since boot, 8 deployed app(s) evaluated, 1 currently down ← v0.222.0
|
||||
events demo-hp app_start_failed warning 2026-08-23 09:27:51 <- v0.223.0
|
||||
demo-hp app_start_failed info 2026-08-23 05:30:14 <- v0.222.0, coerced
|
||||
notification_log demo-hp app_start_failed warning sent operator 2026-08-23 09:27:51
|
||||
(no customer row)
|
||||
```
|
||||
|
||||
### Step 2 — Scenario B: unhealthy with nothing dead ✅
|
||||
**The number that says it all: 91 `app_start_failed` events stored all-time, ZERO `notification_log`
|
||||
rows before 09:00 today.** Not one, ever, on any channel.
|
||||
|
||||
The database was restarted; the front end stayed `unhealthy` with nothing down. Aggregate read
|
||||
**`unhealthy`**, not `degraded`. **0 new `app_start_failed`**, and — the positive observable —
|
||||
**no `restart-policy of down member` line at all**, meaning the supervised path was not entered. That
|
||||
absence is trustworthy because the same line HAD appeared on this box 10 minutes earlier.
|
||||
Banner cleared; `bookstack state=running`.
|
||||
**An honest limit, stated rather than implied:** this run does **not** prove the customer *gate*.
|
||||
`demo-hp` has no `customer_notifications` row at all, so the customer leg could not have delivered
|
||||
regardless. The gate is proven by the unit tests, which configure prefs both ways — and Scenario B
|
||||
there is the positive control showing the customer leg *can* deliver for this event type.
|
||||
|
||||
*Honest limit:* the live window in which docker reported the front end unhealthy with the database up
|
||||
was ~11 s wide, and the controller's cached read was taken at its edge. The three-shape unit test
|
||||
`TestR384_UnhealthyWithNothingDeadDoesNotAlarm` carries the rest of this case.
|
||||
### Step 2 — Scenario D: stopped out of band, intent `running` ✅
|
||||
|
||||
### Step 3 — Scenario D: a full deploy cycle ✅ **0 alarms**
|
||||
`docker compose stop privatebin` at 09:31:27Z → **`app_start_failed (warning)` at 09:31:51Z, 24
|
||||
seconds later.** Heartbeat, new beside last session's:
|
||||
|
||||
`POST /api/stacks/docmost/stop` then `/start` — the exact calls the launcher's buttons make — on a
|
||||
3-container stack, watched for 5 minutes to settled healthy.
|
||||
```
|
||||
2026/08/23 05:47:44 [deadapp] check alive: 40 scans since boot, 8 deployed app(s) evaluated, 0 currently down <- v0.222.0
|
||||
2026/08/23 09:34:51 [deadapp] check alive: 20 scans since boot, 8 deployed app(s) evaluated, 1 currently down <- v0.223.0
|
||||
```
|
||||
|
||||
| Observable | Result |
|
||||
|---|---|
|
||||
| `app_start_failed` across the cycle | **0** |
|
||||
| dead-app scans in the window | **9** |
|
||||
| supervised-down path entered | 1, at 05:42:34Z (docmost itself momentarily down beside two live members) |
|
||||
Hub-side the second alarm was logged `suppressed — operator cooldown 1h` (see §15.1).
|
||||
|
||||
**Alarms that v0.221.1 would NOT have produced: ZERO.** The single `degraded` reading at 05:42:34Z
|
||||
has `running > 0`, so v0.221.1's mixed-case branch reaches the identical verdict. No moment in the
|
||||
cycle had a down supervised member with only non-`running` survivors, which is the only shape where
|
||||
the two versions differ.
|
||||
### Step 3 — Scenario C: the customer presses Stop ✅
|
||||
|
||||
### Step 4 — Scenario E: the quiesce cycle — **DROPPED, and why**
|
||||
Through `POST /api/stacks/privatebin/stop`, the exact call the button makes. Intent moved
|
||||
`running → stopped`. Four minutes and **9 dead-app scans later: 0 alarms, 0 unknown-intent lines.**
|
||||
|
||||
Dropped as a direct consequence of the §4 halt (see the banner at the top), and it is item **(3)** on
|
||||
the task's own drop list. Running it would have meant triggering a real backup cycle on the box
|
||||
*after* a halt condition had already fired. **Covered at unit level instead** by
|
||||
`TestR384_QuiesceSuppressionStillHoldsForDegraded`, which asserts a `degraded` stack inside the
|
||||
quiesce set produces no banner and `Down=false`. **Not proven live in this session — stated plainly
|
||||
rather than implied.**
|
||||
### Step 4 — Scenario E: intent absent ✅ (and the count)
|
||||
|
||||
### Step 5 — §4's measurement ✅ (it reproduced — see §10)
|
||||
demo-hp has **zero** such apps, so one was created: `desired_state` removed from `privatebin`'s
|
||||
`app.yaml`, as a pre-R-166 box would look (backed up, restored, no code writer added). Result —
|
||||
evaluated (`deployed=True state=stopped`, 8 apps evaluated), **suppressed (0 alarms)**, and:
|
||||
|
||||
### Step 6 — Part 1's gate, both directions ✅
|
||||
```
|
||||
[deadapp] 1 stopped app(s) have NO recorded customer intent, so their dead-app alarm is suppressed
|
||||
by the unknown-intent fallback (R-386): privatebin. This closes itself as each app is started or
|
||||
stopped through the interface.
|
||||
```
|
||||
|
||||
| Run | Gate | CHANGELOG | Golden | Exit |
|
||||
|---|---|---|---|---|
|
||||
| `gate-01` | **old** | v0.221.0 | 0.221.1 | **0** — the blindness, on the real history |
|
||||
| `gate-02` | **new** | v0.221.0 | 0.221.1 | **1** — convicted, naming the missing heading and the route |
|
||||
| `gate-03` | new | v0.221.1 | 0.221.1 | 0 |
|
||||
| `gate-04` | new | absent clone | — | **2** — INCONCLUSIVE preserved |
|
||||
| `gate-05` | new | v0.222.0 | 0.222.0 | 0 — post-bake |
|
||||
### Step 5 — Scenario G: save changing nothing ✅ byte-identical, **on the third attempt**
|
||||
|
||||
## 10. §4's answer, in plain words
|
||||
`sha256(enabled_events)` = `10840f3a95bac168f0d7c79760998138` **before and after** a real save
|
||||
(7 boxes ticked, refusal banner absent).
|
||||
|
||||
**A single-container app that is stopped out of band raises no alarm at all, and a comment in the
|
||||
code says the opposite.**
|
||||
**The first two attempts were silently REFUSED behind an HTTP 200** — the empty-email wipe guard
|
||||
declines and renders an error page, still `200`. Run 1: no prefs existed. Run 2: the email `<input>`
|
||||
spans **three lines**, so a single-line grep read it as `""`. **Both times the hashes matched —
|
||||
because nothing was saved, not because nothing changed.** Fixed by asserting the refusal banner is
|
||||
absent. *A warning beside a success is read as a success.*
|
||||
|
||||
`aggregateState` folds `StateExited` into the `stopped` counter, so when every member is down it
|
||||
returns `StateStopped` — **`StateExited` never survives aggregation**, which is the path the task
|
||||
suspected and could not find. `classifyRunStates` then whitelists `StateStopped` as a deliberate user
|
||||
stop. So the comment at `cmd/controller/main.go` — *"An out-of-band `docker compose stop` leaves the
|
||||
containers present → StateExited → still alerts, which is correct: out-of-band tampering IS
|
||||
reportable"* — is **false**, and so is the neighbouring I2 claim that a crashing app never comes to
|
||||
rest at `stopped`.
|
||||
### Step 6 — Scenario H: a bad severity to the hub ✅
|
||||
|
||||
**Measured:** `privatebin` (1 container, `unless-stopped`) stopped 05:47:35Z. At 05:51:53Z:
|
||||
`state=stopped`, **9 dead-app scans had run, 0 events, 0 banner lines.**
|
||||
**Positive control first, per standing rule 3:** `app_start_failed` fired for BookStack at 05:30:14Z
|
||||
on the same box 17 minutes earlier, so the detector was demonstrably alive.
|
||||
```
|
||||
[WARN] [api] Event from demo-hp: severity "warn" is not in {info,warning,error,critical} — coercing
|
||||
to "info", which severityNotifies DROPS, so this backup_failed alert will reach NOBODY. Fix the
|
||||
emitting controller; this event is stored but not routed.
|
||||
```
|
||||
|
||||
**Scoped honestly:** a genuine *crash* under `unless-stopped` is restarted by Docker and surfaces as
|
||||
`restarting` → the 5-minute crash-loop path, which does alarm. The silent case is an explicit
|
||||
out-of-band stop of a stack with no surviving member.
|
||||
Both the bad POST and an `error` control returned **200** (nothing lost); the control produced **no**
|
||||
warning.
|
||||
|
||||
**Filed as R-386 (OPEN — MEDIUM). Not fixed here, per §12 and §13.**
|
||||
## 10. The absent-intent count on `demo-hp`, in plain words
|
||||
|
||||
## 11. Evidence
|
||||
**Zero.** All **8** deployed apps carry `desired_state: running`; none is `stopped` and none is absent.
|
||||
The unknown-intent fallback therefore suppresses nothing on this box today — the population is already
|
||||
empty on a machine that has been exercised through the interface. It will be larger on a box upgraded
|
||||
and left alone, which is why the log line exists rather than a one-off count.
|
||||
|
||||
`felhom.eu/documentation/audits/DRILL-r384-dead-db-alarm-2026-08-23/evidence/` — 5 gate runs, 4
|
||||
red-proof transcripts, 20 live-walk files including two full controller-log windows (1808 and 2139
|
||||
lines) pulled off the guest. **Both log windows were copied off before the app was restarted**, per
|
||||
standing rule 5.
|
||||
## 11. The dead-branch decision, and the reason
|
||||
|
||||
## 12. Teardown, three layers, and the box's end state
|
||||
**KEPT.** `cmd/hub/main.go` wires `dispatcher.ProcessEvent` **directly** as the
|
||||
`monitor.EventNotifyFunc` for the staleness, host-staleness and offsite-box checkers — those events
|
||||
never pass the ingest handler, so for them that line is the only severity guard there is. Deleting it
|
||||
as "dead" would have removed the live half while the dead half supplied the justification. All 90
|
||||
severity literals in `internal/monitor` were verified already valid, so the guard is silent because
|
||||
the producers are correct.
|
||||
|
||||
1. **Guest 9201 / apps** — nothing provisioned. `bookstack-db` restarted and **`bookstack` confirmed
|
||||
healthy**; `privatebin` restarted and healthy; `docmost` (all 3) healthy. Planted data untouched
|
||||
throughout — no app was rebuilt, redeployed or restored.
|
||||
2. **Bake VM** — the drill VM on DooPlex ran the bake and its build guest 9100 is stopped inside it;
|
||||
the qcow2 reverts to the `virgin` snapshot. No storage was added anywhere, so `pvesm status` has
|
||||
nothing to compare.
|
||||
3. **Hub-side record — stated explicitly even though there is none.** No appliance was registered, no
|
||||
customer created, no config written, no artifact manifest changed. **The hub was READ ONLY**
|
||||
(`GET /configuration`, `GET /events`). Nothing to discard.
|
||||
## 12. Evidence
|
||||
|
||||
**End state:** `demo-hp` guest 9201 runs controller **0.222.0**, all 8 deployed apps healthy, golden
|
||||
0.222.0 baked and published but **NOT vouched** — floor still **0.221.1**.
|
||||
`felhom.eu/documentation/audits/DRILL-r329-r386-2026-08-23/evidence/` — 26 files: 5 red-proof
|
||||
transcripts, 20 live-walk files, two full controller-log windows (1041 and 4044 lines) **pulled off
|
||||
before each revert**, and the hub's own DB queries.
|
||||
|
||||
## 13. Register size
|
||||
## 13. Teardown, three layers, and the end state
|
||||
|
||||
1. **Guest 9201 / apps** — nothing provisioned. **All 17 app containers healthy.** `privatebin`'s
|
||||
`app.yaml` restored from backup and the backup deleted; intent reads `running`. `demo-hp`'s
|
||||
notification settings restored to `enabled_events: null`, no e-mail — their pre-drill state. No app
|
||||
rebuilt, redeployed or restored; planted data untouched.
|
||||
2. **Bake VM** — powered off, Gitea token and runner script **shredded**, `drill.qcow2` reverted to
|
||||
`virgin`. Build guest 9100 exists only inside that reverted snapshot. No storage added anywhere, so
|
||||
`pvesm status` has nothing to compare.
|
||||
3. **Hub-side, stated explicitly.** The hub was **written** this session, unlike last: the deployment
|
||||
is v0.107.0 via the manifest, and **two probe events remain as rows for `demo-hp`** from Scenario H
|
||||
(`backup_failed`, "R-387 scenario H probe" and "…control"). They are inert records; named here
|
||||
rather than left for someone to find. Nothing else: no appliance registered, no customer created,
|
||||
no artifact manifest changed, floor untouched.
|
||||
|
||||
**End state:** controller **0.223.0** and hub **0.107.0** deployed; golden **0.223.0** baked and
|
||||
published but **NOT vouched**; floor still **0.222.0**; all apps running; planted data present.
|
||||
|
||||
## 14. Register size
|
||||
|
||||
| File | Before | After |
|
||||
|---|---|---|
|
||||
| `OPEN-ITEMS.md` | 327,266 B | **328,325 B** |
|
||||
| `CLOSED-ITEMS.md` | 68,464 B | **71,441 B** |
|
||||
| `OPEN-ITEMS.md` | 328,325 B | **328,132 B** |
|
||||
| `CLOSED-ITEMS.md` | 71,441 B | **74,642 B** |
|
||||
|
||||
R-383 and R-384 moved to CLOSED compressed; R-385 (closed) and R-386 (open) filed. OPEN grew by
|
||||
~1 KB despite two closures because R-386 is a substantial new finding — recorded rather than
|
||||
smoothed over.
|
||||
R-329 and R-386 closed and compressed; **R-387** (closed) and **R-388** (the notification-model
|
||||
product decision — open, operator's call) filed.
|
||||
|
||||
## 14. Observations — noticed, documented, NOT acted on
|
||||
## 15. Observations — noticed, documented, NOT acted on
|
||||
|
||||
1. **R-329 is live and now matters much more.** `app_start_failed` is pushed with severity **`warn`**,
|
||||
which is not in the hub's vocabulary (`{info, warning, error, critical}`) and coerces silently to
|
||||
`info`, e-mailing nobody, while the POST still returns 200. Observed again today:
|
||||
`PushEvent: type=app_start_failed severity=warn`. **R-384 makes this event actually fire, so a
|
||||
known-broken severity moved from unreachable to load-bearing.** Not in scope; not touched.
|
||||
2. **Two files carry pre-existing `gofmt` drift** — `internal/backup/offbox.go` and
|
||||
`internal/backup/offbox_recovery_cli.go`. Confirmed pre-existing by stashing this session's work
|
||||
and re-running `gofmt -l`. Not touched (§12 forbids nearby refactors).
|
||||
3. **The runbook's golden-bake step is missing `pveam update`.** On the `virgin` snapshot the template
|
||||
index is stale, so `pveam available` offers `13.1-2` and downloading it fails with
|
||||
`400 Parameter verification failed. template: no such template`. Recorded in the bake evidence
|
||||
README; the runbook itself was not edited.
|
||||
4. **The register's own suggested fix for R-384 was wrong** — it proposed a sustained-`unhealthy`
|
||||
threshold on the `crashLoopAfter` model. The defect needed no threshold at all, only an ordering.
|
||||
Recorded in the CLOSED entry so the next reader sees that a register remedy is a hypothesis.
|
||||
5. **Deliberately left open, untouched:** R-102, R-359, R-361's sibling surfaces.
|
||||
1. **The operator cooldown key has no app identifier, and it now bites.** PrivateBin's alarm four
|
||||
minutes after BookStack's was logged `suppressed — operator cooldown 1h,
|
||||
key=demo-hp:app_start_failed`, so **only the first app-down per hour e-mails the operator**. This is
|
||||
R-182's known cooldown-key shape; it was harmless while `app_start_failed` was undeliverable and is
|
||||
not any more. **Same pattern as R-329 itself: a known-broken thing moved from unreachable to
|
||||
load-bearing.** Not fixed here.
|
||||
2. **The settings page grew 12 → 15 toggles in one session** — one new alarm, plus two compound
|
||||
toggles split into four. Recorded as the argument inside R-388.
|
||||
3. `internal/notify/notifier.go` carries **pre-existing** gofmt drift in an unrelated const block,
|
||||
confirmed by stashing this session's work and re-running `gofmt -l`. Not touched (§12).
|
||||
4. **The golden-bake runbook still lacks `pveam update`** — second consecutive bake to hit the stale
|
||||
index on the `virgin` snapshot, presenting as `400 … no such template`.
|
||||
5. **Deliberately left open, untouched:** R-102, R-359, R-385, and R-388's redesign.
|
||||
|
||||
### CI runs, confirmed by ID
|
||||
|
||||
| Commit | Repo | CI `id` | `run_number` | Result |
|
||||
|---|---|---|---|---|
|
||||
| `9832760` | felhom-controller | **408** | 88 | success |
|
||||
| `68a9f54` | felhom.eu | **409** | 261 | success |
|
||||
| `2f7c9a6` | felhom.eu | **410** | 262 | success |
|
||||
|
||||
(The controller docs commit's run is confirmed after its push and is the next `id` in that repo.)
|
||||
|
||||
Reference in New Issue
Block a user