# REPORT — controller v0.221.1 (record) + v0.222.0 (R-384, R-383) **Session 2026-08-23, UNATTENDED. Live leg on `demo-hp` (Tier 0, disposable).** > ## ⚠ HALT DECLARED — §4's measurement found a real defect, filed as R-386, NOT fixed > > The task's §4 asked for a measurement and named it a halt condition. It reproduced. > **A single-container app stopped out of band raises no alarm at all.** Details in §10 below. > Per §13 the fix was NOT attempted here. Live-walk step 4 (Scenario E) was dropped as a > consequence — it is item (3) on the task's own drop list. Everything else completed. --- ## 1. Baselines used, and the hub's four numbers as read | Repo | `main` at start | Verified | |---|---|---| | felhom-controller | `f7881787f434` | matches the task | | felhom.eu | `1eb64bec5183` | task said `4e488321bfd1+`; it had moved on | | felhom-agent | untouched | — | **Hub's own numbers, read live from `/configuration` (ClusterIP + Basic auth) 2026-08-23:** | Field | Value | |---|---| | `golden_version` | **0.221.1** | | `agent_version` | **0.130.0** | | `min_agent` | **0.129.0** | | controller floor (`min_controller_version`) | **0.221.1** | The task expected golden/floor **0.220.2**; the operator had already vouched **0.221.1** and raised the floor. Live controller on `demo-hp` at session start: **0.221.1** — so the running version, the golden and the floor all agreed, and only the RECORD disagreed. That is exactly R-385's shape. ## 2. Architecture documents read - `documentation/architecture/00-capability-map.md` — its 2026-08-22 paragraph already NAMED R-384 as an open finding, from the held-app measurement. - `felhom-controller/internal/stacks/manager.go` `IsDownState` + `aggregateState` + `supervisedPolicy` - `cmd/controller/main.go` `classifyRunStates` and its three suppressions - `internal/quiesce/suppress.go`, `internal/bootrecon/bootrecon.go`, `internal/stacks/desiredstate.go` - `felhom.eu/scripts/golden_currency_gate.py` (all 133 lines) **§2's conditional applies and is answered: NO document owned the alarm ladder.** That absence is reported as a finding, and `documentation/architecture/08-alarm-ladder.md` now owns it (Part N.4). It is why the ordering defect was legible only by reading one function top to bottom. ## 3. Part 0 — the record, pushed ALONE Exact heading written: ``` ## v0.221.1 — the undo-copy prune stopped running because another fix made its guard reachable (2026-08-23, R-361 follow-on) ``` Commit **`da75603`**, pushed alone before anything else. The reasoning was **moved verbatim** from the v0.221.0 entry (which no longer claims it), not rewritten and not duplicated. ## 4. Files changed, commits, CI runs | Commit | Contents | |---|---| | **`da75603`** | Part 0 — the `v0.221.1` heading, alone | | **`5da11c4`** | v0.222.0 — R-384 + R-383, tests, CHANGELOG | Modified: `CHANGELOG.md`, `internal/stacks/manager.go`, `internal/stacks/degraded_test.go`, `internal/backup/offbox_reconstitute.go`, `cmd/controller/r361_classifier_control_test.go`. Added: `cmd/controller/r384_dead_db_alarm_test.go`, `internal/backup/r383_undo_phrase_test.go`. **CI runs confirmed BY ID** (`id` and `run_number` diverge, both printed): | Commit | CI `id` | `run_number` | Result | |---|---|---|---| | `da75603` | **404** | 85 | success | | `5da11c4` | **405** | 86 | success | ## 5. Red-proofs — four planted, FOUR SEEN FAILING | # | Mutation | Layer the guard sits at | Observed failure | |---|---|---|---| | 1 | hoisted block moved back **below** `unhealthy > 0` | `aggregateState` — the ORDERING | `aggregateState = "unhealthy", want "degraded"` **and** `bookstack state = "unhealthy", want "degraded"` (production-path wiring) | | 2 | `up` narrowed back to `running` alone | `aggregateState` — the GUARD | same subtest, plus `"starting"` and `"restarting"` — all three survivor shapes convict | | 3 | classifier drops `StateDegraded` from `down` | `classifyRunStates` — the CONSEQUENCE | `dead-app banner = [], want exactly one entry for bookstack` | | 4 | `undoCopyPhrase` reverted to the unconditional claim | the phrase builder — where the CLAIM is made | `phrase "…mentése megvan: …mariadb.sql" contains "mentése megvan" — it asserts a file that is not on disk`; the empty set printed `megvan: .`, naming a file that never existed | **Every mutation was asserted to have applied** (the scripts `assert` the pre-fix text is present before rewriting and print `MUTATION APPLIED`). **None passed first time.** Mutations 1 and 2 convict independently, which is what proves the fix genuinely has two halves. ## 6. Test count **1494 → 1504** top-level test functions (measured by `go test ./... -list '.*'` on the stashed and unstashed tree, not estimated). Full green gate `go build && go vet && go test ./...` → **exit 0, zero failures**, run after Part 2 and again after Part 3. ## 7. Deployed version, and the golden ``` gitea.dooplex.hu/admin/felhom-controller:0.222.0 Up 20 seconds (healthy) ``` **Golden BAKED and PUBLISHED: YES — version 0.222.0.** `GOLDEN_SHA256 = 19f5904f53792684f046ec0bc25426645cb87ad73d5cfc6c03639d9f82706037`, `upload OK (HTTP 201)`, round-trip `HTTP 206` from the package URL. All acceptance markers counted and recorded. **VOUCHING IS THE OPERATOR'S ACT AND WAS NOT DONE HERE.** ## 8. The five `IsDownState` consumers, walked and named | Consumer | What changes | |---|---| | `cmd/controller/main.go:2173` `classifyRunStates` | **THE INTENDED CHANGE.** A stack that read `unhealthy` now reads `degraded` → `down=true` → banner + `app_start_failed`. `userStopped` tests `StateStopped` specifically, so the whitelist cannot swallow `degraded`. | | `internal/bootrecon/bootrecon.go:213` (DesiredStateRunning) | **CHANGES, and toward repair.** A half-started stack at boot now reads `degraded` → an orphan → `compose up -d`. Previously it read `unhealthy` → not an orphan → left half-dead. Aligned with the file's own stated intent. | | `internal/bootrecon/bootrecon.go:216` (legacy DesiredStateUnknown) | Same shape, same direction. | | `internal/bootrecon/bootrecon.go:288` (recovery check) | **CHANGES, and toward truth.** A stack that came back with a dead supervised member is no longer counted `Recovered`; it stays pending and is retried, bounded by `r.attempts`. It used to be declared recovered while half-dead. | | `internal/stacks/desiredstate.go:150` `isObservedUp` | **UNAFFECTED — verified, not assumed.** It is an allow-list of `{running, starting}`; neither `unhealthy` nor `degraded` was ever in it, so a stack moving between them does not cross the boundary. | | `internal/quiesce/suppress.go` | **UNAFFECTED.** It does not call `IsDownState` at all — the suppression is cycle-keyed and state-blind, which is precisely why R-97b's guarantee cannot be weakened by a state change. Pinned by `TestR384_QuiesceSuppressionStillHoldsForDegraded`. | | dashboard state badge | **Already handled.** R-51 wired `degraded` through `handlers.go:158/169` (counts with stopped) and `funcmap.go:255` (filters with stopped). Verified live — the badge rendered `(degraded)`. | ## 9. The live walk Method: endpoint-level. No browser exists on DooPlex; every read below is either the exact endpoint the UI calls (`POST /api/stacks//`, `GET /api/stacks`, `GET /dashboard`) or the controller's own log. Guest clock is UTC. ### Step 1 — Scenario A: a database dies behind a healthy-looking app ✅ | Observable | Result | |---|---| | `bookstack-db` stopped out of band | 05:30:07Z | | front end went `unhealthy` | 05:31:21Z — **the state that used to swallow the alarm** | | aggregate state read | **`degraded`** while the front end was `unhealthy` (05:32:03Z) | | `app_start_failed` | **fired at 05:30:14Z**, 7 s after the stop | | new code path visible | `manager.go:703: restart-policy of down member "bookstack-db" = "unless-stopped"` | | banner | *„Telepített alkalmazás nem fut: BookStack (degraded)"* on **both** `/launcher` and `/dashboard` | | edge-triggered, not per-scan | **1 event across 22 scans** | **The heartbeat, old beside new — same 8 apps evaluated:** ``` 2026/08/22 21:13:48 [deadapp] check alive: 580 scans since boot, 8 deployed app(s) evaluated, 0 currently down ← v0.220.2/0.221.1 2026/08/23 05:37:44 [deadapp] check alive: 20 scans since boot, 8 deployed app(s) evaluated, 1 currently down ← v0.222.0 ``` ### Step 2 — Scenario B: unhealthy with nothing dead ✅ The database was restarted; the front end stayed `unhealthy` with nothing down. Aggregate read **`unhealthy`**, not `degraded`. **0 new `app_start_failed`**, and — the positive observable — **no `restart-policy of down member` line at all**, meaning the supervised path was not entered. That absence is trustworthy because the same line HAD appeared on this box 10 minutes earlier. Banner cleared; `bookstack state=running`. *Honest limit:* the live window in which docker reported the front end unhealthy with the database up was ~11 s wide, and the controller's cached read was taken at its edge. The three-shape unit test `TestR384_UnhealthyWithNothingDeadDoesNotAlarm` carries the rest of this case. ### Step 3 — Scenario D: a full deploy cycle ✅ **0 alarms** `POST /api/stacks/docmost/stop` then `/start` — the exact calls the launcher's buttons make — on a 3-container stack, watched for 5 minutes to settled healthy. | Observable | Result | |---|---| | `app_start_failed` across the cycle | **0** | | dead-app scans in the window | **9** | | supervised-down path entered | 1, at 05:42:34Z (docmost itself momentarily down beside two live members) | **Alarms that v0.221.1 would NOT have produced: ZERO.** The single `degraded` reading at 05:42:34Z has `running > 0`, so v0.221.1's mixed-case branch reaches the identical verdict. No moment in the cycle had a down supervised member with only non-`running` survivors, which is the only shape where the two versions differ. ### Step 4 — Scenario E: the quiesce cycle — **DROPPED, and why** Dropped as a direct consequence of the §4 halt (see the banner at the top), and it is item **(3)** on the task's own drop list. Running it would have meant triggering a real backup cycle on the box *after* a halt condition had already fired. **Covered at unit level instead** by `TestR384_QuiesceSuppressionStillHoldsForDegraded`, which asserts a `degraded` stack inside the quiesce set produces no banner and `Down=false`. **Not proven live in this session — stated plainly rather than implied.** ### Step 5 — §4's measurement ✅ (it reproduced — see §10) ### Step 6 — Part 1's gate, both directions ✅ | Run | Gate | CHANGELOG | Golden | Exit | |---|---|---|---|---| | `gate-01` | **old** | v0.221.0 | 0.221.1 | **0** — the blindness, on the real history | | `gate-02` | **new** | v0.221.0 | 0.221.1 | **1** — convicted, naming the missing heading and the route | | `gate-03` | new | v0.221.1 | 0.221.1 | 0 | | `gate-04` | new | absent clone | — | **2** — INCONCLUSIVE preserved | | `gate-05` | new | v0.222.0 | 0.222.0 | 0 — post-bake | ## 10. §4's answer, in plain words **A single-container app that is stopped out of band raises no alarm at all, and a comment in the code says the opposite.** `aggregateState` folds `StateExited` into the `stopped` counter, so when every member is down it returns `StateStopped` — **`StateExited` never survives aggregation**, which is the path the task suspected and could not find. `classifyRunStates` then whitelists `StateStopped` as a deliberate user stop. So the comment at `cmd/controller/main.go` — *"An out-of-band `docker compose stop` leaves the containers present → StateExited → still alerts, which is correct: out-of-band tampering IS reportable"* — is **false**, and so is the neighbouring I2 claim that a crashing app never comes to rest at `stopped`. **Measured:** `privatebin` (1 container, `unless-stopped`) stopped 05:47:35Z. At 05:51:53Z: `state=stopped`, **9 dead-app scans had run, 0 events, 0 banner lines.** **Positive control first, per standing rule 3:** `app_start_failed` fired for BookStack at 05:30:14Z on the same box 17 minutes earlier, so the detector was demonstrably alive. **Scoped honestly:** a genuine *crash* under `unless-stopped` is restarted by Docker and surfaces as `restarting` → the 5-minute crash-loop path, which does alarm. The silent case is an explicit out-of-band stop of a stack with no surviving member. **Filed as R-386 (OPEN — MEDIUM). Not fixed here, per §12 and §13.** ## 11. Evidence `felhom.eu/documentation/audits/DRILL-r384-dead-db-alarm-2026-08-23/evidence/` — 5 gate runs, 4 red-proof transcripts, 20 live-walk files including two full controller-log windows (1808 and 2139 lines) pulled off the guest. **Both log windows were copied off before the app was restarted**, per standing rule 5. ## 12. Teardown, three layers, and the box's end state 1. **Guest 9201 / apps** — nothing provisioned. `bookstack-db` restarted and **`bookstack` confirmed healthy**; `privatebin` restarted and healthy; `docmost` (all 3) healthy. Planted data untouched throughout — no app was rebuilt, redeployed or restored. 2. **Bake VM** — the drill VM on DooPlex ran the bake and its build guest 9100 is stopped inside it; the qcow2 reverts to the `virgin` snapshot. No storage was added anywhere, so `pvesm status` has nothing to compare. 3. **Hub-side record — stated explicitly even though there is none.** No appliance was registered, no customer created, no config written, no artifact manifest changed. **The hub was READ ONLY** (`GET /configuration`, `GET /events`). Nothing to discard. **End state:** `demo-hp` guest 9201 runs controller **0.222.0**, all 8 deployed apps healthy, golden 0.222.0 baked and published but **NOT vouched** — floor still **0.221.1**. ## 13. Register size | File | Before | After | |---|---|---| | `OPEN-ITEMS.md` | 327,266 B | **328,325 B** | | `CLOSED-ITEMS.md` | 68,464 B | **71,441 B** | R-383 and R-384 moved to CLOSED compressed; R-385 (closed) and R-386 (open) filed. OPEN grew by ~1 KB despite two closures because R-386 is a substantial new finding — recorded rather than smoothed over. ## 14. Observations — noticed, documented, NOT acted on 1. **R-329 is live and now matters much more.** `app_start_failed` is pushed with severity **`warn`**, which is not in the hub's vocabulary (`{info, warning, error, critical}`) and coerces silently to `info`, e-mailing nobody, while the POST still returns 200. Observed again today: `PushEvent: type=app_start_failed severity=warn`. **R-384 makes this event actually fire, so a known-broken severity moved from unreachable to load-bearing.** Not in scope; not touched. 2. **Two files carry pre-existing `gofmt` drift** — `internal/backup/offbox.go` and `internal/backup/offbox_recovery_cli.go`. Confirmed pre-existing by stashing this session's work and re-running `gofmt -l`. Not touched (§12 forbids nearby refactors). 3. **The runbook's golden-bake step is missing `pveam update`.** On the `virgin` snapshot the template index is stale, so `pveam available` offers `13.1-2` and downloading it fails with `400 Parameter verification failed. template: no such template`. Recorded in the bake evidence README; the runbook itself was not edited. 4. **The register's own suggested fix for R-384 was wrong** — it proposed a sustained-`unhealthy` threshold on the `crashLoopAfter` model. The defect needed no threshold at all, only an ordering. Recorded in the CLOSED entry so the next reader sees that a register remedy is a hypothesis. 5. **Deliberately left open, untouched:** R-102, R-359, R-361's sibling surfaces.