REPORT.md overwritten with the full run: baselines and the hub's four numbers, the four red-proofs with the mutation and observed text for each, the five IsDownState consumers walked and named, the live walk in full with the old and new heartbeat lines quoted side by side, and the halt. CONTEXT records the decision - a dead supervised member is asked about before a failing healthcheck, because they are different questions and the second was answering the first - plus the fence that IsDownState did not move, the trap that three existing subtests pinned the defect, and R-386. README gains the `degraded` row, which the state table never had, and a note that the ORDER is load-bearing. Points at the new alarm-ladder architecture doc.
15 KiB
REPORT — controller v0.221.1 (record) + v0.222.0 (R-384, R-383)
Session 2026-08-23, UNATTENDED. Live leg on demo-hp (Tier 0, disposable).
⚠ HALT DECLARED — §4's measurement found a real defect, filed as R-386, NOT fixed
The task's §4 asked for a measurement and named it a halt condition. It reproduced. A single-container app stopped out of band raises no alarm at all. Details in §10 below. Per §13 the fix was NOT attempted here. Live-walk step 4 (Scenario E) was dropped as a consequence — it is item (3) on the task's own drop list. Everything else completed.
1. Baselines used, and the hub's four numbers as read
| Repo | main at start |
Verified |
|---|---|---|
| felhom-controller | f7881787f434 |
matches the task |
| felhom.eu | 1eb64bec5183 |
task said 4e488321bfd1+; it had moved on |
| felhom-agent | untouched | — |
Hub's own numbers, read live from /configuration (ClusterIP + Basic auth) 2026-08-23:
| Field | Value |
|---|---|
golden_version |
0.221.1 |
agent_version |
0.130.0 |
min_agent |
0.129.0 |
controller floor (min_controller_version) |
0.221.1 |
The task expected golden/floor 0.220.2; the operator had already vouched 0.221.1 and raised
the floor. Live controller on demo-hp at session start: 0.221.1 — so the running version, the
golden and the floor all agreed, and only the RECORD disagreed. That is exactly R-385's shape.
2. Architecture documents read
documentation/architecture/00-capability-map.md— its 2026-08-22 paragraph already NAMED R-384 as an open finding, from the held-app measurement.felhom-controller/internal/stacks/manager.goIsDownState+aggregateState+supervisedPolicycmd/controller/main.goclassifyRunStatesand its three suppressionsinternal/quiesce/suppress.go,internal/bootrecon/bootrecon.go,internal/stacks/desiredstate.gofelhom.eu/scripts/golden_currency_gate.py(all 133 lines)
§2's conditional applies and is answered: NO document owned the alarm ladder. That absence is
reported as a finding, and documentation/architecture/08-alarm-ladder.md now owns it (Part N.4).
It is why the ordering defect was legible only by reading one function top to bottom.
3. Part 0 — the record, pushed ALONE
Exact heading written:
## v0.221.1 — the undo-copy prune stopped running because another fix made its guard reachable (2026-08-23, R-361 follow-on)
Commit da75603, pushed alone before anything else. The reasoning was moved verbatim from the
v0.221.0 entry (which no longer claims it), not rewritten and not duplicated.
4. Files changed, commits, CI runs
| Commit | Contents |
|---|---|
da75603 |
Part 0 — the v0.221.1 heading, alone |
5da11c4 |
v0.222.0 — R-384 + R-383, tests, CHANGELOG |
Modified: CHANGELOG.md, internal/stacks/manager.go, internal/stacks/degraded_test.go,
internal/backup/offbox_reconstitute.go, cmd/controller/r361_classifier_control_test.go.
Added: cmd/controller/r384_dead_db_alarm_test.go, internal/backup/r383_undo_phrase_test.go.
CI runs confirmed BY ID (id and run_number diverge, both printed):
| Commit | CI id |
run_number |
Result |
|---|---|---|---|
da75603 |
404 | 85 | success |
5da11c4 |
405 | 86 | success |
5. Red-proofs — four planted, FOUR SEEN FAILING
| # | Mutation | Layer the guard sits at | Observed failure |
|---|---|---|---|
| 1 | hoisted block moved back below unhealthy > 0 |
aggregateState — the ORDERING |
aggregateState = "unhealthy", want "degraded" and bookstack state = "unhealthy", want "degraded" (production-path wiring) |
| 2 | up narrowed back to running alone |
aggregateState — the GUARD |
same subtest, plus "starting" and "restarting" — all three survivor shapes convict |
| 3 | classifier drops StateDegraded from down |
classifyRunStates — the CONSEQUENCE |
dead-app banner = [], want exactly one entry for bookstack |
| 4 | undoCopyPhrase reverted to the unconditional claim |
the phrase builder — where the CLAIM is made | phrase "…mentése megvan: …mariadb.sql" contains "mentése megvan" — it asserts a file that is not on disk; the empty set printed megvan: ., naming a file that never existed |
Every mutation was asserted to have applied (the scripts assert the pre-fix text is present
before rewriting and print MUTATION APPLIED). None passed first time. Mutations 1 and 2 convict
independently, which is what proves the fix genuinely has two halves.
6. Test count
1494 → 1504 top-level test functions (measured by go test ./... -list '.*' on the stashed and
unstashed tree, not estimated). Full green gate go build && go vet && go test ./... → exit 0, zero
failures, run after Part 2 and again after Part 3.
7. Deployed version, and the golden
gitea.dooplex.hu/admin/felhom-controller:0.222.0 Up 20 seconds (healthy)
Golden BAKED and PUBLISHED: YES — version 0.222.0.
GOLDEN_SHA256 = 19f5904f53792684f046ec0bc25426645cb87ad73d5cfc6c03639d9f82706037, upload OK (HTTP 201),
round-trip HTTP 206 from the package URL. All acceptance markers counted and recorded.
VOUCHING IS THE OPERATOR'S ACT AND WAS NOT DONE HERE.
8. The five IsDownState consumers, walked and named
| Consumer | What changes |
|---|---|
cmd/controller/main.go:2173 classifyRunStates |
THE INTENDED CHANGE. A stack that read unhealthy now reads degraded → down=true → banner + app_start_failed. userStopped tests StateStopped specifically, so the whitelist cannot swallow degraded. |
internal/bootrecon/bootrecon.go:213 (DesiredStateRunning) |
CHANGES, and toward repair. A half-started stack at boot now reads degraded → an orphan → compose up -d. Previously it read unhealthy → not an orphan → left half-dead. Aligned with the file's own stated intent. |
internal/bootrecon/bootrecon.go:216 (legacy DesiredStateUnknown) |
Same shape, same direction. |
internal/bootrecon/bootrecon.go:288 (recovery check) |
CHANGES, and toward truth. A stack that came back with a dead supervised member is no longer counted Recovered; it stays pending and is retried, bounded by r.attempts. It used to be declared recovered while half-dead. |
internal/stacks/desiredstate.go:150 isObservedUp |
UNAFFECTED — verified, not assumed. It is an allow-list of {running, starting}; neither unhealthy nor degraded was ever in it, so a stack moving between them does not cross the boundary. |
internal/quiesce/suppress.go |
UNAFFECTED. It does not call IsDownState at all — the suppression is cycle-keyed and state-blind, which is precisely why R-97b's guarantee cannot be weakened by a state change. Pinned by TestR384_QuiesceSuppressionStillHoldsForDegraded. |
| dashboard state badge | Already handled. R-51 wired degraded through handlers.go:158/169 (counts with stopped) and funcmap.go:255 (filters with stopped). Verified live — the badge rendered (degraded). |
9. The live walk
Method: endpoint-level. No browser exists on DooPlex; every read below is either the exact endpoint
the UI calls (POST /api/stacks/<name>/<action>, GET /api/stacks, GET /dashboard) or the
controller's own log. Guest clock is UTC.
Step 1 — Scenario A: a database dies behind a healthy-looking app ✅
| Observable | Result |
|---|---|
bookstack-db stopped out of band |
05:30:07Z |
front end went unhealthy |
05:31:21Z — the state that used to swallow the alarm |
| aggregate state read | degraded while the front end was unhealthy (05:32:03Z) |
app_start_failed |
fired at 05:30:14Z, 7 s after the stop |
| new code path visible | manager.go:703: restart-policy of down member "bookstack-db" = "unless-stopped" |
| banner | „Telepített alkalmazás nem fut: BookStack (degraded)" on both /launcher and /dashboard |
| edge-triggered, not per-scan | 1 event across 22 scans |
The heartbeat, old beside new — same 8 apps evaluated:
2026/08/22 21:13:48 [deadapp] check alive: 580 scans since boot, 8 deployed app(s) evaluated, 0 currently down ← v0.220.2/0.221.1
2026/08/23 05:37:44 [deadapp] check alive: 20 scans since boot, 8 deployed app(s) evaluated, 1 currently down ← v0.222.0
Step 2 — Scenario B: unhealthy with nothing dead ✅
The database was restarted; the front end stayed unhealthy with nothing down. Aggregate read
unhealthy, not degraded. 0 new app_start_failed, and — the positive observable —
no restart-policy of down member line at all, meaning the supervised path was not entered. That
absence is trustworthy because the same line HAD appeared on this box 10 minutes earlier.
Banner cleared; bookstack state=running.
Honest limit: the live window in which docker reported the front end unhealthy with the database up
was ~11 s wide, and the controller's cached read was taken at its edge. The three-shape unit test
TestR384_UnhealthyWithNothingDeadDoesNotAlarm carries the rest of this case.
Step 3 — Scenario D: a full deploy cycle ✅ 0 alarms
POST /api/stacks/docmost/stop then /start — the exact calls the launcher's buttons make — on a
3-container stack, watched for 5 minutes to settled healthy.
| Observable | Result |
|---|---|
app_start_failed across the cycle |
0 |
| dead-app scans in the window | 9 |
| supervised-down path entered | 1, at 05:42:34Z (docmost itself momentarily down beside two live members) |
Alarms that v0.221.1 would NOT have produced: ZERO. The single degraded reading at 05:42:34Z
has running > 0, so v0.221.1's mixed-case branch reaches the identical verdict. No moment in the
cycle had a down supervised member with only non-running survivors, which is the only shape where
the two versions differ.
Step 4 — Scenario E: the quiesce cycle — DROPPED, and why
Dropped as a direct consequence of the §4 halt (see the banner at the top), and it is item (3) on
the task's own drop list. Running it would have meant triggering a real backup cycle on the box
after a halt condition had already fired. Covered at unit level instead by
TestR384_QuiesceSuppressionStillHoldsForDegraded, which asserts a degraded stack inside the
quiesce set produces no banner and Down=false. Not proven live in this session — stated plainly
rather than implied.
Step 5 — §4's measurement ✅ (it reproduced — see §10)
Step 6 — Part 1's gate, both directions ✅
| Run | Gate | CHANGELOG | Golden | Exit |
|---|---|---|---|---|
gate-01 |
old | v0.221.0 | 0.221.1 | 0 — the blindness, on the real history |
gate-02 |
new | v0.221.0 | 0.221.1 | 1 — convicted, naming the missing heading and the route |
gate-03 |
new | v0.221.1 | 0.221.1 | 0 |
gate-04 |
new | absent clone | — | 2 — INCONCLUSIVE preserved |
gate-05 |
new | v0.222.0 | 0.222.0 | 0 — post-bake |
10. §4's answer, in plain words
A single-container app that is stopped out of band raises no alarm at all, and a comment in the code says the opposite.
aggregateState folds StateExited into the stopped counter, so when every member is down it
returns StateStopped — StateExited never survives aggregation, which is the path the task
suspected and could not find. classifyRunStates then whitelists StateStopped as a deliberate user
stop. So the comment at cmd/controller/main.go — "An out-of-band docker compose stop leaves the
containers present → StateExited → still alerts, which is correct: out-of-band tampering IS
reportable" — is false, and so is the neighbouring I2 claim that a crashing app never comes to
rest at stopped.
Measured: privatebin (1 container, unless-stopped) stopped 05:47:35Z. At 05:51:53Z:
state=stopped, 9 dead-app scans had run, 0 events, 0 banner lines.
Positive control first, per standing rule 3: app_start_failed fired for BookStack at 05:30:14Z
on the same box 17 minutes earlier, so the detector was demonstrably alive.
Scoped honestly: a genuine crash under unless-stopped is restarted by Docker and surfaces as
restarting → the 5-minute crash-loop path, which does alarm. The silent case is an explicit
out-of-band stop of a stack with no surviving member.
Filed as R-386 (OPEN — MEDIUM). Not fixed here, per §12 and §13.
11. Evidence
felhom.eu/documentation/audits/DRILL-r384-dead-db-alarm-2026-08-23/evidence/ — 5 gate runs, 4
red-proof transcripts, 20 live-walk files including two full controller-log windows (1808 and 2139
lines) pulled off the guest. Both log windows were copied off before the app was restarted, per
standing rule 5.
12. Teardown, three layers, and the box's end state
- Guest 9201 / apps — nothing provisioned.
bookstack-dbrestarted andbookstackconfirmed healthy;privatebinrestarted and healthy;docmost(all 3) healthy. Planted data untouched throughout — no app was rebuilt, redeployed or restored. - Bake VM — the drill VM on DooPlex ran the bake and its build guest 9100 is stopped inside it;
the qcow2 reverts to the
virginsnapshot. No storage was added anywhere, sopvesm statushas nothing to compare. - Hub-side record — stated explicitly even though there is none. No appliance was registered, no
customer created, no config written, no artifact manifest changed. The hub was READ ONLY
(
GET /configuration,GET /events). Nothing to discard.
End state: demo-hp guest 9201 runs controller 0.222.0, all 8 deployed apps healthy, golden
0.222.0 baked and published but NOT vouched — floor still 0.221.1.
13. Register size
| File | Before | After |
|---|---|---|
OPEN-ITEMS.md |
327,266 B | 328,325 B |
CLOSED-ITEMS.md |
68,464 B | 71,441 B |
R-383 and R-384 moved to CLOSED compressed; R-385 (closed) and R-386 (open) filed. OPEN grew by ~1 KB despite two closures because R-386 is a substantial new finding — recorded rather than smoothed over.
14. Observations — noticed, documented, NOT acted on
- R-329 is live and now matters much more.
app_start_failedis pushed with severitywarn, which is not in the hub's vocabulary ({info, warning, error, critical}) and coerces silently toinfo, e-mailing nobody, while the POST still returns 200. Observed again today:PushEvent: type=app_start_failed severity=warn. R-384 makes this event actually fire, so a known-broken severity moved from unreachable to load-bearing. Not in scope; not touched. - Two files carry pre-existing
gofmtdrift —internal/backup/offbox.goandinternal/backup/offbox_recovery_cli.go. Confirmed pre-existing by stashing this session's work and re-runninggofmt -l. Not touched (§12 forbids nearby refactors). - The runbook's golden-bake step is missing
pveam update. On thevirginsnapshot the template index is stale, sopveam availableoffers13.1-2and downloading it fails with400 Parameter verification failed. template: no such template. Recorded in the bake evidence README; the runbook itself was not edited. - The register's own suggested fix for R-384 was wrong — it proposed a sustained-
unhealthythreshold on thecrashLoopAftermodel. The defect needed no threshold at all, only an ordering. Recorded in the CLOSED entry so the next reader sees that a register remedy is a hypothesis. - Deliberately left open, untouched: R-102, R-359, R-361's sibling surfaces.