Files
felhom-controller/REPORT.md
T
admin 14137efac5
gates / gates (push) Successful in 11s
docs(v0.222.0): REPORT, CONTEXT decisions, README state table + the ordering
REPORT.md overwritten with the full run: baselines and the hub's four numbers,
the four red-proofs with the mutation and observed text for each, the five
IsDownState consumers walked and named, the live walk in full with the old and
new heartbeat lines quoted side by side, and the halt.

CONTEXT records the decision - a dead supervised member is asked about before a
failing healthcheck, because they are different questions and the second was
answering the first - plus the fence that IsDownState did not move, the trap
that three existing subtests pinned the defect, and R-386.

README gains the `degraded` row, which the state table never had, and a note
that the ORDER is load-bearing. Points at the new alarm-ladder architecture doc.
2026-08-23 07:58:09 +02:00

15 KiB

REPORT — controller v0.221.1 (record) + v0.222.0 (R-384, R-383)

Session 2026-08-23, UNATTENDED. Live leg on demo-hp (Tier 0, disposable).

⚠ HALT DECLARED — §4's measurement found a real defect, filed as R-386, NOT fixed

The task's §4 asked for a measurement and named it a halt condition. It reproduced. A single-container app stopped out of band raises no alarm at all. Details in §10 below. Per §13 the fix was NOT attempted here. Live-walk step 4 (Scenario E) was dropped as a consequence — it is item (3) on the task's own drop list. Everything else completed.


1. Baselines used, and the hub's four numbers as read

Repo main at start Verified
felhom-controller f7881787f434 matches the task
felhom.eu 1eb64bec5183 task said 4e488321bfd1+; it had moved on
felhom-agent untouched —

Hub's own numbers, read live from /configuration (ClusterIP + Basic auth) 2026-08-23:

Field Value
golden_version 0.221.1
agent_version 0.130.0
min_agent 0.129.0
controller floor (min_controller_version) 0.221.1

The task expected golden/floor 0.220.2; the operator had already vouched 0.221.1 and raised the floor. Live controller on demo-hp at session start: 0.221.1 — so the running version, the golden and the floor all agreed, and only the RECORD disagreed. That is exactly R-385's shape.

2. Architecture documents read

  • documentation/architecture/00-capability-map.md — its 2026-08-22 paragraph already NAMED R-384 as an open finding, from the held-app measurement.
  • felhom-controller/internal/stacks/manager.go IsDownState + aggregateState + supervisedPolicy
  • cmd/controller/main.go classifyRunStates and its three suppressions
  • internal/quiesce/suppress.go, internal/bootrecon/bootrecon.go, internal/stacks/desiredstate.go
  • felhom.eu/scripts/golden_currency_gate.py (all 133 lines)

§2's conditional applies and is answered: NO document owned the alarm ladder. That absence is reported as a finding, and documentation/architecture/08-alarm-ladder.md now owns it (Part N.4). It is why the ordering defect was legible only by reading one function top to bottom.

3. Part 0 — the record, pushed ALONE

Exact heading written:

## v0.221.1 — the undo-copy prune stopped running because another fix made its guard reachable (2026-08-23, R-361 follow-on)

Commit da75603, pushed alone before anything else. The reasoning was moved verbatim from the v0.221.0 entry (which no longer claims it), not rewritten and not duplicated.

4. Files changed, commits, CI runs

Commit Contents
da75603 Part 0 — the v0.221.1 heading, alone
5da11c4 v0.222.0 — R-384 + R-383, tests, CHANGELOG

Modified: CHANGELOG.md, internal/stacks/manager.go, internal/stacks/degraded_test.go, internal/backup/offbox_reconstitute.go, cmd/controller/r361_classifier_control_test.go. Added: cmd/controller/r384_dead_db_alarm_test.go, internal/backup/r383_undo_phrase_test.go.

CI runs confirmed BY ID (id and run_number diverge, both printed):

Commit CI id run_number Result
da75603 404 85 success
5da11c4 405 86 success

5. Red-proofs — four planted, FOUR SEEN FAILING

# Mutation Layer the guard sits at Observed failure
1 hoisted block moved back below unhealthy > 0 aggregateState — the ORDERING aggregateState = "unhealthy", want "degraded" and bookstack state = "unhealthy", want "degraded" (production-path wiring)
2 up narrowed back to running alone aggregateState — the GUARD same subtest, plus "starting" and "restarting" — all three survivor shapes convict
3 classifier drops StateDegraded from down classifyRunStates — the CONSEQUENCE dead-app banner = [], want exactly one entry for bookstack
4 undoCopyPhrase reverted to the unconditional claim the phrase builder — where the CLAIM is made phrase "…mentése megvan: …mariadb.sql" contains "mentése megvan" — it asserts a file that is not on disk; the empty set printed megvan: ., naming a file that never existed

Every mutation was asserted to have applied (the scripts assert the pre-fix text is present before rewriting and print MUTATION APPLIED). None passed first time. Mutations 1 and 2 convict independently, which is what proves the fix genuinely has two halves.

6. Test count

1494 → 1504 top-level test functions (measured by go test ./... -list '.*' on the stashed and unstashed tree, not estimated). Full green gate go build && go vet && go test ./... → exit 0, zero failures, run after Part 2 and again after Part 3.

7. Deployed version, and the golden

gitea.dooplex.hu/admin/felhom-controller:0.222.0 Up 20 seconds (healthy)

Golden BAKED and PUBLISHED: YES — version 0.222.0. GOLDEN_SHA256 = 19f5904f53792684f046ec0bc25426645cb87ad73d5cfc6c03639d9f82706037, upload OK (HTTP 201), round-trip HTTP 206 from the package URL. All acceptance markers counted and recorded. VOUCHING IS THE OPERATOR'S ACT AND WAS NOT DONE HERE.

8. The five IsDownState consumers, walked and named

Consumer What changes
cmd/controller/main.go:2173 classifyRunStates THE INTENDED CHANGE. A stack that read unhealthy now reads degraded → down=true → banner + app_start_failed. userStopped tests StateStopped specifically, so the whitelist cannot swallow degraded.
internal/bootrecon/bootrecon.go:213 (DesiredStateRunning) CHANGES, and toward repair. A half-started stack at boot now reads degraded → an orphan → compose up -d. Previously it read unhealthy → not an orphan → left half-dead. Aligned with the file's own stated intent.
internal/bootrecon/bootrecon.go:216 (legacy DesiredStateUnknown) Same shape, same direction.
internal/bootrecon/bootrecon.go:288 (recovery check) CHANGES, and toward truth. A stack that came back with a dead supervised member is no longer counted Recovered; it stays pending and is retried, bounded by r.attempts. It used to be declared recovered while half-dead.
internal/stacks/desiredstate.go:150 isObservedUp UNAFFECTED — verified, not assumed. It is an allow-list of {running, starting}; neither unhealthy nor degraded was ever in it, so a stack moving between them does not cross the boundary.
internal/quiesce/suppress.go UNAFFECTED. It does not call IsDownState at all — the suppression is cycle-keyed and state-blind, which is precisely why R-97b's guarantee cannot be weakened by a state change. Pinned by TestR384_QuiesceSuppressionStillHoldsForDegraded.
dashboard state badge Already handled. R-51 wired degraded through handlers.go:158/169 (counts with stopped) and funcmap.go:255 (filters with stopped). Verified live — the badge rendered (degraded).

9. The live walk

Method: endpoint-level. No browser exists on DooPlex; every read below is either the exact endpoint the UI calls (POST /api/stacks/<name>/<action>, GET /api/stacks, GET /dashboard) or the controller's own log. Guest clock is UTC.

Step 1 — Scenario A: a database dies behind a healthy-looking app ✅

Observable Result
bookstack-db stopped out of band 05:30:07Z
front end went unhealthy 05:31:21Z — the state that used to swallow the alarm
aggregate state read degraded while the front end was unhealthy (05:32:03Z)
app_start_failed fired at 05:30:14Z, 7 s after the stop
new code path visible manager.go:703: restart-policy of down member "bookstack-db" = "unless-stopped"
banner „Telepített alkalmazás nem fut: BookStack (degraded)" on both /launcher and /dashboard
edge-triggered, not per-scan 1 event across 22 scans

The heartbeat, old beside new — same 8 apps evaluated:

2026/08/22 21:13:48  [deadapp] check alive: 580 scans since boot, 8 deployed app(s) evaluated, 0 currently down   ← v0.220.2/0.221.1
2026/08/23 05:37:44  [deadapp] check alive:  20 scans since boot, 8 deployed app(s) evaluated, 1 currently down   ← v0.222.0

Step 2 — Scenario B: unhealthy with nothing dead ✅

The database was restarted; the front end stayed unhealthy with nothing down. Aggregate read unhealthy, not degraded. 0 new app_start_failed, and — the positive observable — no restart-policy of down member line at all, meaning the supervised path was not entered. That absence is trustworthy because the same line HAD appeared on this box 10 minutes earlier. Banner cleared; bookstack state=running.

Honest limit: the live window in which docker reported the front end unhealthy with the database up was ~11 s wide, and the controller's cached read was taken at its edge. The three-shape unit test TestR384_UnhealthyWithNothingDeadDoesNotAlarm carries the rest of this case.

Step 3 — Scenario D: a full deploy cycle ✅ 0 alarms

POST /api/stacks/docmost/stop then /start — the exact calls the launcher's buttons make — on a 3-container stack, watched for 5 minutes to settled healthy.

Observable Result
app_start_failed across the cycle 0
dead-app scans in the window 9
supervised-down path entered 1, at 05:42:34Z (docmost itself momentarily down beside two live members)

Alarms that v0.221.1 would NOT have produced: ZERO. The single degraded reading at 05:42:34Z has running > 0, so v0.221.1's mixed-case branch reaches the identical verdict. No moment in the cycle had a down supervised member with only non-running survivors, which is the only shape where the two versions differ.

Step 4 — Scenario E: the quiesce cycle — DROPPED, and why

Dropped as a direct consequence of the §4 halt (see the banner at the top), and it is item (3) on the task's own drop list. Running it would have meant triggering a real backup cycle on the box after a halt condition had already fired. Covered at unit level instead by TestR384_QuiesceSuppressionStillHoldsForDegraded, which asserts a degraded stack inside the quiesce set produces no banner and Down=false. Not proven live in this session — stated plainly rather than implied.

Step 5 — §4's measurement ✅ (it reproduced — see §10)

Step 6 — Part 1's gate, both directions ✅

Run Gate CHANGELOG Golden Exit
gate-01 old v0.221.0 0.221.1 0 — the blindness, on the real history
gate-02 new v0.221.0 0.221.1 1 — convicted, naming the missing heading and the route
gate-03 new v0.221.1 0.221.1 0
gate-04 new absent clone — 2 — INCONCLUSIVE preserved
gate-05 new v0.222.0 0.222.0 0 — post-bake

10. §4's answer, in plain words

A single-container app that is stopped out of band raises no alarm at all, and a comment in the code says the opposite.

aggregateState folds StateExited into the stopped counter, so when every member is down it returns StateStopped — StateExited never survives aggregation, which is the path the task suspected and could not find. classifyRunStates then whitelists StateStopped as a deliberate user stop. So the comment at cmd/controller/main.go — "An out-of-band docker compose stop leaves the containers present → StateExited → still alerts, which is correct: out-of-band tampering IS reportable" — is false, and so is the neighbouring I2 claim that a crashing app never comes to rest at stopped.

Measured: privatebin (1 container, unless-stopped) stopped 05:47:35Z. At 05:51:53Z: state=stopped, 9 dead-app scans had run, 0 events, 0 banner lines. Positive control first, per standing rule 3: app_start_failed fired for BookStack at 05:30:14Z on the same box 17 minutes earlier, so the detector was demonstrably alive.

Scoped honestly: a genuine crash under unless-stopped is restarted by Docker and surfaces as restarting → the 5-minute crash-loop path, which does alarm. The silent case is an explicit out-of-band stop of a stack with no surviving member.

Filed as R-386 (OPEN — MEDIUM). Not fixed here, per §12 and §13.

11. Evidence

felhom.eu/documentation/audits/DRILL-r384-dead-db-alarm-2026-08-23/evidence/ — 5 gate runs, 4 red-proof transcripts, 20 live-walk files including two full controller-log windows (1808 and 2139 lines) pulled off the guest. Both log windows were copied off before the app was restarted, per standing rule 5.

12. Teardown, three layers, and the box's end state

  1. Guest 9201 / apps — nothing provisioned. bookstack-db restarted and bookstack confirmed healthy; privatebin restarted and healthy; docmost (all 3) healthy. Planted data untouched throughout — no app was rebuilt, redeployed or restored.
  2. Bake VM — the drill VM on DooPlex ran the bake and its build guest 9100 is stopped inside it; the qcow2 reverts to the virgin snapshot. No storage was added anywhere, so pvesm status has nothing to compare.
  3. Hub-side record — stated explicitly even though there is none. No appliance was registered, no customer created, no config written, no artifact manifest changed. The hub was READ ONLY (GET /configuration, GET /events). Nothing to discard.

End state: demo-hp guest 9201 runs controller 0.222.0, all 8 deployed apps healthy, golden 0.222.0 baked and published but NOT vouched — floor still 0.221.1.

13. Register size

File Before After
OPEN-ITEMS.md 327,266 B 328,325 B
CLOSED-ITEMS.md 68,464 B 71,441 B

R-383 and R-384 moved to CLOSED compressed; R-385 (closed) and R-386 (open) filed. OPEN grew by ~1 KB despite two closures because R-386 is a substantial new finding — recorded rather than smoothed over.

14. Observations — noticed, documented, NOT acted on

  1. R-329 is live and now matters much more. app_start_failed is pushed with severity warn, which is not in the hub's vocabulary ({info, warning, error, critical}) and coerces silently to info, e-mailing nobody, while the POST still returns 200. Observed again today: PushEvent: type=app_start_failed severity=warn. R-384 makes this event actually fire, so a known-broken severity moved from unreachable to load-bearing. Not in scope; not touched.
  2. Two files carry pre-existing gofmt drift — internal/backup/offbox.go and internal/backup/offbox_recovery_cli.go. Confirmed pre-existing by stashing this session's work and re-running gofmt -l. Not touched (§12 forbids nearby refactors).
  3. The runbook's golden-bake step is missing pveam update. On the virgin snapshot the template index is stale, so pveam available offers 13.1-2 and downloading it fails with 400 Parameter verification failed. template: no such template. Recorded in the bake evidence README; the runbook itself was not edited.
  4. The register's own suggested fix for R-384 was wrong — it proposed a sustained-unhealthy threshold on the crashLoopAfter model. The defect needed no threshold at all, only an ordering. Recorded in the CLOSED entry so the next reader sees that a register remedy is a hypothesis.
  5. Deliberately left open, untouched: R-102, R-359, R-361's sibling surfaces.