diff --git a/CONTEXT.md b/CONTEXT.md index 4e4ae76..13db583 100644 --- a/CONTEXT.md +++ b/CONTEXT.md @@ -7,7 +7,61 @@ > > Ask Claude Code: "Please update CONTEXT.md with what we did today" -Last updated: 2026-08-23 (v0.221.1 — R-361: taking the undo copy destroyed the app's own backup) +Last updated: 2026-08-23 (v0.222.0 — R-384: a dead database raised no alarm; R-383; and R-386 filed) + +> **2026-08-23 — v0.222.0 (R-384 + R-383), and a bigger hole found by a measurement that was told not to fix it.** +> +> **[DECISION] A dead SUPERVISED member is asked about BEFORE a failing healthcheck, because they are +> different questions and the second was answering the first.** `aggregateState` returned +> `StateUnhealthy` the moment `unhealthy > 0`, and the R-51 mixed-case block that asks "is a supervised +> member dead?" sat below it. A two-container app whose database exits goes `unhealthy` seconds later +> *because it cannot reach that database* — so **the symptom the fault causes was what suppressed the +> alarm for it.** The supervised test is now hoisted above the unhealthy/starting/restarting returns. +> +> **[DECISION] "Some members are up" means ANY member not in the down bucket** — running, unhealthy, +> starting or restarting. The old guard was `running > 0` counting `StateRunning` alone, which made the +> R-51 block **unreachable in precisely the case it was written for**: an unhealthy survivor beside a +> dead database counted as nothing being up. Either half alone leaves the defect standing, and the two +> red-proofs convict independently. +> +> **[FENCE, unchanged] `IsDownState` is byte-identical and `unhealthy` stays excluded.** An unhealthy +> container is RUNNING; folding it in reintroduces the flapping that exclusion exists to stop. **No new +> state was minted** — `StateDegraded` already meant this and every consumer already handled it. The +> fix is an ORDER, not a widening. +> +> **[FACT] The register's own suggested remedy was wrong.** R-384's row proposed a sustained-`unhealthy` +> threshold on the `crashLoopAfter` model. The defect needed no threshold at all. **A register's +> "recommended fix" is a hypothesis written before the diagnosis, and must be re-derived from source.** +> +> **[TRAP] Three existing subtests pinned the DEFECT as settled behaviour.** +> `TestAggregateState_UnchangedBranches` asserted an unhealthy/starting/restarting member beat an +> `exited` peer that was on `unless-stopped`. They were amended (down member given a benign policy, +> which is the only case where that sentence was ever true) and the change is reported, not buried. +> **A green suite can be green about the wrong thing.** +> +> **[FINDING — R-386, OPEN, NOT FIXED] A single-container app stopped out of band raises NO alarm, and +> a comment states the opposite.** `aggregateState` folds `StateExited` into the `stopped` counter, so +> an all-down stack returns `StateStopped` and **`StateExited` never survives aggregation**; +> `classifyRunStates` then whitelists `stopped` as a deliberate user stop. The comment at +> `cmd/controller/main.go` claiming an out-of-band `docker compose stop` "still alerts" is **measured +> false** — `privatebin`, 9 scans, 0 events, 0 banner. **Case #10 of "a comment asserting an invariant +> the code does not provide".** The task asked for this as a MEASUREMENT and forbade a fix; it is filed. +> +> **[DECISION, R-383] A message may not assert a file exists without asking the disk.** The +> double-failure sentence named the undo copy as present, built from the returned path — and a missing +> file is one of the two ways that rollback fails. `undoCopyPhrase` now reads from disk; a zero-length +> dump counts as MISSING; and the absent case still names WHERE the file should have been, because +> R-351's lesson is that a refusal naming nothing forces someone to remember what the product knows. +> +> **[DECISION] The alarm ladder now has an owning document** — +> `felhom.eu/documentation/architecture/08-alarm-ladder.md`. Until 2026-08-23 no document owned it; the +> rules lived as comments in four packages, each locally correct, with the ordering between them legible +> only by reading one function top to bottom. **That absence is why R-384 survived review.** +> +> **[GOTCHA] `app_start_failed` still ships severity `warn` (R-329)**, which is not in the hub's +> vocabulary and coerces silently to `info`, e-mailing nobody, while the POST returns 200. R-384 moved +> this event from unreachable to load-bearing, so the severity bug now matters. + > **2026-08-23 — v0.221.0/.1 (R-361), and two negatives worth as much as the fix.** > diff --git a/REPORT.md b/REPORT.md index c1d840b..5a6893d 100644 --- a/REPORT.md +++ b/REPORT.md @@ -1,90 +1,267 @@ -# REPORT — R-361: taking the undo copy destroyed the app's own database backup +# REPORT — controller v0.221.1 (record) + v0.222.0 (R-384, R-383) -**Controller v0.221.0 → v0.221.1**, proven live on `demo-hp`. -Full record: `felhom.eu/documentation/audits/DRILL-r361-2026-08-22/`. +**Session 2026-08-23, UNATTENDED. Live leg on `demo-hp` (Tier 0, disposable).** -## Baselines and subjects +> ## ⚠ HALT DECLARED — §4's measurement found a real defect, filed as R-386, NOT fixed +> +> The task's §4 asked for a measurement and named it a halt condition. It reproduced. +> **A single-container app stopped out of band raises no alarm at all.** Details in §10 below. +> Per §13 the fix was NOT attempted here. Live-walk step 4 (Scenario E) was dropped as a +> consequence — it is item (3) on the task's own drop list. Everything else completed. -| repo | HEAD at start | version | +--- + +## 1. Baselines used, and the hub's four numbers as read + +| Repo | `main` at start | Verified | |---|---|---| -| felhom-controller | `2024ed99826717bf765803e159c074909b28bf83` | v0.220.2 → **v0.221.1** | -| felhom-agent | `40d857b52711dd8c9c88bdd21ffcaad819a33a84` | v0.130.0, unchanged | -| felhom.eu | `a8caa0fdde7c678bb88b5b25727044695bbdeaa9` | unchanged | -| app-catalog | `459766cb16395fd1d1a66282f5cc6da59ead5924` | unchanged | +| felhom-controller | `f7881787f434` | matches the task | +| felhom.eu | `1eb64bec5183` | task said `4e488321bfd1+`; it had moved on | +| felhom-agent | untouched | — | -Hub read: golden **0.220.2**, agent **0.130.0**, min_agent **0.129.0**, floor **0.220.2** — as stated. -**Subjects FOUND, not rebuilt:** `docmost` and `bookstack`, healthy, with two sessions of data. +**Hub's own numbers, read live from `/configuration` (ClusterIP + Basic auth) 2026-08-23:** -**Architecture read:** `07-backup-architecture.md` §6.1–§6.3 including the v0.220.x failure-ladder -`[DESIGN]`; and `audits/DRILL-r379-rollback-2026-08-22/README.md`. +| Field | Value | +|---|---| +| `golden_version` | **0.221.1** | +| `agent_version` | **0.130.0** | +| `min_agent` | **0.129.0** | +| controller floor (`min_controller_version`) | **0.221.1** | -## R-361 — the whole thing is one comparison +The task expected golden/floor **0.220.2**; the operator had already vouched **0.221.1** and raised +the floor. Live controller on `demo-hp` at session start: **0.221.1** — so the running version, the +golden and the floor all agreed, and only the RECORD disagreed. That is exactly R-385's shape. -| app | canonical dump sha256, before | after a restore | +## 2. Architecture documents read + +- `documentation/architecture/00-capability-map.md` — its 2026-08-22 paragraph already NAMED R-384 as + an open finding, from the held-app measurement. +- `felhom-controller/internal/stacks/manager.go` `IsDownState` + `aggregateState` + `supervisedPolicy` +- `cmd/controller/main.go` `classifyRunStates` and its three suppressions +- `internal/quiesce/suppress.go`, `internal/bootrecon/bootrecon.go`, `internal/stacks/desiredstate.go` +- `felhom.eu/scripts/golden_currency_gate.py` (all 133 lines) + +**§2's conditional applies and is answered: NO document owned the alarm ladder.** That absence is +reported as a finding, and `documentation/architecture/08-alarm-ladder.md` now owns it (Part N.4). +It is why the ordering defect was legible only by reading one function top to bottom. + +## 3. Part 0 — the record, pushed ALONE + +Exact heading written: + +``` +## v0.221.1 — the undo-copy prune stopped running because another fix made its guard reachable (2026-08-23, R-361 follow-on) +``` + +Commit **`da75603`**, pushed alone before anything else. The reasoning was **moved verbatim** from the +v0.221.0 entry (which no longer claims it), not rewritten and not duplicated. + +## 4. Files changed, commits, CI runs + +| Commit | Contents | +|---|---| +| **`da75603`** | Part 0 — the `v0.221.1` heading, alone | +| **`5da11c4`** | v0.222.0 — R-384 + R-383, tests, CHANGELOG | + +Modified: `CHANGELOG.md`, `internal/stacks/manager.go`, `internal/stacks/degraded_test.go`, +`internal/backup/offbox_reconstitute.go`, `cmd/controller/r361_classifier_control_test.go`. +Added: `cmd/controller/r384_dead_db_alarm_test.go`, `internal/backup/r383_undo_phrase_test.go`. + +**CI runs confirmed BY ID** (`id` and `run_number` diverge, both printed): + +| Commit | CI `id` | `run_number` | Result | +|---|---|---|---| +| `da75603` | **404** | 85 | success | +| `5da11c4` | **405** | 86 | success | + +## 5. Red-proofs — four planted, FOUR SEEN FAILING + +| # | Mutation | Layer the guard sits at | Observed failure | +|---|---|---|---| +| 1 | hoisted block moved back **below** `unhealthy > 0` | `aggregateState` — the ORDERING | `aggregateState = "unhealthy", want "degraded"` **and** `bookstack state = "unhealthy", want "degraded"` (production-path wiring) | +| 2 | `up` narrowed back to `running` alone | `aggregateState` — the GUARD | same subtest, plus `"starting"` and `"restarting"` — all three survivor shapes convict | +| 3 | classifier drops `StateDegraded` from `down` | `classifyRunStates` — the CONSEQUENCE | `dead-app banner = [], want exactly one entry for bookstack` | +| 4 | `undoCopyPhrase` reverted to the unconditional claim | the phrase builder — where the CLAIM is made | `phrase "…mentése megvan: …mariadb.sql" contains "mentése megvan" — it asserts a file that is not on disk`; the empty set printed `megvan: .`, naming a file that never existed | + +**Every mutation was asserted to have applied** (the scripts `assert` the pre-fix text is present +before rewriting and print `MUTATION APPLIED`). **None passed first time.** Mutations 1 and 2 convict +independently, which is what proves the fix genuinely has two halves. + +## 6. Test count + +**1494 → 1504** top-level test functions (measured by `go test ./... -list '.*'` on the stashed and +unstashed tree, not estimated). Full green gate `go build && go vet && go test ./...` → **exit 0, zero +failures**, run after Part 2 and again after Part 3. + +## 7. Deployed version, and the golden + +``` +gitea.dooplex.hu/admin/felhom-controller:0.222.0 Up 20 seconds (healthy) +``` + +**Golden BAKED and PUBLISHED: YES — version 0.222.0.** +`GOLDEN_SHA256 = 19f5904f53792684f046ec0bc25426645cb87ad73d5cfc6c03639d9f82706037`, `upload OK (HTTP 201)`, +round-trip `HTTP 206` from the package URL. All acceptance markers counted and recorded. +**VOUCHING IS THE OPERATOR'S ACT AND WAS NOT DONE HERE.** + +## 8. The five `IsDownState` consumers, walked and named + +| Consumer | What changes | +|---|---| +| `cmd/controller/main.go:2173` `classifyRunStates` | **THE INTENDED CHANGE.** A stack that read `unhealthy` now reads `degraded` → `down=true` → banner + `app_start_failed`. `userStopped` tests `StateStopped` specifically, so the whitelist cannot swallow `degraded`. | +| `internal/bootrecon/bootrecon.go:213` (DesiredStateRunning) | **CHANGES, and toward repair.** A half-started stack at boot now reads `degraded` → an orphan → `compose up -d`. Previously it read `unhealthy` → not an orphan → left half-dead. Aligned with the file's own stated intent. | +| `internal/bootrecon/bootrecon.go:216` (legacy DesiredStateUnknown) | Same shape, same direction. | +| `internal/bootrecon/bootrecon.go:288` (recovery check) | **CHANGES, and toward truth.** A stack that came back with a dead supervised member is no longer counted `Recovered`; it stays pending and is retried, bounded by `r.attempts`. It used to be declared recovered while half-dead. | +| `internal/stacks/desiredstate.go:150` `isObservedUp` | **UNAFFECTED — verified, not assumed.** It is an allow-list of `{running, starting}`; neither `unhealthy` nor `degraded` was ever in it, so a stack moving between them does not cross the boundary. | +| `internal/quiesce/suppress.go` | **UNAFFECTED.** It does not call `IsDownState` at all — the suppression is cycle-keyed and state-blind, which is precisely why R-97b's guarantee cannot be weakened by a state change. Pinned by `TestR384_QuiesceSuppressionStillHoldsForDegraded`. | +| dashboard state badge | **Already handled.** R-51 wired `degraded` through `handlers.go:158/169` (counts with stopped) and `funcmap.go:255` (filters with stopped). Verified live — the badge rendered `(degraded)`. | + +## 9. The live walk + +Method: endpoint-level. No browser exists on DooPlex; every read below is either the exact endpoint +the UI calls (`POST /api/stacks//`, `GET /api/stacks`, `GET /dashboard`) or the +controller's own log. Guest clock is UTC. + +### Step 1 — Scenario A: a database dies behind a healthy-looking app ✅ + +| Observable | Result | +|---|---| +| `bookstack-db` stopped out of band | 05:30:07Z | +| front end went `unhealthy` | 05:31:21Z — **the state that used to swallow the alarm** | +| aggregate state read | **`degraded`** while the front end was `unhealthy` (05:32:03Z) | +| `app_start_failed` | **fired at 05:30:14Z**, 7 s after the stop | +| new code path visible | `manager.go:703: restart-policy of down member "bookstack-db" = "unless-stopped"` | +| banner | *„Telepített alkalmazás nem fut: BookStack (degraded)"* on **both** `/launcher` and `/dashboard` | +| edge-triggered, not per-scan | **1 event across 22 scans** | + +**The heartbeat, old beside new — same 8 apps evaluated:** + +``` +2026/08/22 21:13:48 [deadapp] check alive: 580 scans since boot, 8 deployed app(s) evaluated, 0 currently down ← v0.220.2/0.221.1 +2026/08/23 05:37:44 [deadapp] check alive: 20 scans since boot, 8 deployed app(s) evaluated, 1 currently down ← v0.222.0 +``` + +### Step 2 — Scenario B: unhealthy with nothing dead ✅ + +The database was restarted; the front end stayed `unhealthy` with nothing down. Aggregate read +**`unhealthy`**, not `degraded`. **0 new `app_start_failed`**, and — the positive observable — +**no `restart-policy of down member` line at all**, meaning the supervised path was not entered. That +absence is trustworthy because the same line HAD appeared on this box 10 minutes earlier. +Banner cleared; `bookstack state=running`. + +*Honest limit:* the live window in which docker reported the front end unhealthy with the database up +was ~11 s wide, and the controller's cached read was taken at its edge. The three-shape unit test +`TestR384_UnhealthyWithNothingDeadDoesNotAlarm` carries the rest of this case. + +### Step 3 — Scenario D: a full deploy cycle ✅ **0 alarms** + +`POST /api/stacks/docmost/stop` then `/start` — the exact calls the launcher's buttons make — on a +3-container stack, watched for 5 minutes to settled healthy. + +| Observable | Result | +|---|---| +| `app_start_failed` across the cycle | **0** | +| dead-app scans in the window | **9** | +| supervised-down path entered | 1, at 05:42:34Z (docmost itself momentarily down beside two live members) | + +**Alarms that v0.221.1 would NOT have produced: ZERO.** The single `degraded` reading at 05:42:34Z +has `running > 0`, so v0.221.1's mixed-case branch reaches the identical verdict. No moment in the +cycle had a down supervised member with only non-`running` survivors, which is the only shape where +the two versions differ. + +### Step 4 — Scenario E: the quiesce cycle — **DROPPED, and why** + +Dropped as a direct consequence of the §4 halt (see the banner at the top), and it is item **(3)** on +the task's own drop list. Running it would have meant triggering a real backup cycle on the box +*after* a halt condition had already fired. **Covered at unit level instead** by +`TestR384_QuiesceSuppressionStillHoldsForDegraded`, which asserts a `degraded` stack inside the +quiesce set produces no banner and `Down=false`. **Not proven live in this session — stated plainly +rather than implied.** + +### Step 5 — §4's measurement ✅ (it reproduced — see §10) + +### Step 6 — Part 1's gate, both directions ✅ + +| Run | Gate | CHANGELOG | Golden | Exit | +|---|---|---|---|---| +| `gate-01` | **old** | v0.221.0 | 0.221.1 | **0** — the blindness, on the real history | +| `gate-02` | **new** | v0.221.0 | 0.221.1 | **1** — convicted, naming the missing heading and the route | +| `gate-03` | new | v0.221.1 | 0.221.1 | 0 | +| `gate-04` | new | absent clone | — | **2** — INCONCLUSIVE preserved | +| `gate-05` | new | v0.222.0 | 0.222.0 | 0 — post-bake | + +## 10. §4's answer, in plain words + +**A single-container app that is stopped out of band raises no alarm at all, and a comment in the +code says the opposite.** + +`aggregateState` folds `StateExited` into the `stopped` counter, so when every member is down it +returns `StateStopped` — **`StateExited` never survives aggregation**, which is the path the task +suspected and could not find. `classifyRunStates` then whitelists `StateStopped` as a deliberate user +stop. So the comment at `cmd/controller/main.go` — *"An out-of-band `docker compose stop` leaves the +containers present → StateExited → still alerts, which is correct: out-of-band tampering IS +reportable"* — is **false**, and so is the neighbouring I2 claim that a crashing app never comes to +rest at `stopped`. + +**Measured:** `privatebin` (1 container, `unless-stopped`) stopped 05:47:35Z. At 05:51:53Z: +`state=stopped`, **9 dead-app scans had run, 0 events, 0 banner lines.** +**Positive control first, per standing rule 3:** `app_start_failed` fired for BookStack at 05:30:14Z +on the same box 17 minutes earlier, so the detector was demonstrably alive. + +**Scoped honestly:** a genuine *crash* under `unless-stopped` is restarted by Docker and surfaces as +`restarting` → the 5-minute crash-loop path, which does alarm. The silent case is an explicit +out-of-band stop of a stack with no surviving member. + +**Filed as R-386 (OPEN — MEDIUM). Not fixed here, per §12 and §13.** + +## 11. Evidence + +`felhom.eu/documentation/audits/DRILL-r384-dead-db-alarm-2026-08-23/evidence/` — 5 gate runs, 4 +red-proof transcripts, 20 live-walk files including two full controller-log windows (1808 and 2139 +lines) pulled off the guest. **Both log windows were copied off before the app was restarted**, per +standing rule 5. + +## 12. Teardown, three layers, and the box's end state + +1. **Guest 9201 / apps** — nothing provisioned. `bookstack-db` restarted and **`bookstack` confirmed + healthy**; `privatebin` restarted and healthy; `docmost` (all 3) healthy. Planted data untouched + throughout — no app was rebuilt, redeployed or restored. +2. **Bake VM** — the drill VM on DooPlex ran the bake and its build guest 9100 is stopped inside it; + the qcow2 reverts to the `virgin` snapshot. No storage was added anywhere, so `pvesm status` has + nothing to compare. +3. **Hub-side record — stated explicitly even though there is none.** No appliance was registered, no + customer created, no config written, no artifact manifest changed. **The hub was READ ONLY** + (`GET /configuration`, `GET /events`). Nothing to discard. + +**End state:** `demo-hp` guest 9201 runs controller **0.222.0**, all 8 deployed apps healthy, golden +0.222.0 baked and published but **NOT vouched** — floor still **0.221.1**. + +## 13. Register size + +| File | Before | After | |---|---|---| -| `docmost` | `5d35678349bbbdb318ac22656d4f96436b5751ef5b3b1659d48305a526df1aed` | **identical** | -| `bookstack` | `7837aa5de2955dcf3125534f015f43df42debe13c765793d3b75173abcc55e7b` | **identical** | +| `OPEN-ITEMS.md` | 327,266 B | **328,325 B** | +| `CLOSED-ITEMS.md` | 68,464 B | **71,441 B** | -**Before the fix, on the same box: neither app had a canonical dump at all.** +R-383 and R-384 moved to CLOSED compressed; R-385 (closed) and R-386 (open) filed. OPEN grew by +~1 KB despite two closures because R-386 is a substantial new finding — recorded rather than +smoothed over. -`DumpOneTo` takes an explicit final path and derives its `.tmp` from it. `DumpOne`'s signature did not -move. The comment that asserted the old behaviour was safe now states the invariant and how it is -enforced. +## 14. Observations — noticed, documented, NOT acted on -## Manifest — every consumer named, as required - -`Manifest.DBDumps` has **three** references, all in `internal/backup/recovery_unit.go`: the -declaration (`:55`), the enumeration, and the already-current compare. **None reads it for recovery; -no hub or agent consumer exists.** Undo copies are therefore excluded from `db_dumps`. Verified live: -`db_dumps: ['docmost-postgres.sql']` with **4 undo copies on disk** as the positive control. - -## Part 3 — the measurement that cancelled Part 2 - -**The reading did not reproduce.** A held app aggregates to `unhealthy`; `IsDownState` is -`{stopped, exited, degraded}`; the dead-app heartbeat read `0 currently down` across scans 600 and -620 while a hold was in force. **Part 2 dropped in full**, no register row opened, negative recorded -in the capability map. - -**The positive control took three attempts and that is the finding.** Two live attempts failed — -`privatebin` went `stopped` (whitelisted by design) and `bookstack` went `degraded` then `unhealthy`. -The control that works is at the detector's own layer: `classifyRunStates` is pure and raises the -banner for `degraded`/`exited` while staying silent for the states measured live. - -## Findings filed - -- **R-383 (MEDIUM)** — the double-failure message says the previous state's backup **exists** while - naming the very file whose absence caused the failure. Observed twice, on v0.220.2 and v0.221.1. - **R-361's own class**, one surface over. -- **R-384 (MEDIUM)** — an app whose **database** has died reads `unhealthy` and raises no dead-app - alarm. `bookstack-db` stopped at 21:27:01; `0 currently down` throughout. - -**Register: `OPEN-ITEMS.md` 325 236 → 327 266 bytes; `CLOSED-ITEMS.md` 66 777 → 68 464.** R-361 -compressed; full text `git show a8caa0fdde7c:documentation/backlog/OPEN-ITEMS.md`. - -## Two defects in this session's own work - -1. **Two red-proofs passed first time**, both reported. The behavioural tests inject the dump seam, so - a mutation *inside* `DumpOneTo` was invisible; and Part 1.3 had no test at all. Guards added at the - layer each defect lives in; both mutations then convicted. -2. **One change made another unreachable.** A stable `db_dumps` let the already-current early return - fire, and the undo prune sat after it — **four copies against a cap of three, counted on the box**. - Fixed in v0.221.1; the cap now holds at 3 on both apps, verified live. - -**Test count 1485 → 1493.** Green gate clean; controller gates all OK; **no push used `--no-verify`.** - -## Part 4 — the double-failure path on what ships - -Trigger: **the undo copy is lost between being written and being needed** — realistic, and one of only -two ways a rollback can fail. Verified on 0.221.1: held, refused by the customer button, refused after -a full controller restart, then listed/cleared/started via the operator route. Message captured -verbatim (369 bytes, hex in the drill record) with **no engine output**. - -## Teardown - -Nothing provisioned; 20 containers healthy; `docmost` still holds its 4 pages and 1 user; both -canonical dumps present; **no app left held**; no hub-side record created. - -## Operator follow-up - -Vouch a golden carrying **0.221.1** (baked in this session), then raise the floor to 0.221.1 last, in -its own save. +1. **R-329 is live and now matters much more.** `app_start_failed` is pushed with severity **`warn`**, + which is not in the hub's vocabulary (`{info, warning, error, critical}`) and coerces silently to + `info`, e-mailing nobody, while the POST still returns 200. Observed again today: + `PushEvent: type=app_start_failed severity=warn`. **R-384 makes this event actually fire, so a + known-broken severity moved from unreachable to load-bearing.** Not in scope; not touched. +2. **Two files carry pre-existing `gofmt` drift** — `internal/backup/offbox.go` and + `internal/backup/offbox_recovery_cli.go`. Confirmed pre-existing by stashing this session's work + and re-running `gofmt -l`. Not touched (§12 forbids nearby refactors). +3. **The runbook's golden-bake step is missing `pveam update`.** On the `virgin` snapshot the template + index is stale, so `pveam available` offers `13.1-2` and downloading it fails with + `400 Parameter verification failed. template: no such template`. Recorded in the bake evidence + README; the runbook itself was not edited. +4. **The register's own suggested fix for R-384 was wrong** — it proposed a sustained-`unhealthy` + threshold on the `crashLoopAfter` model. The defect needed no threshold at all, only an ordering. + Recorded in the CLOSED entry so the next reader sees that a register remedy is a hypothesis. +5. **Deliberately left open, untouched:** R-102, R-359, R-361's sibling surfaces. diff --git a/controller/README.md b/controller/README.md index 9324b9c..2547ba5 100644 --- a/controller/README.md +++ b/controller/README.md @@ -702,10 +702,19 @@ When app templates are updated (e.g., a new `APP_KEY` secret is added to `.felho | Running + starting | Orange | "Indulas..." | Healthcheck not yet passed | | Deploying | Orange | "Telepítés..." | Compose up in progress (image pull, container creation) | | Running + unhealthy | Yellow | "Nem egeszseges" | Docker or controller-side healthcheck failing | +| **Degraded** | Red | "Leallitva" (counts with stopped) | **A SUPERVISED member of a multi-container app is dead** — e.g. the app's database — while other members are still up. Counts as DOWN: raises the dead-app banner and `app_start_failed` (R-51, R-384) | | Stopped/exited | Red | "Leallitva" | All containers stopped | -| Restarting | Yellow | "Ujrainditas..." | Restart loop | +| Restarting | Yellow | "Ujrainditas..." | Restart loop; becomes down only after 5 minutes sustained (crash loop) | | Not deployed | Gray | "Nincs telepitve" | Compose file exists, not deployed | +**The order these are decided in is load-bearing (v0.222.0, R-384).** `degraded` is evaluated +**before** `unhealthy`/`starting`/`restarting`. An app whose database dies drags its own front end +`unhealthy` seconds later — so if `unhealthy` were decided first (as it was until v0.222.0), the +symptom would mask the fault and the app would be silently down. A down member whose restart policy +is `no`/`on-failure` is a finished one-shot init/migrate container and stays benign. The full ladder, +including which states deliberately do NOT alarm and why, is documented in +`felhom.eu/documentation/architecture/08-alarm-ladder.md`. + **Route-unpublished indicator (F5, v0.61.0).** Traefik's Docker provider only publishes a route to a container that is healthy (or has no healthcheck), so an `unhealthy`/`restarting` deployed app returns a hard **404** at its URL even though the container is running. The `routeUnpublished` template helper