docs(v0.222.0): REPORT, CONTEXT decisions, README state table + the ordering
gates / gates (push) Successful in 11s

REPORT.md overwritten with the full run: baselines and the hub's four numbers,
the four red-proofs with the mutation and observed text for each, the five
IsDownState consumers walked and named, the live walk in full with the old and
new heartbeat lines quoted side by side, and the halt.

CONTEXT records the decision - a dead supervised member is asked about before a
failing healthcheck, because they are different questions and the second was
answering the first - plus the fence that IsDownState did not move, the trap
that three existing subtests pinned the defect, and R-386.

README gains the `degraded` row, which the state table never had, and a note
that the ORDER is load-bearing. Points at the new alarm-ladder architecture doc.
This commit is contained in:
2026-08-23 07:58:09 +02:00
parent 5da11c4480
commit 14137efac5
3 changed files with 320 additions and 80 deletions
+55 -1
View File
@@ -7,7 +7,61 @@
> >
> Ask Claude Code: "Please update CONTEXT.md with what we did today" > Ask Claude Code: "Please update CONTEXT.md with what we did today"
Last updated: 2026-08-23 (v0.221.1 — R-361: taking the undo copy destroyed the app's own backup) Last updated: 2026-08-23 (v0.222.0 — R-384: a dead database raised no alarm; R-383; and R-386 filed)
> **2026-08-23 — v0.222.0 (R-384 + R-383), and a bigger hole found by a measurement that was told not to fix it.**
>
> **[DECISION] A dead SUPERVISED member is asked about BEFORE a failing healthcheck, because they are
> different questions and the second was answering the first.** `aggregateState` returned
> `StateUnhealthy` the moment `unhealthy > 0`, and the R-51 mixed-case block that asks "is a supervised
> member dead?" sat below it. A two-container app whose database exits goes `unhealthy` seconds later
> *because it cannot reach that database* — so **the symptom the fault causes was what suppressed the
> alarm for it.** The supervised test is now hoisted above the unhealthy/starting/restarting returns.
>
> **[DECISION] "Some members are up" means ANY member not in the down bucket** — running, unhealthy,
> starting or restarting. The old guard was `running > 0` counting `StateRunning` alone, which made the
> R-51 block **unreachable in precisely the case it was written for**: an unhealthy survivor beside a
> dead database counted as nothing being up. Either half alone leaves the defect standing, and the two
> red-proofs convict independently.
>
> **[FENCE, unchanged] `IsDownState` is byte-identical and `unhealthy` stays excluded.** An unhealthy
> container is RUNNING; folding it in reintroduces the flapping that exclusion exists to stop. **No new
> state was minted** — `StateDegraded` already meant this and every consumer already handled it. The
> fix is an ORDER, not a widening.
>
> **[FACT] The register's own suggested remedy was wrong.** R-384's row proposed a sustained-`unhealthy`
> threshold on the `crashLoopAfter` model. The defect needed no threshold at all. **A register's
> "recommended fix" is a hypothesis written before the diagnosis, and must be re-derived from source.**
>
> **[TRAP] Three existing subtests pinned the DEFECT as settled behaviour.**
> `TestAggregateState_UnchangedBranches` asserted an unhealthy/starting/restarting member beat an
> `exited` peer that was on `unless-stopped`. They were amended (down member given a benign policy,
> which is the only case where that sentence was ever true) and the change is reported, not buried.
> **A green suite can be green about the wrong thing.**
>
> **[FINDING — R-386, OPEN, NOT FIXED] A single-container app stopped out of band raises NO alarm, and
> a comment states the opposite.** `aggregateState` folds `StateExited` into the `stopped` counter, so
> an all-down stack returns `StateStopped` and **`StateExited` never survives aggregation**;
> `classifyRunStates` then whitelists `stopped` as a deliberate user stop. The comment at
> `cmd/controller/main.go` claiming an out-of-band `docker compose stop` "still alerts" is **measured
> false** — `privatebin`, 9 scans, 0 events, 0 banner. **Case #10 of "a comment asserting an invariant
> the code does not provide".** The task asked for this as a MEASUREMENT and forbade a fix; it is filed.
>
> **[DECISION, R-383] A message may not assert a file exists without asking the disk.** The
> double-failure sentence named the undo copy as present, built from the returned path — and a missing
> file is one of the two ways that rollback fails. `undoCopyPhrase` now reads from disk; a zero-length
> dump counts as MISSING; and the absent case still names WHERE the file should have been, because
> R-351's lesson is that a refusal naming nothing forces someone to remember what the product knows.
>
> **[DECISION] The alarm ladder now has an owning document** —
> `felhom.eu/documentation/architecture/08-alarm-ladder.md`. Until 2026-08-23 no document owned it; the
> rules lived as comments in four packages, each locally correct, with the ordering between them legible
> only by reading one function top to bottom. **That absence is why R-384 survived review.**
>
> **[GOTCHA] `app_start_failed` still ships severity `warn` (R-329)**, which is not in the hub's
> vocabulary and coerces silently to `info`, e-mailing nobody, while the POST returns 200. R-384 moved
> this event from unreachable to load-bearing, so the severity bug now matters.
> **2026-08-23 — v0.221.0/.1 (R-361), and two negatives worth as much as the fix.** > **2026-08-23 — v0.221.0/.1 (R-361), and two negatives worth as much as the fix.**
> >
+255 -78
View File
@@ -1,90 +1,267 @@
# REPORT — R-361: taking the undo copy destroyed the app's own database backup # REPORT — controller v0.221.1 (record) + v0.222.0 (R-384, R-383)
**Controller v0.221.0 → v0.221.1**, proven live on `demo-hp`. **Session 2026-08-23, UNATTENDED. Live leg on `demo-hp` (Tier 0, disposable).**
Full record: `felhom.eu/documentation/audits/DRILL-r361-2026-08-22/`.
## Baselines and subjects > ## ⚠ HALT DECLARED — §4's measurement found a real defect, filed as R-386, NOT fixed
>
> The task's §4 asked for a measurement and named it a halt condition. It reproduced.
> **A single-container app stopped out of band raises no alarm at all.** Details in §10 below.
> Per §13 the fix was NOT attempted here. Live-walk step 4 (Scenario E) was dropped as a
> consequence — it is item (3) on the task's own drop list. Everything else completed.
| repo | HEAD at start | version | ---
## 1. Baselines used, and the hub's four numbers as read
| Repo | `main` at start | Verified |
|---|---|---| |---|---|---|
| felhom-controller | `2024ed99826717bf765803e159c074909b28bf83` | v0.220.2 → **v0.221.1** | | felhom-controller | `f7881787f434` | matches the task |
| felhom-agent | `40d857b52711dd8c9c88bdd21ffcaad819a33a84` | v0.130.0, unchanged | | felhom.eu | `1eb64bec5183` | task said `4e488321bfd1+`; it had moved on |
| felhom.eu | `a8caa0fdde7c678bb88b5b25727044695bbdeaa9` | unchanged | | felhom-agent | untouched | — |
| app-catalog | `459766cb16395fd1d1a66282f5cc6da59ead5924` | unchanged |
Hub read: golden **0.220.2**, agent **0.130.0**, min_agent **0.129.0**, floor **0.220.2** — as stated. **Hub's own numbers, read live from `/configuration` (ClusterIP + Basic auth) 2026-08-23:**
**Subjects FOUND, not rebuilt:** `docmost` and `bookstack`, healthy, with two sessions of data.
**Architecture read:** `07-backup-architecture.md` §6.1–§6.3 including the v0.220.x failure-ladder | Field | Value |
`[DESIGN]`; and `audits/DRILL-r379-rollback-2026-08-22/README.md`. |---|---|
| `golden_version` | **0.221.1** |
| `agent_version` | **0.130.0** |
| `min_agent` | **0.129.0** |
| controller floor (`min_controller_version`) | **0.221.1** |
## R-361 — the whole thing is one comparison The task expected golden/floor **0.220.2**; the operator had already vouched **0.221.1** and raised
the floor. Live controller on `demo-hp` at session start: **0.221.1** — so the running version, the
golden and the floor all agreed, and only the RECORD disagreed. That is exactly R-385's shape.
| app | canonical dump sha256, before | after a restore | ## 2. Architecture documents read
- `documentation/architecture/00-capability-map.md` — its 2026-08-22 paragraph already NAMED R-384 as
an open finding, from the held-app measurement.
- `felhom-controller/internal/stacks/manager.go` `IsDownState` + `aggregateState` + `supervisedPolicy`
- `cmd/controller/main.go` `classifyRunStates` and its three suppressions
- `internal/quiesce/suppress.go`, `internal/bootrecon/bootrecon.go`, `internal/stacks/desiredstate.go`
- `felhom.eu/scripts/golden_currency_gate.py` (all 133 lines)
**§2's conditional applies and is answered: NO document owned the alarm ladder.** That absence is
reported as a finding, and `documentation/architecture/08-alarm-ladder.md` now owns it (Part N.4).
It is why the ordering defect was legible only by reading one function top to bottom.
## 3. Part 0 — the record, pushed ALONE
Exact heading written:
```
## v0.221.1 — the undo-copy prune stopped running because another fix made its guard reachable (2026-08-23, R-361 follow-on)
```
Commit **`da75603`**, pushed alone before anything else. The reasoning was **moved verbatim** from the
v0.221.0 entry (which no longer claims it), not rewritten and not duplicated.
## 4. Files changed, commits, CI runs
| Commit | Contents |
|---|---|
| **`da75603`** | Part 0 — the `v0.221.1` heading, alone |
| **`5da11c4`** | v0.222.0 — R-384 + R-383, tests, CHANGELOG |
Modified: `CHANGELOG.md`, `internal/stacks/manager.go`, `internal/stacks/degraded_test.go`,
`internal/backup/offbox_reconstitute.go`, `cmd/controller/r361_classifier_control_test.go`.
Added: `cmd/controller/r384_dead_db_alarm_test.go`, `internal/backup/r383_undo_phrase_test.go`.
**CI runs confirmed BY ID** (`id` and `run_number` diverge, both printed):
| Commit | CI `id` | `run_number` | Result |
|---|---|---|---|
| `da75603` | **404** | 85 | success |
| `5da11c4` | **405** | 86 | success |
## 5. Red-proofs — four planted, FOUR SEEN FAILING
| # | Mutation | Layer the guard sits at | Observed failure |
|---|---|---|---|
| 1 | hoisted block moved back **below** `unhealthy > 0` | `aggregateState` — the ORDERING | `aggregateState = "unhealthy", want "degraded"` **and** `bookstack state = "unhealthy", want "degraded"` (production-path wiring) |
| 2 | `up` narrowed back to `running` alone | `aggregateState` — the GUARD | same subtest, plus `"starting"` and `"restarting"` — all three survivor shapes convict |
| 3 | classifier drops `StateDegraded` from `down` | `classifyRunStates` — the CONSEQUENCE | `dead-app banner = [], want exactly one entry for bookstack` |
| 4 | `undoCopyPhrase` reverted to the unconditional claim | the phrase builder — where the CLAIM is made | `phrase "…mentése megvan: …mariadb.sql" contains "mentése megvan" — it asserts a file that is not on disk`; the empty set printed `megvan: .`, naming a file that never existed |
**Every mutation was asserted to have applied** (the scripts `assert` the pre-fix text is present
before rewriting and print `MUTATION APPLIED`). **None passed first time.** Mutations 1 and 2 convict
independently, which is what proves the fix genuinely has two halves.
## 6. Test count
**1494 → 1504** top-level test functions (measured by `go test ./... -list '.*'` on the stashed and
unstashed tree, not estimated). Full green gate `go build && go vet && go test ./...` → **exit 0, zero
failures**, run after Part 2 and again after Part 3.
## 7. Deployed version, and the golden
```
gitea.dooplex.hu/admin/felhom-controller:0.222.0 Up 20 seconds (healthy)
```
**Golden BAKED and PUBLISHED: YES — version 0.222.0.**
`GOLDEN_SHA256 = 19f5904f53792684f046ec0bc25426645cb87ad73d5cfc6c03639d9f82706037`, `upload OK (HTTP 201)`,
round-trip `HTTP 206` from the package URL. All acceptance markers counted and recorded.
**VOUCHING IS THE OPERATOR'S ACT AND WAS NOT DONE HERE.**
## 8. The five `IsDownState` consumers, walked and named
| Consumer | What changes |
|---|---|
| `cmd/controller/main.go:2173` `classifyRunStates` | **THE INTENDED CHANGE.** A stack that read `unhealthy` now reads `degraded` → `down=true` → banner + `app_start_failed`. `userStopped` tests `StateStopped` specifically, so the whitelist cannot swallow `degraded`. |
| `internal/bootrecon/bootrecon.go:213` (DesiredStateRunning) | **CHANGES, and toward repair.** A half-started stack at boot now reads `degraded` → an orphan → `compose up -d`. Previously it read `unhealthy` → not an orphan → left half-dead. Aligned with the file's own stated intent. |
| `internal/bootrecon/bootrecon.go:216` (legacy DesiredStateUnknown) | Same shape, same direction. |
| `internal/bootrecon/bootrecon.go:288` (recovery check) | **CHANGES, and toward truth.** A stack that came back with a dead supervised member is no longer counted `Recovered`; it stays pending and is retried, bounded by `r.attempts`. It used to be declared recovered while half-dead. |
| `internal/stacks/desiredstate.go:150` `isObservedUp` | **UNAFFECTED — verified, not assumed.** It is an allow-list of `{running, starting}`; neither `unhealthy` nor `degraded` was ever in it, so a stack moving between them does not cross the boundary. |
| `internal/quiesce/suppress.go` | **UNAFFECTED.** It does not call `IsDownState` at all — the suppression is cycle-keyed and state-blind, which is precisely why R-97b's guarantee cannot be weakened by a state change. Pinned by `TestR384_QuiesceSuppressionStillHoldsForDegraded`. |
| dashboard state badge | **Already handled.** R-51 wired `degraded` through `handlers.go:158/169` (counts with stopped) and `funcmap.go:255` (filters with stopped). Verified live — the badge rendered `(degraded)`. |
## 9. The live walk
Method: endpoint-level. No browser exists on DooPlex; every read below is either the exact endpoint
the UI calls (`POST /api/stacks/<name>/<action>`, `GET /api/stacks`, `GET /dashboard`) or the
controller's own log. Guest clock is UTC.
### Step 1 — Scenario A: a database dies behind a healthy-looking app ✅
| Observable | Result |
|---|---|
| `bookstack-db` stopped out of band | 05:30:07Z |
| front end went `unhealthy` | 05:31:21Z — **the state that used to swallow the alarm** |
| aggregate state read | **`degraded`** while the front end was `unhealthy` (05:32:03Z) |
| `app_start_failed` | **fired at 05:30:14Z**, 7 s after the stop |
| new code path visible | `manager.go:703: restart-policy of down member "bookstack-db" = "unless-stopped"` |
| banner | *„Telepített alkalmazás nem fut: BookStack (degraded)"* on **both** `/launcher` and `/dashboard` |
| edge-triggered, not per-scan | **1 event across 22 scans** |
**The heartbeat, old beside new — same 8 apps evaluated:**
```
2026/08/22 21:13:48 [deadapp] check alive: 580 scans since boot, 8 deployed app(s) evaluated, 0 currently down ← v0.220.2/0.221.1
2026/08/23 05:37:44 [deadapp] check alive: 20 scans since boot, 8 deployed app(s) evaluated, 1 currently down ← v0.222.0
```
### Step 2 — Scenario B: unhealthy with nothing dead ✅
The database was restarted; the front end stayed `unhealthy` with nothing down. Aggregate read
**`unhealthy`**, not `degraded`. **0 new `app_start_failed`**, and — the positive observable —
**no `restart-policy of down member` line at all**, meaning the supervised path was not entered. That
absence is trustworthy because the same line HAD appeared on this box 10 minutes earlier.
Banner cleared; `bookstack state=running`.
*Honest limit:* the live window in which docker reported the front end unhealthy with the database up
was ~11 s wide, and the controller's cached read was taken at its edge. The three-shape unit test
`TestR384_UnhealthyWithNothingDeadDoesNotAlarm` carries the rest of this case.
### Step 3 — Scenario D: a full deploy cycle ✅ **0 alarms**
`POST /api/stacks/docmost/stop` then `/start` — the exact calls the launcher's buttons make — on a
3-container stack, watched for 5 minutes to settled healthy.
| Observable | Result |
|---|---|
| `app_start_failed` across the cycle | **0** |
| dead-app scans in the window | **9** |
| supervised-down path entered | 1, at 05:42:34Z (docmost itself momentarily down beside two live members) |
**Alarms that v0.221.1 would NOT have produced: ZERO.** The single `degraded` reading at 05:42:34Z
has `running > 0`, so v0.221.1's mixed-case branch reaches the identical verdict. No moment in the
cycle had a down supervised member with only non-`running` survivors, which is the only shape where
the two versions differ.
### Step 4 — Scenario E: the quiesce cycle — **DROPPED, and why**
Dropped as a direct consequence of the §4 halt (see the banner at the top), and it is item **(3)** on
the task's own drop list. Running it would have meant triggering a real backup cycle on the box
*after* a halt condition had already fired. **Covered at unit level instead** by
`TestR384_QuiesceSuppressionStillHoldsForDegraded`, which asserts a `degraded` stack inside the
quiesce set produces no banner and `Down=false`. **Not proven live in this session — stated plainly
rather than implied.**
### Step 5 — §4's measurement ✅ (it reproduced — see §10)
### Step 6 — Part 1's gate, both directions ✅
| Run | Gate | CHANGELOG | Golden | Exit |
|---|---|---|---|---|
| `gate-01` | **old** | v0.221.0 | 0.221.1 | **0** — the blindness, on the real history |
| `gate-02` | **new** | v0.221.0 | 0.221.1 | **1** — convicted, naming the missing heading and the route |
| `gate-03` | new | v0.221.1 | 0.221.1 | 0 |
| `gate-04` | new | absent clone | — | **2** — INCONCLUSIVE preserved |
| `gate-05` | new | v0.222.0 | 0.222.0 | 0 — post-bake |
## 10. §4's answer, in plain words
**A single-container app that is stopped out of band raises no alarm at all, and a comment in the
code says the opposite.**
`aggregateState` folds `StateExited` into the `stopped` counter, so when every member is down it
returns `StateStopped` — **`StateExited` never survives aggregation**, which is the path the task
suspected and could not find. `classifyRunStates` then whitelists `StateStopped` as a deliberate user
stop. So the comment at `cmd/controller/main.go` — *"An out-of-band `docker compose stop` leaves the
containers present → StateExited → still alerts, which is correct: out-of-band tampering IS
reportable"* — is **false**, and so is the neighbouring I2 claim that a crashing app never comes to
rest at `stopped`.
**Measured:** `privatebin` (1 container, `unless-stopped`) stopped 05:47:35Z. At 05:51:53Z:
`state=stopped`, **9 dead-app scans had run, 0 events, 0 banner lines.**
**Positive control first, per standing rule 3:** `app_start_failed` fired for BookStack at 05:30:14Z
on the same box 17 minutes earlier, so the detector was demonstrably alive.
**Scoped honestly:** a genuine *crash* under `unless-stopped` is restarted by Docker and surfaces as
`restarting` → the 5-minute crash-loop path, which does alarm. The silent case is an explicit
out-of-band stop of a stack with no surviving member.
**Filed as R-386 (OPEN — MEDIUM). Not fixed here, per §12 and §13.**
## 11. Evidence
`felhom.eu/documentation/audits/DRILL-r384-dead-db-alarm-2026-08-23/evidence/` — 5 gate runs, 4
red-proof transcripts, 20 live-walk files including two full controller-log windows (1808 and 2139
lines) pulled off the guest. **Both log windows were copied off before the app was restarted**, per
standing rule 5.
## 12. Teardown, three layers, and the box's end state
1. **Guest 9201 / apps** — nothing provisioned. `bookstack-db` restarted and **`bookstack` confirmed
healthy**; `privatebin` restarted and healthy; `docmost` (all 3) healthy. Planted data untouched
throughout — no app was rebuilt, redeployed or restored.
2. **Bake VM** — the drill VM on DooPlex ran the bake and its build guest 9100 is stopped inside it;
the qcow2 reverts to the `virgin` snapshot. No storage was added anywhere, so `pvesm status` has
nothing to compare.
3. **Hub-side record — stated explicitly even though there is none.** No appliance was registered, no
customer created, no config written, no artifact manifest changed. **The hub was READ ONLY**
(`GET /configuration`, `GET /events`). Nothing to discard.
**End state:** `demo-hp` guest 9201 runs controller **0.222.0**, all 8 deployed apps healthy, golden
0.222.0 baked and published but **NOT vouched** — floor still **0.221.1**.
## 13. Register size
| File | Before | After |
|---|---|---| |---|---|---|
| `docmost` | `5d35678349bbbdb318ac22656d4f96436b5751ef5b3b1659d48305a526df1aed` | **identical** | | `OPEN-ITEMS.md` | 327,266 B | **328,325 B** |
| `bookstack` | `7837aa5de2955dcf3125534f015f43df42debe13c765793d3b75173abcc55e7b` | **identical** | | `CLOSED-ITEMS.md` | 68,464 B | **71,441 B** |
**Before the fix, on the same box: neither app had a canonical dump at all.** R-383 and R-384 moved to CLOSED compressed; R-385 (closed) and R-386 (open) filed. OPEN grew by
~1 KB despite two closures because R-386 is a substantial new finding — recorded rather than
smoothed over.
`DumpOneTo` takes an explicit final path and derives its `.tmp` from it. `DumpOne`'s signature did not ## 14. Observations — noticed, documented, NOT acted on
move. The comment that asserted the old behaviour was safe now states the invariant and how it is
enforced.
## Manifest — every consumer named, as required 1. **R-329 is live and now matters much more.** `app_start_failed` is pushed with severity **`warn`**,
which is not in the hub's vocabulary (`{info, warning, error, critical}`) and coerces silently to
`Manifest.DBDumps` has **three** references, all in `internal/backup/recovery_unit.go`: the `info`, e-mailing nobody, while the POST still returns 200. Observed again today:
declaration (`:55`), the enumeration, and the already-current compare. **None reads it for recovery; `PushEvent: type=app_start_failed severity=warn`. **R-384 makes this event actually fire, so a
no hub or agent consumer exists.** Undo copies are therefore excluded from `db_dumps`. Verified live: known-broken severity moved from unreachable to load-bearing.** Not in scope; not touched.
`db_dumps: ['docmost-postgres.sql']` with **4 undo copies on disk** as the positive control. 2. **Two files carry pre-existing `gofmt` drift** — `internal/backup/offbox.go` and
`internal/backup/offbox_recovery_cli.go`. Confirmed pre-existing by stashing this session's work
## Part 3 — the measurement that cancelled Part 2 and re-running `gofmt -l`. Not touched (§12 forbids nearby refactors).
3. **The runbook's golden-bake step is missing `pveam update`.** On the `virgin` snapshot the template
**The reading did not reproduce.** A held app aggregates to `unhealthy`; `IsDownState` is index is stale, so `pveam available` offers `13.1-2` and downloading it fails with
`{stopped, exited, degraded}`; the dead-app heartbeat read `0 currently down` across scans 600 and `400 Parameter verification failed. template: no such template`. Recorded in the bake evidence
620 while a hold was in force. **Part 2 dropped in full**, no register row opened, negative recorded README; the runbook itself was not edited.
in the capability map. 4. **The register's own suggested fix for R-384 was wrong** — it proposed a sustained-`unhealthy`
threshold on the `crashLoopAfter` model. The defect needed no threshold at all, only an ordering.
**The positive control took three attempts and that is the finding.** Two live attempts failed — Recorded in the CLOSED entry so the next reader sees that a register remedy is a hypothesis.
`privatebin` went `stopped` (whitelisted by design) and `bookstack` went `degraded` then `unhealthy`. 5. **Deliberately left open, untouched:** R-102, R-359, R-361's sibling surfaces.
The control that works is at the detector's own layer: `classifyRunStates` is pure and raises the
banner for `degraded`/`exited` while staying silent for the states measured live.
## Findings filed
- **R-383 (MEDIUM)** — the double-failure message says the previous state's backup **exists** while
naming the very file whose absence caused the failure. Observed twice, on v0.220.2 and v0.221.1.
**R-361's own class**, one surface over.
- **R-384 (MEDIUM)** — an app whose **database** has died reads `unhealthy` and raises no dead-app
alarm. `bookstack-db` stopped at 21:27:01; `0 currently down` throughout.
**Register: `OPEN-ITEMS.md` 325 236 → 327 266 bytes; `CLOSED-ITEMS.md` 66 777 → 68 464.** R-361
compressed; full text `git show a8caa0fdde7c:documentation/backlog/OPEN-ITEMS.md`.
## Two defects in this session's own work
1. **Two red-proofs passed first time**, both reported. The behavioural tests inject the dump seam, so
a mutation *inside* `DumpOneTo` was invisible; and Part 1.3 had no test at all. Guards added at the
layer each defect lives in; both mutations then convicted.
2. **One change made another unreachable.** A stable `db_dumps` let the already-current early return
fire, and the undo prune sat after it — **four copies against a cap of three, counted on the box**.
Fixed in v0.221.1; the cap now holds at 3 on both apps, verified live.
**Test count 1485 → 1493.** Green gate clean; controller gates all OK; **no push used `--no-verify`.**
## Part 4 — the double-failure path on what ships
Trigger: **the undo copy is lost between being written and being needed** — realistic, and one of only
two ways a rollback can fail. Verified on 0.221.1: held, refused by the customer button, refused after
a full controller restart, then listed/cleared/started via the operator route. Message captured
verbatim (369 bytes, hex in the drill record) with **no engine output**.
## Teardown
Nothing provisioned; 20 containers healthy; `docmost` still holds its 4 pages and 1 user; both
canonical dumps present; **no app left held**; no hub-side record created.
## Operator follow-up
Vouch a golden carrying **0.221.1** (baked in this session), then raise the floor to 0.221.1 last, in
its own save.
+10 -1
View File
@@ -702,10 +702,19 @@ When app templates are updated (e.g., a new `APP_KEY` secret is added to `.felho
| Running + starting | Orange | "Indulas..." | Healthcheck not yet passed | | Running + starting | Orange | "Indulas..." | Healthcheck not yet passed |
| Deploying | Orange | "Telepítés..." | Compose up in progress (image pull, container creation) | | Deploying | Orange | "Telepítés..." | Compose up in progress (image pull, container creation) |
| Running + unhealthy | Yellow | "Nem egeszseges" | Docker or controller-side healthcheck failing | | Running + unhealthy | Yellow | "Nem egeszseges" | Docker or controller-side healthcheck failing |
| **Degraded** | Red | "Leallitva" (counts with stopped) | **A SUPERVISED member of a multi-container app is dead** — e.g. the app's database — while other members are still up. Counts as DOWN: raises the dead-app banner and `app_start_failed` (R-51, R-384) |
| Stopped/exited | Red | "Leallitva" | All containers stopped | | Stopped/exited | Red | "Leallitva" | All containers stopped |
| Restarting | Yellow | "Ujrainditas..." | Restart loop | | Restarting | Yellow | "Ujrainditas..." | Restart loop; becomes down only after 5 minutes sustained (crash loop) |
| Not deployed | Gray | "Nincs telepitve" | Compose file exists, not deployed | | Not deployed | Gray | "Nincs telepitve" | Compose file exists, not deployed |
**The order these are decided in is load-bearing (v0.222.0, R-384).** `degraded` is evaluated
**before** `unhealthy`/`starting`/`restarting`. An app whose database dies drags its own front end
`unhealthy` seconds later — so if `unhealthy` were decided first (as it was until v0.222.0), the
symptom would mask the fault and the app would be silently down. A down member whose restart policy
is `no`/`on-failure` is a finished one-shot init/migrate container and stays benign. The full ladder,
including which states deliberately do NOT alarm and why, is documented in
`felhom.eu/documentation/architecture/08-alarm-ladder.md`.
**Route-unpublished indicator (F5, v0.61.0).** Traefik's Docker provider only publishes a route to a **Route-unpublished indicator (F5, v0.61.0).** Traefik's Docker provider only publishes a route to a
container that is healthy (or has no healthcheck), so an `unhealthy`/`restarting` deployed app returns a container that is healthy (or has no healthcheck), so an `unhealthy`/`restarting` deployed app returns a
hard **404** at its URL even though the container is running. The `routeUnpublished` template helper hard **404** at its URL even though the container is running. The `routeUnpublished` template helper