R-385: make an UNRECORDED golden fail the currency gate; file R-386; own the alarm ladder
gates / gates (push) Successful in 17s
gates / gates (push) Successful in 17s
The gate failed only on `released > baked`, so it could catch a forgotten bake and nothing else. A golden AHEAD of the record passed silently - and that is how controller 0.221.1 was built, baked AND vouched while the newest CHANGELOG heading still read v0.221.0, with every gate green. Reproduced on the real history: newest released 0.221.0 / newest golden baked 0.221.1 -> exit 0. The gate now asks whether the version being shipped is WRITTEN DOWN: the baked version must have its own `## vX.Y.Z` heading anywhere in the CHANGELOG. Membership rather than `baked > released` deliberately - a comparison against the newest heading alone goes green the moment any later entry is written, leaving the unrecorded version permanently unrecorded. INCONCLUSIVE (exit 2) preserved; every refusal names a reason and a route. Red-proofed both directions: old gate/old record exit 0, new gate/old record exit 1, new gate/fixed record exit 0, absent clone exit 2, post-bake exit 0. 08-alarm-ladder.md is new, and its absence was itself the finding: no document owned "when does a broken app raise an alarm?". The rules lived as comments in four packages, each locally correct, with the ordering between them legible only by reading one function top to bottom - which is how R-384 survived review. R-383 and R-384 closed into CLOSED-ITEMS with their rules kept. R-385 filed closed. R-386 filed OPEN: a single-container app stopped out of band raises no alarm, and a comment claims the opposite - measured live, 9 scans, 0 events, against a positive control from the same box 17 minutes earlier. Not fixed here. Golden 0.222.0 baked and published; vouching is the operator's act.
This commit is contained in:
@@ -67,6 +67,23 @@ the banner, pinned by `TestClassifyRunStates_PositiveControl_ADownStackDoesAlarm
|
||||
measurement DID expose is **R-384**: an app whose database has died is `unhealthy` too, and is
|
||||
likewise silent. Evidence: `audits/DRILL-r361-2026-08-22/evidence/06-part3-decision.txt`.
|
||||
|
||||
> **R-384 CLOSED in controller v0.222.0 (2026-08-23), proven live.** The defect was the ORDER of two
|
||||
> questions, not the `unhealthy` exclusion: `aggregateState` now asks *"is a supervised member dead?"*
|
||||
> **before** the `unhealthy`/`starting`/`restarting` returns, and "some members are up" counts any
|
||||
> member not in the down bucket rather than `running` alone. `IsDownState` is byte-identical.
|
||||
> **Measured on `demo-hp` 2026-08-23** with the same fixture that read `0 currently down` the day
|
||||
> before: `bookstack-db` stopped 05:30:07Z → `app_start_failed` fired at **05:30:14Z**, the banner read
|
||||
> *„Telepített alkalmazás nem fut: BookStack (degraded)"*, the stack read `state=degraded` **while its
|
||||
> front end was `unhealthy`**, and the heartbeat printed **`1 currently down`** against the previous
|
||||
> day's `0`. Evidence: `audits/DRILL-r384-dead-db-alarm-2026-08-23/`.
|
||||
>
|
||||
> **The HELD-app half of the paragraph above is now also covered** — a held app keeps its database
|
||||
> container, so it is the same shape and reaches the same `degraded` verdict.
|
||||
>
|
||||
> **The alarm ladder that decides all of this now has an owning document:** see
|
||||
> `08-alarm-ladder.md` (written 2026-08-23 — before that date no document owned it, and that absence
|
||||
> is why the ordering defect was legible only from source).
|
||||
|
||||
**WHAT IS STILL NOT CLAIMED:** the FAILURE path is where this class is weak, not the success path — a corrupt dump leaves the customer with an emptied or partially-applied database and an undo copy **no product action can apply** (**R-379**), and on MariaDB it does so behind an app that reports `health=healthy` (**R-380**). The success story is proven; the recovery-from-a-bad-restore story is not.
|
||||
|
||||
**NARROWED 2026-08-21 by the backup-truth drill — kept as history; both halves have since closed, see the entry immediately above.** Proven that night on `demo-hp` with planted, hash-recorded files: the **declared-userdata** leg of a **drive-declaring** app does come back byte-identical (`calibre-web`, 5/5 including two Hungarian accented filenames). **Two legs of the same story do NOT:** (a) the off-site restore has **no named-volume leg at all**, so an app's volume tar sits in the unit, in the snapshot and in the checking folder and is never replayed (**R-354**); (b) for the **40 of 53** apps that declare no data drive the off-site restore **refuses outright**, saying a running app „nincs telepítve" (**R-356**). Since the 40-class keeps ALL its data in named volumes, the end-to-end story is **unproven for that class and disproven for the volume leg generally**. The escrow/key half of this row is untouched by that and still stands. Evidence: `audits/DRILL-backup-truth-2026-08-21/evidence/` and `REPORT.md` (2026-08-21). |
|
||||
|
||||
Reference in New Issue
Block a user