R-385: make an UNRECORDED golden fail the currency gate; file R-386; own the alarm ladder
gates / gates (push) Successful in 17s
gates / gates (push) Successful in 17s
The gate failed only on `released > baked`, so it could catch a forgotten bake and nothing else. A golden AHEAD of the record passed silently - and that is how controller 0.221.1 was built, baked AND vouched while the newest CHANGELOG heading still read v0.221.0, with every gate green. Reproduced on the real history: newest released 0.221.0 / newest golden baked 0.221.1 -> exit 0. The gate now asks whether the version being shipped is WRITTEN DOWN: the baked version must have its own `## vX.Y.Z` heading anywhere in the CHANGELOG. Membership rather than `baked > released` deliberately - a comparison against the newest heading alone goes green the moment any later entry is written, leaving the unrecorded version permanently unrecorded. INCONCLUSIVE (exit 2) preserved; every refusal names a reason and a route. Red-proofed both directions: old gate/old record exit 0, new gate/old record exit 1, new gate/fixed record exit 0, absent clone exit 2, post-bake exit 0. 08-alarm-ladder.md is new, and its absence was itself the finding: no document owned "when does a broken app raise an alarm?". The rules lived as comments in four packages, each locally correct, with the ordering between them legible only by reading one function top to bottom - which is how R-384 survived review. R-383 and R-384 closed into CLOSED-ITEMS with their rules kept. R-385 filed closed. R-386 filed OPEN: a single-container app stopped out of band raises no alarm, and a comment claims the opposite - measured live, 9 scans, 0 events, against a positive control from the same box 17 minutes earlier. Not fixed here. Golden 0.222.0 baked and published; vouching is the operator's act.
This commit is contained in:
@@ -0,0 +1,96 @@
|
||||
# DRILL — R-384: an app whose database dies raised no alarm (2026-08-23)
|
||||
|
||||
**Controller v0.221.1 → v0.222.0. Live leg on `demo-hp` (Tier 0, disposable), guest 9201.**
|
||||
**UNATTENDED.** Method: endpoint-level — no browser exists on DooPlex, so every read is either the
|
||||
exact endpoint the UI calls or the controller's own log. Guest clock is UTC.
|
||||
|
||||
## Verdict
|
||||
|
||||
| Part | Outcome |
|
||||
|---|---|
|
||||
| Part 0 — the record | ✅ `v0.221.1` given its own heading, pushed ALONE (`da75603`) |
|
||||
| Part 1 — the blind gate | ✅ fixed; both directions red-proofed against the real history |
|
||||
| Part 2 — R-384 | ✅ shipped v0.222.0, **proven live** |
|
||||
| Part 3 — R-383 | ✅ shipped v0.222.0 |
|
||||
| §4 — the measurement | ⚠ **reproduced. Filed as R-386. NOT fixed — that was the instruction.** |
|
||||
| Live walk step 4 (Scenario E) | **DROPPED** — consequence of the halt; drop-list item (3) |
|
||||
|
||||
## The one-line result
|
||||
|
||||
The same fixture that printed **`0 currently down`** on 2026-08-22 printed **`1 currently down`** on
|
||||
2026-08-23, with 8 apps evaluated both times:
|
||||
|
||||
```
|
||||
2026/08/22 21:13:48 [deadapp] check alive: 580 scans since boot, 8 deployed app(s) evaluated, 0 currently down
|
||||
2026/08/23 05:37:44 [deadapp] check alive: 20 scans since boot, 8 deployed app(s) evaluated, 1 currently down
|
||||
```
|
||||
|
||||
## What was actually wrong
|
||||
|
||||
**The ORDER of two questions**, not the `unhealthy` exclusion. "Is a supervised member dead?" and "is
|
||||
a running member failing its healthcheck?" are different questions, and the second was answering the
|
||||
first — because a dying database drags its own front end `unhealthy`, **the symptom the fault causes
|
||||
was what suppressed the alarm for it.** `IsDownState` was not touched; no state was minted.
|
||||
|
||||
The fix has **two halves and either alone leaves the defect standing**: the hoist, and widening "some
|
||||
members are up" from `running > 0` to *any member not in the down bucket*. The old guard made the
|
||||
R-51 block unreachable in precisely the case R-51 was written for.
|
||||
|
||||
## Evidence index (`evidence/`)
|
||||
|
||||
| File | What it shows |
|
||||
|---|---|
|
||||
| `gate-01-old-gate-old-changelog.txt` | the blindness: old gate, real history, **exit 0** |
|
||||
| `gate-02-new-gate-old-changelog.txt` | new gate on the same history, **exit 1**, naming the fix |
|
||||
| `gate-03-new-gate-new-changelog.txt` | with Part 0's heading, **exit 0** |
|
||||
| `gate-04-inconclusive.txt` | INCONCLUSIVE (**exit 2**) preserved |
|
||||
| `gate-05-new-gate-post-bake.txt` | v0.222.0 + golden 0.222.0, **exit 0** |
|
||||
| `redproof-R384-1-order.txt` | mutation: hoist reverted → `"unhealthy", want "degraded"` |
|
||||
| `redproof-R384-2-upguard.txt` | mutation: `up` narrowed → all three survivor shapes convict |
|
||||
| `redproof-R384-3-classifier.txt` | mutation: classifier ignores `degraded` → empty banner |
|
||||
| `redproof-R383-undo-phrase.txt` | mutation: unconditional claim → prints the false sentence |
|
||||
| `live-00-the-old-heartbeat-2026-08-22.txt` | yesterday's `0 currently down`, carried in for contrast |
|
||||
| `live-01`…`live-06` | Scenario A: pre-state, stop, log, decisive read, heartbeat, banner |
|
||||
| `live-08-golden-bake-markers.txt` | the bake's acceptance markers, each counted |
|
||||
| `live-09-…-full-controller-log.txt` | 1808 lines, pulled off **before** the app was restarted |
|
||||
| `live-10`…`live-13` | Scenario B: unhealthy with nothing dead, no alarm, banner cleared |
|
||||
| `live-14`…`live-16` | Scenario D: full stop→start cycle, **0 alarms across 9 scans** |
|
||||
| `live-17`, `live-18` | §4: `privatebin` out-of-band stop, 9 scans, **0 events, 0 banner** |
|
||||
| `live-19-…-full-log.txt` | 2139 lines covering Scenario D and §4 |
|
||||
|
||||
## §4 — what it found, in plain words
|
||||
|
||||
`aggregateState` folds `StateExited` into the `stopped` counter, so an all-down stack returns
|
||||
`StateStopped` and **`StateExited` never survives aggregation** — that is the path the task suspected
|
||||
and could not find in source. `classifyRunStates` then whitelists `StateStopped` as a deliberate user
|
||||
stop. So this comment in `cmd/controller/main.go` is **false**:
|
||||
|
||||
> *"(An out-of-band `docker compose stop` leaves the containers present → StateExited → still alerts,
|
||||
> which is correct: out-of-band tampering IS reportable.)"*
|
||||
|
||||
Measured: `privatebin` (1 container, `unless-stopped`) stopped 05:47:35Z; at 05:51:53Z it read
|
||||
`state=stopped` with 9 dead-app scans behind it, **zero events and zero banner lines**.
|
||||
|
||||
**The absence is trustworthy because the detector was shown alive first** (standing rule 3):
|
||||
`app_start_failed` fired for BookStack at 05:30:14Z on the same box 17 minutes earlier.
|
||||
|
||||
**Scoped honestly:** a genuine crash under `unless-stopped` is restarted by Docker and surfaces as
|
||||
`restarting` → the 5-minute crash-loop path, which does alarm. The silent case is an explicit
|
||||
out-of-band stop of a stack with no surviving member.
|
||||
|
||||
## Two things noticed that are NOT this drill's work
|
||||
|
||||
1. **R-329 moved from unreachable to load-bearing.** `app_start_failed` ships severity `warn`, which
|
||||
is not in the hub's vocabulary and coerces silently to `info` — e-mailing nobody, POST still 200.
|
||||
Observed again today. While the event never fired, this was harmless; it no longer is.
|
||||
2. **The golden-bake runbook is missing `pveam update`.** On the `virgin` snapshot the template index
|
||||
is stale, so `pveam available` offers `13.1-2` and downloading it fails with
|
||||
`400 Parameter verification failed. template: no such template` — a confusing 400 rather than a
|
||||
legible "your index is old". Recorded in `documentation/tests/golden-0.222.0-2026-08-23/README.md`.
|
||||
|
||||
## Teardown
|
||||
|
||||
Nothing was provisioned. All apps restored and confirmed healthy (`bookstack` + `bookstack-db`,
|
||||
`docmost` ×3, `privatebin`); planted data untouched; no app rebuilt or restored. **Hub-side: nothing
|
||||
to discard — the hub was READ ONLY this session** (`GET /configuration`, `GET /events`); no appliance
|
||||
registered, no config written, no artifact manifest changed.
|
||||
+5
@@ -0,0 +1,5 @@
|
||||
### OLD gate + OLD changelog (v0.221.0 top, golden 0.221.1 baked) ###
|
||||
newest released controller : 0.221.0 (## v0.221.0 — taking the undo copy destroyed the app's own database backup (2026-08-22, R-)
|
||||
newest golden baked : 0.221.1 (documentation/tests/golden-0.221.1-2026-08-23)
|
||||
golden currency gate OK — the newest released controller has a golden (NOTE: this checks the BAKE, not the vouch — see the module docstring)
|
||||
EXIT=0
|
||||
+9
@@ -0,0 +1,9 @@
|
||||
### NEW gate + OLD changelog (v0.221.0 top, golden 0.221.1 baked) -> must FAIL ###
|
||||
newest released controller : 0.221.0 (## v0.221.0 — taking the undo copy destroyed the app's own database backup (2026-08-22, R-)
|
||||
newest golden baked : 0.221.1 (documentation/tests/golden-0.221.1-2026-08-23)
|
||||
|
||||
GOLDEN CURRENCY GATE FAILED: golden 0.221.1 is baked but UNRECORDED — the controller CHANGELOG has no '## v0.221.1' heading.
|
||||
The newest heading is 0.221.0. A golden ahead of the record was built from a version nobody wrote down, so no one can read what the fleet is running.
|
||||
Fix: give v0.221.1 its own '## v0.221.1 — <what changed>' heading in felhom-controller/CHANGELOG.md, above the entries it supersedes. If its fix is currently described inside another version's entry, MOVE that text — do not duplicate it, and do not delete the reasoning.
|
||||
If this bake was a throwaway that must never be delivered, delete its documentation/tests/golden-<VER>-<DATE>/ directory — never leave it to read as shipped.
|
||||
EXIT=1
|
||||
+5
@@ -0,0 +1,5 @@
|
||||
### NEW gate + FIXED changelog (v0.221.1 heading present) -> must PASS ###
|
||||
newest released controller : 0.221.1 (## v0.221.1 — the undo-copy prune stopped running because another fix made its guard reach)
|
||||
newest golden baked : 0.221.1 (documentation/tests/golden-0.221.1-2026-08-23)
|
||||
golden currency gate OK — the newest released controller has a golden (NOTE: this checks the BAKE, not the vouch — see the module docstring)
|
||||
EXIT=0
|
||||
+3
@@ -0,0 +1,3 @@
|
||||
### INCONCLUSIVE preserved: unreadable CHANGELOG path ###
|
||||
GOLDEN CURRENCY GATE INCONCLUSIVE: controller clone not found at /nonexistent/CHANGELOG.md
|
||||
EXIT=2
|
||||
+5
@@ -0,0 +1,5 @@
|
||||
### NEW gate, post-bake: CHANGELOG v0.222.0 + golden 0.222.0 -> must PASS ###
|
||||
newest released controller : 0.222.0 (## v0.222.0 — an app whose database dies raised no alarm, because the wrong question answe)
|
||||
newest golden baked : 0.222.0 (documentation/tests/golden-0.222.0-2026-08-23)
|
||||
golden currency gate OK — the newest released controller has a golden (NOTE: this checks the BAKE, not the vouch — see the module docstring)
|
||||
EXIT=0
|
||||
+18
@@ -0,0 +1,18 @@
|
||||
=== PART 3 MEASUREMENT — v0.220.2, no code change
|
||||
measured at: 21:13:51Z (hold created 21:11:19Z)
|
||||
elapsed: ~2m20s = 5 deadapp scans at 30s cadence (heartbeat lines confirm 9 scans in the log window)
|
||||
|
||||
-- 3a. WHICH RUN STATE did docmost aggregate to?
|
||||
bookstack state=running deployed=True
|
||||
docmost state=unhealthy deployed=True
|
||||
privatebin state=running deployed=True
|
||||
|
||||
-- 3b. IS THE DATABASE STILL UP?
|
||||
docmost-postgres | Up 2 minutes (healthy)
|
||||
|
||||
-- 3c. DID A CUSTOMER-FACING EVENT FIRE for docmost?
|
||||
2026/08/22 21:11:19 notifier.go:234: [INFO] Event pushed: backup_run_failures (error) — App "docmost" is HELD STOPPED: its off-site database restore failed AND the rollback to the customer's own pre-restore copy also failed. The app will not start from any path until the hold is cleared. Replay error: importing postgres dump for docmost: postgres import into docmost-postgres failed: exit status 3. Rollback error: a visszavonáshoz szükséges mentés nem található (pre-restore-20260822T211114Z-docmost-postgres.sql): stat /mnt/sys_drive/felhom-data/backups/primary/docmost/db-dumps/pre-restore-20260822T211114Z-docmost-postgres.sql: no such file or directory
|
||||
(none above = no event fired)
|
||||
|
||||
-- 3d. DEAD-APP BANNER STATE:
|
||||
2026/08/22 21:13:48 main.go:1730: [INFO] [deadapp] check alive: 580 scans since boot, 8 deployed app(s) evaluated, 0 currently down
|
||||
+11
@@ -0,0 +1,11 @@
|
||||
=== SCENARIO A — a database dies behind a healthy-looking app (v0.222.0) ===
|
||||
controller version:
|
||||
gitea.dooplex.hu/admin/felhom-controller:0.222.0 Up 3 minutes (healthy)
|
||||
|
||||
-- PRE-STATE (guest UTC) --
|
||||
2026-08-23T05:30:02Z
|
||||
bookstack Up 3 hours (healthy)
|
||||
bookstack-db Up 3 hours (healthy)
|
||||
|
||||
/bookstack restart=unless-stopped
|
||||
/bookstack-db restart=unless-stopped
|
||||
+6
@@ -0,0 +1,6 @@
|
||||
-- STOPPING bookstack-db OUT OF BAND --
|
||||
2026-08-23T05:30:07Z
|
||||
bookstack-db
|
||||
2026-08-23T05:30:07Z
|
||||
bookstack Up 3 hours (healthy)
|
||||
bookstack-db Exited (0) Less than a second ago
|
||||
+18
@@ -0,0 +1,18 @@
|
||||
-- controller log since the stop (05:30:00Z) --
|
||||
2026/08/23 05:30:04 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping bookstack — last check 3m10s ago, effective interval 5m0s, healthy=true
|
||||
2026/08/23 05:30:14 scheduler.go:67: [DEBUG] [scheduler] job deadapp-check: execution starting
|
||||
2026/08/23 05:30:14 manager.go:703: [DEBUG] [stacks] restart-policy of down member "bookstack-db" = "unless-stopped"
|
||||
2026/08/23 05:30:14 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping bookstack — last check 3m20s ago, effective interval 5m0s, healthy=true
|
||||
2026/08/23 05:30:14 notifier.go:206: [DEBUG] PushEvent: type=app_start_failed severity=warn url=https://hub.felhom.eu/api/v1/event
|
||||
2026/08/23 05:30:14 notifier.go:232: [DEBUG] PushEvent: app_start_failed pushed OK (HTTP 200)
|
||||
2026/08/23 05:30:14 notifier.go:234: [INFO] Event pushed: app_start_failed (warn) — Telepített alkalmazás nem fut: BookStack
|
||||
2026/08/23 05:30:24 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping bookstack — last check 3m30s ago, effective interval 5m0s, healthy=true
|
||||
2026/08/23 05:30:34 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping bookstack — last check 3m40s ago, effective interval 5m0s, healthy=true
|
||||
2026/08/23 05:30:44 scheduler.go:67: [DEBUG] [scheduler] job deadapp-check: execution starting
|
||||
2026/08/23 05:30:44 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping bookstack — last check 3m50s ago, effective interval 5m0s, healthy=true
|
||||
2026/08/23 05:30:44 manager.go:502: [DEBUG] [stacks] ScanStacks: found stack "bookstack" deployed=true composePath=/opt/docker/stacks/bookstack/docker-compose.yml
|
||||
2026/08/23 05:30:54 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping bookstack — last check 4m0s ago, effective interval 5m0s, healthy=true
|
||||
2026/08/23 05:31:04 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping bookstack — last check 4m10s ago, effective interval 5m0s, healthy=true
|
||||
2026/08/23 05:31:14 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping bookstack — last check 4m20s ago, effective interval 5m0s, healthy=true
|
||||
2026/08/23 05:31:14 scheduler.go:67: [DEBUG] [scheduler] job deadapp-check: execution starting
|
||||
2026/08/23 05:31:24 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping bookstack — last check 4m30s ago, effective interval 5m0s, healthy=true
|
||||
+13
@@ -0,0 +1,13 @@
|
||||
-- DECISIVE READ: front end UNHEALTHY, database EXITED --
|
||||
2026-08-23T05:32:03Z
|
||||
bookstack Up 3 hours (unhealthy)
|
||||
bookstack-db Exited (0) About a minute ago
|
||||
|
||||
bookstack state=degraded deployed=True
|
||||
calibre-web state=running deployed=True
|
||||
docmost state=running deployed=True
|
||||
kimai state=running deployed=True
|
||||
opengist state=running deployed=True
|
||||
paperless-ngx state=running deployed=True
|
||||
privatebin state=running deployed=True
|
||||
romm state=running deployed=True
|
||||
+2
@@ -0,0 +1,2 @@
|
||||
-- waiting for the F-OBS heartbeat (every 20 scans = ~10 min) --
|
||||
2026/08/23 05:37:44 main.go:1730: [INFO] [deadapp] check alive: 20 scans since boot, 8 deployed app(s) evaluated, 1 currently down
|
||||
+9
@@ -0,0 +1,9 @@
|
||||
-- the alert surface: /api/alerts + the launcher page --
|
||||
### /api/alerts
|
||||
http=404 bytes=42
|
||||
### /launcher
|
||||
http=200 bytes=42201
|
||||
<span class="alert-message">Telepített alkalmazás nem fut: BookStack (degraded)</span>
|
||||
### /dashboard
|
||||
http=200 bytes=56437
|
||||
<span class="alert-message">Telepített alkalmazás nem fut: BookStack (degraded)</span>
|
||||
+2
@@ -0,0 +1,2 @@
|
||||
-- hub-side: the event as the hub STORED it (R-329 severity check) --
|
||||
http=404
|
||||
+13
@@ -0,0 +1,13 @@
|
||||
=== GOLDEN 0.222.0 BAKE — acceptance markers ===
|
||||
docker OK (overlay2 : 1
|
||||
including mount point: 2
|
||||
upload OK (HTTP 201) : 1
|
||||
excluding (must be 0): 0
|
||||
FATAL (must be 0): 0
|
||||
--- the marker lines ---
|
||||
docker OK (overlay2; data-root /var/lib/docker)
|
||||
INFO: including mount point rootfs ('/') in backup
|
||||
INFO: including mount point mp0 ('/var/lib/felhom') in backup
|
||||
[golden] upload OK (HTTP 201)
|
||||
GOLDEN_VERSION=0.222.0
|
||||
GOLDEN_SHA256=19f5904f53792684f046ec0bc25426645cb87ad73d5cfc6c03639d9f82706037
|
||||
+1808
File diff suppressed because it is too large
Load Diff
+7
@@ -0,0 +1,7 @@
|
||||
=== SCENARIO B — unhealthy with NOTHING dead (the flapping case) ===
|
||||
-- restarting bookstack-db; the front end stays unhealthy for a while with nothing down --
|
||||
2026-08-23T05:40:01Z
|
||||
bookstack-db
|
||||
2026-08-23T05:40:09Z
|
||||
bookstack Up 3 hours (unhealthy)
|
||||
bookstack-db Up 8 seconds (healthy)
|
||||
+6
@@ -0,0 +1,6 @@
|
||||
-- SCENARIO B decisive read: front end unhealthy, database UP, nothing dead --
|
||||
2026-08-23T05:40:20Z
|
||||
bookstack Up 3 hours (healthy)
|
||||
bookstack-db Up 18 seconds (healthy)
|
||||
|
||||
bookstack state=unhealthy
|
||||
+7
@@ -0,0 +1,7 @@
|
||||
-- SCENARIO B: any NEW alarm during the recovery window? --
|
||||
app_start_failed events since 05:40:00Z: 0
|
||||
--- what the aggregate read during the window ---
|
||||
2026/08/23 05:40:04 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping bookstack — last check 3m0s ago, effective interval 5m0s, healthy=true
|
||||
2026/08/23 05:40:14 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping bookstack — last check 3m10s ago, effective interval 5m0s, healthy=true
|
||||
2026/08/23 05:40:24 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping bookstack — last check 3m20s ago, effective interval 5m0s, healthy=true
|
||||
2026/08/23 05:40:34 healthprobe.go:53: [DEBUG] [stacks] RunHealthProbes: skipping bookstack — last check 3m30s ago, effective interval 5m0s, healthy=true
|
||||
+6
@@ -0,0 +1,6 @@
|
||||
-- banner cleared + bookstack healthy again --
|
||||
2026-08-23T05:40:53Z
|
||||
bookstack Up 3 hours (healthy)
|
||||
bookstack-db Up 52 seconds (healthy)
|
||||
banner lines: 0
|
||||
bookstack state=running
|
||||
+13
@@ -0,0 +1,13 @@
|
||||
=== SCENARIO D — a full stop -> start cycle through the PRODUCTION endpoint ===
|
||||
method: POST /api/stacks/docmost/{stop,start} — the exact call the launcher's buttons make
|
||||
subject: docmost (3 containers: docmost, docmost-postgres, docmost-redis)
|
||||
csrf len=64
|
||||
BASELINE alarms before: 0
|
||||
T0=2026-08-23T05:42:15Z --- STOP ---
|
||||
{"ok":true,"message":"Stack docmost stop completed"}
|
||||
http=200
|
||||
|
||||
after stop: 2026-08-23T05:42:28Z
|
||||
--- START ---
|
||||
{"ok":true,"message":"Stack docmost start completed"}
|
||||
http=200
|
||||
+11
@@ -0,0 +1,11 @@
|
||||
-- SCENARIO D: watching a FULL cycle settle (5 min past the start) --
|
||||
05:42:47 docmost:Up 7 seconds (health: starting) docmost-postgres:Up 18 seconds (healthy) docmost-redis:Up 18 seconds (healthy)
|
||||
05:43:12 docmost:Up 33 seconds (healthy) docmost-postgres:Up 43 seconds (healthy) docmost-redis:Up 43 seconds (healthy)
|
||||
05:43:37 docmost:Up 58 seconds (healthy) docmost-postgres:Up About a minute (healthy) docmost-redis:Up About a minute (healthy)
|
||||
05:44:02 docmost:Up About a minute (healthy) docmost-postgres:Up About a minute (healthy) docmost-redis:Up About a minute (healthy)
|
||||
05:44:27 docmost:Up About a minute (healthy) docmost-postgres:Up About a minute (healthy) docmost-redis:Up About a minute (healthy)
|
||||
05:44:52 docmost:Up 2 minutes (healthy) docmost-postgres:Up 2 minutes (healthy) docmost-redis:Up 2 minutes (healthy)
|
||||
05:45:17 docmost:Up 2 minutes (healthy) docmost-postgres:Up 2 minutes (healthy) docmost-redis:Up 2 minutes (healthy)
|
||||
05:45:42 docmost:Up 3 minutes (healthy) docmost-postgres:Up 3 minutes (healthy) docmost-redis:Up 3 minutes (healthy)
|
||||
05:46:07 docmost:Up 3 minutes (healthy) docmost-postgres:Up 3 minutes (healthy) docmost-redis:Up 3 minutes (healthy)
|
||||
05:46:32 docmost:Up 3 minutes (healthy) docmost-postgres:Up 4 minutes (healthy) docmost-redis:Up 4 minutes (healthy)
|
||||
+8
@@ -0,0 +1,8 @@
|
||||
-- SCENARIO D VERDICT: alarms across the whole cycle (T0=05:42:15Z) --
|
||||
app_start_failed events since T0 : 0
|
||||
deadapp scans in the window : 9
|
||||
supervised-down path entered : 1
|
||||
|
||||
--- any docmost down-state reading? ---
|
||||
2026/08/23 05:42:34 manager.go:703: [DEBUG] [stacks] restart-policy of down member "docmost" = "unless-stopped"
|
||||
(no lines above = none)
|
||||
+11
@@ -0,0 +1,11 @@
|
||||
=== §4 MEASUREMENT — a SINGLE-container app crashes out of band ===
|
||||
POSITIVE CONTROL, established on this box today: app_start_failed fired at 05:30:14Z for
|
||||
BookStack (see live-09). The detector demonstrably works here and now, so an absence below
|
||||
is a real absence, not a dead detector.
|
||||
|
||||
subject: privatebin (1 container, restart policy below). Stopped OUT OF BAND — the controller
|
||||
did not do it, so no Deploying flag and no quiesce key.
|
||||
privatebin restart=unless-stopped
|
||||
T0=2026-08-23T05:47:35Z
|
||||
2026-08-23T05:47:40Z
|
||||
privatebin Exited (0) 5 seconds ago
|
||||
+12
@@ -0,0 +1,12 @@
|
||||
-- §4: waiting 4 minutes, well past every grace window (deadapp scan 30s, quiesce grace 180s) --
|
||||
2026-08-23T05:51:53Z
|
||||
privatebin Exited (0) 4 minutes ago
|
||||
|
||||
-- the aggregate state the controller reads --
|
||||
privatebin state=stopped deployed=True
|
||||
|
||||
-- DID ANY ALARM FIRE? --
|
||||
app_start_failed since T0 : 0
|
||||
deadapp scans since T0 : 9
|
||||
-- banner? --
|
||||
banner lines: 0
|
||||
+2139
File diff suppressed because it is too large
Load Diff
+18
@@ -0,0 +1,18 @@
|
||||
### RED-PROOF R-383 — mutation: undoCopyPhrase reverted to the unconditional pre-fix claim ###
|
||||
--- FAIL: TestR383_AbsentUndoCopyIsNotClaimedToExist (0.00s)
|
||||
--- FAIL: TestR383_AbsentUndoCopyIsNotClaimedToExist/absent_—_must_NOT_claim_it_exists,_must_still_name_where_it_should_be (0.00s)
|
||||
r383_undo_phrase_test.go:92: phrase "a korábbi állapot mentése megvan: pre-restore-20260823T120000Z-app-mariadb.sql" does not contain "NEM találjuk"
|
||||
r383_undo_phrase_test.go:97: phrase "a korábbi állapot mentése megvan: pre-restore-20260823T120000Z-app-mariadb.sql" contains "mentése megvan" — it asserts a file that is not on disk
|
||||
--- FAIL: TestR383_AbsentUndoCopyIsNotClaimedToExist/zero-length_—_counts_as_missing (0.00s)
|
||||
r383_undo_phrase_test.go:92: phrase "a korábbi állapot mentése megvan: pre-restore-20260823T120000Z-app-empty.sql" does not contain "NEM találjuk"
|
||||
r383_undo_phrase_test.go:97: phrase "a korábbi állapot mentése megvan: pre-restore-20260823T120000Z-app-empty.sql" contains "mentése megvan" — it asserts a file that is not on disk
|
||||
--- FAIL: TestR383_AbsentUndoCopyIsNotClaimedToExist/partial_—_both_halves_named,_neither_hidden (0.00s)
|
||||
r383_undo_phrase_test.go:92: phrase "a korábbi állapot mentése megvan: pre-restore-20260823T120000Z-app-postgres.sql" does not contain "RÉSZBEN"
|
||||
r383_undo_phrase_test.go:92: phrase "a korábbi állapot mentése megvan: pre-restore-20260823T120000Z-app-postgres.sql" does not contain "HIÁNYZIK"
|
||||
r383_undo_phrase_test.go:92: phrase "a korábbi állapot mentése megvan: pre-restore-20260823T120000Z-app-postgres.sql" does not contain "pre-restore-20260823T120000Z-app-mariadb.sql"
|
||||
--- FAIL: TestR383_AbsentUndoCopyIsNotClaimedToExist/no_undo_was_ever_written_—_said_plainly,_not_silently (0.00s)
|
||||
r383_undo_phrase_test.go:92: phrase "a korábbi állapot mentése megvan: ." does not contain "nem készült"
|
||||
r383_undo_phrase_test.go:97: phrase "a korábbi állapot mentése megvan: ." contains "megvan" — it asserts a file that is not on disk
|
||||
FAIL
|
||||
FAIL gitea.dooplex.hu/admin/felhom-controller/internal/backup 0.008s
|
||||
FAIL
|
||||
+9
@@ -0,0 +1,9 @@
|
||||
### RED-PROOF R-384 #1 — the ORDER (mutation: hoist moved back below `unhealthy > 0`) ###
|
||||
--- FAIL: TestR384_DeadSupervisedMemberIsAskedAboutFirst (0.00s)
|
||||
--- FAIL: TestR384_DeadSupervisedMemberIsAskedAboutFirst/unhealthy_survivor_—_the_bookstack_case_measured_live (0.00s)
|
||||
degraded_test.go:165: survivor "unhealthy" beside a dead SUPERVISED member: aggregateState = "unhealthy", want "degraded" — a dead database must not hide behind it
|
||||
--- FAIL: TestR384_WiresTheDeadDatabaseThroughTheRealPath (0.00s)
|
||||
degraded_test.go:458: bookstack state = "unhealthy", want "degraded" — a dead database must not hide behind its own unhealthy front end (measured live 2026-08-22: this read "unhealthy" and nothing alarmed)
|
||||
FAIL
|
||||
FAIL gitea.dooplex.hu/admin/felhom-controller/internal/stacks 0.007s
|
||||
FAIL
|
||||
+13
@@ -0,0 +1,13 @@
|
||||
### RED-PROOF R-384 #2 — the GUARD (mutation: up = running only, the old `running > 0`) ###
|
||||
--- FAIL: TestR384_DeadSupervisedMemberIsAskedAboutFirst (0.00s)
|
||||
--- FAIL: TestR384_DeadSupervisedMemberIsAskedAboutFirst/unhealthy_survivor_—_the_bookstack_case_measured_live (0.00s)
|
||||
degraded_test.go:165: survivor "unhealthy" beside a dead SUPERVISED member: aggregateState = "unhealthy", want "degraded" — a dead database must not hide behind it
|
||||
--- FAIL: TestR384_DeadSupervisedMemberIsAskedAboutFirst/starting_survivor (0.00s)
|
||||
degraded_test.go:165: survivor "starting" beside a dead SUPERVISED member: aggregateState = "starting", want "degraded" — a dead database must not hide behind it
|
||||
--- FAIL: TestR384_DeadSupervisedMemberIsAskedAboutFirst/restarting_survivor (0.00s)
|
||||
degraded_test.go:165: survivor "restarting" beside a dead SUPERVISED member: aggregateState = "restarting", want "degraded" — a dead database must not hide behind it
|
||||
--- FAIL: TestR384_WiresTheDeadDatabaseThroughTheRealPath (0.00s)
|
||||
degraded_test.go:458: bookstack state = "unhealthy", want "degraded" — a dead database must not hide behind its own unhealthy front end (measured live 2026-08-22: this read "unhealthy" and nothing alarmed)
|
||||
FAIL
|
||||
FAIL gitea.dooplex.hu/admin/felhom-controller/internal/stacks 0.007s
|
||||
FAIL
|
||||
+6
@@ -0,0 +1,6 @@
|
||||
### RED-PROOF R-384 #3 — the CONSEQUENCE layer (mutation: classifier ignores StateDegraded) ###
|
||||
--- FAIL: TestR384_ADeadDatabaseBehindAnUnhealthyAppAlarms (0.00s)
|
||||
r384_dead_db_alarm_test.go:29: dead-app banner = [], want exactly one entry for bookstack
|
||||
FAIL
|
||||
FAIL gitea.dooplex.hu/admin/felhom-controller/cmd/controller 0.008s
|
||||
FAIL
|
||||
Reference in New Issue
Block a user