From 45b52b6ed50b81098d86f5fcda771112d635f8d7 Mon Sep 17 00:00:00 2001 From: kisfenyo Date: Sun, 30 Aug 2026 18:12:26 +0200 Subject: [PATCH] =?UTF-8?q?R-330:=20live=20validation=20on=20demo-hp=20?= =?UTF-8?q?=E2=80=94=203=20scans=20inside=20the=20window,=200=20events?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit v0.224.0 deployed to both demo boxes (both `0.224.0 … (healthy)`). Proven through POST /api/backup/run, the endpoint the UI button invokes: 8 stacks stopped and restarted over 87s, three dead-app scans ran INSIDE that window (16:09:26 docmost, 16:09:56 paperless-ngx, 16:10:26 romm -- the same three apps that alarmed the night before on 0.223.0), zero app_start_failed pushed. The scan count is the positive control, not decoration: an absent alarm is equally consistent with "suppressed correctly" and "the scanner stopped". A first run is discarded IN THE REPORT rather than quietly dropped -- it fired 52s after a controller restart, inside deadAppBootGrace (90s), where the scan returns early and could not have alarmed whatever the code did. demo-felhom is deployed but NOT independently proven and says so: its single app cycles in ~1s, too fast for any 30s scan to land inside. Co-Authored-By: Claude Opus 5 (1M context) Claude-Session: https://claude.ai/code/session_01LB8FmJaGd2cyjvy6dbEjpM --- CHANGELOG.md | 13 +++++++++++++ REPORT.md | 52 +++++++++++++++++++++++++++++++++++++++++++++++----- 2 files changed, 60 insertions(+), 5 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index f5b1c96..9a97682 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -81,6 +81,19 @@ directly would have passed against the shipped defect. `cmd/controller/r330_back walks main.go's AST rather than matching a string, because a commented-out call satisfies `strings.Contains` — a sibling test in that package records paying for exactly that. +### Proven live on `demo-hp` + +`POST /api/backup/run` (the endpoint the "Mentés indítása" button invokes), 8 stacks stopped and +restarted over 87 s, **3 dead-app scans ran inside that window** — 16:09:26 on `docmost`, 16:09:56 on +`paperless-ngx`, 16:10:26 on `romm`, the same three apps that alarmed the night before — and **zero** +`app_start_failed` events were pushed. The scan count is the positive control: an absent alarm is +equally consistent with "suppressed correctly" and "the scanner stopped". + +A first attempt was **discarded and said so**: it fired 52 s after a controller restart, inside +`deadAppBootGrace` (90 s), where the scan returns early and could not have alarmed whatever the code +did. `demo-felhom` is deployed but not independently proven — its single app cycles in ~1 s, too fast +for any 30 s scan to land inside. + ## v0.223.0 — the alarm we had just built reached nobody, and the stop nobody heard (2026-08-23, R-329 + R-386) **MinAgent: 0.129.0** (unchanged — no new agent coupling) diff --git a/REPORT.md b/REPORT.md index 7bbb2cc..6199a9f 100644 --- a/REPORT.md +++ b/REPORT.md @@ -141,10 +141,52 @@ against the shipped defect. **Green gate:** `go build ./... && go vet ./... && go test ./...` in `felhom-controller/controller/` — clean, no failures. -## 6. Not done, and why +## 6. Live validation — v0.224.0 deployed and PROVEN on `demo-hp` -- **No live deploy yet.** The build/deploy of 0.224.0 to the two boxes is the next step and is - reported separately; this report covers the change and its unit-land proof only. +Built, pushed and deployed to both boxes; both report `0.224.0 … (healthy)`. + +**Method:** `POST /api/backup/run` — the exact endpoint the "Mentés indítása" button invokes — driven +headlessly from inside guest 9201 (no browser on DooPlex; the residual is client-side rendering only). +No state was hand-set: the real `RunDBDumps` ran and really stopped and started all eight stacks. + +**The first run is discarded and the reason is recorded rather than quietly dropped.** The controller +had restarted at 16:00:25 UTC and the trigger landed at 16:01:17 — inside `deadAppBootGrace` (90 s), +during which the scan returns early. Half that window could not have alarmed whatever the code did, so +it proves nothing. A second run was taken at **16:08:56**, 8.5 minutes past the grace. + +**Run 2 — the numbers (demo-hp, UTC):** + +| | | +|---|---| +| stacks stopped and restarted | 8 (bookstack, calibre-web, docmost, kimai, opengist, paperless-ngx, privatebin, romm) | +| window | 16:08:59 → 16:10:26 (87 s) | +| **dead-app scans that ran INSIDE the window** | **3** — 16:09:26, 16:09:56, 16:10:26 | +| `Event pushed: app_start_failed` | **0** (0 across the whole uptime) | + +**The positive control is the point** (standing rule 3 — an absent alarm is equally consistent with +"suppressed correctly" and "the scanner stopped"). The scanner was demonstrably alive and evaluating +throughout, and each of the three scans landed on an app that was actually down or mid-restart: + +| scan | app in its stop/start window at that moment | alarmed on 0.223.0 last night? | +|---|---|---| +| 16:09:26 | `docmost` (down 16:09:16 → 16:09:31) | **yes** | +| 16:09:56 | `paperless-ngx` (down 16:09:48 → 16:10:08) | **yes** | +| 16:10:26 | `romm` (restarting, back at 16:10:26) | **yes** | + +That is a true A/B on the same box, the same job and the **same three apps** that produced last night's +e-mails: 0.223.0 → 3 events, 0.224.0 → 0 events, with the scanner proven running in both. + +**`demo-felhom` is deployed but NOT independently proven, and this is stated rather than implied.** Its +single app (`opengist`) cycles in **~1 s** (16:11:07 → 16:11:08), so no 30 s scan could land inside the +window — the run produced zero events, but zero events was the expected result either way. Its scanner +is confirmed alive (`[deadapp] check alive: 20 scans since boot, 1 deployed app(s) evaluated, 0 currently +down`). The code path is identical to the one proven on demo-hp; the narrower race is also why that box +sent 2 mails a night rather than 5. + +**The real acceptance test is tonight's unattended 02:30 and 04:15 CEST runs.** Zero +`app_start_failed` mails from either box tomorrow morning closes this; any mail is a regression. + +## 7. Not done, and why - **The `restore-hold` path** (`offbox_reconstitute.go`, an app deliberately held down after a failed replay) calls `End()`, so it gets the 180 s grace and then alarms. That is **today's behaviour plus 180 s** and is deliberate: the app really is down, the customer should learn that, and the hold has @@ -154,7 +196,7 @@ clean, no failures. needed. The new in-memory set is a superset for alarm purposes; the durable one stays the recovery record. -## 7. Found while diagnosing — a second, still-open defect (Option B) +## 8. Found while diagnosing — a second, still-open defect (Option B) **The hub's customer Backup card is inert for every customer.** It reads `Snapshots 0 · Repo Size 0 MB · Integrity Unknown` while the same box's log says @@ -170,7 +212,7 @@ never copied into the hub report. **Until it is fixed, that card must not be read as evidence of a missing backup.** It is a two-repo change (controller report builder + hub) and is Option B of this request. -## 8. Also observed (not changed) +## 9. Also observed (not changed) - **`ssh demo-hp` no longer works** — the tailnet peer `100.76.96.79` has been offline 8 days (`tailscale status`: `offline, last seen 8d ago`). The box is reachable on the home LAN as