R-330: live validation on demo-hp — 3 scans inside the window, 0 events
gates / gates (push) Successful in 12s

v0.224.0 deployed to both demo boxes (both `0.224.0 … (healthy)`). Proven
through POST /api/backup/run, the endpoint the UI button invokes: 8 stacks
stopped and restarted over 87s, three dead-app scans ran INSIDE that window
(16:09:26 docmost, 16:09:56 paperless-ngx, 16:10:26 romm -- the same three apps
that alarmed the night before on 0.223.0), zero app_start_failed pushed.

The scan count is the positive control, not decoration: an absent alarm is
equally consistent with "suppressed correctly" and "the scanner stopped".

A first run is discarded IN THE REPORT rather than quietly dropped -- it fired
52s after a controller restart, inside deadAppBootGrace (90s), where the scan
returns early and could not have alarmed whatever the code did. demo-felhom is
deployed but NOT independently proven and says so: its single app cycles in ~1s,
too fast for any 30s scan to land inside.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LB8FmJaGd2cyjvy6dbEjpM
This commit is contained in:
2026-08-30 18:12:26 +02:00
parent 92cebb8c95
commit 45b52b6ed5
2 changed files with 60 additions and 5 deletions
+13
View File
@@ -81,6 +81,19 @@ directly would have passed against the shipped defect. `cmd/controller/r330_back
walks main.go's AST rather than matching a string, because a commented-out call satisfies
`strings.Contains` — a sibling test in that package records paying for exactly that.
### Proven live on `demo-hp`
`POST /api/backup/run` (the endpoint the "Mentés indítása" button invokes), 8 stacks stopped and
restarted over 87 s, **3 dead-app scans ran inside that window** — 16:09:26 on `docmost`, 16:09:56 on
`paperless-ngx`, 16:10:26 on `romm`, the same three apps that alarmed the night before — and **zero**
`app_start_failed` events were pushed. The scan count is the positive control: an absent alarm is
equally consistent with "suppressed correctly" and "the scanner stopped".
A first attempt was **discarded and said so**: it fired 52 s after a controller restart, inside
`deadAppBootGrace` (90 s), where the scan returns early and could not have alarmed whatever the code
did. `demo-felhom` is deployed but not independently proven — its single app cycles in ~1 s, too fast
for any 30 s scan to land inside.
## v0.223.0 — the alarm we had just built reached nobody, and the stop nobody heard (2026-08-23, R-329 + R-386)
**MinAgent: 0.129.0** (unchanged — no new agent coupling)
+47 -5
View File
@@ -141,10 +141,52 @@ against the shipped defect.
**Green gate:** `go build ./... && go vet ./... && go test ./...` in `felhom-controller/controller/` —
clean, no failures.
## 6. Not done, and why
## 6. Live validation — v0.224.0 deployed and PROVEN on `demo-hp`
- **No live deploy yet.** The build/deploy of 0.224.0 to the two boxes is the next step and is
reported separately; this report covers the change and its unit-land proof only.
Built, pushed and deployed to both boxes; both report `0.224.0 … (healthy)`.
**Method:** `POST /api/backup/run` — the exact endpoint the "Mentés indítása" button invokes — driven
headlessly from inside guest 9201 (no browser on DooPlex; the residual is client-side rendering only).
No state was hand-set: the real `RunDBDumps` ran and really stopped and started all eight stacks.
**The first run is discarded and the reason is recorded rather than quietly dropped.** The controller
had restarted at 16:00:25 UTC and the trigger landed at 16:01:17 — inside `deadAppBootGrace` (90 s),
during which the scan returns early. Half that window could not have alarmed whatever the code did, so
it proves nothing. A second run was taken at **16:08:56**, 8.5 minutes past the grace.
**Run 2 — the numbers (demo-hp, UTC):**
| | |
|---|---|
| stacks stopped and restarted | 8 (bookstack, calibre-web, docmost, kimai, opengist, paperless-ngx, privatebin, romm) |
| window | 16:08:59 → 16:10:26 (87 s) |
| **dead-app scans that ran INSIDE the window** | **3** — 16:09:26, 16:09:56, 16:10:26 |
| `Event pushed: app_start_failed` | **0** (0 across the whole uptime) |
**The positive control is the point** (standing rule 3 — an absent alarm is equally consistent with
"suppressed correctly" and "the scanner stopped"). The scanner was demonstrably alive and evaluating
throughout, and each of the three scans landed on an app that was actually down or mid-restart:
| scan | app in its stop/start window at that moment | alarmed on 0.223.0 last night? |
|---|---|---|
| 16:09:26 | `docmost` (down 16:09:16 → 16:09:31) | **yes** |
| 16:09:56 | `paperless-ngx` (down 16:09:48 → 16:10:08) | **yes** |
| 16:10:26 | `romm` (restarting, back at 16:10:26) | **yes** |
That is a true A/B on the same box, the same job and the **same three apps** that produced last night's
e-mails: 0.223.0 → 3 events, 0.224.0 → 0 events, with the scanner proven running in both.
**`demo-felhom` is deployed but NOT independently proven, and this is stated rather than implied.** Its
single app (`opengist`) cycles in **~1 s** (16:11:07 → 16:11:08), so no 30 s scan could land inside the
window — the run produced zero events, but zero events was the expected result either way. Its scanner
is confirmed alive (`[deadapp] check alive: 20 scans since boot, 1 deployed app(s) evaluated, 0 currently
down`). The code path is identical to the one proven on demo-hp; the narrower race is also why that box
sent 2 mails a night rather than 5.
**The real acceptance test is tonight's unattended 02:30 and 04:15 CEST runs.** Zero
`app_start_failed` mails from either box tomorrow morning closes this; any mail is a regression.
## 7. Not done, and why
- **The `restore-hold` path** (`offbox_reconstitute.go`, an app deliberately held down after a failed
replay) calls `End()`, so it gets the 180 s grace and then alarms. That is **today's behaviour plus
180 s** and is deliberate: the app really is down, the customer should learn that, and the hold has
@@ -154,7 +196,7 @@ clean, no failures.
needed. The new in-memory set is a superset for alarm purposes; the durable one stays the recovery
record.
## 7. Found while diagnosing — a second, still-open defect (Option B)
## 8. Found while diagnosing — a second, still-open defect (Option B)
**The hub's customer Backup card is inert for every customer.** It reads
`Snapshots 0 · Repo Size 0 MB · Integrity Unknown` while the same box's log says
@@ -170,7 +212,7 @@ never copied into the hub report.
**Until it is fixed, that card must not be read as evidence of a missing backup.** It is a two-repo
change (controller report builder + hub) and is Option B of this request.
## 8. Also observed (not changed)
## 9. Also observed (not changed)
- **`ssh demo-hp` no longer works** — the tailnet peer `100.76.96.79` has been offline 8 days
(`tailscale status`: `offline, last seen 8d ago`). The box is reachable on the home LAN as