R-330: live validation on demo-hp — 3 scans inside the window, 0 events
gates / gates (push) Successful in 12s
gates / gates (push) Successful in 12s
v0.224.0 deployed to both demo boxes (both `0.224.0 … (healthy)`). Proven through POST /api/backup/run, the endpoint the UI button invokes: 8 stacks stopped and restarted over 87s, three dead-app scans ran INSIDE that window (16:09:26 docmost, 16:09:56 paperless-ngx, 16:10:26 romm -- the same three apps that alarmed the night before on 0.223.0), zero app_start_failed pushed. The scan count is the positive control, not decoration: an absent alarm is equally consistent with "suppressed correctly" and "the scanner stopped". A first run is discarded IN THE REPORT rather than quietly dropped -- it fired 52s after a controller restart, inside deadAppBootGrace (90s), where the scan returns early and could not have alarmed whatever the code did. demo-felhom is deployed but NOT independently proven and says so: its single app cycles in ~1s, too fast for any 30s scan to land inside. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01LB8FmJaGd2cyjvy6dbEjpM
This commit is contained in:
@@ -81,6 +81,19 @@ directly would have passed against the shipped defect. `cmd/controller/r330_back
|
||||
walks main.go's AST rather than matching a string, because a commented-out call satisfies
|
||||
`strings.Contains` — a sibling test in that package records paying for exactly that.
|
||||
|
||||
### Proven live on `demo-hp`
|
||||
|
||||
`POST /api/backup/run` (the endpoint the "Mentés indítása" button invokes), 8 stacks stopped and
|
||||
restarted over 87 s, **3 dead-app scans ran inside that window** — 16:09:26 on `docmost`, 16:09:56 on
|
||||
`paperless-ngx`, 16:10:26 on `romm`, the same three apps that alarmed the night before — and **zero**
|
||||
`app_start_failed` events were pushed. The scan count is the positive control: an absent alarm is
|
||||
equally consistent with "suppressed correctly" and "the scanner stopped".
|
||||
|
||||
A first attempt was **discarded and said so**: it fired 52 s after a controller restart, inside
|
||||
`deadAppBootGrace` (90 s), where the scan returns early and could not have alarmed whatever the code
|
||||
did. `demo-felhom` is deployed but not independently proven — its single app cycles in ~1 s, too fast
|
||||
for any 30 s scan to land inside.
|
||||
|
||||
## v0.223.0 — the alarm we had just built reached nobody, and the stop nobody heard (2026-08-23, R-329 + R-386)
|
||||
**MinAgent: 0.129.0** (unchanged — no new agent coupling)
|
||||
|
||||
|
||||
@@ -141,10 +141,52 @@ against the shipped defect.
|
||||
**Green gate:** `go build ./... && go vet ./... && go test ./...` in `felhom-controller/controller/` —
|
||||
clean, no failures.
|
||||
|
||||
## 6. Not done, and why
|
||||
## 6. Live validation — v0.224.0 deployed and PROVEN on `demo-hp`
|
||||
|
||||
- **No live deploy yet.** The build/deploy of 0.224.0 to the two boxes is the next step and is
|
||||
reported separately; this report covers the change and its unit-land proof only.
|
||||
Built, pushed and deployed to both boxes; both report `0.224.0 … (healthy)`.
|
||||
|
||||
**Method:** `POST /api/backup/run` — the exact endpoint the "Mentés indítása" button invokes — driven
|
||||
headlessly from inside guest 9201 (no browser on DooPlex; the residual is client-side rendering only).
|
||||
No state was hand-set: the real `RunDBDumps` ran and really stopped and started all eight stacks.
|
||||
|
||||
**The first run is discarded and the reason is recorded rather than quietly dropped.** The controller
|
||||
had restarted at 16:00:25 UTC and the trigger landed at 16:01:17 — inside `deadAppBootGrace` (90 s),
|
||||
during which the scan returns early. Half that window could not have alarmed whatever the code did, so
|
||||
it proves nothing. A second run was taken at **16:08:56**, 8.5 minutes past the grace.
|
||||
|
||||
**Run 2 — the numbers (demo-hp, UTC):**
|
||||
|
||||
| | |
|
||||
|---|---|
|
||||
| stacks stopped and restarted | 8 (bookstack, calibre-web, docmost, kimai, opengist, paperless-ngx, privatebin, romm) |
|
||||
| window | 16:08:59 → 16:10:26 (87 s) |
|
||||
| **dead-app scans that ran INSIDE the window** | **3** — 16:09:26, 16:09:56, 16:10:26 |
|
||||
| `Event pushed: app_start_failed` | **0** (0 across the whole uptime) |
|
||||
|
||||
**The positive control is the point** (standing rule 3 — an absent alarm is equally consistent with
|
||||
"suppressed correctly" and "the scanner stopped"). The scanner was demonstrably alive and evaluating
|
||||
throughout, and each of the three scans landed on an app that was actually down or mid-restart:
|
||||
|
||||
| scan | app in its stop/start window at that moment | alarmed on 0.223.0 last night? |
|
||||
|---|---|---|
|
||||
| 16:09:26 | `docmost` (down 16:09:16 → 16:09:31) | **yes** |
|
||||
| 16:09:56 | `paperless-ngx` (down 16:09:48 → 16:10:08) | **yes** |
|
||||
| 16:10:26 | `romm` (restarting, back at 16:10:26) | **yes** |
|
||||
|
||||
That is a true A/B on the same box, the same job and the **same three apps** that produced last night's
|
||||
e-mails: 0.223.0 → 3 events, 0.224.0 → 0 events, with the scanner proven running in both.
|
||||
|
||||
**`demo-felhom` is deployed but NOT independently proven, and this is stated rather than implied.** Its
|
||||
single app (`opengist`) cycles in **~1 s** (16:11:07 → 16:11:08), so no 30 s scan could land inside the
|
||||
window — the run produced zero events, but zero events was the expected result either way. Its scanner
|
||||
is confirmed alive (`[deadapp] check alive: 20 scans since boot, 1 deployed app(s) evaluated, 0 currently
|
||||
down`). The code path is identical to the one proven on demo-hp; the narrower race is also why that box
|
||||
sent 2 mails a night rather than 5.
|
||||
|
||||
**The real acceptance test is tonight's unattended 02:30 and 04:15 CEST runs.** Zero
|
||||
`app_start_failed` mails from either box tomorrow morning closes this; any mail is a regression.
|
||||
|
||||
## 7. Not done, and why
|
||||
- **The `restore-hold` path** (`offbox_reconstitute.go`, an app deliberately held down after a failed
|
||||
replay) calls `End()`, so it gets the 180 s grace and then alarms. That is **today's behaviour plus
|
||||
180 s** and is deliberate: the app really is down, the customer should learn that, and the hold has
|
||||
@@ -154,7 +196,7 @@ clean, no failures.
|
||||
needed. The new in-memory set is a superset for alarm purposes; the durable one stays the recovery
|
||||
record.
|
||||
|
||||
## 7. Found while diagnosing — a second, still-open defect (Option B)
|
||||
## 8. Found while diagnosing — a second, still-open defect (Option B)
|
||||
|
||||
**The hub's customer Backup card is inert for every customer.** It reads
|
||||
`Snapshots 0 · Repo Size 0 MB · Integrity Unknown` while the same box's log says
|
||||
@@ -170,7 +212,7 @@ never copied into the hub report.
|
||||
**Until it is fixed, that card must not be read as evidence of a missing backup.** It is a two-repo
|
||||
change (controller report builder + hub) and is Option B of this request.
|
||||
|
||||
## 8. Also observed (not changed)
|
||||
## 9. Also observed (not changed)
|
||||
|
||||
- **`ssh demo-hp` no longer works** — the tailnet peer `100.76.96.79` has been offline 8 days
|
||||
(`tailscale status`: `offline, last seen 8d ago`). The box is reachable on the home LAN as
|
||||
|
||||
Reference in New Issue
Block a user