9c69b3ff07
gates / gates (push) Successful in 27s
Evidence off the machine at the end of the phases that produced it (R-320). Teardown follows. PHASE 2 — the two database engines, through the REAL Update button: - MariaDB 11.6 -> 12.3 on nextcloud: PROVEN, and pressed through the button for the first time. All four SPIKE-r459 observables: the datadir's own record moved 11.6.2 -> 12.3.3; the engine itself says "already upgraded ... no need to run mariadb-upgrade again"; the entrypoint says "Major version upgrade detected ... Check required!" and then STARTED and FINISHED it (not the `skipped due to $MARIADB_AUTO_UPGRADE` line R-459 feared); and the engine took its own pre-upgrade backup, 631 905 B. The seeded Nextcloud account read back. - PostgreSQL 16 -> 17 on docmost: FAILED exactly as R-463 predicted and nobody had measured. 5.1 s to held; the pin named 17 while nothing ran; the restore brought it back in 29.1 s. The engine's REFUSAL LINE was destroyed by failAndHold before any probe could read it, so it was REPRODUCED INDEPENDENTLY with a control on every step (R-320). PHASE 3 — the bad days. B1 produced THE UNATTENDED HOLD, which this project has never had: the caller pressed once with nobody watching, the app held after 312.9 s, and passes 2 and 3 pressed nothing. B2 put the pin back on a pull failure in 1.0 s. B3 refused `busy` six times. B4 showed there is NO single-flight — 5 of 5 updates ran at once and all ended honest. B5 cut the power in `backing-up` and the box recovered itself and said so. B7 refused under the 2 GB floor. B9 found R-458's risk narrower than the row states. PHASE 4 — every badge on the box is TRUE, and the held app answers all four of Q4's questions. FINDINGS, five new and three corrections to existing rows. The one that matters: R-618 is P1 — two templates name a health probe the app does not answer, and because the guarded update waits on that same probe, a SUCCESSFUL update ends by STOPPING a working app. Measured: tandoor served HTTP 200 on the new version at four samples across five minutes and was then stopped. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
36 lines
2.1 KiB
Markdown
36 lines
2.1 KiB
Markdown
# The alarm truth table for the update night, scored against `08-alarm-ladder.md`
|
|
|
|
**READ THE LIMIT FIRST, because it bounds every row.** Guest 9202 runs with `hub.enabled: false`.
|
|
Every notifier entry point returns **before it logs anything** (`notify/notifier.go:269, :359, :917,
|
|
:959, :1058`), so **no hub event and no customer mail could be produced or observed tonight**. The
|
|
hub was not enabled on 9202 to get around this: that would register an unclaimed host at the live
|
|
hub and could mail a real address, and the brief fences the hub. Filed as **R-620**.
|
|
|
|
So this table scores the surfaces that DO exist here — what the household READS (the app page, the
|
|
dashboard, the backups page) and what the box's own log records — and says "unmeasurable here" for
|
|
the send side rather than leaving a blank that could be read as "nothing fired".
|
|
|
|
| # | what happened | should it alarm, per `08` | did the box classify it that way | what the household could READ | verdict |
|
|
|---|---|---|---|---|---|
|
|
<!-- ROWS -->
|
|
|
|
## The three things this table is for
|
|
|
|
**1. Which alarm fired and was it true.** <FIRED>
|
|
|
|
**2. Which should have fired and did not.** <MISSING>
|
|
|
|
**3. The one that is structurally invisible on this venue.** Every row's "who was told" column is
|
|
answered only for the SCREEN. `08` §6.1's severity contract and §6.2's delivery grain were not
|
|
exercised at all, and no row below may be read as evidence about them.
|
|
|
|
## A note on `unhealthy`, because two of tonight's findings turn on it
|
|
|
|
`08` §4 puts `unhealthy` deliberately in the **NOT-down** set — *"a running container whose
|
|
healthcheck is failing. Folding it in reintroduces the flapping fix-3 was written to stop."* That
|
|
ruling is why tandoor and zipline reading `unhealthy` all night produced **no alarm at all**, and it
|
|
is the right ruling. **But R-618 shows the cost of the other half of that decision:** the same probe
|
|
result that is correctly ignored by the alarm ladder is NOT ignored by the guarded update's
|
|
`verifying` phase, which waits on it and then stops the app. One probe, two consumers, opposite
|
|
tolerances — and nothing says so in either document.
|