9c69b3ff07
gates / gates (push) Successful in 27s
Evidence off the machine at the end of the phases that produced it (R-320). Teardown follows. PHASE 2 — the two database engines, through the REAL Update button: - MariaDB 11.6 -> 12.3 on nextcloud: PROVEN, and pressed through the button for the first time. All four SPIKE-r459 observables: the datadir's own record moved 11.6.2 -> 12.3.3; the engine itself says "already upgraded ... no need to run mariadb-upgrade again"; the entrypoint says "Major version upgrade detected ... Check required!" and then STARTED and FINISHED it (not the `skipped due to $MARIADB_AUTO_UPGRADE` line R-459 feared); and the engine took its own pre-upgrade backup, 631 905 B. The seeded Nextcloud account read back. - PostgreSQL 16 -> 17 on docmost: FAILED exactly as R-463 predicted and nobody had measured. 5.1 s to held; the pin named 17 while nothing ran; the restore brought it back in 29.1 s. The engine's REFUSAL LINE was destroyed by failAndHold before any probe could read it, so it was REPRODUCED INDEPENDENTLY with a control on every step (R-320). PHASE 3 — the bad days. B1 produced THE UNATTENDED HOLD, which this project has never had: the caller pressed once with nobody watching, the app held after 312.9 s, and passes 2 and 3 pressed nothing. B2 put the pin back on a pull failure in 1.0 s. B3 refused `busy` six times. B4 showed there is NO single-flight — 5 of 5 updates ran at once and all ended honest. B5 cut the power in `backing-up` and the box recovered itself and said so. B7 refused under the 2 GB floor. B9 found R-458's risk narrower than the row states. PHASE 4 — every badge on the box is TRUE, and the held app answers all four of Q4's questions. FINDINGS, five new and three corrections to existing rows. The one that matters: R-618 is P1 — two templates name a health probe the app does not answer, and because the guarded update waits on that same probe, a SUCCESSFUL update ends by STOPPING a working app. Measured: tandoor served HTTP 200 on the new version at four samples across five minutes and was then stopped. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
64 lines
3.5 KiB
Plaintext
64 lines
3.5 KiB
Plaintext
R-618 ESCALATION — measured live, 2026-09-21, guest 9202, controller v0.261.0
|
|
=============================================================================
|
|
|
|
tandoor's guarded Update from 2.6.13 to 2.6.15 was pressed at 21:16:47 and entered `verifying` at
|
|
21:17:46. While it sat there, the app was asked the same question the household would ask, through
|
|
the household's own front door:
|
|
|
|
21:18:28 curl -H 'Host: recipes.enkisfelhom.hu' https://192.168.0.114/accounts/login/
|
|
-> HTTP 200
|
|
|
|
21:18:28 docker ps: tandoor ghcr.io/tandoorrecipes/recipes:2.6.15 Up 25 seconds
|
|
|
|
So the NEW version was installed and SERVING. The controller nevertheless could not see it, because
|
|
`.felhom.yml` probes `type: http, port: 8080` and the tandoor container listens on **80 only**.
|
|
|
|
The consequence is the part that matters, and it is bigger than a wrong label:
|
|
|
|
* `verifying` cannot pass, so the full `update.health_timeout` (5 minutes) is spent.
|
|
* `Manager.failAndHold` then runs `compose down` and the app is STOPPED.
|
|
* The household is told the update failed and is sent to a restore they do not need.
|
|
* Their recipe app — which was working, on the new version, with its data intact — is DOWN.
|
|
|
|
Nothing is lost: the data is in the volumes and the restore works. But a single wrong port number
|
|
in a template turns a SUCCESSFUL update into an outage plus an unnecessary restore, every time,
|
|
for every household running that app.
|
|
|
|
The same shape applies to `zipline` (probe `type: api` + `expect: status 200` on `/api/health`,
|
|
which that app answers 404) and, unmeasured, to `wger`.
|
|
|
|
|
|
A SECOND CHECK, five minutes later in the same `verifying` window
|
|
-----------------------------------------------------------------
|
|
21:19:52 front door=200 docker=healthy
|
|
21:20:18 front door=200 docker=healthy
|
|
21:20:45 front door=200 docker=healthy
|
|
|
|
THE END OF IT, 2026-09-21 21:22:49 (+361.9 s)
|
|
---------------------------------------------
|
|
phase=failed
|
|
„A(z) tandoor frissítése 2026-09-21 21:22-kor nem sikerult, es az alkalmazas nem indult el az uj
|
|
verzioval. Az alkalmazas biztonsagi okbol leallitva marad, hogy az adatai ne serüljenek.
|
|
Visszaallithato a Mentesek oldalon ebbol a biztonsagi mentesbol: sajat meghajto, 2026-09-21 21:15
|
|
— ez a masolat a beallitasokat, az adatbazist es az adatkoteteket tartalmazza."
|
|
(quoted with ASCII fragments; the live sentence carries its accents — see apps/tandoor/)
|
|
|
|
docker ps -a : NO tandoor container at all
|
|
front door : HTTP 404 (it was 200 four times in the preceding five minutes)
|
|
pinned_images : ghcr.io/tandoorrecipes/recipes:2.6.15
|
|
installed_images : ghcr.io/tandoorrecipes/recipes:2.6.13 <- DISAGREES with the pin
|
|
live compose image: : ghcr.io/tandoorrecipes/recipes:2.6.15
|
|
docker inspect : []
|
|
|
|
The four version observables DISAGREE after a hold, and that is CORRECT — `09` §5.2: the pin is a
|
|
DECISION and `installed_images` is an OBSERVATION, and "when they disagree that is a signal, not a
|
|
bug to paper over". Here the signal reads exactly right: we decided 2.6.15, the last thing actually
|
|
observed running was 2.6.13, and nothing is running now.
|
|
|
|
TIMELINE, one line
|
|
------------------
|
|
21:16:47 Update pressed 21:17:46 verifying begins
|
|
21:18:28 NEW VERSION SERVING 200, docker healthcheck green
|
|
21:19:52 200 / healthy 21:20:18 200 / healthy 21:20:45 200 / healthy
|
|
21:22:49 health timeout -> failAndHold -> compose down -> app STOPPED, front door 404
|