Files
felhom.eu/documentation/audits/update-night-2026-09-21/14-tandoor-serving-while-verifying.txt
T
admin 9c69b3ff07
gates / gates (push) Successful in 27s
Update night: Phases 2-4 evidence — both engines, the unattended HOLD, and five new findings
Evidence off the machine at the end of the phases that produced it (R-320). Teardown follows.

PHASE 2 — the two database engines, through the REAL Update button:
- MariaDB 11.6 -> 12.3 on nextcloud: PROVEN, and pressed through the button for the first time.
  All four SPIKE-r459 observables: the datadir's own record moved 11.6.2 -> 12.3.3; the engine
  itself says "already upgraded ... no need to run mariadb-upgrade again"; the entrypoint says
  "Major version upgrade detected ... Check required!" and then STARTED and FINISHED it (not the
  `skipped due to $MARIADB_AUTO_UPGRADE` line R-459 feared); and the engine took its own
  pre-upgrade backup, 631 905 B. The seeded Nextcloud account read back.
- PostgreSQL 16 -> 17 on docmost: FAILED exactly as R-463 predicted and nobody had measured.
  5.1 s to held; the pin named 17 while nothing ran; the restore brought it back in 29.1 s.
  The engine's REFUSAL LINE was destroyed by failAndHold before any probe could read it, so it
  was REPRODUCED INDEPENDENTLY with a control on every step (R-320).

PHASE 3 — the bad days. B1 produced THE UNATTENDED HOLD, which this project has never had: the
caller pressed once with nobody watching, the app held after 312.9 s, and passes 2 and 3 pressed
nothing. B2 put the pin back on a pull failure in 1.0 s. B3 refused `busy` six times. B4 showed
there is NO single-flight — 5 of 5 updates ran at once and all ended honest. B5 cut the power in
`backing-up` and the box recovered itself and said so. B7 refused under the 2 GB floor. B9 found
R-458's risk narrower than the row states.

PHASE 4 — every badge on the box is TRUE, and the held app answers all four of Q4's questions.

FINDINGS, five new and three corrections to existing rows. The one that matters: R-618 is P1 —
two templates name a health probe the app does not answer, and because the guarded update waits
on that same probe, a SUCCESSFUL update ends by STOPPING a working app. Measured: tandoor served
HTTP 200 on the new version at four samples across five minutes and was then stopped.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-21 22:13:57 +02:00

64 lines
3.5 KiB
Plaintext

R-618 ESCALATION — measured live, 2026-09-21, guest 9202, controller v0.261.0
=============================================================================
tandoor's guarded Update from 2.6.13 to 2.6.15 was pressed at 21:16:47 and entered `verifying` at
21:17:46. While it sat there, the app was asked the same question the household would ask, through
the household's own front door:
21:18:28 curl -H 'Host: recipes.enkisfelhom.hu' https://192.168.0.114/accounts/login/
-> HTTP 200
21:18:28 docker ps: tandoor ghcr.io/tandoorrecipes/recipes:2.6.15 Up 25 seconds
So the NEW version was installed and SERVING. The controller nevertheless could not see it, because
`.felhom.yml` probes `type: http, port: 8080` and the tandoor container listens on **80 only**.
The consequence is the part that matters, and it is bigger than a wrong label:
* `verifying` cannot pass, so the full `update.health_timeout` (5 minutes) is spent.
* `Manager.failAndHold` then runs `compose down` and the app is STOPPED.
* The household is told the update failed and is sent to a restore they do not need.
* Their recipe app — which was working, on the new version, with its data intact — is DOWN.
Nothing is lost: the data is in the volumes and the restore works. But a single wrong port number
in a template turns a SUCCESSFUL update into an outage plus an unnecessary restore, every time,
for every household running that app.
The same shape applies to `zipline` (probe `type: api` + `expect: status 200` on `/api/health`,
which that app answers 404) and, unmeasured, to `wger`.
A SECOND CHECK, five minutes later in the same `verifying` window
-----------------------------------------------------------------
21:19:52 front door=200 docker=healthy
21:20:18 front door=200 docker=healthy
21:20:45 front door=200 docker=healthy
THE END OF IT, 2026-09-21 21:22:49 (+361.9 s)
---------------------------------------------
phase=failed
„A(z) tandoor frissítése 2026-09-21 21:22-kor nem sikerult, es az alkalmazas nem indult el az uj
verzioval. Az alkalmazas biztonsagi okbol leallitva marad, hogy az adatai ne serüljenek.
Visszaallithato a Mentesek oldalon ebbol a biztonsagi mentesbol: sajat meghajto, 2026-09-21 21:15
— ez a masolat a beallitasokat, az adatbazist es az adatkoteteket tartalmazza."
(quoted with ASCII fragments; the live sentence carries its accents — see apps/tandoor/)
docker ps -a : NO tandoor container at all
front door : HTTP 404 (it was 200 four times in the preceding five minutes)
pinned_images : ghcr.io/tandoorrecipes/recipes:2.6.15
installed_images : ghcr.io/tandoorrecipes/recipes:2.6.13 <- DISAGREES with the pin
live compose image: : ghcr.io/tandoorrecipes/recipes:2.6.15
docker inspect : []
The four version observables DISAGREE after a hold, and that is CORRECT — `09` §5.2: the pin is a
DECISION and `installed_images` is an OBSERVATION, and "when they disagree that is a signal, not a
bug to paper over". Here the signal reads exactly right: we decided 2.6.15, the last thing actually
observed running was 2.6.13, and nothing is running now.
TIMELINE, one line
------------------
21:16:47 Update pressed 21:17:46 verifying begins
21:18:28 NEW VERSION SERVING 200, docker healthcheck green
21:19:52 200 / healthy 21:20:18 200 / healthy 21:20:45 200 / healthy
21:22:49 health timeout -> failAndHold -> compose down -> app STOPPED, front door 404