Files
felhom.eu/documentation/audits/the-28-2026-09-22/BODY-r630.md
T
admin 186546d562
gates / gates (push) Successful in 27s
THE TWENTY-EIGHT: every app no drill had touched, walked in one night
All 28 walked on scratch guest 9202 against the private drill catalog. 26 deployed, 6 proven,
5 inconclusive, 14 with no upstream edge, 1 failed honestly (outline 1.9.1->1.10.1, HELD with the
right sentence), 2 undeployable - one (plant-it) by design, refused by the lifecycle gate, proven
live for the first time. Each app also got the half the update night skipped: a restore from its own
copy with the seed read back again - 21 restored, 2 correctly REFUSED per 07 6.2.

R-630 RAISED TO P1 by measurement: a stack with NO probe container does not skip verifying - it
waits out the full health timeout and HOLDS, stopping an app whose three containers read healthy.
The controller's own words: "not healthy within 5m0s (last: no probe container)".

R-633 opened: a remove sent during a restore reports success and leaves a container restarting with
a live public route. The product already refuses that clash for update and for restore, naming the
blocker; remove has no such guard.

R-634 opened: an app can be running, healthy and serving while recorded as deployed=false, and is
then unremovable. Reproducible alone on sparkyfitness; concurrency-linked on two others.

R-631 and R-632 CLOSED. Register 321 -> 323. Seven interventions, six of them my own harness -
named, with what each cost. No product code. The live catalog's image: lines are byte-identical to
the start of the night.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-22 16:00:56 +02:00

1.9 KiB
Raw Blame History


5.1 — what the guarded Update does when NO probe exists (R-630)

paperless-ngx has no container whose name equals or begins with its stack name, so findProbeContainer returns "" and RunHealthProbes skips the stack silently. Its probe has never run on any box. The open question was what verifying — which waits on that same probe — does when there is nothing to wait on: pass at once, wait out the timeout, or hold.

It waits out the full timeout and then HOLDS, stopping a working app.

Deployed on 9202, all three containers reported healthy, the controller read running, the front door answered 302. No upstream edge exists for paperless-ngx tonight, so the Update was pressed on the same version — which is what a household does on an up-to-date app, and it still walks the whole phase machine. That difference is stated, not glossed.

phase at
checking → safety-dump → pinning → pulling 0.0–1.1 s
starting +2.1 s
verifying +3.1 s
failed +313.0 s — the app STOPPED

Afterwards: controller state stopped, front door 404.

The controller names the cause itself, so no inference was needed:

update paperless-ngx FAILED after the new version was started: not healthy: not healthy within 5m0s (last: no probe container) — stopping and HOLDING the app; the pin stays on the new version (its migration may have run)

no probe container. And the hold sentence is correct about the route back — for this class-A app it warns „csak a beállításokat és az adatbázist tartalmazza, a fájlokat nem".

This is R-618's outcome reached by the opposite road. There a probe named a port the app does not answer; here no probe exists at all — and the static gate cannot see it, because there is nothing to compare. The gate does print it as a WARNING on every push, which is how it was found. R-630 is raised P2 → P1.