All 28 walked on scratch guest 9202 against the private drill catalog. 26 deployed, 6 proven, 5 inconclusive, 14 with no upstream edge, 1 failed honestly (outline 1.9.1->1.10.1, HELD with the right sentence), 2 undeployable - one (plant-it) by design, refused by the lifecycle gate, proven live for the first time. Each app also got the half the update night skipped: a restore from its own copy with the seed read back again - 21 restored, 2 correctly REFUSED per 07 6.2. R-630 RAISED TO P1 by measurement: a stack with NO probe container does not skip verifying - it waits out the full health timeout and HOLDS, stopping an app whose three containers read healthy. The controller's own words: "not healthy within 5m0s (last: no probe container)". R-633 opened: a remove sent during a restore reports success and leaves a container restarting with a live public route. The product already refuses that clash for update and for restore, naming the blocker; remove has no such guard. R-634 opened: an app can be running, healthy and serving while recorded as deployed=false, and is then unremovable. Reproducible alone on sparkyfitness; concurrency-linked on two others. R-631 and R-632 CLOSED. Register 321 -> 323. Seven interventions, six of them my own harness - named, with what each cost. No product code. The live catalog's image: lines are byte-identical to the start of the night. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
5.5 KiB
What this night is, in three lines
- Interventions: SEVEN — over the brief's limit of five, and six of the seven were my own harness,
not the product. Three were bugs in tonight's driver that cost apps their walk and were fixed
mid-run (a missing
import re; a variable that shadowed the app's metadata; a crash when the Update was refused before any phase existed). Two were deliberate method changes (recording a REFUSED restore as its own verdict; making the harness wait for a restore to settle before removing). One was a waiter that deadlocked on its own command line. The seventh was the product's: three leftovers it could not clear, which a shell had to. - 26 of 28 deployed; 6 proven; 5 inconclusive; 14 with no upstream edge; 1 failed honestly; 2 that could not be deployed — one of those by design.
- The one result that matters most: an app with no health probe at all has its working
installation stopped by a successful update.
paperless-ngxwas healthy on all three containers; the Update ran the full five-minute health wait and then held the app, and the controller named the reason itself —no probe container. R-630 is raised to P1.
The two findings that are not about any single app
R-633 — a remove sent while a restore is still running reports success and leaves an orphan
gokapi was restored from its own local copy at 11:34:07 and removed at 11:34:22.
POST /backup/restore answers 302 and does its work in the background; the remove tore down
what existed, and the restore's own compose up then re-created the container at 11:34:24.
Both calls returned success.
Twenty-five minutes later, GET /api/stacks/gokapi reads deployed: false while docker ps -a
shows gokapi Restarting (1) with its full Traefik label set still attached — including
traefik.http.routers.gokapi.rule: Host(.enkisfelhom.hu), a rule with an empty subdomain,
because the deploy values that filled it were deleted with the app. Its own log loops
„Salt for admin password invalid… password does not appear to be a SHA-1 hash" — the volume
holding its config was removed correctly, so the binary can never start.
A household can press exactly those two buttons in that order. The product accepted both and
the remove reported success while leaving the orphan; nothing in the alarm ladder can fire,
because 08 §4 keys on stacks the controller still knows about. This is R-626's class with the
mechanism finally visible — that row saw a removed navidrome come back and could not diagnose
it, because the controller had restarted and its log no longer reached the moment. Here the window
is seventeen seconds and both halves are in the evidence.
The harness was then fenced against its own race, so every app after gokapi measures the product.
Evidence: apps/gokapi/came-back-evidence.txt.
A restore that is REFUSED is the product being right, and nearly went down as a failure
calibre-web is a class-A app (07 §6.2): it has a readable file leg, and the local Tier-1 copy
does not hold it. The restore was refused, with this sentence:
„Ez a mentés nem tartalmazza az alkalmazás fájljait, ezért nem állítjuk vissza az adatbázist föléjük — a fájlok így a helyükön maradnak. A fájlok a távoli másolatból állíthatók vissza: Biztonsági mentés → Visszaállítás, „Teljes visszaállítás (fájlok + adatbázis)"."
That is exactly what 07 §6.2 predicts, it names the action that does work, and it refuses
before touching anything. The first version of tonight's harness recorded it as a failed
restore. A harness that calls a correct refusal a failure buries the best result of the night,
so refusals are now recorded as their own verdict and the sentence is quoted.
One real upstream edge HELD honestly — outline 1.9.1 → 1.10.1
The most valuable single result after R-630, because it is the guarded update's own promise exercised on a real upstream version rather than a staged one.
| phase | at |
|---|---|
safety-dump |
0.0 s |
pulling |
+1.1 s |
starting |
+64.8 s |
verifying |
+65.8 s |
failed |
+368.6 s |
The app was stopped and held, and the sentence the household reads names the tier, the date and what the copy contains:
„A(z) outline frissítése 2026-09-22 14:50-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-22 14:44 — ez a másolat a beállításokat, az adatbázist és az adatköteteket tartalmazza."
The restore named in that sentence was then walked and the app came back. outline must not be
promoted.
Two refusals that are the product guarding itself, and both name the blocker
- The Update refuses while a backup or restore runs: „A frissítés most nem indítható: mentés/visszaállítás folyamatban. Próbáld újra, ha befejeződött."
- A second restore refuses and NAMES the app that is blocking it: „Egy visszaállítási művelet (jellyfin) már fut, ezért most nem indítható újabb."
That second guard is exactly the fence R-633 is missing. The product already knows how to refuse
a conflicting operation and how to say which one — for update and for restore. remove has no
such guard, which is why a remove sent during a restore reports success and leaves an orphan.