Both from their own records, not re-run. calibre-web: the flash_error refusal is in its log; that walk predated the refusal-capture code. The report's own prose already said so while the table disagreed. calcom: the restore was accepted and the app read running; the classifier then saw "starting", a settling state it did not list beside running/unhealthy, and fell through to failed. The brief supposed a read-back artefact - that is wrong, and restore_state_seen says so. Restores correctly refused: 2 -> 3. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
25 KiB
THE TWENTY-EIGHT — every app no update drill had ever touched, 2026-09-22
Evidence: the-28-2026-09-22/. What came before: DRILL-update-night-2026-09-21.md (21 edges,
19 apps) and PROBE-FIX-2026-09-22.md (the probe fix and the fifteen moves).
Not done, or changed from the brief
Read this first.
Claims in the brief that turned out wrong
-
There is no
requires:key in.felhom.yml. The brief said "A.felhom.ymlrequires:(HDD, x86) that 9202 cannot meet → recorded, not forced." No template has such a key. The constraints live underresources:asneeds_hddandpi_compatible. And none of them excluded anything: 9202 isx86_64with/mnt/felhom-drives/scratch_hddmounted, so every one of the 28 was installable on those grounds. Nothing was skipped for a resource reason. -
The brief's file-leg list is wrong, and it told me where to check. It grouped "immich, jellyfin, plex, emby, komga, calibre-web, gokapi, homebox, gramps-web" as file-leg apps. Read from
07-backup-architecture.md§6.2 — which the brief itself says to read rather than re-derive — only four of the 28 are class A (at least one readable file leg):calibre-web,immich,komga,paperless-ngx.jellyfin,plexandembyare class B precisely because their only bind is a:romedia mount, whichClassifyBindsexcludes;gokapi,homeboxandgramps-webkeep everything in named volumes. The table below uses07§6.2's classes, not the brief's. -
The brief's database list is incomplete. It named calcom, claper, outline, paperless-ngx, rallly and sparkyfitness for PostgreSQL and kimai for MariaDB.
immichalso carries PostgreSQL (a vectorchord variant) and redis, andwanderercarries meilisearch — two apps the brief put in the "file-leg" group actually run their own datastore. -
One of the 28 cannot be installed at all, by design.
plant-itislifecycle: "abandoned", and the product's lifecycle gate refused the deploy with 409 „Ez az alkalmazás jelenleg nem telepíthető." That is correct behaviour, measured live for the first time. It is the only lifecycle-gated template in the whole catalog of 53.
Claims in the brief that were verified true, by looking
- The drill repo's Actions are off (R-629). Read from the API:
has_actions: false, private: true. And measured rather than trusted: 47 CI jobs before the reset push, 47 after — the push produced no run and no mail. repoint_drill.pystill works after the catalog moved. 9202's cache now readsorigin …/app-catalog-drill.gitat1ad1f34.- 9202 has the capacity. 25.9 GB RAM (23.5 free), 28 GB free on
/, 842 GB on the scratch drive. The root disk is the binding constraint, so each app's images are removed by name after its verdict — neverprune(rule 3). Disk held at 1.9 GB used throughout.
Changed method, named
-
A restore that is REFUSED is recorded as
refused-with-a-sentence, not as a failed restore. The first version of the harness collapsed the two and mislabelledcalibre-web. The product had in fact done the right thing — see the finding below — and a harness that calls a correct refusal a failure would have buried it. -
The harness now waits for a restore to settle before removing. It did not at first, and that race produced R-633, a real defect. The race was left in the record for
gokapiand fenced out afterwards, so the remaining apps measure the product rather than the harness. -
Three bugs in tonight's own harness, each named with what it cost. A missing
import rein the restore step killed the restore half forwanderer,claper,sparkyfitnessandcalcom. A variable namedmshadowed the app's metadata and brokepaperless-ngx's teardown. An empty phase list crashed on[-1]when the Update was refused before any phase existed, which costralllyits whole walk. All four apps, and five more, were re-walked serially afterwards — and that re-walk is what corrected R-634 and producedghost's proof. An instrument that can drop results silently is not a measurement; these dropped them loudly and were re-run. -
Concurrency is part of the method and it changed two results. Three walks ran at once to fit 28 apps in one night.
POST /api/backup/runis box-wide, so a second caller gets409 „Mentés már folyamatban", and the Update refuses while a backup or restore is in flight (409 busy). Both refusals are the product being right and both are quoted below. But they costghostandralllytheir edge on the first pass, and they are implicated in two of R-634's three instances. The nine re-walks were serial for exactly this reason.
Corrected 2026-09-22 (evening), from the records rather than by re-running
-
calibre-web's restore wasrefused-with-a-sentence, notfailed. The refusal is in this app's ownlog.txtline 12 as aflash_erroron the redirect, and its state stayedrunningthroughout. It was recordedfailedbecause that walk ran before the refusal-capture code was added later the same night — the document's own section "A restore that is REFUSED" already said so while the table and the record disagreed with it. -
calcom's restore wasinconclusive, notfailed— and that was my harness, not the product. The restore was accepted (flash=flash.restore.started, noflash_error), the app readrunningat +45 s with no hold and no phase, and the classifier looked once more and sawstarting— a settling state it did not list besiderunning/unhealthy, so it fell through tofailed. The task brief supposed a different cause — "a read-back of data that was never seeded" — and that is wrong: the seed half is recorded separately and was alreadyno route. The record's ownrestore_state_seen: "starting"is the evidence.Totals move with it: restores correctly REFUSED go from 2 to 3.
What this night is, in three lines
- Interventions: SEVEN — over the brief's limit of five, and six of the seven were my own harness,
not the product. Three were bugs in tonight's driver that cost apps their walk and were fixed
mid-run (a missing
import re; a variable that shadowed the app's metadata; a crash when the Update was refused before any phase existed). Two were deliberate method changes (recording a REFUSED restore as its own verdict; making the harness wait for a restore to settle before removing). One was a waiter that deadlocked on its own command line. The seventh was the product's: three leftovers it could not clear, which a shell had to. - 26 of 28 deployed; 6 proven; 5 inconclusive; 14 with no upstream edge; 1 failed honestly; 2 that could not be deployed — one of those by design.
- The one result that matters most: an app with no health probe at all has its working
installation stopped by a successful update.
paperless-ngxwas healthy on all three containers; the Update ran the full five-minute health wait and then held the app, and the controller named the reason itself —no probe container. R-630 is raised to P1.
The two findings that are not about any single app
R-633 — a remove sent while a restore is still running reports success and leaves an orphan
gokapi was restored from its own local copy at 11:34:07 and removed at 11:34:22.
POST /backup/restore answers 302 and does its work in the background; the remove tore down
what existed, and the restore's own compose up then re-created the container at 11:34:24.
Both calls returned success.
Twenty-five minutes later, GET /api/stacks/gokapi reads deployed: false while docker ps -a
shows gokapi Restarting (1) with its full Traefik label set still attached — including
traefik.http.routers.gokapi.rule: Host(.enkisfelhom.hu), a rule with an empty subdomain,
because the deploy values that filled it were deleted with the app. Its own log loops
„Salt for admin password invalid… password does not appear to be a SHA-1 hash" — the volume
holding its config was removed correctly, so the binary can never start.
A household can press exactly those two buttons in that order. The product accepted both and
the remove reported success while leaving the orphan; nothing in the alarm ladder can fire,
because 08 §4 keys on stacks the controller still knows about. This is R-626's class with the
mechanism finally visible — that row saw a removed navidrome come back and could not diagnose
it, because the controller had restarted and its log no longer reached the moment. Here the window
is seventeen seconds and both halves are in the evidence.
The harness was then fenced against its own race, so every app after gokapi measures the product.
Evidence: apps/gokapi/came-back-evidence.txt.
A restore that is REFUSED is the product being right, and nearly went down as a failure
calibre-web is a class-A app (07 §6.2): it has a readable file leg, and the local Tier-1 copy
does not hold it. The restore was refused, with this sentence:
„Ez a mentés nem tartalmazza az alkalmazás fájljait, ezért nem állítjuk vissza az adatbázist föléjük — a fájlok így a helyükön maradnak. A fájlok a távoli másolatból állíthatók vissza: Biztonsági mentés → Visszaállítás, „Teljes visszaállítás (fájlok + adatbázis)"."
That is exactly what 07 §6.2 predicts, it names the action that does work, and it refuses
before touching anything. The first version of tonight's harness recorded it as a failed
restore. A harness that calls a correct refusal a failure buries the best result of the night,
so refusals are now recorded as their own verdict and the sentence is quoted.
One real upstream edge HELD honestly — outline 1.9.1 → 1.10.1
The most valuable single result after R-630, because it is the guarded update's own promise exercised on a real upstream version rather than a staged one.
| phase | at |
|---|---|
safety-dump |
0.0 s |
pulling |
+1.1 s |
starting |
+64.8 s |
verifying |
+65.8 s |
failed |
+368.6 s |
The app was stopped and held, and the sentence the household reads names the tier, the date and what the copy contains:
„A(z) outline frissítése 2026-09-22 14:50-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-22 14:44 — ez a másolat a beállításokat, az adatbázist és az adatköteteket tartalmazza."
The restore named in that sentence was then walked and the app came back. outline must not be
promoted.
Two refusals that are the product guarding itself, and both name the blocker
- The Update refuses while a backup or restore runs: „A frissítés most nem indítható: mentés/visszaállítás folyamatban. Próbáld újra, ha befejeződött."
- A second restore refuses and NAMES the app that is blocking it: „Egy visszaállítási művelet (jellyfin) már fut, ezért most nem indítható újabb."
That second guard is exactly the fence R-633 is missing. The product already knows how to refuse
a conflicting operation and how to say which one — for update and for restore. remove has no
such guard, which is why a remove sent during a restore reports success and leaves an orphan.
The table — all twenty-eight
Classes are 07-backup-architecture.md §6.2's, not re-derived.
| app | class | deployed | seeded | backup | edge | update | restore | removed clean | s | evidence |
|---|---|---|---|---|---|---|---|---|---|---|
calcom |
B volumes-only + postgres | yes | no route | 1 copy | none upstream | — | inconclusive | yes | 675.8 | apps/calcom/ |
calibre-web |
A file-leg | yes | no route | 1 copy | none upstream | — | refused-with-a-sentence | yes | 187.5 | apps/calibre-web/ |
claper |
B volumes-only + postgres | yes | no route | 1 copy | none upstream | — | ok | yes | 255.3 | apps/claper/ |
code-server |
B volumes-only | yes | route failed | 1 copy | 4.129.0 → 4.138.0 |
done | ok | yes | 299.9 | apps/code-server/ |
crafty-controller |
B volumes-only | yes | no route | 1 copy | 4.10.7 → 4.11.0 |
done | ok | yes | 288.5 | apps/crafty-controller/ |
emby |
B volumes-only | yes | yes | 1 copy | 4.10.0.20 → 4.11.0.1 |
done | ok | yes | 198.7 | apps/emby/ |
ghost |
B volumes-only | yes | yes | 1 copy | 6.53.0-alpine → 6.64.0-alpine |
done | ok | yes | 289.6 | apps/ghost/ |
gokapi |
B volumes-only | yes | no route | 1 copy | none upstream | — | ok | no | 132.3 | apps/gokapi/ |
gramps-web |
B volumes-only | yes | route failed | 1 copy | none upstream | — | ok | yes | 334.1 | apps/gramps-web/ |
homebox |
B volumes-only | yes | route failed | 1 copy | none upstream | — | ok | yes | 127.9 | apps/homebox/ |
homepage |
B volumes-only | yes | no route | 1 copy | none upstream | — | ok | yes | 137.1 | apps/homepage/ |
immich |
A file-leg + postgres+redis | yes | yes | 1 copy | v3.0.3 → v3.2.2 |
done | refused-with-a-sentence | yes | 375.0 | apps/immich/ |
jellyfin |
B volumes-only | yes | route failed | 1 copy | none upstream | — | ok | yes | 220.7 | apps/jellyfin/ |
kimai |
B volumes-only + mariadb | yes | route failed | 1 copy | none upstream | — | ok | yes | 311.3 | apps/kimai/ |
komga |
A file-leg | yes | route failed | 1 copy | 1.25.0 → 1.27.1 |
done | ok | yes | 251.9 | apps/komga/ |
onlyoffice |
B volumes-only | yes | route failed | 1 copy | none upstream | — | ok | yes | 227.3 | apps/onlyoffice/ |
outline |
B volumes-only + postgres+redis | yes | route failed | 1 copy | 1.9.1 → 1.10.1 |
failed | failed | yes | 1389.6 | apps/outline/ |
paperless-ngx |
A file-leg + postgres+redis | yes | no route | 1 copy | none upstream | — | refused-with-a-sentence | yes | 253.2 | apps/paperless-ngx/ |
plant-it |
B volumes-only | no | no route | — | none upstream | — | not-attempted | yes | 78.2 | apps/plant-it/ |
plex |
B volumes-only | yes | route failed | 1 copy | 1.41.4.9463-630c9f557 → 1.43.4.10903-e5521bd8c |
done | ok | yes | 262.1 | apps/plex/ |
radarr |
B volumes-only | yes | yes | 1 copy | 6.3.0 → 6.4.4 |
done | ok | yes | 227.8 | apps/radarr/ |
rallly |
B volumes-only + postgres | yes | route failed | 1 copy | 4.11.1 → 4.15.2 |
done | ok | yes | 230.1 | apps/rallly/ |
recipe-importer |
B volumes-only | yes | no route | 1 copy | none upstream | — | ok | yes | 123.0 | apps/recipe-importer/ |
seerr |
B volumes-only | yes | no route | 1 copy | none upstream | — | ok | yes | 218.8 | apps/seerr/ |
sonarr |
B volumes-only | yes | yes | 1 copy | 4.0.19 → 4.0.20 |
done | ok | yes | 216.6 | apps/sonarr/ |
sparkyfitness |
B volumes-only + postgres | no | no route | — | none upstream | — | not-attempted | no | 534.0 | apps/sparkyfitness/ |
termix |
B volumes-only | yes | yes | 1 copy | 2.5.0 → 2.8.0 |
done | ok | no | 232.5 | apps/termix/ |
wanderer |
B volumes-only + meilisearch | yes | no route | 1 copy | none upstream | — | ok | yes | 400.5 | apps/wanderer/ |
Totals: 14 no-edge · 6 proven · 5 inconclusive · 2 could-not-deploy · 1 failed — 28 of 28 recorded.
Read across the walk rather than down one column: 26 of 28 deployed, 6 had a non-browser route that seeded AND read back, 21 restored from their own copy, 3 were correctly REFUSED a restore, 6 proven / 1 failed on the apps that had an upstream edge, and 2 left a container behind (R-633).
5.2 — the templates the static gate cannot judge (R-631)
The probe gate's oracle is the probed service's own compose healthcheck. For five templates there is no such oracle, or the paths differ on a check that cannot fail. A static rule cannot settle any of them; asking the running container can. Each was deployed on 9202, its listening sockets read from inside the container, and the probe's own target dialled on the compose network — the same call the controller makes.
| app | probe | what it listens on | the probe's own dial | verdict |
|---|---|---|---|---|
mealie |
tcp 9000 |
0.0.0.0:9000 |
200 | correct |
uptime-kuma |
http 3001 |
*:3001 |
302 | correct — http calls any response healthy, and 302 proves something answers |
vikunja |
api 3456 /api/v1/info expect 200 |
(busybox: no ss, no netstat) |
200 | correct — and this is the one that could have failed, because its expect block compares the code |
home-assistant |
api 8123 /api/ no expect |
0.0.0.0:8123 |
401 | correct today, and the 401 is the measurement that proves the warning |
crafty-controller |
tcp 8443 |
(not read — the app never reached deployed; see R-634) |
ok (1 ms) | correct |
home-assistant is the one to carry forward. Its probe dials /api/ and gets 401 — not 200.
It reads healthy only because probeHTTP treats any response as healthy when the type is api with
no expect block (healthprobe.go:253-262). Add expect: {status: 200} to that template — a
change that looks like a tightening — and home-assistant goes permanently unhealthy, and every
successful update of it starts stopping it. That is R-618's failure exactly, one edit away, and it
is now a measured number rather than a caution.
All five are settled. crafty-controller's reading came from its own walk rather than the side job: the controller's log shows Health probe crafty-controller: TCP :8443 -> ok (1ms) twice, six minutes apart, while the app was running. The gate's WARN list is therefore not a backlog of suspects: it is
four correct templates the gate honestly cannot prove, and one that is correct by accident.
5.1 — what the guarded Update does when NO probe exists (R-630)
paperless-ngx has no container whose name equals or begins with its stack name, so
findProbeContainer returns "" and RunHealthProbes skips the stack silently. Its probe has
never run on any box. The open question was what verifying — which waits on that same probe —
does when there is nothing to wait on: pass at once, wait out the timeout, or hold.
It waits out the full timeout and then HOLDS, stopping a working app.
Deployed on 9202, all three containers reported healthy, the controller read running, the
front door answered 302. No upstream edge exists for paperless-ngx tonight, so the Update was
pressed on the same version — which is what a household does on an up-to-date app, and it still
walks the whole phase machine. That difference is stated, not glossed.
| phase | at |
|---|---|
checking → safety-dump → pinning → pulling |
0.0–1.1 s |
starting |
+2.1 s |
verifying |
+3.1 s |
failed |
+313.0 s — the app STOPPED |
Afterwards: controller state stopped, front door 404.
The controller names the cause itself, so no inference was needed:
update paperless-ngx FAILED after the new version was started: not healthy: not healthy within 5m0s (last: no probe container) — stopping and HOLDING the app; the pin stays on the new version (its migration may have run)
no probe container. And the hold sentence is correct about the route back — for this class-A
app it warns „csak a beállításokat és az adatbázist tartalmazza, a fájlokat nem".
This is R-618's outcome reached by the opposite road. There a probe named a port the app does not answer; here no probe exists at all — and the static gate cannot see it, because there is nothing to compare. The gate does print it as a WARNING on every push, which is how it was found. R-630 is raised P2 → P1.
The second promotion list — for the operator, not for me
CC promotes nothing. These are proposals with the evidence beside them.
Proposed to move (6)
| app | move | update took | its own migration line |
|---|---|---|---|
emby |
4.10.0.20 → 4.11.0.1 |
198.7 s | yes — emby | Info SqliteUserRepository: Sqlite compiler options: ATOMIC_INTRINSICS=1,COMPILER=g |
ghost |
6.53.0-alpine → 6.64.0-alpine |
289.6 s | yes — ghost | [2026-09-22 15:13:37] [36mINFO[39m Stripe members-migrations skipped because it |
immich |
v3.0.3 → v3.2.2 |
375.0 s | none printed |
radarr |
6.3.0 → 6.4.4 |
227.8 s | yes — radarr | [migrations] started |
sonarr |
4.0.19 → 4.0.20 |
216.6 s | yes — sonarr | [migrations] started |
termix |
2.5.0 → 2.8.0 |
232.5 s | yes — termix | [1:37:29 PM] [INFO] [🗄️] Database layer pre-upgrade backup created [op:database_ |
Must NOT move, with why (6)
| app | edge | why not |
|---|---|---|
code-server |
4.129.0 → 4.138.0 |
the update reached done, but this app has no non-browser data route, so nothing proves the household's data survived it |
crafty-controller |
4.10.7 → 4.11.0 |
the update reached done, but this app has no non-browser data route, so nothing proves the household's data survived it |
komga |
1.25.0 → 1.27.1 |
the update reached done, but this app has no non-browser data route, so nothing proves the household's data survived it |
outline |
1.9.1 → 1.10.1 |
the update ended failed — fixture ran and found no non-browser seed route: sign-in requires an external identity provider (OIDC/Slack/Google); no local sign-up route exists |
plex |
1.41.4.9463-630c9f557 → 1.43.4.10903-e5521bd8c |
the update reached done, but this app has no non-browser data route, so nothing proves the household's data survived it |
rallly |
4.11.1 → 4.15.2 |
the update reached done, but this app has no non-browser data route, so nothing proves the household's data survived it |
No upstream edge tonight, so nothing to propose (16): calcom, calibre-web, claper, gokapi, gramps-web, homebox, homepage, jellyfin, kimai, onlyoffice, paperless-ngx, plant-it, recipe-importer, seerr, sparkyfitness, wanderer.
Teardown — three layers plus Gitea, every claim READ BACK
The machine (9202). controller.yaml restored from controller.yaml.pre-28; git.repo_url
reads back as the live catalog with an empty token; and — the one that actually decides which
remote is followed (R-615) — the cache reads
origin https://gitea.dooplex.hu/admin/app-catalog-felhom.eu.git at 1ad1f34. No drill
images. Disk 1.9 GB used of 32 GB, unchanged from the start.
Three things the product could NOT clear, and a shell had to. This is not tidy-up, it is the
finding: the termix and gokapi containers left by R-633, and sparkyfitness's app.yaml left by
R-634. gokapi was still Restarting two hours later. They were removed by name
(docker rm -f termix gokapi, rm .../sparkyfitness/app.yaml) — never a prune. Afterwards 9202
runs exactly felhom-controller, filebrowser, traefik, and no app.yaml exists anywhere.
A household has no shell. Evidence: teardown/manual-cleanup.txt.
The host (demo-hp). pct list before and after: 9201 demo-hp and 9202 demo-hp-scratch, both
running, unchanged. pvesm status unchanged but for expected scratch growth (nvme-scratch
5.19% → 6.95%). Guest 9201 was never touched.
The hub. Nothing provisioned, nothing changed. 9202 runs hub.enabled: false (R-620).
Gitea. The drill repo is reset to the live main (1ad1f34b6e51). git diff of the live
catalog's templates/ against the night's baseline: 0 lines, and image: lines changed: NONE.
The fences, each read back rather than asserted:
| fence | at the start | at the end |
|---|---|---|
live catalog origin/main |
1ad1f34b6e51 |
1ad1f34b6e51 |
| demo-hp guest 9201 catalog cache | 1ad1f34, live remote |
1ad1f34, live remote |
| demo-felhom guest 9201 catalog cache | 1ad1f34, live remote |
1ad1f34, live remote |
| drill repo CI jobs | 47 | 47 — no run, no mail, all night (R-629 holds) |
Peti's box is parked and received nothing. Nothing ran on DooPlex beyond ordinary pushes, and
nothing on ep0. felhom-controller, felhom-agent and the hub were read only — no product code
was written. No golden, no bake, no vouch, no --no-verify, no branch. local-lvm untouched, no
prune, tester-1 never reset, drill-r50 untouched.