# THE TWENTY-EIGHT — every app no update drill had ever touched, 2026-09-22 Evidence: `the-28-2026-09-22/`. What came before: `DRILL-update-night-2026-09-21.md` (21 edges, 19 apps) and `PROBE-FIX-2026-09-22.md` (the probe fix and the fifteen moves). --- ## Not done, or changed from the brief **Read this first.** ### Claims in the brief that turned out wrong 1. **There is no `requires:` key in `.felhom.yml`.** The brief said *"A `.felhom.yml` `requires:` (HDD, x86) that 9202 cannot meet → recorded, not forced."* No template has such a key. The constraints live under **`resources:`** as `needs_hdd` and `pi_compatible`. **And none of them excluded anything:** 9202 is `x86_64` with `/mnt/felhom-drives/scratch_hdd` mounted, so every one of the 28 was installable on those grounds. Nothing was skipped for a resource reason. 2. **The brief's file-leg list is wrong, and it told me where to check.** It grouped *"immich, jellyfin, plex, emby, komga, calibre-web, gokapi, homebox, gramps-web"* as file-leg apps. Read from `07-backup-architecture.md` §6.2 — which the brief itself says to read rather than re-derive — only **four of the 28 are class A** (at least one readable file leg): `calibre-web`, `immich`, `komga`, `paperless-ngx`. **`jellyfin`, `plex` and `emby` are class B precisely because their only bind is a `:ro` media mount**, which `ClassifyBinds` excludes; `gokapi`, `homebox` and `gramps-web` keep everything in named volumes. The table below uses `07` §6.2's classes, not the brief's. 3. **The brief's database list is incomplete.** It named calcom, claper, outline, paperless-ngx, rallly and sparkyfitness for PostgreSQL and kimai for MariaDB. **`immich` also carries PostgreSQL (a vectorchord variant) and redis, and `wanderer` carries meilisearch** — two apps the brief put in the "file-leg" group actually run their own datastore. 4. **One of the 28 cannot be installed at all, by design.** `plant-it` is `lifecycle: "abandoned"`, and the product's lifecycle gate refused the deploy with **409 „Ez az alkalmazás jelenleg nem telepíthető."** That is correct behaviour, measured live for the first time. It is the only lifecycle-gated template in the whole catalog of 53. ### Claims in the brief that were verified true, by looking - **The drill repo's Actions are off (R-629).** Read from the API: `has_actions: false, private: true`. And measured rather than trusted: **47 CI jobs before the reset push, 47 after** — the push produced no run and no mail. - **`repoint_drill.py` still works after the catalog moved.** 9202's cache now reads `origin …/app-catalog-drill.git` at `1ad1f34`. - **9202 has the capacity.** 25.9 GB RAM (23.5 free), 28 GB free on `/`, 842 GB on the scratch drive. The root disk is the binding constraint, so **each app's images are removed by name after its verdict** — never `prune` (rule 3). Disk held at 1.9 GB used throughout. ### Changed method, named 5. **A restore that is REFUSED is recorded as `refused-with-a-sentence`, not as a failed restore.** The first version of the harness collapsed the two and mislabelled `calibre-web`. The product had in fact done the right thing — see the finding below — and a harness that calls a correct refusal a failure would have buried it. 6. **The harness now waits for a restore to settle before removing.** It did not at first, and that race produced **R-633**, a real defect. The race was left in the record for `gokapi` and fenced out afterwards, so the remaining apps measure the product rather than the harness. 7. **Three bugs in tonight's own harness, each named with what it cost.** A missing `import re` in the restore step killed the restore half for `wanderer`, `claper`, `sparkyfitness` and `calcom`. A variable named `m` shadowed the app's metadata and broke `paperless-ngx`'s teardown. An empty phase list crashed on `[-1]` when the Update was **refused** before any phase existed, which cost `rallly` its whole walk. **All four apps, and five more, were re-walked serially afterwards** — and that re-walk is what corrected R-634 and produced `ghost`'s proof. An instrument that can drop results silently is not a measurement; these dropped them loudly and were re-run. 8. **Concurrency is part of the method and it changed two results.** Three walks ran at once to fit 28 apps in one night. `POST /api/backup/run` is **box-wide**, so a second caller gets `409 „Mentés már folyamatban"`, and the Update refuses while a backup or restore is in flight (`409 busy`). **Both refusals are the product being right** and both are quoted below. But they cost `ghost` and `rallly` their edge on the first pass, and they are implicated in two of R-634's three instances. The nine re-walks were serial for exactly this reason. ### Corrected 2026-09-22 (evening), from the records rather than by re-running 9. **`calibre-web`'s restore was `refused-with-a-sentence`, not `failed`.** The refusal is in this app's own `log.txt` line 12 as a `flash_error` on the redirect, and its state stayed `running` throughout. It was recorded `failed` because that walk ran **before** the refusal-capture code was added later the same night — the document's own section *"A restore that is REFUSED"* already said so while the table and the record disagreed with it. 10. **`calcom`'s restore was `inconclusive`, not `failed` — and that was my harness, not the product.** The restore was accepted (`flash=flash.restore.started`, no `flash_error`), the app read `running` at +45 s with no hold and no phase, and the classifier looked once more and saw **`starting`** — a settling state it did not list beside `running`/`unhealthy`, so it fell through to `failed`. **The task brief supposed a different cause** — "a read-back of data that was never seeded" — and that is wrong: the seed half is recorded separately and was already `no route`. The record's own `restore_state_seen: "starting"` is the evidence. **Totals move with it:** restores correctly REFUSED go from 2 to **3**. --- ## What this night is, in three lines - **Interventions: SEVEN — over the brief's limit of five, and six of the seven were my own harness, not the product.** Three were bugs in tonight's driver that cost apps their walk and were fixed mid-run (a missing `import re`; a variable that shadowed the app's metadata; a crash when the Update was refused before any phase existed). Two were deliberate method changes (recording a REFUSED restore as its own verdict; making the harness wait for a restore to settle before removing). One was a waiter that deadlocked on its own command line. **The seventh was the product's:** three leftovers it could not clear, which a shell had to. - **26 of 28 deployed; 6 proven; 5 inconclusive; 14 with no upstream edge; 1 failed honestly; 2 that could not be deployed** — one of those by design. - **The one result that matters most:** an app with **no health probe at all** has its working installation **stopped by a successful update**. `paperless-ngx` was healthy on all three containers; the Update ran the full five-minute health wait and then held the app, and the controller named the reason itself — **`no probe container`**. R-630 is raised to P1. --- ## The two findings that are not about any single app ### R-633 — a remove sent while a restore is still running reports success and leaves an orphan `gokapi` was restored from its own local copy at **11:34:07** and removed at **11:34:22**. `POST /backup/restore` answers **302 and does its work in the background**; the remove tore down what existed, and the restore's own `compose up` then **re-created the container at 11:34:24**. Both calls returned success. Twenty-five minutes later, `GET /api/stacks/gokapi` reads **`deployed: false`** while `docker ps -a` shows `gokapi` **`Restarting (1)`** with its full Traefik label set still attached — including `traefik.http.routers.gokapi.rule: Host(`.enkisfelhom.hu`)`, **a rule with an empty subdomain**, because the deploy values that filled it were deleted with the app. Its own log loops *„Salt for admin password invalid… password does not appear to be a SHA-1 hash"* — the volume holding its config was removed correctly, so the binary can never start. **A household can press exactly those two buttons in that order.** The product accepted both and **the remove reported success while leaving the orphan**; nothing in the alarm ladder can fire, because `08` §4 keys on stacks the controller still knows about. This is **R-626's class with the mechanism finally visible** — that row saw a removed `navidrome` come back and could not diagnose it, because the controller had restarted and its log no longer reached the moment. Here the window is **seventeen seconds** and both halves are in the evidence. The harness was then fenced against its own race, so every app after `gokapi` measures the product. Evidence: `apps/gokapi/came-back-evidence.txt`. ### A restore that is REFUSED is the product being right, and nearly went down as a failure `calibre-web` is a class-A app (`07` §6.2): it has a readable file leg, and the local Tier-1 copy does not hold it. The restore was **refused**, with this sentence: > „Ez a mentés nem tartalmazza az alkalmazás fájljait, ezért nem állítjuk vissza az adatbázist > föléjük — a fájlok így a helyükön maradnak. A fájlok a távoli másolatból állíthatók vissza: > Biztonsági mentés → Visszaállítás, „Teljes visszaállítás (fájlok + adatbázis)"." That is exactly what `07` §6.2 predicts, it names the action that does work, and it refuses **before** touching anything. The first version of tonight's harness recorded it as a failed restore. **A harness that calls a correct refusal a failure buries the best result of the night**, so refusals are now recorded as their own verdict and the sentence is quoted. ### One real upstream edge HELD honestly — `outline 1.9.1 → 1.10.1` The most valuable single result after R-630, because it is the guarded update's own promise exercised on a real upstream version rather than a staged one. | phase | at | |---|---| | `safety-dump` | 0.0 s | | `pulling` | +1.1 s | | `starting` | +64.8 s | | `verifying` | +65.8 s | | **`failed`** | **+368.6 s** | The app was stopped and held, and the sentence the household reads names the tier, the date **and what the copy contains**: > „A(z) outline frissítése 2026-09-22 14:50-kor nem sikerült, és az alkalmazás nem indult el az új > verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. > Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-22 14:44 > — ez a másolat a beállításokat, az adatbázist és az adatköteteket tartalmazza." The restore named in that sentence was then walked and the app came back. **`outline` must not be promoted.** ### Two refusals that are the product guarding itself, and both name the blocker - **The Update refuses while a backup or restore runs:** „A frissítés most nem indítható: mentés/visszaállítás folyamatban. Próbáld újra, ha befejeződött." - **A second restore refuses and NAMES the app that is blocking it:** „Egy visszaállítási művelet **(jellyfin)** már fut, ezért most nem indítható újabb." **That second guard is exactly the fence R-633 is missing.** The product already knows how to refuse a conflicting operation and how to say which one — for `update` and for `restore`. **`remove` has no such guard**, which is why a remove sent during a restore reports success and leaves an orphan. --- ## The table — all twenty-eight Classes are `07-backup-architecture.md` §6.2's, not re-derived. | app | class | deployed | seeded | backup | edge | update | restore | removed clean | s | evidence | |---|---|---|---|---|---|---|---|---|---|---| | `calcom` | B volumes-only + postgres | yes | no route | 1 copy | none upstream | — | inconclusive | yes | 675.8 | `apps/calcom/` | | `calibre-web` | A file-leg | yes | no route | 1 copy | none upstream | — | refused-with-a-sentence | yes | 187.5 | `apps/calibre-web/` | | `claper` | B volumes-only + postgres | yes | no route | 1 copy | none upstream | — | ok | yes | 255.3 | `apps/claper/` | | `code-server` | B volumes-only | yes | route failed | 1 copy | `4.129.0` → `4.138.0` | done | ok | yes | 299.9 | `apps/code-server/` | | `crafty-controller` | B volumes-only | yes | no route | 1 copy | `4.10.7` → `4.11.0` | done | ok | yes | 288.5 | `apps/crafty-controller/` | | `emby` | B volumes-only | yes | yes | 1 copy | `4.10.0.20` → `4.11.0.1` | done | ok | yes | 198.7 | `apps/emby/` | | `ghost` | B volumes-only | yes | yes | 1 copy | `6.53.0-alpine` → `6.64.0-alpine` | done | ok | yes | 289.6 | `apps/ghost/` | | `gokapi` | B volumes-only | yes | no route | 1 copy | none upstream | — | ok | **no** | 132.3 | `apps/gokapi/` | | `gramps-web` | B volumes-only | yes | route failed | 1 copy | none upstream | — | ok | yes | 334.1 | `apps/gramps-web/` | | `homebox` | B volumes-only | yes | route failed | 1 copy | none upstream | — | ok | yes | 127.9 | `apps/homebox/` | | `homepage` | B volumes-only | yes | no route | 1 copy | none upstream | — | ok | yes | 137.1 | `apps/homepage/` | | `immich` | A file-leg + postgres+redis | yes | yes | 1 copy | `v3.0.3` → `v3.2.2` | done | refused-with-a-sentence | yes | 375.0 | `apps/immich/` | | `jellyfin` | B volumes-only | yes | route failed | 1 copy | none upstream | — | ok | yes | 220.7 | `apps/jellyfin/` | | `kimai` | B volumes-only + mariadb | yes | route failed | 1 copy | none upstream | — | ok | yes | 311.3 | `apps/kimai/` | | `komga` | A file-leg | yes | route failed | 1 copy | `1.25.0` → `1.27.1` | done | ok | yes | 251.9 | `apps/komga/` | | `onlyoffice` | B volumes-only | yes | route failed | 1 copy | none upstream | — | ok | yes | 227.3 | `apps/onlyoffice/` | | `outline` | B volumes-only + postgres+redis | yes | route failed | 1 copy | `1.9.1` → `1.10.1` | failed | failed | yes | 1389.6 | `apps/outline/` | | `paperless-ngx` | A file-leg + postgres+redis | yes | no route | 1 copy | none upstream | — | refused-with-a-sentence | yes | 253.2 | `apps/paperless-ngx/` | | `plant-it` | B volumes-only | **no** | no route | — | none upstream | — | not-attempted | yes | 78.2 | `apps/plant-it/` | | `plex` | B volumes-only | yes | route failed | 1 copy | `1.41.4.9463-630c9f557` → `1.43.4.10903-e5521bd8c` | done | ok | yes | 262.1 | `apps/plex/` | | `radarr` | B volumes-only | yes | yes | 1 copy | `6.3.0` → `6.4.4` | done | ok | yes | 227.8 | `apps/radarr/` | | `rallly` | B volumes-only + postgres | yes | route failed | 1 copy | `4.11.1` → `4.15.2` | done | ok | yes | 230.1 | `apps/rallly/` | | `recipe-importer` | B volumes-only | yes | no route | 1 copy | none upstream | — | ok | yes | 123.0 | `apps/recipe-importer/` | | `seerr` | B volumes-only | yes | no route | 1 copy | none upstream | — | ok | yes | 218.8 | `apps/seerr/` | | `sonarr` | B volumes-only | yes | yes | 1 copy | `4.0.19` → `4.0.20` | done | ok | yes | 216.6 | `apps/sonarr/` | | `sparkyfitness` | B volumes-only + postgres | **no** | no route | — | none upstream | — | not-attempted | **no** | 534.0 | `apps/sparkyfitness/` | | `termix` | B volumes-only | yes | yes | 1 copy | `2.5.0` → `2.8.0` | done | ok | **no** | 232.5 | `apps/termix/` | | `wanderer` | B volumes-only + meilisearch | yes | no route | 1 copy | none upstream | — | ok | yes | 400.5 | `apps/wanderer/` | **Totals:** **14** no-edge · **6** proven · **5** inconclusive · **2** could-not-deploy · **1** failed — 28 of 28 recorded. **Read across the walk rather than down one column:** **26 of 28 deployed**, **6** had a non-browser route that seeded AND read back, **21** restored from their own copy, **3** were correctly REFUSED a restore, **6 proven / 1 failed** on the apps that had an upstream edge, and **2** left a container behind (R-633). --- ## 5.2 — the templates the static gate cannot judge (R-631) The probe gate's oracle is the probed service's own compose healthcheck. For five templates there is no such oracle, or the paths differ on a check that cannot fail. A static rule cannot settle any of them; asking the running container can. Each was deployed on 9202, its listening sockets read from inside the container, and the probe's own target dialled **on the compose network** — the same call the controller makes. | app | probe | what it listens on | the probe's own dial | verdict | |---|---|---|---|---| | `mealie` | `tcp` 9000 | `0.0.0.0:9000` | 200 | **correct** | | `uptime-kuma` | `http` 3001 | `*:3001` | 302 | **correct** — `http` calls any response healthy, and 302 proves something answers | | `vikunja` | `api` 3456 `/api/v1/info` **expect 200** | (busybox: no `ss`, no `netstat`) | **200** | **correct** — and this is the one that could have failed, because its `expect` block compares the code | | `home-assistant` | `api` 8123 `/api/` **no expect** | `0.0.0.0:8123` | **401** | **correct today, and the 401 is the measurement that proves the warning** | | `crafty-controller` | `tcp` 8443 | (not read — the app never reached `deployed`; see R-634) | **ok (1 ms)** | **correct** | **home-assistant is the one to carry forward.** Its probe dials `/api/` and gets **401** — not 200. It reads healthy only because `probeHTTP` treats any response as healthy when the type is `api` with no `expect` block (`healthprobe.go:253-262`). **Add `expect: {status: 200}` to that template — a change that looks like a tightening — and home-assistant goes permanently unhealthy, and every successful update of it starts stopping it.** That is R-618's failure exactly, one edit away, and it is now a measured number rather than a caution. **All five are settled.** crafty-controller's reading came from its own walk rather than the side job: the controller's log shows `Health probe crafty-controller: TCP :8443 -> ok (1ms)` twice, six minutes apart, while the app was running. The gate's WARN list is therefore not a backlog of suspects: it is four correct templates the gate honestly cannot prove, and one that is correct by accident. --- ## 5.1 — what the guarded Update does when NO probe exists (R-630) `paperless-ngx` has no container whose name equals or begins with its stack name, so `findProbeContainer` returns `""` and `RunHealthProbes` skips the stack silently. **Its probe has never run on any box.** The open question was what `verifying` — which waits on that same probe — does when there is nothing to wait on: pass at once, wait out the timeout, or hold. **It waits out the full timeout and then HOLDS, stopping a working app.** Deployed on 9202, all three containers reported **`healthy`**, the controller read **`running`**, the front door answered **302**. No upstream edge exists for paperless-ngx tonight, so the Update was pressed on the **same version** — which is what a household does on an up-to-date app, and it still walks the whole phase machine. That difference is stated, not glossed. | phase | at | |---|---| | `checking` → `safety-dump` → `pinning` → `pulling` | 0.0–1.1 s | | `starting` | +2.1 s | | `verifying` | +3.1 s | | **`failed`** | **+313.0 s — the app STOPPED** | Afterwards: controller state **`stopped`**, front door **404**. **The controller names the cause itself, so no inference was needed:** > `update paperless-ngx FAILED after the new version was started: not healthy: not healthy within > 5m0s (last: no probe container) — stopping and HOLDING the app; the pin stays on the new version > (its migration may have run)` **`no probe container`.** And the hold sentence is correct about the route back — for this class-A app it warns *„csak a beállításokat és az adatbázist tartalmazza, a fájlokat nem"*. **This is R-618's outcome reached by the opposite road.** There a probe named a port the app does not answer; here no probe exists at all — and the static gate cannot see it, because there is nothing to compare. The gate does print it as a WARNING on every push, which is how it was found. **R-630 is raised P2 → P1.** ## The second promotion list — for the operator, not for me **CC promotes nothing.** These are proposals with the evidence beside them. ### Proposed to move (6) | app | move | update took | its own migration line | |---|---|---|---| | `emby` | `4.10.0.20` → `4.11.0.1` | 198.7 s | yes — `emby \| Info SqliteUserRepository: Sqlite compiler options: ATOMIC_INTRINSICS=1,COMPILER=g` | | `ghost` | `6.53.0-alpine` → `6.64.0-alpine` | 289.6 s | yes — `ghost \| [2026-09-22 15:13:37] INFO Stripe members-migrations skipped because it` | | `immich` | `v3.0.3` → `v3.2.2` | 375.0 s | none printed | | `radarr` | `6.3.0` → `6.4.4` | 227.8 s | yes — `radarr \| [migrations] started` | | `sonarr` | `4.0.19` → `4.0.20` | 216.6 s | yes — `sonarr \| [migrations] started` | | `termix` | `2.5.0` → `2.8.0` | 232.5 s | yes — `termix \| [1:37:29 PM] [INFO] [🗄️] Database layer pre-upgrade backup created [op:database_` | ### Must NOT move, with why (6) | app | edge | why not | |---|---|---| | `code-server` | `4.129.0` → `4.138.0` | the update reached `done`, but this app has no non-browser data route, so nothing proves the household's data survived it | | `crafty-controller` | `4.10.7` → `4.11.0` | the update reached `done`, but this app has no non-browser data route, so nothing proves the household's data survived it | | `komga` | `1.25.0` → `1.27.1` | the update reached `done`, but this app has no non-browser data route, so nothing proves the household's data survived it | | `outline` | `1.9.1` → `1.10.1` | the update ended `failed` — fixture ran and found no non-browser seed route: sign-in requires an external identity provider (OIDC/Slack/Google); no local sign-up route exists | | `plex` | `1.41.4.9463-630c9f557` → `1.43.4.10903-e5521bd8c` | the update reached `done`, but this app has no non-browser data route, so nothing proves the household's data survived it | | `rallly` | `4.11.1` → `4.15.2` | the update reached `done`, but this app has no non-browser data route, so nothing proves the household's data survived it | **No upstream edge tonight, so nothing to propose (16):** `calcom`, `calibre-web`, `claper`, `gokapi`, `gramps-web`, `homebox`, `homepage`, `jellyfin`, `kimai`, `onlyoffice`, `paperless-ngx`, `plant-it`, `recipe-importer`, `seerr`, `sparkyfitness`, `wanderer`. --- ## Teardown — three layers plus Gitea, every claim READ BACK **The machine (9202).** `controller.yaml` restored from `controller.yaml.pre-28`; `git.repo_url` reads back as the **live** catalog with an empty token; and — the one that actually decides which remote is followed (R-615) — the **cache** reads `origin https://gitea.dooplex.hu/admin/app-catalog-felhom.eu.git` at **`1ad1f34`**. No drill images. Disk 1.9 GB used of 32 GB, unchanged from the start. **Three things the product could NOT clear, and a shell had to.** This is not tidy-up, it is the finding: the `termix` and `gokapi` containers left by R-633, and `sparkyfitness`'s `app.yaml` left by R-634. `gokapi` was still `Restarting` two hours later. They were removed **by name** (`docker rm -f termix gokapi`, `rm .../sparkyfitness/app.yaml`) — never a `prune`. Afterwards 9202 runs exactly `felhom-controller`, `filebrowser`, `traefik`, and no `app.yaml` exists anywhere. **A household has no shell.** Evidence: `teardown/manual-cleanup.txt`. **The host (demo-hp).** `pct list` before and after: `9201 demo-hp` and `9202 demo-hp-scratch`, both running, unchanged. `pvesm status` unchanged but for expected scratch growth (`nvme-scratch` 5.19% → 6.95%). **Guest 9201 was never touched.** **The hub.** Nothing provisioned, nothing changed. 9202 runs `hub.enabled: false` (R-620). **Gitea.** The drill repo is reset to the live `main` (`1ad1f34b6e51`). **`git diff` of the live catalog's `templates/` against the night's baseline: 0 lines, and `image:` lines changed: NONE.** **The fences, each read back rather than asserted:** | fence | at the start | at the end | |---|---|---| | live catalog `origin/main` | `1ad1f34b6e51` | **`1ad1f34b6e51`** | | demo-hp guest 9201 catalog cache | `1ad1f34`, live remote | **`1ad1f34`, live remote** | | demo-felhom guest 9201 catalog cache | `1ad1f34`, live remote | **`1ad1f34`, live remote** | | drill repo CI jobs | **47** | **47** — no run, no mail, all night (R-629 holds) | **Peti's box is parked and received nothing.** Nothing ran on DooPlex beyond ordinary pushes, and nothing on `ep0`. `felhom-controller`, `felhom-agent` and the hub were read only — **no product code was written.** No golden, no bake, no vouch, no `--no-verify`, no branch. `local-lvm` untouched, no `prune`, `tester-1` never reset, `drill-r50` untouched.