Files
felhom.eu/documentation/audits/DRILL-the-28-2026-09-22.md
T
admin 1de6aaf904
gates / gates (push) Successful in 25s
corrections: calibre-web restore was REFUSED, calcom was INCONCLUSIVE
Both from their own records, not re-run.

calibre-web: the flash_error refusal is in its log; that walk predated the refusal-capture code.
The report's own prose already said so while the table disagreed.

calcom: the restore was accepted and the app read running; the classifier then saw "starting", a
settling state it did not list beside running/unhealthy, and fell through to failed. The brief
supposed a read-back artefact - that is wrong, and restore_state_seen says so.

Restores correctly refused: 2 -> 3.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-22 20:18:37 +02:00

25 KiB
Raw Blame History

THE TWENTY-EIGHT — every app no update drill had ever touched, 2026-09-22

Evidence: the-28-2026-09-22/. What came before: DRILL-update-night-2026-09-21.md (21 edges, 19 apps) and PROBE-FIX-2026-09-22.md (the probe fix and the fifteen moves).


Not done, or changed from the brief

Read this first.

Claims in the brief that turned out wrong

  1. There is no requires: key in .felhom.yml. The brief said "A .felhom.yml requires: (HDD, x86) that 9202 cannot meet → recorded, not forced." No template has such a key. The constraints live under resources: as needs_hdd and pi_compatible. And none of them excluded anything: 9202 is x86_64 with /mnt/felhom-drives/scratch_hdd mounted, so every one of the 28 was installable on those grounds. Nothing was skipped for a resource reason.

  2. The brief's file-leg list is wrong, and it told me where to check. It grouped "immich, jellyfin, plex, emby, komga, calibre-web, gokapi, homebox, gramps-web" as file-leg apps. Read from 07-backup-architecture.md §6.2 — which the brief itself says to read rather than re-derive — only four of the 28 are class A (at least one readable file leg): calibre-web, immich, komga, paperless-ngx. jellyfin, plex and emby are class B precisely because their only bind is a :ro media mount, which ClassifyBinds excludes; gokapi, homebox and gramps-web keep everything in named volumes. The table below uses 07 §6.2's classes, not the brief's.

  3. The brief's database list is incomplete. It named calcom, claper, outline, paperless-ngx, rallly and sparkyfitness for PostgreSQL and kimai for MariaDB. immich also carries PostgreSQL (a vectorchord variant) and redis, and wanderer carries meilisearch — two apps the brief put in the "file-leg" group actually run their own datastore.

  4. One of the 28 cannot be installed at all, by design. plant-it is lifecycle: "abandoned", and the product's lifecycle gate refused the deploy with 409 „Ez az alkalmazás jelenleg nem telepíthető." That is correct behaviour, measured live for the first time. It is the only lifecycle-gated template in the whole catalog of 53.

Claims in the brief that were verified true, by looking

  • The drill repo's Actions are off (R-629). Read from the API: has_actions: false, private: true. And measured rather than trusted: 47 CI jobs before the reset push, 47 after — the push produced no run and no mail.
  • repoint_drill.py still works after the catalog moved. 9202's cache now reads origin …/app-catalog-drill.git at 1ad1f34.
  • 9202 has the capacity. 25.9 GB RAM (23.5 free), 28 GB free on /, 842 GB on the scratch drive. The root disk is the binding constraint, so each app's images are removed by name after its verdict — never prune (rule 3). Disk held at 1.9 GB used throughout.

Changed method, named

  1. A restore that is REFUSED is recorded as refused-with-a-sentence, not as a failed restore. The first version of the harness collapsed the two and mislabelled calibre-web. The product had in fact done the right thing — see the finding below — and a harness that calls a correct refusal a failure would have buried it.

  2. The harness now waits for a restore to settle before removing. It did not at first, and that race produced R-633, a real defect. The race was left in the record for gokapi and fenced out afterwards, so the remaining apps measure the product rather than the harness.

  3. Three bugs in tonight's own harness, each named with what it cost. A missing import re in the restore step killed the restore half for wanderer, claper, sparkyfitness and calcom. A variable named m shadowed the app's metadata and broke paperless-ngx's teardown. An empty phase list crashed on [-1] when the Update was refused before any phase existed, which cost rallly its whole walk. All four apps, and five more, were re-walked serially afterwards — and that re-walk is what corrected R-634 and produced ghost's proof. An instrument that can drop results silently is not a measurement; these dropped them loudly and were re-run.

  4. Concurrency is part of the method and it changed two results. Three walks ran at once to fit 28 apps in one night. POST /api/backup/run is box-wide, so a second caller gets 409 „Mentés már folyamatban", and the Update refuses while a backup or restore is in flight (409 busy). Both refusals are the product being right and both are quoted below. But they cost ghost and rallly their edge on the first pass, and they are implicated in two of R-634's three instances. The nine re-walks were serial for exactly this reason.

Corrected 2026-09-22 (evening), from the records rather than by re-running

  1. calibre-web's restore was refused-with-a-sentence, not failed. The refusal is in this app's own log.txt line 12 as a flash_error on the redirect, and its state stayed running throughout. It was recorded failed because that walk ran before the refusal-capture code was added later the same night — the document's own section "A restore that is REFUSED" already said so while the table and the record disagreed with it.

  2. calcom's restore was inconclusive, not failed — and that was my harness, not the product. The restore was accepted (flash=flash.restore.started, no flash_error), the app read running at +45 s with no hold and no phase, and the classifier looked once more and saw starting — a settling state it did not list beside running/unhealthy, so it fell through to failed. The task brief supposed a different cause — "a read-back of data that was never seeded" — and that is wrong: the seed half is recorded separately and was already no route. The record's own restore_state_seen: "starting" is the evidence.

    Totals move with it: restores correctly REFUSED go from 2 to 3.


What this night is, in three lines

  • Interventions: SEVEN — over the brief's limit of five, and six of the seven were my own harness, not the product. Three were bugs in tonight's driver that cost apps their walk and were fixed mid-run (a missing import re; a variable that shadowed the app's metadata; a crash when the Update was refused before any phase existed). Two were deliberate method changes (recording a REFUSED restore as its own verdict; making the harness wait for a restore to settle before removing). One was a waiter that deadlocked on its own command line. The seventh was the product's: three leftovers it could not clear, which a shell had to.
  • 26 of 28 deployed; 6 proven; 5 inconclusive; 14 with no upstream edge; 1 failed honestly; 2 that could not be deployed — one of those by design.
  • The one result that matters most: an app with no health probe at all has its working installation stopped by a successful update. paperless-ngx was healthy on all three containers; the Update ran the full five-minute health wait and then held the app, and the controller named the reason itself — no probe container. R-630 is raised to P1.

The two findings that are not about any single app

R-633 — a remove sent while a restore is still running reports success and leaves an orphan

gokapi was restored from its own local copy at 11:34:07 and removed at 11:34:22. POST /backup/restore answers 302 and does its work in the background; the remove tore down what existed, and the restore's own compose up then re-created the container at 11:34:24. Both calls returned success.

Twenty-five minutes later, GET /api/stacks/gokapi reads deployed: false while docker ps -a shows gokapi Restarting (1) with its full Traefik label set still attached — including traefik.http.routers.gokapi.rule: Host(.enkisfelhom.hu), a rule with an empty subdomain, because the deploy values that filled it were deleted with the app. Its own log loops „Salt for admin password invalid… password does not appear to be a SHA-1 hash" — the volume holding its config was removed correctly, so the binary can never start.

A household can press exactly those two buttons in that order. The product accepted both and the remove reported success while leaving the orphan; nothing in the alarm ladder can fire, because 08 §4 keys on stacks the controller still knows about. This is R-626's class with the mechanism finally visible — that row saw a removed navidrome come back and could not diagnose it, because the controller had restarted and its log no longer reached the moment. Here the window is seventeen seconds and both halves are in the evidence.

The harness was then fenced against its own race, so every app after gokapi measures the product. Evidence: apps/gokapi/came-back-evidence.txt.

A restore that is REFUSED is the product being right, and nearly went down as a failure

calibre-web is a class-A app (07 §6.2): it has a readable file leg, and the local Tier-1 copy does not hold it. The restore was refused, with this sentence:

„Ez a mentés nem tartalmazza az alkalmazás fájljait, ezért nem állítjuk vissza az adatbázist föléjük — a fájlok így a helyükön maradnak. A fájlok a távoli másolatból állíthatók vissza: Biztonsági mentés → Visszaállítás, „Teljes visszaállítás (fájlok + adatbázis)"."

That is exactly what 07 §6.2 predicts, it names the action that does work, and it refuses before touching anything. The first version of tonight's harness recorded it as a failed restore. A harness that calls a correct refusal a failure buries the best result of the night, so refusals are now recorded as their own verdict and the sentence is quoted.

One real upstream edge HELD honestly — outline 1.9.1 → 1.10.1

The most valuable single result after R-630, because it is the guarded update's own promise exercised on a real upstream version rather than a staged one.

phase at
safety-dump 0.0 s
pulling +1.1 s
starting +64.8 s
verifying +65.8 s
failed +368.6 s

The app was stopped and held, and the sentence the household reads names the tier, the date and what the copy contains:

„A(z) outline frissítése 2026-09-22 14:50-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-22 14:44 — ez a másolat a beállításokat, az adatbázist és az adatköteteket tartalmazza."

The restore named in that sentence was then walked and the app came back. outline must not be promoted.

Two refusals that are the product guarding itself, and both name the blocker

  • The Update refuses while a backup or restore runs: „A frissítés most nem indítható: mentés/visszaállítás folyamatban. Próbáld újra, ha befejeződött."
  • A second restore refuses and NAMES the app that is blocking it: „Egy visszaállítási művelet (jellyfin) már fut, ezért most nem indítható újabb."

That second guard is exactly the fence R-633 is missing. The product already knows how to refuse a conflicting operation and how to say which one — for update and for restore. remove has no such guard, which is why a remove sent during a restore reports success and leaves an orphan.


The table — all twenty-eight

Classes are 07-backup-architecture.md §6.2's, not re-derived.

app class deployed seeded backup edge update restore removed clean s evidence
calcom B volumes-only + postgres yes no route 1 copy none upstream — inconclusive yes 675.8 apps/calcom/
calibre-web A file-leg yes no route 1 copy none upstream — refused-with-a-sentence yes 187.5 apps/calibre-web/
claper B volumes-only + postgres yes no route 1 copy none upstream — ok yes 255.3 apps/claper/
code-server B volumes-only yes route failed 1 copy 4.129.0 → 4.138.0 done ok yes 299.9 apps/code-server/
crafty-controller B volumes-only yes no route 1 copy 4.10.7 → 4.11.0 done ok yes 288.5 apps/crafty-controller/
emby B volumes-only yes yes 1 copy 4.10.0.20 → 4.11.0.1 done ok yes 198.7 apps/emby/
ghost B volumes-only yes yes 1 copy 6.53.0-alpine → 6.64.0-alpine done ok yes 289.6 apps/ghost/
gokapi B volumes-only yes no route 1 copy none upstream — ok no 132.3 apps/gokapi/
gramps-web B volumes-only yes route failed 1 copy none upstream — ok yes 334.1 apps/gramps-web/
homebox B volumes-only yes route failed 1 copy none upstream — ok yes 127.9 apps/homebox/
homepage B volumes-only yes no route 1 copy none upstream — ok yes 137.1 apps/homepage/
immich A file-leg + postgres+redis yes yes 1 copy v3.0.3 → v3.2.2 done refused-with-a-sentence yes 375.0 apps/immich/
jellyfin B volumes-only yes route failed 1 copy none upstream — ok yes 220.7 apps/jellyfin/
kimai B volumes-only + mariadb yes route failed 1 copy none upstream — ok yes 311.3 apps/kimai/
komga A file-leg yes route failed 1 copy 1.25.0 → 1.27.1 done ok yes 251.9 apps/komga/
onlyoffice B volumes-only yes route failed 1 copy none upstream — ok yes 227.3 apps/onlyoffice/
outline B volumes-only + postgres+redis yes route failed 1 copy 1.9.1 → 1.10.1 failed failed yes 1389.6 apps/outline/
paperless-ngx A file-leg + postgres+redis yes no route 1 copy none upstream — refused-with-a-sentence yes 253.2 apps/paperless-ngx/
plant-it B volumes-only no no route — none upstream — not-attempted yes 78.2 apps/plant-it/
plex B volumes-only yes route failed 1 copy 1.41.4.9463-630c9f557 → 1.43.4.10903-e5521bd8c done ok yes 262.1 apps/plex/
radarr B volumes-only yes yes 1 copy 6.3.0 → 6.4.4 done ok yes 227.8 apps/radarr/
rallly B volumes-only + postgres yes route failed 1 copy 4.11.1 → 4.15.2 done ok yes 230.1 apps/rallly/
recipe-importer B volumes-only yes no route 1 copy none upstream — ok yes 123.0 apps/recipe-importer/
seerr B volumes-only yes no route 1 copy none upstream — ok yes 218.8 apps/seerr/
sonarr B volumes-only yes yes 1 copy 4.0.19 → 4.0.20 done ok yes 216.6 apps/sonarr/
sparkyfitness B volumes-only + postgres no no route — none upstream — not-attempted no 534.0 apps/sparkyfitness/
termix B volumes-only yes yes 1 copy 2.5.0 → 2.8.0 done ok no 232.5 apps/termix/
wanderer B volumes-only + meilisearch yes no route 1 copy none upstream — ok yes 400.5 apps/wanderer/

Totals: 14 no-edge · 6 proven · 5 inconclusive · 2 could-not-deploy · 1 failed — 28 of 28 recorded.

Read across the walk rather than down one column: 26 of 28 deployed, 6 had a non-browser route that seeded AND read back, 21 restored from their own copy, 3 were correctly REFUSED a restore, 6 proven / 1 failed on the apps that had an upstream edge, and 2 left a container behind (R-633).


5.2 — the templates the static gate cannot judge (R-631)

The probe gate's oracle is the probed service's own compose healthcheck. For five templates there is no such oracle, or the paths differ on a check that cannot fail. A static rule cannot settle any of them; asking the running container can. Each was deployed on 9202, its listening sockets read from inside the container, and the probe's own target dialled on the compose network — the same call the controller makes.

app probe what it listens on the probe's own dial verdict
mealie tcp 9000 0.0.0.0:9000 200 correct
uptime-kuma http 3001 *:3001 302 correct — http calls any response healthy, and 302 proves something answers
vikunja api 3456 /api/v1/info expect 200 (busybox: no ss, no netstat) 200 correct — and this is the one that could have failed, because its expect block compares the code
home-assistant api 8123 /api/ no expect 0.0.0.0:8123 401 correct today, and the 401 is the measurement that proves the warning
crafty-controller tcp 8443 (not read — the app never reached deployed; see R-634) ok (1 ms) correct

home-assistant is the one to carry forward. Its probe dials /api/ and gets 401 — not 200. It reads healthy only because probeHTTP treats any response as healthy when the type is api with no expect block (healthprobe.go:253-262). Add expect: {status: 200} to that template — a change that looks like a tightening — and home-assistant goes permanently unhealthy, and every successful update of it starts stopping it. That is R-618's failure exactly, one edit away, and it is now a measured number rather than a caution.

All five are settled. crafty-controller's reading came from its own walk rather than the side job: the controller's log shows Health probe crafty-controller: TCP :8443 -> ok (1ms) twice, six minutes apart, while the app was running. The gate's WARN list is therefore not a backlog of suspects: it is four correct templates the gate honestly cannot prove, and one that is correct by accident.


5.1 — what the guarded Update does when NO probe exists (R-630)

paperless-ngx has no container whose name equals or begins with its stack name, so findProbeContainer returns "" and RunHealthProbes skips the stack silently. Its probe has never run on any box. The open question was what verifying — which waits on that same probe — does when there is nothing to wait on: pass at once, wait out the timeout, or hold.

It waits out the full timeout and then HOLDS, stopping a working app.

Deployed on 9202, all three containers reported healthy, the controller read running, the front door answered 302. No upstream edge exists for paperless-ngx tonight, so the Update was pressed on the same version — which is what a household does on an up-to-date app, and it still walks the whole phase machine. That difference is stated, not glossed.

phase at
checking → safety-dump → pinning → pulling 0.0–1.1 s
starting +2.1 s
verifying +3.1 s
failed +313.0 s — the app STOPPED

Afterwards: controller state stopped, front door 404.

The controller names the cause itself, so no inference was needed:

update paperless-ngx FAILED after the new version was started: not healthy: not healthy within 5m0s (last: no probe container) — stopping and HOLDING the app; the pin stays on the new version (its migration may have run)

no probe container. And the hold sentence is correct about the route back — for this class-A app it warns „csak a beállításokat és az adatbázist tartalmazza, a fájlokat nem".

This is R-618's outcome reached by the opposite road. There a probe named a port the app does not answer; here no probe exists at all — and the static gate cannot see it, because there is nothing to compare. The gate does print it as a WARNING on every push, which is how it was found. R-630 is raised P2 → P1.

The second promotion list — for the operator, not for me

CC promotes nothing. These are proposals with the evidence beside them.

Proposed to move (6)

app move update took its own migration line
emby 4.10.0.20 → 4.11.0.1 198.7 s yes — emby | Info SqliteUserRepository: Sqlite compiler options: ATOMIC_INTRINSICS=1,COMPILER=g
ghost 6.53.0-alpine → 6.64.0-alpine 289.6 s yes — ghost | [2026-09-22 15:13:37] INFO Stripe members-migrations skipped because it
immich v3.0.3 → v3.2.2 375.0 s none printed
radarr 6.3.0 → 6.4.4 227.8 s yes — radarr | [migrations] started
sonarr 4.0.19 → 4.0.20 216.6 s yes — sonarr | [migrations] started
termix 2.5.0 → 2.8.0 232.5 s yes — termix | [1:37:29 PM] [INFO] [🗄️] Database layer pre-upgrade backup created [op:database_

Must NOT move, with why (6)

app edge why not
code-server 4.129.0 → 4.138.0 the update reached done, but this app has no non-browser data route, so nothing proves the household's data survived it
crafty-controller 4.10.7 → 4.11.0 the update reached done, but this app has no non-browser data route, so nothing proves the household's data survived it
komga 1.25.0 → 1.27.1 the update reached done, but this app has no non-browser data route, so nothing proves the household's data survived it
outline 1.9.1 → 1.10.1 the update ended failed — fixture ran and found no non-browser seed route: sign-in requires an external identity provider (OIDC/Slack/Google); no local sign-up route exists
plex 1.41.4.9463-630c9f557 → 1.43.4.10903-e5521bd8c the update reached done, but this app has no non-browser data route, so nothing proves the household's data survived it
rallly 4.11.1 → 4.15.2 the update reached done, but this app has no non-browser data route, so nothing proves the household's data survived it

No upstream edge tonight, so nothing to propose (16): calcom, calibre-web, claper, gokapi, gramps-web, homebox, homepage, jellyfin, kimai, onlyoffice, paperless-ngx, plant-it, recipe-importer, seerr, sparkyfitness, wanderer.


Teardown — three layers plus Gitea, every claim READ BACK

The machine (9202). controller.yaml restored from controller.yaml.pre-28; git.repo_url reads back as the live catalog with an empty token; and — the one that actually decides which remote is followed (R-615) — the cache reads origin https://gitea.dooplex.hu/admin/app-catalog-felhom.eu.git at 1ad1f34. No drill images. Disk 1.9 GB used of 32 GB, unchanged from the start.

Three things the product could NOT clear, and a shell had to. This is not tidy-up, it is the finding: the termix and gokapi containers left by R-633, and sparkyfitness's app.yaml left by R-634. gokapi was still Restarting two hours later. They were removed by name (docker rm -f termix gokapi, rm .../sparkyfitness/app.yaml) — never a prune. Afterwards 9202 runs exactly felhom-controller, filebrowser, traefik, and no app.yaml exists anywhere. A household has no shell. Evidence: teardown/manual-cleanup.txt.

The host (demo-hp). pct list before and after: 9201 demo-hp and 9202 demo-hp-scratch, both running, unchanged. pvesm status unchanged but for expected scratch growth (nvme-scratch 5.19% → 6.95%). Guest 9201 was never touched.

The hub. Nothing provisioned, nothing changed. 9202 runs hub.enabled: false (R-620).

Gitea. The drill repo is reset to the live main (1ad1f34b6e51). git diff of the live catalog's templates/ against the night's baseline: 0 lines, and image: lines changed: NONE.

The fences, each read back rather than asserted:

fence at the start at the end
live catalog origin/main 1ad1f34b6e51 1ad1f34b6e51
demo-hp guest 9201 catalog cache 1ad1f34, live remote 1ad1f34, live remote
demo-felhom guest 9201 catalog cache 1ad1f34, live remote 1ad1f34, live remote
drill repo CI jobs 47 47 — no run, no mail, all night (R-629 holds)

Peti's box is parked and received nothing. Nothing ran on DooPlex beyond ordinary pushes, and nothing on ep0. felhom-controller, felhom-agent and the hub were read only — no product code was written. No golden, no bake, no vouch, no --no-verify, no branch. local-lvm untouched, no prune, tester-1 never reset, drill-r50 untouched.