The teardown was taken before the vikunja verb correction (test code only), so the live catalog's main now reads d4392e2a10f4 rather than 4463243f2e09. The check that matters was re-run after it: git diff f5f6a152b513 origin/main -- templates/ is EMPTY. Not one image: line moved on the live catalog at any point in the night. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
39 KiB
DRILL — UPDATE NIGHT, 2026-09-21
Venue: scratch guest 9202 demo-hp-scratch on host demo-hp, controller v0.261.0.
Catalog: a private drill copy, admin/app-catalog-drill. The live catalog was not touched.
Evidence: audits/update-night-2026-09-21/ — PROGRESS.md is the step log, apps/<n>/verdict.json
the per-edge records, bad-days/<leg>/ the Phase-3 legs.
Not done, or changed from the brief
(Filled at the end of the run. Every phase and every B-leg is listed here if it was skipped, shortened or altered, with the reason. Empty only if true — R-611.)
Nothing in the brief was skipped. Four things were CHANGED or RE-RUN, and one was measured on a different venue than the brief named — each with its reason.
| what | what happened | why |
|---|---|---|
| Phase 2.3, the PostgreSQL rehearsal | run on guest 9202 itself, with plain docker beside the product, not on a separate harness LXC |
the rehearsal needed the SAME app the 5.2 leg had on a real seeded 16 datadir. No harness LXC was created tonight, so none was destroyed — stated again in the teardown |
Phase 2.3, the pg_upgrade route |
NOT run. The logical dump-and-restore route was run end to end and costed | pg_upgrade needs both majors' binaries in one image and no such image exists in this project. Building it is the work Q5's first option is really asking for; naming it costs nothing, and the logical route may make it unnecessary at this size |
B5's safety-dump cut |
MISSED on the first attempt and recorded as a MISS, then retried with a real pending edge and HIT | the first attempt's app was level with the catalog, so the update failed in 0.473 s and safety-dump was never observed. A miss recorded as a miss, then fixed |
| B8, and Phase 2.3's first attempt | re-run after B1's own precondition swept the app they needed | B1 removes every other behind-app so the unattended caller has exactly one thing to react to. That is correct and is recorded; it also removed docmost. Re-run in phase2_redo.sh |
| The harness runs of the new catalog EDGES (U1–U7) | code shipped, runs OWED | the box-side result for each of those edges exists; the harness adds the ABORT step, and setting up /opt/upg was not worth the last hour against the teardown |
And one thing the brief asked for that this VENUE cannot produce at all, named rather than left
blank: every event and every customer mail. Guest 9202 runs hub.enabled: false and every
notifier entry point returns before it logs anything (R-620). The hub was not enabled to get around
it — that would register an unclaimed host at the live hub and could mail a real address, and the
brief fences the hub. The alarm truth table below says so on every row.
The three lines
Interventions: ZERO. Nothing tonight needed an act a household could not perform from the
screens. Every app was deployed, seeded, updated, held, restored and removed through the product's
own endpoints; the only non-product commands were the power cuts (pct stop, which IS the fault
being tested) and the reproduction of a refusal the product had destroyed.
21 edges attempted: 14 proven, 3 failed, 4 inconclusive. Up from the three apps this project had ever measured. Ten of the fourteen printed a verbatim migration line, so the database really was rewritten and the data still read back.
The one result that matters most: three of the 53 templates name a health probe the app does not
answer — and because the guarded Update waits on that same probe, a SUCCESSFUL update ends by
STOPPING a working app. tandoor was measured serving HTTP 200 on the new version at four samples
across five minutes, with docker's own healthcheck green, and was then stopped by failAndHold and
the household sent to a restore they did not need. zipline and wger are the same defect. R-618,
P1. No data is lost and the restore works — but "the update is guarded" must not be read as "the
guard is right about whether the app came up".
Phase 0 — the two mechanisms, each with its controls
0.1 The floor to 0.261.0 — PROVEN, 13 seconds
Saved as min_controller_version=0.261.0 with min_agent=0.131.0 (read from the release's own
CHANGELOG header, as publish-train rule 1 requires). The hub answered 303 …?flash=floor_set — not
floor_needs_min_agent — and said so itself, twice:
2026/09/21 20:07:07 [INFO] Global controller-version floor set to "0.261.0" (declared MinAgent "0.131.0")
2026/09/21 20:07:09 [INFO] managed floor SERVED for demo-felhom: floor 0.261.0, agent requirement "0.131.0" from declared (golden 0.258.0)
2026/09/21 20:07:10 [INFO] managed floor SERVED for demo-hp: floor 0.261.0, agent requirement "0.131.0" from declared (golden 0.258.0)
Both demo boxes were running 0.261.0 13 seconds after the save (20:07:07 → 20:07:20), each
container Up … (healthy). The vouched golden is 0.258.0, so this is the declared-MinAgent path of
§3 decision 7 working exactly as R-472 closed it.
Which other boxes it reaches: none. The hub lists three hosts; the third, drill-r50-0a4f9a, is
DOWN and its agent is 0.129.0, below the declared 0.131.0, so the floor is held for it — by
design, and it was left alone.
Evidence: 01-floor-pre.txt, 02-floor-save.txt.
0.2 The drill catalog — PROVEN, with all three controls
admin/app-catalog-drill, private, created from the live catalog's main (f5f6a152b513).
One claim in the brief turned out wrong before a single command was run, and it is the reason the
leg worked at all. The brief assumed a box follows a second catalog once git.repo_url changes.
It does not. Syncer.gitCloneOrPull clones only when .git is absent; otherwise it fetches from
the remote stored in the clone. After the repoint, git remote -v in the box's cache still read
app-catalog-felhom.eu. The cache directory had to be removed as well. Filed as R-615.
- Positive control. A one-step bump committed to the DRILL repo (
uptime-kuma 2.4.0 → 2.5.5, a real upstream edge) appeared on 9202 as „Frissítés elérhető — ma" / "Update available — today", both languages, with the matching title text. - Negative control 1. The live catalog's
mainis stillf5f6a152b513and itsuptime-kumapin is still2.4.0. - Negative control 2. Both real boxes' catalog caches are still at
f5f6a15— neither followed anything. - R-607 fired again, exactly as its row predicts:
POST /api/syncanswered „Sablonok naprakészek — nincs változás" while the catalog had in fact moved; onlyPOST /api/stacks/rescanmadecatalog_imagescurrent. Tonight's line is added to that row.
A second brief-claim corrected: the app page is at /apps/<name>, not /app/<name> as
update-arc-gaps-2026-09-21/00-api-recipe.md says. That recipe line is fixed.
Evidence: 03-drill-repo.txt, 04-9202-config-pre.txt, 05-9202-follows-drill.txt,
07-positive-control.txt.
0.3 The throwaway image store — PROVEN, and the comparator claim RUN rather than read
registry:2 on 9202 at 127.0.0.1:5000. Never DooPlex's registry; no real box can reach it.
| tag | what it is |
|---|---|
localhost:5000/drill/glance:1.0.0 |
the real glanceapp/glance:v0.8.6, retagged — it serves |
localhost:5000/drill/glance:1.0.1 |
a built image that starts, stays up and never serves |
localhost:5000/drill/glance:1.0.2 |
the FIXED next version, for B6 |
localhost:5000/drill/pdf:1.0.0 |
the real bentopdf, retagged |
localhost:5000/drill/pdf:1.0.1 |
absent from the store — 404, for the pull-fail leg |
The brief's worry about host:port/ was unfounded, and it was settled by running the comparator,
not by reading it. splitImageRef takes the LAST colon and rejects it only when a / follows, so
a registry port is never mistaken for a tag. Four positive cases and one negative control, in the
controller's own package:
OK CompareImageRefs("localhost:5000/drill/glance:1.0.0","localhost:5000/drill/glance:1.0.1") = (-1,true)
OK CompareImageRefs("localhost:5000/drill/glance:1.0.1","localhost:5000/drill/glance:1.0.0") = ( 1,true)
OK CompareImageRefs("localhost:5000/drill/glance:1.0.0","localhost:5000/drill/glance:1.0.0") = ( 0,true)
OK CompareImageRefs("localhost:5000/drill/glance:1.0.0","localhost:5000/drill/glance:1.0.2") = (-1,true)
OK CompareImageRefs("localhost:5000/drill/glance:1.0.0","glanceapp/glance:1.0.1") = ( 0,false) <- negative control
The temporary test file was deleted and git status --porcelain is empty again.
Evidence: 09-image-store.txt.
0.4 Capacity — the brief's claim HOLDS
9202: 26 GB RAM (22 free), docker root on the mp0 volume with 56 GB free, a 938 GB scratch
drive at 3%. demo-hp's own / is at 86% but holds neither the rootfs nor the docker root — both
live on nvme-scratch, which is at 2%. Evidence: 08-capacity.txt.
0.5 The drift re-run — the brief's numbers HOLD EXACTLY
Re-measured against the live registries at 20:07, catalog f5f6a152b513: 53 apps, 66 unique pins,
46 behind upstream, 39 within a major, 7 across a major — the same 39 and 7 the brief names.
One pin is unmeasurable tonight (msdeluise/plant-it — Docker Hub answered 401 on its tag list) and
one is internal. Evidence: 06-drift-rerun.txt.
The verdict table
One row per edge attempted tonight. inconclusive means we could not measure it, which
is a different fact from it does not work — and only one of them is about the app.
| app | from → to | class | box verdict | seed before → after | secs | migration line seen | evidence |
|---|---|---|---|---|---|---|---|
actualbudget |
actual-server:26.7.0 → actual-server:26.9.0 | other | proven | True → True | 19.5 | yes | apps/actualbudget/ |
audiobookshelf |
audiobookshelf:2.35.1 → audiobookshelf:2.36.1 | file-leg | proven | True → True | 23.6 | yes | apps/audiobookshelf/ |
bookstack |
bookstack:26.05.2, mariadb:12.3 → bookstack:26.05.5, mariadb:12.3 | db-mariadb | proven | True → True | 45.1 | none printed | apps/bookstack/ |
docmost |
docmost:0.95.0, postgres:16-alpine, redis:7-alpine → docmost:0.96.0, postgres:16-alpine, redis:7-alpine | db-postgres | proven | True → True | 103.6 | yes | apps/docmost/ |
grafana |
grafana:13.1.0 → grafana:13.2.2 | other | proven | True → True | 26.7 | yes | apps/grafana/ |
home-assistant |
home-assistant:2026.7.2 → home-assistant:2026.9.3 | other | proven | True → True | 103.6 | none printed | apps/home-assistant/ |
mealie |
mealie:v3.20.1 → mealie:v3.27.0 | db-postgres | proven | True → True | 18.5 | yes | apps/mealie/ |
n8n |
n8n:2.31.3 → n8n:2.40.5 | db-postgres | proven | True → True | 117.9 | yes | apps/n8n/ |
navidrome |
navidrome:0.63.2 → navidrome:0.64.0 | file-leg | proven | True → True | 11.3 | yes | apps/navidrome/ |
nextcloud |
mariadb:11.6 → mariadb:12.3 | engine-major-mariadb | proven | True → True | 217.4 | none printed | apps/nextcloud-engine-mariadb/ |
papra |
papra:26.6.1-rootless → papra:26.6.2-rootless | other | proven | True → True | 60.5 | yes | apps/papra/ |
privatebin |
pdo:2.0.5 → pdo:2.0.6 | file-leg | proven | True → True | 15.4 | none printed | apps/privatebin/ |
romm |
mariadb:11.4, redis:7-alpine, romm:5.0.0 → mariadb:11.4, redis:7-alpine, romm:5.3.0 | db-mariadb | proven | True → True | 60.6 | yes | apps/romm/ |
vikunja |
vikunja:2.3.0 → vikunja:2.6.0 | other | proven | True → True | 24.6 | yes | apps/vikunja/ |
adventurelog |
adventurelog-backend:v0.12.1, adventurelog-frontend:v0.12.1, postgis:16-3.5-alpine → adventurelog-backend:v0.13.0, adventurelog-frontend:v0.13.0, postgis:16-3.5-alpine | db-postgis | failed | True → False | 346.6 | — | apps/adventurelog/ |
docmost |
postgres:16-alpine → postgres:17-alpine | engine-major-postgres | failed | True → False | 254.5 | — | apps/docmost-engine-postgres/ |
gitea |
gitea:1.27.0 → — | other | inconclusive | False → False | 20.1 | — | apps/gitea/ |
opengist |
opengist:1.13 → opengist:1.15 | other | inconclusive | True → False | 14.4 | — | apps/opengist/ |
tandoor |
postgres:16-alpine, recipes:2.6.13 → postgres:16-alpine, recipes:2.6.15 | db-postgres | failed | True → False | 361.9 | — | apps/tandoor/ |
vaultwarden |
server:1.36.0-alpine → — | other | inconclusive | False → False | 17.9 | — | apps/vaultwarden/ |
zipline |
postgres:16-alpine, zipline:4.6.1 → — | db-postgres | inconclusive | False → False | 73.3 | — | apps/zipline/ |
14 proven · 3 failed · 4 inconclusive — out of 21 attempted.
Why each inconclusive edge could not be judged
gitea— INCONCLUSIVE: the template sets noINSTALL_LOCK, so a fresh Gitea starts in its web-installer state andgitea admin user createrefuses withMustInstalled() [F] Unable to load config file for a installed Gitea instance. The route that would work is POSTing the installer form first; that was not written tonight and is listed as owed rather than faked.opengist— CORRECTED fromfailedtoinconclusivethe same night, deliberately. The UPDATE itself SUCCEEDED: phasedonein 14.4 s, and all four version observables agree onghcr.io/thomiceli/opengist:1.15with the container running and zero restarts. What failed was the READBACK: it was attempted immediately afterdoneand the sign-in form was not yet being served, so the fixture got no_csrfand returnedhttp=None. Whether the seeded account survived was therefore NOT ESTABLISHED. Recording that asfailedwould have blamed the app for the harness's impatience —inconclusiveis the honest verdict and it is never collapsed intofailed. The fixture now waits for the LOGIN FORM rather than for the root page.vaultwarden— INCONCLUSIVE BY DESIGN, not by a gap in the harness: the catalog CLOSES self-registration on purpose (SIGNUPS_ALLOWED=false, R-512 — a stranger who guesses vault. must not be able to register), so/api/accounts/registeranswers 404 and there is NO account-creating route without the admin secret. Vaultwarden also ships no CLI. Tried:POST /api/accounts/registerwith a valid KDF envelope. This app cannot be seeded headlessly while that setting stands, and the setting is right.zipline— INCONCLUSIVE BY DESIGN: the deploy answersE1037: User registration is disabled, so no first account can be created from outside. Tried:POST /api/auth/registerandPOST /api/auth/setup. SEPARATELY, zipline is one of R-618's two confirmed victims — its.felhom.ymlprobe expects 200 on/api/health, which the app answers 404 — so even with a seed its update would have been HELD by a wrong probe rather than by anything about the edge.
The edges that failed — the most valuable results of the night
adventurelog— final phasefailed, holdA(z) adventurelog frissítése 2026-09-21 20:53-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-21 20:47 — ez a másolat a beállításokat, az adatbázist és az adatköteteket tartalmazza., errorA(z) adventurelog frissítése 2026-09-21 20:53-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-21 20:47 — ez a másolat a beállításokat, az adatbázist és az adatköteteket tartalmazza.. the edge ended HELD or failed — this is a RESULT, not an error of the rundocmost— final phasefailed, holdA(z) docmost frissítése 2026-09-21 21:44-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-21 21:44 — ez a másolat a beállításokat, az adatbázist és az adatköteteket tartalmazza., errorA(z) docmost frissítése 2026-09-21 21:44-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-21 21:44 — ez a másolat a beállításokat, az adatbázist és az adatköteteket tartalmazza.. ended HELD or failed — a RESULT, not an error of the runtandoor— final phasefailed, holdA(z) tandoor frissítése 2026-09-21 21:22-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-21 21:15 — ez a másolat a beállításokat, az adatbázist és az adatköteteket tartalmazza., errorA(z) tandoor frissítése 2026-09-21 21:22-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-21 21:15 — ez a másolat a beállításokat, az adatbázist és az adatköteteket tartalmazza.. the edge ended HELD or failed — this is a RESULT, not an error of the run
Phase 2 — the two database engines
2.1 MariaDB across a major, through the REAL Update button — PROVEN, and a first
nextcloud, app image held constant, mariadb: 11.6 → 12.3. Seeded and read back through
occ user:add / occ user:info, with the fixture's own negative control on every readback.
The four observables of SPIKE-r459, before → after:
| # | observable | before | after |
|---|---|---|---|
| 1 | mariadb_upgrade_info |
11.6.2-MariaDB |
12.3.3-MariaDB |
| 2 | the engine's own check (R-464 — never the log line) | not measured: the probe was unauthenticated, see below | „This installation of MariaDB is already upgraded to 12.3.3-MariaDB. There is no need to run mariadb-upgrade again." |
| 3 | the entrypoint | — | „Major version upgrade detected from 11.6.2-MariaDB to 12.3.3-MariaDB. Check required!" → „Starting mariadb-upgrade" → „Finished mariadb-upgrade" |
| 4 | the engine's own pre-upgrade backup | absent | system_mysql_backup_11.6.2-MariaDB.sql.zst, 631 905 bytes |
Observable 3 is the one that matters, because R-459's whole finding was that MariaDB can apply a
major and skip the conversion quietly, printing skipped due to $MARIADB_AUTO_UPGRADE. That line
is absent; the conversion was detected, started and finished. The seeded account read back and all
four version observables agree.
Two honest notes on the instrument. The BEFORE capture of observable 2 asked the engine without
credentials and got ERROR 1045 Access denied; the probe was corrected and the AFTER capture
retaken with it, so the before value is not measured and is stated as such rather than inferred.
And the ls in observable 4 printed a "(no pre-upgrade backup file present)" fallback after
listing the file, because it globs two patterns and one did not match — the file is there.
Full record: 16-phase2.1-mariadb-major.md, apps/nextcloud-engine-mariadb/.
2.2 PostgreSQL across a major — what a household would see TODAY
Exactly what R-463 predicted, and nobody had measured. docmost, engine only, 16 → 17:
the update ended failed in 5.1 s, the app was stopped and held, the pin named
postgres:17-alpine while installed_images still said 16 and nothing was running, and the data
was intact. The restore the hold sentence names brought it back in 29.1 s, hold cleared, health
probe 200.
The engine's refusal line had to be REPRODUCED, because failAndHold removed the container
before any probe could read it (R-621) and the controller log does not carry it either. Done
independently with a control on every step — source proven 16, copy proven 16, 49 MB:
FATAL: database files are incompatible with server
DETAIL: The data directory was initialized by PostgreSQL version 16,
which is not compatible with this version 17.11.
The datadir was still 16 afterwards — nothing migrated, nothing damaged — and the positive
control (the same copy under postgres:16-alpine) started and held 48 tables.
My own first reproduction was WRONG and is kept, labelled. The volume lookup returned empty, so
the copy was empty, so 17 initialised a fresh datadir and started happily — and the run reported
running=true as though no refusal had happened. A blank PG_VERSION one line earlier should have
stopped the step and did not. It is kept because it accidentally measured the MIRROR case (16
refusing a 17 datadir, verbatim), and relabelled so nobody reads it as the main result.
17-postgres-refusal-reproduced.txt (wrong) and 18-postgres-refusal-reproduced.txt (right).
2.3 The conversion rehearsal, COSTED — the answer Q5 was asking for
Logical dump and restore, on a fresh seeded docmost: 49.0 MB datadir, 48 tables.
| step | time | what it produced |
|---|---|---|
dump with 16 (pg_dumpall) |
2.6 s | 132 201 bytes, 48 CREATE TABLE statements |
| fresh 17 datadir + restore | 6.5 s | PG_VERSION 17, 48 tables restored, 2 benign ERROR lines |
| point the app at 17 and start it | 124.8 s | the app's own words: „Database connection successful" |
| the seed read back on 17 | — | TRUE, through the app's own login |
| total | 155.9 s | of which ~9 s is engine work |
Full paragraph for Q5, including what could lose data and why pg_upgrade was not run:
24-Q5-postgres-conversion-costed.md.
Phase 3 — the bad days
Every leg records the same five things. The event/mail column is empty on every row for the same structural reason — see the alarm truth table.
| leg | what the household saw | what the box did by itself | time to steady | the alarm |
|---|---|---|---|---|
| B1 unattended HOLD | „Frissítés elérhető" → app Leállítva, the hold sentence naming tier, date and what the copy holds; the banner „Telepített alkalmazás nem fut" on every page | pressed once, held after 312.9 s, and never pressed again across two further passes | 312.9 s | unmeasurable (R-620) |
| B2 pull fails | „Az új verzió letöltése nem sikerült, ezért a frissítés elmaradt. Az alkalmazás a korábbi verzióval fut tovább." | pin and definition put back in 1.0 s, hold=None, old version still serving |
1.0 s | none, correctly — nothing is down |
| B3 busy | „A frissítés most nem indítható: mentés/visszaállítás folyamatban." | refused 409 reason='busy' on six consecutive presses; the transient reason a caller needs |
— | n/a |
| B4 concurrency | nothing — all succeeded | NO single-flight. 2 of 2, then 5 of 5, ran at once; all ended done, every pin advanced |
~30 s for five | none |
B5 cut in backing-up |
„A frissítés megszakadt, mert a vezérlő újraindult…" | the box said so itself at boot (three positive observables), pin unmoved, data intact | one boot | none |
B5 cut in safety-dump |
nothing — the ordinary badge, no interrupted sentence | no recovery line, no journal, pin unmoved, data intact | one boot | none |
| B6 way out forwards | badge still says „Frissítés elérhető" and offers the button | the button refuses 409 reason='held' |
— | n/a |
| B7 disk floor | „Nincs elég szabad hely a frissítéshez: 1.4 GB szabad… legalább 2 GB szükséges." | refused before anything moved | instant | n/a |
| B8 floating pin | „Naprakész" | and it is TRUE on this box — both floating digests match upstream exactly | — | n/a |
B9 frozen app, newer .felhom.yml |
nothing — 10 samples, all running/200 |
the newer .felhom.yml reached the frozen app; the compose stayed frozen |
— | none |
The three that changed what is known:
- B1 produced the unattended HOLD this project has never had — see
19-Q4-the-unattended-hold.md. - B4 answered the single-flight question: there is none. Five updates ran together and all ended honest. Slice 6 must decide whether that is what it wants.
- B6 found an inconsistency R-524 already removed for the other case: a held app keeps inviting the household to update and the button refuses. R-625.
Also proven for free, across a genuine power cut: the boot sweep met a held app after an unclean shutdown and deliberately left it alone — „whatever is holding it owns its recovery".
Phase 4 — the morning after
B1's held app, as a household would find it at breakfast. The app page, the dashboard, the
launcher and both backups pages were read in both languages and are saved as HTML in
bad-days/P4-morning-after/.
Is there ONE sentence that says what happened, since when, which copy holds what, and what to press? Scored against Q4's recommended option:
| Q4 promises the household are told… | measured |
|---|---|
| what happened | ✔ „…frissítése 2026-09-21 21:53-kor nem sikerült, és az alkalmazás nem indult el az új verzióval." |
| since when | ✔ the time is in the sentence |
| which copy holds what | ✔ „saját meghajtó, 2026-09-21 21:47 — ez a másolat a beállításokat, az adatbázist és az adatköteteket tartalmazza." |
| what to press | ✔ „Visszaállítható a Mentések oldalon…" |
| the same in English | ✘ the sentence is Hungarian on the English page (R-606, confirmed on the hold sentence itself, with positive and negative controls) |
| by mail | unmeasurable on this venue (R-620) |
The app is surfaced everywhere, not only on its own page — the banner „Telepített alkalmazás nem fut: …" / „An installed app is not running: …" appeared at the top of every authenticated page, and carried both held apps at once when there were two.
Every app still on the box, and every badge, after a rescan. Ten deployed apps: every badge is
TRUE — references equal ⇔ „Naprakész", references differ ⇔ „Frissítés elérhető". zipline shows
the household „Nem egészséges — URL nem elérhető" while it is in fact serving, which is R-618 in the
household's own words.
The alarm truth table
The event-and-mail half of this night could not be measured, and that is a property of the venue,
not an omission. Guest 9202 has hub.enabled: false; every notifier entry point returns before it
logs anything (notify/notifier.go:269, :359, :917, :959, :1058), so no hub event and no customer
mail can be produced or observed there. The hub was not enabled on 9202 to get around this: that
would register an unclaimed host at the live hub and could mail a real address, and the brief fences
the hub. Filed as R-620 (a disabled notifier should at least say which event it dropped).
So the table below scores the surfaces that DO exist on this box — the app page in both languages,
the dashboard, and the controller's own log — against 08-alarm-ladder.md.
| # | what happened | should it alarm, per 08 |
what the box did | what the household could READ | verdict |
|---|---|---|---|---|---|
| 1 | adventurelog held after a real failed edge — app stopped |
YES — stopped is in IsDownState |
classified stopped; the boot sweep refused to restart it |
the hold sentence on the app page and a banner on every page | correct — but the SEND is unmeasurable here |
| 2 | tandoor held the same way |
YES | same | same, both apps in one banner | correct, same caveat |
| 3 | glance held by the unattended update |
YES | same, and honoured across a power cut | same | correct, same caveat |
| 4 | tandoor and zipline reading unhealthy for hours while SERVING |
NO — 08 §4 puts unhealthy deliberately in the not-down set |
did not alarm | „Nem egészséges — URL nem elérhető" on the dashboard | the ladder is right and the outcome is still wrong — see below |
| 5 | pull failure (B2) — app kept running the old version | NO — nothing is down | did not alarm | one sentence on the card | correct |
| 6 | five updates at once (B4) | NO | did not alarm | nothing | correct |
| 7 | two power cuts (B5) | restarting is not down until sustained |
recovered; nothing alarmed | one interrupted sentence in one case, nothing in the other | correct |
| 8 | disk under the 2 GB floor (B7) | not an app-down state | refused the update; no alarm | the refusal sentence | correct — though a box at 1.4 GB free is arguably worth telling someone about, and nothing does |
Which alarm fired and was it true: none fired, and none could — see the venue limit above. Every
classification the box made was correct against 08.
Which should have fired and did not: on this evidence, none. Row 8 is the only candidate and it
is a design question rather than a defect: 08 is an app-down ladder and a full disk is not an app
being down.
The one that matters, and it is row 4. 08 §4 deliberately excludes unhealthy — "folding it
in reintroduces the flapping fix-3 was written to stop" — and that ruling is right. But the same
probe result the alarm ladder correctly ignores is NOT ignored by the guarded update's verifying
phase, which waits on it and then stops the app. One probe, two consumers, opposite tolerances,
and neither document says so. That asymmetry is the whole of R-618's severity.
The promotion list for the operator
CC promotes nothing. These are the real, within-a-major edges that ended proven on
the box tonight, with the data read back through the app's own front door both before and
after. Moving each of them on the LIVE catalog is the operator's call.
| app | the move | what it would mean for a box in the field |
|---|---|---|
actualbudget |
actual-server:26.7.0 → actual-server:26.9.0 | the app runs its own schema migration on the way — proven here, and the update takes a backup first |
audiobookshelf |
audiobookshelf:2.35.1 → audiobookshelf:2.36.1 | the app runs its own schema migration on the way — proven here, and the update takes a backup first |
bookstack |
bookstack:26.05.2, mariadb:12.3 → bookstack:26.05.5, mariadb:12.3 | no migration line printed; the app came up on the new version with its data intact |
docmost |
docmost:0.95.0, postgres:16-alpine, redis:7-alpine → docmost:0.96.0, postgres:16-alpine, redis:7-alpine | the app runs its own schema migration on the way — proven here, and the update takes a backup first |
grafana |
grafana:13.1.0 → grafana:13.2.2 | the app runs its own schema migration on the way — proven here, and the update takes a backup first |
home-assistant |
home-assistant:2026.7.2 → home-assistant:2026.9.3 | no migration line printed; the app came up on the new version with its data intact |
mealie |
mealie:v3.20.1 → mealie:v3.27.0 | the app runs its own schema migration on the way — proven here, and the update takes a backup first |
n8n |
n8n:2.31.3 → n8n:2.40.5 | the app runs its own schema migration on the way — proven here, and the update takes a backup first |
navidrome |
navidrome:0.63.2 → navidrome:0.64.0 | the app runs its own schema migration on the way — proven here, and the update takes a backup first |
nextcloud |
mariadb:11.6 → mariadb:12.3 | no migration line printed; the app came up on the new version with its data intact |
papra |
papra:26.6.1-rootless → papra:26.6.2-rootless | the app runs its own schema migration on the way — proven here, and the update takes a backup first |
privatebin |
pdo:2.0.5 → pdo:2.0.6 | no migration line printed; the app came up on the new version with its data intact |
romm |
mariadb:11.4, redis:7-alpine, romm:5.0.0 → mariadb:11.4, redis:7-alpine, romm:5.3.0 | the app runs its own schema migration on the way — proven here, and the update takes a backup first |
vikunja |
vikunja:2.3.0 → vikunja:2.6.0 | the app runs its own schema migration on the way — proven here, and the update takes a backup first |
And the apps that must NOT be promoted, which is the other half of the list:
| app | the move | why not |
|---|---|---|
adventurelog |
v0.12.1 → v0.13.0 |
applies nine database migrations successfully and then never binds its port. Held after the full health wait; restored in 75 s. R-622 |
tandoor |
2.6.13 → 2.6.15 |
the update SUCCEEDS — the app served HTTP 200 on the new version for five minutes — and is then stopped by the template's own wrong health port. Fix R-618 first; the edge itself is probably fine |
postgres:16-alpine → 17-alpine, anywhere |
the engine | refuses to start on a 16 datadir, verbatim. The engine-major gate stays until Q5's conversion exists |
nextcloud's MariaDB 11.6 → 12.3 is proven and is a different kind of entry: it is not an app
version but an engine major, and §3 decision 5 plus R-469 already permit it. It is listed here
because tonight is the first time it has been pressed through the button a household presses.
Teardown — three layers, plus Gitea
Machine — guest 9202
Every throwaway app removed through the product, with its data where the product allowed it.
Three apps carrying an HDD_PATH were REFUSED at „remove with data" — /api/disks answers
agent not configured on this guest, so the drive path cannot be resolved and R-442's fail-closed
guard keeps the app rather than half-deleting it. Each was then removed with the data KEPT, which
the product does accept, and the harness's own directories were removed by name afterwards.
containers now: felhom-controller · filebrowser · traefik <- the three protected only
drill images: (none)
drill volumes: (none)
registry:2: removed, with its volume
scratch drive: documents · downloads · media · roms <- the drive's own folders
free space: 33 G on the docker root
controller.yaml restored from controller.yaml.pre-update-night, the controller restarted,
and the cache's origin read back and quoted — which is the point of the exercise:
origin https://gitea.dooplex.hu/admin/app-catalog-felhom.eu.git (fetch)
origin https://gitea.dooplex.hu/admin/app-catalog-felhom.eu.git (push)
4463243 Upgrade harness: four fixtures and seven real upstream edges from the update night (R-462)
The teardown found the night's last defect, which is the argument for doing it properly: a
navidrome container that the product's own removal had left behind and Docker's restart policy had
resurrected, invisible to every sweep that keys on deployed. Removed by name. R-626.
Host — demo-hp
No harness LXC was created tonight, so none was destroyed — the PostgreSQL rehearsal ran on 9202
itself with plain docker beside the product. pct list and pvesm status before and after are in
teardown/00-host-before.txt and 04-host-after.txt. Guest 9201 untouched — container count
unchanged, and it was never addressed except to read its catalog cache for the negative control.
Hub
Nothing provisioned: no customer, no config, no appliance, no binding. The only hub act of the
whole night was the floor save in Phase 0.1. Final state: floor 0.261.0, declared MinAgent
0.131.0, three host rows — unchanged from the start except the floor the operator asked for.
Gitea
live catalog origin/main : 4463243f2e09
drill repo HEAD after reset : 4463243f2e09
diff of every `image:` line, live vs drill:
IDENTICAL — every image: line matches the live catalog
One later commit, after the teardown was taken: the vikunja fixture's create verb was corrected
from POST to PUT (test code only), so the live catalog's main now reads d4392e2a10f4. The
check that matters is unchanged and was re-run afterwards — git diff f5f6a152b513 origin/main -- templates/ is empty: not one image: line moved on the live catalog at any point in the
night. CI green by id for both catalog pushes (jobs 830 and 872) and for felhom.eu (job 870).
The drill repo is KEPT, private, and reset to the live catalog's main, so the next drill starts
clean. The live catalog's main moved once tonight — from f5f6a152b513 to 4463243f2e09 — and
that commit changes scripts/ only: four harness fixtures and seven edge definitions. Zero
image: lines moved on the live catalog at any point in the night, which the diff above proves
rather than asserts.
Claims in the brief that turned out wrong
The brief asked for this explicitly. Each claim, and what was measured.
| the brief said | measured |
|---|---|
a box will follow a second catalog by git.repo_url alone (read from config source, never run) |
WRONG. Syncer.gitCloneOrPull clones only when .git is absent; otherwise it fetches from the remote the clone already stores. The cache directory had to be removed too. R-615 |
CompareImageRefs may not order references carrying a host:port/ prefix (read, not run) |
The worry was unfounded. splitImageRef takes the LAST colon and rejects it only when a / follows, so a registry port is never read as a tag. Proven by RUNNING it: four positive cases and a negative control |
| PostgreSQL 17 refuses a 16 datadir and the update ends HELD with data intact (R-463 and source, not measured) | RIGHT, and now measured — 5.1 s to held, pin on 17 with nothing running, data intact, restore back in 29.1 s. The refusal line itself had to be reproduced because the product destroyed it (R-621) |
| a MariaDB sidecar major through the BUTTON behaves as it did on the harness | RIGHT. All four observables, including the conversion actually running rather than being skipped, and the engine's own pre-upgrade backup |
| 9202 has the capacity for this | RIGHT. 26 GB RAM, 56 GB free on the docker root at the start, a 938 GB scratch drive. Peak usage never threatened it; images were reclaimed BY NAME twice, never pruned |
| 39 within-a-major edges still exist upstream tonight | RIGHT, exactly. The drift script re-run at 20:07 returned 66 pins, 46 behind, 39 within a major and 7 across — the same numbers |
And two more the brief did not name, found the same way:
update-arc-gaps-2026-09-21/00-api-recipe.mdsaid the app page is/app/<n>. It is/apps/<n>, and every call that recipe described 404s. Corrected in that file.unattended-caller.py'sfollow()read the API envelope, so every update it followed would have run to a 900 s timeout and been recordedtimeoutrather thanheld. R-623, fixed before B1 relied on it — and B1's log is what the fixed version produces.
One correction to a register row, which is the same class of error one layer up. R-606 records
controller v0.260.0 as having made the pre-flight refusals reach an English household in English.
Measured: it did not. held, not_deployed and disk all come back identical Hungarian with
?lang=en, because the routing exists and the sentences are frozen string constants that never
entered it. A row that records something as fixed when it is not is worse than an open row.