# DRILL — UPDATE NIGHT, 2026-09-21 **Venue:** scratch guest **9202 `demo-hp-scratch`** on host `demo-hp`, controller **v0.261.0**. **Catalog:** a private drill copy, `admin/app-catalog-drill`. **The live catalog was not touched.** **Evidence:** `audits/update-night-2026-09-21/` — `PROGRESS.md` is the step log, `apps//verdict.json` the per-edge records, `bad-days//` the Phase-3 legs. --- ## Not done, or changed from the brief *(Filled at the end of the run. Every phase and every B-leg is listed here if it was skipped, shortened or altered, with the reason. Empty only if true — R-611.)* **Nothing in the brief was skipped. Four things were CHANGED or RE-RUN, and one was measured on a different venue than the brief named — each with its reason.** | what | what happened | why | |---|---|---| | **Phase 2.3, the PostgreSQL rehearsal** | run on **guest 9202 itself**, with plain `docker` beside the product, not on a separate harness LXC | the rehearsal needed the SAME app the 5.2 leg had on a real seeded 16 datadir. No harness LXC was created tonight, so none was destroyed — stated again in the teardown | | **Phase 2.3, the `pg_upgrade` route** | **NOT run.** The logical dump-and-restore route was run end to end and costed | `pg_upgrade` needs both majors' binaries in one image and no such image exists in this project. Building it is the work Q5's first option is really asking for; naming it costs nothing, and the logical route may make it unnecessary at this size | | **B5's `safety-dump` cut** | **MISSED on the first attempt and recorded as a MISS**, then retried with a real pending edge and HIT | the first attempt's app was level with the catalog, so the update failed in 0.473 s and `safety-dump` was never observed. A miss recorded as a miss, then fixed | | **B8, and Phase 2.3's first attempt** | **re-run** after B1's own precondition swept the app they needed | B1 removes every other behind-app so the unattended caller has exactly one thing to react to. That is correct and is recorded; it also removed `docmost`. Re-run in `phase2_redo.sh` | | **The harness runs of the new catalog EDGES (U1–U7)** | **code shipped, runs OWED** | the box-side result for each of those edges exists; the harness adds the ABORT step, and setting up `/opt/upg` was not worth the last hour against the teardown | **And one thing the brief asked for that this VENUE cannot produce at all, named rather than left blank: every event and every customer mail.** Guest 9202 runs `hub.enabled: false` and every notifier entry point returns before it logs anything (R-620). The hub was not enabled to get around it — that would register an unclaimed host at the live hub and could mail a real address, and the brief fences the hub. The alarm truth table below says so on every row. --- ## The three lines **Interventions: ZERO.** Nothing tonight needed an act a household could not perform from the screens. Every app was deployed, seeded, updated, held, restored and removed through the product's own endpoints; the only non-product commands were the power cuts (`pct stop`, which IS the fault being tested) and the reproduction of a refusal the product had destroyed. **21 edges attempted: 14 proven, 3 failed, 4 inconclusive.** Up from the **three** apps this project had ever measured. Ten of the fourteen printed a verbatim migration line, so the database really was rewritten and the data still read back. **The one result that matters most: three of the 53 templates name a health probe the app does not answer — and because the guarded Update waits on that same probe, a SUCCESSFUL update ends by STOPPING a working app.** `tandoor` was measured serving HTTP 200 on the new version at four samples across five minutes, with docker's own healthcheck green, and was then stopped by `failAndHold` and the household sent to a restore they did not need. `zipline` and `wger` are the same defect. **R-618, P1.** No data is lost and the restore works — but "the update is guarded" must not be read as "the guard is right about whether the app came up". --- ## Phase 0 — the two mechanisms, each with its controls ### 0.1 The floor to 0.261.0 — PROVEN, 13 seconds Saved as `min_controller_version=0.261.0` with `min_agent=0.131.0` (read from the release's own CHANGELOG header, as publish-train rule 1 requires). The hub answered `303 …?flash=floor_set` — not `floor_needs_min_agent` — and said so itself, twice: ``` 2026/09/21 20:07:07 [INFO] Global controller-version floor set to "0.261.0" (declared MinAgent "0.131.0") 2026/09/21 20:07:09 [INFO] managed floor SERVED for demo-felhom: floor 0.261.0, agent requirement "0.131.0" from declared (golden 0.258.0) 2026/09/21 20:07:10 [INFO] managed floor SERVED for demo-hp: floor 0.261.0, agent requirement "0.131.0" from declared (golden 0.258.0) ``` Both demo boxes were running `0.261.0` **13 seconds after the save** (`20:07:07` → `20:07:20`), each container `Up … (healthy)`. The vouched golden is `0.258.0`, so this is the declared-MinAgent path of §3 decision 7 working exactly as R-472 closed it. **Which other boxes it reaches:** none. The hub lists three hosts; the third, `drill-r50-0a4f9a`, is `DOWN` and its agent is `0.129.0`, below the declared `0.131.0`, so the floor is held for it — by design, and it was left alone. Evidence: `01-floor-pre.txt`, `02-floor-save.txt`. ### 0.2 The drill catalog — PROVEN, with all three controls `admin/app-catalog-drill`, private, created from the live catalog's `main` (`f5f6a152b513`). **One claim in the brief turned out wrong before a single command was run, and it is the reason the leg worked at all.** The brief assumed a box follows a second catalog once `git.repo_url` changes. It does not. `Syncer.gitCloneOrPull` clones **only when `.git` is absent**; otherwise it fetches from the remote stored in the clone. After the repoint, `git remote -v` in the box's cache still read `app-catalog-felhom.eu`. The cache directory had to be removed as well. Filed as **R-615**. - **Positive control.** A one-step bump committed to the DRILL repo (`uptime-kuma 2.4.0 → 2.5.5`, a real upstream edge) appeared on 9202 as „**Frissítés elérhető — ma**" / "**Update available — today**", both languages, with the matching title text. - **Negative control 1.** The live catalog's `main` is still `f5f6a152b513` and its `uptime-kuma` pin is still `2.4.0`. - **Negative control 2.** Both real boxes' catalog caches are still at `f5f6a15` — neither followed anything. - **R-607 fired again**, exactly as its row predicts: `POST /api/sync` answered „Sablonok naprakészek — nincs változás" while the catalog had in fact moved; only `POST /api/stacks/rescan` made `catalog_images` current. Tonight's line is added to that row. **A second brief-claim corrected:** the app page is at **`/apps/`**, not `/app/` as `update-arc-gaps-2026-09-21/00-api-recipe.md` says. That recipe line is fixed. Evidence: `03-drill-repo.txt`, `04-9202-config-pre.txt`, `05-9202-follows-drill.txt`, `07-positive-control.txt`. ### 0.3 The throwaway image store — PROVEN, and the comparator claim RUN rather than read `registry:2` on 9202 at `127.0.0.1:5000`. Never DooPlex's registry; no real box can reach it. | tag | what it is | |---|---| | `localhost:5000/drill/glance:1.0.0` | the real `glanceapp/glance:v0.8.6`, retagged — it serves | | `localhost:5000/drill/glance:1.0.1` | a built image that starts, stays up and **never serves** | | `localhost:5000/drill/glance:1.0.2` | the FIXED next version, for B6 | | `localhost:5000/drill/pdf:1.0.0` | the real bentopdf, retagged | | `localhost:5000/drill/pdf:1.0.1` | **absent from the store** — 404, for the pull-fail leg | **The brief's worry about `host:port/` was unfounded, and it was settled by running the comparator, not by reading it.** `splitImageRef` takes the LAST colon and rejects it only when a `/` follows, so a registry port is never mistaken for a tag. Four positive cases and one negative control, in the controller's own package: ``` OK CompareImageRefs("localhost:5000/drill/glance:1.0.0","localhost:5000/drill/glance:1.0.1") = (-1,true) OK CompareImageRefs("localhost:5000/drill/glance:1.0.1","localhost:5000/drill/glance:1.0.0") = ( 1,true) OK CompareImageRefs("localhost:5000/drill/glance:1.0.0","localhost:5000/drill/glance:1.0.0") = ( 0,true) OK CompareImageRefs("localhost:5000/drill/glance:1.0.0","localhost:5000/drill/glance:1.0.2") = (-1,true) OK CompareImageRefs("localhost:5000/drill/glance:1.0.0","glanceapp/glance:1.0.1") = ( 0,false) <- negative control ``` The temporary test file was deleted and `git status --porcelain` is empty again. Evidence: `09-image-store.txt`. ### 0.4 Capacity — the brief's claim HOLDS 9202: **26 GB RAM** (22 free), docker root on the `mp0` volume with **56 GB free**, a 938 GB scratch drive at 3%. `demo-hp`'s own `/` is at 86% but holds neither the rootfs nor the docker root — both live on `nvme-scratch`, which is at 2%. Evidence: `08-capacity.txt`. ### 0.5 The drift re-run — the brief's numbers HOLD EXACTLY Re-measured against the live registries at 20:07, catalog `f5f6a152b513`: 53 apps, 66 unique pins, **46 behind upstream, 39 within a major, 7 across a major** — the same 39 and 7 the brief names. One pin is unmeasurable tonight (`msdeluise/plant-it` — Docker Hub answered 401 on its tag list) and one is internal. Evidence: `06-drift-rerun.txt`. --- ## The verdict table One row per edge attempted tonight. `inconclusive` means *we could not measure it*, which is a different fact from *it does not work* — and only one of them is about the app. | app | from → to | class | box verdict | seed before → after | secs | migration line seen | evidence | |---|---|---|---|---|---|---|---| | `actualbudget` | actual-server:26.7.0 → actual-server:26.9.0 | other | **proven** | True → True | 19.5 | yes | `apps/actualbudget/` | | `audiobookshelf` | audiobookshelf:2.35.1 → audiobookshelf:2.36.1 | file-leg | **proven** | True → True | 23.6 | yes | `apps/audiobookshelf/` | | `bookstack` | bookstack:26.05.2, mariadb:12.3 → bookstack:26.05.5, mariadb:12.3 | db-mariadb | **proven** | True → True | 45.1 | none printed | `apps/bookstack/` | | `docmost` | docmost:0.95.0, postgres:16-alpine, redis:7-alpine → docmost:0.96.0, postgres:16-alpine, redis:7-alpine | db-postgres | **proven** | True → True | 103.6 | yes | `apps/docmost/` | | `grafana` | grafana:13.1.0 → grafana:13.2.2 | other | **proven** | True → True | 26.7 | yes | `apps/grafana/` | | `home-assistant` | home-assistant:2026.7.2 → home-assistant:2026.9.3 | other | **proven** | True → True | 103.6 | none printed | `apps/home-assistant/` | | `mealie` | mealie:v3.20.1 → mealie:v3.27.0 | db-postgres | **proven** | True → True | 18.5 | yes | `apps/mealie/` | | `n8n` | n8n:2.31.3 → n8n:2.40.5 | db-postgres | **proven** | True → True | 117.9 | yes | `apps/n8n/` | | `navidrome` | navidrome:0.63.2 → navidrome:0.64.0 | file-leg | **proven** | True → True | 11.3 | yes | `apps/navidrome/` | | `nextcloud` | mariadb:11.6 → mariadb:12.3 | engine-major-mariadb | **proven** | True → True | 217.4 | none printed | `apps/nextcloud-engine-mariadb/` | | `papra` | papra:26.6.1-rootless → papra:26.6.2-rootless | other | **proven** | True → True | 60.5 | yes | `apps/papra/` | | `privatebin` | pdo:2.0.5 → pdo:2.0.6 | file-leg | **proven** | True → True | 15.4 | none printed | `apps/privatebin/` | | `romm` | mariadb:11.4, redis:7-alpine, romm:5.0.0 → mariadb:11.4, redis:7-alpine, romm:5.3.0 | db-mariadb | **proven** | True → True | 60.6 | yes | `apps/romm/` | | `vikunja` | vikunja:2.3.0 → vikunja:2.6.0 | other | **proven** | True → True | 24.6 | yes | `apps/vikunja/` | | `adventurelog` | adventurelog-backend:v0.12.1, adventurelog-frontend:v0.12.1, postgis:16-3.5-alpine → adventurelog-backend:v0.13.0, adventurelog-frontend:v0.13.0, postgis:16-3.5-alpine | db-postgis | **failed** | True → False | 346.6 | — | `apps/adventurelog/` | | `docmost` | postgres:16-alpine → postgres:17-alpine | engine-major-postgres | **failed** | True → False | 254.5 | — | `apps/docmost-engine-postgres/` | | `gitea` | gitea:1.27.0 → — | other | **inconclusive** | False → False | 20.1 | — | `apps/gitea/` | | `opengist` | opengist:1.13 → opengist:1.15 | other | **inconclusive** | True → False | 14.4 | — | `apps/opengist/` | | `tandoor` | postgres:16-alpine, recipes:2.6.13 → postgres:16-alpine, recipes:2.6.15 | db-postgres | **failed** | True → False | 361.9 | — | `apps/tandoor/` | | `vaultwarden` | server:1.36.0-alpine → — | other | **inconclusive** | False → False | 17.9 | — | `apps/vaultwarden/` | | `zipline` | postgres:16-alpine, zipline:4.6.1 → — | db-postgres | **inconclusive** | False → False | 73.3 | — | `apps/zipline/` | **14 proven · 3 failed · 4 inconclusive — out of 21 attempted.** ### Why each inconclusive edge could not be judged - **`gitea`** — INCONCLUSIVE: the template sets no `INSTALL_LOCK`, so a fresh Gitea starts in its web-installer state and `gitea admin user create` refuses with `MustInstalled() [F] Unable to load config file for a installed Gitea instance`. The route that would work is POSTing the installer form first; that was not written tonight and is listed as owed rather than faked. - **`opengist`** — CORRECTED from `failed` to `inconclusive` the same night, deliberately. The UPDATE itself SUCCEEDED: phase `done` in 14.4 s, and all four version observables agree on `ghcr.io/thomiceli/opengist:1.15` with the container running and zero restarts. What failed was the READBACK: it was attempted immediately after `done` and the sign-in form was not yet being served, so the fixture got no `_csrf` and returned `http=None`. Whether the seeded account survived was therefore NOT ESTABLISHED. Recording that as `failed` would have blamed the app for the harness's impatience — `inconclusive` is the honest verdict and it is never collapsed into `failed`. The fixture now waits for the LOGIN FORM rather than for the root page. - **`vaultwarden`** — INCONCLUSIVE BY DESIGN, not by a gap in the harness: the catalog CLOSES self-registration on purpose (`SIGNUPS_ALLOWED=false`, R-512 — *a stranger who guesses vault. must not be able to register*), so `/api/accounts/register` answers 404 and there is NO account-creating route without the admin secret. Vaultwarden also ships no CLI. Tried: `POST /api/accounts/register` with a valid KDF envelope. This app cannot be seeded headlessly while that setting stands, and the setting is right. - **`zipline`** — INCONCLUSIVE BY DESIGN: the deploy answers `E1037: User registration is disabled`, so no first account can be created from outside. Tried: `POST /api/auth/register` and `POST /api/auth/setup`. SEPARATELY, zipline is one of R-618's two confirmed victims — its `.felhom.yml` probe expects 200 on `/api/health`, which the app answers 404 — so even with a seed its update would have been HELD by a wrong probe rather than by anything about the edge. ### The edges that failed — the most valuable results of the night - **`adventurelog`** — final phase `failed`, hold `A(z) adventurelog frissítése 2026-09-21 20:53-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-21 20:47 — ez a másolat a beállításokat, az adatbázist és az adatköteteket tartalmazza.`, error `A(z) adventurelog frissítése 2026-09-21 20:53-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-21 20:47 — ez a másolat a beállításokat, az adatbázist és az adatköteteket tartalmazza.`. the edge ended HELD or failed — this is a RESULT, not an error of the run - **`docmost`** — final phase `failed`, hold `A(z) docmost frissítése 2026-09-21 21:44-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-21 21:44 — ez a másolat a beállításokat, az adatbázist és az adatköteteket tartalmazza.`, error `A(z) docmost frissítése 2026-09-21 21:44-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-21 21:44 — ez a másolat a beállításokat, az adatbázist és az adatköteteket tartalmazza.`. ended HELD or failed — a RESULT, not an error of the run - **`tandoor`** — final phase `failed`, hold `A(z) tandoor frissítése 2026-09-21 21:22-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-21 21:15 — ez a másolat a beállításokat, az adatbázist és az adatköteteket tartalmazza.`, error `A(z) tandoor frissítése 2026-09-21 21:22-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-21 21:15 — ez a másolat a beállításokat, az adatbázist és az adatköteteket tartalmazza.`. the edge ended HELD or failed — this is a RESULT, not an error of the run --- ## Phase 2 — the two database engines ### 2.1 MariaDB across a major, through the REAL Update button — PROVEN, and a first `nextcloud`, app image held constant, `mariadb: 11.6 → 12.3`. Seeded and read back through `occ user:add` / `occ user:info`, with the fixture's own negative control on every readback. **The four observables of `SPIKE-r459`, before → after:** | # | observable | before | after | |---|---|---|---| | 1 | `mariadb_upgrade_info` | `11.6.2-MariaDB` | **`12.3.3-MariaDB`** | | 2 | the engine's own check (R-464 — never the log line) | *not measured: the probe was unauthenticated, see below* | **„This installation of MariaDB is already upgraded to 12.3.3-MariaDB. There is no need to run mariadb-upgrade again."** | | 3 | the entrypoint | — | **„Major version upgrade detected from 11.6.2-MariaDB to 12.3.3-MariaDB. Check required!"** → „Starting mariadb-upgrade" → **„Finished mariadb-upgrade"** | | 4 | the engine's own pre-upgrade backup | absent | **`system_mysql_backup_11.6.2-MariaDB.sql.zst`, 631 905 bytes** | **Observable 3 is the one that matters**, because R-459's whole finding was that MariaDB can apply a major and *skip* the conversion quietly, printing `skipped due to $MARIADB_AUTO_UPGRADE`. **That line is absent**; the conversion was detected, started and finished. The seeded account read back and all four version observables agree. **Two honest notes on the instrument.** The BEFORE capture of observable 2 asked the engine without credentials and got `ERROR 1045 Access denied`; the probe was corrected and the AFTER capture retaken with it, so the before value is **not measured** and is stated as such rather than inferred. And the `ls` in observable 4 printed a "(no pre-upgrade backup file present)" fallback *after* listing the file, because it globs two patterns and one did not match — the file is there. Full record: `16-phase2.1-mariadb-major.md`, `apps/nextcloud-engine-mariadb/`. ### 2.2 PostgreSQL across a major — what a household would see TODAY **Exactly what R-463 predicted, and nobody had measured.** `docmost`, engine only, `16 → 17`: the update ended **`failed` in 5.1 s**, the app was stopped and held, **the pin named `postgres:17-alpine` while `installed_images` still said 16 and nothing was running**, and the data was intact. The restore the hold sentence names brought it back in **29.1 s**, hold cleared, health probe 200. **The engine's refusal line had to be REPRODUCED**, because `failAndHold` removed the container before any probe could read it (**R-621**) and the controller log does not carry it either. Done independently with a control on every step — source proven 16, copy proven 16, 49 MB: ``` FATAL: database files are incompatible with server DETAIL: The data directory was initialized by PostgreSQL version 16, which is not compatible with this version 17.11. ``` **The datadir was still `16` afterwards** — nothing migrated, nothing damaged — and the positive control (the same copy under `postgres:16-alpine`) started and held **48 tables**. **My own first reproduction was WRONG and is kept, labelled.** The volume lookup returned empty, so the copy was empty, so 17 initialised a fresh datadir and started happily — and the run reported `running=true` as though no refusal had happened. A blank `PG_VERSION` one line earlier should have stopped the step and did not. It is kept because it accidentally measured the MIRROR case (16 refusing a 17 datadir, verbatim), and relabelled so nobody reads it as the main result. `17-postgres-refusal-reproduced.txt` (wrong) and `18-postgres-refusal-reproduced.txt` (right). ### 2.3 The conversion rehearsal, COSTED — the answer Q5 was asking for Logical dump and restore, on a fresh seeded `docmost`: **49.0 MB datadir, 48 tables.** | step | time | what it produced | |---|---|---| | dump with 16 (`pg_dumpall`) | **2.6 s** | **132 201 bytes**, 48 `CREATE TABLE` statements | | fresh 17 datadir + restore | **6.5 s** | `PG_VERSION` 17, **48 tables restored**, 2 benign ERROR lines | | point the app at 17 and start it | 124.8 s | the app's own words: *„Database connection successful"* | | **the seed read back on 17** | — | **TRUE**, through the app's own login | | **total** | **155.9 s** | of which **~9 s is engine work** | Full paragraph for Q5, including what could lose data and why `pg_upgrade` was not run: `24-Q5-postgres-conversion-costed.md`. --- ## Phase 3 — the bad days Every leg records the same five things. **The event/mail column is empty on every row for the same structural reason — see the alarm truth table.** | leg | what the household saw | what the box did by itself | time to steady | the alarm | |---|---|---|---|---| | **B1 unattended HOLD** | „Frissítés elérhető" → app **Leállítva**, the hold sentence naming tier, date and what the copy holds; the banner „Telepített alkalmazás nem fut" on **every** page | pressed **once**, held after **312.9 s**, and **never pressed again** across two further passes | 312.9 s | unmeasurable (R-620) | | **B2 pull fails** | „Az új verzió letöltése nem sikerült, ezért a frissítés elmaradt. Az alkalmazás a korábbi verzióval fut tovább." | pin **and** definition put back in **1.0 s**, `hold=None`, old version still serving | 1.0 s | none, correctly — nothing is down | | **B3 busy** | „A frissítés most nem indítható: mentés/visszaállítás folyamatban." | refused `409 reason='busy'` on six consecutive presses; the transient reason a caller needs | — | n/a | | **B4 concurrency** | nothing — all succeeded | **NO single-flight.** 2 of 2, then **5 of 5**, ran at once; all ended `done`, every pin advanced | ~30 s for five | none | | **B5 cut in `backing-up`** | „A frissítés megszakadt, mert a vezérlő újraindult…" | the box said so itself at boot (three positive observables), **pin unmoved**, data intact | one boot | none | | **B5 cut in `safety-dump`** | nothing — the ordinary badge, **no interrupted sentence** | no recovery line, no journal, **pin unmoved**, data intact | one boot | none | | **B6 way out forwards** | badge still says „Frissítés elérhető" and offers the button | the button **refuses `409 reason='held'`** | — | n/a | | **B7 disk floor** | „Nincs elég szabad hely a frissítéshez: 1.4 GB szabad… legalább 2 GB szükséges." | refused before anything moved | instant | n/a | | **B8 floating pin** | „Naprakész" | and it is **TRUE on this box** — both floating digests match upstream exactly | — | n/a | | **B9 frozen app, newer `.felhom.yml`** | nothing — 10 samples, all `running`/200 | the newer `.felhom.yml` reached the frozen app; the compose stayed frozen | — | none | **The three that changed what is known:** 1. **B1 produced the unattended HOLD** this project has never had — see `19-Q4-the-unattended-hold.md`. 2. **B4 answered the single-flight question: there is none.** Five updates ran together and all ended honest. Slice 6 must decide whether that is what it wants. 3. **B6 found an inconsistency R-524 already removed for the other case:** a held app keeps inviting the household to update and the button refuses. **R-625.** Also proven for free, across a genuine power cut: the boot sweep met a held app after an unclean shutdown and deliberately left it alone — *„whatever is holding it owns its recovery"*. --- ## Phase 4 — the morning after **B1's held app, as a household would find it at breakfast.** The app page, the dashboard, the launcher and both backups pages were read in **both languages** and are saved as HTML in `bad-days/P4-morning-after/`. **Is there ONE sentence that says what happened, since when, which copy holds what, and what to press?** Scored against Q4's recommended option: | Q4 promises the household are told… | measured | |---|---| | **what happened** | ✔ „…frissítése 2026-09-21 21:53-kor nem sikerült, és az alkalmazás nem indult el az új verzióval." | | **since when** | ✔ the time is in the sentence | | **which copy holds what** | ✔ „saját meghajtó, 2026-09-21 21:47 — ez a másolat a beállításokat, az adatbázist és az adatköteteket tartalmazza." | | **what to press** | ✔ „Visszaállítható a Mentések oldalon…" | | **the same in English** | ✘ **the sentence is Hungarian on the English page** (R-606, confirmed on the hold sentence itself, with positive and negative controls) | | by mail | **unmeasurable on this venue** (R-620) | **The app is surfaced everywhere, not only on its own page** — the banner „Telepített alkalmazás nem fut: …" / „An installed app is not running: …" appeared at the top of every authenticated page, and carried both held apps at once when there were two. **Every app still on the box, and every badge, after a rescan.** Ten deployed apps: every badge is **TRUE** — references equal ⇔ „Naprakész", references differ ⇔ „Frissítés elérhető". `zipline` shows the household „Nem egészséges — URL nem elérhető" while it is in fact serving, which is R-618 in the household's own words. --- ## The alarm truth table **The event-and-mail half of this night could not be measured, and that is a property of the venue, not an omission.** Guest 9202 has `hub.enabled: false`; every notifier entry point returns before it logs anything (`notify/notifier.go:269, :359, :917, :959, :1058`), so no hub event and no customer mail can be produced or observed there. The hub was **not** enabled on 9202 to get around this: that would register an unclaimed host at the live hub and could mail a real address, and the brief fences the hub. Filed as **R-620** (a disabled notifier should at least say which event it dropped). So the table below scores the surfaces that DO exist on this box — the app page in both languages, the dashboard, and the controller's own log — against `08-alarm-ladder.md`. | # | what happened | should it alarm, per `08` | what the box did | what the household could READ | verdict | |---|---|---|---|---|---| | 1 | `adventurelog` held after a real failed edge — app **stopped** | **YES** — `stopped` is in `IsDownState` | classified `stopped`; the boot sweep refused to restart it | the hold sentence on the app page **and** a banner on every page | **correct** — but the SEND is unmeasurable here | | 2 | `tandoor` held the same way | **YES** | same | same, both apps in one banner | **correct**, same caveat | | 3 | `glance` held by the **unattended** update | **YES** | same, and honoured across a **power cut** | same | **correct**, same caveat | | 4 | `tandoor` and `zipline` reading `unhealthy` for hours while SERVING | **NO** — `08` §4 puts `unhealthy` deliberately in the not-down set | did not alarm | „Nem egészséges — URL nem elérhető" on the dashboard | **the ladder is right and the outcome is still wrong** — see below | | 5 | pull failure (B2) — app kept running the old version | **NO** — nothing is down | did not alarm | one sentence on the card | **correct** | | 6 | five updates at once (B4) | **NO** | did not alarm | nothing | **correct** | | 7 | two power cuts (B5) | `restarting` is not down until sustained | recovered; nothing alarmed | one interrupted sentence in one case, nothing in the other | **correct** | | 8 | disk under the 2 GB floor (B7) | not an app-down state | refused the update; no alarm | the refusal sentence | **correct** — though a box at 1.4 GB free is arguably worth telling someone about, and nothing does | **Which alarm fired and was it true:** none fired, and none could — see the venue limit above. Every classification the box made was correct against `08`. **Which should have fired and did not:** on this evidence, none. Row 8 is the only candidate and it is a design question rather than a defect: `08` is an *app-down* ladder and a full disk is not an app being down. **The one that matters, and it is row 4.** `08` §4 deliberately excludes `unhealthy` — *"folding it in reintroduces the flapping fix-3 was written to stop"* — and that ruling is right. **But the same probe result the alarm ladder correctly ignores is NOT ignored by the guarded update's `verifying` phase, which waits on it and then stops the app.** One probe, two consumers, opposite tolerances, and neither document says so. That asymmetry is the whole of R-618's severity. --- ## The promotion list for the operator **CC promotes nothing.** These are the real, within-a-major edges that ended `proven` on the box tonight, with the data read back through the app's own front door both before and after. Moving each of them on the LIVE catalog is the operator's call. | app | the move | what it would mean for a box in the field | |---|---|---| | `actualbudget` | actual-server:26.7.0 → actual-server:26.9.0 | the app runs its own schema migration on the way — proven here, and the update takes a backup first | | `audiobookshelf` | audiobookshelf:2.35.1 → audiobookshelf:2.36.1 | the app runs its own schema migration on the way — proven here, and the update takes a backup first | | `bookstack` | bookstack:26.05.2, mariadb:12.3 → bookstack:26.05.5, mariadb:12.3 | no migration line printed; the app came up on the new version with its data intact | | `docmost` | docmost:0.95.0, postgres:16-alpine, redis:7-alpine → docmost:0.96.0, postgres:16-alpine, redis:7-alpine | the app runs its own schema migration on the way — proven here, and the update takes a backup first | | `grafana` | grafana:13.1.0 → grafana:13.2.2 | the app runs its own schema migration on the way — proven here, and the update takes a backup first | | `home-assistant` | home-assistant:2026.7.2 → home-assistant:2026.9.3 | no migration line printed; the app came up on the new version with its data intact | | `mealie` | mealie:v3.20.1 → mealie:v3.27.0 | the app runs its own schema migration on the way — proven here, and the update takes a backup first | | `n8n` | n8n:2.31.3 → n8n:2.40.5 | the app runs its own schema migration on the way — proven here, and the update takes a backup first | | `navidrome` | navidrome:0.63.2 → navidrome:0.64.0 | the app runs its own schema migration on the way — proven here, and the update takes a backup first | | `nextcloud` | mariadb:11.6 → mariadb:12.3 | no migration line printed; the app came up on the new version with its data intact | | `papra` | papra:26.6.1-rootless → papra:26.6.2-rootless | the app runs its own schema migration on the way — proven here, and the update takes a backup first | | `privatebin` | pdo:2.0.5 → pdo:2.0.6 | no migration line printed; the app came up on the new version with its data intact | | `romm` | mariadb:11.4, redis:7-alpine, romm:5.0.0 → mariadb:11.4, redis:7-alpine, romm:5.3.0 | the app runs its own schema migration on the way — proven here, and the update takes a backup first | | `vikunja` | vikunja:2.3.0 → vikunja:2.6.0 | the app runs its own schema migration on the way — proven here, and the update takes a backup first | **And the apps that must NOT be promoted, which is the other half of the list:** | app | the move | why not | |---|---|---| | `adventurelog` | `v0.12.1 → v0.13.0` | applies **nine database migrations successfully** and then never binds its port. Held after the full health wait; restored in 75 s. **R-622** | | `tandoor` | `2.6.13 → 2.6.15` | the update SUCCEEDS — the app served HTTP 200 on the new version for five minutes — and is then stopped by the template's own wrong health port. Fix **R-618** first; the edge itself is probably fine | | `postgres:16-alpine → 17-alpine`, anywhere | the engine | refuses to start on a 16 datadir, verbatim. The engine-major gate stays until Q5's conversion exists | **`nextcloud`'s MariaDB `11.6 → 12.3` is proven and is a different kind of entry**: it is not an app version but an engine major, and §3 decision 5 plus R-469 already permit it. It is listed here because tonight is the first time it has been pressed through the button a household presses. --- ## Teardown — three layers, plus Gitea ### Machine — guest 9202 Every throwaway app removed **through the product**, with its data where the product allowed it. Three apps carrying an `HDD_PATH` were REFUSED at „remove with data" — `/api/disks` answers `agent not configured` on this guest, so the drive path cannot be resolved and R-442's fail-closed guard keeps the app rather than half-deleting it. Each was then removed with the data KEPT, which the product does accept, and the harness's own directories were removed by name afterwards. ``` containers now: felhom-controller · filebrowser · traefik <- the three protected only drill images: (none) drill volumes: (none) registry:2: removed, with its volume scratch drive: documents · downloads · media · roms <- the drive's own folders free space: 33 G on the docker root ``` **`controller.yaml` restored** from `controller.yaml.pre-update-night`, the controller restarted, and the cache's origin **read back and quoted** — which is the point of the exercise: ``` origin https://gitea.dooplex.hu/admin/app-catalog-felhom.eu.git (fetch) origin https://gitea.dooplex.hu/admin/app-catalog-felhom.eu.git (push) 4463243 Upgrade harness: four fixtures and seven real upstream edges from the update night (R-462) ``` **The teardown found the night's last defect**, which is the argument for doing it properly: a `navidrome` container that the product's own removal had left behind and Docker's restart policy had resurrected, invisible to every sweep that keys on `deployed`. Removed by name. **R-626.** ### Host — demo-hp **No harness LXC was created tonight, so none was destroyed** — the PostgreSQL rehearsal ran on 9202 itself with plain `docker` beside the product. `pct list` and `pvesm status` before and after are in `teardown/00-host-before.txt` and `04-host-after.txt`. **Guest 9201 untouched** — container count unchanged, and it was never addressed except to read its catalog cache for the negative control. ### Hub **Nothing provisioned: no customer, no config, no appliance, no binding.** The only hub act of the whole night was the floor save in Phase 0.1. Final state: floor `0.261.0`, declared MinAgent `0.131.0`, three host rows — unchanged from the start except the floor the operator asked for. ### Gitea ``` live catalog origin/main : 4463243f2e09 drill repo HEAD after reset : 4463243f2e09 diff of every `image:` line, live vs drill: IDENTICAL — every image: line matches the live catalog ``` **One later commit, after the teardown was taken:** the vikunja fixture's create verb was corrected from `POST` to `PUT` (test code only), so the live catalog's `main` now reads `d4392e2a10f4`. The check that matters is unchanged and was re-run afterwards — `git diff f5f6a152b513 origin/main -- templates/` is **empty**: **not one `image:` line moved on the live catalog at any point in the night.** CI green by id for both catalog pushes (jobs 830 and 872) and for felhom.eu (job 870). The drill repo is **KEPT**, private, and reset to the live catalog's `main`, so the next drill starts clean. **The live catalog's `main` moved once tonight** — from `f5f6a152b513` to `4463243f2e09` — and that commit changes `scripts/` only: four harness fixtures and seven edge definitions. **Zero `image:` lines moved on the live catalog at any point in the night**, which the diff above proves rather than asserts. --- ## Claims in the brief that turned out wrong The brief asked for this explicitly. Each claim, and what was measured. | the brief said | measured | |---|---| | **a box will follow a second catalog by `git.repo_url` alone** (read from config source, never run) | **WRONG.** `Syncer.gitCloneOrPull` clones only when `.git` is absent; otherwise it fetches from the remote the clone already stores. The cache directory had to be removed too. **R-615** | | **`CompareImageRefs` may not order references carrying a `host:port/` prefix** (read, not run) | **The worry was unfounded.** `splitImageRef` takes the LAST colon and rejects it only when a `/` follows, so a registry port is never read as a tag. Proven by RUNNING it: four positive cases and a negative control | | **PostgreSQL 17 refuses a 16 datadir and the update ends HELD with data intact** (R-463 and source, not measured) | **RIGHT, and now measured** — 5.1 s to held, pin on 17 with nothing running, data intact, restore back in 29.1 s. The refusal line itself had to be reproduced because the product destroyed it (**R-621**) | | **a MariaDB sidecar major through the BUTTON behaves as it did on the harness** | **RIGHT.** All four observables, including the conversion actually running rather than being skipped, and the engine's own pre-upgrade backup | | **9202 has the capacity for this** | **RIGHT.** 26 GB RAM, 56 GB free on the docker root at the start, a 938 GB scratch drive. Peak usage never threatened it; images were reclaimed BY NAME twice, never pruned | | **39 within-a-major edges still exist upstream tonight** | **RIGHT, exactly.** The drift script re-run at 20:07 returned 66 pins, 46 behind, **39 within a major and 7 across** — the same numbers | **And two more the brief did not name, found the same way:** - `update-arc-gaps-2026-09-21/00-api-recipe.md` said the app page is `/app/`. **It is `/apps/`**, and every call that recipe described 404s. Corrected in that file. - `unattended-caller.py`'s `follow()` read the API **envelope**, so every update it followed would have run to a 900 s timeout and been recorded `timeout` rather than `held`. **R-623**, fixed before B1 relied on it — and B1's log is what the fixed version produces. **One correction to a register row, which is the same class of error one layer up.** R-606 records controller v0.260.0 as having made the pre-flight refusals reach an English household in English. **Measured: it did not.** `held`, `not_deployed` and `disk` all come back identical Hungarian with `?lang=en`, because the routing exists and the sentences are frozen string constants that never entered it. A row that records something as fixed when it is not is worse than an open row.