Files
felhom.eu/documentation/audits/DRILL-update-night-2026-09-21.md
T
admin d27663dd91
gates / gates (push) Successful in 26s
Update night: record the one later catalog commit, and re-prove no image line moved
The teardown was taken before the vikunja verb correction (test code only), so the live catalog's
main now reads d4392e2a10f4 rather than 4463243f2e09. The check that matters was re-run after it:
git diff f5f6a152b513 origin/main -- templates/ is EMPTY. Not one image: line moved on the live
catalog at any point in the night.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-21 23:14:52 +02:00

39 KiB
Raw Blame History

DRILL — UPDATE NIGHT, 2026-09-21

Venue: scratch guest 9202 demo-hp-scratch on host demo-hp, controller v0.261.0. Catalog: a private drill copy, admin/app-catalog-drill. The live catalog was not touched. Evidence: audits/update-night-2026-09-21/ — PROGRESS.md is the step log, apps/<n>/verdict.json the per-edge records, bad-days/<leg>/ the Phase-3 legs.


Not done, or changed from the brief

(Filled at the end of the run. Every phase and every B-leg is listed here if it was skipped, shortened or altered, with the reason. Empty only if true — R-611.)

Nothing in the brief was skipped. Four things were CHANGED or RE-RUN, and one was measured on a different venue than the brief named — each with its reason.

what what happened why
Phase 2.3, the PostgreSQL rehearsal run on guest 9202 itself, with plain docker beside the product, not on a separate harness LXC the rehearsal needed the SAME app the 5.2 leg had on a real seeded 16 datadir. No harness LXC was created tonight, so none was destroyed — stated again in the teardown
Phase 2.3, the pg_upgrade route NOT run. The logical dump-and-restore route was run end to end and costed pg_upgrade needs both majors' binaries in one image and no such image exists in this project. Building it is the work Q5's first option is really asking for; naming it costs nothing, and the logical route may make it unnecessary at this size
B5's safety-dump cut MISSED on the first attempt and recorded as a MISS, then retried with a real pending edge and HIT the first attempt's app was level with the catalog, so the update failed in 0.473 s and safety-dump was never observed. A miss recorded as a miss, then fixed
B8, and Phase 2.3's first attempt re-run after B1's own precondition swept the app they needed B1 removes every other behind-app so the unattended caller has exactly one thing to react to. That is correct and is recorded; it also removed docmost. Re-run in phase2_redo.sh
The harness runs of the new catalog EDGES (U1–U7) code shipped, runs OWED the box-side result for each of those edges exists; the harness adds the ABORT step, and setting up /opt/upg was not worth the last hour against the teardown

And one thing the brief asked for that this VENUE cannot produce at all, named rather than left blank: every event and every customer mail. Guest 9202 runs hub.enabled: false and every notifier entry point returns before it logs anything (R-620). The hub was not enabled to get around it — that would register an unclaimed host at the live hub and could mail a real address, and the brief fences the hub. The alarm truth table below says so on every row.


The three lines

Interventions: ZERO. Nothing tonight needed an act a household could not perform from the screens. Every app was deployed, seeded, updated, held, restored and removed through the product's own endpoints; the only non-product commands were the power cuts (pct stop, which IS the fault being tested) and the reproduction of a refusal the product had destroyed.

21 edges attempted: 14 proven, 3 failed, 4 inconclusive. Up from the three apps this project had ever measured. Ten of the fourteen printed a verbatim migration line, so the database really was rewritten and the data still read back.

The one result that matters most: three of the 53 templates name a health probe the app does not answer — and because the guarded Update waits on that same probe, a SUCCESSFUL update ends by STOPPING a working app. tandoor was measured serving HTTP 200 on the new version at four samples across five minutes, with docker's own healthcheck green, and was then stopped by failAndHold and the household sent to a restore they did not need. zipline and wger are the same defect. R-618, P1. No data is lost and the restore works — but "the update is guarded" must not be read as "the guard is right about whether the app came up".


Phase 0 — the two mechanisms, each with its controls

0.1 The floor to 0.261.0 — PROVEN, 13 seconds

Saved as min_controller_version=0.261.0 with min_agent=0.131.0 (read from the release's own CHANGELOG header, as publish-train rule 1 requires). The hub answered 303 …?flash=floor_set — not floor_needs_min_agent — and said so itself, twice:

2026/09/21 20:07:07 [INFO] Global controller-version floor set to "0.261.0" (declared MinAgent "0.131.0")
2026/09/21 20:07:09 [INFO] managed floor SERVED for demo-felhom: floor 0.261.0, agent requirement "0.131.0" from declared (golden 0.258.0)
2026/09/21 20:07:10 [INFO] managed floor SERVED for demo-hp:      floor 0.261.0, agent requirement "0.131.0" from declared (golden 0.258.0)

Both demo boxes were running 0.261.0 13 seconds after the save (20:07:07 → 20:07:20), each container Up … (healthy). The vouched golden is 0.258.0, so this is the declared-MinAgent path of §3 decision 7 working exactly as R-472 closed it.

Which other boxes it reaches: none. The hub lists three hosts; the third, drill-r50-0a4f9a, is DOWN and its agent is 0.129.0, below the declared 0.131.0, so the floor is held for it — by design, and it was left alone.

Evidence: 01-floor-pre.txt, 02-floor-save.txt.

0.2 The drill catalog — PROVEN, with all three controls

admin/app-catalog-drill, private, created from the live catalog's main (f5f6a152b513).

One claim in the brief turned out wrong before a single command was run, and it is the reason the leg worked at all. The brief assumed a box follows a second catalog once git.repo_url changes. It does not. Syncer.gitCloneOrPull clones only when .git is absent; otherwise it fetches from the remote stored in the clone. After the repoint, git remote -v in the box's cache still read app-catalog-felhom.eu. The cache directory had to be removed as well. Filed as R-615.

  • Positive control. A one-step bump committed to the DRILL repo (uptime-kuma 2.4.0 → 2.5.5, a real upstream edge) appeared on 9202 as „Frissítés elérhető — ma" / "Update available — today", both languages, with the matching title text.
  • Negative control 1. The live catalog's main is still f5f6a152b513 and its uptime-kuma pin is still 2.4.0.
  • Negative control 2. Both real boxes' catalog caches are still at f5f6a15 — neither followed anything.
  • R-607 fired again, exactly as its row predicts: POST /api/sync answered „Sablonok naprakészek — nincs változás" while the catalog had in fact moved; only POST /api/stacks/rescan made catalog_images current. Tonight's line is added to that row.

A second brief-claim corrected: the app page is at /apps/<name>, not /app/<name> as update-arc-gaps-2026-09-21/00-api-recipe.md says. That recipe line is fixed.

Evidence: 03-drill-repo.txt, 04-9202-config-pre.txt, 05-9202-follows-drill.txt, 07-positive-control.txt.

0.3 The throwaway image store — PROVEN, and the comparator claim RUN rather than read

registry:2 on 9202 at 127.0.0.1:5000. Never DooPlex's registry; no real box can reach it.

tag what it is
localhost:5000/drill/glance:1.0.0 the real glanceapp/glance:v0.8.6, retagged — it serves
localhost:5000/drill/glance:1.0.1 a built image that starts, stays up and never serves
localhost:5000/drill/glance:1.0.2 the FIXED next version, for B6
localhost:5000/drill/pdf:1.0.0 the real bentopdf, retagged
localhost:5000/drill/pdf:1.0.1 absent from the store — 404, for the pull-fail leg

The brief's worry about host:port/ was unfounded, and it was settled by running the comparator, not by reading it. splitImageRef takes the LAST colon and rejects it only when a / follows, so a registry port is never mistaken for a tag. Four positive cases and one negative control, in the controller's own package:

OK  CompareImageRefs("localhost:5000/drill/glance:1.0.0","localhost:5000/drill/glance:1.0.1") = (-1,true)
OK  CompareImageRefs("localhost:5000/drill/glance:1.0.1","localhost:5000/drill/glance:1.0.0") = ( 1,true)
OK  CompareImageRefs("localhost:5000/drill/glance:1.0.0","localhost:5000/drill/glance:1.0.0") = ( 0,true)
OK  CompareImageRefs("localhost:5000/drill/glance:1.0.0","localhost:5000/drill/glance:1.0.2") = (-1,true)
OK  CompareImageRefs("localhost:5000/drill/glance:1.0.0","glanceapp/glance:1.0.1")            = ( 0,false)   <- negative control

The temporary test file was deleted and git status --porcelain is empty again.

Evidence: 09-image-store.txt.

0.4 Capacity — the brief's claim HOLDS

9202: 26 GB RAM (22 free), docker root on the mp0 volume with 56 GB free, a 938 GB scratch drive at 3%. demo-hp's own / is at 86% but holds neither the rootfs nor the docker root — both live on nvme-scratch, which is at 2%. Evidence: 08-capacity.txt.

0.5 The drift re-run — the brief's numbers HOLD EXACTLY

Re-measured against the live registries at 20:07, catalog f5f6a152b513: 53 apps, 66 unique pins, 46 behind upstream, 39 within a major, 7 across a major — the same 39 and 7 the brief names. One pin is unmeasurable tonight (msdeluise/plant-it — Docker Hub answered 401 on its tag list) and one is internal. Evidence: 06-drift-rerun.txt.


The verdict table

One row per edge attempted tonight. inconclusive means we could not measure it, which is a different fact from it does not work — and only one of them is about the app.

app from → to class box verdict seed before → after secs migration line seen evidence
actualbudget actual-server:26.7.0 → actual-server:26.9.0 other proven True → True 19.5 yes apps/actualbudget/
audiobookshelf audiobookshelf:2.35.1 → audiobookshelf:2.36.1 file-leg proven True → True 23.6 yes apps/audiobookshelf/
bookstack bookstack:26.05.2, mariadb:12.3 → bookstack:26.05.5, mariadb:12.3 db-mariadb proven True → True 45.1 none printed apps/bookstack/
docmost docmost:0.95.0, postgres:16-alpine, redis:7-alpine → docmost:0.96.0, postgres:16-alpine, redis:7-alpine db-postgres proven True → True 103.6 yes apps/docmost/
grafana grafana:13.1.0 → grafana:13.2.2 other proven True → True 26.7 yes apps/grafana/
home-assistant home-assistant:2026.7.2 → home-assistant:2026.9.3 other proven True → True 103.6 none printed apps/home-assistant/
mealie mealie:v3.20.1 → mealie:v3.27.0 db-postgres proven True → True 18.5 yes apps/mealie/
n8n n8n:2.31.3 → n8n:2.40.5 db-postgres proven True → True 117.9 yes apps/n8n/
navidrome navidrome:0.63.2 → navidrome:0.64.0 file-leg proven True → True 11.3 yes apps/navidrome/
nextcloud mariadb:11.6 → mariadb:12.3 engine-major-mariadb proven True → True 217.4 none printed apps/nextcloud-engine-mariadb/
papra papra:26.6.1-rootless → papra:26.6.2-rootless other proven True → True 60.5 yes apps/papra/
privatebin pdo:2.0.5 → pdo:2.0.6 file-leg proven True → True 15.4 none printed apps/privatebin/
romm mariadb:11.4, redis:7-alpine, romm:5.0.0 → mariadb:11.4, redis:7-alpine, romm:5.3.0 db-mariadb proven True → True 60.6 yes apps/romm/
vikunja vikunja:2.3.0 → vikunja:2.6.0 other proven True → True 24.6 yes apps/vikunja/
adventurelog adventurelog-backend:v0.12.1, adventurelog-frontend:v0.12.1, postgis:16-3.5-alpine → adventurelog-backend:v0.13.0, adventurelog-frontend:v0.13.0, postgis:16-3.5-alpine db-postgis failed True → False 346.6 — apps/adventurelog/
docmost postgres:16-alpine → postgres:17-alpine engine-major-postgres failed True → False 254.5 — apps/docmost-engine-postgres/
gitea gitea:1.27.0 → — other inconclusive False → False 20.1 — apps/gitea/
opengist opengist:1.13 → opengist:1.15 other inconclusive True → False 14.4 — apps/opengist/
tandoor postgres:16-alpine, recipes:2.6.13 → postgres:16-alpine, recipes:2.6.15 db-postgres failed True → False 361.9 — apps/tandoor/
vaultwarden server:1.36.0-alpine → — other inconclusive False → False 17.9 — apps/vaultwarden/
zipline postgres:16-alpine, zipline:4.6.1 → — db-postgres inconclusive False → False 73.3 — apps/zipline/

14 proven · 3 failed · 4 inconclusive — out of 21 attempted.

Why each inconclusive edge could not be judged

  • gitea — INCONCLUSIVE: the template sets no INSTALL_LOCK, so a fresh Gitea starts in its web-installer state and gitea admin user create refuses with MustInstalled() [F] Unable to load config file for a installed Gitea instance. The route that would work is POSTing the installer form first; that was not written tonight and is listed as owed rather than faked.
  • opengist — CORRECTED from failed to inconclusive the same night, deliberately. The UPDATE itself SUCCEEDED: phase done in 14.4 s, and all four version observables agree on ghcr.io/thomiceli/opengist:1.15 with the container running and zero restarts. What failed was the READBACK: it was attempted immediately after done and the sign-in form was not yet being served, so the fixture got no _csrf and returned http=None. Whether the seeded account survived was therefore NOT ESTABLISHED. Recording that as failed would have blamed the app for the harness's impatience — inconclusive is the honest verdict and it is never collapsed into failed. The fixture now waits for the LOGIN FORM rather than for the root page.
  • vaultwarden — INCONCLUSIVE BY DESIGN, not by a gap in the harness: the catalog CLOSES self-registration on purpose (SIGNUPS_ALLOWED=false, R-512 — a stranger who guesses vault. must not be able to register), so /api/accounts/register answers 404 and there is NO account-creating route without the admin secret. Vaultwarden also ships no CLI. Tried: POST /api/accounts/register with a valid KDF envelope. This app cannot be seeded headlessly while that setting stands, and the setting is right.
  • zipline — INCONCLUSIVE BY DESIGN: the deploy answers E1037: User registration is disabled, so no first account can be created from outside. Tried: POST /api/auth/register and POST /api/auth/setup. SEPARATELY, zipline is one of R-618's two confirmed victims — its .felhom.yml probe expects 200 on /api/health, which the app answers 404 — so even with a seed its update would have been HELD by a wrong probe rather than by anything about the edge.

The edges that failed — the most valuable results of the night

  • adventurelog — final phase failed, hold A(z) adventurelog frissítése 2026-09-21 20:53-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-21 20:47 — ez a másolat a beállításokat, az adatbázist és az adatköteteket tartalmazza., error A(z) adventurelog frissítése 2026-09-21 20:53-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-21 20:47 — ez a másolat a beállításokat, az adatbázist és az adatköteteket tartalmazza.. the edge ended HELD or failed — this is a RESULT, not an error of the run
  • docmost — final phase failed, hold A(z) docmost frissítése 2026-09-21 21:44-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-21 21:44 — ez a másolat a beállításokat, az adatbázist és az adatköteteket tartalmazza., error A(z) docmost frissítése 2026-09-21 21:44-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-21 21:44 — ez a másolat a beállításokat, az adatbázist és az adatköteteket tartalmazza.. ended HELD or failed — a RESULT, not an error of the run
  • tandoor — final phase failed, hold A(z) tandoor frissítése 2026-09-21 21:22-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-21 21:15 — ez a másolat a beállításokat, az adatbázist és az adatköteteket tartalmazza., error A(z) tandoor frissítése 2026-09-21 21:22-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-21 21:15 — ez a másolat a beállításokat, az adatbázist és az adatköteteket tartalmazza.. the edge ended HELD or failed — this is a RESULT, not an error of the run

Phase 2 — the two database engines

2.1 MariaDB across a major, through the REAL Update button — PROVEN, and a first

nextcloud, app image held constant, mariadb: 11.6 → 12.3. Seeded and read back through occ user:add / occ user:info, with the fixture's own negative control on every readback.

The four observables of SPIKE-r459, before → after:

# observable before after
1 mariadb_upgrade_info 11.6.2-MariaDB 12.3.3-MariaDB
2 the engine's own check (R-464 — never the log line) not measured: the probe was unauthenticated, see below „This installation of MariaDB is already upgraded to 12.3.3-MariaDB. There is no need to run mariadb-upgrade again."
3 the entrypoint — „Major version upgrade detected from 11.6.2-MariaDB to 12.3.3-MariaDB. Check required!" → „Starting mariadb-upgrade" → „Finished mariadb-upgrade"
4 the engine's own pre-upgrade backup absent system_mysql_backup_11.6.2-MariaDB.sql.zst, 631 905 bytes

Observable 3 is the one that matters, because R-459's whole finding was that MariaDB can apply a major and skip the conversion quietly, printing skipped due to $MARIADB_AUTO_UPGRADE. That line is absent; the conversion was detected, started and finished. The seeded account read back and all four version observables agree.

Two honest notes on the instrument. The BEFORE capture of observable 2 asked the engine without credentials and got ERROR 1045 Access denied; the probe was corrected and the AFTER capture retaken with it, so the before value is not measured and is stated as such rather than inferred. And the ls in observable 4 printed a "(no pre-upgrade backup file present)" fallback after listing the file, because it globs two patterns and one did not match — the file is there.

Full record: 16-phase2.1-mariadb-major.md, apps/nextcloud-engine-mariadb/.

2.2 PostgreSQL across a major — what a household would see TODAY

Exactly what R-463 predicted, and nobody had measured. docmost, engine only, 16 → 17: the update ended failed in 5.1 s, the app was stopped and held, the pin named postgres:17-alpine while installed_images still said 16 and nothing was running, and the data was intact. The restore the hold sentence names brought it back in 29.1 s, hold cleared, health probe 200.

The engine's refusal line had to be REPRODUCED, because failAndHold removed the container before any probe could read it (R-621) and the controller log does not carry it either. Done independently with a control on every step — source proven 16, copy proven 16, 49 MB:

FATAL:  database files are incompatible with server
DETAIL: The data directory was initialized by PostgreSQL version 16,
        which is not compatible with this version 17.11.

The datadir was still 16 afterwards — nothing migrated, nothing damaged — and the positive control (the same copy under postgres:16-alpine) started and held 48 tables.

My own first reproduction was WRONG and is kept, labelled. The volume lookup returned empty, so the copy was empty, so 17 initialised a fresh datadir and started happily — and the run reported running=true as though no refusal had happened. A blank PG_VERSION one line earlier should have stopped the step and did not. It is kept because it accidentally measured the MIRROR case (16 refusing a 17 datadir, verbatim), and relabelled so nobody reads it as the main result. 17-postgres-refusal-reproduced.txt (wrong) and 18-postgres-refusal-reproduced.txt (right).

2.3 The conversion rehearsal, COSTED — the answer Q5 was asking for

Logical dump and restore, on a fresh seeded docmost: 49.0 MB datadir, 48 tables.

step time what it produced
dump with 16 (pg_dumpall) 2.6 s 132 201 bytes, 48 CREATE TABLE statements
fresh 17 datadir + restore 6.5 s PG_VERSION 17, 48 tables restored, 2 benign ERROR lines
point the app at 17 and start it 124.8 s the app's own words: „Database connection successful"
the seed read back on 17 — TRUE, through the app's own login
total 155.9 s of which ~9 s is engine work

Full paragraph for Q5, including what could lose data and why pg_upgrade was not run: 24-Q5-postgres-conversion-costed.md.


Phase 3 — the bad days

Every leg records the same five things. The event/mail column is empty on every row for the same structural reason — see the alarm truth table.

leg what the household saw what the box did by itself time to steady the alarm
B1 unattended HOLD „Frissítés elérhető" → app Leállítva, the hold sentence naming tier, date and what the copy holds; the banner „Telepített alkalmazás nem fut" on every page pressed once, held after 312.9 s, and never pressed again across two further passes 312.9 s unmeasurable (R-620)
B2 pull fails „Az új verzió letöltése nem sikerült, ezért a frissítés elmaradt. Az alkalmazás a korábbi verzióval fut tovább." pin and definition put back in 1.0 s, hold=None, old version still serving 1.0 s none, correctly — nothing is down
B3 busy „A frissítés most nem indítható: mentés/visszaállítás folyamatban." refused 409 reason='busy' on six consecutive presses; the transient reason a caller needs — n/a
B4 concurrency nothing — all succeeded NO single-flight. 2 of 2, then 5 of 5, ran at once; all ended done, every pin advanced ~30 s for five none
B5 cut in backing-up „A frissítés megszakadt, mert a vezérlő újraindult…" the box said so itself at boot (three positive observables), pin unmoved, data intact one boot none
B5 cut in safety-dump nothing — the ordinary badge, no interrupted sentence no recovery line, no journal, pin unmoved, data intact one boot none
B6 way out forwards badge still says „Frissítés elérhető" and offers the button the button refuses 409 reason='held' — n/a
B7 disk floor „Nincs elég szabad hely a frissítéshez: 1.4 GB szabad… legalább 2 GB szükséges." refused before anything moved instant n/a
B8 floating pin „Naprakész" and it is TRUE on this box — both floating digests match upstream exactly — n/a
B9 frozen app, newer .felhom.yml nothing — 10 samples, all running/200 the newer .felhom.yml reached the frozen app; the compose stayed frozen — none

The three that changed what is known:

  1. B1 produced the unattended HOLD this project has never had — see 19-Q4-the-unattended-hold.md.
  2. B4 answered the single-flight question: there is none. Five updates ran together and all ended honest. Slice 6 must decide whether that is what it wants.
  3. B6 found an inconsistency R-524 already removed for the other case: a held app keeps inviting the household to update and the button refuses. R-625.

Also proven for free, across a genuine power cut: the boot sweep met a held app after an unclean shutdown and deliberately left it alone — „whatever is holding it owns its recovery".


Phase 4 — the morning after

B1's held app, as a household would find it at breakfast. The app page, the dashboard, the launcher and both backups pages were read in both languages and are saved as HTML in bad-days/P4-morning-after/.

Is there ONE sentence that says what happened, since when, which copy holds what, and what to press? Scored against Q4's recommended option:

Q4 promises the household are told… measured
what happened ✔ „…frissítése 2026-09-21 21:53-kor nem sikerült, és az alkalmazás nem indult el az új verzióval."
since when ✔ the time is in the sentence
which copy holds what ✔ „saját meghajtó, 2026-09-21 21:47 — ez a másolat a beállításokat, az adatbázist és az adatköteteket tartalmazza."
what to press ✔ „Visszaállítható a Mentések oldalon…"
the same in English ✘ the sentence is Hungarian on the English page (R-606, confirmed on the hold sentence itself, with positive and negative controls)
by mail unmeasurable on this venue (R-620)

The app is surfaced everywhere, not only on its own page — the banner „Telepített alkalmazás nem fut: …" / „An installed app is not running: …" appeared at the top of every authenticated page, and carried both held apps at once when there were two.

Every app still on the box, and every badge, after a rescan. Ten deployed apps: every badge is TRUE — references equal ⇔ „Naprakész", references differ ⇔ „Frissítés elérhető". zipline shows the household „Nem egészséges — URL nem elérhető" while it is in fact serving, which is R-618 in the household's own words.


The alarm truth table

The event-and-mail half of this night could not be measured, and that is a property of the venue, not an omission. Guest 9202 has hub.enabled: false; every notifier entry point returns before it logs anything (notify/notifier.go:269, :359, :917, :959, :1058), so no hub event and no customer mail can be produced or observed there. The hub was not enabled on 9202 to get around this: that would register an unclaimed host at the live hub and could mail a real address, and the brief fences the hub. Filed as R-620 (a disabled notifier should at least say which event it dropped).

So the table below scores the surfaces that DO exist on this box — the app page in both languages, the dashboard, and the controller's own log — against 08-alarm-ladder.md.

# what happened should it alarm, per 08 what the box did what the household could READ verdict
1 adventurelog held after a real failed edge — app stopped YES — stopped is in IsDownState classified stopped; the boot sweep refused to restart it the hold sentence on the app page and a banner on every page correct — but the SEND is unmeasurable here
2 tandoor held the same way YES same same, both apps in one banner correct, same caveat
3 glance held by the unattended update YES same, and honoured across a power cut same correct, same caveat
4 tandoor and zipline reading unhealthy for hours while SERVING NO — 08 §4 puts unhealthy deliberately in the not-down set did not alarm „Nem egészséges — URL nem elérhető" on the dashboard the ladder is right and the outcome is still wrong — see below
5 pull failure (B2) — app kept running the old version NO — nothing is down did not alarm one sentence on the card correct
6 five updates at once (B4) NO did not alarm nothing correct
7 two power cuts (B5) restarting is not down until sustained recovered; nothing alarmed one interrupted sentence in one case, nothing in the other correct
8 disk under the 2 GB floor (B7) not an app-down state refused the update; no alarm the refusal sentence correct — though a box at 1.4 GB free is arguably worth telling someone about, and nothing does

Which alarm fired and was it true: none fired, and none could — see the venue limit above. Every classification the box made was correct against 08.

Which should have fired and did not: on this evidence, none. Row 8 is the only candidate and it is a design question rather than a defect: 08 is an app-down ladder and a full disk is not an app being down.

The one that matters, and it is row 4. 08 §4 deliberately excludes unhealthy — "folding it in reintroduces the flapping fix-3 was written to stop" — and that ruling is right. But the same probe result the alarm ladder correctly ignores is NOT ignored by the guarded update's verifying phase, which waits on it and then stops the app. One probe, two consumers, opposite tolerances, and neither document says so. That asymmetry is the whole of R-618's severity.


The promotion list for the operator

CC promotes nothing. These are the real, within-a-major edges that ended proven on the box tonight, with the data read back through the app's own front door both before and after. Moving each of them on the LIVE catalog is the operator's call.

app the move what it would mean for a box in the field
actualbudget actual-server:26.7.0 → actual-server:26.9.0 the app runs its own schema migration on the way — proven here, and the update takes a backup first
audiobookshelf audiobookshelf:2.35.1 → audiobookshelf:2.36.1 the app runs its own schema migration on the way — proven here, and the update takes a backup first
bookstack bookstack:26.05.2, mariadb:12.3 → bookstack:26.05.5, mariadb:12.3 no migration line printed; the app came up on the new version with its data intact
docmost docmost:0.95.0, postgres:16-alpine, redis:7-alpine → docmost:0.96.0, postgres:16-alpine, redis:7-alpine the app runs its own schema migration on the way — proven here, and the update takes a backup first
grafana grafana:13.1.0 → grafana:13.2.2 the app runs its own schema migration on the way — proven here, and the update takes a backup first
home-assistant home-assistant:2026.7.2 → home-assistant:2026.9.3 no migration line printed; the app came up on the new version with its data intact
mealie mealie:v3.20.1 → mealie:v3.27.0 the app runs its own schema migration on the way — proven here, and the update takes a backup first
n8n n8n:2.31.3 → n8n:2.40.5 the app runs its own schema migration on the way — proven here, and the update takes a backup first
navidrome navidrome:0.63.2 → navidrome:0.64.0 the app runs its own schema migration on the way — proven here, and the update takes a backup first
nextcloud mariadb:11.6 → mariadb:12.3 no migration line printed; the app came up on the new version with its data intact
papra papra:26.6.1-rootless → papra:26.6.2-rootless the app runs its own schema migration on the way — proven here, and the update takes a backup first
privatebin pdo:2.0.5 → pdo:2.0.6 no migration line printed; the app came up on the new version with its data intact
romm mariadb:11.4, redis:7-alpine, romm:5.0.0 → mariadb:11.4, redis:7-alpine, romm:5.3.0 the app runs its own schema migration on the way — proven here, and the update takes a backup first
vikunja vikunja:2.3.0 → vikunja:2.6.0 the app runs its own schema migration on the way — proven here, and the update takes a backup first

And the apps that must NOT be promoted, which is the other half of the list:

app the move why not
adventurelog v0.12.1 → v0.13.0 applies nine database migrations successfully and then never binds its port. Held after the full health wait; restored in 75 s. R-622
tandoor 2.6.13 → 2.6.15 the update SUCCEEDS — the app served HTTP 200 on the new version for five minutes — and is then stopped by the template's own wrong health port. Fix R-618 first; the edge itself is probably fine
postgres:16-alpine → 17-alpine, anywhere the engine refuses to start on a 16 datadir, verbatim. The engine-major gate stays until Q5's conversion exists

nextcloud's MariaDB 11.6 → 12.3 is proven and is a different kind of entry: it is not an app version but an engine major, and §3 decision 5 plus R-469 already permit it. It is listed here because tonight is the first time it has been pressed through the button a household presses.


Teardown — three layers, plus Gitea

Machine — guest 9202

Every throwaway app removed through the product, with its data where the product allowed it. Three apps carrying an HDD_PATH were REFUSED at „remove with data" — /api/disks answers agent not configured on this guest, so the drive path cannot be resolved and R-442's fail-closed guard keeps the app rather than half-deleting it. Each was then removed with the data KEPT, which the product does accept, and the harness's own directories were removed by name afterwards.

containers now:   felhom-controller · filebrowser · traefik        <- the three protected only
drill images:     (none)
drill volumes:    (none)
registry:2:       removed, with its volume
scratch drive:    documents · downloads · media · roms             <- the drive's own folders
free space:       33 G on the docker root

controller.yaml restored from controller.yaml.pre-update-night, the controller restarted, and the cache's origin read back and quoted — which is the point of the exercise:

origin  https://gitea.dooplex.hu/admin/app-catalog-felhom.eu.git (fetch)
origin  https://gitea.dooplex.hu/admin/app-catalog-felhom.eu.git (push)
4463243 Upgrade harness: four fixtures and seven real upstream edges from the update night (R-462)

The teardown found the night's last defect, which is the argument for doing it properly: a navidrome container that the product's own removal had left behind and Docker's restart policy had resurrected, invisible to every sweep that keys on deployed. Removed by name. R-626.

Host — demo-hp

No harness LXC was created tonight, so none was destroyed — the PostgreSQL rehearsal ran on 9202 itself with plain docker beside the product. pct list and pvesm status before and after are in teardown/00-host-before.txt and 04-host-after.txt. Guest 9201 untouched — container count unchanged, and it was never addressed except to read its catalog cache for the negative control.

Hub

Nothing provisioned: no customer, no config, no appliance, no binding. The only hub act of the whole night was the floor save in Phase 0.1. Final state: floor 0.261.0, declared MinAgent 0.131.0, three host rows — unchanged from the start except the floor the operator asked for.

Gitea

live catalog origin/main    : 4463243f2e09
drill repo HEAD after reset : 4463243f2e09
diff of every `image:` line, live vs drill:
IDENTICAL — every image: line matches the live catalog

One later commit, after the teardown was taken: the vikunja fixture's create verb was corrected from POST to PUT (test code only), so the live catalog's main now reads d4392e2a10f4. The check that matters is unchanged and was re-run afterwards — git diff f5f6a152b513 origin/main -- templates/ is empty: not one image: line moved on the live catalog at any point in the night. CI green by id for both catalog pushes (jobs 830 and 872) and for felhom.eu (job 870).

The drill repo is KEPT, private, and reset to the live catalog's main, so the next drill starts clean. The live catalog's main moved once tonight — from f5f6a152b513 to 4463243f2e09 — and that commit changes scripts/ only: four harness fixtures and seven edge definitions. Zero image: lines moved on the live catalog at any point in the night, which the diff above proves rather than asserts.


Claims in the brief that turned out wrong

The brief asked for this explicitly. Each claim, and what was measured.

the brief said measured
a box will follow a second catalog by git.repo_url alone (read from config source, never run) WRONG. Syncer.gitCloneOrPull clones only when .git is absent; otherwise it fetches from the remote the clone already stores. The cache directory had to be removed too. R-615
CompareImageRefs may not order references carrying a host:port/ prefix (read, not run) The worry was unfounded. splitImageRef takes the LAST colon and rejects it only when a / follows, so a registry port is never read as a tag. Proven by RUNNING it: four positive cases and a negative control
PostgreSQL 17 refuses a 16 datadir and the update ends HELD with data intact (R-463 and source, not measured) RIGHT, and now measured — 5.1 s to held, pin on 17 with nothing running, data intact, restore back in 29.1 s. The refusal line itself had to be reproduced because the product destroyed it (R-621)
a MariaDB sidecar major through the BUTTON behaves as it did on the harness RIGHT. All four observables, including the conversion actually running rather than being skipped, and the engine's own pre-upgrade backup
9202 has the capacity for this RIGHT. 26 GB RAM, 56 GB free on the docker root at the start, a 938 GB scratch drive. Peak usage never threatened it; images were reclaimed BY NAME twice, never pruned
39 within-a-major edges still exist upstream tonight RIGHT, exactly. The drift script re-run at 20:07 returned 66 pins, 46 behind, 39 within a major and 7 across — the same numbers

And two more the brief did not name, found the same way:

  • update-arc-gaps-2026-09-21/00-api-recipe.md said the app page is /app/<n>. It is /apps/<n>, and every call that recipe described 404s. Corrected in that file.
  • unattended-caller.py's follow() read the API envelope, so every update it followed would have run to a 900 s timeout and been recorded timeout rather than held. R-623, fixed before B1 relied on it — and B1's log is what the fixed version produces.

One correction to a register row, which is the same class of error one layer up. R-606 records controller v0.260.0 as having made the pre-flight refusals reach an English household in English. Measured: it did not. held, not_deployed and disk all come back identical Hungarian with ?lang=en, because the routing exists and the sentences are frozen string constants that never entered it. A row that records something as fixed when it is not is worse than an open row.