Found by the live run, not by reading: POST /api/sync answered 'nincs valtozas' while the box's catalog cache HAD moved, and catalog_images stayed stale until a separate rescan. Since CatalogImages is the one input CatalogOrder compares against, the badge answers from a stale catalog for that window — and the session nearly recorded a stale tag-ok badge as proof of the R-524 ahead arm. Neither half is isolated, so the row records the observation, not a diagnosis. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
20 KiB
UPDATE ARC — the state, measured (2026-09-21)
Phase 0 of the resumed update arc. Nothing here is an estimate. Every number was read from a
live box, a live registry or live source on 2026-09-21. Raw captures:
audits/update-arc-2026-09-21/.
Baselines at the start: controller 19ef0329ab66 v0.259.0 · agent d9864a94bf62 v0.132.0 ·
felhom.eu bcdd5b205875 hub v0.119.0 · catalog 18a6d2d8243e. All four trees clean and level with
origin/main.
1. Three claims in the brief that turned out wrong — named first
1.1 R-589 is NOT open. It shipped in v0.258.0. The brief said, reviewer-verified, that
updatebadge.go L82–L99 builds the badge from four raw Hungarian literals and that R-589 is
therefore open. The literals are real and the conclusion does not follow. Those literals are the
Hungarian form, and they are deliberately frozen — that IS the localisation parity guarantee. The
ENGLISH form has been rebuilt from the bundle in web.localeFuncs (internal/web/i18n_web.go,
the "updateBadge" entry) since v0.258.0, with badge.update.* in both hu.json and en.json,
pinned by TestUpdateBadgeFollowsTheLanguage, and proven live on a fresh box the same morning —
DRILL-first-hour-en-0258-2026-09-20.md item 9 reads "PASS — 'Up to date' in English (R-589, fixed
this morning, proven on a fresh box)", and R-561's own row records R-589 among the three defects
Part 0 of v0.258.0 fixed.
The register row is stale, not the code. It is closed in this session with that citation. The
lesson is the general one: a reviewer who reads one producer cannot see a second producer that
overrides it. Reading updatebadge.go alone gives exactly the wrong answer, and the file now says
so in its own comment.
1.2 The chaos-night canary is NOT a defect to fix. The brief asked why
check-image-resolvable and check-volume-persistence returned INCONCLUSIVE on 2026-09-17 and
whether the fix is bounded. They behaved exactly as designed. Both scripts' own headers state the
rule: a detector that cannot prove itself must refuse to report rather than guess —
check-image-resolvable.py cites the 2026-07-21 incident where treating a Docker Hub throttle as
failure swept 24 of 65 pins as falsely dead, and the phrase in the drill doc, "its own canary
failed", is check-volume-persistence.py's own literal error text (its self_test, which needs to
docker build a canary image and therefore needs registry access). The most-supported reading is
that registry access was constrained at that moment; that is inferred from the code plus the
documented throttle precedent, not observed — no raw stdout of the gate run survives in either
evidence directory. Round 7's response — substitute use for update, log the deviation, leave the
tree clean — is what the scripts' own documentation asks for.
No register row named it, and there is one real gap, which is diagnostic rather than behavioural:
catalog_gates.py collapses "the harness refused to run at all" and "a per-app result is
genuinely undetermined" into one INCONCLUSIVE label, so a reader cannot tell which happened
without re-running with output captured. Filed as R-605.
1.3 "The hub report carries no image field" — CONFIRMED, on both sides, with one useful nuance.
The brief confirmed the controller half (internal/report/types.go L98–103: name, state, CPU,
memory). The hub half was this session's to check. Store.SaveReport
(hub/internal/store/store.go:965) stores the raw report JSON whole and denormalises only a
fixed list out of it — controller version, URL, language, CPU/memory percent, container total and
running counts, last snapshot, health status. Containers is two integers; there is no per-app
structure anywhere, and no hub code reads an image tag.
The nuance matters for costing Slice 7: because the raw payload is kept whole, a new
controller-side field would already be stored the day the controller sends it — what is missing is
the denormalisation and the page, not the transport. That makes Slice 7 additive on both sides, as
09 §8.9 says, and slightly cheaper than "a hub-side change" suggests.
2. How far behind are the two demo boxes? Not at all.
Read on-disk on both boxes (app.yaml installed_images against the syncer's own catalog clone at
<data>/catalog-cache/templates/<app>/docker-compose.yml — the exact two inputs
stacks.CatalogOrder reads), controller v0.259.0 on both, catalog cache at 18a6d2d8 on both.
| box | deployed apps | behind | unknown (no record) | oldest catalog_since |
|---|---|---|---|---|
N100 demo-felhom-8363b5 |
1 (opengist) | 0 | 0 | 2026-07-18 |
HP t740 demo-hp-bb76ea |
9 | 0 | 0 | 2026-07-12 |
Every deployed app on the fleet reads „Naprakész" today, and every one of them has a record. The v0.234.0 backfill has done its job: the "unknown on a quiet box" residue of §8.3 is zero here.
This is the honest caveat, and it is the whole of §8.1: „Naprakész" on these boxes means the reference matches, not the image has not moved. Section 3 measures exactly how false that can be.
3. Upstream drift — the number that decides how urgent Slice 6 is
53 apps · 79 image: lines · 66 unique pins. Measured from DooPlex
against the public registries, never from a box (§8.1's rule applies to the product, not to an
audit). The dead-image sweep used the catalog's own check-image-resolvable.py --all, unmodified.
3.1 The exact pins
| count | |
|---|---|
| exact pins evaluated | 58 |
| behind upstream today | 46 |
| …of those, a MAJOR step | 7 |
| …of those, a minor/patch step | 39 |
| already newest | 10 |
| Felhom's own internal image (not upstream) | 1 |
| inconclusive | 1 |
39 of 46 are within a major — that is the population Slice 6's automatic update would serve. Seven cross a major and are human work by the 2026-09-02 ruling: nextcloud 34→35, paperless-ngx 2.20→3.2, claper 2.5→3.0, gokapi 1.9→2.2, homepage 1.13→2.4, sparkyfitness 0.17→1.7 (two images). All seven were pinned 2026-07-18 — 65 days ago.
3.2 By class
| class | apps | pins | floating | behind (of exact) | major | minor/patch |
|---|---|---|---|---|---|---|
| database | 14 | 23 | 7 | 14 of 16 | 5 | 9 |
| file-leg | 9 | 9 | 0 | 7 of 9 | 0 | 7 |
| other | 30 | 34 | 1 | 25 of 33 | 2 | 23 |
Classification caveat, stated rather than folded in: adventurelog runs a postgis/postgis
sidecar, which is PostgreSQL, but the literal substring rule puts it in other. It is the one app
that moves if postgis counts as a database (other 30→29, database 14→15). v0.240.0's R-484
already treats PostGIS/pgvector/TimescaleDB as Postgres for the logical dump, so the product is
right and only this table's rule is literal.
3.3 The floating pins — §8.1 is now MEASURED, and its COUNT was stale
First, the count itself. 09 §8.1, REUSE.md, the badge's own comment and R-446 all said
"23 of the catalog's 66 distinct pins float". That number matches no definition the catalog
supports today, and it has been corrected in all four places. Recounted at catalog 18a6d2d8,
with the definition written down so it can be rechecked: a pin FLOATS when its tag names a version
LINE rather than an exact release.
| shape | count | examples |
|---|---|---|
full X.Y.Z |
48 | privatebin/pdo:2.0.5 |
| two-part line | 6 | mariadb:11.4/11.6/12.3, claper:2.5, opengist:1.13, wger/server:2.6 |
| major line | 4 | postgres:15-alpine, postgres:16-alpine, redis:7-alpine, postgis:16-3.5-alpine |
| exact version + variant suffix | 8 | ghost:6.53.0-alpine, nextcloud:34.0.1-apache, kimai/kimai2:apache-2.57.0 |
10 of 66 float. The last row does not, by this definition — those name an exact release and only wear a flavour.
Then, the movement. The sweep measured the 8 database and cache engine pins, of which 7 were measurable:
For each floating tag, the registry's own last-push date against the date the catalog set the pin:
| pin | apps | pinned | last repushed upstream | moved? |
|---|---|---|---|---|
postgres:16-alpine |
8 apps | 2026-07-18 | 2026-09-21 | yes |
postgres:15-alpine |
sparkyfitness | 2026-07-18 | 2026-09-21 | yes |
redis:7-alpine |
6 apps | 2026-07-18 | 2026-09-18 | yes |
mariadb:11.4 |
romm | 2026-07-18 | 2026-09-18 | yes |
mariadb:12.3 |
bookstack | 2026-07-18 | 2026-09-18 | yes |
postgis/postgis:16-3.5-alpine |
adventurelog | 2026-07-18 | 2026-08-31 | yes |
mariadb:11.6 |
kimai, nextcloud | 2026-07-18 | 2025-02-04 | no |
ghcr.io/immich-app/postgres:16-vectorchord… |
immich | 2026-07-18 | — | UNMEASURED |
Six of the seven measurable floating tags have been repushed since the catalog pinned them. So on
demo-hp today, adventurelog, bookstack, docmost and romm read „Naprakész" over a database
engine image that has demonstrably moved. The badge is not lying — it answers the question it was
built to answer — but the question a household hears is the other one. R-446 is no longer a
theoretical blind spot: it is six pins out of seven. The eighth (immich's own ghcr build) could not
be measured: ghcr's anonymous API exposes no last-modified timestamp and no authenticated token was
used. Stated unmeasured rather than guessed.
3.4 One image is gone upstream
msdeluise/plant-it:0.10.0 — the gate says INCONCLUSIVE (a denied/unauthorized stderr, which it
refuses to read as "dead"), and two independent signals say it really is gone: Docker Hub's catalog
API answers object not found for the whole repository, and the upstream GitHub repo 404s. The app
is already lifecycle: abandoned, so this is the expected end state, not an incident. A deployed
plant-it would survive — nothing re-pulls a running image — but it can never be redeployed.
4. What this session then did
- controller v0.260.0 — R-524 closed: a box AHEAD of the catalog reads „Naprakész", and the
guarded Update refuses to move a pin backwards (
downgrade, 409). The comparison moved tostacks.CatalogOrderso the badge and the refusal cannot drift. Three red-proofs, each seen to fail. See the repo's CHANGELOG. - catalog — R-469: the MariaDB half of the engine-major rule lifted, with R-450's own-edge clause enforced in its place; PostgreSQL and MySQL stay refused, now citing R-463 rather than the shipped R-448. Two new decoys, two red-proofs.
- R-589 closed as already-shipped (§1.1). R-605 filed (§1.2).
Live proof of v0.260.0 and the R-520 power cut: audits/update-arc-2026-09-21/.
4a. R-524 — PROVEN LIVE, both languages, both surfaces
The state was built the honest way, not by editing a file on the box: a throwaway uptime-kuma on
the scratch guest, a real catalog move 2.4.0 → 2.5.0, the guarded Update pressed and run to
completion, then the catalog reverted to 2.4.0 — exactly the BIGNIGHT Phase 6 shape that R-524
recorded. The box then runs 2.5.0 against a 2.4.0 catalog.
The badge — app page AND apps list, both languages:
[hu] <span class="tag tag-ok" title="Ez az alkalmazás a katalógusnál újabb változatot futtat,
ezért nincs teendőd.">Naprakész</span>
[en] <span class="tag tag-ok" title="This app is running a version newer than the catalog offers,
so there is nothing for you to do.">Up to date</span>
tag-warn is absent on all four captures (negative control), and the AHEAD title is a different
sentence from the ordinary up-to-date title captured at baseline — which is the v0.260.0 behaviour,
not a coincidence of wording.
The refusal — the endpoint the button invokes:
| request | code | body |
|---|---|---|
POST /api/stacks/uptime-kuma/update |
409 | „Ez a változat újabb a katalógusban lévőnél — visszalépés csak az üzemeltető kérésére." |
the same with ?lang=en |
409 | "This version is newer than the one in the catalog — moving back needs the operator." |
the same with ?lang=hu (control) |
409 | the Hungarian sentence |
The English one is the half that was nearly missed. The new refusal was born as a bundle key, but
api.Router rendered update refusals from ref.Message — the Hungarian fallback — so without the
one-line change to errText in the same release the key would have been a seam built and never
wired. The ?lang=en row above is the proof that it is wired.
The log names both maps, so the verdict can be audited without guessing:
[ERROR] [stacks] update uptime-kuma REFUSED (downgrade): installed is provably NEWER than the
catalog on every differing service (installed=map[uptime-kuma:{louislam/uptime-kuma:2.5.0 …}]
catalog=map[uptime-kuma:louislam/uptime-kuma:2.4.0])
And NOTHING MOVED — which is the property that makes the refusal safe to add. After three refused POSTs the app is still up and healthy on 2.5.0, the live compose line still reads 2.5.0, and no update journal exists (with a positive control proving the directory searched was the real one). That is §6.1's phase-0 contract intact: a refusal taken before the intent is recorded leaves no trace, so a household that presses a refused button has changed nothing.
4b. R-520 — a power cut during a REAL version change. CLOSED, and it behaved.
What R-520 asked for, and why the 2026-09-14 attempt could not answer it: that night the only Update the catalog allowed was a same-version one, so nothing could have gone wrong and nothing did. The row asked for the same cut on a real bump.
The measurement, on the scratch guest 9202 (demo-hp), controller v0.260.0, with a throwaway
uptime-kuma and a real one-step catalog move 2.4.0 → 2.5.0 (reverted in the same session):
| 11:04:54.436Z | poller sees updating: true, update_phase: pulling |
| 11:04:54.439Z | pct stop 9202 — the cut |
| 11:04:58.207Z | stopped |
Honest limit on the instrument, recorded rather than smoothed: pct stop took 3.8 s to return, so
the phase at the moment of the decision is observed and the phase at the moment the kernel
froze is inferred.
After pct start, the box said so itself — the POSITIVE observable, not an absence:
update recovery: uptime-kuma was interrupted in pulling (started 2026-09-21T11:04:52Z)
— nothing had run; putting the pin back
pin uptime-kuma: uptime-kuma=louislam/uptime-kuma:2.4.0
update uptime-kuma: pin and definition PUT BACK to the pre-update version
pinned_imagesandinstalled_imagesboth 2.4.0; the live compose line back to 2.4.0.- No hold (correctly — nothing ran),
update_phase: failed. - The app is running and healthy on 2.4.0, 31 s after boot.
- The household is told, and the sentence is true: „A frissítés megszakadt, mert a vezérlő újraindult, mielőtt az új verzió elindult volna. Az alkalmazás a korábbi verzióval fut tovább."
R-520 closes. A power cut in pulling during a real version change is recovered honestly: the
pin goes back, the old version runs, and the page says what happened.
The instrument trap this run found, and it is worth more than the result
The first post-crash read, taken from the host against the stopped guest, reported the update journal ABSENT — which would have made this a defect finding. It was a false negative.
pct mount maps the guest's ROOTFS ONLY and does not apply the guest's own internal mounts.
Guest 9202 keeps /var/lib/docker on its own ext4 mount under the mp0 volume; from the host that
path exists and is an empty stub. The journal was at
/var/lib/lxc/9202/rootfs/var/lib/felhom/docker/volumes/felhom-controller-data/_data/data/, and the
box's own recovery read it at the next boot — which is the proof that it was there all along.
It was caught by running find over the whole rootfs instead of trusting one constructed path,
and by demanding a positive control for the directory being searched. Read from the right place, the
journal held exactly what §6.1 promises: phase: "pulling", prev_pin: …2.4.0. An empty directory
is not evidence of an absent file — it is equally consistent with "you are looking at an unmounted
stub". That is R-96 rule 3 in a new place.
A second instrument trap in the same run — and it nearly bought a false proof
POST /api/sync answered „nincs változás" while the box's catalog cache file HAD moved to the new
tag, and the API's catalog_images stayed at the old value until a separate POST /api/stacks/rescan.
Since CatalogImages is the one input stacks.CatalogOrder compares against, the badge answers
from a stale catalog for that window — and the session nearly recorded a tag-ok badge as proof of
the R-524 ahead arm when the badge was merely stale. The honest reading came only after the rescan.
An instrument that can report an old value as a current one is not a measurement. Filed as
R-607; neither half of it was isolated, so the row records the observation, not a diagnosis.
4c. A defect the live run found — filed, not fixed today
On the English app page, the badge is English and the sentence under it is Hungarian. Captured verbatim at step 5, directly above one another:
[en] <span class="tag tag-warn" title="A newer version of this app is available…">Update available — today</span>
A frissítés megszakadt, mert a vezérlő újraindult, mielőtt az új verzió elindult volna.
Az alkalmazás a korábbi verzióval fut tovább.
Stack.UpdateError is a finished Hungarian STRING, not a key. Manager.finishUpdate stores it from
eight raw literals in update.go, and both app_info.html and stacks.html render it verbatim.
The same path carries UpdatePhaseLabel and the HOLD sentence — including
backup.UpdateCopyHolds's „csak a beállításokat és az adatbázist tartalmazza, a fájlokat nem",
which is a promise about whether the customer's files come back.
v0.260.0 did not close this. That release routed the 409 REFUSAL through errText, and a refusal
is the path where nothing happened; these are the sentences for when something did. Filed as
R-606 (P2), not fixed — one release per repo per session, and today's had already shipped.
4d. Teardown — three layers, stated
| layer | state |
|---|---|
| machine (guest 9202) | throwaway uptime-kuma removed through the product — the API refused the remove while it ran (409, "stop it first"), so it was stopped through the product and then removed, taking its volume and its 24K backup. Only the three protected infra containers remain, as at baseline. No update journal. |
| host (demo-hp) | guest 9202 left running, as it was found. Guest 9201 untouched — 24 containers before and after. felhom-pve not involved. |
| hub | nothing provisioned, nothing enrolled, nothing discarded. Guest 9202 reports to no hub by design. |
The catalog is back where it started: louislam/uptime-kuma:2.4.0, tree clean and level with
origin/main. catalog_since reads 2026-09-21 rather than 2026-07-18 — the gate requires an image
move to carry the day's date, and the bump and its revert are two moves that net to zero. No box is
behind on that app, so no age count is affected.
5. What the numbers say about the two open slices
Slice 6 (automatic within a major) is worth more than its P2 suggests, and the reason is §3.1: 39 of 58 exact pins are one minor step behind, every one of them a change the 2026-09-02 ruling already says may happen without a human. Nobody presses 39 buttons. The support window (§3 decision 2) runs on how far behind the catalog a box is, so today every box is inside it only because the catalog has not moved since 2026-09-15 — the moment the catalog catches up with upstream, every box on the fleet is 46 pins behind at once.
Slice 7 (the fleet view) is still correctly P3-LOW at two enrolled boxes — but §2 shows why it will not stay cheap: the only reason a person could answer "is the fleet current?" today is that a session read both boxes' files by hand.
R-446 (digests) is the cheapest real improvement on this list. §3.3 measured the blind spot at
six pins of seven, and the catalog's check-image-resolvable.py already resolves a digest for every
pin at push time. Recording it costs one field and closes the gap without the box ever reaching a
registry.