Files
felhom.eu/documentation/audits/UPDATE-ARC-STATE-2026-09-21.md
T
admin d19f07ea04
gates / gates (push) Successful in 26s
File R-607 (a sync that says 'no change' while the cache moves) and sharpen the audit
Found by the live run, not by reading: POST /api/sync answered 'nincs valtozas'
while the box's catalog cache HAD moved, and catalog_images stayed stale until a
separate rescan. Since CatalogImages is the one input CatalogOrder compares
against, the badge answers from a stale catalog for that window — and the session
nearly recorded a stale tag-ok badge as proof of the R-524 ahead arm.

Neither half is isolated, so the row records the observation, not a diagnosis.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-21 13:14:10 +02:00

20 KiB
Raw Blame History

UPDATE ARC — the state, measured (2026-09-21)

Phase 0 of the resumed update arc. Nothing here is an estimate. Every number was read from a live box, a live registry or live source on 2026-09-21. Raw captures: audits/update-arc-2026-09-21/.

Baselines at the start: controller 19ef0329ab66 v0.259.0 · agent d9864a94bf62 v0.132.0 · felhom.eu bcdd5b205875 hub v0.119.0 · catalog 18a6d2d8243e. All four trees clean and level with origin/main.


1. Three claims in the brief that turned out wrong — named first

1.1 R-589 is NOT open. It shipped in v0.258.0. The brief said, reviewer-verified, that updatebadge.go L82–L99 builds the badge from four raw Hungarian literals and that R-589 is therefore open. The literals are real and the conclusion does not follow. Those literals are the Hungarian form, and they are deliberately frozen — that IS the localisation parity guarantee. The ENGLISH form has been rebuilt from the bundle in web.localeFuncs (internal/web/i18n_web.go, the "updateBadge" entry) since v0.258.0, with badge.update.* in both hu.json and en.json, pinned by TestUpdateBadgeFollowsTheLanguage, and proven live on a fresh box the same morning — DRILL-first-hour-en-0258-2026-09-20.md item 9 reads "PASS — 'Up to date' in English (R-589, fixed this morning, proven on a fresh box)", and R-561's own row records R-589 among the three defects Part 0 of v0.258.0 fixed.

The register row is stale, not the code. It is closed in this session with that citation. The lesson is the general one: a reviewer who reads one producer cannot see a second producer that overrides it. Reading updatebadge.go alone gives exactly the wrong answer, and the file now says so in its own comment.

1.2 The chaos-night canary is NOT a defect to fix. The brief asked why check-image-resolvable and check-volume-persistence returned INCONCLUSIVE on 2026-09-17 and whether the fix is bounded. They behaved exactly as designed. Both scripts' own headers state the rule: a detector that cannot prove itself must refuse to report rather than guess — check-image-resolvable.py cites the 2026-07-21 incident where treating a Docker Hub throttle as failure swept 24 of 65 pins as falsely dead, and the phrase in the drill doc, "its own canary failed", is check-volume-persistence.py's own literal error text (its self_test, which needs to docker build a canary image and therefore needs registry access). The most-supported reading is that registry access was constrained at that moment; that is inferred from the code plus the documented throttle precedent, not observed — no raw stdout of the gate run survives in either evidence directory. Round 7's response — substitute use for update, log the deviation, leave the tree clean — is what the scripts' own documentation asks for.

No register row named it, and there is one real gap, which is diagnostic rather than behavioural: catalog_gates.py collapses "the harness refused to run at all" and "a per-app result is genuinely undetermined" into one INCONCLUSIVE label, so a reader cannot tell which happened without re-running with output captured. Filed as R-605.

1.3 "The hub report carries no image field" — CONFIRMED, on both sides, with one useful nuance. The brief confirmed the controller half (internal/report/types.go L98–103: name, state, CPU, memory). The hub half was this session's to check. Store.SaveReport (hub/internal/store/store.go:965) stores the raw report JSON whole and denormalises only a fixed list out of it — controller version, URL, language, CPU/memory percent, container total and running counts, last snapshot, health status. Containers is two integers; there is no per-app structure anywhere, and no hub code reads an image tag.

The nuance matters for costing Slice 7: because the raw payload is kept whole, a new controller-side field would already be stored the day the controller sends it — what is missing is the denormalisation and the page, not the transport. That makes Slice 7 additive on both sides, as 09 §8.9 says, and slightly cheaper than "a hub-side change" suggests.


2. How far behind are the two demo boxes? Not at all.

Read on-disk on both boxes (app.yaml installed_images against the syncer's own catalog clone at <data>/catalog-cache/templates/<app>/docker-compose.yml — the exact two inputs stacks.CatalogOrder reads), controller v0.259.0 on both, catalog cache at 18a6d2d8 on both.

box deployed apps behind unknown (no record) oldest catalog_since
N100 demo-felhom-8363b5 1 (opengist) 0 0 2026-07-18
HP t740 demo-hp-bb76ea 9 0 0 2026-07-12

Every deployed app on the fleet reads „Naprakész" today, and every one of them has a record. The v0.234.0 backfill has done its job: the "unknown on a quiet box" residue of §8.3 is zero here.

This is the honest caveat, and it is the whole of §8.1: „Naprakész" on these boxes means the reference matches, not the image has not moved. Section 3 measures exactly how false that can be.


3. Upstream drift — the number that decides how urgent Slice 6 is

53 apps · 79 image: lines · 66 unique pins. Measured from DooPlex against the public registries, never from a box (§8.1's rule applies to the product, not to an audit). The dead-image sweep used the catalog's own check-image-resolvable.py --all, unmodified.

3.1 The exact pins

count
exact pins evaluated 58
behind upstream today 46
…of those, a MAJOR step 7
…of those, a minor/patch step 39
already newest 10
Felhom's own internal image (not upstream) 1
inconclusive 1

39 of 46 are within a major — that is the population Slice 6's automatic update would serve. Seven cross a major and are human work by the 2026-09-02 ruling: nextcloud 34→35, paperless-ngx 2.20→3.2, claper 2.5→3.0, gokapi 1.9→2.2, homepage 1.13→2.4, sparkyfitness 0.17→1.7 (two images). All seven were pinned 2026-07-18 — 65 days ago.

3.2 By class

class apps pins floating behind (of exact) major minor/patch
database 14 23 7 14 of 16 5 9
file-leg 9 9 0 7 of 9 0 7
other 30 34 1 25 of 33 2 23

Classification caveat, stated rather than folded in: adventurelog runs a postgis/postgis sidecar, which is PostgreSQL, but the literal substring rule puts it in other. It is the one app that moves if postgis counts as a database (other 30→29, database 14→15). v0.240.0's R-484 already treats PostGIS/pgvector/TimescaleDB as Postgres for the logical dump, so the product is right and only this table's rule is literal.

3.3 The floating pins — §8.1 is now MEASURED, and its COUNT was stale

First, the count itself. 09 §8.1, REUSE.md, the badge's own comment and R-446 all said "23 of the catalog's 66 distinct pins float". That number matches no definition the catalog supports today, and it has been corrected in all four places. Recounted at catalog 18a6d2d8, with the definition written down so it can be rechecked: a pin FLOATS when its tag names a version LINE rather than an exact release.

shape count examples
full X.Y.Z 48 privatebin/pdo:2.0.5
two-part line 6 mariadb:11.4/11.6/12.3, claper:2.5, opengist:1.13, wger/server:2.6
major line 4 postgres:15-alpine, postgres:16-alpine, redis:7-alpine, postgis:16-3.5-alpine
exact version + variant suffix 8 ghost:6.53.0-alpine, nextcloud:34.0.1-apache, kimai/kimai2:apache-2.57.0

10 of 66 float. The last row does not, by this definition — those name an exact release and only wear a flavour.

Then, the movement. The sweep measured the 8 database and cache engine pins, of which 7 were measurable:

For each floating tag, the registry's own last-push date against the date the catalog set the pin:

pin apps pinned last repushed upstream moved?
postgres:16-alpine 8 apps 2026-07-18 2026-09-21 yes
postgres:15-alpine sparkyfitness 2026-07-18 2026-09-21 yes
redis:7-alpine 6 apps 2026-07-18 2026-09-18 yes
mariadb:11.4 romm 2026-07-18 2026-09-18 yes
mariadb:12.3 bookstack 2026-07-18 2026-09-18 yes
postgis/postgis:16-3.5-alpine adventurelog 2026-07-18 2026-08-31 yes
mariadb:11.6 kimai, nextcloud 2026-07-18 2025-02-04 no
ghcr.io/immich-app/postgres:16-vectorchord… immich 2026-07-18 — UNMEASURED

Six of the seven measurable floating tags have been repushed since the catalog pinned them. So on demo-hp today, adventurelog, bookstack, docmost and romm read „Naprakész" over a database engine image that has demonstrably moved. The badge is not lying — it answers the question it was built to answer — but the question a household hears is the other one. R-446 is no longer a theoretical blind spot: it is six pins out of seven. The eighth (immich's own ghcr build) could not be measured: ghcr's anonymous API exposes no last-modified timestamp and no authenticated token was used. Stated unmeasured rather than guessed.

3.4 One image is gone upstream

msdeluise/plant-it:0.10.0 — the gate says INCONCLUSIVE (a denied/unauthorized stderr, which it refuses to read as "dead"), and two independent signals say it really is gone: Docker Hub's catalog API answers object not found for the whole repository, and the upstream GitHub repo 404s. The app is already lifecycle: abandoned, so this is the expected end state, not an incident. A deployed plant-it would survive — nothing re-pulls a running image — but it can never be redeployed.


4. What this session then did

  • controller v0.260.0 — R-524 closed: a box AHEAD of the catalog reads „Naprakész", and the guarded Update refuses to move a pin backwards (downgrade, 409). The comparison moved to stacks.CatalogOrder so the badge and the refusal cannot drift. Three red-proofs, each seen to fail. See the repo's CHANGELOG.
  • catalog — R-469: the MariaDB half of the engine-major rule lifted, with R-450's own-edge clause enforced in its place; PostgreSQL and MySQL stay refused, now citing R-463 rather than the shipped R-448. Two new decoys, two red-proofs.
  • R-589 closed as already-shipped (§1.1). R-605 filed (§1.2).

Live proof of v0.260.0 and the R-520 power cut: audits/update-arc-2026-09-21/.


4a. R-524 — PROVEN LIVE, both languages, both surfaces

The state was built the honest way, not by editing a file on the box: a throwaway uptime-kuma on the scratch guest, a real catalog move 2.4.0 → 2.5.0, the guarded Update pressed and run to completion, then the catalog reverted to 2.4.0 — exactly the BIGNIGHT Phase 6 shape that R-524 recorded. The box then runs 2.5.0 against a 2.4.0 catalog.

The badge — app page AND apps list, both languages:

[hu] <span class="tag tag-ok" title="Ez az alkalmazás a katalógusnál újabb változatot futtat,
                                     ezért nincs teendőd.">Naprakész</span>
[en] <span class="tag tag-ok" title="This app is running a version newer than the catalog offers,
                                     so there is nothing for you to do.">Up to date</span>

tag-warn is absent on all four captures (negative control), and the AHEAD title is a different sentence from the ordinary up-to-date title captured at baseline — which is the v0.260.0 behaviour, not a coincidence of wording.

The refusal — the endpoint the button invokes:

request code body
POST /api/stacks/uptime-kuma/update 409 „Ez a változat újabb a katalógusban lévőnél — visszalépés csak az üzemeltető kérésére."
the same with ?lang=en 409 "This version is newer than the one in the catalog — moving back needs the operator."
the same with ?lang=hu (control) 409 the Hungarian sentence

The English one is the half that was nearly missed. The new refusal was born as a bundle key, but api.Router rendered update refusals from ref.Message — the Hungarian fallback — so without the one-line change to errText in the same release the key would have been a seam built and never wired. The ?lang=en row above is the proof that it is wired.

The log names both maps, so the verdict can be audited without guessing:

[ERROR] [stacks] update uptime-kuma REFUSED (downgrade): installed is provably NEWER than the
catalog on every differing service (installed=map[uptime-kuma:{louislam/uptime-kuma:2.5.0 …}]
catalog=map[uptime-kuma:louislam/uptime-kuma:2.4.0])

And NOTHING MOVED — which is the property that makes the refusal safe to add. After three refused POSTs the app is still up and healthy on 2.5.0, the live compose line still reads 2.5.0, and no update journal exists (with a positive control proving the directory searched was the real one). That is §6.1's phase-0 contract intact: a refusal taken before the intent is recorded leaves no trace, so a household that presses a refused button has changed nothing.


4b. R-520 — a power cut during a REAL version change. CLOSED, and it behaved.

What R-520 asked for, and why the 2026-09-14 attempt could not answer it: that night the only Update the catalog allowed was a same-version one, so nothing could have gone wrong and nothing did. The row asked for the same cut on a real bump.

The measurement, on the scratch guest 9202 (demo-hp), controller v0.260.0, with a throwaway uptime-kuma and a real one-step catalog move 2.4.0 → 2.5.0 (reverted in the same session):

11:04:54.436Z poller sees updating: true, update_phase: pulling
11:04:54.439Z pct stop 9202 — the cut
11:04:58.207Z stopped

Honest limit on the instrument, recorded rather than smoothed: pct stop took 3.8 s to return, so the phase at the moment of the decision is observed and the phase at the moment the kernel froze is inferred.

After pct start, the box said so itself — the POSITIVE observable, not an absence:

update recovery: uptime-kuma was interrupted in pulling (started 2026-09-21T11:04:52Z)
                 — nothing had run; putting the pin back
pin uptime-kuma: uptime-kuma=louislam/uptime-kuma:2.4.0
update uptime-kuma: pin and definition PUT BACK to the pre-update version
  • pinned_images and installed_images both 2.4.0; the live compose line back to 2.4.0.
  • No hold (correctly — nothing ran), update_phase: failed.
  • The app is running and healthy on 2.4.0, 31 s after boot.
  • The household is told, and the sentence is true: „A frissítés megszakadt, mert a vezérlő újraindult, mielőtt az új verzió elindult volna. Az alkalmazás a korábbi verzióval fut tovább."

R-520 closes. A power cut in pulling during a real version change is recovered honestly: the pin goes back, the old version runs, and the page says what happened.

The instrument trap this run found, and it is worth more than the result

The first post-crash read, taken from the host against the stopped guest, reported the update journal ABSENT — which would have made this a defect finding. It was a false negative.

pct mount maps the guest's ROOTFS ONLY and does not apply the guest's own internal mounts. Guest 9202 keeps /var/lib/docker on its own ext4 mount under the mp0 volume; from the host that path exists and is an empty stub. The journal was at /var/lib/lxc/9202/rootfs/var/lib/felhom/docker/volumes/felhom-controller-data/_data/data/, and the box's own recovery read it at the next boot — which is the proof that it was there all along.

It was caught by running find over the whole rootfs instead of trusting one constructed path, and by demanding a positive control for the directory being searched. Read from the right place, the journal held exactly what §6.1 promises: phase: "pulling", prev_pin: …2.4.0. An empty directory is not evidence of an absent file — it is equally consistent with "you are looking at an unmounted stub". That is R-96 rule 3 in a new place.

A second instrument trap in the same run — and it nearly bought a false proof

POST /api/sync answered „nincs változás" while the box's catalog cache file HAD moved to the new tag, and the API's catalog_images stayed at the old value until a separate POST /api/stacks/rescan. Since CatalogImages is the one input stacks.CatalogOrder compares against, the badge answers from a stale catalog for that window — and the session nearly recorded a tag-ok badge as proof of the R-524 ahead arm when the badge was merely stale. The honest reading came only after the rescan. An instrument that can report an old value as a current one is not a measurement. Filed as R-607; neither half of it was isolated, so the row records the observation, not a diagnosis.


4c. A defect the live run found — filed, not fixed today

On the English app page, the badge is English and the sentence under it is Hungarian. Captured verbatim at step 5, directly above one another:

[en] <span class="tag tag-warn" title="A newer version of this app is available…">Update available — today</span>
     A frissítés megszakadt, mert a vezérlő újraindult, mielőtt az új verzió elindult volna.
     Az alkalmazás a korábbi verzióval fut tovább.

Stack.UpdateError is a finished Hungarian STRING, not a key. Manager.finishUpdate stores it from eight raw literals in update.go, and both app_info.html and stacks.html render it verbatim. The same path carries UpdatePhaseLabel and the HOLD sentence — including backup.UpdateCopyHolds's „csak a beállításokat és az adatbázist tartalmazza, a fájlokat nem", which is a promise about whether the customer's files come back.

v0.260.0 did not close this. That release routed the 409 REFUSAL through errText, and a refusal is the path where nothing happened; these are the sentences for when something did. Filed as R-606 (P2), not fixed — one release per repo per session, and today's had already shipped.


4d. Teardown — three layers, stated

layer state
machine (guest 9202) throwaway uptime-kuma removed through the product — the API refused the remove while it ran (409, "stop it first"), so it was stopped through the product and then removed, taking its volume and its 24K backup. Only the three protected infra containers remain, as at baseline. No update journal.
host (demo-hp) guest 9202 left running, as it was found. Guest 9201 untouched — 24 containers before and after. felhom-pve not involved.
hub nothing provisioned, nothing enrolled, nothing discarded. Guest 9202 reports to no hub by design.

The catalog is back where it started: louislam/uptime-kuma:2.4.0, tree clean and level with origin/main. catalog_since reads 2026-09-21 rather than 2026-07-18 — the gate requires an image move to carry the day's date, and the bump and its revert are two moves that net to zero. No box is behind on that app, so no age count is affected.


5. What the numbers say about the two open slices

Slice 6 (automatic within a major) is worth more than its P2 suggests, and the reason is §3.1: 39 of 58 exact pins are one minor step behind, every one of them a change the 2026-09-02 ruling already says may happen without a human. Nobody presses 39 buttons. The support window (§3 decision 2) runs on how far behind the catalog a box is, so today every box is inside it only because the catalog has not moved since 2026-09-15 — the moment the catalog catches up with upstream, every box on the fleet is 46 pins behind at once.

Slice 7 (the fleet view) is still correctly P3-LOW at two enrolled boxes — but §2 shows why it will not stay cheap: the only reason a person could answer "is the fleet current?" today is that a session read both boxes' files by hand.

R-446 (digests) is the cheapest real improvement on this list. §3.3 measured the blind spot at six pins of seven, and the catalog's check-image-resolvable.py already resolves a digest for every pin at push time. Recording it costs one field and closes the gap without the box ever reaching a registry.