Update night 2026-09-21: the full record, twelve rows, and the answers to five of the seven questions
gates / gates (push) Successful in 28s
gates / gates (push) Successful in 28s
The drill is complete. Teardown done in three layers plus Gitea; the live catalog's every `image:` line is proven identical to before. WHAT WAS MEASURED. 21 edges across 19 apps, on scratch guest 9202 through the product's own guarded Update, against a PRIVATE DRILL CATALOG so the live catalog carried no test reference at any point: 14 proven, 3 failed, 4 inconclusive. Each app seeded and read back through its OWN front door, with a negative control on every readback. Ten of the fourteen printed a verbatim migration line. Up from the three apps this project had ever measured. THE RESULT THAT MATTERS. R-618, P1: three of the 53 templates name a health probe the app does not answer, and because the guarded update WAITS on that same probe, a SUCCESSFUL update ends by STOPPING a working app. tandoor was measured serving HTTP 200 on the new version at four samples across five minutes, docker's own healthcheck green, and was then stopped and the household sent to a restore they did not need. zipline and wger are the same defect, both confirmed live. The gate that catches all three is static and cheap: both health checks already sit in the same file. WHAT THE NIGHT ANSWERED that was open. The UNATTENDED HOLD (312.9 s, pressed once, never again) — which needed a purpose-built image store, because the rule that makes automatic updates safe is the same rule that refuses the obvious way to break one. MariaDB across a major through the real button, all four observables, first time. PostgreSQL across a major, refusing exactly as predicted, with the conversion costed at ~9 s of engine work. There is NO single-flight: five updates ran at once and all ended honest. And the two EARLY power-cut phases nobody had cut in. TWELVE NEW ROWS (R-615..R-626), register 303 -> 315, and eight existing rows updated with what was measured — including two CORRECTIONS: R-606 records the pre-flight refusals as reaching an English household in English and they do not, and R-446/R-458 are both narrower than their rows state. Two instrument fixes were needed before anything could be trusted: the unattended caller turned every success into a timeout (R-623), and one of my own reproductions was wrong and is kept labelled with what it actually measured. Interventions: zero. No controller, agent or hub code written. The hub was never touched beyond the floor the operator asked for. Gates: repo_gates.py --fast, all 15 OK. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
@@ -102,7 +102,7 @@ likewise silent. Evidence: `audits/DRILL-r361-2026-08-22/evidence/06-part3-decis
|
||||
|---|---|---|---|---|
|
||||
| Deploy an app from the catalog (env config, memory guard, health-aware progress) | controller, catalog (~52 apps, images pinned) | **PROVEN-LIVE** | `CAMPAIGN-2` T-DEPLOY-SET (7 apps, env config, health-aware); `RERUN-p1p3` (×4 PASS) | Memory-guard FIRING is not live-shown (T-RES-MEMGUARD never fired: ample RAM / auth-walled) — implemented + unit-level only |
|
||||
| App lifecycle: start/stop/restart/update/logs/remove/redeploy | controller | **PROVEN-LIVE — the ACTIONS work. NARROWED 2026-09-13: `CAMPAIGN-3` proved `remove` removes the APP, not the DATA — the "delete my data" half was INERT on every box until controller v0.236.0 (R-442). RE-PROVEN 2026-09-13 on demo-hp: data written by the app itself (63 MB) gone after removal and listed; an unresolvable data location is REFUSED (409) with the app kept; an SSD app gets `[]` and a note.** | `CAMPAIGN-2` T-LIFECYCLE (stop/start/restart/update/logs); remove (app only) live in `CAMPAIGN-3`; **remove WITH data: `audits/R442-2026-09-13/`**; **data behaviour: `audits/SPIKE-app-update-2026-09-01.md` (2026-09-01)** | Redeploy-after-remove edge remains open (T-REMOVE-REDEPLOY never cleanly passed — stale dryrun journal); non-pilot-critical |
|
||||
| **Update is GUARDED: it refuses without a restorable backup, backs up first when the copy is stale, and HOLDS an app that does not come up — on ANY backup tier, and the release itself arrives by the managed floor** | controller **v0.237.0 + v0.238.0 + v0.238.1 + v0.239.0**, hub **v0.112.0** | **PROVEN-LIVE (2026-09-13, and again the same afternoon for any tier + floor delivery)** — **afternoon (`audits/rulings-r472-r475-2026-09-13/`):** an undeclared floor above the golden refused with nothing stored (02); a declared floor 0.239.0 / MinAgent 0.129.0 served `from declared` and both demo boxes self-updated in 14 s and 15 s (03); nothing on any tier → backed up first, Tier 1 chosen, done (04); gokapi updated on its Tier-1 unit alone (05); a never-healthy update held naming „saját meghajtó" (07); restored from „helyi", hold cleared (08). **Morning:** scenarios A (real upgrade, success only after health), B (stale copy → backup first), E (pull failure → pin back, app untouched), F (never healthy → held, hold text on API and page), H (start/restart/update and the boot sweep all refuse the held app) and **the restore walk** (Mentések unit restore → back on the old version, hold cleared), on demo-hp with a throwaway app | **`audits/slice4-2026-09-13/`** (live/, redproofs/, gates/); design `architecture/09-update-architecture.md` §6.1 | ~~**Tier-2-only precondition**~~ — superseded by v0.239.0 (any tier, R-475 CLOSED); a Tier-1 route back restores only what the unit holds (R-479); the card keeps the failure sentence after a successful restore (R-480); no automatic rollback, by measurement; a release does not reach the fleet by floor between golden bakes (R-472) |
|
||||
| **Update is GUARDED: it refuses without a restorable backup, backs up first when the copy is stale, and HOLDS an app that does not come up — on ANY backup tier, and the release itself arrives by the managed floor** | controller **v0.237.0 + v0.238.0 + v0.238.1 + v0.239.0**, hub **v0.112.0** | **PROVEN-LIVE (2026-09-13, and again the same afternoon for any tier + floor delivery)** — **afternoon (`audits/rulings-r472-r475-2026-09-13/`):** an undeclared floor above the golden refused with nothing stored (02); a declared floor 0.239.0 / MinAgent 0.129.0 served `from declared` and both demo boxes self-updated in 14 s and 15 s (03); nothing on any tier → backed up first, Tier 1 chosen, done (04); gokapi updated on its Tier-1 unit alone (05); a never-healthy update held naming „saját meghajtó" (07); restored from „helyi", hold cleared (08). **Morning:** scenarios A (real upgrade, success only after health), B (stale copy → backup first), E (pull failure → pin back, app untouched), F (never healthy → held, hold text on API and page), H (start/restart/update and the boot sweep all refuse the held app) and **the restore walk** (Mentések unit restore → back on the old version, hold cleared), on demo-hp with a throwaway app | **`audits/slice4-2026-09-13/`** (live/, redproofs/, gates/); design `architecture/09-update-architecture.md` §6.1 | ~~**Tier-2-only precondition**~~ — superseded by v0.239.0 (any tier, R-475 CLOSED); a Tier-1 route back restores only what the unit holds (R-479); the card keeps the failure sentence after a successful restore (R-480); no automatic rollback, by measurement; a release does not reach the fleet by floor between golden bakes (R-472) **WIDENED 2026-09-21 (the update night) from 3 apps to 21 edges across 19 apps, and NARROWED in one place by the same run.** `audits/DRILL-update-night-2026-09-21.md`. On scratch guest 9202 (controller v0.261.0), against a **private drill catalog** so the live catalog carried no test reference at any point, 21 edges across 19 apps real within-a-major upstream edges were walked through the product's own guarded Update, each app seeded and read back **through its own front door** (R-156) with a negative control on every readback: **14 proven, 3 failed, 4 inconclusive.** **What the PROVEN edges prove, precisely:** the app moved, the four version observables agreed, and the data the app itself was given came back through the app's own interface afterwards. Ten of them printed a verbatim migration line. **What the FAILED edges prove, and they are the more valuable half.** `adventurelog` (a real upstream edge that migrates and then never serves), `tandoor` (an update that SUCCEEDED and was stopped by its own wrong health port), and the PostgreSQL engine major, which refused exactly as predicted. `adventurelog v0.12.1 → v0.13.0` applied **nine database migrations successfully** and then never bound its port; the update held after the full health wait, the hold sentence named the tier, the date and what the copy holds, and the restore the sentence names brought the app back. **That is this row's own promise, exercised on a real upstream edge rather than a staged one.** **AND THE NARROWING, which this row must carry because it is the same mechanism:** the `verifying` phase trusts the `.felhom.yml` probe absolutely, and **two of the 53 templates name a probe the app does not answer** — `tandoor` (port 8080; it listens on 80) and `zipline` (`/api/health`; it answers 404 there, while the compose healthcheck in the same file uses `/api/healthcheck` and is green). For those apps a **successful** update is stopped by its own health wait: tandoor was measured **serving HTTP 200 on the new version at four samples across five minutes**, with docker's own healthcheck green, and was then stopped by `failAndHold` and the household sent to a restore they did not need. **R-618, P1.** No data was lost and the restore works — but "the update is guarded" must not be read as "the guard is right about whether the app came up". **Still true and unchanged:** no automatic rollback (by measurement); the route back is the restore; a multi-major jump ends held honestly. **Not measured on this venue, and named rather than assumed:** every event and every customer mail. Guest 9202 runs `hub.enabled: false` and the notifier returns before it logs (**R-620**), so the whole "who was told" half of `08` was structurally unobservable tonight. |
|
||||
| **What `restart` and `update` do to a deployed app whose compose file the catalog already moved** | controller **v0.235.0** | **CHANGED 2026-09-06 — they NO LONGER upgrade it.** The row below records what shipped; this text records what it replaced, because every box under v0.235.0 still behaves the old way. **Up to v0.234.0: PROVEN-LIVE (2026-09-01) — they UPGRADE it.** Every lifecycle action ends in `docker compose up -d`, which makes the container match the file and PULLS the image itself when it is missing (measured: 18.3 s with a pull, 0.5 s without; negative control with an unchanged file did not even recreate the container). This is DELIBERATE on the restart path — `Manager.RestartStack` says so in a comment — but the syncer moves the file under a deployed app on a 15-minute cycle with no deployed check (R-438), and NOTHING tells the customer. | `audits/SPIKE-app-update-2026-09-01.md` §2, §3 | **No safety copy is taken by any of them** — `writeSafetyDump` is DATABASE-ONLY and is not on the update path at all. R-438, R-440, R-443. |
|
||||
| **Whether the box UPGRADES an app by itself, with nobody pressing anything** | controller | **PROVEN-LIVE (2026-09-01) — YES, but only when an app fails to come back.** A plain power cut does NOT upgrade: Docker's `restart: unless-stopped` restores the old containers and the reconciler logs `no boot-orphaned apps (nothing to start)`. When an app does NOT return, `Reconciler.Run` (`bootrecon.go:269`) calls `StartStack` -> `compose up -d` and the app comes back on the NEW version, unattended (measured). **13 non-API call sites across 9 files reach `up -d` this way** — not the five previously believed. | `audits/SPIKE-app-update-2026-09-01.md` §2, §8 | The drive-return gate (`intermediary.go:222`) and `AppStopGuard.Recover` (`appstop_marker.go:283`) call the same function; located by reading, **not exercised live** — stated as such. |
|
||||
| **Whether an app UPGRADE can be undone** | controller + catalog | **PROVEN-LIVE (2026-09-01) — NO, and "rollback" is the wrong word for it.** Once a migration has RUN, putting the old image tag back yields a container that refuses to start: Nextcloud — *"the version of the data (32.0.9.2) is higher than the docker image version (31.0.14.1) and downgrading is not supported"*. A 3-major jump is refused outright (*"only possible to upgrade one major version at a time"*) and IS recoverable, precisely because nothing migrated. Positive control: the data is not destroyed — returning to 32.0.9 restored both seeded markers byte-identical. | `audits/SPIKE-app-update-2026-09-01.md` §7 | The only route back is restoring DATA from a copy taken BEFORE the update — which no update path takes. And a restore's image-level rollback is itself overwritten by the syncer within 15 minutes (R-441). R-40 is confirmed live by the same measurement. |
|
||||
|
||||
@@ -266,6 +266,25 @@ update **proceeds**, and the household is told what the copy holds.
|
||||
**If nothing is decided:** Slice 6 must be built for the safe subset only, and the file-leg apps stay
|
||||
manual — which is the third option by default, without anyone choosing it.
|
||||
|
||||
|
||||
**MEASURED 2026-09-21 (update night).** The hold sentence this question turns on was read verbatim
|
||||
off a REAL failure rather than from source. `adventurelog v0.12.1 -> v0.13.0` applied nine database
|
||||
migrations successfully, never bound its port, and held:
|
||||
|
||||
> „A(z) adventurelog frissitese 2026-09-21 20:53-kor nem sikerult, es az alkalmazas nem indult el az
|
||||
> uj verzioval. Az alkalmazas biztonsagi okbol leallitva marad, hogy az adatai ne serüljenek.
|
||||
> Visszaallithato a Mentesek oldalon ebbol a biztonsagi mentesbol: **sajat meghajto, 2026-09-21 20:47
|
||||
> — ez a masolat a beallitasokat, az adatbazist es az adatkoteteket tartalmazza.**"
|
||||
|
||||
(ASCII fragments here; the live page carries its accents.) So the machinery this question's first
|
||||
option would key on **exists and works**: the sentence names the tier, the date and **what the copy
|
||||
holds**, unprompted, on a real edge. Whether the AUTOMATIC rule should differ from the button's is
|
||||
untouched by that and remains the operator's.
|
||||
|
||||
**And one thing Q2 did not ask, which tonight makes urgent: after the hold, nobody can find out WHY.**
|
||||
`failAndHold` removes the containers, so the failing version's own output is gone within seconds
|
||||
(**R-621**). With a person pressing, they at least watched it happen.
|
||||
|
||||
### Q3 — What counts as "within a major" when the tag is not a version number?
|
||||
|
||||
*§3 decision 3 says automatic within a major, never across. What about `postgres:16-alpine`,
|
||||
@@ -289,6 +308,15 @@ one comparator, never a second one.**
|
||||
**If nothing is decided:** Slice 6 would have to invent a rule under time pressure, which is how a
|
||||
major gets automated by accident.
|
||||
|
||||
|
||||
**MEASURED 2026-09-21 by RUNNING the comparator rather than reading it.** `CompareImageRefs` orders a
|
||||
reference carrying a `host:port/` prefix correctly — `splitImageRef` takes the LAST colon and rejects
|
||||
it only when a `/` follows, so a registry port is never mistaken for a tag. Four positive cases and
|
||||
one negative control (different repositories are not orderable). **This is what made the unattended
|
||||
hold measurable at all**: the drill edge `localhost:5000/drill/glance:1.0.0 -> :1.0.1` PASSES the
|
||||
within-a-major test and still fails, which no real catalog move does. The recommendation is
|
||||
unchanged; the *same major?* extension it already names is still owed.
|
||||
|
||||
### Q4 — A held app: who is told, when, and does the box try again?
|
||||
|
||||
*An automatic update that ends HELD happened while everyone was asleep.*
|
||||
@@ -319,6 +347,33 @@ within-a-major test and still fails its health check — same repository, same m
|
||||
starts and does not serve — which probably means a purpose-built image rather than a catalog move.
|
||||
**So this question still rests on the ATTENDED hold measured in slice 4 (v0.238.0, Scenario F).**
|
||||
|
||||
|
||||
**MEASURED 2026-09-21 (update night) — and this is the half that was missing.** The caller pressed
|
||||
ONCE with nobody watching; the app held after **312.9 s**; passes 2 and 3 pressed nothing at all
|
||||
(`outcomes={'glance': ('held', 312.9)} never_again=['glance']`).
|
||||
|
||||
| the question | the answer, measured |
|
||||
|---|---|
|
||||
| does an unattended update ever produce a HOLD? | **yes** — 312.9 s, the full health wait plus the phases |
|
||||
| does the box try again? | **no** — two further passes pressed nothing |
|
||||
| is the household told? | **on the screen, yes** — the app page, and a banner on EVERY authenticated page carrying every held app at once |
|
||||
| told what? | what happened, when, **which copy** and **what that copy holds** — all four scored True |
|
||||
| by MAIL? | **still unmeasured** — the scratch guest runs `hub.enabled: false` and the notifier returns before it logs (**R-620**) |
|
||||
| in ENGLISH? | **no** — the sentence is Hungarian on the English page (**R-606**, confirmed on the hold sentence itself) |
|
||||
|
||||
**So the mechanism this question's recommended option rests on is already there and already behaves
|
||||
that way.** What remains in Q4 is the MAIL and the ENGLISH, not the hold.
|
||||
|
||||
**Two further facts this measurement produced, neither of which the question anticipated.**
|
||||
**(1) There is NO single-flight** — five Updates pressed within 0.45 s all ran at once and all ended
|
||||
honest, so a caller pressing N apps runs N updates simultaneously. **(2) A held app keeps inviting
|
||||
the household to update it and the button then refuses** (`409 reason='held'`), even after the
|
||||
catalog publishes a FIXED newer version — the household's only route out is the restore. Correct per
|
||||
§6.1, and the page says otherwise (**R-625**).
|
||||
|
||||
**Also proven across a genuine power cut:** the boot sweep met a held app after an unclean shutdown
|
||||
and deliberately left it alone — *„whatever is holding it owns its recovery"*.
|
||||
|
||||
### Q5 — PostgreSQL: what has to exist before the catalog may move `postgres:16` to `17`?
|
||||
|
||||
*Eleven templates, and the image performs no conversion — it refuses to start on an older major's
|
||||
@@ -335,6 +390,35 @@ rule's MariaDB half is LIFTED (R-469 — MariaDB has both a backup in front of i
|
||||
`MARIADB_AUTO_UPGRADE=1`); this half is exactly what stays. **If nothing is decided:** nothing breaks
|
||||
— the gate refuses the move — but the eleven apps drift further from upstream every month.
|
||||
|
||||
|
||||
**MEASURED 2026-09-21, both halves, on a real seeded datadir.**
|
||||
|
||||
**(a) What a household would see today — as predicted, and now observed.** The guarded Update of
|
||||
`postgres:16-alpine -> 17-alpine` ended **`failed` in 5.1 s**; the app was stopped and held; **the pin
|
||||
named 17 while `installed_images` still said 16 and nothing was running**; the data was intact; and
|
||||
the restore the hold sentence names brought it back in **29.1 s**. The engine's refusal had to be
|
||||
REPRODUCED independently, because `failAndHold` destroyed it before any probe could read it
|
||||
(**R-621**) — *FATAL: database files are incompatible with server / DETAIL: The data directory was
|
||||
initialized by PostgreSQL version 16, which is not compatible with this version 17.11.* The datadir
|
||||
was still `16` afterwards; the positive control (the same copy under 16) started and held 48 tables.
|
||||
|
||||
**(b) The conversion rehearsal, COSTED.** Logical dump and restore, 49 MB / 48 tables:
|
||||
`pg_dumpall` **2.6 s / 132 201 B**; fresh 17 datadir plus replay **6.5 s / 48 tables restored**; the
|
||||
app up on 17 saying *Database connection successful*; **the seeded account read back**; **total
|
||||
155.9 s, of which ~9 s is engine work.** For eleven apps that is a maintenance window, not a project.
|
||||
`pg_upgrade` was NOT run — it needs both majors' binaries in one image and no such image exists in
|
||||
this project; the logical route may make it unnecessary at this size. Full paragraph:
|
||||
`audits/update-night-2026-09-21/24-Q5-postgres-conversion-costed.md`.
|
||||
|
||||
**(c) A fact about the INSTRUMENT, not the engine.** `upgrade-test.py`'s PostgreSQL probe is
|
||||
`cat /var/lib/postgresql/data/PG_VERSION` **inside the container**. Against the converted datadir it
|
||||
answered `17`, exit 0 — it works. **But it is blind in exactly the case that matters**: when
|
||||
PostgreSQL refuses, the container is not running, so `docker exec` cannot ask it anything. Tonight it
|
||||
recorded `No such container`, which its own honesty rule covers — but it must never be read as *the
|
||||
engine is content*.
|
||||
|
||||
**The recommendation is unchanged.** Tonight gives it a price rather than a new opinion.
|
||||
|
||||
### Q6 — Should the catalog record each pin's DIGEST at push time?
|
||||
|
||||
*So the box can tell a moved floating tag from an unmoved one without ever reaching a registry.*
|
||||
@@ -355,6 +439,23 @@ that has demonstrably moved.
|
||||
**If nothing is decided:** „Naprakész" keeps meaning "the reference matches", which is measurably not
|
||||
what it sounds like.
|
||||
|
||||
|
||||
**MEASURED ON A BOX 2026-09-21 (update night, leg B8), and it refines the picture in two ways.**
|
||||
§8.1's numbers came from a registry sweep on DooPlex; this is the same question asked of a
|
||||
customer-shaped box, where the badge actually renders. On guest 9202, `docmost`'s two floating pins
|
||||
were read as `installed_images` records them and compared with the upstream digests measured the same
|
||||
night: `postgres:16-alpine` -> `sha256:721873c34ceb9…` **on both sides**; `redis:7-alpine` ->
|
||||
`sha256:858f009f9709c…` **on both sides**. **Identical — so „Naprakesz" is TRUE for this box.**
|
||||
|
||||
**(1) The defect's size is set by INSTALL AGE, not by the catalog.** A floating pin is wrong only for
|
||||
a box that pulled BEFORE the tag moved. R-446's six repushed pins measure the tag against the date
|
||||
the CATALOG set it, which is the right measure for the catalog and not for a box.
|
||||
|
||||
**(2) The producer this question needs ALREADY EXISTS on the box.** `installed_images` records a real
|
||||
`digest` per service — the box knows exactly what it is running. What it cannot do is COMPARE,
|
||||
because the catalog carries no digest. That is precisely this question's proposal, and only the
|
||||
catalog half is missing. **The recommendation is unchanged.**
|
||||
|
||||
### Q7 — What does the hub's report need to carry for a fleet view?
|
||||
|
||||
*Slice 7 lets the operator SEE and MOVE how far behind every box is.*
|
||||
@@ -767,6 +868,22 @@ headlessly (R-460).
|
||||
| E | ~~the automatic night~~ **MOSTLY DONE 2026-09-21 (R-611)** — the success night and the no-retry proof both measured. **What remains: the unattended HOLD**, which needs an edge that passes the within-a-major test and still fails health (see Q4) | ~1 CC-hour + a purpose-built image |
|
||||
| F | the remaining **38 apps**, through the nightly rotation as decision 6 directs | ~1 app/night; fixtures amortised |
|
||||
|
||||
**RE-COSTED 2026-09-21 FROM THE NIGHT'S REAL NUMBERS** (`audits/DRILL-update-night-2026-09-21.md`):
|
||||
|
||||
| leg | status after the update night |
|
||||
|---|---|
|
||||
| A — the 15 database services | **LARGELY DONE.** 21 edges across 19 apps walked box-side in one night, including both engines and 8 database-carrying apps. **Machine time was never the cost and is now known: a proven edge took 11-218 s, median ~45 s.** The cost was fixtures, exactly as costed — and the real surprise is that two apps can NEVER be seeded headlessly while the catalog rightly closes their sign-up (R-624) |
|
||||
| B — the power cut | **COMPLETE.** The two EARLY phases nobody had cut in were cut tonight: `backing-up` (the box recovered and said so) and `safety-dump` (nothing moved, nothing to say). Only a cut inside `starting` itself remains, and it still needs an in-process fault injector |
|
||||
| C — the PostgreSQL rehearsal | **DONE and COSTED**: ~9 s of engine work, 155.9 s end to end for 49 MB / 48 tables. `pg_upgrade` still owed and may prove unnecessary |
|
||||
| D — the downgrade refusal | already done, v0.260.0 |
|
||||
| E — the automatic night | **COMPLETE.** The unattended HOLD was produced at last (312.9 s), with no retry across two further passes. It needed the image store of §6.5 |
|
||||
| F — the remaining apps | ~34 still unwalked. The fixtures for 20 exist and amortise |
|
||||
|
||||
**What the night ADDED to this table, which none of the legs anticipated:** the `verifying` phase
|
||||
trusts the `.felhom.yml` probe absolutely, and three of 53 templates name a probe the app does not
|
||||
answer — so a SUCCESSFUL update of those apps ends by STOPPING a working app (**R-618**, P1). That is
|
||||
now the first thing Slice 6 has to be safe against, ahead of everything in this table.
|
||||
|
||||
**Total for legs A–E: roughly 21–34 CC-hours**, plus ~25–30 GB of images on a scratch host. Legs C
|
||||
and E are the ones that unblock a decision; leg A is the one that takes the time.
|
||||
|
||||
@@ -774,6 +891,64 @@ and E are the ones that unblock a decision; leg A is the one that takes the time
|
||||
harness on DooPlex for the image-side ones.
|
||||
|
||||
|
||||
## 6.5 The drill catalog and the image store — the standing method for update drills
|
||||
|
||||
**Why this section exists.** On 2026-09-21 an afternoon session put a deliberately broken image into
|
||||
the LIVE catalog for thirteen minutes to produce a failing edge. It was reverted and nothing reached
|
||||
a customer, but the method was wrong and the brief that asked for it said so. This is the method that
|
||||
replaces it, proven the same night.
|
||||
|
||||
**The rule, and it has no exception:** *nothing broken, dummy, cross-repo or engine-major ever enters
|
||||
the live catalog — not as a fallback, not for thirteen minutes. If a leg cannot be done without that,
|
||||
the leg is skipped and named.*
|
||||
|
||||
### The two mechanisms
|
||||
|
||||
| | what it is | what it makes possible |
|
||||
|---|---|---|
|
||||
| **the drill catalog** | `admin/app-catalog-drill` on Gitea — private, a copy of the live catalog's `main` | a scratch box can be pointed at a catalog where a failing edge is *committable*, because it carries none of the live repo's gates |
|
||||
| **the image store** | a `registry:2` container on the scratch guest at `127.0.0.1:5000` | an edge that **passes the within-a-major test and still fails** — the one shape a real catalog move cannot produce |
|
||||
|
||||
**The image store is not a convenience.** `09` §3b Q4 could not be measured for a year of drills
|
||||
because the only failing edges available were across-a-major, and the within-a-major rule — correctly
|
||||
— refuses those before the guarded update is ever reached. *The rule that makes automatic updates
|
||||
safe is the same rule that refuses the obvious way to break one.* Measuring an unattended HOLD needs
|
||||
`drill/<app>:X.Y.Z` (the real image, retagged) against `drill/<app>:X.Y.(Z+1)` (a built image that
|
||||
starts, stays up and never serves) — same repository, same major, plain version tags. A third
|
||||
flavour, a tag simply **absent** from the store, gives the pull-failure leg.
|
||||
|
||||
`stacks.CompareImageRefs` orders a `host:port/` reference correctly: `splitImageRef` takes the last
|
||||
colon and rejects it only when a `/` follows, so a registry port is never read as a tag. **Proven by
|
||||
running it**, four positive cases and a negative control, 2026-09-21.
|
||||
|
||||
### Pointing a box at the drill catalog — the step that is NOT obvious
|
||||
|
||||
**`git.repo_url` alone is inert.** `Syncer.gitCloneOrPull` clones only when the cache has no `.git`;
|
||||
otherwise it fetches from the remote the clone already stores. The cache directory must be removed as
|
||||
well, or the box goes on following the live catalog and reports success. Filed as **R-615**; until it
|
||||
is fixed, the drill procedure is:
|
||||
|
||||
1. save `controller.yaml` as `controller.yaml.pre-update-night`;
|
||||
2. set `git.repo_url` (and `username`/`token` — the drill repo is private);
|
||||
3. **remove `<data>/catalog-cache`**;
|
||||
4. restart the controller, sync, **rescan** (R-607: a sync can answer „nincs változás" while the
|
||||
catalog has moved, and the badge answers from the stale value until the rescan);
|
||||
5. **three controls, all quoted in the report** — the drill bump appears on the scratch box; the
|
||||
other boxes' caches are unchanged; the live catalog's `main` hash is unchanged.
|
||||
|
||||
### What the drill must leave behind
|
||||
|
||||
- `controller.yaml` restored from the saved copy, the controller restarted, and `git.repo_url` **read
|
||||
back and quoted** as the live catalog.
|
||||
- The registry container and its volume removed; drill images removed **by name**. Never `prune`.
|
||||
- The drill repo **kept**, private, reset to the live catalog's `main`, so the next drill starts clean.
|
||||
- A diff of every `image:` line against the live catalog's `main` — expected: identical.
|
||||
|
||||
### The fence
|
||||
|
||||
Only a scratch guest is ever pointed at the drill catalog. The drill repo's README says so, and no
|
||||
customer box has credentials for it. The store listens on the guest's loopback only.
|
||||
|
||||
## 7. What slices 1 and 2 actually built
|
||||
|
||||
### 7.1 The record (slice 1)
|
||||
@@ -887,8 +1062,20 @@ Version strings stay in the logs, the API and the hub.
|
||||
PostgreSQL half (R-463) has no equivalent — the image performs no `pg_upgrade` — and the
|
||||
engine-major rule (§3 precaution 3, R-469) is what keeps both engines inside their major until
|
||||
Slice 4 gives the Update button a backup.
|
||||
8. **Only three of 53 apps have ever had an upgrade measured**, and one of them (bookstack) can only
|
||||
be half-proven headlessly (**R-460**). The widening is **R-462**, costed with real numbers.
|
||||
8. ~~**Only three of 53 apps have ever had an upgrade measured.**~~ **WIDENED 2026-09-21 to 21
|
||||
EDGES ACROSS 19 APPS** (`audits/DRILL-update-night-2026-09-21.md`), on scratch guest 9202
|
||||
through the product's own guarded Update, against a PRIVATE DRILL CATALOG so the live catalog
|
||||
carried no test reference at any point: **14 proven, 3 failed, 4 inconclusive**, each app
|
||||
seeded and read back through its OWN front door with a negative control on every readback.
|
||||
Ten of the fourteen printed a verbatim migration line. **What stays true:** bookstack is still
|
||||
only half-provable headlessly (**R-460**), and **two apps cannot be seeded AT ALL** while the
|
||||
catalog rightly closes their sign-up — vaultwarden (`SIGNUPS_ALLOWED=false`, R-512) and
|
||||
zipline — which is a permanent ceiling on R-462's scope rather than a fixture nobody has
|
||||
written (**R-624**). **And one thing this widening FOUND that no count would have:** the
|
||||
`verifying` phase trusts the `.felhom.yml` probe absolutely, and **three of the 53 templates
|
||||
name a probe the app does not answer**, so a SUCCESSFUL update of those apps ends by STOPPING
|
||||
a working app (**R-618**, P1 — tandoor measured serving HTTP 200 on the new version at four
|
||||
samples across five minutes, then stopped).
|
||||
9. **The hub does not record image tags at all.** Its report's container payload carries name, state,
|
||||
CPU and memory, and no image field (spike §5). So the fleet view of §6 slice 7 needs a hub-side
|
||||
change; it is not derivable from what is already reported.
|
||||
|
||||
Reference in New Issue
Block a user