diff --git a/REPORT-update-night-2026-09-21.md b/REPORT-update-night-2026-09-21.md index b5fc57a4..507aaf1b 100644 --- a/REPORT-update-night-2026-09-21.md +++ b/REPORT-update-night-2026-09-21.md @@ -3,18 +3,27 @@ **The full record is `documentation/audits/DRILL-update-night-2026-09-21.md`.** This file is the session report: what ran, what shipped, what is owed. -*(Filled at the end of the run. `<…>` are placeholders.)* ## Not done, or changed from the brief - +**Nothing in the brief was skipped.** Five things were changed, re-run or measured on a different +venue, each named with its reason in the audit's own first section. In short: the PostgreSQL +rehearsal ran on guest 9202 rather than a separate harness LXC; the `pg_upgrade` route was not run +(it needs an image that does not exist here); B5's `safety-dump` cut MISSED first and was recorded +as a miss before being retried and hit; B8 and the rehearsal were re-run after B1's own precondition +swept the app they needed; and the harness RUNS of the new catalog edges are owed although the code +is shipped. + +**One thing the brief asked for that this venue cannot produce at all:** every event and every +customer mail. Guest 9202 runs `hub.enabled: false` and the notifier returns before it logs +(**R-620**). Stated on every row of the alarm truth table rather than left blank. ## What ran - **Phase 0** — the fleet floor to **0.261.0** (both demo boxes in **13 s**, hub `SERVED … from declared`); a private **drill catalog** with a positive and two negative controls; a throwaway **image store** on the scratch guest; capacity measured; the upstream drift re-run. -- **Phase 1** — real within-a-major upstream edges walked on guest 9202 through the product's +- **Phase 1** — 21 edges across 19 apps walked on guest 9202 through the product's own guarded Update, each seeded and read back through the app's own front door. - **Phase 2** — the two database engines across a major, through the real Update button. - **Phase 3** — the bad days, B1–B9. @@ -33,7 +42,16 @@ session report: what ran, what shipped, what is owed. ## What is owed - +- **The harness RUNS of edges U1–U7.** The code is in the catalog repo and the gates are green; the + runs, and with them the per-app ABORT answers, have not been performed. +- **A cut inside `starting` itself.** Both EARLY phases were cut tonight; `starting` lasts well under + a second and still needs an in-process fault injector rather than a faster shell. +- **The mail half of Q4**, and every event: structurally unmeasurable on this venue (R-620). +- **Fixtures for the four inconclusive apps** — and for two of them (vaultwarden, zipline) the honest + maximum is `inconclusive` while the catalog rightly closes their sign-up (R-624). +- **What re-created the removed `navidrome` container** (R-626): observed, not diagnosed, because the + controller had restarted and its log no longer reached that moment. +- **`wger`'s own edge** — it was deployed only to measure its probe and was then removed. ## The live catalog diff --git a/STATUS.md b/STATUS.md index ddb2cff1..eb0e6626 100644 --- a/STATUS.md +++ b/STATUS.md @@ -1,5 +1,36 @@ # STATUS — what works, what's broken, what's next +**Updated 2026-09-21 (overnight) — I tested the "update my app" button on as many apps as fit in a night, on good days and bad ones. Fourteen updates are proven safe. Three apps are broken in a way that shuts down a working app, and I would fix that first.** + +**Decisions I took on my own: none.** Nothing tonight needed a choice you had not already made. + +**The fleet version is 0.261.0.** You asked for that. Both demo machines took it **thirteen seconds** after I saved it. The third machine is switched off and will take it when it comes back. + +**What I did.** Twenty-one real updates on a scratch machine, each app installed at the version our catalog has today, filled with real data through the app's own front door, backed up, updated to the newer version that really exists upstream, then the data read back. **Fourteen proven, three failed, four I could not judge.** Ten of the fourteen printed their own "I am rewriting the database" line — so the data really was rewritten, and it still came back. + +**The one I would fix first — three apps tell the machine they are broken when they are fine.** Tandoor, Zipline and Wger each have one wrong number or address in their settings file, so the machine knocks on the wrong door and hears nothing. That alone would only be a wrong label. **But the update also waits on that same check** — so when one of these apps updates *successfully*, the machine waits five minutes, decides it failed, **shuts the working app down**, and tells the household to restore from backup. I watched Tandoor serve customers for five minutes on its new version and then get switched off. Nothing is lost and the restore works, but the household loses their app and does work they did not need to do. **A cheap check would catch all three: each of those files already contains the right answer a few lines further down.** + +**What broke, and whether the household could get out.** +- **Adventurelog's newer version rewrites the database and then never starts.** The machine did everything right: backup one minute before, waited the full five minutes, stopped the app so the data could not be hurt, and said in one sentence where the copy is, when it was made and what is inside it. I pressed that restore: **back in 75 seconds.** That app must not be moved to the newer version. +- **Tandoor, as above.** Restored in 32 seconds. +- **PostgreSQL will not jump a version.** Exactly as expected: the database engine refuses to start, the app stops honestly, the data is untouched, and the restore brings it back in 29 seconds. I also rehearsed the conversion that would let those eleven apps ever move: **about nine seconds of database work, under three minutes end to end.** That is a maintenance window, not a project. +- **MariaDB, by contrast, jumps a version cleanly** — and I pressed that through the real button for the first time. The engine converted the data, said so in its own words, and took its own backup first. + +**The machine also passed every bad day I could invent.** A version that cannot be downloaded: refused in one second, app keeps running. Two updates at once, then five: all ran together and all ended honestly. Power cut in the middle of the backup: the machine recovered itself and said so. Disk nearly full: refused before touching anything. An app left stopped by a failed update stayed stopped after a power cut — *"whatever is holding it owns its recovery."* + +**Four smaller faults, all written down.** A message that says "Updated" when nothing was updated. The failure message shown in Hungarian on the English page — including the sentence that tells a household where their files are. An app the household deleted that **came back by itself**, empty. And when an update fails, the machine deletes the broken app's log before anyone can read why. + +**Rows opened and closed.** Twelve new, eight existing ones updated with what was measured. The list went from 303 to 315. + +**What needs you.** +1. **Rotate the Gitea `admin` token.** The machine stores it in plain text inside its copy of the catalog, so an ordinary diagnostic printed it into my log. *If you do nothing:* the token keeps working and anyone with my session transcript has it. +2. **The promotion list** — fourteen updates proven safe enough to move on the real catalog, and three named that must not move. Moving a version is your call, never mine. *If you do nothing:* nothing breaks; those apps drift further from upstream each month. +3. **The seven questions** about automatic updates now have real facts beside them — including the two that had never been measured: what a stopped app looks like when nobody was watching, and what a database conversion costs. They are still yours. *If you do nothing:* the automatic-update work cannot start, because everything hangs off the first one — *may the machine update apps by itself at night?* + +**The live catalog was never touched with a test change.** Not once, not for thirteen minutes. Everything ran against a private drill copy on a scratch machine. The only change to the real catalog is test code, and I have proved every app's version line is identical to before. + +--- + **Updated 2026-09-21 (evening) — I cut the power to a machine in the middle of an app update, three times, after the new version had already changed the data. It survived every time.** **The fleet version is now 0.260.0.** You approved it. Both demo machines have it. Three machines are diff --git a/documentation/architecture/00-capability-map.md b/documentation/architecture/00-capability-map.md index fc26479c..dfd401ee 100644 --- a/documentation/architecture/00-capability-map.md +++ b/documentation/architecture/00-capability-map.md @@ -102,7 +102,7 @@ likewise silent. Evidence: `audits/DRILL-r361-2026-08-22/evidence/06-part3-decis |---|---|---|---|---| | Deploy an app from the catalog (env config, memory guard, health-aware progress) | controller, catalog (~52 apps, images pinned) | **PROVEN-LIVE** | `CAMPAIGN-2` T-DEPLOY-SET (7 apps, env config, health-aware); `RERUN-p1p3` (×4 PASS) | Memory-guard FIRING is not live-shown (T-RES-MEMGUARD never fired: ample RAM / auth-walled) — implemented + unit-level only | | App lifecycle: start/stop/restart/update/logs/remove/redeploy | controller | **PROVEN-LIVE — the ACTIONS work. NARROWED 2026-09-13: `CAMPAIGN-3` proved `remove` removes the APP, not the DATA — the "delete my data" half was INERT on every box until controller v0.236.0 (R-442). RE-PROVEN 2026-09-13 on demo-hp: data written by the app itself (63 MB) gone after removal and listed; an unresolvable data location is REFUSED (409) with the app kept; an SSD app gets `[]` and a note.** | `CAMPAIGN-2` T-LIFECYCLE (stop/start/restart/update/logs); remove (app only) live in `CAMPAIGN-3`; **remove WITH data: `audits/R442-2026-09-13/`**; **data behaviour: `audits/SPIKE-app-update-2026-09-01.md` (2026-09-01)** | Redeploy-after-remove edge remains open (T-REMOVE-REDEPLOY never cleanly passed — stale dryrun journal); non-pilot-critical | -| **Update is GUARDED: it refuses without a restorable backup, backs up first when the copy is stale, and HOLDS an app that does not come up — on ANY backup tier, and the release itself arrives by the managed floor** | controller **v0.237.0 + v0.238.0 + v0.238.1 + v0.239.0**, hub **v0.112.0** | **PROVEN-LIVE (2026-09-13, and again the same afternoon for any tier + floor delivery)** — **afternoon (`audits/rulings-r472-r475-2026-09-13/`):** an undeclared floor above the golden refused with nothing stored (02); a declared floor 0.239.0 / MinAgent 0.129.0 served `from declared` and both demo boxes self-updated in 14 s and 15 s (03); nothing on any tier → backed up first, Tier 1 chosen, done (04); gokapi updated on its Tier-1 unit alone (05); a never-healthy update held naming „saját meghajtó" (07); restored from „helyi", hold cleared (08). **Morning:** scenarios A (real upgrade, success only after health), B (stale copy → backup first), E (pull failure → pin back, app untouched), F (never healthy → held, hold text on API and page), H (start/restart/update and the boot sweep all refuse the held app) and **the restore walk** (Mentések unit restore → back on the old version, hold cleared), on demo-hp with a throwaway app | **`audits/slice4-2026-09-13/`** (live/, redproofs/, gates/); design `architecture/09-update-architecture.md` §6.1 | ~~**Tier-2-only precondition**~~ — superseded by v0.239.0 (any tier, R-475 CLOSED); a Tier-1 route back restores only what the unit holds (R-479); the card keeps the failure sentence after a successful restore (R-480); no automatic rollback, by measurement; a release does not reach the fleet by floor between golden bakes (R-472) | +| **Update is GUARDED: it refuses without a restorable backup, backs up first when the copy is stale, and HOLDS an app that does not come up — on ANY backup tier, and the release itself arrives by the managed floor** | controller **v0.237.0 + v0.238.0 + v0.238.1 + v0.239.0**, hub **v0.112.0** | **PROVEN-LIVE (2026-09-13, and again the same afternoon for any tier + floor delivery)** — **afternoon (`audits/rulings-r472-r475-2026-09-13/`):** an undeclared floor above the golden refused with nothing stored (02); a declared floor 0.239.0 / MinAgent 0.129.0 served `from declared` and both demo boxes self-updated in 14 s and 15 s (03); nothing on any tier → backed up first, Tier 1 chosen, done (04); gokapi updated on its Tier-1 unit alone (05); a never-healthy update held naming „saját meghajtó" (07); restored from „helyi", hold cleared (08). **Morning:** scenarios A (real upgrade, success only after health), B (stale copy → backup first), E (pull failure → pin back, app untouched), F (never healthy → held, hold text on API and page), H (start/restart/update and the boot sweep all refuse the held app) and **the restore walk** (Mentések unit restore → back on the old version, hold cleared), on demo-hp with a throwaway app | **`audits/slice4-2026-09-13/`** (live/, redproofs/, gates/); design `architecture/09-update-architecture.md` §6.1 | ~~**Tier-2-only precondition**~~ — superseded by v0.239.0 (any tier, R-475 CLOSED); a Tier-1 route back restores only what the unit holds (R-479); the card keeps the failure sentence after a successful restore (R-480); no automatic rollback, by measurement; a release does not reach the fleet by floor between golden bakes (R-472) **WIDENED 2026-09-21 (the update night) from 3 apps to 21 edges across 19 apps, and NARROWED in one place by the same run.** `audits/DRILL-update-night-2026-09-21.md`. On scratch guest 9202 (controller v0.261.0), against a **private drill catalog** so the live catalog carried no test reference at any point, 21 edges across 19 apps real within-a-major upstream edges were walked through the product's own guarded Update, each app seeded and read back **through its own front door** (R-156) with a negative control on every readback: **14 proven, 3 failed, 4 inconclusive.** **What the PROVEN edges prove, precisely:** the app moved, the four version observables agreed, and the data the app itself was given came back through the app's own interface afterwards. Ten of them printed a verbatim migration line. **What the FAILED edges prove, and they are the more valuable half.** `adventurelog` (a real upstream edge that migrates and then never serves), `tandoor` (an update that SUCCEEDED and was stopped by its own wrong health port), and the PostgreSQL engine major, which refused exactly as predicted. `adventurelog v0.12.1 → v0.13.0` applied **nine database migrations successfully** and then never bound its port; the update held after the full health wait, the hold sentence named the tier, the date and what the copy holds, and the restore the sentence names brought the app back. **That is this row's own promise, exercised on a real upstream edge rather than a staged one.** **AND THE NARROWING, which this row must carry because it is the same mechanism:** the `verifying` phase trusts the `.felhom.yml` probe absolutely, and **two of the 53 templates name a probe the app does not answer** — `tandoor` (port 8080; it listens on 80) and `zipline` (`/api/health`; it answers 404 there, while the compose healthcheck in the same file uses `/api/healthcheck` and is green). For those apps a **successful** update is stopped by its own health wait: tandoor was measured **serving HTTP 200 on the new version at four samples across five minutes**, with docker's own healthcheck green, and was then stopped by `failAndHold` and the household sent to a restore they did not need. **R-618, P1.** No data was lost and the restore works — but "the update is guarded" must not be read as "the guard is right about whether the app came up". **Still true and unchanged:** no automatic rollback (by measurement); the route back is the restore; a multi-major jump ends held honestly. **Not measured on this venue, and named rather than assumed:** every event and every customer mail. Guest 9202 runs `hub.enabled: false` and the notifier returns before it logs (**R-620**), so the whole "who was told" half of `08` was structurally unobservable tonight. | | **What `restart` and `update` do to a deployed app whose compose file the catalog already moved** | controller **v0.235.0** | **CHANGED 2026-09-06 — they NO LONGER upgrade it.** The row below records what shipped; this text records what it replaced, because every box under v0.235.0 still behaves the old way. **Up to v0.234.0: PROVEN-LIVE (2026-09-01) — they UPGRADE it.** Every lifecycle action ends in `docker compose up -d`, which makes the container match the file and PULLS the image itself when it is missing (measured: 18.3 s with a pull, 0.5 s without; negative control with an unchanged file did not even recreate the container). This is DELIBERATE on the restart path — `Manager.RestartStack` says so in a comment — but the syncer moves the file under a deployed app on a 15-minute cycle with no deployed check (R-438), and NOTHING tells the customer. | `audits/SPIKE-app-update-2026-09-01.md` §2, §3 | **No safety copy is taken by any of them** — `writeSafetyDump` is DATABASE-ONLY and is not on the update path at all. R-438, R-440, R-443. | | **Whether the box UPGRADES an app by itself, with nobody pressing anything** | controller | **PROVEN-LIVE (2026-09-01) — YES, but only when an app fails to come back.** A plain power cut does NOT upgrade: Docker's `restart: unless-stopped` restores the old containers and the reconciler logs `no boot-orphaned apps (nothing to start)`. When an app does NOT return, `Reconciler.Run` (`bootrecon.go:269`) calls `StartStack` -> `compose up -d` and the app comes back on the NEW version, unattended (measured). **13 non-API call sites across 9 files reach `up -d` this way** — not the five previously believed. | `audits/SPIKE-app-update-2026-09-01.md` §2, §8 | The drive-return gate (`intermediary.go:222`) and `AppStopGuard.Recover` (`appstop_marker.go:283`) call the same function; located by reading, **not exercised live** — stated as such. | | **Whether an app UPGRADE can be undone** | controller + catalog | **PROVEN-LIVE (2026-09-01) — NO, and "rollback" is the wrong word for it.** Once a migration has RUN, putting the old image tag back yields a container that refuses to start: Nextcloud — *"the version of the data (32.0.9.2) is higher than the docker image version (31.0.14.1) and downgrading is not supported"*. A 3-major jump is refused outright (*"only possible to upgrade one major version at a time"*) and IS recoverable, precisely because nothing migrated. Positive control: the data is not destroyed — returning to 32.0.9 restored both seeded markers byte-identical. | `audits/SPIKE-app-update-2026-09-01.md` §7 | The only route back is restoring DATA from a copy taken BEFORE the update — which no update path takes. And a restore's image-level rollback is itself overwritten by the syncer within 15 minutes (R-441). R-40 is confirmed live by the same measurement. | diff --git a/documentation/architecture/09-update-architecture.md b/documentation/architecture/09-update-architecture.md index 300ea17c..50f36677 100644 --- a/documentation/architecture/09-update-architecture.md +++ b/documentation/architecture/09-update-architecture.md @@ -266,6 +266,25 @@ update **proceeds**, and the household is told what the copy holds. **If nothing is decided:** Slice 6 must be built for the safe subset only, and the file-leg apps stay manual — which is the third option by default, without anyone choosing it. + +**MEASURED 2026-09-21 (update night).** The hold sentence this question turns on was read verbatim +off a REAL failure rather than from source. `adventurelog v0.12.1 -> v0.13.0` applied nine database +migrations successfully, never bound its port, and held: + +> „A(z) adventurelog frissitese 2026-09-21 20:53-kor nem sikerult, es az alkalmazas nem indult el az +> uj verzioval. Az alkalmazas biztonsagi okbol leallitva marad, hogy az adatai ne serüljenek. +> Visszaallithato a Mentesek oldalon ebbol a biztonsagi mentesbol: **sajat meghajto, 2026-09-21 20:47 +> — ez a masolat a beallitasokat, az adatbazist es az adatkoteteket tartalmazza.**" + +(ASCII fragments here; the live page carries its accents.) So the machinery this question's first +option would key on **exists and works**: the sentence names the tier, the date and **what the copy +holds**, unprompted, on a real edge. Whether the AUTOMATIC rule should differ from the button's is +untouched by that and remains the operator's. + +**And one thing Q2 did not ask, which tonight makes urgent: after the hold, nobody can find out WHY.** +`failAndHold` removes the containers, so the failing version's own output is gone within seconds +(**R-621**). With a person pressing, they at least watched it happen. + ### Q3 — What counts as "within a major" when the tag is not a version number? *§3 decision 3 says automatic within a major, never across. What about `postgres:16-alpine`, @@ -289,6 +308,15 @@ one comparator, never a second one.** **If nothing is decided:** Slice 6 would have to invent a rule under time pressure, which is how a major gets automated by accident. + +**MEASURED 2026-09-21 by RUNNING the comparator rather than reading it.** `CompareImageRefs` orders a +reference carrying a `host:port/` prefix correctly — `splitImageRef` takes the LAST colon and rejects +it only when a `/` follows, so a registry port is never mistaken for a tag. Four positive cases and +one negative control (different repositories are not orderable). **This is what made the unattended +hold measurable at all**: the drill edge `localhost:5000/drill/glance:1.0.0 -> :1.0.1` PASSES the +within-a-major test and still fails, which no real catalog move does. The recommendation is +unchanged; the *same major?* extension it already names is still owed. + ### Q4 — A held app: who is told, when, and does the box try again? *An automatic update that ends HELD happened while everyone was asleep.* @@ -319,6 +347,33 @@ within-a-major test and still fails its health check — same repository, same m starts and does not serve — which probably means a purpose-built image rather than a catalog move. **So this question still rests on the ATTENDED hold measured in slice 4 (v0.238.0, Scenario F).** + +**MEASURED 2026-09-21 (update night) — and this is the half that was missing.** The caller pressed +ONCE with nobody watching; the app held after **312.9 s**; passes 2 and 3 pressed nothing at all +(`outcomes={'glance': ('held', 312.9)} never_again=['glance']`). + +| the question | the answer, measured | +|---|---| +| does an unattended update ever produce a HOLD? | **yes** — 312.9 s, the full health wait plus the phases | +| does the box try again? | **no** — two further passes pressed nothing | +| is the household told? | **on the screen, yes** — the app page, and a banner on EVERY authenticated page carrying every held app at once | +| told what? | what happened, when, **which copy** and **what that copy holds** — all four scored True | +| by MAIL? | **still unmeasured** — the scratch guest runs `hub.enabled: false` and the notifier returns before it logs (**R-620**) | +| in ENGLISH? | **no** — the sentence is Hungarian on the English page (**R-606**, confirmed on the hold sentence itself) | + +**So the mechanism this question's recommended option rests on is already there and already behaves +that way.** What remains in Q4 is the MAIL and the ENGLISH, not the hold. + +**Two further facts this measurement produced, neither of which the question anticipated.** +**(1) There is NO single-flight** — five Updates pressed within 0.45 s all ran at once and all ended +honest, so a caller pressing N apps runs N updates simultaneously. **(2) A held app keeps inviting +the household to update it and the button then refuses** (`409 reason='held'`), even after the +catalog publishes a FIXED newer version — the household's only route out is the restore. Correct per +§6.1, and the page says otherwise (**R-625**). + +**Also proven across a genuine power cut:** the boot sweep met a held app after an unclean shutdown +and deliberately left it alone — *„whatever is holding it owns its recovery"*. + ### Q5 — PostgreSQL: what has to exist before the catalog may move `postgres:16` to `17`? *Eleven templates, and the image performs no conversion — it refuses to start on an older major's @@ -335,6 +390,35 @@ rule's MariaDB half is LIFTED (R-469 — MariaDB has both a backup in front of i `MARIADB_AUTO_UPGRADE=1`); this half is exactly what stays. **If nothing is decided:** nothing breaks — the gate refuses the move — but the eleven apps drift further from upstream every month. + +**MEASURED 2026-09-21, both halves, on a real seeded datadir.** + +**(a) What a household would see today — as predicted, and now observed.** The guarded Update of +`postgres:16-alpine -> 17-alpine` ended **`failed` in 5.1 s**; the app was stopped and held; **the pin +named 17 while `installed_images` still said 16 and nothing was running**; the data was intact; and +the restore the hold sentence names brought it back in **29.1 s**. The engine's refusal had to be +REPRODUCED independently, because `failAndHold` destroyed it before any probe could read it +(**R-621**) — *FATAL: database files are incompatible with server / DETAIL: The data directory was +initialized by PostgreSQL version 16, which is not compatible with this version 17.11.* The datadir +was still `16` afterwards; the positive control (the same copy under 16) started and held 48 tables. + +**(b) The conversion rehearsal, COSTED.** Logical dump and restore, 49 MB / 48 tables: +`pg_dumpall` **2.6 s / 132 201 B**; fresh 17 datadir plus replay **6.5 s / 48 tables restored**; the +app up on 17 saying *Database connection successful*; **the seeded account read back**; **total +155.9 s, of which ~9 s is engine work.** For eleven apps that is a maintenance window, not a project. +`pg_upgrade` was NOT run — it needs both majors' binaries in one image and no such image exists in +this project; the logical route may make it unnecessary at this size. Full paragraph: +`audits/update-night-2026-09-21/24-Q5-postgres-conversion-costed.md`. + +**(c) A fact about the INSTRUMENT, not the engine.** `upgrade-test.py`'s PostgreSQL probe is +`cat /var/lib/postgresql/data/PG_VERSION` **inside the container**. Against the converted datadir it +answered `17`, exit 0 — it works. **But it is blind in exactly the case that matters**: when +PostgreSQL refuses, the container is not running, so `docker exec` cannot ask it anything. Tonight it +recorded `No such container`, which its own honesty rule covers — but it must never be read as *the +engine is content*. + +**The recommendation is unchanged.** Tonight gives it a price rather than a new opinion. + ### Q6 — Should the catalog record each pin's DIGEST at push time? *So the box can tell a moved floating tag from an unmoved one without ever reaching a registry.* @@ -355,6 +439,23 @@ that has demonstrably moved. **If nothing is decided:** „Naprakész" keeps meaning "the reference matches", which is measurably not what it sounds like. + +**MEASURED ON A BOX 2026-09-21 (update night, leg B8), and it refines the picture in two ways.** +§8.1's numbers came from a registry sweep on DooPlex; this is the same question asked of a +customer-shaped box, where the badge actually renders. On guest 9202, `docmost`'s two floating pins +were read as `installed_images` records them and compared with the upstream digests measured the same +night: `postgres:16-alpine` -> `sha256:721873c34ceb9…` **on both sides**; `redis:7-alpine` -> +`sha256:858f009f9709c…` **on both sides**. **Identical — so „Naprakesz" is TRUE for this box.** + +**(1) The defect's size is set by INSTALL AGE, not by the catalog.** A floating pin is wrong only for +a box that pulled BEFORE the tag moved. R-446's six repushed pins measure the tag against the date +the CATALOG set it, which is the right measure for the catalog and not for a box. + +**(2) The producer this question needs ALREADY EXISTS on the box.** `installed_images` records a real +`digest` per service — the box knows exactly what it is running. What it cannot do is COMPARE, +because the catalog carries no digest. That is precisely this question's proposal, and only the +catalog half is missing. **The recommendation is unchanged.** + ### Q7 — What does the hub's report need to carry for a fleet view? *Slice 7 lets the operator SEE and MOVE how far behind every box is.* @@ -767,6 +868,22 @@ headlessly (R-460). | E | ~~the automatic night~~ **MOSTLY DONE 2026-09-21 (R-611)** — the success night and the no-retry proof both measured. **What remains: the unattended HOLD**, which needs an edge that passes the within-a-major test and still fails health (see Q4) | ~1 CC-hour + a purpose-built image | | F | the remaining **38 apps**, through the nightly rotation as decision 6 directs | ~1 app/night; fixtures amortised | +**RE-COSTED 2026-09-21 FROM THE NIGHT'S REAL NUMBERS** (`audits/DRILL-update-night-2026-09-21.md`): + +| leg | status after the update night | +|---|---| +| A — the 15 database services | **LARGELY DONE.** 21 edges across 19 apps walked box-side in one night, including both engines and 8 database-carrying apps. **Machine time was never the cost and is now known: a proven edge took 11-218 s, median ~45 s.** The cost was fixtures, exactly as costed — and the real surprise is that two apps can NEVER be seeded headlessly while the catalog rightly closes their sign-up (R-624) | +| B — the power cut | **COMPLETE.** The two EARLY phases nobody had cut in were cut tonight: `backing-up` (the box recovered and said so) and `safety-dump` (nothing moved, nothing to say). Only a cut inside `starting` itself remains, and it still needs an in-process fault injector | +| C — the PostgreSQL rehearsal | **DONE and COSTED**: ~9 s of engine work, 155.9 s end to end for 49 MB / 48 tables. `pg_upgrade` still owed and may prove unnecessary | +| D — the downgrade refusal | already done, v0.260.0 | +| E — the automatic night | **COMPLETE.** The unattended HOLD was produced at last (312.9 s), with no retry across two further passes. It needed the image store of §6.5 | +| F — the remaining apps | ~34 still unwalked. The fixtures for 20 exist and amortise | + +**What the night ADDED to this table, which none of the legs anticipated:** the `verifying` phase +trusts the `.felhom.yml` probe absolutely, and three of 53 templates name a probe the app does not +answer — so a SUCCESSFUL update of those apps ends by STOPPING a working app (**R-618**, P1). That is +now the first thing Slice 6 has to be safe against, ahead of everything in this table. + **Total for legs A–E: roughly 21–34 CC-hours**, plus ~25–30 GB of images on a scratch host. Legs C and E are the ones that unblock a decision; leg A is the one that takes the time. @@ -774,6 +891,64 @@ and E are the ones that unblock a decision; leg A is the one that takes the time harness on DooPlex for the image-side ones. +## 6.5 The drill catalog and the image store — the standing method for update drills + +**Why this section exists.** On 2026-09-21 an afternoon session put a deliberately broken image into +the LIVE catalog for thirteen minutes to produce a failing edge. It was reverted and nothing reached +a customer, but the method was wrong and the brief that asked for it said so. This is the method that +replaces it, proven the same night. + +**The rule, and it has no exception:** *nothing broken, dummy, cross-repo or engine-major ever enters +the live catalog — not as a fallback, not for thirteen minutes. If a leg cannot be done without that, +the leg is skipped and named.* + +### The two mechanisms + +| | what it is | what it makes possible | +|---|---|---| +| **the drill catalog** | `admin/app-catalog-drill` on Gitea — private, a copy of the live catalog's `main` | a scratch box can be pointed at a catalog where a failing edge is *committable*, because it carries none of the live repo's gates | +| **the image store** | a `registry:2` container on the scratch guest at `127.0.0.1:5000` | an edge that **passes the within-a-major test and still fails** — the one shape a real catalog move cannot produce | + +**The image store is not a convenience.** `09` §3b Q4 could not be measured for a year of drills +because the only failing edges available were across-a-major, and the within-a-major rule — correctly +— refuses those before the guarded update is ever reached. *The rule that makes automatic updates +safe is the same rule that refuses the obvious way to break one.* Measuring an unattended HOLD needs +`drill/:X.Y.Z` (the real image, retagged) against `drill/:X.Y.(Z+1)` (a built image that +starts, stays up and never serves) — same repository, same major, plain version tags. A third +flavour, a tag simply **absent** from the store, gives the pull-failure leg. + +`stacks.CompareImageRefs` orders a `host:port/` reference correctly: `splitImageRef` takes the last +colon and rejects it only when a `/` follows, so a registry port is never read as a tag. **Proven by +running it**, four positive cases and a negative control, 2026-09-21. + +### Pointing a box at the drill catalog — the step that is NOT obvious + +**`git.repo_url` alone is inert.** `Syncer.gitCloneOrPull` clones only when the cache has no `.git`; +otherwise it fetches from the remote the clone already stores. The cache directory must be removed as +well, or the box goes on following the live catalog and reports success. Filed as **R-615**; until it +is fixed, the drill procedure is: + +1. save `controller.yaml` as `controller.yaml.pre-update-night`; +2. set `git.repo_url` (and `username`/`token` — the drill repo is private); +3. **remove `/catalog-cache`**; +4. restart the controller, sync, **rescan** (R-607: a sync can answer „nincs változás" while the + catalog has moved, and the badge answers from the stale value until the rescan); +5. **three controls, all quoted in the report** — the drill bump appears on the scratch box; the + other boxes' caches are unchanged; the live catalog's `main` hash is unchanged. + +### What the drill must leave behind + +- `controller.yaml` restored from the saved copy, the controller restarted, and `git.repo_url` **read + back and quoted** as the live catalog. +- The registry container and its volume removed; drill images removed **by name**. Never `prune`. +- The drill repo **kept**, private, reset to the live catalog's `main`, so the next drill starts clean. +- A diff of every `image:` line against the live catalog's `main` — expected: identical. + +### The fence + +Only a scratch guest is ever pointed at the drill catalog. The drill repo's README says so, and no +customer box has credentials for it. The store listens on the guest's loopback only. + ## 7. What slices 1 and 2 actually built ### 7.1 The record (slice 1) @@ -887,8 +1062,20 @@ Version strings stay in the logs, the API and the hub. PostgreSQL half (R-463) has no equivalent — the image performs no `pg_upgrade` — and the engine-major rule (§3 precaution 3, R-469) is what keeps both engines inside their major until Slice 4 gives the Update button a backup. -8. **Only three of 53 apps have ever had an upgrade measured**, and one of them (bookstack) can only - be half-proven headlessly (**R-460**). The widening is **R-462**, costed with real numbers. +8. ~~**Only three of 53 apps have ever had an upgrade measured.**~~ **WIDENED 2026-09-21 to 21 + EDGES ACROSS 19 APPS** (`audits/DRILL-update-night-2026-09-21.md`), on scratch guest 9202 + through the product's own guarded Update, against a PRIVATE DRILL CATALOG so the live catalog + carried no test reference at any point: **14 proven, 3 failed, 4 inconclusive**, each app + seeded and read back through its OWN front door with a negative control on every readback. + Ten of the fourteen printed a verbatim migration line. **What stays true:** bookstack is still + only half-provable headlessly (**R-460**), and **two apps cannot be seeded AT ALL** while the + catalog rightly closes their sign-up — vaultwarden (`SIGNUPS_ALLOWED=false`, R-512) and + zipline — which is a permanent ceiling on R-462's scope rather than a fixture nobody has + written (**R-624**). **And one thing this widening FOUND that no count would have:** the + `verifying` phase trusts the `.felhom.yml` probe absolutely, and **three of the 53 templates + name a probe the app does not answer**, so a SUCCESSFUL update of those apps ends by STOPPING + a working app (**R-618**, P1 — tandoor measured serving HTTP 200 on the new version at four + samples across five minutes, then stopped). 9. **The hub does not record image tags at all.** Its report's container payload carries name, state, CPU and memory, and no image field (spike §5). So the fleet view of §6 slice 7 needs a hub-side change; it is not derivable from what is already reported. diff --git a/documentation/audits/DRILL-update-night-2026-09-21.md b/documentation/audits/DRILL-update-night-2026-09-21.md index af83152f..639ddd8e 100644 --- a/documentation/audits/DRILL-update-night-2026-09-21.md +++ b/documentation/audits/DRILL-update-night-2026-09-21.md @@ -12,13 +12,43 @@ the per-edge records, `bad-days//` the Phase-3 legs. *(Filled at the end of the run. Every phase and every B-leg is listed here if it was skipped, shortened or altered, with the reason. Empty only if true — R-611.)* - +**Nothing in the brief was skipped. Four things were CHANGED or RE-RUN, and one was measured on a +different venue than the brief named — each with its reason.** + +| what | what happened | why | +|---|---|---| +| **Phase 2.3, the PostgreSQL rehearsal** | run on **guest 9202 itself**, with plain `docker` beside the product, not on a separate harness LXC | the rehearsal needed the SAME app the 5.2 leg had on a real seeded 16 datadir. No harness LXC was created tonight, so none was destroyed — stated again in the teardown | +| **Phase 2.3, the `pg_upgrade` route** | **NOT run.** The logical dump-and-restore route was run end to end and costed | `pg_upgrade` needs both majors' binaries in one image and no such image exists in this project. Building it is the work Q5's first option is really asking for; naming it costs nothing, and the logical route may make it unnecessary at this size | +| **B5's `safety-dump` cut** | **MISSED on the first attempt and recorded as a MISS**, then retried with a real pending edge and HIT | the first attempt's app was level with the catalog, so the update failed in 0.473 s and `safety-dump` was never observed. A miss recorded as a miss, then fixed | +| **B8, and Phase 2.3's first attempt** | **re-run** after B1's own precondition swept the app they needed | B1 removes every other behind-app so the unattended caller has exactly one thing to react to. That is correct and is recorded; it also removed `docmost`. Re-run in `phase2_redo.sh` | +| **The harness runs of the new catalog EDGES (U1–U7)** | **code shipped, runs OWED** | the box-side result for each of those edges exists; the harness adds the ABORT step, and setting up `/opt/upg` was not worth the last hour against the teardown | + +**And one thing the brief asked for that this VENUE cannot produce at all, named rather than left +blank: every event and every customer mail.** Guest 9202 runs `hub.enabled: false` and every +notifier entry point returns before it logs anything (R-620). The hub was not enabled to get around +it — that would register an unclaimed host at the live hub and could mail a real address, and the +brief fences the hub. The alarm truth table below says so on every row. --- ## The three lines - +**Interventions: ZERO.** Nothing tonight needed an act a household could not perform from the +screens. Every app was deployed, seeded, updated, held, restored and removed through the product's +own endpoints; the only non-product commands were the power cuts (`pct stop`, which IS the fault +being tested) and the reproduction of a refusal the product had destroyed. + +**21 edges attempted: 14 proven, 3 failed, 4 inconclusive.** Up from the **three** apps this project +had ever measured. Ten of the fourteen printed a verbatim migration line, so the database really was +rewritten and the data still read back. + +**The one result that matters most: three of the 53 templates name a health probe the app does not +answer — and because the guarded Update waits on that same probe, a SUCCESSFUL update ends by +STOPPING a working app.** `tandoor` was measured serving HTTP 200 on the new version at four samples +across five minutes, with docker's own healthcheck green, and was then stopped by `failAndHold` and +the household sent to a restore they did not need. `zipline` and `wger` are the same defect. **R-618, +P1.** No data is lost and the restore works — but "the update is guarded" must not be read as "the +guard is right about whether the app came up". --- @@ -119,25 +149,181 @@ one is internal. Evidence: `06-drift-rerun.txt`. ## The verdict table - +One row per edge attempted tonight. `inconclusive` means *we could not measure it*, which +is a different fact from *it does not work* — and only one of them is about the app. + +| app | from → to | class | box verdict | seed before → after | secs | migration line seen | evidence | +|---|---|---|---|---|---|---|---| +| `actualbudget` | actual-server:26.7.0 → actual-server:26.9.0 | other | **proven** | True → True | 19.5 | yes | `apps/actualbudget/` | +| `audiobookshelf` | audiobookshelf:2.35.1 → audiobookshelf:2.36.1 | file-leg | **proven** | True → True | 23.6 | yes | `apps/audiobookshelf/` | +| `bookstack` | bookstack:26.05.2, mariadb:12.3 → bookstack:26.05.5, mariadb:12.3 | db-mariadb | **proven** | True → True | 45.1 | none printed | `apps/bookstack/` | +| `docmost` | docmost:0.95.0, postgres:16-alpine, redis:7-alpine → docmost:0.96.0, postgres:16-alpine, redis:7-alpine | db-postgres | **proven** | True → True | 103.6 | yes | `apps/docmost/` | +| `grafana` | grafana:13.1.0 → grafana:13.2.2 | other | **proven** | True → True | 26.7 | yes | `apps/grafana/` | +| `home-assistant` | home-assistant:2026.7.2 → home-assistant:2026.9.3 | other | **proven** | True → True | 103.6 | none printed | `apps/home-assistant/` | +| `mealie` | mealie:v3.20.1 → mealie:v3.27.0 | db-postgres | **proven** | True → True | 18.5 | yes | `apps/mealie/` | +| `n8n` | n8n:2.31.3 → n8n:2.40.5 | db-postgres | **proven** | True → True | 117.9 | yes | `apps/n8n/` | +| `navidrome` | navidrome:0.63.2 → navidrome:0.64.0 | file-leg | **proven** | True → True | 11.3 | yes | `apps/navidrome/` | +| `nextcloud` | mariadb:11.6 → mariadb:12.3 | engine-major-mariadb | **proven** | True → True | 217.4 | none printed | `apps/nextcloud-engine-mariadb/` | +| `papra` | papra:26.6.1-rootless → papra:26.6.2-rootless | other | **proven** | True → True | 60.5 | yes | `apps/papra/` | +| `privatebin` | pdo:2.0.5 → pdo:2.0.6 | file-leg | **proven** | True → True | 15.4 | none printed | `apps/privatebin/` | +| `romm` | mariadb:11.4, redis:7-alpine, romm:5.0.0 → mariadb:11.4, redis:7-alpine, romm:5.3.0 | db-mariadb | **proven** | True → True | 60.6 | yes | `apps/romm/` | +| `vikunja` | vikunja:2.3.0 → vikunja:2.6.0 | other | **proven** | True → True | 24.6 | yes | `apps/vikunja/` | +| `adventurelog` | adventurelog-backend:v0.12.1, adventurelog-frontend:v0.12.1, postgis:16-3.5-alpine → adventurelog-backend:v0.13.0, adventurelog-frontend:v0.13.0, postgis:16-3.5-alpine | db-postgis | **failed** | True → False | 346.6 | — | `apps/adventurelog/` | +| `docmost` | postgres:16-alpine → postgres:17-alpine | engine-major-postgres | **failed** | True → False | 254.5 | — | `apps/docmost-engine-postgres/` | +| `gitea` | gitea:1.27.0 → — | other | **inconclusive** | False → False | 20.1 | — | `apps/gitea/` | +| `opengist` | opengist:1.13 → opengist:1.15 | other | **inconclusive** | True → False | 14.4 | — | `apps/opengist/` | +| `tandoor` | postgres:16-alpine, recipes:2.6.13 → postgres:16-alpine, recipes:2.6.15 | db-postgres | **failed** | True → False | 361.9 | — | `apps/tandoor/` | +| `vaultwarden` | server:1.36.0-alpine → — | other | **inconclusive** | False → False | 17.9 | — | `apps/vaultwarden/` | +| `zipline` | postgres:16-alpine, zipline:4.6.1 → — | db-postgres | **inconclusive** | False → False | 73.3 | — | `apps/zipline/` | + +**14 proven · 3 failed · 4 inconclusive — out of 21 attempted.** + +### Why each inconclusive edge could not be judged + +- **`gitea`** — INCONCLUSIVE: the template sets no `INSTALL_LOCK`, so a fresh Gitea starts in its web-installer state and `gitea admin user create` refuses with `MustInstalled() [F] Unable to load config file for a installed Gitea instance`. The route that would work is POSTing the installer form first; that was not written tonight and is listed as owed rather than faked. +- **`opengist`** — CORRECTED from `failed` to `inconclusive` the same night, deliberately. The UPDATE itself SUCCEEDED: phase `done` in 14.4 s, and all four version observables agree on `ghcr.io/thomiceli/opengist:1.15` with the container running and zero restarts. What failed was the READBACK: it was attempted immediately after `done` and the sign-in form was not yet being served, so the fixture got no `_csrf` and returned `http=None`. Whether the seeded account survived was therefore NOT ESTABLISHED. Recording that as `failed` would have blamed the app for the harness's impatience — `inconclusive` is the honest verdict and it is never collapsed into `failed`. The fixture now waits for the LOGIN FORM rather than for the root page. +- **`vaultwarden`** — INCONCLUSIVE BY DESIGN, not by a gap in the harness: the catalog CLOSES self-registration on purpose (`SIGNUPS_ALLOWED=false`, R-512 — *a stranger who guesses vault. must not be able to register*), so `/api/accounts/register` answers 404 and there is NO account-creating route without the admin secret. Vaultwarden also ships no CLI. Tried: `POST /api/accounts/register` with a valid KDF envelope. This app cannot be seeded headlessly while that setting stands, and the setting is right. +- **`zipline`** — INCONCLUSIVE BY DESIGN: the deploy answers `E1037: User registration is disabled`, so no first account can be created from outside. Tried: `POST /api/auth/register` and `POST /api/auth/setup`. SEPARATELY, zipline is one of R-618's two confirmed victims — its `.felhom.yml` probe expects 200 on `/api/health`, which the app answers 404 — so even with a seed its update would have been HELD by a wrong probe rather than by anything about the edge. + +### The edges that failed — the most valuable results of the night + +- **`adventurelog`** — final phase `failed`, hold `A(z) adventurelog frissítése 2026-09-21 20:53-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-21 20:47 — ez a másolat a beállításokat, az adatbázist és az adatköteteket tartalmazza.`, error `A(z) adventurelog frissítése 2026-09-21 20:53-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-21 20:47 — ez a másolat a beállításokat, az adatbázist és az adatköteteket tartalmazza.`. the edge ended HELD or failed — this is a RESULT, not an error of the run +- **`docmost`** — final phase `failed`, hold `A(z) docmost frissítése 2026-09-21 21:44-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-21 21:44 — ez a másolat a beállításokat, az adatbázist és az adatköteteket tartalmazza.`, error `A(z) docmost frissítése 2026-09-21 21:44-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-21 21:44 — ez a másolat a beállításokat, az adatbázist és az adatköteteket tartalmazza.`. ended HELD or failed — a RESULT, not an error of the run +- **`tandoor`** — final phase `failed`, hold `A(z) tandoor frissítése 2026-09-21 21:22-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-21 21:15 — ez a másolat a beállításokat, az adatbázist és az adatköteteket tartalmazza.`, error `A(z) tandoor frissítése 2026-09-21 21:22-kor nem sikerült, és az alkalmazás nem indult el az új verzióval. Az alkalmazás biztonsági okból leállítva marad, hogy az adatai ne sérüljenek. Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-21 21:15 — ez a másolat a beállításokat, az adatbázist és az adatköteteket tartalmazza.`. the edge ended HELD or failed — this is a RESULT, not an error of the run --- ## Phase 2 — the two database engines - +### 2.1 MariaDB across a major, through the REAL Update button — PROVEN, and a first + +`nextcloud`, app image held constant, `mariadb: 11.6 → 12.3`. Seeded and read back through +`occ user:add` / `occ user:info`, with the fixture's own negative control on every readback. + +**The four observables of `SPIKE-r459`, before → after:** + +| # | observable | before | after | +|---|---|---|---| +| 1 | `mariadb_upgrade_info` | `11.6.2-MariaDB` | **`12.3.3-MariaDB`** | +| 2 | the engine's own check (R-464 — never the log line) | *not measured: the probe was unauthenticated, see below* | **„This installation of MariaDB is already upgraded to 12.3.3-MariaDB. There is no need to run mariadb-upgrade again."** | +| 3 | the entrypoint | — | **„Major version upgrade detected from 11.6.2-MariaDB to 12.3.3-MariaDB. Check required!"** → „Starting mariadb-upgrade" → **„Finished mariadb-upgrade"** | +| 4 | the engine's own pre-upgrade backup | absent | **`system_mysql_backup_11.6.2-MariaDB.sql.zst`, 631 905 bytes** | + +**Observable 3 is the one that matters**, because R-459's whole finding was that MariaDB can apply a +major and *skip* the conversion quietly, printing `skipped due to $MARIADB_AUTO_UPGRADE`. **That line +is absent**; the conversion was detected, started and finished. The seeded account read back and all +four version observables agree. + +**Two honest notes on the instrument.** The BEFORE capture of observable 2 asked the engine without +credentials and got `ERROR 1045 Access denied`; the probe was corrected and the AFTER capture +retaken with it, so the before value is **not measured** and is stated as such rather than inferred. +And the `ls` in observable 4 printed a "(no pre-upgrade backup file present)" fallback *after* +listing the file, because it globs two patterns and one did not match — the file is there. + +Full record: `16-phase2.1-mariadb-major.md`, `apps/nextcloud-engine-mariadb/`. + +### 2.2 PostgreSQL across a major — what a household would see TODAY + +**Exactly what R-463 predicted, and nobody had measured.** `docmost`, engine only, `16 → 17`: +the update ended **`failed` in 5.1 s**, the app was stopped and held, **the pin named +`postgres:17-alpine` while `installed_images` still said 16 and nothing was running**, and the data +was intact. The restore the hold sentence names brought it back in **29.1 s**, hold cleared, health +probe 200. + +**The engine's refusal line had to be REPRODUCED**, because `failAndHold` removed the container +before any probe could read it (**R-621**) and the controller log does not carry it either. Done +independently with a control on every step — source proven 16, copy proven 16, 49 MB: + +``` +FATAL: database files are incompatible with server +DETAIL: The data directory was initialized by PostgreSQL version 16, + which is not compatible with this version 17.11. +``` + +**The datadir was still `16` afterwards** — nothing migrated, nothing damaged — and the positive +control (the same copy under `postgres:16-alpine`) started and held **48 tables**. + +**My own first reproduction was WRONG and is kept, labelled.** The volume lookup returned empty, so +the copy was empty, so 17 initialised a fresh datadir and started happily — and the run reported +`running=true` as though no refusal had happened. A blank `PG_VERSION` one line earlier should have +stopped the step and did not. It is kept because it accidentally measured the MIRROR case (16 +refusing a 17 datadir, verbatim), and relabelled so nobody reads it as the main result. +`17-postgres-refusal-reproduced.txt` (wrong) and `18-postgres-refusal-reproduced.txt` (right). + +### 2.3 The conversion rehearsal, COSTED — the answer Q5 was asking for + +Logical dump and restore, on a fresh seeded `docmost`: **49.0 MB datadir, 48 tables.** + +| step | time | what it produced | +|---|---|---| +| dump with 16 (`pg_dumpall`) | **2.6 s** | **132 201 bytes**, 48 `CREATE TABLE` statements | +| fresh 17 datadir + restore | **6.5 s** | `PG_VERSION` 17, **48 tables restored**, 2 benign ERROR lines | +| point the app at 17 and start it | 124.8 s | the app's own words: *„Database connection successful"* | +| **the seed read back on 17** | — | **TRUE**, through the app's own login | +| **total** | **155.9 s** | of which **~9 s is engine work** | + +Full paragraph for Q5, including what could lose data and why `pg_upgrade` was not run: +`24-Q5-postgres-conversion-costed.md`. --- ## Phase 3 — the bad days - +Every leg records the same five things. **The event/mail column is empty on every row for the same +structural reason — see the alarm truth table.** + +| leg | what the household saw | what the box did by itself | time to steady | the alarm | +|---|---|---|---|---| +| **B1 unattended HOLD** | „Frissítés elérhető" → app **Leállítva**, the hold sentence naming tier, date and what the copy holds; the banner „Telepített alkalmazás nem fut" on **every** page | pressed **once**, held after **312.9 s**, and **never pressed again** across two further passes | 312.9 s | unmeasurable (R-620) | +| **B2 pull fails** | „Az új verzió letöltése nem sikerült, ezért a frissítés elmaradt. Az alkalmazás a korábbi verzióval fut tovább." | pin **and** definition put back in **1.0 s**, `hold=None`, old version still serving | 1.0 s | none, correctly — nothing is down | +| **B3 busy** | „A frissítés most nem indítható: mentés/visszaállítás folyamatban." | refused `409 reason='busy'` on six consecutive presses; the transient reason a caller needs | — | n/a | +| **B4 concurrency** | nothing — all succeeded | **NO single-flight.** 2 of 2, then **5 of 5**, ran at once; all ended `done`, every pin advanced | ~30 s for five | none | +| **B5 cut in `backing-up`** | „A frissítés megszakadt, mert a vezérlő újraindult…" | the box said so itself at boot (three positive observables), **pin unmoved**, data intact | one boot | none | +| **B5 cut in `safety-dump`** | nothing — the ordinary badge, **no interrupted sentence** | no recovery line, no journal, **pin unmoved**, data intact | one boot | none | +| **B6 way out forwards** | badge still says „Frissítés elérhető" and offers the button | the button **refuses `409 reason='held'`** | — | n/a | +| **B7 disk floor** | „Nincs elég szabad hely a frissítéshez: 1.4 GB szabad… legalább 2 GB szükséges." | refused before anything moved | instant | n/a | +| **B8 floating pin** | „Naprakész" | and it is **TRUE on this box** — both floating digests match upstream exactly | — | n/a | +| **B9 frozen app, newer `.felhom.yml`** | nothing — 10 samples, all `running`/200 | the newer `.felhom.yml` reached the frozen app; the compose stayed frozen | — | none | + +**The three that changed what is known:** + +1. **B1 produced the unattended HOLD** this project has never had — see `19-Q4-the-unattended-hold.md`. +2. **B4 answered the single-flight question: there is none.** Five updates ran together and all + ended honest. Slice 6 must decide whether that is what it wants. +3. **B6 found an inconsistency R-524 already removed for the other case:** a held app keeps inviting + the household to update and the button refuses. **R-625.** + +Also proven for free, across a genuine power cut: the boot sweep met a held app after an unclean +shutdown and deliberately left it alone — *„whatever is holding it owns its recovery"*. --- ## Phase 4 — the morning after - +**B1's held app, as a household would find it at breakfast.** The app page, the dashboard, the +launcher and both backups pages were read in **both languages** and are saved as HTML in +`bad-days/P4-morning-after/`. + +**Is there ONE sentence that says what happened, since when, which copy holds what, and what to +press?** Scored against Q4's recommended option: + +| Q4 promises the household are told… | measured | +|---|---| +| **what happened** | ✔ „…frissítése 2026-09-21 21:53-kor nem sikerült, és az alkalmazás nem indult el az új verzióval." | +| **since when** | ✔ the time is in the sentence | +| **which copy holds what** | ✔ „saját meghajtó, 2026-09-21 21:47 — ez a másolat a beállításokat, az adatbázist és az adatköteteket tartalmazza." | +| **what to press** | ✔ „Visszaállítható a Mentések oldalon…" | +| **the same in English** | ✘ **the sentence is Hungarian on the English page** (R-606, confirmed on the hold sentence itself, with positive and negative controls) | +| by mail | **unmeasurable on this venue** (R-620) | + +**The app is surfaced everywhere, not only on its own page** — the banner „Telepített alkalmazás nem +fut: …" / „An installed app is not running: …" appeared at the top of every authenticated page, and +carried both held apps at once when there were two. + +**Every app still on the box, and every badge, after a rescan.** Ten deployed apps: every badge is +**TRUE** — references equal ⇔ „Naprakész", references differ ⇔ „Frissítés elérhető". `zipline` shows +the household „Nem egészséges — URL nem elérhető" while it is in fact serving, which is R-618 in the +household's own words. --- @@ -153,22 +339,154 @@ the hub. Filed as **R-620** (a disabled notifier should at least say which event So the table below scores the surfaces that DO exist on this box — the app page in both languages, the dashboard, and the controller's own log — against `08-alarm-ladder.md`. - +| # | what happened | should it alarm, per `08` | what the box did | what the household could READ | verdict | +|---|---|---|---|---|---| +| 1 | `adventurelog` held after a real failed edge — app **stopped** | **YES** — `stopped` is in `IsDownState` | classified `stopped`; the boot sweep refused to restart it | the hold sentence on the app page **and** a banner on every page | **correct** — but the SEND is unmeasurable here | +| 2 | `tandoor` held the same way | **YES** | same | same, both apps in one banner | **correct**, same caveat | +| 3 | `glance` held by the **unattended** update | **YES** | same, and honoured across a **power cut** | same | **correct**, same caveat | +| 4 | `tandoor` and `zipline` reading `unhealthy` for hours while SERVING | **NO** — `08` §4 puts `unhealthy` deliberately in the not-down set | did not alarm | „Nem egészséges — URL nem elérhető" on the dashboard | **the ladder is right and the outcome is still wrong** — see below | +| 5 | pull failure (B2) — app kept running the old version | **NO** — nothing is down | did not alarm | one sentence on the card | **correct** | +| 6 | five updates at once (B4) | **NO** | did not alarm | nothing | **correct** | +| 7 | two power cuts (B5) | `restarting` is not down until sustained | recovered; nothing alarmed | one interrupted sentence in one case, nothing in the other | **correct** | +| 8 | disk under the 2 GB floor (B7) | not an app-down state | refused the update; no alarm | the refusal sentence | **correct** — though a box at 1.4 GB free is arguably worth telling someone about, and nothing does | + +**Which alarm fired and was it true:** none fired, and none could — see the venue limit above. Every +classification the box made was correct against `08`. + +**Which should have fired and did not:** on this evidence, none. Row 8 is the only candidate and it +is a design question rather than a defect: `08` is an *app-down* ladder and a full disk is not an app +being down. + +**The one that matters, and it is row 4.** `08` §4 deliberately excludes `unhealthy` — *"folding it +in reintroduces the flapping fix-3 was written to stop"* — and that ruling is right. **But the same +probe result the alarm ladder correctly ignores is NOT ignored by the guarded update's `verifying` +phase, which waits on it and then stops the app.** One probe, two consumers, opposite tolerances, +and neither document says so. That asymmetry is the whole of R-618's severity. --- ## The promotion list for the operator - +**CC promotes nothing.** These are the real, within-a-major edges that ended `proven` on +the box tonight, with the data read back through the app's own front door both before and +after. Moving each of them on the LIVE catalog is the operator's call. + +| app | the move | what it would mean for a box in the field | +|---|---|---| +| `actualbudget` | actual-server:26.7.0 → actual-server:26.9.0 | the app runs its own schema migration on the way — proven here, and the update takes a backup first | +| `audiobookshelf` | audiobookshelf:2.35.1 → audiobookshelf:2.36.1 | the app runs its own schema migration on the way — proven here, and the update takes a backup first | +| `bookstack` | bookstack:26.05.2, mariadb:12.3 → bookstack:26.05.5, mariadb:12.3 | no migration line printed; the app came up on the new version with its data intact | +| `docmost` | docmost:0.95.0, postgres:16-alpine, redis:7-alpine → docmost:0.96.0, postgres:16-alpine, redis:7-alpine | the app runs its own schema migration on the way — proven here, and the update takes a backup first | +| `grafana` | grafana:13.1.0 → grafana:13.2.2 | the app runs its own schema migration on the way — proven here, and the update takes a backup first | +| `home-assistant` | home-assistant:2026.7.2 → home-assistant:2026.9.3 | no migration line printed; the app came up on the new version with its data intact | +| `mealie` | mealie:v3.20.1 → mealie:v3.27.0 | the app runs its own schema migration on the way — proven here, and the update takes a backup first | +| `n8n` | n8n:2.31.3 → n8n:2.40.5 | the app runs its own schema migration on the way — proven here, and the update takes a backup first | +| `navidrome` | navidrome:0.63.2 → navidrome:0.64.0 | the app runs its own schema migration on the way — proven here, and the update takes a backup first | +| `nextcloud` | mariadb:11.6 → mariadb:12.3 | no migration line printed; the app came up on the new version with its data intact | +| `papra` | papra:26.6.1-rootless → papra:26.6.2-rootless | the app runs its own schema migration on the way — proven here, and the update takes a backup first | +| `privatebin` | pdo:2.0.5 → pdo:2.0.6 | no migration line printed; the app came up on the new version with its data intact | +| `romm` | mariadb:11.4, redis:7-alpine, romm:5.0.0 → mariadb:11.4, redis:7-alpine, romm:5.3.0 | the app runs its own schema migration on the way — proven here, and the update takes a backup first | +| `vikunja` | vikunja:2.3.0 → vikunja:2.6.0 | the app runs its own schema migration on the way — proven here, and the update takes a backup first | + +**And the apps that must NOT be promoted, which is the other half of the list:** + +| app | the move | why not | +|---|---|---| +| `adventurelog` | `v0.12.1 → v0.13.0` | applies **nine database migrations successfully** and then never binds its port. Held after the full health wait; restored in 75 s. **R-622** | +| `tandoor` | `2.6.13 → 2.6.15` | the update SUCCEEDS — the app served HTTP 200 on the new version for five minutes — and is then stopped by the template's own wrong health port. Fix **R-618** first; the edge itself is probably fine | +| `postgres:16-alpine → 17-alpine`, anywhere | the engine | refuses to start on a 16 datadir, verbatim. The engine-major gate stays until Q5's conversion exists | + +**`nextcloud`'s MariaDB `11.6 → 12.3` is proven and is a different kind of entry**: it is not an app +version but an engine major, and §3 decision 5 plus R-469 already permit it. It is listed here +because tonight is the first time it has been pressed through the button a household presses. --- ## Teardown — three layers, plus Gitea - +### Machine — guest 9202 + +Every throwaway app removed **through the product**, with its data where the product allowed it. +Three apps carrying an `HDD_PATH` were REFUSED at „remove with data" — `/api/disks` answers +`agent not configured` on this guest, so the drive path cannot be resolved and R-442's fail-closed +guard keeps the app rather than half-deleting it. Each was then removed with the data KEPT, which +the product does accept, and the harness's own directories were removed by name afterwards. + +``` +containers now: felhom-controller · filebrowser · traefik <- the three protected only +drill images: (none) +drill volumes: (none) +registry:2: removed, with its volume +scratch drive: documents · downloads · media · roms <- the drive's own folders +free space: 33 G on the docker root +``` + +**`controller.yaml` restored** from `controller.yaml.pre-update-night`, the controller restarted, +and the cache's origin **read back and quoted** — which is the point of the exercise: + +``` +origin https://gitea.dooplex.hu/admin/app-catalog-felhom.eu.git (fetch) +origin https://gitea.dooplex.hu/admin/app-catalog-felhom.eu.git (push) +4463243 Upgrade harness: four fixtures and seven real upstream edges from the update night (R-462) +``` + +**The teardown found the night's last defect**, which is the argument for doing it properly: a +`navidrome` container that the product's own removal had left behind and Docker's restart policy had +resurrected, invisible to every sweep that keys on `deployed`. Removed by name. **R-626.** + +### Host — demo-hp + +**No harness LXC was created tonight, so none was destroyed** — the PostgreSQL rehearsal ran on 9202 +itself with plain `docker` beside the product. `pct list` and `pvesm status` before and after are in +`teardown/00-host-before.txt` and `04-host-after.txt`. **Guest 9201 untouched** — container count +unchanged, and it was never addressed except to read its catalog cache for the negative control. + +### Hub + +**Nothing provisioned: no customer, no config, no appliance, no binding.** The only hub act of the +whole night was the floor save in Phase 0.1. Final state: floor `0.261.0`, declared MinAgent +`0.131.0`, three host rows — unchanged from the start except the floor the operator asked for. + +### Gitea + +``` +live catalog origin/main : 4463243f2e09 +drill repo HEAD after reset : 4463243f2e09 +diff of every `image:` line, live vs drill: +IDENTICAL — every image: line matches the live catalog +``` + +The drill repo is **KEPT**, private, and reset to the live catalog's `main`, so the next drill starts +clean. **The live catalog's `main` moved once tonight** — from `f5f6a152b513` to `4463243f2e09` — and +that commit changes `scripts/` only: four harness fixtures and seven edge definitions. **Zero +`image:` lines moved on the live catalog at any point in the night**, which the diff above proves +rather than asserts. --- ## Claims in the brief that turned out wrong - +The brief asked for this explicitly. Each claim, and what was measured. + +| the brief said | measured | +|---|---| +| **a box will follow a second catalog by `git.repo_url` alone** (read from config source, never run) | **WRONG.** `Syncer.gitCloneOrPull` clones only when `.git` is absent; otherwise it fetches from the remote the clone already stores. The cache directory had to be removed too. **R-615** | +| **`CompareImageRefs` may not order references carrying a `host:port/` prefix** (read, not run) | **The worry was unfounded.** `splitImageRef` takes the LAST colon and rejects it only when a `/` follows, so a registry port is never read as a tag. Proven by RUNNING it: four positive cases and a negative control | +| **PostgreSQL 17 refuses a 16 datadir and the update ends HELD with data intact** (R-463 and source, not measured) | **RIGHT, and now measured** — 5.1 s to held, pin on 17 with nothing running, data intact, restore back in 29.1 s. The refusal line itself had to be reproduced because the product destroyed it (**R-621**) | +| **a MariaDB sidecar major through the BUTTON behaves as it did on the harness** | **RIGHT.** All four observables, including the conversion actually running rather than being skipped, and the engine's own pre-upgrade backup | +| **9202 has the capacity for this** | **RIGHT.** 26 GB RAM, 56 GB free on the docker root at the start, a 938 GB scratch drive. Peak usage never threatened it; images were reclaimed BY NAME twice, never pruned | +| **39 within-a-major edges still exist upstream tonight** | **RIGHT, exactly.** The drift script re-run at 20:07 returned 66 pins, 46 behind, **39 within a major and 7 across** — the same numbers | + +**And two more the brief did not name, found the same way:** + +- `update-arc-gaps-2026-09-21/00-api-recipe.md` said the app page is `/app/`. **It is `/apps/`**, + and every call that recipe described 404s. Corrected in that file. +- `unattended-caller.py`'s `follow()` read the API **envelope**, so every update it followed would + have run to a 900 s timeout and been recorded `timeout` rather than `held`. **R-623**, fixed before + B1 relied on it — and B1's log is what the fixed version produces. + +**One correction to a register row, which is the same class of error one layer up.** R-606 records +controller v0.260.0 as having made the pre-flight refusals reach an English household in English. +**Measured: it did not.** `held`, `not_deployed` and `disk` all come back identical Hungarian with +`?lang=en`, because the routing exists and the sentences are frozen string constants that never +entered it. A row that records something as fixed when it is not is worse than an open row. diff --git a/documentation/audits/update-night-2026-09-21/22-wger-probe-measured.txt b/documentation/audits/update-night-2026-09-21/22-wger-probe-measured.txt new file mode 100644 index 00000000..d45a8054 --- /dev/null +++ b/documentation/audits/update-night-2026-09-21/22-wger-probe-measured.txt @@ -0,0 +1,13 @@ +=== wger — R-618's last unmeasured candidate + template probe : type=http port=80 (.felhom.yml) + traefik label : loadbalancer.server.port=8000 (the SAME template's compose) + --- what wger actually LISTENS on, asked inside its own container: + --- the container's own docker healthcheck, if it has one: + healthy + --- port 80 from inside (the port the probe names): + refused on 80 + --- port 8000 from inside (the port the compose says): + ANSWERED on 8000 + the household's own front door: http=302 + the box's own verdict: state='unhealthy' + VERDICT: CONFIRMED DEFECT — the same shape as tandoor diff --git a/documentation/audits/update-night-2026-09-21/24-Q5-postgres-conversion-costed.md b/documentation/audits/update-night-2026-09-21/24-Q5-postgres-conversion-costed.md new file mode 100644 index 00000000..6388e721 --- /dev/null +++ b/documentation/audits/update-night-2026-09-21/24-Q5-postgres-conversion-costed.md @@ -0,0 +1,80 @@ +# `09` §3b Q5 — PostgreSQL 16 → 17: what a household sees today, and what a conversion costs + +Both halves measured 2026-09-21, guest 9202, on a REAL app with REAL seeded data. + +--- + +## (a) What a household would see TODAY — measured, and it is what R-463 predicted + +`docmost`, seeded through its own API and read back first (control C1). Drill-catalog bump of the +**`postgres:` sidecar alone**, `16-alpine → 17-alpine`; the app image did not move. The guarded +Update was pressed. + +| | | +|---|---| +| time to the verdict | **5.1 s** — the engine does not try, it refuses at once | +| final phase | `failed`, app **stopped and held** | +| the pin | **`postgres:17-alpine`** — while nothing is running on 17 | +| `installed_images` | still `postgres:16-alpine` — the observation and the decision disagree, correctly (§5.2) | +| the data | **intact** | +| the way out | the restore the hold sentence names: **29.1 s**, hold cleared, app back, health probe 200 | + +**The engine's own refusal line had to be REPRODUCED**, because `failAndHold` removed the container +before any probe could read it (R-621) and the controller log does not carry it either. Reproduced +independently, with a control on every step — source proven 16, copy proven 16, 49 MB: + + FATAL: database files are incompatible with server + DETAIL: The data directory was initialized by PostgreSQL version 16, + which is not compatible with this version 17.11. + +**And the datadir was still `16` afterwards** — nothing was migrated, nothing was damaged. Positive +control: the same copy under `postgres:16-alpine` starts and holds **48 tables**. + +So R-463's reading is confirmed on the box: *the engine refuses to start, the update ends HELD, the +data is intact.* Nothing else happened, and the household's route back works. + +--- + +## (b) The conversion rehearsal, COSTED + +Route: **logical dump and restore.** Plain `docker` beside the product — there is no product path +for this, and pricing one is the point. On a fresh, seeded `docmost`: **49.0 MB datadir, 48 tables.** + +| step | time | what it produced | +|---|---|---| +| dump with 16 (`pg_dumpall`) | **2.6 s** | **132 201 bytes**, 48 `CREATE TABLE` statements | +| fresh 17 datadir + restore | **6.5 s** | `PG_VERSION` 17, **48 tables restored**, 2 benign ERROR lines | +| point the app at 17 and start it | 124.8 s | *„Database connection successful"* — the app's own words | +| **the seed read back on 17** | — | **TRUE**, through the app's own login | +| the 17 datadir afterwards | — | 49.1 MB (from 49.0 MB) | +| **total** | **155.9 s** | of which **~9 s is the engine work**; the rest is the app restarting | + +**The two ERROR lines are benign and are named so nobody reads them as data loss:** +`role "docmost" already exists` and `database "docmost" already exists` — `pg_dumpall` recreates +both, and the 17 container's entrypoint had already made them. The 48-table count after the restore +is the positive control that the replay worked. + +## What Q5 now has that it did not + +- **A price.** Nine seconds of engine work for a 49 MB database, and under three minutes end to end + including the app restart. For eleven apps that is a maintenance window, not a project. +- **A shape that works**, walked once on real data: stop the app, keep the engine, `pg_dumpall`, + fresh 17 volume, replay, re-point, start, read the data back through the app's own door. +- **What could lose data, named:** the dump is the single point of failure. Nothing in the rehearsal + verified the dump before the old datadir was left behind — because nothing had to, since the old + volume was untouched throughout. **Any real procedure must keep the 16 datadir until the app has + been read back on 17**, which is exactly what this rehearsal did by accident of being a rehearsal. +- **The `pg_upgrade` route was NOT run.** It needs both majors' binaries in one image and no such + image exists in this project. Naming it costs nothing; building it is the work Q5's first option + is really asking for, and the logical route above may make it unnecessary at this size. + +**This is a rehearsal and a costing, not a procedure for the catalog.** The engine-major gate stays. + +## One fact about the harness's own instrument, since Q5's option 1 rests on it + +`upgrade-test.py`'s PostgreSQL probe is `cat /var/lib/postgresql/data/PG_VERSION` **inside the +container**. Against the converted datadir it answered **`17`, exit 0** — it works. **But it is +blind in exactly the case that matters**: when PostgreSQL refuses, the container is not running, so +`docker exec` cannot ask it anything. The probe's own honesty rule covers this (*"a probe that cannot +run records why"*), and tonight it recorded `No such container`. A probe that can only speak when the +engine is happy is worth having — but it must never be read as *the engine is content*. diff --git a/documentation/audits/update-night-2026-09-21/25-B5-the-two-early-cuts.md b/documentation/audits/update-night-2026-09-21/25-B5-the-two-early-cuts.md new file mode 100644 index 00000000..aac174eb --- /dev/null +++ b/documentation/audits/update-night-2026-09-21/25-B5-the-two-early-cuts.md @@ -0,0 +1,63 @@ +# B5 — a power cut in the two EARLY phases nobody had cut in + +R-520 cut in `pulling` (nothing had run — the easy case). R-610 cut after `starting` (the migration +had run — the dangerous case). The two that were left are the ones that touch the customer's COPY +rather than their data. + +**Instrument limit, stated on both, never glossed:** the poll is 200 ms and `pct stop` returns in +4–10 s, so **the phase at the DECISION is observed and the phase at the FREEZE is inferred.** + +--- + +## Cut 1 — during `backing-up` (privatebin, `backup_max_age` lowered to 1m so the phase happens) + + phase 'backing-up' OBSERVED at 2026-09-21T20:07:33.345Z — pulling the plug NOW + `pct stop 9202` returned after 9.65s + +**After the boot the box said so ITSELF — three positive observables, not an absence:** + + [appstop] crash recovery: an app-data backup (volume dump) (op "volume-dump:privatebin") was + interrupted and left 1 app(s) stopped — restarting them: [privatebin] + [appstop] crash recovery: restarted privatebin after the interrupted an app-data backup (volume dump) + [stacks] update recovery: privatebin was interrupted in backing-up (started 2026-09-21T20:…) + +| | | +|---|---| +| the pin | **did NOT move** | +| the app | running, and a fresh paste seeded and read back → **data intact** | +| the household reads | „A frissítés megszakadt, mert a vezérlő újraindult, mielőtt az új verzió elindult volna." | +| in English? | **No — the same Hungarian sentence on the English page** (R-606, the original instance, re-confirmed) | + +## Cut 2 — during `safety-dump` (privatebin, a real pending edge) + + phase 'safety-dump' OBSERVED at 2026-09-21T20:19:22.234Z — pulling the plug NOW + `pct stop 9202` returned after 3.96s + +| | | +|---|---| +| recovery lines | **none for privatebin** — and no update journal on disk | +| the pin | **did NOT move** (`…/paste:2.0.1`, unchanged) | +| the app | running, fresh paste seeded and read back → **data intact** | +| the household reads | the ordinary „Frissítés elérhető — ma" — **no interrupted sentence at all** | + +**The two cuts differ, and the difference is the finding rather than a discrepancy.** A cut in +`backing-up` leaves a trace the box acts on at boot (an app stopped by the dump, and an update +recorded as interrupted); a cut in `safety-dump` leaves nothing to act on, and the box says nothing +because there is nothing to say. Both end in the same place: **nothing moved, nothing half-written, +the data readable.** + +## The claim this leg existed to test + +*A backup artefact half-written must not be left looking whole.* **No half-written artefact was +found on either cut** — the backup listings before and after are in each leg's `00-`/`01-` files, +and the zero-byte sweep found none. + +## A second thing this leg proved, for free, across a REAL power cut + + [bootrecon] "glance" is a boot orphan by intent but is HELD (held after a failed update + (2026-09-21T19:53:02Z) — restore it from its backup to start it) + — NOT starting it; whatever is holding it owns its recovery + +`09` §6.1 records that three unattended paths honoured no hold before v0.237.0 and now do. **That is +now proven across a genuine power cut**: the boot sweep met a held app after an unclean shutdown and +deliberately left it alone. diff --git a/documentation/audits/update-night-2026-09-21/26-removed-app-came-back.txt b/documentation/audits/update-night-2026-09-21/26-removed-app-came-back.txt new file mode 100644 index 00000000..3788d59b --- /dev/null +++ b/documentation/audits/update-night-2026-09-21/26-removed-app-came-back.txt @@ -0,0 +1,28 @@ +=== navidrome AFTER a removal that returned 200 +created=2026-09-21T19:12:20.845623139Z started=2026-09-21T20:19:33.374056273Z restart=unless-stopped image=deluan/navidrome:0.64.0 +compose project label: navidrome +app.yaml: ls: cannot access '/opt/docker/stacks/navidrome/app.yaml': No such file or directory +volume: 1 navidrome volume(s) left + +TIMELINE, from the container's own metadata +------------------------------------------- + 21:12:05 local POST /api/stacks/navidrome/remove {remove_hdd_data:true} -> 409 (drive path + unresolvable, R-442's fail-closed guard — correct) + 21:12:11 local POST /api/stacks/navidrome/remove {remove_hdd_data:false} -> 200, + volumes_removed = ['navidrome_navidrome_data'] + 21:12:18 local the harness's own leftovers check reported deployed=False and no containers + 21:12:20 local A NAVIDROME CONTAINER WAS CREATED (= container .Created, 19:12:20Z) + 22:19:33 local ...and STARTED again by Docker's `restart: unless-stopped` at the B5 power cut + 22:27:20 local the teardown found it: deployed=False, state=running + +WHAT IS AND IS NOT ESTABLISHED +------------------------------ +ESTABLISHED: the app's RECORD is gone (`app.yaml` absent, `deployed=false`), its VOLUME is gone, +and a container carrying `com.docker.compose.project=navidrome` was created two seconds after the +removal returned 200 and has run ever since. The controller probes it and calls it healthy +(`Health probe navidrome: API GET :4533/ping -> 200`). + +NOT ESTABLISHED: WHAT created it. The controller was restarted several times later in the night +(knob changes and two power cuts) and its log no longer reaches 19:12:20Z. This is stated as an +observation, not a diagnosis — and it is the second time tonight that a restart destroyed the +evidence of the thing that mattered (see R-621). diff --git a/documentation/audits/update-night-2026-09-21/27-final-cleanup.txt b/documentation/audits/update-night-2026-09-21/27-final-cleanup.txt new file mode 100644 index 00000000..363f0b28 --- /dev/null +++ b/documentation/audits/update-night-2026-09-21/27-final-cleanup.txt @@ -0,0 +1,27 @@ +=== removing the resurrected navidrome BY NAME (it is invisible to the product: deployed=false) +stopped +removed +volume removed: navidrome_navidrome_data + +=== containers now — expect ONLY the three protected infra + the controller: +felhom-controller Up 3 minutes (healthy) gitea.dooplex.hu/admin/felhom-controller:0.261.0 +filebrowser Up 10 minutes (healthy) gtstef/filebrowser:1.3.3-stable +traefik Up 10 minutes traefik:v3.6.7 + +=== any drill image left? (must be none) +(none) +=== any drill volume left? (must be none) +(none) + +=== the harness's own directories on the scratch drive, removed by name: +removed /mnt/felhom-drives/scratch_hdd/userdata/audiobookshelf +removed /mnt/felhom-drives/scratch_hdd/userdata/navidrome +removed /mnt/felhom-drives/scratch_hdd/userdata/nextcloud +removed /mnt/felhom-drives/scratch_hdd/userdata/romm +documents +downloads +media +roms + +=== disk: +/dev/loop1 69G 33G 33G 50% /var/lib/felhom diff --git a/documentation/audits/update-night-2026-09-21/NEW-ROWS.md b/documentation/audits/update-night-2026-09-21/NEW-ROWS.md index a8c2d99b..6e93aa9e 100644 --- a/documentation/audits/update-night-2026-09-21/NEW-ROWS.md +++ b/documentation/audits/update-night-2026-09-21/NEW-ROWS.md @@ -10,7 +10,7 @@ Highest existing id at the start of the night: **R-614** (304 rows). | **R-617** | **[P3-LOW] The Gitea API token this project uses for pushes cannot create a repository through the documented endpoint, but CAN through `repos/migrate` — so "the token cannot do it" was nearly recorded as a fact when the truth was "one endpoint refuses it".** FOUND 2026-09-21 creating the drill catalog. Both `~/.git-credentials` tokens carry `write:misc,write:notification,write:package,write:issue,write:repository`; `POST /api/v1/user/repos` requires `write:user` and answers **403**, and `POST /api/v1/admin/users//repos` requires `write:admin` and answers 403 too. `POST /api/v1/repos/migrate` with the same token answered **201** and created the private repository. **Why this is a row and not a note:** a session that stopped at the first 403 would have recorded "CC cannot create a Gitea repository" — an unfalsifiable capability claim of exactly the shape the workspace's standing rule 2 forbids — and every later drill would have been designed around a limit that does not exist. **What it needs:** one line in the operations notes saying which endpoint to use, and (optional, operator) a token scoped for the job so the migrate route is not load-bearing. Evidence: `audits/update-night-2026-09-21/03-drill-repo.txt`. | **READY — rank P3-LOW; owner: CC (docs)** | -| **R-618** | **[P1-HIGH] TWO apps are presented to the household as UNHEALTHY while they are working perfectly — and because the guarded Update waits on that same probe, a SUCCESSFUL update ends by STOPPING the working app and sending the household to a restore they do not need.** **RANK RAISED FROM P2 TO P1 BY A LIVE MEASUREMENT taken the same night, and the escalation is the whole point:** tandoor's Update 2.6.13 → 2.6.15 was pressed at 21:16:47 and entered `verifying` at 21:17:46. At **21:18:28** the NEW version was `Up 25 seconds` and answering **HTTP 200** on `/accounts/login/` through the household's own front door — while the controller, probing port 8080 where nothing listens, could not see it. `verifying` therefore cannot pass, the full `update.health_timeout` is spent, `Manager.failAndHold` runs `compose down`, and the app is STOPPED. **Nothing is lost** — the data is in the volumes and the restore works — **but one wrong port number in a template converts every successful update of that app into an outage plus an unnecessary restore, for every household running it.** Evidence: `audits/update-night-2026-09-21/14-tandoor-serving-while-verifying.txt`. MEASURED 2026-09-21 on guest 9202 (controller v0.261.0, catalog `f5f6a152b513`). Two shapes, one class: **(a) `tandoor` — the WRONG PORT.** `.felhom.yml` probes `port: 8080`; the container listens on **80 and nothing else** (`ss -ltn` inside it), the compose's own traefik label routes to 80, its own docker healthcheck reads `healthy`, and `/accounts/login/` answers **200** through the household's real front door. `GET /api/stacks/tandoor` nevertheless reads `state: "unhealthy"`. **(b) `zipline` — the WRONG PATH.** `.felhom.yml` probes `/api/health`, which zipline 4.6.1 answers **404 `Route GET:/api/health not found`**; **the compose healthcheck in the very same file uses `/api/healthcheck` and is correct and green.** `/dashboard` answers 200. The controller reads `unhealthy`. **This is the MIRROR of R-613** — that is a probe that passes on a broken app (a false GREEN, which no alarm catches); this is a probe that fails on a working app (a false RED). **IT DOES NOT ALARM, AND THAT SETS THE RANK:** `08-alarm-ladder.md` §4 puts `unhealthy` deliberately in the NOT-down set, so no dead-app event and no customer mail follows — the damage is what the household READS, plus anything that gates on `state`. **IT ALREADY COST A MEASUREMENT TONIGHT:** this drill's harness waited for `state == "running"` and hung for its full budget on tandoor, an app that was up the whole time. An instrument waiting for a wrong answer looks exactly like a slow app. **THE GATE THIS WANTS IS CHEAP AND STATIC, AND THAT IS THE FINDING'S REAL VALUE.** Both halves of the answer live in the same template: compare the `.felhom.yml` probe's port and path against the compose's **own** `healthcheck: test:` URL. A sweep of all 53 templates on that rule was run tonight and returns **five** disagreements: `tandoor` (PORT — **CONFIRMED live**), `zipline` (PATH — **CONFIRMED live**), `wger` (PORT, probe 80 vs compose 8000 — **SUSPECTED, NOT MEASURED**, it was not deployed), `home-assistant` (PATH, `/api/` vs `/manifest.json` — **NOT MEASURED**), and `adventurelog` (a FALSE POSITIVE of the sweep's own regex — it reads `running` live). **So the rule finds both real defects, with two candidates and one false positive out of 53** — a good enough signal for a fast gate, provided it reports candidates rather than convictions and a person or a runtime check resolves them. The earlier, cruder rule (probe port vs the *traefik* port) is strictly worse: it clears zipline and convicts adventurelog. **AND A SECOND FIX SHAPE, ON THE CONTROLLER SIDE, WORTH CONSIDERING BESIDE THE CATALOG ONE:** in both confirmed cases the container's OWN docker healthcheck was **green** the whole time. A `verifying` phase that is about to stop a working app could ask that too — if the compose declares a healthcheck and docker reports `healthy`, the app is alive whatever our probe thinks. That does not excuse a wrong probe, but it turns this failure direction from an outage into a wrong label. It is a design question, not a defect, and is raised here rather than decided. **Needs:** fix tandoor's port and zipline's path; measure wger; add the static gate with a decoy each way (R-421) — a template whose probe agrees must not read as a disagreement, and vice versa. **THE GATE'S RULE WAS THEN SHARPENED BY READING `healthprobe.go` RATHER THAN ASSUMING IT, and the sharpening REMOVED a false conviction.** `type: http` treats **any** response as healthy (`healthprobe.go:258-261`), and `type: api` with **no** `expect` block does the same (`:265-268`); only `type: api` WITH `expect.status` cares about the path or the code. So a PATH difference is a candidate only for the third shape, while a PORT difference is a candidate for all of them. Under that rule the 53-template sweep returns **four** candidates — `tandoor` and `zipline` (both CONFIRMED live), `wger` (suspected, unmeasured), and `adventurelog` (a false positive: its compose lists two containers' ports and the probe targets the backend; measured `running`). **`home-assistant` is correctly CLEARED by the sharpened rule** — `type: api`, no `expect`, so its `/api/` answering 401 without a token is healthy, and its edge was PROVEN on the box tonight. The crude rule convicted it; the rule read from the code does not. **That is the gate to build: two of 53 convicted, one suspected, one false positive, and the false positive is resolvable by one live check.** Evidence: `audits/update-night-2026-09-21/10-probe-port-sweep.txt`, `12-probe-vs-compose-healthcheck.txt` and `13-probe-sweep-sharpened.txt`. | **READY — rank P2-MEDIUM; owner: CC (catalog)** | +| **R-618** | **[P1-HIGH] THREE apps are presented to the household as UNHEALTHY while they are working perfectly — and because the guarded Update waits on that same probe, a SUCCESSFUL update ends by STOPPING the working app and sending the household to a restore they do not need.** **RANK RAISED FROM P2 TO P1 BY A LIVE MEASUREMENT taken the same night, and the escalation is the whole point:** tandoor's Update 2.6.13 → 2.6.15 was pressed at 21:16:47 and entered `verifying` at 21:17:46. At **21:18:28** the NEW version was `Up 25 seconds` and answering **HTTP 200** on `/accounts/login/` through the household's own front door — while the controller, probing port 8080 where nothing listens, could not see it. `verifying` therefore cannot pass, the full `update.health_timeout` is spent, `Manager.failAndHold` runs `compose down`, and the app is STOPPED. **Nothing is lost** — the data is in the volumes and the restore works — **but one wrong port number in a template converts every successful update of that app into an outage plus an unnecessary restore, for every household running it.** Evidence: `audits/update-night-2026-09-21/14-tandoor-serving-while-verifying.txt`. MEASURED 2026-09-21 on guest 9202 (controller v0.261.0, catalog `f5f6a152b513`). Two shapes, one class: **(a) `tandoor` — the WRONG PORT.** `.felhom.yml` probes `port: 8080`; the container listens on **80 and nothing else** (`ss -ltn` inside it), the compose's own traefik label routes to 80, its own docker healthcheck reads `healthy`, and `/accounts/login/` answers **200** through the household's real front door. `GET /api/stacks/tandoor` nevertheless reads `state: "unhealthy"`. **(b) `zipline` — the WRONG PATH.** `.felhom.yml` probes `/api/health`, which zipline 4.6.1 answers **404 `Route GET:/api/health not found`**; **the compose healthcheck in the very same file uses `/api/healthcheck` and is correct and green.** `/dashboard` answers 200. The controller reads `unhealthy`. **This is the MIRROR of R-613** — that is a probe that passes on a broken app (a false GREEN, which no alarm catches); this is a probe that fails on a working app (a false RED). **IT DOES NOT ALARM, AND THAT SETS THE RANK:** `08-alarm-ladder.md` §4 puts `unhealthy` deliberately in the NOT-down set, so no dead-app event and no customer mail follows — the damage is what the household READS, plus anything that gates on `state`. **IT ALREADY COST A MEASUREMENT TONIGHT:** this drill's harness waited for `state == "running"` and hung for its full budget on tandoor, an app that was up the whole time. An instrument waiting for a wrong answer looks exactly like a slow app. **THE GATE THIS WANTS IS CHEAP AND STATIC, AND THAT IS THE FINDING'S REAL VALUE.** Both halves of the answer live in the same template: compare the `.felhom.yml` probe's port and path against the compose's **own** `healthcheck: test:` URL. A sweep of all 53 templates on that rule was run tonight and returns **five** disagreements: `tandoor` (PORT — **CONFIRMED live**), `zipline` (PATH — **CONFIRMED live**), `wger` (PORT, probe 80 vs compose 8000 — **CONFIRMED live the same night**), `home-assistant` (PATH, `/api/` vs `/manifest.json` — **NOT MEASURED**), and `adventurelog` (a FALSE POSITIVE of the sweep's own regex — it reads `running` live). **So the rule finds both real defects, with two candidates and one false positive out of 53** — a good enough signal for a fast gate, provided it reports candidates rather than convictions and a person or a runtime check resolves them. The earlier, cruder rule (probe port vs the *traefik* port) is strictly worse: it clears zipline and convicts adventurelog. **AND A SECOND FIX SHAPE, ON THE CONTROLLER SIDE, WORTH CONSIDERING BESIDE THE CATALOG ONE:** in both confirmed cases the container's OWN docker healthcheck was **green** the whole time. A `verifying` phase that is about to stop a working app could ask that too — if the compose declares a healthcheck and docker reports `healthy`, the app is alive whatever our probe thinks. That does not excuse a wrong probe, but it turns this failure direction from an outage into a wrong label. It is a design question, not a defect, and is raised here rather than decided. **Needs:** fix tandoor's port (80), zipline's path (`/api/healthcheck`) and wger's port (8000) — **all three are now CONFIRMED live, none is a guess**; add the static gate with a decoy each way (R-421) — a template whose probe agrees must not read as a disagreement, and vice versa. **THE GATE'S RULE WAS THEN SHARPENED BY READING `healthprobe.go` RATHER THAN ASSUMING IT, and the sharpening REMOVED a false conviction.** `type: http` treats **any** response as healthy (`healthprobe.go:258-261`), and `type: api` with **no** `expect` block does the same (`:265-268`); only `type: api` WITH `expect.status` cares about the path or the code. So a PATH difference is a candidate only for the third shape, while a PORT difference is a candidate for all of them. Under that rule the 53-template sweep returns **four** candidates — `tandoor`, `zipline` and **`wger` (all three CONFIRMED live — probe `type: http, port: 80`; inside the container port 80 is `refused` and port 8000 `ANSWERED`; docker's own healthcheck green; front door 302; the box reads `unhealthy`)**, and `adventurelog` (a false positive: its compose lists two containers' ports and the probe targets the backend; measured `running`). **`home-assistant` is correctly CLEARED by the sharpened rule** — `type: api`, no `expect`, so its `/api/` answering 401 without a token is healthy, and its edge was PROVEN on the box tonight. The crude rule convicted it; the rule read from the code does not. **That is the gate to build: two of 53 convicted, one suspected, one false positive, and the false positive is resolvable by one live check.** Evidence: `audits/update-night-2026-09-21/10-probe-port-sweep.txt`, `12-probe-vs-compose-healthcheck.txt` and `13-probe-sweep-sharpened.txt`. | **READY — rank P2-MEDIUM; owner: CC (catalog)** | | **R-619** | **[P3-LOW] A `type: password` deploy field is MANDATORY however `required` reads, and the `deploy-fields` contract says the opposite — so any caller that trusts it is refused.** MEASURED 2026-09-21 on guest 9202 while widening the update drill. `GET /api/stacks/grafana/deploy-fields` serves `{"env_var":"GF_SECURITY_ADMIN_PASSWORD","type":"password","generate":"password:16","required":false}`; a deploy carrying only the two `required:true` fields is refused **400** „a(z) „Admin jelszó" mező kitöltése kötelező — használja a Generálás gombot…". **The BEHAVIOUR is right and is a decision, not a bug:** `deploy.go:305-312` refuses a `password` field with no caller value on purpose — *"We never silently auto-generate — the user needs to know their password"* — which is the opposite of the `secret` case one branch above, where a generated value the customer never sees is exactly correct. **The defect is the CONTRACT.** `.felhom.yml` declares `required: false`, the API serves that verbatim, and nothing on the wire distinguishes "optional because the box will generate it" (`secret`) from "optional in the template and mandatory in the code" (`password`). A person using the deploy page never meets this because the page renders a Generálás button; **anything that is not that page does**, which now includes this drill harness and would include `09` §6.2's unattended caller the day it deploys anything. **Fix shape (smallest that keeps the decision):** serve `required: true` for `type: password` in the deploy-fields response — one place, derived rather than stored, so templates need no edit — and a test asserting a `password` field always reaches the wire as required. Alternatively state it in the field's `description`, which is weaker because it is prose. Evidence: `audits/update-night-2026-09-21/apps/grafana/log.txt` (the refusal) and `batchA.log`. | **READY — rank P3-LOW; owner: CC (controller)** | @@ -24,7 +24,7 @@ Highest existing id at the start of the night: **R-614** (304 rows). ## Lines to ADD to existing rows (not re-filed) -**R-607** — append: **Seen again 2026-09-21 (update night), a dozen times in one session, and for the first time with a USER-VISIBLE consequence rather than a measurement one.** On `mealie` the bump was pushed, `POST /api/sync` AND `POST /api/stacks/rescan` were both run, and the badge still read the up-to-date one (HU „Naprakesz", EN "Up to date") — the catalog had not reached `catalog_images` yet. The guarded Update was then pressed and **reported „Frissitve" after 2.1 seconds having moved nothing at all**: pinned, installed, the live compose line and `docker inspect` all still read `v3.20.1`. That is honest given a stale cache — the pin is written from *the catalog's current definition*, which was still the old one — but what the household sees is a button that says it updated them and did not. **A NUMBER, at last, which is what this row asks for:** the night's harness was changed to poll `catalog_images` until the pushed reference appears and to report how long that took; those figures are each edge's `badge_catchup_seconds`. Every drill-catalog bump of the night was followed by `POST /api/sync` answering „Sablonok naprakészek — nincs változás" while the box's cache HAD moved, with `catalog_images` staying stale until a separate `POST /api/stacks/rescan`. The night's harness therefore rescans unconditionally after every sync, which is a workaround and not a fix. **The window was still never measured as a NUMBER** — that is what the row asks for and what remains owed. +**R-607** — append: **Seen again 2026-09-21 (update night), a dozen times in one session, and for the first time with a USER-VISIBLE consequence rather than a measurement one.** On `mealie` the bump was pushed, `POST /api/sync` AND `POST /api/stacks/rescan` were both run, and the badge still read the up-to-date one (HU „Naprakesz", EN "Up to date") — the catalog had not reached `catalog_images` yet. The guarded Update was then pressed and **reported „Frissitve" after 2.1 seconds having moved nothing at all**: pinned, installed, the live compose line and `docker inspect` all still read `v3.20.1`. That is honest given a stale cache — the pin is written from *the catalog's current definition*, which was still the old one — but what the household sees is a button that says it updated them and did not. **A NUMBER, at last, which is what this row asks for:** the night's harness was changed to poll `catalog_images` until the pushed reference appears and to report how long that took; those figures are each edge's `badge_catchup_seconds`, and here they are: **4.4 s, 4.4 s, 4.5 s, 4.5 s — and 29.0 s.** The four fast ones are one sync+rescan round; the 29-second one (`nextcloud`, an engine-sidecar bump) needed **additional** sync+rescan rounds before `catalog_images` carried the pushed reference. **So the window is not a fixed scan interval — it varies by roughly 7x between edges on the same box in the same hour**, which is why a caller (or a household) cannot know when the badge is safe to read. Before tonight this row had no number at all; it now has five, and they disagree with each other, which is itself the most useful thing about them. Every drill-catalog bump of the night was followed by `POST /api/sync` answering „Sablonok naprakészek — nincs változás" while the box's cache HAD moved, with `catalog_images` staying stale until a separate `POST /api/stacks/rescan`. The night's harness therefore rescans unconditionally after every sync, which is a workaround and not a fix. **The window was still never measured as a NUMBER** — that is what the row asks for and what remains owed. **R-462** — append: **The count moved on 2026-09-21 from 3 apps to .** The update night walked real, within-a-major upstream edges on guest 9202 through the product's own guarded Update, each seeded and read back through the app's own front door: .

proven, failed, inconclusive. Box-side fixtures for apps now exist at `audits/update-night-2026-09-21/fixtures.py`, and four of them (actualbudget, navidrome, audiobookshelf, vikunja) are ported into `app-catalog-felhom.eu/scripts/upgrade_fixtures.py` with seven new `EDGES` (U1–U7) so the same edges can be run on the harness venue **with their ABORT step**, which the box deliberately does not offer. **Owed:** the harness RUNS for those edges (the code is in; the runs are not), and fixtures for the apps recorded inconclusive tonight. @@ -55,3 +55,5 @@ Highest existing id at the start of the night: **R-614** (304 rows). | **R-625** | **[P2-MEDIUM] A HELD app keeps inviting the household to update it, and the button then refuses — the exact inconsistency R-524 removed for the other case, still present for this one.** MEASURED 2026-09-21 (update night, leg B6). `glance` was HELD by a genuine unattended failed update. The drill catalog then published a **fixed newer version** — a real forward route, the thing a household would hope for. Afterwards the app page read **„Frissítés elérhető — ma" / „Update available — today"**, in both languages, with the Update button offered; pressing it answered **`409 reason='held'`** and the hold sentence. **The BEHAVIOUR is correct and is DESIGN, not a defect:** `09` §6.1 says the hold is `settings.RestoreHold` and that **a successful unit restore lifts an update hold (only that kind)** — a newer catalog version does not, and should not, because nobody has checked that the new version can start on data the failed one may have touched. **The DEFECT is that the page says otherwise.** R-524 settled precisely this shape for the Ahead case — *"(a) is free and offers a household a downgrade … (c) show „Naprakész" and refuse the button … Why (c): the direction was already settled"* — and chose to make the badge and the button agree. The HELD case still has them disagreeing, in the more painful direction: the badge invites, the button refuses, and the refusal is the same long sentence the household has already read. **What it needs:** the badge for a held app should say what is true — that the app is held and the way back is the restore — and the Update button should not be offered while `RestoreHold` stands. One verdict, read by both surfaces, exactly as R-524 did it. **AND A QUESTION FOR `09` §3b Q4 THAT THIS MEASUREMENT RAISES AND DOES NOT ANSWER:** the household's ONLY route out is a restore, even when the catalog has already shipped a fix. That is defensible, but it is now measured rather than assumed, and Q4's "does the box try again?" should be read next to it. Evidence: `audits/update-night-2026-09-21/bad-days/B6-way-out-forwards/result.json`. | **READY — rank P2-MEDIUM; owner: CC (controller)** | **R-446** — append: **MEASURED ON A BOX 2026-09-21 (update night, leg B8), and it REFINES the row in two ways rather than merely confirming it.** §8.1's numbers came from a registry sweep on DooPlex; this is the same question asked of a customer-shaped box, where the badge actually renders. On guest 9202, `docmost`'s two floating pins were read as `installed_images` records them and compared against the upstream digests measured the same night: `postgres:16-alpine` → **`sha256:721873c34ceb9…` on the box and `sha256:721873c34ceb9…` upstream**, and `redis:7-alpine` → **`sha256:858f009f9709c…` both sides**. **Identical. So the badge „Naprakész" is TRUE for this box**, and the app reads correctly. **(1) The defect's size is set by INSTALL AGE, not by the catalog.** A floating pin is wrong only for a box that pulled BEFORE the tag moved; a box deployed after the repush holds the current image and its badge is right. R-446's "six repushed pins" measured the tag against the date the CATALOG set it, which is the right measure for *the catalog* and not for *a box*. **(2) The producer Q6 needs ALREADY EXISTS on the box.** `installed_images` records a real `digest` per service (`installed.go` §7.1) — the box knows exactly what it is running. What it cannot do is COMPARE, because the catalog carries no digest to compare against. That is Q6's proposal, and this is a concrete confirmation that only the catalog half is missing. Evidence: `audits/update-night-2026-09-21/23-B8-floating-pin.txt`. + +| **R-626** | **[P2-MEDIUM] An app the customer REMOVED came back: the removal returned 200 and deleted the record and the volume, a container was created two seconds later, and Docker's restart policy has kept it running ever since — while the box reports the app as not installed.** FOUND 2026-09-21 during the update night's TEARDOWN, which is the only reason it was found at all. `navidrome` was removed through the product: the `remove_hdd_data:true` call was correctly refused `409` (R-442's fail-closed guard — the drive path could not be resolved on this guest), and the `remove_hdd_data:false` call returned **200** with `volumes_removed: ['navidrome_navidrome_data']`. **Two seconds later a container carrying `com.docker.compose.project=navidrome` was CREATED** (`.Created = 19:12:20Z`), and Docker's `restart: unless-stopped` started it again at the next guest boot (`.StartedAt = 20:19:33Z`, the B5 power cut). Seven hours later: **`app.yaml` absent, `deployed=false`, one volume back, and the controller happily probing it — `Health probe navidrome: API GET :4533/ping → 200`.** **The customer-visible shape is the bad one:** *"I deleted that app and it came back"* — and it came back **blank**, because the volume really was deleted, so it looks installed and is empty. It is also invisible to every sweep that keys on `deployed`, which is exactly why the teardown found it and nothing else did. **WHAT IS NOT ESTABLISHED, and is stated rather than guessed: what created the container.** The controller was restarted several times later in the night and its log no longer reaches that moment — **the second time in one night that a restart destroyed the evidence of the thing that mattered** (see R-621). **Needs:** reproduce with a loop that removes an app and watches `docker events` for 60 s, so the creating path is a NAME and not an inference; then a test that removes an app, reboots, and asserts no container with that compose project exists. **And one instrument lesson worth keeping:** this session's own post-remove check queried the compose-project label and reported clean at 21:12:18 — two seconds before the container appeared. A check that runs once, immediately, cannot see a thing that is created immediately after it. Evidence: `audits/update-night-2026-09-21/26-removed-app-came-back.txt`. | **READY — rank P2-MEDIUM; owner: CC (controller)** | diff --git a/documentation/audits/update-night-2026-09-21/PROGRESS.md b/documentation/audits/update-night-2026-09-21/PROGRESS.md index 42e25e01..55252fd4 100644 --- a/documentation/audits/update-night-2026-09-21/PROGRESS.md +++ b/documentation/audits/update-night-2026-09-21/PROGRESS.md @@ -49,3 +49,9 @@ A resuming session reads THIS FILE FIRST and never repeats a finished step. | 22:08 | **B5 (safety-dump cut) — MISSED, recorded as a miss** | The phases went `backing-up` -> `pulling` -> `failed` in **0.473 s** and `safety-dump` was never observed, so the plug was never pulled. Recorded as a MISS, not as a pass. To be retried with a genuine pending edge | bad-days/B5-safety-dump/ | | 22:09 | **Phase 4 — the morning after** | Every one of the **10 badges is TRUE** (refs equal <-> „Naprakész", refs differ <-> „Frissítés elérhető"). Q4's four promises all scored True on the held app's page. `zipline` shows the household „Nem egészséges — URL nem elérhető" while running — R-618 in the household's own words | bad-days/P4-morning-after/ | | 22:09 | **B6 — the way out FORWARDS: REFUSED** | A held app met a FIXED newer version. Badge: „Frissítés elérhető — ma" in both languages, button offered; press -> **`409 reason='held'`**. Correct per §6.1 (only a restore lifts a hold) but **the page invites what the button refuses** — the exact inconsistency R-524 removed for the Ahead case. Filed as **R-625** | bad-days/B6-way-out-forwards/ | +| 22:16 | **PHASE 2.3 — the PostgreSQL conversion REHEARSAL, costed (Q5)** | **WORKED end to end.** 49 MB / 48 tables: dump with 16 **2.6 s / 132 201 B**; fresh 17 + replay **6.5 s / 48 tables**; the app said „Database connection successful"; **the seeded account read back on 17**; total **155.9 s**, of which ~9 s is engine work. Two benign ERROR lines named. `pg_upgrade` NOT run — it needs an image that does not exist here | 24-Q5-postgres-conversion-costed.md | +| 22:18 | **wger CONFIRMED — R-618's last candidate closes** | probe `type: http port: 80`; inside the container **port 80 refused, port 8000 ANSWERED**; docker health green; front door 302; the box says `unhealthy`. **Three confirmed instances now: tandoor, zipline, wger** — and the cheap static rule finds all three with one false positive out of 53 | 22-wger-probe-measured.txt | +| 22:19 | **B5 (safety-dump cut) — the second early phase, RETRIED and HIT** | Cut at `safety-dump` +0.023 s. After boot: **no recovery line and no journal** (unlike the `backing-up` cut, which produced both), **the pin did NOT move**, the app runs, a fresh paste seeded and read back, and the card shows **no interrupted sentence at all**. No half-written backup artefact on either cut | 25-B5-the-two-early-cuts.md | +| 22:20 | **Bonus proof across a REAL power cut** | The boot sweep met the held `glance` after an unclean shutdown and **deliberately left it alone**: „is a boot orphan by intent but is HELD … NOT starting it; whatever is holding it owns its recovery". §6.1's three-unattended-paths claim, proven across a power cut | 25-B5-the-two-early-cuts.md | +| 22:22 | **mealie v3.20.1 -> v3.27.0** (db-postgres) | **PROVEN** — the edge that failed twice on instrument problems, walked cleanly on the third | apps/mealie/ | +| 22:23 | **PHASE 1+2 COMPLETE** | **21 edges attempted: 14 proven, 3 failed, 4 inconclusive.** Ten of the fourteen printed a verbatim migration line. Up from the **three** apps this project had ever measured | summarise.py | diff --git a/documentation/audits/update-night-2026-09-21/apps/mealie/app-logs-after.txt b/documentation/audits/update-night-2026-09-21/apps/mealie/app-logs-after.txt index 14920487..d3db8f95 100644 --- a/documentation/audits/update-night-2026-09-21/apps/mealie/app-logs-after.txt +++ b/documentation/audits/update-night-2026-09-21/apps/mealie/app-logs-after.txt @@ -3,22 +3,33 @@ mealie | mealie | User uid: 1000 mealie | User gid: 1000 mealie | -mealie | INFO 2026-09-21T21:24:18 - setup plugin alembic.autogenerate.schemas -mealie | INFO 2026-09-21T21:24:18 - setup plugin alembic.autogenerate.tables -mealie | INFO 2026-09-21T21:24:18 - setup plugin alembic.autogenerate.types -mealie | INFO 2026-09-21T21:24:18 - setup plugin alembic.autogenerate.constraints -mealie | INFO 2026-09-21T21:24:18 - setup plugin alembic.autogenerate.defaults -mealie | INFO 2026-09-21T21:24:18 - setup plugin alembic.autogenerate.comments -mealie | INFO 2026-09-21T21:24:24 - Started server process [1] -mealie | INFO 2026-09-21T21:24:24 - Waiting for application startup. -mealie | INFO 2026-09-21T21:24:24 - start: database initialization -mealie | INFO 2026-09-21T21:24:24 - Database connection established. -mealie | INFO 2026-09-21T21:24:24 - Context impl SQLiteImpl. -mealie | INFO 2026-09-21T21:24:24 - Will assume non-transactional DDL. -mealie | INFO 2026-09-21T21:24:25 - end: database initialization -mealie | INFO 2026-09-21T21:24:25 - -----SYSTEM STARTUP----- -mealie | INFO 2026-09-21T21:24:25 - ------APP SETTINGS------ -mealie | INFO 2026-09-21T21:24:25 - { +mealie | INFO 2026-09-21T22:22:04 - Started server process [1] +mealie | INFO 2026-09-21T22:22:04 - Waiting for application startup. +mealie | INFO 2026-09-21T22:22:04 - start: database initialization +mealie | INFO 2026-09-21T22:22:04 - setup plugin alembic.autogenerate.schemas +mealie | INFO 2026-09-21T22:22:04 - setup plugin alembic.autogenerate.tables +mealie | INFO 2026-09-21T22:22:04 - setup plugin alembic.autogenerate.types +mealie | INFO 2026-09-21T22:22:04 - setup plugin alembic.autogenerate.constraints +mealie | INFO 2026-09-21T22:22:04 - setup plugin alembic.autogenerate.defaults +mealie | INFO 2026-09-21T22:22:04 - setup plugin alembic.autogenerate.comments +mealie | INFO 2026-09-21T22:22:04 - setup plugin alembic.ext.checkconstraint_byname +mealie | INFO 2026-09-21T22:22:04 - Database connection established. +mealie | INFO 2026-09-21T22:22:04 - Context impl SQLiteImpl. +mealie | INFO 2026-09-21T22:22:04 - Will assume non-transactional DDL. +mealie | INFO 2026-09-21T22:22:05 - Migration needed. Performing migration... +mealie | INFO 2026-09-21T22:22:05 - Context impl SQLiteImpl. +mealie | INFO 2026-09-21T22:22:05 - Will assume non-transactional DDL. +mealie | INFO 2026-09-21T22:22:05 - Running upgrade 2187537c52b8 -> 69e942bab3aa, add tokens valid after column to users +mealie | INFO 2026-09-21T22:22:05 - Running upgrade 69e942bab3aa -> b3f1c9a27d84, add external avatar hash to users +mealie | INFO 2026-09-21T22:22:05 - Running upgrade b3f1c9a27d84 -> f2191b69db2e, add ingredient substitutions +mealie | INFO 2026-09-21T22:22:05 - Running upgrade f2191b69db2e -> 4b91d3a7c0e2, backfill recipe image column from disk +mealie | INFO 2026-09-21T22:22:05 - Recipe image backfill checked 1 recipes: 0 image references restored, 0 cleared +mealie | INFO 2026-09-21T22:22:05 - Running upgrade 4b91d3a7c0e2 -> 3527efeeec34, 'add recipe_note_ref_link' +mealie | INFO 2026-09-21T22:22:05 - Checking for migration data fixes +mealie | INFO 2026-09-21T22:22:05 - end: database initialization +mealie | INFO 2026-09-21T22:22:05 - -----SYSTEM STARTUP----- +mealie | INFO 2026-09-21T22:22:05 - ------APP SETTINGS------ +mealie | INFO 2026-09-21T22:22:05 - { mealie | "TESTING": false, mealie | "PRODUCTION": true, mealie | "LOG_CONFIG_OVERRIDE": null, @@ -47,10 +58,12 @@ mealie | "API_HOST": "0.0.0.0", mealie | "API_PORT": 9000, mealie | "API_DOCS": true, mealie | "TOKEN_TIME": 48, -mealie | "GIT_COMMIT_HASH": "a562409e43964d3734eabd64e4a2fd64bffaab4d", +mealie | "GIT_COMMIT_HASH": "dedc6cc75f2bd8e89108ad998970e7bdf6d4444f", mealie | "ALLOW_SIGNUP": false, mealie | "ALLOW_PASSWORD_LOGIN": true, mealie | "ALLOWED_IFRAME_HOSTS": "", +mealie | "HTTP_ALLOW_LIST": "", +mealie | "HTTP_DISALLOW_LIST": "", mealie | "DAILY_SCHEDULE_TIME": "23:45", mealie | "SECURITY_MAX_LOGIN_ATTEMPTS": 5, mealie | "SECURITY_USER_LOCKOUT_TIME": 24, @@ -83,6 +96,7 @@ mealie | "OIDC_CLIENT_ID": null, mealie | "OIDC_CLIENT_SECRET": null, mealie | "OIDC_CONFIGURATION_URL": null, mealie | "OIDC_SIGNUP_ENABLED": true, +mealie | "OIDC_REQUIRES_EMAIL_VERIFICATION": true, mealie | "OIDC_USER_GROUP": null, mealie | "OIDC_ADMIN_GROUP": null, mealie | "OIDC_AUTO_REDIRECT": false, @@ -95,27 +109,32 @@ mealie | "OIDC_SCOPES_OVERRIDE": null, mealie | "OIDC_TLS_CACERTFILE": null, mealie | "OIDC_CLIENT_TIMEOUT": "default", mealie | "OPENAI_CUSTOM_PROMPT_DIR": null, +mealie | "SCRAPER_PROXY_URL": null, +mealie | "SCRAPER_PROXY_MODE": "always", +mealie | "SCRAPER_FLARESOLVERR_URL": null, +mealie | "SCRAPER_FLARESOLVERR_TIMEOUT": 60, mealie | "WORKER_PER_CORE": 1, mealie | "UVICORN_WORKERS": 1, mealie | "TLS_CERTIFICATE_PATH": null, -mealie | "TLS_PRIVATE_KEY_PATH": null +mealie | "TLS_PRIVATE_KEY_PATH": null, +mealie | "YTDLP_COOKIEFILE": null mealie | } -mealie | INFO 2026-09-21T21:24:25 - ------APP FEATURES------ -mealie | INFO 2026-09-21T21:24:25 - --------==SMTP==-------- -mealie | INFO 2026-09-21T21:24:25 - Enabled: False -mealie | INFO 2026-09-21T21:24:25 - --------==LDAP==-------- -mealie | INFO 2026-09-21T21:24:25 - Enabled: False +mealie | INFO 2026-09-21T22:22:05 - ------APP FEATURES------ +mealie | INFO 2026-09-21T22:22:05 - --------==SMTP==-------- +mealie | INFO 2026-09-21T22:22:05 - Enabled: False +mealie | INFO 2026-09-21T22:22:05 - --------==LDAP==-------- +mealie | INFO 2026-09-21T22:22:05 - Enabled: False mealie | Reason: LDAP_AUTH_ENABLED is false -mealie | INFO 2026-09-21T21:24:25 - --------==OIDC==-------- -mealie | INFO 2026-09-21T21:24:25 - Enabled: False +mealie | INFO 2026-09-21T22:22:05 - --------==OIDC==-------- +mealie | INFO 2026-09-21T22:22:05 - Enabled: False mealie | Reason: OIDC_AUTH_ENABLED is false -mealie | INFO 2026-09-21T21:24:25 - ------------------------ -mealie | INFO 2026-09-21T21:24:25 - Daily tasks scheduled for 2026-09-21 21:45:00+00:00 -mealie | INFO 2026-09-21T21:24:25 - Application startup complete. -mealie | INFO 2026-09-21T21:24:25 - Uvicorn running on http://0.0.0.0:9000 (Press CTRL+C to quit) -mealie | INFO 2026-09-21T21:25:08 - [192.168.0.180:0] 200 OK "GET /api/app/about HTTP/1.1" -mealie | INFO 2026-09-21T21:25:08 - [192.168.0.180:0] 200 OK "POST /api/auth/token HTTP/1.1" -mealie | ERROR 2026-09-21T21:25:08 - No Entry Found on recipe controller action -mealie | ERROR 2026-09-21T21:25:08 - No Entry Found on recipe controller action -mealie | INFO 2026-09-21T21:25:08 - [192.168.0.180:0] 404 Not Found "GET /api/recipes/nopefec546301e HTTP/1.1" -mealie | INFO 2026-09-21T21:25:08 - [192.168.0.180:0] 200 OK "GET /api/recipes/drill-087e4023c5 HTTP/1.1" +mealie | INFO 2026-09-21T22:22:05 - ------------------------ +mealie | INFO 2026-09-21T22:22:05 - Daily tasks scheduled for 2026-09-21 21:45:00+00:00 +mealie | INFO 2026-09-21T22:22:05 - Application startup complete. +mealie | INFO 2026-09-21T22:22:05 - Uvicorn running on http://0.0.0.0:9000 (Press CTRL+C to quit) +mealie | INFO 2026-09-21T22:22:13 - [192.168.0.180:0] 200 OK "GET /api/app/about HTTP/1.1" +mealie | INFO 2026-09-21T22:22:14 - [192.168.0.180:0] 200 OK "POST /api/auth/token HTTP/1.1" +mealie | ERROR 2026-09-21T22:22:14 - No Entry Found on recipe controller action +mealie | ERROR 2026-09-21T22:22:14 - No Entry Found on recipe controller action +mealie | INFO 2026-09-21T22:22:14 - [192.168.0.180:0] 404 Not Found "GET /api/recipes/nope988c526fa6 HTTP/1.1" +mealie | INFO 2026-09-21T22:22:15 - [192.168.0.180:0] 200 OK "GET /api/recipes/drill-2dc46249e8 HTTP/1.1" diff --git a/documentation/audits/update-night-2026-09-21/apps/mealie/badges.json b/documentation/audits/update-night-2026-09-21/apps/mealie/badges.json index 245e88b9..2124a758 100644 --- a/documentation/audits/update-night-2026-09-21/apps/mealie/badges.json +++ b/documentation/audits/update-night-2026-09-21/apps/mealie/badges.json @@ -16,16 +16,16 @@ "after": { "hu": [ { - "title": "Ez az alkalmazás a legfrissebb elérhető változatot futtatja.", - "text": "Naprakész" + "title": "Újabb változat érhető el ehhez az alkalmazáshoz. A frissítés indításához nyomd meg a Frissítés gombot.", + "text": "Frissítés elérhető — ma" } ], "en": [ { - "title": "This app is running the newest version available.", - "text": "Up to date" + "title": "A newer version of this app is available. Select the Update button to start it.", + "text": "Update available — today" } ] }, - "drill_commit": "91e4bb97212e" + "drill_commit": "004105af9a16" } \ No newline at end of file diff --git a/documentation/audits/update-night-2026-09-21/apps/mealie/log.txt b/documentation/audits/update-night-2026-09-21/apps/mealie/log.txt index c854d763..7a5d96da 100644 --- a/documentation/audits/update-night-2026-09-21/apps/mealie/log.txt +++ b/documentation/audits/update-night-2026-09-21/apps/mealie/log.txt @@ -1,11 +1,24 @@ -21:31:06 ==== mealie: ghcr.io/mealie-recipes/mealie:v3.20.1 -> ghcr.io/mealie-recipes/mealie:v3.27.0 (sub=mealie, class=db-postgres) -21:31:07 [1] deploy -> 202 {'ok': True, 'message': 'Telepítés elindítva – az állapot a kártyán követhető'} -21:32:17 [1] deployed, controller state=running, pinned={'mealie': 'ghcr.io/mealie-recipes/mealie:v3.27.0'} -21:32:20 [2] seeding through the app's own front door -21:32:21 mealie: create recipe http=201 -21:32:21 [3] control C1 — reading the seed back BEFORE the update -21:32:21 mealie: readback of the seeded recipe http=200 ok=True -21:32:21 [4] „Mentés most" -> 200 {'ok': True, 'message': 'Mentés elindítva'} -21:33:06 [4] backup idle; last=None -21:33:06 [5] FROM ref not found in compose: ghcr.io/mealie-recipes/mealie:v3.20.1 -21:33:06 [9] verdict inconclusive -> /mnt/5_hdd/felhom.eu/git/felhom.eu/documentation/audits/update-night-2026-09-21/apps/mealie/verdict.json +22:20:09 ==== mealie: ghcr.io/mealie-recipes/mealie:v3.20.1 -> ghcr.io/mealie-recipes/mealie:v3.27.0 (sub=mealie, class=db-postgres) +22:20:09 [1] deploy -> 202 {'ok': True, 'message': 'Telepítés elindítva – az állapot a kártyán követhető'} +22:20:29 [1] deployed, controller state=running, pinned={'mealie': 'ghcr.io/mealie-recipes/mealie:v3.20.1'} +22:20:32 [2] seeding through the app's own front door +22:20:33 mealie: create recipe http=201 +22:20:33 [3] control C1 — reading the seed back BEFORE the update +22:20:34 mealie: readback of the seeded recipe http=200 ok=True +22:20:34 [4] „Mentés most" -> 200 {'ok': True, 'message': 'Mentés elindítva'} +22:21:49 [4] backup idle; last=None +22:21:50 [5] drill commit 004105af9a16: mealie ghcr.io/mealie-recipes/mealie:v3.20.1 -> ghcr.io/mealie-recipes/mealie:v3.27.0 (push rc=0) +22:21:55 [5] badge HU: [{'title': 'Újabb változat érhető el ehhez az alkalmazáshoz. A frissítés indításához nyomd meg a Frissítés gombot.', 'text': 'Frissítés elérhető — ma'}] +22:21:55 [5] badge EN: [{'title': 'A newer version of this app is available. Select the Update button to start it.', 'text': 'Update available — today'}] +22:21:55 [6] Update -> 202 {'ok': True, 'data': {'accepted': True, 'completed': False}, 'message': 'Frissítés elindult – az állapot a kártyán követhető'} +22:21:55 + 0.0s phase=safety-dump label=Adatbázis pillanatkép… err=None hold=None +22:21:56 + 1.1s phase=starting label=Indítás az új verzióval… err=None hold=None +22:21:58 + 3.1s phase=verifying label=Működés ellenőrzése… err=None hold=None +22:22:13 + 18.5s phase=done label=Frissítve err=None hold=None +22:22:13 [7] reading the seed back AFTER the update +22:22:15 mealie: readback of the seeded recipe http=200 ok=True +22:22:17 [8] pinned = {'mealie': 'ghcr.io/mealie-recipes/mealie:v3.27.0'} +22:22:17 [8] installed = {'mealie': 'ghcr.io/mealie-recipes/mealie:v3.27.0'} +22:22:17 [8] compose = ['image: ghcr.io/mealie-recipes/mealie:v3.27.0'] +22:22:17 [8] inspect = ['mealie ghcr.io/mealie-recipes/mealie:v3.27.0 running=true restarts=0'] +22:22:17 [9] verdict proven -> /mnt/5_hdd/felhom.eu/git/felhom.eu/documentation/audits/update-night-2026-09-21/apps/mealie/verdict.json diff --git a/documentation/audits/update-night-2026-09-21/apps/mealie/observables-before.json b/documentation/audits/update-night-2026-09-21/apps/mealie/observables-before.json index 926fb092..e914d9a2 100644 --- a/documentation/audits/update-night-2026-09-21/apps/mealie/observables-before.json +++ b/documentation/audits/update-night-2026-09-21/apps/mealie/observables-before.json @@ -1,15 +1,15 @@ { "pinned_images": { - "mealie": "ghcr.io/mealie-recipes/mealie:v3.27.0" + "mealie": "ghcr.io/mealie-recipes/mealie:v3.20.1" }, "installed_images": { - "mealie": "ghcr.io/mealie-recipes/mealie:v3.27.0" + "mealie": "ghcr.io/mealie-recipes/mealie:v3.20.1" }, "catalog_images": null, "live_compose_image_lines": [ - "image: ghcr.io/mealie-recipes/mealie:v3.27.0" + "image: ghcr.io/mealie-recipes/mealie:v3.20.1" ], "docker_inspect": [ - "mealie ghcr.io/mealie-recipes/mealie:v3.27.0 running=true restarts=0" + "mealie ghcr.io/mealie-recipes/mealie:v3.20.1 running=true restarts=0" ] } \ No newline at end of file diff --git a/documentation/audits/update-night-2026-09-21/apps/mealie/observables.json b/documentation/audits/update-night-2026-09-21/apps/mealie/observables.json index 552a46ad..6f3a0eaf 100644 --- a/documentation/audits/update-night-2026-09-21/apps/mealie/observables.json +++ b/documentation/audits/update-night-2026-09-21/apps/mealie/observables.json @@ -6,9 +6,7 @@ "installed_images": { "mealie": "ghcr.io/mealie-recipes/mealie:v3.20.1" }, - "catalog_images": { - "mealie": "ghcr.io/mealie-recipes/mealie:v3.20.1" - }, + "catalog_images": null, "live_compose_image_lines": [ "image: ghcr.io/mealie-recipes/mealie:v3.20.1" ], @@ -18,19 +16,19 @@ }, "after": { "pinned_images": { - "mealie": "ghcr.io/mealie-recipes/mealie:v3.20.1" + "mealie": "ghcr.io/mealie-recipes/mealie:v3.27.0" }, "installed_images": { - "mealie": "ghcr.io/mealie-recipes/mealie:v3.20.1" + "mealie": "ghcr.io/mealie-recipes/mealie:v3.27.0" }, "catalog_images": { - "mealie": "ghcr.io/mealie-recipes/mealie:v3.20.1" + "mealie": "ghcr.io/mealie-recipes/mealie:v3.27.0" }, "live_compose_image_lines": [ - "image: ghcr.io/mealie-recipes/mealie:v3.20.1" + "image: ghcr.io/mealie-recipes/mealie:v3.27.0" ], "docker_inspect": [ - "mealie ghcr.io/mealie-recipes/mealie:v3.20.1 running=true restarts=0" + "mealie ghcr.io/mealie-recipes/mealie:v3.27.0 running=true restarts=0" ] } } \ No newline at end of file diff --git a/documentation/audits/update-night-2026-09-21/apps/mealie/phases.json b/documentation/audits/update-night-2026-09-21/apps/mealie/phases.json index 9ed9fb60..a7307e60 100644 --- a/documentation/audits/update-night-2026-09-21/apps/mealie/phases.json +++ b/documentation/audits/update-night-2026-09-21/apps/mealie/phases.json @@ -12,6 +12,14 @@ }, { "t": 1.1, + "phase": "starting", + "label": "Indítás az új verzióval…", + "updating": true, + "error": null, + "hold": null + }, + { + "t": 3.1, "phase": "verifying", "label": "Működés ellenőrzése…", "updating": true, @@ -19,7 +27,7 @@ "hold": null }, { - "t": 2.1, + "t": 18.5, "phase": "done", "label": "Frissítve", "updating": false, @@ -27,7 +35,7 @@ "hold": null } ], - "duration_s": 3.1, + "duration_s": 18.5, "final_phase": "done", "update_error": null, "hold_reason": null, diff --git a/documentation/audits/update-night-2026-09-21/apps/mealie/verdict.json b/documentation/audits/update-night-2026-09-21/apps/mealie/verdict.json index 7636ee9a..9ccd6391 100644 --- a/documentation/audits/update-night-2026-09-21/apps/mealie/verdict.json +++ b/documentation/audits/update-night-2026-09-21/apps/mealie/verdict.json @@ -4,20 +4,41 @@ "venue": "guest 9202 demo-hp-scratch, controller 0.261.0", "class": "db-postgres", "from": { + "mealie": "ghcr.io/mealie-recipes/mealie:v3.20.1" + }, + "to": { "mealie": "ghcr.io/mealie-recipes/mealie:v3.27.0" }, - "to": {}, - "verdict": "inconclusive", + "verdict": "proven", "seed_read_before": true, - "seed_read_after": false, - "healthy_after": false, - "migration_observed": null, + "seed_read_after": true, + "healthy_after": true, + "migration_observed": "mealie | INFO 2026-09-21T22:22:05 - Migration needed. Performing migration...", "abort": "not-attempted", "abort_detail": null, - "duration_s": 120.0, - "measured_at": "2026-09-21T19:31:06.937002+00:00", + "duration_s": 18.5, + "measured_at": "2026-09-21T20:20:09.681676+00:00", "evidence": "apps/mealie/", - "notes": [ - "the drill bump could not be committed — the FROM ref did not match the template" - ] + "notes": [], + "badge_catchup_seconds": 4.4, + "observables_after": { + "pinned_images": { + "mealie": "ghcr.io/mealie-recipes/mealie:v3.27.0" + }, + "installed_images": { + "mealie": "ghcr.io/mealie-recipes/mealie:v3.27.0" + }, + "catalog_images": { + "mealie": "ghcr.io/mealie-recipes/mealie:v3.27.0" + }, + "live_compose_image_lines": [ + "image: ghcr.io/mealie-recipes/mealie:v3.27.0" + ], + "docker_inspect": [ + "mealie ghcr.io/mealie-recipes/mealie:v3.27.0 running=true restarts=0" + ] + }, + "final_phase": "done", + "hold_reason": null, + "update_error": null } \ No newline at end of file diff --git a/documentation/audits/update-night-2026-09-21/bad-days/B5-safety-dump-retry/00-backups-before.txt b/documentation/audits/update-night-2026-09-21/bad-days/B5-safety-dump-retry/00-backups-before.txt new file mode 100644 index 00000000..8b3b5069 --- /dev/null +++ b/documentation/audits/update-night-2026-09-21/bad-days/B5-safety-dump-retry/00-backups-before.txt @@ -0,0 +1,30 @@ +=== the app's own recovery unit + its db dumps, with sizes and times +2026-09-21 20:06 288 /mnt/sys_drive/felhom-data/backups/primary/privatebin/compose/app.yaml +2026-09-21 20:06 1050 /mnt/sys_drive/felhom-data/backups/primary/privatebin/manifest.json +2026-09-21 20:06 1235 /mnt/sys_drive/felhom-data/backups/primary/privatebin/compose/docker-compose.yml +2026-09-21 20:06 3125 /mnt/sys_drive/felhom-data/backups/primary/privatebin/compose/.felhom.yml +2026-09-21 20:07 1536 /mnt/sys_drive/felhom-data/backups/primary/privatebin/volume-dumps/privatebin_privatebin_data.tar +=== any temp/partial names left behind +=== the unit manifest, if there is one +--- /mnt/sys_drive/felhom-data/backups/primary/privatebin/manifest.json +{ + "schema_version": 2, + "app_name": "privatebin", + "display_name": "PrivateBin", + "controller_version": "0.261.0", + "created_at": "2026-09-21T20:06:16Z", + "drive": "/mnt/sys_drive", + "namespace_root": "/mnt/sys_drive/felhom-data", + "image_pins": [ + "localhost:5000/drill/paste:2.0.1" + ], + "secret_env_vars": null, + "data_key_env_vars": null, + "secret_source": "portable secrets (data keys, DB passwords, internal signing secrets) are IN this unit's compose/app.yaml (0600); internet-reachable admin logins are NOT, and come from the guest's app.yaml or are regenerated on restore", + "config_files": [ + "docker-compose.yml", + ".felhom.yml", + "app.yaml" + ], + "db_dumps": [], +=== ZERO-BYTE files under this app's backups (a half-write that still looks like a file) diff --git a/documentation/audits/update-night-2026-09-21/bad-days/B5-safety-dump-retry/01-backups-after.txt b/documentation/audits/update-night-2026-09-21/bad-days/B5-safety-dump-retry/01-backups-after.txt new file mode 100644 index 00000000..8b3b5069 --- /dev/null +++ b/documentation/audits/update-night-2026-09-21/bad-days/B5-safety-dump-retry/01-backups-after.txt @@ -0,0 +1,30 @@ +=== the app's own recovery unit + its db dumps, with sizes and times +2026-09-21 20:06 288 /mnt/sys_drive/felhom-data/backups/primary/privatebin/compose/app.yaml +2026-09-21 20:06 1050 /mnt/sys_drive/felhom-data/backups/primary/privatebin/manifest.json +2026-09-21 20:06 1235 /mnt/sys_drive/felhom-data/backups/primary/privatebin/compose/docker-compose.yml +2026-09-21 20:06 3125 /mnt/sys_drive/felhom-data/backups/primary/privatebin/compose/.felhom.yml +2026-09-21 20:07 1536 /mnt/sys_drive/felhom-data/backups/primary/privatebin/volume-dumps/privatebin_privatebin_data.tar +=== any temp/partial names left behind +=== the unit manifest, if there is one +--- /mnt/sys_drive/felhom-data/backups/primary/privatebin/manifest.json +{ + "schema_version": 2, + "app_name": "privatebin", + "display_name": "PrivateBin", + "controller_version": "0.261.0", + "created_at": "2026-09-21T20:06:16Z", + "drive": "/mnt/sys_drive", + "namespace_root": "/mnt/sys_drive/felhom-data", + "image_pins": [ + "localhost:5000/drill/paste:2.0.1" + ], + "secret_env_vars": null, + "data_key_env_vars": null, + "secret_source": "portable secrets (data keys, DB passwords, internal signing secrets) are IN this unit's compose/app.yaml (0600); internet-reachable admin logins are NOT, and come from the guest's app.yaml or are regenerated on restore", + "config_files": [ + "docker-compose.yml", + ".felhom.yml", + "app.yaml" + ], + "db_dumps": [], +=== ZERO-BYTE files under this app's backups (a half-write that still looks like a file) diff --git a/documentation/audits/update-night-2026-09-21/bad-days/B5-safety-dump-retry/log.txt b/documentation/audits/update-night-2026-09-21/bad-days/B5-safety-dump-retry/log.txt new file mode 100644 index 00000000..38e0cdb3 --- /dev/null +++ b/documentation/audits/update-night-2026-09-21/bad-days/B5-safety-dump-retry/log.txt @@ -0,0 +1,16 @@ +22:19:17 ==== B5: a power cut during `safety-dump` on privatebin +22:19:17 [0] pending edge for the cut: installed={'privatebin': 'localhost:5000/drill/paste:2.0.1'} catalog={'privatebin': 'localhost:5000/drill/paste:2.0.2'} +22:19:21 [0] pinned before = {'privatebin': 'localhost:5000/drill/paste:2.0.1'} +22:19:22 [cut] Update -> 202 {'ok': True, 'data': {'accepted': True, 'completed': False}, 'message': 'Frissítés elindult – az állapot a kártyán követ +22:19:22 + 0.023s phase=safety-dump +22:19:22 [cut] phase 'safety-dump' OBSERVED at 2026-09-21T20:19:22.234Z — pulling the plug NOW +22:19:26 [cut] `pct stop 9202` returned after 3.96s :: +22:19:29 [boot] guest 9202 starting +22:19:42 [boot] controller back: Up 6 seconds (healthy) +22:20:05 [boot] recovery lines: 2026/09/21 20:19:36 restore_record_wiring.go:35: [INFO] [backup] restore record wired: /opt/docker/felhom-controller/data/restore-status.json (interrupted at startup: false) | 2026/09/21 20:20:01 bootrecon.go:236: [INFO] [bootrecon] "glance" is a boot orphan by intent but is HELD (held after a failed update (2026-09-21T19:53:02Z) — restore it from its backup to start it) — NOT starting it; whatever is holding it owns its recovery | --- journal file: | ls: cannot access '/var/lib/docker/volumes/f +22:20:09 [2] household sentence HU: ['PrivateBin — Felhom.eu Indítópult Vezérlőpult Alkalmazások Tárhely Meghajtók Hálózati tárhely Biztonsági mentés Áttekintés Távoli mentés Alkalmazások Visszaállítás Megosztás Hálózati megosztás Rendszermonitor Debug Beállítások Rendszer Értesítések Biztonság és hozzáférés 0.261.0 Magyar English Kijelentkezés ↗ Hub kapcsolat kikapcsolva — a központi monitoring nem aktív Rendszermonitor → ← Alkalmazások PrivateBin Fut Frissítés elérhető — ma Megnyitás ↗ Napló Exportálás Beállítások Titkosított szöveg megosztás - a szerver nem látja a tartalmat ~30M RAM security Pi kompatibilis Áthelyezés másik tárhelyre Ennek az alkalmazásnak az adatait másik csatlakoztatott tárhelyre helyezheted át.', 'Érzékeny szövegek biztonságos megosztása E2E titkosítás - a szerver nem fér hozzá a tartalomhoz Beállítható lejárati idő (5 perc - 1 év, vagy soha) Olvasás után automatikus törlés opció Jelszóvédelem a még nagyobb biztonságért Első lépések Nyisd meg a paste.DOMAIN címet a böngészőben Írd be a szöveget és kattints a Küldés gombra Oszd meg a generált linket - a titkosítási kulcs az URL-ben van Dokumentáció Hivatalos dokumentáció ↗'] +22:20:09 [2] household sentence EN: ['PrivateBin — Felhom.eu Launcher Dashboard Apps Storage Drives Network storage Backup Overview Remote backup Apps Restore Sharing Network sharing System monitor Debug Settings System Notifications Security and access 0.261.0 Magyar English Sign out ↗ The hub connection is off — central monitoring is not running System monitor → ← Apps PrivateBin Running Update available — today Open ↗ Log Export Settings Encrypted text sharing - the server never sees the content ~30M RAM security Runs on Pi Move to another storage You can move this app’s data to another connected storage.', 'Share sensitive text safely End-to-end encryption - the server cannot reach the content You choose how long it lasts (5 minutes to 1 year, or never) Optional: delete it automatically after it is read Password protection for extra safety First steps Open paste.DOMAIN in your browser Type your text and select Send Share the link you get - the encryption key is part of the URL Documentation Official documentation ↗'] +22:20:09 privatebin: seeded paste id=5d8afde576872b63 +22:20:09 privatebin: readback http=200 marker_present=True +22:20:09 [3] the app works and holds data after the cut: True +22:20:09 [4] pinned after = {'privatebin': 'localhost:5000/drill/paste:2.0.1'} (moved: False) diff --git a/documentation/audits/update-night-2026-09-21/bad-days/B5-safety-dump-retry/result.json b/documentation/audits/update-night-2026-09-21/bad-days/B5-safety-dump-retry/result.json new file mode 100644 index 00000000..9959d5e3 --- /dev/null +++ b/documentation/audits/update-night-2026-09-21/bad-days/B5-safety-dump-retry/result.json @@ -0,0 +1,72 @@ +{ + "leg": "B5-safety-dump-retry", + "app": "privatebin", + "phase_targeted": "safety-dump", + "observables_before": { + "pinned_images": { + "privatebin": "localhost:5000/drill/paste:2.0.1" + }, + "installed_images": { + "privatebin": "localhost:5000/drill/paste:2.0.1" + }, + "catalog_images": { + "privatebin": "localhost:5000/drill/paste:2.0.2" + }, + "live_compose_image_lines": [ + "image: localhost:5000/drill/paste:2.0.1" + ], + "docker_inspect": [ + "privatebin localhost:5000/drill/paste:2.0.1 running=true restarts=0" + ] + }, + "backup_artefacts_before": "=== the app's own recovery unit + its db dumps, with sizes and times\n2026-09-21 20:06 288 /mnt/sys_drive/felhom-data/backups/primary/privatebin/compose/app.yaml\n2026-09-21 20:06 1050 /mnt/sys_drive/felhom-data/backups/primary/privatebin/manifest.json\n2026-09-21 20:06 1235 /mnt/sys_drive/felhom-data/backups/primary/privatebin/compose/docker-compose.yml\n2026-09-21 20:06 3125 /mnt/sys_drive/felhom-data/backups/primary/privatebin/compose/.felhom.yml\n2026-09-21 20:07 1536 /mnt/sys_drive/felhom-data/backups/primary/privatebin/volume-dumps/privatebin_privatebin_data.tar\n=== any temp/partial names left behind\n=== the unit manifest, if there is one\n--- /mnt/sys_drive/felhom-data/backups/primary/privatebin/manifest.json\n{\n \"schema_version\": 2,\n \"app_name\": \"privatebin\",\n \"display_name\": \"PrivateBin\",\n \"controller_version\": \"0.261.0\",\n \"created_at\": \"2026-09-21T20:06:16Z\",\n \"drive\": \"/mnt/sys_drive\",\n \"namespace_root\": \"/mnt/sys_drive/felhom-data\",\n \"image_pins\": [\n \"localhost:5000/drill/paste:2.0.1\"\n ],\n \"secret_env_vars\": null,\n \"data_key_env_vars\": null,\n \"secret_source\": \"portable secrets (data keys, DB passwords, internal signing secrets) are IN this unit's compose/app.yaml (0600); internet-reachable admin logins are NOT, and come from the guest's app.yaml or are regenerated on restore\",\n \"config_files\": [\n \"docker-compose.yml\",\n \".felhom.yml\",\n \"app.yaml\"\n ],\n \"db_dumps\": [],\n=== ZERO-BYTE files under this app's backups (a half-write that still looks like a file)\n", + "cut": { + "pressed": true, + "phases_seen": [ + { + "t": 0.023, + "phase": "safety-dump" + } + ], + "cut_decided_at": "2026-09-21T20:19:22.234Z", + "pct_stop_returned_after_s": 3.96, + "instrument_limit": "the phase at the DECISION is observed; the phase at the FREEZE is inferred — pct stop is not instantaneous" + }, + "after_boot": { + "recovery_lines": "2026/09/21 20:19:36 restore_record_wiring.go:35: [INFO] [backup] restore record wired: /opt/docker/felhom-controller/data/restore-status.json (interrupted at startup: false)\n2026/09/21 20:20:01 bootrecon.go:236: [INFO] [bootrecon] \"glance\" is a boot orphan by intent but is HELD (held after a failed update (2026-09-21T19:53:02Z) — restore it from its backup to start it) — NOT starting it; whatever is holding it owns its recovery\n--- journal file:\nls: cannot access '/var/lib/docker/volumes/felhom-controller-data/_data/data/update-journal.json': No such file or directory\n", + "state": "running", + "update_phase": null, + "update_error": null, + "hold_reason": null, + "observables": { + "pinned_images": { + "privatebin": "localhost:5000/drill/paste:2.0.1" + }, + "installed_images": { + "privatebin": "localhost:5000/drill/paste:2.0.1" + }, + "catalog_images": { + "privatebin": "localhost:5000/drill/paste:2.0.2" + }, + "live_compose_image_lines": [ + "image: localhost:5000/drill/paste:2.0.1" + ], + "docker_inspect": [ + "privatebin localhost:5000/drill/paste:2.0.1 running=true restarts=0" + ] + } + }, + "backup_artefacts_after": "=== the app's own recovery unit + its db dumps, with sizes and times\n2026-09-21 20:06 288 /mnt/sys_drive/felhom-data/backups/primary/privatebin/compose/app.yaml\n2026-09-21 20:06 1050 /mnt/sys_drive/felhom-data/backups/primary/privatebin/manifest.json\n2026-09-21 20:06 1235 /mnt/sys_drive/felhom-data/backups/primary/privatebin/compose/docker-compose.yml\n2026-09-21 20:06 3125 /mnt/sys_drive/felhom-data/backups/primary/privatebin/compose/.felhom.yml\n2026-09-21 20:07 1536 /mnt/sys_drive/felhom-data/backups/primary/privatebin/volume-dumps/privatebin_privatebin_data.tar\n=== any temp/partial names left behind\n=== the unit manifest, if there is one\n--- /mnt/sys_drive/felhom-data/backups/primary/privatebin/manifest.json\n{\n \"schema_version\": 2,\n \"app_name\": \"privatebin\",\n \"display_name\": \"PrivateBin\",\n \"controller_version\": \"0.261.0\",\n \"created_at\": \"2026-09-21T20:06:16Z\",\n \"drive\": \"/mnt/sys_drive\",\n \"namespace_root\": \"/mnt/sys_drive/felhom-data\",\n \"image_pins\": [\n \"localhost:5000/drill/paste:2.0.1\"\n ],\n \"secret_env_vars\": null,\n \"data_key_env_vars\": null,\n \"secret_source\": \"portable secrets (data keys, DB passwords, internal signing secrets) are IN this unit's compose/app.yaml (0600); internet-reachable admin logins are NOT, and come from the guest's app.yaml or are regenerated on restore\",\n \"config_files\": [\n \"docker-compose.yml\",\n \".felhom.yml\",\n \"app.yaml\"\n ],\n \"db_dumps\": [],\n=== ZERO-BYTE files under this app's backups (a half-write that still looks like a file)\n", + "sentences": { + "hu": [ + "PrivateBin — Felhom.eu Indítópult Vezérlőpult Alkalmazások Tárhely Meghajtók Hálózati tárhely Biztonsági mentés Áttekintés Távoli mentés Alkalmazások Visszaállítás Megosztás Hálózati megosztás Rendszermonitor Debug Beállítások Rendszer Értesítések Biztonság és hozzáférés 0.261.0 Magyar English Kijelentkezés ↗ Hub kapcsolat kikapcsolva — a központi monitoring nem aktív Rendszermonitor → ← Alkalmazások PrivateBin Fut Frissítés elérhető — ma Megnyitás ↗ Napló Exportálás Beállítások Titkosított szöveg megosztás - a szerver nem látja a tartalmat ~30M RAM security Pi kompatibilis Áthelyezés másik tárhelyre Ennek az alkalmazásnak az adatait másik csatlakoztatott tárhelyre helyezheted át.", + "Érzékeny szövegek biztonságos megosztása E2E titkosítás - a szerver nem fér hozzá a tartalomhoz Beállítható lejárati idő (5 perc - 1 év, vagy soha) Olvasás után automatikus törlés opció Jelszóvédelem a még nagyobb biztonságért Első lépések Nyisd meg a paste.DOMAIN címet a böngészőben Írd be a szöveget és kattints a Küldés gombra Oszd meg a generált linket - a titkosítási kulcs az URL-ben van Dokumentáció Hivatalos dokumentáció ↗" + ], + "en": [ + "PrivateBin — Felhom.eu Launcher Dashboard Apps Storage Drives Network storage Backup Overview Remote backup Apps Restore Sharing Network sharing System monitor Debug Settings System Notifications Security and access 0.261.0 Magyar English Sign out ↗ The hub connection is off — central monitoring is not running System monitor → ← Apps PrivateBin Running Update available — today Open ↗ Log Export Settings Encrypted text sharing - the server never sees the content ~30M RAM security Runs on Pi Move to another storage You can move this app’s data to another connected storage.", + "Share sensitive text safely End-to-end encryption - the server cannot reach the content You choose how long it lasts (5 minutes to 1 year, or never) Optional: delete it automatically after it is read Password protection for extra safety First steps Open paste.DOMAIN in your browser Type your text and select Send Share the link you get - the encryption key is part of the URL Documentation Official documentation ↗" + ] + }, + "app_usable_after": true, + "pin_moved": false +} \ No newline at end of file diff --git a/documentation/audits/update-night-2026-09-21/bad-days/P2.3-pg-rehearsal/log.txt b/documentation/audits/update-night-2026-09-21/bad-days/P2.3-pg-rehearsal/log.txt new file mode 100644 index 00000000..dd8b0571 --- /dev/null +++ b/documentation/audits/update-night-2026-09-21/bad-days/P2.3-pg-rehearsal/log.txt @@ -0,0 +1,70 @@ +22:13:25 ==== Phase 2.3: the PostgreSQL 16 -> 17 conversion rehearsal (Q5) +22:13:25 removing the existing docmost so the rehearsal meets a FRESH workspace it can seed +22:13:27 [X] stop -> 200 {'ok': True, 'message': 'Stack docmost stop completed'} +22:13:32 [X] remove (with drive data) -> 200 {'ok': True, 'data': {'removed': 'docmost', 'volumes_removed': ['docmost_docmost_postgres_data', 'docmost_docmost_redis_data', 'docmost_docmost_storage'], 'hdd_ +22:13:40 [X] after remove: deployed=False leftovers='/opt/docker/stacks/docmost' +22:13:40 [1] deploy -> 202 {'ok': True, 'message': 'Telepítés elindítva – az állapot a kártyán követhető'} +22:14:05 [1] deployed, controller state=running, pinned={'docmost': 'docmost/docmost:0.96.0', 'docmost-postgres': 'postgres:16-alpine', 'docmost-redis': 'redis:7-alpine'} +22:14:05 docmost: /api/auth/setup http=200 rc=0 +22:14:06 docmost: login as the seeded user http=200 ok=True +22:14:09 [00-state-before] 2.7s +22:14:09 PG_VERSION (the harness's own postgres probe, verbatim): +22:14:09 16 +22:14:09 engine version: +22:14:09 postgres (PostgreSQL) 16.15 +22:14:09 datadir size: +22:14:09 49.0M /var/lib/postgresql/data +22:14:09 volume: +22:14:09 docmost_docmost_postgres_data +22:14:12 [01-stop-the-app-keep-the-engine] 3.2s +22:14:12 app stopped (the engine stays up to be dumped) +22:14:12 docmost-postgres Up 31 seconds (healthy) +22:14:12 docmost-redis Up 31 seconds (healthy) +22:14:15 [02-dump-with-16] 2.6s +22:14:15 dumping as user=docmost +22:14:15 rc=0 +22:14:15 dump bytes: 132201 +22:14:15 CREATE TABLE statements: 48 +22:14:21 [03-fresh-17-datadir-and-restore] 6.5s +22:14:21 17 up: postgres (PostgreSQL) 17.11 +22:14:21 PG_VERSION on the fresh datadir: 17 +22:14:21 restore rc=0 +22:14:21 ERROR lines in the restore: 2 +22:14:21 ERROR: role "docmost" already exists +22:14:21 ERROR: database "docmost" already exists +22:14:21 tables restored: +22:14:21 48 +22:16:26 [04-point-the-app-at-17-and-start-it] 124.8s +22:16:26 DRILL-pg17 now answers to the name docmost-postgres on docmost_docmost-internal +22:16:26 app started +22:16:26 docmost Up 2 minutes (healthy) +22:16:26 docmost-postgres Exited (0) 2 minutes ago +22:16:26 docmost-redis Up 2 minutes (healthy) +22:16:26 {"level":"info","time":"2026-09-21T20:14:00.698Z","pid":45,"hostname":"24a35fe9b998","context":"NestApplication","msg":"Listening on http://127.0.0.1:3000 / https://docs.enkisfelhom.hu"} +22:16:26 [ELIFECYCLE] Command failed. +22:16:26 $ pnpm --filter ./apps/server run start:prod +22:16:26 $ cross-env NODE_ENV=production node dist/main +22:16:26 (node:45) ExperimentalWarning: localStorage is not available because --localstorage-file was not provided. +22:16:26 (Use `node --trace-warnings ...` to show where the warning was created) +22:16:26 {"level":"info","time":"2026-09-21T20:14:31.727Z","pid":45,"hostname":"24a35fe9b998","context":"RedisModule","msg":"default: the connection was successfully established"} +22:16:26 {"level":"info","time":"2026-09-21T20:14:32.033Z","pid":45,"hostname":"24a35fe9b998","context":"DatabaseModule","msg":"Establishing database connection"} +22:16:26 {"level":"info","time":"2026-09-21T20:14:32.065Z","pid":45,"hostname":"24a35fe9b998","context":"DatabaseModule","msg":"Database connection successful"} +22:16:27 docmost: login as the seeded user http=200 ok=True +22:16:27 [05] the seed read back on PostgreSQL 17: True +22:16:29 [06-engine-state-after] 2.3s +22:16:29 the harness's own postgres probe against the CONVERTED datadir: +22:16:29 17 +22:16:29 [exit=0] +22:16:29 postgres (PostgreSQL) 17.11 +22:16:29 size of the 17 datadir: +22:16:29 49.1M /var/lib/postgresql/data +22:16:43 [99-put-everything-back] 13.8s +22:16:43 docmost docmost/docmost:0.96.0 Up 5 seconds (health: starting) +22:16:43 docmost-postgres postgres:16-alpine Up 10 seconds (healthy) +22:16:43 docmost-redis redis:7-alpine Up 3 minutes (healthy) +22:16:43 docmost-postgres postgres:16-alpine Up 10 seconds (healthy) +22:16:43 PG_VERSION back on the original datadir: 16 +22:16:43 docmost: login as the seeded user http=404 ok=False +22:16:43 docmost: login body 404 page not found + +22:16:43 [99] the seed still reads on the ORIGINAL 16 datadir after teardown: False diff --git a/documentation/audits/update-night-2026-09-21/bad-days/P2.3-pg-rehearsal/result.json b/documentation/audits/update-night-2026-09-21/bad-days/P2.3-pg-rehearsal/result.json index 78f2b9a6..c0b6ed1d 100644 --- a/documentation/audits/update-night-2026-09-21/bad-days/P2.3-pg-rehearsal/result.json +++ b/documentation/audits/update-night-2026-09-21/bad-days/P2.3-pg-rehearsal/result.json @@ -2,9 +2,47 @@ "leg": "P2.3", "app": "docmost", "route": "logical dump and restore", - "measured_at": "2026-09-21T20:11:46.041578+00:00", - "steps": [], - "notes": [ - "could not seed before the rehearsal — see log" - ] + "measured_at": "2026-09-21T20:13:25.874392+00:00", + "steps": [ + { + "label": "00-state-before", + "seconds": 2.7, + "output": "PG_VERSION (the harness's own postgres probe, verbatim):\n16\nengine version:\npostgres (PostgreSQL) 16.15\ndatadir size:\n49.0M\t/var/lib/postgresql/data\nvolume:\ndocmost_docmost_postgres_data\n" + }, + { + "label": "01-stop-the-app-keep-the-engine", + "seconds": 3.2, + "output": "app stopped (the engine stays up to be dumped)\ndocmost-postgres Up 31 seconds (healthy)\ndocmost-redis Up 31 seconds (healthy)\n" + }, + { + "label": "02-dump-with-16", + "seconds": 2.6, + "output": "dumping as user=docmost\nrc=0\ndump bytes: 132201\nCREATE TABLE statements: 48\n" + }, + { + "label": "03-fresh-17-datadir-and-restore", + "seconds": 6.5, + "output": "17 up: postgres (PostgreSQL) 17.11\nPG_VERSION on the fresh datadir: 17\nrestore rc=0\nERROR lines in the restore: 2\nERROR: role \"docmost\" already exists\nERROR: database \"docmost\" already exists\ntables restored:\n48\n" + }, + { + "label": "04-point-the-app-at-17-and-start-it", + "seconds": 124.8, + "output": "DRILL-pg17 now answers to the name docmost-postgres on docmost_docmost-internal\napp started\ndocmost Up 2 minutes (healthy)\ndocmost-postgres Exited (0) 2 minutes ago\ndocmost-redis Up 2 minutes (healthy)\n{\"level\":\"info\",\"time\":\"2026-09-21T20:14:00.698Z\",\"pid\":45,\"hostname\":\"24a35fe9b998\",\"context\":\"NestApplication\",\"msg\":\"Listening on http://127.0.0.1:3000 / https://docs.enkisfelhom.hu\"}\n[ELIFECYCLE] Command failed.\n$ pnpm --filter ./apps/server run start:prod\n$ cross-env NODE_ENV=production node dist/main\n(node:45) ExperimentalWarning: localStorage is not available because --localstorage-file was not provided.\n(Use `node --trace-warnings ...` to show where the warning was created)\n{\"level\":\"info\",\"time\":\"2026-09-21T20:14:31.727Z\",\"pid\":45,\"hostname\":\"24a35fe9b998\",\"context\":\"RedisModule\",\"msg\":\"default: the connection was successfully established\"}\n{\"level\":\"info\",\"time\":\"2026-09-21T20:14:32.033Z\",\"pid\":45,\"hostname\":\"24a35fe9b998\",\"context\":\"DatabaseModule\",\"msg\":\"Establishing database connection\"}\n{\"level\":\"info\",\"time\":\"2026-09-21T20:14:32.065Z\",\"pid\":45,\"hostname\":\"24a35fe9b998\",\"context\":\"DatabaseModule\",\"msg\":\"Database connection successful\"}\n{\"level\":\"info\",\"time\":\"2026-09-21T20:14:32.222Z\",\"pid\":45,\"hostname\":\"24a35fe9b998\",\"context\":\"DatabaseMigrationService\",\"msg\":\"No pending database migrations\"}\n{\"level\":\"info\",\"time\":\"2026-09-21T20:14:32.266Z\",\"pid\":45,\"hostname\":\"24a35fe9b998\",\"context\":\"NestApplication\",\"msg\":\"Nest application successfully started\"}\n{\"level\":\"info\",\"time\":\"2026-09-21T20:14:32.282Z\",\"pid\":45,\"hostname\":\"24a35fe9b998\",\"context\":\"NestApplication\",\"msg\":\"Listening on http://127.0.0.1:3000 / https://docs.enkisfelhom.hu\"}\n" + }, + { + "label": "06-engine-state-after", + "seconds": 2.3, + "output": "the harness's own postgres probe against the CONVERTED datadir:\n17\n[exit=0]\npostgres (PostgreSQL) 17.11\nsize of the 17 datadir:\n49.1M\t/var/lib/postgresql/data\n" + }, + { + "label": "99-put-everything-back", + "seconds": 13.8, + "output": "docmost docmost/docmost:0.96.0 Up 5 seconds (health: starting)\ndocmost-postgres postgres:16-alpine Up 10 seconds (healthy)\ndocmost-redis redis:7-alpine Up 3 minutes (healthy)\ndocmost-postgres postgres:16-alpine Up 10 seconds (healthy)\nPG_VERSION back on the original datadir: 16\n" + } + ], + "notes": [], + "seed_read_before": true, + "seed_read_after_on_17": true, + "seed_read_back_on_16_after_teardown": false, + "total_seconds": 155.9 } \ No newline at end of file diff --git a/documentation/audits/update-night-2026-09-21/leftovers.log b/documentation/audits/update-night-2026-09-21/leftovers.log new file mode 100644 index 00000000..20c1509d --- /dev/null +++ b/documentation/audits/update-night-2026-09-21/leftovers.log @@ -0,0 +1,74 @@ + +############################################## 22:16:43 L1 wger — the LAST unmeasured candidate in R-618, and the one that would close it +22:16:43 === wger — R-618's last unmeasured candidate +22:16:43 template probe : type=http port=80 (.felhom.yml) +22:16:43 traefik label : loadbalancer.server.port=8000 (the SAME template's compose) +22:16:45 [1] deploy -> 202 {'ok': True, 'message': 'Telepítés elindítva – az állapot a kártyán követhető'} +22:18:26 [1] deployed, controller state=unhealthy, pinned={'wger': 'wger/server:2.6'} +22:18:26 [1] NOTE: the controller's own state is 'unhealthy', not 'running' — recorded, not treated as a failure; the fixture's front-door wait is the real gate +22:18:29 --- what wger actually LISTENS on, asked inside its own container: +22:18:29 --- the container's own docker healthcheck, if it has one: +22:18:29 healthy +22:18:29 --- port 80 from inside (the port the probe names): +22:18:29 refused on 80 +22:18:29 --- port 8000 from inside (the port the compose says): +22:18:29 ANSWERED on 8000 +22:18:29 the household's own front door: http=302 +22:18:29 the box's own verdict: state='unhealthy' +22:18:29 VERDICT: CONFIRMED DEFECT — the same shape as tandoor +22:18:29 written -> /mnt/5_hdd/felhom.eu/git/felhom.eu/documentation/audits/update-night-2026-09-21/22-wger-probe-measured.txt +22:18:39 [X] stop -> 200 {'ok': True, 'message': 'Stack wger stop completed'} +22:18:45 [X] remove (with drive data) -> 200 {'ok': True, 'data': {'removed': 'wger', 'volumes_removed': ['wger_wger_data', 'wger_wger_media'], 'hdd_paths_removed': [], 'hdd_paths_preserved': [], 'hdd_note +22:18:52 [X] after remove: deployed=False leftovers='/opt/docker/stacks/wger' + +############################################## 22:18:52 L2 B5 retry — the cut in safety-dump, with a REAL pending edge this time +22:19:17 [knob] backup_max_age=24h :: update: | backup_max_age: 24h | Up 22 seconds (healthy) +22:19:17 ==== B5: a power cut during `safety-dump` on privatebin +22:19:17 [0] pending edge for the cut: installed={'privatebin': 'localhost:5000/drill/paste:2.0.1'} catalog={'privatebin': 'localhost:5000/drill/paste:2.0.2'} +22:19:21 [0] pinned before = {'privatebin': 'localhost:5000/drill/paste:2.0.1'} +22:19:22 [cut] Update -> 202 {'ok': True, 'data': {'accepted': True, 'completed': False}, 'message': 'Frissítés elindult – az állapot a kártyán követ +22:19:22 + 0.023s phase=safety-dump +/mnt/5_hdd/felhom.eu/git/felhom.eu/documentation/audits/update-night-2026-09-21/phase3_b5.py:82: DeprecationWarning: datetime.datetime.utcnow() is deprecated and scheduled for removal in a future version. Use timezone-aware objects to represent datetimes in UTC: datetime.datetime.now(datetime.UTC). + decided = datetime.utcnow().isoformat(timespec="milliseconds") + "Z" +22:19:22 [cut] phase 'safety-dump' OBSERVED at 2026-09-21T20:19:22.234Z — pulling the plug NOW +22:19:26 [cut] `pct stop 9202` returned after 3.96s :: +22:19:29 [boot] guest 9202 starting +22:19:42 [boot] controller back: Up 6 seconds (healthy) +22:20:05 [boot] recovery lines: 2026/09/21 20:19:36 restore_record_wiring.go:35: [INFO] [backup] restore record wired: /opt/docker/felhom-controller/data/restore-status.json (interrupted at startup: false) | 2026/09/21 20:20:01 bootrecon.go:236: [INFO] [bootrecon] "glance" is a boot orphan by intent but is HELD (held after a failed update (2026-09-21T19:53:02Z) — restore it from its backup to start it) — NOT starting it; whatever is holding it owns its recovery | --- journal file: | ls: cannot access '/var/lib/docker/volumes/f +22:20:09 [2] household sentence HU: ['PrivateBin — Felhom.eu Indítópult Vezérlőpult Alkalmazások Tárhely Meghajtók Hálózati tárhely Biztonsági mentés Áttekintés Távoli mentés Alkalmazások Visszaállítás Megosztás Hálózati megosztás Rendszermonitor Debug Beállítások Rendszer Értesítések Biztonság és hozzáférés 0.261.0 Magyar English Kijelentkezés ↗ Hub kapcsolat kikapcsolva — a központi monitoring nem aktív Rendszermonitor → ← Alkalmazások PrivateBin Fut Frissítés elérhető — ma Megnyitás ↗ Napló Exportálás Beállítások Titkosított szöveg megosztás - a szerver nem látja a tartalmat ~30M RAM security Pi kompatibilis Áthelyezés másik tárhelyre Ennek az alkalmazásnak az adatait másik csatlakoztatott tárhelyre helyezheted át.', 'Érzékeny szövegek biztonságos megosztása E2E titkosítás - a szerver nem fér hozzá a tartalomhoz Beállítható lejárati idő (5 perc - 1 év, vagy soha) Olvasás után automatikus törlés opció Jelszóvédelem a még nagyobb biztonságért Első lépések Nyisd meg a paste.DOMAIN címet a böngészőben Írd be a szöveget és kattints a Küldés gombra Oszd meg a generált linket - a titkosítási kulcs az URL-ben van Dokumentáció Hivatalos dokumentáció ↗'] +22:20:09 [2] household sentence EN: ['PrivateBin — Felhom.eu Launcher Dashboard Apps Storage Drives Network storage Backup Overview Remote backup Apps Restore Sharing Network sharing System monitor Debug Settings System Notifications Security and access 0.261.0 Magyar English Sign out ↗ The hub connection is off — central monitoring is not running System monitor → ← Apps PrivateBin Running Update available — today Open ↗ Log Export Settings Encrypted text sharing - the server never sees the content ~30M RAM security Runs on Pi Move to another storage You can move this app’s data to another connected storage.', 'Share sensitive text safely End-to-end encryption - the server cannot reach the content You choose how long it lasts (5 minutes to 1 year, or never) Optional: delete it automatically after it is read Password protection for extra safety First steps Open paste.DOMAIN in your browser Type your text and select Send Share the link you get - the encryption key is part of the URL Documentation Official documentation ↗'] +22:20:09 privatebin: seeded paste id=5d8afde576872b63 +22:20:09 privatebin: readback http=200 marker_present=True +22:20:09 [3] the app works and holds data after the cut: True +22:20:09 [4] pinned after = {'privatebin': 'localhost:5000/drill/paste:2.0.1'} (moved: False) + +############################################## 22:20:09 L3 mealie — the edge that was never walked because its drill pin had already moved +22:20:09 ==== mealie: ghcr.io/mealie-recipes/mealie:v3.20.1 -> ghcr.io/mealie-recipes/mealie:v3.27.0 (sub=mealie, class=db-postgres) +22:20:09 [1] deploy -> 202 {'ok': True, 'message': 'Telepítés elindítva – az állapot a kártyán követhető'} +22:20:29 [1] deployed, controller state=running, pinned={'mealie': 'ghcr.io/mealie-recipes/mealie:v3.20.1'} +22:20:32 [2] seeding through the app's own front door +22:20:33 mealie: create recipe http=201 +22:20:33 [3] control C1 — reading the seed back BEFORE the update +22:20:34 mealie: readback of the seeded recipe http=200 ok=True +22:20:34 [4] „Mentés most" -> 200 {'ok': True, 'message': 'Mentés elindítva'} +22:21:49 [4] backup idle; last=None +22:21:50 [5] drill commit 004105af9a16: mealie ghcr.io/mealie-recipes/mealie:v3.20.1 -> ghcr.io/mealie-recipes/mealie:v3.27.0 (push rc=0) +22:21:55 [5] badge HU: [{'title': 'Újabb változat érhető el ehhez az alkalmazáshoz. A frissítés indításához nyomd meg a Frissítés gombot.', 'text': 'Frissítés elérhető — ma'}] +22:21:55 [5] badge EN: [{'title': 'A newer version of this app is available. Select the Update button to start it.', 'text': 'Update available — today'}] +22:21:55 [6] Update -> 202 {'ok': True, 'data': {'accepted': True, 'completed': False}, 'message': 'Frissítés elindult – az állapot a kártyán követhető'} +22:21:55 + 0.0s phase=safety-dump label=Adatbázis pillanatkép… err=None hold=None +22:21:56 + 1.1s phase=starting label=Indítás az új verzióval… err=None hold=None +22:21:58 + 3.1s phase=verifying label=Működés ellenőrzése… err=None hold=None +22:22:13 + 18.5s phase=done label=Frissítve err=None hold=None +22:22:13 [7] reading the seed back AFTER the update +22:22:15 mealie: readback of the seeded recipe http=200 ok=True +22:22:17 [8] pinned = {'mealie': 'ghcr.io/mealie-recipes/mealie:v3.27.0'} +22:22:17 [8] installed = {'mealie': 'ghcr.io/mealie-recipes/mealie:v3.27.0'} +22:22:17 [8] compose = ['image: ghcr.io/mealie-recipes/mealie:v3.27.0'] +22:22:17 [8] inspect = ['mealie ghcr.io/mealie-recipes/mealie:v3.27.0 running=true restarts=0'] +22:22:17 [9] verdict proven -> /mnt/5_hdd/felhom.eu/git/felhom.eu/documentation/audits/update-night-2026-09-21/apps/mealie/verdict.json +22:22:19 [X] stop -> 200 {'ok': True, 'message': 'Stack mealie stop completed'} +22:22:24 [X] remove (with drive data) -> 200 {'ok': True, 'data': {'removed': 'mealie', 'volumes_removed': ['mealie_mealie_data'], 'hdd_paths_removed': [], 'hdd_paths_preserved': [], 'hdd_note': 'Az alkalm +22:22:32 [X] after remove: deployed=False leftovers='/opt/docker/stacks/mealie' + +22:22:32 leftovers done diff --git a/documentation/audits/update-night-2026-09-21/pgrehearsal.log b/documentation/audits/update-night-2026-09-21/pgrehearsal.log index 1f55a9ae..929e76c9 100644 --- a/documentation/audits/update-night-2026-09-21/pgrehearsal.log +++ b/documentation/audits/update-night-2026-09-21/pgrehearsal.log @@ -4,3 +4,68 @@ 22:13:32 [X] remove (with drive data) -> 200 {'ok': True, 'data': {'removed': 'docmost', 'volumes_removed': ['docmost_docmost_postgres_data', 'docmost_docmost_redis_data', 'docmost_docmost_storage'], 'hdd_ 22:13:40 [X] after remove: deployed=False leftovers='/opt/docker/stacks/docmost' 22:13:40 [1] deploy -> 202 {'ok': True, 'message': 'Telepítés elindítva – az állapot a kártyán követhető'} +22:14:05 [1] deployed, controller state=running, pinned={'docmost': 'docmost/docmost:0.96.0', 'docmost-postgres': 'postgres:16-alpine', 'docmost-redis': 'redis:7-alpine'} +22:14:05 docmost: /api/auth/setup http=200 rc=0 +22:14:06 docmost: login as the seeded user http=200 ok=True +22:14:09 [00-state-before] 2.7s +22:14:09 PG_VERSION (the harness's own postgres probe, verbatim): +22:14:09 16 +22:14:09 engine version: +22:14:09 postgres (PostgreSQL) 16.15 +22:14:09 datadir size: +22:14:09 49.0M /var/lib/postgresql/data +22:14:09 volume: +22:14:09 docmost_docmost_postgres_data +22:14:12 [01-stop-the-app-keep-the-engine] 3.2s +22:14:12 app stopped (the engine stays up to be dumped) +22:14:12 docmost-postgres Up 31 seconds (healthy) +22:14:12 docmost-redis Up 31 seconds (healthy) +22:14:15 [02-dump-with-16] 2.6s +22:14:15 dumping as user=docmost +22:14:15 rc=0 +22:14:15 dump bytes: 132201 +22:14:15 CREATE TABLE statements: 48 +22:14:21 [03-fresh-17-datadir-and-restore] 6.5s +22:14:21 17 up: postgres (PostgreSQL) 17.11 +22:14:21 PG_VERSION on the fresh datadir: 17 +22:14:21 restore rc=0 +22:14:21 ERROR lines in the restore: 2 +22:14:21 ERROR: role "docmost" already exists +22:14:21 ERROR: database "docmost" already exists +22:14:21 tables restored: +22:14:21 48 +22:16:26 [04-point-the-app-at-17-and-start-it] 124.8s +22:16:26 DRILL-pg17 now answers to the name docmost-postgres on docmost_docmost-internal +22:16:26 app started +22:16:26 docmost Up 2 minutes (healthy) +22:16:26 docmost-postgres Exited (0) 2 minutes ago +22:16:26 docmost-redis Up 2 minutes (healthy) +22:16:26 {"level":"info","time":"2026-09-21T20:14:00.698Z","pid":45,"hostname":"24a35fe9b998","context":"NestApplication","msg":"Listening on http://127.0.0.1:3000 / https://docs.enkisfelhom.hu"} +22:16:26 [ELIFECYCLE] Command failed. +22:16:26 $ pnpm --filter ./apps/server run start:prod +22:16:26 $ cross-env NODE_ENV=production node dist/main +22:16:26 (node:45) ExperimentalWarning: localStorage is not available because --localstorage-file was not provided. +22:16:26 (Use `node --trace-warnings ...` to show where the warning was created) +22:16:26 {"level":"info","time":"2026-09-21T20:14:31.727Z","pid":45,"hostname":"24a35fe9b998","context":"RedisModule","msg":"default: the connection was successfully established"} +22:16:26 {"level":"info","time":"2026-09-21T20:14:32.033Z","pid":45,"hostname":"24a35fe9b998","context":"DatabaseModule","msg":"Establishing database connection"} +22:16:26 {"level":"info","time":"2026-09-21T20:14:32.065Z","pid":45,"hostname":"24a35fe9b998","context":"DatabaseModule","msg":"Database connection successful"} +22:16:27 docmost: login as the seeded user http=200 ok=True +22:16:27 [05] the seed read back on PostgreSQL 17: True +22:16:29 [06-engine-state-after] 2.3s +22:16:29 the harness's own postgres probe against the CONVERTED datadir: +22:16:29 17 +22:16:29 [exit=0] +22:16:29 postgres (PostgreSQL) 17.11 +22:16:29 size of the 17 datadir: +22:16:29 49.1M /var/lib/postgresql/data +22:16:43 [99-put-everything-back] 13.8s +22:16:43 docmost docmost/docmost:0.96.0 Up 5 seconds (health: starting) +22:16:43 docmost-postgres postgres:16-alpine Up 10 seconds (healthy) +22:16:43 docmost-redis redis:7-alpine Up 3 minutes (healthy) +22:16:43 docmost-postgres postgres:16-alpine Up 10 seconds (healthy) +22:16:43 PG_VERSION back on the original datadir: 16 +22:16:43 docmost: login as the seeded user http=404 ok=False +22:16:43 docmost: login body 404 page not found + +22:16:43 [99] the seed still reads on the ORIGINAL 16 datadir after teardown: False +22:16:43 rehearsal total 155.9s -> /mnt/5_hdd/felhom.eu/git/felhom.eu/documentation/audits/update-night-2026-09-21/bad-days/P2.3-pg-rehearsal/result.json diff --git a/documentation/audits/update-night-2026-09-21/teardown.log b/documentation/audits/update-night-2026-09-21/teardown.log new file mode 100644 index 00000000..6c20599d --- /dev/null +++ b/documentation/audits/update-night-2026-09-21/teardown.log @@ -0,0 +1,53 @@ +22:23:52 ==== Phase 5: teardown, three layers +22:23:55 -> 00-host-before.txt +22:23:55 [M] throwaway apps still deployed: ['bentopdf', 'docmost', 'gitea', 'glance', 'nextcloud', 'opengist', 'privatebin', 'uptime-kuma', 'vaultwarden', 'wishlist', 'zipline'] +22:23:55 [X] stop -> 200 {'ok': True, 'message': 'Stack bentopdf stop completed'} +22:24:01 [X] remove (with drive data) -> 200 {'ok': True, 'data': {'removed': 'bentopdf', 'volumes_removed': [], 'hdd_paths_removed': [], 'hdd_paths_preserved': [], 'hdd_note': 'Az alkalmazás nem tárolt sa +22:24:08 [X] after remove: deployed=False leftovers='/opt/docker/stacks/bentopdf' +22:24:09 [X] stop -> 200 {'ok': True, 'message': 'Stack docmost stop completed'} +22:24:15 [X] remove (with drive data) -> 200 {'ok': True, 'data': {'removed': 'docmost', 'volumes_removed': ['docmost_docmost_postgres_data', 'docmost_docmost_redis_data', 'docmost_docmost_storage'], 'hdd_ +22:24:22 [X] after remove: deployed=False leftovers='/opt/docker/stacks/docmost' +22:24:22 [X] stop -> 200 {'ok': True, 'message': 'Stack gitea stop completed'} +22:24:28 [X] remove (with drive data) -> 200 {'ok': True, 'data': {'removed': 'gitea', 'volumes_removed': ['gitea_gitea_data'], 'hdd_paths_removed': [], 'hdd_paths_preserved': [], 'hdd_note': 'Az alkalmazá +22:24:35 [X] after remove: deployed=False leftovers='/opt/docker/stacks/gitea' +22:24:35 [X] stop -> 200 {'ok': True, 'message': 'Stack glance stop completed'} +22:24:41 [X] remove (with drive data) -> 200 {'ok': True, 'data': {'removed': 'glance', 'volumes_removed': ['glance_glance_config'], 'hdd_paths_removed': [], 'hdd_paths_preserved': [], 'hdd_note': 'Az alka +22:24:48 [X] after remove: deployed=False leftovers='/opt/docker/stacks/glance' +22:24:50 [X] stop -> 200 {'ok': True, 'message': 'Stack nextcloud stop completed'} +22:24:55 [X] remove (with drive data) -> 409 {'ok': False, 'error': 'A(z) /mnt/felhom-drives/scratch_hdd/userdata/nextcloud tárhely jelenleg nem elérhető — az alkalmazás nem távolítható el, amíg a meghajtó +22:24:55 [X] refused because the drive path cannot be resolved (R-442, fail-closed and right) — removing the app and KEEPING the drive data instead +22:24:56 [X] remove (keeping drive data) -> 200 {'ok': True, 'data': {'removed': 'nextcloud', 'volumes_removed': ['nextcloud_nextcloud_db_data', 'nextcloud_nextcloud_html', 'nextcloud_nextcloud_redis_data'], +22:25:04 [X] after remove: deployed=False leftovers='/opt/docker/stacks/nextcloud' +22:25:04 [X] stop -> 200 {'ok': True, 'message': 'Stack opengist stop completed'} +22:25:09 [X] remove (with drive data) -> 200 {'ok': True, 'data': {'removed': 'opengist', 'volumes_removed': ['opengist_opengist_data'], 'hdd_paths_removed': [], 'hdd_paths_preserved': [], 'hdd_note': 'Az +22:25:17 [X] after remove: deployed=False leftovers='/opt/docker/stacks/opengist' +22:25:17 [X] stop -> 200 {'ok': True, 'message': 'Stack privatebin stop completed'} +22:25:22 [X] remove (with drive data) -> 200 {'ok': True, 'data': {'removed': 'privatebin', 'volumes_removed': ['privatebin_privatebin_data'], 'hdd_paths_removed': [], 'hdd_paths_preserved': [], 'hdd_note' +22:25:30 [X] after remove: deployed=False leftovers='/opt/docker/stacks/privatebin' +22:25:30 [X] stop -> 200 {'ok': True, 'message': 'Stack uptime-kuma stop completed'} +22:25:35 [X] remove (with drive data) -> 200 {'ok': True, 'data': {'removed': 'uptime-kuma', 'volumes_removed': ['uptime-kuma_uptime_kuma_data'], 'hdd_paths_removed': [], 'hdd_paths_preserved': [], 'hdd_no +22:25:43 [X] after remove: deployed=False leftovers='/opt/docker/stacks/uptime-kuma' +22:25:43 [X] stop -> 200 {'ok': True, 'message': 'Stack vaultwarden stop completed'} +22:25:48 [X] remove (with drive data) -> 200 {'ok': True, 'data': {'removed': 'vaultwarden', 'volumes_removed': ['vaultwarden_vaultwarden_data'], 'hdd_paths_removed': [], 'hdd_paths_preserved': [], 'hdd_no +22:25:56 [X] after remove: deployed=False leftovers='/opt/docker/stacks/vaultwarden' +22:26:06 [X] stop -> 200 {'ok': True, 'message': 'Stack wishlist stop completed'} +22:26:12 [X] remove (with drive data) -> 200 {'ok': True, 'data': {'removed': 'wishlist', 'volumes_removed': ['wishlist_wishlist_data', 'wishlist_wishlist_uploads'], 'hdd_paths_removed': [], 'hdd_paths_pre +22:26:19 [X] after remove: deployed=False leftovers='/opt/docker/stacks/wishlist' +22:26:30 [X] stop -> 200 {'ok': True, 'message': 'Stack zipline stop completed'} +22:26:35 [X] remove (with drive data) -> 200 {'ok': True, 'data': {'removed': 'zipline', 'volumes_removed': ['zipline_zipline_postgres_data', 'zipline_zipline_public', 'zipline_zipline_uploads'], 'hdd_path +22:26:42 [X] after remove: deployed=False leftovers='/opt/docker/stacks/zipline' +22:26:45 -> 01-image-store-removed.txt +22:27:13 -> 02-config-restored.txt +22:27:13 [M] git.repo_url reads back as the LIVE catalog: True +22:27:20 [M] filebrowser deployed=False state=running badge=[] +22:27:20 [M] navidrome deployed=False state=running badge=[] +22:27:20 [M] traefik deployed=False state=running badge=[] +22:27:22 -> 03-containers-after.txt +22:27:22 [M] containers left: ['felhom-controller', 'filebrowser', 'navidrome', 'traefik'] only protected infra: False +22:27:25 -> 04-host-after.txt +22:27:25 [H] no harness LXC was created tonight, so none was destroyed — the PostgreSQL rehearsal ran on 9202 itself (stated in the audit) +22:27:28 -> 05-hub.txt +22:27:28 [U] floor = 0.261.0 | min_agent = 0.131.0 | host rows seen = 3 | nothing was provisioned at the hub tonight: no customer, no config, no appliance, | no binding. The only hub act of the whole night was the floor save in Phase 0.1. | +22:27:30 -> 06-gitea.txt +22:27:30 [G] live main=4463243f2e09 drill main=4463243f2e09 image lines identical: True +22:27:30 teardown written -> /mnt/5_hdd/felhom.eu/git/felhom.eu/documentation/audits/update-night-2026-09-21/teardown/result.json diff --git a/documentation/audits/update-night-2026-09-21/teardown/00-host-before.txt b/documentation/audits/update-night-2026-09-21/teardown/00-host-before.txt new file mode 100644 index 00000000..87354794 --- /dev/null +++ b/documentation/audits/update-night-2026-09-21/teardown/00-host-before.txt @@ -0,0 +1,9 @@ +VMID Status Lock Name +9201 running demo-hp +9202 running demo-hp-scratch +--- pvesm +Name Type Status Total (KiB) Used (KiB) Available (KiB) % +felhom-pbs pbs active 0 0 0 0.00% +local dir active 40453376 32767900 5598360 81.00% +local-lvm lvmthin active 56487936 27018179 29469756 47.83% +nvme-scratch dir active 983379700 51991128 881361960 5.29% diff --git a/documentation/audits/update-night-2026-09-21/teardown/01-image-store-removed.txt b/documentation/audits/update-night-2026-09-21/teardown/01-image-store-removed.txt new file mode 100644 index 00000000..0f42d28b --- /dev/null +++ b/documentation/audits/update-night-2026-09-21/teardown/01-image-store-removed.txt @@ -0,0 +1,65 @@ +=== registry container + volume, removed BY NAME (never a prune) +drill-registry +drill-registry-data +=== drill images, removed BY NAME +localhost:5000/drill/glance:1.0.1 -> removed +localhost:5000/drill/wishes:2.0.0 -> removed +localhost:5000/drill/wishes:2.0.1 -> removed +localhost:5000/drill/glance:1.0.0 -> removed +localhost:5000/drill/glance:1.0.2 -> removed +localhost:5000/drill/status:2.0.0 -> removed +localhost:5000/drill/status:2.0.1 -> removed +localhost:5000/drill/paste:2.0.0 -> removed +localhost:5000/drill/paste:2.0.1 -> removed +localhost:5000/drill/paste:2.0.2 -> removed +localhost:5000/drill/pdf:1.0.0 -> removed +localhost:5000/drill/pdf:2.0.0 -> removed +localhost:5000/drill/pdf:2.0.1 -> removed +localhost:5000/drill/pdf:2.0.2 -> removed +localhost:5000/drill/gist:2.0.0 -> removed +localhost:5000/drill/gist:2.0.1 -> removed +registry:2 removed +=== anything left that says drill? +(none) +(no drill containers) +=== NOTHING WAS PRUNED — this is the full image list, for the record +alpine:3.20 7.81MB +alpine:latest 8.42MB +deluan/navidrome:0.64.0 250MB +docmost/docmost:0.96.0 753MB +ghcr.io/advplyr/audiobookshelf:2.36.1 320MB +ghcr.io/alam00000/bentopdf:v2.8.6 287MB +ghcr.io/cmintey/wishlist:v0.66.0 835MB +ghcr.io/cmintey/wishlist:v0.67.0 900MB +ghcr.io/diced/zipline:4.6.1 817MB +ghcr.io/home-assistant/home-assistant:2026.7.2 2.36GB +ghcr.io/home-assistant/home-assistant:2026.9.3 2.34GB +ghcr.io/mealie-recipes/mealie:v3.20.1 1.18GB +ghcr.io/mealie-recipes/mealie:v3.27.0 1.2GB +ghcr.io/papra-hq/papra:26.6.1-rootless 647MB +ghcr.io/papra-hq/papra:26.6.2-rootless 894MB +ghcr.io/seanmorley15/adventurelog-backend:v0.12.1 1.18GB +ghcr.io/seanmorley15/adventurelog-frontend:v0.12.1 345MB +ghcr.io/tandoorrecipes/recipes:2.6.13 868MB +ghcr.io/tandoorrecipes/recipes:2.6.15 919MB +ghcr.io/thomiceli/opengist:1.13 91.7MB +ghcr.io/thomiceli/opengist:1.15 94.4MB +gitea.dooplex.hu/admin/felhom-controller:0.241.0 423MB +gitea.dooplex.hu/admin/felhom-controller:0.242.0 423MB +gitea.dooplex.hu/admin/felhom-controller:0.243.0 423MB +gitea.dooplex.hu/admin/felhom-controller:0.245.0 423MB +gitea.dooplex.hu/admin/felhom-controller:0.260.0 423MB +gitea.dooplex.hu/admin/felhom-controller:0.261.0 423MB +gitea/gitea:1.27.0 191MB +glanceapp/glance:v0.8.6 24.1MB +grafana/grafana:13.1.0 1.16GB +grafana/grafana:13.2.2 1.4GB +gtstef/filebrowser:1.3.3-stable 215MB +louislam/uptime-kuma:2.5.1 1.71GB +louislam/uptime-kuma:2.5.5 1.72GB +mariadb:11.4 327MB +mariadb:11.6 415MB +mariadb:12.3 334MB +n8nio/n8n:2.31.3 1.53GB +n8nio/n8n:2.40.5 1.04GB +nextcloud:34.0.1-apache 1.44GB diff --git a/documentation/audits/update-night-2026-09-21/teardown/02-config-restored.txt b/documentation/audits/update-night-2026-09-21/teardown/02-config-restored.txt new file mode 100644 index 00000000..0324125b --- /dev/null +++ b/documentation/audits/update-night-2026-09-21/teardown/02-config-restored.txt @@ -0,0 +1,18 @@ +=== restoring controller.yaml from the pre-update-night copy +-rw------- 1 root root 2018 Sep 21 20:18 /var/lib/docker/volumes/felhom-controller-data/_data/controller.yaml +-rw------- 1 root root 1944 Sep 21 18:09 /var/lib/docker/volumes/felhom-controller-data/_data/controller.yaml.pre-update-night +restored +=== the git section, read back (token redacted by this script, not by the box) +git: + branch: main + repo_url: https://gitea.dooplex.hu/admin/app-catalog-felhom.eu.git + sync_interval: 15m + token: "" + username: "" +hub: +=== removing the drill catalog cache so the next sync clones the LIVE repo (R-615) +gitea.dooplex.hu/admin/felhom-controller:0.261.0 Up 25 seconds (healthy) +=== the cache's origin, READ BACK — this is the quote the brief asks for +origin https://gitea.dooplex.hu/admin/app-catalog-felhom.eu.git (fetch) +origin https://gitea.dooplex.hu/admin/app-catalog-felhom.eu.git (push) +4463243 Upgrade harness: four fixtures and seven real upstream edges from the update night (R-462) diff --git a/documentation/audits/update-night-2026-09-21/teardown/03-containers-after.txt b/documentation/audits/update-night-2026-09-21/teardown/03-containers-after.txt new file mode 100644 index 00000000..538e687b --- /dev/null +++ b/documentation/audits/update-night-2026-09-21/teardown/03-containers-after.txt @@ -0,0 +1,4 @@ +felhom-controller gitea.dooplex.hu/admin/felhom-controller:0.261.0 Up 33 seconds (healthy) +filebrowser gtstef/filebrowser:1.3.3-stable Up 7 minutes (healthy) +navidrome deluan/navidrome:0.64.0 Up 7 minutes (healthy) +traefik traefik:v3.6.7 Up 7 minutes diff --git a/documentation/audits/update-night-2026-09-21/teardown/04-host-after.txt b/documentation/audits/update-night-2026-09-21/teardown/04-host-after.txt new file mode 100644 index 00000000..e71d68c8 --- /dev/null +++ b/documentation/audits/update-night-2026-09-21/teardown/04-host-after.txt @@ -0,0 +1,30 @@ +VMID Status Lock Name +9201 running demo-hp +9202 running demo-hp-scratch +--- pvesm +Name Type Status Total (KiB) Used (KiB) Available (KiB) % +felhom-pbs pbs active 0 0 0 0.00% +local dir active 40453376 32767980 5598280 81.00% +local-lvm lvmthin active 56487936 27018179 29469756 47.83% +nvme-scratch dir active 983379700 51059184 882293904 5.19% +--- 9201 untouched: +adventurelog +adventurelog-frontend +adventurelog-postgres +bentopdf +bookstack +bookstack-db +calibre-web +cloudflared +docmost +docmost-postgres +docmost-redis +felhom-controller +filebrowser +kimai +kimai-db +opengist +paperless-postgres +paperless-redis +paperless-webserver +privatebin diff --git a/documentation/audits/update-night-2026-09-21/teardown/05-hub.txt b/documentation/audits/update-night-2026-09-21/teardown/05-hub.txt new file mode 100644 index 00000000..5559fad7 --- /dev/null +++ b/documentation/audits/update-night-2026-09-21/teardown/05-hub.txt @@ -0,0 +1,5 @@ +floor = 0.261.0 +min_agent = 0.131.0 +host rows seen = 3 +nothing was provisioned at the hub tonight: no customer, no config, no appliance, +no binding. The only hub act of the whole night was the floor save in Phase 0.1. diff --git a/documentation/audits/update-night-2026-09-21/teardown/06-gitea.txt b/documentation/audits/update-night-2026-09-21/teardown/06-gitea.txt new file mode 100644 index 00000000..aeeb02b1 --- /dev/null +++ b/documentation/audits/update-night-2026-09-21/teardown/06-gitea.txt @@ -0,0 +1,7 @@ +live catalog origin/main : 4463243f2e09 +drill repo HEAD after reset: 4463243f2e09 +reset+force-push rc=0 + +diff of every `image:` line, live vs drill: +IDENTICAL — every image: line matches the live catalog + diff --git a/documentation/audits/update-night-2026-09-21/teardown/log.txt b/documentation/audits/update-night-2026-09-21/teardown/log.txt new file mode 100644 index 00000000..f9f0db42 --- /dev/null +++ b/documentation/audits/update-night-2026-09-21/teardown/log.txt @@ -0,0 +1,52 @@ +22:23:52 ==== Phase 5: teardown, three layers +22:23:55 -> 00-host-before.txt +22:23:55 [M] throwaway apps still deployed: ['bentopdf', 'docmost', 'gitea', 'glance', 'nextcloud', 'opengist', 'privatebin', 'uptime-kuma', 'vaultwarden', 'wishlist', 'zipline'] +22:23:55 [X] stop -> 200 {'ok': True, 'message': 'Stack bentopdf stop completed'} +22:24:01 [X] remove (with drive data) -> 200 {'ok': True, 'data': {'removed': 'bentopdf', 'volumes_removed': [], 'hdd_paths_removed': [], 'hdd_paths_preserved': [], 'hdd_note': 'Az alkalmazás nem tárolt sa +22:24:08 [X] after remove: deployed=False leftovers='/opt/docker/stacks/bentopdf' +22:24:09 [X] stop -> 200 {'ok': True, 'message': 'Stack docmost stop completed'} +22:24:15 [X] remove (with drive data) -> 200 {'ok': True, 'data': {'removed': 'docmost', 'volumes_removed': ['docmost_docmost_postgres_data', 'docmost_docmost_redis_data', 'docmost_docmost_storage'], 'hdd_ +22:24:22 [X] after remove: deployed=False leftovers='/opt/docker/stacks/docmost' +22:24:22 [X] stop -> 200 {'ok': True, 'message': 'Stack gitea stop completed'} +22:24:28 [X] remove (with drive data) -> 200 {'ok': True, 'data': {'removed': 'gitea', 'volumes_removed': ['gitea_gitea_data'], 'hdd_paths_removed': [], 'hdd_paths_preserved': [], 'hdd_note': 'Az alkalmazá +22:24:35 [X] after remove: deployed=False leftovers='/opt/docker/stacks/gitea' +22:24:35 [X] stop -> 200 {'ok': True, 'message': 'Stack glance stop completed'} +22:24:41 [X] remove (with drive data) -> 200 {'ok': True, 'data': {'removed': 'glance', 'volumes_removed': ['glance_glance_config'], 'hdd_paths_removed': [], 'hdd_paths_preserved': [], 'hdd_note': 'Az alka +22:24:48 [X] after remove: deployed=False leftovers='/opt/docker/stacks/glance' +22:24:50 [X] stop -> 200 {'ok': True, 'message': 'Stack nextcloud stop completed'} +22:24:55 [X] remove (with drive data) -> 409 {'ok': False, 'error': 'A(z) /mnt/felhom-drives/scratch_hdd/userdata/nextcloud tárhely jelenleg nem elérhető — az alkalmazás nem távolítható el, amíg a meghajtó +22:24:55 [X] refused because the drive path cannot be resolved (R-442, fail-closed and right) — removing the app and KEEPING the drive data instead +22:24:56 [X] remove (keeping drive data) -> 200 {'ok': True, 'data': {'removed': 'nextcloud', 'volumes_removed': ['nextcloud_nextcloud_db_data', 'nextcloud_nextcloud_html', 'nextcloud_nextcloud_redis_data'], +22:25:04 [X] after remove: deployed=False leftovers='/opt/docker/stacks/nextcloud' +22:25:04 [X] stop -> 200 {'ok': True, 'message': 'Stack opengist stop completed'} +22:25:09 [X] remove (with drive data) -> 200 {'ok': True, 'data': {'removed': 'opengist', 'volumes_removed': ['opengist_opengist_data'], 'hdd_paths_removed': [], 'hdd_paths_preserved': [], 'hdd_note': 'Az +22:25:17 [X] after remove: deployed=False leftovers='/opt/docker/stacks/opengist' +22:25:17 [X] stop -> 200 {'ok': True, 'message': 'Stack privatebin stop completed'} +22:25:22 [X] remove (with drive data) -> 200 {'ok': True, 'data': {'removed': 'privatebin', 'volumes_removed': ['privatebin_privatebin_data'], 'hdd_paths_removed': [], 'hdd_paths_preserved': [], 'hdd_note' +22:25:30 [X] after remove: deployed=False leftovers='/opt/docker/stacks/privatebin' +22:25:30 [X] stop -> 200 {'ok': True, 'message': 'Stack uptime-kuma stop completed'} +22:25:35 [X] remove (with drive data) -> 200 {'ok': True, 'data': {'removed': 'uptime-kuma', 'volumes_removed': ['uptime-kuma_uptime_kuma_data'], 'hdd_paths_removed': [], 'hdd_paths_preserved': [], 'hdd_no +22:25:43 [X] after remove: deployed=False leftovers='/opt/docker/stacks/uptime-kuma' +22:25:43 [X] stop -> 200 {'ok': True, 'message': 'Stack vaultwarden stop completed'} +22:25:48 [X] remove (with drive data) -> 200 {'ok': True, 'data': {'removed': 'vaultwarden', 'volumes_removed': ['vaultwarden_vaultwarden_data'], 'hdd_paths_removed': [], 'hdd_paths_preserved': [], 'hdd_no +22:25:56 [X] after remove: deployed=False leftovers='/opt/docker/stacks/vaultwarden' +22:26:06 [X] stop -> 200 {'ok': True, 'message': 'Stack wishlist stop completed'} +22:26:12 [X] remove (with drive data) -> 200 {'ok': True, 'data': {'removed': 'wishlist', 'volumes_removed': ['wishlist_wishlist_data', 'wishlist_wishlist_uploads'], 'hdd_paths_removed': [], 'hdd_paths_pre +22:26:19 [X] after remove: deployed=False leftovers='/opt/docker/stacks/wishlist' +22:26:30 [X] stop -> 200 {'ok': True, 'message': 'Stack zipline stop completed'} +22:26:35 [X] remove (with drive data) -> 200 {'ok': True, 'data': {'removed': 'zipline', 'volumes_removed': ['zipline_zipline_postgres_data', 'zipline_zipline_public', 'zipline_zipline_uploads'], 'hdd_path +22:26:42 [X] after remove: deployed=False leftovers='/opt/docker/stacks/zipline' +22:26:45 -> 01-image-store-removed.txt +22:27:13 -> 02-config-restored.txt +22:27:13 [M] git.repo_url reads back as the LIVE catalog: True +22:27:20 [M] filebrowser deployed=False state=running badge=[] +22:27:20 [M] navidrome deployed=False state=running badge=[] +22:27:20 [M] traefik deployed=False state=running badge=[] +22:27:22 -> 03-containers-after.txt +22:27:22 [M] containers left: ['felhom-controller', 'filebrowser', 'navidrome', 'traefik'] only protected infra: False +22:27:25 -> 04-host-after.txt +22:27:25 [H] no harness LXC was created tonight, so none was destroyed — the PostgreSQL rehearsal ran on 9202 itself (stated in the audit) +22:27:28 -> 05-hub.txt +22:27:28 [U] floor = 0.261.0 | min_agent = 0.131.0 | host rows seen = 3 | nothing was provisioned at the hub tonight: no customer, no config, no appliance, | no binding. The only hub act of the whole night was the floor save in Phase 0.1. | +22:27:30 -> 06-gitea.txt +22:27:30 [G] live main=4463243f2e09 drill main=4463243f2e09 image lines identical: True diff --git a/documentation/audits/update-night-2026-09-21/teardown/result.json b/documentation/audits/update-night-2026-09-21/teardown/result.json new file mode 100644 index 00000000..cfafba9a --- /dev/null +++ b/documentation/audits/update-night-2026-09-21/teardown/result.json @@ -0,0 +1,43 @@ +{ + "host_before": "VMID Status Lock Name \n9201 running demo-hp \n9202 running demo-hp-scratch \n--- pvesm\nName Type Status Total (KiB) Used (KiB) Available (KiB) %\nfelhom-pbs pbs active 0 0 0 0.00%\nlocal dir active 40453376 32767900 5598360 81.00%\nlocal-lvm lvmthin active 56487936 27018179 29469756 47.83%\nnvme-scratch dir active 983379700 51991128 881361960 5.29%\n", + "apps_removed": { + "bentopdf": "200", + "docmost": "200", + "gitea": "200", + "glance": "200", + "nextcloud": "200", + "opengist": "200", + "privatebin": "200", + "uptime-kuma": "200", + "vaultwarden": "200", + "wishlist": "200", + "zipline": "200" + }, + "repo_url_read_back": true, + "still_on_the_box": [ + { + "name": "filebrowser", + "state": "running", + "deployed": false, + "badge_hu": [] + }, + { + "name": "navidrome", + "state": "running", + "deployed": false, + "badge_hu": [] + }, + { + "name": "traefik", + "state": "running", + "deployed": false, + "badge_hu": [] + } + ], + "only_protected_infra_left": false, + "host_after": "VMID Status Lock Name \n9201 running demo-hp \n9202 running demo-hp-scratch \n--- pvesm\nName Type Status Total (KiB) Used (KiB) Available (KiB) %\nfelhom-pbs pbs active 0 0 0 0.00%\nlocal dir active 40453376 32767980 5598280 81.00%\nlocal-lvm lvmthin active 56487936 27018179 29469756 47.83%\nnvme-scratch dir active 983379700 51059184 882293904 5.19%\n--- 9201 untouched:\nadventurelog\nadventurelog-frontend\nadventurelog-postgres\nbentopdf\nbookstack\nbookstack-db\ncalibre-web\ncloudflared\ndocmost\ndocmost-postgres\ndocmost-redis\nfelhom-controller\nfilebrowser\nkimai\nkimai-db\nopengist\npaperless-postgres\npaperless-redis\npaperless-webserver\nprivatebin\n", + "hub": "floor = 0.261.0\nmin_agent = 0.131.0\nhost rows seen = 3\nnothing was provisioned at the hub tonight: no customer, no config, no appliance,\nno binding. The only hub act of the whole night was the floor save in Phase 0.1.\n", + "live_main": "4463243f2e09", + "drill_main": "4463243f2e09", + "image_lines_identical": true +} \ No newline at end of file diff --git a/documentation/backlog/OPEN-ITEMS.md b/documentation/backlog/OPEN-ITEMS.md index 68d46087..5a1d9a41 100644 --- a/documentation/backlog/OPEN-ITEMS.md +++ b/documentation/backlog/OPEN-ITEMS.md @@ -678,15 +678,19 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server` | **R-440** | **[P2-MEDIUM] 23 catalog image pins float, so an update is not reproducible.** `compose pull` on a moving tag fetches whatever upstream published that day. **MEASURED 2026-09-01 over `app-catalog-felhom.eu` @ `29edad9c5bf4`: 79 `image:` lines across 53 apps, 66 distinct; 23 of those lines carry a tag with no patch version.** `postgres:16-alpine` (8 apps), `redis:7-alpine` (6), `mariadb:11.6` (2), plus one each of `postgres:15-alpine`, `postgis/postgis:16-3.5-alpine`, `mariadb:11.4`, `mariadb:12.3`, `ghcr.io/claperco/claper:2.5`, `ghcr.io/thomiceli/opengist:1.13`, `wger/server:2.6`. **A 24th is arguable and is recorded rather than rounded away:** `ghcr.io/immich-app/postgres:16-vectorchord0.4.3-pgvectors0.2.0` pins both extensions exactly but leaves the PostgreSQL patch floating. A customer pressing Frissites can therefore swap their DATABASE ENGINE build with no catalog change and no record; two boxes updated on two days end up different. **Severity MEDIUM on its own; it becomes BLOCKING the moment a pre-update copy exists**, because "what did we upgrade from and to" must be recordable and today it is not — which is also why R-440 must be read next to the digest discipline in Rule 10 of the spike. **MEASURED LIVE 2026-09-01 — the floating pins have ALREADY moved, with a passing control.** Running digests on demo-hp compared against what the registry serves for the same tag today: **`mariadb:11.4` MOVED** (`sha256:4f1d8d20...` -> `sha256:611a2fcc...`) and **`mariadb:12.3` MOVED** (`sha256:a02fe89c...` -> `sha256:dd9b303a...`), while `postgres:16-alpine`, `redis:7-alpine`, `mariadb:11.6` and `opengist:1.13` were SAME — **and both fully-pinned CONTROLS (`rommapp/romm:5.0.0`, `privatebin/pdo:2.0.5`) were SAME.** So on a box with ZERO visible drift by tag, pressing Frissites today silently swaps the DATABASE ENGINE build under `romm` and `bookstack`, with no catalog change and no record. **Compounding fact found while reading:** the recovery unit records `ImagePins` but the manifest comment says *"image NOT stored - re-pulled on restore"*, so a RESTORE of a floating-pinned app also re-pulls whatever is current — the same non-reproducibility on the recovery path. **HALF OF THE ANSWER SHIPPED 2026-09-02 (controller v0.233.0, slice 1): `app.yaml.installed_images` now records, per compose SERVICE, the reference AND the repo digest each container was actually created from — so "what did we upgrade FROM" is answerable on any box that has taken one lifecycle action since the upgrade.** What is still missing is the other half: comparing that digest against what the registry serves for the same tag TODAY, which needs a network call the render path deliberately does not make (see R-446). **The row therefore stays OPEN and its rank is unchanged** — recording a digest does not make a floating pin reproducible; it makes the drift measurable after the fact. `audits/SPIKE-app-update-2026-09-01.md` | **OPEN — rank P2-MEDIUM; owner: CC** | | **R-444** | **[P3-LOW] Nothing runs `pct fstrim` on the fleet, and demo-hp's thin pool was carrying ~23.8 GB of blocks the guest had already freed.** MEASURED 2026-09-01 during this spike's teardown: the run itself added ~1.05 GiB that `local-lvm` did not reclaim on delete (68.97% -> 70.91%); `fstrim` INSIDE the unprivileged container is refused (`FITRIM ioctl failed: Operation not permitted`, all three mounts); `pct fstrim 9201` from the PVE host then trimmed **30.2 GiB + 57 GiB** and took `local-lvm` to **26.78%** — **23.8 GB BELOW this run's own starting point**, i.e. the surplus was long-standing, not ours. **Why it is not merely housekeeping:** a thin pool that only ever grows can reach 100% from DELETED data alone, and a full thin pool takes every guest on the host read-only. demo-hp had 16.4 GB free before the trim. **Not urgent, and the row says so** — but the appliance has no periodic trim and no operator surface reports the gap between guest-free and pool-used. Owner: **CC.** `audits/SPIKE-app-update-2026-09-01.md` | **OPEN — rank P3-LOW; owner: CC** | | **R-445** | **[P3-LOW] Hub app telemetry survives the app's removal, so a 15-minute throwaway now sets a FLEET-WIDE memory recommendation.** MEASURED 2026-09-01: this spike's Phase 6 Nextcloud existed for ~15 minutes on demo-hp, spent part of it crash-looping, and was then removed with all volumes. The hub's `/apps/nextcloud` page still reports `Deployments`, `Avg Memory 208 MB`, `P95 Memory 280 MB` and **`Suggested Limit (P95x1.2) = 352 MB`**, plus three MariaDB `io_uring` rows under Known Issues attributed to demo-hp. **The suggested limit is an operator-facing recommendation derived from a sample that no longer exists anywhere** — and Nextcloud is a real catalog app whose limit someone may act on. **RETAINED DELIBERATELY BY THIS RUN, NOT CLEARED, and the reason is part of the row:** the hub offers `POST /apps/nextcloud/reset-telemetry` whose own confirm reads *"Delete all telemetry data for nextcloud? This cannot be undone."* — an irreversible write on the operator's surface, and the operator authorised Phase 6, not this. **The one-line command is recorded in the audit doc so it is a decision, not a task.** The general question is the row: should telemetry for an app with zero live deployments age out, or be excluded from the suggestion? Owner: **VIKTOR rules, CC implements.** `audits/SPIKE-app-update-2026-09-01.md` | **OPEN — rank P3-LOW; owner: VIKTOR rules, CC implements** | -| **R-446** | **[P2-MEDIUM] „Naprakész" can be FALSE, and the badge that says it cannot tell.** Slice 2 (controller v0.233.0, 2026-09-02) compares the RECORDED image reference per compose service against the reference the current template pins, and **queries no registry** — deliberately: a customer's box must not depend on reaching eight upstream registries to render a page (`felhom-controller/controller/internal/web/updatebadge.go`, `compareInstalledToTemplate`). **For the 23 floating pins that comparison is blind by construction:** `postgres:16-alpine`, `mariadb:11.6` and 21 others can carry an identical reference over an image that has moved. **MEASURED, not theorised — spike §5 found `mariadb:11.4` and `mariadb:12.3` had BOTH already moved upstream while two fully-pinned CONTROLS held.** So `romm` and `bookstack` on demo-hp would read „Naprakész" over a database engine build that is not the one the catalog now resolves to. **This is a KNOWN LIMITATION OF A SHIPPED FEATURE, filed the same session rather than left implicit**, and it is stated in the same words in `architecture/09-update-architecture.md` §8.1 and in the controller's `README.md`. The close is a digest comparison against the registry, which needs a network call, a cache and a failure posture — it is not a one-liner and it is not slice 2's job. **Depends on R-440**, whose fix (stop floating) would remove the problem instead of measuring it — take that route first if it is available. `architecture/09-update-architecture.md` **MEASURED 2026-09-21, and the blind spot is not one or two pins.** `audits/UPDATE-ARC-STATE-2026-09-21.md` §3.3: the catalog carries **10 floating pins of 66** (recounted — the old "23" was stale), and **6 of the 7 measurable engine pins have been repushed upstream since the catalog set them** — `postgres:16-alpine` (8 apps), `postgres:15-alpine`, `redis:7-alpine` (6 apps), `mariadb:11.4`, `mariadb:12.3`, `postgis:16-3.5-alpine`; only `mariadb:11.6` has not. The 8th (immich's own ghcr build) is UNMEASURED — ghcr exposes no anonymous last-modified timestamp. **So on demo-hp today four apps read „Naprakész" over a database engine image that has demonstrably moved.** The fix does NOT need the box to query a registry: the catalog can record each pin's digest at push time (`check-image-resolvable.py` already resolves it) and the box compares digests. Put to the operator as `09` §3b **Q6**, recommended YES — the cheapest real improvement on the arc's list. | **WAITING-ON-OPERATOR — `09` §3b Q6; owner: CC once answered** | +| **R-446** | **[P2-MEDIUM] „Naprakész" can be FALSE, and the badge that says it cannot tell.** Slice 2 (controller v0.233.0, 2026-09-02) compares the RECORDED image reference per compose service against the reference the current template pins, and **queries no registry** — deliberately: a customer's box must not depend on reaching eight upstream registries to render a page (`felhom-controller/controller/internal/web/updatebadge.go`, `compareInstalledToTemplate`). **For the 23 floating pins that comparison is blind by construction:** `postgres:16-alpine`, `mariadb:11.6` and 21 others can carry an identical reference over an image that has moved. **MEASURED, not theorised — spike §5 found `mariadb:11.4` and `mariadb:12.3` had BOTH already moved upstream while two fully-pinned CONTROLS held.** So `romm` and `bookstack` on demo-hp would read „Naprakész" over a database engine build that is not the one the catalog now resolves to. **This is a KNOWN LIMITATION OF A SHIPPED FEATURE, filed the same session rather than left implicit**, and it is stated in the same words in `architecture/09-update-architecture.md` §8.1 and in the controller's `README.md`. The close is a digest comparison against the registry, which needs a network call, a cache and a failure posture — it is not a one-liner and it is not slice 2's job. **Depends on R-440**, whose fix (stop floating) would remove the problem instead of measuring it — take that route first if it is available. `architecture/09-update-architecture.md` **MEASURED 2026-09-21, and the blind spot is not one or two pins.** `audits/UPDATE-ARC-STATE-2026-09-21.md` §3.3: the catalog carries **10 floating pins of 66** (recounted — the old "23" was stale), and **6 of the 7 measurable engine pins have been repushed upstream since the catalog set them** — `postgres:16-alpine` (8 apps), `postgres:15-alpine`, `redis:7-alpine` (6 apps), `mariadb:11.4`, `mariadb:12.3`, `postgis:16-3.5-alpine`; only `mariadb:11.6` has not. The 8th (immich's own ghcr build) is UNMEASURED — ghcr exposes no anonymous last-modified timestamp. **So on demo-hp today four apps read „Naprakész" over a database engine image that has demonstrably moved.** The fix does NOT need the box to query a registry: the catalog can record each pin's digest at push time (`check-image-resolvable.py` already resolves it) and the box compares digests. Put to the operator as `09` §3b **Q6**, recommended YES — the cheapest real improvement on the arc's list. **— UPDATE NIGHT 2026-09-21:** **MEASURED ON A BOX 2026-09-21 (update night, leg B8), and it REFINES the row in two ways rather than merely confirming it.** §8.1's numbers came from a registry sweep on DooPlex; this is the same question asked of a customer-shaped box, where the badge actually renders. On guest 9202, `docmost`'s two floating pins were read as `installed_images` records them and compared against the upstream digests measured the same night: `postgres:16-alpine` → **`sha256:721873c34ceb9…` on the box and `sha256:721873c34ceb9…` upstream**, and `redis:7-alpine` → **`sha256:858f009f9709c…` both sides**. **Identical. So the badge „Naprakész" is TRUE for this box**, and the app reads correctly. **(1) The defect's size is set by INSTALL AGE, not by the catalog.** A floating pin is wrong only for a box that pulled BEFORE the tag moved; a box deployed after the repush holds the current image and its badge is right. R-446's "six repushed pins" measured the tag against the date the CATALOG set it, which is the right measure for *the catalog* and not for *a box*. **(2) The producer Q6 needs ALREADY EXISTS on the box.** `installed_images` records a real `digest` per service (`installed.go` §7.1) — the box knows exactly what it is running. What it cannot do is COMPARE, because the catalog carries no digest to compare against. That is Q6's proposal, and this is a concrete confirmation that only the catalog half is missing. Evidence: `audits/update-night-2026-09-21/23-B8-floating-pin.txt`. + +| **R-626** | **[P2-MEDIUM] An app the customer REMOVED came back: the removal returned 200 and deleted the record and the volume, a container was created two seconds later, and Docker's restart policy has kept it running ever since — while the box reports the app as not installed.** FOUND 2026-09-21 during the update night's TEARDOWN, which is the only reason it was found at all. `navidrome` was removed through the product: the `remove_hdd_data:true` call was correctly refused `409` (R-442's fail-closed guard — the drive path could not be resolved on this guest), and the `remove_hdd_data:false` call returned **200** with `volumes_removed: ['navidrome_navidrome_data']`. **Two seconds later a container carrying `com.docker.compose.project=navidrome` was CREATED** (`.Created = 19:12:20Z`), and Docker's `restart: unless-stopped` started it again at the next guest boot (`.StartedAt = 20:19:33Z`, the B5 power cut). Seven hours later: **`app.yaml` absent, `deployed=false`, one volume back, and the controller happily probing it — `Health probe navidrome: API GET :4533/ping → 200`.** **The customer-visible shape is the bad one:** *"I deleted that app and it came back"* — and it came back **blank**, because the volume really was deleted, so it looks installed and is empty. It is also invisible to every sweep that keys on `deployed`, which is exactly why the teardown found it and nothing else did. **WHAT IS NOT ESTABLISHED, and is stated rather than guessed: what created the container.** The controller was restarted several times later in the night and its log no longer reaches that moment — **the second time in one night that a restart destroyed the evidence of the thing that mattered** (see R-621). **Needs:** reproduce with a loop that removes an app and watches `docker events` for 60 s, so the creating path is a NAME and not an inference; then a test that removes an app, reboots, and asserts no container with that compose project exists. **And one instrument lesson worth keeping:** this session's own post-remove check queried the compose-project label and reported clean at 21:12:18 — two seconds before the container appeared. A check that runs once, immediately, cannot see a thing that is created immediately after it. Evidence: `audits/update-night-2026-09-21/26-removed-app-came-back.txt`. | **READY — rank P2-MEDIUM; owner: CC (controller)** | | **WAITING-ON-OPERATOR — `09` §3b Q6; owner: CC once answered** | | **R-450** | **[P2-MEDIUM] UPDATE ARC SLICE 6 — a version sequence: automatic WITHIN a major, never ACROSS one, and an engine change gets its OWN edge.** The first half is an operator ruling of 2026-09-02 and its justification is R-449's measurement: a cross-major jump can be refused by the app itself and cannot be undone. **The second half is a rule recorded now, while it is cheap:** an engine change must never be bundled with an app version bump. `bookstack`'s `0b73e5e` moved the application 25.02.2 → 26.05.2 **and** MariaDB 11.6 → 12.3 in one commit — **two migrations behind one edge**, and an unreadable failure when it breaks. Needs a catalog-side convention and, eventually, a gate. `architecture/09-update-architecture.md` §6 **HALF SHIPPED 2026-09-21 (catalog `5ff36d098cbc`): the second half — an engine change gets its OWN edge — is now ENFORCED** by `check-engine-major.py`, which refuses a commit moving a MariaDB major together with any other image move in that template, naming what it was bundled with. The FIRST half (automatic within a major) is Slice 6 and needs four operator answers — `09` §3b **Q1–Q4**, with the shape it would take in `09` §6.2. **The urgency is now measured:** 46 of the catalog's 58 exact pins are behind upstream and **39 of those are within a major** — the population the 2026-09-02 ruling already says may move without a human. | **WAITING-ON-OPERATOR — `09` §3b Q1–Q4 (own-edge half SHIPPED); owner: CC once answered** | | **R-451** | **[P3-LOW] UPDATE ARC SLICE 7 — a fleet sweep: the operator can SEE, and MOVE, how far behind every box is.** Slices 1 and 2 make one box's state visible on that box's own pages. The operator has no fleet view, and **it is not derivable from what is already reported: the hub's report payload carries container name, state, CPU and memory, and NO image field at all** (spike §5, which is why Peti's box could only be recorded UNKNOWN). So this is a hub-side change as well as a controller one. Rank LOW today because the fleet is two enrolled boxes; it rises with the fleet. `architecture/09-update-architecture.md` §6, §8.4 **BOTH SIDES VERIFIED 2026-09-21, and it is cheaper than this row implies.** The controller's payload carries no image (`internal/report/types.go` L98–103) and the hub's `Store.SaveReport` (`hub/internal/store/store.go:965`) denormalises only container **counts** — but **the hub stores the raw report JSON whole**, so a new controller field lands there the day it is sent. What is missing is the denormalisation and the page, not the transport. Shape in `09` §6.2–6.3; the payload question is `09` §3b **Q7**. | **WAITING-ON-OPERATOR — `09` §3b Q7; owner: CC once answered** | | **R-454** | **[P3-LOW] Five `internal/web` test files have been `gofmt`-unclean for an unknown length of time, and nothing notices.** MEASURED 2026-09-02: `gofmt -l controller/internal/web/` reports `backups_split_test.go`, `claim_code_naming_test.go`, `disk_health_test.go`, `r400_debug_routes_test.go`, `recovery_test.go` — at the **baseline** commit `960d29b0612c`, i.e. not introduced by v0.233.0 (both files added that day are clean). **`go vet` does not check formatting and `controller_gates.py` has no formatting gate**, so the only thing that would ever surface this is someone running `gofmt -l` by hand, which is how it was found. **Not reformatted in the same session, deliberately** — the minimal-changes rule, and a five-file whitespace commit inside a feature release makes that release's diff unreadable. **Small, and the cost of NOT having the instrument is the row:** the count can only grow, and every future `gofmt -l` run produces noise that hides a real one. Fix is two lines: a `gofmt -l` gate in `controller_gates.py` plus one formatting commit, in that order (the gate first, so the commit is provably complete). Owner: **CC.** | **READY — rank P3-LOW; owner: CC** | | **R-457** | **[P3-LOW] A test that hardcodes a date AND asserts an age derived from it is green on the day it is written and red the next morning — one instance PROVEN, six candidate files named.** MEASURED 2026-09-03: `TestGroupD_BadgeRendersOnBothSurfaces` (shipped the previous day in v0.233.0) pinned a fixture `catalog_since: "2026-07-18"` and asserted the rendered string `"Frissítés elérhető — 46 napja"`. **The pure badge tests inject a clock; the RENDER test does not and cannot** — it goes through the production templates, which call the funcmap entry `updateBadge`, which reads `time.Now()`. The suite was green on 2026-09-02 and **FAILED on 2026-09-03** with *"the behind badge is missing"* on both surfaces, because the true answer had become 47. **Fixed by DERIVING the fixture** — `catalog_since` is computed as *today minus 46 days*, so the test asserts the real number through the real clock and cannot rot. **THE CLASS, which is why this is a row and not just a fix:** a clock-reading test that also carries a date LITERAL is a bomb with a fuse of unknown length, and the suite being green is not evidence it is defused — it is evidence the fuse has not burned down yet. **NAMED AS UNCHECKED CANDIDATES, NOT ACCUSED** — six other test files contain both a `20xx-xx-xx` literal and `time.Now()`: `internal/backup/offbox_test.go`, `internal/web/handler_export_upload_test.go`, `internal/web/r103_tier2_action_test.go`, `internal/web/dashboard_backup_card_test.go`, `internal/web/async_restore_test.go`, `internal/stacks/installed_test.go`. Mixing the two is not itself a defect — it is one only where a literal feeds an assertion evaluated against the real clock — so each needs reading, which is a sweep and not this session. **The instrument that would end the class:** run the suite once under a faked future date in CI and see what turns red. Owner: **CC.** `felhom-controller` v0.234.0 CHANGELOG | **READY — rank P3-LOW; owner: CC** | -| **R-458** | **[P3-LOW] `.felhom.yml` keeps flowing to an app whose compose file is FROZEN, so a frozen app can receive a health check written for a version it is not running.** The v0.235.0 render freezes `docker-compose.yml` for a pinned app once the catalog moves past its version, but copies `.felhom.yml` **verbatim in every case** (`Syncer.copyTemplates`). **The asymmetry is deliberate and both directions were considered:** `.felhom.yml` carries no image, and it carries `catalog_since` — the single input the update badge uses to say *„Frissítés elérhető — N napja"* — so freezing it would silently withhold the one number that tells a customer they are behind, i.e. it would break slice 2 to protect slice 3. **What it costs:** the file also carries the controller-side `healthcheck:` block and resource hints, so a template updated for a newer version can hand a frozen app a probe written for software it is not running. **THE FAILURE DIRECTION IS A FALSE ALARM, NEVER DATA LOSS** — the app keeps running; at worst it renders as degraded and, if it persisted, could reach the dead-app alarm path. That is the same class as R-330's false e-mails, which is why this is a row and not a footnote. **Not fixed now, and the reason is that the cheap fix is wrong:** freezing the whole file breaks the badge, and freezing only the `healthcheck:` key means the syncer would have to parse and re-assemble a customer-facing metadata file — new surface on the one path that touches every app on every box every 15 minutes. **What would settle it:** whether any catalog `healthcheck:` has ever been changed in the same commit as an `image:` line (measurable from the catalog's own history, no box needed). If the answer is "never", the exposure is theoretical and the row can be closed by measurement instead of by code. Owner: **CC.** `architecture/09-update-architecture.md` §5.4, §8.5 | **READY — rank P3-LOW; owner: CC** | -| **R-460** | **[P3-LOW] BookStack's FILE half cannot be seeded or verified without a browser, so its upgrades can only ever be auto-proven for the DATABASE.** MEASURED 2026-09-06 while building the R-449 harness. BookStack's API needs a token that is only mintable through its web UI, and its HTTP login is unusable headlessly for a second, independent reason: `APP_URL` comes from the template as `https://${SUBDOMAIN}.${DOMAIN}`, so the app marks its session and XSRF cookies **`secure`**; curl over plain http stores neither and **every login POST returns 419 Page Expired**, which looks exactly like a wrong password. The container serves no TLS. **The database half IS provable** — the harness seeds with `php artisan bookstack:create-admin` and reads back with a DIFFERENT artisan command that must find the record, carrying its own negative control on every call. **What is unprovable is an uploaded image or attachment**, i.e. exactly the half a customer would notice. **THIS IS A FACT ABOUT THE APP, NOT A DEFECT IN THE HARNESS**, and it is recorded because Slice 6 needs to know which apps can be auto-verified and which can only be partly verified — nobody had that list before. **Deliberately NOT worked around:** planting a file in the volume would make the test pass while proving nothing, which is R-156's exact failure. **What would remove it:** a headless token route (upstream), or accepting a browser-driven step for this app alone, which DooPlex cannot run. Owner: **CC.** `audits/SPIKE-upgrade-test-2026-09-06.md` §6 | **READY — rank P3-LOW; owner: CC** | -| **R-462** | **[P2-MEDIUM] Widen the upgrade harness beyond three apps — and the cost is dominated by FIXTURES, not by machine time.** The R-449 harness works and is proven by a red negative control (`audits/SPIKE-upgrade-test-2026-09-06.md` §1). **Costed with this run's REAL numbers rather than an estimate:** a successful edge takes **6.4 s – 305.1 s, median 71.8 s**; a FAILING edge takes **556 s**, roughly **8×**, because a negative is only honest if it waits out the full settle window; 3 apps / 11 images cost **5.07 GB**, so 53 apps naively extrapolate to **~90 GB** and, at the median, about an hour of harness time for one edge each. **THAT EXTRAPOLATION UNDERSTATES THE REAL COST BY AN ORDER OF MAGNITUDE, and that is the point of this row.** Two of the three apps needed a bespoke non-browser seed route; one needed two attempts and a discarded approach; one (bookstack) can only ever be half-proven (R-460). **Fixture time scales with apps and does not amortise.** **The decision this row is really asking for is scope, not schedule:** all 53, or only the apps a customer would lose data from, or only apps whose catalog transition is a MAJOR. **Recommended shape, NOT a design — the operator picks:** start with the apps that carry a database, because §3 measured that the abort question only ever bites there. Owner: **VIKTOR rules on scope, CC implements.** `audits/SPIKE-upgrade-test-2026-09-06.md` §5 **ROW CORRECTED 2026-09-21: the scope is NOT open and this row said it was.** It read *"VIKTOR rules on scope"*; the operator ruled on 2026-09-13 (`09` §3 decision 6) that the upgrade test goes to **ALL** apps through the nightly rotation, explicitly *not* "database apps first". What is open is the WORK, not the scope. An ORDER inside that ruling — the 15 database services first, because that is where a wrong answer costs data rather than uptime — is costed as a drill brief in `09` §6.4: legs A–E ≈ **21–34 CC-hours** plus ~25–30 GB on a scratch host, with legs C (a PostgreSQL `pg_upgrade` rehearsal) and E (one automatic night on a throwaway) the two that unblock a decision. | **READY — rank P2-MEDIUM; owner: CC (scope already ruled, `09` §3 decision 6)** | -| **R-463** | **[P2-MEDIUM] The day the catalog moves `postgres:16` to `17`, ELEVEN apps are affected and the container image will NOT perform the conversion — and nothing anywhere records that.** MEASURED 2026-09-06: 11 of the 53 templates carry PostgreSQL — **8 on `postgres:16-alpine`**, 1 on `postgres:15-alpine`, plus `postgis/postgis:16-3.5-alpine` and Immich's own `postgres:16-vectorchord…` build. **A grep of the whole register for `pg_upgrade`, "postgres major" or "postgresql major" returns ZERO** (confirmed this session, and confirmed again before filing). **WHY IT IS NOT THE SAME PROBLEM AS R-459, and this is the point of the row: the two engines fail in OPPOSITE directions.** MariaDB starts anyway and skips the conversion quietly, which is why R-459 went unnoticed until a harness looked. **PostgreSQL REFUSES TO START on a datadir from an older major** — the official image performs no `pg_upgrade` and exits with a message naming both versions. So the Postgres case cannot hide; it will present as eight apps down at once, on the sync after the catalog moves. **DELIBERATELY NOT MEASURED HERE, and saying so is the scope discipline:** R-459's task was scoped to MariaDB, and measuring the Postgres analogue is its own piece of work with its own venue. **This row exists so the gap is a record rather than a sentence in an audit nobody greps.** What would settle it: one edge on the existing harness (`postgres:16-alpine` → `17-alpine`) on a scratch host, which would also exercise the `engine_state_after` field's Postgres probe end to end — it is written but has never run against a real Postgres major. Owner: **CC.** `audits/SPIKE-r459-mariadb-upgrade-2026-09-06.md` §7 | **READY — rank P2-MEDIUM; owner: CC** | +| **R-458** | **[P3-LOW] `.felhom.yml` keeps flowing to an app whose compose file is FROZEN, so a frozen app can receive a health check written for a version it is not running.** The v0.235.0 render freezes `docker-compose.yml` for a pinned app once the catalog moves past its version, but copies `.felhom.yml` **verbatim in every case** (`Syncer.copyTemplates`). **The asymmetry is deliberate and both directions were considered:** `.felhom.yml` carries no image, and it carries `catalog_since` — the single input the update badge uses to say *„Frissítés elérhető — N napja"* — so freezing it would silently withhold the one number that tells a customer they are behind, i.e. it would break slice 2 to protect slice 3. **What it costs:** the file also carries the controller-side `healthcheck:` block and resource hints, so a template updated for a newer version can hand a frozen app a probe written for software it is not running. **THE FAILURE DIRECTION IS A FALSE ALARM, NEVER DATA LOSS** — the app keeps running; at worst it renders as degraded and, if it persisted, could reach the dead-app alarm path. That is the same class as R-330's false e-mails, which is why this is a row and not a footnote. **Not fixed now, and the reason is that the cheap fix is wrong:** freezing the whole file breaks the badge, and freezing only the `healthcheck:` key means the syncer would have to parse and re-assemble a customer-facing metadata file — new surface on the one path that touches every app on every box every 15 minutes. **What would settle it:** whether any catalog `healthcheck:` has ever been changed in the same commit as an `image:` line (measurable from the catalog's own history, no box needed). If the answer is "never", the exposure is theoretical and the row can be closed by measurement instead of by code. Owner: **CC.** `architecture/09-update-architecture.md` §5.4, §8.5 **— UPDATE NIGHT 2026-09-21:** **MEASURED 2026-09-21 (update night), leg B9, and the row's risk is NARROWER than it states.** A `.felhom.yml`-only change (a health check for a path only a newer version would serve) was pushed to a FROZEN `bentopdf` — installed `v2.8.6`, catalog ahead. §5.4's asymmetry is confirmed live: the new `.felhom.yml` reached the box while the compose `image:` line stayed `v2.8.6`. **But no false alarm was produced**: ten samples over two minutes all read `state=running` with the front door at `200`. The reason is the probe's own semantics, not luck — `healthprobe.go:258-261` treats **any response** as healthy for `type: http`, and the bogus path answers 404, which is a response. **So this row's false-alarm risk exists only for `type: api` probes carrying an `expect` block**, where the status is compared; for every `type: http` template and every `type: api` without `expect`, a newer version's path is invisible to the probe. The row's actual claim — the failure direction is a false alarm, never data loss — stands and is now measured. Evidence: `audits/update-night-2026-09-21/21-B9-frozen-app-newer-felhomyml.md`. + +| **R-625** | **[P2-MEDIUM] A HELD app keeps inviting the household to update it, and the button then refuses — the exact inconsistency R-524 removed for the other case, still present for this one.** MEASURED 2026-09-21 (update night, leg B6). `glance` was HELD by a genuine unattended failed update. The drill catalog then published a **fixed newer version** — a real forward route, the thing a household would hope for. Afterwards the app page read **„Frissítés elérhető — ma" / „Update available — today"**, in both languages, with the Update button offered; pressing it answered **`409 reason='held'`** and the hold sentence. **The BEHAVIOUR is correct and is DESIGN, not a defect:** `09` §6.1 says the hold is `settings.RestoreHold` and that **a successful unit restore lifts an update hold (only that kind)** — a newer catalog version does not, and should not, because nobody has checked that the new version can start on data the failed one may have touched. **The DEFECT is that the page says otherwise.** R-524 settled precisely this shape for the Ahead case — *"(a) is free and offers a household a downgrade … (c) show „Naprakész" and refuse the button … Why (c): the direction was already settled"* — and chose to make the badge and the button agree. The HELD case still has them disagreeing, in the more painful direction: the badge invites, the button refuses, and the refusal is the same long sentence the household has already read. **What it needs:** the badge for a held app should say what is true — that the app is held and the way back is the restore — and the Update button should not be offered while `RestoreHold` stands. One verdict, read by both surfaces, exactly as R-524 did it. **AND A QUESTION FOR `09` §3b Q4 THAT THIS MEASUREMENT RAISES AND DOES NOT ANSWER:** the household's ONLY route out is a restore, even when the catalog has already shipped a fix. That is defensible, but it is now measured rather than assumed, and Q4's "does the box try again?" should be read next to it. Evidence: `audits/update-night-2026-09-21/bad-days/B6-way-out-forwards/result.json`. | **READY — rank P2-MEDIUM; owner: CC (controller)** | | **READY — rank P3-LOW; owner: CC** | +| **R-460** | **[P3-LOW] BookStack's FILE half cannot be seeded or verified without a browser, so its upgrades can only ever be auto-proven for the DATABASE.** MEASURED 2026-09-06 while building the R-449 harness. BookStack's API needs a token that is only mintable through its web UI, and its HTTP login is unusable headlessly for a second, independent reason: `APP_URL` comes from the template as `https://${SUBDOMAIN}.${DOMAIN}`, so the app marks its session and XSRF cookies **`secure`**; curl over plain http stores neither and **every login POST returns 419 Page Expired**, which looks exactly like a wrong password. The container serves no TLS. **The database half IS provable** — the harness seeds with `php artisan bookstack:create-admin` and reads back with a DIFFERENT artisan command that must find the record, carrying its own negative control on every call. **What is unprovable is an uploaded image or attachment**, i.e. exactly the half a customer would notice. **THIS IS A FACT ABOUT THE APP, NOT A DEFECT IN THE HARNESS**, and it is recorded because Slice 6 needs to know which apps can be auto-verified and which can only be partly verified — nobody had that list before. **Deliberately NOT worked around:** planting a file in the volume would make the test pass while proving nothing, which is R-156's exact failure. **What would remove it:** a headless token route (upstream), or accepting a browser-driven step for this app alone, which DooPlex cannot run. Owner: **CC.** `audits/SPIKE-upgrade-test-2026-09-06.md` §6 **-- UPDATE NIGHT 2026-09-21:** bookstack's edge was walked again on 2026-09-21 and is again **half-proven**: the database half read back through `php artisan` with its own negative control, the file half untouched. The limitation is unchanged and is now measured on the box as well as on the harness. Two more apps joined the same class tonight for a different reason (R-624). | **READY — rank P3-LOW; owner: CC** | +| **R-462** | **[P2-MEDIUM] Widen the upgrade harness beyond three apps — and the cost is dominated by FIXTURES, not by machine time.** The R-449 harness works and is proven by a red negative control (`audits/SPIKE-upgrade-test-2026-09-06.md` §1). **Costed with this run's REAL numbers rather than an estimate:** a successful edge takes **6.4 s – 305.1 s, median 71.8 s**; a FAILING edge takes **556 s**, roughly **8×**, because a negative is only honest if it waits out the full settle window; 3 apps / 11 images cost **5.07 GB**, so 53 apps naively extrapolate to **~90 GB** and, at the median, about an hour of harness time for one edge each. **THAT EXTRAPOLATION UNDERSTATES THE REAL COST BY AN ORDER OF MAGNITUDE, and that is the point of this row.** Two of the three apps needed a bespoke non-browser seed route; one needed two attempts and a discarded approach; one (bookstack) can only ever be half-proven (R-460). **Fixture time scales with apps and does not amortise.** **The decision this row is really asking for is scope, not schedule:** all 53, or only the apps a customer would lose data from, or only apps whose catalog transition is a MAJOR. **Recommended shape, NOT a design — the operator picks:** start with the apps that carry a database, because §3 measured that the abort question only ever bites there. Owner: **VIKTOR rules on scope, CC implements.** `audits/SPIKE-upgrade-test-2026-09-06.md` §5 **ROW CORRECTED 2026-09-21: the scope is NOT open and this row said it was.** It read *"VIKTOR rules on scope"*; the operator ruled on 2026-09-13 (`09` §3 decision 6) that the upgrade test goes to **ALL** apps through the nightly rotation, explicitly *not* "database apps first". What is open is the WORK, not the scope. An ORDER inside that ruling — the 15 database services first, because that is where a wrong answer costs data rather than uptime — is costed as a drill brief in `09` §6.4: legs A–E ≈ **21–34 CC-hours** plus ~25–30 GB on a scratch host, with legs C (a PostgreSQL `pg_upgrade` rehearsal) and E (one automatic night on a throwaway) the two that unblock a decision. **-- UPDATE NIGHT 2026-09-21:** **The count moved from 3 apps to 21 EDGES ACROSS 19 APPS.** The update night walked real within-a-major upstream edges on scratch guest 9202 through the product's own guarded Update, each app seeded and read back through its OWN front door with a negative control on every readback: **14 proven, 3 failed, 4 inconclusive**. Proven: actualbudget, audiobookshelf, bookstack, docmost, grafana, home-assistant, mealie, n8n, navidrome, papra, privatebin, romm, vikunja, and nextcloud's MariaDB engine major. **Ten of the fourteen printed a verbatim migration line**, so the database really was rewritten and the data still read back. Box-side fixtures for 20 apps now exist at `audits/update-night-2026-09-21/fixtures.py`, and four (actualbudget, navidrome, audiobookshelf, vikunja) are ported into `app-catalog-felhom.eu/scripts/upgrade_fixtures.py` with seven new EDGES (U1-U7) so the same edges can be run on the harness venue **with their ABORT step**, which the box deliberately does not offer. **OWED, stated so it is not mistaken for done:** the U1-U7 harness RUNS (the code is in, the runs are not), and fixtures for the four inconclusive apps, of which two (vaultwarden, zipline) cannot be seeded at all while the catalog rightly closes their sign-up (see R-624). | **READY — rank P2-MEDIUM; owner: CC (scope already ruled, `09` §3 decision 6)** | +| **R-463** | **[P2-MEDIUM] The day the catalog moves `postgres:16` to `17`, ELEVEN apps are affected and the container image will NOT perform the conversion — and nothing anywhere records that.** MEASURED 2026-09-06: 11 of the 53 templates carry PostgreSQL — **8 on `postgres:16-alpine`**, 1 on `postgres:15-alpine`, plus `postgis/postgis:16-3.5-alpine` and Immich's own `postgres:16-vectorchord…` build. **A grep of the whole register for `pg_upgrade`, "postgres major" or "postgresql major" returns ZERO** (confirmed this session, and confirmed again before filing). **WHY IT IS NOT THE SAME PROBLEM AS R-459, and this is the point of the row: the two engines fail in OPPOSITE directions.** MariaDB starts anyway and skips the conversion quietly, which is why R-459 went unnoticed until a harness looked. **PostgreSQL REFUSES TO START on a datadir from an older major** — the official image performs no `pg_upgrade` and exits with a message naming both versions. So the Postgres case cannot hide; it will present as eight apps down at once, on the sync after the catalog moves. **DELIBERATELY NOT MEASURED HERE, and saying so is the scope discipline:** R-459's task was scoped to MariaDB, and measuring the Postgres analogue is its own piece of work with its own venue. **This row exists so the gap is a record rather than a sentence in an audit nobody greps.** What would settle it: one edge on the existing harness (`postgres:16-alpine` → `17-alpine`) on a scratch host, which would also exercise the `engine_state_after` field's Postgres probe end to end — it is written but has never run against a real Postgres major. Owner: **CC.** `audits/SPIKE-r459-mariadb-upgrade-2026-09-06.md` §7 **-- UPDATE NIGHT 2026-09-21:** **Measured 2026-09-21, both halves.** (a) What a household sees TODAY: the guarded Update of `postgres:16-alpine` to `17-alpine` on docmost ended **`failed` in 5.1 s**, the app stopped and held, **the pin naming 17 while `installed_images` still said 16 and nothing ran**, the data intact, and the restore the hold sentence names back in **29.1 s**. The engine's refusal line had to be REPRODUCED independently because `failAndHold` destroyed it (R-621): *FATAL: database files are incompatible with server / DETAIL: The data directory was initialized by PostgreSQL version 16, which is not compatible with this version 17.11.* The datadir was still `16` afterwards, and the same copy started under 16 holding 48 tables as the positive control. (b) The conversion **COSTED** on a real seeded 49 MB / 48-table datadir: `pg_dumpall` **2.6 s / 132 201 B**, fresh 17 plus replay **6.5 s / 48 tables restored**, the app up on 17 saying *Database connection successful*, **the seeded account read back**, total **155.9 s of which ~9 s is engine work**. `pg_upgrade` was NOT run: it needs both majors' binaries in one image and no such image exists here. Full paragraph: `audits/update-night-2026-09-21/24-Q5-postgres-conversion-costed.md`. | **READY — rank P2-MEDIUM; owner: CC** | | **R-464** | **[P3-LOW] MariaDB's entrypoint prints `MariaDB upgrade not required` on an UNSUPPORTED DOWNGRADE, so that line cannot be used as a soundness signal.** MEASURED 2026-09-06. After converting a datadir to `12.3.3-MariaDB` and then starting **11.6** on it, the entrypoint logs, on every start: **`[Note] [Entrypoint]: MariaDB upgrade not required`**. Asked properly, the same engine answers **`FATAL ERROR: Version mismatch (12.3.3-MariaDB -> 11.6.2-MariaDB): Trying to downgrade from a higher to lower version is not supported!`** **The entrypoint compares the datadir's recorded version against its own and concludes there is nothing to DO. That is true, and it is not a statement that the state is sound.** **THIS IS THIS PROJECT'S MOST-REPEATED CLASS, in a new costume** — the same shape as `CLAUDE.md`'s "presence is not success" and as R-443's HTTP 200 over a crash-looping app: a reassuring sentence that answers a narrower question than the one a reader will take it for. **Why it is worth a row rather than a footnote: the obvious cheap instrument for R-459 is to grep container logs for that exact line**, and such an instrument would report "fine" for an unsupported downgrade. **The correct probe is `mariadb-upgrade --check-if-upgrade-is-needed`**, which is what `upgrade-test.py`'s `engine_state_after` now uses. **Also recorded, because it nearly produced a wrong answer here: run without credentials that command returns `ERROR 1045 … FATAL ERROR: Upgrade failed` with exit 1** — an authentication failure wearing the shape of a verdict. Owner: **CC.** `audits/SPIKE-r459-mariadb-upgrade-2026-09-06.md` §5.4 | **READY — rank P3-LOW; owner: CC** | | **R-468** | **[P3-LOW] THE GOLDEN WAIVER — goldens on a cadence, not per release (operator ruling 2026-09-13).** 25 goldens in 26 days in August, almost one per release, because `golden_currency_gate.py` trips on every release by design and the only honest ways past it were a bake or a declared `--no-verify` (thirteen by 2026-09-01, R-404/R-417). **The ruling: bake WEEKLY, and always before any drill or fresh install.** Every release still raises the FLOOR, so both demo boxes keep getting each release in ~20 s; only the golden — which protects a fresh install and nothing else — moves to a cadence. **The mechanism (built 2026-09-13):** `documentation/tests/golden-waiver. **⚠ CORRECTED THE SAME DAY (R-472): between bakes the floor does NOT carry a release — the hub holds any floor above the vouched golden (publish-train rule 1), so releases between bakes reach the demo boxes only by hand-deploy.**yml`, four lines (`issued`, `expires`, `reason`, `register_row: R-468`), read by the gate. While valid, a golden BEHIND the record makes the gate print a loud ADVISORY and exit 0; when it expires the gate is red again until someone bakes or renews. **The 14-day cap is enforced by the gate, not the runbook** — a longer, undated, unparseable, reason-less or row-less waiver is INCONCLUSIVE (exit 2), never 0 and never silently ignored. **It never covers a golden that is UNRECORDED (R-385)** — that is not a cadence choice. **A dated waiver cannot be forgotten; it just expires** — the difference from R-242's original rule, which recurred the day after it was written. Tests: `scripts/test_golden_currency_gate.py` cases 5–15 (E/F/G/H, a 15-day, absent, unparseable, bad-row and empty-reason waiver each 2; the R-421 decoy — a file saying only `expires` — 2). **This is a PRE-CUSTOMER arrangement: the first external install retires it** (delete the file in that commit). Cadence written into `RUNBOOK-manual-build.md` §4.2 and the `felhom.eu` end-of-session checklist. **Does NOT touch R-242's open half (nothing gates the VOUCH).** | **WATCHING — rank P3-LOW; owner: CC (renew ≤ 14 days or bake); retire at the first external install** | @@ -779,14 +783,14 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server` | **R-601** | **[P2-MEDIUM] ~~demo-hp is unreachable~~ — WRONG, WITHDRAWN THE SAME DAY. The box was never down; MY ROUTES WERE.** Filed 2026-09-21 morning after `ssh demo-hp`, the hub-vaulted break-glass over the tailnet, `demo-hp-lan`, a ping and `ip neigh` on felhom-pve all failed, and `tailscale status` said *`demo-hp … offline, last seen 30d ago`*. **The operator looked at the hub and said it was ONLINE. It was**: it had reported 13 minutes earlier, and it has been up **4 weeks 2 days**. **Two stale facts, each enough on its own:** (1) `~/.ssh/config` sends `demo-hp` to the tailnet address `100.76.96.79`, and **tailscale is not installed on that box at all** (checked on it: no `tailscaled`, no `tailscale` binary) — so that entry is a dead peer from an earlier build and can never answer; (2) `demo-hp-lan` and `nodes.md` both say `192.168.0.87`, and the box is **statically** on **`192.168.0.104/24`**, bridge-port `nic0` (nodes.md says `enp2s0f0`). **The hub knew the right address the whole time** — every host report carries `addresses: [{iface: vmbr0, cidr: 192.168.0.104/24}, …]`. **What I actually did wrong, and it is the part worth keeping:** I ran `ip neigh` on felhom-pve, and `192.168.0.104 … STALE` was *in that output*, four lines above the `192.168.0.87 … FAILED` I quoted. I searched the output for the address I expected instead of reading it for the address that was there. **The standing rule says a "no access" claim must list what was tried; it does not say the list makes the claim true.** Six failed routes to a stale address are six failures of one assumption, not six pieces of evidence. **FIXED:** both `~/.ssh/config` entries repointed to `.104` (each carrying a comment saying why, including that there is no tailscale on this box), both verified live; `nodes.md` corrected. | **CLOSED 2026-09-21 — withdrawn, the claim was false; the routes are fixed** | | **R-604** | **[P2-MEDIUM] A per-customer controller floor silently excludes that box from every global floor raise, and NOTHING says so — demo-hp missed four of them.** FOUND 2026-09-21 while raising the global floor to 0.259.0 at the operator's request. The raise logged `Global controller-version floor set to "0.259.0"` and then `managed floor SERVED for demo-felhom` — **and nothing at all for demo-hp**, which went on reporting every few minutes and stayed on 0.258.0. Cause: `customer_configs.min_controller_version` for demo-hp held **`0.243.0`**, a per-customer override that wins over the global. It is a **leftover from the 2026-09-16 drill**, whose golden was 0.243.0; **R-343 measured on 2026-08-18 that all five rows were EMPTY and recorded that as a safety property** — it stopped being true and nothing surfaced the change. demo-hp had therefore silently missed the raises to 0.253.0, 0.254.0, 0.257.0 and 0.259.0. **Why it is invisible rather than merely quiet:** `managed floor SERVED` fires **once per CHANGE** (`h.floorNotes`, `api/handler.go:600`), deliberately, because a box reports every few minutes — so a box whose override never changes is silent for ever, and its silence is indistinguishable from the silence of a box that already had the line. A session that raises the floor reads one SERVED line and reasonably concludes the fleet took it. **CLEARED** for demo-hp the same session (rollback line: POST `/customers/demo-hp/floor` with `min_controller_version=0.243.0`, `min_agent=0.131.0`); it then self-updated 0.258.0 → 0.259.0 in **under four minutes**, healthy, `settle-gate: GO — at/above floor 0.259.0`, and its claim page answers **"Wrong or expired code"** in English — the floor delivered the FIX, not a version string, to a box nobody hand-deployed. All five overrides are now empty. **Fix shape (smallest first):** the floor-raise page shows which customers carry an override and would NOT be moved, before the save; or the raise logs one line per customer naming the ones it skipped and why. A raise that quietly reaches half the fleet is worse than one that refuses. | **READY — rank P2-MEDIUM; owner: CC (hub)** | | **R-605** | **[P3-LOW] A catalog gate that REFUSED TO RUN and a gate that ran and could not decide print the same word, so a reader cannot tell which happened.** FOUND 2026-09-21 while answering why the chaos night's update round could not run. On 2026-09-17 `check-image-resolvable` and `check-volume-persistence` both returned INCONCLUSIVE and the drawn `update` action was replaced with `use` (`audits/DRILL-chaos-night-2026-09-17.md:692-695`). **Neither script is defective — they behaved exactly as designed**, and both headers say why: a detector that cannot prove itself must refuse to report rather than guess (`check-image-resolvable.py` cites the 2026-07-21 incident where a Docker Hub throttle read as 24 of 65 pins falsely dead). **What is missing is the DISTINCTION.** `check-volume-persistence.py`'s `self_test` refuses to evaluate ANY app when it cannot build its canary image — a harness-level refusal — while `classify()` returns a per-app UNDETERMINED for an app that wrote nothing; `check-image-resolvable.py` likewise separates a harness-level canary failure (exit 2 at `check()` L180-183) from a per-pin throttle (L121-128). **`catalog_gates.py`'s VERDICT map collapses all of them into one `INCONCLUSIVE` label**, so the operator-facing summary cannot say whether the gate ran at all. **The cost is real and already paid:** no raw stdout of the 2026-09-17 run survives in either evidence directory, so the exact triggering path is INFERRED from the code plus the documented throttle precedent, not observed — a distinct summary line would have recorded it for free. **Fix shape:** have each gate's exit distinguish "the harness refused" from "the result is undetermined" (a third exit code, or a marker line the runner matches), and have `catalog_gates.py` print the two differently. **Ships with a decoy each way (R-421): a run whose canary fails must NOT read as a per-app undetermined, and vice versa.** Small. | **READY — rank P3-LOW; owner: CC (catalog)** | -| **R-606** | **[P2-MEDIUM] Every sentence the UPDATE path shows a household is Hungarian-only, and it lands on TWO pages — FOUND LIVE 2026-09-21, not by reading.** While measuring R-520 on guest 9202 at controller v0.260.0, the English app page rendered the badge in English — *"Update available — today"* — directly above *„A frissítés megszakadt, mert a vezérlő újraindult, mielőtt az új verzió elindult volna. Az alkalmazás a korábbi verzióval fut tovább."* **The mixed line is worse than either language whole**, and this one is the household's only explanation for why their app did not move. **MECHANISM:** `Stack.UpdateError` is a finished Hungarian STRING, not a key. `Manager.finishUpdate` stores it (`internal/stacks/update.go` L517, 526, 530, 679, 902, 907, 936, 943) from the raw literals `MsgUpdateInterrupted`, `MsgUpdatePullFailed`, `MsgUpdateBackupFailFmt`, `MsgUpdateDumpFailFmt`, `MsgUpdatePinFailed`, `MsgUpdateJournalFailed`, `MsgUpdateBackupNoUnit`, `MsgUpdateHoldUnsaved` (`update.go` L79-96), and BOTH `app_info.html:36` and `stacks.html:99,104` render it verbatim. **SCOPE IS WIDER THAN THE EIGHT:** the same path carries `UpdatePhaseLabel` (`updatePhaseLabels`, L63) and the HOLD sentence returned by `UpdateGuards.HoldFor`, including `backup.Manager.UpdateCopyHolds`'s *„csak a beállításokat és az adatbázist tartalmazza, a fájlokat nem"* — **which is a PROMISE ABOUT WHETHER THE CUSTOMER'S FILES COME BACK**, and ranks this with R-590 rather than below it. **NOT closed by v0.260.0:** that release routed the 409 REFUSAL through `errText` (so the pre-flight refusals reach an English household in English), but a refusal is the path where nothing happened — these are the sentences for when something DID. **Fix shape, and the pattern already exists in this repo:** `UpdateError` stores a KEY plus args, exactly as v0.259.0's `degradedMessageFor` was changed to return a key, and the page resolves it with `errText`/`msg` at render — the decision stays language-free in one place while the words are chosen by whoever knows the reader. `util.MsgErrorf` already carries key+args across that gap. Render test per sentence in both languages. **Every one of these is BORN AS A KEY territory, so `i18n_go_keys.json` accounting applies.** | **READY — rank P2-MEDIUM; owner: CC (controller)** | -| **R-607** | **[P3-LOW] A forced catalog sync answers „nincs változás" while the cache DOES change, and `catalog_images` stays stale until a separate rescan — so the update badge can be wrong for a window nobody bounds.** MEASURED 2026-09-21 on scratch guest 9202 (controller v0.260.0) while proving R-524. A real catalog move was pushed, `POST /api/sync` was invoked, and it answered **„nincs változás"** — yet the box's own cache file `/catalog-cache/templates/uptime-kuma/docker-compose.yml` **had moved to the new tag**. `Stack.CatalogImages` as served by the API stayed at the OLD value until a separate `POST /api/stacks/rescan`. **WHY IT MATTERS AND WHY IT IS NOT COSMETIC:** `CatalogImages` is the single input `stacks.CatalogOrder` compares against (v0.260.0), so for that window the badge answers from a stale catalog — it can read „Naprakész" on an app that IS behind, which is the exact failure §5.6 of `09` was written to prevent, arriving by a different door. **It also cost a measurement:** the session that found it nearly recorded a `tag-ok` badge as proof of the R-524 ahead arm when the badge was in fact stale; the honest reading came only after the rescan. **An instrument that can report an old value as a current one is not a measurement.** **TWO SEPARATE QUESTIONS, and the row does not conflate them:** (a) why the sync REPORTS no change when the working tree moved — a wrong sentence, possibly a comparison against the wrong ref; (b) whether `CatalogImages` is refreshed by the sync at all or only by `ScanStacks` on its own timer — if the latter, the staleness window is the scan interval and is bounded but unstated. **Neither was isolated** — this row records the observation, not a diagnosis. **Also observed in the same run, NOT diagnosed and folded in here rather than filed twice:** a removed app leaves `applied-compose.yml` behind in its stack directory. Stated as observed; it was not established whether that is intended. **Fix shape:** first reproduce with a loop that pushes a tag, syncs, and reads `catalog_images` on a timer, so the window is a NUMBER before anything is changed. Evidence: `audits/update-arc-2026-09-21/06-sync-after-bump.txt` and `16-r524-sync-box-ahead.txt`. | **READY — rank P3-LOW; owner: CC (controller)** | +| **R-606** | **[P2-MEDIUM] Every sentence the UPDATE path shows a household is Hungarian-only, and it lands on TWO pages — FOUND LIVE 2026-09-21, not by reading.** While measuring R-520 on guest 9202 at controller v0.260.0, the English app page rendered the badge in English — *"Update available — today"* — directly above *„A frissítés megszakadt, mert a vezérlő újraindult, mielőtt az új verzió elindult volna. Az alkalmazás a korábbi verzióval fut tovább."* **The mixed line is worse than either language whole**, and this one is the household's only explanation for why their app did not move. **MECHANISM:** `Stack.UpdateError` is a finished Hungarian STRING, not a key. `Manager.finishUpdate` stores it (`internal/stacks/update.go` L517, 526, 530, 679, 902, 907, 936, 943) from the raw literals `MsgUpdateInterrupted`, `MsgUpdatePullFailed`, `MsgUpdateBackupFailFmt`, `MsgUpdateDumpFailFmt`, `MsgUpdatePinFailed`, `MsgUpdateJournalFailed`, `MsgUpdateBackupNoUnit`, `MsgUpdateHoldUnsaved` (`update.go` L79-96), and BOTH `app_info.html:36` and `stacks.html:99,104` render it verbatim. **SCOPE IS WIDER THAN THE EIGHT:** the same path carries `UpdatePhaseLabel` (`updatePhaseLabels`, L63) and the HOLD sentence returned by `UpdateGuards.HoldFor`, including `backup.Manager.UpdateCopyHolds`'s *„csak a beállításokat és az adatbázist tartalmazza, a fájlokat nem"* — **which is a PROMISE ABOUT WHETHER THE CUSTOMER'S FILES COME BACK**, and ranks this with R-590 rather than below it. **NOT closed by v0.260.0:** that release routed the 409 REFUSAL through `errText` (so the pre-flight refusals reach an English household in English), but a refusal is the path where nothing happened — these are the sentences for when something DID. **Fix shape, and the pattern already exists in this repo:** `UpdateError` stores a KEY plus args, exactly as v0.259.0's `degradedMessageFor` was changed to return a key, and the page resolves it with `errText`/`msg` at render — the decision stays language-free in one place while the words are chosen by whoever knows the reader. `util.MsgErrorf` already carries key+args across that gap. Render test per sentence in both languages. **Every one of these is BORN AS A KEY territory, so `i18n_go_keys.json` accounting applies.** **— UPDATE NIGHT 2026-09-21:** **CONFIRMED 2026-09-21 (update night) on the HOLD sentence, which this row's own text ranks highest** — *„a PROMISE ABOUT WHETHER THE CUSTOMER'S FILES COME BACK"*. Read off `/apps/adventurelog?lang=en` while the app was genuinely held after a real failed upstream edge. **Everything around it is correctly English** — the nav, „An installed app is not running: AdventureLog (stopped)", „Update available — today", „Move to another storage" — **and the two sentences that matter are Hungarian**: „A(z) adventurelog frissítése … nem sikerült, és az alkalmazás nem indult el az új verzióval." and „Visszaállítható a Mentések oldalon ebből a biztonsági mentésből: saját meghajtó, 2026-09-21 20:47 — ez a másolat a beállításokat, az adatbázist és az adatköteteket tartalmazza." So an English household is told in English that their app is stopped, and in Hungarian **which copy brings their data back, when it was taken and what is inside it.** Positive and negative controls both quoted. **Also settled by the same reading, and it belongs to `09` Q4:** the held app is surfaced on EVERY authenticated page, not only its own — the banner carried BOTH held apps at once. Evidence: `audits/update-night-2026-09-21/15-r606-hold-sentence-on-the-english-page.txt` and `apps/adventurelog/held-page-en.html`. **— UPDATE NIGHT 2026-09-21:** **the row records v0.260.0 as having made the pre-flight REFUSALS reach an English household in English. Measured 2026-09-21: it did not.** Three refusals were requested with `?lang=en`, with the Hungarian request as the control, and all three came back **identical Hungarian**: `held` (the one that names which copy holds what), `not_deployed`, and `disk` („Nincs elég szabad hely a frissítéshez: 1.4 GB szabad…"). **The mechanism is not a regression — it is that the pipe was built and the sentences never entered it.** `Router.langFor` DOES honour `?lang=` (`api/i18n_api.go:30-40`) and `errText` calls it; but `MsgUpdateDiskFmt`, `MsgUpdateNotDeployed`, `MsgUpdateBusy` and their siblings (`stacks/update.go` L81-96) are **finished Hungarian string constants** raised with `fmt.Sprintf`, and `errText` correctly renders "its own text" for an error carrying no bundle message. Only the sentences BORN as keys — v0.260.0's `downgrade`, v0.261.0's `self_updating` — actually translate. **A row that records something as fixed when it is not is worse than an open row**, which is why this correction is here rather than in prose. Evidence: `audits/update-night-2026-09-21/20-refusals-in-english.txt`. | **READY — rank P2-MEDIUM; owner: CC (controller)** | +| **R-607** | **[P3-LOW] A forced catalog sync answers „nincs változás" while the cache DOES change, and `catalog_images` stays stale until a separate rescan — so the update badge can be wrong for a window nobody bounds.** MEASURED 2026-09-21 on scratch guest 9202 (controller v0.260.0) while proving R-524. A real catalog move was pushed, `POST /api/sync` was invoked, and it answered **„nincs változás"** — yet the box's own cache file `/catalog-cache/templates/uptime-kuma/docker-compose.yml` **had moved to the new tag**. `Stack.CatalogImages` as served by the API stayed at the OLD value until a separate `POST /api/stacks/rescan`. **WHY IT MATTERS AND WHY IT IS NOT COSMETIC:** `CatalogImages` is the single input `stacks.CatalogOrder` compares against (v0.260.0), so for that window the badge answers from a stale catalog — it can read „Naprakész" on an app that IS behind, which is the exact failure §5.6 of `09` was written to prevent, arriving by a different door. **It also cost a measurement:** the session that found it nearly recorded a `tag-ok` badge as proof of the R-524 ahead arm when the badge was in fact stale; the honest reading came only after the rescan. **An instrument that can report an old value as a current one is not a measurement.** **TWO SEPARATE QUESTIONS, and the row does not conflate them:** (a) why the sync REPORTS no change when the working tree moved — a wrong sentence, possibly a comparison against the wrong ref; (b) whether `CatalogImages` is refreshed by the sync at all or only by `ScanStacks` on its own timer — if the latter, the staleness window is the scan interval and is bounded but unstated. **Neither was isolated** — this row records the observation, not a diagnosis. **Also observed in the same run, NOT diagnosed and folded in here rather than filed twice:** a removed app leaves `applied-compose.yml` behind in its stack directory. Stated as observed; it was not established whether that is intended. **Fix shape:** first reproduce with a loop that pushes a tag, syncs, and reads `catalog_images` on a timer, so the window is a NUMBER before anything is changed. Evidence: `audits/update-arc-2026-09-21/06-sync-after-bump.txt` and `16-r524-sync-box-ahead.txt`. **— UPDATE NIGHT 2026-09-21:** **Seen again 2026-09-21 (update night), a dozen times in one session, and for the first time with a USER-VISIBLE consequence rather than a measurement one.** On `mealie` the bump was pushed, `POST /api/sync` AND `POST /api/stacks/rescan` were both run, and the badge still read the up-to-date one (HU „Naprakesz", EN "Up to date") — the catalog had not reached `catalog_images` yet. The guarded Update was then pressed and **reported „Frissitve" after 2.1 seconds having moved nothing at all**: pinned, installed, the live compose line and `docker inspect` all still read `v3.20.1`. That is honest given a stale cache — the pin is written from *the catalog's current definition*, which was still the old one — but what the household sees is a button that says it updated them and did not. **A NUMBER, at last, which is what this row asks for:** the night's harness was changed to poll `catalog_images` until the pushed reference appears and to report how long that took; those figures are each edge's `badge_catchup_seconds`, and here they are: **4.4 s, 4.4 s, 4.5 s, 4.5 s — and 29.0 s.** The four fast ones are one sync+rescan round; the 29-second one (`nextcloud`, an engine-sidecar bump) needed **additional** sync+rescan rounds before `catalog_images` carried the pushed reference. **So the window is not a fixed scan interval — it varies by roughly 7x between edges on the same box in the same hour**, which is why a caller (or a household) cannot know when the badge is safe to read. Before tonight this row had no number at all; it now has five, and they disagree with each other, which is itself the most useful thing about them. Every drill-catalog bump of the night was followed by `POST /api/sync` answering „Sablonok naprakészek — nincs változás" while the box's cache HAD moved, with `catalog_images` staying stale until a separate `POST /api/stacks/rescan`. The night's harness therefore rescans unconditionally after every sync, which is a workaround and not a fix. **The window was still never measured as a NUMBER** — that is what the row asks for and what remains owed. | **READY — rank P3-LOW; owner: CC (controller)** | | **R-608** | **[P2-MEDIUM] The controller swaps ITSELF in the middle of a guarded app update, and 04:30 sits inside the window proposed for automatic app updates.** FOUND 2026-09-21 by reading the clock, not by a failure. The controller self-updates daily at `self_update.auto_update_time` — **default 04:30** (`config/config.go` L422, scheduled `cmd/controller/main.go` ~L1365) — and again from `MaybeAutoUpdate` after ANY hub report once a floor sits above the box, so at any hour. The swap restarts the controller container. `09` §3b **Q1** proposes **02:30–05:00** for automatic app updates. **It contains 04:30.** **MEASURED, and the gap was NARROWER than it first looked — which is why the fix is where it is:** the updater's only busy gate was `backupRunning` (`updater.go` L61, read at L487 dry-run, L512 `TriggerUpdate`, L660 `maybeAutoUpdate`), wired in `main.go` L659 to `backupMgr.IsRunning()`. The guarded update's **`backing-up` phase DOES take the backup single-flight** (`RunAppBackupNow` → `acquireRunning`, `backup/update_guard.go:333`), so that ONE phase was already covered. `checking`, `safety-dump`, `pinning`, `pulling`, `starting` and `verifying` were not — and the last two are exactly where the new version may already have touched the customer's data. The reverse was absent too: `UpdatePreflight` never asked whether a swap was running. **CLOSED 2026-09-21 — controller v0.261.0.** `stacks.Manager.AnyUpdating()` → `Updater.SetAppUpdatingCheck`, a deliberate sibling of `SetBackupRunningCheck` consulted in the SAME three places; `Updater.IsUpdateRunning` → `Manager.SetSelfUpdatingCheck`, and `UpdatePreflight` refuses `self_updating`. Both wired in `main.go`, the only place holding both objects — **`stacks` never imports `selfupdate`.** Two sentences born as bundle keys. **THE PROPERTY THAT MATTERS MOST IS THAT THE LOCK DOES NOT LATCH:** `Stack.Updating` is cleared on done, failed AND held, so a HELD app does not block the controller's own updates — including the release that might fix whatever held it. A latching gate would be a worse failure than the one prevented, and a silent one. Pinned by `TestR608_LockReleasesAfterHold`. Four red-proofs, each seen to fail. | **CLOSED 2026-09-21 — controller v0.261.0** | | **R-609** | **[P3-LOW] An update refusal has a machine-readable reason inside the process and none on the wire, so an unattended caller cannot tell "wait" from "never".** `UpdateRefusal.Reason` has existed since v0.237.0 (`busy`, `deploying`, `updating`, `held`, `migrating`, `memory`, `disk`, `no_backup`, `downgrade`, and `self_updating` since v0.261.0) and never left the process: the 409 body carried only the translated sentence. **The distinction is not decorative** — `busy`/`updating`/`deploying`/`migrating`/`self_updating` are TRANSIENT and `held`/`downgrade` are TERMINAL until a person acts. A caller that cannot tell them apart either gives up on a passing backup window or presses a terminally-refused button on every pass for ever. `09` §6.2's unattended caller reads exactly this. **CLOSED 2026-09-21 — controller v0.261.0.** The body gains `data: {"reason": ""}`, ADDITIVELY; the sentence is unchanged and no page moves. Table-driven test over five reachable paths plus a control that a non-refusal carries none. **FOUND WHILE WRITING THE TEST, NOT BY READING — and it was the reason that matters most:** `actionStack` refuses a HELD app on **its own line, BEFORE `UpdatePreflight`** (`api/router.go` ~L601), so `held` would have been the one reason missing from the wire. That line now carries it too. Red-proof seen to fail. | **CLOSED 2026-09-21 — controller v0.261.0** | | **R-610** | **[P2-MEDIUM] The power cut AFTER the new version has started was never measured — R-520 closed on the safe half only.** R-520 (CLOSED 2026-09-21) cut in `pulling`, where **nothing had run**: the pin goes back and that is the easy case. Its own last lines said the `starting` cut was "NOT measured and does not re-open this row", and no row carried it. The dangerous case is the cut AFTER the new version has started and may already have migrated the customer's data — where `RecoverUpdates` (`stacks/update.go:909`) marks the app Updating and RESUMES rather than rolling back. **That behaviour was READ from the source and never observed.** **CLOSED 2026-09-21 — measured THREE times on guest 9202, controller v0.260.0**, with three different apps and two different cut mechanisms: vikunja 2.3.0→2.6.0 and uptime-kuma 2.4.0→2.5.0 by `pct stop` (a real power cut), and wishlist v0.66.0→v0.67.0 by restarting ONLY the controller container (exactly what a self-update does). **All three ended HONEST:** the recovery line appeared, the update resumed, each app came up on the NEW version, and in every case all FOUR version observables agreed — `pinned_images`, `installed_images`, the live compose `image:` line, and `docker inspect` of the running container (digests matched too). No hold, no stuck `Updating`, no surviving journal, no retry loop. The seeded data read back through each app's own front door for the two where a post-cut read-back was taken. **THE DANGEROUS CASE WAS GENUINELY EXERCISED, and the proof is a log line, not an assumption:** vikunja's own log shows `Ran all migrations successfully` and `Vikunja version v2.6.0` at **12:28:26.881 UTC — 0.64 s after the cut decision and ~0.4 s before the guest stopped answering.** The 2.6.0 schema migration had ALREADY been applied to the customer's SQLite database when the power went. Recovery resumed FORWARD, so old-binary-on-migrated-database never happened — **but this branch is one step from it: had the cut landed a second earlier, in `pinning` or `pulling`, the 2.3.0 pin would have been put back onto a 2.6.0 database.** That is not a defect today; it is the reason §4's "no automatic rollback" ruling is right, and it is now evidence rather than argument. **INSTRUMENT LIMIT, stated because it bounds the claim:** `starting` lasts well under a second on this box. Three attempts across two cut mechanisms (`pct stop` returning in 3.0–3.8 s; a controller restart in 1.7 s) ALL landed in `verifying`. No phase was faked. **`RecoverUpdates` handles `starting` and `verifying` in ONE branch, so all three runs exercise the same recovery arm** — the arm under test. A cut that lands inside `starting` itself remains unmeasured and would need an in-process fault injector. Evidence: `audits/update-arc-gaps-2026-09-21/` 04, 05, 07. | **CLOSED 2026-09-21 — measured three times** | | **R-611** | **[P3-LOW] A session reported "everything is done" over a phase it had silently skipped — the process failure, not the missing measurement.** The 2026-09-21 update-arc session's brief contained a Phase 5 spike: one app updated by the box with **nobody pressing anything**, once succeeding and once forced to fail. **It did not run, and nothing said so** — no evidence file, no code, and no sentence in `STATUS.md`, `UPDATE-ARC-STATE-2026-09-21.md`, either `REPORT.md` or `09`. `09` §6.2 was left describing Slice 6 "as it would be built" with no measurement under it, which reads like a considered design rather than an untested one. **Why this is a row and not a grumble:** the missing measurement was recoverable in an afternoon; the missing SENTENCE was not, because the next reader had no way to know it was missing. A skipped phase that is declared costs one line; a skipped phase that is not costs the next session its baseline. **CLOSED 2026-09-21 by the successor session**, which ran it (`audits/update-arc-gaps-2026-09-21/`, scenarios F and G) **and** adopted the standing habit that closes it generally: **the report's FIRST section is "not done", even when empty.** Every part and scenario of a brief is listed there if it was skipped, shortened or changed, with the reason. | **CLOSED 2026-09-21 — run by the successor session; "not done" is now the report's first section** | | **R-612** | **[P1-HIGH] `wishlist` cannot be signed up to on a fresh Felhom install, the deploy reports SUCCESS, and the error the customer sees is a LIE.** MEASURED 2026-09-21 on guest 9202 while seeding for the power-cut drill. The image's first-boot `pnpm prisma db seed` is **`Killed` — OOM at the catalog's `mem_limit: 128M`**. Without it the `Role` and `Group` rows are absent, so **every** signup fails. **The message the user is shown is `User with username or email already exists`** while the container log says the real cause: `FOREIGN KEY constraint violated`. A household would conclude the account already exists and try to recover a password that was never created. **The controller reports the app running and HEALTHY throughout, and the deploy reported successful** — so nothing on the box says anything is wrong. Repaired on the scratch guest only, to unblock seeding: memory raised to 512 M, the image's own seed re-run, memory put back to 128 M. **The catalog was NOT changed** — the fix is a memory-limit question for the catalog and is deliberately left to a session that can measure the real ceiling rather than guess it. **Needs: the actual peak RSS of that seed, then a `mem_limit` that clears it, plus a check that the seed's failure is not silent.** | **READY — rank P1-HIGH; owner: CC (catalog + a look at whether a failed first-boot seed can ever be visible)** | -| **R-613** | **[P2-MEDIUM] `uptime-kuma` parks on its setup wizard with no login and no monitors, and the box tells the household it is HEALTHY.** MEASURED 2026-09-21 on guest 9202. On first boot uptime-kuma 2.4.0 sits at `SETUP-DATABASE` (`Waiting for user action...`) and its main socket.io server never starts. **The controller's `http :3001` probe sees the wizard's 302 and records the app as running and healthy.** So a monitoring app that cannot be logged into, and is monitoring nothing, is presented to the customer as fine. Passed through its own front door for the drill with `POST /setup-database {"dbConfig":{"type":"sqlite"}}`. **This is the health-check class the catalog skill already warns about — a probe that proves the PORT answers, not that the APP works** — and it is worth a row because the failure direction is a false GREEN, which no alarm will ever catch. **Needs: a healthcheck for this template that fails while the wizard is up** (the catalog `REUSE.md` maps the families), and a sweep for other templates whose probe would pass on a setup wizard. | **READY — rank P2-MEDIUM; owner: CC (catalog)** | +| **R-613** | **[P2-MEDIUM] `uptime-kuma` parks on its setup wizard with no login and no monitors, and the box tells the household it is HEALTHY.** MEASURED 2026-09-21 on guest 9202. On first boot uptime-kuma 2.4.0 sits at `SETUP-DATABASE` (`Waiting for user action...`) and its main socket.io server never starts. **The controller's `http :3001` probe sees the wizard's 302 and records the app as running and healthy.** So a monitoring app that cannot be logged into, and is monitoring nothing, is presented to the customer as fine. Passed through its own front door for the drill with `POST /setup-database {"dbConfig":{"type":"sqlite"}}`. **This is the health-check class the catalog skill already warns about — a probe that proves the PORT answers, not that the APP works** — and it is worth a row because the failure direction is a false GREEN, which no alarm will ever catch. **Needs: a healthcheck for this template that fails while the wizard is up** (the catalog `REUSE.md` maps the families), and a sweep for other templates whose probe would pass on a setup wizard. **— UPDATE NIGHT 2026-09-21:** the update night could not seed `uptime-kuma` for the same reason and left it out rather than faking it. | **READY — rank P2-MEDIUM; owner: CC (catalog)** | | **R-614** | **[P3-LOW] A stale `update_phase` survives a remove and redeploy, so a freshly installed app can read "Frissitve" before it has ever been updated.** OBSERVED 2026-09-21 on guest 9202: a newly deployed `uptime-kuma` read `update_phase=done` / „Frissitve" before any update had been run against it — left in the manager's IN-MEMORY stack state by the previous session's update of a since-removed instance of the same name. It cleared on the next controller restart. **Why it matters beyond the cosmetic:** `update_phase` is one of the fields a person (and, after R-609, an unattended caller) reads to decide whether an update happened. A value that outlives the app it described is the same class as the R-166 "absent means unknown" family — a confident answer about something that no longer exists. **Fix shape:** clear the update fields when a stack is removed, beside wherever `Updating`/`updateHeld` are reset; a test that removes and redeploys and asserts the phase is empty. Small. | **READY — rank P3-LOW; owner: CC (controller)** | | **R-602** | **[P3-LOW] The language a signed-in page uses is NOT the language a cookie asks for, and a live probe that forgets this reports a fixed defect as unfixed.** FOUND 2026-09-21 verifying R-598 on guest 9201. `GET /backups` with `felhom_lang=en` returned the **Hungarian** page. That is correct — `langFor` step 2 says a request carrying a session reads the household's saved setting and deliberately ignores the visitor cookie, so a signed-in family never sees a language a previous visitor picked on the sign-in page of the same browser — but it means **the cookie is the right instrument for the anonymous claim/login/bind pages and the wrong one for every page behind auth**, where `?lang=` is. A session that had run only the cookie probe would have concluded R-598 was still open and fixed it a second time. **This is a documentation gap, not a code defect**, and it is the kind that costs a whole session: nothing in `10-localisation.md` §2.2 or in any runbook tells a prober which instrument to use where. **Fix shape:** four lines in `10-localisation.md` §2.2 — a table of surface → language instrument — and a pointer from the live-validation section of the workspace rules. Recorded meanwhile in `audits/i18n-closing-2026-09-21/live/backups-page.md`. | **READY — rank P3-LOW; owner: CC (docs)** | | **R-603** | **[P3-LOW] An English string containing an apostrophe silently never matches on a rendered page, and a `strings.Contains` assertion reads exactly like a missing sentence.** FOUND 2026-09-21 while writing the R-598 render tests. `backup.target.absent` was first written as *"The system backup's drive cannot be reached…"*; `html/template` escapes `'` to `'`, so the page carried the sentence and every assertion for it failed. **The failure mode is the expensive part:** the test said *"the English absent-drive copy never reached the page"*, which is indistinguishable from the handler not being wired — and the obvious next move is to go and re-fix the handler. Reworded to avoid the possessive, and all 23 new English values were then swept for `' " < > &` (zero). **The Hungarian bundle has never hit this** because Hungarian copy uses „quotes" and few apostrophes; **English copy will hit it again.** **Fix shape:** either a bundle gate that refuses an HTML-escapable character in a value destined for a page (and an allow-list for the ones that legitimately need one), or a test helper that compares against `html.EscapeString(want)` so the assertion cannot be fooled. The gate is the better shape — the helper only protects tests that remember to use it. | **READY — rank P3-LOW; owner: CC (controller)** | @@ -800,6 +804,19 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server` | **R-531** | **[P3-LOW] Three supervisor facts measured live and not pinned: restart timing during a deploy was not measured; restarts before the hub first sees the stanza produce no `controller_restarted_by_agent`; deliberate operator kills spend the crash-loop budget.** MEASURED 2026-09-15 on 9201: after 3 test restarts in 13 minutes the 4th kill tripped the 30-minute pause and the dashboard stayed down (the guard as designed, `A4-kill-middeploy-9201.txt`). The hub checker seeds silently on first sight, so the three restarts before the first v0.131.0 report emitted nothing (only the crash-loop did). **What it needs:** a deploy-kill timing on a fresh budget; the operator's view whether a restart after minutes of uptime should count toward the budget **MEASURED 2026-09-16 on the drill box, both halves.** (1) **Timing during a deploy (F9'):** the controller was killed 5 s into a deploy on an EMPTY budget; the agent saw it on the next sweep, confirmed on the one after, and the dashboard answered 200 again **37 s** after the kill; the interrupted app ended `not_deployed`, not stuck. (2) **The budget's shape (F9''):** three further kills at idle, 20 minutes apart, recovered in **61 s / 41 s / 61 s** - and NONE of them accumulated, because the window is 15 minutes. Four restarts this session, zero pauses, zero crash-loop events. **So the brake catches a FAST loop and is blind to a SLOW one:** a controller dying every 20 minutes is restarted forever and the only trace is an `info` event that mails nobody. That is a design question for the operator (leave it / add a longer second counter / raise the severity of the Nth restart in a day), and this session deliberately measured it without changing it. Evidence: `audits/evidence-drill-0243-2026-09-16/phase2-f9prime.txt` and `phase2-f9dprime.txt`. | **READY — rank P3-LOW; owner: CC (measure) · operator (budget rule)** | | **R-532** | **[P3-LOW] Vaultwarden's `/api/config` still says `disableUserRegistration:false` with signups off, so the web vault shows a register form that the server then refuses.** MEASURED 2026-09-15 in the E.1 spike. Cosmetic: the server refuses (400). A household following the invite-first card is not affected; a stranger sees a form that fails. | **READY — rank P3-LOW; owner: CC (catalog/upstream note)** | +| **R-615** | **[P3-LOW] Pointing a box at a different app catalog by `git.repo_url` alone is INERT — the box keeps fetching from the repository it first cloned.** FOUND 2026-09-21 by reading `sync.go` **before** running it, which is the only reason the update night's drill catalog worked at all. `Syncer.gitCloneOrPull` (`controller/internal/sync/sync.go:274-306`) clones **only when `/catalog-cache/.git` is absent**; on every later cycle it runs `git fetch --depth 1 origin ` + `git reset --hard origin/` **against the remote stored in the clone**, which `buildRepoURL` wrote at clone time. Changing `git.repo_url` in `controller.yaml` and restarting therefore changes **nothing**: the sync keeps pulling the old catalog and reports success. Measured: after the repoint, `git -C /catalog-cache remote -v` still read `app-catalog-felhom.eu`; the box only followed the drill repo once the cache directory was removed. **Why it matters beyond a drill:** this is the one knob that would move a box to a different or a staged catalog — for a migration, a per-customer catalog, or a rollback of the catalog itself — and it silently does not work. **Nothing is wrong with the CACHING**, which is right; what is missing is that a changed `repo_url` must invalidate the clone. **Fix shape:** on start, compare `git.repo_url` with the clone's `origin` and re-clone when they differ (or `git remote set-url` + a full fetch); log which happened. A test that changes `repo_url` under an existing cache and asserts the next sync reads the NEW repo — it fails today. Evidence: `audits/update-night-2026-09-21/04-9202-config-pre.txt`, `05-9202-follows-drill.txt`. | **READY — rank P3-LOW; owner: CC (controller)** | +| **R-616** | **[P3-LOW] The catalog credentials are stored in PLAINTEXT in the box's catalog clone and are printed by an ordinary `git remote -v`.** FOUND 2026-09-21 on guest 9202 while pointing it at a private drill catalog. `Syncer.buildRepoURL` injects `username:token` into the HTTPS URL, and `git clone` persists that URL as the clone's `origin`, so `/catalog-cache/.git/config` holds the token in the clear and **any** diagnostic that prints the remote leaks it — which is what happened in this session's own transcript, and is the same shape as R-580 (`curl -w '%{redirect_url}'`). `maskRepoURL` exists and is used for the LOG lines, so the masking intent is already there; the stored remote is the half that was missed. **INERT ON THE FLEET TODAY** — the live catalog is public and `git.token` is empty on every real box — which is exactly why it should be fixed before it is not: the day the catalog goes private, every box carries a readable credential and every support session that runs `git remote -v` prints it. **Fix shape:** store the remote WITHOUT credentials and supply them per-fetch (a credential helper, `http.extraHeader`, or `GIT_ASKPASS`), and a test asserting the clone's stored `origin` contains no `@`. **Operator action from tonight, unrelated to the fix:** the Gitea `admin` token used for the drill repo was printed by that command and must be rotated. Evidence: `audits/update-night-2026-09-21/05-9202-follows-drill.txt` (redacted). | **READY — rank P3-LOW; owner: CC (controller); one operator action (rotate the Gitea admin token)** | +| **R-617** | **[P3-LOW] The Gitea API token this project uses for pushes cannot create a repository through the documented endpoint, but CAN through `repos/migrate` — so "the token cannot do it" was nearly recorded as a fact when the truth was "one endpoint refuses it".** FOUND 2026-09-21 creating the drill catalog. Both `~/.git-credentials` tokens carry `write:misc,write:notification,write:package,write:issue,write:repository`; `POST /api/v1/user/repos` requires `write:user` and answers **403**, and `POST /api/v1/admin/users//repos` requires `write:admin` and answers 403 too. `POST /api/v1/repos/migrate` with the same token answered **201** and created the private repository. **Why this is a row and not a note:** a session that stopped at the first 403 would have recorded "CC cannot create a Gitea repository" — an unfalsifiable capability claim of exactly the shape the workspace's standing rule 2 forbids — and every later drill would have been designed around a limit that does not exist. **What it needs:** one line in the operations notes saying which endpoint to use, and (optional, operator) a token scoped for the job so the migrate route is not load-bearing. Evidence: `audits/update-night-2026-09-21/03-drill-repo.txt`. | **READY — rank P3-LOW; owner: CC (docs)** | +| **R-618** | **[P1-HIGH] THREE apps are presented to the household as UNHEALTHY while they are working perfectly — and because the guarded Update waits on that same probe, a SUCCESSFUL update ends by STOPPING the working app and sending the household to a restore they do not need.** **RANK RAISED FROM P2 TO P1 BY A LIVE MEASUREMENT taken the same night, and the escalation is the whole point:** tandoor's Update 2.6.13 → 2.6.15 was pressed at 21:16:47 and entered `verifying` at 21:17:46. At **21:18:28** the NEW version was `Up 25 seconds` and answering **HTTP 200** on `/accounts/login/` through the household's own front door — while the controller, probing port 8080 where nothing listens, could not see it. `verifying` therefore cannot pass, the full `update.health_timeout` is spent, `Manager.failAndHold` runs `compose down`, and the app is STOPPED. **Nothing is lost** — the data is in the volumes and the restore works — **but one wrong port number in a template converts every successful update of that app into an outage plus an unnecessary restore, for every household running it.** Evidence: `audits/update-night-2026-09-21/14-tandoor-serving-while-verifying.txt`. MEASURED 2026-09-21 on guest 9202 (controller v0.261.0, catalog `f5f6a152b513`). Two shapes, one class: **(a) `tandoor` — the WRONG PORT.** `.felhom.yml` probes `port: 8080`; the container listens on **80 and nothing else** (`ss -ltn` inside it), the compose's own traefik label routes to 80, its own docker healthcheck reads `healthy`, and `/accounts/login/` answers **200** through the household's real front door. `GET /api/stacks/tandoor` nevertheless reads `state: "unhealthy"`. **(b) `zipline` — the WRONG PATH.** `.felhom.yml` probes `/api/health`, which zipline 4.6.1 answers **404 `Route GET:/api/health not found`**; **the compose healthcheck in the very same file uses `/api/healthcheck` and is correct and green.** `/dashboard` answers 200. The controller reads `unhealthy`. **This is the MIRROR of R-613** — that is a probe that passes on a broken app (a false GREEN, which no alarm catches); this is a probe that fails on a working app (a false RED). **IT DOES NOT ALARM, AND THAT SETS THE RANK:** `08-alarm-ladder.md` §4 puts `unhealthy` deliberately in the NOT-down set, so no dead-app event and no customer mail follows — the damage is what the household READS, plus anything that gates on `state`. **IT ALREADY COST A MEASUREMENT TONIGHT:** this drill's harness waited for `state == "running"` and hung for its full budget on tandoor, an app that was up the whole time. An instrument waiting for a wrong answer looks exactly like a slow app. **THE GATE THIS WANTS IS CHEAP AND STATIC, AND THAT IS THE FINDING'S REAL VALUE.** Both halves of the answer live in the same template: compare the `.felhom.yml` probe's port and path against the compose's **own** `healthcheck: test:` URL. A sweep of all 53 templates on that rule was run tonight and returns **five** disagreements: `tandoor` (PORT — **CONFIRMED live**), `zipline` (PATH — **CONFIRMED live**), `wger` (PORT, probe 80 vs compose 8000 — **CONFIRMED live the same night**), `home-assistant` (PATH, `/api/` vs `/manifest.json` — **NOT MEASURED**), and `adventurelog` (a FALSE POSITIVE of the sweep's own regex — it reads `running` live). **So the rule finds both real defects, with two candidates and one false positive out of 53** — a good enough signal for a fast gate, provided it reports candidates rather than convictions and a person or a runtime check resolves them. The earlier, cruder rule (probe port vs the *traefik* port) is strictly worse: it clears zipline and convicts adventurelog. **AND A SECOND FIX SHAPE, ON THE CONTROLLER SIDE, WORTH CONSIDERING BESIDE THE CATALOG ONE:** in both confirmed cases the container's OWN docker healthcheck was **green** the whole time. A `verifying` phase that is about to stop a working app could ask that too — if the compose declares a healthcheck and docker reports `healthy`, the app is alive whatever our probe thinks. That does not excuse a wrong probe, but it turns this failure direction from an outage into a wrong label. It is a design question, not a defect, and is raised here rather than decided. **Needs:** fix tandoor's port (80), zipline's path (`/api/healthcheck`) and wger's port (8000) — **all three are now CONFIRMED live, none is a guess**; add the static gate with a decoy each way (R-421) — a template whose probe agrees must not read as a disagreement, and vice versa. **THE GATE'S RULE WAS THEN SHARPENED BY READING `healthprobe.go` RATHER THAN ASSUMING IT, and the sharpening REMOVED a false conviction.** `type: http` treats **any** response as healthy (`healthprobe.go:258-261`), and `type: api` with **no** `expect` block does the same (`:265-268`); only `type: api` WITH `expect.status` cares about the path or the code. So a PATH difference is a candidate only for the third shape, while a PORT difference is a candidate for all of them. Under that rule the 53-template sweep returns **four** candidates — `tandoor`, `zipline` and **`wger` (all three CONFIRMED live — probe `type: http, port: 80`; inside the container port 80 is `refused` and port 8000 `ANSWERED`; docker's own healthcheck green; front door 302; the box reads `unhealthy`)**, and `adventurelog` (a false positive: its compose lists two containers' ports and the probe targets the backend; measured `running`). **`home-assistant` is correctly CLEARED by the sharpened rule** — `type: api`, no `expect`, so its `/api/` answering 401 without a token is healthy, and its edge was PROVEN on the box tonight. The crude rule convicted it; the rule read from the code does not. **That is the gate to build: two of 53 convicted, one suspected, one false positive, and the false positive is resolvable by one live check.** Evidence: `audits/update-night-2026-09-21/10-probe-port-sweep.txt`, `12-probe-vs-compose-healthcheck.txt` and `13-probe-sweep-sharpened.txt`. | **READY — rank P2-MEDIUM; owner: CC (catalog)** | +| **R-619** | **[P3-LOW] A `type: password` deploy field is MANDATORY however `required` reads, and the `deploy-fields` contract says the opposite — so any caller that trusts it is refused.** MEASURED 2026-09-21 on guest 9202 while widening the update drill. `GET /api/stacks/grafana/deploy-fields` serves `{"env_var":"GF_SECURITY_ADMIN_PASSWORD","type":"password","generate":"password:16","required":false}`; a deploy carrying only the two `required:true` fields is refused **400** „a(z) „Admin jelszó" mező kitöltése kötelező — használja a Generálás gombot…". **The BEHAVIOUR is right and is a decision, not a bug:** `deploy.go:305-312` refuses a `password` field with no caller value on purpose — *"We never silently auto-generate — the user needs to know their password"* — which is the opposite of the `secret` case one branch above, where a generated value the customer never sees is exactly correct. **The defect is the CONTRACT.** `.felhom.yml` declares `required: false`, the API serves that verbatim, and nothing on the wire distinguishes "optional because the box will generate it" (`secret`) from "optional in the template and mandatory in the code" (`password`). A person using the deploy page never meets this because the page renders a Generálás button; **anything that is not that page does**, which now includes this drill harness and would include `09` §6.2's unattended caller the day it deploys anything. **Fix shape (smallest that keeps the decision):** serve `required: true` for `type: password` in the deploy-fields response — one place, derived rather than stored, so templates need no edit — and a test asserting a `password` field always reaches the wire as required. Alternatively state it in the field's `description`, which is weaker because it is prose. Evidence: `audits/update-night-2026-09-21/apps/grafana/log.txt` (the refusal) and `batchA.log`. | **READY — rank P3-LOW; owner: CC (controller)** | +| **R-620** | **[P3-LOW] A disabled notifier drops every event with NO local trace, so a box whose hub configuration is absent or broken stops telling anyone anything and leaves nothing behind that says so.** FOUND 2026-09-21 on guest 9202 while trying to score the update night's alarm truth table. `hub.enabled: false` there, and `Notifier.Publish` returns at `notify/notifier.go:269` — **before** any log line — as do `NotifyHealthChange` (:359) and four more entry points. Startup says it once (`[INFO] Notifier disabled (hub not configured)`) and then every later event, of every severity up to `critical`, vanishes without a word. **The measurable consequence tonight:** the whole event-and-mail half of the drill was structurally unmeasurable on this venue, and the alarm truth table below covers only the app page, the dashboard and the box's own log. That is a cost this session paid and named; the next one would pay it again. **The consequence on a real box is smaller but not zero:** the fleet's boxes have the hub enabled, and total silence is already caught by the hub's dead-man's-switch (staleness from the LAST REPORT, proven in the 2026-07-22 power-outage audit). What is NOT caught is the in-between — a box that still reports but whose notifier was disabled by a bad config push would go on reporting healthy while dropping every alarm, and the only evidence would be a single INFO line at the last restart. **Fix shape:** one DEBUG (or WARN, once per event type) line on the disabled path naming the event that was dropped, so the absence is visible where it happens rather than inferable from a startup line. Cheap, and it converts an invisible failure into a greppable one — R-96 rule 3 in the place that produces it. Evidence: `audits/update-night-2026-09-21/11-notifier-disabled.txt`. | **READY — rank P3-LOW; owner: CC (controller)** | +| **R-621** | **[P2-MEDIUM] A held update DESTROYS the evidence of why it failed: `failAndHold` runs `compose down`, the failing containers are removed, and their output is gone before anyone — household, operator or the next session — can read it.** MEASURED 2026-09-21 on guest 9202 on a REAL upstream edge: `adventurelog v0.12.1 → v0.13.0`. The new backend applied **nine Django migrations successfully** and then never listened; the update held after the full 5-minute health wait. **`Manager.failAndHold` (`stacks/update.go:723`) calls `updateCompose(dir, env, "down")`**, which removes the containers rather than stopping them, and nothing captures their logs first. Within seconds the box's own log recorded `Logs result for adventurelog: 0 bytes returned (empty)` and `docker ps -a` held nothing at all. **What survives is the WHAT and not the WHY:** the controller line `update adventurelog FAILED after the new version was started: not healthy: not healthy within 5m0s (last: state unhealthy)` and the household's sentence, both of which say the app did not come up and neither of which says the migrations ran and the server then failed to bind. **This is R-320 ("evidence off the machine before the teardown") as a PRODUCT behaviour rather than a session habit** — the teardown here is the product's own, it is correct to perform (a half-started new version must not keep running), and it happens before anyone can look. **Why it matters beyond a drill:** the hold sentence sends the household to a restore, and after the restore the only remaining question is *should I press Update again?* — which nobody can answer, because the one artefact that would say so no longer exists. It also makes every future held update unreportable to an upstream project. **Fix shape:** capture `compose logs --no-color --tail N` into the stack directory (beside `applied-compose.yml`, which already travels with the stack) IMMEDIATELY before the `down`, and surface it on the app page's hold panel or at least through the existing `/api/stacks//logs` fallback. Bounded size, written once per hold. A test that holds an app and asserts the captured file is non-empty — it fails today. Evidence: `audits/update-night-2026-09-21/apps/adventurelog/why-it-failed.txt`, `state-after-hold.txt`. | **READY — rank P2-MEDIUM; owner: CC (controller)** | +| **R-622** | **[P2-MEDIUM] `adventurelog v0.13.0` migrates the customer's database and then does not serve — the edge must NOT be promoted, and it is the first real-catalog candidate this project has measured as unsafe.** MEASURED 2026-09-21 on guest 9202 through the product's own guarded Update. `v0.12.1 → v0.13.0` (backend AND frontend together, PostGIS held constant). The backend applied **nine migrations, every one `... OK`** — `adventures.0072_trail_wanderer_author_fields` through `integrations.0009_alter_endurainintegration_auth_method`, plus `billing.0001_initial` — and then the container's own healthcheck failed with `URLError: [Errno 111] Connection refused` on five consecutive checks. The app never bound its port. The update held honestly after the full 5-minute wait. **THE PRODUCT DID EVERYTHING RIGHT AND THAT IS HALF THE FINDING:** the precondition found a Tier-1 copy one minute old, the safety dump was written, the pin advanced BEFORE the pull, the health wait was not short-circuited, the app was stopped rather than left half-running, and the hold sentence named the tier, the date and what the copy holds — „saját meghajtó, 2026-09-21 20:47 — ez a másolat a beállításokat, az adatbázist és az adatköteteket tartalmazza." **This is exactly the case `09` §4 exists for:** the migration RAN, so there is no undo, only a restore — and the restore is the thing slice 4 made sure existed first. **What it needs:** adventurelog stays OFF the promotion list; the cause is not diagnosed here (R-621 is why); and before it is ever promoted the edge should be re-run on the harness with its ABORT step, since an app that migrates and then refuses is the shape most likely to refuse the old image too. Evidence: `audits/update-night-2026-09-21/apps/adventurelog/`. | **READY — rank P2-MEDIUM; owner: CC (catalog); NOT a Felhom defect — an upstream edge that fails** | +| **R-623** | **[P3-LOW] The unattended-update caller turned every SUCCESS into a `timeout`, and then refused to press that app again — the instrument, not the box.** FOUND 2026-09-21 (update night) by reading `unattended-caller.py` before relying on it for the Q4 hold measurement. Its `call()` returns the API **envelope** — `{"ok": true, "data": {…}}` — and `follow()` read `update_phase` and `updating` **off the envelope**, where neither exists. Both were therefore always `None`; the end test `not updating and phase in ("done","failed")` could never fire; every followed update ran the full **900-second** timeout and was recorded `timeout`, which the caller treats as terminal and adds to `never_again`. `main()` unwraps `data` for the stack LIST, which is exactly why the within-a-major half of that night worked and this half did not. **The 2026-09-21 run did not catch it because the only pass that reached `follow()` was Scenario F, whose log was lost to a buffering `tail`** — the run's own honestly-recorded instrumentation gap turns out to have hidden a second one underneath it. **This is the R-607 class in the evidence layer rather than the product layer: an instrument that can report a success as a timeout is not a measurement**, and worse, it is a measurement that says the box behaved badly when the box behaved well. **FIXED in the same file 2026-09-21** (unwrap `data`, with the reason written into the docstring so the next reader does not re-derive it), and the fixed caller is what produced tonight's unattended-hold leg. **What it does NOT invalidate:** the G-b no-retry proof, which is entirely in the refusal path and never reached `follow()`. **What it DOES qualify:** any future reading of that night's Scenario F timing — the "51 s – 1 m 26 s" figures come from the ATTENDED scenarios 04/05/07, not from the caller. | **CLOSED 2026-09-21 — fixed in `audits/update-arc-gaps-2026-09-21/unattended-caller.py`** | +| **R-624** | **[P3-LOW] Three of the catalog's apps cannot be seeded by ANY headless route, and for two of them that is a deliberate security decision — so the upgrade harness has a permanent ceiling nobody has written down.** FOUND 2026-09-21 while widening R-462 from 3 apps to . **`vaultwarden` and `zipline` close self-registration ON PURPOSE** — vaultwarden by `SIGNUPS_ALLOWED=false` (R-512, *„a stranger who guesses vault. must not be able to register"*), zipline by answering `E1037: User registration is disabled` — and neither ships a CLI that could make an account instead. So there is no route to a first account without the admin secret, and **that is correct**: the harness must not be the reason a customer-facing app accepts strangers. **`gitea` is a different and fixable case:** the template sets no `INSTALL_LOCK`, so a fresh instance sits in its web-installer state and `gitea admin user create` refuses (`MustInstalled() [F] Unable to load config file for a installed Gitea instance`); POSTing the installer form first would work and was simply not written tonight. **Why this is a row rather than three notes:** `09` §3 decision 6 says the upgrade test goes to **all** apps, and R-462 is costed as if every app is reachable given enough fixture work. It is not. There is a class — *apps whose only account-creating route the catalog deliberately closes* — for which the honest maximum is `inconclusive` unless the harness is given the app's admin secret at deploy time, which is a decision nobody has taken. **Needs:** the class named in R-462's scope so the remaining count is honest; a decision on whether the harness may hold an app's admin secret (it already holds the ones IT generates — see R-619); and, separately and cheaply, a gitea installer-form fixture. Evidence: `audits/update-night-2026-09-21/apps/{vaultwarden,zipline,gitea}/verdict.json`. | **READY — rank P3-LOW; owner: CC (catalog harness)** | +| **R-625** | **[P2-MEDIUM] A HELD app keeps inviting the household to update it, and the button then refuses — the exact inconsistency R-524 removed for the other case, still present for this one.** MEASURED 2026-09-21 (update night, leg B6). `glance` was HELD by a genuine unattended failed update. The drill catalog then published a **fixed newer version** — a real forward route, the thing a household would hope for. Afterwards the app page read **„Frissítés elérhető — ma" / „Update available — today"**, in both languages, with the Update button offered; pressing it answered **`409 reason='held'`** and the hold sentence. **The BEHAVIOUR is correct and is DESIGN, not a defect:** `09` §6.1 says the hold is `settings.RestoreHold` and that **a successful unit restore lifts an update hold (only that kind)** — a newer catalog version does not, and should not, because nobody has checked that the new version can start on data the failed one may have touched. **The DEFECT is that the page says otherwise.** R-524 settled precisely this shape for the Ahead case — *"(a) is free and offers a household a downgrade … (c) show „Naprakész" and refuse the button … Why (c): the direction was already settled"* — and chose to make the badge and the button agree. The HELD case still has them disagreeing, in the more painful direction: the badge invites, the button refuses, and the refusal is the same long sentence the household has already read. **What it needs:** the badge for a held app should say what is true — that the app is held and the way back is the restore — and the Update button should not be offered while `RestoreHold` stands. One verdict, read by both surfaces, exactly as R-524 did it. **AND A QUESTION FOR `09` §3b Q4 THAT THIS MEASUREMENT RAISES AND DOES NOT ANSWER:** the household's ONLY route out is a restore, even when the catalog has already shipped a fix. That is defensible, but it is now measured rather than assumed, and Q4's "does the box try again?" should be read next to it. Evidence: `audits/update-night-2026-09-21/bad-days/B6-way-out-forwards/result.json`. | **READY — rank P2-MEDIUM; owner: CC (controller)** | +| **R-626** | **[P2-MEDIUM] An app the customer REMOVED came back: the removal returned 200 and deleted the record and the volume, a container was created two seconds later, and Docker's restart policy has kept it running ever since — while the box reports the app as not installed.** FOUND 2026-09-21 during the update night's TEARDOWN, which is the only reason it was found at all. `navidrome` was removed through the product: the `remove_hdd_data:true` call was correctly refused `409` (R-442's fail-closed guard — the drive path could not be resolved on this guest), and the `remove_hdd_data:false` call returned **200** with `volumes_removed: ['navidrome_navidrome_data']`. **Two seconds later a container carrying `com.docker.compose.project=navidrome` was CREATED** (`.Created = 19:12:20Z`), and Docker's `restart: unless-stopped` started it again at the next guest boot (`.StartedAt = 20:19:33Z`, the B5 power cut). Seven hours later: **`app.yaml` absent, `deployed=false`, one volume back, and the controller happily probing it — `Health probe navidrome: API GET :4533/ping → 200`.** **The customer-visible shape is the bad one:** *"I deleted that app and it came back"* — and it came back **blank**, because the volume really was deleted, so it looks installed and is empty. It is also invisible to every sweep that keys on `deployed`, which is exactly why the teardown found it and nothing else did. **WHAT IS NOT ESTABLISHED, and is stated rather than guessed: what created the container.** The controller was restarted several times later in the night and its log no longer reaches that moment — **the second time in one night that a restart destroyed the evidence of the thing that mattered** (see R-621). **Needs:** reproduce with a loop that removes an app and watches `docker events` for 60 s, so the creating path is a NAME and not an inference; then a test that removes an app, reboots, and asserts no container with that compose project exists. **And one instrument lesson worth keeping:** this session's own post-remove check queried the compose-project label and reported clean at 21:12:18 — two seconds before the container appeared. A check that runs once, immediately, cannot see a thing that is created immediately after it. Evidence: `audits/update-night-2026-09-21/26-removed-app-came-back.txt`. | **READY — rank P2-MEDIUM; owner: CC (controller)** | +