diff --git a/documentation/backlog/OPEN-ITEMS.md b/documentation/backlog/OPEN-ITEMS.md index b4e35c17..a9b8e5a0 100644 --- a/documentation/backlog/OPEN-ITEMS.md +++ b/documentation/backlog/OPEN-ITEMS.md @@ -801,7 +801,7 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server` | **R-631** | **[P3-LOW] Five templates cannot be judged by the probe gate at all, and one is correct only by accident.** FOUND 2026-09-22 when `check-probe-matches-compose.py` was run over all 53. The gate's oracle is the probed service's own compose `healthcheck.test`; where that dials no loopback URL the gate has nothing to compare and reports a WARNING rather than a pass. **`crafty-controller`, `mealie` and `uptime-kuma`** run their healthcheck through a python or script helper (`ssl._create_unverified_context()`, `socket.create_connection`, `extra/healthcheck`), so the port is inside code the gate does not execute. **`vikunja`** has no compose healthcheck on its probed service at all. **`home-assistant`** is the interesting one: its probe path `/api/` differs from the compose's `/manifest.json`, and it is healthy today **only because its check is `type: api` with no `expect` block, which `probeHTTP` treats as "any response is healthy"** — add `expect: {status: 200}` to that template, a change that looks like a tightening, and the app goes permanently unhealthy and every successful update of it starts stopping it. **The gate warns on it for exactly that reason and refuses to call it a pass.** **Needs:** a live probe reading for each of the five on a scratch guest — deploy, read `GET /api/stacks/`, compare with the front door — which is one rotation night's work and closes the last gap R-618 left. **CLOSED 2026-09-22 — all five read live on guest 9202, and all five probes are CORRECT.** Each app was deployed, its listening sockets read from inside the probed container, and the probe's own target dialled **on the compose network**, which is the call the controller makes. `mealie` `tcp 9000` → listens `0.0.0.0:9000`, dial 200. `uptime-kuma` `http 3001` → listens `*:3001`, dial 302 (and `http` calls any response healthy, so 302 passes and proves something answers). `vikunja` `api 3456 /api/v1/info` **with `expect: {status: 200}`** → dial **200** — the one that could have failed, because its expect block compares the code. `crafty-controller` `tcp 8443` → the controller's own log reads `Health probe crafty-controller: TCP :8443 -> ok (1ms)` twice, six minutes apart. **`home-assistant` is the one to carry forward:** `api 8123 /api/` with NO expect → the dial returns **401**, not 200. It reads healthy only because `probeHTTP` treats any response as healthy for that shape (`healthprobe.go:253-262`). **Add `expect: {status: 200}` to that template — a change that looks like a tightening — and home-assistant goes permanently unhealthy and every successful update of it starts stopping it.** That is R-618 one edit away, now a measured number rather than a caution. **So the gate's WARN list is not a backlog of suspects: it is four correct templates the gate honestly cannot prove, plus one correct by accident.** Evidence: `audits/the-28-2026-09-22/sidejobs/r631.json`. | **CLOSED 2026-09-22 — five live readings, five correct probes; home-assistant's fragility is now a number** | | **R-632** | **[P3-LOW] Twenty-eight of the 53 templates have never been deployed by any update drill, so nothing is known about whether their updates work.** COUNTED 2026-09-22 against the 2026-09-21 sweep, which is the widest one ever run. **20 apps have a verdict record** (14 proven, 3 failed, 3 inconclusive, plus tandoor re-walked to proven on 2026-09-22); **4 more were deployed as props in the bad-days legs with no edge walked** (`bentopdf`, `glance`, `uptime-kuma`, `wishlist`); **1 was deployed only to measure its probe** (`wger`); and **28 have never been deployed at all**: `calcom`, `calibre-web`, `claper`, `code-server`, `crafty-controller`, `emby`, `ghost`, `gokapi`, `gramps-web`, `homebox`, `homepage`, `immich`, `jellyfin`, `kimai`, `komga`, `onlyoffice`, `outline`, `paperless-ngx`, `plant-it`, `plex`, `radarr`, `rallly`, `recipe-importer`, `seerr`, `sonarr`, `sparkyfitness`, `termix`, `wanderer`. **THIS IS NOT A COMPLAINT ABOUT THE SWEEP** — it took one app-catalog-wide night to go from 3 apps ever measured to 21, and a night is the unit available. It is a record of what the catalog's update promise currently rests on: **for 28 of 53 apps, nothing.** **The list is the nightly rotation's queue**, smallest and least stateful first; `paperless-ngx` should be early because R-630 needs a live reading from it anyway, and `crafty-controller`, `mealie` and `uptime-kuma` should be early because R-631 needs one from each. Machine-readable copy: `audits/probe-fix-2026-09-22/not-judged.json`. **WORKED IN ONE NIGHT, 2026-09-22 — all 28 walked, so this row CLOSES and hands its findings to others.** Every one was installed on guest 9202 against the private drill catalog and taken through the same walk: deploy at the live pin, seed through the app's own front door, read it back, „Mentés most”, the guarded Update where a real within-a-major edge exists upstream, **restore from that copy and read the seed back a second time** (the half the update night skipped), then remove and a 60-second check that nothing came back. **26 of 28 deployed; 6 proven; the rest inconclusive, no-edge or refused.** **What the night produced that this row could not have predicted:** R-630 raised to P1 by measurement (a stack with no probe container has its working app STOPPED by a successful update), R-633 (a remove during a restore leaves an orphan with a live public route), R-634 (an app running and healthy while recorded as not deployed, and then unremovable), and one real upstream edge that HELD honestly (`outline 1.9.1 → 1.10.1`). **Also settled:** `plant-it` is `lifecycle: abandoned` and the product refuses to install it — the only lifecycle-gated template in the catalog, and its gate is now proven live. Full record: `audits/DRILL-the-28-2026-09-22.md`. | **CLOSED 2026-09-22 — all 28 walked in one night; the findings live in R-630, R-633, R-634** | | **R-633** | **[P2-MEDIUM] A remove sent while a restore is still running reports success, deletes the app's record, and leaves a container restarting forever with a live public route.** MEASURED 2026-09-22 on guest 9202, controller v0.261.0, during the twenty-eight walk. `gokapi` was restored from its own local copy at 11:34:07 and removed at 11:34:22. `POST /backup/restore` answers **302 and works in the background**; the remove tore down what existed and the restore's own `compose up` then RE-CREATED the container at **11:34:24**. **Both calls returned success.** Twenty-five minutes later: `GET /api/stacks/gokapi` reads **`deployed: false`**, and `docker ps -a` shows `gokapi` **`Restarting (1)`** with `RestartCount` climbing, carrying its full traefik label set — including `traefik.http.routers.gokapi.rule: Host(`.enkisfelhom.hu`)`, **a rule with an empty subdomain**, because the deploy values that filled it were deleted with the app. Its own log loops *„Salt for admin password invalid, generating new salt… password does not appear to be a SHA-1 hash"* — the volume holding its config was removed correctly, so the binary can never start. **WHY THIS IS A ROW AND NOT A HARNESS ARTEFACT:** a household can press exactly these two buttons in exactly this order, the product accepted both, and **the remove reported success while leaving the orphan**. Presence of a success message is not evidence of a result. **WHAT IT COSTS:** an app the household believes is gone keeps a container in a restart loop, keeps a route registered on the public reverse proxy, and is invisible to every product surface because the controller no longer records the stack. Nothing in the alarm ladder fires: `08` §4 keys on stacks the controller KNOWS about. **This is R-626's class with the mechanism finally visible** — that row recorded a removed `navidrome` coming back and could not diagnose it because the controller had restarted; here the window is 17 seconds and both halves are in the evidence. **Needs:** the remove path to refuse, or to wait, while a restore for the same stack is in flight — and, either way, to verify the teardown rather than report success without looking. Evidence: `audits/the-28-2026-09-22/apps/gokapi/came-back-evidence.txt`, `apps/gokapi/log.txt`. **FIXED in v0.262.0, both halves.** **(a) The guard the product already had everywhere else:** `RemoveStack` now consults `UpdateGuards.Busy` (backup single-flight, restore status, app-data holds) and `IsUpdating`, and refuses with the app's own sentence in both languages. The product refused exactly this clash for `update` and for `restore` — the restore refusal even NAMES the blocking app — and `remove` was the one door with no lock on it. **(b) `down` returning 0 is a request, not a result:** the compose project is now watched for 25 s after `down`, anything carrying its label is removed BY NAME with its labels logged, and the API answer carries `verified: true/false`. That is the instrument R-626 asked for, in the product rather than in a drill script. **PROVEN LIVE ON 9202, and the live proof found a defect a test had not.** The guard fires: with `privatebin` STOPPED (so the pre-existing "still running" check could not answer first) and a box-wide backup in flight, the remove was refused and the controller said `RemoveStack privatebin REFUSED (busy): a backup or restore is running (single-flight held)`, with the household's own sentence *„Az alkalmazáson mentés vagy visszaállítás fut. Várd meg, amíg befejeződik.”* **But it answered HTTP 500.** `router.go`'s status mapping greps the error TEXT for `not deployed` / `still running` / `not found` / `protected`, and the busy sentence contains none of them, so it fell through to the default. **A 500 tells the UI something broke; this is “wait a moment”.** Fixed in **v0.262.1**: a typed `*stacks.RemoveBusyError` matched with `errors.As`, answered **409**, carrying both the Hungarian bytes and the key — and its test asserts the sentence contains none of the words the text mapping greps for, so the type is load-bearing rather than decorative. **The verification half is proven too:** a permitted remove answered `verified: true` with no reappearance, and no container carrying the project label existed 60 s later. **AND A FIRST ATTEMPT THAT PROVED NOTHING, recorded because it nearly went down as a pass:** the first run refused the remove with `409 still running — stop it first`, which is the PRE-EXISTING running check, not this guard — the restore had already finished by then. A refusal from the wrong rule is not evidence for the new one. | **CLOSED 2026-09-22 — v0.262.0 + v0.262.1; refusal and verification both proven live, and the live proof corrected the status code — **re-proven on v0.262.1: HTTP 409 with the household's sentence**, `RemoveStack privatebin REFUSED (busy)`** | -| **R-634** | **[P1-HIGH] An app can be RUNNING, HEALTHY and serving while the controller records it as not deployed — and in that state the household cannot remove it through the product at all.** MEASURED 2026-09-22 on guest 9202, controller v0.261.0, on **two independent apps in one night**: `outline` and `sparkyfitness`. **The controller's own words, in order.** `outline` deploy accepted 11:54:47; **`11:55:51 StopStack outline: current state=deploying deployed=true containers=0`** — a stop while the stack is still deploying; `11:56:09 SaveAppConfig: saving /opt/docker/stacks/outline — 5 env vars, **0 encrypted**, 3 sensitive fields` (the two saves before it both read `3 encrypted`); then **`11:57:16` and `12:02:16 Health probe outline: API GET :3000/_health -> 200`** — the app is up and answering its own health endpoint; and **`12:02:19 StopStack outline: current state=running deployed=false containers=3`**. Three containers, state `running`, health 200, **`deployed=false`**. The remove then answers **`RemoveStack outline: state=not_deployed, deployed=false, orphaned=false, deploying=false` -> `[ERROR] Remove failed for outline: stack "outline" is not deployed`** for BOTH the remove-with-data and the remove-keeping-data call, while `ScanStacks` goes on finding the stack every ten seconds. `app.yaml` survives with `desired_state: stopped` and an **empty `installed_images`**. `sparkyfitness` produced the identical shape 13 minutes earlier. **WHAT THIS COSTS A HOUSEHOLD:** an app that works is invisible to the product as an installation — no badge, no update, no backup selection, and **no way to delete it**; the only exit is a shell. It is the mirror of R-633 (there the record is gone and the container remains; here the container is fine and the record is gone) and it is the **worse** of the two, because the app is serving customer traffic the whole time. **THE MECHANISM IS NOT DIAGNOSED, and this row says so rather than guessing.** What was tried: the controller's full container log for both apps (the sequence above), `app.yaml` on disk, `docker ps -a`, and `GET /api/stacks/`. What was NOT done: reading `runComposeDeploy`'s pin-write path — this was an unattended run and the brief forbade product code. **The one discriminator worth running first:** both apps were walked while two other walks ran concurrently, and `POST /api/backup/run` is box-wide, so a backup or restore for a NEIGHBOURING app was in flight. A serial re-walk is queued tonight; if it reproduces alone, concurrency is not the cause. **A THIRD THING THE SAME LOG SHOWS, recorded here because it is one line away:** at `11:55:57` the health probe dialled **`http://outline-postgres:3000/_health`** — during startup, with the exactly-named container not yet running, `findProbeContainer`'s PREFIX fallback latched onto the POSTGRES sidecar and probed port 3000 on it. Transient, and it resolved once `outline` came up, but it is the same function R-630 is about. Evidence: `audits/the-28-2026-09-22/apps/half-state-outline-sparkyfitness.txt`. **THE SERIAL RE-WALK WAS RUN THE SAME NIGHT AND IT SPLITS THIS ROW IN TWO — recorded here rather than left as the first reading.** Walked again one at a time, with no other walk running: **`outline` deployed normally and removed clean**, and **`crafty-controller` deployed normally, updated `4.10.7 → 4.11.0` to `done`, restored and removed clean.** So for those two the half-state did NOT reproduce alone, and concurrency — a box-wide `POST /api/backup/run` or a restore in flight for a NEIGHBOURING app — is implicated rather than the deploy path itself. **`sparkyfitness` reproduced EXACTLY, alone, in 534 s**: deploy accepted, never reached `deployed` with a pin, `app.yaml` left with `desired_state: stopped` and an empty `installed_images`, and **both remove calls refused with `stack "sparkyfitness" is not deployed`.** **So the row stands, at one reproducible app instead of three, and the honest split is:** (a) `sparkyfitness` has a deploy that does not finish and leaves a record the product cannot clear — reproducible, P1; (b) under concurrent work the same unremovable half-state can be reached by apps that are otherwise fine, which is the more alarming half because those apps were **running, healthy and serving** while recorded as not deployed. **Neither half is diagnosed** — `runComposeDeploy`'s pin write was not read, because the brief forbade product code. **THE BOUNDED HALF IS FIXED in v0.262.0; the MECHANISM IS STILL NOT DIAGNOSED, and this row stays open for it.** `RemoveStack` refused on `!stack.Deployed` — a FLAG — while the machine plainly had containers, a compose file and an `app.yaml`. It now asks whether anything EXISTS (`halfStateEvidence`): containers, a compose file, or an `app.yaml` are each enough, and it removes what exists with every other fence unchanged (R-442 stays fail-closed). **The household must always be able to remove what the box shows them** — that is true whatever the cause of the bad record. Red-proof seen failing. **What is NOT done:** why `deployed` stays false while containers run. `runComposeDeploy`'s pin-write path was not read against a fresh reproduction in this session. **THE UNREMOVABLE HALF IS PROVEN LIVE ON 9202:** `sparkyfitness` rebuilt in its exact measured shape — `app.yaml` and a compose file on disk, no containers, `deployed: false`, `state: not_deployed` — and the remove answered **200** with `leftovers: NONE`. Under v0.261.0 the identical call answered `stack "sparkyfitness" is not deployed`. | **OPEN — P1 for the MECHANISM only; the unremovable half is CLOSED in v0.262.0 and proven live** | +| **R-634** | **[P1-HIGH] An app can be RUNNING, HEALTHY and serving while the controller records it as not deployed — and in that state the household cannot remove it through the product at all.** MEASURED 2026-09-22 on guest 9202, controller v0.261.0, on **two independent apps in one night**: `outline` and `sparkyfitness`. **The controller's own words, in order.** `outline` deploy accepted 11:54:47; **`11:55:51 StopStack outline: current state=deploying deployed=true containers=0`** — a stop while the stack is still deploying; `11:56:09 SaveAppConfig: saving /opt/docker/stacks/outline — 5 env vars, **0 encrypted**, 3 sensitive fields` (the two saves before it both read `3 encrypted`); then **`11:57:16` and `12:02:16 Health probe outline: API GET :3000/_health -> 200`** — the app is up and answering its own health endpoint; and **`12:02:19 StopStack outline: current state=running deployed=false containers=3`**. Three containers, state `running`, health 200, **`deployed=false`**. The remove then answers **`RemoveStack outline: state=not_deployed, deployed=false, orphaned=false, deploying=false` -> `[ERROR] Remove failed for outline: stack "outline" is not deployed`** for BOTH the remove-with-data and the remove-keeping-data call, while `ScanStacks` goes on finding the stack every ten seconds. `app.yaml` survives with `desired_state: stopped` and an **empty `installed_images`**. `sparkyfitness` produced the identical shape 13 minutes earlier. **WHAT THIS COSTS A HOUSEHOLD:** an app that works is invisible to the product as an installation — no badge, no update, no backup selection, and **no way to delete it**; the only exit is a shell. It is the mirror of R-633 (there the record is gone and the container remains; here the container is fine and the record is gone) and it is the **worse** of the two, because the app is serving customer traffic the whole time. **THE MECHANISM IS NOT DIAGNOSED, and this row says so rather than guessing.** What was tried: the controller's full container log for both apps (the sequence above), `app.yaml` on disk, `docker ps -a`, and `GET /api/stacks/`. What was NOT done: reading `runComposeDeploy`'s pin-write path — this was an unattended run and the brief forbade product code. **The one discriminator worth running first:** both apps were walked while two other walks ran concurrently, and `POST /api/backup/run` is box-wide, so a backup or restore for a NEIGHBOURING app was in flight. A serial re-walk is queued tonight; if it reproduces alone, concurrency is not the cause. **A THIRD THING THE SAME LOG SHOWS, recorded here because it is one line away:** at `11:55:57` the health probe dialled **`http://outline-postgres:3000/_health`** — during startup, with the exactly-named container not yet running, `findProbeContainer`'s PREFIX fallback latched onto the POSTGRES sidecar and probed port 3000 on it. Transient, and it resolved once `outline` came up, but it is the same function R-630 is about. Evidence: `audits/the-28-2026-09-22/apps/half-state-outline-sparkyfitness.txt`. **THE SERIAL RE-WALK WAS RUN THE SAME NIGHT AND IT SPLITS THIS ROW IN TWO — recorded here rather than left as the first reading.** Walked again one at a time, with no other walk running: **`outline` deployed normally and removed clean**, and **`crafty-controller` deployed normally, updated `4.10.7 → 4.11.0` to `done`, restored and removed clean.** So for those two the half-state did NOT reproduce alone, and concurrency — a box-wide `POST /api/backup/run` or a restore in flight for a NEIGHBOURING app — is implicated rather than the deploy path itself. **`sparkyfitness` reproduced EXACTLY, alone, in 534 s**: deploy accepted, never reached `deployed` with a pin, `app.yaml` left with `desired_state: stopped` and an empty `installed_images`, and **both remove calls refused with `stack "sparkyfitness" is not deployed`.** **So the row stands, at one reproducible app instead of three, and the honest split is:** (a) `sparkyfitness` has a deploy that does not finish and leaves a record the product cannot clear — reproducible, P1; (b) under concurrent work the same unremovable half-state can be reached by apps that are otherwise fine, which is the more alarming half because those apps were **running, healthy and serving** while recorded as not deployed. **Neither half is diagnosed** — `runComposeDeploy`'s pin write was not read, because the brief forbade product code. **THE BOUNDED HALF IS FIXED in v0.262.0; the MECHANISM IS STILL NOT DIAGNOSED, and this row stays open for it.** `RemoveStack` refused on `!stack.Deployed` — a FLAG — while the machine plainly had containers, a compose file and an `app.yaml`. It now asks whether anything EXISTS (`halfStateEvidence`): containers, a compose file, or an `app.yaml` are each enough, and it removes what exists with every other fence unchanged (R-442 stays fail-closed). **The household must always be able to remove what the box shows them** — that is true whatever the cause of the bad record. Red-proof seen failing. **What is NOT done:** why `deployed` stays false while containers run. `runComposeDeploy`'s pin-write path was not read against a fresh reproduction in this session. **THE UNREMOVABLE HALF IS PROVEN LIVE ON 9202:** `sparkyfitness` rebuilt in its exact measured shape — `app.yaml` and a compose file on disk, no containers, `deployed: false`, `state: not_deployed` — and the remove answered **200** with `leftovers: NONE`. Under v0.261.0 the identical call answered `stack "sparkyfitness" is not deployed`. **THE MECHANISM, DIAGNOSED 2026-09-23 (evening) on 9202, controller v0.264.0, before any code** (`audits/cleanup-2026-09-23/12-*`, `13-*`). **Reproduced on demand:** deploy `outline`, press the whole-box backup 20 s later. The chain, at `file:line`: (1) `DeployStack` sets the in-memory `Deployed=true` at ACCEPT (`stacks/deploy.go:395`, the no-stale-button UX); (2) `stackAdapter.ListDeployedStacks` (`cmd/controller/main.go:2556`) filters on that flag only, so a DEPLOYING app is in the backup's list; (3) `runVolumeDumps` (`backup/backup.go:693`) calls `DumpAppVolumesSafe` on it — **`StopStack` runs `docker compose down` in the middle of the deploy** (`14:39:32 StopStack outline: current state=deploying deployed=true containers=0`), tars half-made volumes (1.5 KB each), and **`StartStack` runs a SECOND `compose up -d` while the deploy's own is still running** (`backup.go:918`, `14:39:33`); (4) both `compose up` calls fail at the same second (`14:39:48`, exit 1), and `runComposeDeploy`'s failure branch (`stacks/deploy.go:421-435`) writes `Deployed=false` to memory and disk without asking whether containers exist. **Which of the two `up` calls wins decides the end state:** on 2026-09-22 the backup's restart won → containers RUNNING under `deployed=false` (outline's 11:56:09 `installed_images`, recorded by `StartStack`'s `recordInstalledImages`); tonight both lost → containers `Created` under `deployed=false`. **Shape (i) of the brief is ruled out:** there is no deploy time limit (`composeExecCustomEnv` has no timeout). **`sparkyfitness` did NOT reproduce alone** at today's pin on v0.264.0: deployed in 47.3 s (`10-*`) — the 2026-09-22 "alone" reading is not reproduced and is most likely the same backup race inside the walk (the walk presses the whole-box backup). **Not a design choice:** the brief itself names the fix (a whole-box backup skips a deploying stack, as the update refuses `busy`). **One design question remains and goes to the operator:** when `compose up` fails for its OWN reasons and leaves containers behind, should the deploy remove what it started, or keep the record "failed" with the containers visible? | **OPEN — P1 for the MECHANISM only; the unremovable half is CLOSED in v0.262.0 and proven live** | | **R-635** | **[P1-HIGH] `romm 5.3.0` does not fit the memory the template gives it, and the guarded Update called that a success — the app has been OOM-crash-looping on demo-hp for six hours at ~500% CPU.** FOUND 2026-09-22 17:37 **because the operator heard the fans**, which is the only reason it was found at all. `romm` was promoted `5.0.0 -> 5.3.0` on the live catalog that morning (`audits/PROBE-FIX-2026-09-22.md`) after the edge was PROVEN on scratch guest 9202, and the guarded Update was then pressed on demo-hp guest 9201 at **09:08:22Z**, reaching **`done` in 74.8 s** with the app `running`. **It ran clean for two hours.** The first worker kill is at **11:09:20Z**; by 15:38Z there had been **4,530** of them — `Worker (pid:…) was sent SIGKILL! Perhaps out of memory?` — with `docker inspect` reading **`OOMKilled: true`**, the container pinned at **457 MiB of its 512 MiB limit**, and `docker stats` showing **499.51% CPU**. The host's load average sat at **5.2 while otherwise idle**. `rq_cron` is killed and restarted every few seconds in a permanent storm. **THREE THINGS THIS ESTABLISHES, and the third is the one that changes how promotions are judged.** (1) The template's own comment says *`RAM: ~300MB (mem_limit: 1024M total — romm 512M + mariadb 384M + redis 128M)`*; **5.3.0 needs more than 512M and the template was not re-sized when the version moved.** A version move is not only an `image:` line. (2) **The update reported `done` and the app reads `running`**, because nginx answers `GET /` with 200 while the gunicorn workers behind it are being killed — a THIRD variant of the R-618/R-630 theme: the probe is right, the port is right, and the answer is still a false green. (3) **THE ALARM DID FIRE, AND MY FIRST WRITE-UP OF THIS ROW SAID IT DID NOT — CORRECTED 2026-09-22 BY THE OPERATOR, WHO PRODUCED THE MAILS.** The controller HAS an OOM detector (`main.go:821`, `[deadapp] romm: container romm was OOM-killed`, evaluated every 30 s), it emits **`app_oom`** (`notifier.go:726`), the hub allow-lists it (`dispatcher.go:636`) and dispatched it to the OPERATOR channel as **SENT** — `admin@felhom.eu` received **`[Felhom] ⚠️ demo-hp: app_oom`** at **11:09 CEST** and again at **17:48 CEST**, each carrying the app name, the Hungarian sentence and the dashboard link. The CUSTOMER channel is **SKIPPED**, which is correct. **I asserted an absence without looking at the instrument** — the hub's own Events and Notifications tabs show all of it — and I did it by reasoning from a memory note (`lxc-docker-oom-signals-unreliable`, R-528) instead of reading the hub. **That is R-628's shape exactly, four days on, and from the same hand.** **WHAT IS ACTUALLY WRONG, and it is narrower and real:** `notifier.go:715-726` keys the alarm on `container|startedAt` in an `oomSeen` map and emits **once per container lifetime**. So **six hours of continuous thrashing — 4,530 worker kills — produced exactly ONE mail**, severity `warning`, never escalating, while the app went on reading `running`. **A six-hour storm is indistinguishable from a single transient kill.** The hub's App Telemetry panel did carry the magnitude (RomM: **5,023 errors, 632 warnings**, peak 855 MB) but nothing turns that into a second, louder signal. So the fix worth having is not a detector — there is one — but an ESCALATION: a warning that repeats for hours should stop looking like a warning that happened once. **AND THE LESSON FOR R-462's METHOD, which is the real cost:** every `proven` verdict in the update night and in the twenty-eight measures the app for the **minutes of the walk**, not for a day of running. `romm` passed its walk, was seeded, read back and restored — and broke two hours later. **`proven` currently means "the update applied and the data survived", NOT "the new version runs".** Needs: decide between rolling the catalog back to 5.0.0 and raising romm's `mem_limit` (measured, not guessed); and a soak longer than a walk before any future promotion. **FIXED AND MEASURED 2026-09-22, in two steps, and the FIRST step was still a guess.** *Step 1 (operator's choice):* the limit was raised 512M → **768M** (`app-catalog-felhom.eu@886956d`). It slowed the kills from ~12/min to ~7/min and **stopped nothing** — 37 SIGKILLs in five minutes, `OOMKilled` still true, and the cgroup's own `memory.peak` read **exactly 768 MiB**: it hit the new ceiling and died there. *Step 2, from a MEASUREMENT instead:* the per-process RSS inside the container reads **~216 MiB per warm uvicorn worker**, so the image's default of four workers plus the master needs **~882 MiB** before nginx and the job runner — more than any sensible limit for this box. **The lever was in the image all along:** `/init:143` runs `--workers "${WEB_SERVER_CONCURRENCY:-4}"`. **Four workers is a SERVER default on an appliance serving one household.** Setting **`WEB_SERVER_CONCURRENCY=2`** (`app-catalog-felhom.eu@f4eb94f`, limit left at 768M) and applying it through the product's own Update button fixed it. **PROVEN UNDER LOAD, not just at idle:** 6 concurrent callers driven at romm through the household's own route for 300 s — **26,645 requests** (9,687 × 200, 16,958 × 401 on the auth-gated endpoints), CPU a steady **~200%** (exactly two workers saturated, by design), memory oscillating **416–614 MiB against the 768 MiB limit and trending DOWN**, and **zero** SIGKILLs, `OOMKilled: false`, `RestartCount: 0` throughout. At idle afterwards: **1.64% CPU**, 610 MiB, host load falling from 5.2 to 2.4. **AND THE FIRST SOAK MEASURED NOTHING, which is worth more than the second one:** it was pointed at `arcade.enkisfelhom.hu` — the scratch-guest fixture's default subdomain — while this box deployed romm at `jatek`. Every request 404'd at traefik in 9 ms, romm sat idle at 0.64% CPU, and the counter cheerfully reported **14,026 successful requests**. It was caught only because 0.64% CPU under load is not believable. **A positive AND a negative control are now asserted before any load is driven** (`jatek` must not 404; a nonsense host must). R-96 rule 3, in a new surface. **WHAT STAYS OPEN, and it is the part that outlives romm:** 610 MiB of 768 MiB is **79%** — it works with ~158 MiB of headroom and the soak never exceeded 614 MiB, but it is not generous, and nothing watches it. **And the method lesson for R-462:** a version move is not only an `image:` line — the new version's SHAPE (worker counts, per-worker footprint) has to be measured too, and a walk lasting minutes cannot see a ceiling reached in two hours. Every `proven` verdict in the update night and in the twenty-eight means *"the update applied and the data survived"*, **not** *"the new version runs"*. Evidence: `audits/probe-fix-2026-09-22/romm-soak.json`, `romm-soak.out`. | **CLOSED 2026-09-22 — two workers, 768M, proven under 26,645 requests; the 79% headroom and the `proven`-means-minutes lesson are carried into R-462** | | **R-636** | **[P2-MEDIUM] An app that has been out of memory for six hours sends the same single warning an app that hiccuped once sends.** FOUND 2026-09-22 when the operator produced the alarm mails I had wrongly written off as absent (R-635). **The detector is correct and works.** `main.go:821` re-checks every 30 s and logs `[deadapp] romm: container romm was OOM-killed`; `notifier.go:726` emits **`app_oom`** at severity `warning`; the hub allow-lists it (`dispatcher.go:636`) and delivered it to the OPERATOR channel — two mails, 11:09 and 17:48 CEST, each naming the app and linking the dashboard. CUSTOMER is SKIPPED, correctly. **The defect is the SHAPE of the signal, not its absence.** `notifier.go:715-724` keys on `container|startedAt` in an `oomSeen` map and returns early on a repeat, so the alarm fires **once per container lifetime**. romm's 09:08 container produced **4,530 worker kills over six hours and exactly one mail**; the 15:47 container produced one more. Severity never escalates, the app goes on reading `running`, and `08` §4 rightly does not put it in `IsDownState`. **So a six-hour storm that pins five cores is indistinguishable, in the operator's inbox, from one transient kill at 3 a.m.** — and the operator, who had two correct mails, still found the fault by hearing the fans. **The once-per-lifetime rule is right in itself** (it is what stops a crash loop from mailing 4,530 times, which is R-629's lesson) — what is missing is the second, louder signal when the same key keeps re-firing. **The magnitude IS already collected:** the hub's App Telemetry panel read **RomM: 5,023 errors, 632 warnings, peak 855 MB against a 1280M limit** while the same panel showed every other app at 0. Nothing turns that into an event. **Candidate shapes, none chosen here:** escalate `app_oom` to `error` when the same `container|startedAt` key re-fires past a threshold; or a periodic digest for a key still firing after N minutes; or let App Telemetry raise its own event when an app's error count crosses a bound. All are hub/controller code. Evidence: the operator's screenshots of the Events, Notifications and App Telemetry tabs plus the two mails; `felhom-controller/controller/internal/notify/notifier.go:715-726`. | **OPEN — P2; owner: CC; product code, so not fixed unattended** | | **R-638** | **[P2-MEDIUM] The product's database loader cannot replay a copy over a NEWER schema: on PostgreSQL it FAILS, on MariaDB it leaves the newer version's tables behind.** MEASURED 2026-09-23 on 9202. `ImportDump` (`appbackup/dbdump.go:719`, `psql -v ON_ERROR_STOP=1 --single-transaction`) replays a `pg_dump --clean --if-exists` file over the live database. After docmost 0.95.0 → 0.96.0 migrated, the replay of the pre-update copy was refused in 0.40 s, rc 3: *cannot drop constraint workspaces_pkey on table public.workspaces because other objects depend on it / DETAIL: constraint oauth_clients_workspace_id_fkey …* — the new version created six tables whose foreign keys point at old ones, and `--clean` only drops what the dump knows. Database unchanged (the transaction rolled back). On MariaDB (`mariadb-dump`, `FOREIGN_KEY_CHECKS=0`) the same replay after romm 5.0.0 → 5.3.0 returned rc 0 in 1.25 s and left **12 base tables** of the new version behind; RomM 5.0.0 happened to ignore them. **What worked:** `DROP SCHEMA public CASCADE; CREATE SCHEMA public;` + the dump in ONE transaction — rc 0 in 1.38 s, every table, index and extension back. **Why this is a row of its own and not only part of R-637:** the SAME loader backs shipped paths — `rollbackSafetyDump` (off-site restore's undo) and the dump replay of the restores — so **any restore of a copy taken BEFORE an update that migrated, replayed over the migrated database, may fail the same way. NOT MEASURED:** whether the unit restore the hold sentence names does this (it also carries the data VOLUME tar, which may make the replay moot). That is the measurement owed, on 9202, before anyone relies on it. Evidence: `audits/update-rulings-2026-09-23/README.md` Part 1, `docmost-45`, `romm-44`. **-- NARROWED 2026-09-23:** the undo no longer touches this loader — it copies folders (decision 19, controller v0.263.0). **What stays open is the part about SHIPPED paths:** `rollbackSafetyDump` and the restores' dump replay still replay over whatever schema is live, and whether the unit restore the hold sentence names works after a real schema migration is STILL UNMEASURED. | **OPEN — P2, narrowed to the restore paths; owner: CC; measure the named restore after a real schema migration first** |