diff --git a/STATUS.md b/STATUS.md index 318ad0e0..32c2ee24 100644 --- a/STATUS.md +++ b/STATUS.md @@ -1,44 +1,27 @@ # STATUS — what works, what's broken, what's next -**Updated 2026-09-22 (overnight) — I installed and tested all 28 apps that no test had ever touched. Every app in our catalogue has now been tried at least once. Three things are quietly wrong, and one of them stops a working app.** +**Updated 2026-09-22 (late) — I fixed the six faults the two drill nights found in the update, delete and hold machinery, and shipped the six app versions you approved. One thing needs your word: whether the fleet moves to the new controller.** **Decisions I took on my own: none.** -**What I did.** Each of the twenty-eight got the same walk: install it at the version our catalogue offers today, put real data in through the app's own front door, back it up, update it if a newer version really exists, **restore it from that backup and read the data back again**, then delete it and check a minute later that nothing came back. That restore step is new — the update night skipped it. **Twenty-six of the twenty-eight installed. Six are proven end to end.** Fourteen had no newer version to move to tonight. One failed honestly. Two would not install, and one of those is meant not to. +**The one that mattered most is fixed and proven.** An app with no health check used to be **shut down by a successful update** — the machine waited five minutes for a check that could never arrive, then stopped a working app. Paperless-ngx, same app, same button: **before, it failed after 5 minutes and the app went dark. Now it finishes in 53 seconds and keeps running.** -**The thing I would fix first — an app with no health check gets shut down by a successful update.** Paperless-ngx is never health-checked at all: its containers are named differently from the app, so the machine looks for one, finds nothing, and moves on without a word. I always thought that was just a missing badge. It is not. **The update waits five minutes for a health check that can never arrive, then declares failure and shuts the working app down.** All three of its containers were healthy the whole time. The machine says so in its own words: *"not healthy within 5m0s (last: no probe container) — stopping and HOLDING the app"*. Every household running Paperless who presses Update loses their app and is sent to a restore they do not need. +**Five more, all proven on the test machine.** +- **Deleting an app while it is being backed up or restored is now refused**, with a plain sentence telling you to wait — instead of quietly tearing it down and leaving a ghost behind. +- **A delete now checks its own work.** The machine watches for 25 seconds afterwards and removes anything that comes back, and says whether it verified. +- **An app the machine has lost track of can now be deleted.** Before, if its record went wrong, no button worked and only a command line could clear it. +- **A failed update now keeps the app's own log** before shutting it down. Twice we lost the only evidence of why. +- **Deleting an app clears its old update status**, so a fresh install of the same app no longer shows a stale "Updated". -**Two more, both about the machine losing track of an app rather than its health.** -- **Deleting an app while it is being restored leaves a ghost.** Both buttons say they worked. The app vanishes from every screen, and a container keeps restarting on the machine, still holding a public web address. **The machine already knows how to refuse this** — it refuses an *update* while a backup runs, and refuses a second *restore* while one is going, and it even names which app is blocking. Delete has no such guard. -- **An app can be running perfectly while the machine records it as not installed — and then it cannot be deleted.** I saw this three times. Two only happened when several jobs ran at once; one happened on its own, repeatably. In that state there is no button that works. +**The six versions you approved are live on the catalogue** — Emby, Ghost, Immich, Radarr, Sonarr, Termix. **None of them is installed on either demo machine**, so nothing updated; they simply show as available. -**In all three cases I needed a command line to clean up what the product could not. A household has none.** +**What I did not do, and it is on purpose.** Two items from the plan are untouched and named rather than half-finished: finding out *why* an app's record goes wrong in the first place (I fixed the consequence, not the cause), and making a held app stop offering an Update button it will refuse. -**The best thing I saw.** Two apps keep their files outside the database, and their local copy does not hold those files. When I asked to restore them, the machine **refused** — and said, in plain Hungarian, that it will not put an old database on top of files it does not have, that the files stay where they are, and which button does work. That is exactly right. +**What went wrong on my side.** I lost **44 minutes** to my own progress-watchers: they waited for a build that had already succeeded, because each was watching for a name its own command contained. The same bug cost me a pile of stuck watchers earlier in the day. It is now written down as a rule so it does not happen a third time. I also nearly recorded one test as passing when it had proved nothing — the refusal I saw came from an older rule, not the new one. I caught it and re-ran it properly. -**What I got wrong.** My own test script had three bugs that cost nine apps their walk. I found them, fixed them, and walked those nine again one at a time — and that second pass is what corrected my conclusions and produced one of the six proofs. The report names all three. +**Rows opened and closed.** Four closed, one narrowed to what is still unknown. The list stands at 325. -**Rows opened and closed.** Three new (the two above, plus the one that raised Paperless to urgent). Two closed. The list went from 321 to 323. +**What needs you — one question.** +1. **Shall the fleet move to controller 0.262.1?** Right now only the test machine has it. **My recommendation is yes:** every change here only refuses, waits, records, or removes what someone already asked to remove — none of them makes the machine do more on its own. *If you do nothing:* both demo machines stay on 0.261.0 and keep all six faults, including the one that shuts down a working app. Peti's machine is parked and would take it only if it ever comes back online. -**What needs you.** -1. **The second promotion list** — six app versions this night proved safe enough to move on the real catalogue, and six named that must not move, each with the reason. It is in the report. Moving a version is your call, never mine. *If you do nothing:* nothing breaks; those apps drift further from upstream each month. - -**Nothing on your own machine, the tester's machine, or the off-site box was touched. No product code was written. The real catalogue was never changed — I checked its version lines against the start of the night and not one differs.** - ---- - -**Evening addition, 2026-09-22 — you heard the fans, and you were right.** - -**RomM was cooking the HP box and it was my fault.** I moved it to a new version that morning. The update said it worked, and it did — for two hours. Then it ran out of memory and spent six hours killing and restarting its own workers, about seven times a minute, burning five processor cores. - -**CORRECTION — the machine DID warn you, and I was wrong to say it did not.** You showed me the two e-mails: *"Alkalmazás memóriája elfogyott: romm"*, at 11:09 and again at 17:48. They are in the hub's Events and Notifications tabs too. **I wrote "nothing warned anyone" without opening either tab** — I went by an old note saying this signal was unproven and turned that into "it did not happen". That is the same mistake I made with the Hetzner tickets four days ago. - -**What is actually wrong is smaller and real:** the machine sends **one** warning per app start. Six hours of trouble and 4,530 worker deaths produced **one e-mail** — the same e-mail a single harmless hiccup would send. It never gets louder, and the app keeps showing as running. The hub *did* have the full picture on the App Telemetry page (RomM: 5,023 errors, 632 warnings, while every other app showed zero), but nothing turns that into a second, louder alert. **That is why a correct warning still got missed, and it is now written down as its own item.** - -**Fixed, and proven under real load.** Giving it more memory was not enough — I measured that rather than assuming it. The real cause was that the app starts **four web workers**, which is a server setting, on a box serving one household. It now starts **two**. I then drove **26,645 real requests** at it for five minutes: memory stayed between 416 and 614 MB against its 768 MB ceiling and **trended down**, with **zero** worker deaths. At rest it now uses **1.6% of a core** instead of 500%. - -**Two things I got wrong on the way, both worth knowing.** My first memory number (768 MB) was a guess and it did not fix anything. And my first load test was pointed at the wrong web address — every request bounced off the front door in 9 milliseconds while the counter reported 14,026 successes. I caught that only because the app looked *too* idle. Both are written up. - -**The lesson that outlives RomM.** Every "proven" result this week measured an app for the **minutes of the test**. RomM passed everything and broke two hours later. **Proven has meant "the update worked and the data survived", not "the new version runs".** That gap is now on the record. - -**Still worth an eye:** RomM sits at 610 MB of its 768 MB, about 79%. It works, but it is not roomy, and nothing watches it. +**Nothing on your own machine or the off-site box was touched. The demo machines were not touched — they only see the six new version badges.** diff --git a/documentation/architecture/00-capability-map.md b/documentation/architecture/00-capability-map.md index e060c614..4269c44b 100644 --- a/documentation/architecture/00-capability-map.md +++ b/documentation/architecture/00-capability-map.md @@ -102,7 +102,7 @@ likewise silent. Evidence: `audits/DRILL-r361-2026-08-22/evidence/06-part3-decis |---|---|---|---|---| | Deploy an app from the catalog (env config, memory guard, health-aware progress) | controller, catalog (~52 apps, images pinned) | **PROVEN-LIVE** | `CAMPAIGN-2` T-DEPLOY-SET (7 apps, env config, health-aware); `RERUN-p1p3` (×4 PASS) | Memory-guard FIRING is not live-shown (T-RES-MEMGUARD never fired: ample RAM / auth-walled) — implemented + unit-level only | | App lifecycle: start/stop/restart/update/logs/remove/redeploy | controller | **PROVEN-LIVE — the ACTIONS work. NARROWED 2026-09-13: `CAMPAIGN-3` proved `remove` removes the APP, not the DATA — the "delete my data" half was INERT on every box until controller v0.236.0 (R-442). RE-PROVEN 2026-09-13 on demo-hp: data written by the app itself (63 MB) gone after removal and listed; an unresolvable data location is REFUSED (409) with the app kept; an SSD app gets `[]` and a note.** | `CAMPAIGN-2` T-LIFECYCLE (stop/start/restart/update/logs); remove (app only) live in `CAMPAIGN-3`; **remove WITH data: `audits/R442-2026-09-13/`**; **data behaviour: `audits/SPIKE-app-update-2026-09-01.md` (2026-09-01)** | Redeploy-after-remove edge remains open (T-REMOVE-REDEPLOY never cleanly passed — stale dryrun journal); non-pilot-critical | -| **Update is GUARDED: it refuses without a restorable backup, backs up first when the copy is stale, and HOLDS an app that does not come up — on ANY backup tier, and the release itself arrives by the managed floor** | controller **v0.237.0 + v0.238.0 + v0.238.1 + v0.239.0**, hub **v0.112.0** | **PROVEN-LIVE (2026-09-13, and again the same afternoon for any tier + floor delivery)** — **afternoon (`audits/rulings-r472-r475-2026-09-13/`):** an undeclared floor above the golden refused with nothing stored (02); a declared floor 0.239.0 / MinAgent 0.129.0 served `from declared` and both demo boxes self-updated in 14 s and 15 s (03); nothing on any tier → backed up first, Tier 1 chosen, done (04); gokapi updated on its Tier-1 unit alone (05); a never-healthy update held naming „saját meghajtó" (07); restored from „helyi", hold cleared (08). **Morning:** scenarios A (real upgrade, success only after health), B (stale copy → backup first), E (pull failure → pin back, app untouched), F (never healthy → held, hold text on API and page), H (start/restart/update and the boot sweep all refuse the held app) and **the restore walk** (Mentések unit restore → back on the old version, hold cleared), on demo-hp with a throwaway app | **`audits/slice4-2026-09-13/`** (live/, redproofs/, gates/); design `architecture/09-update-architecture.md` §6.1 | ~~**Tier-2-only precondition**~~ — superseded by v0.239.0 (any tier, R-475 CLOSED); a Tier-1 route back restores only what the unit holds (R-479); the card keeps the failure sentence after a successful restore (R-480); no automatic rollback, by measurement; a release does not reach the fleet by floor between golden bakes (R-472) **WIDENED 2026-09-21 (the update night) from 3 apps to 21 edges across 19 apps, and NARROWED in one place by the same run.** `audits/DRILL-update-night-2026-09-21.md`. On scratch guest 9202 (controller v0.261.0), against a **private drill catalog** so the live catalog carried no test reference at any point, 21 edges across 19 apps real within-a-major upstream edges were walked through the product's own guarded Update, each app seeded and read back **through its own front door** (R-156) with a negative control on every readback: **14 proven, 3 failed, 4 inconclusive.** **What the PROVEN edges prove, precisely:** the app moved, the four version observables agreed, and the data the app itself was given came back through the app's own interface afterwards. Ten of them printed a verbatim migration line. **What the FAILED edges prove, and they are the more valuable half.** `adventurelog` (a real upstream edge that migrates and then never serves), `tandoor` (an update that SUCCEEDED and was stopped by its own wrong health port), and the PostgreSQL engine major, which refused exactly as predicted. `adventurelog v0.12.1 → v0.13.0` applied **nine database migrations successfully** and then never bound its port; the update held after the full health wait, the hold sentence named the tier, the date and what the copy holds, and the restore the sentence names brought the app back. **That is this row's own promise, exercised on a real upstream edge rather than a staged one.** **AND THE NARROWING, which this row must carry because it is the same mechanism:** the `verifying` phase trusts the `.felhom.yml` probe absolutely, and **two of the 53 templates name a probe the app does not answer** — `tandoor` (port 8080; it listens on 80) and `zipline` (`/api/health`; it answers 404 there, while the compose healthcheck in the same file uses `/api/healthcheck` and is green). For those apps a **successful** update is stopped by its own health wait: tandoor was measured **serving HTTP 200 on the new version at four samples across five minutes**, with docker's own healthcheck green, and was then stopped by `failAndHold` and the household sent to a restore they did not need. **R-618, P1.** No data was lost and the restore works — but "the update is guarded" must not be read as "the guard is right about whether the app came up". **Still true and unchanged:** no automatic rollback (by measurement); the route back is the restore; a multi-major jump ends held honestly. **Not measured on this venue, and named rather than assumed:** every event and every customer mail. Guest 9202 runs `hub.enabled: false` and the notifier returns before it logs (**R-620**), so the whole "who was told" half of `08` was structurally unobservable tonight. **THE NARROWING ABOVE WAS CLOSED THE NEXT DAY, 2026-09-22 — and re-widened the row.** `audits/PROBE-FIX-2026-09-22.md`. All three wrong probes were corrected in the catalog (`app-catalog-felhom.eu@793c4fb`: tandoor `8080→80`, wger `80→8000`, zipline `/api/health→/api/healthcheck`) and **red-proofed live on 9202 through the product in both directions**: at the live pin all three read `Nem egészséges` / `Not healthy` on their own app page while docker reported every container healthy and the front door served a real page; after the real sync all three read `Fut` / `Running` with no redeploy. **tandoor's edge was then re-walked with nothing else changed and ended `done` at +41.1 s**, seed read back, where the identical edge had ended `failed` at +361.9 s with the app stopped — so the tally is now **15 proven, 2 failed, 4 inconclusive**, and all fifteen are on the live catalog. A `--fast` catalog gate (`check-probe-matches-compose.py`) now refuses a probe that does not match the same service's own compose healthcheck, with four red-proofs and ten decoys including the no-PyYAML mode CI actually runs. **WHAT THIS ROW STILL CANNOT CLAIM, and the reason is exactly R-96 rule 3:** the guard is now shown correct for **47 of 53** templates. `paperless-ngx`'s probe has **never run on any box** — no container name matches its stack name, so it is silently skipped and its badge can never go red (**R-630**); and five more cannot be judged statically, one of which (`home-assistant`) is right only because its check type cannot fail (**R-631**). An absent alarm is equally consistent with healthy and with never checked. **And the sweep's ceiling, counted: 28 of the 53 templates have never been deployed by any drill (R-632).** **THAT CEILING WAS REMOVED THE SAME NIGHT, 2026-09-22 — all 28 walked (`audits/DRILL-the-28-2026-09-22.md`), so every template in the catalog has now been attempted at least once.** 26 of 28 deployed, **6 proven**, 5 inconclusive, 14 with no within-a-major edge upstream, 1 failed honestly and 2 undeployable — one of those (`plant-it`) **by design**, refused by the product's lifecycle gate, proven live for the first time. Each app also got the half the update night skipped: a **restore from its own copy, with the seed read back again** — 21 restored, and **2 were correctly REFUSED** with the sentence `07` §6.2 predicts for a class-A app whose local copy holds no file leg. **AND THE NIGHT NARROWED THIS ROW AGAIN, in the place the probe work could not reach.** `paperless-ngx` has no container matching its stack name, so no probe is ever built for it — and `verifying` does not skip: it waits out the full `update.health_timeout` and **HOLDS**, stopping an app whose three containers all read `healthy`. The controller's own words: *`not healthy within 5m0s (last: no probe container) — stopping and HOLDING the app`*, at **+313.0 s**. **R-630, raised to P1.** So "the update is guarded" is now shown correct for 47 of 53 templates, wrong for none, and **actively harmful for the one template that has no probe at all**. **Two further limits on what this row may claim, both about STATE rather than health:** a `remove` sent while a restore is still running reports success and leaves a container restarting with a live public route (**R-633**) — while the product already refuses exactly that clash for `update` and for `restore`, naming the blocking operation; and an app can be **running, healthy and serving while recorded as `deployed: false`**, in which state the product refuses to remove it at all (**R-634**). In both, a person needed a shell to clear what the product could not. | +| **Update is GUARDED: it refuses without a restorable backup, backs up first when the copy is stale, and HOLDS an app that does not come up — on ANY backup tier, and the release itself arrives by the managed floor** | controller **v0.237.0 + v0.238.0 + v0.238.1 + v0.239.0**, hub **v0.112.0** | **PROVEN-LIVE (2026-09-13, and again the same afternoon for any tier + floor delivery)** — **afternoon (`audits/rulings-r472-r475-2026-09-13/`):** an undeclared floor above the golden refused with nothing stored (02); a declared floor 0.239.0 / MinAgent 0.129.0 served `from declared` and both demo boxes self-updated in 14 s and 15 s (03); nothing on any tier → backed up first, Tier 1 chosen, done (04); gokapi updated on its Tier-1 unit alone (05); a never-healthy update held naming „saját meghajtó" (07); restored from „helyi", hold cleared (08). **Morning:** scenarios A (real upgrade, success only after health), B (stale copy → backup first), E (pull failure → pin back, app untouched), F (never healthy → held, hold text on API and page), H (start/restart/update and the boot sweep all refuse the held app) and **the restore walk** (Mentések unit restore → back on the old version, hold cleared), on demo-hp with a throwaway app | **`audits/slice4-2026-09-13/`** (live/, redproofs/, gates/); design `architecture/09-update-architecture.md` §6.1 | ~~**Tier-2-only precondition**~~ — superseded by v0.239.0 (any tier, R-475 CLOSED); a Tier-1 route back restores only what the unit holds (R-479); the card keeps the failure sentence after a successful restore (R-480); no automatic rollback, by measurement; a release does not reach the fleet by floor between golden bakes (R-472) **WIDENED 2026-09-21 (the update night) from 3 apps to 21 edges across 19 apps, and NARROWED in one place by the same run.** `audits/DRILL-update-night-2026-09-21.md`. On scratch guest 9202 (controller v0.261.0), against a **private drill catalog** so the live catalog carried no test reference at any point, 21 edges across 19 apps real within-a-major upstream edges were walked through the product's own guarded Update, each app seeded and read back **through its own front door** (R-156) with a negative control on every readback: **14 proven, 3 failed, 4 inconclusive.** **What the PROVEN edges prove, precisely:** the app moved, the four version observables agreed, and the data the app itself was given came back through the app's own interface afterwards. Ten of them printed a verbatim migration line. **What the FAILED edges prove, and they are the more valuable half.** `adventurelog` (a real upstream edge that migrates and then never serves), `tandoor` (an update that SUCCEEDED and was stopped by its own wrong health port), and the PostgreSQL engine major, which refused exactly as predicted. `adventurelog v0.12.1 → v0.13.0` applied **nine database migrations successfully** and then never bound its port; the update held after the full health wait, the hold sentence named the tier, the date and what the copy holds, and the restore the sentence names brought the app back. **That is this row's own promise, exercised on a real upstream edge rather than a staged one.** **AND THE NARROWING, which this row must carry because it is the same mechanism:** the `verifying` phase trusts the `.felhom.yml` probe absolutely, and **two of the 53 templates name a probe the app does not answer** — `tandoor` (port 8080; it listens on 80) and `zipline` (`/api/health`; it answers 404 there, while the compose healthcheck in the same file uses `/api/healthcheck` and is green). For those apps a **successful** update is stopped by its own health wait: tandoor was measured **serving HTTP 200 on the new version at four samples across five minutes**, with docker's own healthcheck green, and was then stopped by `failAndHold` and the household sent to a restore they did not need. **R-618, P1.** No data was lost and the restore works — but "the update is guarded" must not be read as "the guard is right about whether the app came up". **Still true and unchanged:** no automatic rollback (by measurement); the route back is the restore; a multi-major jump ends held honestly. **Not measured on this venue, and named rather than assumed:** every event and every customer mail. Guest 9202 runs `hub.enabled: false` and the notifier returns before it logs (**R-620**), so the whole "who was told" half of `08` was structurally unobservable tonight. **THE NARROWING ABOVE WAS CLOSED THE NEXT DAY, 2026-09-22 — and re-widened the row.** `audits/PROBE-FIX-2026-09-22.md`. All three wrong probes were corrected in the catalog (`app-catalog-felhom.eu@793c4fb`: tandoor `8080→80`, wger `80→8000`, zipline `/api/health→/api/healthcheck`) and **red-proofed live on 9202 through the product in both directions**: at the live pin all three read `Nem egészséges` / `Not healthy` on their own app page while docker reported every container healthy and the front door served a real page; after the real sync all three read `Fut` / `Running` with no redeploy. **tandoor's edge was then re-walked with nothing else changed and ended `done` at +41.1 s**, seed read back, where the identical edge had ended `failed` at +361.9 s with the app stopped — so the tally is now **15 proven, 2 failed, 4 inconclusive**, and all fifteen are on the live catalog. A `--fast` catalog gate (`check-probe-matches-compose.py`) now refuses a probe that does not match the same service's own compose healthcheck, with four red-proofs and ten decoys including the no-PyYAML mode CI actually runs. **WHAT THIS ROW STILL CANNOT CLAIM, and the reason is exactly R-96 rule 3:** the guard is now shown correct for **47 of 53** templates. `paperless-ngx`'s probe has **never run on any box** — no container name matches its stack name, so it is silently skipped and its badge can never go red (**R-630**); and five more cannot be judged statically, one of which (`home-assistant`) is right only because its check type cannot fail (**R-631**). An absent alarm is equally consistent with healthy and with never checked. **And the sweep's ceiling, counted: 28 of the 53 templates have never been deployed by any drill (R-632).** **THAT CEILING WAS REMOVED THE SAME NIGHT, 2026-09-22 — all 28 walked (`audits/DRILL-the-28-2026-09-22.md`), so every template in the catalog has now been attempted at least once.** 26 of 28 deployed, **6 proven**, 5 inconclusive, 14 with no within-a-major edge upstream, 1 failed honestly and 2 undeployable — one of those (`plant-it`) **by design**, refused by the product's lifecycle gate, proven live for the first time. Each app also got the half the update night skipped: a **restore from its own copy, with the seed read back again** — 21 restored, and **2 were correctly REFUSED** with the sentence `07` §6.2 predicts for a class-A app whose local copy holds no file leg. **AND THE NIGHT NARROWED THIS ROW AGAIN, in the place the probe work could not reach.** `paperless-ngx` has no container matching its stack name, so no probe is ever built for it — and `verifying` does not skip: it waits out the full `update.health_timeout` and **HOLDS**, stopping an app whose three containers all read `healthy`. The controller's own words: *`not healthy within 5m0s (last: no probe container) — stopping and HOLDING the app`*, at **+313.0 s**. **R-630, raised to P1.** So "the update is guarded" is now shown correct for 47 of 53 templates, wrong for none, and **actively harmful for the one template that has no probe at all**. **Two further limits on what this row may claim, both about STATE rather than health:** a `remove` sent while a restore is still running reports success and leaves a container restarting with a live public route (**R-633**) — while the product already refuses exactly that clash for `update` and for `restore`, naming the blocking operation; and an app can be **running, healthy and serving while recorded as `deployed: false`**, in which state the product refuses to remove it at all (**R-634**). In both, a person needed a shell to clear what the product could not. **ALL THREE ARE FIXED IN CONTROLLER v0.262.0 (2026-09-22), and the first is PROVEN LIVE.** *The stopped app:* `verifying` no longer loops on a probe that resolves to nothing — it settles on container state, the way an app declaring no check is judged, and says which it did. **Measured on paperless-ngx: the identical Update that ended `failed` at +313.0 s with the app stopped now ends `done` at +53.4 s**, with no `no probe container` warning in the log because the explicit `healthcheck.container` resolved the target. *The ghost:* `RemoveStack` consults the backup side's `Busy` guard — which the product already applied to `update` and to `restore` — and then WATCHES the compose project for 25 s after `down`, removing anything that carries its label and answering `verified: true/false`, because `down` returning 0 is a request rather than a result. *The unremovable app:* the refusal now asks whether anything EXISTS (containers, a compose file, an `app.yaml`) instead of reading a flag. **WHAT THIS ROW STILL MAY NOT CLAIM:** R-634's MECHANISM — why `deployed` goes false while containers run — **is not diagnosed**; only the consequence is fixed. And a probe can be right about the port and still wrong about what a 200 means: `romm` answered 200 from nginx for six hours while its workers were OOM-killed behind it (**R-635**). **"The update is guarded" has never meant "the new version runs".** | | **What `restart` and `update` do to a deployed app whose compose file the catalog already moved** | controller **v0.235.0** | **CHANGED 2026-09-06 — they NO LONGER upgrade it.** The row below records what shipped; this text records what it replaced, because every box under v0.235.0 still behaves the old way. **Up to v0.234.0: PROVEN-LIVE (2026-09-01) — they UPGRADE it.** Every lifecycle action ends in `docker compose up -d`, which makes the container match the file and PULLS the image itself when it is missing (measured: 18.3 s with a pull, 0.5 s without; negative control with an unchanged file did not even recreate the container). This is DELIBERATE on the restart path — `Manager.RestartStack` says so in a comment — but the syncer moves the file under a deployed app on a 15-minute cycle with no deployed check (R-438), and NOTHING tells the customer. | `audits/SPIKE-app-update-2026-09-01.md` §2, §3 | **No safety copy is taken by any of them** — `writeSafetyDump` is DATABASE-ONLY and is not on the update path at all. R-438, R-440, R-443. | | **Whether the box UPGRADES an app by itself, with nobody pressing anything** | controller | **PROVEN-LIVE (2026-09-01) — YES, but only when an app fails to come back.** A plain power cut does NOT upgrade: Docker's `restart: unless-stopped` restores the old containers and the reconciler logs `no boot-orphaned apps (nothing to start)`. When an app does NOT return, `Reconciler.Run` (`bootrecon.go:269`) calls `StartStack` -> `compose up -d` and the app comes back on the NEW version, unattended (measured). **13 non-API call sites across 9 files reach `up -d` this way** — not the five previously believed. | `audits/SPIKE-app-update-2026-09-01.md` §2, §8 | The drive-return gate (`intermediary.go:222`) and `AppStopGuard.Recover` (`appstop_marker.go:283`) call the same function; located by reading, **not exercised live** — stated as such. | | **Whether an app UPGRADE can be undone** | controller + catalog | **PROVEN-LIVE (2026-09-01) — NO, and "rollback" is the wrong word for it.** Once a migration has RUN, putting the old image tag back yields a container that refuses to start: Nextcloud — *"the version of the data (32.0.9.2) is higher than the docker image version (31.0.14.1) and downgrading is not supported"*. A 3-major jump is refused outright (*"only possible to upgrade one major version at a time"*) and IS recoverable, precisely because nothing migrated. Positive control: the data is not destroyed — returning to 32.0.9 restored both seeded markers byte-identical. | `audits/SPIKE-app-update-2026-09-01.md` §7 | The only route back is restoring DATA from a copy taken BEFORE the update — which no update path takes. And a restore's image-level rollback is itself overwritten by the syncer within 15 minutes (R-441). R-40 is confirmed live by the same measurement. | diff --git a/documentation/architecture/09-update-architecture.md b/documentation/architecture/09-update-architecture.md index 5b0b19be..25b5f3d9 100644 --- a/documentation/architecture/09-update-architecture.md +++ b/documentation/architecture/09-update-architecture.md @@ -653,7 +653,7 @@ closed by construction: nothing reports an update complete on the compose exit c | 4 | `pinning` | the previous definition is copied aside and journaled, then the pin advances | pin put back | | 5 | `pulling` | `compose pull` | **pin and definition PUT BACK** — nothing ran (Scenario E) | | 6 | `starting` | `compose up -d --remove-orphans` | stop + HOLD | -| 7 | `verifying` | the `.felhom.yml` health check through the existing probe, or 60 s of every container running and none restarting; bounded by `update.health_timeout` | **stop + HOLD; the pin STAYS** — the migration may have run (Scenario F) | +| 7 | `verifying` | the `.felhom.yml` health check through the existing probe **when it resolves to a container**, else 60 s of every container running and none restarting; bounded by `update.health_timeout` | **stop + HOLD; the pin STAYS** — the migration may have run (Scenario F) | | 8 | `done` | installed images recorded, journal cleared | — | **The two knobs** (`controller.yaml`, operator-owned): `update.backup_max_age` (default `24h`) and @@ -879,6 +879,18 @@ headlessly (R-460). | E — the automatic night | **COMPLETE.** The unattended HOLD was produced at last (312.9 s), with no retry across two further passes. It needed the image store of §6.5 | | F — the remaining apps | **DONE 2026-09-22 (`audits/DRILL-the-28-2026-09-22.md`): the 28 apps no drill had ever touched were walked in ONE night, so the catalog is now **53 of 53 attempted**, not 25.** 26 of the 28 deployed; 6 proven; 5 inconclusive; 14 had no within-a-major edge upstream that night; 1 failed honestly (`outline 1.9.1 -> 1.10.1`, HELD with the right sentence); 2 could not be deployed, one of them (`plant-it`) **by design** — it is `lifecycle: abandoned` and the product's lifecycle gate refused it, proven live for the first time. **The cost is now known and it is not machine time:** three concurrent walks did 28 apps in about four hours, and the binding constraints were FIXTURES (only 6 of 28 had a non-browser route that both seeded and read back) and the fact that `POST /api/backup/run` is BOX-WIDE, so concurrent walks serialise on it. **This leg also added the night's biggest finding**, which no count would have produced: R-630 | +**The words "when it resolves to a container" are v0.262.0's, and they are the whole of R-630.** +Before it, the probe branch had no exit: `findProbeContainer` returning `""` set a message and +looped, while the settle path that judges an app declaring NO check sat in the outer `else`, +unreachable. So a stack whose probe resolved to nothing could only ever time out — and +`failAndHold` then stopped an app whose containers were all healthy. **A stack with no probe is not +healthy and not failing; it is settled on container state (§3), and never a reason to stop a running +app.** Proven live on paperless-ngx 2026-09-22: the identical Update that ended **`failed` at ++313.0 s with the app stopped** now ends **`done` at +53.4 s** +(`audits/v0262-live-2026-09-22/live262.json`). Which container is probed is now decidable too — +exact stack name, then `healthcheck.container`, then a UNIQUE prefix, else nothing with the +candidates logged; the old rule took the FIRST prefix match. + **What the night ADDED to this table, which none of the legs anticipated:** the `verifying` phase trusts the `.felhom.yml` probe absolutely, and three of 53 templates name a probe the app does not answer — so a SUCCESSFUL update of those apps ends by STOPPING a working app (**R-618**, P1). That is @@ -1121,6 +1133,16 @@ Version strings stay in the logs, the API and the hub. **So limitation 8 now has two shapes, not one:** a probe that names the wrong target (R-618, fixed) and **no probe at all** (**R-630, raised to P1**), and the static gate can see the first but not the second, because there is nothing to compare. + **FIXED 2026-09-22 in controller v0.262.0**, and the fix is in the phase table above: the + no-probe case now settles on container state instead of looping, and the probe TARGET is + decidable (exact name → `healthcheck.container` → a UNIQUE prefix → nothing, candidates logged). + `paperless-ngx` and `immich` carry the explicit field, and the catalog gate REFUSES a probe that + resolves to nothing rather than warning about it. **So limitation 8's two shapes are both closed + in the product**; what remains is that six templates still cannot be judged STATICALLY (R-631, + all five read live and correct) and that a probe can still be right about the port and wrong + about what a 200 means — `romm` answered 200 from nginx while its workers were being OOM-killed + for six hours (**R-635**), which is a third shape again and the reason "the update is guarded" + must never be read as "the new version runs". **Two more things the same night measured, both about state rather than health:** a `remove` sent while a restore is still running reports success and leaves a container restarting with a live public route (**R-633**) — and the product already has exactly that guard for `update` and for diff --git a/documentation/audits/v0262-live-2026-09-22/live262.json b/documentation/audits/v0262-live-2026-09-22/live262.json new file mode 100644 index 00000000..5b9bb98e --- /dev/null +++ b/documentation/audits/v0262-live-2026-09-22/live262.json @@ -0,0 +1,133 @@ +{ + "A": { + "scenario": "A (R-630)", + "app": "paperless-ngx", + "before": { + "state": "running", + "front_door": "302", + "containers": [ + "paperless-webserver|Up About a minute (healthy)", + "paperless-redis|Up About a minute (healthy)", + "paperless-postgres|Up About a minute (healthy)" + ] + }, + "phases": { + "accepted": true, + "http": "202", + "phases": [ + { + "t": 0.0, + "phase": "backing-up", + "label": "Biztonsági mentés készül a frissítés előtt…", + "updating": true, + "error": null, + "hold": null + }, + { + "t": 21.5, + "phase": "pulling", + "label": "Új verzió letöltése…", + "updating": true, + "error": null, + "hold": null + }, + { + "t": 22.6, + "phase": "starting", + "label": "Indítás az új verzióval…", + "updating": true, + "error": null, + "hold": null + }, + { + "t": 23.6, + "phase": "verifying", + "label": "Működés ellenőrzése…", + "updating": true, + "error": null, + "hold": null + }, + { + "t": 53.4, + "phase": "done", + "label": "Frissítve", + "updating": false, + "error": null, + "hold": null + } + ], + "duration_s": 53.4, + "final_phase": "done", + "update_error": null, + "hold_reason": null, + "state": "running" + }, + "wall_s": 53.7, + "after": { + "state": "running", + "front_door": "302" + }, + "controller_says": "2026/09/22 19:33:55 update.go:996: [INFO] [stacks] update paperless-ngx: phase checking\n2026/09/22 19:33:55 update.go:996: [INFO] [stacks] update paperless-ngx: phase backing-up\n2026/09/22 19:34:16 update.go:996: [INFO] [stacks] update paperless-ngx: phase safety-dump\n2026/09/22 19:34:16 update.go:996: [INFO] [stacks] update paperless-ngx: phase pinning\n2026/09/22 19:34:16 update.go:996: [INFO] [stacks] update paperless-ngx: phase pulling\n2026/09/22 19:34:17 update.go:996: [INFO] [stacks] update paperless-ngx: phase starting\n2026/09/22 19:34:18 update.go:996: [INFO] [stacks] update paperless-ngx: phase verifying\n2026/09/22 19:34:48 healthprobe.go:181: [DEBUG] Health probe paperless-ngx: HTTP GET :8000/ → 302 (4ms)" + }, + "F": { + "scenario": "F (R-614)", + "app": "paperless-ngx", + "phase_before_remove": "done", + "after_remove_deployed": false, + "phase_after_redeploy": null, + "verdict": "clean" + }, + "B": { + "scenario": "B (R-633/R-626)", + "app": "privatebin", + "snapshots": 1, + "restore_started": "{'ok': True, 'snapshot_id': 'helyi', 'snapshots': [{'time': '2026-09-22T19:41:55Z', 'short_id': 'helyi', 'tier': 1, 'drive_label': 'Belső SSD (rendszer)'}], 'http': 'HTTP/2 302', 'location': ['locatio", + "remove_during_restore": { + "http": "409", + "answer": { + "ok": false, + "error": "stack \"privatebin\" is still running — stop it first before removing" + } + }, + "refused_as_expected": true, + "remove_after_restore": { + "http": "200", + "verified": true, + "reappeared": null + }, + "containers_60s_later": "" + }, + "C": { + "scenario": "C (R-634)", + "app": "sparkyfitness", + "setup": "total 28\ndrwxr-xr-x 2 root root 4096 Sep 22 19:43 .\ndrwxr-xr-x 57 root root 4096 Sep 13 20:22 ..\n-rw-r--r-- 1 root root 4737 Sep 20 14:32 .felhom.yml\n-rw-r--r-- 1 root root 41 Sep 22 19:43 app.yaml\n-rw-r--r-- 1 root root 4936 Sep 22 19:43 docker-compose.yml", + "before": { + "deployed": false, + "state": "not_deployed" + }, + "remove": { + "http": "200", + "answer": "{'ok': True, 'data': {'removed': 'sparkyfitness', 'volumes_removed': [], 'hdd_paths_removed': [], 'hdd_paths_preserved': [], 'backup_paths_removed': ['/mnt/sys_drive/felhom-data/backups/primary/sparkyfitness (48K)'], 'verified': True}, 'message': 'Stack sparkyfitness removed'}" + }, + "accepted": true, + "leftovers": "NONE" + }, + "B2": { + "scenario": "B2 (R-633) — the BUSY guard, reached deliberately", + "app": "privatebin", + "state_before": "stopped", + "backup_started": { + "http": "200", + "answer": "{'ok': True, 'message': 'Mentés elindítva'}" + }, + "remove_during_backup": { + "http": "409", + "answer": { + "ok": false, + "error": "Az alkalmazáson mentés vagy visszaállítás fut. Várd meg, amíg befejeződik." + } + }, + "refused_by_busy_guard": true, + "controller_says": "2026/09/22 19:49:46 delete.go:557: [ERROR] [stacks] RemoveStack privatebin REFUSED (busy): a backup or restore is running (single-flight held)\n2026/09/22 19:51:32 delete.go:557: [ERROR] [stacks] RemoveStack privatebin REFUSED (busy): a backup or restore is running (single-flight held)" + } +} \ No newline at end of file diff --git a/documentation/audits/v0262-live-2026-09-22/live262.py b/documentation/audits/v0262-live-2026-09-22/live262.py new file mode 100644 index 00000000..ac7f96e2 --- /dev/null +++ b/documentation/audits/v0262-live-2026-09-22/live262.py @@ -0,0 +1,66 @@ +#!/usr/bin/env python3 +"""live262.py — the v0.262.0 scenarios, on scratch guest 9202, through the product's endpoints. + +A: R-630 — paperless-ngx has a probe that resolves to no container WITHOUT the explicit field, and + an explicit one WITH it. Both are exercised: the catalog now carries the field, so the Update + must reach `done`; the drill catalog lets the field be removed to show the OTHER half. +B: R-633 — a remove sent during a restore is refused; a remove after it is VERIFIED clean. +C: R-634 — a half-state (app.yaml, no deployed flag) is removable. +F: R-614 — a redeploy reads no stale phase. +""" +import json, os, sys, time +HERE = os.path.dirname(os.path.abspath(__file__)) +sys.path.insert(0, os.path.join(os.path.dirname(HERE), "update-night-2026-09-21")) +import walk as w # noqa: E402 + +OUT = {} + + +def save(tag, rec): + OUT[tag] = rec + json.dump(OUT, open(os.path.join(HERE, "live262.json"), "w"), ensure_ascii=False, indent=2) + print(json.dumps({tag: rec}, ensure_ascii=False, indent=2)[:1400], flush=True) + + +def scenario_A(): + """paperless-ngx: the update must now END `done`, where v0.261.0 held it at +313 s.""" + app, sub = "paperless-ngx", "paperless" + rec = {"scenario": "A (R-630)", "app": app} + w.deploy(app, sub) + w.wait_app(sub, "/", tries=60) + st = w.stack(app) + rec["before"] = {"state": st.get("state"), + "front_door": w.app_curl(sub, "/")[1], + "containers": w.guest("docker ps --format '{{.Names}}|{{.Status}}' | grep -i paperless").strip().split("\n")} + # the probe now resolves — the controller's own log says which container + t0 = time.time() + rec["phases"] = w.press_update(app) + rec["wall_s"] = round(time.time() - t0, 1) + rec["after"] = {"state": w.stack(app).get("state"), "front_door": w.app_curl(sub, "/")[1]} + rec["controller_says"] = w.guest( + "docker logs --since 15m felhom-controller 2>&1 | grep -iE 'paperless' | " + "grep -iE 'no probe container|settling|phase |health probe|FAILED after' | tail -8").strip() + save("A", rec) + return app + + +def scenario_F(app): + """R-614: remove clears the phase; a redeploy reads none.""" + rec = {"scenario": "F (R-614)", "app": app} + rec["phase_before_remove"] = w.stack(app).get("update_phase") + w.remove(app) + time.sleep(5) + rec["after_remove_deployed"] = w.stack(app).get("deployed") + w.deploy(app, "paperless") + time.sleep(5) + st = w.stack(app) + rec["phase_after_redeploy"] = st.get("update_phase") + rec["verdict"] = "clean" if not st.get("update_phase") else "STALE PHASE SURVIVED" + save("F", rec) + + +if __name__ == "__main__": + w.login() + which = sys.argv[1] if len(sys.argv) > 1 else "A" + if which == "A": + scenario_F(scenario_A()) diff --git a/documentation/audits/v0262-live-2026-09-22/liveA.log b/documentation/audits/v0262-live-2026-09-22/liveA.log new file mode 100644 index 00000000..b372df07 --- /dev/null +++ b/documentation/audits/v0262-live-2026-09-22/liveA.log @@ -0,0 +1,88 @@ +21:32:42 [1] made the drive paths this app requires: ['/mnt/felhom-drives/scratch_hdd/userdata/paperless-ngx'] +21:32:42 [1] required fields filled beyond DOMAIN/SUBDOMAIN: ['PAPERLESS_ADMIN_PASSWORD', 'HDD_PATH'] +21:32:42 [1] deploy -> 202 {'ok': True, 'message': 'Telepítés elindítva – az állapot a kártyán követhető'} +21:33:52 [1] deployed, controller state=running, pinned={'paperless-postgres': 'postgres:16-alpine', 'paperless-redis': 'redis:7-alpine', 'paperless-webserver': 'ghcr.io/paperless-ngx/paperless-ngx:2.20.15'} +21:33:55 [6] Update -> 202 {'ok': True, 'data': {'accepted': True, 'completed': False}, 'message': 'Frissítés elindult – az állapot a kártyán követhető'} +21:33:55 + 0.0s phase=backing-up label=Biztonsági mentés készül a frissítés előtt… err=None hold=None +21:34:16 + 21.5s phase=pulling label=Új verzió letöltése… err=None hold=None +21:34:18 + 22.6s phase=starting label=Indítás az új verzióval… err=None hold=None +21:34:19 + 23.6s phase=verifying label=Működés ellenőrzése… err=None hold=None +21:34:48 + 53.4s phase=done label=Frissítve err=None hold=None +{ + "A": { + "scenario": "A (R-630)", + "app": "paperless-ngx", + "before": { + "state": "running", + "front_door": "302", + "containers": [ + "paperless-webserver|Up About a minute (healthy)", + "paperless-redis|Up About a minute (healthy)", + "paperless-postgres|Up About a minute (healthy)" + ] + }, + "phases": { + "accepted": true, + "http": "202", + "phases": [ + { + "t": 0.0, + "phase": "backing-up", + "label": "Biztonsági mentés készül a frissítés előtt…", + "updating": true, + "error": null, + "hold": null + }, + { + "t": 21.5, + "phase": "pulling", + "label": "Új verzió letöltése…", + "updating": true, + "error": null, + "hold": null + }, + { + "t": 22.6, + "phase": "starting", + "label": "Indítás az új verzióval…", + "updating": true, + "error": null, + "hold": null + }, + { + "t": 23.6, + "phase": "verifying", + "label": "Működés ellenőrzése…", + "updating": true, + "error": null, + "hold": null + }, + { + "t": 53.4, + "phase": "done", + "label": "Frissítve", + "updating": false, + "error": null, + "hold": null + } + +21:34:57 [X] stop -> 200 {'ok': True, 'message': 'Stack paperless-ngx stop completed'} +21:35:03 [X] remove (with drive data) -> 409 {'ok': False, 'error': 'A(z) /mnt/felhom-drives/scratch_hdd/userdata/paperless-ngx tárhely jelenleg nem elérhető — az alkalmazás nem távolítható el, amíg a megh +21:35:03 [X] refused because the drive path cannot be resolved (R-442, fail-closed and right) — removing the app and KEEPING the drive data instead +21:35:29 [X] remove (keeping drive data) -> 200 {'ok': True, 'data': {'removed': 'paperless-ngx', 'volumes_removed': ['paperless-ngx_paperless_data', 'paperless-ngx_paperless_postgres_data', 'paperless-ngx_pa +21:35:37 [X] after remove: deployed=False leftovers='/opt/docker/stacks/paperless-ngx' +21:35:44 [1] made the drive paths this app requires: ['/mnt/felhom-drives/scratch_hdd/userdata/paperless-ngx'] +21:35:44 [1] required fields filled beyond DOMAIN/SUBDOMAIN: ['PAPERLESS_ADMIN_PASSWORD', 'HDD_PATH'] +21:35:44 [1] deploy -> 202 {'ok': True, 'message': 'Telepítés elindítva – az állapot a kártyán követhető'} +21:36:49 [1] deployed, controller state=running, pinned={'paperless-postgres': 'postgres:16-alpine', 'paperless-redis': 'redis:7-alpine', 'paperless-webserver': 'ghcr.io/paperless-ngx/paperless-ngx:2.20.15'} +{ + "F": { + "scenario": "F (R-614)", + "app": "paperless-ngx", + "phase_before_remove": "done", + "after_remove_deployed": false, + "phase_after_redeploy": null, + "verdict": "clean" + } +} +RC=0 diff --git a/documentation/audits/v0262-live-2026-09-22/liveB2.log b/documentation/audits/v0262-live-2026-09-22/liveB2.log new file mode 100644 index 00000000..4e8f2ea8 --- /dev/null +++ b/documentation/audits/v0262-live-2026-09-22/liveB2.log @@ -0,0 +1,21 @@ +21:44:56 [1] deploy -> 202 {'ok': True, 'message': 'Telepítés elindítva – az állapot a kártyán követhető'} +21:45:11 [1] deployed, controller state=running, pinned={'privatebin': 'privatebin/pdo:2.0.6'} +{ + "scenario": "B2 (R-633) — the BUSY guard, reached deliberately", + "app": "privatebin", + "state_before": "stopped", + "backup_started": { + "http": "200", + "answer": "{'ok': True, 'message': 'Mentés elindítva'}" + }, + "remove_during_backup": { + "http": "500", + "answer": { + "ok": false, + "error": "Az alkalmazáson mentés vagy visszaállítás fut. Várd meg, amíg befejeződik." + } + }, + "refused_by_busy_guard": false, + "controller_says": "2026/09/22 19:45:20 delete.go:542: [ERROR] [stacks] RemoveStack privatebin REFUSED (busy): a backup or restore is running (single-flight held)" +} +RC=0 diff --git a/documentation/audits/v0262-live-2026-09-22/liveB2.py b/documentation/audits/v0262-live-2026-09-22/liveB2.py new file mode 100644 index 00000000..909e8c65 --- /dev/null +++ b/documentation/audits/v0262-live-2026-09-22/liveB2.py @@ -0,0 +1,39 @@ +#!/usr/bin/env python3 +"""B2 — prove the BUSY guard specifically (R-633). + +B's first run refused the remove with `409 still running — stop it first`, which is the PRE-EXISTING +running check, not v0.262.0's new guard: the restore had already finished. To reach the new guard the +app must be STOPPED (so the running check passes) while the backup side is busy. +""" +import json, os, sys, time +HERE = os.path.dirname(os.path.abspath(__file__)) +sys.path.insert(0, os.path.join(os.path.dirname(HERE), "update-night-2026-09-21")) +import walk as w # noqa: E402 + +OUT = json.load(open(os.path.join(HERE, "live262.json"))) +app, sub = "privatebin", "paste" +rec = {"scenario": "B2 (R-633) — the BUSY guard, reached deliberately", "app": app} + +w.login() +w.deploy(app, sub) +w.wait_app(sub, "/", tries=30) +# stop it, so the pre-existing "still running" check cannot answer first +w.ctl("POST", f"/api/stacks/{app}/stop") +for _ in range(24): + time.sleep(5) + if w.stack(app).get("state") != "running": + break +rec["state_before"] = w.stack(app).get("state") + +# box-wide backup: `POST /api/backup/run` holds the single-flight the guard consults +code, d = w.ctl("POST", "/api/backup/run") +rec["backup_started"] = {"http": code, "answer": str(d)[:120]} +time.sleep(3) +code, d = w.ctl("POST", f"/api/stacks/{app}/remove", {"remove_hdd_data": False, "remove_backups": False}) +rec["remove_during_backup"] = {"http": code, "answer": d} +rec["refused_by_busy_guard"] = (code == "409" and "ment" in json.dumps(d, ensure_ascii=False).lower()) +rec["controller_says"] = w.guest( + f"docker logs --since 5m felhom-controller 2>&1 | grep -iE 'RemoveStack {app}.*(REFUSED|busy)' | tail -3").strip() +OUT["B2"] = rec +json.dump(OUT, open(os.path.join(HERE, "live262.json"), "w"), ensure_ascii=False, indent=2) +print(json.dumps(rec, ensure_ascii=False, indent=2)) diff --git a/documentation/audits/v0262-live-2026-09-22/liveB2b.log b/documentation/audits/v0262-live-2026-09-22/liveB2b.log new file mode 100644 index 00000000..bf432ced --- /dev/null +++ b/documentation/audits/v0262-live-2026-09-22/liveB2b.log @@ -0,0 +1,19 @@ +21:51:24 [1] privatebin already deployed — reusing +{ + "scenario": "B2 (R-633) — the BUSY guard, reached deliberately", + "app": "privatebin", + "state_before": "stopped", + "backup_started": { + "http": "200", + "answer": "{'ok': True, 'message': 'Mentés elindítva'}" + }, + "remove_during_backup": { + "http": "409", + "answer": { + "ok": false, + "error": "Az alkalmazáson mentés vagy visszaállítás fut. Várd meg, amíg befejeződik." + } + }, + "refused_by_busy_guard": true, + "controller_says": "2026/09/22 19:49:46 delete.go:557: [ERROR] [stacks] RemoveStack privatebin REFUSED (busy): a backup or restore is running (single-flight held)\n2026/09/22 19:51:32 delete.go:557: [ERROR] [stacks] RemoveStack privatebin REFUSED (busy): a backup or restore is running (single-flight held)" +} diff --git a/documentation/audits/v0262-live-2026-09-22/liveBC.log b/documentation/audits/v0262-live-2026-09-22/liveBC.log new file mode 100644 index 00000000..fe05eea7 --- /dev/null +++ b/documentation/audits/v0262-live-2026-09-22/liveBC.log @@ -0,0 +1,49 @@ +21:41:22 [1] deploy -> 202 {'ok': True, 'message': 'Telepítés elindítva – az állapot a kártyán követhető'} +21:41:32 [1] deployed, controller state=running, pinned={'privatebin': 'privatebin/pdo:2.0.6'} +21:41:32 [4] „Mentés most" -> 200 {'ok': True, 'message': 'Mentés elindítva'} +21:41:57 [4] backup idle; last=None +21:41:57 [R] restoring privatebin from snapshot 'helyi' (of 1 offered) +21:41:57 [R] POST /backup/restore -> HTTP/2 302 ['location: /backups/restore?flash=flash.restore.started'] +21:41:57 + 0.0s restore (True, None, None) +21:42:07 + 10.1s restore (False, None, None) +21:42:07 [R] after restore: state=running hold=None phase=None + +== B +{ + "scenario": "B (R-633/R-626)", + "app": "privatebin", + "snapshots": 1, + "restore_started": "{'ok': True, 'snapshot_id': 'helyi', 'snapshots': [{'time': '2026-09-22T19:41:55Z', 'short_id': 'helyi', 'tier': 1, 'drive_label': 'Belső SSD (rendszer)'}], 'http': 'HTTP/2 302', 'location': ['locatio", + "remove_during_restore": { + "http": "409", + "answer": { + "ok": false, + "error": "stack \"privatebin\" is still running — stop it first before removing" + } + }, + "refused_as_expected": true, + "remove_after_restore": { + "http": "200", + "verified": true, + "reappeared": null + }, + "containers_60s_later": "" +} + +== C +{ + "scenario": "C (R-634)", + "app": "sparkyfitness", + "setup": "total 28\ndrwxr-xr-x 2 root root 4096 Sep 22 19:43 .\ndrwxr-xr-x 57 root root 4096 Sep 13 20:22 ..\n-rw-r--r-- 1 root root 4737 Sep 20 14:32 .felhom.yml\n-rw-r--r-- 1 root root 41 Sep 22 19:43 app.yaml\n-rw-r--r-- 1 root root 4936 Sep 22 19:43 docker-compose.yml", + "before": { + "deployed": false, + "state": "not_deployed" + }, + "remove": { + "http": "200", + "answer": "{'ok': True, 'data': {'removed': 'sparkyfitness', 'volumes_removed': [], 'hdd_paths_removed': [], 'hdd_paths_preserved': [], 'backup_paths_removed': ['/mnt/sys_drive/felhom-data/backups/primary/sparkyfitness (48K)'], 'verified': True}, 'message': 'Stack sparkyfitness removed'}" + }, + "accepted": true, + "leftovers": "NONE" +} +RC=0 diff --git a/documentation/audits/v0262-live-2026-09-22/liveBC.py b/documentation/audits/v0262-live-2026-09-22/liveBC.py new file mode 100644 index 00000000..bcb20eec --- /dev/null +++ b/documentation/audits/v0262-live-2026-09-22/liveBC.py @@ -0,0 +1,85 @@ +#!/usr/bin/env python3 +"""B (R-633) and C (R-634) live on 9202, against v0.262.0.""" +import json, os, sys, time +HERE = os.path.dirname(os.path.abspath(__file__)) +sys.path.insert(0, os.path.join(os.path.dirname(HERE), "update-night-2026-09-21")) +import walk as w # noqa: E402 + +OUT = json.load(open(os.path.join(HERE, "live262.json"))) if os.path.exists( + os.path.join(HERE, "live262.json")) else {} + + +def save(tag, rec): + OUT[tag] = rec + json.dump(OUT, open(os.path.join(HERE, "live262.json"), "w"), ensure_ascii=False, indent=2) + print("\n== " + tag + "\n" + json.dumps(rec, ensure_ascii=False, indent=2)[:1500], flush=True) + + +def B(): + """A remove sent DURING a restore must be refused; a remove after it must be VERIFIED.""" + # gokapi is crash-looping from this morning's R-633 artefact — its config volume was removed, so + # its binary can never start and the stack never settles. A subject that cannot reach a steady + # state proves nothing about a guard that fires between steady states. privatebin is light, + # reliable and was walked all night. + app, sub = "privatebin", "paste" + rec = {"scenario": "B (R-633/R-626)", "app": app} + w.deploy(app, sub) + w.wait_app(sub, "/", tries=40) + w.backup_now(app) + snaps = w.snapshots(app) + rec["snapshots"] = len(snaps) + # start a restore and, while it runs, ask to remove + r = w.restore(app) + rec["restore_started"] = str(r)[:200] + time.sleep(2) + code, d = w.ctl("POST", f"/api/stacks/{app}/remove", {"remove_hdd_data": False, "remove_backups": False}) + rec["remove_during_restore"] = {"http": code, "answer": d} + rec["refused_as_expected"] = code not in ("200", "202") + # let the restore settle, then remove properly + for _ in range(60): + st = w.stack(app) + if st.get("state") in ("running", "unhealthy", "stopped", "not_deployed"): + break + time.sleep(5) + time.sleep(10) + code, d = w.ctl("POST", f"/api/stacks/{app}/stop") + time.sleep(8) + code, d = w.ctl("POST", f"/api/stacks/{app}/remove", {"remove_hdd_data": True, "remove_backups": True}) + rec["remove_after_restore"] = {"http": code, "verified": (d.get("data") or {}).get("verified"), + "reappeared": (d.get("data") or {}).get("reappeared_removed")} + time.sleep(30) + rec["containers_60s_later"] = w.guest( + f"docker ps -a --filter label=com.docker.compose.project={app} --format '{{{{.Names}}}}'").strip() + save("B", rec) + + +def C(): + """A half-state — app.yaml on disk, deployed=false, no containers — must be REMOVABLE.""" + app = "sparkyfitness" + rec = {"scenario": "C (R-634)", "app": app} + # build the exact measured shape by hand on the scratch guest: the stack dir with an app.yaml + # and no containers, which is what two failed deploys left behind twice. + rec["setup"] = w.guest(f'''set -e +mkdir -p /opt/docker/stacks/{app} +cp /var/lib/docker/volumes/felhom-controller-data/_data/data/catalog-cache/templates/{app}/docker-compose.yml /opt/docker/stacks/{app}/ 2>/dev/null || true +printf 'desired_state: stopped\\ninstalled_images:\\n' > /opt/docker/stacks/{app}/app.yaml +ls -la /opt/docker/stacks/{app}/''').strip()[-300:] + w.ctl("POST", "/api/stacks/rescan") + time.sleep(6) + st = w.stack(app) + rec["before"] = {"deployed": st.get("deployed"), "state": st.get("state")} + code, d = w.ctl("POST", f"/api/stacks/{app}/remove", {"remove_hdd_data": False, "remove_backups": True}) + rec["remove"] = {"http": code, "answer": str(d)[:300]} + rec["accepted"] = code == "200" + time.sleep(5) + rec["leftovers"] = w.guest(f"ls /opt/docker/stacks/{app}/app.yaml 2>/dev/null || echo NONE").strip() + save("C", rec) + + +if __name__ == "__main__": + w.login() + for fn in (B, C): + try: + fn() + except Exception as e: + save(fn.__name__ + "-ERROR", {"error": f"{type(e).__name__}: {e}"}) diff --git a/documentation/backlog/OPEN-ITEMS.md b/documentation/backlog/OPEN-ITEMS.md index 8cafd925..8cae2ce8 100644 --- a/documentation/backlog/OPEN-ITEMS.md +++ b/documentation/backlog/OPEN-ITEMS.md @@ -772,7 +772,7 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server` | **R-611** | **[P3-LOW] A session reported "everything is done" over a phase it had silently skipped — the process failure, not the missing measurement.** The 2026-09-21 update-arc session's brief contained a Phase 5 spike: one app updated by the box with **nobody pressing anything**, once succeeding and once forced to fail. **It did not run, and nothing said so** — no evidence file, no code, and no sentence in `STATUS.md`, `UPDATE-ARC-STATE-2026-09-21.md`, either `REPORT.md` or `09`. `09` §6.2 was left describing Slice 6 "as it would be built" with no measurement under it, which reads like a considered design rather than an untested one. **Why this is a row and not a grumble:** the missing measurement was recoverable in an afternoon; the missing SENTENCE was not, because the next reader had no way to know it was missing. A skipped phase that is declared costs one line; a skipped phase that is not costs the next session its baseline. **CLOSED 2026-09-21 by the successor session**, which ran it (`audits/update-arc-gaps-2026-09-21/`, scenarios F and G) **and** adopted the standing habit that closes it generally: **the report's FIRST section is "not done", even when empty.** Every part and scenario of a brief is listed there if it was skipped, shortened or changed, with the reason. | **CLOSED 2026-09-21 — run by the successor session; "not done" is now the report's first section** | | **R-612** | **[P1-HIGH] `wishlist` cannot be signed up to on a fresh Felhom install, the deploy reports SUCCESS, and the error the customer sees is a LIE.** MEASURED 2026-09-21 on guest 9202 while seeding for the power-cut drill. The image's first-boot `pnpm prisma db seed` is **`Killed` — OOM at the catalog's `mem_limit: 128M`**. Without it the `Role` and `Group` rows are absent, so **every** signup fails. **The message the user is shown is `User with username or email already exists`** while the container log says the real cause: `FOREIGN KEY constraint violated`. A household would conclude the account already exists and try to recover a password that was never created. **The controller reports the app running and HEALTHY throughout, and the deploy reported successful** — so nothing on the box says anything is wrong. Repaired on the scratch guest only, to unblock seeding: memory raised to 512 M, the image's own seed re-run, memory put back to 128 M. **The catalog was NOT changed** — the fix is a memory-limit question for the catalog and is deliberately left to a session that can measure the real ceiling rather than guess it. **Needs: the actual peak RSS of that seed, then a `mem_limit` that clears it, plus a check that the seed's failure is not silent.** | **READY — rank P1-HIGH; owner: CC (catalog + a look at whether a failed first-boot seed can ever be visible)** | | **R-613** | **[P2-MEDIUM] `uptime-kuma` parks on its setup wizard with no login and no monitors, and the box tells the household it is HEALTHY.** MEASURED 2026-09-21 on guest 9202. On first boot uptime-kuma 2.4.0 sits at `SETUP-DATABASE` (`Waiting for user action...`) and its main socket.io server never starts. **The controller's `http :3001` probe sees the wizard's 302 and records the app as running and healthy.** So a monitoring app that cannot be logged into, and is monitoring nothing, is presented to the customer as fine. Passed through its own front door for the drill with `POST /setup-database {"dbConfig":{"type":"sqlite"}}`. **This is the health-check class the catalog skill already warns about — a probe that proves the PORT answers, not that the APP works** — and it is worth a row because the failure direction is a false GREEN, which no alarm will ever catch. **Needs: a healthcheck for this template that fails while the wizard is up** (the catalog `REUSE.md` maps the families), and a sweep for other templates whose probe would pass on a setup wizard. **— UPDATE NIGHT 2026-09-21:** the update night could not seed `uptime-kuma` for the same reason and left it out rather than faking it. | **READY — rank P2-MEDIUM; owner: CC (catalog)** | -| **R-614** | **[P3-LOW] A stale `update_phase` survives a remove and redeploy, so a freshly installed app can read "Frissitve" before it has ever been updated.** OBSERVED 2026-09-21 on guest 9202: a newly deployed `uptime-kuma` read `update_phase=done` / „Frissitve" before any update had been run against it — left in the manager's IN-MEMORY stack state by the previous session's update of a since-removed instance of the same name. It cleared on the next controller restart. **Why it matters beyond the cosmetic:** `update_phase` is one of the fields a person (and, after R-609, an unattended caller) reads to decide whether an update happened. A value that outlives the app it described is the same class as the R-166 "absent means unknown" family — a confident answer about something that no longer exists. **Fix shape:** clear the update fields when a stack is removed, beside wherever `Updating`/`updateHeld` are reset; a test that removes and redeploys and asserts the phase is empty. Small. | **READY — rank P3-LOW; owner: CC (controller)** | +| **R-614** | **[P3-LOW] A stale `update_phase` survives a remove and redeploy, so a freshly installed app can read "Frissitve" before it has ever been updated.** OBSERVED 2026-09-21 on guest 9202: a newly deployed `uptime-kuma` read `update_phase=done` / „Frissitve" before any update had been run against it — left in the manager's IN-MEMORY stack state by the previous session's update of a since-removed instance of the same name. It cleared on the next controller restart. **Why it matters beyond the cosmetic:** `update_phase` is one of the fields a person (and, after R-609, an unattended caller) reads to decide whether an update happened. A value that outlives the app it described is the same class as the R-166 "absent means unknown" family — a confident answer about something that no longer exists. **Fix shape:** clear the update fields when a stack is removed, beside wherever `Updating`/`updateHeld` are reset; a test that removes and redeploys and asserts the phase is empty. Small. **FIXED in v0.262.0.** `RemoveStack` calls `ClearUpdateState`, which wipes `Updating`, `UpdatePhase`, `UpdatePhaseLabel`, `UpdateError`, `updateHeld` and `HealthProbe`. The name is the only thing a new install shares with the old one, so the record has to go when the app does. Red-proof seen failing. | **CLOSED 2026-09-22 — v0.262.0: a remove clears the update record** | | **R-602** | **[P3-LOW] The language a signed-in page uses is NOT the language a cookie asks for, and a live probe that forgets this reports a fixed defect as unfixed.** FOUND 2026-09-21 verifying R-598 on guest 9201. `GET /backups` with `felhom_lang=en` returned the **Hungarian** page. That is correct — `langFor` step 2 says a request carrying a session reads the household's saved setting and deliberately ignores the visitor cookie, so a signed-in family never sees a language a previous visitor picked on the sign-in page of the same browser — but it means **the cookie is the right instrument for the anonymous claim/login/bind pages and the wrong one for every page behind auth**, where `?lang=` is. A session that had run only the cookie probe would have concluded R-598 was still open and fixed it a second time. **This is a documentation gap, not a code defect**, and it is the kind that costs a whole session: nothing in `10-localisation.md` §2.2 or in any runbook tells a prober which instrument to use where. **Fix shape:** four lines in `10-localisation.md` §2.2 — a table of surface → language instrument — and a pointer from the live-validation section of the workspace rules. Recorded meanwhile in `audits/i18n-closing-2026-09-21/live/backups-page.md`. | **READY — rank P3-LOW; owner: CC (docs)** | | **R-603** | **[P3-LOW] An English string containing an apostrophe silently never matches on a rendered page, and a `strings.Contains` assertion reads exactly like a missing sentence.** FOUND 2026-09-21 while writing the R-598 render tests. `backup.target.absent` was first written as *"The system backup's drive cannot be reached…"*; `html/template` escapes `'` to `'`, so the page carried the sentence and every assertion for it failed. **The failure mode is the expensive part:** the test said *"the English absent-drive copy never reached the page"*, which is indistinguishable from the handler not being wired — and the obvious next move is to go and re-fix the handler. Reworded to avoid the possessive, and all 23 new English values were then swept for `' " < > &` (zero). **The Hungarian bundle has never hit this** because Hungarian copy uses „quotes" and few apostrophes; **English copy will hit it again.** **Fix shape:** either a bundle gate that refuses an HTML-escapable character in a value destined for a page (and an allow-list for the ones that legitimately need one), or a test helper that compares against `html.EscapeString(want)` so the assertion cannot be fooled. The gate is the better shape — the helper only protects tests that remember to use it. | **READY — rank P3-LOW; owner: CC (controller)** | | **R-537** | **[P1-HIGH] The app-backup page labels the tier-1 backup „DB + Konfig + Adatok" and prints the app's data-drive size next to it — but the tier-1 unit contains NO drive-side app data at all.** MEASURED 2026-09-16 on the drill box (fresh install, controller 0.243.0, one drive, tier 2 and tier 3 both „Nincs beállítva"): five photos (3 000 000 B) were uploaded into Nextcloud through its own WebDAV interface, then the customer-visible „Mentés most" was pressed (`POST /api/backup/run` → 200, the unit grew 25 337 B → 978 MB). The resulting unit's `manifest.json` lists `db-dumps` + three **docker volume** dumps and nothing else; listing the 781 MB `nextcloud_nextcloud_html.tar` (29 346 entries, positive control `version.php` = 3 hits) gives **`Fotok` = 0 and `nyaralas` = 0**, and `./data/` is the empty bind-mount point. A `find` over the whole `backups/` tree for `*appdata*` / `*Fotok*` returns nothing. The page nevertheless renders „1. mentés … DB + Konfig + Adatok" and „Nextcloud Adatlemez 65.1 MB" — a size measured on exactly the data it does not copy (`internal/web/handlers.go:1176-1178`, `BackupContents`). **This is a truth defect, not a design defect:** `07-backup-architecture.md` §6.2 places nextcloud's file leg at **Tier 2 and Tier 3 only**, and its „[FACT] What the whole-guest tiers do NOT carry" says `mp8 /mnt/felhom-drives` is out of vzdump scope (confirmed live: „excluding bind mount point mp8 … (not a volume)"). So on a one-drive box with no off-site tier — the state every fresh install starts in — the household's files are in **no backup**, while the page says „Adatok". Same family as R-517/R-518. **Fix shape:** render tier-1 contents from the capture set actually written (`ComputeCaptureSet`), so a unit with no file leg reads „DB + Konfig" and the drive size is not shown beside it; and say on the page that the app's files need tier 2 or tier 3. Evidence: `audits/evidence-drill-0243-2026-09-16/phase2-f10.txt`. **CLOSED 2026-09-16 — controller v0.244.0, proven live.** The contents label is computed PER TIER from what that tier captures: Tier 1 says „Adatok" only when the app's data really is in the volumes the unit captured, and a class-A app carries one sentence saying where its files ARE protected. Proven on demo-hp through the page the customer opens: Paperless-ngx reads „1. mentés … DB + Konfig" with „Az alkalmazás fájljait a távoli másolat (és a második meghajtó) védi …", while its „2. mentés" row still reads „DB + Konfig + Adatok". Red-proof: restoring the old app-shaped label fails `TestAppBackupRows_Tier1LabelDoesNotClaimFilesItCannotHold`. **RE-PROVEN 2026-09-16 on a FRESH box** (installed from the built ISO 1.28.0, controller 0.244.0, off-site on by default): the Nextcloud row read „1. mentés … DB + Konfig" with the new sentence, „2. mentés … Nincs 2. (off-drive) másolat", „3. mentés Sikeres restic → …your-storagebox.de"; „DB + Konfig + Adatok" appeared ZERO times while the local unit held no file leg. | **CLOSED 2026-09-16 — controller v0.244.0 (proven live on demo-hp)** | @@ -790,7 +790,7 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server` | **R-618** | **[P1-HIGH] THREE apps are presented to the household as UNHEALTHY while they are working perfectly — and because the guarded Update waits on that same probe, a SUCCESSFUL update ends by STOPPING the working app and sending the household to a restore they do not need.** **RANK RAISED FROM P2 TO P1 BY A LIVE MEASUREMENT taken the same night, and the escalation is the whole point:** tandoor's Update 2.6.13 → 2.6.15 was pressed at 21:16:47 and entered `verifying` at 21:17:46. At **21:18:28** the NEW version was `Up 25 seconds` and answering **HTTP 200** on `/accounts/login/` through the household's own front door — while the controller, probing port 8080 where nothing listens, could not see it. `verifying` therefore cannot pass, the full `update.health_timeout` is spent, `Manager.failAndHold` runs `compose down`, and the app is STOPPED. **Nothing is lost** — the data is in the volumes and the restore works — **but one wrong port number in a template converts every successful update of that app into an outage plus an unnecessary restore, for every household running it.** Evidence: `audits/update-night-2026-09-21/14-tandoor-serving-while-verifying.txt`. MEASURED 2026-09-21 on guest 9202 (controller v0.261.0, catalog `f5f6a152b513`). Two shapes, one class: **(a) `tandoor` — the WRONG PORT.** `.felhom.yml` probes `port: 8080`; the container listens on **80 and nothing else** (`ss -ltn` inside it), the compose's own traefik label routes to 80, its own docker healthcheck reads `healthy`, and `/accounts/login/` answers **200** through the household's real front door. `GET /api/stacks/tandoor` nevertheless reads `state: "unhealthy"`. **(b) `zipline` — the WRONG PATH.** `.felhom.yml` probes `/api/health`, which zipline 4.6.1 answers **404 `Route GET:/api/health not found`**; **the compose healthcheck in the very same file uses `/api/healthcheck` and is correct and green.** `/dashboard` answers 200. The controller reads `unhealthy`. **This is the MIRROR of R-613** — that is a probe that passes on a broken app (a false GREEN, which no alarm catches); this is a probe that fails on a working app (a false RED). **IT DOES NOT ALARM, AND THAT SETS THE RANK:** `08-alarm-ladder.md` §4 puts `unhealthy` deliberately in the NOT-down set, so no dead-app event and no customer mail follows — the damage is what the household READS, plus anything that gates on `state`. **IT ALREADY COST A MEASUREMENT TONIGHT:** this drill's harness waited for `state == "running"` and hung for its full budget on tandoor, an app that was up the whole time. An instrument waiting for a wrong answer looks exactly like a slow app. **THE GATE THIS WANTS IS CHEAP AND STATIC, AND THAT IS THE FINDING'S REAL VALUE.** Both halves of the answer live in the same template: compare the `.felhom.yml` probe's port and path against the compose's **own** `healthcheck: test:` URL. A sweep of all 53 templates on that rule was run tonight and returns **five** disagreements: `tandoor` (PORT — **CONFIRMED live**), `zipline` (PATH — **CONFIRMED live**), `wger` (PORT, probe 80 vs compose 8000 — **CONFIRMED live the same night**), `home-assistant` (PATH, `/api/` vs `/manifest.json` — **NOT MEASURED**), and `adventurelog` (a FALSE POSITIVE of the sweep's own regex — it reads `running` live). **So the rule finds both real defects, with two candidates and one false positive out of 53** — a good enough signal for a fast gate, provided it reports candidates rather than convictions and a person or a runtime check resolves them. The earlier, cruder rule (probe port vs the *traefik* port) is strictly worse: it clears zipline and convicts adventurelog. **AND A SECOND FIX SHAPE, ON THE CONTROLLER SIDE, WORTH CONSIDERING BESIDE THE CATALOG ONE:** in both confirmed cases the container's OWN docker healthcheck was **green** the whole time. A `verifying` phase that is about to stop a working app could ask that too — if the compose declares a healthcheck and docker reports `healthy`, the app is alive whatever our probe thinks. That does not excuse a wrong probe, but it turns this failure direction from an outage into a wrong label. It is a design question, not a defect, and is raised here rather than decided. **Needs:** fix tandoor's port (80), zipline's path (`/api/healthcheck`) and wger's port (8000) — **all three are now CONFIRMED live, none is a guess**; add the static gate with a decoy each way (R-421) — a template whose probe agrees must not read as a disagreement, and vice versa. **THE GATE'S RULE WAS THEN SHARPENED BY READING `healthprobe.go` RATHER THAN ASSUMING IT, and the sharpening REMOVED a false conviction.** `type: http` treats **any** response as healthy (`healthprobe.go:258-261`), and `type: api` with **no** `expect` block does the same (`:265-268`); only `type: api` WITH `expect.status` cares about the path or the code. So a PATH difference is a candidate only for the third shape, while a PORT difference is a candidate for all of them. Under that rule the 53-template sweep returns **four** candidates — `tandoor`, `zipline` and **`wger` (all three CONFIRMED live — probe `type: http, port: 80`; inside the container port 80 is `refused` and port 8000 `ANSWERED`; docker's own healthcheck green; front door 302; the box reads `unhealthy`)**, and `adventurelog` (a false positive: its compose lists two containers' ports and the probe targets the backend; measured `running`). **`home-assistant` is correctly CLEARED by the sharpened rule** — `type: api`, no `expect`, so its `/api/` answering 401 without a token is healthy, and its edge was PROVEN on the box tonight. The crude rule convicted it; the rule read from the code does not. **That is the gate to build: two of 53 convicted, one suspected, one false positive, and the false positive is resolvable by one live check.** Evidence: `audits/update-night-2026-09-21/10-probe-port-sweep.txt`, `12-probe-vs-compose-healthcheck.txt` and `13-probe-sweep-sharpened.txt`. **CLOSED 2026-09-22.** All three fixed in one commit (`app-catalog-felhom.eu@793c4fb`): tandoor `8080 -> 80`, wger `80 -> 8000`, zipline `/api/health -> /api/healthcheck`. No `image:` line moved, so no `catalog_since` moved. **RED-PROOFED LIVE ON 9202 THROUGH THE PRODUCT, BOTH DIRECTIONS.** Before the fix, at the LIVE pin, all three read **`Nem egészséges` / `Not healthy`** on their own app page while docker reported every container healthy and each front door served a real page through the household's own route — tandoor 200 `Login / Sign In`, zipline 200 `Zipline`, wger 200 `wger Workout Manager`. The fix was applied through the REAL sync (`POST /api/sync` answered *frissítve: tandoor, wger, zipline*) and all three read **`Fut` / `Running`** at the next poll, with no redeploy and no restart. **AND THE EDGE THAT FAILED WAS RE-WALKED AND PASSED:** tandoor `2.6.13 -> 2.6.15` via the drill catalog ended **`done` at +41.1 s** with the seed read back through tandoor's own front door and both containers running with zero restarts — where the identical edge on 2026-09-21 entered `verifying` at +58.4 s and ended `failed` at **+361.9 s** with the app stopped. **Same app, same versions, same button; the only change is one port number.** tandoor's verdict moved `failed -> proven` and it is now on the live catalog. **THE GATE SHIPPED WITH IT:** `scripts/check-probe-matches-compose.py`, a `--fast` row in `catalog_gates.py`, comparing the probe against the SAME service's own compose healthcheck. The rule was read out of `healthprobe.go` rather than guessed: a wrong PORT refuses for every check type; a wrong PATH refuses only for `type: api` WITH an `expect` block and WARNS otherwise, which is why `home-assistant` is warned about and not convicted. Four red-proofs and five decoys, plus five more with PyYAML shadowed out (R-630's sibling problem — see below). **Residual, filed separately:** R-630 (paperless-ngx's probe can never run at all) and R-631 (five templates the static rule cannot judge). | **CLOSED 2026-09-22 — three probes fixed, red-proofed live both ways, gate shipped with decoys; tandoor re-walked `failed -> proven`** | | **R-619** | **[P3-LOW] A `type: password` deploy field is MANDATORY however `required` reads, and the `deploy-fields` contract says the opposite — so any caller that trusts it is refused.** MEASURED 2026-09-21 on guest 9202 while widening the update drill. `GET /api/stacks/grafana/deploy-fields` serves `{"env_var":"GF_SECURITY_ADMIN_PASSWORD","type":"password","generate":"password:16","required":false}`; a deploy carrying only the two `required:true` fields is refused **400** „a(z) „Admin jelszó" mező kitöltése kötelező — használja a Generálás gombot…". **The BEHAVIOUR is right and is a decision, not a bug:** `deploy.go:305-312` refuses a `password` field with no caller value on purpose — *"We never silently auto-generate — the user needs to know their password"* — which is the opposite of the `secret` case one branch above, where a generated value the customer never sees is exactly correct. **The defect is the CONTRACT.** `.felhom.yml` declares `required: false`, the API serves that verbatim, and nothing on the wire distinguishes "optional because the box will generate it" (`secret`) from "optional in the template and mandatory in the code" (`password`). A person using the deploy page never meets this because the page renders a Generálás button; **anything that is not that page does**, which now includes this drill harness and would include `09` §6.2's unattended caller the day it deploys anything. **Fix shape (smallest that keeps the decision):** serve `required: true` for `type: password` in the deploy-fields response — one place, derived rather than stored, so templates need no edit — and a test asserting a `password` field always reaches the wire as required. Alternatively state it in the field's `description`, which is weaker because it is prose. Evidence: `audits/update-night-2026-09-21/apps/grafana/log.txt` (the refusal) and `batchA.log`. | **READY — rank P3-LOW; owner: CC (controller)** | | **R-620** | **[P3-LOW] A disabled notifier drops every event with NO local trace, so a box whose hub configuration is absent or broken stops telling anyone anything and leaves nothing behind that says so.** FOUND 2026-09-21 on guest 9202 while trying to score the update night's alarm truth table. `hub.enabled: false` there, and `Notifier.Publish` returns at `notify/notifier.go:269` — **before** any log line — as do `NotifyHealthChange` (:359) and four more entry points. Startup says it once (`[INFO] Notifier disabled (hub not configured)`) and then every later event, of every severity up to `critical`, vanishes without a word. **The measurable consequence tonight:** the whole event-and-mail half of the drill was structurally unmeasurable on this venue, and the alarm truth table below covers only the app page, the dashboard and the box's own log. That is a cost this session paid and named; the next one would pay it again. **The consequence on a real box is smaller but not zero:** the fleet's boxes have the hub enabled, and total silence is already caught by the hub's dead-man's-switch (staleness from the LAST REPORT, proven in the 2026-07-22 power-outage audit). What is NOT caught is the in-between — a box that still reports but whose notifier was disabled by a bad config push would go on reporting healthy while dropping every alarm, and the only evidence would be a single INFO line at the last restart. **Fix shape:** one DEBUG (or WARN, once per event type) line on the disabled path naming the event that was dropped, so the absence is visible where it happens rather than inferable from a startup line. Cheap, and it converts an invisible failure into a greppable one — R-96 rule 3 in the place that produces it. Evidence: `audits/update-night-2026-09-21/11-notifier-disabled.txt`. | **READY — rank P3-LOW; owner: CC (controller)** | -| **R-621** | **[P2-MEDIUM] A held update DESTROYS the evidence of why it failed: `failAndHold` runs `compose down`, the failing containers are removed, and their output is gone before anyone — household, operator or the next session — can read it.** MEASURED 2026-09-21 on guest 9202 on a REAL upstream edge: `adventurelog v0.12.1 → v0.13.0`. The new backend applied **nine Django migrations successfully** and then never listened; the update held after the full 5-minute health wait. **`Manager.failAndHold` (`stacks/update.go:723`) calls `updateCompose(dir, env, "down")`**, which removes the containers rather than stopping them, and nothing captures their logs first. Within seconds the box's own log recorded `Logs result for adventurelog: 0 bytes returned (empty)` and `docker ps -a` held nothing at all. **What survives is the WHAT and not the WHY:** the controller line `update adventurelog FAILED after the new version was started: not healthy: not healthy within 5m0s (last: state unhealthy)` and the household's sentence, both of which say the app did not come up and neither of which says the migrations ran and the server then failed to bind. **This is R-320 ("evidence off the machine before the teardown") as a PRODUCT behaviour rather than a session habit** — the teardown here is the product's own, it is correct to perform (a half-started new version must not keep running), and it happens before anyone can look. **Why it matters beyond a drill:** the hold sentence sends the household to a restore, and after the restore the only remaining question is *should I press Update again?* — which nobody can answer, because the one artefact that would say so no longer exists. It also makes every future held update unreportable to an upstream project. **Fix shape:** capture `compose logs --no-color --tail N` into the stack directory (beside `applied-compose.yml`, which already travels with the stack) IMMEDIATELY before the `down`, and surface it on the app page's hold panel or at least through the existing `/api/stacks//logs` fallback. Bounded size, written once per hold. A test that holds an app and asserts the captured file is non-empty — it fails today. Evidence: `audits/update-night-2026-09-21/apps/adventurelog/why-it-failed.txt`, `state-after-hold.txt`. | **READY — rank P2-MEDIUM; owner: CC (controller)** | +| **R-621** | **[P2-MEDIUM] A held update DESTROYS the evidence of why it failed: `failAndHold` runs `compose down`, the failing containers are removed, and their output is gone before anyone — household, operator or the next session — can read it.** MEASURED 2026-09-21 on guest 9202 on a REAL upstream edge: `adventurelog v0.12.1 → v0.13.0`. The new backend applied **nine Django migrations successfully** and then never listened; the update held after the full 5-minute health wait. **`Manager.failAndHold` (`stacks/update.go:723`) calls `updateCompose(dir, env, "down")`**, which removes the containers rather than stopping them, and nothing captures their logs first. Within seconds the box's own log recorded `Logs result for adventurelog: 0 bytes returned (empty)` and `docker ps -a` held nothing at all. **What survives is the WHAT and not the WHY:** the controller line `update adventurelog FAILED after the new version was started: not healthy: not healthy within 5m0s (last: state unhealthy)` and the household's sentence, both of which say the app did not come up and neither of which says the migrations ran and the server then failed to bind. **This is R-320 ("evidence off the machine before the teardown") as a PRODUCT behaviour rather than a session habit** — the teardown here is the product's own, it is correct to perform (a half-started new version must not keep running), and it happens before anyone can look. **Why it matters beyond a drill:** the hold sentence sends the household to a restore, and after the restore the only remaining question is *should I press Update again?* — which nobody can answer, because the one artefact that would say so no longer exists. It also makes every future held update unreportable to an upstream project. **Fix shape:** capture `compose logs --no-color --tail N` into the stack directory (beside `applied-compose.yml`, which already travels with the stack) IMMEDIATELY before the `down`, and surface it on the app page's hold panel or at least through the existing `/api/stacks//logs` fallback. Bounded size, written once per hold. A test that holds an app and asserts the captured file is non-empty — it fails today. Evidence: `audits/update-night-2026-09-21/apps/adventurelog/why-it-failed.txt`, `state-after-hold.txt`. **FIXED in v0.262.0.** `failAndHold` now writes each service's log into `/hold-logs//compose-logs.txt` (`compose logs --no-color --tail 400`) **before** the `down` that destroys them, best-effort by design: a hold must never fail because its evidence could not be written. This is what `adventurelog` cost — nine migrations ran, the app never bound its port, and the only log that could have said why was gone before anyone looked. | **CLOSED 2026-09-22 — v0.262.0: the hold keeps the app's own log before stopping it** | | **R-622** | **[P2-MEDIUM] `adventurelog v0.13.0` migrates the customer's database and then does not serve — the edge must NOT be promoted, and it is the first real-catalog candidate this project has measured as unsafe.** MEASURED 2026-09-21 on guest 9202 through the product's own guarded Update. `v0.12.1 → v0.13.0` (backend AND frontend together, PostGIS held constant). The backend applied **nine migrations, every one `... OK`** — `adventures.0072_trail_wanderer_author_fields` through `integrations.0009_alter_endurainintegration_auth_method`, plus `billing.0001_initial` — and then the container's own healthcheck failed with `URLError: [Errno 111] Connection refused` on five consecutive checks. The app never bound its port. The update held honestly after the full 5-minute wait. **THE PRODUCT DID EVERYTHING RIGHT AND THAT IS HALF THE FINDING:** the precondition found a Tier-1 copy one minute old, the safety dump was written, the pin advanced BEFORE the pull, the health wait was not short-circuited, the app was stopped rather than left half-running, and the hold sentence named the tier, the date and what the copy holds — „saját meghajtó, 2026-09-21 20:47 — ez a másolat a beállításokat, az adatbázist és az adatköteteket tartalmazza." **This is exactly the case `09` §4 exists for:** the migration RAN, so there is no undo, only a restore — and the restore is the thing slice 4 made sure existed first. **What it needs:** adventurelog stays OFF the promotion list; the cause is not diagnosed here (R-621 is why); and before it is ever promoted the edge should be re-run on the harness with its ABORT step, since an app that migrates and then refuses is the shape most likely to refuse the old image too. Evidence: `audits/update-night-2026-09-21/apps/adventurelog/`. | **READY — rank P2-MEDIUM; owner: CC (catalog); NOT a Felhom defect — an upstream edge that fails** | | **R-623** | **[P3-LOW] The unattended-update caller turned every SUCCESS into a `timeout`, and then refused to press that app again — the instrument, not the box.** FOUND 2026-09-21 (update night) by reading `unattended-caller.py` before relying on it for the Q4 hold measurement. Its `call()` returns the API **envelope** — `{"ok": true, "data": {…}}` — and `follow()` read `update_phase` and `updating` **off the envelope**, where neither exists. Both were therefore always `None`; the end test `not updating and phase in ("done","failed")` could never fire; every followed update ran the full **900-second** timeout and was recorded `timeout`, which the caller treats as terminal and adds to `never_again`. `main()` unwraps `data` for the stack LIST, which is exactly why the within-a-major half of that night worked and this half did not. **The 2026-09-21 run did not catch it because the only pass that reached `follow()` was Scenario F, whose log was lost to a buffering `tail`** — the run's own honestly-recorded instrumentation gap turns out to have hidden a second one underneath it. **This is the R-607 class in the evidence layer rather than the product layer: an instrument that can report a success as a timeout is not a measurement**, and worse, it is a measurement that says the box behaved badly when the box behaved well. **FIXED in the same file 2026-09-21** (unwrap `data`, with the reason written into the docstring so the next reader does not re-derive it), and the fixed caller is what produced tonight's unattended-hold leg. **What it does NOT invalidate:** the G-b no-retry proof, which is entirely in the refusal path and never reached `follow()`. **What it DOES qualify:** any future reading of that night's Scenario F timing — the "51 s – 1 m 26 s" figures come from the ATTENDED scenarios 04/05/07, not from the caller. | **CLOSED 2026-09-21 — fixed in `audits/update-arc-gaps-2026-09-21/unattended-caller.py`** | | **R-624** | **[P3-LOW] Three of the catalog's apps cannot be seeded by ANY headless route, and for two of them that is a deliberate security decision — so the upgrade harness has a permanent ceiling nobody has written down.** FOUND 2026-09-21 while widening R-462 from 3 apps to . **`vaultwarden` and `zipline` close self-registration ON PURPOSE** — vaultwarden by `SIGNUPS_ALLOWED=false` (R-512, *„a stranger who guesses vault. must not be able to register"*), zipline by answering `E1037: User registration is disabled` — and neither ships a CLI that could make an account instead. So there is no route to a first account without the admin secret, and **that is correct**: the harness must not be the reason a customer-facing app accepts strangers. **`gitea` is a different and fixable case:** the template sets no `INSTALL_LOCK`, so a fresh instance sits in its web-installer state and `gitea admin user create` refuses (`MustInstalled() [F] Unable to load config file for a installed Gitea instance`); POSTing the installer form first would work and was simply not written tonight. **Why this is a row rather than three notes:** `09` §3 decision 6 says the upgrade test goes to **all** apps, and R-462 is costed as if every app is reachable given enough fixture work. It is not. There is a class — *apps whose only account-creating route the catalog deliberately closes* — for which the honest maximum is `inconclusive` unless the harness is given the app's admin secret at deploy time, which is a decision nobody has taken. **Needs:** the class named in R-462's scope so the remaining count is honest; a decision on whether the harness may hold an app's admin secret (it already holds the ones IT generates — see R-619); and, separately and cheaply, a gitea installer-form fixture. Evidence: `audits/update-night-2026-09-21/apps/{vaultwarden,zipline,gitea}/verdict.json`. | **READY — rank P3-LOW; owner: CC (catalog harness)** | @@ -799,11 +799,11 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server` | **R-628** | **[P2-MEDIUM] An empty search of a mailbox I do not control was turned into a claim about what a THIRD PARTY had done, and it went into the register as fact.** FOUND 2026-09-22; the operator caught it within minutes by producing the thread. The `due_checks_gate` fired R-433 (*have Hetzner answered?*). Two searches of the felhom catch-all came back empty, and the emptiness was written into R-433 as **"Hetzner has not answered"** and **"there is no evidence the tickets were ever opened"**. Both were false: ticket **#2026090103040671** had been opened and answered. **THE PRECISE FAULT IS THE INFERENCE, NOT THE QUERY — and that distinction is the whole value of this row.** Re-run afterwards WITH a positive control: `from:monitoring@felhom.eu` returns **201 threads**, `in:anywhere … includeTrash` reaches SENT, TRASH and mail back to January — **the instrument works.** And the exact ticket number, the exact subject `Storage Box issue`, and `from:hetzner` each still return **nothing**. So the literal finding — *this correspondence is not in this mailbox* — was CORRECT. What was invented was the step from there to *Hetzner has not answered* and *the tickets were never opened*. **A mailbox I can read is not the only place a reply can be**, and the operator's own account is exactly where a support ticket he opened would land. An absent record in ONE place can never answer a question about what SOMEONE ELSE did. **This is R-96 rule 3 in a new surface, one step further out than R-607**: there the instrument reported a stale value as current; here a correct observation was promoted to a conclusion it could not carry. **It is sharper still because the same session, the night before, gave every fixture a negative control and every gate a red-proof — and then reached for a search with neither.** **THE RULE, in one line:** *an empty search may be reported as "absent from the place I looked", never as "it did not happen" — and only after a control query that MUST hit has been seen to hit.* **Done:** the rule is written into the Gmail-access memory, where the next session meets it before it searches rather than after. | **CLOSED 2026-09-22 — rule recorded; R-433 corrected with the real answers** | | **R-627** | **[P2-MEDIUM] Nothing checked that the register is a well-formed table, so an append that ate two rows' state cells went unnoticed until a person read the file — and one row had been broken the same way for 45 days.** FOUND 2026-09-22. The 2026-09-21 update night appended measured results to eight rows with a regex that matched each row's trailing state cell; on **R-446** and **R-458** it consumed the cell and did not restore it, the cell reappeared as a stray FOURTH cell on a DUPLICATED copy of **R-626** and **R-625**, and a blank line was left between each pair. The register then reported **317 rows for 315 findings**, two rows carried no state at all, and two findings existed twice with contradictory state cells. **Nothing caught it:** `one_register_gate.py` compares this file against ROADMAP and `closed_register_gate.py` forbids an id in BOTH files — neither asks whether the file is a well-formed table, and neither notices an id duplicated WITHIN it. **THE RED-PROOF THEN FOUND AN OLDER INSTANCE NOBODY HAD SEEN: R-254 lost its state cell on 2026-08-08 (commit `59527d0`) and had rendered without a State column for 45 days.** **Closed the same day:** `scripts/register_shape_gate.py`, registered in `repo_gates.py` as gate 14 and reached by the pre-push hook, refusing a row that does not end with `|` (an eaten state cell), a duplicated id, or a blank line splitting the table; four decoys in `test_gate_decoys.py` — three convicting on the exact damage shapes and one asserting a healthy register still passes. **TWO THINGS THE RED-PROOF CORRECTED IN THE GATE ITSELF, kept because they are the finding's real content:** a first draft counted CELLS and convicted **125 innocent rows** — register cells carry literal `|` inside prose and shell snippets (`owner: CC | …`), so a row cannot be split on `|`, and a count that cannot be computed is not a check; and it skipped malformed rows before counting ids, reporting **5 duplicates where there were 2**. **Repaired:** both state cells restored from the stray cells that carried them, the two duplicate rows deleted, R-254's verdict sentence given its cell back, and **15 blank lines that split the register into 12 separate markdown tables** removed — every row's text byte-identical afterwards, proven by diff. **And the gate immediately earned itself:** the very next row inserted in this session (R-628) left a blank line behind and the gate refused it. | **CLOSED 2026-09-22 — gate 14, four decoys, register repaired 317→315** | | **R-629** | **[P2-MEDIUM] The drill catalog sent the operator 47 CI-failure alarms in one night, and the drill method that created it did not mention CI at all.** FOUND 2026-09-22 while checking a different mailbox question — which is the only reason it was found. `admin/app-catalog-drill` was created on 2026-09-21 by `POST /api/v1/repos/migrate` from the live catalog, and a migrated repo inherits `has_actions: true`. Every drill push therefore ran the catalog's CI workflow, which failed immediately (the drill repo carries the workflow but the run has no meaningful gate context), and **each failure mailed `admin@felhom.eu`**: *"[felhom CI] gates FAILED in admin/app-catalog-drill"*. **47 runs, 47 alarms, all overnight, all unread and flagged IMPORTANT.** **WHY THIS MATTERS MORE THAN THE NOISE:** that mailbox is the operator's alarm channel, and R-168 made CI mail the thing that notices a bypassed gate. A night of throwaway failures from a repo nobody must ever act on trains the reader to skim exactly the sender that must never be skimmed — and it did it on the night the same mailbox was also carrying real `offsite_snapshots_dropped` and `offsite_proof_empty` alarms. **The drill method wrote down the fences it needed** — private repo, no customer box may follow it, reset at teardown — **and said nothing about CI, because nobody had run a drill repo through a CI-enabled Gitea before.** **FIXED 2026-09-22:** `has_actions` set to **false** on the drill repo (verified by re-reading the repo: `has_actions: False, private: True`), and `09` §6.5 now carries it as a step of creating a drill repo rather than as a thing to notice afterwards. **What is NOT done:** the 47 mails are still in the operator's inbox, unread — deleting another person's mail is not mine to do, and they are named here so they can be cleared in one search: `subject:"gates FAILED in admin/app-catalog-drill"`. | **CLOSED 2026-09-22 — actions disabled, method updated; the 47 mails are the operator's to clear** | -| **R-630** | **[P2-MEDIUM] `paperless-ngx`'s health probe has never run, on any box, and the badge can never go red.** FOUND 2026-09-22 by the new `probe-matches-compose` gate, which reported it as a WARNING while looking for something else. `findProbeContainer` (`controller/internal/stacks/healthprobe.go:297`) takes the container whose name EQUALS the stack name, else the first whose name has it as a PREFIX; paperless-ngx's containers are `paperless-webserver`, `paperless-postgres` and `paperless-redis`, and **none of them begins with `paperless-ngx`**. The function returns `""`, `:62` counts the stack in `skippedNoContainer` and `continue`s, and no probe is ever built for it. **WHY THIS IS WORSE THAN A WRONG PROBE, WHICH IS WHAT R-618 WAS:** a wrong probe is a FALSE RED — loud, visible, and it stopped an app, which is how it was found within one night. This is a SILENT ABSENCE. The app page shows whatever the container state alone says, nothing contradicts it, and **R-96 rule 3 is the exact shape: an absent alarm is equally consistent with `healthy` and with `never checked`.** **NOT MEASURED, and said so rather than assumed:** what the guarded update's `verifying` phase does for a stack with no probe target — whether it passes immediately or waits out `update.health_timeout` — has not been run. `paperless-ngx` IS installed on demo-hp, so it is answerable on a real box. **Two candidate fixes, neither taken here because both are product code and this session was forbidden it:** give the template a `container_name: paperless-ngx` on its webserver service (catalog-only, one line, but it renames a running container on every existing box); or let the probe fall back to the service the compose declares first, and SAY SO in the health detail. **What is safe to say today:** the gate names it on every push, so it cannot go back to being invisible. **MEASURED 2026-09-22 (the twenty-eight), and the answer is the WORST of the three possibilities, so this row is RAISED P2 → P1.** paperless-ngx was deployed on guest 9202 with all three containers **`healthy`**, the controller reading **`running`** and the front door answering **302**. The Update was then pressed (no upstream edge exists tonight, so it was pressed on the same version — which is what a household does on an up-to-date app and still walks the whole phase machine; stated rather than glossed). Phases: `checking` → `safety-dump` → `pinning` → `pulling` → `starting` (+2.1 s) → `verifying` (+3.1 s) → **`failed` at +313.0 s, the app STOPPED** (`state_after: stopped`, front door **404**). **THE CONTROLLER'S OWN WORDS NAME THE CAUSE, and no inference was needed:** *`update paperless-ngx FAILED after the new version was started: not healthy: not healthy within 5m0s (last: no probe container) — stopping and HOLDING the app; the pin stays on the new version (its migration may have run)`*. **`no probe container`.** So `verifying` does not pass when there is no probe and it does not skip — **it waits out the full `update.health_timeout` and then HOLDS.** **WHAT THIS COSTS:** every paperless-ngx household that presses Update has their working app **stopped for five minutes and then left stopped**, and is sent to a restore they do not need — and the hold sentence correctly warns that the copy *„csak a beállításokat és az adatbázist tartalmazza, a fájlokat nem"*, so for this class-A app the route back is the off-site copy. **This is R-618's outcome reached by the opposite road:** there a probe named the wrong port; here no probe exists at all, and the static gate cannot see it because there is nothing to compare. The gate does name it on every push as a WARNING, which is how it was found. **Needs (unchanged, and now urgent):** give the template a `container_name: paperless-ngx` on its webserver service, or make `verifying` treat "no probe target" as something other than a failure — both are product/catalog decisions. Evidence: `audits/the-28-2026-09-22/sidejobs/r630-controller-words.txt`, `sidejobs/r630.json`. | **OPEN — RAISED TO P1 2026-09-22 by measurement; owner: CC; a successful update STOPS the app** | +| **R-630** | **[P2-MEDIUM] `paperless-ngx`'s health probe has never run, on any box, and the badge can never go red.** FOUND 2026-09-22 by the new `probe-matches-compose` gate, which reported it as a WARNING while looking for something else. `findProbeContainer` (`controller/internal/stacks/healthprobe.go:297`) takes the container whose name EQUALS the stack name, else the first whose name has it as a PREFIX; paperless-ngx's containers are `paperless-webserver`, `paperless-postgres` and `paperless-redis`, and **none of them begins with `paperless-ngx`**. The function returns `""`, `:62` counts the stack in `skippedNoContainer` and `continue`s, and no probe is ever built for it. **WHY THIS IS WORSE THAN A WRONG PROBE, WHICH IS WHAT R-618 WAS:** a wrong probe is a FALSE RED — loud, visible, and it stopped an app, which is how it was found within one night. This is a SILENT ABSENCE. The app page shows whatever the container state alone says, nothing contradicts it, and **R-96 rule 3 is the exact shape: an absent alarm is equally consistent with `healthy` and with `never checked`.** **NOT MEASURED, and said so rather than assumed:** what the guarded update's `verifying` phase does for a stack with no probe target — whether it passes immediately or waits out `update.health_timeout` — has not been run. `paperless-ngx` IS installed on demo-hp, so it is answerable on a real box. **Two candidate fixes, neither taken here because both are product code and this session was forbidden it:** give the template a `container_name: paperless-ngx` on its webserver service (catalog-only, one line, but it renames a running container on every existing box); or let the probe fall back to the service the compose declares first, and SAY SO in the health detail. **What is safe to say today:** the gate names it on every push, so it cannot go back to being invisible. **MEASURED 2026-09-22 (the twenty-eight), and the answer is the WORST of the three possibilities, so this row is RAISED P2 → P1.** paperless-ngx was deployed on guest 9202 with all three containers **`healthy`**, the controller reading **`running`** and the front door answering **302**. The Update was then pressed (no upstream edge exists tonight, so it was pressed on the same version — which is what a household does on an up-to-date app and still walks the whole phase machine; stated rather than glossed). Phases: `checking` → `safety-dump` → `pinning` → `pulling` → `starting` (+2.1 s) → `verifying` (+3.1 s) → **`failed` at +313.0 s, the app STOPPED** (`state_after: stopped`, front door **404**). **THE CONTROLLER'S OWN WORDS NAME THE CAUSE, and no inference was needed:** *`update paperless-ngx FAILED after the new version was started: not healthy: not healthy within 5m0s (last: no probe container) — stopping and HOLDING the app; the pin stays on the new version (its migration may have run)`*. **`no probe container`.** So `verifying` does not pass when there is no probe and it does not skip — **it waits out the full `update.health_timeout` and then HOLDS.** **WHAT THIS COSTS:** every paperless-ngx household that presses Update has their working app **stopped for five minutes and then left stopped**, and is sent to a restore they do not need — and the hold sentence correctly warns that the copy *„csak a beállításokat és az adatbázist tartalmazza, a fájlokat nem"*, so for this class-A app the route back is the off-site copy. **This is R-618's outcome reached by the opposite road:** there a probe named the wrong port; here no probe exists at all, and the static gate cannot see it because there is nothing to compare. The gate does name it on every push as a WARNING, which is how it was found. **Needs (unchanged, and now urgent):** give the template a `container_name: paperless-ngx` on its webserver service, or make `verifying` treat "no probe target" as something other than a failure — both are product/catalog decisions. Evidence: `audits/the-28-2026-09-22/sidejobs/r630-controller-words.txt`, `sidejobs/r630.json`. **FIXED in controller v0.262.0 and in the catalog, 2026-09-22.** **(a) The branch that caused it.** `waitUpdateHealthy` (`update.go`) had the probe inside `if hc != nil && len(hc.Checks) > 0`, and when `findProbeContainer` returned `""` its `else` set `last = "no probe container"` and **looped** — the settle path that judges an app with NO declared check sat in the outer `else` and was unreachable. So `verifying` could only ever time out. It now falls through to that same settle path with a WARN naming the candidates, and the journal says `no probe container — settled on container state` so it is distinguishable from `no health check declared` without reading the log. **A stack with no probe is not healthy and not failing — it is SETTLED ON CONTAINER STATE (`09` §3), and never a reason to stop a running app.** **(b) The target is now decidable.** `HealthCheckConfig.Container` (precedent: `InitialCredentials.Container`) and `findProbeContainerMeta` resolve by **exact stack name → explicit `container` → a UNIQUE prefix → nothing, with the candidates returned**. The old rule took the FIRST prefix match, which for `immich` — four `immich-*` containers, no exact match — meant whichever the container list yielded, and which was seen live on `outline` probing `outline-postgres:3000` during startup. `paperless-ngx` and `immich` now carry the explicit field. **(c) The silence is gone:** a skipped stack gets a `HealthProbe` result carrying `health.no_probe_container` in both languages instead of nothing. **(d) The gate convicts it now rather than warning**, by the same four rules, with eight decoy cases including the no-PyYAML mode CI runs. **Red-proofs, each seen failing:** the old first-prefix rule (immich resolved to `immich-server` by luck, the explicit field ignored); `settleReason` collapsed to one sentence. | **CLOSED 2026-09-22 — ctrl v0.262.0 + catalog; the no-probe wait settles, the target is decidable, the gate refuses an ambiguity** | | **R-631** | **[P3-LOW] Five templates cannot be judged by the probe gate at all, and one is correct only by accident.** FOUND 2026-09-22 when `check-probe-matches-compose.py` was run over all 53. The gate's oracle is the probed service's own compose `healthcheck.test`; where that dials no loopback URL the gate has nothing to compare and reports a WARNING rather than a pass. **`crafty-controller`, `mealie` and `uptime-kuma`** run their healthcheck through a python or script helper (`ssl._create_unverified_context()`, `socket.create_connection`, `extra/healthcheck`), so the port is inside code the gate does not execute. **`vikunja`** has no compose healthcheck on its probed service at all. **`home-assistant`** is the interesting one: its probe path `/api/` differs from the compose's `/manifest.json`, and it is healthy today **only because its check is `type: api` with no `expect` block, which `probeHTTP` treats as "any response is healthy"** — add `expect: {status: 200}` to that template, a change that looks like a tightening, and the app goes permanently unhealthy and every successful update of it starts stopping it. **The gate warns on it for exactly that reason and refuses to call it a pass.** **Needs:** a live probe reading for each of the five on a scratch guest — deploy, read `GET /api/stacks/`, compare with the front door — which is one rotation night's work and closes the last gap R-618 left. **CLOSED 2026-09-22 — all five read live on guest 9202, and all five probes are CORRECT.** Each app was deployed, its listening sockets read from inside the probed container, and the probe's own target dialled **on the compose network**, which is the call the controller makes. `mealie` `tcp 9000` → listens `0.0.0.0:9000`, dial 200. `uptime-kuma` `http 3001` → listens `*:3001`, dial 302 (and `http` calls any response healthy, so 302 passes and proves something answers). `vikunja` `api 3456 /api/v1/info` **with `expect: {status: 200}`** → dial **200** — the one that could have failed, because its expect block compares the code. `crafty-controller` `tcp 8443` → the controller's own log reads `Health probe crafty-controller: TCP :8443 -> ok (1ms)` twice, six minutes apart. **`home-assistant` is the one to carry forward:** `api 8123 /api/` with NO expect → the dial returns **401**, not 200. It reads healthy only because `probeHTTP` treats any response as healthy for that shape (`healthprobe.go:253-262`). **Add `expect: {status: 200}` to that template — a change that looks like a tightening — and home-assistant goes permanently unhealthy and every successful update of it starts stopping it.** That is R-618 one edit away, now a measured number rather than a caution. **So the gate's WARN list is not a backlog of suspects: it is four correct templates the gate honestly cannot prove, plus one correct by accident.** Evidence: `audits/the-28-2026-09-22/sidejobs/r631.json`. | **CLOSED 2026-09-22 — five live readings, five correct probes; home-assistant's fragility is now a number** | | **R-632** | **[P3-LOW] Twenty-eight of the 53 templates have never been deployed by any update drill, so nothing is known about whether their updates work.** COUNTED 2026-09-22 against the 2026-09-21 sweep, which is the widest one ever run. **20 apps have a verdict record** (14 proven, 3 failed, 3 inconclusive, plus tandoor re-walked to proven on 2026-09-22); **4 more were deployed as props in the bad-days legs with no edge walked** (`bentopdf`, `glance`, `uptime-kuma`, `wishlist`); **1 was deployed only to measure its probe** (`wger`); and **28 have never been deployed at all**: `calcom`, `calibre-web`, `claper`, `code-server`, `crafty-controller`, `emby`, `ghost`, `gokapi`, `gramps-web`, `homebox`, `homepage`, `immich`, `jellyfin`, `kimai`, `komga`, `onlyoffice`, `outline`, `paperless-ngx`, `plant-it`, `plex`, `radarr`, `rallly`, `recipe-importer`, `seerr`, `sonarr`, `sparkyfitness`, `termix`, `wanderer`. **THIS IS NOT A COMPLAINT ABOUT THE SWEEP** — it took one app-catalog-wide night to go from 3 apps ever measured to 21, and a night is the unit available. It is a record of what the catalog's update promise currently rests on: **for 28 of 53 apps, nothing.** **The list is the nightly rotation's queue**, smallest and least stateful first; `paperless-ngx` should be early because R-630 needs a live reading from it anyway, and `crafty-controller`, `mealie` and `uptime-kuma` should be early because R-631 needs one from each. Machine-readable copy: `audits/probe-fix-2026-09-22/not-judged.json`. **WORKED IN ONE NIGHT, 2026-09-22 — all 28 walked, so this row CLOSES and hands its findings to others.** Every one was installed on guest 9202 against the private drill catalog and taken through the same walk: deploy at the live pin, seed through the app's own front door, read it back, „Mentés most”, the guarded Update where a real within-a-major edge exists upstream, **restore from that copy and read the seed back a second time** (the half the update night skipped), then remove and a 60-second check that nothing came back. **26 of 28 deployed; 6 proven; the rest inconclusive, no-edge or refused.** **What the night produced that this row could not have predicted:** R-630 raised to P1 by measurement (a stack with no probe container has its working app STOPPED by a successful update), R-633 (a remove during a restore leaves an orphan with a live public route), R-634 (an app running and healthy while recorded as not deployed, and then unremovable), and one real upstream edge that HELD honestly (`outline 1.9.1 → 1.10.1`). **Also settled:** `plant-it` is `lifecycle: abandoned` and the product refuses to install it — the only lifecycle-gated template in the catalog, and its gate is now proven live. Full record: `audits/DRILL-the-28-2026-09-22.md`. | **CLOSED 2026-09-22 — all 28 walked in one night; the findings live in R-630, R-633, R-634** | -| **R-633** | **[P2-MEDIUM] A remove sent while a restore is still running reports success, deletes the app's record, and leaves a container restarting forever with a live public route.** MEASURED 2026-09-22 on guest 9202, controller v0.261.0, during the twenty-eight walk. `gokapi` was restored from its own local copy at 11:34:07 and removed at 11:34:22. `POST /backup/restore` answers **302 and works in the background**; the remove tore down what existed and the restore's own `compose up` then RE-CREATED the container at **11:34:24**. **Both calls returned success.** Twenty-five minutes later: `GET /api/stacks/gokapi` reads **`deployed: false`**, and `docker ps -a` shows `gokapi` **`Restarting (1)`** with `RestartCount` climbing, carrying its full traefik label set — including `traefik.http.routers.gokapi.rule: Host(`.enkisfelhom.hu`)`, **a rule with an empty subdomain**, because the deploy values that filled it were deleted with the app. Its own log loops *„Salt for admin password invalid, generating new salt… password does not appear to be a SHA-1 hash"* — the volume holding its config was removed correctly, so the binary can never start. **WHY THIS IS A ROW AND NOT A HARNESS ARTEFACT:** a household can press exactly these two buttons in exactly this order, the product accepted both, and **the remove reported success while leaving the orphan**. Presence of a success message is not evidence of a result. **WHAT IT COSTS:** an app the household believes is gone keeps a container in a restart loop, keeps a route registered on the public reverse proxy, and is invisible to every product surface because the controller no longer records the stack. Nothing in the alarm ladder fires: `08` §4 keys on stacks the controller KNOWS about. **This is R-626's class with the mechanism finally visible** — that row recorded a removed `navidrome` coming back and could not diagnose it because the controller had restarted; here the window is 17 seconds and both halves are in the evidence. **Needs:** the remove path to refuse, or to wait, while a restore for the same stack is in flight — and, either way, to verify the teardown rather than report success without looking. Evidence: `audits/the-28-2026-09-22/apps/gokapi/came-back-evidence.txt`, `apps/gokapi/log.txt`. | **OPEN — P2; owner: CC; product code, so not fixed in this unattended run** | -| **R-634** | **[P1-HIGH] An app can be RUNNING, HEALTHY and serving while the controller records it as not deployed — and in that state the household cannot remove it through the product at all.** MEASURED 2026-09-22 on guest 9202, controller v0.261.0, on **two independent apps in one night**: `outline` and `sparkyfitness`. **The controller's own words, in order.** `outline` deploy accepted 11:54:47; **`11:55:51 StopStack outline: current state=deploying deployed=true containers=0`** — a stop while the stack is still deploying; `11:56:09 SaveAppConfig: saving /opt/docker/stacks/outline — 5 env vars, **0 encrypted**, 3 sensitive fields` (the two saves before it both read `3 encrypted`); then **`11:57:16` and `12:02:16 Health probe outline: API GET :3000/_health -> 200`** — the app is up and answering its own health endpoint; and **`12:02:19 StopStack outline: current state=running deployed=false containers=3`**. Three containers, state `running`, health 200, **`deployed=false`**. The remove then answers **`RemoveStack outline: state=not_deployed, deployed=false, orphaned=false, deploying=false` -> `[ERROR] Remove failed for outline: stack "outline" is not deployed`** for BOTH the remove-with-data and the remove-keeping-data call, while `ScanStacks` goes on finding the stack every ten seconds. `app.yaml` survives with `desired_state: stopped` and an **empty `installed_images`**. `sparkyfitness` produced the identical shape 13 minutes earlier. **WHAT THIS COSTS A HOUSEHOLD:** an app that works is invisible to the product as an installation — no badge, no update, no backup selection, and **no way to delete it**; the only exit is a shell. It is the mirror of R-633 (there the record is gone and the container remains; here the container is fine and the record is gone) and it is the **worse** of the two, because the app is serving customer traffic the whole time. **THE MECHANISM IS NOT DIAGNOSED, and this row says so rather than guessing.** What was tried: the controller's full container log for both apps (the sequence above), `app.yaml` on disk, `docker ps -a`, and `GET /api/stacks/`. What was NOT done: reading `runComposeDeploy`'s pin-write path — this was an unattended run and the brief forbade product code. **The one discriminator worth running first:** both apps were walked while two other walks ran concurrently, and `POST /api/backup/run` is box-wide, so a backup or restore for a NEIGHBOURING app was in flight. A serial re-walk is queued tonight; if it reproduces alone, concurrency is not the cause. **A THIRD THING THE SAME LOG SHOWS, recorded here because it is one line away:** at `11:55:57` the health probe dialled **`http://outline-postgres:3000/_health`** — during startup, with the exactly-named container not yet running, `findProbeContainer`'s PREFIX fallback latched onto the POSTGRES sidecar and probed port 3000 on it. Transient, and it resolved once `outline` came up, but it is the same function R-630 is about. Evidence: `audits/the-28-2026-09-22/apps/half-state-outline-sparkyfitness.txt`. **THE SERIAL RE-WALK WAS RUN THE SAME NIGHT AND IT SPLITS THIS ROW IN TWO — recorded here rather than left as the first reading.** Walked again one at a time, with no other walk running: **`outline` deployed normally and removed clean**, and **`crafty-controller` deployed normally, updated `4.10.7 → 4.11.0` to `done`, restored and removed clean.** So for those two the half-state did NOT reproduce alone, and concurrency — a box-wide `POST /api/backup/run` or a restore in flight for a NEIGHBOURING app — is implicated rather than the deploy path itself. **`sparkyfitness` reproduced EXACTLY, alone, in 534 s**: deploy accepted, never reached `deployed` with a pin, `app.yaml` left with `desired_state: stopped` and an empty `installed_images`, and **both remove calls refused with `stack "sparkyfitness" is not deployed`.** **So the row stands, at one reproducible app instead of three, and the honest split is:** (a) `sparkyfitness` has a deploy that does not finish and leaves a record the product cannot clear — reproducible, P1; (b) under concurrent work the same unremovable half-state can be reached by apps that are otherwise fine, which is the more alarming half because those apps were **running, healthy and serving** while recorded as not deployed. **Neither half is diagnosed** — `runComposeDeploy`'s pin write was not read, because the brief forbade product code. | **OPEN — P1; owner: CC; reproducible alone on `sparkyfitness`, concurrency-linked on the other two; next step is `runComposeDeploy`'s pin write** | +| **R-633** | **[P2-MEDIUM] A remove sent while a restore is still running reports success, deletes the app's record, and leaves a container restarting forever with a live public route.** MEASURED 2026-09-22 on guest 9202, controller v0.261.0, during the twenty-eight walk. `gokapi` was restored from its own local copy at 11:34:07 and removed at 11:34:22. `POST /backup/restore` answers **302 and works in the background**; the remove tore down what existed and the restore's own `compose up` then RE-CREATED the container at **11:34:24**. **Both calls returned success.** Twenty-five minutes later: `GET /api/stacks/gokapi` reads **`deployed: false`**, and `docker ps -a` shows `gokapi` **`Restarting (1)`** with `RestartCount` climbing, carrying its full traefik label set — including `traefik.http.routers.gokapi.rule: Host(`.enkisfelhom.hu`)`, **a rule with an empty subdomain**, because the deploy values that filled it were deleted with the app. Its own log loops *„Salt for admin password invalid, generating new salt… password does not appear to be a SHA-1 hash"* — the volume holding its config was removed correctly, so the binary can never start. **WHY THIS IS A ROW AND NOT A HARNESS ARTEFACT:** a household can press exactly these two buttons in exactly this order, the product accepted both, and **the remove reported success while leaving the orphan**. Presence of a success message is not evidence of a result. **WHAT IT COSTS:** an app the household believes is gone keeps a container in a restart loop, keeps a route registered on the public reverse proxy, and is invisible to every product surface because the controller no longer records the stack. Nothing in the alarm ladder fires: `08` §4 keys on stacks the controller KNOWS about. **This is R-626's class with the mechanism finally visible** — that row recorded a removed `navidrome` coming back and could not diagnose it because the controller had restarted; here the window is 17 seconds and both halves are in the evidence. **Needs:** the remove path to refuse, or to wait, while a restore for the same stack is in flight — and, either way, to verify the teardown rather than report success without looking. Evidence: `audits/the-28-2026-09-22/apps/gokapi/came-back-evidence.txt`, `apps/gokapi/log.txt`. **FIXED in v0.262.0, both halves.** **(a) The guard the product already had everywhere else:** `RemoveStack` now consults `UpdateGuards.Busy` (backup single-flight, restore status, app-data holds) and `IsUpdating`, and refuses with the app's own sentence in both languages. The product refused exactly this clash for `update` and for `restore` — the restore refusal even NAMES the blocking app — and `remove` was the one door with no lock on it. **(b) `down` returning 0 is a request, not a result:** the compose project is now watched for 25 s after `down`, anything carrying its label is removed BY NAME with its labels logged, and the API answer carries `verified: true/false`. That is the instrument R-626 asked for, in the product rather than in a drill script. **PROVEN LIVE ON 9202, and the live proof found a defect a test had not.** The guard fires: with `privatebin` STOPPED (so the pre-existing "still running" check could not answer first) and a box-wide backup in flight, the remove was refused and the controller said `RemoveStack privatebin REFUSED (busy): a backup or restore is running (single-flight held)`, with the household's own sentence *„Az alkalmazáson mentés vagy visszaállítás fut. Várd meg, amíg befejeződik.”* **But it answered HTTP 500.** `router.go`'s status mapping greps the error TEXT for `not deployed` / `still running` / `not found` / `protected`, and the busy sentence contains none of them, so it fell through to the default. **A 500 tells the UI something broke; this is “wait a moment”.** Fixed in **v0.262.1**: a typed `*stacks.RemoveBusyError` matched with `errors.As`, answered **409**, carrying both the Hungarian bytes and the key — and its test asserts the sentence contains none of the words the text mapping greps for, so the type is load-bearing rather than decorative. **The verification half is proven too:** a permitted remove answered `verified: true` with no reappearance, and no container carrying the project label existed 60 s later. **AND A FIRST ATTEMPT THAT PROVED NOTHING, recorded because it nearly went down as a pass:** the first run refused the remove with `409 still running — stop it first`, which is the PRE-EXISTING running check, not this guard — the restore had already finished by then. A refusal from the wrong rule is not evidence for the new one. | **CLOSED 2026-09-22 — v0.262.0 + v0.262.1; refusal and verification both proven live, and the live proof corrected the status code — **re-proven on v0.262.1: HTTP 409 with the household's sentence**, `RemoveStack privatebin REFUSED (busy)`** | +| **R-634** | **[P1-HIGH] An app can be RUNNING, HEALTHY and serving while the controller records it as not deployed — and in that state the household cannot remove it through the product at all.** MEASURED 2026-09-22 on guest 9202, controller v0.261.0, on **two independent apps in one night**: `outline` and `sparkyfitness`. **The controller's own words, in order.** `outline` deploy accepted 11:54:47; **`11:55:51 StopStack outline: current state=deploying deployed=true containers=0`** — a stop while the stack is still deploying; `11:56:09 SaveAppConfig: saving /opt/docker/stacks/outline — 5 env vars, **0 encrypted**, 3 sensitive fields` (the two saves before it both read `3 encrypted`); then **`11:57:16` and `12:02:16 Health probe outline: API GET :3000/_health -> 200`** — the app is up and answering its own health endpoint; and **`12:02:19 StopStack outline: current state=running deployed=false containers=3`**. Three containers, state `running`, health 200, **`deployed=false`**. The remove then answers **`RemoveStack outline: state=not_deployed, deployed=false, orphaned=false, deploying=false` -> `[ERROR] Remove failed for outline: stack "outline" is not deployed`** for BOTH the remove-with-data and the remove-keeping-data call, while `ScanStacks` goes on finding the stack every ten seconds. `app.yaml` survives with `desired_state: stopped` and an **empty `installed_images`**. `sparkyfitness` produced the identical shape 13 minutes earlier. **WHAT THIS COSTS A HOUSEHOLD:** an app that works is invisible to the product as an installation — no badge, no update, no backup selection, and **no way to delete it**; the only exit is a shell. It is the mirror of R-633 (there the record is gone and the container remains; here the container is fine and the record is gone) and it is the **worse** of the two, because the app is serving customer traffic the whole time. **THE MECHANISM IS NOT DIAGNOSED, and this row says so rather than guessing.** What was tried: the controller's full container log for both apps (the sequence above), `app.yaml` on disk, `docker ps -a`, and `GET /api/stacks/`. What was NOT done: reading `runComposeDeploy`'s pin-write path — this was an unattended run and the brief forbade product code. **The one discriminator worth running first:** both apps were walked while two other walks ran concurrently, and `POST /api/backup/run` is box-wide, so a backup or restore for a NEIGHBOURING app was in flight. A serial re-walk is queued tonight; if it reproduces alone, concurrency is not the cause. **A THIRD THING THE SAME LOG SHOWS, recorded here because it is one line away:** at `11:55:57` the health probe dialled **`http://outline-postgres:3000/_health`** — during startup, with the exactly-named container not yet running, `findProbeContainer`'s PREFIX fallback latched onto the POSTGRES sidecar and probed port 3000 on it. Transient, and it resolved once `outline` came up, but it is the same function R-630 is about. Evidence: `audits/the-28-2026-09-22/apps/half-state-outline-sparkyfitness.txt`. **THE SERIAL RE-WALK WAS RUN THE SAME NIGHT AND IT SPLITS THIS ROW IN TWO — recorded here rather than left as the first reading.** Walked again one at a time, with no other walk running: **`outline` deployed normally and removed clean**, and **`crafty-controller` deployed normally, updated `4.10.7 → 4.11.0` to `done`, restored and removed clean.** So for those two the half-state did NOT reproduce alone, and concurrency — a box-wide `POST /api/backup/run` or a restore in flight for a NEIGHBOURING app — is implicated rather than the deploy path itself. **`sparkyfitness` reproduced EXACTLY, alone, in 534 s**: deploy accepted, never reached `deployed` with a pin, `app.yaml` left with `desired_state: stopped` and an empty `installed_images`, and **both remove calls refused with `stack "sparkyfitness" is not deployed`.** **So the row stands, at one reproducible app instead of three, and the honest split is:** (a) `sparkyfitness` has a deploy that does not finish and leaves a record the product cannot clear — reproducible, P1; (b) under concurrent work the same unremovable half-state can be reached by apps that are otherwise fine, which is the more alarming half because those apps were **running, healthy and serving** while recorded as not deployed. **Neither half is diagnosed** — `runComposeDeploy`'s pin write was not read, because the brief forbade product code. **THE BOUNDED HALF IS FIXED in v0.262.0; the MECHANISM IS STILL NOT DIAGNOSED, and this row stays open for it.** `RemoveStack` refused on `!stack.Deployed` — a FLAG — while the machine plainly had containers, a compose file and an `app.yaml`. It now asks whether anything EXISTS (`halfStateEvidence`): containers, a compose file, or an `app.yaml` are each enough, and it removes what exists with every other fence unchanged (R-442 stays fail-closed). **The household must always be able to remove what the box shows them** — that is true whatever the cause of the bad record. Red-proof seen failing. **What is NOT done:** why `deployed` stays false while containers run. `runComposeDeploy`'s pin-write path was not read against a fresh reproduction in this session. **THE UNREMOVABLE HALF IS PROVEN LIVE ON 9202:** `sparkyfitness` rebuilt in its exact measured shape — `app.yaml` and a compose file on disk, no containers, `deployed: false`, `state: not_deployed` — and the remove answered **200** with `leftovers: NONE`. Under v0.261.0 the identical call answered `stack "sparkyfitness" is not deployed`. | **OPEN — P1 for the MECHANISM only; the unremovable half is CLOSED in v0.262.0 and proven live** | | **R-635** | **[P1-HIGH] `romm 5.3.0` does not fit the memory the template gives it, and the guarded Update called that a success — the app has been OOM-crash-looping on demo-hp for six hours at ~500% CPU.** FOUND 2026-09-22 17:37 **because the operator heard the fans**, which is the only reason it was found at all. `romm` was promoted `5.0.0 -> 5.3.0` on the live catalog that morning (`audits/PROBE-FIX-2026-09-22.md`) after the edge was PROVEN on scratch guest 9202, and the guarded Update was then pressed on demo-hp guest 9201 at **09:08:22Z**, reaching **`done` in 74.8 s** with the app `running`. **It ran clean for two hours.** The first worker kill is at **11:09:20Z**; by 15:38Z there had been **4,530** of them — `Worker (pid:…) was sent SIGKILL! Perhaps out of memory?` — with `docker inspect` reading **`OOMKilled: true`**, the container pinned at **457 MiB of its 512 MiB limit**, and `docker stats` showing **499.51% CPU**. The host's load average sat at **5.2 while otherwise idle**. `rq_cron` is killed and restarted every few seconds in a permanent storm. **THREE THINGS THIS ESTABLISHES, and the third is the one that changes how promotions are judged.** (1) The template's own comment says *`RAM: ~300MB (mem_limit: 1024M total — romm 512M + mariadb 384M + redis 128M)`*; **5.3.0 needs more than 512M and the template was not re-sized when the version moved.** A version move is not only an `image:` line. (2) **The update reported `done` and the app reads `running`**, because nginx answers `GET /` with 200 while the gunicorn workers behind it are being killed — a THIRD variant of the R-618/R-630 theme: the probe is right, the port is right, and the answer is still a false green. (3) **THE ALARM DID FIRE, AND MY FIRST WRITE-UP OF THIS ROW SAID IT DID NOT — CORRECTED 2026-09-22 BY THE OPERATOR, WHO PRODUCED THE MAILS.** The controller HAS an OOM detector (`main.go:821`, `[deadapp] romm: container romm was OOM-killed`, evaluated every 30 s), it emits **`app_oom`** (`notifier.go:726`), the hub allow-lists it (`dispatcher.go:636`) and dispatched it to the OPERATOR channel as **SENT** — `admin@felhom.eu` received **`[Felhom] ⚠️ demo-hp: app_oom`** at **11:09 CEST** and again at **17:48 CEST**, each carrying the app name, the Hungarian sentence and the dashboard link. The CUSTOMER channel is **SKIPPED**, which is correct. **I asserted an absence without looking at the instrument** — the hub's own Events and Notifications tabs show all of it — and I did it by reasoning from a memory note (`lxc-docker-oom-signals-unreliable`, R-528) instead of reading the hub. **That is R-628's shape exactly, four days on, and from the same hand.** **WHAT IS ACTUALLY WRONG, and it is narrower and real:** `notifier.go:715-726` keys the alarm on `container|startedAt` in an `oomSeen` map and emits **once per container lifetime**. So **six hours of continuous thrashing — 4,530 worker kills — produced exactly ONE mail**, severity `warning`, never escalating, while the app went on reading `running`. **A six-hour storm is indistinguishable from a single transient kill.** The hub's App Telemetry panel did carry the magnitude (RomM: **5,023 errors, 632 warnings**, peak 855 MB) but nothing turns that into a second, louder signal. So the fix worth having is not a detector — there is one — but an ESCALATION: a warning that repeats for hours should stop looking like a warning that happened once. **AND THE LESSON FOR R-462's METHOD, which is the real cost:** every `proven` verdict in the update night and in the twenty-eight measures the app for the **minutes of the walk**, not for a day of running. `romm` passed its walk, was seeded, read back and restored — and broke two hours later. **`proven` currently means "the update applied and the data survived", NOT "the new version runs".** Needs: decide between rolling the catalog back to 5.0.0 and raising romm's `mem_limit` (measured, not guessed); and a soak longer than a walk before any future promotion. **FIXED AND MEASURED 2026-09-22, in two steps, and the FIRST step was still a guess.** *Step 1 (operator's choice):* the limit was raised 512M → **768M** (`app-catalog-felhom.eu@886956d`). It slowed the kills from ~12/min to ~7/min and **stopped nothing** — 37 SIGKILLs in five minutes, `OOMKilled` still true, and the cgroup's own `memory.peak` read **exactly 768 MiB**: it hit the new ceiling and died there. *Step 2, from a MEASUREMENT instead:* the per-process RSS inside the container reads **~216 MiB per warm uvicorn worker**, so the image's default of four workers plus the master needs **~882 MiB** before nginx and the job runner — more than any sensible limit for this box. **The lever was in the image all along:** `/init:143` runs `--workers "${WEB_SERVER_CONCURRENCY:-4}"`. **Four workers is a SERVER default on an appliance serving one household.** Setting **`WEB_SERVER_CONCURRENCY=2`** (`app-catalog-felhom.eu@f4eb94f`, limit left at 768M) and applying it through the product's own Update button fixed it. **PROVEN UNDER LOAD, not just at idle:** 6 concurrent callers driven at romm through the household's own route for 300 s — **26,645 requests** (9,687 × 200, 16,958 × 401 on the auth-gated endpoints), CPU a steady **~200%** (exactly two workers saturated, by design), memory oscillating **416–614 MiB against the 768 MiB limit and trending DOWN**, and **zero** SIGKILLs, `OOMKilled: false`, `RestartCount: 0` throughout. At idle afterwards: **1.64% CPU**, 610 MiB, host load falling from 5.2 to 2.4. **AND THE FIRST SOAK MEASURED NOTHING, which is worth more than the second one:** it was pointed at `arcade.enkisfelhom.hu` — the scratch-guest fixture's default subdomain — while this box deployed romm at `jatek`. Every request 404'd at traefik in 9 ms, romm sat idle at 0.64% CPU, and the counter cheerfully reported **14,026 successful requests**. It was caught only because 0.64% CPU under load is not believable. **A positive AND a negative control are now asserted before any load is driven** (`jatek` must not 404; a nonsense host must). R-96 rule 3, in a new surface. **WHAT STAYS OPEN, and it is the part that outlives romm:** 610 MiB of 768 MiB is **79%** — it works with ~158 MiB of headroom and the soak never exceeded 614 MiB, but it is not generous, and nothing watches it. **And the method lesson for R-462:** a version move is not only an `image:` line — the new version's SHAPE (worker counts, per-worker footprint) has to be measured too, and a walk lasting minutes cannot see a ceiling reached in two hours. Every `proven` verdict in the update night and in the twenty-eight means *"the update applied and the data survived"*, **not** *"the new version runs"*. Evidence: `audits/probe-fix-2026-09-22/romm-soak.json`, `romm-soak.out`. | **CLOSED 2026-09-22 — two workers, 768M, proven under 26,645 requests; the 79% headroom and the `proven`-means-minutes lesson are carried into R-462** | | **R-636** | **[P2-MEDIUM] An app that has been out of memory for six hours sends the same single warning an app that hiccuped once sends.** FOUND 2026-09-22 when the operator produced the alarm mails I had wrongly written off as absent (R-635). **The detector is correct and works.** `main.go:821` re-checks every 30 s and logs `[deadapp] romm: container romm was OOM-killed`; `notifier.go:726` emits **`app_oom`** at severity `warning`; the hub allow-lists it (`dispatcher.go:636`) and delivered it to the OPERATOR channel — two mails, 11:09 and 17:48 CEST, each naming the app and linking the dashboard. CUSTOMER is SKIPPED, correctly. **The defect is the SHAPE of the signal, not its absence.** `notifier.go:715-724` keys on `container|startedAt` in an `oomSeen` map and returns early on a repeat, so the alarm fires **once per container lifetime**. romm's 09:08 container produced **4,530 worker kills over six hours and exactly one mail**; the 15:47 container produced one more. Severity never escalates, the app goes on reading `running`, and `08` §4 rightly does not put it in `IsDownState`. **So a six-hour storm that pins five cores is indistinguishable, in the operator's inbox, from one transient kill at 3 a.m.** — and the operator, who had two correct mails, still found the fault by hearing the fans. **The once-per-lifetime rule is right in itself** (it is what stops a crash loop from mailing 4,530 times, which is R-629's lesson) — what is missing is the second, louder signal when the same key keeps re-firing. **The magnitude IS already collected:** the hub's App Telemetry panel read **RomM: 5,023 errors, 632 warnings, peak 855 MB against a 1280M limit** while the same panel showed every other app at 0. Nothing turns that into an event. **Candidate shapes, none chosen here:** escalate `app_oom` to `error` when the same `container|startedAt` key re-fires past a threshold; or a periodic digest for a key still firing after N minutes; or let App Telemetry raise its own event when an app's error count crosses a bound. All are hub/controller code. Evidence: the operator's screenshots of the Events, Notifications and App Telemetry tabs plus the two mails; `felhom-controller/controller/internal/notify/notifier.go:715-726`. | **OPEN — P2; owner: CC; product code, so not fixed unattended** |