# ADDENDUM (same session) — R-649, the operator's ruling: a failed install removes what it started (controller v0.266.0) Ruling: "On failed install, the controller should clean up the containers." Built in `964ae75`: `runComposeDeploy`'s failure branch runs `compose down` (no `-v` — volumes kept) before recording not-deployed. Red-proof: the `down` removed → `TestR649_AFailedInstallRemovesWhatItStarted` fails with `calls=["up -d"]`. Live on 9202 (`audits/r649-2026-09-23/`): outline with a never-healthy redis (drill template only) failed after 102.7 s → 0 containers left, 3 volumes kept, log line `the failed deploy's containers were removed (volumes kept) — R-649`. Floor **0.266.0** read back; both demo boxes on it in 20 s. R-649 closed; **open rows 331 → 330.** Teardown: outline removed, images by name, config identical to the saved copy, live catalog, drill reset. Nothing else touched. --- # REPORT — clean-up evening: R-634's cause, held apps, the OOM storm (controller v0.265.0, hub v0.121.0) 2026-09-23 (late evening). Baselines verified live: controller `0a3026180ae8` (v0.264.0), agent `d9864a94bf62`, felhom.eu `21f17ed32bdc` (hub v0.120.0), catalog `cfcfe5278428` — as the brief said. Repos touched: **felhom-controller** (v0.265.0 `0054d4b`), **felhom.eu** (hub v0.121.0, manifest, docs, register, evidence), **admin/app-catalog-drill** (drill commits, reset to live `main`). **felhom-agent, app-catalog-felhom.eu: untouched.** Architecture read and named: `09-update-architecture.md` §6.1/§6.1a/§6.4, `08-alarm-ladder.md` §4/§6.2, `02-controller-module-map.md` (stacks ↔ backup seam). Evidence: `documentation/audits/cleanup-2026-09-23/README.md`. **Method: endpoint-level** (the endpoints the UI invokes; no browser). --- ## 1. Not done, or changed | item | state | |---|---| | Part 1 — R-634 | **cause found in ~25 min (cap 2 h), fixed, proven live.** One design question left → **R-649** (operator) | | Part 2 — R-625 | **done.** The Update button was ALREADY hidden for a held app on the list (since v0.238.0); the defect was the badge only. Now pinned by a test with a positive control. The badge applies to **both** hold kinds (a failed restore holds the app stopped too; the way back is also a restore) — the brief said `update_failed` only | | Part 3 — R-636 | **done, and the brief's rule changed shape:** "the same key re-fires ≥ 20 times" cannot work — `OOMKilled` is sticky, so the key re-fires every 30 s after ONE kill (a hiccup would storm after 10 min). The controller now counts real kills (kernel `oom_kill`); the 20-in-30-min threshold is applied to THAT | | Part 3 — hub live `curl` | **not done:** posting a synthetic event with a real customer's key would put a fabricated alarm in the operator's record; the hub side is proven by unit tests | | Part 4 — R-647 | **done** | | Part 5 — R-648 | **done — no per-app endpoint exists**, so the drill presses nothing | | Part 6 — floor | **raised to 0.265.0** — every live proof passed | | a mistake of mine | my first draft of the R-634 backup test drove the real dump and **created an empty volume `outline_outline_data` on DooPlex**; my inspection of it pulled `alpine:3.20`. Both verified new and unused, removed by name. Tests rewritten on seams. **R-650** filed (the class is unguarded) | **Claims in the brief, checked:** 1. *"sparkyfitness reproduces alone at today's pin"* — **wrong.** Deployed alone in 47.3 s on v0.264.0 (`10-*`). The 2026-09-22 "alone" walk pressed the whole-box backup itself — the same race. 2. *"a whole-box backup touches a DEPLOYING stack at all (never read)"* — **true, and it is THE cause.** 3. *"20 kills in 30 minutes separates a storm from a hiccup"* — **true only if kills are counted, and they were not.** RomM's real rate: 4,530 kills in 6 h ≈ **375 per 30 min**; a hiccup is 1–3. 20 kept, applied to the kernel's counter. Live: 21 kills 2.5 min after start → one storm. 4. *"a per-app backup endpoint exists"* — **wrong.** `router.go` has `/backup/run` and `/backup/tier2` only; the per-app backup (`RunAppBackupNow`) is reachable only through the guarded update. ## 2. R-634 — the mechanism, at `file:line` 1. `stacks/deploy.go:395` — `DeployStack` sets the in-memory `Deployed=true` at ACCEPT (no-stale-button UX). 2. `cmd/controller/main.go:2556` — `ListDeployedStacks` filtered on that flag only → a deploying app is in every backup leg's list. 3. `backup/backup.go:693` → `DumpAppVolumesSafe` — `StopStack` (`compose down`) in the middle of the deploy, tar of half-made volumes, `StartStack` (a SECOND `compose up -d`, `backup.go:918`). 4. Both `up` calls fail together; `stacks/deploy.go:421-435` writes `Deployed=false` without asking whether containers exist. Which `up` wins decides the end: running under „not deployed" (2026-09-22) or `Created` (tonight, 14:39:48). **No deploy time limit exists** — shape (i) ruled out. **Fix:** deploying apps leave the list; the volume leg re-asks right before the stop (SKIP, never FAIL); `StopStack`/`StartStack` refuse a deploying stack (`ErrStackDeploying`) — the backstop for quiesce, restore, export, storage and the Stop button, which all had the same exposure. **Live (`32-*`, `33-*`):** outline's image removed so the deploy must pull, backup at +5 s: the backup ran 15:23:07–15:23:31, stopped gokapi, paperless-ngx and privatebin, **never outline**; `deployed successfully (took 48.9s)`. The first fix run is kept and labelled NOT A RACE (images cached, 6.3 s). ## 3. Red-proofs (each seen failing, mutation restored, `=== RUN` checked) | # | mutation | test → failure | |---|---|---| | 1 | backup skip removed | `TestR634_VolumeLegNeverStopsADeployingApp` — `dumped=[outline]` | | 2 | StopStack guard removed | `TestR634_StopAndStartRefuseADeployingStack` — `exit code -1` instead of `ErrStackDeploying` | | 3 | held badge removed (v0.264.0's shape) | `TestR625_AHeldAppSaysStoppedRestoreNeeded` — badge missing, „Frissítés elérhető" present | | 4 | reader funcs off | `TestR647_AReaderInTheOtherLanguageReadsTheHoldInTheirs` — English hold for the Hungarian reader | | 5 | Hungarian phrase wired back | `TestR647_HeldEventCarriesTheCopyHoldsKey` | | 6 | health status passed as severity | `TestR647_DisabledHealthChangeNamesTheSeverity` — `severity warn` | | 7 | escalation removed | `TestR636_TwentyKillsInThirtyMinutesIsOneStorm` — `kills=20 … storms=0, want 1` | | 8 | counter never read | `TestR636_ScanReadsTheCounterOfFlaggedContainersOnly` — `Kills:-1` | | 9 | hub: not operator-only | `TestAppOOMStormIsAllowlistedAndOperatorOnly` — "must be operator-only" | | 10 | hub: not allow-listed | same test — "must be in allowedEventTypes" | | 11 | hub: no per-app cooldown | `TestR389_TheAllowListHasExactlyOneMember` | Also pinned without a mutation run: 19 kills → no storm, 200 → one, one hiccup over 6 h → none, 20 kills over 3 h → none, an unreadable counter → never. Gates: controller `go test ./...` rc=0 + `controller_gates.py` rc=0; hub `go test ./...` rc=0 + `repo_gates.py` rc=0. Parity: `app_info_deployed` and `stacks_full` regenerated — measured diff **exactly one line each, the new held badge**. ## 4. Live proofs in both languages (9202) - **R-625 + R-647 (1)** (`44-*`), every box/reader pair: hu reader „Megállítva — visszaállítás szükséges", title „A frissítés nem sikerült, és az automatikus visszaállítás sem."; en reader "Stopped — restore needed", title "The update did not succeed, and the automatic undo did not either."; with the box in ENGLISH the Hungarian reader reads Hungarian (page and API). No vikunja Update button; other apps keep theirs (control). `POST …/update` → **409 `held`**. - **R-636** (`45-*`, `46-*`): romm at 320M — kills 8, 13, **21 → `OOM STORM — 21 kills in 30 min (limit 320M, peak 320M)`** + `DROPPED event app_oom_storm (severity error)`; one storm line at 49 kills. - **R-648** (`43-*`): vikunja's update ran its own `backing-up` phase. ## 5. Rows **Closed (5):** R-634, R-625, R-636, R-647, R-648. **Opened (2):** R-649 (operator question — a deploy that fails on its own), R-650 (tests that reach real docker act on DooPlex). **Open rows 334 → 331.** `09` §6.4 parts 8 and 9 marked SHIPPED; `08` §6.2 has the storm rung and the two per-app grains. ## 6. Floor 0.265.0 / MinAgent 0.131.0, read back from the hub (`min_controller_version=0.265.0`, `min_agent=0.131.0`; `managed floor SERVED` for demo-felhom and demo-hp); both on 0.265.0 in 20 s. ## 7. Teardown — three layers - **machine (9202):** all four test apps removed through the product; romm's drive folder by name (R-442 fence, as always); 0 volumes, 0 undo copies; test images removed by name where unused; config identical to the saved copy; live catalog; language `hu`; recorders stopped. 9201 not touched. - **host:** nothing provisioned. **hub:** v0.121.0 + floor (planned). **DooPlex:** the volume and image above, removed. **drill repo:** reset to live `main`. Password files deleted from the scratchpad.