Files
felhom.eu/documentation/audits/cleanup-2026-09-23

Clean-up evening — R-634's cause, held apps, the OOM storm (controller v0.265.0, hub v0.121.0), 2026-09-23

Method: endpoint-level on scratch guest 9202 (the endpoints the UI invokes; pages as HTML with ?lang=hu|en); docker events + the controller log recorded inside the guest from before the first call (setsid docker events/logs -f to a file — never piped through a buffering command). No browser. Harness copied from undo-fleet-2026-09-23/; walk.py backup_now presses nothing now (R-648).

Part 1 — R-634, diagnosed before any code

# what result evidence
1a sparkyfitness alone, live pin, v0.264.0 did NOT reproduce: deployed in 47.3 s, running 10-*, 11-*
1b outline + whole-box backup at +20 s, v0.264.0 reproduced: StopStack outline: state=deploying deployed=true containers=0 (14:39:32) → Restarting outline after volume dump (14:39:33) → both compose up fail at 14:39:48 → deployed=false, containers Created 12-*, 13-*
fix, try 1 same, v0.265.0 not a race — images cached, deploy done in 6.3 s before the backup came; kept as a record, not as proof 30-fix-attempt1-*
fix, try 2 outline image removed by name, backup at +5 s, v0.265.0 the backup ran 15:23:07–15:23:31 across the deploy, stopped gokapi, paperless-ngx, privatebin and never outline; deploy successfully (took 48.9s), deployed=true, running 32-*, 33-*

Mechanism at file:line: stacks/deploy.go:395 (in-memory Deployed=true at accept) → cmd/controller/main.go:2556 ListDeployedStacks (flag only) → backup/backup.go:693 runVolumeDumps → DumpAppVolumesSafe StopStack + StartStack (backup.go:908/918) → stacks/deploy.go:421-435 failure branch writes Deployed=false. There is no deploy time limit (the brief's shape (i) is ruled out).

Parts 2–4 (live, 9202, drill catalog: romm 320M/4 workers, vikunja at 2.3.0)

proof result evidence
R-625 badge, 4 box/reader pairs hu reader: „Megállítva — visszaállítás szükséges", title „A frissítés nem sikerült, és az automatikus visszaállítás sem."; en reader: "Stopped — restore needed", title "The update did not succeed, and the automatic undo did not either."; no „Frissítés elérhető"/"Update available"; no vikunja Update button on the list while other apps keep theirs (control); API POST …/update → 409 held 44-*
R-647 (1) box en, ?lang=hu: the hold in Hungarian on the app page, and update_error Hungarian in the API 43-*, 44-*
R-648 vikunja's update ran its own backing-up phase; no whole-box press 42-*, 43-*
R-636 romm: kills 8 → 13 → 21 → OOM STORM — 21 kills in 30 min (limit 320M, peak 320M) + DROPPED event app_oom_storm (severity error) (R-620 witness, 9202 has no hub); still ONE storm at 49 kills 45-*, 46-*

Hub v0.121.0 deployed first (20-*). The hub side of R-636 is proven by unit tests only: posting a synthetic event with a real customer's key would put a fabricated alarm into the operator's record.

Floor

0.265.0 / MinAgent 0.131.0, read back; demo-hp and demo-felhom both on 0.265.0 within 20 s (50-*).

Teardown — three layers

  • machine (9202): outline, sparkyfitness, vikunja, romm removed through the product (romm's with-data remove refused 409 on the drive path, R-442 as in every drill — its drive folder removed by name); 0 volumes, 0 undo copies; test images removed by name only where no container used them (postgres:16-alpine kept — in use); controller.yaml identical to .pre-cleanup; catalog on live cfcfe52; language hu; recorders stopped and /root/r634 removed; standing apps as at the start.
  • host: nothing provisioned on demo-hp. 9201 not touched (no backup press; it reached 0.265.0 by the floor).
  • hub: v0.121.0 (planned); floor 0.265.0.
  • DooPlex: one empty volume outline_outline_data and one alpine:3.20 image were created by MY first test draft / my inspection of it; both verified unused and removed by name (R-650).
  • drill repo: reset to live main, read back from the remote.