# REPORT — the design build: restore order, out-of-memory alarm, one stop per tier, the family window, wger's files (2026-10-06 evening) | Part | Result | |---|---| | **A** — housekeeping | **done** — decisions 151–156 recorded first; R-469 closed by ruling; the agent repo has the shared rule file (identical: `diff` against the controller's copy empty, one md5 across five copies); the Tester 1 walk route is NOT a 30-minute fix → R-892 | | **B** — restore order (R-638) | **done** — measured first on 9202: both main paths SAFE (docmost, romm); two side paths fixed by order, red-proved; exposure 3 not fixable by order → known limit in `07` §6.3 + R-893; R-638 closed | | **C** — out-of-memory alarm (R-528) | **STOPPED by its measurement** — Docker 29.8.2 set `OOMKilled` correctly in all four shapes; nothing built; a decision for you | | **D** — one stop per backup tier (R-518) | **built and delivered, not shown live** — measured first (read-only): the off-site part of a stop is ~2 s; red-proved; the live button test was not possible on 9202 (no agent connection) and the demo boxes take deliveries only → R-518 open for the first night read-back | | **E** — reopen a sign-up for the family window (R-717) | **done for wishlist, proven live; opengist stopped** (no tool in its container) → R-717 narrowed | | **F** — wger's file server (R-762) | **done and proven on both venues; wger stays hidden** — R-762 narrowed to the `runserver` half it carries from R-755 | | Rows before | Rows after | Opened | Closed | |---|---|---|---| | **137** | **137** | **2** (R-892, R-893) | **2** (R-469, R-638) | ## Baselines and rulings Verified at the start: felhom.eu `01d5dbd3e0`, controller `31242786af` (v0.300.0), agent `3e8ebeb96c` (v0.149.0), catalog `ecb8552ee9`, register 137. Read: the three designs in `audits/night-burndown-2026-10-05/`, rows R-638, R-528, R-518, R-717, R-762, R-469, `07` §6, `08` (the OOM rung), `09` §3 decisions 47/49. Rulings recorded first as `09` §3 151–156 (ruling 6 is carried as A by the brief; the operator's chat answer was not visible to this session). ## Part B — R-638 **Slice 0, measurement only** (scratch 9202, controller 0.299.0, drill catalog; `audits/design-build-2026-10-06/B/`). Install the OLD definition → seed → the guarded Update migrates (its backing-up phase makes the copy at the old version) → the household's restore of that copy (`POST /backup/restore`, the only copy offered: "helyi"). | app | migration | after restore | `Imported DB dump` | seed | the copy's own dump (control) | |---|---|---|---|---|---| | docmost 0.95.0 → 0.96.0 (pg16) | 42 → 48 tables (+6 `oauth_*`, `public_spaces`, `siem_destinations`) | 42, none of the 6 | yes | read back | 42, equal | | romm 5.0.0 → 5.3.0 (mariadb 11.4) | 27 → 39 (+12) | 27, none of the 12 | yes | read back | 27, equal (views counted) | Both went back to the old pinned version and ran. **Said plainly:** romm needed four runs — run 1 refused the install (its data path must be a drive), run 2 hit an empty drill commit, run 3 measured correctly but my control read the wrong file (a `pre-restore-*` dump) and missed three VIEWS; run 4 fixed both. docmost's run 1 installed the live version because the box had not read the drill yet (the script then waited for the old reference). **Slices 1 and 2** (controller `9c94568`): `RestoreApp` starts only the database services, replays, then the app; a dump with no identifiable database service is refused before anything is touched; a failed volume step skips the replay (fallback and unit restore). Red: `B/red-slice1-fallback-order.txt` ("the FULL stack was already up when the replay fired (calls before the replay: stop,start)"), `B/red-slice2-no-replay-after-volume-failure.txt`. **Exposure 3** (the off-site rollback loads the newer pre-restore dump over the older volume; the older definition is not put back) cannot be fixed by order — known limit in `07` §6.3, row R-893. ## Part C — R-528, stopped `C/C0-oom-signal-measure-9202.txt`: four runs in a throwaway alpine container under a cap. Runs 1–3: the kernel killed the hog AND the main process (the box cannot protect it: `--oom-score-adj -1000` left pid 1 at 0) — exit 137, `OOMKilled=true`. Run 4 (256 MB, `sleep infinity`, one 400 MB block): the container kept running, `oom_kill` 0 → 3, `OOMKilled=true`, three `oom` events. The design's premise — the flag stays false while the counter rises — did not reproduce on Docker 29.8.2 (the false flags of 2026-09-15 were on 29.8.0). By the brief's measure-first rule, nothing was built. Cost read for the record: one exec read of demo-hp's 21 containers 1.6 s wall (`C/C1-…`). `08` records it. ## Part D — R-518 **Measured first, read-only:** demo-felhom's night off-site job (2026-10-06): started 06:21:08, `snapshotted` 06:21:10, its one app running at 06:21:18. demo-hp had no successful off-site run in 10 days. A new off-site run from 9202 was not possible (it would write to ep0). **Built** (helper, then corrected by me): one window per tier, resume at that tier's `snapshotted`, the next tier in a later cycle; the button makes the LOCAL copy only. **My correction:** the helper let the button back up the off-site tier when the local storage is absent; the ruling says local only, so it now backs up nothing and stops nothing (red: `D/red-press-no-local-backs-up-nothing.txt`); an agent that flags no tier primary keeps the old first tier (the breaker test's fake agent showed that case). Red: `D/red-first-tier-resume.txt` (restarts sampled at each poll `[0 0 0 0]`), `D/red-manual-local-only.txt`. Page text, hu + en: about 1–1.5 minutes — an estimate from measured parts, not a measured press. **Not shown live** (see the Part table). ## Part E — R-717 Controller (`5b86569`): `after_setup` `open_command` / `open_success`; the window marks `opening` on disk first; a failed open closes again and the page says so; the close runs at the window's end, at every start and after a successful update; a failed close retries every 2 min. Six red-proofs in `E/`. Catalog (`a5a5b51`, pushed AFTER the controller reached the three boxes): wishlist closes/opens `system_config.enableSignup` (group `global`) with Node's own sqlite module, `min_controller: "0.301.0"`. **Live on 9202** (the 0.301.0 test image, `E/live.txt`): after setup `false`, a stranger straight at the app 401 "invite only", users 1; window `true`, a family member 200, users 2; window end 13:47:34Z, close 13:47:44Z, `false`, a stranger 401, users 2. **Opengist:** no sqlite tool, no script runtime, and its CLI has only `create-user`, `reset-password`, `toggle-admin` — not doable with this mechanism. ## Part F — R-762 Upstream's production compose read first (its `nginx` serves `/static/` and `/media/` from the shared volumes; `prod.env` sets `DJANGO_DEBUG=False`). Here traefik sends only those two paths to a `wger-files` nginx (`nginx:1.30.5-alpine`, 32 MB, stock config, volumes read-only) — every box gate wraps every router, so it is gated like wger. The bench tool could not express a step that adds a service, so I added `--move-to` and `--write-ladder --to-definition` (catalog `c48a0db`, red-proved). **Bench:** proven, wger anon peak 50.3 %, wger-files 15.9 %, the page's hashed CSS 200 from wger-files and 404 straight at wger. **9202:** the guarded Update done in 59.4 s, the seed read back, a photo posted before the step 404 → 200 / 178 B / image/png after, CSS 404 → 200, wger-files 7.6 MiB. **Said plainly:** my box script crashed after the step when it decoded the PNG as text; the after-checks were re-read binary-safe and the box verdict was written by hand from them (it says so). **Ready to show?** Files and sign-up yes; it still runs Django's `runserver` (upstream's gunicorn runs 3 workers, which do not fit 384 MB) — not ready. Cost: 283 MB of collected static files in a named volume, in every wger backup (the checks refuse an unnamed volume). ## Release and delivery Controller **v0.301.0** (`0b1b8d3`, image `sha256:0e80f6af…`), MinAgent 0.131.0. Golden **0.301.0** (`GOLDEN_SHA256=96e94fed…`, Docker pinned to the approved set, registry round trip equal, token grep 0 with a working control). Hub: vouched agent 0.149.0 / golden 0.301.0 / min_agent 0.131.0; floors 0.301.0 for demo-hp, demo-felhom, tester-1 — first saved without a declared MinAgent (allowed: floor = golden), then re-saved with 0.131.0. demo-hp and demo-felhom run 0.301.0 (healthy); tester-1's hub page "Controller elindult (0.301.0)". Global floor and Tester 2 not touched. No agent release. ## Instruction-file edits (decision 150) - `.claude/rules/unprompted-work.md`, all five copies, line 8–9: before „…in `felhom.eu`, `felhom-controller` and `app-catalog-felhom.eu` `.claude/rules/`, and in the workspace root's unversioned `.claude/rules/`; change all four or none." → after: names `felhom-agent` too, „change all five or none". Why: decision 152 added the agent's copy. - `felhom-agent/.claude/rules/unprompted-work.md`: new file (decision 152), identical to the other copies. ## Fixed without a row - `OPEN-ITEMS.md` section headings: every count was stale (e.g. „Install & onboarding — 11 rows" over 5); recomputed. ## CI, last commit of every repo controller `0b1b8d3` → 1439 success; catalog `cf1ed43` → 1438 success (`a5a5b51` checked after its push); agent `de812bc` (checked after its push); felhom.eu `8db4425` → 1440 (checked), and this commit (checked after the push). ## Teardown Machine: 9202 back on the live catalog (`controller.yaml` byte-identical), on the released 0.301.0 (was 0.299.0), every test app removed through the product (no volume left), the test images removed by name, the same six containers as at the start. Bench 9401 stopped. Drill VM: build guest destroyed, files shredded, off, `virgin`. Drill catalog = live. Host: nothing else. Hub: the vouch and three floors; two read-only DB copies, deleted.