Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
9.8 KiB
REPORT — the design build: restore order, out-of-memory alarm, one stop per tier, the family window, wger's files (2026-10-06 evening)
| Part | Result |
|---|---|
| A — housekeeping | done — decisions 151–156 recorded first; R-469 closed by ruling; the agent repo has the shared rule file (identical: diff against the controller's copy empty, one md5 across five copies); the Tester 1 walk route is NOT a 30-minute fix → R-892 |
| B — restore order (R-638) | done — measured first on 9202: both main paths SAFE (docmost, romm); two side paths fixed by order, red-proved; exposure 3 not fixable by order → known limit in 07 §6.3 + R-893; R-638 closed |
| C — out-of-memory alarm (R-528) | STOPPED by its measurement — Docker 29.8.2 set OOMKilled correctly in all four shapes; nothing built; a decision for you |
| D — one stop per backup tier (R-518) | built and delivered, not shown live — measured first (read-only): the off-site part of a stop is ~2 s; red-proved; the live button test was not possible on 9202 (no agent connection) and the demo boxes take deliveries only → R-518 open for the first night read-back |
| E — reopen a sign-up for the family window (R-717) | done for wishlist, proven live; opengist stopped (no tool in its container) → R-717 narrowed |
| F — wger's file server (R-762) | done and proven on both venues; wger stays hidden — R-762 narrowed to the runserver half it carries from R-755 |
| Rows before | Rows after | Opened | Closed |
|---|---|---|---|
| 137 | 137 | 2 (R-892, R-893) | 2 (R-469, R-638) |
Baselines and rulings
Verified at the start: felhom.eu 01d5dbd3e0, controller 31242786af (v0.300.0), agent 3e8ebeb96c (v0.149.0), catalog
ecb8552ee9, register 137. Read: the three designs in audits/night-burndown-2026-10-05/, rows R-638, R-528, R-518, R-717,
R-762, R-469, 07 §6, 08 (the OOM rung), 09 §3 decisions 47/49. Rulings recorded first as 09 §3 151–156 (ruling 6 is
carried as A by the brief; the operator's chat answer was not visible to this session).
Part B — R-638
Slice 0, measurement only (scratch 9202, controller 0.299.0, drill catalog; audits/design-build-2026-10-06/B/).
Install the OLD definition → seed → the guarded Update migrates (its backing-up phase makes the copy at the old version) →
the household's restore of that copy (POST /backup/restore, the only copy offered: "helyi").
| app | migration | after restore | Imported DB dump |
seed | the copy's own dump (control) |
|---|---|---|---|---|---|
| docmost 0.95.0 → 0.96.0 (pg16) | 42 → 48 tables (+6 oauth_*, public_spaces, siem_destinations) |
42, none of the 6 | yes | read back | 42, equal |
| romm 5.0.0 → 5.3.0 (mariadb 11.4) | 27 → 39 (+12) | 27, none of the 12 | yes | read back | 27, equal (views counted) |
Both went back to the old pinned version and ran. Said plainly: romm needed four runs — run 1 refused the install
(its data path must be a drive), run 2 hit an empty drill commit, run 3 measured correctly but my control read the wrong
file (a pre-restore-* dump) and missed three VIEWS; run 4 fixed both. docmost's run 1 installed the live version
because the box had not read the drill yet (the script then waited for the old reference).
Slices 1 and 2 (controller 9c94568): RestoreApp starts only the database services, replays, then the app; a dump
with no identifiable database service is refused before anything is touched; a failed volume step skips the replay
(fallback and unit restore). Red: B/red-slice1-fallback-order.txt ("the FULL stack was already up when the replay fired
(calls before the replay: stop,start)"), B/red-slice2-no-replay-after-volume-failure.txt. Exposure 3 (the off-site
rollback loads the newer pre-restore dump over the older volume; the older definition is not put back) cannot be fixed by
order — known limit in 07 §6.3, row R-893.
Part C — R-528, stopped
C/C0-oom-signal-measure-9202.txt: four runs in a throwaway alpine container under a cap. Runs 1–3: the kernel killed the
hog AND the main process (the box cannot protect it: --oom-score-adj -1000 left pid 1 at 0) — exit 137,
OOMKilled=true. Run 4 (256 MB, sleep infinity, one 400 MB block): the container kept running, oom_kill 0 → 3,
OOMKilled=true, three oom events. The design's premise — the flag stays false while the counter rises — did not
reproduce on Docker 29.8.2 (the false flags of 2026-09-15 were on 29.8.0). By the brief's measure-first rule, nothing was
built. Cost read for the record: one exec read of demo-hp's 21 containers 1.6 s wall (C/C1-…). 08 records it.
Part D — R-518
Measured first, read-only: demo-felhom's night off-site job (2026-10-06): started 06:21:08, snapshotted 06:21:10,
its one app running at 06:21:18. demo-hp had no successful off-site run in 10 days. A new off-site run from 9202 was not
possible (it would write to ep0). Built (helper, then corrected by me): one window per tier, resume at that tier's
snapshotted, the next tier in a later cycle; the button makes the LOCAL copy only. My correction: the helper let the
button back up the off-site tier when the local storage is absent; the ruling says local only, so it now backs up
nothing and stops nothing (red: D/red-press-no-local-backs-up-nothing.txt); an agent that flags no tier primary keeps
the old first tier (the breaker test's fake agent showed that case). Red: D/red-first-tier-resume.txt (restarts
sampled at each poll [0 0 0 0]), D/red-manual-local-only.txt. Page text, hu + en: about 1–1.5 minutes — an estimate
from measured parts, not a measured press. Not shown live (see the Part table).
Part E — R-717
Controller (5b86569): after_setup open_command / open_success; the window marks opening on disk first; a failed
open closes again and the page says so; the close runs at the window's end, at every start and after a successful
update; a failed close retries every 2 min. Six red-proofs in E/. Catalog (a5a5b51, pushed AFTER the controller
reached the three boxes): wishlist closes/opens system_config.enableSignup (group global) with Node's own sqlite
module, min_controller: "0.301.0". Live on 9202 (the 0.301.0 test image, E/live.txt): after setup false, a
stranger straight at the app 401 "invite only", users 1; window true, a family member 200, users 2; window end 13:47:34Z,
close 13:47:44Z, false, a stranger 401, users 2. Opengist: no sqlite tool, no script runtime, and its CLI has only
create-user, reset-password, toggle-admin — not doable with this mechanism.
Part F — R-762
Upstream's production compose read first (its nginx serves /static/ and /media/ from the shared volumes;
prod.env sets DJANGO_DEBUG=False). Here traefik sends only those two paths to a wger-files nginx
(nginx:1.30.5-alpine, 32 MB, stock config, volumes read-only) — every box gate wraps every router, so it is gated
like wger. The bench tool could not express a step that adds a service, so I added --move-to and
--write-ladder --to-definition (catalog c48a0db, red-proved). Bench: proven, wger anon peak 50.3 %, wger-files
15.9 %, the page's hashed CSS 200 from wger-files and 404 straight at wger. 9202: the guarded Update done in 59.4 s,
the seed read back, a photo posted before the step 404 → 200 / 178 B / image/png after, CSS 404 → 200, wger-files
7.6 MiB. Said plainly: my box script crashed after the step when it decoded the PNG as text; the after-checks were
re-read binary-safe and the box verdict was written by hand from them (it says so). Ready to show? Files and sign-up
yes; it still runs Django's runserver (upstream's gunicorn runs 3 workers, which do not fit 384 MB) — not ready. Cost:
283 MB of collected static files in a named volume, in every wger backup (the checks refuse an unnamed volume).
Release and delivery
Controller v0.301.0 (0b1b8d3, image sha256:0e80f6af…), MinAgent 0.131.0. Golden 0.301.0
(GOLDEN_SHA256=96e94fed…, Docker pinned to the approved set, registry round trip equal, token grep 0 with a working
control). Hub: vouched agent 0.149.0 / golden 0.301.0 / min_agent 0.131.0; floors 0.301.0 for demo-hp, demo-felhom,
tester-1 — first saved without a declared MinAgent (allowed: floor = golden), then re-saved with 0.131.0. demo-hp and
demo-felhom run 0.301.0 (healthy); tester-1's hub page "Controller elindult (0.301.0)". Global floor and Tester 2 not
touched. No agent release.
Instruction-file edits (decision 150)
.claude/rules/unprompted-work.md, all five copies, line 8–9: before „…infelhom.eu,felhom-controllerandapp-catalog-felhom.eu.claude/rules/, and in the workspace root's unversioned.claude/rules/; change all four or none." → after: namesfelhom-agenttoo, „change all five or none". Why: decision 152 added the agent's copy.felhom-agent/.claude/rules/unprompted-work.md: new file (decision 152), identical to the other copies.
Fixed without a row
OPEN-ITEMS.mdsection headings: every count was stale (e.g. „Install & onboarding — 11 rows" over 5); recomputed.
CI, last commit of every repo
controller 0b1b8d3 → 1439 success; catalog cf1ed43 → 1438 success (a5a5b51 checked after its push); agent
de812bc (checked after its push); felhom.eu 8db4425 → 1440 (checked), and this commit (checked after the push).
Teardown
Machine: 9202 back on the live catalog (controller.yaml byte-identical), on the released 0.301.0 (was 0.299.0), every
test app removed through the product (no volume left), the test images removed by name, the same six containers as at
the start. Bench 9401 stopped. Drill VM: build guest destroyed, files shredded, off, virgin. Drill catalog = live.
Host: nothing else. Hub: the vouch and three floors; two read-only DB copies, deleted.