3e58c184f6
gates / gates (push) Successful in 27s
DRILL-night-2026-09-23.md: Parts A-E. 09 §3 decisions 21 (operator word), 22 and 23 (CC unattended, operator may reverse); §6.4 parts 4 and 6 (catalog half) shipped; §6.1a residuals R-658/R-659. Register 330 -> 336: R-651..R-660 opened (R-658 and R-659 P1), R-650/R-640/R-499/R-626 closed. Capability map, nightly rotation (opengist), STATUS (one question: the floor), CONTEXT, REPORT. The floor stays 0.266.0. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
42 lines
7.8 KiB
Markdown
42 lines
7.8 KiB
Markdown
# PROGRESS — night 2026-09-23 (one line per finished step, written when it finishes)
|
||
|
||
| time (CEST; rows before 20:53 are approximate) | step | result | evidence |
|
||
|---|---|---|---|
|
||
| 19:16 | P0 baselines | controller `15e630aa02ed` v0.266.0 · agent `d9864a94bf62` v0.132.0 · felhom.eu `e443841b752a` · catalog `cfcfe5278428` — all match the brief; trees clean; disk 37 % / 53 % | this line |
|
||
| ~19:40 | P0 drift re-run | 66 pins; **26 within-a-major edges**, 11 across-a-major (listed, not moved), 1 error (plant-it, abandoned) | 02-drift.json, 02-drift.log |
|
||
| 19:20 | P0 capacity | demo-hp 29 GB RAM (23 free), nvme-scratch 7 %; 9202 25.9 GB, /var/lib/docker 64 % (24 GB free); 9201 standing apps listed | 00-demo-hp-before.txt, 01-capacity-9202-9201.txt |
|
||
| 20:00 | A1 R-650 | `internal/dockerexec` guard + repo-wide sweep test; 8 api tests + web/stacks fixtures were reaching real docker → package TestMain stub; red-proofs both FAIL as expected | A1-r650-redproof.txt, A1-r650-sweep.txt |
|
||
| 20:05 | A2 R-640 | completion-marker check on all three restore paths, before the first mutation; 3 red-proofs FAIL as expected | A2-r640-redproof.txt |
|
||
| 20:17 | A3 R-626 | **NOT reproduced** on v0.266.0: remove → 390 s of events (destroy seen, no create) + controller restart + guest reboot → 0 containers, 0 volumes. Leftover applied-compose.yml/applied-meta (row) | A3-r626-*.txt/json |
|
||
| 20:15 | A4 R-499 | 4-branch system-disk sentence; 2 red-proofs FAIL as expected; parity fixtures per branch | A4-r499-redproof.txt |
|
||
| 20:20 | A5 R-518 | the "few seconds" promise was ALREADY gone (v0.243.0) — brief claim wrong; copy now states the measured ~8 min; red-proof FAILs as expected | A5-r518-redproof.txt |
|
||
| 20:26 | A6 v0.267.0 | pushed `80e6ad8c4772`, gates + full suite rc 0; image built; deployed to **9202 only** (Up, healthy). Floor waits for Part E | A6-deploy-0.267.0-on-9202.txt |
|
||
| 20:28 | A7 pages live | 9202 (no agent): R-499 renders the `unknown` branch in hu+en, matching `/api/storage/backup-target` `known:false`; R-518 text on /backups in hu+en | A7-pages-live-0.267.0.txt |
|
||
| 20:28 | B1 SPIKE | **PASS on v0.266.0 AND v0.267.0**: drill navidrome with a top-level `update_ladder:` (JSON flow lines) synced, deploy-fields read, deployed `running`, probe `API GET :4533/ping → 200`, badge „Naprakész"/"Up to date", no yaml errors. Source agrees: no `KnownFields` anywhere | B1-spike-v0.26{6,7}.0.{json,log} |
|
||
| 20:55 | B2 gate | `check-test-record.py` (static, CI) + `check-test-record-move.py` (history + registry for moved refs); 16 decoys both ways; 3 red-proofs FAIL as expected | B2-gate-redproofs.txt |
|
||
| 21:05 | B3 writer | `upgrade-test.py --write-ladder` (+ `--move`, harness v3: box fixtures on the bench, files_may_change); `test_ladder_writer.py` 5/5 | catalog `scripts/` |
|
||
| 21:10 | B4 backfill | 21 of 21 proven from their cited records (nextcloud's commit cited none — its record named, entry says so); romm carries M1's 80.9 % memory_tight | catalog `6db08a5` |
|
||
| 21:12 | B5 pushed | catalog `6db08a5`, hook gates OK, **CI job 920 success** (matched on head_sha) | — |
|
||
| 20:58 | C0 venue | bench LXC **9401** on demo-hp (nvme-scratch, 6c/10G), docker 26.1.5, Docker Hub login; **C3 negative control → failed** (bench measures) | C0-harness-venue-create.txt |
|
||
| 21:06 | C R-612 | bench: 128M → seed OOM-killed (oom_kill 3), 1024M peak 345M, 512M peak 312M seed completes; 9202 before: seed SIGKILLed but after the rows (sign-up worked — lie NOT reproduced, timing-dependent); after 512M: seed completes. Live catalog `a5a729a` | apps/wishlist/r612-*.json |
|
||
| 21:15 | C R-613 | UPTIME_KUMA_DB_TYPE=sqlite: before `setup-database` while box says running; after `entryPage`, /metrics 401; db in the volume. First "after" run used a stale template (R-607 lag) — kept, renamed. Live catalog `a5a729a` | apps/uptime-kuma/r613-*.json |
|
||
| 20:53 | C queues | bench (9401) and box (9202) queues started, database apps first; each edge's evidence copied off the bench when it ends | benchq-*.txt, boxq-*.txt |
|
||
| 21:18 | C fix | bench load generator followed redirects (every request `err`) → fixed; anon memory sampled (R-652/R-653); ghost re-run | — |
|
||
| 21:30 | C decision 22 | `memory_tight` reads the app's own memory (anon) where measured — CC unattended, recorded in `09` §3 | 09-update-architecture.md |
|
||
| 21:32 | C move | **romm 5.3.0→5.3.1** catalog `7ff4e68` (CI 922) — gate judged it on the real registry; wrong-digest control REFUSED | C-moves.txt |
|
||
| 21:34 | C move | **ghost 6.64.0→6.65.0** `e6afab9` (CI 923) — after a re-run with real load | apps/ghost |
|
||
| 21:41 | C move | **wishlist v0.66.0→v0.67.1** + **opengist 1.13→1.15** `657c83f` (CI 925) | apps/wishlist, apps/opengist |
|
||
| 21:45 | C move | **navidrome 0.64.0→0.64.1** `76cc1f5` — second step on its ladder | apps/navidrome |
|
||
| 21:40 | C not moved | **adventurelog v0.13.0**: backend + data fine, the NEW frontend's own healthcheck (`/health` → backend `/health/`) never passes in the template's network shape — box UNDID the update (under the 90 s drill timeout), bench never healthy in 420 s | apps/adventurelog |
|
||
| 21:58–22:18 | C moves | kimai-db 11.6→11.8 `376b2c3`, romm-db 11.4→11.8 `431cdec` (each its own edge; the engine says "already upgraded to 11.8.9"), immich-ML v3.0.3→v3.2.2 `b82b7c6`, emby `8d573ad`, n8n `04b63db`, komga + 768M `fc1becc`, nextcloud (files_may_change) `5d49a11` — **Part C done: 12 steps, 11 apps** | C-moves.txt |
|
||
| 22:05 | D0 | chaos schedule + app table committed BEFORE round 1 (`cf8dec8`); round 11's second update changed to romm's engine step before round 1 | D0-schedule.txt |
|
||
| 22:25 | D setup | six deployed + seeded (A), whole-box backup 95 s, seed B (docmost, vikunja; romm's B route 404 on 5.3.0); docmost made big with ONE login (412 spaces, 11 MB — its login is rate-limited); nextcloud rebuilt clean after **R-657** | chaos/00-* |
|
||
| 22:29–23:57 | D rounds | 12 rounds run (chaos/round-NN.*); rounds 2–4 accidents landed early (the runner waited on the wrong file for the catalog — fixed: armed at the press); round 8 held → household restore brought it back, A read; round 9 kill mid-verify → resumed → UNDONE, A+B read; round 11 held → **the named restore REFUSED → R-659 (P1), the stop rule** | chaos/, D-rounds.md |
|
||
| 23:12 | D finding | restored apps' volumes carry NO compose label → the undo copies nothing (**R-658, P1**), with a control (never-restored apps labeled) | chaos/09-volumes-unlabeled-after-restore.txt |
|
||
| 23:21 | D 9201 | guarded Update on the HP box's standing apps moved tonight: opengist → 1.15 `done`, kimai-db → 11.8 `done`, romm → 5.3.1 + 11.8 in ONE press `done`; front doors 200; romm 0 OOM kills after | press9201/ |
|
||
| 00:10 | E1 R-640 live | the household's restore refused a half-length docmost copy („…csonka… nem indult el — az alkalmazás és az adatai érintetlenek"); container start times unchanged; A read; the whole copy then restored | E1-r640-live.* |
|
||
| 00:15 | E teardown | 9202: all apps removed through the product (0 containers/volumes/copies; restored apps' remove answered `volumes_removed: []` though the volumes went — R-658), kept drive folders removed by name, config identical, live catalog; drill reset `5d49a11`, has_actions false; bench 9401 destroyed + template removed; 9201 24/24 healthy | E2–E7 |
|
||
| 00:20 | E events | the debug ring missed both `app_update_held` — rebuilt from the controller's full log (`E8-*`); R-660 (a held app also raises app_start_failed) | E8-* |
|
||
| 00:30 | E floor | **NOT raised** (brief's rule: Part D found R-658/R-659, both pre-existing); recommendation to raise → the operator's one question in STATUS | STATUS.md |
|
||
| 00:45 | E docs | DRILL record, `09` §3 21–23 + §6.4 + §6.1a, capability map, rotation tick (opengist), register 330 → 336 (R-651..R-660 opened; R-650, R-640, R-499, R-626 closed), CONTEXT, REPORT, STATUS | — |
|