Files
felhom.eu/documentation/audits/night-2026-09-23/PROGRESS.md
T
admin 3e58c184f6
gates / gates (push) Successful in 27s
night shift 2026-09-23: the record, the register, the morning note
DRILL-night-2026-09-23.md: Parts A-E. 09 §3 decisions 21 (operator word),
22 and 23 (CC unattended, operator may reverse); §6.4 parts 4 and 6
(catalog half) shipped; §6.1a residuals R-658/R-659. Register 330 -> 336:
R-651..R-660 opened (R-658 and R-659 P1), R-650/R-640/R-499/R-626 closed.
Capability map, nightly rotation (opengist), STATUS (one question: the
floor), CONTEXT, REPORT. The floor stays 0.266.0.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-24 00:24:04 +02:00

7.8 KiB
Raw Blame History

PROGRESS — night 2026-09-23 (one line per finished step, written when it finishes)

time (CEST; rows before 20:53 are approximate) step result evidence
19:16 P0 baselines controller 15e630aa02ed v0.266.0 · agent d9864a94bf62 v0.132.0 · felhom.eu e443841b752a · catalog cfcfe5278428 — all match the brief; trees clean; disk 37 % / 53 % this line
~19:40 P0 drift re-run 66 pins; 26 within-a-major edges, 11 across-a-major (listed, not moved), 1 error (plant-it, abandoned) 02-drift.json, 02-drift.log
19:20 P0 capacity demo-hp 29 GB RAM (23 free), nvme-scratch 7 %; 9202 25.9 GB, /var/lib/docker 64 % (24 GB free); 9201 standing apps listed 00-demo-hp-before.txt, 01-capacity-9202-9201.txt
20:00 A1 R-650 internal/dockerexec guard + repo-wide sweep test; 8 api tests + web/stacks fixtures were reaching real docker → package TestMain stub; red-proofs both FAIL as expected A1-r650-redproof.txt, A1-r650-sweep.txt
20:05 A2 R-640 completion-marker check on all three restore paths, before the first mutation; 3 red-proofs FAIL as expected A2-r640-redproof.txt
20:17 A3 R-626 NOT reproduced on v0.266.0: remove → 390 s of events (destroy seen, no create) + controller restart + guest reboot → 0 containers, 0 volumes. Leftover applied-compose.yml/applied-meta (row) A3-r626-*.txt/json
20:15 A4 R-499 4-branch system-disk sentence; 2 red-proofs FAIL as expected; parity fixtures per branch A4-r499-redproof.txt
20:20 A5 R-518 the "few seconds" promise was ALREADY gone (v0.243.0) — brief claim wrong; copy now states the measured ~8 min; red-proof FAILs as expected A5-r518-redproof.txt
20:26 A6 v0.267.0 pushed 80e6ad8c4772, gates + full suite rc 0; image built; deployed to 9202 only (Up, healthy). Floor waits for Part E A6-deploy-0.267.0-on-9202.txt
20:28 A7 pages live 9202 (no agent): R-499 renders the unknown branch in hu+en, matching /api/storage/backup-target known:false; R-518 text on /backups in hu+en A7-pages-live-0.267.0.txt
20:28 B1 SPIKE PASS on v0.266.0 AND v0.267.0: drill navidrome with a top-level update_ladder: (JSON flow lines) synced, deploy-fields read, deployed running, probe API GET :4533/ping → 200, badge „Naprakész"/"Up to date", no yaml errors. Source agrees: no KnownFields anywhere B1-spike-v0.26{6,7}.0.{json,log}
20:55 B2 gate check-test-record.py (static, CI) + check-test-record-move.py (history + registry for moved refs); 16 decoys both ways; 3 red-proofs FAIL as expected B2-gate-redproofs.txt
21:05 B3 writer upgrade-test.py --write-ladder (+ --move, harness v3: box fixtures on the bench, files_may_change); test_ladder_writer.py 5/5 catalog scripts/
21:10 B4 backfill 21 of 21 proven from their cited records (nextcloud's commit cited none — its record named, entry says so); romm carries M1's 80.9 % memory_tight catalog 6db08a5
21:12 B5 pushed catalog 6db08a5, hook gates OK, CI job 920 success (matched on head_sha) —
20:58 C0 venue bench LXC 9401 on demo-hp (nvme-scratch, 6c/10G), docker 26.1.5, Docker Hub login; C3 negative control → failed (bench measures) C0-harness-venue-create.txt
21:06 C R-612 bench: 128M → seed OOM-killed (oom_kill 3), 1024M peak 345M, 512M peak 312M seed completes; 9202 before: seed SIGKILLed but after the rows (sign-up worked — lie NOT reproduced, timing-dependent); after 512M: seed completes. Live catalog a5a729a apps/wishlist/r612-*.json
21:15 C R-613 UPTIME_KUMA_DB_TYPE=sqlite: before setup-database while box says running; after entryPage, /metrics 401; db in the volume. First "after" run used a stale template (R-607 lag) — kept, renamed. Live catalog a5a729a apps/uptime-kuma/r613-*.json
20:53 C queues bench (9401) and box (9202) queues started, database apps first; each edge's evidence copied off the bench when it ends benchq-.txt, boxq-.txt
21:18 C fix bench load generator followed redirects (every request err) → fixed; anon memory sampled (R-652/R-653); ghost re-run —
21:30 C decision 22 memory_tight reads the app's own memory (anon) where measured — CC unattended, recorded in 09 §3 09-update-architecture.md
21:32 C move romm 5.3.0→5.3.1 catalog 7ff4e68 (CI 922) — gate judged it on the real registry; wrong-digest control REFUSED C-moves.txt
21:34 C move ghost 6.64.0→6.65.0 e6afab9 (CI 923) — after a re-run with real load apps/ghost
21:41 C move wishlist v0.66.0→v0.67.1 + opengist 1.13→1.15 657c83f (CI 925) apps/wishlist, apps/opengist
21:45 C move navidrome 0.64.0→0.64.1 76cc1f5 — second step on its ladder apps/navidrome
21:40 C not moved adventurelog v0.13.0: backend + data fine, the NEW frontend's own healthcheck (/health → backend /health/) never passes in the template's network shape — box UNDID the update (under the 90 s drill timeout), bench never healthy in 420 s apps/adventurelog
21:58–22:18 C moves kimai-db 11.6→11.8 376b2c3, romm-db 11.4→11.8 431cdec (each its own edge; the engine says "already upgraded to 11.8.9"), immich-ML v3.0.3→v3.2.2 b82b7c6, emby 8d573ad, n8n 04b63db, komga + 768M fc1becc, nextcloud (files_may_change) 5d49a11 — Part C done: 12 steps, 11 apps C-moves.txt
22:05 D0 chaos schedule + app table committed BEFORE round 1 (cf8dec8); round 11's second update changed to romm's engine step before round 1 D0-schedule.txt
22:25 D setup six deployed + seeded (A), whole-box backup 95 s, seed B (docmost, vikunja; romm's B route 404 on 5.3.0); docmost made big with ONE login (412 spaces, 11 MB — its login is rate-limited); nextcloud rebuilt clean after R-657 chaos/00-*
22:29–23:57 D rounds 12 rounds run (chaos/round-NN.*); rounds 2–4 accidents landed early (the runner waited on the wrong file for the catalog — fixed: armed at the press); round 8 held → household restore brought it back, A read; round 9 kill mid-verify → resumed → UNDONE, A+B read; round 11 held → the named restore REFUSED → R-659 (P1), the stop rule chaos/, D-rounds.md
23:12 D finding restored apps' volumes carry NO compose label → the undo copies nothing (R-658, P1), with a control (never-restored apps labeled) chaos/09-volumes-unlabeled-after-restore.txt
23:21 D 9201 guarded Update on the HP box's standing apps moved tonight: opengist → 1.15 done, kimai-db → 11.8 done, romm → 5.3.1 + 11.8 in ONE press done; front doors 200; romm 0 OOM kills after press9201/
00:10 E1 R-640 live the household's restore refused a half-length docmost copy („…csonka… nem indult el — az alkalmazás és az adatai érintetlenek"); container start times unchanged; A read; the whole copy then restored E1-r640-live.*
00:15 E teardown 9202: all apps removed through the product (0 containers/volumes/copies; restored apps' remove answered volumes_removed: [] though the volumes went — R-658), kept drive folders removed by name, config identical, live catalog; drill reset 5d49a11, has_actions false; bench 9401 destroyed + template removed; 9201 24/24 healthy E2–E7
00:20 E events the debug ring missed both app_update_held — rebuilt from the controller's full log (E8-*); R-660 (a held app also raises app_start_failed) E8-*
00:30 E floor NOT raised (brief's rule: Part D found R-658/R-659, both pre-existing); recommendation to raise → the operator's one question in STATUS STATUS.md
00:45 E docs DRILL record, 09 §3 21–23 + §6.4 + §6.1a, capability map, rotation tick (opengist), register 330 → 336 (R-651..R-660 opened; R-650, R-640, R-499, R-626 closed), CONTEXT, REPORT, STATUS —