1a72fef87f
gates / gates (push) Successful in 26s
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
7.4 KiB
7.4 KiB
PROGRESS — night 2026-09-24 (run 2026-09-24 from 11:07 CEST)
Step log. Newest at the bottom. Times CEST unless marked Z.
- 11:07 baselines verified: controller 7c3b3a969475 (v0.268.0), agent d9864a94bf62, felhom.eu
86c4b9a0d8(hub 0.122.0), catalog 635c7e68f716. Register 335 rows / 679,393 B. - 11:08
09§3 decisions 26–28 recorded. - 11:09 9202 repointed to the drill catalog (controller.yaml saved as .pre-night0924)
- 11:20 A1 (whole restore + R-668) and A2 (decision 27) built, tested, red-proofed
- 11:34 A3 built: detector (6 restarts/10 min — measured gokapi 1/min), stop+hold+event, Start lifts; red-proofs detector + skips
- 11:49 A4 (R-664 box+catalog, R-665, R-662 removed) and Part B (digests) built, red-proofed; hub 0.123.0 live
- 12:21 9202: drill setting update.backup_max_age 1m (forces the update's own backup + Tier 2); restored at teardown from .pre-night0924
- 12:25 A1 LIVE PASS (9202, v0.269.0): whole restore from the SSD Tier-2 copy — files restored=1 kept-newer=1 unchanged=137, unit 3/3 volumes 1/1 db, 38 s; account + all three files read back (A1/21-readback.json). R-668 live: the new copy went to the SSD, not the same disk.
- 12:29 A3 LIVE trip 1: gokapi stopped by the box „crash_loop — 6 in 10m0s", page hu/en + Start, no restore badge (A3/21).
- 12:37 A3 LIVE Start + trip 2: Start LIFTED the hold; Docker's back-off reset on the new start → 9 restarts in 32 s → stopped again, trip 2, page hu/en says support is told; 0 app_start_failed lines (A3/30-start-then-trip2.*).
- 12:41 A1 LIVE hold names the second drive: failing edge (redis 7.4-alpine + probe 8999), undo copy cut → HELD, page hu/en „második meghajtó … beállításokat, az adatbázist és a fájlokat" / "second drive … settings, the database and the files"; hdd-data keep_data_only absent (A2 positive control); the page's whole restore → hold CLEARED, 33.8 s, unit 3/3 1/1, account + files read back (A1/30-held-then-whole.*).
- 12:42 Part B seen live in passing: nextcloud's compose on 9202 runs redis:7-alpine@sha256:858f… (floating tag, digest from the ladder entry).
- 12:43 FINDING: after hold + whole restore the probe stays on the FAILED step's .felhom.yml (applied-meta 10:39:21 = port 8999; stack .felhom.yml restored from the unit, which the sync had already flowed). Steady probe heals on the next sync; applied-meta does not until a pin advance → a later failed update's undo would probe the old version with the wrong check. ROW to file.
- 12:43 FINDING:
[ERROR] .felhom.yml backup block rejected in …/pre-update-meta: docker-compose.yml unreadableon an undo (10:24:41). ROW to file. - 12:44 Part E schedule drawn (tools/chaos_schedule.py, seed 20260924, md5 of table b5ccf895…).
- 12:50 A3/40: first-start restarts in all drill evidence — no healthy app ≥ 6; immich 12 (broken, first-start import). ROW (watch). ROWS TO FILE: R-668 (Tier-2 target on a missing path, fixed); applied-meta after restore; pre-update-meta ERROR noise; ladder log "older than the ladder" at head; missingFileLegsRefusal text names file restore not whole; immich first-start watch; R-644 gokapi disposition.
- 12:48 A2 LIVE PASS: a no-whole-copy hold (second-drive mirror moved aside = a one-drive box) → page hu/en „…ügyfélszolgálatát értesítettük — addig ne indítsd újra, és ne töröld az adatait." ; hdd-data keep_data_only=true; Remove with data → 409 hu/en, with backups → 409; app still deployed; mirror back → whole restore cleared it; account + files read back (A2/10-held-no-copy-keep-data.*).
- 12:48 FINDING CONSEQUENCE PROVEN: that hold was reached because the undo judged the (correct) old version with the FAILED step's probe (applied-meta 8999 left by the previous hold+restore) → a false hold (A2/11). Undo-copy volumes of the earlier hold (…103924Z, ~0.9 GiB) still present after its hold was cleared by a restore. ROWS.
- 12:55 Part B: compose config accepts tag@digest for every template (25 apps, 37 digests, 0 missing, 0 differ) — B/01.
- 12:51 Part B LIVE on 0.269.0: badge Behind for a newer tested digest of a floating tag, Update pulled it, record + badge current (B/10). FINDING: the SYNC wrote the new digest into the running app's compose BEFORE Update → a restart would pull it unguarded. FIXED v0.269.1 (CarryDigests), red-proofed (redproofs/B-sync-carry.txt), pushed 5b1b191, image built, 9202 on 0.269.1 (B/20).
- 13:05 Part B LIVE PASS on 0.269.1: sync KEPT the running digest before Update (True), badge Behind hu/en, Update done 44 s, running digest = the tested one, badge current (B/21-floating-tag-0.269.1.*).
- DECISION (CC, unattended): a second controller release tonight (0.269.1) instead of shipping the bypass with the floor — first in the morning note.
- 13:07 FOUND (Part C read of the chain): demo-hp's thin pool
local-lvm100 % FULL (out_of_data_space, error_if_no_space) since 10:35 CEST. Cause (agent journal, C-04): the SCHEDULED restore-test at 10:29 restored 9201's 22 GiB archive into the SAME pool as scratch guest 990000 with no free-space check; start failed; teardown failed („filesystem in use") → "left for Recover", and Recover runs only at agent start. The agent logged „a full pool corrupts every guest on it" every few seconds and did nothing else. 9201's controller: „no space left on device" from 10:38 CEST, its docker log stops 10:39. 9201 also carried a stalesnapshot-deletelock; its whole-box backups failed 06:59–08:22 (CT is locked). Before tonight's run; not caused by it (9202 lives on nvme-scratch). - 13:09 INTERVENTION 1:
systemctl restart felhom-agenton demo-hp so the agent's own Recover tears the scratch down (a directpct destroywas refused by the session's permission check). Result (C-05): „recover: destroyed leaked restore-test scratch guest" vmid=990000; pool 100 % → 58.91 %, state rw; 9201's lock gone. A further read of the morning vzdump task logs was refused by the permission check — left for the operator. ROWS TO FILE (agent, P1): restore-test fills the production pool (no space check) and its leaked scratch waits for an agent restart; the pool-fill WARN has no action; 9201 stale snapshot-delete lock blocked whole-box backups all morning. - 13:45 Part E round 1 (oom_app × power cut): OOM storm counted again after the boot, stopped +185 s; events app_oom_storm (error) + app_stopped_unhealthy (warning). PASS.
- 13:49 Part E round 2 (install × docker restart): the restart killed the controller mid-deploy; n8n came back NOT installed, no event, no page message; leftovers app.yaml/applied files in the stack dir. FINDING (row). n8n reinstalled through the product (211 s, leftovers did not block) so round 10 has its step.
- 14:04 Part E round 3 (cutoff file app × power cut at verifying): after the boot the update RESUMED, undo found the cut copy → HELD naming the second drive (whole) → whole restore 37 s → account + files read back. PASS. BUT the hold named the Tier-2 copy from 13:04, not a fresh one — the update's own backup did not refresh Tier 2 (to reproduce in round 11); the pre-cut controller log was lost with the container (runner now saves it at arm time).
- 14:07–14:40 Part E rounds 3–12 (see the findings doc table). Kills: 2 of 3, 20 min 32 s apart.
- 14:41–14:50 Part F: 9202 apps removed through the product, drive folders by name, window 02:30, config restored; drill reset; host read; floor 0.269.1 → N100 arrived, demo-hp 9201 not (R-672).
- 15:2x CI green by head_sha: controller 7c6140d (job 958), catalog c802509 (959), felhom.eu
e924a17(961). Scratch secrets shredded.