Files
felhom.eu/documentation/audits/night-2026-09-24/PROGRESS.md
T
admin 1a72fef87f
gates / gates (push) Successful in 26s
night 2026-09-24: CI by head_sha, final progress line
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-24 14:56:48 +02:00

37 lines
7.4 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# PROGRESS — night 2026-09-24 (run 2026-09-24 from 11:07 CEST)
Step log. Newest at the bottom. Times CEST unless marked Z.
- 11:07 baselines verified: controller 7c3b3a969475 (v0.268.0), agent d9864a94bf62, felhom.eu 86c4b9a0d8bf (hub 0.122.0), catalog 635c7e68f716. Register 335 rows / 679,393 B.
- 11:08 `09` §3 decisions 26–28 recorded.
- 11:09 9202 repointed to the drill catalog (controller.yaml saved as .pre-night0924)
- 11:20 A1 (whole restore + R-668) and A2 (decision 27) built, tested, red-proofed
- 11:34 A3 built: detector (6 restarts/10 min — measured gokapi 1/min), stop+hold+event, Start lifts; red-proofs detector + skips
- 11:49 A4 (R-664 box+catalog, R-665, R-662 removed) and Part B (digests) built, red-proofed; hub 0.123.0 live
- 12:21 9202: drill setting update.backup_max_age 1m (forces the update's own backup + Tier 2); restored at teardown from .pre-night0924
- 12:25 A1 LIVE PASS (9202, v0.269.0): whole restore from the SSD Tier-2 copy — files restored=1 kept-newer=1 unchanged=137, unit 3/3 volumes 1/1 db, 38 s; account + all three files read back (A1/21-readback.json). R-668 live: the new copy went to the SSD, not the same disk.
- 12:29 A3 LIVE trip 1: gokapi stopped by the box „crash_loop — 6 in 10m0s", page hu/en + Start, no restore badge (A3/21).
- 12:37 A3 LIVE Start + trip 2: Start LIFTED the hold; Docker's back-off reset on the new start → 9 restarts in 32 s → stopped again, trip 2, page hu/en says support is told; 0 app_start_failed lines (A3/30-start-then-trip2.*).
- 12:41 A1 LIVE hold names the second drive: failing edge (redis 7.4-alpine + probe 8999), undo copy cut → HELD, page hu/en „második meghajtó … beállításokat, az adatbázist és a fájlokat" / "second drive … settings, the database and the files"; hdd-data keep_data_only absent (A2 positive control); the page's whole restore → hold CLEARED, 33.8 s, unit 3/3 1/1, account + files read back (A1/30-held-then-whole.*).
- 12:42 Part B seen live in passing: nextcloud's compose on 9202 runs redis:7-alpine@sha256:858f… (floating tag, digest from the ladder entry).
- 12:43 FINDING: after hold + whole restore the probe stays on the FAILED step's .felhom.yml (applied-meta 10:39:21 = port 8999; stack .felhom.yml restored from the unit, which the sync had already flowed). Steady probe heals on the next sync; applied-meta does not until a pin advance → a later failed update's undo would probe the old version with the wrong check. ROW to file.
- 12:43 FINDING: `[ERROR] .felhom.yml backup block rejected in …/pre-update-meta: docker-compose.yml unreadable` on an undo (10:24:41). ROW to file.
- 12:44 Part E schedule drawn (tools/chaos_schedule.py, seed 20260924, md5 of table b5ccf895…).
- 12:50 A3/40: first-start restarts in all drill evidence — no healthy app ≥ 6; immich 12 (broken, first-start import). ROW (watch).
ROWS TO FILE: R-668 (Tier-2 target on a missing path, fixed); applied-meta after restore; pre-update-meta ERROR noise; ladder log "older than the ladder" at head; missingFileLegsRefusal text names file restore not whole; immich first-start watch; R-644 gokapi disposition.
- 12:48 A2 LIVE PASS: a no-whole-copy hold (second-drive mirror moved aside = a one-drive box) → page hu/en „…ügyfélszolgálatát értesítettük — addig ne indítsd újra, és ne töröld az adatait." ; hdd-data keep_data_only=true; Remove with data → 409 hu/en, with backups → 409; app still deployed; mirror back → whole restore cleared it; account + files read back (A2/10-held-no-copy-keep-data.*).
- 12:48 FINDING CONSEQUENCE PROVEN: that hold was reached because the undo judged the (correct) old version with the FAILED step's probe (applied-meta 8999 left by the previous hold+restore) → a false hold (A2/11). Undo-copy volumes of the earlier hold (…103924Z, ~0.9 GiB) still present after its hold was cleared by a restore. ROWS.
- 12:55 Part B: compose config accepts tag@digest for every template (25 apps, 37 digests, 0 missing, 0 differ) — B/01.
- 12:51 Part B LIVE on 0.269.0: badge Behind for a newer tested digest of a floating tag, Update pulled it, record + badge current (B/10). FINDING: the SYNC wrote the new digest into the running app's compose BEFORE Update → a restart would pull it unguarded. FIXED v0.269.1 (CarryDigests), red-proofed (redproofs/B-sync-carry.txt), pushed 5b1b191, image built, 9202 on 0.269.1 (B/20).
- 13:05 Part B LIVE PASS on 0.269.1: sync KEPT the running digest before Update (True), badge Behind hu/en, Update done 44 s, running digest = the tested one, badge current (B/21-floating-tag-0.269.1.*).
- DECISION (CC, unattended): a second controller release tonight (0.269.1) instead of shipping the bypass with the floor — first in the morning note.
- 13:07 FOUND (Part C read of the chain): demo-hp's thin pool `local-lvm` 100 % FULL (out_of_data_space, error_if_no_space) since 10:35 CEST. Cause (agent journal, C-04): the SCHEDULED restore-test at 10:29 restored 9201's 22 GiB archive into the SAME pool as scratch guest 990000 with no free-space check; start failed; teardown failed („filesystem in use") → "left for Recover", and Recover runs only at agent start. The agent logged „a full pool corrupts every guest on it" every few seconds and did nothing else. 9201's controller: „no space left on device" from 10:38 CEST, its docker log stops 10:39. 9201 also carried a stale `snapshot-delete` lock; its whole-box backups failed 06:59–08:22 (CT is locked). Before tonight's run; not caused by it (9202 lives on nvme-scratch).
- 13:09 INTERVENTION 1: `systemctl restart felhom-agent` on demo-hp so the agent's own Recover tears the scratch down (a direct `pct destroy` was refused by the session's permission check). Result (C-05): „recover: destroyed leaked restore-test scratch guest" vmid=990000; pool 100 % → 58.91 %, state rw; 9201's lock gone. A further read of the morning vzdump task logs was refused by the permission check — left for the operator.
ROWS TO FILE (agent, P1): restore-test fills the production pool (no space check) and its leaked scratch waits for an agent restart; the pool-fill WARN has no action; 9201 stale snapshot-delete lock blocked whole-box backups all morning.
- 13:45 Part E round 1 (oom_app × power cut): OOM storm counted again after the boot, stopped +185 s; events app_oom_storm (error) + app_stopped_unhealthy (warning). PASS.
- 13:49 Part E round 2 (install × docker restart): the restart killed the controller mid-deploy; n8n came back NOT installed, no event, no page message; leftovers app.yaml/applied files in the stack dir. FINDING (row). n8n reinstalled through the product (211 s, leftovers did not block) so round 10 has its step.
- 14:04 Part E round 3 (cutoff file app × power cut at verifying): after the boot the update RESUMED, undo found the cut copy → HELD naming the second drive (whole) → whole restore 37 s → account + files read back. PASS. BUT the hold named the Tier-2 copy from 13:04, not a fresh one — the update's own backup did not refresh Tier 2 (to reproduce in round 11); the pre-cut controller log was lost with the container (runner now saves it at arm time).
- 14:07–14:40 Part E rounds 3–12 (see the findings doc table). Kills: 2 of 3, 20 min 32 s apart.
- 14:41–14:50 Part F: 9202 apps removed through the product, drive folders by name, window 02:30, config restored; drill reset; host read; floor 0.269.1 → N100 arrived, demo-hp 9201 not (R-672).
- 15:2x CI green by head_sha: controller 7c6140d (job 958), catalog c802509 (959), felhom.eu e924a17 (961). Scratch secrets shredded.