Files
felhom.eu/documentation/audits/update-night-2026-09-21/PROGRESS.md
T
admin 1ee14ce166
gates / gates (push) Successful in 24s
Update night: the final PROGRESS lines — teardown, documents, CI green by id
Both CI runs confirmed by job id: app-catalog-felhom.eu job 830 and felhom.eu job 870, each
'completed' / 'success', matched on head_sha. unproven.py --summary did not move and that is
stated rather than left for the reader to notice.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-21 22:35:42 +02:00

62 lines
18 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# UPDATE NIGHT 2026-09-21 — progress log (one line per finished step)
Evidence dir: `documentation/audits/update-night-2026-09-21/`
A resuming session reads THIS FILE FIRST and never repeats a finished step.
| time (CEST) | step | verdict | evidence |
|---|---|---|---|
| 20:07 | P0.1 floor → 0.261.0 (declared MinAgent 0.131.0) | **PROVEN** — both demo boxes arrived in **13 s**; hub `managed floor SERVED … from declared (golden 0.258.0)`; drill-r50 stays held (agent 0.129.0 < 0.131.0, host DOWN) | 01-floor-pre.txt, 02-floor-save.txt |
| 20:08 | P0.2a drill repo `admin/app-catalog-drill` created (private) | **DONE** — Gitea token lacks `write:user`, so `POST /api/v1/repos/migrate` was used instead of `/user/repos`; main = `f5f6a152b513` = live | 03-drill-repo.txt |
| 20:10 | P0.2b 9202 repointed at the drill repo | **PROVEN** — **`git.repo_url` ALONE IS INERT**: `gitCloneOrPull` only clones when `.git` is absent, else fetches from the old `origin`. Cache dir had to be removed too. Config saved as `controller.yaml.pre-update-night` | 04-9202-config-pre.txt, 05-9202-follows-drill.txt |
| 20:12 | P0.2c positive + negative controls | **PROVEN** — drill bump uptime-kuma 2.4.0→2.5.5 shows on 9202 as „Frissítés elérhető — ma" / "Update available — today"; live catalog still `f5f6a152b513` with pin 2.4.0; both real boxes' caches still at `f5f6a15`. R-607 fired again (sync said „nincs változás"). App route is **`/apps/<n>`**, not `/app/<n>` | 06-drift-rerun.txt, 07-positive-control.txt |
| 20:07 | Drift re-run | **CONFIRMED** — 66 pins, 46 behind, **39 within-major, 7 across-major** — the brief's numbers hold exactly | 06-drift-rerun.txt |
| 20:13 | P0.4 capacity measured | **OK** — 9202: 26 GB RAM (22 free), docker root on mp0 with **56 GB free**, drive 872 GB. demo-hp `/` is 86% but holds neither. Brief's capacity claim HOLDS | 08-capacity.txt |
| 20:18 | P0.3 throwaway image store | **PROVEN** — `registry:2` on 9202 `127.0.0.1:5000`; `drill/glance:1.0.0` = real v0.8.6 retagged, `:1.0.1` = starts/stays up/never serves, `:1.0.2` absent (404). **`CompareImageRefs` DOES order `host:port/` refs** — 4 positive + 1 negative control, run not read; tmp test deleted, tree clean | 09-image-store.txt |
| 20:20 | **EDGE privatebin 2.0.5 -> 2.0.6** (file-leg) | **PROVEN** — 15.4 s; seed read back both sides; all four observables agree | apps/privatebin/ |
| 20:25 | **EDGE docmost 0.95.0 -> 0.96.0** (db-postgres, engine constant) | **PROVEN** — 103.5 s; seeded account authenticated after; all four observables agree | apps/docmost/ |
| 20:28 | **EDGE bookstack 26.05.2 -> 26.05.5** (db-mariadb, engine constant) | **PROVEN** — 45.1 s; artisan readback with its own negative control; DATABASE HALF ONLY (R-460) | apps/bookstack/ |
| 20:37 | **FINDING R-618** tandoor probe port | **DEFECT, measured** — probe 8080, app listens on 80 only; docker healthy + front door 200 + controller `unhealthy`. 53-template sweep run: wger suspected, adventurelog cleared | 10-probe-port-sweep.txt |
| 20:38 | **FINDINGS R-615/616/617/619** | filed — repo_url inert; catalog token plaintext in the clone; Gitea migrate endpoint; `type: password` mandatory though the wire says optional | NEW-ROWS.md |
| 20:39 | **EDGE actualbudget 26.7.0 -> 26.9.0** | **PROVEN** — 19.5 s | apps/actualbudget/ |
| 20:44 | **EDGE navidrome 0.63.2 -> 0.64.0** (file-leg, HDD_PATH) | **PROVEN** — 11.3 s | apps/navidrome/ |
| 20:46 | **EDGE audiobookshelf 2.35.1 -> 2.36.1** (file-leg) | **PROVEN** — 23.6 s | apps/audiobookshelf/ |
| 20:54 | **EDGE adventurelog v0.12.1 -> v0.13.0** (db-postgis) | **FAILED — the most valuable result so far.** Nine migrations applied OK, then the app never bound its port; held after the full 5-min wait; the sentence names tier, date and what the copy holds. Rows R-621 (the hold destroys the failure evidence) and R-622 (do not promote this edge) | apps/adventurelog/ |
| 20:59 | **Harness code pushed to the catalog** — 4 fixtures + edges U1..U7 | **DONE, gates green** — live catalog main moves `f5f6a152b513` -> `4463243f2e09`, **`scripts/` ONLY, zero `image:` lines** (the teardown diff still expects every image line identical). Harness RUNS owed | app-catalog-felhom.eu@4463243f2e09 |
| 20:59 | image reclaim BY NAME (no prune) | 34 unused images removed by exact reference; 40 GB -> 55 GB free | reclaim.sh |
| 21:08 | CI check for the catalog push | **GREEN** — job **id=830**, `name='gates'`, `status='completed'`, `conclusion='success'`, matched on `head_sha=4463243f2e09`. Confirms CLAUDE.md's warning: page 1's newest id was 52, the real newest was 830 — job ids are NOT page-ordered | scratchpad/ci.sh |
| 21:12 | **Observation, NOT a new row** — an app with an `HDD_PATH` cannot have its DATA removed on 9202 | R-442's guard working as designed: `/api/disks` answers `agent not configured` on this guest, so the drive path cannot be RESOLVED and the removal is **refused with the app kept** rather than half-deleted. The household's other choice — remove the app, KEEP the data — is accepted. The harness now takes that route and tidies its own directories by name at teardown | apps/navidrome/, apps/romm/ |
| 21:13 | **FINDING R-623** — the unattended caller turned every SUCCESS into a `timeout` | Read before use, not after: `follow()` read `update_phase`/`updating` off the API ENVELOPE, so both were always `None`, every followed update hit the 900 s timeout and was then marked never-press-again. **Fixed in that file** before B1 relied on it. The earlier night missed it because its only `follow()` pass was the one whose log was lost | update-arc-gaps-2026-09-21/unattended-caller.py |
| 21:18 | **Interim commit pushed** — felhom.eu `da20722e76a7` | Phase 0 + Phase 1 evidence off the machine before Phases 2-5 (R-320). repo_gates.py --fast: all 15 OK. Includes the two instrument fixes (api-recipe route, unattended-caller `follow()`) | felhom.eu@da20722e76a7 |
| 21:18 | **R-618 RAISED TO P1 by a live measurement** | tandoor 2.6.13→2.6.15 was **serving HTTP 200 on the NEW version** at 21:18:28 while the update sat in `verifying`, because the probe names port 8080 and the app listens on 80. The health wait then times out and `failAndHold` STOPS the working app. A wrong port turns every successful update into an outage plus a needless restore | 14-tandoor-serving-while-verifying.txt |
| 21:22 | **EDGE tandoor 2.6.13 -> 2.6.15** (db-postgres) | **FAILED — and the failure is OURS, not the app's.** The new version was SERVING 200 with a green docker healthcheck at four samples across five minutes; the controller's probe names port 8080 and the app listens on 80, so `verifying` timed out and `failAndHold` STOPPED a working app. The four observables disagree exactly as `09` §5.2 predicts. **R-618 is now P1** | 14-tandoor-serving-while-verifying.txt, apps/tandoor/ |
| 21:35 | **failwalk adventurelog — the household's WAY OUT, on a real failed edge** | **PROVEN** — the restore the hold sentence names („helyi", 1 copy offered) ran in **75 s**; app back on v0.12.1, all three containers running, hold CLEARED, state `running`. R-606 CONFIRMED on the hold sentence itself (English page, Hungarian sentence) with positive and negative controls | apps/adventurelog/failwalk.json, 15-r606-*.txt |
| 21:36 | Observation after the restore | The hold sentence is GONE from the card (R-480 working) and the app reads „Fut". The badge briefly read „Frissítés elérhető — **65 napja**" rather than „ma", because the restore brought back the recovery unit's own `.felhom.yml` with its pre-drill `catalog_since`; the syncer copies that file verbatim every cycle, so it heals within one sync. Recorded as an observation, not a row | apps/adventurelog/failwalk.json |
| 21:37 | **failwalk tandoor — the way out of the WRONG-PROBE hold** | **PROVEN** — restore in **32.4 s**, hold cleared, app running, data reads back. Note what this means for R-618: the household's route back works, so nothing is LOST — but they had to walk it for an update that had actually succeeded | apps/tandoor/failwalk.json |
| 21:40 | **PHASE 2.1 — MariaDB 11.6 -> 12.3 through the REAL Update button** | **PROVEN, a first.** All four SPIKE-r459 observables: datadir `11.6.2`->`12.3.3`; the engine's own check says „already upgraded … no need to run mariadb-upgrade again"; the entrypoint says „Major version upgrade detected … Check required!" then started and **finished** (NOT the `skipped due to $MARIADB_AUTO_UPGRADE` line R-459 feared); the engine's own pre-upgrade backup `system_mysql_backup_11.6.2-MariaDB.sql.zst` 631 905 B. Seeded Nextcloud account read back. All four version observables agree | 16-phase2.1-mariadb-major.md, apps/nextcloud-engine-mariadb/ |
| 21:45 | **PHASE 2.2 — PostgreSQL 16 -> 17 through the real Update button** | **FAILED, exactly as R-463 predicted and nobody had measured.** The update ended in **5.1 s**, the app was stopped and held, and **`pinned` named `postgres:17-alpine` while `installed` still said 16 and nothing was running** — the brief's own question answered. The restore then brought it back in **29.1 s** (hold CLEARED, health probe 200) | apps/docmost-engine-postgres/ |
| 21:45 | **R-621 demonstrated itself on this leg's KEY artefact** | The engine's refusal line was GONE before any probe could read it — `failAndHold` had removed the container, and the controller log does not carry it either. Reproduced independently instead (R-320) | 18-postgres-refusal-reproduced.txt |
| 21:45 | **My own first reproduction was WRONG, and is kept labelled** | The volume lookup returned empty -> the copy was empty -> postgres:17 initialised a FRESH datadir and started happily. A blank `PG_VERSION` should have stopped the step and did not. Rewritten so every step proves itself first. The accidental result (16 refusing a 17 datadir, verbatim) is real and kept | 17-postgres-refusal-reproduced.txt |
| 21:56 | **B1 — THE UNATTENDED HOLD. MEASURED, and it is the thing `09` Q4 has never had.** | The caller pressed ONCE with nobody watching; the drill image started, stayed up and never served; `safety-dump` -> `verifying` -> **`held after 312.9 s`**; then **pass 2 and pass 3 pressed NOTHING** — `outcomes={'glance': ('held', 312.9)} never_again=['glance']`. The hold sentence names the tier, the date and what the copy holds. **My `follow()` fix (R-623) was load-bearing: without it this would have read `timeout` after 900 s** | bad-days/B1-unattended-hold/ |
| 21:57 | **B2 — the new tag cannot be pulled** | **Exactly §6.1 Scenario E.** `phase=failed` in **1.0 s**, `hold=None`, „Az új verzió letöltése nem sikerült, ezért a frissítés elmaradt. Az alkalmazás a korábbi verzióval fut tovább." Nothing held; the app keeps running the old version | bad-days/B2-pull-fails/ |
| 21:57 | **B3 — Update pressed while a backup runs** | `409 reason='busy'` on six consecutive presses with „A frissítés most nem indítható: mentés/visszaállítás folyamatban. Próbáld újra, ha befejeződött." — the transient reason an unattended caller needs (R-609) | bad-days/B3-busy-during-backup/ |
| 21:58 | **B7 — the 2 GB disk floor** | **REFUSED correctly**: filled to 1.4 GB free, `409 reason='disk'`, „Nincs elég szabad hely a frissítéshez: 1.4 GB szabad… legalább 2 GB szükséges." Nothing moved. Space released immediately after | bad-days/B7-disk-and-memory-refusals/ |
| 21:58 | **CORRECTION to R-606** — the refusals are NOT English | Three refusals requested with `?lang=en`, Hungarian request as control: `held`, `not_deployed` and `disk` all come back **identical Hungarian**. The pipe exists (`langFor` honours `?lang=`) but the sentences are frozen string constants, so routing them through `errText` changes nothing | 20-refusals-in-english.txt |
| 22:00 | **B9 — a frozen app gets a newer `.felhom.yml`** | §5.4's asymmetry confirmed live; **no false alarm** — 10 samples, all `running`/200. The reason NARROWS R-458: `type: http` treats any response as healthy, so a wrong path is invisible; the risk is real only for `type: api` + `expect` | 21-B9-frozen-app-newer-felhomyml.md |
| 22:00 | B8 skipped here — docmost was swept by B1's precondition | re-run in `phase2_redo.sh` with a freshly deployed docmost | phase2_redo.sh |
| 22:02 | **B4 (two) — two Updates within one second** | **THERE IS NO SINGLE-FLIGHT.** Both accepted `202` within **0.268 s**, and **2 of 2 ran in flight at the same time**. Both ended `done`, no error, no hold, both pins advanced. That is a real Slice-6 design input: an unattended caller pressing N apps would run N updates at once | bad-days/B4-concurrent-updates/two.json |
| 22:06 | **B4 (five) — five Updates within half a second** | **5 of 5 ran concurrently and ALL ended honest**: accepted within **0.448 s**, every one `done`, no error, no hold, every pin advanced, whole batch ~30 s. So the box does not serialise updates at all — and on this box it coped. Memory samples in the record | bad-days/B4-concurrent-updates/five.json |
| 22:08 | **B5 (backing-up cut) — the phase nobody had cut in** | **BEHAVED.** Cut at `backing-up` +0.027 s (`pct stop` returned 9.65 s later — the usual instrument limit, stated). After boot the box said so ITSELF: „crash recovery: an app-data backup (volume dump) … was interrupted and left 1 app(s) stopped — restarting them: [privatebin]" and „update recovery: privatebin was interrupted in backing-up". **The pin did NOT move**, the app runs, and a fresh paste seeded+read back | bad-days/B5-backing-up/ |
| 22:08 | **B5 (safety-dump cut) — MISSED, recorded as a miss** | The phases went `backing-up` -> `pulling` -> `failed` in **0.473 s** and `safety-dump` was never observed, so the plug was never pulled. Recorded as a MISS, not as a pass. To be retried with a genuine pending edge | bad-days/B5-safety-dump/ |
| 22:09 | **Phase 4 — the morning after** | Every one of the **10 badges is TRUE** (refs equal <-> „Naprakész", refs differ <-> „Frissítés elérhető"). Q4's four promises all scored True on the held app's page. `zipline` shows the household „Nem egészséges — URL nem elérhető" while running — R-618 in the household's own words | bad-days/P4-morning-after/ |
| 22:09 | **B6 — the way out FORWARDS: REFUSED** | A held app met a FIXED newer version. Badge: „Frissítés elérhető — ma" in both languages, button offered; press -> **`409 reason='held'`**. Correct per §6.1 (only a restore lifts a hold) but **the page invites what the button refuses** — the exact inconsistency R-524 removed for the Ahead case. Filed as **R-625** | bad-days/B6-way-out-forwards/ |
| 22:16 | **PHASE 2.3 — the PostgreSQL conversion REHEARSAL, costed (Q5)** | **WORKED end to end.** 49 MB / 48 tables: dump with 16 **2.6 s / 132 201 B**; fresh 17 + replay **6.5 s / 48 tables**; the app said „Database connection successful"; **the seeded account read back on 17**; total **155.9 s**, of which ~9 s is engine work. Two benign ERROR lines named. `pg_upgrade` NOT run — it needs an image that does not exist here | 24-Q5-postgres-conversion-costed.md |
| 22:18 | **wger CONFIRMED — R-618's last candidate closes** | probe `type: http port: 80`; inside the container **port 80 refused, port 8000 ANSWERED**; docker health green; front door 302; the box says `unhealthy`. **Three confirmed instances now: tandoor, zipline, wger** — and the cheap static rule finds all three with one false positive out of 53 | 22-wger-probe-measured.txt |
| 22:19 | **B5 (safety-dump cut) — the second early phase, RETRIED and HIT** | Cut at `safety-dump` +0.023 s. After boot: **no recovery line and no journal** (unlike the `backing-up` cut, which produced both), **the pin did NOT move**, the app runs, a fresh paste seeded and read back, and the card shows **no interrupted sentence at all**. No half-written backup artefact on either cut | 25-B5-the-two-early-cuts.md |
| 22:20 | **Bonus proof across a REAL power cut** | The boot sweep met the held `glance` after an unclean shutdown and **deliberately left it alone**: „is a boot orphan by intent but is HELD … NOT starting it; whatever is holding it owns its recovery". §6.1's three-unattended-paths claim, proven across a power cut | 25-B5-the-two-early-cuts.md |
| 22:22 | **mealie v3.20.1 -> v3.27.0** (db-postgres) | **PROVEN** — the edge that failed twice on instrument problems, walked cleanly on the third | apps/mealie/ |
| 22:23 | **PHASE 1+2 COMPLETE** | **21 edges attempted: 14 proven, 3 failed, 4 inconclusive.** Ten of the fourteen printed a verbatim migration line. Up from the **three** apps this project had ever measured | summarise.py |
| 22:27 | **PHASE 5 — teardown, three layers plus Gitea** | **Machine:** every throwaway app removed through the product; three with an `HDD_PATH` refused at „remove with data" (R-442, correct) and removed keeping it; registry, drill images and drill volumes gone BY NAME, never pruned; `controller.yaml` restored and `git.repo_url` **read back as the LIVE catalog with an empty token**; only `felhom-controller`, `filebrowser`, `traefik` left. **Host:** no harness LXC was created, so none was destroyed; guest 9201 untouched. **Hub:** nothing provisioned; floor 0.261.0. **Gitea:** drill repo KEPT, private, reset to live `main`; **every `image:` line IDENTICAL** | teardown/ |
| 22:28 | **The teardown found the night's LAST defect** | A `navidrome` container the product's own removal left behind, resurrected by Docker's restart policy at a power cut, invisible to every sweep keying on `deployed`. Removed by name. **R-626** | 26-removed-app-came-back.txt |
| 22:4x | **Documents, register, CI** | 12 new rows (register **303 -> 315**), 8 existing rows updated, `09` §3b answered for Q2–Q6, §6.4 re-costed, §6.5 (the drill method) added, limitation 8 replaced, capability map widened, rotation **2 -> 16** fully-walked apps, STATUS morning note. **CI green by id: catalog job 830, felhom.eu job 870, both `success`** | felhom.eu@8d786f7940b0 |
| — | **`unproven.py --summary` did NOT move** | 55 claims, 35 not walked — unchanged. Tonight measured the UPDATE arc, which that file tracks as one claim already marked walked; no claim's status changed, so no number moved. Stated because the checklist asks | scripts/unproven.py |