1ee14ce166
gates / gates (push) Successful in 24s
Both CI runs confirmed by job id: app-catalog-felhom.eu job 830 and felhom.eu job 870, each 'completed' / 'success', matched on head_sha. unproven.py --summary did not move and that is stated rather than left for the reader to notice. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
62 lines
18 KiB
Markdown
62 lines
18 KiB
Markdown
# UPDATE NIGHT 2026-09-21 — progress log (one line per finished step)
|
||
|
||
Evidence dir: `documentation/audits/update-night-2026-09-21/`
|
||
A resuming session reads THIS FILE FIRST and never repeats a finished step.
|
||
|
||
| time (CEST) | step | verdict | evidence |
|
||
|---|---|---|---|
|
||
| 20:07 | P0.1 floor → 0.261.0 (declared MinAgent 0.131.0) | **PROVEN** — both demo boxes arrived in **13 s**; hub `managed floor SERVED … from declared (golden 0.258.0)`; drill-r50 stays held (agent 0.129.0 < 0.131.0, host DOWN) | 01-floor-pre.txt, 02-floor-save.txt |
|
||
| 20:08 | P0.2a drill repo `admin/app-catalog-drill` created (private) | **DONE** — Gitea token lacks `write:user`, so `POST /api/v1/repos/migrate` was used instead of `/user/repos`; main = `f5f6a152b513` = live | 03-drill-repo.txt |
|
||
| 20:10 | P0.2b 9202 repointed at the drill repo | **PROVEN** — **`git.repo_url` ALONE IS INERT**: `gitCloneOrPull` only clones when `.git` is absent, else fetches from the old `origin`. Cache dir had to be removed too. Config saved as `controller.yaml.pre-update-night` | 04-9202-config-pre.txt, 05-9202-follows-drill.txt |
|
||
| 20:12 | P0.2c positive + negative controls | **PROVEN** — drill bump uptime-kuma 2.4.0→2.5.5 shows on 9202 as „Frissítés elérhető — ma" / "Update available — today"; live catalog still `f5f6a152b513` with pin 2.4.0; both real boxes' caches still at `f5f6a15`. R-607 fired again (sync said „nincs változás"). App route is **`/apps/<n>`**, not `/app/<n>` | 06-drift-rerun.txt, 07-positive-control.txt |
|
||
| 20:07 | Drift re-run | **CONFIRMED** — 66 pins, 46 behind, **39 within-major, 7 across-major** — the brief's numbers hold exactly | 06-drift-rerun.txt |
|
||
| 20:13 | P0.4 capacity measured | **OK** — 9202: 26 GB RAM (22 free), docker root on mp0 with **56 GB free**, drive 872 GB. demo-hp `/` is 86% but holds neither. Brief's capacity claim HOLDS | 08-capacity.txt |
|
||
| 20:18 | P0.3 throwaway image store | **PROVEN** — `registry:2` on 9202 `127.0.0.1:5000`; `drill/glance:1.0.0` = real v0.8.6 retagged, `:1.0.1` = starts/stays up/never serves, `:1.0.2` absent (404). **`CompareImageRefs` DOES order `host:port/` refs** — 4 positive + 1 negative control, run not read; tmp test deleted, tree clean | 09-image-store.txt |
|
||
| 20:20 | **EDGE privatebin 2.0.5 -> 2.0.6** (file-leg) | **PROVEN** — 15.4 s; seed read back both sides; all four observables agree | apps/privatebin/ |
|
||
| 20:25 | **EDGE docmost 0.95.0 -> 0.96.0** (db-postgres, engine constant) | **PROVEN** — 103.5 s; seeded account authenticated after; all four observables agree | apps/docmost/ |
|
||
| 20:28 | **EDGE bookstack 26.05.2 -> 26.05.5** (db-mariadb, engine constant) | **PROVEN** — 45.1 s; artisan readback with its own negative control; DATABASE HALF ONLY (R-460) | apps/bookstack/ |
|
||
| 20:37 | **FINDING R-618** tandoor probe port | **DEFECT, measured** — probe 8080, app listens on 80 only; docker healthy + front door 200 + controller `unhealthy`. 53-template sweep run: wger suspected, adventurelog cleared | 10-probe-port-sweep.txt |
|
||
| 20:38 | **FINDINGS R-615/616/617/619** | filed — repo_url inert; catalog token plaintext in the clone; Gitea migrate endpoint; `type: password` mandatory though the wire says optional | NEW-ROWS.md |
|
||
| 20:39 | **EDGE actualbudget 26.7.0 -> 26.9.0** | **PROVEN** — 19.5 s | apps/actualbudget/ |
|
||
| 20:44 | **EDGE navidrome 0.63.2 -> 0.64.0** (file-leg, HDD_PATH) | **PROVEN** — 11.3 s | apps/navidrome/ |
|
||
| 20:46 | **EDGE audiobookshelf 2.35.1 -> 2.36.1** (file-leg) | **PROVEN** — 23.6 s | apps/audiobookshelf/ |
|
||
| 20:54 | **EDGE adventurelog v0.12.1 -> v0.13.0** (db-postgis) | **FAILED — the most valuable result so far.** Nine migrations applied OK, then the app never bound its port; held after the full 5-min wait; the sentence names tier, date and what the copy holds. Rows R-621 (the hold destroys the failure evidence) and R-622 (do not promote this edge) | apps/adventurelog/ |
|
||
| 20:59 | **Harness code pushed to the catalog** — 4 fixtures + edges U1..U7 | **DONE, gates green** — live catalog main moves `f5f6a152b513` -> `4463243f2e09`, **`scripts/` ONLY, zero `image:` lines** (the teardown diff still expects every image line identical). Harness RUNS owed | app-catalog-felhom.eu@4463243f2e09 |
|
||
| 20:59 | image reclaim BY NAME (no prune) | 34 unused images removed by exact reference; 40 GB -> 55 GB free | reclaim.sh |
|
||
| 21:08 | CI check for the catalog push | **GREEN** — job **id=830**, `name='gates'`, `status='completed'`, `conclusion='success'`, matched on `head_sha=4463243f2e09`. Confirms CLAUDE.md's warning: page 1's newest id was 52, the real newest was 830 — job ids are NOT page-ordered | scratchpad/ci.sh |
|
||
| 21:12 | **Observation, NOT a new row** — an app with an `HDD_PATH` cannot have its DATA removed on 9202 | R-442's guard working as designed: `/api/disks` answers `agent not configured` on this guest, so the drive path cannot be RESOLVED and the removal is **refused with the app kept** rather than half-deleted. The household's other choice — remove the app, KEEP the data — is accepted. The harness now takes that route and tidies its own directories by name at teardown | apps/navidrome/, apps/romm/ |
|
||
| 21:13 | **FINDING R-623** — the unattended caller turned every SUCCESS into a `timeout` | Read before use, not after: `follow()` read `update_phase`/`updating` off the API ENVELOPE, so both were always `None`, every followed update hit the 900 s timeout and was then marked never-press-again. **Fixed in that file** before B1 relied on it. The earlier night missed it because its only `follow()` pass was the one whose log was lost | update-arc-gaps-2026-09-21/unattended-caller.py |
|
||
| 21:18 | **Interim commit pushed** — felhom.eu `da20722e76a7` | Phase 0 + Phase 1 evidence off the machine before Phases 2-5 (R-320). repo_gates.py --fast: all 15 OK. Includes the two instrument fixes (api-recipe route, unattended-caller `follow()`) | felhom.eu@da20722e76a7 |
|
||
| 21:18 | **R-618 RAISED TO P1 by a live measurement** | tandoor 2.6.13→2.6.15 was **serving HTTP 200 on the NEW version** at 21:18:28 while the update sat in `verifying`, because the probe names port 8080 and the app listens on 80. The health wait then times out and `failAndHold` STOPS the working app. A wrong port turns every successful update into an outage plus a needless restore | 14-tandoor-serving-while-verifying.txt |
|
||
| 21:22 | **EDGE tandoor 2.6.13 -> 2.6.15** (db-postgres) | **FAILED — and the failure is OURS, not the app's.** The new version was SERVING 200 with a green docker healthcheck at four samples across five minutes; the controller's probe names port 8080 and the app listens on 80, so `verifying` timed out and `failAndHold` STOPPED a working app. The four observables disagree exactly as `09` §5.2 predicts. **R-618 is now P1** | 14-tandoor-serving-while-verifying.txt, apps/tandoor/ |
|
||
| 21:35 | **failwalk adventurelog — the household's WAY OUT, on a real failed edge** | **PROVEN** — the restore the hold sentence names („helyi", 1 copy offered) ran in **75 s**; app back on v0.12.1, all three containers running, hold CLEARED, state `running`. R-606 CONFIRMED on the hold sentence itself (English page, Hungarian sentence) with positive and negative controls | apps/adventurelog/failwalk.json, 15-r606-*.txt |
|
||
| 21:36 | Observation after the restore | The hold sentence is GONE from the card (R-480 working) and the app reads „Fut". The badge briefly read „Frissítés elérhető — **65 napja**" rather than „ma", because the restore brought back the recovery unit's own `.felhom.yml` with its pre-drill `catalog_since`; the syncer copies that file verbatim every cycle, so it heals within one sync. Recorded as an observation, not a row | apps/adventurelog/failwalk.json |
|
||
| 21:37 | **failwalk tandoor — the way out of the WRONG-PROBE hold** | **PROVEN** — restore in **32.4 s**, hold cleared, app running, data reads back. Note what this means for R-618: the household's route back works, so nothing is LOST — but they had to walk it for an update that had actually succeeded | apps/tandoor/failwalk.json |
|
||
| 21:40 | **PHASE 2.1 — MariaDB 11.6 -> 12.3 through the REAL Update button** | **PROVEN, a first.** All four SPIKE-r459 observables: datadir `11.6.2`->`12.3.3`; the engine's own check says „already upgraded … no need to run mariadb-upgrade again"; the entrypoint says „Major version upgrade detected … Check required!" then started and **finished** (NOT the `skipped due to $MARIADB_AUTO_UPGRADE` line R-459 feared); the engine's own pre-upgrade backup `system_mysql_backup_11.6.2-MariaDB.sql.zst` 631 905 B. Seeded Nextcloud account read back. All four version observables agree | 16-phase2.1-mariadb-major.md, apps/nextcloud-engine-mariadb/ |
|
||
| 21:45 | **PHASE 2.2 — PostgreSQL 16 -> 17 through the real Update button** | **FAILED, exactly as R-463 predicted and nobody had measured.** The update ended in **5.1 s**, the app was stopped and held, and **`pinned` named `postgres:17-alpine` while `installed` still said 16 and nothing was running** — the brief's own question answered. The restore then brought it back in **29.1 s** (hold CLEARED, health probe 200) | apps/docmost-engine-postgres/ |
|
||
| 21:45 | **R-621 demonstrated itself on this leg's KEY artefact** | The engine's refusal line was GONE before any probe could read it — `failAndHold` had removed the container, and the controller log does not carry it either. Reproduced independently instead (R-320) | 18-postgres-refusal-reproduced.txt |
|
||
| 21:45 | **My own first reproduction was WRONG, and is kept labelled** | The volume lookup returned empty -> the copy was empty -> postgres:17 initialised a FRESH datadir and started happily. A blank `PG_VERSION` should have stopped the step and did not. Rewritten so every step proves itself first. The accidental result (16 refusing a 17 datadir, verbatim) is real and kept | 17-postgres-refusal-reproduced.txt |
|
||
| 21:56 | **B1 — THE UNATTENDED HOLD. MEASURED, and it is the thing `09` Q4 has never had.** | The caller pressed ONCE with nobody watching; the drill image started, stayed up and never served; `safety-dump` -> `verifying` -> **`held after 312.9 s`**; then **pass 2 and pass 3 pressed NOTHING** — `outcomes={'glance': ('held', 312.9)} never_again=['glance']`. The hold sentence names the tier, the date and what the copy holds. **My `follow()` fix (R-623) was load-bearing: without it this would have read `timeout` after 900 s** | bad-days/B1-unattended-hold/ |
|
||
| 21:57 | **B2 — the new tag cannot be pulled** | **Exactly §6.1 Scenario E.** `phase=failed` in **1.0 s**, `hold=None`, „Az új verzió letöltése nem sikerült, ezért a frissítés elmaradt. Az alkalmazás a korábbi verzióval fut tovább." Nothing held; the app keeps running the old version | bad-days/B2-pull-fails/ |
|
||
| 21:57 | **B3 — Update pressed while a backup runs** | `409 reason='busy'` on six consecutive presses with „A frissítés most nem indítható: mentés/visszaállítás folyamatban. Próbáld újra, ha befejeződött." — the transient reason an unattended caller needs (R-609) | bad-days/B3-busy-during-backup/ |
|
||
| 21:58 | **B7 — the 2 GB disk floor** | **REFUSED correctly**: filled to 1.4 GB free, `409 reason='disk'`, „Nincs elég szabad hely a frissítéshez: 1.4 GB szabad… legalább 2 GB szükséges." Nothing moved. Space released immediately after | bad-days/B7-disk-and-memory-refusals/ |
|
||
| 21:58 | **CORRECTION to R-606** — the refusals are NOT English | Three refusals requested with `?lang=en`, Hungarian request as control: `held`, `not_deployed` and `disk` all come back **identical Hungarian**. The pipe exists (`langFor` honours `?lang=`) but the sentences are frozen string constants, so routing them through `errText` changes nothing | 20-refusals-in-english.txt |
|
||
| 22:00 | **B9 — a frozen app gets a newer `.felhom.yml`** | §5.4's asymmetry confirmed live; **no false alarm** — 10 samples, all `running`/200. The reason NARROWS R-458: `type: http` treats any response as healthy, so a wrong path is invisible; the risk is real only for `type: api` + `expect` | 21-B9-frozen-app-newer-felhomyml.md |
|
||
| 22:00 | B8 skipped here — docmost was swept by B1's precondition | re-run in `phase2_redo.sh` with a freshly deployed docmost | phase2_redo.sh |
|
||
| 22:02 | **B4 (two) — two Updates within one second** | **THERE IS NO SINGLE-FLIGHT.** Both accepted `202` within **0.268 s**, and **2 of 2 ran in flight at the same time**. Both ended `done`, no error, no hold, both pins advanced. That is a real Slice-6 design input: an unattended caller pressing N apps would run N updates at once | bad-days/B4-concurrent-updates/two.json |
|
||
| 22:06 | **B4 (five) — five Updates within half a second** | **5 of 5 ran concurrently and ALL ended honest**: accepted within **0.448 s**, every one `done`, no error, no hold, every pin advanced, whole batch ~30 s. So the box does not serialise updates at all — and on this box it coped. Memory samples in the record | bad-days/B4-concurrent-updates/five.json |
|
||
| 22:08 | **B5 (backing-up cut) — the phase nobody had cut in** | **BEHAVED.** Cut at `backing-up` +0.027 s (`pct stop` returned 9.65 s later — the usual instrument limit, stated). After boot the box said so ITSELF: „crash recovery: an app-data backup (volume dump) … was interrupted and left 1 app(s) stopped — restarting them: [privatebin]" and „update recovery: privatebin was interrupted in backing-up". **The pin did NOT move**, the app runs, and a fresh paste seeded+read back | bad-days/B5-backing-up/ |
|
||
| 22:08 | **B5 (safety-dump cut) — MISSED, recorded as a miss** | The phases went `backing-up` -> `pulling` -> `failed` in **0.473 s** and `safety-dump` was never observed, so the plug was never pulled. Recorded as a MISS, not as a pass. To be retried with a genuine pending edge | bad-days/B5-safety-dump/ |
|
||
| 22:09 | **Phase 4 — the morning after** | Every one of the **10 badges is TRUE** (refs equal <-> „Naprakész", refs differ <-> „Frissítés elérhető"). Q4's four promises all scored True on the held app's page. `zipline` shows the household „Nem egészséges — URL nem elérhető" while running — R-618 in the household's own words | bad-days/P4-morning-after/ |
|
||
| 22:09 | **B6 — the way out FORWARDS: REFUSED** | A held app met a FIXED newer version. Badge: „Frissítés elérhető — ma" in both languages, button offered; press -> **`409 reason='held'`**. Correct per §6.1 (only a restore lifts a hold) but **the page invites what the button refuses** — the exact inconsistency R-524 removed for the Ahead case. Filed as **R-625** | bad-days/B6-way-out-forwards/ |
|
||
| 22:16 | **PHASE 2.3 — the PostgreSQL conversion REHEARSAL, costed (Q5)** | **WORKED end to end.** 49 MB / 48 tables: dump with 16 **2.6 s / 132 201 B**; fresh 17 + replay **6.5 s / 48 tables**; the app said „Database connection successful"; **the seeded account read back on 17**; total **155.9 s**, of which ~9 s is engine work. Two benign ERROR lines named. `pg_upgrade` NOT run — it needs an image that does not exist here | 24-Q5-postgres-conversion-costed.md |
|
||
| 22:18 | **wger CONFIRMED — R-618's last candidate closes** | probe `type: http port: 80`; inside the container **port 80 refused, port 8000 ANSWERED**; docker health green; front door 302; the box says `unhealthy`. **Three confirmed instances now: tandoor, zipline, wger** — and the cheap static rule finds all three with one false positive out of 53 | 22-wger-probe-measured.txt |
|
||
| 22:19 | **B5 (safety-dump cut) — the second early phase, RETRIED and HIT** | Cut at `safety-dump` +0.023 s. After boot: **no recovery line and no journal** (unlike the `backing-up` cut, which produced both), **the pin did NOT move**, the app runs, a fresh paste seeded and read back, and the card shows **no interrupted sentence at all**. No half-written backup artefact on either cut | 25-B5-the-two-early-cuts.md |
|
||
| 22:20 | **Bonus proof across a REAL power cut** | The boot sweep met the held `glance` after an unclean shutdown and **deliberately left it alone**: „is a boot orphan by intent but is HELD … NOT starting it; whatever is holding it owns its recovery". §6.1's three-unattended-paths claim, proven across a power cut | 25-B5-the-two-early-cuts.md |
|
||
| 22:22 | **mealie v3.20.1 -> v3.27.0** (db-postgres) | **PROVEN** — the edge that failed twice on instrument problems, walked cleanly on the third | apps/mealie/ |
|
||
| 22:23 | **PHASE 1+2 COMPLETE** | **21 edges attempted: 14 proven, 3 failed, 4 inconclusive.** Ten of the fourteen printed a verbatim migration line. Up from the **three** apps this project had ever measured | summarise.py |
|
||
| 22:27 | **PHASE 5 — teardown, three layers plus Gitea** | **Machine:** every throwaway app removed through the product; three with an `HDD_PATH` refused at „remove with data" (R-442, correct) and removed keeping it; registry, drill images and drill volumes gone BY NAME, never pruned; `controller.yaml` restored and `git.repo_url` **read back as the LIVE catalog with an empty token**; only `felhom-controller`, `filebrowser`, `traefik` left. **Host:** no harness LXC was created, so none was destroyed; guest 9201 untouched. **Hub:** nothing provisioned; floor 0.261.0. **Gitea:** drill repo KEPT, private, reset to live `main`; **every `image:` line IDENTICAL** | teardown/ |
|
||
| 22:28 | **The teardown found the night's LAST defect** | A `navidrome` container the product's own removal left behind, resurrected by Docker's restart policy at a power cut, invisible to every sweep keying on `deployed`. Removed by name. **R-626** | 26-removed-app-came-back.txt |
|
||
| 22:4x | **Documents, register, CI** | 12 new rows (register **303 -> 315**), 8 existing rows updated, `09` §3b answered for Q2–Q6, §6.4 re-costed, §6.5 (the drill method) added, limitation 8 replaced, capability map widened, rotation **2 -> 16** fully-walked apps, STATUS morning note. **CI green by id: catalog job 830, felhom.eu job 870, both `success`** | felhom.eu@8d786f7940b0 |
|
||
| — | **`unproven.py --summary` did NOT move** | 55 claims, 35 not walked — unchanged. Tonight measured the UPDATE arc, which that file tracks as one claim already marked walked; no claim's status changed, so no number moved. Stated because the checklist asks | scripts/unproven.py |
|