Files
felhom.eu/documentation/audits/update-night-2026-09-21/PROGRESS.md
T
admin 1ee14ce166
gates / gates (push) Successful in 24s
Update night: the final PROGRESS lines — teardown, documents, CI green by id
Both CI runs confirmed by job id: app-catalog-felhom.eu job 830 and felhom.eu job 870, each
'completed' / 'success', matched on head_sha. unproven.py --summary did not move and that is
stated rather than left for the reader to notice.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-21 22:35:42 +02:00

18 KiB
Raw Blame History

UPDATE NIGHT 2026-09-21 — progress log (one line per finished step)

Evidence dir: documentation/audits/update-night-2026-09-21/ A resuming session reads THIS FILE FIRST and never repeats a finished step.

time (CEST) step verdict evidence
20:07 P0.1 floor → 0.261.0 (declared MinAgent 0.131.0) PROVEN — both demo boxes arrived in 13 s; hub managed floor SERVED … from declared (golden 0.258.0); drill-r50 stays held (agent 0.129.0 < 0.131.0, host DOWN) 01-floor-pre.txt, 02-floor-save.txt
20:08 P0.2a drill repo admin/app-catalog-drill created (private) DONE — Gitea token lacks write:user, so POST /api/v1/repos/migrate was used instead of /user/repos; main = f5f6a152b513 = live 03-drill-repo.txt
20:10 P0.2b 9202 repointed at the drill repo PROVEN — git.repo_url ALONE IS INERT: gitCloneOrPull only clones when .git is absent, else fetches from the old origin. Cache dir had to be removed too. Config saved as controller.yaml.pre-update-night 04-9202-config-pre.txt, 05-9202-follows-drill.txt
20:12 P0.2c positive + negative controls PROVEN — drill bump uptime-kuma 2.4.0→2.5.5 shows on 9202 as „Frissítés elérhető — ma" / "Update available — today"; live catalog still f5f6a152b513 with pin 2.4.0; both real boxes' caches still at f5f6a15. R-607 fired again (sync said „nincs változás"). App route is /apps/<n>, not /app/<n> 06-drift-rerun.txt, 07-positive-control.txt
20:07 Drift re-run CONFIRMED — 66 pins, 46 behind, 39 within-major, 7 across-major — the brief's numbers hold exactly 06-drift-rerun.txt
20:13 P0.4 capacity measured OK — 9202: 26 GB RAM (22 free), docker root on mp0 with 56 GB free, drive 872 GB. demo-hp / is 86% but holds neither. Brief's capacity claim HOLDS 08-capacity.txt
20:18 P0.3 throwaway image store PROVEN — registry:2 on 9202 127.0.0.1:5000; drill/glance:1.0.0 = real v0.8.6 retagged, :1.0.1 = starts/stays up/never serves, :1.0.2 absent (404). CompareImageRefs DOES order host:port/ refs — 4 positive + 1 negative control, run not read; tmp test deleted, tree clean 09-image-store.txt
20:20 EDGE privatebin 2.0.5 -> 2.0.6 (file-leg) PROVEN — 15.4 s; seed read back both sides; all four observables agree apps/privatebin/
20:25 EDGE docmost 0.95.0 -> 0.96.0 (db-postgres, engine constant) PROVEN — 103.5 s; seeded account authenticated after; all four observables agree apps/docmost/
20:28 EDGE bookstack 26.05.2 -> 26.05.5 (db-mariadb, engine constant) PROVEN — 45.1 s; artisan readback with its own negative control; DATABASE HALF ONLY (R-460) apps/bookstack/
20:37 FINDING R-618 tandoor probe port DEFECT, measured — probe 8080, app listens on 80 only; docker healthy + front door 200 + controller unhealthy. 53-template sweep run: wger suspected, adventurelog cleared 10-probe-port-sweep.txt
20:38 FINDINGS R-615/616/617/619 filed — repo_url inert; catalog token plaintext in the clone; Gitea migrate endpoint; type: password mandatory though the wire says optional NEW-ROWS.md
20:39 EDGE actualbudget 26.7.0 -> 26.9.0 PROVEN — 19.5 s apps/actualbudget/
20:44 EDGE navidrome 0.63.2 -> 0.64.0 (file-leg, HDD_PATH) PROVEN — 11.3 s apps/navidrome/
20:46 EDGE audiobookshelf 2.35.1 -> 2.36.1 (file-leg) PROVEN — 23.6 s apps/audiobookshelf/
20:54 EDGE adventurelog v0.12.1 -> v0.13.0 (db-postgis) FAILED — the most valuable result so far. Nine migrations applied OK, then the app never bound its port; held after the full 5-min wait; the sentence names tier, date and what the copy holds. Rows R-621 (the hold destroys the failure evidence) and R-622 (do not promote this edge) apps/adventurelog/
20:59 Harness code pushed to the catalog — 4 fixtures + edges U1..U7 DONE, gates green — live catalog main moves f5f6a152b513 -> 4463243f2e09, scripts/ ONLY, zero image: lines (the teardown diff still expects every image line identical). Harness RUNS owed app-catalog-felhom.eu@4463243f2e09
20:59 image reclaim BY NAME (no prune) 34 unused images removed by exact reference; 40 GB -> 55 GB free reclaim.sh
21:08 CI check for the catalog push GREEN — job id=830, name='gates', status='completed', conclusion='success', matched on head_sha=4463243f2e09. Confirms CLAUDE.md's warning: page 1's newest id was 52, the real newest was 830 — job ids are NOT page-ordered scratchpad/ci.sh
21:12 Observation, NOT a new row — an app with an HDD_PATH cannot have its DATA removed on 9202 R-442's guard working as designed: /api/disks answers agent not configured on this guest, so the drive path cannot be RESOLVED and the removal is refused with the app kept rather than half-deleted. The household's other choice — remove the app, KEEP the data — is accepted. The harness now takes that route and tidies its own directories by name at teardown apps/navidrome/, apps/romm/
21:13 FINDING R-623 — the unattended caller turned every SUCCESS into a timeout Read before use, not after: follow() read update_phase/updating off the API ENVELOPE, so both were always None, every followed update hit the 900 s timeout and was then marked never-press-again. Fixed in that file before B1 relied on it. The earlier night missed it because its only follow() pass was the one whose log was lost update-arc-gaps-2026-09-21/unattended-caller.py
21:18 Interim commit pushed — felhom.eu da20722e76a7 Phase 0 + Phase 1 evidence off the machine before Phases 2-5 (R-320). repo_gates.py --fast: all 15 OK. Includes the two instrument fixes (api-recipe route, unattended-caller follow()) felhom.eu@da20722e76a7
21:18 R-618 RAISED TO P1 by a live measurement tandoor 2.6.13→2.6.15 was serving HTTP 200 on the NEW version at 21:18:28 while the update sat in verifying, because the probe names port 8080 and the app listens on 80. The health wait then times out and failAndHold STOPS the working app. A wrong port turns every successful update into an outage plus a needless restore 14-tandoor-serving-while-verifying.txt
21:22 EDGE tandoor 2.6.13 -> 2.6.15 (db-postgres) FAILED — and the failure is OURS, not the app's. The new version was SERVING 200 with a green docker healthcheck at four samples across five minutes; the controller's probe names port 8080 and the app listens on 80, so verifying timed out and failAndHold STOPPED a working app. The four observables disagree exactly as 09 §5.2 predicts. R-618 is now P1 14-tandoor-serving-while-verifying.txt, apps/tandoor/
21:35 failwalk adventurelog — the household's WAY OUT, on a real failed edge PROVEN — the restore the hold sentence names („helyi", 1 copy offered) ran in 75 s; app back on v0.12.1, all three containers running, hold CLEARED, state running. R-606 CONFIRMED on the hold sentence itself (English page, Hungarian sentence) with positive and negative controls apps/adventurelog/failwalk.json, 15-r606-*.txt
21:36 Observation after the restore The hold sentence is GONE from the card (R-480 working) and the app reads „Fut". The badge briefly read „Frissítés elérhető — 65 napja" rather than „ma", because the restore brought back the recovery unit's own .felhom.yml with its pre-drill catalog_since; the syncer copies that file verbatim every cycle, so it heals within one sync. Recorded as an observation, not a row apps/adventurelog/failwalk.json
21:37 failwalk tandoor — the way out of the WRONG-PROBE hold PROVEN — restore in 32.4 s, hold cleared, app running, data reads back. Note what this means for R-618: the household's route back works, so nothing is LOST — but they had to walk it for an update that had actually succeeded apps/tandoor/failwalk.json
21:40 PHASE 2.1 — MariaDB 11.6 -> 12.3 through the REAL Update button PROVEN, a first. All four SPIKE-r459 observables: datadir 11.6.2->12.3.3; the engine's own check says „already upgraded … no need to run mariadb-upgrade again"; the entrypoint says „Major version upgrade detected … Check required!" then started and finished (NOT the skipped due to $MARIADB_AUTO_UPGRADE line R-459 feared); the engine's own pre-upgrade backup system_mysql_backup_11.6.2-MariaDB.sql.zst 631 905 B. Seeded Nextcloud account read back. All four version observables agree 16-phase2.1-mariadb-major.md, apps/nextcloud-engine-mariadb/
21:45 PHASE 2.2 — PostgreSQL 16 -> 17 through the real Update button FAILED, exactly as R-463 predicted and nobody had measured. The update ended in 5.1 s, the app was stopped and held, and pinned named postgres:17-alpine while installed still said 16 and nothing was running — the brief's own question answered. The restore then brought it back in 29.1 s (hold CLEARED, health probe 200) apps/docmost-engine-postgres/
21:45 R-621 demonstrated itself on this leg's KEY artefact The engine's refusal line was GONE before any probe could read it — failAndHold had removed the container, and the controller log does not carry it either. Reproduced independently instead (R-320) 18-postgres-refusal-reproduced.txt
21:45 My own first reproduction was WRONG, and is kept labelled The volume lookup returned empty -> the copy was empty -> postgres:17 initialised a FRESH datadir and started happily. A blank PG_VERSION should have stopped the step and did not. Rewritten so every step proves itself first. The accidental result (16 refusing a 17 datadir, verbatim) is real and kept 17-postgres-refusal-reproduced.txt
21:56 B1 — THE UNATTENDED HOLD. MEASURED, and it is the thing 09 Q4 has never had. The caller pressed ONCE with nobody watching; the drill image started, stayed up and never served; safety-dump -> verifying -> held after 312.9 s; then pass 2 and pass 3 pressed NOTHING — outcomes={'glance': ('held', 312.9)} never_again=['glance']. The hold sentence names the tier, the date and what the copy holds. My follow() fix (R-623) was load-bearing: without it this would have read timeout after 900 s bad-days/B1-unattended-hold/
21:57 B2 — the new tag cannot be pulled Exactly §6.1 Scenario E. phase=failed in 1.0 s, hold=None, „Az új verzió letöltése nem sikerült, ezért a frissítés elmaradt. Az alkalmazás a korábbi verzióval fut tovább." Nothing held; the app keeps running the old version bad-days/B2-pull-fails/
21:57 B3 — Update pressed while a backup runs 409 reason='busy' on six consecutive presses with „A frissítés most nem indítható: mentés/visszaállítás folyamatban. Próbáld újra, ha befejeződött." — the transient reason an unattended caller needs (R-609) bad-days/B3-busy-during-backup/
21:58 B7 — the 2 GB disk floor REFUSED correctly: filled to 1.4 GB free, 409 reason='disk', „Nincs elég szabad hely a frissítéshez: 1.4 GB szabad… legalább 2 GB szükséges." Nothing moved. Space released immediately after bad-days/B7-disk-and-memory-refusals/
21:58 CORRECTION to R-606 — the refusals are NOT English Three refusals requested with ?lang=en, Hungarian request as control: held, not_deployed and disk all come back identical Hungarian. The pipe exists (langFor honours ?lang=) but the sentences are frozen string constants, so routing them through errText changes nothing 20-refusals-in-english.txt
22:00 B9 — a frozen app gets a newer .felhom.yml §5.4's asymmetry confirmed live; no false alarm — 10 samples, all running/200. The reason NARROWS R-458: type: http treats any response as healthy, so a wrong path is invisible; the risk is real only for type: api + expect 21-B9-frozen-app-newer-felhomyml.md
22:00 B8 skipped here — docmost was swept by B1's precondition re-run in phase2_redo.sh with a freshly deployed docmost phase2_redo.sh
22:02 B4 (two) — two Updates within one second THERE IS NO SINGLE-FLIGHT. Both accepted 202 within 0.268 s, and 2 of 2 ran in flight at the same time. Both ended done, no error, no hold, both pins advanced. That is a real Slice-6 design input: an unattended caller pressing N apps would run N updates at once bad-days/B4-concurrent-updates/two.json
22:06 B4 (five) — five Updates within half a second 5 of 5 ran concurrently and ALL ended honest: accepted within 0.448 s, every one done, no error, no hold, every pin advanced, whole batch ~30 s. So the box does not serialise updates at all — and on this box it coped. Memory samples in the record bad-days/B4-concurrent-updates/five.json
22:08 B5 (backing-up cut) — the phase nobody had cut in BEHAVED. Cut at backing-up +0.027 s (pct stop returned 9.65 s later — the usual instrument limit, stated). After boot the box said so ITSELF: „crash recovery: an app-data backup (volume dump) … was interrupted and left 1 app(s) stopped — restarting them: [privatebin]" and „update recovery: privatebin was interrupted in backing-up". The pin did NOT move, the app runs, and a fresh paste seeded+read back bad-days/B5-backing-up/
22:08 B5 (safety-dump cut) — MISSED, recorded as a miss The phases went backing-up -> pulling -> failed in 0.473 s and safety-dump was never observed, so the plug was never pulled. Recorded as a MISS, not as a pass. To be retried with a genuine pending edge bad-days/B5-safety-dump/
22:09 Phase 4 — the morning after Every one of the 10 badges is TRUE (refs equal <-> „Naprakész", refs differ <-> „Frissítés elérhető"). Q4's four promises all scored True on the held app's page. zipline shows the household „Nem egészséges — URL nem elérhető" while running — R-618 in the household's own words bad-days/P4-morning-after/
22:09 B6 — the way out FORWARDS: REFUSED A held app met a FIXED newer version. Badge: „Frissítés elérhető — ma" in both languages, button offered; press -> 409 reason='held'. Correct per §6.1 (only a restore lifts a hold) but the page invites what the button refuses — the exact inconsistency R-524 removed for the Ahead case. Filed as R-625 bad-days/B6-way-out-forwards/
22:16 PHASE 2.3 — the PostgreSQL conversion REHEARSAL, costed (Q5) WORKED end to end. 49 MB / 48 tables: dump with 16 2.6 s / 132 201 B; fresh 17 + replay 6.5 s / 48 tables; the app said „Database connection successful"; the seeded account read back on 17; total 155.9 s, of which ~9 s is engine work. Two benign ERROR lines named. pg_upgrade NOT run — it needs an image that does not exist here 24-Q5-postgres-conversion-costed.md
22:18 wger CONFIRMED — R-618's last candidate closes probe type: http port: 80; inside the container port 80 refused, port 8000 ANSWERED; docker health green; front door 302; the box says unhealthy. Three confirmed instances now: tandoor, zipline, wger — and the cheap static rule finds all three with one false positive out of 53 22-wger-probe-measured.txt
22:19 B5 (safety-dump cut) — the second early phase, RETRIED and HIT Cut at safety-dump +0.023 s. After boot: no recovery line and no journal (unlike the backing-up cut, which produced both), the pin did NOT move, the app runs, a fresh paste seeded and read back, and the card shows no interrupted sentence at all. No half-written backup artefact on either cut 25-B5-the-two-early-cuts.md
22:20 Bonus proof across a REAL power cut The boot sweep met the held glance after an unclean shutdown and deliberately left it alone: „is a boot orphan by intent but is HELD … NOT starting it; whatever is holding it owns its recovery". §6.1's three-unattended-paths claim, proven across a power cut 25-B5-the-two-early-cuts.md
22:22 mealie v3.20.1 -> v3.27.0 (db-postgres) PROVEN — the edge that failed twice on instrument problems, walked cleanly on the third apps/mealie/
22:23 PHASE 1+2 COMPLETE 21 edges attempted: 14 proven, 3 failed, 4 inconclusive. Ten of the fourteen printed a verbatim migration line. Up from the three apps this project had ever measured summarise.py
22:27 PHASE 5 — teardown, three layers plus Gitea Machine: every throwaway app removed through the product; three with an HDD_PATH refused at „remove with data" (R-442, correct) and removed keeping it; registry, drill images and drill volumes gone BY NAME, never pruned; controller.yaml restored and git.repo_url read back as the LIVE catalog with an empty token; only felhom-controller, filebrowser, traefik left. Host: no harness LXC was created, so none was destroyed; guest 9201 untouched. Hub: nothing provisioned; floor 0.261.0. Gitea: drill repo KEPT, private, reset to live main; every image: line IDENTICAL teardown/
22:28 The teardown found the night's LAST defect A navidrome container the product's own removal left behind, resurrected by Docker's restart policy at a power cut, invisible to every sweep keying on deployed. Removed by name. R-626 26-removed-app-came-back.txt
22:4x Documents, register, CI 12 new rows (register 303 -> 315), 8 existing rows updated, 09 §3b answered for Q2–Q6, §6.4 re-costed, §6.5 (the drill method) added, limitation 8 replaced, capability map widened, rotation 2 -> 16 fully-walked apps, STATUS morning note. CI green by id: catalog job 830, felhom.eu job 870, both success felhom.eu@8d786f7940b0
— unproven.py --summary did NOT move 55 claims, 35 not walked — unchanged. Tonight measured the UPDATE arc, which that file tracks as one claim already marked walked; no claim's status changed, so no number moved. Stated because the checklist asks scripts/unproven.py