1ee14ce166
gates / gates (push) Successful in 24s
Both CI runs confirmed by job id: app-catalog-felhom.eu job 830 and felhom.eu job 870, each 'completed' / 'success', matched on head_sha. unproven.py --summary did not move and that is stated rather than left for the reader to notice. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
18 KiB
18 KiB
UPDATE NIGHT 2026-09-21 — progress log (one line per finished step)
Evidence dir: documentation/audits/update-night-2026-09-21/
A resuming session reads THIS FILE FIRST and never repeats a finished step.
| time (CEST) | step | verdict | evidence |
|---|---|---|---|
| 20:07 | P0.1 floor → 0.261.0 (declared MinAgent 0.131.0) | PROVEN — both demo boxes arrived in 13 s; hub managed floor SERVED … from declared (golden 0.258.0); drill-r50 stays held (agent 0.129.0 < 0.131.0, host DOWN) |
01-floor-pre.txt, 02-floor-save.txt |
| 20:08 | P0.2a drill repo admin/app-catalog-drill created (private) |
DONE — Gitea token lacks write:user, so POST /api/v1/repos/migrate was used instead of /user/repos; main = f5f6a152b513 = live |
03-drill-repo.txt |
| 20:10 | P0.2b 9202 repointed at the drill repo | PROVEN — git.repo_url ALONE IS INERT: gitCloneOrPull only clones when .git is absent, else fetches from the old origin. Cache dir had to be removed too. Config saved as controller.yaml.pre-update-night |
04-9202-config-pre.txt, 05-9202-follows-drill.txt |
| 20:12 | P0.2c positive + negative controls | PROVEN — drill bump uptime-kuma 2.4.0→2.5.5 shows on 9202 as „Frissítés elérhető — ma" / "Update available — today"; live catalog still f5f6a152b513 with pin 2.4.0; both real boxes' caches still at f5f6a15. R-607 fired again (sync said „nincs változás"). App route is /apps/<n>, not /app/<n> |
06-drift-rerun.txt, 07-positive-control.txt |
| 20:07 | Drift re-run | CONFIRMED — 66 pins, 46 behind, 39 within-major, 7 across-major — the brief's numbers hold exactly | 06-drift-rerun.txt |
| 20:13 | P0.4 capacity measured | OK — 9202: 26 GB RAM (22 free), docker root on mp0 with 56 GB free, drive 872 GB. demo-hp / is 86% but holds neither. Brief's capacity claim HOLDS |
08-capacity.txt |
| 20:18 | P0.3 throwaway image store | PROVEN — registry:2 on 9202 127.0.0.1:5000; drill/glance:1.0.0 = real v0.8.6 retagged, :1.0.1 = starts/stays up/never serves, :1.0.2 absent (404). CompareImageRefs DOES order host:port/ refs — 4 positive + 1 negative control, run not read; tmp test deleted, tree clean |
09-image-store.txt |
| 20:20 | EDGE privatebin 2.0.5 -> 2.0.6 (file-leg) | PROVEN — 15.4 s; seed read back both sides; all four observables agree | apps/privatebin/ |
| 20:25 | EDGE docmost 0.95.0 -> 0.96.0 (db-postgres, engine constant) | PROVEN — 103.5 s; seeded account authenticated after; all four observables agree | apps/docmost/ |
| 20:28 | EDGE bookstack 26.05.2 -> 26.05.5 (db-mariadb, engine constant) | PROVEN — 45.1 s; artisan readback with its own negative control; DATABASE HALF ONLY (R-460) | apps/bookstack/ |
| 20:37 | FINDING R-618 tandoor probe port | DEFECT, measured — probe 8080, app listens on 80 only; docker healthy + front door 200 + controller unhealthy. 53-template sweep run: wger suspected, adventurelog cleared |
10-probe-port-sweep.txt |
| 20:38 | FINDINGS R-615/616/617/619 | filed — repo_url inert; catalog token plaintext in the clone; Gitea migrate endpoint; type: password mandatory though the wire says optional |
NEW-ROWS.md |
| 20:39 | EDGE actualbudget 26.7.0 -> 26.9.0 | PROVEN — 19.5 s | apps/actualbudget/ |
| 20:44 | EDGE navidrome 0.63.2 -> 0.64.0 (file-leg, HDD_PATH) | PROVEN — 11.3 s | apps/navidrome/ |
| 20:46 | EDGE audiobookshelf 2.35.1 -> 2.36.1 (file-leg) | PROVEN — 23.6 s | apps/audiobookshelf/ |
| 20:54 | EDGE adventurelog v0.12.1 -> v0.13.0 (db-postgis) | FAILED — the most valuable result so far. Nine migrations applied OK, then the app never bound its port; held after the full 5-min wait; the sentence names tier, date and what the copy holds. Rows R-621 (the hold destroys the failure evidence) and R-622 (do not promote this edge) | apps/adventurelog/ |
| 20:59 | Harness code pushed to the catalog — 4 fixtures + edges U1..U7 | DONE, gates green — live catalog main moves f5f6a152b513 -> 4463243f2e09, scripts/ ONLY, zero image: lines (the teardown diff still expects every image line identical). Harness RUNS owed |
app-catalog-felhom.eu@4463243f2e09 |
| 20:59 | image reclaim BY NAME (no prune) | 34 unused images removed by exact reference; 40 GB -> 55 GB free | reclaim.sh |
| 21:08 | CI check for the catalog push | GREEN — job id=830, name='gates', status='completed', conclusion='success', matched on head_sha=4463243f2e09. Confirms CLAUDE.md's warning: page 1's newest id was 52, the real newest was 830 — job ids are NOT page-ordered |
scratchpad/ci.sh |
| 21:12 | Observation, NOT a new row — an app with an HDD_PATH cannot have its DATA removed on 9202 |
R-442's guard working as designed: /api/disks answers agent not configured on this guest, so the drive path cannot be RESOLVED and the removal is refused with the app kept rather than half-deleted. The household's other choice — remove the app, KEEP the data — is accepted. The harness now takes that route and tidies its own directories by name at teardown |
apps/navidrome/, apps/romm/ |
| 21:13 | FINDING R-623 — the unattended caller turned every SUCCESS into a timeout |
Read before use, not after: follow() read update_phase/updating off the API ENVELOPE, so both were always None, every followed update hit the 900 s timeout and was then marked never-press-again. Fixed in that file before B1 relied on it. The earlier night missed it because its only follow() pass was the one whose log was lost |
update-arc-gaps-2026-09-21/unattended-caller.py |
| 21:18 | Interim commit pushed — felhom.eu da20722e76a7 |
Phase 0 + Phase 1 evidence off the machine before Phases 2-5 (R-320). repo_gates.py --fast: all 15 OK. Includes the two instrument fixes (api-recipe route, unattended-caller follow()) |
felhom.eu@da20722e76a7 |
| 21:18 | R-618 RAISED TO P1 by a live measurement | tandoor 2.6.13→2.6.15 was serving HTTP 200 on the NEW version at 21:18:28 while the update sat in verifying, because the probe names port 8080 and the app listens on 80. The health wait then times out and failAndHold STOPS the working app. A wrong port turns every successful update into an outage plus a needless restore |
14-tandoor-serving-while-verifying.txt |
| 21:22 | EDGE tandoor 2.6.13 -> 2.6.15 (db-postgres) | FAILED — and the failure is OURS, not the app's. The new version was SERVING 200 with a green docker healthcheck at four samples across five minutes; the controller's probe names port 8080 and the app listens on 80, so verifying timed out and failAndHold STOPPED a working app. The four observables disagree exactly as 09 §5.2 predicts. R-618 is now P1 |
14-tandoor-serving-while-verifying.txt, apps/tandoor/ |
| 21:35 | failwalk adventurelog — the household's WAY OUT, on a real failed edge | PROVEN — the restore the hold sentence names („helyi", 1 copy offered) ran in 75 s; app back on v0.12.1, all three containers running, hold CLEARED, state running. R-606 CONFIRMED on the hold sentence itself (English page, Hungarian sentence) with positive and negative controls |
apps/adventurelog/failwalk.json, 15-r606-*.txt |
| 21:36 | Observation after the restore | The hold sentence is GONE from the card (R-480 working) and the app reads „Fut". The badge briefly read „Frissítés elérhető — 65 napja" rather than „ma", because the restore brought back the recovery unit's own .felhom.yml with its pre-drill catalog_since; the syncer copies that file verbatim every cycle, so it heals within one sync. Recorded as an observation, not a row |
apps/adventurelog/failwalk.json |
| 21:37 | failwalk tandoor — the way out of the WRONG-PROBE hold | PROVEN — restore in 32.4 s, hold cleared, app running, data reads back. Note what this means for R-618: the household's route back works, so nothing is LOST — but they had to walk it for an update that had actually succeeded | apps/tandoor/failwalk.json |
| 21:40 | PHASE 2.1 — MariaDB 11.6 -> 12.3 through the REAL Update button | PROVEN, a first. All four SPIKE-r459 observables: datadir 11.6.2->12.3.3; the engine's own check says „already upgraded … no need to run mariadb-upgrade again"; the entrypoint says „Major version upgrade detected … Check required!" then started and finished (NOT the skipped due to $MARIADB_AUTO_UPGRADE line R-459 feared); the engine's own pre-upgrade backup system_mysql_backup_11.6.2-MariaDB.sql.zst 631 905 B. Seeded Nextcloud account read back. All four version observables agree |
16-phase2.1-mariadb-major.md, apps/nextcloud-engine-mariadb/ |
| 21:45 | PHASE 2.2 — PostgreSQL 16 -> 17 through the real Update button | FAILED, exactly as R-463 predicted and nobody had measured. The update ended in 5.1 s, the app was stopped and held, and pinned named postgres:17-alpine while installed still said 16 and nothing was running — the brief's own question answered. The restore then brought it back in 29.1 s (hold CLEARED, health probe 200) |
apps/docmost-engine-postgres/ |
| 21:45 | R-621 demonstrated itself on this leg's KEY artefact | The engine's refusal line was GONE before any probe could read it — failAndHold had removed the container, and the controller log does not carry it either. Reproduced independently instead (R-320) |
18-postgres-refusal-reproduced.txt |
| 21:45 | My own first reproduction was WRONG, and is kept labelled | The volume lookup returned empty -> the copy was empty -> postgres:17 initialised a FRESH datadir and started happily. A blank PG_VERSION should have stopped the step and did not. Rewritten so every step proves itself first. The accidental result (16 refusing a 17 datadir, verbatim) is real and kept |
17-postgres-refusal-reproduced.txt |
| 21:56 | B1 — THE UNATTENDED HOLD. MEASURED, and it is the thing 09 Q4 has never had. |
The caller pressed ONCE with nobody watching; the drill image started, stayed up and never served; safety-dump -> verifying -> held after 312.9 s; then pass 2 and pass 3 pressed NOTHING — outcomes={'glance': ('held', 312.9)} never_again=['glance']. The hold sentence names the tier, the date and what the copy holds. My follow() fix (R-623) was load-bearing: without it this would have read timeout after 900 s |
bad-days/B1-unattended-hold/ |
| 21:57 | B2 — the new tag cannot be pulled | Exactly §6.1 Scenario E. phase=failed in 1.0 s, hold=None, „Az új verzió letöltése nem sikerült, ezért a frissítés elmaradt. Az alkalmazás a korábbi verzióval fut tovább." Nothing held; the app keeps running the old version |
bad-days/B2-pull-fails/ |
| 21:57 | B3 — Update pressed while a backup runs | 409 reason='busy' on six consecutive presses with „A frissítés most nem indítható: mentés/visszaállítás folyamatban. Próbáld újra, ha befejeződött." — the transient reason an unattended caller needs (R-609) |
bad-days/B3-busy-during-backup/ |
| 21:58 | B7 — the 2 GB disk floor | REFUSED correctly: filled to 1.4 GB free, 409 reason='disk', „Nincs elég szabad hely a frissítéshez: 1.4 GB szabad… legalább 2 GB szükséges." Nothing moved. Space released immediately after |
bad-days/B7-disk-and-memory-refusals/ |
| 21:58 | CORRECTION to R-606 — the refusals are NOT English | Three refusals requested with ?lang=en, Hungarian request as control: held, not_deployed and disk all come back identical Hungarian. The pipe exists (langFor honours ?lang=) but the sentences are frozen string constants, so routing them through errText changes nothing |
20-refusals-in-english.txt |
| 22:00 | B9 — a frozen app gets a newer .felhom.yml |
§5.4's asymmetry confirmed live; no false alarm — 10 samples, all running/200. The reason NARROWS R-458: type: http treats any response as healthy, so a wrong path is invisible; the risk is real only for type: api + expect |
21-B9-frozen-app-newer-felhomyml.md |
| 22:00 | B8 skipped here — docmost was swept by B1's precondition | re-run in phase2_redo.sh with a freshly deployed docmost |
phase2_redo.sh |
| 22:02 | B4 (two) — two Updates within one second | THERE IS NO SINGLE-FLIGHT. Both accepted 202 within 0.268 s, and 2 of 2 ran in flight at the same time. Both ended done, no error, no hold, both pins advanced. That is a real Slice-6 design input: an unattended caller pressing N apps would run N updates at once |
bad-days/B4-concurrent-updates/two.json |
| 22:06 | B4 (five) — five Updates within half a second | 5 of 5 ran concurrently and ALL ended honest: accepted within 0.448 s, every one done, no error, no hold, every pin advanced, whole batch ~30 s. So the box does not serialise updates at all — and on this box it coped. Memory samples in the record |
bad-days/B4-concurrent-updates/five.json |
| 22:08 | B5 (backing-up cut) — the phase nobody had cut in | BEHAVED. Cut at backing-up +0.027 s (pct stop returned 9.65 s later — the usual instrument limit, stated). After boot the box said so ITSELF: „crash recovery: an app-data backup (volume dump) … was interrupted and left 1 app(s) stopped — restarting them: [privatebin]" and „update recovery: privatebin was interrupted in backing-up". The pin did NOT move, the app runs, and a fresh paste seeded+read back |
bad-days/B5-backing-up/ |
| 22:08 | B5 (safety-dump cut) — MISSED, recorded as a miss | The phases went backing-up -> pulling -> failed in 0.473 s and safety-dump was never observed, so the plug was never pulled. Recorded as a MISS, not as a pass. To be retried with a genuine pending edge |
bad-days/B5-safety-dump/ |
| 22:09 | Phase 4 — the morning after | Every one of the 10 badges is TRUE (refs equal <-> „Naprakész", refs differ <-> „Frissítés elérhető"). Q4's four promises all scored True on the held app's page. zipline shows the household „Nem egészséges — URL nem elérhető" while running — R-618 in the household's own words |
bad-days/P4-morning-after/ |
| 22:09 | B6 — the way out FORWARDS: REFUSED | A held app met a FIXED newer version. Badge: „Frissítés elérhető — ma" in both languages, button offered; press -> 409 reason='held'. Correct per §6.1 (only a restore lifts a hold) but the page invites what the button refuses — the exact inconsistency R-524 removed for the Ahead case. Filed as R-625 |
bad-days/B6-way-out-forwards/ |
| 22:16 | PHASE 2.3 — the PostgreSQL conversion REHEARSAL, costed (Q5) | WORKED end to end. 49 MB / 48 tables: dump with 16 2.6 s / 132 201 B; fresh 17 + replay 6.5 s / 48 tables; the app said „Database connection successful"; the seeded account read back on 17; total 155.9 s, of which ~9 s is engine work. Two benign ERROR lines named. pg_upgrade NOT run — it needs an image that does not exist here |
24-Q5-postgres-conversion-costed.md |
| 22:18 | wger CONFIRMED — R-618's last candidate closes | probe type: http port: 80; inside the container port 80 refused, port 8000 ANSWERED; docker health green; front door 302; the box says unhealthy. Three confirmed instances now: tandoor, zipline, wger — and the cheap static rule finds all three with one false positive out of 53 |
22-wger-probe-measured.txt |
| 22:19 | B5 (safety-dump cut) — the second early phase, RETRIED and HIT | Cut at safety-dump +0.023 s. After boot: no recovery line and no journal (unlike the backing-up cut, which produced both), the pin did NOT move, the app runs, a fresh paste seeded and read back, and the card shows no interrupted sentence at all. No half-written backup artefact on either cut |
25-B5-the-two-early-cuts.md |
| 22:20 | Bonus proof across a REAL power cut | The boot sweep met the held glance after an unclean shutdown and deliberately left it alone: „is a boot orphan by intent but is HELD … NOT starting it; whatever is holding it owns its recovery". §6.1's three-unattended-paths claim, proven across a power cut |
25-B5-the-two-early-cuts.md |
| 22:22 | mealie v3.20.1 -> v3.27.0 (db-postgres) | PROVEN — the edge that failed twice on instrument problems, walked cleanly on the third | apps/mealie/ |
| 22:23 | PHASE 1+2 COMPLETE | 21 edges attempted: 14 proven, 3 failed, 4 inconclusive. Ten of the fourteen printed a verbatim migration line. Up from the three apps this project had ever measured | summarise.py |
| 22:27 | PHASE 5 — teardown, three layers plus Gitea | Machine: every throwaway app removed through the product; three with an HDD_PATH refused at „remove with data" (R-442, correct) and removed keeping it; registry, drill images and drill volumes gone BY NAME, never pruned; controller.yaml restored and git.repo_url read back as the LIVE catalog with an empty token; only felhom-controller, filebrowser, traefik left. Host: no harness LXC was created, so none was destroyed; guest 9201 untouched. Hub: nothing provisioned; floor 0.261.0. Gitea: drill repo KEPT, private, reset to live main; every image: line IDENTICAL |
teardown/ |
| 22:28 | The teardown found the night's LAST defect | A navidrome container the product's own removal left behind, resurrected by Docker's restart policy at a power cut, invisible to every sweep keying on deployed. Removed by name. R-626 |
26-removed-app-came-back.txt |
| 22:4x | Documents, register, CI | 12 new rows (register 303 -> 315), 8 existing rows updated, 09 §3b answered for Q2–Q6, §6.4 re-costed, §6.5 (the drill method) added, limitation 8 replaced, capability map widened, rotation 2 -> 16 fully-walked apps, STATUS morning note. CI green by id: catalog job 830, felhom.eu job 870, both success |
felhom.eu@8d786f7940b0 |
| — | unproven.py --summary did NOT move |
55 claims, 35 not walked — unchanged. Tonight measured the UPDATE arc, which that file tracks as one claim already marked walked; no claim's status changed, so no number moved. Stated because the checklist asks | scripts/unproven.py |