Files
felhom.eu/documentation/audits/update-night-2026-09-21/PROGRESS.md
T
admin 8d786f7940
gates / gates (push) Successful in 28s
Update night 2026-09-21: the full record, twelve rows, and the answers to five of the seven questions
The drill is complete. Teardown done in three layers plus Gitea; the live catalog's every `image:`
line is proven identical to before.

WHAT WAS MEASURED. 21 edges across 19 apps, on scratch guest 9202 through the product's own
guarded Update, against a PRIVATE DRILL CATALOG so the live catalog carried no test reference at
any point: 14 proven, 3 failed, 4 inconclusive. Each app seeded and read back through its OWN
front door, with a negative control on every readback. Ten of the fourteen printed a verbatim
migration line. Up from the three apps this project had ever measured.

THE RESULT THAT MATTERS. R-618, P1: three of the 53 templates name a health probe the app does not
answer, and because the guarded update WAITS on that same probe, a SUCCESSFUL update ends by
STOPPING a working app. tandoor was measured serving HTTP 200 on the new version at four samples
across five minutes, docker's own healthcheck green, and was then stopped and the household sent
to a restore they did not need. zipline and wger are the same defect, both confirmed live. The
gate that catches all three is static and cheap: both health checks already sit in the same file.

WHAT THE NIGHT ANSWERED that was open. The UNATTENDED HOLD (312.9 s, pressed once, never again) —
which needed a purpose-built image store, because the rule that makes automatic updates safe is
the same rule that refuses the obvious way to break one. MariaDB across a major through the real
button, all four observables, first time. PostgreSQL across a major, refusing exactly as predicted,
with the conversion costed at ~9 s of engine work. There is NO single-flight: five updates ran at
once and all ended honest. And the two EARLY power-cut phases nobody had cut in.

TWELVE NEW ROWS (R-615..R-626), register 303 -> 315, and eight existing rows updated with what was
measured — including two CORRECTIONS: R-606 records the pre-flight refusals as reaching an English
household in English and they do not, and R-446/R-458 are both narrower than their rows state.

Two instrument fixes were needed before anything could be trusted: the unattended caller turned
every success into a timeout (R-623), and one of my own reproductions was wrong and is kept
labelled with what it actually measured.

Interventions: zero. No controller, agent or hub code written. The hub was never touched beyond
the floor the operator asked for.

Gates: repo_gates.py --fast, all 15 OK.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-21 22:34:24 +02:00

16 KiB

UPDATE NIGHT 2026-09-21 — progress log (one line per finished step)

Evidence dir: documentation/audits/update-night-2026-09-21/ A resuming session reads THIS FILE FIRST and never repeats a finished step.

time (CEST) step verdict evidence
20:07 P0.1 floor → 0.261.0 (declared MinAgent 0.131.0) PROVEN — both demo boxes arrived in 13 s; hub managed floor SERVED … from declared (golden 0.258.0); drill-r50 stays held (agent 0.129.0 < 0.131.0, host DOWN) 01-floor-pre.txt, 02-floor-save.txt
20:08 P0.2a drill repo admin/app-catalog-drill created (private) DONE — Gitea token lacks write:user, so POST /api/v1/repos/migrate was used instead of /user/repos; main = f5f6a152b513 = live 03-drill-repo.txt
20:10 P0.2b 9202 repointed at the drill repo PROVEN — git.repo_url ALONE IS INERT: gitCloneOrPull only clones when .git is absent, else fetches from the old origin. Cache dir had to be removed too. Config saved as controller.yaml.pre-update-night 04-9202-config-pre.txt, 05-9202-follows-drill.txt
20:12 P0.2c positive + negative controls PROVEN — drill bump uptime-kuma 2.4.0→2.5.5 shows on 9202 as „Frissítés elérhető — ma" / "Update available — today"; live catalog still f5f6a152b513 with pin 2.4.0; both real boxes' caches still at f5f6a15. R-607 fired again (sync said „nincs változás"). App route is /apps/<n>, not /app/<n> 06-drift-rerun.txt, 07-positive-control.txt
20:07 Drift re-run CONFIRMED — 66 pins, 46 behind, 39 within-major, 7 across-major — the brief's numbers hold exactly 06-drift-rerun.txt
20:13 P0.4 capacity measured OK — 9202: 26 GB RAM (22 free), docker root on mp0 with 56 GB free, drive 872 GB. demo-hp / is 86% but holds neither. Brief's capacity claim HOLDS 08-capacity.txt
20:18 P0.3 throwaway image store PROVEN — registry:2 on 9202 127.0.0.1:5000; drill/glance:1.0.0 = real v0.8.6 retagged, :1.0.1 = starts/stays up/never serves, :1.0.2 absent (404). CompareImageRefs DOES order host:port/ refs — 4 positive + 1 negative control, run not read; tmp test deleted, tree clean 09-image-store.txt
20:20 EDGE privatebin 2.0.5 -> 2.0.6 (file-leg) PROVEN — 15.4 s; seed read back both sides; all four observables agree apps/privatebin/
20:25 EDGE docmost 0.95.0 -> 0.96.0 (db-postgres, engine constant) PROVEN — 103.5 s; seeded account authenticated after; all four observables agree apps/docmost/
20:28 EDGE bookstack 26.05.2 -> 26.05.5 (db-mariadb, engine constant) PROVEN — 45.1 s; artisan readback with its own negative control; DATABASE HALF ONLY (R-460) apps/bookstack/
20:37 FINDING R-618 tandoor probe port DEFECT, measured — probe 8080, app listens on 80 only; docker healthy + front door 200 + controller unhealthy. 53-template sweep run: wger suspected, adventurelog cleared 10-probe-port-sweep.txt
20:38 FINDINGS R-615/616/617/619 filed — repo_url inert; catalog token plaintext in the clone; Gitea migrate endpoint; type: password mandatory though the wire says optional NEW-ROWS.md
20:39 EDGE actualbudget 26.7.0 -> 26.9.0 PROVEN — 19.5 s apps/actualbudget/
20:44 EDGE navidrome 0.63.2 -> 0.64.0 (file-leg, HDD_PATH) PROVEN — 11.3 s apps/navidrome/
20:46 EDGE audiobookshelf 2.35.1 -> 2.36.1 (file-leg) PROVEN — 23.6 s apps/audiobookshelf/
20:54 EDGE adventurelog v0.12.1 -> v0.13.0 (db-postgis) FAILED — the most valuable result so far. Nine migrations applied OK, then the app never bound its port; held after the full 5-min wait; the sentence names tier, date and what the copy holds. Rows R-621 (the hold destroys the failure evidence) and R-622 (do not promote this edge) apps/adventurelog/
20:59 Harness code pushed to the catalog — 4 fixtures + edges U1..U7 DONE, gates green — live catalog main moves f5f6a152b513 -> 4463243f2e09, scripts/ ONLY, zero image: lines (the teardown diff still expects every image line identical). Harness RUNS owed app-catalog-felhom.eu@4463243f2e09
20:59 image reclaim BY NAME (no prune) 34 unused images removed by exact reference; 40 GB -> 55 GB free reclaim.sh
21:08 CI check for the catalog push GREEN — job id=830, name='gates', status='completed', conclusion='success', matched on head_sha=4463243f2e09. Confirms CLAUDE.md's warning: page 1's newest id was 52, the real newest was 830 — job ids are NOT page-ordered scratchpad/ci.sh
21:12 Observation, NOT a new row — an app with an HDD_PATH cannot have its DATA removed on 9202 R-442's guard working as designed: /api/disks answers agent not configured on this guest, so the drive path cannot be RESOLVED and the removal is refused with the app kept rather than half-deleted. The household's other choice — remove the app, KEEP the data — is accepted. The harness now takes that route and tidies its own directories by name at teardown apps/navidrome/, apps/romm/
21:13 FINDING R-623 — the unattended caller turned every SUCCESS into a timeout Read before use, not after: follow() read update_phase/updating off the API ENVELOPE, so both were always None, every followed update hit the 900 s timeout and was then marked never-press-again. Fixed in that file before B1 relied on it. The earlier night missed it because its only follow() pass was the one whose log was lost update-arc-gaps-2026-09-21/unattended-caller.py
21:18 Interim commit pushed — felhom.eu da20722e76a7 Phase 0 + Phase 1 evidence off the machine before Phases 2-5 (R-320). repo_gates.py --fast: all 15 OK. Includes the two instrument fixes (api-recipe route, unattended-caller follow()) felhom.eu@da20722e76a7
21:18 R-618 RAISED TO P1 by a live measurement tandoor 2.6.13→2.6.15 was serving HTTP 200 on the NEW version at 21:18:28 while the update sat in verifying, because the probe names port 8080 and the app listens on 80. The health wait then times out and failAndHold STOPS the working app. A wrong port turns every successful update into an outage plus a needless restore 14-tandoor-serving-while-verifying.txt
21:22 EDGE tandoor 2.6.13 -> 2.6.15 (db-postgres) FAILED — and the failure is OURS, not the app's. The new version was SERVING 200 with a green docker healthcheck at four samples across five minutes; the controller's probe names port 8080 and the app listens on 80, so verifying timed out and failAndHold STOPPED a working app. The four observables disagree exactly as 09 §5.2 predicts. R-618 is now P1 14-tandoor-serving-while-verifying.txt, apps/tandoor/
21:35 failwalk adventurelog — the household's WAY OUT, on a real failed edge PROVEN — the restore the hold sentence names („helyi", 1 copy offered) ran in 75 s; app back on v0.12.1, all three containers running, hold CLEARED, state running. R-606 CONFIRMED on the hold sentence itself (English page, Hungarian sentence) with positive and negative controls apps/adventurelog/failwalk.json, 15-r606-*.txt
21:36 Observation after the restore The hold sentence is GONE from the card (R-480 working) and the app reads „Fut". The badge briefly read „Frissítés elérhető — 65 napja" rather than „ma", because the restore brought back the recovery unit's own .felhom.yml with its pre-drill catalog_since; the syncer copies that file verbatim every cycle, so it heals within one sync. Recorded as an observation, not a row apps/adventurelog/failwalk.json
21:37 failwalk tandoor — the way out of the WRONG-PROBE hold PROVEN — restore in 32.4 s, hold cleared, app running, data reads back. Note what this means for R-618: the household's route back works, so nothing is LOST — but they had to walk it for an update that had actually succeeded apps/tandoor/failwalk.json
21:40 PHASE 2.1 — MariaDB 11.6 -> 12.3 through the REAL Update button PROVEN, a first. All four SPIKE-r459 observables: datadir 11.6.2->12.3.3; the engine's own check says „already upgraded … no need to run mariadb-upgrade again"; the entrypoint says „Major version upgrade detected … Check required!" then started and finished (NOT the skipped due to $MARIADB_AUTO_UPGRADE line R-459 feared); the engine's own pre-upgrade backup system_mysql_backup_11.6.2-MariaDB.sql.zst 631 905 B. Seeded Nextcloud account read back. All four version observables agree 16-phase2.1-mariadb-major.md, apps/nextcloud-engine-mariadb/
21:45 PHASE 2.2 — PostgreSQL 16 -> 17 through the real Update button FAILED, exactly as R-463 predicted and nobody had measured. The update ended in 5.1 s, the app was stopped and held, and pinned named postgres:17-alpine while installed still said 16 and nothing was running — the brief's own question answered. The restore then brought it back in 29.1 s (hold CLEARED, health probe 200) apps/docmost-engine-postgres/
21:45 R-621 demonstrated itself on this leg's KEY artefact The engine's refusal line was GONE before any probe could read it — failAndHold had removed the container, and the controller log does not carry it either. Reproduced independently instead (R-320) 18-postgres-refusal-reproduced.txt
21:45 My own first reproduction was WRONG, and is kept labelled The volume lookup returned empty -> the copy was empty -> postgres:17 initialised a FRESH datadir and started happily. A blank PG_VERSION should have stopped the step and did not. Rewritten so every step proves itself first. The accidental result (16 refusing a 17 datadir, verbatim) is real and kept 17-postgres-refusal-reproduced.txt
21:56 B1 — THE UNATTENDED HOLD. MEASURED, and it is the thing 09 Q4 has never had. The caller pressed ONCE with nobody watching; the drill image started, stayed up and never served; safety-dump -> verifying -> held after 312.9 s; then pass 2 and pass 3 pressed NOTHING — outcomes={'glance': ('held', 312.9)} never_again=['glance']. The hold sentence names the tier, the date and what the copy holds. My follow() fix (R-623) was load-bearing: without it this would have read timeout after 900 s bad-days/B1-unattended-hold/
21:57 B2 — the new tag cannot be pulled Exactly §6.1 Scenario E. phase=failed in 1.0 s, hold=None, „Az új verzió letöltése nem sikerült, ezért a frissítés elmaradt. Az alkalmazás a korábbi verzióval fut tovább." Nothing held; the app keeps running the old version bad-days/B2-pull-fails/
21:57 B3 — Update pressed while a backup runs 409 reason='busy' on six consecutive presses with „A frissítés most nem indítható: mentés/visszaállítás folyamatban. Próbáld újra, ha befejeződött." — the transient reason an unattended caller needs (R-609) bad-days/B3-busy-during-backup/
21:58 B7 — the 2 GB disk floor REFUSED correctly: filled to 1.4 GB free, 409 reason='disk', „Nincs elég szabad hely a frissítéshez: 1.4 GB szabad… legalább 2 GB szükséges." Nothing moved. Space released immediately after bad-days/B7-disk-and-memory-refusals/
21:58 CORRECTION to R-606 — the refusals are NOT English Three refusals requested with ?lang=en, Hungarian request as control: held, not_deployed and disk all come back identical Hungarian. The pipe exists (langFor honours ?lang=) but the sentences are frozen string constants, so routing them through errText changes nothing 20-refusals-in-english.txt
22:00 B9 — a frozen app gets a newer .felhom.yml §5.4's asymmetry confirmed live; no false alarm — 10 samples, all running/200. The reason NARROWS R-458: type: http treats any response as healthy, so a wrong path is invisible; the risk is real only for type: api + expect 21-B9-frozen-app-newer-felhomyml.md
22:00 B8 skipped here — docmost was swept by B1's precondition re-run in phase2_redo.sh with a freshly deployed docmost phase2_redo.sh
22:02 B4 (two) — two Updates within one second THERE IS NO SINGLE-FLIGHT. Both accepted 202 within 0.268 s, and 2 of 2 ran in flight at the same time. Both ended done, no error, no hold, both pins advanced. That is a real Slice-6 design input: an unattended caller pressing N apps would run N updates at once bad-days/B4-concurrent-updates/two.json
22:06 B4 (five) — five Updates within half a second 5 of 5 ran concurrently and ALL ended honest: accepted within 0.448 s, every one done, no error, no hold, every pin advanced, whole batch ~30 s. So the box does not serialise updates at all — and on this box it coped. Memory samples in the record bad-days/B4-concurrent-updates/five.json
22:08 B5 (backing-up cut) — the phase nobody had cut in BEHAVED. Cut at backing-up +0.027 s (pct stop returned 9.65 s later — the usual instrument limit, stated). After boot the box said so ITSELF: „crash recovery: an app-data backup (volume dump) … was interrupted and left 1 app(s) stopped — restarting them: [privatebin]" and „update recovery: privatebin was interrupted in backing-up". The pin did NOT move, the app runs, and a fresh paste seeded+read back bad-days/B5-backing-up/
22:08 B5 (safety-dump cut) — MISSED, recorded as a miss The phases went backing-up -> pulling -> failed in 0.473 s and safety-dump was never observed, so the plug was never pulled. Recorded as a MISS, not as a pass. To be retried with a genuine pending edge bad-days/B5-safety-dump/
22:09 Phase 4 — the morning after Every one of the 10 badges is TRUE (refs equal <-> „Naprakész", refs differ <-> „Frissítés elérhető"). Q4's four promises all scored True on the held app's page. zipline shows the household „Nem egészséges — URL nem elérhető" while running — R-618 in the household's own words bad-days/P4-morning-after/
22:09 B6 — the way out FORWARDS: REFUSED A held app met a FIXED newer version. Badge: „Frissítés elérhető — ma" in both languages, button offered; press -> 409 reason='held'. Correct per §6.1 (only a restore lifts a hold) but the page invites what the button refuses — the exact inconsistency R-524 removed for the Ahead case. Filed as R-625 bad-days/B6-way-out-forwards/
22:16 PHASE 2.3 — the PostgreSQL conversion REHEARSAL, costed (Q5) WORKED end to end. 49 MB / 48 tables: dump with 16 2.6 s / 132 201 B; fresh 17 + replay 6.5 s / 48 tables; the app said „Database connection successful"; the seeded account read back on 17; total 155.9 s, of which ~9 s is engine work. Two benign ERROR lines named. pg_upgrade NOT run — it needs an image that does not exist here 24-Q5-postgres-conversion-costed.md
22:18 wger CONFIRMED — R-618's last candidate closes probe type: http port: 80; inside the container port 80 refused, port 8000 ANSWERED; docker health green; front door 302; the box says unhealthy. Three confirmed instances now: tandoor, zipline, wger — and the cheap static rule finds all three with one false positive out of 53 22-wger-probe-measured.txt
22:19 B5 (safety-dump cut) — the second early phase, RETRIED and HIT Cut at safety-dump +0.023 s. After boot: no recovery line and no journal (unlike the backing-up cut, which produced both), the pin did NOT move, the app runs, a fresh paste seeded and read back, and the card shows no interrupted sentence at all. No half-written backup artefact on either cut 25-B5-the-two-early-cuts.md
22:20 Bonus proof across a REAL power cut The boot sweep met the held glance after an unclean shutdown and deliberately left it alone: „is a boot orphan by intent but is HELD … NOT starting it; whatever is holding it owns its recovery". §6.1's three-unattended-paths claim, proven across a power cut 25-B5-the-two-early-cuts.md
22:22 mealie v3.20.1 -> v3.27.0 (db-postgres) PROVEN — the edge that failed twice on instrument problems, walked cleanly on the third apps/mealie/
22:23 PHASE 1+2 COMPLETE 21 edges attempted: 14 proven, 3 failed, 4 inconclusive. Ten of the fourteen printed a verbatim migration line. Up from the three apps this project had ever measured summarise.py