Files
felhom.eu/documentation/audits/update-night-2026-09-21/pgrehearsal.log
T
admin 8d786f7940
gates / gates (push) Successful in 28s
Update night 2026-09-21: the full record, twelve rows, and the answers to five of the seven questions
The drill is complete. Teardown done in three layers plus Gitea; the live catalog's every `image:`
line is proven identical to before.

WHAT WAS MEASURED. 21 edges across 19 apps, on scratch guest 9202 through the product's own
guarded Update, against a PRIVATE DRILL CATALOG so the live catalog carried no test reference at
any point: 14 proven, 3 failed, 4 inconclusive. Each app seeded and read back through its OWN
front door, with a negative control on every readback. Ten of the fourteen printed a verbatim
migration line. Up from the three apps this project had ever measured.

THE RESULT THAT MATTERS. R-618, P1: three of the 53 templates name a health probe the app does not
answer, and because the guarded update WAITS on that same probe, a SUCCESSFUL update ends by
STOPPING a working app. tandoor was measured serving HTTP 200 on the new version at four samples
across five minutes, docker's own healthcheck green, and was then stopped and the household sent
to a restore they did not need. zipline and wger are the same defect, both confirmed live. The
gate that catches all three is static and cheap: both health checks already sit in the same file.

WHAT THE NIGHT ANSWERED that was open. The UNATTENDED HOLD (312.9 s, pressed once, never again) —
which needed a purpose-built image store, because the rule that makes automatic updates safe is
the same rule that refuses the obvious way to break one. MariaDB across a major through the real
button, all four observables, first time. PostgreSQL across a major, refusing exactly as predicted,
with the conversion costed at ~9 s of engine work. There is NO single-flight: five updates ran at
once and all ended honest. And the two EARLY power-cut phases nobody had cut in.

TWELVE NEW ROWS (R-615..R-626), register 303 -> 315, and eight existing rows updated with what was
measured — including two CORRECTIONS: R-606 records the pre-flight refusals as reaching an English
household in English and they do not, and R-446/R-458 are both narrower than their rows state.

Two instrument fixes were needed before anything could be trusted: the unattended caller turned
every success into a timeout (R-623), and one of my own reproductions was wrong and is kept
labelled with what it actually measured.

Interventions: zero. No controller, agent or hub code written. The hub was never touched beyond
the floor the operator asked for.

Gates: repo_gates.py --fast, all 15 OK.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-21 22:34:24 +02:00

72 lines
4.6 KiB
Plaintext
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
22:13:25 ==== Phase 2.3: the PostgreSQL 16 -> 17 conversion rehearsal (Q5)
22:13:25 removing the existing docmost so the rehearsal meets a FRESH workspace it can seed
22:13:27 [X] stop -> 200 {'ok': True, 'message': 'Stack docmost stop completed'}
22:13:32 [X] remove (with drive data) -> 200 {'ok': True, 'data': {'removed': 'docmost', 'volumes_removed': ['docmost_docmost_postgres_data', 'docmost_docmost_redis_data', 'docmost_docmost_storage'], 'hdd_
22:13:40 [X] after remove: deployed=False leftovers='/opt/docker/stacks/docmost'
22:13:40 [1] deploy -> 202 {'ok': True, 'message': 'Telepítés elindítva – az állapot a kártyán követhető'}
22:14:05 [1] deployed, controller state=running, pinned={'docmost': 'docmost/docmost:0.96.0', 'docmost-postgres': 'postgres:16-alpine', 'docmost-redis': 'redis:7-alpine'}
22:14:05 docmost: /api/auth/setup http=200 rc=0
22:14:06 docmost: login as the seeded user http=200 ok=True
22:14:09 [00-state-before] 2.7s
22:14:09 PG_VERSION (the harness's own postgres probe, verbatim):
22:14:09 16
22:14:09 engine version:
22:14:09 postgres (PostgreSQL) 16.15
22:14:09 datadir size:
22:14:09 49.0M /var/lib/postgresql/data
22:14:09 volume:
22:14:09 docmost_docmost_postgres_data
22:14:12 [01-stop-the-app-keep-the-engine] 3.2s
22:14:12 app stopped (the engine stays up to be dumped)
22:14:12 docmost-postgres Up 31 seconds (healthy)
22:14:12 docmost-redis Up 31 seconds (healthy)
22:14:15 [02-dump-with-16] 2.6s
22:14:15 dumping as user=docmost
22:14:15 rc=0
22:14:15 dump bytes: 132201
22:14:15 CREATE TABLE statements: 48
22:14:21 [03-fresh-17-datadir-and-restore] 6.5s
22:14:21 17 up: postgres (PostgreSQL) 17.11
22:14:21 PG_VERSION on the fresh datadir: 17
22:14:21 restore rc=0
22:14:21 ERROR lines in the restore: 2
22:14:21 ERROR: role "docmost" already exists
22:14:21 ERROR: database "docmost" already exists
22:14:21 tables restored:
22:14:21 48
22:16:26 [04-point-the-app-at-17-and-start-it] 124.8s
22:16:26 DRILL-pg17 now answers to the name docmost-postgres on docmost_docmost-internal
22:16:26 app started
22:16:26 docmost Up 2 minutes (healthy)
22:16:26 docmost-postgres Exited (0) 2 minutes ago
22:16:26 docmost-redis Up 2 minutes (healthy)
22:16:26 {"level":"info","time":"2026-09-21T20:14:00.698Z","pid":45,"hostname":"24a35fe9b998","context":"NestApplication","msg":"Listening on http://127.0.0.1:3000 / https://docs.enkisfelhom.hu"}
22:16:26 [ELIFECYCLE] Command failed.
22:16:26 $ pnpm --filter ./apps/server run start:prod
22:16:26 $ cross-env NODE_ENV=production node dist/main
22:16:26 (node:45) ExperimentalWarning: localStorage is not available because --localstorage-file was not provided.
22:16:26 (Use `node --trace-warnings ...` to show where the warning was created)
22:16:26 {"level":"info","time":"2026-09-21T20:14:31.727Z","pid":45,"hostname":"24a35fe9b998","context":"RedisModule","msg":"default: the connection was successfully established"}
22:16:26 {"level":"info","time":"2026-09-21T20:14:32.033Z","pid":45,"hostname":"24a35fe9b998","context":"DatabaseModule","msg":"Establishing database connection"}
22:16:26 {"level":"info","time":"2026-09-21T20:14:32.065Z","pid":45,"hostname":"24a35fe9b998","context":"DatabaseModule","msg":"Database connection successful"}
22:16:27 docmost: login as the seeded user http=200 ok=True
22:16:27 [05] the seed read back on PostgreSQL 17: True
22:16:29 [06-engine-state-after] 2.3s
22:16:29 the harness's own postgres probe against the CONVERTED datadir:
22:16:29 17
22:16:29 [exit=0]
22:16:29 postgres (PostgreSQL) 17.11
22:16:29 size of the 17 datadir:
22:16:29 49.1M /var/lib/postgresql/data
22:16:43 [99-put-everything-back] 13.8s
22:16:43 docmost docmost/docmost:0.96.0 Up 5 seconds (health: starting)
22:16:43 docmost-postgres postgres:16-alpine Up 10 seconds (healthy)
22:16:43 docmost-redis redis:7-alpine Up 3 minutes (healthy)
22:16:43 docmost-postgres postgres:16-alpine Up 10 seconds (healthy)
22:16:43 PG_VERSION back on the original datadir: 16
22:16:43 docmost: login as the seeded user http=404 ok=False
22:16:43 docmost: login body 404 page not found
22:16:43 [99] the seed still reads on the ORIGINAL 16 datadir after teardown: False
22:16:43 rehearsal total 155.9s -> /mnt/5_hdd/felhom.eu/git/felhom.eu/documentation/audits/update-night-2026-09-21/bad-days/P2.3-pg-rehearsal/result.json