Files
felhom.eu/documentation/audits/update-night-2026-09-21/phase2_pgrehearsal.py
T
admin 9c69b3ff07
gates / gates (push) Successful in 27s
Update night: Phases 2-4 evidence — both engines, the unattended HOLD, and five new findings
Evidence off the machine at the end of the phases that produced it (R-320). Teardown follows.

PHASE 2 — the two database engines, through the REAL Update button:
- MariaDB 11.6 -> 12.3 on nextcloud: PROVEN, and pressed through the button for the first time.
  All four SPIKE-r459 observables: the datadir's own record moved 11.6.2 -> 12.3.3; the engine
  itself says "already upgraded ... no need to run mariadb-upgrade again"; the entrypoint says
  "Major version upgrade detected ... Check required!" and then STARTED and FINISHED it (not the
  `skipped due to $MARIADB_AUTO_UPGRADE` line R-459 feared); and the engine took its own
  pre-upgrade backup, 631 905 B. The seeded Nextcloud account read back.
- PostgreSQL 16 -> 17 on docmost: FAILED exactly as R-463 predicted and nobody had measured.
  5.1 s to held; the pin named 17 while nothing ran; the restore brought it back in 29.1 s.
  The engine's REFUSAL LINE was destroyed by failAndHold before any probe could read it, so it
  was REPRODUCED INDEPENDENTLY with a control on every step (R-320).

PHASE 3 — the bad days. B1 produced THE UNATTENDED HOLD, which this project has never had: the
caller pressed once with nobody watching, the app held after 312.9 s, and passes 2 and 3 pressed
nothing. B2 put the pin back on a pull failure in 1.0 s. B3 refused `busy` six times. B4 showed
there is NO single-flight — 5 of 5 updates ran at once and all ended honest. B5 cut the power in
`backing-up` and the box recovered itself and said so. B7 refused under the 2 GB floor. B9 found
R-458's risk narrower than the row states.

PHASE 4 — every badge on the box is TRUE, and the held app answers all four of Q4's questions.

FINDINGS, five new and three corrections to existing rows. The one that matters: R-618 is P1 —
two templates name a health probe the app does not answer, and because the guarded update waits
on that same probe, a SUCCESSFUL update ends by STOPPING a working app. Measured: tandoor served
HTTP 200 on the new version at four samples across five minutes and was then stopped.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-21 22:13:57 +02:00

171 lines
8.8 KiB
Python

#!/usr/bin/env python3
"""Phase 2.3 — the PostgreSQL conversion rehearsal (`09` §3b Q5).
A REHEARSAL AND A COSTING, not a procedure for the catalog. It answers one question with numbers:
*what would it actually take to move one of the eleven PostgreSQL apps from 16 to 17, and what
could lose data?*
Route rehearsed: **logical dump and restore** — dump with 16, fresh 17 datadir, restore, app up,
seed read back. The `pg_upgrade` route is named at the end with what it would need; it is not run
unless time remains, because it needs BOTH majors' binaries in one image and that image does not
exist in this project.
VENUE, stated because it differs from the brief: this runs on guest 9202 itself, against the same
app the 5.2 leg left on a real 16 datadir with real seeded data — NOT on a separate harness LXC.
The rehearsal is deliberately performed with plain `docker` commands beside the product, never
through the product, because no product path for this exists and inventing one is what this
rehearsal is meant to COST rather than to build.
"""
import json, os, sys, time
from datetime import datetime, timezone
HERE = os.path.dirname(os.path.abspath(__file__))
sys.path.insert(0, HERE)
import walk as w # noqa: E402
from fixtures import FIXTURES # noqa: E402
APP = "docmost"
SUB = "docs"
PG = "docmost-postgres"
OUT = os.path.join(HERE, "bad-days", "P2.3-pg-rehearsal")
def step(label, script, timeout=900):
t0 = time.time()
out = w.guest(script, timeout=timeout)
dt = round(time.time() - t0, 1)
w.say(f" [{label}] {dt}s")
for line in out.strip().split("\n")[:14]:
if line.strip():
w.say(f" {line[:200]}")
return {"label": label, "seconds": dt, "output": out}
def main():
os.makedirs(OUT, exist_ok=True)
w.login()
w.say("==== Phase 2.3: the PostgreSQL 16 -> 17 conversion rehearsal (Q5)")
rec = {"leg": "P2.3", "app": APP, "route": "logical dump and restore",
"measured_at": datetime.now(timezone.utc).isoformat(), "steps": [], "notes": []}
# THE REHEARSAL DEPLOYS AND SEEDS IN ONE PROCESS. Measured twice tonight: a docmost that was
# already seeded answers `403 Workspace setup already completed`, and the seed token cannot
# cross a process boundary, so a rehearsal that assumes someone else seeded it cannot verify
# anything. It therefore starts from a FRESH app of its own.
st = w.stack(APP)
if st.get("deployed"):
w.say(" removing the existing docmost so the rehearsal meets a FRESH workspace it can seed")
w.remove(APP)
if not w.deploy(APP, SUB):
rec["notes"].append("docmost never came up for the rehearsal")
json.dump(rec, open(f"{OUT}/result.json", "w"), indent=2, ensure_ascii=False)
return
fx = FIXTURES[APP]
tok = fx.seed(w, SUB, w.say)
if tok is None:
rec["notes"].append("could not seed before the rehearsal — see log")
json.dump(rec, open(f"{OUT}/result.json", "w"), indent=2, ensure_ascii=False)
return
if not fx.verify(w, SUB, tok, w.say):
rec["notes"].append("C1 failed before the rehearsal")
json.dump(rec, open(f"{OUT}/result.json", "w"), indent=2, ensure_ascii=False)
return
rec["seed_read_before"] = True
rec["steps"].append(step("00-state-before", f"""
echo "PG_VERSION (the harness's own postgres probe, verbatim):"
docker exec {PG} sh -c 'cat /var/lib/postgresql/data/PG_VERSION 2>&1'
echo "engine version:"; docker exec {PG} postgres --version
echo "datadir size:"; docker exec {PG} sh -c 'du -sh /var/lib/postgresql/data 2>/dev/null'
echo "volume:"; docker inspect {PG} --format '{{{{range .Mounts}}}}{{{{if eq .Destination "/var/lib/postgresql/data"}}}}{{{{.Name}}}}{{{{end}}}}{{{{end}}}}'
"""))
rec["steps"].append(step("01-stop-the-app-keep-the-engine", f"""
docker stop {APP} >/dev/null 2>&1 && echo "app stopped (the engine stays up to be dumped)"
docker ps --filter name={APP} --format '{{{{.Names}}}} {{{{.Status}}}}'
"""))
rec["steps"].append(step("02-dump-with-16", f"""
PW=$(docker inspect {PG} --format '{{{{range .Config.Env}}}}{{{{println .}}}}{{{{end}}}}' | sed -n 's/^POSTGRES_PASSWORD=//p')
U=$(docker inspect {PG} --format '{{{{range .Config.Env}}}}{{{{println .}}}}{{{{end}}}}' | sed -n 's/^POSTGRES_USER=//p')
echo "dumping as user=$U"
time docker exec -e PGPASSWORD="$PW" {PG} pg_dumpall -U "$U" > /var/lib/felhom/DRILL-pg16.sql 2>/var/lib/felhom/DRILL-pg16.err
echo "rc=$?"
ls -l --block-size=1 /var/lib/felhom/DRILL-pg16.sql | awk '{{print "dump bytes:", $5}}'
head -3 /var/lib/felhom/DRILL-pg16.err 2>/dev/null
grep -c 'CREATE TABLE' /var/lib/felhom/DRILL-pg16.sql | sed 's/^/CREATE TABLE statements: /'
"""))
rec["steps"].append(step("03-fresh-17-datadir-and-restore", f"""
PW=$(docker inspect {PG} --format '{{{{range .Config.Env}}}}{{{{println .}}}}{{{{end}}}}' | sed -n 's/^POSTGRES_PASSWORD=//p')
U=$(docker inspect {PG} --format '{{{{range .Config.Env}}}}{{{{println .}}}}{{{{end}}}}' | sed -n 's/^POSTGRES_USER=//p')
DB=$(docker inspect {PG} --format '{{{{range .Config.Env}}}}{{{{println .}}}}{{{{end}}}}' | sed -n 's/^POSTGRES_DB=//p')
NET=$(docker inspect {PG} --format '{{{{range $k,$v := .NetworkSettings.Networks}}}}{{{{$k}}}}{{{{end}}}}' | head -1)
docker rm -f DRILL-pg17 >/dev/null 2>&1; docker volume rm DRILL-pg17-data >/dev/null 2>&1
docker volume create DRILL-pg17-data >/dev/null
docker run -d --name DRILL-pg17 --network "$NET" \\
-e POSTGRES_USER="$U" -e POSTGRES_PASSWORD="$PW" -e POSTGRES_DB="$DB" \\
-v DRILL-pg17-data:/var/lib/postgresql/data postgres:17-alpine >/dev/null
for i in $(seq 1 60); do docker exec DRILL-pg17 pg_isready -U "$U" >/dev/null 2>&1 && break; sleep 2; done
echo "17 up: $(docker exec DRILL-pg17 postgres --version)"
echo "PG_VERSION on the fresh datadir: $(docker exec DRILL-pg17 cat /var/lib/postgresql/data/PG_VERSION)"
time docker exec -i -e PGPASSWORD="$PW" DRILL-pg17 psql -U "$U" -d postgres < /var/lib/felhom/DRILL-pg16.sql > /var/lib/felhom/DRILL-restore.log 2>&1
echo "restore rc=$?"
grep -ciE '^ERROR' /var/lib/felhom/DRILL-restore.log | sed 's/^/ERROR lines in the restore: /'
grep -iE '^ERROR' /var/lib/felhom/DRILL-restore.log | head -5
echo "tables restored:"; docker exec -e PGPASSWORD="$PW" DRILL-pg17 psql -U "$U" -d "$DB" -tAc "select count(*) from information_schema.tables where table_schema='public'"
"""))
rec["steps"].append(step("04-point-the-app-at-17-and-start-it", f"""
NET=$(docker inspect {PG} --format '{{{{range $k,$v := .NetworkSettings.Networks}}}}{{{{$k}}}}{{{{end}}}}' | head -1)
docker stop {PG} >/dev/null 2>&1
docker network disconnect "$NET" DRILL-pg17 >/dev/null 2>&1
docker network connect --alias {PG} "$NET" DRILL-pg17
echo "DRILL-pg17 now answers to the name {PG} on $NET"
docker start {APP} >/dev/null && echo "app started"
for i in $(seq 1 60); do
S=$(docker inspect {APP} --format '{{{{.State.Running}}}}'); [ "$S" = true ] || break; sleep 2
done
docker ps -a --filter name={APP} --format '{{{{.Names}}}} {{{{.Status}}}}'
docker logs --tail 12 {APP} 2>&1 | tail -12
"""))
ok = fx.verify(w, SUB, tok, w.say)
rec["seed_read_after_on_17"] = ok
w.say(f" [05] the seed read back on PostgreSQL 17: {ok}")
rec["steps"].append(step("06-engine-state-after", f"""
echo "the harness's own postgres probe against the CONVERTED datadir:"
docker exec DRILL-pg17 sh -c 'cat /var/lib/postgresql/data/PG_VERSION 2>&1'; echo "[exit=$?]"
docker exec DRILL-pg17 postgres --version
echo "size of the 17 datadir:"; docker exec DRILL-pg17 sh -c 'du -sh /var/lib/postgresql/data'
"""))
rec["steps"].append(step("99-put-everything-back", f"""
NET=$(docker inspect {APP} --format '{{{{range $k,$v := .NetworkSettings.Networks}}}}{{{{$k}}}}{{{{end}}}}' | head -1)
docker stop {APP} >/dev/null 2>&1
docker network disconnect "$NET" DRILL-pg17 >/dev/null 2>&1
docker rm -f DRILL-pg17 >/dev/null 2>&1
docker volume rm DRILL-pg17-data >/dev/null 2>&1
docker start {PG} >/dev/null 2>&1; sleep 5
docker start {APP} >/dev/null 2>&1; sleep 5
rm -f /var/lib/felhom/DRILL-pg16.sql /var/lib/felhom/DRILL-pg16.err /var/lib/felhom/DRILL-restore.log
docker ps --filter name={APP} --format '{{{{.Names}}}} {{{{.Image}}}} {{{{.Status}}}}'
docker ps --filter name={PG} --format '{{{{.Names}}}} {{{{.Image}}}} {{{{.Status}}}}'
echo "PG_VERSION back on the original datadir: $(docker exec {PG} cat /var/lib/postgresql/data/PG_VERSION 2>&1)"
"""))
back = fx.verify(w, SUB, tok, w.say)
rec["seed_read_back_on_16_after_teardown"] = back
w.say(f" [99] the seed still reads on the ORIGINAL 16 datadir after teardown: {back}")
rec["total_seconds"] = sum(s["seconds"] for s in rec["steps"])
json.dump(rec, open(f"{OUT}/result.json", "w"), indent=2, ensure_ascii=False)
open(f"{OUT}/log.txt", "w").write("\n".join(w.LOG) + "\n")
w.say(f" rehearsal total {rec['total_seconds']}s -> {OUT}/result.json")
if __name__ == "__main__":
main()