9c69b3ff07
gates / gates (push) Successful in 27s
Evidence off the machine at the end of the phases that produced it (R-320). Teardown follows. PHASE 2 — the two database engines, through the REAL Update button: - MariaDB 11.6 -> 12.3 on nextcloud: PROVEN, and pressed through the button for the first time. All four SPIKE-r459 observables: the datadir's own record moved 11.6.2 -> 12.3.3; the engine itself says "already upgraded ... no need to run mariadb-upgrade again"; the entrypoint says "Major version upgrade detected ... Check required!" and then STARTED and FINISHED it (not the `skipped due to $MARIADB_AUTO_UPGRADE` line R-459 feared); and the engine took its own pre-upgrade backup, 631 905 B. The seeded Nextcloud account read back. - PostgreSQL 16 -> 17 on docmost: FAILED exactly as R-463 predicted and nobody had measured. 5.1 s to held; the pin named 17 while nothing ran; the restore brought it back in 29.1 s. The engine's REFUSAL LINE was destroyed by failAndHold before any probe could read it, so it was REPRODUCED INDEPENDENTLY with a control on every step (R-320). PHASE 3 — the bad days. B1 produced THE UNATTENDED HOLD, which this project has never had: the caller pressed once with nobody watching, the app held after 312.9 s, and passes 2 and 3 pressed nothing. B2 put the pin back on a pull failure in 1.0 s. B3 refused `busy` six times. B4 showed there is NO single-flight — 5 of 5 updates ran at once and all ended honest. B5 cut the power in `backing-up` and the box recovered itself and said so. B7 refused under the 2 GB floor. B9 found R-458's risk narrower than the row states. PHASE 4 — every badge on the box is TRUE, and the held app answers all four of Q4's questions. FINDINGS, five new and three corrections to existing rows. The one that matters: R-618 is P1 — two templates name a health probe the app does not answer, and because the guarded update waits on that same probe, a SUCCESSFUL update ends by STOPPING a working app. Measured: tandoor served HTTP 200 on the new version at four samples across five minutes and was then stopped. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
51 lines
2.6 KiB
Bash
Executable File
51 lines
2.6 KiB
Bash
Executable File
#!/bin/bash
|
|
# The three pieces that need docmost back on the box, run AFTER run_rest.sh.
|
|
#
|
|
# WHY THEY ARE HERE AND NOT IN LINE:
|
|
# * Phase 2.2's own verdict is already recorded and solid — the PostgreSQL 16 -> 17 edge FAILED,
|
|
# the app held, the pin named 17 while nothing ran, and the restore brought it back in 29 s.
|
|
# What was LOST is the engine's REFUSAL LINE, because `failAndHold` removed the container
|
|
# before any probe could read it. That is R-621 demonstrating itself on the single artefact
|
|
# this leg existed to collect. R-320 says: say so plainly, and REPRODUCE IT INDEPENDENTLY.
|
|
# * Phase 2.3's rehearsal could not seed, because the restored docmost already had its workspace
|
|
# (`403 Workspace setup already completed`), so it needs a FRESH deploy.
|
|
# * B8 needs a deployed app carrying a floating engine pin; docmost is that app.
|
|
#
|
|
# The first attempt at the reproduction is kept in 17-postgres-refusal-reproduced.txt and is WRONG:
|
|
# the volume lookup returned empty, the copy was therefore empty, and postgres:17 initialised a
|
|
# fresh datadir instead of meeting a 16 one. It printed an EMPTY PG_VERSION and I let it run on —
|
|
# an instrument that reports nothing should have stopped the step. It is kept, labelled, because
|
|
# the accidental result (a 17 datadir refused by 16, verbatim) is itself worth having.
|
|
cd /mnt/5_hdd/felhom.eu/git/felhom.eu/documentation/audits/update-night-2026-09-21 || exit 1
|
|
step() { echo; echo "############################################## $(date +%H:%M:%S) $*"; }
|
|
|
|
step "R1 redeploy docmost fresh on postgres:16-alpine, and seed it"
|
|
timeout 1800 python3 - <<'PY'
|
|
import sys; sys.path.insert(0, ".")
|
|
import walk as w
|
|
from fixtures import FIXTURES
|
|
w.login()
|
|
st = w.stack("docmost")
|
|
if st.get("deployed"):
|
|
w.say(" docmost is deployed — removing it so the rehearsal meets a FRESH workspace")
|
|
w.remove("docmost")
|
|
w.sync_rescan(expect_app="docmost", expect_ref="postgres:16-alpine")
|
|
if not w.deploy("docmost", "docs"):
|
|
sys.exit("docmost never came up")
|
|
tok = FIXTURES["docmost"].seed(w, "docs", w.say)
|
|
w.say(f" seeded: {bool(tok)}")
|
|
if tok:
|
|
w.say(f" C1 readback: {FIXTURES['docmost'].verify(w, 'docs', tok, w.say)}")
|
|
PY
|
|
|
|
step "R2 reproduce the PostgreSQL refusal INDEPENDENTLY, with a control on every step"
|
|
timeout 1800 bash ./reproduce_pg_refusal.sh
|
|
|
|
step "R3 Phase 2.3 — the conversion rehearsal, costed"
|
|
timeout 4500 python3 phase2_pgrehearsal.py
|
|
|
|
step "R4 B8 — the floating pin, on docmost's postgres"
|
|
timeout 1800 python3 phase3_legs2.py b8 docmost docmost-postgres
|
|
|
|
echo; echo "$(date +%H:%M:%S) phase2_redo done"
|