Files
felhom.eu/documentation/audits/update-night-2026-09-21/run_rest.sh
T
admin 9c69b3ff07
gates / gates (push) Successful in 27s
Update night: Phases 2-4 evidence — both engines, the unattended HOLD, and five new findings
Evidence off the machine at the end of the phases that produced it (R-320). Teardown follows.

PHASE 2 — the two database engines, through the REAL Update button:
- MariaDB 11.6 -> 12.3 on nextcloud: PROVEN, and pressed through the button for the first time.
  All four SPIKE-r459 observables: the datadir's own record moved 11.6.2 -> 12.3.3; the engine
  itself says "already upgraded ... no need to run mariadb-upgrade again"; the entrypoint says
  "Major version upgrade detected ... Check required!" and then STARTED and FINISHED it (not the
  `skipped due to $MARIADB_AUTO_UPGRADE` line R-459 feared); and the engine took its own
  pre-upgrade backup, 631 905 B. The seeded Nextcloud account read back.
- PostgreSQL 16 -> 17 on docmost: FAILED exactly as R-463 predicted and nobody had measured.
  5.1 s to held; the pin named 17 while nothing ran; the restore brought it back in 29.1 s.
  The engine's REFUSAL LINE was destroyed by failAndHold before any probe could read it, so it
  was REPRODUCED INDEPENDENTLY with a control on every step (R-320).

PHASE 3 — the bad days. B1 produced THE UNATTENDED HOLD, which this project has never had: the
caller pressed once with nobody watching, the app held after 312.9 s, and passes 2 and 3 pressed
nothing. B2 put the pin back on a pull failure in 1.0 s. B3 refused `busy` six times. B4 showed
there is NO single-flight — 5 of 5 updates ran at once and all ended honest. B5 cut the power in
`backing-up` and the box recovered itself and said so. B7 refused under the 2 GB floor. B9 found
R-458's risk narrower than the row states.

PHASE 4 — every badge on the box is TRUE, and the held app answers all four of Q4's questions.

FINDINGS, five new and three corrections to existing rows. The one that matters: R-618 is P1 —
two templates name a health probe the app does not answer, and because the guarded update waits
on that same probe, a SUCCESSFUL update ends by STOPPING a working app. Measured: tandoor served
HTTP 200 on the new version at four samples across five minutes and was then stopped.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-21 22:13:57 +02:00

69 lines
3.6 KiB
Bash
Executable File

#!/bin/bash
# The rest of the night, in DEPENDENCY order. Each step writes its own evidence before the next
# starts (R-320), and each is individually re-runnable.
#
# The order is not arbitrary:
# * B2/B3/B7/B9 all use bentopdf and must run BEFORE B4, which re-points it at the 2.0.x pair.
# * B8 needs a deployed app with a floating engine pin, which Phase 2.2's restore leaves behind.
# * B5 stops and starts the whole guest, so it comes after everything that does not want a reboot.
# * B6 is the LAST thing done to B1's held app, because Phase 4's morning-after look must see the
# app exactly as a household would find it at breakfast — held, and not yet interfered with.
cd /mnt/5_hdd/felhom.eu/git/felhom.eu/documentation/audits/update-night-2026-09-21 || exit 1
step() { echo; echo "############################################## $(date +%H:%M:%S) $*"; }
step "1/15 failwalk — the household's way out of tonight's TWO real held edges"
timeout 1800 python3 failwalk.py adventurelog travel
timeout 1800 python3 failwalk.py tandoor recipes
step "2/15 Phase 2.1 — MariaDB across a major, through the real Update BUTTON"
timeout 2700 python3 engine_edge.py nextcloud mariadb:11.6 mariadb:12.3 \
--engine mariadb --container nextcloud-db --sub cloud --restore-after
step "3/15 Phase 2.2 — PostgreSQL across a major: what a household would see TODAY"
timeout 2700 python3 engine_edge.py docmost postgres:16-alpine postgres:17-alpine \
--engine postgres --container docmost-postgres --sub docs --restore-after
step "4/15 Phase 2.3 — the PostgreSQL conversion rehearsal (Q5), costed on a real seeded datadir"
timeout 4500 python3 phase2_pgrehearsal.py
step "5/15 B1 — the UNATTENDED HOLD (the app is then LEFT HELD until Phase 4)"
timeout 2700 python3 phase3_b1.py
step "6/15 B2 — the new tag cannot be pulled"
timeout 1800 python3 phase3_legs.py b2
step "7/15 B3 — Update pressed while a backup runs"
timeout 1800 python3 phase3_legs.py b3 bentopdf
step "8/15 B7 — the 2 GB disk floor, and the refusal in both languages"
timeout 1800 python3 phase3_legs.py b7 bentopdf
step "9/15 B9 — a FROZEN app receives a newer .felhom.yml (R-458)"
timeout 1800 python3 phase3_legs2.py b9 bentopdf pdf
step "10/15 B8 — a FLOATING pin: the up-to-date badge over an engine image that has moved upstream"
timeout 1800 python3 phase3_legs2.py b8 docmost docmost-postgres
step "11/15 B4 — two Updates within one second, then five"
timeout 1800 python3 phase3_b4.py prep privatebin,bentopdf
timeout 1800 python3 phase3_b4.py publish privatebin,bentopdf
timeout 1800 python3 phase3_b4.py fire privatebin,bentopdf two
timeout 2700 python3 phase3_b4.py prep privatebin,bentopdf,wishlist,uptime-kuma,opengist
timeout 2700 python3 phase3_b4.py publish privatebin,bentopdf,wishlist,uptime-kuma,opengist
timeout 2700 python3 phase3_b4.py fire privatebin,bentopdf,wishlist,uptime-kuma,opengist five
step "12/15 B5 — a power cut in backing-up, then in safety-dump (the two phases nobody has cut in)"
timeout 900 python3 phase3_b5.py knob backup_max_age 1m
timeout 2700 python3 phase3_b5.py privatebin paste backing-up backing-up
timeout 2700 python3 phase3_b5.py bentopdf pdf safety-dump safety-dump
timeout 900 python3 phase3_b5.py knob backup_max_age 24h
step "13/15 Phase 4 — the morning after, with B1's app still held"
timeout 2700 python3 phase4.py
step "14/15 B6 — the way out FORWARDS: a held app meets a fixed newer version"
timeout 2700 python3 phase3_legs2.py b6 glance
step "15/15 DONE — teardown is run separately and deliberately (phase5_teardown.py)"
echo "$(date +%H:%M:%S) all steps attempted"