Files
felhom.eu/documentation/audits/night-2026-09-26/A/README.md
T

89 lines
6.3 KiB
Markdown

# Part A — the PostgreSQL conversion spike on docmost (9202, 2026-09-25 midday)
Written BEFORE any build. docmost 0.96.0 / `postgres:16-alpine` installed on 9202 through the product
(`POST /api/stacks/docmost/deploy`), seeded through its own API (a workspace + an account), read back
(`data-readback.txt`: `A-read0`). Every measurement below ran on a COPY of its datadir.
## A1 — the target major: 18 (CC-unattended decision, `09` §3 decision 37)
**Measured (`A1-images.txt`):**
- `postgres:16-alpine` and `17-alpine`: `PGDATA=/var/lib/postgresql/data`, `VOLUME /var/lib/postgresql/data`.
- `postgres:18-alpine` (18.6): `PGDATA=/var/lib/postgresql/18/docker`, `VOLUME /var/lib/postgresql`.
- **18 with an EMPTY volume at `/var/lib/postgresql/data` (where all eleven templates mount it) REFUSES:
exit 1** — *"in 18+, these Docker images are configured to store database data in a format which is
compatible with pg_ctlcluster … there appears to be PostgreSQL data in: /var/lib/postgresql/data (unused
mount/volume)"*. It refuses even when that mount is empty. Control: the same image with the volume at
`/var/lib/postgresql` initialises and serves, `PG_VERSION` 18.
- **docmost's upstream compose ships `image: postgres:18` with `db_data:/var/lib/postgresql`**
(github docmost/docmost `docker-compose.yml`, read 2026-09-25). Its docs name no other version.
**So the brief's claim is TRUE**, and stronger than stated: 18 refuses at the old mount, even empty. A move to
18 therefore needs the mount moved to `/var/lib/postgresql` in the same step, which the step's own definition
carries (the template's compose / `steps/<key>.yml`).
**Decision 37 — docmost converts 16 → 18, not 16 → 17.**
*One sentence:* which major does the first conversion target? **Options:** (a) 17 — no mount change;
(b) 18 — the mount moves to `/var/lib/postgresql` in the same step. **Costs:** (a) is a second conversion
later, and upstream already ships 18, so every box converts twice; (b) the step changes the volume's mount
point, so an undo must put the OLD mount back with the old data — which the undo does anyway (it restores the
old definition and the volume's bytes). **Why (b):** one conversion is better than two when the app's own
upstream runs the newer major; the mount change rides the step's definition and the undo's existing
definition restore. Reversible (the catalog can pin 17 instead before any box moves).
## A2 — what to load from: `pg_dumpall` from the old engine (decision 38)
`A2-A3-A5-measure.txt`. On the seeded datadir (48 tables, 62 rows, 49 MB):
| | size | time | what it carries |
|---|---|---|---|
| `pg_dumpall` (old engine) | 132 184 B | 0.59 s | roles WITH their password hashes, `CREATE DATABASE … LOCALE`, owners, grants |
| `pg_dump --no-owner --no-privileges docmost` (today's safety dump) | 123 545 B | 0.27 s | one database's schema + data; no roles, no owners, no database settings |
Both loaded into a fresh 18 gave **identical** databases, owners, encodings, collations, extensions
(`pg_trgm`, `plpgsql`, `unaccent`) and row counts for docmost — because docmost has ONE role (the bootstrap
superuser the new engine's entrypoint recreates from the same env) and ONE database.
**So the brief's claim is TRUE** (the safety dump is per-database `pg_dump` with `--no-owner --no-privileges`
— `appbackup/dbdump.go`), **and for docmost it would have been enough.** It would not be for an app with a
second role or a second database, which the other ten have not been measured for.
**Decision 38 — the box loads from a `pg_dumpall` taken from the OLD engine at conversion time, and the load
tolerates NO error.** *One sentence:* what does the box load into the new engine? **Options:** (a) the
existing safety dump; (b) a `pg_dumpall` of the old engine, loaded with `ON_ERROR_STOP`. **Costs:** (a) is
free, and silently drops any role, grant or database setting beyond the bootstrap ones; (b) one more dump
(0.6 s here) and a load that must not trip on the two objects the new engine's entrypoint already made (the
bootstrap role and database: measured, `pg_dumpall` loaded raw gives exactly two `already exists` ERRORs).
**Why (b):** decision 16 says "save everything". The two known collisions are removed precisely — the
entrypoint's empty databases are dropped before the load (only when they hold no table), and the dump's
`CREATE ROLE <name>;` line is skipped for a role that already exists (its `ALTER ROLE … PASSWORD` still
runs) — so any OTHER error stops the load and the undo runs. Proven by a test and live.
## A3 — the check
Old engine, before it stops; new engine, after the load — the same query set (0.5 s each here):
per database: owner, encoding, collation; every role (superuser, login, has a password); every extension by
name; every table's exact row count (`count(*)` per table). After: plus `PG_VERSION` of the new datadir
(`$PGDATA/PG_VERSION`, which follows 17's and 18's different layouts). All IDENTICAL for docmost on both
load routes.
## A4 — the undo after the volume was EMPTIED: proven by hand
`A4-undo-after-empty.txt`: copy with the product's own helper command (+ marker) → volume emptied (0
entries) → the product's `Restore` command → `PG_VERSION` 16 → the old engine "ready to accept
connections" → **the seeded account logged in through docmost's front door** (`A4-after-undo-productstart`).
**The brief's claim is TRUE.**
*Recorded as it happened:* my first start after the restore was a bare `docker compose up -d` in the stack
dir, which has no `.env` (the controller passes the app's env itself). docmost started without its secrets,
crash-looped, and **the box stopped it after 7 restarts in 10 min (decision 28) — the product working**, not
a fault. Started again with the product's own Start (`POST /api/stacks/docmost/start`); the seed read back.
## A5 — space
The conversion adds, on top of today's undo copy (the old datadir, 68.7 MB here): the dump (132 KB — 0.2 % of
the datadir) and nothing for the new datadir (it is built IN the emptied volume: 52.3 MB). The box's check:
**the dump's bound is the DB volume's own size** (a logical dump of live rows is smaller than the datadir
holding them, measured 0.2 %), with a 25 % margin, on the filesystem that holds the stack directory (where the
dump is written), plus the fixed 2 GB floor the update already keeps. Refused before anything moves.