Files
felhom.eu/documentation/audits/version-travel-2026-09-26/A1/README.md
T

92 lines
7.1 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# A1 — the window between an update and the next backup, measured (9202, controller 0.274.0, drill catalog)
2026-09-26 07:24–07:44Z. Guest 9202 on demo-hp, pointed at the drill catalog (`tools/repoint.py drill`, saved
config `controller.yaml.pre-vt0926`). docmost deployed at `docmost/docmost:0.95.0` + `postgres:16-alpine` from a
one-commit drill template (`85af330`, reverted by `8c3d099` right after the deploy, so the drill equals live
`main` again), seeded through docmost's own front door (a user + login, `fixtures.py`), then climbed the live
ladder with two presses of the guarded Update. Every file here is the raw record; this page is the reading.
Method: endpoint-level (the buttons' own endpoints: `POST /api/stacks/docmost/update`, the backups page's
`POST /backup/restore` form with `snapshot_id=helyi`, the second-drive „Teljes visszaállítás" form
`POST /backup/tier2/unit-restore`), and the unit read off the disk (`tools/a1.py unit`).
## Shape 1 — an app-image step (0.95.0 → 0.96.0, PostgreSQL 16 both sides)
| when (UTC) | unit `compose/` + `image_pins` | unit `db-dumps/docmost-postgres.sql` | restore point offered |
|---|---|---|---|
| 07:25:57 (the update's own "back up first") | 0.95.0 / 16 | 07:25:43, 0.95.0 schema | — |
| 07:26:33 update `done` | 0.95.0 / 16 (unchanged) | 07:25:43 | 07:25:57 |
| **07:27:53 the 5-minute status refresh** | **0.96.0 / 16** (`created_at` 07:27:53) | **07:25:43, 0.95.0 schema** | **07:27:53** |
`S1a-right-after-step1.txt`, `S1b-after-refresh.txt`. **The claim holds:** `CaptureRecoveryUnit`, run by the
refresh (`RefreshCache` → `captureAllRecoveryUnits`), rewrote `compose/` and `image_pins` from the stack dir
the first time it ran after the pin moved, and the dumps stayed as the backup legs left them. Tier 1's time
(`ListRestorePoints` = newest of manifest + dumps) moved to the refresh: **07:27:53 for data from 07:25:43**.
**Restore in the window (own unit, `helyi`) — `S2-restore-window-step1.txt`:** volumes first (3 tars,
including the database's own volume), then the definition from the unit's `compose/` (0.96.0), then the SQL
with only the database up. It completed ("3 volume(s) of 3 listed, 1 database(s) of 1 listed"), the seed read
back — **and docmost 0.96.0 ran four schema migrations on 0.95.0's data at its first start**
(`20260824T211732-page-title-trgm-index` … `20260904T171920-public-spaces`). So the restore performed the
update step itself, with no precondition, no safety copy of its own and no undo. Whole here; unguarded always.
## Shape 2 — an engine step (PostgreSQL 16 → 18, docmost-shaped, with the box's conversion)
| when (UTC) | unit `compose/` | unit `.sql` | pg volume tar |
|---|---|---|---|
| 07:29:45 back up first | 0.96.0 / **16**, mount `/var/lib/postgresql/data` | 07:29:45, `Dumped from database version 16.15` | 07:29:47, a 16 datadir at the volume root |
| 07:30:39 update `done` (converting 7 s) | unchanged | unchanged | unchanged |
| **07:32:53 refresh** | **0.96.0 / 18, mount `/var/lib/postgresql`** | **07:29:45, 16.15** | **07:29:47, 16 datadir** |
`S3a-right-after-step2.txt`, `S3b-after-refresh-step2.txt`. `conversion_copy` recorded (at 07:30:39).
**Restore in the window (own unit) — `S4-restore-window-step2.txt`. The claim is TRUE and the outcome is the
worst of the three the brief listed: REFUSED by the engine, and the app left DOWN.** The 16 datadir tar was
poured into the volume first (the volume is removed and recreated, so the live 18 datadir is GONE at that
moment); the unit's 18 definition was then written; the DB-only start of `postgres:18-alpine` refused the
old layout (`Error: in 18+, these Docker images are configured to store database data in a format which is
compatible with "pg_ctlcluster" … Counter to that, there appears to be PostgreSQL data in:
/var/lib/postgresql`) and restarted in a loop; the replay gave up (`waiting for docmost-postgres (postgres)
readiness: timeout after 30s`); the full start failed (`docker compose up -d … exit code 1`). The page said
the restore failed; docmost stayed `restarting`; the seed did NOT read back (the app's router answered 404).
Neither "a fresh empty cluster with the SQL loaded into it" nor "the app comes back whole".
## A1.3 — the other two copies
**Second drive (Tier 2) — `S5-restore-tier2-window.txt`.** The mirror had been written by the update's own
"back up first" (07:29) and was NOT touched by the refresh: its `compose/` still said 0.96.0 / **16** with
the 16 dump and tar. „Teljes visszaállítás" from it brought docmost back **whole at PostgreSQL 16**: 3 volumes,
1 database, the seed read back, `PG_VERSION` 16, and the pin moved back to `postgres:16-alpine` (the restore
pins to the unit's definition, `09` §5.3). This is option A's shape — by timing, not by design: the mirror
is rewritten on the next Tier-2 run from whatever the primary unit then holds, so after the nightly Tier 2
it would carry the same mix as the primary.
**Off-site (Tier 3) — NOT measured live**: 9202 has no off-site target, and the only boxes that have one
(9201, demo-felhom) are fenced read-only for this session. **Read from source instead
(`offbox_reconstitute.go` `ReconstituteFromOffsite`):** the off-site restore **never writes the definition**
— it resolves the database service from the LIVE compose ("reconstitution never overwrites the stack dir"),
replays the snapshot's volume tars (the database's raw volume included), then the SQL. The snapshot is taken
nightly after its own dump leg, so it is internally coherent — but an update during the day moves the live
definition, and until the next night's off-site run **every** off-site restore puts the old data under the
new definition. For an engine step that is S4's mechanism exactly, with one difference: this path takes a
safety dump and rolls back on a replay failure — but the DB-only start fails before any replay, so the path
goes to `restartStack` with the live datadir already replaced. Inferred from source + S4, not observed.
## Also found
1. **A restore drops the `conversion_copy` record but not the volume** (`S4`, `S5`: `conversion_copy=None`
while `docmost_docmost_postgres_data.pre-update-20260926T073003Z` still exists). `PersistUnitRedeployConfig`
builds a fresh `AppConfig`, so the record the hourly release reads is gone and the copy is never released —
disk only, never data. Filed.
2. **The refresh also re-captures `.felhom.yml`** (6426 B → 3998 B, the applied record of the pinned version,
R-669) — expected; it belongs with the definition and is kept with it by the fix.
3. **The automatic leg already refuses an app older than the ladder** (`unattended.go` `legCandidate`,
`LegSkipOlderThanLadder`); a PERSON's press jumps to the template (`ladder.go` `nextLadderStep` Index −1,
measured live by part 5 on vikunja 2.5.0). So the brief's claim "a restore of a version older than the
ladder's first entry would then jump" is **true for a person's press and false for the automatic leg**.
## State left (for Part A's live proof)
docmost on 9202 at 0.96.0 / PostgreSQL 16, pinned there, seed readable; the orphaned pre-conversion copy
volume stays until teardown. The drill repo equals live `main` (`8c3d099` is a revert of `85af330`).