92 lines
7.1 KiB
Markdown
92 lines
7.1 KiB
Markdown
# A1 — the window between an update and the next backup, measured (9202, controller 0.274.0, drill catalog)
|
||
|
||
2026-09-26 07:24–07:44Z. Guest 9202 on demo-hp, pointed at the drill catalog (`tools/repoint.py drill`, saved
|
||
config `controller.yaml.pre-vt0926`). docmost deployed at `docmost/docmost:0.95.0` + `postgres:16-alpine` from a
|
||
one-commit drill template (`85af330`, reverted by `8c3d099` right after the deploy, so the drill equals live
|
||
`main` again), seeded through docmost's own front door (a user + login, `fixtures.py`), then climbed the live
|
||
ladder with two presses of the guarded Update. Every file here is the raw record; this page is the reading.
|
||
|
||
Method: endpoint-level (the buttons' own endpoints: `POST /api/stacks/docmost/update`, the backups page's
|
||
`POST /backup/restore` form with `snapshot_id=helyi`, the second-drive „Teljes visszaállítás" form
|
||
`POST /backup/tier2/unit-restore`), and the unit read off the disk (`tools/a1.py unit`).
|
||
|
||
## Shape 1 — an app-image step (0.95.0 → 0.96.0, PostgreSQL 16 both sides)
|
||
|
||
| when (UTC) | unit `compose/` + `image_pins` | unit `db-dumps/docmost-postgres.sql` | restore point offered |
|
||
|---|---|---|---|
|
||
| 07:25:57 (the update's own "back up first") | 0.95.0 / 16 | 07:25:43, 0.95.0 schema | — |
|
||
| 07:26:33 update `done` | 0.95.0 / 16 (unchanged) | 07:25:43 | 07:25:57 |
|
||
| **07:27:53 the 5-minute status refresh** | **0.96.0 / 16** (`created_at` 07:27:53) | **07:25:43, 0.95.0 schema** | **07:27:53** |
|
||
|
||
`S1a-right-after-step1.txt`, `S1b-after-refresh.txt`. **The claim holds:** `CaptureRecoveryUnit`, run by the
|
||
refresh (`RefreshCache` → `captureAllRecoveryUnits`), rewrote `compose/` and `image_pins` from the stack dir
|
||
the first time it ran after the pin moved, and the dumps stayed as the backup legs left them. Tier 1's time
|
||
(`ListRestorePoints` = newest of manifest + dumps) moved to the refresh: **07:27:53 for data from 07:25:43**.
|
||
|
||
**Restore in the window (own unit, `helyi`) — `S2-restore-window-step1.txt`:** volumes first (3 tars,
|
||
including the database's own volume), then the definition from the unit's `compose/` (0.96.0), then the SQL
|
||
with only the database up. It completed ("3 volume(s) of 3 listed, 1 database(s) of 1 listed"), the seed read
|
||
back — **and docmost 0.96.0 ran four schema migrations on 0.95.0's data at its first start**
|
||
(`20260824T211732-page-title-trgm-index` … `20260904T171920-public-spaces`). So the restore performed the
|
||
update step itself, with no precondition, no safety copy of its own and no undo. Whole here; unguarded always.
|
||
|
||
## Shape 2 — an engine step (PostgreSQL 16 → 18, docmost-shaped, with the box's conversion)
|
||
|
||
| when (UTC) | unit `compose/` | unit `.sql` | pg volume tar |
|
||
|---|---|---|---|
|
||
| 07:29:45 back up first | 0.96.0 / **16**, mount `/var/lib/postgresql/data` | 07:29:45, `Dumped from database version 16.15` | 07:29:47, a 16 datadir at the volume root |
|
||
| 07:30:39 update `done` (converting 7 s) | unchanged | unchanged | unchanged |
|
||
| **07:32:53 refresh** | **0.96.0 / 18, mount `/var/lib/postgresql`** | **07:29:45, 16.15** | **07:29:47, 16 datadir** |
|
||
|
||
`S3a-right-after-step2.txt`, `S3b-after-refresh-step2.txt`. `conversion_copy` recorded (at 07:30:39).
|
||
|
||
**Restore in the window (own unit) — `S4-restore-window-step2.txt`. The claim is TRUE and the outcome is the
|
||
worst of the three the brief listed: REFUSED by the engine, and the app left DOWN.** The 16 datadir tar was
|
||
poured into the volume first (the volume is removed and recreated, so the live 18 datadir is GONE at that
|
||
moment); the unit's 18 definition was then written; the DB-only start of `postgres:18-alpine` refused the
|
||
old layout (`Error: in 18+, these Docker images are configured to store database data in a format which is
|
||
compatible with "pg_ctlcluster" … Counter to that, there appears to be PostgreSQL data in:
|
||
/var/lib/postgresql`) and restarted in a loop; the replay gave up (`waiting for docmost-postgres (postgres)
|
||
readiness: timeout after 30s`); the full start failed (`docker compose up -d … exit code 1`). The page said
|
||
the restore failed; docmost stayed `restarting`; the seed did NOT read back (the app's router answered 404).
|
||
Neither "a fresh empty cluster with the SQL loaded into it" nor "the app comes back whole".
|
||
|
||
## A1.3 — the other two copies
|
||
|
||
**Second drive (Tier 2) — `S5-restore-tier2-window.txt`.** The mirror had been written by the update's own
|
||
"back up first" (07:29) and was NOT touched by the refresh: its `compose/` still said 0.96.0 / **16** with
|
||
the 16 dump and tar. „Teljes visszaállítás" from it brought docmost back **whole at PostgreSQL 16**: 3 volumes,
|
||
1 database, the seed read back, `PG_VERSION` 16, and the pin moved back to `postgres:16-alpine` (the restore
|
||
pins to the unit's definition, `09` §5.3). This is option A's shape — by timing, not by design: the mirror
|
||
is rewritten on the next Tier-2 run from whatever the primary unit then holds, so after the nightly Tier 2
|
||
it would carry the same mix as the primary.
|
||
|
||
**Off-site (Tier 3) — NOT measured live**: 9202 has no off-site target, and the only boxes that have one
|
||
(9201, demo-felhom) are fenced read-only for this session. **Read from source instead
|
||
(`offbox_reconstitute.go` `ReconstituteFromOffsite`):** the off-site restore **never writes the definition**
|
||
— it resolves the database service from the LIVE compose ("reconstitution never overwrites the stack dir"),
|
||
replays the snapshot's volume tars (the database's raw volume included), then the SQL. The snapshot is taken
|
||
nightly after its own dump leg, so it is internally coherent — but an update during the day moves the live
|
||
definition, and until the next night's off-site run **every** off-site restore puts the old data under the
|
||
new definition. For an engine step that is S4's mechanism exactly, with one difference: this path takes a
|
||
safety dump and rolls back on a replay failure — but the DB-only start fails before any replay, so the path
|
||
goes to `restartStack` with the live datadir already replaced. Inferred from source + S4, not observed.
|
||
|
||
## Also found
|
||
|
||
1. **A restore drops the `conversion_copy` record but not the volume** (`S4`, `S5`: `conversion_copy=None`
|
||
while `docmost_docmost_postgres_data.pre-update-20260926T073003Z` still exists). `PersistUnitRedeployConfig`
|
||
builds a fresh `AppConfig`, so the record the hourly release reads is gone and the copy is never released —
|
||
disk only, never data. Filed.
|
||
2. **The refresh also re-captures `.felhom.yml`** (6426 B → 3998 B, the applied record of the pinned version,
|
||
R-669) — expected; it belongs with the definition and is kept with it by the fix.
|
||
3. **The automatic leg already refuses an app older than the ladder** (`unattended.go` `legCandidate`,
|
||
`LegSkipOlderThanLadder`); a PERSON's press jumps to the template (`ladder.go` `nextLadderStep` Index −1,
|
||
measured live by part 5 on vikunja 2.5.0). So the brief's claim "a restore of a version older than the
|
||
ladder's first entry would then jump" is **true for a person's press and false for the automatic leg**.
|
||
|
||
## State left (for Part A's live proof)
|
||
|
||
docmost on 9202 at 0.96.0 / PostgreSQL 16, pinned there, seed readable; the orphaned pre-conversion copy
|
||
volume stays until teardown. The drill repo equals live `main` (`8c3d099` is a revert of `85af330`).
|