a8caa0fdde
gates / gates (push) Successful in 17s
07-backup-architecture.md 6.3 gains a dated [DESIGN] paragraph on replay -> rollback -> hold, including why no engine flag closes it: --single-transaction makes Postgres atomic, MariaDB DDL is not transactional, so the rollback is the fix and the flag is a belt. Drill record for the live walk, including the TWO defects the walk found in the fix itself (a rollback into a re-created container; an operator route that cleared the file while the running controller kept refusing) and the ONE red-proof that PASSED, which is reported rather than omitted. R-379..R-382 compressed into CLOSED-ITEMS.md. OPEN-ITEMS 330683 -> 325236 bytes. STATUS.md restates the outcome and names the next operator step.
115 lines
7.1 KiB
Markdown
115 lines
7.1 KiB
Markdown
# DRILL — R-379/R-380: the rollback, proven live (2026-08-22)
|
|
|
|
**Subject:** `demo-hp` (Tier 0), guest 9201. **Controller v0.220.0 → v0.220.1 → v0.220.2** during the
|
|
walk — the walk itself found two defects and both were fixed and re-proven.
|
|
**Method:** endpoint-level, the exact endpoints the UI's forms post to. No browser on DooPlex.
|
|
**Subjects were FOUND, not rebuilt:** `docmost` (Postgres 16) and `bookstack` (MariaDB 12.3), left
|
|
running with their planted data by the 2026-08-22 R-356b drill.
|
|
|
|
## The ladder, proven rung by rung
|
|
|
|
| step | subject | result |
|
|
|---|---|---|
|
|
| 1 | `docmost` (Postgres) | replay forced to fail → **rollback succeeded** → app healthy → data byte-identical |
|
|
| 2 | `bookstack` (MariaDB) | same → **`migrations` back at 102 rows**, the exact cell R-380 was measured in |
|
|
| 3 | `docmost` | both failed → **app held, not started**; every start path refused; survived a controller restart; cleared via the operator route |
|
|
| 4 | `docmost` | clean restore → **no rollback, no hold**, proven with a working positive control |
|
|
| 5 | `/backups/apps` | no phantom rows; the undo cap holds at 3 per app |
|
|
|
|
**Step 1 data check.** Pre-restore: 4 pages, `ALTERED-VALUE-B`, titles sha256
|
|
`8ec1fa8710c5ca08589161a2d930ede24a98c6b8ad3395308a90def21b68be19`. Post-restore: **identical on all
|
|
three**.
|
|
|
|
**Step 2 data check.** Pre: entities 1, users 2, migrations 102, book name hex
|
|
`C3817276C3AD7A74C5B172C591206BC3B66E79766573706F6C6320E28094205233353662`, disc `ORIGINAL-VALUE-A`.
|
|
Post: **identical on all five**.
|
|
|
|
## The customer message, verbatim
|
|
|
|
**257 bytes** (was 407 on this path, of which ~250 were engine output):
|
|
|
|
> A teljes visszaállítás sikertelen: a(z) docmost adatbázisának visszaállítása sikertelen — az adataid
|
|
> visszakerültek a visszaállítás előtti állapotba, az alkalmazás fut tovább. Ha újra megpróbálnád,
|
|
> előbb vedd fel velünk a kapcsolatot
|
|
|
|
Hex in `10-step1-message.txt`. **Judged, not just recorded:** it states the failure AND the recovery.
|
|
A message reporting only the failure would leave a customer believing their data was gone when it is
|
|
not. Checked for `ERROR:`, `LINE 1:`, `COPY public.`, `exit status`, `psql`, `INSERT INTO` — **all
|
|
absent**.
|
|
|
|
The held-app message (Scenario C), **verbatim**:
|
|
|
|
> a(z) docmost adatbázisának visszaállítása sikertelen, és a korábbi állapot visszatöltése sem
|
|
> sikerült. Az alkalmazást biztonsági okból LEÁLLÍTVA hagytuk, hogy az adatai ne sérüljenek tovább.
|
|
> Vedd fel velünk a kapcsolatot — a korábbi állapot mentése megvan: pre-restore-…-docmost-postgres.sql
|
|
|
|
## Scenario C in plain words
|
|
|
|
**What a customer sees:** the app is stopped and stays stopped. Pressing start returns, in Hungarian,
|
|
that the restore broke, the previous state could not be put back, the app is deliberately stopped so
|
|
the data cannot be damaged further, and to contact us. The app's row is **red**, not green.
|
|
|
|
**What an operator does:** `docker exec felhom-controller /usr/local/bin/felhom-controller
|
|
--restore-holds` lists the app with both errors. After checking the data (the undo copies are in the
|
|
app's unit `db-dumps/`), `--clear-restore-hold <app>` clears it — **and then the controller must be
|
|
restarted**, which the command now says.
|
|
|
|
## TWO DEFECTS THE WALK FOUND IN THE FIX ITSELF
|
|
|
|
**1. The rollback used a dead container (v0.220.1).** `writeSafetyDump` captures its `DiscoveredDB`
|
|
before the stop; the DB-only start re-creates the container. Measured: `docmost-postgres` captured as
|
|
`9adbc14f9af6` at 16:05:44, re-created as `309795897b82` at 16:05:47, rollback's `docker exec` against
|
|
the dead id sat in `waitDBReady` for 30 s. **So the app was HELD for an infrastructure reason while
|
|
its data was perfectly recoverable** — the hold behaved correctly on a case that should never have
|
|
reached it. **No unit test could see this: they all inject the import seam and never look at container
|
|
identity.** Fixed by re-discovering and matching on `{stack, engine}`; a new test asserts the identity
|
|
handed to the import, and its red-proof convicts.
|
|
|
|
**2. The operator route did not take effect (v0.220.2).** `--clear-restore-hold` runs as a second
|
|
process: it cleared `settings.json` correctly and the running controller went on refusing, because it
|
|
holds its own in-memory settings. Found by using the route, not by reading it. The command now prints
|
|
the restart it needs. The lost-update window between the two processes is recorded rather than hidden.
|
|
|
|
## Red-proofs — eight planned, one PASSED
|
|
|
|
| # | mutation | observed |
|
|
|---|---|---|
|
|
| 1 | remove the rollback entirely | 3 tests failed: "the undo copy must be re-applied exactly once, got 0" |
|
|
| 2 | roll back only the first database | "BOTH databases must be rolled back, got 1" |
|
|
| 3 | start the app anyway after a failed rollback | "the app was STARTED onto a half-written database" |
|
|
| 4 | move the hold check below the driveless early return | "the hold check (line 1822) is AFTER the driveless early return (line 1820)" |
|
|
| 5 | paste engine stderr back into the customer error | **PASSED — reported, not omitted.** See below |
|
|
| 6 | delete the operator log line | "the engine stderr is no longer logged either" |
|
|
| 7 | revert the undo naming | "must resolve to the app it belongs to, not a phantom; got `pre-restore-…-docmost`" |
|
|
| 8 | prune keeps the oldest | "the newest 3 must survive; 20260401T000000Z is gone" |
|
|
| 9 | use the captured container id | "the rollback used the CAPTURED container id 9adbc14f9af6" |
|
|
|
|
**Red-proof 5 passed and that is a finding.** The R-381 behavioural test injects at `m.importDBDump`,
|
|
i.e. BELOW `ImportDump`, so re-adding the stderr inside `ImportDump` could not fail it — the test was
|
|
hollow for the layer the leak lives in. A guard was added at that layer (an AST assertion that the
|
|
returned error does not carry the captured stderr, plus a positive control that it is still logged),
|
|
and the mutation then convicted.
|
|
|
|
## What did NOT reproduce
|
|
|
|
**The undo copies rendering as apps on the customer's backup page.** The live page was read BEFORE any
|
|
change and contained **zero** `pre-restore` strings while four such files sat on disk — with a
|
|
positive control showing 8 real app rows and a database section. `buildAppBackupRows` iterates
|
|
DEPLOYED apps and only reads the derived-name map by key, so a phantom name becomes a KEY and never a
|
|
ROW. The phantom key was real and is fixed; the visible row was not. Their visibility is unchanged and
|
|
remains deliberate.
|
|
|
|
## Teardown — three layers
|
|
|
|
1. **Nothing was provisioned.** `docmost` and `bookstack` were already on the box. Both are healthy at
|
|
the end, with their data byte-identical to the start.
|
|
2. **No storage was added.** The undo copies are now capped at **3 per app** (was 4 and 2 growing),
|
|
and the cap was observed firing live: *"pruned old undo copy … (keeping the newest 3)"*.
|
|
3. **No hub-side record was created.** No customer, no appliance. Nothing to dispose of.
|
|
|
|
## Deliberately left open
|
|
|
|
R-102 (the Tier-2 unit mirror read by nothing), R-359 (no readability check on the off-site store),
|
|
R-361 (the app's own dump vanishing from the unit after a reconstitution — reproduced again here and
|
|
still not fixed). Separate rows, untouched.
|