Files
felhom.eu/documentation/audits/DRILL-r379-rollback-2026-08-22
admin a8caa0fdde
gates / gates (push) Successful in 17s
R-379/R-380 docs: the failure ladder, the drill record, register housekeeping
07-backup-architecture.md 6.3 gains a dated [DESIGN] paragraph on replay ->
rollback -> hold, including why no engine flag closes it: --single-transaction
makes Postgres atomic, MariaDB DDL is not transactional, so the rollback is the
fix and the flag is a belt.

Drill record for the live walk, including the TWO defects the walk found in the
fix itself (a rollback into a re-created container; an operator route that
cleared the file while the running controller kept refusing) and the ONE
red-proof that PASSED, which is reported rather than omitted.

R-379..R-382 compressed into CLOSED-ITEMS.md. OPEN-ITEMS 330683 -> 325236 bytes.

STATUS.md restates the outcome and names the next operator step.
2026-08-22 18:40:22 +02:00
..

DRILL — R-379/R-380: the rollback, proven live (2026-08-22)

Subject: demo-hp (Tier 0), guest 9201. Controller v0.220.0 → v0.220.1 → v0.220.2 during the walk — the walk itself found two defects and both were fixed and re-proven. Method: endpoint-level, the exact endpoints the UI's forms post to. No browser on DooPlex. Subjects were FOUND, not rebuilt: docmost (Postgres 16) and bookstack (MariaDB 12.3), left running with their planted data by the 2026-08-22 R-356b drill.

The ladder, proven rung by rung

step subject result
1 docmost (Postgres) replay forced to fail → rollback succeeded → app healthy → data byte-identical
2 bookstack (MariaDB) same → migrations back at 102 rows, the exact cell R-380 was measured in
3 docmost both failed → app held, not started; every start path refused; survived a controller restart; cleared via the operator route
4 docmost clean restore → no rollback, no hold, proven with a working positive control
5 /backups/apps no phantom rows; the undo cap holds at 3 per app

Step 1 data check. Pre-restore: 4 pages, ALTERED-VALUE-B, titles sha256 8ec1fa8710c5ca08589161a2d930ede24a98c6b8ad3395308a90def21b68be19. Post-restore: identical on all three.

Step 2 data check. Pre: entities 1, users 2, migrations 102, book name hex C3817276C3AD7A74C5B172C591206BC3B66E79766573706F6C6320E28094205233353662, disc ORIGINAL-VALUE-A. Post: identical on all five.

The customer message, verbatim

257 bytes (was 407 on this path, of which ~250 were engine output):

A teljes visszaállítás sikertelen: a(z) docmost adatbázisának visszaállítása sikertelen — az adataid visszakerültek a visszaállítás előtti állapotba, az alkalmazás fut tovább. Ha újra megpróbálnád, előbb vedd fel velünk a kapcsolatot

Hex in 10-step1-message.txt. Judged, not just recorded: it states the failure AND the recovery. A message reporting only the failure would leave a customer believing their data was gone when it is not. Checked for ERROR:, LINE 1:, COPY public., exit status, psql, INSERT INTO — all absent.

The held-app message (Scenario C), verbatim:

a(z) docmost adatbázisának visszaállítása sikertelen, és a korábbi állapot visszatöltése sem sikerült. Az alkalmazást biztonsági okból LEÁLLÍTVA hagytuk, hogy az adatai ne sérüljenek tovább. Vedd fel velünk a kapcsolatot — a korábbi állapot mentése megvan: pre-restore-…-docmost-postgres.sql

Scenario C in plain words

What a customer sees: the app is stopped and stays stopped. Pressing start returns, in Hungarian, that the restore broke, the previous state could not be put back, the app is deliberately stopped so the data cannot be damaged further, and to contact us. The app's row is red, not green.

What an operator does: docker exec felhom-controller /usr/local/bin/felhom-controller --restore-holds lists the app with both errors. After checking the data (the undo copies are in the app's unit db-dumps/), --clear-restore-hold <app> clears it — and then the controller must be restarted, which the command now says.

TWO DEFECTS THE WALK FOUND IN THE FIX ITSELF

1. The rollback used a dead container (v0.220.1). writeSafetyDump captures its DiscoveredDB before the stop; the DB-only start re-creates the container. Measured: docmost-postgres captured as 9adbc14f9af6 at 16:05:44, re-created as 309795897b82 at 16:05:47, rollback's docker exec against the dead id sat in waitDBReady for 30 s. So the app was HELD for an infrastructure reason while its data was perfectly recoverable — the hold behaved correctly on a case that should never have reached it. No unit test could see this: they all inject the import seam and never look at container identity. Fixed by re-discovering and matching on {stack, engine}; a new test asserts the identity handed to the import, and its red-proof convicts.

2. The operator route did not take effect (v0.220.2). --clear-restore-hold runs as a second process: it cleared settings.json correctly and the running controller went on refusing, because it holds its own in-memory settings. Found by using the route, not by reading it. The command now prints the restart it needs. The lost-update window between the two processes is recorded rather than hidden.

Red-proofs — eight planned, one PASSED

# mutation observed
1 remove the rollback entirely 3 tests failed: "the undo copy must be re-applied exactly once, got 0"
2 roll back only the first database "BOTH databases must be rolled back, got 1"
3 start the app anyway after a failed rollback "the app was STARTED onto a half-written database"
4 move the hold check below the driveless early return "the hold check (line 1822) is AFTER the driveless early return (line 1820)"
5 paste engine stderr back into the customer error PASSED — reported, not omitted. See below
6 delete the operator log line "the engine stderr is no longer logged either"
7 revert the undo naming "must resolve to the app it belongs to, not a phantom; got pre-restore-…-docmost"
8 prune keeps the oldest "the newest 3 must survive; 20260401T000000Z is gone"
9 use the captured container id "the rollback used the CAPTURED container id 9adbc14f9af6"

Red-proof 5 passed and that is a finding. The R-381 behavioural test injects at m.importDBDump, i.e. BELOW ImportDump, so re-adding the stderr inside ImportDump could not fail it — the test was hollow for the layer the leak lives in. A guard was added at that layer (an AST assertion that the returned error does not carry the captured stderr, plus a positive control that it is still logged), and the mutation then convicted.

What did NOT reproduce

The undo copies rendering as apps on the customer's backup page. The live page was read BEFORE any change and contained zero pre-restore strings while four such files sat on disk — with a positive control showing 8 real app rows and a database section. buildAppBackupRows iterates DEPLOYED apps and only reads the derived-name map by key, so a phantom name becomes a KEY and never a ROW. The phantom key was real and is fixed; the visible row was not. Their visibility is unchanged and remains deliberate.

Teardown — three layers

  1. Nothing was provisioned. docmost and bookstack were already on the box. Both are healthy at the end, with their data byte-identical to the start.
  2. No storage was added. The undo copies are now capped at 3 per app (was 4 and 2 growing), and the cap was observed firing live: "pruned old undo copy … (keeping the newest 3)".
  3. No hub-side record was created. No customer, no appliance. Nothing to dispose of.

Deliberately left open

R-102 (the Tier-2 unit mirror read by nothing), R-359 (no readability check on the off-site store), R-361 (the app's own dump vanishing from the unit after a reconstitution — reproduced again here and still not fixed). Separate rows, untouched.