CONTEXT + README: the failure ladder and the operator ruling (v0.220.x)
gates / gates (push) Successful in 11s

This commit is contained in:
2026-08-22 18:40:33 +02:00
parent 1a2405e86f
commit 2024ed9982
2 changed files with 59 additions and 1 deletions
+45 -1
View File
@@ -7,7 +7,51 @@
>
> Ask Claude Code: "Please update CONTEXT.md with what we did today"
Last updated: 2026-08-22 (v0.219.0 — R-356: the off-site restore refused every app that has no data drive)
Last updated: 2026-08-22 (v0.220.2 — R-379/R-380: the undo copy goes back when a database restore fails)
> **2026-08-22 — v0.220.0/.1/.2 (R-379, R-380, R-381, R-382).**
>
> **[DECISION — the OPERATOR's, 2026-08-22] When a database replay fails AND the rollback to the
> customer's own pre-restore copy also fails, the app is HELD STOPPED rather than started.** A
> running app on a half-written database lets the customer type into it, and that turns a recoverable
> state into a permanent one. The alternative — start it and mark it — was put to the operator and
> declined. If that judgement is ever revisited, this is the sentence to revisit.
>
> **[DESIGN] R-379 and R-380 were ONE failure with ONE fix.** Both ended with a half-restored
> database; the only difference was whether it looked broken (Postgres emptied and crash-looping,
> MariaDB partly applied behind `health=healthy`). No engine flag closes that: MariaDB's DDL is not
> transactional. Putting the customer's own copy back is what removes the half state, and it is the
> same `ImportDump` call a person ran by hand on 2026-08-22 to recover both apps.
>
> **THE UNDO SET IS MATCHED ON THE RUN'S OWN STAMP, never on the `pre-restore-` prefix.** Four such
> files accumulated on one app in one afternoon; a prefix match would replay an arbitrary older
> state. And `writeSafetyDump` returns the SET — it used to return the first path, which for a
> two-database app would have restored one and left the other half-written.
>
> **THE ROLLBACK RE-DISCOVERS THE CONTAINER.** The undo FILE is stable; the container is not. Found
> by v0.220.0's own live walk on its first real run: `docmost-postgres` was captured as
> `9adbc14f9af6`, re-created as `309795897b82` by the DB-only start, and the rollback's `docker exec`
> against the dead id timed out — so the app was held for an infrastructure reason while its data was
> recoverable. Fixed in v0.220.1. **No unit test saw it because they all inject the import seam and
> never look at container identity.**
>
> **THE HOLD IS NOT `DesiredState`.** That field is the customer's stated intent; writing our failure
> into it makes our fault indistinguishable from their choice. It is not the app-stop marker either —
> that means "owed a restart", and a held app is not owed one; leaving it would have `Recover()` start
> the broken app at the next boot. It is `Settings.RestoreHolds`, consulted by the shared
> `driveStartGate` **above** its driveless early return, because the apps this exists for have no
> drive.
>
> **The way out is `--clear-restore-hold <app>`, and it REQUIRES A CONTROLLER RESTART** — it runs as a
> second process and the running controller keeps its in-memory settings. v0.220.2 makes the command
> say so. Clearing through the running controller is the right shape later; it needs an operator tier
> the HTTP surface does not have (it authenticates as the customer, and a customer clearing their own
> hold is what the hold prevents).
>
> **Proven live on `demo-hp`**: Postgres and MariaDB both rolled back to byte-identical prior state
> (docmost titles sha256 `8ec1fa87…` unchanged; bookstack `migrations` 102, the exact cell R-380 was
> measured in). Evidence: `felhom.eu/documentation/audits/DRILL-r379-rollback-2026-08-22/`.
> **2026-08-22 — v0.219.0 (R-356). One predicate was answering two questions.**
>