From 2024ed99826717bf765803e159c074909b28bf83 Mon Sep 17 00:00:00 2001 From: kisfenyo Date: Sat, 22 Aug 2026 18:40:33 +0200 Subject: [PATCH] CONTEXT + README: the failure ladder and the operator ruling (v0.220.x) --- CONTEXT.md | 46 +++++++++++++++++++++++++++++++++++++++++++- controller/README.md | 14 ++++++++++++++ 2 files changed, 59 insertions(+), 1 deletion(-) diff --git a/CONTEXT.md b/CONTEXT.md index 8d903f8..6de6106 100644 --- a/CONTEXT.md +++ b/CONTEXT.md @@ -7,7 +7,51 @@ > > Ask Claude Code: "Please update CONTEXT.md with what we did today" -Last updated: 2026-08-22 (v0.219.0 — R-356: the off-site restore refused every app that has no data drive) +Last updated: 2026-08-22 (v0.220.2 — R-379/R-380: the undo copy goes back when a database restore fails) + +> **2026-08-22 — v0.220.0/.1/.2 (R-379, R-380, R-381, R-382).** +> +> **[DECISION — the OPERATOR's, 2026-08-22] When a database replay fails AND the rollback to the +> customer's own pre-restore copy also fails, the app is HELD STOPPED rather than started.** A +> running app on a half-written database lets the customer type into it, and that turns a recoverable +> state into a permanent one. The alternative — start it and mark it — was put to the operator and +> declined. If that judgement is ever revisited, this is the sentence to revisit. +> +> **[DESIGN] R-379 and R-380 were ONE failure with ONE fix.** Both ended with a half-restored +> database; the only difference was whether it looked broken (Postgres emptied and crash-looping, +> MariaDB partly applied behind `health=healthy`). No engine flag closes that: MariaDB's DDL is not +> transactional. Putting the customer's own copy back is what removes the half state, and it is the +> same `ImportDump` call a person ran by hand on 2026-08-22 to recover both apps. +> +> **THE UNDO SET IS MATCHED ON THE RUN'S OWN STAMP, never on the `pre-restore-` prefix.** Four such +> files accumulated on one app in one afternoon; a prefix match would replay an arbitrary older +> state. And `writeSafetyDump` returns the SET — it used to return the first path, which for a +> two-database app would have restored one and left the other half-written. +> +> **THE ROLLBACK RE-DISCOVERS THE CONTAINER.** The undo FILE is stable; the container is not. Found +> by v0.220.0's own live walk on its first real run: `docmost-postgres` was captured as +> `9adbc14f9af6`, re-created as `309795897b82` by the DB-only start, and the rollback's `docker exec` +> against the dead id timed out — so the app was held for an infrastructure reason while its data was +> recoverable. Fixed in v0.220.1. **No unit test saw it because they all inject the import seam and +> never look at container identity.** +> +> **THE HOLD IS NOT `DesiredState`.** That field is the customer's stated intent; writing our failure +> into it makes our fault indistinguishable from their choice. It is not the app-stop marker either — +> that means "owed a restart", and a held app is not owed one; leaving it would have `Recover()` start +> the broken app at the next boot. It is `Settings.RestoreHolds`, consulted by the shared +> `driveStartGate` **above** its driveless early return, because the apps this exists for have no +> drive. +> +> **The way out is `--clear-restore-hold `, and it REQUIRES A CONTROLLER RESTART** — it runs as a +> second process and the running controller keeps its in-memory settings. v0.220.2 makes the command +> say so. Clearing through the running controller is the right shape later; it needs an operator tier +> the HTTP surface does not have (it authenticates as the customer, and a customer clearing their own +> hold is what the hold prevents). +> +> **Proven live on `demo-hp`**: Postgres and MariaDB both rolled back to byte-identical prior state +> (docmost titles sha256 `8ec1fa87…` unchanged; bookstack `migrations` 102, the exact cell R-380 was +> measured in). Evidence: `felhom.eu/documentation/audits/DRILL-r379-rollback-2026-08-22/`. + > **2026-08-22 — v0.219.0 (R-356). One predicate was answering two questions.** > diff --git a/controller/README.md b/controller/README.md index 2c0c13d..8c39d61 100644 --- a/controller/README.md +++ b/controller/README.md @@ -571,6 +571,20 @@ Each app can define rich metadata in `.felhom.yml`: (they are the undo). The live recovery unit is still never overwritten, which is why the replay source is the scratch. Honesty surfaces (`OffsiteScratchPair`): dump age, an unstamped-pair warning, and the R-44 empty-dump sniff — all warn-level, none of them gates. + - **What happens when the database replay FAILS (v0.220.0–.2, R-379/R-380).** A ladder, and every + rung is observable: **replay → rollback → hold.** The pre-restore undo copy has always been + taken; since v0.220.0 it is also **put back** when the replay fails — the whole set for this run, + matched on the run's own stamp (never on the `pre-restore-` prefix, and never just the first + file), re-applied with the DB service still up and before any restart, into a **re-discovered** + container (the DB-only start re-creates it, so the captured id is dead by then — R-379, + v0.220.1). The app then starts and the message says both that the restore failed and that the + data is back. If the rollback ALSO fails the app is **held stopped** — the operator's ruling — + the hold is persisted in `Settings.RestoreHolds`, every start path refuses it (customer button, + app-stop `Recover()`, boot sweep, via `driveStartGate` **above** its driveless early return), the + app-stop marker is ended so nothing auto-restarts it, and the row goes red. Cleared with + `--clear-restore-hold `, **which requires a controller restart**. `--single-transaction` on + the Postgres import is a belt only; MariaDB DDL is not transactional, which is why the rollback + is the fix. - **Where an off-site restore puts the data (v0.219.0, R-356).** `ReconstituteFromOffsite` and `PlaceOffsiteRestore` resolve the destination with `Manager.GetAppDrivePath` — **the same resolver `CaptureRecoveryUnit` wrote the snapshot with**: the app's `HDD_PATH` if it declares