CONTEXT + README: the failure ladder and the operator ruling (v0.220.x)
gates / gates (push) Successful in 11s
gates / gates (push) Successful in 11s
This commit is contained in:
+45
-1
@@ -7,7 +7,51 @@
|
||||
>
|
||||
> Ask Claude Code: "Please update CONTEXT.md with what we did today"
|
||||
|
||||
Last updated: 2026-08-22 (v0.219.0 — R-356: the off-site restore refused every app that has no data drive)
|
||||
Last updated: 2026-08-22 (v0.220.2 — R-379/R-380: the undo copy goes back when a database restore fails)
|
||||
|
||||
> **2026-08-22 — v0.220.0/.1/.2 (R-379, R-380, R-381, R-382).**
|
||||
>
|
||||
> **[DECISION — the OPERATOR's, 2026-08-22] When a database replay fails AND the rollback to the
|
||||
> customer's own pre-restore copy also fails, the app is HELD STOPPED rather than started.** A
|
||||
> running app on a half-written database lets the customer type into it, and that turns a recoverable
|
||||
> state into a permanent one. The alternative — start it and mark it — was put to the operator and
|
||||
> declined. If that judgement is ever revisited, this is the sentence to revisit.
|
||||
>
|
||||
> **[DESIGN] R-379 and R-380 were ONE failure with ONE fix.** Both ended with a half-restored
|
||||
> database; the only difference was whether it looked broken (Postgres emptied and crash-looping,
|
||||
> MariaDB partly applied behind `health=healthy`). No engine flag closes that: MariaDB's DDL is not
|
||||
> transactional. Putting the customer's own copy back is what removes the half state, and it is the
|
||||
> same `ImportDump` call a person ran by hand on 2026-08-22 to recover both apps.
|
||||
>
|
||||
> **THE UNDO SET IS MATCHED ON THE RUN'S OWN STAMP, never on the `pre-restore-` prefix.** Four such
|
||||
> files accumulated on one app in one afternoon; a prefix match would replay an arbitrary older
|
||||
> state. And `writeSafetyDump` returns the SET — it used to return the first path, which for a
|
||||
> two-database app would have restored one and left the other half-written.
|
||||
>
|
||||
> **THE ROLLBACK RE-DISCOVERS THE CONTAINER.** The undo FILE is stable; the container is not. Found
|
||||
> by v0.220.0's own live walk on its first real run: `docmost-postgres` was captured as
|
||||
> `9adbc14f9af6`, re-created as `309795897b82` by the DB-only start, and the rollback's `docker exec`
|
||||
> against the dead id timed out — so the app was held for an infrastructure reason while its data was
|
||||
> recoverable. Fixed in v0.220.1. **No unit test saw it because they all inject the import seam and
|
||||
> never look at container identity.**
|
||||
>
|
||||
> **THE HOLD IS NOT `DesiredState`.** That field is the customer's stated intent; writing our failure
|
||||
> into it makes our fault indistinguishable from their choice. It is not the app-stop marker either —
|
||||
> that means "owed a restart", and a held app is not owed one; leaving it would have `Recover()` start
|
||||
> the broken app at the next boot. It is `Settings.RestoreHolds`, consulted by the shared
|
||||
> `driveStartGate` **above** its driveless early return, because the apps this exists for have no
|
||||
> drive.
|
||||
>
|
||||
> **The way out is `--clear-restore-hold <app>`, and it REQUIRES A CONTROLLER RESTART** — it runs as a
|
||||
> second process and the running controller keeps its in-memory settings. v0.220.2 makes the command
|
||||
> say so. Clearing through the running controller is the right shape later; it needs an operator tier
|
||||
> the HTTP surface does not have (it authenticates as the customer, and a customer clearing their own
|
||||
> hold is what the hold prevents).
|
||||
>
|
||||
> **Proven live on `demo-hp`**: Postgres and MariaDB both rolled back to byte-identical prior state
|
||||
> (docmost titles sha256 `8ec1fa87…` unchanged; bookstack `migrations` 102, the exact cell R-380 was
|
||||
> measured in). Evidence: `felhom.eu/documentation/audits/DRILL-r379-rollback-2026-08-22/`.
|
||||
|
||||
|
||||
> **2026-08-22 — v0.219.0 (R-356). One predicate was answering two questions.**
|
||||
>
|
||||
|
||||
@@ -571,6 +571,20 @@ Each app can define rich metadata in `.felhom.yml`:
|
||||
(they are the undo). The live recovery unit is still never overwritten, which is why the replay
|
||||
source is the scratch. Honesty surfaces (`OffsiteScratchPair`): dump age, an unstamped-pair
|
||||
warning, and the R-44 empty-dump sniff — all warn-level, none of them gates.
|
||||
- **What happens when the database replay FAILS (v0.220.0–.2, R-379/R-380).** A ladder, and every
|
||||
rung is observable: **replay → rollback → hold.** The pre-restore undo copy has always been
|
||||
taken; since v0.220.0 it is also **put back** when the replay fails — the whole set for this run,
|
||||
matched on the run's own stamp (never on the `pre-restore-` prefix, and never just the first
|
||||
file), re-applied with the DB service still up and before any restart, into a **re-discovered**
|
||||
container (the DB-only start re-creates it, so the captured id is dead by then — R-379,
|
||||
v0.220.1). The app then starts and the message says both that the restore failed and that the
|
||||
data is back. If the rollback ALSO fails the app is **held stopped** — the operator's ruling —
|
||||
the hold is persisted in `Settings.RestoreHolds`, every start path refuses it (customer button,
|
||||
app-stop `Recover()`, boot sweep, via `driveStartGate` **above** its driveless early return), the
|
||||
app-stop marker is ended so nothing auto-restarts it, and the row goes red. Cleared with
|
||||
`--clear-restore-hold <app>`, **which requires a controller restart**. `--single-transaction` on
|
||||
the Postgres import is a belt only; MariaDB DDL is not transactional, which is why the rollback
|
||||
is the fix.
|
||||
- **Where an off-site restore puts the data (v0.219.0, R-356).** `ReconstituteFromOffsite` and
|
||||
`PlaceOffsiteRestore` resolve the destination with `Manager.GetAppDrivePath` — **the same
|
||||
resolver `CaptureRecoveryUnit` wrote the snapshot with**: the app's `HDD_PATH` if it declares
|
||||
|
||||
Reference in New Issue
Block a user