CONTEXT + README: the failure ladder and the operator ruling (v0.220.x)
gates / gates (push) Successful in 11s

This commit is contained in:
2026-08-22 18:40:33 +02:00
parent 1a2405e86f
commit 2024ed9982
2 changed files with 59 additions and 1 deletions
+45 -1
View File
@@ -7,7 +7,51 @@
>
> Ask Claude Code: "Please update CONTEXT.md with what we did today"
Last updated: 2026-08-22 (v0.219.0 — R-356: the off-site restore refused every app that has no data drive)
Last updated: 2026-08-22 (v0.220.2 — R-379/R-380: the undo copy goes back when a database restore fails)
> **2026-08-22 — v0.220.0/.1/.2 (R-379, R-380, R-381, R-382).**
>
> **[DECISION — the OPERATOR's, 2026-08-22] When a database replay fails AND the rollback to the
> customer's own pre-restore copy also fails, the app is HELD STOPPED rather than started.** A
> running app on a half-written database lets the customer type into it, and that turns a recoverable
> state into a permanent one. The alternative — start it and mark it — was put to the operator and
> declined. If that judgement is ever revisited, this is the sentence to revisit.
>
> **[DESIGN] R-379 and R-380 were ONE failure with ONE fix.** Both ended with a half-restored
> database; the only difference was whether it looked broken (Postgres emptied and crash-looping,
> MariaDB partly applied behind `health=healthy`). No engine flag closes that: MariaDB's DDL is not
> transactional. Putting the customer's own copy back is what removes the half state, and it is the
> same `ImportDump` call a person ran by hand on 2026-08-22 to recover both apps.
>
> **THE UNDO SET IS MATCHED ON THE RUN'S OWN STAMP, never on the `pre-restore-` prefix.** Four such
> files accumulated on one app in one afternoon; a prefix match would replay an arbitrary older
> state. And `writeSafetyDump` returns the SET — it used to return the first path, which for a
> two-database app would have restored one and left the other half-written.
>
> **THE ROLLBACK RE-DISCOVERS THE CONTAINER.** The undo FILE is stable; the container is not. Found
> by v0.220.0's own live walk on its first real run: `docmost-postgres` was captured as
> `9adbc14f9af6`, re-created as `309795897b82` by the DB-only start, and the rollback's `docker exec`
> against the dead id timed out — so the app was held for an infrastructure reason while its data was
> recoverable. Fixed in v0.220.1. **No unit test saw it because they all inject the import seam and
> never look at container identity.**
>
> **THE HOLD IS NOT `DesiredState`.** That field is the customer's stated intent; writing our failure
> into it makes our fault indistinguishable from their choice. It is not the app-stop marker either —
> that means "owed a restart", and a held app is not owed one; leaving it would have `Recover()` start
> the broken app at the next boot. It is `Settings.RestoreHolds`, consulted by the shared
> `driveStartGate` **above** its driveless early return, because the apps this exists for have no
> drive.
>
> **The way out is `--clear-restore-hold <app>`, and it REQUIRES A CONTROLLER RESTART** — it runs as a
> second process and the running controller keeps its in-memory settings. v0.220.2 makes the command
> say so. Clearing through the running controller is the right shape later; it needs an operator tier
> the HTTP surface does not have (it authenticates as the customer, and a customer clearing their own
> hold is what the hold prevents).
>
> **Proven live on `demo-hp`**: Postgres and MariaDB both rolled back to byte-identical prior state
> (docmost titles sha256 `8ec1fa87…` unchanged; bookstack `migrations` 102, the exact cell R-380 was
> measured in). Evidence: `felhom.eu/documentation/audits/DRILL-r379-rollback-2026-08-22/`.
> **2026-08-22 — v0.219.0 (R-356). One predicate was answering two questions.**
>
+14
View File
@@ -571,6 +571,20 @@ Each app can define rich metadata in `.felhom.yml`:
(they are the undo). The live recovery unit is still never overwritten, which is why the replay
source is the scratch. Honesty surfaces (`OffsiteScratchPair`): dump age, an unstamped-pair
warning, and the R-44 empty-dump sniff — all warn-level, none of them gates.
- **What happens when the database replay FAILS (v0.220.0–.2, R-379/R-380).** A ladder, and every
rung is observable: **replay → rollback → hold.** The pre-restore undo copy has always been
taken; since v0.220.0 it is also **put back** when the replay fails — the whole set for this run,
matched on the run's own stamp (never on the `pre-restore-` prefix, and never just the first
file), re-applied with the DB service still up and before any restart, into a **re-discovered**
container (the DB-only start re-creates it, so the captured id is dead by then — R-379,
v0.220.1). The app then starts and the message says both that the restore failed and that the
data is back. If the rollback ALSO fails the app is **held stopped** — the operator's ruling —
the hold is persisted in `Settings.RestoreHolds`, every start path refuses it (customer button,
app-stop `Recover()`, boot sweep, via `driveStartGate` **above** its driveless early return), the
app-stop marker is ended so nothing auto-restarts it, and the row goes red. Cleared with
`--clear-restore-hold <app>`, **which requires a controller restart**. `--single-transaction` on
the Postgres import is a belt only; MariaDB DDL is not transactional, which is why the rollback
is the fix.
- **Where an off-site restore puts the data (v0.219.0, R-356).** `ReconstituteFromOffsite` and
`PlaceOffsiteRestore` resolve the destination with `Manager.GetAppDrivePath` — **the same
resolver `CaptureRecoveryUnit` wrote the snapshot with**: the app's `HDD_PATH` if it declares