CONTEXT + README: the failure ladder and the operator ruling (v0.220.x)
gates / gates (push) Successful in 11s

This commit is contained in:
2026-08-22 18:40:33 +02:00
parent 1a2405e86f
commit 2024ed9982
2 changed files with 59 additions and 1 deletions
+45 -1
View File
@@ -7,7 +7,51 @@
> >
> Ask Claude Code: "Please update CONTEXT.md with what we did today" > Ask Claude Code: "Please update CONTEXT.md with what we did today"
Last updated: 2026-08-22 (v0.219.0 — R-356: the off-site restore refused every app that has no data drive) Last updated: 2026-08-22 (v0.220.2 — R-379/R-380: the undo copy goes back when a database restore fails)
> **2026-08-22 — v0.220.0/.1/.2 (R-379, R-380, R-381, R-382).**
>
> **[DECISION — the OPERATOR's, 2026-08-22] When a database replay fails AND the rollback to the
> customer's own pre-restore copy also fails, the app is HELD STOPPED rather than started.** A
> running app on a half-written database lets the customer type into it, and that turns a recoverable
> state into a permanent one. The alternative — start it and mark it — was put to the operator and
> declined. If that judgement is ever revisited, this is the sentence to revisit.
>
> **[DESIGN] R-379 and R-380 were ONE failure with ONE fix.** Both ended with a half-restored
> database; the only difference was whether it looked broken (Postgres emptied and crash-looping,
> MariaDB partly applied behind `health=healthy`). No engine flag closes that: MariaDB's DDL is not
> transactional. Putting the customer's own copy back is what removes the half state, and it is the
> same `ImportDump` call a person ran by hand on 2026-08-22 to recover both apps.
>
> **THE UNDO SET IS MATCHED ON THE RUN'S OWN STAMP, never on the `pre-restore-` prefix.** Four such
> files accumulated on one app in one afternoon; a prefix match would replay an arbitrary older
> state. And `writeSafetyDump` returns the SET — it used to return the first path, which for a
> two-database app would have restored one and left the other half-written.
>
> **THE ROLLBACK RE-DISCOVERS THE CONTAINER.** The undo FILE is stable; the container is not. Found
> by v0.220.0's own live walk on its first real run: `docmost-postgres` was captured as
> `9adbc14f9af6`, re-created as `309795897b82` by the DB-only start, and the rollback's `docker exec`
> against the dead id timed out — so the app was held for an infrastructure reason while its data was
> recoverable. Fixed in v0.220.1. **No unit test saw it because they all inject the import seam and
> never look at container identity.**
>
> **THE HOLD IS NOT `DesiredState`.** That field is the customer's stated intent; writing our failure
> into it makes our fault indistinguishable from their choice. It is not the app-stop marker either —
> that means "owed a restart", and a held app is not owed one; leaving it would have `Recover()` start
> the broken app at the next boot. It is `Settings.RestoreHolds`, consulted by the shared
> `driveStartGate` **above** its driveless early return, because the apps this exists for have no
> drive.
>
> **The way out is `--clear-restore-hold <app>`, and it REQUIRES A CONTROLLER RESTART** — it runs as a
> second process and the running controller keeps its in-memory settings. v0.220.2 makes the command
> say so. Clearing through the running controller is the right shape later; it needs an operator tier
> the HTTP surface does not have (it authenticates as the customer, and a customer clearing their own
> hold is what the hold prevents).
>
> **Proven live on `demo-hp`**: Postgres and MariaDB both rolled back to byte-identical prior state
> (docmost titles sha256 `8ec1fa87…` unchanged; bookstack `migrations` 102, the exact cell R-380 was
> measured in). Evidence: `felhom.eu/documentation/audits/DRILL-r379-rollback-2026-08-22/`.
> **2026-08-22 — v0.219.0 (R-356). One predicate was answering two questions.** > **2026-08-22 — v0.219.0 (R-356). One predicate was answering two questions.**
> >
+14
View File
@@ -571,6 +571,20 @@ Each app can define rich metadata in `.felhom.yml`:
(they are the undo). The live recovery unit is still never overwritten, which is why the replay (they are the undo). The live recovery unit is still never overwritten, which is why the replay
source is the scratch. Honesty surfaces (`OffsiteScratchPair`): dump age, an unstamped-pair source is the scratch. Honesty surfaces (`OffsiteScratchPair`): dump age, an unstamped-pair
warning, and the R-44 empty-dump sniff — all warn-level, none of them gates. warning, and the R-44 empty-dump sniff — all warn-level, none of them gates.
- **What happens when the database replay FAILS (v0.220.0–.2, R-379/R-380).** A ladder, and every
rung is observable: **replay → rollback → hold.** The pre-restore undo copy has always been
taken; since v0.220.0 it is also **put back** when the replay fails — the whole set for this run,
matched on the run's own stamp (never on the `pre-restore-` prefix, and never just the first
file), re-applied with the DB service still up and before any restart, into a **re-discovered**
container (the DB-only start re-creates it, so the captured id is dead by then — R-379,
v0.220.1). The app then starts and the message says both that the restore failed and that the
data is back. If the rollback ALSO fails the app is **held stopped** — the operator's ruling —
the hold is persisted in `Settings.RestoreHolds`, every start path refuses it (customer button,
app-stop `Recover()`, boot sweep, via `driveStartGate` **above** its driveless early return), the
app-stop marker is ended so nothing auto-restarts it, and the row goes red. Cleared with
`--clear-restore-hold <app>`, **which requires a controller restart**. `--single-transaction` on
the Postgres import is a belt only; MariaDB DDL is not transactional, which is why the rollback
is the fix.
- **Where an off-site restore puts the data (v0.219.0, R-356).** `ReconstituteFromOffsite` and - **Where an off-site restore puts the data (v0.219.0, R-356).** `ReconstituteFromOffsite` and
`PlaceOffsiteRestore` resolve the destination with `Manager.GetAppDrivePath` — **the same `PlaceOffsiteRestore` resolve the destination with `Manager.GetAppDrivePath` — **the same
resolver `CaptureRecoveryUnit` wrote the snapshot with**: the app's `HDD_PATH` if it declares resolver `CaptureRecoveryUnit` wrote the snapshot with**: the app's `HDD_PATH` if it declares