Records the db_dumps decision with every consumer named, the trap that a stable db_dumps lets CaptureRecoveryUnit's already-current early return fire (so per-capture housekeeping must sit above it), and the NEGATIVE that a held app does not raise the dead-app alarm - measured, not reasoned, so nobody re-derives it.
This commit is contained in:
+36
-1
@@ -7,7 +7,42 @@
|
||||
>
|
||||
> Ask Claude Code: "Please update CONTEXT.md with what we did today"
|
||||
|
||||
Last updated: 2026-08-22 (v0.220.2 — R-379/R-380: the undo copy goes back when a database restore fails)
|
||||
Last updated: 2026-08-23 (v0.221.1 — R-361: taking the undo copy destroyed the app's own backup)
|
||||
|
||||
> **2026-08-23 — v0.221.0/.1 (R-361), and two negatives worth as much as the fix.**
|
||||
>
|
||||
> **[DECISION] `db_dumps` lists the app's OWN dumps, not the `pre-restore-*` undo copies.** They are
|
||||
> local material for a restore that went wrong, not part of the app's recovery set. **Every consumer
|
||||
> of `Manifest.DBDumps` was grepped and named — three, all inside `recovery_unit.go`** (the
|
||||
> declaration, the enumeration, the change-detection compare); none reads it for recovery, and no hub
|
||||
> or agent consumer exists. Three copies per app were being pushed off-site permanently for no
|
||||
> recovery value. **The files are neither deleted nor hidden** — their visibility is a separate
|
||||
> recorded decision and it stands.
|
||||
>
|
||||
> **[TRAP, and it bit within minutes] A stable `db_dumps` lets `CaptureRecoveryUnit`'s already-current
|
||||
> early return fire.** Anything that must happen on EVERY capture — bounding the undo copies — has to
|
||||
> sit ABOVE that check. It did not, and the cap silently stopped applying: four copies against a cap
|
||||
> of three, counted on the box. Fixed in v0.221.1. **One change made another unreachable, and only
|
||||
> counting files on a real machine showed it.**
|
||||
>
|
||||
> **[FACT] The comment was the defect.** `writeSafetyDump` called `DumpOne` into the app's own unit
|
||||
> and renamed afterwards; `DumpOne` writes the canonical `<stack>-<dbtype>.sql`, so every safety dump
|
||||
> destroyed the app's real backup. The comment said the rename meant it "can never overwrite the app's
|
||||
> real dump" — false as written, for four months. The fix is a DESTINATION (`DumpOneTo`), not a
|
||||
> rename, and the `.tmp` derives from the final path so a nightly dump beside it cannot collide.
|
||||
> **`DumpOne`'s signature did not move.**
|
||||
>
|
||||
> **[NEGATIVE — do not re-derive this] A HELD app does NOT raise the dead-app alarm.** It was read
|
||||
> from source that it would, because it keeps its database container and so is not `StateStopped`.
|
||||
> Measured on the shipped v0.220.2: it aggregates to `unhealthy`, `aggregateState` checks
|
||||
> `unhealthy > 0` before the mixed-case degraded branch, and `IsDownState` excludes `unhealthy`.
|
||||
> Heartbeat read `0 currently down` throughout. **No suppression was built.** The same measurement
|
||||
> exposed **R-384**: an app whose database has died is `unhealthy` too, and is likewise silent.
|
||||
>
|
||||
> **Proven live on `demo-hp`:** the canonical dump's sha256 unchanged across a restore on both engines
|
||||
> — `docmost` `5d35678349bb…`, `bookstack` `7837aa5de295…`. Evidence:
|
||||
> `felhom.eu/documentation/audits/DRILL-r361-2026-08-22/`.
|
||||
|
||||
|
||||
> **2026-08-22 — v0.220.0/.1/.2 (R-379, R-380, R-381, R-382).**
|
||||
>
|
||||
|
||||
Reference in New Issue
Block a user