Records the db_dumps decision with every consumer named, the trap that a stable db_dumps lets CaptureRecoveryUnit's already-current early return fire (so per-capture housekeeping must sit above it), and the NEGATIVE that a held app does not raise the dead-app alarm - measured, not reasoned, so nobody re-derives it.
This commit is contained in:
+36
-1
@@ -7,7 +7,42 @@
|
|||||||
>
|
>
|
||||||
> Ask Claude Code: "Please update CONTEXT.md with what we did today"
|
> Ask Claude Code: "Please update CONTEXT.md with what we did today"
|
||||||
|
|
||||||
Last updated: 2026-08-22 (v0.220.2 — R-379/R-380: the undo copy goes back when a database restore fails)
|
Last updated: 2026-08-23 (v0.221.1 — R-361: taking the undo copy destroyed the app's own backup)
|
||||||
|
|
||||||
|
> **2026-08-23 — v0.221.0/.1 (R-361), and two negatives worth as much as the fix.**
|
||||||
|
>
|
||||||
|
> **[DECISION] `db_dumps` lists the app's OWN dumps, not the `pre-restore-*` undo copies.** They are
|
||||||
|
> local material for a restore that went wrong, not part of the app's recovery set. **Every consumer
|
||||||
|
> of `Manifest.DBDumps` was grepped and named — three, all inside `recovery_unit.go`** (the
|
||||||
|
> declaration, the enumeration, the change-detection compare); none reads it for recovery, and no hub
|
||||||
|
> or agent consumer exists. Three copies per app were being pushed off-site permanently for no
|
||||||
|
> recovery value. **The files are neither deleted nor hidden** — their visibility is a separate
|
||||||
|
> recorded decision and it stands.
|
||||||
|
>
|
||||||
|
> **[TRAP, and it bit within minutes] A stable `db_dumps` lets `CaptureRecoveryUnit`'s already-current
|
||||||
|
> early return fire.** Anything that must happen on EVERY capture — bounding the undo copies — has to
|
||||||
|
> sit ABOVE that check. It did not, and the cap silently stopped applying: four copies against a cap
|
||||||
|
> of three, counted on the box. Fixed in v0.221.1. **One change made another unreachable, and only
|
||||||
|
> counting files on a real machine showed it.**
|
||||||
|
>
|
||||||
|
> **[FACT] The comment was the defect.** `writeSafetyDump` called `DumpOne` into the app's own unit
|
||||||
|
> and renamed afterwards; `DumpOne` writes the canonical `<stack>-<dbtype>.sql`, so every safety dump
|
||||||
|
> destroyed the app's real backup. The comment said the rename meant it "can never overwrite the app's
|
||||||
|
> real dump" — false as written, for four months. The fix is a DESTINATION (`DumpOneTo`), not a
|
||||||
|
> rename, and the `.tmp` derives from the final path so a nightly dump beside it cannot collide.
|
||||||
|
> **`DumpOne`'s signature did not move.**
|
||||||
|
>
|
||||||
|
> **[NEGATIVE — do not re-derive this] A HELD app does NOT raise the dead-app alarm.** It was read
|
||||||
|
> from source that it would, because it keeps its database container and so is not `StateStopped`.
|
||||||
|
> Measured on the shipped v0.220.2: it aggregates to `unhealthy`, `aggregateState` checks
|
||||||
|
> `unhealthy > 0` before the mixed-case degraded branch, and `IsDownState` excludes `unhealthy`.
|
||||||
|
> Heartbeat read `0 currently down` throughout. **No suppression was built.** The same measurement
|
||||||
|
> exposed **R-384**: an app whose database has died is `unhealthy` too, and is likewise silent.
|
||||||
|
>
|
||||||
|
> **Proven live on `demo-hp`:** the canonical dump's sha256 unchanged across a restore on both engines
|
||||||
|
> — `docmost` `5d35678349bb…`, `bookstack` `7837aa5de295…`. Evidence:
|
||||||
|
> `felhom.eu/documentation/audits/DRILL-r361-2026-08-22/`.
|
||||||
|
|
||||||
|
|
||||||
> **2026-08-22 — v0.220.0/.1/.2 (R-379, R-380, R-381, R-382).**
|
> **2026-08-22 — v0.220.0/.1/.2 (R-379, R-380, R-381, R-382).**
|
||||||
>
|
>
|
||||||
|
|||||||
@@ -1,105 +1,90 @@
|
|||||||
# REPORT — R-356: the off-site restore refused every app that has no data drive
|
# REPORT — R-361: taking the undo copy destroyed the app's own database backup
|
||||||
|
|
||||||
**Controller v0.219.0**, shipped and proven live on `demo-hp` 2026-08-22.
|
**Controller v0.221.0 → v0.221.1**, proven live on `demo-hp`.
|
||||||
|
Full record: `felhom.eu/documentation/audits/DRILL-r361-2026-08-22/`.
|
||||||
|
|
||||||
## Baselines used (re-verified at session start)
|
## Baselines and subjects
|
||||||
|
|
||||||
| repo | `main` @ | version | → |
|
| repo | HEAD at start | version |
|
||||||
|---|---|---|---|
|
|
||||||
| felhom-controller | `2da259af38cc` | v0.218.0 | **v0.219.0** |
|
|
||||||
| felhom-agent | `40d857b52711` | v0.130.0 | unchanged |
|
|
||||||
| felhom.eu (hub) | `f2edf7e54595` | v0.106.0 | unchanged |
|
|
||||||
| app-catalog | `459766cb1639` | — | unchanged |
|
|
||||||
|
|
||||||
**MinAgent stays 0.129.0.** All four hashes matched the prompt.
|
|
||||||
|
|
||||||
**The catalogue count, measured in this session and not taken from the prompt:** 53 templates,
|
|
||||||
**13 declare `needs_hdd: true`**, **40 declare `false`**, and exactly 13 carry a `backup:` block.
|
|
||||||
Agrees with §1.
|
|
||||||
|
|
||||||
## The defect
|
|
||||||
|
|
||||||
`ReconstituteFromOffsite` and `PlaceOffsiteRestore` both resolved the restore destination with the
|
|
||||||
**raw** `stackProvider.GetStackHDDPath(stack)` and refused when it was empty, with a message saying
|
|
||||||
the app is not installed. For the 40 driveless apps that value is never set, so both refused
|
|
||||||
permanently for a running, healthy app — and the remedy they offered ("reinstall it in the same
|
|
||||||
place") cannot be carried out, because those apps offer no place to pick.
|
|
||||||
|
|
||||||
The capture side never had this: `CaptureRecoveryUnit` resolves via `GetAppDrivePath`, which falls
|
|
||||||
back to `systemDataPath`. That is why the 2026-08-21 refusal could print `/mnt/sys_drive` — a
|
|
||||||
destination the backup had recorded and the restore refused to use.
|
|
||||||
|
|
||||||
## The change
|
|
||||||
|
|
||||||
- `internal/backup/backup.go` — new `Manager.isStackDeployed`, built on `knownStackNames()` →
|
|
||||||
`ListDeployedStacks()`. **Nil provider ⇒ false**, documented with its reason: the caller's next act
|
|
||||||
is a write, so "cannot tell" must fail into the recoverable refusal, not into a copy.
|
|
||||||
- `internal/backup/offbox_reconstitute.go` — the refusal now asks *installed?* first (R-351 sentence
|
|
||||||
and the recorded-drive half intact), then resolves the destination with `GetAppDrivePath`, then has
|
|
||||||
a **third, distinct** refusal for installed-but-no-resolvable-data-root. R-253 and R-351 comments
|
|
||||||
kept and extended with what the refusal stopped covering and why.
|
|
||||||
- `internal/backup/offbox_restore.go` — the same four steps in `PlaceOffsiteRestore`. The headroom
|
|
||||||
gate below it then measures the resolved namespace, which is the correct disk in both cases.
|
|
||||||
|
|
||||||
**Not changed, deliberately:** `offboxCaptureSet`'s raw `GetStackHDDPath` (`offbox_capture.go:43`) and
|
|
||||||
`offboxRestoreScratchDir`.
|
|
||||||
|
|
||||||
**`ListDeployedStacks()` semantics were verified, not assumed.** Source: `stackAdapter` skips
|
|
||||||
`!s.Deployed` (`cmd/controller/main.go:2138`). Test: positive and negative controls in
|
|
||||||
`TestR356_IsStackDeployedIsExactAndFailsClosed`. Seam: an AST walk in
|
|
||||||
`cmd/controller/r356_deployed_seam_test.go` pins both the `SetStackProvider` wiring in `main()` and
|
|
||||||
the `!s.Deployed` guard itself.
|
|
||||||
|
|
||||||
## Tests
|
|
||||||
|
|
||||||
New: `internal/backup/r356_hot_only_restore_test.go` (10 tests, Scenarios A–E across both entry
|
|
||||||
points) and `cmd/controller/r356_deployed_seam_test.go` (2 AST tests). **Test count in
|
|
||||||
`internal/backup` 315 → 325, `cmd/controller` 42 → 44.**
|
|
||||||
|
|
||||||
**Every red-proof was SEEN failing.** Mutation → observed failure:
|
|
||||||
|
|
||||||
| # | mutation applied | observed |
|
|
||||||
|---|---|---|
|
|---|---|---|
|
||||||
| A | destination back to `stackProvider.GetStackHDDPath` | `ScenarioA` failed: *"a(z) immich telepítve van, de a vezérlő nem tudja megállapítani, hová tartoznak az adatai…"* |
|
| felhom-controller | `2024ed99826717bf765803e159c074909b28bf83` | v0.220.2 → **v0.221.1** |
|
||||||
| B | `GetAppDrivePath` always returns `systemDataPath` | 3 tests failed, incl. *"placement …/sys/felhom-data/… left the app's own drive"* — this is also the **positive control** for the "13-class unchanged" absence claim |
|
| felhom-agent | `40d857b52711dd8c9c88bdd21ffcaad819a33a84` | v0.130.0, unchanged |
|
||||||
| C | `isStackDeployed` always true | 3 tests failed; the not-installed refusal was replaced by a placement-mismatch prompt |
|
| felhom.eu | `a8caa0fdde7c678bb88b5b25727044695bbdeaa9` | unchanged |
|
||||||
| D | widen `nincs telepítve` to cover the no-data-root case | both ScenarioD tests failed: *"a running app must not be told it is not installed"* |
|
| app-catalog | `459766cb16395fd1d1a66282f5cc6da59ead5924` | unchanged |
|
||||||
| E | revert `PlaceOffsiteRestore` only (fix one entry point) | `ScenarioE_PlaceDriveless` failed |
|
|
||||||
|
|
||||||
**No red-proof passed.** Green gate: `go build ./... && go vet ./... && go test ./...` — all green.
|
Hub read: golden **0.220.2**, agent **0.130.0**, min_agent **0.129.0**, floor **0.220.2** — as stated.
|
||||||
|
**Subjects FOUND, not rebuilt:** `docmost` and `bookstack`, healthy, with two sessions of data.
|
||||||
|
|
||||||
**Fixture correction, stated rather than hidden:** several existing fixtures marked an app "installed"
|
**Architecture read:** `07-backup-architecture.md` §6.1–§6.3 including the v0.220.x failure-ladder
|
||||||
by giving it an HDD path — the exact conflation this change removes. 17 tests failed on that alone
|
`[DESIGN]`; and `audits/DRILL-r379-rollback-2026-08-22/README.md`.
|
||||||
after the change; they now state deployment as its own fact (`prov.deployed`). No assertion was
|
|
||||||
weakened, and `TestPlace_UndeployedRefused` was passing for the wrong reason before.
|
|
||||||
|
|
||||||
## Live walk on `demo-hp` (controller v0.219.0, healthy)
|
## R-361 — the whole thing is one comparison
|
||||||
|
|
||||||
Endpoint-level — `claude-in-chrome` is not available on DooPlex. Full record and every artefact:
|
| app | canonical dump sha256, before | after a restore |
|
||||||
`felhom.eu/documentation/audits/DRILL-r356-hot-only-restore-2026-08-22/`.
|
|---|---|---|
|
||||||
|
| `docmost` | `5d35678349bbbdb318ac22656d4f96436b5751ef5b3b1659d48305a526df1aed` | **identical** |
|
||||||
|
| `bookstack` | `7837aa5de2955dcf3125534f015f43df42debe13c765793d3b75173abcc55e7b` | **identical** |
|
||||||
|
|
||||||
**Driveless (`privatebin`):** planted through the app's own JSON API + a direct file plant with a
|
**Before the fix, on the same box: neither app had a canonical dump at all.**
|
||||||
Hungarian accented name; findability proved by plant→find→remove→fail-to-find; off-sited; **11 files
|
|
||||||
deleted outright**; scratch prepared (snapshot `306accff`, **unit-only** — the shape never driven to
|
|
||||||
completion before); `POST /backup/offbox/reconstitute` — **no refusal**; **15/15 files back byte for
|
|
||||||
byte**. Message (160 bytes): *„A(z) privatebin: 0 fájl és 1 adatkötet visszaállítva (mentés:
|
|
||||||
2026-08-22 13:14) — az alkalmazás újraindult. Ennek az alkalmazásnak nincs adatbázisa."* It **names
|
|
||||||
the volume**, so it is not the R-354 shape.
|
|
||||||
|
|
||||||
**Drive app (`calibre-web`):** same walk, **3/3 byte-identical**, placed on `/mnt/felhom-drives/hdd_1`;
|
`DumpOneTo` takes an explicit final path and derives its `.tmp` from it. `DumpOne`'s signature did not
|
||||||
`/mnt/sys_drive/felhom-data/userdata/media` does not exist, so nothing was misplaced onto the SSD.
|
move. The comment that asserted the old behaviour was safe now states the invariant and how it is
|
||||||
|
enforced.
|
||||||
|
|
||||||
**Every downstream leg held on first run.** No finding raised against them.
|
## Manifest — every consumer named, as required
|
||||||
|
|
||||||
## Also in this session
|
`Manifest.DBDumps` has **three** references, all in `internal/backup/recovery_unit.go`: the
|
||||||
|
declaration (`:55`), the enumeration, and the already-current compare. **None reads it for recovery;
|
||||||
|
no hub or agent consumer exists.** Undo copies are therefore excluded from `db_dumps`. Verified live:
|
||||||
|
`db_dumps: ['docmost-postgres.sql']` with **4 undo copies on disk** as the positive control.
|
||||||
|
|
||||||
CI run **387** (job 386) failed on `instructions_gate` — **a real gate bug, not mine**: `ef6ac6f` in
|
## Part 3 — the measurement that cancelled Part 2
|
||||||
`felhom.eu` compressed closed rows into `CLOSED-ITEMS.md`, which the gate does not read, so every
|
|
||||||
citation of a compressed item became "a reference to nothing". Fixed in
|
**The reading did not reproduce.** A held app aggregates to `unhealthy`; `IsDownState` is
|
||||||
`felhom.eu/scripts/instructions_gate.py` (commit `e18668f`); both controls still convict.
|
`{stopped, exited, degraded}`; the dead-app heartbeat read `0 currently down` across scans 600 and
|
||||||
|
620 while a hold was in force. **Part 2 dropped in full**, no register row opened, negative recorded
|
||||||
|
in the capability map.
|
||||||
|
|
||||||
|
**The positive control took three attempts and that is the finding.** Two live attempts failed —
|
||||||
|
`privatebin` went `stopped` (whitelisted by design) and `bookstack` went `degraded` then `unhealthy`.
|
||||||
|
The control that works is at the detector's own layer: `classifyRunStates` is pure and raises the
|
||||||
|
banner for `degraded`/`exited` while staying silent for the states measured live.
|
||||||
|
|
||||||
|
## Findings filed
|
||||||
|
|
||||||
|
- **R-383 (MEDIUM)** — the double-failure message says the previous state's backup **exists** while
|
||||||
|
naming the very file whose absence caused the failure. Observed twice, on v0.220.2 and v0.221.1.
|
||||||
|
**R-361's own class**, one surface over.
|
||||||
|
- **R-384 (MEDIUM)** — an app whose **database** has died reads `unhealthy` and raises no dead-app
|
||||||
|
alarm. `bookstack-db` stopped at 21:27:01; `0 currently down` throughout.
|
||||||
|
|
||||||
|
**Register: `OPEN-ITEMS.md` 325 236 → 327 266 bytes; `CLOSED-ITEMS.md` 66 777 → 68 464.** R-361
|
||||||
|
compressed; full text `git show a8caa0fdde7c:documentation/backlog/OPEN-ITEMS.md`.
|
||||||
|
|
||||||
|
## Two defects in this session's own work
|
||||||
|
|
||||||
|
1. **Two red-proofs passed first time**, both reported. The behavioural tests inject the dump seam, so
|
||||||
|
a mutation *inside* `DumpOneTo` was invisible; and Part 1.3 had no test at all. Guards added at the
|
||||||
|
layer each defect lives in; both mutations then convicted.
|
||||||
|
2. **One change made another unreachable.** A stable `db_dumps` let the already-current early return
|
||||||
|
fire, and the undo prune sat after it — **four copies against a cap of three, counted on the box**.
|
||||||
|
Fixed in v0.221.1; the cap now holds at 3 on both apps, verified live.
|
||||||
|
|
||||||
|
**Test count 1485 → 1493.** Green gate clean; controller gates all OK; **no push used `--no-verify`.**
|
||||||
|
|
||||||
|
## Part 4 — the double-failure path on what ships
|
||||||
|
|
||||||
|
Trigger: **the undo copy is lost between being written and being needed** — realistic, and one of only
|
||||||
|
two ways a rollback can fail. Verified on 0.221.1: held, refused by the customer button, refused after
|
||||||
|
a full controller restart, then listed/cleared/started via the operator route. Message captured
|
||||||
|
verbatim (369 bytes, hex in the drill record) with **no engine output**.
|
||||||
|
|
||||||
|
## Teardown
|
||||||
|
|
||||||
|
Nothing provisioned; 20 containers healthy; `docmost` still holds its 4 pages and 1 user; both
|
||||||
|
canonical dumps present; **no app left held**; no hub-side record created.
|
||||||
|
|
||||||
## Operator follow-up
|
## Operator follow-up
|
||||||
|
|
||||||
1. Bake and **vouch** a golden carrying controller 0.219.0. The `golden-currency` gate is red until
|
Vouch a golden carrying **0.221.1** (baked in this session), then raise the floor to 0.221.1 last, in
|
||||||
then, correctly: a machine installed now still receives 0.218.0.
|
its own save.
|
||||||
2. **Then** raise the update floor to 0.219.0 — last, in a separate save.
|
|
||||||
|
|||||||
@@ -571,6 +571,14 @@ Each app can define rich metadata in `.felhom.yml`:
|
|||||||
(they are the undo). The live recovery unit is still never overwritten, which is why the replay
|
(they are the undo). The live recovery unit is still never overwritten, which is why the replay
|
||||||
source is the scratch. Honesty surfaces (`OffsiteScratchPair`): dump age, an unstamped-pair
|
source is the scratch. Honesty surfaces (`OffsiteScratchPair`): dump age, an unstamped-pair
|
||||||
warning, and the R-44 empty-dump sniff — all warn-level, none of them gates.
|
warning, and the R-44 empty-dump sniff — all warn-level, none of them gates.
|
||||||
|
- **Where the pre-restore undo copy is written (v0.221.0–.1, R-361).** `DumpOneTo` takes an
|
||||||
|
EXPLICIT final path; `DumpOne` keeps its signature and calls it with the canonical
|
||||||
|
`<stack>-<dbtype>.sql`. The safety dump asks for `pre-restore-<stamp>-…` **directly** — it used to
|
||||||
|
dump to the canonical name and rename afterwards, which destroyed the app's own backup on every
|
||||||
|
restore. The `.tmp` derives from the final path, so a nightly dump and a safety dump in the same
|
||||||
|
directory cannot share a scratch file. `db_dumps` in the manifest lists the app's own dumps only;
|
||||||
|
the undo copies stay on disk and stay visible, capped at 3 per app, pruned from the capture side
|
||||||
|
**above** the already-current early return.
|
||||||
- **What happens when the database replay FAILS (v0.220.0–.2, R-379/R-380).** A ladder, and every
|
- **What happens when the database replay FAILS (v0.220.0–.2, R-379/R-380).** A ladder, and every
|
||||||
rung is observable: **replay → rollback → hold.** The pre-restore undo copy has always been
|
rung is observable: **replay → rollback → hold.** The pre-restore undo copy has always been
|
||||||
taken; since v0.220.0 it is also **put back** when the replay fails — the whole set for this run,
|
taken; since v0.220.0 it is also **put back** when the replay fails — the whole set for this run,
|
||||||
|
|||||||
Reference in New Issue
Block a user