R-361 docs: CONTEXT decision, README, REPORT
gates / gates (push) Successful in 11s

Records the db_dumps decision with every consumer named, the trap that a stable
db_dumps lets CaptureRecoveryUnit's already-current early return fire (so
per-capture housekeeping must sit above it), and the NEGATIVE that a held app
does not raise the dead-app alarm - measured, not reasoned, so nobody re-derives
it.
This commit is contained in:
2026-08-23 00:26:32 +02:00
parent 810b18ab8e
commit f7881787f4
3 changed files with 118 additions and 90 deletions
+36 -1
View File
@@ -7,7 +7,42 @@
> >
> Ask Claude Code: "Please update CONTEXT.md with what we did today" > Ask Claude Code: "Please update CONTEXT.md with what we did today"
Last updated: 2026-08-22 (v0.220.2 — R-379/R-380: the undo copy goes back when a database restore fails) Last updated: 2026-08-23 (v0.221.1 — R-361: taking the undo copy destroyed the app's own backup)
> **2026-08-23 — v0.221.0/.1 (R-361), and two negatives worth as much as the fix.**
>
> **[DECISION] `db_dumps` lists the app's OWN dumps, not the `pre-restore-*` undo copies.** They are
> local material for a restore that went wrong, not part of the app's recovery set. **Every consumer
> of `Manifest.DBDumps` was grepped and named — three, all inside `recovery_unit.go`** (the
> declaration, the enumeration, the change-detection compare); none reads it for recovery, and no hub
> or agent consumer exists. Three copies per app were being pushed off-site permanently for no
> recovery value. **The files are neither deleted nor hidden** — their visibility is a separate
> recorded decision and it stands.
>
> **[TRAP, and it bit within minutes] A stable `db_dumps` lets `CaptureRecoveryUnit`'s already-current
> early return fire.** Anything that must happen on EVERY capture — bounding the undo copies — has to
> sit ABOVE that check. It did not, and the cap silently stopped applying: four copies against a cap
> of three, counted on the box. Fixed in v0.221.1. **One change made another unreachable, and only
> counting files on a real machine showed it.**
>
> **[FACT] The comment was the defect.** `writeSafetyDump` called `DumpOne` into the app's own unit
> and renamed afterwards; `DumpOne` writes the canonical `<stack>-<dbtype>.sql`, so every safety dump
> destroyed the app's real backup. The comment said the rename meant it "can never overwrite the app's
> real dump" — false as written, for four months. The fix is a DESTINATION (`DumpOneTo`), not a
> rename, and the `.tmp` derives from the final path so a nightly dump beside it cannot collide.
> **`DumpOne`'s signature did not move.**
>
> **[NEGATIVE — do not re-derive this] A HELD app does NOT raise the dead-app alarm.** It was read
> from source that it would, because it keeps its database container and so is not `StateStopped`.
> Measured on the shipped v0.220.2: it aggregates to `unhealthy`, `aggregateState` checks
> `unhealthy > 0` before the mixed-case degraded branch, and `IsDownState` excludes `unhealthy`.
> Heartbeat read `0 currently down` throughout. **No suppression was built.** The same measurement
> exposed **R-384**: an app whose database has died is `unhealthy` too, and is likewise silent.
>
> **Proven live on `demo-hp`:** the canonical dump's sha256 unchanged across a restore on both engines
> — `docmost` `5d35678349bb…`, `bookstack` `7837aa5de295…`. Evidence:
> `felhom.eu/documentation/audits/DRILL-r361-2026-08-22/`.
> **2026-08-22 — v0.220.0/.1/.2 (R-379, R-380, R-381, R-382).** > **2026-08-22 — v0.220.0/.1/.2 (R-379, R-380, R-381, R-382).**
> >
+74 -89
View File
@@ -1,105 +1,90 @@
# REPORT — R-356: the off-site restore refused every app that has no data drive # REPORT — R-361: taking the undo copy destroyed the app's own database backup
**Controller v0.219.0**, shipped and proven live on `demo-hp` 2026-08-22. **Controller v0.221.0 → v0.221.1**, proven live on `demo-hp`.
Full record: `felhom.eu/documentation/audits/DRILL-r361-2026-08-22/`.
## Baselines used (re-verified at session start) ## Baselines and subjects
| repo | `main` @ | version | → | | repo | HEAD at start | version |
|---|---|---|---|
| felhom-controller | `2da259af38cc` | v0.218.0 | **v0.219.0** |
| felhom-agent | `40d857b52711` | v0.130.0 | unchanged |
| felhom.eu (hub) | `f2edf7e54595` | v0.106.0 | unchanged |
| app-catalog | `459766cb1639` | — | unchanged |
**MinAgent stays 0.129.0.** All four hashes matched the prompt.
**The catalogue count, measured in this session and not taken from the prompt:** 53 templates,
**13 declare `needs_hdd: true`**, **40 declare `false`**, and exactly 13 carry a `backup:` block.
Agrees with §1.
## The defect
`ReconstituteFromOffsite` and `PlaceOffsiteRestore` both resolved the restore destination with the
**raw** `stackProvider.GetStackHDDPath(stack)` and refused when it was empty, with a message saying
the app is not installed. For the 40 driveless apps that value is never set, so both refused
permanently for a running, healthy app — and the remedy they offered ("reinstall it in the same
place") cannot be carried out, because those apps offer no place to pick.
The capture side never had this: `CaptureRecoveryUnit` resolves via `GetAppDrivePath`, which falls
back to `systemDataPath`. That is why the 2026-08-21 refusal could print `/mnt/sys_drive` — a
destination the backup had recorded and the restore refused to use.
## The change
- `internal/backup/backup.go` — new `Manager.isStackDeployed`, built on `knownStackNames()` →
`ListDeployedStacks()`. **Nil provider ⇒ false**, documented with its reason: the caller's next act
is a write, so "cannot tell" must fail into the recoverable refusal, not into a copy.
- `internal/backup/offbox_reconstitute.go` — the refusal now asks *installed?* first (R-351 sentence
and the recorded-drive half intact), then resolves the destination with `GetAppDrivePath`, then has
a **third, distinct** refusal for installed-but-no-resolvable-data-root. R-253 and R-351 comments
kept and extended with what the refusal stopped covering and why.
- `internal/backup/offbox_restore.go` — the same four steps in `PlaceOffsiteRestore`. The headroom
gate below it then measures the resolved namespace, which is the correct disk in both cases.
**Not changed, deliberately:** `offboxCaptureSet`'s raw `GetStackHDDPath` (`offbox_capture.go:43`) and
`offboxRestoreScratchDir`.
**`ListDeployedStacks()` semantics were verified, not assumed.** Source: `stackAdapter` skips
`!s.Deployed` (`cmd/controller/main.go:2138`). Test: positive and negative controls in
`TestR356_IsStackDeployedIsExactAndFailsClosed`. Seam: an AST walk in
`cmd/controller/r356_deployed_seam_test.go` pins both the `SetStackProvider` wiring in `main()` and
the `!s.Deployed` guard itself.
## Tests
New: `internal/backup/r356_hot_only_restore_test.go` (10 tests, Scenarios A–E across both entry
points) and `cmd/controller/r356_deployed_seam_test.go` (2 AST tests). **Test count in
`internal/backup` 315 → 325, `cmd/controller` 42 → 44.**
**Every red-proof was SEEN failing.** Mutation → observed failure:
| # | mutation applied | observed |
|---|---|---| |---|---|---|
| A | destination back to `stackProvider.GetStackHDDPath` | `ScenarioA` failed: *"a(z) immich telepítve van, de a vezérlő nem tudja megállapítani, hová tartoznak az adatai…"* | | felhom-controller | `2024ed99826717bf765803e159c074909b28bf83` | v0.220.2 → **v0.221.1** |
| B | `GetAppDrivePath` always returns `systemDataPath` | 3 tests failed, incl. *"placement …/sys/felhom-data/… left the app's own drive"* — this is also the **positive control** for the "13-class unchanged" absence claim | | felhom-agent | `40d857b52711dd8c9c88bdd21ffcaad819a33a84` | v0.130.0, unchanged |
| C | `isStackDeployed` always true | 3 tests failed; the not-installed refusal was replaced by a placement-mismatch prompt | | felhom.eu | `a8caa0fdde7c678bb88b5b25727044695bbdeaa9` | unchanged |
| D | widen `nincs telepítve` to cover the no-data-root case | both ScenarioD tests failed: *"a running app must not be told it is not installed"* | | app-catalog | `459766cb16395fd1d1a66282f5cc6da59ead5924` | unchanged |
| E | revert `PlaceOffsiteRestore` only (fix one entry point) | `ScenarioE_PlaceDriveless` failed |
**No red-proof passed.** Green gate: `go build ./... && go vet ./... && go test ./...` — all green. Hub read: golden **0.220.2**, agent **0.130.0**, min_agent **0.129.0**, floor **0.220.2** — as stated.
**Subjects FOUND, not rebuilt:** `docmost` and `bookstack`, healthy, with two sessions of data.
**Fixture correction, stated rather than hidden:** several existing fixtures marked an app "installed" **Architecture read:** `07-backup-architecture.md` §6.1–§6.3 including the v0.220.x failure-ladder
by giving it an HDD path — the exact conflation this change removes. 17 tests failed on that alone `[DESIGN]`; and `audits/DRILL-r379-rollback-2026-08-22/README.md`.
after the change; they now state deployment as its own fact (`prov.deployed`). No assertion was
weakened, and `TestPlace_UndeployedRefused` was passing for the wrong reason before.
## Live walk on `demo-hp` (controller v0.219.0, healthy) ## R-361 — the whole thing is one comparison
Endpoint-level — `claude-in-chrome` is not available on DooPlex. Full record and every artefact: | app | canonical dump sha256, before | after a restore |
`felhom.eu/documentation/audits/DRILL-r356-hot-only-restore-2026-08-22/`. |---|---|---|
| `docmost` | `5d35678349bbbdb318ac22656d4f96436b5751ef5b3b1659d48305a526df1aed` | **identical** |
| `bookstack` | `7837aa5de2955dcf3125534f015f43df42debe13c765793d3b75173abcc55e7b` | **identical** |
**Driveless (`privatebin`):** planted through the app's own JSON API + a direct file plant with a **Before the fix, on the same box: neither app had a canonical dump at all.**
Hungarian accented name; findability proved by plant→find→remove→fail-to-find; off-sited; **11 files
deleted outright**; scratch prepared (snapshot `306accff`, **unit-only** — the shape never driven to
completion before); `POST /backup/offbox/reconstitute` — **no refusal**; **15/15 files back byte for
byte**. Message (160 bytes): *„A(z) privatebin: 0 fájl és 1 adatkötet visszaállítva (mentés:
2026-08-22 13:14) — az alkalmazás újraindult. Ennek az alkalmazásnak nincs adatbázisa."* It **names
the volume**, so it is not the R-354 shape.
**Drive app (`calibre-web`):** same walk, **3/3 byte-identical**, placed on `/mnt/felhom-drives/hdd_1`; `DumpOneTo` takes an explicit final path and derives its `.tmp` from it. `DumpOne`'s signature did not
`/mnt/sys_drive/felhom-data/userdata/media` does not exist, so nothing was misplaced onto the SSD. move. The comment that asserted the old behaviour was safe now states the invariant and how it is
enforced.
**Every downstream leg held on first run.** No finding raised against them. ## Manifest — every consumer named, as required
## Also in this session `Manifest.DBDumps` has **three** references, all in `internal/backup/recovery_unit.go`: the
declaration (`:55`), the enumeration, and the already-current compare. **None reads it for recovery;
no hub or agent consumer exists.** Undo copies are therefore excluded from `db_dumps`. Verified live:
`db_dumps: ['docmost-postgres.sql']` with **4 undo copies on disk** as the positive control.
CI run **387** (job 386) failed on `instructions_gate` — **a real gate bug, not mine**: `ef6ac6f` in ## Part 3 — the measurement that cancelled Part 2
`felhom.eu` compressed closed rows into `CLOSED-ITEMS.md`, which the gate does not read, so every
citation of a compressed item became "a reference to nothing". Fixed in **The reading did not reproduce.** A held app aggregates to `unhealthy`; `IsDownState` is
`felhom.eu/scripts/instructions_gate.py` (commit `e18668f`); both controls still convict. `{stopped, exited, degraded}`; the dead-app heartbeat read `0 currently down` across scans 600 and
620 while a hold was in force. **Part 2 dropped in full**, no register row opened, negative recorded
in the capability map.
**The positive control took three attempts and that is the finding.** Two live attempts failed —
`privatebin` went `stopped` (whitelisted by design) and `bookstack` went `degraded` then `unhealthy`.
The control that works is at the detector's own layer: `classifyRunStates` is pure and raises the
banner for `degraded`/`exited` while staying silent for the states measured live.
## Findings filed
- **R-383 (MEDIUM)** — the double-failure message says the previous state's backup **exists** while
naming the very file whose absence caused the failure. Observed twice, on v0.220.2 and v0.221.1.
**R-361's own class**, one surface over.
- **R-384 (MEDIUM)** — an app whose **database** has died reads `unhealthy` and raises no dead-app
alarm. `bookstack-db` stopped at 21:27:01; `0 currently down` throughout.
**Register: `OPEN-ITEMS.md` 325 236 → 327 266 bytes; `CLOSED-ITEMS.md` 66 777 → 68 464.** R-361
compressed; full text `git show a8caa0fdde7c:documentation/backlog/OPEN-ITEMS.md`.
## Two defects in this session's own work
1. **Two red-proofs passed first time**, both reported. The behavioural tests inject the dump seam, so
a mutation *inside* `DumpOneTo` was invisible; and Part 1.3 had no test at all. Guards added at the
layer each defect lives in; both mutations then convicted.
2. **One change made another unreachable.** A stable `db_dumps` let the already-current early return
fire, and the undo prune sat after it — **four copies against a cap of three, counted on the box**.
Fixed in v0.221.1; the cap now holds at 3 on both apps, verified live.
**Test count 1485 → 1493.** Green gate clean; controller gates all OK; **no push used `--no-verify`.**
## Part 4 — the double-failure path on what ships
Trigger: **the undo copy is lost between being written and being needed** — realistic, and one of only
two ways a rollback can fail. Verified on 0.221.1: held, refused by the customer button, refused after
a full controller restart, then listed/cleared/started via the operator route. Message captured
verbatim (369 bytes, hex in the drill record) with **no engine output**.
## Teardown
Nothing provisioned; 20 containers healthy; `docmost` still holds its 4 pages and 1 user; both
canonical dumps present; **no app left held**; no hub-side record created.
## Operator follow-up ## Operator follow-up
1. Bake and **vouch** a golden carrying controller 0.219.0. The `golden-currency` gate is red until Vouch a golden carrying **0.221.1** (baked in this session), then raise the floor to 0.221.1 last, in
then, correctly: a machine installed now still receives 0.218.0. its own save.
2. **Then** raise the update floor to 0.219.0 — last, in a separate save.
+8
View File
@@ -571,6 +571,14 @@ Each app can define rich metadata in `.felhom.yml`:
(they are the undo). The live recovery unit is still never overwritten, which is why the replay (they are the undo). The live recovery unit is still never overwritten, which is why the replay
source is the scratch. Honesty surfaces (`OffsiteScratchPair`): dump age, an unstamped-pair source is the scratch. Honesty surfaces (`OffsiteScratchPair`): dump age, an unstamped-pair
warning, and the R-44 empty-dump sniff — all warn-level, none of them gates. warning, and the R-44 empty-dump sniff — all warn-level, none of them gates.
- **Where the pre-restore undo copy is written (v0.221.0–.1, R-361).** `DumpOneTo` takes an
EXPLICIT final path; `DumpOne` keeps its signature and calls it with the canonical
`<stack>-<dbtype>.sql`. The safety dump asks for `pre-restore-<stamp>-…` **directly** — it used to
dump to the canonical name and rename afterwards, which destroyed the app's own backup on every
restore. The `.tmp` derives from the final path, so a nightly dump and a safety dump in the same
directory cannot share a scratch file. `db_dumps` in the manifest lists the app's own dumps only;
the undo copies stay on disk and stay visible, capped at 3 per app, pruned from the capture side
**above** the already-current early return.
- **What happens when the database replay FAILS (v0.220.0–.2, R-379/R-380).** A ladder, and every - **What happens when the database replay FAILS (v0.220.0–.2, R-379/R-380).** A ladder, and every
rung is observable: **replay → rollback → hold.** The pre-restore undo copy has always been rung is observable: **replay → rollback → hold.** The pre-restore undo copy has always been
taken; since v0.220.0 it is also **put back** when the replay fails — the whole set for this run, taken; since v0.220.0 it is also **put back** when the replay fails — the whole set for this run,