Records the db_dumps decision with every consumer named, the trap that a stable db_dumps lets CaptureRecoveryUnit's already-current early return fire (so per-capture housekeeping must sit above it), and the NEGATIVE that a held app does not raise the dead-app alarm - measured, not reasoned, so nobody re-derives it.
4.8 KiB
REPORT — R-361: taking the undo copy destroyed the app's own database backup
Controller v0.221.0 → v0.221.1, proven live on demo-hp.
Full record: felhom.eu/documentation/audits/DRILL-r361-2026-08-22/.
Baselines and subjects
| repo | HEAD at start | version |
|---|---|---|
| felhom-controller | 2024ed99826717bf765803e159c074909b28bf83 |
v0.220.2 → v0.221.1 |
| felhom-agent | 40d857b52711dd8c9c88bdd21ffcaad819a33a84 |
v0.130.0, unchanged |
| felhom.eu | a8caa0fdde7c678bb88b5b25727044695bbdeaa9 |
unchanged |
| app-catalog | 459766cb16395fd1d1a66282f5cc6da59ead5924 |
unchanged |
Hub read: golden 0.220.2, agent 0.130.0, min_agent 0.129.0, floor 0.220.2 — as stated.
Subjects FOUND, not rebuilt: docmost and bookstack, healthy, with two sessions of data.
Architecture read: 07-backup-architecture.md §6.1–§6.3 including the v0.220.x failure-ladder
[DESIGN]; and audits/DRILL-r379-rollback-2026-08-22/README.md.
R-361 — the whole thing is one comparison
| app | canonical dump sha256, before | after a restore |
|---|---|---|
docmost |
5d35678349bbbdb318ac22656d4f96436b5751ef5b3b1659d48305a526df1aed |
identical |
bookstack |
7837aa5de2955dcf3125534f015f43df42debe13c765793d3b75173abcc55e7b |
identical |
Before the fix, on the same box: neither app had a canonical dump at all.
DumpOneTo takes an explicit final path and derives its .tmp from it. DumpOne's signature did not
move. The comment that asserted the old behaviour was safe now states the invariant and how it is
enforced.
Manifest — every consumer named, as required
Manifest.DBDumps has three references, all in internal/backup/recovery_unit.go: the
declaration (:55), the enumeration, and the already-current compare. None reads it for recovery;
no hub or agent consumer exists. Undo copies are therefore excluded from db_dumps. Verified live:
db_dumps: ['docmost-postgres.sql'] with 4 undo copies on disk as the positive control.
Part 3 — the measurement that cancelled Part 2
The reading did not reproduce. A held app aggregates to unhealthy; IsDownState is
{stopped, exited, degraded}; the dead-app heartbeat read 0 currently down across scans 600 and
620 while a hold was in force. Part 2 dropped in full, no register row opened, negative recorded
in the capability map.
The positive control took three attempts and that is the finding. Two live attempts failed —
privatebin went stopped (whitelisted by design) and bookstack went degraded then unhealthy.
The control that works is at the detector's own layer: classifyRunStates is pure and raises the
banner for degraded/exited while staying silent for the states measured live.
Findings filed
- R-383 (MEDIUM) — the double-failure message says the previous state's backup exists while naming the very file whose absence caused the failure. Observed twice, on v0.220.2 and v0.221.1. R-361's own class, one surface over.
- R-384 (MEDIUM) — an app whose database has died reads
unhealthyand raises no dead-app alarm.bookstack-dbstopped at 21:27:01;0 currently downthroughout.
Register: OPEN-ITEMS.md 325 236 → 327 266 bytes; CLOSED-ITEMS.md 66 777 → 68 464. R-361
compressed; full text git show a8caa0fdde7c:documentation/backlog/OPEN-ITEMS.md.
Two defects in this session's own work
- Two red-proofs passed first time, both reported. The behavioural tests inject the dump seam, so
a mutation inside
DumpOneTowas invisible; and Part 1.3 had no test at all. Guards added at the layer each defect lives in; both mutations then convicted. - One change made another unreachable. A stable
db_dumpslet the already-current early return fire, and the undo prune sat after it — four copies against a cap of three, counted on the box. Fixed in v0.221.1; the cap now holds at 3 on both apps, verified live.
Test count 1485 → 1493. Green gate clean; controller gates all OK; no push used --no-verify.
Part 4 — the double-failure path on what ships
Trigger: the undo copy is lost between being written and being needed — realistic, and one of only two ways a rollback can fail. Verified on 0.221.1: held, refused by the customer button, refused after a full controller restart, then listed/cleared/started via the operator route. Message captured verbatim (369 bytes, hex in the drill record) with no engine output.
Teardown
Nothing provisioned; 20 containers healthy; docmost still holds its 4 pages and 1 user; both
canonical dumps present; no app left held; no hub-side record created.
Operator follow-up
Vouch a golden carrying 0.221.1 (baked in this session), then raise the floor to 0.221.1 last, in its own save.