Files
felhom-controller/REPORT.md
T
admin f7881787f4
gates / gates (push) Successful in 11s
R-361 docs: CONTEXT decision, README, REPORT
Records the db_dumps decision with every consumer named, the trap that a stable
db_dumps lets CaptureRecoveryUnit's already-current early return fire (so
per-capture housekeeping must sit above it), and the NEGATIVE that a held app
does not raise the dead-app alarm - measured, not reasoned, so nobody re-derives
it.
2026-08-23 00:26:32 +02:00

4.8 KiB

REPORT — R-361: taking the undo copy destroyed the app's own database backup

Controller v0.221.0 → v0.221.1, proven live on demo-hp. Full record: felhom.eu/documentation/audits/DRILL-r361-2026-08-22/.

Baselines and subjects

repo HEAD at start version
felhom-controller 2024ed99826717bf765803e159c074909b28bf83 v0.220.2 → v0.221.1
felhom-agent 40d857b52711dd8c9c88bdd21ffcaad819a33a84 v0.130.0, unchanged
felhom.eu a8caa0fdde7c678bb88b5b25727044695bbdeaa9 unchanged
app-catalog 459766cb16395fd1d1a66282f5cc6da59ead5924 unchanged

Hub read: golden 0.220.2, agent 0.130.0, min_agent 0.129.0, floor 0.220.2 — as stated. Subjects FOUND, not rebuilt: docmost and bookstack, healthy, with two sessions of data.

Architecture read: 07-backup-architecture.md §6.1–§6.3 including the v0.220.x failure-ladder [DESIGN]; and audits/DRILL-r379-rollback-2026-08-22/README.md.

R-361 — the whole thing is one comparison

app canonical dump sha256, before after a restore
docmost 5d35678349bbbdb318ac22656d4f96436b5751ef5b3b1659d48305a526df1aed identical
bookstack 7837aa5de2955dcf3125534f015f43df42debe13c765793d3b75173abcc55e7b identical

Before the fix, on the same box: neither app had a canonical dump at all.

DumpOneTo takes an explicit final path and derives its .tmp from it. DumpOne's signature did not move. The comment that asserted the old behaviour was safe now states the invariant and how it is enforced.

Manifest — every consumer named, as required

Manifest.DBDumps has three references, all in internal/backup/recovery_unit.go: the declaration (:55), the enumeration, and the already-current compare. None reads it for recovery; no hub or agent consumer exists. Undo copies are therefore excluded from db_dumps. Verified live: db_dumps: ['docmost-postgres.sql'] with 4 undo copies on disk as the positive control.

Part 3 — the measurement that cancelled Part 2

The reading did not reproduce. A held app aggregates to unhealthy; IsDownState is {stopped, exited, degraded}; the dead-app heartbeat read 0 currently down across scans 600 and 620 while a hold was in force. Part 2 dropped in full, no register row opened, negative recorded in the capability map.

The positive control took three attempts and that is the finding. Two live attempts failed — privatebin went stopped (whitelisted by design) and bookstack went degraded then unhealthy. The control that works is at the detector's own layer: classifyRunStates is pure and raises the banner for degraded/exited while staying silent for the states measured live.

Findings filed

  • R-383 (MEDIUM) — the double-failure message says the previous state's backup exists while naming the very file whose absence caused the failure. Observed twice, on v0.220.2 and v0.221.1. R-361's own class, one surface over.
  • R-384 (MEDIUM) — an app whose database has died reads unhealthy and raises no dead-app alarm. bookstack-db stopped at 21:27:01; 0 currently down throughout.

Register: OPEN-ITEMS.md 325 236 → 327 266 bytes; CLOSED-ITEMS.md 66 777 → 68 464. R-361 compressed; full text git show a8caa0fdde7c:documentation/backlog/OPEN-ITEMS.md.

Two defects in this session's own work

  1. Two red-proofs passed first time, both reported. The behavioural tests inject the dump seam, so a mutation inside DumpOneTo was invisible; and Part 1.3 had no test at all. Guards added at the layer each defect lives in; both mutations then convicted.
  2. One change made another unreachable. A stable db_dumps let the already-current early return fire, and the undo prune sat after it — four copies against a cap of three, counted on the box. Fixed in v0.221.1; the cap now holds at 3 on both apps, verified live.

Test count 1485 → 1493. Green gate clean; controller gates all OK; no push used --no-verify.

Part 4 — the double-failure path on what ships

Trigger: the undo copy is lost between being written and being needed — realistic, and one of only two ways a rollback can fail. Verified on 0.221.1: held, refused by the customer button, refused after a full controller restart, then listed/cleared/started via the operator route. Message captured verbatim (369 bytes, hex in the drill record) with no engine output.

Teardown

Nothing provisioned; 20 containers healthy; docmost still holds its 4 pages and 1 user; both canonical dumps present; no app left held; no hub-side record created.

Operator follow-up

Vouch a golden carrying 0.221.1 (baked in this session), then raise the floor to 0.221.1 last, in its own save.