R-404 CLOSED with the ruling; R-417 CLOSED by cause removal; R-418/419/420 filed
gates / gates (push) Successful in 17s

THE RULING WAS NEITHER OPTION AS FRAMED. Both offered answers - narrow the gate, or leave it and
write waivers - argued about the gate, and the gate was never the problem.

DIAGNOSIS, from live source: golden_currency_gate.py never looks at the push. It compares the
controller's newest CHANGELOG heading against this repo's bake evidence and returns the same
verdict whatever you are pushing - correct for a standing invariant, wrong as a push gate. And
controller_gates.py had NO golden-currency entry at all. So the repo where a release happens never
checked, and the repo that cannot create the debt was refused on every push. 18 of the last 24
pushes here touched no code - measured, and the new classifier agrees EXACTLY - most of them by
construction, because the controller's code is in one repo and its register lives in this one. SIX
of those 18 were bake records, so the push that PAYS the debt is itself documents-only: the gate
was blocking its own cure.

Not the waiver its docstring prescribes: that clause was written for a release nobody wants a
golden for. R-417 was a release we DID want a golden for, on a night the runbook forbade baking. A
waiver would have recorded a lie.

RULING: block the push that can create the debt, notify the push that cannot.

The gate's logic, exit codes and wording are BYTE-IDENTICAL. Only the consequence changed, for one
gate, on one kind of push, with a loud ADVISORY block so nothing goes quiet.

R-242's vouch half is amended in place to say it is UNTOUCHED and still open - a baked-but-unvouched
golden still passes both the gate and the new notice. Do not read R-404's closure as closing it.

FILED: R-418 - this runner's docstring listed ELEVEN gates while THIRTEEN were registered;
one-register and closed-register ran undocumented since 2026-08-24. Enumeration fixed here, the
correspondence is still unenforced. R-419 - observations_gate.py accepts an item whose body merely
CONTAINS "NOT-A-FINDING", even in prose disclaiming it; found by accident when a planted test
observation passed and my live validation proved nothing. R-420 - controller_gates.py could not
express a non-blocking gate at all before today.

Register: OPEN 171 -> 172, CLOSED 158 -> 160.
This commit is contained in:
2026-09-01 12:01:10 +02:00
parent 1f74427fd2
commit 1e6c387a0b
6 changed files with 116 additions and 18 deletions
+2
View File
@@ -26,6 +26,8 @@
---
| **R-404** | **DECISION — should a documents-only push be subject to the golden-currency gate? RULED 2026-09-01: NEITHER option as framed. Block the push that can create the debt; notify the push that cannot.** The two options on the table were *narrow the gate* and *leave it and build a waiver*, and both were wrong for the same reason: they argued about the GATE, and the gate was never the problem. **The DIAGNOSIS, measured from live source, is that the check was aimed at the wrong repository.** `golden_currency_gate.py` never looks at the push at all — it compares the controller's newest CHANGELOG heading against this repo's bake evidence and returns the same verdict whatever you are pushing, which is correct for a standing invariant and wrong as a push gate. Meanwhile `controller_gates.py` had NO golden-currency entry, so **the repo where a release happens never checked, and the repo that cannot create the debt enforced it on every push.** 18 of the last 24 pushes here touched no code — MEASURED, and the classifier agrees exactly — most of them for a structural reason: the controller's code is in one repo and its register, architecture and status live in this one, so **every controller change produces a documents-only push here by construction.** Six of those 18 were bake records — **the push that PAYS the debt is itself documents-only, so the gate was blocking its own cure.** **WHY NOT THE WAIVER** the gate's own docstring prescribes: that clause was written for *a release nobody wants a golden for*. The case that actually occurred (R-417) was *a release we did want a golden for, on a night the runbook forbade baking*. A waiver would have recorded a lie. **SHIPPED:** `scripts/push_scope.py` (allow-list; every uncertainty answers `code`), a fifth `exemptible` field in `repo_gates.py` + `--scope`, a new **ADVISORY** verdict printed in its own block, the pre-push hook reading git's stdin, the same rule in CI from the push event payload, and `felhom-controller/controller/scripts/golden_notice.py` — a NON-BLOCKING notice at the moment a release is committed. **The gate's own logic, exit codes and wording are byte-identical**; only the consequence changed. The exemption is ONE gate wide and `TestR3`/Scenario C pins it, red-proved by widening it. Proven live on the real hook: docs+debt → ADVISORY, pushed; code+debt → refused; docs+debt+a second gate → refused for that gate alone | **CLOSED 2026-09-01** — ruled and shipped |
| **R-417** | **A drill night that forbids baking a golden made `golden_currency_gate.py` red, so pushing the drill's own evidence needed `--no-verify` — the very signal CI e-mails about.** Measured 2026-09-01: five consecutive felhom.eu CI runs red (jobs 469/470/471/473/476), all mine, all on step 3 `Run the gate entry point`; job 478 green the moment the golden-0.232.0 evidence was committed. Cause confirmed by isolation — moving that directory aside reproduces exit=1, restoring it gives exit=0. **The gate was right every time**: 0.231.0 and 0.232.0 were released with no golden carrying them. **CAUSE REMOVED, not worked around** (R-404): a drill's pushes are documents-only, so the conviction now prints as a loud ADVISORY and the push proceeds — in the hook AND in CI, so a drill night no longer produces red runs indistinguishable from real ones. The expectation is now written where the next drill author reads it (`documentation/runbooks/target-selection.md`), which is the half I had left out. | **CLOSED 2026-09-01** — by R-404 |
| **R-361** | **The pre-restore safety dump overwrote the app's own DB dump, and the comment beside it said it could not.** Shipped in controller v0.221.0 (+v0.221.1). Evidence: `audits/DRILL-r361-2026-08-22/evidence/`. **Reasoning kept:** *`DumpOne` writes `<stack>-<dbtype>.sql` — the app's canonical dump, the name the replay loop matches EXACTLY — so nothing else may ever be written to it.* The fix is a DESTINATION, not a rename: `DumpOneTo` takes the final path and derives its own `.tmp` from it, so neither the destination nor the scratch file can collide with a nightly dump running beside it. **`DumpOne`'s signature did not move** — it has callers outside this concern. **The manifest no longer lists the undo copies:** every consumer of `Manifest.DBDumps` was grepped and named — three, all inside `recovery_unit.go`, none reading it for recovery. **AND THAT CHANGE MADE ANOTHER UNREACHABLE:** a stable `db_dumps` let `CaptureRecoveryUnit`'s already-current early return fire, and the undo-copy prune sat after it — four copies on disk against a cap of three, counted live. The prune now runs ABOVE the check; it is housekeeping on the dump directory and is independent of whether the manifest needs rewriting. **PROVEN LIVE the only way it can be:** the canonical dump's sha256, unchanged across a restore — `docmost` `5d35678349bb…`, `bookstack` `7837aa5de295…`, both byte-identical before and after. A test asserting merely that the undo copy exists passes just as well when the app's backup was destroyed. | **CLOSED — SHIPPED + PROVEN-LIVE** (controller v0.221.1, 2026-08-23) | full text: `git show a8caa0fdde7c:documentation/backlog/OPEN-ITEMS.md` |
| **R-379** | **The pre-restore undo copy was valid, was named to the customer, and no product action could apply it.** Shipped in controller v0.220.0 (+v0.220.1, v0.220.2). Evidence: `audits/DRILL-r379-rollback-2026-08-22/evidence/`. **Reasoning kept:** *R-379 and R-380 were ONE failure with ONE fix — both ended with a half-restored database and the only difference was whether it looked broken.* **The undo set is matched on THE RUN'S OWN STAMP, never on the `pre-restore-` prefix** (four copies coexisted on one app in one afternoon; a prefix match replays an arbitrary older state) **and never just the first file** (a two-database app would have had one restored and the other left half-written). **The rollback RE-DISCOVERS the container** — the undo file is stable, the container is not: the DB-only start re-creates it, and v0.220.0's own first live run held an app for 30 s of `waitDBReady` against a dead id while its data was recoverable. **No unit test saw that: they all inject the import seam and never look at container identity.** | **CLOSED — SHIPPED + PROVEN-LIVE** (controller v0.220.1, 2026-08-22; docmost and bookstack both rolled back to byte-identical prior state) | full text: `git show 4e488321bfd1:documentation/backlog/OPEN-ITEMS.md` |
| **R-380** | **A failed MariaDB replay left a partially-applied database behind an app reporting `health=healthy`.** Shipped in controller v0.220.0. Evidence: `audits/DRILL-r379-rollback-2026-08-22/evidence/13-step2-verify.txt`. **Reasoning kept:** **no engine flag closes this** — `--single-transaction` was added to the Postgres import and does make it all-or-nothing, but **MariaDB's DDL is not transactional**, so a partial apply there is unavoidable at the engine. The flag is a belt; the rollback is the fix, and this row must not be read as saying otherwise. Proven live: `bookstack`'s `migrations` table back at **102 rows**, the exact cell the defect was measured in. | **CLOSED — SHIPPED + PROVEN-LIVE** (controller v0.220.0, 2026-08-22) | full text: `git show 4e488321bfd1:documentation/backlog/OPEN-ITEMS.md` |
File diff suppressed because one or more lines are too long