R-87 put back in the register; closed_register_gate.py is the 12th gate (R-405, R-406)
gates / gates (push) Failing after 17s

Records and process only. No machine contacted. No product code, no version bump,
no build, no deploy.

R-87 was moved into CLOSED-ITEMS.md by the 2026-08-22 compression sweep ef6ac6f while
its own state cell read "READY - RE-RANKED UP 2026-08-03 (R-86 closed)". R-378 records
that sweep moving six still-open rows and restoring them in the same session; this was
a seventh it missed. Nine days in the wrong file, with the register's ranking paragraph
ranking it fourth and pointing at nothing. Restored verbatim from ef6ac6f^, beside R-95
where it sat before.

The predicate is the LEADING VERDICT of the state cell, which is R-378's lesson and
decides the answer here. Measured on the file as pushed: an open word anywhere in the
state cell convicts 3 of 151 rows, two of them genuinely closed (R-224 and R-260 carry
"open"/"OPEN" inside long prose verdicts); the leading verdict convicts exactly 1; the
whole row convicts 144.

closed_register_gate.py, two rules: no open state word leading a CLOSED-ITEMS.md row's
verdict, and no R- id with a row in both registers. Red-proofed both, and negative-
controlled against the pushed pre-fix files where it convicts R-87 by name, rc=1;
restoring the planted rows leaves the file byte-identical. Registered as the 12th gate
in repo_gates.py, --fast, after it was green. Four residual holes in its docstring.

R-398 was also in both registers - a deliberate cross-reference stub. Now prose beneath
the table rather than a table row, because a row in both files is what rule 2 convicts on.

R-406 filed: two unrelated findings in OPEN-ITEMS.md both numbered R-133. The only such
collision in either register. Deliberately NOT gated - a within-register duplicate rule
would fail on a pre-existing row, and a registered-but-failing gate refuses every push.

golden-currency is RED at this commit and was already red at dddcc80 - controller
v0.230.0 released, newest golden 0.229.0. Pre-existing, not this session's debt.

Ceiling R-404 -> R-406.
This commit is contained in:
2026-08-31 15:32:40 +02:00
parent dddcc808be
commit 6e550aedd3
5 changed files with 177 additions and 2 deletions
+2 -2
View File
@@ -69,7 +69,6 @@
| **R-114** | **On target-drive loss the customer is told the wrong story and offered the drive that just vanished.** Shipped in 0.113.0. Evidence: `audits/E2D-fresh-vm-2026-07-29.md`, `audits/SESSION-C-2026-07-29.md`. | **SHIPPED + PROVEN-LIVE** (controller v0.186.0, 2026-07-29) | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` |
| **R-29** | **The green gates are not enforced anywhere — one was RED for 16 releases before anyone ran it.** Shipped in 1.19.0, 1.22.0, v0.129.0. | **CLOSED — both halves shipped** (2026-08-02) | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` |
| **R-86** | **Restore-tests are interval-scheduled, not backup-aligned** | **CLOSED — SHIPPED + PROVEN-LIVE 2026-08-03** (agent **v0.121.0**, hub **v0.91.0**) | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` |
| **R-87** | **The restic tier is never restore-tested** **Reasoning kept:** **And R-95 still applies:** that credential can delete, so a restic restore-test must never be able to write to the repo CC | **READY — RE-RANKED UP 2026-08-03 (R-86 closed)** | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` |
| **R-185** | **The agent cannot see the host backup tier's archives on demo-felhom — the PVE token has no ACL on `/storage/felhom-backup`, so the content listing returns EMPTY where root sees three archives.** Shipped in v0.123.0. | **CLOSED — SHIPPED + PROVEN-LIVE 2026-08-03** (agent **v0.123.0**, installer **1.24.0**) | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` |
| **R-186** | **A released agent binary's sha256 cannot be reproduced from its tag.** Shipped in v0.120.1, v0.121.0, v0.121.2. | **CLOSED — SHIPPED + MEASURED 2026-08-03** (agent **v0.122.0**) | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` |
| **R-187** | **R-115's one-command release had never actually run its publish leg — the first real use died there.** Shipped in v0.121.0. | **CLOSED — SHIPPED 2026-08-03** (`felhom-agent`) | full text: `git show fddfe00ce268:documentation/backlog/OPEN-ITEMS.md` |
@@ -176,7 +175,8 @@ Two rows closed, one CORRECTED and deliberately left open.
|---|---|---|---|
| **R-359** | The off-site restic store was never verified by anything, ever | controller v0.227.0/v0.227.1 | `documentation/tests/r359-integrity-2026-08-30/` |
| **R-397** | `NotifyIntegrityOK`/`NotifyIntegrityFailed` had no caller, and the product advertised a weekly check that did not exist | controller v0.227.0 | `documentation/tests/r359-integrity-2026-08-30/` |
| R-398 | **NOT closed — CORRECTED.** The premise was wrong; the seam already existed | — | stays in OPEN-ITEMS as the record |
**R-398 is deliberately NOT a row here.** It was proposed for closure in the same pass and was **CORRECTED instead** — the premise was wrong, the seam already existed — so it stays in `OPEN-ITEMS.md` as the record. It is written as prose rather than a table row because a register row in both files is exactly what `closed_register_gate.py` convicts on.
**The rules these leave behind:**
+3
View File
@@ -225,6 +225,7 @@ unconditional promise, which CAMPAIGN-11 §7 step 7 measured the customer-facing
| **R-121** | **A BOX's installed agent can sit releases behind the vouched one and nothing notices — the R-120 gate does not cover it.** demo-hp ran agent **0.113.0** while the hub vouched **0.116.0**, through the whole R-116/R-117 arc, and no signal existed on any channel | **READY (S) — NEW 2026-07-30** | — | **Fourth instance of the drift family** (R-111 golden's agent 17 releases behind, R-115 built+deployed but never published, R-120 golden a controller behind — and now installed-vs-vouched on a live box). **Confirmed at source that R-120's gate cannot catch it:** `hub/internal/web/configs.go:1165-1169` compares `goldenVer` against `store.NewestReportedControllerVersion()` — it is a **golden-artifact vs fleet-CONTROLLER** check and says nothing about the agent installed on a box. **`MinAgent` does not cover it either:** it is used to HOLD the controller floor for a box whose agent is too old (`hub/internal/api/handler.go:530-538`, `store.go:1857`) — protective, not an alarm — and demo-hp's 0.113.0 **equalled** `min_agent` 0.113.0, so even a floor comparison was satisfied. **The cost, measured:** R-117's whole subject is the R-113 conjunction, which landed in **0.114.0** — so the designated drill host could not exercise the code under investigation at all, and the R-117 spike had to route every predicate result through an out-of-repo probe built from `main` instead of the installed agent (`audits/SPIKE-r117-bind-liveness-2026-07-30.md` §1, §2.3). Discovered because the R-117 task made bringing the box current an explicit prerequisite. **Fix shape (not implemented):** the hub already receives `AgentVersion` on every host report, and already has semver comparison in Go — the missing piece is a checker comparing reported agent vs the **vouched** agent and surfacing it, operator-tier. Note the honest tension: a box legitimately lags between publish and deploy, so this wants a staleness window rather than an instant alarm | CC |
| **R-118** | **An absent drive's union row advertises the ROOT filesystem's capacity as its own.** In the absent-state payload the registry-union row reports `total_bytes: 49675956224 / used_bytes: 4584579072` — **byte-identical to the `local` row** (`durable_id: path:/var/lib/vz`, i.e. `pve-root`) in the same response. The real drive is **4 GB** | **READY (XS) — NEW 2026-07-30** | — | Cause: `statfsCapacity(d.MountPath)` (`disks.go:335-338`) statfs's `/mnt/cel`, which with the device gone is a **bare directory on the root filesystem**. `observe.go:176-183`'s comment warns about exactly this trap and guards the Observe path ("*an unmounted removable dir-storage's mountpoint reverts to a bare directory on root … catastrophic DR mis-id*"); **the union path has no equivalent guard.** **Not a DR mis-id** — `durable_id` on that row is still the correct `uuid:…`, so re-attach identity is safe. It is a **false capacity** reaching every consumer of `total_bytes`/`used_fraction` (fill monitors, storage cards): a detached 4 GB drive advertises 46 GiB at 9.2 % used. Same class as `role.go:180-181` — an absent drive's fields decaying to the root filesystem's. Evidence: `audits/DIAG-r116-disks-payload-2026-07-30.md` §12 | CC |
| **R-95** | restic offsite credential **can delete** (`readonly=False`, `forget --prune` runs from the box); SFTP cannot express append-only | **READY** | — | Root exposure still open. Mitigation now ARMED — split prune off-box or move to REST `--append-only` | CC |
| **R-87** | The restic tier is never restore-tested | **READY — RE-RANKED UP 2026-08-03 (R-86 closed)** | — | Design a controller-side test (no scratch-guest analogue transfers). **Most of what this row needed now exists.** R-86 built the piece that was missing: a tier is proved **per archive generation**, on its own rhythm, with the proof recorded as *which archive* — which is exactly the shape a weekly-ish restic tier needs, and the reason this row could not simply reuse the whole-guest scheduler before. What remains is genuinely restic-specific and is NOT a scheduling problem: there is no scratch-guest analogue, so the test has to be a controller-side restore of a bounded sample into a throwaway path, with its own definition of "proved". **Two things to carry over rather than re-derive:** the proof must record the SNAPSHOT it proved (not a timestamp), and the hub's staleness window must learn this tier's rhythm the way `restoreProvenWindow` now does — a restic tier on a weekly cadence lands on the same false-alarm line the flat 7 days did. **And R-95 still applies:** that credential can delete, so a restic restore-test must never be able to write to the repo | CC |
| **R-191** | **Every weekly offsite backup UPLOADS successfully and then FAILS the job on a prune the box is deliberately not allowed to do — on both demo boxes.** Measured on demo-felhom 2026-08-04 06:49–06:53: the upload completed (223 s, 629 MiB of 1.874 GiB, **67.2 % reused incrementally**), then `ERROR: prune 'ct/9201': proxmox-backup-client failed: Error: permission check failed - missing Datastore.Modify\|Datastore.Prune on /datastore/felhom-offsite/demo-felhom` → `ERROR: Backup of VM 9201 failed - error pruning backups` → `TASK ERROR: job errors`. The hub raised `whole_guest_backup_failed` | **CLOSED — SHIPPED 2026-08-04** (installer **1.25.0**; both live boxes corrected) | — | **This is R-89's rule not reaching the config.** R-89 moved PBS pruning SERVER-SIDE — *"boxes set `keep_last: 0`, ep0 runs prune jobs; box tokens stay write-only, never widen the grant"*. The token behaves exactly as designed: it refuses. But **both** demo boxes still arm the offsite tier with `keep_last=2 prune_pbs_allowed=true` (`backup_targets: [{target_id: felhom-pbs, cadence_seconds: 604800, keep_last: 2}]`), so every run asks for a prune that must fail. **The data is SAFE and that is why this is not a P1:** the snapshot lands before the prune is attempted; what is wrong is the job's VERDICT and the weekly operator e-mail it produces. **But it is corrosive in the specific way this project keeps finding:** a backup that reports FAILED while succeeding trains the operator to discount `whole_guest_backup_failed`, which is the same alert that would carry a real one — and it is exactly the failure the R-100 corollary warns about, an alarm whose text is true and whose trigger is not the thing you would act on. **Fix is one config line per box** (`keep_last: 0` on the PBS tier) plus whatever writes it on a fresh install; **deliberately NOT applied in this session** — the session was a runbook with an explicit "change nothing, and if a change appears necessary, stop and report" rule, and a retention field on a live backup tier is not a change to slip into an observation run. **Check before fixing:** whether ep0's prune jobs actually cover these two namespaces, or the snapshots simply accumulate once the box stops asking **THE GATE WAS RUN FIRST, AND IT MATTERED.** Before disabling anything, ep0 was read (read-only, Tier 2): prune jobs `prune-demo-felhom` and `prune-demo-hp` exist on datastore `felhom-offsite`, one per namespace, `schedule 03:30`, `keep-last 2`, comment *"R-82 retention keep-last=2, server-side (box tokens are write-only)"* — and they have run **every day since 2026-07-27: 18 tasks, all `status=OK`**. The newest task log reads `retention options: --ns demo-felhom --max-depth 0 --keep-last 2` / `keep ct/9201/2026-07-27…` / `keep ct/9201/2026-07-28…` / `TASK OK`. Retention happens, and it happens there. **A METHODOLOGICAL WARNING WORTH MORE THAN THE FIX.** Three separate queries said the OPPOSITE — *no prune jobs have ever run* — and **all three were broken instruments**: `worker-type` where the field is `worker_type`; the value `prune` where the worker type is `prunejob`; and `journalctl -u proxmox-backup` where the unit is `proxmox-backup-proxy`. A fourth reading (3 snapshots under keep-last 2) was mis-framed by CC and self-corrected — the third snapshot had landed AFTER that day's 03:30 window. Acting on any of them would have disabled the only pruning ATTEMPT while reporting that nothing prunes: a weekly false alarm traded for unbounded growth on the protected endpoint, invisible for months. **The gate is what caught it, and only because it demanded evidence rather than a verdict.** **Shipped:** installer **1.25.0** writes `keep_last: 0` on the offsite tier (the agent's existing guard `allowPBSPrune = !primary && keep_last > 0` already reads that as *never prune from the box* — no agent change), the justifying paragraph is rewritten to say where retention lives and cite R-89, and `hostinstall_gates.py` asserts it (red-proved: pinning `keep_last: 2` back fails the gate). **Both live boxes corrected in their own config** — `backup tier armed target=felhom-pbs … keep_last=0 … prune_pbs_allowed=false` on demo-felhom and demo-hp, with the local tier untouched at `keep_last=3`. Served over HTTPS at `1.25.0` with `"keep_last":0` in the served bytes. **STILL TO OBSERVE:** the next weekly offsite run completing OK end-to-end. The change removes the failing step; the *schedule* proving it is next week's event, and this row should carry that line when it happens. | CC |
| **R-194** | **PVE's permission cache delays every grant-state verdict by an unknown amount, so "the agent can read it" and "the ACL exists" are not the same measurement.** Observed twice while validating R-190's self-repair on demo-felhom 2026-08-04: both ACL rows for `/storage/felhom-backup` were deleted, and `GET /access/permissions` continued to report `Datastore.AllocateSpace` present — for **~40 s** in one run and **~16 minutes** in another. During that window the capability probe reads healthy and the self-repair does not fire | **OPEN** | — | **Why it matters beyond the delay:** it puts a floor under how fast a lost grant can be noticed, it makes any single permission read a lagging indicator, and — the interesting part — **it is a candidate contributor to R-190's own timeline**: a grant removed at an unknown moment could keep working until a cache expiry, which is exactly the shape of *worked at 04:44, refused at 09:24*. That does not explain what removed it, but it may explain when the refusal SURFACED, and the two have been treated as the same instant. **Not a defect in our code** — it is PVE behaviour, and the mitigation already tolerates it (the repair fires on the next probe after the cache clears). **What is worth deciding:** whether the store-grant probe should ALSO consult the storage content listing as a second signal, since that appeared to reflect the loss immediately (`{"data":[]}` while the permission read still said present) — two signals disagreeing is itself information, and today only one of them is read | CC |
| **R-200** | **The DR password-injection seam has a handler, a route and tests — and no form.** `POST /backup/offbox/inject-password` is routed (`controller/internal/web/server.go:510`) to `offboxInjectPasswordHandler` (`offbox_handlers.go:174-196`) → `InjectOffboxPassword` (`backup/offbox.go:541`). **No template in the repository contains that path or any form posting to it** (grep over `internal/web/templates/`: one unrelated hit, an XSS comment) | **PLUMBING COMPLETE** (controller **v0.196.0**); **the FORM is not built — still open** | — | **The tenth instance of this project's built-but-never-wired class, and the exact shape `CLAUDE.md` and `felhom.eu/CLAUDE.md`'s seam-wiring rule were written for:** handler tests that POST directly (`offbox_escrow_test.go:167,180`) prove nothing about reachability. To use the only implemented recovery seam today, a person must hand-craft an authenticated POST with a session cookie and CSRF token. **Note the layering while fixing it:** this form takes a **64-hex repo password**, not a recovery code (`offboxRepoPwPattern`, `offbox.go:543`) — they are different secrets at different layers, and the operator's 2026-08-04 ruling asks for a form that takes **R**. Build the R form and treat this one as the operator/DR fallback it was written as, but **ship it with a render test per branch of whatever gate it sits behind**. Source: `audits/RECON-offsite-dr-chain-2026-08-04.md` §3 link 9 **THE DIAGNOSTIC HALF IS DONE AND IT ANSWERED THE QUESTION.** `--recover-offsite-check` is a `docker exec` escape hatch in the shape of `--print-reset-code`: R on **STDIN** (never argv, never `ps`, never shell history, never a transcript), fetch+unseal via the agent, and a verdict of **two sha256 hashes**. **It compares and never installs** — the recovered password is not written to `offbox/repo_password`; a test asserts the data dir is byte-unchanged and its red-proof (adding the install call) fails it. Confirmed live: `repo_password` mtime still `2026-08-03 07:18:02` after the successful check at `2026-08-04 11:49`. **Exit codes are load-bearing** — `0` match, `2` a clean MISMATCH, `1` a step failed; "it failed" and "it worked and disagreed" must never share a status because only one is a finding about the system. A box with no local password reports distinctly (the rebuilt-box shape, where the next step is to INSTALL rather than compare). **WHAT IS NOT BUILT, deliberately:** no card, no form, no preview, no customer-facing text — building an interface on top of a chain nobody had walked is how the preceding three weeks went wrong. **What remains for this row:** link 9 (the recovered password placed so `WriteOffboxSecrets` keeps it) and the customer-facing shape the operator ruled on 2026-08-04 (yell → R form → preview → proceed), which is now priced against a chain that exists rather than one that is assumed **PART 0 SHIPPED 2026-08-04 (v0.196.0):** `--recover-offsite-install` is the sibling of the check — same fetch/unseal path, same STDIN discipline for R — and it **places** the recovered password via `InjectOffboxPassword`. The confirmation is a SECOND invocation (`--confirm-install`): without it, both hashes print and nothing is written, so the operator sees the comparison before a write is possible. Three outcomes named distinctly: **installed** (no local password — the rebuilt-box shape), **unchanged** (identical key, nothing written), **refused** (a DIFFERENT key present — installing would clobber the key the current repository is encrypted under; exit 2, no force offered). It re-reads the file after writing rather than trusting the call. **Red-proof observed:** removing the confirmation gate makes the dry run write the password. The R-persistence test carries a **positive control** (a planted copy found, then removed and not found). **NOT YET EXERCISED AGAINST A LIVE RECOVERY** — the R-201 drill halted before step 9, so this is unit-proven only. **What remains for this row:** the customer-facing shape the operator ruled on 2026-08-04 (yell → recovery-code form → preview → proceed) | CC |
@@ -581,6 +582,8 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server`
| **R-105** | **Three hub-held DR records are empty on the entire live fleet.** `hosts.dr_record_json` = `{}` on all 3 hosts; `host_escrow.directive_json` = `{}` on both escrowed hosts; `dr_recipe.host_half.drives` = `[]` on every customer **including two with enrolled data drives** (916 GB USB on demo-felhom, 938 GB NVMe on demo-hp) **MIGRATED FROM `ROADMAP.md` 2026-08-22 (R-369) — originally filed 2026-07-28, size M, roadmap state `READY — 2026-07-28`.** Moved verbatim; nothing added or reinterpreted. The roadmap keeps its copy as history, marked moved. ****PARTLY FIXED BY ITS OWN UPDATE.** The `drives` third was traced and populated on both demo boxes on 2026-07-28 (the enrolled drives were never PVE storages, so `isUserDataDrive` never saw them). The other two thirds — `hosts.dr_record_json` and `host_escrow.directive_json` — were NOT re-verified this session and are carried as written.** | **OPEN — migrated from ROADMAP 2026-08-22, rank unchanged** | — | These are exactly the fields a host-loss recovery reads: `05-hub-architecture.md:175-176,186` names the slim DR record as one of four durable sources; `06-offsite-connectivity.md:148-150` says the escrow upload carried the DR directive; `felhom-agent/internal/dr/plan.go:34-35` makes `PlannedDrive` the re-attach-by-`durable_id` wrong-disk guard. **The three may have different causes** — `isUserDataDrive` (`internal/hub/dr_recipe.go:129-136`) requires type `usb`/`local-dir` **and** a non-empty `DurableID` **and** `MountPath`, and which of the three fails was not traced. Evidence: `architecture/_recovery-inventory-2026-07-28.md` Part D2.3. **UPDATE 2026-07-28 (vzdump-target move): the `drives` third is TRACED and now POPULATED on both demo boxes.** Cause: the enrolled data drives were never PVE storages at all — only agent-generated systemd mounts — so they never entered `report.StorageTargets` and `isUserDataDrive` never saw them. Giving each drive a `dir` storage at its own mountpoint supplied all three required fields at once (type `local-dir`, fs-UUID durable id, mount path), and the recipe now emits `uuid:91d2dc2d-…`/`/mnt/nvme-1tb` on demo-hp and `uuid:47a3361a-…`/`/mnt/hdd_1` o | CC |
| **R-340** | **The new reachability check does not touch the surface that actually failed.** R-339 reports when the hub cannot READ ep0 — but the read it performs is the `usage` op, which is `proxmox-backup-manager` plus `df` over SSH, and therefore rides the **local API daemon**. The 2026-08-18 incident explicitly CLEARED that daemon: `proxmox-backup.service` was healthy throughout, and it was the **HTTPS proxy on 8007** that was wedged with a full accept queue. **So R-339's check would have returned green for all 9 h 37 m of that outage.** It closes the case where ep0 is unreachable *as a host*; it does not close the case that actually happened. This is not a defect in R-339 — it is the honest boundary of what it watches, recorded so a future reader does not mistake a green box gauge for a working off-site tier | **READY (M) — NEW 2026-08-18** | a tenantsync endpoint-script version bump (the op is added on ep0, so it needs the same version-gated rollout `ErrUsageUnsupported` already models) | Add a **health op** to `scripts/felhom-tenantsync.sh` that probes `https://127.0.0.1:8007/` on ep0 and reports the proxy's fd count and listen-queue depth, then surface it as a third signal. **Overlaps the connections spike (R-336's remaining half): both want the same observations from ep0, so whichever runs SECOND must reuse the first's evidence rather than re-measuring a protected machine twice** **REUSE, per this row's own instruction — the connections spike ran FIRST (2026-08-20) and already produced most of what the health op wants; do not re-measure a protected machine a third time.** Available in `audits/evidence-ep0-established-connections-2026-08-20/`: the proxy **fd count** and its type breakdown (`lsof` + `/proc/<pid>/fd`), the **listen-queue depth** (`ss -lnt` — `Recv-Q 0`, `Send-Q 1024`), the **ESTAB/CLOSE-WAIT split**, the **per-peer** connection histogram, a **31-minute persistence diff** of full 4-tuples, and a **46.18 h** slope with Poisson bounds. What the health op would still add beyond these: a loopback `GET https://127.0.0.1:8007/` probe — the observation that distinguished "process problem" from "network problem" on 2026-08-18 and the one thing this spike did NOT take, because it is the surface R-339 cannot see. **And this spike sharpens what the op should report:** a rising **ESTAB** count is the live signal (CLOSE-WAIT was **0**, not merely flat), and per **R-344** the fd ceiling that matters may be the **agent's**, not only ep0's. | CC |
| **R-404** | **DECISION FOR VIKTOR — should a documentation-only push be subject to the golden-currency gate?** `git push --no-verify` has now been used **six times**, each with a recorded reason, because `repo_gates.py --fast` runs `golden_currency_gate.py` on every push to `felhom.eu` including pushes that touch only `documentation/`. **A guard that is correctly bypassed six times is training everyone to bypass it**, and the seventh bypass will be faster to reach for than the sixth | **OPEN — a DECISION, not work. Filed 2026-08-31, deliberately NOT acted on** | — | **FOR narrowing it:** a docs-only push cannot be the push that finishes a release, so scoping the gate to pushes that touch `controller/`, `agent/` or `hub/` code is arguably not a weakening at all — it would fire on exactly the pushes that can create the gap and on no others. It would also end the habit, which is the real cost being paid now. **AGAINST narrowing it:** the gate was earned by a real recurrence — v0.206.0 shipped while the vouched golden carried 0.205.0, and v0.204.0/v0.205.0 before it — and every narrowing of a guard risks the thing it was built for coming back. The gate is also deliberately `--fast` so that BOTH the pre-push hook and CI run it (R-29's census failure); a scoped version must stay in both or it runs in neither. **IF VIKTOR DOES NOTHING:** the bypass stays routine and the count keeps rising; nothing breaks, and the guard quietly stops being one. **Owner: Viktor decides, CC builds. This task did NOT change the gate.** | Viktor |
| **R-405** | **R-87 sat in `CLOSED-ITEMS.md` for nine days while it was still open, and the register's own ranking paragraph ranked it fourth pointing at nothing.** Established from history, not inferred: it was moved by the 2026-08-22 compression sweep, commit `ef6ac6f` (*One register, enforced by a gate; closed work compressed into siblings, R-376..R-378*) — the same commit and the same defect class R-378 records. **R-378 caught six — R-123, R-190, R-214, R-264, R-295, R-352 — and missed a seventh.** R-87 escaped because its state cell read `READY — RE-RANKED UP 2026-08-03 (R-86 closed)`: the leading verdict is `READY` and the word `closed` later in the same cell describes a **different** row. **Count reproduced independently 2026-08-31, and the predicate decides the answer:** matching an open word anywhere in the state cell convicts **three** of 151 rows (R-87, plus R-224 and R-260, both genuinely closed with the words "open"/"OPEN" inside long prose verdicts); matching the **leading verdict** convicts exactly **one**, R-87; matching the whole row convicts **144**. **Fixed this session:** the row is back in `OPEN-ITEMS.md` verbatim from `ef6ac6f^`, next to R-95 where it sat before, and `scripts/closed_register_gate.py` is the 12th gate. Red-proofed both rules and negative-controlled against the pushed pre-fix files, where it convicts R-87 by name. **`R-398` was ALSO in both registers** — a deliberate cross-reference stub — and is now prose beneath the table rather than a row, because a row in both files is what rule 2 convicts on. | **CLOSED 2026-08-31 — corrected + gated in the same session** | R-378 | Nothing further. The gate's four residual holes are named in its docstring; hole 4 is R-406. | CC |
| **R-406** | **Two unrelated findings in `OPEN-ITEMS.md` share the identifier R-133.** `OPEN-ITEMS.md:267` is *the hub enforces uniqueness on `customer_id` only* (`domain` is `TEXT NOT NULL DEFAULT ''` with no UNIQUE/CHECK); `OPEN-ITEMS.md:273` is *the vaulted break-glass console credential is PLAINTEXT AT REST*. Different subjects, different owners, one number. Found 2026-08-31 while measuring duplicate ids for R-405's gate — **this is the only such collision in either register** (measured: no id appears twice in `CLOSED-ITEMS.md`, and R-88a/R-88b and R-209/R-209a are distinct suffixed ids, not duplicates). **Why the gate does NOT check for it:** a within-register duplicate rule would fail on this pre-existing row, and a registered-but-failing gate refuses every push. **Renumbering is not obviously safe** — `R-133` is cited elsewhere and a blind renumber breaks whichever citation meant the other one. | **OPEN — LOW** | R-405 | Establish which of the two `R-133` citations exist outside the register, then renumber the one with fewer (or none) and add the within-register duplicate rule to `closed_register_gate.py`. Do NOT renumber before grepping the citations. | CC |
<!-- DUE-CHECKS-BEGIN — machine-readable. Parsed by scripts/due_checks_gate.py.
One row per dated check. The R-number must have a row above. Dates are UTC.
+22
View File
@@ -1,3 +1,25 @@
## closed_register_gate.py v1.0.0 — CLOSED-ITEMS.md holds closed work only (2026-08-31, R-405)
**One row corrected, one gate added, no product code touched.** R-87 ("the restic tier is never
restore-tested") was moved into `CLOSED-ITEMS.md` by the 2026-08-22 compression sweep (`ef6ac6f`)
while its own state cell read `READY — RE-RANKED UP 2026-08-03 (R-86 closed)`. R-378 records that
sweep moving six still-open rows and restoring them in the same session; **this was a seventh it
missed**, and it sat in the wrong file for nine days while `OPEN-ITEMS.md`'s ranking paragraph ranked
it fourth and pointed at nothing. The row is back beside R-95, verbatim from `ef6ac6f^`.
- **The predicate is the LEADING VERDICT of the state cell** — R-378's whole lesson, and it decides
the answer here. Measured on the file as pushed: an open word anywhere in the state cell convicts
**3 of 151** rows, two of them genuinely closed (R-224, R-260 carry "open"/"OPEN" inside long prose
verdicts); the leading verdict convicts exactly **1**; the whole row convicts **144**.
- **Two rules only** — no open state word leading a `CLOSED-ITEMS.md` row's verdict, and no `R-`
identifier with a row in both registers. Its four residual holes are in its docstring.
- **Red-proofed both rules**, and negative-controlled against the pushed pre-fix files, where it
convicts R-87 by name and rc=1. Restoring the planted rows leaves the file byte-identical.
- **`R-398` was also in both registers** — a deliberate cross-reference stub. It is now prose beneath
the table rather than a table row, because a row in both files is what rule 2 convicts on.
- Registered as the **12th** gate in `repo_gates.py`, `--fast` (two file reads), **after** it was
green.
## check_skills.py v1.0.0 + five process-domain skills (2026-08-25)
**Five new skills under `skills/`, one new checker under `scripts/`, no product code touched and
+145
View File
@@ -0,0 +1,145 @@
#!/usr/bin/env python3
# -*- coding: utf-8 -*-
"""closed_register_gate.py — CLOSED-ITEMS.md holds closed work ONLY (R-405, 2026-08-31).
WHAT IT CONVICTS ON (a FAIL, exit 1), two rules and nothing else:
RULE 1 — no row in `CLOSED-ITEMS.md` may carry an OPEN state word (READY, OPEN, BLOCKED,
WATCHING, WAITING-ON-OPERATOR) as the LEADING VERDICT of its state cell.
RULE 2 — no `R-` identifier may have a table row in BOTH `CLOSED-ITEMS.md` and `OPEN-ITEMS.md`.
WHY IT EXISTS, and the cost that bought it. The 2026-08-22 compression sweep (commit `ef6ac6f`,
R-376..R-378) moved finished work out of `OPEN-ITEMS.md`. Its classifier matched a status word
ANYWHERE in the row, so rows that were not closed went with it. R-378 caught six of them in the same
session — R-123, R-190, R-214, R-264, R-295, R-352 — and restored them verbatim. **It missed a
seventh.** R-87 ("the restic tier is never restore-tested") went into the closed file with its own
state cell reading `READY — RE-RANKED UP 2026-08-03 (R-86 closed)`: a LEADING verdict of `READY`, and
the word `closed` later in the same cell describing a DIFFERENT row. It sat in the wrong file for
nine days while `OPEN-ITEMS.md`'s own ranking paragraph ranked it fourth and pointed at nothing.
READ THE LEADING VERDICT, NOT THE CELL, NOT THE ROW. This is R-378's whole lesson and this gate would
be wrong without it. Measured on `CLOSED-ITEMS.md` at 2026-08-31, matching an open word anywhere in
the state cell convicts THREE rows; two of them — R-224 and R-260 — are genuinely closed and merely
contain the words "open" and "OPEN" inside long prose verdicts. Matching the leading verdict convicts
exactly ONE, which is the real one. Matching the whole ROW convicts 144 of 151.
WHAT THIS GATE CANNOT SEE — the residual holes, named rather than implied:
1. **A row whose body contains a `|` shifts its own cells and is read in the wrong place.** Two
rows do this today (R-309, R-351: a pipe inside inline code). The gate cannot tell a shifted
cell from a real one, so a mis-filed row that also contains a pipe escapes. It prints such rows
as a WARNING when the leading verdict is not a verdict word AND an open word sits somewhere in
the cell, which is the best signal available without a real Markdown parser. A leading verdict
that is merely a version string is NOT warned on: this table's `Shipped` column legitimately
holds `controller v0.227.0`, and ten warnings per run is how a gate gets ignored.
2. **A row with no state cell at all escapes.** Two exist today (R-399, R-400 — two columns where
the table declares four). An empty state cell is not an open word, so RULE 1 passes it. They are
printed as a WARNING.
3. **A closed-sounding verdict that is not true escapes.** `PARTLY CLOSED` leads with no open word.
This gate checks where a row FILED, never whether the verdict is honest.
4. **A duplicate id WITHIN one register escapes.** `OPEN-ITEMS.md` carries two unrelated findings
both numbered R-133 (filed as R-406). Adding that rule would fail the gate on a pre-existing
defect, and a registered-but-failing gate refuses every push, so it was deliberately left out.
5. Nothing here reads audits, spikes or inventories. A finding that never reaches either register
is invisible to this gate, as it is to `one_register_gate.py`.
Run: python3 scripts/closed_register_gate.py [--fast]
"""
import io
import os
import re
import sys
ROOT = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
CLOSED = os.path.join(ROOT, "documentation", "backlog", "CLOSED-ITEMS.md")
OPEN = os.path.join(ROOT, "documentation", "backlog", "OPEN-ITEMS.md")
# The id keeps its suffix letter: R-88a and R-88b are different rows, and R-209/R-209a are a closed
# row and its open successor. Dropping the letter invents duplicates that are not there.
ROW = re.compile(r"^\|\s*\*{0,2}(R-\d+[a-z]?)\b")
# A verdict ends at the first separator. `READY — RE-RANKED UP (R-86 closed)` leads with READY;
# `CLOSED 2026-08-06 — ... did not open the sealed bundle` leads with CLOSED.
SEPARATOR = re.compile(u"[—–(:,.]|\\s-\\s")
OPEN_WORD = re.compile(r"\b(READY|OPEN|BLOCKED|WATCHING|WAITING-ON-OPERATOR)\b", re.I)
DONE_WORD = re.compile(r"\b(CLOSED|SHIPPED|FIXED|KILLED|DONE|SUPERSEDED|OBSOLETE|WITHDRAWN|MERGED|"
r"DISCHARGED|RULED|MOVED|BANKED|PROVEN-LIVE|EXECUTED)\b", re.I)
def leading_verdict(cell):
"""The verdict word(s) before the first separator, with Markdown emphasis stripped."""
text = cell.replace("*", "").replace("~", "").replace("`", "").strip()
return SEPARATOR.split(text)[0].strip()
def rows(path):
"""Yield (line_no, id, state_cell_or_None, line) for every R- row in a pipe table.
CLOSED-ITEMS.md declares `| ID | Title | Shipped | Evidence |` — four columns, so a well-formed
row splits into six fields and the state cell is the third. A row of any other width has no
state cell this function is willing to guess at, and says so with None.
"""
for n, line in enumerate(io.open(path, encoding="utf-8"), 1):
line = line.rstrip("\n")
m = ROW.match(line)
if not m:
continue
cells = line.split("|")
state = cells[3] if len(cells) == 6 else None
yield n, m.group(1), state, line
def main():
for path in (CLOSED, OPEN):
if not os.path.exists(path):
print("closed-register gate: %s is missing — FAILURE, never a skip" % path)
return 1
convicted, warnings, checked = [], [], 0
for n, rid, state, line in rows(CLOSED):
if state is None:
warnings.append((n, rid, "row is not four columns — no state cell to read"))
continue
checked += 1
verdict = leading_verdict(state)
if OPEN_WORD.search(verdict):
convicted.append((n, rid, verdict, "open state word in the leading verdict"))
elif not DONE_WORD.search(verdict) and OPEN_WORD.search(state):
# Ambiguous and worth a human look: the leading verdict is not a verdict word at all
# (so the cells may be shifted by a `|` in the body) AND an open word sits somewhere in
# the cell. Deliberately NOT warned on a leading verdict that is merely a version
# string — this table's `Shipped` column legitimately holds `controller v0.227.0`, and
# ten such warnings per run is how a gate teaches people to stop reading it.
warnings.append((n, rid, "leading verdict %r is not a verdict word and an open word "
"sits in the cell — the row's cells may be shifted by a `|` "
"in its body" % verdict[:60]))
closed_ids = set(rid for _, rid, _, _ in rows(CLOSED))
open_ids = {}
for n, rid, _, _ in rows(OPEN):
open_ids.setdefault(rid, n)
for rid in sorted(closed_ids & set(open_ids), key=lambda r: (int(re.sub(r"\D", "", r)), r)):
convicted.append((open_ids[rid], rid, "", "has a row in BOTH registers"))
print("closed-register gate — %d closed rows with a readable state cell, %d open rows, "
"%d convicted, %d warnings" % (checked, len(open_ids), len(convicted), len(warnings)))
for n, rid, why in warnings:
print(" WARNING L%-5d %-7s %s" % (n, rid, why))
if convicted:
print()
print("CONVICTED — CLOSED-ITEMS.md is for closed work only:")
for n, rid, verdict, why in convicted:
print(" L%-5d %-7s %s%s" % (n, rid, why, (" — verdict %r" % verdict) if verdict else ""))
print()
print("Move the row back into OPEN-ITEMS.md keeping its text, state, owner and rank, or —")
print("if it really is finished — write the verdict that says so at the START of its state")
print("cell. R-87 sat in the wrong file for nine days because nobody could see it there.")
return 1
print("closed-register gate OK — no open work filed as closed, no id in both registers.")
return 0
sys.exit(main())
+5
View File
@@ -106,6 +106,11 @@ GATES = [
# finding sat in ROADMAP.md for 25 days invisible to every "grep the register" rule and was
# rediscovered by an overnight drill. Fast: two file reads.
("one-register", os.path.join(SCRIPTS, "one_register_gate.py"), [], True),
# R-405 — the 2026-08-22 compression sweep moved R-87 into CLOSED-ITEMS.md while its own state
# cell read READY; R-378 caught six siblings in the same session and missed this one, so it sat
# in the wrong file for nine days while the ranking paragraph pointed at nothing. Fast: two
# file reads.
("closed-register", os.path.join(SCRIPTS, "closed_register_gate.py"), [], True),
# R-389 — a live finding lived in a REPORT.md observations paragraph and nowhere else, and
# REPORT.md is overwritten every session. Fast: stdlib file reads.
("observations", os.path.join(SCRIPTS, "observations_gate.py"), [ROOT], True),