Campaign 10: full write-up of the run-2a harness defects

The previous commit message was truncated by an unescaped paren in the shell, so
the fix detail and the product observations were lost from the record. This adds
them as evidence, where they belong.

Covers: the cc_proof table showing no C010-A row at all (the seed never landed);
both harness defects; why an ambiguous I7 justified stopping a 10-cycle run; the
red-proofed controls; and two transient product observations recorded but NOT
filed as findings - the health probe naming the DB container on the app's port for
about 70s during recovery, and a ValidateDump WARN on a dump taken while the app
was down.
This commit is contained in:
2026-08-01 19:52:51 +02:00
parent 9ca57e591b
commit 80db2c103a
@@ -0,0 +1,78 @@
# Run 2a — the first two violations were HARNESS defects, not product defects (2026-08-01)
Recorded in full, because **a check that fails for the wrong reason is as corrosive as one that
passes for the wrong reason** — and this arc already has six of the latter.
## What fired
cycle 10 HARD-RESET VIOLATION VM returned=True canaries_intact=False
cycle 10 I7 VIOLATION want=C10-C010-A-194530 got=C10-C009-A-192929
snap=helyi restore_ok=True rto=41.0s
I7 reporting a **stale restore** is the campaign's headline invariant failing. It was not real.
## Root cause — evidenced, not reasoned
Cycle 10's permutation was
`detach_target → detach_nontarget → hard_reset_mid_write → redeploy_app → restore_verify → …`,
so the hard reset ran **before** the restore in the same cycle. It left the app container
`rallly | Exited (255)` while `rallly-postgres` stayed healthy.
The decisive evidence is the canary table itself — every generation the campaign ever wrote, in order:
39|C10-C008-A-185905
40|C10-C009-A-192929 <-- highest row
**There is no `C10-C010-A` row at all.** The seed never landed, because `_psql` runs
`docker run --network container:rallly …` and there was no `rallly` container to attach to. The
restore then correctly returned the newest generation that *had* been written — C009-A — and the
harness called that "stale".
Two harness defects, both mine:
1. `atom_restore_verify` called `L.seed(genA)` and **never checked the return value**, so an unwritten
generation was indistinguishable from a stale restore.
2. `atom_hard_reset_mid_write` waited only for `felhom-controller` and then read canaries out of app
containers that were still starting — it measured startup latency, not corruption.
## Why this justified stopping a 10-cycle run
A **real** stale restore would have produced a byte-identical journal row. Leaving the flaw in place
would have made I7 — the one invariant the whole campaign is built around — unable to tell a genuine
data-loss event from a slow container start. Depth is worthless if the verdict is ambiguous.
## The fixes, red-proofed in both directions before restarting
- `L.apps_ready()` waits for the **app** containers, health-aware, not the controller.
- `L.seed(verify=True)` **reads the value back**; a write that reported success but did not land now
returns False.
- `atom_restore_verify` gates on both and records **`I7-SKIP`** — a PASS row carrying the reason —
when preconditions fail. A skipped check is never silently a green one.
- `atom_hard_reset_mid_write` judges canaries only once `apps_ready`; **apps failing to return within
10 minutes is itself the violation**, which is the thing actually worth alarming on.
| Control | Result |
|---|---|
| positive — all apps up | `apps_ready=True` in 2 s; all four seeds land |
| negative — grafana stopped | `seed → {grafana: False, …}`; `apps_ready=False` |
## Product observations from the same window — NOT filed as findings
Both are recorded so a later session does not rediscover them as novel, and neither is claimed as a
defect: the app recovered on its own and the behaviour was transient.
- **The health probe named the wrong container.** During recovery the controller logged
`Health probe rallly: HTTP GET :3000/ → Get "http://rallly-postgres:3000/"` — the **database**
container's hostname on the **app's** port. **5 occurrences inside a ~70 s window**, then correct.
rallly's own log showed `▲ Next.js 16.2.6 ✓ Ready` throughout. An operator reading only this log
would chase a DNS fault on postgres. Transient mis-resolution while the app container was absent.
- **`ValidateDump` WARN:** *"structurally valid (32 tables) but its accounts table has NO rows — the
dump may predate the customer's data"* — the guard firing as designed, on a dump taken while the
app was down.
## Disposition
Run 2a stopped at cycle 10. Its evidence is preserved as `journal-run2a.tsv`, `rto-run2a.tsv`,
`runner-run2a.log`. **The 10 cycles' other invariant rows remain valid** — I1, I2, I3, I4, I5, I6,
I10 and I11 are independent of the seed path. Run 2b restarted from cycle 1 with the fixed harness.