Files
felhom.eu/documentation/tests/campaign10-evidence-2026-07-31/run2a-violations-were-harness.md
T
admin 80db2c103a Campaign 10: full write-up of the run-2a harness defects
The previous commit message was truncated by an unescaped paren in the shell, so
the fix detail and the product observations were lost from the record. This adds
them as evidence, where they belong.

Covers: the cc_proof table showing no C010-A row at all (the seed never landed);
both harness defects; why an ambiguous I7 justified stopping a 10-cycle run; the
red-proofed controls; and two transient product observations recorded but NOT
filed as findings - the health probe naming the DB container on the app's port for
about 70s during recovery, and a ValidateDump WARN on a dump taken while the app
was down.
2026-08-01 19:52:51 +02:00

4.2 KiB

Run 2a — the first two violations were HARNESS defects, not product defects (2026-08-01)

Recorded in full, because a check that fails for the wrong reason is as corrosive as one that passes for the wrong reason — and this arc already has six of the latter.

What fired

cycle 10  HARD-RESET  VIOLATION  VM returned=True canaries_intact=False
cycle 10  I7          VIOLATION  want=C10-C010-A-194530 got=C10-C009-A-192929
                                 snap=helyi restore_ok=True rto=41.0s

I7 reporting a stale restore is the campaign's headline invariant failing. It was not real.

Root cause — evidenced, not reasoned

Cycle 10's permutation was detach_target → detach_nontarget → hard_reset_mid_write → redeploy_app → restore_verify → …, so the hard reset ran before the restore in the same cycle. It left the app container rallly | Exited (255) while rallly-postgres stayed healthy.

The decisive evidence is the canary table itself — every generation the campaign ever wrote, in order:

…
39|C10-C008-A-185905
40|C10-C009-A-192929        <-- highest row

There is no C10-C010-A row at all. The seed never landed, because _psql runs docker run --network container:rallly … and there was no rallly container to attach to. The restore then correctly returned the newest generation that had been written — C009-A — and the harness called that "stale".

Two harness defects, both mine:

  1. atom_restore_verify called L.seed(genA) and never checked the return value, so an unwritten generation was indistinguishable from a stale restore.
  2. atom_hard_reset_mid_write waited only for felhom-controller and then read canaries out of app containers that were still starting — it measured startup latency, not corruption.

Why this justified stopping a 10-cycle run

A real stale restore would have produced a byte-identical journal row. Leaving the flaw in place would have made I7 — the one invariant the whole campaign is built around — unable to tell a genuine data-loss event from a slow container start. Depth is worthless if the verdict is ambiguous.

The fixes, red-proofed in both directions before restarting

  • L.apps_ready() waits for the app containers, health-aware, not the controller.
  • L.seed(verify=True) reads the value back; a write that reported success but did not land now returns False.
  • atom_restore_verify gates on both and records I7-SKIP — a PASS row carrying the reason — when preconditions fail. A skipped check is never silently a green one.
  • atom_hard_reset_mid_write judges canaries only once apps_ready; apps failing to return within 10 minutes is itself the violation, which is the thing actually worth alarming on.
Control Result
positive — all apps up apps_ready=True in 2 s; all four seeds land
negative — grafana stopped seed → {grafana: False, …}; apps_ready=False

Product observations from the same window — NOT filed as findings

Both are recorded so a later session does not rediscover them as novel, and neither is claimed as a defect: the app recovered on its own and the behaviour was transient.

  • The health probe named the wrong container. During recovery the controller logged Health probe rallly: HTTP GET :3000/ → Get "http://rallly-postgres:3000/" — the database container's hostname on the app's port. 5 occurrences inside a ~70 s window, then correct. rallly's own log showed ▲ Next.js 16.2.6 ✓ Ready throughout. An operator reading only this log would chase a DNS fault on postgres. Transient mis-resolution while the app container was absent.
  • ValidateDump WARN: "structurally valid (32 tables) but its accounts table has NO rows — the dump may predate the customer's data" — the guard firing as designed, on a dump taken while the app was down.

Disposition

Run 2a stopped at cycle 10. Its evidence is preserved as journal-run2a.tsv, rto-run2a.tsv, runner-run2a.log. The 10 cycles' other invariant rows remain valid — I1, I2, I3, I4, I5, I6, I10 and I11 are independent of the seed path. Run 2b restarted from cycle 1 with the fixed harness.