CB-REV-0003: round 3, and three of four FATAL came from round 2's fixes
Some checks failed
ci / check (push) Failing after 3s
Some checks failed
ci / check (push) Failing after 3s
The pattern is now measured over three rounds: 5 fatal, then 3 (2 from the previous round's corrections), then 4 (3 from them). The corrections are not getting safer. FATAL 1: round 2's short-cell assertion went into regulation.rs only. attack-value.rs — which produced every number in CB-EV-0030's DARVO table — still just warned, and the gate registered to close the finding claimed the property for both. FATAL 2, the sharpest of the three rounds: counting games proves they STARTED. Stopping the engine after one round gives 200 games, all-zero columns and exit 0 — byte for byte the signature CB-EV-0030 says the instrumentation distinguishes from a real result. Both harnesses now require every counted game to have reached an outcome over five rounds. FATAL 3: round 2's `.csv` filter was applied to all three loops, so catalog.yaml and rules_delta.yaml — whose missing digests were round 1's finding — were recorded and then never compared, and never checked against upstream at all. Only the parser loop filters now. FATAL 4: five of six tiebreak comparators had no coverage. GR-E04's tiebreak never executes in any scenario. All four are now covered and mutation-verified; the Blame key needed compensating claims to be reachable at all, since Blame also lowers the coalition score. SERIOUS: "peak held" computed the same number as "peak assigned" for every possible input — the real gap was that START_STRESS was an unchecked constant, now read off the dealt state; cadence="none" was a pure loophole, removed; sibling discovery swapped a hand-written list for hand-written globs and missed metadata.json and VARIANT.md, both named in the package's own changed_files — now walked, and it found them immediately; and "~72,000 games" was unsourced, make panels runs 17,600. Also separated two kinds of number that were presented alike: seats×games is invariant, 363 and 29 vary 7.1%-11.5% across samples. Round 4 owed. The conclusion is not that the work is nearly right — it is that author-made corrections to measurement work should be assumed defective until a fresh reader has attacked them. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
parent
feb68027f9
commit
c4a8a227c0
9 changed files with 386 additions and 57 deletions
|
|
@ -316,11 +316,13 @@ def check_gate_registry(root=REPO):
|
|||
if not target:
|
||||
continue
|
||||
cadence = g.get("cadence")
|
||||
if cadence not in ("all", "manual", "none"):
|
||||
if cadence not in ("all", "manual"):
|
||||
out.append(Finding(
|
||||
"gates", "gates.toml",
|
||||
f"{target!r} declares no cadence — say whether `make all` runs "
|
||||
f"it, or a gate can exist without ever running"))
|
||||
f"{target!r} declares no cadence — say `all` or `manual`. "
|
||||
f"`none` was a pure loophole: a real target that runs "
|
||||
f"nowhere and passes, which is the condition this rule "
|
||||
f"exists to prevent (CB-REV-0003 #7)"))
|
||||
elif cadence == "all" and target not in deps:
|
||||
out.append(Finding(
|
||||
"gates", "gates.toml",
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue