CB-WP-0006 T08: control loop — 4 of 14 to 8 of 14, and a cost regression
Test 1: the enforced count rose. AM-2, AM-6, AM-9 and AM-11 moved from unmutatable to red; AM-7's hash clause was re-earned so it is 2/3 rather than 1/3. Kernel spec->code link 15/18 -> 18/18, names only. Both denominators are stated. 8 of 14 is 57%, but four rows cannot be enforced — AM-3 blocked on an artifact, AM-4c withdrawn, AM-5 declared ungated by the spec, AM-10 withdrawn — so it is 8 of 10 enforceable. The 14 stays the headline and AM-4c stays in it on purpose: a score improved by deleting the question is not an improvement. Test 2: one row regressed and was caught. Moving AM-6's gate from debug to release turned its mutation SURVIVED, because 4,000 black_box iterations were calibrated against debug's 3.4x headroom and are invisible against release's 20x. The generalizable finding is that a weak mutation is not a fixed property of a row — it can become weak when the row's measurement conditions change, without the row, the mutation or the code being touched. Final SURVIVED count: 0. Test 3 is now mechanical rather than asserted. mutation-check gained an EXPECT-VACUOUS verdict: if a row's expect string appears in PASSING output, the FA guard would accept any failure at all, so the row is reported broken rather than red. Final run: 0 vacuous expects across 14 rows. The control exists because the failure happened — my first expect for AM-2 was "AM-2", which appears in the passing report and would have accepted a compile error as proof of enforcement. The cost result is a refutation, not a win. Mechanical share rose to 50%, the highest ever recorded and above the 38% baseline that motivated CB-WP-0004. That is not a tooling regression: environment setup and task closes are still at zero two passes on. It is the other half of CB-WP-0004 T06's finding arriving in force — text patching (45 turns, $13.79) and orientation (19 turns, $10.97) never had their manual path removed, and a code-heavy pass is exactly where that spends. Mean context 493,486 against a 200,000 target, up from 315,170. SessionShape has stated SS-01..SS-05 since CB-WP-0003 and none has ever been enforced — the only acceptance-adjacent numbers in this project with no gate at all, in a pass whose entire subject was ungated numbers. Numbering corrected: the workplan said CB-EV-0004, which CB-WP-0005 already used. This is CB-EV-0005. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
parent
c51c7c9b47
commit
ce353adde8
2 changed files with 178 additions and 4 deletions
|
|
@ -282,10 +282,20 @@ def check_row(row):
|
|||
|
||||
# Positive control 2: the baseline must be green, or "mutant red"
|
||||
# proves nothing.
|
||||
base_ok, base_tail, _ = run(row.verify)
|
||||
base_ok, base_tail, base_out = run(row.verify)
|
||||
if not base_ok:
|
||||
return "inconclusive", f"baseline already red: {base_tail}"
|
||||
|
||||
# CB-WP-0006 T08: the FA guard is only a guard if its `expect` string
|
||||
# cannot appear in PASSING output. An expect of "AM-2" would match the
|
||||
# normal report and accept any failure at all — which is how the guard
|
||||
# goes vacuous without anyone noticing. My first attempt on AM-2 did
|
||||
# exactly that.
|
||||
if row.expect and row.expect in base_out:
|
||||
return "EXPECT-VACUOUS", (
|
||||
f"expect string {row.expect!r} appears in PASSING output, so it "
|
||||
f"would accept any failure — the FA guard is inert for this row")
|
||||
|
||||
try:
|
||||
mutated = original.replace(old, new, 1)
|
||||
# Positive control 3: the file content must actually differ.
|
||||
|
|
@ -337,12 +347,13 @@ def report(only=None):
|
|||
detail = (f"{len(r.clauses) - len(unmet)}/{len(r.clauses)} "
|
||||
f"clauses enforced")
|
||||
tally[verdict] = tally.get(verdict, 0) + 1
|
||||
if verdict == "HARNESS-BROKEN":
|
||||
if verdict in ("HARNESS-BROKEN", "EXPECT-VACUOUS"):
|
||||
broken.append(r.id)
|
||||
mark = {"red": "red ", "SURVIVED": "SURVIVED ",
|
||||
"unmutatable": "unmutatable", "inconclusive": "inconclusive",
|
||||
"PARTIAL": "PARTIAL ",
|
||||
"HARNESS-BROKEN": "BROKEN ",
|
||||
"EXPECT-VACUOUS": "EXPECT-VOID",
|
||||
"WRONG-REASON": "WRONG-REASON"}[verdict]
|
||||
print(f" [{mark}] {r.id:<6} {r.claim[:52]}")
|
||||
if detail:
|
||||
|
|
@ -366,8 +377,8 @@ def report(only=None):
|
|||
print(f" ({withdrawn} withdrawn row(s) retained in the denominator "
|
||||
f"on purpose — a score\n improved by deleting the question "
|
||||
f"is not an improvement)")
|
||||
for k in ("PARTIAL", "SURVIVED", "WRONG-REASON", "unmutatable",
|
||||
"inconclusive"):
|
||||
for k in ("PARTIAL", "SURVIVED", "WRONG-REASON", "EXPECT-VACUOUS",
|
||||
"unmutatable", "inconclusive"):
|
||||
if tally.get(k):
|
||||
print(f" {k:<13} {tally[k]}")
|
||||
if only:
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue