CB-WP-0022 T02: the separate reviewer found the showcase finding was false

First adversarial review in this repo run by a genuinely separate agent.
CB-RES-0006's reviewer opened by conceding it could not be, and called its
own findings "a lower bound on what a genuinely separate reviewer would
find." This is the measurement: the separate reviewer ran git log against
the survey's central example and found 2da19a4 had falsified it four days
earlier, while the author -- who wrote that commit -- quoted the dead
number twice.

Seven challenges: four conceded, two conceded in part, one answered.

C1 changes the design. "GR-E01 is admissible because 4/6/9 against 5/7/9
is a computation anyone can rerun" was a computation already rerun: the
edition import measured 6/9/12, the conclusion inverted, and the scenario
was renamed -unreachable- to -reachable-. That finding was one of the TWO
that passed the reproduction rule. So three wrong premises have now
reached ground-game and the third satisfied an existence test -- existence
is not the property that was missing. The rule gains shape (ground-game's
row-level deal table, promoted from a T04 addendum) and a clause the
survey never contemplated: a reproduction must be able to fail. Ours went
green and stayed admissible.

C1 also caught a defect in flight. T06's payload, status todo, still named
4/6/9 and was queued to send it to ground-game as "no dataset reconciles
them." Withdrawn before sending -- the fourth wrong premise, and the only
one stopped.

C2 withdraws the baseline's precision: design-baseline.py is a
hand-maintained dict counting itself, has_reproduction never checks the
file exists (its YES-control is green against a deleted path), and
Makefile:127 runs only --self-test so the reporting path has no CI. The
direction survives; 33% is not a measured rate and T05 must not build on
it. C3: "six provisional defaults" is five, GR-E01 double-counted. C4:
GroundRules §Underdetermined was never evaluated as a candidate and
already delivers four of five benchmarks -- T03's burden flips to arguing
extension over replacement.

Survived: the rule's affordability, and reuse of the provisional
machinery.

loop-lint: no findings.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
tegwick 2026-08-05 15:03:12 +02:00
parent 129ed03492
commit 04c3a4977f
3 changed files with 617 additions and 9 deletions

View file

@ -0,0 +1,361 @@
# 260805 — challenge to CB-RES-0007
Adversarial review, one round, per InnerLoop §Step 2. Target: the survey
`research/CB-RES-0007-design-instrument.md`, the harness
`tools/design-baseline.py`, and the §judgment in `CB-WP-0022`.
**Fidelity note, first.** Unlike `260802-render-port-challenge.md`, this
review *was* run in a separate agent session with no access to the author's
reasoning trail — only the files. So it has the property §Step 2 asks for.
It does **not** have a different sample: there is exactly one baseline
harness and one repo, so every number below was re-derived from the
author's instrument. Per §Step 2 that is stated rather than reported as a
clean verify. What this review substitutes for a second sample is
**mutation**: each quoted number was traced to the assertion that produces
it and the assertion was inverted. Four of the five headline numbers do not
move when the world they claim to measure moves.
Ranked. C1 and C2 land hardest. C6 and C7 are marked weak — do not spend
the one response round on them.
---
## C1 — The rule's own showcase example is a false finding that shipped a green reproduction
**This is the challenge that lands, and it is worse than the two cases the
task description already knows about.**
`CB-WP-0022:61` states the admissibility rule and then gives its canonical
example:
> *"**GR-E01 is admissible because 4/6/9 against 5/7/9 is a computation
> anyone can rerun**; the SOLVE inertness is admissible because a recorded
> session shows three no-ops."*
Both halves of that sentence are now known to be wrong.
- The SOLVE half: `CB-WP-0023:87` — *"the reported case was not the one I
reported… `validate` already rejected face-down, so it never was
offered."* Already conceded by the author.
- **The GR-E01 half has not been conceded, and it is the more damaging
one.** `4/6/9` is stale. Commit `2da19a4` (CB-WP-0021, 2026-08-04)
measured **6/9/12** against thresholds 5/7/9, `ADR-0011:110` records the
ruled 6/9/12, and the scenario was renamed
`gr-e01-threshold-**unreachable**-2p.yaml`
`gr-e01-threshold-**reachable**-2p.yaml`. The finding did not get
refined; **its conclusion inverted**.
Now the part that breaks the proposal's central claim. GR-E01 is one of the
**two** findings in the survey's `2/6 = 33% with a runnable reproduction`
(`design-baseline.py:26-29`). It is a *success case* of the reproduction
rule. It carried a scenario. The scenario was green. And the scenario
asserted the wrong conclusion for four days, was reported to ground-game on
that basis, and was renamed only when someone re-did the arithmetic by
hand.
So the count is not two bad premises. **It is three, and the third one is
the exemplar the workplan uses to define what a good finding looks like.**
The survey's §3 argument — *"in formal methods a claim without a trace is
not a result, it is a hypothesis"* — is true and irrelevant here. A model
checker's counterexample is generated *by* the checker from the property;
it cannot disagree with the property, because it is the property's witness.
Our scenarios are **hand-authored assertions of an expected value**. A
hand-written scenario that encodes a wrong expectation is a trace that
witnesses the author's belief, not the system's behaviour. §3 imports the
prestige of counterexample traces onto an artifact class that does not have
their guarantee, and it does so in the section that supplies the whole
justification for the rule.
**Required:** the survey cannot claim the reproduction rule is the load-
bearing control while its own two positive instances of that rule both
described the wrong thing. Either (a) name the *additional* control that
would have caught GR-E01 — the ground-game row-level deal table
requirement is the obvious candidate and it is a **shape** requirement, not
an existence one — and demote "carries its reproduction" from *the* rule to
*half* the rule, or (b) withdraw §3's claim that a reproduction is
sufficient to make a finding admissible.
Secondary, and checkable: `4/6/9` is still live in three places —
`workplans/CB-WP-0022:61`, `workplans/CB-WP-0022:283` (task T06, still
`todo`, still planning to report the stale number to ground-game), and
`evidence/CB-EV-0018:153`. **The pass that exists to stop wrong premises
reaching ground-game currently has a wrong premise queued for transmission
to ground-game.** `grep -rn "4/6/9" --include=*.md .` settles it.
## C2 — The baseline is a hand-maintained list reporting on itself, and four of its five numbers are mutation-proof
The survey's §1 heading is *"The baseline is us, and **it is measured**"*.
It is not measured. It is a dict literal at `tools/design-baseline.py:16-36`
with one globbed row, printed back with arithmetic.
Per §Step 2 (v1.4), each number was traced to its assertion and mutated.
**(a) `findings 6`** is `len(FINDINGS)` (`:89`) — the number of keys typed
into a dict at `:16-36`. It is a count of hardcoded rows, not of findings.
The self-test that "pins" it (`:58`) asserts `len(FINDINGS) >= 6` against
the same literal. Nothing in the repo can change this number; only editing
line 16-36 can.
**(b) `2/6 = 33%`** rests on `has_reproduction()` (`:38-43`), which is a
**string prefix test on `p.startswith("scenarios/")`**. It never touches
the filesystem. Mutation:
```
has_reproduction(['scenarios/COMPLETELY-MADE-UP.yaml']) -> True
has_reproduction(['scenarios/']) -> True
has_reproduction(['crates/cb-engine/tests/foo.rs']) -> False
```
Two consequences. First, the self-test's YES-control at `:63` cites
`scenarios/ground/gr-e01-threshold-unreachable-2p.yaml` — **a file that no
longer exists** (renamed in `2da19a4`; `ls scenarios/ground/ | grep e01`).
It passes anyway, because the detector never looks. This is exactly the
render-port C6 class: a control that cannot distinguish "found a
reproduction" from "found a string". The `every cited location exists`
check at `:57` iterates `FINDINGS.values()` only and never sees the
self-test's own literal, so the one guard that would have caught it is
pointed at the wrong list.
Second, the detector **contradicts the rule it is measuring**. `CB-WP-0022:55-56`
admits *"a scenario that fails, an arithmetic check that prints the
contradiction, a recorded game the reader can replay, or **a named test**."*
`has_reproduction` recognises exactly one of those four. A finding backed by
a named Rust test scores 0. The 33% is therefore not the baseline for the
proposed rule; it is the baseline for a narrower rule nobody proposed, and
it is biased **downward**, which flatters the "nobody was ever asked for
one" story in §1.
**(c) `4 days`** is `datetime.date(2026, 8, 3) - datetime.date(2026, 7, 30)`
— two literals on `:96-97`, subtracted on `:98`. No data source. Mutated
`:96` to `2020-01-01`:
```
U1..U10: raised 2020-01-01, first READ 2026-08-03 — 2406 days
--self-test: all 5 checks [ok]
```
The headline number can be set to anything and every control stays green.
Worse, **it disagrees with the repo**: `specs/GroundRules.md:224` states
*"Owner: ground-game. **Raised: 2026-07-31**"*, and
`scenarios/ground/gr-a04-bond-support.yaml:9` carries
`provisional_raised: 2026-07-31`. Under the spec's own date the interval is
**3 days, not 4**. Two sources of the same fact, disagreeing, with zero
`<!-- fact: -->` tags in either the survey or the workplan
(`grep -c "fact:"` returns 0 for both) — the drift surface InnerLoop
§Single-source-of-fact names explicitly.
`4 days` is also the row the survey's §5 calls *"the number this exists to
fix; any mechanism that does not move it has failed"*. It is the least
defensible number in the document.
**(d) `11 files`** is `len(places)` — the deduplicated union of the
hardcoded paths. It is derived, so it is the healthiest of the four, but it
measures **citation sites the author typed**, not files holding findings.
Add a citation, the "findability" benchmark worsens; nothing about the
repo changed.
**(e)** The one number with real provenance is the 5-element glob at
`:30-34` — and see C3, it disagrees with the label above it.
**What the harness would report if the work silently stopped** (§Step 2
requires this be stated): the reporting path, `:80-99`, performs **no
existence check at all**. Delete every file in `FINDINGS` and it still
prints `6 findings, 2/6 = 33%, 11 files, 4 days`. The existence guard lives
only in `--self-test`, and `Makefile:127` runs **only** `--self-test`
the numbers quoted in the survey come from a code path CI never executes.
**Required:** either the survey stops calling §1 "measured" and labels the
baseline what it is — a hand-curated snapshot, which is a legitimate thing
to have but caps at the same directional status the survey correctly
assigns to the external rows — or the harness derives its finding set from
the repo (the `provisional:` marks, `evidence/` frontmatter, the U-list)
so that mutating the repo moves the number.
## C3 — "six provisional defaults" is five, and the survey's own instrument says so
`research/CB-RES-0007:46` lists `six provisional defaults | 6 | yes`.
```
$ python3 tools/design-baseline.py
repro 5 location(s) six provisional defaults
$ make coverage
provisional U-item defaults: 5
```
Six became five in `2da19a4` when `gr-e01-threshold-unreachable-2p.yaml`
lost its `provisional: true` (verified:
`git grep -l "provisional: true" 2da19a4^ -- scenarios/ground/ | wc -l`
**6**; at HEAD → **5**). The author edited `design-baseline.py:22-29` in
that same window to add the `[RESOLVED]` comment on the GR-E01 row, and did
not notice that the *next* row's glob had silently dropped a member and its
hardcoded label `"six provisional defaults"` had gone stale.
Two things follow.
1. **The survey's table row 5 (`locations: 6`) is a number the survey's own
tool contradicts on stdout**, and no self-test compares the word "six"
to `len()`. §4 of the survey argues the register *"must reuse [the
provisional machinery], not compete with it"* — and the very row that
reads the provisional machinery is the one that drifted from it.
2. **GR-E01 was double-counted.** It is row 4 *and* was a member of row 5's
six. The denominator of `2/6` contained the same finding twice, and one
of the two numerator hits was that finding.
**Required:** correct the row, and state whether the `2/6` survives the
double-count. If rows 4 and 5 overlap, the honest denominator at survey
time was 5 distinct findings, and the headline rate changes.
## C4 — `GroundRules.md §Underdetermined` was never evaluated as the incumbent, and it already implements four of the five benchmarks
T02 requires an attempt at *"whether a register is needed at all, rather
than one more section in `GroundRules.md §Underdetermined`."* The survey
does not attempt it. §4 mentions the provisional machinery only as
*machinery to reuse*; §5's benchmark table has a `today` column that never
names §Underdetermined; and §6's "did not settle" list does not include it.
Per §Step 2's row for surveys, **that is an omitted candidate, and it is
the strongest one, because it is the incumbent.**
Read against the survey's own five benchmarks (`specs/GroundRules.md:222-277`):
| benchmark (§5) | §Underdetermined today |
|---|---|
| findability: *one register with an index* | **has it** — one section, ten numbered items, `grep U7` |
| reproducibility: *runnable artifact* | **has the obligation***"every U-item's default must be covered by a scenario tagged `provisional: true`"* (`:274-277`) |
| closure: *closes when the source changed* | **has it***"A ruling flips the **scenario**, not the kernel; a correction is a data change"* (`:231-233`). This is the Magic Oracle property the survey went to Magic to discover. It has been in the spec since 2026-07-31. |
| raised → read, with age | **has it, executable** — owner + raise date per item, `make coverage` prints `age=5d` per item and warns past 30 days |
| taxonomy fits reality | not addressed — genuinely absent |
The survey travels to Magic: the Gathering to extract *"a finding closes
when the source changes, not when an annotation is added"* and presents it
as a corrected assumption (§2, and `CB-WP-0022:136-139` treats it as the
pass's main external yield). **The repo already states that property in
those terms, one directory over.** That is not a small omission: it is the
survey's headline external finding being rediscovered from the incumbent it
failed to survey.
**Where this challenge stops, honestly.** §Underdetermined has a real gap
and the review found it: it obliges ten U-items to carry provisional
scenarios and **five exist**. `make coverage` prints `provisional U-item
defaults: 5` and does *not* flag the five missing — it reports what exists,
never what is owed. That is a genuine argument for something with a
denominator, and it is stronger than any argument the survey actually
makes. It is also an argument for *one gate in `rule-coverage.py`*, not for
a new register.
**Required:** evaluate §Underdetermined as a candidate row with a `today`
column, and state what a register does that §Underdetermined + one
completeness check cannot. If the answer is "holds findings that are not
U-items" — say which of the six, and note that four of the six are
single-file prose notes in `evidence/`.
## C5 — The engine-evolution "third thing" exists, and its homelessness is visible in the filesystem
`CB-WP-0022:73-84` argues the engine-evolution register is redundant
because *"`evidence/CB-EV-*` records what each pass found, `decisions/ADR-*`
records what was decided…, `gates.toml` records what every control has
caught…, and workplans record what was attempted."*
The list is incomplete, and the omission is the whole question. The third
thing — *what we learned about building engines* — is
**`specs/InnerLoopReference.md`, 383 lines**, and its section list is
nothing but that:
```
§ Implementation rules the first pass earned
§ Why review is adversarial execution, not re-derivation
§ Numbers you inherit — two cases from CB-WP-0002
§ Quoting a cost — six for six, never high
§ Chaos roll — window 1's verdict and the d4 → d8 change
§ Retargeting — the case, and the four corrections it must not forbid
```
None of those is a decision (ADR), a per-pass finding (evidence), or a
control (gates.toml). They are cross-pass lessons about method. Plus six
`history/*-retrospective.md` files, 769 lines, in the same category.
**And the homelessness is on the record in the headings themselves.** Two
sections are literally titled *"(moved from InnerLoop v1.6)"* — the content
grew inside the normative spec until it had to be evicted into an
appendix-shaped file that is not in the author's list of homes. That is
precisely the signature T02 asked the reviewer to look for: *content that
had nowhere to go*.
**Where this challenge stops.** It does not establish that a *register* is
needed. It establishes that the author's redundancy argument is made
against the wrong inventory. `InnerLoopReference.md` and `history/` **are**
a home, so the conclusion "no second register" probably survives — but it
survives for a different reason than the one given, and the difference
matters for T03: the asymmetry claimed at `:83` (*"engine evolution has a
home and game design does not"*) is real, but the home is an unindexed
383-line spec appendix plus six retrospectives, i.e. **exactly the
`11 files, no index` condition the survey calls a defect when it is game
findings**. The survey applies one standard to game findings and another to
engine lessons and does not state the distinction.
**Required:** name `InnerLoopReference.md` and `history/*-retrospective.md`
in the redundancy argument, and say why "several files, no index" is
disqualifying for one subject and acceptable for the other.
## C6 — (weak) The survey's own §2 conclusion contradicts the harness's stated design
§2: *"A register that accumulates permanent rulings is a system that has
given up on fixing its rules. Ours should be a **queue that empties**, not
an archive that grows."*
`design-baseline.py:23-25`: *"Kept in the baseline because… a register that
drops findings when they close cannot report a close rate."*
Both are defensible and they are not strictly contradictory — a queue can
retain closed items out-of-band. But T03 has to pick, and the two artifacts
of this pass currently point opposite ways, with the tool's comment being
the more considered of the two. Marked **weak**: this is a clarification,
not a defect.
## C7 — (weak) `n=6`, all engine-surfaced, and §6 already concedes it
§6 concedes the sample is six and all engine-surfaced, and flags the
play-sourced finding (*"DARVO feels punishing at 2 players"*) as unsettled.
That concession is correct and I have nothing to add to it, so this is not
a challenge — it is noted only so the response does not spend a round
defending something already conceded. Marked **weak**.
---
## Verdict
**Not approvable as written.** C1, C2 and C3 each require a change to the
survey, not a clarification. C4 requires an added candidate row. C5
requires a corrected inventory in the workplan's §judgment.
| # | verdict |
|---|---|
| **C1** | **lands hardest.** Requires demoting "carries its reproduction" from sufficient to necessary-but-not-sufficient, and adding the shape requirement. GR-E01 is a rule-satisfying finding that was wrong. |
| **C2** | **lands.** Four of five headline numbers do not move under mutation; `4 days` is a literal and disagrees with `GroundRules.md:224` (3 days). §1's word "measured" must go or the harness must derive. |
| **C3** | **lands.** "six" is five; the survey's table contradicts its own tool's stdout; GR-E01 double-counted in the `2/6` denominator. |
| **C4** | **lands, moderate.** The incumbent was never surveyed, and it already holds the Magic Oracle property the survey travelled to Magic to find. Weakened by a real gap the review found for the author: 10 U-items owe scenarios, 5 exist. |
| **C5** | **partial.** The third thing exists and is visibly homeless; the redundancy *conclusion* probably survives, the *argument* does not. |
| **C6** | **weak.** Clarification for T03. |
| **C7** | **weak / already conceded.** |
**What survives.** Two claims attacked and not broken:
- **That the reproduction rule is affordable.** §1's four cheap
reproductions were checked and the estimates are plausible; `bot::play`
exists, the scenario harness exists, and the SOLVE case in fact got its
reproduction in CB-WP-0023 T01 for a fraction of a pass. What would have
falsified it: a finding among the four whose reproduction needed engine
work that does not exist. There is none. **Affordability is not the
problem; sufficiency is (C1).**
- **That the register must reuse the provisional machinery rather than
compete with it** (§4). Checked against `rule-coverage.py:98-108,285-308`
— owner, raise date, age and a warn threshold all already exist and are
executable. What would have falsified it: a second age-tracking mechanism
already in the repo. There is none.
**The single challenge that forces a change to the design: C1.** The
proposal's load-bearing rule is stated as an existence test on a
reproduction. Three findings have now reached the record on wrong premises,
and **the third one satisfied the existence test**. The rule as written
does not catch its own worst observed failure. T03 cannot adopt it in the
form `CB-WP-0022:53` states it.

View file

@ -0,0 +1,187 @@
# 260805 — response to the challenge on CB-RES-0007
One round, per InnerLoop §Step 2. **This review was run by a separate
agent**, unlike CB-RES-0006's, whose challenge opened with a fidelity note
saying it could not be. That note said *"treat the challenges below as a
lower bound on what a genuinely separate reviewer would find."* It was
right, and this is what the difference looks like: the separate reviewer
went to `git log` and found that the survey's showcase example had been
falsified four days earlier by our own commit.
Four challenges conceded, two conceded in part, one answered. **The
load-bearing rule survives but does not survive intact** — it is
insufficient as stated, and C1 is the reason.
---
## C1 — the showcase finding was itself a wrong premise, and it passed the rule — **conceded, and it is larger than the challenge says**
The challenge is correct and I verified every step of it independently
rather than accepting it.
`CB-WP-0022:61` defines admissibility by example:
> *"GR-E01 is admissible because 4/6/9 against 5/7/9 is a computation
> anyone can rerun."*
Commit `2da19a4` (2026-08-04, CB-WP-0021 T01/T02/T05) measured the deal
after the edition import: **6/9/12 against thresholds 5/7/9 — the game is
reachable at every seat count.** The scenario was renamed
`gr-e01-threshold-unreachable-2p.yaml``gr-e01-threshold-reachable-2p.yaml`.
The finding did not merely get a new number. **Its conclusion inverted.**
So the count is not two wrong premises reaching `ground-game`. It is
three, and **the third one was in the 2-of-6 set that passed the
reproduction rule.** The rule's own success case is a false finding that
shipped a green artifact.
**And it was still loaded.** `CB-WP-0022:283`, task T06, status `todo`:
> *"GR-E01 vs GR-S01 — the deal count puts 4/6/9 points in play against
> thresholds of 5/7/9, so either the count or the thresholds are wrong and
> no dataset reconciles them"*
That is queued for transmission to `ground-game` as a finding. Had T06 run
before this review, it would have sent a **fourth** wrong premise
downstream — one asserting no dataset reconciles numbers our own dataset
reconciled on 2026-08-04. The review did not catch a stale citation. It
caught a defect in flight.
**What this forces.** The rule at `CB-WP-0022:53` is an *existence* test:
a finding is admissible if a reproduction exists. Three of three wrong
premises would have passed an existence test, and one did. Existence is
therefore not the property that was missing. T03 must adopt the shape
requirement — `ground-game`'s row-level deal table, GROUND-WP-0004 T02 —
as **part of the rule**, not as the T04 addendum it currently is. And it
needs a second clause the survey did not contemplate: **a reproduction
must be able to fail.** A finding whose artifact has gone green is not a
closed finding until someone says which it is; ours silently became a
counterexample to itself.
`4/6/9` is corrected wherever it appears in this pass's artifacts, and
`evidence/CB-EV-0018:153` is left alone — it is a dated record of what was
believed then, and rewriting evidence is worse than a stale number in it.
## C2 — the baseline is a hand-maintained inventory reporting on itself — **conceded**
`tools/design-baseline.py:16-36` is a literal dict. `n = len(FINDINGS)`
(`:89`). **"6 findings" is a count of rows someone typed.** It is not a
measurement of anything, and the survey presented it as a measured
baseline, in a repo whose §Single source of fact rule exists precisely to
stop typed numbers.
Worse, and the part I would not have found:
- `has_reproduction` (`:38-43`) is `p.startswith("scenarios/")` and
nothing else. **It never checks the file exists.** The self-test's
YES-control (`:63`) passes a path — `gr-e01-threshold-unreachable-2p.yaml`
that `2da19a4` deleted. The positive control for "detection can say
YES" is green against a file that is not there. That is a positive
control that controls nothing, which is the exact class `gates.toml`
exists to catch.
- The same detector **rejects a named test**, which `CB-WP-0022:55-56`
explicitly admits as a reproduction. So `33%` is biased down by the
detector's own narrowness.
- The 4-day figure is two date literals subtracted (`:96-98`). The
reviewer mutated them to 2020-01-01, got **2406 days**, and all five
self-test checks stayed green.
- `Makefile:127` runs `--self-test` only. **The reporting path is never
executed by CI**, so none of the above was ever going to be caught.
**What survives, and it matters that something does.** The *direction* of
every headline number is checkable without the tool: 11 files with no
index is verifiable by reading, and U1U10 being open with 0 of 10 ruled
is verifiable against the hub. The baseline's conclusion — *we have
findings scattered across the repo with no register and no reproductions* —
is not in doubt. **What is withdrawn is its precision**, and any claim
that the 33% is a measured rate.
T05 must not build on this tool. The register itself is the measurement,
and `make design` reports over the register's contents, not over a dict.
## C3 — the arithmetic in the survey's own table — **conceded**
"Six provisional defaults" is **five**. The tool's own line prints
`5 location(s)`; `make coverage` prints 5; the survey table (`:46`) says
6. It was 6 at `2da19a4^` and 5 at HEAD — the survey table was written
against a repo state that had already moved.
And **GR-E01 is double-counted**: it is row 4 in its own right and was a
member of row 5's six. So the `2/6` denominator contains one finding
twice, and both copies are in the numerator. The real figure, on the
survey's own definitions, is not 2/6.
I am not restating a corrected rate here. **The tool that would produce it
is the one C2 just withdrew**, and quoting a new number from a discredited
instrument is how the first one got in.
## C4 — §Underdetermined was never evaluated as a candidate — **conceded in part**
Correct, and it is the omission that embarrasses the survey most, because
§Step 1 asks for the best implementation *in existence* and this one is in
our own repo. The survey went to Magic to discover the Oracle property —
*the fix lands in the source, the annotation is not authoritative* — and
`specs/GroundRules.md:231-233` already states it in our own words:
*"a ruling flips the scenario, not the kernel."*
**Conceded in part, not whole,** and the reviewer marked the honest limit
themselves: 10 U-items owe provisional scenarios, 5 exist, and
`make coverage` never flags the missing five. A section that cannot report
its own gaps is not yet the register. But T03's burden has changed
direction — it must now argue why §Underdetermined should be **extended**
rather than replaced, and the default answer is that it should be.
## C5 — the third thing exists — **conceded as to the argument, not the conclusion**
`specs/InnerLoopReference.md` is 383 lines with two sections literally
headed *"moved from InnerLoop v1.6"*, and `history/` holds 769 lines of
retrospectives. That is content about *how we learn to build this engine*
which was placed where it fit rather than where it belonged. My redundancy
inventory named `evidence/`, `decisions/`, `gates.toml` and workplans, and
**named neither of those** — so the inventory was incomplete and the
argument as written does not stand.
The conclusion probably does, and for a reason the challenge concedes: the
homeless content found homes, and both are load-bearing and read. A third
register would be a third place to look for material that is currently in
two. **T03 records this as argued-and-narrowed rather than settled**, which
is the outcome the declaration asked for when it said this was the
judgment it wanted reviewed rather than assumed.
## C6 — "queue that empties" vs keeping resolved findings — **answered**
Marked weak by the reviewer and it is weak, but it has a real answer.
Magic's property is that a finding **closes when the source changes**. It
is not that the record is deleted. `design-baseline.py:23-25`'s comment —
keep resolved findings or you cannot report a close rate — is compatible:
the *queue* of open findings empties, the *log* of closed ones does not.
T03's lifecycle should say this in one sentence so the next reader does
not re-litigate it.
Falsifier: if the register's default view shows closed findings mixed with
open ones, the queue property is gone regardless of what the lifecycle
says.
## C7 — n=6, all engine-surfaced — **already conceded in §6**
The survey states this limit itself. Nothing to add.
---
## What the review changed
| | before | after |
|---|---|---|
| the rule | reproduction must **exist** | must exist, **have the ruled shape**, and **be able to fail** |
| wrong premises shipped | 2 | **3** — and one passed the rule |
| the baseline | a measured 33% | a hand-maintained inventory; direction stands, precision withdrawn |
| §Underdetermined | not considered | the incumbent T03 must argue against |
| engine-evolution register | redundant | still redundant, on a corrected inventory |
| T06's payload | a finding | **a fourth wrong premise, caught before sending** |
**The strongest thing here is not any single challenge.** It is that a
separate reviewer with no stake ran `git log` against the survey's central
example and found it had been falsified by our own commit four days
earlier — while the author, who wrote that commit, quoted the dead number
twice. CB-RES-0006's review said a same-session reviewer would be a lower
bound. This is how much lower.

View file

@ -57,9 +57,21 @@ register that collects opinions would reproduce it in a new medium.
Concretely: a scenario that fails, an arithmetic check that prints the
contradiction, a recorded game the reader can replay, or a named test.
*"This feels unbalanced"* is a note, not a finding. **GR-E01 is admissible
because 4/6/9 against 5/7/9 is a computation anyone can rerun; the SOLVE
inertness is admissible because a recorded session shows three no-ops.**
*"This feels unbalanced"* is a note, not a finding. **The SOLVE inertness
is admissible because a recorded session shows three no-ops.**
> **The example that stood here was GR-E01, and the adversarial review
> killed it (C1, 2026-08-05).** *"4/6/9 against 5/7/9 is a computation
> anyone can rerun"* was a computation that had already been rerun:
> `2da19a4` measured **6/9/12 against 5/7/9** and renamed the scenario
> `-unreachable-``-reachable-`. The finding's conclusion inverted, and
> it was one of the two findings that **passed** this rule.
>
> So existence is not the property that was missing — three wrong premises
> have now reached `ground-game`, and the third satisfied an existence
> test. T03 must adopt the shape requirement as part of the rule, plus a
> clause the survey never contemplated: **a reproduction must be able to
> fail.** Ours went green and stayed admissible.
This is what would make clay-borg a design tool rather than a suggestion
box, and it is the one part of this proposal that must not be traded away
@ -147,7 +159,7 @@ than duplicate.
```task
id: CB-WP-0022-T02
status: todo
status: done
priority: high
state_hub_task_id: "d2597895-fe2e-4f1e-a3db-a5fb8833434c"
```
@ -173,6 +185,46 @@ above, and require an attempt at:
Record the trail in `history/`, unpolished.
**Done 2026-08-05.** Trail:
[challenge](../history/260805-design-instrument-challenge.md),
[response](../history/260805-design-instrument-response.md).
**Run by a separate agent** — the first in this repo that was. CB-RES-0006's
review opened by conceding it could not be, and called its own findings
*"a lower bound on what a genuinely separate reviewer would find."* That
was measurable, and this is the measurement: the separate reviewer ran
`git log` against the survey's central example and found our own commit
had falsified it four days earlier, while the author — who wrote that
commit — quoted the dead number twice.
**Seven challenges: four conceded, two conceded in part, one answered.**
- **C1 lands hardest and changed the design.** The rule's showcase finding
was false and had *passed* the rule. Existence is not the missing
property; **shape** and **falsifiability** are. Folded into §The
load-bearing rule above, and it is T03's to settle.
- **C1 also caught a defect in flight** — T06's payload, `todo`, still
named the dead number. Withdrawn above before sending.
- **C2 withdrew the baseline's precision.** `tools/design-baseline.py` is a
hand-maintained dict counting itself (`:16-36`, `:89`); `has_reproduction`
(`:38-43`) never checks the file exists, so the self-test's YES-control
(`:63`) is green against a path `2da19a4` deleted. `Makefile:127` runs
only `--self-test`, so the reporting path has no CI. The direction
stands — 11 files, no index, 0 of 10 ruled are all checkable without the
tool — but **33% is not a measured rate and T05 must not build on it.**
- **C3**: "six provisional defaults" is five, and GR-E01 is double-counted
in the `2/6`. No corrected rate is quoted here; the instrument that would
produce it is the one C2 withdrew.
- **C4**: `§Underdetermined` was never evaluated as a candidate, and it
already delivers four of five benchmarks including the Oracle property
the survey went to Magic to find. **T03's burden flips: argue why it is
extended, not replaced.**
- **C5**: the engine-evolution "third thing" is visible in
`specs/InnerLoopReference.md` and `history/`'s retrospectives, neither of
which my redundancy inventory named. Conclusion narrowed, not settled.
- **Survived**: the reproduction rule's *affordability*, and §4's reuse of
the provisional machinery. Both with stated falsifiers.
## Task: decide
```task
@ -278,12 +330,20 @@ So the report must land somewhere that persists: a file in `ground-game`
under its own workplan, not only an inbox entry. GROUND-WP-0002 already
holds the ten U-items; this should extend it rather than duplicate it.
Include the two sharpened findings this pass has already produced:
Include the findings this pass has sharpened:
- **GR-E01 vs GR-S01** — the deal count puts 4/6/9 points in play against
thresholds of 5/7/9, so either the count or the thresholds are wrong and
no dataset reconciles them;
- **SOLVE's legality** against a face-down Problem or an unmatchable suit.
- **SOLVE's legality** against a face-down Problem or an unmatchable suit —
and note that the case we *reported* was not the case that fired
(CB-WP-0023 T01).
- ~~**GR-E01 vs GR-S01** — 4/6/9 against 5/7/9, no dataset reconciles
them~~ — **withdrawn 2026-08-05, before sending.** The adversarial
review (C1) found `2da19a4` had already measured **6/9/12 against
5/7/9**: the dataset reconciles them and the scenario is now
`-reachable-`. Sending this would have been the **fourth** wrong premise
to reach `ground-game`, and the only one caught before transmission.
**Report the withdrawal, not the finding** — GROUND-WP-0002 holds the
original, and a claim retracted silently is how the first three
survived.
## Task: evidence