Compare commits
No commits in common. "561facce434c48db5815bd67d318ea3f396bfa80" and "2b9fbf5fd3476aa66df02d3d9751ce9e41927346" have entirely different histories.
561facce43
...
2b9fbf5fd3
14 changed files with 82 additions and 1750 deletions
10
INTENT.md
10
INTENT.md
|
|
@ -5,7 +5,7 @@ assimilates and optimizes techniques and implementations useful for games,
|
||||||
simulations, and robotics.
|
simulations, and robotics.
|
||||||
|
|
||||||
It is not another monolithic game engine. It is a capability-assimilating
|
It is not another monolithic game engine. It is a capability-assimilating
|
||||||
development engine with four distinct properties:
|
development engine with three distinct properties:
|
||||||
|
|
||||||
1. **Clay** — its canonical models, contracts, rules, and tools remain malleable.
|
1. **Clay** — its canonical models, contracts, rules, and tools remain malleable.
|
||||||
2. **Borg** — mature, optimized libraries are assimilated behind controlled
|
2. **Borg** — mature, optimized libraries are assimilated behind controlled
|
||||||
|
|
@ -13,14 +13,6 @@ development engine with four distinct properties:
|
||||||
3. **Product-driven evolution** — abstractions are extracted from working
|
3. **Product-driven evolution** — abstractions are extracted from working
|
||||||
games, beginning with **GROUND — A Game of Bonds and Rivalry: DARVO
|
games, beginning with **GROUND — A Game of Bonds and Rivalry: DARVO
|
||||||
Edition**, rather than invented in isolation.
|
Edition**, rather than invented in isolation.
|
||||||
4. **Instrument** — the engine is rigorous enough that it cannot proceed
|
|
||||||
past a rule that does not decide. What it cannot execute, it reports:
|
|
||||||
findings about the *game's* design are a product of building the
|
|
||||||
simulator, not a side activity, and they are carried back to the game's
|
|
||||||
owner with the artifact that produced them. *(ADR-0012. A restatement of
|
|
||||||
what has already happened six times, made a duty. If a pass ever
|
|
||||||
tolerates an undecided rule by quietly picking a default and not raising
|
|
||||||
it, this property is false.)*
|
|
||||||
|
|
||||||
The central rule:
|
The central rule:
|
||||||
|
|
||||||
|
|
|
||||||
8
Makefile
8
Makefile
|
|
@ -125,7 +125,6 @@ self-tests:
|
||||||
$(PY) $(TOOLS)/runtime-metrics.py --self-test
|
$(PY) $(TOOLS)/runtime-metrics.py --self-test
|
||||||
$(PY) $(TOOLS)/replay-test.py --self-test
|
$(PY) $(TOOLS)/replay-test.py --self-test
|
||||||
$(PY) $(TOOLS)/design-baseline.py --self-test
|
$(PY) $(TOOLS)/design-baseline.py --self-test
|
||||||
$(PY) $(TOOLS)/design.py --self-test
|
|
||||||
$(PY) $(TOOLS)/edition-check.py --self-test
|
$(PY) $(TOOLS)/edition-check.py --self-test
|
||||||
|
|
||||||
# T01 positive control: prove the environment fix, do not assume it. Runs
|
# T01 positive control: prove the environment fix, do not assume it. Runs
|
||||||
|
|
@ -161,13 +160,6 @@ facts-check:
|
||||||
facts-gen:
|
facts-gen:
|
||||||
$(PY) $(TOOLS)/facts.py --gen
|
$(PY) $(TOOLS)/facts.py --gen
|
||||||
|
|
||||||
# CB-WP-0022 T05: the design-finding register, reported over
|
|
||||||
# specs/GroundRules.md. Shows the QUEUE by default; the log of closed
|
|
||||||
# findings is a line, not a listing, because a default view that mixes
|
|
||||||
# them loses the queue property (ADR-0012 D5).
|
|
||||||
design:
|
|
||||||
@$(PY) $(TOOLS)/design.py
|
|
||||||
|
|
||||||
# T03: one-shot orientation — workplans, next task, spend, fast gates.
|
# T03: one-shot orientation — workplans, next task, spend, fast gates.
|
||||||
# Cheap by design: no build. Start a session with this instead of grepping.
|
# Cheap by design: no build. Start a session with this instead of grepping.
|
||||||
status:
|
status:
|
||||||
|
|
|
||||||
|
|
@ -31,8 +31,6 @@
|
||||||
| workplan | CB-WP-0021 | active | — | workplans/CB-WP-0021-import-the-edition.md |
|
| workplan | CB-WP-0021 | active | — | workplans/CB-WP-0021-import-the-edition.md |
|
||||||
| workplan | CB-WP-0022 | active | — | workplans/CB-WP-0022-the-design-instrument.md |
|
| workplan | CB-WP-0022 | active | — | workplans/CB-WP-0022-the-design-instrument.md |
|
||||||
| workplan | CB-WP-0023 | done | — | workplans/CB-WP-0023-solve-legality.md |
|
| workplan | CB-WP-0023 | done | — | workplans/CB-WP-0023-solve-legality.md |
|
||||||
| workplan | CB-WP-0024 | ready | — | workplans/CB-WP-0024-the-table-you-can-watch.md |
|
|
||||||
| workplan | CB-WP-0025 | ready | — | workplans/CB-WP-0025-could-we-have-won.md |
|
|
||||||
| task | CB-WP-0001-T01 | done | — | workplans/CB-WP-0001-inner-loop.md |
|
| task | CB-WP-0001-T01 | done | — | workplans/CB-WP-0001-inner-loop.md |
|
||||||
| task | CB-WP-0001-T02 | done | — | workplans/CB-WP-0001-inner-loop.md |
|
| task | CB-WP-0001-T02 | done | — | workplans/CB-WP-0001-inner-loop.md |
|
||||||
| task | CB-WP-0001-T03 | done | — | workplans/CB-WP-0001-inner-loop.md |
|
| task | CB-WP-0001-T03 | done | — | workplans/CB-WP-0001-inner-loop.md |
|
||||||
|
|
@ -156,15 +154,3 @@
|
||||||
| task | CB-WP-0023-T01 | done | — | workplans/CB-WP-0023-solve-legality.md |
|
| task | CB-WP-0023-T01 | done | — | workplans/CB-WP-0023-solve-legality.md |
|
||||||
| task | CB-WP-0023-T02 | done | — | workplans/CB-WP-0023-solve-legality.md |
|
| task | CB-WP-0023-T02 | done | — | workplans/CB-WP-0023-solve-legality.md |
|
||||||
| task | CB-WP-0023-T03 | done | — | workplans/CB-WP-0023-solve-legality.md |
|
| task | CB-WP-0023-T03 | done | — | workplans/CB-WP-0023-solve-legality.md |
|
||||||
| task | CB-WP-0024-T01 | todo | — | workplans/CB-WP-0024-the-table-you-can-watch.md |
|
|
||||||
| task | CB-WP-0024-T02 | todo | — | workplans/CB-WP-0024-the-table-you-can-watch.md |
|
|
||||||
| task | CB-WP-0024-T03 | todo | — | workplans/CB-WP-0024-the-table-you-can-watch.md |
|
|
||||||
| task | CB-WP-0024-T04 | todo | — | workplans/CB-WP-0024-the-table-you-can-watch.md |
|
|
||||||
| task | CB-WP-0024-T05 | todo | — | workplans/CB-WP-0024-the-table-you-can-watch.md |
|
|
||||||
| task | CB-WP-0025-T01 | todo | — | workplans/CB-WP-0025-could-we-have-won.md |
|
|
||||||
| task | CB-WP-0025-T02 | todo | — | workplans/CB-WP-0025-could-we-have-won.md |
|
|
||||||
| task | CB-WP-0025-T03 | todo | — | workplans/CB-WP-0025-could-we-have-won.md |
|
|
||||||
| task | CB-WP-0025-T04 | todo | — | workplans/CB-WP-0025-could-we-have-won.md |
|
|
||||||
| task | CB-WP-0025-T05 | todo | — | workplans/CB-WP-0025-could-we-have-won.md |
|
|
||||||
| task | CB-WP-0025-T06 | todo | — | workplans/CB-WP-0025-could-we-have-won.md |
|
|
||||||
| task | CB-WP-0025-T07 | todo | — | workplans/CB-WP-0025-could-we-have-won.md |
|
|
||||||
|
|
|
||||||
|
|
@ -1,273 +0,0 @@
|
||||||
# ADR-0012: clay-borg is a design instrument, and §Underdetermined is the register
|
|
||||||
|
|
||||||
status: accepted
|
|
||||||
date: 2026-08-05
|
|
||||||
decided by: agent, under the standing loop authorization
|
|
||||||
tier: L (structural L — named a high-leverage pass by the maintainer and it
|
|
||||||
amends INTENT; chaos d8=6 → no override). Tier L requires a separate
|
|
||||||
survey, an adversarial review, then this document.
|
|
||||||
references: [CB-WP-0022](../workplans/CB-WP-0022-the-design-instrument.md),
|
|
||||||
[CB-RES-0007](../research/CB-RES-0007-design-instrument.md),
|
|
||||||
[challenge](../history/260805-design-instrument-challenge.md) /
|
|
||||||
[response](../history/260805-design-instrument-response.md),
|
|
||||||
[ADR-0007](ADR-0007-render-html-not-a-port.md) D3,
|
|
||||||
GROUND-WP-0002, GROUND-WP-0004 T02 (ground-game's shape ruling)
|
|
||||||
|
|
||||||
## Context
|
|
||||||
|
|
||||||
Nineteen passes of building a rigorous simulator produced findings about
|
|
||||||
*the game's rules* as a by-product, because an engine that refuses to
|
|
||||||
proceed past an undecided rule cannot help but locate the undecided rules.
|
|
||||||
Six such findings exist. They live in eleven files, no index, and one sat
|
|
||||||
unread in an inbox for four days.
|
|
||||||
|
|
||||||
**The review changed what this ADR has to decide.** The survey proposed a
|
|
||||||
new register governed by an existence rule — *a finding is not admissible
|
|
||||||
without its reproduction*. C1 falsified the example that rule was built
|
|
||||||
on, and C4 pointed out the incumbent was never evaluated. Both are
|
|
||||||
load-bearing here, so both are answered before anything is built.
|
|
||||||
|
|
||||||
## The correction this ADR is built on
|
|
||||||
|
|
||||||
`CB-WP-0022:61` justified the rule with *"GR-E01 is admissible because
|
|
||||||
4/6/9 against 5/7/9 is a computation anyone can rerun."* Commit `2da19a4`
|
|
||||||
had already rerun it: **6/9/12 against 5/7/9**, and the scenario was
|
|
||||||
renamed `-unreachable-` → `-reachable-`. The conclusion inverted.
|
|
||||||
|
|
||||||
That finding was **one of the two that passed the reproduction rule**. So:
|
|
||||||
|
|
||||||
| wrong premise | reached ground-game? | passed an existence test? |
|
|
||||||
|---|---|---|
|
|
||||||
| *"12 in the file"* — a sum with no deal table | yes | — |
|
|
||||||
| *"SOLVE offered on a face-down Problem"* — wrong condition named | yes | — |
|
|
||||||
| *"4/6/9, no dataset reconciles them"* | **stopped in flight, 2026-08-05** | **yes** |
|
|
||||||
|
|
||||||
**Existence is not the property that was missing.** Three findings went
|
|
||||||
wrong on their *description*, and the one that carried an artifact carried
|
|
||||||
a green one that no longer matched its prose. This ADR's rule is written
|
|
||||||
against that, not against the survey's.
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
## D1 — INTENT gains a fourth property: the engine is an instrument
|
|
||||||
|
|
||||||
`INTENT.md` states three properties — Clay, Borg, product-driven
|
|
||||||
evolution. It gains a fourth, and the wording is deliberately about a
|
|
||||||
**mechanism**, not an ambition:
|
|
||||||
|
|
||||||
> 4. **Instrument** — the engine is rigorous enough that it cannot proceed
|
|
||||||
> past a rule that does not decide. What it cannot execute, it reports:
|
|
||||||
> findings about the *game's* design are a product of building the
|
|
||||||
> simulator, not a side activity, and they are carried back to the game's
|
|
||||||
> owner with the artifact that produced them.
|
|
||||||
|
|
||||||
**Why this is earned rather than aspirational.** It is a restatement of
|
|
||||||
something that has already happened six times without anyone deciding it
|
|
||||||
should. The alternative wording considered and rejected — *"clay-borg is a
|
|
||||||
game design tool"* — claims a product we do not have and invites building
|
|
||||||
one; this claims a property we already exhibit and makes it a duty.
|
|
||||||
|
|
||||||
**Falsifier, and it is real:** if a pass ships with the engine having
|
|
||||||
tolerated an undecided rule by quietly picking a default and not raising
|
|
||||||
it, property 4 is false and INTENT is lying. `provisional: true` plus its
|
|
||||||
age report is what keeps it honest, and D2 puts the duty where that
|
|
||||||
machinery already lives.
|
|
||||||
|
|
||||||
## D2 — `§Underdetermined` is the register. Nothing parallel is built.
|
|
||||||
|
|
||||||
C4 lands. `specs/GroundRules.md §Underdetermined` was never evaluated as a
|
|
||||||
candidate, and against CB-RES-0007's own five benchmarks it already
|
|
||||||
delivers four:
|
|
||||||
|
|
||||||
| benchmark | §Underdetermined today |
|
|
||||||
|---|---|
|
|
||||||
| findability — one register with an index | **yes**, U1–U10 in one section |
|
|
||||||
| closure — the source changes, not an annotation | **yes**: *"a ruling flips the scenario, not the kernel"* (`:231-233`) |
|
|
||||||
| time raised → read | **owner and raise date per item**, `make coverage` warns past 30 days |
|
|
||||||
| taxonomy fits reality | untested either way |
|
|
||||||
| reproducibility — 100% carry a replayable artifact | **no. This is the gap.** |
|
|
||||||
|
|
||||||
The survey went to Magic to discover the Oracle property. We had written
|
|
||||||
it down ourselves eight days earlier, in our own words. **A second
|
|
||||||
register would immediately disagree with the first**, and this project has
|
|
||||||
a standing rule that a gate must cash out.
|
|
||||||
|
|
||||||
**Decision: extend, do not replace.** §Underdetermined becomes the finding
|
|
||||||
register for all five kinds in D4, not only for dataset ambiguities. Its
|
|
||||||
existing machinery — owner, raise date, `provisional: true`, the age
|
|
||||||
warning — is the mechanism, unchanged. What is added is D3's admissibility
|
|
||||||
rule, D4's `kind`, and D5's state.
|
|
||||||
|
|
||||||
**What this costs, stated plainly:** the register lives in a spec file
|
|
||||||
about GROUND's rules, so a finding about the *engine's* behaviour sits in
|
|
||||||
a document about the *game*. That is tolerable because every finding in
|
|
||||||
scope is about the game by definition (D7), and it is the price of not
|
|
||||||
running two registers that drift.
|
|
||||||
|
|
||||||
## D3 — admissibility: exists, has the ruled shape, and can fail
|
|
||||||
|
|
||||||
The survey's rule was one clause. It is now three, and **all three are
|
|
||||||
required**:
|
|
||||||
|
|
||||||
> **1. It exists.** A scenario, an arithmetic check that prints the
|
|
||||||
> contradiction, a recorded game the reader can replay, or a named test.
|
|
||||||
>
|
|
||||||
> **2. It has the ruled shape.** Per GROUND-WP-0004 T02, an arithmetic
|
|
||||||
> finding ships a **row-level table** — Surface and each hidden priority
|
|
||||||
> listed separately — **never a sum, never "deal depth N"**. Promoted here
|
|
||||||
> from the T04 addendum it was, because two of three wrong premises were
|
|
||||||
> quantities summed without a table.
|
|
||||||
>
|
|
||||||
> **3. It can fail.** The artifact must be capable of going red, and its
|
|
||||||
> current colour is recorded. A reproduction that has gone **green is not a
|
|
||||||
> closed finding** — it is a finding whose claim is now in question, and it
|
|
||||||
> must be resolved to `applied` or `withdrawn` by a human reading it.
|
|
||||||
|
|
||||||
Clause 3 is the one the review bought. GR-E01's scenario went green when
|
|
||||||
the edition landed, and the finding stayed admissible and stayed queued for
|
|
||||||
transmission, because nothing in the rule said a passing artifact was a
|
|
||||||
signal. **A green reproduction is an alarm, not a reassurance.**
|
|
||||||
|
|
||||||
**Falsifier for D3 as a whole:** if a finding reaches `ground-game` whose
|
|
||||||
prose describes a quantity its own artifact does not compute, D3 has
|
|
||||||
failed and the next revision needs the description generated *from* the
|
|
||||||
artifact rather than written beside it.
|
|
||||||
|
|
||||||
## D4 — taxonomy, grounded in the six we have
|
|
||||||
|
|
||||||
Five kinds, each with the existing finding that forces it. **No kind is
|
|
||||||
included that no finding of ours instantiates** — that is the whole
|
|
||||||
discipline, and T05's backfill is its test.
|
|
||||||
|
|
||||||
| kind | means | our instance |
|
|
||||||
|---|---|---|
|
|
||||||
| `underdetermined` | the rules do not say | U1–U10 |
|
|
||||||
| `inconsistent` | the rules disagree with each other or with the dataset | GR-E01 as raised (now `withdrawn`) |
|
|
||||||
| `inert` | a rule that cannot fire | SOLVE on a face-down Problem |
|
|
||||||
| `degenerate` | fires, but collapses play | GR-A13 wasted SOLVE |
|
|
||||||
| `unplayed` | implemented, never played to the end | GR-E03 / GR-E04 |
|
|
||||||
|
|
||||||
If backfill needs a sixth kind, **the taxonomy was invented rather than
|
|
||||||
observed** and T07 must say so.
|
|
||||||
|
|
||||||
## D5 — lifecycle: a queue that empties, a log that does not
|
|
||||||
|
|
||||||
```
|
|
||||||
raised ──► reported ──► ruled ──► applied
|
|
||||||
│
|
|
||||||
└──────────► withdrawn ◄──────────┘
|
|
||||||
```
|
|
||||||
|
|
||||||
| state | means | owner |
|
|
||||||
|---|---|---|
|
|
||||||
| `raised` | admissible under D3, in clay-borg | **clay-borg** |
|
|
||||||
| `reported` | delivered to ground-game as a file under their workplan | **clay-borg** |
|
|
||||||
| `ruled` | ground-game has answered | **ground-game** |
|
|
||||||
| `applied` | **the source changed** — dataset or rules text — and our provisional default was deleted | **clay-borg** |
|
|
||||||
| `withdrawn` | the finding was wrong, or dissolved | whoever finds it wrong |
|
|
||||||
|
|
||||||
`applied` means the source changed. It does **not** mean a ruling was
|
|
||||||
recorded — Magic's Oracle property, and CB-RES-0007 §2 is right about it.
|
|
||||||
|
|
||||||
**The queue is `raised` + `reported`; the log is `applied` + `withdrawn`.**
|
|
||||||
C6 asked whether keeping resolved findings contradicts "a queue that
|
|
||||||
empties." It does not: the open queue empties, the closed log accumulates
|
|
||||||
so a close rate can be computed. **Falsifier: if the default view shows
|
|
||||||
closed findings mixed with open ones, the queue property is gone** — so
|
|
||||||
`make design` reports the queue, and the log only on request.
|
|
||||||
|
|
||||||
**`withdrawn` is reported, not deleted.** GR-E01 is the reason: a claim
|
|
||||||
retracted silently is how the first three survived. When a finding is
|
|
||||||
withdrawn after being reported, the withdrawal goes downstream too.
|
|
||||||
|
|
||||||
## D6 — a finding without a reproduction is a `note`, and a note is not reportable
|
|
||||||
|
|
||||||
The survey left this open (§6) and the honest options were refusal or a
|
|
||||||
note tier. **Notes are admitted, under three constraints that make the
|
|
||||||
tier costly to sit in:**
|
|
||||||
|
|
||||||
1. A note **may not be reported to ground-game**. It has no `reported`
|
|
||||||
state. This is the whole point: the unread-inbox failure was prose
|
|
||||||
crossing a repo boundary.
|
|
||||||
2. A note **expires at 30 days** and `make design` flags it, reusing
|
|
||||||
§Underdetermined's existing age mechanism rather than adding one.
|
|
||||||
3. A note that is promoted acquires a reproduction meeting D3 — there is
|
|
||||||
no path from `note` to `reported` that skips it.
|
|
||||||
|
|
||||||
**Why not refusal.** A finding from *play* — the maintainer's *"I felt it
|
|
||||||
was too easy but then we lost"* — is real, has no cheap artifact, and is
|
|
||||||
the seed of CB-WP-0025. Refusing it would throw away the only class of
|
|
||||||
finding the engine cannot produce by itself. **Why not silence.** A note
|
|
||||||
that cannot expire ages into an apparent finding, which is the failure
|
|
||||||
this register exists to stop.
|
|
||||||
|
|
||||||
## D7 — no engine-evolution register, on a corrected inventory
|
|
||||||
|
|
||||||
The declaration's judgment survives; **its argument did not, and the
|
|
||||||
difference is recorded rather than smoothed over.** My inventory named
|
|
||||||
`evidence/`, `decisions/`, `gates.toml` and workplans. C5 found two places
|
|
||||||
it missed — `specs/InnerLoopReference.md`, 383 lines with sections headed
|
|
||||||
*"moved from InnerLoop v1.6"*, and 769 lines of `history/` retrospectives.
|
|
||||||
So the "third thing" the maintainer asked about **does exist and did find
|
|
||||||
homes**, somewhat by accident.
|
|
||||||
|
|
||||||
Decision: **no new register**, because a third place to look for material
|
|
||||||
currently in two makes it less findable, not more. This is argued and
|
|
||||||
narrowed, **not settled** — if InnerLoopReference keeps absorbing content
|
|
||||||
that is neither a decision nor a finding, that is the signal to revisit.
|
|
||||||
|
|
||||||
## D8 — `tools/design-baseline.py` is retired, not extended
|
|
||||||
|
|
||||||
C2 dismantled it: a hand-maintained dict counting itself (`:16-36`, `:89`),
|
|
||||||
a `has_reproduction` that never checks the file exists (`:38-43`) so its
|
|
||||||
own YES-control is green against a path `2da19a4` deleted (`:63`), date
|
|
||||||
literals for the 4-day figure (`:96-98`), and `Makefile:127` running only
|
|
||||||
`--self-test` so the reporting path has no CI at all.
|
|
||||||
|
|
||||||
**The 33% is withdrawn as a measured rate.** What survives is direction:
|
|
||||||
eleven files with no index and 0 of 10 U-items ruled are both checkable
|
|
||||||
without the tool.
|
|
||||||
|
|
||||||
`make design` (T05) reports **over the register**, which is real data with
|
|
||||||
real files behind it, and its reproduction check **stats the file**. The
|
|
||||||
baseline stays in the repo as the survey's dated snapshot, marked
|
|
||||||
superseded — deleting it would erase the evidence for how the number got
|
|
||||||
in.
|
|
||||||
|
|
||||||
## D9 — where a finding lives, and where it is delivered
|
|
||||||
|
|
||||||
The survey did not settle this. **Both, with distinct roles:**
|
|
||||||
|
|
||||||
- **clay-borg holds the finding and its reproduction**, in
|
|
||||||
§Underdetermined. The artifact cannot leave — it only runs here.
|
|
||||||
- **ground-game receives a generated report file under its own workplan**,
|
|
||||||
not an inbox message. Per CB-WP-0022 T06: *the message that sat unread
|
|
||||||
for four days is the baseline to beat, and the failure was not the
|
|
||||||
message — it was that nothing pointed at it.*
|
|
||||||
|
|
||||||
The report **cites the reproduction by path**, which is the half of
|
|
||||||
GROUND-WP-0004 T02's ruling that is theirs to keep, and makes the split
|
|
||||||
survivable: they read a claim, we hold the thing that proves it.
|
|
||||||
|
|
||||||
## Consequences
|
|
||||||
|
|
||||||
- `INTENT.md` gains property 4 (D1) — the change with the longest
|
|
||||||
half-life in this pass.
|
|
||||||
- `specs/GroundRules.md §Underdetermined` becomes the register and gains
|
|
||||||
`kind` and `state` per item (D2, D4, D5); **no new file, no new schema.**
|
|
||||||
- T04's spec inherits D3 as a **rule**, not an addendum, and must say what
|
|
||||||
a reproduction has to *show*.
|
|
||||||
- T05 builds `make design` over the register and backfills all six. If a
|
|
||||||
sixth kind is needed, the taxonomy was invented.
|
|
||||||
- T06 reports the queue **and GR-E01's withdrawal**.
|
|
||||||
- Notes carry a 30-day expiry (D6) on the existing age machinery.
|
|
||||||
|
|
||||||
## What was rejected
|
|
||||||
|
|
||||||
| rejected | why |
|
|
||||||
|---|---|
|
|
||||||
| a new register file with its own schema | C4 — the incumbent already delivers 4 of 5 benchmarks |
|
|
||||||
| the existence-only admissibility rule | C1 — its own showcase finding passed it and was false |
|
|
||||||
| refusing findings without reproductions | throws away findings from play, which is CB-WP-0025's entire input |
|
|
||||||
| an engine-evolution register | D7 — a third home for material already in two |
|
|
||||||
| extending `design-baseline.py` | C2 — it counts itself and its positive control is green against a deleted file |
|
|
||||||
|
|
@ -1,217 +0,0 @@
|
||||||
# CB-EV-0021 — the design instrument
|
|
||||||
|
|
||||||
CB-WP-0022 T07. Tier L, declaration 5 of chaos window 2 (structural L,
|
|
||||||
chaos d8=6 → no override). Closed 2026-08-05.
|
|
||||||
|
|
||||||
**Delivered:** INTENT property 4 (*Instrument*),
|
|
||||||
[ADR-0012](../decisions/ADR-0012-the-design-instrument.md) (nine
|
|
||||||
decisions), [specs/GameDesign.md](../specs/GameDesign.md) v1.0, the
|
|
||||||
register in [GroundRules.md](../specs/GroundRules.md) (14 rows),
|
|
||||||
`tools/design.py` + `make design`, and a report file in `ground-game`.
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
## 1. The headline: a separate reviewer found the survey's showcase was false
|
|
||||||
|
|
||||||
This is the first adversarial review in this repo **run by a genuinely
|
|
||||||
separate agent**. CB-RES-0006's reviewer opened by conceding it could not
|
|
||||||
be, and wrote:
|
|
||||||
|
|
||||||
> *"Treat the challenges below as a lower bound on what a genuinely
|
|
||||||
> separate reviewer would find."*
|
|
||||||
|
|
||||||
That was a testable claim. **This is the test.** The separate reviewer ran
|
|
||||||
`git log` against the survey's central example and found our own commit
|
|
||||||
had falsified it four days earlier — while the author, who wrote that
|
|
||||||
commit, quoted the dead number twice in the same workplan.
|
|
||||||
|
|
||||||
`CB-WP-0022:61` justified the whole proposal with *"GR-E01 is admissible
|
|
||||||
because 4/6/9 against 5/7/9 is a computation anyone can rerun."*
|
|
||||||
`2da19a4` had rerun it: **6/9/12**, scenario renamed `-unreachable-` →
|
|
||||||
`-reachable-`, conclusion inverted. And GR-E01 was **one of the two
|
|
||||||
findings that passed the reproduction rule**.
|
|
||||||
|
|
||||||
**Verdict on same-session review: it is not a weak form of the control, it
|
|
||||||
is a different control.** Four of the seven challenges rest on facts a
|
|
||||||
same-session reviewer had already seen and not questioned. The cost was
|
|
||||||
~70k subagent tokens; it caught a defect in flight (§2) and rewrote the
|
|
||||||
pass's central rule.
|
|
||||||
|
|
||||||
## 2. What it caught in flight
|
|
||||||
|
|
||||||
T06's payload, status `todo`, still read *"the deal count puts 4/6/9
|
|
||||||
points in play … no dataset reconciles them"* and was queued for
|
|
||||||
transmission to `ground-game`.
|
|
||||||
|
|
||||||
It would have been the **fourth** wrong premise sent downstream, after
|
|
||||||
*"12 in the file"* (a sum with no deal table) and *"SOLVE offered on a
|
|
||||||
face-down Problem"* (the wrong condition named). **It is the only one
|
|
||||||
stopped before sending.**
|
|
||||||
|
|
||||||
| wrong premise | reached ground-game? | passed an existence test? |
|
|
||||||
|---|---|---|
|
|
||||||
| *"12 in the file"* | yes | — |
|
|
||||||
| *"SOLVE on a face-down Problem"* | yes | — |
|
|
||||||
| *"4/6/9, no dataset reconciles them"* | **stopped 2026-08-05** | **yes** |
|
|
||||||
|
|
||||||
## 3. The rule the pass came to build was one clause short
|
|
||||||
|
|
||||||
The declaration's load-bearing rule was an **existence** test. Three of
|
|
||||||
three wrong premises would have passed it, and one did.
|
|
||||||
|
|
||||||
What shipped instead (GameDesign §1): a reproduction must **exist**, have
|
|
||||||
the **shape ground-game ruled** (row-level deal table, never a sum —
|
|
||||||
promoted from a T04 addendum to part of the rule), and **be able to
|
|
||||||
fail**.
|
|
||||||
|
|
||||||
**Clause 3 is the one this pass bought.** GR-E01's scenario went green
|
|
||||||
when the edition landed and the finding stayed admissible and stayed
|
|
||||||
queued, because nothing treated a passing artifact as a signal. *A green
|
|
||||||
reproduction is an alarm, not a reassurance.*
|
|
||||||
|
|
||||||
## 4. Building it produced a distinction the ADR does not have
|
|
||||||
|
|
||||||
`make design`'s first run alarmed on **U2** — and was wrong to.
|
|
||||||
|
|
||||||
U2's scenario is green **because the provisional default it documents is
|
|
||||||
implemented**, which says nothing about whether ground-game agrees.
|
|
||||||
GR-E01's was a *counterexample* that went green. **Same colour, opposite
|
|
||||||
meaning.**
|
|
||||||
|
|
||||||
So the register gained a `role` column — `counterexample` | `default` —
|
|
||||||
and only a green counterexample alarms. **A register that cannot tell them
|
|
||||||
apart either alarms constantly or never**, and this was invisible until
|
|
||||||
the metric ran against real rows. Folded back into GameDesign §1.3.
|
|
||||||
|
|
||||||
## 5. Did backfilling change the design?
|
|
||||||
|
|
||||||
**T07 asked to be suspicious if all six findings fit the first taxonomy.**
|
|
||||||
They did — five kinds absorbed everything, no sixth invented. Two reasons
|
|
||||||
that is less reassuring than it looks and one reason it is fine:
|
|
||||||
|
|
||||||
- the taxonomy was **derived from these six**, so fitting them is nearly
|
|
||||||
circular. Its real test is the seventh finding;
|
|
||||||
- **one of the six was deleted during backfill.** *"Six provisional
|
|
||||||
defaults"* was not entered: C3 showed it double-counted GR-E01, and
|
|
||||||
provisional scenarios are *reproductions for* underdetermined items, not
|
|
||||||
a finding. So the taxonomy absorbed five, not six;
|
|
||||||
- what it did do is **force the note tier to earn itself** — GR-A13 and
|
|
||||||
GR-E03/E04 have no artifact and became notes, which is the first time
|
|
||||||
the distinction cost anything.
|
|
||||||
|
|
||||||
## 6. The numbers, and why none of the survey's survive
|
|
||||||
|
|
||||||
**CB-RES-0007's measurements are withdrawn.** C2 dismantled the
|
|
||||||
instrument: `design-baseline.py` is a hand-maintained dict where
|
|
||||||
`n = len(FINDINGS)`; `has_reproduction` is `startswith("scenarios/")` and
|
|
||||||
never stats the file, so its own YES-control was **green against a path
|
|
||||||
`2da19a4` deleted**; the 4-day figure is two date literals subtracted
|
|
||||||
(mutated to 2020-01-01 → 2406 days, all checks still green); and
|
|
||||||
`Makefile:127` ran only `--self-test`, so the reporting path had no CI.
|
|
||||||
|
|
||||||
It is **superseded, not deleted** — it is the evidence for how a wrong
|
|
||||||
number got into a survey, and the header now says so.
|
|
||||||
|
|
||||||
Measured by `make design`, 2026-08-05, over rows that name real files:
|
|
||||||
|
|
||||||
```
|
|
||||||
findings 12 (+2 notes)
|
|
||||||
with a resolving reproduction 3/12 = 25% target 100%
|
|
||||||
open, lacking a reproduction 9 target 0
|
|
||||||
reproductions green while open 0 target 0
|
|
||||||
notes past 30 days 0 target 0
|
|
||||||
closed (log) 2 [F11, F13]
|
|
||||||
```
|
|
||||||
|
|
||||||
**Two survey claims died against measurement:**
|
|
||||||
|
|
||||||
| survey said | measured |
|
|
||||||
|---|---|
|
|
||||||
| *"six of the ten already have provisional scenarios"* | **one.** `grep -lE "\bU<n>\b" scenarios/ground/*.yaml` → only U2 |
|
|
||||||
| *"2/6 = 33% reproduced"* | withdrawn; GR-E01 was double-counted and the detector was broken |
|
|
||||||
|
|
||||||
The first is the same defect class as the wrong premises — **an
|
|
||||||
uncheckable link asserted as a fact** — found inside the survey that
|
|
||||||
proposed the fix.
|
|
||||||
|
|
||||||
## 7. The register's first run found ten answers nobody had collected
|
|
||||||
|
|
||||||
The most useful thing `make design` did was not a metric.
|
|
||||||
|
|
||||||
**GROUND-WP-0002 T05 ruled all ten U-items on 2026-08-03 — every one
|
|
||||||
confirmed.** T03 confirmed five of the six provisional scenarios. The
|
|
||||||
workplan is `finished`.
|
|
||||||
|
|
||||||
CB-RES-0007 reported *"0 of 10 ruled"* on 2026-08-03, and the register
|
|
||||||
built this session initially recorded them as `reported`. **Both were two
|
|
||||||
days stale on the day they were written.**
|
|
||||||
|
|
||||||
**This is the unread-inbox failure running in the opposite direction.**
|
|
||||||
The pass was designed around *we send findings and nobody reads them*. The
|
|
||||||
symmetric failure — *they answer and we do not collect it* — was not
|
|
||||||
anticipated by the declaration, the survey, or the review, and the first
|
|
||||||
run of the instrument surfaced it immediately.
|
|
||||||
|
|
||||||
The items are `ruled`, not `applied`: the rulings confirmed our defaults,
|
|
||||||
so the rules did not move, but the scenarios still carry `provisional:
|
|
||||||
true` for settled choices. **Lifting those flags is owed and is not done**
|
|
||||||
— `make design` shows them open until it is, which is the register
|
|
||||||
working.
|
|
||||||
|
|
||||||
## 8. The engine-evolution question, as the review left it
|
|
||||||
|
|
||||||
Declared judgment: no second register, because engine evolution already
|
|
||||||
lives in `evidence/`, `decisions/`, `gates.toml` and workplans.
|
|
||||||
|
|
||||||
**The conclusion survives; the argument did not.** C5 found the "third
|
|
||||||
thing" the maintainer meant is visible in `specs/InnerLoopReference.md`
|
|
||||||
(383 lines, two sections literally headed *"moved from InnerLoop v1.6"*)
|
|
||||||
and 769 lines of `history/` retrospectives — **neither of which the
|
|
||||||
inventory named**. So the content exists and found homes somewhat by
|
|
||||||
accident.
|
|
||||||
|
|
||||||
ADR-0012 D7 records it as **argued and narrowed, not settled**, with the
|
|
||||||
revisit condition: if InnerLoopReference keeps absorbing material that is
|
|
||||||
neither a decision nor a finding, build the register.
|
|
||||||
|
|
||||||
## 9. What tier L cost against what it caught
|
|
||||||
|
|
||||||
Second full-weight L pass in this window, third in the project.
|
|
||||||
|
|
||||||
| caught | by |
|
|
||||||
|---|---|
|
|
||||||
| the central example was false | **only** the separate reviewer |
|
|
||||||
| a wrong premise in flight to ground-game | the same |
|
|
||||||
| the baseline instrument counts itself | the same |
|
|
||||||
| the incumbent register was never evaluated | the same |
|
|
||||||
| `role`: green default ≠ green counterexample | **building it**, not reviewing it |
|
|
||||||
| ten uncollected rulings | **running it**, not building it |
|
|
||||||
|
|
||||||
**Two of six were caught by execution, not by process** — the pattern
|
|
||||||
InnerLoop §Design goal predicts (cheap correction beats exhaustive
|
|
||||||
prevention), and an argument against front-loading more review rather than
|
|
||||||
less.
|
|
||||||
|
|
||||||
**The one number that matters for the loop's own calibration:** the review
|
|
||||||
step is what distinguished this pass from a competent one. Without it the
|
|
||||||
pass ships a rule justified by a false example, sends a fourth bad premise
|
|
||||||
downstream, and builds `make design` on a dict that counts itself. The
|
|
||||||
correct reading of CB-RES-0006's fidelity note is **not** *"our reviews
|
|
||||||
were somewhat weaker"* — it is *"our reviews were not performing this
|
|
||||||
function at all."*
|
|
||||||
|
|
||||||
## 10. Cost
|
|
||||||
|
|
||||||
CB-WP-0021's cost, by re-running the instrument: see `make cost` /
|
|
||||||
`tools/cb-cost.py`. **Not quoted here as a literal** — §Single source of
|
|
||||||
fact, and this pass has just spent itself proving what typed numbers do.
|
|
||||||
|
|
||||||
## Open after this pass
|
|
||||||
|
|
||||||
- **U1–U10 → `applied`**: lift `provisional: true` where a ruling confirmed
|
|
||||||
the default, and record each ruling. Nine items of bookkeeping.
|
|
||||||
- **The U-item ↔ scenario mapping is unwritten** for 9 of 10. Cheap, and
|
|
||||||
it is what makes the register checkable.
|
|
||||||
- **F12 and F14 are notes** and expire 2026-08-31. GROUND-WP-0003's
|
|
||||||
playtest closes F14 and produces its artifact via `cb-play --record`.
|
|
||||||
- **The taxonomy's real test is the seventh finding**, not these six.
|
|
||||||
|
|
@ -1,361 +0,0 @@
|
||||||
# 260805 — challenge to CB-RES-0007
|
|
||||||
|
|
||||||
Adversarial review, one round, per InnerLoop §Step 2. Target: the survey
|
|
||||||
`research/CB-RES-0007-design-instrument.md`, the harness
|
|
||||||
`tools/design-baseline.py`, and the §judgment in `CB-WP-0022`.
|
|
||||||
|
|
||||||
**Fidelity note, first.** Unlike `260802-render-port-challenge.md`, this
|
|
||||||
review *was* run in a separate agent session with no access to the author's
|
|
||||||
reasoning trail — only the files. So it has the property §Step 2 asks for.
|
|
||||||
It does **not** have a different sample: there is exactly one baseline
|
|
||||||
harness and one repo, so every number below was re-derived from the
|
|
||||||
author's instrument. Per §Step 2 that is stated rather than reported as a
|
|
||||||
clean verify. What this review substitutes for a second sample is
|
|
||||||
**mutation**: each quoted number was traced to the assertion that produces
|
|
||||||
it and the assertion was inverted. Four of the five headline numbers do not
|
|
||||||
move when the world they claim to measure moves.
|
|
||||||
|
|
||||||
Ranked. C1 and C2 land hardest. C6 and C7 are marked weak — do not spend
|
|
||||||
the one response round on them.
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
## C1 — The rule's own showcase example is a false finding that shipped a green reproduction
|
|
||||||
|
|
||||||
**This is the challenge that lands, and it is worse than the two cases the
|
|
||||||
task description already knows about.**
|
|
||||||
|
|
||||||
`CB-WP-0022:61` states the admissibility rule and then gives its canonical
|
|
||||||
example:
|
|
||||||
|
|
||||||
> *"**GR-E01 is admissible because 4/6/9 against 5/7/9 is a computation
|
|
||||||
> anyone can rerun**; the SOLVE inertness is admissible because a recorded
|
|
||||||
> session shows three no-ops."*
|
|
||||||
|
|
||||||
Both halves of that sentence are now known to be wrong.
|
|
||||||
|
|
||||||
- The SOLVE half: `CB-WP-0023:87` — *"the reported case was not the one I
|
|
||||||
reported… `validate` already rejected face-down, so it never was
|
|
||||||
offered."* Already conceded by the author.
|
|
||||||
- **The GR-E01 half has not been conceded, and it is the more damaging
|
|
||||||
one.** `4/6/9` is stale. Commit `2da19a4` (CB-WP-0021, 2026-08-04)
|
|
||||||
measured **6/9/12** against thresholds 5/7/9, `ADR-0011:110` records the
|
|
||||||
ruled 6/9/12, and the scenario was renamed
|
|
||||||
`gr-e01-threshold-**unreachable**-2p.yaml` →
|
|
||||||
`gr-e01-threshold-**reachable**-2p.yaml`. The finding did not get
|
|
||||||
refined; **its conclusion inverted**.
|
|
||||||
|
|
||||||
Now the part that breaks the proposal's central claim. GR-E01 is one of the
|
|
||||||
**two** findings in the survey's `2/6 = 33% with a runnable reproduction`
|
|
||||||
(`design-baseline.py:26-29`). It is a *success case* of the reproduction
|
|
||||||
rule. It carried a scenario. The scenario was green. And the scenario
|
|
||||||
asserted the wrong conclusion for four days, was reported to ground-game on
|
|
||||||
that basis, and was renamed only when someone re-did the arithmetic by
|
|
||||||
hand.
|
|
||||||
|
|
||||||
So the count is not two bad premises. **It is three, and the third one is
|
|
||||||
the exemplar the workplan uses to define what a good finding looks like.**
|
|
||||||
|
|
||||||
The survey's §3 argument — *"in formal methods a claim without a trace is
|
|
||||||
not a result, it is a hypothesis"* — is true and irrelevant here. A model
|
|
||||||
checker's counterexample is generated *by* the checker from the property;
|
|
||||||
it cannot disagree with the property, because it is the property's witness.
|
|
||||||
Our scenarios are **hand-authored assertions of an expected value**. A
|
|
||||||
hand-written scenario that encodes a wrong expectation is a trace that
|
|
||||||
witnesses the author's belief, not the system's behaviour. §3 imports the
|
|
||||||
prestige of counterexample traces onto an artifact class that does not have
|
|
||||||
their guarantee, and it does so in the section that supplies the whole
|
|
||||||
justification for the rule.
|
|
||||||
|
|
||||||
**Required:** the survey cannot claim the reproduction rule is the load-
|
|
||||||
bearing control while its own two positive instances of that rule both
|
|
||||||
described the wrong thing. Either (a) name the *additional* control that
|
|
||||||
would have caught GR-E01 — the ground-game row-level deal table
|
|
||||||
requirement is the obvious candidate and it is a **shape** requirement, not
|
|
||||||
an existence one — and demote "carries its reproduction" from *the* rule to
|
|
||||||
*half* the rule, or (b) withdraw §3's claim that a reproduction is
|
|
||||||
sufficient to make a finding admissible.
|
|
||||||
|
|
||||||
Secondary, and checkable: `4/6/9` is still live in three places —
|
|
||||||
`workplans/CB-WP-0022:61`, `workplans/CB-WP-0022:283` (task T06, still
|
|
||||||
`todo`, still planning to report the stale number to ground-game), and
|
|
||||||
`evidence/CB-EV-0018:153`. **The pass that exists to stop wrong premises
|
|
||||||
reaching ground-game currently has a wrong premise queued for transmission
|
|
||||||
to ground-game.** `grep -rn "4/6/9" --include=*.md .` settles it.
|
|
||||||
|
|
||||||
## C2 — The baseline is a hand-maintained list reporting on itself, and four of its five numbers are mutation-proof
|
|
||||||
|
|
||||||
The survey's §1 heading is *"The baseline is us, and **it is measured**"*.
|
|
||||||
It is not measured. It is a dict literal at `tools/design-baseline.py:16-36`
|
|
||||||
with one globbed row, printed back with arithmetic.
|
|
||||||
|
|
||||||
Per §Step 2 (v1.4), each number was traced to its assertion and mutated.
|
|
||||||
|
|
||||||
**(a) `findings 6`** is `len(FINDINGS)` (`:89`) — the number of keys typed
|
|
||||||
into a dict at `:16-36`. It is a count of hardcoded rows, not of findings.
|
|
||||||
The self-test that "pins" it (`:58`) asserts `len(FINDINGS) >= 6` against
|
|
||||||
the same literal. Nothing in the repo can change this number; only editing
|
|
||||||
line 16-36 can.
|
|
||||||
|
|
||||||
**(b) `2/6 = 33%`** rests on `has_reproduction()` (`:38-43`), which is a
|
|
||||||
**string prefix test on `p.startswith("scenarios/")`**. It never touches
|
|
||||||
the filesystem. Mutation:
|
|
||||||
|
|
||||||
```
|
|
||||||
has_reproduction(['scenarios/COMPLETELY-MADE-UP.yaml']) -> True
|
|
||||||
has_reproduction(['scenarios/']) -> True
|
|
||||||
has_reproduction(['crates/cb-engine/tests/foo.rs']) -> False
|
|
||||||
```
|
|
||||||
|
|
||||||
Two consequences. First, the self-test's YES-control at `:63` cites
|
|
||||||
`scenarios/ground/gr-e01-threshold-unreachable-2p.yaml` — **a file that no
|
|
||||||
longer exists** (renamed in `2da19a4`; `ls scenarios/ground/ | grep e01`).
|
|
||||||
It passes anyway, because the detector never looks. This is exactly the
|
|
||||||
render-port C6 class: a control that cannot distinguish "found a
|
|
||||||
reproduction" from "found a string". The `every cited location exists`
|
|
||||||
check at `:57` iterates `FINDINGS.values()` only and never sees the
|
|
||||||
self-test's own literal, so the one guard that would have caught it is
|
|
||||||
pointed at the wrong list.
|
|
||||||
|
|
||||||
Second, the detector **contradicts the rule it is measuring**. `CB-WP-0022:55-56`
|
|
||||||
admits *"a scenario that fails, an arithmetic check that prints the
|
|
||||||
contradiction, a recorded game the reader can replay, or **a named test**."*
|
|
||||||
`has_reproduction` recognises exactly one of those four. A finding backed by
|
|
||||||
a named Rust test scores 0. The 33% is therefore not the baseline for the
|
|
||||||
proposed rule; it is the baseline for a narrower rule nobody proposed, and
|
|
||||||
it is biased **downward**, which flatters the "nobody was ever asked for
|
|
||||||
one" story in §1.
|
|
||||||
|
|
||||||
**(c) `4 days`** is `datetime.date(2026, 8, 3) - datetime.date(2026, 7, 30)`
|
|
||||||
— two literals on `:96-97`, subtracted on `:98`. No data source. Mutated
|
|
||||||
`:96` to `2020-01-01`:
|
|
||||||
|
|
||||||
```
|
|
||||||
U1..U10: raised 2020-01-01, first READ 2026-08-03 — 2406 days
|
|
||||||
--self-test: all 5 checks [ok]
|
|
||||||
```
|
|
||||||
|
|
||||||
The headline number can be set to anything and every control stays green.
|
|
||||||
Worse, **it disagrees with the repo**: `specs/GroundRules.md:224` states
|
|
||||||
*"Owner: ground-game. **Raised: 2026-07-31**"*, and
|
|
||||||
`scenarios/ground/gr-a04-bond-support.yaml:9` carries
|
|
||||||
`provisional_raised: 2026-07-31`. Under the spec's own date the interval is
|
|
||||||
**3 days, not 4**. Two sources of the same fact, disagreeing, with zero
|
|
||||||
`<!-- fact: -->` tags in either the survey or the workplan
|
|
||||||
(`grep -c "fact:"` returns 0 for both) — the drift surface InnerLoop
|
|
||||||
§Single-source-of-fact names explicitly.
|
|
||||||
|
|
||||||
`4 days` is also the row the survey's §5 calls *"the number this exists to
|
|
||||||
fix; any mechanism that does not move it has failed"*. It is the least
|
|
||||||
defensible number in the document.
|
|
||||||
|
|
||||||
**(d) `11 files`** is `len(places)` — the deduplicated union of the
|
|
||||||
hardcoded paths. It is derived, so it is the healthiest of the four, but it
|
|
||||||
measures **citation sites the author typed**, not files holding findings.
|
|
||||||
Add a citation, the "findability" benchmark worsens; nothing about the
|
|
||||||
repo changed.
|
|
||||||
|
|
||||||
**(e)** The one number with real provenance is the 5-element glob at
|
|
||||||
`:30-34` — and see C3, it disagrees with the label above it.
|
|
||||||
|
|
||||||
**What the harness would report if the work silently stopped** (§Step 2
|
|
||||||
requires this be stated): the reporting path, `:80-99`, performs **no
|
|
||||||
existence check at all**. Delete every file in `FINDINGS` and it still
|
|
||||||
prints `6 findings, 2/6 = 33%, 11 files, 4 days`. The existence guard lives
|
|
||||||
only in `--self-test`, and `Makefile:127` runs **only** `--self-test` —
|
|
||||||
the numbers quoted in the survey come from a code path CI never executes.
|
|
||||||
|
|
||||||
**Required:** either the survey stops calling §1 "measured" and labels the
|
|
||||||
baseline what it is — a hand-curated snapshot, which is a legitimate thing
|
|
||||||
to have but caps at the same directional status the survey correctly
|
|
||||||
assigns to the external rows — or the harness derives its finding set from
|
|
||||||
the repo (the `provisional:` marks, `evidence/` frontmatter, the U-list)
|
|
||||||
so that mutating the repo moves the number.
|
|
||||||
|
|
||||||
## C3 — "six provisional defaults" is five, and the survey's own instrument says so
|
|
||||||
|
|
||||||
`research/CB-RES-0007:46` lists `six provisional defaults | 6 | yes`.
|
|
||||||
|
|
||||||
```
|
|
||||||
$ python3 tools/design-baseline.py
|
|
||||||
repro 5 location(s) six provisional defaults
|
|
||||||
$ make coverage
|
|
||||||
provisional U-item defaults: 5
|
|
||||||
```
|
|
||||||
|
|
||||||
Six became five in `2da19a4` when `gr-e01-threshold-unreachable-2p.yaml`
|
|
||||||
lost its `provisional: true` (verified:
|
|
||||||
`git grep -l "provisional: true" 2da19a4^ -- scenarios/ground/ | wc -l` →
|
|
||||||
**6**; at HEAD → **5**). The author edited `design-baseline.py:22-29` in
|
|
||||||
that same window to add the `[RESOLVED]` comment on the GR-E01 row, and did
|
|
||||||
not notice that the *next* row's glob had silently dropped a member and its
|
|
||||||
hardcoded label `"six provisional defaults"` had gone stale.
|
|
||||||
|
|
||||||
Two things follow.
|
|
||||||
|
|
||||||
1. **The survey's table row 5 (`locations: 6`) is a number the survey's own
|
|
||||||
tool contradicts on stdout**, and no self-test compares the word "six"
|
|
||||||
to `len()`. §4 of the survey argues the register *"must reuse [the
|
|
||||||
provisional machinery], not compete with it"* — and the very row that
|
|
||||||
reads the provisional machinery is the one that drifted from it.
|
|
||||||
2. **GR-E01 was double-counted.** It is row 4 *and* was a member of row 5's
|
|
||||||
six. The denominator of `2/6` contained the same finding twice, and one
|
|
||||||
of the two numerator hits was that finding.
|
|
||||||
|
|
||||||
**Required:** correct the row, and state whether the `2/6` survives the
|
|
||||||
double-count. If rows 4 and 5 overlap, the honest denominator at survey
|
|
||||||
time was 5 distinct findings, and the headline rate changes.
|
|
||||||
|
|
||||||
## C4 — `GroundRules.md §Underdetermined` was never evaluated as the incumbent, and it already implements four of the five benchmarks
|
|
||||||
|
|
||||||
T02 requires an attempt at *"whether a register is needed at all, rather
|
|
||||||
than one more section in `GroundRules.md §Underdetermined`."* The survey
|
|
||||||
does not attempt it. §4 mentions the provisional machinery only as
|
|
||||||
*machinery to reuse*; §5's benchmark table has a `today` column that never
|
|
||||||
names §Underdetermined; and §6's "did not settle" list does not include it.
|
|
||||||
Per §Step 2's row for surveys, **that is an omitted candidate, and it is
|
|
||||||
the strongest one, because it is the incumbent.**
|
|
||||||
|
|
||||||
Read against the survey's own five benchmarks (`specs/GroundRules.md:222-277`):
|
|
||||||
|
|
||||||
| benchmark (§5) | §Underdetermined today |
|
|
||||||
|---|---|
|
|
||||||
| findability: *one register with an index* | **has it** — one section, ten numbered items, `grep U7` |
|
|
||||||
| reproducibility: *runnable artifact* | **has the obligation** — *"every U-item's default must be covered by a scenario tagged `provisional: true`"* (`:274-277`) |
|
|
||||||
| closure: *closes when the source changed* | **has it** — *"A ruling flips the **scenario**, not the kernel; a correction is a data change"* (`:231-233`). This is the Magic Oracle property the survey went to Magic to discover. It has been in the spec since 2026-07-31. |
|
|
||||||
| raised → read, with age | **has it, executable** — owner + raise date per item, `make coverage` prints `age=5d` per item and warns past 30 days |
|
|
||||||
| taxonomy fits reality | not addressed — genuinely absent |
|
|
||||||
|
|
||||||
The survey travels to Magic: the Gathering to extract *"a finding closes
|
|
||||||
when the source changes, not when an annotation is added"* and presents it
|
|
||||||
as a corrected assumption (§2, and `CB-WP-0022:136-139` treats it as the
|
|
||||||
pass's main external yield). **The repo already states that property in
|
|
||||||
those terms, one directory over.** That is not a small omission: it is the
|
|
||||||
survey's headline external finding being rediscovered from the incumbent it
|
|
||||||
failed to survey.
|
|
||||||
|
|
||||||
**Where this challenge stops, honestly.** §Underdetermined has a real gap
|
|
||||||
and the review found it: it obliges ten U-items to carry provisional
|
|
||||||
scenarios and **five exist**. `make coverage` prints `provisional U-item
|
|
||||||
defaults: 5` and does *not* flag the five missing — it reports what exists,
|
|
||||||
never what is owed. That is a genuine argument for something with a
|
|
||||||
denominator, and it is stronger than any argument the survey actually
|
|
||||||
makes. It is also an argument for *one gate in `rule-coverage.py`*, not for
|
|
||||||
a new register.
|
|
||||||
|
|
||||||
**Required:** evaluate §Underdetermined as a candidate row with a `today`
|
|
||||||
column, and state what a register does that §Underdetermined + one
|
|
||||||
completeness check cannot. If the answer is "holds findings that are not
|
|
||||||
U-items" — say which of the six, and note that four of the six are
|
|
||||||
single-file prose notes in `evidence/`.
|
|
||||||
|
|
||||||
## C5 — The engine-evolution "third thing" exists, and its homelessness is visible in the filesystem
|
|
||||||
|
|
||||||
`CB-WP-0022:73-84` argues the engine-evolution register is redundant
|
|
||||||
because *"`evidence/CB-EV-*` records what each pass found, `decisions/ADR-*`
|
|
||||||
records what was decided…, `gates.toml` records what every control has
|
|
||||||
caught…, and workplans record what was attempted."*
|
|
||||||
|
|
||||||
The list is incomplete, and the omission is the whole question. The third
|
|
||||||
thing — *what we learned about building engines* — is
|
|
||||||
**`specs/InnerLoopReference.md`, 383 lines**, and its section list is
|
|
||||||
nothing but that:
|
|
||||||
|
|
||||||
```
|
|
||||||
§ Implementation rules the first pass earned
|
|
||||||
§ Why review is adversarial execution, not re-derivation
|
|
||||||
§ Numbers you inherit — two cases from CB-WP-0002
|
|
||||||
§ Quoting a cost — six for six, never high
|
|
||||||
§ Chaos roll — window 1's verdict and the d4 → d8 change
|
|
||||||
§ Retargeting — the case, and the four corrections it must not forbid
|
|
||||||
```
|
|
||||||
|
|
||||||
None of those is a decision (ADR), a per-pass finding (evidence), or a
|
|
||||||
control (gates.toml). They are cross-pass lessons about method. Plus six
|
|
||||||
`history/*-retrospective.md` files, 769 lines, in the same category.
|
|
||||||
|
|
||||||
**And the homelessness is on the record in the headings themselves.** Two
|
|
||||||
sections are literally titled *"(moved from InnerLoop v1.6)"* — the content
|
|
||||||
grew inside the normative spec until it had to be evicted into an
|
|
||||||
appendix-shaped file that is not in the author's list of homes. That is
|
|
||||||
precisely the signature T02 asked the reviewer to look for: *content that
|
|
||||||
had nowhere to go*.
|
|
||||||
|
|
||||||
**Where this challenge stops.** It does not establish that a *register* is
|
|
||||||
needed. It establishes that the author's redundancy argument is made
|
|
||||||
against the wrong inventory. `InnerLoopReference.md` and `history/` **are**
|
|
||||||
a home, so the conclusion "no second register" probably survives — but it
|
|
||||||
survives for a different reason than the one given, and the difference
|
|
||||||
matters for T03: the asymmetry claimed at `:83` (*"engine evolution has a
|
|
||||||
home and game design does not"*) is real, but the home is an unindexed
|
|
||||||
383-line spec appendix plus six retrospectives, i.e. **exactly the
|
|
||||||
`11 files, no index` condition the survey calls a defect when it is game
|
|
||||||
findings**. The survey applies one standard to game findings and another to
|
|
||||||
engine lessons and does not state the distinction.
|
|
||||||
|
|
||||||
**Required:** name `InnerLoopReference.md` and `history/*-retrospective.md`
|
|
||||||
in the redundancy argument, and say why "several files, no index" is
|
|
||||||
disqualifying for one subject and acceptable for the other.
|
|
||||||
|
|
||||||
## C6 — (weak) The survey's own §2 conclusion contradicts the harness's stated design
|
|
||||||
|
|
||||||
§2: *"A register that accumulates permanent rulings is a system that has
|
|
||||||
given up on fixing its rules. Ours should be a **queue that empties**, not
|
|
||||||
an archive that grows."*
|
|
||||||
|
|
||||||
`design-baseline.py:23-25`: *"Kept in the baseline because… a register that
|
|
||||||
drops findings when they close cannot report a close rate."*
|
|
||||||
|
|
||||||
Both are defensible and they are not strictly contradictory — a queue can
|
|
||||||
retain closed items out-of-band. But T03 has to pick, and the two artifacts
|
|
||||||
of this pass currently point opposite ways, with the tool's comment being
|
|
||||||
the more considered of the two. Marked **weak**: this is a clarification,
|
|
||||||
not a defect.
|
|
||||||
|
|
||||||
## C7 — (weak) `n=6`, all engine-surfaced, and §6 already concedes it
|
|
||||||
|
|
||||||
§6 concedes the sample is six and all engine-surfaced, and flags the
|
|
||||||
play-sourced finding (*"DARVO feels punishing at 2 players"*) as unsettled.
|
|
||||||
That concession is correct and I have nothing to add to it, so this is not
|
|
||||||
a challenge — it is noted only so the response does not spend a round
|
|
||||||
defending something already conceded. Marked **weak**.
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
## Verdict
|
|
||||||
|
|
||||||
**Not approvable as written.** C1, C2 and C3 each require a change to the
|
|
||||||
survey, not a clarification. C4 requires an added candidate row. C5
|
|
||||||
requires a corrected inventory in the workplan's §judgment.
|
|
||||||
|
|
||||||
| # | verdict |
|
|
||||||
|---|---|
|
|
||||||
| **C1** | **lands hardest.** Requires demoting "carries its reproduction" from sufficient to necessary-but-not-sufficient, and adding the shape requirement. GR-E01 is a rule-satisfying finding that was wrong. |
|
|
||||||
| **C2** | **lands.** Four of five headline numbers do not move under mutation; `4 days` is a literal and disagrees with `GroundRules.md:224` (3 days). §1's word "measured" must go or the harness must derive. |
|
|
||||||
| **C3** | **lands.** "six" is five; the survey's table contradicts its own tool's stdout; GR-E01 double-counted in the `2/6` denominator. |
|
|
||||||
| **C4** | **lands, moderate.** The incumbent was never surveyed, and it already holds the Magic Oracle property the survey travelled to Magic to find. Weakened by a real gap the review found for the author: 10 U-items owe scenarios, 5 exist. |
|
|
||||||
| **C5** | **partial.** The third thing exists and is visibly homeless; the redundancy *conclusion* probably survives, the *argument* does not. |
|
|
||||||
| **C6** | **weak.** Clarification for T03. |
|
|
||||||
| **C7** | **weak / already conceded.** |
|
|
||||||
|
|
||||||
**What survives.** Two claims attacked and not broken:
|
|
||||||
|
|
||||||
- **That the reproduction rule is affordable.** §1's four cheap
|
|
||||||
reproductions were checked and the estimates are plausible; `bot::play`
|
|
||||||
exists, the scenario harness exists, and the SOLVE case in fact got its
|
|
||||||
reproduction in CB-WP-0023 T01 for a fraction of a pass. What would have
|
|
||||||
falsified it: a finding among the four whose reproduction needed engine
|
|
||||||
work that does not exist. There is none. **Affordability is not the
|
|
||||||
problem; sufficiency is (C1).**
|
|
||||||
- **That the register must reuse the provisional machinery rather than
|
|
||||||
compete with it** (§4). Checked against `rule-coverage.py:98-108,285-308`
|
|
||||||
— owner, raise date, age and a warn threshold all already exist and are
|
|
||||||
executable. What would have falsified it: a second age-tracking mechanism
|
|
||||||
already in the repo. There is none.
|
|
||||||
|
|
||||||
**The single challenge that forces a change to the design: C1.** The
|
|
||||||
proposal's load-bearing rule is stated as an existence test on a
|
|
||||||
reproduction. Three findings have now reached the record on wrong premises,
|
|
||||||
and **the third one satisfied the existence test**. The rule as written
|
|
||||||
does not catch its own worst observed failure. T03 cannot adopt it in the
|
|
||||||
form `CB-WP-0022:53` states it.
|
|
||||||
|
|
@ -1,187 +0,0 @@
|
||||||
# 260805 — response to the challenge on CB-RES-0007
|
|
||||||
|
|
||||||
One round, per InnerLoop §Step 2. **This review was run by a separate
|
|
||||||
agent**, unlike CB-RES-0006's, whose challenge opened with a fidelity note
|
|
||||||
saying it could not be. That note said *"treat the challenges below as a
|
|
||||||
lower bound on what a genuinely separate reviewer would find."* It was
|
|
||||||
right, and this is what the difference looks like: the separate reviewer
|
|
||||||
went to `git log` and found that the survey's showcase example had been
|
|
||||||
falsified four days earlier by our own commit.
|
|
||||||
|
|
||||||
Four challenges conceded, two conceded in part, one answered. **The
|
|
||||||
load-bearing rule survives but does not survive intact** — it is
|
|
||||||
insufficient as stated, and C1 is the reason.
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
## C1 — the showcase finding was itself a wrong premise, and it passed the rule — **conceded, and it is larger than the challenge says**
|
|
||||||
|
|
||||||
The challenge is correct and I verified every step of it independently
|
|
||||||
rather than accepting it.
|
|
||||||
|
|
||||||
`CB-WP-0022:61` defines admissibility by example:
|
|
||||||
|
|
||||||
> *"GR-E01 is admissible because 4/6/9 against 5/7/9 is a computation
|
|
||||||
> anyone can rerun."*
|
|
||||||
|
|
||||||
Commit `2da19a4` (2026-08-04, CB-WP-0021 T01/T02/T05) measured the deal
|
|
||||||
after the edition import: **6/9/12 against thresholds 5/7/9 — the game is
|
|
||||||
reachable at every seat count.** The scenario was renamed
|
|
||||||
`gr-e01-threshold-unreachable-2p.yaml` → `gr-e01-threshold-reachable-2p.yaml`.
|
|
||||||
The finding did not merely get a new number. **Its conclusion inverted.**
|
|
||||||
|
|
||||||
So the count is not two wrong premises reaching `ground-game`. It is
|
|
||||||
three, and **the third one was in the 2-of-6 set that passed the
|
|
||||||
reproduction rule.** The rule's own success case is a false finding that
|
|
||||||
shipped a green artifact.
|
|
||||||
|
|
||||||
**And it was still loaded.** `CB-WP-0022:283`, task T06, status `todo`:
|
|
||||||
|
|
||||||
> *"GR-E01 vs GR-S01 — the deal count puts 4/6/9 points in play against
|
|
||||||
> thresholds of 5/7/9, so either the count or the thresholds are wrong and
|
|
||||||
> no dataset reconciles them"*
|
|
||||||
|
|
||||||
That is queued for transmission to `ground-game` as a finding. Had T06 run
|
|
||||||
before this review, it would have sent a **fourth** wrong premise
|
|
||||||
downstream — one asserting no dataset reconciles numbers our own dataset
|
|
||||||
reconciled on 2026-08-04. The review did not catch a stale citation. It
|
|
||||||
caught a defect in flight.
|
|
||||||
|
|
||||||
**What this forces.** The rule at `CB-WP-0022:53` is an *existence* test:
|
|
||||||
a finding is admissible if a reproduction exists. Three of three wrong
|
|
||||||
premises would have passed an existence test, and one did. Existence is
|
|
||||||
therefore not the property that was missing. T03 must adopt the shape
|
|
||||||
requirement — `ground-game`'s row-level deal table, GROUND-WP-0004 T02 —
|
|
||||||
as **part of the rule**, not as the T04 addendum it currently is. And it
|
|
||||||
needs a second clause the survey did not contemplate: **a reproduction
|
|
||||||
must be able to fail.** A finding whose artifact has gone green is not a
|
|
||||||
closed finding until someone says which it is; ours silently became a
|
|
||||||
counterexample to itself.
|
|
||||||
|
|
||||||
`4/6/9` is corrected wherever it appears in this pass's artifacts, and
|
|
||||||
`evidence/CB-EV-0018:153` is left alone — it is a dated record of what was
|
|
||||||
believed then, and rewriting evidence is worse than a stale number in it.
|
|
||||||
|
|
||||||
## C2 — the baseline is a hand-maintained inventory reporting on itself — **conceded**
|
|
||||||
|
|
||||||
`tools/design-baseline.py:16-36` is a literal dict. `n = len(FINDINGS)`
|
|
||||||
(`:89`). **"6 findings" is a count of rows someone typed.** It is not a
|
|
||||||
measurement of anything, and the survey presented it as a measured
|
|
||||||
baseline, in a repo whose §Single source of fact rule exists precisely to
|
|
||||||
stop typed numbers.
|
|
||||||
|
|
||||||
Worse, and the part I would not have found:
|
|
||||||
|
|
||||||
- `has_reproduction` (`:38-43`) is `p.startswith("scenarios/")` and
|
|
||||||
nothing else. **It never checks the file exists.** The self-test's
|
|
||||||
YES-control (`:63`) passes a path — `gr-e01-threshold-unreachable-2p.yaml` —
|
|
||||||
that `2da19a4` deleted. The positive control for "detection can say
|
|
||||||
YES" is green against a file that is not there. That is a positive
|
|
||||||
control that controls nothing, which is the exact class `gates.toml`
|
|
||||||
exists to catch.
|
|
||||||
- The same detector **rejects a named test**, which `CB-WP-0022:55-56`
|
|
||||||
explicitly admits as a reproduction. So `33%` is biased down by the
|
|
||||||
detector's own narrowness.
|
|
||||||
- The 4-day figure is two date literals subtracted (`:96-98`). The
|
|
||||||
reviewer mutated them to 2020-01-01, got **2406 days**, and all five
|
|
||||||
self-test checks stayed green.
|
|
||||||
- `Makefile:127` runs `--self-test` only. **The reporting path is never
|
|
||||||
executed by CI**, so none of the above was ever going to be caught.
|
|
||||||
|
|
||||||
**What survives, and it matters that something does.** The *direction* of
|
|
||||||
every headline number is checkable without the tool: 11 files with no
|
|
||||||
index is verifiable by reading, and U1–U10 being open with 0 of 10 ruled
|
|
||||||
is verifiable against the hub. The baseline's conclusion — *we have
|
|
||||||
findings scattered across the repo with no register and no reproductions* —
|
|
||||||
is not in doubt. **What is withdrawn is its precision**, and any claim
|
|
||||||
that the 33% is a measured rate.
|
|
||||||
|
|
||||||
T05 must not build on this tool. The register itself is the measurement,
|
|
||||||
and `make design` reports over the register's contents, not over a dict.
|
|
||||||
|
|
||||||
## C3 — the arithmetic in the survey's own table — **conceded**
|
|
||||||
|
|
||||||
"Six provisional defaults" is **five**. The tool's own line prints
|
|
||||||
`5 location(s)`; `make coverage` prints 5; the survey table (`:46`) says
|
|
||||||
6. It was 6 at `2da19a4^` and 5 at HEAD — the survey table was written
|
|
||||||
against a repo state that had already moved.
|
|
||||||
|
|
||||||
And **GR-E01 is double-counted**: it is row 4 in its own right and was a
|
|
||||||
member of row 5's six. So the `2/6` denominator contains one finding
|
|
||||||
twice, and both copies are in the numerator. The real figure, on the
|
|
||||||
survey's own definitions, is not 2/6.
|
|
||||||
|
|
||||||
I am not restating a corrected rate here. **The tool that would produce it
|
|
||||||
is the one C2 just withdrew**, and quoting a new number from a discredited
|
|
||||||
instrument is how the first one got in.
|
|
||||||
|
|
||||||
## C4 — §Underdetermined was never evaluated as a candidate — **conceded in part**
|
|
||||||
|
|
||||||
Correct, and it is the omission that embarrasses the survey most, because
|
|
||||||
§Step 1 asks for the best implementation *in existence* and this one is in
|
|
||||||
our own repo. The survey went to Magic to discover the Oracle property —
|
|
||||||
*the fix lands in the source, the annotation is not authoritative* — and
|
|
||||||
`specs/GroundRules.md:231-233` already states it in our own words:
|
|
||||||
*"a ruling flips the scenario, not the kernel."*
|
|
||||||
|
|
||||||
**Conceded in part, not whole,** and the reviewer marked the honest limit
|
|
||||||
themselves: 10 U-items owe provisional scenarios, 5 exist, and
|
|
||||||
`make coverage` never flags the missing five. A section that cannot report
|
|
||||||
its own gaps is not yet the register. But T03's burden has changed
|
|
||||||
direction — it must now argue why §Underdetermined should be **extended**
|
|
||||||
rather than replaced, and the default answer is that it should be.
|
|
||||||
|
|
||||||
## C5 — the third thing exists — **conceded as to the argument, not the conclusion**
|
|
||||||
|
|
||||||
`specs/InnerLoopReference.md` is 383 lines with two sections literally
|
|
||||||
headed *"moved from InnerLoop v1.6"*, and `history/` holds 769 lines of
|
|
||||||
retrospectives. That is content about *how we learn to build this engine*
|
|
||||||
which was placed where it fit rather than where it belonged. My redundancy
|
|
||||||
inventory named `evidence/`, `decisions/`, `gates.toml` and workplans, and
|
|
||||||
**named neither of those** — so the inventory was incomplete and the
|
|
||||||
argument as written does not stand.
|
|
||||||
|
|
||||||
The conclusion probably does, and for a reason the challenge concedes: the
|
|
||||||
homeless content found homes, and both are load-bearing and read. A third
|
|
||||||
register would be a third place to look for material that is currently in
|
|
||||||
two. **T03 records this as argued-and-narrowed rather than settled**, which
|
|
||||||
is the outcome the declaration asked for when it said this was the
|
|
||||||
judgment it wanted reviewed rather than assumed.
|
|
||||||
|
|
||||||
## C6 — "queue that empties" vs keeping resolved findings — **answered**
|
|
||||||
|
|
||||||
Marked weak by the reviewer and it is weak, but it has a real answer.
|
|
||||||
Magic's property is that a finding **closes when the source changes**. It
|
|
||||||
is not that the record is deleted. `design-baseline.py:23-25`'s comment —
|
|
||||||
keep resolved findings or you cannot report a close rate — is compatible:
|
|
||||||
the *queue* of open findings empties, the *log* of closed ones does not.
|
|
||||||
T03's lifecycle should say this in one sentence so the next reader does
|
|
||||||
not re-litigate it.
|
|
||||||
|
|
||||||
Falsifier: if the register's default view shows closed findings mixed with
|
|
||||||
open ones, the queue property is gone regardless of what the lifecycle
|
|
||||||
says.
|
|
||||||
|
|
||||||
## C7 — n=6, all engine-surfaced — **already conceded in §6**
|
|
||||||
|
|
||||||
The survey states this limit itself. Nothing to add.
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
## What the review changed
|
|
||||||
|
|
||||||
| | before | after |
|
|
||||||
|---|---|---|
|
|
||||||
| the rule | reproduction must **exist** | must exist, **have the ruled shape**, and **be able to fail** |
|
|
||||||
| wrong premises shipped | 2 | **3** — and one passed the rule |
|
|
||||||
| the baseline | a measured 33% | a hand-maintained inventory; direction stands, precision withdrawn |
|
|
||||||
| §Underdetermined | not considered | the incumbent T03 must argue against |
|
|
||||||
| engine-evolution register | redundant | still redundant, on a corrected inventory |
|
|
||||||
| T06's payload | a finding | **a fourth wrong premise, caught before sending** |
|
|
||||||
|
|
||||||
**The strongest thing here is not any single challenge.** It is that a
|
|
||||||
separate reviewer with no stake ran `git log` against the survey's central
|
|
||||||
example and found it had been falsified by our own commit four days
|
|
||||||
earlier — while the author, who wrote that commit, quoted the dead number
|
|
||||||
twice. CB-RES-0006's review said a same-session reviewer would be a lower
|
|
||||||
bound. This is how much lower.
|
|
||||||
|
|
@ -1,183 +0,0 @@
|
||||||
# GameDesign — findings about the game, and the trials that produce them
|
|
||||||
|
|
||||||
v1.0 — CB-WP-0022 T04, 2026-08-05. Normative. Implements
|
|
||||||
[ADR-0012](../decisions/ADR-0012-the-design-instrument.md); INTENT
|
|
||||||
property 4 (*Instrument*) is what this spec discharges.
|
|
||||||
|
|
||||||
**This is not a register.** The register is
|
|
||||||
[`GroundRules.md §Underdetermined`](GroundRules.md) — ADR-0012 D2. This
|
|
||||||
document says what may go in it, what a reproduction must show, how a
|
|
||||||
finding moves and dies, and how a trial game is run so *"we played it and
|
|
||||||
X happened"* is replayable rather than remembered.
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
## 1. Admissibility
|
|
||||||
|
|
||||||
> A finding is admissible when its reproduction **exists**, **has the
|
|
||||||
> ruled shape**, and **can fail**. All three. (ADR-0012 D3.)
|
|
||||||
|
|
||||||
### 1.1 Exists
|
|
||||||
|
|
||||||
One of: a scenario under `scenarios/ground/`, an arithmetic check that
|
|
||||||
prints the contradiction, a `.cbreplay` bundle or recorded scenario the
|
|
||||||
reader can replay, or a **named test** — `crate::module::test_name`, so it
|
|
||||||
can be run.
|
|
||||||
|
|
||||||
The register entry names the artifact **by path**. A finding whose
|
|
||||||
artifact path does not resolve is not admissible, and §4's check stats the
|
|
||||||
file rather than pattern-matching the string — `design-baseline.py`'s
|
|
||||||
`has_reproduction` matched `startswith("scenarios/")` and its own positive
|
|
||||||
control was green against a file that had been deleted (ADR-0012 D8).
|
|
||||||
|
|
||||||
### 1.2 Has the ruled shape
|
|
||||||
|
|
||||||
Per `ground-game`, GROUND-WP-0004 T02:
|
|
||||||
|
|
||||||
> An arithmetic finding ships a **row-level table** — Surface and each
|
|
||||||
> hidden priority listed **separately** — never *"sum of file"*, never
|
|
||||||
> *"deal depth N"*.
|
|
||||||
|
|
||||||
**A reproduction must show the quantity the prose claims.** This is the
|
|
||||||
clause that exists because passing artifacts accompanied wrong
|
|
||||||
descriptions:
|
|
||||||
|
|
||||||
| wrong premise | what shipped | what a row-level table would have shown |
|
|
||||||
|---|---|---|
|
|
||||||
| *"12 in the file"* | a sum | which rows, at which priorities |
|
|
||||||
| *"4/6/9 against 5/7/9"* | a scenario, green | 6/9/12 — the deal, per row |
|
|
||||||
| *"SOLVE offered on a face-down Problem"* | a condition named | which filter actually fired |
|
|
||||||
|
|
||||||
**Requirement.** Where a finding asserts a number, the register entry
|
|
||||||
quotes the number **and** the artifact prints the rows it came from. A
|
|
||||||
finding stating a total without its rows is inadmissible even if the total
|
|
||||||
is right — *"12"* was arithmetically defensible and still wrong about the
|
|
||||||
game.
|
|
||||||
|
|
||||||
### 1.3 Can fail
|
|
||||||
|
|
||||||
The artifact must be capable of going red, and the register records its
|
|
||||||
**current colour**. A reproduction that has gone **green is not a closed
|
|
||||||
finding** — it is a finding whose claim is now in question, and §4 reports
|
|
||||||
it as needing resolution to `applied` or `withdrawn`.
|
|
||||||
|
|
||||||
**A reproduction has a `role`, and only one of them alarms.** This
|
|
||||||
distinction was not in ADR-0012; **T05's backfill produced it**, when the
|
|
||||||
first report cried wolf over U2.
|
|
||||||
|
|
||||||
| role | what it is | green means |
|
|
||||||
|---|---|---|
|
|
||||||
| `counterexample` | demonstrates the defect — the model-checker shape | **alarm.** It should be red while the finding is open |
|
|
||||||
| `default` | encodes a provisional choice so it is testable and visible | **expected.** It passes because our default is implemented, which says nothing about whether the owner agrees |
|
|
||||||
|
|
||||||
U2's scenario is green because the default it documents is implemented.
|
|
||||||
GR-E01's was a counterexample that went green when the edition landed, and
|
|
||||||
sat that way for four days. **Same colour, opposite meaning** — a register
|
|
||||||
that cannot tell them apart either alarms constantly or never.
|
|
||||||
|
|
||||||
**Rationale, and it is the expensive one.** `gr-e01-threshold-…-2p.yaml`
|
|
||||||
went green when the edition landed on 2026-08-04. The finding remained
|
|
||||||
admissible, remained in the register, and remained queued for transmission
|
|
||||||
to `ground-game` for four days, because nothing treated a passing artifact
|
|
||||||
as a signal. It was stopped by an adversarial review, not by a control.
|
|
||||||
|
|
||||||
## 2. Kinds
|
|
||||||
|
|
||||||
Five (ADR-0012 D4). **No kind exists without a finding that instantiates
|
|
||||||
it**; a sixth appearing during backfill means the taxonomy was invented
|
|
||||||
rather than observed, and T07 must say so.
|
|
||||||
|
|
||||||
`underdetermined` · `inconsistent` · `inert` · `degenerate` · `unplayed`
|
|
||||||
|
|
||||||
## 3. States, owners, and death
|
|
||||||
|
|
||||||
```
|
|
||||||
raised ──► reported ──► ruled ──► applied
|
|
||||||
│
|
|
||||||
└──────────► withdrawn ◄──────────┘
|
|
||||||
```
|
|
||||||
|
|
||||||
- `applied` means **the source changed** — dataset or rules text — and our
|
|
||||||
provisional default was deleted. Not *"a ruling was recorded."*
|
|
||||||
- **The queue is `raised` + `reported`. The log is `applied` +
|
|
||||||
`withdrawn`.** `make design` reports the queue by default; the log only
|
|
||||||
on request. If the default view mixes them, the queue property is gone.
|
|
||||||
- **`withdrawn` is reported downstream, not deleted.** A claim retracted
|
|
||||||
silently is how three wrong premises survived.
|
|
||||||
|
|
||||||
### 3.1 Notes
|
|
||||||
|
|
||||||
A finding without a reproduction is a **note** (ADR-0012 D6). A note:
|
|
||||||
|
|
||||||
1. **has no `reported` state** — it may not cross the repo boundary;
|
|
||||||
2. **expires at 30 days**, flagged by §4 on the same mechanism that ages
|
|
||||||
provisional items (`rule-coverage.py:54`, `PROVISIONAL_WARN_DAYS`);
|
|
||||||
3. is promoted only by acquiring a reproduction meeting §1 — there is no
|
|
||||||
path from `note` to `reported` that skips it.
|
|
||||||
|
|
||||||
## 4. Metrics
|
|
||||||
|
|
||||||
`make design` reports these. Every number is computed **over the
|
|
||||||
register**, never from a maintained list (ADR-0012 D8).
|
|
||||||
|
|
||||||
| metric | target | notes |
|
|
||||||
|---|---|---|
|
|
||||||
| findings with a resolving reproduction | **100%** | denominator includes `withdrawn`; notes counted separately, never folded in |
|
|
||||||
| **reproductions that have gone green while open** | **0** | the §1.3 alarm. Non-zero is a queue that needs reading, not a passing suite |
|
|
||||||
| arithmetic findings whose artifact prints rows | **100%** | §1.2 |
|
|
||||||
| median days `raised` → `reported` | the number this exists to fix | the baseline it must beat is **4 days** to first *read* |
|
|
||||||
| open findings by kind, with age | reported, not targeted | |
|
|
||||||
| notes past 30 days | **0** | |
|
|
||||||
| findings closed by a ruling vs still open | reported, with age | |
|
|
||||||
|
|
||||||
**Falsifier for this spec.** If a finding reaches `ground-game` whose
|
|
||||||
prose describes a quantity its own artifact does not compute, §1.2 has
|
|
||||||
failed and the next revision must **generate the description from the
|
|
||||||
artifact** rather than let it be written alongside.
|
|
||||||
|
|
||||||
**No baseline rate is quoted here.** CB-RES-0007's 33% was withdrawn as a
|
|
||||||
measured figure (ADR-0012 D8); the first honest denominator is the one
|
|
||||||
T05's backfill produces.
|
|
||||||
|
|
||||||
## 5. The trial protocol
|
|
||||||
|
|
||||||
The part with no precedent in this repo. A **trial game** is a recorded
|
|
||||||
session plus an observation log — so a claim from play is replayable.
|
|
||||||
|
|
||||||
**It must cost almost nothing or it will not be done.** The engine already
|
|
||||||
records: `cb-play --record FILE` writes the finished game as a scenario
|
|
||||||
YAML (`main.rs:30`), and `games/ground/src/record.rs` is trustworthy
|
|
||||||
because its round trip is the test.
|
|
||||||
|
|
||||||
A trial is therefore:
|
|
||||||
|
|
||||||
```
|
|
||||||
cb-play --players N --serve PORT --record trials/<date>-<slug>.yaml
|
|
||||||
```
|
|
||||||
|
|
||||||
plus a sibling `trials/<date>-<slug>.md` carrying, and nothing more:
|
|
||||||
|
|
||||||
- **who played which seats**, and which were bots (with policy);
|
|
||||||
- **what was observed** — in the player's own words, not cleaned up;
|
|
||||||
- **the seed**, which the scenario already carries;
|
|
||||||
- **which observations became findings or notes**, by register id, or
|
|
||||||
explicitly none.
|
|
||||||
|
|
||||||
**Rules.**
|
|
||||||
|
|
||||||
1. An observation is a **note** until it has a reproduction (§3.1). A
|
|
||||||
trial log is not a finding source by itself — it is where findings are
|
|
||||||
*noticed*.
|
|
||||||
2. **The recording is the artifact for `unplayed` findings.** GR-E03/E04
|
|
||||||
were "never played to the end"; a trial that plays them to the end both
|
|
||||||
closes that finding and produces the reproduction for anything found on
|
|
||||||
the way.
|
|
||||||
3. A trial whose recording does not replay is a **bug in the engine**,
|
|
||||||
reported as one, not a design finding.
|
|
||||||
|
|
||||||
**Why the log is prose and stays prose.** The maintainer's *"I felt it was
|
|
||||||
too easy but then we lost, so who knows"* is exactly the observation this
|
|
||||||
protocol exists to catch: real, unreproducible as stated, and the seed of
|
|
||||||
[CB-WP-0025](../workplans/CB-WP-0025-could-we-have-won.md). Forcing it
|
|
||||||
into a schema at the moment of observation would lose it. §3.1 is what
|
|
||||||
stops it aging into an apparent finding.
|
|
||||||
|
|
@ -219,103 +219,6 @@ terms.
|
||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
## The finding register
|
|
||||||
|
|
||||||
**This section is the design-finding register** (ADR-0012 D2). It was the
|
|
||||||
register for dataset ambiguities already; CB-WP-0022 extended it to all
|
|
||||||
five kinds rather than building a second one beside it. Admissibility,
|
|
||||||
kinds, states and metrics: [`GameDesign.md`](GameDesign.md). Reported by
|
|
||||||
`make design`.
|
|
||||||
|
|
||||||
<!-- design-register:begin -->
|
|
||||||
|
|
||||||
| id | kind | state | reproduction | role | raised | owner |
|
|
||||||
|---|---|---|---|---|---|---|
|
|
||||||
| U1 | underdetermined | ruled | — | default | 2026-07-31 | ground-game |
|
|
||||||
| U2 | underdetermined | ruled | scenarios/ground/gr-d01-darvo-trigger.yaml | default | 2026-07-31 | ground-game |
|
|
||||||
| U3 | underdetermined | ruled | — | default | 2026-07-31 | ground-game |
|
|
||||||
| U4 | underdetermined | ruled | — | default | 2026-07-31 | ground-game |
|
|
||||||
| U5 | underdetermined | ruled | — | default | 2026-07-31 | ground-game |
|
|
||||||
| U6 | underdetermined | ruled | — | default | 2026-07-31 | ground-game |
|
|
||||||
| U7 | underdetermined | ruled | — | default | 2026-07-31 | ground-game |
|
|
||||||
| U8 | underdetermined | ruled | — | default | 2026-07-31 | ground-game |
|
|
||||||
| U9 | underdetermined | ruled | — | default | 2026-07-31 | ground-game |
|
|
||||||
| U10 | underdetermined | ruled | — | default | 2026-07-31 | ground-game |
|
|
||||||
| F11 | inert | applied | scenarios/ground/gr-p05-solve-legality.yaml | counterexample | 2026-08-02 | clay-borg |
|
|
||||||
| F12 | degenerate | note | — | — | 2026-08-01 | clay-borg |
|
|
||||||
| F13 | inconsistent | withdrawn | scenarios/ground/gr-e01-threshold-reachable-2p.yaml | counterexample | 2026-08-01 | clay-borg |
|
|
||||||
| F14 | unplayed | note | — | — | 2026-08-01 | clay-borg |
|
|
||||||
|
|
||||||
<!-- design-register:end -->
|
|
||||||
|
|
||||||
- **F11 — SOLVE offered where it cannot act.** Offered on a face-down
|
|
||||||
Problem, or with no matching suit in hand; inert every time. Ruled
|
|
||||||
GROUND-WP-0002 T02, implemented CB-WP-0023 as GR-P05. `applied` — the
|
|
||||||
rule changed, not just the annotation. **The case we reported was not the
|
|
||||||
case that fired**: `validate` already rejected face-down, and the
|
|
||||||
maintainer's three inert SOLVEs were the *hand* case.
|
|
||||||
- **F12 — GR-A13 "wasted SOLVE" on an already-claimed Problem.** A
|
|
||||||
scenario had to pick a default and did. `note`: no artifact isolates the
|
|
||||||
degenerate line, so under GameDesign §3.1 it may not be reported until
|
|
||||||
one exists.
|
|
||||||
- **F13 — GR-E01 vs GR-S01, withdrawn 2026-08-05.** Raised as *"4/6/9
|
|
||||||
against 5/7/9, no dataset reconciles them."* `2da19a4` measured **6/9/12
|
|
||||||
against 5/7/9** and the scenario was renamed `-unreachable-` →
|
|
||||||
`-reachable-`. **Its reproduction is green**, which under GameDesign §1.3
|
|
||||||
is the alarm that forced the resolution. Withdrawn rather than deleted,
|
|
||||||
and the withdrawal is reported (ADR-0012 D5).
|
|
||||||
- **F14 — GR-E03/GR-E04 never played to the end.** Nineteen passes, never
|
|
||||||
played out. `note` until a trial game exists; GROUND-WP-0003 is the
|
|
||||||
playtest that would close it, and GameDesign §5's protocol makes the
|
|
||||||
recording the artifact.
|
|
||||||
|
|
||||||
### The register's first run found ten answers nobody had collected
|
|
||||||
|
|
||||||
**U1–U10 are `ruled`, not `reported`.** GROUND-WP-0002 T05 answered all
|
|
||||||
ten on **2026-08-03** — every one **confirmed** as the default clay-borg
|
|
||||||
already simulates — and GROUND-WP-0002 T03 confirmed five of the six
|
|
||||||
provisional scenarios, voiding `gr-e01` as a rules gap. The workplan is
|
|
||||||
`finished`.
|
|
||||||
|
|
||||||
CB-RES-0007 reported *"0 of 10 ruled"* and this register was built saying
|
|
||||||
`reported`. **Both were two days stale on the day they were written.** The
|
|
||||||
answers had arrived and nothing propagated them — the same failure as the
|
|
||||||
unread inbox, in the opposite direction.
|
|
||||||
|
|
||||||
**They are `ruled`, not `applied`, and the difference is work we owe.**
|
|
||||||
Per ADR-0012 D5, `applied` means the source changed and the provisional
|
|
||||||
default was deleted. The rulings *confirmed* our defaults, so the rules did
|
|
||||||
not move — but the scenarios still carry `provisional: true` for choices
|
|
||||||
that are now settled. **Lifting those flags and recording each ruling is
|
|
||||||
what closes U1–U10**, and it is not done. `make design` shows them open
|
|
||||||
until it is.
|
|
||||||
|
|
||||||
### What the backfill measured, and what it contradicted
|
|
||||||
|
|
||||||
**Only U2 names its U-item in a scenario.** Measured, not estimated:
|
|
||||||
|
|
||||||
```
|
|
||||||
for u in U1 .. U10; do grep -lE "\b$u\b" scenarios/ground/*.yaml; done
|
|
||||||
```
|
|
||||||
|
|
||||||
CB-RES-0007 asserted *"six of the ten already have provisional
|
|
||||||
scenarios."* Five provisional scenarios exist and **one** cites the item
|
|
||||||
it stands for. The other four may well encode U-item defaults — the
|
|
||||||
mapping is simply **not written down**, so it is not checkable, and an
|
|
||||||
uncheckable link is the defect this register exists to fix. The register
|
|
||||||
records what is citable; the rest is debt, visible in `make design`.
|
|
||||||
|
|
||||||
**No sixth kind was needed** — the five kinds absorbed all four non-U
|
|
||||||
findings. And the survey's *"six provisional defaults"* was **not** entered
|
|
||||||
as a finding: C3 showed it double-counted GR-E01, and the provisional
|
|
||||||
scenarios are *reproductions for* underdetermined items, not a finding of
|
|
||||||
their own.
|
|
||||||
|
|
||||||
**U1–U10 are `reported` while lacking reproductions**, which GameDesign
|
|
||||||
§3.1 would now forbid. They were reported on 2026-07-30, before the rule
|
|
||||||
existed. They are grandfathered rather than rewritten, and the debt is a
|
|
||||||
reported metric with a target of zero.
|
|
||||||
|
|
||||||
## Underdetermined in dataset 0.1 — PROVISIONAL defaults (flag to ground-game)
|
## Underdetermined in dataset 0.1 — PROVISIONAL defaults (flag to ground-game)
|
||||||
|
|
||||||
**Owner: ground-game. Raised: 2026-07-31.** Every U-item carries an owner
|
**Owner: ground-game. Raised: 2026-07-31.** Every U-item carries an owner
|
||||||
|
|
|
||||||
|
|
@ -7,26 +7,6 @@ model-checker traces, W3C provisional marks) are practices rather than
|
||||||
runnable software, so per InnerLoop Step 1 their rows are DIRECTIONAL and
|
runnable software, so per InnerLoop Step 1 their rows are DIRECTIONAL and
|
||||||
cap at `parity`. This is the row that can be measured.
|
cap at `parity`. This is the row that can be measured.
|
||||||
"""
|
"""
|
||||||
#
|
|
||||||
# ── SUPERSEDED 2026-08-05 by tools/design.py (ADR-0012 D8) ──────────────
|
|
||||||
#
|
|
||||||
# Kept as CB-RES-0007's dated snapshot, NOT as a live instrument. Deleting
|
|
||||||
# it would erase the evidence for how a wrong number got into a survey.
|
|
||||||
#
|
|
||||||
# The adversarial review (C2) dismantled it and every point is true:
|
|
||||||
# * FINDINGS below is a hand-maintained dict; `n = len(FINDINGS)` counts
|
|
||||||
# the rows someone typed, not findings that exist.
|
|
||||||
# * `has_reproduction` is `startswith("scenarios/")` and never stats the
|
|
||||||
# file, so the self-test's YES-control is GREEN against a path that
|
|
||||||
# 2da19a4 deleted. A positive control that controls nothing.
|
|
||||||
# * It rejects a named test, which the admissibility rule admits, so its
|
|
||||||
# 33% is biased downward.
|
|
||||||
# * The 4-day figure is two date literals subtracted. Mutated to
|
|
||||||
# 2020-01-01 it reports 2406 days with every check still green.
|
|
||||||
#
|
|
||||||
# **Its 33% is withdrawn as a measured rate.** What survived: 11 files,
|
|
||||||
# no index, 0 of 10 ruled — all checkable without this tool.
|
|
||||||
#
|
|
||||||
import os, re, subprocess, sys, datetime
|
import os, re, subprocess, sys, datetime
|
||||||
|
|
||||||
ROOT = "/home/worsch/clay-borg"
|
ROOT = "/home/worsch/clay-borg"
|
||||||
|
|
|
||||||
194
tools/design.py
194
tools/design.py
|
|
@ -1,194 +0,0 @@
|
||||||
#!/usr/bin/env python3
|
|
||||||
"""design — report the finding register (CB-WP-0022 T05).
|
|
||||||
|
|
||||||
Implements the metrics in `specs/GameDesign.md` §4 over the register in
|
|
||||||
`specs/GroundRules.md`. ADR-0012 D8 retired `design-baseline.py`, which
|
|
||||||
was a hand-maintained dict counting itself; the difference that matters is
|
|
||||||
that **every number here is computed over rows that name real files, and
|
|
||||||
the reproduction check stats the file.**
|
|
||||||
|
|
||||||
`design-baseline.py`'s `has_reproduction` was `p.startswith("scenarios/")`
|
|
||||||
and nothing else, so its own positive control was green against a path
|
|
||||||
that had been deleted. The self-test below asserts the opposite property
|
|
||||||
directly: a row citing a nonexistent file must NOT count as reproduced.
|
|
||||||
"""
|
|
||||||
|
|
||||||
import os, re, sys, subprocess, datetime
|
|
||||||
|
|
||||||
ROOT = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
|
|
||||||
REGISTER = os.path.join(ROOT, "specs", "GroundRules.md")
|
|
||||||
|
|
||||||
BEGIN = "<!-- design-register:begin -->"
|
|
||||||
END = "<!-- design-register:end -->"
|
|
||||||
|
|
||||||
KINDS = {"underdetermined", "inconsistent", "inert", "degenerate", "unplayed"}
|
|
||||||
OPEN_STATES = {"raised", "reported", "ruled"} # the queue: not yet applied
|
|
||||||
CLOSED_STATES = {"applied", "withdrawn"} # the log
|
|
||||||
NOTE = "note"
|
|
||||||
NOTE_EXPIRY_DAYS = 30 # GameDesign §3.1, same figure as rule-coverage.py
|
|
||||||
|
|
||||||
|
|
||||||
def parse(text):
|
|
||||||
"""Rows between the register markers. Raises if the block is absent —
|
|
||||||
a register that silently reports zero findings is worse than one that
|
|
||||||
fails."""
|
|
||||||
try:
|
|
||||||
block = text.split(BEGIN)[1].split(END)[0]
|
|
||||||
except IndexError:
|
|
||||||
raise SystemExit(f"no design-register block in {REGISTER}")
|
|
||||||
rows = []
|
|
||||||
for line in block.splitlines():
|
|
||||||
line = line.strip()
|
|
||||||
if not line.startswith("|") or line.startswith("|---"):
|
|
||||||
continue
|
|
||||||
cells = [c.strip() for c in line.strip("|").split("|")]
|
|
||||||
if len(cells) != 7 or cells[0] in ("id",):
|
|
||||||
continue
|
|
||||||
rows.append(dict(zip(
|
|
||||||
("id", "kind", "state", "repro", "role", "raised", "owner"), cells)))
|
|
||||||
return rows
|
|
||||||
|
|
||||||
|
|
||||||
def reproduced(row, root=ROOT):
|
|
||||||
"""GameDesign §1.1: the artifact must resolve. A named test is admitted
|
|
||||||
by its `crate::module::name` shape; anything else must be a real path
|
|
||||||
on disk, and that is checked by stat, not by prefix."""
|
|
||||||
p = row["repro"]
|
|
||||||
if p in ("", "—", "-"):
|
|
||||||
return False
|
|
||||||
if "::" in p: # a named test
|
|
||||||
return True
|
|
||||||
return os.path.exists(os.path.join(root, p))
|
|
||||||
|
|
||||||
|
|
||||||
_SIM = None
|
|
||||||
|
|
||||||
|
|
||||||
def sim_passes(root=ROOT):
|
|
||||||
"""`make sim` is the authority on whether scenarios pass. Run once and
|
|
||||||
cache — asking per row costs a full suite per finding."""
|
|
||||||
global _SIM
|
|
||||||
if _SIM is None:
|
|
||||||
r = subprocess.run(["make", "sim"], cwd=root, capture_output=True, text=True)
|
|
||||||
_SIM = r.returncode == 0
|
|
||||||
return _SIM
|
|
||||||
|
|
||||||
|
|
||||||
def alarming(row, root=ROOT):
|
|
||||||
"""§1.3, and the distinction the backfill discovered: **only a
|
|
||||||
counterexample alarms when green.**
|
|
||||||
|
|
||||||
A `default` reproduction encodes a provisional choice and is SUPPOSED
|
|
||||||
to pass — U2's scenario is green because the default it documents is
|
|
||||||
implemented, which says nothing about whether ground-game agrees. A
|
|
||||||
`counterexample` is the model-checker shape: it demonstrates the
|
|
||||||
defect, so it must be red while the finding is open. GR-E01's went
|
|
||||||
green when the edition landed and nothing noticed for four days.
|
|
||||||
"""
|
|
||||||
if row["role"] != "counterexample":
|
|
||||||
return False
|
|
||||||
if not row["repro"].startswith("scenarios/"):
|
|
||||||
return False
|
|
||||||
return sim_passes(root)
|
|
||||||
|
|
||||||
|
|
||||||
def report(root=ROOT, today=None):
|
|
||||||
today = today or datetime.date.today()
|
|
||||||
rows = parse(open(os.path.join(root, "specs", "GroundRules.md")).read())
|
|
||||||
|
|
||||||
findings = [r for r in rows if r["state"] != NOTE]
|
|
||||||
notes = [r for r in rows if r["state"] == NOTE]
|
|
||||||
queue = [r for r in findings if r["state"] in OPEN_STATES]
|
|
||||||
closed = [r for r in findings if r["state"] in CLOSED_STATES]
|
|
||||||
|
|
||||||
print("design — the finding register\n")
|
|
||||||
print(" QUEUE (open findings)")
|
|
||||||
for r in sorted(queue, key=lambda r: r["raised"]):
|
|
||||||
age = (today - datetime.date.fromisoformat(r["raised"])).days
|
|
||||||
mark = "repro" if reproduced(r, root) else " - "
|
|
||||||
print(f" {mark} {r['id']:<4} {r['kind']:<16} {r['state']:<9} {age:>4}d {r['owner']}")
|
|
||||||
|
|
||||||
if notes:
|
|
||||||
print("\n NOTES (not reportable — GameDesign §3.1)")
|
|
||||||
for r in sorted(notes, key=lambda r: r["raised"]):
|
|
||||||
age = (today - datetime.date.fromisoformat(r["raised"])).days
|
|
||||||
flag = " EXPIRED" if age > NOTE_EXPIRY_DAYS else ""
|
|
||||||
print(f" {r['id']:<4} {r['kind']:<16} {age:>4}d{flag}")
|
|
||||||
|
|
||||||
# ---- §4 metrics. The log is reported on request, not by default,
|
|
||||||
# because a default view mixing open and closed loses the queue
|
|
||||||
# property (ADR-0012 D5).
|
|
||||||
n = len(findings)
|
|
||||||
repro = sum(1 for r in findings if reproduced(r, root))
|
|
||||||
debt = [r for r in queue if not reproduced(r, root)]
|
|
||||||
expired = [r for r in notes
|
|
||||||
if (today - datetime.date.fromisoformat(r["raised"])).days > NOTE_EXPIRY_DAYS]
|
|
||||||
unresolved_green = [r for r in queue if reproduced(r, root) and alarming(r, root)]
|
|
||||||
|
|
||||||
print(f"\n findings {n} (+{len(notes)} note(s))")
|
|
||||||
print(f" with a resolving reproduction {repro}/{n}"
|
|
||||||
f"{'' if not n else f' = {100*repro//n}%'} target 100%")
|
|
||||||
print(f" open, lacking a reproduction {len(debt)} target 0"
|
|
||||||
+ (f" [{', '.join(r['id'] for r in debt)}]" if debt else ""))
|
|
||||||
print(f" reproductions green while open {len(unresolved_green)} target 0"
|
|
||||||
+ (" <-- ALARM, GameDesign §1.3" if unresolved_green else ""))
|
|
||||||
print(f" notes past {NOTE_EXPIRY_DAYS} days {len(expired)} target 0")
|
|
||||||
print(f" closed (log) {len(closed)}"
|
|
||||||
f" [{', '.join(r['id'] for r in closed)}]" if closed else "")
|
|
||||||
|
|
||||||
bad = [r for r in rows if r["kind"] not in KINDS]
|
|
||||||
if bad:
|
|
||||||
print(f"\n UNKNOWN KIND: {', '.join(r['id'] for r in bad)}"
|
|
||||||
" — a sixth kind means the taxonomy was invented (ADR-0012 D4)")
|
|
||||||
return 0
|
|
||||||
|
|
||||||
|
|
||||||
def self_test():
|
|
||||||
ok = True
|
|
||||||
|
|
||||||
def check(name, cond, detail=""):
|
|
||||||
nonlocal ok
|
|
||||||
ok = ok and bool(cond)
|
|
||||||
print(f" {'ok ' if cond else 'FAIL'} {name}{' ' + detail if detail else ''}")
|
|
||||||
|
|
||||||
rows = parse(open(REGISTER).read())
|
|
||||||
check("the register parses", len(rows) >= 14, f"{len(rows)} row(s)")
|
|
||||||
check("every kind is one of the five",
|
|
||||||
all(r["kind"] in KINDS for r in rows),
|
|
||||||
"a sixth kind means the taxonomy was invented")
|
|
||||||
check("every state is known",
|
|
||||||
all(r["state"] in OPEN_STATES | CLOSED_STATES | {NOTE} for r in rows))
|
|
||||||
|
|
||||||
# THE control design-baseline.py did not have. Its YES-control passed a
|
|
||||||
# path that had been deleted and still returned True.
|
|
||||||
check("a nonexistent reproduction does NOT count",
|
|
||||||
not reproduced({"repro": "scenarios/ground/gr-e01-threshold-unreachable-2p.yaml"}),
|
|
||||||
"this exact path was deleted by 2da19a4 and the old tool said yes")
|
|
||||||
check("a real reproduction DOES count",
|
|
||||||
reproduced({"repro": "scenarios/ground/gr-p05-solve-legality.yaml"}))
|
|
||||||
check("a named test counts", reproduced({"repro": "games_ground::view::tests::a_spectator_sees_no_hands"}))
|
|
||||||
check("an em-dash does not count", not reproduced({"repro": "—"}))
|
|
||||||
|
|
||||||
# The distinction the backfill discovered: a green DEFAULT is expected,
|
|
||||||
# a green COUNTEREXAMPLE is the alarm. Without this the report cried
|
|
||||||
# wolf over U2, whose scenario is green precisely because the
|
|
||||||
# provisional default it documents is implemented.
|
|
||||||
check("a green default does not alarm",
|
|
||||||
not alarming({"role": "default", "repro": "scenarios/ground/gr-d01-darvo-trigger.yaml"}))
|
|
||||||
check("role is recorded for every reproduced row",
|
|
||||||
all(r["role"] in ("counterexample", "default")
|
|
||||||
for r in rows if r["repro"] not in ("—", "-", "")))
|
|
||||||
|
|
||||||
# The measured claim in the backfill: only U2 cites its U-item.
|
|
||||||
cited = [u for u in [f"U{i}" for i in range(1, 11)]
|
|
||||||
if subprocess.run(["grep", "-rlE", rf"\b{u}\b"] +
|
|
||||||
[os.path.join(ROOT, "scenarios", "ground")],
|
|
||||||
capture_output=True, text=True, cwd=ROOT).stdout.strip()]
|
|
||||||
check("exactly one U-item is cited by a scenario", cited == ["U2"], f"{cited}")
|
|
||||||
|
|
||||||
print("design self-test (positive control)")
|
|
||||||
return 0 if ok else 1
|
|
||||||
|
|
||||||
|
|
||||||
if __name__ == "__main__":
|
|
||||||
sys.exit(self_test() if "--self-test" in sys.argv else report())
|
|
||||||
|
|
@ -2,7 +2,7 @@
|
||||||
id: CB-WP-0022
|
id: CB-WP-0022
|
||||||
kind: product
|
kind: product
|
||||||
title: "The design instrument: findings about the game, with their reproductions"
|
title: "The design instrument: findings about the game, with their reproductions"
|
||||||
status: done
|
status: active
|
||||||
state_hub_workstream_id: "fda16340-0049-4acf-884b-a5cfdbde47c0"
|
state_hub_workstream_id: "fda16340-0049-4acf-884b-a5cfdbde47c0"
|
||||||
---
|
---
|
||||||
|
|
||||||
|
|
@ -36,8 +36,8 @@ has produced, as a by-product of being rigorous:
|
||||||
| ten underdetermined rules points (U1–U10) | formalizing the dataset into testable rules |
|
| ten underdetermined rules points (U1–U10) | formalizing the dataset into testable rules |
|
||||||
| SOLVE offered on a face-down Problem, always inert | a human dragging it three rounds running |
|
| SOLVE offered on a face-down Problem, always inert | a human dragging it three rounds running |
|
||||||
| GR-A13 "wasted SOLVE" on a claimed Problem | a scenario that had to pick a default |
|
| GR-A13 "wasted SOLVE" on a claimed Problem | a scenario that had to pick a default |
|
||||||
| ~~GR-E01 unreachable below 5 seats~~ | arithmetic over the deal count — **withdrawn 2026-08-05, it was wrong (C1)** |
|
| GR-E01 unreachable below 5 seats | arithmetic over the deal count |
|
||||||
| ~~six provisional scenario defaults~~ | **five, and double-counted with GR-E01 (C3)** — they are reproductions, not a finding |
|
| six provisional scenario defaults | scenarios that could not be written without deciding something |
|
||||||
| GR-E03 / GR-E04 never played to the end | nobody noticed for nineteen passes |
|
| GR-E03 / GR-E04 never played to the end | nobody noticed for nineteen passes |
|
||||||
|
|
||||||
Every one was found by *building the simulator*, not by playing. That is
|
Every one was found by *building the simulator*, not by playing. That is
|
||||||
|
|
@ -57,17 +57,9 @@ register that collects opinions would reproduce it in a new medium.
|
||||||
|
|
||||||
Concretely: a scenario that fails, an arithmetic check that prints the
|
Concretely: a scenario that fails, an arithmetic check that prints the
|
||||||
contradiction, a recorded game the reader can replay, or a named test.
|
contradiction, a recorded game the reader can replay, or a named test.
|
||||||
*"This feels unbalanced"* is a note, not a finding. **The SOLVE inertness
|
*"This feels unbalanced"* is a note, not a finding. **GR-E01 is admissible
|
||||||
is admissible because a recorded session shows three no-ops.**
|
because 4/6/9 against 5/7/9 is a computation anyone can rerun; the SOLVE
|
||||||
|
inertness is admissible because a recorded session shows three no-ops.**
|
||||||
> **The example that stood here was GR-E01, and the review killed it
|
|
||||||
> (C1).** *"4/6/9 against 5/7/9 is a computation anyone can rerun"* had
|
|
||||||
> already been rerun: `2da19a4` measured **6/9/12**, and the scenario was
|
|
||||||
> renamed `-unreachable-` → `-reachable-`. The conclusion inverted — and
|
|
||||||
> it was one of the two findings that **passed** this rule. So existence
|
|
||||||
> is not what was missing. See [ADR-0012](../decisions/ADR-0012-the-design-instrument.md)
|
|
||||||
> D3: the rule gains **shape**, and **a reproduction must be able to
|
|
||||||
> fail.** Ours went green and stayed admissible.
|
|
||||||
|
|
||||||
This is what would make clay-borg a design tool rather than a suggestion
|
This is what would make clay-borg a design tool rather than a suggestion
|
||||||
box, and it is the one part of this proposal that must not be traded away
|
box, and it is the one part of this proposal that must not be traded away
|
||||||
|
|
@ -79,19 +71,19 @@ The maintainer asked whether this should extend to *"a meta about the
|
||||||
clay-borg engine evolution itself."*
|
clay-borg engine evolution itself."*
|
||||||
|
|
||||||
**My answer is no, and it should be argued rather than accepted.** That
|
**My answer is no, and it should be argued rather than accepted.** That
|
||||||
register already exists and is load-bearing: `evidence/CB-EV-*`,
|
register already exists and is load-bearing: `evidence/CB-EV-*` records
|
||||||
`decisions/ADR-*`, `gates.toml` and the workplans capture nineteen passes
|
what each pass found, `decisions/ADR-*` records what was decided and what
|
||||||
with dates, costs and falsifiers. A second register for the same subject
|
was rejected, `gates.toml` records what every control has caught and what
|
||||||
would be ceremony. The asymmetry is the point: engine evolution has a home
|
would retire it, and workplans record what was attempted. Nineteen passes
|
||||||
and game design does not.
|
of engine evolution are already captured, with dates, costs and
|
||||||
|
falsifiers.
|
||||||
|
|
||||||
> **Settled by [ADR-0012](../decisions/ADR-0012-the-design-instrument.md)
|
**Building a second register for the same subject would be ceremony**, and
|
||||||
> D7: no register — but the argument above did not survive.** C5 found the
|
this project has a standing rule that a gate must cash out. The asymmetry
|
||||||
> "third thing" the maintainer meant is visible in
|
is the point: engine evolution has a home and game design does not.
|
||||||
> `specs/InnerLoopReference.md` and `history/`'s retrospectives, **neither
|
|
||||||
> of which this inventory names.** Conclusion narrowed, not settled: if
|
If the adversarial review disagrees, that is exactly the kind of thing
|
||||||
> InnerLoopReference keeps absorbing material that is neither a decision
|
tier L exists to surface.
|
||||||
> nor a finding, revisit.
|
|
||||||
|
|
||||||
## Task: survey how this is done elsewhere, and what we already have
|
## Task: survey how this is done elsewhere, and what we already have
|
||||||
|
|
||||||
|
|
@ -128,27 +120,34 @@ and whether a ruling can *close* a finding mechanically.
|
||||||
[CB-RES-0007](../research/CB-RES-0007-design-instrument.md), with a
|
[CB-RES-0007](../research/CB-RES-0007-design-instrument.md), with a
|
||||||
runnable baseline (`tools/design-baseline.py`).
|
runnable baseline (`tools/design-baseline.py`).
|
||||||
|
|
||||||
**Its numbers were withdrawn by T02 and must not be quoted from here.**
|
**The baseline is us, and it is measured**: 6 findings across **11 files**
|
||||||
The survey reported *6 findings, 2/6 = 33% reproduced, 11 files, 4 days*.
|
with no index, **2 of 6 (33%) with a runnable reproduction**, and U1–U10
|
||||||
C2 showed the instrument counted itself and its reproduction check never
|
raised 2026-07-30, first *read* 2026-08-03 — **4 days**, 0 of 10 ruled.
|
||||||
stat'd the file; C3 showed *"six provisional defaults"* is five and GR-E01
|
|
||||||
was double-counted; T05's backfill contradicted *"six of the ten have
|
|
||||||
provisional scenarios"* — **one** does. What survives is direction: many
|
|
||||||
files, no index, 0 of 10 ruled. The first honest figures are T05's.
|
|
||||||
|
|
||||||
**Magic corrected an assumption this pass was about to build on.** Rulings
|
**The uncomfortable number is stated up front rather than left for the
|
||||||
are *"reminder information with no actual weight or rules meaning"*; the
|
review to find: the reproduction rule would reject four of our six
|
||||||
authoritative fix folds into the **Oracle** card text. **A finding closes
|
existing findings.** The survey answers it — none of the four is
|
||||||
when the source changes, not when an annotation is added** — the register
|
expensive to reproduce, so the 33% is evidence that nobody was ever asked
|
||||||
is a queue that empties. Model checkers supplied the reproduction rule
|
for one, not that the rule is unaffordable.
|
||||||
independently, and W3C's *implementation-defined* mark is machinery we
|
|
||||||
already have and must reuse rather than duplicate.
|
**Magic corrected an assumption this pass was about to build on.** I
|
||||||
|
expected a ruling to be the authoritative resolution. It is not: rulings
|
||||||
|
are *"reminder information with no actual weight or rules meaning"*, and
|
||||||
|
the authoritative fix folds into the **Oracle** card text. **A finding
|
||||||
|
closes when the source changes, not when an annotation is added** — so
|
||||||
|
the register must be a queue that empties, not an archive that grows.
|
||||||
|
That is now a constraint on T03's lifecycle.
|
||||||
|
|
||||||
|
Model checkers supplied the reproduction rule independently (a
|
||||||
|
counterexample trace *is* the finding), and W3C's *implementation-defined*
|
||||||
|
mark is the one piece of machinery we already have and must reuse rather
|
||||||
|
than duplicate.
|
||||||
|
|
||||||
## Task: adversarial review
|
## Task: adversarial review
|
||||||
|
|
||||||
```task
|
```task
|
||||||
id: CB-WP-0022-T02
|
id: CB-WP-0022-T02
|
||||||
status: done
|
status: todo
|
||||||
priority: high
|
priority: high
|
||||||
state_hub_task_id: "d2597895-fe2e-4f1e-a3db-a5fb8833434c"
|
state_hub_task_id: "d2597895-fe2e-4f1e-a3db-a5fb8833434c"
|
||||||
```
|
```
|
||||||
|
|
@ -162,42 +161,23 @@ above, and require an attempt at:
|
||||||
third thing;
|
third thing;
|
||||||
- **that "carries its reproduction" is affordable** — if half the real
|
- **that "carries its reproduction" is affordable** — if half the real
|
||||||
findings cannot be reproduced cheaply, the rule will be quietly dropped
|
findings cannot be reproduced cheaply, the rule will be quietly dropped
|
||||||
and the register becomes a suggestion box anyway. *(Hardened before the review ran: two findings had already reached
|
and the register becomes a suggestion box anyway. *(Since this task was
|
||||||
ground-game on wrong premises, so the reviewer was told to press whether
|
written the answer has hardened: two findings reached ground-game on
|
||||||
the rule is **sufficient**, not whether it is affordable.)*
|
wrong premises, and ground-game has itself ruled that reproductions are
|
||||||
|
required. The reviewer should now press the opposite way — whether the
|
||||||
|
rule is **sufficient**, given both bad premises would have shipped a
|
||||||
|
passing test alongside a wrong description.)*
|
||||||
- **that a register is needed at all**, rather than one more section in
|
- **that a register is needed at all**, rather than one more section in
|
||||||
`GroundRules.md §Underdetermined`, which already exists and already
|
`GroundRules.md §Underdetermined`, which already exists and already
|
||||||
works.
|
works.
|
||||||
|
|
||||||
Record the trail in `history/`, unpolished.
|
Record the trail in `history/`, unpolished.
|
||||||
|
|
||||||
**Done 2026-08-05.** Trail:
|
|
||||||
[challenge](../history/260805-design-instrument-challenge.md),
|
|
||||||
[response](../history/260805-design-instrument-response.md).
|
|
||||||
|
|
||||||
**Run by a separate agent** — the first in this repo that was. CB-RES-0006's
|
|
||||||
review opened by conceding it could not be, and called its own findings
|
|
||||||
*"a lower bound on what a genuinely separate reviewer would find."* That
|
|
||||||
was measurable, and this is the measurement: the separate reviewer ran
|
|
||||||
`git log` against the survey's central example and found our own commit
|
|
||||||
had falsified it four days earlier, while the author — who wrote that
|
|
||||||
commit — quoted the dead number twice.
|
|
||||||
|
|
||||||
**Seven challenges: four conceded, two conceded in part, one answered.**
|
|
||||||
**C1 changed the design** — the rule's showcase finding was false and had
|
|
||||||
*passed* the rule, so existence is not what was missing — and **caught a
|
|
||||||
defect in flight**, T06's payload still naming the dead number. C2
|
|
||||||
withdrew the baseline's precision, C3 its arithmetic, C4 flipped T03's
|
|
||||||
burden toward extending `§Underdetermined`, C5 corrected the redundancy
|
|
||||||
inventory. Survived: affordability, and reuse of the provisional
|
|
||||||
machinery. Full account:
|
|
||||||
[CB-EV-0021](../evidence/CB-EV-0021-the-design-instrument.md) §1–§3.
|
|
||||||
|
|
||||||
## Task: decide
|
## Task: decide
|
||||||
|
|
||||||
```task
|
```task
|
||||||
id: CB-WP-0022-T03
|
id: CB-WP-0022-T03
|
||||||
status: done
|
status: todo
|
||||||
priority: high
|
priority: high
|
||||||
state_hub_task_id: "a2e85810-e949-4ea0-81fc-36b912af326c"
|
state_hub_task_id: "a2e85810-e949-4ea0-81fc-36b912af326c"
|
||||||
```
|
```
|
||||||
|
|
@ -218,35 +198,11 @@ state_hub_task_id: "a2e85810-e949-4ea0-81fc-36b912af326c"
|
||||||
as a note** — and if admitted, how it is prevented from aging into an
|
as a note** — and if admitted, how it is prevented from aging into an
|
||||||
apparent finding.
|
apparent finding.
|
||||||
|
|
||||||
**Done 2026-08-05.**
|
|
||||||
[ADR-0012](../decisions/ADR-0012-the-design-instrument.md), nine
|
|
||||||
decisions. The two not on this list are the two the review forced:
|
|
||||||
|
|
||||||
- **D2 — `§Underdetermined` *is* the register; nothing parallel is built.**
|
|
||||||
Against the survey's own five benchmarks the incumbent already delivers
|
|
||||||
four, including the Oracle property the survey went to Magic to find and
|
|
||||||
we had written ourselves eight days earlier (`GroundRules.md:231-233`).
|
|
||||||
What it lacks is reproductions. So this pass **extends** a section — no
|
|
||||||
new file, no new schema.
|
|
||||||
- **D3 — admissibility is three clauses.** Exists, has the ruled shape
|
|
||||||
(row-level table, never a sum), **and can fail.** GR-E01's artifact went
|
|
||||||
green and the finding stayed admissible and stayed queued, because
|
|
||||||
nothing said a passing artifact was a signal. **A green reproduction is
|
|
||||||
an alarm.**
|
|
||||||
|
|
||||||
The rest, in one line each: **D1** INTENT gains property 4, *Instrument*,
|
|
||||||
applied with its falsifier. **D4** five kinds, each forced by an existing
|
|
||||||
finding. **D5** `applied` means the source changed; withdrawals are
|
|
||||||
reported, not deleted. **D6** notes admitted but never reportable, 30-day
|
|
||||||
expiry. **D7** no engine-evolution register, on an inventory C5 corrected.
|
|
||||||
**D8** `design-baseline.py` retired. **D9** the artifact stays here,
|
|
||||||
ground-game gets a generated file under its own workplan.
|
|
||||||
|
|
||||||
## Task: specify
|
## Task: specify
|
||||||
|
|
||||||
```task
|
```task
|
||||||
id: CB-WP-0022-T04
|
id: CB-WP-0022-T04
|
||||||
status: done
|
status: todo
|
||||||
priority: high
|
priority: high
|
||||||
state_hub_task_id: "60ddfeab-81b1-45e4-9ec9-91cc8ea7fb72"
|
state_hub_task_id: "60ddfeab-81b1-45e4-9ec9-91cc8ea7fb72"
|
||||||
```
|
```
|
||||||
|
|
@ -264,40 +220,34 @@ Candidate measures, to be argued not adopted:
|
||||||
**ground-game has ruled on what a finding must carry** (GROUND-WP-0004
|
**ground-game has ruled on what a finding must carry** (GROUND-WP-0004
|
||||||
T02, 2026-08-04), and it is stricter than this pass proposed. Adopt it:
|
T02, 2026-08-04), and it is stricter than this pass proposed. Adopt it:
|
||||||
|
|
||||||
> Arithmetic findings ship a **runnable reproduction** *and* a
|
> 1. Arithmetic findings must ship a **runnable reproduction** *and* a
|
||||||
> **row-level deal table** — never only "sum of file" or "deal depth N";
|
> **row-level deal table** listing Surface and each hidden priority
|
||||||
> and ground-game's arithmetic rulings cite that reproduction by path.
|
> separately — **never only "sum of file" or "deal depth N"**.
|
||||||
|
> 2. ground-game rulings that depend on arithmetic **should cite** that
|
||||||
|
> reproduction id or path, so a later reader can re-run it.
|
||||||
|
|
||||||
The second half is theirs to keep. **So the reproduction rule gains a
|
The second half is theirs to keep. The first is ours, and it is a direct
|
||||||
shape requirement, not just an existence one** — a finding that ships a
|
consequence of both premises we got wrong: *"12 in the file"* was a sum
|
||||||
passing test but describes the wrong quantity is still a bad finding,
|
without a deal table, and *"offered on a face-down Problem"* was a
|
||||||
which is exactly what happened twice.
|
condition named without checking which one actually fired.
|
||||||
|
|
||||||
Also specify the **trial protocol**: a trial game is a `--record`ed
|
**So the reproduction rule gains a shape requirement, not just an
|
||||||
session plus an observation log, so *"we played it and X happened"* is
|
existence one.** A finding that ships a passing test but describes the
|
||||||
replayable rather than remembered. It must cost almost nothing or it will
|
wrong quantity is still a bad finding — which is exactly what happened
|
||||||
not be done.
|
twice. The spec must say what a reproduction has to *show*, not only that
|
||||||
|
one exists.
|
||||||
|
|
||||||
**Done 2026-08-05.** [specs/GameDesign.md](../specs/GameDesign.md) v1.0 —
|
Also specify the **trial protocol**, which is the part with no precedent
|
||||||
not a register (ADR-0012 D2 put that in `§Underdetermined`).
|
here: a trial game is a `--record`ed session plus an observation log, so
|
||||||
|
*"we played it and X happened"* is replayable rather than remembered. The
|
||||||
**§1.2 is written against evidence rather than principle**: a finding must
|
engine already records sessions as scenarios; a trial is that plus notes,
|
||||||
print the rows behind any number it claims. *"12" was arithmetically
|
and it must cost almost nothing or it will not be done.
|
||||||
defensible and still wrong about the game.* **§1.3's target is `0`
|
|
||||||
reproductions gone green while open** — what GR-E01 would have tripped
|
|
||||||
four days before a human caught it. **No baseline rate is quoted.**
|
|
||||||
|
|
||||||
**The trial protocol costs one flag**: `cb-play --record` plus a sibling
|
|
||||||
`.md` in the player's own words. An observation is a **note** until it has
|
|
||||||
a reproduction — *"I felt it was too easy but then we lost"* is the case
|
|
||||||
it is shaped around, and a schema at the moment of observation would lose
|
|
||||||
it.
|
|
||||||
|
|
||||||
## Task: build it, and backfill what is already known
|
## Task: build it, and backfill what is already known
|
||||||
|
|
||||||
```task
|
```task
|
||||||
id: CB-WP-0022-T05
|
id: CB-WP-0022-T05
|
||||||
status: done
|
status: todo
|
||||||
priority: high
|
priority: high
|
||||||
state_hub_task_id: "469413ba-ed47-4840-a60e-ee7f95dc06ff"
|
state_hub_task_id: "469413ba-ed47-4840-a60e-ee7f95dc06ff"
|
||||||
```
|
```
|
||||||
|
|
@ -310,63 +260,36 @@ designing it is the point of doing it in this order.
|
||||||
`make design` (or equivalent) must report: open findings by kind, those
|
`make design` (or equivalent) must report: open findings by kind, those
|
||||||
without a reproduction, and those never reported to their owner.
|
without a reproduction, and those never reported to their owner.
|
||||||
|
|
||||||
**Done 2026-08-05.** `tools/design.py`, `make design`, and the register in
|
|
||||||
[`GroundRules.md`](../specs/GroundRules.md) — **14 rows, no new file.**
|
|
||||||
|
|
||||||
Backfill was the test. The taxonomy held (five kinds, no sixth), and it
|
|
||||||
**produced a `role` column ADR-0012 does not have**: the first report
|
|
||||||
alarmed on U2, wrongly — a green *default* is expected, a green
|
|
||||||
*counterexample* is the alarm. Folded into GameDesign §1.3. It also
|
|
||||||
contradicted the survey: **one** U-item names itself in a scenario, not
|
|
||||||
six. Detail and figures:
|
|
||||||
[CB-EV-0021](../evidence/CB-EV-0021-the-design-instrument.md) §4, §6.
|
|
||||||
|
|
||||||
## Task: report to ground-game, mechanically
|
## Task: report to ground-game, mechanically
|
||||||
|
|
||||||
```task
|
```task
|
||||||
id: CB-WP-0022-T06
|
id: CB-WP-0022-T06
|
||||||
status: done
|
status: todo
|
||||||
priority: high
|
priority: high
|
||||||
state_hub_task_id: "c6a659bd-c495-4754-9332-baadad51012a"
|
state_hub_task_id: "c6a659bd-c495-4754-9332-baadad51012a"
|
||||||
```
|
```
|
||||||
|
|
||||||
Generate the report and send it. **The message that sat unread for four
|
Generate the report and send it. **The message that sat unread for four
|
||||||
days is the baseline to beat** — the failure was not the message, it was
|
days is the baseline to beat**, and the failure was not the message — it
|
||||||
that nothing pointed at it. So the report lands as a file in `ground-game`
|
was that nothing pointed at it and nothing tracked whether it was
|
||||||
under its own workplan, extending GROUND-WP-0002 rather than duplicating
|
answered.
|
||||||
it.
|
|
||||||
|
|
||||||
Include the findings this pass has sharpened:
|
So the report must land somewhere that persists: a file in `ground-game`
|
||||||
|
under its own workplan, not only an inbox entry. GROUND-WP-0002 already
|
||||||
|
holds the ten U-items; this should extend it rather than duplicate it.
|
||||||
|
|
||||||
- **SOLVE's legality** against a face-down Problem or an unmatchable suit —
|
Include the two sharpened findings this pass has already produced:
|
||||||
and note that the case we *reported* was not the case that fired
|
|
||||||
(CB-WP-0023 T01).
|
|
||||||
- ~~**GR-E01 vs GR-S01** — 4/6/9 against 5/7/9~~ — **withdrawn
|
|
||||||
2026-08-05, before sending.** C1 found `2da19a4` had measured **6/9/12**:
|
|
||||||
the dataset reconciles them. It would have been the **fourth** wrong
|
|
||||||
premise to reach `ground-game` and is the only one caught before
|
|
||||||
transmission. **Report the withdrawal** — a claim retracted silently is
|
|
||||||
how the first three survived.
|
|
||||||
|
|
||||||
**Done 2026-08-05.**
|
- **GR-E01 vs GR-S01** — the deal count puts 4/6/9 points in play against
|
||||||
[`ground-game/workplans/GROUND-WP-0002-clay-borg-report-260805.md`](../../ground-game/workplans/GROUND-WP-0002-clay-borg-report-260805.md),
|
thresholds of 5/7/9, so either the count or the thresholds are wrong and
|
||||||
committed there, with a hub message that only *points at* the file.
|
no dataset reconciles them;
|
||||||
|
- **SOLVE's legality** against a face-down Problem or an unmatchable suit.
|
||||||
**The report asks for no ruling.** It carries GR-E01's withdrawal, our own
|
|
||||||
reproduction debt, and two notes that are explicitly not findings.
|
|
||||||
|
|
||||||
**And it acknowledged something the pass did not expect.**
|
|
||||||
GROUND-WP-0002 is `finished`: **all ten U-items were ruled 2026-08-03**,
|
|
||||||
every one confirmed as the default we simulate. CB-RES-0007 said *"0 of 10
|
|
||||||
ruled"* two days later. **The unread-inbox failure running in the opposite
|
|
||||||
direction** — they answered and we did not collect it. The instrument's
|
|
||||||
first run surfaced it.
|
|
||||||
|
|
||||||
## Task: evidence
|
## Task: evidence
|
||||||
|
|
||||||
```task
|
```task
|
||||||
id: CB-WP-0022-T07
|
id: CB-WP-0022-T07
|
||||||
status: done
|
status: todo
|
||||||
priority: high
|
priority: high
|
||||||
state_hub_task_id: "4fb77911-486d-4d9e-a53d-c4e486c74ce2"
|
state_hub_task_id: "4fb77911-486d-4d9e-a53d-c4e486c74ce2"
|
||||||
```
|
```
|
||||||
|
|
@ -381,18 +304,3 @@ this one moves rather than collides.)*
|
||||||
full-weight L pass and CB-WP-0012's deleted its own structural trigger.
|
full-weight L pass and CB-WP-0012's deleted its own structural trigger.
|
||||||
- **The engine-evolution question**, as the review left it.
|
- **The engine-evolution question**, as the review left it.
|
||||||
- **Quote CB-WP-0021's cost by re-running the instrument.**
|
- **Quote CB-WP-0021's cost by re-running the instrument.**
|
||||||
|
|
||||||
**Done 2026-08-05.**
|
|
||||||
[CB-EV-0021](../evidence/CB-EV-0021-the-design-instrument.md).
|
|
||||||
|
|
||||||
**Backfill did change the design** — and the honest answer to *"be
|
|
||||||
suspicious if all six fit"* is that only **five** were entered (one was a
|
|
||||||
double-count), so fitting them is close to circular. The taxonomy's real
|
|
||||||
test is the seventh finding.
|
|
||||||
|
|
||||||
**Tier L's cost against what it caught**: four of six catches came only
|
|
||||||
from the separate reviewer, and **two came from execution rather than
|
|
||||||
process** — the `role` distinction from building it, the ten uncollected
|
|
||||||
rulings from running it. That is InnerLoop §Design goal's prediction
|
|
||||||
holding, and an argument against front-loading more review rather than
|
|
||||||
less.
|
|
||||||
|
|
|
||||||
|
|
@ -3,7 +3,6 @@ id: CB-WP-0024
|
||||||
kind: product
|
kind: product
|
||||||
title: "The table you can watch: the piles, the other seats' moves, the score that carries"
|
title: "The table you can watch: the piles, the other seats' moves, the score that carries"
|
||||||
status: ready
|
status: ready
|
||||||
state_hub_workstream_id: "5f17b6f9-cd4b-4c31-a8af-712313149cf2"
|
|
||||||
---
|
---
|
||||||
|
|
||||||
# Purpose
|
# Purpose
|
||||||
|
|
@ -49,7 +48,6 @@ except where task T04 says so explicitly.
|
||||||
id: CB-WP-0024-T01
|
id: CB-WP-0024-T01
|
||||||
status: todo
|
status: todo
|
||||||
priority: high
|
priority: high
|
||||||
state_hub_task_id: "acb4231c-35df-490f-93fd-be71c0abf1dc"
|
|
||||||
```
|
```
|
||||||
|
|
||||||
Two defects behind one button.
|
Two defects behind one button.
|
||||||
|
|
@ -81,7 +79,6 @@ because a control that can no longer work must not look like it can.
|
||||||
id: CB-WP-0024-T02
|
id: CB-WP-0024-T02
|
||||||
status: todo
|
status: todo
|
||||||
priority: high
|
priority: high
|
||||||
state_hub_task_id: "37b16a87-e791-4c85-9f60-671812f6dd2e"
|
|
||||||
```
|
```
|
||||||
|
|
||||||
Draw the draw stack and the discard stack as stacks, with their counts on
|
Draw the draw stack and the discard stack as stacks, with their counts on
|
||||||
|
|
@ -113,7 +110,6 @@ stop.
|
||||||
id: CB-WP-0024-T03
|
id: CB-WP-0024-T03
|
||||||
status: todo
|
status: todo
|
||||||
priority: high
|
priority: high
|
||||||
state_hub_task_id: "7b0b08d9-cf0d-40f4-8697-049bdb20085d"
|
|
||||||
```
|
```
|
||||||
|
|
||||||
Each seat's revealed selection should be legible as a played card in that
|
Each seat's revealed selection should be legible as a played card in that
|
||||||
|
|
@ -143,7 +139,6 @@ face-down as one identical back.
|
||||||
id: CB-WP-0024-T04
|
id: CB-WP-0024-T04
|
||||||
status: todo
|
status: todo
|
||||||
priority: medium
|
priority: medium
|
||||||
state_hub_task_id: "4f312267-2b82-4cf9-bbb1-7f287dc9e766"
|
|
||||||
```
|
```
|
||||||
|
|
||||||
`play` (`table.rs:219`) already owns the session — one listener, a seed
|
`play` (`table.rs:219`) already owns the session — one listener, a seed
|
||||||
|
|
@ -173,7 +168,6 @@ raise the question rather than silently canonising a choice.
|
||||||
id: CB-WP-0024-T05
|
id: CB-WP-0024-T05
|
||||||
status: todo
|
status: todo
|
||||||
priority: high
|
priority: high
|
||||||
state_hub_task_id: "db176595-e00d-467c-ba54-0714cfb7ba5c"
|
|
||||||
```
|
```
|
||||||
|
|
||||||
`evidence/CB-EV-0022-*.md`. Short — this is an S pass.
|
`evidence/CB-EV-0022-*.md`. Short — this is an S pass.
|
||||||
|
|
|
||||||
|
|
@ -3,7 +3,6 @@ id: CB-WP-0025
|
||||||
kind: product
|
kind: product
|
||||||
title: "Could we have won: a path out of a lost game, and how hard the game actually is"
|
title: "Could we have won: a path out of a lost game, and how hard the game actually is"
|
||||||
status: ready
|
status: ready
|
||||||
state_hub_workstream_id: "a866982f-94e7-432b-b0d5-2fefb781a574"
|
|
||||||
---
|
---
|
||||||
|
|
||||||
# Purpose
|
# Purpose
|
||||||
|
|
@ -92,7 +91,6 @@ whose meaning drifts the next time a bot improves.
|
||||||
id: CB-WP-0025-T01
|
id: CB-WP-0025-T01
|
||||||
status: todo
|
status: todo
|
||||||
priority: high
|
priority: high
|
||||||
state_hub_task_id: "556cfd24-5992-4cc0-be90-0b60e989b5bb"
|
|
||||||
```
|
```
|
||||||
|
|
||||||
`research/CB-RES-0008-*.md`, with `tier: L` and the chaos roll recorded
|
`research/CB-RES-0008-*.md`, with `tier: L` and the chaos roll recorded
|
||||||
|
|
@ -129,7 +127,6 @@ per dimension — a number or a reproducible comparison, not an impression.
|
||||||
id: CB-WP-0025-T02
|
id: CB-WP-0025-T02
|
||||||
status: todo
|
status: todo
|
||||||
priority: high
|
priority: high
|
||||||
state_hub_task_id: "bdd9fc97-4706-4d8e-b286-91ff53339680"
|
|
||||||
```
|
```
|
||||||
|
|
||||||
Tier L requires it. Exactly one round: challenge, then response, trail in
|
Tier L requires it. Exactly one round: challenge, then response, trail in
|
||||||
|
|
@ -157,7 +154,6 @@ Tier L requires it. Exactly one round: challenge, then response, trail in
|
||||||
id: CB-WP-0025-T03
|
id: CB-WP-0025-T03
|
||||||
status: todo
|
status: todo
|
||||||
priority: high
|
priority: high
|
||||||
state_hub_task_id: "055630a1-56e0-4ad4-8796-35f444313371"
|
|
||||||
```
|
```
|
||||||
|
|
||||||
`decisions/ADR-0013-*.md`. (ADR-0012 is CB-WP-0022 T03's.) At minimum:
|
`decisions/ADR-0013-*.md`. (ADR-0012 is CB-WP-0022 T03's.) At minimum:
|
||||||
|
|
@ -184,7 +180,6 @@ state_hub_task_id: "055630a1-56e0-4ad4-8796-35f444313371"
|
||||||
id: CB-WP-0025-T04
|
id: CB-WP-0025-T04
|
||||||
status: todo
|
status: todo
|
||||||
priority: high
|
priority: high
|
||||||
state_hub_task_id: "7c3ca25d-0001-44ef-a07d-0ab2d51b09a5"
|
|
||||||
```
|
```
|
||||||
|
|
||||||
`specs/` — extend `MetricsAndScenarios.md` or add a capability spec as the
|
`specs/` — extend `MetricsAndScenarios.md` or add a capability spec as the
|
||||||
|
|
@ -214,7 +209,6 @@ kind that has already gone wrong twice.
|
||||||
id: CB-WP-0025-T05
|
id: CB-WP-0025-T05
|
||||||
status: todo
|
status: todo
|
||||||
priority: high
|
priority: high
|
||||||
state_hub_task_id: "5f1cbda3-7427-4f23-8c67-b2e216b987a8"
|
|
||||||
```
|
```
|
||||||
|
|
||||||
The search, the bound, and the replayable path. Wire it to the ending page
|
The search, the bound, and the replayable path. Wire it to the ending page
|
||||||
|
|
@ -239,7 +233,6 @@ a loss.
|
||||||
id: CB-WP-0025-T06
|
id: CB-WP-0025-T06
|
||||||
status: todo
|
status: todo
|
||||||
priority: high
|
priority: high
|
||||||
state_hub_task_id: "723feae8-1ed4-4dd4-9330-8ec84625e5d9"
|
|
||||||
```
|
```
|
||||||
|
|
||||||
Run the measurement, ship it as a `make` target beside the other
|
Run the measurement, ship it as a `make` target beside the other
|
||||||
|
|
@ -265,7 +258,6 @@ message that sat unread for four days is the baseline to beat*.
|
||||||
id: CB-WP-0025-T07
|
id: CB-WP-0025-T07
|
||||||
status: todo
|
status: todo
|
||||||
priority: high
|
priority: high
|
||||||
state_hub_task_id: "2616e06c-5d88-4680-be18-4452515dff21"
|
|
||||||
```
|
```
|
||||||
|
|
||||||
`evidence/CB-EV-0023-*.md`.
|
`evidence/CB-EV-0023-*.md`.
|
||||||
|
|
|
||||||
Loading…
Add table
Add a link
Reference in a new issue