The question was whether M-D1-MUT is a real instrument or a name-counter with extra steps, given that writing a weak mutation is as easy as writing a strong one. It is real, but only because it was hardened three times in one pass. Five controls now stand between a mutation and a red verdict — the mutation must apply, the baseline must be green, the tree must be restored and verified, the failure must match a stated reason, and that stated reason must be absent from passing output — and every one of them exists because its failure actually occurred. The last is the sharpest: the FA guard needed a guard, because my first AM-2 expect was "AM-2", which the passing report contains. Generalizable: an instrument that measures whether other instruments work needs more controls than the instruments it measures. M-D1-MUT carries five; dep-weight and rule-coverage carry one each. That asymmetry is the cost of a meta-instrument, and a project adding one should budget for it. A worse failure mode than CB-WP-0005 predicted: a mutation can become weak without anyone touching it. AM-6's went SURVIVED when T04 moved its gate from debug to release — nothing about the row, the mutation or the code changed, only the headroom. Mutation strength is coupled to measurement conditions, so a mutation is not a write-once artifact. CB-WP-0005 T08's stronger remedy — mutations written by someone other than the author — was NOT tested and should not be assumed unnecessary: EXPECT-VACUOUS covers the cheap failure, not the expensive one AM-6 demonstrated. The "removes the manual path" test is settled as a predictor of cost, not of worth. mutation-check fails it outright and produced six defects nothing else would have found. Prediction error collapsed: 4-5x, then 2.5x, now small — because this pass predicted per task, as a mechanism, with the alternative named. Both branches are outcomes someone must defend, so the prediction cannot be dodged. AM-3 and AM-4c took the second branch and are better resolved for it than if a number had been forced. No InnerLoop change. v1.4's mutation requirement is one pass old and changing it before a second use would be the invention-in-isolation INTENT warns about — the same argument used to amend K14 four hours earlier. Named next candidate: specs/SessionShape.md. SS-01..SS-05 have been stated since CB-WP-0003 and none has ever been enforced. This pass ran at 2.5x the context ceiling its own spec sets and nothing said a word. CB-WP-0006 status -> done, 9/9. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|---|---|---|
| .forgejo/workflows | ||
| benchmarks | ||
| crates | ||
| decisions | ||
| evidence | ||
| games/ground | ||
| history | ||
| research | ||
| scenarios | ||
| specs | ||
| tools | ||
| workplans | ||
| .custodian-brief.md | ||
| .gitignore | ||
| Cargo.lock | ||
| Cargo.toml | ||
| clippy.toml | ||
| facts.toml | ||
| INTENT.md | ||
| LICENSE | ||
| Makefile | ||
| README.md | ||
| rust-toolchain.toml | ||
| WORK-RECORDS.md | ||
clay-borg
A rebuild from scratch simulation and games engine framework set up to assimilate and optimize techniques and implementations useful for games, simulations, robotics.
Licensed under the Target Revenue Source License (TRSL V1C1) — see LICENSE; canonical text lives in the org's target-revenue repository.
Running the gates
make all # every gate, from a clean shell
make -C /path/to/clay-borg all # …or from any other directory
There is no environment setup step. No cd, no export PATH, no
activation script. make locates the repo from its own path and cargo
from the standard rustup locations; the Python tools do the same via
tools/repo.py. The only prerequisites are a rustup toolchain and Python
3.11+.
This is deliberate and enforced: make env-test runs every tool from /
with a PATH containing no cargo, and make all includes it. CB-RES-0003
measured 84 agent turns and $15.33 spent prefixing commands with cd and
export PATH before that friction was fixed at the root (CB-WP-0004 T01).
Other useful targets: make cost (spend per task), make cost-budget
(spend since the last commit), make cost-mix (mechanical vs judgment
turns), make loop-lint (executable InnerLoop rules), make self-tests
(every tool's positive control).
GROUND
The first product vertical is a virtual tabletop implementation of GROUND — A Game of Bonds and Rivalry: DARVO Edition. The boardgame itself (rules, editions, content) is at home in the sister repository ground-game — that repo is authoritative for what GROUND is; clay-borg implements the engine that runs it.