From 1c3019c7e81e7fd6c156ebee29474dbb018c8262 Mon Sep 17 00:00:00 2001 From: tegwick Date: Fri, 31 Jul 2026 18:38:15 +0200 Subject: [PATCH] CB-WP-0006 T02: instrument AM-2; report AM-3 blocked, with the argument MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Both rows were unmutatable for the same stated reason. They resolved differently, and the difference is the point. AM-2 is instrumented and enforced — tools/size-metrics.py, in `make all`: AM-2: 27.2 LOC/rule [ok target <= 40] (1.47x headroom) 1,575 impl lines / 58 rules Tests are excluded because AM-2 asks what a rule costs, not how much it is exercised; lib.rs is ~18% test code and including it would have flattered the number. This matters because AM-2 is AM-1's anti-gaming pair: 100% rule coverage means nothing if the rules are trivially small, and AM-1 has been reported met since CB-WP-0001 with its pair uninstrumented. Verified red by a property mutation — ~800 lines of filler injected into the impl, pushing the ratio past 40 — not a threshold tweak. The expect string is the precise failure signature "FAIL target <= 40"; my first attempt used "AM-2", which also matches passing output and would have made the FA guard vacuous. AM-3 is BLOCKED, not uninstrumented, and that is a finding rather than a deferral. It measures LOC to express the CB-RES-0001 synthetic game on our kernel, against a boardgame.io baseline of ~36 LOC for a declarative 3p commit/reveal game object. That artifact has never been built: games/ contains only ground, and benches/synthetic.rs drives GROUND rather than defining a synthetic game. Measuring GROUND's 1,575 impl lines against a 36-line synthetic game object would compare two different games and call the difference a D1 result. So the tool ships the measurement — a marker-delimited region, self-tested — and reports the row blocked, naming the missing artifact. A number would have been worse than a blank. It stays unmutatable and still counts against M-D1-MUT per ADR-0005 §1: a row that cannot fail asserts nothing, however good the reason. M-D1-MUT: 5 -> 6 of 14. Co-Authored-By: Claude Opus 5 --- Makefile | 9 +- facts.toml | 4 +- tools/__pycache__/cb-cost.cpython-312.pyc | Bin 37064 -> 37064 bytes tools/__pycache__/dep-weight.cpython-312.pyc | Bin 9295 -> 9295 bytes .../mutation-check.cpython-312.pyc | Bin 19987 -> 20597 bytes tools/mutation-check.py | 29 ++- tools/size-metrics.py | 230 ++++++++++++++++++ workplans/CB-WP-0006-instrument-the-table.md | 39 +++ 8 files changed, 301 insertions(+), 10 deletions(-) create mode 100644 tools/size-metrics.py diff --git a/Makefile b/Makefile index 17f3117..55a76bd 100644 --- a/Makefile +++ b/Makefile @@ -24,7 +24,7 @@ TOOLS := $(REPO)/tools # Every cargo recipe runs at the repo root; the shell does not persist cd. IN_REPO := cd $(REPO) && -.PHONY: check test sim bench bench-test coverage dep-weight cost cost-test cost-pin cost-budget cost-mix loop-lint self-tests env-test task-done status facts-check facts-gen mutation-check loc all +.PHONY: check test sim bench bench-test coverage dep-weight cost cost-test cost-pin cost-budget cost-mix loop-lint self-tests env-test task-done status facts-check facts-gen mutation-check size-metrics loc all ## fmt + clippy (deny warnings) + HashMap deny-lint check: @@ -42,6 +42,10 @@ dep-weight: coverage: $(PY) $(TOOLS)/rule-coverage.py +# AM-2 (M-D1-SPL, AM-1's anti-gaming pair) and AM-3. +size-metrics: + $(PY) $(TOOLS)/size-metrics.py + # M-D2-CST (specs/CostAccounting.md). cost-test is the positive control and # runs first: a cost number from an unverified collector is void. cost: cost-test @@ -65,6 +69,7 @@ self-tests: $(PY) $(TOOLS)/status.py --self-test $(PY) $(TOOLS)/facts.py --self-test $(PY) $(TOOLS)/mutation-check.py --self-test + $(PY) $(TOOLS)/size-metrics.py --self-test # T01 positive control: prove the environment fix, do not assume it. Runs # every tool from a foreign working directory with a PATH that has no @@ -142,4 +147,4 @@ loc: printf '%-28s %s\n' $$d "$$(find $$d/src -name '*.rs' | xargs cat | grep -vcE '^\s*(//|$$)')"; \ done -all: check test sim coverage dep-weight self-tests env-test facts-check loop-lint bench-test +all: check test sim coverage size-metrics dep-weight self-tests env-test facts-check loop-lint bench-test diff --git a/facts.toml b/facts.toml index f14f17f..1f0bc28 100644 --- a/facts.toml +++ b/facts.toml @@ -40,8 +40,8 @@ fmt = "{:,}" by = "tools/mutation-check.py" [am_unmutatable] -value = 7 -text = "7" +value = 6 +text = "6" fmt = "{:,}" by = "tools/mutation-check.py" diff --git a/tools/__pycache__/cb-cost.cpython-312.pyc b/tools/__pycache__/cb-cost.cpython-312.pyc index 03aa62eab3234a2851f79ba519a0abdb5a293682..2d6315cdaa4df74b631cd1dfb92c8f33c1ca5bd7 100644 GIT binary patch delta 21 bcmX@Hkm6n=LGre*jklmUuAa9XG^I5W(YLZy^aehRHZYf6cln$SD%-g&(8=Do?i zcR;4vacNvx$GCD~y3q(5eR1WcaiJzA4Y(pMtZ|{Hae)yxt~~F}RBa*gF7D*rbH4AK z?>lEoACNykBn{uz*M~Jc5AInv7B79;uu)55r7vn*XM*g|M&mlZc}G=#F~5l&X3ZOs_28aJ3u{H*whHf&jq3HPJvr^{3Ct-zOnP$P zq^H=^tYhC2oebE~eIv(MmkO23X_}$2?mK09$JsM`WfW0I59@sd5Ul|E5d7>NvJu-y zdz>X6!Sq)!NtR+KjKIE_0d|rNKGyOzzfm^CPO;}63zZ|o>@*wMH*$uJqJd}Gm=W-Y zb?#pWjfQzGI_`c`{}m~=D95p6zcdRxp5;oIx=iHO{%SL5mdk!aIC-1(OQH8$W;!8c z0nBX9f^owT+@Oj>#t|^%XU!|``(~jzW5XiNo>PH6mEtKjey7E3Kn03QAm!2 zO~Ab)AaboSx{6MNN{EpP9{rtBdw*L#=c^HEud)=vzI4dI6oBw+fk6W zOWT{eI$OK%^(q#_+)qWl1l9`h}B+IGH3+@p? zFGZAC<|59xWoAu}8XYJiT@e_GCXa* z9ji{>VqT#f6=>F~gA>bSxbV#Pz3@fr)Uqn$^n*oP=h*I43jLRPOC?}l;npEf-o{}( zPkGBOb)`AAEl!Ha!~Nbg%&B$sqJl0hME>oDY-XKd{{yeF<5d| z{UI5S!~7Nmu1DCh6Hy`#;h(YeUE?e{TIy`dk)!TAO&0?PL?=EMCQHMSp;BMu{nA!s zw)+@9i7xNz_O9c2eP4^aW6`z6R5H074T?^8yhR2`m-~6!tq{3iI@xjJ<-=M{-CNU} zk+JROmd*C$=8@D^b80&h-;NAzNBTD7=k5+{#izF-)7!11zXwUn*`4ZOWNhc47U|io z57qSUHdY8mcLH_o7s+n0zUFzd+u`2sm?hzvstbX-(2jN?;QrF-1jvu>OUFJwKz{qI z?{+PUd3SK3ioELnI{ZU5$+};UjSwQ;;CP6T6*oL_pOAw4U@{jV;9i^lnGn;RxG+;i WK5^fkJxf9_21w|SOzGa-`M&`n+R()S delta 1030 zcmZuvOH5Ni6n(da3N0<=V=JOE{Hf4V5GjhEKmi3RYWzeK6M6Jb9#7i$nt87fh_PE2 zZp=$ej3yZ4LPIw?D`VWcF-F&#=*|UUL8NY6Ipq@|d9yf^x#yfaXKv>23$UJp@3YtI zmc+XHC3Wx5_*-Az17+3+&+QzR<;z!`MM{#)(yd*R(2X9fuw}gHyDwp-0!u=S{(R|5 z$(C!6y9%rKXliUVwRjK%`JCP5K|F+YwwE8qdegNNQjze;%dHqk@mRiLWjp?!AOsut z>6^Co&BELA67sD%8&ey$?=y95n?iUTI~7O1)yC6>VT{;D?#3wg*fPC%0%Lu))Jg0Y z1B~H-;@BMXsofDOzH3tGG+*<6g4`obC#q(cESsj9Ns@XxlQpS9R5L=R)AAf8DLpw& zGECJ>gCs7;iL4>fbaPhK6oTr!ib6fJNLV&Sg*_{p=xno(%micVq%S&1;=|$5o8f3Q z+Cvyk>I@AsbSd0Rs%3*@HAXW!Gbu(wz$&e*zc;p&+m|asuFq0pENbQ~HPs|p(Aiu{ zmyxIjnUxtPb)@2;S<6*{z7B&@GTAc_+Kj4h#}-PB&QnIxR5r3q$aiC6#OkU^K!ded zlYs`lS~Kpb{})u`G&Q;vCf9Xi+6jX(YcV5wThax9Bksxf>RD?9Hl)w9KFf3X-VxcLT2fSWoOpu)8-EdU>H(jD-Xzi)Yc0Dkb9 zuE!qm^Fr)K2~6^dGe1k=4i60v0?cq4cL7kI8u 100, f"{loc:,} LOC") + check("real rule count measures non-zero", rules > 10, f"{rules} rules") + + # The test region must actually be excluded — AM-2 asks what a rule + # costs, not how much it is exercised, and lib.rs is ~30% tests. + full = code_lines(open(RULES_SRC).read().split("\n")) + check("the #[cfg(test)] region is excluded", loc < full, + f"{loc:,} impl vs {full:,} whole file") + + # AM-3's marker mechanism must work, so that a future artifact is + # measured automatically rather than needing this tool changed. + import tempfile + with tempfile.TemporaryDirectory() as d: + os.makedirs(os.path.join(d, "games")) + p = os.path.join(d, "games", "x.rs") + with open(p, "w") as fh: + fh.write(f"pre\n{AM3_BEGIN}\nlet a = 1;\n// c\n\nlet b = 2;\n" + f"{AM3_END}\npost\n") + here = os.getcwd() + try: + os.chdir(d) + got = am3_region() + check("AM-3 markers delimit a region and count only code", + got is not None and code_lines(got[1]) == 2, + f"{code_lines(got[1]) if got else 'none'} lines") + finally: + os.chdir(here) + check("AM-3 reports blocked when no region is marked", + am3_region() is None, + "the synthetic game has not been built; a number here would be wrong") + + print("size-metrics self-test (positive control)") + ok = True + for name, passed, det in results: + print(f" [{'ok ' if passed else 'FAIL'}] {name}" + + (f" — {det}" if det else "")) + ok &= passed + return 0 if ok else 1 + + +def main(): + enter_root() + if "--self-test" in sys.argv: + return self_test() + return report() + + +if __name__ == "__main__": + sys.exit(main()) diff --git a/workplans/CB-WP-0006-instrument-the-table.md b/workplans/CB-WP-0006-instrument-the-table.md index 6e5fc90..7bd1dd3 100644 --- a/workplans/CB-WP-0006-instrument-the-table.md +++ b/workplans/CB-WP-0006-instrument-the-table.md @@ -124,6 +124,45 @@ AM-3 depends on K18 (benches driven from scenario files); if K18 resolves toward amending the spec rather than implementing it, AM-3 must be restated or withdrawn with an argument rather than left unmeasured. +**Delivered — but the two rows resolved differently, and the difference is +the point.** + +**AM-2 is instrumented and enforced.** `tools/size-metrics.py` + +`make size-metrics`, in `make all`: + +```text +AM-2: 27.2 LOC/rule [ok target <= 40] (1.47x headroom) + 1,575 code lines before the first #[cfg(test)] / 58 numbered rules +``` + +Tests are excluded because AM-2 asks what a rule *costs*, not how much it +is exercised — `lib.rs` is ~18% test code and including it would have +flattered the number. Verified red by a **property** mutation: ~800 lines +of filler injected into the impl, pushing the ratio past 40. `expect` is +the precise failure signature `FAIL target <= 40`, not the row name, which +would have matched passing output too. + +**AM-3 is BLOCKED, not uninstrumented — and this is a finding, not a +deferral.** The row measures "LOC to express the CB-RES-0001 synthetic +game on our kernel" against a boardgame.io baseline of ~36 LOC for a +declarative **3p commit/reveal game object**. That artifact has never been +built: `games/` contains only `ground`, and `benches/synthetic.rs` +*drives* GROUND rather than *defining* a synthetic game. + +Measuring GROUND's 1,575 impl lines against a 36-line synthetic game +object would compare **two different games** and call the difference a D1 +result. So the tool ships the measurement mechanism — a marker-delimited +`// AM-3:BEGIN` / `// AM-3:END` region, self-tested — and **reports the row +blocked, naming the missing artifact**. A number here would have been +worse than a blank. + +It therefore stays `unmutatable` and **still counts against M-D1-MUT**, per +ADR-0005 §1: a row that cannot fail asserts nothing, however good the +reason. Resolving it needs an artifact, not a metric tweak — carried +forward, not silently dropped. + +**M-D1-MUT: 5 → 6 of 14.** + ## Task: AM-5, AM-9 — measure or withdraw, but stop leaving them blank ```task