clay-borg/Makefile
tegwick ff89288581
Some checks failed
ci / check (push) Failing after 4s
CB-WP-0042 T05: H2 measured — it largely succeeds where H1 failed
Measured against ground-game's own §3 criteria, read from their design
note rather than reused from H1.

Criterion 1 met with room: greedy SHARED wins 120 at 3p and 175 at 4p,
against H1's 0 and 0, restoring 73% and 92% of baseline.

Criterion 3 met, and it was H1's clearest failure. Under H1 the
unregulated seat armed DARVO constantly and never won; under H2 it arms
and wins 13/13/57. "Non-zero for some policy that still sometimes wins" is
exactly the shape H1 could not produce.

Criterion 2 met at 2p/4p/6p and missed at 3p — 1.50 against baseline's
1.57 — reported as a miss because that is what this sample says. The
mechanism is visible: H1-greedy's spread is 0.00 at 3p+, because a flat
tax on every seat creates no variance at all. That is the clearest
statement of why scoping was the right correction.

Criterion 5 is the best evidence in the pass. Forcing every scope to
global and changing nothing else reproduces H1's collapse exactly — 120 to
0 at 3p, 175 to 0 at 4p — so the scoping is what saves it, not any other
difference between the packages.

Criterion 4 came out backwards and the prediction held. The workplan said
this panel might be unable to test it, because no policy here models
another seat or knows what a scope is; bond claim rates are LOWER than
personal at 3p and 4p, driven by suit availability rather than incentive.
Reported as untested with an incidental figure pointing the wrong way,
not as a refutation.

Wired into make panels. First pass declared after ADR-0021, so no chaos
roll is recorded.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-08 16:09:16 +02:00

304 lines
13 KiB
Makefile

# One command surface (InnerLoop §agentic-efficiency #3). Deterministic,
# greppable output; precursor of the `cb` CLI.
#
# CB-WP-0004 T01: every target here runs from a clean shell, from any
# directory, with no prefix. Invoke as `make -C <repo> <target>` from
# elsewhere. No target requires `cd` or `export PATH` — CB-RES-0003
# measured 84 turns and $15.33 spent on exactly those two prefixes.
# Absolute path to this Makefile's directory, so recipes never depend on
# the caller's working directory.
REPO := $(patsubst %/,%,$(dir $(abspath $(lastword $(MAKEFILE_LIST)))))
# Locate cargo instead of requiring it on the inherited PATH. Mirrors
# tools/repo.py:cargo_bin() — kept in sync by `make env-test`.
CARGO := $(firstword $(shell command -v cargo 2>/dev/null) \
$(wildcard $(HOME)/.cargo/bin/cargo) \
$(wildcard /usr/local/cargo/bin/cargo) \
cargo)
export PATH := $(dir $(CARGO)):$(PATH)
PY := python3
TOOLS := $(REPO)/tools
# `make` with no target lists what there is, rather than running the
# heaviest thing in the file. The Makefile's own header calls itself "one
# command surface" -- a surface you have to read the source of is not one.
.DEFAULT_GOAL := help
# A trial's name. Timestamped so two sessions on one day cannot overwrite
# each other's notes, and prefixed with the date because `tools/trials.py`
# reads the age from there (falling back to mtime).
TRIAL_NAME := $(shell date +%Y-%m-%d-%H%M)$(if $(SLUG),-$(SLUG),)
PLAYERS ?= 3
PORT ?= 0
# Every cargo recipe runs at the repo root; the shell does not persist cd.
IN_REPO := cd $(REPO) &&
.PHONY: help ground check test sim bench bench-test coverage dep-weight cost cost-test cost-pin cost-budget shape-budget cost-mix loop-lint self-tests env-test task-done status facts-check facts-gen mutation-check size-metrics runtime-metrics build-time am6 am7 am8 edition-check replay-test loc play gate-review panels all \
design difficulty trials
# `design`, `difficulty` and `trials` were added by CB-WP-0022, CB-WP-0025
# and CB-WP-0027 and none was declared here. Only `trials` revealed it, by
# colliding with the trials/ DIRECTORY -- Make saw an up-to-date file and
# ran nothing. The other two work by luck: no file happens to share their
# name. A target that is a command, not a file, belongs on this line.
## list every target, with what it does
## `make help` is also what bare `make` runs.
help:
@echo "clay-borg — make <target> (bare \`make\` shows this)"
@echo
@awk '/^## /{ sub(/^## /,""); d[++n]=$$0; next } \
/^[a-zA-Z][a-zA-Z0-9_-]*:/{ \
if (n) { split($$1,a,":"); printf " %-16s %s\n", a[1], d[1]; \
for (i=2;i<=n;i++) printf " %-16s %s\n", "", d[i]; print "" } \
n=0; next } \
{ n=0 }' $(MAKEFILE_LIST)
@echo " Variables: PLAYERS=3 PORT=0 SLUG=<name> ARGS=\"...\""
@echo
@echo " Undocumented (instruments and gate internals):"
@awk '/^## /{ d=1; next } \
/^[a-zA-Z][a-zA-Z0-9_-]*:/{ if (!d) { split($$1,a,":"); print a[1] } d=0; next } \
{ d=0 }' $(MAKEFILE_LIST) | sort -u | tr "\n" " " | fold -s -w 66 | sed "s/^/ /"
@echo
## play GROUND in a browser, recording a trial (the usual way in)
## Opens a URL; game on the left, notes on the right. Everything you
## type in the notes panel is bound to the position you typed it at.
## `make ground PLAYERS=2 SLUG=darvo-confusion`
## Read the notes back afterwards with `make trials`.
ground:
@mkdir -p $(REPO)/trials
@echo " trial: trials/$(TRIAL_NAME).md (notes) + .yaml (the game)"
$(IN_REPO) $(CARGO) run -q -p cb-play -- \
--players $(PLAYERS) --serve $(PORT) \
--record trials/$(TRIAL_NAME).yaml \
--trial trials/$(TRIAL_NAME).md $(ARGS)
## fmt + clippy (deny warnings) + HashMap deny-lint
check:
$(IN_REPO) $(CARGO) fmt --all --check
$(IN_REPO) $(CARGO) clippy --workspace --all-targets -- -D warnings
## play GROUND from the terminal (INTENT stage 0's CLI player).
## `make play ARGS="--all-bots --bot random"` to watch one instead.
play:
$(IN_REPO) $(CARGO) run -q -p cb-play -- $(ARGS)
## which control gates are due for a keep-or-kill argument (ADR-0006 D3)
gate-review:
$(PY) $(TOOLS)/gate-review.py
## unit + scenario-format tests
test:
$(IN_REPO) $(CARGO) test --workspace
## AM-4 dependency weight: third-party lines behind each budget
dep-weight:
$(PY) $(TOOLS)/dep-weight.py
## AM-1 rule coverage: which GROUND rules a scenario exercises
coverage:
$(PY) $(TOOLS)/rule-coverage.py
# K10: a bundle must re-execute, and must be able to fail. Four controls
# from ADR-0005 §6, including truncate-by-one-byte and mutated-seed.
replay-test:
$(PY) $(TOOLS)/replay-test.py
# AM-6 gate. Runs in RELEASE, where headroom is ~24x; the same assertion
# in debug has ~3.4x and flaked under load. The target is unchanged — only
# where it is measured.
am6:
$(IN_REPO) $(CARGO) test --release -p games-ground --all-features \
am6_throughput -- --ignored --nocapture --test-threads=1
# ADR-0011 D2: is the vendored edition still what ground-game published?
# Reports "upstream not checked out" as its own outcome, never a pass.
edition-check:
$(PY) $(TOOLS)/edition-check.py
# AM-8 N=10 determinism gate. One scenario, ten same-seed replays, all
# compared to the first. Not all 25: `make sim` already runs K8's double-
# run over every scenario, and repeating that eight more times costs 47 s
# per build to re-answer a question the second run already answered. The
# extra runs exist for the probabilistic class, and one workload gives
# that class its ten samples — see scenario::run_n.
am8:
$(IN_REPO) $(CARGO) run -q -p cb-sim -- --runs 10 \
$(REPO)/scenarios/ground/gr-r06-round-resolve.yaml
# AM-7 scaling gate. Same release / single-thread reasoning as am6 and more
# so: a *ratio* of two timings taken under varying contention is worse than
# one reading, because the noise multiplies rather than cancels.
am7:
$(IN_REPO) $(CARGO) test --release -p games-ground --all-features \
am7_cost_per_event -- --ignored --nocapture --test-threads=1
# AM-9 peak RSS (fast, gated). AM-5 needs a clean build — see build-time.
runtime-metrics:
$(PY) $(TOOLS)/runtime-metrics.py --fast
# AM-5: clean release build, ~90s, into a throwaway CARGO_TARGET_DIR so the
# working cache survives. Reported not gated per GameKernel §5.
build-time:
$(PY) $(TOOLS)/runtime-metrics.py
# AM-2 (M-D1-SPL, AM-1's anti-gaming pair) and AM-3.
size-metrics:
$(PY) $(TOOLS)/size-metrics.py
# M-D2-CST (specs/CostAccounting.md). cost-test is the positive control and
# runs first: a cost number from an unverified collector is void.
cost: cost-test
$(PY) $(TOOLS)/cb-cost.py --composition --by-task
cost-test:
$(PY) $(TOOLS)/cb-cost.py --self-test
# InnerLoop rules that are mechanically checkable (CB-WP-0003 T01).
loop-lint:
$(PY) $(TOOLS)/loop-lint.py
# Positive control for every reporting tool, per InnerLoop v1.1 Step 5.
## positive control for every reporting tool
self-tests:
$(PY) $(TOOLS)/cb-cost.py --self-test
$(PY) $(TOOLS)/loop-lint.py --self-test
$(PY) $(TOOLS)/rule-coverage.py --self-test
$(PY) $(TOOLS)/dep-weight.py --self-test
$(PY) $(TOOLS)/repo.py --self-test
$(PY) $(TOOLS)/task-done.py --self-test
$(PY) $(TOOLS)/status.py --self-test
$(PY) $(TOOLS)/facts.py --self-test
$(PY) $(TOOLS)/mutation-check.py --self-test
$(PY) $(TOOLS)/size-metrics.py --self-test
$(PY) $(TOOLS)/runtime-metrics.py --self-test
$(PY) $(TOOLS)/replay-test.py --self-test
$(PY) $(TOOLS)/design.py --self-test
$(PY) $(TOOLS)/trials.py --self-test
cargo run --release -q -p games-ground --example difficulty -- --self-test
$(PY) $(TOOLS)/edition-check.py --self-test
# T01 positive control: prove the environment fix, do not assume it. Runs
# every tool from a foreign working directory with a PATH that has no
# cargo on it. Before T01 this failed; if it fails again, the friction is
# back and CB-EV-0003's measurement is invalid.
env-test:
@cd / && env PATH=/usr/bin:/bin $(PY) $(TOOLS)/repo.py --self-test
@cd / && env PATH=/usr/bin:/bin $(PY) $(TOOLS)/rule-coverage.py --self-test >/dev/null \
&& echo " [ok ] rule-coverage runs from / with no cargo on PATH"
@cd / && env PATH=/usr/bin:/bin $(PY) $(TOOLS)/dep-weight.py --self-test >/dev/null \
&& echo " [ok ] dep-weight runs from / with no cargo on PATH"
@cd / && env PATH=/usr/bin:/bin $(PY) $(TOOLS)/cb-cost.py --self-test >/dev/null \
&& echo " [ok ] cb-cost runs from / with no cargo on PATH"
@cd / && env PATH=/usr/bin:/bin $(PY) $(TOOLS)/loop-lint.py --self-test >/dev/null \
&& echo " [ok ] loop-lint runs from / with no cargo on PATH"
@$(MAKE) -C $(REPO) coverage >/dev/null \
&& echo " [ok ] make -C <repo> works from any directory"
# M-D1-MUT (CB-WP-0005 T02): invert each acceptance row's property and
# require the verifying command to go red. Deliberately NOT in `make all`:
# it rebuilds the workspace once per mutated row. Run it on demand and in
# CI, not in the inner loop.
mutation-check:
$(PY) $(TOOLS)/mutation-check.py $(ARGS)
# T04: single source of fact (InnerLoop v1.2) — the DFD gate.
# facts.toml is GENERATED; facts-check fails if it disagrees with the
# instruments, or if a tagged artifact disagrees with it.
facts-check:
$(PY) $(TOOLS)/facts.py --check
facts-gen:
$(PY) $(TOOLS)/facts.py --gen
# CB-WP-0027 T04: what the players said, and where. Surfacing is the
# deliverable -- a commentary feature nobody can read is this project's
# signature failure in a new medium (ADR-0014 D6).
## what the players said while playing, and where
trials:
@$(PY) $(TOOLS)/trials.py
# CB-WP-0025 T06: the difficulty table (specs/RetrospectiveAnalysis.md §4).
# Winnable fraction from the solver plus a PLURAL policy panel -- a single
# policy's win rate may not be reported as a difficulty (§4.1).
## winnable fraction + a plural policy panel (never one bot's win rate)
difficulty:
@cargo run --release -q -p games-ground --example difficulty
# CB-REV-0002 #7: the H1 measurement harnesses were run by NO gate. Every
# number in CB-EV-0030 and CB-EV-0031 came from a manual invocation of an
# ungated binary -- so the assertion added to catch short cells was
# unreachable from `make`, and "what would the harness report if the work
# silently stopped" answered: green, and nothing else.
## the variant panels: ATTACK's value, and regulation under H1
panels:
@cargo run --release -q -p games-ground --example attack-value
@cargo run --release -q -p games-ground --example regulation
@cargo run --release -q -p games-ground --example perfect-recall
@cargo run --release -q -p games-ground --example h2-panel
# CB-WP-0022 T05: the design-finding register, reported over
# specs/GroundRules.md. Shows the QUEUE by default; the log of closed
# findings is a line, not a listing, because a default view that mixes
# them loses the queue property (ADR-0012 D5).
## the design-finding register: what is open, and what lacks a reproduction
design:
@$(PY) $(TOOLS)/design.py
# T03: one-shot orientation — workplans, next task, spend, fast gates.
# Cheap by design: no build. Start a session with this instead of grepping.
## one-shot orientation: workplans, next task, spend, fast gates
status:
@$(PY) $(TOOLS)/status.py
# T02: close a task — flip the workplan file, read the *measured* cost
# from the transcripts, push the hub event with real numbers. Refuses on an
# unknown or already-done task, and refuses to report an estimate.
# make task-done T=CB-WP-0004-T02
task-done:
@test -n "$(T)" || { echo "usage: make task-done T=CB-WP-0004-T02" >&2; exit 2; }
$(PY) $(TOOLS)/task-done.py $(T) $(ARGS)
# CB-WP-0007 T03 / InnerLoop v1.5: session shape for the window since the
# last commit. Deliberately NOT in `make all` — failing the build on
# context would block committing, and committing is the natural point to
# compact. A gate that blocks the remedy is a trap.
shape-budget: cost-test
$(PY) $(TOOLS)/cb-cost.py --shape-budget
# CB-01/CB-02: live spend since the last commit.
cost-budget: cost-test
$(PY) $(TOOLS)/cb-cost.py --budget
# CB-RES-0003 baseline: mechanical vs judgment turns.
cost-mix: cost-test
$(PY) $(TOOLS)/cb-cost.py --composition
cost-pin: cost-test
$(PY) $(TOOLS)/cb-cost.py --pin fc76445 --composition --by-task
## run all GROUND scenarios through cb-sim
sim:
$(IN_REPO) $(CARGO) run -q -p cb-sim -- $(REPO)/scenarios/ground/*.yaml
## Criterion benches (AM-6/AM-7)
bench:
$(IN_REPO) $(CARGO) bench -p games-ground
## InnerLoop positive control: run every bench once, no measurement.
## Fails if a workload stalls or produces the wrong event count.
bench-test:
$(IN_REPO) $(CARGO) bench -p games-ground --bench synthetic -- --test
## AM-2/AM-3 input: source LOC per crate (excludes tests would need tokei)
loc:
@$(IN_REPO) for d in crates/cb-kernel crates/cb-events crates/cb-game-runtime games/ground tools/cb-sim; do \
printf '%-28s %s\n' $$d "$$(find $$d/src -name '*.rs' | xargs cat | grep -vcE '^\s*(//|$$)')"; \
done
## every gate, in order. The one CI would run.
all: check test sim coverage size-metrics runtime-metrics am6 am7 am8 edition-check replay-test dep-weight self-tests env-test facts-check loop-lint bench-test panels