clay-borg/Makefile
tegwick d30938b259
Some checks failed
ci / check (push) Failing after 3s
CB-WP-0047: all four boards, and every mode named on the page
The modes were already implemented; nothing had ever COMPARED them. The
scenarios were not implemented at all: edition::deal has taken a
scenario_id since it was written and the only caller passed the literal
"SCN_01", so 15 of 20 Problem cards had never been dealt by anything.
The seam was the whole mechanism and it sat unused, with nothing red
because nothing asked.

Scenario is now state (serde default SCN_01, so all 26 recordings replay
unchanged), selected by preset `scn-03-4p` with `standard-Np` still
meaning SCN_01, and by --scenario/SCENARIO= accepting ids, numbers or
titles, validated against the edition rather than a pattern.

The threshold now comes off the Scenario card, closing F25's hardcoded
5/7/9. The first version of that control was worthless and mutation said
so: all four scenarios print 5/7/9, so reverting to the bands left it
green. Split threshold_from() so it can be handed a card that disagrees.

The header read `scoring CommonProblem` where the Mode card is titled
COMMON PROBLEM, PERSONAL EDGE -- the defect CB-WP-0034 deleted from the
move buttons, still standing on the line that says what winning means.
The coverage probe was matching that Debug output and went red when it
was fixed: third instance (CB-WP-0024, CB-WP-0034). Page now carries the
premise, the mode's rules text, and the tiebreak.

scenario-panel plays 4x3x3. Findings: SCN_01 and SCN_02 are the same
board (identical cells, pinned by a characterisation test); SCN_04 is
the hard board at 2p (52% vs 67/73%, the only deck needing two Repair);
and group success is EXACTLY equal across all three modes in all 36
cells, because greedy never reads state.mode -- filed F27, the two
competitive modes are scoring lenses over cooperative play.

F28: SHARED GROUND's mastery subtracts penalties from the claimed COUNT
where the mode card's shared score is claimed VALUE. Raised, not fixed;
scoring is ground-game's to rule on.

Also fixes design.py reporting a backticked path as no reproduction.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-08 20:51:46 +02:00

317 lines
14 KiB
Makefile

# One command surface (InnerLoop §agentic-efficiency #3). Deterministic,
# greppable output; precursor of the `cb` CLI.
#
# CB-WP-0004 T01: every target here runs from a clean shell, from any
# directory, with no prefix. Invoke as `make -C <repo> <target>` from
# elsewhere. No target requires `cd` or `export PATH` — CB-RES-0003
# measured 84 turns and $15.33 spent on exactly those two prefixes.
# Absolute path to this Makefile's directory, so recipes never depend on
# the caller's working directory.
REPO := $(patsubst %/,%,$(dir $(abspath $(lastword $(MAKEFILE_LIST)))))
# Locate cargo instead of requiring it on the inherited PATH. Mirrors
# tools/repo.py:cargo_bin() — kept in sync by `make env-test`.
CARGO := $(firstword $(shell command -v cargo 2>/dev/null) \
$(wildcard $(HOME)/.cargo/bin/cargo) \
$(wildcard /usr/local/cargo/bin/cargo) \
cargo)
export PATH := $(dir $(CARGO)):$(PATH)
PY := python3
TOOLS := $(REPO)/tools
# `make` with no target lists what there is, rather than running the
# heaviest thing in the file. The Makefile's own header calls itself "one
# command surface" -- a surface you have to read the source of is not one.
.DEFAULT_GOAL := help
# A trial's name. Timestamped so two sessions on one day cannot overwrite
# each other's notes, and prefixed with the date because `tools/trials.py`
# reads the age from there (falling back to mtime).
TRIAL_NAME := $(shell date +%Y-%m-%d-%H%M)$(if $(SLUG),-$(SLUG),)
PLAYERS ?= 3
# CB-WP-0044: which rules package to play. The page names it too --
# a session that could not say which variant it was is what produced
# the report "I did a session and noted no changes".
VARIANT ?= ground-darvo-r0
# CB-WP-0047: the other three boards and the other two modes were
# unreachable from the way anyone actually plays. Defaults are what was
# always played, so `make ground` is unchanged.
SCENARIO ?= 1
MODE ?= shared
PORT ?= 0
# Every cargo recipe runs at the repo root; the shell does not persist cd.
IN_REPO := cd $(REPO) &&
.PHONY: help ground check test sim bench bench-test coverage dep-weight cost cost-test cost-pin cost-budget shape-budget cost-mix loop-lint self-tests env-test task-done status facts-check facts-gen mutation-check size-metrics runtime-metrics build-time am6 am7 am8 edition-check replay-test loc play gate-review panels all \
design difficulty trials
# `design`, `difficulty` and `trials` were added by CB-WP-0022, CB-WP-0025
# and CB-WP-0027 and none was declared here. Only `trials` revealed it, by
# colliding with the trials/ DIRECTORY -- Make saw an up-to-date file and
# ran nothing. The other two work by luck: no file happens to share their
# name. A target that is a command, not a file, belongs on this line.
## list every target, with what it does
## `make help` is also what bare `make` runs.
help:
@echo "clay-borg — make <target> (bare \`make\` shows this)"
@echo
@awk '/^## /{ sub(/^## /,""); d[++n]=$$0; next } \
/^[a-zA-Z][a-zA-Z0-9_-]*:/{ \
if (n) { split($$1,a,":"); printf " %-16s %s\n", a[1], d[1]; \
for (i=2;i<=n;i++) printf " %-16s %s\n", "", d[i]; print "" } \
n=0; next } \
{ n=0 }' $(MAKEFILE_LIST)
@echo " Variables: PLAYERS=3 PORT=0 SLUG=<name> ARGS=\"...\""
@echo
@echo " Undocumented (instruments and gate internals):"
@awk '/^## /{ d=1; next } \
/^[a-zA-Z][a-zA-Z0-9_-]*:/{ if (!d) { split($$1,a,":"); print a[1] } d=0; next } \
{ d=0 }' $(MAKEFILE_LIST) | sort -u | tr "\n" " " | fold -s -w 66 | sed "s/^/ /"
@echo
## play GROUND in a browser, recording a trial (the usual way in)
## Opens a URL; game on the left, notes on the right. Everything you
## type in the notes panel is bound to the position you typed it at.
## `make ground PLAYERS=2 SLUG=darvo-confusion`
## Read the notes back afterwards with `make trials`.
## play a session (VARIANT=ground-darvo-r0|h1|h2, SCENARIO=1|2|3|4, MODE=shared|common|coalitions)
ground:
@echo " rules: $(VARIANT) scenario: $(SCENARIO) scoring: $(MODE)"
@mkdir -p $(REPO)/trials
@echo " trial: trials/$(TRIAL_NAME).md (notes) + .yaml (the game)"
$(IN_REPO) $(CARGO) run -q -p cb-play -- \
--players $(PLAYERS) --serve $(PORT) --variant $(VARIANT) \
--scenario $(SCENARIO) --mode $(MODE) \
--record trials/$(TRIAL_NAME).yaml \
--trial trials/$(TRIAL_NAME).md $(ARGS)
## fmt + clippy (deny warnings) + HashMap deny-lint
check:
$(IN_REPO) $(CARGO) fmt --all --check
$(IN_REPO) $(CARGO) clippy --workspace --all-targets -- -D warnings
## play GROUND from the terminal (INTENT stage 0's CLI player).
## `make play ARGS="--all-bots --bot random"` to watch one instead.
play:
$(IN_REPO) $(CARGO) run -q -p cb-play -- $(ARGS)
## which control gates are due for a keep-or-kill argument (ADR-0006 D3)
gate-review:
$(PY) $(TOOLS)/gate-review.py
## unit + scenario-format tests
test:
$(IN_REPO) $(CARGO) test --workspace
## AM-4 dependency weight: third-party lines behind each budget
dep-weight:
$(PY) $(TOOLS)/dep-weight.py
## AM-1 rule coverage: which GROUND rules a scenario exercises
coverage:
$(PY) $(TOOLS)/rule-coverage.py
# K10: a bundle must re-execute, and must be able to fail. Four controls
# from ADR-0005 §6, including truncate-by-one-byte and mutated-seed.
replay-test:
$(PY) $(TOOLS)/replay-test.py
# AM-6 gate. Runs in RELEASE, where headroom is ~24x; the same assertion
# in debug has ~3.4x and flaked under load. The target is unchanged — only
# where it is measured.
am6:
$(IN_REPO) $(CARGO) test --release -p games-ground --all-features \
am6_throughput -- --ignored --nocapture --test-threads=1
# ADR-0011 D2: is the vendored edition still what ground-game published?
# Reports "upstream not checked out" as its own outcome, never a pass.
edition-check:
$(PY) $(TOOLS)/edition-check.py
# AM-8 N=10 determinism gate. One scenario, ten same-seed replays, all
# compared to the first. Not all 25: `make sim` already runs K8's double-
# run over every scenario, and repeating that eight more times costs 47 s
# per build to re-answer a question the second run already answered. The
# extra runs exist for the probabilistic class, and one workload gives
# that class its ten samples — see scenario::run_n.
am8:
$(IN_REPO) $(CARGO) run -q -p cb-sim -- --runs 10 \
$(REPO)/scenarios/ground/gr-r06-round-resolve.yaml
# AM-7 scaling gate. Same release / single-thread reasoning as am6 and more
# so: a *ratio* of two timings taken under varying contention is worse than
# one reading, because the noise multiplies rather than cancels.
am7:
$(IN_REPO) $(CARGO) test --release -p games-ground --all-features \
am7_cost_per_event -- --ignored --nocapture --test-threads=1
# AM-9 peak RSS (fast, gated). AM-5 needs a clean build — see build-time.
runtime-metrics:
$(PY) $(TOOLS)/runtime-metrics.py --fast
# AM-5: clean release build, ~90s, into a throwaway CARGO_TARGET_DIR so the
# working cache survives. Reported not gated per GameKernel §5.
build-time:
$(PY) $(TOOLS)/runtime-metrics.py
# AM-2 (M-D1-SPL, AM-1's anti-gaming pair) and AM-3.
size-metrics:
$(PY) $(TOOLS)/size-metrics.py
# M-D2-CST (specs/CostAccounting.md). cost-test is the positive control and
# runs first: a cost number from an unverified collector is void.
cost: cost-test
$(PY) $(TOOLS)/cb-cost.py --composition --by-task
cost-test:
$(PY) $(TOOLS)/cb-cost.py --self-test
# InnerLoop rules that are mechanically checkable (CB-WP-0003 T01).
loop-lint:
$(PY) $(TOOLS)/loop-lint.py
# Positive control for every reporting tool, per InnerLoop v1.1 Step 5.
## positive control for every reporting tool
self-tests:
$(PY) $(TOOLS)/cb-cost.py --self-test
$(PY) $(TOOLS)/loop-lint.py --self-test
$(PY) $(TOOLS)/rule-coverage.py --self-test
$(PY) $(TOOLS)/dep-weight.py --self-test
$(PY) $(TOOLS)/repo.py --self-test
$(PY) $(TOOLS)/task-done.py --self-test
$(PY) $(TOOLS)/status.py --self-test
$(PY) $(TOOLS)/facts.py --self-test
$(PY) $(TOOLS)/mutation-check.py --self-test
$(PY) $(TOOLS)/size-metrics.py --self-test
$(PY) $(TOOLS)/runtime-metrics.py --self-test
$(PY) $(TOOLS)/replay-test.py --self-test
$(PY) $(TOOLS)/design.py --self-test
$(PY) $(TOOLS)/trials.py --self-test
cargo run --release -q -p games-ground --example difficulty -- --self-test
$(PY) $(TOOLS)/edition-check.py --self-test
# T01 positive control: prove the environment fix, do not assume it. Runs
# every tool from a foreign working directory with a PATH that has no
# cargo on it. Before T01 this failed; if it fails again, the friction is
# back and CB-EV-0003's measurement is invalid.
env-test:
@cd / && env PATH=/usr/bin:/bin $(PY) $(TOOLS)/repo.py --self-test
@cd / && env PATH=/usr/bin:/bin $(PY) $(TOOLS)/rule-coverage.py --self-test >/dev/null \
&& echo " [ok ] rule-coverage runs from / with no cargo on PATH"
@cd / && env PATH=/usr/bin:/bin $(PY) $(TOOLS)/dep-weight.py --self-test >/dev/null \
&& echo " [ok ] dep-weight runs from / with no cargo on PATH"
@cd / && env PATH=/usr/bin:/bin $(PY) $(TOOLS)/cb-cost.py --self-test >/dev/null \
&& echo " [ok ] cb-cost runs from / with no cargo on PATH"
@cd / && env PATH=/usr/bin:/bin $(PY) $(TOOLS)/loop-lint.py --self-test >/dev/null \
&& echo " [ok ] loop-lint runs from / with no cargo on PATH"
@$(MAKE) -C $(REPO) coverage >/dev/null \
&& echo " [ok ] make -C <repo> works from any directory"
# M-D1-MUT (CB-WP-0005 T02): invert each acceptance row's property and
# require the verifying command to go red. Deliberately NOT in `make all`:
# it rebuilds the workspace once per mutated row. Run it on demand and in
# CI, not in the inner loop.
mutation-check:
$(PY) $(TOOLS)/mutation-check.py $(ARGS)
# T04: single source of fact (InnerLoop v1.2) — the DFD gate.
# facts.toml is GENERATED; facts-check fails if it disagrees with the
# instruments, or if a tagged artifact disagrees with it.
facts-check:
$(PY) $(TOOLS)/facts.py --check
facts-gen:
$(PY) $(TOOLS)/facts.py --gen
# CB-WP-0027 T04: what the players said, and where. Surfacing is the
# deliverable -- a commentary feature nobody can read is this project's
# signature failure in a new medium (ADR-0014 D6).
## what the players said while playing, and where
trials:
@$(PY) $(TOOLS)/trials.py
# CB-WP-0025 T06: the difficulty table (specs/RetrospectiveAnalysis.md §4).
# Winnable fraction from the solver plus a PLURAL policy panel -- a single
# policy's win rate may not be reported as a difficulty (§4.1).
## winnable fraction + a plural policy panel (never one bot's win rate)
difficulty:
@cargo run --release -q -p games-ground --example difficulty
# CB-REV-0002 #7: the H1 measurement harnesses were run by NO gate. Every
# number in CB-EV-0030 and CB-EV-0031 came from a manual invocation of an
# ungated binary -- so the assertion added to catch short cells was
# unreachable from `make`, and "what would the harness report if the work
# silently stopped" answered: green, and nothing else.
## the variant panels: ATTACK's value, and regulation under H1
panels:
@cargo run --release -q -p games-ground --example attack-value
@cargo run --release -q -p games-ground --example regulation
@cargo run --release -q -p games-ground --example perfect-recall
@cargo run --release -q -p games-ground --example h2-panel
@cargo run --release -q -p games-ground --example scenario-panel
# CB-WP-0022 T05: the design-finding register, reported over
# specs/GroundRules.md. Shows the QUEUE by default; the log of closed
# findings is a line, not a listing, because a default view that mixes
# them loses the queue property (ADR-0012 D5).
## the design-finding register: what is open, and what lacks a reproduction
design:
@$(PY) $(TOOLS)/design.py
# T03: one-shot orientation — workplans, next task, spend, fast gates.
# Cheap by design: no build. Start a session with this instead of grepping.
## one-shot orientation: workplans, next task, spend, fast gates
status:
@$(PY) $(TOOLS)/status.py
# T02: close a task — flip the workplan file, read the *measured* cost
# from the transcripts, push the hub event with real numbers. Refuses on an
# unknown or already-done task, and refuses to report an estimate.
# make task-done T=CB-WP-0004-T02
task-done:
@test -n "$(T)" || { echo "usage: make task-done T=CB-WP-0004-T02" >&2; exit 2; }
$(PY) $(TOOLS)/task-done.py $(T) $(ARGS)
# CB-WP-0007 T03 / InnerLoop v1.5: session shape for the window since the
# last commit. Deliberately NOT in `make all` — failing the build on
# context would block committing, and committing is the natural point to
# compact. A gate that blocks the remedy is a trap.
shape-budget: cost-test
$(PY) $(TOOLS)/cb-cost.py --shape-budget
# CB-01/CB-02: live spend since the last commit.
cost-budget: cost-test
$(PY) $(TOOLS)/cb-cost.py --budget
# CB-RES-0003 baseline: mechanical vs judgment turns.
cost-mix: cost-test
$(PY) $(TOOLS)/cb-cost.py --composition
cost-pin: cost-test
$(PY) $(TOOLS)/cb-cost.py --pin fc76445 --composition --by-task
## run all GROUND scenarios through cb-sim
sim:
$(IN_REPO) $(CARGO) run -q -p cb-sim -- $(REPO)/scenarios/ground/*.yaml
## Criterion benches (AM-6/AM-7)
bench:
$(IN_REPO) $(CARGO) bench -p games-ground
## InnerLoop positive control: run every bench once, no measurement.
## Fails if a workload stalls or produces the wrong event count.
bench-test:
$(IN_REPO) $(CARGO) bench -p games-ground --bench synthetic -- --test
## AM-2/AM-3 input: source LOC per crate (excludes tests would need tokei)
loc:
@$(IN_REPO) for d in crates/cb-kernel crates/cb-events crates/cb-game-runtime games/ground tools/cb-sim; do \
printf '%-28s %s\n' $$d "$$(find $$d/src -name '*.rs' | xargs cat | grep -vcE '^\s*(//|$$)')"; \
done
## every gate, in order. The one CI would run.
all: check test sim coverage size-metrics runtime-metrics am6 am7 am8 edition-check replay-test dep-weight self-tests env-test facts-check loop-lint bench-test panels