clay-borg/Makefile
tegwick 3f2ba2ec5d
Some checks failed
ci / check (push) Failing after 4s
CB-WP-0050: every configuration, faceted by aspect — and F31
module-panel sweeps points in aspect space the way scenario-panel sweeps
boards: one module at a time with every other aspect held, then the
profiles where an interaction is the hypothesis.

F31: attack_relief.self_soothe_ge4 HAS NEVER FIRED. Zero ATTACKs in
every cell, every seat band, both bots. The module acts only on an
uncancelled ATTACK by a seat at Stress >= 4, and greedy ranks Ground 100
at the gate against Attack's 10 -- so at exactly the position where
self-soothe would pay, GROUND wins. ModuleAwarePolicy adds 55 and that
is still not enough. `unplayed`, not `inert`: the kernel implements it
correctly and nothing has ever given it an opportunity.

This reaches backwards: CB-EV-0030's H1 verdict rests on its
flat-pressure half alone, because H1's other half never had an
opportunity in those games either.

Also: problem_stress.scoped is the only module with a measured effect
(3p 85->60 greedy, 86->63 module-aware, peak Stress 2->4);
flat_any_open drives group success to 0 at every band, reproducing the
H1 rejection from the module side; and scoped_plus_attack_soothe is
EXACTLY scoped alone -- necessarily, given F31 -- so the catalog's first
intentional multi-aspect combination cannot currently be evaluated as a
combination.

THE PANEL'S OWN DEFECT, first run: it printed 85/86 for a module that
never fired -- real-looking numbers inviting "measured, no effect" when
no seat ever created the precondition. Its docstring already said it
would not do that; the claim was written before the behaviour was.
First fix marked whole rows unmeasured, which threw away h1's real
flat-pressure result; the shipped fix names the specific module and
keeps the row's numbers, which are real for the modules that did fire.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-09 01:45:53 +02:00

326 lines
14 KiB
Makefile

# One command surface (InnerLoop §agentic-efficiency #3). Deterministic,
# greppable output; precursor of the `cb` CLI.
#
# CB-WP-0004 T01: every target here runs from a clean shell, from any
# directory, with no prefix. Invoke as `make -C <repo> <target>` from
# elsewhere. No target requires `cd` or `export PATH` — CB-RES-0003
# measured 84 turns and $15.33 spent on exactly those two prefixes.
# Absolute path to this Makefile's directory, so recipes never depend on
# the caller's working directory.
REPO := $(patsubst %/,%,$(dir $(abspath $(lastword $(MAKEFILE_LIST)))))
# Locate cargo instead of requiring it on the inherited PATH. Mirrors
# tools/repo.py:cargo_bin() — kept in sync by `make env-test`.
CARGO := $(firstword $(shell command -v cargo 2>/dev/null) \
$(wildcard $(HOME)/.cargo/bin/cargo) \
$(wildcard /usr/local/cargo/bin/cargo) \
cargo)
export PATH := $(dir $(CARGO)):$(PATH)
PY := python3
TOOLS := $(REPO)/tools
# `make` with no target lists what there is, rather than running the
# heaviest thing in the file. The Makefile's own header calls itself "one
# command surface" -- a surface you have to read the source of is not one.
.DEFAULT_GOAL := help
# A trial's name. Timestamped so two sessions on one day cannot overwrite
# each other's notes, and prefixed with the date because `tools/trials.py`
# reads the age from there (falling back to mtime).
TRIAL_NAME := $(shell date +%Y-%m-%d-%H%M)$(if $(SLUG),-$(SLUG),)
PLAYERS ?= 3
# CB-WP-0044: which rules package to play. The page names it too --
# a session that could not say which variant it was is what produced
# the report "I did a session and noted no changes".
VARIANT ?= ground-darvo-r0
# CB-WP-0047: the other three boards and the other two modes were
# unreachable from the way anyone actually plays. Defaults are what was
# always played, so `make ground` is unchanged.
SCENARIO ?= 1
MODE ?= shared
PORT ?= 0
# Every cargo recipe runs at the repo root; the shell does not persist cd.
IN_REPO := cd $(REPO) &&
.PHONY: help ground check test sim bench bench-test coverage dep-weight cost cost-test cost-pin cost-budget shape-budget cost-mix loop-lint self-tests env-test task-done status facts-check facts-gen mutation-check size-metrics runtime-metrics build-time am6 am7 am8 edition-check replay-test loc play gate-review panels all \
design difficulty trials
# `design`, `difficulty` and `trials` were added by CB-WP-0022, CB-WP-0025
# and CB-WP-0027 and none was declared here. Only `trials` revealed it, by
# colliding with the trials/ DIRECTORY -- Make saw an up-to-date file and
# ran nothing. The other two work by luck: no file happens to share their
# name. A target that is a command, not a file, belongs on this line.
## list every target, with what it does
## `make help` is also what bare `make` runs.
help:
@echo "clay-borg — make <target> (bare \`make\` shows this)"
@echo
@awk '/^## /{ sub(/^## /,""); d[++n]=$$0; next } \
/^[a-zA-Z][a-zA-Z0-9_-]*:/{ \
if (n) { split($$1,a,":"); printf " %-16s %s\n", a[1], d[1]; \
for (i=2;i<=n;i++) printf " %-16s %s\n", "", d[i]; print "" } \
n=0; next } \
{ n=0 }' $(MAKEFILE_LIST)
@echo " Variables: PLAYERS=3 PORT=0 SLUG=<name> ARGS=\"...\""
@echo
@echo " Undocumented (instruments and gate internals):"
@awk '/^## /{ d=1; next } \
/^[a-zA-Z][a-zA-Z0-9_-]*:/{ if (!d) { split($$1,a,":"); print a[1] } d=0; next } \
{ d=0 }' $(MAKEFILE_LIST) | sort -u | tr "\n" " " | fold -s -w 66 | sed "s/^/ /"
@echo
## play GROUND in a browser, recording a trial (the usual way in)
## Opens a URL; game on the left, notes on the right. Everything you
## type in the notes panel is bound to the position you typed it at.
## `make ground PLAYERS=2 SLUG=darvo-confusion`
## Read the notes back afterwards with `make trials`.
## play a session (VARIANT=ground-darvo-r0|h1|h2, SCENARIO=1|2|3|4, MODE=shared|common|coalitions)
ground:
@echo " rules: $(VARIANT) scenario: $(SCENARIO) scoring: $(MODE)"
@mkdir -p $(REPO)/trials
@echo " trial: trials/$(TRIAL_NAME).md (notes) + .yaml (the game)"
$(IN_REPO) $(CARGO) run -q -p cb-play -- \
--players $(PLAYERS) --serve $(PORT) --variant $(VARIANT) \
--scenario $(SCENARIO) --mode $(MODE) \
--record trials/$(TRIAL_NAME).yaml \
--trial trials/$(TRIAL_NAME).md $(ARGS)
## re-vendor the edition mirror from ground-game and record its digests
## The mirror goes stale whenever ground-game is edited. This syncs it
## and REGENERATES the digest block by walking editions/ — never by
## hand, which is what drifts. It reports files only one side has
## rather than resolving them: that is a decision, not a sync.
vendor:
@$(PY) $(TOOLS)/vendor-editions.py
## fmt + clippy (deny warnings) + HashMap deny-lint
check:
$(IN_REPO) $(CARGO) fmt --all --check
$(IN_REPO) $(CARGO) clippy --workspace --all-targets -- -D warnings
## play GROUND from the terminal (INTENT stage 0's CLI player).
## `make play ARGS="--all-bots --bot random"` to watch one instead.
play:
$(IN_REPO) $(CARGO) run -q -p cb-play -- $(ARGS)
## which control gates are due for a keep-or-kill argument (ADR-0006 D3)
gate-review:
$(PY) $(TOOLS)/gate-review.py
## unit + scenario-format tests
test:
$(IN_REPO) $(CARGO) test --workspace
## AM-4 dependency weight: third-party lines behind each budget
dep-weight:
$(PY) $(TOOLS)/dep-weight.py
## AM-1 rule coverage: which GROUND rules a scenario exercises
coverage:
$(PY) $(TOOLS)/rule-coverage.py
# K10: a bundle must re-execute, and must be able to fail. Four controls
# from ADR-0005 §6, including truncate-by-one-byte and mutated-seed.
replay-test:
$(PY) $(TOOLS)/replay-test.py
# AM-6 gate. Runs in RELEASE, where headroom is ~24x; the same assertion
# in debug has ~3.4x and flaked under load. The target is unchanged — only
# where it is measured.
am6:
$(IN_REPO) $(CARGO) test --release -p games-ground --all-features \
am6_throughput -- --ignored --nocapture --test-threads=1
# ADR-0011 D2: is the vendored edition still what ground-game published?
# Reports "upstream not checked out" as its own outcome, never a pass.
edition-check:
$(PY) $(TOOLS)/edition-check.py
# AM-8 N=10 determinism gate. One scenario, ten same-seed replays, all
# compared to the first. Not all 25: `make sim` already runs K8's double-
# run over every scenario, and repeating that eight more times costs 47 s
# per build to re-answer a question the second run already answered. The
# extra runs exist for the probabilistic class, and one workload gives
# that class its ten samples — see scenario::run_n.
am8:
$(IN_REPO) $(CARGO) run -q -p cb-sim -- --runs 10 \
$(REPO)/scenarios/ground/gr-r06-round-resolve.yaml
# AM-7 scaling gate. Same release / single-thread reasoning as am6 and more
# so: a *ratio* of two timings taken under varying contention is worse than
# one reading, because the noise multiplies rather than cancels.
am7:
$(IN_REPO) $(CARGO) test --release -p games-ground --all-features \
am7_cost_per_event -- --ignored --nocapture --test-threads=1
# AM-9 peak RSS (fast, gated). AM-5 needs a clean build — see build-time.
runtime-metrics:
$(PY) $(TOOLS)/runtime-metrics.py --fast
# AM-5: clean release build, ~90s, into a throwaway CARGO_TARGET_DIR so the
# working cache survives. Reported not gated per GameKernel §5.
build-time:
$(PY) $(TOOLS)/runtime-metrics.py
# AM-2 (M-D1-SPL, AM-1's anti-gaming pair) and AM-3.
size-metrics:
$(PY) $(TOOLS)/size-metrics.py
# M-D2-CST (specs/CostAccounting.md). cost-test is the positive control and
# runs first: a cost number from an unverified collector is void.
cost: cost-test
$(PY) $(TOOLS)/cb-cost.py --composition --by-task
cost-test:
$(PY) $(TOOLS)/cb-cost.py --self-test
# InnerLoop rules that are mechanically checkable (CB-WP-0003 T01).
loop-lint:
$(PY) $(TOOLS)/loop-lint.py
# Positive control for every reporting tool, per InnerLoop v1.1 Step 5.
## positive control for every reporting tool
self-tests:
$(PY) $(TOOLS)/cb-cost.py --self-test
$(PY) $(TOOLS)/loop-lint.py --self-test
$(PY) $(TOOLS)/rule-coverage.py --self-test
$(PY) $(TOOLS)/dep-weight.py --self-test
$(PY) $(TOOLS)/repo.py --self-test
$(PY) $(TOOLS)/task-done.py --self-test
$(PY) $(TOOLS)/status.py --self-test
$(PY) $(TOOLS)/facts.py --self-test
$(PY) $(TOOLS)/mutation-check.py --self-test
$(PY) $(TOOLS)/size-metrics.py --self-test
$(PY) $(TOOLS)/runtime-metrics.py --self-test
$(PY) $(TOOLS)/replay-test.py --self-test
$(PY) $(TOOLS)/design.py --self-test
$(PY) $(TOOLS)/trials.py --self-test
cargo run --release -q -p games-ground --example difficulty -- --self-test
$(PY) $(TOOLS)/edition-check.py --self-test
# T01 positive control: prove the environment fix, do not assume it. Runs
# every tool from a foreign working directory with a PATH that has no
# cargo on it. Before T01 this failed; if it fails again, the friction is
# back and CB-EV-0003's measurement is invalid.
env-test:
@cd / && env PATH=/usr/bin:/bin $(PY) $(TOOLS)/repo.py --self-test
@cd / && env PATH=/usr/bin:/bin $(PY) $(TOOLS)/rule-coverage.py --self-test >/dev/null \
&& echo " [ok ] rule-coverage runs from / with no cargo on PATH"
@cd / && env PATH=/usr/bin:/bin $(PY) $(TOOLS)/dep-weight.py --self-test >/dev/null \
&& echo " [ok ] dep-weight runs from / with no cargo on PATH"
@cd / && env PATH=/usr/bin:/bin $(PY) $(TOOLS)/cb-cost.py --self-test >/dev/null \
&& echo " [ok ] cb-cost runs from / with no cargo on PATH"
@cd / && env PATH=/usr/bin:/bin $(PY) $(TOOLS)/loop-lint.py --self-test >/dev/null \
&& echo " [ok ] loop-lint runs from / with no cargo on PATH"
@$(MAKE) -C $(REPO) coverage >/dev/null \
&& echo " [ok ] make -C <repo> works from any directory"
# M-D1-MUT (CB-WP-0005 T02): invert each acceptance row's property and
# require the verifying command to go red. Deliberately NOT in `make all`:
# it rebuilds the workspace once per mutated row. Run it on demand and in
# CI, not in the inner loop.
mutation-check:
$(PY) $(TOOLS)/mutation-check.py $(ARGS)
# T04: single source of fact (InnerLoop v1.2) — the DFD gate.
# facts.toml is GENERATED; facts-check fails if it disagrees with the
# instruments, or if a tagged artifact disagrees with it.
facts-check:
$(PY) $(TOOLS)/facts.py --check
facts-gen:
$(PY) $(TOOLS)/facts.py --gen
# CB-WP-0027 T04: what the players said, and where. Surfacing is the
# deliverable -- a commentary feature nobody can read is this project's
# signature failure in a new medium (ADR-0014 D6).
## what the players said while playing, and where
trials:
@$(PY) $(TOOLS)/trials.py
# CB-WP-0025 T06: the difficulty table (specs/RetrospectiveAnalysis.md §4).
# Winnable fraction from the solver plus a PLURAL policy panel -- a single
# policy's win rate may not be reported as a difficulty (§4.1).
## winnable fraction + a plural policy panel (never one bot's win rate)
difficulty:
@cargo run --release -q -p games-ground --example difficulty
# CB-REV-0002 #7: the H1 measurement harnesses were run by NO gate. Every
# number in CB-EV-0030 and CB-EV-0031 came from a manual invocation of an
# ungated binary -- so the assertion added to catch short cells was
# unreachable from `make`, and "what would the harness report if the work
# silently stopped" answered: green, and nothing else.
## the variant panels: ATTACK's value, and regulation under H1
panels:
@cargo run --release -q -p games-ground --example attack-value
@cargo run --release -q -p games-ground --example regulation
@cargo run --release -q -p games-ground --example perfect-recall
@cargo run --release -q -p games-ground --example h2-panel
@cargo run --release -q -p games-ground --example scenario-panel
@cargo run --release -q -p games-ground --example module-panel
# CB-WP-0022 T05: the design-finding register, reported over
# specs/GroundRules.md. Shows the QUEUE by default; the log of closed
# findings is a line, not a listing, because a default view that mixes
# them loses the queue property (ADR-0012 D5).
## the design-finding register: what is open, and what lacks a reproduction
design:
@$(PY) $(TOOLS)/design.py
# T03: one-shot orientation — workplans, next task, spend, fast gates.
# Cheap by design: no build. Start a session with this instead of grepping.
## one-shot orientation: workplans, next task, spend, fast gates
status:
@$(PY) $(TOOLS)/status.py
# T02: close a task — flip the workplan file, read the *measured* cost
# from the transcripts, push the hub event with real numbers. Refuses on an
# unknown or already-done task, and refuses to report an estimate.
# make task-done T=CB-WP-0004-T02
task-done:
@test -n "$(T)" || { echo "usage: make task-done T=CB-WP-0004-T02" >&2; exit 2; }
$(PY) $(TOOLS)/task-done.py $(T) $(ARGS)
# CB-WP-0007 T03 / InnerLoop v1.5: session shape for the window since the
# last commit. Deliberately NOT in `make all` — failing the build on
# context would block committing, and committing is the natural point to
# compact. A gate that blocks the remedy is a trap.
shape-budget: cost-test
$(PY) $(TOOLS)/cb-cost.py --shape-budget
# CB-01/CB-02: live spend since the last commit.
cost-budget: cost-test
$(PY) $(TOOLS)/cb-cost.py --budget
# CB-RES-0003 baseline: mechanical vs judgment turns.
cost-mix: cost-test
$(PY) $(TOOLS)/cb-cost.py --composition
cost-pin: cost-test
$(PY) $(TOOLS)/cb-cost.py --pin fc76445 --composition --by-task
## run all GROUND scenarios through cb-sim
sim:
$(IN_REPO) $(CARGO) run -q -p cb-sim -- $(REPO)/scenarios/ground/*.yaml
## Criterion benches (AM-6/AM-7)
bench:
$(IN_REPO) $(CARGO) bench -p games-ground
## InnerLoop positive control: run every bench once, no measurement.
## Fails if a workload stalls or produces the wrong event count.
bench-test:
$(IN_REPO) $(CARGO) bench -p games-ground --bench synthetic -- --test
## AM-2/AM-3 input: source LOC per crate (excludes tests would need tokei)
loc:
@$(IN_REPO) for d in crates/cb-kernel crates/cb-events crates/cb-game-runtime games/ground tools/cb-sim; do \
printf '%-28s %s\n' $$d "$$(find $$d/src -name '*.rs' | xargs cat | grep -vcE '^\s*(//|$$)')"; \
done
## every gate, in order. The one CI would run.
all: check test sim coverage size-metrics runtime-metrics am6 am7 am8 edition-check replay-test dep-weight self-tests env-test facts-check loop-lint bench-test panels