Some checks failed
ci / check (push) Failing after 4s
Reported as "I did a session and noted no changes" — correct, and the fault was ours twice. make ground passes no --variant, so the session was baseline, and CB-WP-0043 deliberately leaves the baseline layout untouched because a variant that redraws the baseline invalidates every prior look at it. But the deeper defect is that nothing said so. The page named the round, the step, the Lead, the scoring mode and the viewer, and never which rules it was playing — GroundView did not even carry the variant. There was no way from the screen to tell baseline from H1 from H2. A session that cannot say which rules it is playing cannot report a rules change; the player did the right thing and the instrument had nothing to tell them. Now: variant on GroundView, `rules <id>` in the page header beside the scoring mode, `rules <id>` in the inspector because a replay that cannot say which rules produced it is the same defect in the other tool, and make ground VARIANT=h2 so the capability is reachable. Both coverage probes caught the new field independently — the render crate's and cb-play's — the second time in two passes that they have turned an addition into a legibility requirement instead of letting it be silent state. Verified by fetching the served page rather than by reading the code: make ground VARIANT=h2 prints "rules: h2" and the page carries "rules h2-scoped-problem-stress" with the scope labels; the baseline says "rules ground-darvo-r0" and keeps its row. Still open: nothing explains what a scope DOES, and the trial log header does not record the variant either — the same defect one artifact along. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
310 lines
13 KiB
Makefile
310 lines
13 KiB
Makefile
# One command surface (InnerLoop §agentic-efficiency #3). Deterministic,
|
|
# greppable output; precursor of the `cb` CLI.
|
|
#
|
|
# CB-WP-0004 T01: every target here runs from a clean shell, from any
|
|
# directory, with no prefix. Invoke as `make -C <repo> <target>` from
|
|
# elsewhere. No target requires `cd` or `export PATH` — CB-RES-0003
|
|
# measured 84 turns and $15.33 spent on exactly those two prefixes.
|
|
|
|
# Absolute path to this Makefile's directory, so recipes never depend on
|
|
# the caller's working directory.
|
|
REPO := $(patsubst %/,%,$(dir $(abspath $(lastword $(MAKEFILE_LIST)))))
|
|
|
|
# Locate cargo instead of requiring it on the inherited PATH. Mirrors
|
|
# tools/repo.py:cargo_bin() — kept in sync by `make env-test`.
|
|
CARGO := $(firstword $(shell command -v cargo 2>/dev/null) \
|
|
$(wildcard $(HOME)/.cargo/bin/cargo) \
|
|
$(wildcard /usr/local/cargo/bin/cargo) \
|
|
cargo)
|
|
export PATH := $(dir $(CARGO)):$(PATH)
|
|
|
|
PY := python3
|
|
TOOLS := $(REPO)/tools
|
|
|
|
# `make` with no target lists what there is, rather than running the
|
|
# heaviest thing in the file. The Makefile's own header calls itself "one
|
|
# command surface" -- a surface you have to read the source of is not one.
|
|
.DEFAULT_GOAL := help
|
|
|
|
# A trial's name. Timestamped so two sessions on one day cannot overwrite
|
|
# each other's notes, and prefixed with the date because `tools/trials.py`
|
|
# reads the age from there (falling back to mtime).
|
|
TRIAL_NAME := $(shell date +%Y-%m-%d-%H%M)$(if $(SLUG),-$(SLUG),)
|
|
PLAYERS ?= 3
|
|
# CB-WP-0044: which rules package to play. The page names it too --
|
|
# a session that could not say which variant it was is what produced
|
|
# the report "I did a session and noted no changes".
|
|
VARIANT ?= ground-darvo-r0
|
|
PORT ?= 0
|
|
|
|
# Every cargo recipe runs at the repo root; the shell does not persist cd.
|
|
IN_REPO := cd $(REPO) &&
|
|
|
|
.PHONY: help ground check test sim bench bench-test coverage dep-weight cost cost-test cost-pin cost-budget shape-budget cost-mix loop-lint self-tests env-test task-done status facts-check facts-gen mutation-check size-metrics runtime-metrics build-time am6 am7 am8 edition-check replay-test loc play gate-review panels all \
|
|
design difficulty trials
|
|
# `design`, `difficulty` and `trials` were added by CB-WP-0022, CB-WP-0025
|
|
# and CB-WP-0027 and none was declared here. Only `trials` revealed it, by
|
|
# colliding with the trials/ DIRECTORY -- Make saw an up-to-date file and
|
|
# ran nothing. The other two work by luck: no file happens to share their
|
|
# name. A target that is a command, not a file, belongs on this line.
|
|
|
|
## list every target, with what it does
|
|
## `make help` is also what bare `make` runs.
|
|
help:
|
|
@echo "clay-borg — make <target> (bare \`make\` shows this)"
|
|
@echo
|
|
@awk '/^## /{ sub(/^## /,""); d[++n]=$$0; next } \
|
|
/^[a-zA-Z][a-zA-Z0-9_-]*:/{ \
|
|
if (n) { split($$1,a,":"); printf " %-16s %s\n", a[1], d[1]; \
|
|
for (i=2;i<=n;i++) printf " %-16s %s\n", "", d[i]; print "" } \
|
|
n=0; next } \
|
|
{ n=0 }' $(MAKEFILE_LIST)
|
|
@echo " Variables: PLAYERS=3 PORT=0 SLUG=<name> ARGS=\"...\""
|
|
@echo
|
|
@echo " Undocumented (instruments and gate internals):"
|
|
@awk '/^## /{ d=1; next } \
|
|
/^[a-zA-Z][a-zA-Z0-9_-]*:/{ if (!d) { split($$1,a,":"); print a[1] } d=0; next } \
|
|
{ d=0 }' $(MAKEFILE_LIST) | sort -u | tr "\n" " " | fold -s -w 66 | sed "s/^/ /"
|
|
@echo
|
|
|
|
## play GROUND in a browser, recording a trial (the usual way in)
|
|
## Opens a URL; game on the left, notes on the right. Everything you
|
|
## type in the notes panel is bound to the position you typed it at.
|
|
## `make ground PLAYERS=2 SLUG=darvo-confusion`
|
|
## Read the notes back afterwards with `make trials`.
|
|
## play a session (VARIANT=ground-darvo-r0|h1|h2)
|
|
ground:
|
|
@echo " rules: $(VARIANT)"
|
|
@mkdir -p $(REPO)/trials
|
|
@echo " trial: trials/$(TRIAL_NAME).md (notes) + .yaml (the game)"
|
|
$(IN_REPO) $(CARGO) run -q -p cb-play -- \
|
|
--players $(PLAYERS) --serve $(PORT) --variant $(VARIANT) \
|
|
--record trials/$(TRIAL_NAME).yaml \
|
|
--trial trials/$(TRIAL_NAME).md $(ARGS)
|
|
|
|
## fmt + clippy (deny warnings) + HashMap deny-lint
|
|
check:
|
|
$(IN_REPO) $(CARGO) fmt --all --check
|
|
$(IN_REPO) $(CARGO) clippy --workspace --all-targets -- -D warnings
|
|
|
|
## play GROUND from the terminal (INTENT stage 0's CLI player).
|
|
## `make play ARGS="--all-bots --bot random"` to watch one instead.
|
|
play:
|
|
$(IN_REPO) $(CARGO) run -q -p cb-play -- $(ARGS)
|
|
|
|
## which control gates are due for a keep-or-kill argument (ADR-0006 D3)
|
|
gate-review:
|
|
$(PY) $(TOOLS)/gate-review.py
|
|
|
|
## unit + scenario-format tests
|
|
test:
|
|
$(IN_REPO) $(CARGO) test --workspace
|
|
|
|
## AM-4 dependency weight: third-party lines behind each budget
|
|
dep-weight:
|
|
$(PY) $(TOOLS)/dep-weight.py
|
|
|
|
## AM-1 rule coverage: which GROUND rules a scenario exercises
|
|
coverage:
|
|
$(PY) $(TOOLS)/rule-coverage.py
|
|
|
|
# K10: a bundle must re-execute, and must be able to fail. Four controls
|
|
# from ADR-0005 §6, including truncate-by-one-byte and mutated-seed.
|
|
replay-test:
|
|
$(PY) $(TOOLS)/replay-test.py
|
|
|
|
# AM-6 gate. Runs in RELEASE, where headroom is ~24x; the same assertion
|
|
# in debug has ~3.4x and flaked under load. The target is unchanged — only
|
|
# where it is measured.
|
|
am6:
|
|
$(IN_REPO) $(CARGO) test --release -p games-ground --all-features \
|
|
am6_throughput -- --ignored --nocapture --test-threads=1
|
|
|
|
# ADR-0011 D2: is the vendored edition still what ground-game published?
|
|
# Reports "upstream not checked out" as its own outcome, never a pass.
|
|
edition-check:
|
|
$(PY) $(TOOLS)/edition-check.py
|
|
|
|
# AM-8 N=10 determinism gate. One scenario, ten same-seed replays, all
|
|
# compared to the first. Not all 25: `make sim` already runs K8's double-
|
|
# run over every scenario, and repeating that eight more times costs 47 s
|
|
# per build to re-answer a question the second run already answered. The
|
|
# extra runs exist for the probabilistic class, and one workload gives
|
|
# that class its ten samples — see scenario::run_n.
|
|
am8:
|
|
$(IN_REPO) $(CARGO) run -q -p cb-sim -- --runs 10 \
|
|
$(REPO)/scenarios/ground/gr-r06-round-resolve.yaml
|
|
|
|
# AM-7 scaling gate. Same release / single-thread reasoning as am6 and more
|
|
# so: a *ratio* of two timings taken under varying contention is worse than
|
|
# one reading, because the noise multiplies rather than cancels.
|
|
am7:
|
|
$(IN_REPO) $(CARGO) test --release -p games-ground --all-features \
|
|
am7_cost_per_event -- --ignored --nocapture --test-threads=1
|
|
|
|
# AM-9 peak RSS (fast, gated). AM-5 needs a clean build — see build-time.
|
|
runtime-metrics:
|
|
$(PY) $(TOOLS)/runtime-metrics.py --fast
|
|
|
|
# AM-5: clean release build, ~90s, into a throwaway CARGO_TARGET_DIR so the
|
|
# working cache survives. Reported not gated per GameKernel §5.
|
|
build-time:
|
|
$(PY) $(TOOLS)/runtime-metrics.py
|
|
|
|
# AM-2 (M-D1-SPL, AM-1's anti-gaming pair) and AM-3.
|
|
size-metrics:
|
|
$(PY) $(TOOLS)/size-metrics.py
|
|
|
|
# M-D2-CST (specs/CostAccounting.md). cost-test is the positive control and
|
|
# runs first: a cost number from an unverified collector is void.
|
|
cost: cost-test
|
|
$(PY) $(TOOLS)/cb-cost.py --composition --by-task
|
|
|
|
cost-test:
|
|
$(PY) $(TOOLS)/cb-cost.py --self-test
|
|
|
|
# InnerLoop rules that are mechanically checkable (CB-WP-0003 T01).
|
|
loop-lint:
|
|
$(PY) $(TOOLS)/loop-lint.py
|
|
|
|
# Positive control for every reporting tool, per InnerLoop v1.1 Step 5.
|
|
## positive control for every reporting tool
|
|
self-tests:
|
|
$(PY) $(TOOLS)/cb-cost.py --self-test
|
|
$(PY) $(TOOLS)/loop-lint.py --self-test
|
|
$(PY) $(TOOLS)/rule-coverage.py --self-test
|
|
$(PY) $(TOOLS)/dep-weight.py --self-test
|
|
$(PY) $(TOOLS)/repo.py --self-test
|
|
$(PY) $(TOOLS)/task-done.py --self-test
|
|
$(PY) $(TOOLS)/status.py --self-test
|
|
$(PY) $(TOOLS)/facts.py --self-test
|
|
$(PY) $(TOOLS)/mutation-check.py --self-test
|
|
$(PY) $(TOOLS)/size-metrics.py --self-test
|
|
$(PY) $(TOOLS)/runtime-metrics.py --self-test
|
|
$(PY) $(TOOLS)/replay-test.py --self-test
|
|
$(PY) $(TOOLS)/design.py --self-test
|
|
$(PY) $(TOOLS)/trials.py --self-test
|
|
cargo run --release -q -p games-ground --example difficulty -- --self-test
|
|
$(PY) $(TOOLS)/edition-check.py --self-test
|
|
|
|
# T01 positive control: prove the environment fix, do not assume it. Runs
|
|
# every tool from a foreign working directory with a PATH that has no
|
|
# cargo on it. Before T01 this failed; if it fails again, the friction is
|
|
# back and CB-EV-0003's measurement is invalid.
|
|
env-test:
|
|
@cd / && env PATH=/usr/bin:/bin $(PY) $(TOOLS)/repo.py --self-test
|
|
@cd / && env PATH=/usr/bin:/bin $(PY) $(TOOLS)/rule-coverage.py --self-test >/dev/null \
|
|
&& echo " [ok ] rule-coverage runs from / with no cargo on PATH"
|
|
@cd / && env PATH=/usr/bin:/bin $(PY) $(TOOLS)/dep-weight.py --self-test >/dev/null \
|
|
&& echo " [ok ] dep-weight runs from / with no cargo on PATH"
|
|
@cd / && env PATH=/usr/bin:/bin $(PY) $(TOOLS)/cb-cost.py --self-test >/dev/null \
|
|
&& echo " [ok ] cb-cost runs from / with no cargo on PATH"
|
|
@cd / && env PATH=/usr/bin:/bin $(PY) $(TOOLS)/loop-lint.py --self-test >/dev/null \
|
|
&& echo " [ok ] loop-lint runs from / with no cargo on PATH"
|
|
@$(MAKE) -C $(REPO) coverage >/dev/null \
|
|
&& echo " [ok ] make -C <repo> works from any directory"
|
|
|
|
# M-D1-MUT (CB-WP-0005 T02): invert each acceptance row's property and
|
|
# require the verifying command to go red. Deliberately NOT in `make all`:
|
|
# it rebuilds the workspace once per mutated row. Run it on demand and in
|
|
# CI, not in the inner loop.
|
|
mutation-check:
|
|
$(PY) $(TOOLS)/mutation-check.py $(ARGS)
|
|
|
|
# T04: single source of fact (InnerLoop v1.2) — the DFD gate.
|
|
# facts.toml is GENERATED; facts-check fails if it disagrees with the
|
|
# instruments, or if a tagged artifact disagrees with it.
|
|
facts-check:
|
|
$(PY) $(TOOLS)/facts.py --check
|
|
|
|
facts-gen:
|
|
$(PY) $(TOOLS)/facts.py --gen
|
|
|
|
# CB-WP-0027 T04: what the players said, and where. Surfacing is the
|
|
# deliverable -- a commentary feature nobody can read is this project's
|
|
# signature failure in a new medium (ADR-0014 D6).
|
|
## what the players said while playing, and where
|
|
trials:
|
|
@$(PY) $(TOOLS)/trials.py
|
|
|
|
# CB-WP-0025 T06: the difficulty table (specs/RetrospectiveAnalysis.md §4).
|
|
# Winnable fraction from the solver plus a PLURAL policy panel -- a single
|
|
# policy's win rate may not be reported as a difficulty (§4.1).
|
|
## winnable fraction + a plural policy panel (never one bot's win rate)
|
|
difficulty:
|
|
@cargo run --release -q -p games-ground --example difficulty
|
|
|
|
# CB-REV-0002 #7: the H1 measurement harnesses were run by NO gate. Every
|
|
# number in CB-EV-0030 and CB-EV-0031 came from a manual invocation of an
|
|
# ungated binary -- so the assertion added to catch short cells was
|
|
# unreachable from `make`, and "what would the harness report if the work
|
|
# silently stopped" answered: green, and nothing else.
|
|
## the variant panels: ATTACK's value, and regulation under H1
|
|
panels:
|
|
@cargo run --release -q -p games-ground --example attack-value
|
|
@cargo run --release -q -p games-ground --example regulation
|
|
@cargo run --release -q -p games-ground --example perfect-recall
|
|
@cargo run --release -q -p games-ground --example h2-panel
|
|
|
|
# CB-WP-0022 T05: the design-finding register, reported over
|
|
# specs/GroundRules.md. Shows the QUEUE by default; the log of closed
|
|
# findings is a line, not a listing, because a default view that mixes
|
|
# them loses the queue property (ADR-0012 D5).
|
|
## the design-finding register: what is open, and what lacks a reproduction
|
|
design:
|
|
@$(PY) $(TOOLS)/design.py
|
|
|
|
# T03: one-shot orientation — workplans, next task, spend, fast gates.
|
|
# Cheap by design: no build. Start a session with this instead of grepping.
|
|
## one-shot orientation: workplans, next task, spend, fast gates
|
|
status:
|
|
@$(PY) $(TOOLS)/status.py
|
|
|
|
# T02: close a task — flip the workplan file, read the *measured* cost
|
|
# from the transcripts, push the hub event with real numbers. Refuses on an
|
|
# unknown or already-done task, and refuses to report an estimate.
|
|
# make task-done T=CB-WP-0004-T02
|
|
task-done:
|
|
@test -n "$(T)" || { echo "usage: make task-done T=CB-WP-0004-T02" >&2; exit 2; }
|
|
$(PY) $(TOOLS)/task-done.py $(T) $(ARGS)
|
|
|
|
# CB-WP-0007 T03 / InnerLoop v1.5: session shape for the window since the
|
|
# last commit. Deliberately NOT in `make all` — failing the build on
|
|
# context would block committing, and committing is the natural point to
|
|
# compact. A gate that blocks the remedy is a trap.
|
|
shape-budget: cost-test
|
|
$(PY) $(TOOLS)/cb-cost.py --shape-budget
|
|
|
|
# CB-01/CB-02: live spend since the last commit.
|
|
cost-budget: cost-test
|
|
$(PY) $(TOOLS)/cb-cost.py --budget
|
|
|
|
# CB-RES-0003 baseline: mechanical vs judgment turns.
|
|
cost-mix: cost-test
|
|
$(PY) $(TOOLS)/cb-cost.py --composition
|
|
|
|
cost-pin: cost-test
|
|
$(PY) $(TOOLS)/cb-cost.py --pin fc76445 --composition --by-task
|
|
|
|
## run all GROUND scenarios through cb-sim
|
|
sim:
|
|
$(IN_REPO) $(CARGO) run -q -p cb-sim -- $(REPO)/scenarios/ground/*.yaml
|
|
|
|
## Criterion benches (AM-6/AM-7)
|
|
bench:
|
|
$(IN_REPO) $(CARGO) bench -p games-ground
|
|
|
|
## InnerLoop positive control: run every bench once, no measurement.
|
|
## Fails if a workload stalls or produces the wrong event count.
|
|
bench-test:
|
|
$(IN_REPO) $(CARGO) bench -p games-ground --bench synthetic -- --test
|
|
|
|
|
|
## AM-2/AM-3 input: source LOC per crate (excludes tests would need tokei)
|
|
loc:
|
|
@$(IN_REPO) for d in crates/cb-kernel crates/cb-events crates/cb-game-runtime games/ground tools/cb-sim; do \
|
|
printf '%-28s %s\n' $$d "$$(find $$d/src -name '*.rs' | xargs cat | grep -vcE '^\s*(//|$$)')"; \
|
|
done
|
|
|
|
## every gate, in order. The one CI would run.
|
|
all: check test sim coverage size-metrics runtime-metrics am6 am7 am8 edition-check replay-test dep-weight self-tests env-test facts-check loop-lint bench-test panels
|