Some checks failed
ci / check (push) Failing after 4s
The answer is "it depends what you call an information set", and the distinction is the result. 44,938 decision points, random play, 2/3/4/6 seats. Reading A — information set = the seat's current projection, which is what project(Viewer::Player(seat)) returns and what the page renders: 22 violations. Reading B — information set = the seat's observation history, every view seen and action taken in order: 0. The Reading A witness is concrete. Two histories reach a byte-identical view — round 3, Select step, same hand, same claimed Problem — where the seat had played SOLVE then GROUND-OU(protect) in one and SUPPORT then SOLVE in the other. The view does not tell the seat what it did, because our state is a snapshot rather than a history: selections clear each round and effects coincide, so a player cannot reconstruct their own past from the present. In a real game the player's memory supplies it; in the state, nothing does. That is precisely OpenSpiel's ObservationString vs InformationStateString split, arrived at here by measurement rather than read off. project() is an observation, not an information state. So Track B is not closed, it is constrained, and usefully: an extensive-form game built from this engine must key information sets on observation histories, never on project(). Both directions are asserted — Reading B empty AND Reading A non-empty — because if the sample stops finding Reading A violations the conclusion is unsupported and must be re-derived rather than quietly kept. And the check samples, so it can falsify perfect recall and cannot establish it: Reading B's zero means no counterexample was drawn, which is printed as such. Wired into make panels, so it is re-derived by the gate rather than by hand. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
303 lines
13 KiB
Makefile
303 lines
13 KiB
Makefile
# One command surface (InnerLoop §agentic-efficiency #3). Deterministic,
|
|
# greppable output; precursor of the `cb` CLI.
|
|
#
|
|
# CB-WP-0004 T01: every target here runs from a clean shell, from any
|
|
# directory, with no prefix. Invoke as `make -C <repo> <target>` from
|
|
# elsewhere. No target requires `cd` or `export PATH` — CB-RES-0003
|
|
# measured 84 turns and $15.33 spent on exactly those two prefixes.
|
|
|
|
# Absolute path to this Makefile's directory, so recipes never depend on
|
|
# the caller's working directory.
|
|
REPO := $(patsubst %/,%,$(dir $(abspath $(lastword $(MAKEFILE_LIST)))))
|
|
|
|
# Locate cargo instead of requiring it on the inherited PATH. Mirrors
|
|
# tools/repo.py:cargo_bin() — kept in sync by `make env-test`.
|
|
CARGO := $(firstword $(shell command -v cargo 2>/dev/null) \
|
|
$(wildcard $(HOME)/.cargo/bin/cargo) \
|
|
$(wildcard /usr/local/cargo/bin/cargo) \
|
|
cargo)
|
|
export PATH := $(dir $(CARGO)):$(PATH)
|
|
|
|
PY := python3
|
|
TOOLS := $(REPO)/tools
|
|
|
|
# `make` with no target lists what there is, rather than running the
|
|
# heaviest thing in the file. The Makefile's own header calls itself "one
|
|
# command surface" -- a surface you have to read the source of is not one.
|
|
.DEFAULT_GOAL := help
|
|
|
|
# A trial's name. Timestamped so two sessions on one day cannot overwrite
|
|
# each other's notes, and prefixed with the date because `tools/trials.py`
|
|
# reads the age from there (falling back to mtime).
|
|
TRIAL_NAME := $(shell date +%Y-%m-%d-%H%M)$(if $(SLUG),-$(SLUG),)
|
|
PLAYERS ?= 3
|
|
PORT ?= 0
|
|
|
|
# Every cargo recipe runs at the repo root; the shell does not persist cd.
|
|
IN_REPO := cd $(REPO) &&
|
|
|
|
.PHONY: help ground check test sim bench bench-test coverage dep-weight cost cost-test cost-pin cost-budget shape-budget cost-mix loop-lint self-tests env-test task-done status facts-check facts-gen mutation-check size-metrics runtime-metrics build-time am6 am7 am8 edition-check replay-test loc play gate-review panels all \
|
|
design difficulty trials
|
|
# `design`, `difficulty` and `trials` were added by CB-WP-0022, CB-WP-0025
|
|
# and CB-WP-0027 and none was declared here. Only `trials` revealed it, by
|
|
# colliding with the trials/ DIRECTORY -- Make saw an up-to-date file and
|
|
# ran nothing. The other two work by luck: no file happens to share their
|
|
# name. A target that is a command, not a file, belongs on this line.
|
|
|
|
## list every target, with what it does
|
|
## `make help` is also what bare `make` runs.
|
|
help:
|
|
@echo "clay-borg — make <target> (bare \`make\` shows this)"
|
|
@echo
|
|
@awk '/^## /{ sub(/^## /,""); d[++n]=$$0; next } \
|
|
/^[a-zA-Z][a-zA-Z0-9_-]*:/{ \
|
|
if (n) { split($$1,a,":"); printf " %-16s %s\n", a[1], d[1]; \
|
|
for (i=2;i<=n;i++) printf " %-16s %s\n", "", d[i]; print "" } \
|
|
n=0; next } \
|
|
{ n=0 }' $(MAKEFILE_LIST)
|
|
@echo " Variables: PLAYERS=3 PORT=0 SLUG=<name> ARGS=\"...\""
|
|
@echo
|
|
@echo " Undocumented (instruments and gate internals):"
|
|
@awk '/^## /{ d=1; next } \
|
|
/^[a-zA-Z][a-zA-Z0-9_-]*:/{ if (!d) { split($$1,a,":"); print a[1] } d=0; next } \
|
|
{ d=0 }' $(MAKEFILE_LIST) | sort -u | tr "\n" " " | fold -s -w 66 | sed "s/^/ /"
|
|
@echo
|
|
|
|
## play GROUND in a browser, recording a trial (the usual way in)
|
|
## Opens a URL; game on the left, notes on the right. Everything you
|
|
## type in the notes panel is bound to the position you typed it at.
|
|
## `make ground PLAYERS=2 SLUG=darvo-confusion`
|
|
## Read the notes back afterwards with `make trials`.
|
|
ground:
|
|
@mkdir -p $(REPO)/trials
|
|
@echo " trial: trials/$(TRIAL_NAME).md (notes) + .yaml (the game)"
|
|
$(IN_REPO) $(CARGO) run -q -p cb-play -- \
|
|
--players $(PLAYERS) --serve $(PORT) \
|
|
--record trials/$(TRIAL_NAME).yaml \
|
|
--trial trials/$(TRIAL_NAME).md $(ARGS)
|
|
|
|
## fmt + clippy (deny warnings) + HashMap deny-lint
|
|
check:
|
|
$(IN_REPO) $(CARGO) fmt --all --check
|
|
$(IN_REPO) $(CARGO) clippy --workspace --all-targets -- -D warnings
|
|
|
|
## play GROUND from the terminal (INTENT stage 0's CLI player).
|
|
## `make play ARGS="--all-bots --bot random"` to watch one instead.
|
|
play:
|
|
$(IN_REPO) $(CARGO) run -q -p cb-play -- $(ARGS)
|
|
|
|
## which control gates are due for a keep-or-kill argument (ADR-0006 D3)
|
|
gate-review:
|
|
$(PY) $(TOOLS)/gate-review.py
|
|
|
|
## unit + scenario-format tests
|
|
test:
|
|
$(IN_REPO) $(CARGO) test --workspace
|
|
|
|
## AM-4 dependency weight: third-party lines behind each budget
|
|
dep-weight:
|
|
$(PY) $(TOOLS)/dep-weight.py
|
|
|
|
## AM-1 rule coverage: which GROUND rules a scenario exercises
|
|
coverage:
|
|
$(PY) $(TOOLS)/rule-coverage.py
|
|
|
|
# K10: a bundle must re-execute, and must be able to fail. Four controls
|
|
# from ADR-0005 §6, including truncate-by-one-byte and mutated-seed.
|
|
replay-test:
|
|
$(PY) $(TOOLS)/replay-test.py
|
|
|
|
# AM-6 gate. Runs in RELEASE, where headroom is ~24x; the same assertion
|
|
# in debug has ~3.4x and flaked under load. The target is unchanged — only
|
|
# where it is measured.
|
|
am6:
|
|
$(IN_REPO) $(CARGO) test --release -p games-ground --all-features \
|
|
am6_throughput -- --ignored --nocapture --test-threads=1
|
|
|
|
# ADR-0011 D2: is the vendored edition still what ground-game published?
|
|
# Reports "upstream not checked out" as its own outcome, never a pass.
|
|
edition-check:
|
|
$(PY) $(TOOLS)/edition-check.py
|
|
|
|
# AM-8 N=10 determinism gate. One scenario, ten same-seed replays, all
|
|
# compared to the first. Not all 25: `make sim` already runs K8's double-
|
|
# run over every scenario, and repeating that eight more times costs 47 s
|
|
# per build to re-answer a question the second run already answered. The
|
|
# extra runs exist for the probabilistic class, and one workload gives
|
|
# that class its ten samples — see scenario::run_n.
|
|
am8:
|
|
$(IN_REPO) $(CARGO) run -q -p cb-sim -- --runs 10 \
|
|
$(REPO)/scenarios/ground/gr-r06-round-resolve.yaml
|
|
|
|
# AM-7 scaling gate. Same release / single-thread reasoning as am6 and more
|
|
# so: a *ratio* of two timings taken under varying contention is worse than
|
|
# one reading, because the noise multiplies rather than cancels.
|
|
am7:
|
|
$(IN_REPO) $(CARGO) test --release -p games-ground --all-features \
|
|
am7_cost_per_event -- --ignored --nocapture --test-threads=1
|
|
|
|
# AM-9 peak RSS (fast, gated). AM-5 needs a clean build — see build-time.
|
|
runtime-metrics:
|
|
$(PY) $(TOOLS)/runtime-metrics.py --fast
|
|
|
|
# AM-5: clean release build, ~90s, into a throwaway CARGO_TARGET_DIR so the
|
|
# working cache survives. Reported not gated per GameKernel §5.
|
|
build-time:
|
|
$(PY) $(TOOLS)/runtime-metrics.py
|
|
|
|
# AM-2 (M-D1-SPL, AM-1's anti-gaming pair) and AM-3.
|
|
size-metrics:
|
|
$(PY) $(TOOLS)/size-metrics.py
|
|
|
|
# M-D2-CST (specs/CostAccounting.md). cost-test is the positive control and
|
|
# runs first: a cost number from an unverified collector is void.
|
|
cost: cost-test
|
|
$(PY) $(TOOLS)/cb-cost.py --composition --by-task
|
|
|
|
cost-test:
|
|
$(PY) $(TOOLS)/cb-cost.py --self-test
|
|
|
|
# InnerLoop rules that are mechanically checkable (CB-WP-0003 T01).
|
|
loop-lint:
|
|
$(PY) $(TOOLS)/loop-lint.py
|
|
|
|
# Positive control for every reporting tool, per InnerLoop v1.1 Step 5.
|
|
## positive control for every reporting tool
|
|
self-tests:
|
|
$(PY) $(TOOLS)/cb-cost.py --self-test
|
|
$(PY) $(TOOLS)/loop-lint.py --self-test
|
|
$(PY) $(TOOLS)/rule-coverage.py --self-test
|
|
$(PY) $(TOOLS)/dep-weight.py --self-test
|
|
$(PY) $(TOOLS)/repo.py --self-test
|
|
$(PY) $(TOOLS)/task-done.py --self-test
|
|
$(PY) $(TOOLS)/status.py --self-test
|
|
$(PY) $(TOOLS)/facts.py --self-test
|
|
$(PY) $(TOOLS)/mutation-check.py --self-test
|
|
$(PY) $(TOOLS)/size-metrics.py --self-test
|
|
$(PY) $(TOOLS)/runtime-metrics.py --self-test
|
|
$(PY) $(TOOLS)/replay-test.py --self-test
|
|
$(PY) $(TOOLS)/design.py --self-test
|
|
$(PY) $(TOOLS)/trials.py --self-test
|
|
cargo run --release -q -p games-ground --example difficulty -- --self-test
|
|
$(PY) $(TOOLS)/edition-check.py --self-test
|
|
|
|
# T01 positive control: prove the environment fix, do not assume it. Runs
|
|
# every tool from a foreign working directory with a PATH that has no
|
|
# cargo on it. Before T01 this failed; if it fails again, the friction is
|
|
# back and CB-EV-0003's measurement is invalid.
|
|
env-test:
|
|
@cd / && env PATH=/usr/bin:/bin $(PY) $(TOOLS)/repo.py --self-test
|
|
@cd / && env PATH=/usr/bin:/bin $(PY) $(TOOLS)/rule-coverage.py --self-test >/dev/null \
|
|
&& echo " [ok ] rule-coverage runs from / with no cargo on PATH"
|
|
@cd / && env PATH=/usr/bin:/bin $(PY) $(TOOLS)/dep-weight.py --self-test >/dev/null \
|
|
&& echo " [ok ] dep-weight runs from / with no cargo on PATH"
|
|
@cd / && env PATH=/usr/bin:/bin $(PY) $(TOOLS)/cb-cost.py --self-test >/dev/null \
|
|
&& echo " [ok ] cb-cost runs from / with no cargo on PATH"
|
|
@cd / && env PATH=/usr/bin:/bin $(PY) $(TOOLS)/loop-lint.py --self-test >/dev/null \
|
|
&& echo " [ok ] loop-lint runs from / with no cargo on PATH"
|
|
@$(MAKE) -C $(REPO) coverage >/dev/null \
|
|
&& echo " [ok ] make -C <repo> works from any directory"
|
|
|
|
# M-D1-MUT (CB-WP-0005 T02): invert each acceptance row's property and
|
|
# require the verifying command to go red. Deliberately NOT in `make all`:
|
|
# it rebuilds the workspace once per mutated row. Run it on demand and in
|
|
# CI, not in the inner loop.
|
|
mutation-check:
|
|
$(PY) $(TOOLS)/mutation-check.py $(ARGS)
|
|
|
|
# T04: single source of fact (InnerLoop v1.2) — the DFD gate.
|
|
# facts.toml is GENERATED; facts-check fails if it disagrees with the
|
|
# instruments, or if a tagged artifact disagrees with it.
|
|
facts-check:
|
|
$(PY) $(TOOLS)/facts.py --check
|
|
|
|
facts-gen:
|
|
$(PY) $(TOOLS)/facts.py --gen
|
|
|
|
# CB-WP-0027 T04: what the players said, and where. Surfacing is the
|
|
# deliverable -- a commentary feature nobody can read is this project's
|
|
# signature failure in a new medium (ADR-0014 D6).
|
|
## what the players said while playing, and where
|
|
trials:
|
|
@$(PY) $(TOOLS)/trials.py
|
|
|
|
# CB-WP-0025 T06: the difficulty table (specs/RetrospectiveAnalysis.md §4).
|
|
# Winnable fraction from the solver plus a PLURAL policy panel -- a single
|
|
# policy's win rate may not be reported as a difficulty (§4.1).
|
|
## winnable fraction + a plural policy panel (never one bot's win rate)
|
|
difficulty:
|
|
@cargo run --release -q -p games-ground --example difficulty
|
|
|
|
# CB-REV-0002 #7: the H1 measurement harnesses were run by NO gate. Every
|
|
# number in CB-EV-0030 and CB-EV-0031 came from a manual invocation of an
|
|
# ungated binary -- so the assertion added to catch short cells was
|
|
# unreachable from `make`, and "what would the harness report if the work
|
|
# silently stopped" answered: green, and nothing else.
|
|
## the variant panels: ATTACK's value, and regulation under H1
|
|
panels:
|
|
@cargo run --release -q -p games-ground --example attack-value
|
|
@cargo run --release -q -p games-ground --example regulation
|
|
@cargo run --release -q -p games-ground --example perfect-recall
|
|
|
|
# CB-WP-0022 T05: the design-finding register, reported over
|
|
# specs/GroundRules.md. Shows the QUEUE by default; the log of closed
|
|
# findings is a line, not a listing, because a default view that mixes
|
|
# them loses the queue property (ADR-0012 D5).
|
|
## the design-finding register: what is open, and what lacks a reproduction
|
|
design:
|
|
@$(PY) $(TOOLS)/design.py
|
|
|
|
# T03: one-shot orientation — workplans, next task, spend, fast gates.
|
|
# Cheap by design: no build. Start a session with this instead of grepping.
|
|
## one-shot orientation: workplans, next task, spend, fast gates
|
|
status:
|
|
@$(PY) $(TOOLS)/status.py
|
|
|
|
# T02: close a task — flip the workplan file, read the *measured* cost
|
|
# from the transcripts, push the hub event with real numbers. Refuses on an
|
|
# unknown or already-done task, and refuses to report an estimate.
|
|
# make task-done T=CB-WP-0004-T02
|
|
task-done:
|
|
@test -n "$(T)" || { echo "usage: make task-done T=CB-WP-0004-T02" >&2; exit 2; }
|
|
$(PY) $(TOOLS)/task-done.py $(T) $(ARGS)
|
|
|
|
# CB-WP-0007 T03 / InnerLoop v1.5: session shape for the window since the
|
|
# last commit. Deliberately NOT in `make all` — failing the build on
|
|
# context would block committing, and committing is the natural point to
|
|
# compact. A gate that blocks the remedy is a trap.
|
|
shape-budget: cost-test
|
|
$(PY) $(TOOLS)/cb-cost.py --shape-budget
|
|
|
|
# CB-01/CB-02: live spend since the last commit.
|
|
cost-budget: cost-test
|
|
$(PY) $(TOOLS)/cb-cost.py --budget
|
|
|
|
# CB-RES-0003 baseline: mechanical vs judgment turns.
|
|
cost-mix: cost-test
|
|
$(PY) $(TOOLS)/cb-cost.py --composition
|
|
|
|
cost-pin: cost-test
|
|
$(PY) $(TOOLS)/cb-cost.py --pin fc76445 --composition --by-task
|
|
|
|
## run all GROUND scenarios through cb-sim
|
|
sim:
|
|
$(IN_REPO) $(CARGO) run -q -p cb-sim -- $(REPO)/scenarios/ground/*.yaml
|
|
|
|
## Criterion benches (AM-6/AM-7)
|
|
bench:
|
|
$(IN_REPO) $(CARGO) bench -p games-ground
|
|
|
|
## InnerLoop positive control: run every bench once, no measurement.
|
|
## Fails if a workload stalls or produces the wrong event count.
|
|
bench-test:
|
|
$(IN_REPO) $(CARGO) bench -p games-ground --bench synthetic -- --test
|
|
|
|
|
|
## AM-2/AM-3 input: source LOC per crate (excludes tests would need tokei)
|
|
loc:
|
|
@$(IN_REPO) for d in crates/cb-kernel crates/cb-events crates/cb-game-runtime games/ground tools/cb-sim; do \
|
|
printf '%-28s %s\n' $$d "$$(find $$d/src -name '*.rs' | xargs cat | grep -vcE '^\s*(//|$$)')"; \
|
|
done
|
|
|
|
## every gate, in order. The one CI would run.
|
|
all: check test sim coverage size-metrics runtime-metrics am6 am7 am8 edition-check replay-test dep-weight self-tests env-test facts-check loop-lint bench-test panels
|