Some checks failed
ci / check (push) Failing after 4s
The answer is "it depends what you call an information set", and the distinction is the result. 44,938 decision points, random play, 2/3/4/6 seats. Reading A — information set = the seat's current projection, which is what project(Viewer::Player(seat)) returns and what the page renders: 22 violations. Reading B — information set = the seat's observation history, every view seen and action taken in order: 0. The Reading A witness is concrete. Two histories reach a byte-identical view — round 3, Select step, same hand, same claimed Problem — where the seat had played SOLVE then GROUND-OU(protect) in one and SUPPORT then SOLVE in the other. The view does not tell the seat what it did, because our state is a snapshot rather than a history: selections clear each round and effects coincide, so a player cannot reconstruct their own past from the present. In a real game the player's memory supplies it; in the state, nothing does. That is precisely OpenSpiel's ObservationString vs InformationStateString split, arrived at here by measurement rather than read off. project() is an observation, not an information state. So Track B is not closed, it is constrained, and usefully: an extensive-form game built from this engine must key information sets on observation histories, never on project(). Both directions are asserted — Reading B empty AND Reading A non-empty — because if the sample stops finding Reading A violations the conclusion is unsupported and must be re-derived rather than quietly kept. And the check samples, so it can falsify perfect recall and cannot establish it: Reading B's zero means no counterexample was drawn, which is printed as such. Wired into make panels, so it is re-derived by the gate rather than by hand. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
197 lines
11 KiB
TOML
197 lines
11 KiB
TOML
# The gate registry (ADR-0006 D3, CB-WP-0009 T02).
|
|
#
|
|
# A gate without an expiry is a permanent tax justified once. Every
|
|
# standing control mechanism gets an entry here saying what it checks,
|
|
# what it has actually **caught**, when its keep-or-kill argument is due,
|
|
# and what would retire it.
|
|
#
|
|
# `make gate-review` reports what is overdue and what has caught nothing.
|
|
# It reports; it does not fail the build — CB-RES-0005 §4: a gate that
|
|
# blocks the remedy when the metric breaches is a trap, not a gate.
|
|
#
|
|
# `caught` is the load-bearing field. An empty `caught` is not proof a
|
|
# gate is useless — it may be preventing rather than missing — but it
|
|
# means the argument has to be made out loud on `review_by`.
|
|
|
|
# Targets in `make all` that are build or acceptance steps rather than
|
|
# *control* gates — they measure the product, not how we work. Listed
|
|
# explicitly so a new target has to be classified rather than ignored;
|
|
# `loop-lint` fails when a target is in neither list.
|
|
not_control_gates = [
|
|
"check", "test", "sim", "bench-test", "size-metrics", "runtime-metrics",
|
|
"am6", "am7", "am8", "edition-check", "replay-test", "dep-weight", "self-tests", "env-test",
|
|
]
|
|
|
|
[[gate]]
|
|
id = "CB-01/CB-02"
|
|
name = "cost budget"
|
|
target = "cost-budget"
|
|
cadence = "manual"
|
|
checks = "spend since the last commit; soft $10, hard $22"
|
|
added = "2026-07-30"
|
|
review_by = "2026-11-30"
|
|
caught = [
|
|
"CB-WP-0005: hard breach forced the Phase C re-plan",
|
|
"CB-WP-0006 T07: $12.06 in one task, the pass's most expensive",
|
|
]
|
|
retire_if = "two consecutive passes never approach the soft line, or commits get small enough that the window is always trivial"
|
|
|
|
[[gate]]
|
|
id = "SH-1/SH-2"
|
|
name = "session-shape budget"
|
|
target = "shape-budget"
|
|
cadence = "manual"
|
|
checks = "mean and p90 context since the last commit (SH-3 retired as a gate, ADR-0008 D1 — still reported as a diagnostic)"
|
|
added = "2026-08-01"
|
|
review_by = "2026-11-30"
|
|
caught = [
|
|
"first run fired HARD at 656,574 against a 300,000 ceiling, which is what prompted the compaction before CB-WP-0008",
|
|
"CB-WP-0013: SH-3 retired from this gate. Its window is *since the last commit*, which during a pass is implementation work, and batching needs two calls whose inputs are known at once — which is orientation work. Measured: 0 batched turns in 150 in-pass responses across three passes, against 37.5% in the gap between two of them. A floor the window structurally excludes is not a target (ADR-0008 D1)",
|
|
]
|
|
retire_if = "context stops correlating with cost, or the model's context handling makes the number unactionable"
|
|
|
|
[[gate]]
|
|
id = "M-D1-MUT"
|
|
name = "mutation coverage of acceptance rows"
|
|
target = "mutation-check"
|
|
cadence = "manual"
|
|
checks = "each acceptance row's assertion must go red for a stated reason when mutated"
|
|
added = "2026-07-31"
|
|
review_by = "2026-12-31"
|
|
caught = [
|
|
"AM-6 measuring contention, not throughput",
|
|
"AM-5's 61% measurement error under load",
|
|
"peak RSS over-reported 3x",
|
|
"K10's first round trip not reproducing",
|
|
"the bench workload existing twice",
|
|
"AM-2's expect matching its own passing output (EXPECT-VACUOUS)",
|
|
"CB-WP-0015: the two clauses it had reported inert since CB-WP-0005 — AM-7's scaling ratio (a test named `replay_100k_events_is_linear_and_fast` that computed both throughputs, printed both, and never divided one by the other) and AM-8's N=10 (the runner did two). Both now red, 10/14, and the >=10-of-14 prediction met for the first time",
|
|
"CB-WP-0015: AM-4a's own mutation, stale since ADR-0008 D3 moved the target 250,000 -> 161,000 in CB-WP-0013. Reported HARNESS-BROKEN and refused to publish a score, which is the harness catching itself. The build-free half of that check is now a --self-test assertion, so it runs in `make all` instead of only on a full run",
|
|
]
|
|
retire_if = "a full pass adds rows without finding anything, twice running — the harness costs real money per run"
|
|
|
|
[[gate]]
|
|
id = "DFD"
|
|
name = "single source of fact"
|
|
target = "facts-check"
|
|
cadence = "all"
|
|
checks = "every tagged number in the docs matches the tool that measures it"
|
|
added = "2026-07-31"
|
|
review_by = "2026-12-31"
|
|
caught = [
|
|
"gr_scenarios stale at 21 after CB-WP-0008 T03 added three scenarios",
|
|
]
|
|
retire_if = "the untagged-literal count reaches zero and stays there, meaning the docs stopped restating measured numbers"
|
|
|
|
[[gate]]
|
|
id = "AM-1b"
|
|
name = "kernel spec->code link"
|
|
target = "coverage"
|
|
cadence = "all"
|
|
checks = "every numbered K-rule is named in the source; binds 2026-08-31"
|
|
added = "2026-07-31"
|
|
review_by = "2026-08-31"
|
|
caught = [
|
|
"3 unlinked K-rules at introduction (15/18); 18/18 today",
|
|
]
|
|
retire_if = "it stays at 100% through two passes that add kernel rules — at that point it is measuring a habit, not enforcing one"
|
|
|
|
[[gate]]
|
|
id = "META-20"
|
|
name = "meta budget"
|
|
target = "status"
|
|
cadence = "manual"
|
|
checks = "share of the trailing 5 passes spent on the loop itself; soft 20%. Purpose (InnerLoop v1.7): most spend on the task at hand, some on control, review and improving the process that carries the work forward"
|
|
added = "2026-08-01"
|
|
review_by = "2026-11-30"
|
|
caught = [
|
|
"its own cumulative-window defect, reported in CB-EV-0007 §3 and fixed by CB-WP-0009 T01",
|
|
"CB-WP-0019: it had a threshold and no stated purpose, which is why the number was argued three times. The maintainer set 80/20; writing that down exposed that the ratio and the window are a pair — one meta pass among three at parity cost already reads 33%, so 20% over a trailing 3 would have silently also demanded it be half-price. Now 20% over a trailing 5, which is one pass in five at normal cost",
|
|
]
|
|
retire_if = "product and meta stop being separable, or the share sits under the line for four passes without anyone consulting it"
|
|
|
|
# The phase setting (CB-WP-0019 T05). Uncomment, argue, and date it to move
|
|
# the split for a phase — stage 0 and a stabilisation phase do not deserve
|
|
# the same ratio. It REVERTS to META_SOFT_PCT on `review_by` unless
|
|
# re-argued, and `status.py --self-test` refuses one with no reason or no
|
|
# expiry, because a threshold anyone may move is not a threshold.
|
|
#
|
|
# meta_phase = { pct = 35, reason = "why this phase differs", review_by = "YYYY-MM-DD" }
|
|
|
|
[[gate]]
|
|
id = "LOOP-LINT"
|
|
name = "executable InnerLoop rules"
|
|
target = "loop-lint"
|
|
cadence = "all"
|
|
checks = "loadability, unmeasured verdicts, tier and chaos declarations, review trails, self-test entry points, and this registry"
|
|
added = "2026-07-30"
|
|
review_by = "2026-12-31"
|
|
caught = [
|
|
"four loadability breaches (401, 427, 406, 409 lines), each fixed structurally rather than by raising the limit",
|
|
"a reporting tool with no --self-test entry point (tools/repo.py)",
|
|
"CB-WP-0019: two new checks, and both fired on the pass that wrote them. `own-cost` caught CB-EV-0017 quoting its own pass's cost; `lifecycle` caught CB-WP-0019 still saying `active` with every task closed. The lifecycle check's FIRST version was itself wrong and its own self-test found it — it stripped the leading `status:` assuming frontmatter, which silently dropped a real task once the frontmatter said `ready` or `active`",
|
|
]
|
|
retire_if = "two passes run with no finding while artifacts keep growing — that would mean it is measuring the wrong properties"
|
|
|
|
[[gate]]
|
|
id = "CHAOS"
|
|
name = "the chaos roll"
|
|
target = ""
|
|
cadence = "manual"
|
|
checks = "d8 on each tier declaration, 12-declaration calibration window (window 2, opened 2026-08-03; window 1 ran at d4)"
|
|
added = "2026-07-30"
|
|
review_by = "2026-11-30"
|
|
caught = [
|
|
"CB-WP-0011: first fire in 6 declarations — d4=4 rolled stage 1 from structural L to S; the deleted survey would have opened on 2D toolkits while the existing text renderer was showing 24 of 41 view fields (CB-EV-0009 §1)",
|
|
"CB-WP-0017: d4=4 — second override in twelve, and the first to roll UP (structural S → M). It bought ADR-0010: the script's widening from 'it does one thing' to holding a drag, following the pointer and marking other elements would otherwise have landed under a tier-S provenance paragraph, silently outgrowing ADR-0007 D5. The ADR's own finding is that the permitted and forbidden designs are indistinguishable from outside, which demoted the vocabulary grep to a cheap first line and produced the two behavioural controls that replaced it",
|
|
"CB-WP-0012: d4=1, no override — and the contrast is the entry. Tier L at full weight deleted its own structural trigger: adversarial review withdrew the capability port the declaration was made to build (ADR-0007 D2), and corrected the survey's headline claim by 85x (128x -> 1.5x, CB-EV-0010 §2). Two passes on one subject at two tiers, priced: 0.123 $/response at L against 0.099 at S (CB-EV-0010 §5)",
|
|
]
|
|
retire_if = "an override changes nothing twice running (window 2 condition, CB-WP-0018 T04). Window 1's condition — no override changing the outcome — was NOT met: both did, so the mechanism was kept and the rate dropped d4 → d8 instead"
|
|
# VERDICT, CB-EV-0015 §5 (window closed 2026-08-02, 12 declarations, 2 overrides).
|
|
# Not retired: both overrides changed the outcome. CB-WP-0011 (L→S) bought a
|
|
# defect in the existing renderer that the deleted survey would have walked
|
|
# past, at 0.099 $/response against 0.123. CB-WP-0017 (S→M) bought the
|
|
# restatement of control 5 — at tier S the script would have outgrown
|
|
# ADR-0007 D5 under a one-paragraph commit note.
|
|
# RECOMMENDED, and owed to the next declaration as tier-M work: drop the rate
|
|
# d4 → d8 and open a second window of 12. Both overrides were informative
|
|
# BECAUSE they were rare; a mechanism firing on a quarter of declarations
|
|
# stops being a calibration and becomes the tier system.
|
|
# New retirement condition proposed: retire if an override changes nothing
|
|
# twice running.
|
|
|
|
[[gate]]
|
|
id = "GATE-REVIEW"
|
|
name = "this registry"
|
|
target = "gate-review"
|
|
cadence = "manual"
|
|
checks = "gates past their review date, and gates that have caught nothing"
|
|
added = "2026-08-01"
|
|
review_by = "2026-12-31"
|
|
caught = [
|
|
"CB-WP-0013: forced SH-3's re-justification and then its retirement. `make gate-review` had reported SH-3 as a standing breach for seven passes with zero actions taken, which is this registry's own ritual test (ADR-0006 D4); asking what it had ever caught is what exposed that the metric could not read above ~0% in the window it was gated on (ADR-0008 D1)",
|
|
]
|
|
retire_if = "it has retired, tightened, or forced the re-justification of nothing by its review date — then it is a ritual, and ADR-0006 D4 says rituals cash out or go"
|
|
|
|
[[gate]]
|
|
id = "CB-REV-0002/7"
|
|
name = "variant panels"
|
|
target = "panels"
|
|
cadence = "all"
|
|
checks = "the measurement harnesses run; a short cell fails, a cell whose games did not reach an outcome fails, and perfect recall is re-derived on both readings"
|
|
notes = """
|
|
Registered because the second adversarial review asked what the harness
|
|
would report if the work silently stopped, and the answer was "green, and
|
|
nothing else": `attack-value` and `regulation` were wired into no target,
|
|
so every figure in CB-EV-0030 and CB-EV-0031 came from a manual run of an
|
|
ungated binary -- including the assertion added to catch short cells,
|
|
which was unreachable from `make`.
|
|
|
|
CB-REV-0003 #1 then found this `checks` line claimed a property that held
|
|
for only one of the two harnesses: `attack-value`, which produced every
|
|
number in CB-EV-0030's table, still only warned. And #2 found that
|
|
counting games proves they STARTED: stopping the engine after one round
|
|
gave 200 games, all-zero columns and exit 0 -- the signature CB-EV-0030
|
|
says the instrumentation distinguishes from a result. Both now assert an
|
|
outcome was reached.
|
|
"""
|