clay-borg/gates.toml
tegwick c4a8a227c0
Some checks failed
ci / check (push) Failing after 3s
CB-REV-0003: round 3, and three of four FATAL came from round 2's fixes
The pattern is now measured over three rounds: 5 fatal, then 3 (2 from the
previous round's corrections), then 4 (3 from them). The corrections are
not getting safer.

FATAL 1: round 2's short-cell assertion went into regulation.rs only.
attack-value.rs — which produced every number in CB-EV-0030's DARVO table
— still just warned, and the gate registered to close the finding claimed
the property for both.

FATAL 2, the sharpest of the three rounds: counting games proves they
STARTED. Stopping the engine after one round gives 200 games, all-zero
columns and exit 0 — byte for byte the signature CB-EV-0030 says the
instrumentation distinguishes from a real result. Both harnesses now
require every counted game to have reached an outcome over five rounds.

FATAL 3: round 2's `.csv` filter was applied to all three loops, so
catalog.yaml and rules_delta.yaml — whose missing digests were round 1's
finding — were recorded and then never compared, and never checked against
upstream at all. Only the parser loop filters now.

FATAL 4: five of six tiebreak comparators had no coverage. GR-E04's
tiebreak never executes in any scenario. All four are now covered and
mutation-verified; the Blame key needed compensating claims to be
reachable at all, since Blame also lowers the coalition score.

SERIOUS: "peak held" computed the same number as "peak assigned" for every
possible input — the real gap was that START_STRESS was an unchecked
constant, now read off the dealt state; cadence="none" was a pure
loophole, removed; sibling discovery swapped a hand-written list for
hand-written globs and missed metadata.json and VARIANT.md, both named in
the package's own changed_files — now walked, and it found them
immediately; and "~72,000 games" was unsourced, make panels runs 17,600.

Also separated two kinds of number that were presented alike: seats×games
is invariant, 363 and 29 vary 7.1%-11.5% across samples.

Round 4 owed. The conclusion is not that the work is nearly right — it is
that author-made corrections to measurement work should be assumed
defective until a fresh reader has attacked them.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-08 10:26:25 +02:00

197 lines
11 KiB
TOML

# The gate registry (ADR-0006 D3, CB-WP-0009 T02).
#
# A gate without an expiry is a permanent tax justified once. Every
# standing control mechanism gets an entry here saying what it checks,
# what it has actually **caught**, when its keep-or-kill argument is due,
# and what would retire it.
#
# `make gate-review` reports what is overdue and what has caught nothing.
# It reports; it does not fail the build — CB-RES-0005 §4: a gate that
# blocks the remedy when the metric breaches is a trap, not a gate.
#
# `caught` is the load-bearing field. An empty `caught` is not proof a
# gate is useless — it may be preventing rather than missing — but it
# means the argument has to be made out loud on `review_by`.
# Targets in `make all` that are build or acceptance steps rather than
# *control* gates — they measure the product, not how we work. Listed
# explicitly so a new target has to be classified rather than ignored;
# `loop-lint` fails when a target is in neither list.
not_control_gates = [
"check", "test", "sim", "bench-test", "size-metrics", "runtime-metrics",
"am6", "am7", "am8", "edition-check", "replay-test", "dep-weight", "self-tests", "env-test",
]
[[gate]]
id = "CB-01/CB-02"
name = "cost budget"
target = "cost-budget"
cadence = "manual"
checks = "spend since the last commit; soft $10, hard $22"
added = "2026-07-30"
review_by = "2026-11-30"
caught = [
"CB-WP-0005: hard breach forced the Phase C re-plan",
"CB-WP-0006 T07: $12.06 in one task, the pass's most expensive",
]
retire_if = "two consecutive passes never approach the soft line, or commits get small enough that the window is always trivial"
[[gate]]
id = "SH-1/SH-2"
name = "session-shape budget"
target = "shape-budget"
cadence = "manual"
checks = "mean and p90 context since the last commit (SH-3 retired as a gate, ADR-0008 D1 — still reported as a diagnostic)"
added = "2026-08-01"
review_by = "2026-11-30"
caught = [
"first run fired HARD at 656,574 against a 300,000 ceiling, which is what prompted the compaction before CB-WP-0008",
"CB-WP-0013: SH-3 retired from this gate. Its window is *since the last commit*, which during a pass is implementation work, and batching needs two calls whose inputs are known at once — which is orientation work. Measured: 0 batched turns in 150 in-pass responses across three passes, against 37.5% in the gap between two of them. A floor the window structurally excludes is not a target (ADR-0008 D1)",
]
retire_if = "context stops correlating with cost, or the model's context handling makes the number unactionable"
[[gate]]
id = "M-D1-MUT"
name = "mutation coverage of acceptance rows"
target = "mutation-check"
cadence = "manual"
checks = "each acceptance row's assertion must go red for a stated reason when mutated"
added = "2026-07-31"
review_by = "2026-12-31"
caught = [
"AM-6 measuring contention, not throughput",
"AM-5's 61% measurement error under load",
"peak RSS over-reported 3x",
"K10's first round trip not reproducing",
"the bench workload existing twice",
"AM-2's expect matching its own passing output (EXPECT-VACUOUS)",
"CB-WP-0015: the two clauses it had reported inert since CB-WP-0005 — AM-7's scaling ratio (a test named `replay_100k_events_is_linear_and_fast` that computed both throughputs, printed both, and never divided one by the other) and AM-8's N=10 (the runner did two). Both now red, 10/14, and the >=10-of-14 prediction met for the first time",
"CB-WP-0015: AM-4a's own mutation, stale since ADR-0008 D3 moved the target 250,000 -> 161,000 in CB-WP-0013. Reported HARNESS-BROKEN and refused to publish a score, which is the harness catching itself. The build-free half of that check is now a --self-test assertion, so it runs in `make all` instead of only on a full run",
]
retire_if = "a full pass adds rows without finding anything, twice running — the harness costs real money per run"
[[gate]]
id = "DFD"
name = "single source of fact"
target = "facts-check"
cadence = "all"
checks = "every tagged number in the docs matches the tool that measures it"
added = "2026-07-31"
review_by = "2026-12-31"
caught = [
"gr_scenarios stale at 21 after CB-WP-0008 T03 added three scenarios",
]
retire_if = "the untagged-literal count reaches zero and stays there, meaning the docs stopped restating measured numbers"
[[gate]]
id = "AM-1b"
name = "kernel spec->code link"
target = "coverage"
cadence = "all"
checks = "every numbered K-rule is named in the source; binds 2026-08-31"
added = "2026-07-31"
review_by = "2026-08-31"
caught = [
"3 unlinked K-rules at introduction (15/18); 18/18 today",
]
retire_if = "it stays at 100% through two passes that add kernel rules — at that point it is measuring a habit, not enforcing one"
[[gate]]
id = "META-20"
name = "meta budget"
target = "status"
cadence = "manual"
checks = "share of the trailing 5 passes spent on the loop itself; soft 20%. Purpose (InnerLoop v1.7): most spend on the task at hand, some on control, review and improving the process that carries the work forward"
added = "2026-08-01"
review_by = "2026-11-30"
caught = [
"its own cumulative-window defect, reported in CB-EV-0007 §3 and fixed by CB-WP-0009 T01",
"CB-WP-0019: it had a threshold and no stated purpose, which is why the number was argued three times. The maintainer set 80/20; writing that down exposed that the ratio and the window are a pair — one meta pass among three at parity cost already reads 33%, so 20% over a trailing 3 would have silently also demanded it be half-price. Now 20% over a trailing 5, which is one pass in five at normal cost",
]
retire_if = "product and meta stop being separable, or the share sits under the line for four passes without anyone consulting it"
# The phase setting (CB-WP-0019 T05). Uncomment, argue, and date it to move
# the split for a phase — stage 0 and a stabilisation phase do not deserve
# the same ratio. It REVERTS to META_SOFT_PCT on `review_by` unless
# re-argued, and `status.py --self-test` refuses one with no reason or no
# expiry, because a threshold anyone may move is not a threshold.
#
# meta_phase = { pct = 35, reason = "why this phase differs", review_by = "YYYY-MM-DD" }
[[gate]]
id = "LOOP-LINT"
name = "executable InnerLoop rules"
target = "loop-lint"
cadence = "all"
checks = "loadability, unmeasured verdicts, tier and chaos declarations, review trails, self-test entry points, and this registry"
added = "2026-07-30"
review_by = "2026-12-31"
caught = [
"four loadability breaches (401, 427, 406, 409 lines), each fixed structurally rather than by raising the limit",
"a reporting tool with no --self-test entry point (tools/repo.py)",
"CB-WP-0019: two new checks, and both fired on the pass that wrote them. `own-cost` caught CB-EV-0017 quoting its own pass's cost; `lifecycle` caught CB-WP-0019 still saying `active` with every task closed. The lifecycle check's FIRST version was itself wrong and its own self-test found it — it stripped the leading `status:` assuming frontmatter, which silently dropped a real task once the frontmatter said `ready` or `active`",
]
retire_if = "two passes run with no finding while artifacts keep growing — that would mean it is measuring the wrong properties"
[[gate]]
id = "CHAOS"
name = "the chaos roll"
target = ""
cadence = "manual"
checks = "d8 on each tier declaration, 12-declaration calibration window (window 2, opened 2026-08-03; window 1 ran at d4)"
added = "2026-07-30"
review_by = "2026-11-30"
caught = [
"CB-WP-0011: first fire in 6 declarations — d4=4 rolled stage 1 from structural L to S; the deleted survey would have opened on 2D toolkits while the existing text renderer was showing 24 of 41 view fields (CB-EV-0009 §1)",
"CB-WP-0017: d4=4 — second override in twelve, and the first to roll UP (structural S → M). It bought ADR-0010: the script's widening from 'it does one thing' to holding a drag, following the pointer and marking other elements would otherwise have landed under a tier-S provenance paragraph, silently outgrowing ADR-0007 D5. The ADR's own finding is that the permitted and forbidden designs are indistinguishable from outside, which demoted the vocabulary grep to a cheap first line and produced the two behavioural controls that replaced it",
"CB-WP-0012: d4=1, no override — and the contrast is the entry. Tier L at full weight deleted its own structural trigger: adversarial review withdrew the capability port the declaration was made to build (ADR-0007 D2), and corrected the survey's headline claim by 85x (128x -> 1.5x, CB-EV-0010 §2). Two passes on one subject at two tiers, priced: 0.123 $/response at L against 0.099 at S (CB-EV-0010 §5)",
]
retire_if = "an override changes nothing twice running (window 2 condition, CB-WP-0018 T04). Window 1's condition — no override changing the outcome — was NOT met: both did, so the mechanism was kept and the rate dropped d4 → d8 instead"
# VERDICT, CB-EV-0015 §5 (window closed 2026-08-02, 12 declarations, 2 overrides).
# Not retired: both overrides changed the outcome. CB-WP-0011 (L→S) bought a
# defect in the existing renderer that the deleted survey would have walked
# past, at 0.099 $/response against 0.123. CB-WP-0017 (S→M) bought the
# restatement of control 5 — at tier S the script would have outgrown
# ADR-0007 D5 under a one-paragraph commit note.
# RECOMMENDED, and owed to the next declaration as tier-M work: drop the rate
# d4 → d8 and open a second window of 12. Both overrides were informative
# BECAUSE they were rare; a mechanism firing on a quarter of declarations
# stops being a calibration and becomes the tier system.
# New retirement condition proposed: retire if an override changes nothing
# twice running.
[[gate]]
id = "GATE-REVIEW"
name = "this registry"
target = "gate-review"
cadence = "manual"
checks = "gates past their review date, and gates that have caught nothing"
added = "2026-08-01"
review_by = "2026-12-31"
caught = [
"CB-WP-0013: forced SH-3's re-justification and then its retirement. `make gate-review` had reported SH-3 as a standing breach for seven passes with zero actions taken, which is this registry's own ritual test (ADR-0006 D4); asking what it had ever caught is what exposed that the metric could not read above ~0% in the window it was gated on (ADR-0008 D1)",
]
retire_if = "it has retired, tightened, or forced the re-justification of nothing by its review date — then it is a ritual, and ADR-0006 D4 says rituals cash out or go"
[[gate]]
id = "CB-REV-0002/7"
name = "variant panels"
target = "panels"
cadence = "all"
checks = "both H1 measurement harnesses run; a short cell fails, and so does a cell whose games did not reach an outcome"
notes = """
Registered because the second adversarial review asked what the harness
would report if the work silently stopped, and the answer was "green, and
nothing else": `attack-value` and `regulation` were wired into no target,
so every figure in CB-EV-0030 and CB-EV-0031 came from a manual run of an
ungated binary -- including the assertion added to catch short cells,
which was unreachable from `make`.
CB-REV-0003 #1 then found this `checks` line claimed a property that held
for only one of the two harnesses: `attack-value`, which produced every
number in CB-EV-0030's table, still only warned. And #2 found that
counting games proves they STARTED: stopping the engine after one round
gave 200 games, all-zero columns and exit 0 -- the signature CB-EV-0030
says the instrumentation distinguishes from a result. Both now assert an
outcome was reached.
"""