clay-borg/gates.toml

165 lines
9.7 KiB
TOML
Raw Normal View History

# The gate registry (ADR-0006 D3, CB-WP-0009 T02).
#
# A gate without an expiry is a permanent tax justified once. Every
# standing control mechanism gets an entry here saying what it checks,
# what it has actually **caught**, when its keep-or-kill argument is due,
# and what would retire it.
#
# `make gate-review` reports what is overdue and what has caught nothing.
# It reports; it does not fail the build — CB-RES-0005 §4: a gate that
# blocks the remedy when the metric breaches is a trap, not a gate.
#
# `caught` is the load-bearing field. An empty `caught` is not proof a
# gate is useless — it may be preventing rather than missing — but it
# means the argument has to be made out loud on `review_by`.
# Targets in `make all` that are build or acceptance steps rather than
# *control* gates — they measure the product, not how we work. Listed
# explicitly so a new target has to be classified rather than ignored;
# `loop-lint` fails when a target is in neither list.
not_control_gates = [
"check", "test", "sim", "bench-test", "size-metrics", "runtime-metrics",
CB-WP-0015: the two inert clauses, AM-7 scaling and AM-8 N=10 Provenance (tier S, one paragraph in lieu of survey and ADR): the two clauses mutation-check has reported inert since CB-WP-0005. AM-7's scaling ratio was held up by a test literally named replay_100k_events_is_linear_and_fast that computed both throughputs, printed both, and never divided one by the other. AM-8's N=10 was held up by a runner that does two. Both are now red. AM-7 3/3, AM-8 2/2, M-D1-MUT 10/14, and ADR-0005's >=10-of-14 prediction MET for the first time. Neither was closed by amending the question away, which was the live risk: the denominator is unchanged and the four unenforced rows are the four already unenforceable. AM-7 needed three estimators. Best-of-N per leg then divide (AM-6's, correct for a floor on one number) gave 0.581-1.085 on an unchanged binary; legs back-to-back gave medians 0.931-1.004; legs interleaved at fold granularity give 0.987/0.991/0.989, and 0.989 under 8-way CPU contention while absolute throughput fell 4x. The INDETERMINATE guard demanded unanimity and failed a good measurement over one sample 0.001 under the floor; it now requires a two-thirds majority. The control that matters: AM-6's constant-cost mutation halves throughput and leaves this ratio at 0.999x green, so AM-7 is not a second AM-6. AM-8 kept N=10 because the measurement said so. Perturbing the RNG only from its fourth construction on: --runs 2 PASSES, --runs 10 fails. A late-onset divergence is deterministic, not flaky, so it is a control rather than a coin flip. Ten runs live on one scenario (make am8, ~2s) rather than all 25 (47s a build). GameKernel 5b records it. The full run also found AM-4a's own mutation stale since ADR-0008 D3 moved the target 250,000 -> 161,000 in CB-WP-0013 -- reported HARNESS-BROKEN, no score published. The build-free half of that check is now a --self-test assertion, so make all catches the next one. mutation-check clauses may now carry their own verify and mutation, and then the enforced flag is measured rather than declared; a declaration disagreeing with its measurement is refused. make all exits 0. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-02 14:07:08 +02:00
"am6", "am7", "am8", "replay-test", "dep-weight", "self-tests", "env-test",
]
[[gate]]
id = "CB-01/CB-02"
name = "cost budget"
target = "cost-budget"
checks = "spend since the last commit; soft $10, hard $22"
added = "2026-07-30"
review_by = "2026-11-30"
caught = [
"CB-WP-0005: hard breach forced the Phase C re-plan",
"CB-WP-0006 T07: $12.06 in one task, the pass's most expensive",
]
retire_if = "two consecutive passes never approach the soft line, or commits get small enough that the window is always trivial"
[[gate]]
CB-WP-0013-T02/T03: retire SH-3 as a gate; correct AM-4a and its target ADR-0008, tier M (survey and ADR merged). D1 — SH-3 retired as a gate, kept as a diagnostic. Investigating it found a third defect, deeper than the two this pass was declared on. Re-deriving batching from the raw transcripts, independently of cb-cost: CB-WP-0011 pass 54 with tools 0 batched 0.0% gap -> next decl 16 with tools 6 batched 37.5% CB-WP-0012 pass 86 with tools 0 batched 0.0% gap -> next decl 10 with tools 1 batched 10.0% CB-WP-0013 so far 10 with tools 0 batched 0.0% Zero batched turns in 150 in-pass responses; 37.5% in one gap, above the 20% floor. Batching needs two calls whose inputs are known at once — orientation work. Implementation consumes each step's result before the next. SH-3's window is since the last commit, which during a pass is always implementation. The metric could not read above ~0% in the window it was gated on. A floor the window structurally excludes is not a target. This pass's own declaration was also wrong: it claimed batching "has got worse" (7.8-8.6% vs 1.1-6.3%). Differently-placed windows, not different behaviour. Withdrawn — the same class of error, in the pass written to correct it. Not retargeting to match the measurement: the floor was not moved to 6%, the gate was removed on an argument about what the quantity is worth. The number is still reported; only the verdict is gone. D2/D3 — AM-4a counts --edges normal,no-proc-macro: 157,202, not 246,250. The target moves down with it, 250,000 -> 161,000, so the correction hands back essentially nothing (headroom 3,750 -> 3,798). Three controls: the exclusion drops exactly the five expected crates, only removes and never adds, and is not a no-op. The DFD gate then caught the follow-on it exists for — three historical documents carrying live fact tags for a number that had changed. Not rewritten; untagged, with a supersession banner. AM-4b is deliberately not corrected: its proc-macro share is unmeasured. gate-review now reads 0 due, 0 silent, 0 drifted — GATE-REVIEW earns its first caught entry by forcing SH-3's re-justification, and the registry has no silent gates left. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-02 07:30:14 +02:00
id = "SH-1/SH-2"
name = "session-shape budget"
target = "shape-budget"
CB-WP-0013-T02/T03: retire SH-3 as a gate; correct AM-4a and its target ADR-0008, tier M (survey and ADR merged). D1 — SH-3 retired as a gate, kept as a diagnostic. Investigating it found a third defect, deeper than the two this pass was declared on. Re-deriving batching from the raw transcripts, independently of cb-cost: CB-WP-0011 pass 54 with tools 0 batched 0.0% gap -> next decl 16 with tools 6 batched 37.5% CB-WP-0012 pass 86 with tools 0 batched 0.0% gap -> next decl 10 with tools 1 batched 10.0% CB-WP-0013 so far 10 with tools 0 batched 0.0% Zero batched turns in 150 in-pass responses; 37.5% in one gap, above the 20% floor. Batching needs two calls whose inputs are known at once — orientation work. Implementation consumes each step's result before the next. SH-3's window is since the last commit, which during a pass is always implementation. The metric could not read above ~0% in the window it was gated on. A floor the window structurally excludes is not a target. This pass's own declaration was also wrong: it claimed batching "has got worse" (7.8-8.6% vs 1.1-6.3%). Differently-placed windows, not different behaviour. Withdrawn — the same class of error, in the pass written to correct it. Not retargeting to match the measurement: the floor was not moved to 6%, the gate was removed on an argument about what the quantity is worth. The number is still reported; only the verdict is gone. D2/D3 — AM-4a counts --edges normal,no-proc-macro: 157,202, not 246,250. The target moves down with it, 250,000 -> 161,000, so the correction hands back essentially nothing (headroom 3,750 -> 3,798). Three controls: the exclusion drops exactly the five expected crates, only removes and never adds, and is not a no-op. The DFD gate then caught the follow-on it exists for — three historical documents carrying live fact tags for a number that had changed. Not rewritten; untagged, with a supersession banner. AM-4b is deliberately not corrected: its proc-macro share is unmeasured. gate-review now reads 0 due, 0 silent, 0 drifted — GATE-REVIEW earns its first caught entry by forcing SH-3's re-justification, and the registry has no silent gates left. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-02 07:30:14 +02:00
checks = "mean and p90 context since the last commit (SH-3 retired as a gate, ADR-0008 D1 — still reported as a diagnostic)"
added = "2026-08-01"
review_by = "2026-11-30"
caught = [
"first run fired HARD at 656,574 against a 300,000 ceiling, which is what prompted the compaction before CB-WP-0008",
CB-WP-0013-T02/T03: retire SH-3 as a gate; correct AM-4a and its target ADR-0008, tier M (survey and ADR merged). D1 — SH-3 retired as a gate, kept as a diagnostic. Investigating it found a third defect, deeper than the two this pass was declared on. Re-deriving batching from the raw transcripts, independently of cb-cost: CB-WP-0011 pass 54 with tools 0 batched 0.0% gap -> next decl 16 with tools 6 batched 37.5% CB-WP-0012 pass 86 with tools 0 batched 0.0% gap -> next decl 10 with tools 1 batched 10.0% CB-WP-0013 so far 10 with tools 0 batched 0.0% Zero batched turns in 150 in-pass responses; 37.5% in one gap, above the 20% floor. Batching needs two calls whose inputs are known at once — orientation work. Implementation consumes each step's result before the next. SH-3's window is since the last commit, which during a pass is always implementation. The metric could not read above ~0% in the window it was gated on. A floor the window structurally excludes is not a target. This pass's own declaration was also wrong: it claimed batching "has got worse" (7.8-8.6% vs 1.1-6.3%). Differently-placed windows, not different behaviour. Withdrawn — the same class of error, in the pass written to correct it. Not retargeting to match the measurement: the floor was not moved to 6%, the gate was removed on an argument about what the quantity is worth. The number is still reported; only the verdict is gone. D2/D3 — AM-4a counts --edges normal,no-proc-macro: 157,202, not 246,250. The target moves down with it, 250,000 -> 161,000, so the correction hands back essentially nothing (headroom 3,750 -> 3,798). Three controls: the exclusion drops exactly the five expected crates, only removes and never adds, and is not a no-op. The DFD gate then caught the follow-on it exists for — three historical documents carrying live fact tags for a number that had changed. Not rewritten; untagged, with a supersession banner. AM-4b is deliberately not corrected: its proc-macro share is unmeasured. gate-review now reads 0 due, 0 silent, 0 drifted — GATE-REVIEW earns its first caught entry by forcing SH-3's re-justification, and the registry has no silent gates left. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-02 07:30:14 +02:00
"CB-WP-0013: SH-3 retired from this gate. Its window is *since the last commit*, which during a pass is implementation work, and batching needs two calls whose inputs are known at once — which is orientation work. Measured: 0 batched turns in 150 in-pass responses across three passes, against 37.5% in the gap between two of them. A floor the window structurally excludes is not a target (ADR-0008 D1)",
]
retire_if = "context stops correlating with cost, or the model's context handling makes the number unactionable"
[[gate]]
id = "M-D1-MUT"
name = "mutation coverage of acceptance rows"
target = "mutation-check"
checks = "each acceptance row's assertion must go red for a stated reason when mutated"
added = "2026-07-31"
review_by = "2026-12-31"
caught = [
"AM-6 measuring contention, not throughput",
"AM-5's 61% measurement error under load",
"peak RSS over-reported 3x",
"K10's first round trip not reproducing",
"the bench workload existing twice",
"AM-2's expect matching its own passing output (EXPECT-VACUOUS)",
CB-WP-0015: the two inert clauses, AM-7 scaling and AM-8 N=10 Provenance (tier S, one paragraph in lieu of survey and ADR): the two clauses mutation-check has reported inert since CB-WP-0005. AM-7's scaling ratio was held up by a test literally named replay_100k_events_is_linear_and_fast that computed both throughputs, printed both, and never divided one by the other. AM-8's N=10 was held up by a runner that does two. Both are now red. AM-7 3/3, AM-8 2/2, M-D1-MUT 10/14, and ADR-0005's >=10-of-14 prediction MET for the first time. Neither was closed by amending the question away, which was the live risk: the denominator is unchanged and the four unenforced rows are the four already unenforceable. AM-7 needed three estimators. Best-of-N per leg then divide (AM-6's, correct for a floor on one number) gave 0.581-1.085 on an unchanged binary; legs back-to-back gave medians 0.931-1.004; legs interleaved at fold granularity give 0.987/0.991/0.989, and 0.989 under 8-way CPU contention while absolute throughput fell 4x. The INDETERMINATE guard demanded unanimity and failed a good measurement over one sample 0.001 under the floor; it now requires a two-thirds majority. The control that matters: AM-6's constant-cost mutation halves throughput and leaves this ratio at 0.999x green, so AM-7 is not a second AM-6. AM-8 kept N=10 because the measurement said so. Perturbing the RNG only from its fourth construction on: --runs 2 PASSES, --runs 10 fails. A late-onset divergence is deterministic, not flaky, so it is a control rather than a coin flip. Ten runs live on one scenario (make am8, ~2s) rather than all 25 (47s a build). GameKernel 5b records it. The full run also found AM-4a's own mutation stale since ADR-0008 D3 moved the target 250,000 -> 161,000 in CB-WP-0013 -- reported HARNESS-BROKEN, no score published. The build-free half of that check is now a --self-test assertion, so make all catches the next one. mutation-check clauses may now carry their own verify and mutation, and then the enforced flag is measured rather than declared; a declaration disagreeing with its measurement is refused. make all exits 0. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-02 14:07:08 +02:00
"CB-WP-0015: the two clauses it had reported inert since CB-WP-0005 — AM-7's scaling ratio (a test named `replay_100k_events_is_linear_and_fast` that computed both throughputs, printed both, and never divided one by the other) and AM-8's N=10 (the runner did two). Both now red, 10/14, and the >=10-of-14 prediction met for the first time",
"CB-WP-0015: AM-4a's own mutation, stale since ADR-0008 D3 moved the target 250,000 -> 161,000 in CB-WP-0013. Reported HARNESS-BROKEN and refused to publish a score, which is the harness catching itself. The build-free half of that check is now a --self-test assertion, so it runs in `make all` instead of only on a full run",
]
retire_if = "a full pass adds rows without finding anything, twice running — the harness costs real money per run"
[[gate]]
id = "DFD"
name = "single source of fact"
target = "facts-check"
checks = "every tagged number in the docs matches the tool that measures it"
added = "2026-07-31"
review_by = "2026-12-31"
caught = [
"gr_scenarios stale at 21 after CB-WP-0008 T03 added three scenarios",
]
retire_if = "the untagged-literal count reaches zero and stays there, meaning the docs stopped restating measured numbers"
[[gate]]
id = "AM-1b"
name = "kernel spec->code link"
target = "coverage"
checks = "every numbered K-rule is named in the source; binds 2026-08-31"
added = "2026-07-31"
review_by = "2026-08-31"
caught = [
"3 unlinked K-rules at introduction (15/18); 18/18 today",
]
retire_if = "it stays at 100% through two passes that add kernel rules — at that point it is measuring a habit, not enforcing one"
[[gate]]
id = "META-20"
name = "meta budget"
target = "status"
checks = "share of the trailing 5 passes spent on the loop itself; soft 20%. Purpose (InnerLoop v1.7): most spend on the task at hand, some on control, review and improving the process that carries the work forward"
added = "2026-08-01"
review_by = "2026-11-30"
caught = [
"its own cumulative-window defect, reported in CB-EV-0007 §3 and fixed by CB-WP-0009 T01",
"CB-WP-0019: it had a threshold and no stated purpose, which is why the number was argued three times. The maintainer set 80/20; writing that down exposed that the ratio and the window are a pair — one meta pass among three at parity cost already reads 33%, so 20% over a trailing 3 would have silently also demanded it be half-price. Now 20% over a trailing 5, which is one pass in five at normal cost",
]
retire_if = "product and meta stop being separable, or the share sits under the line for four passes without anyone consulting it"
# The phase setting (CB-WP-0019 T05). Uncomment, argue, and date it to move
# the split for a phase — stage 0 and a stabilisation phase do not deserve
# the same ratio. It REVERTS to META_SOFT_PCT on `review_by` unless
# re-argued, and `status.py --self-test` refuses one with no reason or no
# expiry, because a threshold anyone may move is not a threshold.
#
# meta_phase = { pct = 35, reason = "why this phase differs", review_by = "YYYY-MM-DD" }
[[gate]]
id = "LOOP-LINT"
name = "executable InnerLoop rules"
target = "loop-lint"
checks = "loadability, unmeasured verdicts, tier and chaos declarations, review trails, self-test entry points, and this registry"
added = "2026-07-30"
review_by = "2026-12-31"
caught = [
"four loadability breaches (401, 427, 406, 409 lines), each fixed structurally rather than by raising the limit",
"a reporting tool with no --self-test entry point (tools/repo.py)",
]
retire_if = "two passes run with no finding while artifacts keep growing — that would mean it is measuring the wrong properties"
[[gate]]
id = "CHAOS"
name = "the chaos roll"
target = ""
CB-WP-0018 T03/T04: explanations, and window 1's verdict T03: input::describe writes a sentence per legal command; data-descs carries them in step with data-targets; the ghost already following the pointer shows the one for whatever legal target is under it, so the explanation lands beside the target with no overlay layer to keep aligned. ADR-0010 D1 binds -- the page renders it, never composes it. Both mutations INITIALLY SURVIVED because the fixture's Attack card had exactly one target, where an off-by-one shift and a truncation are both no-ops. CB-EV-0014's lesson one level in: a fixture too thin to express a failure is how the failure survives. Two attack targets now, both red. T04: chaos rate d4 -> d8, window 2 open at 12 declarations, retiring if an override changes nothing twice running. Window 1's condition was NOT met -- both overrides changed the outcome -- so the mechanism is kept. The weakest part of the decision is that it is a rate change argued from n=2, so window 2 carries a falsifier: no override at all is evidence the rate went too far, not that the mechanism is healthy. InnerLoop.md hit 401 lines and the loadability gate fired; the rationale moved to InnerLoopReference.md, structurally, per the standing precedent that limits are not raised. CB-WP-0017 settled at $9.48/40 against $5.19/23 reported mid-flight, 83% higher. Six for six, always low -- read by re-running the instrument at the moment of quoting, which is CB-EV-0015's correction applied for the first time. make all exits 0. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-03 02:24:13 +02:00
checks = "d8 on each tier declaration, 12-declaration calibration window (window 2, opened 2026-08-03; window 1 ran at d4)"
added = "2026-07-30"
CB-WP-0018 T03/T04: explanations, and window 1's verdict T03: input::describe writes a sentence per legal command; data-descs carries them in step with data-targets; the ghost already following the pointer shows the one for whatever legal target is under it, so the explanation lands beside the target with no overlay layer to keep aligned. ADR-0010 D1 binds -- the page renders it, never composes it. Both mutations INITIALLY SURVIVED because the fixture's Attack card had exactly one target, where an off-by-one shift and a truncation are both no-ops. CB-EV-0014's lesson one level in: a fixture too thin to express a failure is how the failure survives. Two attack targets now, both red. T04: chaos rate d4 -> d8, window 2 open at 12 declarations, retiring if an override changes nothing twice running. Window 1's condition was NOT met -- both overrides changed the outcome -- so the mechanism is kept. The weakest part of the decision is that it is a rate change argued from n=2, so window 2 carries a falsifier: no override at all is evidence the rate went too far, not that the mechanism is healthy. InnerLoop.md hit 401 lines and the loadability gate fired; the rationale moved to InnerLoopReference.md, structurally, per the standing precedent that limits are not raised. CB-WP-0017 settled at $9.48/40 against $5.19/23 reported mid-flight, 83% higher. Six for six, always low -- read by re-running the instrument at the moment of quoting, which is CB-EV-0015's correction applied for the first time. make all exits 0. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-03 02:24:13 +02:00
review_by = "2026-11-30"
caught = [
"CB-WP-0011: first fire in 6 declarations — d4=4 rolled stage 1 from structural L to S; the deleted survey would have opened on 2D toolkits while the existing text renderer was showing 24 of 41 view fields (CB-EV-0009 §1)",
CB-WP-0017: legible interaction, and the chaos window's verdict Provenance (tier M, structural S, chaos d4=4 -> OVERRIDE drawn M): the maintainer could drag after CB-WP-0016 but could not tell what was pickable, held, or droppable. Underneath that, the page was WRONG about which moves exist: 9 legal commands rendered as 5 cards each claiming all three target kinds, from a const string in the emitter. Investigate is legal on problems 2 and 3 but not 1; Solve on 1 but not 2 or 3. The live page now says 'Solve onto problem 1'. ADR-0010 restates control 5, which this work would otherwise have outgrown in silence: every game fact the page acts on must arrive from Rust as data; the script may read, match and render it, never compute, infer, filter or default one. The survey's real finding is that the permitted and forbidden designs are indistinguishable from outside, so the vocabulary grep is demoted to a cheap first line and two behavioural properties become the controls -- the highlighted set EQUALS the set Rust emitted, and anything the page marks legal must resolve. Both mutation-proven; the derive-legality mutation produces a plausible highlight (seat-0,1,2 where only seat-1 is legal) and is caught. Visible now: .pick resting shadow, .held on the grabbed element, .dropok on every legal target including BOTH drawings of a seat, and a ghost following the pointer. Nothing perceptual is verified and ADR-0010 D5 says so. The DOM stub now models classList/querySelectorAll/createElement and builds its node set from the real emitted page. Trap recorded: QuickJS fixes its stack limit at Context creation relative to that frame, so a helper returning a Context makes every later eval report 'SyntaxError: stack overflow'. CHAOS WINDOW CLOSED, 12 declarations, 2 overrides, one each way. Both changed the outcome, so the retirement condition is not met. Verdict: keep, and recommend d4 -> d8 with a second window of 12 -- that is a change to the loop's own constraints and is owed to the next declaration as tier-M work, not made here. CB-EV-0014 corrected: it quoted CB-WP-0015 at $15.14/136 and called it the first settled figure quoted. Now $22.70/166. The number had been read during CB-WP-0015 itself, so there are two defects -- the boundary, and quoting from memory instead of re-running the instrument. make all exits 0. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-02 22:42:45 +02:00
"CB-WP-0017: d4=4 — second override in twelve, and the first to roll UP (structural S → M). It bought ADR-0010: the script's widening from 'it does one thing' to holding a drag, following the pointer and marking other elements would otherwise have landed under a tier-S provenance paragraph, silently outgrowing ADR-0007 D5. The ADR's own finding is that the permitted and forbidden designs are indistinguishable from outside, which demoted the vocabulary grep to a cheap first line and produced the two behavioural controls that replaced it",
CB-WP-0012-T05: evidence — tier L deleted its own deliverable CB-EV-0010. The pass's own verdict on the tier it ran at. Full-weight review withdrew the capability port the declaration was made to build. A tier-S pass has no step 2 and would have shipped it, and stage 2 would have found it unimplementable — which is what CommitWindow is already on record in this repo for doing. Two corrections of the survey's own numbers, compounding: AM-4a headroom 3,750 claimed -> 92,798 measured (25x) cheapest windowed 480,501 claimed -> 140,079 measured (3.4x) headline ratio 128x -> 1.5x (85x) The prediction from CB-EV-0009 §4 held: meta budget reads 0%, published in advance and unfalsified. A correction that is now a pattern: CB-EV-0009 reported CB-WP-0011 at 45 responses / $4.23 / 0.094; final is 71 / $7.02 / 0.099. Still the cheapest pass, so the conclusion stands. But that is the second consecutive evidence file to report its own pass's cost low — a pass cannot measure its own cost, and one quoting its own is quoting a floor. First priced tier comparison on a single subject: 0.123 $/response at L against 0.099 at S — 24% more, for a pass that found the two errors above. On one data point, step 2 is cheap. Not shipped, and said plainly: the emitted JavaScript has never been executed. The socket loop is tested end to end with synthetic HTTP and the page is asserted against as a parsed document, but no browser engine has run it. INTENT stage 1 therefore stays open even though all four of its named deliverables now exist. SH-3 reads 0.0% for a sixth consecutive pass and remains the oldest unargued number in the project. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-02 04:30:10 +02:00
"CB-WP-0012: d4=1, no override — and the contrast is the entry. Tier L at full weight deleted its own structural trigger: adversarial review withdrew the capability port the declaration was made to build (ADR-0007 D2), and corrected the survey's headline claim by 85x (128x -> 1.5x, CB-EV-0010 §2). Two passes on one subject at two tiers, priced: 0.123 $/response at L against 0.099 at S (CB-EV-0010 §5)",
]
CB-WP-0018 T03/T04: explanations, and window 1's verdict T03: input::describe writes a sentence per legal command; data-descs carries them in step with data-targets; the ghost already following the pointer shows the one for whatever legal target is under it, so the explanation lands beside the target with no overlay layer to keep aligned. ADR-0010 D1 binds -- the page renders it, never composes it. Both mutations INITIALLY SURVIVED because the fixture's Attack card had exactly one target, where an off-by-one shift and a truncation are both no-ops. CB-EV-0014's lesson one level in: a fixture too thin to express a failure is how the failure survives. Two attack targets now, both red. T04: chaos rate d4 -> d8, window 2 open at 12 declarations, retiring if an override changes nothing twice running. Window 1's condition was NOT met -- both overrides changed the outcome -- so the mechanism is kept. The weakest part of the decision is that it is a rate change argued from n=2, so window 2 carries a falsifier: no override at all is evidence the rate went too far, not that the mechanism is healthy. InnerLoop.md hit 401 lines and the loadability gate fired; the rationale moved to InnerLoopReference.md, structurally, per the standing precedent that limits are not raised. CB-WP-0017 settled at $9.48/40 against $5.19/23 reported mid-flight, 83% higher. Six for six, always low -- read by re-running the instrument at the moment of quoting, which is CB-EV-0015's correction applied for the first time. make all exits 0. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-03 02:24:13 +02:00
retire_if = "an override changes nothing twice running (window 2 condition, CB-WP-0018 T04). Window 1's condition — no override changing the outcome — was NOT met: both did, so the mechanism was kept and the rate dropped d4 → d8 instead"
CB-WP-0017: legible interaction, and the chaos window's verdict Provenance (tier M, structural S, chaos d4=4 -> OVERRIDE drawn M): the maintainer could drag after CB-WP-0016 but could not tell what was pickable, held, or droppable. Underneath that, the page was WRONG about which moves exist: 9 legal commands rendered as 5 cards each claiming all three target kinds, from a const string in the emitter. Investigate is legal on problems 2 and 3 but not 1; Solve on 1 but not 2 or 3. The live page now says 'Solve onto problem 1'. ADR-0010 restates control 5, which this work would otherwise have outgrown in silence: every game fact the page acts on must arrive from Rust as data; the script may read, match and render it, never compute, infer, filter or default one. The survey's real finding is that the permitted and forbidden designs are indistinguishable from outside, so the vocabulary grep is demoted to a cheap first line and two behavioural properties become the controls -- the highlighted set EQUALS the set Rust emitted, and anything the page marks legal must resolve. Both mutation-proven; the derive-legality mutation produces a plausible highlight (seat-0,1,2 where only seat-1 is legal) and is caught. Visible now: .pick resting shadow, .held on the grabbed element, .dropok on every legal target including BOTH drawings of a seat, and a ghost following the pointer. Nothing perceptual is verified and ADR-0010 D5 says so. The DOM stub now models classList/querySelectorAll/createElement and builds its node set from the real emitted page. Trap recorded: QuickJS fixes its stack limit at Context creation relative to that frame, so a helper returning a Context makes every later eval report 'SyntaxError: stack overflow'. CHAOS WINDOW CLOSED, 12 declarations, 2 overrides, one each way. Both changed the outcome, so the retirement condition is not met. Verdict: keep, and recommend d4 -> d8 with a second window of 12 -- that is a change to the loop's own constraints and is owed to the next declaration as tier-M work, not made here. CB-EV-0014 corrected: it quoted CB-WP-0015 at $15.14/136 and called it the first settled figure quoted. Now $22.70/166. The number had been read during CB-WP-0015 itself, so there are two defects -- the boundary, and quoting from memory instead of re-running the instrument. make all exits 0. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-02 22:42:45 +02:00
# VERDICT, CB-EV-0015 §5 (window closed 2026-08-02, 12 declarations, 2 overrides).
# Not retired: both overrides changed the outcome. CB-WP-0011 (L→S) bought a
# defect in the existing renderer that the deleted survey would have walked
# past, at 0.099 $/response against 0.123. CB-WP-0017 (S→M) bought the
# restatement of control 5 — at tier S the script would have outgrown
# ADR-0007 D5 under a one-paragraph commit note.
# RECOMMENDED, and owed to the next declaration as tier-M work: drop the rate
# d4 → d8 and open a second window of 12. Both overrides were informative
# BECAUSE they were rare; a mechanism firing on a quarter of declarations
# stops being a calibration and becomes the tier system.
# New retirement condition proposed: retire if an override changes nothing
# twice running.
[[gate]]
id = "GATE-REVIEW"
name = "this registry"
target = "gate-review"
checks = "gates past their review date, and gates that have caught nothing"
added = "2026-08-01"
review_by = "2026-12-31"
CB-WP-0013-T02/T03: retire SH-3 as a gate; correct AM-4a and its target ADR-0008, tier M (survey and ADR merged). D1 — SH-3 retired as a gate, kept as a diagnostic. Investigating it found a third defect, deeper than the two this pass was declared on. Re-deriving batching from the raw transcripts, independently of cb-cost: CB-WP-0011 pass 54 with tools 0 batched 0.0% gap -> next decl 16 with tools 6 batched 37.5% CB-WP-0012 pass 86 with tools 0 batched 0.0% gap -> next decl 10 with tools 1 batched 10.0% CB-WP-0013 so far 10 with tools 0 batched 0.0% Zero batched turns in 150 in-pass responses; 37.5% in one gap, above the 20% floor. Batching needs two calls whose inputs are known at once — orientation work. Implementation consumes each step's result before the next. SH-3's window is since the last commit, which during a pass is always implementation. The metric could not read above ~0% in the window it was gated on. A floor the window structurally excludes is not a target. This pass's own declaration was also wrong: it claimed batching "has got worse" (7.8-8.6% vs 1.1-6.3%). Differently-placed windows, not different behaviour. Withdrawn — the same class of error, in the pass written to correct it. Not retargeting to match the measurement: the floor was not moved to 6%, the gate was removed on an argument about what the quantity is worth. The number is still reported; only the verdict is gone. D2/D3 — AM-4a counts --edges normal,no-proc-macro: 157,202, not 246,250. The target moves down with it, 250,000 -> 161,000, so the correction hands back essentially nothing (headroom 3,750 -> 3,798). Three controls: the exclusion drops exactly the five expected crates, only removes and never adds, and is not a no-op. The DFD gate then caught the follow-on it exists for — three historical documents carrying live fact tags for a number that had changed. Not rewritten; untagged, with a supersession banner. AM-4b is deliberately not corrected: its proc-macro share is unmeasured. gate-review now reads 0 due, 0 silent, 0 drifted — GATE-REVIEW earns its first caught entry by forcing SH-3's re-justification, and the registry has no silent gates left. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-02 07:30:14 +02:00
caught = [
"CB-WP-0013: forced SH-3's re-justification and then its retirement. `make gate-review` had reported SH-3 as a standing breach for seven passes with zero actions taken, which is this registry's own ritual test (ADR-0006 D4); asking what it had ever caught is what exposed that the metric could not read above ~0% in the window it was gated on (ADR-0008 D1)",
]
retire_if = "it has retired, tightened, or forced the re-justification of nothing by its review date — then it is a ritual, and ADR-0006 D4 says rituals cash out or go"