CB-WP-0013-T04: evidence — and a rule about quoting your own cost
Some checks failed
ci / check (push) Failing after 4s
Some checks failed
ci / check (push) Failing after 4s
CB-EV-0011. Three instruments were wrong; the pass found the third defect itself and committed the fourth. The rule this pass should leave behind: an evidence file quotes the PREVIOUS pass's final cost, not its own. Three for three, all low by a strikingly consistent ~2/3: CB-WP-0009 $6.73 reported -> $11.31 final CB-WP-0011 $4.23 reported -> $7.02 final CB-WP-0012 $8.82 reported -> $14.52 final And this time it changed a conclusion. CB-EV-0010 reported the first priced tier comparison as "0.123 at L against 0.099 at S — 24% more". The final figures are 0.150 against 0.099 — 51% more. Tier L costs twice what that comparison claimed. It still deleted a port that would have been rewritten and found errors of 25x and 85x, so it still looks worth it, but the price doubled and the file that made the argument could not have known. Also closed rather than deferred again: cb-play's "third mode" never arrived. --serve is a play flag selecting which Policy fills a human seat; the binary still has two modes, and CB-EV-0010 §6 was wrong. Carried for two passes, resolved, stopped. Recorded as a near-miss: an instrument reading that looked impossible (53 tool calls in 53 responses) turned out to be correct. One independent re-derivation before writing anything stopped a false accusation — in the pass whose whole subject is instruments reporting things that are not so. gate-review: 0 due, 0 silent, 0 drifted. Every gate in the registry has a caught entry for the first time. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
parent
d2c3f340f9
commit
768eee7d27
3 changed files with 205 additions and 2 deletions
|
|
@ -95,3 +95,7 @@
|
|||
| task | CB-WP-0012-T03 | done | — | workplans/CB-WP-0012-render-port.md |
|
||||
| task | CB-WP-0012-T04 | done | — | workplans/CB-WP-0012-render-port.md |
|
||||
| task | CB-WP-0012-T05 | done | — | workplans/CB-WP-0012-render-port.md |
|
||||
| task | CB-WP-0013-T01 | done | — | workplans/CB-WP-0013-instrument-corrections.md |
|
||||
| task | CB-WP-0013-T02 | done | — | workplans/CB-WP-0013-instrument-corrections.md |
|
||||
| task | CB-WP-0013-T03 | done | — | workplans/CB-WP-0013-instrument-corrections.md |
|
||||
| task | CB-WP-0013-T04 | done | — | workplans/CB-WP-0013-instrument-corrections.md |
|
||||
|
|
|
|||
181
evidence/CB-EV-0011-instrument-corrections.md
Normal file
181
evidence/CB-EV-0011-instrument-corrections.md
Normal file
|
|
@ -0,0 +1,181 @@
|
|||
# CB-EV-0011 — three instruments, and a rule about quoting yourself
|
||||
|
||||
CB-WP-0013 T04. Measured 2026-08-02 at `d2c3f34`. Pass kind `meta`,
|
||||
tier **M** (chaos d4=3, no override). Declaration 8 of 12.
|
||||
|
||||
The pass was declared to fix two instruments. It found a third defect in
|
||||
one of them, and a fourth in itself.
|
||||
|
||||
---
|
||||
|
||||
## 1. What was actually wrong with SH-3
|
||||
|
||||
The declaration named one defect. There were three, each deeper.
|
||||
|
||||
**1. The sample.** `--shape-budget` windows since the last commit and is
|
||||
read while writing the evidence file — right after a commit, when the
|
||||
window holds one or two responses. SH-1/SH-2 are location statistics and
|
||||
survive that; SH-3 is a rate, and at n=2 its only values are 0%, 50%,
|
||||
100%. Six evidence files quoted the resulting `0.0%` as a measurement.
|
||||
|
||||
**2. The claimed decline — which was this pass's own error.** The
|
||||
declaration stated batching *"has got worse"*, citing 7.8–8.6% pinned
|
||||
history against 1.1–6.3% per-pass windows. Those are differently-*placed*
|
||||
windows, not different behaviour. **Withdrawn.** The pass declared to
|
||||
correct a trusted-arithmetic error opened by committing one.
|
||||
|
||||
**3. The one that decided it.** Re-deriving batching from the raw
|
||||
transcripts, independently of `cb-cost`:
|
||||
|
||||
| window | with tools | batched | rate |
|
||||
|---|---:|---:|---:|
|
||||
| CB-WP-0011 pass | 54 | 0 | **0.0%** |
|
||||
| gap → next declaration | 16 | 6 | **37.5%** |
|
||||
| CB-WP-0012 pass | 86 | 0 | **0.0%** |
|
||||
| gap → next declaration | 10 | 1 | **10.0%** |
|
||||
| CB-WP-0013 so far | 10 | 0 | **0.0%** |
|
||||
|
||||
**Zero batched turns in 150 in-pass responses. 37.5% in one gap** — above
|
||||
the floor the gate had been failing for seven passes.
|
||||
|
||||
Batching needs two tool calls whose inputs are known at once. That is
|
||||
orientation: read four files, run three probes. Implementation has almost
|
||||
none, because each step consumes the previous step's result. And SH-3's
|
||||
window is *since the last commit*, which during a pass is implementation.
|
||||
**The metric could not read above ~0% in the window it was gated on.**
|
||||
|
||||
A gate meetable only by not working is not a target. SH-3 is retired as a
|
||||
gate (ADR-0008 D1) and kept as a diagnostic over the whole session, where
|
||||
it reads 5.8%.
|
||||
|
||||
**On the appearance of burying it:** retiring a metric that shows a
|
||||
regression looks bad, and should. The answer is defect 2 — the regression
|
||||
was an artifact and has been withdrawn, so there is no established decline
|
||||
to bury. The number is still printed; only the verdict is gone. And the
|
||||
floor was *not* moved to 6%, which is the thing `specs/SessionShape.md`
|
||||
actually forbids.
|
||||
|
||||
## 2. A near-miss worth recording: I almost reported a defect that was not there
|
||||
|
||||
Midway through, a bounded window reported `53 tool calls in 53 responses;
|
||||
0 in batched turns` — exactly one call per response, which looked like the
|
||||
counter had collapsed. I had batched repeatedly in this session, so the
|
||||
reading looked impossible.
|
||||
|
||||
It was correct. Re-deriving from the transcripts independently gave the
|
||||
same 0, and the 131 batched responses that do exist all fall *outside*
|
||||
those windows. The instrument was right and my expectation was wrong.
|
||||
|
||||
Recorded because the **false-accusation class** is already on this
|
||||
project's record, and because the thing that stopped it was cheap: one
|
||||
independent re-derivation before writing anything down. The cost of
|
||||
checking was two minutes; the cost of not checking would have been an
|
||||
evidence file alleging a defect in a tool that did not have one — and this
|
||||
pass's whole subject is instruments reporting things that are not so.
|
||||
|
||||
## 3. AM-4a, corrected, gives back nothing
|
||||
|
||||
```
|
||||
before 23 crates 246,250 lines [ok target 250,000] headroom 3,750
|
||||
after 18 crates 157,202 lines [ok target 161,000] headroom 3,798
|
||||
```
|
||||
|
||||
The instrument dropped 89,048 lines of proc-macro code that never reaches
|
||||
a binary; **the target moved down with it** (ADR-0008 D3). Leaving the
|
||||
target at 250,000 would have handed this project 89,048 lines of headroom
|
||||
it did not earn, in the same change that revealed the error.
|
||||
|
||||
Three controls, all green: the exclusion drops **exactly** the five
|
||||
expected crates, only ever removes and never adds, and is not a no-op.
|
||||
|
||||
**The DFD gate then caught the follow-on**, which is exactly what it is
|
||||
for: three historical documents carried live `fact:am4a_loc` tags for a
|
||||
number that had changed. They are history and were not rewritten — the
|
||||
tags came off and both carry a supersession banner. A gate catching the
|
||||
second-order consequence of a correction, unprompted, is the best evidence
|
||||
yet that the registry is doing something.
|
||||
|
||||
**AM-4b is deliberately not corrected.** Its proc-macro share is
|
||||
unmeasured, and correcting a second instrument on the strength of the
|
||||
first one's ratio is the error this pass exists to fix. Owed.
|
||||
|
||||
## 4. The rule this pass should leave behind: stop quoting your own cost
|
||||
|
||||
| pass | reported mid-pass | final | understated by |
|
||||
|---|---|---|---|
|
||||
| CB-WP-0009 | $6.73 (CB-EV-0008) | $11.31 | 68% |
|
||||
| CB-WP-0011 | $4.23 / 0.094 (CB-EV-0009) | $7.02 / 0.099 | 66% |
|
||||
| CB-WP-0012 | $8.82 / 0.123 (CB-EV-0010) | **$14.52 / 0.150** | 65% |
|
||||
|
||||
Three for three, always low, by a strikingly consistent ~2/3. And this
|
||||
time it **changed a conclusion**, not just a figure:
|
||||
|
||||
> CB-EV-0010 §5 reported the first priced tier comparison as *"0.123
|
||||
> $/response at L against 0.099 at S — 24% more."* The final figures are
|
||||
> **0.150 against 0.099 — 51% more.** The direction holds; the magnitude
|
||||
> was understated by half.
|
||||
|
||||
Tier L is twice as expensive as that comparison claimed. It still found
|
||||
errors of 25× and 85× in CB-WP-0012 and deleted a port that would have
|
||||
been rewritten, so it still looks worth it — but the price of the argument
|
||||
doubled and the evidence file that made it could not have known.
|
||||
|
||||
**Proposed rule, for the next pass to adopt or reject:** an evidence file
|
||||
quotes the *previous* pass's final cost and does not quote its own. Its
|
||||
own is knowable only after the commit that closes it, so any figure it
|
||||
prints is a floor, and this project has now printed three of them and let
|
||||
one carry a conclusion.
|
||||
|
||||
## 5. `cb-play`'s third mode never arrived
|
||||
|
||||
CB-EV-0009 §5 said a third mode would be the second use at which
|
||||
`cb-play`'s single-binary shape should be reconsidered. CB-EV-0010 §6
|
||||
recorded that it had arrived, with `--serve`, and that the pass had not
|
||||
acted. **CB-EV-0010 was wrong.**
|
||||
|
||||
`--serve` is registered in `parse_args` as a *play flag*: it selects which
|
||||
`Policy` fills a human seat, terminal or browser. The binary still has
|
||||
exactly two modes — play a game, or read a recording — and they still take
|
||||
disjoint flags. The trigger did not fire.
|
||||
|
||||
Carried for two passes, resolved here, and **stopped**: there is no split
|
||||
to do, and the standing item is closed rather than deferred again.
|
||||
|
||||
## 6. Cost and shape
|
||||
|
||||
| pass | kind | responses | cost | $/response |
|
||||
|---|---|---|---|---|
|
||||
| CB-WP-0011 | product | 71 | $7.02 | 0.099 |
|
||||
| CB-WP-0012 | product | 97 | $14.52 | 0.150 |
|
||||
| CB-WP-0013 | meta | 34+ | $5.78+ | provisional, see §4 |
|
||||
|
||||
**Meta budget: 21%**, inside the soft 25%, without needing ADR-0006 D2's
|
||||
repair exemption — though both halves would have qualified for it.
|
||||
|
||||
**SH-1 290,475 `[SOFT]`**, drifting toward the hard line; the remedy is a
|
||||
compaction and it is free. **SH-3 now reads `insufficient sample (1
|
||||
response with tool calls, need 14)`** — the fix working in production on
|
||||
its first real reading, where the old instrument would have printed a
|
||||
confident `0.0%`.
|
||||
|
||||
**`gate-review`: 0 due, 0 silent, 0 drifted.** Every gate in the registry
|
||||
now has a `caught` entry, for the first time. GATE-REVIEW earned its own
|
||||
by forcing SH-3's re-justification — asking what a gate had ever caught is
|
||||
what exposed that it could not read what it gated on.
|
||||
|
||||
## 7. Open
|
||||
|
||||
- **AM-4b's proc-macro share is unmeasured.** §3.
|
||||
- **The self-quoting rule in §4 is proposed, not adopted.**
|
||||
- **AM-4a still cannot survive INTENT stage 2**, and this pass made the
|
||||
ratio worse: `wgpu` + `winit` is 1,741,979 against a target now 161,000
|
||||
— **10.8×**, up from 7×. Reserved for the maintainer, twice over now.
|
||||
- **ADR-0007 Decision 3's acquisition rule** remains an argument awaiting
|
||||
ratification.
|
||||
- **INTENT stage 1 stays open**; the emitted JavaScript has still never
|
||||
been executed.
|
||||
- **Tier M's merged survey-and-ADR was the right weight.** One document,
|
||||
no separate review, and the pass still found a defect deeper than the
|
||||
one it was declared on — because the finding came from *measuring*, not
|
||||
from an extra review round. That is one data point for the tier table
|
||||
and against the assumption that depth comes from process weight.
|
||||
|
|
@ -2,7 +2,7 @@
|
|||
id: CB-WP-0013
|
||||
kind: meta
|
||||
title: "Two instruments that were wrong, and the record they wrote"
|
||||
status: todo
|
||||
status: done
|
||||
---
|
||||
|
||||
# Purpose
|
||||
|
|
@ -235,7 +235,7 @@ first one's ratio is the error this task exists to fix. Filed as owed.
|
|||
|
||||
```task
|
||||
id: CB-WP-0013-T04
|
||||
status: todo
|
||||
status: done
|
||||
priority: medium
|
||||
```
|
||||
|
||||
|
|
@ -251,3 +251,21 @@ Also carry the one that is now overdue: **`cb-play` is three modes in one
|
|||
binary.** CB-EV-0009 §5 named the third mode as the trigger to reconsider
|
||||
the shape, CB-EV-0010 §6 recorded that it had arrived and that the pass
|
||||
did not act. Either do it or state why not, and stop carrying it.
|
||||
|
||||
**Done 2026-08-02.**
|
||||
[CB-EV-0011](../evidence/CB-EV-0011-instrument-corrections.md).
|
||||
|
||||
- **The self-quoting rule is proposed.** Three passes have now reported
|
||||
their own cost mid-flight and all three were low by ~2/3 — and this
|
||||
time it changed a conclusion: CB-EV-0010's tier comparison said L cost
|
||||
*24% more* than S; the final figures say **51%**.
|
||||
- **`cb-play`'s third mode never arrived.** `--serve` is a *play flag*
|
||||
selecting which `Policy` fills a human seat. The binary still has two
|
||||
modes. CB-EV-0010 §6 was wrong; the item is **closed, not deferred**.
|
||||
- **A near-miss recorded**: an instrument reading that looked impossible
|
||||
turned out to be correct, and one independent re-derivation before
|
||||
writing anything stopped a false accusation.
|
||||
- **Tier M was the right weight** — one merged document, no separate
|
||||
review, and the pass still found a defect deeper than the one it was
|
||||
declared on, because the finding came from measuring rather than from
|
||||
an extra review round.
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue