Declare CB-WP-0013: two instruments that were wrong, tier M
Declaration 8 of 12. Structural tier M (changes how the loop constrains its own operation — a budget instrument and a gate floor); d4 = 3, no override. Meta budget reads 0% of the trailing three, so this is in budget without the ADR-0006 D2 repair exemption. The SH-3 half is larger than CB-WP-0012 filed it as. "SH-3 has read 0.0% for six consecutive passes" is an artifact: --shape-budget measures the window since the last commit, and it is read while writing the evidence file, immediately after a commit, when that window holds one or two responses. SH-1 and SH-2 are location statistics and survive n=2; SH-3 is a rate, and at n=2 its only possible values are 0%, 50%, 100%. Measured over real per-pass windows it is 1.1%-6.3%, not 0.0%. Which exposes what the artifact hid: the pinned historical rate is 7.8%-8.6%, so batching has got WORSE, and a metric stuck at a constant could not have shown it. That is the ninth recorded instance of the trusted-arithmetic class, and the second one in this same metric — SessionShape.md §4 already records "0 of 330 tool calls were batched" as the eighth. Also owed: AM-4a counts proc-macro crates that never reach a binary (89,048 lines, 36.2%). The task is explicit that the target must move down with the instrument, or the pass must argue why 89,048 lines of unearned headroom are legitimate. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
parent
e482ab9f8f
commit
7105027173
1 changed files with 171 additions and 0 deletions
171
workplans/CB-WP-0013-instrument-corrections.md
Normal file
171
workplans/CB-WP-0013-instrument-corrections.md
Normal file
|
|
@ -0,0 +1,171 @@
|
||||||
|
---
|
||||||
|
id: CB-WP-0013
|
||||||
|
kind: meta
|
||||||
|
title: "Two instruments that were wrong, and the record they wrote"
|
||||||
|
status: todo
|
||||||
|
---
|
||||||
|
|
||||||
|
# Purpose
|
||||||
|
|
||||||
|
```
|
||||||
|
structural tier M (changes how the loop constrains its own operation:
|
||||||
|
a budget instrument and a gate floor — InnerLoop
|
||||||
|
v1.6, ADR-0006 D5)
|
||||||
|
chaos d4 = 3 → no override
|
||||||
|
declared tier M
|
||||||
|
```
|
||||||
|
|
||||||
|
Declaration 8 of 12 in the calibration window. Tier M means the survey and
|
||||||
|
the ADR merge into one document and the adversarial review is optional.
|
||||||
|
|
||||||
|
**Meta budget: 0% of the trailing three**, so this pass is in budget
|
||||||
|
without the ADR-0006 D2 repair exemption — though it would qualify for it,
|
||||||
|
since both halves repair an instrument that reports a breach.
|
||||||
|
|
||||||
|
## Why now
|
||||||
|
|
||||||
|
CB-WP-0012 ended owing two corrections, and one of them turns out to be
|
||||||
|
larger than it was filed as.
|
||||||
|
|
||||||
|
### SH-3 has never read what six evidence files say it read
|
||||||
|
|
||||||
|
CB-EV-0010 §5 states *"SH-3 has now read 0.0% for six consecutive passes"*
|
||||||
|
and calls it the oldest unargued number in the project. The number is an
|
||||||
|
artifact.
|
||||||
|
|
||||||
|
`--shape-budget` measures **the window since the last commit**, and it is
|
||||||
|
read while writing the evidence file — immediately *after* a commit, when
|
||||||
|
that window holds one or two responses. SH-1 and SH-2 are location
|
||||||
|
statistics and survive a sample of two. **SH-3 is a rate**, and at n=2 its
|
||||||
|
only possible values are 0%, 50% and 100%. One window, three metrics, and
|
||||||
|
it is only wrong for one of them.
|
||||||
|
|
||||||
|
Measured over real per-pass windows instead:
|
||||||
|
|
||||||
|
| window since | responses | SH-3 |
|
||||||
|
|---|---:|---:|
|
||||||
|
| `0d2ab22` | 115 | 6.2% |
|
||||||
|
| `6fb0aea` | 113 | 6.3% |
|
||||||
|
| `331e7e9` | 95 | 1.1% |
|
||||||
|
| `c5fa610` | 74 | 1.4% |
|
||||||
|
| `883b608` | 70 | 1.4% |
|
||||||
|
| `84d6886` | 24 | 4.3% |
|
||||||
|
|
||||||
|
So the substantive claim survives — SH-3 is badly unmet — but **the figure
|
||||||
|
quoted in six evidence files was never the measurement**. Worse, the real
|
||||||
|
numbers say something the artifact hid: the historical rate pinned in
|
||||||
|
`specs/SessionShape.md` is **7.8–8.6%**, and these windows are **1.1–6.3%**.
|
||||||
|
Batching has got *worse*, and a metric stuck at a constant 0.0% could not
|
||||||
|
show a trend at all.
|
||||||
|
|
||||||
|
### AM-4a overstates every figure it has ever reported by 36%
|
||||||
|
|
||||||
|
ADR-0007 Decision 4, filed and unimplemented. AM-4a counts
|
||||||
|
`--edges normal`, which includes proc-macro crates that run in the compiler
|
||||||
|
and never reach a binary: **89,048 lines, 36.2%** of the shipped-runtime
|
||||||
|
figure, `syn` alone 66,916.
|
||||||
|
|
||||||
|
## Task: fix SH-3's instrument, and correct the record
|
||||||
|
|
||||||
|
```task
|
||||||
|
id: CB-WP-0013-T01
|
||||||
|
status: todo
|
||||||
|
priority: high
|
||||||
|
```
|
||||||
|
|
||||||
|
A rate needs a sample. `--shape-budget` must **refuse to report SH-3**
|
||||||
|
below a stated minimum rather than print a number the window cannot
|
||||||
|
support — `insufficient sample (n responses)`, not `0.0%`.
|
||||||
|
|
||||||
|
**The choice to argue, not assume:** whether SH-3 should instead be
|
||||||
|
windowed differently from SH-1/SH-2 — over the pass rather than since the
|
||||||
|
last commit — since it is the only one of the three that a two-response
|
||||||
|
window destroys. Pick one and say why the other was rejected.
|
||||||
|
|
||||||
|
**Controls.** A minimum-sample rule is exactly the kind of guard that can
|
||||||
|
be satisfied vacuously: assert that a window *above* the floor still
|
||||||
|
reports a rate, that one below reports the refusal, and that the refusal
|
||||||
|
is distinguishable from a genuine 0.0%. Mutate each.
|
||||||
|
|
||||||
|
**Correct the record.** Six evidence files quote 0.0%. They are history and
|
||||||
|
are not rewritten; the correction is recorded once, where a reader will
|
||||||
|
find it, and named as what it is — the same class as CB-WP-0002's
|
||||||
|
"0 of 330 tool calls were batched", which `specs/SessionShape.md` §4
|
||||||
|
already records as the eighth trusted-arithmetic instance. **This is the
|
||||||
|
ninth, and it is in the same metric.** That coincidence is worth a
|
||||||
|
sentence: SH-3 is the number this project keeps getting wrong.
|
||||||
|
|
||||||
|
## Task: argue SH-3's floor, or move it
|
||||||
|
|
||||||
|
```task
|
||||||
|
id: CB-WP-0013-T02
|
||||||
|
status: todo
|
||||||
|
priority: high
|
||||||
|
```
|
||||||
|
|
||||||
|
The declaration CB-EV-0010 said was owed. `specs/SessionShape.md` is
|
||||||
|
explicit that the 20% floor was set above the measured rate deliberately
|
||||||
|
and that *"whether 20% is reachable or arbitrary is unknown until a pass
|
||||||
|
tries"*. Seven passes have now reported it and none has tried.
|
||||||
|
|
||||||
|
Decide, with the real numbers rather than the artifact:
|
||||||
|
|
||||||
|
- is 20% reachable, and at what cost to how the work is actually done;
|
||||||
|
- is the measured decline from 7.8–8.6% to 1.1–6.3% a behavioural change
|
||||||
|
worth reversing, or an artifact of tier-L passes doing more sequential
|
||||||
|
reading than tier-S ones;
|
||||||
|
- **and the honest option: retire SH-3.** `specs/SessionShape.md` §4 puts
|
||||||
|
the ceiling at **$2–4 on a $93 pass** and says *"batch because it is
|
||||||
|
free, not because it is the lever."* A floor that is 3–20× the measured
|
||||||
|
rate, cannot be moved by the party measured, and is worth single-digit
|
||||||
|
dollars is a candidate for `retire_if`, not a standing breach.
|
||||||
|
|
||||||
|
**Do not retarget to match the measurement.** `specs/SessionShape.md`
|
||||||
|
already forbids exactly that, and CB-WP-0003 T07 exists to prevent it. If
|
||||||
|
the floor moves, it moves on an argument about what batching is worth, and
|
||||||
|
the argument is written down before the number changes.
|
||||||
|
|
||||||
|
## Task: correct AM-4a to count what ships
|
||||||
|
|
||||||
|
```task
|
||||||
|
id: CB-WP-0013-T03
|
||||||
|
status: todo
|
||||||
|
priority: high
|
||||||
|
```
|
||||||
|
|
||||||
|
ADR-0007 Decision 4. `tools/dep-weight.py` counts
|
||||||
|
`--edges normal,no-proc-macro`.
|
||||||
|
|
||||||
|
**The targets do not move with the instrument.** Correcting a metric that
|
||||||
|
reports 246,250 down to 157,202 while leaving the target at 250,000 hands
|
||||||
|
this project 89,048 lines of headroom it did not earn, in the same commit
|
||||||
|
that reveals the error. Either the target moves down by the same amount
|
||||||
|
the instrument does, or the pass states plainly why the extra headroom is
|
||||||
|
legitimate. It is not obviously legitimate; argue it or give it back.
|
||||||
|
|
||||||
|
**Positive control.** `dep-weight.py` already has the guard that an
|
||||||
|
unlocatable crate must not silently count as zero. The new flag must not
|
||||||
|
be able to drop crates it should have counted: assert that the proc-macro
|
||||||
|
set is exactly the five expected crates and that removing the flag
|
||||||
|
restores the old figure.
|
||||||
|
|
||||||
|
## Task: evidence
|
||||||
|
|
||||||
|
```task
|
||||||
|
id: CB-WP-0013-T04
|
||||||
|
status: todo
|
||||||
|
priority: medium
|
||||||
|
```
|
||||||
|
|
||||||
|
`evidence/CB-EV-0011-instrument-corrections.md`.
|
||||||
|
|
||||||
|
Carry: whether tier M's merged survey-and-ADR was the right weight for two
|
||||||
|
instrument corrections; the cost of a meta pass immediately after two
|
||||||
|
product passes; and the standing question of whether **three** consecutive
|
||||||
|
evidence files reporting their own pass's cost low means the loop should
|
||||||
|
simply stop quoting it.
|
||||||
|
|
||||||
|
Also carry the one that is now overdue: **`cb-play` is three modes in one
|
||||||
|
binary.** CB-EV-0009 §5 named the third mode as the trigger to reconsider
|
||||||
|
the shape, CB-EV-0010 §6 recorded that it had arrived and that the pass
|
||||||
|
did not act. Either do it or state why not, and stop carrying it.
|
||||||
Loading…
Add table
Add a link
Reference in a new issue