T11: dated rates and staleness become data, ahead of the 2026-08-31 flip

The price sheet had two defects of one shape -- a schema that could not
hold the fact it needed, the same criticism the cost survey levelled at
the State Hub.

  CA-16  time-boxed rates are DATA. Sonnet's intro price lived in a
         `# intro ...` comment and was invisible to the collector that
         reads the file. Now promo_input/promo_output/promo_until,
         applied per response at its own timestamp.
  CA-17  the 90-day staleness rule was prose in MetricsAndScenarios 1a
         that every M-D2-CST verdict silently inherited. Now `recorded`
         + `max_age_days` in the sheet, and a stale sheet ABORTS.

Both are exercised by make cost-test: the promo rate must apply before
2026-08-31 and lapse after, and a 102-day-old sheet must trip.

Applying CA-16 moved AC-1 from $93.32 to $93.15 -- the $0.17 CB-EV-0002
predicted, now collected rather than noted. That is a legitimate
retarget under T07's distinction: the instrument disproved the target,
and its output is in this commit. The number has now been stated five
times ($248.46, $92.21, $92.87, $93.32, $93.15), each correction from a
different mechanism.

Evidence tables regenerated from the tool rather than hand-patched,
per CA-15 -- which is the rule that exists because hand-typed tables
were the only thing the adversarial review found wrong.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
tegwick 2026-07-31 09:23:29 +02:00
parent 06628e83e1
commit db731aa0bc
6 changed files with 138 additions and 32 deletions

View file

@ -17,9 +17,9 @@ transcribed by hand (CA-15).
| ID | Metric | Target | Measured | Verdict |
|---|---|---|---|---|
| AC-1 | pinned total, as two components | $93.32 = $92.21 + $1.11 | **$92.21 main + $1.11 subagent = $93.32** | **met** |
| AC-1 | pinned total, as two components | $93.15 = $92.03 + $1.11 | **$92.03 main + $1.11 subagent = $93.15** | **met** |
| AC-2 | reconciliation residual | $0.00 | **$0.000000** | **met** |
| AC-3 | unattributed share reported | present, 33% | **32.5%, own line** | **met** |
| AC-3 | unattributed share reported | present, 33% | **32.4%, own line** | **met** |
| AC-4 | composition reported | 5 components | **5 of 5** | **met** |
| AC-5 | dedup violation aborts | non-zero exit | **abort raised** | **met** |
| AC-6 | zero responses refuses to report | non-zero exit | **0 rows, no number emitted** | **met** |
@ -27,7 +27,7 @@ transcribed by hand (CA-15).
| AC-8 | 5m cache priced at 1.25× | $1.25/100k @ fable | **$1.2500 (1h would be $2.0000)** | **met** |
| AC-9 | streamed partial output → final | 5,5,195 → 195 | **195** | **met** |
No `unmeasured` rows. AC-1's target was corrected three times before this
No `unmeasured` rows. AC-1's target was corrected four times before and after this
run; §4 records why, because the sequence is more useful than the final
number.
@ -35,23 +35,23 @@ number.
```text
input 690 tok $ 0.00 0.0%
output 323,643 tok $ 11.15 11.9%
cache_read 80,611,798 tok $ 59.75 64.0%
output 323,643 tok $ 11.13 11.9%
cache_read 80,611,798 tok $ 59.67 64.1%
write_5m 37,467 tok $ 0.47 0.5%
write_1h 1,672,854 tok $ 21.95 23.5%
TOTAL $ 93.32
write_1h 1,672,854 tok $ 21.87 23.5%
TOTAL $ 93.15
```
**88.0% of spend is cache; 11.9% is output.** The ratio of context re-read
to text written is 249:1. A single total would have shown $93.32 and
to text written is 249:1. A single total would have shown $93.15 and
concealed all of it — which is exactly what CB-WP-0001's `M-D2-TOK` would
have done, and why that metric is now demoted.
## 3. Per-task attribution
```text
UNATTRIBUTED $ 30.32 32.5%
T08 $ 21.02 22.5%
UNATTRIBUTED $ 30.14 32.4%
T08 $ 21.02 22.6%
T07 $ 12.01 12.9%
T03 $ 9.91 10.6%
T04 $ 8.00 8.6%
@ -60,7 +60,7 @@ have done, and why that metric is now demoted.
T09 $ 1.29 1.4%
```
**Stated limit (CA-10):** 32.5% of cost sits in commits whose subject
**Stated limit (CA-10):** 32.4% of cost sits in commits whose subject
carries no `T##` tag, so this table is a view over 67.5% of spend. That is
a property of commit hygiene, not of the collector.
@ -75,14 +75,15 @@ is worth watching, not yet a conclusion from n=1.
- **The per-task figures are not comparable across passes.** They mix
models (opus/fable/sonnet) at different price points and different cache
states. The dollar figure is comparable; a token count is not.
- **AC-3's 32.5% is a fixture pin, not a quality target.** Improving commit
- **AC-3's 32.4% is a fixture pin, not a quality target.** Improving commit
tagging will move it, and that is the desired direction.
- **This is one session.** Every ratio here (cache share, $/turn, the
compaction effect in §5) is n=1 and should be treated as a hypothesis
until a second pass reproduces it.
- **The price sheet cannot express a time-boxed rate.** Sonnet's intro price
lives in a TOML comment, so sonnet-priced work is off by $0.17 here
(0.19%). This becomes a real error on 2026-08-31.
- ~~**The price sheet cannot express a time-boxed rate.**~~ **Fixed
2026-07-31 by CB-WP-0003 T11** (CA-16/CA-17). Applying the sonnet
promotional rate moved AC-1 from $93.32 to **$93.15** — the $0.17 this
section predicted, now collected rather than merely noted.
- **AC-1 is not invoice-verified.** No admin key exists, so the Anthropic
billing API could not independently confirm the total. The transcript
counters are the same ones billing uses, but that is an argument, not a
@ -145,7 +146,7 @@ largest sample was assumed to hold on the smallest one.** The main
transcript is 338 of 346 responses, so 206/206 felt conclusive; the
violation lives entirely in the 8 responses nobody checked separately.
Cost of the adversarial review this pass: **$1.11**, against a $93.32 pass.
Cost of the adversarial review this pass: **$1.11**, against a $93.15 pass.
It found three approval-blocking defects, one of which (the subagent
exclusion) would have made this evidence file certify a broken collector.
Second consecutive pass where a ~1% spend on review changed the outcome.