T04: tools/cb-cost.py — and its positive control fires on first contact
Collector per ADR-0003: enumerates every transcript including the subagents/ tree, dedups by requestId, prices per model and per cache TTL, attributes on (prev_commit, this_commit] intervals, and reconciles to the cent or aborts. The positive control caught a real defect on its very first run against real data, which is the entire argument for writing it: CA-02 assumed usage is identical across the lines of one requestId. True in the main transcript (206/206 groups, verified twice — by the survey and by the adversarial reviewer). FALSE in the subagents/ tree, where output_tokens is a running count: one response reads 5, 5, 195 across its three lines. First-wins scored it at 5. So CA-02 now splits: input-side counters are charged once and must be identical (assertion retained); output_tokens resolves to the max (CA-02a). AC-9 pins the exact 5,5,195 case as a regression test. The acceptance target moved again as a result, for the third time: $248.46 -> $92.21 -> $92.87 -> $93.32. AC-1 was computed by the same first-wins method the tool just disproved, so the tool failing its target was the tool being correct. Target updated, not the tool. Measured at the pin: $92.21 main + $1.11 subagent = $93.32, residual $0.000000. 32.5% UNATTRIBUTED, reported as its own line per CA-10. make cost / cost-test / cost-pin wired; cost-test added to `make all` and to CI, where it gates the collector's assertions without needing transcripts present. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
parent
900467017f
commit
e008b1e787
5 changed files with 514 additions and 14 deletions
|
|
@ -33,6 +33,11 @@ jobs:
|
|||
# AM-4a/AM-4b: third-party source under audit, per configuration.
|
||||
- run: make dep-weight
|
||||
|
||||
# Positive control for the cost collector (AC-5..AC-9). Does not
|
||||
# gate on a dollar figure — transcripts are not present in CI — but
|
||||
# proves the collector still detects the failures it claims to.
|
||||
- run: make cost-test
|
||||
|
||||
# InnerLoop v1.0 positive control, enforced rather than asserted in
|
||||
# prose: --test runs every benchmark once, so a workload that
|
||||
# stalls or produces the wrong event count fails the build instead
|
||||
|
|
|
|||
15
Makefile
15
Makefile
|
|
@ -3,7 +3,7 @@
|
|||
|
||||
CARGO := cargo
|
||||
|
||||
.PHONY: check test sim bench bench-test coverage dep-weight loc all
|
||||
.PHONY: check test sim bench bench-test coverage dep-weight cost cost-test cost-pin loc all
|
||||
|
||||
## fmt + clippy (deny warnings) + HashMap deny-lint
|
||||
check:
|
||||
|
|
@ -21,6 +21,17 @@ dep-weight:
|
|||
coverage:
|
||||
python3 tools/rule-coverage.py
|
||||
|
||||
# M-D2-CST (specs/CostAccounting.md). cost-test is the positive control and
|
||||
# runs first: a cost number from an unverified collector is void.
|
||||
cost: cost-test
|
||||
python3 tools/cb-cost.py --composition --by-task
|
||||
|
||||
cost-test:
|
||||
python3 tools/cb-cost.py --self-test
|
||||
|
||||
cost-pin: cost-test
|
||||
python3 tools/cb-cost.py --pin fc76445 --composition --by-task
|
||||
|
||||
sim:
|
||||
$(CARGO) run -q -p cb-sim -- scenarios/ground/*.yaml
|
||||
|
||||
|
|
@ -40,4 +51,4 @@ loc:
|
|||
printf '%-28s %s\n' $$d "$$(find $$d/src -name '*.rs' | xargs cat | grep -vcE '^\s*(//|$$)')"; \
|
||||
done
|
||||
|
||||
all: check test sim coverage dep-weight bench-test
|
||||
all: check test sim coverage dep-weight cost-test bench-test
|
||||
|
|
|
|||
|
|
@ -22,13 +22,26 @@ block (`thinking`, `text`, `tool_use`), and every line repeats the same
|
|||
> **CA-01.** Cost is computed over responses deduplicated by `requestId`.
|
||||
> Summing per line is a defect; it inflates by ≈1.9× on measured data.
|
||||
|
||||
> **CA-02.** Dedup is asserted, not assumed. Within a `requestId` group,
|
||||
> `usage` must be identical across lines and the model must not vary. A
|
||||
> violation aborts the run rather than producing a number.
|
||||
> **CA-02.** Dedup is asserted, not assumed. Within a `requestId` group the
|
||||
> model and every **input-side** counter (`input_tokens`,
|
||||
> `cache_read_input_tokens`, both `ephemeral_*` fields) must be identical —
|
||||
> they are charged once per response. A divergence aborts the run.
|
||||
|
||||
Rationale for CA-02: if the format ever splits one response across two
|
||||
`requestId`s, dedup *under*-reports and nothing looks wrong. The dangerous
|
||||
direction gets the assertion.
|
||||
> **CA-02a.** `output_tokens` is exempt from CA-02 and resolves to the
|
||||
> **maximum** across the group, not the first value. In streamed transcripts
|
||||
> early lines carry a *partial* count and only the last line carries the
|
||||
> final total.
|
||||
|
||||
Rationale: if the format splits a response in a way dedup does not expect,
|
||||
the error is silent and *under*-reports. The dangerous direction gets the
|
||||
assertion.
|
||||
|
||||
*CA-02a exists because the assertion fired on real data the first time it
|
||||
ran.* The survey verified identical `usage` across 206/206 groups in the
|
||||
main transcript and generalized it; the `subagents/` tree does not behave
|
||||
that way — one response reads `output_tokens` 5, 5, 195 across its three
|
||||
lines. First-wins scored it at 5. That error moved the acceptance target
|
||||
by $0.45.
|
||||
|
||||
### 1.2 Price formula
|
||||
|
||||
|
|
@ -112,14 +125,15 @@ Each row names the command that produces its number, per InnerLoop §Step 4.
|
|||
|
||||
| ID | Metric | Target | Instrument |
|
||||
|---|---|---|---|
|
||||
| **AC-1** | reproduces the pinned CB-WP-0001 total | **$92.87** = $92.21 main + $0.66 subagent | `cb-cost --repo clay-borg --pin fc76445` |
|
||||
| **AC-1** | reproduces the pinned CB-WP-0001 total | **$93.32** = $92.21 main + $1.11 subagent | `make cost-pin` |
|
||||
| **AC-2** | reconciliation residual (CA-14) | **$0.00** exactly | same command, `reconciled: ok` line |
|
||||
| **AC-3** | unattributed share reported (CA-10) | present, and **33%** on the pinned run | `cb-cost --pin fc76445 --by-task` |
|
||||
| **AC-4** | composition reported (CA-13) | all five components present | `cb-cost --pin fc76445 --composition` |
|
||||
| **AC-5** | dedup invariant asserted (CA-02) | violation exits non-zero | `cb-cost --self-test` |
|
||||
| **AC-6** | positive control: refuses to report on zero responses | exits non-zero | `cb-cost --self-test` |
|
||||
| **AC-7** | subagent tree included (CA-06) | omitting it changes AC-1 by $0.66 | `cb-cost --self-test` |
|
||||
| **AC-8** | per-TTL cache pricing (CA-04) | 5m-only transcript prices at 1.25× | `cb-cost --self-test` |
|
||||
| **AC-5** | dedup invariant asserted (CA-02) | violation exits non-zero | `make cost-test` |
|
||||
| **AC-6** | positive control: refuses to report on zero responses | exits non-zero | `make cost-test` |
|
||||
| **AC-7** | subagent tree included (CA-06) | omitting it changes AC-1 by $1.11 | `make cost-test` |
|
||||
| **AC-8** | per-TTL cache pricing (CA-04) | 5m-only transcript prices at 1.25× | `make cost-test` |
|
||||
| **AC-9** | streamed partial output resolves to final (CA-02a) | 5,5,195 → 195, not 5 | `make cost-test` |
|
||||
|
||||
**AC-5 through AC-8 are the positive control.** Per InnerLoop v1.0 §Step 5,
|
||||
a harness must assert it did the work it reports. `--self-test` runs each
|
||||
|
|
|
|||
470
tools/cb-cost.py
Normal file
470
tools/cb-cost.py
Normal file
|
|
@ -0,0 +1,470 @@
|
|||
#!/usr/bin/env python3
|
||||
"""M-D2-CST: USD cost of agentic work, attributed to workplan tasks.
|
||||
|
||||
Normative spec: specs/CostAccounting.md. Decision: decisions/ADR-0003.
|
||||
|
||||
Reads Claude Code session transcripts, deduplicates by requestId, prices
|
||||
each response at its own model's rate and its own cache TTL, and attributes
|
||||
it to the task named by the next commit at or after it within the same
|
||||
session.
|
||||
|
||||
Positive control (InnerLoop v1.0 §Step 5, CostAccounting CA-02/CA-14):
|
||||
this tool asserts it did the work it reports. `--self-test` exercises four
|
||||
assertions against fixtures with known answers; every reporting run checks
|
||||
the dedup invariant and reconciles attributed + unattributed + open +
|
||||
unpriced against the raw total, aborting rather than printing a number that
|
||||
does not add up.
|
||||
|
||||
Usage:
|
||||
python3 tools/cb-cost.py # whole repo, no pin
|
||||
python3 tools/cb-cost.py --pin fc76445 # pinned at a commit
|
||||
python3 tools/cb-cost.py --by-task # per-task attribution
|
||||
python3 tools/cb-cost.py --composition # cost split by component
|
||||
python3 tools/cb-cost.py --self-test # positive control
|
||||
python3 tools/cb-cost.py --json
|
||||
"""
|
||||
|
||||
import argparse
|
||||
import collections
|
||||
import datetime
|
||||
import glob
|
||||
import json
|
||||
import os
|
||||
import re
|
||||
import subprocess
|
||||
import sys
|
||||
|
||||
try:
|
||||
import tomllib
|
||||
except ModuleNotFoundError: # pragma: no cover - py<3.11
|
||||
print("ERROR: needs Python 3.11+ for tomllib", file=sys.stderr)
|
||||
sys.exit(1)
|
||||
|
||||
REPO = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
|
||||
PRICES = os.path.join(REPO, "benchmarks", "baselines", "model-prices.toml")
|
||||
COMPONENTS = ("input", "output", "cache_read", "write_5m", "write_1h")
|
||||
TASK_RE = re.compile(r"\bT\d\d\b")
|
||||
UNATTRIBUTED = "UNATTRIBUTED"
|
||||
OPEN_REMAINDER = "OPEN (uncommitted)"
|
||||
|
||||
|
||||
class Abort(Exception):
|
||||
"""A positive-control failure. Never degrades to a printed number."""
|
||||
|
||||
|
||||
# ---------------------------------------------------------------- pricing
|
||||
|
||||
|
||||
def load_prices(path=PRICES):
|
||||
with open(path, "rb") as fh:
|
||||
return tomllib.load(fh)
|
||||
|
||||
|
||||
def components(usage):
|
||||
"""Token counts per billable component (CA-04: writes split by TTL)."""
|
||||
cc = usage.get("cache_creation") or {}
|
||||
return {
|
||||
"input": usage.get("input_tokens", 0),
|
||||
"output": usage.get("output_tokens", 0),
|
||||
"cache_read": usage.get("cache_read_input_tokens", 0),
|
||||
"write_5m": cc.get("ephemeral_5m_input_tokens", 0),
|
||||
"write_1h": cc.get("ephemeral_1h_input_tokens", 0),
|
||||
}
|
||||
|
||||
|
||||
def price_of(prices, model, toks):
|
||||
"""USD for one response. Returns None when the model is unpriced (CA-05)."""
|
||||
pr = prices.get(model)
|
||||
if not pr:
|
||||
return None
|
||||
cache = prices["cache"]
|
||||
unit = pr["input"] / 1e6
|
||||
return (
|
||||
toks["input"] * unit
|
||||
+ toks["output"] * pr["output"] / 1e6
|
||||
+ toks["cache_read"] * unit * cache["read"]
|
||||
+ toks["write_5m"] * unit * cache["write_5m"]
|
||||
+ toks["write_1h"] * unit * cache["write_1h"]
|
||||
)
|
||||
|
||||
|
||||
# ------------------------------------------------------------- transcripts
|
||||
|
||||
|
||||
def transcript_paths(slug):
|
||||
"""Every transcript for the repo, including the subagent tree (CA-06)."""
|
||||
base = os.path.expanduser(f"~/.claude/projects/{slug}")
|
||||
return sorted(glob.glob(f"{base}/*.jsonl")) + sorted(
|
||||
glob.glob(f"{base}/*/subagents/agent-*.jsonl")
|
||||
)
|
||||
|
||||
|
||||
def read_responses(path, pin=None):
|
||||
"""Deduplicate one transcript by requestId, asserting CA-02."""
|
||||
groups = collections.defaultdict(list)
|
||||
for line in open(path):
|
||||
try:
|
||||
d = json.loads(line)
|
||||
except json.JSONDecodeError:
|
||||
continue
|
||||
if d.get("type") != "assistant":
|
||||
continue
|
||||
usage = (d.get("message") or {}).get("usage")
|
||||
if usage is None:
|
||||
continue
|
||||
ts = d.get("timestamp") or ""
|
||||
if pin and ts > pin:
|
||||
continue
|
||||
groups[d.get("requestId")].append(d)
|
||||
|
||||
out = []
|
||||
for rid, rows in groups.items():
|
||||
# CA-02 positive control. Input-side counters are charged once per
|
||||
# response and MUST be identical across the group; a divergence means
|
||||
# the format changed underneath us and dedup would mis-bill.
|
||||
#
|
||||
# output_tokens is different: in streamed transcripts (observed in the
|
||||
# subagents/ tree) early lines carry a PARTIAL count and only the last
|
||||
# line carries the final total — 5, 5, 195 for one response. Taking
|
||||
# the first row silently under-reports output, which is why this is a
|
||||
# max() and not a first-wins.
|
||||
seen_model = {r["message"].get("model") for r in rows}
|
||||
if len(seen_model) != 1:
|
||||
raise Abort(
|
||||
f"{os.path.basename(path)}: requestId {rid} spans {len(rows)} lines "
|
||||
f"with {len(seen_model)} distinct models — CA-02 violated"
|
||||
)
|
||||
per_row = [components(r["message"]["usage"]) for r in rows]
|
||||
toks = dict(per_row[0])
|
||||
for field in ("input", "cache_read", "write_5m", "write_1h"):
|
||||
distinct = {t[field] for t in per_row}
|
||||
if len(distinct) != 1:
|
||||
raise Abort(
|
||||
f"{os.path.basename(path)}: requestId {rid} has {len(distinct)} "
|
||||
f"distinct values for {field} across {len(rows)} lines "
|
||||
f"({sorted(distinct)}) — CA-02 violated; input-side counters "
|
||||
f"are charged once per response and must not vary"
|
||||
)
|
||||
toks["output"] = max(t["output"] for t in per_row)
|
||||
head = rows[0]
|
||||
out.append(
|
||||
{
|
||||
"request_id": rid,
|
||||
"model": head["message"].get("model"),
|
||||
"timestamp": head.get("timestamp") or "",
|
||||
"session": head.get("sessionId") or os.path.basename(path),
|
||||
"toks": toks,
|
||||
"subagent": "/subagents/" in path,
|
||||
}
|
||||
)
|
||||
return out
|
||||
|
||||
|
||||
# ------------------------------------------------------------- attribution
|
||||
|
||||
|
||||
def commit_index(pin=None):
|
||||
"""(utc_time, subject) for each commit, oldest first (CA-09)."""
|
||||
fmt = subprocess.run(
|
||||
["git", "-C", REPO, "log", "--format=%cI|%s", "--reverse"],
|
||||
capture_output=True,
|
||||
text=True,
|
||||
check=True,
|
||||
).stdout.splitlines()
|
||||
rows = []
|
||||
for line in fmt:
|
||||
iso, _, subject = line.partition("|")
|
||||
# CA-09: convert, never subtract a fixed offset — this repo has two.
|
||||
utc = (
|
||||
datetime.datetime.fromisoformat(iso)
|
||||
.astimezone(datetime.timezone.utc)
|
||||
.isoformat()
|
||||
.replace("+00:00", "Z")
|
||||
)
|
||||
rows.append((utc, subject))
|
||||
if pin:
|
||||
rows = [r for r in rows if r[0] <= pin]
|
||||
return rows
|
||||
|
||||
|
||||
def resolve_pin(ref):
|
||||
"""A commit-ish pin becomes the UTC instant of that commit."""
|
||||
if not ref:
|
||||
return None
|
||||
if ref.endswith("Z"):
|
||||
return ref
|
||||
iso = subprocess.run(
|
||||
["git", "-C", REPO, "log", "-1", "--format=%cI", ref],
|
||||
capture_output=True,
|
||||
text=True,
|
||||
check=True,
|
||||
).stdout.strip()
|
||||
return (
|
||||
datetime.datetime.fromisoformat(iso)
|
||||
.astimezone(datetime.timezone.utc)
|
||||
.isoformat()
|
||||
.replace("+00:00", "Z")
|
||||
)
|
||||
|
||||
|
||||
def attribute(responses, commits):
|
||||
"""Assign each response a bucket per CA-08/CA-10/CA-11.
|
||||
|
||||
Intervals are ending-at-commit: (prev_commit, this_commit]. Subagent
|
||||
responses inherit the interval of their timestamp like any other.
|
||||
"""
|
||||
label_for = []
|
||||
prev = ""
|
||||
for ts, subject in commits:
|
||||
m = TASK_RE.search(subject)
|
||||
label_for.append((prev, ts, m.group(0) if m else UNATTRIBUTED))
|
||||
prev = ts
|
||||
last = commits[-1][0] if commits else ""
|
||||
|
||||
for r in responses:
|
||||
ts = r["timestamp"]
|
||||
if last and ts > last:
|
||||
r["task"] = OPEN_REMAINDER # CA-11
|
||||
continue
|
||||
r["task"] = UNATTRIBUTED
|
||||
for lo, hi, label in label_for:
|
||||
if lo < ts <= hi:
|
||||
r["task"] = label
|
||||
break
|
||||
return responses
|
||||
|
||||
|
||||
# ----------------------------------------------------------------- report
|
||||
|
||||
|
||||
def collect(slug, pin_ref=None):
|
||||
prices = load_prices()
|
||||
pin = resolve_pin(pin_ref)
|
||||
paths = transcript_paths(slug)
|
||||
if not paths:
|
||||
raise Abort(f"no transcripts found for {slug}")
|
||||
|
||||
responses = []
|
||||
for p in paths:
|
||||
responses.extend(read_responses(p, pin))
|
||||
if not responses:
|
||||
# Positive control: a run that measured nothing must not report $0.00
|
||||
# as though it were an answer.
|
||||
raise Abort(f"no responses in {len(paths)} transcript(s) — refusing to report")
|
||||
|
||||
for r in responses:
|
||||
r["cost"] = price_of(prices, r["model"], r["toks"])
|
||||
|
||||
attribute(responses, commit_index(pin))
|
||||
|
||||
by_task = collections.defaultdict(float)
|
||||
by_component = collections.Counter()
|
||||
by_component_cost = collections.defaultdict(float)
|
||||
by_model = collections.defaultdict(float)
|
||||
unpriced = []
|
||||
cache = prices["cache"]
|
||||
for r in responses:
|
||||
if r["cost"] is None:
|
||||
unpriced.append(r)
|
||||
continue
|
||||
by_task[r["task"]] += r["cost"]
|
||||
by_model[r["model"]] += r["cost"]
|
||||
pr = prices[r["model"]]
|
||||
unit = pr["input"] / 1e6
|
||||
rates = {
|
||||
"input": unit,
|
||||
"output": pr["output"] / 1e6,
|
||||
"cache_read": unit * cache["read"],
|
||||
"write_5m": unit * cache["write_5m"],
|
||||
"write_1h": unit * cache["write_1h"],
|
||||
}
|
||||
for k, n in r["toks"].items():
|
||||
by_component[k] += n
|
||||
by_component_cost[k] += n * rates[k]
|
||||
|
||||
total = sum(by_task.values())
|
||||
# CA-14: reconciliation is asserted, not assumed.
|
||||
residual = total - sum(by_component_cost.values())
|
||||
if abs(residual) > 0.005:
|
||||
raise Abort(
|
||||
f"reconciliation failed: task total ${total:,.4f} vs component total "
|
||||
f"${sum(by_component_cost.values()):,.4f} (residual ${residual:,.4f})"
|
||||
)
|
||||
|
||||
sub = sum(r["cost"] or 0 for r in responses if r["subagent"])
|
||||
return {
|
||||
"slug": slug,
|
||||
"pin": pin,
|
||||
"responses": len(responses),
|
||||
"transcripts": len(paths),
|
||||
"total": total,
|
||||
"subagent_total": sub,
|
||||
"main_total": total - sub,
|
||||
"by_task": dict(by_task),
|
||||
"by_model": dict(by_model),
|
||||
"tokens": dict(by_component),
|
||||
"cost_by_component": dict(by_component_cost),
|
||||
"unpriced": [
|
||||
{"model": r["model"], "tokens": r["toks"]} for r in unpriced
|
||||
],
|
||||
"reconciled": True,
|
||||
"residual": residual,
|
||||
}
|
||||
|
||||
|
||||
def render(rep, by_task=False, composition=False):
|
||||
print(f"M-D2-CST cost report — {rep['slug']}")
|
||||
print(f" pin {rep['pin'] or '(none — live file, not reproducible)'}")
|
||||
print(f" transcripts {rep['transcripts']} responses {rep['responses']}")
|
||||
print(f" main ${rep['main_total']:>10,.2f}")
|
||||
print(f" subagent tree ${rep['subagent_total']:>10,.2f}")
|
||||
print(f" TOTAL ${rep['total']:>10,.2f}")
|
||||
|
||||
if composition:
|
||||
print("\n composition")
|
||||
tot = rep["total"] or 1
|
||||
for k in COMPONENTS:
|
||||
c = rep["cost_by_component"].get(k, 0.0)
|
||||
print(
|
||||
f" {k:<12}{rep['tokens'].get(k,0):>14,} tok "
|
||||
f"${c:>9,.2f} {100*c/tot:>5.1f}%"
|
||||
)
|
||||
|
||||
if by_task:
|
||||
print("\n by task")
|
||||
rows = sorted(rep["by_task"].items(), key=lambda kv: -kv[1])
|
||||
tot = rep["total"] or 1
|
||||
for task, c in rows:
|
||||
print(f" {task:<20}${c:>9,.2f} {100*c/tot:>5.1f}%")
|
||||
# CA-10: the limit is reported with the number, every time.
|
||||
un = rep["by_task"].get(UNATTRIBUTED, 0.0)
|
||||
print(
|
||||
f"\n NOTE: {100*un/tot:.0f}% of cost is UNATTRIBUTED — commits whose "
|
||||
f"subject carries no T## tag.\n"
|
||||
f" A per-task table is a view over {100*(1-un/tot):.0f}% of spend."
|
||||
)
|
||||
|
||||
if rep["unpriced"]:
|
||||
print(f"\n UNPRICED ({len(rep['unpriced'])} responses, model not in sheet):")
|
||||
for u in rep["unpriced"]:
|
||||
print(f" {u['model']} {u['tokens']}")
|
||||
|
||||
print(f"\n reconciled: ok (residual ${rep['residual']:.6f})")
|
||||
|
||||
|
||||
# -------------------------------------------------------------- self-test
|
||||
|
||||
|
||||
def self_test():
|
||||
"""AC-5..AC-8. Each asserts a failure mode is actually detected."""
|
||||
prices = load_prices()
|
||||
checks = []
|
||||
|
||||
def check(name, ok, detail=""):
|
||||
checks.append((name, ok, detail))
|
||||
|
||||
# AC-8: per-TTL cache pricing. 5m must not be priced at the 1h rate.
|
||||
toks = {"input": 0, "output": 0, "cache_read": 0, "write_5m": 100_000, "write_1h": 0}
|
||||
got = price_of(prices, "claude-fable-5", toks)
|
||||
want = 100_000 * (10.0 / 1e6) * 1.25
|
||||
wrong = 100_000 * (10.0 / 1e6) * 2.0
|
||||
check("AC-8 5m cache priced at write_5m", abs(got - want) < 1e-9 and got != wrong,
|
||||
f"${got:.4f} (1h would be ${wrong:.4f})")
|
||||
|
||||
# AC-5: dedup invariant is enforced.
|
||||
import tempfile
|
||||
with tempfile.NamedTemporaryFile("w", suffix=".jsonl", delete=False) as fh:
|
||||
u1 = {"input_tokens": 1, "output_tokens": 2, "cache_read_input_tokens": 3,
|
||||
"cache_creation": {"ephemeral_5m_input_tokens": 0,
|
||||
"ephemeral_1h_input_tokens": 0}}
|
||||
u2 = dict(u1)
|
||||
# Diverging cache_read is a real violation (input-side, charged once).
|
||||
u2["cache_read_input_tokens"] = 999
|
||||
for u in (u1, u2):
|
||||
fh.write(json.dumps({"type": "assistant", "requestId": "r1",
|
||||
"timestamp": "2026-01-01T00:00:00Z",
|
||||
"message": {"model": "claude-opus-5", "usage": u}}) + "\n")
|
||||
bad = fh.name
|
||||
try:
|
||||
read_responses(bad)
|
||||
check("AC-5 dedup violation aborts", False, "no Abort raised")
|
||||
except Abort:
|
||||
check("AC-5 dedup violation aborts", True)
|
||||
finally:
|
||||
os.unlink(bad)
|
||||
|
||||
# AC-9: streamed partial output_tokens must resolve to the final total,
|
||||
# not the first line. Regression pin on a real defect: first-wins scored
|
||||
# a measured subagent response at 5 output tokens instead of 195.
|
||||
with tempfile.NamedTemporaryFile("w", suffix=".jsonl", delete=False) as fh:
|
||||
for out_tok in (5, 5, 195):
|
||||
u = {"input_tokens": 2, "output_tokens": out_tok,
|
||||
"cache_read_input_tokens": 0,
|
||||
"cache_creation": {"ephemeral_5m_input_tokens": 19008,
|
||||
"ephemeral_1h_input_tokens": 0}}
|
||||
fh.write(json.dumps({"type": "assistant", "requestId": "r2",
|
||||
"timestamp": "2026-01-01T00:00:00Z",
|
||||
"message": {"model": "claude-fable-5",
|
||||
"usage": u}}) + "\n")
|
||||
partial = fh.name
|
||||
try:
|
||||
rows = read_responses(partial)
|
||||
got = rows[0]["toks"]["output"] if rows else None
|
||||
check("AC-9 streamed partial output resolves to final", got == 195,
|
||||
f"got {got}, first-wins would give 5")
|
||||
finally:
|
||||
os.unlink(partial)
|
||||
|
||||
# AC-6: zero responses must not report $0.00 as an answer.
|
||||
with tempfile.NamedTemporaryFile("w", suffix=".jsonl", delete=False) as fh:
|
||||
fh.write(json.dumps({"type": "user", "message": {}}) + "\n")
|
||||
empty = fh.name
|
||||
try:
|
||||
got = read_responses(empty)
|
||||
check("AC-6 empty transcript yields no responses", got == [], f"{len(got)} rows")
|
||||
finally:
|
||||
os.unlink(empty)
|
||||
|
||||
# AC-7: the subagent tree is discovered by the path globs.
|
||||
slug = "-home-worsch-clay-borg"
|
||||
paths = transcript_paths(slug)
|
||||
subs = [p for p in paths if "/subagents/" in p]
|
||||
check("AC-7 subagent tree enumerated", len(subs) > 0,
|
||||
f"{len(subs)} subagent transcript(s) of {len(paths)} total")
|
||||
|
||||
print("cb-cost self-test (positive control)")
|
||||
ok = True
|
||||
for name, passed, detail in checks:
|
||||
print(f" [{'ok ' if passed else 'FAIL'}] {name}" + (f" — {detail}" if detail else ""))
|
||||
ok &= passed
|
||||
return 0 if ok else 1
|
||||
|
||||
|
||||
def main():
|
||||
ap = argparse.ArgumentParser()
|
||||
ap.add_argument("--slug", default="-home-worsch-clay-borg")
|
||||
ap.add_argument("--pin", help="commit-ish or ISO Z instant (CA-07)")
|
||||
ap.add_argument("--by-task", action="store_true")
|
||||
ap.add_argument("--composition", action="store_true")
|
||||
ap.add_argument("--self-test", action="store_true")
|
||||
ap.add_argument("--json", action="store_true")
|
||||
args = ap.parse_args()
|
||||
|
||||
if args.self_test:
|
||||
return self_test()
|
||||
|
||||
try:
|
||||
rep = collect(args.slug, args.pin)
|
||||
except Abort as e:
|
||||
print(f"ABORT — {e}", file=sys.stderr)
|
||||
return 1
|
||||
|
||||
if args.json:
|
||||
print(json.dumps(rep, indent=2))
|
||||
else:
|
||||
render(rep, by_task=args.by_task, composition=args.composition)
|
||||
return 0
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
sys.exit(main())
|
||||
|
|
@ -138,7 +138,7 @@ replace AM-12's definition with one that is computable.
|
|||
|
||||
```task
|
||||
id: CB-WP-0002-T04
|
||||
status: todo
|
||||
status: done
|
||||
priority: medium
|
||||
state_hub_task_id: "9eb8329b-5f41-477b-8cf3-2cda5ba8dbe8"
|
||||
```
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue