infospace-bench/tests
tegwick 3ca891de4a fix: review findings from Lefevre live smoke
Two small fixes informed by the 2026-05-18 live OpenRouter chapter-I run.

1. extract-entities templates (trading-literature and general-knowledge):
   the # Entity Title placeholder was interpreted by gpt-4o-mini as a
   literal heading prefix, so every entity came back as "# Entity Title:
   Bucket Shop" etc. The instruction now spells the placeholder out
   with concrete examples and an explicit "not the literal string"
   note, so smaller models hit the intended shape.

2. generate plan grows --model <id>. When supplied, the cost estimate
   pulls per-prompt and per-completion rates from the bundled
   model_rates.yaml instead of multiplying a single blended
   --cost-per-1k value across all tokens. The summary now also returns
   a separate estimated_completion_tokens field plus a cost_source tag
   ("rate_table:<model>" | "cost_per_1k_blended" | None).

This is a stopgap. LLM-WP-0005 (proposed in llm-connect this round)
will move the rate registry and token-shape problem classes upstream
so consumers stop re-implementing them.

The live smoke ran 28k prompt tokens / 7.5k completion / $0.0088
actual. With --model openai/gpt-4o-mini the plan estimate now lands at
$0.0076 (within 14% of actual) versus the prior $8.40 estimate at
--cost-per-1k 0.30.

181 tests pass, 2 skipped (both live OpenRouter smokes).

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-05-19 04:30:33 +02:00
..
fixtures/lefevre IB-WP-0016-T05: deterministic Lefevre acceptance fixture 2026-05-17 22:31:17 +02:00
test_agentic_memory_profile.py Agentic memory profile 2026-05-15 16:01:35 +02:00
test_archive.py archive: include contracts/, schemas/; report skipped top-level dirs 2026-05-17 12:21:19 +02:00
test_budget_registry.py IB-WP-0020-T03: routing CLI flags 2026-05-18 22:08:51 +02:00
test_cli.py IB-WP-0014: archive-list, restore, retention annotation, docs (T03-T05) 2026-05-17 11:46:23 +02:00
test_engine_boundary.py engine and lifecycle 2026-05-14 16:26:42 +02:00
test_epub3_intake.py IB-WP-0016-T02: chapter-aware chunking and stable IDs 2026-05-17 15:52:47 +02:00
test_evaluation.py Initial implementation 2026-05-14 11:32:25 +02:00
test_evaluation_history.py eval history and metrics 2026-05-14 15:35:04 +02:00
test_generic_generator.py generic source-to-infospace generator 2026-05-14 19:33:22 +02:00
test_inspection.py Initial implementation 2026-05-14 11:32:25 +02:00
test_lefevre_fixture.py IB-WP-0016-T07: review report and output policy; close IB-WP-0016 2026-05-18 01:22:41 +02:00
test_legacy_pilot.py Kontextual Engine Integration Boundary 2026-05-14 16:43:29 +02:00
test_lifecycle.py Initial implementation 2026-05-14 11:32:25 +02:00
test_markdown_adapter.py markitect-tool integration 2026-05-14 14:53:16 +02:00
test_openrouter_live.py IB-WP-0020-T04: example routing config + live routing smoke 2026-05-18 22:19:54 +02:00
test_plan_scale.py fix: review findings from Lefevre live smoke 2026-05-19 04:30:33 +02:00
test_reference_pilot.py Initial implementation 2026-05-14 11:32:25 +02:00
test_replacement_readiness.py command parity and migration guide 2026-05-14 17:16:39 +02:00
test_routing_adapter.py IB-WP-0018-T03+T04: shadow sampling + report/CLI surfacing; close IB-WP-0018 2026-05-18 11:52:05 +02:00
test_routing_cli.py IB-WP-0020-T03: routing CLI flags 2026-05-18 22:08:51 +02:00
test_routing_config.py IB-WP-0020-T05: shadow-mode CLI flags; close IB-WP-0020 2026-05-18 23:30:36 +02:00
test_semantics.py entity relationship model 2026-05-14 15:06:17 +02:00
test_trading_literature_profile.py IB-WP-0016-T04: trading-literature profile 2026-05-17 18:59:45 +02:00
test_wealth_vsm_generation.py infospace pipeline for wealth of nations example 2026-05-14 18:04:38 +02:00
test_workflow.py acceptance matrix and workflow generation 2026-05-14 16:01:32 +02:00