markitect-main/markitect/infospace
tegwick c0615c2d50 feat(infospace,llm): stabilize free-tier eval workflow
Five improvements that eliminate most of the agent-in-the-loop friction
observed while closing out the 988-entity WoN evaluation (C.1):

1. Gemini adapter now retries on 429 + 5xx with exponential backoff
   (same pattern already used by OpenRouter/OpenAI). Removes the need
   for shell-level retry wrappers when hitting free-tier rate limits.

2. evaluate CLI prints the underlying error ("ERROR — HTTP 503 …")
   instead of a bare "ERROR", so agents don't have to drop into Python
   to diagnose transient failures.

3. --entity/--chapter now respect existing evaluation files by default
   (previously only the full-collection pass did). New --force flag
   opts into re-evaluation. Stops silently burning free-tier quota on
   re-runs of the same slug.

4. --entity accepts hyphenated slugs (matching entity filenames) and
   normalizes them to the underscore form used on disk. On a miss the
   CLI suggests near matches instead of a bare "not found".

5. eval-summary --update-metrics is no longer destructive:
   read_metrics_file/write_metrics_file preserve structured values
   (type_distribution) and don't flatten ints to floats. Fixes a
   silent data loss observed on every run.

Bonus: the evaluator field in written evaluation frontmatter now
falls back from run_config.model_name to the adapter's resolved model
(or the model echoed back in the API response), so rows no longer
show `evaluator: null` when --model is omitted.

Tests: new tests/unit/llm/test_gemini.py covers retry behavior;
tests/unit/infospace/test_history.py gains a round-trip test that
pins the type_distribution / int-preservation invariants.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-04-22 00:51:00 +02:00
..
checks docs(metrics): clarify C2 coverage — domain×chapter matrix, not domain×VSM 2026-02-20 00:08:46 +01:00
__init__.py feat(infospace): add infospace configuration model and state (S2.1) 2026-02-19 01:44:14 +01:00
classification.py feat(example): add L2 classifications for 823/988 WoN entities (S3.4) 2026-02-23 12:49:11 +01:00
classification_io.py feat(example): add L2 classifications for 823/988 WoN entities (S3.4) 2026-02-23 12:49:11 +01:00
classifier.py feat(example): add L2 classifications for 823/988 WoN entities (S3.4) 2026-02-23 12:49:11 +01:00
cli.py feat(infospace,llm): stabilize free-tier eval workflow 2026-04-22 00:51:00 +02:00
composition.py feat(infospace): add composition model for discipline binding (S2.6) 2026-02-19 02:03:54 +01:00
config.py feat(infospace): add L2 entity classification with type × VSM matrix (S2.9) 2026-02-23 09:35:58 +01:00
entity_parser.py feat(example): add supply-chain-vsm composition demo (S3.5) 2026-02-23 00:08:51 +01:00
evaluate.py feat(infospace,llm): stabilize free-tier eval workflow 2026-04-22 00:51:00 +02:00
evaluation.py feat(infospace): add structured evaluation output with history and diffing (S1.5) 2026-02-19 01:35:22 +01:00
evaluation_io.py feat(infospace): add structured evaluation output with history and diffing (S1.5) 2026-02-19 01:35:22 +01:00
graph_export.py feat(infospace): add entity-relation graph export (Mermaid + DOT) 2026-02-23 13:14:25 +01:00
history.py feat(infospace,llm): stabilize free-tier eval workflow 2026-04-22 00:51:00 +02:00
models.py feat(infospace): add entity metadata parser (S1.1) 2026-02-19 00:27:45 +01:00
pipeline.py fix(pipeline): retry on all LLM errors (not just rate limits) 2026-02-19 20:32:23 +01:00
relation_models.py feat(infospace): add L3 relation graph with VSM-aware triplets (S2.8) 2026-02-23 06:04:28 +01:00
relation_parser.py feat(infospace): add L3 relation graph with VSM-aware triplets (S2.8) 2026-02-23 06:04:28 +01:00
schema.py feat(infospace): add schema compliance validator (S1.2) 2026-02-19 00:48:57 +01:00
state.py feat(infospace): add infospace configuration model and state (S2.1) 2026-02-19 01:44:14 +01:00
validator.py feat(infospace): add schema compliance validator (S1.2) 2026-02-19 00:48:57 +01:00