CANP-WP-0002 T04: eval schema, render checks and output criteria
`evals/` was a reserved path holding unvalidated blobs: section 12 named the directory and gave an illustrative snippet, but nothing was specified, so no tool could act on an eval file. Every eval file must now declare a `schema`, and CPF defines exactly one — `canned-prompts/eval-rubric/v0.1`. Unrecognized schemas stay legal and are skipped rather than rejected, so the format gains something actionable without becoming an evaluation language, which remains a non-goal. The schema splits along the same seam as T01 and T03. Render checks (`contains`, `not_contains`, `resolves_all`) assert properties of the rendered prompt text, need no model, and are therefore run by the reference CLI. Output criteria describe a good result and are declared but not run, because judging them requires a model. That division is now the format's consistent answer to "deterministic locally, or not". An eval references a fixture already declared in the manifest's `examples` rather than carrying its own copy, so an example that is also an eval fixture stays honest — both break together. An eval declares assessment and must not record outcomes. Results are run evidence and live outside the immutable package, per INTENT.md and section 17. Spec: 12 rewritten with 12.1, 18 (rules 17-18), 21 (`eval` verb). Reference CLI: read_eval, validate_eval, load_example_values, run_render_checks, cmd_eval; a failed render check exits non-zero. Tests 42 -> 51. examples/pqrst-estimate/evals/quality.yaml is a real eval with four render checks and four output criteria, and it passes. Fixes a latent bug reaching a fixture exposed: coerce_value assumed every value was a command-line string, so a YAML fixture carrying a real type (include_rationale: true) crashed on .lower(). Typed values are now validated but not re-parsed. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Bjefh8NUiEiahN4JLwoSKM Assistant: claude-code Assistant-Model: opus Assistant-Process: 388925@bnt-lap001 Assistant-Session: 3507023f-e0fd-4a1e-9d90-a0d4217d1502
This commit is contained in:
parent
cac6bc135f
commit
c5f1640454
8 changed files with 444 additions and 18 deletions
26
README.md
26
README.md
|
|
@ -114,6 +114,32 @@ There is no template inheritance. A package never extends another or overrides
|
|||
its parts; composition is by reference only, so a package's content stays
|
||||
readable without chasing ancestors.
|
||||
|
||||
## Evals
|
||||
|
||||
An eval file declares what to assess. `evals/` used to hold whatever an author
|
||||
put there; it now has one schema the tooling understands, split along the same
|
||||
line as composition:
|
||||
|
||||
- **render checks** — deterministic assertions about the *rendered prompt*
|
||||
(`contains`, `not_contains`, `resolves_all`). No model needed, so the
|
||||
reference CLI runs them.
|
||||
- **output criteria** — statements about a good *result*. Declared, not run.
|
||||
|
||||
```bash
|
||||
python canned_prompts.py eval practice/pqrst-estimate
|
||||
```
|
||||
|
||||
```text
|
||||
local:practice/pqrst-estimate@0.2.0
|
||||
evals/quality.yaml (pqrst-estimate-quality)
|
||||
render PASS contains "must sum to exactly 100%"
|
||||
render PASS resolves_all
|
||||
output -- 4 criteria declared (not run: judging output needs a model)
|
||||
```
|
||||
|
||||
An eval declares assessment, never results. Results are run evidence and live
|
||||
outside the immutable package.
|
||||
|
||||
## Registries and identity
|
||||
|
||||
An id names a package *within a registry* (`CannedPromptFormat-v0.1.md`
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue