canned-prompts/reference/README.md
tegwick c5f1640454 CANP-WP-0002 T04: eval schema, render checks and output criteria
`evals/` was a reserved path holding unvalidated blobs: section 12 named the
directory and gave an illustrative snippet, but nothing was specified, so no
tool could act on an eval file.

Every eval file must now declare a `schema`, and CPF defines exactly one —
`canned-prompts/eval-rubric/v0.1`. Unrecognized schemas stay legal and are
skipped rather than rejected, so the format gains something actionable without
becoming an evaluation language, which remains a non-goal.

The schema splits along the same seam as T01 and T03. Render checks
(`contains`, `not_contains`, `resolves_all`) assert properties of the rendered
prompt text, need no model, and are therefore run by the reference CLI. Output
criteria describe a good result and are declared but not run, because judging
them requires a model. That division is now the format's consistent answer to
"deterministic locally, or not".

An eval references a fixture already declared in the manifest's `examples`
rather than carrying its own copy, so an example that is also an eval fixture
stays honest — both break together.

An eval declares assessment and must not record outcomes. Results are run
evidence and live outside the immutable package, per INTENT.md and section 17.

Spec: 12 rewritten with 12.1, 18 (rules 17-18), 21 (`eval` verb).

Reference CLI: read_eval, validate_eval, load_example_values,
run_render_checks, cmd_eval; a failed render check exits non-zero.
Tests 42 -> 51.

examples/pqrst-estimate/evals/quality.yaml is a real eval with four render
checks and four output criteria, and it passes.

Fixes a latent bug reaching a fixture exposed: coerce_value assumed every
value was a command-line string, so a YAML fixture carrying a real type
(include_rationale: true) crashed on .lower(). Typed values are now validated
but not re-parsed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Bjefh8NUiEiahN4JLwoSKM

Assistant: claude-code
Assistant-Model: opus
Assistant-Process: 388925@bnt-lap001
Assistant-Session: 3507023f-e0fd-4a1e-9d90-a0d4217d1502
2026-09-06 01:38:11 +02:00

87 lines
3.1 KiB
Markdown

# canned-prompts reference CLI
This is intentionally a **small reference implementation**, not the intended final architecture.
It demonstrates eight verbs:
```text
add PATH
search QUERY
show ID
resolve ID --set key=value
render ID --set key=value
eval ID
install ID [--version VERSION]
publish PATH
```
The implementation uses a local catalog plus a filesystem registry and performs no model calls.
`resolve` and `render` are separate because the specification separates them
(§ 5.1): resolution decides each value and may be non-deterministic, rendering
substitutes and always is. `resolve` prints where every value came from —
supplied, default, or fallback — before any prompt is produced.
## Stores
Default locations:
```text
~/.canned-prompts/catalog
~/.canned-prompts/registry
```
A **registry** stores packages flat, because an id is unambiguous within one
registry:
```text
<registry>/<id path>/<version>/...
```
A **catalog** is namespaced by registry, because identity is registry-scoped
(§ 3.2) and the same id may be installed from more than one place:
```text
<catalog>/<registry name>/<id path>/<version>/...
```
For example:
```text
~/.canned-prompts/catalog/house/practice/pqrst-estimate/0.1.0/
~/.canned-prompts/catalog/local/practice/pqrst-estimate/0.1.0/
```
A registry's name comes from its optional `registry.yaml`, and otherwise from
its directory basename. `add` takes a package from a path rather than a
registry, so it files it under `local` (override with `--as`).
Commands that take an ID accept a bare id or a qualified `<registry>:<id>`.
A bare id installed from more than one registry is reported as ambiguous
rather than resolved by guessing.
## Design choices
- YAML manifest via PyYAML.
- `{{ name }}` template substitution only.
- No arbitrary expression/code execution.
- Published versions are immutable by default.
- `install` copies from registry to catalog.
- `add` copies a package directly to catalog.
- `search`, `show`, `resolve`, and `render` operate on catalog packages, and
print qualified `<registry>:<id>` references.
- `include` defaults are satisfied (inclusion is deterministic); `derive`
defaults are reported, not run. Inclusion cycles are detected and named.
- `eval` runs the deterministic render checks of any eval declaring the
`canned-prompts/eval-rubric/v0.1` schema, and reports output criteria as
declared but not run. Unrecognized schemas are skipped, not rejected. A
failed render check exits non-zero.
- An optional `registry.yaml` names a registry and records namespace claims.
`publish` warns when a namespace is declared `closed` — it cannot
authenticate a publisher, and says so rather than implying it checked.
- Static input defaults are applied; **derived** defaults (§ 6.1) are not. This
tool never calls a model, so a derived default is satisfied only by its
static fallback `value`. Without one, `resolve` reports the input as
unresolved and `render` refuses rather than substituting empty text.
Use this implementation to challenge the format. Replace it once real usage reveals the right architecture.