canned-prompts/reference
tegwick c5f1640454 CANP-WP-0002 T04: eval schema, render checks and output criteria
`evals/` was a reserved path holding unvalidated blobs: section 12 named the
directory and gave an illustrative snippet, but nothing was specified, so no
tool could act on an eval file.

Every eval file must now declare a `schema`, and CPF defines exactly one —
`canned-prompts/eval-rubric/v0.1`. Unrecognized schemas stay legal and are
skipped rather than rejected, so the format gains something actionable without
becoming an evaluation language, which remains a non-goal.

The schema splits along the same seam as T01 and T03. Render checks
(`contains`, `not_contains`, `resolves_all`) assert properties of the rendered
prompt text, need no model, and are therefore run by the reference CLI. Output
criteria describe a good result and are declared but not run, because judging
them requires a model. That division is now the format's consistent answer to
"deterministic locally, or not".

An eval references a fixture already declared in the manifest's `examples`
rather than carrying its own copy, so an example that is also an eval fixture
stays honest — both break together.

An eval declares assessment and must not record outcomes. Results are run
evidence and live outside the immutable package, per INTENT.md and section 17.

Spec: 12 rewritten with 12.1, 18 (rules 17-18), 21 (`eval` verb).

Reference CLI: read_eval, validate_eval, load_example_values,
run_render_checks, cmd_eval; a failed render check exits non-zero.
Tests 42 -> 51.

examples/pqrst-estimate/evals/quality.yaml is a real eval with four render
checks and four output criteria, and it passes.

Fixes a latent bug reaching a fixture exposed: coerce_value assumed every
value was a command-line string, so a YAML fixture carrying a real type
(include_rationale: true) crashed on .lower(). Typed values are now validated
but not re-parsed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Bjefh8NUiEiahN4JLwoSKM

Assistant: claude-code
Assistant-Model: opus
Assistant-Process: 388925@bnt-lap001
Assistant-Session: 3507023f-e0fd-4a1e-9d90-a0d4217d1502
2026-09-06 01:38:11 +02:00
..
tests CANP-WP-0002 T04: eval schema, render checks and output criteria 2026-09-06 01:38:11 +02:00
canned_prompts.py CANP-WP-0002 T04: eval schema, render checks and output criteria 2026-09-06 01:38:11 +02:00
pyproject.toml Register with Custodian State Hub and seed format open-questions workplan 2026-09-06 00:45:28 +02:00
README.md CANP-WP-0002 T04: eval schema, render checks and output criteria 2026-09-06 01:38:11 +02:00
requirements-dev.txt Register with Custodian State Hub and seed format open-questions workplan 2026-09-06 00:45:28 +02:00
requirements.txt Register with Custodian State Hub and seed format open-questions workplan 2026-09-06 00:45:28 +02:00

canned-prompts reference CLI

This is intentionally a small reference implementation, not the intended final architecture.

It demonstrates eight verbs:

add      PATH
search   QUERY
show     ID
resolve  ID --set key=value
render   ID --set key=value
eval     ID
install  ID [--version VERSION]
publish  PATH

The implementation uses a local catalog plus a filesystem registry and performs no model calls.

resolve and render are separate because the specification separates them (§ 5.1): resolution decides each value and may be non-deterministic, rendering substitutes and always is. resolve prints where every value came from — supplied, default, or fallback — before any prompt is produced.

Stores

Default locations:

~/.canned-prompts/catalog
~/.canned-prompts/registry

A registry stores packages flat, because an id is unambiguous within one registry:

<registry>/<id path>/<version>/...

A catalog is namespaced by registry, because identity is registry-scoped (§ 3.2) and the same id may be installed from more than one place:

<catalog>/<registry name>/<id path>/<version>/...

For example:

~/.canned-prompts/catalog/house/practice/pqrst-estimate/0.1.0/
~/.canned-prompts/catalog/local/practice/pqrst-estimate/0.1.0/

A registry's name comes from its optional registry.yaml, and otherwise from its directory basename. add takes a package from a path rather than a registry, so it files it under local (override with --as).

Commands that take an ID accept a bare id or a qualified <registry>:<id>. A bare id installed from more than one registry is reported as ambiguous rather than resolved by guessing.

Design choices

  • YAML manifest via PyYAML.
  • {{ name }} template substitution only.
  • No arbitrary expression/code execution.
  • Published versions are immutable by default.
  • install copies from registry to catalog.
  • add copies a package directly to catalog.
  • search, show, resolve, and render operate on catalog packages, and print qualified <registry>:<id> references.
  • include defaults are satisfied (inclusion is deterministic); derive defaults are reported, not run. Inclusion cycles are detected and named.
  • eval runs the deterministic render checks of any eval declaring the canned-prompts/eval-rubric/v0.1 schema, and reports output criteria as declared but not run. Unrecognized schemas are skipped, not rejected. A failed render check exits non-zero.
  • An optional registry.yaml names a registry and records namespace claims. publish warns when a namespace is declared closed — it cannot authenticate a publisher, and says so rather than implying it checked.
  • Static input defaults are applied; derived defaults (§ 6.1) are not. This tool never calls a model, so a derived default is satisfied only by its static fallback value. Without one, resolve reports the input as unresolved and render refuses rather than substituting empty text.

Use this implementation to challenge the format. Replace it once real usage reveals the right architecture.