`evals/` was a reserved path holding unvalidated blobs: section 12 named the directory and gave an illustrative snippet, but nothing was specified, so no tool could act on an eval file. Every eval file must now declare a `schema`, and CPF defines exactly one — `canned-prompts/eval-rubric/v0.1`. Unrecognized schemas stay legal and are skipped rather than rejected, so the format gains something actionable without becoming an evaluation language, which remains a non-goal. The schema splits along the same seam as T01 and T03. Render checks (`contains`, `not_contains`, `resolves_all`) assert properties of the rendered prompt text, need no model, and are therefore run by the reference CLI. Output criteria describe a good result and are declared but not run, because judging them requires a model. That division is now the format's consistent answer to "deterministic locally, or not". An eval references a fixture already declared in the manifest's `examples` rather than carrying its own copy, so an example that is also an eval fixture stays honest — both break together. An eval declares assessment and must not record outcomes. Results are run evidence and live outside the immutable package, per INTENT.md and section 17. Spec: 12 rewritten with 12.1, 18 (rules 17-18), 21 (`eval` verb). Reference CLI: read_eval, validate_eval, load_example_values, run_render_checks, cmd_eval; a failed render check exits non-zero. Tests 42 -> 51. examples/pqrst-estimate/evals/quality.yaml is a real eval with four render checks and four output criteria, and it passes. Fixes a latent bug reaching a fixture exposed: coerce_value assumed every value was a command-line string, so a YAML fixture carrying a real type (include_rationale: true) crashed on .lower(). Typed values are now validated but not re-parsed. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Bjefh8NUiEiahN4JLwoSKM Assistant: claude-code Assistant-Model: opus Assistant-Process: 388925@bnt-lap001 Assistant-Session: 3507023f-e0fd-4a1e-9d90-a0d4217d1502
163 lines
6 KiB
Markdown
163 lines
6 KiB
Markdown
# canned-prompts seed
|
|
|
|
> **Collect, reuse and share prompts and prompt templates.**
|
|
|
|
This bundle contains a first project seed for `canned-prompts`:
|
|
|
|
- [`INTENT.md`](INTENT.md) — project mission, boundaries, principles, and success criteria.
|
|
- [`CannedPromptFormat-v0.1.md`](CannedPromptFormat-v0.1.md) — experimental package-format specification.
|
|
- [`reference/`](reference/) — deliberately small Python CLI implementing the basic lifecycle.
|
|
- [`examples/pqrst-estimate/`](examples/pqrst-estimate/) — a real package that can be used to exercise the implementation.
|
|
- [`examples/house-style/`](examples/house-style/) — a `type: fragment` package that `pqrst-estimate` composes.
|
|
|
|
## Try the reference implementation
|
|
|
|
```bash
|
|
cd reference
|
|
python -m venv .venv
|
|
. .venv/bin/activate # Windows: .venv\Scripts\activate
|
|
pip install -r requirements.txt
|
|
|
|
# add the included examples to your local catalog
|
|
python canned_prompts.py add ../examples/house-style
|
|
python canned_prompts.py add ../examples/pqrst-estimate
|
|
|
|
# find and inspect it
|
|
python canned_prompts.py search pqrst
|
|
python canned_prompts.py show practice/pqrst-estimate
|
|
|
|
# see how each input and parameter resolves, and where the value came from
|
|
python canned_prompts.py resolve practice/pqrst-estimate \
|
|
--set session_summary="Implemented feature X, read unfamiliar code, added tests."
|
|
|
|
# render it
|
|
python canned_prompts.py render practice/pqrst-estimate \
|
|
--set session_summary="Implemented feature X, read unfamiliar code, added tests."
|
|
|
|
# publish it to the local filesystem registry
|
|
python canned_prompts.py publish ../examples/pqrst-estimate
|
|
|
|
# remove the local catalog if you want to simulate another machine, then install
|
|
python canned_prompts.py install practice/pqrst-estimate --version 0.1.0
|
|
```
|
|
|
|
Packages added from a path are filed under the registry name `local`; packages
|
|
installed from a registry are filed under that registry's name. `search` prints
|
|
qualified references:
|
|
|
|
```text
|
|
house:practice/pqrst-estimate@0.1.0 PQRST Estimate
|
|
local:practice/pqrst-estimate@0.1.0 PQRST Estimate
|
|
```
|
|
|
|
By default the reference tool uses:
|
|
|
|
```text
|
|
~/.canned-prompts/catalog
|
|
~/.canned-prompts/registry
|
|
```
|
|
|
|
Override them with:
|
|
|
|
```text
|
|
CANNED_PROMPTS_HOME=/some/path
|
|
```
|
|
|
|
or command-level `--catalog` / `--registry` options.
|
|
|
|
## Resolution vs rendering
|
|
|
|
The format separates the two steps (`CannedPromptFormat-v0.1.md` § 5.1).
|
|
Resolution decides a value for every input and parameter and may be
|
|
non-deterministic; rendering substitutes those values and always is. An input
|
|
may declare a default that is either a static value or a *derived* one — a
|
|
prompt that a capable consumer may run to produce the value, declared without
|
|
naming any resolver or model.
|
|
|
|
This reference tool never calls a model, so it resolves supplied values and
|
|
static defaults only, and reports anything it cannot derive instead of
|
|
rendering a prompt with a silent hole in it.
|
|
|
|
## Composition
|
|
|
|
A package composes another by declaring it as a dependency and binding it to an
|
|
input default. There are two kinds:
|
|
|
|
- **`include`** — inline the other package's *rendered template* as text.
|
|
Deterministic, needs no model, and the reference CLI performs it.
|
|
- **`derive`** — use the other package's *result*, obtained by running it.
|
|
Only a consumer able to run it can supply one.
|
|
|
|
`examples/pqrst-estimate` includes `examples/house-style`, so a shared style
|
|
block is a versioned package rather than copied text:
|
|
|
|
```yaml
|
|
dependencies:
|
|
prompts:
|
|
- id: practice/house-style
|
|
version: ">= 0.1.0"
|
|
requirement: required
|
|
|
|
inputs:
|
|
- name: house_style
|
|
required: false
|
|
default:
|
|
include: practice/house-style
|
|
```
|
|
|
|
A dependency pins an exact version by default; `any`, `newest` and a `>= X.Y.Z`
|
|
lower bound are explicit opt-ins. These are per-dependency selectors, not
|
|
version ranges — there is no solver, and constraint resolution across a
|
|
dependency graph remains a non-goal.
|
|
|
|
There is no template inheritance. A package never extends another or overrides
|
|
its parts; composition is by reference only, so a package's content stays
|
|
readable without chasing ancestors.
|
|
|
|
## Evals
|
|
|
|
An eval file declares what to assess. `evals/` used to hold whatever an author
|
|
put there; it now has one schema the tooling understands, split along the same
|
|
line as composition:
|
|
|
|
- **render checks** — deterministic assertions about the *rendered prompt*
|
|
(`contains`, `not_contains`, `resolves_all`). No model needed, so the
|
|
reference CLI runs them.
|
|
- **output criteria** — statements about a good *result*. Declared, not run.
|
|
|
|
```bash
|
|
python canned_prompts.py eval practice/pqrst-estimate
|
|
```
|
|
|
|
```text
|
|
local:practice/pqrst-estimate@0.2.0
|
|
evals/quality.yaml (pqrst-estimate-quality)
|
|
render PASS contains "must sum to exactly 100%"
|
|
render PASS resolves_all
|
|
output -- 4 criteria declared (not run: judging output needs a model)
|
|
```
|
|
|
|
An eval declares assessment, never results. Results are run evidence and live
|
|
outside the immutable package.
|
|
|
|
## Registries and identity
|
|
|
|
An id names a package *within a registry* (`CannedPromptFormat-v0.1.md`
|
|
§ 3.2). The same id obtained from two registries may be two different
|
|
packages, so the catalog keeps them apart and a bare id that matches more than
|
|
one is reported as ambiguous rather than guessed. Qualify it when you need to:
|
|
|
|
```bash
|
|
python canned_prompts.py render house:practice/pqrst-estimate --set ...
|
|
```
|
|
|
|
A registry may describe itself with an optional `registry.yaml` naming it and
|
|
recording which namespaces are claimed and under what policy. Those claims are
|
|
descriptive: a filesystem registry cannot authenticate a publisher, and
|
|
signing and trust scoring are explicit non-goals. Ownership lives with the
|
|
registry rather than in the package, so no package carries an unverifiable
|
|
assertion of authority.
|
|
|
|
## Deliberate limitations
|
|
|
|
This seed has no hosted registry, model execution, authentication, network access, dependency resolver, or social features. `publish` and `install` operate on a filesystem registry so that the package semantics can be tested before infrastructure is built around them.
|