canned-prompts/reference/README.md

96 lines
3.6 KiB
Markdown
Raw Normal View History

# canned-prompts reference CLI
This is intentionally a **small reference implementation**, not the intended final architecture.
CANP-WP-0002 T04: eval schema, render checks and output criteria `evals/` was a reserved path holding unvalidated blobs: section 12 named the directory and gave an illustrative snippet, but nothing was specified, so no tool could act on an eval file. Every eval file must now declare a `schema`, and CPF defines exactly one — `canned-prompts/eval-rubric/v0.1`. Unrecognized schemas stay legal and are skipped rather than rejected, so the format gains something actionable without becoming an evaluation language, which remains a non-goal. The schema splits along the same seam as T01 and T03. Render checks (`contains`, `not_contains`, `resolves_all`) assert properties of the rendered prompt text, need no model, and are therefore run by the reference CLI. Output criteria describe a good result and are declared but not run, because judging them requires a model. That division is now the format's consistent answer to "deterministic locally, or not". An eval references a fixture already declared in the manifest's `examples` rather than carrying its own copy, so an example that is also an eval fixture stays honest — both break together. An eval declares assessment and must not record outcomes. Results are run evidence and live outside the immutable package, per INTENT.md and section 17. Spec: 12 rewritten with 12.1, 18 (rules 17-18), 21 (`eval` verb). Reference CLI: read_eval, validate_eval, load_example_values, run_render_checks, cmd_eval; a failed render check exits non-zero. Tests 42 -> 51. examples/pqrst-estimate/evals/quality.yaml is a real eval with four render checks and four output criteria, and it passes. Fixes a latent bug reaching a fixture exposed: coerce_value assumed every value was a command-line string, so a YAML fixture carrying a real type (include_rationale: true) crashed on .lower(). Typed values are now validated but not re-parsed. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Bjefh8NUiEiahN4JLwoSKM Assistant: claude-code Assistant-Model: opus Assistant-Process: 388925@bnt-lap001 Assistant-Session: 3507023f-e0fd-4a1e-9d90-a0d4217d1502
2026-09-06 01:38:11 +02:00
It demonstrates eight verbs:
```text
add PATH
search QUERY
show ID
CANP-WP-0002 T01: input defaults, static and derived Closes the gap that made optional inputs unusable: rendering rule 5.1(4) made any unresolved placeholder an error while inputs had no `default`, so an input marked `required: false` and referenced from the template failed every render in which the caller omitted it — including the spec's own section 4 example. Section 10 already carried the derivation mechanism (`requirement: generate`, resolution deliberately undefined), so a derived default needed a binding rather than a new concept: the input's default names a declared prompt dependency. Spec: - 5.1 rewritten as "Resolution and rendering". Resolution may be non-deterministic and must report what it derived; rendering is deterministic and must not derive. A tool that handles only supplied values and static defaults is stated to be conforming. - 6.1 (new) covers both declaration forms. Reference form is preferred, with the reason stated — an inline prompt is anonymous, so unversioned, unprovenanced and un-evaluable — and validators should warn when a published package derives inline. - A derived default may declare a static fallback `value`. Without one the input stays unresolved, which is an error; derivation never silently yields empty content. - 4, 10, 18 (rules 11-13), 19 (two new MUST NOTs), 21 updated accordingly. Reference CLI: - New `resolve` verb reporting the origin of every value. - `Resolution` dataclass and `resolve_inputs`; `resolve_values` kept as a wrapper so existing callers are unaffected. - `render` refuses with a specific error naming underivable inputs rather than substituting empty text. - Tests 3 -> 11. Example package lifecycle re-verified end to end. INTENT.md is unchanged: splitting resolve from render preserves success criterion 4 (deterministic rendering) as written. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Bjefh8NUiEiahN4JLwoSKM Assistant: claude-code Assistant-Model: opus Assistant-Process: 388925@bnt-lap001 Assistant-Session: 3507023f-e0fd-4a1e-9d90-a0d4217d1502
2026-09-06 00:59:20 +02:00
resolve ID --set key=value
render ID --set key=value
CANP-WP-0002 T04: eval schema, render checks and output criteria `evals/` was a reserved path holding unvalidated blobs: section 12 named the directory and gave an illustrative snippet, but nothing was specified, so no tool could act on an eval file. Every eval file must now declare a `schema`, and CPF defines exactly one — `canned-prompts/eval-rubric/v0.1`. Unrecognized schemas stay legal and are skipped rather than rejected, so the format gains something actionable without becoming an evaluation language, which remains a non-goal. The schema splits along the same seam as T01 and T03. Render checks (`contains`, `not_contains`, `resolves_all`) assert properties of the rendered prompt text, need no model, and are therefore run by the reference CLI. Output criteria describe a good result and are declared but not run, because judging them requires a model. That division is now the format's consistent answer to "deterministic locally, or not". An eval references a fixture already declared in the manifest's `examples` rather than carrying its own copy, so an example that is also an eval fixture stays honest — both break together. An eval declares assessment and must not record outcomes. Results are run evidence and live outside the immutable package, per INTENT.md and section 17. Spec: 12 rewritten with 12.1, 18 (rules 17-18), 21 (`eval` verb). Reference CLI: read_eval, validate_eval, load_example_values, run_render_checks, cmd_eval; a failed render check exits non-zero. Tests 42 -> 51. examples/pqrst-estimate/evals/quality.yaml is a real eval with four render checks and four output criteria, and it passes. Fixes a latent bug reaching a fixture exposed: coerce_value assumed every value was a command-line string, so a YAML fixture carrying a real type (include_rationale: true) crashed on .lower(). Typed values are now validated but not re-parsed. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Bjefh8NUiEiahN4JLwoSKM Assistant: claude-code Assistant-Model: opus Assistant-Process: 388925@bnt-lap001 Assistant-Session: 3507023f-e0fd-4a1e-9d90-a0d4217d1502
2026-09-06 01:38:11 +02:00
eval ID
install ID [--version VERSION]
publish PATH
```
The implementation uses a local catalog plus a filesystem registry and performs no model calls.
CANP-WP-0002 T01: input defaults, static and derived Closes the gap that made optional inputs unusable: rendering rule 5.1(4) made any unresolved placeholder an error while inputs had no `default`, so an input marked `required: false` and referenced from the template failed every render in which the caller omitted it — including the spec's own section 4 example. Section 10 already carried the derivation mechanism (`requirement: generate`, resolution deliberately undefined), so a derived default needed a binding rather than a new concept: the input's default names a declared prompt dependency. Spec: - 5.1 rewritten as "Resolution and rendering". Resolution may be non-deterministic and must report what it derived; rendering is deterministic and must not derive. A tool that handles only supplied values and static defaults is stated to be conforming. - 6.1 (new) covers both declaration forms. Reference form is preferred, with the reason stated — an inline prompt is anonymous, so unversioned, unprovenanced and un-evaluable — and validators should warn when a published package derives inline. - A derived default may declare a static fallback `value`. Without one the input stays unresolved, which is an error; derivation never silently yields empty content. - 4, 10, 18 (rules 11-13), 19 (two new MUST NOTs), 21 updated accordingly. Reference CLI: - New `resolve` verb reporting the origin of every value. - `Resolution` dataclass and `resolve_inputs`; `resolve_values` kept as a wrapper so existing callers are unaffected. - `render` refuses with a specific error naming underivable inputs rather than substituting empty text. - Tests 3 -> 11. Example package lifecycle re-verified end to end. INTENT.md is unchanged: splitting resolve from render preserves success criterion 4 (deterministic rendering) as written. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Bjefh8NUiEiahN4JLwoSKM Assistant: claude-code Assistant-Model: opus Assistant-Process: 388925@bnt-lap001 Assistant-Session: 3507023f-e0fd-4a1e-9d90-a0d4217d1502
2026-09-06 00:59:20 +02:00
`resolve` and `render` are separate because the specification separates them
(§ 5.1): resolution decides each value and may be non-deterministic, rendering
substitutes and always is. `resolve` prints where every value came from —
supplied, default, or fallback — before any prompt is produced.
## Stores
Default locations:
```text
~/.canned-prompts/catalog
~/.canned-prompts/registry
```
CANP-WP-0002 T02: registry-scoped identity and namespace ownership Section 17 already defined what a registry does when the same id@version is republished; nothing defined what a *consumer* does. That is where namespace conflict actually bites — in the catalog, after installing from two registries. Identity is now registry-scoped: an id names a package within a registry, the way a path names a file within a repository. A local-first format with no signing, no federation and no central authority cannot enforce global uniqueness, and an unenforceable guarantee is worse than none — it invites consumers to conflate two packages that merely share a name. A qualified `<registry>:<id>` reference distinguishes them, and `:` is now barred from ids so the separator stays available. Installing the same id from two registries is therefore not a conflict. The catalog is namespaced by registry and keeps both. Ownership is registry policy, not package data. An optional `registry.yaml` names a registry and records namespace claims. Those claims are explicitly descriptive — a filesystem registry cannot authenticate a publisher, and `publish` says so rather than implying it checked. Keeping the claim out of packages leaves artifacts free of unverifiable assertions of authority, and means package semantics do not change when a hosted registry appears later. Spec: 3.2 (registry-scoped identity, qualified references), 17 (immutability scoped to a registry), 20.1 and 20.2 (new), 18 (registry-manifest validation), 21 (qualified references, reserved `local` name). Reference CLI: parse_reference, check_registry_name, read_registry_manifest, registry_name, namespace_policy; registry_package_path and catalog_package_path split; resolve_installed reports ambiguity and returns the source registry; iter_catalog; `add --as`; closed-namespace warning on publish. Tests 11 -> 21. The catalog layout changed. An existing catalog is detected and reported with instructions rather than failing as "package not found". Signing, trust scoring and federation remain non-goals and were not touched. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Bjefh8NUiEiahN4JLwoSKM Assistant: claude-code Assistant-Model: opus Assistant-Process: 388925@bnt-lap001 Assistant-Session: 3507023f-e0fd-4a1e-9d90-a0d4217d1502
2026-09-06 01:14:54 +02:00
A **registry** stores packages flat, because an id is unambiguous within one
registry:
```text
CANP-WP-0002 T02: registry-scoped identity and namespace ownership Section 17 already defined what a registry does when the same id@version is republished; nothing defined what a *consumer* does. That is where namespace conflict actually bites — in the catalog, after installing from two registries. Identity is now registry-scoped: an id names a package within a registry, the way a path names a file within a repository. A local-first format with no signing, no federation and no central authority cannot enforce global uniqueness, and an unenforceable guarantee is worse than none — it invites consumers to conflate two packages that merely share a name. A qualified `<registry>:<id>` reference distinguishes them, and `:` is now barred from ids so the separator stays available. Installing the same id from two registries is therefore not a conflict. The catalog is namespaced by registry and keeps both. Ownership is registry policy, not package data. An optional `registry.yaml` names a registry and records namespace claims. Those claims are explicitly descriptive — a filesystem registry cannot authenticate a publisher, and `publish` says so rather than implying it checked. Keeping the claim out of packages leaves artifacts free of unverifiable assertions of authority, and means package semantics do not change when a hosted registry appears later. Spec: 3.2 (registry-scoped identity, qualified references), 17 (immutability scoped to a registry), 20.1 and 20.2 (new), 18 (registry-manifest validation), 21 (qualified references, reserved `local` name). Reference CLI: parse_reference, check_registry_name, read_registry_manifest, registry_name, namespace_policy; registry_package_path and catalog_package_path split; resolve_installed reports ambiguity and returns the source registry; iter_catalog; `add --as`; closed-namespace warning on publish. Tests 11 -> 21. The catalog layout changed. An existing catalog is detected and reported with instructions rather than failing as "package not found". Signing, trust scoring and federation remain non-goals and were not touched. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Bjefh8NUiEiahN4JLwoSKM Assistant: claude-code Assistant-Model: opus Assistant-Process: 388925@bnt-lap001 Assistant-Session: 3507023f-e0fd-4a1e-9d90-a0d4217d1502
2026-09-06 01:14:54 +02:00
<registry>/<id path>/<version>/...
```
A **catalog** is namespaced by registry, because identity is registry-scoped
(§ 3.2) and the same id may be installed from more than one place:
```text
<catalog>/<registry name>/<id path>/<version>/...
```
For example:
```text
CANP-WP-0002 T02: registry-scoped identity and namespace ownership Section 17 already defined what a registry does when the same id@version is republished; nothing defined what a *consumer* does. That is where namespace conflict actually bites — in the catalog, after installing from two registries. Identity is now registry-scoped: an id names a package within a registry, the way a path names a file within a repository. A local-first format with no signing, no federation and no central authority cannot enforce global uniqueness, and an unenforceable guarantee is worse than none — it invites consumers to conflate two packages that merely share a name. A qualified `<registry>:<id>` reference distinguishes them, and `:` is now barred from ids so the separator stays available. Installing the same id from two registries is therefore not a conflict. The catalog is namespaced by registry and keeps both. Ownership is registry policy, not package data. An optional `registry.yaml` names a registry and records namespace claims. Those claims are explicitly descriptive — a filesystem registry cannot authenticate a publisher, and `publish` says so rather than implying it checked. Keeping the claim out of packages leaves artifacts free of unverifiable assertions of authority, and means package semantics do not change when a hosted registry appears later. Spec: 3.2 (registry-scoped identity, qualified references), 17 (immutability scoped to a registry), 20.1 and 20.2 (new), 18 (registry-manifest validation), 21 (qualified references, reserved `local` name). Reference CLI: parse_reference, check_registry_name, read_registry_manifest, registry_name, namespace_policy; registry_package_path and catalog_package_path split; resolve_installed reports ambiguity and returns the source registry; iter_catalog; `add --as`; closed-namespace warning on publish. Tests 11 -> 21. The catalog layout changed. An existing catalog is detected and reported with instructions rather than failing as "package not found". Signing, trust scoring and federation remain non-goals and were not touched. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Bjefh8NUiEiahN4JLwoSKM Assistant: claude-code Assistant-Model: opus Assistant-Process: 388925@bnt-lap001 Assistant-Session: 3507023f-e0fd-4a1e-9d90-a0d4217d1502
2026-09-06 01:14:54 +02:00
~/.canned-prompts/catalog/house/practice/pqrst-estimate/0.1.0/
~/.canned-prompts/catalog/local/practice/pqrst-estimate/0.1.0/
```
CANP-WP-0002 T02: registry-scoped identity and namespace ownership Section 17 already defined what a registry does when the same id@version is republished; nothing defined what a *consumer* does. That is where namespace conflict actually bites — in the catalog, after installing from two registries. Identity is now registry-scoped: an id names a package within a registry, the way a path names a file within a repository. A local-first format with no signing, no federation and no central authority cannot enforce global uniqueness, and an unenforceable guarantee is worse than none — it invites consumers to conflate two packages that merely share a name. A qualified `<registry>:<id>` reference distinguishes them, and `:` is now barred from ids so the separator stays available. Installing the same id from two registries is therefore not a conflict. The catalog is namespaced by registry and keeps both. Ownership is registry policy, not package data. An optional `registry.yaml` names a registry and records namespace claims. Those claims are explicitly descriptive — a filesystem registry cannot authenticate a publisher, and `publish` says so rather than implying it checked. Keeping the claim out of packages leaves artifacts free of unverifiable assertions of authority, and means package semantics do not change when a hosted registry appears later. Spec: 3.2 (registry-scoped identity, qualified references), 17 (immutability scoped to a registry), 20.1 and 20.2 (new), 18 (registry-manifest validation), 21 (qualified references, reserved `local` name). Reference CLI: parse_reference, check_registry_name, read_registry_manifest, registry_name, namespace_policy; registry_package_path and catalog_package_path split; resolve_installed reports ambiguity and returns the source registry; iter_catalog; `add --as`; closed-namespace warning on publish. Tests 11 -> 21. The catalog layout changed. An existing catalog is detected and reported with instructions rather than failing as "package not found". Signing, trust scoring and federation remain non-goals and were not touched. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Bjefh8NUiEiahN4JLwoSKM Assistant: claude-code Assistant-Model: opus Assistant-Process: 388925@bnt-lap001 Assistant-Session: 3507023f-e0fd-4a1e-9d90-a0d4217d1502
2026-09-06 01:14:54 +02:00
A registry's name comes from its optional `registry.yaml`, and otherwise from
its directory basename. `add` takes a package from a path rather than a
registry, so it files it under `local` (override with `--as`).
Commands that take an ID accept a bare id or a qualified `<registry>:<id>`.
A bare id installed from more than one registry is reported as ambiguous
rather than resolved by guessing.
## Design choices
- YAML manifest via PyYAML.
- `{{ name }}` template substitution only.
- No arbitrary expression/code execution.
- Published versions are immutable by default.
CANP-WP-0002 T07: semver precedence and strict packaging Two defects from the original review of the seed. Prerelease ordering was worse than first recorded. parse_semver returned (major, minor, patch, raw_string), so 1.0.0-rc1 and 1.0.0 tied on the numeric fields and then compared as strings — "1.0.0-rc1" > "1.0.0". A release candidate therefore shadowed its own release for `newest` and for `>=`, not just for the no-version case. parse_semver now returns a SemVer section 11 precedence key: numeric fields, a release/prerelease rank, then dot-separated prerelease identifiers with numeric ones compared numerically. Build metadata is ignored. Beyond ordering, prereleases are excluded from `any`, `newest` and `>=` entirely; only an exact pin selects one, so publishing a release candidate never changes what existing consumers resolve to. A package holding only prereleases now says so rather than reporting a bare not-found. Packaging copied the whole source directory, so a stray .git, virtualenv or scratch file landed in the catalog and registry. Section 2 already required otherwise — tools MUST ignore unknown non-reserved files unless a manifest field references them — so this is conformance rather than a new rule. What is new is that omissions are reported instead of silent: not packaged (not a reserved path, not referenced by the manifest): .git/, .venv/, notes.txt LICENSE joins the reserved paths. Strict packaging would otherwise drop a package's license text while faithfully copying its `license` field, which contradicts section 14's instruction to surface licensing on publish and install. Spec: 2 (LICENSE, packaging obligation, reporting), 17.1 new, 10.3 note. Reference CLI: parse_semver rewritten with is_prerelease; select_version and pick_version updated; copy_package and report_skipped replace copy_immutable. Tests 65 -> 78. examples/pqrst-estimate carries a LICENSE and a license field, exercising the new reserved path. Also fixes a leaked loop variable in package_members that would have reported a bad `template` path as an `evals` error. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Bjefh8NUiEiahN4JLwoSKM Assistant: claude-code Assistant-Model: opus Assistant-Process: 388925@bnt-lap001 Assistant-Session: 3507023f-e0fd-4a1e-9d90-a0d4217d1502
2026-09-06 09:32:42 +02:00
- `add`, `publish` and `install` copy only reserved paths and manifest-
referenced files. A working directory's `.git`, virtualenv or scratch files
never ship; whatever is left out is named on stderr.
- Version precedence follows SemVer: a prerelease ranks below its release, and
`any`, `newest` and `>=` skip prereleases entirely. Only an exact pin
selects one.
- `install` copies from registry to catalog.
- `add` copies a package directly to catalog.
CANP-WP-0002 T02: registry-scoped identity and namespace ownership Section 17 already defined what a registry does when the same id@version is republished; nothing defined what a *consumer* does. That is where namespace conflict actually bites — in the catalog, after installing from two registries. Identity is now registry-scoped: an id names a package within a registry, the way a path names a file within a repository. A local-first format with no signing, no federation and no central authority cannot enforce global uniqueness, and an unenforceable guarantee is worse than none — it invites consumers to conflate two packages that merely share a name. A qualified `<registry>:<id>` reference distinguishes them, and `:` is now barred from ids so the separator stays available. Installing the same id from two registries is therefore not a conflict. The catalog is namespaced by registry and keeps both. Ownership is registry policy, not package data. An optional `registry.yaml` names a registry and records namespace claims. Those claims are explicitly descriptive — a filesystem registry cannot authenticate a publisher, and `publish` says so rather than implying it checked. Keeping the claim out of packages leaves artifacts free of unverifiable assertions of authority, and means package semantics do not change when a hosted registry appears later. Spec: 3.2 (registry-scoped identity, qualified references), 17 (immutability scoped to a registry), 20.1 and 20.2 (new), 18 (registry-manifest validation), 21 (qualified references, reserved `local` name). Reference CLI: parse_reference, check_registry_name, read_registry_manifest, registry_name, namespace_policy; registry_package_path and catalog_package_path split; resolve_installed reports ambiguity and returns the source registry; iter_catalog; `add --as`; closed-namespace warning on publish. Tests 11 -> 21. The catalog layout changed. An existing catalog is detected and reported with instructions rather than failing as "package not found". Signing, trust scoring and federation remain non-goals and were not touched. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Bjefh8NUiEiahN4JLwoSKM Assistant: claude-code Assistant-Model: opus Assistant-Process: 388925@bnt-lap001 Assistant-Session: 3507023f-e0fd-4a1e-9d90-a0d4217d1502
2026-09-06 01:14:54 +02:00
- `search`, `show`, `resolve`, and `render` operate on catalog packages, and
print qualified `<registry>:<id>` references.
CANP-WP-0002 T03: composition by reference, two kinds T01 had already delivered half of composition without naming it: a derived default binds an input to a prompt dependency, which is transclusion — run package B, use its output. What was missing was the deterministic half. `include` inlines another package's rendered template as text. No model is involved, so the reference CLI can actually perform it, and a shared preamble, rubric or style block becomes a versioned package instead of copied text. This is the concrete way to honor INTENT principle 9 without any runtime. `derive` stays as it was. Both are input defaults, so composition reuses the resolution machinery rather than adding a second one. No template inheritance. Four of this repo's own documents argue against it: INTENT principle 3 (hidden context defeats reuse), section 19's "make package contents visible before execution", section 17's requirement that behavior changes produce a new version, and the non-goal on range resolution. Version selectors: an exact pin is the expected form, with `any`, `newest` and `>= X.Y.Z` as explicit opt-ins so looseness is written rather than implied by absence. Selectors are evaluated per dependency against what is available — no solver, no cross-dependency constraint satisfaction — which is what keeps them outside the range-resolution non-goal, and the spec says so. Also defines `type` (template | fragment), which appeared once in the section 4 manifest surface and was specified nowhere. Spec: 3.2 (type), 5.1 (inclusion resolution rule, renumbered), 6.1 (included default), 10.1 and 10.2 (new), 18 (rules 14-16), 21. Reference CLI: validate_version_selector, select_version, prompt_dependencies replacing prompt_dependency_ids, check_composition_reference, CatalogComposer with cycle detection, and resolve_inputs gaining composer= and inherited=. Tests 21 -> 42. Examples: house-style is a real fragment package; pqrst-estimate composes it and is bumped 0.1.0 -> 0.2.0 per section 17. Fixes an ordering bug found while testing: inputs resolved before parameters, so an included package could not see the including package's parameters and silently fell back to its own defaults — the fragment rendered tone=neutral where the including package said blunt. Parameters now resolve first; the report still lists inputs first. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Bjefh8NUiEiahN4JLwoSKM Assistant: claude-code Assistant-Model: opus Assistant-Process: 388925@bnt-lap001 Assistant-Session: 3507023f-e0fd-4a1e-9d90-a0d4217d1502
2026-09-06 01:31:58 +02:00
- `include` defaults are satisfied (inclusion is deterministic); `derive`
defaults are reported, not run. Inclusion cycles are detected and named.
CANP-WP-0002 T04: eval schema, render checks and output criteria `evals/` was a reserved path holding unvalidated blobs: section 12 named the directory and gave an illustrative snippet, but nothing was specified, so no tool could act on an eval file. Every eval file must now declare a `schema`, and CPF defines exactly one — `canned-prompts/eval-rubric/v0.1`. Unrecognized schemas stay legal and are skipped rather than rejected, so the format gains something actionable without becoming an evaluation language, which remains a non-goal. The schema splits along the same seam as T01 and T03. Render checks (`contains`, `not_contains`, `resolves_all`) assert properties of the rendered prompt text, need no model, and are therefore run by the reference CLI. Output criteria describe a good result and are declared but not run, because judging them requires a model. That division is now the format's consistent answer to "deterministic locally, or not". An eval references a fixture already declared in the manifest's `examples` rather than carrying its own copy, so an example that is also an eval fixture stays honest — both break together. An eval declares assessment and must not record outcomes. Results are run evidence and live outside the immutable package, per INTENT.md and section 17. Spec: 12 rewritten with 12.1, 18 (rules 17-18), 21 (`eval` verb). Reference CLI: read_eval, validate_eval, load_example_values, run_render_checks, cmd_eval; a failed render check exits non-zero. Tests 42 -> 51. examples/pqrst-estimate/evals/quality.yaml is a real eval with four render checks and four output criteria, and it passes. Fixes a latent bug reaching a fixture exposed: coerce_value assumed every value was a command-line string, so a YAML fixture carrying a real type (include_rationale: true) crashed on .lower(). Typed values are now validated but not re-parsed. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Bjefh8NUiEiahN4JLwoSKM Assistant: claude-code Assistant-Model: opus Assistant-Process: 388925@bnt-lap001 Assistant-Session: 3507023f-e0fd-4a1e-9d90-a0d4217d1502
2026-09-06 01:38:11 +02:00
- `eval` runs the deterministic render checks of any eval declaring the
`canned-prompts/eval-rubric/v0.1` schema, and reports output criteria as
declared but not run. Unrecognized schemas are skipped, not rejected. A
failed render check exits non-zero.
CANP-WP-0002 T05: required versus observed, and typed context dependencies The answer to the capability question was conditional: keep both fields if they carry the required/observed distinction, fix the terminology if they do not. They did not. Section 9 opened with "records known requirements or observations", mixing both in one field — `models` was observational ("known to be compatible or evaluated") while `capabilities` was prescriptive ("expected from the execution environment"). Section 10 then described dependencies as what a prompt "expects". Both fields said expected, so the overlap was real ambiguity rather than redundancy, and the fix is terminology. `dependencies` now means **required**; `compatibility` means **observed**. A consumer must not refuse to run a package because its environment is absent from a compatibility list. The same capability name may legitimately appear in both: required to run at all, and separately observed to work well on particular models. `compatibility.aliases` records the same capability under other names, so a consumer can recognize a requirement its environment labels differently. Dependencies now have three kinds, separated by what the format can do about them: `prompts` it resolves by id and version; `context` names what it does not package at all; `capabilities` are what the environment must be able to do. Context entries use `name` rather than `id`, because nothing can look them up, and `description` is required because nothing else can explain an unpackaged dependency. A capability takes no version and no `requirement: generate` — it is not an artifact and cannot be fetched, pinned or generated. Capability names are free-form kebab-case, validated for shape and not membership, exactly as tags are. Section 10.1 also draws the line the format had never stated: an input is content the caller passes for one use; a context dependency is a standing fact about the environment. Spec: 9 rewritten, 9.1 and 10.1 and 10.2 new, 10 reframed, 18 (rules 19-20), 4 updated. Former 10.1/10.2 renumbered to 10.3/10.4 with cross-references. Reference CLI: validate_capabilities, validate_context_dependencies, and a `resolve` section listing required capabilities and context under "this tool cannot verify these" rather than implying it checked. Tests 51 -> 65. Also drops an invented `session-review` capability from the example package in favour of an honest `long-context` observation. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Bjefh8NUiEiahN4JLwoSKM Assistant: claude-code Assistant-Model: opus Assistant-Process: 388925@bnt-lap001 Assistant-Session: 3507023f-e0fd-4a1e-9d90-a0d4217d1502
2026-09-06 08:11:13 +02:00
- `resolve` lists required capabilities and context dependencies, which this
tool cannot verify, rather than implying it checked them.
CANP-WP-0002 T02: registry-scoped identity and namespace ownership Section 17 already defined what a registry does when the same id@version is republished; nothing defined what a *consumer* does. That is where namespace conflict actually bites — in the catalog, after installing from two registries. Identity is now registry-scoped: an id names a package within a registry, the way a path names a file within a repository. A local-first format with no signing, no federation and no central authority cannot enforce global uniqueness, and an unenforceable guarantee is worse than none — it invites consumers to conflate two packages that merely share a name. A qualified `<registry>:<id>` reference distinguishes them, and `:` is now barred from ids so the separator stays available. Installing the same id from two registries is therefore not a conflict. The catalog is namespaced by registry and keeps both. Ownership is registry policy, not package data. An optional `registry.yaml` names a registry and records namespace claims. Those claims are explicitly descriptive — a filesystem registry cannot authenticate a publisher, and `publish` says so rather than implying it checked. Keeping the claim out of packages leaves artifacts free of unverifiable assertions of authority, and means package semantics do not change when a hosted registry appears later. Spec: 3.2 (registry-scoped identity, qualified references), 17 (immutability scoped to a registry), 20.1 and 20.2 (new), 18 (registry-manifest validation), 21 (qualified references, reserved `local` name). Reference CLI: parse_reference, check_registry_name, read_registry_manifest, registry_name, namespace_policy; registry_package_path and catalog_package_path split; resolve_installed reports ambiguity and returns the source registry; iter_catalog; `add --as`; closed-namespace warning on publish. Tests 11 -> 21. The catalog layout changed. An existing catalog is detected and reported with instructions rather than failing as "package not found". Signing, trust scoring and federation remain non-goals and were not touched. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Bjefh8NUiEiahN4JLwoSKM Assistant: claude-code Assistant-Model: opus Assistant-Process: 388925@bnt-lap001 Assistant-Session: 3507023f-e0fd-4a1e-9d90-a0d4217d1502
2026-09-06 01:14:54 +02:00
- An optional `registry.yaml` names a registry and records namespace claims.
`publish` warns when a namespace is declared `closed` — it cannot
authenticate a publisher, and says so rather than implying it checked.
CANP-WP-0002 T01: input defaults, static and derived Closes the gap that made optional inputs unusable: rendering rule 5.1(4) made any unresolved placeholder an error while inputs had no `default`, so an input marked `required: false` and referenced from the template failed every render in which the caller omitted it — including the spec's own section 4 example. Section 10 already carried the derivation mechanism (`requirement: generate`, resolution deliberately undefined), so a derived default needed a binding rather than a new concept: the input's default names a declared prompt dependency. Spec: - 5.1 rewritten as "Resolution and rendering". Resolution may be non-deterministic and must report what it derived; rendering is deterministic and must not derive. A tool that handles only supplied values and static defaults is stated to be conforming. - 6.1 (new) covers both declaration forms. Reference form is preferred, with the reason stated — an inline prompt is anonymous, so unversioned, unprovenanced and un-evaluable — and validators should warn when a published package derives inline. - A derived default may declare a static fallback `value`. Without one the input stays unresolved, which is an error; derivation never silently yields empty content. - 4, 10, 18 (rules 11-13), 19 (two new MUST NOTs), 21 updated accordingly. Reference CLI: - New `resolve` verb reporting the origin of every value. - `Resolution` dataclass and `resolve_inputs`; `resolve_values` kept as a wrapper so existing callers are unaffected. - `render` refuses with a specific error naming underivable inputs rather than substituting empty text. - Tests 3 -> 11. Example package lifecycle re-verified end to end. INTENT.md is unchanged: splitting resolve from render preserves success criterion 4 (deterministic rendering) as written. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Bjefh8NUiEiahN4JLwoSKM Assistant: claude-code Assistant-Model: opus Assistant-Process: 388925@bnt-lap001 Assistant-Session: 3507023f-e0fd-4a1e-9d90-a0d4217d1502
2026-09-06 00:59:20 +02:00
- Static input defaults are applied; **derived** defaults (§ 6.1) are not. This
tool never calls a model, so a derived default is satisfied only by its
static fallback `value`. Without one, `resolve` reports the input as
unresolved and `render` refuses rather than substituting empty text.
Use this implementation to challenge the format. Replace it once real usage reveals the right architecture.