canned-prompts/reference/tests/test_canned_prompts.py

788 lines
24 KiB
Python
Raw Normal View History

from pathlib import Path
import pytest
import canned_prompts as cp
@pytest.fixture()
def package(tmp_path: Path) -> Path:
pkg = tmp_path / "pkg"
pkg.mkdir()
(pkg / "prompt.yaml").write_text(
"""\
format: canned-prompt/v0.1
id: demo/hello
name: Hello
version: 1.0.0
summary: Say hello.
template: prompt.md
inputs:
- name: person
required: true
parameters:
tone:
type: enum
values: [warm, formal]
default: warm
""",
encoding="utf-8",
)
(pkg / "prompt.md").write_text(
"Say hello to {{ person }} in a {{ tone }} tone.\n", encoding="utf-8"
)
return pkg
def test_validate_and_render(package: Path) -> None:
manifest = cp.validate_package(package)
values = cp.resolve_values(manifest, {"person": "Ada"})
rendered = cp.render_template((package / "prompt.md").read_text(), values)
assert rendered == "Say hello to Ada in a warm tone.\n"
def test_missing_required_input_fails(package: Path) -> None:
manifest = cp.validate_package(package)
with pytest.raises(cp.CannedPromptError, match="missing required input"):
cp.resolve_values(manifest, {})
def test_undeclared_placeholder_fails(package: Path) -> None:
(package / "prompt.md").write_text("{{ missing }}\n", encoding="utf-8")
with pytest.raises(cp.CannedPromptError, match="undeclared placeholders"):
cp.validate_package(package)
CANP-WP-0002 T01: input defaults, static and derived Closes the gap that made optional inputs unusable: rendering rule 5.1(4) made any unresolved placeholder an error while inputs had no `default`, so an input marked `required: false` and referenced from the template failed every render in which the caller omitted it — including the spec's own section 4 example. Section 10 already carried the derivation mechanism (`requirement: generate`, resolution deliberately undefined), so a derived default needed a binding rather than a new concept: the input's default names a declared prompt dependency. Spec: - 5.1 rewritten as "Resolution and rendering". Resolution may be non-deterministic and must report what it derived; rendering is deterministic and must not derive. A tool that handles only supplied values and static defaults is stated to be conforming. - 6.1 (new) covers both declaration forms. Reference form is preferred, with the reason stated — an inline prompt is anonymous, so unversioned, unprovenanced and un-evaluable — and validators should warn when a published package derives inline. - A derived default may declare a static fallback `value`. Without one the input stays unresolved, which is an error; derivation never silently yields empty content. - 4, 10, 18 (rules 11-13), 19 (two new MUST NOTs), 21 updated accordingly. Reference CLI: - New `resolve` verb reporting the origin of every value. - `Resolution` dataclass and `resolve_inputs`; `resolve_values` kept as a wrapper so existing callers are unaffected. - `render` refuses with a specific error naming underivable inputs rather than substituting empty text. - Tests 3 -> 11. Example package lifecycle re-verified end to end. INTENT.md is unchanged: splitting resolve from render preserves success criterion 4 (deterministic rendering) as written. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Bjefh8NUiEiahN4JLwoSKM Assistant: claude-code Assistant-Model: opus Assistant-Process: 388925@bnt-lap001 Assistant-Session: 3507023f-e0fd-4a1e-9d90-a0d4217d1502
2026-09-06 00:59:20 +02:00
def write_pkg(pkg: Path, manifest: str, template: str = "{{ greeting }}\n") -> Path:
pkg.mkdir(exist_ok=True)
(pkg / "prompt.yaml").write_text(manifest, encoding="utf-8")
(pkg / "prompt.md").write_text(template, encoding="utf-8")
return pkg
BASE = """\
format: canned-prompt/v0.1
id: demo/defaults
name: Defaults
version: 1.0.0
summary: Exercise input defaults.
template: prompt.md
"""
def test_static_default_fills_optional_input(tmp_path: Path) -> None:
pkg = write_pkg(
tmp_path / "p",
BASE
+ """\
inputs:
- name: greeting
required: false
default: "hello there"
""",
)
manifest = cp.validate_package(pkg)
resolution = cp.resolve_inputs(manifest, {})
assert resolution.values["greeting"] == "hello there"
assert resolution.origins["greeting"] == "default"
assert cp.render_template("{{ greeting }}", resolution.values) == "hello there"
def test_supplied_value_overrides_static_default(tmp_path: Path) -> None:
pkg = write_pkg(
tmp_path / "p",
BASE
+ """\
inputs:
- name: greeting
required: false
default: "hello there"
""",
)
manifest = cp.validate_package(pkg)
resolution = cp.resolve_inputs(manifest, {"greeting": "hi"})
assert resolution.values["greeting"] == "hi"
assert resolution.origins["greeting"] == "supplied"
def test_derived_default_uses_static_fallback(tmp_path: Path) -> None:
pkg = write_pkg(
tmp_path / "p",
BASE
+ """\
dependencies:
prompts:
- id: context/greeting
version: 1.0.0
requirement: generate
inputs:
- name: greeting
required: false
default:
derive: context/greeting
value: "(none)"
""",
)
manifest = cp.validate_package(pkg)
resolution = cp.resolve_inputs(manifest, {})
assert resolution.values["greeting"] == "(none)"
assert resolution.origins["greeting"] == "fallback (not derived)"
assert resolution.underivable == []
def test_derived_default_without_fallback_is_underivable(tmp_path: Path) -> None:
pkg = write_pkg(
tmp_path / "p",
BASE
+ """\
dependencies:
prompts:
- id: context/greeting
version: 1.0.0
requirement: generate
inputs:
- name: greeting
required: false
default:
derive: context/greeting
""",
)
manifest = cp.validate_package(pkg)
resolution = cp.resolve_inputs(manifest, {})
assert resolution.underivable == ["greeting"]
assert "greeting" not in resolution.values
with pytest.raises(cp.CannedPromptError, match="unresolved placeholder"):
cp.render_template("{{ greeting }}", resolution.values)
def test_inline_derive_is_valid(tmp_path: Path) -> None:
pkg = write_pkg(
tmp_path / "p",
BASE
+ """\
inputs:
- name: greeting
required: false
default:
derive:
prompt: Produce a greeting suited to the audience.
value: "(none)"
""",
)
manifest = cp.validate_package(pkg)
assert cp.resolve_inputs(manifest, {}).values["greeting"] == "(none)"
def test_default_with_required_true_fails(tmp_path: Path) -> None:
pkg = write_pkg(
tmp_path / "p",
BASE
+ """\
inputs:
- name: greeting
required: true
default: "hello"
""",
)
with pytest.raises(cp.CannedPromptError, match="required: true"):
cp.validate_package(pkg)
def test_derive_reference_must_be_declared(tmp_path: Path) -> None:
pkg = write_pkg(
tmp_path / "p",
BASE
+ """\
inputs:
- name: greeting
required: false
default:
derive: context/greeting
""",
)
with pytest.raises(cp.CannedPromptError, match="not declared in dependencies"):
cp.validate_package(pkg)
def test_inline_derive_cannot_also_reference(tmp_path: Path) -> None:
pkg = write_pkg(
tmp_path / "p",
BASE
+ """\
inputs:
- name: greeting
required: false
default:
derive:
prompt: Produce a greeting.
id: context/greeting
""",
)
with pytest.raises(cp.CannedPromptError, match="cannot also reference"):
cp.validate_package(pkg)
CANP-WP-0002 T02: registry-scoped identity and namespace ownership Section 17 already defined what a registry does when the same id@version is republished; nothing defined what a *consumer* does. That is where namespace conflict actually bites — in the catalog, after installing from two registries. Identity is now registry-scoped: an id names a package within a registry, the way a path names a file within a repository. A local-first format with no signing, no federation and no central authority cannot enforce global uniqueness, and an unenforceable guarantee is worse than none — it invites consumers to conflate two packages that merely share a name. A qualified `<registry>:<id>` reference distinguishes them, and `:` is now barred from ids so the separator stays available. Installing the same id from two registries is therefore not a conflict. The catalog is namespaced by registry and keeps both. Ownership is registry policy, not package data. An optional `registry.yaml` names a registry and records namespace claims. Those claims are explicitly descriptive — a filesystem registry cannot authenticate a publisher, and `publish` says so rather than implying it checked. Keeping the claim out of packages leaves artifacts free of unverifiable assertions of authority, and means package semantics do not change when a hosted registry appears later. Spec: 3.2 (registry-scoped identity, qualified references), 17 (immutability scoped to a registry), 20.1 and 20.2 (new), 18 (registry-manifest validation), 21 (qualified references, reserved `local` name). Reference CLI: parse_reference, check_registry_name, read_registry_manifest, registry_name, namespace_policy; registry_package_path and catalog_package_path split; resolve_installed reports ambiguity and returns the source registry; iter_catalog; `add --as`; closed-namespace warning on publish. Tests 11 -> 21. The catalog layout changed. An existing catalog is detected and reported with instructions rather than failing as "package not found". Signing, trust scoring and federation remain non-goals and were not touched. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Bjefh8NUiEiahN4JLwoSKM Assistant: claude-code Assistant-Model: opus Assistant-Process: 388925@bnt-lap001 Assistant-Session: 3507023f-e0fd-4a1e-9d90-a0d4217d1502
2026-09-06 01:14:54 +02:00
MINIMAL = """\
format: canned-prompt/v0.1
id: practice/thing
name: Thing
version: 1.0.0
summary: A thing.
template: prompt.md
inputs:
- name: greeting
required: false
default: hi
"""
def test_parse_reference() -> None:
assert cp.parse_reference("practice/thing") == (None, "practice/thing")
assert cp.parse_reference("house:practice/thing") == ("house", "practice/thing")
with pytest.raises(cp.CannedPromptError, match="malformed reference"):
cp.parse_reference("house:")
def test_registry_name_falls_back_to_basename(tmp_path: Path) -> None:
registry = tmp_path / "upstream"
registry.mkdir()
assert cp.registry_name(registry) == "upstream"
def test_registry_name_from_manifest(tmp_path: Path) -> None:
registry = tmp_path / "some-dir"
registry.mkdir()
(registry / "registry.yaml").write_text(
"format: canned-prompt-registry/v0.1\nname: house\n", encoding="utf-8"
)
assert cp.registry_name(registry) == "house"
def test_registry_manifest_rejects_bad_format(tmp_path: Path) -> None:
registry = tmp_path / "r"
registry.mkdir()
(registry / "registry.yaml").write_text(
"format: something-else\nname: house\n", encoding="utf-8"
)
with pytest.raises(cp.CannedPromptError, match="unsupported registry format"):
cp.read_registry_manifest(registry)
def test_registry_manifest_rejects_bad_policy(tmp_path: Path) -> None:
registry = tmp_path / "r"
registry.mkdir()
(registry / "registry.yaml").write_text(
"format: canned-prompt-registry/v0.1\n"
"name: house\n"
"namespaces:\n practice:\n policy: maybe\n",
encoding="utf-8",
)
with pytest.raises(cp.CannedPromptError, match="policy must be"):
cp.read_registry_manifest(registry)
def test_namespace_policy(tmp_path: Path) -> None:
registry = tmp_path / "r"
registry.mkdir()
(registry / "registry.yaml").write_text(
"format: canned-prompt-registry/v0.1\n"
"name: house\n"
"namespaces:\n practice:\n owner: Ada\n policy: closed\n",
encoding="utf-8",
)
assert cp.namespace_policy(registry, "practice/thing")[0] == "closed"
assert cp.namespace_policy(registry, "scratch/thing")[0] == "open"
def install_into(catalog: Path, registry: str) -> None:
"""Place a package in the catalog under a given registry name."""
dst = cp.catalog_package_path(catalog, registry, "practice/thing", "1.0.0")
dst.mkdir(parents=True)
(dst / "prompt.yaml").write_text(MINIMAL, encoding="utf-8")
(dst / "prompt.md").write_text("{{ greeting }}\n", encoding="utf-8")
def test_same_id_from_two_registries_coexists(tmp_path: Path) -> None:
catalog = tmp_path / "catalog"
install_into(catalog, "house")
install_into(catalog, "upstream")
assert cp.catalog_registries(catalog) == ["house", "upstream"]
package_dir, registry = cp.resolve_installed(catalog, "house:practice/thing", None)
assert registry == "house"
assert package_dir.is_dir()
def test_bare_id_in_two_registries_is_ambiguous(tmp_path: Path) -> None:
catalog = tmp_path / "catalog"
install_into(catalog, "house")
install_into(catalog, "upstream")
with pytest.raises(cp.CannedPromptError, match="more than one registry"):
cp.resolve_installed(catalog, "practice/thing", None)
def test_bare_id_in_one_registry_resolves(tmp_path: Path) -> None:
catalog = tmp_path / "catalog"
install_into(catalog, "house")
_, registry = cp.resolve_installed(catalog, "practice/thing", None)
assert registry == "house"
def test_legacy_catalog_layout_is_reported(tmp_path: Path) -> None:
catalog = tmp_path / "catalog"
legacy = catalog / "practice" / "thing" / "1.0.0"
legacy.mkdir(parents=True)
(legacy / "prompt.yaml").write_text(MINIMAL, encoding="utf-8")
(legacy / "prompt.md").write_text("{{ greeting }}\n", encoding="utf-8")
with pytest.raises(cp.CannedPromptError, match="pre-registry-scoped"):
cp.resolve_installed(catalog, "practice/thing", None)
CANP-WP-0002 T03: composition by reference, two kinds T01 had already delivered half of composition without naming it: a derived default binds an input to a prompt dependency, which is transclusion — run package B, use its output. What was missing was the deterministic half. `include` inlines another package's rendered template as text. No model is involved, so the reference CLI can actually perform it, and a shared preamble, rubric or style block becomes a versioned package instead of copied text. This is the concrete way to honor INTENT principle 9 without any runtime. `derive` stays as it was. Both are input defaults, so composition reuses the resolution machinery rather than adding a second one. No template inheritance. Four of this repo's own documents argue against it: INTENT principle 3 (hidden context defeats reuse), section 19's "make package contents visible before execution", section 17's requirement that behavior changes produce a new version, and the non-goal on range resolution. Version selectors: an exact pin is the expected form, with `any`, `newest` and `>= X.Y.Z` as explicit opt-ins so looseness is written rather than implied by absence. Selectors are evaluated per dependency against what is available — no solver, no cross-dependency constraint satisfaction — which is what keeps them outside the range-resolution non-goal, and the spec says so. Also defines `type` (template | fragment), which appeared once in the section 4 manifest surface and was specified nowhere. Spec: 3.2 (type), 5.1 (inclusion resolution rule, renumbered), 6.1 (included default), 10.1 and 10.2 (new), 18 (rules 14-16), 21. Reference CLI: validate_version_selector, select_version, prompt_dependencies replacing prompt_dependency_ids, check_composition_reference, CatalogComposer with cycle detection, and resolve_inputs gaining composer= and inherited=. Tests 21 -> 42. Examples: house-style is a real fragment package; pqrst-estimate composes it and is bumped 0.1.0 -> 0.2.0 per section 17. Fixes an ordering bug found while testing: inputs resolved before parameters, so an included package could not see the including package's parameters and silently fell back to its own defaults — the fragment rendered tone=neutral where the including package said blunt. Parameters now resolve first; the report still lists inputs first. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Bjefh8NUiEiahN4JLwoSKM Assistant: claude-code Assistant-Model: opus Assistant-Process: 388925@bnt-lap001 Assistant-Session: 3507023f-e0fd-4a1e-9d90-a0d4217d1502
2026-09-06 01:31:58 +02:00
# --- version selectors (§ 10.1) ---
AVAILABLE = ["2.1.0", "2.0.0", "1.5.0", "1.0.0"]
@pytest.mark.parametrize(
"available,selector,expected",
[
(AVAILABLE, None, "2.1.0"),
(AVAILABLE, "any", "2.1.0"),
(AVAILABLE, "newest", "2.1.0"),
(AVAILABLE, "1.5.0", "1.5.0"),
(AVAILABLE, ">=2.0.0", "2.1.0"),
(AVAILABLE, "9.9.9", None),
(AVAILABLE, ">=9.9.9", None),
(["1.5.0", "1.0.0"], ">=1.2.0", "1.5.0"),
([], "newest", None),
],
)
def test_select_version(available, selector, expected) -> None:
assert cp.select_version(available, selector) == expected
@pytest.mark.parametrize("bad", ["^1.0.0", "~1.2", "1.x", ">=nope", "", "latest"])
def test_range_syntax_is_rejected(bad) -> None:
with pytest.raises(cp.CannedPromptError):
cp.validate_version_selector(bad, "dep")
# --- composition (§ 10.2) ---
FRAGMENT = """\
format: canned-prompt/v0.1
id: style/house
name: House Style
version: 1.0.0
summary: Shared style block.
type: fragment
template: prompt.md
parameters:
tone:
type: enum
values: [neutral, blunt]
default: neutral
"""
COMPOSER = """\
format: canned-prompt/v0.1
id: review/change
name: Change Review
version: 1.0.0
summary: Composes the shared style block.
template: prompt.md
dependencies:
prompts:
- id: style/house
version: 1.0.0
inputs:
- name: house_style
required: false
default:
include: style/house
parameters:
tone:
type: enum
values: [neutral, blunt]
default: blunt
"""
def place(catalog: Path, registry: str, package_id: str, version: str,
manifest: str, template: str) -> None:
dst = cp.catalog_package_path(catalog, registry, package_id, version)
dst.mkdir(parents=True)
(dst / "prompt.yaml").write_text(manifest, encoding="utf-8")
(dst / "prompt.md").write_text(template, encoding="utf-8")
@pytest.fixture()
def composed(tmp_path: Path) -> Path:
catalog = tmp_path / "catalog"
place(catalog, "local", "style/house", "1.0.0", FRAGMENT, "Tone is {{ tone }}.")
place(catalog, "local", "review/change", "1.0.0", COMPOSER, "S: {{ house_style }}")
return catalog
def test_include_inlines_rendered_template(composed: Path) -> None:
package_dir, _ = cp.resolve_installed(composed, "review/change", None)
manifest = cp.validate_package(package_dir)
resolution = cp.resolve_inputs(manifest, {}, composer=cp.CatalogComposer(composed))
assert resolution.values["house_style"] == "Tone is blunt."
assert resolution.origins["house_style"] == "included from style/house"
def test_include_inherits_outer_parameters(composed: Path) -> None:
"""The fragment's own default is neutral; the including package says blunt."""
package_dir, _ = cp.resolve_installed(composed, "review/change", None)
manifest = cp.validate_package(package_dir)
resolution = cp.resolve_inputs(
manifest, {"tone": "neutral"}, composer=cp.CatalogComposer(composed)
)
assert resolution.values["house_style"] == "Tone is neutral."
def test_include_without_composer_falls_back_to_unresolved(composed: Path) -> None:
package_dir, _ = cp.resolve_installed(composed, "review/change", None)
manifest = cp.validate_package(package_dir)
resolution = cp.resolve_inputs(manifest, {})
assert resolution.underivable == ["house_style"]
def test_inclusion_cycle_is_detected(tmp_path: Path) -> None:
catalog = tmp_path / "catalog"
for this, other in (("cyc/a", "cyc/b"), ("cyc/b", "cyc/a")):
manifest = (
"format: canned-prompt/v0.1\n"
f"id: {this}\nname: X\nversion: 1.0.0\nsummary: s\ntemplate: prompt.md\n"
f"dependencies:\n prompts:\n - id: {other}\n version: 1.0.0\n"
f"inputs:\n - name: other\n required: false\n"
f" default:\n include: {other}\n"
)
place(catalog, "local", this, "1.0.0", manifest, "{{ other }}")
package_dir, _ = cp.resolve_installed(catalog, "cyc/a", None)
manifest = cp.validate_package(package_dir)
with pytest.raises(cp.CannedPromptError, match="inclusion cycle"):
cp.resolve_inputs(manifest, {}, composer=cp.CatalogComposer(catalog))
def test_include_and_derive_together_is_rejected(tmp_path: Path) -> None:
pkg = write_pkg(
tmp_path / "p",
BASE
+ """\
dependencies:
prompts:
- id: style/house
version: 1.0.0
inputs:
- name: greeting
required: false
default:
include: style/house
derive: style/house
""",
)
with pytest.raises(cp.CannedPromptError, match="at most one of"):
cp.validate_package(pkg)
def test_composed_dependency_must_declare_a_version(tmp_path: Path) -> None:
pkg = write_pkg(
tmp_path / "p",
BASE
+ """\
dependencies:
prompts:
- id: style/house
inputs:
- name: greeting
required: false
default:
include: style/house
""",
)
with pytest.raises(cp.CannedPromptError, match="must declare a version"):
cp.validate_package(pkg)
CANP-WP-0002 T04: eval schema, render checks and output criteria `evals/` was a reserved path holding unvalidated blobs: section 12 named the directory and gave an illustrative snippet, but nothing was specified, so no tool could act on an eval file. Every eval file must now declare a `schema`, and CPF defines exactly one — `canned-prompts/eval-rubric/v0.1`. Unrecognized schemas stay legal and are skipped rather than rejected, so the format gains something actionable without becoming an evaluation language, which remains a non-goal. The schema splits along the same seam as T01 and T03. Render checks (`contains`, `not_contains`, `resolves_all`) assert properties of the rendered prompt text, need no model, and are therefore run by the reference CLI. Output criteria describe a good result and are declared but not run, because judging them requires a model. That division is now the format's consistent answer to "deterministic locally, or not". An eval references a fixture already declared in the manifest's `examples` rather than carrying its own copy, so an example that is also an eval fixture stays honest — both break together. An eval declares assessment and must not record outcomes. Results are run evidence and live outside the immutable package, per INTENT.md and section 17. Spec: 12 rewritten with 12.1, 18 (rules 17-18), 21 (`eval` verb). Reference CLI: read_eval, validate_eval, load_example_values, run_render_checks, cmd_eval; a failed render check exits non-zero. Tests 42 -> 51. examples/pqrst-estimate/evals/quality.yaml is a real eval with four render checks and four output criteria, and it passes. Fixes a latent bug reaching a fixture exposed: coerce_value assumed every value was a command-line string, so a YAML fixture carrying a real type (include_rationale: true) crashed on .lower(). Typed values are now validated but not re-parsed. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Bjefh8NUiEiahN4JLwoSKM Assistant: claude-code Assistant-Model: opus Assistant-Process: 388925@bnt-lap001 Assistant-Session: 3507023f-e0fd-4a1e-9d90-a0d4217d1502
2026-09-06 01:38:11 +02:00
# --- evals (§ 12) ---
EVAL_BASE = """\
format: canned-prompt/v0.1
id: demo/evaluated
name: Evaluated
version: 1.0.0
summary: Exercise eval files.
template: prompt.md
examples:
- examples/basic.yaml
evals:
- evals/quality.yaml
"""
def write_evaluated(pkg: Path, eval_body: str, template: str = "Sum to 100%. {{ topic }}\n") -> Path:
pkg.mkdir(exist_ok=True)
(pkg / "prompt.yaml").write_text(
EVAL_BASE + "inputs:\n - name: topic\n required: false\n default: cats\n",
encoding="utf-8",
)
(pkg / "prompt.md").write_text(template, encoding="utf-8")
(pkg / "examples").mkdir(exist_ok=True)
(pkg / "examples" / "basic.yaml").write_text(
"name: basic\nvalues:\n topic: dogs\n", encoding="utf-8"
)
(pkg / "evals").mkdir(exist_ok=True)
(pkg / "evals" / "quality.yaml").write_text(eval_body, encoding="utf-8")
return pkg
RUBRIC = """\
schema: canned-prompts/eval-rubric/v0.1
name: quality
example: examples/basic.yaml
render:
- contains: "Sum to 100%"
- not_contains: "{{"
- resolves_all: true
output:
criteria:
- Answers the question.
"""
def test_valid_eval_passes_validation(tmp_path: Path) -> None:
pkg = write_evaluated(tmp_path / "p", RUBRIC)
assert cp.validate_package(pkg)["id"] == "demo/evaluated"
def test_eval_must_declare_a_schema(tmp_path: Path) -> None:
pkg = write_evaluated(tmp_path / "p", "name: quality\nrender: []\n")
with pytest.raises(cp.CannedPromptError, match="must declare a schema"):
cp.validate_package(pkg)
def test_unknown_eval_schema_is_ignored(tmp_path: Path) -> None:
pkg = write_evaluated(tmp_path / "p", "schema: someone/else/v1\nwhatever: true\n")
assert cp.validate_package(pkg)["id"] == "demo/evaluated"
def test_eval_asserting_nothing_is_rejected(tmp_path: Path) -> None:
pkg = write_evaluated(
tmp_path / "p", "schema: canned-prompts/eval-rubric/v0.1\nname: empty\n"
)
with pytest.raises(cp.CannedPromptError, match="asserts nothing"):
cp.validate_package(pkg)
def test_unknown_render_check_is_rejected(tmp_path: Path) -> None:
pkg = write_evaluated(
tmp_path / "p",
"schema: canned-prompts/eval-rubric/v0.1\nname: q\nrender:\n - matches: 'x.*'\n",
)
with pytest.raises(cp.CannedPromptError, match="unknown render check"):
cp.validate_package(pkg)
def test_eval_example_must_be_declared(tmp_path: Path) -> None:
pkg = write_evaluated(
tmp_path / "p",
"schema: canned-prompts/eval-rubric/v0.1\nname: q\n"
"example: examples/missing.yaml\nrender:\n - contains: x\n",
)
with pytest.raises(cp.CannedPromptError, match="not declared in the manifest"):
cp.validate_package(pkg)
def test_render_checks_evaluate(tmp_path: Path) -> None:
resolution = cp.Resolution(values={"topic": "dogs"}, origins={"topic": "supplied"})
checks = [{"contains": "dogs"}, {"contains": "cats"}, {"not_contains": "cats"}]
outcomes = cp.run_render_checks("about dogs", resolution, checks)
assert [ok for ok, _ in outcomes] == [True, False, True]
def test_resolves_all_reports_unresolved_names() -> None:
resolution = cp.Resolution(
values={"a": 1}, origins={"a": "supplied", "b": "unresolved (no default)"}
)
outcomes = cp.run_render_checks("text", resolution, [{"resolves_all": True}])
assert outcomes[0][0] is False
assert "b" in outcomes[0][1]
def test_typed_fixture_values_are_not_reparsed(tmp_path: Path) -> None:
"""A YAML fixture carries real types; only CLI strings need parsing."""
manifest = {"parameters": {"flag": {"type": "boolean", "default": False}}}
assert cp.resolve_inputs(manifest, {"flag": True}).values["flag"] is True
CANP-WP-0002 T05: required versus observed, and typed context dependencies The answer to the capability question was conditional: keep both fields if they carry the required/observed distinction, fix the terminology if they do not. They did not. Section 9 opened with "records known requirements or observations", mixing both in one field — `models` was observational ("known to be compatible or evaluated") while `capabilities` was prescriptive ("expected from the execution environment"). Section 10 then described dependencies as what a prompt "expects". Both fields said expected, so the overlap was real ambiguity rather than redundancy, and the fix is terminology. `dependencies` now means **required**; `compatibility` means **observed**. A consumer must not refuse to run a package because its environment is absent from a compatibility list. The same capability name may legitimately appear in both: required to run at all, and separately observed to work well on particular models. `compatibility.aliases` records the same capability under other names, so a consumer can recognize a requirement its environment labels differently. Dependencies now have three kinds, separated by what the format can do about them: `prompts` it resolves by id and version; `context` names what it does not package at all; `capabilities` are what the environment must be able to do. Context entries use `name` rather than `id`, because nothing can look them up, and `description` is required because nothing else can explain an unpackaged dependency. A capability takes no version and no `requirement: generate` — it is not an artifact and cannot be fetched, pinned or generated. Capability names are free-form kebab-case, validated for shape and not membership, exactly as tags are. Section 10.1 also draws the line the format had never stated: an input is content the caller passes for one use; a context dependency is a standing fact about the environment. Spec: 9 rewritten, 9.1 and 10.1 and 10.2 new, 10 reframed, 18 (rules 19-20), 4 updated. Former 10.1/10.2 renumbered to 10.3/10.4 with cross-references. Reference CLI: validate_capabilities, validate_context_dependencies, and a `resolve` section listing required capabilities and context under "this tool cannot verify these" rather than implying it checked. Tests 51 -> 65. Also drops an invented `session-review` capability from the example package in favour of an honest `long-context` observation. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Bjefh8NUiEiahN4JLwoSKM Assistant: claude-code Assistant-Model: opus Assistant-Process: 388925@bnt-lap001 Assistant-Session: 3507023f-e0fd-4a1e-9d90-a0d4217d1502
2026-09-06 08:11:13 +02:00
# --- context and capability dependencies (§ 10.1, § 10.2) ---
def ctx_pkg(tmp_path: Path, body: str) -> Path:
return write_pkg(tmp_path / "p", BASE + body, "hello\n")
def test_context_dependency_requires_name_and_description(tmp_path: Path) -> None:
pkg = ctx_pkg(
tmp_path,
"dependencies:\n context:\n - name: repository-tree\n"
" description: A listing of the repository.\n",
)
assert cp.validate_package(pkg)["id"] == "demo/defaults"
def test_context_dependency_without_description_is_rejected(tmp_path: Path) -> None:
pkg = ctx_pkg(tmp_path, "dependencies:\n context:\n - name: repository-tree\n")
with pytest.raises(cp.CannedPromptError, match="must declare a description"):
cp.validate_package(pkg)
def test_context_dependency_using_id_is_rejected(tmp_path: Path) -> None:
"""Context entries use `name`; `id` would imply resolvable package identity."""
pkg = ctx_pkg(
tmp_path,
"dependencies:\n context:\n - id: policy/security\n"
" description: A policy.\n",
)
with pytest.raises(cp.CannedPromptError, match="not 'id'"):
cp.validate_package(pkg)
def test_context_requirement_generate_is_rejected(tmp_path: Path) -> None:
pkg = ctx_pkg(
tmp_path,
"dependencies:\n context:\n - name: repository-tree\n"
" description: A listing.\n requirement: generate\n",
)
with pytest.raises(cp.CannedPromptError, match="required' or 'optional"):
cp.validate_package(pkg)
@pytest.mark.parametrize("name", ["web-search", "code-analysis", "vision", "a1-b2"])
def test_valid_capability_names(name) -> None:
assert cp.validate_capabilities([name], "dependencies.capabilities") == [name]
@pytest.mark.parametrize("name", ["Web-Search", "web_search", "-lead", "trail-", ""])
def test_invalid_capability_names(name) -> None:
with pytest.raises(cp.CannedPromptError, match="kebab-case"):
cp.validate_capabilities([name], "dependencies.capabilities")
def test_capability_may_be_both_required_and_observed(tmp_path: Path) -> None:
"""§ 9.1: required to run at all, and separately observed to work well."""
pkg = ctx_pkg(
tmp_path,
"dependencies:\n capabilities:\n - web-search\n"
"compatibility:\n capabilities:\n - web-search\n - long-context\n",
)
manifest = cp.validate_package(pkg)
assert manifest["dependencies"]["capabilities"] == ["web-search"]
assert "long-context" in manifest["compatibility"]["capabilities"]
CANP-WP-0002 T07: semver precedence and strict packaging Two defects from the original review of the seed. Prerelease ordering was worse than first recorded. parse_semver returned (major, minor, patch, raw_string), so 1.0.0-rc1 and 1.0.0 tied on the numeric fields and then compared as strings — "1.0.0-rc1" > "1.0.0". A release candidate therefore shadowed its own release for `newest` and for `>=`, not just for the no-version case. parse_semver now returns a SemVer section 11 precedence key: numeric fields, a release/prerelease rank, then dot-separated prerelease identifiers with numeric ones compared numerically. Build metadata is ignored. Beyond ordering, prereleases are excluded from `any`, `newest` and `>=` entirely; only an exact pin selects one, so publishing a release candidate never changes what existing consumers resolve to. A package holding only prereleases now says so rather than reporting a bare not-found. Packaging copied the whole source directory, so a stray .git, virtualenv or scratch file landed in the catalog and registry. Section 2 already required otherwise — tools MUST ignore unknown non-reserved files unless a manifest field references them — so this is conformance rather than a new rule. What is new is that omissions are reported instead of silent: not packaged (not a reserved path, not referenced by the manifest): .git/, .venv/, notes.txt LICENSE joins the reserved paths. Strict packaging would otherwise drop a package's license text while faithfully copying its `license` field, which contradicts section 14's instruction to surface licensing on publish and install. Spec: 2 (LICENSE, packaging obligation, reporting), 17.1 new, 10.3 note. Reference CLI: parse_semver rewritten with is_prerelease; select_version and pick_version updated; copy_package and report_skipped replace copy_immutable. Tests 65 -> 78. examples/pqrst-estimate carries a LICENSE and a license field, exercising the new reserved path. Also fixes a leaked loop variable in package_members that would have reported a bad `template` path as an `evals` error. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Bjefh8NUiEiahN4JLwoSKM Assistant: claude-code Assistant-Model: opus Assistant-Process: 388925@bnt-lap001 Assistant-Session: 3507023f-e0fd-4a1e-9d90-a0d4217d1502
2026-09-06 09:32:42 +02:00
# --- semver precedence and prereleases (§ 17.1) ---
def test_release_outranks_its_prerelease() -> None:
assert cp.parse_semver("1.0.0") > cp.parse_semver("1.0.0-rc1")
def test_prerelease_ordering_follows_semver() -> None:
ordered = sorted(
["1.0.0", "1.0.0-rc.2", "1.0.0-rc.10", "1.0.0-alpha", "0.9.0"],
key=cp.parse_semver,
reverse=True,
)
assert ordered == ["1.0.0", "1.0.0-rc.10", "1.0.0-rc.2", "1.0.0-alpha", "0.9.0"]
def test_build_metadata_is_ignored_for_precedence() -> None:
assert cp.parse_semver("1.0.0+build.1") == cp.parse_semver("1.0.0+build.2")
@pytest.mark.parametrize("selector", ["newest", "any", ">=1.0.0", None])
def test_prerelease_is_not_selected_implicitly(selector) -> None:
available = ["1.1.0-rc1", "1.0.0"]
assert cp.select_version(available, selector) == "1.0.0"
def test_exact_pin_selects_a_prerelease() -> None:
assert cp.select_version(["1.1.0-rc1", "1.0.0"], "1.1.0-rc1") == "1.1.0-rc1"
def test_only_prereleases_available_selects_nothing() -> None:
assert cp.select_version(["1.0.0-rc1"], "newest") is None
def test_prerelease_only_package_reports_why(tmp_path: Path) -> None:
catalog = tmp_path / "catalog"
place(catalog, "local", "practice/thing", "1.0.0-rc1", MINIMAL.replace(
"version: 1.0.0", "version: 1.0.0-rc1"), "{{ greeting }}\n")
with pytest.raises(cp.CannedPromptError, match="only prerelease versions"):
cp.resolve_installed(catalog, "practice/thing", None)
# --- strict packaging (§ 2) ---
def test_only_reserved_and_referenced_paths_are_packaged(tmp_path: Path) -> None:
src = write_evaluated(tmp_path / "src", RUBRIC)
(src / "LICENSE").write_text("MIT\n", encoding="utf-8")
(src / "notes.txt").write_text("scratch\n", encoding="utf-8")
(src / ".git").mkdir()
(src / ".git" / "config").write_text("junk\n", encoding="utf-8")
manifest = cp.validate_package(src)
dst = tmp_path / "out"
skipped = cp.copy_package(src, dst, manifest, "package")
packaged = sorted(p.relative_to(dst).as_posix() for p in dst.rglob("*") if p.is_file())
assert packaged == [
"LICENSE",
"evals/quality.yaml",
"examples/basic.yaml",
"prompt.md",
"prompt.yaml",
]
assert skipped == [".git/", "notes.txt"]
def test_packaged_copy_is_still_valid(tmp_path: Path) -> None:
src = write_evaluated(tmp_path / "src", RUBRIC)
manifest = cp.validate_package(src)
dst = tmp_path / "out"
cp.copy_package(src, dst, manifest, "package")
assert cp.validate_package(dst)["id"] == "demo/evaluated"
def test_copy_refuses_to_overwrite(tmp_path: Path) -> None:
src = write_evaluated(tmp_path / "src", RUBRIC)
manifest = cp.validate_package(src)
dst = tmp_path / "out"
cp.copy_package(src, dst, manifest, "package")
with pytest.raises(cp.CannedPromptError, match="already exists"):
cp.copy_package(src, dst, manifest, "package")
CANP-WP-0002 T06: revision v0.2, and section 23 rewritten Closes the workplan. The format becomes `canned-prompt/v0.2`, and packages declaring v0.1 remain valid — everything added across T01-T05 is additive, so a v0.1 package means exactly what it always meant. That is the MINOR case section 17 itself describes. The spec file loses its version suffix: CannedPromptFormat-v0.1.md becomes CannedPromptFormat.md, with the revision stated inside. One stable path that never breaks a link, and no rename per revision; the version belongs in the `format` string where tools actually read it. Section 23 is rewritten into three parts rather than the planned two. "Settled since v0.1" tables the five resolved questions against where each rule now lives. "Still deferred" carries the eight unpromoted items plus pattern-matching render checks. "Decided against" holds template inheritance alone, because calling it deferred would misdescribe it — reopening it means overturning a decision and answering four recorded objections, not filling a gap. Section 23 also names the two habits the five decisions turned out to share, so later revisions follow them rather than rediscover them: separate the deterministic half from the rest, and a package never asserts what it cannot back. The eval-rubric and registry-manifest schemas keep their own v0.1. They are new in this revision and sit on their own version lines. Reference CLI: ACCEPTED_FORMATS; an unknown revision is rejected naming what is accepted. Tests 78 -> 81. Example packages declare v0.2 and are bumped 0.1.0 -> 0.1.1 and 0.2.0 -> 0.2.1 as section 17 PATCH — metadata corrections with behavior unchanged. Also refreshes section 22's worked example, which had drifted: it showed pqrst-estimate at 0.1.0 with no composition, contradicting the package actually in the repo. It now mirrors the real package and doubles as a composition illustration. CANP-WP-0002 is finished. CANP-WP-0003 carries forward the one residual: the default registry's basename-derived name reads as `registry:`. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Bjefh8NUiEiahN4JLwoSKM Assistant: claude-code Assistant-Model: opus Assistant-Process: 388925@bnt-lap001 Assistant-Session: 3507023f-e0fd-4a1e-9d90-a0d4217d1502
2026-09-06 14:22:45 +02:00
# --- format revisions (§ 3.2) ---
@pytest.mark.parametrize("declared", ["canned-prompt/v0.2", "canned-prompt/v0.1"])
def test_both_format_revisions_are_accepted(tmp_path: Path, declared) -> None:
"""v0.2 is additive, so a v0.1 package means exactly what it always meant."""
pkg = write_pkg(
tmp_path / "p",
BASE.replace("canned-prompt/v0.1", declared)
+ "inputs:\n - name: greeting\n required: false\n default: hi\n",
)
assert cp.validate_package(pkg)["format"] == declared
def test_unknown_format_is_rejected(tmp_path: Path) -> None:
pkg = write_pkg(
tmp_path / "p",
BASE.replace("canned-prompt/v0.1", "canned-prompt/v9.9")
+ "inputs:\n - name: greeting\n required: false\n default: hi\n",
)
with pytest.raises(cp.CannedPromptError, match="unsupported format"):
cp.validate_package(pkg)