`evals/` was a reserved path holding unvalidated blobs: section 12 named the directory and gave an illustrative snippet, but nothing was specified, so no tool could act on an eval file. Every eval file must now declare a `schema`, and CPF defines exactly one — `canned-prompts/eval-rubric/v0.1`. Unrecognized schemas stay legal and are skipped rather than rejected, so the format gains something actionable without becoming an evaluation language, which remains a non-goal. The schema splits along the same seam as T01 and T03. Render checks (`contains`, `not_contains`, `resolves_all`) assert properties of the rendered prompt text, need no model, and are therefore run by the reference CLI. Output criteria describe a good result and are declared but not run, because judging them requires a model. That division is now the format's consistent answer to "deterministic locally, or not". An eval references a fixture already declared in the manifest's `examples` rather than carrying its own copy, so an example that is also an eval fixture stays honest — both break together. An eval declares assessment and must not record outcomes. Results are run evidence and live outside the immutable package, per INTENT.md and section 17. Spec: 12 rewritten with 12.1, 18 (rules 17-18), 21 (`eval` verb). Reference CLI: read_eval, validate_eval, load_example_values, run_render_checks, cmd_eval; a failed render check exits non-zero. Tests 42 -> 51. examples/pqrst-estimate/evals/quality.yaml is a real eval with four render checks and four output criteria, and it passes. Fixes a latent bug reaching a fixture exposed: coerce_value assumed every value was a command-line string, so a YAML fixture carrying a real type (include_rationale: true) crashed on .lower(). Typed values are now validated but not re-parsed. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Bjefh8NUiEiahN4JLwoSKM Assistant: claude-code Assistant-Model: opus Assistant-Process: 388925@bnt-lap001 Assistant-Session: 3507023f-e0fd-4a1e-9d90-a0d4217d1502
361 lines
17 KiB
Markdown
361 lines
17 KiB
Markdown
---
|
||
id: CANP-WP-0002
|
||
type: workplan
|
||
title: "Resolve CPF v0.1 open questions promoted for v0.2"
|
||
domain: agents
|
||
repo: canned-prompts
|
||
status: proposed
|
||
owner: codex
|
||
topic_slug: practice
|
||
created: "2026-09-06"
|
||
updated: "2026-09-06"
|
||
reviewed_at: "2026-09-06"
|
||
reviewed_by: "claude"
|
||
context_paths:
|
||
- "INTENT.md"
|
||
- "CannedPromptFormat-v0.1.md"
|
||
- "reference/canned_prompts.py"
|
||
- "examples/pqrst-estimate/"
|
||
state_hub_workstream_id: "c87c8e27-8b11-5306-b687-86779ef3f1be"
|
||
---
|
||
|
||
# Resolve CPF v0.1 open questions promoted for v0.2
|
||
|
||
`CannedPromptFormat-v0.1.md` § 23 currently lists twelve items as "experience
|
||
should determine". A review of the seed on 2026-09-06 (spec + reference CLI +
|
||
`examples/pqrst-estimate`; 3/3 tests pass, full local lifecycle smokes clean)
|
||
plus an operator interview promoted five of them from "wait and see" to
|
||
"decide, with a stated leaning". This workplan carries those decisions.
|
||
|
||
Unpromoted § 23 items stay deferred and unchanged: content macros, cryptographic
|
||
integrity/signing, model capability vocabularies, run manifests and evidence
|
||
formats, deterministic compilation manifests, trust/reputation signals,
|
||
federated discovery, richer template syntax.
|
||
|
||
**Guardrail for every task below:** `INTENT.md` § Deliberate boundary. The
|
||
format may *declare* a requirement; it must not specify the resolver, runtime,
|
||
or execution engine that satisfies it.
|
||
|
||
## Optional inputs need defaults
|
||
|
||
```task
|
||
id: CANP-WP-0002-T01
|
||
status: done
|
||
priority: high
|
||
state_hub_task_id: "e49bf036-0aee-5bb2-abbd-3d3f683df298"
|
||
```
|
||
|
||
**Gap found in review.** Rendering rule § 5.1(4) makes any unresolved
|
||
placeholder an error, and § 6 gives `inputs` no `default` field. An input
|
||
declared `required: false` and referenced from the template therefore makes
|
||
`render` fail unless a value is supplied — optional inputs are effectively
|
||
unusable. The spec's own § 4 manifest example hits this with
|
||
`repository_context`. `reference/canned_prompts.py` is conformant here; the
|
||
defect is in the spec.
|
||
|
||
**Decision (operator, 2026-09-06):** add a `default` field to `inputs`, and
|
||
allow that default to be either
|
||
|
||
- a **static** value, or
|
||
- a **derived** default — a prompt that produces the value from available
|
||
context.
|
||
|
||
Unresolved-with-no-default remains an error.
|
||
|
||
**Design note.** § 10 already carried the derivation mechanism:
|
||
`dependencies` accepts `requirement: generate`, defined as "a resolver MAY
|
||
satisfy a missing dependency by invoking an appropriate generation process",
|
||
with resolution left undefined. A derived default therefore did not need a new
|
||
concept, only a binding — the input's default names the dependency, and the
|
||
dependency records what may be generated and at which version.
|
||
|
||
**Follow-up decisions (operator, 2026-09-06):**
|
||
|
||
- *Both declaration forms are allowed.* A derived default may reference a
|
||
declared prompt dependency, or carry an inline `prompt`. The reference form
|
||
is preferred and the spec says why (an inline prompt is anonymous —
|
||
unversioned, unprovenanced, un-evaluable), and a validator SHOULD warn when a
|
||
**published** package derives inline. Inline is for local and draft packages.
|
||
- *Error unless static fallback.* A derived default MAY declare a static
|
||
`value`. A consumer that does not derive uses it; with no fallback the input
|
||
stays unresolved, which is an error. Derivation never silently yields empty
|
||
content.
|
||
- *Split render from resolve.* Resolution decides values and may be
|
||
non-deterministic; rendering substitutes and always is. Rendering MUST NOT
|
||
derive. This preserves INTENT success criterion 4 verbatim — no change to
|
||
`INTENT.md` was needed.
|
||
|
||
Work:
|
||
|
||
Delivered:
|
||
|
||
1. § 5.1 rewritten as "Resolution and rendering" — six resolution rules and
|
||
four rendering rules, with rendering explicitly deterministic and forbidden
|
||
from deriving (rule 10). Resolution must report which values were derived
|
||
(rule 6). A tool that resolves only supplied values and static defaults is
|
||
stated to be conforming.
|
||
2. § 6 gains `default` in the field table plus a new § 6.1 covering both
|
||
declaration forms, the static fallback, and why the reference form is
|
||
preferred.
|
||
3. § 4 manifest surface updated: `repository_context` — the field that
|
||
demonstrated the original defect — now carries a derived default with a
|
||
fallback, backed by a `requirement: generate` prompt dependency.
|
||
4. § 10 states the binding between `requirement: generate` and a referenced
|
||
derived default.
|
||
5. § 18 gains validation rules 11–13 and the publish-time inline warning.
|
||
6. § 19 gains two MUST NOTs: a derived default's prompt text is not an
|
||
instruction addressed to the consuming tool, and derivation must be visible
|
||
to the caller.
|
||
7. § 21 documents the new `resolve` verb.
|
||
8. `reference/canned_prompts.py`: `Resolution` dataclass, `resolve_inputs`
|
||
(with `resolve_values` kept as a wrapper), `input_default_kind`,
|
||
`prompt_dependency_ids`, `validate_input_default`, a `resolve` command, and
|
||
a specific render-time error naming the underivable inputs. Tests: 3 → 11,
|
||
all passing. Both READMEs document the split.
|
||
|
||
## Registry namespaces and ownership
|
||
|
||
```task
|
||
id: CANP-WP-0002-T02
|
||
status: done
|
||
priority: high
|
||
state_hub_task_id: "a52cd41e-2bcc-50ee-96a1-d645a06df922"
|
||
```
|
||
|
||
An id like `practice/pqrst-estimate` has a namespace prefix with no owner. Two
|
||
authors publishing `practice/…` into one registry collide today; `add` and
|
||
`publish` only refuse an exact `id@version` that already exists.
|
||
|
||
**Sharpened during the work.** § 17 already said what a *registry* does on
|
||
collision; nothing said what a *consumer* does. That is where the conflict
|
||
actually bites — in the catalog, after installing from two registries.
|
||
|
||
**Decisions (operator, 2026-09-06):**
|
||
|
||
- *Identity is registry-scoped.* An id names a package within a registry, the
|
||
way a path names a file within a repository. The same id from two registries
|
||
may be two different packages. A local-first format with no signing, no
|
||
federation and no central authority cannot enforce global uniqueness, and an
|
||
unenforceable guarantee is worse than none — it invites consumers to
|
||
conflate two packages that merely share a name. A qualified
|
||
`<registry>:<id>` reference distinguishes them.
|
||
- *Ownership is registry policy.* An optional `registry.yaml` names the
|
||
registry and records namespace claims (`owner`, `policy: open | closed`).
|
||
Packages never assert who owns their namespace, keeping unverifiable
|
||
authority claims out of artifacts (§ 19) and leaving package semantics
|
||
unchanged when a filesystem registry is later replaced by a hosted one.
|
||
- *Installing the same id from two registries is allowed, not a conflict.* The
|
||
catalog is namespaced by registry, so both are retained. This follows from
|
||
registry-scoped identity rather than working around it: two packages that
|
||
merely share a name were never in conflict to begin with.
|
||
|
||
Delivered:
|
||
|
||
1. § 3.2 gains "Identity is registry-scoped" — the reasoning, the qualified
|
||
reference form, and a ban on `:` in ids so the separator stays available.
|
||
A tool finding one id in several registries MUST report the ambiguity.
|
||
2. § 17 scopes immutability to a registry, consistent with identity.
|
||
3. § 20.1 (new) specifies the optional `registry.yaml`: `format`, `name`,
|
||
`description`, `namespaces`. Claims are explicitly descriptive — a
|
||
filesystem registry cannot authenticate a publisher. A bare directory
|
||
remains a valid registry, named by its basename.
|
||
4. § 20.2 (new) requires consumers to keep registries distinct and documents
|
||
the reference catalog layout.
|
||
5. § 18 gains registry-manifest validation; § 21 documents qualified
|
||
references and the reserved `local` registry name.
|
||
6. `reference/canned_prompts.py`: `parse_reference`, `check_registry_name`,
|
||
`read_registry_manifest`, `registry_name`, `namespace_policy`, split
|
||
`registry_package_path` / `catalog_package_path`, a `resolve_installed`
|
||
that reports ambiguity and returns the source registry, `iter_catalog`,
|
||
`--as` on `add`, a publish-time warning on closed namespaces, and a
|
||
specific error for the pre-registry-scoped catalog layout. Tests 11 → 21.
|
||
|
||
**Migration note:** the catalog layout changed. An existing catalog from
|
||
before this change is detected and reported with instructions rather than
|
||
failing as "package not found". Signing, trust scoring and federation remain
|
||
non-goals and were not touched.
|
||
|
||
## Prompt composition and inheritance
|
||
|
||
```task
|
||
id: CANP-WP-0002-T03
|
||
status: done
|
||
priority: high
|
||
state_hub_task_id: "0505c790-e57e-568c-9c8c-376957c9b08e"
|
||
```
|
||
|
||
Largest gap between `INTENT.md` and the spec. Principle 9 ("composition without
|
||
capture") and the `dependencies.prompts` field both promise composition, but
|
||
v0.1 defines no mechanism — `dependencies` is a declared field with no
|
||
semantics.
|
||
|
||
**Found during the work.** T01 had already delivered half of composition
|
||
without it being named as such: a derived default binds an input to a prompt
|
||
dependency, which is transclusion — run package B, use its output. What was
|
||
missing was the *deterministic* half. Separately, `type: template` appeared
|
||
once in the § 4 manifest surface and was defined nowhere in the specification.
|
||
|
||
**Decisions (operator, 2026-09-06):**
|
||
|
||
- *No inheritance.* A package never extends another, overrides its sections,
|
||
or inherits its inputs. Four of this repo's own documents argue against it:
|
||
INTENT principle 3 (hidden context defeats reuse); § 19's "make package
|
||
contents visible before execution", which an inheritance chain prevents;
|
||
§ 17's requirement that behavior changes produce a new version, which an
|
||
inherited change bypasses; and the non-goal on range resolution, which an
|
||
override chain would drag in.
|
||
- *Two composition kinds.* `include` inlines another package's **rendered
|
||
template** as text — deterministic, no model, and the reference CLI performs
|
||
it. `derive` (T01) uses another package's **result** — non-deterministic and
|
||
only available to a consumer that can run it. Both are expressed as input
|
||
defaults, so composition reuses the resolution machinery instead of adding a
|
||
second one.
|
||
- *Version selectors.* An exact pin is the expected form; `any`, `newest` and
|
||
`>= X.Y.Z` are explicit opt-ins, so looseness is always written rather than
|
||
implied by absence. These are per-dependency selectors evaluated
|
||
independently against what is available. There is no solver and no
|
||
cross-dependency constraint satisfaction, which is the line that keeps this
|
||
outside the range-resolution non-goal — the spec says so explicitly.
|
||
|
||
Delivered:
|
||
|
||
1. § 3.2 defines `type` (`template` | `fragment`), closing the undefined-field
|
||
gap. It is advisory: a fragment is a valid package and may be rendered
|
||
alone; the field records intent so a consumer can warn.
|
||
2. § 10.1 (new) covers naming and pinning: qualified dependency ids, the four
|
||
version selectors, `version` required for anything composed, and an
|
||
explicit statement of why this is not range resolution.
|
||
3. § 10.2 (new) defines both composition kinds, parameter pass-through into an
|
||
included package, mandatory cycle detection, and the reasoned refusal of
|
||
inheritance.
|
||
4. § 6.1 gains the included-default form; § 5.1 gains resolution rule 6 for
|
||
inclusion (rules renumbered); § 18 gains rules 14–16; § 21 notes that the
|
||
reference tool satisfies inclusions and reports derivations.
|
||
5. `reference/canned_prompts.py`: `validate_version_selector`,
|
||
`select_version`, `prompt_dependencies` (replacing
|
||
`prompt_dependency_ids`), `check_composition_reference`, `CatalogComposer`
|
||
with cycle detection, and `resolve_inputs(composer=, inherited=)`.
|
||
Tests 21 → 42.
|
||
6. `examples/house-style/` (new) is a real `type: fragment` package, and
|
||
`examples/pqrst-estimate` composes it — bumped 0.1.0 → 0.2.0 per § 17,
|
||
since including a style block changes intended behavior.
|
||
|
||
**Ordering bug found and fixed during the work.** Inputs were resolved before
|
||
parameters, so an included package could not see the including package's
|
||
parameters and silently fell back to its own defaults — the fragment rendered
|
||
`tone: neutral` where the including package said `blunt`. Parameters now
|
||
resolve first; the report still lists inputs before parameters.
|
||
|
||
## Canonical eval schemas
|
||
|
||
```task
|
||
id: CANP-WP-0002-T04
|
||
status: done
|
||
priority: medium
|
||
state_hub_task_id: "0daf1097-6599-563f-8965-739cbc874ddf"
|
||
```
|
||
|
||
Cheapest item to pin down. `evals/` is a reserved path and `evals:` is a
|
||
manifest list, but § 12 defines no schema, so an eval file is an unvalidated
|
||
blob that no tool can act on.
|
||
|
||
**Decisions (operator, 2026-09-06):**
|
||
|
||
- *Envelope plus one canonical schema.* Every eval file MUST declare a
|
||
`schema`; CPF defines exactly one, `canned-prompts/eval-rubric/v0.1`. Other
|
||
schemas stay legal and are skipped rather than rejected, so `evals/` holds
|
||
something tools can act on without CPF becoming an evaluation language.
|
||
- *Two kinds of check.* The same seam as T01 and T03: a **render check** is a
|
||
deterministic assertion about the rendered prompt text (`contains`,
|
||
`not_contains`, `resolves_all`) that any implementation able to render can
|
||
run, and **output criteria** are statements about a good result, declared but
|
||
not run because judging them needs a model.
|
||
- *Fixtures are referenced, not duplicated.* An eval names a path already
|
||
declared in the manifest's `examples`, tying two reserved paths together and
|
||
keeping one copy of each fixture.
|
||
|
||
Delivered:
|
||
|
||
1. § 12 rewritten: the envelope rule, the unknown-schema ignore rule, and
|
||
§ 12.1 specifying the rubric schema field by field.
|
||
2. Render checks and output criteria are specified separately, each with the
|
||
reason it does or does not run locally.
|
||
3. "Results are not part of the package" states that an eval declares
|
||
assessment and MUST NOT record outcomes; results are run evidence living
|
||
outside the immutable package, per `INTENT.md` and § 17.
|
||
4. § 18 gains validation rules 17–18; § 21 documents the `eval` verb.
|
||
5. `reference/canned_prompts.py`: `read_eval`, `validate_eval`,
|
||
`load_example_values`, `run_render_checks`, `cmd_eval`. A failed render
|
||
check exits non-zero. Tests 42 → 51.
|
||
6. `examples/pqrst-estimate/evals/quality.yaml` (new) is a real eval with four
|
||
render checks and four output criteria, and it passes.
|
||
|
||
**Fixed while implementing.** `coerce_value` assumed every incoming value was
|
||
a command-line string, so a YAML fixture carrying a real type
|
||
(`include_rationale: true`) crashed on `.lower()`. Typed values are now
|
||
validated but not re-parsed — reaching a fixture through `eval` was the first
|
||
code path that supplied them.
|
||
|
||
**Deferred deliberately.** No regex render check. It would add a matching
|
||
language and a backtracking hazard for little gain over `contains` at this
|
||
stage; `contains`, `not_contains` and `resolves_all` cover the cases the seed
|
||
actually has.
|
||
|
||
## Typed context and dependency contracts
|
||
|
||
```task
|
||
id: CANP-WP-0002-T05
|
||
status: todo
|
||
priority: medium
|
||
state_hub_task_id: "5beec7ba-e6c9-58b1-a065-b06aba015034"
|
||
```
|
||
|
||
`dependencies.context` and `dependencies.capabilities` appear in the § 4
|
||
manifest surface with no semantics whatsoever in v0.1, and § 9
|
||
`compatibility.capabilities` overlaps them without a stated relationship.
|
||
|
||
Decide: what a context dependency declares, how it differs from an input, how
|
||
it relates to `compatibility.capabilities`, and whether capability names are
|
||
free strings in v0.2 (model capability vocabularies stay deferred). T01 is now settled and partly answers this: a derived default is a
|
||
consumer-resolved context requirement expressed through `dependencies.prompts`
|
||
rather than through `dependencies.context`. Decide whether `context` is still a
|
||
distinct concept or collapses into the prompt-dependency mechanism.
|
||
|
||
## Rewrite specification section 23
|
||
|
||
```task
|
||
id: CANP-WP-0002-T06
|
||
status: todo
|
||
priority: medium
|
||
state_hub_task_id: "06032689-fc19-5fcd-a76e-5dbc87d02cd4"
|
||
```
|
||
|
||
After T01–T05 land, replace § 23's flat twelve-item list with two sections:
|
||
questions **being decided for v0.2** (each with its scoped question and stated
|
||
leaning, referencing this workplan) and questions **still deferred** (the eight
|
||
unpromoted items listed at the top of this file). Bump the spec status line if
|
||
the format revision warrants it.
|
||
|
||
## Reference implementation conformance fixes
|
||
|
||
```task
|
||
id: CANP-WP-0002-T07
|
||
status: todo
|
||
priority: low
|
||
state_hub_task_id: "f642abe6-8222-5958-88ee-7a53454f2344"
|
||
```
|
||
|
||
Two defects found in the same review, independent of the format questions:
|
||
|
||
1. **Prerelease versions sort as newest.** `parse_semver` in
|
||
`reference/canned_prompts.py` returns `(major, minor, patch, raw_string)`,
|
||
so `1.0.0-rc1` and `1.0.0` tie on the numeric fields and then compare as
|
||
strings — `"1.0.0-rc1" > "1.0.0"`. `resolve_installed` with no `--version`
|
||
therefore selects a prerelease over its own release. Order prerelease below
|
||
release, or state in § 17 that v0.1 ignores prerelease ordering.
|
||
2. **`copy_immutable` copies everything.** `add`/`publish` use
|
||
`shutil.copytree` over the whole source directory, so a stray `.git`,
|
||
`.venv`, or scratch file lands in the catalog and registry. § 2 says tools
|
||
MUST ignore unknown non-reserved files unless a manifest field references
|
||
them. Decide whether packaging is reserved-paths-only or an explicit
|
||
ignore list, then align the implementation and § 2.
|