canned-prompts/workplans/CANP-WP-0002-format-open-questions.md
tegwick 8376a00a88 CANP-WP-0002 T06: revision v0.2, and section 23 rewritten
Closes the workplan. The format becomes `canned-prompt/v0.2`, and packages
declaring v0.1 remain valid — everything added across T01-T05 is additive, so
a v0.1 package means exactly what it always meant. That is the MINOR case
section 17 itself describes.

The spec file loses its version suffix: CannedPromptFormat-v0.1.md becomes
CannedPromptFormat.md, with the revision stated inside. One stable path that
never breaks a link, and no rename per revision; the version belongs in the
`format` string where tools actually read it.

Section 23 is rewritten into three parts rather than the planned two.
"Settled since v0.1" tables the five resolved questions against where each
rule now lives. "Still deferred" carries the eight unpromoted items plus
pattern-matching render checks. "Decided against" holds template inheritance
alone, because calling it deferred would misdescribe it — reopening it means
overturning a decision and answering four recorded objections, not filling a
gap.

Section 23 also names the two habits the five decisions turned out to share,
so later revisions follow them rather than rediscover them: separate the
deterministic half from the rest, and a package never asserts what it cannot
back.

The eval-rubric and registry-manifest schemas keep their own v0.1. They are
new in this revision and sit on their own version lines.

Reference CLI: ACCEPTED_FORMATS; an unknown revision is rejected naming what is
accepted. Tests 78 -> 81.

Example packages declare v0.2 and are bumped 0.1.0 -> 0.1.1 and 0.2.0 -> 0.2.1
as section 17 PATCH — metadata corrections with behavior unchanged.

Also refreshes section 22's worked example, which had drifted: it showed
pqrst-estimate at 0.1.0 with no composition, contradicting the package
actually in the repo. It now mirrors the real package and doubles as a
composition illustration.

CANP-WP-0002 is finished. CANP-WP-0003 carries forward the one residual: the
default registry's basename-derived name reads as `registry:`.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Bjefh8NUiEiahN4JLwoSKM

Assistant: claude-code
Assistant-Model: opus
Assistant-Process: 388925@bnt-lap001
Assistant-Session: 3507023f-e0fd-4a1e-9d90-a0d4217d1502
2026-09-06 14:22:45 +02:00

481 lines
24 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

---
id: CANP-WP-0002
type: workplan
title: "Resolve CPF v0.1 open questions promoted for v0.2"
domain: agents
repo: canned-prompts
status: finished
owner: codex
topic_slug: practice
created: "2026-09-06"
updated: "2026-09-06"
reviewed_at: "2026-09-06"
reviewed_by: "claude"
context_paths:
- "INTENT.md"
- "CannedPromptFormat.md"
- "reference/canned_prompts.py"
- "examples/pqrst-estimate/"
state_hub_workstream_id: "c87c8e27-8b11-5306-b687-86779ef3f1be"
---
# Resolve CPF v0.1 open questions promoted for v0.2
`CannedPromptFormat.md` § 23 currently lists twelve items as "experience
should determine". A review of the seed on 2026-09-06 (spec + reference CLI +
`examples/pqrst-estimate`; 3/3 tests pass, full local lifecycle smokes clean)
plus an operator interview promoted five of them from "wait and see" to
"decide, with a stated leaning". This workplan carries those decisions.
Unpromoted § 23 items stay deferred and unchanged: content macros, cryptographic
integrity/signing, model capability vocabularies, run manifests and evidence
formats, deterministic compilation manifests, trust/reputation signals,
federated discovery, richer template syntax.
**Guardrail for every task below:** `INTENT.md` § Deliberate boundary. The
format may *declare* a requirement; it must not specify the resolver, runtime,
or execution engine that satisfies it.
**Outcome.** All seven tasks are done. The format is now revision v0.2
(`CannedPromptFormat.md`), the reference CLI covers eight verbs with 81 passing
tests, and § 23 records what was settled, what stays deferred, and what was
decided against. One residual is carried forward as `CANP-WP-0003`.
## Optional inputs need defaults
```task
id: CANP-WP-0002-T01
status: done
priority: high
state_hub_task_id: "e49bf036-0aee-5bb2-abbd-3d3f683df298"
```
**Gap found in review.** Rendering rule § 5.1(4) makes any unresolved
placeholder an error, and § 6 gives `inputs` no `default` field. An input
declared `required: false` and referenced from the template therefore makes
`render` fail unless a value is supplied — optional inputs are effectively
unusable. The spec's own § 4 manifest example hits this with
`repository_context`. `reference/canned_prompts.py` is conformant here; the
defect is in the spec.
**Decision (operator, 2026-09-06):** add a `default` field to `inputs`, and
allow that default to be either
- a **static** value, or
- a **derived** default — a prompt that produces the value from available
context.
Unresolved-with-no-default remains an error.
**Design note.** § 10 already carried the derivation mechanism:
`dependencies` accepts `requirement: generate`, defined as "a resolver MAY
satisfy a missing dependency by invoking an appropriate generation process",
with resolution left undefined. A derived default therefore did not need a new
concept, only a binding — the input's default names the dependency, and the
dependency records what may be generated and at which version.
**Follow-up decisions (operator, 2026-09-06):**
- *Both declaration forms are allowed.* A derived default may reference a
declared prompt dependency, or carry an inline `prompt`. The reference form
is preferred and the spec says why (an inline prompt is anonymous —
unversioned, unprovenanced, un-evaluable), and a validator SHOULD warn when a
**published** package derives inline. Inline is for local and draft packages.
- *Error unless static fallback.* A derived default MAY declare a static
`value`. A consumer that does not derive uses it; with no fallback the input
stays unresolved, which is an error. Derivation never silently yields empty
content.
- *Split render from resolve.* Resolution decides values and may be
non-deterministic; rendering substitutes and always is. Rendering MUST NOT
derive. This preserves INTENT success criterion 4 verbatim — no change to
`INTENT.md` was needed.
Work:
Delivered:
1. § 5.1 rewritten as "Resolution and rendering" — six resolution rules and
four rendering rules, with rendering explicitly deterministic and forbidden
from deriving (rule 10). Resolution must report which values were derived
(rule 6). A tool that resolves only supplied values and static defaults is
stated to be conforming.
2. § 6 gains `default` in the field table plus a new § 6.1 covering both
declaration forms, the static fallback, and why the reference form is
preferred.
3. § 4 manifest surface updated: `repository_context` — the field that
demonstrated the original defect — now carries a derived default with a
fallback, backed by a `requirement: generate` prompt dependency.
4. § 10 states the binding between `requirement: generate` and a referenced
derived default.
5. § 18 gains validation rules 1113 and the publish-time inline warning.
6. § 19 gains two MUST NOTs: a derived default's prompt text is not an
instruction addressed to the consuming tool, and derivation must be visible
to the caller.
7. § 21 documents the new `resolve` verb.
8. `reference/canned_prompts.py`: `Resolution` dataclass, `resolve_inputs`
(with `resolve_values` kept as a wrapper), `input_default_kind`,
`prompt_dependency_ids`, `validate_input_default`, a `resolve` command, and
a specific render-time error naming the underivable inputs. Tests: 3 → 11,
all passing. Both READMEs document the split.
## Registry namespaces and ownership
```task
id: CANP-WP-0002-T02
status: done
priority: high
state_hub_task_id: "a52cd41e-2bcc-50ee-96a1-d645a06df922"
```
An id like `practice/pqrst-estimate` has a namespace prefix with no owner. Two
authors publishing `practice/…` into one registry collide today; `add` and
`publish` only refuse an exact `id@version` that already exists.
**Sharpened during the work.** § 17 already said what a *registry* does on
collision; nothing said what a *consumer* does. That is where the conflict
actually bites — in the catalog, after installing from two registries.
**Decisions (operator, 2026-09-06):**
- *Identity is registry-scoped.* An id names a package within a registry, the
way a path names a file within a repository. The same id from two registries
may be two different packages. A local-first format with no signing, no
federation and no central authority cannot enforce global uniqueness, and an
unenforceable guarantee is worse than none — it invites consumers to
conflate two packages that merely share a name. A qualified
`<registry>:<id>` reference distinguishes them.
- *Ownership is registry policy.* An optional `registry.yaml` names the
registry and records namespace claims (`owner`, `policy: open | closed`).
Packages never assert who owns their namespace, keeping unverifiable
authority claims out of artifacts (§ 19) and leaving package semantics
unchanged when a filesystem registry is later replaced by a hosted one.
- *Installing the same id from two registries is allowed, not a conflict.* The
catalog is namespaced by registry, so both are retained. This follows from
registry-scoped identity rather than working around it: two packages that
merely share a name were never in conflict to begin with.
Delivered:
1. § 3.2 gains "Identity is registry-scoped" — the reasoning, the qualified
reference form, and a ban on `:` in ids so the separator stays available.
A tool finding one id in several registries MUST report the ambiguity.
2. § 17 scopes immutability to a registry, consistent with identity.
3. § 20.1 (new) specifies the optional `registry.yaml`: `format`, `name`,
`description`, `namespaces`. Claims are explicitly descriptive — a
filesystem registry cannot authenticate a publisher. A bare directory
remains a valid registry, named by its basename.
4. § 20.2 (new) requires consumers to keep registries distinct and documents
the reference catalog layout.
5. § 18 gains registry-manifest validation; § 21 documents qualified
references and the reserved `local` registry name.
6. `reference/canned_prompts.py`: `parse_reference`, `check_registry_name`,
`read_registry_manifest`, `registry_name`, `namespace_policy`, split
`registry_package_path` / `catalog_package_path`, a `resolve_installed`
that reports ambiguity and returns the source registry, `iter_catalog`,
`--as` on `add`, a publish-time warning on closed namespaces, and a
specific error for the pre-registry-scoped catalog layout. Tests 11 → 21.
**Migration note:** the catalog layout changed. An existing catalog from
before this change is detected and reported with instructions rather than
failing as "package not found". Signing, trust scoring and federation remain
non-goals and were not touched.
## Prompt composition and inheritance
```task
id: CANP-WP-0002-T03
status: done
priority: high
state_hub_task_id: "0505c790-e57e-568c-9c8c-376957c9b08e"
```
Largest gap between `INTENT.md` and the spec. Principle 9 ("composition without
capture") and the `dependencies.prompts` field both promise composition, but
v0.1 defines no mechanism — `dependencies` is a declared field with no
semantics.
**Found during the work.** T01 had already delivered half of composition
without it being named as such: a derived default binds an input to a prompt
dependency, which is transclusion — run package B, use its output. What was
missing was the *deterministic* half. Separately, `type: template` appeared
once in the § 4 manifest surface and was defined nowhere in the specification.
**Decisions (operator, 2026-09-06):**
- *No inheritance.* A package never extends another, overrides its sections,
or inherits its inputs. Four of this repo's own documents argue against it:
INTENT principle 3 (hidden context defeats reuse); § 19's "make package
contents visible before execution", which an inheritance chain prevents;
§ 17's requirement that behavior changes produce a new version, which an
inherited change bypasses; and the non-goal on range resolution, which an
override chain would drag in.
- *Two composition kinds.* `include` inlines another package's **rendered
template** as text — deterministic, no model, and the reference CLI performs
it. `derive` (T01) uses another package's **result** — non-deterministic and
only available to a consumer that can run it. Both are expressed as input
defaults, so composition reuses the resolution machinery instead of adding a
second one.
- *Version selectors.* An exact pin is the expected form; `any`, `newest` and
`>= X.Y.Z` are explicit opt-ins, so looseness is always written rather than
implied by absence. These are per-dependency selectors evaluated
independently against what is available. There is no solver and no
cross-dependency constraint satisfaction, which is the line that keeps this
outside the range-resolution non-goal — the spec says so explicitly.
Delivered:
1. § 3.2 defines `type` (`template` | `fragment`), closing the undefined-field
gap. It is advisory: a fragment is a valid package and may be rendered
alone; the field records intent so a consumer can warn.
2. § 10.3 (new) covers naming and pinning: qualified dependency ids, the four
version selectors, `version` required for anything composed, and an
explicit statement of why this is not range resolution.
3. § 10.4 (new) defines both composition kinds, parameter pass-through into an
included package, mandatory cycle detection, and the reasoned refusal of
inheritance.
4. § 6.1 gains the included-default form; § 5.1 gains resolution rule 6 for
inclusion (rules renumbered); § 18 gains rules 1416; § 21 notes that the
reference tool satisfies inclusions and reports derivations.
5. `reference/canned_prompts.py`: `validate_version_selector`,
`select_version`, `prompt_dependencies` (replacing
`prompt_dependency_ids`), `check_composition_reference`, `CatalogComposer`
with cycle detection, and `resolve_inputs(composer=, inherited=)`.
Tests 21 → 42.
6. `examples/house-style/` (new) is a real `type: fragment` package, and
`examples/pqrst-estimate` composes it — bumped 0.1.0 → 0.2.0 per § 17,
since including a style block changes intended behavior.
**Ordering bug found and fixed during the work.** Inputs were resolved before
parameters, so an included package could not see the including package's
parameters and silently fell back to its own defaults — the fragment rendered
`tone: neutral` where the including package said `blunt`. Parameters now
resolve first; the report still lists inputs before parameters.
## Canonical eval schemas
```task
id: CANP-WP-0002-T04
status: done
priority: medium
state_hub_task_id: "0daf1097-6599-563f-8965-739cbc874ddf"
```
Cheapest item to pin down. `evals/` is a reserved path and `evals:` is a
manifest list, but § 12 defines no schema, so an eval file is an unvalidated
blob that no tool can act on.
**Decisions (operator, 2026-09-06):**
- *Envelope plus one canonical schema.* Every eval file MUST declare a
`schema`; CPF defines exactly one, `canned-prompts/eval-rubric/v0.1`. Other
schemas stay legal and are skipped rather than rejected, so `evals/` holds
something tools can act on without CPF becoming an evaluation language.
- *Two kinds of check.* The same seam as T01 and T03: a **render check** is a
deterministic assertion about the rendered prompt text (`contains`,
`not_contains`, `resolves_all`) that any implementation able to render can
run, and **output criteria** are statements about a good result, declared but
not run because judging them needs a model.
- *Fixtures are referenced, not duplicated.* An eval names a path already
declared in the manifest's `examples`, tying two reserved paths together and
keeping one copy of each fixture.
Delivered:
1. § 12 rewritten: the envelope rule, the unknown-schema ignore rule, and
§ 12.1 specifying the rubric schema field by field.
2. Render checks and output criteria are specified separately, each with the
reason it does or does not run locally.
3. "Results are not part of the package" states that an eval declares
assessment and MUST NOT record outcomes; results are run evidence living
outside the immutable package, per `INTENT.md` and § 17.
4. § 18 gains validation rules 1718; § 21 documents the `eval` verb.
5. `reference/canned_prompts.py`: `read_eval`, `validate_eval`,
`load_example_values`, `run_render_checks`, `cmd_eval`. A failed render
check exits non-zero. Tests 42 → 51.
6. `examples/pqrst-estimate/evals/quality.yaml` (new) is a real eval with four
render checks and four output criteria, and it passes.
**Fixed while implementing.** `coerce_value` assumed every incoming value was
a command-line string, so a YAML fixture carrying a real type
(`include_rationale: true`) crashed on `.lower()`. Typed values are now
validated but not re-parsed — reaching a fixture through `eval` was the first
code path that supplied them.
**Deferred deliberately.** No regex render check. It would add a matching
language and a backtracking hazard for little gain over `contains` at this
stage; `contains`, `not_contains` and `resolves_all` cover the cases the seed
actually has.
## Typed context and dependency contracts
```task
id: CANP-WP-0002-T05
status: done
priority: medium
state_hub_task_id: "5beec7ba-e6c9-58b1-a065-b06aba015034"
```
`dependencies.context` and `dependencies.capabilities` appear in the § 4
manifest surface with no semantics whatsoever in v0.1, and § 9
`compatibility.capabilities` overlaps them without a stated relationship.
Decide: what a context dependency declares, how it differs from an input, how
it relates to `compatibility.capabilities`, and whether capability names are
free strings in v0.2 (model capability vocabularies stay deferred). **Decisions (operator, 2026-09-06):**
- *`context` names what CPF cannot package.* `prompts` names artifacts the
format resolves by id and version; `context` names what it does not and will
not package — a live information space, an API, a corpus, a document the
caller supplies. Anything CPF *can* package is a package: a reusable policy
or style block belongs in `prompts` as a `type: fragment`.
- *Capability names are free strings*, lowercase kebab-case, validated for
shape and not for membership, exactly as `tags` are. Model capability
vocabularies stay deferred.
- *Terminology, not deduplication, for the capability overlap.* The operator's
answer was conditional — keep both fields if they carry the required/observed
distinction, and fix the terminology if they do not. **They did not.** § 9
opened with "records known requirements **or** observations", mixing both in
one field: `models` was observational ("known to be compatible or
evaluated") while `capabilities` was prescriptive ("expected from the
execution environment"). § 10 then described dependencies as what a prompt
"expects". Both fields said *expected*, so the ambiguity was real and the
fix was terminology.
Delivered:
1. § 9 rewritten: `compatibility` records **observations, never
requirements**, and a consumer MUST NOT refuse to run a package because its
environment is absent from those lists. Adds `aliases`, recording the same
capability under other names — the operator's point that a capability can
be "known by another name".
2. § 9.1 (new) states the distinction in one word each — `dependencies` means
*required*, `compatibility` means *observed* — with what absence implies
for each, and notes that the same name may legitimately appear in both.
3. § 10 opens with the three dependency kinds separated by what CPF can do
about them, and § 10.1 (new) specifies context dependencies: `name` rather
than `id` because nothing can look them up, a **required** `description`
because nothing else can explain an unpackaged dependency, optional `kind`,
and `requirement` limited to `required`/`optional`.
4. § 10.2 (new) specifies capability dependencies: free-form kebab-case, no
version and no `requirement: generate`, because a capability is not an
artifact and cannot be fetched, pinned or generated.
5. § 10.1 also draws the input/context line: an input is content passed for one
use; a context dependency is a standing fact about the environment.
6. § 18 gains validation rules 1920; § 4's manifest surface updated. Former
§ 10.1/10.2 renumbered to § 10.3/10.4 with all cross-references updated.
7. `reference/canned_prompts.py`: `validate_capabilities`,
`validate_context_dependencies`, and a `resolve` section listing required
capabilities and context under "this tool cannot verify these" rather than
implying it checked. Tests 51 → 65.
**Caught in review of my own change.** I first added a `session-record`
context dependency to `examples/pqrst-estimate` to demonstrate the feature,
then removed it: it described the `session_summary` *input*, which § 10.1
explicitly says a context dependency is not. The example now declares only
what it genuinely has. Illustrations live in the spec; examples stay honest.
## Rewrite specification section 23
```task
id: CANP-WP-0002-T06
status: done
priority: medium
state_hub_task_id: "06032689-fc19-5fcd-a76e-5dbc87d02cd4"
```
**Decisions (operator, 2026-09-06):**
- *Bump to `canned-prompt/v0.2`, accept both.* Everything added is additive, so
a v0.1 package means exactly what it always meant — the MINOR case § 17
itself describes. Tools accept both strings; new packages declare v0.2.
- *Drop the version from the spec filename.* `CannedPromptFormat-v0.1.md`
`CannedPromptFormat.md`, with the revision stated in the document. One stable
path that never breaks a link, and no rename per revision; the version lives
in the `format` string, where tools actually read it.
Delivered:
1. § 23 rewritten into three parts rather than the planned two. **Settled since
v0.1** tables the five resolved questions against where each rule now lives.
**Still deferred** carries the eight unpromoted items plus pattern-matching
render checks, deferred during T04. **Decided against** holds template
inheritance on its own, because "deferred" would misdescribe it: reopening
it means overturning a decision and answering four recorded objections, not
filling a gap.
2. § 23 also names the two habits the five decisions share, so later revisions
follow them rather than rediscover them: *separate the deterministic half
from the rest*, and *a package never asserts what it cannot back*.
3. Header now carries revision, status and a compatibility line; § 3.2 states
the acceptance rule and why v0.2 is MINOR; § 18 rule 3 checks for a
recognized revision rather than one literal string.
4. Stale "CPF v0.1" phrasings throughout replaced with unversioned wording. The
`canned-prompts/eval-rubric/v0.1` and `canned-prompt-registry/v0.1` schemas
deliberately keep their own v0.1 — they are new in this revision and on
their own version lines.
5. `reference/canned_prompts.py`: `ACCEPTED_FORMATS`; unknown revisions are
rejected with the accepted list. Tests 78 → 81.
6. Example packages declare v0.2, bumped 0.1.0 → 0.1.1 and 0.2.0 → 0.2.1 as
§ 17 PATCH — a metadata correction with intended behavior unchanged.
**Also fixed.** § 22's worked example had drifted: it showed `pqrst-estimate`
at 0.1.0 with no composition, contradicting the package actually in the repo.
It now mirrors the real package and doubles as a § 10.4 illustration.
## Reference implementation conformance fixes
```task
id: CANP-WP-0002-T07
status: done
priority: low
state_hub_task_id: "f642abe6-8222-5958-88ee-7a53454f2344"
```
Two defects found in the same review, independent of the format questions:
1. **Prerelease versions sorted as newest.** Confirmed worse than first
recorded: `1.0.0-rc1` shadowed `1.0.0` for `newest` *and* `>= 1.0.0`, so a
release candidate hid its own release from every selector.
2. **`copy_immutable` copied everything**, so a stray `.git`, `.venv` or
scratch file landed in the catalog and registry.
**Decisions (operator, 2026-09-06):**
- *Prereleases are excluded unless named.* They sort below their release and
are skipped by `any`, `newest` and `>=`; only an exact pin selects one.
Publishing a release candidate therefore never changes what existing
consumers resolve to — the point of marking it a candidate. Matches npm and
cargo.
- *`LICENSE` joins the reserved paths.* Strict packaging would otherwise drop
a package's license text while faithfully copying its `license` field, which
contradicts § 14's instruction to surface licensing on publish and install.
The strict-versus-ignore-list question needed no decision: § 2 already
required it — "Tools MUST ignore unknown non-reserved files unless a manifest
field explicitly references them" — so packaging by reserved path is
conformance, not a new choice. What was worth adding is that omissions are
**reported**, never silent.
Delivered:
1. `parse_semver` now returns a SemVer § 11 precedence key: numeric fields,
then a release/prerelease rank, then dot-separated prerelease identifiers
with numeric ones compared numerically. Build metadata is ignored, so
`1.0.0+build.1` and `1.0.0+build.2` tie. `1.0.0-rc.10` correctly outranks
`1.0.0-rc.2`.
2. `select_version` excludes prereleases from `any`, `newest` and `>=`;
`pick_version` reports "only prerelease versions are available" rather than
a bare not-found.
3. `copy_package` replaces `copy_immutable`, copying reserved paths plus
manifest-referenced files only, and returning what it skipped;
`report_skipped` names those on stderr. `skipped_entries` reports top-level
names without walking into an ignored directory, so a stray virtualenv
costs nothing to skip.
4. § 2 gains `LICENSE`, the packaging obligation, and the reporting SHOULD;
§ 17.1 (new) specifies precedence and prerelease selection; § 10.3 notes
the exclusion.
5. `examples/pqrst-estimate` now carries a `LICENSE` and a `license: MIT`
field, exercising the new reserved path. Tests 65 → 78.
**Fixed while implementing.** `package_members` reused a loop variable as the
error label, so a bad `template` path would have been reported as an `evals`
error.