canned-prompts/workplans/CANP-WP-0002-format-open-questions.md
tegwick 95a8bb31d2 CANP-WP-0002 T07: semver precedence and strict packaging
Two defects from the original review of the seed.

Prerelease ordering was worse than first recorded. parse_semver returned
(major, minor, patch, raw_string), so 1.0.0-rc1 and 1.0.0 tied on the numeric
fields and then compared as strings — "1.0.0-rc1" > "1.0.0". A release
candidate therefore shadowed its own release for `newest` and for `>=`, not
just for the no-version case.

parse_semver now returns a SemVer section 11 precedence key: numeric fields, a
release/prerelease rank, then dot-separated prerelease identifiers with
numeric ones compared numerically. Build metadata is ignored. Beyond ordering,
prereleases are excluded from `any`, `newest` and `>=` entirely; only an exact
pin selects one, so publishing a release candidate never changes what existing
consumers resolve to. A package holding only prereleases now says so rather
than reporting a bare not-found.

Packaging copied the whole source directory, so a stray .git, virtualenv or
scratch file landed in the catalog and registry. Section 2 already required
otherwise — tools MUST ignore unknown non-reserved files unless a manifest
field references them — so this is conformance rather than a new rule. What is
new is that omissions are reported instead of silent:

  not packaged (not a reserved path, not referenced by the manifest): .git/, .venv/, notes.txt

LICENSE joins the reserved paths. Strict packaging would otherwise drop a
package's license text while faithfully copying its `license` field, which
contradicts section 14's instruction to surface licensing on publish and
install.

Spec: 2 (LICENSE, packaging obligation, reporting), 17.1 new, 10.3 note.

Reference CLI: parse_semver rewritten with is_prerelease; select_version and
pick_version updated; copy_package and report_skipped replace copy_immutable.
Tests 65 -> 78.

examples/pqrst-estimate carries a LICENSE and a license field, exercising the
new reserved path.

Also fixes a leaked loop variable in package_members that would have reported
a bad `template` path as an `evals` error.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Bjefh8NUiEiahN4JLwoSKM

Assistant: claude-code
Assistant-Model: opus
Assistant-Process: 388925@bnt-lap001
Assistant-Session: 3507023f-e0fd-4a1e-9d90-a0d4217d1502
2026-09-06 09:32:42 +02:00

22 KiB
Raw Blame History

id type title domain repo status owner topic_slug created updated reviewed_at reviewed_by context_paths state_hub_workstream_id
CANP-WP-0002 workplan Resolve CPF v0.1 open questions promoted for v0.2 agents canned-prompts proposed codex practice 2026-09-06 2026-09-06 2026-09-06 claude
INTENT.md
CannedPromptFormat-v0.1.md
reference/canned_prompts.py
examples/pqrst-estimate/
c87c8e27-8b11-5306-b687-86779ef3f1be

Resolve CPF v0.1 open questions promoted for v0.2

CannedPromptFormat-v0.1.md § 23 currently lists twelve items as "experience should determine". A review of the seed on 2026-09-06 (spec + reference CLI + examples/pqrst-estimate; 3/3 tests pass, full local lifecycle smokes clean) plus an operator interview promoted five of them from "wait and see" to "decide, with a stated leaning". This workplan carries those decisions.

Unpromoted § 23 items stay deferred and unchanged: content macros, cryptographic integrity/signing, model capability vocabularies, run manifests and evidence formats, deterministic compilation manifests, trust/reputation signals, federated discovery, richer template syntax.

Guardrail for every task below: INTENT.md § Deliberate boundary. The format may declare a requirement; it must not specify the resolver, runtime, or execution engine that satisfies it.

Optional inputs need defaults

id: CANP-WP-0002-T01
status: done
priority: high
state_hub_task_id: "e49bf036-0aee-5bb2-abbd-3d3f683df298"

Gap found in review. Rendering rule § 5.1(4) makes any unresolved placeholder an error, and § 6 gives inputs no default field. An input declared required: false and referenced from the template therefore makes render fail unless a value is supplied — optional inputs are effectively unusable. The spec's own § 4 manifest example hits this with repository_context. reference/canned_prompts.py is conformant here; the defect is in the spec.

Decision (operator, 2026-09-06): add a default field to inputs, and allow that default to be either

  • a static value, or
  • a derived default — a prompt that produces the value from available context.

Unresolved-with-no-default remains an error.

Design note. § 10 already carried the derivation mechanism: dependencies accepts requirement: generate, defined as "a resolver MAY satisfy a missing dependency by invoking an appropriate generation process", with resolution left undefined. A derived default therefore did not need a new concept, only a binding — the input's default names the dependency, and the dependency records what may be generated and at which version.

Follow-up decisions (operator, 2026-09-06):

  • Both declaration forms are allowed. A derived default may reference a declared prompt dependency, or carry an inline prompt. The reference form is preferred and the spec says why (an inline prompt is anonymous — unversioned, unprovenanced, un-evaluable), and a validator SHOULD warn when a published package derives inline. Inline is for local and draft packages.
  • Error unless static fallback. A derived default MAY declare a static value. A consumer that does not derive uses it; with no fallback the input stays unresolved, which is an error. Derivation never silently yields empty content.
  • Split render from resolve. Resolution decides values and may be non-deterministic; rendering substitutes and always is. Rendering MUST NOT derive. This preserves INTENT success criterion 4 verbatim — no change to INTENT.md was needed.

Work:

Delivered:

  1. § 5.1 rewritten as "Resolution and rendering" — six resolution rules and four rendering rules, with rendering explicitly deterministic and forbidden from deriving (rule 10). Resolution must report which values were derived (rule 6). A tool that resolves only supplied values and static defaults is stated to be conforming.
  2. § 6 gains default in the field table plus a new § 6.1 covering both declaration forms, the static fallback, and why the reference form is preferred.
  3. § 4 manifest surface updated: repository_context — the field that demonstrated the original defect — now carries a derived default with a fallback, backed by a requirement: generate prompt dependency.
  4. § 10 states the binding between requirement: generate and a referenced derived default.
  5. § 18 gains validation rules 1113 and the publish-time inline warning.
  6. § 19 gains two MUST NOTs: a derived default's prompt text is not an instruction addressed to the consuming tool, and derivation must be visible to the caller.
  7. § 21 documents the new resolve verb.
  8. reference/canned_prompts.py: Resolution dataclass, resolve_inputs (with resolve_values kept as a wrapper), input_default_kind, prompt_dependency_ids, validate_input_default, a resolve command, and a specific render-time error naming the underivable inputs. Tests: 3 → 11, all passing. Both READMEs document the split.

Registry namespaces and ownership

id: CANP-WP-0002-T02
status: done
priority: high
state_hub_task_id: "a52cd41e-2bcc-50ee-96a1-d645a06df922"

An id like practice/pqrst-estimate has a namespace prefix with no owner. Two authors publishing practice/… into one registry collide today; add and publish only refuse an exact id@version that already exists.

Sharpened during the work. § 17 already said what a registry does on collision; nothing said what a consumer does. That is where the conflict actually bites — in the catalog, after installing from two registries.

Decisions (operator, 2026-09-06):

  • Identity is registry-scoped. An id names a package within a registry, the way a path names a file within a repository. The same id from two registries may be two different packages. A local-first format with no signing, no federation and no central authority cannot enforce global uniqueness, and an unenforceable guarantee is worse than none — it invites consumers to conflate two packages that merely share a name. A qualified <registry>:<id> reference distinguishes them.
  • Ownership is registry policy. An optional registry.yaml names the registry and records namespace claims (owner, policy: open | closed). Packages never assert who owns their namespace, keeping unverifiable authority claims out of artifacts (§ 19) and leaving package semantics unchanged when a filesystem registry is later replaced by a hosted one.
  • Installing the same id from two registries is allowed, not a conflict. The catalog is namespaced by registry, so both are retained. This follows from registry-scoped identity rather than working around it: two packages that merely share a name were never in conflict to begin with.

Delivered:

  1. § 3.2 gains "Identity is registry-scoped" — the reasoning, the qualified reference form, and a ban on : in ids so the separator stays available. A tool finding one id in several registries MUST report the ambiguity.
  2. § 17 scopes immutability to a registry, consistent with identity.
  3. § 20.1 (new) specifies the optional registry.yaml: format, name, description, namespaces. Claims are explicitly descriptive — a filesystem registry cannot authenticate a publisher. A bare directory remains a valid registry, named by its basename.
  4. § 20.2 (new) requires consumers to keep registries distinct and documents the reference catalog layout.
  5. § 18 gains registry-manifest validation; § 21 documents qualified references and the reserved local registry name.
  6. reference/canned_prompts.py: parse_reference, check_registry_name, read_registry_manifest, registry_name, namespace_policy, split registry_package_path / catalog_package_path, a resolve_installed that reports ambiguity and returns the source registry, iter_catalog, --as on add, a publish-time warning on closed namespaces, and a specific error for the pre-registry-scoped catalog layout. Tests 11 → 21.

Migration note: the catalog layout changed. An existing catalog from before this change is detected and reported with instructions rather than failing as "package not found". Signing, trust scoring and federation remain non-goals and were not touched.

Prompt composition and inheritance

id: CANP-WP-0002-T03
status: done
priority: high
state_hub_task_id: "0505c790-e57e-568c-9c8c-376957c9b08e"

Largest gap between INTENT.md and the spec. Principle 9 ("composition without capture") and the dependencies.prompts field both promise composition, but v0.1 defines no mechanism — dependencies is a declared field with no semantics.

Found during the work. T01 had already delivered half of composition without it being named as such: a derived default binds an input to a prompt dependency, which is transclusion — run package B, use its output. What was missing was the deterministic half. Separately, type: template appeared once in the § 4 manifest surface and was defined nowhere in the specification.

Decisions (operator, 2026-09-06):

  • No inheritance. A package never extends another, overrides its sections, or inherits its inputs. Four of this repo's own documents argue against it: INTENT principle 3 (hidden context defeats reuse); § 19's "make package contents visible before execution", which an inheritance chain prevents; § 17's requirement that behavior changes produce a new version, which an inherited change bypasses; and the non-goal on range resolution, which an override chain would drag in.
  • Two composition kinds. include inlines another package's rendered template as text — deterministic, no model, and the reference CLI performs it. derive (T01) uses another package's result — non-deterministic and only available to a consumer that can run it. Both are expressed as input defaults, so composition reuses the resolution machinery instead of adding a second one.
  • Version selectors. An exact pin is the expected form; any, newest and >= X.Y.Z are explicit opt-ins, so looseness is always written rather than implied by absence. These are per-dependency selectors evaluated independently against what is available. There is no solver and no cross-dependency constraint satisfaction, which is the line that keeps this outside the range-resolution non-goal — the spec says so explicitly.

Delivered:

  1. § 3.2 defines type (template | fragment), closing the undefined-field gap. It is advisory: a fragment is a valid package and may be rendered alone; the field records intent so a consumer can warn.
  2. § 10.3 (new) covers naming and pinning: qualified dependency ids, the four version selectors, version required for anything composed, and an explicit statement of why this is not range resolution.
  3. § 10.4 (new) defines both composition kinds, parameter pass-through into an included package, mandatory cycle detection, and the reasoned refusal of inheritance.
  4. § 6.1 gains the included-default form; § 5.1 gains resolution rule 6 for inclusion (rules renumbered); § 18 gains rules 1416; § 21 notes that the reference tool satisfies inclusions and reports derivations.
  5. reference/canned_prompts.py: validate_version_selector, select_version, prompt_dependencies (replacing prompt_dependency_ids), check_composition_reference, CatalogComposer with cycle detection, and resolve_inputs(composer=, inherited=). Tests 21 → 42.
  6. examples/house-style/ (new) is a real type: fragment package, and examples/pqrst-estimate composes it — bumped 0.1.0 → 0.2.0 per § 17, since including a style block changes intended behavior.

Ordering bug found and fixed during the work. Inputs were resolved before parameters, so an included package could not see the including package's parameters and silently fell back to its own defaults — the fragment rendered tone: neutral where the including package said blunt. Parameters now resolve first; the report still lists inputs before parameters.

Canonical eval schemas

id: CANP-WP-0002-T04
status: done
priority: medium
state_hub_task_id: "0daf1097-6599-563f-8965-739cbc874ddf"

Cheapest item to pin down. evals/ is a reserved path and evals: is a manifest list, but § 12 defines no schema, so an eval file is an unvalidated blob that no tool can act on.

Decisions (operator, 2026-09-06):

  • Envelope plus one canonical schema. Every eval file MUST declare a schema; CPF defines exactly one, canned-prompts/eval-rubric/v0.1. Other schemas stay legal and are skipped rather than rejected, so evals/ holds something tools can act on without CPF becoming an evaluation language.
  • Two kinds of check. The same seam as T01 and T03: a render check is a deterministic assertion about the rendered prompt text (contains, not_contains, resolves_all) that any implementation able to render can run, and output criteria are statements about a good result, declared but not run because judging them needs a model.
  • Fixtures are referenced, not duplicated. An eval names a path already declared in the manifest's examples, tying two reserved paths together and keeping one copy of each fixture.

Delivered:

  1. § 12 rewritten: the envelope rule, the unknown-schema ignore rule, and § 12.1 specifying the rubric schema field by field.
  2. Render checks and output criteria are specified separately, each with the reason it does or does not run locally.
  3. "Results are not part of the package" states that an eval declares assessment and MUST NOT record outcomes; results are run evidence living outside the immutable package, per INTENT.md and § 17.
  4. § 18 gains validation rules 1718; § 21 documents the eval verb.
  5. reference/canned_prompts.py: read_eval, validate_eval, load_example_values, run_render_checks, cmd_eval. A failed render check exits non-zero. Tests 42 → 51.
  6. examples/pqrst-estimate/evals/quality.yaml (new) is a real eval with four render checks and four output criteria, and it passes.

Fixed while implementing. coerce_value assumed every incoming value was a command-line string, so a YAML fixture carrying a real type (include_rationale: true) crashed on .lower(). Typed values are now validated but not re-parsed — reaching a fixture through eval was the first code path that supplied them.

Deferred deliberately. No regex render check. It would add a matching language and a backtracking hazard for little gain over contains at this stage; contains, not_contains and resolves_all cover the cases the seed actually has.

Typed context and dependency contracts

id: CANP-WP-0002-T05
status: done
priority: medium
state_hub_task_id: "5beec7ba-e6c9-58b1-a065-b06aba015034"

dependencies.context and dependencies.capabilities appear in the § 4 manifest surface with no semantics whatsoever in v0.1, and § 9 compatibility.capabilities overlaps them without a stated relationship.

Decide: what a context dependency declares, how it differs from an input, how it relates to compatibility.capabilities, and whether capability names are free strings in v0.2 (model capability vocabularies stay deferred). Decisions (operator, 2026-09-06):

  • context names what CPF cannot package. prompts names artifacts the format resolves by id and version; context names what it does not and will not package — a live information space, an API, a corpus, a document the caller supplies. Anything CPF can package is a package: a reusable policy or style block belongs in prompts as a type: fragment.
  • Capability names are free strings, lowercase kebab-case, validated for shape and not for membership, exactly as tags are. Model capability vocabularies stay deferred.
  • Terminology, not deduplication, for the capability overlap. The operator's answer was conditional — keep both fields if they carry the required/observed distinction, and fix the terminology if they do not. They did not. § 9 opened with "records known requirements or observations", mixing both in one field: models was observational ("known to be compatible or evaluated") while capabilities was prescriptive ("expected from the execution environment"). § 10 then described dependencies as what a prompt "expects". Both fields said expected, so the ambiguity was real and the fix was terminology.

Delivered:

  1. § 9 rewritten: compatibility records observations, never requirements, and a consumer MUST NOT refuse to run a package because its environment is absent from those lists. Adds aliases, recording the same capability under other names — the operator's point that a capability can be "known by another name".
  2. § 9.1 (new) states the distinction in one word each — dependencies means required, compatibility means observed — with what absence implies for each, and notes that the same name may legitimately appear in both.
  3. § 10 opens with the three dependency kinds separated by what CPF can do about them, and § 10.1 (new) specifies context dependencies: name rather than id because nothing can look them up, a required description because nothing else can explain an unpackaged dependency, optional kind, and requirement limited to required/optional.
  4. § 10.2 (new) specifies capability dependencies: free-form kebab-case, no version and no requirement: generate, because a capability is not an artifact and cannot be fetched, pinned or generated.
  5. § 10.1 also draws the input/context line: an input is content passed for one use; a context dependency is a standing fact about the environment.
  6. § 18 gains validation rules 1920; § 4's manifest surface updated. Former § 10.1/10.2 renumbered to § 10.3/10.4 with all cross-references updated.
  7. reference/canned_prompts.py: validate_capabilities, validate_context_dependencies, and a resolve section listing required capabilities and context under "this tool cannot verify these" rather than implying it checked. Tests 51 → 65.

Caught in review of my own change. I first added a session-record context dependency to examples/pqrst-estimate to demonstrate the feature, then removed it: it described the session_summary input, which § 10.1 explicitly says a context dependency is not. The example now declares only what it genuinely has. Illustrations live in the spec; examples stay honest.

Rewrite specification section 23

id: CANP-WP-0002-T06
status: todo
priority: medium
state_hub_task_id: "06032689-fc19-5fcd-a76e-5dbc87d02cd4"

After T01T05 land, replace § 23's flat twelve-item list with two sections: questions being decided for v0.2 (each with its scoped question and stated leaning, referencing this workplan) and questions still deferred (the eight unpromoted items listed at the top of this file). Bump the spec status line if the format revision warrants it.

Reference implementation conformance fixes

id: CANP-WP-0002-T07
status: done
priority: low
state_hub_task_id: "f642abe6-8222-5958-88ee-7a53454f2344"

Two defects found in the same review, independent of the format questions:

  1. Prerelease versions sorted as newest. Confirmed worse than first recorded: 1.0.0-rc1 shadowed 1.0.0 for newest and >= 1.0.0, so a release candidate hid its own release from every selector.
  2. copy_immutable copied everything, so a stray .git, .venv or scratch file landed in the catalog and registry.

Decisions (operator, 2026-09-06):

  • Prereleases are excluded unless named. They sort below their release and are skipped by any, newest and >=; only an exact pin selects one. Publishing a release candidate therefore never changes what existing consumers resolve to — the point of marking it a candidate. Matches npm and cargo.
  • LICENSE joins the reserved paths. Strict packaging would otherwise drop a package's license text while faithfully copying its license field, which contradicts § 14's instruction to surface licensing on publish and install.

The strict-versus-ignore-list question needed no decision: § 2 already required it — "Tools MUST ignore unknown non-reserved files unless a manifest field explicitly references them" — so packaging by reserved path is conformance, not a new choice. What was worth adding is that omissions are reported, never silent.

Delivered:

  1. parse_semver now returns a SemVer § 11 precedence key: numeric fields, then a release/prerelease rank, then dot-separated prerelease identifiers with numeric ones compared numerically. Build metadata is ignored, so 1.0.0+build.1 and 1.0.0+build.2 tie. 1.0.0-rc.10 correctly outranks 1.0.0-rc.2.
  2. select_version excludes prereleases from any, newest and >=; pick_version reports "only prerelease versions are available" rather than a bare not-found.
  3. copy_package replaces copy_immutable, copying reserved paths plus manifest-referenced files only, and returning what it skipped; report_skipped names those on stderr. skipped_entries reports top-level names without walking into an ignored directory, so a stray virtualenv costs nothing to skip.
  4. § 2 gains LICENSE, the packaging obligation, and the reporting SHOULD; § 17.1 (new) specifies precedence and prerelease selection; § 10.3 notes the exclusion.
  5. examples/pqrst-estimate now carries a LICENSE and a license: MIT field, exercising the new reserved path. Tests 65 → 78.

Fixed while implementing. package_members reused a loop variable as the error label, so a bad template path would have been reported as an evals error.