Commit graph

2 commits

Author SHA1 Message Date
f55ef13c75 Package the canonical PQRST prompt faithfully
examples/pqrst-estimate was not the canonical PQRST prompt. The canonical one
is ~/pqrst-practice/PqrstPrompt.md, normatively specified in
spec/PqrstEstimationPractice.md, and hall-of-helix CLOSING.md requires pasting
that block unmodified. Ours was a paraphrase with a different output shape — no
Confidence, no Signature, no Dominant factors — and CANP-WP-0002-T03 made it
worse by prepending a house_style inclusion to a prompt whose governing
document says do not modify it.

A copied prompt that silently forked its source is the exact failure INTENT.md
opens with, sitting in this repo's own examples directory.

The template is now the canonical block extracted programmatically rather than
retyped, and the default render is byte-identical to it: 2886 bytes both ways.

The source documents two optional add-ons appended after the block. CPF has no
conditionals, so they are not a flag: `add_ons` is an input defaulting to the
empty string, and each sanctioned add-on is an example fixture. This is worth
noting as evidence about section 23's deferred "richer template syntax" — the
workaround is adequate here, but it is a workaround.

evals/canonical-fidelity.yaml guards the property with sixteen render checks,
including not_contains checks naming the paraphrase this package used to be, so
the drift cannot silently recur.

Version 0.2.1 -> 1.0.0: changed inputs and a materially different intended
output is the MAJOR case in section 17.

Section 22's worked example and the house-style README both claimed
pqrst-estimate composes the style fragment. It no longer does, by design, so
both are corrected; composition is illustrated in section 10.4 instead.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Bjefh8NUiEiahN4JLwoSKM

Assistant: claude-code
Assistant-Model: opus
Assistant-Process: 388925@bnt-lap001
Assistant-Session: 3507023f-e0fd-4a1e-9d90-a0d4217d1502
2026-09-06 15:26:36 +02:00
c5f1640454 CANP-WP-0002 T04: eval schema, render checks and output criteria
`evals/` was a reserved path holding unvalidated blobs: section 12 named the
directory and gave an illustrative snippet, but nothing was specified, so no
tool could act on an eval file.

Every eval file must now declare a `schema`, and CPF defines exactly one —
`canned-prompts/eval-rubric/v0.1`. Unrecognized schemas stay legal and are
skipped rather than rejected, so the format gains something actionable without
becoming an evaluation language, which remains a non-goal.

The schema splits along the same seam as T01 and T03. Render checks
(`contains`, `not_contains`, `resolves_all`) assert properties of the rendered
prompt text, need no model, and are therefore run by the reference CLI. Output
criteria describe a good result and are declared but not run, because judging
them requires a model. That division is now the format's consistent answer to
"deterministic locally, or not".

An eval references a fixture already declared in the manifest's `examples`
rather than carrying its own copy, so an example that is also an eval fixture
stays honest — both break together.

An eval declares assessment and must not record outcomes. Results are run
evidence and live outside the immutable package, per INTENT.md and section 17.

Spec: 12 rewritten with 12.1, 18 (rules 17-18), 21 (`eval` verb).

Reference CLI: read_eval, validate_eval, load_example_values,
run_render_checks, cmd_eval; a failed render check exits non-zero.
Tests 42 -> 51.

examples/pqrst-estimate/evals/quality.yaml is a real eval with four render
checks and four output criteria, and it passes.

Fixes a latent bug reaching a fixture exposed: coerce_value assumed every
value was a command-line string, so a YAML fixture carrying a real type
(include_rationale: true) crashed on .lower(). Typed values are now validated
but not re-parsed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Bjefh8NUiEiahN4JLwoSKM

Assistant: claude-code
Assistant-Model: opus
Assistant-Process: 388925@bnt-lap001
Assistant-Session: 3507023f-e0fd-4a1e-9d90-a0d4217d1502
2026-09-06 01:38:11 +02:00