canned-prompts/examples/pqrst-estimate/prompt.yaml
tegwick c5f1640454 CANP-WP-0002 T04: eval schema, render checks and output criteria
`evals/` was a reserved path holding unvalidated blobs: section 12 named the
directory and gave an illustrative snippet, but nothing was specified, so no
tool could act on an eval file.

Every eval file must now declare a `schema`, and CPF defines exactly one —
`canned-prompts/eval-rubric/v0.1`. Unrecognized schemas stay legal and are
skipped rather than rejected, so the format gains something actionable without
becoming an evaluation language, which remains a non-goal.

The schema splits along the same seam as T01 and T03. Render checks
(`contains`, `not_contains`, `resolves_all`) assert properties of the rendered
prompt text, need no model, and are therefore run by the reference CLI. Output
criteria describe a good result and are declared but not run, because judging
them requires a model. That division is now the format's consistent answer to
"deterministic locally, or not".

An eval references a fixture already declared in the manifest's `examples`
rather than carrying its own copy, so an example that is also an eval fixture
stays honest — both break together.

An eval declares assessment and must not record outcomes. Results are run
evidence and live outside the immutable package, per INTENT.md and section 17.

Spec: 12 rewritten with 12.1, 18 (rules 17-18), 21 (`eval` verb).

Reference CLI: read_eval, validate_eval, load_example_values,
run_render_checks, cmd_eval; a failed render check exits non-zero.
Tests 42 -> 51.

examples/pqrst-estimate/evals/quality.yaml is a real eval with four render
checks and four output criteria, and it passes.

Fixes a latent bug reaching a fixture exposed: coerce_value assumed every
value was a command-line string, so a YAML fixture carrying a real type
(include_rationale: true) crashed on .lower(). Typed values are now validated
but not re-parsed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Bjefh8NUiEiahN4JLwoSKM

Assistant: claude-code
Assistant-Model: opus
Assistant-Process: 388925@bnt-lap001
Assistant-Session: 3507023f-e0fd-4a1e-9d90-a0d4217d1502
2026-09-06 01:38:11 +02:00

53 lines
1.2 KiB
YAML

format: canned-prompt/v0.1
id: practice/pqrst-estimate
name: PQRST Estimate
version: 0.2.0
summary: 'Produce a post-session estimate of effort distributed across the PQRST categories
for an agentic coding session.
'
type: template
template: prompt.md
dependencies:
prompts:
- id: practice/house-style
version: '>= 0.1.0'
requirement: required
inputs:
- name: house_style
type: content
required: false
description: 'Shared house-style block, composed from practice/house-style.
'
default:
include: practice/house-style
- name: session_summary
type: content
required: true
description: 'Session transcript, summary, or sufficiently detailed account of the
work performed during the coding session.
'
parameters:
include_rationale:
type: boolean
default: true
description: Explain the evidence behind the estimate.
output:
format: markdown
description: A 100% PQRST effort allocation with concise interpretation.
compatibility:
capabilities:
- session-review
tags:
- pqrst
- retrospective
- agentic-coding
- effort-estimation
provenance:
author: canned-prompts seed
examples:
- examples/basic.yaml
evals:
- evals/quality.yaml