`evals/` was a reserved path holding unvalidated blobs: section 12 named the directory and gave an illustrative snippet, but nothing was specified, so no tool could act on an eval file. Every eval file must now declare a `schema`, and CPF defines exactly one — `canned-prompts/eval-rubric/v0.1`. Unrecognized schemas stay legal and are skipped rather than rejected, so the format gains something actionable without becoming an evaluation language, which remains a non-goal. The schema splits along the same seam as T01 and T03. Render checks (`contains`, `not_contains`, `resolves_all`) assert properties of the rendered prompt text, need no model, and are therefore run by the reference CLI. Output criteria describe a good result and are declared but not run, because judging them requires a model. That division is now the format's consistent answer to "deterministic locally, or not". An eval references a fixture already declared in the manifest's `examples` rather than carrying its own copy, so an example that is also an eval fixture stays honest — both break together. An eval declares assessment and must not record outcomes. Results are run evidence and live outside the immutable package, per INTENT.md and section 17. Spec: 12 rewritten with 12.1, 18 (rules 17-18), 21 (`eval` verb). Reference CLI: read_eval, validate_eval, load_example_values, run_render_checks, cmd_eval; a failed render check exits non-zero. Tests 42 -> 51. examples/pqrst-estimate/evals/quality.yaml is a real eval with four render checks and four output criteria, and it passes. Fixes a latent bug reaching a fixture exposed: coerce_value assumed every value was a command-line string, so a YAML fixture carrying a real type (include_rationale: true) crashed on .lower(). Typed values are now validated but not re-parsed. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Bjefh8NUiEiahN4JLwoSKM Assistant: claude-code Assistant-Model: opus Assistant-Process: 388925@bnt-lap001 Assistant-Session: 3507023f-e0fd-4a1e-9d90-a0d4217d1502
53 lines
1.2 KiB
YAML
53 lines
1.2 KiB
YAML
format: canned-prompt/v0.1
|
|
id: practice/pqrst-estimate
|
|
name: PQRST Estimate
|
|
version: 0.2.0
|
|
summary: 'Produce a post-session estimate of effort distributed across the PQRST categories
|
|
for an agentic coding session.
|
|
|
|
'
|
|
type: template
|
|
template: prompt.md
|
|
dependencies:
|
|
prompts:
|
|
- id: practice/house-style
|
|
version: '>= 0.1.0'
|
|
requirement: required
|
|
inputs:
|
|
- name: house_style
|
|
type: content
|
|
required: false
|
|
description: 'Shared house-style block, composed from practice/house-style.
|
|
|
|
'
|
|
default:
|
|
include: practice/house-style
|
|
- name: session_summary
|
|
type: content
|
|
required: true
|
|
description: 'Session transcript, summary, or sufficiently detailed account of the
|
|
work performed during the coding session.
|
|
|
|
'
|
|
parameters:
|
|
include_rationale:
|
|
type: boolean
|
|
default: true
|
|
description: Explain the evidence behind the estimate.
|
|
output:
|
|
format: markdown
|
|
description: A 100% PQRST effort allocation with concise interpretation.
|
|
compatibility:
|
|
capabilities:
|
|
- session-review
|
|
tags:
|
|
- pqrst
|
|
- retrospective
|
|
- agentic-coding
|
|
- effort-estimation
|
|
provenance:
|
|
author: canned-prompts seed
|
|
examples:
|
|
- examples/basic.yaml
|
|
evals:
|
|
- evals/quality.yaml
|