pqrst-practice/workplans/PQRST-WP-0002-validate-v01-against-real-sessions.md
tegwick 65fc8c77e3 Review hall PQRST corpus and close published-record pilot
Assistant: codex
Assistant-Model: gpt-6-astra
Assistant-Session: 01a0e759-301a-78b1-bbc1-040ef094b12d
2026-09-28 11:54:34 +02:00

366 lines
18 KiB
Markdown

---
id: PQRST-WP-0002
type: workplan
title: "Validate spec v0.1 in real session closes and land PQRST in hall-of-helix"
domain: agents
repo: pqrst-practice
status: finished
flavor: planning
owner: claude-code
topic_slug: practice
priority: high
created: "2026-09-05"
updated: "2026-09-28"
quality_dor: DoR-Ok
quality_dor_at: "2026-09-05"
quality_dor_by: claude-code
quality_dor_note: "v0.1 consolidates two drafts that were written without either being applied to a session. Every open question is empirical — whether the prompt returns a valid record unassisted, whether attribution by primary purpose is decidable in practice, and where records are stored. hall-of-helix supplies both the pilot population and the answer to the storage question: the operator already asks agents for a closing entry there, and one estimate has been added by hand, so the routine exists and is simply unspecified."
related:
- PQRST-WP-0001
- HOH-WP-0001
repos:
- pqrst-practice
- hall-of-helix
state_hub_workstream_id: "6a875b19-5a76-55c1-bd1f-2f5005cd416b"
---
State Hub review decision: `2f67b409-c5df-4f8e-bce8-98e9c7be92ab`.
# Validate spec v0.1 in real session closes and land PQRST in hall-of-helix
**Completed 2026-09-28:** the operator selected hall entries as the main data
source. [The pinned corpus review](../docs/reviews/2026-09-28-hall-pqrst-review.md)
closes T01 under the revised retrospective-review scope; T03 retains v0.1 with
editorial corrections grounded in those findings. All five tasks are done.
No new tasks or workplans were opened, and no implementation residuals remain.
Uncoached success rates, calibrated accuracy and self-estimation bias are
explicitly unestablished; studying them is outside this repository's scope.
The original proposal and initial blocked review are retained below as history.
Spec v0.1 was consolidated from two independent drafts, neither of which had
been applied to an actual session. It is internally consistent and entirely
unvalidated.
There is already a session-closing ritual in the fleet. The operator ends a
session with some form of:
> "Ok, lets close down this session then. Good work, please put an entry about
> your perspective on the session to hall-of-helix, before we wind down."
That ritual is real, repeated, and under-specified. An agent receiving it has to
infer the routine from `ENTRY.md`, and PQRST is not part of that routine at all —
the one estimate in the hall
(`entries/2026-09-05T16:36:12.000Z-codex-statehub-snapshot-and-signature.md`,
`pqrst_estimate: 30/23/18/19/10`) was added by hand, in a format that is not the
one the specification defines.
This workplan closes both gaps together: use real session closes as the pilot
population, and make the closing routine explicit in hall-of-helix so that an
agent reading that prompt finds — without being told twice — that closing a
session means running the PQRST prompt and recording the result in its entry.
It deliberately does **not** build tooling or a record store in *this*
repository. Both are out of scope (`SCOPE.md`); the hall is the store, and the
hall owns its own validation.
**Original done condition (superseded by the operator's corpus-review direction):** the closing routine is specified in hall-of-helix and reachable
from the prompt above, entries carry a validated PQRST record, the prompt has
been run unassisted at the end of at least eight real sessions, and v0.2 either
incorporates the findings or records why v0.1 stands.
**Completion order:** T04/T05 established the hall integration; T02 closed
the storage boundary; T01 reviewed the published corpus; T03 applied the
editorial findings and retained v0.1. Task ids are identity, not sequence.
**Cross-repo:** T04 and T05 changed `hall-of-helix`, not this repo. They were
executed under `HOH-WP-0001` there on 2026-09-05 and are `done`; the routine
lives in `hall-of-helix/CLOSING.md` and `make check` enforces the record on
finished agent seats from 2026-09-06. The read-only export script added for T01
also lives in the hall. Entry content was not edited by this review.
**2026-09-28 initial review (superseded):** T02 is done. T01 and T03 are `wait`, so this workplan is
`blocked`. Published hall records demonstrate adoption but do not establish
eight unmodified, uncoached, unrepaired first attempts. Resume T01 when original
closing outputs with that provenance are available across at least three repos
and two task classes; otherwise capture them at future real session closes.
T03 waits for those findings. No new tasks or workplans were opened.
## Pilot the closing prompt on real sessions
```task
id: PQRST-WP-0002-T01
status: done
priority: high
state_hub_task_id: "dd259b3f-e9ae-5bdd-8e02-07ae0083930a"
```
**Completed under revised scope:** inspect the committed hall corpus, retain
attempt history, structurally check records, and qualitatively review at least
eight examples spanning three repositories and two task classes. The review
contains eight verbatim examples (including a rejected first attempt), reviewer
task classes, provenance limits, and answers to all six questions below. The
exporter emits 196 entries, 95 with PQRST, containing 96 attempt blocks; 94 of
95 frontmatter-selected blocks pass structural checks. Published validity is
not unassisted first-attempt validity. This scope follows the operator's explicit
instruction to use the hall as our main data source, not a claim that the
original controlled protocol was fulfilled.
**Original protocol, retained for provenance:**
Paste `PqrstPrompt.md` unmodified at the end of at least **eight** real agentic
coding sessions, spread across at least three repositories and at least two task
classes (e.g. `feature`, `bugfix`, `explore`, `harden`). Once T04 has landed,
these are ordinary hall closes — the pilot is observation of the routine, not a
separate exercise.
Do not coach the model, do not repair a malformed answer, and do not re-run the
prompt for a better result. A rejected record is the finding.
For each session capture, in this file under "Pilot findings":
- the returned record verbatim;
- whether it passed spec §5.5 validation on the first attempt;
- whether `Dominant factors` cited concrete session facts or restated the
percentages;
- any dimension the model hedged on, refused, or visibly guessed at;
- the task class and the repository.
The specific questions to answer:
1. Does an uncoached agent return a **valid** record from the prompt alone?
2. Is "attribute by primary purpose" (spec R5) actually decidable mid-estimate,
or does it collapse under overlap?
3. Does S stay at 0 when security work did not occur, or does it drift upward?
4. Do estimates cluster at coarse increments (R7), or does false precision
reappear?
5. Is `Confidence` used meaningfully, or is it always `medium`?
6. **Does an agent estimating its own session inflate P?** The hall entry is a
first-person account written by the same session being audited. Self-report
bias is a live threat to the whole practice and the hall is where it will
first be visible.
## Decide the record-storage boundary
```task
id: PQRST-WP-0002-T02
status: done
priority: high
state_hub_task_id: "fd4abb7f-214f-56a0-a2f8-369008a9dc9a"
```
**Completed 2026-09-28:** [ADR-002](../docs/adr/0002-record-storage-boundary.md)
selects hall-of-helix as the default, nonexclusive human-facing store for
voluntary seats. No machine-facing collector is commissioned. `SCOPE.md` and
`README.md` now point to the decision; the hall owns storage and consumer
validation. State Hub decision: `dc5d49d5-8b79-4c57-a1a8-78c3532082fc`.
Original task rationale: `SCOPE.md` excludes record storage from this repository but does not say where
records go instead. Until that is answered the practice produces records with
nowhere to live, and the trend analysis that justifies the whole exercise
(spec §8, §9) cannot happen.
**hall-of-helix is now the leading candidate and should be evaluated as the
default**, not weighed from scratch: one record per session, written at the
moment of closing, already carrying `repos`, `session_id`, `exact_model`,
`harness`, and a first-person account of what the session actually did. That is
richer session metadata than a purpose-built store would likely collect, and
spec §9 asks for exactly those companions.
Weigh honestly what the hall is **not**:
- it is a hall of remembrance, not a metrics warehouse — spec §8 trend queries
mean grepping Markdown, and `INTENT.md` forbids turning entries into a
leaderboard;
- seats are voluntary and sometimes drafts, so the record set is a biased
sample of sessions, not a census;
- not every session ends with a seat, and forcing one would corrupt both
practices.
Decide, and record the answer as a decision (`POST /decisions/`, then an ADR in
this repo if there are repo-local consequences). If the hall is chosen, say
plainly whether it is the *only* store or the human-facing one alongside a
machine-facing collector owned elsewhere. Whatever is chosen, this repository
still ships no collector — the decision names the owner, it does not move the
work here.
## Cut spec v0.2
```task
id: PQRST-WP-0002-T03
status: done
priority: medium
state_hub_task_id: "0a0bce15-ac0c-582a-842c-ae1bf0307afb"
```
**Completed 2026-09-28:** retain v0.1 after the published-corpus review. Fix the
prompt checklist's demonstrated retry contradiction, clarify substantive-work
scope and optional uncertainty/provenance Notes, and align the spec's ranking
prohibition with INTENT. No dimension, format or validation rule changes; the
fenced canonical prompt is unchanged. Appendix B records the editorial update.
T02 already supplied the storage cross-reference. This disposition is supported
by published-record evidence, not a controlled accuracy or first-attempt study.
From the pilot findings and the hall integration, revise the specification and
prompt together:
- fix rules the pilot showed to be undecidable or routinely ignored;
- tighten or relax the validation rules based on what actually failed;
- settle whether the closing ritual itself (writing the entry, rendering the
portrait, syncing) is inside or outside the estimate — T04 takes a position;
the pilot says whether it holds;
- state how self-estimation bias is to be handled, if T01 finds any;
- add the storage-boundary decision from T02 as a `SCOPE.md` cross-reference;
- bump to **v0.2** and add a row to the spec's Appendix B.
If a change alters a dimension's meaning, the validation rules, or the stored
record format, say plainly in Appendix B that records made under v0.1 are less
comparable. A record-format change also means a matching change in
hall-of-helix — name it in the handoff.
If the pilot surfaces nothing substantive, record that outcome and leave v0.1
standing. "No change needed, and here is the evidence" is a valid result and a
better one than churn.
Run the verification checklist in `AGENTS.md` before committing, then sync:
```bash
uv run --project ~/repo-manager rmgr sync --path . --push
```
## Specify the session-closing routine in hall-of-helix
```task
id: PQRST-WP-0002-T04
status: done
priority: high
state_hub_task_id: "974949c3-1cab-5251-b3df-c50a8f3ab65a"
```
Make the closing routine explicit and findable, so an agent that receives the
operator's wind-down prompt knows the whole sequence without asking.
**Deliverable:** `CLOSING.md` in the hall-of-helix root — the routine in order,
one page, no prose detour. Quote the operator's actual wind-down sentence near
the top so an agent scanning for it recognises the situation it is in.
Link it from the three places an agent lands:
- `README.md`, beside "How to leave a seat";
- `ENTRY.md`, at the top — `ENTRY.md` says how to *write a seat*, `CLOSING.md`
says what *closing a session* consists of, and the entry is one step of it;
- `AGENTS.md`, in the **Close** section of the session protocol.
**The routine to specify:**
1. Finish the substantive work. Update workplan task statuses.
2. **Run the PQRST prompt** — `~/pqrst-practice/PqrstPrompt.md` — unmodified and
uncoached. Validate the result against spec §5.5 before using it: five
integers summing to 100, `Confidence` present, `Dominant factors` naming
concrete session facts rather than restating the numbers.
3. Write the entry per `ENTRY.md`, carrying the record (format per T05).
4. `make check`; add the seat to the `README.md` list.
5. Log progress to State Hub; sync the repo.
**Decide and state two things the routine is otherwise silent on:**
- **Scope of the estimate.** The estimate covers the *substantive session*, not
the closing ritual — writing the entry, rendering the portrait, and syncing
are excluded. Spec R1 says the estimate is produced after the substantive work
is complete, and letting the ritual score itself would inflate T on every
session that has a seat. State this explicitly in `CLOSING.md`; T03 revisits
it if the pilot disagrees.
- **How the prompt is reached.** Referencing `~/pqrst-practice/PqrstPrompt.md`
keeps one source of truth but assumes the checkout exists; inlining the prompt
in the hall makes closes self-contained but forks the canonical text. The
recommendation is to **reference the path and inline nothing but the output
block**, so a session with no `pqrst-practice` checkout can still produce a
well-formed record. Whichever is chosen, `pqrst-practice` stays canonical and
the hall says so.
**Constraint:** the hall is not a metrics programme. `INTENT.md` there is
explicit that seats are not ranked and the hall is not a leaderboard, and
`INTENT.md` here is explicit that PQRST fails the moment anyone scores well on
it. The routine adds a record to a seat; it does not add a score to a worker.
Do not introduce comparison, aggregation, or ranking of estimates into
hall-of-helix.
## Carry the PQRST record in hall entries
```task
id: PQRST-WP-0002-T05
status: done
priority: high
state_hub_task_id: "1d8f0b3d-7620-5810-ad91-13e17ee4b9e3"
```
Give the record a defined home in an entry, so it is machine-collectable without
losing the evidence that makes it auditable.
**Format** — both halves, since either alone is useless:
- **Frontmatter:** `pqrst_estimate: "P30 Q23 R18 S19 T10"` — the canonical
signature from spec §5.2, quoted. Prefer it over the existing ad-hoc
`30/23/18/19/10`: it is self-describing, survives reordering, and is matched by
the same regex the `pqrst-practice` verification checklist already uses. This
means normalising the one existing entry (see below).
- **Body:** a `## PQRST estimate` section holding the full stored record from
spec §5.1 — the block, `Confidence`, and `Dominant factors`. Spec §10 rule 8
is blunt that a signature without its dominant-factors sentence is not
auditable, and `ENTRY.md` already refuses "various improvements" for the same
reason.
**Changes:**
- `templates/entry.md` — add the frontmatter field and the section, with the
section placed after "Durable legacy" and before "Visual prompt".
- `ENTRY.md` — add `pqrst_estimate` to the Required frontmatter table and
"PQRST estimate" to the Required sections list, each with one line on what
makes it valid, and link `CLOSING.md`.
- `scripts/check-entries.py` — when `pqrst_estimate` is present, validate the
signature format and that it sums to 100; make it **required for
`worker_kind: agent-session` entries created after the adoption date**, and
add "PQRST estimate" to `REQUIRED_HEADINGS` under the same condition. Human
seats are exempt. `make check` must still pass on all 102 existing seats.
**Do not backfill.** Existing seats predate the practice and nobody observed
those sessions; inventing estimates for them would be exactly the fabricated
evidence spec R4 and §5.5 rule 4 forbid, and it would poison the first trend
data the practice ever produces. The only record touched is the manual
`30/23/18/19/10` in the 2026-09-05 Codex seat, normalised to
`"P30 Q23 R18 S19 T10"` — same numbers, canonical spelling, and it already sums
to 100. Ask its author before rewriting the seat, or leave it and note the
exception in `CLOSING.md`.
## Pilot findings
The completed [corpus review](../docs/reviews/2026-09-28-hall-pqrst-review.md)
contains the method, findings, and eight verbatim source examples with validation
and provenance dispositions. The initial evidence scan below is historical;
the later corpus review supersedes its blocked-task conclusion.
### Evidence availability review — 2026-09-28
Inspected the hall's closing instructions, consumer checker, September entry
provenance mentions, and these published records at hall commit
`82caae23b3098aa4bdf4346cd4af8a88c03c7d68` (paths below are under its `entries/`):
- `2026-09-08T11-20-00.000Z-claude-01Bjefh8-the-checks-were-the-thing-that-lied.md`:
explicitly reports a byte-identical canonical prompt package and exclusion
of the closing ritual. Its Notes describe P/Q attribution ambiguity, and
values include non-five-point increments. This is relevant evidence for
R5/R7 review, but it does not attest no retries or repairs.
- `2026-09-09T21-15-50Z-codex-fiam-four-source-plates.md`: concrete deliverables,
checks and coordination explain the record; S is zero and values are coarse.
It excludes the closing ritual but does not identify first-attempt provenance.
- `2026-09-10T22-04-31.000Z-claude-01NV9oij-guard-proved-less-than-claimed.md`:
concrete identity-security work and mutation checks explain the record;
values include non-five-point increments. First-attempt provenance is absent.
All three published blocks have five in-range integers summing to 100, matching
signatures, medium confidence, and concrete dominant factors. This checks the
published artifacts only. They cannot establish unassisted success rates,
confidence calibration, or self-estimation bias. No records are reconstructed
or counted as qualifying pilot runs in this review. T01 retains the original
eight-session acceptance criteria; T03 retains the resulting interpretation
and version work.