Review hall PQRST corpus and close published-record pilot

Assistant: codex
Assistant-Model: gpt-6-astra
Assistant-Session: 01a0e759-301a-78b1-bbc1-040ef094b12d
This commit is contained in:
tegwick 2026-09-28 11:54:34 +02:00
parent 37b49a9517
commit 65fc8c77e3
7 changed files with 413 additions and 18 deletions

View file

@ -138,6 +138,12 @@ Some T is productive coordination. Excessive or repeated T may signal unstable s
**R5 — Attribute overlapping work by primary purpose.** Categories overlap by design. Assign effort by the *primary reason the activity was undertaken at that moment*, and do not double-count.
This is a judgment, not a uniquely recoverable split. When another attribution
would materially change the estimate, use optional `Notes` to name the boundary
(for example P/S) and lower confidence as appropriate. Security-specific analysis
can count as S without changing a control or handling a credential; merely reading
that a task waits on credentials may instead be R or T, depending on purpose.
| Activity | Category |
| --- | --- |
| Reading a module to understand how it works | **R** |
@ -163,6 +169,15 @@ Some T is productive coordination. Excessive or repeated T may signal unstable s
Run the estimate **once**, at the natural end of a session or a clearly bounded work unit — one ticket, one vertical slice, one deliberate "stop here" checkpoint.
The substantive work is the unit being estimated. The subsequent hall closing
ritual — writing the seat, rendering its portrait, and final hall sync — is
excluded. Coordination and verification that belong to the substantive task
remain included. Preserve a rejected output with its reason rather than silently
repairing it or rerunning for a better result. Existing retries remain separate,
ordered attempts, not additional sessions. Optional `Notes` can disclose prior
allocation hints or limited session context; published validity does not prove
an uncoached first attempt.
Do not collect mid-stream unless closing a phase on purpose. Mid-session estimates mix unfinished work with planning residue and are hard to compare.
If a session spanned several distinct modes (explore, then implement, then harden), emit **one overall estimate** plus optional phase notes. The stored record remains a single 5-tuple unless phases are explicitly versioned.
@ -343,7 +358,7 @@ Questions the joined data can answer:
7. **Compare like with like.** An explore session and a one-line fix do not share a healthy template.
8. **Store the signature line plus the dominant-factors sentence.** Numbers without the sentence are not auditable.
9. **Preserve the rationales, not only the percentages.**
10. **Do not rank developers, agents, or teams on raw percentages** without additional outcome evidence.
10. **Do not rank developers, agents, models, or teams on PQRST percentages.** Outcome evidence supports process review, not a worker leaderboard.
11. **Do not discuss a stored estimate during the next session** unless changing process on purpose.
---
@ -408,3 +423,4 @@ At session end:
| Version | Date | Change |
| --- | --- | --- |
| 0.1 | 2026-09-05 | Initial consolidation of the two independent 2026-09-05 drafts into one specification. The block/signature record was adopted as the source of truth; the per-dimension table was retained as an optional presentation form. |
| 0.1 (editorial) | 2026-09-28 | [Hall review](../docs/reviews/2026-09-28-hall-pqrst-review.md): retain the dimensions, format and validation rules; align the prompt checklist with once-only collection, clarify the existing substantive-work boundary and optional uncertainty/provenance notes, and align the anti-ranking rule with INTENT. Published records do not establish first-attempt success or calibrated accuracy. |