Review hall PQRST corpus and close published-record pilot

Assistant: codex
Assistant-Model: gpt-6-astra
Assistant-Session: 01a0e759-301a-78b1-bbc1-040ef094b12d
This commit is contained in:
tegwick 2026-09-28 11:54:34 +02:00
parent 37b49a9517
commit 65fc8c77e3
7 changed files with 413 additions and 18 deletions

View file

@ -6,6 +6,10 @@ Specification: [`spec/PqrstEstimationPractice.md`](spec/PqrstEstimationPractice.
Do not add scoring dimensions inside PQRST. Do not run this prompt mid-session unless deliberately closing a phase.
Estimate the substantive work before the closing ritual. Writing the hall entry,
rendering its portrait, and the final hall sync are excluded. Coordination and
verification performed as part of the substantive task still count.
---
```text
@ -73,7 +77,7 @@ After the block, also emit the equivalent YAML record with keys p, q, r, s, t, t
## Accepting or rejecting a result
Reject and re-run if any of the following is true:
Reject the result as a valid record if any of the following is true:
- the five values are not integers summing to 100;
- `Confidence` is missing or not one of `low` / `medium` / `high`;
@ -81,3 +85,9 @@ Reject and re-run if any of the following is true:
- `Dominant factors` merely restates the numbers or says something like "mixed work across several areas";
- a sixth dimension has been introduced inside the 5-tuple;
- S is non-zero but no security-specific work actually occurred.
Preserve the original output and the rejection reason. Do not silently repair
it or re-run for a better answer; a rejected first attempt is useful evidence.
If an existing record already has a retry, retain both and label their order.
Use optional `Notes` to preserve attribution uncertainty, limited session
context, or any prior allocation hint; do not call those records uncoached.

View file

@ -51,3 +51,6 @@ validators, dashboards, and record storage belong elsewhere — see
When closing with a voluntary hall-of-helix entry, retain the record there.
The hall is the default human-facing store, not a mandatory or exclusive one;
see [the storage decision](docs/adr/0002-record-storage-boundary.md).
The [2026-09-28 hall review](docs/reviews/2026-09-28-hall-pqrst-review.md)
documents the extraction command, corpus findings, and why v0.1 stands.

View file

@ -47,6 +47,12 @@ or commissioned. Seats are a biased sample, not a session census; no seat is
required just to retain an estimate. See
[ADR-002 — Record storage boundary](docs/adr/0002-record-storage-boundary.md).
`hall-of-helix/entries` is the main data source for this practice's record-format
and prompt reviews. Its read-only `scripts/export-pqrst.py` exports a pinned Git
snapshot for analysis elsewhere; the hall retains the originals. Bounded review
excerpts and findings may accompany a workplan here without becoming a session
store or an empirical accuracy study.
### Integration with any particular agent or harness
PQRST is deliberately agent-agnostic and model-agnostic. Wiring it into a specific harness, plugin, skill, or CI pipeline belongs to that harness's own repository.

View file

@ -39,5 +39,12 @@ validation rule, or record format changes, so v0.1 remains in force. The
eight-session first-attempt pilot and the subsequent version decision remain
under the existing T01 and T03; this decision does not establish their outcome.
**2026-09-28 follow-up:** the operator selected hall entries as the main data
source. T01 now reviews that published corpus, with first-attempt uncertainty
explicit rather than a completion gate. The hall owns a read-only Git-snapshot
exporter; no machine-facing collection service or second record store is
commissioned. The [review](../reviews/2026-09-28-hall-pqrst-review.md) closes T01
and supplies T03's decision to retain v0.1 with editorial alignment.
See [SCOPE.md](../../SCOPE.md) and the
[existing workplan](../../workplans/PQRST-WP-0002-validate-v01-against-real-sessions.md).

View file

@ -0,0 +1,323 @@
# Hall PQRST review — 2026-09-28
State Hub review decision: `2f67b409-c5df-4f8e-bce8-98e9c7be92ab`.
## Source and method
The operator selected `hall-of-helix/entries` as the main data source for
PQRST-WP-0002 on 2026-09-28. This replaces the original eight-session controlled
first-attempt pilot with a retrospective review of published records. It does
not relabel repaired or coached estimates as unassisted. The original pilot's
success-rate and bias questions remain unanswerable from this corpus and are
not release gates for the specification-and-prompt repository.
Reviewed hall commit `82caae23b3098aa4bdf4346cd4af8a88c03c7d68`, covering entries
recorded through 2026-09-28. Six local entry edits were excluded by reading Git
objects rather than working files. The hall remains the record store; examples
below are a bounded review sample, not a second session corpus.
The standard-library exporter lives in `hall-of-helix/scripts/export-pqrst.py`.
Run from a workstation with that checkout:
```bash
python3 ~/hall-of-helix/scripts/export-pqrst.py ~/hall-of-helix \
--revision 82caae23b3098aa4bdf4346cd4af8a88c03c7d68 > /tmp/hall-pqrst.jsonl
```
Omit `--revision` to review committed HEAD next time. Each JSONL row preserves
source commit, entry path, content SHA-256, raw frontmatter, all PQRST sections,
and every fenced record in order. Each block includes raw text and a field map;
field values are lists so duplicate fields remain visible. Missing sections
produce empty lists, not silently dropped seats. The exporter does not certify
validity, select a successful retry, rank workers, or modify entries.
For this review, parse the raw frontmatter as YAML, select the block whose
Signature matches `pqrst_estimate`, and retain every other block as attempt
history. Check five unique integer percentage fields in 0..100, sum 100, a
matching numeric signature (leading zeros allowed), one allowed confidence,
and present dominant factors. Check Sum and frontmatter/body agreement too.
Concrete evidence and purpose attribution require human reading; a nonempty
Dominant factors field alone cannot establish spec §5.5 validity.
## Corpus findings
| Observation | Count and denominator |
| --- | --- |
| Committed entries | 196: 195 agent-session, 1 human |
| Entries carrying PQRST | 95: 58 draft, 32 handed-forward, 5 complete |
| Fenced PQRST blocks | 96 across those 95 entries |
| Unique block matching frontmatter | 95/95 entries |
| Matching blocks passing structural checks | 94/95 |
| Confidence on matching blocks | 91 medium, 3 high, 0 low, 1 absent |
| All five values in multiples of five | 87/94 structurally passing matching blocks |
| S explicitly zero | 29/94 structurally passing matching blocks |
| Repository labels represented | 73, including the hall itself; multi-repo sessions overlap |
| Explicit task-class/outcome frontmatter fields | 0/95 |
All 101 entries without PQRST predate the 2026-09-06 adoption boundary by
`recorded_at`; they are not failures to comply with the adopted routine.
The 95 records span 2026-09-05 through 2026-09-28. Drafts are included because
these are already published estimates; portrait status is not estimate quality.
No model/worker comparison, performance score, or average effort target is
inferred from these counts.
The lone structurally incomplete selected record is the grandfathered
`2026-09-05T16:36:12.000Z-codex-statehub-snapshot-and-signature.md`. It lacks
Confidence and Dominant factors. Its hall note calls it valid, but that means
less than spec §5.5: it is incomplete under that contract. Preserve it as
historical evidence; do not invent missing rationale.
`2026-09-22T07-59-35.000Z-claude-62534cdf-warnings-were-the-work.md` contains
an invalid first block and a valid retry. The original has Q=15 in the values
but Q=30 in the signature, and no Dominant factors. Flattening the section into
one dictionary would mix the two attempts and corrupt the result. The author's
explanation directly identifies the prompt's “reject and re-run” checklist as
the reason for retrying, contradicting the spec's once-only collection rule and
the hall's preserve-failures routine. This is a demonstrated instruction defect.
## Answers to the pilot questions
1. **Uncoached first-attempt validity:** not measurable from published seats.
There is one explicitly failed first attempt, and another record explicitly
discloses an allocation hint in a continuation summary (sample E). Therefore
94/95 is published structural completeness, not a first-attempt success rate.
2. **Primary-purpose attribution:** usable as a judgment, not uniquely decidable.
Samples A and H explain P/Q and P/S ambiguity. Sample H says a different
defensible classification would reverse much of P and S. Keep R5 and preserve
that explanation in Notes; do not interpret a small percentage difference
as a measured change in effort.
3. **Zero security:** 29 structurally passing records use S=0. Samples B and G
ground this in the work described; G distinguishes reading about a credential
gate from doing security work. This shows the zero is usable, not that every
security allocation is accurate. Security-specific analysis can count as S
without a code or credential change; classify by purpose, not write activity.
4. **Coarse increments:** 87/94 are all multiples of five. The seven exceptions
are not validation failures: R7 is a preference. Samples A, C and H show
finer splits without evidence that 1% resolution is reliable. Keep the coarse
preference; do not round stored originals or make coarse increments mandatory.
5. **Confidence:** medium dominates (91/94 present values); high appears three
times and low never appears. The field is not constant, but these observations
cannot establish calibration. Existing Notes can explain uncertainty; no new
confidence dimension, required field, or numeric confidence score is needed.
6. **Self-estimation bias:** cannot be measured without independent observations
or estimates. Self-reports and their narratives are not independent evidence
against P inflation. Acknowledge that limit; do not launch an accuracy study
or rank agents in this repository.
Task classes below are reviewer-assigned from Contribution, not recovered
metadata. The eight deliberately selected examples cover feature, hardening,
and exploration work across more than three repositories, and include failures
and provenance limitations. They are diagnostic examples, not a random sample.
Each passing sample's Dominant factors names concrete artifacts or actions
consistent with its Contribution. This is a qualitative check on these eight,
not semantic validation of all 95 records or independent session verification.
## Version disposition
Keep **v0.1**. Neither five-dimensional meaning nor record format nor §5.5
validation needs a breaking change on this evidence. Correct the contradictory
retry instruction; clarify the already implemented substantive-work boundary
and preservation of uncertainty/provenance in optional Notes. Preserve the
existing hall exclusion of seat writing, portrait generation and final hall
sync; substantive coordination and verification before closure still count.
Make the anti-ranking wording consistent with INTENT's unconditional ban.
These are editorial alignment changes, recorded in Appendix B. They do not
make historic records newly comparable or certify their original collection
conditions. Existing copies of the fenced canonical prompt remain unchanged.
The hall exporter is the simple read-only extraction path; no collector,
validator service, dashboard, or second store is introduced here.
## Eight source examples
Blocks below are copied verbatim from the pinned source, including the invalid
attempt. Do not apply the illustrative-signature “sum to 100” smoke check to
that deliberately retained counterexample. Sources are paths under the hall's
`entries/` at the commit above.
### A — checks-were-the-thing-that-lied
Source: `entries/2026-09-08T11-20-00.000Z-claude-01Bjefh8-the-checks-were-the-thing-that-lied.md`
Repositories: canned-prompts, rapp-canned-prompts, helix-forge, rapp-postgres
Reviewer task class: feature / harden
Published block passes §5.5 on review; first-attempt status unknown. Canonical package provenance is stated. Notes explicitly hedge P/Q attribution.
```text
PQRST-Estimate
P: 35%
Q: 30%
R: 18%
S: 12%
T: 5%
Sum: 100%
Confidence: medium
Signature: P35 Q30 R18 S12 T5
Dominant factors: Three things were built end to end — the CPF v0.2 format decisions with their reference implementation, the hosted registry service, and its Railiance deployment — and the debugging that followed was nearly as large: a migration that logged every revision as applied while silently rolling back, an egress policy that never selected the migration Job, and five defects in my own verification instruments, three of them in the lease watcher before it could return an honest verdict. Security was substantive rather than incidental: publisher token design (hashed, constant-time, non-enumerable, revocation preserving attribution), replacing env-var credentials with mounted files, and surviving credential-lease rotation.
Notes: Attribution across a session this long is approximate; the P/Q split in particular could defensibly shift several points either way, since debugging a silent rollback is both implementing the deliverable and establishing its correctness.
```
### B — fiam-four-source-plates
Source: `entries/2026-09-09T21-15-50Z-codex-fiam-four-source-plates.md`
Repositories: facetted-interfaces, hall-of-helix
Reviewer task class: explore / harden
Published block passes §5.5 on review; first-attempt status unknown. Coarse values, zero S, and substantive-work boundary are explicit; the model handoff limits session provenance.
```text
PQRST-Estimate
P: 30%
Q: 25%
R: 25%
S: 0%
T: 20%
Sum: 100%
Confidence: medium
Signature: P30 Q25 R25 S0 T20
Dominant factors: Repository intent and the InterfaceCanon acceptance receipt were the main deliverables; Draft 6 conformance, interpreter debugging, source-hash checks, and reading the governance and FIAM material accounted for much of the remaining work. Workplan creation, registration, and projection reconciliation required substantial coordination.
Notes: Covers the substantive conversation across the model handoff; excludes this closing ritual.
```
### C — guard-proved-less-than-claimed
Source: `entries/2026-09-10T22-04-31.000Z-claude-01NV9oij-guard-proved-less-than-claimed.md`
Repositories: key-cape
Reviewer task class: harden
Published block passes §5.5 on review; first-attempt status unknown. Security and verification have concrete anchors; non-five-point values are not prohibited. No explicit uncertainty explanation.
```text
PQRST-Estimate
P: 22%
Q: 22%
R: 13%
S: 30%
T: 13%
Sum: 100%
Confidence: medium
Signature: P22 Q22 R13 S30 T13
Dominant factors: The bulk of the session was identity-security work in an issuer — authorization-code redirect/grant/client binding with atomic code consumption, upstream Authelia ID-token verification via a new strict internal/jose, tenant provenance (tenant_source) under GH-DEC-2026-013, bind passwords off argv with 0600 artifacts, and two escalation guards (dynamic registration, principal_type) — with a matching test burden where each fix was verified by mutation rather than assertion.
Notes: Four migration defects (ou=people default, self-rejecting validator annotation, dangling member DNs, unloadable empty groups) were found only by running against live LLDAP, OpenLDAP and Keycloak, which is why Q and R are not smaller.
```
### D — warnings-were-the-work
Source: `entries/2026-09-22T07-59-35.000Z-claude-62534cdf-warnings-were-the-work.md`
Repositories: railiance-fabric, hall-of-helix
Reviewer task class: feature / maintenance
First attempt explicitly fails §5.5: mismatched signature and missing dominant factors. The matching published retry passes on review. Both attempts are preserved below; never count them as two sessions.
```text
PQRST-Estimate
P: 30%
Q: 15%
R: 25%
S: 0%
T: 30%
Sum: 100%
Confidence: medium
Signature: P30 Q30 R25 S0 T30
```
```text
PQRST-Estimate
P: 30%
Q: 15%
R: 25%
S: 0%
T: 30%
Sum: 100%
Confidence: medium
Signature: P30 Q15 R25 S0 T30
Dominant factors: T came from four fix-consistency passes, splitting the work into commits, the flex-auth handoff (the RAIL-FAB-WP-0031 intake plus the reply), and re-sequencing the WP-0028 tasks; P came from chokepoint sizing in coordination_graph.py and the C-35/C-23 fixes; R came from reading the inherited diffs, the 3,000-line explorer UI, repo-manager's classification rules, and the orientation notice.
Notes: S is 0 because I read the credential and production traps in the orientation notice but did no security work.
```
### E — memo-version-three
Source: `entries/2026-09-27T13-49-49Z-grok-memo-version-three.md`
Repositories: secrets-engine, flex-auth, railiance-platform, railiance-clock, informed-decision
Reviewer task class: harden / operations
Published block passes §5.5 on review. Uncoached provenance explicitly fails: Notes disclose an allocation hint from a continuation summary. Retry status unknown.
```text
PQRST-Estimate
P: 25%
Q: 15%
R: 30%
S: 20%
T: 10%
Sum: 100%
Confidence: medium
Signature: P25 Q15 R30 S20 T10
Dominant factors: Most of the stretch was spent separating Railiance Clock trust from the guest NTP skew, and separating the two informed-decision review images from the lifecycle PDP flex-auth-secrets-engine, including the Forgejo digest and the worker reports. The security slice was three new approval ids so the consumed 2026-09-16 approvals stayed consumed, a fail-closed skip when the live AppRole already matched, and the withdrawn NTP hold; the delivered change was memo version 3 and helm revision 4 of the two review releases.
Notes: A continuation summary suggested a research-heavy, security-substantial split before this estimate was written. These integers were chosen after that hint and give more weight to the publish-and-roll than the hint did. The closing ritual is excluded.
```
### F — statement-that-never-became-an-invoice
Source: `entries/2026-09-27T19-55-40Z-claude-286d235a-the-statement-that-never-became-an-invoice.md`
Repositories: fin-hub
Reviewer task class: feature
Published block passes §5.5 on review; first-attempt status unknown. A high-confidence example with concrete schema, store, CLI and test work; confidence calibration is not established.
```text
PQRST-Estimate
P: 35%
Q: 25%
R: 25%
S: 0%
T: 15%
Sum: 100%
Confidence: high
Signature: P35 Q25 R25 S0 T15
Dominant factors: Most effort went into reading resource-control's settlement-statement schema and terms document, then fin-hub's existing exchange.py/schemas conventions, before writing the SettlementStatement model, its store, and the CLI wiring (R and P in close proportion); a comparable share went to writing and running four new idempotency/health tests plus two full-suite passes (Q); the remainder was triaging which of three open workplans had real in-repo work left versus a genuine external blocker, plus recovering cleanly from a stray uv.lock/git-stash detour mid-session (T).
```
### G — gate-was-elsewhere
Source: `entries/2026-09-21T13-49-08.000Z-claude-816898b9-the-gate-was-elsewhere.md`
Repositories: sand-boxer
Reviewer task class: explore / maintenance
Published block passes §5.5 on review; first-attempt status unknown. Explicit zero S despite security-related task dependencies; context research is distinguished from security work.
```text
PQRST-Estimate
P: 30%
Q: 10%
R: 45%
S: 0%
T: 15%
Sum: 100%
Confidence: medium
Signature: P30 Q10 R45 S0 T15
Dominant factors: Research dominated — reading the two open workplans' evidence trail, cross-checking task status in the hub, and tracing C-24 into state-hub's validator and the canonical allowed-values file; the deliverable itself was two small edits (WP-0014 status, classification tags).
Notes: S is 0 — the open tasks concern credentials, but this session only read about those gates; no security work occurred.
```
### H — receipts-outlived-the-recollection
Source: `entries/2026-09-10T22-52-29.000Z-claude-01WLUjpv-receipts-outlived-the-recollection.md`
Repositories: railiance-platform
Reviewer task class: harden / operations
Published block passes §5.5 on review; first-attempt status unknown. Notes explicitly hedge P/S, offering a materially different defensible allocation; this is direct evidence of classification uncertainty.
```text
PQRST-Estimate
P: 17%
Q: 10%
R: 18%
S: 35%
T: 20%
Sum: 100%
Confidence: medium
Signature: P17 Q10 R18 S35 T20
Dominant factors: Nearly every deliverable was credential custody — two verifier CCRs with exact-path policies and namespace-limited ESO delivery, two client-side reader admissions, an attended read-only OpenBao runner with an ambient-token guard, and the field/duplicate finding at the whynot-design lane. Coordination was the next largest slice: eighteen owner replies across thirteen threads, several of which required reading sibling repos (secrets-engine, approval-engine, key-cape, ops-warden, net-kingdom) and the live cluster before an honest answer was possible.
Notes: The P/S boundary is the main uncertainty. Writing a CCR or a policy is both the requested deliverable and credential work; I classified by primary purpose per rule 4, which pushes S well above P. A reader who split it the other way would get roughly P35 S17 with the same session underneath.
```

View file

@ -138,6 +138,12 @@ Some T is productive coordination. Excessive or repeated T may signal unstable s
**R5 — Attribute overlapping work by primary purpose.** Categories overlap by design. Assign effort by the *primary reason the activity was undertaken at that moment*, and do not double-count.
This is a judgment, not a uniquely recoverable split. When another attribution
would materially change the estimate, use optional `Notes` to name the boundary
(for example P/S) and lower confidence as appropriate. Security-specific analysis
can count as S without changing a control or handling a credential; merely reading
that a task waits on credentials may instead be R or T, depending on purpose.
| Activity | Category |
| --- | --- |
| Reading a module to understand how it works | **R** |
@ -163,6 +169,15 @@ Some T is productive coordination. Excessive or repeated T may signal unstable s
Run the estimate **once**, at the natural end of a session or a clearly bounded work unit — one ticket, one vertical slice, one deliberate "stop here" checkpoint.
The substantive work is the unit being estimated. The subsequent hall closing
ritual — writing the seat, rendering its portrait, and final hall sync — is
excluded. Coordination and verification that belong to the substantive task
remain included. Preserve a rejected output with its reason rather than silently
repairing it or rerunning for a better result. Existing retries remain separate,
ordered attempts, not additional sessions. Optional `Notes` can disclose prior
allocation hints or limited session context; published validity does not prove
an uncoached first attempt.
Do not collect mid-stream unless closing a phase on purpose. Mid-session estimates mix unfinished work with planning residue and are hard to compare.
If a session spanned several distinct modes (explore, then implement, then harden), emit **one overall estimate** plus optional phase notes. The stored record remains a single 5-tuple unless phases are explicitly versioned.
@ -343,7 +358,7 @@ Questions the joined data can answer:
7. **Compare like with like.** An explore session and a one-line fix do not share a healthy template.
8. **Store the signature line plus the dominant-factors sentence.** Numbers without the sentence are not auditable.
9. **Preserve the rationales, not only the percentages.**
10. **Do not rank developers, agents, or teams on raw percentages** without additional outcome evidence.
10. **Do not rank developers, agents, models, or teams on PQRST percentages.** Outcome evidence supports process review, not a worker leaderboard.
11. **Do not discuss a stored estimate during the next session** unless changing process on purpose.
---
@ -408,3 +423,4 @@ At session end:
| Version | Date | Change |
| --- | --- | --- |
| 0.1 | 2026-09-05 | Initial consolidation of the two independent 2026-09-05 drafts into one specification. The block/signature record was adopted as the source of truth; the per-dimension table was retained as an optional presentation form. |
| 0.1 (editorial) | 2026-09-28 | [Hall review](../docs/reviews/2026-09-28-hall-pqrst-review.md): retain the dimensions, format and validation rules; align the prompt checklist with once-only collection, clarify the existing substantive-work boundary and optional uncertainty/provenance notes, and align the anti-ranking rule with INTENT. Published records do not establish first-attempt success or calibrated accuracy. |

View file

@ -4,7 +4,7 @@ type: workplan
title: "Validate spec v0.1 in real session closes and land PQRST in hall-of-helix"
domain: agents
repo: pqrst-practice
status: blocked
status: finished
flavor: planning
owner: claude-code
topic_slug: practice
@ -24,8 +24,20 @@ repos:
state_hub_workstream_id: "6a875b19-5a76-55c1-bd1f-2f5005cd416b"
---
State Hub review decision: `2f67b409-c5df-4f8e-bce8-98e9c7be92ab`.
# Validate spec v0.1 in real session closes and land PQRST in hall-of-helix
**Completed 2026-09-28:** the operator selected hall entries as the main data
source. [The pinned corpus review](../docs/reviews/2026-09-28-hall-pqrst-review.md)
closes T01 under the revised retrospective-review scope; T03 retains v0.1 with
editorial corrections grounded in those findings. All five tasks are done.
No new tasks or workplans were opened, and no implementation residuals remain.
Uncoached success rates, calibrated accuracy and self-estimation bias are
explicitly unestablished; studying them is outside this repository's scope.
The original proposal and initial blocked review are retained below as history.
Spec v0.1 was consolidated from two independent drafts, neither of which had
been applied to an actual session. It is internally consistent and entirely
unvalidated.
@ -52,23 +64,22 @@ It deliberately does **not** build tooling or a record store in *this*
repository. Both are out of scope (`SCOPE.md`); the hall is the store, and the
hall owns its own validation.
**Done when:** the closing routine is specified in hall-of-helix and reachable
**Original done condition (superseded by the operator's corpus-review direction):** the closing routine is specified in hall-of-helix and reachable
from the prompt above, entries carry a validated PQRST record, the prompt has
been run unassisted at the end of at least eight real sessions, and v0.2 either
incorporates the findings or records why v0.1 stands.
**Execution order:** T04/T05 are complete; T02 closed independently on
2026-09-28 using the existing hall integration. T01 must supply the pilot
evidence before T03 can settle the version decision. Task ids are identity,
not sequence.
**Completion order:** T04/T05 established the hall integration; T02 closed
the storage boundary; T01 reviewed the published corpus; T03 applied the
editorial findings and retained v0.1. Task ids are identity, not sequence.
**Cross-repo:** T04 and T05 changed `hall-of-helix`, not this repo. They were
executed under `HOH-WP-0001` there on 2026-09-05 and are `done`; the routine
lives in `hall-of-helix/CLOSING.md` and `make check` enforces the record on
finished agent seats from 2026-09-06. What remains here is first-attempt pilot
evidence (T01) and the resulting version decision (T03).
finished agent seats from 2026-09-06. The read-only export script added for T01
also lives in the hall. Entry content was not edited by this review.
**2026-09-28 review:** T02 is done. T01 and T03 are `wait`, so this workplan is
**2026-09-28 initial review (superseded):** T02 is done. T01 and T03 are `wait`, so this workplan is
`blocked`. Published hall records demonstrate adoption but do not establish
eight unmodified, uncoached, unrepaired first attempts. Resume T01 when original
closing outputs with that provenance are available across at least three repos
@ -79,11 +90,24 @@ T03 waits for those findings. No new tasks or workplans were opened.
```task
id: PQRST-WP-0002-T01
status: wait
status: done
priority: high
state_hub_task_id: "dd259b3f-e9ae-5bdd-8e02-07ae0083930a"
```
**Completed under revised scope:** inspect the committed hall corpus, retain
attempt history, structurally check records, and qualitatively review at least
eight examples spanning three repositories and two task classes. The review
contains eight verbatim examples (including a rejected first attempt), reviewer
task classes, provenance limits, and answers to all six questions below. The
exporter emits 196 entries, 95 with PQRST, containing 96 attempt blocks; 94 of
95 frontmatter-selected blocks pass structural checks. Published validity is
not unassisted first-attempt validity. This scope follows the operator's explicit
instruction to use the hall as our main data source, not a claim that the
original controlled protocol was fulfilled.
**Original protocol, retained for provenance:**
Paste `PqrstPrompt.md` unmodified at the end of at least **eight** real agentic
coding sessions, spread across at least three repositories and at least two task
classes (e.g. `feature`, `bugfix`, `explore`, `harden`). Once T04 has landed,
@ -164,14 +188,18 @@ work here.
```task
id: PQRST-WP-0002-T03
status: wait
status: done
priority: medium
state_hub_task_id: "0a0bce15-ac0c-582a-842c-ae1bf0307afb"
```
**2026-09-28:** blocked on T01's first-attempt evidence. The storage cross-reference
is already landed through T02. Leave v0.1 unchanged pending the pilot; this is
not a finding that no substantive revision is needed.
**Completed 2026-09-28:** retain v0.1 after the published-corpus review. Fix the
prompt checklist's demonstrated retry contradiction, clarify substantive-work
scope and optional uncertainty/provenance Notes, and align the spec's ranking
prohibition with INTENT. No dimension, format or validation rule changes; the
fenced canonical prompt is unchanged. Appendix B records the editorial update.
T02 already supplied the storage cross-reference. This disposition is supported
by published-record evidence, not a controlled accuracy or first-attempt study.
From the pilot findings and the hall integration, revise the specification and
prompt together:
@ -306,8 +334,10 @@ exception in `CLOSING.md`.
## Pilot findings
_Populated by T01. One subsection per session: repository, task class, the
verbatim record, and whether it validated unassisted._
The completed [corpus review](../docs/reviews/2026-09-28-hall-pqrst-review.md)
contains the method, findings, and eight verbatim source examples with validation
and provenance dispositions. The initial evidence scan below is historical;
the later corpus review supersedes its blocked-task conclusion.
### Evidence availability review — 2026-09-28