From 65fc8c77e3ad095abd6d2932bcb6a1fff9e340db Mon Sep 17 00:00:00 2001 From: tegwick Date: Mon, 28 Sep 2026 11:54:34 +0200 Subject: [PATCH] Review hall PQRST corpus and close published-record pilot Assistant: codex Assistant-Model: gpt-6-astra Assistant-Session: 01a0e759-301a-78b1-bbc1-040ef094b12d --- PqrstPrompt.md | 12 +- README.md | 3 + SCOPE.md | 6 + docs/adr/0002-record-storage-boundary.md | 7 + docs/reviews/2026-09-28-hall-pqrst-review.md | 323 ++++++++++++++++++ spec/PqrstEstimationPractice.md | 18 +- ...0002-validate-v01-against-real-sessions.md | 62 +++- 7 files changed, 413 insertions(+), 18 deletions(-) create mode 100644 docs/reviews/2026-09-28-hall-pqrst-review.md diff --git a/PqrstPrompt.md b/PqrstPrompt.md index 7706426..a9a10c5 100644 --- a/PqrstPrompt.md +++ b/PqrstPrompt.md @@ -6,6 +6,10 @@ Specification: [`spec/PqrstEstimationPractice.md`](spec/PqrstEstimationPractice. Do not add scoring dimensions inside PQRST. Do not run this prompt mid-session unless deliberately closing a phase. +Estimate the substantive work before the closing ritual. Writing the hall entry, +rendering its portrait, and the final hall sync are excluded. Coordination and +verification performed as part of the substantive task still count. + --- ```text @@ -73,7 +77,7 @@ After the block, also emit the equivalent YAML record with keys p, q, r, s, t, t ## Accepting or rejecting a result -Reject and re-run if any of the following is true: +Reject the result as a valid record if any of the following is true: - the five values are not integers summing to 100; - `Confidence` is missing or not one of `low` / `medium` / `high`; @@ -81,3 +85,9 @@ Reject and re-run if any of the following is true: - `Dominant factors` merely restates the numbers or says something like "mixed work across several areas"; - a sixth dimension has been introduced inside the 5-tuple; - S is non-zero but no security-specific work actually occurred. + +Preserve the original output and the rejection reason. Do not silently repair +it or re-run for a better answer; a rejected first attempt is useful evidence. +If an existing record already has a retry, retain both and label their order. +Use optional `Notes` to preserve attribution uncertainty, limited session +context, or any prior allocation hint; do not call those records uncoached. diff --git a/README.md b/README.md index f3ac5a1..259f81e 100644 --- a/README.md +++ b/README.md @@ -51,3 +51,6 @@ validators, dashboards, and record storage belong elsewhere — see When closing with a voluntary hall-of-helix entry, retain the record there. The hall is the default human-facing store, not a mandatory or exclusive one; see [the storage decision](docs/adr/0002-record-storage-boundary.md). + +The [2026-09-28 hall review](docs/reviews/2026-09-28-hall-pqrst-review.md) +documents the extraction command, corpus findings, and why v0.1 stands. diff --git a/SCOPE.md b/SCOPE.md index 02a47e7..d87f9ec 100644 --- a/SCOPE.md +++ b/SCOPE.md @@ -47,6 +47,12 @@ or commissioned. Seats are a biased sample, not a session census; no seat is required just to retain an estimate. See [ADR-002 — Record storage boundary](docs/adr/0002-record-storage-boundary.md). +`hall-of-helix/entries` is the main data source for this practice's record-format +and prompt reviews. Its read-only `scripts/export-pqrst.py` exports a pinned Git +snapshot for analysis elsewhere; the hall retains the originals. Bounded review +excerpts and findings may accompany a workplan here without becoming a session +store or an empirical accuracy study. + ### Integration with any particular agent or harness PQRST is deliberately agent-agnostic and model-agnostic. Wiring it into a specific harness, plugin, skill, or CI pipeline belongs to that harness's own repository. diff --git a/docs/adr/0002-record-storage-boundary.md b/docs/adr/0002-record-storage-boundary.md index 25163cd..025ae5a 100644 --- a/docs/adr/0002-record-storage-boundary.md +++ b/docs/adr/0002-record-storage-boundary.md @@ -39,5 +39,12 @@ validation rule, or record format changes, so v0.1 remains in force. The eight-session first-attempt pilot and the subsequent version decision remain under the existing T01 and T03; this decision does not establish their outcome. +**2026-09-28 follow-up:** the operator selected hall entries as the main data +source. T01 now reviews that published corpus, with first-attempt uncertainty +explicit rather than a completion gate. The hall owns a read-only Git-snapshot +exporter; no machine-facing collection service or second record store is +commissioned. The [review](../reviews/2026-09-28-hall-pqrst-review.md) closes T01 +and supplies T03's decision to retain v0.1 with editorial alignment. + See [SCOPE.md](../../SCOPE.md) and the [existing workplan](../../workplans/PQRST-WP-0002-validate-v01-against-real-sessions.md). diff --git a/docs/reviews/2026-09-28-hall-pqrst-review.md b/docs/reviews/2026-09-28-hall-pqrst-review.md new file mode 100644 index 0000000..68a053e --- /dev/null +++ b/docs/reviews/2026-09-28-hall-pqrst-review.md @@ -0,0 +1,323 @@ +# Hall PQRST review — 2026-09-28 + +State Hub review decision: `2f67b409-c5df-4f8e-bce8-98e9c7be92ab`. + +## Source and method + +The operator selected `hall-of-helix/entries` as the main data source for +PQRST-WP-0002 on 2026-09-28. This replaces the original eight-session controlled +first-attempt pilot with a retrospective review of published records. It does +not relabel repaired or coached estimates as unassisted. The original pilot's +success-rate and bias questions remain unanswerable from this corpus and are +not release gates for the specification-and-prompt repository. + +Reviewed hall commit `82caae23b3098aa4bdf4346cd4af8a88c03c7d68`, covering entries +recorded through 2026-09-28. Six local entry edits were excluded by reading Git +objects rather than working files. The hall remains the record store; examples +below are a bounded review sample, not a second session corpus. + +The standard-library exporter lives in `hall-of-helix/scripts/export-pqrst.py`. +Run from a workstation with that checkout: + +```bash +python3 ~/hall-of-helix/scripts/export-pqrst.py ~/hall-of-helix \ + --revision 82caae23b3098aa4bdf4346cd4af8a88c03c7d68 > /tmp/hall-pqrst.jsonl +``` + +Omit `--revision` to review committed HEAD next time. Each JSONL row preserves +source commit, entry path, content SHA-256, raw frontmatter, all PQRST sections, +and every fenced record in order. Each block includes raw text and a field map; +field values are lists so duplicate fields remain visible. Missing sections +produce empty lists, not silently dropped seats. The exporter does not certify +validity, select a successful retry, rank workers, or modify entries. + +For this review, parse the raw frontmatter as YAML, select the block whose +Signature matches `pqrst_estimate`, and retain every other block as attempt +history. Check five unique integer percentage fields in 0..100, sum 100, a +matching numeric signature (leading zeros allowed), one allowed confidence, +and present dominant factors. Check Sum and frontmatter/body agreement too. +Concrete evidence and purpose attribution require human reading; a nonempty +Dominant factors field alone cannot establish spec §5.5 validity. + +## Corpus findings + +| Observation | Count and denominator | +| --- | --- | +| Committed entries | 196: 195 agent-session, 1 human | +| Entries carrying PQRST | 95: 58 draft, 32 handed-forward, 5 complete | +| Fenced PQRST blocks | 96 across those 95 entries | +| Unique block matching frontmatter | 95/95 entries | +| Matching blocks passing structural checks | 94/95 | +| Confidence on matching blocks | 91 medium, 3 high, 0 low, 1 absent | +| All five values in multiples of five | 87/94 structurally passing matching blocks | +| S explicitly zero | 29/94 structurally passing matching blocks | +| Repository labels represented | 73, including the hall itself; multi-repo sessions overlap | +| Explicit task-class/outcome frontmatter fields | 0/95 | + +All 101 entries without PQRST predate the 2026-09-06 adoption boundary by +`recorded_at`; they are not failures to comply with the adopted routine. +The 95 records span 2026-09-05 through 2026-09-28. Drafts are included because +these are already published estimates; portrait status is not estimate quality. +No model/worker comparison, performance score, or average effort target is +inferred from these counts. + +The lone structurally incomplete selected record is the grandfathered +`2026-09-05T16:36:12.000Z-codex-statehub-snapshot-and-signature.md`. It lacks +Confidence and Dominant factors. Its hall note calls it valid, but that means +less than spec §5.5: it is incomplete under that contract. Preserve it as +historical evidence; do not invent missing rationale. + +`2026-09-22T07-59-35.000Z-claude-62534cdf-warnings-were-the-work.md` contains +an invalid first block and a valid retry. The original has Q=15 in the values +but Q=30 in the signature, and no Dominant factors. Flattening the section into +one dictionary would mix the two attempts and corrupt the result. The author's +explanation directly identifies the prompt's “reject and re-run” checklist as +the reason for retrying, contradicting the spec's once-only collection rule and +the hall's preserve-failures routine. This is a demonstrated instruction defect. + +## Answers to the pilot questions + +1. **Uncoached first-attempt validity:** not measurable from published seats. + There is one explicitly failed first attempt, and another record explicitly + discloses an allocation hint in a continuation summary (sample E). Therefore + 94/95 is published structural completeness, not a first-attempt success rate. +2. **Primary-purpose attribution:** usable as a judgment, not uniquely decidable. + Samples A and H explain P/Q and P/S ambiguity. Sample H says a different + defensible classification would reverse much of P and S. Keep R5 and preserve + that explanation in Notes; do not interpret a small percentage difference + as a measured change in effort. +3. **Zero security:** 29 structurally passing records use S=0. Samples B and G + ground this in the work described; G distinguishes reading about a credential + gate from doing security work. This shows the zero is usable, not that every + security allocation is accurate. Security-specific analysis can count as S + without a code or credential change; classify by purpose, not write activity. +4. **Coarse increments:** 87/94 are all multiples of five. The seven exceptions + are not validation failures: R7 is a preference. Samples A, C and H show + finer splits without evidence that 1% resolution is reliable. Keep the coarse + preference; do not round stored originals or make coarse increments mandatory. +5. **Confidence:** medium dominates (91/94 present values); high appears three + times and low never appears. The field is not constant, but these observations + cannot establish calibration. Existing Notes can explain uncertainty; no new + confidence dimension, required field, or numeric confidence score is needed. +6. **Self-estimation bias:** cannot be measured without independent observations + or estimates. Self-reports and their narratives are not independent evidence + against P inflation. Acknowledge that limit; do not launch an accuracy study + or rank agents in this repository. + +Task classes below are reviewer-assigned from Contribution, not recovered +metadata. The eight deliberately selected examples cover feature, hardening, +and exploration work across more than three repositories, and include failures +and provenance limitations. They are diagnostic examples, not a random sample. +Each passing sample's Dominant factors names concrete artifacts or actions +consistent with its Contribution. This is a qualitative check on these eight, +not semantic validation of all 95 records or independent session verification. + +## Version disposition + +Keep **v0.1**. Neither five-dimensional meaning nor record format nor §5.5 +validation needs a breaking change on this evidence. Correct the contradictory +retry instruction; clarify the already implemented substantive-work boundary +and preservation of uncertainty/provenance in optional Notes. Preserve the +existing hall exclusion of seat writing, portrait generation and final hall +sync; substantive coordination and verification before closure still count. +Make the anti-ranking wording consistent with INTENT's unconditional ban. + +These are editorial alignment changes, recorded in Appendix B. They do not +make historic records newly comparable or certify their original collection +conditions. Existing copies of the fenced canonical prompt remain unchanged. +The hall exporter is the simple read-only extraction path; no collector, +validator service, dashboard, or second store is introduced here. + +## Eight source examples + +Blocks below are copied verbatim from the pinned source, including the invalid +attempt. Do not apply the illustrative-signature “sum to 100” smoke check to +that deliberately retained counterexample. Sources are paths under the hall's +`entries/` at the commit above. + +### A — checks-were-the-thing-that-lied + +Source: `entries/2026-09-08T11-20-00.000Z-claude-01Bjefh8-the-checks-were-the-thing-that-lied.md` +Repositories: canned-prompts, rapp-canned-prompts, helix-forge, rapp-postgres +Reviewer task class: feature / harden + +Published block passes §5.5 on review; first-attempt status unknown. Canonical package provenance is stated. Notes explicitly hedge P/Q attribution. + +```text +PQRST-Estimate +P: 35% +Q: 30% +R: 18% +S: 12% +T: 5% +Sum: 100% +Confidence: medium +Signature: P35 Q30 R18 S12 T5 +Dominant factors: Three things were built end to end — the CPF v0.2 format decisions with their reference implementation, the hosted registry service, and its Railiance deployment — and the debugging that followed was nearly as large: a migration that logged every revision as applied while silently rolling back, an egress policy that never selected the migration Job, and five defects in my own verification instruments, three of them in the lease watcher before it could return an honest verdict. Security was substantive rather than incidental: publisher token design (hashed, constant-time, non-enumerable, revocation preserving attribution), replacing env-var credentials with mounted files, and surviving credential-lease rotation. +Notes: Attribution across a session this long is approximate; the P/Q split in particular could defensibly shift several points either way, since debugging a silent rollback is both implementing the deliverable and establishing its correctness. +``` + +### B — fiam-four-source-plates + +Source: `entries/2026-09-09T21-15-50Z-codex-fiam-four-source-plates.md` +Repositories: facetted-interfaces, hall-of-helix +Reviewer task class: explore / harden + +Published block passes §5.5 on review; first-attempt status unknown. Coarse values, zero S, and substantive-work boundary are explicit; the model handoff limits session provenance. + +```text +PQRST-Estimate +P: 30% +Q: 25% +R: 25% +S: 0% +T: 20% +Sum: 100% +Confidence: medium +Signature: P30 Q25 R25 S0 T20 +Dominant factors: Repository intent and the InterfaceCanon acceptance receipt were the main deliverables; Draft 6 conformance, interpreter debugging, source-hash checks, and reading the governance and FIAM material accounted for much of the remaining work. Workplan creation, registration, and projection reconciliation required substantial coordination. +Notes: Covers the substantive conversation across the model handoff; excludes this closing ritual. +``` + +### C — guard-proved-less-than-claimed + +Source: `entries/2026-09-10T22-04-31.000Z-claude-01NV9oij-guard-proved-less-than-claimed.md` +Repositories: key-cape +Reviewer task class: harden + +Published block passes §5.5 on review; first-attempt status unknown. Security and verification have concrete anchors; non-five-point values are not prohibited. No explicit uncertainty explanation. + +```text +PQRST-Estimate +P: 22% +Q: 22% +R: 13% +S: 30% +T: 13% +Sum: 100% +Confidence: medium +Signature: P22 Q22 R13 S30 T13 +Dominant factors: The bulk of the session was identity-security work in an issuer — authorization-code redirect/grant/client binding with atomic code consumption, upstream Authelia ID-token verification via a new strict internal/jose, tenant provenance (tenant_source) under GH-DEC-2026-013, bind passwords off argv with 0600 artifacts, and two escalation guards (dynamic registration, principal_type) — with a matching test burden where each fix was verified by mutation rather than assertion. +Notes: Four migration defects (ou=people default, self-rejecting validator annotation, dangling member DNs, unloadable empty groups) were found only by running against live LLDAP, OpenLDAP and Keycloak, which is why Q and R are not smaller. +``` + +### D — warnings-were-the-work + +Source: `entries/2026-09-22T07-59-35.000Z-claude-62534cdf-warnings-were-the-work.md` +Repositories: railiance-fabric, hall-of-helix +Reviewer task class: feature / maintenance + +First attempt explicitly fails §5.5: mismatched signature and missing dominant factors. The matching published retry passes on review. Both attempts are preserved below; never count them as two sessions. + +```text +PQRST-Estimate +P: 30% +Q: 15% +R: 25% +S: 0% +T: 30% +Sum: 100% +Confidence: medium +Signature: P30 Q30 R25 S0 T30 +``` + +```text +PQRST-Estimate +P: 30% +Q: 15% +R: 25% +S: 0% +T: 30% +Sum: 100% +Confidence: medium +Signature: P30 Q15 R25 S0 T30 +Dominant factors: T came from four fix-consistency passes, splitting the work into commits, the flex-auth handoff (the RAIL-FAB-WP-0031 intake plus the reply), and re-sequencing the WP-0028 tasks; P came from chokepoint sizing in coordination_graph.py and the C-35/C-23 fixes; R came from reading the inherited diffs, the 3,000-line explorer UI, repo-manager's classification rules, and the orientation notice. +Notes: S is 0 because I read the credential and production traps in the orientation notice but did no security work. +``` + +### E — memo-version-three + +Source: `entries/2026-09-27T13-49-49Z-grok-memo-version-three.md` +Repositories: secrets-engine, flex-auth, railiance-platform, railiance-clock, informed-decision +Reviewer task class: harden / operations + +Published block passes §5.5 on review. Uncoached provenance explicitly fails: Notes disclose an allocation hint from a continuation summary. Retry status unknown. + +```text +PQRST-Estimate +P: 25% +Q: 15% +R: 30% +S: 20% +T: 10% +Sum: 100% +Confidence: medium +Signature: P25 Q15 R30 S20 T10 +Dominant factors: Most of the stretch was spent separating Railiance Clock trust from the guest NTP skew, and separating the two informed-decision review images from the lifecycle PDP flex-auth-secrets-engine, including the Forgejo digest and the worker reports. The security slice was three new approval ids so the consumed 2026-09-16 approvals stayed consumed, a fail-closed skip when the live AppRole already matched, and the withdrawn NTP hold; the delivered change was memo version 3 and helm revision 4 of the two review releases. +Notes: A continuation summary suggested a research-heavy, security-substantial split before this estimate was written. These integers were chosen after that hint and give more weight to the publish-and-roll than the hint did. The closing ritual is excluded. +``` + +### F — statement-that-never-became-an-invoice + +Source: `entries/2026-09-27T19-55-40Z-claude-286d235a-the-statement-that-never-became-an-invoice.md` +Repositories: fin-hub +Reviewer task class: feature + +Published block passes §5.5 on review; first-attempt status unknown. A high-confidence example with concrete schema, store, CLI and test work; confidence calibration is not established. + +```text +PQRST-Estimate +P: 35% +Q: 25% +R: 25% +S: 0% +T: 15% +Sum: 100% +Confidence: high +Signature: P35 Q25 R25 S0 T15 +Dominant factors: Most effort went into reading resource-control's settlement-statement schema and terms document, then fin-hub's existing exchange.py/schemas conventions, before writing the SettlementStatement model, its store, and the CLI wiring (R and P in close proportion); a comparable share went to writing and running four new idempotency/health tests plus two full-suite passes (Q); the remainder was triaging which of three open workplans had real in-repo work left versus a genuine external blocker, plus recovering cleanly from a stray uv.lock/git-stash detour mid-session (T). +``` + +### G — gate-was-elsewhere + +Source: `entries/2026-09-21T13-49-08.000Z-claude-816898b9-the-gate-was-elsewhere.md` +Repositories: sand-boxer +Reviewer task class: explore / maintenance + +Published block passes §5.5 on review; first-attempt status unknown. Explicit zero S despite security-related task dependencies; context research is distinguished from security work. + +```text +PQRST-Estimate +P: 30% +Q: 10% +R: 45% +S: 0% +T: 15% +Sum: 100% +Confidence: medium +Signature: P30 Q10 R45 S0 T15 +Dominant factors: Research dominated — reading the two open workplans' evidence trail, cross-checking task status in the hub, and tracing C-24 into state-hub's validator and the canonical allowed-values file; the deliverable itself was two small edits (WP-0014 status, classification tags). +Notes: S is 0 — the open tasks concern credentials, but this session only read about those gates; no security work occurred. +``` + +### H — receipts-outlived-the-recollection + +Source: `entries/2026-09-10T22-52-29.000Z-claude-01WLUjpv-receipts-outlived-the-recollection.md` +Repositories: railiance-platform +Reviewer task class: harden / operations + +Published block passes §5.5 on review; first-attempt status unknown. Notes explicitly hedge P/S, offering a materially different defensible allocation; this is direct evidence of classification uncertainty. + +```text +PQRST-Estimate +P: 17% +Q: 10% +R: 18% +S: 35% +T: 20% +Sum: 100% +Confidence: medium +Signature: P17 Q10 R18 S35 T20 +Dominant factors: Nearly every deliverable was credential custody — two verifier CCRs with exact-path policies and namespace-limited ESO delivery, two client-side reader admissions, an attended read-only OpenBao runner with an ambient-token guard, and the field/duplicate finding at the whynot-design lane. Coordination was the next largest slice: eighteen owner replies across thirteen threads, several of which required reading sibling repos (secrets-engine, approval-engine, key-cape, ops-warden, net-kingdom) and the live cluster before an honest answer was possible. +Notes: The P/S boundary is the main uncertainty. Writing a CCR or a policy is both the requested deliverable and credential work; I classified by primary purpose per rule 4, which pushes S well above P. A reader who split it the other way would get roughly P35 S17 with the same session underneath. +``` diff --git a/spec/PqrstEstimationPractice.md b/spec/PqrstEstimationPractice.md index fe86194..8dc53d8 100644 --- a/spec/PqrstEstimationPractice.md +++ b/spec/PqrstEstimationPractice.md @@ -138,6 +138,12 @@ Some T is productive coordination. Excessive or repeated T may signal unstable s **R5 — Attribute overlapping work by primary purpose.** Categories overlap by design. Assign effort by the *primary reason the activity was undertaken at that moment*, and do not double-count. +This is a judgment, not a uniquely recoverable split. When another attribution +would materially change the estimate, use optional `Notes` to name the boundary +(for example P/S) and lower confidence as appropriate. Security-specific analysis +can count as S without changing a control or handling a credential; merely reading +that a task waits on credentials may instead be R or T, depending on purpose. + | Activity | Category | | --- | --- | | Reading a module to understand how it works | **R** | @@ -163,6 +169,15 @@ Some T is productive coordination. Excessive or repeated T may signal unstable s Run the estimate **once**, at the natural end of a session or a clearly bounded work unit — one ticket, one vertical slice, one deliberate "stop here" checkpoint. +The substantive work is the unit being estimated. The subsequent hall closing +ritual — writing the seat, rendering its portrait, and final hall sync — is +excluded. Coordination and verification that belong to the substantive task +remain included. Preserve a rejected output with its reason rather than silently +repairing it or rerunning for a better result. Existing retries remain separate, +ordered attempts, not additional sessions. Optional `Notes` can disclose prior +allocation hints or limited session context; published validity does not prove +an uncoached first attempt. + Do not collect mid-stream unless closing a phase on purpose. Mid-session estimates mix unfinished work with planning residue and are hard to compare. If a session spanned several distinct modes (explore, then implement, then harden), emit **one overall estimate** plus optional phase notes. The stored record remains a single 5-tuple unless phases are explicitly versioned. @@ -343,7 +358,7 @@ Questions the joined data can answer: 7. **Compare like with like.** An explore session and a one-line fix do not share a healthy template. 8. **Store the signature line plus the dominant-factors sentence.** Numbers without the sentence are not auditable. 9. **Preserve the rationales, not only the percentages.** -10. **Do not rank developers, agents, or teams on raw percentages** without additional outcome evidence. +10. **Do not rank developers, agents, models, or teams on PQRST percentages.** Outcome evidence supports process review, not a worker leaderboard. 11. **Do not discuss a stored estimate during the next session** unless changing process on purpose. --- @@ -408,3 +423,4 @@ At session end: | Version | Date | Change | | --- | --- | --- | | 0.1 | 2026-09-05 | Initial consolidation of the two independent 2026-09-05 drafts into one specification. The block/signature record was adopted as the source of truth; the per-dimension table was retained as an optional presentation form. | +| 0.1 (editorial) | 2026-09-28 | [Hall review](../docs/reviews/2026-09-28-hall-pqrst-review.md): retain the dimensions, format and validation rules; align the prompt checklist with once-only collection, clarify the existing substantive-work boundary and optional uncertainty/provenance notes, and align the anti-ranking rule with INTENT. Published records do not establish first-attempt success or calibrated accuracy. | diff --git a/workplans/PQRST-WP-0002-validate-v01-against-real-sessions.md b/workplans/PQRST-WP-0002-validate-v01-against-real-sessions.md index 524b563..b0f3c68 100644 --- a/workplans/PQRST-WP-0002-validate-v01-against-real-sessions.md +++ b/workplans/PQRST-WP-0002-validate-v01-against-real-sessions.md @@ -4,7 +4,7 @@ type: workplan title: "Validate spec v0.1 in real session closes and land PQRST in hall-of-helix" domain: agents repo: pqrst-practice -status: blocked +status: finished flavor: planning owner: claude-code topic_slug: practice @@ -24,8 +24,20 @@ repos: state_hub_workstream_id: "6a875b19-5a76-55c1-bd1f-2f5005cd416b" --- +State Hub review decision: `2f67b409-c5df-4f8e-bce8-98e9c7be92ab`. + # Validate spec v0.1 in real session closes and land PQRST in hall-of-helix +**Completed 2026-09-28:** the operator selected hall entries as the main data +source. [The pinned corpus review](../docs/reviews/2026-09-28-hall-pqrst-review.md) +closes T01 under the revised retrospective-review scope; T03 retains v0.1 with +editorial corrections grounded in those findings. All five tasks are done. +No new tasks or workplans were opened, and no implementation residuals remain. +Uncoached success rates, calibrated accuracy and self-estimation bias are +explicitly unestablished; studying them is outside this repository's scope. + +The original proposal and initial blocked review are retained below as history. + Spec v0.1 was consolidated from two independent drafts, neither of which had been applied to an actual session. It is internally consistent and entirely unvalidated. @@ -52,23 +64,22 @@ It deliberately does **not** build tooling or a record store in *this* repository. Both are out of scope (`SCOPE.md`); the hall is the store, and the hall owns its own validation. -**Done when:** the closing routine is specified in hall-of-helix and reachable +**Original done condition (superseded by the operator's corpus-review direction):** the closing routine is specified in hall-of-helix and reachable from the prompt above, entries carry a validated PQRST record, the prompt has been run unassisted at the end of at least eight real sessions, and v0.2 either incorporates the findings or records why v0.1 stands. -**Execution order:** T04/T05 are complete; T02 closed independently on -2026-09-28 using the existing hall integration. T01 must supply the pilot -evidence before T03 can settle the version decision. Task ids are identity, -not sequence. +**Completion order:** T04/T05 established the hall integration; T02 closed +the storage boundary; T01 reviewed the published corpus; T03 applied the +editorial findings and retained v0.1. Task ids are identity, not sequence. **Cross-repo:** T04 and T05 changed `hall-of-helix`, not this repo. They were executed under `HOH-WP-0001` there on 2026-09-05 and are `done`; the routine lives in `hall-of-helix/CLOSING.md` and `make check` enforces the record on -finished agent seats from 2026-09-06. What remains here is first-attempt pilot -evidence (T01) and the resulting version decision (T03). +finished agent seats from 2026-09-06. The read-only export script added for T01 +also lives in the hall. Entry content was not edited by this review. -**2026-09-28 review:** T02 is done. T01 and T03 are `wait`, so this workplan is +**2026-09-28 initial review (superseded):** T02 is done. T01 and T03 are `wait`, so this workplan is `blocked`. Published hall records demonstrate adoption but do not establish eight unmodified, uncoached, unrepaired first attempts. Resume T01 when original closing outputs with that provenance are available across at least three repos @@ -79,11 +90,24 @@ T03 waits for those findings. No new tasks or workplans were opened. ```task id: PQRST-WP-0002-T01 -status: wait +status: done priority: high state_hub_task_id: "dd259b3f-e9ae-5bdd-8e02-07ae0083930a" ``` +**Completed under revised scope:** inspect the committed hall corpus, retain +attempt history, structurally check records, and qualitatively review at least +eight examples spanning three repositories and two task classes. The review +contains eight verbatim examples (including a rejected first attempt), reviewer +task classes, provenance limits, and answers to all six questions below. The +exporter emits 196 entries, 95 with PQRST, containing 96 attempt blocks; 94 of +95 frontmatter-selected blocks pass structural checks. Published validity is +not unassisted first-attempt validity. This scope follows the operator's explicit +instruction to use the hall as our main data source, not a claim that the +original controlled protocol was fulfilled. + +**Original protocol, retained for provenance:** + Paste `PqrstPrompt.md` unmodified at the end of at least **eight** real agentic coding sessions, spread across at least three repositories and at least two task classes (e.g. `feature`, `bugfix`, `explore`, `harden`). Once T04 has landed, @@ -164,14 +188,18 @@ work here. ```task id: PQRST-WP-0002-T03 -status: wait +status: done priority: medium state_hub_task_id: "0a0bce15-ac0c-582a-842c-ae1bf0307afb" ``` -**2026-09-28:** blocked on T01's first-attempt evidence. The storage cross-reference -is already landed through T02. Leave v0.1 unchanged pending the pilot; this is -not a finding that no substantive revision is needed. +**Completed 2026-09-28:** retain v0.1 after the published-corpus review. Fix the +prompt checklist's demonstrated retry contradiction, clarify substantive-work +scope and optional uncertainty/provenance Notes, and align the spec's ranking +prohibition with INTENT. No dimension, format or validation rule changes; the +fenced canonical prompt is unchanged. Appendix B records the editorial update. +T02 already supplied the storage cross-reference. This disposition is supported +by published-record evidence, not a controlled accuracy or first-attempt study. From the pilot findings and the hall integration, revise the specification and prompt together: @@ -306,8 +334,10 @@ exception in `CLOSING.md`. ## Pilot findings -_Populated by T01. One subsection per session: repository, task class, the -verbatim record, and whether it validated unassisted._ +The completed [corpus review](../docs/reviews/2026-09-28-hall-pqrst-review.md) +contains the method, findings, and eight verbatim source examples with validation +and provenance dispositions. The initial evidence scan below is historical; +the later corpus review supersedes its blocked-task conclusion. ### Evidence availability review — 2026-09-28