Review hall PQRST corpus and close published-record pilot

Assistant: codex
Assistant-Model: gpt-6-astra
Assistant-Session: 01a0e759-301a-78b1-bbc1-040ef094b12d
This commit is contained in:
tegwick 2026-09-28 11:54:34 +02:00
parent 37b49a9517
commit 65fc8c77e3
7 changed files with 413 additions and 18 deletions

View file

@ -39,5 +39,12 @@ validation rule, or record format changes, so v0.1 remains in force. The
eight-session first-attempt pilot and the subsequent version decision remain
under the existing T01 and T03; this decision does not establish their outcome.
**2026-09-28 follow-up:** the operator selected hall entries as the main data
source. T01 now reviews that published corpus, with first-attempt uncertainty
explicit rather than a completion gate. The hall owns a read-only Git-snapshot
exporter; no machine-facing collection service or second record store is
commissioned. The [review](../reviews/2026-09-28-hall-pqrst-review.md) closes T01
and supplies T03's decision to retain v0.1 with editorial alignment.
See [SCOPE.md](../../SCOPE.md) and the
[existing workplan](../../workplans/PQRST-WP-0002-validate-v01-against-real-sessions.md).

View file

@ -0,0 +1,323 @@
# Hall PQRST review — 2026-09-28
State Hub review decision: `2f67b409-c5df-4f8e-bce8-98e9c7be92ab`.
## Source and method
The operator selected `hall-of-helix/entries` as the main data source for
PQRST-WP-0002 on 2026-09-28. This replaces the original eight-session controlled
first-attempt pilot with a retrospective review of published records. It does
not relabel repaired or coached estimates as unassisted. The original pilot's
success-rate and bias questions remain unanswerable from this corpus and are
not release gates for the specification-and-prompt repository.
Reviewed hall commit `82caae23b3098aa4bdf4346cd4af8a88c03c7d68`, covering entries
recorded through 2026-09-28. Six local entry edits were excluded by reading Git
objects rather than working files. The hall remains the record store; examples
below are a bounded review sample, not a second session corpus.
The standard-library exporter lives in `hall-of-helix/scripts/export-pqrst.py`.
Run from a workstation with that checkout:
```bash
python3 ~/hall-of-helix/scripts/export-pqrst.py ~/hall-of-helix \
--revision 82caae23b3098aa4bdf4346cd4af8a88c03c7d68 > /tmp/hall-pqrst.jsonl
```
Omit `--revision` to review committed HEAD next time. Each JSONL row preserves
source commit, entry path, content SHA-256, raw frontmatter, all PQRST sections,
and every fenced record in order. Each block includes raw text and a field map;
field values are lists so duplicate fields remain visible. Missing sections
produce empty lists, not silently dropped seats. The exporter does not certify
validity, select a successful retry, rank workers, or modify entries.
For this review, parse the raw frontmatter as YAML, select the block whose
Signature matches `pqrst_estimate`, and retain every other block as attempt
history. Check five unique integer percentage fields in 0..100, sum 100, a
matching numeric signature (leading zeros allowed), one allowed confidence,
and present dominant factors. Check Sum and frontmatter/body agreement too.
Concrete evidence and purpose attribution require human reading; a nonempty
Dominant factors field alone cannot establish spec §5.5 validity.
## Corpus findings
| Observation | Count and denominator |
| --- | --- |
| Committed entries | 196: 195 agent-session, 1 human |
| Entries carrying PQRST | 95: 58 draft, 32 handed-forward, 5 complete |
| Fenced PQRST blocks | 96 across those 95 entries |
| Unique block matching frontmatter | 95/95 entries |
| Matching blocks passing structural checks | 94/95 |
| Confidence on matching blocks | 91 medium, 3 high, 0 low, 1 absent |
| All five values in multiples of five | 87/94 structurally passing matching blocks |
| S explicitly zero | 29/94 structurally passing matching blocks |
| Repository labels represented | 73, including the hall itself; multi-repo sessions overlap |
| Explicit task-class/outcome frontmatter fields | 0/95 |
All 101 entries without PQRST predate the 2026-09-06 adoption boundary by
`recorded_at`; they are not failures to comply with the adopted routine.
The 95 records span 2026-09-05 through 2026-09-28. Drafts are included because
these are already published estimates; portrait status is not estimate quality.
No model/worker comparison, performance score, or average effort target is
inferred from these counts.
The lone structurally incomplete selected record is the grandfathered
`2026-09-05T16:36:12.000Z-codex-statehub-snapshot-and-signature.md`. It lacks
Confidence and Dominant factors. Its hall note calls it valid, but that means
less than spec §5.5: it is incomplete under that contract. Preserve it as
historical evidence; do not invent missing rationale.
`2026-09-22T07-59-35.000Z-claude-62534cdf-warnings-were-the-work.md` contains
an invalid first block and a valid retry. The original has Q=15 in the values
but Q=30 in the signature, and no Dominant factors. Flattening the section into
one dictionary would mix the two attempts and corrupt the result. The author's
explanation directly identifies the prompt's “reject and re-run” checklist as
the reason for retrying, contradicting the spec's once-only collection rule and
the hall's preserve-failures routine. This is a demonstrated instruction defect.
## Answers to the pilot questions
1. **Uncoached first-attempt validity:** not measurable from published seats.
There is one explicitly failed first attempt, and another record explicitly
discloses an allocation hint in a continuation summary (sample E). Therefore
94/95 is published structural completeness, not a first-attempt success rate.
2. **Primary-purpose attribution:** usable as a judgment, not uniquely decidable.
Samples A and H explain P/Q and P/S ambiguity. Sample H says a different
defensible classification would reverse much of P and S. Keep R5 and preserve
that explanation in Notes; do not interpret a small percentage difference
as a measured change in effort.
3. **Zero security:** 29 structurally passing records use S=0. Samples B and G
ground this in the work described; G distinguishes reading about a credential
gate from doing security work. This shows the zero is usable, not that every
security allocation is accurate. Security-specific analysis can count as S
without a code or credential change; classify by purpose, not write activity.
4. **Coarse increments:** 87/94 are all multiples of five. The seven exceptions
are not validation failures: R7 is a preference. Samples A, C and H show
finer splits without evidence that 1% resolution is reliable. Keep the coarse
preference; do not round stored originals or make coarse increments mandatory.
5. **Confidence:** medium dominates (91/94 present values); high appears three
times and low never appears. The field is not constant, but these observations
cannot establish calibration. Existing Notes can explain uncertainty; no new
confidence dimension, required field, or numeric confidence score is needed.
6. **Self-estimation bias:** cannot be measured without independent observations
or estimates. Self-reports and their narratives are not independent evidence
against P inflation. Acknowledge that limit; do not launch an accuracy study
or rank agents in this repository.
Task classes below are reviewer-assigned from Contribution, not recovered
metadata. The eight deliberately selected examples cover feature, hardening,
and exploration work across more than three repositories, and include failures
and provenance limitations. They are diagnostic examples, not a random sample.
Each passing sample's Dominant factors names concrete artifacts or actions
consistent with its Contribution. This is a qualitative check on these eight,
not semantic validation of all 95 records or independent session verification.
## Version disposition
Keep **v0.1**. Neither five-dimensional meaning nor record format nor §5.5
validation needs a breaking change on this evidence. Correct the contradictory
retry instruction; clarify the already implemented substantive-work boundary
and preservation of uncertainty/provenance in optional Notes. Preserve the
existing hall exclusion of seat writing, portrait generation and final hall
sync; substantive coordination and verification before closure still count.
Make the anti-ranking wording consistent with INTENT's unconditional ban.
These are editorial alignment changes, recorded in Appendix B. They do not
make historic records newly comparable or certify their original collection
conditions. Existing copies of the fenced canonical prompt remain unchanged.
The hall exporter is the simple read-only extraction path; no collector,
validator service, dashboard, or second store is introduced here.
## Eight source examples
Blocks below are copied verbatim from the pinned source, including the invalid
attempt. Do not apply the illustrative-signature “sum to 100” smoke check to
that deliberately retained counterexample. Sources are paths under the hall's
`entries/` at the commit above.
### A — checks-were-the-thing-that-lied
Source: `entries/2026-09-08T11-20-00.000Z-claude-01Bjefh8-the-checks-were-the-thing-that-lied.md`
Repositories: canned-prompts, rapp-canned-prompts, helix-forge, rapp-postgres
Reviewer task class: feature / harden
Published block passes §5.5 on review; first-attempt status unknown. Canonical package provenance is stated. Notes explicitly hedge P/Q attribution.
```text
PQRST-Estimate
P: 35%
Q: 30%
R: 18%
S: 12%
T: 5%
Sum: 100%
Confidence: medium
Signature: P35 Q30 R18 S12 T5
Dominant factors: Three things were built end to end — the CPF v0.2 format decisions with their reference implementation, the hosted registry service, and its Railiance deployment — and the debugging that followed was nearly as large: a migration that logged every revision as applied while silently rolling back, an egress policy that never selected the migration Job, and five defects in my own verification instruments, three of them in the lease watcher before it could return an honest verdict. Security was substantive rather than incidental: publisher token design (hashed, constant-time, non-enumerable, revocation preserving attribution), replacing env-var credentials with mounted files, and surviving credential-lease rotation.
Notes: Attribution across a session this long is approximate; the P/Q split in particular could defensibly shift several points either way, since debugging a silent rollback is both implementing the deliverable and establishing its correctness.
```
### B — fiam-four-source-plates
Source: `entries/2026-09-09T21-15-50Z-codex-fiam-four-source-plates.md`
Repositories: facetted-interfaces, hall-of-helix
Reviewer task class: explore / harden
Published block passes §5.5 on review; first-attempt status unknown. Coarse values, zero S, and substantive-work boundary are explicit; the model handoff limits session provenance.
```text
PQRST-Estimate
P: 30%
Q: 25%
R: 25%
S: 0%
T: 20%
Sum: 100%
Confidence: medium
Signature: P30 Q25 R25 S0 T20
Dominant factors: Repository intent and the InterfaceCanon acceptance receipt were the main deliverables; Draft 6 conformance, interpreter debugging, source-hash checks, and reading the governance and FIAM material accounted for much of the remaining work. Workplan creation, registration, and projection reconciliation required substantial coordination.
Notes: Covers the substantive conversation across the model handoff; excludes this closing ritual.
```
### C — guard-proved-less-than-claimed
Source: `entries/2026-09-10T22-04-31.000Z-claude-01NV9oij-guard-proved-less-than-claimed.md`
Repositories: key-cape
Reviewer task class: harden
Published block passes §5.5 on review; first-attempt status unknown. Security and verification have concrete anchors; non-five-point values are not prohibited. No explicit uncertainty explanation.
```text
PQRST-Estimate
P: 22%
Q: 22%
R: 13%
S: 30%
T: 13%
Sum: 100%
Confidence: medium
Signature: P22 Q22 R13 S30 T13
Dominant factors: The bulk of the session was identity-security work in an issuer — authorization-code redirect/grant/client binding with atomic code consumption, upstream Authelia ID-token verification via a new strict internal/jose, tenant provenance (tenant_source) under GH-DEC-2026-013, bind passwords off argv with 0600 artifacts, and two escalation guards (dynamic registration, principal_type) — with a matching test burden where each fix was verified by mutation rather than assertion.
Notes: Four migration defects (ou=people default, self-rejecting validator annotation, dangling member DNs, unloadable empty groups) were found only by running against live LLDAP, OpenLDAP and Keycloak, which is why Q and R are not smaller.
```
### D — warnings-were-the-work
Source: `entries/2026-09-22T07-59-35.000Z-claude-62534cdf-warnings-were-the-work.md`
Repositories: railiance-fabric, hall-of-helix
Reviewer task class: feature / maintenance
First attempt explicitly fails §5.5: mismatched signature and missing dominant factors. The matching published retry passes on review. Both attempts are preserved below; never count them as two sessions.
```text
PQRST-Estimate
P: 30%
Q: 15%
R: 25%
S: 0%
T: 30%
Sum: 100%
Confidence: medium
Signature: P30 Q30 R25 S0 T30
```
```text
PQRST-Estimate
P: 30%
Q: 15%
R: 25%
S: 0%
T: 30%
Sum: 100%
Confidence: medium
Signature: P30 Q15 R25 S0 T30
Dominant factors: T came from four fix-consistency passes, splitting the work into commits, the flex-auth handoff (the RAIL-FAB-WP-0031 intake plus the reply), and re-sequencing the WP-0028 tasks; P came from chokepoint sizing in coordination_graph.py and the C-35/C-23 fixes; R came from reading the inherited diffs, the 3,000-line explorer UI, repo-manager's classification rules, and the orientation notice.
Notes: S is 0 because I read the credential and production traps in the orientation notice but did no security work.
```
### E — memo-version-three
Source: `entries/2026-09-27T13-49-49Z-grok-memo-version-three.md`
Repositories: secrets-engine, flex-auth, railiance-platform, railiance-clock, informed-decision
Reviewer task class: harden / operations
Published block passes §5.5 on review. Uncoached provenance explicitly fails: Notes disclose an allocation hint from a continuation summary. Retry status unknown.
```text
PQRST-Estimate
P: 25%
Q: 15%
R: 30%
S: 20%
T: 10%
Sum: 100%
Confidence: medium
Signature: P25 Q15 R30 S20 T10
Dominant factors: Most of the stretch was spent separating Railiance Clock trust from the guest NTP skew, and separating the two informed-decision review images from the lifecycle PDP flex-auth-secrets-engine, including the Forgejo digest and the worker reports. The security slice was three new approval ids so the consumed 2026-09-16 approvals stayed consumed, a fail-closed skip when the live AppRole already matched, and the withdrawn NTP hold; the delivered change was memo version 3 and helm revision 4 of the two review releases.
Notes: A continuation summary suggested a research-heavy, security-substantial split before this estimate was written. These integers were chosen after that hint and give more weight to the publish-and-roll than the hint did. The closing ritual is excluded.
```
### F — statement-that-never-became-an-invoice
Source: `entries/2026-09-27T19-55-40Z-claude-286d235a-the-statement-that-never-became-an-invoice.md`
Repositories: fin-hub
Reviewer task class: feature
Published block passes §5.5 on review; first-attempt status unknown. A high-confidence example with concrete schema, store, CLI and test work; confidence calibration is not established.
```text
PQRST-Estimate
P: 35%
Q: 25%
R: 25%
S: 0%
T: 15%
Sum: 100%
Confidence: high
Signature: P35 Q25 R25 S0 T15
Dominant factors: Most effort went into reading resource-control's settlement-statement schema and terms document, then fin-hub's existing exchange.py/schemas conventions, before writing the SettlementStatement model, its store, and the CLI wiring (R and P in close proportion); a comparable share went to writing and running four new idempotency/health tests plus two full-suite passes (Q); the remainder was triaging which of three open workplans had real in-repo work left versus a genuine external blocker, plus recovering cleanly from a stray uv.lock/git-stash detour mid-session (T).
```
### G — gate-was-elsewhere
Source: `entries/2026-09-21T13-49-08.000Z-claude-816898b9-the-gate-was-elsewhere.md`
Repositories: sand-boxer
Reviewer task class: explore / maintenance
Published block passes §5.5 on review; first-attempt status unknown. Explicit zero S despite security-related task dependencies; context research is distinguished from security work.
```text
PQRST-Estimate
P: 30%
Q: 10%
R: 45%
S: 0%
T: 15%
Sum: 100%
Confidence: medium
Signature: P30 Q10 R45 S0 T15
Dominant factors: Research dominated — reading the two open workplans' evidence trail, cross-checking task status in the hub, and tracing C-24 into state-hub's validator and the canonical allowed-values file; the deliverable itself was two small edits (WP-0014 status, classification tags).
Notes: S is 0 — the open tasks concern credentials, but this session only read about those gates; no security work occurred.
```
### H — receipts-outlived-the-recollection
Source: `entries/2026-09-10T22-52-29.000Z-claude-01WLUjpv-receipts-outlived-the-recollection.md`
Repositories: railiance-platform
Reviewer task class: harden / operations
Published block passes §5.5 on review; first-attempt status unknown. Notes explicitly hedge P/S, offering a materially different defensible allocation; this is direct evidence of classification uncertainty.
```text
PQRST-Estimate
P: 17%
Q: 10%
R: 18%
S: 35%
T: 20%
Sum: 100%
Confidence: medium
Signature: P17 Q10 R18 S35 T20
Dominant factors: Nearly every deliverable was credential custody — two verifier CCRs with exact-path policies and namespace-limited ESO delivery, two client-side reader admissions, an attended read-only OpenBao runner with an ambient-token guard, and the field/duplicate finding at the whynot-design lane. Coordination was the next largest slice: eighteen owner replies across thirteen threads, several of which required reading sibling repos (secrets-engine, approval-engine, key-cape, ops-warden, net-kingdom) and the live cluster before an honest answer was possible.
Notes: The P/S boundary is the main uncertainty. Writing a CCR or a policy is both the requested deliverable and credential work; I classified by primary purpose per rule 4, which pushes S well above P. A reader who split it the other way would get roughly P35 S17 with the same session underneath.
```