Preserve oracle semantics in generated regression judgments
Assistant: codex Assistant-Model: gpt-6-astra Assistant-Session: 01a0e76f-be98-7ae3-965d-e0b31290a4c4
This commit is contained in:
parent
10077edb8c
commit
3ce772466b
8 changed files with 224 additions and 46 deletions
|
|
@ -147,3 +147,20 @@ Validation: 12 new regressions, of which 10 failed before the fixes; **131 focus
|
|||
tests passed**, and the full suite passed **440 tests in 169.07 seconds**.
|
||||
`git diff --check` is clean. Decision: `aa6c3e31-614c-4030-990c-945acef08515`.
|
||||
No new tasks/workplans; T05 remains waiting and this workplan remains blocked.
|
||||
|
||||
|
||||
**2026-09-28 generated-judgment follow-up — done (T04).** Generated regression
|
||||
modules now embed the actual predicate evaluator used by Oracle, retaining strict
|
||||
boolean and missing/exception/invalid-result INCONCLUSIVE semantics. The adapter
|
||||
evaluates all protected assertions so FAIL dominates any INCONCLUSIVE result;
|
||||
otherwise an inconclusive result becomes a stdlib SkipTest explicitly labeled
|
||||
INCONCLUSIVE in pytest. Callers must inspect skipped outcomes, not treat exit zero
|
||||
as complete verification. The checked-in descendant was regenerated with its
|
||||
original predicate set, frozen route and lineage. No runtime test-driver import
|
||||
or additional dependency is introduced into generated artifacts.
|
||||
|
||||
Validation: the initial 30 parity cases reproduced 24 failures and six controls;
|
||||
all 31 final parity cases pass, including native pytest reporting in a subprocess.
|
||||
The related subset passed 93 tests. Full suite: **471 passed in 167.82 seconds**.
|
||||
`git diff --check` is clean. Decision: `c06c8752-80db-4458-9eaf-6321a0ca710c`.
|
||||
No new task/workplan; T05 remains waiting and this workplan remains blocked.
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue