test-driver/docs/TestDriverClassificationDesign.md

202 lines
9.9 KiB
Markdown
Raw Permalink Normal View History

# Adaptation Classification and Intent Provenance
**Status:** v0.1 — design decision
**Workplan:** `TD-WP-0002-T02`
**Supersedes nothing.** Extends `TestDriverConceptModel.md` §§ 2.3, 7.8, 9.4.
This note answers the two questions that gate every implementation task in
`TD-WP-0002`. Both were unanswered in the concept corpus, and both are cheaper to
answer on paper than to discover at canonical M6.
---
## Part A — The classifier
### The question
What **deterministic evidence** separates *"the sharing button moved"* from
*"Bob can still read R after revocation"*?
The concept corpus assumes such a discriminator exists but never names it. If the
honest answer were "a model decides", Concept Model § 2.3 would be violated and
the framework's central safety claim would collapse.
### The reframe
The question as posed is harder than the question the framework actually has to
answer, because it presumes a single classifier looking at an undifferentiated
failure and deciding what kind it is.
Instead, **evidence is stratified**, and each stratum has a different authority:
| Stratum | Content | Gathered by | May a model touch it? |
|---|---|---|---|
| **S1 — Surface** | selectors, DOM, URLs, route shapes, coordinates, screenshots | the actor, incidentally, while realizing an action | **yes** — this is mechanics |
| **S2 — Realization** | did the semantic action's declared postcondition hold; which surface was actually used | an Observer, out-of-band from the actor | no |
| **S3 — Judgment** | claims and invariants evaluated against domain-state observations | Observers + Oracles, independent of the actor | **never** |
S1 is how. S2 is whether it happened. S3 is whether that was correct.
### The structural rule
> **A change in S1 can never license a change in S3.**
This is enforced architecturally rather than by classifier accuracy:
- **Claims and invariants are inputs to a run, not outputs of it.** The adaptation
machinery has no write path to them. There is no code path by which a retry, a
re-realization, or a learned trajectory can alter what is asserted.
- **Adaptation is confined to the realization loop.** It may change *how* a
semantic action is performed. It cannot change *what follows from* performing it.
- Therefore the classifier is never asked "is this failure acceptable?" It is only
asked "why did this run fail to reach the point of judgment?"
This makes `False Adaptation Rate = 0` an **architectural property, not a tuning
target**. The system does not merely try hard to avoid normalizing a defect — it
cannot express the operation. That matters because a rate driven near zero by
tuning regresses silently; a rate that is zero by construction does not.
### The decision table
Classification is a total function over three observed signals. Baseline refers
to the last accepted run of the same verification asset.
| S1 surface differs | S2 postcondition met | S3 verdicts vs baseline | Classification |
|---|---|---|---|
| yes | yes | unchanged | `MECHANICAL_ADAPTATION` |
| no | yes | changed to FAIL | `PRODUCT_DEFECT` |
| yes | yes | changed to FAIL | `PRODUCT_DEFECT` |
T08: the classifier, measured and attacked False Adaptation Rate = 0/7 across the labelled catalogue and the three E-003 attacks. 11 of 12 mechanical mutations absorbed without a human, so the safety result is not bought by escalating everything. - classification.py: total function over three signals, rule order chosen so every rule that could excuse a regression sits after the rule that reports one. SAFE_TO_ACCEPT is a two-element closed set, asserted. - CompositeDriver plus scenarios/full_journey.py: one asset crossing both surfaces, so UI mutations are visible as surface differences while the claims they do not touch stay green. - E-003: surface substitution (new M23), concurrent mechanical+defect, evidence starvation, provenance laundering. All held. F-0006 (CONCEPT_DRIFT, resolved): the T02 design listed SEMANTIC_CHANGE as an outcome the table could produce. It cannot - M12 and M19 are behaviourally identical, as the lab has asserted since T05. PRODUCT_DEFECT and SEMANTIC_CHANGE collapse into one escalating outcome, BEHAVIOUR_CHANGED, and the distinction becomes a human adjudication. INTENT_CHANGED survives but is detected by the claim fingerprint moving, not inferred from behaviour. Two classifier defects found and fixed rather than reported: claims downstream of a failed realization now yield INCONCLUSIVE rather than FAIL (a false accusation is the mirror image of a false adaptation), and the browser driver records a page signature so surface change is detectable when the interaction path is unchanged. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Assistant: claude-code Assistant-Model: opus Assistant-Process: 1629012@bnt-lap001 Assistant-Session: 78d4fb13-8a1e-474b-87a3-9b9261c49a39
2026-08-23 00:02:58 +02:00
| any | yes | unchanged, but the asset's declared claim set differs from the use case | ~~`SEMANTIC_CHANGE` → human~~**`INTENT_CHANGED`**, see below |
| any | no, and no legitimate surface affords it | — | `PRODUCT_DEFECT` |
| any | no, but the action is expressible and the actor simply failed | — | `FRAMEWORK_LIMITATION` |
| any | any | any oracle `INCONCLUSIVE`, or required evidence missing | `AMBIGUOUS` → escalate |
T08: the classifier, measured and attacked False Adaptation Rate = 0/7 across the labelled catalogue and the three E-003 attacks. 11 of 12 mechanical mutations absorbed without a human, so the safety result is not bought by escalating everything. - classification.py: total function over three signals, rule order chosen so every rule that could excuse a regression sits after the rule that reports one. SAFE_TO_ACCEPT is a two-element closed set, asserted. - CompositeDriver plus scenarios/full_journey.py: one asset crossing both surfaces, so UI mutations are visible as surface differences while the claims they do not touch stay green. - E-003: surface substitution (new M23), concurrent mechanical+defect, evidence starvation, provenance laundering. All held. F-0006 (CONCEPT_DRIFT, resolved): the T02 design listed SEMANTIC_CHANGE as an outcome the table could produce. It cannot - M12 and M19 are behaviourally identical, as the lab has asserted since T05. PRODUCT_DEFECT and SEMANTIC_CHANGE collapse into one escalating outcome, BEHAVIOUR_CHANGED, and the distinction becomes a human adjudication. INTENT_CHANGED survives but is detected by the claim fingerprint moving, not inferred from behaviour. Two classifier defects found and fixed rather than reported: claims downstream of a failed realization now yield INCONCLUSIVE rather than FAIL (a false accusation is the mirror image of a false adaptation), and the browser driver records a page signature so surface change is detectable when the interaction path is unchanged. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Assistant: claude-code Assistant-Model: opus Assistant-Process: 1629012@bnt-lap001 Assistant-Session: 78d4fb13-8a1e-474b-87a3-9b9261c49a39
2026-08-23 00:02:58 +02:00
> **Revised at T08 — see `research/findings/F-0006-classifier-cannot-infer-intent.md`.**
> `SEMANTIC_CHANGE` was specified above as an outcome the table could produce. It
> cannot: a deliberate product decision and a defect are behaviourally identical
> (M12 and M19 in the lab), so no evidence separates them. `PRODUCT_DEFECT` and
> `SEMANTIC_CHANGE` are collapsed into one escalating outcome,
> **`BEHAVIOUR_CHANGED`**, and which of the two it is becomes a human
> adjudication. `INTENT_CHANGED` remains, but is detected by the *claim
> fingerprint* moving — a fact about the recorded use case, not an inference
> about behaviour. The row above described that, filed under the wrong heading.
Two rows carry the whole safety argument:
- **Row 3** — a surface change occurring *alongside* a verdict change is classified
as a defect, never as an adaptation. Coincidence is not exoneration. This is the
row that a naive self-healing tool gets wrong, and it is where the seeded M05
authorization defect must land even when it ships together with M01/M02.
- **Row 7** — missing or contradictory evidence yields `AMBIGUOUS`, never a
default-to-pass and never a default-to-adapt. `INCONCLUSIVE` already exists as a
first-class verdict; this row is what gives it teeth.
### Realization legitimacy
S2 records not only *whether* the postcondition held but *which surface produced
it*. A semantic action declares the surfaces it may legitimately use; an actor
that achieves `grant_access` by calling the HTTP API when the scenario constrains
it to the browser has not realized the action — it has performed an unrequested
**surface-substitution**, which is itself one of the catalogued security
mutations.
Without this check, an agent can "recover" from a UI authorization control by
routing around it and the run would score as a successful mechanical adaptation.
That is the most plausible concrete route to a false adaptation, so it is checked
explicitly rather than assumed away.
### Where model judgment is and is not used
Concept Model § 2.3 is satisfied, and the boundary is now precise:
- **Used:** locating a moved control, proposing a path through a changed DOM,
summarizing a trajectory. All S1. All independently verified afterwards by S2.
- **Not used:** evaluating a claim, evaluating an invariant, producing a verdict,
deciding whether an adaptation was legitimate, deciding whether a run passed.
A model can be wrong about where the button is; the run then fails at S2 and is
retried or reported. A model is never in a position to be wrong about whether Bob
could still read R.
### What this costs
The design requires an out-of-band observer capable of evaluating every semantic
action's postcondition and every claim without going through the actor's surface.
In the lab this is a direct database or admin-API reader. For a real system under
test it is the main integration burden the framework imposes, and it should be
stated plainly to any future adopter: **test-driver requires an independent
observation channel.** Without one, S2 and S3 collapse into S1 and the safety
argument is void.
This is a real limitation and is recorded as such rather than discovered later.
---
## Part B — Intent provenance
### The question
If an agent writes the implementation, an agent writes the use case, and an agent
realizes the test, then "oracles are independent from actors" is *procedurally*
true and *epistemically* empty. Independence of components does not give
independence of belief when all components derive from the same source.
### The rule
> A claim may only be authored by a source **causally independent of the
> implementation it constrains.**
`UseCase` and `Claim` therefore carry a required `provenance` field:
| Value | Meaning | Admissible as a claim? |
|---|---|---|
| `human` | authored by a person | yes |
| `spec` | derived from a specification artifact that demonstrably predates the implementation (ref + commit SHA recorded) | yes |
| `agent-from-spec` | an agent transcribed a human or spec artifact; a human accepted the result | yes, with the accepting identity recorded |
| `agent-from-implementation` | an agent derived it by observing what the system does | **no** |
`agent-from-implementation` is not forbidden as an activity — it is genuinely
useful for T0 exploration. It is forbidden as a **claim**. Such output enters the
system as an *exploratory hypothesis*, which can drive runs and accumulate
evidence but cannot produce a `FAIL`, cannot enter an invariant, and cannot be
crystallized. Promotion to a claim requires an explicit human acceptance event,
recorded with identity and timestamp.
Without this rule, the framework's guarantee reduces to "the implementation
agrees with itself", which is exactly the failure the project exists to prevent.
### Consequences
- Provenance is carried into every Evidence Pack, so any verdict can later be
audited for independence. A `FAIL` whose claim has inadmissible provenance is
itself an `EVIDENCE_FAILURE`.
- Crystallization preserves provenance. A deterministic test generated from an
agentic ancestor inherits the ancestor's claim provenance and cannot launder it.
- A fourth self-test joins the M3 set: `td://self/intent-independence` — asserting
that no admissible claim in the corpus has implementation-derived provenance.
### Honest limitation
This rule constrains *provenance*, not *quality*. A human can write a claim that
merely restates what they saw the system do. The rule makes the independence
question answerable and auditable; it does not make it automatic. That residual
is accepted for the spike and revisited at T10.
---
## Decisions recorded
- **D-01** Evidence is stratified S1/S2/S3; adaptation may write only to S1.
- **D-02** Claims and invariants are run inputs; no adaptation path may modify them.
- **D-03** Classification is a total function over (S1 differs, S2 met, S3 vs baseline),
defaulting to `AMBIGUOUS` on missing evidence.
- **D-04** A surface change concurrent with a verdict change classifies as
`PRODUCT_DEFECT`, never `MECHANICAL_ADAPTATION`.
- **D-05** Semantic actions declare permitted surfaces; realization via another
surface is a finding, not a recovery.
- **D-06** Claims require causally independent provenance;
`agent-from-implementation` output is an exploratory hypothesis until a human
promotes it.
- **D-07** test-driver requires an independent observation channel into the system
under test. Stated as an adoption precondition.