test-driver/research/findings/F-0002-scenario-coverage-gap.md
tegwick 4ddb2f896c T05: the lab and its labelled mutation catalogue
lab/app.py (users, tenants, auth, resources, sharing, read/write, revoke,
audit), lab/http_api.py (JSON API + browser UI, stdlib only), 20 labelled
composable version-stamped mutations, ground-truth matrix. 48 tests pass.

Detection against the reference scenario: MECHANICAL 0/10 flagged (correct),
DEFECT 6/6, SEMANTIC 2/4 with both inert cases declared.

- F-0002: M16 and M18 initially escaped detection entirely. A use case
  protects exactly what it asserts. Resolved by adding two claims already
  stated as intent in INTENT.md; the six-mutation catalogue would never have
  surfaced this.
- test-id axis added: stable selectors survive most UI mutations, which would
  make H-001 trivially false. Mutations now vary on preserves_test_ids so the
  hypothesis is analysed split by that axis rather than rigged.
- M12 (semantic deferred revoke) and M19 (defect race) are behaviourally
  identical and asserted as such - the discrimination problem as a test.

lab/minimal.py removed; superseded by lab/app.py.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

Assistant: claude-code
Assistant-Model: opus
Assistant-Process: 1629012@bnt-lap001
Assistant-Session: 78d4fb13-8a1e-474b-87a3-9b9261c49a39
2026-08-22 23:31:22 +02:00

3.2 KiB

id type class status discovered resolved discovered_by workplan task
F-0002 framework-finding FRAMEWORK_LIMITATION resolved 2026-08-22 2026-08-22 TD-WP-0002-T05 TD-WP-0002 TD-WP-0002-T05

F-0002 — Two seeded defects were invisible to the reference scenario

Observation

On first running the reference scenario against the twenty labelled mutations, six of six DEFECT mutations should have failed; only four did.

Mutation Defect Verdict before Why it escaped
M16 A READ grant confers WRITE PASS No claim mentioned writing. The scenario never attempted one.
M18 Revocation is not audited PASS i-audit-append-only checks ordering, not completeness. A trail missing an entry is still ordered.

Both passed cleanly. Nothing was flaky, nothing was ambiguous, and no oracle reported INCONCLUSIVE — the framework simply had nothing to say, confidently.

Why it matters

This is the failure mode most likely to be mistaken for success. A green run against a lab carrying a seeded privilege escalation looks exactly like a green run against a correct system. Had the catalogue been the six mutations the milestones document originally sketched, this would not have surfaced at all — which is the concrete argument for the larger catalogue, now made from evidence rather than from assertion.

It also sharpens what a verification asset is: a use case protects exactly what it asserts, and not one thing more. Coverage is a property of the claim set, not of the framework. No amount of adaptation, crystallization or energy scoring compensates for an assertion nobody wrote.

Resolution

Path 1 — the implementation changes to match the concept.

Two claims were added to the reference use case, both Provenance.HUMAN and both derivable from INTENT.md rather than from watching the lab:

  • c-bob-cannot-write — "A READ grant does not let Bob write R". INTENT.md § Security by Use-Case Mutation already derives this exact question from the reference use case: "Can Bob write when only read permission was granted?" It was always part of what sharing means; it had simply never been written down as an assertion.
  • c-revoke-audited — "Revocation is recorded in the audit trail". Enforcement being correct is not sufficient: an access change nobody can later evidence is a compliance failure even when the access itself is right.

Supporting changes: the lab gained a write path and the observation channel a non-destructive probe_write.

Both defects are now detected. DEFECT detection is 6/6, and test_every_defect_is_detected fails the suite if that ever regresses.

Note on provenance

These claims were added after observing that mutations escaped, which is uncomfortably close to fitting assertions to the lab. They are admissible because both were already stated as intent in INTENT.md before any lab existed — the finding revealed a transcription gap, not a new requirement. Had the intended behaviour not already been on record, the correct resolution would have been to escalate to a human, not to write the claim.

That distinction is exactly what Provenance exists to make checkable, and this finding is the first case where it did real work.