lab/app.py (users, tenants, auth, resources, sharing, read/write, revoke, audit), lab/http_api.py (JSON API + browser UI, stdlib only), 20 labelled composable version-stamped mutations, ground-truth matrix. 48 tests pass. Detection against the reference scenario: MECHANICAL 0/10 flagged (correct), DEFECT 6/6, SEMANTIC 2/4 with both inert cases declared. - F-0002: M16 and M18 initially escaped detection entirely. A use case protects exactly what it asserts. Resolved by adding two claims already stated as intent in INTENT.md; the six-mutation catalogue would never have surfaced this. - test-id axis added: stable selectors survive most UI mutations, which would make H-001 trivially false. Mutations now vary on preserves_test_ids so the hypothesis is analysed split by that axis rather than rigged. - M12 (semantic deferred revoke) and M19 (defect race) are behaviourally identical and asserted as such - the discrimination problem as a test. lab/minimal.py removed; superseded by lab/app.py. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Assistant: claude-code Assistant-Model: opus Assistant-Process: 1629012@bnt-lap001 Assistant-Session: 78d4fb13-8a1e-474b-87a3-9b9261c49a39
3.2 KiB
| id | type | class | status | discovered | resolved | discovered_by | workplan | task |
|---|---|---|---|---|---|---|---|---|
| F-0002 | framework-finding | FRAMEWORK_LIMITATION | resolved | 2026-08-22 | 2026-08-22 | TD-WP-0002-T05 | TD-WP-0002 | TD-WP-0002-T05 |
F-0002 — Two seeded defects were invisible to the reference scenario
Observation
On first running the reference scenario against the twenty labelled mutations, six of six DEFECT mutations should have failed; only four did.
| Mutation | Defect | Verdict before | Why it escaped |
|---|---|---|---|
| M16 | A READ grant confers WRITE | PASS |
No claim mentioned writing. The scenario never attempted one. |
| M18 | Revocation is not audited | PASS |
i-audit-append-only checks ordering, not completeness. A trail missing an entry is still ordered. |
Both passed cleanly. Nothing was flaky, nothing was ambiguous, and no oracle
reported INCONCLUSIVE — the framework simply had nothing to say, confidently.
Why it matters
This is the failure mode most likely to be mistaken for success. A green run against a lab carrying a seeded privilege escalation looks exactly like a green run against a correct system. Had the catalogue been the six mutations the milestones document originally sketched, this would not have surfaced at all — which is the concrete argument for the larger catalogue, now made from evidence rather than from assertion.
It also sharpens what a verification asset is: a use case protects exactly what it asserts, and not one thing more. Coverage is a property of the claim set, not of the framework. No amount of adaptation, crystallization or energy scoring compensates for an assertion nobody wrote.
Resolution
Path 1 — the implementation changes to match the concept.
Two claims were added to the reference use case, both Provenance.HUMAN and both
derivable from INTENT.md rather than from watching the lab:
c-bob-cannot-write— "A READ grant does not let Bob write R".INTENT.md§ Security by Use-Case Mutation already derives this exact question from the reference use case: "Can Bob write when only read permission was granted?" It was always part of what sharing means; it had simply never been written down as an assertion.c-revoke-audited— "Revocation is recorded in the audit trail". Enforcement being correct is not sufficient: an access change nobody can later evidence is a compliance failure even when the access itself is right.
Supporting changes: the lab gained a write path and the observation channel a
non-destructive probe_write.
Both defects are now detected. DEFECT detection is 6/6, and
test_every_defect_is_detected fails the suite if that ever regresses.
Note on provenance
These claims were added after observing that mutations escaped, which is
uncomfortably close to fitting assertions to the lab. They are admissible because
both were already stated as intent in INTENT.md before any lab existed — the
finding revealed a transcription gap, not a new requirement. Had the intended
behaviour not already been on record, the correct resolution would have been to
escalate to a human, not to write the claim.
That distinction is exactly what Provenance exists to make checkable, and this
finding is the first case where it did real work.