T01: audit every InnerLoop rule, and make the checkable ones executable
41 rules classified executable / checkable / decorative, each tagged with
the failure class it catches. Counts: 11 executable, 22 checkable, 4
decorative (one of them dead policy).
Audit: history/260731-inner-loop-rule-audit.md
New tools/loop-lint.py makes 7 rules executable (tier declared, chaos
roll recorded, tier-L review trail, unmeasured-in-evidence, whole-file
loadability, reporting tools expose --self-test). It found three real
violations on its first run, none previously visible:
- specs/ArchitectureBlueprint.md was 543 lines against a ~400 limit
the loop has stated since v0.2 and never measured. Split at its own
section boundaries into Blueprint (1-8) + Runtime (9-15).
- tools/dep-weight.py and tools/rule-coverage.py had positive-control
logic and no --self-test, so nothing verified the control worked.
Adding rule-coverage's self-test exposed a latent instance of the exact
class this workplan is about: if the spec regex stopped matching, rules
was empty, missing was empty, and the tool exited 0 reporting "0/0" --
a silent pass, in the tool that reports our headline AM-1 number. Both
tools now assert they found something before reporting.
Two demotions applied in the spec rather than left implicit: "structured
over prose" is marked guidance (nothing can check it), and the 8k/10k
token budget is struck through and marked DEAD POLICY pointing at T05.
The audit's uncomfortable finding: rule 13 (re-derive inherited numbers)
has no mechanical form, is deliberately left decorative, and caught the
LARGEST error in CB-WP-0002. That is a counter-example to this
workplan's own hypothesis. "A rule that cannot be executed is not a
rule" is wrong as stated; the defensible version is that such a rule
cannot be relied on to fire, so it must not be the only defence for a
class that matters.
Class coverage: harness-does-nothing has five executable rules;
trusted-arithmetic has ZERO and produced the largest single error.
make loop-lint and make self-tests wired into `make all` and CI.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
parent
ee1ee485b2
commit
fed422a3a3
10 changed files with 827 additions and 242 deletions
|
|
@ -5,30 +5,107 @@ Compares the rule IDs declared in specs/GroundRules.md against the
|
|||
`covers:` lists in scenarios/ground/*.yaml. Exits non-zero when a
|
||||
scenario claims a rule the spec does not define, so coverage can never
|
||||
be inflated by a typo'd or invented rule ID.
|
||||
|
||||
Stated limit (InnerLoop implementation rule 4): this gate counts tags. It
|
||||
proves no rule is unclaimed and no claimed rule is invented. It does NOT
|
||||
prove a scenario exercises the rule it names.
|
||||
|
||||
Positive control (InnerLoop v1.1 §Step 5): the run asserts it actually
|
||||
found rules and scenarios. Before this was added, a broken spec regex
|
||||
yielded rules=[] and missing=[] and the tool exited 0 reporting "0/0" —
|
||||
the harness-does-nothing class, in the tool that reports our headline
|
||||
coverage number.
|
||||
|
||||
Usage:
|
||||
python3 tools/rule-coverage.py
|
||||
python3 tools/rule-coverage.py --self-test
|
||||
"""
|
||||
import glob
|
||||
import re
|
||||
import sys
|
||||
|
||||
spec = open("specs/GroundRules.md").read()
|
||||
rules = sorted(set(re.findall(r"\*\*(GR-[A-Z]+\d+)", spec)))
|
||||
RULE_RE = r"\*\*(GR-[A-Z]+\d+)"
|
||||
COVERS_RE = r"covers: \[(.*?)\]"
|
||||
|
||||
covered = set()
|
||||
for path in sorted(glob.glob("scenarios/ground/*.yaml")):
|
||||
match = re.search(r"covers: \[(.*?)\]", open(path).read(), re.S)
|
||||
if match:
|
||||
covered |= {c.strip() for c in match.group(1).split(",") if c.strip()}
|
||||
|
||||
known = set(rules)
|
||||
hit = sorted(known & covered)
|
||||
missing = [r for r in rules if r not in covered]
|
||||
invented = sorted(covered - known)
|
||||
def parse_rules(spec_text):
|
||||
return sorted(set(re.findall(RULE_RE, spec_text)))
|
||||
|
||||
pct = 100 * len(hit) // len(rules) if rules else 0
|
||||
print(f"AM-1 rule coverage: {len(hit)}/{len(rules)} ({pct}%)")
|
||||
if missing:
|
||||
print(" uncovered:", " ".join(missing))
|
||||
if invented:
|
||||
print(" ERROR — claimed but not defined in the spec:", " ".join(invented))
|
||||
sys.exit(1)
|
||||
sys.exit(0 if not missing else 2)
|
||||
|
||||
def parse_covers(text):
|
||||
match = re.search(COVERS_RE, text, re.S)
|
||||
if not match:
|
||||
return set()
|
||||
return {c.strip() for c in match.group(1).split(",") if c.strip()}
|
||||
|
||||
|
||||
def self_test():
|
||||
"""Each assertion pins a failure this tool must detect."""
|
||||
results = []
|
||||
|
||||
def check(name, ok, detail=""):
|
||||
results.append((name, ok, detail))
|
||||
|
||||
# The defect that motivated this control: a spec that parses to zero
|
||||
# rules must not be reportable as coverage.
|
||||
check("zero rules detected as a failure", parse_rules("no rules here") == [],
|
||||
"empty spec yields no rules; main() now aborts on this")
|
||||
# The matcher must actually match the real format.
|
||||
check("rule matcher works on real spec format",
|
||||
parse_rules("**GR-R06** something\n**GR-A12** other")
|
||||
== ["GR-A12", "GR-R06"])
|
||||
# covers: parsing, including the empty case.
|
||||
check("covers matcher works", parse_covers("covers: [GR-R06, GR-A12]")
|
||||
== {"GR-R06", "GR-A12"})
|
||||
check("missing covers yields empty set", parse_covers("no covers key") == set())
|
||||
|
||||
print("rule-coverage self-test (positive control)")
|
||||
ok = True
|
||||
for name, passed, detail in results:
|
||||
print(f" [{'ok ' if passed else 'FAIL'}] {name}"
|
||||
+ (f" — {detail}" if detail else ""))
|
||||
ok &= passed
|
||||
return 0 if ok else 1
|
||||
|
||||
|
||||
def main():
|
||||
if "--self-test" in sys.argv:
|
||||
return self_test()
|
||||
|
||||
rules = parse_rules(open("specs/GroundRules.md").read())
|
||||
paths = sorted(glob.glob("scenarios/ground/*.yaml"))
|
||||
|
||||
# Positive control: refuse to report a percentage over nothing.
|
||||
if not rules:
|
||||
print("ERROR — no GR-rules parsed from specs/GroundRules.md; "
|
||||
"refusing to report coverage", file=sys.stderr)
|
||||
return 1
|
||||
if not paths:
|
||||
print("ERROR — no scenarios found in scenarios/ground/; "
|
||||
"refusing to report coverage", file=sys.stderr)
|
||||
return 1
|
||||
|
||||
covered = set()
|
||||
for path in paths:
|
||||
covered |= parse_covers(open(path).read())
|
||||
|
||||
known = set(rules)
|
||||
hit = sorted(known & covered)
|
||||
missing = [r for r in rules if r not in covered]
|
||||
invented = sorted(covered - known)
|
||||
|
||||
pct = 100 * len(hit) // len(rules)
|
||||
print(f"AM-1 rule coverage: {len(hit)}/{len(rules)} ({pct}%) "
|
||||
f"over {len(paths)} scenarios")
|
||||
print(" NOTE: counts tags; does not prove a scenario exercises what it names")
|
||||
if missing:
|
||||
print(" uncovered:", " ".join(missing))
|
||||
if invented:
|
||||
print(" ERROR — claimed but not defined in the spec:", " ".join(invented),
|
||||
file=sys.stderr)
|
||||
return 1
|
||||
return 0 if not missing else 2
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
sys.exit(main())
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue