clay-borg/history/260731-inner-loop-rule-audit.md
tegwick 0c1eb9ecba ADR-0004: ratify AM-4a and AM-4b (maintainer decision)
Discharges the open item T07 raised. Values unchanged at 250,000 and
350,000; what was missing was a reviewed decision behind them, since
they had been set by the implementer in the commit that measured them
and that also changed the feature gating being measured.

The ADR supplies the argument T07's test requires -- why the targets
bind on FUTURE work rather than merely passing present work:

  AM-4a leaves 3,750 lines of headroom (1.5%), about one small crate,
  so any new shipped-runtime dependency breaches it almost immediately.
  That is intended: the shipped runtime should be effectively frozen.

  AM-4b leaves 32,979 lines (10.4%), deliberately looser -- dev tooling
  should absorb one moderate dependency without a spec change, not two.

  Both are ceilings on a quantity that only grows by choice. Nothing
  drifts across them; only adding a dependency does.

Falsification condition stated: if a later pass raises AM-4a to
accommodate a dependency it wants, that is the failure the ceiling
exists to catch, and the answer is an ADR arguing for the dependency.

First ADR written under the correction/retarget test; sets the shape.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-31 09:45:35 +02:00

9.5 KiB
Raw Blame History

2026-07-31 — InnerLoop v1.1 rule enforceability audit

CB-WP-0003 T01. Every rule in specs/InnerLoop.md v1.1 classified as executable (a command fails when it is violated), checkable (a human or agent can verify it cheaply and objectively in review), or decorative (neither) — and, per the revised task, tagged with the failure class it catches.

Failure classes on record

Seven error instances across three classes, from history/260731-inner-loop-retrospective.md and history/260731-cost-accounting-retrospective.md:

tag class instances
HDN harness-does-nothing — a run succeeds while performing no work 4
TA trusted arithmetic — a correct-looking sum over data that really exists 2
SSB same-sample blind spot — a property verified on the large sample, assumed on the small 1

The reason for tagging: a rule that catches a class no observed error belongs to is a candidate for deletion even when it is perfectly executable, and a class with no rule covering it is where the next rule should go.

The audit

Status after this pass. loop-lint = tools/loop-lint.py, new here.

# Rule Class Status Enforced by
1 No implementation code before the ADR is committed checkable git history vs ADR date; not automated (see §Deferred)
2 Tier declared before work starts executable loop-lint tier-declared
3 Chaos roll recorded every time, even when it changes nothing executable loop-lint chaos-recorded
4 Invariants bind at every tier regardless of roll executable (elsewhere) make all — determinism tests, disallowed_types
5 Survey names a benchmark-to-beat per dimension checkable review; a regex cannot judge whether a number is a benchmark
6 Cited-only baselines cap the verdict at parity TA decorative → checkable see §Changes
7 Tier-L survey gets one round of adversarial review SSB executable loop-lint review-trail
8 Review trail: research/challenge/response in history/ SSB executable loop-lint review-trail
9 ADR states expected advantage per dimension checkable review
10 Spec carries an acceptance-metrics table checkable review
11 Every metric names its instrument HDN checkable review; naming is textual, existence is not
12 v1.1 The instrument must exist and emit its own target TA checkable review; see §Deferred for why not yet executable
13 v1.1 Inherited numbers are re-derived before use as a target TA decorative nothing can detect a quoted number
14 A metric is checked against contracts in its own spec checkable review
15 Harness asserts it performed the work it reports HDN executable cargo bench -- --test, cb-sim empty-run, --self-test
16 Fail loudly rather than report when the assertion fails HDN executable same
17 Divisor pinned by a test HDN executable bench_shape test module
18 A number that cannot prove its work is void HDN executable same as 15
19 v1.1 Every reporting tool exposes --self-test HDN, SSB executable loop-lint self-test
20 --self-test names a failure it detects, not the happy path HDN checkable review; see §Deferred
21 Evidence names its disanalogies checkable review
22 unmeasured illegal in an evidence file executable loop-lint evidence-unmeasured
23 No silently-ignored input HDN checkable review
24 Decisions get commands, not defaults checkable review
25 Scaffolds are exercised or marked HDN checkable review
26 Coverage gates that count tags say so checkable review
27 Whole-file loadability (~400 lines) executable loop-lint loadability
28 Structured over prose decorative kept as guidance, marked
29 One command surface checkable review
30 Self-contained tasks checkable review
31 Evidence or it didn't happen checkable review
32 Token discipline per the global budget policy decorative — dead nothing; CB-WP-0003 T05 replaces it
3341 Definition-of-done checklist (9 items) mixed checkable review; each maps to an artifact whose existence is testable — see §Deferred

Counts: 11 executable, 22 checkable, 4 decorative (one of which is dead policy).

What the audit found by running

tools/loop-lint.py was written to make rules 2, 3, 7, 8, 19, 22 and 27 executable. On its first run it produced three findings, all real, none previously visible:

  1. specs/ArchitectureBlueprint.md is 543 lines against a ~400-line limit the loop has stated since v0.2. Nobody noticed because nothing measured it. Disposition: split (§Changes).
  2. tools/dep-weight.py has no --self-test. It has positive-control logic — it refuses to report when a crate cannot be located — but nothing verifies that control still works.
  3. tools/rule-coverage.py has no --self-test. Same shape.

Findings 2 and 3 are exactly the recursion this workplan is about: the positive-control rule was applied to benchmarks and to the newest tool, and not to the two older tools that report AM-1 and AM-4 numbers into evidence files.

Changes made in this task

  • Rule 27 made executable and the violation fixed. ArchitectureBlueprint.md split at its own section boundaries into ArchitectureBlueprint.md (§18, the stack) and ArchitectureRuntime.md (§915, runtime/tooling/process), linked both ways.
  • Rules 2, 3, 7, 8, 19, 22 made executable via tools/loop-lint.py, wired as make loop-lint and into make all and CI.
  • Rule 6 (parity cap) demoted from decorative to checkable by stating the check explicitly: an evidence row citing a baseline whose provenance is cited may not carry verdict better. That is mechanical against the survey's provenance column and is queued for loop-lint once a second evidence file exists to test it against — writing a matcher with one sample is the SSB error this pass is trying to stop making.
  • Rule 28 (structured over prose) demoted to guidance, with that status stated in the spec. It is a style preference; nothing can judge it, and leaving it phrased as a requirement is the false assurance this audit exists to remove.
  • Rule 32 (token discipline) marked dead in place, pointing at CB-WP-0003 T05. It is not deleted yet because deleting it is T05's decision, but it now reads as dead rather than as a live control.

Deferred, with reasons

  • Rule 1 (ADR gate) is checkable but not automated. A mechanical check needs a map from capability → source paths, which does not exist. Cheap version worth doing later: require every specs/<X>.md to reference an ADR, and every ADR to precede the first commit touching its capability's directory.
  • Rule 12 (instrument emits its target) resists automation for now. Detecting "this number was typed rather than emitted" requires the tool's output to be committed alongside the spec. The tractable form is to require acceptance targets to appear verbatim in a committed tool output file; deferred to T02 rather than guessed at here.
  • Rule 20 (self-test names a failure) is genuinely checkable only. A self-test that asserts True passes any structural check. Reviewing the assertions is the control, and CB-WP-0002's AC-9 is the model: pin the exact defect that occurred.
  • Rule 13 (re-derive inherited numbers) has no mechanical form at all and is left decorative deliberately — with its status stated. It is the rule that caught the largest error in CB-WP-0002 ($248.46 → $92.21), which is the counter-example to this workplan's own hypothesis: an unenforceable rule was the most valuable one in the pass. The hypothesis "a rule that cannot be executed is not a rule" is therefore wrong as stated. The correct version is narrower: a rule that cannot be executed cannot be relied on to fire, so it must not be the only defence for a class that matters.

Open items raised by this audit

  • AM-4a / AM-4b are unratified retargets. Discharged 2026-07-31 by ADR-0004, ratified by the maintainer. Values unchanged; the ADR supplies the future-binding argument the test requires (AM-4a leaves 1.5% headroom, AM-4b 10.4%, and both measure a quantity that only grows by deliberate choice).

Class coverage — where the gaps are

class executable rules covering it assessment
HDN 15, 16, 17, 18, 19 well covered. Five executable rules; the class that started this.
TA none uncovered by any executable rule. Rules 6, 12, 13 are checkable or decorative. This is the largest gap, and it is the class with the second-most instances.
SSB 7, 8, 19 partly covered. Rule 19 caught the one instance, by accident of running over all data rather than by design. Rules 7/8 enforce that a review happened, not that it sampled differently. T03 addresses the design gap.

The honest read: the loop is hardened against the class it has already suffered most from, and has no executable defence against the class that produced its largest single error. Trusted arithmetic is caught today only by re-derivation, which is a discipline, not a gate.