clay-borg/history/260731-inner-loop-rule-audit.md
tegwick fed422a3a3 T01: audit every InnerLoop rule, and make the checkable ones executable
41 rules classified executable / checkable / decorative, each tagged with
the failure class it catches. Counts: 11 executable, 22 checkable, 4
decorative (one of them dead policy).
Audit: history/260731-inner-loop-rule-audit.md

New tools/loop-lint.py makes 7 rules executable (tier declared, chaos
roll recorded, tier-L review trail, unmeasured-in-evidence, whole-file
loadability, reporting tools expose --self-test). It found three real
violations on its first run, none previously visible:

  - specs/ArchitectureBlueprint.md was 543 lines against a ~400 limit
    the loop has stated since v0.2 and never measured. Split at its own
    section boundaries into Blueprint (1-8) + Runtime (9-15).
  - tools/dep-weight.py and tools/rule-coverage.py had positive-control
    logic and no --self-test, so nothing verified the control worked.

Adding rule-coverage's self-test exposed a latent instance of the exact
class this workplan is about: if the spec regex stopped matching, rules
was empty, missing was empty, and the tool exited 0 reporting "0/0" --
a silent pass, in the tool that reports our headline AM-1 number. Both
tools now assert they found something before reporting.

Two demotions applied in the spec rather than left implicit: "structured
over prose" is marked guidance (nothing can check it), and the 8k/10k
token budget is struck through and marked DEAD POLICY pointing at T05.

The audit's uncomfortable finding: rule 13 (re-derive inherited numbers)
has no mechanical form, is deliberately left decorative, and caught the
LARGEST error in CB-WP-0002. That is a counter-example to this
workplan's own hypothesis. "A rule that cannot be executed is not a
rule" is wrong as stated; the defensible version is that such a rule
cannot be relied on to fire, so it must not be the only defence for a
class that matters.

Class coverage: harness-does-nothing has five executable rules;
trusted-arithmetic has ZERO and produced the largest single error.

make loop-lint and make self-tests wired into `make all` and CI.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-31 09:16:00 +02:00

148 lines
9.1 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# 2026-07-31 — InnerLoop v1.1 rule enforceability audit
CB-WP-0003 T01. Every rule in `specs/InnerLoop.md` v1.1 classified as
**executable** (a command fails when it is violated), **checkable** (a
human or agent can verify it cheaply and objectively in review), or
**decorative** (neither) — and, per the revised task, tagged with the
**failure class it catches**.
## Failure classes on record
Seven error instances across three classes, from
`history/260731-inner-loop-retrospective.md` and
`history/260731-cost-accounting-retrospective.md`:
| tag | class | instances |
|---|---|---|
| **HDN** | harness-does-nothing — a run succeeds while performing no work | 4 |
| **TA** | trusted arithmetic — a correct-looking sum over data that really exists | 2 |
| **SSB** | same-sample blind spot — a property verified on the large sample, assumed on the small | 1 |
The reason for tagging: a rule that catches a class no observed error
belongs to is a candidate for deletion even when it is perfectly
executable, and a class with no rule covering it is where the next rule
should go.
## The audit
Status after this pass. `loop-lint` = `tools/loop-lint.py`, new here.
| # | Rule | Class | Status | Enforced by |
|---|---|---|---|---|
| 1 | No implementation code before the ADR is committed | — | **checkable** | git history vs ADR date; not automated (see §Deferred) |
| 2 | Tier declared before work starts | — | **executable** | `loop-lint` tier-declared |
| 3 | Chaos roll recorded every time, even when it changes nothing | — | **executable** | `loop-lint` chaos-recorded |
| 4 | Invariants bind at every tier regardless of roll | — | executable (elsewhere) | `make all` — determinism tests, `disallowed_types` |
| 5 | Survey names a benchmark-to-beat per dimension | — | **checkable** | review; a regex cannot judge whether a number is a benchmark |
| 6 | Cited-only baselines cap the verdict at `parity` | TA | **decorative → checkable** | see §Changes |
| 7 | Tier-L survey gets one round of adversarial review | SSB | **executable** | `loop-lint` review-trail |
| 8 | Review trail: research/challenge/response in `history/` | SSB | **executable** | `loop-lint` review-trail |
| 9 | ADR states expected advantage per dimension | — | **checkable** | review |
| 10 | Spec carries an acceptance-metrics table | — | **checkable** | review |
| 11 | Every metric names its instrument | HDN | **checkable** | review; naming is textual, existence is not |
| 12 | **v1.1** The instrument must exist and emit its own target | TA | **checkable** | review; see §Deferred for why not yet executable |
| 13 | **v1.1** Inherited numbers are re-derived before use as a target | TA | **decorative** | nothing can detect a quoted number |
| 14 | A metric is checked against contracts in its own spec | — | **checkable** | review |
| 15 | Harness asserts it performed the work it reports | HDN | **executable** | `cargo bench -- --test`, cb-sim empty-run, `--self-test` |
| 16 | Fail loudly rather than report when the assertion fails | HDN | **executable** | same |
| 17 | Divisor pinned by a test | HDN | **executable** | `bench_shape` test module |
| 18 | A number that cannot prove its work is void | HDN | **executable** | same as 15 |
| 19 | **v1.1** Every reporting tool exposes `--self-test` | HDN, SSB | **executable** | `loop-lint` self-test |
| 20 | `--self-test` names a failure it detects, not the happy path | HDN | **checkable** | review; see §Deferred |
| 21 | Evidence names its disanalogies | — | **checkable** | review |
| 22 | `unmeasured` illegal in an evidence file | — | **executable** | `loop-lint` evidence-unmeasured |
| 23 | No silently-ignored input | HDN | **checkable** | review |
| 24 | Decisions get commands, not defaults | — | **checkable** | review |
| 25 | Scaffolds are exercised or marked | HDN | **checkable** | review |
| 26 | Coverage gates that count tags say so | — | **checkable** | review |
| 27 | Whole-file loadability (~400 lines) | — | **executable** | `loop-lint` loadability |
| 28 | Structured over prose | — | **decorative** | kept as guidance, marked |
| 29 | One command surface | — | **checkable** | review |
| 30 | Self-contained tasks | — | **checkable** | review |
| 31 | Evidence or it didn't happen | — | **checkable** | review |
| 32 | Token discipline per the global budget policy | — | **decorative — dead** | nothing; CB-WP-0003 T05 replaces it |
| 3341 | Definition-of-done checklist (9 items) | mixed | **checkable** | review; each maps to an artifact whose existence is testable — see §Deferred |
Counts: **11 executable**, **22 checkable**, **4 decorative** (one of
which is dead policy).
## What the audit found by running
`tools/loop-lint.py` was written to make rules 2, 3, 7, 8, 19, 22 and 27
executable. On its first run it produced **three findings, all real, none
previously visible**:
1. **`specs/ArchitectureBlueprint.md` is 543 lines** against a ~400-line
limit the loop has stated since v0.2. Nobody noticed because nothing
measured it. Disposition: split (§Changes).
2. **`tools/dep-weight.py` has no `--self-test`.** It has positive-control
logic — it refuses to report when a crate cannot be located — but
nothing verifies that control still works.
3. **`tools/rule-coverage.py` has no `--self-test`.** Same shape.
Findings 2 and 3 are exactly the recursion this workplan is about: the
positive-control rule was applied to benchmarks and to the newest tool,
and not to the two older tools that report AM-1 and AM-4 numbers into
evidence files.
## Changes made in this task
- **Rule 27 made executable and the violation fixed.**
`ArchitectureBlueprint.md` split at its own section boundaries into
`ArchitectureBlueprint.md` (§18, the stack) and
`ArchitectureRuntime.md` (§915, runtime/tooling/process), linked both
ways.
- **Rules 2, 3, 7, 8, 19, 22 made executable** via `tools/loop-lint.py`,
wired as `make loop-lint` and into `make all` and CI.
- **Rule 6 (parity cap) demoted from decorative to checkable** by stating
the check explicitly: an evidence row citing a baseline whose provenance
is `cited` may not carry verdict `better`. That is mechanical against the
survey's provenance column and is queued for `loop-lint` once a second
evidence file exists to test it against — writing a matcher with one
sample is the SSB error this pass is trying to stop making.
- **Rule 28 (structured over prose) demoted to guidance**, with that
status stated in the spec. It is a style preference; nothing can judge
it, and leaving it phrased as a requirement is the false assurance this
audit exists to remove.
- **Rule 32 (token discipline) marked dead** in place, pointing at
CB-WP-0003 T05. It is not deleted yet because deleting it is T05's
decision, but it now reads as dead rather than as a live control.
## Deferred, with reasons
- **Rule 1 (ADR gate) is checkable but not automated.** A mechanical check
needs a map from capability → source paths, which does not exist. Cheap
version worth doing later: require every `specs/<X>.md` to reference an
ADR, and every ADR to precede the first commit touching its capability's
directory.
- **Rule 12 (instrument emits its target) resists automation for now.**
Detecting "this number was typed rather than emitted" requires the tool's
output to be committed alongside the spec. The tractable form is to
require acceptance targets to appear verbatim in a committed tool output
file; deferred to T02 rather than guessed at here.
- **Rule 20 (self-test names a failure) is genuinely checkable only.** A
self-test that asserts `True` passes any structural check. Reviewing the
assertions is the control, and CB-WP-0002's AC-9 is the model: pin the
exact defect that occurred.
- **Rule 13 (re-derive inherited numbers) has no mechanical form at all**
and is left decorative *deliberately* — with its status stated. It is
the rule that caught the largest error in CB-WP-0002 ($248.46 → $92.21),
which is the counter-example to this workplan's own hypothesis: **an
unenforceable rule was the most valuable one in the pass.** The
hypothesis "a rule that cannot be executed is not a rule" is therefore
wrong as stated. The correct version is narrower: *a rule that cannot be
executed cannot be relied on to fire, so it must not be the only defence
for a class that matters.*
## Class coverage — where the gaps are
| class | executable rules covering it | assessment |
|---|---|---|
| HDN | 15, 16, 17, 18, 19 | **well covered.** Five executable rules; the class that started this. |
| TA | none | **uncovered by any executable rule.** Rules 6, 12, 13 are checkable or decorative. This is the largest gap, and it is the class with the second-most instances. |
| SSB | 7, 8, 19 | **partly covered.** Rule 19 caught the one instance, by accident of running over all data rather than by design. Rules 7/8 enforce that a review *happened*, not that it sampled differently. T03 addresses the design gap. |
The honest read: **the loop is hardened against the class it has already
suffered most from, and has no executable defence against the class that
produced its largest single error.** Trusted arithmetic is caught today
only by re-derivation, which is a discipline, not a gate.