41 rules classified executable / checkable / decorative, each tagged with
the failure class it catches. Counts: 11 executable, 22 checkable, 4
decorative (one of them dead policy).
Audit: history/260731-inner-loop-rule-audit.md
New tools/loop-lint.py makes 7 rules executable (tier declared, chaos
roll recorded, tier-L review trail, unmeasured-in-evidence, whole-file
loadability, reporting tools expose --self-test). It found three real
violations on its first run, none previously visible:
- specs/ArchitectureBlueprint.md was 543 lines against a ~400 limit
the loop has stated since v0.2 and never measured. Split at its own
section boundaries into Blueprint (1-8) + Runtime (9-15).
- tools/dep-weight.py and tools/rule-coverage.py had positive-control
logic and no --self-test, so nothing verified the control worked.
Adding rule-coverage's self-test exposed a latent instance of the exact
class this workplan is about: if the spec regex stopped matching, rules
was empty, missing was empty, and the tool exited 0 reporting "0/0" --
a silent pass, in the tool that reports our headline AM-1 number. Both
tools now assert they found something before reporting.
Two demotions applied in the spec rather than left implicit: "structured
over prose" is marked guidance (nothing can check it), and the 8k/10k
token budget is struck through and marked DEAD POLICY pointing at T05.
The audit's uncomfortable finding: rule 13 (re-derive inherited numbers)
has no mechanical form, is deliberately left decorative, and caught the
LARGEST error in CB-WP-0002. That is a counter-example to this
workplan's own hypothesis. "A rule that cannot be executed is not a
rule" is wrong as stated; the defensible version is that such a rule
cannot be relied on to fire, so it must not be the only defence for a
class that matters.
Class coverage: harness-does-nothing has five executable rules;
trusted-arithmetic has ZERO and produced the largest single error.
make loop-lint and make self-tests wired into `make all` and CI.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
148 lines
9.1 KiB
Markdown
148 lines
9.1 KiB
Markdown
# 2026-07-31 — InnerLoop v1.1 rule enforceability audit
|
||
|
||
CB-WP-0003 T01. Every rule in `specs/InnerLoop.md` v1.1 classified as
|
||
**executable** (a command fails when it is violated), **checkable** (a
|
||
human or agent can verify it cheaply and objectively in review), or
|
||
**decorative** (neither) — and, per the revised task, tagged with the
|
||
**failure class it catches**.
|
||
|
||
## Failure classes on record
|
||
|
||
Seven error instances across three classes, from
|
||
`history/260731-inner-loop-retrospective.md` and
|
||
`history/260731-cost-accounting-retrospective.md`:
|
||
|
||
| tag | class | instances |
|
||
|---|---|---|
|
||
| **HDN** | harness-does-nothing — a run succeeds while performing no work | 4 |
|
||
| **TA** | trusted arithmetic — a correct-looking sum over data that really exists | 2 |
|
||
| **SSB** | same-sample blind spot — a property verified on the large sample, assumed on the small | 1 |
|
||
|
||
The reason for tagging: a rule that catches a class no observed error
|
||
belongs to is a candidate for deletion even when it is perfectly
|
||
executable, and a class with no rule covering it is where the next rule
|
||
should go.
|
||
|
||
## The audit
|
||
|
||
Status after this pass. `loop-lint` = `tools/loop-lint.py`, new here.
|
||
|
||
| # | Rule | Class | Status | Enforced by |
|
||
|---|---|---|---|---|
|
||
| 1 | No implementation code before the ADR is committed | — | **checkable** | git history vs ADR date; not automated (see §Deferred) |
|
||
| 2 | Tier declared before work starts | — | **executable** | `loop-lint` tier-declared |
|
||
| 3 | Chaos roll recorded every time, even when it changes nothing | — | **executable** | `loop-lint` chaos-recorded |
|
||
| 4 | Invariants bind at every tier regardless of roll | — | executable (elsewhere) | `make all` — determinism tests, `disallowed_types` |
|
||
| 5 | Survey names a benchmark-to-beat per dimension | — | **checkable** | review; a regex cannot judge whether a number is a benchmark |
|
||
| 6 | Cited-only baselines cap the verdict at `parity` | TA | **decorative → checkable** | see §Changes |
|
||
| 7 | Tier-L survey gets one round of adversarial review | SSB | **executable** | `loop-lint` review-trail |
|
||
| 8 | Review trail: research/challenge/response in `history/` | SSB | **executable** | `loop-lint` review-trail |
|
||
| 9 | ADR states expected advantage per dimension | — | **checkable** | review |
|
||
| 10 | Spec carries an acceptance-metrics table | — | **checkable** | review |
|
||
| 11 | Every metric names its instrument | HDN | **checkable** | review; naming is textual, existence is not |
|
||
| 12 | **v1.1** The instrument must exist and emit its own target | TA | **checkable** | review; see §Deferred for why not yet executable |
|
||
| 13 | **v1.1** Inherited numbers are re-derived before use as a target | TA | **decorative** | nothing can detect a quoted number |
|
||
| 14 | A metric is checked against contracts in its own spec | — | **checkable** | review |
|
||
| 15 | Harness asserts it performed the work it reports | HDN | **executable** | `cargo bench -- --test`, cb-sim empty-run, `--self-test` |
|
||
| 16 | Fail loudly rather than report when the assertion fails | HDN | **executable** | same |
|
||
| 17 | Divisor pinned by a test | HDN | **executable** | `bench_shape` test module |
|
||
| 18 | A number that cannot prove its work is void | HDN | **executable** | same as 15 |
|
||
| 19 | **v1.1** Every reporting tool exposes `--self-test` | HDN, SSB | **executable** | `loop-lint` self-test |
|
||
| 20 | `--self-test` names a failure it detects, not the happy path | HDN | **checkable** | review; see §Deferred |
|
||
| 21 | Evidence names its disanalogies | — | **checkable** | review |
|
||
| 22 | `unmeasured` illegal in an evidence file | — | **executable** | `loop-lint` evidence-unmeasured |
|
||
| 23 | No silently-ignored input | HDN | **checkable** | review |
|
||
| 24 | Decisions get commands, not defaults | — | **checkable** | review |
|
||
| 25 | Scaffolds are exercised or marked | HDN | **checkable** | review |
|
||
| 26 | Coverage gates that count tags say so | — | **checkable** | review |
|
||
| 27 | Whole-file loadability (~400 lines) | — | **executable** | `loop-lint` loadability |
|
||
| 28 | Structured over prose | — | **decorative** | kept as guidance, marked |
|
||
| 29 | One command surface | — | **checkable** | review |
|
||
| 30 | Self-contained tasks | — | **checkable** | review |
|
||
| 31 | Evidence or it didn't happen | — | **checkable** | review |
|
||
| 32 | Token discipline per the global budget policy | — | **decorative — dead** | nothing; CB-WP-0003 T05 replaces it |
|
||
| 33–41 | Definition-of-done checklist (9 items) | mixed | **checkable** | review; each maps to an artifact whose existence is testable — see §Deferred |
|
||
|
||
Counts: **11 executable**, **22 checkable**, **4 decorative** (one of
|
||
which is dead policy).
|
||
|
||
## What the audit found by running
|
||
|
||
`tools/loop-lint.py` was written to make rules 2, 3, 7, 8, 19, 22 and 27
|
||
executable. On its first run it produced **three findings, all real, none
|
||
previously visible**:
|
||
|
||
1. **`specs/ArchitectureBlueprint.md` is 543 lines** against a ~400-line
|
||
limit the loop has stated since v0.2. Nobody noticed because nothing
|
||
measured it. Disposition: split (§Changes).
|
||
2. **`tools/dep-weight.py` has no `--self-test`.** It has positive-control
|
||
logic — it refuses to report when a crate cannot be located — but
|
||
nothing verifies that control still works.
|
||
3. **`tools/rule-coverage.py` has no `--self-test`.** Same shape.
|
||
|
||
Findings 2 and 3 are exactly the recursion this workplan is about: the
|
||
positive-control rule was applied to benchmarks and to the newest tool,
|
||
and not to the two older tools that report AM-1 and AM-4 numbers into
|
||
evidence files.
|
||
|
||
## Changes made in this task
|
||
|
||
- **Rule 27 made executable and the violation fixed.**
|
||
`ArchitectureBlueprint.md` split at its own section boundaries into
|
||
`ArchitectureBlueprint.md` (§1–8, the stack) and
|
||
`ArchitectureRuntime.md` (§9–15, runtime/tooling/process), linked both
|
||
ways.
|
||
- **Rules 2, 3, 7, 8, 19, 22 made executable** via `tools/loop-lint.py`,
|
||
wired as `make loop-lint` and into `make all` and CI.
|
||
- **Rule 6 (parity cap) demoted from decorative to checkable** by stating
|
||
the check explicitly: an evidence row citing a baseline whose provenance
|
||
is `cited` may not carry verdict `better`. That is mechanical against the
|
||
survey's provenance column and is queued for `loop-lint` once a second
|
||
evidence file exists to test it against — writing a matcher with one
|
||
sample is the SSB error this pass is trying to stop making.
|
||
- **Rule 28 (structured over prose) demoted to guidance**, with that
|
||
status stated in the spec. It is a style preference; nothing can judge
|
||
it, and leaving it phrased as a requirement is the false assurance this
|
||
audit exists to remove.
|
||
- **Rule 32 (token discipline) marked dead** in place, pointing at
|
||
CB-WP-0003 T05. It is not deleted yet because deleting it is T05's
|
||
decision, but it now reads as dead rather than as a live control.
|
||
|
||
## Deferred, with reasons
|
||
|
||
- **Rule 1 (ADR gate) is checkable but not automated.** A mechanical check
|
||
needs a map from capability → source paths, which does not exist. Cheap
|
||
version worth doing later: require every `specs/<X>.md` to reference an
|
||
ADR, and every ADR to precede the first commit touching its capability's
|
||
directory.
|
||
- **Rule 12 (instrument emits its target) resists automation for now.**
|
||
Detecting "this number was typed rather than emitted" requires the tool's
|
||
output to be committed alongside the spec. The tractable form is to
|
||
require acceptance targets to appear verbatim in a committed tool output
|
||
file; deferred to T02 rather than guessed at here.
|
||
- **Rule 20 (self-test names a failure) is genuinely checkable only.** A
|
||
self-test that asserts `True` passes any structural check. Reviewing the
|
||
assertions is the control, and CB-WP-0002's AC-9 is the model: pin the
|
||
exact defect that occurred.
|
||
- **Rule 13 (re-derive inherited numbers) has no mechanical form at all**
|
||
and is left decorative *deliberately* — with its status stated. It is
|
||
the rule that caught the largest error in CB-WP-0002 ($248.46 → $92.21),
|
||
which is the counter-example to this workplan's own hypothesis: **an
|
||
unenforceable rule was the most valuable one in the pass.** The
|
||
hypothesis "a rule that cannot be executed is not a rule" is therefore
|
||
wrong as stated. The correct version is narrower: *a rule that cannot be
|
||
executed cannot be relied on to fire, so it must not be the only defence
|
||
for a class that matters.*
|
||
|
||
## Class coverage — where the gaps are
|
||
|
||
| class | executable rules covering it | assessment |
|
||
|---|---|---|
|
||
| HDN | 15, 16, 17, 18, 19 | **well covered.** Five executable rules; the class that started this. |
|
||
| TA | none | **uncovered by any executable rule.** Rules 6, 12, 13 are checkable or decorative. This is the largest gap, and it is the class with the second-most instances. |
|
||
| SSB | 7, 8, 19 | **partly covered.** Rule 19 caught the one instance, by accident of running over all data rather than by design. Rules 7/8 enforce that a review *happened*, not that it sampled differently. T03 addresses the design gap. |
|
||
|
||
The honest read: **the loop is hardened against the class it has already
|
||
suffered most from, and has no executable defence against the class that
|
||
produced its largest single error.** Trusted arithmetic is caught today
|
||
only by re-derivation, which is a discipline, not a gate.
|