149 lines
9.1 KiB
Markdown
149 lines
9.1 KiB
Markdown
|
|
# 2026-07-31 — InnerLoop v1.1 rule enforceability audit
|
|||
|
|
|
|||
|
|
CB-WP-0003 T01. Every rule in `specs/InnerLoop.md` v1.1 classified as
|
|||
|
|
**executable** (a command fails when it is violated), **checkable** (a
|
|||
|
|
human or agent can verify it cheaply and objectively in review), or
|
|||
|
|
**decorative** (neither) — and, per the revised task, tagged with the
|
|||
|
|
**failure class it catches**.
|
|||
|
|
|
|||
|
|
## Failure classes on record
|
|||
|
|
|
|||
|
|
Seven error instances across three classes, from
|
|||
|
|
`history/260731-inner-loop-retrospective.md` and
|
|||
|
|
`history/260731-cost-accounting-retrospective.md`:
|
|||
|
|
|
|||
|
|
| tag | class | instances |
|
|||
|
|
|---|---|---|
|
|||
|
|
| **HDN** | harness-does-nothing — a run succeeds while performing no work | 4 |
|
|||
|
|
| **TA** | trusted arithmetic — a correct-looking sum over data that really exists | 2 |
|
|||
|
|
| **SSB** | same-sample blind spot — a property verified on the large sample, assumed on the small | 1 |
|
|||
|
|
|
|||
|
|
The reason for tagging: a rule that catches a class no observed error
|
|||
|
|
belongs to is a candidate for deletion even when it is perfectly
|
|||
|
|
executable, and a class with no rule covering it is where the next rule
|
|||
|
|
should go.
|
|||
|
|
|
|||
|
|
## The audit
|
|||
|
|
|
|||
|
|
Status after this pass. `loop-lint` = `tools/loop-lint.py`, new here.
|
|||
|
|
|
|||
|
|
| # | Rule | Class | Status | Enforced by |
|
|||
|
|
|---|---|---|---|---|
|
|||
|
|
| 1 | No implementation code before the ADR is committed | — | **checkable** | git history vs ADR date; not automated (see §Deferred) |
|
|||
|
|
| 2 | Tier declared before work starts | — | **executable** | `loop-lint` tier-declared |
|
|||
|
|
| 3 | Chaos roll recorded every time, even when it changes nothing | — | **executable** | `loop-lint` chaos-recorded |
|
|||
|
|
| 4 | Invariants bind at every tier regardless of roll | — | executable (elsewhere) | `make all` — determinism tests, `disallowed_types` |
|
|||
|
|
| 5 | Survey names a benchmark-to-beat per dimension | — | **checkable** | review; a regex cannot judge whether a number is a benchmark |
|
|||
|
|
| 6 | Cited-only baselines cap the verdict at `parity` | TA | **decorative → checkable** | see §Changes |
|
|||
|
|
| 7 | Tier-L survey gets one round of adversarial review | SSB | **executable** | `loop-lint` review-trail |
|
|||
|
|
| 8 | Review trail: research/challenge/response in `history/` | SSB | **executable** | `loop-lint` review-trail |
|
|||
|
|
| 9 | ADR states expected advantage per dimension | — | **checkable** | review |
|
|||
|
|
| 10 | Spec carries an acceptance-metrics table | — | **checkable** | review |
|
|||
|
|
| 11 | Every metric names its instrument | HDN | **checkable** | review; naming is textual, existence is not |
|
|||
|
|
| 12 | **v1.1** The instrument must exist and emit its own target | TA | **checkable** | review; see §Deferred for why not yet executable |
|
|||
|
|
| 13 | **v1.1** Inherited numbers are re-derived before use as a target | TA | **decorative** | nothing can detect a quoted number |
|
|||
|
|
| 14 | A metric is checked against contracts in its own spec | — | **checkable** | review |
|
|||
|
|
| 15 | Harness asserts it performed the work it reports | HDN | **executable** | `cargo bench -- --test`, cb-sim empty-run, `--self-test` |
|
|||
|
|
| 16 | Fail loudly rather than report when the assertion fails | HDN | **executable** | same |
|
|||
|
|
| 17 | Divisor pinned by a test | HDN | **executable** | `bench_shape` test module |
|
|||
|
|
| 18 | A number that cannot prove its work is void | HDN | **executable** | same as 15 |
|
|||
|
|
| 19 | **v1.1** Every reporting tool exposes `--self-test` | HDN, SSB | **executable** | `loop-lint` self-test |
|
|||
|
|
| 20 | `--self-test` names a failure it detects, not the happy path | HDN | **checkable** | review; see §Deferred |
|
|||
|
|
| 21 | Evidence names its disanalogies | — | **checkable** | review |
|
|||
|
|
| 22 | `unmeasured` illegal in an evidence file | — | **executable** | `loop-lint` evidence-unmeasured |
|
|||
|
|
| 23 | No silently-ignored input | HDN | **checkable** | review |
|
|||
|
|
| 24 | Decisions get commands, not defaults | — | **checkable** | review |
|
|||
|
|
| 25 | Scaffolds are exercised or marked | HDN | **checkable** | review |
|
|||
|
|
| 26 | Coverage gates that count tags say so | — | **checkable** | review |
|
|||
|
|
| 27 | Whole-file loadability (~400 lines) | — | **executable** | `loop-lint` loadability |
|
|||
|
|
| 28 | Structured over prose | — | **decorative** | kept as guidance, marked |
|
|||
|
|
| 29 | One command surface | — | **checkable** | review |
|
|||
|
|
| 30 | Self-contained tasks | — | **checkable** | review |
|
|||
|
|
| 31 | Evidence or it didn't happen | — | **checkable** | review |
|
|||
|
|
| 32 | Token discipline per the global budget policy | — | **decorative — dead** | nothing; CB-WP-0003 T05 replaces it |
|
|||
|
|
| 33–41 | Definition-of-done checklist (9 items) | mixed | **checkable** | review; each maps to an artifact whose existence is testable — see §Deferred |
|
|||
|
|
|
|||
|
|
Counts: **11 executable**, **22 checkable**, **4 decorative** (one of
|
|||
|
|
which is dead policy).
|
|||
|
|
|
|||
|
|
## What the audit found by running
|
|||
|
|
|
|||
|
|
`tools/loop-lint.py` was written to make rules 2, 3, 7, 8, 19, 22 and 27
|
|||
|
|
executable. On its first run it produced **three findings, all real, none
|
|||
|
|
previously visible**:
|
|||
|
|
|
|||
|
|
1. **`specs/ArchitectureBlueprint.md` is 543 lines** against a ~400-line
|
|||
|
|
limit the loop has stated since v0.2. Nobody noticed because nothing
|
|||
|
|
measured it. Disposition: split (§Changes).
|
|||
|
|
2. **`tools/dep-weight.py` has no `--self-test`.** It has positive-control
|
|||
|
|
logic — it refuses to report when a crate cannot be located — but
|
|||
|
|
nothing verifies that control still works.
|
|||
|
|
3. **`tools/rule-coverage.py` has no `--self-test`.** Same shape.
|
|||
|
|
|
|||
|
|
Findings 2 and 3 are exactly the recursion this workplan is about: the
|
|||
|
|
positive-control rule was applied to benchmarks and to the newest tool,
|
|||
|
|
and not to the two older tools that report AM-1 and AM-4 numbers into
|
|||
|
|
evidence files.
|
|||
|
|
|
|||
|
|
## Changes made in this task
|
|||
|
|
|
|||
|
|
- **Rule 27 made executable and the violation fixed.**
|
|||
|
|
`ArchitectureBlueprint.md` split at its own section boundaries into
|
|||
|
|
`ArchitectureBlueprint.md` (§1–8, the stack) and
|
|||
|
|
`ArchitectureRuntime.md` (§9–15, runtime/tooling/process), linked both
|
|||
|
|
ways.
|
|||
|
|
- **Rules 2, 3, 7, 8, 19, 22 made executable** via `tools/loop-lint.py`,
|
|||
|
|
wired as `make loop-lint` and into `make all` and CI.
|
|||
|
|
- **Rule 6 (parity cap) demoted from decorative to checkable** by stating
|
|||
|
|
the check explicitly: an evidence row citing a baseline whose provenance
|
|||
|
|
is `cited` may not carry verdict `better`. That is mechanical against the
|
|||
|
|
survey's provenance column and is queued for `loop-lint` once a second
|
|||
|
|
evidence file exists to test it against — writing a matcher with one
|
|||
|
|
sample is the SSB error this pass is trying to stop making.
|
|||
|
|
- **Rule 28 (structured over prose) demoted to guidance**, with that
|
|||
|
|
status stated in the spec. It is a style preference; nothing can judge
|
|||
|
|
it, and leaving it phrased as a requirement is the false assurance this
|
|||
|
|
audit exists to remove.
|
|||
|
|
- **Rule 32 (token discipline) marked dead** in place, pointing at
|
|||
|
|
CB-WP-0003 T05. It is not deleted yet because deleting it is T05's
|
|||
|
|
decision, but it now reads as dead rather than as a live control.
|
|||
|
|
|
|||
|
|
## Deferred, with reasons
|
|||
|
|
|
|||
|
|
- **Rule 1 (ADR gate) is checkable but not automated.** A mechanical check
|
|||
|
|
needs a map from capability → source paths, which does not exist. Cheap
|
|||
|
|
version worth doing later: require every `specs/<X>.md` to reference an
|
|||
|
|
ADR, and every ADR to precede the first commit touching its capability's
|
|||
|
|
directory.
|
|||
|
|
- **Rule 12 (instrument emits its target) resists automation for now.**
|
|||
|
|
Detecting "this number was typed rather than emitted" requires the tool's
|
|||
|
|
output to be committed alongside the spec. The tractable form is to
|
|||
|
|
require acceptance targets to appear verbatim in a committed tool output
|
|||
|
|
file; deferred to T02 rather than guessed at here.
|
|||
|
|
- **Rule 20 (self-test names a failure) is genuinely checkable only.** A
|
|||
|
|
self-test that asserts `True` passes any structural check. Reviewing the
|
|||
|
|
assertions is the control, and CB-WP-0002's AC-9 is the model: pin the
|
|||
|
|
exact defect that occurred.
|
|||
|
|
- **Rule 13 (re-derive inherited numbers) has no mechanical form at all**
|
|||
|
|
and is left decorative *deliberately* — with its status stated. It is
|
|||
|
|
the rule that caught the largest error in CB-WP-0002 ($248.46 → $92.21),
|
|||
|
|
which is the counter-example to this workplan's own hypothesis: **an
|
|||
|
|
unenforceable rule was the most valuable one in the pass.** The
|
|||
|
|
hypothesis "a rule that cannot be executed is not a rule" is therefore
|
|||
|
|
wrong as stated. The correct version is narrower: *a rule that cannot be
|
|||
|
|
executed cannot be relied on to fire, so it must not be the only defence
|
|||
|
|
for a class that matters.*
|
|||
|
|
|
|||
|
|
## Class coverage — where the gaps are
|
|||
|
|
|
|||
|
|
| class | executable rules covering it | assessment |
|
|||
|
|
|---|---|---|
|
|||
|
|
| HDN | 15, 16, 17, 18, 19 | **well covered.** Five executable rules; the class that started this. |
|
|||
|
|
| TA | none | **uncovered by any executable rule.** Rules 6, 12, 13 are checkable or decorative. This is the largest gap, and it is the class with the second-most instances. |
|
|||
|
|
| SSB | 7, 8, 19 | **partly covered.** Rule 19 caught the one instance, by accident of running over all data rather than by design. Rules 7/8 enforce that a review *happened*, not that it sampled differently. T03 addresses the design gap. |
|
|||
|
|
|
|||
|
|
The honest read: **the loop is hardened against the class it has already
|
|||
|
|
suffered most from, and has no executable defence against the class that
|
|||
|
|
produced its largest single error.** Trusted arithmetic is caught today
|
|||
|
|
only by re-derivation, which is a discipline, not a gate.
|