clay-borg/specs/InnerLoop.md

343 lines
18 KiB
Markdown
Raw Normal View History

# The Inner Loop — Assimilate and Surpass
T10: InnerLoop v1.2 — hardening does not converge, so optimize correction The retrospective question was whether the mechanism set is complete or each pass still finds a new class. This pass produced both an eighth instance AND a fourth class, so the answer is the uncomfortable one. Ledger: 10 instances, 4 classes, across 3 workplans. Every pass has produced at least one class the previous pass had not seen. HDN harness-does-nothing 5 executable assertions TA trusted arithmetic 3 re-derivation SSB same-sample blind spot 1 assertions over ALL the data DFD duplicated-fact drift 2 NEW -- reading a copy against source DFD is genuinely distinct: no positive control catches it, because both copies are internally consistent, and re-derivation does not either, because the copy faithfully reproduces what it was copied from. Found when an inlined price sheet went stale within an hour of T11 changing the real one. So v1.2 stops trying to enumerate classes in advance. Every error in three passes was corrected in-session for under ~1% of the pass, so the stated design goal is now cheap CORRECTION: keep raw data so numbers are re-derivable, keep artifacts small and committed so a wrong number is one grep from everywhere quoting it, give every number a command. Plus the one rule the new class earns: single source of fact. The original hypothesis is revised rather than confirmed. "A rule that cannot be executed is not a rule" is wrong -- the two most valuable corrections in the project came from a decorative rule that cannot be automated (re-derive inherited numbers). An executable rule fires reliably and catches one class; a decorative one fires unreliably and can catch any class, including unnamed ones. Keep both. Gates this pass: loop-lint caught 3 real violations on first run, then failed on its own author within the hour when a T07 edit pushed InnerLoop.md to 407 lines against its own 400 limit. CB-WP-0003 complete: 11 of 11 tasks done. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-31 09:31:19 +02:00
Status: **v1.2** — corrected from CB-WP-0003 (loop hardening) on
2026-07-31. Changes from v1.1: single source of fact; review targets the
harness and states its sampling limit; correction vs retarget; the chaos
roll's calibration window; the live cost budget. The design goal is now
stated: **optimize for cheap correction, not for exhaustive prevention.**
Rationale: `history/260731-loop-hardening-retrospective.md`.
v1.1 — corrected from CB-WP-0002 (cost accounting) on
T07: InnerLoop v1.1 — the instrument must emit its own target Answers the question the workplan posed: did v1.0's "every metric names its instrument" rule stop a metric being written without a working instrument? No. AC-1 named `cb-cost --pin` before that tool existed and set a hand-computed target of $92.87; the tool returned $93.32. The rule was satisfied completely and the metric was still wrong. v1.1 adds: - the instrument must exist and the target must come out of it; targets are provisional until the tool emits them - a number inherited from earlier work is re-derived before use as a target, or cited as unverified - every reporting tool exposes --self-test, run before the number - cost is in the definition of done; M-D2-CST may not be uncomputable The cost of CB-WP-0001 was stated four times before it was right -- $248.46, $92.21, $92.87, $93.32 -- and each correction came from a different mechanism: re-derivation, adversarial review, and the positive control. None found more than one. That is the case for keeping all three. CB-WP-0003 T10 predicted the next error would be harness-does-nothing. It was not, twice. Trusted arithmetic over real data would pass a positive control; and a property verified on 206/206 groups of the main transcript is false in the 8-response subagent tree that neither the survey nor the reviewer examined separately. Review structurally cannot catch the second -- re-deriving on the same sample reproduces the same blind spot. Workplan status: done. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-31 08:53:52 +02:00
2026-07-31. Changes from v1.0: the instrument must exist and emit its own
target; inherited numbers are re-derived before use; every reporting tool
exposes `--self-test`; cost is in the definition of done. Rationale:
`history/260731-cost-accounting-retrospective.md`.
v1.0 — survived its first full pass (CB-WP-0001, the GROUND game kernel)
and was corrected from it on 2026-07-31. Changes from v0.2: measurement
validity (the positive control), metric feasibility and instrument naming,
four implementation rules the pass earned, and the requirement that
evidence state what it does not support. Rationale and the failures behind
each: `history/260731-inner-loop-retrospective.md`.
Normative process for building every Clay-Borg capability. Referenced by
T10: InnerLoop v1.2 — hardening does not converge, so optimize correction The retrospective question was whether the mechanism set is complete or each pass still finds a new class. This pass produced both an eighth instance AND a fourth class, so the answer is the uncomfortable one. Ledger: 10 instances, 4 classes, across 3 workplans. Every pass has produced at least one class the previous pass had not seen. HDN harness-does-nothing 5 executable assertions TA trusted arithmetic 3 re-derivation SSB same-sample blind spot 1 assertions over ALL the data DFD duplicated-fact drift 2 NEW -- reading a copy against source DFD is genuinely distinct: no positive control catches it, because both copies are internally consistent, and re-derivation does not either, because the copy faithfully reproduces what it was copied from. Found when an inlined price sheet went stale within an hour of T11 changing the real one. So v1.2 stops trying to enumerate classes in advance. Every error in three passes was corrected in-session for under ~1% of the pass, so the stated design goal is now cheap CORRECTION: keep raw data so numbers are re-derivable, keep artifacts small and committed so a wrong number is one grep from everywhere quoting it, give every number a command. Plus the one rule the new class earns: single source of fact. The original hypothesis is revised rather than confirmed. "A rule that cannot be executed is not a rule" is wrong -- the two most valuable corrections in the project came from a decorative rule that cannot be automated (re-derive inherited numbers). An executable rule fires reliably and catches one class; a decorative one fires unreliably and can catch any class, including unnamed ones. Keep both. Gates this pass: loop-lint caught 3 real violations on first run, then failed on its own author within the hour when a T07 edit pushed InnerLoop.md to 407 lines against its own 400 limit. CB-WP-0003 complete: 11 of 11 tasks done. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-31 09:31:19 +02:00
all workplans.
**Design goal (v1.2).** Three passes produced ten error instances across
four classes, and every pass produced a class the previous one had not
seen — prevention is not converging. Every one of those errors was
corrected inside the same session for under ~1% of the pass. The loop
therefore optimizes for **cheap correction**: keep the raw data so numbers
can be re-derived, keep artifacts small and committed so a wrong number is
one grep from everywhere that quotes it, and give every reported number a
command so re-running is free.
> **Single source of fact.** A number, rate, or target lives in exactly
> one place; everywhere else links to it. Where a copy is unavoidable it
> is generated by a command, not typed. *(v1.2: a price sheet inlined into
> a spec went stale within the hour of the real sheet changing, and one
> acceptance figure had to be chased across three artifacts each time it
> moved. No positive control catches this — both copies are internally
> consistent — and re-derivation does not either, because the copy
> reproduces whatever it was copied from.)* The loop's own optimization target is **agentic efficiency**:
every artifact it produces must be small enough to load whole, structured
enough to act on without interpretation, and falsifiable enough that an
agent can judge its own work without a human in the iteration.
---
## The five steps
```text
1 RESEARCH → research/CB-RES-NNNN-<slug>.md (survey, baselines)
2 APPROVE → decision recorded in the ADR (gate: survey complete?)
3 DECIDE → decisions/ADR-NNNN-<slug>.md (assimilate/reimplement/hybrid)
4 SPECIFY → specs/<Capability>.md (contracts + acceptance metrics)
5 CODE LOOP → code + scenarios + benchmarks (iterate until metrics beat baseline)
```
**Hard gate: no implementation code for a capability exists before its ADR
(step 3) is committed.**
## Loop tiers and the chaos roll
Every work packet declares a tier before work starts. The tier sets how
heavy steps 13 are; steps 45 (spec with metrics, code loop with evidence)
are never skipped for code-producing work.
| Tier | Weight of steps 13 | Structural trigger (forces at least this tier) |
|---|---|---|
| **L** | Full: separate survey, adversarial review, ADR | Creates a new capability port, or is named a high-leverage pass by the maintainer |
| **M** | Survey and ADR merged into one document; review optional | Touches a canonical interface, or adds/updates an external dependency |
| **S** | One provenance paragraph in the commit message | Everything else (utilities, fixes, refactors inside a boundary) |
**The chaos roll.** After deriving the structural tier, roll **d4**
(`shuf -i 1-4 -n 1`). On a **4**, the tier is instead picked uniformly at
random (`shuf -e S M L -n 1`), overriding the structural derivation — up or
down.
> **Calibration window, opened 2026-07-31 (CB-WP-0003 T06).** The rate was
> d10 and the mechanism **never fired**: two rolls across two workplans
> (CB-WP-0001: 9, CB-WP-0002: 2), against ~0.2 expected firings. At d10 and
> ~2 tier decisions per workplan it would take roughly twenty workplans to
> observe four overrides, so the mechanism was set at a rate that prevented
> its own evaluation — the one option T06 ruled out.
>
> Raised to **d4 (25%) for the next 12 tier declarations**, then evaluated
> and either kept, returned to d10, or deleted. Expected ~3 firings in the
> window, which is enough to see whether an overridden tier produces a
> different outcome than the argued one.
>
> **Stated cost:** a chaos-L override on work that would have been S buys a
> full survey, adversarial review, and ADR. Measured comparable: CB-WP-0001
> T03 (a tier-L survey) cost **$9.91**. At 25% over 12 declarations the
> window is expected to cost **$2030**. That is the price of finding out
> whether the mechanism is worth keeping, and it is cheaper than carrying an
> unevaluated ritual indefinitely. Both rolls are recorded in the tier declaration
(`tier: M (structural L, chaos 10→M)`). **Record the roll every time,
including when it changes nothing** (`tier: L (structural L, chaos 4)`),
so a mechanism that never fires is visible rather than assumed. Purpose:
an occasional random
reweighting keeps the classification honest — arguing everything into S
stops paying off when audits can compare argued tiers against the random
sample — and occasionally forces a deep look at something "obviously
trivial", which is where local optima hide.
Chaos limits: a rolled-down tier relaxes *process* weight only. Invariants
(zero foreign types in canonical interfaces, determinism, passing
conformance suites) bind at every tier, and a rolled-down pass touching a
canonical interface still requires the interface change to be flagged in
the commit for retrospective review.
### Step 1 — Research
Identify the best implementation in existence for this capability. Produce
`research/CB-RES-NNNN-<slug>.md` following the survey template (below).
The survey is done when it can name, per dimension, a concrete
**benchmark-to-beat**: a number, a property, or a reproducible comparison —
not an impression.
**Runnable-baseline option.** For passes judged high-leverage (declared by
the maintainer or proposed in the survey and confirmed in the ADR), cited
numbers are not enough: the survey must ship a reproducible **baseline
harness** that runs the leading candidate on our machine against our
workload — the same scenario files where feasible. The harness ships with a
*fidelity note* stating what was and wasn't faithfully reproduced, so a
hastily wired competitor setup cannot silently inflate our advantage.
Where the option is not invoked (or the candidate isn't practically
runnable), comparisons against cited-only numbers are **directional**: the
evidence verdict for those rows caps at `parity`, never `better`.
### Step 2 — Approve (adversarial review)
For tier-L passes, approval is earned through an **adversarial review**: a
separate session (or agent) attempts to break the work. Exactly **one
round**: challenge, then response. The work is approvable only when every
challenge is either answered with evidence or conceded and folded in.
**The review target follows the risk.** Reviewing the survey was the
original rule, and it is the wrong target for a capability whose claim
rests on numbers — every serious error in this project has been in
measurement or build configuration, not in prose.
| the claim rests on | the reviewer is given | and must attempt |
|---|---|---|
| a survey of candidates | the survey | an omitted candidate; a stale or unverifiable benchmark; an unmeasured claim presented as measured |
| **numbers** | the survey **and the harness and the evidence file** | **reproduce the number independently**; state what the harness would report if the work silently stopped |
**What review cannot do — state this, do not discover it.** A reviewer
re-derives the author's claims and therefore inherits the author's
sampling. So:
> **The reviewer re-derives on a different sample than the author used.**
> Where only one sample exists, the review says so rather than reporting a
> clean verify.
*(v1.1, from CB-WP-0002: the dedup invariant was verified on the main
transcript by the survey — 206/206 groups — and independently re-verified
by the reviewer, who used the same transcript. It is false in the
8-response `subagents/` tree neither examined. Two independent checks,
one blind spot, because both sampled the same way. Only an assertion
running over all the data at execution time caught it.)*
**What review demonstrably does do.** Measured across two passes:
**$0.66** and **$1.11**, roughly 1% of each pass, each finding
approval-blocking defects — in CB-WP-0002 a target that would have made
the evidence file certify a broken collector. The step pays for itself by
a wide margin and the cost is not a reason to skip it.
**Documentation requirement:** the research process, the challenge, and the
resulting improvements to the research are each documented in timestamped
markdown files under `history/`:
```text
history/YYMMDD-<slug>-research.md # how the survey was conducted: sources,
# queries, what was measured vs cited, dead ends
history/YYMMDD-<slug>-challenge.md # the adversarial attack, verbatim
history/YYMMDD-<slug>-response.md # answers/concessions and what changed in the survey
```
The polished survey artifact remains `research/CB-RES-NNNN-<slug>.md`; the
history files preserve the unpolished trail so a later reader can judge how
hard the survey was actually tested. For tier-M passes the review is
optional but, when performed, follows the same format. If not approvable
after the round, the loop returns to step 1 with the named gaps.
### Step 3 — Decide
`decisions/ADR-NNNN-<slug>.md`: assimilate behind a port, reimplement, or
hybrid — with the **expected advantage stated per dimension** (see rubric).
An honest "worse here, better there, and why that trade is right" beats a
claimed sweep of all four dimensions.
### Step 4 — Specify
`specs/<Capability>.md`: the contracts, invariants, and — mandatory — the
**acceptance metrics table**, each row tied to a baseline from step 1.
A spec without measurable acceptance criteria is not done. Metrics follow
the conventions in [MetricsAndScenarios.md](MetricsAndScenarios.md),
including the rule that metric selection itself passes through a mini
research step (metric provenance).
**Every metric names its instrument, and is checked reachable.** A row
in the acceptance table carries the command that produces its number.
T07: InnerLoop v1.1 — the instrument must emit its own target Answers the question the workplan posed: did v1.0's "every metric names its instrument" rule stop a metric being written without a working instrument? No. AC-1 named `cb-cost --pin` before that tool existed and set a hand-computed target of $92.87; the tool returned $93.32. The rule was satisfied completely and the metric was still wrong. v1.1 adds: - the instrument must exist and the target must come out of it; targets are provisional until the tool emits them - a number inherited from earlier work is re-derived before use as a target, or cited as unverified - every reporting tool exposes --self-test, run before the number - cost is in the definition of done; M-D2-CST may not be uncomputable The cost of CB-WP-0001 was stated four times before it was right -- $248.46, $92.21, $92.87, $93.32 -- and each correction came from a different mechanism: re-derivation, adversarial review, and the positive control. None found more than one. That is the case for keeping all three. CB-WP-0003 T10 predicted the next error would be harness-does-nothing. It was not, twice. Trusted arithmetic over real data would pass a positive control; and a property verified on 206/206 groups of the main transcript is false in the 8-response subagent tree that neither the survey nor the reviewer examined separately. Review structurally cannot catch the second -- re-deriving on the same sample reproduces the same blind spot. Workplan status: done. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-31 08:53:52 +02:00
A metric with no named instrument is a wish, not a metric.
**The instrument must exist, and the target must come out of it.**
Naming a command is not the same as running one. A target computed by
hand and merely *labelled* with a command is the same defect the rule
was written to stop, one level down. Where the instrument is built later
in the pass, the target is marked `provisional:` until the instrument
emits it, and the spec is amended to whatever the instrument returns.
*(v1.1, from CB-WP-0002: `specs/CostAccounting.md` AC-1 named
`cb-cost --pin fc76445` before that tool existed, and set the target to
a hand-computed $92.87. When the tool was built it returned $93.32 —
the hand computation carried a dedup bug the tool's own positive control
caught. The metric satisfied v1.0's rule completely and was still
wrong.)*
**A number inherited from earlier work is re-derived before it is used
as a target, or it is cited as unverified.** Quoting is not measuring.
*(v1.1, from CB-WP-0002: the workplan opened with $248.46, inherited
from a prior pass. Re-derivation put it at $92.21 — the quoted figure
double-counted transcript lines and priced a three-model session at one
model's rate. Neither error was of the harness-does-nothing class; both
T01: audit every InnerLoop rule, and make the checkable ones executable 41 rules classified executable / checkable / decorative, each tagged with the failure class it catches. Counts: 11 executable, 22 checkable, 4 decorative (one of them dead policy). Audit: history/260731-inner-loop-rule-audit.md New tools/loop-lint.py makes 7 rules executable (tier declared, chaos roll recorded, tier-L review trail, unmeasured-in-evidence, whole-file loadability, reporting tools expose --self-test). It found three real violations on its first run, none previously visible: - specs/ArchitectureBlueprint.md was 543 lines against a ~400 limit the loop has stated since v0.2 and never measured. Split at its own section boundaries into Blueprint (1-8) + Runtime (9-15). - tools/dep-weight.py and tools/rule-coverage.py had positive-control logic and no --self-test, so nothing verified the control worked. Adding rule-coverage's self-test exposed a latent instance of the exact class this workplan is about: if the spec regex stopped matching, rules was empty, missing was empty, and the tool exited 0 reporting "0/0" -- a silent pass, in the tool that reports our headline AM-1 number. Both tools now assert they found something before reporting. Two demotions applied in the spec rather than left implicit: "structured over prose" is marked guidance (nothing can check it), and the 8k/10k token budget is struck through and marked DEAD POLICY pointing at T05. The audit's uncomfortable finding: rule 13 (re-derive inherited numbers) has no mechanical form, is deliberately left decorative, and caught the LARGEST error in CB-WP-0002. That is a counter-example to this workplan's own hypothesis. "A rule that cannot be executed is not a rule" is wrong as stated; the defensible version is that such a rule cannot be relied on to fire, so it must not be the only defence for a class that matters. Class coverage: harness-does-nothing has five executable rules; trusted-arithmetic has ZERO and produced the largest single error. make loop-lint and make self-tests wired into `make all` and CI. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-31 09:16:00 +02:00
sums ran over real data, and a positive control would have passed them.)*
T07: separate a correction from a retarget with a mechanical test A blanket "no retargeting in the measuring commit" rule would have been wrong. CB-WP-0002 moved AC-1 three times in exactly that shape and every move was correct -- each time a new instrument disproved the old figure. Four legitimate corrections would have been forbidden to catch one bad retarget. The test is mechanical rather than a statement of intent: correction the target moves and the implementation does not; legal in the same commit provided the instrument's output is there retarget the same commit changes both the target and the code the target measures; requires an ADR stating why the new target binds on future work Applied retroactively: AM-4a/AM-4b are UNRATIFIED. They were set after seeing the measurement, in the commit that produced it, with the implementation changing too -- a retarget by this test. make dep-weight is currently enforcing a target no reviewed decision stands behind. Recorded as an open item; ratifying or changing them is a maintainer decision, not an implementer's. Also: specs/InnerLoop.md split into InnerLoop.md (process) and InnerLoopReference.md (rubric, template, rules, definition of done). Not a stylistic choice -- `make loop-lint` failed on the commit that pushed the file to 407 lines against its own ~400 limit. The gate added this morning to make that rule executable caught its own author within the hour, which is the cheapest possible demonstration that it works. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-31 09:25:22 +02:00
**Retargeting: the instrument may move a target, the implementation may
not.** A metric's target changes in only two ways, and they are not
treated alike:
| | trigger | requirement |
|---|---|---|
| **corrected** | the instrument disproved the target — the target was computed by hand, or by an earlier tool with a defect | legitimate in the same commit, **provided the instrument's output is in that commit** |
| **retargeted** | the implementation missed the target and the target moves to accommodate it | requires an **ADR**: old target, the measurement, and why the new target binds on *future* work rather than merely passing present work |
The distinction is not the implementer's self-report of intent. It is
mechanical: **a correction is one where the target moves and the
implementation does not.** If the same commit changes both the target and
the code the target measures, it is a retarget and needs the ADR.
*(v1.1, from CB-WP-0002/0003: AM-4's targets were measured at 246,250 and
set at 250,000 in one commit by the implementer after seeing the number —
the structure this rule exists to stop. But CB-WP-0002 then moved AC-1
three times, correctly, each time because a new instrument disproved the
old figure ($92.21 → $92.87 → $93.32 → $93.15). A blanket prohibition
would have forbidden four legitimate corrections to catch one bad
retarget.)*
**Applied retroactively:** AM-4a and AM-4b are **unratified** until an ADR
is written or they are changed. They were set by the implementer after
seeing the measurement, in the commit that produced it, and the
implementation changed in that same commit — a retarget by the test above.
Tracked as an open item in `history/260731-inner-loop-rule-audit.md`.
T01: audit every InnerLoop rule, and make the checkable ones executable 41 rules classified executable / checkable / decorative, each tagged with the failure class it catches. Counts: 11 executable, 22 checkable, 4 decorative (one of them dead policy). Audit: history/260731-inner-loop-rule-audit.md New tools/loop-lint.py makes 7 rules executable (tier declared, chaos roll recorded, tier-L review trail, unmeasured-in-evidence, whole-file loadability, reporting tools expose --self-test). It found three real violations on its first run, none previously visible: - specs/ArchitectureBlueprint.md was 543 lines against a ~400 limit the loop has stated since v0.2 and never measured. Split at its own section boundaries into Blueprint (1-8) + Runtime (9-15). - tools/dep-weight.py and tools/rule-coverage.py had positive-control logic and no --self-test, so nothing verified the control worked. Adding rule-coverage's self-test exposed a latent instance of the exact class this workplan is about: if the spec regex stopped matching, rules was empty, missing was empty, and the tool exited 0 reporting "0/0" -- a silent pass, in the tool that reports our headline AM-1 number. Both tools now assert they found something before reporting. Two demotions applied in the spec rather than left implicit: "structured over prose" is marked guidance (nothing can check it), and the 8k/10k token budget is struck through and marked DEAD POLICY pointing at T05. The audit's uncomfortable finding: rule 13 (re-derive inherited numbers) has no mechanical form, is deliberately left decorative, and caught the LARGEST error in CB-WP-0002. That is a counter-example to this workplan's own hypothesis. "A rule that cannot be executed is not a rule" is wrong as stated; the defensible version is that such a rule cannot be relied on to fire, so it must not be the only defence for a class that matters. Class coverage: harness-does-nothing has five executable rules; trusted-arithmetic has ZERO and produced the largest single error. make loop-lint and make self-tests wired into `make all` and CI. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-31 09:16:00 +02:00
**A metric is checked against the contracts in its own spec.** If a
contract makes a target unreachable, one of the two is wrong and the
conflict is resolved when it is noticed, not at the acceptance run.
Re-check the table whenever a contract is added.
*(v1.0, from CB-WP-0001: AM-4's ≤20-crate target was made unreachable by
the K5 and K7 contracts written after it, and AM-12's cost metric was
fully specified and never instrumented, so it could not be computed.)*
### Step 5 — Code loop
Implement iteratively. Each iteration:
```text
change → cb-check (fmt, clippy, tests) → scenarios → benchmarks
→ compare against acceptance table → evidence row appended
```
Done when every acceptance metric meets or beats its baseline and the
comparison numbers are committed as an evidence file
(`evidence/CB-EV-NNNN-<slug>.md`). A failed scenario must yield a replay
artifact an agent can re-execute locally.
#### Measurement validity — the positive control
**Every benchmark and harness must assert that it performed the work it
reports.** Completing without error is not evidence of having done
anything: a loop whose commands are all rejected runs fast and reports a
throughput for work that never happened.
Concretely, a measurement harness must, on every run:
- assert the unit of work produced its expected effect (events applied,
rows written, moves accepted) — not merely that the call returned;
- fail loudly rather than report a number when that assertion fails;
- state the divisor used to convert raw timings into the metric's unit,
pinned by a test so a workload change cannot silently rescale it.
**A number from a run that cannot prove it did the work is void** and
must not reach an evidence file.
**Every tool that reports a number exposes `--self-test`**, and that
self-test runs before the number is produced (`make cost` depends on
`make cost-test`). The assertion must name a failure it detects, not
merely exercise the happy path.
*(v1.0+, from CB-WP-0002: `cb-cost`'s dedup assertion fired on its first
run against real data and aborted, catching a rule that was verified on
206/206 groups of the main transcript and false in the 8-response
subagent tree. The generalization that failed — a property confirmed on
the largest sample assumed to hold on the smallest — is not one review
catches, because both the survey and the adversarial reviewer checked
the same large sample.)*
*(v1.0, from CB-WP-0001: both serious errors in the first pass were of
exactly this shape. A JS harness reported 8.4s for 100k moves while
every move was being rejected, and a Rust benchmark reported 9.3M
events/s — a 93× beat — while most rounds never completed because a
stress gate rejected one player's action. The corrected figure was 5.6×
lower. Adversarial review caught neither; both were claims about
numbers, and review reads prose.)*
#### Evidence states what it does not support
An evidence file that compares across runtimes, languages, or feature
sets **names the disanalogies explicitly**, in the same section as the
number. The reader must not have to infer that a ratio is not
like-for-like. This is the parity-cap rule applied to the write-up:
state the claim you will defend, and the claim you are not making.
---
---
T07: separate a correction from a retarget with a mechanical test A blanket "no retargeting in the measuring commit" rule would have been wrong. CB-WP-0002 moved AC-1 three times in exactly that shape and every move was correct -- each time a new instrument disproved the old figure. Four legitimate corrections would have been forbidden to catch one bad retarget. The test is mechanical rather than a statement of intent: correction the target moves and the implementation does not; legal in the same commit provided the instrument's output is there retarget the same commit changes both the target and the code the target measures; requires an ADR stating why the new target binds on future work Applied retroactively: AM-4a/AM-4b are UNRATIFIED. They were set after seeing the measurement, in the commit that produced it, with the implementation changing too -- a retarget by this test. make dep-weight is currently enforcing a target no reviewed decision stands behind. Recorded as an open item; ratifying or changing them is a maintainer decision, not an implementer's. Also: specs/InnerLoop.md split into InnerLoop.md (process) and InnerLoopReference.md (rubric, template, rules, definition of done). Not a stylistic choice -- `make loop-lint` failed on the commit that pushed the file to 407 lines against its own ~400 limit. The gate added this morning to make that rule executable caught its own author within the hour, which is the cheapest possible demonstration that it works. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-31 09:25:22 +02:00
## Reference material
T07: separate a correction from a retarget with a mechanical test A blanket "no retargeting in the measuring commit" rule would have been wrong. CB-WP-0002 moved AC-1 three times in exactly that shape and every move was correct -- each time a new instrument disproved the old figure. Four legitimate corrections would have been forbidden to catch one bad retarget. The test is mechanical rather than a statement of intent: correction the target moves and the implementation does not; legal in the same commit provided the instrument's output is there retarget the same commit changes both the target and the code the target measures; requires an ADR stating why the new target binds on future work Applied retroactively: AM-4a/AM-4b are UNRATIFIED. They were set after seeing the measurement, in the commit that produced it, with the implementation changing too -- a retarget by this test. make dep-weight is currently enforcing a target no reviewed decision stands behind. Recorded as an open item; ratifying or changing them is a maintainer decision, not an implementer's. Also: specs/InnerLoop.md split into InnerLoop.md (process) and InnerLoopReference.md (rubric, template, rules, definition of done). Not a stylistic choice -- `make loop-lint` failed on the commit that pushed the file to 407 lines against its own ~400 limit. The gate added this morning to make that rule executable caught its own author within the hour, which is the cheapest possible demonstration that it works. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-31 09:25:22 +02:00
The four-dimension rubric, the survey template, the implementation rules
each pass earned, the agentic-efficiency requirements, and the
definition of done live in
**[InnerLoopReference.md](InnerLoopReference.md)**. They are normative;
they are separated only so both files load whole.
T07: separate a correction from a retarget with a mechanical test A blanket "no retargeting in the measuring commit" rule would have been wrong. CB-WP-0002 moved AC-1 three times in exactly that shape and every move was correct -- each time a new instrument disproved the old figure. Four legitimate corrections would have been forbidden to catch one bad retarget. The test is mechanical rather than a statement of intent: correction the target moves and the implementation does not; legal in the same commit provided the instrument's output is there retarget the same commit changes both the target and the code the target measures; requires an ADR stating why the new target binds on future work Applied retroactively: AM-4a/AM-4b are UNRATIFIED. They were set after seeing the measurement, in the commit that produced it, with the implementation changing too -- a retarget by this test. make dep-weight is currently enforcing a target no reviewed decision stands behind. Recorded as an open item; ratifying or changing them is a maintainer decision, not an implementer's. Also: specs/InnerLoop.md split into InnerLoop.md (process) and InnerLoopReference.md (rubric, template, rules, definition of done). Not a stylistic choice -- `make loop-lint` failed on the commit that pushed the file to 407 lines against its own ~400 limit. The gate added this morning to make that rule executable caught its own author within the hour, which is the cheapest possible demonstration that it works. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-31 09:25:22 +02:00
*(Split 2026-07-31: this file reached 407 lines against its own ~400-line
loadability rule, and `make loop-lint` — added the same day to make that
rule executable — failed on the commit that pushed it over. The rule
caught its own author within an hour of being written.)*