clay-borg/specs/InnerLoop.md
tegwick 4fd6322e17 T03: point review at the harness, and state what review cannot catch
The review step targeted the survey. Every serious error in this project
has been in measurement or build configuration, so the loop was
adversarially reviewing the artifact cheapest to fix and leaving
unreviewed the one where errors occur.

Step 2 now routes by risk: when the claim rests on numbers, the reviewer
gets the harness and the evidence file too, and must reproduce the
number independently rather than read about it.

The addition that matters more, because it was learned the hard way: a
reviewer re-derives the author's claims and therefore inherits the
author's SAMPLING. CB-WP-0002's dedup invariant was checked twice --
survey 206/206 groups, then the reviewer independently -- and both used
the main transcript. It is false in the 8-response subagent tree neither
looked at. Two independent verifications, one shared blind spot.

  Rule: the reviewer re-derives on a different sample than the author
  used, and where only one sample exists, says so rather than reporting
  a clean verify.

Also recorded: what review demonstrably DOES do. $0.66 and $1.11 across
two passes, ~1% of each, both finding approval-blocking defects. Cost is
not a reason to skip it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-07-31 09:26:01 +02:00

16 KiB
Raw Blame History

The Inner Loop — Assimilate and Surpass

Status: v1.1 — corrected from CB-WP-0002 (cost accounting) on 2026-07-31. Changes from v1.0: the instrument must exist and emit its own target; inherited numbers are re-derived before use; every reporting tool exposes --self-test; cost is in the definition of done. Rationale: history/260731-cost-accounting-retrospective.md.

v1.0 — survived its first full pass (CB-WP-0001, the GROUND game kernel) and was corrected from it on 2026-07-31. Changes from v0.2: measurement validity (the positive control), metric feasibility and instrument naming, four implementation rules the pass earned, and the requirement that evidence state what it does not support. Rationale and the failures behind each: history/260731-inner-loop-retrospective.md.

Normative process for building every Clay-Borg capability. Referenced by all workplans. The loop's own optimization target is agentic efficiency: every artifact it produces must be small enough to load whole, structured enough to act on without interpretation, and falsifiable enough that an agent can judge its own work without a human in the iteration.


The five steps

1 RESEARCH  → research/CB-RES-NNNN-<slug>.md     (survey, baselines)
2 APPROVE   → decision recorded in the ADR        (gate: survey complete?)
3 DECIDE    → decisions/ADR-NNNN-<slug>.md        (assimilate/reimplement/hybrid)
4 SPECIFY   → specs/<Capability>.md               (contracts + acceptance metrics)
5 CODE LOOP → code + scenarios + benchmarks       (iterate until metrics beat baseline)

Hard gate: no implementation code for a capability exists before its ADR (step 3) is committed.

Loop tiers and the chaos roll

Every work packet declares a tier before work starts. The tier sets how heavy steps 13 are; steps 45 (spec with metrics, code loop with evidence) are never skipped for code-producing work.

Tier Weight of steps 13 Structural trigger (forces at least this tier)
L Full: separate survey, adversarial review, ADR Creates a new capability port, or is named a high-leverage pass by the maintainer
M Survey and ADR merged into one document; review optional Touches a canonical interface, or adds/updates an external dependency
S One provenance paragraph in the commit message Everything else (utilities, fixes, refactors inside a boundary)

The chaos roll. After deriving the structural tier, roll d4 (shuf -i 1-4 -n 1). On a 4, the tier is instead picked uniformly at random (shuf -e S M L -n 1), overriding the structural derivation — up or down.

Calibration window, opened 2026-07-31 (CB-WP-0003 T06). The rate was d10 and the mechanism never fired: two rolls across two workplans (CB-WP-0001: 9, CB-WP-0002: 2), against ~0.2 expected firings. At d10 and ~2 tier decisions per workplan it would take roughly twenty workplans to observe four overrides, so the mechanism was set at a rate that prevented its own evaluation — the one option T06 ruled out.

Raised to d4 (25%) for the next 12 tier declarations, then evaluated and either kept, returned to d10, or deleted. Expected ~3 firings in the window, which is enough to see whether an overridden tier produces a different outcome than the argued one.

Stated cost: a chaos-L override on work that would have been S buys a full survey, adversarial review, and ADR. Measured comparable: CB-WP-0001 T03 (a tier-L survey) cost $9.91. At 25% over 12 declarations the window is expected to cost $2030. That is the price of finding out whether the mechanism is worth keeping, and it is cheaper than carrying an unevaluated ritual indefinitely. Both rolls are recorded in the tier declaration (tier: M (structural L, chaos 10→M)). Record the roll every time, including when it changes nothing (tier: L (structural L, chaos 4)), so a mechanism that never fires is visible rather than assumed. Purpose: an occasional random reweighting keeps the classification honest — arguing everything into S stops paying off when audits can compare argued tiers against the random sample — and occasionally forces a deep look at something "obviously trivial", which is where local optima hide.

Chaos limits: a rolled-down tier relaxes process weight only. Invariants (zero foreign types in canonical interfaces, determinism, passing conformance suites) bind at every tier, and a rolled-down pass touching a canonical interface still requires the interface change to be flagged in the commit for retrospective review.

Step 1 — Research

Identify the best implementation in existence for this capability. Produce research/CB-RES-NNNN-<slug>.md following the survey template (below). The survey is done when it can name, per dimension, a concrete benchmark-to-beat: a number, a property, or a reproducible comparison — not an impression.

Runnable-baseline option. For passes judged high-leverage (declared by the maintainer or proposed in the survey and confirmed in the ADR), cited numbers are not enough: the survey must ship a reproducible baseline harness that runs the leading candidate on our machine against our workload — the same scenario files where feasible. The harness ships with a fidelity note stating what was and wasn't faithfully reproduced, so a hastily wired competitor setup cannot silently inflate our advantage. Where the option is not invoked (or the candidate isn't practically runnable), comparisons against cited-only numbers are directional: the evidence verdict for those rows caps at parity, never better.

Step 2 — Approve (adversarial review)

For tier-L passes, approval is earned through an adversarial review: a separate session (or agent) attempts to break the work. Exactly one round: challenge, then response. The work is approvable only when every challenge is either answered with evidence or conceded and folded in.

The review target follows the risk. Reviewing the survey was the original rule, and it is the wrong target for a capability whose claim rests on numbers — every serious error in this project has been in measurement or build configuration, not in prose.

the claim rests on the reviewer is given and must attempt
a survey of candidates the survey an omitted candidate; a stale or unverifiable benchmark; an unmeasured claim presented as measured
numbers the survey and the harness and the evidence file reproduce the number independently; state what the harness would report if the work silently stopped

What review cannot do — state this, do not discover it. A reviewer re-derives the author's claims and therefore inherits the author's sampling. So:

The reviewer re-derives on a different sample than the author used. Where only one sample exists, the review says so rather than reporting a clean verify.

(v1.1, from CB-WP-0002: the dedup invariant was verified on the main transcript by the survey — 206/206 groups — and independently re-verified by the reviewer, who used the same transcript. It is false in the 8-response subagents/ tree neither examined. Two independent checks, one blind spot, because both sampled the same way. Only an assertion running over all the data at execution time caught it.)

What review demonstrably does do. Measured across two passes: $0.66 and $1.11, roughly 1% of each pass, each finding approval-blocking defects — in CB-WP-0002 a target that would have made the evidence file certify a broken collector. The step pays for itself by a wide margin and the cost is not a reason to skip it.

Documentation requirement: the research process, the challenge, and the resulting improvements to the research are each documented in timestamped markdown files under history/:

history/YYMMDD-<slug>-research.md    # how the survey was conducted: sources,
                                     # queries, what was measured vs cited, dead ends
history/YYMMDD-<slug>-challenge.md   # the adversarial attack, verbatim
history/YYMMDD-<slug>-response.md    # answers/concessions and what changed in the survey

The polished survey artifact remains research/CB-RES-NNNN-<slug>.md; the history files preserve the unpolished trail so a later reader can judge how hard the survey was actually tested. For tier-M passes the review is optional but, when performed, follows the same format. If not approvable after the round, the loop returns to step 1 with the named gaps.

Step 3 — Decide

decisions/ADR-NNNN-<slug>.md: assimilate behind a port, reimplement, or hybrid — with the expected advantage stated per dimension (see rubric). An honest "worse here, better there, and why that trade is right" beats a claimed sweep of all four dimensions.

Step 4 — Specify

specs/<Capability>.md: the contracts, invariants, and — mandatory — the acceptance metrics table, each row tied to a baseline from step 1. A spec without measurable acceptance criteria is not done. Metrics follow the conventions in MetricsAndScenarios.md, including the rule that metric selection itself passes through a mini research step (metric provenance).

Every metric names its instrument, and is checked reachable. A row in the acceptance table carries the command that produces its number. A metric with no named instrument is a wish, not a metric.

The instrument must exist, and the target must come out of it. Naming a command is not the same as running one. A target computed by hand and merely labelled with a command is the same defect the rule was written to stop, one level down. Where the instrument is built later in the pass, the target is marked provisional: until the instrument emits it, and the spec is amended to whatever the instrument returns.

(v1.1, from CB-WP-0002: specs/CostAccounting.md AC-1 named cb-cost --pin fc76445 before that tool existed, and set the target to a hand-computed $92.87. When the tool was built it returned $93.32 — the hand computation carried a dedup bug the tool's own positive control caught. The metric satisfied v1.0's rule completely and was still wrong.)

A number inherited from earlier work is re-derived before it is used as a target, or it is cited as unverified. Quoting is not measuring.

(v1.1, from CB-WP-0002: the workplan opened with $248.46, inherited from a prior pass. Re-derivation put it at $92.21 — the quoted figure double-counted transcript lines and priced a three-model session at one model's rate. Neither error was of the harness-does-nothing class; both sums ran over real data, and a positive control would have passed them.)

Retargeting: the instrument may move a target, the implementation may not. A metric's target changes in only two ways, and they are not treated alike:

trigger requirement
corrected the instrument disproved the target — the target was computed by hand, or by an earlier tool with a defect legitimate in the same commit, provided the instrument's output is in that commit
retargeted the implementation missed the target and the target moves to accommodate it requires an ADR: old target, the measurement, and why the new target binds on future work rather than merely passing present work

The distinction is not the implementer's self-report of intent. It is mechanical: a correction is one where the target moves and the implementation does not. If the same commit changes both the target and the code the target measures, it is a retarget and needs the ADR.

(v1.1, from CB-WP-0002/0003: AM-4's targets were measured at 246,250 and set at 250,000 in one commit by the implementer after seeing the number — the structure this rule exists to stop. But CB-WP-0002 then moved AC-1 three times, correctly, each time because a new instrument disproved the old figure ($92.21 → $92.87 → $93.32 → $93.15). A blanket prohibition would have forbidden four legitimate corrections to catch one bad retarget.)

Applied retroactively: AM-4a and AM-4b are unratified until an ADR is written or they are changed. They were set by the implementer after seeing the measurement, in the commit that produced it, and the implementation changed in that same commit — a retarget by the test above. Tracked as an open item in history/260731-inner-loop-rule-audit.md.

A metric is checked against the contracts in its own spec. If a contract makes a target unreachable, one of the two is wrong and the conflict is resolved when it is noticed, not at the acceptance run. Re-check the table whenever a contract is added.

(v1.0, from CB-WP-0001: AM-4's ≤20-crate target was made unreachable by the K5 and K7 contracts written after it, and AM-12's cost metric was fully specified and never instrumented, so it could not be computed.)

Step 5 — Code loop

Implement iteratively. Each iteration:

change → cb-check (fmt, clippy, tests) → scenarios → benchmarks
       → compare against acceptance table → evidence row appended

Done when every acceptance metric meets or beats its baseline and the comparison numbers are committed as an evidence file (evidence/CB-EV-NNNN-<slug>.md). A failed scenario must yield a replay artifact an agent can re-execute locally.

Measurement validity — the positive control

Every benchmark and harness must assert that it performed the work it reports. Completing without error is not evidence of having done anything: a loop whose commands are all rejected runs fast and reports a throughput for work that never happened.

Concretely, a measurement harness must, on every run:

  • assert the unit of work produced its expected effect (events applied, rows written, moves accepted) — not merely that the call returned;
  • fail loudly rather than report a number when that assertion fails;
  • state the divisor used to convert raw timings into the metric's unit, pinned by a test so a workload change cannot silently rescale it.

A number from a run that cannot prove it did the work is void and must not reach an evidence file.

Every tool that reports a number exposes --self-test, and that self-test runs before the number is produced (make cost depends on make cost-test). The assertion must name a failure it detects, not merely exercise the happy path.

(v1.0+, from CB-WP-0002: cb-cost's dedup assertion fired on its first run against real data and aborted, catching a rule that was verified on 206/206 groups of the main transcript and false in the 8-response subagent tree. The generalization that failed — a property confirmed on the largest sample assumed to hold on the smallest — is not one review catches, because both the survey and the adversarial reviewer checked the same large sample.)

(v1.0, from CB-WP-0001: both serious errors in the first pass were of exactly this shape. A JS harness reported 8.4s for 100k moves while every move was being rejected, and a Rust benchmark reported 9.3M events/s — a 93× beat — while most rounds never completed because a stress gate rejected one player's action. The corrected figure was 5.6× lower. Adversarial review caught neither; both were claims about numbers, and review reads prose.)

Evidence states what it does not support

An evidence file that compares across runtimes, languages, or feature sets names the disanalogies explicitly, in the same section as the number. The reader must not have to infer that a ratio is not like-for-like. This is the parity-cap rule applied to the write-up: state the claim you will defend, and the claim you are not making.



Reference material

The four-dimension rubric, the survey template, the implementation rules each pass earned, the agentic-efficiency requirements, and the definition of done live in InnerLoopReference.md. They are normative; they are separated only so both files load whole.

(Split 2026-07-31: this file reached 407 lines against its own ~400-line loadability rule, and make loop-lint — added the same day to make that rule executable — failed on the commit that pushed it over. The rule caught its own author within an hour of being written.)