# The Inner Loop — Assimilate and Surpass Status: **v0.2 draft** — becomes v1.0 only after surviving its first full pass (CB-WP-0001-T09 retrospective). v0.2 adds loop tiers with the chaos roll, adversarial survey review, and the runnable-baseline option (maintainer decision, 2026-07-31). Normative process for building every Clay-Borg capability. Referenced by all workplans. The loop's own optimization target is **agentic efficiency**: every artifact it produces must be small enough to load whole, structured enough to act on without interpretation, and falsifiable enough that an agent can judge its own work without a human in the iteration. --- ## The five steps ```text 1 RESEARCH → research/CB-RES-NNNN-.md (survey, baselines) 2 APPROVE → decision recorded in the ADR (gate: survey complete?) 3 DECIDE → decisions/ADR-NNNN-.md (assimilate/reimplement/hybrid) 4 SPECIFY → specs/.md (contracts + acceptance metrics) 5 CODE LOOP → code + scenarios + benchmarks (iterate until metrics beat baseline) ``` **Hard gate: no implementation code for a capability exists before its ADR (step 3) is committed.** ## Loop tiers and the chaos roll Every work packet declares a tier before work starts. The tier sets how heavy steps 1–3 are; steps 4–5 (spec with metrics, code loop with evidence) are never skipped for code-producing work. | Tier | Weight of steps 1–3 | Structural trigger (forces at least this tier) | |---|---|---| | **L** | Full: separate survey, adversarial review, ADR | Creates a new capability port, or is named a high-leverage pass by the maintainer | | **M** | Survey and ADR merged into one document; review optional | Touches a canonical interface, or adds/updates an external dependency | | **S** | One provenance paragraph in the commit message | Everything else (utilities, fixes, refactors inside a boundary) | **The chaos roll.** After deriving the structural tier, roll d10 (`shuf -i 1-10 -n 1`). On a **10**, the tier is instead picked uniformly at random (`shuf -e S M L -n 1`), overriding the structural derivation — up or down. Both rolls are recorded in the tier declaration (`tier: M (structural L, chaos 10→M)`). Purpose: an occasional random reweighting keeps the classification honest — arguing everything into S stops paying off when audits can compare argued tiers against the random sample — and occasionally forces a deep look at something "obviously trivial", which is where local optima hide. Chaos limits: a rolled-down tier relaxes *process* weight only. Invariants (zero foreign types in canonical interfaces, determinism, passing conformance suites) bind at every tier, and a rolled-down pass touching a canonical interface still requires the interface change to be flagged in the commit for retrospective review. ### Step 1 — Research Identify the best implementation in existence for this capability. Produce `research/CB-RES-NNNN-.md` following the survey template (below). The survey is done when it can name, per dimension, a concrete **benchmark-to-beat**: a number, a property, or a reproducible comparison — not an impression. **Runnable-baseline option.** For passes judged high-leverage (declared by the maintainer or proposed in the survey and confirmed in the ADR), cited numbers are not enough: the survey must ship a reproducible **baseline harness** that runs the leading candidate on our machine against our workload — the same scenario files where feasible. The harness ships with a *fidelity note* stating what was and wasn't faithfully reproduced, so a hastily wired competitor setup cannot silently inflate our advantage. Where the option is not invoked (or the candidate isn't practically runnable), comparisons against cited-only numbers are **directional**: the evidence verdict for those rows caps at `parity`, never `better`. ### Step 2 — Approve (adversarial review) For tier-L passes, approval is earned through an **adversarial review**: a separate session (or agent), given only the survey document, attempts to break it — an omitted candidate, a stale or unverifiable benchmark, an unmeasured claim presented as measured. Exactly **one round**: challenge, then response. The survey is approvable only when every challenge is either answered with evidence or conceded and folded into the survey. **Documentation requirement:** the research process, the challenge, and the resulting improvements to the research are each documented in timestamped markdown files under `history/`: ```text history/YYMMDD--research.md # how the survey was conducted: sources, # queries, what was measured vs cited, dead ends history/YYMMDD--challenge.md # the adversarial attack, verbatim history/YYMMDD--response.md # answers/concessions and what changed in the survey ``` The polished survey artifact remains `research/CB-RES-NNNN-.md`; the history files preserve the unpolished trail so a later reader can judge how hard the survey was actually tested. For tier-M passes the review is optional but, when performed, follows the same format. If not approvable after the round, the loop returns to step 1 with the named gaps. ### Step 3 — Decide `decisions/ADR-NNNN-.md`: assimilate behind a port, reimplement, or hybrid — with the **expected advantage stated per dimension** (see rubric). An honest "worse here, better there, and why that trade is right" beats a claimed sweep of all four dimensions. ### Step 4 — Specify `specs/.md`: the contracts, invariants, and — mandatory — the **acceptance metrics table**, each row tied to a baseline from step 1. A spec without measurable acceptance criteria is not done. Metrics follow the conventions in [MetricsAndScenarios.md](MetricsAndScenarios.md), including the rule that metric selection itself passes through a mini research step (metric provenance). ### Step 5 — Code loop Implement iteratively. Each iteration: ```text change → cb-check (fmt, clippy, tests) → scenarios → benchmarks → compare against acceptance table → evidence row appended ``` Done when every acceptance metric meets or beats its baseline and the comparison numbers are committed as an evidence file (`evidence/CB-EV-NNNN-.md`). A failed scenario must yield a replay artifact an agent can re-execute locally. --- ## The four-dimension rubric Every survey, ADR, and acceptance table is organized by these dimensions: | Dimension | Question | Example measurable proxies | |---|---|---| | **D1 Ease of specification** | How simply can behavior be stated, tested, understood? | rules-to-scenario coverage %, spec lines per rule, time-to-first-correct-scenario for a fresh agent session | | **D2 Efficiency of implementation** | How cheap to build and keep building? | source LOC, dependency count/weight, clean-build and incremental-build time, tokens-per-completed-task | | **D3 Speed of execution** | How fast does it run? | benchmark wall-time vs baseline, events/sec, memory footprint, determinism overhead | | **D4 Optionality** | How cleanly does it integrate, extend, get replaced? | public API surface size, count of leaked foreign types (must be 0), effort-to-swap measured by null/reference impl existence, WIT-expressibility | Scoring is always **relative to the step-1 baseline**, never absolute: `better / parity / worse / unmeasured` per proxy, with the number attached. `unmeasured` is legal in a survey, illegal in an evidence file. --- ## Survey template (research/CB-RES-NNNN-.md) ```markdown # CB-RES-NNNN: capability: status: draft | approved ## Candidates Per candidate: origin, license, maturity, adoption; data model; mutation mechanism; determinism/replay story; relevant performance (measured if runnable locally, cited with source otherwise). ## Baselines (benchmark-to-beat) | Dimension | Baseline holder | Metric | Value | Provenance | (one row minimum per dimension; provenance = measured / cited / estimated) ## Verdict Which candidate leads per dimension; what none of them do well (the surpass opportunity); risks in the baselines themselves. ``` --- ## Agentic-efficiency requirements The loop exists to be driven by agents. Therefore: 1. **Whole-file loadability** — every loop artifact stays under ~400 lines; split before exceeding, link with relative paths. 2. **Structured over prose** — tables and fenced blocks for anything a later step must parse (baselines, acceptance metrics, evidence rows). 3. **One command surface** — all checks runnable through repo-root commands (eventually `cb *`; until then, `make`/`cargo` aliases declared in one place), each supporting deterministic, greppable output. 4. **Self-contained tasks** — a workplan task names its input artifacts and output artifacts; a fresh session must be able to execute it from the task text plus linked files alone. 5. **Evidence or it didn't happen** — claims of "better" live in committed evidence files with numbers, never only in commit messages or chat. 6. **Token discipline** — per the global budget policy, a loop iteration that exceeds its budget without measurable progress is stopped and decomposed, not pushed through. --- ## Definition of done — one loop pass A capability has completed the loop when all of the following are committed: - [ ] research/CB-RES-NNNN with approved status and full baseline table - [ ] decisions/ADR-NNNN with per-dimension expected advantage - [ ] specs/.md with acceptance-metrics table - [ ] passing scenarios covering every numbered spec rule - [ ] evidence/CB-EV-NNNN with final comparison vs baseline, no `unmeasured` - [ ] retrospective note (may be one paragraph appended to the evidence file): what the loop itself should change ```