9.7 KiB
The Inner Loop — Assimilate and Surpass
Status: v0.2 draft — becomes v1.0 only after surviving its first full pass (CB-WP-0001-T09 retrospective). v0.2 adds loop tiers with the chaos roll, adversarial survey review, and the runnable-baseline option (maintainer decision, 2026-07-31).
Normative process for building every Clay-Borg capability. Referenced by all workplans. The loop's own optimization target is agentic efficiency: every artifact it produces must be small enough to load whole, structured enough to act on without interpretation, and falsifiable enough that an agent can judge its own work without a human in the iteration.
The five steps
1 RESEARCH → research/CB-RES-NNNN-<slug>.md (survey, baselines)
2 APPROVE → decision recorded in the ADR (gate: survey complete?)
3 DECIDE → decisions/ADR-NNNN-<slug>.md (assimilate/reimplement/hybrid)
4 SPECIFY → specs/<Capability>.md (contracts + acceptance metrics)
5 CODE LOOP → code + scenarios + benchmarks (iterate until metrics beat baseline)
Hard gate: no implementation code for a capability exists before its ADR (step 3) is committed.
Loop tiers and the chaos roll
Every work packet declares a tier before work starts. The tier sets how heavy steps 1–3 are; steps 4–5 (spec with metrics, code loop with evidence) are never skipped for code-producing work.
| Tier | Weight of steps 1–3 | Structural trigger (forces at least this tier) |
|---|---|---|
| L | Full: separate survey, adversarial review, ADR | Creates a new capability port, or is named a high-leverage pass by the maintainer |
| M | Survey and ADR merged into one document; review optional | Touches a canonical interface, or adds/updates an external dependency |
| S | One provenance paragraph in the commit message | Everything else (utilities, fixes, refactors inside a boundary) |
The chaos roll. After deriving the structural tier, roll d10
(shuf -i 1-10 -n 1). On a 10, the tier is instead picked uniformly at
random (shuf -e S M L -n 1), overriding the structural derivation — up or
down. Both rolls are recorded in the tier declaration
(tier: M (structural L, chaos 10→M)). Purpose: an occasional random
reweighting keeps the classification honest — arguing everything into S
stops paying off when audits can compare argued tiers against the random
sample — and occasionally forces a deep look at something "obviously
trivial", which is where local optima hide.
Chaos limits: a rolled-down tier relaxes process weight only. Invariants (zero foreign types in canonical interfaces, determinism, passing conformance suites) bind at every tier, and a rolled-down pass touching a canonical interface still requires the interface change to be flagged in the commit for retrospective review.
Step 1 — Research
Identify the best implementation in existence for this capability. Produce
research/CB-RES-NNNN-<slug>.md following the survey template (below).
The survey is done when it can name, per dimension, a concrete
benchmark-to-beat: a number, a property, or a reproducible comparison —
not an impression.
Runnable-baseline option. For passes judged high-leverage (declared by
the maintainer or proposed in the survey and confirmed in the ADR), cited
numbers are not enough: the survey must ship a reproducible baseline
harness that runs the leading candidate on our machine against our
workload — the same scenario files where feasible. The harness ships with a
fidelity note stating what was and wasn't faithfully reproduced, so a
hastily wired competitor setup cannot silently inflate our advantage.
Where the option is not invoked (or the candidate isn't practically
runnable), comparisons against cited-only numbers are directional: the
evidence verdict for those rows caps at parity, never better.
Step 2 — Approve (adversarial review)
For tier-L passes, approval is earned through an adversarial review: a separate session (or agent), given only the survey document, attempts to break it — an omitted candidate, a stale or unverifiable benchmark, an unmeasured claim presented as measured. Exactly one round: challenge, then response. The survey is approvable only when every challenge is either answered with evidence or conceded and folded into the survey.
Documentation requirement: the research process, the challenge, and the
resulting improvements to the research are each documented in timestamped
markdown files under history/:
history/YYMMDD-<slug>-research.md # how the survey was conducted: sources,
# queries, what was measured vs cited, dead ends
history/YYMMDD-<slug>-challenge.md # the adversarial attack, verbatim
history/YYMMDD-<slug>-response.md # answers/concessions and what changed in the survey
The polished survey artifact remains research/CB-RES-NNNN-<slug>.md; the
history files preserve the unpolished trail so a later reader can judge how
hard the survey was actually tested. For tier-M passes the review is
optional but, when performed, follows the same format. If not approvable
after the round, the loop returns to step 1 with the named gaps.
Step 3 — Decide
decisions/ADR-NNNN-<slug>.md: assimilate behind a port, reimplement, or
hybrid — with the expected advantage stated per dimension (see rubric).
An honest "worse here, better there, and why that trade is right" beats a
claimed sweep of all four dimensions.
Step 4 — Specify
specs/<Capability>.md: the contracts, invariants, and — mandatory — the
acceptance metrics table, each row tied to a baseline from step 1.
A spec without measurable acceptance criteria is not done. Metrics follow
the conventions in MetricsAndScenarios.md,
including the rule that metric selection itself passes through a mini
research step (metric provenance).
Step 5 — Code loop
Implement iteratively. Each iteration:
change → cb-check (fmt, clippy, tests) → scenarios → benchmarks
→ compare against acceptance table → evidence row appended
Done when every acceptance metric meets or beats its baseline and the
comparison numbers are committed as an evidence file
(evidence/CB-EV-NNNN-<slug>.md). A failed scenario must yield a replay
artifact an agent can re-execute locally.
The four-dimension rubric
Every survey, ADR, and acceptance table is organized by these dimensions:
| Dimension | Question | Example measurable proxies |
|---|---|---|
| D1 Ease of specification | How simply can behavior be stated, tested, understood? | rules-to-scenario coverage %, spec lines per rule, time-to-first-correct-scenario for a fresh agent session |
| D2 Efficiency of implementation | How cheap to build and keep building? | source LOC, dependency count/weight, clean-build and incremental-build time, tokens-per-completed-task |
| D3 Speed of execution | How fast does it run? | benchmark wall-time vs baseline, events/sec, memory footprint, determinism overhead |
| D4 Optionality | How cleanly does it integrate, extend, get replaced? | public API surface size, count of leaked foreign types (must be 0), effort-to-swap measured by null/reference impl existence, WIT-expressibility |
Scoring is always relative to the step-1 baseline, never absolute:
better / parity / worse / unmeasured per proxy, with the number attached.
unmeasured is legal in a survey, illegal in an evidence file.
Survey template (research/CB-RES-NNNN-.md)
# CB-RES-NNNN: <capability>
capability: <canonical.capability.id>
status: draft | approved
## Candidates
Per candidate: origin, license, maturity, adoption; data model; mutation
mechanism; determinism/replay story; relevant performance (measured if
runnable locally, cited with source otherwise).
## Baselines (benchmark-to-beat)
| Dimension | Baseline holder | Metric | Value | Provenance |
(one row minimum per dimension; provenance = measured / cited / estimated)
## Verdict
Which candidate leads per dimension; what none of them do well
(the surpass opportunity); risks in the baselines themselves.
Agentic-efficiency requirements
The loop exists to be driven by agents. Therefore:
- Whole-file loadability — every loop artifact stays under ~400 lines; split before exceeding, link with relative paths.
- Structured over prose — tables and fenced blocks for anything a later step must parse (baselines, acceptance metrics, evidence rows).
- One command surface — all checks runnable through repo-root
commands (eventually
cb *; until then,make/cargoaliases declared in one place), each supporting deterministic, greppable output. - Self-contained tasks — a workplan task names its input artifacts and output artifacts; a fresh session must be able to execute it from the task text plus linked files alone.
- Evidence or it didn't happen — claims of "better" live in committed evidence files with numbers, never only in commit messages or chat.
- Token discipline — per the global budget policy, a loop iteration that exceeds its budget without measurable progress is stopped and decomposed, not pushed through.
Definition of done — one loop pass
A capability has completed the loop when all of the following are committed:
- research/CB-RES-NNNN with approved status and full baseline table
- decisions/ADR-NNNN with per-dimension expected advantage
- specs/.md with acceptance-metrics table
- passing scenarios covering every numbered spec rule
- evidence/CB-EV-NNNN with final comparison vs baseline, no
unmeasured - retrospective note (may be one paragraph appended to the evidence file): what the loop itself should change