clay-borg/specs/InnerLoop.md

207 lines
9.7 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# The Inner Loop — Assimilate and Surpass
Status: **v0.2 draft** — becomes v1.0 only after surviving its first full
pass (CB-WP-0001-T09 retrospective). v0.2 adds loop tiers with the chaos
roll, adversarial survey review, and the runnable-baseline option
(maintainer decision, 2026-07-31).
Normative process for building every Clay-Borg capability. Referenced by
all workplans. The loop's own optimization target is **agentic efficiency**:
every artifact it produces must be small enough to load whole, structured
enough to act on without interpretation, and falsifiable enough that an
agent can judge its own work without a human in the iteration.
---
## The five steps
```text
1 RESEARCH → research/CB-RES-NNNN-<slug>.md (survey, baselines)
2 APPROVE → decision recorded in the ADR (gate: survey complete?)
3 DECIDE → decisions/ADR-NNNN-<slug>.md (assimilate/reimplement/hybrid)
4 SPECIFY → specs/<Capability>.md (contracts + acceptance metrics)
5 CODE LOOP → code + scenarios + benchmarks (iterate until metrics beat baseline)
```
**Hard gate: no implementation code for a capability exists before its ADR
(step 3) is committed.**
## Loop tiers and the chaos roll
Every work packet declares a tier before work starts. The tier sets how
heavy steps 13 are; steps 45 (spec with metrics, code loop with evidence)
are never skipped for code-producing work.
| Tier | Weight of steps 13 | Structural trigger (forces at least this tier) |
|---|---|---|
| **L** | Full: separate survey, adversarial review, ADR | Creates a new capability port, or is named a high-leverage pass by the maintainer |
| **M** | Survey and ADR merged into one document; review optional | Touches a canonical interface, or adds/updates an external dependency |
| **S** | One provenance paragraph in the commit message | Everything else (utilities, fixes, refactors inside a boundary) |
**The chaos roll.** After deriving the structural tier, roll d10
(`shuf -i 1-10 -n 1`). On a **10**, the tier is instead picked uniformly at
random (`shuf -e S M L -n 1`), overriding the structural derivation — up or
down. Both rolls are recorded in the tier declaration
(`tier: M (structural L, chaos 10→M)`). Purpose: an occasional random
reweighting keeps the classification honest — arguing everything into S
stops paying off when audits can compare argued tiers against the random
sample — and occasionally forces a deep look at something "obviously
trivial", which is where local optima hide.
Chaos limits: a rolled-down tier relaxes *process* weight only. Invariants
(zero foreign types in canonical interfaces, determinism, passing
conformance suites) bind at every tier, and a rolled-down pass touching a
canonical interface still requires the interface change to be flagged in
the commit for retrospective review.
### Step 1 — Research
Identify the best implementation in existence for this capability. Produce
`research/CB-RES-NNNN-<slug>.md` following the survey template (below).
The survey is done when it can name, per dimension, a concrete
**benchmark-to-beat**: a number, a property, or a reproducible comparison —
not an impression.
**Runnable-baseline option.** For passes judged high-leverage (declared by
the maintainer or proposed in the survey and confirmed in the ADR), cited
numbers are not enough: the survey must ship a reproducible **baseline
harness** that runs the leading candidate on our machine against our
workload — the same scenario files where feasible. The harness ships with a
*fidelity note* stating what was and wasn't faithfully reproduced, so a
hastily wired competitor setup cannot silently inflate our advantage.
Where the option is not invoked (or the candidate isn't practically
runnable), comparisons against cited-only numbers are **directional**: the
evidence verdict for those rows caps at `parity`, never `better`.
### Step 2 — Approve (adversarial review)
For tier-L passes, approval is earned through an **adversarial review**: a
separate session (or agent), given only the survey document, attempts to
break it — an omitted candidate, a stale or unverifiable benchmark, an
unmeasured claim presented as measured. Exactly **one round**: challenge,
then response. The survey is approvable only when every challenge is either
answered with evidence or conceded and folded into the survey.
**Documentation requirement:** the research process, the challenge, and the
resulting improvements to the research are each documented in timestamped
markdown files under `history/`:
```text
history/YYMMDD-<slug>-research.md # how the survey was conducted: sources,
# queries, what was measured vs cited, dead ends
history/YYMMDD-<slug>-challenge.md # the adversarial attack, verbatim
history/YYMMDD-<slug>-response.md # answers/concessions and what changed in the survey
```
The polished survey artifact remains `research/CB-RES-NNNN-<slug>.md`; the
history files preserve the unpolished trail so a later reader can judge how
hard the survey was actually tested. For tier-M passes the review is
optional but, when performed, follows the same format. If not approvable
after the round, the loop returns to step 1 with the named gaps.
### Step 3 — Decide
`decisions/ADR-NNNN-<slug>.md`: assimilate behind a port, reimplement, or
hybrid — with the **expected advantage stated per dimension** (see rubric).
An honest "worse here, better there, and why that trade is right" beats a
claimed sweep of all four dimensions.
### Step 4 — Specify
`specs/<Capability>.md`: the contracts, invariants, and — mandatory — the
**acceptance metrics table**, each row tied to a baseline from step 1.
A spec without measurable acceptance criteria is not done. Metrics follow
the conventions in [MetricsAndScenarios.md](MetricsAndScenarios.md),
including the rule that metric selection itself passes through a mini
research step (metric provenance).
### Step 5 — Code loop
Implement iteratively. Each iteration:
```text
change → cb-check (fmt, clippy, tests) → scenarios → benchmarks
→ compare against acceptance table → evidence row appended
```
Done when every acceptance metric meets or beats its baseline and the
comparison numbers are committed as an evidence file
(`evidence/CB-EV-NNNN-<slug>.md`). A failed scenario must yield a replay
artifact an agent can re-execute locally.
---
## The four-dimension rubric
Every survey, ADR, and acceptance table is organized by these dimensions:
| Dimension | Question | Example measurable proxies |
|---|---|---|
| **D1 Ease of specification** | How simply can behavior be stated, tested, understood? | rules-to-scenario coverage %, spec lines per rule, time-to-first-correct-scenario for a fresh agent session |
| **D2 Efficiency of implementation** | How cheap to build and keep building? | source LOC, dependency count/weight, clean-build and incremental-build time, tokens-per-completed-task |
| **D3 Speed of execution** | How fast does it run? | benchmark wall-time vs baseline, events/sec, memory footprint, determinism overhead |
| **D4 Optionality** | How cleanly does it integrate, extend, get replaced? | public API surface size, count of leaked foreign types (must be 0), effort-to-swap measured by null/reference impl existence, WIT-expressibility |
Scoring is always **relative to the step-1 baseline**, never absolute:
`better / parity / worse / unmeasured` per proxy, with the number attached.
`unmeasured` is legal in a survey, illegal in an evidence file.
---
## Survey template (research/CB-RES-NNNN-<slug>.md)
```markdown
# CB-RES-NNNN: <capability>
capability: <canonical.capability.id>
status: draft | approved
## Candidates
Per candidate: origin, license, maturity, adoption; data model; mutation
mechanism; determinism/replay story; relevant performance (measured if
runnable locally, cited with source otherwise).
## Baselines (benchmark-to-beat)
| Dimension | Baseline holder | Metric | Value | Provenance |
(one row minimum per dimension; provenance = measured / cited / estimated)
## Verdict
Which candidate leads per dimension; what none of them do well
(the surpass opportunity); risks in the baselines themselves.
```
---
## Agentic-efficiency requirements
The loop exists to be driven by agents. Therefore:
1. **Whole-file loadability** — every loop artifact stays under ~400 lines;
split before exceeding, link with relative paths.
2. **Structured over prose** — tables and fenced blocks for anything a
later step must parse (baselines, acceptance metrics, evidence rows).
3. **One command surface** — all checks runnable through repo-root
commands (eventually `cb *`; until then, `make`/`cargo` aliases declared
in one place), each supporting deterministic, greppable output.
4. **Self-contained tasks** — a workplan task names its input artifacts and
output artifacts; a fresh session must be able to execute it from the
task text plus linked files alone.
5. **Evidence or it didn't happen** — claims of "better" live in committed
evidence files with numbers, never only in commit messages or chat.
6. **Token discipline** — per the global budget policy, a loop iteration
that exceeds its budget without measurable progress is stopped and
decomposed, not pushed through.
---
## Definition of done — one loop pass
A capability has completed the loop when all of the following are committed:
- [ ] research/CB-RES-NNNN with approved status and full baseline table
- [ ] decisions/ADR-NNNN with per-dimension expected advantage
- [ ] specs/<Capability>.md with acceptance-metrics table
- [ ] passing scenarios covering every numbered spec rule
- [ ] evidence/CB-EV-NNNN with final comparison vs baseline, no `unmeasured`
- [ ] retrospective note (may be one paragraph appended to the evidence
file): what the loop itself should change
```