Scope cut first, on the maintainer's decision after a spend review: the project is 38% product / 62% loop-meta, cost per response is 2.9x worse than its best window, and INTENT stage 0 still lacks a CLI player and bots. CB-WP-0005 and CB-WP-0006 cost ~$74 — 31% of all spend — for zero measured efficiency gain. T02 and T04 are cancelled unstarted. T01: SH-1/SH-2/SH-3 now report over the window since the last commit, and the cumulative figure is retained but labelled "history, NOT the metric". The prediction held decisively — window 655,744 mean context against cumulative 255,307, a 2.6x gap against a 20% refutation threshold. A cumulative mean over 1,094 responses cannot detect a worsening trend because the history outvotes the present. T03: `make shape-budget`, modelled on CB-01/CB-02. Soft thresholds are the existing SessionShape targets; hard is 1.5x, set before the next measurement per §Step 4. Deliberately not in `make all` — failing the build on context would block committing, and committing is what closes the attribution window and is the natural point to compact, so a gate that blocks the remedy is a trap. It fires HARD on its first run: 656,574 against a 300,000 ceiling. InnerLoop v1.5 establishes the soft 25% meta budget. Workplans declare kind: product|meta|mixed and `make status` reports the share; mixed splits 50/50 and says so. Soft on purpose — a task already started may be finished, because stopping mid-task to satisfy a ratio wastes the work. What it forbids is opening new meta work above the line. A pass that exceeds it must say so in its evidence and name the product work displaced. First reading: 68% OVER, of $74.22 attributed. Product reads $0.00 because the only product workplan, CB-WP-0001, predates qualified task ids and its bare T## labels collide across passes — stated in the output rather than papered over. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
399 lines
20 KiB
Markdown
399 lines
20 KiB
Markdown
# The Inner Loop — Assimilate and Surpass
|
||
|
||
Status: **v1.5** — corrected from CB-WP-0007 (session shape) on
|
||
2026-08-01. Change from v1.4: a **soft 25% meta budget** (below), after a
|
||
spend review found the project 38% product / 62% loop-meta with cost per
|
||
response degraded 2.9x from its best window.
|
||
|
||
> **Meta budget — soft, 25% of spend per pass.** Work on the loop's own
|
||
> instruments and process is capped at a quarter of a pass. Measured by
|
||
> `make status` from each workplan's `kind:` frontmatter
|
||
> (`product` | `meta` | `mixed`).
|
||
>
|
||
> **Soft on purpose.** A task already started may be finished — stopping
|
||
> mid-task to satisfy a ratio wastes the work and leaves the tree in a
|
||
> worse state than either finishing or never starting. What the budget
|
||
> forbids is *opening* new meta work above the line.
|
||
>
|
||
> A pass that exceeds it **says so in its evidence file and names the
|
||
> product work displaced**. That is the whole enforcement: this is a
|
||
> reporting budget, not a gate, for the same reason the session-shape
|
||
> budget is (CB-RES-0005 §4) — it constrains judgment, not artifacts.
|
||
>
|
||
> *(v1.5, from the CB-WP-0007 spend review: CB-WP-0005 and CB-WP-0006 cost
|
||
> ~$74, 31% of all spend, for zero measured efficiency gain. Their return
|
||
> was correctness of claims, which is real and is not optimization. The
|
||
> budget exists so that distinction has to be made out loud.)*
|
||
|
||
v1.4 — corrected from CB-WP-0005 (assertion coverage) on
|
||
2026-07-31. Change from v1.3: where a claim rests on numbers, the
|
||
adversarial reviewer must read the assertion behind each quoted number and
|
||
**mutate it** — re-running the command that prints a number is not
|
||
verification of that number (§Step 2).
|
||
|
||
v1.3 changed from v1.2: single source of fact is now executable
|
||
(`make facts-check`, CB-WP-0004 T04), giving the duplicated-fact-drift
|
||
class its first gate.
|
||
|
||
v1.2 changes from v1.1: single source of fact; review targets the
|
||
harness and states its sampling limit; correction vs retarget; the chaos
|
||
roll's calibration window; the live cost budget. The design goal is now
|
||
stated: **optimize for cheap correction, not for exhaustive prevention.**
|
||
Rationale: `history/260731-loop-hardening-retrospective.md`.
|
||
|
||
v1.1 — corrected from CB-WP-0002 (cost accounting) on
|
||
2026-07-31. Changes from v1.0: the instrument must exist and emit its own
|
||
target; inherited numbers are re-derived before use; every reporting tool
|
||
exposes `--self-test`; cost is in the definition of done. Rationale:
|
||
`history/260731-cost-accounting-retrospective.md`.
|
||
|
||
v1.0 — survived its first full pass (CB-WP-0001, the GROUND game kernel)
|
||
and was corrected from it on 2026-07-31. Changes from v0.2: measurement
|
||
validity (the positive control), metric feasibility and instrument naming,
|
||
four implementation rules the pass earned, and the requirement that
|
||
evidence state what it does not support. Rationale and the failures behind
|
||
each: `history/260731-inner-loop-retrospective.md`.
|
||
|
||
Normative process for building every Clay-Borg capability. Referenced by
|
||
all workplans.
|
||
|
||
**Design goal (v1.2).** Three passes produced ten error instances across
|
||
four classes, and every pass produced a class the previous one had not
|
||
seen — prevention is not converging. Every one of those errors was
|
||
corrected inside the same session for under ~1% of the pass. The loop
|
||
therefore optimizes for **cheap correction**: keep the raw data so numbers
|
||
can be re-derived, keep artifacts small and committed so a wrong number is
|
||
one grep from everywhere that quotes it, and give every reported number a
|
||
command so re-running is free.
|
||
|
||
> **Single source of fact.** A number, rate, or target lives in exactly
|
||
> one place; everywhere else links to it. Where a copy is unavoidable it
|
||
> is generated by a command, not typed. *(v1.2: a price sheet inlined into
|
||
> a spec went stale within the hour of the real sheet changing, and one
|
||
> acceptance figure had to be chased across three artifacts each time it
|
||
> moved. No positive control catches this — both copies are internally
|
||
> consistent — and re-derivation does not either, because the copy
|
||
> reproduces whatever it was copied from.)*
|
||
>
|
||
> **Now executable (v1.3, CB-WP-0004 T04):** `make facts-check`. `facts.toml`
|
||
> is generated from the instruments, never edited; an artifact quoting a
|
||
> registry value tags it `<!-- fact:<key> -->` and the gate fails when the
|
||
> two disagree. Untagged literal copies are reported, not failed — that is
|
||
> the drift surface still uncovered, and naming it is more useful than
|
||
> pretending it is closed.
|
||
|
||
The loop's own optimization target is **agentic efficiency**:
|
||
every artifact it produces must be small enough to load whole, structured
|
||
enough to act on without interpretation, and falsifiable enough that an
|
||
agent can judge its own work without a human in the iteration.
|
||
|
||
---
|
||
|
||
## The five steps
|
||
|
||
```text
|
||
1 RESEARCH → research/CB-RES-NNNN-<slug>.md (survey, baselines)
|
||
2 APPROVE → decision recorded in the ADR (gate: survey complete?)
|
||
3 DECIDE → decisions/ADR-NNNN-<slug>.md (assimilate/reimplement/hybrid)
|
||
4 SPECIFY → specs/<Capability>.md (contracts + acceptance metrics)
|
||
5 CODE LOOP → code + scenarios + benchmarks (iterate until metrics beat baseline)
|
||
```
|
||
|
||
**Hard gate: no implementation code for a capability exists before its ADR
|
||
(step 3) is committed.**
|
||
|
||
## Loop tiers and the chaos roll
|
||
|
||
Every work packet declares a tier before work starts. The tier sets how
|
||
heavy steps 1–3 are; steps 4–5 (spec with metrics, code loop with evidence)
|
||
are never skipped for code-producing work.
|
||
|
||
| Tier | Weight of steps 1–3 | Structural trigger (forces at least this tier) |
|
||
|---|---|---|
|
||
| **L** | Full: separate survey, adversarial review, ADR | Creates a new capability port, or is named a high-leverage pass by the maintainer |
|
||
| **M** | Survey and ADR merged into one document; review optional | Touches a canonical interface, or adds/updates an external dependency |
|
||
| **S** | One provenance paragraph in the commit message | Everything else (utilities, fixes, refactors inside a boundary) |
|
||
|
||
**The chaos roll.** After deriving the structural tier, roll **d4**
|
||
(`shuf -i 1-4 -n 1`). On a **4**, the tier is instead picked uniformly at
|
||
random (`shuf -e S M L -n 1`), overriding the structural derivation — up or
|
||
down.
|
||
|
||
> **Calibration window, opened 2026-07-31 (CB-WP-0003 T06).** The rate was
|
||
> d10 and the mechanism **never fired**: two rolls across two workplans
|
||
> (CB-WP-0001: 9, CB-WP-0002: 2), against ~0.2 expected firings. At d10 and
|
||
> ~2 tier decisions per workplan it would take roughly twenty workplans to
|
||
> observe four overrides, so the mechanism was set at a rate that prevented
|
||
> its own evaluation — the one option T06 ruled out.
|
||
>
|
||
> Raised to **d4 (25%) for the next 12 tier declarations**, then evaluated
|
||
> and either kept, returned to d10, or deleted. Expected ~3 firings in the
|
||
> window, which is enough to see whether an overridden tier produces a
|
||
> different outcome than the argued one.
|
||
>
|
||
> **Stated cost:** a chaos-L override on work that would have been S buys a
|
||
> full survey, adversarial review, and ADR. Measured comparable: CB-WP-0001
|
||
> T03 (a tier-L survey) cost **$9.91**. At 25% over 12 declarations the
|
||
> window is expected to cost **$20–30**. That is the price of finding out
|
||
> whether the mechanism is worth keeping, and it is cheaper than carrying an
|
||
> unevaluated ritual indefinitely. Both rolls are recorded in the tier declaration
|
||
(`tier: M (structural L, chaos 10→M)`). **Record the roll every time,
|
||
including when it changes nothing** (`tier: L (structural L, chaos 4)`),
|
||
so a mechanism that never fires is visible rather than assumed. Purpose:
|
||
an occasional random
|
||
reweighting keeps the classification honest — arguing everything into S
|
||
stops paying off when audits can compare argued tiers against the random
|
||
sample — and occasionally forces a deep look at something "obviously
|
||
trivial", which is where local optima hide.
|
||
|
||
Chaos limits: a rolled-down tier relaxes *process* weight only. Invariants
|
||
(zero foreign types in canonical interfaces, determinism, passing
|
||
conformance suites) bind at every tier, and a rolled-down pass touching a
|
||
canonical interface still requires the interface change to be flagged in
|
||
the commit for retrospective review.
|
||
|
||
### Step 1 — Research
|
||
|
||
Identify the best implementation in existence for this capability. Produce
|
||
`research/CB-RES-NNNN-<slug>.md` following the survey template (below).
|
||
The survey is done when it can name, per dimension, a concrete
|
||
**benchmark-to-beat**: a number, a property, or a reproducible comparison —
|
||
not an impression.
|
||
|
||
**Runnable-baseline option.** For passes judged high-leverage (declared by
|
||
the maintainer or proposed in the survey and confirmed in the ADR), cited
|
||
numbers are not enough: the survey must ship a reproducible **baseline
|
||
harness** that runs the leading candidate on our machine against our
|
||
workload — the same scenario files where feasible. The harness ships with a
|
||
*fidelity note* stating what was and wasn't faithfully reproduced, so a
|
||
hastily wired competitor setup cannot silently inflate our advantage.
|
||
Where the option is not invoked (or the candidate isn't practically
|
||
runnable), comparisons against cited-only numbers are **directional**: the
|
||
evidence verdict for those rows caps at `parity`, never `better`.
|
||
|
||
### Step 2 — Approve (adversarial review)
|
||
|
||
For tier-L passes, approval is earned through an **adversarial review**: a
|
||
separate session (or agent) attempts to break the work. Exactly **one
|
||
round**: challenge, then response. The work is approvable only when every
|
||
challenge is either answered with evidence or conceded and folded in.
|
||
|
||
**The review target follows the risk.** Reviewing the survey was the
|
||
original rule, and it is the wrong target for a capability whose claim
|
||
rests on numbers — every serious error in this project has been in
|
||
measurement or build configuration, not in prose.
|
||
|
||
| the claim rests on | the reviewer is given | and must attempt |
|
||
|---|---|---|
|
||
| a survey of candidates | the survey | an omitted candidate; a stale or unverifiable benchmark; an unmeasured claim presented as measured |
|
||
| **numbers** | the survey **and the harness and the evidence file** | **reproduce the number independently**; state what the harness would report if the work silently stopped |
|
||
|
||
**What review cannot do — state this, do not discover it.** A reviewer
|
||
re-derives the author's claims and therefore inherits the author's
|
||
sampling. So:
|
||
|
||
> **The reviewer re-derives on a different sample than the author used.**
|
||
> Where only one sample exists, the review says so rather than reporting a
|
||
> clean verify.
|
||
>
|
||
> **And re-derivation is not enough (v1.4).** Where the claim rests on
|
||
> numbers, the reviewer must **read the assertion behind each quoted
|
||
> number and mutate it**: invert the property and require the suite to go
|
||
> red. Re-running the command that prints a number satisfies "reproduce
|
||
> independently" and finds nothing of this class.
|
||
>
|
||
> *(v1.4, from CB-WP-0005: `evidence/CB-EV-0001` reported `AM-7 replay |
|
||
> met, 2,290×` for a clause that asserts nothing — the hash reaches only a
|
||
> `println!`. The reviewer found it by opening a test out of curiosity and
|
||
> said so; no systematic step pointed there. Mutating it settled it in one
|
||
> command. This is the second verification step to inherit the author's
|
||
> blindness — the first was CB-WP-0002's dedup sample — and both fixes
|
||
> replace re-derivation with **adversarial execution**.)*
|
||
|
||
*(v1.1, from CB-WP-0002: the dedup invariant was verified on the main
|
||
transcript by the survey — 206/206 groups — and independently re-verified
|
||
by the reviewer, who used the same transcript. It is false in the
|
||
8-response `subagents/` tree neither examined. Two independent checks,
|
||
one blind spot, because both sampled the same way. Only an assertion
|
||
running over all the data at execution time caught it.)*
|
||
|
||
**What review demonstrably does do.** Measured across two passes:
|
||
**$0.66** and **$1.11**, roughly 1% of each pass, each finding
|
||
approval-blocking defects — in CB-WP-0002 a target that would have made
|
||
the evidence file certify a broken collector. The step pays for itself by
|
||
a wide margin and the cost is not a reason to skip it.
|
||
|
||
**Documentation requirement:** the research process, the challenge, and the
|
||
resulting improvements to the research are each documented in timestamped
|
||
markdown files under `history/`:
|
||
|
||
```text
|
||
history/YYMMDD-<slug>-research.md # how the survey was conducted: sources,
|
||
# queries, what was measured vs cited, dead ends
|
||
history/YYMMDD-<slug>-challenge.md # the adversarial attack, verbatim
|
||
history/YYMMDD-<slug>-response.md # answers/concessions and what changed in the survey
|
||
```
|
||
|
||
The polished survey artifact remains `research/CB-RES-NNNN-<slug>.md`; the
|
||
history files preserve the unpolished trail so a later reader can judge how
|
||
hard the survey was actually tested. For tier-M passes the review is
|
||
optional but, when performed, follows the same format. If not approvable
|
||
after the round, the loop returns to step 1 with the named gaps.
|
||
|
||
### Step 3 — Decide
|
||
|
||
`decisions/ADR-NNNN-<slug>.md`: assimilate behind a port, reimplement, or
|
||
hybrid — with the **expected advantage stated per dimension** (see rubric).
|
||
An honest "worse here, better there, and why that trade is right" beats a
|
||
claimed sweep of all four dimensions.
|
||
|
||
### Step 4 — Specify
|
||
|
||
`specs/<Capability>.md`: the contracts, invariants, and — mandatory — the
|
||
**acceptance metrics table**, each row tied to a baseline from step 1.
|
||
A spec without measurable acceptance criteria is not done. Metrics follow
|
||
the conventions in [MetricsAndScenarios.md](MetricsAndScenarios.md),
|
||
including the rule that metric selection itself passes through a mini
|
||
research step (metric provenance).
|
||
|
||
**Every metric names its instrument, and is checked reachable.** A row
|
||
in the acceptance table carries the command that produces its number.
|
||
A metric with no named instrument is a wish, not a metric.
|
||
|
||
**The instrument must exist, and the target must come out of it.**
|
||
Naming a command is not the same as running one. A target computed by
|
||
hand and merely *labelled* with a command is the same defect the rule
|
||
was written to stop, one level down. Where the instrument is built later
|
||
in the pass, the target is marked `provisional:` until the instrument
|
||
emits it, and the spec is amended to whatever the instrument returns.
|
||
|
||
*(v1.1, from CB-WP-0002: `specs/CostAccounting.md` AC-1 named
|
||
`cb-cost --pin fc76445` before that tool existed, and set the target to
|
||
a hand-computed $92.87. When the tool was built it returned $93.32 —
|
||
the hand computation carried a dedup bug the tool's own positive control
|
||
caught. The metric satisfied v1.0's rule completely and was still
|
||
wrong.)*
|
||
|
||
**A number inherited from earlier work is re-derived before it is used
|
||
as a target, or it is cited as unverified.** Quoting is not measuring.
|
||
|
||
*(v1.1, from CB-WP-0002: the workplan opened with $248.46, inherited
|
||
from a prior pass. Re-derivation put it at $92.21 — the quoted figure
|
||
double-counted transcript lines and priced a three-model session at one
|
||
model's rate. Neither error was of the harness-does-nothing class; both
|
||
sums ran over real data, and a positive control would have passed them.)*
|
||
|
||
**Retargeting: the instrument may move a target, the implementation may
|
||
not.** A metric's target changes in only two ways, and they are not
|
||
treated alike:
|
||
|
||
| | trigger | requirement |
|
||
|---|---|---|
|
||
| **corrected** | the instrument disproved the target — the target was computed by hand, or by an earlier tool with a defect | legitimate in the same commit, **provided the instrument's output is in that commit** |
|
||
| **retargeted** | the implementation missed the target and the target moves to accommodate it | requires an **ADR**: old target, the measurement, and why the new target binds on *future* work rather than merely passing present work |
|
||
|
||
The distinction is not the implementer's self-report of intent. It is
|
||
mechanical: **a correction is one where the target moves and the
|
||
implementation does not.** If the same commit changes both the target and
|
||
the code the target measures, it is a retarget and needs the ADR.
|
||
|
||
*(v1.1, from CB-WP-0002/0003: AM-4's targets were measured at 246,250 and
|
||
set at 250,000 in one commit by the implementer after seeing the number —
|
||
the structure this rule exists to stop. But CB-WP-0002 then moved AC-1
|
||
three times, correctly, each time because a new instrument disproved the
|
||
old figure ($92.21 → $92.87 → $93.32 → $93.15). A blanket prohibition
|
||
would have forbidden four legitimate corrections to catch one bad
|
||
retarget.)*
|
||
|
||
**Applied retroactively:** AM-4a and AM-4b are **unratified** until an ADR
|
||
is written or they are changed. They were set by the implementer after
|
||
seeing the measurement, in the commit that produced it, and the
|
||
implementation changed in that same commit — a retarget by the test above.
|
||
Tracked as an open item in `history/260731-inner-loop-rule-audit.md`.
|
||
|
||
**A metric is checked against the contracts in its own spec.** If a
|
||
contract makes a target unreachable, one of the two is wrong and the
|
||
conflict is resolved when it is noticed, not at the acceptance run.
|
||
Re-check the table whenever a contract is added.
|
||
|
||
*(v1.0, from CB-WP-0001: AM-4's ≤20-crate target was made unreachable by
|
||
the K5 and K7 contracts written after it, and AM-12's cost metric was
|
||
fully specified and never instrumented, so it could not be computed.)*
|
||
|
||
### Step 5 — Code loop
|
||
|
||
Implement iteratively. Each iteration:
|
||
|
||
```text
|
||
change → cb-check (fmt, clippy, tests) → scenarios → benchmarks
|
||
→ compare against acceptance table → evidence row appended
|
||
```
|
||
|
||
Done when every acceptance metric meets or beats its baseline and the
|
||
comparison numbers are committed as an evidence file
|
||
(`evidence/CB-EV-NNNN-<slug>.md`). A failed scenario must yield a replay
|
||
artifact an agent can re-execute locally.
|
||
|
||
#### Measurement validity — the positive control
|
||
|
||
**Every benchmark and harness must assert that it performed the work it
|
||
reports.** Completing without error is not evidence of having done
|
||
anything: a loop whose commands are all rejected runs fast and reports a
|
||
throughput for work that never happened.
|
||
|
||
Concretely, a measurement harness must, on every run:
|
||
|
||
- assert the unit of work produced its expected effect (events applied,
|
||
rows written, moves accepted) — not merely that the call returned;
|
||
- fail loudly rather than report a number when that assertion fails;
|
||
- state the divisor used to convert raw timings into the metric's unit,
|
||
pinned by a test so a workload change cannot silently rescale it.
|
||
|
||
**A number from a run that cannot prove it did the work is void** and
|
||
must not reach an evidence file.
|
||
|
||
**Every tool that reports a number exposes `--self-test`**, and that
|
||
self-test runs before the number is produced (`make cost` depends on
|
||
`make cost-test`). The assertion must name a failure it detects, not
|
||
merely exercise the happy path.
|
||
|
||
*(v1.0+, from CB-WP-0002: `cb-cost`'s dedup assertion fired on its first
|
||
run against real data and aborted, catching a rule that was verified on
|
||
206/206 groups of the main transcript and false in the 8-response
|
||
subagent tree. The generalization that failed — a property confirmed on
|
||
the largest sample assumed to hold on the smallest — is not one review
|
||
catches, because both the survey and the adversarial reviewer checked
|
||
the same large sample.)*
|
||
|
||
*(v1.0, from CB-WP-0001: both serious errors in the first pass were of
|
||
exactly this shape. A JS harness reported 8.4s for 100k moves while
|
||
every move was being rejected, and a Rust benchmark reported 9.3M
|
||
events/s — a 93× beat — while most rounds never completed because a
|
||
stress gate rejected one player's action. The corrected figure was 5.6×
|
||
lower. Adversarial review caught neither; both were claims about
|
||
numbers, and review reads prose.)*
|
||
|
||
#### Evidence states what it does not support
|
||
|
||
An evidence file that compares across runtimes, languages, or feature
|
||
sets **names the disanalogies explicitly**, in the same section as the
|
||
number. The reader must not have to infer that a ratio is not
|
||
like-for-like. This is the parity-cap rule applied to the write-up:
|
||
state the claim you will defend, and the claim you are not making.
|
||
|
||
---
|
||
|
||
---
|
||
|
||
## Reference material
|
||
|
||
The four-dimension rubric, the survey template, the implementation rules
|
||
each pass earned, the agentic-efficiency requirements, and the
|
||
definition of done live in
|
||
**[InnerLoopReference.md](InnerLoopReference.md)**. They are normative;
|
||
they are separated only so both files load whole.
|
||
|
||
*(Split 2026-07-31: this file reached 407 lines against its own ~400-line
|
||
loadability rule, and `make loop-lint` — added the same day to make that
|
||
rule executable — failed on the commit that pushed it over. The rule
|
||
caught its own author within an hour of being written.)*
|