T10: InnerLoop v1.2 — hardening does not converge, so optimize correction
Some checks failed
ci / check (push) Failing after 4s
Some checks failed
ci / check (push) Failing after 4s
The retrospective question was whether the mechanism set is complete or each pass still finds a new class. This pass produced both an eighth instance AND a fourth class, so the answer is the uncomfortable one. Ledger: 10 instances, 4 classes, across 3 workplans. Every pass has produced at least one class the previous pass had not seen. HDN harness-does-nothing 5 executable assertions TA trusted arithmetic 3 re-derivation SSB same-sample blind spot 1 assertions over ALL the data DFD duplicated-fact drift 2 NEW -- reading a copy against source DFD is genuinely distinct: no positive control catches it, because both copies are internally consistent, and re-derivation does not either, because the copy faithfully reproduces what it was copied from. Found when an inlined price sheet went stale within an hour of T11 changing the real one. So v1.2 stops trying to enumerate classes in advance. Every error in three passes was corrected in-session for under ~1% of the pass, so the stated design goal is now cheap CORRECTION: keep raw data so numbers are re-derivable, keep artifacts small and committed so a wrong number is one grep from everywhere quoting it, give every number a command. Plus the one rule the new class earns: single source of fact. The original hypothesis is revised rather than confirmed. "A rule that cannot be executed is not a rule" is wrong -- the two most valuable corrections in the project came from a decorative rule that cannot be automated (re-derive inherited numbers). An executable rule fires reliably and catches one class; a decorative one fires unreliably and can catch any class, including unnamed ones. Keep both. Gates this pass: loop-lint caught 3 real violations on first run, then failed on its own author within the hour when a T07 edit pushed InnerLoop.md to 407 lines against its own 400 limit. CB-WP-0003 complete: 11 of 11 tasks done. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
parent
e3d0df6690
commit
0da376f7df
3 changed files with 163 additions and 4 deletions
134
history/260731-loop-hardening-retrospective.md
Normal file
134
history/260731-loop-hardening-retrospective.md
Normal file
|
|
@ -0,0 +1,134 @@
|
|||
# 2026-07-31 — retrospective: does hardening converge?
|
||||
|
||||
CB-WP-0003 T10. Pass produced `tools/loop-lint.py`,
|
||||
`specs/SessionShape.md`, the rule audit, the live cost budget, the
|
||||
retarget test, dated price rates, M-D1-LNK, provisional expiry, and
|
||||
InnerLoop **v1.2**.
|
||||
|
||||
## The question this task was given
|
||||
|
||||
> Is the set of mechanisms now complete, or is each new pass still finding
|
||||
> a new class? If this pass produces an eighth instance in a fourth class,
|
||||
> the honest conclusion is that class-by-class hardening does not
|
||||
> converge, and the loop should optimize for *cheap detection and
|
||||
> correction* rather than for prevention — which is a different design.
|
||||
|
||||
**It produced both. The answer is: hardening does not converge.**
|
||||
|
||||
## The ledger
|
||||
|
||||
| # | class | instances | found by |
|
||||
|---|---|---|---|
|
||||
| 1 | **HDN** harness-does-nothing | **5** | executable assertions |
|
||||
| 2 | **TA** trusted arithmetic | **3** | re-derivation |
|
||||
| 3 | **SSB** same-sample blind spot | **1** | an assertion running over *all* data |
|
||||
| 4 | **DFD** duplicated-fact drift *(new)* | **2** | reading the copy against its source |
|
||||
|
||||
Ten instances, four classes. Three workplans in, **every pass has
|
||||
produced at least one instance of a class the previous pass had not
|
||||
seen.**
|
||||
|
||||
### New this pass
|
||||
|
||||
- **HDN #5 (latent).** `tools/rule-coverage.py` would have exited 0
|
||||
reporting `0/0` if its spec regex ever stopped matching — a silent pass
|
||||
in the tool that reports our headline AM-1 number. Never fired; found
|
||||
only because T02 required it to carry a `--self-test` and writing one
|
||||
meant asking what failure it should detect.
|
||||
- **TA #3.** *"0 of 330 tool calls were batched"* — 330 was the count of
|
||||
single-call responses, not the total. Carried unverified from
|
||||
CB-WP-0001's retrospective into CB-WP-0003 T04. Real figure: 7.8–8.6%
|
||||
of responses batched.
|
||||
- **DFD #1–2 (new class).** `specs/MetricsAndScenarios.md` §1a inlined a
|
||||
copy of the price sheet; T11 changed the real sheet and the copy went
|
||||
stale **within the same hour**. Separately, the acceptance figure
|
||||
$92.87/$0.66 had to be chased across a survey, a workplan, and an
|
||||
evidence file each time it moved. The defect is not arithmetic and not
|
||||
a harness — it is **a fact stored in two places, where updating one
|
||||
does not update the other**.
|
||||
|
||||
DFD is genuinely distinct: no positive control catches it (both copies
|
||||
are internally consistent), and re-derivation does not either (the copy
|
||||
reproduces whatever it was copied from). It is caught only by reading a
|
||||
copy against its source, which nothing in the loop required.
|
||||
|
||||
## What this implies, and what changed because of it
|
||||
|
||||
Prevention has not converged: four passes, four classes, no sign of
|
||||
saturation. But the *correction* side has been consistently cheap and
|
||||
fast:
|
||||
|
||||
| correction | mechanism | cost |
|
||||
|---|---|---|
|
||||
| $248.46 → $92.21 | re-derivation | ~0 |
|
||||
| $92.21 → $92.87 | adversarial review | $1.11 |
|
||||
| $92.87 → $93.32 | positive control | ~0 |
|
||||
| $93.32 → $93.15 | dated-rate fix | ~0 |
|
||||
|
||||
Every error found in three passes was corrected inside the same session
|
||||
for under ~1% of the pass. **So the loop should stop trying to enumerate
|
||||
failure classes in advance and invest in the properties that make
|
||||
correction cheap** — which it turns out already has, accidentally:
|
||||
|
||||
1. **Numbers are re-derivable** because the raw data (transcripts, git,
|
||||
the price sheet) is kept, not just the conclusions.
|
||||
2. **Artifacts are small and committed**, so a wrong number is one grep
|
||||
from every place that quotes it.
|
||||
3. **Every reported number has a command**, so re-running is free.
|
||||
|
||||
v1.2 makes this the stated design goal rather than an accident, and adds
|
||||
the one rule the new class earns:
|
||||
|
||||
> **Single source of fact.** A number, rate, or target lives in exactly
|
||||
> one place. Everywhere else links to it. Where a copy is unavoidable, it
|
||||
> is generated by a command, not typed. *(CA-15 said this for evidence
|
||||
> tables; DFD shows it applies to specs quoting other specs.)*
|
||||
|
||||
## What the gates did this pass
|
||||
|
||||
Worth recording because it is the cheapest possible evidence that the
|
||||
executable-rule thesis is at least partly right:
|
||||
|
||||
- `loop-lint` found three real violations on its **first** run —
|
||||
including a 543-line spec against a ~400-line rule the loop had stated
|
||||
since v0.2 and never measured.
|
||||
- `loop-lint` then **failed on its own author within the hour**, when a
|
||||
T07 edit pushed `specs/InnerLoop.md` to 407 lines. The file was split
|
||||
because a gate said so, not because anyone noticed.
|
||||
- `cb-cost`'s dedup assertion fired on real data on its first run and
|
||||
refused to print a number.
|
||||
|
||||
Three gates, three catches, all on first contact. Against that: the two
|
||||
most valuable corrections of the whole project ($248.46 → $92.21, and the
|
||||
batching claim) came from a rule that is **decorative and cannot be made
|
||||
executable** — *re-derive inherited numbers*. Both facts are true and the
|
||||
audit records them side by side.
|
||||
|
||||
## Revised hypothesis
|
||||
|
||||
The workplan opened with: *a loop rule that cannot be executed is not a
|
||||
rule.* Three passes of evidence say that is wrong as stated. The version
|
||||
this pass supports:
|
||||
|
||||
> An executable rule fires reliably and catches one class. A decorative
|
||||
> rule fires unreliably and can catch any class, including ones not yet
|
||||
> named. **Keep both, and never let a decorative rule be the only defence
|
||||
> for a class that matters** — but do not delete it merely because it
|
||||
> cannot be automated, which is what this workplan originally proposed.
|
||||
|
||||
## Open, not closed
|
||||
|
||||
- **AM-4a/AM-4b are unratified retargets** (T07). `make dep-weight`
|
||||
enforces targets no reviewed decision stands behind. Maintainer
|
||||
decision.
|
||||
- **9 of 58 rules are claimed by a scenario and named nowhere in the
|
||||
aggregate** (T09, M-D1-LNK). Reported, not failing, because tagging
|
||||
them without confirming implementation is the overclaim the loop
|
||||
already committed once.
|
||||
- **5 provisional U-item defaults** now have an owner and an age, and
|
||||
ground-game has not ruled (T08).
|
||||
- **The chaos roll is in a 12-declaration calibration window** at d4
|
||||
(T06), expected to cost $20–30 and to be decided on evidence.
|
||||
- **DFD has no gate.** The single-source-of-fact rule added in v1.2 is
|
||||
prose. Making it executable — detecting a number duplicated across
|
||||
artifacts — is the obvious next candidate and is not attempted here.
|
||||
|
|
@ -1,6 +1,13 @@
|
|||
# The Inner Loop — Assimilate and Surpass
|
||||
|
||||
Status: **v1.1** — corrected from CB-WP-0002 (cost accounting) on
|
||||
Status: **v1.2** — corrected from CB-WP-0003 (loop hardening) on
|
||||
2026-07-31. Changes from v1.1: single source of fact; review targets the
|
||||
harness and states its sampling limit; correction vs retarget; the chaos
|
||||
roll's calibration window; the live cost budget. The design goal is now
|
||||
stated: **optimize for cheap correction, not for exhaustive prevention.**
|
||||
Rationale: `history/260731-loop-hardening-retrospective.md`.
|
||||
|
||||
v1.1 — corrected from CB-WP-0002 (cost accounting) on
|
||||
2026-07-31. Changes from v1.0: the instrument must exist and emit its own
|
||||
target; inherited numbers are re-derived before use; every reporting tool
|
||||
exposes `--self-test`; cost is in the definition of done. Rationale:
|
||||
|
|
@ -14,7 +21,25 @@ evidence state what it does not support. Rationale and the failures behind
|
|||
each: `history/260731-inner-loop-retrospective.md`.
|
||||
|
||||
Normative process for building every Clay-Borg capability. Referenced by
|
||||
all workplans. The loop's own optimization target is **agentic efficiency**:
|
||||
all workplans.
|
||||
|
||||
**Design goal (v1.2).** Three passes produced ten error instances across
|
||||
four classes, and every pass produced a class the previous one had not
|
||||
seen — prevention is not converging. Every one of those errors was
|
||||
corrected inside the same session for under ~1% of the pass. The loop
|
||||
therefore optimizes for **cheap correction**: keep the raw data so numbers
|
||||
can be re-derived, keep artifacts small and committed so a wrong number is
|
||||
one grep from everywhere that quotes it, and give every reported number a
|
||||
command so re-running is free.
|
||||
|
||||
> **Single source of fact.** A number, rate, or target lives in exactly
|
||||
> one place; everywhere else links to it. Where a copy is unavoidable it
|
||||
> is generated by a command, not typed. *(v1.2: a price sheet inlined into
|
||||
> a spec went stale within the hour of the real sheet changing, and one
|
||||
> acceptance figure had to be chased across three artifacts each time it
|
||||
> moved. No positive control catches this — both copies are internally
|
||||
> consistent — and re-derivation does not either, because the copy
|
||||
> reproduces whatever it was copied from.)* The loop's own optimization target is **agentic efficiency**:
|
||||
every artifact it produces must be small enough to load whole, structured
|
||||
enough to act on without interpretation, and falsifiable enough that an
|
||||
agent can judge its own work without a human in the iteration.
|
||||
|
|
|
|||
|
|
@ -1,7 +1,7 @@
|
|||
---
|
||||
id: CB-WP-0003
|
||||
title: "Harden the inner loop: executable rules, session economics, dead policy"
|
||||
status: active
|
||||
status: done
|
||||
state_hub_workstream_id: "39d61dc0-870d-45c1-a595-bcf91f289dce"
|
||||
---
|
||||
|
||||
|
|
@ -350,7 +350,7 @@ State Hub: a schema that cannot hold the fact it needs.
|
|||
|
||||
```task
|
||||
id: CB-WP-0003-T10
|
||||
status: todo
|
||||
status: done
|
||||
priority: low
|
||||
state_hub_task_id: "f43de208-92b1-4f0e-9236-a9d27f3ec451"
|
||||
```
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue