clay-borg/tools/design-baseline.py
tegwick 503966bce9
Some checks failed
ci / check (push) Has been cancelled
CB-WP-0021 T06: fix AM-7's measurement, not its floor
The row folded a 5,000-event log against a 100,000-event log and compared
throughputs, which confounds 'does cost per event grow with history'
(the property it claims) with 'does streaming a 20x longer Vec cost more
per element' (a memory-hierarchy fact true of any program). It measured
the second and reported it as the first: importing the edition enlarged
the aggregate and the ratio fell to 0.845 with the state bounded.

Corrected to time the SAME 5,000 events on a state at depth 0 and on a
state at depth 100,000. Equal windows, equal event mix, so the only
difference left is history depth.

  corrected: clean 1.004, mutated 0.589 (red)
  old:       clean 0.845 (red on healthy code), mutated 0.751

It also runs in 8.5s instead of timing out: the first version re-walked
the 100k prefix every repetition, 200M untimed folds per sample, which
under the mutation never finished. A control that cannot be run is not a
control. It now advances to depth once per sample and clones.

Two of my own measurements here were wrong and both were caught by
measuring again. A 2-minute timeout killed the shell line before its
restoring cp ran, so three readings were taken on MUTATED code -- I
diagnosed an event-mix confound that did not exist and 'fixed' it. The
fix is kept on its merits; the justification was fiction. And the probe
that proved state was bounded had checked four of eleven collections.

make all exits 0.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-04 01:10:38 +02:00

99 lines
4.4 KiB
Python
Executable file

#!/usr/bin/env python3
"""Baseline harness: how findable, reproducible and answered are the
design findings this project has already produced?
The comparator is US, today. The external candidates (rulings databases,
model-checker traces, W3C provisional marks) are practices rather than
runnable software, so per InnerLoop Step 1 their rows are DIRECTIONAL and
cap at `parity`. This is the row that can be measured.
"""
import os, re, subprocess, sys, datetime
ROOT = "/home/worsch/clay-borg"
os.chdir(ROOT)
# The findings this project has actually produced, and where each lives.
FINDINGS = {
"U1..U10 underdetermined points": ["specs/GroundRules.md"],
"SOLVE on a face-down Problem": ["workplans/CB-WP-0018-the-browser-is-a-client.md",
"evidence/CB-EV-0016-the-browser-is-a-client.md"],
"GR-A13 wasted SOLVE": ["evidence/CB-EV-0007-stage-0.md"],
# RESOLVED 2026-08-04: ground-game ruled GR-S01's deal, the engine
# imports the edition, and the scenario was renamed from
# `-unreachable-` to `-reachable-`. Kept in the baseline because the
# baseline is a snapshot of what the survey measured, and a register
# that drops findings when they close cannot report a close rate.
"GR-E01 unreachable below 5 seats [RESOLVED]": [
"evidence/CB-EV-0007-stage-0.md",
"scenarios/ground/gr-e01-threshold-reachable-2p.yaml",
"workplans/CB-WP-0021-import-the-edition.md"],
"six provisional defaults": sorted(
os.path.join("scenarios/ground", f)
for f in os.listdir("scenarios/ground")
if f.endswith(".yaml")
and "provisional: true" in open(os.path.join("scenarios/ground", f)).read()),
"GR-E03/GR-E04 never played": ["evidence/CB-EV-0007-stage-0.md"],
}
def has_reproduction(paths):
"""A runnable thing: a scenario file, or a named test/command."""
for p in paths:
if p.startswith("scenarios/"):
return True
return False
def self_test():
"""The control that matters: a harness that read nothing must not
report a clean baseline. Every path this survey cites must exist, and
the reproduction test must be able to say NO — one that answered yes
for everything would report 100% and look excellent."""
results = []
def check(name, ok, detail=""):
results.append((name, ok, detail))
missing = [p for paths in FINDINGS.values() for p in paths
if not os.path.exists(p)]
check("every cited location exists", not missing, ", ".join(missing[:3]))
check("the finding set is not empty", len(FINDINGS) >= 6, f"{len(FINDINGS)}")
check("reproduction detection can say NO",
not has_reproduction(["evidence/CB-EV-0007-stage-0.md"]),
"a detector that always says yes would report 100%")
check("reproduction detection can say YES",
has_reproduction(["scenarios/ground/gr-e01-threshold-unreachable-2p.yaml"]))
# The number this survey turns on, pinned so a later edit cannot move
# it silently: 2 of 6 today.
repro_now = sum(has_reproduction(v) for v in FINDINGS.values())
check("the measured baseline is 2 of 6", repro_now == 2, f"{repro_now}/6")
print("design-baseline self-test (positive control)")
ok = True
for name, passed, det in results:
print(f" [{'ok ' if passed else 'FAIL'}] {name}" + (f"{det}" if det else ""))
ok &= passed
return 0 if ok else 1
if "--self-test" in sys.argv:
raise SystemExit(self_test())
print("BASELINE — design findings as they stand, 2026-08-03\n")
places = set()
repro = 0
for name, paths in FINDINGS.items():
places.update(paths)
r = has_reproduction(paths)
repro += r
print(f" {'repro' if r else ' - '} {len(paths)} location(s) {name}")
n = len(FINDINGS)
print(f"\n findings {n}")
print(f" with a runnable reproduction {repro}/{n} = {100*repro//n}%")
print(f" distinct files holding them {len(places)}")
print(f" single register? NO — {len(places)} files, no index")
# Time from raised to READ, for the one finding with a timestamp trail.
raised = datetime.date(2026, 7, 30) # hub message from clay-borg-custodian
read = datetime.date(2026, 8, 3) # marked read this session
print(f"\n U1..U10: raised {raised}, first READ {read}{(read-raised).days} days")
print(f" U1..U10: answered? NO — {(read-raised).days}+ days open, 0 of 10 ruled")