5 commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
| 302fc95c97 |
GR-E03 and GR-E04 played to the end — F14 closed, and the reason they were
Some checks failed
ci / check (push) Failing after 4s
unplayed was ours Tier S (a fix and a measurement inside a boundary; chaos d8=4 from CB-WP-0029's roll, no override). cb-play built EVERY game with ScoringMode::SharedGround and passed an empty patch. The mode was settable in scenarios and not from the driver, so two of the three shipped modes were unreachable from the only way anyone actually plays. F14 sat open for a week because nobody could reach the thing it was about. --mode added. All three now play out and give DIFFERENT WINNERS FROM IDENTICAL PLAY: shared -> all four seats (mastery 4), common -> P3 alone (top personal scorer), coalitions -> P1+P2 (best Bond network, 4>3>2). Same 37 commands, three answers. AND THEY ANSWER F17'S OPEN QUESTION. I had flagged that ATTACK might earn its place where Blame costs personal score. It does not, in any mode: SHARED GROUND 132/165/190/200 -> identical free but pointless COMMON PROBLEM 59/52/48/44 -> 59/52/48/34 a cost at six seats BONDED COALITIONS 131/134/132/116 -> 59/52/48/34 roughly halved The coalitions row has a mechanism and the data confirms it unprompted. GR-A07 flips a Bond to a Rivalry on Attack, and GR-E04 scores Bond NETWORKS -- so attacking destroys the thing that scores. And the attacking numbers in E04 are IDENTICAL to E03's, which is exactly what that predicts: break every Bond and each seat is a coalition of one, so GR-E04 degenerates into GR-E03. That check was not designed; it fell out. F14 -> applied. F17 strengthened and no longer bounded to co-op: ATTACK has no mode in which it helps, and one where it actively destroys your score. Still framed as a question rather than a verdict. DARVO is the pattern the game is about not falling into, so a self-destructive ATTACK may be the design. What ground-game has to decide is whether the namesake mechanic being unreachable in competent play -- in all three modes -- is intended. make all: exit 0. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
|||
| f4eeddd726 |
25 test games, no faults — and F17 gets the artifact that changes what it
Some checks failed
ci / check (push) Failing after 4s
says Five games per seat count, 2-6 players. NO ANOMALIES: every game reaches 5 rounds with an outcome, no stalls, no stress above the cap, no over-claimed Problems. But the series showed something a crash never would. DARVO NEVER FIRED IN 25 GAMES and stress never exceeded 2. Measured wider: GreedyPolicy plays ATTACK exactly ZERO times in 10,000 selections across 500 games. THAT NUMBER IS ABOUT OUR BOT, NOT THE GAME. bot.rs ranks `Action::Attack => 10`, below everything. Reporting "the game gives no incentive to attack" from a policy we programmed to rank attack last would have been CB-WP-0025's C4 error committed again -- a single policy's behaviour presented as the game's. So the artifact varies exactly one number: ATTACK's rank in an otherwise identical policy, 200 games per cell. rank 10 (below all): 132/165/190/200/200 wins, 0 attacks, 0 DARVO rank 75 (above SUPPORT): 132/165/190/200/200 wins, 315-923, 13-218 rank 95 (above SOLVE): 0/0/0/0/0 wins, 1400-5170, 400-1000 THE MIDDLE ROW IS THE FINDING. Identical win counts at every seat count, while attacking hundreds of times and arming DARVO repeatedly. Attacking is not punished -- it is INERT with respect to the goal. Group success is a function of SOLVE alone, and ATTACK costs anything only when it ranks above SOLVE and displaces it. The maintainer was right and the reason is sharper than his phrasing: there is no incentive because there is no PATH. ATTACK's effects (Stress, Rivalry, DARVO) feed nothing that decides group_success. Bounded honestly to SHARED GROUND. Blame costs PERSONAL score, so ATTACK may earn its place in GR-E03 and GR-E04 -- which have never been played to the end (F14), and that is where to ask next. And this is NOT a claim the game is broken: DARVO is the pattern the game is about not falling into, so a self-destructive ATTACK may be the design. The question for ground-game is whether the namesake mechanic being unreachable in competent co-op play is intended. F17 promoted from note to raised, with games/ground/examples/attack-value.rs as its reproduction. Register: 18 findings, 8 with a resolving reproduction. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
|||
| 7ed9fc730a |
CB-WP-0025 T06/T07: a difficulty baseline that reports its own confound
Some checks failed
ci / check (push) Has been cancelled
T06. games/ground/examples/difficulty.rs, make difficulty, wired into make self-tests, and a report file in ground-game under GROUND-WP-0005 with a hub message pointing at it. THE REPORT OPENS WITH THE RETRACTION, because what this task was written to send was withdrawn by T02 and GROUND-WP-0005 is blocked on exactly that number. They are told, in the first section, that we nearly sent them "the game is too easy at 5-6 seats" and why it was wrong. seats winnable greedy random first-legal spread 2p 60% 60.0% 5.0% 76.7% 71.7 3p 93% 88.3% 6.7% 25.0% 81.7 4p 100% 93.3% 6.7% 30.0% 86.7 5p 100% 100.0% 3.3% 0.0% 100.0 6p 100% 100.0% 3.3% 0.0% 100.0 SPREAD justifies the whole redesign: 71.7 to 100.0 points between three trivial policies. The table now shows why no single rate is a difficulty rather than asserting it. And the 5-6 rows point the OPPOSITE way from the withdrawn claim -- first-legal 0% against greedy 100% is the widest spread in the table, which suggests play matters MORE there, not less. Neither reading is established and the report says so. The confound is stated in the tool's own output, not only in prose: `winnable` is conditioned on greedy's play up to the final round, because searching from round 1 is unaffordable. Presenting it as a property of the deal would repeat this pass's error in a subtler form -- which is exactly how a corrected project reintroduces a defect. NO THRESHOLD CHANGES ARE PROPOSED. The instrument can fail (spec §5): a witness must replay to a win, an unwinnable position must report searched-out rather than a budget cut, a one-node budget must not claim exhaustion, and the policy panel must actually disagree. difficulty-baseline.rs marked superseded, kept as the survey's dated snapshot. Registered as F16, inconsistent / withdrawn. T07. evidence/CB-EV-0024. Five of nine defects came only from the review; four from execution, and all four of those were in work written after it. The wrong-denominator family now has five instances and still no control -- facts-check catches copies that disagree, nothing catches a number computed correctly against the wrong base. Tier L was an over-declaration (no port, structurally M) and paid for itself anyway, because the review is L-only. Chaos window 2 will close with zero overrides, making its retirement condition untestable. Named as open rather than implied done: the witness is NOT wired to the ending page. The search works; the browser cannot ask it yet. make all: exit 0. loop-lint clean. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
|||
| 1f0f652920 |
CB-WP-0025 T02: the review withdrew the finding, and the fifth wrong
Some checks are pending
ci / check (push) Waiting to run
premise never left the repo Separate agent, second tier-L review in this project. Six of seven challenges conceded. The survey's headline finding is WITHDRAWN, not softened. C4 kills it, and the reviewer ranked it fourth. A FirstLegal policy -- take legal[0], no heuristic at all -- scores 0% at five and six seats where GreedyPolicy scores 100%, and 77.5% at two seats where greedy scores 66%. Two unsophisticated agents span the entire range at the same seat count. "The game is too easy at 5-6 seats" is therefore a statement about GreedyPolicy, not about GROUND. The rescue the reviewer offered -- greedy hits the 12-point ceiling in 200/200 deals, so the 6p row is a rules claim -- dies on the same data: FirstLegal reaches that ceiling never. C1: the per-node cost was wrong by 30-50x. The timer started before the seed loop, so "us/node" included two setups, an entire greedy game and a full validate+fold replay, divided by player-decision count. The tell was in my own published output and I did not look at it: the figure FELL (161/139/112) as branching ROSE (4.7/7.4/9.1), which no per-enumeration cost can do. Re-measured with the clock around legal_commands alone: 3.0/3.5/4.1 us, now rising with branching. The reviewer measured 15.6-20.4 by a different isolation; we disagree by ~5x and neither has established which is right, so T04 must benchmark it with criterion rather than adopt either number. C6: "exhaustive search is out at any seat count" is false -- ~3 seconds over the last two rounds at 3p. With C1's correction the budget is ~10^5-10^6 nodes and bounded endgame search fits, so ADR-0013 cannot open with "exhaustive is impossible, therefore determinized sampling" -- especially as sampling carries strategy fusion that exhaustive search does not. C3: the finding failed the admissibility rule this project wrote nine hours earlier. 6/9/12 are sums where GROUND-WP-0004 T02 requires per-priority rows, and the harness has no assertions, no --self-test and no make target, so nothing can turn it red -- a `default` artifact wearing a `counterexample` label, by CB-WP-0022 T05's own distinction. C2: the ratio story explains nothing; 3p and 4p share deal, threshold and ratio and differ by 12.5 points of win rate. C5: "explains the maintainer's report" is contradicted by lib.rs:2487, which records his losses as 3-player games on the pre-ruling deal, arithmetically unwinnable at 6 against 7. T06 exists to report to GROUND-WP-0005, which is BLOCKED waiting on a difficulty baseline. Had this proceeded they would have been invited to move thresholds on the strength of one bot's behaviour. That is the fifth wrong premise this project would have sent them, and the second stopped by an adversarial review rather than by a control. Both tier-L reviews here have now caught a false headline that every gate passed. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
|||
| 469d00d679 |
CB-WP-0025 T01: survey -- the baseline found the game before the solver did
CB-RES-0008 with a runnable baseline (games/ground/examples/difficulty-baseline.rs), and the measurement produced a finding before any solver exists. A GREEDY BOT WINS 200 OF 200 GAMES AT FIVE AND SIX SEATS. Median margin +3, 11.8-12.0 points available against a threshold of 9. The curve is 66% / 82.5% / 95% / 100% / 100% across 2/3/4/5/6 seats. Row-level, as GROUND-WP-0004 T02 requires: available points 6 / 9 / 12 against thresholds 5 / 7 / 9, so the ratio RISES with seat count (1.20, 1.29, 1.33) while the table also gains actions per round to clear it with. Three multipliers pointing the same direction. It also explains the maintainer's report without needing a solver at all: "I felt it was too easy but then we lost" is two true statements about different seat counts. Cost measured and it rules out the obvious approach. Branching is small (mean 4.7-9.1) but legal_commands costs 112-161 us per call because it filters candidates through full validate. Exhaustive search is out at every seat count; 10^4-10^5 nodes is 1.4-14 seconds, which is the budget the ADR must design inside. Prior art names the trap: determinized search (PIMC) suffers strategy fusion (Frank, Basin & Matsubara 1998) -- the search picks different actions in states a real player cannot distinguish, so the witness may require knowing what was on top of the deck. Such a line still replays green, so the checkability benchmark does not catch it. Honesty and checkability are different properties; stated explicitly so T03 cannot conflate them. The survey states its own most likely killer up front (§6): a view-only search cannot fold events, so making the information boundary structural rather than a promise may not be affordable. Better found here than in T05. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |