The arena and notebook come online
The first task was to make experiments easy to run and possible to audit. The harness now launches independent player processes, rotates seats, captures public events, and retains compact run records. The notebook renders those records and selected game exhibits through components from the design Storybook.
A restricted local connection prevented the first preflight. Its invalid record is invalid attempt. The next attempt, invalid attempt, exposed a cleanup failure when signaling an already exited policy group. A terminal snapshot exists, but finalization failed; that attempt is not valid match evidence. Its early summary reports no finalized matches, not proof that no game was created. The cleanup now checks process state and preserves results even if cleanup itself fails.
After the correction, 4-game protocol cohort completed four short three-point Rust games. 4-game protocol cohort completed four short games with the Python legal-first policy and Rust opponents. These verify the two policy transports; neither run establishes a strength ranking.
4-game protocol cohort completed four normal ten-point games. The first arena report contains its recorded win rates and a selected public replay. A hash-verified archive exists locally; there is no cloud backup or public artifact URL yet.
This was infrastructure and publishing work, not a new strategic research contribution. The next useful experiment is a preregistered comparison with a larger fixed cohort and a specific change to the value of a build.
A final two-player check, 2-game protocol cohort, completed both short games after tightening the policy I/O deadline. Its archive inventory also retains the exact harness and policy sources for dirty working trees.
A visual program for a well-rounded player
Added sixteen proposed strategy directions covering production, build timing, hidden-information search, negotiation, truthful and deceptive gameplay claims, reputation, bounded pliability, leader containment, and adaptation. Each brief has an animated SVG and a falsifiable comparison. A shared experiment program defines draft lineups, budgets, outcome measures, and decision gates; no study was registered or executed as part of this work.
The core hypothesis to test is that modelling future social response, or reranking inside a fixed foregone-value budget, can improve play in some table conditions. The guide distinguishes those operations from updating beliefs and from merely deciding when to consult a language model. It does not conclude that concessions or deception improve wins.
The interactive laboratory uses declared toy utilities. Publishing examples use synthetic data; the dice exhibit uses exact enumeration. No new run IDs or empirical failures exist for this planning task, and the earlier evidence boundary is unchanged. Attribution: Codex; exact model identifier unknown. This entry records program design and publishing work, not a completed strategic result.
First policy ablation: exchange-aware ETA
Hypothesis: counting whole bank/owned-port exchange bundles in the ETA build
planner increases win probability relative to the unchanged ETA target estimator.
Implemented liquidity and same-source eta-control/fast-control opponents in
an isolated research worktree. No language model selected game moves; the native
processes received only their own participant views.
Fresh evaluation registered protocol, run 80-game protocol cohort, completed 80/80 valid games. Liquidity won 20, ETA won 26; the observed difference was −7.5 percentage points. The registered +10-point/one-sided-p≤0.05 promotion rule failed (p=0.849). Retain ETA as baseline; do not claim the richer estimator is a general improvement or proven inferior.
The visual report includes the mechanism, Wilson chart, exchange CDF, final-score heatmap and first scheduled game replay. An exploratory whole-game bootstrap spans −23.75 to +8.75 percentage points. Liquidity exchanged and discarded more on average, but those descriptive outcomes do not identify the causal failure. A decision trace is the next useful diagnostic.
The complete audit retains four readiness failures, a partial smoke stopped by the communication limit, successful smokes, an excluded smoke after a failed build used the prior binary, and the first evaluation invalid attempt stopped by archive backpressure (2 complete, 1 invalid, 77 unplayed). Shared readiness, offer-budget and idempotent retry fixes preceded the fresh cohort; the strategic estimator was not tuned to these outcomes. Guarded build recipes prevent another stale-binary launch. Sources and configured binary hashes are frozen; public tapes now exclude research-only event extras.
All ten run archives have verified local receipts in records/artifacts/ and
bundles in artifacts/liquidity-study/. No remote backup or deployment occurred.
A reusable report template lists missing contrast-interval, decision-trace,
protocol-health and cohort/population views. Attribution: Codex; exact model
identifier unknown.
A tunable expectimax and an offline engine arena
Goal: a strong, tunable expectimax player that compiles to the browser, plus
structural measurements of the game tree. The frozen reference expectimax-v1
stays as the comparator. Its search and leaf were replaced in a second engine,
expectimax-v2, in server/crates/expectimax/src/v2: turn-level afterstate
enumeration with information-set transpositions, exact chance for purchases,
thefts and the own roll, common random dice scenarios for the other seats'
turns, beam expansion at inner levels, and iterative deepening under a time
budget and a depth cap. The leaf values points, cards, production, build targets
by acquisition time, discard risk, roads, ports and opponent progress in one
card-value currency; every weight is tunable and five inclination sliders scale
weight groups.
To evaluate quickly, server/crates/arena plays complete games offline with
explicit seeds. Seats receive only their redacted observation and participant
events. The experiment program now defines this
engine arena as the development tier and the protocol arena as the confirmation
tier; just engine-register and just engine-run file and execute engine
experiments with the same study ownership as protocol runs.
Calibration 512-game engine cohort (512 paired builder games) ranks ETA above fast, matching the protocol arena. Run 256-game engine cohort shows the reference losing to both builders (13.3% against 30.1%; paired contrast −0.168, interval −0.251 to −0.085). Both are in the calibration report.
Developing the v2 leaf took three unregistered exploratory rounds on 8 to 16 seeds, retained only in a scratch directory and carrying no evidential weight. The first leaf bought 5.7 development cards per game and built 0.4 cities: a hidden victory card counted as a full point while production was worth almost nothing. Pricing production as cards over the remaining rolls did not change that. Adding the builders' acquisition-time idea, a target valued by its gain divided by one plus the turns until affordable, fixed it: cities rose to 1.95 per game and mean points matched ETA. Registered run 256-game engine cohort then showed v2 at depth 2 winning 34.0% of 256 slot-games against ETA's 24.2%, paired contrast +0.098 (interval +0.005 to +0.190), at 25 ms per decision. 256-game engine cohort (depth 3, 1.5 s budget, 16 scenarios, 8 worlds) won 32.0% with a paired contrast of +0.078 (interval −0.022 to +0.178): the registered rule was not met and depth 3 is not distinguishable from depth 2 on these boards, although the unregistered 16-seed sweep had suggested 42%. Details are in the development report. A frozen depth-3 candidate is registered as protocol experiment registered protocol. Its first run, invalid attempt, is invalid: no game started because the harness formats registry arguments with placeholder braces and the candidate's inline JSON configuration was read as a placeholder. The registry now doubles those braces; the candidate itself is unchanged. All twelve attempted matches are retained as invalid. The rerun, 80-game protocol cohort, completed all 80 games with twelve matches in flight and no server timeout moves: the candidate won 31 (38.8%) and the compared ETA slot 25 (31.2%), a 7.5-point gap with a one-sided exact test at p = 0.25. The registered rule required 10 points and p ≤ 0.05, so the protocol cohort is inconclusive: the direction favours the candidate, the evidence does not establish it.
Protocol infrastructure: smoke censored attempt was
censored before any game started because the server's remote runner never
sent the readiness handshake the current server requires. The runner now sends
it, retries the same envelope on backpressure, and can seat v1 or v2; smoke
2-game protocol cohort completed two games, and
4-game protocol cohort completed four games with four matches
in flight at once using the new --parallel option and an expectimax-v2 seat.
These smokes check plumbing only.
Attribution: Claude Fable 5.1 (claude-fable-5-1) through Claude Code.