Current evidence
Use ntuple-leaf as the reference for a new candidate. It is the latest
policy to meet its registered rule through the authoritative server. Engine
experiments point to a quarter-strength blend as the next candidate; that blend
has not yet passed a protocol cohort.
The protocol baseline
With the same depth-3 search and 1.5-second budget, the learned leaf won ntuple-leaf57.5% (46/80) against expectimax-v2-plan31.3% (25/80). All 80 games on 20 deterministic seeds completed validly without timeout moves. The paired effect was +0.263 wins/game.
Win rate by competitor
95% intervalThis established the baseline for one lineup and one rotation scheme. It does not establish strength against humans or every opponent population. The protocol report includes the registration, complete denominator, and public replay.
The next candidate
The quarter-strength blend adds a small contribution from the hand-written evaluator to the learned tables. Its effect survives both engine lineups: builders and a table with three searches.
A small hand-written contribution helps at both tables
Measured evidenceScroll the chart horizontally to inspect all values.
Only quarter strength has pooled intervals above zero at both tables.
View data table
Source: Source: reports/leaf-blend.mdx, The blend arms; analysis/tables_retest.py. After-the-fact pools of seating-corrected swapped pairs over 192 L1 boards and 128 L2 boards, four rotations in each half, all cohorts 256 of 256 completed. 95% intervals over boards. Full-strength partial/withdrawn runs and unregistered screens are excluded from this figure and retained in the report.
| Series | Comparison | Change in wins per game | Low | High |
|---|---|---|---|---|
| Builders: 192 boards | Quarter strength | 0.066 | 0.024 | 0.107 |
| Builders: 192 boards | Half strength | 0.016 | -0.024 | 0.057 |
| Population: 128 boards | Quarter strength | 0.064 | 0.021 | 0.108 |
| Population: 128 boards | Half strength | -0.019 | -0.064 | 0.027 |
The after-the-fact pools cover 192 builder boards and 128 population boards.
Decision time rises by about 40%. Register a fresh protocol cohort against
ntuple-leaf before changing the server default. The browser still uses the
hand-written leaf; these table-based results do not transfer directly to it.
What the other comparisons changed
| Result | Decision |
|---|---|
| Deeper learned-leaf search: 27 wins against 43 in the protocol cohort | Keep the shallower baseline |
| Transposition cache: identical moves in 11.7% less time at equal nodes | Treat it as a computation saving; the protocol strength interval crosses zero |
| Search-switch retest: 145 swapped pairs, 74,240 games | Keep bargaining, eight dice scenarios, and four sampled worlds; the other tactical switches do not justify a default change |
| Seating diagnosis: identical adjacent searches can differ | Measure two-search comparisons in both seatings; rotating seats alone preserves adjacency |
The experiment timeline links the individual studies. The client guide explains how to implement and run a candidate, and the experiment program explains how to measure it.
Earlier comparisons and complete evidence history
The entries below preserve the sequence of findings. Statements about a new or
retained baseline describe that point in the research history; the current
reference is ntuple-leaf as stated above.
This repository starts with a headless protocol harness, existing Rust fast/ETA opponents, and a minimal Python policy adapter. Setup and smoke validation do not establish a strategy ranking. No claim of optimal play has been made.
Consult records/runs/ for actual run outcomes and reports/ for selected
publications. Promote a finding here only when a fixed candidate, retained
artifacts, complete denominators, and the stated decision rule support it.
The strategy atlas now defines sixteen proposed directions, with animated mechanisms, study briefs, and a shared experiment program. No new strategy evaluation was run to create this guide. Interactive teaching values and synthetic publishing fixtures remain explicitly separate from recorded evidence.
First strategy comparison: retain ETA
The liquidity experiment completed all 80 registered evaluation games in 80-game protocol cohort. Liquidity won 20/80 (25%); ETA control won 26/80 (32.5%). The observed contrast was −7.5 percentage points and the preregistered improvement rule failed. ETA remains the baseline; the result does not prove inherent inferiority or establish general strength across opponent populations.
The candidate changes only the target acquisition-time estimator, with shared
protocol adaptations and a 16-offer budget across all policies. Both native
policies and the comparator source are retained under policies/liquidity/.
The audit preserves the startup failures, rate-limit
failure, stopped first evaluation and all smoke runs. None was pooled into the
fresh cohort. All archives are verified local copies.
Engine arena and the tunable search
An offline engine arena now plays paired-seed games with only redacted
observations reaching each seat. Its calibration
reproduces the protocol ordering of the two builders over 512 slot-games and
shows the frozen expectimax-v1 reference losing to both (13.3% against 30.1%).
Engine-arena results are development-tier evidence; they do not replace a
protocol cohort.
expectimax-v2, a tunable turn-level search with an acquisition-time leaf,
won 34.0% of 256 paired slot-games against ETA's 24.2% at depth 2 (paired
contrast +0.098, 95% interval +0.005 to +0.190). See the
development report. The frozen
depth-3 candidate then won 38.8% of an 80-game protocol cohort against the
compared ETA slot's 31.2%, short of its registered rule, so the protocol-arena
baseline remains ETA and that candidate's advantage is not confirmed.
Those cohorts ran with a turn-planning defect. After the correction, depth 2 wins 50.4% of 256 paired slot-games against ETA's 20.3% (paired contrast +0.301) and depth 3 under 1.5 s wins 57.4% (contrast +0.414).
Deterministic mode and ablations
Both tiers can now play deterministic paired cohorts: the k-th roll, theft, and development draw come from seed-derived tapes, so two policies on one seed face identical luck. A protocol smoke confirmed identical rolls across rotations through the server. An ablation program of eighteen registered cohorts (one change per arm against an unchanged depth-2 control, 256 deterministic games each) found two decisive levers, lookahead (−0.324 wins per game when removed) and dice scenarios (−0.410 with two instead of eight), and no single leaf term or style slider worth more than 0.1 wins per game on 64 seeds. On 192 fresh seeds one belief mechanism held up: sampling hidden hands from the event history rather than public counts alone is worth about 0.07 wins per game; the expansion terms, sampled-world count, trade margin, and plan term showed no measurable effect at depth 2.
Two searches at one table
The one-change arm design seats the changed search in slot 0 and the control in slot 1, and slot 1 always sits one seat after slot 0. That order matters: two identical default searches differ by +0.176 wins per game (95% interval +0.083 to +0.268) on boards 0 to 63 (256-game engine cohort) and by +0.051 (−0.046 to +0.148) on boards 64 to 127 (256-game engine cohort). The cause is the game's turn order, not the arena's trade handling: the seat directly before another acts first after three of the four rolls in a round, the term survives every trade switch and identical builders, and it disappears with a builder between the two searches (seating report). The bargaining arms report reads its arms against the first null and confirms the counters effect with a swapped pair on fresh boards, −0.096 (−0.156 to −0.035) for the seat without counters. Any contrast between two searches at one table must be measured in both seatings before it counts as evidence; the +0.219 for threat-aware bargaining over plain acceptance (256-game engine cohort) is withdrawn as a strength claim, corrected estimate +0.043 (−0.082 to +0.168) on the same boards.
The build-prediction study corroborates the seating term with a different lineup: two identical depth-2 searches with ETA and fast split +0.172 wins per game (95% interval +0.083 to +0.261) for slot 0 on seeds 0 to 63, and evenly on seeds 64 to 79. The single-order ablation cohorts on the development seeds therefore lean positive by about that much, and their near-zero arms may be slightly negative. That study ran its three prediction switches as mirrored pairs in both seatings; none was confirmed on fresh seeds.
A learned leaf beats the hand-written leaf
Lookup tables of hex and portfolio patterns, learned by self-play in under two
hours (patterns/ntuple-leaf, report), replace
the leaf of expectimax-v2 with nothing else changed. On 64 fresh deterministic
boards the tables at depth 2 won 130 of 256 slot-games against 74 for the
hand-written leaf in the same search (paired contrast +0.219, 95% interval
+0.124 to +0.314); played one ply they beat the two-turn hand-written search,
at depth 2 they beat the hand-written leaf at depth 3, and at depth 3 against
depth 3 the margin is +0.242. This is engine-arena evidence on one lineup with
two fixed builders; the tables are a server-side policy (ntuple-leaf) and the
browser keeps the hand-written leaf.
The protocol-arena baseline is now ntuple-leaf
Through the authoritative server, ntuple-leaf (the expectimax-v2-plan
search with the learned tables as its leaf) won 46 of 80
games against 25 for expectimax-v2-plan
on 20 deterministic seeds played once per rotation, a paired contrast of +0.263
wins per game (95% interval +0.094 to +0.431), all games valid. The registered
rule is met. New candidates should be compared against ntuple-leaf in a
four-player table with the fixed builders. As before, the comparison is one
lineup and one seat rotation scheme; strength against other searches, humans,
or negotiation-aware opponents is not established.
A deeper configuration of the same player, ntuple-leaf-deep (depth cap 4
under a 4 s budget with decision-aware deepening), lost its protocol
cohort against ntuple-leaf 27 to 43 (paired
contrast −0.200, 95% interval −0.357 to −0.043), so it is not the baseline.
The search's transposition tables were rebuilt with a verified key and a
per-decision or per-turn lifetime (report):
identical decisions from 11.7% less time at equal nodes, and no strength
change at equal time or through the protocol (27 to 36, interval crossing
zero), so ntuple-leaf remains the baseline and the tables are a cost saving.
Every search switch under the learned leaf
Nine engine-arena studies retested the search's switches with the learned
tables as the leaf, as seating-corrected swapped pairs at a table of builders
and a table of three searches (145 pairs, 74,240 games; synthesis).
Bargaining and the search's information inputs hold; the one improvement is the
hand-written terms blended into the tables at a quarter of their weight, +0.066
wins per game at the builder table and +0.064 at the table of searches, at 1.4
times the decision time. This is development-tier evidence: the protocol-arena
baseline stays ntuple-leaf until the blend wins a protocol cohort, and the
browser keeps the hand-written leaf.
The previous protocol-arena baseline: expectimax-v2-plan
The corrected depth-3 candidate, frozen as expectimax-v2-plan, won 47 of 80
authoritative games against 11 for
the compared ETA slot, with all 80 games valid; the registered rule (a 10-point
gap at p ≤ 0.05) is met at p = 1.0 × 10⁻⁶. New candidates should be compared
against this policy in a four-player table with the fixed builders. The
comparison is one lineup and one seat rotation scheme; strength against other
searches, humans, or negotiation-aware opponents is not established.