Settlers / Research

75 pages · Search titles and descriptions

↑ ↓ to navigate · Enter to open · Esc to closeLocal search
Play the game

Blending the learned leaf, and the endgame switches under it

A small hand-written contribution helps at both tables

Measured evidence
-0.09-0.0350.020.0750.13Quarter strengthHalf strengthComparisonChange in wins per gameNo difference

Scroll the chart horizontally to inspect all values.

Builders: 192 boardsPopulation: 128 boards95% intervalNo difference

Only quarter strength has pooled intervals above zero at both tables.

View data table

Source: Source: reports/leaf-blend.mdx, The blend arms; analysis/tables_retest.py. After-the-fact pools of seating-corrected swapped pairs over 192 L1 boards and 128 L2 boards, four rotations in each half, all cohorts 256 of 256 completed. 95% intervals over boards. Full-strength partial/withdrawn runs and unregistered screens are excluded from this figure and retained in the report.

SeriesComparisonChange in wins per gameLowHigh
Builders: 192 boardsQuarter strength0.0660.0240.107
Builders: 192 boardsHalf strength0.016-0.0240.057
Population: 128 boardsQuarter strength0.0640.0210.108
Population: 128 boardsHalf strength-0.019-0.0640.027

The arms

The learned tables won their cohort as the only leaf (the learned leaf report); the search can also add the hand-written terms on top of them. LeafConfig exposes two knobs for that. leaf.hand scales the hand-written terms that are added to the tables' value (0 by default), and leaf.scale sets how many hand-written points one win is worth (1000 by default, so a win is ten points). The scale does not change the search's preference between moves; it changes how the tables' win probability compares with thresholds that are fixed in points: the trade margin and threat premium in bargaining, the deepen_gap stop rule, and the hand terms when they are blended.

ArmSeat changeWhat it tests
hand 0.25 / 0.5 / 1.0leaf.hand set on the tables seatDo the hand-written terms still carry information the tables missed, and at what weight do they stop helping
scale 500 / 2000leaf.scale halved or doubled on the tables seatDo the fixed thresholds that compare against the leaf value sit at the right win probability
raceendgame.race on top of the hand 0.25 blendThe race leaf from the endgame report, inert under the tables alone, measured on the blend that may become the default
hidden pointsendgame.hidden_points on top of the hand 0.25 blendExpected hidden victory points, same condition

Both endgame switches are hand-leaf terms, so they are inert under the tables alone; they were measured with leaf.hand 0.25 in both the candidate and the control seat. They were first registered on the hand 1.0 blend; when an unregistered screen showed the full blend much weaker than the tables alone, the unstarted hand 1.0 registrations were withdrawn and the arms were re-registered on the 0.25 blend, which is the configuration a switch must justify itself on. One full-blend pair (hidden points) had already run and is reported below as what it is.

The design

Every contrast is a swapped pair on the same seeds: half A seats the candidate in slot 0, half B seats it in slot 1, and the effect is the per-seed half-difference of the two slot-0-minus-slot-1 contrasts, which cancels the arena's seating term (the seating report); half their sum is the term. The builder lineup L1 puts eta and fast in slots 2 and 3, the population lineup L2 puts the tables baseline in slot 2 and eta in slot 3, so trades, blocks, and robber choices come from opponents that trade and block. Stages: development on seeds 0 to 63, confirmation on 800 to 863 (never used before), population on 800 to 863 in L2, and, when the pooled L1 effect stayed positive with its interval within 0.03 of zero, an extension on 864 to 927 in both lineups. Every cohort is 64 seeds, 4 rotations, 256 deterministic games, and every one completed all 256.

The stage 1 pair of the hand 0.25 arm reproduces the orchestrator's unregistered screen exactly (+0.074, +0.011 to +0.138): the same seats on the same deterministic boards are the same games, so the registered pair confirms the screen's arithmetic rather than independently replicating it. The fresh seeds of the later stages are the test.

The blend arms

Seating-corrected effect per stage, wins per game; pooled rows are named after-the-fact pooling.

ArmStage 1 (0-63, L1)Stage 2 (800-863, L1)Stage 3 (800-863, L2)Extension (864-927)Pooled L1Pooled L2
hand 0.25+0.074 (+0.011 to +0.138)+0.035 (−0.034 to +0.105)+0.049 (−0.008 to +0.105)L1 +0.088 (+0.006 to +0.170), L2 +0.080 (+0.015 to +0.146)+0.066 (+0.024 to +0.107)+0.064 (+0.021 to +0.108)
hand 0.5+0.045 (−0.028 to +0.118)+0.051 (−0.018 to +0.120)−0.027 (−0.094 to +0.040)L1 −0.047 (−0.113 to +0.019), L2 −0.010 (−0.072 to +0.053)+0.016 (−0.024 to +0.057)−0.019 (−0.064 to +0.027)
hand 1.0half A only, −0.117 (−0.233 to −0.002); pair withdrawn, screen −0.207 (−0.284 to −0.130)withdrawnwithdrawn

The quarter-strength blend is the only arm whose pooled intervals exclude zero, in both lineups. Its stage 2 number (+0.035) crosses zero, and the extension seeds came back at +0.088, so the pooled L1 estimate rests on 192 seeds and the pooled L2 on 128. The dose-response is monotone in the negative direction: the full blend loses to the tables even from the seat that wins the seating term, half strength is inside the noise, and a quarter is a real improvement.

The scale arms

ArmStage 1 (0-63, L1)Stage 2 (800-863, L1)Stage 3 (800-863, L2)Extension (864-927)Pooled
scale 500−0.059 (−0.121 to +0.004)−0.135 (−0.191 to −0.078), refutedwithdrawn after the refutation−0.097 (−0.139 to −0.054) over 128 seeds
scale 2000+0.051 (−0.009 to +0.111)−0.010 (−0.069 to +0.050)+0.035 (−0.024 to +0.095)L1 +0.047 (−0.026 to +0.120), L2 +0.033 (−0.034 to +0.101)+0.029 (−0.008 to +0.066) over 192 L1 seeds; +0.034 (−0.011 to +0.079) over 128 L2 seeds

Halving the scale makes every fixed threshold twice as large in win probability, and the arm shows it in play: the halved seat completed 2.3 player trades per game against the control's 3.7 and lost ground everywhere. Doubling the scale makes the thresholds relatively cheaper, the doubled seat traded more (5.1 against 4.1) and bought more development cards (4.07 against 3.88), and the effect on wins never separated from zero in six pairs. The default scale of 1000 is on the right side of both changes; the doubling is free in decision time but buys nothing.

The endgame arms

Both were measured on the hand 0.25 blend in both seats, with the blend alone as the control.

ArmStage 1 (0-63, L1)Stage 2 (800-863, L1)Stage 3 (800-863, L2)Extension (864-927)Pooled L1Under the hand-written leaf
race−0.039 (−0.076 to −0.002)−0.035 (−0.080 to +0.010)+0.000 (−0.036 to +0.036)not triggered, pooled L1 negative−0.037 (−0.066 to −0.008)+0.018 (−0.018 to +0.053)
hidden points+0.025 (−0.010 to +0.060)−0.010 (−0.043 to +0.024)−0.025 (−0.067 to +0.016)L1 −0.006 (−0.042 to +0.031), L2 +0.008 (−0.029 to +0.044)+0.003 (−0.017 to +0.023)+0.006 (−0.030 to +0.041)

The race leaf, which replaced the opponent terms with the own-minus-leader margin and the leader's expected points, is a small loss under the blend: its stage 1 interval excludes zero on the wrong side and the pooled L1 estimate does too, so the registered rule reads refutation. Under the hand-written leaf the same switch read +0.018 with an interval crossing zero (the endgame report), so the mechanism does nothing there and slightly hurts the blend. Hidden points is the same story at smaller size: nothing under the hand-written leaf, and nothing under the blend either: its extension was triggered by the letter of the rule (pooled L1 +0.008 after two stages) and the extension seeds brought the pooled 192-seed L1 estimate down to +0.003, with the population lineup slightly negative. The hidden-points pair on the full hand 1.0 blend, which ran before the withdrawal, read −0.029 (−0.075 to +0.016), also nothing.

What the blend changes at the table

Per seat per game over both halves of stage 1 (512 games each side):

ArmCitiesDevelopment cardsSettlementsRoadsPlayer tradesDiscardsMean pointsMean decision
tables (control of the blend arms)1.234.031.815.373.885.207.47667 ms
hand 0.251.503.952.055.454.295.927.95929 ms
hand 0.51.623.742.085.734.707.628.10880 ms
hand 1.01.863.342.126.195.689.028.11895 ms

The hand-written terms push toward cities and settlements, as expected: at a quarter strength the blend builds a quarter more cities and settles more without giving up development cards; at half strength it starts trading development cards for cities; at full strength it abandons the cards the tables were winning with (3.34 against 4.27 for its control) and carries two discards more per game, and the points it gains do not become wins. The quarter-strength blend keeps the tables' style and corrects its build choices.

Decision time is the blend's cost. Compared inside each cohort, where both seats shared the same machine load, the hand 0.25 seat spent 929 ms per decision against its control's 667 ms, about 40 percent more, and the hand 0.5 seat about the same. The two evaluators are both evaluated at every leaf. The scale arms cost nothing (the doubled-scale seat even ran marginally faster, 573 ms against 599 ms, inside noise). The endgame switches cost at most a few percent, and only because the race leaf's tempo estimate runs at the leaf.

Every cohort

All cohorts are 256 deterministic games; every one completed. "Contrast" is the registered slot-0-minus-slot-1 paired contrast of that half; the seating term is half the sum of a pair's two contrasts.

Full results and cohort ledger
ArmStageHalfCandidate wins / control winsContrast (95% interval)ExperimentRun
hand 0.251 (0-63, L1)A138 / 87+0.199 (+0.107 to +0.291)registered protocol256-game engine cohort
hand 0.251 (0-63, L1)B109 / 122+0.051 (-0.051 to +0.153) (control minus candidate)registered protocol256-game engine cohort
hand 0.252 (800-863, L1)A120 / 102+0.070 (-0.032 to +0.173)registered protocol256-game engine cohort
hand 0.252 (800-863, L1)B116 / 116+0.000 (-0.096 to +0.096) (control minus candidate)registered protocol256-game engine cohort
hand 0.253 (800-863, L2)A92 / 74+0.070 (-0.012 to +0.153)registered protocol256-game engine cohort
hand 0.253 (800-863, L2)B83 / 76-0.027 (-0.115 to +0.061) (control minus candidate)registered protocol256-game engine cohort
hand 0.25ext L1 (864-927)A141 / 98+0.168 (+0.047 to +0.289)registered protocol256-game engine cohort
hand 0.25ext L1 (864-927)B120 / 118-0.008 (-0.119 to +0.103) (control minus candidate)registered protocol256-game engine cohort
hand 0.25ext L2 (864-927)A107 / 75+0.125 (+0.031 to +0.219)registered protocol256-game engine cohort
hand 0.25ext L2 (864-927)B84 / 75-0.035 (-0.124 to +0.054) (control minus candidate)registered protocol256-game engine cohort
hand 0.51 (0-63, L1)A133 / 95+0.148 (+0.046 to +0.251)registered protocol256-game engine cohort
hand 0.51 (0-63, L1)B104 / 119+0.059 (-0.045 to +0.163) (control minus candidate)registered protocol256-game engine cohort
hand 0.52 (800-863, L1)A128 / 103+0.098 (-0.010 to +0.205)registered protocol256-game engine cohort
hand 0.52 (800-863, L1)B116 / 115-0.004 (-0.115 to +0.107) (control minus candidate)registered protocol256-game engine cohort
hand 0.53 (800-863, L2)A87 / 83+0.016 (-0.079 to +0.110)registered protocol256-game engine cohort
hand 0.53 (800-863, L2)B77 / 95+0.070 (-0.029 to +0.169) (control minus candidate)registered protocol256-game engine cohort
hand 0.5ext L1 (864-927)A120 / 114+0.023 (-0.084 to +0.131)registered protocol256-game engine cohort
hand 0.5ext L1 (864-927)B105 / 135+0.117 (+0.002 to +0.233) (control minus candidate)registered protocol256-game engine cohort
hand 0.5ext L2 (864-927)A86 / 80+0.023 (-0.068 to +0.115)registered protocol256-game engine cohort
hand 0.5ext L2 (864-927)B79 / 90+0.043 (-0.051 to +0.137) (control minus candidate)registered protocol256-game engine cohort
hand 1.01 (0-63, L1), half A onlyA93 / 123-0.117 (-0.233 to -0.002)registered protocol256-game engine cohort
scale 5001 (0-63, L1)A117 / 108+0.035 (-0.046 to +0.116)registered protocol256-game engine cohort
scale 5001 (0-63, L1)B91 / 130+0.152 (+0.052 to +0.253) (control minus candidate)registered protocol256-game engine cohort
scale 5002 (800-863, L1)A107 / 117-0.039 (-0.143 to +0.065)registered protocol256-game engine cohort
scale 5002 (800-863, L1)B87 / 146+0.230 (+0.147 to +0.314) (control minus candidate)registered protocol256-game engine cohort
scale 20001 (0-63, L1)A141 / 94+0.184 (+0.089 to +0.278)registered protocol256-game engine cohort
scale 20001 (0-63, L1)B101 / 122+0.082 (-0.016 to +0.180) (control minus candidate)registered protocol256-game engine cohort
scale 20002 (800-863, L1)A118 / 108+0.039 (-0.069 to +0.147)registered protocol256-game engine cohort
scale 20002 (800-863, L1)B111 / 126+0.059 (-0.056 to +0.173) (control minus candidate)registered protocol256-game engine cohort
scale 20003 (800-863, L2)A87 / 85+0.008 (-0.089 to +0.104)registered protocol256-game engine cohort
scale 20003 (800-863, L2)B91 / 75-0.062 (-0.148 to +0.023) (control minus candidate)registered protocol256-game engine cohort
scale 2000ext L1 (864-927)A135 / 100+0.137 (+0.029 to +0.244)registered protocol256-game engine cohort
scale 2000ext L1 (864-927)B110 / 121+0.043 (-0.057 to +0.143) (control minus candidate)registered protocol256-game engine cohort
scale 2000ext L2 (864-927)A100 / 76+0.094 (+0.001 to +0.187)registered protocol256-game engine cohort
scale 2000ext L2 (864-927)B83 / 90+0.027 (-0.072 to +0.126) (control minus candidate)registered protocol256-game engine cohort
race1 (0-63, L1)A129 / 100+0.113 (+0.024 to +0.203)registered protocol256-game engine cohort
race1 (0-63, L1)B92 / 141+0.191 (+0.099 to +0.284) (control minus candidate)registered protocol256-game engine cohort
race2 (800-863, L1)A123 / 104+0.074 (-0.037 to +0.186)registered protocol256-game engine cohort
race2 (800-863, L1)B95 / 132+0.145 (+0.035 to +0.254) (control minus candidate)registered protocol256-game engine cohort
race3 (800-863, L2)A84 / 83+0.004 (-0.086 to +0.094)registered protocol256-game engine cohort
race3 (800-863, L2)B81 / 82+0.004 (-0.075 to +0.083) (control minus candidate)registered protocol256-game engine cohort
hidden points1 (0-63, L1)A137 / 97+0.156 (+0.064 to +0.249)registered protocol256-game engine cohort
hidden points1 (0-63, L1)B102 / 129+0.105 (+0.012 to +0.199) (control minus candidate)registered protocol256-game engine cohort
hidden points2 (800-863, L1)A125 / 104+0.082 (-0.025 to +0.189)registered protocol256-game engine cohort
hidden points2 (800-863, L1)B102 / 128+0.102 (-0.009 to +0.212) (control minus candidate)registered protocol256-game engine cohort
hidden points3 (800-863, L2)A83 / 84-0.004 (-0.092 to +0.084)registered protocol256-game engine cohort
hidden points3 (800-863, L2)B74 / 86+0.047 (-0.044 to +0.138) (control minus candidate)registered protocol256-game engine cohort
hidden pointsext L1 (864-927)A130 / 107+0.090 (-0.011 to +0.191)registered protocol256-game engine cohort
hidden pointsext L1 (864-927)B103 / 129+0.102 (-0.003 to +0.206) (control minus candidate)registered protocol256-game engine cohort
hidden pointsext L2 (864-927)A91 / 82+0.035 (-0.046 to +0.117)registered protocol256-game engine cohort
hidden pointsext L2 (864-927)B85 / 90+0.020 (-0.073 to +0.112) (control minus candidate)registered protocol256-game engine cohort
hidden points1 on hand 1.0 (0-63, L1)A95 / 104-0.035 (-0.148 to +0.078)registered protocol256-game engine cohort
hidden points1 on hand 1.0 (0-63, L1)B101 / 107+0.023 (-0.101 to +0.148) (control minus candidate)registered protocol256-game engine cohort

Withdrawn before running, with reasons in the daily log: the hand 1.0 half B (registered protocol) and the race halves on the full blend (registered protocol, registered protocol), because the full blend measured much weaker than the tables alone and a switch on top of it says nothing about a configuration anyone would use; and the scale 500 stage 3 halves (registered protocol, registered protocol), because stage 2 refuted the arm.

What could still be wrong

Everything here is engine-arena evidence at depth 2 on deterministic boards: two fixed builders for L1, one extra search and one builder for L2. No protocol cohort has confirmed the quarter-strength blend, and nothing tests it at depth 3, against negotiating opponents, or with the tables the search would train next. The blend costs about 40 percent more decision time under machine load; the wall-clock numbers are distorted by a machine shared with other agents, and only the within-cohort ratios are meaningful. The 0.25 weight sits between two tested neighbours and was not itself tuned, so it is a coarse reading of where the optimum is; a cheaper subset of the hand terms might buy the same gain for less time. The extension rule was applied by its letter to three arms, and the pooled rows are after-the-fact poolings across stages that were registered separately. The stage 1 pair of hand 0.25 is the same games as the orchestrator's screen by construction.

What this means for the browser

The browser build runs the hand-written leaf in WASM at depth 3 under a one second budget and cannot load the tables, so leaf.hand and leaf.scale are not browser switches; they configure the server-side ntuple-leaf policy, and the quarter-strength blend is a candidate for that policy, worth a protocol cohort against ntuple-leaf before any default changes. The endgame race and hidden-points switches do run under the browser's hand-written leaf, and each now has two swapped-pair readings: under the hand-written leaf +0.018 and +0.006 (seeds 64 to 127), under the tables blend −0.037 pooled over 128 L1 seeds (negative) and +0.003 pooled over 192, all with intervals crossing zero except the race reading under the blend, which is negative. Neither switch reaches the +0.03 wins per game bar under either leaf, so neither is worth enabling in the browser; both stay off. The blend result does say the hand-written terms carry complementary information next to a learned evaluator, which is a reason to keep them sharp in the browser's own leaf, not a change to make today.

Methods and reproduction

Study: patterns/leaf-blend under Learned pattern values. All arms are seat specs on the unchanged expectimax-v2 search: the baseline is v2:{"depth":2,"leaf":{"tables":"/home/keshav/settlers/research/artifacts/ntuple/hex-portfolio-main.bin"}}, the blend arms add "hand":0.25 (or 0.5, 1.0) or replace the scale with "scale":500 (or 2000) inside the same leaf object, and the endgame arms add "endgame":{"race":true} or "endgame":{"hidden_points":true} on top of the 0.25 blend in both search seats. The knobs are documented in the server's docs/expectimax.md under "Learned n-tuple leaf" and "Endgame"; no server code changed in this study.

Reproduction: python3 -m harness.engine run EXPERIMENT --threads 6 in the research checkout with SETTLERS_SERVER_DIR pointing at a server worktree at dev 38ae537. Effects and seating terms: analysis/seating_pair.py HALF_A_HALF_A HALF_B, per-arm diagnostics and pooled estimates: analysis/leaf_blend.py --pair LABEL A B. Lineup L1 seats eta and fast in slots 2 and 3; L2 seats the tables baseline in slot 2 and eta in slot 3. Every run directory retains its manifest, games.jsonl, summary, and inventory under runs/, and a compact record under records/runs/.

All studies · Learned pattern values · Log 2026-09-11

Detailed result and cohort context

The learned tables price end-of-turn positions in win probability, and the hand-written terms price them in points. Added on top of the tables, the hand-written terms at quarter strength (leaf.hand 0.25) beat the tables alone by +0.074 wins per game (95% interval +0.011 to +0.138) on seeds 0 to 63, +0.088 (+0.006 to +0.170) on the extension seeds 864 to 927, and +0.080 (+0.015 to +0.146) there at a three-search population table; pooled after-the-fact over 192 L1 seeds the effect is +0.066 (+0.024 to +0.107) and over 128 L2 seeds +0.064 (+0.021 to +0.108). Half strength is weaker (+0.016 over 192 L1 seeds, interval crossing zero), and full strength is much worse than the tables alone: the withdrawn full-blend half lost from the favorable seat (−0.117) and an unregistered screen of the pair read −0.207. Halving the tables scale to 500 is refuted (−0.135 on fresh seeds, −0.097 pooled), doubling it to 2000 is a cost-free nothing (+0.029 pooled over 192 L1 seeds, crossing zero). The endgame race leaf is −0.037 (−0.066 to −0.008) pooled over its two L1 stages under the blend, and expected hidden points reads +0.003 (−0.017 to +0.023) pooled over its 192 L1 seeds and −0.009 (−0.037 to +0.019) over its 128 L2 seeds, every interval crossing zero. The race and hidden-points switches stay off; the quarter-strength blend is the one candidate worth a protocol cohort.