Settlers / Research

75 pages · Search titles and descriptions

↑ ↓ to navigate · Enter to open · Esc to closeLocal search
Play the game

Robber targeting under the learned leaf

Robber-switch effects under the learned leaf

Measured evidence
-0.1-0.0500.050.1leader, buildersleader, populationneed, buildersneed, populationthreat, buildersthreat, populationblock leader, buildersblock leader, populationneed + block leader, buildersneed + block leader, populationArm and lineupWins per game, arm minus controlno difference

Scroll the chart horizontally to inspect all values.

Seating-corrected effect (95% interval)95% intervalno difference

Wins per game of each arm over the unchanged tables-leaf control, pooled over its stages in each lineup; the builders lineup seats eta and fast, the population lineup a third search.

View data table

Source: Swapped-pair engine-arena cohorts under containment/robber-tables: each arm against the unchanged tables-leaf control, half A with the arm in slot 0 and half B with it in slot 1, both halves on the same 64 deterministic seeds per stage. The point is the pooled seating-corrected effect, half the per-seed difference of the two contrasts, over every stage of the arm in that lineup; run IDs are listed in the robber-tables report.

SeriesArm and lineupWins per game, arm minus controlLowHigh
Seating-corrected effect (95% interval)leader, builders0.0068-0.03490.0486
Seating-corrected effect (95% interval)leader, population0.0312-0.030.0925
Seating-corrected effect (95% interval)need, builders0.0163-0.00830.0408
Seating-corrected effect (95% interval)need, population0.0088-0.01810.0357
Seating-corrected effect (95% interval)threat, builders-0.0049-0.04240.0326
Seating-corrected effect (95% interval)threat, population0.043-0.00470.0906
Seating-corrected effect (95% interval)block leader, builders0.0111-0.00520.0273
Seating-corrected effect (95% interval)block leader, population0.0146-0.00690.0362
Seating-corrected effect (95% interval)need + block leader, builders0.0352-0.0150.0853
Seating-corrected effect (95% interval)need + block leader, population-0.0391-0.09020.0121

The arms and the seats

The baseline seat is the depth-2 search evaluating with the learned tables (ntuple-leaf, artifact hex-portfolio-main.bin). Each arm is that seat plus exactly one switch, and the control is the unchanged baseline. All four switches change candidate generation, so they are active under the tables alone: no arm needed the blended leaf.

ArmChangeEarlier reading under the hand-written leaf
robber.leaderThe hex score counts the points leader's blocked pips in full, every other seat's at a tenth−0.023 (−0.073 to +0.026) on seeds 64-127
robber.needThe victim bonus replaces hand size with 2.8 times the expected share of the resource this seat most needs−0.045 (−0.086 to −0.004), the one refuted arm
robber.threatBlocked pips weighted by 1 + 2 times the blocked seat's bargain threat, low-threat victims spared−0.033 (−0.084 to +0.018)
endgame.block_leaderAfter two thirds of the estimated game length, robber placements and blocking roads weight the leader+0.012 (−0.018 to +0.041)

The hand-written-leaf numbers above are the swapped-pair corrections from containment/robber-targeting-bias and containment/endgame-seating; the single-seating contrasts that first made these arms look strong were the arena's seating term.

Design

Every contrast is a swapped pair on the same seeds: half A seats the arm in slot 0 with the control in slot 1, half B the reverse, each half a separately registered experiment of 64 deterministic seeds played once per rotation (256 games). The seating-corrected effect is half the per-seed difference of the two slot-0-minus-slot-1 contrasts, the seating term half their sum (analysis/seating_pair.py). Two lineups: builders (slots 2 and 3 are eta and fast, matching the earlier hand-leaf cohorts) and population (slot 2 is a third tables search, slot 3 eta, so robber choices are answered by opponents that trade and block). Three stages ran for every arm: development on seeds 0-63 (builders), confirmation on fresh seeds 800-863 (builders), and population on seeds 800-863. The preregistered extension rule then added seeds 864-927 in both lineups for the arms whose pooled builders effect was positive with its interval coming within 0.03 of zero: need and block_leader. Pooled estimates over stages are after-the-fact pooling, named as such.

Every registered decision rule reads the same: all games complete, and the seating-corrected effect's 95 percent seed interval either excludes zero in the hypothesized direction (support), excludes zero the other way (refutation), or neither (inconclusive). Every pair is inconclusive on its own stages; the one cohort whose interval excluded zero is named below.

The cohorts

All 36 cohorts completed all 256 games: 9,216 games, no invalid moves, stalls, or timeouts. One half (leader, development, half B) was interrupted at 143 of 256 games by a shell timeout on the controlling session and was rerun from scratch; the partial interrupted attempt is filed as interrupted and its games were not pooled.

Full results and cohort ledger
ArmCohortHalfExperimentRunSlot 0Slot 1Slot 2Slot 3
leaderdev 0-63, buildersAregistered protocol256-game engine cohort129991315
leaderdev 0-63, buildersBregistered protocol256-game engine cohort13393219
needdev 0-63, buildersAregistered protocol256-game engine cohort1271001910
needdev 0-63, buildersBregistered protocol256-game engine cohort129100189
threatdev 0-63, buildersAregistered protocol256-game engine cohort1221031912
threatdev 0-63, buildersBregistered protocol256-game engine cohort1201042111
block leaderdev 0-63, buildersAregistered protocol256-game engine cohort1251011713
block leaderdev 0-63, buildersBregistered protocol256-game engine cohort1231041613
leaderconf 800-863, buildersAregistered protocol256-game engine cohort127931719
leaderconf 800-863, buildersBregistered protocol256-game engine cohort125108158
needconf 800-863, buildersAregistered protocol256-game engine cohort126952510
needconf 800-863, buildersBregistered protocol256-game engine cohort127103206
threatconf 800-863, buildersAregistered protocol256-game engine cohort124108159
threatconf 800-863, buildersBregistered protocol256-game engine cohort127103197
block leaderconf 800-863, buildersAregistered protocol256-game engine cohort125104207
block leaderconf 800-863, buildersBregistered protocol256-game engine cohort127106185
leaderpop 800-863, populationAregistered protocol256-game engine cohort89807611
leaderpop 800-863, populationBregistered protocol256-game engine cohort80877316
needpop 800-863, populationAregistered protocol256-game engine cohort83867512
needpop 800-863, populationBregistered protocol256-game engine cohort87867211
threatpop 800-863, populationAregistered protocol256-game engine cohort87847015
threatpop 800-863, populationBregistered protocol256-game engine cohort72918112
block leaderpop 800-863, populationAregistered protocol256-game engine cohort78897613
block leaderpop 800-863, populationBregistered protocol256-game engine cohort80897611
needext 864-927, buildersAregistered protocol256-game engine cohort128106148
needext 864-927, buildersBregistered protocol256-game engine cohort118116175
block leaderext 864-927, buildersAregistered protocol256-game engine cohort13299169
block leaderext 864-927, buildersBregistered protocol256-game engine cohort127106158
needext 864-927, populationAregistered protocol256-game engine cohort94816615
needext 864-927, populationBregistered protocol256-game engine cohort88886515
block leaderext 864-927, populationAregistered protocol256-game engine cohort97796515
block leaderext 864-927, populationBregistered protocol256-game engine cohort88876615
need + block leaderfinal 928-991, buildersAregistered protocol256-game engine cohort131941516
need + block leaderfinal 928-991, buildersBregistered protocol256-game engine cohort1221031714
need + block leaderfinal 928-991, populationAregistered protocol256-game engine cohort87867211
need + block leaderfinal 928-991, populationBregistered protocol256-game engine cohort9473836

Seating-corrected effects

ArmDevelopment 0-63Confirmation 800-863Extension 864-927Population 800-863Population extension 864-927
leader−0.020 (−0.079 to +0.040)+0.033 (−0.025 to +0.091)not extended+0.031 (−0.030 to +0.092)not extended
need−0.004 (−0.047 to +0.039)+0.014 (−0.031 to +0.058)+0.039 (−0.001 to +0.079)−0.008 (−0.045 to +0.030)+0.025 (−0.013 to +0.064)
threat+0.006 (−0.051 to +0.063)−0.016 (−0.065 to +0.034)not extended+0.043 (−0.005 to +0.091)not extended
block leader+0.010 (−0.020 to +0.039)+0.000 (−0.029 to +0.029)+0.023 (−0.003 to +0.050)−0.004 (−0.038 to +0.030)+0.033 (+0.007 to +0.060)

The development, confirmation, and extension columns are the builders lineup; the two population columns are the population lineup. The bold cell is the only cohort of the program whose interval excludes zero, on the positive side.

Pooled over stages (after-the-fact pooling, equal weight per seed):

ArmBuilders, pooledSeedsPopulation, pooledSeeds
leader+0.007 (−0.035 to +0.049)128+0.031 (−0.030 to +0.092)64
need+0.016 (−0.008 to +0.041)192+0.009 (−0.018 to +0.036)128
threat−0.005 (−0.042 to +0.033)128+0.043 (−0.005 to +0.091)64
block leader+0.011 (−0.005 to +0.027)192+0.015 (−0.007 to +0.036)128

The seating terms of the builders pairs reproduce the known adjacency term (+0.047 to +0.137, each interval containing the values measured by the null pairs); the population pairs measure seating terms near zero (−0.039 to +0.043), consistent with three searches breaking the two-seat adjacency.

Compared with the hand-written leaf

The development stage uses the same seeds 0-63 and the same builders lineup as the first hand-leaf cohorts, so the raw slot-0 contrasts are comparable. With the arm in slot 0, the hand leaf read leader +0.125, need +0.160, threat +0.117, and block leader +0.168; under the tables the same design reads leader +0.117, need +0.105, threat +0.074, and block leader +0.094. Both sets are seating-contaminated, and both shrink to about zero once the pair is swapped: the earlier studies corrected to −0.023, −0.045, −0.033, and +0.012, and this study's builders pools read +0.007, +0.016, −0.005, and +0.011. The pattern that looked like an effect under the hand leaf was the seating term, and under the tables it is the seating term again. One difference is real: the need rule's hand-leaf refutation (−0.045, the only interval excluding zero in the earlier program) does not reproduce under the tables, where its builders pool reads +0.016 (−0.008 to +0.041). Under the learned leaf, choosing the victim by the card this seat needs is not measurably worse than choosing by hand size, but it is not measurably better either.

Where the robber goes

Where the robber goes under each arm

Measured evidence
00.150.30.450.6leaderneedthreatblock leaderneed + block leaderArmShare of robber moves

Scroll the chart horizontally to inspect all values.

On the leader best hex, arm seatOn the leader best hex, control seatRobbed the leader, arm seatRobbed the leader, control seat

Share of robber moves landing on the leader best hex and robbing the leader, arm seats against control seats.

View data table

Source: Robber-move records of the same swapped-pair cohorts, pooled over both halves and every stage of each arm: the share of the arm seats and control seats placing the robber on the points leader highest-pip hex (cities double) and robbing the points leader among the mover rivals.

SeriesArmShare of robber moves
On the leader best hex, arm seatleader0.3112
On the leader best hex, arm seatneed0.2695
On the leader best hex, arm seatthreat0.2673
On the leader best hex, arm seatblock leader0.2742
On the leader best hex, arm seatneed + block leader0.2827
On the leader best hex, control seatleader0.2701
On the leader best hex, control seatneed0.2719
On the leader best hex, control seatthreat0.2696
On the leader best hex, control seatblock leader0.2701
On the leader best hex, control seatneed + block leader0.2759
Robbed the leader, arm seatleader0.4993
Robbed the leader, arm seatneed0.4236
Robbed the leader, arm seatthreat0.4729
Robbed the leader, arm seatblock leader0.4317
Robbed the leader, arm seatneed + block leader0.4234
Robbed the leader, control seatleader0.4221
Robbed the leader, control seatneed0.4237
Robbed the leader, control seatthreat0.4222
Robbed the leader, control seatblock leader0.4228
Robbed the leader, control seatneed + block leader0.4145

Pooled over both halves of every stage of each arm, the mechanisms do what they claim. The leader arm lands the robber on the points leader's highest-pip hex on 31.1 percent of its robber moves against 27.0 percent for its control seats, and robs the leader on 49.9 percent against 42.2 percent. The threat arm robs the leader on 47.3 percent against 42.2 percent and shifts theft toward the other search (37.2 percent of victims against 34.2 percent), consistent with preferring threatening victims over quiet builders. The need arm barely moves these observables, as expected: it changes which non-leader seat holds the stolen card, a distinction the arena records do not carry. The block leader arm nudges late-game placements (43.2 percent leader victims against 42.3 percent).

Victim distribution by opponent seat, pooled the same way, is nearly flat for every arm and control (about 34 percent each to the other search, the eta builder, and the fast or third-search seat), and the builders themselves land on the leader's best hex on 27.9 percent of their robber moves with 43.0 percent leader victims across all cohorts. The arm seats steal from each opponent category in almost the shares the control seats do; only the leader arm's hex placement and everyone's leader-victim rate move visibly.

The other game statistics agree with no mechanism change: counters (about 2.3 offers countered per game), declines (about 7.6), development bought (2.8 to 2.9), cities built (1.5), roads built (4.5), cards discarded (about 6.0), and player trades (about 4.1) differ between arm and control seats by under a twentieth in every arm.

Decision cost

No switch costs decision time. Within each pair the arm seat and the control seat decide within two percent of each other in every stage (for example 576 against 582 ms for the leader pair under load, 69 against 70 ms for the need extension on an idle machine). The cohorts ran at machine loads from about 150 to 6, so absolute milliseconds are not comparable between stages; the within-pair ratios are. The switches score the same shortlist the search already enumerates, so nothing new is evaluated.

The combination

Need and block leader are the two arms the extension rule marked positive with near-zero intervals, and both read positive in every lineup pooled, so they were combined as a final swapped pair on fresh seeds 928-991 in both lineups. The combination reads +0.035 (−0.015 to +0.085) in the builders lineup and −0.039 (−0.090 to +0.012) in the population lineup. The two readings disagree, both cross zero, and the pooled combination is −0.002. Two small effects that do not individually clear their intervals do not add to one that does.

What could still be wrong

The extension arms were chosen after seeing their first two stages, so their final estimates carry that selection; the block leader population extension that excluded zero is one of ten cells, and the study's own bar asked for the pooled estimate to exclude zero, which it does not. Effects under the tables are development-tier engine evidence at depth 2 against one family of opponents; the tables leaf was trained by self-play against these builders and may already price blocking, which would hide exactly the differences these switches try to create. Population-lineup opponents are three copies of the same search, not distinct retaliators. Decision times are wall-clock and the stages ran at very different machine loads, so only within-pair comparisons are usable. Seeds 0-63 were used by the earlier hand-leaf development cohorts, so the development stage is a comparison on shared boards, not fresh ones, by design.

What this means for the browser

The browser build searches with the hand-written leaf at depth 3 under a one second budget, so the deciding evidence is the hand-leaf measurements: leader −0.023, need −0.045 (refuted), threat −0.033, block leader +0.012, all corrected swapped pairs. The tables evidence agrees: every switch is inconclusive, none reaches +0.03 pooled, and the one positive arm combination disagrees with itself across lineups. Since no switch costs decision time, the reason to keep them off is not cost but the absence of a measured effect in either leaf. None of the four switches is worth enabling in the browser. If one is revisited later, block leader is the candidate: it is the only arm positive in every lineup pooled under the tables and the only cohort to exclude zero, but its hand-leaf reading was +0.012 and its pooled tables effects stay under the bar, so a revisit needs a different mechanism or a sharper lineup, not a bigger cohort of the same one.

Methods and reproduction

Implementation: no server change. The arms reuse SearchConfig.robber (leader, need, threat) and endgame.block_leader in crates/expectimax/src/v2, documented in docs/expectimax.md under "Robber targeting" and the endgame section, on top of the learned leaf (leaf.tables = artifacts/ntuple/hex-portfolio-main.bin, SHA-256 f8e61e66…). Arm seats are that baseline plus one switch, for example v2:{"depth":2,"leaf":{"tables":"/home/keshav/settlers/research/artifacts/ntuple/hex-portfolio-main.bin"},"robber":{"leader":true}}. All 36 experiments are registered under containment/robber-tables with the hypothesis and decision rule quoted per cohort; the analysis is analysis/robber_tables.py (per-stage effects, decision times, robber placement and victim shares) and analysis/robber_tables_assets.py (the two plots from the retained runs). Archives with verified receipts are under artifacts/, one per run.

Reproduction: in the research checkout with SETTLERS_SERVER_DIR=~/settlers/.worktrees/server-robber-tables, python3 -m harness.engine run EXPERIMENT --threads 6 replays any cohort; deterministic tapes make every run reproducible. The interrupted partial is interrupted attempt (143 of 256 games, replaced by 256-game engine cohort). Seating-corrected effects: python3 analysis/seating_pair.py runs/HALF_A/games.jsonl runs/HALF_B/games.jsonl.

All studies · Leader containment · Experiment log

Detailed result and cohort context

The three robber-targeting switches and the late-game leader bias, refuted or zeroed under the hand-written leaf, were retested one at a time against the search with the learned n-tuple tables, each as a swapped pair in a builders lineup (eta and fast) and a population lineup (a third tables search and eta). Every cohort completed. The seating-corrected effects pooled over all stages are leader +0.007 (95% interval −0.035 to +0.049) and +0.031 (−0.030 to +0.092), need +0.016 (−0.008 to +0.041) and +0.009 (−0.018 to +0.036), threat −0.005 (−0.042 to +0.033) and +0.043 (−0.005 to +0.091), and block leader +0.011 (−0.005 to +0.027) and +0.015 (−0.007 to +0.036) wins per game, builders and population lineups in that order. Only one cohort of the ten cells excluded zero: the block leader extension in the population lineup, +0.033 (+0.007 to +0.060). No switch reaches the registered bar of +0.03 pooled with an interval excluding zero, and combining the two arms that read positive (need plus block leader) gave +0.035 (−0.015 to +0.085) in the builders lineup and −0.039 (−0.090 to +0.012) in the population lineup. All switches stay off by default.