Road blocking under the learned leaf
Road switches under the learned leaf
Measured evidenceScroll the chart horizontally to inspect all values.
Seating-corrected effect of each road switch per stage, under the tables alone and on the quarter blend.
View data table
Source: Engine-arena swapped pairs under expansion/roads-tables: the switch against the unchanged tables-leaf control, each half 64 deterministic seeds played once per rotation (256 games), seating-corrected as the per-seed half-difference of the two halves. Dev, confirm, and extension builders use the builders lineup; population stages seat a third tables search in slot 2. Run IDs are listed in the roads-tables report.
| Series | Switch and stage | Wins per game, switch minus control | Low | High |
|---|---|---|---|---|
| Tables leaf alone (95% interval) | block dev | -0.0078 | -0.0186 | 0.0029 |
| Tables leaf alone (95% interval) | block confirm | 0.002 | -0.0096 | 0.0135 |
| Tables leaf alone (95% interval) | block population | 0.002 | -0.0067 | 0.0106 |
| Tables leaf alone (95% interval) | gate dev | -0.0195 | -0.0622 | 0.0232 |
| Tables leaf alone (95% interval) | gate confirm | -0.0176 | -0.0663 | 0.0312 |
| Tables leaf alone (95% interval) | gate population | -0.0234 | -0.0596 | 0.0127 |
| Quarter blend, hand 0.25 in both seats (95% interval) | block dev | 0.0273 | -0.0106 | 0.0653 |
| Quarter blend, hand 0.25 in both seats (95% interval) | block confirm | 0.0059 | -0.0488 | 0.0605 |
| Quarter blend, hand 0.25 in both seats (95% interval) | block population | -0.0586 | -0.096 | -0.0212 |
| Quarter blend, hand 0.25 in both seats (95% interval) | block ext builders | -0.0039 | -0.0512 | 0.0433 |
| Quarter blend, hand 0.25 in both seats (95% interval) | block ext population | -0.0078 | -0.0456 | 0.0299 |
| Quarter blend, hand 0.25 in both seats (95% interval) | gate dev | -0.0605 | -0.1072 | -0.0139 |
| Quarter blend, hand 0.25 in both seats (95% interval) | gate confirm | -0.0586 | -0.1123 | -0.0049 |
| Quarter blend, hand 0.25 in both seats (95% interval) | gate population | 0.0215 | -0.0259 | 0.0689 |
The two switches and the two leaf contexts
Both switches are leaf weights with a candidate-generation part. With
opponent_expansion set, the search charges the leading opponent's best
reachable site, discounted by road distance, and the road shortlist gains the
roads that raise that distance (v2/search.rs). With road_contention_gate
set to 1, the longest-road contention credit, in the leaf and in the road
shortlist, applies only while the seat ties or leads the table in brick plus
lumber production. Under the tables alone (leaf.hand 0) the hand-written leaf
terms are not evaluated, so only the shortlist part of each switch acts. The
leaf parts were measured on top of the blended leaf, with "hand":0.25 in both
the candidate and the control seat: the tables plus a quarter-strength
hand-written leaf, the blend the leaf-blend study program
found ahead of the tables alone. The weight values, 0.5 and 1.0, are the ones
the earlier study registered.
The blend arms were first registered at hand:1.0. An unregistered screen by
the leaf-blend agent then measured the full blend at −0.207 wins per game
against the tables alone, far weaker, and the quarter blend at +0.074, so the
ten full-blend registrations that had not started were withdrawn before running
and re-registered at 0.25. One full-blend cohort was already running and
completed; it is reported below as a single seating, not as evidence of the
switch's effect.
The cohorts
Every cohort is one half of a swapped pair: 64 deterministic seeds played once per rotation (256 games), the candidate and the control in slots 0 and 1, half B swapping them, 6 threads per tournament. Stages 1 and 2 use the builders lineup (an ETA and a fast builder in slots 2 and 3, matching the hand-leaf cohorts); stage 3 and the extension's second pair use the population lineup (the plain tables baseline in slot 2 and ETA in slot 3). Stage 1 uses seeds 0 to 63, stages 2 and 3 use seeds 800 to 863, and the extension uses seeds 864 to 927. Every cohort completed 256 of 256 games.
opponent_expansion 0.5 under the tables alone (shortlist only):
Full results and cohort ledger
| Stage | Half | Run | Slot 0 wins | Slot 1 wins | Contrast (95% interval) |
|---|---|---|---|---|---|
| 1 dev, builders | A | 256-game engine cohort | 124 | 101 | +0.090 (−0.006 to +0.186) |
| 1 dev, builders | B | 256-game engine cohort | 126 | 99 | +0.105 (+0.014 to +0.197) |
| 2 confirm, builders | A | 256-game engine cohort | 126 | 104 | +0.086 (−0.021 to +0.192) |
| 2 confirm, builders | B | 256-game engine cohort | 126 | 105 | +0.082 (−0.024 to +0.188) |
| 3 population | A | 256-game engine cohort | 81 | 91 | −0.039 (−0.127 to +0.048) |
| 3 population | B | 256-game engine cohort | 80 | 91 | −0.043 (−0.126 to +0.040) |
road_contention_gate 1.0 under the tables alone (shortlist gating only):
Full results and cohort ledger
| Stage | Half | Run | Slot 0 wins | Slot 1 wins | Contrast (95% interval) |
|---|---|---|---|---|---|
| 1 dev, builders | A | 256-game engine cohort | 120 | 104 | +0.063 (−0.034 to +0.159) |
| 1 dev, builders | B | 256-game engine cohort | 126 | 100 | +0.102 (+0.010 to +0.193) |
| 2 confirm, builders | A | 256-game engine cohort | 117 | 113 | +0.016 (−0.086 to +0.117) |
| 2 confirm, builders | B | 256-game engine cohort | 122 | 109 | +0.051 (−0.051 to +0.153) |
| 3 population | A | 256-game engine cohort | 76 | 94 | −0.070 (−0.159 to +0.019) |
| 3 population | B | 256-game engine cohort | 78 | 84 | −0.023 (−0.109 to +0.062) |
opponent_expansion 0.5 on the 0.25 blend (shortlist plus leaf term):
Full results and cohort ledger
| Stage | Half | Run | Slot 0 wins | Slot 1 wins | Contrast (95% interval) |
|---|---|---|---|---|---|
| 1 dev, builders | A | 256-game engine cohort | 140 | 94 | +0.180 (+0.086 to +0.273) |
| 1 dev, builders | B | 256-game engine cohort | 129 | 97 | +0.125 (+0.023 to +0.227) |
| 2 confirm, builders | A | 256-game engine cohort | 125 | 108 | +0.066 (−0.036 to +0.169) |
| 2 confirm, builders | B | 256-game engine cohort | 122 | 108 | +0.055 (−0.054 to +0.164) |
| 3 population | A | 256-game engine cohort | 88 | 85 | +0.012 (−0.075 to +0.098) |
| 3 population | B | 256-game engine cohort | 103 | 70 | +0.129 (+0.039 to +0.219) |
| 4 ext, builders | A | 256-game engine cohort | 131 | 104 | +0.105 (−0.001 to +0.212) |
| 4 ext, builders | B | 256-game engine cohort | 128 | 99 | +0.113 (+0.015 to +0.212) |
| 4 ext, population | A | 256-game engine cohort | 84 | 79 | +0.020 (−0.071 to +0.110) |
| 4 ext, population | B | 256-game engine cohort | 89 | 80 | +0.035 (−0.049 to +0.120) |
road_contention_gate 1.0 on the 0.25 blend (shortlist gating plus leaf term):
Full results and cohort ledger
| Stage | Half | Run | Slot 0 wins | Slot 1 wins | Contrast (95% interval) |
|---|---|---|---|---|---|
| 1 dev, builders | A | 256-game engine cohort | 122 | 109 | +0.051 (−0.046 to +0.147) |
| 1 dev, builders | B | 256-game engine cohort | 138 | 94 | +0.172 (+0.083 to +0.260) |
| 2 confirm, builders | A | 256-game engine cohort | 116 | 115 | +0.004 (−0.115 to +0.123) |
| 2 confirm, builders | B | 256-game engine cohort | 131 | 100 | +0.121 (+0.009 to +0.233) |
| 3 population | A | 256-game engine cohort | 92 | 81 | +0.043 (−0.053 to +0.139) |
| 3 population | B | 256-game engine cohort | 85 | 85 | 0.000 (−0.089 to +0.089) |
The one full-blend cohort that was already running when the blend moved to
0.25, stage 1 half A at hand:1.0 (256-game engine cohort,
the registered protocol), finished with the
candidate ahead 106 to 96, a single-seating contrast of +0.039 (−0.066 to
+0.144). Its partner registration was withdrawn, so it has no swapped half and
carries the seating term unmeasured.
Seating-corrected effects
The effect of a switch is the per-seed half-difference of the two halves'
slot-0-minus-slot-1 contrasts and the seating term their half-sum
(analysis/seating_pair.py). Per stage:
| Arm | Stage 1 dev | Stage 2 confirm | Stage 3 population | Stage 4 extension |
|---|---|---|---|---|
opponent_expansion, tables | −0.008 (−0.019 to +0.003) | +0.002 (−0.010 to +0.014) | +0.002 (−0.007 to +0.011) | not triggered |
road_contention_gate, tables | −0.020 (−0.062 to +0.023) | −0.018 (−0.066 to +0.031) | −0.023 (−0.060 to +0.013) | not triggered |
opponent_expansion, blend | +0.027 (−0.011 to +0.065) | +0.006 (−0.049 to +0.061) | −0.059 (−0.096 to −0.021) | builders −0.004 (−0.051 to +0.043), population −0.008 (−0.046 to +0.030) |
road_contention_gate, blend | −0.061 (−0.107 to −0.014) | −0.059 (−0.112 to −0.005) | +0.022 (−0.026 to +0.069) | not triggered |
After-the-fact pooling over stages: opponent_expansion under the tables is
−0.003 (−0.011 to +0.005) over the 128 builders seeds and −0.001 (−0.007 to
+0.005) over all 192 boards; on the blend it is +0.010 (−0.017 to +0.037) over
the 192 builders seeds (the extension fired because the 128-seed read of +0.017
had an interval reaching −0.017, inside the 0.03 gate) and −0.033 (−0.060 to
−0.006) over the 128 population seeds. road_contention_gate under the tables
is −0.019 (−0.051 to +0.014) over the builders seeds and −0.020 (−0.045 to
+0.004) over all 192 boards; on the blend it is −0.060 (−0.095 to −0.024) over
the builders seeds, refuted, with the population stage at +0.022 (−0.026 to
+0.069).
Measured seating terms sit between −0.047 and +0.152 per pair, largest under the blend in the builders lineup, matching the standing diagnosis that adjacent searches race. The stage-3 pairs, where a third search sits between the contestants in strength but not in seat order, show the term near zero or negative (−0.041 and −0.047 under the tables).
Comparison with the hand-written leaf
Under the hand-written leaf the corrected effects were +0.008 (−0.053 to +0.069) for the blocking weight and −0.004 (−0.061 to +0.053) for the gate on boards 64 to 127. The single-seating development contrasts on the same boards 0 to 63 as stage 1 here were +0.168 and +0.113; under the tables the same orientation gave +0.090 and +0.063, and under the blend +0.180 and +0.051, all carrying their pair's seating term (+0.098, +0.082, +0.152, +0.111). Neither leaf shows a real effect of the blocking weight. The gate differs by context: inert under the hand-written leaf and under the tables alone, but a measurable loss under the blend, where its leaf term is live at strength and the gated seat gives up award races it would otherwise contest.
What the switches did to the board
Diagnostics pool each arm's two halves per stage; candidate first, control
second. Roads built per game barely move anywhere (5.2 to 5.7 in both seats),
and opponent_expansion barely finds blocking roads to build: 0.04 to 0.08 per
game against the control's 0.02 to 0.05, with denied-site value near zero. The
learned leaf already prices expansion pressure, so the shortlist's extra
candidates change almost nothing. The gate does exactly what it says: the gated
seat holds the longest-road award at game end in 38% of games against the
control's 53% on the blended development boards (42% against 50% on the
confirmation boards, 29% against 35% in the population stage), builds the same
number of roads, and loses the points the award and the contested races were
worth. Sites lost stay within 0.05 per game between the seats in every arm
(1.7 to 1.9), and settlements built move by hundredths at most.
Decision cost is the blocking weight's other price: the candidate averages about 14% more per decision than the control under the tables alone (the shortlist part) and about 30% more under the blend, where the leaf term runs the opponent reach walk at every leaf. The gate is free or slightly cheaper, since it removes candidates. Absolute decision times in these cohorts are inflated by a heavily shared machine; the ratios within each run are the fair comparison.
What could still be wrong
Everything here is engine-arena development tier on deterministic tapes; no protocol cohort was run, and the opponents never adjust beyond their fixed policies. The 0.25 blend choice rests on another agent's unregistered screen, not a registered cohort. Weight values other than 0.5 and 1.0 were not tried, so a weaker gate could cost less; the interval under the blend says the registered value is harmful, not that every value would be. The population lineup's third search is the plain tables baseline, and the blocking weight's refutation there (−0.033 pooled, after the fact) comes from 128 seeds in one such table. Decision-time ratios are measured under a machine shared with nine other agents and are noisy, though both seats of a pair share the same load. The hand-1.0 half-A cohort is a single seating and says nothing on its own.
What this means for the browser
Do not enable either switch. The browser build runs the hand-written leaf at depth 3 under a one-second budget, and under that leaf both switches already measured zero (+0.008 and −0.004, intervals crossing zero). This study adds that under the learned tables the blocking shortlist is also worth nothing (−0.001 pooled over 192 boards), that the blocking leaf term on the blend is worth nothing against builders (+0.010 over 192 seeds) and is refuted in the population lineup, and that the gate's leaf term is refuted under the blend at −0.060. Cost decides any remaining doubt for the blocking weight: about 30% more decision time under the blend and about a third more under the hand-written leaf (314 ms against 232 ms per decision in the earlier study), both over the 25% bound, for an effect that is zero. The gate is free but harmful where it acts. Neither switch transfers.
Methods and reproduction
Implementation: the switches are weights.opponent_expansion and
weights.road_contention_gate in crates/expectimax/src/v2 (leaf terms in
leaf.rs, road-shortlist parts in search.rs), documented in the server's
docs/expectimax.md under "Leaf". No server change was needed: both switches
behaved sanely under the tables leaf. Seats: control
v2:{"depth":2,"leaf":{"tables":"/home/keshav/settlers/research/artifacts/ntuple/hex-portfolio-main.bin"}},
tables-only candidates add "weights":{"opponent_expansion":0.5} or
"weights":{"road_contention_gate":1.0}, blend candidates and controls add
"hand":0.25 inside the leaf object. Lineups: builders (slots 2 and 3 eta,
fast) and population (slot 2 the tables baseline, slot 3 eta).
Reproduction: python3 -m harness.engine run EXPERIMENT --threads 6 in a
research checkout with SETTLERS_SERVER_DIR pointing at a server build of
settlers-arena. Stage 1 pairs (seeds 0-63, builders): experiments
registered protocol / registered protocol
(blocking, tables), registered protocol /
registered protocol (gate, tables),
registered protocol / registered protocol
(blocking, blend), registered protocol /
registered protocol (gate, blend). Stage 2 (seeds 800-863,
builders): registered protocol /
registered protocol, registered protocol
/ registered protocol, registered protocol
/ registered protocol, registered protocol
/ registered protocol. Stage 3 (seeds 800-863, population):
registered protocol / registered protocol,
registered protocol / registered protocol,
registered protocol / registered protocol,
registered protocol / registered protocol.
Stage 4 extension (seeds 864-927): registered protocol /
registered protocol (builders), registered protocol
/ registered protocol (population). The ten withdrawn
hand:1.0 registrations are listed in the 2026-09-11 log. Analysis:
analysis/roads_tables.py (per-pair effects and diagnostics),
analysis/roads_tables_pool.py (after-the-fact pooling),
analysis/roads_tables_assets.py (the plot above from the retained runs).
Effects are computed from runs/<id>/games.jsonl; archive receipts for every
run are under records/artifacts/.
All studies · Expansion and races · Experiment log
Detailed result and cohort context
The two road switches of the road-blocking study were
re-measured with the learned n-tuple tables as the search's leaf, each as a
registered swapped pair in three lineups and two leaf contexts (7,424 games, all
completed). opponent_expansion 0.5, which under the tables acts only through
the road shortlist, measures −0.003 wins per game (95% interval −0.011 to
+0.005) pooled over the 128 builders-lineup seeds and +0.002 (−0.007 to +0.011)
with three searches at the table. On top of the blended leaf (tables plus
quarter-strength hand terms) it reads +0.010 (−0.017 to +0.037) over 192
builders seeds and −0.033 (−0.060 to −0.006) over 128 population seeds.
road_contention_gate 1.0 reads −0.020 (−0.045 to +0.004) pooled over all
192 tables-leaf boards and −0.060 (−0.095 to −0.024) over the 128
builders seeds under the blend, the one context where its leaf term is live at
strength. Under the hand-written leaf both switches had already corrected to
+0.008 and −0.004. Both stay off everywhere, the browser included.