Catan — 47,000 Games of Empirical Findings

Ioana Roman analyzed 47,000 recorded Catan games from twosheep.io and turned the board game into a stochastic dynamic system with real numbers attached. Headline findings: turn order is essentially balanced, opening placement predicts the winner barely above random (27% vs 25% baseline), Longest Road and Largest Army are symptoms not causes of victory, missing a resource barely matters because the market corrects, Monopoly cards are gigantic timing weapons (one played → 40% win rate; two played → 55%), and winners systematically overpay on trades. The last one is the empirical confirmation, across 47k games, of the bilateral trade valuation thesis: trade evaluation must be trajectory-based, not isolated NPV. Important caveat: despite the title, the video delivers descriptive statistics about how games tend to play out, not a strategic prescription for how to win. The empirical landmarks are real and useful as constraints, but the prescription layer — state-conditional opening policies, timing policies, trajectory valuation — is the work the Monopoly project does and Roman doesn’t.

Links: Gaming, LLM Agents Across Strategic Games, LLM Game Benchmark — Outline, Bilateral Trade Valuation, The Multiplayer Coalition Problem, BattleValue, Planner-LM Composites, The LLM Grounding Problem, Monopoly Theory, Energy-Based Models, Catan 50-Game Validation — proof-of-method run against the public 50-game Kaggle dataset; aggregate city-engine direction confirmed, board-conditional refinement only partially supported, theory’s working form updated to “universal city-engine bias + layered archetype effect”, Randomness as Termination (N≥3) — Catan’s dice + dev cards as the parity-breaker / termination layer

Primary source: I Analysed 47,000 Games of Catan. Here’s How to Win Every Time (mathematically) — Ioana Roman, 2026-05-10. Transcript. Data: twosheep.io.


The mathematical model

Roman frames Catan as a multi-agent stochastic dynamic optimization problem. The skeleton:

The framing matters: this is “probabilistic income streams + endogenous market dynamics + asymmetric information + stopping-time decision,” which is much closer to the bilateral trade valuation and Monopoly project framings than to anything that “luck vs skill” captures. Catan is a small economy under uncertainty.

The seven empirical findings

1. Turn order is balanced

Position Win rate
1st 24.9%
2nd ~25%
3rd slightly >25%
4th ~25%

The snake-order placement (last player places twice before order reverses) genuinely works. There’s no exploitable structural disadvantage to seating. If you lost from 4th, it wasn’t the reason.

2. Longest Road and Largest Army are symptoms, not causes

These are large lifts over baseline, but the holder still loses 40% / 46% of the time. Roman’s read: they reflect underlying structural advantage rather than create it. Longest Road tracks spatial dominance over the graph; Largest Army tracks ore/wheat investment plus robber control. Both are downstream of the actual strategic position.

This is the structural-realism pattern ([[BattleValue]], [[D&D Spell Damage Model]]): the headline number is a proxy. The thing that produces both the bonus and the win is the underlying configuration.

3. Pip-weighting reveals winners prefer the city engine

Each number token has a weight proportional to its dice-roll probability. Summing the weights of a settlement’s three adjacent tiles gives a scalar expected production intensity.

Winners had slightly higher exposure to:

(combined: the “city engine” — what powers cities and development cards)

Winners and losers had nearly identical exposure to:

(the “road engine”)

The finding is small per game but consistent across 47k. Catan is not fundamentally a road game; it’s a compounding city game. Cities double production, development cards provide hidden leverage and Largest Army eligibility, and the late-game acceleration runs on ore+wheat+sheep. The road-heavy intuition that beginners get from “Longest Road feels decisive” is misleading at the strategy-selection layer.

4. Missing a resource barely matters

Across 47k games, missing any single resource at the opening leaves win rate at ~25% with small dips for missing ore or wood. The trading and port system is enough to correct the imbalance.

This is self-balancing market behavior, the same pattern that makes N≥3 multiplayer coalitions game-theoretically interesting. The game has built-in correction mechanisms; outcomes don’t collapse on one bad draw. Resilience comes from the trade graph, not from the production graph.

5. The opening predicts almost nothing

Roman trained a classifier on initial-placement features (pip exposure, resource diversity, balance metrics), tested on 20% held-out games. Accuracy: 27%. Random baseline: 25%.

The opening matters by two percentage points.

The other 73% of outcome variance lives in the post-opening sequence: trades, development card draws, robber placements, timing, interaction. Catan is path-dependent, not predetermined at turn zero. This is the empirical answer to “where does the game actually happen?” — not in the opening, but in the running game.

6. Monopoly cards are timing weapons, not just resource cards

Monopoly cards played Win rate
0 25%
1 40%
2 (max) 55%

Roman’s framing: “systemic liquidity extraction.” A late-game Monopoly when opponents have accumulated inventory is a much bigger shock than an early one — the card redistributes the inventory state, and the larger the state, the bigger the move.

Two architectural implications:

7. The trade-efficiency paradox — winners overpay

Trade efficiency = resources received / resources given.

Trade efficiency Win rate
~1.0 (perfectly even) 8%
> 1.0 (favorable) (not the winners)
< 1.0 (overpaying) (where the winners cluster)

Players who make perfectly fair trades win at 8% — three times worse than baseline. Players who systematically appear to overpay are the ones who win.

The mechanism Roman names: they’re not optimizing the trade; they’re optimizing their position. Overpaying to complete a city before a key roll. Overpaying to secure Longest Road before someone else gets it. Overpaying to block an opponent’s expansion. Short-term inefficiency for long-term positional gain.

This is the empirical confirmation of bilateral trade valuation across 47k games. That page argued: a trade’s value is where both players end up 20 turns later, not the static resource delta. Roman’s data is the smoking gun. Players treating trade as a transaction-level optimization (even ratio = “fair”) lose almost three times the rate of players treating it as a position-level move.

The pattern generalizes well past Catan:

The principle is one layer above game theory: valuation requires a trajectory, not a snapshot.

How this lands in the vault

Confirms

Sharpens

Connects

The descriptive / prescriptive gap — what the analysis doesn’t do

Roman’s title promises “How to Win Every Time (mathematically).” The video doesn’t deliver that. What it delivers is descriptive statistics about how games tend to play out. That’s a different (and lesser) artifact than strategic prescription. Worth naming the distinction because the vault’s Monopoly project lives on the other side of this line, and the comparison clarifies what an actual Catan strategy paper would have to add.

What Roman did (the empirical / descriptive layer):

What she didn’t do (the prescription / policy layer):

The gap is the difference between “winners tend to have X” and “given the current state, do Y”. Roman is honest about this implicitly (her language is “structural advantage,” “tendencies,” “compounding”) but doesn’t bridge to the second layer. Headline tendencies → strategy is a separate exercise the video didn’t take on.

What a Catan prescription layer would look like

The Monopoly project has the pattern. For Monopoly:

The Catan equivalents would be:

Empirical finding (Roman) What the prescription layer would add
Winners favor ore/wheat/sheep A placement frontier over (production weight, resource diversity, port adjacency, blocking value) with Pareto-optimal openings labeled by board configuration
Monopoly cards lift win rate to 40%/55% A state-conditional play policy: hold Monopoly until opponents’ aggregate inventory of resource R exceeds threshold T(R, turn, opponents’ VP)
Winners overpay for position A trajectory-valuation function — when offered trade T at state S, evaluate ΔV(state at turn S+20 with T) − ΔV(state at S+20 without T)
Opening predicts 27% Implies the running game dominates outcome variance; that’s where the planner work needs to go (state-tracking, opponent modeling, robber timing — none of which Roman touches)

This is the same shape as Chris’s frontier trade theory for Monopoly: the empirical observation that “winners overpay” becomes the prescriptive question “by how much, for which positions, against which opponents.” Roman’s data has the inputs; the optimization formalism is missing.

Why the resource-vector framing matters here

Roman models per-player state as a 5D resource vector evolving over time. That formalism is one Pareto-frontier away from a strategy theory — the vault’s Monopoly work already lives in that adjacent space. Once you have state vectors, you can:

Roman has the vectors but stops at descriptive statistics over them. The prescription layer is what to do with the vectors — and the Monopoly project shows the methodology transfers cleanly.

The honest verdict on the video: good empirical landmarks, good intuition-builders, useful as constraints for any future prescriptive theory (a Catan strategy paper that contradicts the 27% classifier accuracy is wrong). But the title overpromises. “How to win every time” would require the prescription layer the video doesn’t build.

Toward the prescription — capability value over face value

Roman’s tokens are sharp because she stops at correlation. The vault-side reading of what those correlations mean is sharper, and it changes the strategy implications.

The general principle: Longest Road, Largest Army, and the Monopoly card don’t lift win rate because of the points they grant. They lift win rate because they unlock or signal strategic capability.

The face value is the wrong unit of analysis. The right unit is what the holder can now do that opponents can’t.

Largest Army → robber control

The 2 victory points are nominal. The real asset is agency over the robber. A player with Largest Army has played enough knights that they have:

The 54% win rate isn’t “Largest Army gives 2 VP.” It’s “Largest Army gives a four-channel intervention mechanism over the production graph.” Roman observes the lift; the lift is the capability rent.

Longest Road → map control

Similarly, the 2 VP from Longest Road is the cheaper reading. The expensive reading is spatial dominance:

The 59% win rate is what map control buys, not what 2 VP buys. The actual VP from Longest Road would be roughly worth 10-15% win-rate by raw point-share math; the other 30+ points of lift is the capability.

Monopoly card → timed shock

Monopoly’s “play one for +15 win rate” effect is essentially capability stacking too:

The 25 → 40 → 55 ladder isn’t “Monopoly cards are good.” It’s “Monopoly cards are a timing weapon whose strike value grows with the game’s accumulated state.” A player who plays it when their opponent has just spent everything is wasting the card.

This is the exact pattern an EBM or planner would discover and a bare LM would miss ([[planner-lm-composites]]): the action is fixed, but its value depends on a global property of the state at the moment of play. Action-bias LMs play the card because it’s playable; state-aware planners hold the card.

The lesson, restated

All three “factors” Roman highlights are the same kind of thing. They’re not levers that grant points. They’re levers that grant the ability to act differently than opponents can. The win-rate lift is option-value rent — the holder’s strategy space is strictly larger than the non-holder’s. That’s the structural advantage Roman names but doesn’t decompose.

This pattern recurs widely:

Catan’s information model — hidden trackable vs hidden random

Roman doesn’t analyze Catan’s information design. It’s worth doing, because the design has a structural inconsistency that affects which strategies pay.

Two kinds of “hidden” coexist in Catan:

What’s hidden How hidden Adversary’s recourse
Resource hands Face down on the table Memory + arithmetic — fully recoverable
Development cards Face down + content unknown until played None — true unpredictability until reveal

Hidden trackable — a soft design flaw

A player’s resource hand is technically face down, but it is fully reconstructible from public information: every dice roll, every settlement/city adjacency, every trade, every steal, every build. A good-memory player can know every other player’s hand at every moment of the game.

This means the “hidden” face-down convention is not actually hiding information — it’s hiding it from players who aren’t bothering to track. Two players at the same table can have radically different information states from the same public game.

That’s a soft design flaw. Skill at the game shouldn’t reduce to skill at bookkeeping. The convention exists mostly to keep play moving — if every trade required players to publicly enumerate hands first, games would drag. So the design accepts an information-asymmetry between memory-good and memory-average players as the price of pace.

The honest reading: resource hiding is a UX optimization, not a strategic mechanic. It looks like hidden information but functionally rewards bookkeeping skill, which is a separate axis from strategic skill.

This shows up in Roman’s data implicitly: the trade-efficiency paradox depends on opponents being unaware that you “overpaid” because the resource state is opaque to them. Against perfect-memory opponents, the paradox might shrink — they’d know exactly what you gained and re-rank the trade. The 8% / overpay finding is partly about exploiting other players’ incomplete bookkeeping, not just about trajectory valuation.

Hidden random — the good design layer

Development cards, in contrast, are genuinely unpredictable:

Opponents can see your roads, settlements, cities, and resource flow. They can model your strategy and react. But they cannot model what cards you’ve drawn or when you’ll play them. This is the surprise channel — and it’s the only piece of Catan that’s robust to a perfect-information opponent.

This is why dev-card strategies tend to win even when straight-development is more raw-efficient per resource invested. The points are roughly equivalent, but dev card points are uninterceptable. A road network gets blocked. A settlement gets robbered. A city gets sandbagged. A face-down VP card sits in your hand until you cross 10 and announce it.

Compounding effects:

The good half / bad half framing

Catan’s information design has two halves that pull in different directions:

This is a useful game-design lens: when a game uses hidden information, ask which kind. Hidden but inferrable is usually a band-aid for play-time. Hidden and random is the genuine strategic mechanic. The two have very different design implications, and players who recognize the difference adjust strategy accordingly — under-investing in tactics opponents can perfectly track, over-investing in the channels they can’t.

The principle generalizes:

The instinct generalizes from Catan: bet on the genuinely-hidden channels, not the bookkeeping-hidden ones. Anything in the second category is a tax on opponents’ attention, not a real strategic asset.

Board as efficient frontier — a prescription theory (sketch)

Chris’s theory of the opening, articulated as the prescription layer Roman didn’t reach. Marked as sketch — the formalism isn’t fully worked out, but the shape is sharp enough to capture and refine.

The thesis

A Catan board, post-random-setup, is a vector space whose structure determines the optimal path to victory. The structure has three observable dimensions:

Together these collapse onto a strategy preference landscape: some build paths are well-supported by this board, others are uphill. The optimal opening isn’t an absolute — it’s conditional on the realized board.

The resource-strategy correspondence

Each resource biases toward specific build paths:

Resource abundance Favored build path Why
Wood + Brick (high pip exposure on both) Road expansion, Longest Road, frontier settlement The road-engine is feasible because the inputs are cheap; aggressive expansion outruns competition
Wheat + Ore (high pip exposure on both) City scaling, dev cards, Largest Army The city-engine compounds; ore+wheat fuels both cities (which double production) and dev cards (which include knights → Largest Army)
Sheep (over-represented) Dev cards + settlement expansion Sheep is the limiting reagent of dev cards (1 sheep per card) and of settlements (1 sheep each); abundant sheep enables the hidden-card strategy
Mixed / balanced Hybrid; flexibility-led No single path dominates; the strategy is adaptability and trade-based correction

Roman’s empirical finding (winners have slightly higher ore/wheat/sheep exposure) is the average of this correspondence over 47k random boards. The interesting claim isn’t “favor city engine in general” — it’s “the strategy should be conditioned on the realized board’s distribution.” A board with abundant wood/brick should produce road-engine winners; a board with abundant ore/wheat should produce city-engine winners.

This is the prescription Roman didn’t deliver: not “favor X” but “the right strategy is a function of the board state — and here’s the function.”

Efficient frontier framing

This shape is the same one the Monopoly project already formalizes:

The Catan equivalent is a frontier curve over (strategy archetype × board configuration):

A player choosing their opening is picking a point on this frontier — and the choice is constrained by which board they’re playing.

This formalism is exactly what would convert Roman’s correlations into a decision procedure. Roman observed that winners had slightly higher ore/wheat/sheep exposure on average; the frontier theory says of course they did — the average board mildly favors the city-engine archetype, but the right move on a brick-heavy board is roads, and on a balanced board is hybrid. The 27% classifier accuracy is what you get when you don’t condition on board structure; conditioning collapses the variance.

Sequential constraint — the frontier shrinks as picks happen

The frontier isn’t static. Snake-order placement creates a stateful auction:

The 24.9% / ~25% / >25% / ~25% win-rate distribution is the empirical signature of this design working: the second-pick advantage in round 2 roughly compensates for the late-pick disadvantage in round 1.

This is the same dynamic as a sequential auction with declining inventory. Each pick is best-responding to the current frontier, not the original. Players who treat opening placement as a static optimization (pick the highest-pip 2:1 ore + wheat regardless) are leaving value on the table; the right play accounts for what the frontier will look like when it’s your next turn.

Common-knowledge competition — the counter-positioning move

Here’s the second-order strategic insight, and the one Roman’s data couldn’t reach without much more sophisticated analysis:

The board is common-knowledge information to all players. Everyone sees the same hex layout, the same number tokens, the same port positions. Therefore:

  1. If the board has a clearly-optimal resource path (say, abundant ore + wheat), every player can see it
  2. Every rational player wants those vertices
  3. Snake order awards them to whoever picks first/last in the snake
  4. Players who don’t get those vertices are forced into the second-best strategy
  5. The popular strategy gets overbid — too many players competing for too few good vertices

The contrarian move: the popular resources are overbid; the less-popular configurations leave open cheaper routes.

If the board obviously favors cities (ore/wheat clustered with high pips), but four players are bidding for those vertices, the third-best ore/wheat vertex might be worse than the best wood/brick vertex on the same board — even though the average ore/wheat opening beats the average wood/brick opening across 47k boards.

This is the same dynamic as crowded trades in finance: the consensus opportunity gets bid up to a point where its risk-adjusted return is lower than the unfashionable alternative. The right play is to best-respond to opponents’ expected bids, not to the nominal frontier.

The strategic implication:

Player’s situation Right play
Picks early (round 1) on a clearly-optimal board Take the obvious vertex; you have the bidding advantage
Picks late (round 1) on a clearly-optimal board The obvious vertices are gone; counter-position toward the second-tier path if its third-best vertex is better than the obvious path’s fifth-best vertex
Picks early on a balanced board The frontier is wide; pick for flexibility and adjacency to ports / future trade partners
Picks late on a balanced board Same flexibility logic; the bidding asymmetry is smaller

This is a Keynesian beauty contest layered on the bilateral trade valuation thesis: you’re not maximizing trade efficiency, you’re maximizing position; and what counts as good position depends on what other players are bidding for.

What this theory would predict (testable claims for Roman’s data)

The frontier-and-bidding theory predicts patterns Roman’s aggregate stats wouldn’t see, but her dataset would:

None of these are proven. They’re the testable form of the theory. Roman’s dataset has the inputs to test all four.

Data access — what’s available and what isn’t (audited 2026-05-18)

The 47k-game dataset Roman analyzed is not publicly downloadable. Twosheep.io collected it on their platform and provided it to Roman directly. The video description has no link; Roman has published no code, GitHub, or blog post associated with the analysis (only an Instagram). Twosheep’s own blog has done smaller public analyses (108 Champs games, 500+ 1v1 games) but never publishes the underlying data.

Source Scale Public? Notes
Twosheep.io 47k dataset (Roman’s data) ~47,000 games No Private; provided to Roman by twosheep directly
Twosheep.io replay browser at twosheep.io/games?to=... All games on platform Replays viewable, not exportable Per-game replay only; no bulk export
Colonist.io game logs 3M / month, 600GB total No Explicitly withheld; community feature request unfulfilled
colonist.inpolen.nl community archive Currently 0 stored Per-game JSON download mentioned Personal use only; bulk scraping IP-banned
Kaggle lumins/settlers-of-catan-games 50 four-player games, 200 rows Yes The standard small public dataset; everyone’s analysis uses this
Kaggle thedevastator/... and koftezz/... Similarly small Yes Alternative small datasets, similar shape

The gap: Roman’s analysis at N=47k is ~1000× larger than the largest publicly downloadable Catan dataset. Her empirical landmarks (24.9% turn-order, 59% Longest Road, 25/40/55 Monopoly card, 8% even-trade win rate) cannot be reproduced or extended at scale without:

  1. Direct outreach — contacting twosheep.io for research-grade data access, or asking Roman to share her code/data. Roman has an ORCID (active research scholar); plausible she’d share post-publication.
  2. DIY collection — Chrome extension scraping Colonist + accepting the legal/IP-ban risk; multi-month accumulation timeline.
  3. Settling for 50 games — the Kaggle dataset is statistically underpowered for the four testable predictions but is enough for proof-of-method work (does the direction of each prediction hold even at small N?).

Strategic takeaway: The frontier-and-bidding theory’s empirical validation is gated by data access that doesn’t currently exist. The theory remains internally coherent and connected to the Monopoly project’s frontier work, where Chris already has access to the simulator and can test analogous predictions. Monopoly is the better proving ground for now — same theoretical structure, full access to the data-generating process.

The exact-vs-frontier architectural distinction

A genealogy point worth recording: the Monopoly project went through two distinct generations of trade-valuation tools, and the lessons map directly onto what a Catan frontier tool would and wouldn’t be.

Generation 1 — Exact bilateral valuation. The bilateral trade valuation page is the trajectory-based exact calculator: simulate both players’ development paths under hypothetical trades, compute who wins what 20 turns later, score the trade against position deltas. It works brilliantly for apples-to-apples trades — “I’ll give you my orange to complete your monopoly, you give me your dark blue to complete mine.” In that situation, both players are completing monopolies, both dominate the rest of the field, and the math closes cleanly. It also was known at the time to say almost nothing about 1:1 trades of non-completion properties — where the strategic value of any single property depends on what else you might acquire, and the exact NPV calculation can’t see that combinatorial structure.

Generation 2 — Efficient frontier subgraphs. The frontier trade theory page is the structural tool: each player has a random portfolio (the properties they hold from random landing), and the frontier subgraph shows the Pareto-optimal expansion paths from that portfolio — which trades or developments would move them toward dominance, which are strictly dominated. The frontier doesn’t give exact valuations; it gives strategic-value visibility (“this property is on my critical path”; “that property is dominated for me but on Player B’s frontier — so they’ll trade hard for it”) and eliminates frivolous exchanges by structural truth rather than enumeration. This is discrete-game Modern Portfolio Theory — the Markowitz/CAPM Pareto-frontier construction adapted from continuous risk-return space to a discrete game’s state space. The combinatorial-bloat reduction works for the same reason MPT works in continuous space: structural dominance means most options are strictly worse than other options on the same axes and can be pruned without evaluation. See the frontier page’s methodological lineage section.

The two tools are complementary, not competing:

Tool What it does When to use
Exact bilateral simulation Precise NPV-delta on a specific trade between two specific players Apples-to-apples completion trades; bidding decisions on auctioned properties; high-stakes single-move evaluations
Frontier subgraph Map of reachable strategic positions and their dominance relations Long-horizon planning; deciding which trades to even consider; identifying counterparty leverage points

The bilateral evaluator says “here’s exactly what this trade is worth.” The frontier subgraph says “here are the trades worth thinking about, here’s what each player values, and here’s the shape of the reachable strategy space.”

What a Catan frontier graph would look like

The Catan equivalent of the Monopoly frontier subgraph would map possible expansion paths from each player’s current state — which is closer to the right shape for Catan than a bilateral evaluator would be (Catan has fewer crisp apples-to-apples trades than Monopoly’s monopoly-completion structure).

Inputs the Catan frontier tool would need (and what we have / lack from the 50-game dataset):

Input Per-game cost to compute Have in 50-game dataset?
Full hex layout (19 tiles with number tokens and resource types) Read from game setup ❌ Only the tiles a player’s settlements touched
Port layout (9 ports around the coast) Read from game setup ⚠️ Partial — only ports adjacent to player settlements
Current settlements and cities for each player Read from game state ✅ Starting only; no mid-game expansion data
Current road network for each player Read from game state ❌ Not in dataset
Each player’s current resource hand Read from game state ❌ Not in dataset

The 50-game dataset is not enough to build the frontier tool — it has starting positions but no mid-game state, and no full board layout. But the algorithmic shape can be specified independently:

The frontier algorithm (sketch):

  1. For each player, compute reachable expansion vertices. A vertex is reachable if you have a road network connecting to it, or if you have the resources + road-build chain to extend to it.
  2. For each reachable vertex, compute the build cost. Number of roads + 1 settlement = (R × brick + R × wood) + (1 brick + 1 wood + 1 wheat + 1 sheep), where R is the number of new roads needed.
  3. For each reachable vertex, compute the production gain. Sum of pip-weights of the three adjacent hexes for the new settlement (plus city upgrades available there).
  4. Pareto-filter the reachable positions. A position is on the frontier if no other reachable position has both lower cost AND higher production gain.
  5. Frontier subgraph output: the set of (position, cost, gain) triples plus the build-order constraints (you need road A before road B before settlement C).

What the frontier graph shows:

What the frontier graph doesn’t show (and intentionally doesn’t):

The “you need X, Y, Z to get here, but the graph doesn’t know how to get X, Y, Z” structure is the same as Monopoly’s frontier subgraph. Both tools answer “what’s the destination worth pursuing” and leave “how to procure the inputs” to other components. In Monopoly the procurement is mostly trade; in Catan it’s trade + production timing + dice + robber play. The split keeps both tools focused on what they’re good at.

Why this matters for the theory line

The whole board-frontier theory sketch earlier in this page can now be more precisely typed:

The vault’s two Monopoly-project tool generations map directly onto two complementary Catan tools. We don’t have the data to build either yet, but the architectural shape is now specified.

Connection back to vault frameworks

This is the Monopoly project’s frontier methodology applied to Catan opening placement — same shape, different game. The instinct generalizes:

The Catan theory and the Monopoly theory are two instances of the same underlying frontier-and-bidding analysis. Formalizing one helps formalize the other; both are open vault threads worth continuing.

Rubber-banding and the skill-cap tension

Catan has very little rubber-banding by design. The main lever is the robber — when 7 is rolled, every player holding more than 7 cards must discard half, and the rolling player may move the robber to a specific tile, blocking production and stealing a card from one adjacent player. This is mild anti-hoarding pressure aimed loosely at the leader: high-producers tend to accumulate more cards, so the discard-on-7 rule strips them disproportionately. It’s transparent, symmetric (applies to anyone over 7 cards), and rule-level — exactly the good-rubber-banding profile the M.U.L.E. analysis named ([[feedback_rubber_banding_and_friction_patterns]]).

Monopoly has effectively zero rubber-banding. No mechanism redistributes from leader to laggards; the snowball just snowballs. Once one player owns three monopolies and the others own none, the game is structurally over — even if it takes 20 more turns to finalize. This is one reason Monopoly is widely disliked despite its iconic status: the loser-player experience drags long after the outcome is determined.

The design tension

The case for more rubber-banding (egalitarian dynamics):

The case for less rubber-banding (skill-rewarding dynamics):

Too little rubber-banding → runaway games where the early-game determines the late-game; losing players have nothing to do. Too much rubber-banding → all players bunched at the finish line, and the winner is whoever rolls the right number on the last turn. Both extremes collapse to a less interesting game, just in different directions:

Rubber-banding level Failure mode Skill cap Example
None Runaway snowball, dead late-game for laggards Maximum Chess, Monopoly
Mild Lead is defensible but pressure exists High Catan (robber), M.U.L.E. (auction + events), Power Grid
Heavy Lead is hard to keep; the last turn matters most Reduced Mario Kart with blue shells
Total Everyone tied at the end; dice/luck decides Floor (≈ random) Some racing-game catch-up AI; “everyone gets a trophy”

The interesting question isn’t “is rubber-banding good or bad” — it’s where on this axis is the game designed to land, and is the choice intentional? Bunten’s M.U.L.E. (1983) made a clear deliberate choice for mild, transparent, rule-level rubber-banding — that’s the design master move the M.U.L.E. project documented. Catan made a similar choice with the robber. Monopoly made the opposite choice (no rubber-banding) — defensible at the time of design, but the snowball-then-finish-it problem is what makes modern players prefer game-night Catan.

Skill cap as a design dial

A subtler version of the tension: rubber-banding caps the skill ceiling. When a leader can be reliably dogpiled or random-evented back to the pack, the skill premium of “playing well from a good position” collapses. Skilled players cannot translate their lead into reliable victory; the game becomes about reading rubber-band moments rather than about executing strategy. This is fun for newcomers but disincentivizes investment in skill development.

Chess at the extreme: grandmaster vs. amateur is a deterministic outcome. The skill premium is total. This is why chess is the canonical competitive game and not a popular casual one.

Most well-designed Eurogames (Catan, Terra Mystica, Brass) sit in the mild rubber-banding zone — enough leveling to keep games interesting without flattening skill. The trick is that the leveling is at the rule level (everyone faces the same discards-on-7, the same auction structure) rather than at the mechanic level (some hidden code targets the leader). Rule-level rubber-banding preserves competitive integrity; mechanic-level rubber-banding feels like the game cheating against the leader.

Why this matters for AI play

A planner-LM composite or an EBM playing a rubber-banding-aware game needs to model how much its lead is worth. In Chess, a +2 advantage is worth ~certainty. In Catan with the robber, a +2 VP advantage is worth maybe 60% — opponents will robber-target you, dogpile via trade refusal, and the dev-card deck remains a wild card. A naive Catan AI that maxes for instantaneous expected VPs will systematically over-value early leads and under-invest in the defensive moves that protect them from rubber-band reversal. Skill-cap calibration is part of opponent modeling.

N≥3 termination — interference completeness is the real mechanism

The Multiplayer Coalition Problem page already documents the three-player stable state: in any game where N≥3 and players can interfere with each other’s development, weaker players gang up on the leader. The leader gets restricted; a new leader emerges; the cycle repeats. No coordination required — each weaker player independently arrives at the same conclusion (the leader is the biggest threat to my survival).

What makes the cycle terminate — that’s where games differ structurally. The first reading of this thread put weight on hidden-random information as the kill switch. The sharper reading puts weight on interference completeness: can the dogpile actually stop progress, or only slow it?

The two axes

Interference completeness:

Hidden-random information:

The interaction matrix for whether N≥3 games terminate under ideal play:

Interference Hidden-random Terminates under ideal play? Example
Complete None ❌ Never — perfect-AI dogpile is stable Hypothetical perfect-info Risk
Complete Present ⚠️ Only via persuasion / asymmetry exploit; fails vs. perfect AI Risk, Diplomacy
Incomplete None ✅ Yes — un-suppressible progress accumulates Hypothetical no-dev-card Catan
Incomplete Present ✅ Yes — production stream forces it; hidden-random adds timing Catan

The hidden-random kill switch helps games like Risk almost terminate, but it’s not enough — Chris’s sharpest point: in Risk, high-level players treat 3+ cards AS-IF the player already has the turn-in armies. That converts the hidden-random into common knowledge, neutralizing it. Good Risk players assume worst-case and plan accordingly, which collapses the asymmetry and restores the perfect-information dogpile. The card turn-in surprises a beginner; it doesn’t surprise a planner.

The same move works in Diplomacy. Good players treat ambiguous-intent moves as if they’re the worst plausible interpretation. Hidden-random becomes assumed-worst, and the dogpile re-stabilizes.

This is why Risk and Diplomacy at high level require out-of-game persuasion to terminate. The strategic stalemate is broken not by mechanical state but by social negotiation — convincing an opponent to defect from the optimal dogpile coalition. Diplomacy is the genre name for this and the literal mechanic of the game. Risk’s endgame is full of meta-talk (“if you attack me I’ll throw the game to him”) because the on-board state cannot resolve the equilibrium by itself.

Against a perfect AI that doesn’t believe out-of-game promises, neither Risk nor Diplomacy would end. This is the structural reason Cicero needed both a planner AND a language model — the language model is what does the persuasion the planner can’t fake. [[cicero-press-diplomacy-captain-meme]] and [[Gunboat Diplomacy and Diplodocus]] together bracket this exact result: Cicero needs the language layer to handle the persuasion side; Diplodocus (pure planner, no language) wins only in Gunboat where persuasion isn’t allowed.

Catan is structurally different — and that’s the design genius

Catan’s interference is incomplete in a specific way: the production stream. Every turn, dice are rolled, and every adjacent settlement/city produces resources. You can rob one tile. You can refuse to trade. You can play knights to displace the robber to a defender’s tile. But you cannot stop the dice from coming up 6 and 8, and as long as a player has high-pip settlements, they will accumulate resources regardless of what their opponents do.

This is why Chris’s sharpest claim holds: “In ideal play, Risk will never end, but Catan will always end.”

The dev cards add a timing element — they let a particular player end the game sooner than visible-state accumulation alone would — but they’re not the structural termination mechanism. The structural mechanism is the un-suppressible production stream. Even in a hypothetical no-dev-card Catan, games would still end; they’d just end slower, with the winner being whichever player’s pip exposure best survived the dogpile-throttle.

The “feels like chance” complaint is correctly diagnosed once you separate the two factors:

Both elements are partly luck. But the structural fact remains: Catan ends because production keeps happening, not because dev cards eventually reveal. The hidden-random is the timing layer; the un-suppressible production is the engine.

Game-class comparison

Game N Interference Hidden-random Ends under perfect AI? Termination mechanism
Chess 2 Complete None Yes (zero-sum) Direct
Risk 3-6 Complete Card turn-ins No — needs persuasion Out-of-game diplomacy
Diplomacy 7 Complete Hidden orders No — needs persuasion Out-of-game diplomacy
Coup 3-6 Complete (player elimination) Hidden roles Partly — hidden roles + bluffing Bluff resolution
Poker (multi-way) 3+ Incomplete (you can only fold, can’t take their stack except by winning hands) Hole cards Yes Forced bet-stream + reveal
Catan 3-4 Incomplete (production stream) Dev cards Yes Production accumulation
Monopoly 2-8 Incomplete (no direct attack; only refuse-trade and cash drain) None significant Yes (snowball) Cash drain to zero
MOO1 (mp) 3+ Mostly complete (fleets can kill) Hidden tech/fleets Partly — hidden tech enables breakouts Tech breakthrough + military

The pattern: incomplete interference is what makes a multiplayer game tractable for perfect AI. Complete-interference games (Risk, Diplomacy) require the language/persuasion layer to terminate at all — which is exactly why Cicero needed an LM and why [[planner-lm-composites]] are the right architecture for that game class.

Catan, Poker, and Monopoly terminate without persuasion because their interference is structurally incomplete. The dice / cards / cash flow can’t be fully throttled.

Why this matters

For AI design: the right architecture depends on whether the game is interference-complete. A pure planner (Diplodocus-style) works in incomplete-interference games where mechanical state dominates. A planner-LM composite (Cicero-style) is required in complete-interference games where the language layer is the termination mechanism.

For game design: interference completeness is a more load-bearing knob than hidden information. A designer choosing N≥3 has to decide whether their game will terminate via mechanical accumulation (incomplete interference, like Catan) or via social negotiation (complete interference, like Risk/Diplomacy). The first works for casual / family / closed-system contexts; the second requires players who treat persuasion as a legitimate game move and willingly defect from optimal coalitions.

For the rubber-banding analysis above: the rubber-banding dial interacts with interference completeness. Heavy rubber-banding on top of complete interference produces deadlock (everyone bunched, can’t escape). Mild rubber-banding on top of incomplete interference produces the Catan zone — skill-driven progress with leveling pressure but guaranteed termination. The design space isn’t one-dimensional.

The vault claim: Chris’s “Risk will never end, Catan will always end” is the cleanest one-line summary of the difference between these two N≥3 design families. It deserves to be stated explicitly: interference completeness, not hidden information, is the structural termination determinant in N≥3 multiplayer games.

How complete-interference games actually terminate — bounded rationality

The perfect-AI result says Risk and Diplomacy should never end. In practice they end every weekend. What gives?

The answer is bounded human rationality, and Risk’s two card-variant designs make this visible:

Fixed turn-ins (4-6-8-10-style, or fixed army counts per card type) — the turn-in value is bounded and predictable. Players can apply Chris’s AS-IF rule cleanly: 3+ cards = treat as if the player already has roughly N armies coming. The hidden-random asymmetry collapses to common-knowledge worst-case, and the dogpile re-stabilizes. Fixed turn-in Risk preserves the deadlock, which is why high-level fixed-turn-in games end via persuasion / kingmaker plays rather than via mechanical state.

Escalation turn-ins — values grow linearly but unboundedly. Typical schedule is +5 per turn-in (e.g., …, 20, 25, 30, 35, 40, 45, 50, …). It’s not exponential; each step is small. But there’s no cap. After ~200 turn-ins in a long game, the value is ~1000 armies (by which point all players should have 2000+ total armies on the board). Mechanically this is just scale, and a perfect AI would track it perfectly. But humans systematically under-update their threat models for two reasons:

  1. Anchoring on the recent past, not the absolute scale. Each +5 looks small in context — 30 → 35 is a 17% increase, feels like a tweak. But the cumulative effect over 100 turn-ins is a 500-army jump, which is several times the typical reinforcement budget. Players normalize to the recent turn-in value rather than re-anchoring on absolute board state.
  2. Long-game distraction and cognitive fatigue. Escalation games run long; by turn 400+, players have spent hours tracking state across the board, and attention to single-channel threats (card-holders) degrades. The 1000-army-card-holder problem isn’t a math problem — it’s a bandwidth problem.

These aren’t the same as the “exponential exceeds cognition” failure I initially described — they’re closer to anchoring bias + base-rate neglect + attention erosion over long games. The math itself is trivially linear; the cognitive failure is in re-pricing the cumulative.

The termination mechanism in escalation Risk is human cognitive failure to re-anchor, not strategy. A perfect AI playing escalation Risk would dogpile card-holders much more aggressively than humans do, because the AI would correctly anchor on the current turn-in value rather than the early-game one. Against perfect AI, escalation Risk would also stall — just at a different scale than fixed Risk. The variant doesn’t change the structural result; it changes how easily humans fall out of the deadlock.

Why this matters

Risk and Diplomacy are designed for humans, not perfect agents. The “tractable termination via persuasion” mechanism and the “tractable termination via cognitive-scaling failure” mechanism are both bounded-rationality exploits. Strip the bound and the games revert to their structural form: indefinite deadlock.

This connects to several existing vault threads:

Generalizable lens: when a multiplayer game terminates under perfect-AI theory it terminates “structurally” (production stream, forced bets, etc.). When it terminates only under bounded-rationality real players it terminates “experientially.” Most successful complete-interference multiplayer games are experiential — Risk, Diplomacy, Coup, Werewolf, social-deduction generally. They are designed for the bound, not despite it. Strip the human and you strip the game.

This is one reason game-AI for these titles is genuinely hard: the AI must model not just optimal play but the opponents’ rationality bound. Cicero’s LM is doing exactly that — modeling what humans will believe and feel persuaded by. Diplodocus avoids the problem by playing Gunboat (no persuasion channel) — but that’s a deliberately bounded variant where the rationality-failure exploit isn’t available.

A third bounded-rationality termination mode — the emotional leveler

The two mechanisms above (persuasion-for-defection, anchoring + fatigue) don’t exhaust the failure modes humans bring to complete-interference games. A third mode worth naming is emotional leveling — players who break the dogpile equilibrium not because the math says they should, and not because their threat models are mis-anchored, but because they have an aesthetic / emotional preference for board balance that overrides strategy.

A worked example from a real play group: a player who didn’t like it when army counts got “too high” would, on a whim, systematically attack across multiple opponents to “even things out.” Mechanically suboptimal — he exhausted his own troops across several players, weakening himself. But the result was exactly what breaks the deadlock: he created a new asymmetric state where the next player’s turn-in produced overwhelming advantage against a specific weakened target, often ending the game in 1-2 turns. The leveler player became an unintentional kingmaker for the next-up player.

This is a distinct bounded-rationality failure from the others:

Failure mode Trigger Mechanism Result
Persuasion-for-defection (fixed Risk, Diplomacy) Out-of-game social pressure Player abandons mathematically-optimal coalition for personal/social reasons Defector becomes kingmaker
Anchoring + fatigue (escalation Risk) Long-game cumulative scale Player under-updates threat model for late-game card values Card-holder accumulates uninterceptable advantage
Emotional leveling (any complete-interference game) Aesthetic discomfort with imbalance Player burns own position to “even” perceived inequality Creates exploitable asymmetry for next-up player

The Diplomacy parallel: there’s a documented “longest tournament Diplomacy game ever” video series — high-tournament Diplomacy genuinely does stall for many hours / multiple game-sessions before terminating. The empirical signature of “perfect play → infinite” shows up exactly as you’d predict: among the skillest available humans (who don’t make obvious anchoring or emotional-leveling mistakes), games trend toward indefinite. Termination requires some opponent to make some sub-optimal move — and the longer the tournament, the more reliably someone eventually does.

The structural conclusion sharpens: complete-interference games terminate in practice because real players have multiple independent bounded-rationality failures any of which can break the deadlock. Persuasion, anchoring, emotional leveling, plus boredom, ego-driven plays, alliance-as-friendship-test, etc. The game design doesn’t need to exploit any specific bound — it just needs the set of human bounds to be large enough that, given N≥3 players over enough turns, at least one will fire. Multiple-independent-bounds redundancy is how human play terminates structurally-infinite games reliably.

This is also the empirical answer to “why hasn’t a perfect Diplomacy AI just dominated the field by being unexploitable” — the answer is that there are some games (the longest ones) where the AI just would not finish, and human tournament play would beat it on time-budget alone. Cicero handled this by being able to cause the bounded-rationality failures (via persuasion) rather than just exploiting ones that happened to occur — which is the structural reason its language model is doing real strategic work, not just performative chatter.

Where the randomness lives — Monopoly vs. Catan as a structural contrast

Both games are resource-allocation problems with random elements; both are vault research threads (Monopoly with deep work, Catan now sketched). But they are mathematically different kinds of problems, and naming the difference clarifies why the same frontier framework looks different in each, and why Catan was groundbreaking when it shipped.

The core distinction

Where in the game does the randomness actually live?

Game Random axis Deterministic axis What the player optimizes
Monopoly Access (dice rolls determine which fixed properties you land on, when) Board configuration (same every game) Convert randomly-dealt holdings into a winning portfolio via trade and development
Catan Terrain (hex layout + number tokens + ports are different every setup) Play sequence once board is set (snake-order + standard turn structure) Read the board and commit to the strategy archetype the terrain supports

This is two different problems sharing the same resource-allocation skeleton:

Chris’s compression

“Monopoly’s efficient frontier is how to get from your random slice of the pie to a winning one. Catan is what are the high-probability resources that will be in play so I know what to play for or around.”

That’s the whole structural difference in one line. Both are “frontier” problems, but the frontiers operate at different time scales:

Why Monopoly is Markov-friendly and Catan isn’t

This explains why the Monopoly project leans heavily on Markov chains and Catan-style analysis doesn’t:

Catan trades one kind of math (Markov-on-time) for another (combinatorics-on-configuration plus per-board state-conditional analysis). It’s actually a harder mathematical problem in the sense that it requires conditioning on the board archetype before any inference is valid. Roman’s aggregate 47k-game stats blur exactly this conditioning, which is why her findings feel descriptive rather than prescriptive — she averaged across board-archetypes that demanded different strategies.

Catan’s modular board as a game design milestone

The modular hex-tile setup was Catan’s first-major-mainstream shift to a different design axis when it shipped in 1995. Modular-board games did exist before Catan in niche / hobbyist circles — Avalon Hill’s Galactic Conquest and other lesser-known titles experimented with variable setup — but the dominant 20th-century mass-market resource-allocation games (Monopoly, Risk, Axis & Allies, Diplomacy) all had fixed boards, and the design instinct was “fix the geometry, randomize the play through it.” Catan was the first crossover hit to reverse this at scale: randomize the geometry, deterministic the play through it. The credit isn’t strict invention; it’s the move from niche-hobby experiment to mainstream-defining standard.

Game-design implications:

Catan-style modular setup spread across the hobby after 1995 — see virtually every modern Eurogame (Terra Mystica, Scythe, Wingspan, Brass: Birmingham). Roman’s analysis is best read as one of the first systematic empirical studies of the modular-board class of games rather than just of Catan specifically.

Why this routes back to Monopoly first

The board-frontier theory’s predictions need a data-generating process to validate against. The vault has full access to one for Monopoly (the simulator) and zero for Catan (Roman’s 47k is private; Kaggle is 50 games). The structural contrast above explains why this is fine rather than a compromise:

Cleaner framing: Monopoly and Catan are not competitor analyses — they’re complementary instances of the same frontier theory at different randomization layers. Validating one strengthens the other.

Open questions the data could answer (but the video didn’t)

Tags

games, strategy, game-theory, mathematics, simulation, economics