Ioana Roman analyzed 47,000 recorded Catan games from twosheep.io and turned the board game into a stochastic dynamic system with real numbers attached. Headline findings: turn order is essentially balanced, opening placement predicts the winner barely above random (27% vs 25% baseline), Longest Road and Largest Army are symptoms not causes of victory, missing a resource barely matters because the market corrects, Monopoly cards are gigantic timing weapons (one played → 40% win rate; two played → 55%), and winners systematically overpay on trades. The last one is the empirical confirmation, across 47k games, of the bilateral trade valuation thesis: trade evaluation must be trajectory-based, not isolated NPV. Important caveat: despite the title, the video delivers descriptive statistics about how games tend to play out, not a strategic prescription for how to win. The empirical landmarks are real and useful as constraints, but the prescription layer — state-conditional opening policies, timing policies, trajectory valuation — is the work the Monopoly project does and Roman doesn’t.
Links: Gaming, LLM Agents Across Strategic Games, LLM Game Benchmark — Outline, Bilateral Trade Valuation, The Multiplayer Coalition Problem, BattleValue, Planner-LM Composites, The LLM Grounding Problem, Monopoly Theory, Energy-Based Models, Catan 50-Game Validation — proof-of-method run against the public 50-game Kaggle dataset; aggregate city-engine direction confirmed, board-conditional refinement only partially supported, theory’s working form updated to “universal city-engine bias + layered archetype effect”, Randomness as Termination (N≥3) — Catan’s dice + dev cards as the parity-breaker / termination layer
Primary source: I Analysed 47,000 Games of Catan. Here’s How to Win Every Time (mathematically) — Ioana Roman, 2026-05-10. Transcript. Data: twosheep.io.
Roman frames Catan as a multi-agent stochastic dynamic optimization problem. The skeleton:
The framing matters: this is “probabilistic income streams + endogenous market dynamics + asymmetric information + stopping-time decision,” which is much closer to the bilateral trade valuation and Monopoly project framings than to anything that “luck vs skill” captures. Catan is a small economy under uncertainty.
| Position | Win rate |
|---|---|
| 1st | 24.9% |
| 2nd | ~25% |
| 3rd | slightly >25% |
| 4th | ~25% |
The snake-order placement (last player places twice before order reverses) genuinely works. There’s no exploitable structural disadvantage to seating. If you lost from 4th, it wasn’t the reason.
These are large lifts over baseline, but the holder still loses 40% / 46% of the time. Roman’s read: they reflect underlying structural advantage rather than create it. Longest Road tracks spatial dominance over the graph; Largest Army tracks ore/wheat investment plus robber control. Both are downstream of the actual strategic position.
This is the structural-realism pattern ([[BattleValue]], [[D&D Spell Damage Model]]): the headline number is a proxy. The thing that produces both the bonus and the win is the underlying configuration.
Each number token has a weight proportional to its dice-roll probability. Summing the weights of a settlement’s three adjacent tiles gives a scalar expected production intensity.
Winners had slightly higher exposure to:
(combined: the “city engine” — what powers cities and development cards)
Winners and losers had nearly identical exposure to:
(the “road engine”)
The finding is small per game but consistent across 47k. Catan is not fundamentally a road game; it’s a compounding city game. Cities double production, development cards provide hidden leverage and Largest Army eligibility, and the late-game acceleration runs on ore+wheat+sheep. The road-heavy intuition that beginners get from “Longest Road feels decisive” is misleading at the strategy-selection layer.
Across 47k games, missing any single resource at the opening leaves win rate at ~25% with small dips for missing ore or wood. The trading and port system is enough to correct the imbalance.
This is self-balancing market behavior, the same pattern that makes N≥3 multiplayer coalitions game-theoretically interesting. The game has built-in correction mechanisms; outcomes don’t collapse on one bad draw. Resilience comes from the trade graph, not from the production graph.
Roman trained a classifier on initial-placement features (pip exposure, resource diversity, balance metrics), tested on 20% held-out games. Accuracy: 27%. Random baseline: 25%.
The opening matters by two percentage points.
The other 73% of outcome variance lives in the post-opening sequence: trades, development card draws, robber placements, timing, interaction. Catan is path-dependent, not predetermined at turn zero. This is the empirical answer to “where does the game actually happen?” — not in the opening, but in the running game.
| Monopoly cards played | Win rate |
|---|---|
| 0 | 25% |
| 1 | 40% |
| 2 (max) | 55% |
Roman’s framing: “systemic liquidity extraction.” A late-game Monopoly when opponents have accumulated inventory is a much bigger shock than an early one — the card redistributes the inventory state, and the larger the state, the bigger the move.
Two architectural implications:
Trade efficiency = resources received / resources given.
| Trade efficiency | Win rate |
|---|---|
| ~1.0 (perfectly even) | 8% |
| > 1.0 (favorable) | (not the winners) |
| < 1.0 (overpaying) | (where the winners cluster) |
Players who make perfectly fair trades win at 8% — three times worse than baseline. Players who systematically appear to overpay are the ones who win.
The mechanism Roman names: they’re not optimizing the trade; they’re optimizing their position. Overpaying to complete a city before a key roll. Overpaying to secure Longest Road before someone else gets it. Overpaying to block an opponent’s expansion. Short-term inefficiency for long-term positional gain.
This is the empirical confirmation of bilateral trade valuation across 47k games. That page argued: a trade’s value is where both players end up 20 turns later, not the static resource delta. Roman’s data is the smoking gun. Players treating trade as a transaction-level optimization (even ratio = “fair”) lose almost three times the rate of players treating it as a position-level move.
The pattern generalizes well past Catan:
The principle is one layer above game theory: valuation requires a trajectory, not a snapshot.
Roman’s title promises “How to Win Every Time (mathematically).” The video doesn’t deliver that. What it delivers is descriptive statistics about how games tend to play out. That’s a different (and lesser) artifact than strategic prescription. Worth naming the distinction because the vault’s Monopoly project lives on the other side of this line, and the comparison clarifies what an actual Catan strategy paper would have to add.
What Roman did (the empirical / descriptive layer):
What she didn’t do (the prescription / policy layer):
The gap is the difference between “winners tend to have X” and “given the current state, do Y”. Roman is honest about this implicitly (her language is “structural advantage,” “tendencies,” “compounding”) but doesn’t bridge to the second layer. Headline tendencies → strategy is a separate exercise the video didn’t take on.
The Monopoly project has the pattern. For Monopoly:
The Catan equivalents would be:
| Empirical finding (Roman) | What the prescription layer would add |
|---|---|
| Winners favor ore/wheat/sheep | A placement frontier over (production weight, resource diversity, port adjacency, blocking value) with Pareto-optimal openings labeled by board configuration |
| Monopoly cards lift win rate to 40%/55% | A state-conditional play policy: hold Monopoly until opponents’ aggregate inventory of resource R exceeds threshold T(R, turn, opponents’ VP) |
| Winners overpay for position | A trajectory-valuation function — when offered trade T at state S, evaluate ΔV(state at turn S+20 with T) − ΔV(state at S+20 without T) |
| Opening predicts 27% | Implies the running game dominates outcome variance; that’s where the planner work needs to go (state-tracking, opponent modeling, robber timing — none of which Roman touches) |
This is the same shape as Chris’s frontier trade theory for Monopoly: the empirical observation that “winners overpay” becomes the prescriptive question “by how much, for which positions, against which opponents.” Roman’s data has the inputs; the optimization formalism is missing.
Roman models per-player state as a 5D resource vector evolving over time. That formalism is one Pareto-frontier away from a strategy theory — the vault’s Monopoly work already lives in that adjacent space. Once you have state vectors, you can:
Roman has the vectors but stops at descriptive statistics over them. The prescription layer is what to do with the vectors — and the Monopoly project shows the methodology transfers cleanly.
The honest verdict on the video: good empirical landmarks, good intuition-builders, useful as constraints for any future prescriptive theory (a Catan strategy paper that contradicts the 27% classifier accuracy is wrong). But the title overpromises. “How to win every time” would require the prescription layer the video doesn’t build.
Roman’s tokens are sharp because she stops at correlation. The vault-side reading of what those correlations mean is sharper, and it changes the strategy implications.
The general principle: Longest Road, Largest Army, and the Monopoly card don’t lift win rate because of the points they grant. They lift win rate because they unlock or signal strategic capability.
The face value is the wrong unit of analysis. The right unit is what the holder can now do that opponents can’t.
The 2 victory points are nominal. The real asset is agency over the robber. A player with Largest Army has played enough knights that they have:
The 54% win rate isn’t “Largest Army gives 2 VP.” It’s “Largest Army gives a four-channel intervention mechanism over the production graph.” Roman observes the lift; the lift is the capability rent.
Similarly, the 2 VP from Longest Road is the cheaper reading. The expensive reading is spatial dominance:
The 59% win rate is what map control buys, not what 2 VP buys. The actual VP from Longest Road would be roughly worth 10-15% win-rate by raw point-share math; the other 30+ points of lift is the capability.
Monopoly’s “play one for +15 win rate” effect is essentially capability stacking too:
The 25 → 40 → 55 ladder isn’t “Monopoly cards are good.” It’s “Monopoly cards are a timing weapon whose strike value grows with the game’s accumulated state.” A player who plays it when their opponent has just spent everything is wasting the card.
This is the exact pattern an EBM or planner would discover and a bare LM would miss ([[planner-lm-composites]]): the action is fixed, but its value depends on a global property of the state at the moment of play. Action-bias LMs play the card because it’s playable; state-aware planners hold the card.
All three “factors” Roman highlights are the same kind of thing. They’re not levers that grant points. They’re levers that grant the ability to act differently than opponents can. The win-rate lift is option-value rent — the holder’s strategy space is strictly larger than the non-holder’s. That’s the structural advantage Roman names but doesn’t decompose.
This pattern recurs widely:
Roman doesn’t analyze Catan’s information design. It’s worth doing, because the design has a structural inconsistency that affects which strategies pay.
Two kinds of “hidden” coexist in Catan:
| What’s hidden | How hidden | Adversary’s recourse |
|---|---|---|
| Resource hands | Face down on the table | Memory + arithmetic — fully recoverable |
| Development cards | Face down + content unknown until played | None — true unpredictability until reveal |
A player’s resource hand is technically face down, but it is fully reconstructible from public information: every dice roll, every settlement/city adjacency, every trade, every steal, every build. A good-memory player can know every other player’s hand at every moment of the game.
This means the “hidden” face-down convention is not actually hiding information — it’s hiding it from players who aren’t bothering to track. Two players at the same table can have radically different information states from the same public game.
That’s a soft design flaw. Skill at the game shouldn’t reduce to skill at bookkeeping. The convention exists mostly to keep play moving — if every trade required players to publicly enumerate hands first, games would drag. So the design accepts an information-asymmetry between memory-good and memory-average players as the price of pace.
The honest reading: resource hiding is a UX optimization, not a strategic mechanic. It looks like hidden information but functionally rewards bookkeeping skill, which is a separate axis from strategic skill.
This shows up in Roman’s data implicitly: the trade-efficiency paradox depends on opponents being unaware that you “overpaid” because the resource state is opaque to them. Against perfect-memory opponents, the paradox might shrink — they’d know exactly what you gained and re-rank the trade. The 8% / overpay finding is partly about exploiting other players’ incomplete bookkeeping, not just about trajectory valuation.
Development cards, in contrast, are genuinely unpredictable:
Opponents can see your roads, settlements, cities, and resource flow. They can model your strategy and react. But they cannot model what cards you’ve drawn or when you’ll play them. This is the surprise channel — and it’s the only piece of Catan that’s robust to a perfect-information opponent.
This is why dev-card strategies tend to win even when straight-development is more raw-efficient per resource invested. The points are roughly equivalent, but dev card points are uninterceptable. A road network gets blocked. A settlement gets robbered. A city gets sandbagged. A face-down VP card sits in your hand until you cross 10 and announce it.
Compounding effects:
Catan’s information design has two halves that pull in different directions:
This is a useful game-design lens: when a game uses hidden information, ask which kind. Hidden but inferrable is usually a band-aid for play-time. Hidden and random is the genuine strategic mechanic. The two have very different design implications, and players who recognize the difference adjust strategy accordingly — under-investing in tactics opponents can perfectly track, over-investing in the channels they can’t.
The principle generalizes:
The instinct generalizes from Catan: bet on the genuinely-hidden channels, not the bookkeeping-hidden ones. Anything in the second category is a tax on opponents’ attention, not a real strategic asset.
Chris’s theory of the opening, articulated as the prescription layer Roman didn’t reach. Marked as sketch — the formalism isn’t fully worked out, but the shape is sharp enough to capture and refine.
A Catan board, post-random-setup, is a vector space whose structure determines the optimal path to victory. The structure has three observable dimensions:
Together these collapse onto a strategy preference landscape: some build paths are well-supported by this board, others are uphill. The optimal opening isn’t an absolute — it’s conditional on the realized board.
Each resource biases toward specific build paths:
| Resource abundance | Favored build path | Why |
|---|---|---|
| Wood + Brick (high pip exposure on both) | Road expansion, Longest Road, frontier settlement | The road-engine is feasible because the inputs are cheap; aggressive expansion outruns competition |
| Wheat + Ore (high pip exposure on both) | City scaling, dev cards, Largest Army | The city-engine compounds; ore+wheat fuels both cities (which double production) and dev cards (which include knights → Largest Army) |
| Sheep (over-represented) | Dev cards + settlement expansion | Sheep is the limiting reagent of dev cards (1 sheep per card) and of settlements (1 sheep each); abundant sheep enables the hidden-card strategy |
| Mixed / balanced | Hybrid; flexibility-led | No single path dominates; the strategy is adaptability and trade-based correction |
Roman’s empirical finding (winners have slightly higher ore/wheat/sheep exposure) is the average of this correspondence over 47k random boards. The interesting claim isn’t “favor city engine in general” — it’s “the strategy should be conditioned on the realized board’s distribution.” A board with abundant wood/brick should produce road-engine winners; a board with abundant ore/wheat should produce city-engine winners.
This is the prescription Roman didn’t deliver: not “favor X” but “the right strategy is a function of the board state — and here’s the function.”
This shape is the same one the Monopoly project already formalizes:
The Catan equivalent is a frontier curve over (strategy archetype × board configuration):
A player choosing their opening is picking a point on this frontier — and the choice is constrained by which board they’re playing.
This formalism is exactly what would convert Roman’s correlations into a decision procedure. Roman observed that winners had slightly higher ore/wheat/sheep exposure on average; the frontier theory says of course they did — the average board mildly favors the city-engine archetype, but the right move on a brick-heavy board is roads, and on a balanced board is hybrid. The 27% classifier accuracy is what you get when you don’t condition on board structure; conditioning collapses the variance.
The frontier isn’t static. Snake-order placement creates a stateful auction:
The 24.9% / ~25% / >25% / ~25% win-rate distribution is the empirical signature of this design working: the second-pick advantage in round 2 roughly compensates for the late-pick disadvantage in round 1.
This is the same dynamic as a sequential auction with declining inventory. Each pick is best-responding to the current frontier, not the original. Players who treat opening placement as a static optimization (pick the highest-pip 2:1 ore + wheat regardless) are leaving value on the table; the right play accounts for what the frontier will look like when it’s your next turn.
Here’s the second-order strategic insight, and the one Roman’s data couldn’t reach without much more sophisticated analysis:
The board is common-knowledge information to all players. Everyone sees the same hex layout, the same number tokens, the same port positions. Therefore:
The contrarian move: the popular resources are overbid; the less-popular configurations leave open cheaper routes.
If the board obviously favors cities (ore/wheat clustered with high pips), but four players are bidding for those vertices, the third-best ore/wheat vertex might be worse than the best wood/brick vertex on the same board — even though the average ore/wheat opening beats the average wood/brick opening across 47k boards.
This is the same dynamic as crowded trades in finance: the consensus opportunity gets bid up to a point where its risk-adjusted return is lower than the unfashionable alternative. The right play is to best-respond to opponents’ expected bids, not to the nominal frontier.
The strategic implication:
| Player’s situation | Right play |
|---|---|
| Picks early (round 1) on a clearly-optimal board | Take the obvious vertex; you have the bidding advantage |
| Picks late (round 1) on a clearly-optimal board | The obvious vertices are gone; counter-position toward the second-tier path if its third-best vertex is better than the obvious path’s fifth-best vertex |
| Picks early on a balanced board | The frontier is wide; pick for flexibility and adjacency to ports / future trade partners |
| Picks late on a balanced board | Same flexibility logic; the bidding asymmetry is smaller |
This is a Keynesian beauty contest layered on the bilateral trade valuation thesis: you’re not maximizing trade efficiency, you’re maximizing position; and what counts as good position depends on what other players are bidding for.
The frontier-and-bidding theory predicts patterns Roman’s aggregate stats wouldn’t see, but her dataset would:
None of these are proven. They’re the testable form of the theory. Roman’s dataset has the inputs to test all four.
The 47k-game dataset Roman analyzed is not publicly downloadable. Twosheep.io collected it on their platform and provided it to Roman directly. The video description has no link; Roman has published no code, GitHub, or blog post associated with the analysis (only an Instagram). Twosheep’s own blog has done smaller public analyses (108 Champs games, 500+ 1v1 games) but never publishes the underlying data.
| Source | Scale | Public? | Notes |
|---|---|---|---|
| Twosheep.io 47k dataset (Roman’s data) | ~47,000 games | No | Private; provided to Roman by twosheep directly |
Twosheep.io replay browser at twosheep.io/games?to=... |
All games on platform | Replays viewable, not exportable | Per-game replay only; no bulk export |
| Colonist.io game logs | 3M / month, 600GB total | No | Explicitly withheld; community feature request unfulfilled |
colonist.inpolen.nl community archive |
Currently 0 stored | Per-game JSON download mentioned | Personal use only; bulk scraping IP-banned |
Kaggle lumins/settlers-of-catan-games |
50 four-player games, 200 rows | Yes | The standard small public dataset; everyone’s analysis uses this |
Kaggle thedevastator/... and koftezz/... |
Similarly small | Yes | Alternative small datasets, similar shape |
The gap: Roman’s analysis at N=47k is ~1000× larger than the largest publicly downloadable Catan dataset. Her empirical landmarks (24.9% turn-order, 59% Longest Road, 25/40/55 Monopoly card, 8% even-trade win rate) cannot be reproduced or extended at scale without:
Strategic takeaway: The frontier-and-bidding theory’s empirical validation is gated by data access that doesn’t currently exist. The theory remains internally coherent and connected to the Monopoly project’s frontier work, where Chris already has access to the simulator and can test analogous predictions. Monopoly is the better proving ground for now — same theoretical structure, full access to the data-generating process.
A genealogy point worth recording: the Monopoly project went through two distinct generations of trade-valuation tools, and the lessons map directly onto what a Catan frontier tool would and wouldn’t be.
Generation 1 — Exact bilateral valuation. The bilateral trade valuation page is the trajectory-based exact calculator: simulate both players’ development paths under hypothetical trades, compute who wins what 20 turns later, score the trade against position deltas. It works brilliantly for apples-to-apples trades — “I’ll give you my orange to complete your monopoly, you give me your dark blue to complete mine.” In that situation, both players are completing monopolies, both dominate the rest of the field, and the math closes cleanly. It also was known at the time to say almost nothing about 1:1 trades of non-completion properties — where the strategic value of any single property depends on what else you might acquire, and the exact NPV calculation can’t see that combinatorial structure.
Generation 2 — Efficient frontier subgraphs. The frontier trade theory page is the structural tool: each player has a random portfolio (the properties they hold from random landing), and the frontier subgraph shows the Pareto-optimal expansion paths from that portfolio — which trades or developments would move them toward dominance, which are strictly dominated. The frontier doesn’t give exact valuations; it gives strategic-value visibility (“this property is on my critical path”; “that property is dominated for me but on Player B’s frontier — so they’ll trade hard for it”) and eliminates frivolous exchanges by structural truth rather than enumeration. This is discrete-game Modern Portfolio Theory — the Markowitz/CAPM Pareto-frontier construction adapted from continuous risk-return space to a discrete game’s state space. The combinatorial-bloat reduction works for the same reason MPT works in continuous space: structural dominance means most options are strictly worse than other options on the same axes and can be pruned without evaluation. See the frontier page’s methodological lineage section.
The two tools are complementary, not competing:
| Tool | What it does | When to use |
|---|---|---|
| Exact bilateral simulation | Precise NPV-delta on a specific trade between two specific players | Apples-to-apples completion trades; bidding decisions on auctioned properties; high-stakes single-move evaluations |
| Frontier subgraph | Map of reachable strategic positions and their dominance relations | Long-horizon planning; deciding which trades to even consider; identifying counterparty leverage points |
The bilateral evaluator says “here’s exactly what this trade is worth.” The frontier subgraph says “here are the trades worth thinking about, here’s what each player values, and here’s the shape of the reachable strategy space.”
The Catan equivalent of the Monopoly frontier subgraph would map possible expansion paths from each player’s current state — which is closer to the right shape for Catan than a bilateral evaluator would be (Catan has fewer crisp apples-to-apples trades than Monopoly’s monopoly-completion structure).
Inputs the Catan frontier tool would need (and what we have / lack from the 50-game dataset):
| Input | Per-game cost to compute | Have in 50-game dataset? |
|---|---|---|
| Full hex layout (19 tiles with number tokens and resource types) | Read from game setup | ❌ Only the tiles a player’s settlements touched |
| Port layout (9 ports around the coast) | Read from game setup | ⚠️ Partial — only ports adjacent to player settlements |
| Current settlements and cities for each player | Read from game state | ✅ Starting only; no mid-game expansion data |
| Current road network for each player | Read from game state | ❌ Not in dataset |
| Each player’s current resource hand | Read from game state | ❌ Not in dataset |
The 50-game dataset is not enough to build the frontier tool — it has starting positions but no mid-game state, and no full board layout. But the algorithmic shape can be specified independently:
The frontier algorithm (sketch):
What the frontier graph shows:
What the frontier graph doesn’t show (and intentionally doesn’t):
The “you need X, Y, Z to get here, but the graph doesn’t know how to get X, Y, Z” structure is the same as Monopoly’s frontier subgraph. Both tools answer “what’s the destination worth pursuing” and leave “how to procure the inputs” to other components. In Monopoly the procurement is mostly trade; in Catan it’s trade + production timing + dice + robber play. The split keeps both tools focused on what they’re good at.
The whole board-frontier theory sketch earlier in this page can now be more precisely typed:
The vault’s two Monopoly-project tool generations map directly onto two complementary Catan tools. We don’t have the data to build either yet, but the architectural shape is now specified.
This is the Monopoly project’s frontier methodology applied to Catan opening placement — same shape, different game. The instinct generalizes:
The Catan theory and the Monopoly theory are two instances of the same underlying frontier-and-bidding analysis. Formalizing one helps formalize the other; both are open vault threads worth continuing.
Catan has very little rubber-banding by design. The main lever is the robber — when 7 is rolled, every player holding more than 7 cards must discard half, and the rolling player may move the robber to a specific tile, blocking production and stealing a card from one adjacent player. This is mild anti-hoarding pressure aimed loosely at the leader: high-producers tend to accumulate more cards, so the discard-on-7 rule strips them disproportionately. It’s transparent, symmetric (applies to anyone over 7 cards), and rule-level — exactly the good-rubber-banding profile the M.U.L.E. analysis named ([[feedback_rubber_banding_and_friction_patterns]]).
Monopoly has effectively zero rubber-banding. No mechanism redistributes from leader to laggards; the snowball just snowballs. Once one player owns three monopolies and the others own none, the game is structurally over — even if it takes 20 more turns to finalize. This is one reason Monopoly is widely disliked despite its iconic status: the loser-player experience drags long after the outcome is determined.
The case for more rubber-banding (egalitarian dynamics):
The case for less rubber-banding (skill-rewarding dynamics):
Too little rubber-banding → runaway games where the early-game determines the late-game; losing players have nothing to do. Too much rubber-banding → all players bunched at the finish line, and the winner is whoever rolls the right number on the last turn. Both extremes collapse to a less interesting game, just in different directions:
| Rubber-banding level | Failure mode | Skill cap | Example |
|---|---|---|---|
| None | Runaway snowball, dead late-game for laggards | Maximum | Chess, Monopoly |
| Mild | Lead is defensible but pressure exists | High | Catan (robber), M.U.L.E. (auction + events), Power Grid |
| Heavy | Lead is hard to keep; the last turn matters most | Reduced | Mario Kart with blue shells |
| Total | Everyone tied at the end; dice/luck decides | Floor (≈ random) | Some racing-game catch-up AI; “everyone gets a trophy” |
The interesting question isn’t “is rubber-banding good or bad” — it’s where on this axis is the game designed to land, and is the choice intentional? Bunten’s M.U.L.E. (1983) made a clear deliberate choice for mild, transparent, rule-level rubber-banding — that’s the design master move the M.U.L.E. project documented. Catan made a similar choice with the robber. Monopoly made the opposite choice (no rubber-banding) — defensible at the time of design, but the snowball-then-finish-it problem is what makes modern players prefer game-night Catan.
A subtler version of the tension: rubber-banding caps the skill ceiling. When a leader can be reliably dogpiled or random-evented back to the pack, the skill premium of “playing well from a good position” collapses. Skilled players cannot translate their lead into reliable victory; the game becomes about reading rubber-band moments rather than about executing strategy. This is fun for newcomers but disincentivizes investment in skill development.
Chess at the extreme: grandmaster vs. amateur is a deterministic outcome. The skill premium is total. This is why chess is the canonical competitive game and not a popular casual one.
Most well-designed Eurogames (Catan, Terra Mystica, Brass) sit in the mild rubber-banding zone — enough leveling to keep games interesting without flattening skill. The trick is that the leveling is at the rule level (everyone faces the same discards-on-7, the same auction structure) rather than at the mechanic level (some hidden code targets the leader). Rule-level rubber-banding preserves competitive integrity; mechanic-level rubber-banding feels like the game cheating against the leader.
A planner-LM composite or an EBM playing a rubber-banding-aware game needs to model how much its lead is worth. In Chess, a +2 advantage is worth ~certainty. In Catan with the robber, a +2 VP advantage is worth maybe 60% — opponents will robber-target you, dogpile via trade refusal, and the dev-card deck remains a wild card. A naive Catan AI that maxes for instantaneous expected VPs will systematically over-value early leads and under-invest in the defensive moves that protect them from rubber-band reversal. Skill-cap calibration is part of opponent modeling.
The Multiplayer Coalition Problem page already documents the three-player stable state: in any game where N≥3 and players can interfere with each other’s development, weaker players gang up on the leader. The leader gets restricted; a new leader emerges; the cycle repeats. No coordination required — each weaker player independently arrives at the same conclusion (the leader is the biggest threat to my survival).
What makes the cycle terminate — that’s where games differ structurally. The first reading of this thread put weight on hidden-random information as the kill switch. The sharper reading puts weight on interference completeness: can the dogpile actually stop progress, or only slow it?
Interference completeness:
Hidden-random information:
The interaction matrix for whether N≥3 games terminate under ideal play:
| Interference | Hidden-random | Terminates under ideal play? | Example |
|---|---|---|---|
| Complete | None | ❌ Never — perfect-AI dogpile is stable | Hypothetical perfect-info Risk |
| Complete | Present | ⚠️ Only via persuasion / asymmetry exploit; fails vs. perfect AI | Risk, Diplomacy |
| Incomplete | None | ✅ Yes — un-suppressible progress accumulates | Hypothetical no-dev-card Catan |
| Incomplete | Present | ✅ Yes — production stream forces it; hidden-random adds timing | Catan |
The hidden-random kill switch helps games like Risk almost terminate, but it’s not enough — Chris’s sharpest point: in Risk, high-level players treat 3+ cards AS-IF the player already has the turn-in armies. That converts the hidden-random into common knowledge, neutralizing it. Good Risk players assume worst-case and plan accordingly, which collapses the asymmetry and restores the perfect-information dogpile. The card turn-in surprises a beginner; it doesn’t surprise a planner.
The same move works in Diplomacy. Good players treat ambiguous-intent moves as if they’re the worst plausible interpretation. Hidden-random becomes assumed-worst, and the dogpile re-stabilizes.
This is why Risk and Diplomacy at high level require out-of-game persuasion to terminate. The strategic stalemate is broken not by mechanical state but by social negotiation — convincing an opponent to defect from the optimal dogpile coalition. Diplomacy is the genre name for this and the literal mechanic of the game. Risk’s endgame is full of meta-talk (“if you attack me I’ll throw the game to him”) because the on-board state cannot resolve the equilibrium by itself.
Against a perfect AI that doesn’t believe out-of-game promises, neither Risk nor Diplomacy would end. This is the structural reason Cicero needed both a planner AND a language model — the language model is what does the persuasion the planner can’t fake. [[cicero-press-diplomacy-captain-meme]] and [[Gunboat Diplomacy and Diplodocus]] together bracket this exact result: Cicero needs the language layer to handle the persuasion side; Diplodocus (pure planner, no language) wins only in Gunboat where persuasion isn’t allowed.
Catan’s interference is incomplete in a specific way: the production stream. Every turn, dice are rolled, and every adjacent settlement/city produces resources. You can rob one tile. You can refuse to trade. You can play knights to displace the robber to a defender’s tile. But you cannot stop the dice from coming up 6 and 8, and as long as a player has high-pip settlements, they will accumulate resources regardless of what their opponents do.
This is why Chris’s sharpest claim holds: “In ideal play, Risk will never end, but Catan will always end.”
The dev cards add a timing element — they let a particular player end the game sooner than visible-state accumulation alone would — but they’re not the structural termination mechanism. The structural mechanism is the un-suppressible production stream. Even in a hypothetical no-dev-card Catan, games would still end; they’d just end slower, with the winner being whichever player’s pip exposure best survived the dogpile-throttle.
The “feels like chance” complaint is correctly diagnosed once you separate the two factors:
Both elements are partly luck. But the structural fact remains: Catan ends because production keeps happening, not because dev cards eventually reveal. The hidden-random is the timing layer; the un-suppressible production is the engine.
| Game | N | Interference | Hidden-random | Ends under perfect AI? | Termination mechanism |
|---|---|---|---|---|---|
| Chess | 2 | Complete | None | Yes (zero-sum) | Direct |
| Risk | 3-6 | Complete | Card turn-ins | No — needs persuasion | Out-of-game diplomacy |
| Diplomacy | 7 | Complete | Hidden orders | No — needs persuasion | Out-of-game diplomacy |
| Coup | 3-6 | Complete (player elimination) | Hidden roles | Partly — hidden roles + bluffing | Bluff resolution |
| Poker (multi-way) | 3+ | Incomplete (you can only fold, can’t take their stack except by winning hands) | Hole cards | Yes | Forced bet-stream + reveal |
| Catan | 3-4 | Incomplete (production stream) | Dev cards | Yes | Production accumulation |
| Monopoly | 2-8 | Incomplete (no direct attack; only refuse-trade and cash drain) | None significant | Yes (snowball) | Cash drain to zero |
| MOO1 (mp) | 3+ | Mostly complete (fleets can kill) | Hidden tech/fleets | Partly — hidden tech enables breakouts | Tech breakthrough + military |
The pattern: incomplete interference is what makes a multiplayer game tractable for perfect AI. Complete-interference games (Risk, Diplomacy) require the language/persuasion layer to terminate at all — which is exactly why Cicero needed an LM and why [[planner-lm-composites]] are the right architecture for that game class.
Catan, Poker, and Monopoly terminate without persuasion because their interference is structurally incomplete. The dice / cards / cash flow can’t be fully throttled.
For AI design: the right architecture depends on whether the game is interference-complete. A pure planner (Diplodocus-style) works in incomplete-interference games where mechanical state dominates. A planner-LM composite (Cicero-style) is required in complete-interference games where the language layer is the termination mechanism.
For game design: interference completeness is a more load-bearing knob than hidden information. A designer choosing N≥3 has to decide whether their game will terminate via mechanical accumulation (incomplete interference, like Catan) or via social negotiation (complete interference, like Risk/Diplomacy). The first works for casual / family / closed-system contexts; the second requires players who treat persuasion as a legitimate game move and willingly defect from optimal coalitions.
For the rubber-banding analysis above: the rubber-banding dial interacts with interference completeness. Heavy rubber-banding on top of complete interference produces deadlock (everyone bunched, can’t escape). Mild rubber-banding on top of incomplete interference produces the Catan zone — skill-driven progress with leveling pressure but guaranteed termination. The design space isn’t one-dimensional.
The vault claim: Chris’s “Risk will never end, Catan will always end” is the cleanest one-line summary of the difference between these two N≥3 design families. It deserves to be stated explicitly: interference completeness, not hidden information, is the structural termination determinant in N≥3 multiplayer games.
The perfect-AI result says Risk and Diplomacy should never end. In practice they end every weekend. What gives?
The answer is bounded human rationality, and Risk’s two card-variant designs make this visible:
Fixed turn-ins (4-6-8-10-style, or fixed army counts per card type) — the turn-in value is bounded and predictable. Players can apply Chris’s AS-IF rule cleanly: 3+ cards = treat as if the player already has roughly N armies coming. The hidden-random asymmetry collapses to common-knowledge worst-case, and the dogpile re-stabilizes. Fixed turn-in Risk preserves the deadlock, which is why high-level fixed-turn-in games end via persuasion / kingmaker plays rather than via mechanical state.
Escalation turn-ins — values grow linearly but unboundedly. Typical schedule is +5 per turn-in (e.g., …, 20, 25, 30, 35, 40, 45, 50, …). It’s not exponential; each step is small. But there’s no cap. After ~200 turn-ins in a long game, the value is ~1000 armies (by which point all players should have 2000+ total armies on the board). Mechanically this is just scale, and a perfect AI would track it perfectly. But humans systematically under-update their threat models for two reasons:
These aren’t the same as the “exponential exceeds cognition” failure I initially described — they’re closer to anchoring bias + base-rate neglect + attention erosion over long games. The math itself is trivially linear; the cognitive failure is in re-pricing the cumulative.
The termination mechanism in escalation Risk is human cognitive failure to re-anchor, not strategy. A perfect AI playing escalation Risk would dogpile card-holders much more aggressively than humans do, because the AI would correctly anchor on the current turn-in value rather than the early-game one. Against perfect AI, escalation Risk would also stall — just at a different scale than fixed Risk. The variant doesn’t change the structural result; it changes how easily humans fall out of the deadlock.
Risk and Diplomacy are designed for humans, not perfect agents. The “tractable termination via persuasion” mechanism and the “tractable termination via cognitive-scaling failure” mechanism are both bounded-rationality exploits. Strip the bound and the games revert to their structural form: indefinite deadlock.
This connects to several existing vault threads:
Generalizable lens: when a multiplayer game terminates under perfect-AI theory it terminates “structurally” (production stream, forced bets, etc.). When it terminates only under bounded-rationality real players it terminates “experientially.” Most successful complete-interference multiplayer games are experiential — Risk, Diplomacy, Coup, Werewolf, social-deduction generally. They are designed for the bound, not despite it. Strip the human and you strip the game.
This is one reason game-AI for these titles is genuinely hard: the AI must model not just optimal play but the opponents’ rationality bound. Cicero’s LM is doing exactly that — modeling what humans will believe and feel persuaded by. Diplodocus avoids the problem by playing Gunboat (no persuasion channel) — but that’s a deliberately bounded variant where the rationality-failure exploit isn’t available.
The two mechanisms above (persuasion-for-defection, anchoring + fatigue) don’t exhaust the failure modes humans bring to complete-interference games. A third mode worth naming is emotional leveling — players who break the dogpile equilibrium not because the math says they should, and not because their threat models are mis-anchored, but because they have an aesthetic / emotional preference for board balance that overrides strategy.
A worked example from a real play group: a player who didn’t like it when army counts got “too high” would, on a whim, systematically attack across multiple opponents to “even things out.” Mechanically suboptimal — he exhausted his own troops across several players, weakening himself. But the result was exactly what breaks the deadlock: he created a new asymmetric state where the next player’s turn-in produced overwhelming advantage against a specific weakened target, often ending the game in 1-2 turns. The leveler player became an unintentional kingmaker for the next-up player.
This is a distinct bounded-rationality failure from the others:
| Failure mode | Trigger | Mechanism | Result |
|---|---|---|---|
| Persuasion-for-defection (fixed Risk, Diplomacy) | Out-of-game social pressure | Player abandons mathematically-optimal coalition for personal/social reasons | Defector becomes kingmaker |
| Anchoring + fatigue (escalation Risk) | Long-game cumulative scale | Player under-updates threat model for late-game card values | Card-holder accumulates uninterceptable advantage |
| Emotional leveling (any complete-interference game) | Aesthetic discomfort with imbalance | Player burns own position to “even” perceived inequality | Creates exploitable asymmetry for next-up player |
The Diplomacy parallel: there’s a documented “longest tournament Diplomacy game ever” video series — high-tournament Diplomacy genuinely does stall for many hours / multiple game-sessions before terminating. The empirical signature of “perfect play → infinite” shows up exactly as you’d predict: among the skillest available humans (who don’t make obvious anchoring or emotional-leveling mistakes), games trend toward indefinite. Termination requires some opponent to make some sub-optimal move — and the longer the tournament, the more reliably someone eventually does.
The structural conclusion sharpens: complete-interference games terminate in practice because real players have multiple independent bounded-rationality failures any of which can break the deadlock. Persuasion, anchoring, emotional leveling, plus boredom, ego-driven plays, alliance-as-friendship-test, etc. The game design doesn’t need to exploit any specific bound — it just needs the set of human bounds to be large enough that, given N≥3 players over enough turns, at least one will fire. Multiple-independent-bounds redundancy is how human play terminates structurally-infinite games reliably.
This is also the empirical answer to “why hasn’t a perfect Diplomacy AI just dominated the field by being unexploitable” — the answer is that there are some games (the longest ones) where the AI just would not finish, and human tournament play would beat it on time-budget alone. Cicero handled this by being able to cause the bounded-rationality failures (via persuasion) rather than just exploiting ones that happened to occur — which is the structural reason its language model is doing real strategic work, not just performative chatter.
Both games are resource-allocation problems with random elements; both are vault research threads (Monopoly with deep work, Catan now sketched). But they are mathematically different kinds of problems, and naming the difference clarifies why the same frontier framework looks different in each, and why Catan was groundbreaking when it shipped.
Where in the game does the randomness actually live?
| Game | Random axis | Deterministic axis | What the player optimizes |
|---|---|---|---|
| Monopoly | Access (dice rolls determine which fixed properties you land on, when) | Board configuration (same every game) | Convert randomly-dealt holdings into a winning portfolio via trade and development |
| Catan | Terrain (hex layout + number tokens + ports are different every setup) | Play sequence once board is set (snake-order + standard turn structure) | Read the board and commit to the strategy archetype the terrain supports |
This is two different problems sharing the same resource-allocation skeleton:
“Monopoly’s efficient frontier is how to get from your random slice of the pie to a winning one. Catan is what are the high-probability resources that will be in play so I know what to play for or around.”
That’s the whole structural difference in one line. Both are “frontier” problems, but the frontiers operate at different time scales:
This explains why the Monopoly project leans heavily on Markov chains and Catan-style analysis doesn’t:
Catan trades one kind of math (Markov-on-time) for another (combinatorics-on-configuration plus per-board state-conditional analysis). It’s actually a harder mathematical problem in the sense that it requires conditioning on the board archetype before any inference is valid. Roman’s aggregate 47k-game stats blur exactly this conditioning, which is why her findings feel descriptive rather than prescriptive — she averaged across board-archetypes that demanded different strategies.
The modular hex-tile setup was Catan’s first-major-mainstream shift to a different design axis when it shipped in 1995. Modular-board games did exist before Catan in niche / hobbyist circles — Avalon Hill’s Galactic Conquest and other lesser-known titles experimented with variable setup — but the dominant 20th-century mass-market resource-allocation games (Monopoly, Risk, Axis & Allies, Diplomacy) all had fixed boards, and the design instinct was “fix the geometry, randomize the play through it.” Catan was the first crossover hit to reverse this at scale: randomize the geometry, deterministic the play through it. The credit isn’t strict invention; it’s the move from niche-hobby experiment to mainstream-defining standard.
Game-design implications:
Catan-style modular setup spread across the hobby after 1995 — see virtually every modern Eurogame (Terra Mystica, Scythe, Wingspan, Brass: Birmingham). Roman’s analysis is best read as one of the first systematic empirical studies of the modular-board class of games rather than just of Catan specifically.
The board-frontier theory’s predictions need a data-generating process to validate against. The vault has full access to one for Monopoly (the simulator) and zero for Catan (Roman’s 47k is private; Kaggle is 50 games). The structural contrast above explains why this is fine rather than a compromise:
Cleaner framing: Monopoly and Catan are not competitor analyses — they’re complementary instances of the same frontier theory at different randomization layers. Validating one strengthens the other.
games, strategy, game-theory, mathematics, simulation, economics