Gaming
Games as interactive laboratories for systems thinking — strategy, economics, simulation, and AI.
Links: Economics, Computation and Information Theory, Monopoly, Slay, Slay-C, Demolition Man Analysis
Why This Matters
Games aren’t a side interest — they’re how Chris thinks. The same frameworks that drive the vault’s economics and philosophy research show up directly in gameplay:
-
4X games (Master of Orion, Civilization) are resource allocation under uncertainty — exactly what Value and Profit and Risk and Entrepreneurship describe theoretically, and which MIRR — 4X as Capital Allocation makes mechanical. Every turn is an entrepreneurial bet: invest in military or infrastructure? Expand now or consolidate? The 4X loop (explore, expand, exploit, exterminate) is the market cycle compressed into a game state.
-
Simulation games (SimCity, Cities: Skylines) are distributed systems made visible. You can watch the price system fail — traffic jams are information bottlenecks, zoning is central planning, and the emergent behavior of thousands of simulated agents demonstrates computational irreducibility in real time. These games tweak the economics brain because they let you experiment with distribution concepts you can only theorize about otherwise.
-
Strategy games with AI (Slay, Monopoly) are where game theory becomes code. Building an AI player forces you to formalize intuitions about valuation, risk, and decision-making under incomplete information — the same problems the Cyborg Model addresses at the human/AI collaboration level.
Sub-Topics
- The Nash Bargaining Problem — the atomic unit of negotiation: two players, $100, agree or get nothing; Nash’s axiomatic solution, Rubinstein’s patience model, ultimatum game behavior, and why identical agents can’t trade
- The Multiplayer Coalition Problem — why multiplayer games resist solution, the self-balancing three-player dynamic, phase decomposition, the Stockfish architecture as template, the interest rate framework, and the relative position model (EPT as slope)
- Slay — Evaluation & Search (the 1v1 case) — the tractable base case under the coalition wall: eval-term surgery (frontier margin, realizable treasury, consequence-weighted + reachability-gated cut, effective income) and transposition-first search, plus three generalizing theses — the cut/join graph duality, realizable-treasury (use-it-or-lose-it), and eval-beats-depth in wide-branching games (the early-NNUE lesson; handcrafted terms become the NNUE’s features)
- Bilateral Trade Valuation — why trade evaluation requires simulating both players simultaneously; trajectory divergence, the patient predator exploit, and Nash equilibrium pricing
- BattleValue — BV = sqrt(Attack × HP): a universal combat comparison metric derived from Lanchester’s Square Law; BV/Cost as the army composition ROI metric
- Combinatorial vs Generative Design Space — a closed ability vocabulary is checkable by enumeration; per-unit free text is not. Measured on HeroClix (0 → 434 chars of bespoke rules text per figure) — and the same shift makes any analysis method’s coverage decay with the age of its target
- MIRR — 4X Strategy as Capital Allocation Under Uncertainty — the central 4X thesis: each turn is a capital-allocation decision ranked by MIRR (reinvestment at your empire’s own growth rate); the cross-game hub MoO + MoM point up to; option value, time-window “broken” strategies, research-as-portfolio
- Capability Without Leverage — a paid capability is worth zero (or negative) unless downstream leverage lets it shift the outcome; the metric correction (impact-share, not face efficiency) across Monopoly (brown trap), MoM (hard counter to zero), Catan (unleveraged port)
- Randomness as the Termination Mechanism (N≥3) — the gang-up equilibrium never terminates without randomness; the design-side corollary of the coalition problem; calibration is the design quality; NA1 canonical specimen, Diplomacy the contrast case
- The Dominance-Frontier Lens — (cross-domain) map cost vs effect, draw dominance edges + the frontier curve, flag asymmetric counters; random-start viability as the design verdict; the methodology the counter-graph / frontier / BV / MIRR pages all share
- NA1 — A Game-Design Crucible — the Nobunaga RE project as a ground-truth source of game-design theses; index to the repo + the theses it sparked (n3-termination + candidates)
- Diplomacy: 7 AI Models — analysis of 7 LLMs playing Diplomacy; case study for the coalition problem, Nash bargaining, physical grounding of strategy, and why different agents CAN trade
- LLM Agents Across Strategic Games — seven-game cross-study (Monopoly, Diplomacy, Among Us, Mafia ×2, Coup, Catan) + clones control; architectural signatures stable across games; verification as a spectrum with hidden pockets concentrating decisive value; Catan isolates three LLM-general failures (action bias, no mechanical model, rhetorical contagion) and points to external-state engine architecture as the fix
- LLM Game Benchmark — Outline — framework for evaluating new LLMs against the seven-game study; measurement axes, infrastructure, scoring method
- Catan — 47,000 Games of Empirical Findings — Ioana Roman’s quantitative analysis of 47k recorded games: balanced turn order, opening predicts win rate barely above random (27% vs 25%), Longest Road / Largest Army are symptoms not causes, winners favor the city engine over the road engine, Monopoly cards as systemic liquidity extraction, and the trade-efficiency paradox (winners overpay for position) confirming bilateral trade valuation across N=47k
- Catan 50-Game Validation — empirical proof-of-method on the public 50-game Kaggle dataset; aggregate city-engine bias replicated (Δ=+0.016 winners vs losers), board-conditional refinement partially supported, theory updated to “universal city-engine bias + layered archetype effect”; methodological finding — Roman’s “balanced turn order” requires snake-order placement specifically
- Gunboat Diplomacy and Diplodocus — Meta’s Gunboat-only AI won a tournament against expert humans with no language model; moves as costly signals; the Denmark disband; the sharpest available confirmation of the planner-LM composite thesis
- CaptainMeme vs. 6 Cicero (Press Diplomacy) — expert human ties Cicero for board top; the human/AI split observations (no grudge, forward-looking pure, intent-appeals fail); the N≥3 self-balancing advantage of forward-looking-pure agents
- D&D Spell Damage Model — using CLT, proof by induction, and bimodal distributions to build a spell comparison metric; same instinct as BattleValue, different game
- Battleship — 30 Billion Boards — (specimen) a closed-form proof that structure creates norms (symmetry-breaking with the substrate fully enumerated): uniform randomness on the asymmetric grid → forced optimal play; the minimax placement result is the sharpest conditional ≠ arbitrary case; NA1 is the epistemic twin; game-theory mechanism = deep best-response half vs flat minimax half
- D&D Monster Tournament — Exact Markov Chains Instead of Dice — (design phase) a dice-rolled YouTube elimination tournament among same-CR monsters, re-solved exactly. The videos’ unfaithfulness is the enabling assumption: “cast when available” is a fixed policy, which collapses an MDP into a plain absorbing Markov chain with an exact answer. Same complication classes as MoM (immunities, attack types, frozen on-the-fly decisions); the hard part is sourcing + the policy spec, not the math. Three deliverables: ground-truth validation of BattleValue, quantifying the melee bias by running mobility on/off, and hunting non-transitive cycles that would mean “same CR” is not a total order
- Risk — The Attrition Constant, and Why Big Battles Are Predictable — a Risk battle solved exactly as an absorbing Markov chain, with a design headline rather than a math one: the 3-vs-2 dice cap fixes the engagement frontage regardless of stack size, so Risk obeys Lanchester’s linear law and concentration of force buys nothing (200v100 = 114.88 survivors against 115.10 for two 100v50 fights) — the direct counterexample to BattleValue’s Square-Law derivation, and plausibly what keeps a Risk leader killable. Three structural facts collapse the matrix, all traceable to a 3v2 round killing exactly two armies: parity invariance (half the grid unreachable, measured 50.7%), an IID one-dimensional bulk, and a 595-state closed endgame shell. The IID bulk gives closed forms —
E[S] = A − (2387/2797)·D by a martingale argument, and sd ≈ 1.447·√D, so spread grows as √D while force grows linearly and big battles are proportionally more predictable. Players’ “7 for 6” envelope is 0.435% off. The two halves of the strategic picture pull opposite ways: tactically there is no deterrence — an army is worth 1.172× more on offense, matched stacks of 12+ favour the attacker, and the premium needed to hold shrinks from 1.90× to 1.31× as borders grow, so a big border is an offensive asset — yet a won battle costs ~85% of the attacking force, so in a symmetric three-player standoff the winner falls from a third of the board to 13.6% while the bystander rises to 86.4%, and attacking without falling behind costs (1+c) = 1.853× the defender’s stack. Design reading: Risk omits tactical deterrence and recovers stability from the player count — the coalition problem with the arithmetic filled in — and that stability decays as players are eliminated, the required border rising from ~nothing at three equal players to 117% of the attacker’s stack once two remain. Corrects two pieces of table advice: matching an escalating border is the most expensive way to stay unprotected, and interior garrisons should be 2, never 3 (the 2nd army buys the defender’s second die at 1.65× value; the 3rd buys nothing, and 3-stacks are marginally worse per budget than 2-stacks). Four independent oracles agree to 1e-15; tool at tools/risk-battle-odds.py
- Yahtzee — 259 Trillion → 405 Million — (specimen — same creator, the Battleship sequel) backward induction over a state space collapsed ~640,000× by keeping only what affects future points (EV ≈ 255). Chris’s verdict: confirmation, not discovery — folk strategy formally verified. The live thread is the EV-vs-win split (the reduction is sound only for the points objective, so the PvP section can’t represent “am I ahead”), and a toy model showing variance-seeking pays off only as the opponent’s score hardens — separating reachability from variance tuning. Opens the parked “solo-together” question
- Hangman — Solving Both Sides — (specimen — third Ballpark Figures video, set against 3Blue1Brown’s entropy Wordle bot) the first of the three to solve both players, and the one that inverts half of Battleship: hangman’s minimax chooser is not max-entropy but a hard concentration (
-ING at 100% for lengths 7–9), because the indifference set of an equilibrium is bounded by the substrate’s structure — only ~15 of 26 letters can be made competitive, since Z cannot do S’s job. Where Battleship shows structure emerging from apparent nothing, hangman shows a substrate already lumpy enough that the lumpiness can’t be optimized away. Two solvers, one parameter of disagreement — confidence in the word list: the exact tree is optimal for its 34,483-word model and undefined off it, while entropy is suboptimal but degrades gracefully, which is exactly why it “wastes” guesses on impossible words. Headline: the compression built for human legibility (the alphabet tier list) is also the most model-error-robust artifact — it survives frequency re-weighting and British spelling; the exact tree doesn’t. Also: tractability is set by branching and path reconvergence, not nominal size (hangman is 6 orders smaller than Battleship yet Wordle — “five letters” — is the intractable one); state-space collapse deployed as the adversary’s weapon (force a suffix → turn a 7-letter game into a harder 4-letter one), the inverse of the reduction thread in Yahtzee and Monopoly’s frontier; certified α-cap bounds with unsolved lengths plotted in red so the chart carries its own uncertainty; and a rules-design result only a two-sided solve can produce — the fair miss allowance is 1–8 depending on word length, every number except 6
- Arithmetic Scarcity and the 3D Problem — (hub) fast arithmetic was priced, not assumed: optional on the PDP-8 (EAE) and early PDP-11s (EIS), a separate chip on the PC (8087), microcoded-not-silicon on the 8086 (16-bit
MUL ≈ 25 adds), absent on the 6502. Real-time 3D can’t route around it, so games became the forcing function — and four distinct strategies fall out: buy the math (Battlezone), compute it anyway (Elite), constrain the geometry until it collapses (Doom/Wolf3D), precompute at art time (Ultima dungeons, Wing Commander). The thesis under test: the strategy is predicted by which resource was cheapest for that team, not by which was technically best — i.e. a dominance frontier over silicon / cycles / design freedom / art budget. Counted result from Elite’s own two builds: log-table multiply (FMLTU, ~61 cyc) beats shift-add (MULT1, ~170 cyc) by 2.8× for 1 KB of tables, while Stellar 7’s polar meshes halve the multiply count for zero bytes — orthogonal factors of N × C that nobody stacked. Two deeper results: strategies 2/3/4 collapse into one — memoize and index — cut at four pipeline depths, where the deeper the cut the bigger the saving and the more visible the quantization (sprite popping is caching too far down); and the coprocessor cycle has a period (EAE → 8087 → GPU → NPU; discrete-then-integrated, with local inference the current phase-A and quantization/sparsity/KV-cache the same three moves — which is why the vault’s own “accumulated state IS the verification layer” is this pattern, not an analogy to it)
- Stellar 7 (1983) — The Same Game Without the Coprocessor — (platform / RE — specimen #2, the control) Damon Slye’s Battlezone-alike on an Apple II: same genre, same CPU family, no math box, frame budget within 7%. It has the two routines Battlezone contains zero of — a 695-byte
Divide16 (with the fixed-point pre-shift folded in, so zoom is just a different exponent) and a fully-unrolled Multiply16_8. But the real answer isn’t a faster multiply: vertices are stored polar (distance, angle, Y), so rotating the model is one ADC. Vertical edges share a transform, making a cube cost 4 vertices instead of 8. The costs the coprocessor was hiding also surface — UpdateSound called four times inside the vertex loop (bit-banged speaker), and clipping degraded to per-vertex culling with a documented visible artifact
- Battlezone (1980) — 3D Without a Multiply Instruction — (platform / RE — specimen #1 of the above) the 6502 has no MUL and no DIV; Atari’s answer was a 16-bit bit-slice coprocessor (4× AMD 2901) driven through a memory-mapped API where the address you write to is the opcode. From McFadden’s full commented disassembly:
MB_SCREEN_X performs rotate + translate + perspective-divide from a single store; the pipeline is reordered to View-first so it can cull before paying for vertices; sine is a 65-entry quarter-wave table and atan2 is an octant fold plus a 256-byte table; and the fixed-point errors are absorbed into the art rather than fixed in code. Counted cost ≈ 220 cycles/vertex against a ~96,000-cycle frame budget — the number that says the game doesn’t exist without the coprocessor. Elite is the software answer to the same problem; Wing Commander’s sprite banks are the third
- Breaking Down an SNES Cart — the teardown method — (platform / RE) how to look under the hood of any SNES ROM before knowing an opcode: header → mapping → memory map → subsystems. The architecture break the NES KOEI decompilers got to skip (65C816 not 6502, LoROM/HiROM must be re-detected per cart, separate SPC700 audio CPU, coprocessors). Grounded in the SNESdev wiki; the reusable substrate a future SNES title decompiler will import (
snes-decompiler)
- Gemfire (SNES) fully decompiled — the 2nd KOEI SNES title — (platform / RE) the full title reverse-engineered: 591 bytecode routines named across 11 WRAM overlay modules (root shared-library + command/comusr/comcmp = the strategic AI/settei/event/senzen-sensou battle engine-sengo/ending). Key finding: the per-routine native JSR-launcher trampoline (body at +5; SNES cousin of L’Empereur’s
$E2E3). Independently walking it caught a systematic offset-0 bug in the ROTK2-SNES decompiler.
- The one that isn’t a VM — Nobunaga’s Ambition (SNES) compiled native — (platform / RE) the exception to KOEI’s portable SNES VM: the NES original’s Sea-16 bytecode game was, on this SNES port, compiled straight to native 65C816 (no VM — proven 3 ways; a native walk hits 123 KB with 0 decode conflicts). First HiROM KOEI SNES title. Drove a new native 65816→C decompiler (reuses the DREAM structurer) so both platforms lift to C. Fully reversed (2026-07-16): 100% label-walk (715/715), 148-label data-walk, whole-program C listing (0-fallback); three-pillar AI = player’s own formulas + a
difficulty handicap, acting alone — same decision model as the NES. 60-syscall graphics BIOS, native mul/div lib @ $C1:F800, printf text engine.
- NA1 NES↔SNES — grading two blind reverse-engineerings — (platform / RE / method) the payoff of reversing NA1 twice: reconstruct the game from the SNES native code alone, then diff against the NES bytecode work — create, then check. Two binaries, two consoles, two execution models agree on every formula: 26-byte record + byte-identical 1560 scenario data at twin ROM addresses, Grow =
2·amt·(6−skill)/√…, event cadence + illness rng(400)<100−health, weakest-neighbour coin-flip AI (frozen across the hardware leap), the 8-stat combat table {5,5,10,10,10,15,20,25} sum 100 → +40%, 115−15·skill handicap. Every discrepancy was a lossy-decompiler over-read the check caught.
- Subgraph Investment Optimization — build-vs-trade decision via Pareto dominance, graph matching, dynamic programming on reachable states
- Frontier Trade Theory — EPT frontier curves, brown trap, denial value, knockout probability, game horizon, race condition
- Subgraph Trade Engine Spec — complete AI architecture: subgraph analysis, trade search, pruning, 3-way cycles
MOO1 Theory — moo1/
- MOO1 Optimal Strategy — race tiers, opening theory, tech priorities, ship design via BV, diplomacy, endgame paths
- MOO1 MIRR Analysis — MIRR-based investment decision: factory vs colonizer; financial metric for opening theory
- Economic Analysis — BV + MIRR applied to all units, buildings, and spells; “balance through imbalance” tested quantitatively
- Tier System and MIRR — time-gated tiers as investment; Wraith rush, Halfling Slinger stacking, building option value
Genres and What They Teach
4X — Resource Allocation Under Uncertainty
The 4X genre is the purest game-form of entrepreneurial decision-making. You start with incomplete information, make irreversible investments, and compete against agents with different strategies and different information.
Key games:
- Master of Orion (1993) — The original. Colony management, tech tree, ship design, diplomacy. Clean enough to see the underlying systems. Chris’s preferred 4X.
- Remnants of the Precursors (RotP) — Java remake of MOO1 by Ray Greer. Modernizes the gameplay while preserving the design philosophy. Unfortunately Java-only, which makes it inaccessible on constrained platforms (iOS).
- Starbase Orion — iOS native, closer to MOO2 in complexity. Solid mobile 4X but lacks RotP’s modernization of the original MOO1 feel.
- Civilization series — Broader scope (cultural, diplomatic, scientific victory conditions), but the economic core is the same: allocate scarce resources across competing priorities under uncertainty. Competitive multiplayer Civilization demonstrates that turn-based play has the same adversarial interest rate dynamics as RTS — the turn structure discretizes time but doesn’t change the economics.
What 4X reveals:
- The tech tree is a capital investment problem — you’re betting on future returns from research you can’t fully evaluate yet
- Diplomacy is repeated-game theory — trust, betrayal, credible commitments
- Military strategy is risk management — how much to invest in defense vs. growth
- The entire genre demonstrates why central planning fails at scale: the combinatorial explosion of possible strategies means there’s no “optimal” build order, only adaptive judgment
Simulation — Emergent Systems and Distribution
City builders and simulation games make economic distribution tangible. You set rules and watch emergent behavior unfold — which is exactly what markets do, and exactly what central planners try (and fail) to replicate.
Key games:
- SimCity (1989–) — The original urban simulation. Zoning, taxation, infrastructure. You learn quickly that top-down planning creates problems faster than it solves them.
- Cities: Skylines — SimCity’s spiritual successor with much deeper simulation. Traffic modeling alone demonstrates how local decisions (one badly placed intersection) cascade into system-wide failures. The distribution and logistics layer is a playground for economic thinking.
What simulation reveals:
- Small rule changes produce wildly disproportionate outcomes (sensitivity to initial conditions — connects to Computation and Information Theory)
- You cannot optimize a city for a single metric without destroying it on others — the same argument against Cocteau in Demolition Man
- Emergent traffic patterns are computationally irreducible — you can’t predict them without running the simulation, which is exactly the market computation argument
- The best cities aren’t “planned” in detail — they’re structured with good rules and allowed to evolve
Building game-playing AI forces you to express intuitions as algorithms. This is where the vault’s game theory connects to its AI research.
Active projects:
- Slay / Slay-C — Hex territory control. Alpha-beta search, heuristic evaluation, the tradeoff between search depth and evaluation quality.
- Monopoly — Markov chains for positional analysis, EPT valuation, strategic trading. The AI has to model other players’ utilities to trade well — which is the economic framework in code.
- MOO1 Opening Optimizer — Economic sim for the colony ship timing problem. Simulates the first 50–100 turns to derive optimal expansion timing across race, planet quality, and distance.
The multiplayer challenge: Both Slay (simplified to 2 players) and Monopoly (3+ player trade dynamics never fully solved) hit the same wall — the coalition problem. See The Multiplayer Coalition Problem for the full analysis.
What game AI reveals:
- Evaluation functions are subjective value made explicit — you have to decide what “good” means before you can search for it
- Search depth vs. evaluation quality is a compute budget problem — same tradeoff as agent team architecture
- Perfect play is computationally intractable for interesting games (see P vs NP discussion) — good play requires heuristic judgment, which is the AI equivalent of entrepreneurial intuition
Master of Orion 1’s UI design — menus, clicks, simple state displays — maps naturally to touch interfaces. The original game would be a near-perfect tablet experience if someone built a native port. Instead:
- DOSBox on iOS — Technically works, practically painful. Touch controls mapped to a DOS mouse interface with emulation overhead.
- Remnants of the Precursors — Excellent modernization, but Java/JVM means no iOS path. Apple’s restrictions on JIT compilation make JVM apps structurally impossible on iOS.
- Starbase Orion — Native iOS, closer to MOO2. Good but different design philosophy than MOO1.
This is a microcosm of a larger problem: the best classic game designs are trapped in legacy platforms, and the economics of porting don’t justify the effort for niche audiences. The games that best teach systems thinking are often the hardest to access.
Open Questions
- Could a game be designed specifically as an economics teaching tool — making the price system, comparative advantage, and entrepreneurial judgment the core mechanics rather than just emergent properties?
- What would a 4X game look like with AI agents as advisors (cyborg model applied to gaming)?
- Is there a formal relationship between game AI evaluation functions and utility theory? Both are trying to compress complex state into a scalar value.
- The simulation genre demonstrates computational irreducibility intuitively — could this be formalized into a teaching tool for the concepts in Computation and Information Theory?
games, game-ai, economics, ai, simulation