Code World Models for General Game Playing

8 Mar 2026 · 22 min · 8 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Code World Models (CWMs) for General Game Playing—using an LLM to translate natural-language game rules into executable Python “world” simulators, then letting classical planning (MCTS) play within those verified rules.

Guests

No guest names or backgrounds are provided in the transcript; it’s presented as a host-led discussion of Google DeepMind research.

Key claims

Standard LLM game agents act like “System 1” next-move predictors and hallucinate illegal/short-sighted actions. CWMs shift the LLM to “game developer,” producing rule code with legal-move enumeration, state transitions, win/loss checks, and rewards; MCTS then searches safely. For hidden information, they add “inference as code” (inference functions) and use ISMCTS/ISMCPS plus a regularized autoencoder loop to handle open-deck vs closed-deck.

Notable examples

Games include tic-tac-toe, Connect 4, backgammon, generalized tic-tac-toe/chess, quadranto, hand of war, and Laduk poker; failure on Gin Rummy (procedural scoring/accounting), with ~84% code accuracy and ~52% inference accuracy.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Current AI Limitations in Strategy Games

1:30 to 4:40

Discover why traditional AI struggles with deep strategy in board games.

“So to give you a quick roadmap of our journey today, we're going to explore why standard cutting-edge AI fails so spectacularly at deep strategy.”

Introducing Code World Models

4:40 to 7:20

Learn about a new approach where AI writes game rules in code.

“This is such a brilliant paradigm shift.”

Advantages of Code World Models

7:20 to 9:30

Understand the benefits of using Python in AI game simulations.

“It also seems like this approach would make the AI incredibly adaptable.”

Proving Generalization in AI Learning

9:30 to 12:20

Explore how researchers tested AI's ability to learn new games.

“So how is a language model writing this flawless simulator code instantly?”

Handling Imperfect Information Games

12:20 to 14:00

Learn about the challenges AI faces in games with hidden information.

“This is where the code tries to guess the current state of the game directly from the most recent observations.”

Understanding Closed Deck Scenarios in AI

14:00 to 16:44

Learn how AI addresses the challenge of hidden game mechanics.

“It helps the AI write the simulator very quickly because it can see the cause and effect clearly.”

Performance of CWM Agents in Game Tournaments

16:44 to 19:11

Discover the impressive performance of CWM agents against various opponents.

“The DeepMind team set up 100 match tournaments for each of the 10 games.”

The Limitations of AI in Complex Games

19:11 to 21:36

Explore the challenges AI faces with complex game rules like Gin Rummy.

“It highlights the current frontier and limitation for this technology.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Imagine trying to sit down and play a really complex board game. You know, something multi-layered like Settlers of Catan or maybe Twilight Imperium with a friend. Oh, Twilight Imperium is a serious time commitment. Right. It is. But now picture that this friend is just incredibly well read. I mean, they've consumed millions of books, articles, forum posts about high level strategy and optimal plays. But there's a massive catch. They don't actually know the rules. Exactly. They don't know the rules of the specific game sitting on the table in front of you. So they just keep trying to make these completely illegal, physically impossible moves, and they do it with absolute, unwavering confidence.

0:42Like trying to trade a sheep for... For skyscraper. Yeah, a skyscraper that doesn't even exist in the game. It sounds incredibly frustrating, right? Well, for you listening, that scenario perfectly captures the primary problem with modern artificial intelligence when it tries to play complex multiplayer games. It's a great analogy, actually, because it has all the surface level strategy in the world, but it completely lacks a rigid understanding of the underlying boundaries. And today, our mission is to explore a groundbreaking shift in how AI is tackling this exact problem. We've got a really fascinating stack of recent research from Google DeepMind, and we are going to dive into a proposed architecture where the AI is asked to write the actual rules of the world in executable code.

1:25Rather than just blindly guessing the next best move based on vibes. Based on vibes, exactly. So to give you a quick roadmap of our journey today, we're going to explore why standard cutting-edge AI fails so spectacularly at deep strategy. Which is a fun topic on its own. It really is. Then we'll look at how forcing the AI to write Python code actually fixes this fundamental flaw. We'll dive into the mind-bending challenges of games with hidden information like the fog of war in poker. And finally, we'll look at how this applies to the real world. Okay, let's unpack this. Let's do it. Let's start by looking at the current status quo in the industry.

2:00Because if you look at the dominant approach today, researchers generally take a large language model, an LOM. Something highly advanced, right? Like Gemini 2.5 Pro. Exactly. And they treat it as what we call an intuitive player. They feed the model the history of the game, usually just a text description of the moves that have happened so far. And they simply ask the neural network to predict the next best move. So they treat it almost entirely like a policy engine. It's using its massive pre-trained neural network to say, based on the billions of words I've ingested, players in this specific text-based situation usually do action X.

2:38That is the core mechanism, yeah. But while this worked incredibly well for generating a polite email or writing a college essay. It's strategically shallow. Incredibly shallow when applied to a rigid environment. The LLM is relying on implicit, fragile pattern matching. Because it doesn't actually have a hard-coded understanding of the physical or logical boundaries of the game, it just constantly suggests illegal moves. Or worse, right, it makes moves that look really smart in the short term but completely fall apart three steps later. Right, it severely lacks tactical foresight. Which brings up a framework that's incredibly useful for understanding this limitation.

3:13We are talking about the Daniel Kahneman connection. Ah, yes. System 1 and System 2. Right. The Nobel laureate Daniel Kahneman wrote famously about the two systems of human thought in his book Thinking Fast and Slow. System 1 is our fast, automatic, intuitive, highly emotional brain. And System 2 is the slow, deliberate, analytical, logical brain. Exactly. And the way current LLMs play games is entirely System 1. It's a knee-jerk associative reaction. But true strategic mastery, like executing a deep multi-step look ahead in chess, absolutely requires the heavy lifting of System 2. Think about how you process a new game as a player.

3:54You don't just react instantly to the board state the second it's your turn. No, I usually stare at it for five minutes and annoy my friends. Exactly, because you're sitting there simulating future moves in your head. You think, if I move my knight here, they'll probably move their bishop there, and then my rook is exposed. you're running a mental simulation of causality. But an LLM doesn't have that. It's fundamentally designed to guess the next likely word in a sentence. It doesn't have a mental simulator. It cannot look ahead because it doesn't possess a rigid model of the world to test its ideas against.

4:24It can't simulate the future if it doesn't know the physics of the present. Which is the exact flaw this deep mind research targets. The AI needs a system too. It needs a reliable, unbreakable way to simulate the future. And that brings us to the ultimate aha moment of this research. Code world models, or CWMs. This is such a brilliant paradigm shift. It really is. Instead of asking the LLM to pick a move, the researchers essentially gave the AI a completely different job. They said, read these natural language rules, look at a few offline examples of people playing this game, and then translate all of that text into a perfectly executable Python program.

5:05The LLM is no longer acting as the player. It becomes the game developer. I love that framing. Yeah, it creates the CodeWorld model, which is a playable, mathematically approximate copy of the target game. This generated Python script includes specific functions for state transitions. Meaning how the digital board physically changes when a move is made. Right, and it includes functions that strictly enumerate every single legal move available, functions that check if the game has reached a win or loss state, and functions that calculate the rewards. So once that Python code is written and verified, the LLM actually steps back.

5:39It tags in a completely different kind of AI to do the heavy lifting of playing. It hands the controller over to high-performance classical planning algorithms, specifically a technique called Monte Carlo Tree Search, or MCTS. Which is the same underlying search algorithm that helped AI famously conquer Go. The very same. MCTS uses the Python code that the LLM just wrote as its personal simulation engine. Think of it like a chess supercomputer. It rigorously tests thousands and thousands of possible branching future moves, playing out entire games in its head, constrained by the Python rules, before it actually makes a single decision in the real world.

6:17So what does this all mean? Why go through the massive computational trouble of making the AI write Python code just to play a simple game? Well, the paper highlights three very distinct, massive advantages to this hybrid approach. The first is verifiability. Because the game environment is now defined by rigid Python syntax rather than loose text prediction, the MCTS algorithm can perfectly enumerate valid actions. It algorithmically prevents the AI from cheating. Exactly. Or hallucinating non-existent pieces or making those frustrating illegal moves we talked about earlier. It essentially builds a walled garden.

6:53The AI physically cannot break the rules, even if its intuitive side wants to, because the Python code will just reject the input. That leads directly into the second advantage, which is strategic depth. By marrying these two systems, we're combining the semantic understanding of the LLM. Its unique ability to read a plain English rulebook and understand what it means. Right, with the brutal, raw, deep search power of classical planners. It's the perfect integration of intuition and calculation, System 1 and System 2 working in tandem. It also seems like this approach would make the AI incredibly adaptable.

7:26Like if it's just translating rules into a simulator, it doesn't need to have seen a billion games of chess in its training data to be able to play a chess-like game. You've hit on the third major advantage there, generalization. By directing the LLM to focus on the metatask of translating text and data into a functional code simulator rather than trying to memorize specific opening moves, it becomes highly adaptable. It can learn to play entirely new games it's never encountered before simply by writing a new simulator for them. And that brings us perfectly to the proving grounds the researchers set up.

8:00Because how do you actually prove an AI is generalizing and not just, you know, regurgitating data it memorized from Reddit during its massive training phase? It's a huge problem in AI testing. Contamination. Right. So the researchers set up a testing arena with 10 different games. Crucially, they split these games into two distinct categories. Five of them were perfect information games where you can see the entire state of the board at all times. They used tic-tac-toe, Kinect 4, backgammon, generalized tic-tac-toe, and generalized chess. Notice the specific names of some of those games. Yeah, the generalized ones.

8:36Right, because this raises an important question. If an AI plays a perfect game of standard chess, is it actually reasoning? Or is it just reciting a famous match between grandmasters it read about online? You can never be entirely sure. Exactly. So to entirely avoid this contamination, four of the 10 games in the arena were completely novel. Generalized tic-tac-toe, generalized chess, quadranto, and hand of war. They were created specifically for this research. They do not exist on the internet. Which is brilliant. By testing the AI on these out-of-distribution, or ode games, the team proved the AI is genuinely learning new rules from scratch.

9:13It's synthesizing new environments, not just recalling old ones. But I have a major mechanical question here. Writing perfect Python code that covers every single edge case of a game is incredibly difficult. Human software engineers can barely do it on the first try without introducing bugs. Oh, absolutely. It's notoriously hard. So how is a language model writing this flawless simulator code instantly? It absolutely doesn't do it instantly. The beauty of the system is a process called iterative refinement. It's essentially a highly automated hyperspeed debugging loop. Okay, how does that work?

9:47The system takes the offline data of humans playing the game, and it automatically generates unit tests. It then runs the LLM's generated Python code against those tests. If the code fails... Say the code accidentally allows a knight to move diagonally instead of in an L shape. Right, exactly. If it fails, the system takes the stack trace, which is the literal error message the computer spits out, and feeds it directly back into the LLM as a prompt. So it's acting like an automated manager saying, hey, your code crashed on turn four. Here's the specific error message. Rewrite the function to fix it.

10:20And it does this using a highly sophisticated tree search method combined with Thompson sampling. Thompson sampling. Yeah. Rather than just trying one fix at a time linearly, imagine the AI spinning up a dozen parallel universes where it tries different code fixes simultaneously. simultaneously. Thompson sampling is the statistical method it uses to look at those universes and say, this specific debugging path seems the most promising. Let's focus our computational effort here. Wow. It prunes the bad code and refines the good code until the simulator hits 100 % accuracy against the offline data.

10:55The idea of an AI spinning up a multiverse of debugging branches is incredible. But let's pivot the conversation a bit. Simulating backgammon or tic-tac-toe is one thing. You can see the whole board. The state of the world is a known, verified fact. Yes. But what happens when the world is partially obscured? Here's where it gets really interesting. We are entering the fog of war. The imperfect information games. If you're playing a game like Laduck Poker, you cannot see your opponent's cards. You can't run a Monte Carlo tree search to simulate the next 20 moves of the future if you don't even know what the current state of the board is.

11:29The simulation breaks down immediately. Right. If I don't know whether you're holding a pair of aces or a pair of twos, my mental simulation of your future moves is essentially useless. So how does this code world model handle hidden secrets? The team introduced a concept called inference as code. To make their planner work in the dark, a planner they upgraded and named information set MCTS or ISMCPS, the AI needs to make an educated mathematical guess about the hidden state of the world right now. Okay. So the researchers prompted the LLM to synthesize entirely new Python functions alongside the game rules, called inference functions.

12:08It's quite literally writing a block of code designed to guess your cards based on your behavior. And they tried two distinct approaches to building these inference functions. The first was called hidden state inference. This is where the code tries to guess the current state of the game directly from the most recent observations. Which sounds computationally simpler. It is, but it carries a major conceptual flaw. It largely ignores the dependency of past actions. It doesn't look at the chain of events that led to the current moment. Which brings us to the second approach, I assume. Hidden history inference.

12:42Instead of just guessing the current state in a vacuum, the LLM writes code that attempts to deduce the entire sequence of past actions. So it tries to calculate the dealer's original shuffles, the specific cards dealt to each player, every micro-decision made by the opponent over the course of the match to logically arrive at the current hidden state. Yes, and the research found that this historical method was vastly superior. It forces the simulation to adhere to a strict, logical timeline of causality. It's the difference between a detective walking into a room and making a wild guess about what happened based on the furniture versus reconstructing the crime scene step by step, tracking every footprint from the door to the window.

13:24That's a great way to visualize it. Reconstructing the history provides a much more solid foundation for the simulation. But this introduces a massive hurdle in the research, the open-deck versus closed-deck problem. Right, the peaking problem. This is perhaps the most deeply technical and impressive part of the entire framework. Let's clearly define the parameters. Open-deck means that while the AI is learning the game, during that offline training phase where it's doing the iterative refinement to write the Python code, it is allowed to peek at the hidden states in the training data. It's like learning to play poker by watching a game where everyone's cards are face up on the table.

14:01It helps the AI write the simulator very quickly because it can see the cause and effect clearly. But that is not how reality works. Exactly. If I log on to an internet server to play a brand new card game, I never get to see my opponent's hand. Even after the game ends, the server hides their cards from me. I only ever see my own cards and my own localized observations. This is the closed deck scenario. Right. So the researchers are asking the AI to write a perfect Python simulator for a game where it is fundamentally permanently blind to half the mechanics. It seems impossible at first glance.

14:34I mean, how do you code rules for things you literally cannot see? The researchers solved this by forcing the LLM to build a regularized autoencoder entirely out of Python code. We need to pause and make sure that concept is crystal clear for you listening. What is a regularized autoencoder in the context of this game-playing AI? Let's break it down. Think of an autoencoder as a translation system with a very strict bottleneck. The inference function we just talked about, the code guessing the hidden history, acts as the encoder. It takes the limited, fragmented things the player actually saw and maps them to a theoretical hidden history.

15:10Then the CodeWorld model itself acts as the decoder. It takes that theoretical history and runs it forward through its own rules to try and recreate the original observations. So to put that in plain English, the AI says, based on the few things I was allowed to see, here is my detailed theory of what happened in the dark. Now, if I play my own theory forward using the physical rules I wrote, does the outcome perfectly match the few things I actually saw? That is the exact loop. They throw out any unit tests during the debugging phase that require godlike hidden knowledge. They only test the code on this specific loop, observation to theory back to observation.

15:50And if the code can complete that loop without crashing and without contradicting the visible data, the theory is considered solid. What's fascinating here is that the structural rules of the game and the strict programming interface required by the testing environment act as a regularizer. It algorithmically prevents the LLM from creating nonsense logic or hallucinating magical game mechanics to explain the hidden data. The rigid constraints of the Python syntax essentially force the AI to discover the true hidden mechanics of the game. It's like trapping the AI in a pitch-black escape room, and it uses the limitations of its own coding language to map out the exact dimensions of the room it can't even see.

16:29It's brilliant. It really is an elegant solution. So the ultimate question. After building these code world models, going through a multiverse of iterative refinement and writing autoencoders to pierce the fog of war, How did the AI actually perform in the arena? The results were genuinely staggering. The DeepMind team set up 100 match tournaments for each of the 10 games. They pitted their new CWM agent against a purely random agent, against an agent that had access to the ground truth code. Meaning the actual underlying code written by human developers who made the game. Right. And finally, against Gemini 2.5 Pro, acting in its standard intuitive system one policy mode.

17:07The ultimate heavyweight bout. System 2 code generation and deep search versus System 1 pure neural network text prediction. How did it shake out? The CWM method outperformed or matched Gemini 2.5 Pro in 9 out of the 10 games. Wow, 9 out of 10. It completely dismantled the traditional approach to playing these games. And the team didn't stop there. To make the Monte Carlo tree search even faster, the LLM also synthesized value functions. Which are pieces of code that can instantly estimate how good or bad a board state is without having to simulate the game all the way to the very end, saving massive amounts of compute time.

17:45Exactly. They even tried running a reinforcement learning algorithm called PPO, Proximal Policy Optimization. You can think of that as a rigorous trial-and-error training montage where the AI gets rewards for good moves. Right, and they ran it entirely inside the generated code simulator. The PPO method easily beat the random agents, though it's still usually lost to the raw analytical depth of the MCTS method. Winning or matching in 9 out of 10 complex games is an incredible success rate for a brand new architecture. But I have to ask about the one failure. Which game tripped up this super system?

18:18Gin Rummy. Gin Rummy? Really? Why Gin Rummy? On the surface, it doesn't seem that much harder than Laduk poker or backgammon. It comes down to procedural and logical complexity. Gin Rummy isn't just about taking turns moving pieces. It has highly intricate, multi-stage scoring phases. Right, you have knocking, laying off melds, calculating dead wood, checking for undercuts. It is a very dense, logical sequence of accounting subroutines. It's less like playing a tactical war game and more like trying to do your taxes while playing cards. That is a perfect analogy. The LLM struggled to perfectly capture this multi-step procedural accounting in Python code just from observing a few offline games.

18:58The training accuracy for the GenRummy code hovered around 84 % rather than hitting 100. And because the foundation was shaky, the inference accuracy, its ability to guess the hidden state of the opponent's cards, dropped to about 52%. It highlights the current frontier and limitation for this technology. It handles strategy and spatial mechanics beautifully, but dense, multi-step mathematical accounting based on sparse data is still a significant hurdle. That makes a lot of sense. Writing a flawless Python script to calculate Deadwood and undercuts from scratch based purely on observing a few hands of cards is a monumental task.

19:34But stepping back from the specific games, the big takeaway here for you listening is profound. We really are witnessing a fundamental shift in how we build AI. By shifting the burden on the LLM away from producing a good policy, meaning just guessing the next move and moving it toward producing a good world model, writing the physics simulator itself, we unlock massive strategic depth and completely avoid those illegal hallucinated moves that plague current systems. It is the literal activation of AI system two. We are giving it the tools to think slowly and deliberately. Consider how you can apply this framework to your own life.

20:10The next time you're facing a complex problem or you feel overwhelmed by a flood of conflicting information at work or in a negotiation, don't just react intuitively. Don't fall into the system one trap of doing the first thing that feels right. Exactly. Take a step back. Try to explicitly map out the rules of your environment first. understand the constraints, write the mental code of the situation, so you can accurately simulate the outcomes of your decisions before you actually make them. And if we connect this to the bigger picture, it leaves us with a rather provocative thought to mull over.

20:44We have just seen that AI can now successfully deduce and code the rules of structured digital games, even when half the information is permanently hidden from it. Yeah. But what happens when we unleash this architecture into unstructured open world environments? If an AI can reverse engineer the hidden rules of a card game using observation and autoencoders, what happens when it is tasked with reverse engineering the hidden code of human social dynamics, macroeconomics, or high-stakes geopolitical negotiations? That's wild to think about. If it can observe the plays we make as humans and write a perfectly functioning Python simulator of our society, what kind of unbeatable strategies might it simulate next?

21:25That is a simultaneously awe-inspiring and slightly terrifying thought to leave off on. We are moving from simply playing the game to writing the physics engine of reality itself. Thank you so much for joining us on this deep dive today. Keep questioning the rules of the game you're playing, and we will catch you next time.

From the publisher

Researchers at Google DeepMind introduced Code World Models (CWM), a framework that uses Large Language Models to translate natural language game rules and player trajectories into executable Python code. Unlike traditional methods that use LLMs as direct move-generating policies, this approach treats the model as a verifiable simulation engine capable of defining state transitions and legal actions. The generated code serves as a foundation for high-performance planning algorithms like Monte Carlo tree search (MCTS), which provides significantly greater strategic depth. The framework also synthesizes inference functions to estimate hidden states in imperfect information games and heuristic value functions to optimize search efficiency. Evaluated across ten diverse games, the CWM agent consistently matched or outperformed Gemini 2.5 Pro, demonstrating superior generalization on novel, out-of-distribution games. This shift from "intuitive" play to System 2 deliberation allows the agent to maintain formal rule adherence while scaling performance with increased computational power.

More from Best AI papers explained

All 475 episodes
Code World Models for General Game PlayingBest AI papers explained · 22 min
Listen in VO