Probing Foundation Models for World Models

15 Jul 2025 · 12 min · 7 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Tests whether foundation models learn “world models” (underlying rules) or just predict sequences via heuristics, using an “inductive bias probe” rather than accuracy alone.

Guests

No guests mentioned; it’s a solo “Deep Dive” episode discussing a specific paper.

Key claims

High prediction accuracy can coexist with low bias toward true underlying laws; models may use non-parsimonious, task-specific shortcuts that don’t generalize.

Notable examples

Simulated Newtonian solar system: transformer (109M params) gets R2 > 0.9999 for next-position prediction, yet shows low Newtonian inductive bias; fine-tuning to predict force vectors performs poorly, and symbolic regression recovers nonsensical, setup-specific “laws.” Lattice tasks: inductive bias drops as state space grows; transformers worse than RNN/LSTM. Othello: ~90% next-legal-move accuracy, but poor board-state bias; models group states by legal-next-move sets (“coarsened representation”), and better state-inductive-bias models transfer better to tile-majority/balance tasks.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Exploring the Research Paper

0:25 to 1:06

Discussion on the paper that probes foundation models and their understanding.

“Using Inductive Bias to Probe for World Models.”

The Physics Experiment

1:06 to 2:18

Description of an AI model trained on a physics simulation and its predictive performance.

“Well, they introduced this really clever technique, something they call an inductive bias probe.”

Inductive Bias Probe Results

2:18 to 3:20

Insights into the model's low bias towards Newtonian mechanics despite high accuracy.

“Okay, this is where it gets really interesting.”

Symbolic Regression Findings

3:20 to 5:04

Exploration of the results from applying symbolic regression to the model's predictions.

“They fine-tuned the same model on a related but different task, predicting the actual force vectors acting on the planets.”

Further Applications of the Probe

5:04 to 6:40

Discussion on applying the inductive bias probe to simpler tasks and the model's performance.

“It was developing these task-specific heuristics, clever tricks or shortcuts that worked for predicting that particular sequence, but not a general unifying world model of physics.”

Implications of Inductive Bias

6:40 to 9:45

Examination of how inductive biases impact model performance and adaptability across tasks.

“It suggests that, at least in these specific tests.”

Concluding Thoughts

9:45 to 11:46

Reflections on the limitations of AI models in understanding underlying principles.

“It shows that this isn't just a philosophical distinction.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Okay, here's a question. When an AI model predicts something incredibly well-thin, like the exact orbit of a planet, does it actually understand the physics behind it? Hmm. Or is it just, you know, exceptionally good at matching patterns? That's really the core question, isn't it? And it's exactly what some fascinating new research digs into, this difference between just predicting and really truly understanding. Welcome to the Deep Dive. Today, we're looking at a paper called What Has a Foundation Model Found? Using Inductive Bias to Probe for World Models. It focuses on these big foundation models, the ones trained on just massive data sets for things like, well, sequence prediction.

0:40Right. And the central idea they explore, it's quite evocative, actually. They suggest these models might be a bit like, say, Kepler. Oh, the astronomer. Exactly. Brilliant at mapping out how planets move. Incredibly predictive. But maybe not quite grasping the why behind it all, the fundamental laws. That leap was more Newton, wasn't it? Precisely. Kepler described what happened. Newton explained why, with underlying laws. And this paper asks if today's AI is still more Kepler than Newton. So how do they test that? It sounds tricky. Well, they introduced this really clever technique, something they call an inductive bias probe.

1:13Okay. It's designed specifically to see if these models are learning what we might call world models. You know, those deeper rules or principles that explain why things work, not just what's likely to happen next. Interesting. So not just judging by the final prediction accuracy. Not at all. It's about probing the model's internal leanings, its biases towards certain kinds of explanations. Let's start with their big physics experiment. It really sets the scene. Okay. Physics it is. So they basically created a simulation. Yeah. Imagine a whole simulated solar system. You've got planets, a sun, everything obeying Newton's laws of motion, nice and clean.

1:53And they generated tons of data from this, like just tracking where the planets went over time. Exactly. Huge data sets showing these orbital trajectories. It's sequences of positions. And then they trained a big AI model on this data, a transformer, right? Yep. A pretty hefty one. 109 million parameters trained simply to predict the next position of a planet in the sequence. And how well did it do? That's the first big question. Okay, this is where it gets really interesting. The model was incredibly good. Its R2 score, that's a measure of predictive accuracy, was above 0.9999. Wow. So basically perfect prediction.

2:31Almost indistinguishable from the real simulated data. It could even generate really long, stable orbits just by predicting step by step. Okay, so looking purely at that R2 score, my first thought is, yeah, it gets it. It understands Newtonian physics. That's the intuitive leap, right. But that's exactly what the researchers wanted to test with their probe. What did it actually learn internally? And what did the probe find? Did it confirm that intuition? No, quite the opposite, actually. Despite that amazing predictive power, the inductive bias probe showed the model had a low bias towards Newtonian mechanics.

3:04Low bias. So it wasn't naturally leaning towards Newton's laws as the underlying explanation. Exactly. It could mimic the results perfectly, but the fundamental principles, the laws of gravity as Newton described them, weren't strongly represented in how the model was working. So they dug deeper into this? They did. They fine-tuned the same model on a related but different task, predicting the actual force vectors acting on the planets. Ah, so getting it to predict the F in EFTA, essentially, the actual gravitational pull. Precisely. That F equals GM1 meter over our Schwared kind of force, a cornerstone of Newtonian physics.

3:43And the results. Pretty poor, actually. The model struggled to predict those force vectors accurately. Which strongly suggests it hadn't really learned the core concept of gravitational force, even if it could predict the resulting motion. Right. And then they did something even more revealing. They used a technique called symbolic regression. Okay, what's that? It basically tries to look at the model's predictions in this case, its poor force predictions, and figure out if there's any kind of mathematical formula that could explain them. Sort of reverse engineering a physical law from the model's behavior.

4:13And did it find Newton's law hidden in there somewhere? Not even close. The paper describes the recovered laws as nonsensical. Nonsensical. Like what? Well, they give an example in Table 1. Something like force is proportional to sine of 1 over sine of minus 0.24 plus 1.45 times 1 over something involving R and M2. Yeah. Yeah. Okay, definitely not Newton. Sounds like mathematical gibberish almost. It really does. And what's more, it wasn't even consistently nonsensical. What do you mean? When they applied the model to different simulated galaxies, slightly different setups, it would come up with a different but equally bizarre looking law for each one.

4:54Ah, so it wasn't learning one universal, albeit wrong, law. It was cooking up specific weird rules for each specific situation it saw. That's exactly what it suggests. It was developing these task-specific heuristics, clever tricks or shortcuts that worked for predicting that particular sequence, but not a general unifying world model of physics. So it's like building a whole bunch of specific tools for specific jobs rather than understanding the general principles of engineering needed to build any tool. That's a great analogy and it raises a big question. Is that enough? For some tasks, maybe just predicting the next step using a heuristic is sufficient.

5:34But it seems like it would limit the AI's flexibility, right? If the situation changes slightly, the heuristic might break down. Exactly. And that lack of a true world model could really hamper its ability to generalize or adapt, which is why it's great they didn't just stop with physics. Right. They applied this probe idea to other areas, too. Yep. They looked at simpler lattice problems. Think of an agent just moving back and forth on a line according to some rule. Yeah. And also the board game Othello. Okay. What did they find with the lattice problems? For the really simple cases, with just a few possible states or positions on the line, the models showed pretty good inductive bias towards the underlying rule.

6:12Made sense. But as the lattice got more complex, more states, that bias dropped off quite a bit. And did the type of model matter? You mentioned the transformer before. It did, interestingly. The transformer actually performed consistently worse on these lattice tasks in terms of inductive bias compared to older architectures like RNNs or LSTMs. Hmm, okay. So maybe Transformers aren't always the best at picking up simple underlying structures compared to sequence-focused models. It suggests that, at least in these specific tests. Yeah. But the Othello results were, I think, even more surprising.

6:45Othello? The board game? How did that work? They trained models on sequences of Othello games, just like the Planet Trajectories. Task was simple, predict the next legal move. And performance-wise? Again, really good. About 90 % accuracy in predicting a legal move. Which sounds like it understands the game pretty well, right? Yeah, 90 % is solid. You'd think it has a good grasp of the board state and rules. But the probes revealed something else. When they tested its inductive bias towards the actual state of the Othello board, you know, which squares are black, which are white, the bias was poor.

7:19Wait, so it can predict legal moves accurately, but it doesn't seem to have a strong internal model of the board itself. How is that even possible? It's a bit mind-bending, isn't it? The researchers suggest the models learn what they call a coarsened state representation. Coarsened. Meaning it doesn't focus on the fine-grained detail of the entire board. Instead, it seems to focus primarily on the set of legal next moves available from a given position. Ah, okay. So it's like knowing from here I can legally play in squares A, B, or C without necessarily having a perfect internal picture of why those are the legal moves based on the full board layout.

7:59Exactly. It's learning the legal next token partition, as the paper puts it. It groups board states together based on what moves are possible next, rather than grouping them by identical board configurations. Subtle. But does that distinction actually matter in practice if it predicts the moves well? Well, this is where their other metrics come in handy and where we see the practical implications. They introduced R.I.B. and D.I.B. Right. Respecting state and distinguishing state inductive bias. Can you break those down simply? Sure. R.I.B. basically asks, does the model treat two inputs that represent the same underlying state, like the exact same Othello board, similarly?

8:36High score is good. Makes sense. Consistency. And D.I.B. asks, does the model treat two inputs that represent different underlying states, different boards, differently? Again, high score is good. It needs to distinguish. Okay. And what did these show for Othello? They showed that the model often treated boards that were actually different quite similarly if those different boards happened to allow the exact same set of legal next moves. Confirming that next token partition idea, it was grouping by available moves, not by board state. Precisely. But here's the really interesting part connecting back to usefulness.

9:10Models that did have stronger inductive biases, the ones closer to representing the true board state. They performed better when they were later fine-tuned for new tasks that required understanding the board state. Like what kind of tasks? Things like predicting who has more tiles, majority tiles, or the overall strategic balance of the board, board balance. Tasks where just knowing the next legal moves isn't enough. Okay, that's a practical consequence. So the models with a better internal world model of the board could transfer their knowledge more effectively to related problems. Absolutely.

9:45It shows that this isn't just a philosophical distinction. Having a better aligned inductive bias, a better world model, leads to more robust and adaptable capabilities. So bringing it all together then, the big takeaway seems to be revisiting that Kepler versus Newton idea. Exactly. Just because a foundation model spits out incredibly accurate predictions doesn't automatically mean it's learned the underlying laws or structure of the world model. It might just be a very, very sophisticated pattern matcher, a Kepler on steroids, perhaps. Or, as the paper suggests, it might be building up these task-specific heuristics, a whole collection of clever shortcuts.

10:23Non-parsimonious representations, they called it, meaning not the simplest or most elegant underlying explanation. Right. It finds a way to get the prediction right for the specific data it saw, but maybe not the fundamental way that generalizes easily. Which definitely highlights a major challenge for AI development, doesn't it? How do we push these incredibly powerful predictive engines to go beyond just mimicking sequences? How do we encourage them or even force them to build more robust, more generalized world models? That's the million dollar question. Yeah, it really is. And it leaves us with a pretty compelling thought, I think, as these AI models get even better, even more uncannily accurate at predicting things.

11:06Their sheer predictive success might actually hide this deeper limitation, a lack of real transferable understanding. It could. And that has huge implications for trust, especially when we think about deploying AI in critical areas, medicine, science, engineering. If they're operating on a complex set of heuristics rather than fundamental principles, how reliable are they when faced with something genuinely new? What does it mean if we're building AI that creates these intricate bags of heuristics instead of unified theories? It really makes you wonder, doesn't it? Maybe leaves the question for you, the listener, to think about how can we design the next generation of AI systems, systems that don't just reflect reality with stunning accuracy, but perhaps begin to truly comprehend it.

From the publisher

This paper investigates whether foundation models truly acquire a deeper understanding of underlying "world models" beyond mere accurate sequence prediction. Researchers introduce an "inductive bias probe" to evaluate how these models adapt to new tasks based on postulated world models, such as Newtonian mechanics for orbital trajectories or game rules for Othello. The findings suggest that while foundation models excel at their primary training objectives, they often fail to develop strong inductive biases toward the actual governing principles. Instead, they appear to rely on task-specific heuristics or coarsened state representations, leading to a lack of generalizability.

More from Best AI papers explained

All 475 episodes
Probing Foundation Models for World ModelsBest AI papers explained · 12 min
Listen in VO