In short
An “inductive bias probe” tests whether foundation models learn true underlying world rules (Newton) or only learn shortcuts that predict outcomes (Kepler).
Key claims
Models can achieve near-perfect next-step prediction without recovering the correct generative laws; their inductive bias often aligns with “next-token partitions” (grouping states by shared sets of next actions) rather than full state/world models.
Guests
None named; the episode is a host-led discussion with no identifiable guest backgrounds.
Notable examples
Newtonian orbital mechanics with a ~109M-parameter transformer: R² > 0.9999 for orbit prediction, but symbolic regression recovered nonsensical “gravity” formulas. Lattice/grid tasks: inductive bias toward true structure drops as state space grows; Transformers underperform RNN/LSTM. Othello: ~90–100% next-move legality accuracy, yet poor full-board inductive bias; models can predict legal moves even when internal board state is wrong. Language preference synthetic task: perfect RIB, variable DIB, suggesting grouping beyond explicit preference order.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOPrediction vs. Understanding in AI
0:29 to 1:26
Discusses the difference between prediction and true understanding in AI models.
“He could predict planetary orbits perfectly, knew exactly where Mars would be.”
Inductive Bias Probe Explained
1:26 to 2:56
Explains the inductive bias probe and how it reveals AI models' understanding.
“This technique they call an inductive bias probe.”
Testing Models: Newton vs. Prediction
2:56 to 5:13
Describes the tests conducted on a transformer model using Newtonian physics.
“beyond just, you know, getting the next word right.”
Surprising Findings: Understanding vs. Prediction
5:13 to 7:20
Reveals that the model could predict well but failed to learn the underlying laws.
“So it aced the prediction test, but failed the understanding test, at least for physics.”
Implications of Inductive Bias in Other Domains
7:20 to 9:48
Explores how the inductive bias probe applies to other domains and tasks.
“Well, the researchers proposed this really interesting hypothesis.”
The Gap Between Prediction and Understanding
9:48 to 11:06
Discusses the implications of AI's predictive capabilities versus its understanding.
“Good at the next step, but maybe fuzzy on the bigger picture.”
Transcript
Automatic transcript. May contain errors.0:28Welcome to the Deep Dive. amazing scientist. Oh, yeah. He could predict planetary orbits perfectly, knew exactly where Mars would be. But it took Newton to find the deep truths, right? The laws of motion, gravity, the why. That's a fantastic analogy. Prediction versus, well, true understanding. Exactly. And you see that same journey elsewhere, too. Think about biology. For ages, breeders saw patterns and traits. They could predict outcomes. Right, long before genetic. Long before Mendel figured out the actual mechanisms of inheritance. So observing patterns, that's prediction. Powerful, sure. But grasping the underlying rules, that's understanding.
1:06So the big question is, where do today's AI models sit? Are they Keplers or Newtons? Precisely. And that's what this new research paper tries to tackle. They've got this really clever new tool to probe exactly that. Ooh. And honestly, the findings are, well, they're pretty surprising. Okay, let's unpack this then. This technique they call an inductive bias probe. Sounds technical, but the idea is pretty neat, right? Yes. Instead of trying to look inside the AI's black box, which is super hard. Notoriously difficult, yes. They just look at how it acts when it gets a little bit of new information.
1:41It's like watching a scientist, you know? Their understanding shows in how they infer things from limited data. That's the core insight. A model's, let's call it implicit world model, if it has one, is revealed by its inductive bias. That's just its natural tendency, its default way of guessing rules, when it only has a few examples. Okay. So if a model has learned the real rules of a world, its guesses, its inductive bias, should line up with those rules. Makes sense. How do they measure that? They use two main metrics. First is respecting state or R.I.B. R.I.B. Got it. That basically checks.
2:17Does the model give the same prediction for different inputs that actually represent the same underlying situation? Higher ID means it gets that, it respects reality. Okay. And the second? Second is distinguishing state or DIB. Yeah. This checks the opposite. So then it gives different predictions for inputs that represent different underlying situations. Ah, so it needs to tell things apart too. Exactly. High DIB means it can distinguish different realities. You need both. It's easy to get a high RAB by just predicting the same thing all the time. Ah, yeah, that would be very useful. But then your DIB would be terrible.
2:50So getting both right is key. And it really makes you ask, what does understanding actually look like in an AI? beyond just, you know, getting the next word right. So this probe, this RIB and DIB, where did they first aim it? What was the first test case? They started with something where we know the rules perfectly. Orbital mechanics. Newtonian physics. Yeah, they simulated planets moving according to Newton's laws and then trained a big transformer model we're talking, like 109 million parameters, just to predict the next position of the planets. And how good was it at predicting, like Kepler?
3:24Oh, incredibly good. The paper says R squared values over 0.9999. Wow. It believes simple methods out of the water generated long, accurate orbits. Pure prediction. A plus jink. Okay. Prediction. Check. But the real test, did it learn Newton's laws? Did it become Newton? That was the probe's job. So instead of just predicting where the planets went, they fine-tuned the model to predict force vectors. Ah, the Y. Like the actual gravitational force, FPM 1 meter 2 over RU squared. or that stuff. The cornerstone of Newtonian mechanics. Testing the why, not just the way. What happened? This is where it gets surprising, you said.
4:01Yeah. The Transformers force predictions. The paper calls them poor. Poor. After being so good at predicting the orbits. Yep. And get this. They use something called symbolic regression. It's a cool technique, trying to figure out the actual math formula the model seems to be using internally. Like reverse engineering its little rule book. Kind of. Yeah. So they tried to recover the law of gravity the model was implicitly using. Did it look like Newton's law? Not even close. They found nonsensical formulas. Like one example they gave was something like F is proportional to the exponential of 1 over R times sine of E to the power of M2 plus 1.1 times R.
4:41Whoa. OK, that is definitely not Newton. That's bizarre. Right. It's completely unlike Newton's elegant universal law. As you said, if you handed that in for physics homework. Yeah, you'd fail. So what does that mean? It knew where the planets would be, but its reason why was just gibberish. It suggests something really fundamental. The model didn't seem to learn a general universal law. It seemed to develop these task-specific heuristics, shortcuts. Like it learned different weird made-up rules for different sets of planets instead of one rule for everything. It learned a recipe for specific cakes, maybe, but had no clue about the general principles of baking.
5:19Okay, that's counterintuitive. So it aced the prediction test, but failed the understanding test, at least for physics. Oh, pretty much. Did they test this idea elsewhere? Is it just a physics thing? No, they didn't stop there. They used the probe on other domains, too, ones where we also know the underlying rules, like lattice problems. Lattice problems, like an agent moving on a grid or a line. Exactly. Simple worlds with clear states. And there, as the world's got bigger, more states, the model's inductive bias towards the true structure just dropped off. Interesting. And notably, Transformers actually did worse than older models like RNNs and LSTMs on those specific tasks.
5:55Huh. Okay. What else? You mentioned games. Yeah. Othello. The classic board game. Right. Eight by eight board, flipping disks. So they tested various models, Transformers, RNNs, LSTMs, even Mamba. And prediction-wise, right again, predicting the legal next moves. They were hitting like 90, 100 % accuracy. So they knew how to play the game, basically. They knew the next legal move. but the probe showed they had poor inductive bias towards the full board state. Wait, how can that be? How can you know the right move if you don't understand the board? This is the really fascinating part. The paper points out that often, even when the model's internal prediction of what the Othello board looked like was actually wrong, the set of legal moves you could derive from that wrong board still perfectly matched the legal moves from the actual true board.
6:45That out! Seriously. Seriously. It's like it learned just enough about the board configuration to figure out which squares were playable next without necessarily having a complete accurate picture of the whole board. So it's predicting the right moves without really seeing the whole board accurately. That feels weird. Like a clever trick, not deep understanding. That's the interpretation. And it fits the pattern. It suggests they're using a kind of simplified or coarsened view of the world state. Just enough to get the next step right. OK, if they're not building these complete world models, we kind of assumed they were.
7:19But they're still so good at predicting. What are they learning? What's the alternative explanation? Well, the researchers proposed this really interesting hypothesis. Maybe they suggest foundation models develop an inductive bias towards what they call next token partitions of the state. Next token partitions. OK, break that down. It basically means they group situations together, not based on the full underlying state being the same, but based on the set of possible next actions being the same. Ah, OK. So if two different board positions in Othello have the exact same set of legal moves available.
7:57The model might treat them as kind of the same or at least very similar, even if the actual patterns of black and white disks are quite different. Right. Because all that matters for the next prediction is that list of moves. Precisely. So they tested this. They refined that DIB metric, the one for distinguishing states. They split it into DIBQ, which measures predictability for different states that happen to have the same next legal actions. And DIBQ, which measures predictability for different states with different next legal actions. And if the hypothesis is right. If they're using these next co-compartitions, they should find it harder to distinguish between states when the next actions are the same.
8:33So DIBQ should be lower, indicating more predictability or confusion between those states. Was it? It was. For both the lattice problems and for Othello, the results were statistically significant. Models were more predictable, less able to distinguish between different states that shared the same set of legal next moves. It strongly supports this idea they're latching onto these next token partitions. Wow. That's a subtle but really important distinction. It's not about the world. It's about what you can do next in the world. Seems like it, yeah. Did they see this in language models too, the LLMs we use every day?
9:09They did look at LLMs, yes, in a synthetic task involving preferences. Okay. The LLMs showed perfect RIB. They were consistent when the input preferences were identical, so they respected the state in that sense. Good. But their DIB, how well they distinguished different preference sets, was much more varied. It suggested that they weren't just learning a simple, clean ordering of preferences. So like in Ocello? Kind of. It implies they were likely extrapolating based on more than just the pure preference order itself, maybe grouping things based on linguistic similarities or, again, these kinds of next-token outputs leading to predictability across states you wouldn't think were related.
9:49Same pattern again. Good at the next step, but maybe fuzzy on the bigger picture. That seems to be the consistent finding across these different domains. Okay, so let's try and synthesize this. What's the big takeaway from this research and this probe? The big picture seems to be this. Foundation models. Undeniably powerful sequence predictors. Predicting orbits, predicting legal moves, predicting the next word they excel. Really impressive performance. Totally. However, this inductive bias probe suggests a pretty significant gap. They often seem to lack a strong limited inductive bias toward genuine world models.
10:26They don't seem to default to learning the simple underlying rules. Instead, they appear to rely on these Corson's state representations or maybe non-parsimonious representations. Basically, complex, task-specific heuristics. A big bag of tricks that work really well for prediction. They learn what to do next extremely well. Incredibly well. But maybe not how the world fundamentally works in a generalizable way. Which brings us back to Kepler and Newton. They're maybe more like Kepler's right now, amazing predictors, but maybe not quite grasping the underlying physics yet. That seems to be what this evidence points towards, yes.
11:00So this raises a really big kind of final thought-provoking question for you listening. If these incredibly powerful AI models are getting so good by learning these sophisticated shortcuts, these heuristics, rather than the deep underlying world models, what does that actually mean? Especially for situations where we need a deep, generalizable understanding of reality. Can we ever fully trust an AI's judgment or understanding if it might be built on this, well, this bag of heuristics rather than a coherent theory of the world? It's something to really mull over. Think about your own field, the problems you work on.
11:36Does AI need to move beyond just prediction? Is just being a really, really good Kepler enough? Or do we need our AIs to eventually become Newtons? Something to think about.
From the publisher
This academic paper introduces a novel "inductive bias probe" to evaluate whether foundation models truly grasp underlying "world models" or simply excel at predictive tasks through task-specific heuristics. The authors illustrate this by showing that a model trained to predict orbital trajectories, while highly accurate, fails to apply Newtonian mechanics when adapted to related physics problems. The research extends this analysis to other domains like lattice problems and Othello, consistently revealing that these models often develop biases towards simpler, "legal next-token" patterns rather than the full, complex state of the world. Ultimately, the paper suggests that stronger inductive biases toward a known world model correlate with better performance on new, related tasks.




