In short
Dan Roberts (OpenAI) explains why reinforcement learning (RL) and “test-time compute/reasoning” are enabling AI systems to solve hard scientific and mathematical problems, including recent breakthroughs on Erdős problems. He contrasts OpenAI’s informal reasoning approach with DeepMind’s Lean/formal proof approach, and argues RL is now central (“the cake”) for turning compute into intelligence.
Guest background
Dan Roberts is a top AI researcher at OpenAI, leads the Foundations of Reinforcement Learning team. He has a PhD in theoretical physics from MIT (quantum gravity/quantum information; black holes, quantum chaos), did a postdoc at IAS, then moved to FAIR to apply theoretical-physics tools to deep learning; co-authored The Principles of Deep Learning Theory; previously worked at Sequoia as entrepreneur-in-residence.
Key claims
AI science progress is gradual, not a sharp “scientist” switch; RL helps models “think” via long token-based scratchpad-like reasoning; verifiable rewards work best in domains like math; smooth scaling matters more than “emergence.”
Notable examples
OpenAI’s unit distance proof (contrarian disproof via long exploration); OpenAI’s Erdős problem result (lower bound conjecture false); DeepMind uses Lean + auto-formalization for airtight proofs.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOIntroducing Dan Roberts
0:45 to 1:18
Matt introduces Dan Roberts and his background in AI and physics.
“In this conversation, we go deep on what reinforcement learning actually is, why it's the most important paradigm in AI right now, and what's ahead for AI and science.”
Dan's Background and Path to OpenAI
1:18 to 3:16
Dan shares his academic journey from physics to AI at OpenAI.
“You are the lead of the foundations of reinforcement learning team at OpenAI.”
The Evolution of AI in Science
3:16 to 7:04
Discussion on the gradual evolution of AI's role in solving scientific problems.
“I did a PhD in theoretical physics from MIT, thinking about the intersection of quantum gravity and quantum information.”
Contrarian Approaches in Mathematics
7:04 to 10:32
Exploring how AI approaches complex mathematical conjectures.
“Where do you think we are in the evolution of AI being increasingly able to solve difficult scientific problems?”
Comparing AI Approaches to Proofs
10:32 to 12:22
Comparison of OpenAI's and DeepMind's methods for formal proofs in mathematics.
“One of the approaches that GDM takes is to take problems, present them in a formal language called Lean, and then use methods to search for proofs in that language.”
Understanding Reinforcement Learning
12:22 to 14:00
Dan explains reinforcement learning using engaging analogies and examples.
“What is the one, two, three sentences definition for reinforcement learning?”
Understanding Reinforcement Learning Basics
14:00 to 17:04
Learn the fundamental concepts of reinforcement learning, including rewards and environments.
“Maybe the first thing you do is you run, you hit the first bad guy, and probably this example is dated, but you lose a life.”
Application of RL in Language Models
17:04 to 18:20
Discover how reinforcement learning from human feedback (RLHF) is applied to improve language models.
“Now let's talk about how RL has been applied in the context of large language models.”
Self-Play and AI Competitions
18:20 to 19:55
Explore the concept of self-play in AI and its application in poker competitions.
“And then now, because during the training process, you can't just pause your training run to ask some humans for input.”
Explore vs. Exploit in Scientific Discovery
19:55 to 22:26
Learn how exploration and exploitation concepts apply to scientific research and AI.
“This is great actually for me because I learned some really exciting work in AI and I got very excited about this while I was doing physics.”
Show all 22 chapters
The Role of Reinforcement Learning in AI
22:26 to 24:50
Understand the evolving role of reinforcement learning in AI systems and its historical context.
“Presumably, the instinct would be that you need a lot of exploration, not exploitation.”
Efficiency of Reinforcement Learning Methods
24:50 to 28:07
Evaluate the efficiency and effectiveness of reinforcement learning in modern AI systems.
“systems so there used to be a saying which I think comes from Yann LeChan that RL was the cherry on top of the cake but I think you have argued that things have switched now and that RL is the main part, the cake?”
Exploring Reinforcement Learning's Role in AI
28:07 to 29:00
Learn about the significance of reinforcement learning in achieving breakthroughs in AI.
“So I think you can see where that logic comes from.”
Debate on Intelligence and Language in AI
29:01 to 30:49
Uncover the debate between different scientific approaches to AI intelligence and the importance of language.
“Somewhat famously, last year, there was a conversation with Rich Sutton on the Dworkish podcast where his claim, in my best attempt to paraphrase it, was that LLMs were not really intelligence.”
The Interplay of Scale and Ideas in AI
30:50 to 32:53
Discuss the importance of both scaling and conceptual breakthroughs in AI development.
“interested in AI, the sort of grounding, I think, that was needed was to make things really work, is through language, because everything goes through language, right?”
Understanding Test Time Computation
32:54 to 34:50
Gain insights into how models generate answers and the role of language in the thought process.
“What actually happens during test time compute that creates those artifacts?”
Challenges of Reinforcement Learning in Various Domains
34:51 to 36:57
Examine the limitations of reinforcement learning in creative and subjective fields.
“Because that's effectively what you described earlier when you were defining RL.”
Generalizing AI Intelligence Through Reinforcement Learning
36:58 to 38:54
Explore the potential for reinforcement learning to enhance general AI capabilities.
“I mean, clearly there's tremendous progress in those domains, but like what is happening?”
Physics and AI: Lessons from Complex Systems
38:55 to 42:00
Discover how principles from physics can help in understanding and developing AI systems.
“And I'll get to why physics really matters for this in a second.”
The Role of AI in Scientific Discovery
42:00 to 45:24
Exploration of how AI contributes to scientific advancements and predictions on its future capabilities.
“to thermodynamics, meaning a compact theory that predicts behavior without tracking every individual bit?”
The Future of AI Autonomy in Research
45:24 to 47:29
Discussion on the potential for AI to autonomously build other AI and its implications for research.
“like this is clearly, I think the unit distance problem is a great example.”
The Open Questions in Science and AI
47:29 to 48:36
Reflections on the lasting questions in science and the potential for AI to help answer them.
“And obviously we'll turn this sort of thing on AI itself and the models will get a lot more powerful and that'll be fun.”
Transcript
Automatic transcript. May contain errors.0:00One of the things that ChatGPT was able to do was assume it was false. When you go against the grain and do something contrarian like that, you really have to have strong conviction in what you're doing in order to persevere down a really long calculation path. I feel really excited that we will get to really answer a lot of fundamental questions in the field of science that we care about, with the aid or the models being the driving force. And so that's just really thrilling. Hi, I'm Matt Turck. Welcome to the Matt Podcast. It's been yet another extraordinary last few days in AI, with OpenAI, DeepMind, and Anthropic cracking some of the most famous long-insolved questions in mathematics, known as the Erdos problems.
0:37A moment many view as a stunning breakthrough and yet another signal that AI is moving from doing the work we ask of it to autonomously making deep science discoveries. To unpack the moment and the fundamental advances in model reasoning that make it possible, I'm excited to welcome Dan Roberts, a top AI researcher at OpenAI, who comes from a deep background in theoretical physics and has a particular interest in the intersection of science and AI. In this conversation, we go deep on what reinforcement learning actually is, why it's the most important paradigm in AI right now, and what's ahead for AI and science.
1:12Please enjoy my conversation with Dan Roberts. Hey Dan, excited to do this. Thanks for taking the time. Of course, very happy to be here. You are the lead of the foundations of reinforcement learning team at OpenAI. So what does that mean? What does the name mean? The larger team that we're on is called foundations. And we think about reinforcement learning. So very boring foundations of reinforcement learning. But the team comes from a mandate of thinking about the science of reinforcement learning. and a long time ago, which in AI speak is like six months ago, maybe a year, I guess now two years.
1:54So before we released a one and thinking reasoning models, we were studying this internally. And, and one of the advantages to being first, or at least to being forced and spending a lot of resources on scaling things up is that you can empower a group of people to not just work on making the thing work, but work on understanding how it works. And then beyond that, how do we scale? How should we think about scaling reinforcement learning versus scaling pre-training? So what are scaling laws look like? But then going beyond that, what sort of things does this kind of training teach us? What doesn't it teach us?
2:32We're very interested in at the frontier for exploratory scenarios. How do we either improve or understand better what reinforcement learning is doing? We have all this compute that famously we are in the process of acquiring and we would like to turn that compute into intelligence. And to do that, we need to make thinking models and somewhere along the way, we interact with that process, usually at the earlier stage for models, not the next model, but things that are like the next model or the next next model. Great. And quickly, what was your path to OpenAI? So how did you go from studying physics to being where you are today?
3:16I did a PhD in theoretical physics from MIT, thinking about the intersection of quantum gravity and quantum information. Thought a lot about black holes, quantum chaos kind of thing of what if you throw something into a black hole? What happens to the information? Does it come out? If we think about black holes as computers, how fast are they? I was very interested in this fundamental question in theoretical physics, which is how do you find a quantum theory of gravity? I also got very interested in this interplay between computation and the laws of physics. Any computer exists in the universe and behaves according to physical law.
3:56So the sort of computations you can do are bounded by the laws of physics and there's some sort of interesting relationship there black holes are pretty interesting because they sort of saturate some conjectured bounds around processing of information and from there i did a postdoc at the institute for advanced study and around that time i'm pretty old now for at least for this field so that was about 2016 uh was when the dqn atari paper from deep mind happened in 2015 and then alpha go was in 2016 and i got very excited about the possibility of machine learning and then deep learning was statistical science that lived in a similar framework to the sort of frameworks that we use to study the rest of the universe.
4:44And so there's always this question of, how does everything work? This three-year-old question of, I'm curious about everything. And if you look outward and you care enough, you end up philosophy maybe. If you're quantitative, you end up in physics. Very, very crude characterization. And AI that works is very fascinating. Or AI systems that work are fascinating because they are simple examples that do things that humans do. And then if it lives in the same framework that we use to understand everything else, then, you know, it's sort of like you can draw the parallels between how does the universe work and how do I work or how does intelligence work.
5:22So I got extremely interested in AI and deep learning. Then I went to FAIR, Facebook's AI research lab, around 2017 to basically try and use the tools from theoretical physics to understand deep learning. Deep learning was supposed to be this really difficult thing that you couldn't understand. And I thought maybe the tools of physics could be helpful. this actually culminated in a book that I wrote with a collaborator who now is still a collaborator of mine now at OpenAI working on the same thing as Shoyeta but we wrote this book The Principles of Deep Learning Theory that is a culmination of these sets of ideas of can we sort of use the statistical ideas of understanding statistical systems like the gas in the room, we can characterize them with some simple laws of thermodynamics like the ideal gas law, and maybe we can make similar progress in understanding deep networks.
6:24So that was sort of my transition. I also had a startup along the way and spent some time at Sequoia Capital as an entrepreneur in residence. So there's some tension between am I a scientist and am I an entrepreneur, but about two years ago, after thinking about whether I want to start another AI company, I realized that the thing that was most exciting right now was what was happening at the frontier, that there's some amazing scientific progress happening in AI. And to really get at the questions and understand what's going on, you need to be there and you need to participate. And that meant joining lab.
7:01So I joined OpenAI two years ago. Great. Thank you for that. Where do you think we are in the evolution of AI being increasingly able to solve difficult scientific problems? I mean, certainly something that we've been talking as an industry about for a while now, but it seems to be accelerating, perhaps just like everything else in AI. But where do you think we are? I think one of the interesting things is that this process is smooth. there's no sharp point, or I don't think there will be a sharp point where we'll say that systems weren't able to be useful for the scientific process to their fully-fledged scientists.
7:43There will be sort of a gradual shift. If you had to point to one moment, maybe it would be the release of O1 by OpenAI and the sort of paradigm of test time compute and reasoning. But I'm sure if I tried to make that claim, you could go and look at GPT-4 and see that there's glimpses of that sort of useful behavior for the scientific process were already present. As a general point, you know, the models are very good at certain types of things that clearly are amenable to making progress in math. They're not open loop, fully fledged scientists in any domain, although, you know, neither am I. It seems like it's just this really nice gradual process.
8:21So it feels like a particularly fun week to be having this conversation because over the last few days, there were a number of different announcements in the general field of AI and mathematics around the AirDosh problems. OpenAI came out first with this progress, but like almost within a few hours, Google DeepMind had a claim as well on different problems and anthropic had some claims however from what i understand the open ai approach and the deep mind approach were very different and that may be very interesting in terms of what that means for ai as a research scientist this conjecture everyone assumed was true and uh but could not prove it one of the things that chat gpt was able to do was uh assume it was false.
9:15And when you go against the grain and do something contrarian like that, you really have to have strong conviction in what you're doing in order to persevere down a really long calculation path. Because there's a lot of choices that you can make along the path. And if you get any of those choices wrong, if your ideas don't work, then you find out that you didn't make any progress. And so you need this really strong persistence. And then you need expertise in this other field, which is like algebraic number theory, some sort of generalization of number theory on things that sort of generalize the integers and the real numbers.
9:51You go down that path really far, you can refute this conjecture. So that was the big result. The big result was that this conjecture of this lower bound for the number of pairs that you can make is false. Not only is it false, it was false due to a really interesting connection to another field of mathematics. And so you would have to be somebody who is aware of this problem as interesting, which sounds like your expertise is one thing, and then be an expertise in something else, and then also be super contrarian and go down this really long path. And then you would have identified the solution.
10:25The OpenAI approach and the deep mining approach were very different. Do you want to compare and contrast the two approaches? One of the approaches that GDM takes is to take problems, present them in a formal language called Lean, and then use methods to search for proofs in that language. And some problems, for problems to be representable, there's this process called auto-formalization, where you take English version of the problem and you translate it into rigorous formal statements, and then you conduct your proofs there. And it's designed so that the proofs can be airtight. No one has to go and check for some hidden assumption or some weird thing.
11:08I guess it's usually hidden assumptions or definitions that are not airtight. But in that setting, which is a setting that DeepMind has cared a lot about, they were able to formalize some problems and use their system to prove them. So that's one approach. Another approach is to just take the problem in English with mathematical expressions as well, but just the English statement of it, which is informal, and understand what is meant by that and solve that in informal language, presenting a proof much like the way a human mathematician would or a human mathematician who's not using Lean. And then you have to check it.
11:49The verification problem is harder because it's not something that auto-checks. And that second approach was OpenAI. Most of our results that we publicize, as far as I can think, are all in the informal setting. We have language models that we've taught them to reason at test time. And one of the applications or benchmarks for that is reasoning in mathematics. Okay, great. All right. So let's get into reinforcement learning. To make this broadly accessible, let's start from the top. What is the one, two, three sentences definition for reinforcement learning? And perhaps give us a simple non-technical analogy.
12:33for people to understand? Maybe a simple thing to do would be to give you two examples of how you could try to learn something, you as an individual. And maybe we can take a game or even, say, a video game. I'm old enough where I played the original 8-bit Mario Brothers, the Super Mario Brothers. And so here are two ways you could learn how to play. One way you could learn how to play is your dad takes it out and plugs it in and he boots up the game and then he plays for a few hours. And then you just watch him play. That's all you do. So he's demonstrating how to play. And then at the end of that, and he's not very nice, so he doesn't let you play.
13:16But then he goes and runs outside and does something else. And you sneak into his room, you plug it in and you try to play. How good are you going to be? Well, all you've done is tried to memorize what he's done. You haven't gotten to push any of the buttons yourself. You haven't gotten to interact with the game yourself. This is sometimes called expert demonstrations. And you're just trying to memorize what someone else is doing. The version of supervised learning, the supervision being like you just watch what he does and accept that that's the true way of doing the thing. Reinforcement learning would be your dad's like, here, why don't you play?
13:52Maybe he shows you once or maybe he doesn't even need to show you because the game is beautifully designed to sort of take you from not knowing anything to being able to play expertly. There's something called a curriculum. But you play. Maybe the first thing you do is you run, you hit the first bad guy, and probably this example is dated, but you lose a life. But then the second time, you press a button and you jump. And so you're taking actions. There's an environment that's giving you feedback. And there's this close connection between the environment, between actions that you can take and then the responses that you're getting.
14:26And then the final part is there's a reward. And the reward can be something that you get pretty often. For instance, every time you do something, there's some score that goes up. Or it could be just something that you get at the end. So you play a game of chess and at the very end, you get a reward, which is you won or you lost. But in the middle, you don't really know how you're doing until the very end. And so this is called sparse resistance rewards. But I think this is the basic idea in that, and there's obviously lots of variance here and ways to quibble with this. But it's this notion that you interact with an environment, you get a reward, and often it's in a way where you get this sort of feedback as opposed to just trying to learn from data that you don't get to interact with.
15:10And why does it work and why is RL so powerful? It works because of this ability to get feedback from the environment. You can go and learn, you know, if you're doing it right, you can figure out how to learn the things that you don't know. And I also think it's powerful because of this fact that it's much easier to learn when you're learning at the right level for you, right? So if you want to learn, you know, addition, you shouldn't read a calculus textbook. You want to learn by being able to practice and learn at the right level. I'm actually making the choices and learning from my own choices, whether they work or not, then I'm able to place it in a better context for the set of things that I understand.
15:58Great. And then conversely, what's the catch and how does RIL break? The setting where very difficult is the setting that I alluded to before where you don't get much feedback from the environment. You have to take many, many, many, many actions and then you get maybe, yes, that whole set of actions was good or no, it was bad. For instance, you're playing a game of chess and you don't know until you make all the moves. That has an opponent, so it's maybe complicated. Maybe it's you are trying to do a homework problem and it's a research level. Or someone gives you a well-defined problem, like we give our language models.
16:37And it's a problem that requires days and days of thinking. There's so many choices that you can make along the way. And at the end, if you don't get any feedback at all, if you're just hidden in the woods by yourself, scribbling in notebooks, it's very hard to make progress that way because you don't have any sense. If you get a yes at the end or you get a no at the end, you have no sense for which of the actions that you took, which of the things you did were good or bad. Okay, great. Now let's talk about how RL has been applied in the context of large language models. So was the first step historically RLHF?
17:16Yeah, I think that's probably fair, at least in a broad sense, that the first kind of RL that was done on language models was part of this post-training process to turn a model that just tries to predict the next word on the internet into either something that will follow your instructions, be nice to you, or fit the form of a chatbot. So do you want to define for people what RLHF is and sort of hide work quickly? The basic idea is that you could use, collect data from humans. That's so that RLHF is reinforcement learning from human feedback. So you collect data from humans and you train a value function.
18:01So you would show in the language model setting, say two different completions from a language model, ask them to say which is better. This sort of comparisons could be used to train a value function. And then you can use that as a reward for reinforcement learning process. Great. And you do that initially with humans, but then you build that into a reward model. Yeah. So you would train a model for this. And then now, because during the training process, you can't just pause your training run to ask some humans for input. Right. The feedback that would have way too much latency. So instead you need a proxy for what a human would say.
18:41So you train this model based on the human preference data, and then you can optimize against it, or at least a little bit. One of the famous things in the history of RL is Move 37. How do you train a model to encourage the model to do that kind of things and come up with brand new ways while being efficient and exploit known paths? Yeah, so the great thing about Go is that you can just train it. It's a zero-sum two-player game. You can train it in what's called self-play. it plays itself and it can go from playing randomly to expert play and it will find whatever the sort of best strategies are.
19:20So if that means exploring, great. That means exploiting. Actually, I have a funny story about this. So I met Noam Brown in grad school. He went to a different grad school than me, but he wanted to enter MIT's PokerBot competition. And he had a poker bot that was the best in the world, but it wasn't something that would compete against humans yet. He just won in this research competition. He collaborated with me and another friend to enter MIT's poker bot competition. This is great actually for me because I learned some really exciting work in AI and I got very excited about this while I was doing physics.
20:03We were playing essentially this kind of self-play equilibrium strategy. There's some nuances, but essentially we could not lose assuming we did not have any bugs in our code. The way this thing worked was that it was a tournament where you would be paired with, say, another person and play them. Depending on the amount of points you got in some sort of round-robin setup, they would eliminate the bottom half and they would keep going until you got to the final table, which would just be say you versus the other person and so the scores there was the award ceremony and we didn't know what what happened but the there was someone else who was um you know what what did the what was everyone's scores over time look like and there was you know say 64 i think there were 32 actually people playing so it was like around a 32 kind of tournament and 30 people over time you know their scores would were all very negative and going down and then there was one person whose score was like pretty much straight up and then there was another that was like pretty good but not like with a crazy slope and so do you want to guess which one we were?
21:14So we were the lower slope and then there was this other guy that had this crazy slope was just like completely crushing all the other players and then this happened for the round of 16, the round of 8, the round of 4 and then in the round of 2 it's heads up, us versus this guy who's like over the course of this tournament won way more than us uh overall like taking more money from from everyone else and then we crushed him because why because he was exploiting the weaknesses of of of everybody else right it's had some theory of mind to try to figure out oh this this guy you know does this when he bluffs and and so it was very i assume it was very good at like taking um you know exploiting everyone else but we were just playing the best possible thing that you could that you can do given you know so the the criteria was not maximize your your amount that you get from anyone else it was don't lose like put you know and and uh and you know so it's the best response to anyone's strategy and so at the end we had to win assuming we did it right and someone else playing the same strategy would would tie okay fascinating so just uh tying this back to the beginning of the conversation about the erdoge problem and and solving unsolved math problems.
22:26Presumably, the instinct would be that you need a lot of exploration, not exploitation. So how does that work in the context of novel scientific discovery? I think math research or scientific research in general has a lot of versions of both explore and exploit. To give the recent example, the OpenAI unit distance proof, I think, is very much in the explore setting where the model was happy to be contrarian and try to disprove this thing that everyone believed. And it was just looking for, it has this huge repository of understanding all of human math. And so it was spending a very long amount of time.
23:10I forget how many hours, but I think we published a rewritten version of this chain of thought, but like hours and hours trying different things. So it's clearly in the domain of exploration. a lot of times though you can ask these models to compute something that they understand very well and then that has a different structure and might look a lot like exploit. There's a paper that came out recently after the OpenAI result where the unrelated Erdős problem has something to do with if you have a set and you try to add the set to itself or you try to multiply the set with itself so take the elements and add them all together or take the element individualize or multiply them together and how many unique sums or products you get there's some conjecture around that and this this one was also disproved and and that was done by by humans and the core idea was um it's like a totally different problem but there there was inspiration from the from the unit distance one the idea that you you can sort of generalize uh from from the pick a certain type of numbers that had a certain property that the open AI model figured out and that they realized that like this applies in this setting so that's very much an exploit thing but and so so I think the process clearly is like this the actual discovery process I think normally when you talk about explore exploits maybe we're talking about when training reinforcement learning models how should we train them but I think there's this interesting point that in the scientific discovery process there's really this interplay between like exploring exploration and then exploitation in order to totally push the field forward switching to RL in modern LLM systems so there used to be a saying which I think comes from Yann LeChan that RL was the cherry on top of the cake but I think you have argued that things have switched now and that RL is the main part, the cake?
25:12Do you want to just walk us through what you were thinking? Yeah, I said that about a year and a half ago. I had to give a talk that was public and I couldn't say much. So I decided to invert this meme with this cake and the cherry. RL is really exciting. That's what I'm here talking about. And I think that when you have a lot of compute, you want to turn that compute into intelligence in a way that's useful. And RL is one way of doing it. And we just started doing it then. And we're going to do a lot more of it now. Why did RL start working well? It's not an entirely new concept. It's been tried for many years now.
25:56What is different now? Yeah, I'm not sure, to be honest, when people say it wasn't working, what that actually means. There was this 2016, 27, maybe even to 2018 before the transform period where DeepMind was all in on RL and OpenAI had Dota and Rubik's Cube and some other exciting results as well. But a lot of people were all in on RL and then there were language models. And the obvious thing to do was scale up the thing that worked, which was pre-training. and I don't know whether or what people tried for RL as you pointed out RLHF was a central thing that came pretty quickly originally was developed for in the in the context of game environments of like trying to prevent reward hacking by using I think the original paper was about using human feedback to like control like a character for what or something like that but there's an interesting thing to point out here, though, which is that there's this question of how do you get models to think in test time and reason?
27:07And there was a reasoning effort at OpenAI that was quite early and spent some time and came up with some algorithms. I think maybe the simple thing to say is that if you have a powerful enough pre-trained model, then it can start to do well at RL. It can start to think at use test time compute to, for instance, solve math problems that it wouldn't otherwise be able to do. There's a viral analysis from earlier this year, February, I think, that claims that RL produces less than one bit of useful information per 10 ,000 tokens. And then Carpathie called it sucking supervision through a straw. What is your take on this and the overall efficiency of RL?
27:52If you look at the DeepSeq algorithm, which is a public thing that we can talk about, then you train on sequences that are correct. So whether it's correct or not is maybe one bit of information. So I think you can see where that logic comes from. I think the question is, is this doing a kind of thing that you can't otherwise do? Maybe you would want to give more supervision, but how are you going to do that? But I think it's very clear that these methods have led to a bunch of breakthroughs in terms of the explosion of what the models can do, and both in coding and in science. I think broadly it's about getting models to think in test time, to use test time compute and do reasoning.
28:41And there's clearly a lot of the pieces of what the RL process is that's essential to make that work. What's your overall feeling in terms of how far we can go with that current sort of systems model where we have pre-training and then we have RL on top? Somewhat famously, last year, there was a conversation with Rich Sutton on the Dworkish podcast where his claim, in my best attempt to paraphrase it, was that LLMs were not really intelligence. And therefore, RL was the only way to do it. And pure RL, not LLM plus RLs. What is your take on this? I mean, obviously, you're on an RL team at a company that does both pre-training and RL combined.
29:30So what's your take? Let me tell another story. So before I did my PhD, I spent two years in the UK, and I was at Oxford for one of those years. And I was at a pub, as one does, and two of my close friends, one was a cognitive scientist, and one was a linguist. And so we had that sort of argument that you do in those situations when you're that age. and so something like physics is the most fundamental of all the sciences because it explains how the world works and everything is in the world i said this earlier my computer exists in the world i exist in the world we all follow the laws of physics and then the cognitive scientists said something like yes but then you have to you know you have to process it so there's all sorts of cognitive biases about that and and you know the way you collect data and learn something something but then the linguist was was like blah blah blah wichtenstein you know the everything goes through language.
30:24That's the method of communication. That's, you know, that's the way words mean things are the central thing. And when we want to talk about the laws of physics, we have to use language. And I sort of feel like he and Victor, like that, that was correct, right? That's what's what, or at least the path through AI suggests that that is a correct path. I'm conceding now to Kyle. And if he's, if he's listening, he's now a linguistics professor. This whole idea of reinforcement learning that kicked off the previous decades interested in AI, the sort of grounding, I think, that was needed was to make things really work, is through language, because everything goes through language, right?
Read the full transcript
31:02All the internet, right? It incorporates the grounding of the real world, all of our scientific knowledge, all of our mathematical knowledge, all of the human, like, you know, the sum total, basically, of human work is represented on the internet in language. And then so having the model have a prior of language and being able to like think in language and then train on top of that that seems like clearly the right thing to do and like you know seems also well grounded in a way that even before all this somebody might have argued would make sense. It's like an amazing prior to have to start with for an intelligence because it's like very much based on us and our society.
31:42I have other disagreements with Rich Sutton but if you want to poke at that. Yeah, so just give us one or two quick ones. I have a somewhat contrarian take with the better lesson that it's not that scale is all you need. You need to also have good ideas to guide the scaling. So there's a deeper interplay than just scale things up. For instance, if you were just trying to scale pre-training, you wouldn't get anywhere near as far as also trying to scale RL on top of pre-training, which is what we do now. And our models are much more powerful for that very good idea. investing in that good idea. And that the good ideas come from humans?
32:20Well, maybe they'll come from AI in the future, but before we had AI, they came from humans. I mean, scaling was also a good idea that came from humans, but there's like this interplay of illicit new phenomena at scale. You try to understand them at that scale, and then that points you at new directions, and then you develop new ideas, and then you try to apply scale on those ideas. So I think it's not just scale, scale, scale. Since you mentioned test time compute, I think there's something that still puzzles people, which is the whole chain of thought thing, which is so magical from a user perspective, whatever you can see.
32:54What actually happens during test time compute that creates those artifacts? What does the model actually do? I think it does what you see it do. We lightly rewrite it or summarize it, but it just produces tokens. And those tokens are like a running thought process, just like you might have, or maybe it's more akin to if you're solving a math problem, the scratch pad, the collection of notes that you have, but it just keeps generating. The cool thing about generating is that, you know, it's a forward pass the model. So we're using a bunch of computation. So we're, you know, it's a way of leveraging a lot more computation on a problem than you would before.
33:36So my colleague, Noam Brown likes to talk about the Riemann hypothesis a lot and you know wouldn't you want to have a model that runs for years that that that can um resolve resolve that prove that if you present it and you want to produce an answer then it only has the number of flops in a single forward pass to produce one one token if it's forced to answer right away but if it gets to answer after thing you know after a long time it can it can re reuse its weights you know produce um a final answer that is a function of a much, much larger amount of computation. And the natural way it thinks is in language.
34:13It's a language model. And so that's sort of this key insight that you can cause it to do better just by producing a thought process in token space, in language. And this was known before RL, the idea that if you asked the model, if you gave a model examples of thinking things out, it would do this before it produced a final answer. Or if you just told it that, then it would do this sort of thing, right? Going back to this SFT versus supervised learning versus reinforcement learning analogy that I gave earlier, like there's a lot of examples on the internet of people thinking for a long time. And so like it's not completely useless.
34:47It can channel that a bit, but RL really like brings that out. What happens during TSM compute is RL related or created? Because that's effectively what you described earlier when you were defining RL. The model goes in one direction, decides maybe that's not a fruitful one, backtracks, tries something else. Is that correct or not? I think maybe the result of the RL process is that the model can then think at test time. And that's why we have these dials or various companies have labs, have reasoning effort dials, right? So you now created a model that will produce a bunch of tokens before it outputs a final answer.
35:29And like causing that to be good is what RL is doing. or one of the things RL is doing. And so the output of doing RL training is the ability to have a model that thinks. So one of the key questions in the field is whether you can expand and generalize the success that LLM systems have had, particularly in coding and now math, but like domains where you can sort of verify whether what the model comes up with is correct or not. What is your view on that? And perhaps start by explaining what a verifiable reward is. So a verifiable reward is, in principle, a reward that can't be hacked. So if it's a math problem and the answer is an integer, you just string match the integer, and then you verify that it solved the problem correctly.
36:23That abstraction has all sorts of problems with it, but a problem that can't be verified is, is this a good piece of creative writing? that there's not something you can sort of string match against, right? That involves questions of taste and maybe different people ask differently. So maybe it's a distributional kind of thing. And so there's clearly a big gap between those two things. So do you think there is a path for RL to be truly effective at domains without reliable rewards? So, you know, consulting, banking, legal. I mean, clearly there's tremendous progress in those domains, but like what is happening?
37:01I definitely think OpenAI will have amazing products that will be relevant in those domains. And some amount of RL will play a role in there. Does RL generalize, meaning that as you train it against more and more domains, it becomes disproportionately good at learning the next domain? I mean, we want to make a model that is generally intelligent and push that intelligence as far as possible. and to do that we want to make everything part of the distribution and then we also want to make it robust in cases where it encounters things that it was not in the distribution and if RL is part of that process I think there's a vague sense as I was trying to say earlier there's a lot of things that are very fuzzy but clearly the question of generalization in AI is an important central one and there's a bunch of examples I think that support that the processes can do this.
38:00So going back to your physics roots, a lot of what we just described about this interplay between pre-training and RIL and all the various bits that we described, those are clearly pretty complex systems. And you were trained in a discipline that is all about studying complex systems. what can physics teach us about how to understand those AI systems that we're currently building? I think there's a lot of angles to answering that question. I think the maybe most interesting one or the most relevant one to how we work currently, and maybe this is a contrarian take, is that the way to think about scaling and say scaling laws is not small to big, but big to small.
38:55And I'll get to why physics really matters for this in a second. When you have the existence of some really big AI system and some weird things happen, and they didn't happen at the small scale, and so we say, oh, this whatever emerged at scale. Sometimes people use the word grokking. There's something discontinuous about the scaling sequence, or the scaling law is broken. And these are things that people might say. But I think I reject that entirely. I think it means that you didn't understand something about what you were scaling up. Maybe even going back to the reasoning thing. I don't know if this is true.
39:31This is a cartoon. I wasn't at OpenA at the time. But if you imagine trying to get small models to reason, GPT-1, GPT-2, GPT-3, and then GPT-4 to the cartoon, you might say, oh, this emerged at scale. And it doesn't happen for the small models. you know that I reject that instead there's some phenomena that's really exciting that we discovered like reasoning or you know maybe something bad like you know like your model blew up and your earlier models didn't blow up and your job is to then figure out how to restore smoothness to the scaling sequence go back and make smaller and simpler models or or simpler toy examples such that the whole thing is smooth and if you can do that if you can figure out what to put into the small thing, then you understand the thing.
40:20And then you can move forward. This is exactly what we do in theoretical physics. There's the standard model, which is, you know, I have a textbook behind me. The description of all the forces except gravity would take, even in compact notation, like the entire page, it's like completely gross. There's, you know, a lot of different particles. Why? You know, who knows? Some of them, there's reasons for, but they're doing all sorts of different things, different things cancel, whatever. Or, you know, this just happens to be the universe that we live in. But you don't need all of that to, to like study pieces of it, to like study electromagnetism.
40:56You forget about everything else. Or if you want to study the Higgs phenomena, you know, which gives mass to some particles, you can study a simplified version of that. And, and, and so what we do, and I think one of the key moves in, at least in my training in physics is to take really complicated systems. This often gets talked about as physicists just study spherical cows. And I think that kind of misses the point. Like you study if the spherical cow is sufficient to describe the thing that you care about, then you did a good job. And if not, you did a bad job. You don't try to retreat to a setting that's simple enough where you can calculate something.
41:32You try to retreat to the setting that's simple enough that contains the thing that you care about. And then you have no idea whether you can make progress there or not. But once you did, you sort of understand what the problem is. and that's a lot of the work in physics. And the same thing is true in AI. You have these crazy, huge systems that have all sorts of interesting phenomena and if you think about it the right way, they don't grok. There's just this nice continuity. Do you think there could be an equivalent in AI to thermodynamics, meaning a compact theory that predicts behavior without tracking every individual bit?
42:08Yeah, Kaplan-McCandlish scaling OpenAI scaling laws work originally is a version of this where you throw away, you know, all you know about the network is how many parameters it is and how much data you've trained on it, and you can predict the final loss. I think the missing piece is going from all the individual weights and biases and how does that add up to the scaling law. I have some very initial work, and there's some other initial work about like trying to bridge that connection but like i think that's that's the missing piece like the sort of statistical mechanics to thermodynamics of how do we like how do these things emerge but there's definitely a lot of useful effective descriptions of how these systems behave i think the other part of your question is like is it is enough to characterize everything that we care about right there's probably a lot there's a lot that we care about other than just the final loss function and so there's there's more thermodynamics to be worked out in addition to like how does the thermodynamics arise from the microscopic description so at that conference a year ago you jokingly uh predicted nine years to einstein level ai what do you think all jokes aside we are on that on that spectrum of uh just um ai uh creating scientific discovery i mean that's where we started the conversation and curious about where this is going the joke maybe it's helpful to deconstruct a joke as as it always is but the the joke was that taking the doubling time for uh the amount of work a system can do autonomously and figuring out how long it would take us to get to a system that can think eight years on its own because einstein spent eight years discovering general relativity and i projected that out and was like nine years from last year so um that that like something i i hate making predictions but i'm pretty sure something will break before that i mean in in general we're not just going to like set up a system and let it think autonomously for eight years if anything because like the system's eight years after will be so much more powerful that it probably doesn't make sense to let a system think for a certain amount you know there's this amount of time it takes for the system to improve and then there's the amount of time it's thinking and like probably when those cross like all these scaling laws are going to to break in in certain ways i do think that the kind of thing that that i was trying to talk about about like how we as physicists approach problems that the structure and flavor of that is maybe different than here's like a very well-defined thing and go and do a calculation, which is like what these Erdos problems are.
44:47I think probably we'll need to have some ideas to bridge from one to the other. I don't think it's, it's not obvious whether it has to be a discontinuous thing or a smooth thing, but, you know, there's part of the scientific process, I think, that the models haven't been imbued with yet. And I'm sure people are thinking about how to do that. You know, like trying to get to what is the right question as opposed to here's a well defined thing and go calculate. and some of that involves research taste. That's not an easily verifiable thing. Is that why I would convince you that AI is doing genuine original science?
45:22No, I'm convinced. And I think we're going to, like this is clearly, I think the unit distance problem is a great example. And also just being able to take a position that is contrary and think for an extremely long amount of time, explore lots of different options and bring to bear the full weight of disparate fields, like where something, you know, it's very unlikely to find a human that has the exact set of skills to solve some of these problems, right? That's a huge, that's a huge thing. How far do you think we are from AI research actually automating itself? Not just AI researchers using AI, but like AI autonomously building AI.
46:02Yeah, I think it's again one of these smooth things where like it's already doing pieces of it now, it'll do more in the future. And there's, I know there's strong versions of this that people like to think about, but I'm not sure that we'll see a really sharp phase transition versus just more and more pieces of, right now, a lot of coding that would take people weeks can be done very efficiently with models. So like some of these math discovery problems, there's also versions of this where for engineering, the models are playing a more central role. And so I think there'll just be more of that.
46:34I think that there's a kind of scientific thinking that humans still seem to be very useful for doing. And I don't want to make specific predictions about when or how. I can imagine you don't want to be caught on record saying the models won't be good at something because you'll definitely be wrong. Or maybe I should say that and then the models will be good at that immediately. And so I should pick the things that I want the models to do and say that they'll never do that. I think it's also just hard to make predictions because I think the way in which people made predictions before, like the actual ways things shook out often are not in that direction.
47:08And so, you know, it's another sort of credit assignment thing. Like if you have this long chain of things that has to happen for whatever to happen, then anything that breaks that chain means your prediction, like, is just way off. And so, but in the, you know, in the, I can make a very long distance prediction, you know, for the next six months, like, I think we'll see more of of these sorts of math and science breakthroughs. And obviously we'll turn this sort of thing on AI itself and the models will get a lot more powerful and that'll be fun. You could think about, you know, that you could do science of AI and have it feel like doing physics and that's true.
47:42Another really exciting thing is that, like I entered physics thinking that I would, you know, when you first start learning a field and maybe you wanna commit to it, at least the perspective I had is that, oh, by the time I get to the end, I'll know all the answers, right? All the fundamental questions, obviously, like this is a journey and at the end of the journey, it'll resolve. And then, I don't know, maybe it was in grad school or maybe when I switched to AI, I realized, oh, like some of these questions will stay open maybe forever. Maybe I'll never get to learn the answers, you know, watching older colleagues as well, start to retire and realize that they might not get to learn the answers.
48:18But I feel really excited that, you know, we will get to really answer a lot of fundamental questions in the fields of science that we care about with the aid or maybe the models being the driving force. And so that, yeah, that's just really thrilling. Well, that feels like a wonderful place to live it. Dan, you gave us plenty to ponder. Really appreciate you spending time with us today. Thank you. Thanks for inviting me. It was a pleasure. Hi, it's Matt Turk again. Thanks for listening to this episode of the Matt Podcast. If you enjoyed it, we'd be very grateful if you would consider subscribing if you haven't already or leaving a positive review or comment on whichever platform you're watching this or listening to this episode from.
49:00This really helps us build a podcast and get great guests. Thanks and see you at the next episode.
From the publisher
Are we witnessing the first real signs of AI becoming a scientist? In this episode of The MAD Podcast, Matt Turck sits down with Dan Roberts, lead of the Foundations of Reinforcement Learning team at OpenAI, to explore one of the biggest shifts happening in AI: the rise of reasoning models, test-time compute, and reinforcement learning as engines of scientific discovery. Dan brings a rare perspective - from theoretical physics, black holes, quantum information, and deep learning theory - to explain how models are learning to “think,” why language may be such a powerful foundation for intelligence, what recent AI math breakthroughs really mean, and whether we are beginning to see AI systems that can contribute to science itself.
(00:00) Intro: AI's wild week in mathematics
(01:21) What OpenAI's Foundations of RL team does
(03:08) Dan's journey: from black holes and quantum gravity to frontier AI
(07:04) Are AI systems becoming useful for real science?
(08:21) The AI math moment: Erdős, OpenAI, DeepMind, and Anthropic
(08:52) Why the OpenAI result was an act of exploration
(10:25) OpenAI vs. DeepMind: informal reasoning vs. formal proof
(12:13) RL 101: learning by doing, not just watching
(15:10) Why reinforcement learning works
(15:58) How RL breaks: sparse feedback and long-horizon tasks
(17:03) RLHF: how human feedback shaped early language models
(18:48) Move 37, self-play, and the search for novel strategies
(22:16) Explore vs. exploit in scientific discovery
(24:49) Why RL may now be "the cake," not the cherry on top
(25:46) Why RL started working with large language models
(27:29) Is RL "sucking supervision through a straw"?
(28:47) Why language may be the grounding layer for intelligence
(31:46) A contrarian take on the Bitter Lesson
(32:41) What test-time compute actually is
(34:50) How RL gives models the ability to think
(35:40) Verifiable rewards, math, coding, and the messy real world
(38:00) What physics can teach us about AI
(42:08) Is there a thermodynamics of AI?
(43:08) From Erdős problems to Einstein-level AI
(45:16) Is AI already doing original science?
(45:51) How far are we from AI automating AI research?
(47:41) Why Dan is excited about the future of science
