Reflection AI’s Misha Laskin on the AlphaGo Moment for LLMs

16 Jul 2024 · 1 h 7 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Podcast Episode Summary: Reflection AI’s Misha Laskin on the AlphaGo Moment for LLMs

Podcast Overview Title: Training Data Description: Conversations hosted by Sequoia Capital partners discussing AI developments with leading builders and researchers. Focus on understanding the implications of evolving technologies for technology, business, and society. Episode Title: Reflection AI’s Misha Laskin on the AlphaGo Moment for LLMs Episode Description: Exploring the potential of AI agents to transform work and life, the episode highlights the journey towards developing agentic LLMs, the challenges faced, and insights from significant AI research.

---

Episode Highlights

Introduction (00:00 - 01:11)

  • Misha Laskin, CEO and co-founder of Reflection AI, is introduced as a former research scientist at DeepMind.
  • Episode discusses the progress and challenges in developing AI agents.

Misha Laskin’s Background (01:11 - 10:01)

  • Born in Russia, moved to Israel, then to the U.S.
  • Parents inspired his interest in technology and science.
  • Developed a passion for understanding complex systems through the Feynman Lectures.

Entry into AI (10:01 - 15:54)

  • Collaboration with co-founder Ioannis Antonoglou, a key contributor to AlphaGo, sparked Misha’s interest in AI.
  • Identified AlphaGo as a milestone in AI’s ability to create superhuman agents.

Reflection AI and Agents (15:54 - 25:41)

  • Reflection AI's goal is to develop universal superhuman agents.
  • Aim to combine reinforcement learning (RL) capabilities with LLMs.
  • Challenges in AI agents include depth of task execution and issues with current models.

Current State of AI Agents (25:41 - 29:17)

  • Best-in-class coding agents perform significantly better than previous baseline metrics but still face limitations.
  • High-level understanding of agents requires tackling depth and reliability.

Analysis of AlphaGo, AlphaZero, and Gemini (29:17 - 32:58)

  • Insights from AlphaGo and its successors inform current AI development.
  • Emphasis on the need for RL and learning to enhance agent capabilities.

Challenges with LLMs (32:58 - 37:53)

  • LLMs currently lack a ground truth reward and struggle with task reliability.
  • Error accumulation in task execution can lead to decreased performance.

Importance of Post-Training (37:53 - 44:12)

  • Post-training is crucial for hardening good behavior in AI systems.
  • The need for effective task verification mechanisms to ensure agent reliability is emphasized.

Task Categories for Agents (44:12 - 45:54)

  • Discusses the necessity of diverse task categories for training and evaluating agents.
  • Reflection AI focuses on creating generalizable agent frameworks.

Attracting Talent (45:54 - 50:52)

  • Reflection AI has successfully recruited top AI talent.
  • Emphasis on collaboration and learning from experienced leaders in the field.

Future of AI Agents (50:52 - 56:01)

  • Discussion on the timeline for developing capable AI agents.
  • Optimistic outlook on achieving significant advancements within a few years.

Lightning Round (56:01 onwards)

  • Quick-fire questions and answers to wrap up the episode.

---

Key Takeaways

  • Agent Definition: An agent is an AI system capable of reasoning and executing tasks autonomously.
  • Current Limitations: Existing models struggle with reliable performance and depth in task execution; current AI agents are still in the early stages of capability.
  • Importance of Depth: Achieving deeper understanding and execution of tasks is critical for the evolution of AI agents.
  • Role of Feedback: Reinforcement learning from human feedback (RLHF) is crucial for developing reliable agent behavior.
  • Exciting Future: Potential for AI agents to revolutionize work and personal productivity within the next few years.

---

Conclusion Misha Laskin's insights highlight the journey toward creating truly capable AI agents. Reflection AI's vision focuses on blending RL with LLMs to achieve a new era of digital agents that can transform work and life as we know it. The discussion underscores the importance of depth, reliability, and the continuous evolution of AI technologies.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00I think someone needs to solve the kind of depth problem. The feel as a whole, I think, or large labs have been, have been really working on the breadth. That's amazing, and there's a big market for that, and a lot of very useful things that get unlocked, but someone needs to solve the depth problem too.

0:33Hi everyone, welcome to Training Data. Today we're hosting Misha Laskin, CEO and co -founder of Reflection AI. Misha's a former research scientist at D -Mind, and his co -founder, Yannis, was the creator of AlphaGo and RLHF Lead for Gemini. Together they are building universal, superhuman agents. We'll chat about why we're still far from the promise of AI agents even with best -in -class models today. What we need to truly unlock authentic capabilities for LLMs and what we can learn from those who have built both the most powerful agents in the world like AlphaGo and Alpha0 and the most powerful LLMs in the world like Gemini.

1:11So Misha, to kick things off we'd love to learn a little bit more about your personal background. You were born in Russia, moved to Israel when you were one, then the United States in Washington State when you were nine. Your parents we're pushing forward the field of technology and research in chemistry. And I think that inspired a lot of your love for pushing forward the frontier of technology and getting into the world of AI today as well. Can you share a little bit more about what inspired you to get into this field and what has inspired you throughout your childhood and adulthood so far? Yeah, definitely.

1:48You know, when my parents, we immigrated from Russia to Israel. It was when the Soviet Union collapsed and they came to Israel basically, I mean with with nothing they had I think $300 in their pocket, which was then stolen from them as soon as they landed because they put down deposit for an apartment and Well didn't get that just a fear they I don't even know if there was an apartment so They and they didn't speak he brew So they decided to kind of pursue a PhD in chemistry at the Hebrew University of Jerusalem, but that's not because it's not because there was kind of some kind of internal passion for kind of academia at that time.

2:34It was more that Israel was giving stipends to Russian immigrants to get further educated. And it was interesting asking my parents about this because they kind of grew to love their craft as they got excellent at it. So I think from them that might be the thing that I took away most. It's not that they're particularly like impassioned about chemistry to start, but as they kind of learned about it, got curious about it, and really went deeply into it. I think they became kind of masters of their craft, and that's something that I found really important myself. Well, moving from there to the States, well, my parents promised that we're moving to this kind of beautiful state with all these mountains on this Washington.

3:23And I remember like, we're taking the plane flight. I mean, I brag to all my friends in Israel and you think, you know, I was really excited. So yeah, we're flying. I do see mountains in the distance. But then the plane does a sort of U -turn. and if maybe you'll don't know us about Washington State, but it's a kind of half desert and half, you know, mountainous and forest and the plane turned to the desert. And so I see it kind of landing in the middle of nowhere and I ask my parents like, where are the mountains? Like, well, you saw them from the plane. The reason I'm saying this is because I basically moved to a very boring place.

4:06where there's there's this area in Washington state called the tri -cities and it has some pretty interesting history. The reason it exists is because it was one of the sites of Manhattan project. So this is where the plutonium was enriched at this place called the Hanford site, which is a sister site to Los Alamos. So it's a town that was basically built for that in the 1940s and is in the middle of nowhere, kind of like Los Alamos is and there's not much to do. I remember seeing my first tumbleweeds. Like, it was literally saw tumbleweeds, kind of rolling across the highway. I found myself in a place where I didn't really speak the language that well, English.

4:48I was in this kind of very rural place. I was different from where I grew up. Didn't have many friends, and had a lot of time on my hands. And the way I got interested in science and the time it was physics is after getting sort of video games out of my system, I got bored again, and I found these Feynman lectures that my parents had on the Feynman lectures in physics. And they were so interesting because Feynman had this way of explaining incredibly complex things in a way that basically a person who's, I mean, not that educated mathematically at the time can really understand something fundamental about how the world works.

5:28And that is probably the thing that inspired me most. I just got really interested in this idea of understanding how things work at this sort of root level and working on problems that are sort of root node problems and that, I mean, there are all these examples I was reading about like the invention of the transistor which was enacted by Joe Bardeen, a theoretical physicist or how GPS works, it turns out you need to understand, you need to make relativistic calculations which is coming from Einstein's theory of special relativity. And I wanted to work on things like that. So that's why I got into physics.

6:10I pursued it for a while, got educated and it got a PhD in it. And I think maybe the critical bit of information that I did not have in my context then was that you don't just you want to work on root node problems of your time. You want to work on the things that are and be unlocked now. And it's no surprise when you're being trained as a physicist, you're doing these really interesting problems and learning very interesting things about how people thought about these about physics basically 100 years ago. 100 years ago, physics was the root node problem of our time. And that's why I decided to not pursue professionally.

6:52I kind of did a 180 and wanted to do something very practical, so I started up. I started up. But as I was working on that, I started noticing deep learning as a field taking off. And in particular, AlphaGo, like when AlphaGo came out, there was something, something just felt very profound about it. That how do they get this system that is like a computer to not just perform at a high level and a human, but do so creatively. In AlphaGo, there was a famous move called Move 37, where the neural network made this kind of move that looked like, it just looked like a bad move. LisaDoll was really perplexed by it.

7:35Everyone was perplexed by it. It just looked like a mistake. And it turned out that 10 moves later, this was actually the optimal move to kind of put AlphaGo in a winning position for the game. And so you could tell that This is not just a brute force thing. This is not a thing that is just, it's obviously, I mean, the system does a lot of search, but it's able to find creative solutions that people haven't thought of before. And so that kind of made me feel pretty viscerally that solving agency, this was the first real, large -scale, superhuman agent. Yeah, that seemed profound. That's why I got into AI.

8:14And got into AI to build agents from the first day. So there was a kind of non -linear path where, you know, I was an outsider. I wasn't really, I mean, it was competitive back then too. And OpenAI released these requests for research at the time. This is maybe 2018 or 19. Their requests for research were just things that they wanted other people to work on. I think by the time that I was looking at this list, it was actually already stale. So I don't think they really cared about those problems, but it gave me something concrete to work on and I started making progress against one of these problems I felt like I was making progress.

8:52I don't know how much progress I actually made. I was kind of peppering a few research scientists from opening eye with questions and kind of I mean effectively cold emailing them until maybe I got you know too annoying and they start you know well, I guess they responded rather I'd say graciously and I built some some relationships there and one of them introduced me to Peter Abiel who is a PI at Berkeley, one of the greatest I think researchers of our time in the field of reinforcing learning and robotics with his lab kind of does everything they have some of the most impactful generative model research as well.

9:30One of the key diffusion paper, a diffusion model papers came out of there. And honestly got lucky. He took a chance on me and brought me into his group. He really, he had no reason to. After I was on the other side and looking at applicants coming into the group, there was really no reason for him to take someone who was not vetted. So he kind of took a chance and that I think was my foot into the field. You and your co -founder, Janis, have worked on some of the most incredible projects out of deep minds and Google. Maybe can you give the folks here a taste of some of the projects that you both worked on like Gemini and AlphaGo?

10:17Maybe what were the key learnings from each and how they propelled your thinking forward to the present day? Yeah. So Janis was basically the reason I got into AI. He was one of the key engineers on AlphaGo. He was there in Seoul when they played against Lisa Dull. And he's actually, you know, before AlphaGo, he worked on this paper called DeepQ Network, his DQNs. And this was actually the first successful agent of the deep learning era. This was an agent that was able to play Atari video games. And it kind of catalyzed this whole field of deep reinforcement learning back then, which was, hey, I systems are autonomously learning to act in, I mean, mostly video game and robotics environments.

11:10But it was the first agent. It was kind of a proof point that you can learn, you know, to act in an environment in a reliable way from just raw sensory inputs coming in. This was, I think, a big unlock that was completely unclear at the time, the same way that the neural networks working on ImageNet was an unlock in 2012. And then, yeah, Yannis worked on AlphaGo and the kind of subsequent line of work. There was AlphaGo, Alpha0, there's a paper called Mu0. And I think that really showed how far you can take this idea. Like, there's really scales. Relatives of the models we have today, large language models, the like, Alphago models actually really small, and it was so smart at this one thing.

11:53I think the key lessons, at least for me, from Alphago, were kind of encapsulated in this famous essay that Rich Sutton, this Reinforced and Learning Researcher, or I guess kind of a sort of father of a lot of Reinforced and Learning Research, put forth, which is this idea of the bitter lesson. And in that essay, he basically says that you want to, if you're building systems that are based on kind of your internal heuristics, those things will likely get washed away with systems that kind of just learn on their own. And the, or rather systems that leverage compute in a scalable way. And he argued that there are two ways to leverage compute.

12:41One is by learning, so it's training. That's when we think of language models today, they're leveraging compute mostly through learning by training them on the internet. The other way is search, which is leveraging compute to unroll a bunch of plans and then pick the best one. And AlphaGo is actually both Ideas and One. I still think it's the most profound idea in AI that combining learning and search together is the optimal way to leverage compute in a scalable sense. And those things together are the things that produce a superhuman agent that go. The issue with AlphaGo was that it was only good at one thing.

13:28And I remember being in the field then, and it did feel kind of stuck, the field of the deep, the reinforcement learning because the goal everyone set out for themselves is to build general agents, superhuman general agents and where the field landed was superhuman very narrow agents. And there was no clear path to how do we make them general because they were so data inefficient that if it takes six billion steps to train on one task, then where are you going to get the data to train all these others? And that was the big unlock of the language model era. Like, one way to think about the internet is, or all the data on the internet, is a collection of many tasks.

14:14Like you have Wikipedia as a task of kind of describing some historical events, stack overflows, a task of Q &A encoding. You can think of the internet as like a massive multitask data set. And that's interesting. Yeah, and at the reason we get generality from language models is because it's basically a system that's trained on tons of tasks. Those tasks aren't particularly, let's say, directed or there's no notion of reliability or agency on the internet. So it's no surprise that the language models that come out of that aren't particularly good agents. They're obviously incredible and they do incredible things.

14:52but one of these fundamental problems in agency is that you need to think over many steps and you have some error rate over each step and that error accumulates, it's called error accumulation. And so it means if you have, you know, some percent chance that you're wrong the first step. Del compound very quickly over a few steps to the point where it's basically impossible for you to be reliable on a task that is meaningful. The key thing that I think is missing is that language models are systems at leverage learning. They're not yet systems at leverage search or planning in a scalable way. And so that I think is kind of the missing piece.

15:39Okay, now we have general agents that we have general agents that are not very competent. And so you kind of want to move up the competence and the only existence proof for that has been alpha -boh and it's been un -through search. I really love that and the encapsulation of how you just shared that. I think that sets the stage wonderfully for reflection. Can you share a little bit more about the original inspiration, this problem space that you're going after, and your long -term vision for reflection? I mean, the original inspiration came very much from working, Janis and I collaborated very closely on Gemini.

16:14Janis led the Arleachreff effort, and I led a reward model training, which was kind of a key part of Arleachreff. We were working on, and everyone is working on these language models, as you, in post -training, you align them for chat. So you align them to be good, interactive experiences with For some end user. So that was, you know, through a product like ChatGPT, or Bard, which has now been named Gemini, these language models, like the pre -trained ones, are very adaptive. And so with the right data mix, you can adapt them to be highly interactive chat bots. And I think our key insight from working on that was that there's nothing specific that was being done for chat.

17:05It was just, you know, you're just collecting data for chat. But if you collected data for anything for another capability, you'd be able to unlock that as well. Of course, it's not so simple. There are a lot of things change in the sense. I mean, one key thing is that chat is subjective. So the algorithms that you train are different than the algorithms that you train for something that has kind of an objective. Like this task was done or not. There are also sorts of issues, but the main thing was that we think the architectures and models work. A lot of the things that I thought were bottlenecks have been kind of washed away with computing scale.

17:49Like long context links is something I thought would be something that would be you need a research breakthrough. and now all the players are releasing models with extremely long context links relative to what we thought was possible even a year or two ago. The methods for training these things and aligning them and post -training them are pretty stable. And it's really, yeah, it's a data problem. And it's a data problem and a problem of how do you enable planning and search on top of these objects. and we thought we'd move faster against this problem if we did it on our own. I think we just wanted to move very quickly against it.

18:34So you've described agents as kind of the dream both for you and Janis's research was also for reflection. Can we pause on the word agents for a little bit? Because now it's become, now that it's become the term of 2024, everybody is calling themselves an agent and the way they're starting to lose it's meaning a little bit and I imagine that you have a probably a more pure definition of what an agent is. Maybe, could you just explain that? How do you think about what an agent is? And, you know, when we look at some of the agents that everyone's gotten really excited about recently, there's still, it seems like they're still very early in terms of being reliable enough to be kind of colleague -level true agents.

19:12So like, where do you think we are on that curve? What is an agent? And how do we kind of get to the promised land? Yeah, it's an interesting question since I... The term agents has been floating around, you know, within, more of a research community for a while. I mean, I think, well, since kind of the start of AI, but I've barely been thinking more ages in the context of the deep learning era, starting with DQNs. And the definition is pretty simple in that it's an AI system that is able to reason on its own and take, well, however many steps it needs to, to accomplish some goal that's been specified to it.

20:02That's kind of it. And now the way that goal is specified has changed over time in the deep reinforcement and learning era, the goal was usually specified through a reward function. So for AlphaGo, it's whether you won the game of Go or not. There's no one wrote, like, Go, win the game of Go, like the text. So that's how people usually thought about agents. They thought of agents as things that are optimizing a reward function. But there's a whole area of research, even then, before language models on goal -conditioned agents. So these would be either in robotics or video games or you set a goal for the robot, which could be you give an image of an apple having been moved somewhere and you ask it to kind of reproduce that image and it has to act on the world and pick up an apple and move it somewhere in order to accomplish the goal.

20:55So there's a short definition. It's AI systems that have to act in an environment to accomplish some goal. That's an agent. And then I guess as a follow up, if you take, for example, coding agents as one potential domain of agents where there's been a lot of activity recently, you could say the goal is create a calculator app for me. And the agent has to be when I accomplish the task. When I look at what sweet agents and what Devon have done, is that in your mind,

21:31is that us to the promised land or do you think there are kind of different approaches that you need more on the RL side or more on whatever other techniques you might need to use to go to the promised land? Because I think those agents are still in the 13, 14 percent. Have a task completion rate range? I'm curious how we get them to 99 percent. They're definitely by the definition of agents. These are agents. They're just on a spectrum of capability, maybe not at a high level of reliability yet. But I think the way most people think about agents today in context of language models is prompted agents.

22:08So you take a model, you prompt it, or you set up some flow of several prompts to get it to accomplish a task. That allows anyone to kind of take a language model and take it from zero to something that's kind of working somewhere. So I think that's quite interesting. I think it can only go so far. So this, this I think is actually like a very, I mean, this is a kind of example of what I think the bitter lesson would apply to because prompting things and kind of really directing them to go in like these specific ways. That's exactly the kinds of Heuristics that we have that we're baking into these models to kind of try to achieve higher intelligence.

22:47I mean every major advance in agency since the deep learning era has been kind of showing that with learning and search, a lot of that gets washed away. I think that the purpose of a prompt is to specify the goal. So you'll always need a prompt, you always need to tell an agent what to do. But if once you start deviating from that and the purpose of a prompt is actually to put the agent on Rails and where you're kind of doing the thinking for it, you're kind of telling it, okay, now just go here and do this thing. That I think is going to disappear. I think that's a local thing that's happening today.

23:32Future systems, I don't think we'll have that. So the keys of the thinking and the planning these to happen in the AI system, not in the prompt layer in order to kind of not hit a wall? I think you want to offload as much as possible to the AI system itself. Again, these language models have not been trained for agency. They're trained for chat interaction and predicting things on the internet. So it's almost a miracle that you can prompt your way to getting something that kind of works. But what's interesting is that once you're able to prompt your way to something that kind of works, that's actually the best place to start for a reinforcement learning algorithm.

24:12A reinforcement learning algorithm, all it does is it reinforces good behavior and it kind of down ways, bad behavior. If you have an agent that is doing zero, that is just doing nothing, then there's no good behavior to upway, and so the algorithm doesn't work. This is known as a sparse reward problem. If you're not hitting your reward, like if you're not accomplishing your task, then ever, then there's nothing to learn from. But if you've prompted your way to an agent that is kind of working, like, you know, sweet agent or something like this, it's getting 13 % or something like this, then you have something that is like minimally capable where you can reinforce really good behavior.

24:50The challenge becomes a data challenge of where do you get the set of prompts to train on? Where do you get the environment to run these things through? Well, I guess, Sweden does come with an environment, but for many problems, like you need to think about that. And then perhaps the biggest challenge is how do you verify that a thing has been done correctly or not in a scalable way? And if you can solve, you see, where the tasks come from, which usually that's your products, that's solvable, what environment you run them through, what algorithm you use, but it's really kind of what environment you run them through.

25:31And then critically, how do you verify if a thing has been done correctly or not in a I think that's a recipe for agency. I think that gets to the crux of the problem space in the AI agents today. Just to set the stage a little bit for the problem that reflection is going after, what do you think is the current state of the market broadly in AI agents? I think many assume that we are capable of more than we actually are with the models that exist today. So, what do you think the problem is? And what do you think is, why do you think the current attempt around AI agents are feeling us today? One way to categorize or classify what it means to be a general agent.

26:16And maybe I'll use the term universal agent since I'll use the term generality to apply to breadth. So a universal agent needs to be abroad, a very general agent that can do many things, can handle many inputs. But it also needs to have depth in the kind of task complexity it can achieve. And so examples are, AlphaGo is probably the deepest agent that we've that has ever been built. It can do one task. So not that useful. It can play Go, but not TickTackTo. The current systems, language model systems like Gemini, Claude, Chad Gpt, the GPD series of models. Leading the other way, they're very broad.

27:02They're not very capable and depth wise sense. They're extremely impressive and capable broadly. And I think that's one of the things that's been honestly miraculous. Like, as I said, the field felt like we did not have an answer to generality. And then these objects came along. But now we're in the opposite end of the spectrum. where we have, I think, more or less de -risked as a field progress towards breadth. That's especially evident with the latest generations of models like GPT -40 and the latest family of Gemini models, then, are multimodal in the sense that they just understand other modalities at the same base layer that they understand language.

27:50You don't need to translate one modality into language. So that's I'd call it breadth. But nowhere along this process where things train for depth, there's no. The internet doesn't have real data around how to kind of think sequentially. The way people try to solve this problem is like work on data sets that might have this structure and hope it generalizes. So math, data sets, coding, data sets, kind of what people refer to reasoning, which usually is reasoning along lines of can you solve a mathematical problem. But that's still not really addressing the problem head on. I think we need methods that you can take at recipes to say.

Read the full transcript

28:36They're general in that you can take any task category, have a bunch of prompts for it for your training data and make a language model kind of iteratively capable, more capable on those things. I think someone needs to solve the kind of depth problem and the field as a whole, I think, or large labs have been really working on the breadth. That's amazing. and there's a big market for that and a lot of very useful things that get unlocked, but so one needs to solve the death problem too. I think that takes us really nicely into the unique insight that you and you on us have from working on AlphaGo, Alpha0, and on Gemini and the importance of post -training and data.

29:29Can you share a little bit more about how those experiences of shape the unique perspective that you have that gets us to the unlock with the gentic capabilities? One of the things I found very surprising about language models is how close they are to, you know, how often times, even if they're not working on something that you want them to, they're actually quite close. They feel like a nudge away. They feel like they need to be grounded in the thing a bit better. And that's, I think that was the insight that led them to be good and chat. like you could play with them and they're yeah a bit unreliable and they kind of go off the rail sometimes but they're almost Good chat companions.

30:09Yeah, and so Then the recipe there. There's a recipe for how do you take a pre -trained language model and make it a reliable chatbot? So by reliability there. It's just a The way you measure that is with human preferences do People interacting with this chatbot prefer over other chatbots or other versions of it the previous versions of itself so If the current version is much more preferred than the last few iterations ago, then you know you made progress. And that progress is made by collecting data for it. So it's collecting data for the kind of queries that users input into a chat box, the outputs that the models provide, and an effective ranking between those outputs so that you push the model to index over on the more preferred outputs.

31:04So when we say ranking, whereas that ranking comes from humans. So there are either human lablers or it's something that's embedded into the product. You sometimes might see thumbs up or thumbs down in chat GPT. It's harvesting your thumbs to know what your preferences. And that data is used to kind of align the model with the user preferences. That's a very general algorithm. That's a reinforcement learning algorithm. And that's why it's called reinforcement learning from human feedback or RLA chef. You're just up -playing the things that human feedback is pressing preferences for. There's no reason why the same approach would not be possible for enabling more reliable agency.

31:49There's a whole sequence of other problems you need to solve. I think the reason this is so hard is because as soon as you go into kind of agent territory, you have more than just the language outputs. You have the tools that they interact with. And the tools being supposed to send an email or work at an IDE or anything that an agent does, it does an environment and that requires tools. That requires the environment. And everyone who's deploying agents is deploying agents in different environments. And so there's a challenge of how do you integrate with the environment and how do you onboard agency onto them.

32:27So I think that that's why it's a bit of a schlep if you get into this kind of line of work and you have to be careful about the environments and kind of the way you structure because you don't want to overfit to some particular environment. But conceptually, it looks very similar to aligning a model for chat. There are just some more integration challenges that need to be solved along the way. Since you view AlphaGo as the pinnacle of building an agent that was truly capable, I imagine you're trying to usher in an AlphaGo moment with LLAMS. What do you think are the differences? Like to me, you know, with gameplay of a very clear reward function, you have the ability to self -play.

33:16Like is doing kind of the reinforcement running from human feedback? Do you think that's enough to kind of get us to an alpha -go moment in LLM's or like, well, I guess, how should I think of the differences here? I think what you said around not having a ground truth reward is, hey, key and maybe the key thing. But we learned from the previous era of reinforcement learning research that if you have a ground truth reward, your kind of guaranteed success, like that's kind of there've been so many very impressive projects that showed this at really unprecedented scale. I mean, you think aside from AlphaGo, there was opening eyes, Dota 5 or AlphaStar and let's say AlphaStar and Dota 5 are a bit more niche in the sense that you kind of have to play those games.

34:04to understand, but as a former Starcraft player, I was, I still am completely blown away by Alpha Star. Like, the strategies that the AI discovered were, it just looked like a very smart, like a smarter than us alien game, like upon the Earth decided to play this game and completely outcompeted us. So that's due to the existence of a number of things but a ground truth reward is really like is extremely important for tightening that behavior.

34:39Now, both with human preferences and for agency, we don't have ground, we, these are very general objects and we don't have ground truth rewards for whether something is accomplished or not. Like, even for a coding task, what's the ground truth of whether this was done the right way? Like it could pass some unit tests, but it can still be wrong. So it's a really hard problem. And I think it's the fundamental problem for agency. There are others as well, but this is kind of the big one. The way you get around this problem for chat is, again, through Arleigh Chath. Well, you train reward models.

35:18models. Her word models are, it's a language model that predicts whether something was done or not correctly. The challenge with that, so first it works well. The challenge with that is that when you don't, in the absence of a ground truth, when you have this kind of noisy thing, that it can be wrong, your policy or the agent quickly gets smart enough for it finds holes in their reward model and exploits them. To give a concrete example in chat, suppose you noticed that your chat bot was outputting was outputting some, let's say, harmful content or like there are some topics that you don't want it to talk about because they might be sensitive.

36:05And so you put in some data into your data mix saying like where it's examples of the chat bot kind of ignoring, or not ignoring but saying, I'm sorry, is a language model I cannot answer this. What can happen is that, okay, you now train a reward model against this, and suppose in your data mix, you really, you only put in data points that showed instances of this kind of, this happening, but not instances of a chatbot taking something like kind of sensitive and actually answering it. What that means is that the reward model could think that it's actually like a good thing when you just don't answer the user's query ever because I've only seen positive use cases of that.

36:48And when you train against that, the policy with language model will at some point get smart enough and discover that this reward model gives me high reward whenever I just don't answer whenever I punt it the question and it can collapse into a language model is just never answers your questions. And this is why it's very finicky and it's very difficult for this reason. I'm sure a lot of users who have interact with chat GPT or or Gemini or these kinds of models like probably through through interacting with them found sometimes that they kind of degrade and they all of a sudden like don't answer questions you know as often as they used to get slightly worse at something or are politically biased in some way.

37:38I think a lot of that is, well, it's artifacts of the data, but the artifacts in the data get amplified by bad reward functions. That is the hardest problem, I think. If I view the rough large model training pipeline or large AI system training pipeline as pre -training and post -training. I kind of think pre -training seems largely like a solved, like we're in the, you know, the techniques are solved and we're just kind of in the race to scale. Moments on the pre -training side, post -training still feels a little bit like in the kind of research phase of market where people are still figuring out what techniques will work in a general way.

38:21I'm curious to do you agree with that and in an ideal state, like what is pre -training responsible for doing, how should we, as laymen, think about it? And what is post training responsible for accomplishing and how should we think about that from the perspective of five -year -old? Yeah.

38:41I would generally agree with that statement that pre -training has become, there are a lot of details I need to get right and it's, by no means easy, so it's a very hard endeavor, but it's a better understood endeavor at this point. And one way that I think about pre -training is I actually think thinking about it through the lens of something like AlphaGo is quite simple and clear because it kind of, you know, rather than thinking about this massive internet thing, you just think about a very clear setting, which is clean setting, which is this game. You can think about the pre -AlphaGo as two phases.

39:24It has an imitation learning phase where a neural network imitates a bunch of expert amateur, like expert go players. And then it has a free -inforcing learning phase. And you can think about pre -training as the imitation learning phase of AlphaGo. You're just kind of acquiring the basic skill of learning to play the game. You're not maybe your neural network then is not the best in the world, but it's pretty good. it goes from zero to pretty good. And pre -training for a language model is going from zero to pretty good on everything, which is why it's so powerful. Post -training is, I think about it as hardening good behavior.

40:07What that means is, with AlphaGo, you start off with, you did imitation learning. You start off at a place where you have a neural network that can do something. It can, I mean, it can play a game pretty well. Then you apply this other recipe to it, which is reinforcement learning, which is then the network starts generating, you know, its own plans and kind of acting through the game, getting feedback, and that good and basically good actions get reinforced. That is, I would say that's post -training, and you can think about, from a chat perspective, you're hardening the model, like, the good behavior very along the chat axis.

40:46So it's actually quite interesting that the high level recipe for training alpha go and for training Gemini is actually the same. You have the simulation learning phase, and then you have a reinforcement learning phase. The reinforcement learning phase and alpha go is just much more sophisticated than what we have now. And the reason comes back to reward models. If you have a reward model that is fairly noisy and exploitable, and exploitable, then there's only so much work, there's only so much you can do before the policy gets smart and finds a way to trick it. And so even if you through the fanciest like RL algorithm at it, like Monte Carlo tree search with AlphaGo, it may not be that effective because it kind of, it collapses into this kind of degenerate state where the policies have their reward model before it could even do any interesting search.

41:43Like, suppose you're thinking about, like, if you're playing chess and you're thinking about what to do multiple moves ahead, but your kind of judgment is really bad at every move, then there's no point of planning 10 moves ahead. And I think that's where we are with Arleigh Chaff today. There's this wonderful paper that I think is very over or underrated, called Scaling laws for reward model over optimization. This is a paper from OpenAI studying this phenomenon. And what's interesting about, I mean, a number of things, but it's showed that this phenomenon happens at all scales and I mean, to try a couple of different RLHF algorithms and it happens for, at all scales for all algorithms that were tried in that paper.

42:31And I think that's interesting paper because it's the kind of fundamental problem of post -training. That's that paper. Yeah, just to put Paul in the thread a little bit. If you follow the results from alpha zero though, then we may not need pre -training at all. Is that a fair conclusion of what to make of this? I think that the at least my more from a practicality standpoint. When we went from or deep mind went from AlphaGo to AlphaStar, there was no Alpha0 of AlphaStar. There's no AlphaStar 0 that was, you know, released after that or anything like this. And AlphaStar had like a big part of it was imitation learning across a lot of games.

43:30I think with AlphaGo, it was like this kind of special place where you have, you don't only have a zero sum game, but you can get to the end of that game fairly quickly. And so it's again, you can get that feedback about whether what you did was right or not. Got it. Okay, so it's just way too unconstrained that we're problem to throw that out generally. Yeah, I think in practice, alpha zero would work generally for everything if we had Ground Truth or Word Functions or everything, But because we don't, you need to do the imitation running piece. This is just like a practical, we need to get it to the game somehow.

44:10You described earlier the importance from a technical perspective of having an agent in its environment. Also from a product distribution and getting the product in user's hands perspectives, it's important to think about what the right task categories are for users to first interact with the most powerful agents. What are some of the task categories that are on your mind? And what do you imagine are some of the possibilities that users could use these and their daily workflow? If you want to make progress along the depth axis, you could go for like, AlphaGo first, which is like a really hard thing, or you could kind of expand concentrically in the sort of complexity that have the task you're able to handle.

44:51And we are focused on kind of enabling depth, but in this sort of concentric way. And we care a lot about having a general recipe that is not, that does not kind of, you know, inherit heuristics that are special to some tasks. So from a research perspective, we're building general recipes for this. Now you have to ground those recipes in something to show progress. And at least for us, it's important to show diversity of environments. And so we're thinking about a number of different types of agents, web agents, coding agents, I have OS, computer agents. The important thing for us is to just show that you can have a general recipe for enabling agency.

45:51Switching gears a little bit, you've attracted a seller team already. Who else are you looking to recruit on your team? Yeah, we've been fortunate to be able to draw some talent from the top AI labs in the industry.

46:12And I think a lot of that has to do with both of the work that he honest and I did, but definitely, you know, I think a lot of credit goes to Janice and his reputation. You know that there is this, I was watching the Michael Jordan documentary and Michael Jordan was, well, one of the reasons he was so effective is because he was such an incredible kind of individual, like, basically contributor to the game, maybe be best, that he really inspired people on his team to get to his level, even if they couldn't get there. And he honest has this effect on people, like I worked very closely with him on Gemini, and he had that effect on me.

47:00I don't know if I ever got to the honest level, but I aspired to, and I definitely became a much better engineer and researcher through the process. And I think that's a lot of the draw, is that you get to learn a lot from him. We're primarily continuing to look for, so we're not hiring out quickly. We're hiring out, I think, more methodically. We're looking for, yeah, other, definitely interested in other researchers and engineers joining us on this mission. I'd say a commonality between everyone who's joined is we're all very hungry. Maybe that's how I put it. We could have, the honest and I could have stayed and tried pushing agents at DeepFind.

47:57And as I said, I think the reason we decided to do it in our own ways, because we think we can move quickly much faster against the school. And some of this urgency is driven by a real belief that we are three or so years away from something that resembles a digital age I. And by that, that's what I'd been referring to as Universal Agent. Something that has both this kind of breadth and depth of knowledge. And that means we're actually on a very accelerated timeline. Yeah. You're, you know, a few months in, you're kind of 5 % away from hitting that timeline. And maybe so this is also driven by how quickly AlphaGo went from Experts into field doubting this is possible.

48:56Yeah, it's kind of decades Human level or like expert human level go play was decades away and How effectively they were able to solve that problem within months? I Think we're seeing a similar kind of acceleration happening with language models. There's I think one viewpoint point you can have is that we've saturated a lot of what we can. We're on this sort of at the tail end of an escurve and we don't we don't view it that way. We think we're still we're still in an exponential. Part of the reason is that these things are so bulky and slow to train that there's no way collectively as a you know as a field of researchers and engineers that we've optimized it, yet, like, if it takes a few months to run, and a few months and a few billion dollars to run the biggest model, then how many experiments can you really run?

49:55So yeah, we kind of see things going at an accelerated pace, and we think solving the depth and reliability problem is something that is not getting the kind of attention it needs. Like there are groups that are following this as I would call it more like side quests within these big companies, but I think you need a player that is focused entirely on it to who solve this problem. Yeah. I love the framing of main quest versus side quest and I love the hunger and the zero complacency and impatience and a healthy way that you and the rest of the team have. And the only other thing I'd highlight is the revered reputation that you described for your honest with inspiring and motivating other people.

50:45I think is true for you and your honest from everyone we've known. We know it. So three years until I have an agent that will write with memos for me. Hopefully. I think the memos might be coming sooner. Yeah, that was one of my burning questions. Is this like decades away? Is this month's away? It sounds like you're closer to the, you know, months or small number of years away. Think small number of years. Wow. Yeah, it, yeah, it's honestly kind of alarming, I think the speed at which the field is moving. And part of, yeah, part of depth and reliability, it's also like, it is, I mean, reliability is safety.

51:30So you want these systems to be safe. I think that there's a lot of very interesting research in terms of... There's a recent paper from Anthropic on kind of mechanistic interpretability and that whole line of work is really interesting and I think starting to kind of get to the point where there's utility in it as well in terms of fighting like neurons in a model that are like, you know, lying neurons or, you know, that you can kind of suppress. But to me, safety is reliability. If the thing is kind of running around your computer, breaking all sorts of things, that's an unsafe system. Maybe it's like a utilitarian safety.

52:14Like you just want these things to work and do what you intended them, but what you ask them to do. So I have a few years to find another hobby other than my memo writing then. Yeah, well, or maybe you'll just have an army of AI interns that, you know, will do all the research work for you. Get wait. Rapping out our topic around reflection, if everything goes right, what is your dream for reflection? I think there are two angles of this question. One is we're working on this because this is the kind of scientific root node problem of our time. We're scientists. That's why we're so kind of interested in committing it.

53:00And it's really there's a world where you get to be part of one of the most exciting journeys in science ever made. And you've accomplished your goal of building universal agent. You have highly safe, reliable digital agents running around on your computer. Basically doing things that tedious work that you don't necessarily want to do. I think rather than people kind of going and spending less time working, I don't think the human need or like, and the human need to be productive and to contribute is going to change. I just think the capacity of each human's ability to produce and impact the world is going to dramatically increase.

53:51As in my line of work there is researcher, there are so many things that I spend time on that a smarter AI could help me out with to make faster progress towards our own goal. I mean, this is kind of a circular, but if we had something close to a digital AGI, we'd get much faster to solving a problem with digital AGI. That's one angle. I think the other angle is from, I guess we've kind of moved on to the other angle. It's from the user perspective of a lot of the things that we do on a computer are, Maybe you can think about computers like the first digital tool that we've been introduced to as people.

54:38In the same way that there were hammers and chisels and sickles that people used. I think we're moving towards the kind of layer beyond that where instead of you having to learn how to use all these tools with great precision and spend all your time on this, which actually is kind of time taken away from achieving whatever personal goals people have, that you kind of have these incredibly helpful AI agents that can help you bring kind of any goal that you have to fruition. And I think it's very exciting because I think the kind of ambition of our individual goals is going is it's already increasing, you know, in this local sense, like the software engineer can get a lot more done today with these tools, but this is just the beginning.

55:36And I think we'll, yeah, we'll be able to really We set dramatically more ambitious goals for ourselves and for kind of these sorts of things we want to achieve simply because we can offload a lot of the work that's needed to get there to these systems. So these are some things I'm really excited about. We'll close it out with a few questions that we like to ask everybody about this state of AI. First question, what are you most excited about in your field or more broadly in the next one year or five years and ten years? I think there are a number of things that the one, the local one that comes to mind is because the paper is fresh.

56:25Is this kind of work on mechanistic interpretability? And that I mean these models are largely black boxes and It's unclear It's really unclear how to study them as like what what's the neuroscience of language models like if you think about them as brains Yeah, and this seems like a really interesting line of work that Is now starting to see kind of signs of both science of it working beyond toy settings. So maybe like the sort of neuroscience of language models is I think kind of a really interesting. Awesome. Feel an AI to get into it. And more generally, if I was in academia, I'd probably be looking a lot at the science of AI.

57:14So this is the neuroscience of AI is one thing, but they're all sorts of things. all sorts of things one can investigate in terms of, well, what really determines the scaling laws that these models have, both from a theoretical perspective and from an empirical perspective how you change data mixes, maybe I've taken a step back.

57:40We're basically in the equivalent of what the eight late 1800s looked like for physics. This electricity was being discovered. No one knew why it worked or how. There was a lot of empirical results, but there was no theory behind it, which just meant that they were not very well understood. Then this very rich set of theoretical models were developed that were very simple to understand these phenomena and that gave rise to actually basically the next wave of empirical breakthroughs. And so I think the science of AIs kind of in that state right now, and I'm very excited to see where that goes. So interesting.

58:32Who do you admire most in the world of A .I.? I think most people, when getting this question, or some people might kind of put someone, yeah, maybe I'll take that back and say like, I want to emphasize people that I admire based on having worked with them and kind of see them operate because through my last kind I have a number of years in AI. I think they're a handful of people like this who've inspired me. And one of them is, so Peter Abiel is certainly one of them. He is, I think, I've never seen anyone, I think, almost operate as efficiently as Peter. That was something that, and to date, like since meeting him and to date, like I think there's sort of, you think a lot about research as a creative pursuit oftentimes, and I think what I learned from Peter is just sort of operational competence and efficiency around this.

59:50He's very creative as well, and like his lab does a lot of creative work, but I think it's brought, like these things are very hard and they need to be pushed hard and with great focus. And he ran his lab against a Titus trip that I've ever been on. And really helped focus all the projects. And so, yeah, I think I look up to him a lot both in terms of his. The work that he's done, obviously, it's remarkably cross -field. like he's both done well incredible break to work and reinforce some learning and on supervised learning, generative modeling and and a lot of it has been I think from kind of recognizing and enabling like talent.

1:00:41So there is this very much. It was like the group was it was a bunch of independent thinkers, students, PhDs, And people are kind of pursuing, like, what was interesting to them. But the way I saw Peter as a sort of like great amplifier. Yeah. Like, he kind of helped people amplify and focus on the thing that really mattered within their pursuit. I think a few other people that come to mind. So my manager deep -mind, his name is Vladimir. He's, yeah, I think also like a very an incredible, very creative scientist. He was first author of the Deep QN paper and then there's actually two papers at the time.

1:01:30There's like this H2C, A3C papers. These are basically the two algorithms that defined reinforcement learning and he, of a deep reinforcement learning and he kind of really pioneered both. Yeah, I think his strength is like he was Extremely kind and people oriented very humble despite his accomplishments Yannis as well. I mean definitely. I mean Yannis has them the Michael Jordan effect. I think he really Yeah, he just You just wanted to be the best you could be when you work with him and And our early chef team was quite small, and people pushed really hard in order to, like, largely I think inspired by him.

1:02:17Yeah, these are some people I really look up to. Thank you for sharing that. It's so interesting to hear you say about it everyone, and one comment on Peter Beale. I tell him all the time that he's also just creating a mafia of fenders in the last couple of years, and it's probably because he's taught them how to do many things and there's a self -selection that's naturally happening. The creative thinkers and the independent thinkers who come into Islam, but he's also taught them a lot about how to run a tight ship and how to focus incredibly well. So I'm sure that doesn't come with that intention on his part.

1:02:53Maybe last question, any advice you have for founders building an AI, you are, you're just starting your journey right now and I'm sure you've asked others for a lot of advice, what advice would you pass on to the next generation? I think one, I think I'll be in a better position to answer that question in a few years in a way that is much more meaningful. But I'll actually provide a piece of advice that I led through my previous startup, which has nothing to do with AI, to just work on things that are like internally, like really matter to you. In a way that is almost independent from what's happening around you.

1:03:33Like, in a way that when things go bad, like it's still interesting to you, like it's kind of, there's just some fundamental drive around this problem that independent of everything else that's happening is just really interesting to you. Maybe I say that about our eyes because there's this, it's such an interesting, highly capable, cool technology and so there's this sort of I think appeal of taking it and just kind of like, always just see what we can do. I think you never really find yourself in a hard place without having a very strong internal compass of independently of AI, that what it is it's important to you and what you want to do.

1:04:15And so having been in that position and previously that's kind of what I would have done differently and what I would advise to do. I really love that. The line that I like to think about is playing your own stadium and don't get distracted by the glitz and glamour of someone else's stadium. You need that internal drive and grit and obsession with the problem to get you through all of tough times. Yeah. And I think there are things that come with it. with like, if you really care about some problem, like you will care about the customers you're solving the problem for. Like having customers that you don't care about is like a terrible place to be.

1:04:57So I think yeah, it has to come. And it's not like, I think it's kind of actually hard to control like who you cared, who you don't care about. That's like a personal thing. So you can't, it's actually really hard to force yourself to like care about something out of necessity if it's kind of not aligned with something inside you already. So no more shopping and retailers for Misha? Yeah, so I was a building software for inventory prediction for retailers. And for some people, they'd really care about that problem. There's a reason they would. They've seen that problem. Maybe if you're a merchant at one of these retail companies, You've really felt it viscerally.

1:05:41And in my case, you know, I hadn't, who was like we were trying to make, you know, just trying to build a revenue -generating business almost kind of independent of like an internal compass. Be sure, thank you so much for joining us today. You are working on the most ambitious problem of our time. I love your framing of the root node problem of our time. And I think that is today agents. And it's very clear that both you and Yonis' experiences make you the very best team at what you do. Yonis, obviously, from a RLHF perspective and yours, from a reward model training perspective. And the insights and experiences that you've both had working on AlphaGo, AlphaZero, and Gemini were so excited for the future of reflection.

1:06:32Yeah, thank you for having me.

From the publisher

LLMs are democratizing digital intelligence, but we’re all waiting for AI agents to take this to the next level by planning tasks and executing actions to actually transform the way we work and live our lives. 

Yet despite incredible hype around AI agents, we’re still far from that “tipping point” with best in class models today. As one measure: coding agents are now scoring in the high-teens % on the SWE-bench benchmark for resolving GitHub issues, which far exceeds the previous unassisted baseline of 2% and the assisted baseline of 5%, but we’ve still got a long way to go.

Why is that? What do we need to truly unlock agentic capability for LLMs? What can we learn from researchers who have built both the most powerful agents in the world, like AlphaGo, and the most powerful LLMs in the world? 

To find out, we’re talking to Misha Laskin, former research scientist at DeepMind. Misha is embarking on his vision to build the best agent models by bringing the search capabilities of RL together with LLMs at his new company, Reflection AI. He and his cofounder Ioannis Antonoglou, co-creator of AlphaGo and AlphaZero and RLHF lead for Gemini, are leveraging their unique insights to train the most reliable models for developers building agentic workflows.

Hosted by: Stephanie Zhan and Sonya Huang, Sequoia Capital 

00:00 Introduction
01:11 Leaving Russia, discovering science
10:01 Getting into AI with Ioannis Antonoglou
15:54 Reflection AI and agents
25:41 The current state of Ai agents
29:17 AlphaGo, AlphaZero and Gemini
32:58 LLMs don’t have a ground truth reward
37:53 The importance of post-training
44:12 Task categories for agents
45:54 Attracting talent
50:52 How far away are capable agents?
56:01 Lightning round

Mentioned: 

The Feynman Lectures on Physics: The classic text that got Misha interested in science.

Mastering the game of Go with deep neural networks and tree search: The original 2016 AlphaGo paper.

Mastering the game of Go without human knowledge: 2017 AlphaGo Zero paper

Scaling Laws for Reward Model Overoptimization: OpenAI paper on how reward models can be gamed at all scales for all algorithms.

Mapping the Mind of a Large Language Model: Article about Anthropic mechanistic interpretability paper that identifies how millions of concepts are represented inside Claude Sonnet

Pieter Abeel: Berkeley professor and founder of Covariant who Misha studied with

A2C and A3C: Advantage Actor Critic and Asynchronous Advantage Actor Critic, the two algorithms developed by Misha’s manager at DeepMind, Volodymyr Mnih, that defined reinforcement learning and deep reinforcement learning

More from Training Data

All 110 episodes
Reflection AI’s Misha Laskin on the AlphaGo Moment for LLMsTraining Data · 1 h 7 min
Listen in VO