In short
Training Data Podcast Episode Notes
Episode
The Breakthroughs Needed for AGI Have Already Been Made: OpenAI Former Research Head Bob McGrew
Hosts
- Stephanie Zhan
- Sonya Huang
Guest
- Bob McGrew - Former Head of Research at OpenAI
Episode Summary In this episode, Bob McGrew reflects on his experiences at OpenAI, discussing the advancements in AI and the three critical components needed for Artificial General Intelligence (AGI): Transformers, scaled pre-training, and reasoning. He emphasizes that the foundation for the future of AI is already established, predicting significant developments in reasoning by 2025. Bob also addresses the implications of AI on various industries and how emerging technologies will transform the economy, particularly in sectors like law and medicine.
---
Key Topics Discussed
- The Three Legs of AGI
- Transformers: The model architecture that has revolutionized natural language processing.
- Scaled Pre-training: The necessity of massive computational resources and data to train effective models.
- Reasoning: The new focus area where Bob believes significant improvements will occur by 2025.
- Current State of AI and Future Predictions
- Bob argues that while pre-training is still important, it is reaching diminishing returns.
- He predicts that reasoning will be a major focus in 2025, leading to more complex AI capabilities.
- With advancements in reasoning, he foresees AI becoming more integrated into industries, impacting the pricing of services due to an oversupply of computationally driven solutions.
- The Agent Economy
- Bob discusses the concept of an "agent economy," where AI agents can perform tasks at scale, leading to competitive pricing based on compute costs.
- Traditional roles, such as lawyers and medical professionals, will face disruption due to this new supply of AI-driven services, potentially making expert services more accessible.
- Opportunities for Startups
- Bob identifies areas where startups can thrive, particularly in robotics, which he believes is on the verge of commercialization.
- He encourages innovation in enterprise-level applications and emphasizes the importance of understanding specific business contexts to leverage AI effectively.
- Personal Insights and Experiences
- Bob shares anecdotes about how his children engage with AI, specifically ChatGPT, highlighting the role of AI in fostering curiosity and agency.
- He discusses the importance of teaching the next generation to leverage technology while also developing foundational skills.
- Management Lessons from OpenAI
- Bob reflects on his experience managing talented individuals at OpenAI, emphasizing the importance of empathy, loyalty, and the need for collaborative efforts in high-pressure environments.
- He discusses the challenges of navigating personal ambitions within a collaborative research context and how to align team objectives.
- Security in an AI-Driven World
- The episode touches upon the security implications of AI advancements, noting that as offensive capabilities increase, so must defensive strategies.
- Bob highlights opportunities for startups to innovate in cybersecurity, particularly with the application of AI in automating threat detection and response.
---
Key Takeaways
- The foundational elements required for AGI are already established, and the next big leap will center around reasoning capabilities.
- The rise of AI agents will transform traditional service sectors, leading to a more competitive and accessible economy.
- Startups should focus on leveraging technology to create value in specific business contexts rather than directly competing with established AI models.
- Management in high-stakes environments necessitates a focus on supporting individual team members while fostering collaboration.
- Security measures must evolve alongside AI capabilities to address increasing threats effectively.
---
Relevant Mentions
- Robotics Research: OpenAI's work on robots solving complex tasks.
- AI Companies: Mention of startups like Distyl and Skild that are exploring innovative applications of AI.
- Personal Experiences: Bob’s insights on using ChatGPT for learning and exploration in real-world scenarios with his children.
---
Conclusion In this episode, Bob McGrew provides a comprehensive overview of the current landscape of AI, its future trajectory, and the transformative effects it will have on various industries. His insights into the role of reasoning and the agent economy paint a picture of an ever-evolving technological landscape that startups and innovators can capitalize on.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:00I think what's really changed is that now that you have LLMs, you have this language interface to the robot so that now you can describe the tasks much more cheaply and you have really strong vision encoders that are tied into that intelligence. So that gives the robots really a headshot at doing generic tasks. So we spent years solving one specific problem teaching a robot to manipulate a Rubik's cube. And now a company like let's say physical intelligence can spend months solving a huge variety of problems like laundry folding and cardboard and packing egg crates. And that's something that they can only have because they're building on top of existing frontier models and you know the entire tech and research stuff that we've built over the last 10 years.
1:05Welcome to Training Data. Today we're excited to welcome Bob McGrou, former Chief Research Officer at OpenAI for a fascinating look behind the scenes of Frontier AI Development. Bob talks about the trifecta we have in AI, pre -training, post -training, and reasoning, and explains why we may have already discovered all the fundamental concepts needed for AGI. You'll learn why he thinks agents will be priced at the cost of compute, hence eroding traditional economic modes, and why even proprietary data will become less valuable when infinitely patient AI can recreate alternatives. Plus, Bob shares his contrary intake on where startup opportunities really lie, and why robotics is finally having its moment after years of being too early.
1:50Enjoy the show. Bob, thank you so much for joining us today. Oh, it's great to be here. We're at a really interesting time in AI development. We have a beautiful new trifecta, pre -training, post -training, reasoning. Can you help us unpack what else is left? What alpha is there in each left? So I think we're going to continue to see capabilities increase. It's going to continue to feel like it's felt super fast, super exciting over the last, even five years. And I think it's going to continue feeling like that. There's not a wall here. But what is going to be different is that 2025 is going to be the year of reasoning.
2:25So it makes a lot of sense. Reasoning is a new technique. When you have a new technique, you know, there's often an overhang of compute, of data, of algorithmic efficiency improvements that you can make. And so that's something that, you know, if you look at just the incredible progress that we saw from 01 Preview back in September, and then six months later, we go to O3 in April. And then, you know, at the same time, we also see diffusion of reasoning from OpenAI, who we've been working on that for years, out to Google, DeepSeek, Anthropic, again, just in a few months. And so this is really the right place where every lab is gonna focus for the year.
3:09And just like a sort of a fun example of how low -hanging the fruit is right now. If you look at the most interesting difference between 01 preview and 03. 01 preview is not able to use tools. 03 can use tools as part of the chain of thought. And this is pretty obvious, right? When we were training 01, we knew that this was a thing that we wanted to do, but it was difficult to implement. It took time. And so, you know, that took six months to get, you know, done and released. The next step on reasoning is going to be a lot less obvious than that. It's going to be a lot harder. And so, you know, as reasoning continues to mature, we're going to see the overhang get eaten up and it's going to start being slower and slower to make progress.
3:51You said there's not a wall. I think there's this meme and the Twitter sphere right now that pre -training is hitting wall. Can you say more about that dynamic? Yeah, and I think that's great. That's a great question here because pre -training is not going away. But what we're seeing out of pre -training is that we are at the place where it's working really well and we're hitting diminishing returns. And so, you know, diminishing returns are baked in because the intelligence of a model is log linear in the amount of compute that you're using to train it, which means that you have to have exponential increases in compute to get each increment in intelligence.
4:26When you pre -train a model, that's a giant training, it takes all of your data center for, you know, a period of months. And when you go to pre -train the next model, you can't really do it on the same data center. You can rely a little bit on algorithmic efficiency to make it better, but fundamentally you have to wait until you get a new data center. And that's not a matter of like something you can do in six months the way you can do improvements in reasoning right now. That's something that takes years. So that doesn't mean pre -training as useless though because the real lever for pre -training in 2025 is improving architectures.
4:59So even though you're working on reasoning, you want to improve pre -training so that you can have better inference time efficiency or so that you can have longer context or better use of the context. And when you're doing that, you have to start back from the beginning, do pre -training on this new architecture and then go through the whole reasoning process again. So that's the role of pre -training now. It's still important. It's just doing something different in the pipeline. Can you help us unpack what's left in post -training? Yeah. So post -training is pretty interesting because both pre -training and reasoning are about increasing intelligence.
5:32And there's a very clear scaling law that you get where you put in more compute and you get out increasing intelligence and post training isn't like that post training is about model personality and
5:43Poet you know intelligence. It's a sort of a thin problem right if you can get better at it it turns out to be very generalizable and it applies to everything so you can work on math and you find that it makes you better at legal reasoning but model personality it's a thick problem. You actually need a lot of human effort to think about what makes a good personality. How do I want this agent to act? And it's much more of a training process like you would go through over many years of interacting with people. And now it's a very hard research problem to take that specification for what the agent is and turn it into an actual appealing personality.
6:20But when you think about post training, I think about people like Joanne Zhang, at OpenAI or Amanda Askel at Anthropic, who really spend a lot of time crafting these model personalities and they're not research practitioners. Interesting. Right? They're people with their product managers or they're people with a very deep understanding of human nature. And are there more legs to the stool? Well, okay. So I'm going to say something potentially controversial and I think actually there aren't. So I think if you go forward to 2030 or if you go forward to 2035 and you look back and you say, what were the fundamental concepts that you needed in order to create more and more intelligence?
7:01Maybe that's AGI, maybe it's something different at that point. I think you're going to come up with the idea of language models with transformers, the idea of scaling the pre -training on those language models. So GPT -1 and GPT -2 basically, and then the idea of reasoning. And sort of woven throughout that, increasing more and more multimodal capabilities. And I think even in 2035 we're not going to see any new trends beyond those. And the reason I think this is if you go back to 2020, so GPT -3 has just been trained. You know, imagine yourself sitting where it opened AI, we haven't released this thing, but we know something apoccal has happened.
7:40And you know, Daru, Emma, Dye, Ilyas, Satskever, Alec Radford, you know, we're all sitting there in the room looking at this thing. And it was fairly obvious internally what the roadmap was. We knew that at this point, going from GPT -3 to GPT -4 by increasing pre -training, was absolutely critical. We could see that we needed to increase multi -modality, ultimately ending in a model that could use a computer. We were beginning to make experiments with test time compute. And in 2021, after the anthropic people left, we really started developing the idea of reasoning at OpenAI. And it's funny actually, if sometimes my friends ask me after an anthropic release computer use, they're like, did you see that coming?
8:24And I was like, well, we were working on that together back before they left. One of the people that did that project went to anthropic and the other one went to open AI and developed operator. And it just took many years before the multimodality had matured enough to get to that point that was obvious to us way back then. And so that's why I think, you know, we're going to see from here on out, there's very important scaling. There's very important development and refinement of these ideas. That is extremely hard. It takes a lot of brain power. It's not going to be easy. But I think if we look back from 2035, we're not going to see anything new and fundamental.
9:04I think I'm right. I kind of hope I'm wrong. It would be a lot more fun if I'm wrong. But I think we'll have to see. That's a hot take. I'm glad we have it on the record. We'll see in 2035, that's amazing. I'm curious about reasoning. It seems to me that opening up I really leaned big behind this paradigm, probably before the others. Now everybody has reasoning models. What did you see in reasoning that caused you to lean so far in so quickly? Well, effectively it really was sort of this missing piece, where with pre -training, the model has an intuitive sense of how to answer the question. But if I ask you to multiply two five digit numbers, this is something that's completely within your capability.
9:47But if I actually do it right now, you wouldn't be able to do it because it is just natural as a human capability to be able to think about something before we answer to have a scratch pad to be able to work through a problem. And that is something that, you know, the initial models, even GPT -3 really didn't have. And so we began to see, you know, glimmers of this publicly, things like, you know, thinking step by step. And the idea of having a chain of thought that you could train, the model would learn itself how to guide a chain of thought, not just be guided by cloning from publicly available data on how humans think.
10:24That was very powerful. And we knew that it would be more powerful than pre -training because in fact, your thoughts are inside your head. They are not something that the model has access to. And so almost all the data that's out there is actually something that's just the final process. But you don't get to see that chain of thought. And so the model had to figure it out for itself. That's why reasoning mattered. You alluded to earlier that we probably still have to uncover more things in reasoning. Do you think we have a good sense of what those things are today or are we earlier in that R &D stage?
10:57I think at this point with reasoning, if you are at the call face, then you're seeing a lot of ideas and refinements of things that you can do. I think we've gone past the point where if you're on the outside, if you're not at a front -chair lab, you're probably not seeing them anymore. This is the same situation we saw where at one point, academic labs could make huge amounts of progress. Later, I would begin to see academic papers and I think, oh, they rediscovered this thing that we found a long time ago. And so, you know, now the level of effort that's being put into this, I think is actually quite intense.
11:37So there are definitely things to be discovered, but they're not sort of simple ideas that you and I could talk about. Cool. Switching gears a little bit, you tweeted recently about agents, I think a very, very interesting take, that agents will be incredibly powerful, but priced at the cost of compute due to competition. If that's the case, where do you see the opportunities in new startups and companies that are now building agents. Yeah. So I mean, I think that the thing about agents is people think, well, you know, I'm going to go develop an agent and they look at how much the job is worth out there by a human.
12:11So, you know, you want to develop an AI lawyer and you think lawyers get paid a lot of money. So I'll be able to develop an AI lawyer and the lawyers, I'm going to be able to charge huge amounts of money. 10 thousand dollars. Exactly. Exactly. Right. But the reason lawyers are expensive is because their time is scarce because there's only so many people who have undergone that training. But by the time you've made an AI model out of it, well, now there's effectively an infinite number of lawyers. And so it's not scarce at all. And maybe you, with your AI lawyer startup, will be able to have a lead over other people.
12:43But it's the same frontier model underneath. And some other startup can come in and compete that away. And so we should expect to see it priced at some opportunity cost over the cost of compute. You're changing. You now have a lot more supply, infinite supply of the highest capability intelligence in whatever domain you now have. On the one hand, there's a story where this is bad because startups can make money, but this is actually the future we want. We want services that don't require people to be extremely cheap. You want everyone to have access to a lawyer. What you want to be expensive and scares are things that are actually about personal relationships.
13:24So maybe we won't be asking the human lawyers to write contracts because agents will be doing that for us, but we'll be asking them for sort of deep advice on how legal challenges affect the detailed challenges that I'm facing in my business. And I think that's the world we want to live in. You think application companies will make any money selling agents though? Like, where would you tell us to invest? Yes and no. So just to back up for a second, people often talk about where does the value accrue in the stack, right? Is it at the model layer? Is it at the application layer? And if you look at the model layer, it's very competitive.
14:02Every company has a frontier model. Some of the frontier models can do things other frontier models can't, but by and large, they're all really very good. And if you're an enterprise, you can swap them out very easily. And beyond the frontier, you know, all the models that are answering the bulk of questions are distilled are very competitive. And so this isn't a very good business to be in when you consider the cost of training the models. So what's the point of training models in the first place? It's to give you an option. It's to give the frontier labs an option on the valuable places in the application layer that are coming up.
14:34So you know, chat GPT, that's a great business, right? There's a lot of competition over that. I think probably it's too late to replace chat GPT. Maybe not. You'd have to do something very different. coding another place where all the frontier labs are eager to see right now. I think you can compete with the frontier labs, but you want to do something that's different, something that involves more than just you know, you talking to your computer, you doing some sort of personal productivity task on your computer, something that involves other people, something that involves an enterprise. I think that's, you know, I think that the modes that you have for your business are going to be the same motes they always were.
15:13There will be network effects, brand, economies of scale. And so you want to find an agent that allows you to have those network effects, not just something that would be high priced out in the world. Are there particular domains that you think are maybe outside of the scope of what frontier labs want to innovate in and build in that you think are interesting and that you've been mulling about? We've got scientists, lawyers, research analysts, agentics, offer engineers. What are the domains have you been thinking about? I am, so personally, I'm very interested in robotics because I think robotics is something, I wouldn't actually say it's off the roadmap of the Frontier Labs right now, but I think it's something that's far enough away that to me, it feels like where AI was a few years ago.
16:02And so I think this is a very good time to be a company like skilled or a company like physical intelligence, or to start a new robotics company, maybe not one that's competing with those two, but something is doing something different, something on its own. I think it's at the end stages of being a research challenge and a matter of months or years, small digit years away from being commercialized. So I think that's really fun. Why now? What do you think has changed? Open AI famously had a robotics effort for a long time. What do you think has changed? Well, you know, so in between Palantir and OpenAI, I actually wanted to start a robotics company myself.
16:38And I got to the point of teaching a robot to play checkers from Vision back in 2016. Wow. Yeah, it was very cool. It could pick up the checkers pieces. It could pick up the checkers pieces and it could move them to a different place on the board. Nice. And my conclusion from this was that it was very fun and super cool and really far away from any form of commercialization. And when we pursued robotics at OpenAI, we didn't pursue it for commercial moment motives. It was really a demonstration of the power of machine learning and some of the ideas we had there later played into large language models.
17:12But I think what's really changed is that now that you have LLMs, you have this language interface to the robot so that now you can describe the tasks much more cheaply. And you have really strong vision encoders that are tied into that intelligence. So that gives the robots really a headshot at doing generic tasks. So we spent years solving one specific problem, teaching a robot to manipulate a Rubik's cube. And now a company like let's say physical intelligence can spend months solving a huge variety of problems like laundry folding and cardboard and packing egg crates. And that's something that they can only have because they're building on top of existing frontier models and the entire tech and research stack that we've built over the last 10 years.
18:04Yeah, I'm going to go back to this point you had on, you know, where is the value? And I really like to be framing that the foundation models kind of have an option on whichever parts of the application stack they want to own. How much of the application market do you think the foundation models will win? I think I would look at this like in a slightly different direction, which is if you're a startup, you know, where is it safe to play and where is it that you're going to get steamrolled by the Frontary Labs? And so I think I think the areas that that I think are safe to play in are areas where you have to understand something very deeply outside the model.
18:40And so I think a lot of enterprise really has this flavor. So for example, you know, Poundter AIP actually really fits this where it's, you know, it's not a model company, but it's something that sits outside the model that interacts with the rest of the business. There's another company I'm an investor and an advisor and called Distill that builds AI systems that allow you a business to sort of extract the context from within the business, feed that to the models, and then use that to make decisions. And so, you know, these are things that the Frontier Labs don't want to do. The Frontier Labs see business problems as how do I train a model to do something new?
19:22And if you look at all these enterprises, each one of those is a very small problem. It's not worth open AI or in Thropic's time to train a model specifically for each one of them. If you look at the problem and you think about what is the system that goes around the models And how do I use the models to sort of ease the context in and get the outputs out? Then suddenly that's one problem and I think it's a big opportunity. What are the specific use cases and problems that distill and palintea's effort solve for those enterprise companies? So a lot of times right now what you see is you're trying to automate some existing piece of work.
19:59And the easy cases are where that piece of work is in a regulated industry. And you're working on something like healthcare, maybe you're interacting with insurance companies. And you have a workflow that is extremely scripted where the company cares a lot about fidelity to that workflow. And that doesn't mean you can just say, hey, AI, go read the clinical guidelines and make these decisions. But with a process of transformation, you can get it to the point where the AI can do that. And that's sort of the low hanging fruit. And then the next level up though is that if imagine you're working on something that isn't a regulated industry or that isn't extremely scripted.
20:39And you want to ask someone, you want to automate some labor and sense process. Well, the first thing you have to do is make that legible. And if you go to someone and you ask them to describe their job, a lot of times, their manager doesn't know what they do. They don't even really know what they do. They can give you examples. But they can't say like, this is the workflow that I follow. because in practice, they don't follow a single workflow. And so I think that is what a lot of these problems look like. And for example, that's actually what Düsseld is doing is to work with companies, help them take the data they have, interview the people with AI, systematize all of that, and have it be something that an AI model can actually execute.
21:23That's really interesting. So this is also somewhat related to this other question I wanted to ask you about proprietary data. I was surprised to actually see you tweet this, but I very intrigued by this question that you posed, which was, how valuable will your proprietary data be compared to what your competitors infinitely smart, infinitely patient agents can estimate from public data? Can you impact that for us a little bit? Yeah, so, you know, a starting point for this is a few years ago, there was a lot of interest in training industry vertical specific models, you know, that finance companies would say we've got all of this data that no one else has And we're going to train a finance model on top of GPs or on top of Lama And it's going to be so much better and actually all of those were worse than the next generation of GPT because the power of intelligence and the ability to synthesize new information was bigger than the power of sort of memorizing the old information that you have.
22:21So that's I think what this theme looked like a couple years ago. But, you know, fast forward a year or two years. You know, now the story is I have all of this proprietary data, I've accumulated it over years. And in some sense, for a lot of times, you know, if that data is teaching the model of skill, or if it's meant to teach the model of skill, that data is sort of embodied labor, right? Someone worked through all these case studies. Some one called all of these customers and found out all of this information. Well, that embodied labor is now free. AI can do all those things. And so now there's an opportunity you can have AI call all those customers.
22:58Do a big survey, find out what they know. You can have AI work through all the case studies. Just a lot of chats with 03, right? And then now you've replicated that proprietary data, but without needing all of that work. How do you square that with the value of real world proprietary data, say something like what cursor gets from its developer community constantly or Tesla autopilot over the last handful of years? So I think those are in the middle because they're really huge amounts of data. I think there are challenges sometimes to training on the data that you get from your users. A lot of time, you know, one thing models can't do is if you train data and you memorize data about a specific person, maybe that leaks out into the next person.
23:48So that's a real challenge to using these kinds of proprietary data. I think there is a kind of real world for proprietary data that's very useful, which is data, very specific data about very specific customers that they trust you to use on their behalf. So to give an example, my financial advisor knows a lot about me. She knows my entire portfolio, she knows the kind of objectives that I have. And Ristallers, right. And she uses all of that information to give me a better outcome, which is what is the next asset I should buy. And she doesn't do that. Like the data doesn't make her a better financial advisor.
24:21It doesn't teach her skills. But it allows her an opportunity to use the skills she already has. And so that's the place where I think proprietary data is really useful. I want to switch gears a bit to coding. It feels like software engineering has just gone through this fast takeoff moment. And just judging from the pace of how quickly things are changing, there's at least a certain subset of the market that thinks the superintelligence takeoff, probability is a lot higher than the folks thought it was just given how quickly coding has taken off. What's your view of what's happened in the coding space?
24:56So I think, you know, on the one hand, coding has taken off very quickly. On the other hand, way back in January 2020, as soon as we saw GPT -3, we launched a project to train GPT -3 how to code. And so, you know, when you look at an exponential curve, you know, the progress is actually the same the whole time, but the impact of that progress can become very non -linear when it passes a threshold. And that's what's happened with coding in the last couple of years. And so my take on where coding will go is that you're going to continue to see a mix of coding with the user in an IDE, traditional cursor -style work, and coding in the background as an agent, something like Devon -style work.
25:43And this is going to continue for a long time as a year or two, maybe, is a long time in AI adoption. That's forever in AI years. But that, you know, if you think about something like vibe coding, right? Like the story you hear with vibe coding is you can, you know, if you have a PM, right? And you want to create a demonstration project. I think that you're going to see PM's vibe coding really cool prototypes, really cool demos that they can use to get user feedback. But then those things are going to get thrown away and that they're going to get rebuilt with professional software engineers because, you know, if you are given a code base that you don't understand.
Read the full transcript
26:21This is a classic software engineering question. Is that a liability or is it an asset, right? And the classic answer is that it's a liability. Like you have to maintain this thing. You don't know how it works. No one knows how it works. That's terrible. Usually the answer is it's actually cheaper to rewrite it from scratch. And so we don't yet have a way that we're comfortable with agents being the ones that understand the code base right now. I think the liability has gone down, but it's still net a liability. you need humans to do the design to understand the code base at a high level so that when something breaks, when the project itself becomes too complicated for the AI to understand, you can have a human do a problem decomposition and break it down into problems that are small enough for the AI.
27:04What do you think happens after that one or two years though? Oh, I don't know. We're going to have to find out. I love your bifurcation, though, on one side, and Agents Software Engineers that handle tasks autonomously in the background, and on the other side, human programmers who code in an IDE with the help of AI. I don't think that most of the mainstream actually, believe that, realize that, can you maybe unpack that for us a little bit? What would the Agents Software Engineers who handle these tasks autonomously handle? And then where do you see this other end of the spectrum go? Did they collide at some point?
27:39Do you think they remain separate things over time in the long term? I think it is already a spectrum where the things that your agentech software engineers can do, you can say, well, fix a bug, do a refactor, something that requires relatively little taste and has a clear outcome. Another great use case I've heard is translate software from cobalt into Python, right? It's very clear when you've done this correctly, but it's a lot of work. It's very boring and you can't get smart people who want to work on this and do a good job on it. On the flip side, if you're doing something where it requires a lot of taste and taste in how it's implemented, where there will be non -obvious consequences to how the implementation works.
28:27Maybe there's non -obvious performance consequences. Maybe there's non -obvious consequences in how the user interface is going to evolve and therefore how that needs to change the abstractions deeper in the system. Those are places where right now we have no alternative but to have humans do that work. And I do think this is very interesting. Is there a way, you know, is there a sufficiently detailed spec or a sufficiently detailed architecture diagram that the agents can be writing for us. That means that when you take work from one agent and you put it into another agent, which could just be the same agent the next day with a different context window that it's able to actually make progress on the code base.
29:07So these are these are the kinds of questions I want to see the answers to over the next couple years. Love it. It's exactly we're working on that reflection. Perfect. Why is it called member of the technical staff? That's a great yeah that's a great question. So for a long time this was true even before I joined OpenAI by the way I believe this was Greg Brockman's idea but we really wanted there not to be a distinction between engineers and researchers. If you look at a classic lab, a place like Google Brain, for example, where a lot of the people who started opening eye came from, at the time, and maybe still today, there was a big differentiation between whether you had a PhD and you were a researcher, or whether you were a software engineer, and you did data, you did implementation.
29:55And it was bad because the researchers didn't feel like they could get their hands dirty writing data code or writing implementation code. And you can't understand the systems aspects of your research unless you're writing the code. If you think about what makes Alec Radford the genius researcher that he is, it's each time he does something, it's that he looked very closely at the data. And he thought, what are the possibilities of this data? He wrote his own data scraping code from the very beginning. And so if you want to have someone who really understands the full stack, I think Paul Graham has this great analogy to painting where the resistance of the medium dictates the kind of painting that you're able to make.
30:37Research is very much like that. It's very much an artistic endeavor and researchers themselves are artists and should act like artists. And so by not having that to things just by calling everyone member of the technical staff, we were able to have a much more level playing field. And later that really came in handy when we had people who didn't have PhDs. Many of the great researchers at OpenAI, a teacher, Ramash, Al -Qradford, many of these people don't have PhDs and in fact learned their trade by working at OpenAI. That's a great answer. The random throwaway question. I love that answer. So what AI sent recently, Sam Altman left us with some interesting fodder, which was how different generations use ChatGBT.
31:19He said, if you're old, you tend to use it as a Google replacement. If you're in your 20s and 30s, you use Chat GPT as a life coach or a life advisor. And if you're in high school or younger, then you're using it as your operating system. How do you see people use Chat GPT around you? How do you have your kids use Chat GPT? Yeah, so look, let think about that operating system comment for a second. At the very highest level, the total addressable market for Chat GPT is every user intent that requires thought or action that you don't want to do yourself. anything that you wish got done, but you didn't have to do is something that you might want to use AI for.
32:01And so there's, I mean, if you think about that, there's a version of that that feels very scary, right? It's like people don't do anything for themselves anymore. There's a de -skilling. No one learns how to do hard things. We're all just zombies watching our VR headsets, you know, watching movies. But I don't think that's actually what people want out of AI. And I I mean, this isn't the world we live in, I think that's true. But that's not what I want out of my relationship with AI. And that's not what I see people doing now. And partly, this is because the technology for chat GBT is an operating system isn't actually there yet.
32:34Pretty famously, you cannot use chat GBT to control your iPhone. But it's also not what people want. And so I see this with my son. He's eight years old. He's been using chat GBT from pretty young age. I used to ask him to test the models before they were publicly released. He always gave pretty good feedback, actually. And he spends a lot of time with Chad GBT. He knows it is not his friend. It is not his companion. It is an expert. Someone he can talk to to explain things to him. And if you were eight years old, having someone who can explain things to you correctly, in great detail, and with a lot of patience, is a very valuable thing.
33:15And so he has like curiosity, he has enthusiasm. And one day he decided he wanted to be a coin collector. And so he collected all the coins in the house, sorted through all the ones that were from before 1970, went to chat to BT, started typing, and just asked, took pictures, and just asked questions about every single one of the coins from before 1970. And he's, you know, what's this worth? Well, what would make this worth more? You know, how can I test? What is a mint mark? You know, all these different questions. And if you think about this, this is something, you know, when I was a kid, I probably could have learned this.
33:50Maybe there were books, there were magazines, maybe I could have looked it in the cyclopedia. But all of this is just so accessible now. And it's accessible to an eight year old. And so when we went on vacation, we took him to a coin shop. And the staff at the coin shop were just shocked how much this eight year old knew. And the very detailed quite, he was a show me all your coins. No, I don't want that one. I want one that has a San Francisco mint mark. I want one from this year. This is the year that they were all made out of silver. And the coin shop owner was just very surprised. He doesn't deal with kids that had that level of detail, at least not until now.
34:30And so this is, I think, actually what we want out of AI, is that AI should make you an expert at the things you want to do. And it should remove the burden of doing the things. that the boring things that you don't want to have to do. Yeah. On the topic of the next generation, how else are you preparing that next generation for all the capabilities to come in AI? I think this is a super, super tough question. If you think about any particular field, you know, should you teach your son how to code, right? Like I think about my year old, you know, you know, my daughter is writing essays. My eldest son is really excited about math.
35:08all of those things are going to be automated. And so it's clearly not some specific skill that you have to teach them. I think there's really two things that I want my kids to understand. The first is the process of learning and figuring things out. So that's the value and the math and the essay writing and the coding. It's sort of this process of being able to learn to learn. The second thing is the, you know, having ideas and projects and the belief that you can do it, and the ability to use whatever tools are at your disposal to figure it out. So this is agency, right? And so that's where I think that's the right way to have kids use AI right now.
35:51And there's always a tradeoff. I'm often very torn between, so my eight -year -old uses chat -chefity for a lot of things, but I don't let him use it to code, because he's trying to learn to code. And if he sees that he doesn't have to use it to code, then it's going to be very hard for him to do that work to get all the way there. I don't let my other kids use it to do their school assignments, of course. Why would you do that?
36:15But I want them to have those basics, and then once they have the basics, once they understand things one level down, to then be able to use it to extend their capabilities. And here's another fun story about my eight -year -old. Last week, he decided he wanted to build a project where the grandparents who were coming to visit could press a button and he could go, it would ring a buzzer in a different room and he could go get breakfast and bed. And he asked Chad G. B. T for help. Yeah. I mean, he asked Chad G. B. T for help. It said, okay, you need jumper wires, you need two Arduino boards and just a sort of list of things.
36:55And he asked a lot of questions, how is this gonna work? He asked it to give us a list of Amazon links for us to buy. I reviewed this, made sure he wouldn't get electrocuted, bought the items on Amazon for him, and now we're putting it together. And my approach in this is I'm gonna let him put it together, everything he can. I'm gonna install the software, because his computer's locked down, he can install software. And this is gonna be his project. That's amazing. None of us could have done that at eight years old. And he has learned so much in doing this. It's not just that he outsourced it all to Chatubt.
37:32Now he understands what Arduino is. He understands what the circuit board is. What is going what happens when I hit this pin? Why is this pin named, you know, GRP1? You know, these are all, I mean, I don't know the answers. These things either. So it's really, you know, it's just this huge help that Chatubt is able to do all these things for him. That's amazing. Sparking curiosity and then agency, I love it. And it's also just the time to impact. And that just feeds more and more curiosity and agency. Yeah, that's right. I mean, if you think back, you know, well, you want to do this project. Well, here's a book on Arduino and, you know, you're going to have to write the code yourself.
38:10And, you know, what... The parapsychology. Circuit boards, am I supposed to do? I don't even know how to do that. Yeah. You know, probably this project just dies on the vine. And, you know, there's a there's a truism in education theory that when someone asks a question, That's the time when they're ready to learn the thing that they're asking the question about. And so, you want to, you know, it's worth going off script to answer someone's question, because you're doing a huge service to them in teaching them that thing right then. And that's, you know, now you have that. You have the ability to get your questions answered on demand at the right time for you when you are mentally ready to do it not when maybe you're tired and you're in school and you're thinking about all sorts of other things, just right then when you actually want to know the answer.
38:55And I think that's hugely powerful. So how else are you using AI in your daily life? Chat you B .T., deep research I'm sure, how we AI for scheduling, maybe autopilot. What else? So yeah, so I am pretty much exclusively use O3 at this point. Once you use a good model, I think it's very hard to go back. I think I could probably use Gemini 2 .5, I hear it's really good. but of course as we talked about, if it's good enough, why switch? And I use deep research about five times a week. And it's hugely helpful. And I think even one time that it saves you a few hours of doing work sort of repays the cost absolutely makes sense.
39:37What do you use deep research for? It's a mix. One answer is I'm batting around something with my kids. And it's a question that no one has ever asked before, probably. And I want to know the answer. For example, what happens when you compress wood? You know, it starts off, it's elastic compression and then it starts deforming. And then you go a little further and it becomes diamond. And then you go a little further from that and it becomes a black hole. But actually, there's like a dozen steps. And so that's a really fun topic to dive into. And just, you know, this is kind of thing that would have been an XKCD comic 15 years ago and would have taken him weeks to figure out.
40:17And now you can get an answer just in a few seconds. Also, I use it when I'm thinking about a new domain or a new startup opportunity. Well, if I'm interested in robotics, tell me everything there is to know about a particular company or about a particular market. And - That's our daily life. Yeah, yeah, yeah. Any other new products? Well, like you mentioned, I use an AI assistant for scheduling, which is great. I mean, I'm solo right now. You know, I could hire an assistant, but it's just actually more fun to do things myself. But calendaring, it's really boring, and it's just very nice and very pleasant to have be able to see, see, you know, an AI agent and have it do the calendaring for me.
41:03I'd love to hear a little bit about managing, you know, open AI, the research org. get such a collection of insanely smart individuals, creative, I'm sure. And the feedback we have on you is exceptional in terms of what a fair and what a great manager and leader you've been for the organization. I guess what have been some of your lessons leading in organization like that? So the sound sort of boring, but the core thing that you have to do as a manager is you have to really care about the people you're managing. And this maybe isn't relevant a lot of the time. A lot of time as a manager, your day -to -day job, you know, you're coordinating, you're helping people understand things, and loyalty doesn't really matter that much.
41:52But there comes a time as a manager when you have to ask someone to do something hard. Early on in your career, this is when you have to ask someone to come in and work on Sunday when they'd rather be playing basketball. But later in their career, it's, you know, working with someone and you have to ask them to give up a project they really care about and Give it to someone else or share credit for a research breakthrough that they know they could get to by themselves But that they that you know a team of people, you know, not just this one talented person, but two very talented people are three very talented people working together could get done even faster and And one thing I learned from working with Alex Carp at Palantir is that very talented people have superpowers, but they also have debilitating weaknesses.
42:41And for people who are at the very edge of these capabilities, they often don't even understand what their weaknesses are, but it's extremely apparent to everyone around them. And, you know, for me as a manager, it's something that I could see very easily. And at this level of capability, when people fail, it's almost always a form of self -destruction. Wow. That there's a choice that they could have made. And, you know, and I don't mean little failures. I don't mean like, oh, I had a bad day. I mean, you know, when someone makes a career altering choice in a bad way,
43:22it's almost They had to do something that was very difficult for them. They had to confront something that was extremely scary for them to do. That to everyone else is kind of obviously the right answer. It's obviously the right thing for a company, but it's emotionally extremely hard for them. And if you, going back to being a manager, if you as a manager, if people know that you're in it for yourself, when you tell them to do something, they won't trust you. But if they know that you were doing what's best for them, then when you tell them to do that thing that is super hard and extremely scary for them, sometimes you can help them across the chasm and you can solve the problem and prevent them from doing something really stupid and end up with something that works out really well.
44:10And I hold this bar even for firing people. For me, it is always, when I am talking to someone, I have to be talking to them, giving them advice, helping them do the thing that is best for them and for the company. Even if you're firing someone, if they're not going to succeed in this role, and you have, I have invested enough time to make sure that they won't succeed in this role, then it is in their own best interests for me to tell them that they're not succeeding and give them the opportunity to find somewhere else. Loyalty in the end is the thing that I think unlocks all of the other things that you want in management.
44:50I really, really love that. There was a nuance there that you said in the middle around working with a ton of high performing individuals who are really excited about a particular research direction that they want to break through. They know they can get there potentially by themselves, essentially with one or two others. They all have a good dose of confidence, maybe sometimes ego. How do you actually get them or convince them to embrace that effort of working together to get there? Yeah, it's very hard. And I think this is actually one of the things that's very different about a research lab from an engineering culture.
45:27Because in an engineering culture, it's sort of an assumption that we're all working together. We're all building one product. But research often comes out of academia, which has this very negative culture of, it's a PI, it's his team. Who's going to be the first author? Who's going to be the last author? None of the other people in the middle matter. And we struggled with this a lot. And I don't think there is any one answer. One thing we tried, which worked well for a time, We published some papers where we actually had open AI be the first author So that there wouldn't be the fight over who's the first author That was you know one technique.
46:05It didn't know we you know we couldn't always do that. It didn't always make sense But you know in the end the the key is really when you work with people you understand there's something they want and You have to find a way to give them the thing they want and let them do the thing they want to do, the art that they're trying to create, while also letting all the other people do that and having it all add up to one big hole, and just spend time over and over again, making sure you're solving that problem. Security, I know, is an interesting topic to you. In an increasingly agentic world, what kinds of security issues do you think we should be aware of and where do you see potential opportunities?
46:47When I think about how AI impacts security, for me, the first order is the much easier ability to do offensive work than you could do previously. And so the number of threats have gone up. The time to execute on a threat has gone up. And so that then pushes the defense to be much more agentic. So there's a company I'm an investor in. It's called Outtake. I met the team there a group of ex -palentary folks and we also ended up using them very successfully at OpenAI. And what they've done is that they have made an agentic stack for doing cybersecurity that uses very little human input. And I think this is, you know, right now we're at a place where the models can actually do all of these things.
47:38You know, if there's something that a human could do that's sort of these, one of these bulk operations, If you can't make the model do it, that's your fault. It's not the models fault. But the barrier then is that businesses and organizations aren't set up to do this. They have to go change their business processes in order to make this happen. And so I think that's an opportunity for startups where similar, you know, as big as, you know, this shift from web to mobile is just disrupting the existing businesses because it may be faster for you to replicate their technology and their distribution than for them to be able to change the way they operate or reduce the number of humans they need.
48:16Yeah, awesome – Bob, thank you so much for joining us. This has been™es pleasure to have you here.</b></b></b></b></b></b></b></b Dell
From the publisher
As OpenAI's former Head of Research, Bob McGrew witnessed the company's evolution from GPT-3’s breakthrough to today's reasoning models. He argues that there are three legs of the stool for AGI—Transformers, scaled pre-training, and reasoning—and that the fundamentals that will shape the next decade-plus are already in place. He thinks 2025 will be defined by reasoning while pre-training hits diminishing returns. Bob discusses why the agent economy will price services at compute costs due to near-infinite supply, fundamentally disrupting industries like law and medicine, and how his children use ChatGPT to spark curiosity and agency. From robotics breakthroughs to managing brilliant researchers, Bob offers a unique perspective on AI’s trajectory and where startups can still find defensible opportunities.
Hosted by: Stephanie Zhan and Sonya Huang, Sequoia Capital
Mentioned in this episode:
Solving Rubik’s Cube with a robot hand: OpenAI’s original robotics research
Computer Use and Operator: Anthropic and OpenAI reasoning breakthroughs that originated with OpenAi researchers
Skild and Physical Intelligence: Robotics-oriented companies Bob sees as well-positioned now
Distyl: AI company founded by ex-Palintir alums to create enterprise workflows driven by proprietary data
Member of the technical staff: Title at OpenAI designed to break down barriers between AI researchers and engineers
Howie.ai: Scheduling app that Bob uses




