Ep 66: Bob McGrew — the Superstar Palantir Alum Leading OpenAI's Transformative Research Projects

18 Aug 2023 · 43 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Podcast Summary: Joe Lonsdale: American Optimist - Ep 66: Bob McGrew

Episode Overview In this episode, Joe Lonsdale interviews Bob McGrew, the Vice President of Research at OpenAI. The discussion explores McGrew's background, his pivotal role in the development of AI technologies, and OpenAI's mission to advance artificial intelligence responsibly. Topics include the evolution of GPT models, the concept of artificial general intelligence (AGI), and the ethical implications of AI.

Key Themes and Discussions

  1. Bob McGrew's Background
  2. Early Passion for Technology: McGrew's fascination with computers began in childhood, influenced by his father's career in computer science.
  3. Educational Journey: Attended Stanford University, where he studied computer science and engaged in various projects, including an internship at PayPal.
  4. Career Milestones: Joined Palantir as an early engineer, where he contributed significantly to developing products for the intelligence community before moving to OpenAI.
  1. The Evolution of AI Models at OpenAI
  2. Role at OpenAI: McGrew joined OpenAI in 2016, becoming a central figure in research projects, including the development of the GPT series.
  3. GPT Models: Discussion of how GPT-4 was trained and what advancements are anticipated with future models like GPT-5 and beyond.
  4. Neural Networks: Explanation of how neural networks function, emphasizing their iterative feedback mechanisms and the importance of large datasets.
  1. Breakthroughs in AI Research
  2. Dota 2 and AGI: Insight into the Dota 2 project, which used reinforcement learning to develop AI that outperformed human players, showcasing the potential for achieving AGI.
  3. Interdisciplinary Collaboration: McGrew highlights the synergy between various AI projects at OpenAI and the importance of allowing researchers the freedom to explore unconventional ideas.
  1. Addressing AI Bias and Ethics
  2. Bias in AI Models: Importance of ensuring AI behaves professionally and appropriately, as models learn from vast amounts of internet data, which may include biased content.
  3. Civil Rights and AI Safety: OpenAI’s proactive stance on civil rights protections and the ethical implications of AI technology, emphasizing the need for responsible development.
  1. The Future of AI and Society
  2. Potential of AI: Optimism about AI's capacity to enhance productivity and living standards by alleviating labor constraints and generating new ideas.
  3. Application Layer Opportunities: Encouragement for young technologists to engage in creating applications that utilize AI in innovative ways.
  4. Infrastructure Needs: Discussion on the need for robust computing infrastructure to support AI advancements and the challenges associated with scaling AI technology.

Key Takeaways

  • Optimism About AI: McGrew expresses a hopeful outlook on AI's potential to transform society positively, advocating for those in the field to harness its capabilities responsibly.
  • Importance of Collaboration: Ongoing collaboration and exploration of new ideas in AI research are crucial for breakthroughs and advancements.
  • Ethical Responsibilities: As AI technology evolves, it is vital to maintain a focus on ethics, bias mitigation, and civil rights to ensure inclusive and fair outcomes.

Conclusion The episode underscores the transformative impacts of AI, the ongoing research efforts at OpenAI, and the responsibilities that come with creating powerful AI technologies. Bob McGrew's insights provide a glimpse into the future of AI and the opportunities it presents for innovation and societal improvement.

For more information, visit [Joe Lonsdale's Blog](https://blog.joelonsdale.com?utm_medium=podcast).

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00In two weeks we went from not being able to grab a ball with a claw to having a five finger humanoid robot hand solving Rubik's cube in simulation. And I was like, wow, wow. That's a big leap. And so to me, that was what convinced me that this was actually possible, that we actually had found a mechanism that if you could figure out how to pour enough compute and pour enough data into a neural network, there's really no limit to how good it could get.

0:35Bob McGrew is an old friend from Stanford. We were in a couple groups there together. He was ahead of me in computer science. He joined PayPal a year before me. You know, I learned a lot from him over the years. We convinced him to come on early at Palantir. He was a key early leader, helped run engineering, helped run product, really helped make Palantir into a great company. Normally, that'd be enough for a good reason to talk to Bob. But you know what? He joined OpenAI first in 2016, full-time in 2017. This was before the Transformer paper. I have no idea how he knew this was the place to go.

1:04OpenAI has transformed the world. He has the front row seat there, helping run research, helping really run a lot of the efforts going on at OpenAI. This is the guy to talk to, to hear what's happening in the AI world. I'm looking forward to having you meet him. I'm Joe Lonsdale. Welcome to the American Optimist. Bob McGrew, thanks for joining us today. Great to see you again. Bob is an old friend. We were in a couple of groups at Stanford together, including FI-Sci. And you were at PayPal a year before me. 2001, 2002. What were you doing at PayPal? I was working on cryptography. And tell us a little bit about your background.

1:33How did you end up at Stanford, at PayPal? Where are you from, Bob? Well, originally I came from a small town in Oklahoma. And my parents were college professors. And my dad was a computer science professor. And I remember, I think third grade, he bought me a book where two kids use computer programs to solve problems for detectives. And I was hooked. And so my whole childhood, I would buy these books and I would always be writing programs. But again, small town, Oklahoma. and this was during the first web boom. And so I always knew I wanted to go to Stanford. I never actually thought I'd get in, but I did.

2:12That's pretty funny, actually. I didn't realize that you're solving problems for detectives in the book. And then you go to Palantir, which also kind of solved problems to help catch the bad guys. It's a crazy coincidence, but I think it just shows that I knew what I wanted to do from a very young age. So you joined OpenAI what year? 2017. 2017. So for the last six years, you've been helping run things at OpenAI. But let's talk about Palantir a little bit. You know, I interviewed Alex Karp earlier, our good old Dr. Karp. And he talked a lot about how it took a really special type of genius to create what we did in terms of the engineers at, you know, Pounder Gotham.

2:44What was Pounder Gotham? What were you working on? Why was it hard? Well, I think the really funny thing about Gotham was the hardest part was we had no idea what it was supposed to be at the beginning. So the goal was to build this platform for the intelligence community. So basically software for spies. But as you remember, the first hard problem about a spy is actually finding one. And then once you find one, they don't actually want to tell you what they're doing. They want to tell you cool stories about using guns and little devices and stuff. But yeah, not really about the information architecture.

3:15Or you try to ask them what they're doing. They're like, oh, that's classified. And they give you some sort of tortured baseball analogy. And so instead of all the UX books, you interview your customers, you find out what their problems are, and you build them to solve their problems. And we couldn't do that. And so, basically, what we had to do was we had to build the product we thought they needed, which was wrong, and show it to them. And they would tell us we were wrong. They'd be like, oh, this is crap. Why don't you build this other thing instead? And we would do that. And you and Stefan would go into DC and show this to them and bring back that feedback.

3:46It was a combination, Rick. Because I do think there were certain first principles, but we all kind of could figure out they would definitely need. And then there was the iteration of what they actually needed. And so, we had to have a lot of both sides. We knew it was something about data, that somehow we had to get their data in, and that we would let them ask questions about their data. But we were totally wrong at the beginning about what kind of data they had. They had a lot less structured data than we assumed. Exactly. Yeah. It turns out, intelligence community, it's all about reading documents and trying to understand and parse documents and figure out who the people are in the documents.

4:13And if the person in this document is the same as that description they have in the other document. They spent over$30 billion on databases and gathering data, so I assumed that they had more structured data. So that was pretty complicated. Yeah. Yeah. And that was, that was mostly wrong. I think they had it and they didn't know what to do with it. Do you remember any of the characters? One of my favorites was Sir Mark. Do you remember him? I do. Yeah. He was the one who said that if you don't get your gun unless you're going to use it, and if you're going to use it to pull the trigger twice, he's the MI6 guy.

4:39You have any other favorites? I remember, I think I have to stick to first names here, but I remember our first customer at the agency. Her name was Sarah. And she was the only person that believed in us. I think the really crazy thing about Palantir was how long it took to get a product that actually worked. I think it took three years before we had anyone using it to solve. I remember a few people almost quit right around year three, right before that happened. And it was like this thing where like people were so exhausted, like this is just crazy and it's not going to work. And then they kind of, we got to stay on and then we kind of got over the hump and people started using it.

5:13I remember that. I remember, you know, and, and, and this, this is one of those things about the really great engineers is that a really great engineer wants certainty. Right. They don't want to, you know, they've been doing something for three years. All they can see are all the problems that are in it. And, you know, I remember sitting down and having a conversation with our best engineer and just sort of convincing him to stick it out for nine more months. And he built the security system that we ended up working with. It was a critical piece to allowing us to actually deploy. And then he quit.

5:43Yeah. But if he had. Well, thank you for holding the team together. We were both fighting hard to hold it together because it was kind of like, what are these kids doing at that point? I don't know if you know, you joining was really, really critical to Peter giving us our money for the second time. Because he was like, you guys got Bob to join? How did you do that? I did not know that. That was actually, yeah. So thank you. You're a good indicator for Peter Thiel. You've kept that a secret from me for all these years. Wow. I forgot to tell you earlier on. We all did very well from that. But you may even be doing even better financially from OpenAI.

6:14Not that we're talking about finances, but this is amazing. You guys have just crushed it in terms of changing the world the last six, seven years. How did you find OpenAI? What's the story there? Well, when I left Palantir, my PhD before I started Palantir was in AI. And I left because I felt AI wasn't happening. And this was 2005. And in fact, AI was not happening in 2005. Yes, it would have been very hard to work on AI 20 years ago. The key thing about AI was actually just big data. And you didn't need AI to unlock it. The last thing was that everyone who went to work on this, it was just a bad choice forever.

6:45It was exciting, but then it was never the right choice. Yeah, until it was. Until it was. So how'd you know that it was? How'd you have the intuition that it might be? Well, so what happened is, in, I think, 2011, there was a paper that was published by Ilya Sutskiver and some other people who, Ilya Sutskiver, chief scientists at OpenAI, where they basically reinvented neural networks. And they showed that if you ran them on GPUs, you could plug in huge amounts, much bigger amounts of compute and much bigger amounts of data. And neural networks went from being, oh, that one thing that you can kind of use for handwriting to something that was actually the best way to identify images.

7:18And for some of our listeners who aren't just technical, we're talking about a neural network. That's something that gives iterative feedback on things. How would you explain neural networks? So, a neural network, the analogy that most people use is that it's like the brain. And it is, but at a very, very high level. So, a neuron is basically, you can think of it as having connections to other neurons. And the strength of those connections determines when the first neuron, the lower level neuron, fires. Then that makes the higher level neuron fire. And if you train these, you do it by showing them the right answer and then basically doing what's called backpropagating the error from the top of the network where the answer is all the way back down through the whole network.

8:01Is it a power detection type of situation then? Or what is it? Yeah, you're building hierarchical representation. So at the very bottom, think about a vision network because it's easy to visualize. So at the very bottom, the neurons are detecting edges. And then they go up a little bit higher and they're detecting corners. And if you go high enough, they're detecting wheels. And then above that, they're detecting cars. If you have quarters in this way, then it's a square. This could be a wheel, and then it could be a car. Yeah, and at some point, there's a Joe Lonsdale neuron that is out there.

8:27And individual people, you can actually find these neurons in the networks. I remember reading On Intelligence by Hawkins, who didn't actually end up solving all these problems. But he said there's like six layers of the neocortex that is part of the vision system. Is that right, or is there actually a lot more than that? You know, I'm not a cognitive scientist. So the other thing is, I think the analogy between the brain and the neural networks, it's very inspirational, but you can take it too far. It's very imprecise. What OpenAI did is you basically didn't base it on the brain. You based it on building it up from scratch, from just the first principles.

8:54So I think that the thesis behind OpenAI, when I joined in 2017, was that neural networks were the final architecture that could take AI all the way to human-level intelligence, or AGI. But wasn't there a transformer breakthrough that was really important, though, right around 2017? Yeah, that's right. So, you know, the first couple of years of OpenAI, we were using sort of standard neural networks. And then there was a paper at Google. Some people came up with an idea called Transformers that allow you to take better understanding of the context. So, for example, you could look at documents. And now you could really understand the preceding three or four pages when the prior architectures had struggled just to understand, like, the previous 10 words.

9:38What's the intuition for why Transformers worked better? What's going on there? It's basically because they have a notion of attention. So you can pay attention to particular words in a document versus previous architecture had a notion of memory. So you could sort of remember the words that you had seen. But it's much easier to be able to look at a sheet of paper and sort of see the different words and think about reading those words. Pay attention to what matters. Pay attention to what matters. Which is kind of how the brain works. Things jump out at us, obviously, when we look at things. Yeah, it's very intuitive.

10:11The intuition, I know this is not based, you're not a cognitive scientist, but I want to try to build intuition here for people. My understanding, the big breakthrough we had for how the brain works is you're constantly predicting what you're going to see next. Like, this may be why we see ghosts, but either way, your brain's constantly looking at what it thinks it expects to see. And then if something is not unexpected, it kind of jumps out at you. Is there anything like that with transformers or there's nothing like that where it's trying to predict things? Well, that's exactly how you train a model like GPT-4.

10:36So, you know, if you take it to train a model like GPT-4, basically we take all of the text, everything we can scrape off the Internet. And what the model is trying to do is it goes character by character and it's trying to predict the next character. So, you know, if it sees, you know, the rain in Spain falls mainly on the it's going to guess, oh, that's a P and it's going to be plain. And it does this over and over again for huge amounts of documents, like trillions of characters. And over time, it seems as though that yields something that looks like intelligence. The prediction seems very, very similar to intelligence, which may be what we're doing, too.

11:13Yeah. At this point, I think we understand neural networks a lot better than we understand the brain. So whenever someone talks to me about the brain, I always think, well, what does a neural network do? And then probably, actually, the brain does that. Do you think of yourself as a neural network now? I do, actually. I actually think of my kids as neural networks. How does that change your interactions with them? you know as i was starting at open ai first i worked on robotics and then i worked on we did some problems with math and then programming and early on you know the kids would always get the problems before the robots were and now it's the other way around now the neural networks are better are you know better at solving problems than the kids because it given you any intuition on how to train young minds uh yeah you just show them lots and lots of things you just have them read just Show them three trillion things.

11:58I mean, that's probably how you and I learned, right? You know, just read a lot of books when we were kids. I guess there's a thing about attention, too. I've always found that kids do better when they're confident. I don't know if the computers need to be given confidence, although maybe that's a signal for something else like attention. Well, that's actually one of the cool things that we've discovered with GPT-3 and GPT-4 is that the models are trained on the Internet. And they're reading everything, right? Then they try to act like what they've seen, right? They're predicting. And so they don't know whether to predict being someone who's dumb or someone who's really smart.

12:31That's what you have to tell them they're really smart when you're talking to them. It's so funny. Exactly. Like you're a confident, you know, powerful physicist who knows everything. The best physicist in the world. Now answer my question. And you sound like Shakespeare because it's fun if you make him sound cool. That's right. Or you have to do everything in a limerick. That's pretty good. So let's back up a little bit. So you joined OpenAI in 2017. What's your role in OpenAI? So I'm the VP of research. What does that mean?

12:58So, if you're the VP of engineering at a company, so Palantir, what I spent all my time doing was basically, fundamentally, you're trying to figure out, are the right people working on the right problems? And figuring out what's most important and finding people to work on that. You would get in the problems yourself, too, but then you're more focused on the people. Yeah, I mean, especially early on, I read every commit and I wrote a lot of the back end for Palantir. But, you know, as the company scales, you know, you as a leader end up focusing more and more on the people problems. And, you know, in a normal company, you're often trying to figure out what the people should be working on.

13:31And at a research company, it actually doesn't work very well that way. So there's this researcher, Aditya Ramesh. He built a model called DALI, which allowed you to take a description and build an image out of it. It's very fun. He spent two years working on that. Wow. For most of those two years, no one other than him believed it would work. I mean, I sort of believed it would work. I really hoped it would work. I made sure he had the GPUs he needed. But in order to have that kind of drive to really push it through for two years, this is your dream. Yeah, you have to be obsessed with it, basically.

14:08Yeah. And I remember, you know, Karp always used to talk about this, that the best engineers were artists. And if you want to keep an engineer happy, what you actually have to do is just let him do his art. Just get out of his way. Let him do his thing. And your job when you're building a company is to have all these brilliant people doing their art. And somehow, magically, it turns into a product that solves a real problem. It seems very unusual to have one person do it for two years. Usually, there's a small team or something. And they're giving feedback and iterating. Yeah. And that's really what's different about research.

14:37There's a lot of things we do that have teams. And I think that's one of the things that we brought to research at OpenAI, is to have teams that can work on problems together. But a lot of the very early stuff is just one person. Wow. Like, this has to be their hill to die on, or it won't happen. So, you have people right now on your team going for a year or two on things that hopefully will work, but it's not clear. That's still a culture there for some of these things. Yeah, yeah. That's really cool. And as we've gotten bigger, I think we've sort of shifted. As a fraction of the company, there's fewer people doing these sort of hills they have to die on that are really early.

15:08And more, we've taken a startup approach to doing research. So, we have teams of people. We have broken out problems. But again, the difference between research and engineering is that if you're trying to tell someone, hey, can you build this user interface? There's not usually a question about whether you can. No one's like, oh, I can't build that. The engineering team is just like, we're going to build this interface. Whereas the research team is, we're going to explore and see if we could do this thing. What were some of those things early on, if you could tell us, from 2017, 18, 19, 20? Other than Dolly, what were some of the crazy research projects?

15:37Yeah, I think the first crazy research project we did was a project called... Basically, we were trying to beat the game Dota 2. So, there's this long history in AI. You beat chess, you beat Go. Dota 2 doesn't have quite the level of class, but it's actually a much harder game. It's a very hard problem. Are you a gamer? Yes, I am a gamer. Well, you know, obviously, I was chess champion, and then the Go thing totally blew me away and freaked me out because that was really scary. And then my favorite part was when they learned to play chess against itself and learn all the openings. Because we always used to have computers, like we'd tell them the openings because they just don't know them.

16:12And then they'd be able to play. And what was really cool was throughout the day, it learned all the old openings. And shifted from playing 15th century players to 19th century players to modern players to then playing even differently. Which is really creepy as a person to watch the computers do hundreds of years of human history in 24 hours. And that's actually one of the really interesting things. You train a model like GPT-3 or GPT-4. It's basically learning what humans have learned by trying to predict what they were doing. The model we trained for Dota 2 or the models you trained for chess or Go is basically doing something different.

16:42It's doing reinforcement learning. It's trying to win. They probably play very differently than people because they're so fast. They can do things we can't do, right? Well, early on, it plays a lot of the same things. And then over time, you see it develop new approaches that humans have never developed before. Yeah, it's because we're not fast enough. Or it also maybe just tries things we haven't tried. Yeah, I think since it's playing itself, it can sort of end up in equilibria that humans would never find. Interesting. And did it end up crushing Dota 2? Yeah, so we ended up beating Dota 2. And I think for me, what was really crazy about this is that when I wanted to open AI, Ilya Suskev, our chief scientist, sat me down and was like, we're going to build AGI.

17:19It's going to be human level intelligence. We're going to do this in 10, 15 years. And I'm like, that would be cool if it were true. But I have no idea how we're going to get there. you know and i would ask him and he'd say well you know we're gonna figure it out and um i remember one week i was sitting there i had spent months you know i was working on robotics not on the doter 2 problem and i was trying to get a robot claw like just a two-finger claw to grab you know a ball it's a hard problem right and i i just couldn't make it work and i remember thinking myself like why are we talking about agi if we can't even have like a claw grab a ball like This is really easy.

17:51We should be able to do this. And at the same time, the Dota 2 team, Jakub Paczoky, who was doing this, had developed a technique that allowed you to put lots of data and compute into solving the problem. And it was really working for Dota. And he came over and started doing this for robotics. And in two weeks, we went from not being able to grab a ball with a claw to having a five-finger humanoid robot hand solving Rubik's Cube in simulation. And I was like, wow. Wow. Wow. That's a big leap. And so to me, that was what convinced me that this was actually possible. That we actually had found a mechanism that if you could figure out how to pour enough compute and pour enough data into a neural network, there's really no limit to how good it could get.

18:37What games are left right now that you can't solve? I think Dota 2 was the last one. Is there really not? I mean, there must be some other more complex human game. So the last really cool game to get solved was Diplomacy. Yeah, I heard of that. Because, you know, what's smart about it is that you have to do a lot of interaction and they have to work with human players, not just with AI. There's a lot of subtle strategy with how you work together versus how you betray people and when you betray people. Very scary. We're teaching them this, but I guess this is like the natural strategy for these things.

19:08It's fascinating. So you can't think of, I mean, they must not be good at baseball, for example, just because the robots aren't good enough. That's true. Yeah. I mean, I think where we are right now is that if you have a job that is, you know, remote, doable remotely, which at this point is most laptop jobs, right? The computers can maybe get better at that. Yeah. It feels really achievable that we're going to be able to get there. This is very unintuitive because forever technology replaced blue collar work. But now technology is replacing the bureaucrats and the blue collar like plumbers are like, we're fine.

19:34But the like annoying lawyer people are like, they're in trouble. I mean, the thing right now is what we're seeing is that technology is basically eliminating the rote parts of the jobs, right? So when I used to program, I remember I would have to write all these XML configs. And nowadays, you just ask GPT-4 and it'll write the XML configs for you. But you want to do something really interesting, then you need to bring your human intelligence to the problem, even if it's just telling the model how to go about it. So you're closer to this than 99.999 % of people in the world. So I want to ask your intuition because I've talked to Sam, obviously, who's involved in this.

20:10And I think you may even be more technical than him and be really close to the research. So, you know, the next 10, 20 years, I mean, are you going to be able to do everything people can do in five or 10 years? Is this like a total change in all of reality by the 2030s? Is there are there actually a lot more steps potentially still like what's going to happen here? Yeah. I mean, the thing about AGI, artificial general intelligence, which is people call AGI, is that it's very clear to me. and I think it's clear to people who are working in the field that progress is going to continue to happen.

20:39What's hard to predict is when you actually are able to do a particular job. When the the there's so many different things that are wrapped up in any particular job that it may turn out that something we didn't even expect was the last thing that that needed to be automated. So I think we're going to have a long time. We may not have like perfectly realistic Westworld robots walking around in like seven or eight years. Robotics? Robotics are probably going to need AI to write that, to figure out how to build robots for us. But I mean, you're going to get... So, that's the other question. I don't know what public or not, so push back on me.

21:15But the thing I've heard from various people is you can probably do GPT 5 and 6 and maybe 7 with people. And then at some point, you're going to need to get to 8 or 9 or whatever. You're going to need GPT itself to do that. If I was running a research group, I'd already be trying to probably figure out how to get this thing to teach itself in different ways. that's probably like there's lots of concepts there. I'd imagine that you're trying to figure out how far are we away from it teaching itself? And is that a thing you guys are working on that you could say? Yeah, I mean, I think that's what everybody's, I think that's sort of the open problem that everybody sees right now is that, you know, the current, as you say, the current GPTs are all just mimicking what humans do.

Read the full transcript

21:48So how do you unlock creativity, right? And if you think about it, it's that same problem we had with Dota or with games where because you have this structured environment of a game, it was able to be creative within the structured environment by trying new things. So how do you give it, how do you give it the ability to be creative within the structured environment of reality and that that we don't know and i and i think the the interesting thing even now is just how you build a structure of reality for the agent to play in that's interesting i mean i imagine you already have lots of versions of 3d structures of reality obviously for robotics to play in but but i guess i guess it's more complicated to have a structure of like human reality i think the most interesting structure we have now is actually the the compiler the interpreter and being able to write code run code, see what happens.

22:29Interesting. So, rather, so, so that's actually a more dynamic place for them to live is like learning how to write code and iterating on what happens with code it writes. Yeah. I spent, I actually spent a lot of time at OpenAI thinking about 3D worlds and, you know, ultimately the problem with the 3D world is it's very limited because you have to create everything interesting that's in there. That's fair. I guess you could create a bunch of like human and emotional actors or something like this for it to interact with, but you're going to create some really dorky robots. They're all just going to be really good coders.

22:56Hey, you know what? You can be all the non-dorkiness. Humans can bring the non-dorkiness. I will not be able to bring that, but normal people will be able to bring that for you. It's almost like you kind of need more human data. You need more data about how people just do normal human things and how they're feeling and how they're thinking, right? Because that's what we're missing. Well, it's unclear. The funny thing is, if you just think about the internet, it is so huge. There's so much. And so there is a lot of data out there about how, you know, if you think about a math problem, like how to approach the math problem, how to iterate on it, how to work with it.

23:32And then the other thing that we've done and that other companies have done in the last couple of years is hire humans to give us data. So, you know, you start by training these models like GP4 on the Internet. And there's this finishing step that you do, which is called reinforcement learning from human feedback or RLHF. Reinforcement learning. It's like padlogs, dogs. You give it a treat if it does the right thing. And human feedback means that you have humans who are doing it. And so they'll have it try to solve a problem. And if it has a good approach to solving a problem, then they'll say, great, you had a great approach.

24:05Whether it got the right answer or not, that turns out to be really helpful. And just do this over and over again. And you don't have a ton of data compared to what you pre-trained on. One thing that's been unintuitive to a lot of us, we haven't caught up on any of this, is just that you'd think it would get better and better over time, but a lot of people seem to experience it getting worse over time, at least their latest interactions. Are they correct that somehow some parts of it seem to be getting worse over time? Or do you have any thoughts on that? We think it's actually getting better over time.

24:35Some of the reports are sort of confounding various different things. So people are just confused about it? I think people are just confused about this. Because even really smart friends of mine who are close to me feel like it's not answering them as well as it was. But maybe they just got expectations too high or got confused on that somehow, huh? Well, I mean, the models are stochastic. So, you know, it's definitely going to be the case. The answer you remember is the one when it got it right. And it doesn't display that flash of brilliance all the time. And what's your intuition that you could say for, I mean, obviously, GPT-3 was just fundamentally different than 2.

25:04And 4 is much better than 3. Like, are we getting to an asymptote here with your work? Is 5 just, like, way better than 4? Like, how are you feeling about this? The way I think about it is that we are still pretty far away from the level of scale that is the human brain. And so there's no reason for there to be an asymptote. I think it just keeps working. When you say level of scale, what's the intuition for that? Like what's the scale of three or four versus five or six? Yeah. So you can count the number of neurons. And, you know, like I said, this is very rough because a computer neuron and a brain neuron are not at all the same thing.

25:36You know, there's probably an order. You know, it takes, you know, 10 or 100 neurons to simulate a neuron in the brain. But if you look at something like GBD-3, you could say, oh, well, that's maybe a lizard number of neurons. And GBD-4 is like maybe a cat. And you still have many orders of magnitude to go before you get to something that is actually the same size as a human. Is that really relevant, though? Because, I mean, you'd say this is human. You'd say this is a blue whale. And it's like a lot more. And blue whales aren't, I assume, and they're not much smarter than us. Yeah, we don't know, actually.

26:03I guess we don't know. They could be doing really cool philosophy in the ocean that we don't know about. But, you know, I think the other thing is that if you think about the data a blue whale is trained on, It's trained on, you know, where are the fish? When does it open its mouth? And, you know, when does it swallow? And, you know, we're training on the human data. So you want something to size the human brain and train on the human data. So you still think it gets a lot better the next three, four or five years. Like we're on a curve right now. I mean, it makes sense to make that for forever to work on AI because it's going to be so much more important in three years.

26:30So that's your intention. I don't see any reason for it to stop. There's no number of years in which it would start to asymptote intuitively to you. Well, look, I mean, I don't think I can put a number of years on it. But I think when you pass human level intelligence, that's where it gets kind of crazy in lots of different ways. You know, maybe maybe learning by mimicking humans no longer works. Maybe there's an asymptote there. Maybe we have to switch to one of these other techniques. Maybe you'll finally feel understood by a creature. The way I sleep at night on this is I assume that there must be like multiple different S's.

27:00Like everyone always assumes there's like an exponential. It just goes to the moon and the world's a singularity. And I always assume there's like this and then another S and then another S. Is it possible there's these different paradigms you have to figure out that could take a while? Or do you think you guys have it all figured out enough to go away? I think there really are. I think these things are fractal, right? And so in some sense, the paradigm is neural networks. But in another sense, well, you need to figure out transformers. And so each of these S-curves, this is a thing that genuinely happens.

27:28Each one is some new way of making sure that you can fit more compute and more data into a larger network. And every time you sort of see a peaking, you need to have some sort of breakthrough. But it's not as fundamental a breakthrough as it was before. Like the difference between neural networks and linear regression was really big. And the difference between neural networks now and neural networks two years ago is just a series of tricks. Yep. So hopefully there's going to be some new paradigm you guys figure out breakthrough. But until then, there probably would be some limit. We'll see. I think we just got to keep trying.

27:58Let's talk a little bit about civil rights protections and AI safety and all that. At Palantir, we were pretty obsessed early on with civil rights. A lot of people think that it's a BS thing or a woke thing. I think to you and me, it's not at all. It's something that's really serious, where you're creating something really powerful, and governments can use it inappropriately and spy on people if you're not careful, and you want to have limits to watch the watcher. I think Palantir was one of the first people to build a civil rights group. We were obsessed with that. Similarly, with open AI, there's all sorts of concerns with AI.

28:27I interviewed Marc Andreessen. You've probably seen his essay. He's very optimistic. I tend to be an American optimist as well around this, but there are obviously concerns. And I have concerns on both sides. I have concerns that you don't want the AI to only have one political view, because that'd be really bad for everyone. You don't know what that view is. And you also, at the same time, you probably don't want the AI to empower really, really hateful people to figure out how to spread hate amongst all sorts of populists and whatnot. That's, you know, a lot of family died in the Holocaust. I probably don't want people spreading all that nonsense everywhere.

28:57Where do you stand on this and how are you guys working on it? Because obviously, you've taken a lot of flack from both sides for things. How do you approach it? Yeah, I think it is the kind of thing that is very ideological, as you say. But when I think about it, I think it's also just very practical. Because think about ChatGPT. Think about the intelligent agent you want to create. Fundamentally, it has read the entire internet. It has definitely read a lot of rants about the Jews, right? Yes. And it's read a lot of bad things about women, etc. But when you have it talking to you, you do not want it to be ranting about the Jews when you're trying to write up a program.

29:32It's pretty funny, even as a Jew. But yes, that would not be good. Well, this is the kind of thing the original GPT-3 would do all the time. Really? Not necessarily about the Jews, but you're trying to solve a problem. And it would just go off about something. It's like a Reddit comment. It's like, here's the answer, you goddamn.

29:51That's not good. There was this famous rant we had posted on our thing about how recycling was terrible. Did you feed it my emails from Fysi? Maybe so, yeah. I mean, it's not economic to recycle most things, Bob. Let's be honest. It clearly learned that from you. But fundamentally, you have these agents. You just want them to be professional. And so I think a lot of near-term safety, a lot of these discussions about bias, they really are about the fact that the model has learned all of these things, which it's learned a lot of true facts. It's learned a lot of things that are inaccurate. But you just want it to be professional like a coworker, fundamentally.

30:26It's not about being politically correct. But it's also not about making sure that it's not necessarily the most important thing that it always, half the time, talks about male nurses and half the time talks about female nurses. You want it to be able to interact with you like a professional that you work with. Also like a normal person without being too careful. Because if it's like a professional who's always really, really extremely one way, that'd be kind of weird too, I guess. But I guess you have to figure it out. When we initially came out with ChatGBT, it would always say, well, as an AI language model, I can't do this.

30:55Very careful not to get to hedge. Yeah, we've relaxed that a little bit. We've gotten rid of that. We heard people don't like that. Yeah, it was a little too much. It was a little too prissy. I had dinner with Brian Armstrong and friends in LA this week, and he's famous at Coinbase for saying no politics at work, which some people didn't like. But I guess that's kind of what the goal is with the AI, is to not bring its politics on either side to answer you, I guess. Yeah, we don't want the AI to have politics. You know, you're the human. You should have politics. You should be able to ask it things that are political, though, if you want.

31:24That's right. You can ask it, you know, who is the 45th president of the United States. You could probably ask it to give advice to Biden on his reelection. Or you could say, what's the leftist argument for this? What's the argument on the right for this to understand? Or no, you don't want to be able to ask that? I think you can do that. We should go give it a try. What you don't want is you don't want to ask it a question and then have it bring, you know, a particular point of view. You want it to be able to say, well, this is one point of view, and here's another point of view. I think this is totally reasonable.

31:50I have to ask this, and you can not answer if you want, but I'm just curious, because a lot of friends are going to ask. There must be something dynamic and complicated about how to hold off areas that are not okay to talk about and that are okay to talk about. I'll just give you an example. For a while, if you were trying to get to praise President Trump, it wouldn't do it, but then it would write a poem about President Biden. Now, is that just a coincidence that it would do both? or is it pretty complicated for it to figure out what it is or it is not supposed to do? Because it seems like sometimes it'd be asymmetric.

32:16Yeah. I mean, honestly, Joe, this stuff is very mysterious. We sometimes learn about these things from Twitter, too. And when we see people on Twitter posting these examples, we're like, oh. We go back to the team and we're like, okay, guys, look. What do we do? It seems like it's biased again. Let's go get new human data, try to figure out what it should do, and go from there. And there's not a little person behind the scenes trying to make it biased. is just a very hard problem. Well, there's actually thousands of AI trainers behind the scenes. We give them prompts, and they have to score the prompts.

32:50And so, I think for better or worse, they bring their biases to it. But we try to make it so the model itself is the accumulation of those things and ends up not biased. Very interesting. And in terms of research projects going on, if anything you could talk about, is there anything you're really excited about for the next few years that's going to change things? Is there any kind of new use cases that you're really excited about? Well, I think the thing that we have done that I am most excited about right now is something called Code Interpreter, which is launched in beta right now. And it's exactly what we were talking about.

33:18It's the ability for the model to write code and then run that code in a sandboxed environment. And so you can do really cool things with it. You can upload a spreadsheet and you can have it analyze the spreadsheet for you. You can have it make an animated GIF of data visualization. And we're not specifically telling it like, OK, let's make the spreadsheet model. Let's give it this. It's just learn these things organically from everything on the internet. Wow. And we're teaching it to use the tools. But what it can do with the tools is an emergent property. I mean, it's going to be really fascinating if it does inappropriate things emergently with GIFs.

34:00It must at some point. You should try this out on ChatGPT. See how it works. And you must have something internally where it's more jailbroken. It just gets to do what it wants. Then you have to learn how to stop it from doing the bad things, I'd assume. Well, so the pre-trained model that just mimics humans is very jailbroken. And then, yeah, we just add the data to it over time. One thing we talk a lot about is agent models and lots of agents that go around doing things for you. And it seems like you could build crazy things with agent models. Have you played with this a lot yourself in terms of different contexts?

34:30I think the agent stuff is super interesting. And I think there's a lot of cool ideas that are going on right now. But I like to think about it when you take it to the asymptote, actually. Imagine you have 1 ,000 agents. And they're not just interacting with you. They're interacting with each other. They're interacting with all the humans out there. They're generating their own data. They're able to look at their own data. I mean, when you push this to the level where AI is at the human level, I mean, they're basically going to be creating their own civilization. Hopefully jointly with us, right?

34:59I used to play role-playing games as a kid. I don't have time for those as much anymore. There's this game, Dragon Warrior 4, was the best ever role-playing game for Nintendo. It was a great game. I love that game. It's such a big, involved world. Imagine with kids today, but with your guys' technology, making these people more interesting. I think it'd be so cool. Yeah, I think I've seen video games that are out there doing that. The thing that I love about my kids is that kids really have a lot, at least my kids, have a lot of questions. They're always asking me, Dad, can you explain this? Mom, can you explain this?

35:28And at this point, I've just gotten to the point, after three or four of these questions, I talk to my kids. But then I'm like, OK, let's just ask ChatGPT. What if the answer is slightly wrong, though, or inappropriately? I guess it's probably good for them anyway. They can learn how to iterate. You know, they have an uncle. They're not completely inappropriate uncle already. Mind you, too. That's fair. It's always an inappropriate uncle. I guess it's not too dangerous. That's funny. I guess kids have to learn how to use this. Do you think kids should be able to use this for their homework and everything?

35:58because it's going to be how they work in the future anyway, right? Yeah, I don't know. I think I don't want my kids using it for their homework. But what I want is I want teachers figuring out how to make homework that requires you to use AI. Yeah, which could be pretty hard, though, because as AI gets better, it's just going to do a lot of it for you. Well, I mean, I think, you know, you can... But also, like, that's what you're going to do in your job. Yeah, that's my point. But it might get to the point where it's like you could do a lot of things for you. So I guess, but that's good. You just learn how to be really efficient, and then you can go hang out and swim in the pool.

36:26Well, you know, and it doesn't always do it right. Right. And so, you know, you have to be able to figure out. I think this is probably a good way to rewrite your homework is like have the AI do something and then you, you know, grade the AI, figure out whether it does it right or not. And then tell it, have it do a second pass. You might have asked it wrong. You might have misunderstood what it's doing. You might have misunderstood it. That's true. So tell us about the launch of chat GPT, Bob. There's all sorts of competition going on there. What was the scenario? Yeah. Well, the funny thing is when we launched it, there was no competition.

36:52So we had launched GPT-3 two years prior. And then we launched a better model, which is still OK, called GPT-3.5. And at the time, the idea was we'd put these in an API. They weren't really ready for humans to talk to them yet. It'll finish your sentence, but it wouldn't have a conversation with you. And we'd had these models out there for a while. And we had already then secretly trained GPT-4. So this was the first half of last year. And GPT-4 was amazing. We knew that if we could figure out what to do with GPT-4, it would change everything. And so the whole company was focused on GPT-4. What could we do with GPT-4?

37:28And one guy, John Schulman, who ran our reinforcement learning team, said, you know what? What if we just make the models conversational? So we took one of the old models, GPT-3.5, trained it to be able to have a conversation. And we all thought the right ultimate path for the GPT models was for it to be an assistant, to help you do things. but we were like these models are clearly not good enough we've had them out there for forever and john was like look it's not perfect i know gpd4 is going to be the answer to everything but let's just let's just put it out on the internet and you know if we get 10 000 people who use it they'll at least tell us where it's bad and then we can iterate and we can make it better and we sort of went back and forth on whether we should do it seemed like a lot of work and eventually we're like okay we're going to do this and i think we had like a week to launch this and again it was a side project.

38:15We called it a low key research preview. We didn't do any press on it and it just blew up and suddenly everybody started using it. And, you know, again, the whole company was focused on something else. We were like GPT-4, that's the thing. Um, and so we spent six months just sort of chasing it, trying to make it actually work. Wow. And you cause all of us now to have to focus on AI investing for everything we do. Yeah. If only we had known, if only we had known. It's a, it's, It's amazing. That does bring me to question, so many people are trying to invest in AI now. It's like a bubble time again.

38:45There's$100 million rounds here. There's$200 million rounds. Everyone will be two pitches us. There's the AI part of the pitch. All of our old SaaS companies, one of them was a legal tech company, you probably don't even know this, that's partnered with OpenAI and suddenly sold for eight times the last round valuation to Thomson Reuters because it was all of a sudden doing really useful things for paralegals. It's changed a lot of things. In terms of other people doing AI, what else is useful that people are working on that you're excited that you're seeing people do? because this is outside of OpenAI now.

39:11Is there infrastructure you like? Is there a need for infrastructure outside of OpenAI? How do you see this? Yeah, I think I'm always a little worried about the infrastructure work because I think you're solving the problems as they exist today. And when GPT-4.5 and GPT-5 come out, they're going to have fundamentally different use cases and fundamentally different infrastructure. What I like to see is people who are using AI to solve something that wasn't possible with AI. You like the application layers. Yeah. And what I think about is AI now, it's like having an infinite number of interns with very short attention spans.

39:43And so anything you could have solved with interns, you know, and GPT-3 was sort of like a high school student. GPT-3.5 is maybe a college freshman. GPT-4 is maybe a college junior. GPT-5 is going to be something else. I'll take your intern. And so, you know, if you have these interns, what can you do with them? And what new business does that allow you to create? Like, that's the thing I'm really excited about. Lots of free interns. Yeah. That's cool. The app layer is where I've built a ton of things. I'm getting pitched tons of infrastructure. Tons of my friends are getting billions of dollars of loans to get A100s and H100s and doing crazy amounts of hardware infrastructure.

40:21Is that going to be needed still? I assume that's going to be needed still. How do you think about that side? Yeah, I think it's really hard to say. A lot of people have tried building new chips. I think what everybody has seen so far is that NVIDIA has just been able to continually do better and better and better. because they have the big market and a lot of capital to throw at it. So I think it's pretty hard to bet against NVIDIA. On the other hand, a small chance of a really big thing is also something worthwhile. Yeah. When I was at Stanford in computer science and graphics, I was obsessed with NVIDIA at the time.

40:50That's where I wanted to work. And then I ended up with PayPal and Peter. You probably did okay. Probably did fine. But it's pretty funny how the company I was really into ended up pivoting and doing this. But I want to ask, we started at American Optimist to push back on a lot of the fear and criticism in our country. There's a lot of people who are really cynical, a lot of AI doomsayers. What's the best reason to be optimistic about AI? What could the world look like if we get it right? Well, I mean, I think the amazing thing about AI is if it's really at the level of human intelligence, then imagine having coworkers, friends, people that you can talk to that can help you do anything you want to do, that can help you go out and achieve your dreams.

41:31and a world where we're fundamentally not limited by the labor constraint. There's capital, there's labor, there's ideas. And if we have AI out there generating ideas, AI out there helping with human labor, we can have a much higher standard of living for everyone. And if a young person wants to work on AI, what should they be doing right now? Well, if they're a really great programmer, they should come work at OpenAI. But I think actually this is the best time ever to start a company and go after that application layer and figure out something fundamentally new that you can do with AI or skate to where the puck is going to be and think about what you could do with AI that's 10 times or 100 times smarter than what it is now.

42:11And I want to ask, you're at the center of the AI universe. Are there other technologies or breakthroughs you're excited about right now? Or is AI kind of like the dominant factor that's by far the most important thing? Well, I want to see what happens with superconductors, for sure. You know, the interesting thing about AI is, in the end, it's going to be limited by things like power. and chips. And if you're trying to build these really massive clusters in 10 years or whatever. So fusion energy, I think, is pretty interesting. I was an early investor in Oclo, which Sam seems to be SPACing as well.

42:41Oh, nice. Small nuclear fission reactors. And I think fusion would be very good, too. Though hopefully your AI will figure out the fusion for us, right? Maybe that's one of your research projects you can do for us later. That's what you want. That's what you want. You want AI to be pushing the technological frontier. TPTA can solve fusion, maybe? Maybe. Let's give it a try. All right, Bob, thanks for joining us. Good to see you, Jim.

From the publisher

Bob McGrew is at the epicenter of the AI revolution.  As the VP of Research at OpenAI, he's instrumental in breakthroughs that are reshaping the world, from building GPT models and launching ChatGPT to overseeing the Dall-E project.  How was GPT-4 trained and what will GPT-5 look like? Why does ChatGPT respond with certain biases and how do they correct for that? What breakthrough led Bob to believe AGI could be achievable?

Bob and I were in Phi Psi together at Stanford and both interned at PayPal in its early days.  We hired Bob as the second engineer at Palantir, where he built and shipped the first products for the intelligence community and went on to lead engineering and help run the company.  Bob joined OpenAI part-time in 2016 and full-time in 2017, where he's been at the forefront of AI innovation.  In this episode, we discuss the early days of Palantir, how he knew AI's moment had arrived, and the most important research projects underway at OpenAI.  We also look ahead to what GPT-8 could unlock for humanity and what's needed for Large Language Models to move beyond mimicking humans to higher forms of intelligence and creativity.



This is a public episode. If you would like to discuss this with other subscribers or get access to bonus episodes, visit blog.joelonsdale.com

More from Joe Lonsdale: American Optimist

All 113 episodes
Ep 66: Bob McGrew — the Superstar Palantir Alum Leading OpenAI's Transformative Research Projects Joe Lonsdale: American Optimist · 43 min
Listen in VO