Gavin Baker - AI, Semiconductors, and the Robotic Frontier - [Invest Like the Best, EP.385]

27 Aug 2024 · 1 h 31 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Episode Summary: Gavin Baker - AI, Semiconductors, and the Robotic Frontier

Podcast Information

  • Podcast Title: Invest Like the Best
  • Host: Patrick O'Shaughnessy
  • Guest: Gavin Baker, Managing Partner and CIO of Atreides Management
  • Episode: EP.385
  • Released: Not specified
  • Episode Description: A deep dive into AI, semiconductors, and the future of robotics with Gavin Baker, who shares profound insights on technology, investment strategies, and industry evolution.

Key Topics Discussed

  1. Tech Competition and The Magnificent Seven
  2. Overview: Discussion on the competition among major tech companies like Google, Amazon, Microsoft, and their overlapping interests in AI and cloud computing.
  3. Competitive Landscape:
  4. Previously distinct competitive lanes are converging in the realm of generative AI.
  5. The existential threats perceived by tech giants in the race to create advanced AI systems.
  1. Generative AI and Scaling Laws
  2. Generative AI: Explores the implications of generative AI technologies and the belief in ongoing improvements due to scaling laws.
  3. Investment Implications: Discussion on the potential ROI in AI technologies and the challenges faced in AI infrastructure development.
  1. Challenges in AI Infrastructure
  2. Efficiency in AI Models: Analysis of the high marginal costs associated with AI and the importance of effective infrastructure for success.
  3. Synthetic Data: The role of synthetic data in AI training and the ongoing quest for improved data sources.
  1. Future of Robotics
  2. AI Integration: Insights into the integration of AI with robotics and how this could reshape the labor market.
  3. Tesla's FSD: Predictions on how Tesla’s autonomous driving capabilities could evolve and the significance of real-world data in training AI.
  1. Leadership in Tech Giants
  2. Influential Leaders: Exploration of leadership qualities in exceptional CEOs like Elon Musk and Jensen Huang, emphasizing their problem-solving approach and mission orientation.
  3. Impact of Leaders on Innovation: How these leaders foster an environment of innovation, attracting top talent and focusing on critical problems.

Insights and Takeaways

  • Investment Strategies:
  • Focus on investing in companies that improve the AI infrastructure equation (e.g., SFU, MFU, checkpointing).
  • The importance of understanding the operational challenges in AI and semiconductors for future investments.
  • Evolution of Investing:
  • The impact of LLMs (Large Language Models) on the future of investment research and decision-making.
  • Emphasis on combining domain knowledge with AI tools to enhance investment strategies.
  • AI's Societal Impact:
  • The potential for AI and robotics to transform the labor market.
  • The necessity for diverse AI models to ensure a plurality of perspectives and avoid a dystopian future driven by a single dominant AI model.

Conclusion Gavin Baker provides an enlightening perspective on the rapid evolution of AI, semiconductors, and robotics. His insights into the competitive landscape of tech giants, the intricacies of AI infrastructure, and the characteristics of successful leadership offer valuable lessons for investors and industry observers alike. The conversation underscores the importance of adaptability in investment strategies and the potential socio-economic impacts of emerging technologies.

For more episodes and detailed notes, visit [Colossus](https://www.joincolossus.com).

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Something I speak about frequently on Invest like the best is the idea of life's work. A more fun way to think about it is that I'm looking for maniacs on a mission. This is the basis for our investment firm, Positive Sum, and it's the reason why I'm so enthusiastic about our presenting sponsor, Ramp. Not only are the founders, Kareem and Eric, life's work level founders, certainly maniacs on a mission. They have created a product that is effectively an unlock for founders and finance team to do more of their life's work by streamlining financial operations, saving everyone their most precious resource, time.

0:30Ramp has built a command and control system for corporate cards and expense management. You can issue cards, manage approvals, make vendor payments of all kinds, and even automate closing your books all in one place. Speaking from my own experience using Ramp for my business, the product is wildly intuitive, simplistic, and makes life so much easier that you'll feel bad for any company who hasn't yet made the switch. The Ramp team is relentless, and the product continues to evolve to save you time that you would never have dreamed of getting back. To me, there is nothing more interesting than technologies that reduce friction for other entrepreneurs to be able to build the thing that they want to.

1:05So much attention has gone to cloud computing, APIs, and other ways of making life easy for founders. What Ramp has done and is doing is build yet another set of tools in this category. To get started, go to ramp.com. Cards issued by Celtic Bank and Sutton Bank, member FDIC. Terms and conditions apply.

1:29Hello and welcome, everyone. I'm Patrick O'Shaughnessy, and this is Invest Like the Best. This show is an open-ended exploration of markets, ideas, stories, and strategies that will help you better invest both your time and your money. Invest Like the Best is part of the Colossus family of podcasts, and you can access all our podcasts, including edited transcripts, show notes, and other resources to keep learning at joincolossus.com. Patrick O'Shaughnessy is the CEO of Positive Sum. All opinions expressed by Patrick and podcast guests are solely their own opinions and do not reflect the opinion of Positive Sum.

2:05This podcast is for informational purposes only and should not be relied upon as a basis for investment decisions. Clients of Positive Sum may maintain positions in the securities discussed in this podcast. To learn more, visit PSUM.VC. My guest this week is Gavin Baker. Gavin is the managing partner and CIO of Atreides Management, and he has been on the show many times before. He is one of my favorite investors to talk to, and this may be my favorite conversation with him. Gavin first started covering NVIDIA as an investor at the turn of the century, making him the perfect guest to discuss all things AI and investing.

2:44There's so much detail in this conversation, and I'm incredibly grateful to Gavin for sharing his wisdom with us again. Please enjoy this fantastic conversation with Gavin Baker.

3:197 survived the shootout. We've got a new MAG7 today, and we could probably spend the whole time talking about them. We won't, but I thought it would be a fun opening moment just to hear you riff on why these massive companies might actually be in some form of business shootout now with a lot of what's happening in the world of technology. So I think these companies were all in their own discreet swim lanes for a long time, competitive swim lanes. The only place where they really overlapped was cloud computing, where you had Google, Amazon, and Microsoft all competing, but that was a very stable oligopoly.

3:59Google cut prices aggressively, something like 2014, 2015, Amazon matched. And that actually materially impaired Amazon's revenue growth and fed into all of these fears that cloud computing was going to be a commodity, which was a real fear, hotly debated topic, looks deeply ridiculous now. I think that set the stage for like, hey, we're effectively going to agree between the three of us on some markup on cost, but then try to differentiate in other ways. Amazon had e-commerce. Facebook had advertising that was higher up the funnel from Google. Google had search. Apple obviously had the device in the OS and the App Store.

4:38Minor competition from Android. Netflix is doing streaming video. And again, a little bit of competition with Google there, maybe. But they're pretty distinct competitive sets. And then Microsoft, while they had cloud, they also had this massive enterprise software business really focused on productivity. And just with Gen AI, with LLMs to the name, I mean, Gen AI means generative AI, but the G in GPT means general purpose. It's such a general technology that they're all of a sudden in the same swim lane and they all feel like it's existential. Mark Zuckerberg, Satya, and Sundar just told you in different ways, we are not even thinking about ROI.

5:22And the reason they said that is because the people who actually control these companies, you know, the founders, there's either super voting stock or significant influence in the case of Microsoft, believe they're in a race to create a digital god. And if you create that first digital god, we can debate whether it's tens of trillions or hundreds of trillions in value. And we can debate whether or not that's ridiculous. But that is what they believe. And they believe that if they lose that race, losing the race is an existential threat to the company. So Larry Page has evidently said internally at Google many times, I am willing to go bankrupt rather than lose this race.

6:06So everybody's really focused on this ROI equation. But the people making the decisions are not because they so strongly believe that scaling laws will continue. And there's a big debate over whether these emergent properties are just in context learning, et cetera, et cetera. But they believe scaling laws are going to continue. The models are going to get better and more capable, better at reasoning. And because they have that belief, they're going to spend until I think there is irrefutable evidence that scaling laws are slowing. And the only way you get irrefutable evidence, why has progress slowed down since GP24 came out?

6:48Because there hasn't been a new generation of NVIDIA GPUs. It's what you need to enable that next real step function change in capability. It is actually going to be interesting. So everyone, the Blackwell delay plays into this. It's really, really hard to create what's called a coherent trading cluster of tens of thousands of GPUs and coherent just means each GPU, we could say, knows what the other is thinking. Technically, it's more like they have a shared memory space and that cluster has to be coherent to train. And the biggest coherent cluster in the world until very recently was 32 ,000.

7:22So you had 32 ,000 H100s, probably only using like at most 15, 16 ,000 of those H100s at once because of efficiency problems. But XAI decided they were going to build 100 ,000 GPU cluster and I would say with Elon's unique physical engineering mind, which we've seen play out at SpaceX and Tesla, and now at Neuralink, where he figured out how to miniaturize everything along with really capable teams. I think he re-architected the XAI data center from first principles in Memphis. It's very different than other data centers. And because of that, they're able to get enough density that they could effectively make a 100 ,000 GPU cluster coherent, even with Hopper, without next-generation networking technologies from NVIDIA, Broadcom, and others.

8:12And they started trading on that. And that means that I think we're going to see, probably, if scaling was old, the first GPT four-and-a-half class model sometime, I don't know, in the next six, nine months. I don't know the timing. And then after that, you'll have Blackwell. and then you'll go up to 300 ,000 GPUs in a cluster because you have next generation networking technologies that make that easier. And that is going to be a massive step function change. And then that's where you can get maybe a GPT-5, five and a half, six class model. And then the reason it's existential, it's just like if GIN AI is so generalizable and you have this ASI, you're like, what's the value of content in a world where AI can make content that's infinitely better than any human?

8:57What's the value of Netflix in a world where I could say, I want to watch a mashup of Star Trek and Star Wars tonight? Search might just go away and be replaced by agents. It's important to have a lot of humility about this. Chase Coleman, who runs Tiger, who's an exceptional investor, had a really interesting statistic, which was ChatGPT came out in 2022 and was to AI as Netscape Navigator, was to the internet in 1994. only less than 1 % of current global internet market cap was founded, founded in the two years after Netscape Navigator came out. And nobody could imagine the companies that were going to be founded.

9:39So it's just the biggest companies were founded many years later or several years later, five or six years later. So just we're very early, important to have a lot of humility. But as long as scaling laws continue, and the only way they can be disproven is if you have a new generation GPU that has the traditional, whatever it is, 3 to 5x performance improvement, and then you get better networking such that you can link 3 to 5x more of them together, and then you get that 10x, the order of magnitude improvement in compute, and you don't see a massive improvement in model quality, then scaling laws will stop and that'll be a catastrophe for the entire data center infrastructure.

10:21But the people close to this all believe that scaly laws are going to continue. Last night, my daughter on the way home from dinner said to me, hey, can you ask ChatGPT some question? And I said, do you have ChatGPT on your iPad? And she said, no, not yet. You have to get it for me. I said, well, just Google it. She goes, what's Google? Wow. She goes, Google's not a real thing. That's what she said. They just use ChatGPT for everything. My son uses it for everything. She's eight, he's 10. It's just fascinating to think about, wait, what? What did you just say? And they use it all day. They use Dolly all day, every day.

10:56Search in their mind is ChatGPT, which is just a fascinating thing for kids that could use both if they wanted to. And rather than a search engine, they just want the answer engine. And I just find that completely fascinating. Absolutely. What does it mean for humanity? the open internet has been really, really good. And yes, we have these walls, content gardens now, but there's still like a very robust, vibrant open internet where when you Google, you do get divergent opinions, generally, even though we're all in these well-known filter bubbles. But if you just get one answer, I actually think we're at an important moment.

11:33It's very rare that to me, investors can really, really contribute to the world. We contribute in aggregate, in an indirect way, in a massive way, because we fund all these new technologies that are solving cancer, indirectly funding AI. But I think it's rare that you can have a more direct impact as opposed to a diffuse impact. I think it is supremely important for humans that we do not end up in a world where there is just one dominant model. That is the most dystopian future I can imagine. Because that model then, if kids all over the world follow your kids' behavior, whatever values that model has will be imbued to the rest of humanity.

12:21And I think we are at a time in history where, for a lot of reasons, the idea of objective truth is under attack. For a lot of younger people, their feelings are facts. The idea of science and objective truth is really under attack. And I think it's really important to have AIs that are dedicated to the idea of objective truth, no matter how unpopular that may be. And I think the best way to accomplish that is to have them compete. You only have one dominant AI. That's very scary and forget 1984. Hello, who knows what kind of future. But whereas if you have three, four, five of these, that's very different.

13:01They'll compete. They'll have different value systems. And I think we can as investors by lowering and improving the cost of AI infrastructure, whether it's breakthrough networking, storage, technology, software that improves the utilization rate of GPUs and directly funding, I think, some of these labs that are competing with Google and Meta. I think it's very important for the world. Can we talk a little bit about the world of data centers and semiconductors? Because you are one of the most seasoned investors in both of these spaces. I think you started your career covering semis 25 years ago or something, and you just know more about it than just about anyone else I've talked to.

13:46And I often find myself asking you or calling you with a question about these two areas. And most of the attention has been on the models and less on the world of semiconductors. Everyone knows NVIDIA, obviously, but they just sort of assume like, yeah, it makes these amazing chips that power the rest of it. And there's just so much else that's going on. You've hinted at it with like the new clusters that are being built. But just give us like a stated from your perspective, how this has evolved and what aspects of it are most interesting to you, whether that's the data center or the individual chip or the systems or however you want to approach it.

14:16I just think you have probably the most interesting and valuable perspective of anyone I know on these two topics and have been thinking about them for a long time. And now it's like the main event. It's for sure the main event. I started covering NVIDIA in January of the year 2000. I would say watching the early days of NVIDIA as a public company, Tesla as a public company, and being a substantial investor in both have been by far the greatest privileges of my career and the most exciting thing I've ever done as an investor. And Elon and Jitson are for sure the two best CEOs I've ever seen with Lisa Su right behind them at AMD.

14:58And I just say that she took AMD from company five times levered, literally four years behind Intel to essentially utterly dominating Intel in every way. I think they had 20 days of cash when she took over. Everybody talks about Satya, and he's very impressive. He took over a monopoly that had been extremely poorly run with recurring revenue and high margins. Anyways, I have been at Simi's for a long time. It is my first love. Where are we? So first, I think it's been touched on a lot of your podcasts. The number one thing we have all heard as tech investors, one reason tech has been so amazing for 30 or 40 years, is that software has zero marginal costs.

15:35These companies all have extremely high gross margins, recurring revenues. And AI is the exact opposite. AI has extremely high marginal costs because scaling laws literally mean that the only way you get an improvement in quality is by spending a lot more. Full stop. That is what scaling laws mean. If you believe in scaling laws, you believe AI will have very high marginal costs. Now, because of a variety of things that we're going to talk about, those marginal costs go down really, really quickly. But they're still really, really high. at the leading edge for models, especially for trading and for inference, but much less so for inference.

16:15Inference and trading are two totally different markets. So what this means that it has really high marginal cost is that infrastructure, efficiency, and excellence, I think, is going to emerge as the single most important success factor, particularly for the model companies themselves. And this has been measured today in something called MFU, model flops utilization, and that generally runs around 35 to 40%. And that's literally the percentage of compute, theoretical compute flops that you're actually applying to trading. So of the theoretical flops, most companies that have been published, and people stopped publishing this because it's so competitive, but the technical papers for GPT-3, I think what was called Google's Lambda, and NVIDIA's Megatron all showed their MFU.

17:07And it was between high 20s and high 30s for all of them. Google had the highest. If you have a higher MFU, it means that you can choose for the same amount of money you were spending. You have the same amount of GPUs and the same amount of power, presumably. You can choose between faster time to market. If you run a 50 % MFU and your competitor's running 40 for an equivalent amount of trading flops, you can be in market 25 % faster. you can choose between better quality you might just do the trading run run as long as possible or you can make the model lower cost in a variety of ways most of which relate to quantization and we could get into that but it's very technical it's almost like if you have a 25 higher quality model because you have a higher mfu and you build in quantization then you can actually get a almost a 50 reduction in inference cost because if you can quantize one level lower than competitors it's a profound advantage.

18:02And so I think MFU, if I were to pick one metric to evaluating lab success, because now there's been five GPD4 class models from it's Google, OpenAI, Anthropic, XAI, and Meta. So there's five GPD. And does Mistral have one that good? Maybe right under, arguably just as impressive because it gets really good results in all the evals with a much smaller parameter count. By the way, these models are commodities today, but I am suspicious once we get to scaling laws continue and GPT-7 or 8 literally cost$500 billion to trade. I don't think they're going to stay commodities. Scale is the most powerful barrier to entry, and that's a lot of scale.

18:46So MFU is the most important metric because it gives you all of these advantages and ways to differentiate yourself amongst five people, six people who've trained these GPT-4 class models. But I think there's something better than MFU, and I actually came up with it this morning, and it adds to MFU and decomposes MFU. I would think of it as maybe like a unified AI efficiency equation. So MFU, I would decompose into two things. The first thing is something called MAMF, Maximum Achievable Matrix Multiplication Flops. MAMF, MatMol. And what this measures is software efficiency. And this goes to CUDA.

19:29And this guy Stan Bexman came up with it on X. What he did is each chip has a theoretical maximum performance, which is just easily calculatable, you know, flops. And then he looked at what can you get in practice? NVIDIA GPUs run it, per his testing, 83%. And I'm sure that there are some labs that are running them at closer to 90%. But this is just a single GPU. One reason it took AMD so long to break into this market, beyond the fact that it almost always in semiconductors, it almost always takes you until your third generation chip to really, really hit it. That's what it took with Google, with TPUs.

20:08You know, Instinct in my 300 is a good chip. because their RockM open source software was terrible, they didn't have the internal capabilities to really improve it. No one in the community took time to improve it. So the Instinct MI 300 first came out, it ran at something like 25 or 30 % mammoth. Now per stance tests, it's up to 60%, but because it has more flops, it's actually slightly ahead of NVIDIA GPUs. And this was actually really, I thought it was brilliant. and as a way of conceptualizing and quantifying the CUDA advantage. They can run an 83, and AMD is running a 25, and it answers the question, the MI250 was actually a pretty good chip.

20:50Nobody used it. Why? Its mammoth was probably 10%. So first there's mammoth, and that is how good is the software for your chip, and then how well do you, as someone who's using that software, optimize it? Netflix used to say that we know how to use AWS more efficiently than Amazon, then a lot of people have said some of these labs know how to use CUDA more efficiently than NVIDIA. That's a big secret sauce. If they're running at 95, right there, that's 10%. If NVIDIA GPUs are running at 83 % efficiency on a per GPU basis, why are we down at 35%, 40 % for MFU. And the reason is something I would call SFU, system flops efficiency.

21:36And this really captures networking, storage, and memory. In each one of those, and we could decompose SFU into each of those, but I think SFU is like a helpful way to think about it. That's not all. So you need to multiply MAMF times SFU, system flops efficiency. And then you need to multiply that by the percentage of time that is spent in a checkpoint per trading run. And these things all compound. And they compound out to some wild things, just in case, should I explain checkpointing or no? Because, as I referenced earlier, a trading cluster has to be coherent to function. And that means each GPU needs to be aware of what every other GPU is thinking.

22:22if any one GPU fails, you lose everything from the last time you saved the model, which is called the checkpoint. The GPUs fail all the time. Not just GPUs, but GPUs melt. If you read the Lama 3 technical paper, I mean, it's like the list of reasons that GPUs fail is like astonishing. An optical link goes down, a switch goes down. There are so many points of failure in the chain of storage, networking, and memory that feeds each GPU that there are innumerable ways for them to fail. The GPU without storage, memory, and networking is worthless. And if you want to really, really understand this, something I highly recommend to everyone, build a gaming PC.

23:06because every computer, whether it's the iPhone, whether it's a laptop, whether it's a data center, has the same, I would call, four fundamental elements. Three are memory, storage, and compute, and then the networking that connects all those things. And when you assemble a gaming PC, you literally have to plug - Yeah, the PCI Express into the GPU and then into the CPU, and then you have to connect, slot into DRAM so that it'll connect. and then the DRAM has to connect to the flash storage and DRAM is obviously memory. I highly recommend this. But I think the way to think of a data center or any computer is just imagine that it's a restaurant.

23:50In this restaurant, the primary unit of compute and in an AI server, it is the GPU, is the head chef. A head chef can do nothing without food and ingredients and utensils. I would conceptualize storage as like the delivery truck that brings food to the restaurant. And then the way storage connects to the rest of the server is today always over PCI Express, and that's a networking technology. And that is literally the guy who moves the food from the food truck into the restaurant's refrigerator. And then we'll call the restaurant's refrigerator, this is an imperfect analogy, maybe the memory. and then you have to move the data from the memory, actually in this case, into the CPUs, into the big pool of DRAM.

24:41The CPU can do its job and then you move it to the GPU's memory and then the GPU can finally do its job. Maybe the GPU memory is like the stove and what flows between the stove and the sous-chef is the connection between the CPU and the GPU. And we can go into this, But just unless that chef has a stove, cooking utensils, and food, he can do nothing. And the big problem in the data center is over the last five years in particular, these numbers are going to be directionally accurate. GPUs have gotten 50 times faster, and the rest of the data center has only gotten four to five times faster. And that is why MFU is so low, because the GPU is sitting around waiting for all those things to do their job, and it's doing nothing most of the time.

25:27So I do think it is sensible, really sensible to invest in next generation networking, storage and memory technologies, particularly in networking. Because if we're going to get to a million GPU cluster, that'll be a million Rubens, which is the generation after Blackwell, we're going to need profound breakthroughs in every step of that thing. Every single thing, all new stoves, all new refrigerators, robotic sous-shafts, new utensils, everything. Otherwise, it's going to be wasted and MFU is going to be 3 % to 5%. But anyways, coming back to checkpointing, because there's so many points of failure in that chain, the head chef in data centers, being the GPU, goes up in flames a lot.

26:17They literally melt. And then every other component breaks a lot. the cluster is 32 ,000 of these metaphorical kitchens, where each step, if it fails, brings down the whole cluster. So because of this, people checkpoint frequently, and that means save the model. Now, if you, through better networking topologies or better cooling technologies such as GPEs don't meld, if you have a lower failure rate, if you have a more reliable cluster, you need to checkpoint less. So we've had Mammoth, SFU, then checkpointing frequency. And that gives us, and like let's just say there's one company that runs at 90 % MAMF, and then it runs at 50 % SFU.

27:02Now you're at 45 % utilization, and they have to checkpoint, I'm just going to make something up a third of the time. You're down to 30 % MFU. If you have a different company that can run at close to 100 % MAMF, because they're CUDA masters. And they run at, let's call it, 60 % SFU. Now you're at 60, the other, you're at 45. So that's already a massive difference. It's a 33 % difference. And then if you have to checkpoint only 10 % of the time, you're at 54%, your competitor's down at 30. That's not even all of it. The last thing that we need to multiply it by is PUE, which is power utilization efficiency.

Read the full transcript

27:44The power has a cost, and it's going vertical for these big clusters. There's only basically three places in the United States today where you can get a gigawatt of power to a single data center that's reliable enough, i.e. that's nuclear. I think it's like eight cents per kilowatt hour on average in the United States. What they're going to be charging for that gigawatt is like 10x that. But your PUE then really matters because your cost of electricity really matters. So things you do to optimize SFU might actually increase your PUE in a really negative way. And so I think if you're, I just thought of this morning, maybe it's super obvious that every lab already is doing this.

28:30But I just think if you leak it all into an equation, and then you do dollars that it takes to get it, you can really make trade-offs. oh, wow, if I spend 2x more on networking, I get much higher SFU, which is great, but then I have a much higher PUE, which is bad. And what we ultimately want is not just actual exaflops per second, but we want exaflops per second per dollar of capex per watt of electricity. and the equation I just described, which I'm going to write out at least for myself, captures all of that. And everybody's going to be making different design decisions and data center architecture and these design decisions you make are going to be immensely important.

29:23And I think you're going to see, particularly with these 100 ,000 clusters, some of these companies are going to have greater than a 100 % advantage of, and exaflops per dollar of capex per watt consumed and this is when you're going to really separate these labs because that is the difference between gpt 7 or 8 costing a trillion dollars or 500 billion the next class costing 200 billion or 400 billion and then for inference it's the same It's the exact same. Not only is it cost to serve, but it directly affects user experience in terms of tokens per second, which we know from Google is one of the most important things for search UX.

30:14So all of this is also going to apply on inference. Inference is just a lot easier because it really, really just comes down to memory bandwidth and on-chip memory. This reminds me so much of the Vaclav Smeal history of energy stuff where you would get a new, he called them prime movers, some source of energy, fossil fuels, wind, whatever. And then you get this long period of the gains all coming from the efficiency. So if you're spinning a turbine or something, in the early days of coal, the turbine only captured 10 or 15 % of the available energy from coal itself. And now we're at like 98 or something like that.

30:49And basically it sounds like that's the exact same thing that you're describing here. The GPU is the coal and those will keep getting better, which is cool in technology unlike coal. But it sounds like that's basically the story, which is so interesting. That is like turning sunlight into usable energy via different fossil fuel mechanisms and motors. And this is turning sunlight into compute, literally. And now that sunlight can come in the form of actual sunlight, it can come in the form of artificial sunlight. That's nuclear. clear it can come in the form of stored sunlight that's fossil fuels but this is the efficiency at which you turn sunlight into compute and so cool and the dollars you pay for that ratio yeah it by the way it is just interesting like it's one reason like i was never that excited about wind because turbines were really efficient whereas you could look 10 or 15 years ago solar was terribly inefficient i actually think it's awesome you know it's a little sad to me you know All these kids are really worried about global warming, and you have all these things about 20-year-olds, oh, I don't want to bring children into a warming world.

31:54Global warming is a big problem. It is a solved problem, and it is solved because photovoltaic cell efficiency is compounding just under 10 % a year, and battery efficiency is compounding 200 bits less than that. And if you compound that out, over the long term, not even the long term, the world is going to run on sunlight directly. Fossil fuels are just going to go away. It's because of economics and it's because sunlight and storage is going to be cheaper than every other way of providing power. It's just not going to be a problem for the next generation now. Maybe we hit some tipping point and it's irreversible and this and that.

32:33Emissions are going to collapse in my lifetime. Literally collapse. obviously like pre-industrial age humans were actually massive polluters because the pollution per unit of firewood is really high and you needed fires otherwise you'd get eaten by a saber-toothed tiger now there weren't that many of them it is possible that we're going to be back to like neolithic levels of emissions in my lifetime assuming i can live a little bit longer because of AI. So that's an awesome and encouraging thought. That's good. And I just wish somebody was shouting that from the rooftops. Well, here we are. I'd love to keep going up the stack here because every single level is interesting.

33:20So the semiconductors themselves and the potential innovations there are interesting to me. The data piece is really interesting to me all the way up to the application layer. And I'm just curious what you think about all of this stuff. So one of the things that we're assuming in all this is that we're going to have more data to train these things on at GPT 5, 6, and 7. And I think there are some interesting discussions about what available data there will be or how we'll get more data or how we'll create synthetic data and whether or not that will work to create a better model. I'm interested in what you think about all of this stuff that is required if we're going to have the big arms race that you were describing earlier.

33:54So I do think this was a real bear case maybe nine months ago, but I think it was hinted at in the CLAW 3.5 technical paper, and maybe a little more explicit in the NVIDIA-Nemotron technical paper. You know, it's awesome. NVIDIA, whatever, it feels like there's going to be a rate-limiting factor for AI. They've solved it, and that's what Nematron did. But I do think for reasons that no one understands, and no one understands these models. No one understands how they work, why they work, why scaling laws. There's all sorts of theories, and maybe we're getting a little better at understanding them, but no one understands them.

34:29So no one understands, point one, but it does look like synthetic data works. No one understands why, but it looks like it works. Now, again, will it continue working? I don't know. Nobody knows. No one knows. People who have seen, Kevin Scott just did a podcast, and he basically said, look, I've seen some early checkpoints of GPT-5, and scaling laws are continuing. I think in a lot of ways, that's the best indication we have that they're scaling. And I actually think largely in respect to XAI, I do think OpenAI is the combination of what XAI is doing and Blackwell delay. The Blackwell delay means that if you're waiting for Blackwell and you're trying to get 100 ,000 cluster, XAI is going to have arguably a one-year advantage, which is untenable to these other labs.

35:18So they're all now frantically working on stating up their own 100 ,000 cluster, but they don't have Elon designing the data center, redesigning the data center from first principles. Data center architecture was always like kind of nice to have. Now it is must have. It's existential. I think synthetic data does look like it's going to work. This goes to where will the value for these models come from? Why is meta so comfortable open sourcing? You may ultimately see everyone open source. UXAI has open source, Glockwad. But the reason is the value may not come from the model. Now look, if you're running that equation I described and you have a one to 200 % advantage on exa-flops per CapEx dollar per watt and scaling laws hold, you're going to have such a massive advantage of model quality that you're never going to open source it.

36:07You can still effectively kind of steal models, distill them. If you have that compute advantage of scaling laws hold, wow, you're going to have something very valuable. The value clearly comes from distribution and unique data. Meta has open-sourced Lama. They haven't open-sourced all their data, and that data is just going to be for their version of Lama. So it will for sure be better. Google, they have YouTube and then all sorts of other data sources that they've developed when they tried to boil the ocean to create the knowledge graph, which are those little knowledge panels that kind of appear when you do searches.

36:43So between YouTube and the work they did for the knowledge graph and Google Maps, they have crazy data. They don't care because they can use that data to monetize their model. XAI, whether they open source or not, they will always have, for sure, access to X data in a way no one else does because X owns 25 % of XAI. And then I would think over time, XAI will kind of be an intelligence layer that cuts across Elon's ecosystem of companies. And by the way, we should talk about AI and robotics, which I actually think may be the biggest disruption in our lifetime, comparable to artificial superintelligence and these digital gods.

37:20If these models, unless someone develops a compounded advantage of mammoth, SFU, checkpointing frequency, and PUE, such that they have a dramatic difference at exoflops per dollar of capex per watt, I think they're probably all going to converge to roughly the same intelligence. And the reality is, given the way that I think at least Google and Meta are thinking, even if someone else is way ahead of them and more efficient, they will try to solve the problem with money. Oh, wow, it only costs them$300 billion? No problem. We're happy to spend a trillion. Although, ultimately, economics will apply.

37:59These may impact stock prices. And I think it's not inconceivable some of these Magnificent Seven forget buybacks and dividends. They may eliminate their dividends, stop buybacks, and start issuing stock to fund this. in four years, we could be in a really wild world in a lot of ways. But these intelligences will probably converge on kind of a similar what I'll call IQ. And then it's just who has the most differentiated real-time data about the world. It's these unique data sources coupled with every time you rate an answer from one of these models, you're helping it improve. And so if you can couple unique data with internet scale distribution, then you're going to have a winning formula.

38:45And there's only a few companies that have that. You know, it's XAI, it's Google, it's Microsoft, although internet scale, maybe. And OpenAI gets that through Microsoft, probably Amazon and Anthropik gets it through that. And then, you know, obviously Meta. And then Apple is the big wild card. And we should talk about that because I think one of the biggest things is where is the inference going to happen? and compute tends to cycle in between centralization and then decentralization and we're at the end of like a long period of centralization in the cloud where a lot of compute ran in the cloud in these big data centers is just because you could get much higher efficiency in these big data centers now clearly for trading that is going to happen in giant data centers that will be in the cloud although something that i think is underappreciated about these AI data centers and why we shouldn't worry about them over the long-term crushing power demand is you can put them anywhere.

39:38You can put them in Wyoming. They don't need to be near a big city. I think we'll eventually see giant data centers in shale gas fields with power plants on those fields long away from any humans. That will probably be an intermediate-term solution to the lack of nuclear power in America. The U.S., we have all the best models. The U.S. wants these models trained in America for better or worse. Obviously, Mr. All is French. I don't know where the models were trained. They have these cycles of centralization and decentralization based on where can you get the lowest cost compute at the highest utilization, kind of a variant of the equation I described earlier.

40:17And I do think the subcomponents of that equation, not only are they maybe helpful to labs, but they're really helpful to investors. If you can improve SFU by 20%, it all comes down to millimeters of silicon. For only a few more millimeters of silicon or maybe even less millimeters of silicon, they're like, oh my God, those millimeters of silicon are the ultimate end cost. Wow, you have a massively winning formula. So you can just almost go look at each step of that long chain, and where do you have really high costs per square millimeter of silicon, and then that is the opportunity to really optimize.

40:58And whether it's at the infrastructure software, the data center hardware, the semiconductor layer, it's all there. But I think you're going to see inference increasingly done on phones. And this is clearly Apple's play. It is one reason that they are in such an advantaged position. all these other companies are trapped in this prisoner's dilemma i am sure they would like given how much it costs they're all economic animals if they could just like reach an agreement and say you know what let's not kill each other is going to stand up a blackwell cluster until 2026 like they probably all sign it but it is literally a classic prisoner's dilemma and that would be a Nash equilibrium.

41:45But we're not going to reach a Nash equilibrium in the race to create digital God when the states are existential. So there are these prisoners to live. Apple's not. In the same way, Google spins collectively, cubitantly spin, I don't know, hundreds of billions of search between CapEx, OpEx, all this stuff. And Apple monetizes it almost as well as Google because they have this distribution chokehold, toll booth in iOS. They're clearly going to do the same thing with Apple intelligence, then Google will obviously do that with their phone, Android. And so you're going to have a small model running on your phone that for simple questions, think of it as like a 100 IQ model that with like really, really sophisticated knowledge.

42:29There'll probably be two of them, they'll check each other. And a lot of inference will then happen at the edge. And if inference is happening at the edge, then I do think we're going to see super phones because the advantage to you as a human, what bounds inference at most times is memory. And you will be willing to pay. Today, iPhones are sold based on how much flash storage they have. What you will most care about is the amount of DRAM you have on your phone because that will determine the perimeter count of the model you can run locally. And that is therefore the quality of your own local intelligence that has access to all of your data in a privacy safe way.

43:11And when you ask it for something, if it can be done on a compute, it will. And I think from networking, like a truism is route when you can switch when you must. But I think the next thing for inference, it will be local when you can cloud when you must. So if you can inference locally on your phone, you will always do that because it's free. And cloud inference definitely cost money because you're burning GPU hours. Now, it's very efficient, but the inference on your phone is effectively free. And then increasingly, our competitiveness as humans, you know, it's very funny. A lot of really smart people are a little skeptical about AI.

43:51I get it. It's like you have 125 IQ or 135 or whatever it is. AI, particularly your domain, isn't at all impressive to you. And it's like, okay, it's a little bit for search for me learning about other domains. but man a lot of humans don't have 120 iqs and what people are missing is that there's a lot of humans with like i don't know i have no idea 100 iqs yeah so let's just say it's 100 iqs and they can use ai and all of a sudden they're like a 115 and it's like holy shit it's this is cool and that's just gonna keep going until these asis have iq of a thousand and it's like why I'm, as a human, am I bothering to do anything other than work with an AI to create art?

44:34Anyways, if I could have like a$3 ,000 iPhone that has four times the DRAM on it that the$1 ,000 iPhone has, it did maybe more storage so that local model could do rag. And this is my model. It likes me. I pick the voice that talks to me in. It knows me. It's my friend. It's my agent. all at once is thumbs up for me in that process of improving these models. I think they will ultimately be, there will be a way that they could be RLHF on an individual basis on the fly, which we're gonna need a lot of breakthroughs to make that happen. And it's gonna be my model that likes me and knows me and knows what I want.

45:18And then this is where you get this agent future that everybody's talking about, where it's like, I have an agent on my phone. And if I have an agent on my phone, because I have a super phone, that has an IQ of 115 or 120 and someone else's agent only has an IQ of 100, I'm going to be profoundly advantaged as a human. And then that continues all the way up because I think eventually Apple will monetize this. And the way Apple will monetize this clearly is when the on-device LLM isn't smart enough, they're going to send it to the cloud. They're going to use something called a router, which is what every AI application company uses.

45:56you just want to route to the best model for the query per dollar apple will charge companies to be in their router and then oh if you want to be routed to eventually you're gonna have to pay a bigger and bigger toll and everybody will pay that toll the same way google paid that toll but if i as a human i can have a smarter model locally and then i can opt in into maybe apple will just make it easy they'll say okay you can have cloud super intelligence or you can have cloud intelligence and I pay$60 a month for cloud super intelligence. Now, or whatever it is, $1 ,000 a month,$10 ,000 a month, what would you pay for that?

46:36And so I have 20 points of IQ on my phone relative to a lot of people. And then I'm paying$10 ,000 a month for cloud super intelligence or hyper intelligence. Super intelligence is just$1 ,000 a month. And then regular intelligence is $20 a month. I'm going to be really advantaged as a human. And that seems dystopian to me, but it's also kind of hard for me to not see that happening. And many, many investment conclusions flow from this in the same way that they flow from like, hey, okay, Mammoth is at 90%. Okay, so that's probably not a good place to invest. SFU is at 30%. Wow, that's a great place to invest.

47:18PUE is at 1.8 and it could theoretically be at 1.3. That's a great place to invest in that equation. You kind of want to invest at the most inefficient chain. A lot of investment implications flow from where does inference happen? And then a lot of implications for humanity flow from that as well. I'd love to talk a little bit about the commoner part of this whole equation. We talked about like Mount Olympus and like the game of kings fighting each other for dominance. We haven't talked at all about. What about just like a normal new company that's being started to take advantage of these super intelligences at the application layer or to do something else?

47:54Then we'll get to robotics after that. But if I force you to just be a early stage series A, series B investor in the world of startups, and they don't have this kind of resources that all these Magnificent Seven and others have, how would you think about the aspects of companies and opportunities that would get most attractive in the world where GPT-X keeps scaling and getting better? First thing I would just say, it's always, what are people doing? What I am doing is really focusing on companies that improve that equation, elements of that equation I described earlier, because I think - That's the choke point.

48:27That's the highest likelihood of success. And in tech, whenever you find a constraint, if you can invest in something that alleviates that constraint, you usually do well. And right now, I profoundly believe the constraint is at SFU, checkpointing, which goes to reliability, and then PUE. So that is where I am targeting my dollars. I think the application layer is really, really hard. There are people who are killing it. I think you have Sarah, who I've never met, but I admire on your podcast. And she's like absolutely, from my perspective, crushing it at the application layer. but just geez investing there today feels really hard to me and i just think anyone who has a lot of conviction about that needs to be reminded of that one percent chase coleman's stat yeah i was just thinking that all these vcs have really strong views of like there are all these companies that were funded in 95 96 97 and they seemed like they had got in the net but it just takes time I guess they did go into that for the VC.

49:30They went public and they were able to cash out, but they actually didn't go into the net with the fullness of time. So I just think it's important to have a lot of humility at the application layer. And maybe it's just that I am so comfortable with this infrastructure layer. That is where I have concentrated to date. There's really only been one application company that I've been excited about. We've seen a lot of them. maybe I'm like too aware of the Chase Coleman stat and I'm too and I'm being too careful at how I approach this application layer but I'm gonna miss out on a lot of CMGI was incredible you know what were some other besides Yahoo what were Lycos you know there are all these companies that we don't even think of today but they're incredible venture outcomes maybe I do need to be a little more mindful of that, but there are clearly people, whether it's Bichmark, whether it's Cirrus firm, who are succeeding.

50:29And what I was saying in the application layer, I heard this from Vishria, but it really resonated with me. And I think it also goes to this whole spilling AI ROI debate. Not only is it an ROI debate about creating a digital god, but I think there's a super clear ROI, and you can actually do math to show that. and a lot of people are conflating CapEx and OpEx in ways that are just not helpful. Meta went down like 80%, partially because they were overspending on the metaverse, but partially because Apple took away their ability to target through IDFA. Meta is a 5X and the revenue growth is reaccelerated.

51:03Why is it reaccelerated? Because they spent a vast amount of money on AI to figure out how to target. It's like maybe the meta return alone justifies all of the spending. They call that meta advantage. But meta advantage is just part of it. That's just where as an advertiser, you can let meta do the targeting. And then Google has performance max. And all of that is, and that's where the, there used to be ads were created. And we're going to come back to applications, but ads were created. We're going to pay some humans to figure out that we need to show this creative to like white 40-year-old guy in Boston and this type to this type of demographic category in this city.

51:41And we're going to show them at this time of day. No, we're going to show them after the local sports team is one and it's sunny because that's actually the best time to see an ad. But now AI can do all of that. And we're going to have a million creatives and it's going to be optimized on the fly. And that's what AI is enabling Google and Meta and other firms to do. And probably just on that - Just that justifies it, yeah. There has been a massive ROI on AI. What are we talking about? And then the other thing that's really funny to me about this whole AI ROI debate is, okay, we're going to have this abstruse debate about ROI.

52:15These companies are all public. And there is something called return on invested capital. And ROIC has gone up for all of these companies since they ramped CapEx. What are we talking about? If you're an AI ROI skeptic, why has ROIC gone up at these companies? Well, yeah, CapEx is up. Notepad is up more. Why is that? Because they're doing exactly what you would expect to happen in a world of AI. they're trading off human labor against GPU hours. That's why their ROIC has gone up and these GPUs are really efficient. Let's have an AI ROI debate. When the ROIC at the big AI spenders starts to go down, until then, it's just the height of intellectual ridiculousness.

53:04I mean, ridiculous. What Vistria said that I thought was so powerful for the application layer, there's all these SaaS metrics. After your first year, you want to be at million to be on pace. After your second year, you want to be at 5 million. After your third year, you want to be at like 10. And if you're above that with reasonable cash burn, you probably run these curves, you're going to be a successful SaaS company. And all those curves, which everybody knew is one reason SaaS multiples got so inflated because it became almost like quantitative. Oh, wow, this company is at$15 million in year three?

53:35That means we can pencil them in for a billion dollars in year eight. A, that led to multiples expanding ridiculously. B, it led to the industry being overfunded such that the degree of competition in each vertical went up such that the curve's no longer held. And this is one reason why if you made a lot of SaaS investments, particularly in 2021, you're in a world of pain. And then on top of that, here AI comes along, and it fundamentally changes the paradigm for application software, because what application software fundamentally does is makes humans more efficient. And today we're in a state with AI where it's making humans more efficient, but you don't need that big of a wrapper around it.

54:13That's where we're going to come to the mystery comment. And then ultimately, if it's replacing humans, and you're selling application software on a per seat model, and those seats start to go down, you have a problem. These companies are super aware of it. They were going fast in one direction, and now they have to contend almost in every vertical with this next generation of AI-first application companies, which are just these often very thin wrappers around GPT. You'll call it an LLM wrapper. And what Fishery said to me that was so funny is like, all these AI companies are blowing these traditional SaaS metrics out of the water.

54:50And all this isn't being counted in those AI ROI papers. Just that there's all these companies that are going like zero to 30 million in nine months. and even though AI is really high marginal cost, they're being more cash flow efficient relative to a software company. So in almost every vertical, there's multiple companies that are AI first that are just really, I think what Vishri said was they're just paper thread wrappers around pick your LLM of choice. Then yeah, they use a router to find the best LLM, but they're like magic to their customers. They aren't going after software budgets. it's their going after labor budgets.

55:29It's very intellectually interesting to me to watch. How do you get conviction that one of these companies can use what is initially magical to their customer to build defensibility in, all while allowing for the fact that the big models are going to continue to compound at a really high rate? And this goes to things like, maybe you just find a really good sales motion that works. Maybe you find a really easy integration point. Maybe you build some defensibility around that integration. Maybe you make it really easy for a small business to do rag on a really unsophisticated system. How are you building differentiation into that wrapper?

56:11And hopefully, you're not fine tuning, because that does extensively fine tuning, because that locks you into a model generation. How are you creating a compound AI system such that you're serving each query at the lowest cost possible, and you're starting with a small model and treating it to a big model. There's a lot of things you can do, and all of those are really, really important. But I just think at that application layer, there's just, it seems like for any category, any category, it feels like they can be imagined. There's multiple startups that have exploded that are defying all traditional SaaS metrics, and it's not clear to me at this point how much defensibility each one of them has.

56:55Some of them are going to have immense defensibility around one of the axes I described, but I think a lot of them may not. And I think that's what makes it so hard and why I've been approaching that space so carefully. You know the amazing thing? I think the framing around labor, I mean, just take me as a stupid, simple, tiny example. I have a team of three people building one of these, we'll call them like a really light LLM wrapper for doing research on private companies. And if you just think about how often I use this thing in a way that I normally would have used an analyst, it's multiple times a day and it's returning work as good.

57:31So there's so much undifferentiated heavy lifting in every white collar labor market, I guess is like one major learning. A hundred percent. If you can have the answers instantly and for effectively free, It's crazy how much you use these things. We're just getting started. Yeah, and I have the same thing with like an LLM at my firm. It was actually very interesting. We had an intern who was really talented. And he said, listen, I think this internal AI tool is amazing. And I use it every day and more every day. It's gotten much better this summer. We've worked on it for a long time. And then I was like, show me how you use it.

58:04Because I'd been using it a certain way. And then this literally, you know, this 21-year-old kid was like, well, this is what I do. And now all of a sudden my usage of this tool is two hours a day, three hours a day, four hours a day. And I do think Jack Welch had this term called scut work, just like hard, unpleasant work that had to be done really well in a white collar setting. And I think it's all these little rappers or mini rappers, whatever we're going to call them, are going to replace a lot of that work. And initially, it's going to be in combination with humans. With science fiction, Neil Asher's world, there's something called a Hi-Man, which is a human-AI hybrid.

58:44And it's where they're linked through something like a neural link to a computer server that is supported by an exostructure, a robotic exostructure that they walk around on. There's like fusion chess, whatever you want to call it. We'll have that for a while. And then that goes to the, you know, 100 IQ humans performing like 120 IQ humans and then forming like 130 IQ humans, then 130 IQ humans performing like 160 IQ humans. But then eventually it feels like as long as scaling laws continue, which is a big F, it's just going to be the AIs. Maybe we could talk about robotics. I had a really interesting conversation, call it April of this year, with an investor that has been investing in lots of these same things for long periods of time, privately and publicly, and has big positions in lots of the companies that we've been talking about today.

59:32His observation to me was the big underestimation that's happening over, let's say, five years is the role that robotics and robots will have combined with all of this technology we've spent all of today talking about. And I would love to hear you riff on that because in the near term, it feels like a little bit, quite a bit of frothiness, like some crazy funding rounds for these companies that you don't really know what the robots are being designed to do, sort of general purpose, humanoid type robots. There's all sorts of interesting, more specialized ones that are cool too. But what do you think about all this?

1:00:05It does seem kind of under discussed relative to just all the foundation model and semiconductor stuff. I agree. I think it may end up being a bigger near-term disruption than what we were just discussing, the automation of a lot of white collar labor. I think the first robot, the first robot that's really going to impact the world is every Tesla car with what they call their AI for hardware. Because from my perspective, there's a publicly sourced miles between disengagements. And you have to remember for Tesla, Tesla is going to get the same miles between disengagements. Like if you built a new city, on mars and it was populated by entirely different looking cars and streets and everything you could drop a Tesla in that city and it would have the same miles between disengagements that it gets in any other city.

1:00:53Whereas something like Waymo is geofenced, we're really only using it in cities that have nice grids and good weather, et cetera, et cetera, et cetera. It is clear to me, looking at the crowdsourced data, miles between disengagements with different versions of FSD, that when they cut over to 12.3, which is effectively all deep learning, I think eliminated almost all human code, something dramatic changed in the rate of progress. And then when they cut over to 12.5, which runs best on the AI4, which used to be called the HW4, it's just like the local computer of the Tesla, it is now rolling out to AI3, was another step function.

1:01:38And these, going to that same scaling law, those step function improvements were made with a fraction of the compute that Tesla is now installing publicly in their data center at the Gigafactory in Austin. They've actually filed, sometimes I wish they, as an investor, I wish they couldn't file so many of these patents, but they have filed some really innovative patents for data center cooling related to what they're doing with what looks, I think has been publicly said is going to be over 50 ,000 H100s or H200s. FSD is now on the same scaling law, and arguably on a faster scaling law because they have a lot of catch-up to do that GPTs have been on.

1:02:23So I think 12.5 is like GPT-3. And it can consistently drive me most places with no interventions. I'm a seasonal driver. I really only drive my Tesla in the summer. And actually, my wife, Becky, tends to do most of the driving because she likes it more than me. So we kind of get like a seasonal look. It's almost like every May we check in. There was just always continuous progress. This year, it's like when we turned on 12.3, it was like all of the progress over the last 10 years was in that one release. From the first time I had that Tesla, and then we had that again when I went 12.3 to 12.5.

1:03:05and they're at probably like a GPT-2 level of compute. I think they're going to go really fast to GPT-4.5 compute, which means you're going to get, using these orders of magnitude, you're going to get like a 100x improvement really fast. So I think there's all these people who have been skeptical. They're all in for abject humiliation. They just are. And then unlike GPT-2, only Tesla has access to a visual trading data set that is based on miles driven. We could argue whether it's 100x, 1 ,000x, 10 ,000x, bigger than the second biggest trading data set, which is Waymo. So it's like people, oh, how are they going to make money?

1:03:50Well, it's like in this case, in the world of self-driving, from my perspective, it's like they own YouTube, they own all of Meta's properties and the open internet and X. And then other people are like trying to do it using Yahoo. Yeah, using Yahoo. Like, good luck. Like, who's going to win? Now, obviously, that could change. Important to have humility. There may be an algorithmic breakthrough that reduces the importance of that trading data set. And for sure, Waymo is going to try and brute force it and just throw whatever amount of dollars they need to get the data to compete. And they have a different approach using LiDAR, Tesla doesn't.

1:04:34We'll see. Like, I don't think it's a foregone conclusion. Nothing about the future is certain. But just if I look at how amazing 12.5 is on AI4 hardware, and think about the tiny amount of compute that that was trained on, in the mega cluster that they are standing up in Austin using known techniques. We're going to skip, I think, 12.5 at GPT-2. We're going to skip really quickly to GPT-4. Then look, I'm sure Waymo will brute force it. There may be algorithmic breakthroughs such that there are other people, we'll see. But then the other big thing is just using an LLM for FSD. One of the best followers on X is Dr.

1:05:13Jim Fan, who's NVIDIA's head of robotics. And he's had a lot of posts about how there's a fascinating exchange between him and Elon on X. It is amazing the extent to which AI happens on X. The JAX team at Google and the PyTorch team on Meta got into this bitter fight, and it went to Mammoth. Which framework was better for Mammoth? Eventually, the heads of each lab had to step in publicly on X and make peace. But like, wow, you learn so much just following that fight. And like every AI researcher is active on X. AI happens on X. And it's such a great forum for using it. But Jim Fann had this fascinating exchange with Elon, where Jim Fann talked about how LLMs could massively improve FST.

1:06:03And Elon replied, yes, the only two data sources that will scale infinitely are synthetic data and real-world video. And I thought that was interesting. And then that goes to, I think, maybe the biggest risk in which this view that I just described of Tesla's autonomous future is wrong is just if synthetic video data can be used in the same way that synthetic data can be. We know that synthetic written data works. We don't know if synthetic video data works. Nobody knows. And obviously, there's a very high bar for regulator. You know, I think it's something like, I forget whatever it is, like 50 ,000 or 100 ,000 people die in car crashes every year globally.

1:06:42It might even be a million. Obviously, we could take that down dramatically using AI, but we're much less willing to tolerate traffic, fatal accidents from AIs than humans. That is what it is. So, you know, it's going to be heavily regulated. But Dr. Jim Phan posited that the reason LLMs were going to be able to really help with FSD is because of the following. This is the way my, relative to some of the people working on these problems, my comparatively low IQ brain conceptualizes it. Anything that's been trained on real world data just knows what to do, what a really good human driver would do in that real world situation.

1:07:19If there's a novel situation, it may not know what to do. And that's where, from my perspective, the LLM can really help. Because one of the emergent properties of GPT-4 and we can debate whether or not it actually is an emergent property or just in context learning. It has what's called a world model. And that means, I'm sure you know this, but if you ask GPT-3, hey, what happens if you stand a champagne bottle upside down and put like a basketball covered in soap on top of it, GPT-3, no idea. GPT-4 will often get questions like that right. I should actually see if he gets that exact question right.

1:07:59A three-year-old human will say, that's going to fall, the champagne bottle is going to shatter. It's really hard for GPT-3, and that goes to this jagged frontier that people talk about. So if you put a really speed-optimized small LLM in locally on each Tesla, there might be just enough reasoning capability to unlock another step function and FSD capability. And now look, Weibo will have that too, has lots of other people, but they won't have Tesla Vision, all of the proprietary data set. Just think this is going to be a reality in a way that is abjectly humiliating to everyone who is an FSD skeptic in the next 12 to 18 months, maybe in the next six months.

1:08:47And I have never been willing to make a prediction like that before. So then you take that and the same thing goes for humanoid robots. Google showed this with research called TensorRT2, where dropping an LLM into a humanoid robot with a world model that understood what things were and what to do just made it so much easier. Instead of training that humanoid robot how to pick up a tennis ball and a basketball and a football and how each one is different, it could reason. And so this is why putting LLMs into these humanoid robots I think is going to be so transformational for the world and make a lot of blue-collar labor optional.

1:09:28I do think politicians and political systems are utterly unprepared for what may be coming. The one thing I would say that Elon and Jitson profoundly agree on publicly is that humanoid robots are the future, not the specialized robots. And the reason is just, of course, a specialized robot can be better than a humanoid robot at any given task. But the humanoid robot can do any task that a human can. The world is optimized for humans. And there are massive scale efficiencies in manufacturing. And so because you can make, it's almost like humanoid robots are going to be to the field of robotics as GPT was to AI.

1:10:08GPT was a generalizable type of AI. And these humanoid robots are going to be a generalizable form of robotics. And because of that, they're going to be manufactured at such a scale that they have a cost advantage. And then the physical world is going to start to be optimized for them. And that's why they're going to win. So good luck to all these non-humanoid startup robot companies. I hope you get Lycos or CMG or MySpace type venture outcome. But I don't think any of them are going to be Google. And in the same way that so much of GPT advantages incumbents, whether that's Meta, Google, X, XAI, Microsoft, these humanoid robots, the reason it advantages incumbents is because they have the raw ingredients of data, compute, and capital, which is what you need to effectively monetize these.

1:11:01And that's why their ROIC is going up even as they have CapEx. I do think that incumbent manufacturers who have expertise in battery design, actuators, motors, with big data sets are going to be advantaged. I'm reasonably bullish on Optimus. That's just such a giant market. There's going to be so many competitors. And whether it evolves FSD, I could see a world where there's just two or three companies. Maybe it's Tesla, Waymo, and some open-sourced variant. Or it could be synthetic real-world data works. LLMs really improve the efficiency of that specialized visual algorithm. And so there's like thousands of them.

1:11:47I think that's unlikely, but it's possible. I think robotics may end up advantaging incumbents in the same way I think FSD and LLMs, GPT have advantage, generally advantaged incumbents. And then startups willing to take like a new AI-first approach. But I do think robotics is going to change the world. It's super exciting. Like, I can't wait to have my own personal robot. Or 10. Yeah, or 10. Like, sign me up. Like, I just want one everywhere. What time is it? It just tells me. It has, you have an AI that's loaded into my phone is also running on that. And it likes me and is friendly to me. And it laughs at my jokes.

1:12:21Sign me up. Yeah. Yeah. I'd love to ask a couple of closing questions that are big picture, big arching questions. The first is responding to something you said earlier, which was having watched Jensen and Elon and Lisa at AMD, operate for so long and rating them as these like exceptional CEOs. It just seems like such a small handful of leaders at these companies can have such a massive impact on the world. And so their character and way of operating is really important for all of us. Of course, this will recycle and we'll get new ones and up and coming ones or whatever. But what is it about that group of three and maybe throw into the recipe a few others that you've learned a lot from?

1:13:01What is it about those three that match them so well with the modern world and way of company building. I mean, like the crazy stuff like Jensen having 40 direct reports or Elon working on seven companies at once. These are unusual human beings. Maybe just riff a little bit on why those three and sort of the nature of leadership in this era of technology. I guess I would say, above and beyond having clearly unusual IQs and ranges of domain knowledge, you know, and I do think it's hilarious that Lisa and gins and her cousins. It's like crazy. I just want to bet on everyone in that family going forward.

1:13:36It was like a first cousin. Died me up to bet on those genes. What a crazy coincidence. Beyond inherent genetic advantages, I actually think the modern American corporation in the same way, like almost any large organization with its hierarchies is set up to reward people who are charismatic and political. And because those people are charismatic and political, they're not always really smart. And because they're charismatic and political, they get big egos because people like them. They get used to having their way. The same way every, I think, child under four is like a functioning sociopath.

1:14:17Most people who become multi-billionaires are really powerful politicians. The physical part of the brain that deals with empathy shrinks changes you as a human. It is a long way of saying that I think the way that modern corporations have been evolved to be run, you end up with a lot of not CEOs who aren't excellent, but objectively terrible CEOs who were amazing at rising the corporate ranks, flattery, whatever it was it took to get there. And then by the time they got there, their ego had gotten so big, their sense of empathy had gotten so small, they're no longer capable of functioning effectively.

1:15:02And often with people like that, you find there's like one or two people behind them. There's like a key leader who runs a division or whatever, who's kind of arisen with that CEO through each division or group of people. Those are the people with the good judgment and skills to make high quality decisions. So Elon and Jensen in particular are nothing like that. Not only are they the front person, but they are always working on the most critical problems at the company. During my year off, I went to see Jensen. It was a different conversation because I wasn't working. And he just said, and maybe this has changed, by the time he said, I have no fixed schedule.

1:15:48I have no standing meetings. I just find out what is the most important problem at the company. And I go and I sit my desk down in that area. And I pull the best resources to work on that problem. And I love it. So he's working on the most important problem. That is also what Elon does at each of his companies. Whatever is the most important problem is what he is working on. When the Raptor engine was in the critical path for Starship, I'm going to get the details wrong, but I think there was like a standing 1 a.m. on Tuesday morning or Monday morning meeting for Raptor engineering, only the like 12 or 18 or whatever the number is, smartest people were allowed to be there.

1:16:33Everyone at the company wants to be there. So I think that is one, something that ties them together, working on the problem. Second is loving to hear bad news. My dad was a bankruptcy attorney. He always said the number one thing all bankrupt companies had in common was the CEO who didn't like to hear bad news. Just say it's Elon and Jensen love hearing bad news. At those companies, if there's bad news, it must immediately go to them. And that's very differentiating. No hierarchy. Wherever in the company the problem is, that is who Jensen and Elon want to work with, the subject matter expert.

1:17:10whether they're like 23 or 50. There's no hierarchy. You know, it's like the same way, like there was this apocryphal story when, I think it's true, when J.P. Morgan was buying Bear Stearns, the best modeler in the company was 24 years old. And so they set up a desk for him side by side with Jamie. And Jamie would say, change that, change this. You know, 24 year old kid was like the guy because he was the best. You know, Jamie Diamond, another exceptional CEO. He didn't ask for that guy's boss or boss's boss or boss or whatever. He's like, this guy's the best. I want to work with him. And the last thing that I think really, really ties together, particularly Jensen and Elon, is a mission orientation.

1:17:49Jensen started out about like, hey, let's make photorealistic video games in virtual worlds because that'll be exciting. Remove what someone called reality privilege and humans will eventually be able to be in the virtual world. And by the way, the metaverse is still coming. I think it's just clear that it will happen first with augmented reality and then with BCI. Just VR goggles, can't be a skeptic, never going to work. Augmented reality glasses, ambient computing, and then BCI is kind of the end state. It's smart, metabothic company controls CGL Labs and is working on this because I think that may actually be the next truly disruptive form factor.

1:18:29We're going to have super phones. The phones are going to stay as the compute layer for extended reality, but BCI may be the next true way we interact with computers after phones, the next true compute platform, and maybe the ideal way for AI going back to that high-man thing. But that was a compelling vision, photorealistic graphics. A lot of technical people are about to play video games. So I think Jensen hired a lot of really smart people based on that. And then Jensen, in a lot of ways, is one of maybe the single most, I don't know, single most, he's very important in the history of AI. Because very early, 10 years ago, he started hanging out with people like Jeffrey Hinton and Yann LeCoultre.

1:19:08I remember him saying those names to me and talking about the famous ImageNet competition with ResNet 51. And this means, effectively, I would paraphrase it, but we've turned intelligence into an engineering problem. And the way we're going to solve intelligence is just by gluing more and more GPUs together. He saw that. It dedicated all of NVIDIA to that. And then it became about intelligence. That's a mission. And then for Elon, all of his companies are really mission-oriented. PayPal was about making local, at the time, commerce frictionless and reducing the tax from all these crazy fintech systems and payment flows, reducing that tax on the world.

1:19:49And I think that tax has been massively reduced, not necessarily because of PayPal, but because of PayPal and a lot of other companies like them. And then Tesla was really about making the world sustainable. And I think between battery storage and pulling EVs forward, Tesla, the world, was always going to run on sunlight just because of economics. I mean, forget emissions or concern about the environment. Economics were going to dictate an emission-free world, apart from rockets, because you cannot get literally out of the Earth's gravity well without chemical propellant just from a physics of energy density perspective.

1:20:23Tesla's done a lot to, I think, make the world more sustainable. And it was just so striking to me when I first started meeting with Tesla executives back in 2011, 2012. They were and are so mission-oriented about making the world more sustainable and environmentally friendly. SpaceX, too. It's incredible if you talk to all of them. I mean, it's just wild. Yeah, the bartender at SpaceX. If you ask the bartender at the Tiki Bar, she's like, I know each engineer's drink. I have it waiting for them at the end of a long day, not that they're having drinks that often. Like, maybe that makes sense. One step closer to Mars.

1:20:54Yeah. But everybody in each company, all these companies, they can say the mission and they all say it with this messianic zeal. I think you get better employees. And then XAI is about being dedicated to objective truth and new scientific advances. And if you have these missions and you're competing, a lot of these, the world's best minds have spent 20 years trying to make people slightly more likely to click on one link than another. And if you have these missions, you get better employees. So I think that mission focus is a really, and the exceptional teams, and that those exceptional teams want to work with people like Elon and Jensen, because they know if you're 25 years old and exceptional, you go to one of those companies, and you happen to be in the critical path of a critical problem, you're going to work directly with them.

1:21:41And there's probably no other company where that is true that has more than, I don't know, pick a number, 5 ,000 employees, 1 ,000 employees, I don't know. So I think that mission orientation, there's no need for hierarchy. You get really exceptional employees. I cannot tell you how exceptional the teams at Elon and Jensen's companies are. Yeah. I mean, I've met a lot of the SpaceX people for a bunch of different reasons. And all of them have the same zeal, intensity, mission focus. It's absolutely remarkable. And a lot of them left and then came back very quickly because there was no other environment they could find quite like it.

1:22:19maybe like the most fun place to close our conversation today. I could literally do this with you for seven straight hours. I got to like one third of my questions and your passion for markets and technology is just so palpable and amazing. My last question is about our business. How do you think the investing process, investing firms, edge alpha will evolve against this backdrop that we spent the last two hours discussing, whether it's the internal tool you talked about or the one we're building. I'm sure everyone else is building something. It just seems like, holy cow. And this has, I guess, been the history of our business is it keeps getting more competitive.

1:22:55But how do you envision it five, 10 years from now as such a passionate practitioner? Public equity. So we both play video games. The meta of any given competitive video game is always changing based on how the game designers balance it. So if it's a PVP game, you might go from a sniper meta to shotgun meta. Sports are similar, but there's a fundamental truth, which is you want to have the highest KD ratio in a shooter-based PvP. The more PvE environment, you want to have the highest or the fastest clear of the most difficult apex in game activity. And the meta of how to accomplish all those is changing.

1:23:29Sports is the same. The NBA evolved from a mid-range jumper meta when Jordan was in it to a three-point meta. And it'll evolve back again. And this is competitiveness in the meta, but there's still an underlying truth in the NBA, which is you win by scoring more points than your competitor. The meta of investing, both public and private, I think is always changing. It changes much faster in publics. And there have been all these huge meta shifts since I started in the business. ROIC was a revolutionary concept in the mid-90s. Everybody knows about ROIC. ROIC was just a big improvement on ROE, which in some ways, focusing on ROE, might argue is Buffett's greatest contribution to the field of investing, more than the four filters, et cetera, et cetera.

1:24:13But ROIC was a big meta shift. Reg FD and how that changed company communication, that was a meta shift and one that I think was awesome. And from my perspective, Reg FD means there's almost no reason to talk to companies because I had kind of a unique experience. Like I was at Fidelity for 18 years. And it's very funny to me when hedge funds are small asset managers, and I define small as, let's say, under 500 billion, say, oh, we have an access advantage with public companies. No, you don't. No, you don't. If you add up every hedge fund's access, all of it, it's a fraction of what firms like Fidelity Capital and T.

1:24:53Rowe have. It is many orders of magnitude. And what was awesome to me, I had the year off, and I just realized they never say anything that is not in a transcript. And I went back and read all the transcripts and things that I thought were like big insights through transcripts. human beings, we have a big IO problem, hence BCIs. And the I for me of reading, the I being the input, my input speed for reading is probably something like five to 10 times faster than humans can talk. And my input when listening is bound by the output speed of my partner. So extremely low bandwidth form of communication, literally, in terms of bits or bytes of information per second.

1:25:32Like I think that and having transcripts of everything, that was a big meta shift. Credit card data was a big meta shift. Expert transcripts have been a big meta shift. I think LLMs are going to be the biggest meta shift. And I actually think they may be hardest for purely quantitative investors, because I think there's going to be a period of five to 10 years. It's very clear that Renaissance, Bridgewater, Two Sigma, they probably realize the importance of data and compute before a lot of other people. They're the only people who can compete with these tech companies who can throw 20 or 30 million or 40 or 50 million at like a leading AI researcher per year.

1:26:08And they cornered a lot of data sources. They have data no one else has. And they're throwing more compute at it than anyone. But I am optimistic that human fundamental investors, the tool I'm using is called Intel Pro, by the way, if anyone's curious. I think right now the only alpha left in the market for fundamental investors at the edge of probability because these quantitative investors, why is there stop being, why is there no dispersion anymore in value strategies? Well, because value strategies, the alpha from them was based on human emotion and people being embarrassed to buy stuff and afraid to buy stuff and taking career risks.

1:26:44Well, algorithms, they don't have any of those. And that's why there's not as much alpha. They're just really simple value strategies that worked amazingly well until, oh, 20 years ago when quantitative investing took off and they squeezed the alpha out of value strategies. That's not to say that value as a factor cannot perform really well and be the best performing factor, but there's just not a lot of dispersion within the value factor because of algorithms. I am hopeful that there is a five to 10-year period where if you're a fundamental investor like me who can get a slight edge on maybe future probability states over what's discounted in the market through deep domain knowledge and try not to have any biases, that I can combine what I have with this tool, with years of data.

1:27:33I'm just so grateful we have a vector database with years of data. I think it is going to enable me and other fundamentally-based portfolio managers to probably benefit from a lot of the things that these firms like Ritizad have been benefiting for. And this is a long way of saying what LLMs do, what AI does, is it means the human language is the programming language. And now because of these LLMs are very soon, I'm going to be able to program effectively from an effective investment effectiveness at like the same level as one of these$50 million a year people or close enough. And then combine that with my own unique set of data and domain knowledge in my mind.

1:28:16And I'm hopeful that there's like a five to 10 year period here where fundamental investors' share of the alpha in the market goes up at the expense of quantitative investors. That's my hope. We'll see if that happens. I don't know if it's going to happen. Going back to that high man, fusion chest, centaur chest, like I hope I have a five to 10-year period here where I can prosecute that. Very quickly in the world of venture, I think venture is going to evolve outside of pure Series A specialists. Venture is going to evolve. Series C it up is going to evolve the way of private equity. Mainstream private equity, there's no sourcing advantage.

1:28:56There's no pricing advantage. In fact, there's a pricing disadvantage. Because the highest bidder wins. You're just the highest bidder. Everything's an auction. Literally every deal that these big firms do is an auction. So where they compete and differentiate is an operational value add. And I think that is where Series C and up growth equity is heading. And it's not just, it's true operational value add. You know, I've spoken about this, you know, like I think Valor does this. Valor does it, yeah. Really do it. You know, really hard operational problems. It's not just helping with HR, helping with PR.

1:29:32It's not the LP window dressing operational support teams. It's real operational support. And I think that is where the world of growth equity is heading and evolving. But that's almost because of, in some ways, the combination of the rise of crossover funds and megafunds that would have happened absent LLMs. I do think LLMs are just going to make the knowledge part of venture so much more democratized. I think about IQ, EQ, and then there's KQ, knowledge quotient. And now it's going to really help to have a really deep differentiated domain knowledge database, but you're going to have to work hard with an LLM to take your IQ from whatever it is up 30 points because people who you had between your combination of IQ and KQ, you had 30 points on them.

1:30:17Now they're at your level unless you use an LLM. But I just think it will really for a while place the emphasis on JQ, judgment quotient. Almost maybe for A's and B's, the most important skill will simply be assessing team quality. And maybe that's that five to 10-year window where VCs can still differentiate in the same way fundamental investors maybe, I hope, will be able to take some alpha share from quantitative of investors because of LLMs, but just it will be all about JQ at the C, the A, and the B. And some combination of JQ and EQ, is this person exceptional? But I don't know. It's just a hypothesis.

1:30:52Fascinating stuff. Gavin, you're one of my favorite investors to talk to. I have loved today's conversation. I love how specific and detailed it was. You're also one of the most passionate investors that I've ever met about your craft. This has been such a total pleasure. Thanks for your time. Thank you, Patrick. This is awesome. if you enjoyed this episode check out joincolossus.com there you'll find every episode of this podcast complete with transcripts show notes and resources to keep learning you can also sign up for our newsletter colossus weekly where we condense episodes to the big ideas quotations and more as well as share the best content we find on the internet every week

1:31:41Thank you.

From the publisher

My guest this week is Gavin Baker. Gavin is the managing partner and CIO of Atreides Management, and he has been on the show many times before. He is one of my favorite investors to talk to and this may be my favorite conversation with him. Gavin first started covering Nvidia as an investor at the turn of the millennium, making him the perfect guest to discuss all things AI and investing. There is so much detail in this discussion and I’m incredibly grateful to Gavin for sharing his wisdom with us again. Please enjoy this fantastic conversation with Gavin Baker.

For the full show notes, transcript, and links to mentioned content, check out the episode page here.
-----
This episode is brought to you by Ramp. Ramp’s mission is to help companies manage their spend in a way that reduces expenses and frees up time for teams to work on more valuable projects. Ramp is the fastest-growing FinTech company in history and it’s backed by more of my favorite past guests (at least 16 of them!) than probably any other company I’m aware of. It’s also notable that many best-in-class businesses use Ramp—companies like Airbnb, Anduril, and Shopify, as well as investors like Sequoia Capital and Vista Equity. They use Ramp to manage their spending, automate tedious financial processes, and reinvest saved dollars and hours into growth. At Colossus and Positive Sum, we use Ramp for exactly the same reason. Go to Ramp.com/invest to sign up for free and get a $250 welcome bonus.
-----
This episode is brought to you by Tegus, where we're changing the game in investment research. Step away from outdated, inefficient methods and into the future with our platform, proudly hosting over 100,000 transcripts – with over 25,000 transcripts added just this year alone. Our platform grows eight times faster and adds twice as much monthly content as our competitors, putting us at the forefront of the industry. Plus, with 75% of private market transcripts available exclusively on Tegus, we offer insights you simply can't find elsewhere. See the difference a vast, quality-driven transcript library makes. Unlock your free trial at tegus.com/patrick.
-----
Invest Like the Best is a property of Colossus, LLC. For more episodes of Invest Like the Best, visit joincolossus.com/episodes. 
Past guests include Tobi Lutke, Kevin Systrom, Mike Krieger, John Collison, Kat Cole, Marc Andreessen, Matthew Ball, Bill Gurley, Anu Hariharan, Ben Thompson, and many more.
Stay up to date on all our podcasts by signing up to Colossus Weekly, our quick dive every Sunday highlighting the top business and investing concepts from our podcasts and the best of what we read that week. Sign up here.
Follow us on Twitter: @patrick_oshag | @JoinColossus
Editing and post-production work for this episode was provided by The Podcast Consultant (https://thepodcastconsultant.com).
Show Notes:
(00:00:00) Welcome to Invest Like the Best
(00:04:42) The Magnificent Seven and Tech Competition
(00:06:29) Generative AI and Scaling Laws
(00:08:36) Challenges in AI Infrastructure
(00:15:02) The Future of AI and Data Centers
(00:17:51) Efficiency in AI Models
(00:35:14) Synthetic Data and AI Training
(00:42:37) Inference and the Role of Smartphones
(00:48:35) Investment Implications in AI
(00:49:09) Opportunities for New Companies
(00:51:20) Challenges at the Application Layer
(00:52:25) AI's Impact on Advertising
(00:53:40) AI ROI Debate
(00:54:39) SaaS Metrics and AI Disruption
(00:55:59) AI-First Application Companies
(01:00:50) The Future of Robotics
(01:14:01) Leadership in Tech Giants
(01:24:05) The Evolution of Investing

More from Invest Like the Best with Patrick O'Shaughnessy

All 200 episodes
Gavin Baker - AI, Semiconductors, and the Robotic Frontier - [Invest Like the Best, EP.385]Invest Like the Best with Patrick O'Shaughnessy · 1 h 31 min
Listen in VO