20VC: AI Chip Wars: How Cerebras Plans to Topple NVIDIA's Dominance | Why We Have Not Reached Scaling Laws in AI | What Happens to the Cost of Inference | How We Underestimate China and Shouldn't Sell To Them with Andrew Feldman

24 Mar 2025 · 1 h 3 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Podcast Summary: The Twenty Minute VC (20VC)

Episode Title

20VC: AI Chip Wars: How Cerebras Plans to Topple NVIDIA's Dominance | Why We Have Not Reached Scaling Laws in AI | What Happens to the Cost of Inference | How We Underestimate China and Shouldn't Sell To Them with Andrew Feldman

Episode Overview In this episode, Harry Stebbings interviews Andrew Feldman, Co-Founder and CEO of Cerebras, a leading player in AI inference and training technologies. The discussion delves into the shifting dynamics of the AI landscape, particularly focusing on Cerebras's strategies against NVIDIA's dominance, the inefficiencies in current AI algorithms, and the broader implications for AI's future, including the role of synthetic data and international competitiveness, especially concerning China.

Key Discussion Points

  1. The AI Landscape in 2015
  2. Early Recognition: Feldman and his team recognized the need for a new type of processor to handle AI workloads that would arise.
  3. Market Underestimation: Initially underestimated the market size but recognized the importance of memory bandwidth and communication structures in AI processing.
  1. NVIDIA’s Position and Vulnerabilities
  2. Strength and Weakness: NVIDIA's architecture (GPUs) is optimized for graphics but now poses challenges for inference tasks due to off-chip memory limitations.
  3. Market Dynamics: Feldman argues that NVIDIA's previous advantages are becoming weaknesses as AI demands shift.
  1. Cost of Inference
  2. Efficiency Issues: Current AI algorithms are underutilized, with GPUs operating at low efficiency (5-7%).
  3. Projected Changes: Anticipates an increase in efficiency and decrease in inference costs as architecture and algorithms improve.
  1. Scaling Laws in AI
  2. Debate on Scaling: Feldman refutes the idea that scaling laws have been reached, arguing there's still significant room for improvement in AI algorithms and architectures.
  3. Future Predictions: Believes that reliance on transformers will diminish in the next few years.
  1. Synthetic vs. Human Data
  2. Future Trends: Predicts a move towards predominantly synthetic data usage in AI training.
  3. Importance of Quality: Emphasizes the necessity for synthetic data to cover rare and challenging scenarios in training AI models.
  1. NVIDIA's Market Position Over Time
  2. Anticipated Changes: Suggests NVIDIA will see a decline in its market share as new players like Cerebras emerge with more efficient architectures.
  3. Future Competitors: Predicts that the next five years will see significant growth for competitors in the inference space.
  1. International Perspectives
  2. China’s Capabilities: Feldman warns against underestimating China's investment in AI technologies and infrastructure.
  3. Ethical Considerations: Discusses the decision not to sell technology to China based on ethical grounds and potential misuse.
  1. AI's Broader Impact
  2. Industry Growth: Anticipates AI's growing integration into everyday workflows, leading to increased demand for efficient inference solutions.
  3. Environmental Concerns: Acknowledges the power-intensive nature of AI and the industry's responsibility to deliver value while managing energy consumption.

Quickfire Round Highlights

  • Feldman shares quick insights on various topics, including the future of AI and what he believes as the most underrated threats to NVIDIA's market position.

Conclusion Harry Stebbings wraps up the episode, highlighting the insightful conversation with Andrew Feldman, emphasizing the excitement surrounding the evolving AI landscape and the competitive dynamics that will shape the future of AI technologies.

Key Takeaways

  • Technological Innovation: Advances in chip architecture are vital for overcoming current limitations in AI algorithm efficiency.
  • Market Dynamics: Expect significant shifts in AI market shares as competitors emerge and capabilities improve.
  • Importance of Ethical Practices: Conscious decisions about where to sell technology can impact the global landscape.

For more details, visit [20VC](http://www.20vc.com).

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Our AI algorithms today are not particularly efficient. In a GPU, most of the time it's doing inference, it's five or seven percent utilized. That means it's 95 or 93 percent wasted. We won't be as dependent on transformers in three years or five years as we are now, 100 percent. The fundamental architecture of the GPU with off -chip memory is not great for inference. Now, they will continue to do well in inference, but it can be beaten and I think they know it. This is 20VC with me Harry Stubbins. Now we did a show with Jonathan Ross at Grock and it blew all numbers out of the water. Millions of plays, everyone loved it and everyone said that we had to get Andrew Feldman from Cerebrus on the show.

0:44So I'm so excited to make this episode happen today. Joining us in the hot seat is Andrew Feldman, co -found and CEO of Cerebrus, the fastest AI inference and training platform in the world. In September 2024, the company filed a go public off the back of a rumoured $1 billion deal with G42 in the UAE. They challenge Nvidia in the inference market. Andrew is the leading expert for all things inference. This show was incredible. I have the best job in the world. I sit down with the smartest people and learn from them and this show is exactly that. But before we dive in today, turning your back of a napkin idea into a billion dollar startup requires countless hours of collaboration and teamwork.

1:25It can be really difficult to build a team that's aligned on everything from values to workflow. But that's exactly what Coda was made to do. Coda is an all -in -one collaborative workspace that started as a napkin sketch. Now, just five years since launching in beta, Coda has helped 50 ,000 teams all over the world get on the same page. Now at 20VC, we've used Coda to bring structure to our content planning and episode prep, and it's made a huge difference. Instead of bouncing between different tools, we can keep everything from guest research to scheduling and notes all in one place, which saves us so much time.

2:01With Cody, you get the flexibility of docs, the structure of spreadsheets and the power of applications, all built for enterprise, and it's got the intelligence of AI which makes it even more awesome. If you restart up team looking to increase alignment and agility, Cody can help you move from planning to execution in record time. To try it for yourself, go to coder .io slash 2 -0 VC today and get 6 free months of the team plan for startups. That's coder .io slash 2 -0 VC to get started for free and get 6 free months of the team plan. Now that your team is aligned in collaborating, let's tackle those messy expense reports.

2:38You know, those receipts that seem to multiply like rabbits in your wallet, the endless email chains asking, can you approve this? Don't even get me started on a month and planet when you realise you have to reconcile it all. Or Plio offers smart company cards, physical, virtual and vendor specific, so teams can buy what they need while finance stays in control. Automate your expense reports, process invoices seamlessly and manage reimbursements effortlessly, all in one platform. With integrations to tools like Zero, QuickBooks and NetSuite, Plio fits right into your workflow, saving time and giving you full visibility over every entity, payment and subscription.

3:16Join over 37 ,000 companies already using PLEO to streamline their finances. Try PLEO today. It's like magic, but with fewer rabbits, find out more at pleo .io -4 -slash20vc. And don't forget to revolutionize how your team works together. Rome, a company of tomorrow runs at hyperspeed with quick drop -in meetings. A company of tomorrow is globally distributed and fully digitized. A company of tomorrow instantly connects human and AI workers. A company of tomorrow is in a Rome virtual office. See a visualization of your whole company, the live presence, the drop in meetings, the AI summaries, the chats.

3:54It's an incredible view to see. Rome is a breakthrough workplace experience loved by over 500 companies of tomorrow for a fraction of the cost of Zoom and Slack. Visit Rome. That's O -R .am for an instant demo of Rome today. Nobody knows what the future holds, but I do know this. It's going to be built in a Rome virtual office, hopefully by you. That's Rome .ro .am for an instant demo. You have now arrived at your destination. Andrew, it is such a pleasure to meet, man. I've wanted to do this one for a while. I've had so many good things from Eric for a long time, so thank you so much for joining me.

4:30Harry, thank you for having me. I appreciate it. Not at all. This will be a fantastic conversation. I have my pen ready. I feel like this is going to be a learning experience for me. I want to go back to 2015. What did you and the team see in the AI landscape in 2015 that led to the founding of Sir Rebus? We saw the rise of a new workload. This is every computer architect's stream. We saw a new problem to solve. What that means is maybe you can build a new machine better suited to that problem. And so in 2015 and the credit goes to Gary and Sean and JP and Michael, my co -founders, they saw on the horizon the rise of AI.

5:10And what that meant was there'd be a new problem for computers. That what the AI software would ask from the underlying chip processor would be different. We came to believe that we could build a better machine for that problem. That's what we saw. Obviously, we didn't see it exactly right. I underestimated it. This is my fifth startup and the first time I underestimated the size of the market by a lot. But what we did get right was that this was going to be big and it would put a different type of pressure on a processor and that it would put pressure on the memory bandwidth, that it would put pressure on the communication structure.

5:52That's what we saw. We dove in. It's been an extraordinary nine years. How does the movement into an age of AI change the requirements from a chip perspective of what is needed for a provider and how that then resulted in how you built syrepers? The way to think about a chip is it does two things. It does calculations and it moves data. This is what a chip does. Sometimes along the way it stores data. And so what AI presented was a very unusual combination of challenges. First, the underlying calculation is trivial. It's a matrix multiplication. And an F -MAC can be developed by any second -year electrical engineering student.

6:37So you say to yourself, Holy cow, this has a huge number of very, very simple calculations. The hard part with AI work is results and intermediate results have to be moved a lot. Therein is the most complicated part. They have to be moved to memory and from memory and they have to be broken up and moved among GPUs. And what we saw was that this was going to be the hard problem and that if we could solve for that problem, we would build an AI computer that was faster and use less power. When we think about how we're going to build and what we're building for to me kind of a couple of core elements Which is that where you're going to focus are you focusing on you know fine -tuning are you focusing on training?

7:21Are you focusing on inference? It's three you chose all three? Yeah Why and I'm sorry for my base questions, but I thought like GPUs were specialized towards training and they weren't specialized towards inference, can you have a monoarchitecture that does three best? The first step in computer architecture is deciding what you're not going to do. What are we not going to be good at is really the first important question. To answer your question, you say, is the computational work for training from scratch different from fine -tuning? And the answer is it's not different.

8:04It's and generative inference in particular has some very challenging requirements, on exactly the communication dimension that I mentioned. In generative inference, you have to move all the weights from memory to compute, to generate a single word, and you have to move them again to generate the next word. And again, so if you have a 70 billion parameter model, not a giant model, and each weight is 16 bits, you're moving what 140 gigabytes a data to generate one word. This is an enormous amount of data movement across memory and that's called what it's consumed that needs is memory bandwidth.

8:43If you have an architecture like we saw in the GPU, that is your fundamental limitation. It's a fundamental architectural limitation. That was what we went to a wafer scale to solve. They use memory, a memory called HBM to type a DRAM. It is phenomenal memory, but it's slow and high capacity. And when they set the architecture for graphics, that's what you wanted. You didn't have to go back and forth to memory very often. S -RAM, on the other hand, is unbelievably fast, but has low capacity. And so we wanted to use S -RAM, but if you build a normal size chip, you can't hold a model. And so by going to wafer scale, we were able to put down a huge amount of S -RAM and get the benefits of speed and enough capacity.

9:30If you build a normal size chip with SRAM and you want to do a 400 billion parameter model in inference, you might need 4 ,000 chips. Or if you want to do a deep -seek 671, you might need 6 or 8 ,000 chips. What an administrative nightmare. If you can keep it as much as you can on one way for two way for your four or ten, you get all the benefit of the SRAM. And because you've been able to use the wafer, you get this tremendous capacity as well. Can I ask you first, I totally get you on HB app and kind of the slow and the soviet. Why is it then that bluntly so much of the market just continues to use it and 40 % of Nvidia's revenue is using that ships for inference.

10:13Unless you went to wafer scale there wasn't really a credible other choice. This is the way GPUs had always been made. It's called a graphics process in unit. That's the way they were built. It was part of their advantage against a CPU, was they were built this way. But now they're dedicated chips like ours. What used to be their advantage is now their weakness. That's a fun market to be in when over a very short period of time, which good ad becomes your weakness. With a market cap like they do, and with Jansen is good as he is, which I'm sure we both agree with, they must know, they do know this, A, they don't make memory.

10:47So they're a consumer of other people's memory. And that's SK, the high -news guys, or Samsung, I mean, micron, they're only three or four or five companies that make huge amounts of memory, not many choices. But it's part of a complex architectural trade -off. The flip side, you could say it's worked really well for them, right? Look at where it's taken them. But in comparison to those of us who do wafer scale, it's a small set. It's a set of one. Us. We have real advantage against them on inference. How do LPs fit into this? We've got HBM, we've got Ashram with you, and I bought me having many more of them to make it work in scale.

11:25What LPs fit into this mix? In our business, there are a lot of ways to skin a cat. Our way is different than in videos way, it's different than the TPU, it's different from Trainium, it's different. Right now, and every day since August 26th, when we launched inference, our way has been the fastest way across a whole set of models tested by artificial analysis and others. Can I ask when we think about that speed? I am, and you said that you're one of one with wafer and the architecture associated. What does that mean in terms of cost? With such efficiency, is it inherently more expensive? And what does that look like from a cost profile?

12:02This isn't our first dance. We've been building computers for a long time. When you make a choice like wafer scale, you have to weigh the trade -offs. We use less power because one of the most power hungry things on a chip are the IOs are moving data off chip And so if you are moving data off chip frequently you're using more power And if you can keep it in the silicon domain on chip so we knew we would use less power We knew if you went to wafer scale that you had to solve some problems that people said were impossible to solve like yield So we had to invent techniques that allowed us to yield wafer In fact, we invented techniques that allow us to yield as well or better than others who are building much smaller shows.

12:46What is yield and why is it impossible to solve? A wafer begins 12 inch diameter circle, slice of silicon. And your chip is punched out of this the way your mother might take a cookie cutter and cut out cookie dough. During the process, at some point, just like your mom might have done, she lifts up the edges and all the little bits are removed and what's left are just the cookies. Those are your chips. Now what happens is there are a set of naturally occurring flaws and that's like your mother closing her eyes and throwing up a handful of M &Ms. Now the bigger the cookie, the higher probability you hit in M &M, the bigger the chip, the higher the possibility that you have a flaw.

13:32And traditionally what you did when you had a flaw was you threw away the chip or you sold it as a less valuable part. You shut down part of the chip and sold it as a less valuable part, something called bending. So every wafer is going to have flaws, the bigger your chip, the higher probability you hit a flaw and the more part of silicon is wasted when you throw it away. This is what everybody thought was known truth. And one of the things our team realized was that there are other ways to handle flaws. What if instead you built your computer, you built your processor out of hundreds of thousands of identical tiles and say there was a flaw, say you just shut down that tile and worked around it, say you had a row or a column of redundant tiles that when you need it, you could just pull in.

14:21Now that had been traditionally the technique used in memory making and the memory yields are extraordinary. And so it occurred to us that if we could build a computer, build a processor, build to hundreds of thousands of identical tiles, we could use redundancy such that when there was a flaw, we just leave it there, shut it down, work around it, and pull in one of the redundant tiles. And that had never been done in a computer before and that's at the heart of our architecture. That allowed us to yield and deliver whole wafers. Nobody ever been able to do that in the 70 year history of our industry.

14:56Really, really smart people struggled. I mean, Gene Amdol, one of the fathers of our industry, had a company called Trilogy that crashed and burned trying to do this. And we figured it out. When you speak about being the fastest and across all benchmarks being the fastest, what matters the most? Is it being the fastest? Is it being the most efficient? Is it being the least costly? How do you think about the stack of prioritization for your customers? I think it varies. If you go to get a cancer diagnosis, God forbid your mother or your wife, I think 93 % accuracy is just plain not as good as 94 % accuracy.

15:34And you pay a lot and wait another week to understand what the accuracy is, right? You pay a lot. Now on the other hand, if you want Lama 405B to generate data to help you tune Lama 70B, maybe you can wait a few days, three days a week more. There's no urgency there. On the other hand, if you want an answer from Proplexity, you don't want to wait 45 seconds for a certain answer. You don't want to wait in a chat. You don't want to wait three minutes for R1 on GPUs to give you an answer. What we know is that in interactive mode, milliseconds matter. In interactive mode, what Oors holds over a Google years ago showed was that you can destroy your user's attention with milliseconds of delay.

16:18So being the fastest matters everything in that domain. So I think what you have to do is sort of be thoughtful and say in some cases being the fastest doesn't matter. We'll call those batch. Lots of maybe the cheapest matters there. In other domains, there is no search if you got to wait eight minutes to get an answer. That's not a product. When you go fast, the whole set of new opportunities open up. Netflix used to mail DVDs. That's what happened when the internet was slow. they've mailed DVDs. I'm not that young. I remember a blockbuster. Yeah. Well, if you remember blockbuster, right? First, I mean, let's look at the history of that.

16:58You're exactly right. First, we used to drive to blockbuster to get a DVD. Then Netflix was mailing them to us. And then we got Brabef. And they're suddenly Amazon's studio, right? It changed everything. And speed in inference does the same thing. When we sounded before, you gave this great equation inference. What was the equation that you gave for inference? Because it was really helpful for me in understanding. It begins with the following. Training makes AI. That's how we make AI. And inference is how we use or consume AI. And so understanding how big the inference market is, is understanding the number of people who are going to use it, how often they're going to use it, times how much compute each use takes.

17:40And right now we are in this rare time where the number of people using AI is growing, the frequency with which they use it is growing, and the amount of compute used in each instance of use is growing. That's why you're getting this extraordinary growth. And that's why it's after charts right now. When we think about the distribution of resources between training and inference, what will that look like in five years time? Because we've seen all focus go to training on, not all, but a lot of focus go to training and not as much go to inference. What is that like? What we made until the middle of 2024, what we made in AI was a novelty.

18:18It wasn't very useful. Late in 2024, what we made began to be useful. What was the timing point? If you look at the models, they became, I mean, ChatGPT was not really a technical innovation. It was a user interface invention. But it gave more people access. But we didn't really write away, know what to do with it. It was cool, right? That's what I mean by novelty. I was like, whoa, this is cool. Now, if your marketing team isn't on an LLM each person several times a day, they're not doing their jobs. That difference between novelty, it's cool. And this is part of everyday workflow. That's what changed starting some time in Q4 last year and running into this year is AI became useful, not just to a select group in Silicon but to my dad, to my brothers, the doctors, to ordinary people who aren't buried in the Silicon Valley discussion.

19:16And when you get them, then the market is ripping. Do you not still think we are so incredibly early, though, going back to your point of like how many, in five years time, then, where are we? Are we a hundred times bigger? We are a thousand times bigger. In terms of demand. I think we're way over a hundred times bigger. What does that mean in terms of what we need to equip ourselves to deliver? These are incredibly energy utilizing. It is incredibly difficult. Our industry consumes a lot of power. Yeah, and a lot of water. And we're seeing that come down. But are we equipped from an energy and a data center standpoint to deliver the inference requirements for a population that is as AI hungry as we are?

19:57I think a couple things. I think the first thing is to admit that this is a power intensive problem. We consume our industry consumes an enormous amount of power. The second thing to say is therefore the burden is on us to deliver exceptional value as an industry. You take both the good and the bad, right? In order to make it worthwhile from a societal perspective, to expand all this power, you better deliver the goods. We better use AI to find cures for diseases. We better use AI to solve a bunch of different societal problems. That's the macro view. Do I think that we are equipped? I think we are in a very unusual situation in the US where we have plenty of power, but it's in all the wrong places.

20:39We have power in Niagara. What we don't have is power where you want to build data centers where we have good fiber. What we don't have is a national way to relax the local regulations that make getting power. And so when you go to Silicon Valley, if you want to build a data center, you're dealing with local government and installed interests. And that is not an efficient way to decide if you want to build a power plant or put a new data center in, especially if it's large. I think those places that have ripped out some of that burden in Texas, for example, are getting a huge amount of data center.

21:14When I spoke to Jonathan at Grock before, he said there were a huge amount of data centers as being built that were not actually really equipped properly. And then we've seen this massive supply of data centers that have really come done by tourists, so to speak. And that is a massive problem and that the provisioning of these data centers isn't there. Do you agree? A data center is a construction project to begin. It's a access to power. And then it's a construction project and it's got a design engineering component. I think there's been a huge push for new construction data centers. We will see.

21:48We don't know if they're going to be good enough. I think many of them will be fine. The guys who were there early were some of the Bitcoin mining companies, Terroof, the guys of Crusoe and the guys in Europe. They were early in building buildings near low -cost power in order to run compute that used a lot of power. And they are some of the leaders now in some of the largest projects. Now those are certainly not tourists. Those are extremely sophisticated data center builders. Sure, there's some tourists, but there are a lot of very, very knowledgeable data center builders building huge facilities right now.

22:25I mean, gigawatt scale facilities, both domestically and internationally. How do you think about how the cost of inference goes down with the surge of demand that we mentioned, you know, over a hundred ice? Does the price reduce a hundred ice? Does it follow Moore's law continuously? How do we think about the ever -reducing price of inference? The cost of inference is built up of several pieces, right? There's the power and space that is consumed to generate their response. That's a data center cost. That's an op -x idle. Number one. Number two, there's the cost of the computer. We can drive down the cost of the computers with each generation by driving up their performance, et cetera.

23:04The other thing we can do is we can develop more efficient algorithms. Our AI algorithms today are not particularly efficient. There's a tremendous amount of room. In a GPU, most of the time it's doing inference, it's five or seven percent utilized. That means it's 95 or 93 percent wasted. Over time, I think as an industry, we get better at things. We can drive the cost of compute down, we can build more efficient data centers with lower PUEs and our algorithms will get more efficient so that our utilization is on our now cheaper computers or higher, so you get a higher percentage of the maximum number of flaws.

23:42You get more tokens per unit time for the same power. When you look at the inefficiency of the algorithms, as you mentioned that, and what that means for the utilization of the chips, why are people suggesting that we're at scaling laws already? That seems to suggest that there is so much room for improvement. How do you think about what you just said in conjunction with the idea that scaling laws were we're hitting this asymptote point. How do you recognize all this here? I don't think there's a lot of debate among senior ML thinkers that we have tremendous room for algorithmic improvement. I don't think there's a lot of debate there.

24:18There's even debate about whether the scaling laws are over, whether we ran out of mojo to keep making data or gathering data to fill these ever bigger models. But OpenAI's work on O1 shows me that the scaling laws certainly for inference are fully functional, right? The more compute you put on inference, the better answer you get. Many of the leading models are now MOEs. They're not presenting all of the weights to teach token. And that's one way to do it. Present the important stuff, not the unimportant stuff. There are other ways to do it that we will invent and learn over time. But we have human models that aren't all to all connected.

24:54Many of our models today are all to all connected. That's a lot of unnecessary connections. connections that don't produce anything that we still end up doing math over. I'm sorry, what does all to all connected mean? In many of the layers in a neural network, every element is connected to every other one. That's not the way actually the learning happens. Some are more valuable and some are not valuable at all. Imagine you're going to read 50 books, you want to learn something, you can read all 50 books, or you could read three books that are really important, or you could read summaries of the three books that are the most important.

25:31The problem is we don't know which they are at the beginning. And there's a process that you could learn, there's things called dropout and all these other techniques to use sparsity to help solve these problems. We are early in the evolution of AI. Plays right into this point that we'll get better at these algorithms. Transformers aren't the end of the world, right? We'll get better. Better will mean faster, more accurate, and more efficient. That's what's exciting about an ever -changing industry. That's why I'm not in all these other industries that don't change quickly. Same nine years as girls there are today.

26:05But this show is kind of strange to me because I speak to a lot of people and they think about the three pillars and they're like, you know, compute algorithms and data. A lot of the common refrain is that actually we're very far along in all of them and that has been refrain. And when I hear you, it's like, actually, it's very exciting. I think they're wrong. I think they're wrong. I don't think we're very far along. And it's very difficult to say that we are early in an industry, but we're far along on all those underpinnings. I think we are early in all of them. If we just take them on by one, in five years time, how much synthetic versus human data will be used to train models if you would have put a percent on it?

Read the full transcript

26:41Almost all synthetic. And the utility value of synthetic is the same as human. When you teach a pilot to fly in a simulator, there is a lot of potential data that isn't very useful in teaching or heard a fly. They've spent a lot of time going straight doing nothing as a pilot. Now, take off in landings are where you want to spend your time, and that's why we put them in simulators, that's what we have them doing. And in simulators, we can create data where engines blow, where there are a whole set of problems where learning can take place. That's simulated data. And in the same way, as we think about creating data, whether it's for self -driving, whether it's for other forms of AI, what we want is the data that's hard together, right?

27:22Otherwise, we just have a bunch of data of people driving straight on a freeway. Not difficult. We've been able to do that for a decade. What we want is an unprotected left turn in the snow. It's snowing. It's hard to see. You've got an unprotected left turn. That's a difficult thing. And you want that thousands of different ways, millions of different ways. That's where the synthetic data comes along is to use it to fill in the empty parts where it's really expensive or painful to get that type of data. Think of the pilot. You want them spending a huge amount of time on things that are rare in their training.

27:56Same with a surgeon. A huge amount of time on things that are rare. Most of the time, it's carpentry. But their expertise is only when it's rare. Something happens. The unexpected occurs. That's when their metal is shown and we will get better synthetic data by a great deal. I love it. I get it from a consumer perspective and from an expectation perspective. If we move the needle on compute algorithms and data, what does that mean for the experience of AI? Faster and cheaper is the first answer. The second is when things become faster and cheaper, new applications emerge. It's used everywhere, right?

28:32When computers became faster and cheaper, were suddenly they were in cars and then you were in your pocket and then they were in your dishwasher and in your TV. That's what happens. I mean, we 30 years ago, like I need to computer my TV, you kidding me? I need to want in my pocket. Now you've got powerful computers in your pocket, you've got them in your TV, you've got in your kids toys, you've got in the car. That's what happens. Defusion of innovation accelerates when you make things faster and cheaper. This is Jevon's Paradox and such as belief that man. Yeah, I know the VC community, you got a site 19th century English economist.

29:06I'm English. If I'm not allowed to cite an English philosopher, what are my handfuls? You just like, oh, he's fucking VC. He's just being like, oh, traveling's powered off. That's right. It's like, like, make stuff cheaper and faster. There are very few examples in our industry. Actually, none in compute in 50 years, in which by making things cheaper and faster, the market got smaller. Mark always gets better always can ask from an architectural standpoint you mentioned Transformers that is there a world where we move past Transformers there's a world of 100 % Well, we won't be as dependent on Transformers in three years or five years as we are now 100 % They're not the end all be all.

29:45Why is it what were replacing what does that look like? I don't know I don't know whether they're gonna be states of base models I don't know whether they're gonna be other types of models But what I know for sure is that innovation doesn't stop. The transformer has some weaknesses that people are desperate to overcome. There's a quadratic effect in the attention head. There's all sorts of things that could be improved, but it's pretty darn good now. It's the best we have, and that's what you run with. You run with the best you have, and the minute it's not the best you have, you drop it and they were the best you have.

30:16I mean, the number of innovative companies designing models is large. And what DeepSeek showed us is you don't need 5 ,000 people and billions of dollars a year. You can do it 200 smart people. More gear than DeepSeek said they had, but less gear than others had. We vary impressed with DeepSeek and what impressed you most? I think it was a result of focused engineering and that impressed me. It was designed to be better. They weren't confused about being model intellectuals or they weren't confused about whether it was important to break new ground or they were interested in being better. From an invention standpoint, that's a little boring.

30:57But from an engineering standpoint, that was sweet effort. They really built a model that was just playing better at many, many things. And that's cool. I like good engineering projects. Now that they chose to announce it right around Trump's inauguration and the politics of it, that's all a separate matter. And we can talk about that later. But is distillation wrong? No, I don't think distillation is wrong. Is summarization wrong? I'm a VC. Are you kidding me? That's what we do. You summarize. You wouldn't know anything, right? That's exactly right. You're not the same. That's exactly right. No, I don't think distillation is wrong.

31:32And if distillation is wrong, then certainly using people's copyrighted data is wrong. That's the problem. The problem is you got to be a little bit consistent. Well, it's time's been a guess many times, and we hope he will be again, so I'm not going to ask him. I hope he's too. I think neither are wrong, actually. But I think you have to be consistent. Well, the thing was with that bluntly deep sea because open so everything that they did innovate on open A I can learn from and take to I think there are a few examples of an open source anything having this sort of immediate impact that model had I mean that that model had a giant impact in a technical community of really smart people and there are very few examples of other open source software projects that had that type of impact in that amount of time.

32:18And you know, you're in the business of betting on these guys. They ramp up and they, oh look, $10 ,000 ,000, $100 ,000 users. It's now a million users. We better start a company around that, get those grad students. But this had a loud boom in the industry immediately. It's like, whoa. The thing I have to think is the venture masters, where is enduring and defensible value simply? and how do I get in early and build that over time? In hardware. Well, this was my question, mission bike. Well, I mean, you have to be a very small investor at Eric Rissier to do Hobba, to be clear. But on the model side, do you think that there is value when you look at the sheer number of players with relatively comparable models?

33:00To demonstrate enduring value, you need both immediate value and a trajectory for more. I think the problem is in some industries you are capable of demonstrating a leadership position for a short period of time. And then someone else, maybe the next generation, they generate the next and the next generation, the next. I think that ends up in the software being you're competing against other people's release cadences. Your four months ahead, there's six months. If that's really where you are, there's not a lot of value. But if you can stay at the top over years, right, even if you're not the best, even if you're top desi over years, and the people above you are changing constantly.

33:43Very large Silicon Valley companies have been built with not the most compelling technology. It might have started the most compelling technology and then it got to a point where it was good enough. It was easy enough to use. That's when you're at the mature market, but we're a long way from there right now. Now, right now we are in the early phases. You characterize my position exactly, right? Data, compute, algorithm. I think we have a ton of room for improvement on all of them. When we let you said that, compute and hardware, that's where the value is. How does that value distribution shake out?

34:14You know, we've obviously got the 800 pound gorilla that is Nvidia. How do you think about how the distribution of value shakes out in hardware and in compute over the next five years? Historically, one of the barriers to entry was sort of the capital intensity of a project. And in the world of building chips, there's both scarce resources and expertise, and it's very expensive. Historically, it hasn't fit very comfortably in a software company. And the things that modern software companies value are not entirely conducively chip -making. So, when I look down the road, who has endured in much of infrastructure tech, people who build systems, Cisco, Juniper, chip makers have endured.

35:00There's a reason that Apple and Nvidia are among the most valuable companies on earth. There is what they do is hard. That's why it's worth challenging. If it weren't hard, if it wasn't enormous and difficult, why spend time being the underdog and challenging it? A lot of people place to fancibility around and videos can to coodaloc in. To what extent is that real versus hype? In inference, it's not real at all. There's no coodaloc in an inference. None. You can move from open AI on an Nvidia GPU to Cerebrus, to Firework service on something else, to together to perplexity with 10 keystrokes. anybody who actually uses AI knows there's no good locking in it.

35:46I think there is, there was a fundamental effort to disintermediate CUDA. First, by Google with TensorFlow and first by some grad students with cafe and some of these early efforts. But later, by Google with TensorFlow and then Facebook or matter with PyTorch. I think today, most AI is written in PyTorch and you ought to be able to compile it and run it on your hardware. In video, it has many modes. When you are a dominant market share leader, that in itself is a mode. That you're the default solution is a mode. That everybody learns to think about AI in your structures. Those are modes. The software, the compilers are hard, but they're tractable.

36:28I completely agree with you in terms of being the leader is a mode in itself. It's never talked about that way. Would you put open AI in that same, it is the leader. Everyone's mother knows chat GPT. Let's look at Intel, right? Intel has made, until hiring LibBoo prior to that, nearly a decade of catastrophic decisions. And they still own 80 % of the X86 market. 75 % of the market. AMD has worked up to like 25 % or 30%. And after a decade of screwing up, and you ask yourself, that's a moat, right? How big's my moat? I can make a bunch of bad decisions for a decade and only lose 20 % share. But that's extraordinary.

37:10The moat was just unbelievable. We'll see. I mean, I'm a huge fan of lip boos. He's an investor in our company. I wish him well. And I think if anybody can change that company, he can. But I think we rarely talk about what being the market share leader means in terms of a moat in the right context. Because as a challenger, we have to think about it exactly because it's exactly that that we need to bridge for. It's exactly these characteristics of the moat that we need to get over. In five years time, though, is it Uber or is it like AWS and Cloud? And what I mean by that is that Cloud is an interesting market.

37:46Well, like a couple of players, several players have relative segments, 25, 30 % and it's shared relatively even in between them, not exactly, but relatively. Or is it one like Uber where Uber has 90 % lift, does five. And then there's all types of providers with the other five. I think it's going to be between those two. So in five years from now, Nvidia is going to have 60. I think right now they have approximately all of it. I think they will come down over time. Orven Vides usage, what percent will be training versus inference? I think they will continue to have a meaningful business on both sides.

38:21I think they're exceptional at training. They will not roll over and play dead in inference. I think they're a world class company. I mean, they've had one of the great decades of any company in history. right? I mean from 2014 they were worth what 10 billion to where they are right now. It's one of the great decades in corporate history. I don't think they're gonna roll over and oh yeah we're not gonna be in the in the inference market. That's not gonna happen. They're gonna have a meaningful share. But the market's growing and we'll have a piece. I think others will have a piece. There'll be some very big companies made in this 100x growth.

38:55Do you think chip providers will be fall larger than model providers in terms of enterprise value? In the five -year time frame? Yes. How does that prediction change in a different timeline? I think in a shorter timeline, when you price an option, variance and uncertainty increases the options value. If you look at the way Black Sholes works or if you look at any option, price and model, uncertainty is a friend. Variability is a friend of the value of the option. When people are paying these extraordinarily high prices for our model companies right now, I think part of that is is extraordinary uncertainty, is his wild variance.

39:31And so in the shorter run, it might not be the case. But in the longer run, his markets mature as we begin to understand the value of these models, we understand what their businesses look like, what their long term, net profitability looks like. What did Warren Buffett say about markets in the short term, their voting mechanism and the long term, their weighing mechanism? At some point, the weighing kicks in. Usually it's in the public markets and then investors say, which is likely to give me better growth in the future. I mean, you mentioned the world public there. I do want to just hone in on your business.

40:04Your cash flow positive in the world where everyone else literally bleeds cash. Help me understand what do you do to make your cash flow positive when everyone else is bleeding or hammering cash? Traditionally, your gross margins were a measure of your technical differentiation, right? If you're running a negative gross margin business, it speaks for itself. Your your song commodity, your value creation isn't being recognized in the market. And so I think our technology is creating an opportunity for us to maintain margins where some others can't. A lot of your revenue is concentrated to the G42 deal.

40:43To what extent is that a strung through a weakness? It's both. The way you catch three large customers is to catch one first. The way you build three large strategic partners is learn to be a strategic partner. That's a learned skill. We didn't arrive knowing how to be a strategic partner G42. Now that we've worked at it and worked at it some muscle we can replicate, we could be a better partner to any of it does in different companies in the world. What have you learned in the G42 relationship process that may see dialogue good part the way that you learned? We've deployed tens of ex -thlops of compute vastly more than than anybody else that isn't AMD or Nvidia, right?

41:21I mean, I have a huge amount of compute. Our software has been hardened on some of the largest AI clusters in the world. We've gone through the growing pains of increasing manufacturing, 2X and 5X and 2X again, through unbelievable growth in manufacturing. We've worked with our supply chain partners to be sure that they're ready for this extraordinary growth. When you work with a strategic partner of this size, is your organization comes out different on the other side. There are things you've learned and the mistakes you've made. And I hadn't done a big relationship in the Middle East. It was a huge amount to learn.

41:55I think you come out of a much better company and much better prepared to do business with Iberscaler, to do business with another massive partner, to do business with another sovereign that takes real work. And your team has to learn. You said you'd come out, Bada. Why did it populate when you did? when this happened, I was like, it seemed preemptive respectfully. And my question now to companies is why go public at all? There is so much private capital. The callouss have shown, I think, very clearly that you can stay for a lot longer than you plan to. It did, but I certainly shown that, right?

42:29I mean, there, those were historically public market valuations, you know, the valuations that they didn't throw up, they can open AI and some of the others are getting are historically public market only valuations. And like you said, you ask one's life, anyone can read it. I wouldn't want people reading mine. We have nothing to hide, are we, date? No, but you'll compare it as a good asymmetric information. Yeah, we've got asymmetric technology. To be public, you have to be ready organizationally, be ready with your processes. You need to be ready to forecast and predict, to be held accountable in a way that private companies historically haven't been.

43:08We think that there's tremendous value. We think that we will be among the first and the category. We think that some of our largest targets would have a stated preference for doing business with public companies, large enterprises in the US. I've done that historically. Those are some of the reasons that led us to. How many G42 relationships are at G42, where you have in the next 24 months? How fast can you wrap them? That's a good question. Several. Those big numbers. Sorry, remind me. How big is the G 42? It's 87 % of revenue. I know that. It was big. I mean, when we announced it, it was some estimated.

43:48It was north of a billion. All done. That must be a bit of a high five. Isn't that? Look, I think. Come on. No, no, no, no. It's a shame. I'm going to end all that. First. Yeah. There's tremendous excitement. And then there's sort of every entrepreneur's reality is, I gotta make a lot more gear. Right? I need to make it. You make a list of your top 10 vendors and you fly it to them all, say, big orders are coming. Be ready. Right? You work with all your partners to get ready because you need to make a great deal more stuff. And that's one of the real differences between hard ones off where is.

44:28When we grow fast, the number of people you need to work within your supply chain and that the amount of collaboration that needs to happen is truly extraordinary. I don't believe you're going to have a clusterfuck of unhappy customers who bluntly have waited so long for chips. By the time they get them, the chips are outdated and they're going, what? All of that's an opportunity for us and others. That's opportunity. I think being a market share leader isn't easy either, but when you're late, when the bully falls, everybody wants to give them a kick. I mean, that's a lot of that happened at Intel.

45:01They'd been the dominant player and when they fell everybody was Happy to jump in and kick them when they were down. I think there is a real opportunity in The potential for an Nvidia customer and happiness for sure for those of us who are competing with them I mean if you can't get your gear you may as well test somebody else's that's a huge opening I'd even to syribers and use the promo code Harry 20 For your chips today

45:32I'm happy baby. Influencing. No worries. It's fine. I hope if we could do a 20 % take and on the billion deal, I'm happy. That's fine. I know this venture business has been so good to you, Harry. And you got to get shoes for your kids and the like. And yeah, we're happy to donate to the 400 million fund and fees when I have no kids as well. No, no, no. Two and 20s are a rough way to make a living here. Dude, you don't get it, okay? You in hard one. You said about the complexity of hardware. Oh, export controls being implemented properly. Do you think that is a good idea? You know, everyone was going with deep seat.

46:16Wow, how did this happen? They must have stolen chips. How could this be? It turns out that they probably did use chips in Singapore. or I think the following, I think managing software and managing hardware compliance are extremely different things because their vector of diffusion is different. There's different weights. If you sell a server at ways five or six hundred pounds, it arrives on a pallet, you can go visit it. You want to deploy it in Kazakhstan. You can put a data center and you can have somebody from the embassy visit it. Take photos of it once a month. It's not going anywhere. You can't keep track of who uses it and provide logs, and that's much, much harder with software.

46:58And open source is a whole another. That's the first observation. The second is that I got to know the leadership in commerce and the previous administration. I didn't always agree with their policies, but it is a world of unintended consequences. You sought to limit Chinese access to EDA tools, to delay the growth of a Chinese chip market. And so, US venture capitalists back tons of Chinese companies in Shenzhen to build EDA tools. Right. I mean, right. That this is a unbelievably slippery, dynamic, challenging problem. I don't know if it's a tractable problem to delay another nation's progress Yes, on a technical trajectory is an enormously challenging thing.

47:52I certainly came to appreciate just how difficult it was for well -meaning people to predict the impact of policy during the last two years, for sure. Do you think this administration is faster for AI than the prior administration? I don't think there's any doubt that's the case. I think the past administration lined itself up against big tech. That was a mistake. AI is also in a different place, so it's easier to be for it. It's less scary now than it was. We sort of have a better picture of the trajectory, both the risks and the benefits. This administration sort of had the foresight to put in place an AI -ZAR leader, to be a focal point for discussions.

48:36Yeah, I think it's probably net a fair bit better. You said it was very challenging to kind of hinder a nation's development, adoption, and progression of a technology. Respectfully, you chose to not sell to China. Yeah. Why was that? And does that not go against the difficulty in hindering progression? No, I have a very simple rule and I encourage a team to use it. I mean, you don't need a big handbook to help you make good decisions in a company. Just ask yourself would my mother be proud? And would she be proud if I did this? Would she be proud if I explained exactly the situation and would she look at me?

49:13and so I'm proud you're doing this. And I asked myself that, and I came to believe that the deal on the table was wouldn't be used for good, and I wasn't comfortable with that, and I wouldn't have been able to explain to my mother. And that's a moral compass. What do you mean, it wouldn't have been used for good? To do facial recognition, to identify minorities for persecution, build military equipment, to things that I either couldn't see or what I saw I didn't feel right, it's more important than money. Do you think we fundamentally underestimate might the Chinese is capable of these 100 % and it is one of the most obvious and frequent errors in judgment is that you underestimate the other side.

49:52You have to look carefully at what they're doing and their investment in infrastructure has been extraordinary. The rate at which they generate engineering talent is exceptional. The government's ability to have a policy and implement it. That's not a democracy. They weren't designed to have checks and balances there. the funding that flowed into the development of AI technology, that their venture capitalists were backed up by their government. They have national champion companies that they've developed a belt and suspender strategy to sort of make much of the third world dependence on them and their technologies.

50:28I think they're absolutely should not be underestimated. They have a lot of people and we see a tiny fraction of it. They have produced industrial policy that has moved their nation forward. What was most significant do you think? The creation of economic zones like Shenzhen was clearly a visionary move. They knew that their own system was in the way. They created zones that relaxed their own system. Could the US learn from them in that way? We did some of the same things in Trump won administration, right? What do we do? We relaxed our own rules and the development of vaccines. We knew that in this time it would be very difficult to go through the steps that we always go through and we tried to implement some thoughtful workarounds rather.

51:17I think that, you know, why are they committed to trains as a motor transportation and we can't build a decent train system in the US or in California or why we have three different standards for train rails and the rest of the world can build extraordinary high speed trains linking important cities. What are we doing wrong in the building of our infrastructure that our bridges and our freeways are in disarray? Those are our questions we got to ask ourselves when we see other people doing it differently. If you watch a good football team and you say, whoa, that's an interesting office. And you're not thinking to yourself, how could our team learn?

51:54What could we do? Why did that work? What was it about the people they had or the talent or the structure or something that made that a successful series of plays? And what can I take away from that? How can that inspire me to do better? I'm always looking for inspiration in others and competitors and partners. We have some of our partners at G42, I mean the work ethic is unbelievable. It inspires me and the scope of the challenge that are taken inspires me. And I think I'm always looking for that. Andrew, I can talk to you all day. I do want to do a quick fight with you. So I say a short statement.

52:31You ready? Yeah, sure. What do you believe that most around you disbelieve? I think we're closer to peace in the Middle East than people believe. There is a rise of a moderate business focused Arab state that it wasn't there 25 or 30 years ago. If you visit the UAE or Qatar or even KSA, what you see is amazing transformation, a desire to be included in the West in their own way, but also to enjoy the benefits of it. we are closer than people think. What's the most underrated threat to Nvidia's market share dominance? The fundamental architecture of the GPU with off -chip memory is not great for inferences.

53:15Now, they will continue to do well in inference, but it can be beaten and I think they know it. What's a crazy AI prediction you have that most people would call science fiction? Dario'd anthropic sample lived to 150. I don't think we're gonna live to 150. I don't think what at 90 % of our code will be written by machines in this year. But I do think that within a year or two, most people in the US will engage with an AI every single day. In one form or another, whether they know it or not. That AI might be in their mapping program that helps them pick a better out to work. It might be any number of different things within a year or two.

53:56AI's penetration will be approximately the same as telephones. What did you change your mind on in the last 12 months? Many decisions I made turned out to be wrong. What was the most wrong decision? There are two ways you can be wrong. You can actively be wrong or you can fight against what was right. In 2016, JP, my brother, co -founders and chief system architect laid out a plan that would have us doing water cooling and for our systems. And nobody else was doing it. And I fought so hard and I was so wrong. JP was right about year two later, Google announced that the TPUs were gonna be water cooled.

54:33We were first and now Nvidia's only saw water cooled parts. I mean, I was dead wrong and JP was right. Many, many instances when you make a lot of decisions every day where you're wrong. I've been wrong about people. People I thought were pretty good turned out to be extraordinary. People I thought would be extraordinary. We're really smart, but couldn't finish projects and get stuff done. If you're not prepared to be wrong a fair bit, you know, it's not to be making a lot of decisions because it comes with the territory. There's a venture capitalist I'm never wrong, so I didn't really. As a venture capitalist, you're wrong nine times in 10 and everybody forgets as long as you're really right.

55:08And I get a picture of you signing the term sheet with me and then I go, yours is a perfect industry in which nobody cares about the average. On average, you're wrong all the time. and what they care about is the occasional time you're really right. That's what moves a fund. That's different than being a CEO. I think we got to be mostly right most of the time. But if you're making a lot of decisions, you're still making a ton of mistakes. This is your fifth startup. I mean, you are a sucker for punishment, aren't you? I mean, really, I'm five times. I cry, Sandra. Did you not get beaten alive enough?

55:44My question, see, though, is like, I believe in the value of serolentipronatio. I've spoken to many who don't. How do you think about the inherent benefits that you have having done it four times before? I think if you are in a business in which running a business is a benefit, then experience matters a great deal. I think if you are in a business in which you look like your customer, there was a reason why social networks were started by people right out of college or in college which is because dating is top of their mind. And they look like their customers. And that was more important than knowing anything about running a business.

56:23In that environment, it will certainly select for people who are of the demographic that their customers are. They know that backwards and forwards. But if you want to have a business has manufacturing in it, it has a supply chain that has you managing hundreds or thousands of engineers to a timeline, to a schedule. I don't think anybody would turn around your statement and with a straight face say, you know what I'm looking for is an engineering leader with no experience Right, well no, and I don't want somebody who's led a team of four or five hundred who has experienced the challenges of growth What I'm looking for is somebody with no experience.

56:58My eventy is a bonus head Yes, right. I think the people who sell that sometimes are consultants, right? All right. All of my guys have no experience in your industry They're not biased right? Maybe a little bit of experience in the industry would help, right? Come on. Where are people investing today in AI? Across the start, you can choose any part where you're like, why is so much cash going to that part? I'm not saying that company, I don't. I think part of the dynamic in your industry is sometimes money needs to find a home. Some guys have raised really, really big funds and they got to find a home for their money.

57:32And some people don't like to be left out. They're willing to make investments for maybe for some status purposes or other reasons that they don't seem to make sense. There's some underappreciated places of investment, I'd say in the chip world, the sub millawatt, really tiny, tiny little chips that live next to sensors that do inference. These are tiny little things that will only send back useful data is an extremely interesting market. and they will sell enormous volume. Now, it's not a part of the market. I love to play in. I like to build bigger things and sell them to the data center. But I think that part is extremely interesting.

58:11I think they'll be fundamental for robotics. That's an area where extremely underappreciated. Final one. If we think about Sirreebors in 10 years' time, why do you envision the business in 10 years' time if everything goes well? Where are we in business having that conversation? So 10 years ago, in video, it was worth $10 billion. That's a long run in our world right now. I think in three to five years, I would like our technology to have been used to solve two important societal problems. I would like it to be used to have found a therapeutic for an affliction that impacts more than a million people a year.

58:50I would like our inference to be powering a collection of apps that don't exist today. And I would like that a meaningful portion of the population in the US and in Europe inadvertently uses our technology. So use is something that we power and that they don't even know. I think those are things that make me really happy. And I've wanted to make this show happen for a long time. As I said, I had so many good things from Harry for many years. There have been so many requests to have you on the show. My team is just like, just get Andrew on the show Harry. I'm like, okay, okay. I like tweeted it.

59:25Obviously, which is how we got this. But finally, you tweeted it and like 40 people sent me and I'll say, how come you're avoiding Harry? How come he has to go? Tweet it. I was just like, all right. It's calming. It's good. Send me a note. Happy to come on. Really thoughtful questions, Harry. Really thoughtful and interesting, a really fun conversation. So I have wanted to do that show for a while. But frankly, I was just blown away by Andrews' humility, his no BS approach. He was incredible to work within the process, and I just so appreciate his time to stay. If you want to watch the full episode, you can find it on YouTube by searching for 20 VEC, that's 2 -0 VEC on YouTube.

1:00:06But before we leave you today, turning your back of a napkin idea into a billion dollar start -up requires countless hours of collaboration and teamwork. It can be really difficult to build a team that's aligned on everything from values to workflow. But that's exactly what Coda was made to do. Coda is an all -in -one collaborative workspace that started as a napkin sketch. Now, just five years since launching in beta, Coda has helped 50 ,000 teams all over the world get on the same page. Now at 20VC, we've used Coda to bring structure to our content planning and episode prep, and it's made a huge difference.

1:00:41Instead of bouncing between different tools, we can keep everything from guest research to scheduling and notes all in one place, which saves us so much time. With Cody, you get the flexibility of docs, the structure of spreadsheets and the power of applications, all built for enterprise, and has got the intelligence of AI, which makes it even more awesome. If you're a startup team looking to increase alignment and agility, Cody can help you move from planning to execution in record time. To try it for yourself, go to coder .io -20vc today and get six free months of the team plan for startups, that's coder .io, slash 20VC to get started for free and get six free months of the team plan.

1:01:22Now that your team is aligned in collaborating, let's tackle those messy expense reports. You know, those receipts that seem to multiply like rabbits in your wallet, the endless email chains asking, can you approve this? Don't even get me started on the month and planet when you realize you have to reconcile it all. Or Plio offers smart company cards, physical, virtual and vendor -specific, so teams can buy what they need while finance stays in control. Automate your expense reports, process invoices seamlessly, and manage reimbursements effortlessly, all in one platform. With integrations to tools like Zero, QuickBooks, and NetSuite, PLEO fits right into your workflow, saving time and giving you full visibility over every entity, payment, and subscription.

1:02:05Join over 37 ,000 companies already using PLEO to streamline their finances, try Plio today. It's like magic, but with fewer rabbits, find out more at plio .io forward slash 20vc. And don't forget to revolutionize how your team works together. Rome, a company of tomorrow runs at hyper speed, with quick drop -in meetings. A company of tomorrow is globally distributed and fully digitized. A company of tomorrow instantly connects human and AI workers. A company of tomorrow is in a Rome virtual office. See a visualization of your whole company, the live presence, the drop -in meetings, the AI summaries, the chats.

1:02:43It's an incredible view to see. Rome is a breakthrough workplace experience loved by over 500 companies of tomorrow for a fraction of the cost of Zoom and Slack. Visit Rome .OR .AM for an instant demo of Rome today. Nobody knows what the future holds, but I do know this. It's going to be built in a Rome virtual office, hopefully by you. That's RomeRO .AM for an instant demo. As always, I so appreciate all your support and stay tuned for a fantastic episode coming on Wednesday with I think one of the most under -disgust firms in venture capital, lead edge capital and their founder Mitchell.

From the publisher

Andrew Feldman is the Co-Founder and CEO @ Cerebras, the fastest AI inference + training platform in the world. In Sept 2024 the company filed to go public off the back of a rumoured $1BN deal with G42 in the UAE. Andrew is the leading expert for all things inference. 

In Today’s Episode We Discuss:

04:23 Where Was AI Landscape in 2015 When Cerebras Founded

05:57 NVIDIA’s Biggest Strength Has Become Their Biggest Weakness

07:09 What Happens to the Cost of Inference?

08:55 Why Are AI Algorithms So Inefficient?

20:30 Why is it Total BS That We Have Hit Scaling Laws?

23:07 What Will Be the Ratio of Synthetic to Human Data Used in 5 Years?

31:37 What Specifically Was So Impressive About Deepseek?

31:51 Why is Distillation Not Wrong and OpenAI Need to Look in the Mirror?

32:34 Where Will Value Accrue in a World of AI?

34:08 How Will NVIDIA’s Market Position Change Over the Next Five Years?

39:59 Why is the CUDA Lockin for NVIDIA BS? What is Their Weakness?

40:46 Why is Trump Better for Business than Biden?

49:41 Do We Underestimate China in a World of AI?

52:33 What is the Most Underappreciated Segment of AI?

54:00 Quickfire Round

 

More from The Twenty Minute VC (20VC): Venture Capital | Startup Funding | The Pitch

All 521 episodes
20VC: AI Chip Wars: How Cerebras Plans to Topple NVIDIA's DominanceThe Twenty Minute VC (20VC): Venture Capital | Startup Funding | The Pitch · 1 h 3 min
Listen in VO