DeepSeek: America’s Sputnik Moment for AI?

6 Feb 2025 · 43 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Podcast Summary: a16z Podcast - DeepSeek: America’s Sputnik Moment for AI?

Episode Overview In this episode of the a16z Podcast, the hosts delve into the implications of "DeepSeek," a Chinese AI reasoning model that is stirring discussions around its potential to reshape the AI landscape, similar to the impact of Sputnik in the 1950s. The conversation includes insights from a16z General Partner Martin Casado and board partner Steven Sinofsky, both of whom share their experiences from previous technological revolutions.

Key Concepts

  • DeepSeek: A groundbreaking Chinese reasoning model (r1) that has been described as potentially more efficient than current models, with claims of being 45x more efficient and costing around $5.6 million to develop.
  • Sputnik Moment: The term denotes a pivotal moment in innovation and competition, alluding to the launch of Sputnik in 1957 that catalyzed America’s space race efforts. The hosts discuss whether DeepSeek represents a similar wake-up call for the U.S. in the AI domain.
  • Open Source and Licensing: The open-source MIT license of DeepSeek and its reasoning traces may lead to increased proliferation of AI models, contrasting with the more guarded approaches of U.S. companies.
  • AI Geopolitics: The episode touches on the competitive landscape between the U.S. and China, suggesting that U.S. policy should focus on fostering innovation rather than imposing restrictions.

Detailed Discussions

The Rise of DeepSeek

  • Surprise Factor: DeepSeek’s emergence surprised the global AI community, indicating a rapid advancement in capabilities that had gone under the radar.
  • Claims of Efficiency: The claims regarding efficiency and cost raised eyebrows and ignited debates about the true nature of the model's capabilities.

Historical Context

  • Comparison to Internet History: The discussion draws parallels between the current AI landscape and the early days of the internet, suggesting that past lessons should inform present actions.
  • Innovation from Unexpected Places: The episode highlights how major advancements often come from smaller, unexpected sources rather than established giants.

Policy Implications

  • Regulatory Overreach: The hosts criticize current U.S. regulatory policies that restrict AI innovation rather than encourage it, arguing for a shift towards supporting research and development.
  • Need for Agile Policies: Reflecting on past missteps during the internet's rise, the podcast emphasizes the need for flexible, forward-thinking policies that promote technological advancement.

Future of AI Development

  • Distillation and Model Efficiency: The potential for smaller, specialized models to arise from larger models like DeepSeek could democratize access to advanced AI capabilities.
  • Shift Towards Application Layer: The conversation suggests a future where the application layer becomes more important than the model layer, with focus on building apps that leverage AI capabilities effectively.

Key Takeaways

  • A Call to Action: The episode encourages AI founders, researchers, and policymakers to view DeepSeek not just as a competitor, but as a catalyst for action and innovation.
  • Optimism for the Future: Despite the challenges posed by international competition, there is a sense of optimism that the U.S. can maintain its leadership in AI through collaboration and investment in innovation.
  • The Role of Open Source: The permissive licensing of DeepSeek could lead to an explosion of new applications and models, indicating a shift in how AI tools may be developed and utilized.

Resources

  • Steven Sinofsky's Article: [DeepSeek Has Been Inevitable and Here's Why](https://hardcoresoftware.learningbyshipping.com/p/228-deepseek-has-been-inevitable)
  • Alex Rampell's Article: [Why DeepSeek Is a Gift to the American People](https://www.thefp.com/p/why-deepseek-is-a-gift-to-the-american)

Conclusion The episode concludes with reflections on the transformative potential of DeepSeek and other innovations, reinforcing the need for proactive engagement in the AI space to harness its full potential. The hosts emphasize that the future of AI is not just about competing, but about collaborating to ensure that technological progress benefits all.

---

This summary captures the essence of the podcast episode, highlighting key themes and discussions surrounding the rise of DeepSeek and its implications for the AI landscape.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00R1 comes out and it looks pretty good. That's not the best layer to monetize it. In fact, there might not be any money in that layer. I have yet to see the GPT wrapper. Internet is such a great example because there's no way this doesn't play out like the Internet. It's actually a very big step when it comes to the proliferation of this model. It's a good reminder that there are always pockets of people innovating. World Common AT &T did not predict the Internet was going to come out of universities. 2 words have caught the internet by storm. Deep, seek. Specifically, a Chinese reasoning model that seems to rival others at the frontier.

0:40But, that's not all. Alongside there are one model that dropped in late January, came, a fully open source MIT license, a paper outlining its methods that some claim maybe 45 times more efficient than other methods, and a legit $5 .6 million dollar cost, the release of reasoning traces, a follow -on image model and the fact that all of this was released by a hedge fund in China. Since then, there have been so many claims and claims about those claims that many are already referring to this as a Sputnik moment. But if you think about it, the reason that Sputnik, the first satellite launched into lower Earth orbit by Russia in 1957, the reason that Sputnik still matters in 2022 -25 is because America took all the actions that it did in 58, 59, 60, and moon landing speech in 62, all the way up to 1969 when we reached the moon.

1:35Those are the actions that made Sputnik, Sputnik. A wake -up call was responded to. So now that we're here, how should we weather your a casual listener, a founder, a researcher at a top AI lab, or a policymaker, not just react to this message, but act. Joining us to discuss this and tease out the signal from the noise are A16Z General Partner and Pioneer of Software -Defined Networking, Martijn Kassato, plus Steven Sinovsky, longtime Microsoft exec, including being the president of the Windows Division between 2006 and 2012. Steven, by the way, has also been a board partner at A16Z for over a decade and shares his learnings online at Hardcore Software, where he recently wrote a viral article called DeepSeek has been inevitable, and here's why.

2:27Of course, we'll link to that in the show notes. Both Martin and Stephen have been on the front lines of prior computing cycles, from the switching wars to the fiber build out, and have even witnessed the trajectory of companies like Cisco, AOL, AT &T, even WorldComb. So what really drove this DeepSeek frenzy? and more importantly, what should we take away? Have bigger and better frontier models been optimizing for the wrong thing? And where does value and stack a group? Today, we address those questions through the lens of Internet history. I hope you enjoy. As a reminder, the content here is for informational purposes only.

3:06Should not be taken as legal, business, tax, or investment advice or be used to evaluate any investment or security and is not directed at any investors or potential investors in any A16Z fund. Please note that A16Z and its affiliates may also maintain investments in the company's discussed in this podcast. For more details including a link to our investments, please see A16Z .com slash Disclosures.

3:33It's been a busy few weeks. I don't know about you guys, my Twitter feed, podcasts, everything deep -seek everywhere, maybe unsurprisingly, but what's your TLDR in terms of what came out and maybe also your take on why it blew up in the way it did because we've seen lots of releases in the last let's say two years since chat GPT. The quick overview of course is out of essentially nowhere a small hedge fund quasi computer science research organization in China releases a whole model. Now those in the no no it didn't just appear there's a year and a half or so build up and they're really good and they're really good.

4:13Nothing was an act, but it appeared to take the whole rest of the world by surprise. And I think there were two big things about it that really caught everybody's attention. One was how did they go from nothing to this thing? And it seems to be a constant factor of compatibility and capabilities with everybody else. And this number got thrown around that it only cost $5 million. Yeah, $6 million. The number is irrelevant. Because it turns out they wrote a paper and they said, hey, we innovated in this particular set of things on training, which even here was like, oh, well, that was pretty clever.

4:53And then because of the weirdness that we don't need to get into of the financial public markets and how this whole thing happened on a Friday, the whole weekend was like everybody whipping themselves into a frenzy so they could wake up money morning and trade away a trillion dollars of market cap, which seems to be a complete overreaction and craziness, but that's not what we're here to talk about. To your point, there's a lot of moving parts here and there's a lot to consider. It's actually a fairly complicated situation. So there has been this view that the traditional one -shot LLMs were starting to maybe ask some talk, like GPT -4.

5:28There hadn't been a big advancement, but then there's this can be this new breath of life and open and release the reasoning model, which is a one and everybody's very excited about that. And so in this grand tapestry we're considering you have all this excitement about a one and of how that's going to drive compute costs and Nvidia and then R1 comes out and it looks pretty good. And then all of a sudden they're saying, well, if you can do it just as cheap as this can actually drive the next wave and so forth. And so there's a lot of build up to O1, which led to the R1 hype. And then I think to your point, people didn't know really what to think about it.

6:01And I agree with you as a total market over correction. By the way, it's also worth pointing out that, in addition to people saying, wow, this is a great model, there's a lot of like theories and rumor around, oh, well, maybe this is the CCP doing a siaop. Maybe it costs a lot more. Maybe this is very intentional. It was right by Chinese New Year. There's just a ton of rumors. Maybe we'll do our best to dissect everything going on. Yeah, maybe let's just do that because two of your points, There was a lot here, right? There was the performance element. There was these quotes around costs. There's the China element.

6:31There's the virality at hit number one in the App Store. There's also shipping speed. I think Martini shared that they released an image model shortly after and then it was released on a Friday. So there's this huge mixture of people reacting, some people who know what they're talking about and some people who don't quite frankly. And so we're like 10 days or so out from this release, which by the way, as both of you said, that was the R1 release. There was the V3 release what two months ago, which was the base model. So now that we're a little bit further out, what's the signal from the noise?

6:59So maybe I'll give you the lens of Chinese people or smart. There's one lens, the lens that I hold, which is trying to have great researchers. DeepSeek is actually released a number of soda models, including V3, which is actually probably a more impressive fetus. It's almost like a chat GPD4. And old by the way, to create one of these chain of thought models, these reasoning models, you need to have a model like that, which they had done and we had known about. All of the contributions that they've done have been in the public literature somewhere, just nobody had really aggregated. So there's a thought that I hold, which is this is a very smart team that has been executing very, very well for a long time in AI.

7:37They are some of the top researchers. The fact that they spent $6 million just in the chain of thought is actually not out of whack, which what Anthropic has now said, they've spent it open. I said that that they spent. And so this is a meaningful contribution from a good team in China. And so it means something and we should respond to it. So some of the outcry is warranted. I do think that we respond to it, but I don't think for the reasons a lot of people are saying. I completely agree with that. And in fact, you also saw the people outside of that team in China sort of piling on to try to make it more intergalactic than it was.

8:08I mean, my favorite old friend of mine, Kai Fu Li comes out on X and says something about this is why I said two years ago, Chinese engineers are better than American engineers. But the truth is, to your point about reaching some asymptotic level of progress. Yeah, like the previous base models, like the GPT lineage, seem to have asymptoted around GPT -4. Right. But what's super interesting about that is that asymptote was true if you looked at it through the lens of the function that everybody was optimized for. is to my view, this kind of crazy, hyper -scaler view of the world, which is we need more compute and more data, more compute, more data, and we're just on that loop.

8:48And a lot of people from the outside were like, well, you are gonna run out of data. And I was just as, you know, a micro -computer person was like, well, at some point, you're gonna end up breaking the problem up to the seven billion endpoints of the world, which will have vastly more compute than you can ever squeeze into one giant nuclear power data center. And so a lot of what they did was sort of a step function change, not necessarily improvement, just a change in the trajectory. Yes, and that to me is the part where the hyperscalers needed to take a deep breath and say, okay, why did we get to where we were?

9:22Well, because you were Google and Meta and OpenAI funded by Microsoft, which all had like billions and billions of dollars. So you obviously saw the problem through the lens of capital and data. And of course, you had English language data, which there's more of than anybody else, so you could keep going. The way I thought of it is, when Microsoft was small, we used to just decide, is it a small problem, a medium problem, or a large problem? And I remember, at one point, we started joking that we lost the ability to understand small and medium problems and solutions. And we only had, like, large, which was just trivial.

9:57And then huge and like ginormous. And our default was ginormous. because we thought, well, we could do it and no one else could, and that's a strategic advantage. And I feel like that's where the AI community in the West, if you will, got just a little carried away. And it was just like every startup that has too much money, the snacks get a little too good. So I've heard two theories of why they were able to do this. One of them is this constraint one that you've said, I think which is actually very true, which is we've just been using this blunt intraminative compute and blunt instrument of all data.

10:30And we just haven't thought about a lot of engineering under constraints. To the second theory I heard, I don't know if it's true, but it's tantalizing, which is the reason B3 is so good is actually because it has access to the Chinese internet as well as the public internet, which is actually an isolated thing. We don't really have access to the internal Chinese internet and we certainly don't train from it as far as I know which they do. So it could be the case both things are true. They could have had a data advantage. They definitely have the internet constraint. Well, even on the data, their starting point is the Chinese internet per se, that has much more structure to it.

11:03It's a much better training set. That's a great point. And in as much as human annotated data is important here. And for chain of thought, you do want experts saying, here's how I would reason about a problem. I mean, this is what this whole chain of thought is. It's basically, what are the reasoning steps? If you want to look at a place to arbitrage, really smart educated people and relatively low cost, it's hard to be China globally. Right? And so they definitely have access to a bunch of, potentially highly educated annotated data, which is very relevant here. And so I happen to be of the belief that this did not come out of nowhere.

11:32It's not a siop. This is a great team taking advantage of what it has. But there are still things that are very significant about it that are worth talking about. For example, the license is very significant. The fact that they decided to release the reasoning steps is very significant. Those are two things that you're not seeing headlines about, right? You're seeing headlines about all the other things that we just talked about. You said the reasoning traces, those were released, which using the comparable O1 were previously not. Right. And then the open source license. So there's two things that are pretty remarkable about deep C -Gar one that have implications on adoption.

12:06We haven't seen a license this permissive recently for a soda model. It's basically MIT license, which is like one page. You can do anything, right? It's like free isn't free beer for real. Yeah, for real for real. And I think at A60C we have one of the large portfolio of AI companies, both at the model layer and at the app layer. And I will say any company at the app layer is using many models. like I have yet to see the GPT wrapper. They're all using a lot of models. They do use open source models in licenses really matter. And so this is definitely gonna result in a lot of proliferation. The second thing is, so a reasoning model actually thinks through the steps of the problem.

12:41And it uses that chain of reasoning or chain of thought to come up with deeper answers. And when OpenAI released a one, they did not release that chain of thought. Now we don't know why they didn't do it. But it just turns out that that chain of thought, If you have access, it allows you to train smaller models very quickly and very cheaply. And that's called distilling. And so it turns out that you can get very, very high quality smaller models by distilling these public models. And the implications are both that this is just more useful for somebody using R1. But also, you get a lot more models that can run on a lot smaller devices.

13:15So you just get more proliferation that way. So it's actually a very big step when it comes to the proliferation of this model. Absolutely. And I think that there's this tendency to peg yourself at, oh, it should just be open, but without really defining it, which I think is important in this case. And I think because of where they came from and that they don't have a business model, that was part of what was unique about this, was it was a hedge fund, like almost a side project, but not really a side project, it has this effect that like, well, we're just going to give the whole thing away. And the rest of the companies are still trying to figure out their revenue models, which I would argue was probably premature, and it starts to look a little to me like, hey, let's charge for a web server.

13:56And it's like the business of serving HTTP, not a great business. And I think everybody just got focused on the first breakthrough, which was the LLM, which if you look back at the internet, what exactly happened was everybody got very focused on monetizing the first part of the internet with through HTML and HTTP. And then along came, I don't know, Microsoft and a bunch of other companies to say that's not the best layer to monetize it. In fact, there might not be any money in that layer. And the real money is going to be in shopping and in plane tickets and in television. And even other companies, AT &T got wound up trying to monetize even lower layer.

14:33But that's not how you're going to get to 7 billion endpoints. And I think that the licensing model really matters because what's going to happen is that there's going to end up being some level of standardization. Now, I don't know where in this stack or in what level, but there is going to be some level of standardization. And the licensing model for the different layers is going to start to matter a lot. Anyone who was around during the internet remembers the battles over the different canoe, the three, the four, the open this license. And it was very well. Right, well, between a dissertation and turns out, even your dissertation, which part of it, and how you released it was a huge issue because it could make or break a whole approach.

15:15And I think that the US industry lost sight of that importance because they got so used to this model of like open just means we're a business and we pick and choose what we throw out there is evidence that we're an open company. Yeah, totally. And I think that view isn't an aligned with how technology has just shown to evolve in an era where there's no cost for distribution. Before when there was a cost for distribution, it turns out the free model was irrelevant because you still couldn't figure out how to get it to anybody. It's got people to use it. Yeah, totally. I do want to take the other side of this because I actually tend to agree with you.

15:48And so what you just said is it could be the case that the models, the wrong place to focus and everybody thinks there's a lot of value in there and so they're playing all these cute games with openness as opposed to distribution. And that could very well be true, but there's another view which is actually the models really are pretty valuable and in particular, the model itself isn't an app. But it could be the case that if you're building an app you need to vertically integrate into the model. It could be the case. And therefore, like if I'm building the next version of chat, GPT, or we just had today deep research, it could be that the apps actually require you to own the model.

16:21And in that case, DeepSeek is less relevant because they're not building apps. And then this means that the impact to the opening of eyes or anthropics are not as great, right? And so I do think that there's this fork that we don't know the answer. Fork number one is, maybe the models do get come on, it's as you need to focus at the app layer. and then the license doesn't matter or the model's really matter up the stack, in which case the whole deep -seek phenomenon really isn't as impactful an event as people are making it. So I'm going to go on that just because I want to say you're right both times.

16:51And the variable is time. And the internet is such a great example because there's no way this doesn't play out like the internet. Like it just has to. And what we saw was for a while building one app seemed like a crazy thing because you had to own windows and you had to own office. But then a new app came along that didn't own any of those and it was search. And so that's why I think a lot of people also because of age and what they lived through immediately jumped to, oh, these LLMs are gonna replace search. But it turns out that's actually gonna be really, really hard because there's a lot of things that search does that the models are bad at, really bad.

17:30And so what's going to happen is a new app is going to emerge. And then when the new app emerges, that's going to get vertically integrated. And the research app is a super good example of that. And then all of a sudden, other apps are going to spring up. Oh, there's Google Maps. And there's Search. And then there's Chrome. And then it goes back and eats the things that it couldn't do before. And I really feel like that's what the trajectory were on. Now, it's still a matter of where and what integrates. But the thing is is that the apps that ended up mattering on the internet literally didn't exist before the internet.

18:05And I think that's what people are losing side of it. Same with mobile, like you were. Everybody is completely right. There were no social apps. Okay fine, I get it. There was geo cities and a bunch of others. But people get so caught up on new thing, it's gonna replace something. There's something so different. There's zero sum and you can think of everything as this spectrum. And when something new comes along, the whole spectrum gets divided up differently, which is what Google said when they bought it rightly. They said, you know what people would do in the internet? They're going to type stuff.

18:35And what are they going to type? They're going to type it, but they're going to type it with other people. Okay, so this is great. So we're actually seeing this happening now, which was someone will come up with a model that does something like in a consumer space. Let's say like text to image. And then it turns out that over time people are like, oh, it's kind of like Canva. Exactly. It's like slowly do the AI native version. Just like the cloud native version of Word. the AI -native version of these kind of existing apps. The reason it's important is because it looks like Canva or it looked like Word or it looked like PowerPoint or it looked like Excel, but what's important is that they're actually different.

19:07Nothing is gonna ever be PowerPoint again. Why? Because PowerPoint, the whole reason for existing was to be able to render something that couldn't ever be rendered before. And so all of the whole product, it's 3 ,000 different formatting commands. Like literally, that's not an unbridled up. It's 3 ,000 ways to earn and nudge and squiggle and color. And actually, it turns out you don't need to do any of that in AI. So the whole product isn't gonna have any of those things. Yeah, exactly. And then it turns out all those things make it really hard to make it multi -user. And so then when Google comes along and starts to bundle up their competitor that's gonna replace it, they're focused on sharing.

19:47So Steven, let me ask you this. You said something really interesting. I'm good, I got it, I did. which is this has to pin out like the internet. And you guys have used examples of different companies, the mobile wave, cloud era, those are things we can learn from. But I just want to probe you, is there something different here? To bring it back to deep seek. This is very important to realize the capabilities of China. It's a very credible player. But I don't think that R1 itself is a standalone is going to have that deep of an impact. But on the internet, so there's actually these parallels when it comes to capital build out that you see in the AI, which is it takes a lot of investment.

20:22And there's a special parallel that Mark and Drisnext would remind me of it, which people don't tend to see as well, which is in the early days of the internet, like the mid to late 90s, a lot of investors, a lot of big money, think banks or sovereigns, they wanted exposure to the internet, but they had no idea how to invest in software companies. Like what are these new software companies? Who are these people? Like they're all private companies. Like, so what did all of them do? They all invested in fiber infrastructure. So we're starting to see this thing again, right? we see a lot of banks and big investors, let's say we want to build that data centers, because they don't know how to invest in startups, like we know how to invest in startups, right?

20:56So on one hand, you can be like, oh, we're going to see all of this kind of capital expenditure and all this capital expenditure is going to go into physical infrastructure, and therefore we're going to have another fiber glut equivalent by the data center glut. So the counter to that point, where I think is different is, at the time of the fiber build out, you've had one company would happen to be cooking its numbers, where it had a ton of debt to build all of this out. When the price of fiber dropped, that company went out of business and that caused the huge issue. You have a much better foundation for the AI wave.

21:27The primary investors are the big three cloud companies that got hundreds of billions of dollars on the balance sheet. Even if all of this goes away, they'll be fine. In video, I can take a price dip, and video will be fine. So I don't think we're heading to the same type of glut and crash that other people have, which is very appealing to draw parallels to the internet for that I don't think is it? Oh, I am completely with you on that. That part of it is gonna look like the amount that Google invested in the, or the 2000s, or the amount that Facebook invested five years later, or people forget that Microsoft poured, I don't know, $30, $40 billion into Bing.

22:00And it's still number three or whatever, but it still doesn't matter. Yeah, I would bet, I don't know this is a fact. I'll bet Meta's spending more money on VR than it is on AI right now. Yeah, not just to show. It's an Apple too, right? Well, Apple, also because Apple, whatever is bigger than Gargantuan, is how much they're spending. And so it really isn't about the investing profile. And I think that is a super important point that you made to really just hammer home. There's a certainty that nobody's going to come out of this unscathed, but the scathing is not going to be at all when anybody thinks.

22:31And then not like what it was. Oh, yeah. Well, come on. I believe he had $40 billion in debt, right? I mean, it was just one of these things where structurally it was. Oh, and there were companies that we've all forgotten about. That went bankrupt over that era. Actually, there was one in Seattle whose name I'm forgetting, but that was like $20 billion, just poor God. To your point, these companies have had so much cash on their balance sheet. They've been waiting for a moment to invest in the next generation, which also contributes to their willingness to scale up as much as they did. So let's talk about that.

22:57In your article, you talk about the difference between scale up and scale out, and the natural tendency in these early parts of the wave to scale up, when really there tends to be a shift towards software basically going to zero, So, Coss, so, Stephen, what do you mean by that? And are we at that change in directory? Now we'll just switch to make sure we're really talking about the technology now, not the finances. But when you're big, you want to double down on being big. And so you start building bigger and bigger and bigger computers that don't distribute the computation elsewhere. So if your IBM, you just say the next mainframe is another mainframe that's even bigger.

Read the full transcript

23:31If your Sun Microsystems, you just keep building bigger and bigger workstations. Then if your digital equipment, bigger and bigger many computers. And by the way, all along, you're just doing more MIPS in the acronym sense than the previous maker for less money. And then the micro computer comes along. And not only did they do like fewer MIPS, but they cost nothing and they were going to be gazillions of them. And so you went from an era when IBM would lease 100 or 500 new mainframes in a year. And Sun might sell 500 ,000 workstations to like, oh, let's sell 10 million computers in a quarter. And I think that scale out where there's less computing, but in many more endpoints is a deep architectural win as well, because it gives more people more control over what happens.

24:17It reduces the cost. So today, you know, the most expensive MIPS you can get are in like a nuclear powered data center with like liquid cooling and blah, blah, blah. Whereas the MIPS on my phone are free and readily available for use. And I think that, to me, has been a blind spot with the model developers. Now they all do it. I mean, I run Lama on my Mac, and the first time you do it, your mind is blown. And then you start to go, well, now that's just how it should happen. And then you look at Apple and their strategy, which the execution hasn't been great, but the idea that all these things will just surfaces, features popping up all over my phone.

24:53And they're not going to cost anything. My data is not going to go anywhere. That's got to be the way that this evolves. Now, will there be some set of features that are only hyper scale cloud inference? Oh, yeah. Just like most data operations happen in the cloud now. But most databases are still on my device. So I'm smiling because this is the story from like a micro computer guy. I'll tell you the story from an internet guy. There's the perfect parallel, which is, do you remember the switching wars? Oh, yeah. Yeah. I mean, so for the longest time you had the telephone networks and they were perfect.

25:26They would converge in milliseconds. They would never drop anything. You got guaranteed quality of service. And here comes five dimes. That's five dimes. And then here comes the internet. You had none of these things like convergence was minutes like it dropped half. It's all the time. You couldn't enforce quality of service. And there was these crazy wars at the time where like why are you doing this internet stuff at silly? We know how to do networking. But what the switching people, the telephone people, didn't get was what happens when you actually have a best effort delivery. and then how it enabled endpoints.

25:55They needed the value to be in the network and they couldn't think that way. And that really brought the internet. And I think the exact same thing is playing out. I actually see it a lot of the times, like people, they look at these models, oh, they hallucinate. Yeah, yeah. Or, oh, they're not correct at these things, but they enable an entirely new set of stuff, like creativity and coding. And it's an entirely white space and it's gonna grow very quickly. And to assume that somehow, they don't fit the old model is irrelevant to where it's gonna go. So what I do is I just slash QLS to hallucinate.

26:26And because like to now explain what happened was I was going to all these meetings in the 90s with all these pocket protector AT &T people who would just show up and they would yell at Bill Gates like QLS QLS and we had to go all look up what QLS was because not only were not using TCP IP but the network we were using never worked because it was like a PC based network and the IBM people. Like the Nebui style. I am talking to a networking genius. I should like the ping of death. But it was just hilarious because they're telling me about QLS. I didn't know what it was. I walked him over to my office and this was like in the winter of 1994.

27:05And I'm like, oh, look, here is a video of the Lily Hammer Winter Olympics playing on my Mac. Yeah, awesome. And it was like literally it was a postage stamp, the size of an iPhone icon. And they were like, well, that's 15 frames a second. I'm like, I know it's usually like five. And like, where's the audio? I said, well, if I want the audio, I just call up this phone number on your system. And then they just laughed at me. And so here we are, of course, all using Netflix on every device all over the world. And I think that they can't understand that these paradigms where like the liabilities either don't matter or just become features.

27:42And of course, that's what gave birth to Cisco. And they just went, well, this is how we've been doing it. And it all works. It only works in our crazy weird universities and in the defense department. And now that's all we do. And I want to tie this back to DeepSeek because the reason we're getting so excited about this is because we've seen things like DeepSeek come out before. And it's not zero sum. It doesn't replace the old thing. Yeah. Right? It is a component of the new thing. And the new thing, we still haven't even envisioned yet. Right? And it's like, the internet is just coming right now.

28:09And our excitement is for the new thing to come. And so when I saw DeepSeek, I'm like, amazing. This is another step to basically AGI in your pocket. These can run on small models. It shows that we're going forward. My reaction was not, oh shit, I need to like short and video or whatever. I think that's actually the wrong answer. Yeah, I mean, I read the, let's short and video blog posts that flew around that whole weekend. And I was like, are you crazy? I'm like, hey, Jensen is a genius. Their company is filled with geniuses. What about the tam just expanded, don't you like? Yeah, it's exactly.

28:39And so it is super exciting. this is the scale out step just happened. And so now you could see everybody doubling down. And to your point that you made earlier that I think is super insightful and really important, is this enabling of specialized models? Because that's what's gonna end up being on your phone. And that's what's gonna enable the app layer to really exist. To me, this is all the equivalent of the browser getting JavaScript. Yes, I do. Because once the browser got JavaScript, then all of a sudden you could do anything you needed without going to some standards body. or building your own browser.

29:12Yeah. And I think that's where we are right now. One follow up there is, if you think about how this progresses to date, I feel like the benchmarks have always been like, which model has the most parameters? How's it doing on this coding test? That isn't representative necessarily like, what device can this fit on? How much does it cost? Do we expect then a different set of benchmarks or things that we're judging these models by? Or should we just be looking at the app layer? Does there need to be some sort of shift that kind of moves us away from bigger, better as you're saying skill up and something that represents skill out.

29:43Of course, I thought all those benchmarks were just silly to begin with. To me, they all seemed like, remember, the benchmark we used to do with browsers was like, how fast it could finish rendering a whole picture. And so, Mark and Jason invented the image tag in the browser. The neat thing that they did in their implementation was progressively render it. And then what that did is empower stopwatches all over the world of magazines to write, who finishes rendering a picture faster. And of course, here we stand today like that's a thing you can measure, even, that's a time and it doesn't matter.

30:12And so I think those will all go away and we're just very quickly to get to what does it actually do. I do think that the measure that's going to start to really matter will depend on the application that people are going after. Take this research stuff that just appeared like this week. Well, it turns out when you're doing research, the metric that matters is truth. And all of a sudden, you're giving footnote links and you're giving sources, because what's really happening under the covers, it's a little bit less of generative and a little bit more of IR. And all of a sudden, vector databases and looking things up and reproducing them matter.

30:48And so now we're probably along the lines of ImageNet and they're going to start to generate thousands and thousands of routine tests that are like, is this true? This is totally an aside, but you reminded me of a kind of a weird historical errada, which is the fact that Andrews made the image tag. So in a way, he's also the grandfather to some AI because Clip, which is an AI model, basically will take an image and describe it. The way it does is using the meta tags and image tags. So he created the meta data to do this. I will say back on the topic of the images, here's one thing I've noticed working with these companies where these models are actually pretty magic by themselves.

31:24right? If you have a big model, you just expose it people use them, right? Which is very different than computers. Like you just put the model out there. Yeah, yeah. The thing is all the other models catch up very quickly because they distill so well. So it's not defensible in a way. And so the companies that are defensible that I've seen is they'll put out a model that's very compelling. And then once the users engage with the model, they find ways to build an app around that actually is attentive, right? So it'll start converging on like PowerPoint. It's more stateful on the requires configuration.

31:53So that tends to be very defensible. And then the applications that use models, they use lots of models and they do fine tune these models on hold buttons. The last two years have been the story of the large model. It really has been and they've been magic. Like people use it, man. People really like them. And the first time you're in Chattichipu, you like this is amazing. And now I think we're in the era of workflow -around models, which are stateful complex systems, right? And also many models. Many models is a great point to build on that. Now, this is what happened with user interface. The whole notion of user interface that IBM put forward was just derived exactly from their greenscreens in their 3270s and they made a shelf of rules on how for the characters of like exactly how the UI should be and this is the f10 button and this is the whatever.

32:37And then it turns out that people are building all sorts of UI frameworks, actually looks exactly like the browser today where there's a zillion frameworks on the endpoint. You pick and choose, you do what you want to do. you can invent a new calendar drop down if you want or not waste your time. It's really up to you. And I do think that aspect of creativity is extremely important to applications. And then for apps to be differentiatable and to also to use a MBA, have a mode, apps are gonna also embrace the enterprise. And for better or worse, one of the lessons that we keep learning is if you wanna get adoption in the enterprise, you're gonna have to do a bunch of work to turn off parts of your app or to filter parts of your app of disabled or whatever it is.

33:17And I think the smartest entrepreneurs are gonna recognize the need for sign -on, single sign -on at the beginning. Our back and that's the so -or like, every time. Every single time, because it turns out that's also a great way to price. It's not super hard. And I think so much dumb stuff has been done about AI and alignment and censorship. And whose point of view is that in all this other stuff that there's now a whole industry that just wants to show up and tell you all the things that they don't want out of AI. And the smartest entrepreneurs are going to actually get ahead of that. And they'll be there to sell because turns out that is actually enormously sticky in the enterprise.

33:57And I think that we're going to see the smart productivity tools embrace that immediately. And it could be even at the most granular level of turn it off for these users or whatever. Well, we had Scott Belsky at Speedrun recently. And to your point, he talked about Adobe and someone said, Well, you have all these licensed images, right, for Firefly. Do consumers really care about that? And he was like, honestly, not really. But you know who does care? It's the enterprise, right? So to your point, those are two different modalities and founders are gonna have to figure that out. But I do wanna touch on, you know, a lot of people are talking about Deepseek as this butnik moment.

34:30And that can be viewed in the lens of geopolitics, US, China. But also, if you think about Sputnik, that wouldn't have been a moment if Kennedy didn't do his moon landing speech if we didn't actually get there. So in other words, it changes more made. And so let's say you're in a boardroom, you're an advisor. I don't want to talk to the board. I want to talk to the US government, right? And so like for me, actually the biggest aha of deep sea because nothing we've talked about right now, the biggest aha of deep sea because how blind our policies have been around AI. Right? They've been so wrong headed.

35:00So our previous policies around AI have been, we can't open source because it'll enable China. We've got to limit our big labs. We've got to put all of this regulation on top of it. And the reason is for safety and all this other stuff. X -borg controls. All the X -borg controls. So we can't enable other countries. X -borg controls on chips. We've talked about putting X -borg controls on software. Weight limits, all of this other stuff. Like that was our entire policy. And for me, the biggest, biggest takeaway, the whole deep seek thing is, that's the wrong way to do policy. China is kind of a lot of very smart people.

35:34They're incredibly capable. great researchers, they can build stuff as well as we can and they can open source it. We did not enable them. They did this even with export controls on chips. So there's basically all of our activity has been for not. And what we should be doing is funding and investing in our research labs and we should be going as fast as we can. And it really is the AI race just like we went through the space race and we need to win. And we have everything that we need to win. The only thing in our way is our own regulatory. Just to build on that, the lesson is not spot next, the lesson is the internet.

36:04What we learned from the internet, which Al Gore famously claimed to have invented the internet, but what he really did was invent the regulation that allowed the internet to flourish. Yes. And they could have looked at the internet and said, oh my god, this is a spot in a moment, and then tried to turn it into what AT &T and WorldCom wanted. And they were there lobbying, trying to make that happen. And frankly, AOL wanted it to happen that way too. And so they ignored that, and they went with what made the internet strong to begin with. And so what gave us this deep -seek moment was the strength of the worldwide technology community.

36:36And so as much as people want to own it and be the singular provider, it's not gonna work. The biggest difference not to over -analyze the analogy. I think it's a sputnik moment in the sense that it's a wake -up call for half the world. It isn't a geopolitical wake -up call, it's not about war. It's literally just about technology diffusion. And we've had so many misfires since then. I mean, we had the whole encryption war where we tried to put export controls on encryption and all this. And, you know, although people thought we were being silly as an industry when many of us would champion this, well, you can't.

37:09It's like outlying math. It turns out it is out of that law and math. And the fact that it used those chips, well, the world's economy that was we've seen is very, very hard to put export controls on things. Remember when we were going to export control play stations? Oh, yeah. It was. No, we Xbox like the government came to it. or like actually 2048 bit encryption in email. Yes. Because people came to them, well, we can't have bad actors. That's their favorite phrase. Bad actors encrypting their email. Like, well, they're just gonna encrypt the attachment themselves. And then there's nothing we can do about that.

37:42For sure. But in this case, we've actually put export controls on GPUs before. I mean, like a perfect analog. We're like, oh, listen, you can do weapon simulation on these things. Like a PlayStation was the first to actually use the SGI. Right, right, right. You'd remember that. We're gonna export control that. We can't let that into Saddam Hussein's hands, the whole thing, total failure, because it just turns out global markets are global markets, and we're much, much better in investing, which at the time we did, in our own infrastructure, we did a great job of that. And I think it was great analogy with the internet and with Al Gore.

38:09We should be doing exactly that again. And some politician needs to stand up and be the Al Gore of this moment. I think that we will get that. So I do think that there is now a wake up call. I think that the futility of the past four or five years of this kind of stuff is now very, very clear. And I mean that even more broadly than you were saying, like I mean, like the people who wanted to control this technology at this very granular level in all these think tanks and institutes that were all aligned. I mean, the number of books written, the number of academic departments started, the number of assault on technology companies to align.

38:43I mean, whole meetings in Switzerland about aligning, you know, with the world leaders. That's just not how anything evolves. And if the biggest lesson for computing starting in 1981 with the IBM PC or, frankly, 1977 with the Apple, has been the creativity at the edge, just in enabling that. And I think the problem that the regulators had was they had never faced regulating a connected world before. And I think the other lesson from DeepSeek is just, okay, the world is already connected, the world is already native in all of this stuff. So now, the amount of actual calendar time it takes for something to diffuse technically to get to zero.

39:23100 million users. I mean, deep sea, I think was the number I saw this morning is like 35 % of the DAUs of open AI. And that's a giant spike because just all the same people are just trying it out because there's no friction. It takes no time. And so it's so unbelievably exciting to be part of what's going on right now. And we just don't need to throw water on it and be party poopers. So one thing I will say is I personally don't think this is a crisis moment for open AI or I think like apps are hard to build. I think that like right now the apps that they put out are very complex. They actually know their users.

39:58They have very specific use cases. And so I mean, I think for them it's a bit of a wake -up call that they can't Slouch. They got to move very quickly, but I'm still very, very bullish on our labs. I think they can stay ahead too. So again, there is this view of deep seek as a crisis moment for Nvidia, a crisis moment for OpenAI and Anthropic. I don't buy any of that. I think it's more of like a wake -up call for the regulatory environment. And then listen, we should all acknowledge that. Listen, there's going to be global competition. We need to stay ahead. I would also say that what we should see now, the right reaction from all of these frontier folks, is they should all just start building apps.

40:31Because the best feedback loop to build a great platform for other people to use is to be building apps. And there's this whole concentrated conversation over competing with your partners or whatever. Our industry is co -opitation through and through. It's Andy Groves' lesson. So just everybody should be prepared for these big players to compete with you. But history has shown that's no surefire success. And Tam agrees with 10X. There's just a lot of room for a lot of phone. Yeah, I mean Microsoft spent 10 plus years like a distant number three in the applications business. And it was a platform shift that all the other players ignored that caused it to win.

41:09And so I think that the Tam is going to be 100X. it's going to be every endpoint, the revenue is going to come from the apps side of it. And then there'll be a developer side of it. It'll just be a different pricing model for different sets of scenarios, but it's going to be there. So everything is rising right now. Since it is this positive, some growing world, do you have any thoughts just real quick on the fact that this came from an algorithmic hedge fund, a quant? Is that any different to your expectation or does that actually signal that more can participate? It's a good reminder that there are always pockets of people innovating.

41:43World -com and AT &T did not predict the internet was going to come out of universities. They did not think that a physics lab in Switzerland was going to invent the protocols that become foundational. That's so true. And they certainly, and they also didn't expect a failed corporate lab to develop TCP -IP that became the standard. I mean, it wasn't like the IBM lab. It was like literally a lab that they'd all but shut down because it failed just down on the street at Park. And so, you remember like SRI was like, all these places that you don't even think about anywhere. Right, and so most of this isn't gonna be even in any history that's written in five years.

42:19And I think that that is the exciting.

42:24All right, that is all for today. If you did make it this far, first of all, thank you. We put a lot of thought into each of these episodes, whether it's guests, the calendar, Tetris, the cycles, where they're amazing editor Tommy, until the music is just right. So if you'd like what we put together, consider dropping us a line at ratethispodcast .com slash a16z. And let us know what your favorite episode is. It'll make my day, and I'm sure Tommy's too. We'll catch you on the flip side.

From the publisher

Two words have caught the Internet by storm. DeepSeek. 

The Chinese reasoning model r1 is rivaling others at the frontier with an open-source MIT license, methods that some claim may be 45x more efficient, an alleged $5.6m cost, the release of reasoning traces, a follow-on image model, and the fact that all of this was released by a hedge fund China.

Many are already referring to this as a Sputnik moment. If that’s true, how should we – whether founder, researcher, policy maker – not just react, but act? Joining us to tease out the signal from the noise are a16z General Partner Martin Casado and a16z board partner, Steven Sinofsky. Both Martin and Steven have been on the frontlines of prior computing cycles, from the switching wars to the fiber buildout, and have witnessed the trajectories of companies like Cisco to AOL to ATT – even Worldcom.

So what really drove this DeepSeek frenzy and more importantly what should we take away? Today, we answer that question through the lens of Internet history.

 

Resources:

 

Stay Updated: 

Let us know what you think: https://ratethispodcast.com/a16z

Find a16z on Twitter: https://twitter.com/a16z

Find a16z on LinkedIn: https://www.linkedin.com/company/a16z

Subscribe on your favorite podcast app: https://a16z.simplecast.com/

Follow our host: https://twitter.com/stephsmithio

Please note that the content here is for informational purposes only; should NOT be taken as legal, business, tax, or investment advice or be used to evaluate any investment or security; and is not directed at any investors or potential investors in any a16z fund. a16z and its affiliates may maintain investments in the companies discussed. For more details please see a16z.com/disclosures.

Stay Updated:

Find a16z on X

Find a16z on LinkedIn

Listen to the a16z Podcast on Spotify

Listen to the a16z Podcast on Apple Podcasts

Follow our host: https://twitter.com/eriktorenberg

 

Please note that the content here is for informational purposes only; should NOT be taken as legal, business, tax, or investment advice or be used to evaluate any investment or security; and is not directed at any investors or potential investors in any a16z fund. a16z and its affiliates may maintain investments in the companies discussed. For more details please see a16z.com/disclosures.


Hosted by Simplecast, an AdsWizz company. See pcm.adswizz.com for information about our collection and use of personal data for advertising.

More from The a16z Show

All 489 episodes
DeepSeek: America’s Sputnik Moment for AI?The a16z Show · 43 min
Listen in VO