The AI Model That Tanked the Stock Market

28 Jan 2025 · 21 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Podcast Summary: Odd Lots - The AI Model That Tanked the Stock Market

Episode Overview In this emergency session of the Odd Lots podcast, hosts Joe Weisenthal and Tracy Alloway discuss the significant stock market decline triggered by the introduction of the Chinese AI model, DeepSeek. The episode features insights from AI expert Zvi Mowshowitz, who provides clarity on the implications of this development for the tech industry and the broader market.

Key Events

  • Market Reaction: The U.S. stock market experienced a sharp decline, with Nvidia losing $589 billion in market capitalization—the largest one-day wipeout in U.S. market history. Other tech and chipmakers also suffered significant losses.
  • DeepSeek Model: The episode centers around the introduction of DeepSeek, an open-source AI model that has raised concerns regarding competition against established U.S. AI firms like OpenAI and Anthropic.

Discussion Points

Impact of DeepSeek

  • Rapid Market Movement: Despite being announced in December, the real market impact did not manifest until late January, prompting questions about market dynamics.
  • Cost Efficiency: DeepSeek's V3 model was reported to be trained at a remarkably low cost of $5.5 million, which raised eyebrows regarding the feasibility and implications for future AI model development.
  • Open Source Philosophy: DeepSeek positions itself ideologically in favor of open access to AI technologies, contrasting with the more closed approaches of firms like OpenAI.

Expert Insights from Zvi Mowshowitz

  • Training Efficiency: Mowshowitz elaborates on how DeepSeek achieved its low training cost through extensive optimizations, despite needing a significant investment in infrastructure and engineering talent.
  • Consequences of Open Sourcing: The potential risks associated with open-sourcing powerful AI models are discussed, particularly regarding the implications for global AI competition and safety.

Market Dynamics

  • Jevons Paradox: The discussion touches on Jevons Paradox, indicating that increased efficiency in AI models may not lower demand for computational power but rather increase it as capabilities expand.
  • Competitors' Responses: The implications of DeepSeek for major U.S. AI players like OpenAI, Anthropic, Meta, and Google are considered. Each company's position and responses to the emerging competition from DeepSeek are evaluated.

Conclusion The episode concludes with reflections on the rapid evolution of AI technologies and their market implications, underscoring the tension between innovation, competition, and safety in the AI landscape.

Key Takeaways

  • Market Volatility: The introduction of DeepSeek has caused significant disruption in the tech market, highlighting the fragility of investor confidence in AI-related stocks.
  • Open-Source Threats: DeepSeek's success showcases the potential of open-source AI to challenge established corporate giants, reconfiguring the landscape of AI development.
  • Future of AI Models: As the industry evolves, understanding the dynamics of competition, cost, and innovation will be crucial for stakeholders across the tech ecosystem.

Further Reading

  • [AI-Fueled Stock Rally Dealt $1 Trillion Blow by Chinese Upstart](https://bloom.bg/4hdIiV7)
  • [World’s Richest People Lose $108 Billion After DeepSeek Selloff](https://bloom.bg/3PQAFb8)

Podcast Credits

  • Hosts: Joe Weisenthal, Tracy Alloway
  • Guest: Zvi Mowshowitz
  • Production Team: Carmen Rodriguez, Dashiell Bennett, Kale Brooks
  • For More Content: Visit [Odd Lots](https://www.bloomberg.com/oddlots) and join the Discord community for discussions.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Your best bottling plant employs 3 ,300 people. How do you get 3 ,300 people working at peak efficiency? Your best store has reduced waste, water, and energy usage. How do you make every store like your best store? Your best property has every guest raving. How do you make every property like your best property? The answer is Ecolab. Better performance, better outcomes, better impact. Ecolab. Now every location is your best location. Being a small business owner isn't just a career. it's a calling. Chase for Business knows how much heart and effort go into building something of your own. Manage all your business finances, from banking to payments to credit cards, all in one place with their digital tools.

0:44Plus, access online resources designed to help your business thrive. Learn more at chase.com slash business. Chase for Business. Make more of what's yours. The Chase mobile app is available for select mobile devices. Message and data rates may apply. JPMorgan Chase Bank N.A. Member FDIC. Copyright 2025, JPMorgan Chase and Company.

1:09Bloomberg Audio Studios. Podcasts. Radio. News.

1:24Hello and welcome to another episode of the Odd Lots podcast. I'm Joe Weisenthal. And I'm Tracy Alloway. Tracy, the deep seek sell off. That's right. It's pretty deep. Has anyone made that joke yet? We're in deep seek. Yeah. I don't think anyone has made that joke yet. I will say like, you know, it's bad in markets when all the headlines are about standard deviations. Yeah, right. And then and then, you know, it's really bad when you see people start to say it's not a crash. It's a healthy correction. That's the real cope. But just for like real scene setting, you know, we've done some very timely interviews about tech concentration in the market lately and how so much of the market is this big concentrated bet on AI, etc.

2:08Anyway, on Monday, I think people will be listening to this on Tuesday. Markets got clobbered. NVIDIA, one of the big winners as of the time I'm talking about this, 3.30 p.m. on Monday, down 17%. So we're talking major losses really across the tech complex. Basically, it seems to be catalyzed by the introduction of this high-performance, open-source Chinese AI model called DeepSeek. It was born, from what we know, out of a hedge fund. Apparently, it was very cheap to train, very cheap to build. You know, the tech constraints at this point didn't seem to be much of a problem. They may be a problem going forward.

2:44But yes, here is something the entire market betting on a lot of companies making AI and are now concerned about, of course, a cheap Chinese competitor. I just realized, Joe, this is actually your fault, isn't it? Yeah, yeah. Because last week you wrote that you were a deep-seek AI bro, and look what you've done. You've wiped$560 billion off of NVIDIA's market cap. Yeah, my being, my being. That's you. Anyway, one of the interesting questions, though, is that this was sort of announced in a white paper in December. Why did it take until January 27th for really to freak people out? Big questions.

3:18Anyway, let's jump right into it. We really do have the perfect guest. someone who was here for our election eve special a guy who knows all about numbers and ai and quant stuff and he writes a substack that has become for me a daily absolute must read where he writes an extraordinary amount i don't even know how he writes so much on a given day we're going to be speaking with zvi mashevitz he is the author of the don't worry about the vase blog or substack Zvi, you're also a DeepSeek AI bro. You've switched to using that? So I use a wide variety of different AIs. So I will use Claude from Anthropic.

3:55I will use O1 from ChatGPT from OpenAI. I'll use Gemini sometimes and I'll use Perplexity for web searches. But yeah, I'll use R1, the new DeepSeek model for certain types of queries where I want to see how it thinks and like see the logic laid out. And then I can judge like, did that make sense? Do I agree with that? So one of the things that seems to be freaking people out as well as the market is that purportedly this was trained on like a very low cost, something like$5.5 million for DeepSeek V3. Although I've seen people erroneously say that the$5.5 million was for all of its R1 models, and that's not what it says in the technical paper.

4:40It was just for V3. But anyway, oh, I should mention, it also seems like a big chunk of it was built on Lama. So they're sort of piggybacking off of others' investment. But anyway,$5.5 million to train. Is that A, realistic? And then B, do we have any sense of how they were able to do that? So we have a very good sense of exactly what they did because they are unusually open and they gave us technical papers that tell us what they did. They still hid some parts of the process, especially with getting from V3, which was trained for the 5.5 million, to R1, which is the reasoning model for additional millions of dollars, where they tried to make it a little bit harder for us to duplicate it by not sharing their reinforcement learning techniques.

5:22But we shouldn't get over anchored or carried away with the 5.5 million dollar number. It's not that it's not real. It's very real. But in order to get that ability to spend$5.5 million and get the model to pop out, they had to acquire the data. They had to hire the engineers. They had to build their own cluster. They had to over-optimize to the bone their cluster because they're having problems of chip access thanks to our export controls and they're training on H800s. And the way that they did this was they did all these sorts of little optimizations, including just exactly integrating the hardware, the software, everything they were doing.

5:58in order to train as cheaply as possible on 15 trillion tokens and get the same level of performance or close to the same level of performance as other companies have gotten with much, much more compute. But it doesn't mean that you can get your own model for$5.5 million, even though they told you a lot of the information. In total, they're spending hundreds of millions of dollars to get this result. Wait, explain that further. Why does it still take hundreds of millions? And does this mean if it takes hundreds of millions of dollars that the gap between what they're able to do versus the, say, American labs is perhaps not as wide as maybe people think?

6:32Well, what DeepSeek is doing is they have less access to chips. They can't just buy NVIDIA chips the same way that OpenAI or Microsoft or Anthropic can buy NVIDIA chips. So instead, they had to make good use, very, very efficient killer use of the chips that they did have. So they focused on all of these optimizations and all of these ways that they could save on compute. But in order to get there, they had to spend a lot of money to figure out how to do that and to build the infrastructure to do that. And once they knew what to do, it cost them$5.5 million to do it. And they've shared a lot of that information.

7:08And this has dramatically reduced the cost of somebody who wants to follow in their footsteps and train a new model because they've shown the way of many of their optimizations that people didn't realize they could do or didn't realize how to do them that can now very easily be copied. But it does not mean that you are$5.5 million away from your own V3. So the other thing that is freaking people out is the fact that this is open source, right? We all remember the days when open AI was more open and now it's moved to closed source. Why do you think they did that? And like, how big a deal is that?

7:42So this is one of those things where they have a story and you can believe their story or not believe their story, but their story is that they are essentially ideologically in favor of the idea that everyone should have access to the same AI, that AI should be shared with the world, especially that China should help pump out its own ecosystem, and they should help grow all of the AI for the betterment of humanity. And they're going to get artificial general intelligence, and they're going to open source that as well. And this is the main point of DeepSeek. This is why DeepSeek exists. They're disclaiming even having a business model, really.

8:15And they're an outgrowth of a hedge fund, and the hedge fund makes money. And maybe they can just do this if they choose to do that. Or maybe they will end up with a different business model. But it was obviously very concerning from a lot of angles if you open source increasingly capable models because artificial general intelligence means something that's as smart and capable as you and I as a human and perhaps more so. And if you just hand that over in open form to anybody in the world who wants to do anything with it, then we don't know how dangerous that is. But it's existentially risky at some limit to unleash things that are smarter, more capable, more competitive than us that are then going to be free and loose to engage in whatever any human directs them to do.

9:05I have a really dumb question, but I hear people say artificial artificial general intelligence all the time, AGI. What does that actually mean? There is a lot of dispute over exactly what that means. The words are not used consistently, but it stands for artificial general intelligence. Generally, it is understood to mean you can do any task that can be done on a computer, that can be done cognitively only, as well as a human. I mean, most of these things do things much better than me. I don't know how to code. But I get that there are still some things. Maybe they wouldn't be as good as approving some of the are you a human tests.

9:44Everyone's talking about Jevons Paradox. And so we see NVIDIA and Broadcom shares, these chip companies, they're getting crumbled today. And one of the theories like, oh, no, with all these optimizations and so forth, researchers will just use those and they'll still have max demand for compute. And so it won't actually change the ultimate end for compute. How are you thinking about this question? So I'm definitely a Jevons Paradox bro right now from the perspective of this debate. So you don't think it'll have a negative impact and just the amount of compute demanded? The tweet I sent this morning was, NVIDIA down 11 % pre-market on news that his chips are highly useful.

10:22And I believe that what we've shown is that, yes, you can get a lot more, in some sense, out of each NVIDIA chip than you expected. You can get more AI. And if there was a limited amount of stuff to do with AI, and once you did that stuff, you were done, then that would be a different story. But that's very much not the case. As we get further along towards AGI, as these AIs get more capable, we're going to want to use them for more and more things more and more often. And most importantly, the entire revolution of R1 and also OpenAI's O1 is inference time compute. What that means is every time you ask a question, it's going to use more compute, more cycles of GPUs to think for longer, to basically use more tokens or words to figure out what the best possible answer is.

11:09And this scales, not necessarily without limit, but it scales very, very far. So OpenAI's new O3 is capable of thinking for many minutes. It's capable of potentially spending hundreds or even, in theory, thousands of dollars or more on individual query. and if you knock that down by an order of magnitude, that almost certainly gets you to use it more for a given result, not use it less because that is in fact starting to get prohibitive. And over time, if you have the ability to spend remarkably little money and then get things like virtual employees and abilities to answer any question out of the sun, yeah, there's basically unlimited demand to do that or to scale up the quality of the answers as the price drops.

11:51So I basically expect that as fast as NVIDIA can manufacture chips and we can put them into data centers and give them electrical power, people will be happy to buy those chips. At the risk of angering the Jevons Paradox bros, just to push on the NVIDIA point a little bit more. So my understanding of DeepSeek is that one of the reasons it's special is because it doesn't rely on specialized components, custom operators. And so it can work on a variety of GPUs. Is there a scenario where, you know, AI becomes so free and plentiful, which could in theory be good for NVIDIA, but at the same time, because it's easy to run on a bunch of other GPUs, people start using, you know, more like ASIC chips, like customized chips for a specific purpose?

12:41Yes. I mean, in the long run, we will almost certainly see specialized inference chips, whether they're from NVIDIA or they're from someone else. And we will almost certainly see various different advancements. Today's chips are going to be obsolete in a few years. That's how AI works, right? There's all these rapid advancements. But I think NVIDIA is in a very, very good position to take advantage of all of this. I certainly don't think that you'll just use your laptop to run the best AGIs, and therefore we don't have to worry about buying GPUs is a poor position. It's certainly possible that rivals will come up with superior chips.

13:15That's always possible. NVIDIA does not have a monopoly, but NVIDIA certainly seems to be in a dominant position right now. I mean.

13:40How do you make every location like your best location? Your best paper mill has been operating at peak productivity. How do you make every mill like your best mill? Your best data center has optimized every drop of water. How do you make every data center like your best data center? The answer is Ecolab. Better performance, better outcomes, better impact. Ecolab. Now every location is your best location. For enterprise organizations, managing all your food needs is a tall order. But with EasyCater, you get a single workplace food vendor with the tools and resources to make it easy. Giving teams across your organization an easy way to order from a huge variety of restaurants, all on one platform.

14:23All while consolidating your corporate food spend so you can control costs, streamline billing and payment, and simplify reporting. EasyCater, your business tool for food. To learn more, visit easycater.com slash podcast. It seems to me, I mean, I know there's others, but it seems to me in the U.S. there's like three main AI producers and models that people know about. There's OpenAI, there's Claude, and then there's Meta with Llama. And it's worth knowing that Meta is green today, that the stock is actually up as of the time I'm talking about this, 1.1%. And just go through each one real quickly, how the sort of deep seek shock affects them and their viability and where they stand today.

15:08I think the most amazing thing about your question is that you forgot about Google. Oh, yeah, right. Yeah, that's very telling, isn't it? But everyone else has forgotten about Google as well. I know, I never used Gemini. It wasn't that surprising. Gemini Flash Thinking, their version of O1 and R1, got updated a few days ago. And there are many reports that it's actually very good now and potentially competitive. and effectively it's free to use for a lot of people on AI Studio. But nobody I know has taken the time to check and find out how good it is because we've all been too obsessed with being DeepSeek bros.

15:41Google's had its rhetorical lunch eaten over and over and over again. December, OpenAI would come out with advance after advance after advance. Then Google would have advance after advance after advance. And Google's would be seemingly actually, if anything, more impressive. And yet everyone would always just talk about OpenAI. So this is not even new. Something's going on there. So in terms of OpenAI, OpenAI should be very nervous in some sense, of course, because they have the reasoning models. And now their reasoning model has been copied much more effectively than previously. And the competition is a hell of a lot cheaper than what OpenAI is charging.

16:11So it's a direct threat to their business model for obvious reasons. And it looks like their lead in reasoning models is smaller and faster to undo than you would expect. Because if Dipsy can do it, of course, Anthropic and Google can do it and everyone else can do it as well. Anthropic, which produces Claude, has not yet produced their own reasoning model. They clearly are operating under a shortage of compute in some sense. So it's entirely possible that they have chosen not to launch a reasoning model, even though they could, or not focused on training one as quickly as possible until they have addressed this problem.

16:42They're continuously taking investment. We should expect them to solve their problems over time. But they seem like they should be directly concerned because they're less of a directly competitive product in some sense. But also they tend to market to effectively much more aware people. So their people will also know about DeepSeek and they will have a choice to make. If I was Meta, I would be far more worried, especially if I was on their Gen AI team and wanted to keep my job, because Meta's lunch has been eaten massively here, right? Meta with Llama had the best open models and all the best open models were effectively fine tunes of Llama.

17:19And now DeepSeek comes out, and this is absolutely not in any way a fine tune of Lama. This is their own product. And V3 was already blowing everything that Meta had out of the water. R1, there are reports that it's better than their new version that they're training now. It's better than Lava 4, which I would expect to be true. And so there's no point in releasing an inferior open model of everyone on the open model community just being like, why don't I just use DeepSeek? Tracy, it's interesting that as V said, the people who should be nervous are the employees of Meta, not Meta itself, because Meta is up.

17:56And so you got to wonder, it's like, well, maybe they don't, I don't know, maybe they don't need to invest as much in their own open source AI if there's a better one out there. Now the stock is up. Anyway, keep curious. The market has been very strange from my perspective on how it reacts to different things that Meta does. For a while, Meta would announce, we're spending more in AI. We're investing in all these data centers. We're training all of these models. And the market would go, what are you doing? This is another metaverse or something, and we're going to hammer your stock and we're going to drag you down.

18:24And then with the most recent$65 billion announced spend, then meta was up. Presumably, they're going to use it mostly for inference, effectively, in a lot of scenarios, because they had these massive inference costs to want to put AI over Facebook and Instagram. So if anything, I think the market might be speculating that this means that they will know how to train better llamas that are cheaper to operate and their costs will go down and then they'll be in a better position. And that theory isn't crazy. Since we all just collectively remembered Google, I have a question that's sort of been in the back of my mind.

19:01I think Joe has brought this up before as well. But when Google debuted, it took years and years and years for people to sort of catch up to the search function. And actually, no one ever really caught up. So Google has like dominated for years. Why is it when it comes to these chatbots, there aren't like higher, wider boats around these businesses? So one reason is that everyone's training on roughly the same data, meaning the entire internet and all of human knowledge. So it's very hard to get that much of a permanent data edge there unless you're creating synthetic data off of your own models, which is what OpenAI is plausibly doing now.

19:44Another reason is because everybody is scaling as fast as possible and adding zeros to everything on a periodic basis. In calendar time, it doesn't take that long before your rival is going to have access to more compute than you had. And they're copying your techniques more aggressively. There's just a lot less secret sauce. There's only so many algorithms. Fundamentally, everyone is relying on the scaling laws. It's called the bitter lesson. It's the idea that you just scale more. You just use more compute. You just use more data. You just use more parameters. And DeepSeek is saying maybe you can do more optimizations.

20:14You can get around this problem and still get a superior model. But mostly, yeah, there's been a lot of just I can catch up to you by copying what you did. Also, I can see the outputs, right? I can query your model and I can use your model's outputs to actively train my model. And you see this in things like most models that get trained. and you ask them who trained you, and they will often say, oh, I am from OpenAI. The internet has gotten so weird. The internet is so weird. Zvi Masovic, thank you so much for running over to the OddLots and helping us record this emergency pod on the DeepSeek sell-off.

20:53That was fantastic. All right, thank you.

21:07Tracy, I love talking to Zvi. We got to just sort of make him our AI guy. I mean, to be honest, we could probably have him back on again this week because there's going to be stuff happening, right? Maybe we will. And obviously, we could go a lot longer. This is a really exciting story. This is a really exciting story. And things are just getting really weird these days. It is kind of crazy how fast all of this is happening. And then the other thing I would say is just the bitter lesson. Great name for a band. Oh, totally. Totally great. Maybe when we do our AI-themed prog rock band, Tracy, that could be our name.

21:45Yes, let's do that. Okay, shall we leave it there? Let's leave it there. This has been another episode of the Odd Lots podcast. I'm Tracy Alloway. You can follow me at Tracy Alloway. And I'm Jill Weisenthal. You can follow me at The Stalwart. Follow our guest, Zvi Moshavitz. He's at The Zvi. Also, definitely check out his free sub stack. It's a must read for me. Don't worry about the vase. Really great stuff every single day. Follow our producers, Carmen Rodriguez at CarmenArmond, Dashiell Bennett at Dashbot, and Kale Brooks at Kale Brooks. For more OddLots content, go to Bloomberg.com slash OddLots.

22:17We have transcripts, a blog, and a newsletter. And you can chat about all of these topics 24-7 in our Discord, discord.gg slash OddLots. Maybe we'll get Zvi to do a Q &A in there with people. Oh, yeah. I'll ask them. That'd be great. And if you enjoy All Thoughts, if you like it when we roll out these emergency episodes, then please leave us a positive review on your favorite platform. Thanks for listening.

23:11How many vendors does it take to meet all your organization's food needs? Just one. EasyCater, the workplace food platform that lets teams order from a huge variety of restaurants, over 100 ,000 nationwide, all through a single vendor. In addition to all that variety, Easy Cater also gives you full visibility of your organization's food spend with invoicing, centralized reporting, and seamless integration with expense management systems, all on one platform. Easy Cater, your business tool for food. To learn more, visit easycater.com slash podcast. This is Tom Keen inviting you to join me for the Bloomberg Surveillance Podcast.

23:53It's about making you smarter each and every business day. We bring you a recap of what happened overnight in Europe and Asia, the day's economic data, and complete coverage of the U.S. market open. We cover stocks, bonds, commodities, currencies, even crypto, all the information you need to excel. Bloomberg Surveillance also brings you the analysis behind the headlines. We do that with lengthy conversations with our expert guests, the smartest names in economics, finance investment, and international relations. We do all this live each and every weekday that bring you the best analysis in our daily podcast.

Read the full transcript

24:33Search for Bloomberg Surveillance on YouTube, Apple, Spotify, or anywhere else you listen. On the East Coast, listen at lunch, and on the West Coast, when you wake up. That's the Bloomberg Surveillance Podcast with me, Tom Keen, along with Paul Sweeney and Lisa Mateo. Subscribe today wherever you get your podcasts.

From the publisher

On Monday, the stock market tanked, seemingly in reaction to the emergence of DeepSeek, an open source AI model developed in China. Nvidia, the semiconductor giant that has been the largest winner of the AI boom, erased $589 billion in market cap, for the biggest one-day wipeout in US stock-market history. Other chipmakers and big tech giants also swooned. So how did DeepSeek do it? Is it a big threat to the American AI giants like OpenAI and Anthropic? What does this say about export restrictions on US chips? On this special emergency session of the podcast, we spoke with Zvi Mowshowitz, an AI expert who authors the excellent Substack, Don’t Worry About the Vase. He answered all our questions and more to help understand what it means.

Read more: 
AI-Fueled Stock Rally Dealt $1 Trillion Blow by Chinese Upstart
World’s Richest People Lose $108 Billion After DeepSeek Selloff

Only Bloomberg.com subscribers can get the Odd Lots newsletter in their inbox — now delivered every weekday — plus unlimited access to the site and app. Subscribe at bloomberg.com/subscriptions/oddlots

      See omnystudio.com/listener for privacy information.

      More from Odd Lots

      All 683 episodes
      The AI Model That Tanked the Stock MarketOdd Lots · 21 min
      Listen in VO