What The Hell Is DeepSeek?

31 Jan 2025 · 33 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Podcast Summary: Better Offline - Episode: What The Hell Is DeepSeek?

Overview In this episode of *Better Offline*, host Ed Zitron delves into the emergence of DeepSeek, a relatively unknown Chinese AI model developer, which has challenged the existing generative AI landscape dominated by American tech companies. The discussion centers around how DeepSeek's innovative and cost-effective models have sparked chaos within the generative AI sector, prompting a reevaluation of the previously held assumptions about the need for massive investments in AI development.

Key Concepts

Introduction to DeepSeek

  • DeepSeek: A Chinese AI developer incubated in a hedge fund, which has introduced models that not only compete with but also significantly undercut existing models from companies like OpenAI.
  • Market Impact: DeepSeek's efficient models have thrown the U.S. startup scene and markets into disarray, challenging the "bigger is better" narrative prevalent in the AI industry.

Generative AI Bubble

  • Current Landscape: The generative AI industry has been fueled by heavy investments, leading to an expectation that large models require extensive resources.
  • Hubris: The industry has operated under the assumption that expanding resources was the only route to success, which DeepSeek’s entry has called into question.

DeepSeek's Competitive Edge

  • Efficiency: DeepSeek's models are reported to be as much as 30 times cheaper to run than their competitors, such as OpenAI's models.
  • Open Source: Their models are open source, allowing for widespread use without royalty fees, contrasting sharply with the closed-off nature of OpenAI's offerings.

Models Comparison

  • DeepSeek's R1 Model:
  • Competitively priced and designed to run on standard consumer hardware, even a 2021 MacBook Pro.
  • Up to 96% cheaper than OpenAI's offerings.
  • DeepSeek's V3 Model:
  • Comparable to OpenAI's GPT-4O; priced at 0.07 cents per 1 million input tokens, compared to OpenAI's $2.50.

Implications for OpenAI and the Tech Industry

  • Erosion of Competitive Advantage: DeepSeek's models expose the inefficiency of current AI practices, leading to questions about the long-term viability of companies like OpenAI and Anthropic.
  • Crisis for U.S. Companies: The entry of DeepSeek has prompted a panic regarding whether U.S. firms wasted billions on infrastructure that may not be necessary for developing effective AI technologies.

Key Takeaways

  • Disruption: DeepSeek’s innovations represent a significant disruption in the generative AI market, challenging the status quo maintained by prominent U.S. tech firms.
  • Economic Considerations: The conversation raises ethical questions about the sustainability of funding models that have allowed companies to prioritize size over efficiency.
  • Future Outlook: There are concerns about the geopolitical implications of using Chinese models, as well as uncertainties regarding DeepSeek’s funding sources and operational transparency.

Conclusion Ed Zitron emphasizes the need for critical evaluation of the narratives driven by established tech companies and encourages listeners to reconsider the dynamics at play in the AI landscape. The episode serves as a foundational discussion for a two-part exploration of DeepSeek's impact, promising further insights in the follow-up episode.

Additional Resources

  • For further exploration, listeners are encouraged to check out:
  • [Better Offline Links](https://www.tinyurl.com/betterofflinelinks)
  • [Ed Zitron on Twitter](https://twitter.com/edzitron)
  • [Better Offline Reddit](https://www.reddit.com/r/BetterOffline/)

This podcast episode highlights the transformative potential of new entrants in the technology sector and the urgent need for established companies to adapt to an evolving competitive landscape.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00This is an iHeart Podcast.

0:30Thomson Reuters and specialized bikes have since they upgraded to the next generation of the cloud. Oracle Cloud Infrastructure. OCI is the blazing fast platform for your infrastructure, database, application development, and AI needs, where you can run any workload in a high availability, consistently high performance environment, and spend less than you would with other clouds. How is it faster? OCI's block storage gives you more operations per second. Cheaper? Better? OCI costs up to 50 % less for computing, 70 % less for storage, and 80 % less for networking. Better? In test after test, OCI customers report lower latency and higher bandwidth versus other clouds.

1:12This is the cloud built for AI and all your biggest workloads. Right now, with zero commitment, try OCI for free. Head to oracle.com slash strategic. That's oracle.com slash strategic. Run a business and not thinking about podcasting? Think again. More Americans listen to podcasts than ad-supported streaming music from Spotify and Pandora. And as the number one podcaster, iHeart's twice as large as the next two combined. Learn how podcasting can help your business. Call 844-844-iHeart. Did it occur to you that he'd charmed you in any way? Yes, it did. But he was a charming man. It looks like the ingredients of a really grand spy story.

1:53because this ties together the Cold War with the new one. I often ask myself now, did I know the true Jan at all? Listen to Hot Money, Agent of Chaos on the iHeartRadio app, Apple Podcasts or wherever you get your podcasts.

2:16Hello and welcome to Better Offline. I'm your host, Ed Zitron.

2:30A lot of you have been getting in touch. Yes, you're getting your Deep Seek episode. In fact, this is the first of a two-parter. This will come out on Friday, which is when you're listening to this, and then it'll follow up on Monday. I apologize. I spent a lot of Monday writing this and also learning about a lot of this stuff in an attempt to distill it as best I could. This situation is extremely weird, and it's developing. And I think even when I put out this episode, there will be new parts of it that I have yet to really get to. I will do my absolute best to explain in these episodes both what is happening with DeepSeek, what it means, what they've built, and what it's going to do in the future.

3:10But let's begin. So as January came to a close, the entire generative AI industry found itself in a kind of chaos. In short, the recent AI bubble, and in particular the hundreds of billions of dollars being spent on it, hinged on this big idea that we need bigger models, which are both trained and run on bigger and even larger GPUs, almost entirely sold by NVIDIA. And in turn, they're based in bigger and bigger data centers, owned by companies like Microsoft, Oracle, Amazon, and Google. Now, there was also this expectation that this would always be the case. Hubris within this industry is kind of part of the whole deal.

3:49and generative AI was always meant to be this way, at least for the American developers. It was always meant to be energy and compute hungry. Throwing entire zoos worth of animals and boiling lakes was necessary to do this. There was never any other way to do it. And I thought, at least I've thought for a while, that this was because they tried to make them more efficient, but they couldn't. There was just something about transformer-based architecture, like the stuff that underpins ChatGPT, so the GPT model under ChatGPT either. It wasn't the case, though. A Chinese artificial intelligence company that few people had really heard of called DeepSeq came along a few weeks ago with multiple models that aren't merely competitive with open AIs, but actually undercut them in several meaningful ways.

4:33DeepSeq's models are both open source, which means that their source code and research is public, and they're significantly more efficient as well. As much as 30 times cheaper to run in the case of their reasoning model R1, which is competitive with OpenAI's 01, and 15 or more times more efficient than GPT-40. It's actually kind of crazy when you think about it. And as you're going to hear, this whole thing has jokified me all over again. And what's crazy is that some of them can be distilled, which I'll get to later, and run on local devices like a laptop. It's kind of crazy. And as a result, the markets have kind of panicked because the entire narrative of the AI bubble has been that these models have to be expensive because they are the future, and that's why hyperscalers had to burn$200 billion in capital expenditures for infrastructure to support this wonderful boom, and specifically, the ideas of OpenAI and Anthropic.

5:28The idea that there was another way to do this, that in fact we didn't need to spend all this money and that maybe we could find a more efficient way of doing it, well, that would require them to have another idea other than throw as much money at the problem as possible. Yeah, they just didn't consider it, it turns out. And now along has come this outsider that's upended the whole conventional understanding and perhaps even dethroned a member of America's tech royalty. Sam Altman, a man who has crafted, if not a cult of personality, some sort of public image of an unassailable visionary that will lead the vanguard in the biggest technological change since the internet.

6:05Yeah, he's wrong. He never was doing that. I've been saying it for a while. He's never been doing this. But DeepSeek isn't just an outsider. No, they're a company that's emerged as a side project from a tiny, tiny Chinese hedge fund, at least by the standards of hedge funds, like$5.5 billion on assets under management. And their founding team has nowhere near the level of fame and celebrity or even the accolades of Sam Altman. It's distinctly humiliating for everyone involved that isn't DeepSeek. And on top of all of that, DeepSeek's biggest, ugliest insult is that its model, DeepSeek R1, is competitive.

6:41like I said, with OpenAI's incredibly expensive O1 reasoning model, yet significantly, and I mean 96%, cheaper to run. And it can even be run locally, like I said. Speaking to a few developers I know, one was able to run DeepSeq's R1 model on their 2021 MacBook Pro with an M1 chip. That is a four-year-old computer, not a 30 ,000 GPU in sight. It's kind of crazy. Worse still, DeepSeq's models are made freely available to use, with the source code published under the MIT Tech License, along with the research on how they were made, although not the training data, which makes some people say it's not really open source, but for the sake of argument, I'm just going to say open source.

7:21And this means, by the way, that DeepSeq's models can be adapted and used for commercial use without the need for royalties or fees. Anyone can take this and build their own. It's kind of crazy. By contrast, OpenAI is anything but open, and its last LLM to be released under the MIT license was 2019's GPT-2. No, no, wait, wait, shit. Let me correct that. DeepSeq's biggest, ugliest secret is actually that it's obviously taking aim at every element of OpenAI's portfolio. As the company was already dominating headlines this week, it quietly dropped its Janus Pro 7B image generation and analysis model, which the company says outperforms both Stable Diffusion and OpenAI's DALI-3, And those are, by the way, image generation things.

8:04So you type in something like Garfield with boobs, and then out comes a Garfield with juicy cans. And that's probably the first time you'll hear that on the podcast, but probably not the last. And as with its other code, DeepSeek has made this freely available to both commercial and personal users alike, whereas OpenAI is largely paywall DALI 3. This is really, it's a truly crazy situation. And it's also this cynical, vulgar version of David and Goliath, where a tech startup backed by a shadowy Chinese hedge fund with$8 billion under management is somehow the plucky upstart against the lumbering lossy Ofish$150 billion startup backed by multiple public tech companies with a market capitalization of over$3 trillion.

8:47I realize, by the way, I said earlier$5.5 billion under management. This is why you check your notes in advance. But I'm not cutting it. This is fresh. I am inside a closet in New York. the content must flow. Anyway, DeepSeq's V3 model, which is comparable and competitive with both OpenAI's GPT-40 and Anthropik's Claude Sonnet 3.5 models, which, by the way, has some reasoning features, like I said, it's 53 times cheaper to run the R1 when using the company's own cloud services. And as mentioned earlier, said model is effectively free for anyone to use, locally or on their own cloud instances, and can be taken by any commercial enterprise and turned into a product of their own, should they desire to, say, compete with OpenAI, the loudest and most annoying startup of all time.

9:34In essence, DeepSeek, and I'll get into its background and the concerns people might have about its Chinese origins, released two models that perform competitively, and even beat models from both OpenAI and Anthropic, undercut them in price, and then made them open, undermining not just the economics of the biggest generative AI companies, but laying bare exactly how they work. The magic's gone. There's no more voodoo inside Sam Altman's soul. It's all out there. And the last point is extremely important when it comes to OpenAI's reasoning model, which specifically hid its chain of thought for fear of these unsafe thoughts that might manipulate the customer.

10:10And then they added slightly under their breath that the actual reason they did it was a competitive advantage. Now, to explain what that means, when you make a request with OpenAI's O1 model say, give me all the states with the letter R in them. It actually shows you, like, the thinking. And by the way, these things don't fucking think. They're computer bullshit. Like, they don't think at all. But I'm going to use it just for this. So you see it say, okay, here are all the American states. Which ones have that letter? I'm checking all of those. It's effectively having a large language model check a large language model.

10:43Now, the thing is, the steps they were showing you were all cleaned up. They would look nice. They would be formatted nicely. DeepSeq's chain of thought is completely laid bare, which is very interesting because it really takes the wind out of OpenAI's sails. And on top of that, it allows you to see actually how these things think through things. Again, not really thinking. But still, you can see things about how large language models work that these companies didn't want you to have. On top of this, OpenAI's O1 model has something even shittier to it, which is these chain of thought things all cost money.

11:18When you see it generate these thoughts, it's actually generating more thoughts than you see, because they're hiding the chain of thought. So OpenAI is just charging you an indeterminate amount of money, an insane amount of money as I'll get to later, but nevertheless, you don't know what you're being charged for. You don't even know what's really going on under the hood. Or you could use DeepSeq. And let's be completely clear, by the way, OpenAI's literal only competitive advantage against Meta and Anthropic was its reasoning models, O1 and O3. And O3, by the way, is currently in a research preview and is mostly just more of the same.

11:52Although I mentioned earlier in the show that Anthropic's Claude Sonnet 3.5 has some reasoning features, they're comparatively more rudimentary than those in O1 and O3, and I'd argue R1, which is DeepSeq's model. In an AI context, reasoning works by breaking down a prompt into a series of different steps with considerations of different approaches. Like I said earlier, effectively a large language model checking its own homework with no thinking involved because, like I said, they do not think or know things. And OpenAI rushed to launch its O1 reasoning model last year because, and I quote Fortune from last October, Sam Orman was eager to prove to potential investors that in the company's latest funding round, the OpenAI remains at the forefront of AI development.

12:34And, as I've noted in my newsletter at the time, it was not particularly reliable, failing to accurately count the number of times the letter R appeared in the word strawberry, which was the codename 401. Very funny stuff. At this point, it's fairly obvious that OpenAI wasn't anywhere near the forefront of AI development. And now that its competitive advantage is effectively gone, there are genuine doubts about what comes next for the company. As I'll go into, there are many questionable parts of DeepSeek's story. Its funding, what GPUs it has, and how much it actually spent training these models.

13:07But what we definitively understand to be true is bad news for OpenAI. And I would argue every other large US tech firm that's jumped onto the generative AI bandwagon in the past few years.

13:25Parking shouldn't slow you down. ParkWiz gives every driver a shortcut. Book ahead, save up to 50 % and skip the hassle of circling the block. Park smarter, park faster. ParkWiz. Download the ParkWiz app today and save every time you park. Run a business and not thinking about podcasting? Think again. More Americans listen to podcasts than ad-supported streaming music from Spotify and Pandora. And as the number one podcaster, iHeart's twice as large as the next two combined. So whatever your customers listen to, they'll hear your message. Plus, only iHeart can extend your message to audiences across broadcast radio.

14:01Think podcasting can help your business? Think iHeart. Streaming, radio, and podcasting. Let us show you at iHeartAdvertising.com. That's iHeartAdvertising.com. Jan Marsalek was a model of German corporate success. It seemed so damn simple for him. Also, it turned out, a fraudster. Where does the money come from? That was something that I always was questioning myself. But what if I told you that was the least interesting thing about him? His secret office was less than 500 meters down the road. I often ask myself now, did I know the true Jan at all? Certain things in my life since then have gone terribly wrong.

14:45I don't know if they followed me to my home. It looks like the ingredients of a really grand spy story. because this ties together the Cold War with the new one.

14:58Listen to Hot Money, Agent of Chaos on the iHeartRadio app, Apple Podcasts, or wherever you get your podcasts.

15:10Do you want to hear the secrets of serial killers, psychopaths, pedophiles, robbers? They are sitting there waiting for the vulnerable thing. They're waiting for the unprotected. I'm Dr. Leslie, forensic psychologist. I advocate for safety and awareness of predators while wearing pink. When you were described to me as a forensic psychologist, I was like snooze. We ended up talking for hours and I was like, this girl is my best friend. This is a podcast where I cut through the noise with sarcasm, satire and hard truths. I'm not going to fake it and force it. But would you force an orgasm? Because that's like a different layer.

15:46The car accident you didn't want to see but couldn't turn away from. In this episode, I discuss personal safety and self-defense, tools, instincts, and strategies to protect yourself and your loved ones in everyday life and high-risk situations. Listen to Intentionally Disturbing on the iHeartRadio app, Apple Podcasts, or wherever you get your podcasts.

16:14DeepSeq's models actually exist. They work, at least by the standards of hallucination-prone LLMs that don't, at the risk of repeating myself, know anything. They've been independently verified to be competitive in performance, and their magnitude's cheaper in price than those from both hyperscalers, Google's Gemini, MetzLama, Amazon Q, and so on and so forth, and from those released by OpenAI and Anthropic. DeepSeq's models don't require massive new data centers. They run on GPUs currently used to run services like ChatGPT, and even work on more austere hardware. Nor do they require an endless supply of bigger, faster NVIDIA GPUs every single year to progress.

16:51The entire AI bubble was inflated based on the premise that these models were simply impossible to build without burning massive amounts of cash, straining the power grid and blowing past emissions goals, and that these costs were both necessary and really good because they'd lead to creating powerful AI, something that's yet to happen. And it's kind of obvious at this point that that wasn't true. Now the markets are sitting around, they're asking a very reasonable question. Shit, did we just waste$200 billion? Anyway, let's get into the nitty-gritty. What is DeepSeek? First of all, if you want a super deep dive into what it is, I can't recommend VentureBeats write-up enough, I'll link to it in the show notes, as I usually do.

17:34It's really good, and it goes into a lot more detail than I will. But here's the too-long-didn't-read for you. DeepSeek is a spin-off from a Chinese hedge fund called Highflyer Quant. It's a relatively small and young company, and from its inception, it went big on algorithmic and AI-driven trading. Later, it started building its own standalone chatbots, including a chat GPT equivalent for the Chinese market. This is what we know right now. I'm sure some of you will say, oh, well, who knows if that's really true? So, sure, I think that that's fair. I also think that there are parts of Sam Altman's legend that we should question as well.

18:07I think the circumstances under which Sam Altman got made head of Y Combinator are extremely questionable. I'm saying you can question DeepSeek, and indeed you should. We should be more critical of these powerful companies. But don't do it halfway. If we're going to be worried, let's be worried about everyone. Now, DeepSeek did a few things differently, like open sourcing its models, although it likely built upon tech from other companies, like MetasLama and the ML library PyTorch. To train its models, it secured over 10 ,000 NVIDIA GPUs right before the US imposed export restrictions, which sounds like a lot, but it's a fraction of what the big AI labs like Google, OpenAI, and Anthropic have to play with.

18:45I think I've heard estimates of like 100 ,000 to 300 ,000 each, if not more. Now, you've likely seen or heard that DeepSeq trained its latest model for$5.6 million, as opposed to the insane amounts that I'll get to later. and I want to be clear that any and all mentions of this number are estimates. In fact, the provenance of the$5.58 million number appears to be a citation of a post made by an NVIDIA engineer in an article from the South China Morning Post, which links to another article from the South China Morning Post, which simply states that DeepSeq v3 comes with 671 billion parameters and was trained in around two months at a cost of$5.58 million, with no additional citations of any kind.

19:27So you should take it with a pinch of salt, but it's not totally ludicrous. While there are some that have estimated the cost, DeepSeq's V3 model was allegedly trained using 2048 NVIDIA H800 GPUs, according to its paper, and Ben Thompson of StrateTree has made this clear that the$5.5 million number only covers the literal training cost of the official training run, and this is made fairly clear in the paper, by the way, of V3, and that's the one that's competitive with OpenAI's GPT-4O model, meaning that any costs related to prior research or experiments on how to build the model were left out.

20:02Now, big shout-out to Minimax here, the guy on Blue Sky and Twitter, he's great. He is wonderful and also added that this is fairly standard for the industry. Again, you choose how you feel about this, but I want to give you the information. And while it's safe to say that DeepSeq's models are cheaper to train, the actual costs, especially as DeepSeq doesn't share its training data, which some might argue means its models are not really open source, as I said, the numbers get a little harder to guess at. Thompson notes that DeepSeek had to craft a bunch of elegant workarounds to make the model perform, including writing code that ultimately changed how GPUs actually communicated with each other.

20:38This functionality isn't otherwise possible using NVIDIA's developer tools. They really had to get in there, it's kind of cool. DeepSeek's models, V3 and R1, are more efficient and, as a result, cheaper to run, and can be accessed via its API at prices that are astronomically cheaper than OpenAI's. DeepSeek Chat, running DeepSeek's GPT-4O competitive v3 model, costs 0.07 cents per 1 million input tokens, as in commands given to the model, and$1.110 per 1 million output tokens, as in the resulting output from the model. I know that these numbers kind of like just sound like numbers, like maybe you don't have context, so let me give you some.

21:17This is a dramatic price drop from the$2.50 per 1 million input tokens and$10 per 1 million output tokens that OpenAI charges for GPT-40. This isn't just undercutting. This is a bunker buster. Now, there is a side that I'll kind of get into a little bit later in that you are using models hosted in a country that you don't know, probably China. There are data concerns. But again, you can put this on your own server. You could put this in Google Cloud. Both Microsoft and Google are apparently thinking about it. Now, the information reported that Google had added it to Google Cloud. No, they did not.

21:56They didn't do that. They allowed you to connect Hugging Face. This is a whole bunch of technical stuff that if you understand, you'd be like, yeah, right, I know. Long story short, the hyperscalers are already bringing DeepSeq out. And I'll get to why that's bad later in detail. But it's also very funny. Now here's something else that's funny. DeepSeek Reasoner, its reasoning model, costs$0.55 per 1 million input tokens and$2.19 per 1 million output tokens. Now that sounds expensive, maybe it is, whatever. That's goddamn nothing compared to the$15 per 1 million input tokens and$60 per 1 million output tokens of OpenAI.

22:35Oof. If I'm Sam Altman, I'm shitting myself. But there's an obvious part here. We do not know where DeepSeek is hosting its models, who has access to that data, or where that data is coming from or going to. We don't know who funds DeepSeek, other than it's connected to HighFlyer, the hedge fund that I mentioned earlier that it split from in 2023. There are concerns that DeepSeek could be state-funded, and that DeepSeek's low prices are a kind of geopolitical weapon, breaking the back of the generative AI industry in America. I'm not really sure whether that's the case or not. It's certainly true that China has long treated AI as a strategic part of its national industrial policy, and is reported to help companies and sectors where it wants to catch up with the Western world.

23:15The Made in China 2025 initiative saw a reported hundreds of billions of dollars provided to Chinese firms working in industries like chip making, aviation, and yeah, AI. The extent of that support isn't exactly transparent, surprise surprise, and so it's not entirely out of the realm of possibility that DeepSeek is also the recipient of state aid. The good news is that we're going to find out fairly quickly. American AI infrastructure company Grok is already bringing DeepSeek's model online, meaning that we'll get at least a very… some sort of confirmation of whether these prices are realistic, or whether they're heavily subsidised by whoever it is that backs DeepSeek.

Read the full transcript

23:50It's also true that DeepSeek is owned in part by a hedge fund which likely isn't short of cash to pump into them. But as an aside, given that OpenAI is the benefactor of billions of dollars of cloud compute credits and gets reduced pricing for Microsoft's Azure cloud services to run its actual models, it's a bit tough for them to complain about Arrival being subsidized by a larger entity with the ability to absorb the costs of doing business, should that be the case. Same goes for Anthropic, by the way. And yes, I know Microsoft isn't a state, but with a market cap of$3.2 trillion and quarterly revenues larger than the combined GDPs of some EU and NATO nations, it's kind of the next best thing.

24:30But I digress. Whatever concerns there may be about malign Chinese influence are bordering on irrelevant, outside of the low prices of course offered by DeepSeek itself. And even that is speculative at this point. Once these models are hosted elsewhere and once DeepSeek's methods, which I'll get to in a little bit, are recreated, and by the way that's not really going to take very long, I believe we're going to see that these prices are indicative of how cheap these models are to run.

25:03Parking shouldn't slow you down. ParkWiz gives every driver a shortcut. Book ahead, save up to 50 % and skip the hassle of circling the block. Park smarter, park faster. ParkWiz. Download the ParkWiz app today and save every time you park. Run a business and not thinking about podcasting? Think again. More Americans listen to podcasts than ad-supported streaming music from Spotify and Pandora. And as the number one podcaster, iHeart's twice as large as the next two combined. So whatever your customers listen to, they'll hear your message. Plus, only iHeart can extend your message to audiences across broadcast radio.

25:38Think podcasting can help your business? Think iHeart. Streaming, radio, and podcasting. Call 844-844-iHeart to get started. That's 844-844-iHeart. Jan Marselech was a model of German corporate success. It seemed so damn simple for him. Also, it turned out, a fraudster. Where does the money come from? That was something that I always was questioning myself. But what if I told you that was the least interesting thing about him? His secret office was less than 500 meters down the road. I often ask myself now, did I know the true Jan at all? Certain things in my life since then have gone terribly wrong.

26:22I don't know if they followed me to my home. It looks like the ingredients of a really grand spy story. because this ties together the Cold War with the new one.

26:36Listen to Hot Money, Agent of Chaos on the iHeartRadio app, Apple Podcasts, or wherever you get your podcasts.

26:48Do you want to hear the secrets of serial killers, psychopaths, pedophiles, robbers? They are sitting there waiting for the vulnerable thing. They're waiting for the unprotected. I'm Dr. Leslie, forensic psychologist. I advocate for safety and awareness of predators while wearing pink. When you were described to me as a forensic psychologist, I was like snooze. We ended up talking for hours and I was like, this girl is my best friend. This is a podcast where I cut through the noise with sarcasm, satire, and hard truths. I'm not going to fake it and force it. But would you force an orgasm? Because that's like a different layer.

27:24The car accident you didn't want to see but couldn't turn away from. In this episode, I discuss personal safety and self-defense, tools, instincts, and strategies to protect yourself and your loved ones in everyday life and high-risk situations. Listen to Intentionally Disturbing on the iHeartRadio app, Apple Podcasts, or wherever you get your podcasts.

27:51So you might be wondering, how the hell is this so much cheaper? And that's a bloody good question. And because I'm me, I have a hypothesis. I do not believe that the companies making these foundation models, such as Open Air and Anthropic, have actually been incentivized to do more with less. And because their chummy little relationships with hyperscalers like Amazon, Google, and Microsoft were focused almost entirely on making the biggest, most hugest models possible using the biggest, even huger-er-est chips, and because the absence of profitability didn't stop them from raising more money, well they've never had to be fucking efficient have they they've never had to try maybe they should buy less avocado fucking toast anyway let me put it in simpler terms imagine living on fifteen hundred dollars a month and then imagine how you'd live on a hundred and fifty thousand dollars a month and that you have to like brewster's millions spend as much of it as you can to complete a mission a very simple mission live in the former example you concern survival You have a limited amount of money and must make it go as far as possible with real sacrifices to be made with every dollar you spend.

28:55If you want to have fun, you're going to have to eat less potentially, or the food you eat will have to be cheaper. You have to live on a budget. You have to make decisions, and indeed, you might learn to cook at home. You might walk more. You might do things that will help you not spend all your money. In the latter example, where you have$150 ,000 a month that you must spend, you're incentivized to splurge, to lean into excess, to pursue this vague idea of living your life. Your actions are dictated not by any existential threats or indeed any kind of future planning, but by whatever you perceive to be an opportunity to live.

29:30Open AI and Anthropic are emblematic of what happens when survival takes a backseat to living. They have been incentivized by frothy venture capital and public markets desperate for the next big thing the next big growth to build bigger models and sell even bigger dreams like Dario Amadei of Anthropics saying that your AI and I quote could surpass almost all human beings at almost everything shortly after 2027 and I just want to take a fucking second journalists if you're listening to this stop fucking quoting this bullshit stop it you're doing nothing you are failing at your goddamn job every single time you quote this bullshit this nonsense shortly after 2027?

30:10What the fuck does that mean? 2028? 2029? 2030? What does surpassing humans and almost everything even mean? This shit doesn't work. This shit is not good. Oh my god. Anyway, back to the podcast, Ed. Calm down. Both OpenAI and Anthropic have effectively lived their existence with the infinite money cheat from The Sims. And I know some of you might say, by the way, it's not an infinite money. Just add, you go into the console. You get my point. And both companies have been bleeding billions of dollars a year after revenue. And that's, by the way, making billions of dollars and then still losing billions is insane.

30:45And they still operated as if money would never run out, because it kind of wouldn't. If they were actually worried about that happening, they would have certainly tried to do what DeepSeek has done, except they didn't have to, because both of them had the endless cache and access to GPUs from either Microsoft, Amazon, or Google. And the Stargate thing is just, I will mention it later, just long story short, they're not going to put 500 billion dollars into the it it was up to 500 i'm so tired of this shit open ai and anthropic have never been made to sweat unlike me in this closet where i'm recording this and they've received endless amount of free marketing from a tech and business media happy to print whatever vapid bullshit they spout and it's just very frustrating they've raised money at will with and anthropic by the way is currently raising another two billion dollars valuing the company at 60 billion dollars and this was i think happening while deep sea was going on which is really funny and they've done all of this off of a narrative of the we need more money than any company has ever needed ever because the things we're doing have to cost this much there is no other way you must give us more money my name is sam altman i need more money than has ever been made from my huge beautiful company that sucks and needs money to train it help me please my big beautiful sick company is dying but the best and most important company of all time.

32:05It's also normal. Now, do I think that they were aware that there were methods to make their models more efficient? Sure. OpenAI tried and failed in 2023 to deliver a more efficient model to Microsoft called Arrakis. I'm sure there are teams at both Anthropic and OpenAI that are specifically dedicated to making things kind of more efficient, but they didn't have to do it, and so they didn't. And as I've written before in my newsletter and argued on this very podcast, OpenAI simply burns money and have been allowed to burn money, and up until recently likely would have been allowed to burn even more money because everybody, all of the American model developers, appeared to agree that the only way to develop large language models was to make them as big as humanely possible, and work out troublesome stuff like making them profitable or turning them into a useful thing.

32:55later, which is, I presume, when AGI happens, a thing that they're still in the process of defining, let alone doing. DeepSeek, on the other hand, had to work out a way to make its own large language models within the constraints of the hamstrung NVIDIA chips that can be legally sold to China. While there's a whole cottaged industry of selling chips in China using resellers and other parties to get restricted silicon into the country, the entire way in which DeepSeek went about developing its models suggests that it was working around very specific memory bandwidth constraints, meaning that the amount of data that could be fed into it and out of it, and into the chips.

33:30In essence, doing more with less wasn't something it chose, but it's something they had to do. I've touched already on the technical how of these models in greater depth, and you can really read in that in my newsletter, and you can go to where's your at, not at, it's at the end of the episode, but I'll also have show notes to articles like Ben Thompson's from Strategiary, because there are lots of things to read here. I know there are some really technical listeners and I'm sure you're going to flay me in my emails. Please go and read it. I'm not wrong. I've checked with a lot of people too. And by the way, all of this austerity stuff seems to have worked.

34:03There's also the training data situation and another mea culpa. I previously discussed the concept of model collapse and how feeding synthetic data, which is training data created by a generative model into another model, can end up teaching it bad habits, which in turn would destroy the model. But it seems that DeepSeek has succeeded in training its models using generative data. Specifically, though, and I'm quoting GeekWire's John Turow, like mathematics where correctness is unambiguous, and using, and I quote again, highly efficient reward functions that could identify which new training examples would actually improve the model, avoiding wasted compute on redundant data.

34:39And it seems to have worked. Though model collapse may still be a possibility, this approach, extremely precise use of synthetic data, is in line with some of the defenses against model collapse I've heard from LLM developers I've talked to. This is also a situation where we don't know the exact training data, and it doesn't negate any of the previous points I've made about model collapse. Now, we'll see what happens there, but synthetic data might work where the output is something that you could figure out using a calculator, but when you get into anything a bit more fuzzy, like written text or anything with an element of analysis, you'll likely encounter some unhappy side effects.

35:12But I don't know if that's really going to change how good these things are. There's also a little scuttlebutt about where DeepSeq got its data. Ben Thompson at Stratechary suggests that DeepSeq's models are potentially distilling other models' outputs, by which I mean having another model, say, Meta's Lama or OpenAI's GPT-40, which is why DeepSeq identified itself as ChatGPT at one point, spit out outputs specifically to train parts of DeepSeq. This obviously violates the terms of service of these tools, as OpenAI and its rivals would much rather have you not use its technology to create its next rival.

35:46And OpenAI, by the way, has recently reportedly found evidence that DeepSeek used OpenAI's models to train its rivals, and this is from the Financial Times. Although it failed to make any formal allegations, but it did say that using ChatGPT to train a competing model violates its terms of service, and David Sachs, the investor in Trump administration AI and CryptoZar, says it's possible that this occurred. Although he failed to provide evidence, I just want to say how fucking funny it is that OpenAI is going, where? Where? You're stealing my stuff. Don't steal my things. Where? Fucking coward, pansy bastard bitches.

36:23Fucking hell. What a bunch of whiny babies. Oh no, my plagiarism machine got plagiarized. Where? Kiss my entire asshole, Sam Altman, you little worm. You fucking embarrassment to Silicon Valley. You should be ashamed of yourself for many reasons, but so much this though. Where? Oh no, you stole from me. My plagiarism machine that requires me to steal from literally every artist and author on the Internet, the thing where we went on YouTube and transcribed everything and fed it into the machine. That's not stealing. That's good. But you using our model to generate answers, that's just not fair.

37:00What a bunch of babies. You guys, Sam was worth billions of dollars. He has a five million dollar car. Cry more, you little worm! Personally, I genuinely want OpenAI to point a finger at DeepSeek and accuse it of IP theft, mostly for the yucks, but also for the hypocrisy factor. This is a company that, as I've just very cleanly said, exists purely from the wholesale industrial larceny of content produced by literally fucking everyone. And now they're crying. I'm Sam Altman, I'm a big baby, I filled my diaper because someone stole from my plagiarism machine. Kiss my ass. Kiss my ass. These companies haven't got shit.

37:40OpenAI doesn't have shit. They don't have anything. They don't have a next product. Without reasoning, they haven't got anything. And now they don't have that disgusting justification, that overspending the fat, ugly American startup culture of spending as much as you can to build America's next top monopoly. They should be fucking ashamed of themselves. They shouldn't be billionaires. They should be poverty stricken. They should have to pay everyone they stole for. And it's just, it sickens me seeing the reaction from some people on this, seeing the cinephobia, but seeing this level of defensiveness of a company like OpenAI or Anthropic.

38:17And as I'll get into next episode, we are really running out of time here. And I think DeepSeek is really, I think it could be really the end of days for these companies. I don't know how much they've got left, time-wise or even money-wise, and I'm not sure how they even raise money. But in the next episode, I'm going to deep dive into DeepSeek and I'll tell you how they sent the US tech market into a panic and what it actually means for the future of OpenAI, Anthropic and the hyperscalers backing them. This has been a crazy few days. I hope this has helped. And on Monday, you'll find out more.

38:53Thank you so much for listening. The support I've got for the show has been incredible and the emails I've got about DeepSeek, I've been trying, okay? I've really been trying. It's the fastest I could do it. But I'm so happy to do this show, and I'm so grateful for all of you.

39:15Thank you for listening to Better Offline. The editor and composer of the Better Offline theme song is Matt Ossowski. You can check out more of his music and audio projects at matosowski.com. M-A-T-T-O-S-O-W-S-K-I dot com. You can email me at ez at betteroffline.com or visit betteroffline.com to find more podcast links and, of course, my newsletter. I also really recommend you go to chat.wheresyoured.at to visit the Discord and go to r slash betteroffline to check out our Reddit. Thank you so much for listening. Better Offline is a production of Cool Zone Media. For more from Cool Zone Media, visit our website, coolzonemedia.com, or check us out on the iHeartRadio app, Apple Podcasts, or wherever you get your podcasts.

40:15We'll see you next time.

40:30Park smarter, park faster. ParkWiz. Download the ParkWiz app today and save every time you park. Did it occur to you that he'd charmed you in any way? Yes, it did. But he was a charming man. It looks like the ingredients of a really grand spy story. Because this ties together the Cold War with the new one. I often ask myself now, did I know the true Jan at all? Listen to Hot Money, Agent of Chaos on the iHeartRadio app, Apple Podcasts, or wherever you get your podcasts.

41:29real-life physical pain. Learn more about the psychology of everyday life and, of course, your 20s this September. Listen to The Psychology of Your 20s on the iHeartRadio app, Apple Podcasts, or wherever you get your podcasts. Let's start with a quick puzzle. The answer is Ken Jennings' appearance on The Puzzler with AJ Jacobs. The question is, what is the most entertaining listening experience in podcast land. Jeopardy truthers believe in... I guess they would be Ken-spiracy theorists. That's right. They give you the answers and you still blew it. The Puzzler. Listen on the iHeartRadio app, Apple Podcasts, or wherever you get your podcasts.

42:12This is an iHeart Podcast.

From the publisher

In this episode, Ed Zitron explains how DeepSeek, a relatively-unknown Chinese model AI developer incubated in a hedge fund, has punctured the generative AI bubble, throwing the US startup scene (and markets) into disarray.

---

LINKS: https://www.tinyurl.com/betterofflinelinks

Newsletter: https://www.wheresyoured.at/

Reddit: https://www.reddit.com/r/BetterOffline/ 

Discord: chat.wheresyoured.at

Ed's Socials:

https://twitter.com/edzitron

https://www.instagram.com/edzitron

https://bsky.app/profile/edzitron.com

https://www.threads.net/@edzitron

See omnystudio.com/listener for privacy information.

More from Better Offline

All 276 episodes
What The Hell Is DeepSeek?Better Offline · 33 min
Listen in VO