In short
The episode argues against “pacing the frontier” and claims anti–data center politics are largely misinformation driven by China. It also explains why AI inference is memory-bound, how KV caching and compression reduce inference cost, and whether planned data centers will actually be built given energy and regulatory constraints.
Guest
Thomas Sohmers, co-founder and chairman of Positron AI (semiconductor/hardware company building chips, software stack, and rack-scale systems for generative AI inference). Positron raised an $875M Series C at a $5B valuation; investors include Gavin Baker (Atreides).
Key claims
- “Anti–data centres” is framed as a Chinese psyop; China is allegedly building gigawatts of (often dirty) power and data centers while spreading false narratives.
- Inference is memory-bound (forward pass reads weights per token), unlike training which is more compute/FLOPs-bound.
- API providers’ “cash token” pricing yields very high margins (Anthropic cited at ~80% gross margin).
- KV caching is crucial: attention compute grows quadratically with context length, while cached K/V grows linearly; caching/compression enables long-context economics.
Notable examples
- “Single in and out” water use claim vs US data centers; golf courses cited as higher water use.
- AgentX benchmark: ~96% token reuse in traced agentic coding sessions.
- KV cache scaling: user sessions can reach ~100GB, exceeding model-weight sizes at scale.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOOverview of Positron and AI Infrastructure
4:01 to 6:38
Thomas Sohmers explains Positron's role in hardware for generative AI.
“I said to you just before this that I think there are some big questions that the world doesn't know or things that they know that I think we're going to correct today.”
Differences Between Training and Inference
6:39 to 8:06
Discussion on the computational differences between AI training and inference.
“So that's something you can just crush through with a bunch of compute.”
The Memory Bandwidth Challenge in AI
8:07 to 10:55
Thomas discusses the disparity in advancements between compute and memory technology.
“But just like the ratio of the compute to memory bandwidth had this divergence.”
Profitability and Margins in AI Providers
10:56 to 13:16
Insights into the economics of AI providers and their margins.
“I do find it compared to a year or two ago, there are now different prices listed for cash versus uncashed tokens.”
Pacing the Frontier and Regulatory Concerns
13:17 to 14:03
Thomas shares his views on the implications of AI regulation and pacing technology.
“I don't think that's actually the intention or anything.”
Concentration of AI Technology and Its Risks
14:03 to 16:42
Discusses the dangers of AI technology becoming centralized among a few powerful entities and the implications for freedom and progress.
“And that by pushing for pause, it's actually giving ammunition, giving better basis for those that actually just want to stop the technology over completely.”
Global AI Competition and Regulatory Challenges
16:42 to 19:38
Explores the geopolitical implications of AI development, emphasizing the risks of regulatory burdens in the U.S. compared to China.
“Like Dario basically, on one hand, verbally begging for government governments to take over Anthropic.”
Critique of Anti-Data Center Sentiment
19:38 to 22:20
Examines the narrative against data centers and argues it may be influenced by foreign strategies, particularly from China.
“comes to AI technology, there's not the acceptance that export controls and all of the other elements that theoretically would allow us to pace and have people keep paces behind us just aren't good long-term solutions.”
Economic Impact and Misconceptions About Data Centers
22:20 to 26:06
Discusses the economic benefits of data centers and addresses misconceptions regarding their impact on electricity and resources.
“We need to dress them up to be the world wonders that they are technologically.”
Future of Data Center Locations and Innovations
26:06 to 28:00
Speculates on the future locations of data centers and introduces innovative solutions like ocean-based data centers.
“from anything that's already been allocated.”
Show all 23 chapters
Exploring Alternative Data Center Technologies
28:00 to 28:31
Learn about ocean-based data centers and the potential for space economies.
“I also think there's great alternative technologies.”
Energy as a Bottleneck in Computing
28:31 to 29:52
Understand how energy costs influence intelligence and economic limitations.
“If we think about the cost of intelligence being tied to the cost of energy, how should we think about energy as a bottleneck moving forwards?”
Concerns Over National Debt and Economic Risk
29:52 to 31:16
Discuss the implications of sovereign debt and its effects on the financial system.
“I would say we have got way more economic limitations.”
Understanding KV Caching and Its Importance
31:16 to 32:08
Learn about KV caching and its role in enhancing computational efficiency.
“Like the past three, four years have shown, we're accelerating every aspect of these businesses in terms of revenue profits and how they are improving the productivity and value downstream.”
Mechanics of Tokens and Sequence in AI
32:08 to 35:38
Explore how tokens work within AI models and the significance of sequences.
“So there's always, I would say, a give and take relationship with innovation.”
Compression Techniques for Efficient AI
35:38 to 38:24
Delve into compression methods that improve AI model efficiency without losing much accuracy.
“And it comes questions of how long do you want to keep that for?”
Challenges of KV Caching in Large Systems
38:24 to 41:25
Examine the trade-offs and challenges of implementing KV caching for large AI models.
“You'd still have some lobotomy, but thankfully it's kept within like 1 % of unquantized model.”
Future of Memory in AI Inference
41:25 to 42:00
Discuss strategies for improving memory capacity in AI inference systems.
“If you appreciate the importance of saving on compute, but the challenge of memory with KV caching, what's the answer then?”
Understanding Memory Hierarchy in AI Models
42:00 to 44:06
Explore how data is stored and accessed in AI systems and its implications.
“your GPU accelerator memory that is primarily responsible for holding the weights.”
The Future of Model Sizes and Context Links
44:06 to 47:36
Discussion on the evolving size and capabilities of AI models and their contexts.
“I just want to break some of the things you set up there.”
The Breakthroughs of GPT-6 Astra
47:36 to 53:54
Insights into the capabilities and improvements of the latest AI models.
“So I completely hear you in terms of maybe 5 % of them will be in this smaller model enterprise owned kind of model landscape.”
The Role of Chip Development in AI
53:54 to 56:00
Examine how advancements in chip technology impact AI capabilities and industry evolution.
“We see the commoditization of the chip player with everyone building their own chips.”
Innovations in Memory Capacity and AI Models
56:00 to 1:09:03
Learn about the advancements in memory capacity and their implications on AI models.
“I think with traditional linear or quadratic attention, there is going to be limits of scale when it came to what the hardware could provide.”
Transcript
Automatic transcript. May contain errors.0:00I would say I'm overall opposed to, you know, the pace of the frontier direction it's going in. The scariest thing to me on the political spectrum and the way all of this is being treated is that it's now become an almost unifying issue on left and right about being anti-data centres. I think that is almost entirely a Chinese psyop. A single in and out uses, you know, more water than, you know, the largest data centres in the United States.
0:22Harry Stebbings:This is 20VC with me, Harry Stebbings. Now we have probably one of the most important times in history for technology. We have the biggest model provider saying we need to pace the frontier. But what does that actually mean in reality? How possible is it? What does it mean for the threat from China? What does it mean for the infrastructure layer moving forwards? We have a true expert of the space on the show today in the form of Thomas Somers. He's the co-founder and chairman at Positron AI. They just raised an$875 million Series C at a$5 billion valuation. They've got some of the best investors in the business, including the one and only Gavin Baker at Atreides and many more great names.
1:00Harry Stebbings:Thomas did not hold back in this episode and it's this beautiful combination of incredible education on the infrastructure that powers this economy for AI and then also, I don't know how to say it, but analytical gossip would be a more intellectual way of saying incredible discussion about what we can expect in the next few months from the biggest players in this space. But before we dive into the show today, the best model for your application might not exist yet. The most ambitious AI teams are training open models to beat the frontier in their domain. Now, on November 3rd in San Francisco, Forge, by Fireworks, brings those teams together.
1:35Harry Stebbings:People like Jensen Huang, CEO of NVIDIA, Michelle Katasta, President and Head of AI at Replit, Lin Kuo, CEO of Fireworks, and more, to hear how AI leaders are taking control of their differentiation, margins and roadmap by owning their intelligence. Learn how they're building it with inference, training and intelligent routing. Forge is free to attend, but space is limited. Apply today at fireworks.ai forward slash forge. While Fireworks AI powers product intelligence, Asana keeps the work moving. Most companies have tried AI, most aren't seeing results. Not because AI doesn't work, it's because AI hasn't reached the workflows yet.
2:15Harry Stebbings:That's the gap asana is built to close asana is the operating system for human agent teams your easy button for ai productivity across every team ready to go ai teammates pre-built for marketing ops and it no prompt engineering no setup they show up where the work is happening already onboarded in your workflows ready to deliver with asana your whole company can work on the same plan towards the same goal, whether you're a team of 10 or a team of 10 ,000. Asana, where humans and agents workflow together. Try it at asana.com. That's A-S-A-N-A dot com. While Asana organizes the work, superhuman speeds up the day.
2:56Harry Stebbings:Can I be honest with you for a second? Here's what my week actually looks like. Back-to-back meetings all day, and the real pressure builds up in the gaps between them. I'll come out of six founder meetings with a whole stack of follow-ups waiting, and my inbox, It's hundreds of messages, and somewhere in there is the one thing that really matters, which I always tell my mother is her. I've tried AI tools for this, but they always lived in another tab. I'd have to stop, go find it, paste in the context, write the perfect prompt. It was just enough friction that it never really stuck. Superhuman Go, from the makers of Grammarly, totally different.
3:30Harry Stebbings:It's right there where you need it, when you need it, in the doc, when I'm reading, the email I'm writing, when a thread gets dense, I highlight it and ask Go to pull out what matters without losing my place, and after a run of back-to-backs, I ask what was actually decided, and Go turns it into clear action items or drafts, the reply I need to send. It's already up to speed, and I'm always the one in control. Try Superhuman Go from the makers of Grammarly, and find out more at superhuman.com. You have now arrived at your destination. Thomas, I am so excited for this, dude. I said to you just before this that I think there are some big questions that the world doesn't know or things that they know that I think we're going to correct today.
4:11Harry Stebbings:So thank you so much for joining me. Yeah, great to be here. Thanks, Harry. Can we just have a brief description of what Positron is and where does it sit in the stack? Yeah. So Positron is a fabulous semiconductor startup that's building hardware. So really everything from the chips, the software directly running on top of that, all the way up to the full systems and rack scale deployments to power generative AI inference. So effectively everything in the hardware and sort of the low level direct talking to hardware software stack that powers all of the applications, you know, everyone in the world is excited about right now.
4:45So everything from the likes of ChatGPT, Claude, et cetera.
4:49Harry Stebbings:How does the infrastructure stat required for inference, what you're working on, change compared to training? So training, I would say, from the underlying compute level, fundamentally, you know, is a compute bound problem. So it's a workload that's the more flops that you have. And if you look at, you know, be it from the regulatory and some of the, you know, export control frameworks are heavily focused on the just flops required. So how many floating point operations per second can be done. And more or less, you know, the amazing thing that the scaling laws of the past decade have shown is that the more parameters you add to a network and the more flops you dedicate to that during training, the better that model is going to become.
5:32The big difference with inference and so the deployment of those models is the fact that for the actual math and the steps that you're doing is about half of what you're doing during training in terms of that the steps shouldn't be thought of as like the actual compute involved. But what it turns out to be is that that forward pass, that inference portion of it is heavily, heavily memory bound due to the fact that basically for every single token that's generated, every little bit of output, that requires going through the weights, the parameters you could kind of, you know, from a biological sense, think of the neurons.
6:06You have to read the values of that for every single individual token. And so the way that that's, I think, very interesting about this, and I don't know if it says anything about the value or kind of actually saying that the inferences is somehow, I don't want to say more important because, of course, you have to train. But fundamentally, when you're training, you already have the corpus. You already have all of the training data. And so with all that data, you can massively parallelize the token inputs, all of the sequences of words and sentences, paragraphs, et cetera, that are going into that.
6:41So that's something you can just crush through with a bunch of compute. But when you're inferring, because that's actually generative, you don't know what the token is five words down the line. And so you have to generate each and every one auto-aggressively or in order without, you know, foresight. And so that becomes a hugely memory bound problem that can't just be massively paralyzed like training.
7:03Harry Stebbings:So I totally get that in terms of the shift from compute bound to memory bound. Is that what people mean when they say about the memory wall with regards to what you're doing? Partially. I mean, the memory wall as a phrase has been around for a long time before, you know, the hype around AI. And really what it's come down to is if you look at the past 50, 60 years of computing, we've been able to have Moore's law giving us more transistors per square millimeter of silicon consistently. And while that has been able to result in greater raw compute, flops, et cetera, the improvement of the memory technology is not kept up at the same rate.
7:40So roughly speaking, between 2014, just very early innings of the new AI era, till 2024, you had about 120x improvement in the flops of GPUs. So a single NVIDIA GPU had about 120-fold improvement in flops. And that's what enabled a whole lot of the improvements over that decade. The improvement in memory bandwidth is only 17x. So I would say the real embodiment of this is we had massive improvements on a per device basis of the flops and then a whole bunch of elements on the periphery of improving the connectivity, etc, etc. But just like the ratio of the compute to memory bandwidth had this divergence.
8:24And so you had cases where if a problem was memory bound and you couldn't just scale the compute linearly with that, you were getting more and more memory bound as the decade progressed.
8:34Harry Stebbings:Why was there such a misalignment in the progression between the two? One's 100x, one's 17x. Why is that the case? Yeah, well, it comes down to a lot of technical implementation details, like the fact that if you look at the lowest level, the type of memory that is used on the silicon itself is called SRAM, static RAM. And SRAM is made out of six transistors with a bit line, a word line, some other control logic around it. But that SRAM cell has not scaled in terms of the sizing of that with Moore's Law over the past about 15 years. So they have grown or shrunk, I should say, much slower than just a group of transistors that you'll use for other purposes.
9:17And I would say that there's just been a lot more architectural advancements that could happen on the compute side, while an S6TS RAM more or less has not changed in 30 or 40 years from an architectural primitive perspective. And so that's on the input side of raw technical capabilities on fabrication, et cetera, have not been able to improve. But I would also say that there wasn't the right motivations for most of that decade. So with convolutional neural networks, so things that powered like AlexNet, which really launched the deep learning revolution in 2012, and then ResNet and all of the, I would say, the advancements during the 2010s was in the realm of machine learning models that were fundamentally compute bound.
10:00You could just throw more and more flops at CNNs and get better results. And you didn't really need all that much, be it memory capacity or memory bandwidth. But it was really with the transformer. And even though the attention is all you need paper came out in 2017, I would say it did not really get the attention, you know, pun intended. It deserved until 2020. Well, GPT-1 and GPT-2 came out prior to that 2018, 2019. But it was really GPT-3 showing that, OK, you go from a billion-ish parameter up to 175 billion parameters and you actually get this massive improvement in capability. And that's really where I would say the transformer revolution started.
10:42And most people didn't really catch on to that until the end of 2022 when ChatGPT came out.
10:48Harry Stebbings:When you look at token economics and token efficiency today, what does no one know or talk about that you think should be much more front and center? I do find it compared to a year or two ago, there are now different prices listed for cash versus uncashed tokens. But I don't think people realize how any providers that charge the same amount, even with a lot of people's cash prices, how high margin that is. It's like insane. You make all of your money on selling cashed input and output tokens. Why is that? Sorry, just so I understand that. Oh, because for like when we discussed earlier that both processing a cash token is essentially free.
11:28It's one one thousandth of the cost, you know, order of magnitude of actually having to recompute and generate that token. So there is so much you can juice out of selling those cash tokens. And basically all the providers, they charge you to cash a token. and they charge a higher rate than just the normal processing fee for an input token. And then they charge you a lower rate when you read from that. And it's great when you're paying that lower rate, but they're making obscene margin on that cash read. And there's a reason why Anthropic is reported to have 80 points of gross margin right now on API business.
12:09Harry Stebbings:Were you surprised by that 80 points? Not really. I'm impressed by 80 points of margin in basically any industry. It's difficult to get that margin. And the great thing about capitalism is those margins will compress with competition. So I'm confident and happy for that, even though those people are theoretically my customers and my margin is sort of based on their margin. But I care more about a healthy ecosystem long term. It's more surprising to me how many people still today think that these are horribly unprofitable businesses and that the whole market's going to zero. It's absurd to me that the meme of OpenAI, Anthropic, et cetera, are just burning cash and eventually they'll run out of cash that they can burn.
12:49If they stop training, they'd be massively profitable overnight. And there's a ton of other levers that they have without pacing the frontier as Dario just had in his essay. I mean, I jokingly think that, you know, a little bit of the pacing the frontier discussion is, oh, this is a great way to reduce costs ahead of IPO. But I don't think Anthropic or anyone needs to do that.
13:12Harry Stebbings:I think they're amazingly profitable businesses with their scaling rates. And it would reduce costs just so I understand because they would spend less on training because they would be slowing down the speed. Yeah. Right. Yeah. I don't think that's actually the intention or anything. But yes, that's the bulk of their costs. How did you analyze the pacing the frontier? You brought it up. How did you analyze it? I have mixed feelings on the safety topic. I am a human, would like to live old age, and more so than that, have humanity continue to the stars and beyond. But I believe much more in the ability for this technology to revolutionize every part of humanity in a positive way.
13:55I do worry that pacing in a lot of the ways that's being talked about, not necessarily how it will be implemented, has two big risks. One, a major pause and sort of that playing into, for lack of a better word, the bloodhate sentiment that exists. And that by pushing for pause, it's actually giving ammunition, giving better basis for those that actually just want to stop the technology over completely. And I see that as a major risk for humanity. The second piece is my actual PDU, the thing that I'm most worried about of any AI outcomes is that technology and capability being concentrated to relatively few people.
14:35And having a lot of what's being discussed from a regulatory framework and limitations on technology, et cetera, I think is like the modern road to serfdom. It's like the concentration of technological capability, making legal to do matrix multiplications is like the thing that will set us back to, you know, pre not just industrial revolution. It's like pre enlightenment capabilities like that. That is like the biggest attack on classical, like liberal freedom concepts that I can I can think of, because while I do think that the vast majority of tokens are going to be produced by the big players.
15:13If the technology itself is restricted to just those, then they're going to be the new lords and kings and everyone else is back to surfs.
15:23Harry Stebbings:The greatest of respects, is it not just lip service? Great, we'll stick meter in the corner. They can do that compliance. And then we can IPO. Sam can have a reason not to IPO because his numbers aren't as good as Anthropics. It plays into both our desires. And Elon wants time to catch up as well. So it plays into everyone. I completely agree. And I think that is the biggest internal reason for everyone other than Dario. Dario and I would say the vast majority of people in Thropic are true believers, both in all of the promise and capabilities of technology and the risks. And if I were in their shoes, I would also be taking the massive amount of responsibility for that.
16:04But I will say like, yeah, there's a lot of strategic reasons of saying, OK, by having these auditors, et cetera, that removes some potential responsibility, culpability from like legal perspectives, et cetera. The risk that I just don't think that any of I think this is a problem with a lot of very smart people, especially when they've amassed large wealth and power, et cetera, is they think that they're going to be able to keep that. And the scariest thing, kind of my point, like the centralization of technology, like if it gets concentrated with companies, governments, et cetera, is you've got people that think that they're the smartest people in the room, not realize that they are not going to be the ones to actually control it when they put these measures in place.
16:43Like Dario basically, on one hand, verbally begging for government governments to take over Anthropic. He's in part saying that because he doesn't think it will actually happen. I would love to see his reaction if and when that actually happens. And he realizes, oh, shit, I thought that if I was begging for regulations, they would then make me the regulator. And when that doesn't happen and it just becomes a bureaucracy that halts all progress and the capabilities that currently exist basically get squandered to select bureaucrats, that's the worst outcome I can imagine.
17:16Harry Stebbings:When you consider the advancements that China are making, especially with their open ecosystem, that they are an incredibly talented ecosystem right now moving forwards. If we pace and they don't, what happens then? I guess when I said that the worst possible outcome, I wasn't counting the Terminator outcome and I wasn't counting that. So I think, of course, everyone can agree Terminator or similar is very bad, but I think is extremely low probability. And just, I'm not a believer in that doom scenario. For the vast majority of people would result in the same level of serfdom that I worry about with the scenario that I described would be significantly worse for some number of people in a Chinese CCP controlled, you know, super intelligent AI scenario.
18:04On one hand, their strategic angle right now is have technology proliferate through open source, etc. I think as soon as they get into pole position, the latter gets pulled up with them in some way. I don't think they actually want the technology to be easily accessible to everyone. Now, I don't know if they will decide that it's okay if the rest of the world has some access to the technology, but they definitely will not let the billion people that are not CCP party members benefit equally from technology.
18:35Harry Stebbings:So just so I understand, do you agree with it? Because to me, I just didn't get it. You can't pace the frontier unless the global AI community paces the frontier. And I don't see Putin signing up. Agreed. And I think this is a little bit the same naivety that I described by these company leaders and in general, people in the Western world thinking, oh, we're so great. We're so advanced, so far ahead that we can't get caught up to. I mean, on paper, is the U.S. the greatest military force in the world? Yes. If we had to all of a sudden have a drone incursion, the same level of what's happening in Ukraine, Russia, coming up from Mexico.
19:14And if you take Mexico, just say that they developed very naive drone technology, et cetera, like on the level of what's happening in Russia, Ukraine, and Iran, how would we respond to that as a country if we had that coming up across our border? It doesn't matter. Our amazing military might. We built our military to fight the last war. And I think geopolitically, our thinking is, oh, we're the big dog still. And that when it comes to AI technology, there's not the acceptance that export controls and all of the other elements that theoretically would allow us to pace and have people keep paces behind us just aren't good long-term solutions.
19:51Do you think we should have export controls? I am very much a strong believer in free trade and free exchange of ideas. The exception to that is a little bit, I think China has been a free rider of all of the benefits of a liberal free trade order for the rest of the world while they get to keep everything closed off. I am very, very happy and think that any government societies, people that want to embrace free exchange of ideas and trade and everything else, we should have a very vibrant economy and ecosystem. But totalitarian regimes should not be able to participate with that, especially in the case where they get all of the benefits of that and get to export themselves, you know, things that make them better able to have that totalitarian system keep up.
20:36Harry Stebbings:We mentioned pacing the frontier and the different people who supported it. You had Zuck and Jensen say nothing. Well, Zuck actually come out in opposition to it saying that we should continue as planned. What should we take from those two seemingly silence and opposing it? I would say based on my overall beliefs right now, as probably evidenced by the conversation so far, I would say I'm overall opposed to the pace of the frontier direction it's going in. And so I appreciate anyone that is adding to the discussion that is, I think, being realist about the benefits and risks. But you always have to take that with a grain of salt of what are the motives of anyone that's discussing it.
21:22And I would say I probably appreciate Zuck or Dario's comments infinitely more than a random politician and not just random, the quote unquote leading politicians that don't actually understand the technology. The scariest thing to me on the political spectrum and the way all of this being treated is that it's now become a almost unifying issue on left and right about being anti data centers. I think that is almost entirely a Chinese psyop.
21:48Harry Stebbings:Can I ask, why is it a Chinese psyop being anti data centers? Because it does increase. I'm totally with you in the benefits of them. And Gavin Baker said how they're the greatest economic kind of needle mover for large parts of the country. I guess they see increased electricity prices, increased water prices and ugly data centers in their backyard. Why is it a Chinese set up? What am I not seeing? Well, just on the ugly and all that, I totally am supportive. We need to have beautification campaigns and really turn them into centerpieces of our society. I think if thousands of years from now, future history looks and they should see these massive data centers, the really, really massive, impressive ones, should be like the Great Pyramids.
22:34We need to dress them up to be the world wonders that they are technologically. The thing from water usage and the amount of power they consume, et cetera, so much of the early information that went out by unsophisticated writers or things that are just patently false. A single in and out uses more water than the largest data centers in the United States. Golf courses are orders of magnitude more. These are closed loop liquid cool systems that you don't even want to use water in a lot of these cases. So from the Chinese PSYOP perspective, they're not, to your point, they're not pacing the frontier.
Read the full transcript
23:12They're adding gigawatts of new electricity generation capacity, most of it being dirty. They're building massive new data centers, horribly displacing people. It just irks me so much that we have the freedom in the Western world to criticize companies, governments, everything based on false information. And I love the freedom elements of that. But it is a strategic disadvantage when China can just say, yeah, we're going to just bulldoze all these people's homes and do rolling blackouts wherever in order to serve the greater good of new training capacity.
23:49Harry Stebbings:I mean this with the greatest of respects, but I don't understand how anyone thinks the US or Europe can beat China when they have no regulatory or policy restrictions. And I mean, the UK, you can't put up a paper airplane without getting a permit. so like we're fucked but you are getting there and you're getting you're becoming a european state in terms of the regulation and policy requirements am i wrong am i being overly negative i i didn't get it i you're right um i think the greatest advantage the u.s has in that regard is that there's still a lot of land a lot of places that do not have all the same levels restrictions i don't agree with a lot of things of most you know administrations of my lifetime But the current administration gets attacked for saying they're destroying our environments and destroying national parks, et cetera.
24:38The vast, vast majority, 90 plus percent, I don't know the exact numbers, of national federal land is just open, empty desert in the West that is not part of a national park or anything. And the fact that there are so many restrictions to utilizing BLM land for building data centers where it's literally hundreds of miles from any populated area. My great state of Nevada has plentiful geothermal, solar, all these green energy technologies. And we could build nuclear and other things in the middle of the desert where it won't impact anyone. And that there's restrictions to that is completely absurd to me.
25:16And I will say there has been some political will and push to solve these things. But literally just in the past year, you have Republican governors and other politicians that at least had part of their platform to be pro-growth and all of these things backing away because they see from their own political base being anti -data center based on completely false premises. And one of the points I want to go back to that you brought up was that people would have higher electricity costs. This is the most basic supply and demand. If we increase generation capacity and no one's saying we want to be taking energy from what's reserved for people's homes.
25:55One of the regulatory problems I see is power companies have to have this offer of energy availability that is baked into cost and capabilities for everyone. There is absolutely zero cases where a data center could be potentially pulling power from anything that's already been allocated. So that's just an impossibility. And all these data centers that are getting built right now are coming with generation capacity that covers their own use and beyond that. And we're just not allowing them to hook up to the grid where they could actually be lowering the prices for everyone. And then you've got people on the power company side, like they're lobbying against new generation capacity because that will actually, you know, market forces will, more capacity will decrease prices, which would be good for consumers.
26:40So it's a very wonky market.
26:43Harry Stebbings:What percentage of data centers that are planned will be completed, do you think? From the major providers, I think I would say that the capacity that they have planned, they may be in different locations. I mean, you've had some local communities that have successfully stopped facilities going in there, but then those data centers just move. I don't think a year ago, the major data center builders and operators were thinking that the political problems were as bad as they were. And so there is a lot more effort being put into education in those communities now, which I think will turn the tide a bit.
27:18But I mean, it's also just going to mean that those data centers move to locales that aren't going to have those problems as well. And like I said, we've got large tracts of land that can support it. So I'm not too worried that it's going to be like an existential threat and capacity build out. And then, of course, there's space if Elon's successful.
27:37Harry Stebbings:Do you believe that space is a viable alternative truly, or is it conference talk and lip service to justify a market cap? I think something can start as one thing and turn into something else. I would never, ever bet against Elon. I mean, I primarily bet for Elon. If you asked me a year ago, I just would not have thought that there would be a good reason for it in the near term because it's going to be cheaper, easier, et cetera, to build on land. I also think there's great alternative technologies. A company we're partnered with, and I'm good friends with the CEO, is a company called Panthalossa that's building ocean-based data centers, basically a very interesting pumped hydro solution in the middle of the ocean.
28:18So there are alternatives that don't require going into space. I think long-term. Part of the reason I'm a long-term big believer in space data centers is I just think we're going to need to have a space economy for humanity to live up to its long-term potential.
28:31Harry Stebbings:Love that. Totally agree on Everbet against Elon. If we think about the cost of intelligence being tied to the cost of energy, how should we think about energy as a bottleneck moving forwards? To what extent is it? We mentioned policy and regulation being a core bottleneck. Is energy a bottleneck moving forwards or less than people consider? I think there's two pieces to it. I mean, one, Positron is trying to deliver more compute, more capabilities per watt per megawatt. And so, you know, sort of on our base case, we can turn what you would have spent 500 megawatts with NVIDIA equipment and do that 100 megawatt.
29:07I don't think that's actually going to mean that you're only going to build 100 megawatt facility. You're still going to build the maximum amount of compute that you can. You're just getting more tokens, more intelligence per joule. And so if I go back to the long-term thinking, I think assuming humanity continues for hundreds, thousands of years, everything turns into an energy problem. And you can go back thousands of years and just look at the progression of mankind. Fundamentally, that is a perfect track of our ability to produce and use energy, discovery of fire up to nuclear power plants.
29:41The simple tongue-in-cheek answer to your question is all progress is gated by energy. And even if there's energy available, it may not be economical, and so it won't be done. So actually, I would say the bigger limiter than just saying that energy, like our ability to build and produce energy, we've got plenty of technologies and capability to do it. I would say we have got way more economic limitations. It's like how much debt is the world willing to take on to build out everything over the next couple of years? ties into energy, it ties into the infrastructure itself, etc. So I think economics is a much easier sort of scapegoat to pin.
30:16Harry Stebbings:I mean, people are already very concerned by the levels of debt being taken out in the debt cycle. Do you think their concerns are justified and do you share them? I think we've got a major sovereign debt problem that masks a huge amount of second and third order elements in the financial system. Just the inflationary consequences of government that can print infinite amounts of its own currency. And the fact that we are, as we're already seeing the treasuries and the greater bond markets, that there is greater and greater perceived risk of the most quote-unquote risk-free asset, I think will trickle down to all elements of the financial system.
30:57And so when people worry about Oracle's debt and credit rating, I'm like, I believe in Oracle's business model and ability to execute and do everything. a whole lot more than the United States government. It's just the United States government can issue its own currency and also has guns and nukes to take tax revenue. So my biggest economic concern there is that there will be a more acute specific crisis that arises out of the compounding of national debt leading to devaluation of the currency that has all of the consequences down the stream rather than like, I'm really not worried about any of the companies in the AI debt stream, not hitting their revenue targets.
31:38Like the past three, four years have shown, we're accelerating every aspect of these businesses in terms of revenue profits and how they are improving the productivity and value downstream.
31:52Harry Stebbings:I'm jumping around, but fuck it. When I was doing the research, I was reading about KV caching and compression as part of this. And I was honestly getting lost, but I was intrigued and digging deeper and deeper. And I was like, why did I not know this before? And so I don't think many will know this. What should we know about KV caching? Why is it important? Can you explain it to me a little bit? Yeah. So there's always, I would say, a give and take relationship with innovation. I guess one element I'll just have to explain to make all this clear is the concept of a sequence. I kind of already talked about a token.
32:28You know, a token effectively as part of the training process when any of these big model apps are developing a new model, they have a vocabulary that they define. So they take their big giant corpus and they do some statistical work to figure out what is the best encoding method to take all of the text in this and break it into chunks that get reused frequently to have things be more efficient. And so what you end up doing is if you took a English dictionary, you'll find that there are common prefixes and suffixes and groupings of words. And if you just try to think of how would I best compress this if I just had symbols for these prefixes, suffixes, et cetera, compress this into a thing.
33:09And basically what ends up happening is a token, if you're using ChatGPT and you see text streaming out, if that's going particularly slow or you're quite a keen eye, you'll see that it's portions of words that come out at times, sometimes a full word, sometimes a small fraction of word. And each of those little flashes that you see as a token. And roughly speaking, a token is equivalent to a half to like 75 % of a word on average in large English corpuses. So that's token. A sequence or the thing that builds up to be in context in a model is the grouping of all those tokens in N order. And what happens when you're running an inference, you've given a prompt, you have, what is the capital of France as your input prompt?
33:54That is tokenized, that is four or five, six tokens that go into the model. And when you do inference, it's going to say the capital of France is Paris and the City of Lights, you know, some thing after that. And so when you have that entire sequence, when Transformers originally came out, for every token that was generated, you were doing the computation for generating all of those tokens, including the ones that you've already processed. And, you know, clever, I would say kind of obvious based on, you know, all of the developments in the past of computer history, but wasn't done initially, was that, well, you don't actually have to redo the compute of the things that you've already had as inputs and what you've already generated in this turn.
34:36And so the KV cache was born, where within the model, there's these two matrices called K and V, keys and values. And those matrices are fully based on the prompt and whatever is generated during a turn. And so by actually storing those two matrices, you can avoid having to redo computation at the cost of now having to store this thing in memory. And that's the simple example, very, very small, hundreds of kilobytes of data. But the thing is that these things grow with the sequence length. But the interesting thing is that for the attention mechanism, the compute per token grows quadratically with the sequence length.
35:17So you're having to spend more and more compute quadratically. So that's an exponential curve as sequence length grows. But when you store that as just your K and V, that's just a linear growth. What KV caching does is means that you don't have to do that compute, which gets very expensive very quickly at the expense of needing to store these things. Storing that is a complexity in itself because that's a unique KV cache for every single user that you're serving. And it comes questions of how long do you want to keep that for? And how do you manage all of that in a large system?
35:49Harry Stebbings:What does it mean then when we hear about compression and uncompression of KV caching and potential entropy within the system? So there's two different forms of compression, a couple more than that, but the two main ones. So one is quantization. So if you've got each of your values, your weights, your KV caches, activations stored in a particular data type. So before the machine learning revolution, most of the world's computation was done in FP32. So you have 32 bits to represent a floating point number, and that's broken into Mantissa and Exponent. Well, it was pretty quickly realized that having 32 bits of precision was super overkill for the things that you're wanting to represent.
36:30And it's both cost more from a storage and computational perspective than lower precision. So we went to FP16, Google developed BF16, a little rejickering of those bits. Went to FP8, now we're at FP4 in popular systems. And so we've been reducing the precision quite quickly, but that does lead to, you know, for lack of a better word, some brain loss when these models run, just because you are now trying to encode the same information into fewer bits. And so there's been a lot of interesting schemes to say, okay, I'm going to take this group of FP16 values, BF16 values, and I'm going to quantize those.
37:12So I'm going to use a truncate and rounding that down to, let's say, int4 values. So now you actually saved 75 % of your total size of that group of values. You shrunk that down from 16 bits to 4 bits. But just doing that naively will mean that on a lot of benchmark scores, you'll have them get 20, 30 % worse. So you get that 75 % savings in space, but you kind of lobotomize the model. But advanced quantization techniques actually say, okay, these 16 values, I'm able to have a shared bias and a multiplier for it. Let's say for those 16 now int4 or fp4 values, you store one new fp16 value that gets applied to all of those at compute time.
38:01So you get a 75 % compression on all those values at the cost of now adding to add one new fp16. And basically the state of the art here is you're able to get things compressed from fp16, 16 bits per value down to like four and a half bits per value. And that can be applied to weights, the actual parameters and model that could be applied to the KV caches. But there's no such thing as free launch. You'd still have some lobotomy, but thankfully it's kept within like 1 % of unquantized model.
38:31Harry Stebbings:Is KV caching the hardest element of building that inference infrastructure? Or is it, you name it, latency SLOs or load spikes or anything else that we could come up with? Is that the hardest? What do we not see that we should see? You can run a service and do something without having KV caching at all. You're going to economics and performance and everything else wise is going to be much worse. The dark art and magic with it is the workloads that the industry so far has found the most valuable happen to be very, very highly cacheable. So SemiAnalysis has their AgentX benchmark and suite of test data based on taking a whole lot of cloud code sessions and having dozens to hundreds of turns in those cloud sessions with subagents and everything else.
39:20And what they found is over these massive number of interactions of these like real traced code generation agentic coding sessions, about 96 % of all the tokens that go through these entire sessions are cached. So if you know your workload is going to have this extremely high caching rate where you're going to be reusing the same tokens again and again, that drastically shifts the importance of how you can retrieve those caches. Because these things get to be very, very large. We have gone into trillions of parameters. So if we just take the GPT-4 got leaked as 1.8 trillion parameter model. Now, assuming that that is int4 quantized and rounding down a little bit, that's 900 gigabytes of data size for the model weights.
40:09If we take the high expectations of Claude Fable, that's a 10 trillion parameter model, so around 5 terabytes of model weights. But the crazy thing is, at these long context links for these size models, you have the individual user sessions being in the, let's say, in the 100 gigabyte range. So with just 50 users on your service, the user context, just those individual sessions end up being greater than the model weights that you're trying to store. So that's, you know, Cloud and OpenAI have a whole lot more than 50 users. And so it becomes a really interesting trade-off of, okay, how much of the accelerator memory do you want to dedicate to weights, which you need to process every single token generated?
40:53And you want that to be as fast as possible because that sets your SLO. That sets the token latency. But if you don't have their KV caches persistent, you're actually losing a huge amount of efficiency because that was work that you didn't have to actually repeat. so it saves you as an operator money more than anything at some level you know having users kv caches be persistent will give some level of speed improvements that the user perceives but it's mostly an economics thing for the service provider where if you can return to them and use those tokens again and again that saves you money as an operator massively totally get that it saves
41:30Harry Stebbings:us money because we don't have to use as much compute but then it's harder from a memory challenge perspective, how do you think about the right logical next step then? If you appreciate the importance of saving on compute, but the challenge of memory with KV caching, what's the answer then? That we just have bigger and bigger memory stacks on chip? What does that look like? The most common deployed solution, and the vast majority of inferences out there are taking place on GPUs, followed up by TPUs and computer devices. But the most common paradigm today is you've got your GPU accelerator memory that is primarily responsible for holding the weights.
42:07And you will keep some number of user sessions on there, the ones that you're actively processing. But the larger group of users has a tiered hierarchy. So you'll have users that were around, say, in the past couple of seconds, but haven't returned, don't have an active request. That's residing in host memory. And let's say that's on the order of anywhere 4 to 10x more memory on the host than in the accelerators. So you'll be able to store more of those there. And then if someone hasn't been around in a couple of minutes, maybe a couple of hours, that's going to be stored in even further away memory.
42:37So that could be in NVMe, so flash storage, so a lot slower, but a lot larger capacity on that host. It could be in flash storage on a network attached drive. And eventually, I bet the chat GPT sessions that I had six months ago, somewhere residing on a disk, slow SSD or somewhere in a data center, but it would be dumb for them to use expensive memory to store that. So that tiering is the norm, but that introduces a huge amount of complexity of how do you decide when and where you're going to store something for your massive number of users? I would say our solution, kind of how we're trying to go about it, both from our expectation that model sizes are going to drastically increase, the number of users for all these things are going to drastically increase and the context themselves.
43:24Like two, three years ago, the typical context links were on the order of 8 ,000 to like 64 ,000 tokens. Then it got up to 128, 256, you know, a million token context links are the norm now in terms of what the model supports. But a million token context links can only hold, you know, a portion of some of like our internal companies, like largest code repositories, like it will be a fraction of that. And so if you really want an agent that can take over the capabilities of a whole team of programmers, I think the main limiter today isn't like the model capabilities itself and scaling the model size.
43:59It's on how much context can that model have of all of the data it needs to make smart decisions.
44:06Harry Stebbings:I just want to break some of the things you set up there. You said that you think model sizes will increase. I thought we were all moving to owning our own intelligence, every enterprise having their own smaller model with proprietary data. Does that go against what you think in terms of model size is increasing? Can you help me understand? Yeah, I think you can kind of break it into two tiers again. So there's going to be the frontier models and capabilities that are being really at the forefront of the development by OpenAI Anthropic, maybe Google, SpaceX AI, etc. I still think there is a long road to go in terms of getting to pushing the frontier of model capabilities.
44:44And those will continue to grow, continue to get better. And there are a lot of workloads where, let's say internal, just speaking for how Positron uses LLMs, where I don't today care that much about the cost. If I can get 10 times the output value out of a model today, I very gladly pay 10 times more per token. And I really want the frontier to push that. Some of it is cost saving. Some of it is just truly owning your proprietary data. There is a push from enterprises to have inference on site. And it's a lot more difficult to provide that for largest models. And most companies, if they're adapting from open source or developing their own model, don't have the resources to pushing the frontier.
45:25And so that is kind of what's gone to smaller models. And do you not buy that reality? I would say like right now, somewhere around 80, 85 % of all tokens consumed and produced are done by just the top four model companies. And I would say the next five or 10 % is done by the three or four after them. So I can totally buy, believe that 5 % of all tokens consumed will be done by things on-prem, locked into the big guys. But both from Positron's business perspective and just how I see the world evolving, I'm going to care more about the high volume set of things. That being said, I actually think the small model stuff is actually much more interesting for everything happening on your phone.
46:05And the amazing thing about that is I think there's this misconception that, oh, if questions could be answered by your phone or any prompts can be done locally, that actually is meaning that there's less tokens going to be used with the big guys in the cloud. I think it's the opposite. The reality is that if I have an LLM running on my laptop or phone or in my enterprise's secure on-prem cloud, whatever, that is going to be consuming data at such a fast rate of everything coming into it. And it's going to be generating analysis based on that. And it will decide, okay, what is it that I'm going to actually return to the user locally?
46:44And what is it actually requires more intelligence from a better model that isn't self-hosted? I think for any of the tokens that are being quote unquote saved by running locally, that's actually going to generate more things. Because I mean, in some ways for a simple naive use case of a personal user of LLMs, they're only going to prompt chat JBT or quad ever so often. Like they're sort of limited by their thoughts of when to actually ask an LLM something. But if they have a local LLM that is constantly checking their email, their calendar messages, et cetera, deciding to do these lookups to cloud hosted models frequently.
47:21That's now on a per person basis, a massive increase in the number of tokens being consumed and generated by the cloud models, even though there was the naive view is that, oh, there was the shift to this on-device LOM.
47:35Harry Stebbings:So can I just understand? So I completely hear you in terms of maybe 5 % of them will be in this smaller model enterprise owned kind of model landscape. Why are you so bullish then on much larger models and the size is increasing i'll break into a portion so there's the increase of the model sizes which i think of the thing that would be shared by a lot of people in the ai space is it's kind of a gut feel where based on the fact that we have seen these scaling laws like we call them scaling laws the fact that's going from a you know 100 million to a billion to 10 billion 100 billion 1 trillion per annum models we've seen this amazing increase in capabilities with that.
48:13We see that still to this day going from a trillion to five to 10 trillion at the largest end right now. We call it a law because we've observed it, but there's no actual mathematical proof that this will continue. So it's sort of on vibes that, okay, this has continued scale. There's no sign of it slowing down. So is that going to continue to 50 trillion, 100 trillion and beyond? And I don't see any indication that that's going to stop. So I'll be bullish on that.
48:41Harry Stebbings:What does scaling laws look like at three times what it is now? If AGI has been declared now by Janssen, forgive me, but what is three times this? That's a good question. I mean, GPT-6 Astra, my first 24 hours with it were basically as magical as my first experience with GPT-3.5 in November of 22. I was at the chat GPT launch at NeurIPS in 2022. And it was so funny because, you know, Sam and Ilio were there and they, you know, it was a party in New Orleans for the Noreps conference. And basically at the end, they just said, hey, we launched this little fun, you know, experiment called ChatGPT.
49:23Go check it out. And like zero fanfare, you know, it was really just a side mention. And I don't think anyone really gave it a thought at the event. When I went back to the hotel, I loaded it up and I got back at 10, 11 p.m. or whatever. And I was up for four or five hours straight just giving random prompts and that this was the most magical experience that I've ever had with a computer. I would say I got very close when Sora 2 came out that I had similar experience, short amount of time, but just mind blown by the quality of the videos. And especially the weekend that Sora 2 launched when there being no restrictions on what you could generate.
50:00but yeah gpt6 astra i do think is is agi and to to your question of like what does that mean going forward i think my guess is as good as basically anyone's but why was gpt astra so good for you
50:14Harry Stebbings:why was it comparably such a breakthrough because i haven't it's great but honestly kind of the same as before oh um in terms of the things that i've found lms to fail the most out in the past so i'll give case where it is more linear improvement. So just in terms of general coding capabilities, performance, and analyzing problems, et cetera, it is a step function improvement, but not mind bogglingly. So there are a bunch of things that other models have not been able to fix or kind of went in circles and found inelegant solutions. And it's still like a human software architect that was able to come up with a better solution.
50:50With Astra, they're just initially giving it a couple of really hard problems that I've not been able to solve with other LLMs, was able to do it one shot, having it go through a code base and find both performance improvements, bugs, et cetera, basically discovering new spaces that I didn't know existed in our bunch of portions of our code base. So that's one element, step function, but not mind boggling. The second case that was mind boggling just from a like, wow, is the computer use abilities with set of generic tools. So like being able to do blender animations, like, you know, it's become, there's a bunch of memes online of it recreating different videos, et cetera.
51:25But just the fidelity of that and where that was basically impossible with GPT 5.6 Sol was massive increased capability. And like I had it design, you know, do interior design of my house just based on a couple pictures and just like, wow, like I did not think that what's fundamentally a text model could do that. And then finally, like the biggest thing for Positron was I've been trying with every single new model release, to have these models be able to actually take a relatively simple logic design problem, implementing an encryption block in this case, and being able to take that through the full RTL to GDS flow.
52:03So from basically the specification of do this encryption function, implement the Verilog, so the hardware description language for that. So write that code and then be able to take that code and go through all the way until you have got it chip design that theoretically you could go tape out. LMs could do different portions of that and could write the scripts and fail a lot of different midpoints on the way. But a big problem with the electronic design automation tools, the EDA tools for doing chip design, is that they were designed in the 90s, early 2000s. They're really unintuitive. None of the documentation exists out in the public web, so that these models don't have a real good innate view of them.
52:43But GPT-6, with both a combination of computer use and just an ungodly, amazing scripting ability, has been able to take this KKK block and implement it with the TSMC and three PDKs and take that all the way to GDS and do that in a little over 50-something hours and meet timing over a gigahertz, etc. And that as a task, if I was giving to someone similarly new to a thing, like getting the flow mostly working, I would say would take on the order of a week. And getting it optimized to the point that Astra is at with that design would maybe be one or two additional weeks, depending on the person.
53:26So compressing that two to three weeks down to two days and change, it's still mind boggling. It shouldn't be this good at this, as I would naively think about its training set. But obviously, with OpenAI's own chip development in-house, I'm glad that those capabilities are getting added to the models they're releasing to the public and not just being kept inside.
53:48Harry Stebbings:We see Jalapeno, terrible name, I think, personally, but their own chip development, Anthropica developing their own chips, DeepSeek is supposedly developing their own chips. We see the commoditization of the chip player with everyone building their own chips. How should we think about that? As a consumer of all these things, if I take my Positron hat shirt off, I would say that that's a great thing for the industry, having fundamentally that's going to bring costs down and capabilities up and bring it to more people. I think it's such an interesting world where when I got started in the semiconductor space 13 years ago, silicon was a dirty word in Silicon Valley.
54:26And now you have all the biggest companies in the world being somehow connected to the semiconductor industry and most interesting, exciting applications and the companies building them, vertically integrating down to the silicon layer. The interesting thing with all the ones that you mentioned and the broader set is companies have the same macro goals. The implementation details are all unique, though. And that's just as an engineer and technologist is exciting to me that there are a lot of different ways to skin a cat and people can have their own architectural view and go about implementing it and get different results.
55:02And, you know, Hot Chips, the biggest conference for this design space and was where Open AI unveiled Jalapeno last month. I'm happy that the industry is still pretty open and willing to share not as much details as people would have shared, you know, five, six years ago, but still a good amount of in the open discussion of things. And so I think how that applies to Positron is, you know, we have our particular architectural views and way that we've decided to do things and that will evolve in the future as well. everyone else's. But, you know, there's still plenty of space to make bets and go in different directions.
55:40And the great thing about the market is that the market gets to decide what is valuable and those that trade value will receive a reward for that.
55:50Harry Stebbings:We spoke about context window length earlier and the expansion of it. How much does that expand? Is there infinite expansion capability of context window length? And what does that mean we can do that we can't do today? I'm just fascinated. I think with traditional linear or quadratic attention, there is going to be limits of scale when it came to what the hardware could provide. Now, one of the big things for Positron is we're trying to massively increase the memory capacity per device. With our upcoming generation, we're going to have eight times more memory capacity than the highest memory skew from NVIDIA.
56:24And NVIDIA is actually decreasing the amount of memory per device based on the market memory conditions. But I frankly think that context length going from like a million tokens that it is today, going to 10 or even much more than that is really, really hard with that quadratic expansion of memory cost. The algorithmic advancements that have happened over the past year have been very, very interesting in terms of being able to further reduce the amount of storage and compute necessary for that context with linear and sparse attention mechanisms. And those have really, really been innovated by the Chinese model labs.
57:00And this is a great example of when you have constraints of, you know, we had export controls on the chips with the highest memory capacity and flops. And so they innovated on not needing that. And so DeepSeq beginning of 2025 with DeepSeq v3 had made a lot of waves because they were able to get massive decrease in KV cache size with multi-head latent tension. So you were actually spending more flops to be able to have a smaller KV cache. And that's advanced a lot over the past year and a half. And probably the most interesting or my personal favorite right now is gated DeltaNet and its derivative versions where you can have like a 75 % decrease in the total time you're spending on the attention portion with this mechanism.
57:50Harry Stebbings:How important then is new hardware if DeepSeek without it just on architecture innovation alone can cut costs by 80 %? Yeah, I would say that there's no such thing as free launch. So when they have that MLA compression, it does come at cost of model capabilities in some form. And there's a reason why the Chinese labs have really heavily embraced MLA while none of the US labs have. I should say based on rumors, but I also have a very good information and belief that none of the major US companies are – they definitely are not using MLA and not using some of its brethren. I think that that will evolve and change in the future.
58:33But yeah, basically the sword version of that is it's not just a pure savings on that side. But the reality is to and to go back to your previous question of like everyone does want greater context length. Like if I had 10 million token context length, I think that would be enough for holding multiple of our largest code bases and really have that cross-pollination happen between them for an agentic coding model. And I guess one thing I should have explained, it's not just having the full context. There were some early models that advertised a million token context length. But as soon as you went above 64 ,000 tokens, that's two years ago, it's recallability just went garbage.
59:14So just saying that something has this maximum context length is one thing. Can it actually use that context length effectively is an entirely different thing. And that's actually going back to the Astra thing, an amazing thing about it. there's a couple of different benchmarks measuring like the long context performance. One of them is called ruler and there's like these find a needle in a haystack. So you just flood the context window with a bunch of junk, basically like just passages from books and all this. And you just put somewhere randomly in the middle of all of that, you know, a hash, you know, some value that looks out of place and you prompt the model, say, what's the secret value?
59:53and a lot of models have done really poorly on this. GPT 5.6, which is only like six or seven weeks old, only could do this about 70 % of the time. GPT 6 Astra does it like over 95 % of the time correctly. So there's a lot of room for improvement in these things.
1:00:10Harry Stebbings:Speaking of room for improvement, can I ask you, when I was doing the research for the show, I saw that the Silicon Data token price index dropped below$1 per million tokens this month. And five years ago, it was$60. per million. $60 to one. What does a million tokens cost in 2028, say two years from now, do you think? I think the thing I care a lot more about than just that 60 to one is the fact that a$60 token five years ago, no one would pay a cent for today. That was a complete garbage token, relatively speaking, five years ago. And the level of quality for a token that you pay a dollar per million tokens for now is so much astronomically more valuable.
1:00:56Harry Stebbings:And so that's because of token efficiency and what can be done? No, I'm saying just in model capabilities. If you say, okay, so it's 2026. So the best model in the world in 2021 was GPT-3. It's kind of crazy at the rate models get released today that GPT-3 was the best in the world basically from 2020. I think it was August 2021 it released all the way up to they didn't have a new release until chat GPT in November of 2022. So it was two, two and a half years between model releases. And really GPT 3.5 was just doing reinforcement learning with human feedback on the same base model. So if you remember how bad GPT 3.5 was and what was the value of that in terms of economic productivity value of GPT 3.5 versus GPT 6 today, or pick whatever comparison in points you want, the value per token in terms of what can improve a person's life, a company's business practices, et cetera, is orders of magnitude.
1:01:59I would say 100 or 1 ,000 fold. I think there's actually two points to your axis of going from$60 to$1. Yes, that's a decrease in cost. But that token today is, let's just say, I think conservatively, 100 times more valuable. So I would actually be saying that there needs to be some multiplier there as well, where like the value per unit of intelligence is probably closer to a thousand fold, not just the 60 fold you're talking about.
1:02:27Harry Stebbings:What does that mean then if we extrapolate that out to 2028? What does that mean? Does the cost of a token then actually matter? Is that the primary unit that we should measure? Because everyone talks about cost of token. Is there actually a different metric that we should measure? It is interesting that with the GPT-6 launch, Greg Brockman had said that he doesn't think that they're going to be pricing things in tokens much longer. and that they want to be moving to cost per useful result to that. And I don't think that's where it'll end up because that's really difficult to price and qualia, etc.
1:03:00But I think that the price per token is really great because you can easily calculate the cost to generate a token. So determining a margin on that and pricing it in bulk volume to generic customers is really easy. And I think that's going to stick around in large form because that is so easy. We'll see for the largest providers of tokens, how they potentially evolve their business models in terms of if you have a GPT-7 or 8 that is superhuman and can fully function as an employee in an amazing capacity. And OpenAI calculates through whatever method that's running at full tilt, etc. It's only going to cost them however many hundreds of thousands of dollars to produce tokens continuously with that.
1:03:44they may decide that it's actually easier and they'll be able to get more adoption if they just charged a million dollars a year, just using a random number to have full and limited usage of that virtual agent worker. So that may be how things evolve.
1:03:58Harry Stebbings:Can I ask you, I'm always very careful of being like the young, naive one. I'm not that young anymore, but like being the naive one who's not seen cycles. But Gavin Baker says it well when he says, I can't speak to a company that don't have numbers that are parabolically up and to the right. And just everything is better than it's ever been. What would be the first signs of a crack in the chasm, right? A shift from frontier models to open weight models, anthropic and open AI not continuing in the same level, not quite growth rate, because it's impossible to stage that, but level of growth, missing numbers next year, and then the bubble getting burst a little bit, whether the two core leaders are having some form of strife.
1:04:40I agree that that's a possibility. The reason I don't think it's likely is I think that the development of open source models and things happening locally, etc., will actually drive greater token volumes for the big guys.
1:04:53Harry Stebbings:Sorry, how does that work? I thought they were competitive. Yeah, no, I think that the smarter and more capable that Siri is on my phone is going to result on it doing a whole bunch of background tasks and things that remove me from having to be the one that instigates having requests and data be processed by even smarter models. I really do think that in a lot of AI applications right now, the bottleneck is actually a human making some sort of decision. And different tasks have different levels of autonomy that will result in things getting sent to be processed by a model, by OpenAI or Anthropic.
1:05:30But I think the next really big order of magnitude, couple orders of magnitude increase in token volumes is going to come when us humans trust a local LLM that has access to all our data all the time to have it decide to do things on its own that it is not smart enough to do. Right now, I trust Astra a lot more than myself on a whole lot of different things, but I still prompt it to do things and maybe it will go run autonomously for 12 hours or three days. I think the next big leap is when I trust a model running on my laptop or on my phone to prompt the smarter models to do even more wider set of tasks.
1:06:10Harry Stebbings:When we look at the data economy that powers the larger models, which you believe in, we see Mercore, we see Surge hitting$3 billion in revenue. How big do these companies become? Because Anthropic and OpenAI are$4 to$5 trillion businesses, say, feasible. It's wholly feasible, isn't it, that Mercore and Surge are$200 billion businesses, which serve both Frontier Labs and some of the world's biggest enterprises building their own models? I think my only skepticism there is on there being vertical integration by the frontier labs. I think the reason that hasn't happened is because Anthropic and OpenAI have better uses of their capital and mental power, etc.
1:06:55than doing all of the, you know, scale AI, Mercur, etc. work. But I think that won't always exist.
1:07:08large business, why they wouldn't just have agents be taking over those tasks.
1:07:12Harry Stebbings:Before we leave, I'd like to finish on a tone of optimism. What are you most excited about today that you think the world does not spend enough time on that we should spend time on? I'm very sympathetic to the problems that I think the smarter set of the AI alignment, AI safety community are when it comes to thinking about how do you align incentives? A lot of people just talked about AI alignment being a problem. I think there's a key part of that is human alignment. It's like, how do we, society, human race, align ourselves to have a good outcome that I think will be empowered by artificial intelligence.
1:07:49And so we discussed a bunch of the different problems that we're facing geopolitically and socially and how these different things are handled. And I think a lot of smart people are doing good work on the AI alignment problem and thinking through how we solve that, but they may be gated in what can be done there if we don't get better human alignment on regulatory frameworks and energy production, where we're going to put the data centers, et cetera, et cetera. And I think framing it as this being a similar sort of technical problem that smart people can work and reason through will hopefully get more people thinking about it that way.
1:08:29And I think a core element that gets discounted by, I think, a lot of people in that sphere. There are a lot of economic factors. And I think it's the economic factors that actually will drive real decision making and actions. And if we don't look at it from a rational, self-interested actors and all these different things, you're just not going to make progress.
1:08:50Harry Stebbings:Thomas, this has been the most varied discussion ever from education on unbelievable infrastructure evolution to Dario and Sam, You are a star. Thank you so much for joining me today. Thank you so much, Harry. But before we leave you today, the best model for your application might not exist yet. The most ambitious AI teams are training open models to beat the frontier in their domain. Now, on November the 3rd in San Francisco, Forge by Fireworks brings those teams together. People like Jensen Huang, CEO of NVIDIA, Michelle Katasta, President and Head of AI at Replit, Lin Kuo, CEO of Fireworks, and more, to hear how AI leaders are taking control of their differentiation, margins, and roadmap by owning their intelligence.
1:09:35Harry Stebbings:Learn how they're building it with inference, training, and intelligent routing. Forge is free to attend, but space is limited. Apply today at fireworks.ai forward slash forge. While Fireworks AI powers product intelligence, Asana keeps the work moving. Most companies have tried AI. Most aren't seeing results. Not because AI doesn't work, it's because AI hasn't reached the workflows yet. That's the gap Asana is built to close. Asana is the operating system for human agent teams, your easy button for AI productivity across every team. Ready-to-go AI teammates, pre-built for marketing, ops, and IT.
1:10:12Harry Stebbings:No prompt engineering, no setup. They show up where the work is happening, already onboarded in your workflows, ready to deliver. With Asana, your whole company can work on the same plan towards the same goal, whether you're a team of 10 or a team of 10 ,000. Asana, where humans and agents workflow together. Try it at asana.com. That's A-S-A-N-A dot com. While Asana organizes the work, Superhuman speeds up the day. Can I be honest with you for a second? Here's what my week actually looks like. Back-to-back meetings all day, and the real pressure builds up in the gaps between them. I'll come out of six founder meetings with a whole stack of follow-ups waiting.
1:10:51Harry Stebbings:And my inbox, it's hundreds of messages. And somewhere in there is the one thing that really matters, which I always tell my mother is her. I've tried AI tools for this, but they always lived in another tab. I'd have to stop, go find it, paste in the context, write the perfect prompt. It was just enough friction that it never really stuck. Superhuman Go from the makers of Grammarly, totally different. It's right there where you need it, when you need it, in the doc, when I'm reading, the email I'm writing, When a thread gets dense, I highlight it and ask Go to pull out what matters without losing my place.
1:11:24Harry Stebbings:And after a run of back-to-backs, I ask what was actually decided. And Go turns it into clear action items or drafts, the reply I need to send. It's already up to speed and I'm always the one in control. Try Superhuman Go from the makers of Grammarly and find out more at superhuman.com.
From the publisher
Thomas Sohmers is the co-founder and chairman of Positron AI, building chips to make running AI dramatically cheaper and more energy-efficient. The company recently announced an $875 million Series C at a $5 billion valuation, backed by investors including Gavin Baker's Atreides Management, NEA, Valor Equity Partners and Netscape co-founder Jim Clark.
AGENDA:
04:30 Why Does AI Inference Need Different Hardware from Training?
10:25 What Is Nobody Telling You About AI's Token Economics?
13:10 Should We Really Slow Down the AI Frontier?
16:55 What Happens If America Slows Down—and China Doesn't?
19:30 Should We Restrict China's Access to AI Chips?
20:15 Why Aren't Zuckerberg and Jensen Backing an AI Slowdown?
21:25 Are We Being Misled About Data Centers?
23:25 Can the West Beat China While Drowning in Regulation?
26:20 How Many Planned Data Centers Will Actually Get Built?
27:15 Data Centers in Space: Real Opportunity or Elon Hype?
28:10 Is Energy AI's Biggest Bottleneck?
29:55 Could the Debt Boom Derail AI?
31:25 What Is KV Caching—and Why Does It Change AI Economics?
35:25 Does Compressing AI's Memory Make It Less Intelligent?
41:05 How Do We Solve AI's Exploding Memory Demands?
43:45 Will Every Company Own Its Own AI Model?
47:10 Why Bet on Ever-Bigger Models?
48:20 If We Already Have AGI, What Comes Next?
53:25 What Happens When Every AI Lab Builds Its Own Chips?
57:30 If DeepSeek Can Slash Costs with Software, Why Build New Hardware?
59:50 How Cheap Will AI Tokens Be by 2028?
1:03:15 What Would Be the First Warning Sign of an AI Bust?
1:03:55 Could Open Models Break OpenAI and Anthropic's Growth?
1:05:55 Could Mercor and Surge Become $200 Billion Companies?




