In short
Token maxxing as a vanity metric and the real-world “AI compute crunch” caused by hardware and power bottlenecks.
Key claims
AI demand (tokens and compute) is rising faster than supply; multiple bottlenecks overlap and won’t fully ease until ~2028+.
Notable examples
Anthropic’s Claude Code/Claude rate limits (power users running 24/7); OpenAI pausing new Sora signups to redirect scarce GPUs; GitHub stopping new Copilot subscribers; unreliable on-demand H100 access without reserved capacity. Bottlenecks: (1) GPUs: 36–52 week NVIDIA lead times; Blackwell sold out to mid-2026; TSMC CoWoS packaging sold out through 2026; NVIDIA ~60% of CoWoS via 2027. (2) HBM memory: HBM demand ~5x since 2023; SK Hynix/Samsung/Micron sold out into 2026; RTX 50 production cut 30–40%. (3) CPU: agentic AI needs ~1 CPU per GPU vs ~1 CPU per 12 GPUs for chat; Intel cites billions unmet demand. (4) Electricity: grid/transformer lead times 18–36 and 18–24 months; Gartner expects 40% of AI data centers restricted by 2027; local moratoriums.
Guests
No guests mentioned; only host Jon Krohn and references to company executives/analysts.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOUnderstanding Token Maxing and AI Compute Crunch
0:46 to 2:11
Explore the concept of token maxing and its relation to AI compute limitations.
“You might have noticed this personally over the past several months.”
The GPU Bottleneck
2:12 to 4:02
Delve into the issues surrounding GPU availability and the challenges faced by AI developers.
“Lead times on NVIDIA's data center GPUs are now running 36 to 52 weeks.”
Memory Constraints in AI
4:03 to 5:45
Examine the challenges posed by high bandwidth memory shortages affecting AI systems.
“So I mentioned memory in bottleneck number one, and it's the memory itself that forms another bottleneck.”
Rising CPU Demand
5:46 to 7:20
Discuss the increasing need for CPUs as AI systems evolve towards more complex tasks.
“So 12 times as many CPUs per GPU relative to what we were expecting in the Gen AI era up until now that the Gen AI era has morphed into the agentic AI era.”
Electricity as a Critical Bottleneck
7:21 to 9:59
Understand the implications of electricity supply issues on AI data centers and expansion.
“So through 2024, the binding constraint on building a new AI data center was capital and chips.”
Optimism Amidst Challenges
10:00 to 12:15
Discover reasons for optimism in AI development despite current hardware constraints.
“even an aggressive ramp up now wouldn't deliver meaningful relief until 2028 or later.”
Transcript
Automatic transcript. May contain errors.0:00Jon Krohn:This is episode number 992 on token maxing and the AI hardware bottleneck. Welcome back to the Super Data Science Podcast. I'm your host, Jon Krohn. The hottest social media trend in AI today seems to be token maxing, wherein folks use agentic tools like Cloud Code and OpenAI's Codex to maximize the number of tokens they consume. I don't know why someone would think this is an actual indicator of productivity, but apparently it stems, this trend of token maxing stems from companies tracking AI usage via internal dashboards. I hope you and your company aren't engaged in this vanity metric nonsense.
0:46Jon Krohn:And relatedly, today's topic on the podcast is the AI compute crunch, how the breakneck rush to train and deploy ever larger language models with people using ever more tokens from them is now slamming into very real, very physical bottlenecks across the global supply chain. You might have noticed this personally over the past several months. If you're a heavy Claude code user, you probably bumped up against the weekly rate limits Anthropic introduced last August. Anthropic was unusually candid about why. They said that a small fraction of power users were running Claude 24-7 in the background and eating into capacity for everyone else.
1:24Jon Krohn:The company has been blunt that they are, in their own words, very compute constrained. OpenAI also famously paused new Sora signups to redirect scarce GPUs. GitHub at one point stopped accepting new subscribers for its co-pilot bot. And on the developer side, on-demand H100 access, a popular NVIDIA GPU, through AWS, Azure, or Google Cloud has become genuinely unreliable for any team without pre-reserved capacity. So what's actually going on? In short, demand for AI compute is rising substantially faster than the underlying hardware supply chain can scale. And the bottleneck is no longer any one thing.
2:02Jon Krohn:There are at least four overlapping strategies that have to be solved simultaneously, and solving any one of them doesn't actually unstick the system. Bottleneck number one is the most famous, GPUs themselves. Lead times on NVIDIA's data center GPUs are now running 36 to 52 weeks. And Blackwell allocation, the next generation after those popular hopper style H100s, Blackwell allocation is largely sold out through mid-2026 with reported backlogs in the millions of units. Even those older NVIDIA H100 chips, originally launched back in 2022, have actually gotten more expensive to rent on the spot market, somewhere around 30 % more expensive since November because customers can't get hold of newer hardware and are falling back on previous generations.
2:50Jon Krohn:But here's the wrinkle. GPU fabrication itself isn't really the binding constraint anymore. The real choke point sits one layer up at TSMC, the world's largest dedicated independent semiconductor foundry who have a monopoly on state of the art chips. So, yeah, the real choke point is with them with TSMC and something called COWAS, Chip on Wafer on Substrate, the advanced packaging step that bonds GPU dyes to their high bandwidth memory stack. Whether that phrase means anything concrete to you or not, the main takeaway on this is that TSMC's CEO, CCWay, has publicly stated that COWAS capacity chip on WaveRound Substrate is sold out through 2026, and NVIDIA alone has reportedly locked up roughly 60 % of TSMC's total COWAS allocation through 2027.
3:39Jon Krohn:TSMC is racing to nearly quadruple monthly co-auth output by late 2026. But as that TSMC CEO, CC Wade, dryly puts it, there are no shortcuts. Building a new fab, a new fabricator that creates these chips takes two to three years. All right, so that's bottleneck number one, the GPUs themselves. Bottleneck number two is the memory side of that same equation, high bandwidth memory or HBM. So I mentioned memory in bottleneck number one, and it's the memory itself that forms another bottleneck. So every NVIDIA H100 needs 80 gigabytes of HBM3, so high bandwidth memory, third generation, and every Blackwell B200 needs 192 gigabytes, so three times as much of another kind of high bandwidth memory called HBM3e, which is fifth generation and 25 % faster than the HBM three third generation high bandwidth memory.
4:41Jon Krohn:So total HBM demand, high bandwidth memory demand has roughly quintupled since 2023. And there are only three companies on the planet that make this stuff, SK Hynix, Samsung, and Micron. All three say their HBM supply is largely sold out well into 2026. And Nvidia has reportedly cut consumer RTX 50 graphics card production by 30 % to 40 % in the first half of this year because the same memory fabs feeding HBM lines are also responsible for memory on consumer devices. For example, you know, the kinds of consumer devices that are used to render video games on people's personal computers. New HBM fabs, similar to the way that with bottleneck number one, I was talking about how fabs take two to three years to make for GPUs.
5:29With HBM fabs, it's kind of the same story. It takes 18 to 24 months to have a new HBM fabricator come online and demand is expected to outpace supply for at least three more years. All right. So bottleneck number one was GPUs. Number two was memory. Bottleneck number three is one that I find particularly interesting because it's caught a lot of folks off guard
5:51Jon Krohn:And that's CPUs. As workloads shift from training to inference, especially as agentic AI proliferates, that is AI systems that plan multi-step tasks, call tools, and coordinate work across many GPUs, the CPU to GPU ratio in the data center is climbing dramatically. Analysis from Morgan Stanley suggests that chatbot style systems like, you know, the chat GPT interface that we were first introduced to a few years ago, that needs roughly one CPU for every 12 GPUs, whereas agentic systems require a one to one ratio. Whoa. So 12 times as many CPUs per GPU relative to what we were expecting in the Gen AI era up until now that the Gen AI era has morphed into the agentic AI era.
6:38And so this is good news for CPU manufacturers like Intel.
6:44Jon Krohn:Intel said on its Q1 earnings call this year that its server CPU shortfall in CFO David Zinsner's words starts with a B, meaning billions of dollars of unmet demand for CPUs. Server CPU prices, therefore, have jumped 10 to 20 percent in just the past couple of months. That, incidentally, is a big part of why Intel's market cap has more than doubled over the past six months. All right, we're on to the fourth and final bottleneck, and this is arguably the most intractable. That's electricity. So through 2024, the binding constraint on building a new AI data center was capital and chips. But now, every major hyperscaler has reported that their buildouts are gated not by money or hardware, but by grid interconnect timelines, which can run 18 to 36 months, and transformer lead times, which is another 18 to 24 months.
7:45Jon Krohn:On top of that, Gartner now projects that power shortages will restrict 40 % of AI data centers by 2027, and local communities are pushing back hard. As of March, at least 12 U.S. states had filed data center moratorium bills. Maine's legislature, for example, actually passed a statewide moratorium in April. The governor vetoed it, but more than 50 local moratoriums, moratoria, I'm not sure of the plural, but more than 50 local moratoria have already passed across the U.S. Communities near hyperscale clusters in Virginia, Texas, and Georgia are already seeing electricity rate increases of 8 to 15 percent.
8:22Jon Krohn:And that, along with concerns about water consumption and land use, is fueling the backlash. To put the scale of the demand wave in perspective, the five biggest hyperscalers, Alphabet, Amazon, Meta, Microsoft, and Oracle, they're on track to spend something on the order of$725 billion combined on capital expenditure in 2026. roughly three quarters of which is targeting AI infrastructure. That's up from about$120 billion total in 2022. That's a six-fold increase in four years. Anthropic alone has in just the past few weeks committed over$100 billion to AWS for up to 5 gigawatts of GPU capacity or AI compute capacity, locked in roughly another 5 gigawatts from Google and signed a deal to take all of the SpaceX Colossus 1 site for more than 300 megawatts and over 220 ,000 NVIDIA GPUs, which you'll notice is exactly the capacity injection that allowed Anthropic to recently double Claude Code's rate limits.
9:27And yet, perhaps the most remarkable
9:30Jon Krohn:pattern across the data I've been digging through is the asymmetry between hyperscaler capital expenditure and the hardware suppliers who actually have to build the underlying gear. While the hyperscalers have tripled their combined capex over the past two years, the chip makers, equipment vendors, networking gear suppliers, and cooling specialists have only increased theirs by about half. Those hardware suppliers are wary of overbuilding capacity that could sit idle if the AI boom slows. And given the timelines involved, even an aggressive ramp up now wouldn't deliver meaningful relief until 2028 or later.
10:06Jon Krohn:Elon Musk has announced plans for what he's called a terafab, intended to produce more processing power per year than today's entire global semiconductor industry, but even it isn't expected to start production until 2028 at the earliest and at a fraction of the envisioned scale that year. Now, before you write off the whole AI build out as doomed, there are real reasons for optimism as well. The key thing here is algorithmic efficiency, which continues to improve dramatically. Google's TurboQuant, for example, announced in March. I've got a link to that for you in the show notes. Briefly tanked memory stock prices because it promised to materially reduce how much memory inference workloads need.
10:45Jon Krohn:Custom silicon is also proliferating. So Amazon's Tranium 2 cluster powering Anthropics training is already operational with over 500 ,000 chips, and the hyperscalers are increasingly looking to complement, but apparently not replace, NVIDIA. And as I've discussed many times on the podcast, there's still enormous headroom for further LLM efficiency gains in both training and inference through techniques such as mixture of experts approaches, which you can hear more about back in episode number 778. The summary overall for me is this. The gap between AI software's voracious appetite and the messy multi-year physical reality of building chips, packaging them, manufacturing memory, and wiring up substations is going to be a major story for at least the next two to three years.
11:29Improving an LLM takes weeks. Building a fab takes years. Building transmission lines can take a decade. That mismatch will shape who gets to scale, who gets throttled, and quite possibly which AI companies survive the next economic cycle.
11:46Jon Krohn:But for those of us building applications on top of these models, the picture is actually pretty exciting still. Compute scarcity is forcing the entire industry to get dramatically more efficient, better quantization, smarter inference scheduling, smaller specialized models, custom silicon that's economical to run and scale. Every one of those efficiency gains makes the AI applications you can dream up cheaper to deploy and easier to put into the hands of more people. Just don't bloody token max as you do it. all right um and we have time for a apple podcast review here at the end of the episode um we have one here from someone named the liz with a whole bunch of z's or z's depending on what country you're from the liz w who says that this is the only ai podcast you need thanks liz w you.
12:40She, I assume, goes on to say that you know how we all have that one podcast that we go to when we need consistently valuable information, but don't have a lot of time. For me, it's the Super Data Science AI Podcast with John Crone. She goes on to say that hands down, my go-to source for keeping a pulse on what matters. The content is always timely and speaks grounded truth for practitioners. I've learned more from the podcast than any other source available. Fantastic. It continues to go on, but I'll stop there. Thank you very much for that wonderful, positive five-star review. Thanks to everyone for all the recent ratings and feedback on Apple Podcasts, Spotify, and all the other podcasting platforms out there, as well as for likes and comments on our YouTube videos.
13:25Really appreciate all that. You know, I read all the feedback, which is helpful for adapting the show for the future. And that also reminds me that we did. I'm very grateful to all of you who filled in the brand survey that I posted on social media and mentioned once on the podcast as well. Nearly 600 people filled out that 40 question questionnaire around our podcast brand. So we'll have, yeah, we're digging into those data now and we'll have some kind of results to tell you about in the near future. All right. So yeah, thanks for continuing to support the show, uh, interacting with us and yeah, letting other folks know about the podcast through, um, yeah, your comments and that kind of thing.
14:14If you write written feedback, um, on air on Apple podcasts, I will read it on air just like I did today.
14:26Thank you.
From the publisher
While “tokenmaxxing”, the social media trend of maximizing AI token consumption as a vanity metric, takes off online, the physical infrastructure behind AI is slamming into serious bottlenecks. In this Five-Minute Friday, Jon Krohn maps out the four overlapping supply-chain constraints choking AI compute: GPUs (with NVIDIA Blackwell sold out through mid-2026), high-bandwidth memory (quintupled demand since 2023, only three manufacturers worldwide), CPUs (agentic AI requires 12x more CPUs per GPU than chatbots), and electricity (Gartner projects power shortages will restrict 40% of AI data centres by 2027). Find out why the five biggest hyperscalers are on track to spend $725 billion on AI infrastructure in 2026, where the reasons for optimism lie, and why Jon says you should definitely not tokenmaxx.
Additional materials: www.superdatascience.com/992
Interested in sponsoring a SuperDataScience Podcast episode? Email natalie@superdatascience.com for sponsorship information.




