OpenAI’s Compute Chief: We Can’t Build Fast Enough | Sachin Katti

16 Jul 2026 · 44 min · 26 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

OpenAI’s “industrial compute” push to build AI supercomputers fast enough for demand, covering data-center scale (“factories turning electrons into tokens”), liquid cooling, power/grid constraints (and behind-the-meter generation), custom silicon (“Jalapeno” optimizing tokens per watt), and Stargate as an umbrella compute strategy.

Key claims

demand far outstrips compute supply; the main risk is not overbuilding but failing to build enough physical capacity because factories/supply chains can’t scale quickly; AI will accelerate chip and system design via more experiments; data centers are net positives for local communities and use shockingly little water via closed-loop recycling.

Notable examples

liquid-cooled “giant football fields”; investing in grid upgrades; nuclear as a dense energy source; Abilene data center for training; Oracle building Stargate data centers; MRC networking protocol to keep 100,000-GPU training running despite link failures; guaranteed capacity for “tokens of intelligence.”

Guests

Sanchin (Sachin) Katti—Head of Industrial Compute at OpenAI; Stanford professor (networking), multi-time founder, former CTO at Intel; background across academia, startups, and Intel.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Introduction to AI Compute Challenges

0:00 to 0:30

Learn about the current compute supply challenges in AI development.

“Anytime we have thought we have enough compute, we can slow down, always negatively surprises like, oh, we should not have slowed down.”

The Scale of Compute Infrastructure

1:26 to 2:32

Discover the scale and urgency of the infrastructure build-out at OpenAI.

“We are recording this on the sidelines of the RACE conference in Paris.”

Rapid Growth in Compute Spending

2:32 to 3:41

Understand the financial implications and growth of compute spending in AI.

“We compute that historically from my previous role, for example, at Intel, we probably take months to make given the magnitude of those decisions.”

Building New Muscles in AI

3:41 to 4:48

Learn how OpenAI is evolving its approach to building compute capabilities.

“So obviously not quite a pivot because obviously the all AI research is going full speed ahead.”

What is a Data Center?

4:48 to 5:52

Get a detailed understanding of what a data center entails in the AI context.

“So it does absolutely feel like a new muscle that we're building in the company.”

Cooling Technologies in Data Centers

5:52 to 7:10

Explore the innovative cooling technologies used in AI data centers.

“And so I think the best way to visualize data centers is giant factories that are turning electrons into tokens.”

Power and Energy Management

7:10 to 10:13

Learn about the power generation and energy management strategies for data centers.

“And is a cooling technology that is being used something that's well understood and it's just getting deployed or is there like fundamental new things happening in cooling right now?”

Nuclear Power and Its Potential

10:13 to 12:01

Discuss the role and potential of nuclear energy in powering data centers.

“not just for data centers, but also for households.”

Introduction to Jalapeno

12:01 to 13:32

Understand the significance of Jalapeno in OpenAI's chip strategy.

“But I'm curious, and we'll go into some details of Alapino later.”

Inference vs. Training in AI

13:32 to 14:00

Explore the evolving dynamics between training and inference in AI compute usage.

“So we look at it as a very critical ingredient and scaling how we deliver intelligence to the world.”
Show all 26 chapters

The Importance of Inference in Compute

14:00 to 17:45

Learn how inference has become a significant part of AI compute and its implications.

“Because a lot of training is now inference.”

Data Centers: Local Impact and Benefits

17:45 to 19:42

Discover the positive impacts of data centers on rural communities and local economies.

“supply chains, factories don't move that fast, cannot add capacity that fast.”

Myths About Water Consumption in Data Centers

19:42 to 21:16

Understand the misconceptions surrounding water usage in data centers and their actual impact.

“And so I think one of the things we are investing a lot in is explaining local benefits every time we build a data center somewhere and making sure that it is well understood the kind of upside that this has.”

Factors in Site Selection for Data Centers

21:16 to 22:26

Explore the key factors that influence the selection of data center locations.

“Since we talked about data centers at the beginning of this conversation, why do OpenAI and other companies pick rural areas?”

Sachin Katti's Role in OpenAI's Compute Strategy

22:26 to 27:32

Learn about Sachin Katti's responsibilities and the organizational structure of OpenAI's compute efforts.

“We're going to go into all of this in more detail, but let's talk about you a little bit and your journey.”

Current Compute Resources and Partnerships

27:32 to 28:00

Get an overview of OpenAI's compute resources and key partnerships in the industry.

“We are building the largest computer in the world.”

OpenAI's Compute Strategy Overview

28:00 to 28:34

Learn about OpenAI's current compute partnerships and their strategy.

“something which is very attractive, obviously.”

Building on Compute Partnerships

28:34 to 29:16

Discover how OpenAI plans to evolve their compute strategy through partnerships and infrastructure.

“We have compute from effectively many sources.”

Stargate: A New Compute Framework

29:16 to 30:21

Understand the Stargate framework and its role in OpenAI's compute evolution.

“more options where we design the compute the data center itself ourselves or potentially even build the data center ourselves so all of those are ways of scaling the amount of compute that we have.”

Current Data Center Developments

30:21 to 31:58

Explore the current state and future developments of OpenAI's data centers.

“For example, with Oracle Close Partnership, we help them design.”

Oracle’s Role in Data Center Expansion

31:58 to 33:51

Learn how Oracle is expanding data center capacity for OpenAI's needs.

“so it's super threat and you're seeing the results you're seeing how quickly the models are becoming more capable it's because of these kinds of compute.”

Rapid Chip Development: The Jalapeno Case

33:51 to 36:14

Discover the reasons behind the quick design and tap-out of the Jalapeno chip.

“So across the board, like everything, you're building stuff as well.”

Innovations in Networking Protocols

36:14 to 38:47

Learn about the MRC networking protocol and its importance for large cluster fabrics.

“where AI will design the systems it needs to train and run the next generation of AI.”

Addressing Bottlenecks in Computing

38:47 to 40:02

Examine the various bottlenecks faced in the computing industry today.

“So it seems like the nature of the bottleneck keeps changing in the computer industry.”

Guaranteed Capacity for Enterprises

40:02 to 42:00

Understand OpenAI's strategy to offer guaranteed capacity for enterprise customers.

“What's the story behind that and the strategy?”

The Future of Space-Based Compute

42:00 to 43:30

Explore the potential of data centers in space and the future of orbital compute.

“But I think that's what intelligence is becoming.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Anytime we have thought we have enough compute, we can slow down, always negatively surprises like, oh, we should not have slowed down. Demand far outstrips compute supply today. So anything we can bring online, we consume immediately. Our biggest worry is that still, at the scale at which we are trying to get compute and build compute, the physical world does not move that fast. We do believe that the world of recursion is not that far, where AI will design the systems it needs to train and run the next generation of AI, including this. Hi, I'm Matt Turk. Welcome to the Matt Podcast. My guest today is Sartin Kharty, who holds what might be the most fascinating and relevant title in tech right now, head of industrial compute at OpenAI.

0:40Sartin has an incredible background. He was a professor at Stanford, a multi-time founder, and most recently the CTO at Intel. Now he's leading what many are calling the largest infrastructure build out in human history. In this episode, we step away from the model layer and dive deep into compute and the physical reality of the AI boom. We talk about the staggering scale of the data centers being built. We get into the weeds on liquid-cooled supercomputers, power grid constraints, the potential of nuclear energy, and OpenAI's move into custom silicon with Jalapeno. We also discuss the broader Stargate strategy and the slightly self-aware reality that AI is now beginning to help design its own chips.

1:20It is a fantastic look behind the curtain at what it actually takes to power the future of intelligence. Please enjoy my conversation with Sanchin Khati. All right, Sanchin, welcome. Excited to do this. We are recording this on the sidelines of the RACE conference in Paris. So thank you for braving the heat. It's another heat wave here. Thank you. Great to be here. Great to be here. To start, some people describe what's currently happening in the world of compute data centers as the largest infrastructure build out in history, bigger than the highway, bigger than the railroads. And I'm curious, one, if you agree, and two, what it feels like from the inside.

2:08How do you view what you're currently building at OpenAI? It definitely feels like one of the largest things humanity has ever built, effectively. Definitely bigger than many of the things that I've heard of. I'm not old enough to have experienced the highway build-out, But no, it feels exactly like what it sounds. It sounds in the belly of the beast, so to speak. Every day is we are making decisions. We compute that historically from my previous role, for example, at Intel, we probably take months to make given the magnitude of those decisions. but the demand is so insatiable and it is growing so rapidly that we have to move very quickly.

2:57So it's an intense time but it's probably the most exciting thing an engineer would want to be part of. Yeah, and I read somewhere that OpenAI was planning on spending about 50 billion in compute this year. Is that still the rough number directionally? Actually, that sounds about right. And the whole industry itself was going to be$700 billion in compute spend this year as well. So insane numbers, it seems. Yeah, and it's probably continuing to grow, right? And a lot of build happening. So a lot of that is also going to translate to compute usage from people like us in a year or two. Is the right way to think about this that for OpenAI, it's a bit of a new world, right?

3:49So obviously not quite a pivot because obviously the all AI research is going full speed ahead. But like building a whole new business within the company, is that fair? Is that how people think about it? Because obviously building models is one thing, building data centers, it's a whole different world. Yeah, I think OpenAI always has had a fundamental belief that computers are the foundation of everything. Computers are the foundation for intelligence. And the way we keep continuing to scale intelligence and distribute intelligence is by having compute. So that has never been different. It's never always been the belief.

4:29I think what's becoming clear is to build the kind of compute we need and at this scale. We have to not just rely on getting compute from our partners. We increasingly have to take a much more active role in building and getting that compute that we need. So it does absolutely feel like a new muscle that we're building in the company. And maybe to anchor the conversation from the beginning, it would actually be very helpful to talk about what a data center is in reality. So I think everybody knows that data centers are being built. But, you know, gun to one's head, I'm not sure that everybody could say, well, what is actually being built?

5:16Because we've been building data centers of the industry for cloud for decades at this point. So what is politically different and new about the data centers that we're building for AI today? I think the biggest probably is the scale, right? So we are essentially building large supercomputers as we think about AI. And as we build intelligence and deliver intelligence and models become more capable, we use it for more and more complex tasks. we need more and more bigger computers effectively. And so I think the best way to visualize data centers is giant factories that are turning electrons into tokens.

6:00That's a popular phrase nowadays. But it actually has a ring of truth to it. So how do we take power, how do we take those electrons and actually use it to power chips that effectively are delivering intelligence? And the way I visualize it is large football fields. liquid cooled because these chips run really hot. The temperatures on these chips are very, very high. And so you have to cool them with liquids. You can't cool them with air. So a lot of liquid cooled, basically refrigerators, effectively, that are sitting alongside the building. And on that topic, while we're at it, the cooling happens at the data center level?

6:41What does it happen at the chip level or both? Both. So you need to cool the data holes, but you also need to cool the chips individually because it's not going to be enough to do one or the other. And you also have to cool the things that connect chips, right? And so that's why you need cooling pretty much everywhere nowadays. Even the cables that are the transformers that distribute the power become too hot, so they also need to be cooled. So everything that processes energy produces heat. And is a cooling technology that is being used something that's well understood and it's just getting deployed or is there like fundamental new things happening in cooling right now?

7:21I think liquid cooling has been around for some time, but has never been deployed at this scale. And so the innovation is more around how to make it reliable, how to make it cheaper, more scalable, right? So there's a lot of innovation around that. There's also a lot of new innovation, new kinds of liquids, new kinds of materials that can absorb heat better because anything that can improve the efficiency of heat transfer is very important for data centers. So we can then run the chips hotter, right? And there is a direct correlation between running a chip hotter and how powerful the compute is.

7:59So the hotter the chip, the more memory bandwidth you get, the more flops you get. And so there's a strong payoff. If you can cool well, that also means you can produce more intelligence. All right. So giant big factories, lots of cooling. The other part that seems to be very critical to any discussion is power and energy. So how does that work starting at a high level? Do you connect to the grid? Do you have your own power generation? I think the early days, we all connected to the grid. And we still all would want to connect to the grid. At this point, we are beginning to hit, and we are investing in generation infrastructure for the grid, transmission infrastructure for the grid.

8:44So whenever we build a data center anywhere, we make it a hard commitment that we are not taking power away from the grid. In fact, we are investing in the grid to generate new power so that we can consume it for data centers. What does that mean practically? You have a grid somewhere. It has a certain generation and distribution capability, a certain number of megawatts. Obviously, a data center shows up. If there was spare capacity, then of course the data center can use it. But if there isn't spare capacity, then we have to add new gas or solar or hydro generation infrastructure. To the grid.

9:24So we are investing and funding that build-up. And then you have to build transmission lines, invest in transformers, substations to distribute that power. So wherever we're building data centers, we are funding the development of all of that infrastructure. And so that's one of the things that we do want to emphasize, which is this is infrastructure that would otherwise not have been funded, if not for these data centers. And one of the side benefits of this big data center build-out is the grid infrastructure of America, and the whole world for that matter, is getting upgraded very quickly. And so that's the power piece.

10:05So whenever we can do that, we do that, and we consume power from the grid, but it's also good citizens of the grid because we are improving the infrastructure for everyone, not just for data centers, but also for households. In some places, we are beginning to hit the limits of how much grid power we can build and consume. And so there everyone's looking at behind the meter. We are also doing some behind the meter generation where we would have on-site power generation and distribution capability. That does not come from the grid. But in fact, the data center becomes effectively self-sufficient in terms of power.

10:43There's gas turbines. What is it? Today it's gas turbines, especially in the U.S., because that's the most dense transportable form of energy. And also the one that is quite widely available in the U.S. But there will be a bottleneck by the supply chain. Do you think the nuclear conversation is interesting? I think maybe we're recording this in France, which has a bunch of nuclear power generation. Nuclearists have come back to the discussion in the U.S. as well. Is that something that you think about or you think is interesting? Absolutely. It can't come soon enough. I think it is the densest form of energy we can all produce and consume.

11:27And it's also key. So I think definitely would be a good source of massive scalable energy for our data centers. Obviously, outside of France, the rest of the world has a lot of catching up to do and building this infrastructure. But I think it's going to play a very important role to the data center. Okay, so that's a great introduction on data centers. The other interesting bit of news that you guys recently had is Jalapeno. So now for OpenAI, in addition to being in the application business, consumer and enterprise, and then being in the model AI research business, and then the computing data center business, since that OpenAI is a mini chip business, if that's fair.

12:17So a completely full stack. But I'm curious, and we'll go into some details of Alapino later. But into the overall strategy, where does that fit? As we begin to serve a pretty big fraction of the world's population, AI usage is exploding. Inference is obviously becoming a big fraction of our work. It's consuming a lot of compute. And one of the other realizations is because we know what is the work toward exactly, what is the model we want to run, we can co-design the hardware to be super efficient and delivering those models. And so the strategic thesis we have at Halapenia is how do we take advantage of knowing what the end workload is, what the model itself is, and design chips that are very efficient in serving those models.

13:13So it really allows us to drive efficiency advantage, drive more tokens per watt. So the key metric that Jalapeno is optimizing is maximizing the number of tokens you can produce per watt. And because the world is constrained by power today, so the more tokens you can produce for the same number of power, order, it's better for everyone. So we look at it as a very critical ingredient and scaling how we deliver intelligence to the world. Great. So I'll come back to help in a second, but you mentioned, you just mentioned inference and it's such an interesting evolution as well without commenting on necessarily what's going on at OpenAI specifically.

13:55Is inference equally big or much bigger than training these days in terms of like usage of compute as as have we shifted uh from uh you know being uh those are very heavy pre-training runs uh as the major uh use case for compute to now just inference being the majority no inference is a big perhaps even the majority on compute and and i think one of the i we don't like to make a distinction between train training and inference. Because a lot of training is now inference. So when we train a new model, we are generating synthetic data, for example. That's inference. When we train a new model, we are doing post-training.

14:41And that's inference. When you train a model, you're doing test and compute. That's all inference. So when we say training, a lot of the compute actually is inference, even in that phase of the work. So in terms of the fundamentals, building work. Yep. Obviously, I cannot resist asking the inevitable question around the potential risk of overbuilding, given the lag between demand and usage and how long it takes to build a data center. Then you mentioned somewhere that you were deliberately very paranoid about the problems ahead in the next three years, very paranoid about the surprises ahead, which sounds like a very healthy approach.

15:28So how do you think about that? Is there any way to mitigate that? Or is it just like, how do we play a deep belief that this is the future and we should all just go, go, go? We have deep conviction in scaling, right? And history has brought us up. So effectively, our Yetni, for example, has tracked compute. We tripled compute and we tripled revenue. And we believe that, I mean, that continues to be true. Demand far outstrips compute supply today. So anything we can bring online, we consume immediately. So there's no compute that is going fixed for us. So I think that conviction has not changed whatsoever.

16:13And if anything, we are saying that scaling laws on research and training continue to hold. and potentially the pace at which we are doing research is accelerated because of AI itself. So AI is doing a lot of AI research now. And so one of the subtle implications of that is previously our researchers used to run experiments and they needed compute to run experiments. But the number of experiments they could run was limited by the number of human researchers they had, which is a scarce resource on the board. There's not a lot of people who can do AI research. Now, AI itself can do AI research. The number of experiments we can run explodes.

16:57And therefore, the amount of compute you need for research also explodes. So we don't see a world where we will have unused to less compute for the foreseeable future. When I was referring to surprises, my worry is more on the downside of we are not able to actually build all the compute we want. and

17:21because that is consistently the case. Any time we have thought we have enough compute, we can slow down, always negatively surprise us like, oh shit, we should not have slowed down. And so our biggest worry is that still and at the scale at which we are planning to get compute and break compute, the physical world does not move that fast. Physical supply chains, factories don't move that fast, cannot add capacity that fast. So for us, the surprise is more on that direction than the other direction. Very fascinating. You alluded to communities a minute ago, and obviously that's a key debate. So curious about your perspective on a spectrum where, you know, on the one end, one extreme, you'd say, well, the AI industry and computer industry has a PR problem, and there's no problem.

18:18It's just the problems that we cannot explain it well enough. To the other extreme, actually, those communities have a point. What do you think the reality is? I think any time there's new technology, which is as revolutionary as this technology is, there is always disruption that's going to happen. But inevitably, we have learned this over history that this always leads to better outcomes for society. And so how do we draw a line from where we are today to that outcome and explain to the world why this is the trajectory we all need to be on? I mean, it's our responsibility to do that. On the community side, there's a liberalism, local versus a global issue.

19:08The communities, I think data centers are, even today, a net positive to every community. Because we are building these data centers in rural areas of America, for example, where there's nothing else that is being built on this scale. So we show up in rural Texas, we build a data center that produces new property tax receipts for the community, that funds schools, that funds hospitals. we show up and we invest in new grid infrastructure, which otherwise would never have happened because there's no demand. So there's a modernized grid that AA can enjoy. We've seen we produce jobs. And so I think one of the things we are investing a lot in is explaining local benefits every time we build a data center somewhere and making sure that it is well understood the kind of upside that this has.

20:03And data centers, once they're built, are essentially very clean citizens. They don't produce any gases or toxic chemicals or anything like that. They're self-contained. They just produce intelligence. The typical question that comes up is water. And I think that's been debunked quite a bit by research that maybe give us a scholar on how you all think about the water for she. These are liquid cooled, and the liquid is recycled. So we do actually, the water consumption of a data center is shockingly small relative to household water consumption. So I think, as you put it, it's been debugged. It's a misperception that data centers consume a lot of water.

20:51It's anything they consume so little water for what they do, and all of that water is recycled. So we don't net consume new water. Once we get to a particular point, the water just gets recycled as we use to get cold. Yeah. So all those stories of brown water just don't make sense because the water at a data center happens in a contained circuit. It's a closed loop. Yeah, it's a closed loop. It's a closed loop. By the way, you mentioned Texas and rural areas. Since we talked about data centers at the beginning of this conversation, why do OpenAI and other companies pick rural areas? How do you select a site for a data center?

21:30So many factors. So one is, of course, land, like plentiful land. Number two would be permitting, like can we build these things? And we want to build these things such that they are not affecting any neighborhoods. So land that is somewhat removed is their ideal candidate. Of course, access to power, right? So a strong grid, strong gas availability, all of those are important factors. And then four is labor, right? So how quickly can you build these things? So availability of labor, construction labor, qualified electricians, pluggers, all of this can be able. So all of those factors go into every single site selection decision.

22:17And I know we obviously Texas isn't popular because it fits a lot of these criteria, but it's not the only state. I mean, we have data centers all around the back of LA Plus. Okay, great. We're going to go into all of this in more detail, but let's talk about you a little bit and your journey. So you're the head of industrial compute at OpenAI, which, by the way, to the beginning of this convention, just the title industrial compute is so, I mean, such a perfect title for the moment we're in. But what does that mean? What is the role and how is this whole effort organized within OpenAI to the extent you can talk about it?

23:00Yeah, I think think of it as my role and our team's role rather as how do we bring compute online at Industry FK. That's effectively what we're doing. That's the entire life cycle. So how do we find the ingredients that go into compute? Land, power, shells, chips. How do we finance them? Because these are master dollars. And so how do we make sure that we finance the grid infrastructure? How do you finance the construction of the inform shells? How do we finance the chips? Then it's about how do you operationalize all this? So how do you actually make sure these things happen in time? They stay up.

23:44How do you operationalize all of this infrastructure? So it's an entire life cycle. And then, of course, how do you actually use the compute? So a big part of my role is capacity allocation inside OpenAIR. So it is always a scarce resource. You're a very popular guy. I am not very popular. There is always someone who is unhappy with whatever decisions he makes. But, yeah, me, actually, our team provides the input to make the capacity allocation decisions. So we surface what are the different choice points and what are the what-if questions on different allocation choices. We have the capacity planning and then, of course, using that to forecast how much capacity we need where.

24:31Because it's not just more compute, it's also where, what kind, what shape, what chip, what workload you want to run there. So all of those are this team figures out kind of what should be the forecasting and planning. and that informs, that throws us a loop. So that informs, where do I go find the next chunk of land and power and chips to put into? And again, without going into anything confidential, although I guess when you guys go public, all of this will be public, but is that thousands of people at this stage? Is that multiple different teams? Or do you guys outsource a bunch of things and work with a bunch of contractors?

25:11It's a portfolio approach. So we are never going to be in a world where we outsource everything or build everything ourselves. It's always going to be fixed because that's the reasonable thing to do. So you don't want to put your rights all in one basket. So we will have hyperscalers providing a big chunk of a majority of a compute. We will have new clouds, parts of a portfolio. We will be partnering with design firms that can build the compute that we need. And, of course, we will build some of it ourselves. And so we are always going to have a portfolio approach because at the scale which we need, we will need to tap into all sources of compute.

25:56We can't just rely on one particular mechanism. And your background, the default of this, is you are both a professor-centered and an entrepreneur or a founder, or mostly an academic? Like, just walk us through your journey. Bit of all of the above. But yes, I'm at my heart. I'm an academic. So I've been a professor at Stanford since 2010. Recently. What do you focus on there? I was a faculty in computer science and electrical engineering. With a particular interest? Yeah, my area of research was networking. Networking. So I already built networks, both mobile and wireless networks out of the net.

26:38So, there isn't any networks, actually. But three, four years ago, while I was at Stanford, I did a couple of startups. The last startup got acquired by VMware. And that's how I got to know Pat. Pat became Intel CEO. And that's how I ended up at Intel. Most recently, before coming to OpenAI, I was Intel's CTA. And so, I've kind of seen all the different things, academia, startups, corporate in Intel, and then, of course, a mix of all of the above in Opera, because we have a research lab, a startup, and a fast-growing company all mixed into one by Opera. Yes. Is that what you said? What did you say yes to the job when the job came up?

27:25It's actually what I just said. That makes this so unique, and it's hard to find anywhere, right? Because you always have to choose that having a world-class research environment coupled with the hardest technical problems. We are building the largest computer in the world. And so there are a lot of new problems that we need to solve. But also a fast learning business. I like problems that sit at the intersection of business, technology, and strategies. And so there's a very unique time in history and a unique role. something which is very attractive, obviously. Very cool. All right, so going into a bit more specifics about OpenAI's compute strategy.

28:12So maybe let's summarize what you guys currently have. So I think there's some Microsoft, you did this big$20 billion deal with Cerebrus. There's a bunch of things around this target. Maybe just give us the lay of the land of what you currently have, and then we'll talk about what's your equation building next. We have compute from effectively many sources. So Microsoft, obviously, is a big partner, a important partner. We also have compute, as we have announced, from AWS, Ionogogl. So we have compute from all of the IPvascaters, effectively. We also have compute from Corby, for example, so a near cloud.

28:55And then, of course, compute that chip partners are supplying now actually it's up directly pretty pretty compute for us so I think that's the mix roughly today as we go forward obviously there'll be building on all of these relationships but also looking at more options where we design the compute the data center itself ourselves or potentially even build the data center ourselves so all of those are ways of scaling the amount of compute that we have. So, I think, coming back to my earlier answer, it's, the answer always will probably be me try, but it's all of the above. Yeah, yeah. That's helpful, and obviously, diversification makes all of the sins in the world, given the scarcity.

Read the full transcript

29:48And then, it seems that it started as evolved, so from what was going to be a joint venture, with Oracle and SoftBank to what now seems like it's more like an umbrella term for the compute strategy these days. Is that a fair way to describe it? Yeah. I mean, we look at Stargate as our compute strategy, and it is varying degrees of us designing or building the compute ourselves. For example, with Oracle Close Partnership, we help them design. We help them with how to operate AI compute, which is effectively a new kind of compute for us. We work with SoftBank Energy, which is public. We basically have co-designed the Wormshell with them, and they're executing on that Wormshells.

30:46and we will be kind of figuring out how to operate our chips in these data centers ourselves, using the new chips. So Stargate to us is that umbrella strategy across all of these different things. And think of it as an evolution that we continuously be on because it's never going to be tomorrow we wake up and do only one kind of way of building compute. I think Stargate to us is a continuous way of learning how to scale compute, and we're adding on more and more capability. And as part of that, there are data centers being built, right? Like Abilene Techs. Yes. So maybe walk us through that. What is currently being built for people to have a situational awareness?

31:36We obviously have a big partnership with Arthur. That's the Abilene Data Center. that's where for example we are training our newest models so very excited about that that's a very big GB black belt cluster for our needs so that's up and running it's up and running it's being used for training the last two models so it's super threat and you're seeing the results you're seeing how quickly the models are becoming more capable it's because of these kinds of compute. Are there new systems that are currently being built that are good? Yes. Yes. So Oracle is building a number of data centers, all of which are public, so across Michigan and Texas and other places.

32:28So these are coming online in the next couple of years as they get built, and we put whatever chips on that time training who are the latest. but these are really meant to be very big clusters that allow us to do both training but also product inference kind of compete. Great. But the way those deals are structured, the Oracle is the prime, building those. Oracle is the cloud. And you're the core tenant. Interesting. How does the, for the stuff that you're building, how does the financing strategy work? I mean, obviously, you all raised, what was it,$122 billion? Was it the number recently? So there's no shortage of cash, although, you know, given all those expenses, I don't know.

33:20But, you know, the core ways of the world are famous for being very strategic users of debts. Is that part of the financing strategy as well? How do you all think about it? I mean, with all of these compute that we have today, we have amazing partners who actually are handling that for us. So Microsoft, Google, Amazon, Oracle, we are the offtake. So we are the tenants, as you put it. So we commit to consuming that compute, to buying that compute, whether it's online. So across the board, like everything, you're building stuff as well. So the partners are building. Partners are building. And you're always across the board the tenant, not the owner.

34:11Correct. Okay. So therefore, financing is outsourced to apartments. Okay. We talked about Jalapeno. Let's go into a bit more detail there. because in particular it seems that you guys went incredibly quickly in design like I read somewhere in nine months from design to tap out so maybe maybe walk us through that and what was the reason it went so quick yes it was incredibly quick nine months is very very fast probably the fastest I've seen in my career I think several reasons one is it's a team. It's a strong team. Many of the teams have designed TPU chips at Google in the past. So very well experienced team.

34:57We have a great partner in Broadcom that has a very strong track record of delivering XPOs ASICs. So I think a strong partnership with Broadcom in making this happen. Three, I think perhaps an OpenAI unique point. in most chip companies or when you design chips, you don't know what you're designing it for because you're a vendor, the customer who eventually runs the workload is someone different. There's a unique advantage here of us knowing what the future models might look like and therefore being able to short circuit a lot of the decisions you need to make, design decisions you need to make on the chip side.

35:37So that's super helpful. And finally, increasingly, AI itself helping design and optimize the chip, that is usually the one that takes the longest time because you're basically limited by how much human time is there to process all this data and run the experiments. And we can do a lot of those iterations much faster using AI. Yeah. So AI is building its own chips now. Yes. I think that world is not very far. I mean, AI is assisting in chip design, but we do believe that the world of Likershyn is not that far where AI will design the systems it needs to train and run the next generation of AI.

36:20Including chips. Including chips. Including chips. You also released a few weeks ago MRC, which is a networking protocol. Workers show through that. What is it and why is it a big deal? It's a new networking protocol routing technology, if you will, to scale these really large cluster fabrics. So imagine you have 100 ,000 GPUs. They need to be connected together. And when you're doing large training runs, they're constantly communicating with each other because these models are so large that the processing of the models is happening over the entire 100 ,000 GPU cluster, for example. And you can imagine the number of links and switches and NIC cards that need to be there to connect all of these chips together.

37:15At this scale, CDIs are fun, right? It happens all the time. You can't really even enumerate all the ways things could fail. So the strategy behind MRC is how do you design algorithms and protocols that can gracefully mask all these failures and make sure that the training workload does not get impacted. The network is an abstracted system that the training job does not have a way. It's always going to be there. It's always going to find a path, even if a link fails. So it is all about reliability. It is all about availability. So how do we design protocols that contain the complexity of such a big cluster and make sure that we don't get stopped because of failures, which are very common in the system?

38:10So MRC is like a multi-part spraying protocol where you can spray packets or multiple parts. So between any two chips, there are many, many routes to get there. Kind of like between any two points in the city, there are many routes. So instead of picking just one route, we will send traffic across all of them. And whichever one succeeds, you take that. And so that way, even if any one of them fails, it's not a showstopper. So that's kind of the basic intuition behind it. It's obviously a lot more sophisticated than that. But doing this at scale and that speed is hard. And so that's why it's quite innovative.

38:46What are the bottlenecks that you experience these days? So it seems like the nature of the bottleneck keeps changing in the computer industry. Obviously, people talk a lot about memory these days. Is that one of them? What else? I think there are bottlenecks everywhere, to be honest, in this application. So I don't think there's any one, right? We have bottlenecks in the data center building itself around permitting, around availability of gas turbines, transformers. those industries have not added much capacity over the last decade or so ago. And they've suddenly experienced a demand shock. And it takes years before you can add capacity to produce motorbines and transformers.

39:33So they're trying to catch up. And to the jump thing that you mentioned earlier, is there a shortage of electricians and technical people, trade people that know how to build those things? Absolutely. Absolutely. So that's why we should all do as AI replaces a knowledge worker, we should become electricians. No, I think there is definitely a shortage of electricians, plumbers, all kinds of traits, you name it. So anything we can do to train more folks to be able to do those, there are very well-paying jobs that a lot of us, all of the hyperscalers, all of the labs would actively hire you for. if you had the like conversations so that is definitely a bottleneck and i think they become an increasing bottleneck because we are all trying to build more and more and we have a limited number of things these capabilities right let's talk for a minute about the the business side of things another thing you launched recently is a guarantee capacity for customers uh to lock in compute um that which is interesting right it almost feels like open is also becoming an utility company providing compute to others?

40:49What's the story behind that and the strategy? Yeah, so guaranteed capacity is guaranteed tokens, right? So we are effectively saying we will guarantee you a certain dollars worth of tokens of intelligence. I mean, it makes sense, right? So in a world where compute is a shortage, therefore tokens are always going to be at a premium and there's a shortage of tokens that we can produce given the limited compute that we have. And as this becomes a fundamental input to the enterprise, right, so enterprises are going to need intelligence more and more of it to run, right? And so this is a way of enterprises gaining assurance that the tokens of intelligence that they need will be there for them.

41:31And so they don't have to take business risk. And so I think it's good business hygiene. if you have a critical supply resource and every enterprise intelligence is kind of the most important supply item, if you will. It makes business sense to make sure you secure that supply. And so that's the demand we're trying to fulfill. I know it's a new concept, like what does it mean to have a guaranteed capacity for the intelligence? But I think that's what intelligence is becoming. It is becoming a supply unit for every digital enterprise. So maybe to end on a fun one, and you can answer either with your OpenAI hat or not, data centers in space, is that exciting?

42:20Is that science fiction? Is that needed? Is that something that people like to talk about just because it's cool or where do you land? For the geek in me and the engineer in me, it's definitely exciting. It's one of those things that can be super cool to see a constellation of satellites producing compute. I actually think it will become feasible, right, as engineering problems we solve. But they can be solved with enough time and investment. Whether it is needed, I think there is room for orbital compute. I don't think it's going to serve all the compute beats, but it's definitely going to be a component in the arsenal.

43:04I think what we are waiting for is when does the economics of the launching satellite change, and when does the economics of the hardware change, because we need to get to a point where it's cheap to launch the hardware, and if something fails, it's cheap to throw it away. I think we can't go up and fix it, unlike on the ground. And so I think that inflection point, hopefully it happens soon, And at that time, I'm a bit unswilled. Great. So Chile has been wonderful. Thank you so much for spending time with us. We appreciate it. Thank you. It's been great to have this chat. Cool. Hi, it's Matt Turk again.

43:37Thanks for listening to this episode of the Mad Podcast. If you enjoyed it, we'd be very grateful if you would consider subscribing if you haven't already or leaving a positive review or comment on whichever platform you're watching this or listening to this episode from. This really helps us build a podcast and get great guests. Thanks and see you on the next episode.

From the publisher

Is the AI industry actually overbuilding, or is the physical world moving too slowly to keep up? In this episode of the MAD Podcast, OpenAI's Head of Industrial Compute, Sachin Katti, takes us inside the "belly of the beast" of what may be the largest infrastructure project in human history. We explore the staggering physical reality of the AI boom—from $50 billion supercomputers and liquid-cooled data centers that "turn electrons into tokens," to overhauling the U.S. power grid and exploring nuclear energy. Sachin also pulls back the curtain on OpenAI's Stargate strategy, their move into custom silicon with Project Jalapeno, and the mind-bending reality that AI is now beginning to design the very chips that will power its own future.


(00:00) — Cold open: “One of the largest things humanity has ever built”
(00:30) — Welcome: Sachin Katti, Head of Industrial Compute at OpenAI
(01:44) — Is this the biggest infrastructure buildout in history?
(03:41) — Why OpenAI is building a new industrial muscle
(04:54) — What an AI data center actually is
(05:27) — “Factories turning electrons into tokens”
(06:35) — Why AI data centers need liquid cooling everywhere
(08:10) — The power problem: grids, generation, transmission, substations
(10:43) — Behind-the-meter power and gas turbines
(11:02) — Why nuclear “can’t come soon enough”
(11:49) — Jalapeño: why OpenAI is designing its own AI chips
(13:19) — Tokens per watt: the new metric that matters
(13:38) — Why inference may now dominate AI compute
(14:58) — Is OpenAI overbuilding compute?
(16:47) — Why OpenAI thinks the bigger risk is not building fast enough
(17:55) — Communities, jobs, water, and the local data-center debate
(21:16) — How OpenAI chooses data-center sites
(22:25) — What “industrial compute” means inside OpenAI
(25:59) — Sachin’s path: Stanford, startups, Intel, OpenAI
(28:05) — OpenAI’s compute portfolio: Microsoft, hyperscalers, neoclouds
(29:37) — Stargate explained
(31:21) — Abilene, Oracle, and the next wave of AI data centers
(32:48) — How massive AI compute gets financed
(34:05) — How OpenAI designed Jalapeño so quickly
(35:59) — AI is starting to help design AI chips
(36:20) — MRC: the networking problem behind 100,000 GPUs
(38:47) — Bottlenecks: transformers, turbines, electricians, supply chains
(40:29) — Guaranteed capacity: intelligence as a supply unit
(42:08) — Will AI data centers move to space?

More from The MAD Podcast with Matt Turck

All 44 episodes
OpenAI’s Compute Chief: We Can’t Build Fast EnoughThe MAD Podcast with Matt Turck · 44 min
Listen in VO