In short
How Lightning AI’s VP of Infrastructure, Frank Basso, designs and provisions high-density AI data centers end to end—covering GPU/compute nodes, liquid-to-chip cooling, power/cooling constraints, networking (east-west, north-south, management/out-of-band), storage performance, and operational “lights-out” reliability.
Key claims
AI data centers differ from generic ones mainly by density (e.g., moving from ~18kW racks to much higher), heavy cabinet weight, and extreme power/cooling needs. Lightning uses liquid-to-chip (not immersion) and plans capacity using reference architectures plus conservative power derating (D-rating to ~81% safety), contracting for available grid power (e.g., 20MW rooms). They avoid network oversubscription and build multi-tenant segmentation without sharing GPU nodes.
Notable examples
Chicago “liquid-to-chip” GB300 super pod: ~10,000 chips in a ~10,000 sq ft room with ~20MW power; serviceable units sized around GB300 (1,152 GPUs per unit; multiple units per building).
Guest
Frank Basso is VP of Infrastructure at Lightning AI (Los Angeles), responsible for physical data centers, DCO teams, network engineering/operations, infrastructure, and platform engineering. Background: Lightning AI is a US “NeoCloud” with 35,000+ modern GPUs (growing toward ~50,000) and >$500M ARR.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOWhat Makes an AI Data Center Unique?
0:50 to 2:10
Frank Basso explains the differences between AI data centers and traditional data centers.
“This episode of Super Data Science is made possible by Anthropic, Cisco, Excel Data, and Grobi.”
Challenges in Building AI Data Centers
2:10 to 5:40
Frank discusses challenges in building AI data centers, including power consumption and design constraints.
“You know, we've had tons of episodes about open source, open source Python libraries.”
Co-location Process for AI Data Centers
5:40 to 10:48
The conversation dives into the co-location process used by Lightning AI for GPU provisioning.
“And I don't think I've mentioned on air yet that this is a lot of GPUs.”
Future of AI Data Centers and Infrastructure Planning
10:48 to 14:01
Discussing future trends in AI data centers and the complexities of planning infrastructure.
“But they're charging us per megawatt per month as a base fee plus our usage from the utility.”
Challenges of NeoClouds vs. Hyperscalers
14:01 to 15:23
Learn about the difficulties faced by NeoClouds in comparison to hyperscale providers.
“I've heard people in the industry say, NeoClouds have it the worst because you never know what your customers are going to be up to versus people who build it for themselves, like XAI and OpenAI.”
The Internet of Cognition by Cisco
15:23 to 15:47
Explore Cisco's initiative to create infrastructure for AI agents to share knowledge.
“They're publishing the architecture and building reference implementations.”
Understanding NeoClouds and GPU Focus
15:47 to 16:46
Discover what defines a NeoCloud and its emphasis on GPU performance and bandwidth.
“Wow, that is a fascinating insight into how you're doing this.”
Performance Constraints in Data Centers
16:46 to 18:40
Examine the challenges and strategies for managing traffic and storage in data centers.
“through our private network interconnects across our backbone.”
Bespoke Infrastructure for Client Needs
18:40 to 21:18
Learn about the importance of customizing infrastructure based on client demands.
“So the nodes, it depends really on how you're running it.”
Designing GPU Nodes
21:18 to 22:43
Understand the specifications and challenges in designing GPU nodes for data centers.
“Liquid cooling systems are extremely expensive to install and operate.”
Show all 27 chapters
Liquid Cooling Systems Explained
22:43 to 24:48
Gain insights into the workings and benefits of liquid cooling systems for GPUs.
“A node or a GPU node is effectively the actual box itself, the server that is a GPU server.”
Heat Rejection Techniques in Data Centers
24:48 to 28:00
Explore modern heat rejection methods used to maintain temperature in data centers.
“It gets heated by the chips, and then it gets cooled down in another part.”
Heat Rejection in Data Centers
28:00 to 29:43
Learn about heat rejection techniques used in modern data centers and their environmental implications.
“water coming through and the air is cooling the water those techniques are pretty much banned everywhere now.”
Debunking Water Usage Myths
29:43 to 30:12
Discover how misconceptions about water usage in data centers are addressed.
“The nimbyism, the not in my backyard, hey, you're taking all our water.”
Networking in Data Centers: East-West vs. North-South
30:12 to 34:24
Understand the differences between East-West and North-South networking in data centers.
“And maybe now you can also debunk the concerns of people who saw TikTok about all the water usage that the new data center is going to use in their area.”
Management and Out-of-Band Networks
34:24 to 35:39
Learn about the management networks in data centers and their importance for operations.
“If they're having issues getting to some endpoint, you can look at all of that through your out-of-band network.”
Lights-Out Operations and GPU Management
35:39 to 36:56
Explore the concept of lights-out operations in data centers and GPU management strategies.
“things thing about gpus are they get run so hard that they do break and there's a certain percentage of failure.”
Managed Kubernetes Services Explained
36:56 to 38:28
Gain insights into managed Kubernetes services and how they simplify operations for clients.
“Yeah, they can effectively plug into our system directly and it just makes it easier for them.”
Understanding East-West and North-South Connectivity
38:28 to 40:06
Delve deeper into the concepts of East-West and North-South connectivity in data centers.
“this esoteric performance thing that I'm trying to obtain.”
Design Considerations for Data Centers
40:06 to 42:00
Explore design considerations for data centers, including distance limitations and data hall organization.
“So when you're, so when you see those photos, you see these long kind of like hallways that the servers make up.”
Understanding Data Center Connections
42:00 to 44:25
Learn about the intricate connections and design considerations in AI data centers.
“You could have as many as you wanted, basically.”
The Unique Environment of AI Data Centers
44:25 to 47:16
Discover the challenges of noise levels and safety in AI data centers.
“just what it's like to be in one of these AI data centers.”
Personnel Safety and Gear in Data Centers
47:16 to 51:09
Explore the safety measures and protective gear required for data center workers.
“It is a very loud industrial environment.”
Electricity Usage and Regulation in Data Centers
51:09 to 56:03
Analyze the impact of data centers on local electricity usage and regulations.
“And you have to wear a wristband if you ever touch or open a box.”
Data Centers and Power Generation
56:03 to 58:15
Learn how data centers interact with power grids and the role of fuel cells.
“These gas-fired plants are called peaker plants, which you see going in that, like, hey, Goliath is using these peaker plants.”
Energy Future and AI's Role
58:15 to 1:02:30
Explore the potential of AI in advancing sustainable energy solutions.
“That was obviously a well-practiced set of information.”
Book Recommendation and Final Thoughts
1:02:30 to 1:06:36
Discover a book recommendation that ties into the discussion of electricity.
“And so, yeah, AI has for a while been helping reduce consumption, even as we build more of these AI data centers.”
Transcript
Automatic transcript. May contain errors.0:00Jon Krohn:We've done over a thousand episodes of this show on every layer of the AI stack except the one that physically runs all of it, the data center. Today, we finally fixed that. Welcome to episode number 1003 of the Super Data Science Podcast. I'm your host, Jon Krohn. You are in for a treat with an exceptionally interesting episode today with Frank Basso, who is vice president of infrastructure at Lightning AI, the US-based startup that has over 35 ,000 modern GPUs, over$500 million in ARR, and that makes it a lot more. makes it easy to go from AI idea to product lightning fast. In this episode, Frank explains how he builds the physical AI data centers that allow us to do all the mind-blowing things that we do with AI.
0:43Jon Krohn:He digs into GPUs, compute nodes, liquid cooling systems, and much more. Enjoy this special episode. This episode of Super Data Science is made possible by Anthropic, Cisco, Excel Data, and Grobi. Frank, welcome to the Super Data Science Podcast. How are you doing today? I'm doing quite well. Thanks for having me. It is my pleasure to have you on. We are actually colleagues. We should probably get that out of the way for our listeners. We both work at Lightning AI. I'm a fellow there, which is quite a loose role, but you have a very specific role at Lightning AI. What are you up to there? Oh, wow.
1:20Jon Krohn:Yeah, fellow. That's anything we throw your way to get it done in any way we need, right? Exactly. I really love those ambiguity within that titlage. So, no, I'm Frank Bassa. I'm VP of Infrastructure. I'm responsible for basically everything that plugs into the wall from the physical data centers and the DCO teams through network engineering and operations and infrastructure and platform engineering. And where are you based? I'm based in Los Angeles, California. Los Angeles, the data center capital of the world. If only electricity cost half as much here. Yes, it would be. And so I'm so excited to do this episode with you, Frank, because while we've done over a thousand episodes of the show, they've all been on a different part of the AI stack.
2:10Jon Krohn:You know, we've had tons of episodes about open source, open source Python libraries. We've had lots of episodes about tools, software tools, that make it easier to build and deploy AI models, just like Lightning AI does. And we even have done episodes on GPUs and on chips, but we've never talked about the physical centers that actually run all of those chips. So I think potentially an interesting place to start, because this is Lightning AI specialty, What makes an AI data center different from a generic data center for computing? That's a great place to start. I think the differences are density, right?
2:54So traditional data center, even hyperscale data center, upwards of five years ago, the maximum power density you'd see, Microsoft had the highest at 50kW racks, which were kind of insanity. But the industry standard was 18s. and 18 kilowatts. That's one server now. And so that's a huge differentiator. Also, the physical data centers themselves weren't built to handle the densities of not just power and cooling, but in the space itself, raised floor environments are no longer the case because these cabinets weigh upwards of two tons now versus they were eight or 900 pounds before. Now they weigh 4 ,200 pounds.
3:41And so this has presented a lot of design challenges and constraints when it comes to what we knew five years ago versus what we're building for today and into the future. And in the future, it becomes even more exciting.
3:52Jon Krohn:Yeah, even more power consumption in the future, I understand from the research that we did for this episode, which is actually kind of an interesting thing. I was going to ask you this later on, but just kind of seems to fit in nicely here with you talking about the power consumption. How do you build a data center that today, say, needs to run H100s when you know that more B200s are coming and the next generation after that is going to be even more power dense? Well, you run into some physics constraints with these designs, and that's what many are solving for. Different companies are solving for them in different ways.
4:26NVIDIA's way is to basically co-locate more equipment together, meaning the higher density. Right now, we went from eight GPUs in a stack to liquid gold model where you have 72 or 144, 288 and on and on that kind of doubling within the same cabinet or what is a new factor cabinet footprint and getting all the way up to half a megawatt of power. And currently, we've recently deployed GB300s, a 10 ,000 chip super pod in our Chicago data center. That's a liquid to chip cooled solution. It's a 10 ,000 square foot room with 20 megawatts of power running in it. 20 megawatts of power. That used to be more than an entire 100 or 200 ,000 square foot data center.
5:17Now it's in 10 ,000 square feet in a room. Of course, there's 50 ,000 square feet of supporting infrastructure like chillers and cooling pumps and UPSs and generator lineups and all of those things that you need to support the heart of that building, which is that room.
5:32Jon Krohn:Frank, that sounds like a ton of different variables to think about when you're getting a center together. My understanding is that today, in 2026, Lightning AI does something called co-location when it's provisioning the infrastructure for all of its GPUs. And I don't think I've mentioned on air yet that this is a lot of GPUs. So Lightning has over 35 ,000 GPUs. And so when Lightning's building a new AI data center, they use something called co-location. Tell us about that process and how it works. Yeah, that's a great question. Okay. And just so you know, we're always expanding. So by the end of this year, it'll be more like 50 ,000, which is kind of the growth curve that we're going through at Lightning now.
6:16But co-location is same as you, slightly different than you'd seen in the past. Like it used to be, you'd say, hey, I need a cabinet of gear in a data center. You'd go and lease it out with a small amount of power and you'd install your kit and you'd get your providers lined up and you'd be on the air. That was great if you were to co-locate. Co-location at this scale is more hyperscale. When you start talking about a minimum of 10 megawatts of power and upwards, it's a slightly different conversation. We have a number of data center partners that we work with, and that list is getting longer by every day almost.
7:00and basically we work with them and say hey what's the available grid utility power that they have available to the site and whether that's a new build or a brownfield retrofit of an existing facility and how much it load can they support in the building and if it's an existing center there may be some fixed constraints that this is what we have we have 10 megawatts and that's what we have and there's no room for expandability or they might say hey there's another 20 megawatts in the parking lot at the substation. And so then we talk about a build. And so we then take the amount of available power and work with the data center provider.
7:38And in this case, everything moving forward is liquid to chip because it's more efficient overall and consumes less overhead and has better what they call, you know, PUE or efficiency rating of actual workloads in the building because the cooling is directed at chip. It's not going to air to the chip and then back to air, and then you have to cool the air. You lose a lot of efficiency with that if you're using it in a mechanical way like that. So we work with them on a design, and the design of available power will tell us how many chips can we put in there of a certain type, say GB300. We say, okay, we want to build in GB300, a serviceable unit is 1 ,152 GPUs.
8:22That's 16 cabinets of gear. And we want nine serviceable units in a building. That's 10 ,300 and change GPUs. And then we need on top of the GPUs, well, we need all the networking, networking cabinets, and we need the storage and the storage subsystems. And we need all the internet routing and all the compute, all the CPU to support because GPUs don't run on their own. You need lots of VMs to do all kinds of things. I'm sure you talked about on this bad podcast at length. And so there's a ratio and that ratio is increasing for CPU. So there's a lot of other things you have to put into your design constraints.
9:04Then you have to do load calculations based on what is our usage profile going to look like for this gear? Because you have what's called a plate rating. Say the plate rating is 22 megawatts, like based on what it says on the side of the box, you know, you're going to draw 22 megawatts of power. But the reality is that plate rating of the box means the maximum amount of CPU or the maximum amount of memory installed and the maximum amount of drives for our configuration with a lower amount of drives or a lower amount of memory per box or what they're basically designed for the customer. There's no way that you could actually draw that much power.
9:42So we apply a D rating. We do a safety rating at 81%. It's no industry secret. Most companies use 70 % to 75%, which I think is a little bare, but we do a conservative number at 81%. So we come out at, instead of being whatever it is, we come out at under 20%. And we know that power envelope is what's available in the building is 20 megawatts. Even though the plate rating or the delivery and the infrastructure is designed for more, it actually will never exceed 20 megawatts. And that's what we're actually going to be paying for. And that's what we contract for. And so after that point, we have a design and the data center provider says, okay, great.
10:28This is how much it'll cost per megawatt to build. And that varies a lot depending on where you are in the world and how much things cost or whether you have union labor or things of that sort, hugely variable from one to 4 million per megawatt right now. And that's just cost to us. The cost to them is probably six times that. But they're charging us per megawatt per month as a base fee plus our usage from the utility. So the base fee per megawatt is how they make back all their money over the term of a contract. And the contract these days is usually 10 years in length.
11:03Jon Krohn:Wow. It is wild to be planning over those kinds of timeframes. When you think about things like you were talking about how CPU usage is going up, And you can correct me with these kinds of like rough finger in the air numbers. But my understanding is that in the pure Gen AI era, when we were talking about just the chatbot experience, when you're in something like ChatGPT or Claude, that was something like a 12 to 1 GPU to CPU ratio. But now as we're increasingly in this agentic era, agents are doing more processing and kind of coordinating of tasks that then get sent out to GPUs. And that means we're getting closer to a one-to-one ratio of CPUs to GPUs.
11:46Jon Krohn:And that's just over a timeframe of a few years. So it's pretty wild to think how complicated it is to be planning 10 years ahead on what kind of infrastructure, what kinds of power demands you're going to have in a given center when the industry, when the AI industry is changing so quickly. No, that's true. It is. The good news is we have a little help. NVIDIA produces what they call RA or reference architectures. Now, NVIDIA reference architectures are a starting point. They're effectively the minimum viable product or bare minimum you need to do to make this system work in a performant manner.
12:26And it includes what your layout should look like, what serviceable units are, what your network should look like, whether you're InfiniBand or Rocky on your East-West GPU interconnection network or on your North-South network. And then the internet access you may need for that. They have recommendations. But we use those as starting points. A lot of companies build that or build somewhere not near that. It's all over the place, which predictability is becoming a very interesting thing in the consumer side of the market. But we take and we build on top of reference architecture. For instance, we do no oversubscription within our network in any way, shape or form.
13:07And that's common on the East-West network. You don't oversubscribe.
13:10Jon Krohn:What does that mean, no oversubscription? Oh, oversubscription. This is a very old networking term that means say you have two things connected at 800 gigs of speed. I'm using a modern number. And you only have eight from the switch. You only had 800 gigs of uplink. Well, you're oversubscribed two to one. So we're not. If it goes into a switch, you have two ports going 800 gigs each. There's two ports that are going up to the next layer of the network to match that. And so we never oversubscribe. This is a bad thing because of contention within the network. And also, the RAs are mostly designed for single tenant systems, where we're a multi tenant cloud.
13:58And so you never know what the customers are going to do. I've heard people in the industry say, NeoClouds have it the worst because you never know what your customers are going to be up to versus people who build it for themselves, like XAI and OpenAI. Those guys are like, we build to support our engineers and we don't have to do anything like NeoClouds. So they do one specific tasking for their designs and it makes it a lot easier, just like Hyperscalers did. Predictable workloads, predictable outcomes. For us, it's the next customer comes in and says, well, what do you think about X? And we go, Ooh, okay.
14:33Uh, we can do that. And, uh, sometimes they're like, wow, no one said you could do that before. I'm like, well, we designed for future headroom and future performance levels that people aren't thinking about when they built. We have to, because otherwise we don't future-proof our systems.
14:49Jon Krohn:Quick reality check for anyone building with AI agents. Your agents can discover each other. They can pass messages. They can coordinate on tasks, but here's what they can't do. They can't think together. When your agent figures out how to handle a complex workflow, that knowledge stays isolated. The industry has focused on scaling AI vertically, bigger models, more compute. Those breakthroughs matter, but intelligence also scales horizontally. Agents sharing knowledge across a network, coordinating on common intent, reasoning together, the infrastructure for that second horizontal axis doesn't exist yet.
15:23Jon Krohn:Outshift by Cisco is formalizing it. They call it the Internet of Cognition. They're publishing the architecture and building reference implementations. Read Scaling Out Superintelligence. We've got a link to that in the show notes. Then check out episode number 961. In it, Dr. Vijoy Pandey, the head of Outshift by Cisco, walks through how horizontal scaling of intelligence works and why it matters. Wow, that is a fascinating insight into how you're doing this. For people who aren't aware or who haven't listened to the episode that I did about three months ago with the CEO of Lightning AI, Will Falcon, tell us a little bit about what a neocloud is, like what Lightning AI is.
16:04Neocloud meaning unique is the ultimate definition, non-hyperscale, even though we build things at hyperscale. So the NeoClouds are being unique or bespoke industry versus the normal Amazon, you know, the Amazon, Google, those type of systems that we provide. We're focused on the GPU, the GPU and the GPU performance. And then secondarily, the memory bandwidth and performance, and then the storage bandwidth and performance, because it's insane compared to what a normal cloud provider would have to put up with. We have customers that they may be just doing training, but they're doing reinforcement learning and they're doing a distributed basis and they're pushing almost a terabit of traffic in and out the front door on the internet access side.
16:58through our private network interconnects across our backbone. And we've built a substantial network of not just, hey, we're connected to the internet, but we build regional networks within the metro that interconnect our data centers to all the internet exchange points within the region. Plus we have a nationwide backbone that connects all of our data centers together so we can shuffle and move data on our client's behalf and also run them as one cohesive system. I mean, we're effectively running a telecom carrier-style network in North America to tie together all of our sites. Nobody really does that.
17:36But maybe since I had, as a recovering network engineer, I knew better. And so when I joined, I went, oh, this is going to be a problem. So we started on that a year ago, and it's just come to fruition, and customers are very happy about the upgrades. But that same customer that's pushing all that traffic is also pummeling the storage arrays at like per GPU, a gigabit or gigabyte per GPU per second is not unheard of. And any normal storage arrays would just have lots of wait times and hour classes waiting on those apps in the background and not be performant. But we have to design for those kind of performance constraints.
18:18We have lots of demanding clients. And with inference, it becomes even more challenging than with training. Training is fairly predictive for a long run over a period of time where inference is very spotty. And it's, what do you call it, peaky or bursty? Yeah, bursty.
18:36Jon Krohn:And that's a tough one. When you talk about storage arrays, is that something like when you're training a model, you are updating weights on a whole bunch of different GPUs and you need to store those weight updates somewhere so the information is getting sent back and forth between the GPUs and the storage array? Correct. So the nodes, it depends really on how you're running it. Every customer does this differently. Some fetch, they've pre-trained somewhere else and they're fetching it in real time, chunk by chunk. And when they process it, they just use the local storage for scratch space. And when they have a result, they push it back out of the network in real time.
19:16Some download all their pre-training and have it locally, and then all the nodes pummel the storage array and then update, make all their updates. And when they're done, they batch load it. We never know how they're going to use it. Some customers are more demanding than others, but storage arrays and storage performance, along with the CPUs that they're using to interact and load things in and out, the workloads are so random on all the profiles. It's kind of fun to watch, to see what customers are up to next, to see if we've built and designed a data center that can work for them. And if we need to make any adjustments on the next generation that we build or do a tweak or refit to an existing site to support those things.
19:57Jon Krohn:Something that I didn't appreciate before this episode, I guess I had kind of heard conversations around the office that should have tipped me off to this being important. but it sounds absolutely critical that you are listening to each client's needs. And actually, in a lot of cases, it sounds like developing bespoke infrastructure systems for them and their needs. That seems to be essential to being able to deliver the kinds of services the clients are looking for. And it sounds like before, something that you said earlier is that before we have these clients come on as prospective clients, you have to be designing the whole system with extra leeway, with extra headroom to be able to meet whatever client demands come up.
20:42Jon Krohn:Yes. Yes. And that's the challenge that the XAI and OpenAI, my peers over there kind of look at me and say, wow, good luck, man. Because it's a much harder thing. I mean, in a lot of ways, we're taking a swag at it based on what we know, based on trends coming forward, working with the industry, listening to the industry, talking with NVIDIA about future items, and then finding a middle ground that is the best balance between performance and economic cost. These things are not inexpensive to build. GPU data centers are very expensive because of their high density requirements. Liquid cooling systems are extremely expensive to install and operate.
21:28They just don't sit there and run on their own. They take a lot of active tuning to keep everything running at peak efficiency. And same thing with the GPU nodes themselves. GPU nodes are fickle. They're very specific and prescribed. And so if a customer tries to do something different, it may break an entire cluster by doing that. And then, of course, we have to have all the guardrails in place with our VPC networking to isolate the customers so that in a multi-tenant environment where they're not sharing GPU nodes, but they're sharing common north-south network fabrics. The east-west fabrics are large cluster fabrics that are segmented up to based on how many nodes that whether a customer has eight nodes or they have 512 nodes.
22:19Isolating them and keeping them segmented for the highest level of security, highest level of performance, and other key measurements and KPIs that we measure against. Those design constraints have to be done up front, and they have to be discussed openly up front. And then there's knobs you turn for the performance of how much you really want to spend because the sky's the limit on how much you can spend here.
22:42Jon Krohn:For our listeners who aren't aware, Frank, what is a node? A node or a GPU node is effectively the actual box itself, the server that is a GPU server. It's not unlike a traditional, you know, x86 or ARM or other compute server, except for it's built and added on to it's much an air-cooled scenario. It's upwards of 12U tall, so 12 rack units, which are 1.75 inches each. So it's a very tall box. You can put four in a full-size cabinet. And it moves a lot of air through. It has eight GPUs in a typical configuration. And besides that, it's a normal server. But it has eight GPUs in there, plus it has eight network cards, one for each GPU, plus It has two other network cards in it for what we call the north-south network, which is your internet and storage layer access networks.
23:46The same server that would be 12U with the same configuration comes in a 3U as well, but now it's liquid. So instead of four per cabinet, I can put eight per cabinet easily.
23:59Jon Krohn:And so my density is increased. Because you need less space for air to move around because you're using the liquid cooling. Correct. When you talk about liquid cooling, I've built tiny little GPU boxes for myself. And they have been liquid cooled. And something that might surprise people, you hear a lot about, and we're going to talk about energy later on, but when you hear liquid cooling, I think people think that you're hooked up like a water hose to the municipal water system, and you have this cool water running through the surface. But I've had liquid-cooled GPUs in tiny little servers that I built myself, and the water just stays in the system.
24:43Jon Krohn:It's a fixed amount of liquid. I guess it's not even really, I don't even know if it's water. But the liquid is fixed in there, and it just loops around. It gets heated by the chips, and then it gets cooled down in another part. You're not needing to add more liquid on a continuous basis. That's correct. And this is one of those large pieces of misinformation based on designs from 20 years ago in data centers that aren't really used anymore. So your home system, you've got that little water reserve and little radiator, and you're taking heat away from the chip through that interface. And when we say liquid cool, this is not immersion.
Read the full transcript
25:21Like the chips aren't dunked into a vat of water. it's not like uh bitcoin mining where you submerge the entire system in a in a water bath or a dielectric fluid bath is it non-conducting uh binary fluid but this is instead of a heat sink on your chip you've got the heat sink has a water loop interface to it so cool water's coming in or what they call pm25 fluid which is very special fluid that isn't water it's it's like a very specialty glycol fluid that you would put in your, like something you put in a race car versus regular car, very lightweight, very highly performant. And it flows on what's called an SFN or a secondary fluid network.
26:07So the secondary fluid network is the same as the one that's in the case at your house. This circulates this fluid from the chips itself, you know, comes into the cabinet with, you know, these inch and a half to two inch hoses runs down through manifolds in the cabinet. Manifolds have connections or points. The servers plug into the manifolds and now they're on the secondary fluid network. The secondary fluid network goes to a heat exchange point. So that's not where it stops. This is the beginning. So instead of saying, hey, it's going to vent to your, to your den or your room or your office at home, it actually has another radiator, a heat exchanger that exchanges with the primary water system of the building, the hydronic system for the building, and that exchanges the heat at that point.
26:57So the secondary fluid network is super clean. It's sub 2.5 micron filtered water, so it never plugs up the little capillaries on the chips or on the cooling plates and things like that inside the box. There's a number of different technologies in the boxes, depending on who makes them. And then it goes to the building primary loop. Okay, now you have cold water coming in at as low as 7 degrees C to 20 degrees C, so 45 to 68 degree water coming in from the building. That then goes outside the building and goes into giant heat rejection systems. So this is where the misnomer about water happens.
27:40Back in the day, the cooling towers used water they used water sources to basically spray over these radiators and then blow air across them and you'd get this great Bernoulli cooling effect across the surface with the water and or other methods of using direct air cooling could be a waterfall the water coming through and the air is cooling the water those techniques are pretty much banned everywhere now. So what it has outside, because of water waste and water use, everyone's been very concerned. But the industry started going about 10 years ago to what's called heat rejection. What it is, is a giant radiator like that's in your car, but it's huge.
28:25Or there's 30 of them lined out outside with fans. And the fluid in the building, the glycol-based fluid in the building, just like in your car never gets changed once you fill it up once you fill it up it's full so you may need to if you have to fix a broken pipe or change something out or do maintenance then you may have to top it off but you're not using millions of gallons per month you're using no more than a household does in a month and in a giant data center so you reject basically it's called heat rejection you're taking the heat and you're exchanging the heat into a from the secondary to the primary.
29:04And then in the primary, you go through a giant radiator like in your car, but there's bigger ones and they blow air across it and you vent all that heat. And you blow that heat upward. So it blows up and away from the building and blends with the air above the building. And so sometimes even when you drive by data centers, you'll see like on a cold day, if you're in like Chicago or something, you'll be like, wow, what's that plume coming off a building? Same thing off a data center. It's the differential of heat and it condenses and it makes a pretty cloud. But when you do see that, you know you're not using water.
29:34You're doing heat rejection. And that's one of the biggest myths today on why people are turning towards not wanting data centers, right? The nimbyism, the not in my backyard, hey, you're taking all our water. And we'll talk about power, I'm sure, in a minute. But the water part's not true anymore. And I think that's just old information that is no longer being done. But it keeps coming up in the news.
29:59Jon Krohn:Yeah, we'll get to the power situation shortly, but almost like 95 % of what you said there was new information to me. I did not know any of that. And it is really interesting. So I hope a lot of listeners have enjoyed that as well. And maybe now you can also debunk the concerns of people who saw TikTok about all the water usage that the new data center is going to use in their area. So yeah, that is interesting. Another, when you were talking about nodes, I don't know, five minutes or so ago and I asked you about nodes. No, no, not at all. Nothing to apologize for. But one of the things that you were talking about there, and you'd mentioned it earlier in the episode as well, also something that I know nothing about.
30:38Jon Krohn:You were talking about East-West connections and North-South connections. What are those? And they also, they sound different. That is a good question. Networks within the data centers are broken up into two, well, there's actually four networks within the data center itself. So the first one being the one that a lot of us focus on is called East-West. So if you have all the GPU boxes lined up in a row, they need to talk to each other. They travel East and West to talk to their adjacent GPU nodes. And this network is very fast. Every single GPU has a dedicated physical interface to it, but then logically can connect to all the other GPUs.
31:19Jon Krohn:And this is done. All the other GPUs on the East-West connection. Correct. So if you have 255 nodes, say a 2K cluster, 2 ,000 chips in a cluster, then all of them can talk to each other. And they can talk to each other at full line rate. So whether it's a 400 gig interface or an 800 gig interface in the newer deployments and soon to be 1.6 terabits per GPU, leaving the box, you have eight of those connections in an eight configured server. or if it's a water-cooled 4 GPU server, you have that many connections connecting to all the rest. That way, the GPUs can work together, share memory, share bandwidth, and you can cluster the GPUs to load larger models and do other things with them and distribute the work across multiple GPUs together.
32:16That networking east-west is very powerful. That makes all the difference. And the different manufacturers have different versions of that. But, you know, NVIDIA is by far the furthest out, the most performant. The East-West traffic is outstanding. And there's two ways to do East-West. You can do it with Infiniband or IB. That is what NVIDIA purchased from Mellanox years ago. They're the only manufacturer of this. Others had license to it, but they stopped making it. But NVIDIA has built on it. And I think all those other manufacturers are regretting their decision to stop making it. And then there's ROC-E.
32:58That's RDMA over Ethernet and converged Ethernet, R-O-C-E. And that is using traditional Ethernet framing and then traditional Ethernet technology to pass GPU traffic across it. This works very well. It scales incredibly well. But we're going to leave that religious argument alone for this discussion because people are seated in one house or the other, and it's literally one of those religious arguments within technology that people are passionate about.
33:29Jon Krohn:Specifically, InfiniBand versus Rocky. Yes, that is correct. So that's a big one. Then you have the North-South network. The north-south network is it goes up and out and back down. So from the GPUs up to the different layers out through firewalls and border networks and to the internet, that's your north-south network. On the inside, it'll pass through like a services layer. So it can go to get to the CPUs and it can get to the storage networks. And those are tied in as well. So in the north-south network, there are complexities about storage on the east-west network, but we're not going to go there.
34:08Okay, yeah. But traditionally, north-south is for that. And then you have another network on top of that, which is your management network. And then you have a true out-of-band network, which is the break glass network and how you manage these things from the outside in. So if there's a problem on the inside, you can work on it and run all your observability and see the node health and the network health and the traffic and the flows and kind of what the customers are getting at based on from looking at it from the outside in. If they're having issues getting to some endpoint, you can look at all of that through your out-of-band network.
34:44So it's truly out-of-band. And then you can kind of break glass and barge in if you need to from the outside to ensure everything stays running. Now, a lot of people don't build that network, totally segmented. They just have the in-band management. And then if something breaks, you send someone with a laptop to the data center, which doesn't work. So the larger you build, the more kind of guardrails you need. And that true out-of-band management plane is one of those things, an expense that a lot of people don't necessarily buy in on, but it's critical for lights out operations of these site locations.
35:19Jon Krohn:Right, yeah, lights out operation, meaning this is this was a new term that i learned just in the past few days as i was doing research for this episode lights out means that you don't need to have any lights on in the center because there are no humans in there it's a fully autonomous system that's the idea at least that's the notion um i've built lights out systems for years um with modular data centers and other things thing about gpus are they get run so hard that they do break and there's a certain percentage of failure. Some models have more failures than others. And the newer chips are just so much better.
35:56The liquid to chip cooling has really pushed that number down. You don't see a lot of failures within the LTC chipsets, which are great because you can adjust your staffing models in the data centers. We're staffed 24-7 around the clock at our data centers with our own teams to ensure that we can provide that Four Seasons white glove experience for our customers, whether they're on Bare Metal or MKS or anything else, that their nodes are always available and always running. What does MKS mean? Managed Kubernetes. So more than just, hey, we run Kubernetes. No, we run a full managed Kubernetes suite.
36:36So a full platform as a service, control plane management provisioning, slurm on top if you want it, other things. But it's more than just, hey, here you go, kick it over the fence and you can run Kubernetes on your own. Good luck. No, we have a full managed Kubernetes suite and a very robust platform as a service.
36:58Jon Krohn:So basically that means that a client of Lightning can come with their Kubernetes, whatever they want to be scaling up with Kubernetes, and they can just have that configured already, packaged up, and then they can bring it to this MKS managed Kubernetes service and have it scale up easily. Yeah, they can effectively plug into our system directly and it just makes it easier for them. We have customers that range in skill sets all the way from the 10 scale, the smartest, brightest, craziest idea. AI native companies are just doing the things you read about. all the way down to kind of the enterprises which aren't doing those foundational and frontier things that want to be users of the system.
37:52And they want an easy button. They want assurances. They want more security. They want more of the enterprise features. And we maintain all of those things as well for them. And we have an easy button through our provisioning and management systems to do that. They don't know anything about bare metal. They don't know anything about north, south, east, west storage networks, internet access. They don't. And they don't care. And they shouldn't have to. They want to consume it like they would AWS or GCP. They want to click and deploy. And we have those options as well. Then we want some folks who, hey, I want to change the bias settings on the boxes for this esoteric performance thing that I'm trying to obtain.
38:32Can you test that with me? Sure. Of course we can. It's a different customer. with a different understanding. And the amount of customers that know those things in the hardware space are becoming less and less, and they're relying on us more and more to make those things happen.
38:48Jon Krohn:That's really cool. It's nice to know that Lightning does that. Another question that came to me as we've been talking about this, that is pro, I think I know what the answer is, but it'd be nice to hear you say it. When you talk about all this East, West, North, South, obviously that is completely independent of like magnetic North, right? It's not like you need to set up your data center so that all of the rows are like east-west. No, it's Feng Shui. No, just kidding.
39:17No, you could line the cabinets up at a 45-degree angle in the room if that was your groove. And that's actually a joke I use with the providers. I said, if the pipes are running at a 45-degree angle under the floor, I'm not going to make you move them. I'm going to line my cabinets up to them so we can value engineer the deployments as much as possible. But no, it's a logical assignment of east, west, north, south. That's a good question.
39:46Jon Krohn:Yeah, so at some point, people just decided to use that convention because obviously it could have ended up being the other way, presumably. The nomenclature, like everything that you're talking about is east, west. That could have been the north, south. I don't know if that makes sense. Well, east, west, meaning you're looking left to right. So you're looking at things that are, if you had a row of people or a row of GPUs, you're looking left, you're looking right, you're looking east and west versus north and south. You're looking up, looking down. I see. Yeah, that makes a lot of sense. And then so digging into this a little bit more, when people have seen pictures online of any kinds of data centers, like we have been able to see online for decades, probably most of us haven't actually been to a data center.
40:27Jon Krohn:So when you're, so when you see those photos, you see these long kind of like hallways that the servers make up. And so when you're talking about East West, it's like you're, it's like you're standing looking at one of those server racks and you're examining it from the left to the right. That's kind of like looking West to East, regardless of where your compass would actually be pointing. Yeah, a hundred percent. And then the East West is also limited into what we call serviceable units. so there's a logical distance on the cables that you have to maintain um they don't go very far they go between 50 and you know 500 meters and that's it so they can't go down the street they can't go across the to another data hall so data halls have to be contained and the serviceable units are there now how you connect all those serviceable units together that is some magic but there's there's limits and scaling factors that each manufacturer has for the design and constraints of their east-west connectivity within the actual structure i see so basically based on that kind of 500 meter cable limit if you build a big enough data center you would you could potentially have a bunch of different uh what was the what was the term you used there for kind of like like a serviceable units well yeah so yeah there was another term though where kind of if I could potentially have, you know, that east-west, I'd reach my east-west maximum for one data hall.
41:54Jon Krohn:That was the word that I was looking for, data hall. And then so you could have then another data hall to my east and another data hall to my west. You could have as many as you wanted, basically. And then those would be connected not by east-west connections, but by the north-south connections if they need to be connected. No, actually, it's more complicated than that. So within the east-west fabric that connects it together, there are distance limits. So the limits are actually between the GPU nodes for timing in the picosecond level. So from GPU in one to GPU, say, 255, they can't be more than 50, 100, 500 meters, depending on the type of equipment you use apart.
42:35So if your data center is the size of a Costco or a large factory store like that, if you had one at each end of the other, it would be too far. So they have to be grouped together physically. And the data halls are now designed to group for those groupings. And when you design, lay them out. Now you can have one data hall connected to another by going up a tier within East West. So you have your leafs that connect your local connections, just like traditional networking. Then you have your spines that connect all of those leafs together. And then you can have a third tier or super spines, which could allow you to go between data halls.
43:17Now, when you're doing training, that kind of literally adding X picoseconds to go to the next room, it might break your training. The timing may be off, it may mess you up. So while they're all connected so you could share data, certain types of workloads just will function badly unless you tune them very specifically for that data center and that topology and that design. So working with the clients to let them understand what the underlying design is and sharing it with them openly is super important. Otherwise, they'll make an assumption that they're using it another provider and they're like, hey, why doesn't this work?
43:55Or why does this work so much better? It's like, oh, well, we'll share with you and show you why.
44:00Jon Krohn:Wow, so much to think about there. And even in my simple-minded view of how these data halls could connect. I really appreciate you elucidating for me how these spines work and super spines. That is great to know. So I think probably for now, I'm going to take a pause on talking about these physical centers. Unless you think there's something interesting that our audience needs to hear about just what it's like to be in one of these AI data centers. A lot of us can probably picture a photo that we've seen online of these long hallways of server racks. But is there anything else that's kind of interesting, maybe particularly interesting about an AI data center when you're physically standing there?
44:45When you're physically in there, one of the differentiators from a traditional data center is the noise level. These systems are very noisy. We call them screaming banshees.
44:57Jon Krohn:Oh my God, I had no idea. Yeah, so especially within air-cooled, And inside the data hall, the levels range from way beyond what you'd hear at a rock concert if you're in the front row. And so hearing protection for our staff, we require two types of hearing protection at all times. You have both like the buds that go in your ear. You can use molded ones or not, you know, listen to your music or whatever. So something occlusional in your inner ear and then cans, right? and cans being all passive. You cannot wear or use active noise canceling systems within a data hall. This is a big thing that people are like, yeah, I put my noise canceling on, it's great.
45:42Well, if the data hall is 105 to 110 decibels, to cancel the noise, noise canceling generates 105 to 110 decibels. So that does just as much damage. You're not hearing it, but it's damaging your drum.
45:59Jon Krohn:Wow, I had no idea about that. It makes so much sense now you say it, but I had no idea that with my noise-canceling headphones, I'm here thinking, I'm wearing noise-canceling headphones right now. Obviously, it's not canceling 100 decibels of noise. I don't have screaming banshees in my recording studio, believe it or not. But that's actually an interesting take-home tip. For anybody going to a rock concert or whatever, you need to have passive noise, not canceling, but just suppression. Yeah, occlusional or suppressive. And so really it blocks the two different kinds of hearing protection, the inner ear hearing protection, and then the outer ear blocks different frequencies of noise as well.
46:39So the frequencies of noise that cause your hearing damage, the higher frequency noise from the fans and the motors and the power supplies that are humming that you can't really hear to your native ear, they're present. And so you need to block all those out for safety. We take that very seriously in all of our locations. And we actually have OSHA sound studies done and we maintain OSHA compliance and everyone has to get trained to be in the data center. even visitors we warn visitors when they're coming like hey this is a very loud environment and and uh what's funny is that the occlusional blocks out so many things but if i talk loud enough like like that loud grandma talking at you because she can't hear anymore uh to someone in the data center with hearing protection on the frequency of my voice comes clearly through but you don't hear any of the high frequency noises or things that can damage your hearing so you might think that how does anyone work with somebody else well there's a couple ways one we have some systems that are like intercom based like racing radio style that you can talk to each other with or and we also have like you know kind of hand signals that we use in the data center and things like that there's there's a bunch of different combinations depending on which location we're at how loud it is and you think oh the liquid cooled ones don't have as many fans they're just as loud they have they have rear door heat exchangers they have other cooling systems and pumps and things running in the room.
48:03It is a very loud industrial environment. The liquid to chip data centers are more industrial, if that makes sense. The traditional data centers that are pretty and they're really nice, they raise floors and they're super orderly and things like that. You look at some of the pictures you see online, I'll look at your chip centers, and there's hoses and pipes and cables and everywhere. All of ours currently are fed from above. So there's 20 inch water mains running through the room, you know, that are insulated. So they don't make water or sweat because the temperature differential at the room, there's hoses to every machine.
48:42There's, it's just, if you look up, you're like, oh my gosh, what is all this stuff in here? Well, that's the, how the sausage is made. It's very industrial. It's like you're in the reactor room on a submarine or something.
48:55Jon Krohn:It's pretty cool. Do you think that there is a higher rate of sign language fluency among data center workers relative to the general population?
49:07That would be an interesting question. I don't know. I don't know. But you'd think there should be at some point. That's actually a really good idea, even though I'm sure they use their own version of sign language, if you know what I mean.
49:20Jon Krohn:Right, right, right. Well, yeah, that's interesting. I guess it's potentially a good career choice for people with a hearing problem. Oh, yeah. Yeah. There you go. And so for people, Frank used the word OSHA, which people in the U.S. probably will, everyone will know what that means. But if you're outside the U.S., it means occupational safety and health administration as a federal body that, yeah, is keeping workers safe in all kinds of industries. Quick question for you. With these data centers being so large, do you just always get around on foot? or do some people use pedals or motorized vehicles ever?
49:57Depends on how good your insurance is. A lot of data centers, some of these now have scooters and bicycles just sitting all over the place. I'd say don't wear Heelys because you need to wear actually proper work boots in these locations because things are heavy and if something were to drop on your foot, that would be bad. so you need to wear a proper work attire. Do you wear hard hats? During construction phase, we wear personal protective equipment. So hard hats and vests and ceramic toed boots and non-flammable things during constructability and provisioning. And then once it's online, the data center technicians aren't required to wear that, but for hearing protection and, or if they're working in a cabinet, safety glasses.
50:49And then, of course, they need proper ESD projection. Like we issue ESD shirts for our teams. So, you know, they're wearing a shirt that's not polyester that won't spark every time they touch a cabinet. Kind of standard issued uniform stuff that we've been working through.
51:07Jon Krohn:So ESD is like electrosensitivity something? Yes. Electrostatic discharge. Electrostatic discharge. Yeah. And you have to wear a wristband if you ever touch or open a box. So you put it on and then the cabinets literally have like these little light and bolt and plug plugs on either side, every single cabinet front and rear. And you plug yourself into those. So you're now connected to the cabinet because with all the air flowing through the systems, they can generate static electricity. So it's for safety of the worker and for safety of the gear. Wow, that is so cool. Thank you. I didn't anticipate talking about these kinds of physical things like sound and scooters and stuff, but it just kind of occurred to me as we were talking about this more and more.
51:51Jon Krohn:Let's go back. Before we wrap up this episode, there is something that you mentioned earlier that you said we would discuss it later on. So let's make sure we get to that, which is the electricity issue. You mentioned NIMBYs earlier. You hear that there's a lot of NIMBYism all over the world. You and I are both in the US, so we hear it, you know, particularly in the US press, around concerns of electricity usage in a given region where lots of data centers are coming online. Now, I read The Economist every week, and The Economist has a number of times in the past year done articles on how, at least up to this point, if you are in a region that has an increased electricity bill, it is almost certainly not due to data centers or AI-specific data centers being built there.
52:41Jon Krohn:Yeah, I'd love to hear your thoughts on this issue. People probably, yeah, I don't know if you're a cocktail partisan, are like, Frank, I can't believe what you're doing. Yeah. Well, yeah, I get that a lot. Like, oh, you build those things? Like, what do you mean by that? And everyone thinks that they're one of those people, those people, you AI people. I've been building data centers a long time. And this isn't the first time this has come up. everyone thinks that this is why oh my power bill went up well it has nothing to do with anything it has to do with so many different public utility commissions decisions and taxations and other things that drive the cost of electricity to a residence i live in california the power is very expensive here it's the high it's like pretty much the highest in the state but in the country well yeah sorry la is the highest in the state and then california is the highest state no actually Actually, it's not.
53:37I was going to say like this is kind of the highest in the country. But we have multiple power companies here. So within the same state, generating off the same generation, meaning the same natural gas, fire plants. We don't have any coal out here. We do have nuclear plants still. We have one remaining. But the power varies from$0.10 per kilowatt hour all the way to$0.58 a kilowatt hour for residential delivery. Why is that? Well, it's because the state has messed it all up, right? This is state regulation, state changes. Here in Los Angeles is the second cheapest power in the state. Silicon Valley power in Santa Clara, California is the cheapest power in the state.
54:19and they have their own gas-fired plants within the city, just like LADWP, LA Department of Water and Power does here in Los Angeles. Then you have two other power companies. You have Southern California Edison, our wildfire specialists. And then we have Pacific Gas and Electric. Pacific Gas and Electric and Edison, they're some of the most expensive power in the entire country, whereas LADWP is the same as it is in Chicago, right, in the Midwest. and Silicon Valley power is as cheap as it is in Texas, which is super cost-effective, or in Florida where there's plenty of inexpensive power. There's plenty of power out there.
55:00It's all about regulation. So when a data center goes in, a couple of things happen. And let's say you don't have a data center and they come in and your power rates won't go up because the data center itself is going to pay millions to tens of millions of dollars to the utility to improve the power grid to connect to that site. They're going to pay for upgrades. They're going to pay for new transmission lines. They're going to pay for upgrades to the generation near you. They're going to pay for more generation. Or they're going to put gas-fired, what they call behind-the-meter power stations or fuel cells, which are silenced at the site.
55:37And oh, by the way, those fuel cells, the byproducts, they make water. So they make distilled water coming out of the back of the fuel cells, whether they're natural gas-based or hydrogen-based fuel cells, you're getting water output as that. They capture that water and they use it to keep the building full and topped off. Or they use it to put it through a filtration system and they water the grass out in front of the data center. They're making water at these locations with new fuel cell technology. And that's the way forward. These gas-fired plants are called peaker plants, which you see going in that, like, hey, Goliath is using these peaker plants.
56:11Well, they're burning what's called dirty gas. So there's stuff that's coming right from like the oil fields and the gas fields, and they're burning that. But then they burn it and they filter it and they keep it clean versus like, if anyone's seen like an offshore oil rig or a gas field, you see like these big burning torches in the middle of the night. They're burning all that gas, all that excess gas. Now they're capturing that gas, pipelining it over to the data centers, and they're using it to generate power for the data centers. What's burned and ended up being CO2 that was wasted in the atmosphere, now can at least be used for something.
56:47And so there's a lot of information about how those plants work. Sure, those plants need some water. But if you use fuel cells, you make water. And you notice like Project Jupiter down in New Mexico, fuel cells. A couple of companies out there that are building, fuel cells. All of them are using fuel cells now. Fuel cells are great technology, just like heat rejection, fuel cells of the future are data centers. So they're not only upgrading the grid, but they're adding additional capacity. And then here's something people don't know. Data centers have interconnection agreements with the grids or your utilities.
57:24When your utility has a problem or a big storm comes or you have a heat wave event or whatever it is, data centers are required to start their onboard generation, whether it's fuel cells or diesel generators or whatever it is, interconnect and push power back to the grid to support your local grid in time of crisis. And so the more data centers you have, they're like mini power stations all over the country. They can provide stability of the grid when the grid operator can't handle that hundred year ice storm that comes through town. You know, so it's a symbiotic relationship that is just not well understood.
58:01And I have to admit to all the listeners, this is a complex topic, right? And I covered it at a very high level. But at the same time, data centers aren't bad for the grid. Data centers don't cause your power bill to go up. They just don't. That's a misnomer. Just like they use water. No new data center uses water.
58:20Jon Krohn:That was obviously a well-practiced set of information. And although it was high level, you did have a lot of detail there. And I learned a lot of things from what you were saying. I think there will still be some people out there who will say, well, it's bad to be putting gas-fired plants online. We should be using all sustainable power. And I think that that's an ideal. And I think a lot of data centers do get built with commitments to be using nuclear power, maybe worst case, but also to be using as much solar and wind as possible. So that does happen. You see lots of data center contracts being agreed to today where it is entirely regenerated, renewable power sources.
58:59Jon Krohn:And, you know, that's a great way to be moving. I think, you know, there may be some data center scaling happening globally where it's happening so quickly that a gas fire plant is needed in the short term to power that. an argument that I make that kind of, well, at least allows me to sleep well at night as someone who is so deep in the AI world is that AI is helping us increasingly push the frontier of what humans are doing, including things like making nuclear fusion commercially viable, which is hopefully the energy of the future where, you know, we can be splitting water to be, and there's small amounts of water.
59:41Jon Krohn:It's not like we're getting around out of water because of the nuclear fusion plants, but creating huge amounts of energy, like having suns on our planet. And then that allows us to pump carbon dioxide back into the earth. And so anyway, it is a very complex topic. I understand the concerns about energy use in general. And I'm not asking you to necessarily have any comment on what I just said, Frank, but. Yeah, no, data center industry is pioneering this tech. Data centers are why we're looking at small reactors now. It's why reactors are back on the scene. Nuclear power is not evil. Nuclear power is great.
1:00:15It powers our entire Navy and has without incident for many years. You know, I remember when the tsunami hit Indonesia. Well, the USS Ronald Reagan, a nuclear aircraft carrier, pulled up and connected itself to the shore and provided power for the country for like a month. Yeah, like it was in the harbor. Oh, yeah. They provided shore power for emergency services and stuff like that. Nuclear power is not the enemy. It's more uncertainty and unknown, fear of the unknown. But data centers and data center technology companies are, you might have heard, like, we're increasing the voltage. We're going away from AC power because we can make the entire place 5 % more efficient by going back to DC power and going higher voltages, 800 volts DC.
1:01:01and then you're talking about coming out of the fuel cell is DC power. Then you put a grid-tied battery energy system in there, like a giant Tesla grid battery or some other brand, and those are battery-tied. Now you have no diesel generation. You don't need generators. You don't need UPS systems with big batteries in them because you already have the battery bank outside. And then it's straight DC power all the way in. and you have efficiencies of 5 % to 10 % gain on the power lossiness within the building. So your overall efficiencies go up. And the data center business and the supporting industry is pouring, you know, billions of dollars into this technology for cleaner running data centers and always looking for better.
1:01:46I wish they'd spend this much money on cars to make gas-fired cars this efficient. And the battery systems are being driven by data centers, not just the car manufacturers. The data centers are looking for cleaner power, longer-term power, salt, water-based electrolytes and non-chemical electrolytes. There's no fire hazards. And salt batteries, salt chemistry batteries. These things are all being driven by the data center business. They really are. Because they're grid-level storage and for data centers. So you can use renewables, fill those up, run on batteries overnight, and then let the sun come back up and power the data center.
1:02:26Jon Krohn:That's all happening. You're talking about those kind of 5 % gains in efficiency. That reminds me how before we even had the Gen AI era, I believe the first commercial value provided by the DeepMind acquisition that Google made now a decade ago was creating efficiencies in the way that power was being routed within Google data centers, allowing them to get those kinds of single digit or low double digit efficiency improvements. And so, yeah, AI has for a while been helping reduce consumption, even as we build more of these AI data centers. So, yeah, complex topic. And I'm certainly not an expert, but I'm hopeful for sure in the long term with what we're doing with AI.
1:03:10Jon Krohn:And I strongly believe, as any regular listener will know, that I'm techno-optimist. And I think things have never been better. Things are going to continue to get better. Hopefully, that's kind of a nice note to start to end the podcast episode on. Frank, you and I were discussing before we started recording about what book you might choose as your book recommendation for our listeners. And if you are going to go with the one that you recommended, it's kind of perfect because we were just talking about alternating current, direct current. And yeah, is that your pick? Are you going to go with the same one?
1:03:41No, that's my pick. It's a good book. It's been out for a while, but it's a good listen. It's based in truth with a little twist because it's about Edison, Tesla, and Westinghouse in the early days. And it's called The Last Days of Night. So talking about electrical power and electricity to homes and buildings and street lighting and things like that. And it's written from the point of the view of an attorney who's involved with all these things. And so it's very interesting that it's a total outsider watching Westinghouse, Edison, and Tesla argue and discuss these things and things that happened.
1:04:22And very interesting book, The Last Days of Night. Cool.
1:04:27Jon Krohn:Yeah. And a good reminder, actually, now that you mentioned that kind of the last days of night, kind of also ties into my techno optimism and how there's probably not many listeners that would like to go pre-electricity. And in the future, I think people, you know, couldn't imagine going back to a pre-AI society where we have unmetered intelligence, just making everything easier than ever before for us. And hopefully, yeah, allowing lots of positive human outcomes as well. Frank, I actually didn't warn you about this, but my final question that I ask every guest is just how people can follow your thoughts after this episode.
1:05:04Jon Krohn:I don't know if that's, should people be following you on LinkedIn or what? If you post anywhere publicly? No, not really. Occasionally on LinkedIn, I'll share things that I think are salient and what I get exposed to day in and day out, like some of our customers who are working towards cures for human illness and things using AI tech and those things that are not like, oh, I can play this game better or I can do this better. No, real things that move the needle, things that move society and the world forward. That's why we have this AI tech. Those are the things that really drive us. And so I'll post and discuss those.
1:05:41Very interesting. But yeah, I am, of course, always on LinkedIn. Besides that, I'm kind of a recluse online, a typical security background guy that doesn't post a lot online.
1:05:55Jon Krohn:So almost an exclusive here into Frank's brain, unless you are also, if you're a lightning AI employee and you have access to the Slack, Frank, you've recently been a very heavy user of the random thread on Slack and posting things in there. I think you've been the number one boaster. Well, there's lots of fun things to share and random they are. Maybe you're just doing that because you can't shout it loud enough in the data center. So It's in the Slack. Exactly. Frank, it's been really interesting having you on the show. I have learned so much. I'm sure a lot of our listeners have as well. This was an important episode to help us understand how data centers are built that are allowing us to have all the AI capabilities that we're talking about on the show all the time.
1:06:40Jon Krohn:So thank you, Frank, for taking the time out of your super busy schedule. Really appreciate it. Thanks for having me, John. And thanks for everyone who actually tuned in to listen. Love that episode today. In it, Frank detailed how Lightning AI uses co-location to provision its GPUs, sizing each build around available grid power, applying a conservative 81 % D rating, and then contracting per megawatt over typically 10-year terms. He talked about how liquid cooling doesn't waste water at all. Modern data centers run on a sealed glycol loop and reject heat through giant radiators, so a huge facility uses no more power per month than a single household.
1:07:18Jon Krohn:He talked about how GPUs talk to each other over ultra-fast east-west networks limited to a few hundred meters, and how thanks to screaming banshees, AI data halls run at 105 to 110 decibels, louder than the front row of a rock concert. As always, you can get all the show notes, including the transcript for this episode, the video recording, any materials mentioned on the show, the URLs for Frank's social media profiles, as well as my own, at superdatascience.com slash 1003. Thanks to everyone on the Super Data Science Podcast team, our podcast manager Sonja Brejovich, media editor Mario Pombo, partnerships manager Natalie Zajski, researcher Serge Macis, and our founder Kirill Aromenko.
1:08:00Jon Krohn:Thanks to all of them for producing another outstanding episode for us today for enabling that great team to create this free podcast for you. We are deeply grateful to our sponsors. You can support this show by checking out our sponsors' links in the show notes. And if you ever want to sponsor the show yourself, you can see how to do that at johnkrone.com slash podcast. Otherwise, please help us out by sharing this episode with someone who would love to learn about AI data centers, review this episode on whatever podcasting platform you listen to podcasts on or on YouTube, subscribe if you're not already a subscriber.
1:08:36Jon Krohn:But most importantly, I hope you'll just keep on tuning in. I'm so grateful to have you listening. And I hope I can continue to make episodes you love for years and years to come. Till next time, keep on rocking it out there. And I'm looking forward to enjoying another round of the Super Data Science Podcast with you very soon.
From the publisher
Frank Basso, VP of Infrastructure at Lightning AI, joins Jon Krohn for a rare ground-level tour of the one layer of the AI stack the show had never covered in over a thousand episodes: the physical data center. Frank explains how Lightning AI provisions its 35,000-plus GPUs through hyperscale co-location, why everything new is liquid-to-chip cooled, how GPUs talk to each other over ultra-fast east-west networks, and what it’s actually like to stand inside a 110-decibel AI data hall. He also debunks the most persistent myths about data-center water and electricity use, and makes the case for fuel cells, nuclear power, and 800-volt DC distribution as the path forward.
Additional materials: https://www.superdatascience.com/1003
Interested in sponsoring a SuperDataScience Podcast episode? Email natalie@superdatascience.com for sponsorship information.
In this episode you will learn:
(02:47) What actually makes an AI data center different from a traditional one
(06:04) How Lightning AI provisions its 35,000+ GPUs through hyperscale co-location
(24:01) Why liquid cooling doesn’t waste water, debunking the biggest data-center myth
(29:46) East-west vs. north-south networks, explained
(43:47) “Screaming banshees”: why AI data halls run at 105–110 decibels
(51:52) Why data centers don’t actually drive up your power bill




