In short
CoreWeave’s AI infrastructure and how AI changes computing beyond “more GPUs,” focusing on cluster operations, Kubernetes, security/guardrails, agentic workloads, and physical AI.
Guest
Chen (Henn) Goldberg, product engineering leader at CoreWeave; previously helped scale Kubernetes at Google Cloud for 8+ years, taking it from open source to cloud-native foundation.
Key claims
AI-generated code and agent workloads require new testing, review, monitoring, and “self-healing” operations; modern AI clusters behave like a “giant computer” where network/storage/health/security all affect performance. CoreWeave’s Vera Rubin NVL72 multi-rack cluster connects hundreds of GPUs for training, inference, and agentic workloads.
Notable examples
Velvy (liquid cooling valve manager) and RECI/REC manager (signals to “mission control” dashboards); customers building in weeks vs a year; agent “AI loop” (deploy, observe/cure data, iterate, redeploy); 90%+ of CoreWeave AI workloads run on Kubernetes with managed/bare-metal changes; physical AI uses simulation/data to shorten cycles (e.g., vehicle calibration from months to 24 hours).
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOThe Impact of AI on Software Development
0:00 to 1:11
Explore how AI is transforming the software development landscape.
“They have thousands of, many thousands of engineers, and most of the code now is being generated by AI.”
The Evolution of Infrastructure in AI
2:31 to 3:40
Discuss the shift in how infrastructure supports AI workloads.
“I guess to start out, you were part of that cloud native transition and help bring Kubernetes into the mainstream.”
Emerging Bottlenecks in AI Infrastructure
3:40 to 4:58
Learn about the surprising bottlenecks arising in AI infrastructure.
“You know, you've mentioned before, I believe, that every time a hard problem gets solved, a new bottleneck manages to appear.”
Challenges in Building AI Clusters
4:58 to 8:08
Understand the complexities of creating and maintaining AI clusters.
“And those who lean in the most will see better results.”
Innovations with the Vera Rubin Cluster
8:48 to 11:03
Examine the advancements brought by the Vera Rubin cluster.
“You recently announced the big Vera Rubin cluster.”
The Future of AI and Its Challenges
11:03 to 14:00
Analyze the ongoing AI developments and their implications.
“So having that powerful supercomputer is actually, of course, creating more possibilities for our customers.”
The Speed of Innovation in AI Development
14:00 to 15:10
Learn about the rapid pace of AI development and its implications for engineers.
“I think that's one of the things we do very well at CallWeave is that we move fast.”
Transformations in Software Infrastructure
15:10 to 17:46
Explore how AI is changing software infrastructure and project management.
“Maybe as a hobby on the weekends or something, you know.”
Security Challenges with AI Agents
17:46 to 21:57
Understand the security concerns related to powerful AI agents and control measures.
“And of course, you think about a multimodal and how you scale up and down and how close the data is.”
Integrating ML Engineers into Development
21:57 to 24:45
Discover the evolving role of ML engineers within software development teams.
“It looks to me and feels to me more like this is like here's a new cybersecurity type of problem that is just going to be a thing we're always working on for the foreseeable future.”
Show all 16 chapters
Infrastructure Challenges and Solutions
24:45 to 26:15
Learn about the current challenges in AI infrastructure and strategies to overcome them.
“They're not just taking, you know, a frontier model or an open-weight model.”
Adapting to the AI Era in Software Development
26:15 to 28:00
Examine how software development practices are evolving to meet the demands of AI.
“In over time here, you'll have also expanded well beyond renting GPUs and doing things like you've acquired Weights and Biases, OpenPipe, Marimo.”
Evolving AI Workloads on Kubernetes
28:00 to 30:40
Learn how Kubernetes is adapting to the growing demands of AI workloads.
“So all of those tools continue to evolve with customers.”
Exploring Physical AI Applications
30:40 to 33:30
Discover the unique challenges and innovations in the physical AI landscape.
“I'd like to shift to physical AI a little bit for a minute.”
Future of Energy and AI Scaling
33:30 to 36:40
Understand the interplay between energy needs and AI infrastructure growth.
“Something that kind of shifting a little bit, I think I saw that at the end of Q2, CoreWeave had like 1.5 gigawatts, something like that, of active power and roughly 3.7, I think it was contracted.”
Optimism for the Future of AI
36:40 to 37:55
Delve into the hopeful advancements in AI that can transform various sectors.
“And I hope that, you know, the work that we are doing, the thing that excites me is making sure that these tools and these technologies will be accessible.”
Transcript
Automatic transcript. May contain errors.0:00They have thousands of, many thousands of engineers, and most of the code now is being generated by AI. So what does that mean for them? How do you do the testing? How do you do code review? There is a lot of change that is happening, so I don't think it's going to slow down in any way, but we were just going to see it maybe impacting more parts of the world. We keep joking about that identity crisis a little bit. Like if you like writing code, you no longer write code. It's new. Like you need to get used to this new tool and what does that mean? and it's interesting. The flip side, of course, just saw a demo yesterday from one of the folks on my team.
0:36They've built something in three weeks and they said that before that it would have taken them one year. The thing that excites me is making sure that these tools and these technologies will be accessible because if it's kept, no matter where you are in the world, you could be part of this transformation. And I want to make sure that that's what we are also building right now. I'm sure that, you know, like where I live, I have access to some set of problems, but someone maybe that lives somewhere else or have a different profession or is at a different age has access to different set of problems.
1:07And we should have the same tools and we should be able to collaborate. Welcome, humans, to the Neuron AI Explained. I'm Corey Knowles, and today Grant's not here. It's just me flying solo. But we're going to have an exciting conversation, and this feels like a really fitting time to have it. So CoreWeave recently announced it's brought a multi-rack NVIDIA Vera Rubin NVL72 cluster online, connecting hundreds of Rubin GPUs into one scale-out system designed for increasingly demanding training, inference, and agentic workloads. And that gets at the bigger question we want to explore. As AI gets more capable, what actually has to change underneath it?
1:45Our guest has a front-row seat to two enormous shifts in computing. Henn Goldberg helped scale Kubernetes from an open source project into one of the foundations of cloud-native computing during more than eight years at Google Cloud. Today, she leads product engineering at CoreWeave, building the infrastructure and software stack behind some of the world's most demanding AI workloads. Henn, welcome to The Neuron. Thank you, Corey. It's great to be here. It's great to have you. We're really excited about this. We've been watching what you do from afar for a very long time, and we're really familiar.
2:18And we don't talk about compute often enough on our channel. And we were just having that conversation when this opportunity came. And it was like, this is perfect. This is who we should be talking to. So it is the best. It is. It is. I guess to start out, you were part of that cloud native transition and help bring Kubernetes into the mainstream. Back then, you talked about making infrastructure kind of boring so developers could just build. Do you think that's still possible with AI or has the infrastructure become part of the product again? This is what I think is most exciting about the times we are in for folks like me that have been building infrastructure and cloud.
2:58We are no longer behind the curtain, maybe. Right. So making it boring, making sure that just things continue to work. Everything we do impacts the outcome, the systems and has a business impact as well. So that has definitely has changed. And the change are driven by from two different angles, I would say. One, the hardware is more complicated. There are more components. The workloads are bigger, but also how pervasive the technology is and how it is changing how people innovate and build new services. Right. So it's both a people change and a technology change and that that's great that opportunity to reimagine cloud.
3:40You know, you've mentioned before, I believe, that every time a hard problem gets solved, a new bottleneck manages to appear. What bottleneck has surprised you the most over the last year or so? I'm really curious. Or is it a new one every day? I was just thinking about it, you know, coming and, you know, when I prepare for a podcast like that, there's just so many things going on every day. There are news happening. I think that the pace surprised me. And the volume, the magnitude of things is definitely something that you've never seen before. I keep reminding my team what kind of conversations we had a year ago.
4:22They completely changed. What's possible completely changed. So I think that's probably one of the most surprising things that is happening. And maybe another thing that is both surprising but encouraging is that usually when you have this kind of transformation, there's a lot of resistance in the enterprise. And by the way, we are hearing voices that are worried about the change and what does that mean. But at least with the customers that I talk to, I think that what people recognize that this is inevitable. Okay, this is going to happen. So we need to lean in. And this is the opportunity. And those who lean in the most will see better results.
5:01And I'll bet it has changed so much. You know, a lot of people, I would say, still think of AI infrastructure as basically, how many GPUs can I get? And that's kind of what the average person is probably aware of in this space. But when you look at a modern AI cluster, what are the less obvious pieces that you think determine whether those fancy, expensive GPUs are actually doing what you hope they can? So maybe before going directly to the GPU, I'm going to go think first about the applications and what we are actually running when we are talking about AI workloads. So those models are bigger.
5:41They need a lot of storage in memory and a lot of processing in order to do both training and serving, what we call inference. And what it introduced from a system perspective, that many times it's not a single node or single computer that can run them. But you are creating that fabric of both compute and network of storage that all needs to work in harmony in order to achieve the performance, the latency and the accuracy that your application needs. So at a high level, this change between like cloud native where workloads were more portable between different systems. Now the workload requires a much more tightened cluster or loop.
6:25And that's what creates those kind of challenges. So the things that we care about is, of course, that all the GPUs should be healthy. And for that, we need to monitor them all the time. And then you, of course, want to make sure that there is the right network between the GPUs, between GPUs and CPUs. You need to make sure that the storage data set, everything is very much dependent on the data. And also recently, because we see more use cases, security is added into that mix. So how do you create an environment that meets all, check all the boxes from what you expect the production workload? And that's where it's basically a giant computer, right?
7:03It is a supercomputer. It is awesome. It is definitely a supercomputer. You can see me smile. I mean, it is a supercomputer. That's what's so exciting about it. And the way we approach it is that from a system perspective, you know, we start from the GPU. we make sure it's healthy, then we check multiple GPUs, and then RAC, and then multiple RAC, and then how do we run training on it, and how do we run inference, and it has to go for a long time. Because they are using all of those resources, it's enough to have one of those capabilities, one of those components not performing well. It doesn't even have to be broken.
7:37It can be just not at the expected thresholds. And the entire cluster will not perform as well. And you will see it and you will need to troubleshoot it. And that creates the challenge of how you build those systems and how you monitor and how you operate them. That's complex. I never really thought of it that way. And that's such a perfect description that, like, it's not just a GPU floating in the air. This is a whole system. And it's everything you think of as part of a normal computer just on a massive scale. So your agents can talk to each other, but they still can't think together. That's the gap OutShift by Cisco is closing.
8:22OutShift is Cisco's incubation engine, building the internet of cognition, shared intent, shared memory, and guardrails that allow agents to work together from anywhere. Open source protocols and code, a system for intelligence that scales. Read the paper and experience the demo at OutShift.com. That's O-U-T-S-H-I-F-T dot com. Check out the links in the description. Now back to our video. You recently announced the big Vera Rubin cluster. I think that's an exciting thing. What does that change? So from our perspective, you know, some things hasn't changed. Okay, of course, first of all, Vera Rubin are more powerful.
9:03The latest, of course, innovation with NVIDIA that are a great partner, and we are working closely with them on that type of innovation. And there are many things that we were able to bring from our previous experience in building large clusters. And that's the way we do the testing and our performance benchmarks and so on. So a lot of that best practices, we've brought them along. What I like about specifically about Vera Rubin is that it's in that production environment that this time around, we have expanded the innovation surface area. So there are more things that we know that we were able, you know, every time you do it once again, you have an opportunity to do something better.
9:45That's a continuous improvement. So we've announced together with that several innovations in the hardware space. We have a fun name for them, Velvy and Racky. Oh, I love it. So Velvy is a valve manager. It's the way that we are managing the liquid cooling. And the reason why we wanted to improve on that and build something of our own is that we wanted more signals and ability to respond better to what we are seeing. And we have what we call RECI, which is our REC manager. And that really ties all the signals into our mission control. Imagine like, think about like a dashboard that gets all the data from all the systems that allows us to know how the infrastructure is performing and how it is impacting the jobs that are running on it and allowing us to take actions.
10:32So those are just two examples of what it does. The second thing that is really exciting is that Vera Rubin is, of course, optimized for both training and inference. And the timing is excellent because what we are seeing from our customers, more and more use cases, more and more demand. I'm not going to go too much into that, but the demand for capacity continues to rise. Right. Like, I mean, it's that's what blocks innovation, allowing customers to do more experimentation. So having that powerful supercomputer is actually, of course, creating more possibilities for our customers. And I assume that brings the ability to to serve more faster AI in the same square footage over time, I suppose, as well, because because demand is definitely it seems pretty crazy.
11:25When I think about just my personal use, when I think to early last year, and mind you, I'm not an engineer. I was using, I don't know, 20 ,000 tokens a day. It wasn't that much a year and a half ago. I mean, I was still using it heavily, but it was more Q &A. It was back and forth. It was smaller tasks where now, like my codex setup, I'm running, you know, 200 or 300 million a day, which is insane as a growth rate. And I wonder at scale that that has to make such a more complex problem to solve. Like what I think is interesting about it, and this is why I think AI is not an incremental change.
12:07It's transformational because not only the infrastructure changes in the workload, like you said, like as a user, you're building things differently. You're using it differently. There are more use cases and more possibilities of things that you could not do before. and the work that we have, and actually if they haven't talked about it, call we. Why do we exist? We want to support that level of innovation. And early on, we were really working with the most sophisticated users, right? That they wanted to build those huge clusters, foundation model builders, you know, customers like OpenAI and Meta are among our customers and using our infrastructure.
12:47But now there are different type of customers and users. You know, just in our last month, we've talked about Caterpillar. We've been talking about our physical AI engineering. So the use cases has changed, which what it means is that the tools that we are building has to change as well. And that goes together with, of course, build an ecosystem of services. So if before it was really compute, storage and network, and now you can speak about our inference services and training services and how do you think about open weight models? There's a lot of conversation around that. And how do you manage cost?
13:27And that's maybe back to your question, like what are those hard problems? There are plenty of those that we are working on and really with the goal to empower innovation. I assume too that this isn't going to slow down. Like, I generally believe that, like at my level of nerdery, John Q. Public is still a little ways behind that, that there's this, the wave just seems like it's going to keep rolling for some time. Do you, are you all preparing for a year from now? I mean, how do you even predict if you don't mind me asking such a silly question? It's not a silly question. I think that's one of the things we do very well at CallWeave is that we move fast.
14:12We experiment, we iterate. It's okay for us to have some new ideas and throw them away and adjust as we work with customers. Is it going to slow down? Probably not. But I think the pace of where we're going to see innovation is going to change. And I was just talking today with a customer, a large financial service company, and they have thousands of, many thousands of engineers. And most of the code now is being generated by AI. So what does that mean for them? How do you do the testing? How do you do code review? How do you manage that cost? Is it productive? Is it great code quality? So there is a lot of change that is happening.
14:52so I don't think it's going to slow down in any way but we were just going to see it maybe impacting more parts of the world, I would say. I mean, you know, being an engineer we keep joking about that identity crisis a little bit. Like if you like writing code, you no longer write code. That's right. Maybe as a hobby on the weekends or something, you know. Yes, it's new. Like you need to get used to this new tool and what does that mean? And it's interesting. And the flip side, of course, you know, I just saw a demo yesterday from one of the folks on my team. They've built something in three weeks.
15:29And they said that before that, it would have taken them one year. Wow. It's wild how fast that can move. Do you feel like that level of innovation from the AI you serve has been helpful in your ability to design new software infrastructure that you need to scale? 100%. 100%. Wow. First of all, the ability to use all the data that we have. I think that's like probably something that everybody feels immediately in any space that you're in, right? Like it's everybody are using the same models, maybe the same GPUs, but our data is different. And that gives you like, you know, an asset, something that you can differentiate on it.
16:17And we are using it a lot. The things that we've learned and making sure that we're using it in a way that helps us do better and do more work. and also just you know from again you know we do of course a lot of software development that has significantly changed. So something I'm curious about and this is this is a little more technical I guess is I'm curious about how like latency has changed around agents versus traditional inference in a chat bot like we've gone from you know what might have been just one simple API call to now a single task that goes on for hours might have hundreds and hundreds of calls.
16:59How does that change the dynamics of what you build and serve on? First of all, from an infrastructure perspective, exactly like you say, right? We're moving from those one-time, like, stateless things to those long-running jobs that have to continue to perform. And I think what's very interesting, and people probably see it in their own usage, is that there can be a degradation in performance, like from accuracy of how the agent works. So building those self-healing systems and like, can you improve the model and continuously change that while the system is running? That's definitely one thing that we are working on.
17:32We're calling that the AI loop, meaning you want to make sure you start first with deploy your agent, right? You have a working system, but then you need to observe, cure the data, iterate, fix, and then deploy it again. So that kind of process is something that, again, it's both on the infrastructure perspective, but also from the tooling perspective and the systems that we provide our users. And of course, you think about a multimodal and how you scale up and down and how close the data is. We've actually made some announcement today around data processing. How do you make sure that you get the data close to the GPU, wherever it is, as well with new caching mechanism?
18:14So there is a lot of work on that world. I want to say maybe I want to also touch on a point that I'm sure many listeners also think about is that with those agents and how powerful they are. There are concerns around guardrails and security and what they can do and what kind of access they have. and that also creates a set new solutions around how can I control and make sure that the agents do the things I expect them to do and I have the right auditing for that and they cannot impact other systems around them in a non-undesired way. That's something that I'm thinking a lot about, for sure. And our customers do as well.
19:00Like having those controls are critical. Yeah, I think you're right. And I think that's really interesting in terms of like understanding where those kind of controls can be throughout a system. Like are they at the infra level, the model level, every level? Every level. Because actually that's what's the fun, you know, we've been talking about it. And there's a lot of, again, a lot of noise and news around security. But if you want to explain like what has changed, What has changed is those tools that are available to my team to build great things are also available to bad actors. Yes. And when you have malicious intent, you don't care about any damage you create along the way, right?
19:47Like when I make a change and when I build systems, I have to make sure that we're always making the system better and have those controls. But if you want to break things, then it's easier, it's faster, I would say. And I think that's something that is companies like ours, our job is to help bring that confidence. And you know what? Like, I actually want to spend more time on confidence. I think it's, we talk, when I talk with customers, and it's actually everywhere. Like, everybody knows that AI is happening, right? And I need to use it. Yeah. that's the one thing everyone knows everybody knows but it's actually really hard and we don't talk about it as much like am i doing the right things with how the tools i give my engineers and how the tools i give into my systems and what kind of controls and because again we talked about it this is not like an incremental change the entire system has changed yes and i would love to think that you know a company like ours and core is really leading the way that is our job is to help our customers to increase their confidence that they are doing things right in a safe way.
20:56And, you know, we've also last week, I think, yes, you see their news every week. And we talked about the physical AI engineering team. And this is exactly what we're doing, right? We are bringing the experts in that space and we're working together as one team with our customers making sure that this is not just a demo or a POC, we're working with them until it's in production. and hopefully giving that confidence, right? Like the goal is to make sure that this is getting the outcomes they expect. And we will only end that project with the customers when they are, of course, seeing the outcomes they need in production.
21:35And I think that's really important. So that's one angle. But I think confidence and really thinking about that and what kind of metrics we can provide to give that confidence is key. That's wonderful. I think there's a bit of a misconception around there, too, among people that ideas like AI safety and agent security are a thing we're going to just solve and it's going to be over. It looks to me and feels to me more like this is like here's a new cybersecurity type of problem that is just going to be a thing we're always working on for the foreseeable future. Do you think I'm right or am I misunderstanding that?
22:12I think that what we're seeing happening with infrastructure is also the same with security. If you're a security engineer, right, or a CISO team, like your number one job is to make sure that nobody knows you exist. You need to make sure everything is safe, right? Everything is safe all the time. And with this change, it's going to be very hard to do it behind the scenes. So I do think that people should be aware and think about what actions they take. In our team, we are saying no. Like every employee at Coe is a security engineer. That's a good attitude. Yes. We are all in it together. That's a really smart way to look at it because that is the case.
22:57The truth is it matters more than ever that everyone kind of get it. Like in many cases, you know, the best guardrail is that end user understanding what not to ask. What to ask, what not to give it, what not to unlock for it. And I think some level of base understanding of that for everyone could be really valuable. Yeah. And we are just starting with that cycle. Yeah. And it's going to be exciting. And I'm confident we're going to see a lot of innovation happening in this space. I have a question that's a little more elementary now, or at least I think it is. And that is when we talk about infrastructure around AI and we talk about reinforcement learning and evaluation and inference, are these all separate infra?
23:47Is this all happening in one system? Is there a difference in how you serve one versus how you serve another? That's not a basic question. I think it's a very important question to ask. Because what has happened over the last few years, before those ML engineers would be on the sidelines. ML is not a new discipline, right? We've been having those data scientist teams that have been working on hard problems and finding solutions. They just have been using different tools. And with the change right now, what we are seeing is a couple of things. One, you know, like every engineer now has to use more tools of AI and that's part of their work.
24:32But actually ML engineers are now part of that regular software development lifecycle. Some of them, like, again, you know, the customers I was talking with today, they are training their own models. They're not just taking, you know, a frontier model or an open-weight model. They're also building their own models. But then those models are being deployed as part of a system in that organization that developers are building applications on top of. And then, of course, you have reinforcement learning. So that loop already exists today. And one of the things that we are working at CoEvan is we are really making sure that it's easy to iterate an experience.
25:15So from a development perspective, that's definitely something that is changing. But I think what's also interesting is what we started with is that the infrastructure is expensive as well. And capacity is hard to get. So our job is to make sure that the capacity or the infrastructure is fungible. And you want to make sure that you utilize it in the best way possible and giving tools into that. And again, I think that's another change from before where we were not thinking about capacity as a constraint or what the data is and so on. That would be something that was easy to solve, but that's not the case right now.
25:53So the system that we are building, not only all of those capabilities are required to building applications, you better get them running on the same infrastructure. Yeah, because people want it and they demand it. And that's such an interesting thing. And I appreciate you taking the time to break that down. In over time here, you'll have also expanded well beyond renting GPUs and doing things like you've acquired Weights and Biases, OpenPipe, Marimo. Is that the right way to pronounce it? And you've kind of worked your way higher into that developer workflow along the way. And I'm wondering kind of what's the advantage of owning a little more of that stack instead of remaining primarily down at the infrastructure or up at the infrastructure level?
26:46The journey that we have taken was not from a hardware infrastructure software perspective. It is really walking alongside our users. Because even two years ago or three years ago, there was a lot of software involved always in building those large training clusters and how we've done observability. And we had the innovation that we were first to talk about around mission control and how we bring the data and how we actually do this orchestration that you just talked about. Like, how can I get the most out of my cluster? And as our customers expanded and we see more and more use cases, that's what actually drove new capabilities.
27:27I understand, yes, it does look like going up the stack, but actually what it means is it's about solving for different use cases and different personas. Like I've just said before, you know, like weights and biases, for example, the best tools for researchers for doing training. And for that persona, we've recently announced ARIA, which gives you that autonomous loop for researchers. So how can I use the data that I have to automatically suggest what would be the next run for me? So all of those tools continue to evolve with customers. And that was always, I would say, what we were working towards.
Read the full transcript
28:14And that scale requires us to be very intentional with how we build the team and how we build the company and what kind of solutions and who are the customers that we are working with. That's smart. That's a really smart approach. I have a Kubernetes question since you're an authority on the subject. Is Kubernetes still the right abstraction for the AI era? Or as these workloads are getting bigger, are we going to at some point be forced into something that is fundamentally different? Absolutely. I'm asking you to look into a crystal ball a little, I know. I think one of the things that I've learned through my entire career is that you can't force a change on users and developers and technology people.
29:01And the reality is that today, 90 % plus of our AI workloads are running on Kubernetes. That's where we are today. And that became the standard way of running workloads in the cloud, either public cloud or private cloud. That's the way we do it. And there is an amazing ecosystem around. So the goal for us is to make sure that we actually don't throw it away, but there are some things that has to change. and one of the things that I love that we bring along with the Kubernetes paradigm is for example the idea that it's not just that container orchestration or workload orchestration there is a huge ecosystem around that and that's something that we should bring along to the AI ecosystem and you know the problem of orchestration we just talked about it that's exactly what Kubernetes came to solve However, the opportunity, and that's actually what we've done at Coe, we've decided to build our managed Kubernetes differently in a way that will handle the resources more carefully.
30:11That makes sense. So our infrastructure is simpler. We are running Kubernetes on bare metal. We bring much more data and more observability into our machine control to make sure that we can make the right decisions. and we are also expanding the way that we are thinking about workloads. Like I mentioned before, it's not a single node, but it's multiple nodes. So those are the kind of things that we need to evolve. I don't want to lose the great things about Kubernetes that it brings to the table. That's fair. That's fair. I appreciate the perspective too. I think that's really interesting. I'd like to shift to physical AI a little bit for a minute.
30:47Of course. Because it's an exciting area that manages to get more exciting. If you'd have told me five years ago that the first people I would hear talking about physical AI would be in agriculture, I never would have believed it in a million years. But there's so many of the people I talk about, I talk to in the robotic space are doing neat things with tractors and weed reduction and all of this kind of stuff. And what I'm wondering is when it comes to physical AI, why do you think it needs a kind of hands-on model instead of simply like infrastructure and APIs? It's uniquely different, I feel like.
31:25It is uniquely different. Because first of all, the physical world is sometimes decoupled from the technology world. Yes. And definitely AI expertise. So that's like just the first thing that the gap that you need to bridge. The second thing that is interesting is that that field has always been dependent a lot on simulations and data generation. And that's what makes AI powerful. Okay, so this is the opportunity. That's what's amazing about it. So now we have those new tools that were not available before in an industry that is used to running simulation and building a lot of data. Because, you know, every experiment usually in that world is very expensive.
32:10You cannot make decisions very easily. So that's, I think, what we are really seeing companies lean into. They are seeing an opportunity to bring that technology to the field, to where they do the work, to the physical world. I remember an example about trying to reduce the vehicle calibration cycle from, I think it was three months to 24 hours. And I'm wondering, when AI starts touching physical engineering rather than images, pictures, text, what changes in and how you have to prove it's useful and that you can trust it? That's actually a very interesting problem. So first of all, I think that, you know, some of the tools that we are using right now, like you're saying, it's things that used to take much longer.
33:02and how they can be done at a shorter time, but you still have data from the past that you can compare it to what good looks like. So that would be one thing. And of course, there's a lot of effort around accuracy and checking those models. And there's a lot of work around simulation and data simulation. So we're using the models as well to see how things will look like. But that's definitely still early on as a space overall. Okay. Something that kind of shifting a little bit, I think I saw that at the end of Q2, CoreWeave had like 1.5 gigawatts, something like that, of active power and roughly 3.7, I think it was contracted.
33:47When you look several years ahead, do you think energy and physical infrastructure become the constraint on AI scaling? Or do you think maybe software gains are able to kind of help keep that moving and that you continue to sort of grow together the way it's been? 100 % we're going to be growing together on all of those tracks. We're not going to talk about putting data centers in space right now. I think before we go there, there is a lot of levers and a lot of innovation. And maybe where we started this conversation, I was talking about the space of infrastructure and how you can innovate. And then we talked about security and how you can innovate.
34:31And we talked about tooling and physical AI. With power and energy management, it's the same. Again, there is this new set of problems, hard problems that are creating an opportunity of innovation. And we are starting to see that. So I expect that we, of course, we will see demand continue to grow because we find those use cases that are applicable. Like, you know, the technology is applicable almost everywhere. Yeah. But the place you can innovate is almost also everywhere. And I think that's really exciting. It is. And there's more money and attention going into it than there's ever been, I would say, as far as like solving these problems and making sure things can continue working.
35:12And it's a unique time to be alive, I would say.
35:21Yeah, the things that I'm most excited about are the things that I think will materially change the way we live. You know, there's a lot of advancements in health and drug discovery and green energy. So there's things that literally were not possible before. and education. There's just, I love thinking about it. I understand that there's a lot of work to do it the right way, with the right controls and make sure that we are moving in the right direction, 100%. But I love seeing, you know, our customers leaning in and saying, okay, we are going to bring our expertise, our experience, and I'll marry it with this new set of tools and technologies and move forward.
36:10And that's a moment that I think is amazing for us to be in. And that's when you suddenly find yourself curing diseases and solving millennium problems in mathematics and all of these things that are happening. And we're so, with that kind of stuff, just scratching the surface, I think. There's so much that can be unlocked, new problems that will be discovered from everyone that's solved probably as well. And I like to think that's the promise that brought us all here. Couldn't agree more. And I hope that, you know, the work that we are doing, the thing that excites me is making sure that these tools and these technologies will be accessible.
36:54Because if it's kept, and this is not a new idea, like when I was 10 years ago speaking in the context of just Kubernetes. It used to be the most successful, highest velocity open source project. And the thing that excites me the most is that no matter where you are in the world, you could be part of this transformation. Yes. And I want to make sure that that's what we are also building right now. Because I'm sure that where I live, I have access to some set of problems, but someone maybe that lives somewhere else or have a different profession or is at a different age has access to different set of problems.
37:30And we should have the same tools and we should be able to collaborate. Yes. On those kind of problems, exactly like we talked about in this physical AI space, right? How we bring different domains together. And that excites me about the future. It is. It is. Good, clean, open science that where everybody grows together, I think is vital in this, especially in this space, if ever in any space. I love how optimistic this episode is. Right? That's good. Yeah. I try to be. I try to be. I also try to be. I really, I believe the good is out there and that we can get there. And it excites me. There's nothing I love more than reading some random story or just passing a tweet about how some amazing discovery has happened.
38:18And I'm always like, oh, this is so cool because this is just, we're just getting started. I think about space and understanding the universe and about so many things that I hope humanity finds itself able to unlock as time goes. Well, Henn, thank you so much for joining me. This has been an absolute delight. You're wonderful to chat with. And the work you all are doing over there is amazing. Thank you so much. It was a pleasure. So, Henn, where can people go to learn more? That's an easy one. So, of course, co-weave.com is always the best place to start. And we also have our conference coming up next week, fully connected.
38:55And what's amazing about this conference is that all the talks are with our customers and users. So it's all about how people are building AI and using AI. And my hope is that anyone that comes and attends can start doing something the same day, maybe differently, maybe learn something new. That's amazing. Well, to everyone watching, I hope you enjoyed today's video. Please take just a moment to like and subscribe to the channel. Don't forget to pop by the neuron dot a I and check out our daily newsletter read by about 700000 people every morning. We'd love for you to be one of them. And on that note, thanks for joining us.
39:34It's all we have for today. So farewell for now, humans.
39:47Thank you.
From the publisher
AI infrastructure used to be the part developers were supposed to forget about. What happens when the infrastructure itself starts determining what AI can do
Corey Noles sits down with Chen Goldberg, Executive Vice President, Product & Engineering at CoreWeave, to unpack the systems underneath the next wave of AI. Using CoreWeave’s newly announced multi-rack NVIDIA Vera Rubin NVL72 deployment as a jumping-off point, Chen explains why modern AI infrastructure is no longer just a question of how many GPUs you can get. Compute, networking, storage, cooling, observability, security, and orchestration all have to behave like one enormous machine.
The conversation moves from long-running AI agents and latency to Kubernetes, security guardrails, physical AI, energy constraints, and the changing role of software engineers. Chen also explains why CoreWeave sees AI as a transformational rather than incremental shift — and why making this infrastructure accessible matters if the benefits are going to reach more than a small group of companies and researchers.
For more practical AI news, tools, and analysis, subscribe to The Neuron at https://theneuron.ai.
Sponsored by Outshift:“Scaling Out Superintelligence” Vijoy Pandey, January 2026. The technical whitepaper detailing the Internet of Cognition architecture, three-layer stack, and cognition state protocols. https://outshift.cisco.com/internet-of-cognition/whitepaper?utm_campaign=fy27q1_outshift_ww_paid_ioc-neuron-wp_podcast&utm_channel=podcast&utm_source=podcastInternet of Cognition Interactive Demo Clickable walkthrough showing per-agent activity, intent, context, and collective reasoning across a multi-agent SRE system.https://outshift.cisco.com/internet-of-cognition/explore?utm_campaign=fy27q1_outshift_ww_paid_ioc-neuron-explore_podcast&utm_channel=podcast&utm_source=podcast
