In short
The “software infrastructure” layer that feeds AI GPUs—Vast Data’s role in building AI factories (data, storage, networking, databases, and policy) and its new model-management announcement focused on confidential computing (“Data Enclave”) so enterprises can run inference without exposing model weights or sensitive data.
Guest backgrounds
Renan Halleck, founder and CEO of Vast Data (valued around $30B). Previously worked at ExtremeIO (employee #1; acquired by EMC in 2012) and led EMC’s R&E organization. Founded Vast in 2015/2016 after seeing neural nets succeed due to faster access to much more data.
Key claims
AI demand is accelerating (customers expanding from hundreds of petabytes to extra exabytes). Traditional “shared-nothing” storage/database architectures don’t scale for AI; Vast uses “shared everything” (many nodes share all data via fast networks). Enterprises need cloud-like simplicity but must retain control of IP (data, models, weights, agents). Confidential computing enables on-prem inference with encrypted memory.
Notable examples
500 petabytes → +2 exabytes planning story; Singapore camera feeds producing trillions of vectors; inference must be 100% up; encrypted inference via NVIDIA hardware and enterprise enclaves.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOThe AI Layer Cake
0:45 to 1:33
Discussion on the five-layer cake model of AI infrastructure and its significance.
“Please enjoy my conversation with Renan Halleck.”
Building the AI Factory
1:33 to 4:40
Exploration of what constitutes an AI factory and its requirements in modern data centers.
“So at a high level, what does that mean, software infrastructure in the context of AI building?”
Enterprise AI Adoption
4:40 to 8:13
Analysis of how enterprises need to adapt to leverage AI technologies effectively.
“But in all of these cases, enterprises will need AI factories in the medium term.”
Agents and Fine-Tuning
8:13 to 10:40
Examination of the role of agents in AI and how they can evolve through learning.
“and have new companies with AI experts take their place.”
Intelligence Marketplace Concept
10:40 to 13:16
Discussion on the potential for creating a marketplace for AI models and their utilization.
“And within teams, they can learn from each other.”
Model Management in AI
13:16 to 14:00
Introduction to new developments in model management within the software infrastructure layer.
“And I guess that goes to your big announcement of this week.”
Model Management and Confidential Computing
14:00 to 15:32
Learn about the advancements in model management and the role of confidential computing in AI.
“How do we as the operating system now get an understanding of this model is really good at coding and that model is really good at generating images.”
Renen's Journey and Mathematical Background
15:32 to 16:45
Explore Renen's mathematical journey and its connection to AI development.
“But taking a step back, I would love to talk about your journey a little bit and then how it led to you to build Vast, which is this incredible company, which is valued at$30 billion.”
P vs NP Problem and OpenAI's Claims
16:45 to 17:44
Discover the significance of the P vs NP problem and recent claims about its resolution.
“It's much easier for us to review a document and say, yeah, that works than to write that document or to look at a picture and say that's beautiful than to paint that picture.”
Computers Accelerating Problem Solving
17:44 to 18:55
Understand how computers can enhance problem-solving capabilities in mathematics.
“But going back 20 years, it was clear to me that if we could get computers to help with these hard problems, that's the way to solve it because people are slow and the evolution of our knowledge is slow.”
Show all 31 chapters
The Founding of VAST and AI's Evolution
18:55 to 21:32
Learn about the founding of VAST Data and its role in the evolution of AI technologies.
“The last one I was with was a company called EMC, which builds large storage systems.”
Revolutionizing Storage Architecture for AI
21:32 to 25:05
Explore how VAST Data's architecture redefines storage systems for AI applications.
“They were doing medical imaging analysis, genomics analysis, hedge funds.”
Building Databases for Scalable AI Solutions
25:05 to 28:00
Discover the need for new database systems to support the demands of modern AI.
“and we realized over the years, our customers forced us to realize that it's not just storage that needs to be redone, but the entire part of that stack that needs to be redone.”
Challenges of Traditional Databases for AI
28:00 to 29:18
Explore the limitations of traditional database architectures in the age of AI.
“We realized that because our customers told us that.”
The Evolution of AI Cloud Services
29:18 to 31:30
Discuss how AI clouds differ from traditional cloud services and their importance.
“And then, you know, I think some people would say, well, storage was solved.”
Measuring Customer Happiness at VAST
31:30 to 32:44
Learn about how VAST measures customer satisfaction and its significance.
“So obviously data gravity is a wonderful thing if you're vast.”
Data Movement: Training vs. Inference
32:44 to 38:26
Understand the differences in data handling between training and inference in AI.
“Let's get into some of the technicality of how the data interacts with models.”
Security and Collaboration in AI Systems
38:26 to 41:38
Examine the security measures and collaborative potential of AI agents.
“And to what you just said, the concept of agent memory works the same way.”
Data Enclave: Protecting Sensitive Information
41:38 to 42:00
Learn about the concept of Data Enclave and its role in data security.
“And so hopefully that brings us to a level where the AI can help us solve some of the big problems that we've been struggling with over the years.”
Understanding Data Enclaves
42:00 to 43:18
Learn how data enclaves enable secure data processing for enterprises.
“I'm not sure I like the name, but that's what we chose.”
The Evolution of Confidential Computing
43:18 to 44:43
Explore the evolution and importance of confidential computing in AI.
“who allow us to bring encrypted data through the network, into memory, from memory, up to the GPU.”
Building Trust in AI Ecosystems
44:43 to 46:04
Discover the significance of trust and collaboration in AI ecosystems.
“And it also requires many companies, many parts of the ecosystem to work together in order to provide this end-to-end solution.”
Demand Dynamics in AI Computing
46:04 to 49:53
Examine the growing demand for AI compute power and its implications.
“we're very very humbled to be in the middle and in that position.”
Profitability in the AI Stack
49:53 to 51:25
Understand how profitability influences competition in the AI space.
“Or do you think that's just like a means to whatever needs to happen for the ecosystem to be actually built?”
The Future of NeoClouds
51:25 to 56:01
Learn what separates successful NeoClouds from those that falter.
“Speaking of finance, you're in a very interesting and somewhat unique position because based on the little ad, I think you guys are profitable.”
The Innovator's Dilemma in AI
56:01 to 57:32
Learn about the challenges faced by companies managing legacy systems while innovating in AI.
“When you have something new to build and you have this old cash cow to focus on, it's a lot harder than when you have something new to build and you don't have that legacy.”
Sovereign AI: A Basic Need
57:32 to 59:20
Explore the emerging concept of sovereign AI and its implications for nations.
“It seemed to me that the concept of AI factory that we discussed at the very beginning was largely around this concept of sovereign AI like a year or two ago.”
Understanding Value in the AI Stack
59:20 to 1:01:13
Discover how value is distributed across different layers of the AI stack and the potential for commoditization.
“and which part of the stack commoditized?”
Navigating Partnerships with NVIDIA
1:01:13 to 1:04:00
Learn about the dynamics of working with dominant players in the AI ecosystem like NVIDIA.
“In the old stack, we had only four layers and the new stack perhaps that merges as well.”
Accelerating Innovation in AI
1:04:00 to 1:06:55
Gain insights into how companies can keep pace with the rapid changes in AI technology.
“is that you guys seem to be just relentlessly launching new products, expanding new customers, having key relationships.”
Future Predictions for AI and Society
1:06:55 to 1:09:47
Consider the potential transformations in society and industry due to advancements in AI over the next decade.
“projecting ourselves in the future a little bit as possible as it can be today.”
Transcript
Automatic transcript. May contain errors.0:00Sometimes it scares me. We had a customer, one of these AI clouds, they said we're probably going to need about 500 petabytes over the next three years. Last week they came back to us and said we're going to need an extra two exabytes on top of that 500 petabytes. I think in the next 10 years we'll see more difference than we did in the last thousand years. Hi I'm Matt Turck, welcome to the Matt Podcast. We talk a lot about AI models on this show and we talk a lot about compute. But today is about the layer in between, the crucial software infrastructure layer that feeds massive amounts of data to all those GPUs.
0:30My guest today is Renan Halleck, founder and CEO of VastData, a company that is surprisingly under the radar, considering it's been most recently valued at$30 billion and powers top AI players like XAI and some of the biggest AI neoclads in the world. Please enjoy my conversation with Renan Halleck. Renan, welcome. Thank you. All right. So to anchor this whole conversation, I wanted to start with that mental model from Jensen at NVIDIA that I find so helpful, where when he describes AI, he uses the analogy of a five -layer cake where you have power at the very bottom and then hardware, then software infrastructure, then models, then applications.
1:17So where do you sit in that layer cake? Right in the middle. That software infrastructure layer is the one that we do in collaboration, of course, with the hardware vendors underneath us and with the model builders on top of us. Yeah. So at a high level, what does that mean, software infrastructure in the context of AI building? Yeah. So everything you would imagine from the old stack, we need for this new stack. We need to manage the compute. We need to manage the storage. We need to manage the networking. and then as you go up the stack, we need to make sense of all of it, make sense of the data through a database, make sense of the compute through training and inference functions, and as you go even higher up the stack, and we shift from training predominantly to inference and agents and reinforced learning, we need more and more tools to enable this new world, make it simple, make it safe, make it fast and scalable and resilient, all of that sits in that middle layer.
2:21Sometimes I like to call it software infrastructure, and other times I like to call it the operating system for this new era. There is a related concept to that five-layer cake that thing may have originated from NVIDIA again, but it keeps coming back in conversation, which is this idea of AI factory. What does that mean? So if you start with a bunch of GPUs, what do you need to add to create an AI factory? I think even below the GPUs, the data centers are looking very different today than they did five years ago. The old stack, you had 10 kilowatt racks. Today, you need 500 kilowatt racks. And so you need more power.
3:03You need more power density. You need GPUs, of course, versus CPUs. You need extremely fast networks and very large SSDs rather than 10 gig ethernet and old hard drives. And so every layer of this cake needs to be redone. And from our perspective, you need much more data and much faster access to that data because we're no longer analyzing numbers and columns of a database the way you did with big data in the old stack. We're doing pictures and video and sound and natural language, and that takes a lot more space. and that's before new generative AI data that gets generated. And of course, all of this compute, all of these GPUs need very fast access to this information, whether they're training on it or whether they're running low latency inference functions.
3:58And so the entirety of the stack is new and I think the culmination of this stack is what we call the AI factory. And is the concept of AI Factory mostly for the big labs or like this whole like data center plus model kind of thing? Or is it also for enterprises? I think it'll be for everyone eventually. The big labs are starting. They're ahead of the pack and they're leading the way for everybody else. But anybody who stays with the old stack, with old business information systems will get left behind. And so people need to start leveraging these new AI abilities. Some start by leveraging them out of the clouds, out of the hyperscalers, the neoclouds.
4:46Others build within their environments. But in all of these cases, enterprises will need AI factories in the medium term. And then after that, I think as we shift from people-based organizations to augmenting them with agents, these AI factories become a lot more than what we're used to thinking from just our compute infrastructure. To play it back, right now, if you go to a large enterprise, a non-tech enterprise, so like Pfizer or Goldman Sachs or Walmart, I think the dominant view is that leveraging AI basically means, you know, you start partnering with Anthropic or OpenAI and, you know, you have FDs coming into your walls and then, you know, you build some stuff.
5:37Initially, that was chatbots. Now it's agents. Is the concept of an AI factory that you build your own AI? So you have your own data infrastructure, but you also rent your GPUs or buy GPUs, and then you do model work. Is that the concept? I think once you get to a certain level of scale, definitely it's more cost effective to buy than to rent. And so from that perspective, yes, you should build your own AI factories. but beyond that over time I think all of these organizations IP will be distilled in models it will be weights and they don't necessarily need to build the base model I like to give analogies of the human brain we were all born with a base model but then over time as we travel the world we learn new things we infer during the day we fine-tune at night and tomorrow will be a little bit smarter than we were today.
6:35And my model of the universe is slightly different than your model of the universe because we had different experiences. And so I think in the same way, all of these agents will fine-tune models through reinforced learning over time. And every organization's IP will be able to be distilled into the models that they own and the weights that they've accumulated. What does that mean in terms of what those enterprises need to be able to do? You know, like doing model work and like weights and GPUs, that's a whole kind of expertise that ultimately few people in the world already own or already know, have.
7:17And then a lot of those companies just don't have that. So how do they go from here to there? I think, again, I like analogies. The analogy I would give here is you go back to the 1970s, you needed a PhD in computer science to operate a computer. Today, everybody can do it. And the reason everybody can do it is not because we learned. It's because operating a computer became much simpler. The operating system made it intuitive, made it easy for everybody to use. It also made it safe for everybody to use. These enterprises are under a lot of regulation that they need to comply with. And so it is our job as the ones building that software infrastructure layer to make these new technologies accessible to the enterprises such that they don't need to think about it too much.
8:07and that they can start adopting them at pace rather than get left behind and have new companies with AI experts take their place. Maybe to push back a little bit or maybe to just dig in to make sure I understand. That's the whole premise of the cloud is to make things incredibly simple and use those functionality as a service. The concept of an AI factory revolves around owning a lot of this. So it sounds like very kind of like on-prem, private cloud, which historically has been a lot more complicated. So is effectively what you're saying that is that like there's an opportunity to like abstract away all of this in a way that feels cloud-ish in experience?
8:49I think so. I think the physical location of this AI factory can be on-prem, can be in a colo, can be in an AI cloud where you rent it from somebody rather than manage it yourself. Eventually, my guess is that the hyperscalers will also enable these AI factories to be built within their premises. But so long as it's within your control and so long as you're the one that owns it in terms of really, I think, owns the information, owns the models, owns the weights, owns the agents. I think that's the distinction that I'm trying to make here. Doesn't necessarily need to be physical equipment within a building that belongs to you.
9:31Okay, great, great. Do you want to spend more time on what you alluded to a minute ago about agents and agents doing their own fine tuning? I mean, right now, the way I think most people think about agents is that, you know, you do reinforcement learning to enable the agents. But I think you're saying agents do it. I, again, like to think of agents as artificial people. And over time, I think we're all going to have agents that work for us. whether we're people or organizations for different tasks. And we'll have teams of agents that can help us develop software or personalized medicine or whatever it is that we want done, we will have them do on our behalf.
10:16And I think over time, we will get attached to our agents and we will expect them to know things that happened yesterday or that happened a year ago. And that will allow us familiarity with our agents, but that requires them to fine tune models based on experiences that they had with us. And of course, that also makes them smarter and able to do more for us. And within teams, they can learn from each other. And I think that generates the next level of artificial intelligence versus where we are today, where everybody's based on some base model from one big lab or another. And every few weeks we get an upgrade to that model based on what they did in their environment.
11:00in this new way, each one of us, each organization owns the IP generated through these fine-tuning mechanisms rather than letting the big model builders basically own everything over time. Okay. So just to play it back, so that the knowledge loop happens within the enterprise and that becomes the new IP. And so we're saying we cannot do that with OpenAI or Anthropical, whoever. You don't want to because you don't want to expose your proprietary information, your processes, your secret sauce out to somebody else such that they can take it and run with it. Which is super interesting, right? Because people used to say that about data, but now we're saying this about...
11:44Intelligence. Intelligence, processes and memory and therefore intelligence. Okay. Does that mean that open source becomes the play here? Because if you want to... fine-tune or do RL against model, having access to an open source model enables you to do that? So is that your vision of the world? Not necessarily. I think we need to build a way for people to monetize these abilities that they worked so hard and paid so much to get to. And so we need a way where, on the one hand, an enterprise does not expose their data. On the other hand, a model builder does not expose their weights. and again the software infrastructure layer is the right place to enforce that type of trust and over time as the enterprise starts to generate their own models or fine-tune on top of base models they should also be able to monetize the skill sets that they develop and so if I run a carpentry shop and my carpenter robots are the best in the business I should be able to lease to you that brain module such that you can build a chair or whatever it is that you're trying to do.
13:00But I'm only giving it to you for a week. And at the end of that week, you don't have access to it anymore. Or maybe I'm selling it to you, but I still don't want you to see those weights. And so there should be ways for us to build an intelligence marketplace. Okay. And I guess that goes to your big announcement of this week. Do you want to mention the principle of it? And then we'll go into the details in a minute. But at a high level, how does what you guys just announced address this specific problem? Sure. So basically what we're trying to do is fill out that software infrastructure layer.
13:39And what we've realized was that the five-layer cake that we discussed is almost complete. The models don't really live on top of the software infrastructure. Models over time have become a resource, just like a GPU is a resource or an SSD is a resource and resources need to be managed. And so this announcement is all about model management. How do we as the operating system now get an understanding of this model is really good at coding and that model is really good at generating images. And this one is expensive and that one is inexpensive. And that way we can understand which tasks belong under which models.
14:22And that's today when we have four or five of them. As we start to reinforce learning and fine tune, as we start to have hundreds and thousands of agents and over time generating new derivatives of those models, that's going to be a much bigger task. One specific aspect of model management, which we're announcing, and we have a lot of the model builders as partners join the announcement, is around confidential computing. And enabling that ability of enterprises to run inference on-prem without exposing the model builder's weights. And so both sides can rest assured that my data is not exposed and your weights are not exposed to me.
15:08And so with help from NVIDIA and encrypted memory underneath, we're able to make sure that no one sees what they're not allowed to see and to hopefully increase the total addressable market of these model builders into regulated industries, into large enterprises that are careful with their sensitive information. Let's put a pin in this and we'll revisit in a few minutes and go into how that actually works. But taking a step back, I would love to talk about your journey a little bit and then how it led to you to build Vast, which is this incredible company, which is valued at$30 billion. Interestingly, perhaps because you're right in the middle of the cake, not everybody has heard about yet.
16:02So I think you were telling me growing up in Israel that you had a particular affinity for math. So my words, not yours, but you were a bit of a whiz kid in math. Affinity for math, for sure. Whiz kid, definitely not. It's funny, when I came up here, the math museum is downstairs. And we had some big news this week about math. Um, when I was in school 20 years ago, I spent, uh, six months of my life trying to figure out if P equals NP, uh, which I hope your audience knows. But for those who don't, it's the question of, is it easier to, uh, verify a problem than it is to solve the problem, which intuitively it is.
16:49It's much easier for us to review a document and say, yeah, that works than to write that document or to look at a picture and say that's beautiful than to paint that picture. But mathematically, it's unknown if that is the case or not. And interestingly, this is one of seven millennium problems that the Clay Institute back in the year 2000 put a million dollar bounty on. And I remember spending those six months trying to solve it because I thought, if I prove that P does equal NP, then I also solve the other six and I'll get$7 million. And if I prove that it's not equal, then at least I get one.
17:27And after six months, what I realized was that I'm not nearly smart enough to do it. Do you want to provide the full context about what happened this week just for anybody that didn't follow it? Another one of those problems, Navier-Stokes, was solved by OpenAI, or at least they claim that they've solved it. And the rumors are, hopefully by the time that this is aired, some of those rumors will come true, that there's at least two more of the seven problems that OpenAI and Anthropic are trying to solve and are close to solving using AI. and these are problems that have been open for nearly 100 years and for the last 26 years there was a million dollar prize associated with them and still people were not able to solve and so for the first time I think we're at a point today where AI has proven that it can do things that we don't know how to do.
18:22But going back 20 years, it was clear to me that if we could get computers to help with these hard problems, that's the way to solve it because people are slow and the evolution of our knowledge is slow. And if we are to see any of these things get solved within our lifetime, we need computers to speed us up. We didn't know how to do it back in 2006. And so I went along my way and joined a few technology companies. The last one I was with was a company called EMC, which builds large storage systems. And if I recall correctly, so you joined a startup as employee number one. Extreme IO, yes. And that company was acquired by EMC.
19:06That's right. And then you helped run that, the R &E organization within EMC for that And that acquisition happened in 2012. And by 2015, it became clear that there is a shot at computers helping us think and helping us solve problems. Because neural nets, which again, when I was in school 10 years earlier or 15 years earlier, were a curiosity. They didn't really, weren't able to do anything. now they were for the first time starting to show that they can do something they were able to recognize which videos had cats in them and which videos did not have cats in them which was again astonishing to me because suddenly we can try to mimic the human brain without really understanding the human brain and so immediately I tried to figure out how this happened and it was very clear that it wasn't necessarily new algorithms.
20:07It was giving those neural nets much faster access to a lot more data that enabled them to perform these very simple tasks. So that was the intuition behind the founding of VAST. Correct. And that was 2015? 2015, beginning of 2016, we started. Yeah. And then obviously that was at a minimum three years before the Transformers paper came out. and another two years before the chat GPT moment. So what was the idea then? Because there were not a lot of customers using massive amounts of data for AI at the time. There was no generative AI at the time. Google acquired DeepMind a year previous to that.
20:51But it was clear that none of the systems, none of the infrastructure that we had available to us would be scalable enough, would be performant enough to enable this next revolution. The architectures were built on an assumption that things needed to be fast or big. And here you needed something that was both fast and very, very large. And so we needed to build a new underlying layer to enable the success of this next revolution if it were to come to fruition. In the early days, of course, our customers were not doing generative AI. They were doing large-scale analytics. They were doing autonomous driving projects.
21:35They were doing medical imaging analysis, genomics analysis, hedge funds. So traditional scientific computing. Correct. Hedge funds were trying to get signal from news feeds and natural language. And so it was the beginning of AI, but before ChatGPT. And actually, as a way of getting into what VAST actually does, do you want to explain in super simple terms what you just said? So you started from storage, and then storage, as you just said, was either – the alternative at the time was either slow and inexpensive or fast but expensive, and the whole scientific kind of breakthrough is to break this, right?
22:26And was it what you called days? Disaggregated, shared everything. All right, so explain that in super simple terms. So the old way to build scalable systems is based on a concept called sharding, what's known as shared nothing systems, where you have a lot of nodes. Each node is responsible for a piece of the pie. They collaborate with each other in order to serve up application requests. And as you add more nodes, you get more performance, you get more capacity. It works well, up to a certain limit. And so once you have too many nodes, the communication between them starts to show diminishing returns.
23:04And so if you're analyzing text, it may work. If you're analyzing numbers, it works well. If you're analyzing columns of a database, that's fine. Once you grow beyond that, and today the amounts of data are four or five orders of magnitude beyond where they were back then, soon they will be seven, eight orders of magnitude bigger. It doesn't work. It breaks down because the communication grows quadratically within the cluster. And then from a resilience perspective, also, when one part fails, everybody needs to recover. That takes time. You can't really have more than 100 nodes in one of those systems.
23:42To build AI, we need systems with many millions of nodes. We have one system at one of our customer sites that is today multiple exabytes within a single cluster delivering tens of terabytes per second. This is tens of thousands of nodes even today before the agentic revolution. And so to do that, we needed to build an architecture that ended up being the opposite of the way shared nothing systems work. instead of having the SSDs, the media directly attached to the CPU, it's on the other side of the network. And that through a very fast network and through new protocols like NVMe over fabrics and through new media types that allow us to save not just data on the other side of the network, but also metadata, we were able to have all of the nodes see all of the information as if it was directly attached.
24:37And that's what led us to being able to do what we call shared everything, All of the nodes now don't need to communicate with each other because they share access to all of the information. That's the basic idea behind our architecture. And it proved really, really good for a new type of storage system for AI and then for a new type of database for AI and then for a new way to orchestrate compute for AI. and we realized over the years, our customers forced us to realize that it's not just storage that needs to be redone, but the entire part of that stack that needs to be redone. As I'm just saying, I'm smiling.
25:19It must have been a lot of fun to pitch a storage company in 2016 to DVC. The storage was probably considered at the time the least sexy part of the entire stack. I owe you guys a lot of our success because as hard as it was to pitch storage to VCs, now we're reaping the benefits of having a lot less competition than we otherwise would have had. And so, thank you. Yeah. Is that part of why that middle layer is less known? There's just less companies pitching it? I mean, you don't have two comments necessarily, but the only companies that come to mind are like Wicca. And there doesn't seem to be a lot of companies doing what you do.
26:04There aren't. The ones that we're competing with started three years before us or 30 years before us. And when we built this new architecture, we were basing it on forward looking technologies. When we started in 2016, the underlying parts weren't there yet. They became available throughout 2017, 2018, 2019. So you had to believe that they were going to appear. And so it wasn't just difficult to raise money from VCs in those early board meetings when they kept asking me, what do we do if this part doesn't come to life? I kept telling them we don't have a plan B, which they didn't like. And so I was not their favorite person in those days.
26:45Hopefully that changed over time. They made some money. but yes there was no way for a company that started before us to make those forward bets and so that's our advantage today that's our moat the fact that we're the only ones with this architecture and it's true that there aren't many other companies in our layer I think mainly because it's it's not sexy it's a lot more appealing to build a model company or an application company than it is to build the plumbing. And the DAES architecture is still at the very basis of what you do, all products currently. That's right. We have one product that's this operating system, but all of the parts that we're adding are based on DAES.
27:32And that's part of the reason why we can't take open source and bolt it on top. We have to write most of it ourselves. Why did you start building databases on top? was not addressed by the market that you needed to build? Same as with storage. We realized that there was a trade-off between price and performance and scale and resilience and ease of use in the world of database in the same way that there was in the world of storage. We realized that because our customers told us that. We have a big customer, for example, in Singapore, and they have a lot of camera feeds coming in and all of those feeds need to be vectorized as the images come in and that results in trillions of vectors that need to be placed in a database.
28:25There isn't a database that was built for trillions of rows. There isn't a database that was built for thousands of agents It's analyzing different aspects of what's happening in the system at the same time. Old stack databases were built for big amounts of data, which today seem very, very small. They were built for people running queries and doing the analysis. And again, that old architecture, the shared nothing architecture, is not scalable enough to the levels that we need today. So it's scale, right? I mean, just to push back a little bit or probe, you know, I'm sure like Databricks or like Snowflake would say, well, no, no, no, we massively scalable and we can just run massive analytics on data and including some in real time with our newer products.
29:23And then, you know, I think some people would say, well, storage was solved. Like that's S3, that's infinitely scalable. But what we're saying here is that all of this may be true for certain use cases, but not for what we're talking about here. I think both of those companies that you mentioned started, again, about five years before us. And so they, too, are built on that older architecture and they started top down. They are based on underlying cloud services like S3, which were not built for this era of AI. And we're seeing that today. that's a big part of the reason these AI clouds are starting to pop up because the traditional hyperscalers that built S3 built the old stack and these new AI clouds are building the new stack and now the older clouds need to adapt in order to stay relevant.
30:20So is the ask to customer as they build those AI factories to basically move all the data to VAST? They don't have to start by moving all the data to VAST. Usually we start by just serving the new applications, just serving the new AI applications. Usually what happens after that is that the customer, the enterprise, the organization realize that this is actually useful and I want it to access the entirety of my data set rather than just this subset. The good news is that we support all of the old stack interfaces. And so if you're talking storage, S3, we support file systems, we support block devices, we support, if you're talking databases, we support SQL.
31:09If you're talking streaming, we support Kafka. And so it's very easy to plug in those old stack applications onto our platform in parallel to the new agentic workloads and AI applications. And in that sense, we become a bridge between the old world and the new world as companies make that journey. But the end of the journey is that the old world merges into the new world. Yes. So obviously data gravity is a wonderful thing if you're vast. How do customers feel about the idea of like moving all the data ultimately to one vendor? So because the interfaces are all standard, it's just as easy for them to move off of us as it is to move onto us.
Read the full transcript
31:57It's our job to make sure that they're happy. The one chart that I look at every morning is our customer happiness chart. That's the only thing we care about. What does that look like? How do you measure customer happiness? It has three colors. It has green, it has yellow, it has red. And anybody in the company is allowed to move a customer from green to yellow or from yellow to red, but very few people are allowed to move them back. And so... But based on what metrics? The only metric that really matters in order for us to move a customer back is that they say so. That they say we're happy now, everything is good, everything is solved.
32:35And yeah, that's what we work so hard every day to accomplish. I think the numbers show that it's working. the numbers that people like you usually care about are net dollar retention gross dollar retention we don't have any churn in the company our gross dollar retention it's not a hundred percent because a few of our customers have gone out of business over the years but no one's actively decided i'm going to stop using vast which is something i'm very proud of and on On average, our customers double and triple the amount of data that they place on our systems every year. Let's get into some of the technicality of how the data interacts with models.
33:23What does it actually mean to move data to the GPU? And is it different for training compared to inference? Training is super simple. training you just need a lot of gpus and you need to feed them with a lot of training data once in a while they put down a checkpoint so that if something crashes they don't have to go all the way back to the beginning but other than that there's not much to it you need large scale you need fast access that's it uh inference is becoming a whole complex world of its own Inference started by obviously being requiring low latency access because people are on the other end prompting and they want fast responses.
34:11That meant that you need a distributed system close to where those people are at the edge. That also means that you need a system that's always up. Somebody once told me, one of the big model builders, that if training is down, nobody notices. If inference is down, there is no service. And so it needs to be up 100 % of the time. Resilience, distributed nature, low latency, very secure environments. All of that is important in inference versus training. But that's just the ground level. Once we start talking about multiple models, once we start talking about agents, once we start talking about reinforced learning, definitely once we expand outside of the world of data centers into physical AI and actual devices, each one of those adds another layer that means that we need to develop a new set of tools.
35:07For example, now that we have multiple models, each model is good at different things. And so I want to make sure that this prompt makes its way to that model because it's good at software development. I want to make sure that this prompt makes its way to that model because it's inexpensive and this isn't a high value task. I want to make sure that no sensitive information is exposed to this model because they're going to use that information for things that I don't want them to. Whereas here I have a contract that protects me. And so just that aspect of it is starting to get complex. once we start thinking in that direction there's all kinds of optimizations that we need to do around things like heavy caching and saving context windows and making sure that we route the prompt not just to the right model but to the right GPU that has that model loaded and that has my context in cache context is just the last few minutes of conversation we need short-term memory the last few weeks or long-term memory what happened two years ago to also be accessible through things like RAG into these environments.
36:18And so all of that requires us to manage data, the raw data, the generated data, the input data, the weights of the model, the context. And we need to make sure that nobody sees information that they're not allowed to. We need to keep track as we do reinforce learning of what information went into fine-tuning each model and make sure that that information isn't leaked through interacting with somebody who's using that model, whether it's a human or an agent, it becomes very, very interesting as you layer one, two, three, four of these things one on top of the other. So if you have enterprise data to play it back, that data could be used by the model through the context window, through the KV cache, which is the model's memory, or through RAG.
37:11Who makes those decisions? Is your job as the fast data provider to just be available and then the model pulls as needed? How does that work? So we need policies to be set. And today we just connect into your organization's authentication and authorization mechanisms. If you use Active Directory, if you use access control lists, we then know who is allowed to see what from a human perspective. We then extend that to agents as those inherent properties from the people that deploy them or from other agents that spawn them. And based on that, we know what's allowed. What are the rules of the game?
37:51Then we need to enforce those rules, which becomes very difficult when you're talking about multiple organizations. most of these model builders won't let you infer within your environment which means that you just need to send your data to them and hope for the best that i think has been slowing down the adoption of ai in a significant way and so we need to make sure that it's secure we need to make sure that it's safe um we need to make sure that there aren't any obstacles uh for ai to achieve its true potential. And to what you just said, the concept of agent memory works the same way. So accessing that vast data storage.
38:38Did you say you effectively give the agent authentication authorization the way a human would? Very similar, yes. So access control lists that apply to a user ID. Now this user ID is assigned to an agent rather than a person. We have access controls based on roles. We have access controls based on attributes. And so it's very fine granularity. And you can, in a dynamic way, set your policies. It's accessing data, but also what tool calls they can or cannot make. Correct. And who they're allowed to talk to, because now you have multiple agents talking to multiple people. And so that conversation happens over a persistent streaming service that we provide That allows us then observability into it because we can query through the database and ask questions about those conversations.
39:30And so the entirety of this, remember, it's not in one data center anymore. It's across geographies. And so the entirety of this new world gets managed by this operating system from a observability and from a control and from a policy setting perspective. especially when it starts to expand outside the data center into cars and humanoid robots and satellites then the stakes get higher because physical elements can create physical damage and so rather than just data safety we need to think of things like actual safety that robot should not pick up that tool or that car should never accidentally hit a person.
40:18And so you need us to be in the data path to not just be able to report on these things, but also to intervene and enforce the policy. And in the meantime, the concept of agent collaboration, which is all the rage now, you know, highlighted in part through the recent security and safety issues, but the the concept of agent swarm is everywhere. So that collaboration layer would be also enabled the same way they would have access to a group, like the way humans would have access to a group access to data. That's exactly right. And I think for a long time, we've been hoping that interaction between AIs, interaction between agents, will give us that next step function in a level of intelligence in the same way that when we talk, we make each other smarter and we give each other ideas Because I think this solving of the Navier-Stokes problem earlier this week proves that that is actually happening.
41:21They set 10 ,000 agents loose at an age-old math problem, and within a few days, they came back with the result. And so that's one aspect that was missing, I think, which is good for theoretical problems. The other aspect that's missing still is access to the natural world, which is good for physical problems. And so hopefully that brings us to a level where the AI can help us solve some of the big problems that we've been struggling with over the years. All right. Amazing. We said we would go back to that private and sensitive data and exposing it to commercial models. What is that called, by the way?
42:03It's called Data Enclave? Enclave, yes. Enclave. Okay. I'm not sure I like the name, but that's what we chose. So how does that work practically in simple terms, but practically? It's super simple. In fact, it all goes back to encrypted memory. And so if today model builders want to keep their weights very close because that's their IP and they're afraid of giving it to a million customers and having one of them take it and do something that they're not supposed to. On the other hand, a specific class of customers, specifically enterprises, cannot just give their data to the model builders in the way that they're expected to today.
42:49What we're doing is we're letting that inference process, rather than running within the premises of the model builder or within one of their clouds, we're letting it run within the premises of the enterprise. whether it's a physical enterprise, a premises or a VPC, a virtual private cloud that they own. And the way that we make sure the model builders are happy is by encrypting those weights and we encrypt them end to end. We leverage hardware abilities underneath us from companies like NVIDIA, who is a big part of this launch, who allow us to bring encrypted data through the network, into memory, from memory, up to the GPU.
43:31And only the model builder's inference application is allowed to access it. And so we have full custody, chain of control, end to end, such that those waits are safe and the enterprise knows that nobody can ever see their information. And why is it possible now that the concept of confidential computing, I'm trying to remember the name, there are two or three names for that concept. It's been around for like, I don't know, a decade plus. Why is it happening now? I don't think there's anything new in the underlying technology. It had to be plumbed because this was, as you say, available for a long time for CPUs.
44:16Now it's available for GPUs. And so that piece of it needed to be done underneath us. Our layer requires us to manage all of this and make sure we understand which model is which, who does it belong to, who is allowed to see it, where should we block it, those types of things. I don't think there's a technological breakthrough that was required for this. I think it's just an idea whose time has come. And it also requires many companies, many parts of the ecosystem to work together in order to provide this end-to-end solution. And so we're working with the AI clouds to do it. We're working with the OEMs to build physical boxes that can sit in the enterprise's data center.
45:02We're working with the model builders. There is an appliance concept. Exactly. Yeah. Companies like Cisco or Supermicro are now collaborating with us on this. And of course, the chip makers themselves like NVIDIA. And beyond the technology, there has to be a trust and auditability layer, presumably, right? Because now you become the bottleneck, like you're in the middle of like this great confluence of just different players from across the ecosystem. So how do you think about that? I think that's where the trust should be. That's the layer that is, on the one hand, malleable because it's software.
45:43On the other hand, it's low enough in the stack such that it can benefit all of the different applications. historically that's where safety was done in the operating system and so we needed to be done in our layer I hope that we are trustworthy we're working very hard to build that trust with everybody else in the ecosystem and yes we're very very humbled to be in the middle and in that position. That's like another huge milestone for the company. So amazing. I'd love to take a step back and just like think through the macro environment from your perspective. You're in this very privileged position as the middle layer of the cake, as we said a couple of times now.
46:37I'm curious what your take is on a bunch of things. So the inevitable question is around the demand side of the whole compute boom. So you work with a lot of the NeoCloud and NVIDIA. So you're intimately familiar with the supply side. You're part of the supply side yourself. what gives you comfort that the demand side is going to materialize in a way that is not going to get this entire AI ecosystem in trouble? I think the demand is there today. I'm not sure it gives me comfort and sometimes it scares me as to how much demand is there and the acceleration in demand because every quarter, our customers are coming back to us and telling us they need a lot more than they thought they did.
47:33We had a customer, one of these AI clouds, one of the smaller ones, come to us a quarter ago and we told them we need to plan ahead for the next three years because there's supply chain implications and you need to be aware of this. And they said, we're probably going to need about 500 petabytes over the next three years. And last week they came back to us and said, we're going to need an extra two exabytes on top of that 500 petabytes. And my expectation is that from what I'm seeing in other parts of the landscape, that they'll come back again in a quarter or two and say, we need a double digit number of exabytes.
48:11We undershot. And we're seeing that across the board, both in the smaller and medium sized environments, as well as in the big environments where they thought they would need tens of exabytes. So now they're talking in triple digits. And how much of it is big labs building versus end customers and end users? Well, the big labs are building because they have end customers and end users that are using it. And the AI clouds are building because they have both big labs and medium labs that are consuming it. And I don't think you'll find infrastructure that's just lying around and not being utilized.
48:52At the moment, in fact, the opposite is true. There is a lot more that people want to do that they cannot do. Some of these AI clouds have stopped selling more capacity just because they're sold out for the next year and a half. And so I think the limiting factor here is physical. It's land, it's power, it's chips. Building fabs takes a long time. And so as fast as we feel this is moving, and it is moving very, very fast, it's still the build outs are happening much more slowly than the demand requires from them. Will it sustain forever? I don't know. I'm sure over time there will be hiccups. but as far as I can tell we're just at the very very beginning of most of the world running on AI and if AI will really be able to do all of these things that we're starting to see it do then it should continue at least for the next five to ten years at this pace so we have a lot of work to do Do the circular deals and the whole like financing debt aspect of the ecosystem, does that make you nervous?
50:18Or do you think that's just like a means to whatever needs to happen for the ecosystem to be actually built? So I think going back to the beginning, the old stack and the new stack, I think there's a new ecosystem getting built around the new stack. And some of the older companies are joining this ecosystem and some have not yet joined this ecosystem. And so it's lacking in liquidity at some points in this supply chain. We see the build outs that need to happen in some points need money in order to finance them. And sometimes you need that money ahead of when the results appear and ahead of when you actually get returns on that investment.
51:08And I think the companies that are in the ecosystem are a natural spot to get that liquidity because they see what's happening. They believe in it. They understand how big this is going to be versus someone who's on the outside looking in that maybe feels this is riskier for them to finance. And that's, I think, what's happening. Speaking of finance, you're in a very interesting and somewhat unique position because based on the little ad, I think you guys are profitable. We are. Which is not the case of just like a lot of other players in the ecosystem. What does that tell us in terms of where risk in the ecosystem?
51:53Like you happen to be in this layer, which is indispensable, and therefore there's no other competitor. Therefore, you can be profitable. Are you more disciplined? What means what, I guess? It's probably a combination, but I think it goes towards our business model. We sell software and so our gross margins are high. And as we grow on the trajectory of AI growth, which is pretty consistently been about 10x every two years, we get more and more efficient and we generate more cash and we generate more profitability the companies that are higher up the stack have a lot of compute costs that they need to do and the companies that are underneath us in the stack are building hardware and so they too have a very different business model than than the one that we have but yeah i think the combination of fast growth and efficient growth is something that we've built into the way we work from the early days.
53:04And it goes back to VCs not necessarily wanting to invest in the space in the early days, so we needed to be self-sufficient. And it also goes back to VCs not wanting to invest in the space in the early days, so we have less competition than maybe exists in some of these other layers. You have a lot of customers in the NeoCloud ecosystem, and I'm curious about what you view might be on that ecosystem. I mean, there's literally hundreds of NeoClouds. What do you think separates the one that will be around in two to three years as dominant forces in that ecosystem versus the ones that will not? So they need access to power.
53:50They need access to hardware. They need access to money. And I think the ones that will succeed are the ones that build, that know how to build. I find the end users come, what NVIDIA calls off takers, and they want immediate access to compute. and if you have it, then they'll come to you and they're willing to pay. And if you don't have it and you tell them, let's build it together, we'll have it ready for you next year, then they'll go to somebody else. And so I've seen the ones that are most successful be the ones that build in anticipation of demand rather than behind the demand. Build physically or build capacity through rentals or?
54:34Everything. And you don't have to own all of the layers of the stack. There are multiple layers of the stack underneath us and the NeoClouds. But yeah, and I find the ones that have that combination of good relationship with the hardware vendors, access to power and land, and that are savvy in the way that they finance their operations, be the ones that are most successful. Yeah, and to the last point about money, one can debate whether that's still current or not, but the case against Neoclass for a long time was that, well, effectively, they're kind of like a financing vehicle for NVIDIA GPUs.
55:26The primary skill is finance, and eventually hyperscalers will beat them because hyperscalers will always have a lower cost of capital. is that in the three aspects that you mentioned, is financing skills as important as the rest? What do you make of all of this? I think that the reality has proven to be the opposite of what you just said. The fact that we have so many of these new AI clouds and the fact that the hyperscalers have not beaten them yet, I think it goes back to the innovator's dilemma. When you have something new to build and you have this old cash cow to focus on, it's a lot harder than when you have something new to build and you don't have that legacy.
56:16But yes, regardless of where it started, I think these AI clouds have built up a skill set and a specialty. And it's not the same to build this new stack and these AI factories versus to build the old stack and the way that the hyperscalers built clouds five and 10 years ago. And every day that passes, these AI clouds that are actually in the trenches doing the work, they're learning. And the neoclouds that are sitting on very cheap financing, but not yet, the hyperscalers that are sitting on very cheap financing, but are not yet doing this work, are not learning that. And so the gap in skills just keeps increasing.
56:59Now, I think the hyperscalers have realized that. And if two years ago they were saying, yeah, we have this, we have storage, we have compute, we have networking, we know what we're doing, we don't need to worry about this new thing. Today, they're singing a very, very different tune. And I think they're aware of the fact that their lunch is being eaten by someone else. And I think they're reacting to it. So it'll be very interesting to see how that dynamic evolves over the next couple of years. What do you make of sovereign AI? It seemed to me that the concept of AI factory that we discussed at the very beginning was largely around this concept of sovereign AI like a year or two ago.
57:44As we discussed at the beginning of this conversation, it seems to have expanded now to also include enterprises. But from your vantage point, what's the reality of that concept around the world as it applies to nations? I think this becomes a basic need over time in the same way that electricity is a basic need or that water is a basic need. AI intelligence will be a basic need. We will not be able to live in the way that we would like to live without it. And so in the world that we are in today, countries are adversarial one to the other and they don't trust each other. And in the business community, as much as we have competition, I think we collaborate very, very nicely.
58:32But nations like to have their own. And so every nation wants to make sure that the U.S. can't shut down its AI or that somebody else can't shut down its AI. And so they're trying to build sovereign AI solutions. Obviously, not every nation can build the software. Not every nation can build the hardware. Not every nation can build all the different layers of the stack. But I think they get comfort from knowing that it's within their borders, at least. And within their jurisdiction from a legal perspective. Zooming out a little bit from that whole stack that we described with the five layers, where do you think the value accrues over time?
59:20and which part of the stack commoditized? People have been talking about model commoditization for a very long time. Does that commoditize? What about the rest? I think the best way for me to think about this is to look to history. And historically, hardware tends to commoditize relatively quickly. We have, because it's standard, not because it's easy, but maybe it's easier to copy. I don't know. The software layer tends to accrue a lot of value. If you look at the most valuable companies in the world today, if you look at Amazon, if you look at Microsoft, if you look at Apple, if you look at Google, those are the companies that built those software layers of the cloud, of the PC era, of the mobile era, of the internet.
1:00:08And they are worth trillions of dollars as a consequence. The application layer, there tend to be a lot of applications, some of which will have a lot of value and others will not have that much value. But I think consumer loyalty is something that is also valuable in those previous revolutions. And so the software infrastructure layer tends to accrue a lot of value and then on a more selective basis at the application layer. Where the models fit into that is very, very difficult to say because we never had a model layer before. So that would be interesting to see. And you see the model companies have become application companies and the interesting semi-recent trend has been to see the application companies become model companies again, starting to do a lot more work, whether that's a cognition or a cursor or...
1:01:07So I'm not sure those two layers will stay separate over time. Perhaps they merge back into one. In the old stack, we had only four layers and the new stack perhaps that merges as well. I'm curious about how you think about the place of NVIDIA in the ecosystem. The name has come up already a bunch of times during this conversation and it's so central. They're an investor, they're a partner. How does one work with such a dominant player in the space in a way that preserves optionality and preserves your independence and self-reliance as a business long term? So interestingly, other than as an investor, there's no legal document between us and NVIDIA.
1:01:59They don't resell our solution. We don't actually do anything together from a legal perspective. Having said that, they are our best partner by far. And at any given point in time, we have a high teens number of projects that we're collaborating with them on. We have a triple digit number of developers working with them. They have the same on their side with us. Everything ranging from new networking gear to new inference microservices to this agentic confidential computing project to the next generation of GPUs and DPUs and how we take advantage of them. all of that we love because as we build this together then the solution ends up being better for our joint customers and so we really enjoy working with nvidia also from a cultural perspective they move fast they believe that anything that is physically possible is possible they believe that it can be done when others believe that it cannot.
1:03:14And so it's been a lot of fun working with them in parallel to it being also very lucrative because they are the ones that are creating these new environments and putting names on all of these new concepts. And so they're leading the charge and we're trying to, as much as we can, ride their coattails. Having said that, there's no type of exclusivity. We work with AMD, we work with other companies in parallel to working with NVIDIA. In practice, they do have the majority of the market and so the majority of the deployments that we have are with them. One thing that strikes me in a big part of this conversation is that you guys seem to be just relentlessly launching new products, expanding new customers, having key relationships.
1:04:09We mentioned NVIDIA, but XAI has been a very important customer for you guys. What have you learned in terms of moving at the speed of AI? Well, you can't move faster than Elon. I was fortunate to see some of the ways in which he pushes his team to move faster. and it's super simple. You find the limiting factor and you get rid of it and then you find the new limiting factor and you get rid of it and that cycle keeps accelerating you faster and faster and faster. And that's what we like. The reason we keep innovating and keep accelerating and keep demanding more from ourselves is A, because we're paranoid and we're afraid that somebody from somewhere will pop up and try to catch up to us but B, because we love it.
1:05:02We enjoy it. It's the joy of building new things and exploring and figuring stuff out that you didn't know how to do last week or last month. And our customers expect that from us. Our customers are the most demanding in the world. And they're all in a race. And they're racing each other in this quest for better AI. And they need us to never slow them down. We can never be the limiting factor in their progress and in achieving their success. And practically for the founders or leaders listening to this, how does one actually do that? Meaning finding bottlenecks and because the bigger the organization, the more bottlenecks they are.
1:05:49And they may be at the top of the organization or at the very bottom. As a CEO, how do you do that? You have to talk to people and ask them that question. they know as CEO you don't know but the people that are on the ground know and so you need to build an organization that's as flat as you can possibly build it such that you have direct access to everybody and you need to build an organization where people are not afraid of raising those problems there cannot be a chain of command because that slows things down I think I heard this from Elon, bad things should be stated loudly and often and good things once and softly.
1:06:35So we can't get into a mode where we're congratulating ourselves too much. We always have to be in this mode of somebody's going to kill us. We don't know who it is and we have to keep fixing everything so that they don't have a crink in the armor to get in through. All right. So maybe to close some kind of forward thinking, projecting ourselves in the future a little bit as possible as it can be today. But 10 years in from now, whatever, many years, looking back, what do you think that the industry believes today that may turn out to be wrong? Like just one thing that may come to mind. I think 10 years from now, again, assuming AI, and I think we're starting to see proof points, will not plateau at the level of human intelligence, but will continue beyond the level of human intelligence.
1:07:31And assuming it has physical access to all of the things that we have physical access to on this planet and in space, I think everything is different. Nothing is the same 10 years from now. We will not be needed for all of these tasks. And I think the way government is structured will be different. And I think the way money is transferred and ownership of things will be different. And the way we think about how we live our lives will be very different. And so it's impossible for me to imagine that world, except to say that it's probably going to be a lot more different to what we have today versus where we are now and where we were a thousand years ago.
1:08:28So I think in the next 10 years, we'll see more difference than we did in the last thousand years if this plays out in the way that it seems to be playing out. And in that world, and maybe that's not 10 years, that's just five years, in your wildest dreams or perhaps your very pragmatic vision, where does VAST sit? Becoming the operating system for AI, what does that actually mean if everything plays out? We want to, A, enable this so that it actually can happen and that we don't see obstacles in the form of we don't have the required infrastructure for it to happen because we can only build this fast and we want to build a thousand times faster or a thousand times bigger.
1:09:13We need to make sure that it doesn't kill us and so that it's safe and that along the path, we put safeguards in place and enable us as people to understand what's happening at the very least, if not guide what's happening. I tell my team all the time, if all the data is managed by us, we can't ask for anything more than that. And we're still a ways away from that. But that's the ideal that we're marching towards. All the data in the world. All right. Well, that feels like a wonderful place to live it, Renan. Thank you so much. This was a terrific conversation. Thank you. Hi, it's Matt Turk again.
1:09:56Thanks for listening to this episode of the Matt Podcast. If you enjoyed it, we'd be very grateful if you would consider subscribing if you haven't already or leaving a positive review or comment on whichever platform you're watching this or listening to this episode from. This really helps us build a podcast and get great guests. Thanks and see you at the next episode.
From the publisher
Everyone talks about GPUs. Almost nobody talks about the layer that feeds them. Renen Hallak is the founder & CEO of VAST Data — the $30 billion company powering xAI and some of the world's biggest AI clouds — and he sits in the hidden layer of the AI stack.
In this episode, we cover what an AI factory actually is, why every company will eventually own its own AI, the architecture bet behind VAST (DASE, explained simply), KV caches and agent memory, and DataEnclave — VAST's brand-new confidential AI announcement with NVIDIA that lets leading models run on the world's most sensitive data.
Plus: the demand signal that scares even him (a customer went from 500 petabytes to 2 exabytes), circular financing, which neoclouds survive, sovereign AI, working with NVIDIA and Elon Musk's xAI — and why the next 10 years will bring more change than the last 1,000.
(00:00) Intro
(00:51) The hidden software layer in NVIDIA's AI stack
(02:28) What actually makes an "AI factory"?
(05:13) Should Walmart and Goldman Sachs build their own AI?
(06:20) "We infer during the day, fine-tune at night"
(13:16) The announcement: models become a resource to manage
(15:32) From P vs. NP to founding VAST Data
(17:32) OpenAI, Navier–Stokes and 10,000 collaborating agents
(20:25) The pre-transformer insight behind VAST
(21:55) DASE: VAST's "shared everything" architecture explained
(25:18) "Storage was where startups go to die"
(27:39) Trillions of vectors: why old databases break
(29:01) Are S3, Snowflake and Databricks ready for AI?
(31:29) Data gravity, vendor lock-in and zero churn
(33:13) Training vs. inference: why the infrastructure changes
(34:46) Model routing, KV caches, RAG and agent memory
(36:59) Identity, permissions and security for AI agents
(40:27) Can multi-agent systems unlock scientific discovery?
(41:55) DataEnclave: how confidential AI protects data and weights
(45:16) Who should be AI's trust layer?
(46:37) "Sometimes it scares me": 500 petabytes to 2 exabytes
(50:08) Is circular AI financing creating systemic risk?
(51:32) Why VAST is profitable when AI infra isn't
(53:24) What separates the winning neoclouds?
(55:06) "Their lunch is being eaten": why hyperscalers lag
(59:10) Where will the trillions accrue across the AI stack?
(1:01:18) NVIDIA: "There's no legal document between us"
(1:03:59) What VAST learned from xAI and Elon Musk
(1:05:36) "Bad things loudly and often": building at AI speed
(1:06:53) More change in 10 years than the previous 1,000?
(1:08:39) VAST's endgame: all the data in the world
