#279 Matthew Carroll: Immuta’s Approach to Secure, Scalable Data Access in the Age of AI

14 Aug 2025 · 54 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Eye On A.I. Episode #279 Summary: Matthew Carroll - Immuta’s Approach to Secure, Scalable Data Access in the Age of AI

Podcast Overview

  • Title: Eye On A.I.
  • Host: Craig S. Smith, former New York Times correspondent
  • Focus: Discusses innovations and implications of artificial intelligence technology.

Episode Details

  • Episode Title: #279 Matthew Carroll: Immuta’s Approach to Secure, Scalable Data Access in the Age of AI
  • Episode Description: Matthew Carroll, CEO and co-founder of Immuta, explains the evolution of data access governance in the context of generative AI and emphasizes the need for automated compliance and secure access across various enterprises.

Key Themes and Discussions

Introduction to Immuta

  • Company Overview: Immuta focuses on automating compliance and data governance.
  • Market Need: Traditional security models are inadequate due to the increase in both human and non-human data consumers driven by AI technologies.

Evolving Data Governance Needs

  • Rise of Generative AI: The technological shift has expanded the number of data consumers dramatically, complicating security and governance.
  • Role Dynamics: Different roles in organizations require diverse data access levels, complicating the governance landscape.

Automation and Compliance

  • Automating Data Governance: Immuta uses AI to handle the complexities of data access and compliance automatically.
  • Policy Enforcement: Companies can dynamically enforce policies using a graphical user interface (GUI), allowing non-technical staff to manage data access efficiently.

Use Cases in Healthcare

  • Blinded Clinical Trials: Immuta’s platform allows pharmaceutical companies to conduct blinded studies necessary for FDA approval without revealing sensitive patient information.

Data Access and Governance Challenges

  • International Data Laws: Data sovereignty laws differ globally, affecting how organizations handle cross-border data access.
  • Negotiation Process: Immuta implements a marketplace for data governance, allowing consumers to understand their access rights and request additional permissions when necessary.

Technological Architecture of Immuta

  • Native vs. Proxy Integration: Emphasizes the importance of native integration within cloud environments for optimal performance.
  • AI in Scaling Governance: As the demand for data access rises (potentially into millions of requests), incorporating AI is vital for managing permissions at scale.

Future of Data Governance

  • Agentic Architectures: The future will see non-human identities (agents) playing significant roles in automating data access negotiations and compliance, leading to a shift away from deterministic regulations to risk-based approaches.
  • Corporate Data Strategies: Organizations will need robust tagging and auditing systems to manage corporate data usage effectively.

Market Landscape

  • Competitive Environment: Immuta holds a leading position in the niche market of data access governance, with adjacent sectors experiencing convergence due to the integration of data quality, tagging, and auditing.

Conclusion

  • Regulatory Evolution: Future regulations might focus more on data usage contracts rather than strict control over data access.
  • Anticipated Changes: As organizations increasingly rely on automated systems and AI for governance, the landscape of data management is expected to transform significantly over the next few years.

Key Takeaways

  • Data governance must evolve to address the complexities introduced by generative AI.
  • Automation and AI will be critical in scaling governance practices and managing the increasing number of data consumers.
  • Regulatory frameworks may shift towards a more flexible, risk-based approach in response to the rising significance of AI in data management.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00So think of Immuta as a platform that kind of orchestrates the whole provisioning process. We're getting tailwinds around this movement around AI, not because we're magically going to secure AI, but rather because the nexus of what created what we call our first era was this concept of separating policy from platform. They run these very large clinical trials and studies and for primary and secondary analysis. They will use Immuta now as a blinded solution, meaning they blind between the patient, the provider and the researcher. And we act as a blinding system. If that fails, or we can't prove that those blinds were in place, it does not pass through the stages of the FDA cycles.

0:44In business, they say you can have better, cheaper, or faster. But you only get to pick two. What if you could have all three at the same time? That's exactly what Coher, Thomson Reuters, and Specialized Bikes have. since they upgraded to the next generation of the cloud, Oracle Cloud Infrastructure. OCI is the blazing fast platform for your infrastructure, database, application development, and AI needs, where you can run any workload in a high availability, consistently high performance environment, and spend less than you would with other clouds. How is it faster? OCI's block storage gives you more operations per second.

1:35Cheaper? OCI costs up to 50 % less for compute, 70 % less for storage, and 80 % less for networking. Better? In test after test, OCI customers report lower latency and higher bandwidth versus other clouds. This is a cloud built for AI and all your biggest workloads. Right now, with zero commitment, try OCI for free. Head to oracle.com slash IonAI. IonAI, all run together, E-Y-E-O-N-A-I. That's oracle.com slash IonAI. Okay, so let's talk about Immuta. Everyone knows that data is the big proprietary piece for most enterprises when approaching generative AI. And I guess data security is increasingly important and complex because the attack surface seems to have spread with all these generative AI apps.

2:56Is that what you guys address? Yeah, I think data security is definitely one aspect of the lens. I think the challenge is there's different roles when it comes to access to data, right? There's definitely a security aspect, which is adversaries trying to steal your data for intellectual property, steal it to use it against you, right? You know, like many other things, but it's equally just there's regulatory risk, right? So especially like clinical trial management, as you can imagine, a pharma is like, how do we do blinded studies in the cloud, right? you know or if you go into a bank like in the cloud we converge all this data about how do we separate buy and sell sides so there's not conflict of interest so there's this kind of governance element of it so if you think of security you've got governance and you've got legal and so the problem is is I think everyone in the world kind of gets security but not everyone in the world really gets the governance and legal aspects of it right and I think as AI specifically generative AI has taken off, what it's done is lowered the barrier of entry for how many people want to get access to data, right?

4:12So that's really the interesting thing. It's not that the data going into the AI is maybe such a risk. It's actually a lot of the generative AI has allowed everyone from someone on a manufacturing floor to the CEO that underneath the covers can run queries and code against data. So you're going from hundreds of highly technical users to tens, if not hundreds of thousands of data consumers that want access to raw data. And so what used to be a security kind of very deterministic rule set for a very finite set of people is now a kind of a gray area that governance and legal kind of have to manage.

4:55So think of Amuta as a platform that kind of orchestrates the whole provisioning process for those consumers. And we're kind of hitting, we're getting tailwinds around this movement around AI, not because we're magically going to secure AI, but rather because there's a lot more people that want it. So we help the scale of making determinations of who can see what, when, where, why, how. Yeah. So it's a permissioning platform. Correct. Yeah. Yeah, that's how we started. We started, you know, we spun out of the U.S. intelligence community. And believe it or not, there's more lawyers than you could ever imagine in the intelligence community.

5:36The government has no shortage of that. And what kind of the nexus of what created our kind of what we call our first era was this concept of separating policy from platform. And so as a couple of intel agencies started investing in real cloud, right? So Amazon Web Services, you can look it up. It was called the Commercial Cloud Solutions, so C2S. They just threw data in S3 buckets. But the problem was it was the same data, but because it was two different, kind of think of it as business units, right? Same data, but each business unit, the way they collected it was different. So therefore, there's different rules as to who can see it, right?

6:17But a data scientist, they don't think like that, right? They're just like, I just want to merge the data together, get some insight, right? And when they went to do that, the whole policy broke down. Because the question you have to ask yourself is not just what is the right policy, it's who makes the determination when policies conflict, who gets to decide which policy wins, or is it a third policy? And it wasn't clear. There wasn't a business process to handle that. So if you think about what Amuta, when we first started, it was this, how do we build an abstraction tier? Rather than writing security logic and code inside a database of what that policy is, like mask this column, remove this, whatever, depending on this user, we built an abstraction layer where a lawyer or a data owner without knowing how to write code can say, basically do that, where they can just click, click, click and build policy on it rather than writing a memo.

7:13and it would enforce that policy as that end user is querying the data. So that's what it gave birth to us. So that's what we call the separation of policy from platform, and we then created, that's what Immuta started as, yeah. Right, okay. Can you give me a use case to illustrate that? I was following you until the click, click, click. Yeah, yeah. So let's get to the click, click, click. So it's really important when you understand like governance teams, I always say, woe is the data governor, especially in this age of AI. They're very capable people, 10 to 20 years experience in their domain.

7:56I'm going to take you through like a real world example in health care. But like these are potentially PhDs, potentially JDs. So they may have both. So they're very high IQ intellectuals that have worked in a domain for a very long time. And they're experts in their field and they understand the data very well. Right. So in life sciences, these might be eight to 12 of these people in a hundred thousand person pharmaceutical company. So there's a real customer, real world example. And they've got to manage clinical trial data. right which is the bread and butter of how they design medications that they want to take to market right and they have to go through a traditional FDA cycle all the approvals and they run these very large clinical trials and studies and for primary and secondary analysis they will use a muta now as a blinded solution meaning they blind between the patient the provider and the researcher.

9:01And we act as a blinding system. If that fails, we can't prove that those blinds were in place. It does not pass through the stages of FDA cycles. That's billions of dollars. That's it's very bad. And the way they managed this before was they would isolate each study as a separate database in like a SAS, SAS Institute. So like a SAS database. And then they would physically move it around. Now they collate it all into like a snowflake or database. So what they do is in Immuta is we can hook into an identity system, pull all the metadata about the users, who they are, what groups they belong to, what's their role, what attributes they have about them.

9:44Like they're an investigator on this study. We also pull all the metadata off the data, right you know what are these column tags how is it sent what's the sensitivity level which study is this part of right um what are the different types of patients in there we merge that together and based on that metadata rather than that governor the governors are like i said they they know the law they know the regulations they know all these things they use that metadata so rather than instructing someone to write code as to how to enforce these blinds, what they do is they go into our GUI and they can say, for these types of users, let's just say a principal investigator on this drug that they're building can see this type of data based on this metadata, right?

10:32And it'll dynamically filter. It'll dynamically mask. It could do things like canonization where we will nullify certain cells. If like, for instance, someone in the trial is seven foot six and is an outlier, right? It'll self-identify outliers and remove what we call indirect identifiers. So because like height is an indirect identifier, it's not important to the study, but you can figure out who that person is because there's only one person that tall. But the idea is they can go and argue and use the metadata about the the identity and the data to create a policy. And that policy could be three different types of policy.

11:14The first is a subscription. So who can even know about the data? The second is a data policy. How do we put masking and redaction and those types of techniques to remove utility out of the data? And then the third is a purpose-based. So are you asking a question of the data by which the reason you're asking is the same reason we collected the data? And so it's a purpose base. And that's important for like the GDPR and data sovereignty, et cetera. So that's what I mean when I say click, click, click. What I mean is there's literally a GUI to where they don't have to write code. They can be a lawyer, they can be a scientist, and they can build the controls without really having to understand any of the data infrastructure or having to write code or be that technical.

12:04That's what I mean by that. And to finish that point, what I would say is that's important because they can make it global, right? Meaning it's not just one study. It's not just one data source. But you can do this consistently across all of your data landscape, which is ultimately how you have to scale in converged architectures. What happens when someone, a user, runs into a roadblock somewhere, doesn't have permission, or certain data is hidden and they believe they have the right to see it and they need it? then can you go to the governor, the data governor who's managing this dashboard and have a discussion with him?

12:57Yeah, that's a great question because that's the pain, right? So as we move forward in our journey as a company, going from 2015 to like 2020, that's the core pain. The problem has always been they a lot of these consumers had to be extremely technical because the only way for them to understand that they they're not getting all the data and there's actually a policy in place. You would have to be technical. You'd have to write code and basically get a denied entry. And then you would have to reverse engineer. Well, who's the data owner? Who's you know? And you have to figure it out. Recently, we've built a data marketplace for that exact use case where what we now do is express the visibility.

13:51So you as a user can see what do I not have access to? And then we can then bring governance and legal into the same platform. and you can have a dialogue between one another and a discussion. And when you do request, it's policy mediation. So there's a policy exception process. And that policy exception process is a negotiation between the data consumer, governance, legal, and security. It's like, how much utility do you need out of this data? Like, do you really need this unmasked? or do you just want it unmasked? And there's a difference. And so prior to Immuta, like all these companies, they would just create a ticket in service now.

14:42And then IT would then make a determination and they make a copy. So there was no negotiation. There was no like, so you're spot on in the sense of the pain is, again, it was consumer to IT because it was all engineering. But now that the consumer isn't an engineer, in all cases, governance and legal need to have a conversation, a non-technical conversation. And so, yeah, we provide a mechanism for them to learn about what they don't have access to and then create that negotiation to be able to ask for more of it. Yeah, I would imagine that your business is growing because data is becoming increasingly complicated in where it lies and how it's connected.

15:37And presumably, Immuta, you can connect all different data sources to the platform. I mean, there's companies have, you know, everything from vector databases to S3 buckets you mentioned. How does all that get piped into the platform? Yes, it's really important. So there's kind of two elements you hit on that I think are architecturally really important. So the first paradigm is this concept of native versus proxy. So modern middleware is changing quite quickly. So the way we used to operate prior to the age of AI, and what I would say is the era of agents, the agentic architectures is where I'm concerned.

16:33And so when you think about high transactional volume, that basically, whether it's a human or machine, there's just lots and lots and lots of queries on data. performance is tantamount, right? It's kind of like, could you imagine sitting at a cashier waiting five minutes or even three minutes for your credit card to process? It would be awkward. So that's kind of the same feeling. Three minutes isn't a long time, but it is in the sense of high transactionality. So to act as modern middleware between cloud data infrastructure or cloud storage and the end consumer, you have to be native. So the proxy design models where you kind of act as a go-between classic SOA architecture where you have web services and so API to API to API, an app connects to you, you connect to the data store, that doesn't work.

17:26It breaks. So you have to natively integrate inside the compute infrastructure. So inside a Snowflake, inside a Databricks, inside a Redshift, inside S3, right? Like whether it's unstructured or structured, you have to be in the compute pattern or you have to create a compute pattern that can break that query out and scale it. Right. So that's the first part of scale in these. I call them agentic architectures because I think the AI is all leading towards these, you know, human human augmented, but agentic, you know, kind of operating environments. So that's like number one is so we, Amuta, you know, in my opinion, in our domain, kind of were, you know, like the first to go about that.

18:13Everyone else was trying to do a virtualization model, whereas with us, we are native inside and we were very fortunate to partner with the Databricks or Snowflake like early days, like 2018. So that's important for anyone listening. Native is the only way it works because it's the baseline of scale. The second issue, though, is what people don't think about, which is what happens when I have lots and lots and lots of data consumers? So it's not just humans, but also machines, potentially millions of identities, human or non-human identities requesting access to data, to data products, to full domains of data, et cetera.

18:59And that's where things also break. And so you have to be able, because the amount of tickets coming in, like the pharma I talked about, right now with just humans, they average 200 ,000 data requests a year. So that's 750 to 800 tickets per data governor per week. Now, let's just say you made that 2 million. So the problem is you have to add AI into this natively integrated permissioning system to be able to handle that scale. So that's how a Muta, when we talk about how we integrate into the data infrastructure, there's two aspects of that. The first aspect is how do you natively integrate into the compute so it can break and it's performing just like you ran a query directly against it.

19:57The second, though, is how do you manage that volume of tickets? And you have to provide direct access at the speed of the query, which means you need AI to automate those decisions. And that's what Amute is doing. Yeah.

20:15How does this apply to sharing across international borders? I mean, there are a lot of companies and organizations that are global. Yeah. Can you talk about the governance of that kind of data sharing? Yeah, absolutely. So there's two types of international data law, if you will. So there's data sovereignty rights. And so people, I think sometimes just assume when you talk about data sovereignty, that's all kind of, you know, one kind of big bucket, just like one use case. It's actually not right. So like the laws in China, for instance, and data exfiltration out of China, the People's Republic of China sees data that is produced there as government controlled.

21:10So data is not going to leave there. You can extract certain metadata out of there, but raw data needs to stay in the sovereignty of China. Same thing in like we see a similar type of legislation in Brazil. the the eu with you know there's a lack of treaty between you know the eu and now u.s um you know a lot of people think of gdpr right but that's really not data sovereignty right it's just that's more about um a series of permissions around like use so why did you collect it why are you processing it and there's ways to utilize uh statistical de-identification so you can anonymize data to then export it, right?

21:51And so the question isn't, the tech isn't the hard part. The hard part is actually applying policies consistently as a global company to where is this person, where did the data originate, and who wants to access it, right? Who are they? And you have to make real-time policy decisions on that. And so that means you need to categorize your data appropriately based off of geo zones. And you need to establish purposes on that data as to certain purposes need to enact different types of policy. Because like a purpose is like in Europe, great. We just need to anonymize that raw data, which we can do in flight.

22:36In the case of China, no, we need to run that locally. That becomes really important in analytics. So if you're adding AI, if you're training a model, let's just say I'm a manufacturer as a customer of ours, and they obviously collect a lot of telemetry off vehicles in China. That model needs to run and be housed, trained, validated and run in China on Chinese infrastructure. Now, in Europe, that's not necessarily the case. We just need to anonymize certain aspects of it. And so that model can be running in, let's just say, California, as well as UK, as well as Germany. Right. So but again, what Immute is doing as part of that is we're taking the metadata.

23:25So the data like the metadata on the data. So where is this data located? Where did it originate from? And we can look at the metadata like the attributes of the user. So the metadata of the user, and we can make a real-time determination based on that policy, what that data needs to look like depending on the use case. And we can change that dynamically and change the policy dynamically. And what's really interesting is, and this is where audit becomes really important, is your audit records need to tell you that it was, you know, the data was run, processed, and queried out of China and stayed in China.

24:00But in Europe, we did this. In America, we did this. In Brazil, we did this. Each of the compute systems, though, don't share audit records. And your seams aren't very good at this. So what Immuta does is we unify that audit across who access what data, when, where, why. And we can express that into a report in our audit logs. So we kind of help the data infrastructure layer, and we help those governance teams simplify not just the policy, like I said, but also the control and that policy, guaranteeing you met the regulation, if that makes sense. Yeah.

24:42And how you talked about scale, and we're talking about international now, what are the limits to scaling? Is there, I mean, if you have multiple jurisdictions that you're operating in and multiple data sources and a huge community of users in different markets. How do you scale that? Is the platform elastic enough that it can just scale infinitely Or is there some bound at which you've got to set up a new account and split? Yeah, no, no, no. That's actually kind of the interesting thing about this. While sovereignty legislation is complicated, the reality is the tech is not the limiting factor, actually.

25:51This is why you have to have a native paradigm. The performance and compute has just gotten so good that the tech side of the house is actually not the limiting barrier. It can scale at cloud scale. The problem always, this is, again, we go back to woe is the data governor. And no one's title is data governor. But there's so few of these. You can't scale them. We have an exponential growth in those who could potentially consume data. And we have a linear potential growth, but even that would be cost and effective. Like when you're talking to someone who needs to have 10 to 20 years of experience in data governance and working within your domain, it takes them at least a couple of years to ramp in that corporate.

26:44Because these are big companies, right? And understand how they work and how you make decisions and the risk tolerance of security and legal. And then the CEO may change out and you get reorgs. Like it takes a long time to ramp in these companies. Like a Roche is going to take you five years to really be good as a data governor. Right. And they're one of the best, in my opinion, at these data mesh architectures and establishing data product managers. And so it's like really difficult to scale those people. And they, unfortunately, are at the center of all this. And you can't make access decisions and you can't make sovereignty decisions and you can't make security decisions without the governor approving it.

27:36And so when you've got 10 of them for every 100 ,000 employees, that's a problem. And so the only way to solve that scale problem is to build AI for them and to have agents that act on their behalf based off of their pattern of how they make decisions. And that's the only way forward here. And is there an agentic layer to your platform? That's what's coming. That's what we're building. But the problem that here's there's no magical solution. There's no horizontal governor. Every company is going to have a different risk tolerance. Right. Like talk. Let's like like talk about iconic brands like what, you know, like ignore the like like customers or not.

28:33But like a Disney is going to be very cautious about like their brand. And they're not going to just accept, I don't know, the risk tolerance of, let's just say, maybe like a company like, you know, a high flying tech company like, I don't know, Facebook. Right. I don't know. But they want to govern their data and they want to regulate themselves their way. So you can't collate all of the access decisions and negotiations of how they got to those decisions and build one horizontal system. So you have to create the mechanism by which each company can kind of create one mind of how their governance team works and makes decisions.

29:30And so we are building agents for governors. But what we are doing first is creating the engine to collect all that metadata, to understand this one collective mind of data governors in these companies. And we're doing that through a policy exception workflow. As I said, someone requests, but they don't have everything they need. So they then have to negotiate. And through that negotiation of literally chat back and forth, like all this natural language, and then we can collect the metadata on how they got to the decision and the metadata on the data and the metadata of the identity who requested it and the metadata of the identity who approved it.

30:13And then did legal have to get involved? Did security have to get involved? We mashed that all up. And what Gen AI is really good at is making sense of that and creating a pattern out of that. And so we need to do, what we are doing is forcing companies rather than creating lots and lots of little like edge case policies, like we always have, we need to actually create global policies that everyone has a problem with and then force those exceptions. Yeah. Because that is the way we create the data that creates the AI to allow these governance teams to scale. Yeah, that's interesting. The platform, how does somebody use it?

30:59Is it a subscription SaaS product? Yeah, yeah. It's, yeah, classic SaaS, like truly SaaS. So we have, we can also fulfill via self-managed on their cloud infrastructure, but classic web application, you log in, SaaS subscription. We're not cheap. This isn't a$15 per month per user kind of platform. But so we're kind of like the Apple, if you will, of, you know, data provisioning and data access management. And it's build is based on usage or? Users. So we have two different types of users, very similar to like the business intelligence construct. So, you know, in Tableau, you have those who can create dashboards and then those who consume from dashboards.

31:52In our world, we have those who can create policy and then those who can consume data via our protected data. Right. Right. So that's kind of how we we view the world. And so we have our governors, our data governors, which is like data stewards and those who can build policy. And we have data consumers, which, you know, some of our customers, hundreds of thousands. Right. And so, yeah. Yeah. Are there areas? I mean, where do you see data governance and data availability and data generation going? I mean, it's just, you know, people talked about big data 20 years ago, but this is another scale and another level of complexity.

32:44What's happening now with generative AI is, are tools or platforms like Commuta keeping up with the complexity or do you think there's going to be a further layer of abstraction, you know, I don't know, maybe at the government level or across regions or something where people are going to want to know where data lies and where it's going and how sensitive it is and all of that? Yeah, I think you're on to something. I mean, I think it's going to take quite some time. I mean, I was at the center of the big data movement. You know, like when I first started thinking about big data, it was just as like InQTel just had invested in Cloudera.

Read the full transcript

33:43You know, we, you know, I was in Baghdad at the time. And, you know, our biggest issue was we had all sorts of assets that were collecting data volumes we couldn't even comprehend. um because a lot of it was video right so it was the first time we really had like drone operations at massive scale and so you can i mean even today you take a video it's you know they become big real quick luckily they you know internet's gotten really fast um but you know that was a big we call it ped like um you know we couldn't we just ran out of storage and then it was like how do you process it. And so those are petabyte days.

34:23That was 2007. So now it's like you fast forward to today, it's at a totally different scale. Concepts like Hadoop have obviously morphed into, you know, Spark and, you know, the type of distribution on cloud is just a totally different game. I think what people see is one storage is, you know, kind of coming out of that movement today. hey, storage is essentially free, right? That's the way they look at it. So I think, but compute isn't. And I think that's what large corporations, which are leading the charge, right? Government organizations are not leading this charge. It's corporations are driving this innovation.

35:03And I think the compute isn't cheap. And I think they're trying to rationalize because there's a movement around AI, they're willing to expend, they're willing to have high CapEx investments investments um so whether it's buying pre-buying infrastructure buying nvidia chips etc etc right everyone's willing to hit the capex but the question is the roi but but so you got to separate what i'm kind of getting to the point here is is there's a limit to which people are willing to invest and that's coming up pretty soon here now what i do think is happening with ai is there is roi so then the question is is okay where are we getting roi and where are we not getting roi So I think that's the first kind of iteration of what you're going to see.

35:50Now, when they do understand where the ROI is, they're going to double, triple, maybe even 10x down on that, right? Because we all understand the cost savings and the value of AI around these kind of core data sets. That is where I think there will be regulatory oversight, because I think sovereign nations are going to say. It's not just how we process that data that is like a strategic advantage to our country, but it's the data itself becomes a strategic advantage. And but I don't think the government will want to get into the business of owning that data. but I do think they'll want to get into the regulatory oversight of it, reporting around its use, reporting around how transitory it is, and establishing treaties will likely be something I see happen in the next five to 10 years.

36:40What I will also say, though, is this. In between that timeframe of the next five to 10 years and this type of what I would call government oversight of national security use of data for AI, What I will say is more third parties are going to need to understand corporate access to data because I think what's happening is all of our data, it's not to sell it. It's because our economies are so intrinsically tied to data. Right. Economists are like just desire, you know, like right now, obviously, if you're falling inflation and how like the stagflation. problems. They need access to more data faster to understand it.

37:28Our central banks want it. Manufacturers want it with supply chain, like looking at these tariffs. How do we calculate the supply chain issues? Where's our risk? Everyone needs each other's data because it's in the best interest. Like GM working with Stellantis, working with Ford, while they all compete, it may be in their best interest to understand, you know, steel prices and how those things. So companies need to be able to tag their data and be able to expose their data to non-traditional third parties, which means contractual management of that data is going to be critical. So if I could kind of summarize that, it's, I think one is purpose and all data is going to be really important because purpose defines how anyone could access it.

38:17Two is you You need to have really good tagging of the data to understand how to build policy for third-party contract management because it's going to be corporation to corporation. And the last is the auditing of use of that data and where it goes in your AI. It's going to be really important for regulators to be able to understand. And I think all that's going to happen in the next five years. And the auditing, that's auditing, not monitoring. Correct. At the government level. Correct. I don't see now monitoring is definitely necessary. The zero trust concepts around data provisioning and access are really important.

38:56You have to assume there's insider threats. There's adversaries in this ecosystem because the data is the intellectual property rights into AI. The AI is just a foundation model, right? Like an open AI, et cetera. But the data you put into it is what's really key to you. And the data that comes out of it and you store that and make decisions off it is also equally important. So you need to monitor and you need to have SLAs and SLOs around that monitoring so you can have alerting and everything. But I fundamentally do not believe the government or at least the U.S. government, I should say. I think foreign nations, I think, will be different.

39:36Obviously, China has a much different view of how they insert themselves into their economy. But like so you can look up something called Kalea, old legislation around call device registry. So how does NSA potentially hook in via a warrant legally or law enforcement agencies into, you know, telephone networks? Right. You know, we've all watched the movies. Right. You know, I'll get in tap someone. Right. Um, the legal oversight of something that is relatively limited infrastructure, like AT &T, Verizon, like, you know, there's only so many providers. Right. And that is extremely complicated. Right.

40:19Technically and legally. To do that across all corporations in the United States, I just, there's, we don't have enough budget and time and technical know-how to handle that, I think, in my opinion, to monitor. but to get audit to regulate yeah i can see that yeah and where do you guys stand in the market uh i mean frankly i've you know um it's a market that i'm not familiar with are how many players are there in this this particular market the the governance platform market i'm not sure how you define it yeah it's a broader market than you realize um we've kind of built a niche role because we've really focused on data access governments and this whole provisioning concept um so that's a very you know small group of players right there's like three of us um and we are by far the largest um when you get in but but reality what's happening is our adjacent areas are starting to converge.

41:32So when you think of governance, you have to think of data quality. You have to think of cataloging. You have to start thinking of the data tagging aspects of it and then the audit. So data access governance is just one part of it. The whole piece of it, though, is governance, governance ultimately to provision who can use what data, how, and why. So I think these adjacencies are starting to come together and they each kind of have a piece of the TAM and we're all going to likely converge into a couple of big players. And it was just we all kind of were born out of the problem was you couldn't create one super platform because none of us, even though you used to be able to raise a lot, a lot of money.

42:17Now you certainly can't. But the reason we all were able to create sizable enough companies is because we all had to reinvent an existing area, but make it cloud scale. Because Snowflakes and Databricks grew so fast. That type of scale, you couldn't try to build everything and make it scale like that. So you had to pick your poison. But like I said, in our domain of data access governance and this whole concept of data provisioning, very niche, but I think growing really fast. And hopefully that affords us the opportunity to kind of partner in some areas and then, you know, expand into other adjacent areas.

42:56Yeah. So you mentioned storage is almost free. That's something that I don't really understand. And a couple of years ago, I was trying to track, like, what happens to data as it gets older. and some of it literally ends up on tape decks in old mines, like under the earth or something. Yeah, I mean, how could the storage be free? I mean, is it because of compression technologies are getting better? I mean, there's just so much data. Yeah, I mean, first of all, there's no such thing as a free launch. So I think it's all in proportionality, right? I mean, so when you – the problem is, of course, it's not free.

43:55But this generation of engineers that have entered the workplace, they see it as a negligible cost unit relative to everything else. so i think to your point that it's a very very valid point which is i think those have been in this business a very long time are going to start to have to think about how do we yeah you're going to see this spike probably in a couple of years where all of a sudden you know your retention requirements let's say you're required to retain for 10 years you know seven years ago and beyond it's like not that big of a deal but all of a sudden those numbers start going up and up and up.

44:34I think one is, yeah, compression technologies. I do think, I mean, people joke about this, but I don't think it's a joke. Yeah. I do think some of these legacy mainframes and tapes and like legacy warehousing systems, there might be a cost and I'll say, hey, that's good enough. Just leave it there. Just do that. Right. We don't need this in the cloud. It's, it's, you know, but I also think the store, so that's one part of, of storage. I think the second part of storage though, is storage is becoming just as important as the compute because I think AI is especially, so, so there's AI, which is one just directly pulling from unstructured data is highly valuable and there's real use cases now.

45:20It's not just like deep learning where those tend to be edge cases in the big scheme of things um but the second aspect is i don't anyone i'm happy anyone listen here we can go have a beer and have a debate but like federation technologies are really complicated for any really world use case most of the time when you're federating it's just to go reach out pull it and put the data somewhere else So concepts like iceberg are really important. I mean, that's why, you know, Tabular was bought by an absurd amount of money, you know, like equity from Databricks. It was like, you want to be able to lift and shift, you know, your warehousing and your lakes via storage.

46:02Right. And so the idea in a perfect world is this paradigm where I can rip and replace any compute infrastructure because my iceberg, they're all compatible with iceberg or some sort of data format at the storage layer. And that helps me lift and shift, right? And so therefore, I get economy of scale because I have the ability to force those compute vendors to compete and it brings my compute costs down. So ergo, cost savings there are worth the spending on the storage, which is a fraction of the cost, right? So I think that's the paradigm I'm seeing right now. But to your original point, yes, compression is going to get way better.

46:47I think storage is going to get way better. I think there's going to be a huge advancement in the next couple of years. Yeah. And just to unpack some of what you just said, because I wasn't following all the acronyms and things. Yeah. And I'm not going to remember the first one that I missed, but iceberg, what do you mean iceberg? So Apache Iceberg, so that is a storage technology that emerged over the past couple of years where what we're seeing is the community like the Snowflakes and the Databricks and the Starburst like Slash Trino of the world. They're compatible with that storage. And so the idea being is, is like, if I have data in there, lands in there, and there's like a file format to that, I could theoretically have a dump it in there.

47:39Snowflake and Databricks could be in my infrastructure. And they both know how to read from that and then build their tables out of that, let's just say. So if I, let's just say, Snowflake, I'm just making this up, like Snowflake is way cheaper than Databricks. I, as a CIO, could say, you know what? Move it over here. just keep doing what we're doing on our landing zone and i just focus there or i could say you know what i need to rip snowflake out for the next cool thing drop it in there and it just works right so there's um the the migration costs which typically are at the brunt of your your upfront investments um are now negligible in theory right this is all theoretical we don't see this at scale yet but yeah yeah yeah okay well um is there anything we haven't talked about that you think listeners should hear i would just say the like i can finish on just like one big point which is agentic architectures are going to be adopted at a faster rate than any other kind of like era of AI, if you will.

48:45I think non-human identities are very much a game changer for large scale corporations. I actually don't think there's much of a game changer for small companies and startups, even though I think VCs are saying you can create, you know, billion dollar revenue companies on 20 people. I don't believe that. I don't think there's a bunch of agents that are just magically going to be your customer success. But for large corporations, these agents that can be working in the background and doing things for people will make them far more efficient. They can work 24-7. They can be finding things that are issues.

49:29Ergo, what I would say to everyone is the number of data consumers that are going to be non-human is going to go up exponentially. Yeah. Therefore, and this is, I guess, highly academic.

49:46We always think of everything in set of rules. Like because we're a law and order society. And so everything's deterministic of how I can. Can I use someone's HIPAA data, like HIPAA control data or not? And there's a set of rules. We have law. We have legislation that defines that. I think the world's going to go to deregulation because these agents are going to be so important that we're going to take a lot more risk. So everything's going to be non-deterministic. It's going to be risk-based approach to everything. Who I share, whose data with, how I do that, when I do that, why I do that. So I think as regulation goes out the window, because these agents are so important to these companies and to governments, and their economies are dependent on it.

50:40I think regulation goes out, and I think how people get informed as to how their data is going to be used is going to become a big deal. So non-deterministic access rights, AI-driven access determination and sharing are going to become the thing of the next couple of years. and then the government is going to legislate. It's less about the rules of how your data is being used. It's more about a contract, and you'll be notified about how it's being used, just like any other contract that you have. Yeah, that's fascinating. So agents will be asking other agents for data, and there will be some sort of a negotiation over.

51:26Yeah, based off of perceived risk tolerance, And they'll have to be like a circuit breaker in there of some sort to say, yeah, this thing's broke. All right, let's insert a human to figure out what the heck's going on. But the deed's already done. So you'll just, it'll kind of just be like, you'll then inform. It's like when you get a letter in the mail, it's like, hey, your data was part of some breach. You need, you have questions, call this number, right? It'll send you to some chatbot and it'll be agentic AI in there and you're talking to it, right? But that's the world we're going. Expect, quote unquote, lots of these kind of breach notifications on your data relative to an agent.

52:04Talk to another agent and got some of it on you. In business, they say you can have better, cheaper or faster. But you only get to pick two. What if you could have all three at the same time? That's exactly what Cohare, Thomson Reuters and Specialized Bikes have. since they upgraded to the next generation of the cloud, Oracle Cloud Infrastructure. OCI is the blazing fast platform for your infrastructure, database, application development, and AI needs, where you can run any workload in a high availability, consistently high performance environment, and spend less than you would with other clouds.

52:51How is it faster? OCI's block storage gives you more operations per second. Cheaper? OCI costs up to 50 % less for compute, 70 % less for storage, and 80 % less for networking. Better? In test after test, OCI customers report lower latency and higher bandwidth versus other clouds. This is a cloud built for AI and all your biggest workloads. Right now, with zero commitment, try OCI for free. Head to oracle.com slash IonAI. IonAI, all run together, E-Y-E-O-N-A-I. That's oracle.com slash IonAI.

From the publisher

Try OCI for free at http://oracle.com/eyeonai

This episode is sponsored by Oracle. OCI is the next-generation cloud designed for every workload – where you can run any application, including any AI projects, faster and more securely for less. On average, OCI costs 50% less for compute, 70% less for storage, and 80% less for networking. Join Modal, Skydance Animation, and today’s innovative AI tech companies who upgraded to OCI…and saved. 


Matthew Carroll, CEO and co-founder of Immuta, joins the Eye on AI podcast to explore how data access governance is evolving in the age of generative AI. 

As AI drives a surge in both human and non-human data consumers, traditional security models are no longer enough. Matthew explains how Immuta automates compliance, enforces policies across global jurisdictions, and scales secure access for enterprises—from blinded clinical trials in pharma to cross-border data sharing.

 Discover why AI-powered governance agents, risk-based access controls, and native cloud integration are the future of compliant, scalable data use.


Stay Updated:
Craig Smith on X: https://x.com/craigss
Eye on A.I. on X: https://x.com/EyeOn_AI


(00:00) How Immuta Governs Data in the AI Era  
(02:59) Security Alone Isn’t Enough in AI  
(05:24) From U.S. Intel Roots to Immuta’s Platform  
(07:38) Pharma Trials: Protecting Billions in Drug Development  
(12:58) The Data Marketplace & Access Negotiations  
(16:07) Why Native Integration Beats Proxy Models  
(20:33) Managing Data Across Borders & Sovereignty Laws  
(25:33) The Data Governor Bottleneck (and AI Solution)  
(31:05) Inside Immuta’s SaaS Model & User Roles  
(33:26) Big Data Lessons Driving AI ROI  
(41:05) The Growing Data Governance Market  
(45:48) Storage Portability & Apache Iceberg Explained  
(48:32) Future Shift to Risk-Based, AI-Driven Access

More from Eye On A.I.

All 266 episodes
#279 Matthew Carroll: Immuta’s Approach to Secure, Scalable Data Access in the Age of AIEye On A.I. · 54 min
Listen in VO