In short
Agentic data management for enterprises—how AI agents can automate data cataloging, data quality, and self-healing data pipelines across hybrid/multi-cloud “data sprawl,” using Excel Data’s X-Lake reasoning engine and ADM (agentic data management).
Guest backgrounds
Ashwin Rajeeva is co-founder and CTO of Excel Data (Accel Data), a startup founded in 2019. He previously worked at Hortonworks (Hadoop/CloudEra ecosystem). Excel Data has raised over $100M and works with Fortune 500 customers in banking, insurance, and telco. He’s based in the Bay Area part-time and works from India the rest of the year; teams also exist in Canada and London.
Key claims
LLMs need enterprise context to avoid hallucinations; X-Lake provides that context by connecting multiple data lakes (on-prem, AWS, Azure, Snowflake) at petabyte scale. Agents can spin up specialized “data quality” agents (e.g., a data catalog agent) and remediate pipeline errors by cloning code, rewriting Spark/SQL, and redeploying. Human-in-the-loop approval remains required.
Notable examples
An agent detecting anomalous pipeline logs, using metadata (tables/columns/types/owners) to rewrite Spark code, and redeploying to the cluster; catalog agent answering where a “sales order” table lives and its schema details.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOIntroduction to Self-Fixing Data Pipelines
0:00 to 0:11
Explore the potential of self-repairing data pipelines using AI.
“What if your data pipelines could fix themselves, not in some distant future, but right now, detecting errors, rewriting code, and redeploying without waking up an engineer at 3 a.m.?”
Accel Data's Journey and Global Presence
0:45 to 3:20
Ashwin discusses the history of Accel Data and its global operations.
“Ashwin, welcome to the Super Data Science podcast.”
Understanding Agentic Data Management and X-Lake
3:20 to 7:40
An in-depth look at agentic data management and the X-Lake reasoning engine.
“Let's talk a bit more about what you're up to there and the products that you're developing.”
Enterprise Data Agents and Their Applications
7:40 to 11:23
Ashwin outlines four types of enterprise data agents and their roles.
“That was a great explanation of the problems that people are facing.”
Challenges and Solutions in Data Management
11:23 to 14:00
Discussion on the complexities of managing petabytes of enterprise data.
“So it's going to do all this research, hit the Bing API or the Google API and get information back.”
Understanding AI's Interaction with Data
14:00 to 15:09
Learn how AI models handle data and the importance of context in queries.
“this the model no matter which model you use whether you use an open source one or you use the state-of-the-art foundation model right it can only access the information which is compressed in its neural net, right?”
The Mechanism of AI Systems in Data Pipelines
15:39 to 20:02
Explore how AI can optimize data pipelines and prevent errors with automated fixes.
“Obviously, without spilling too much proprietary sauce for our audience, but how does that work in a way that prevents hallucinations or divergent interpretations from cascading through the system?”
Human Oversight in Autonomous Systems
20:02 to 23:56
Understand the balance between AI autonomy and human oversight in data management.
“is how do you provide, number one, AI context about where's the code, what is the pipeline, what is the metadata, and what are the errors which are happening?”
Building a Data Management Platform from Scratch
23:56 to 26:52
Learn about the development and architecture of a robust data management platform.
“but you do most of the work without supervision, but there's somebody at the gate all the time.”
Navigating Data Sprawl in Enterprises
26:52 to 28:00
Discover what data sprawl is and its implications for enterprise data management.
“This is open source compute, up to date with upstream vulnerabilities taken care of, and it all integrates with the accelerator platform.”
Show all 17 chapters
Understanding Data Sprawl in Enterprises
28:00 to 31:46
Learn about data sprawl, its challenges, and the complexities it introduces in enterprise data management.
“So let's go back a little bit to the way that you have built this platform, ADP, to span petabyte scale, complex, hybrid, multi-cloud environments.”
Leadership Insights for Technology Professionals
31:46 to 38:41
Discover key insights on leadership in technology, focusing on customer-centric approaches and fostering innovation within teams.
“And so I'd like to shift gears here a little bit now to kind of maybe some more generic insights that you might have for our listeners.”
The Importance of Passion in Tech Roles
38:41 to 42:00
Explore how to identify passionate candidates for technical roles and the significance of meaningful work in tech careers.
“It's funny to me to think that like somebody who gets a hundred million dollars signing bonus at Meta is just like, oh, man.”
Interview on Passion in Data Management
42:00 to 47:59
Learn how to identify passionate candidates for technical roles.
“I think if we can provide an environment like that, then new ideas come in.”
AI in Enterprise Decision-making
48:00 to 49:22
Understand how AI is changing decision cycles in enterprises.
“but they've started using it internally.”
Embracing AI to Avoid Being Left Behind
49:23 to 49:49
Discover the importance of adopting AI to stay competitive.
“or it is people who are adopting, you know, AI in enterprises, they will realize that unless they do this, they are going to be left behind.”
Book Recommendation and Social Media Presence
49:50 to 51:48
Hear Ashwin's book recommendation and discuss his social media strategy.
“I've really enjoyed listening to you speak.”
Transcript
Automatic transcript. May contain errors.0:00Jon Krohn:What if your data pipelines could fix themselves, not in some distant future, but right now, detecting errors, rewriting code, and redeploying without waking up an engineer at 3 a.m.? Welcome to the Super Data Science Podcast. I'm your host, Jon Krohn. I'm joined today on the show by Ashwin Rajeeva, co-founder and CTO of Excel Data, a startup that's raised over$100 million in venture capital to bring you the agentic data management platform. Ashwin's outstanding communication on this meaty technical topic, involving autonomous scouring over petabytes of enterprise data, makes for a tremendous episode.
0:35Jon Krohn:Enjoy. This episode of Super Data Science is made possible by Dell, Intel, Fabi, and Cisco.
0:45Jon Krohn:Ashwin, welcome to the Super Data Science podcast. It's great to have you on the show. Where are you calling in from today? I'm in the Bay Area. I usually work out of India, but I'm here, you know, four months in a year. And thank you, John, for calling me to the podcast. Happy to be here. Yeah, yeah. So the company you work at, Accel Data, you guys are a big deal. Correct me if I'm wrong on this, but I was looking this up before the episode. It looks like you guys have raised over$100 million in venture capital already. Yes, we have. So we've been around since 2019, founded in 2019. And over the years, we have been fortunate enough to work with some of the best investors, raised over$100 million over three rounds, and also worked with, you know, opportunity to work with some of the biggest Fortune 500 kind of customers.
1:33So yeah, life's been good.
1:35Jon Krohn:Fantastic. Congratulations on all of that early success. And we're going to talk about the Excel data product, obviously. But something I wanted to touch on really quickly before we got going that's kind of interesting is it seems like Excel data is more or less headquartered in the Bay Area, but you mentioned how, you know, you work from the Bay Area four months of the year. And then, you know, most of the rest of the year, I guess you're based out of India. If our research is correct, it looks like maybe three out of your four co-founders are based in India most of the time. Yes. So, you know, all four of us are from a technical background.
2:08So we used to, you know, work at this company called Hortonworks. Hortonworks, you know, may use me, of course, the Hadoop platform along with CloudEra. And, you know, once we started, it, the whole idea was that, hey, let's just put something together quickly, start working with customers, because we've been working in data for such a long time that we kind of know whom do we need to sell to in some sense. So the three of us, of course, started off there in India, and we kind of built the whole tech team and platform over there. And Rohit, who's our CEO, essentially moved to the Bay Area and somewhere in 2020 to set up the GTM and sales and all of it.
2:46And most of our engineering comes from that background of big data specialists and, you know, engineers, data engineers. And so we found that it's the best place to kind of set up an engineer on. But we have teams in the Bay Area as well. We have, you know, some engineers here. We have some engineers in Canada to be close to the customers in the East Coast. And we are setting up a presence in the EMEA region as well because we have a small team in London where we're setting it up.
3:15Jon Krohn:Very nice. Nice. It sounds like you guys have things figured out all over the world. Let's talk a bit more about what you're up to there and the products that you're developing. So as a CTO and co-founder of Excel Data, you've been driving the shift toward agentic data management, which I think you guys call ADM for short. And it seems like a key part of that is something that you call the X-Lake reasoning engine. Do you want to tell us about X-Lake and agentic data management in general? Yeah, absolutely. So I think before we get into what agentic data management is and what this whole X-Lake theory is, I think it's a little useful to talk about where we're coming from, right?
3:56So we work with some, you know, the bigger banks, insurance companies, the telco companies. And usually what's ended up happening is that this field of data management and components in it. So if you have a data catalog, MDM, curation, any of it, any of this stack, the practices which have been around have been now around for 15 years. Most of the companies have been around for that period of time. I'm not talking about the compute engines because they come and go. But let's say the field of metadata and metadata management is essentially now the same for the last 15, 20 years with incremental improvements.
4:38And what we thought is, you know, as AI becomes more and more capable, we found that I think a lot of work, which is manual and, you know, kind of businessy, because in the end, a lot of business software is a lot of forms and a lot of clicking and a lot of, you know, working with interfaces with very well-defined workflows. And we feel that with the advent of AI, a lot of that can essentially be automated, right? And so there are two parts to it. One is the ADM part of it, which is how do you make data management agentic in the sense, let's say you are a data steward and you are doing five things regularly in your day job.
5:23How do you essentially make agents to it? That's where the industry is heading. Whether it's sales or marketing, everybody's trying to do this. So we are trying to apply that to data management, right? So can a data lake essentially classify the data itself, for example, right? Can you create validation rules automatically without having to spend an inordinate amount of time on some of these things? And what we also realized is that to power something like this, you actually need an engine. So a simple example is that when you use an AI, it knows everything about the world because in its brain there is a compression of facts and almost the whole internet but it knows very little about your data so if you ask it a simple question adm a simple question saying okay look at my data and you know let's say create a report about about some you know business topic that you have it really has no way for it to go and figure out what's in your environment it only knows about what's the internal knowledge it has.
6:26And so what we've done is created a cross-lake compute engine, which essentially provides the tools to the AI layer, right? Which it can use to connect to your internal enterprise systems. And we call it an X-Lake reasoning engine because the way our architecture is built is that you can actually connect all your data sources or data lakes together. So whether you're on-prem, whether you're on cloud, right? You might have big companies have all sorts of environments. They will buy Snowflake. They will buy AWS. They also have Azure. They also have a huge on-prem footprint. So why should we call it the X-Lake reasoning engine?
7:06It's because it can provide AI, the context it needs, while operating across different data lakes. And that is the reason why ADM is so powerful, that you can almost ask it anything about your environment and it knows how to get data from your cloud, from your internal environment, from your data centers, put it all together at scale and provide the AI, the real tooling or the context it needs to answer your question. So that's what that's all about.
7:39Jon Krohn:Nice, all of that makes perfect sense to me. That was a great explanation of the problems that people are facing. And it makes so much sense to me that in data management, enterprises are trying to do the same kinds of things as they're doing in marketing and sales and software development and be able to automate as much as they can. This is the future of business. There's no question about it. In our research, Ashwin, we pulled up that, I don't know how accurate this is, but we pulled up that there are four particular enterprise data agents that you guys have that solve existing data problems and unlock new capabilities.
8:13Jon Krohn:Does that make sense? talking about these kinds of four particular types of agents? This is very interesting, right? Every company or everybody who you talk to talks about agents, right? And essentially something with agency which can act by itself, right? That's the whole idea. Now, what we've also realized is that, you know, there are aspects of any enterprise where it's easy for them to adopt new technology like agents. for example let's say you know your sales people have a call and there's an agent recording that call creating summaries and sending it to you know your aes or something like that which makes a lot of sense and it's easy to do but the same thing cannot be applied to maybe you know some sort of discount offers being run on a data lake and then reports being created automatically and then shown because the business problem itself is fundamentally kind of customized to each company so if you're working with a bank, they have different problems.
9:13If you're working with insurers, they have a different problem. So healthcare companies have a different business problem. So it's not easy for us to, you know, create some sort of a stock data management agent and say, okay, you know, it's applicable to everybody. So what we're doing is talking to our customer base and essentially trying to figure out what makes sense in the data management world, right? And distill it down to maybe four, five, six agents. Right now we have kind of centralized, you know, I think some of the material online is old, but we are now almost close to 10 agents. And these are to do with different aspects of data quality.
9:47So I'll give you an example. One of the most important things that a data enterprise does is essentially create a catalog of all your data. So we've created a data catalog agent, which essentially does the same, but it has the knowledge and the context, which is derived from all your metadata. So it knows if you ask for a sales order table, where is it? How many columns it has? What is the data type? It knows almost everything about it. When does the data come in? How often does it come in? Who is the owner of it? So it knows everything. And in the same way we are starting to create all these different agents.
10:28And so when you ask a question to ADM, it can kind of spin off these agents and go and figure out what can be done And how can I best answer what a user wants to know?
10:39Jon Krohn:Right. So it's sort of similar to maybe when I'm in the ChatGPT interface or I'm in the Cloud interface or Gemini and I ask a deep research request and it goes off and it spins up some number of agents. It seems like some of these platforms like ChatGPT seems to spin up just a few agents. Anthropic seems to spin up like hundreds of research agents. and then they go over the public web or if you have connected things like your Google Drive or Microsoft Office or whatever, it'll look over all of those kinds of things as well. So it kind of sounds like a similar kind of idea, but this is operating in a way that is enterprise grade and it's designed for looking over, as you described, X-Lake, all of the data lakes.
11:23Jon Krohn:Exactly, exactly. And so that's a very interesting thought because I think in your example, you just nailed it saying hey you ask a question to claude and and it says that hey i don't have this information in my you know matrix so i'm going to go and look at the internet and look for the top 10 results and try to get you something right now whichever search engine it uses it's done let's say it uses duck duck go right or it uses google now they have done the hard work of integrating all of the internet's information into a searchable index and given anthropic an api which you can hit to say, okay, I'm going to look for, let's say that, you know, which stocks in my portfolio need to be, you know, liquidated because of, you know, whatever reason.
12:09So it's going to do all this research, hit the Bing API or the Google API and get information back. So Google has all done all the hard work to provide you that data, right? But now just imagine an enterprise, right? They have petabytes of data across many, many different data lakes. And so even if you connect a cloud, right? How do you, so the first question, how do you connect a cloud interface or a chat JPD interface with an API, which can actually aggregate data over this petabyte data scale? And how do you do it at scale? So, for example, you ask for some data set saying, hey, aggregate around my sales order, how many different segments I have.
12:53And now in the end, this has to be translated to a cluster level query, which has to be fired at, you know, in your data center in some parquet format, right? And so that's what the X-Lake interface does. It essentially provides you the tooling required to access your data lake and provides it in a way that AI can make sense of it, right? So we do all the hard work to provide you the index to your data store.
13:21Jon Krohn:It's pretty wild. As you described that process to me, I'm trying to wrap my head around it. When you're talking about petabytes of information unstructured across all these different data lakes and being able to bring a query back to your users quickly. I mean, that's pretty wild. Congrats, Ashwin, on getting that all together. Yeah. Yeah. It's a hard, I mean, it's a problem which has to be solved. I think it's going to be solved, not just by us, but over the years, right? because you know this is this this often repeated thing that the ai is as good as the data it has access to thing but i think it's a little confusing because um this the model no matter which model you use whether you use an open source one or you use the state-of-the-art foundation model right it can only access the information which is compressed in its neural net, right?
14:16And that information provides it a lot of effective ways of doing things. For example, you can ask it to write an SQL query. It'll write it. And these days it'll write it absolutely correctly. But once it's written it, it needs to hit somewhere, right? And when it needs to hit it, it needs to hit it with the same, the scale. It has no idea how much data it's going to return. It has no idea whether the user who's using the interface can access a column in that query, right? And so that's what we want to do. We want to make sure that the Xlake engine is like the context provider to AI models over your entire data landscape.
14:58So that's the whole idea behind it.
15:01Jon Krohn:Data scientists, it's time to talk about your tech. With Windows 10 support coming to an end, now is the perfect moment to rethink your setup. Enter Dell AI PCs powered by Intel Core Ultra processors. These devices are built for the demands of modern data science, delivering faster performance, smoother multitasking, and the power to handle even the most complex workflows. Whether you're training machine learning models or analyzing massive datasets, these PCs are designed to keep you ahead of the curve. Don't let outdated tech slow you down. Visit dell.com slash shoppcs to explore how you can upgrade your device and elevate your work.
15:38Jon Krohn:That's dell.com slash SHOPPCS.
16:09Jon Krohn:So how does that kind of system work? Obviously, without spilling too much proprietary sauce for our audience, but how does that work in a way that prevents hallucinations or divergent interpretations from cascading through the system? Yeah, it's a great question. I think it's easiest described with an example, right? So let's take a data pipeline. Most data pipelines are one of two ways, right? You have some sort of a drag and drop interface, which you drag and drop an ETL, and then you create some sort of data pipeline. But usually it's a part of a larger structure. and maybe written in a deep, you know, in a form of in dbt or some sort of airflow.
16:53Let's pick spark, for example, spark technology writes a data pipeline. Now spark in the end is code, right? And so let's assume that you have some piece of code written in a spark, which is representative of your data pipeline, pick data from somewhere aggregated, make some joints and dump it to a different place. Now, what will end up happening is that this spark code is being executed often onto your cluster. Now, let's take one step back. Let's look at how is something like agentic coding working these days. So we can pick up, let's say, Visual Studio Code or C-Line or Cursor or something like that, and we can ask it to generate code.
17:33Let's say we create a small website or a stock follower, anything. And so it's going to generate a lot of code, and it's going to use your laptop or your mac os as the runtime because it's going to say npm run right it's going to run something and it's going to use your laptop and its cpu and its runtime to essentially show up a browser and then you say okay i like it the design is great everything is okay now i'm going to deploy it and then you create the deployment out of it and you put it on a server where that website uses the resources or the runtime of that server, right? Now, just imagine a data pipeline.
18:13The data pipelines in bigger companies are not executed on laptops and servers. These are usually executed on clusters of machines, right? So whether it's a Hadoop cluster or a Trino cluster or a Kubernetes cluster, and then you're running jobs and you're executing a data pipeline solid. Now, what we can do is, let's say you get an alert from your data pipeline, which says that, hey, there's some issue and it requires a small fix. Now, before AI, somebody had to file a bug report and you had to clone this code. And the person who knows about what this code does and what needs to happen can actually make the fix, put it into a CI, and then go deploy it on the cluster.
18:58That's how it's going to work. Now, in this day and age, what you can also do is that you can have an agent detect if there is a log entry, which shows some sort of an anomaly or something is not right. You can actually clone the code automatically, provide the context of this error, provide the metadata of what this data, this pipeline is touching through the Xlake engine, saying which table, how many columns, what data type. And the AI can actually then just rewrite the Spark code. And then you can then go ahead and deploy it back into the execution engine, which is the cluster, right? Theoretically, all of this is possible.
19:41And because we know that AI can generate great code, especially with Gemini 3, almost one shot, most 90 % of what you want to do. And then if you automate the rest of it, which is the metadata context, how do you deploy it into the cluster, then you can imagine an agent which can actually self-heal, right? And that's somewhat the path which we have chosen is how do you provide, number one, AI context about where's the code, what is the pipeline, what is the metadata, and what are the errors which are happening? And then attempt if the AI can actually fix the code or the pipeline by itself. Now, not everything can be fixed.
20:22If you have some proprietary, you know, Informatica pipeline, HVR pipeline, some Oracle standard, you know, stored procedure doing it, not so easy. But more and more, especially with now with AI, because of, you know, code generation, I feel is one of the biggest applications which we'll see, right? It is probably number one in terms of the potential because the code can express a lot more than anything else. For sure. And you can know if it works or not. Exactly. And so more and more, we feel that data pipelines are going to shift from drag and drop interfaces to code because now it's as easy.
21:02It's maybe easier to generate code than actually drag and draw pipelines. And so as more and more data pipelines become code, then it's easier to also mutate that code through AI when you have an error. So all you need is something which gives you the context of your entire data lake, what are the errors happening, what are the issues happening, and then essentially have a pipeline which kind of feeds it. So that's the part to auto-remediation. Now, I won't claim that we can auto-remate everything, but if most of the factors are right, hey, you have a code-based pipeline and it's largely, you know, it's version control and it's Spark or it's, you know, SQL or something, then it is entirely possible to attempt some of these things.
21:42Jon Krohn:That's really exciting, Ashwin. And everything that you were talking about there was fully autonomous, but it sounds like your agentic data management system, ADM, also allows for a human in the loop if people want that. So how does that work? Do you define when a human should be in the loop or does the client define that? Is it on a case-by-case basis? How does that work logistically? Yeah. Yeah, you know, philosophically, I think even the best agentic systems or, you know, even the simpler ones, you would want some control before you actually accept it, right? Whether it's an, you know, you probably won't trust an AI to generate automated emails for you, right?
22:25You would probably want to review it. And I think the more complex the problem is, I think it's going to be very hard for people to accept anything blindly which an AI provides. And so when we talk about autonomous, I think the work differential which you can generate versus a human being, that's the autonomous part, right? so uh you know it's as simple as let's say you know you ask a colleague to write you write you up some code or a module and you do a code review before actually accepting it right for for most every junior engineer or anybody in your org and it's a similar philosophy i think the differential the autonomy the value the economic value comes from the fact that you compress maybe you know one week of work into maybe an hour by doing a lot of autonomous stuff.
23:18But I think the need to approve, disapprove, review, or reject has to come back to somebody in the value chain. And usually, these are the people who are the experts, right? So somebody who knows the system, but now does not need, let's say, an army of engineers, but he needs an army of agents, or they need an army of agents, but they still have the control on approve, reject, review, do better, or just completely ask it to rewrite. I think that's not going to go away. And that's, in the way I think about it, that's what autonomy means. Not just that you do the whole thing without any supervision, but you do most of the work without supervision, but there's somebody at the gate all the time.
24:02Jon Krohn:Of course, that makes a lot of sense. And so this complex system, this autonomous data management platform, ADM that runs across petabytes of data, has agents involved, has these human-in-the-loop capabilities. It sounds like for the most part, if not entirely based on our research, you can correct me where I'm wrong here, but it sounds like a lot of this you guys did from scratch. You didn't want to rely on existing open source technologies and just kind of build on top of that. I think some of it came from the fact that, you know, we've been around since 2019, like I said. And so we've built, talking to customers, we've built a lot of other technology as well, right, which can then go ahead, access data, different engines, how do you reliably run, let's say, validate a million rows every hour, right, for quality, something like that.
24:55and we have essentially ended up using the same technology under the hood and especially in this field, I think what differentiates us is that we're not just a metadata player. A lot of, I think, data management relies a lot on metadata and then providing value on top. Our engine differentiates itself by also having access to the data underneath. So when you deploy the ADOC or the ADM platform, You just don't have something on the cloud and then it sucks all your metadata. You also deploy what we call a data plane, which is in your environment. And we have built the technology to provide like a secure MTLS system between the two.
Read the full transcript
25:38So no matter how much data you ask for, right, you can ask us to validate a hundred rows, a million rows or a hundred million rows. We rely on the data plane, which runs in your environment to actually have access to data. And so over the years we've had to build a lot of these pieces. How do you connect them? How do you make sure the jobs are running reliably? How do you authenticate users down into somebody's environment? And so in some sense, we have kind of put all of that together and reused a lot of that stack. But of course, we wouldn't be there without open source technology. So we rely, we built our stuff on Kubernetes, Spark, and all the good work the community does.
26:20and we've also tried our best to you know give back some of it um so we run an open source platform called odp you can actually download it it's like a full data platform which is available for you um you know if if you're interested and that is something which we have built and self-managed because a lot of work our work internal work is also you know making sure we can run different environments at scale that's really cool odp it's just like a github repo it's just a github repo yeah it's a it's our own version of the whole big data management stack and you know it's uh you it's it's open source and we patch it and we build it we make we put the security vulnerabilities there and so a lot of these kind of primitives get used at the underlying system so if you have let's say you know um a rack which is available to you and you don't have compute, you can use this.
27:14This is open source compute, up to date with upstream vulnerabilities taken care of, and it all integrates with the accelerator platform. So it's more of an end-to-end story that we have built. And when you start up, let's say you're four people, you're 10 people, you don't have the resources, capital. So you pick your battles. So we have done the hard work over the years to release different parts of the component. And now we feel with ADM, we kind of have the whole stack. So we have ADM, we have ADOC, we have the data management, we have the compute as well at the bottom. And now we have AI, which can actually accelerate a lot of this work as well.
27:52Jon Krohn:I see. I found the link to your open data platform, ODP. So I'll have that link for our listeners in the show notes and our viewers in the show notes to check out. Thanks for that. So let's go back a little bit to the way that you have built this platform, ADP, to span petabyte scale, complex, hybrid, multi-cloud environments. In a previous interview, you've described this as something like the captain's deck, where you have all the information you need across this extremely complicated system at your fingertips. And you noted that part of the big problem that you're solving with Excel data is what you call data sprawl.
28:34Jon Krohn:Do you want to tell us about what data sprawl is? That's actually not a term that I think I've ever come across. Yeah, it's very simple. I think, you know, I once did a round table and a lot of these, you know, representatives from most enterprise companies come in. And I keep repeating this term enterprise is because I think it's very important to know the kind of customers and the people that we work with. Technology is only as useful if it's applicable. So, you know, it's very hard for us to sell our software to another startup because the startup is, you know, busy building the business. They're not in a space where they are investing hugely in data technology, right?
29:15That's a game which requires money, requires time, and you need to be in a position where you're using the data for something. For example, a customer of ours like a Nestle is a global organization where they invest a huge amount of money in data to provide competitive advantage, right? In different markets and all the skills they sell. Now, it's not true for, let's say, a company like us, which is, you know,$100 million. So some of these problems are, you know, quote unquote, enterprise problems. And one of the things which an enterprise usually does is that they buy a lot of technology. So if you ever have the chance to talk to some of the data leaders, they will always tell you that they have been doing this since the time, 2010, build data platforms.
30:03And so that environment is never homogenous. You won't find a company which is essentially frozen on one technology. And there are different groups in the organization and they use different set of technologies. So now, usually you end up having some cloud environments where there's some data, there's probably some investment on something like a snowflake or data breaks to drive analytics. Maybe some new leader came in and invested in new technology. Of course, there's like a huge legacy landscape of data sitting in data centers. And that's what essentially data sprawl is, right? You have hundreds of terabytes of data on your on-prem clusters, and you moved some of it for analytics into the cloud, and you moved some of it into Google because you signed a cloud transformation journey with Google.
30:57And now you end up in a situation where you have data everywhere. And somebody needs to find out, right, what's where, how do I access it, and maybe put them together in some sort of a report. And that requires you to, number one, know where all the data is, right? And number two is to be able to access it in a good way, which can make sense to a user. And so that's what really data sprawl is, right? It's the nature of, it's the, let's say, eventual fate of any big enterprise that they have data everywhere. And, you know, they want to centralize, but it's hard. It takes many years. Until then, they need to work with all of this data.
31:44Jon Krohn:Fantastic. Well, this has been a great tour of the Excel data platform, all the exciting things that you have going on there, the agentic data management platform, ADM, as well as the ODP. Yeah, open data platform as well. And so I'd like to shift gears here a little bit now to kind of maybe some more generic insights that you might have for our listeners. Of course, it'll still end up centering a lot on Excel data, I'm sure, given how much it's been a big part of your life for the last six years. In an interview, you mentioned realizing that being a CTO is as much of a people problem as it is a pure technology problem.
32:26Jon Krohn:And that you had to move from an architect slash programming mindset to putting yourself in the shoes of customers, of marketers, and employees. So what kinds of guidance do you have for our listeners? What kinds of habits or frameworks should our listeners be thinking of to have the kind of success that you've had as a technology leader and, you know, growing a huge startup like Excel data? Yeah, you know, we've grown from four people to 100 people. I'm just talking about engineering as such, right? We have, of course, more people in the on. And we have done this through COVID and remote work and then coming back to office.
33:09And the reason why I said it at that time, and I really believe it, is because, you know, most of the time as technologists and programmers, you always believe that your version of what you're building, right, is correct, right? It's the best way of doing something, right? So if I build it, of course, they'll come, right? That's the standard thing. And then there are two aspects to it. One is that why do you need to think of it from your customers? perspective because the customers they operate in a different sort of an environment they don't have the same degrees of freedom that you do and most of the time everybody is coming in from some sort of a guidance right so you know the way i think about is that um a new ceo comes in or a new data leader comes in and they come in with you know some sort of vision that hey you know i can do this because i know i've been successful somewhere and now i'm going to go down this technological path.
34:07And then somebody down the organization gets a guidance. And the guidance is, you know, it has two axes. One is the technology, you know, axis of it, which is saying, okay, we need to modernize, we need to use this technology. And for example, now the whole thing is we need to do AI, right? That's the whole thing, something like that. And the other one is the business, which is that there is a business problem you're trying to solve. This is not like a school project where, you know, you build some sort of a website and say, hey, this, you know, Gemini one-shotted this and it's so cool. So usually it's in those axes.
34:40And unless you actually put yourself in the customer's shoes and see if your technology, whatever you're building, is actually applicable to what they're doing, can bring them some value, can make them successful so that they can show this success to their leadership. And then it actually adds value. So next year, when they come back to you, they're willing to buy more or they're willing to continue with your POC, POV, and then convert it into a contract. And then you get revenue and then you expand from there, you build new offerings, right? So unless you think about it, because your engineers are not going to think about it, right?
35:23So you hire people, you give them a task, you say, hey, you're going to work on AI, you're going to work on data, you're going to work on network. Nobody's going to think about it. So somebody in an engineering org has to actually do it. I've always felt that product managers are great, but they're focused on features and issues and bugs. And it's really difficult to be that good at all of these aspects. So as a technologist, somebody in the org, so if not the CTO, then who else, has to think about the fact that, am I building something which will bring value to a customer? And how do you create an org which kind of gets that information?
36:03and so that's from a customer's standpoint and from you know from an engineer's perspective um most people would not be in this is something which i've read as well right at after some point a salary or you know money is not so interesting to people and so if you want to attract um people who want to stay with you especially during a startup's first few years, you don't want people to come and go all the time because it kind of, you know, hinders your ability to carry on context and build at a certain speed. And so if you want to bring in people at an early stage, they have to have the belief of, hey, what is this company going to do?
36:48Why am I doing it? What's my own, what's my way of expressing, let's say, creativity through code or through engineering. And that is hard to do if all they get is like a bunch of Jira tickets which have to be done through Sprint. Right? So you have to create an environment where people think that they are, in some sense, innovating in whatever domain the company has chosen for itself. And for us, that has to do with enterprise and data management. And so the way we want to run our R &D and engineering is to see how much we can innovate here and not just the fact that, hey, the company roadmap has been decided in the beginning of the year and here is a bunch of plans for the rest of the year and do it.
37:35So we have allowed, or let's say, work with our team to innovate on a lot of things. And some of these examples are like the ODP link. This is completely engineer-led initiative. If you don't have a data platform, we needed one internally for testing. We don't want to pay license fees. We don't want to have security vulnerabilities. So we built it. We also built our own vulnerability management systems because we ship a lot of this enterprise software and everybody will scan it. There's a lot of vulnerabilities. So we built a whole AI agent to fix vulnerabilities. And we've built our own kind of data center stack, like an open stack, because, you know, VMware is expensive.
38:14Open stack is too complex. So we built something. And so this is, I think, what inspires people to say, hey, I'm here to do technology. And this company happens to do technology in this arena. And that's where I'll spend time, right? So you got to put yourself in the shoes of the engineer you're hiring and also the customer, you know, you're talking to, and then try to bring in a plan, which kind of works for both of them. So at least that's the way I think about business.
38:41Jon Krohn:That is such a great answer. Sure. You went into a lot of detail there on even the kinds of things I was going to ask about in my next question, which were because you've said in past interviews that a talented programmer is looking for meaning in their day to day that, you know, highly paid engineers still just complain about their jobs. It's funny to me to think that like somebody who gets a hundred million dollars signing bonus at Meta is just like, oh, man. But I'm sure that happens. You know, it's, I don't know what, where the stat is at today, but, um, when I was doing my PhD, something like 15 years ago, I attended a lecture on the economics of happiness, just, just for fun.
39:19Jon Krohn:I just went to this lecture. And at that time it was showing that in the U S if a household was making over something like 80 ,000 or a hundred thousand us dollars a year, that, that is the happiest you can be making more money beyond that point. didn't make people happier. Now with inflation, the numbers are probably a little bit higher. But directionally, I think this kind of gives the idea that you're explaining, which is that beyond having your basic needs taken care of and knowing that you have security for you and your loved ones, the extra money beyond that could end up being a hassle. Yeah, it is.
39:56And it's also interesting, right? I mean, there are studies on developer productivity, right? And you would see that the numbers are insane. I mean, people talk about how an engineer, a software engineer is productive, maybe four hours a day or three hours a day. And the rest of the time is spent in meeting, planning, whatever it is. Now, I feel that even if you take two hours for meeting, you still are leaving a lot of this time out. And if you think about it, what better privilege can someone have than to sit in a usually a great office on a laptop without having to move. Moving is optional, right?
40:35And then get paid top dollar for it. And most engineers then leave their jobs. You know, it's not just, you know, people leave Accelator, but people leave all sorts of companies. And there has to be a reason for it is that most people would be happy because knowledge work is something, you know, which is, which has to do with creativity. It's hard to sit in one place and realize that the work which was presented to you or asked of you could be done maybe in two hours and then you got to sit and find something to do. And it's good for a few days and you spend some time. But after a while, you start getting this feeling that, hey, what am I doing?
41:13I'm supposed to do something better. Let me find a mission which resonates with me and my work and my philosophy of it. And so I think making sure that no matter what the business environment you're in as an executive, the engineers or the R &D teams believe that fundamentally we are in the business of innovation in this field. And there are very few fields, whether it's, you know, something data management, even something as boring as the enterprise content management. I'm sure that there is innovation that could be done, new ways of doing things. And people should believe that they have the freedom to do it.
41:56And they're not just, you know, dictated by quarterly plans and this. I think if we can provide an environment like that, then new ideas come in. And, you know, for us, it's worked out because for a company of our age and size, we have a lot of, let's say, capability that we've built over the years. Whether it's to do with, you know, ODP, ADOC, We have a pulse monitoring system. We have ADM. We are working on the next version of our platform, which will be released in May. And so that is what allows us to do it is where people believe that, hey, in this field of what the company has chosen, data management, there is innovation that can be driven through pure engineering work.
42:40And that's what drives people.
42:42Jon Krohn:Nice. I like that. How do you say in an interview, do you think you have a way of telling whether somebody is going to be passionate about a technical infrastructure heavy mission like data management at Excel data versus somebody who's just coming to collect a paycheck? I think it is, it's easy to tell in some sense. Of course, we've made mistakes there as well, like everybody else. But I feel once you start talking to people, the, and this is what I felt. I mean, I've always felt that management in the technical field can only be done by people who have been in the trenches to some sense. And I'm sure there are models everywhere else and which which are different.
43:35And people have seen managers who work extremely well without actually being on the field. and so the number one thing at least when i look for um potential hires is to see if they can build things right and it's it doesn't have to be working on some data problem right the question is can you build if you are given a problem right and it could be any problem it could be something like, hey, how do you design, let's say, an e-commerce warehouse? How do you design a logistic system or any other business problem? And then can you translate it to something which you have learned? You probably know Go or Python or Java.
44:27Can you put something together which represents a real world problem. I think if you can, then those are the people who bring the most value, who can actually look at like a business problem and then convert it down to what they know. And that's what technology is a good end. Of course, this is for slightly senior people. I think for people who are just coming out of college, it's purely based on potential. Saying, hey, you know, some of it is your scores and your background and some of it, hey, how interested are you into doing this? And then you take a bet and maybe after a few months you decide.
44:57But for most senior people, I would recommend checking if they can build things.
45:04Jon Krohn:That makes a lot of sense. When you're interviewing today, to what extent do you encourage people or discourage people from using LLMs to support themselves? And how do you assess, you know, if you're in a situation where you're trying to assess somebody's capability in Go or Python, yeah, how do you do that and ensure that they're not using an LLM to cheat? Yeah, so one of the things which we've done now, this year onwards, is not to have remote coding interviews, right? We don't want to do it. There are way too many tools which you can use to cheat, and it's very hard to tell. And I always felt personally very uncomfortable, you know, trying to, when I'm talking to you, but I'm always looking for an evidence that you're doing something wrong.
45:52And so it takes away from that conversation, isn't it? And so, you know, we don't want to do that. So the number one thing which we're doing now is face-to-face. The second thing, you know, is that we tell the candidates that, hey, of course, you're going to use LLM to the job. We have given everybody cursor licenses and, you know, got chat GPT licenses, whatever you need. But this interview is all about problem solving. And so there are two things we have done. One is to go towards more practical approach where we kind of created a problem which represents what you would do day to day. so a bunch of failing tests some missing implementation and you won't do it of course you can google right you nobody memorizes all the api so you can google figure out but don't use an llm uh to answer the question and even if you end up using an llm it's unless you install like cloud code or something it's not possible to cheat because you'll have to just copy paste context from different places find an answer um so you know make some sort of a gentleman's agreement with them and then start off.
46:54But, you know, I feel that more and more code will be generated by LLMs. So it might not be a scalable strategy going forward, but this is what we're doing right now.
47:04Jon Krohn:Yeah. Yeah. It is really cool how much you could be getting done with LLMs today. And on that note, a few months ago, you demonstrated an MCP server that prompted a GDPR compliance scoring agent and an interactive dashboard. And it highlighted how AI could complete in minutes what would have historically taken a team days of coordination, research, estimation, and meetings. So as a CTO, building a long-term product roadmap, how do you reconcile this compressed innovation cycle that's now possible today with the slower risk-averse realities of the large enterprises that are your clients? The enterprises, I feel they're all slow in themselves in the decision cycle, but they want it, right?
47:50Because nobody wants to lose out on innovation. So that's why everybody's invested, even though they probably won't be using it, or they don't have the go-ahead from all their compliance and security people or whatever, but they've started using it internally. And so I feel this trend more and more is going to continue. And at some point, you know, some of these people are going to adopt it faster, they will see a competitive advantage, the other people will be like, okay, we are left behind, let's go figure out what we need to do to get there. And over a period of time, it will become more and more easier.
48:29And so our job is also to anticipate the future in some sense, right? We say, you know, we don't build something because somebody wants it today. But you look at where the industry is going and you realize that ai is coming for these sort of jobs let's say as well or these sort of functions more more appropriately and then plan to meet them there right so that's the whole strategy behind it so the whole mcp demo which i did was kind of a idea right which says that um there's this engine which has access to all your data then you have a coding environment now can you connect the two and create something which would have taken somebody a lot of time to do, but an AI can actually do it.
49:10Now, it's not complete, but it's not incomplete as well, right? So the AI has done, say, 60 % of the work, and you could probably prompt it to get 80 % of the work. And then maybe two engineers can finish the rest instead of having a whole team. So more and more, whether it's, you know, vendors and technology builders like us, or it is people who are adopting, you know, AI in enterprises, they will realize that unless they do this, they are going to be left behind. And we want to be in a situation where we meet our customers when they are ready. And that means we have to do all the hard work of making sure we are ready as well with the solution.
49:48Jon Krohn:Great answer, Ashwin. Let's have all of your answers in today's episode. I've really enjoyed listening to you speak. You're a wise man, if you don't mind me saying. It's really nice, really enjoyable and outstanding communication. So yeah, really enjoyed this episode. Ashwin, before I let my guests go, I always ask for a book recommendation. Do you have anything for us? Yes, I do. So, you know, I don't have like a regular reading habit, right? But I try to do at least a book like in a couple of months or so. And I was super interested in understanding about, especially the the g2 as they're calling it right the whole um u.s china rivalry in some sense which is going on so i read a couple of books but i think the most recent one which i'm almost done with is called house of huawei which is by um a lady called eva 2 or eva lu which i forget the name.
50:53And it's all about how Huawei started and, you know, how it became the technology company it is. And also goes in, in the beginning of China's transformation from, you know, kind of a socialist, completely socialist thing to now a capitalist plus socialist society, and also addresses the challenges around some of the human rights stuff. Very interesting, highly recommended for anybody to read, just to get what's going on.
51:22Jon Krohn:Yeah, that's really interesting. I hadn't heard of that book, but it sounds like a good mix of technology, geopolitics, news, just being able to understand what's going on in the world better. Yes, it's a great book. Nice recommendation, Ashwin. Thank you. All right. So for folks who want to get more of your insights after this episode, where should people follow you? Do you have social media accounts that people should follow? I'm a little bit active on LinkedIn about especially Accelerator and what we're doing. I don't unfortunately have an account on X or anything like that. But LinkedIn probably is the best place.
51:54And then Accelerator has some channels. So, you know, please visit our website. And there's our resources and blogs and stuff we have built and written about, which you can read.
52:03Jon Krohn:Yeah, we'll have links to all of that in the show notes. And I got to say, Ashwin, I don't think you need to apologize for not having an X account. People saying today on this podcast that they just have a LinkedIn account is the most common answer. So, yeah, that's how things have evolved in recent years. Yeah, a few years ago it was different, but now it's crazy. It's LinkedIn almost all the time. Yeah, it's LinkedIn also. I feel people are a little measured in communication over LinkedIn than they are on X. And that has to do with how the algorithm and how, what they want the platform to be, which I think is more of like a tabloid than a source of information.
52:47And somehow it doesn't appeal to me. Yeah, yeah, for sure.
52:51Jon Krohn:I think it's interesting, you know, LinkedIn, it's almost because it doesn't really change very much. I think people get comfortable with that. and it's funny like social media platforms they always feel like they need to be changing and adding this and adding that but i think people have found that you know just being able to see you know often business related information that's relevant but often still also entertaining as well and i think a little bit positive so i think something that's interesting like you talk about um the algorithm there over at x it seems like it wouldn't surprise me if on linkedin i am regularly interacting with people who have very different political views from me, but it doesn't matter at all.
53:34Jon Krohn:You don't even know. Yeah, yeah, exactly. And X is different. Yeah. You know, I think it's a personal preference. Some people are that way that they are, they would like to engage in a debate and maybe it's not just all about, and then they find joy doing it. And for more people, for people, it's more of you know, broadcast where they put out their thoughts and then maybe some people find some, you know, something interesting in it. Yeah. But I think one should be able to, and it's something which, you know, I'm thinking about as well, that you should be present as a representative of the company because you don't know where customers are, what people find interesting.
54:15So nothing to do with X itself. I think it's just that some people like me, just not comfortable with that sort of engagement. That's it. But we should be doing more. Yeah, yeah.
54:30Jon Krohn:All right. Well, thanks for that unexpected little social media conversation at the end, Ashwin. I really enjoyed today's episode. And yeah, wishing Excel data all the best. Seems like you guys are on the right track. and yeah, hope to be catching up with Excel Data again soon on the show. Thanks, Sean. Thanks for having me. I had a great time and hope to be back soon. That was an awesome episode. I learned a ton. I hope you did too. In the episode, Ashwin Rajiva covered how Excel Data's agentic data management platform uses AI agents to automate data quality checks, cataloging and pipeline maintenance across enterprise environments, compressing work that traditionally took weeks into hours while keeping humans in the approval loop if desired.
55:15Jon Krohn:He talked about how their Xlake reasoning engine solves the problem of AI models having no context about your internal data by providing tools that connect to data lakes across on-premise cloud and hybrid environments, enabling queries at petabyte scale. He talked about how self-healing pipelines work by having agents detect errors in logs, access the code repository, understand metadata context through the X-Leg engine, rewrite the Spark or SQL code, and redeploy it automatically to compute clusters. That's pretty crazy to me. And he talked about how engineering retention comes down to meaning and innovation rather than compensation alone, with studies showing engineers are productive only a few hours per day when reduced to executing predefined sprint tasks instead of solving creative problems.
55:59Jon Krohn:As always, you can get all the show notes, including the transcript for this episode, the video recording, any materials mentioned on the show, the URLs for Ashwin and Rajiva's social media profiles, as well as my own, at superdatascience.com slash 957. And thanks, of course, to everyone on the Super Data Science podcast team, podcast manager Sonja Breivich, media editor Mario Pombo, partnerships manager Natalie Jaisky, researcher Serge Massis, writer Dr. Zarqar Shea, and our founder Kirill Aromenko. Thanks to everyone on the Super Data Science podcast team for producing another super episode for us today, for enabling that super team to create this free podcast for you.
56:37Jon Krohn:We're so grateful to our sponsors. They allow this show to happen. I guess you guys do too by listening. They're all, we're all working together here. But you can support the show by checking out our sponsors links, which are in the show notes. And if you yourself are interested in sponsoring an episode, you can find out how at johnkrone.com slash podcast. Otherwise, share, review, subscribe, all that good stuff. But most importantly, just keep on tuning in. I'm so grateful to have you listening. And I hope I can continue to make episodes you love for years and years to come. Until next time, keep on rocking it out there, and I'm looking forward to enjoying another round of the Super Data Science Podcast with you very soon.
From the publisher
AI agents, data lakes, and managing data sprawl: Ashwin Rajeeva, cofounder and CTO of Acceldata, speaks to Jon Krohn about how the agentic data management startup raised over $100 million in venture capital to expand its business in automating data quality assurance as well as cataloguing and pipeline maintenance across enterprise environments. Acceldata utilizes multiple agents to solve enterprise-grade questions with company data. It also uses autonomous data pipelines that can detect and fix issues without human intervention, and the platform’s agentic data management system ADM also lets humans stay in the loop wherever needed.
This episode is brought to you by the Dell, by Intel, by Fabi and by Cisco.
Additional materials: www.superdatascience.com/957
Interested in sponsoring a SuperDataScience Podcast episode? Email natalie@superdatascience.com for sponsorship information.
In this episode you will learn:
(03:25) About Acceldata and xLake
(15:03) Autonomous data pipelines
(21:02) How and when to keep humans in the AI loop
(27:43) How Acceldata solves ‘data sprawl’
(31:53) Habits of successful tech leaders




