1030: Garbage In, Gospel Out: Why Agents Need Better Data, with Salesforce's Gaurav Pathak

25 Sep 2026 · 23 min · 11 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

“Garbage in, Gospel out” explains why enterprise AI agents fail when data context and metadata are wrong. Bad data used to trigger a “data brawl” (humans argue over conflicting dashboard numbers); agents now produce confident answers or make decisions silently, so errors go unnoticed and can be costly. Metadata is framed as “labels on supermarket cans” that help agents find the right data without “opening and smelling” every dataset.

Key claims

context is ~95% of the battle; enterprises must govern data/model access; token burn rises when agents must search many stores (example: PepsiCo has ~800,000 data stores).

Notable examples

support chatbots escalate to humans and make unmet guarantees; profit rules like revenue minus cost illustrate data quality rules.

Guests

Gaurav Pathak, SVP Product Management at Salesforce; previously 13 years at Informatica (metadata and AI products, including the Clare AI Engine).

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Experiencing Dreamforce

0:46 to 2:18

Gaurav shares his experience at Dreamforce, including celebrity appearances.

“Gaurav, welcome to the Super Data Science Podcast.”

Transition from Informatica to Salesforce

2:18 to 3:26

Discussion on Gaurav's transition from Informatica to Salesforce and the evolving data landscape.

“And now you've been inside Salesforce for about 10 months post-acquisition.”

The Importance of Metadata

3:26 to 6:12

Gaurav explains the significance of metadata in enterprise data management and AI.

“So being able to extract metadata, which is nothing but, you know, the example that I gave in my talk is like you imagine going to a supermarket, right?”

The Shift in Data Handling Challenges

6:12 to 7:49

Explaining the shift from data brawls to silent failures with AI agents.

“is a far bigger problem than if, say, the data were just flowing into dashboards or something like you might have done a decade or two ago.”

Garbage In, Gospel Out

7:49 to 8:41

Gaurav highlights the dangers of bad data in AI systems, coining 'garbage in, gospel out.'

“They look at the data and they decide which customer to call based on number of calls logged, for example.”

Building the Case for Data Quality

8:41 to 10:46

Discussion on how to convince organizations to invest in data quality amid a focus on agents.

“And so you need to have this single source of truth for metadata, for data.”

Understanding Data Quality Rules

13:35 to 16:42

Exploration of data quality rules and the advancements in managing them with Clare.

“And then so basically what you're saying is, because my question was kind of, how do you make the ROI case?”

The Evolution of Clare's AI Capabilities

16:42 to 18:10

Discover how Clare's AI engine evolved to handle increased data quality demands.

“And it can do it with an Excel file with a thousand data quality rules that you uploaded in an hour or so.”

Skills for AI Engineers in the Agentic Era

18:10 to 19:56

Explore essential skills AI engineers need to thrive in an era of advanced agents.

“But for those folks listening, what kind of skill or habit now matters more that we're kind of in this agentic era and that agents are consuming so much more data than people are?”

Book Recommendation for Data Practitioners

19:56 to 20:39

Gaurav recommends 'The Mind Illuminated' for data practitioners to manage overwhelm.

“I can't believe you had those all just ready off the top of your head.”
Show all 11 chapters

Follow Gaurav Pathak's Work

20:39 to 21:31

Learn how to connect with Gaurav Pathak on social media for insights and updates.

“Mind Illuminated will teach you how to calm down, get the light back into your mind as well.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Jon Krohn:When humans get bad data, they argue about it in meetings. When AI agents get bad data, they pick a number and present it with total confidence. My guest today calls it garbage in, gospel out. Welcome to another episode of the Super Data Science Podcast. I'm your host, Jon Krohn. Today's guest is Gaurav Pathak, SVP of Product Management at Salesforce. Gaurav spent 13 years at Informatica building its metadata and AI products, including the Clare AI Engine, before Salesforce acquired Informatica last year. In this episode, filmed live at Dreamforce in San Francisco, Gaurav explains why context is 95 % of the battle for enterprise agents, who the sin eaters are that pay for an agent's mistakes, and the three skills that matter most for AI engineers today.

0:46Jon Krohn:Enjoy.

0:49Gaurav, welcome to the Super Data Science Podcast. Thank you for having me, John.

0:54Jon Krohn:And thank you for inviting me to Dreamforce. We're filming live from Dreamforce in San Francisco. And it has been an amazing day so far. It's my first day ever at a Dreamforce. Oh, wow. How has the conference been so far? It's been amazing. The biggest thing for me was Gwen Stefani played before Mark Benioff, the Salesforce CEO, spoke at his keynote. And she's been such an important person to me my whole life. I've never seen her perform before. And she played Don't Speak, which is this heart-wrenching song. I literally burst into tears in public. She's perfect. Yeah. She is. Look forward to the full performance.

1:29Jon Krohn:Yeah. And there were other, you know, people, I was going to say lesser known people, but these days they're almost as famous. Sam Altman's here. Dario Almedeo is here. Obviously, Marbenov is here. Matthew McConaughey. Reese Witherspoon. Almost as famous. I take that. Yes. Great. I love the Dario's interview in the morning as well. And then Jensen was great as well. Jensen Juan. Oh my God, I forgot him. Yeah, maybe the biggest kind of hitter of them all. And it's just, there's so many, how can you even remember them all? Yeah, they were really funny. Jensen and Dario were both really funny. The banter with Mark, you can tell they know each other really well.

2:06Jon Krohn:Anyway, we're not here just to talk about Dreamforce. We're here to talk about some meaty technical stuff for our listeners. So you spent 13 years at Informatica building the metadata and AI products there. And now you've been inside Salesforce for about 10 months post-acquisition. Does that mean it's your first Dreamforce? It is my first Dreamforce as well. Loved all the performances, loved all the media sightings and celebrities as well. But Dreamforce is at such a different level of being a software conference, right? This is the OG of software conferences. So it's amazing to be here, see all the people who are participating and see all the core problems that they bring in, like how to be an agentic enterprise, how to get their data ready for AI.

2:51We love those interactions.

2:53Jon Krohn:Sure, and speaking of problems, what changed about the problem that you're solving when you went from being at Informatica, a data company, to being acquired by Salesforce, which is now an AI agent company, actually? So a lot, and it's not just the acquisition that changed the things. It's also what's happening in the market, with AI coming in at full force. we at Informatica, when we started, we started doing the metadata work all the way back three decades ago, you know, with products that were very technically focused for data engineers and such, right? So being able to extract metadata, which is nothing but, you know, the example that I gave in my talk is like you imagine going to a supermarket, right?

3:41And then you all see our tins without any labels on them, right? And you have to buy a tomato soup without. So, and if you are expected to open each of them and smell them and then get the tomato soup, that's going to be very, very difficult. Metadata is... It's a hygiene nightmare as well. Oh, that's, yeah, if they allow it and so on, absolutely. What metadata really is, is the label on these packs and cans, right? It tells you, you know, what this is about, what are the ingredients, what it is good for, what it is not, and so on. We were in the business of collecting this metadata about every data asset in the enterprise and making it available for humans.

4:24We created the first data catalog way back in 2010. We called it the Enterprise Data Catalog. It was the Google for Enterprise Data Assets, one place where people can come to find the most relevant, trusted data asset for their analytics needs or for their data science needs. But coming into now and then a year that has gone past, we are now seeing more and more AI agents that need the same data with a lot more explicit metadata about what's in there. Because as we have seen these agents getting deployed into the enterprise, they are seeing these data assets, like, you know, it's like the Indiana Jones warehouse, you know, I don't know, I will date myself, Raiders of the Lost Ark, big warehouse with alien artifacts and things like that.

5:13And how do you find the alien artifact? It's like that for agents. If you don't have these labels, it will be every can to be opened, to be smelled, and then you get to the tomato soup. With metadata, you get to find it a lot easier.

5:26Jon Krohn:So it sounds like if you, in the agent world, if you had to be inspecting every grocery store can to find the tomato soup, that would be a lot of wasted tokens, I imagine. Absolutely. So that's the other thing, right? You'll spend a lot of time to first find where this thing is. You're looking for your customer data. You're looking for your supplier data. You're looking for a tomato soup data, right? And it's very, very difficult because you have to go in and look around every database. we have worked with customers who have 800 ,000 data stores. PepsiCo is a good example. And if an agent has to do that again and again in every database, the amount of tokens that will be burned will make every AI company happy.

6:09Jon Krohn:And it seems like having bad data in this agentic world that we're now in is a far bigger problem than if, say, the data were just flowing into dashboards or something like you might have done a decade or two ago. Do you want to tell us why bad data are a bigger problem now than ever before? So the difference I call between these two things is a data brawl that used to happen earlier versus a data brawl. Data brawl. That's right. Oh, B-R-A-W-L. That's right. That's a fight. Right. And I'll talk about how that is versus a silent failure, which is happening now. Now, in a decade before now, when we gave every employee in an organization a self-service BI tool like Tableau or Power BI, that opened up a lot of opportunities for these employees to now work directly with data.

7:02But the problem was, again, the same thing, which data is trusted, which data is of high quality. It was very common to go in a meeting where you'd have some metric that you're looking for, number of employees in my group, and three people coming with different numbers, right? A sales guy will have a different number, a finance guy will have a different number, the org guy will have a different number. And we would call it the data brawl, right? So they would affectionately, that, you know, we talk, oh, no, something looks wrong with your number and so on. And then you'd figure out what really happened.

7:33You know, people got the wrong definition, people got the wrong data, some data quality issue that happened. The good thing with the data brawl, though, was that it was loud. So we knew that there was something wrong. what's happening now is the agent picks one number it's brings that number to the human if you know we are in a dashboarding case and it very confidently tells the human that this is the number we call it garbage in gospel out right because now it is these agents are so persuasive that they can tell you anything and you will believe them right but even worse is cases where these agents are making these decisions themselves.

8:17They look at the data and they decide which customer to call based on number of calls logged, for example. And in those cases, humans are even missing out of the loop. And if they are not getting the right high quality data, they make wrong decisions that land eventually at the doorstep of a human who would, you know, have a very hard time explaining the decisions.

8:39Jon Krohn:Garbage in gospel out, that's going to stick with me for sure. So given that with what you're saying, it's crystal clear that having the plumbing right, having the data quality right, you know, an informatica-like backbone is essential to being able to have an effective enterprise agent system today, especially when you can have lots of different agents from different providers, you know, working each department in a business, or you could imagine even within one department, like a data science team, they might have multiple different agent vendors. And so you need to have this single source of truth for metadata, for data.

9:17Jon Krohn:So it sounds clear to me, actually, to get agentic AI right in an organization, you need to be making a big investment in the data. Absolutely. That is the thing that enterprises bring in to intelligence. It is their business. It is all of their data about the business, which is the context that feeds into these agents themselves. And it's very, very important for enterprises to get that right. You're completely right that there will be a lot of model choices. We have seen large organizations where it's almost become a wild, wild west. People download models from Hugging Face, from other places.

9:56And in some cases, these are not even approved models to be available to employees, et cetera. And governing that has become a big challenge itself. You don't want to be sending data to model providers and vendors who may use it for training the models or even worse things as well. So governing that, making sure that your data is in the right place with all the labels properly put in is very, very important for organizations to invest in right now.

10:24Jon Krohn:And so every CIO, CEO, CTO, they want to be spending money on agents, probably they feel like that's the sexiest thing that they could be spending the money on. How do you convince them? How do you make the ROI case, the return on investment case for good data management when the big shiny thing that they want to be spending on is the agents? That's a great question. I mean, that really describes my job there, John. You know, the good news is - It's a good thing you're so shiny. The good news is as these enterprises, as they have invested in these agents and models, what they have realized is that just investing in the intelligence is not good enough.

11:09We have seen examples of companies where these agents have been made available to their customers, like support bots and things like that you see on the websites. And this very soon realized that all the benefits that they created this agent for are actually backfiring. Take example, in case of support chatbots, they tend to escalate a lot more back to the humans than that was happening before. they would take users down the wrong path, provide guarantees that the organization cannot meet. All of these come back to a human. And then we have word for those humans who actually have to pay for the sins of the model.

11:53These are the sin eaters. So eventually model does the sin. Humans who have deployed the models are the sin eaters and who have to pay for it as well. So organizations in 2026, realize that how important it is to give them right data sets. There are now tools available that can show when the model is going wrong. We are creating those tools in Agent Force 360 that was announced today, the enterprise AI harness, and giving users that transparency. Okay, this is the case where model does not even have information to answer properly. Can you give them the right data sets? We are providing all these kinds of tools to organizations to be in the with them.

12:32Jon Krohn:Regular listeners will already be aware that I'm obsessed with Anthropik's Fable 5 model, and it has taken over my working life. I'm writing a technical book that includes LaTeX files, mathematical notation, Python code examples, and Fable 5 and Cloud Code handles requests I make across whole chapters with accompanying Jupyter Notebooks end-to-end, work that a few short months ago would have been dozens of separate requests with way more manual fiddling required. With Fable 5, it just works, essentially like magic, first time. Claude is the AI for problem solvers. It's the collaborator that understands your entire workflow and thinks with you, not for you.

13:09Jon Krohn:Whether you're debugging code at midnight, building a financial model, or strategizing your next business move, Claude extends your thinking to tackle the problems that matter. For problems worth solving, get started with Claude at claude.ai slash superdata. That's claude.ai slash superdata. And check out Claude Pro. which includes access to all of the features mentioned in today's episode, claude.ai slash superdata. I love that. And then so basically what you're saying is, because my question was kind of, how do you make the ROI case? And basically what you're saying is that in 2026, it's maybe not as hard to make as it used to be.

13:45Jon Krohn:People are realizing that there are data gaps and they need to be spending money on getting it right. Absolutely. Just investing in agents and then only intelligence is not going to be the case. most organizations are realizing context is about 95 % of the battle. 5 % is intelligence and then choosing, but most of the battle for an enterprise is to get the context right. Love it. While researching for this episode, I found this stat that you seem to go to recurringly. You say that customers used to write three to four data quality rules in a good week. and now they use something called Clare to generate around 200 of these data quality rules per day.

14:25Jon Krohn:So first of all, what exactly is it? Is it data quality rule? Cause like, you know, I'm, I'm a data scientist. I've been doing this for a long time, but that isn't even something that's like an obvious term to me. I guess it's like a, it's kind of like a data management term that, you know, isn't, isn't my forte. So what's a data quality rule? how were customers writing these before? And now how has it like gone up two orders of magnitude in terms of the quantity that they're making? Sure. So data quality rule, you can think of them as automated pieces of code that check whether the data that is feeding the agents or feeding an analytics model is of high quality.

15:00Jon Krohn:Basically, those kinds of flanks you were talking about in your previous answer, where like an alarm goes off and says, we don't even have the data for this kind of question. Exactly. So things like that. for example, let's say we are looking at profit, right? And then we have the revenue data and we have the cost data. Profit should always be revenue minus costs, right? So you can give this as a rule to automated system that always checks for all the entities that we have across the world. Profit should always be revenue minus cost. There may be some cases where it is not. And now the data quality rule fails and says, you know, there's something wrong with this data value over here.

15:39Now do this for tens, sometimes hundreds of millions of these data elements like profit that are in the enterprise, number of customers, numbers of employees, all these different metrics and KPIs that organizations track. So, and for each metric, you create 10, 20 different data quality rules because you are looking at all the different dimensions of it as well. So now you have large number of data quality rules to manage as well. So what we've done with Informatica's Clare, which is an AI engine that we actually launched way back in 2018. At that time, there was no generative AI. So, you know, it used to have machine learning algorithms and things like that.

16:19But now it works off these technologies and agents as well. you can now ask it to generate a data quality rule for profit. You can give it subject matter expertise in terms of, oh, you know, profit should always be revenue minus cost. And it generates the data quality rule. It generates test data to validate it. It makes sure that the data quality rule is okay and validates it like an expert human would. And it can do it with an Excel file with a thousand data quality rules that you uploaded in an hour or so. So we're always making it a lot more productive for users to use Clare.

16:53Jon Krohn:Completely new world, which is interesting given that Clare is eight years old. It sounds like something really novel, but you were figuring out how to work the kinks out of it a long time ago. Absolutely. When we launched it, one of the biggest problems was, and it's still a problem, is I don't want to be keeping my sensitive data on these analytics platforms in public outside. I need to be scanning them, making sure that users' SSNs are not on some file that I uploaded to my S3 folder and things like that. So the machine learning algorithms that we added in Clare all the way back then was to scan these things and say, oh, these looks like SSN, right?

17:31Automatically mark them and then classify them so that users have this. Now, the same thing, you don't want to send these SSNs and credit card numbers and God knows all the sensitive data to an agent who can use it in various different ways that we do not know about.

17:45Jon Krohn:Right, right. Yeah, social security numbers, not the thing we want our agents to be processing for sure. And yeah, so Claire, I'll have a link to that in the show notes, but it's C-L-A-I-R-E. And I suspect the AI is part of, what made that such an attractive name choice. For data scientists listening, you know, kind of our core listener is a hands-on practitioner. Today, they're very likely to be an AI engineer, actually. But for those folks listening, what kind of skill or habit now matters more that we're kind of in this agentic era and that agents are consuming so much more data than people are?

18:24It's a great question. I would say three things. As an AI engineer, one should know completely about how evaling these models happen, extracting traces from what evals have been created to understand where these models are failing and making them better. The goal should be that you're creating an automated system that improves on its own. By looking at creating these first set of evals, looking at where the models are going wrong, and getting the right data, you can set the model up for that path. Second, we talked about data itself quite a lot. Data is the lifeblood of all of this, making sure that the right context reaches the right model when it has to answer a question.

19:10And third, very, very important, is also the economics of it, right? And it's not just the accuracy of it, but also the economics. What is the best path for the agent to be able to answer the questions that I'm creating the agent for, right? And that may mean that, you know, it should not open up all the labels. You know, we label the right things so that it reaches to the right data sets very, very fast as well. So the token economics is going to be another big thing that AI engineers need to look about. So - Great answer.

19:40Jon Krohn:Yeah, you're about to reel them off. Perfect. So, you know, the context, context is king, right? And then it's very, very important to get that right evals and traces. And number three, you know, to be able to do token economics better. So that is. I love it. I can't believe you had those all just ready off the top of your head. Gaurav was not prepared for any of the questions that I asked him today. We basically, he had just enough time to kind of like sit down and hit the record button, which means that I didn't get to warn you for my penultimate question, which is always the same. And it's, do you have a book recommendation for us?

20:16Oh, that is a great question. A book recommendation for data scientists and engineers. I would then give you such an esoteric book there, John. It's a book called The Mind Illuminated. The Mind Illuminated. Yes. It's like a meditation book. It's for all the times that you are overwhelmed with all the takeoff of these AI and AGI models that are happening right now. You get overwhelmed by it. Mind Illuminated will teach you how to calm down, get the light back into your mind as well. So it's just amazing.

20:53Jon Krohn:I need that. I need that for sure. I do have a daily meditation practice. And the days that I actually like spend like half an hour sitting doing it, as opposed to like kind of, you know, I'll do it every day. But sometimes it's like five minutes while walking the dog or something. And that's, you know. I'm hoping, you know, these AI models, we get a lot of time walking the dogs and doing meditation as well. I can't wait. And then my final thing is, so how should people follow you after this episode? You gave a great interview, really enjoyed hearing from you. How can people follow your thoughts in the future?

21:22Oh, I'm on Twitter and on LinkedIn. Please search for Gaurav Pathak from Informatic or Salesforce, and you'll be able to find me easily.

21:29Jon Krohn:I'm sure we'll have it in the show notes as well. Gaurav, thank you so much for taking the time out of your busy day here at Dreamforce 2026 in San Francisco. It's been a treat. John Krohn, same thing here as well. Amazing questions. Thank you. It's a great talking to you. Thank you. Awesome. Great episode today. In it, Gaurav Patak detailed how metadata are like the labels on supermarket cans. Without them, an AI agent has to open and smell every tin to find the metaphorical tomato soup, burning tokens across what can be hundreds of thousands of data stores in a large enterprise. He talked about why bad data used to cause a loud data brawl with three people bringing three different numbers to a meeting, whereas today an agent picks one number and delivers it with confidence.

22:10Jon Krohn:That's garbage in, gospel out. And finally, he provided his three priorities for AI engineers. One, evals and traces so that systems improve on their own. Two, getting the context right to the right model. And three, mastering token economics. I hope you enjoyed the conversation today to be sure not to miss any of our exciting upcoming episodes. Subscribe to this podcast if you haven't already. But most importantly, I hope you'll just keep on listening. Until next time, keep on rocking it out there. And I'm looking forward to enjoying another round of the Super Data Science podcast with you very soon.

From the publisher

During their #sponsored discussion, Senior Vice President Product Management AI and Metadata at Salesforce, Gaurav Pathak talks to Jon Krohn about why AI agents need well-labeled, high-quality data to deliver reliable answers in the enterprise. Listen to the episode to hear Gaurav Pathak talk about the difference between a “data brawl” and “garbage in, gospel out”, who the “sin eaters” of enterprise AI are and the three skills that matter most for AI engineers today!

Additional materials: ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠www.superdatascience.com/1030⁠

Interested in sponsoring a SuperDataScience Podcast episode? Email natalie@superdatascience.com for sponsorship information.⁠⁠⁠

In this episode you will learn:

(03:05) Why metadata are the labels AI agents need

(06:31) From “data brawl” to “garbage in, gospel out”

(10:57) Who the “sin eaters” of enterprise AI are

(13:45) What data quality rules are and how CLAIRE generates them

(17:21) Three skills AI engineers need in the agentic era

More from Super Data Science: ML & AI Podcast with Jon Krohn

All 130 episodes
1030: Garbage In, Gospel Out: Why Agents Need Better Data, with Salesforce's Gaurav PathakSuper Data Science: ML & AI Podcast with Jon Krohn · 23 min
Listen in VO