In short
August 2026 “In Case You Missed It” roundup of three prior conversations on AI adoption, agentic decision-making, and enterprise enablement.
Key claims
AI ROI gaps persist despite AI steering committees because teams pick the wrong success metrics and deploy agents without measurable workflow/process acceleration. Employee-facing, human-in-the-loop agent use cases are easier to secure and to attribute ROI. Replace vanity metrics like lines of code/tokens with business outcomes (e.g., faster idea/bug-to-production). In high-stakes decisions, LLMs can violate hard constraints; mathematical optimization should enforce constraints while agents help formulate problems and call solvers. Semantic layers matter for analytics agents to avoid re-deriving definitions and wasting tokens.
Guests
Pete Johnson (MongoDB field CTO of AI; 19 visits across six countries in 2026). Jerry Yurchison (Garobi Optimization manager of decision intelligence strategy). Priyanka Vergadia (Cloud Girl; former Microsoft senior director of AI transformation; Google head of North America developer relations; author). Tristan Handy (DBT Labs founder/CEO).
Notable examples
NRF executive dinner ROI agents; Uber token maxing; Meta token scoreboard; tool-not-called experiments attributed to Sinan Ozdemir; Siemens-scale semantic layer need; Stanford report showing ~8–9% production AI use cases.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOUnderstanding AI ROI Gaps
0:45 to 3:14
Exploring why many organizations struggle to see AI ROI despite having structures in place.
“and he traces his answer back to an executive dinner in New York, where every company seeing real ROI, return on investment, shared two common traits.”
Effective Use Cases for AI
3:14 to 5:28
Discussion on selecting appropriate problems and metrics for AI implementation.
“because if you put bad data into an AI ecosystem, it's not going to solve that.”
The Dangers of Vanity Metrics
5:28 to 6:31
Examining the pitfalls of using ineffective metrics like lines of code for productivity.
“When you measure it in terms of those business metrics, that's really where you see the improvement in overall ROI value.”
The Rapid Evolution of AI Technologies
6:31 to 7:54
Discussing the fast-paced nature of AI development and its implications for engineers.
“I think Meta had a similar kind of story.”
Staying Updated in the AI Field
7:54 to 9:20
Insights into how professionals can keep up with the latest trends and technologies in AI.
“we've just been talking about, I think it might be interesting for listeners.”
Trusting AI for Decision-Making
9:20 to 14:01
A critical look at the reliability of AI in making important business decisions.
“Pete's answer is about picking problems you can measure.”
Understanding Mathematical Optimization
14:01 to 17:31
Explore the importance and structure of mathematical optimization in decision-making.
“And then all of a sudden you put you put into production a solution that that is not good.”
The Role of Skills in AI
18:33 to 24:10
Discover how defining skills can enhance the effectiveness of AI tools.
“My next guest comes at it from the people side.”
The 10-20-70 Framework for AI Success
24:10 to 28:00
Learn about the 10-20-70 framework for budgeting AI resources effectively.
“And I know this sounds crazy, but I work with enterprises that have large number of large teams and have bought the tools and don't see ROI.”
Understanding the Semantic Layer
28:00 to 28:52
Learn about the significance of the semantic layer in data analytics.
“and pour into it, which is why I say, if you spend 70 % on some of this stuff, which is going to be very costly and hard for a CFO to agree to, but that's the only way to build a habit and an effective habit.”
Show all 13 chapters
Organizational Challenges and Solutions
28:52 to 30:14
Explore how the semantic layer addresses organizational data challenges.
“The problem that the semantic layer solves is not a problem for small organizations.”
The Role of AI Agents in Data
30:14 to 31:21
Discover how AI agents interact with the semantic layer for efficient data handling.
“and then stored so that successive people, when they ask those questions, can confidently measure things in the same way twice.”
Layers of Meaning in Data
31:21 to 33:18
Understand the different layers of meaning in data, from raw to ontologies.
“And oftentimes they do one of two things or almost always they do one of two things.”
Transcript
Automatic transcript. May contain errors.0:00Jon Krohn:This is episode number 1024, our In Case You Missed It in August episode.
0:09Jon Krohn:Welcome back to the Super Data Science Podcast. I'm your host, Jon Krohn. This is an In Case You Missed It episode that highlights the best parts of conversations we had on the show over the past month. My first clip is from episode number 1017, where I speak with MongoDB's field CTO of AI, Pete Johnson. Here's the key context. Four out of five organizations have AI steering committees and success metrics in place, and yet only one in five is seeing a return that matches. Pete has an unusually wide view of why, having made 19 stops across six countries this year, and he traces his answer back to an executive dinner in New York, where every company seeing real ROI, return on investment, shared two common traits.
0:55Jon Krohn:In a recent article, you highlight that 83 % of organizations have AI steering committees and success metrics, yet only 19%. So roughly four out of five organizations, and we're probably talking like kind of bigger enterprises in this case, but four to five of them have AI steering committees, success metrics. So they're trying to do something with AI, but only one in five, 19 % see matching ROI. So there's this gap where three out of the five organizations seem to have organizational structures set up to succeed with AI, yet they aren't. So how does that AI ROI gap happen? It comes from a couple of different places.
1:38What I take this back to is last fall, the mainstream media published a series of articles that asked the very fair question, where's the ROI for AI? Collectively, as an industry, we've spent all this capex developing the models. Where's the benefits? And I was at an executive dinner in New York City the first week of January, the NRF, the big retail conference that's there every year. And it was the first time I heard a customer talk about that they had agents deployed in production and they were getting ROI out of them. And they came with two caveats. Caveat number one was they were employee-facing.
2:17And caveat number two is that they were not autonomous, but they were human in the loop. The reason for choosing employee-facing use cases was twofold. Number one, the data security and data quality bar is lower than it would be for customers. Like it's terrible if you accidentally leak someone's salary to another employee, it's way worse if you leak some customer's data to the wrong customer. But the other thing has to do with the ROI part of your question, which is, I know how I'm judging the effectiveness of an employee. If I take an agent and I put it in their workflow and I see those metrics jump, I can attribute that jump to the agents and therefore or back of the envelope, compute some ROI.
3:06So that's a very long way of saying, like picking the right problem is key here. You got to pick a problem that you have good data for, because if you put bad data into an AI ecosystem, it's not going to solve that. And number two, you have to have some metrics of success. You can't just throw AI at a problem that you don't know how difficult it is, or you don't know how to measure the success of it. And that's why these employee-facing use cases were so popular in the first half of 2026 is we already know what metrics we use to bonus people and to evaluate their performance. And like I said, if you put AI in their ecosystem, within their workflow, and notice the change in those metrics, that gives you an ability to then measure that ROI.
3:58So more and more companies are beginning to do this. But when those statistics came out in that survey, that was the problem is people were picking the wrong problems.
4:06Jon Krohn:Right. So instead of having your success metrics be related to how many employees have adopted AI, that kind of thing. Instead, the success metrics should be how much does an existing process get accelerated by having some automation of workflows within the loop? Exactly. And the best example I've seen here in sort of the second quarter of calendar 26 has been software delivery life cycles. So at the beginning of the year in January, we might have measured the success of something like Codex or anti-gravity or Claude Code by the lines of code. How much code are you producing? And if you've been writing software as long as I have, and even if you haven't, you probably know lines of code is a terrible metric to determine the effectiveness of an individual or in this case of some AI.
4:59But the maturity that I've seen really in the last eight to 12 weeks, really since the Uber token maxing story came out, was am I shipping code faster? That it's not just about the effectiveness of writing code or the effectiveness of an individual engineer, but the broader business metric is have I increased the speed it takes me to go from idea or bug report to production solution getting deployed? When you measure it in terms of those business metrics, that's really where you see the improvement in overall ROI value.
5:38Jon Krohn:Makes a lot of sense. And yeah, the token maxing thing was so silly. I did an episode dedicated to that because I'm basically encouraging my listeners and managers not to be using lines of code or number of tokens as a way to measure productivity because it's so easy with agents to just have them spewing things out and be using tons and tons of tokens for sure. Yeah, I think we've seen, it's funny that that, oh yeah, the token overrun that you're talking about, it was something like Uber's expected allotment for tokens for the entire year 2026 ended up being used in three or four months or something like that.
6:18Jon Krohn:13 weeks. 13 weeks. Exactly. Yeah. Wow. Yeah. I'm glad that organizations are waking up to the dangers of having that kind of silly metric, vanity metric to track. I think Meta had a similar kind of story. They did. Meta had their own token scoreboard that was measuring the number of tokens that individual engineers were consuming. And like you said, that turned out to be a bad idea. And among the things that's interesting about this AI wave of technology compared to others that I've seen in my career is how fast the cycles are. Like we, Anthropic dropped model context protocol on the Monday of Thanksgiving week in 2024.
7:02And by March, all of their competitors had embraced it. And here, I mean, token maxing really got its shining moment in March at GTC when Jensen mentioned it on stage. And by the time the Uber story came out like eight weeks later, nobody was talking about token maxing anymore. So like the life cycle of these stories and how quickly we move on to new things is unlike anything I've ever seen in my career.
7:28Jon Krohn:For sure. It seems like being in the AI space, we're going to soon have like a 24-hour news cycle just for what the popular packages and approaches are in our discipline. We're getting there. And as an engineer, that's the hard part is how do I keep up? How do I know where I should invest my learning? And how do I know when to sort of cut bait and move on to the next thing? That's the hard part about being an engineer in this ecosystem right now. Pete, this isn't a question that I had planned for you, but I think it's, yeah, given what we've just been talking about, I think it might be interesting for listeners.
8:01Jon Krohn:It kind of begs this question, what you were just talking about. How do you personally, given that you always need to be right on the cutting edge of what's happening with AI, how do you stay up to date? Well, I'm fortunate in my job that I get to talk to so many different people. I've been calling it my AI world tour this year because I've visited multiple cities. So I've made 19 stops on my world tour in the first six months of 2026 across six countries. I've probably talked to 100 different companies. I get to listen. I mean, I don't spend, I spend a good amount of my time, you know, telling them about how MongoDB can help them, but I get to listen.
8:39Like, what are you doing? What are you seeing? So because of the diversity of the different organizations that I get to talk to, I get to hear not just what one person thinks. I get to hear what a couple dozen people think and use that as a guide for where I go spend my time. And it's things like I think we'll get into a little bit later, things like better agentic memory, things like, you know, not maybe not always calling the generative LLM and being a little bit more creative about how you use rag pipelines and why retrieval quality is important. I get to talk to people about that kind of stuff and let that be the guide for where I spend my time.
9:20Jon Krohn:Pete's answer is about picking problems you can measure. My next clip asks a harder question. Which decisions should you be handing to a language model in the first place? In episode number 1015, Jerry Yurchison, manager of decision intelligence strategy at Garobi Optimization, makes the case for mathematical optimization in the agentic era. An agent will tell you with confidence that it has optimized your business while quietly ignoring the one constraint that could cost you millions. Jerry explains what optimization does differently and where he thinks the division of labor between agents and mathematical solvers ought to fall.
9:55Jon Krohn:Speaking of the times changing rapidly, since you were last on the show last October, we're now constantly talking about agentic AI, obviously on a data science podcast that focuses on AI, which is more and more what data science is all about, I think. Certainly in the way that I've been curating content on the show and I've been experiencing the world. Yeah. And so, yeah, where does mathematical optimization fit into this new agentic world that we're in? From our perspective, that's the million dollar question, billion dollar question. Yeah, it's more than a million. Yeah, a million dollar question because that's the term that people used to use a lot.
10:36I don't know. But yeah, it's billion. It's massive now. But anyways, it's when you think about what and sort of I'm going to be a little bit sort of high level and not absolutely correct. But if you think about what an LLM does, LLM takes all the input tokens and then just produces more output tokens about, you know, text or something like that. Let's say if you're purely natural language type of stuff, like input tokens, your output tokens, that's it. That's all it really cares about is providing the tokens that sort of give you a really good response, a highly likely response. So something that is all about just that sort of input-output flow.
11:20That really doesn't jive with what I was talking about, the types of decisions that optimization can make and what it does and what it and sort of the rigor that it provides is it provides these you know the constraints are what we call hard constraints these are things that cannot be violated it's not like oh my context you know that i i you know i mentioned my constraints just outside of like a context window type of thing and now now the element is
11:45Jon Krohn:sort of forgetting this yeah or even if it's in the context window it's still very a very frequent occurrence that some piece of information that you say, you know, there's an example, Sinan Osdomer. Do you know that guy? No, I don't think so. Sinan Osdomer. I think he's been on this podcast more than anybody else. And he's a crazy prolific author of data science and AI books. He's, I think he's younger than me. He might be in his mid thirties and he's written at least 10 books and he's created tons of online content. Recently, he started working at fireworks AI. But the point that I'm getting to is that Sinan Ozdemir, I've seen him do a talk a couple of times where he shows surprising issues with even frontier LLMs, where he would do something like have a tool available for an agent to call.
12:40Jon Krohn:And in his prompt, he would say, you must use this tool. And it was like single digit percentages, but some single digit percentage of the time, that very simple, very specific instruction, the LLM controlling the agent just wouldn't, just wouldn't do it. It wouldn't call the tool. It would find some other way of doing the approach. And, um, so yeah, Sanon's done lots of, he, he did a whole book on agentic AI, uh, where these kinds of experiments that he was running, he published them in there. Um, I will, while you're talking, I'll look up the name of that book. Yeah. Um, yeah, go ahead. Yeah.
13:17If you think about exactly what you said right there, when it comes to, again, like the, the fate of my business, do I want to trust decision-making to something where that can happen, where it can forget a, let's say you have some sort of environmental constraint where, you know, if, if, if you violate that, then you're going to be fined, you know, millions of dollars, something like that. And then you output a solution that is like, you know, the one thing, you know, all these agentic tools and everything, the confidence is so high. It's like, it's like, I got your boss, like exactly just just what you're looking for.
13:56We're 100 percent good to go. But it misses this environmental constraint. And then all of a sudden you put you put into production a solution that that is not good. and then the bad things happen. And then all because of, of, you know, just forgetting, you know, just doing something that happening. And, and the contrast to mathematical optimization is if that is a constraint in your model saying that here's my, again, I'm going to do the, the arm thing again for everyone listening, uh, just purely video, here's my constraint. And all my decisions are in here. I cannot go past this. I cannot break this environmental constraint, you are guaranteed that.
14:36So that's what we think is a differentiator. First off, is that you have the trust of the model to actually do what it says. And what's really nice about it is what this line represents is something that you talked about. It is a constraint that the people who are designing the problem, who are talking about it, have hopefully agreed upon as an actual constraint. So it's not just like some generated like business rule type or something. It is something that as like if you and I were working on a problem, I'd be like, hey, I think this is an environmental constraint. You'd be like, yes, it is, but it should look more like this.
15:14And then we agree what that is. And then it's represented in there mathematically. So it's just a very different decision-making framework. But where I see this all fitting together is you started to mention that, okay, I have an agent that's going to call a tool. Okay. An agent should be able to, what an agent can do is help you develop the problem statement that you're really trying to solve, help you understand all of the, okay, all the other bits and pieces of all the other regulations. Say, hey, I have this environmental regulation and an LM or an agent can do like, hey, this is what, these are other things that you may want to consider.
15:51And you might be like, holy crap, yeah, I want to consider these. I forgot about them. So it can really help there. It can help you identify the problem, help you actually write the code, help you to come up with the mathematical formulation, do all of that kind of stuff. But it can't do the solving. It can't give you the optimal solution. It can't give you a solution that's close to optimal and have the defendability, the explainability, all that sort of stuff that comes with these high stakes decisions. So how we sort of see it as agents should be able to call, develop all those things and then call an optimization engine like Garobi saying, here is the problem statement that we have.
16:30Here is the model they're trying to solve. And then Garobi runs, gives you the output, sort of gives the solution. And then you can dive deeper into why. Why is this happening? You know, why did I decide to build a new production facility in Atlanta as opposed to Baltimore? or something like that. And those are actual questions that you can get answers to with mathematical optimization. You can, because essentially what can happen in that situation is the, it'll resolve with, with the Baltimore production facility there and say, this is the difference. It is a difference because the cost is going to be this much higher, or you're going to have this much less demand or something, whatever it may be.
17:12You can actually sort of figure those things out and, and get to be able to answer questions that people are going to have when it comes to business problems and decision-making is why this, why not that? Those are all things that can happen with mathematical optimization because of the structure and the rigor that's there.
17:30Jon Krohn:Regular listeners will already be aware that I'm obsessed with Anthropik's Fable 5 model, and it has taken over my working life. I'm writing a technical book that includes LaTeX files, mathematical notation, Python code examples, and Fable 5 and Cloud Code handles requests I make across whole chapters with accompanying Jupyter Notebooks end-to-end. Work that a few short months ago would have been dozens of separate requests with way more manual fiddling required. With Fable 5, it just works, essentially like magic, first time. Claude is the AI for problem solvers. It's the collaborator that understands your entire workflow and thinks with you, not for you.
18:07Jon Krohn:Whether you're debugging code at midnight, building a financial model, or strategizing your next business move, Claude extends your thinking to tackle the problems that matter. For problems worth solving, get started with Claude at claude.ai slash superdata. That's claude.ai slash superdata. And check out Claude Pro, which includes access to all of the features mentioned in today's episode. Claude.ai slash superdata. Those first two clips came at AI adoption from the technology side. My next guest comes at it from the people side. In episode number 1019, the cloud girl Priyanka Vergadia, best-selling author, former senior director of AI transformation at Microsoft, and head of North America developer relations at Google, walks me through how she structures Claude's skills so that her output stops being slop.
18:55Jon Krohn:She then lays out the 10-20-70 framework she uses to guide AI budgets, a split that will make any CFO wince. There's a specific thing now that we're back on Claude. Way back in the content creation section that we were in like 15-20 minutes ago, You were talking about how a big part of your success with using Gen AI tools, with Claude specifically, is having skills set up. And so that's something that I meant to at that time. We ended up going off and talking about something else. But I'd love to come back because I feel like that's something really practical and technical and useful for our listeners.
19:32Jon Krohn:Could you explain for our listeners who don't already know about skills, what those are and how they can make the most of them in Claude? Skill is something that you would define your task to be. So break down your task into small subtasks and you define how you do that. When I write the blog, I do a research. This is how I would do it as a human, right? So think about this when you're writing a skill, Think about it like a human. Now, the way Claude helps you do it is amazing. But before you even get into it and start setting up a skill, think about your task explicitly. I'll give you an example because it'll be more material that way.
20:18I'll do a research first if I'm about to write a blog. And on this topic, I would... And the research prompt has to be really good as well, where I want only high-quality content written by researchers and scientific research from schools and universities and organizations like this, only look for that stuff around this topic and then create a report. Let's say that is the prompt, but that becomes part of my skill as step one. Then the next part of that skill is after I do the research, so the prompt is the how, right? So I've put the how in there as well. Research is the task. How you do it is that prompt.
21:02Then the next step would be to synthesize that research. I usually do a human in the loop thing in there. I don't trust it to make decisions beyond that. So I would do a human in the loop. It sends me a text message. I've got all this set up in Hermes. So it sends me a text message saying, I've done the research, here's the doc. I like to read because I'm trying to do that intentionally so that I don't lose the skill with AI. But you can also, I've also done cases where it would send the audio to me and I can just hear what the research was while I'm running or while I'm going to work out, which is super handy, right?
21:45And then I would give it instructions on, okay, I want to change this or that. And then the next step is like kick off writing the first outline. I don't do, I don't, like most people would just go in, write a blog on this topic. You're going to get slop, right? The whole idea here is how would you approach a blog? I approach with research. Then I would go in and do my analogy and storytelling on top of it. That's my next step. Then I would synthesize after the synthesis, the storytelling, and then putting it into a format that I usually used to write blogs when I was doing it all on my own, right?
22:27Which is, I need to have three images in this blog, and they need to be developer-focused, which usually talk about the flow of movement of tokens or query. And I have some of these examples in there. And I need to have one practical example in there, and that can come from, there's prompts in there where Claude would ask me other practical examples from my experiences. And so that it can take those and do that. So I'm going into too much detail, but the idea of a skill is, how do you do the task? Define it into subtasks. Those are your bullets that go into the skill and also some example prompts that go in there.
23:12That way you will get a much more personalized, the type of outcome you would, it's never perfect, but at least like 80 % there. And now you can start editing it from there.
23:25Jon Krohn:Perfect. Yeah, that did have a lot of examples, a lot of detail, but hopefully it helps us understand just like a lot of your cartoons, your illustrations, you know, go into practical examples. And so we got lots of examples there of how to build effective skills in Claude. So thank you for indulging us with that. Back to the enterprise stuff with your Google Cloud experience, your GitHub Copilot experience at Microsoft. With your work bridging the gap between high-level boardroom strategy and real-world enterprise AI execution, you have something that you call the 10-20-70 framework to help guide budgets for AI success.
24:07Jon Krohn:Could you tell us about that 10-20-70 framework? So 10 % on tools, 20 % on execution with those tools, and 70 % on education and skilling and upskilling. And I know this sounds crazy, but I work with enterprises that have large number of large teams and have bought the tools and don't see ROI. This is exactly the Stanford report that just came out. The impact report is a great example. It has like eight or 9 % of the actual AI use cases in production. In real production use cases are about eight to 9%. Everything else is just like experimentation, right? And this is exactly the reason you're not going to see ROI.
25:03You're not going to see real use cases that are leading to revenue or cost reduction or savings, right? Because you've bought the tools. AI is a habit. And habits don't form in days. They form in an extended period of time. So this whole concept of token maxing and all of these things are just natural evolutions as well. Yes, the idea of token maxing is super weird because we went from, okay, we've got a tool. Now the AI officer in the company is like, nobody's using this. We got to make them use this. So you get into this whole problem of now everybody's using it for writing emails or the dumbest tasks.
Read the full transcript
25:53And now you're in this token maxing situation, which you never thought of, where it's like, oh my God, now we are spending so much money. on this tool and we're not seeing ROI. So you got them to use it, not effectively, and you're in a different problem. But I think it's all a good problem because you at least got them to touch it. When you look at 10, 20, 70-year-old, if you spend that 70 % of your budget and time on upskilling your employees, which means showing them effective ways of using the tool, not just telling them, use it. Showing them effective ways of using the tool, not just giving them training, but actually giving them real use cases of the thing they can do in their job.
26:43And that requires time and effort and energy. my DevRel hat on, that requires building a community and saying, I tried this thing today. Let me share it with you all. If you're a testing team, if you're a coding team, if you're a team that's product managers, got to share those experiences with each other. And you have to bring space for that as a leadership team to allow people to build and form these communities. And after like, this is a whole, a big J curve, right? And after six, eight months, you start to see effective use of tools, actually helping them be productive. And then you get to a point where it's like, now we can write test cases with this.
27:32That's looking really good. How do I write them faster, improve my prompts a little bit more? How do I take an an entire process and make that an agent. Now you have agents in each of the different business units and you can form that into a repository of agents. And now an entire company is becoming efficient. But this is a trajectory and the curve, you have to see the vision for a year or so and pour into it, which is why I say, if you spend 70 % on some of this stuff, which is going to be very costly and hard for a CFO to agree to, but that's the only way to build a habit and an effective habit.
28:16Jon Krohn:We're rounding up a great month with episode number 1021, in which DBT Labs founder and CEO Tristan Handy explains why the semantic layer matters more, not less, now that analytics agents are the ones asking the questions. You mentioned there in your most recent response how DBT adds meaning So that five trends like pipes and dbt adds meeting. There's a term that came up a lot in our research for dbt labs, which is semantic layer. Do you want to explain how dbt acts as a semantic layer for your data? This is a topic that is particularly hot right now as analytics agents are very in view. The problem that the semantic layer solves is not a problem for small organizations.
29:05So if you imagine that you're a part of a whatever, a 20 person company, a 50 person company, like you probably don't need a semantic layer. But now imagine that you are Siemens, you're a global company, you have 2000 data engineers that are serving 300 ,000 employees globally. It is not possible to just know the answer to random questions that you might need to know the answer to without getting meetings together of people that you search for in your Outlook phone book. You get everyone together in a room and you say, how should we be measuring this thing? and there's a lot of conversation and you've got to everybody's got to figure it out and which table should we be using and all this stuff and and literally that's how big companies have for the past whatever 30 years tried to answer questions like this that's that's the process and the semantic layer is tooling that allows those types of decisions to be made and then stored so that successive people, when they ask those questions, can confidently measure things in the same way twice.
30:29They don't have to reconvene the whole group. It is a technically complex problem area, but the problem that it solves is really an organizational problem. It is how do you scale knowledge to increasingly large groups of people? And that is honestly a tale as old as civilization. I mean, we don't have to go too deep on this, but like, as long as people have been organizing together to like figure out how to do stuff, there's been this question of, well, how do we make sure that we know things and that everybody in this organization knows a consistent set of things? And so the semantic layer is an attempt to do that.
31:09And it's particularly relevant today because absent these types of cues, how do you measure X thing, AI agents have to try to re-derive that for themselves at every turn. And oftentimes they do one of two things or almost always they do one of two things. One is that they, in re-deriving how do you measure something, they just take a tremendous number of tokens to do that. They just have to look into a lot of stuff and think a lot, and that becomes slow and expensive. And then the other outcome is that they just get it wrong, or they come up with an answer that may be kind of reasonable, but it's actually not the way that your organization measures these things.
31:53So the semantic layer extends very, very nicely into a world of agents.
31:57Jon Krohn:Yeah, really cool. It seems like one of the key use cases. And if people aren't aware, that word semantic basically just means understanding, just means like the underlying meaning of some data. And so by having this common playing ground, or this common lingua franca, this common agreement on what meaning is across data sets, across agents, there's efficiencies, especially across the large organizations that you were describing. They're like 200 ,000 person companies with 2 ,000 data engineers, that kind of thing. Yes, totally. And there are these successive layers of meaning where you start off with raw data, you go to modeled data, then you go to semantic layers, which are typically, they help you understand how to join these tables together and how to measure certain metrics.
32:55But then you can even go one step further. And this is not an area of particular expertise for me, but it's a very interesting conversation happening in the industry is ontologies. So ontologies are another layer of meaning making on top of data that are not just how do you measure a thing, which is typically how we think about the semantic layer. But it's also how do you model business processes? How do you model causation? And I think that most of us do not operate in an environment where it's appropriate to say we need ontologies for this because the world changes quickly and sometimes it's hard to keep up.
33:39But in very specific high value domains, think like genetic research or like famously, this is employed. Ontologies are employed a lot in defense. So I think about these four layers as kind of like the knowledge or meaning hierarchy. All right.
33:57Jon Krohn:That's it for today's In Case You Missed It episode. Be sure not to miss any of our exciting upcoming episodes. Subscribe to this podcast if you haven't already. But most importantly, I hope you'll just keep on listening. Until next time, keep on rocking it out there. And I'm looking forward to enjoying another round of the Super Data Science Podcast with you very soon.
From the publisher
In ICYMI Episode #1024, Jon Krohn tracks the gap between AI investment and AI return, from the technology side to the people side. Hear from Pete Johnson, Jerry Yurchisin, Priyanka Vergadia and Tristan Handy, discussing why four out of five organizations have the structures for AI success in place while only one in five sees the returns, which decisions should never be handed to a language model however confident it sounds, how to structure Claude skills so that your output stops being slop and why the semantic layer matters more, not less, now that analytics agents are the ones asking the questions.
Additional materials: www.superdatascience.com/1024
Interested in sponsoring a SuperDataScience Podcast episode? Email natalie@superdatascience.com for sponsorship information.
In this episode you will learn:
(00:56) Vector Search, Agentic Memory and Effective RAG
(09:20) Mathematical Optimization in the Agentic AI Era
(17:30) Anyone Can Write Code Now, So What Gets You Hired?
(27:14) How dbt Won Analytics Engineering




