Context-Aware SQL and Metadata with Shinji Kim

4 Sep 2025 · 42 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Podcast Episode Notes: Context-Aware SQL and Metadata with Shinji Kim

Episode Overview

  • Podcast Title: Software Engineering Daily
  • Episode Title: Context-Aware SQL and Metadata with Shinji Kim
  • Episode Description: The episode discusses challenges in data-rich organizations regarding metadata capture and the complexities of data usage. Shinji Kim, founder and CEO of SelectStar, shares insights on their data discovery and metadata platform that creates a knowledge graph, improving data accessibility and AI interaction.

Key Concepts and Discussions

Challenges in Data-Rich Organizations

  • Capturing and maintaining context about data is difficult.
  • As data models become complex, finding relevant datasets can lead to bottlenecks.
  • Reliance on outdated documentation and tribal knowledge hampers data usability.

Introduction to SelectStar

  • SelectStar: A platform for data discovery and metadata management.
  • It constructs a continuously updated knowledge graph by analyzing the data structure and usage.
  • Enhances data with context like popularity, lineage, and semantic models.

Importance of Metadata

  • Metadata provides critical context for understanding data's relevance and trustworthiness.
  • Key components of metadata include:
  • Core Metadata: Basic asset names, descriptions, operational metadata.
  • Usage Signals: Popularity metrics, lineage tracking.
  • Business Context: Semantic modeling, collections, and governance tags.

Use of AI in Metadata Management

  • AI enhances the accuracy of SQL generation by leveraging enriched metadata.
  • SelectStar uses AI to automate the construction of semantic models and the analysis of query logs.

Semantic Layer Explained

  • A semantic layer describes a logical data model, outlining tables, columns, dimensions, and metrics definition.
  • It separates verified data sets, allowing for better governance and usability in reporting.

Data Discovery and Popularity Metrics

  • Popularity metrics help users identify the most trusted and frequently utilized datasets.
  • Popularity and lineage metrics can lead to cost savings by identifying unused data assets.

Future Trends and Challenges

  • Organizations need to address metadata management to leverage AI effectively.
  • The emerging focus on metadata is crucial for operationalizing AI and maintaining data accuracy.

Context in AI and Data Interaction

  • Proper context helps AI generate precise queries and understand data relationships.
  • The evolving landscape of data storage (e.g., cloud services) necessitates solid metadata practices.

Conclusion and Future Insights

  • The episode emphasizes the growing importance of metadata in the era of AI.
  • SelectStar aims to automate and streamline the processes of metadata management, improving overall data governance and usability.

Key Takeaways

  • Effective data usage hinges on capturing and maintaining metadata to provide context.
  • AI can significantly enhance SQL generation and data discovery when guided by robust metadata.
  • The future of data management involves integrating metadata solutions to support AI applications and improve operational efficiencies.

---

Additional Resources

  • For further information, visit the [Software Engineering Daily](https://softwareengineeringdaily.com) website.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00A common challenge in data-rich organizations is that critical context about the data is often hard to capture and even harder to keep up to date. As more people across the organization use data and data models get more complex, simply finding the right data set can be slow and create bottlenecks. SelectStar is a data discovery and metadata platform that builds a continuously updated knowledge graph of an organization's data by analyzing both its structure and how it's actually used. It enriches data with context such as popularity, lineage, and semantic models, making it easier for AI and teams to discover, trust, and use the right data.

0:42These enriched metadata layers are also highly valuable for large language models, significantly improving the accuracy of generated SQL queries. Shinji Kim is the founder and CEO of SelectStar, and she joined Sean Falconer to discuss solving metadata curation challenges, managing data context at scale, using LLMs for SQL generation, emerging trends in metadata management, and more. This episode is hosted by Sean Falconer. Check the show notes for more information on Sean's work and where to find him.

1:29Shinji, welcome to the show. Thanks, Sean. Great to be here. Yeah, I probably should have said welcome back since you've been here before, although it's been a couple of years. Yeah, more than three years ago to introduce SelectStar. But I am really excited to be back. And software engineering daily has always been, yeah, also morphing and changing a lot. Yeah, well, it's been three years. So why don't you catch us up? I mean, three years, especially in the world of tech, the world of startups, and now what's increasingly becoming the world of AI is a lot of time, a lot could happen in three years.

2:02So what's happening with SelectStar today? maybe go back even to the beginning, sort of what's the story behind where you guys started and where are you today? Amazing. Sure. Yes, so much changed. I started SelectStar five years ago after noticing time and time that a lot of enterprises collect, store, and process data. But to try to use the data, it takes days or weeks to find the right data and actually use it properly. You have to rely on outdated documentation. Usually you need to just find somebody else, rely on travel knowledge to understand how to use the data. I mean, this is something that I saw firsthand at Akamai when I was running the product for their IoT data processing, partnering with consumer electronics and automotive enterprises, building their next consumer applications.

2:58they were looking to pull a lot more telematics data. And especially in enterprise perspective, this was an issue. And hence, there are solutions like traditional enterprise data catalogs that are trying to solve this issue. At the same time, I've noticed that there was a lot more demand around this. Also, as more companies are adopting, quote unquote, modern data stack of cloud data warehouses than building their data lakes on the cloud with Snowflake Databricks, data discovery, finding and understanding data has been a lot wider issue in organizations. So that's where SelectStar is really focused on.

3:41We provide a very easy-to-use UI, now MCP server, APIs, Chrome extensions, Slack app, all different places where end users, so whether you are a data scientist, data analyst, software engineer or product managers, whenever you have to touch or see data or data products, you can easily access the context about that data, documentation about their data, where did the data come from, who else is using this inside the company, what other data assets or analysis are already attached or have been built on top of. So there's a lot of, I would say, we are almost like drawing a knowledge graph for you in terms of how your data assets are connected and utilized inside the organization today.

4:32So that's the core of what we do. Why do you think so much of this kind of like metadata has historically been kind of this like tribal knowledge? Why haven't we been focused on capturing that as part of the data we collect? Like we built so much technology for actually collecting data, But then this kind of stuff about why the data exists, how it relates to each other, we've historically, I think, just relied on communicating within the company to ask people, why is it this way, rather than encapsulating that in some sort of piece of technology. Yeah, I mean, correct question. I can just go back to that no one likes documentation, especially, I think, developers.

5:12Most of the databases does not have table column comments. And it just follows the code. A lot of the data tables have descriptive names. But I think today also more so the proliferation of data models and how easy it is to transform and build your own data models, I think it also adds to that. continuing to writing manual documentation doesn't scale. And it's more of a always taken as a after the fact. So in the beginning, when you are starting off from scratch, you will have entity relationship diagrams as part of like modeling the data. But afterwards, as you are, or as you have more people building different types of domain models on top of the data, I think it gets lost very quickly.

6:08Now, I think the metadata collection in that sense of it, a lot of companies have their own internal tools where they just refer to information schema. That's where most of the metadata resides in and where most of the data catalogs really depend on. But the core part of where we focus on at SelectStar and now more modern systems focus on is really what happens in between the data assets. So who is accessing the data? Which query is accessing this data? And how is this accessing the data? These are the parts of the activity information that I would say, if you can parse them through and look at them in aggregate, that analysis of metadata is something that's very valuable.

6:56And that is the full system that we built around. So like any data warehouse that we connect to, we will parse through all of the activity logs or SQL query logs. to understand how is the data actually being created all the way to where it's being used. And also, how is this accessed? Which type of select core is coming from? Which applications are querying the data? And how often is this being queried by how many unique users in the last certain time period? Which helps us to understand the trends of the data usage as well. I think these are the parts that I would say hasn't been looked at as much.

7:37But as there are more consumers of data and there are more usage of data, I think that there is more need to understand this. The other, I think a big part of it is that a lot of companies have now moved to, it's been now easier than ever to have all of your data in one place in your data lake, data warehouse, or to create data mart system. There are so many connectors of all different SaaS tools and business tools that can share that underlying system of record data into one place so that you can join them and then model them on top. So I think, yeah, it comes from multiple places. But I think in the past, when we used to rely on primarily relational data warehouses or more like a Hadoop based systems, there were a lot less number of, I would say, consumers of data directly.

8:33And this is probably maybe why metadata hasn't been as the main highlight that people were looking into primarily. Yeah, I mean, I think your explanation of no one likes documentation is a good one. I think even if someone starts out sort of documenting these things, it's just sort of inevitable that it gets stale over time. It's just every company has the best intentions with a lot of this stuff, even when it comes to like coding or how certain functionality works, internal tools. We all have internal wikis there where we have documentation that's multiple years out of date. So you described kind of looking at the activity log.

9:12So from the activity log, are you kind of like reverse engineering what the relationships are between the data based on how queries are run against it? Yeah. So basically, the way that we look at the metadata is that we look at each of the queries that are coming through, and we attach them to each mentioned assets. And then we run a separate analysis on top in terms of how often this has happened or how many unique users run that through. Did I answer your question? Yeah, it sounded like to me that you're sort of inspecting what the actual behavior of individuals within the organizations or applications that use the data are actually utilizing the data to figure out what the actual knowledge graph behind the data is.

10:02How are different concepts related based on the query execution? Yeah, so we can see like a certain amount of information about the user. We may see the username, but we may not know who that user is, which team they belong to, so on and so forth. That would come from other places, whether if we were to connect to Active Directory or having our customers to group their users and so on and so forth. But the main piece of where we are putting together this knowledge graph just primarily comes from tracking the usage. So which tables are joined together? What's the joint condition look like? What are the most used to the least used tables and within those tables columns look like?

10:48This actually gets a lot more interesting when you connect it to other applications like BI tools. for Power BI and Tableau or Looker for this sales dashboard that a lot of people are relying on, which are the specific fields and tables that really power them? And how are each of KPIs being measured, actually defined or calculated? So I think there are multiple steps of sort of insights that you can get. So we see it kind of in like three levels. So once we connect and ingest the metadata and query logs, there is like first layer, which is the core metadata, just the physical asset names, descriptions, the operational metadata of how big the table is, or things like when's the last updated, things like that.

11:40And then on top of that, there is the second level of usage and behavior signals. So this would include things like popularity. How widely is this being used and trusted? And this also would include entity relationships and lineage. Where did the data come from? Where does it go to? And how is this data model related to one another? What are the common queries and joins that's related to this asset? And then there's the third level that we also see that primarily will be driven by the users, but we will help automate, which would be mostly around business context and semantics. So this will be something like collections.

12:21If you were to group them for a certain business domain, what would that look like? Any tags that we can infer or actually put in so that you can actually govern the data. Having business glossary and metrics definition, a lot of this is what we see also as part of now the metadata context that you can put on top of physical assets so that you have a lot richer context for any of the access or whenever you're trying to leverage data that you have access to. For things like popularity, usage metrics, how are those used and what is the value of tracking those for like an organization that's using SelectStar?

13:05So usually what we see the most popular or most interesting have been leveraging both popularity and lineage. The first thing I can think of is just there's always a lot more added benefit when you have customers starting to realize what is the right data to use. Because it's not just about the semantic relativity, it's the trust score. If I'm looking for data related to active users or our sales regions, the types of data that you want to use would need to come from data sets that other people are also using, right? So popularity really comes in handy for that. And if you were to also leverage lineage with that, then you can also see what other impact that the data also has in other parts of the system or across the system.

13:57This is actually also interesting when you are thinking about cost perspective of running a data infrastructure. We have a number of customers that have a saved cost on their cloud billing, primarily for their warehouse billing, by looking at what is the popularity, meaning they noticed that a lot of models or tables that they have that they thought were being used, but they weren't. They weren't either being queried or they are there to load reports on the BI system, but the BI dashboards actually weren't being viewed by the end business users. So combining both lineage and popularity, there is a big understanding cost implication to that.

14:42Okay. What are some of the other use cases, not necessarily restricted to the popularity and usage, but to SelectStar in general? I would say the number one use case for us just always comes from data discovery. This is something that also correlates really well with how our customers are using SelectStar with their AI agents and doing data work with AI. It's really just providing the right types of results when you are trying to, let's say, build a new model, edit a SQL query, or do exploration around data. The popularity score will allow the agents to be able to find and use the right type of tables and columns.

15:28It will also be able to provide example queries that are relevant so that the AI agents can actually build queries that are a lot more accurate. And I would say this is like a very much of a use case that flew in from or more of a native or next generation from what our end users used to do. Our end users that are in data teams used to come to select our UI to find and understand data so that they can query those data tables directly or build dashboards. Now, new use cases that we're seeing is that it's their agents and AI tools that's using our MCP server to find and create queries and model modification directly.

16:16Okay. Going back to the original problem that we're talking about, the fact that people haven't historically had a good way of really capturing all this metadata and relationship information or at least sort of keeping it up to date. And maybe historically we've been able to get away with that in some capacity, but do you think now when we start to enter this world of people wanting to have AI agents that can, in some capacity, leverage this huge amount of data they're collecting. And part of them being able to effectively leverage that data is they need to be able to understand it, which gets back to like, sort of the knowledge graph, the metadata is associated with it.

16:54Does that make these pain points, like, I don't know, elevate them to a place where this isn't just kind of an admirer annoyance now, this is actually something where a company's like, essentially not going to be able to move forward and leverage all the greatest innovations that are happening in AI until they solve this fundamental problem. Yeah, the way that I see this related to AI or quote-unquote hydrating AI with enterprise data, trying to use AI on top of your own data. Today, the ways that this has worked in the POC environment has always been by putting the very specific schema information, query examples, synonyms, so on and so forth, almost like a build-your-own semantic layer in order for AI to work well.

17:43Or, yeah, actually like to train the model with just those data specifically. This is an approach that I would say kind of gets you to 90%, but is very hard to scale. Without a metadata platform that will continuously evaluate the schema and popularity and leading edge and everything else, there's going to be the manual work of human needing to figure out what should be the part of the metadata that the AI should primarily use. So I would say this is starting to become a lot more important and a lot more companies are starting to look for ways that they can actually scale this part. Can you explain what a semantic layer is and what the components are to it?

18:32Sure. So a semantic layer is usually a separate layer, like a lot of, I think now the definition is starting to get blurred. But usually a semantic layer contains an explanation of data model that's laid out as a logical data model. So it will describe which tables, columns, it should be part of a logical data model and how those fields will make up of a metrics definition, which are the dimensions and measures and facts and how it should be put together, hence by an AI or by semantic, any tools that supports semantic layer. The biggest, I guess, difference or the reason why people have their semantic layer on top of their physical data layer is just so that they can separate out what is considered as verified or certified data sets that should be and can be used by their business users or in their reporting purposes.

19:38Semantic layer and semantic modeling just have gotten a lot more interest recently because that itself can really provide the certification to the AI. And AI can just really follow those definitions to use. And this is a piece that we've noticed how we are starting to see a lot more automation that we can build on top of by just scanning what is being used by your BI dashboards. So for example, if you have a Power BI dashboard and you have a data set or a semantic model defined within Power BI, we can map the lineage for the fields and the calculations that you might have defined within BI tool, and then translate that into a SQL model that your AI can also use to query.

20:32So semantic model and the semantic layer generally is more just focused on defining like how should metric X be calculated and what does that definition look like? Whereas that used to be seen as a way to consolidate or govern the metrics calculations when you're connecting multiple different tools together. But today, this is something that, and from where I'm seeing of the use cases of how the AI can really use that definition to make the queries instead of trying to come up with its own definition for querying the data. For construction of the semantic models and the data lineage and ultimately the knowledge graph, are you leveraging AI internally to automate some of that?

21:21Yeah, we are using a number of different models regarding coming up with, I guess, the generating the queries and also validating the queries related to semantic model. There's also more of the formatting of the files, whether that's a Markdown or YAML, in order to have this integratable with other systems as well. But the core part of, let's say, where that data comes from, when we are defining the logical table or fields or verified queries, those are, I would say, more coming from SelectStar's metadata infrastructure system is something that we've built over the last five years. Mm-hmm. Is there any, I guess, danger with or consequences to using AI to automate some of the construction of this and then AI relying on ultimately the construction to deliver some value?

22:13Some AI system is going to leverage this, what SelectStar provides in order to be able to understand the underlying data better. But since AI is used to construct that, there could be some risk where it's not done 100 % accurate. And does that create a situation where you get a cascading set of inaccuracies that could impact each other? I think that's an interesting question. So every time we're generating a semantic model for our customers or any metrics definition, this is a part where we will have the user to verify. So it can exist and may be used by AI agents readily, but we highly recommend our users to actually take a look at it to actually validate the model.

22:57And then the other side of this that I think is also really important is the evaluation side. So for the business, when someone is considering building any text to SQL bot or agent, having this set of business questions that are likely to be asked and all the definitions to be correct. I think this is more of you're trying to build a product that you do need a set of tests to go along with it. So I guess to answer to that, I don't think it's something that you should 100 % trust. The way that we see this really helps the customers is that you can really kickstart the journey of being able to focus on actually the important part, which is testing and iterating, rather than trying to manually create the YAML files and pick the tables and figure out which tables and columns make sense, what their relationship should look like.

23:55A lot of the times we see companies going back to doing a ton of data modeling on top of their data mart. And a big part of that is almost like rewriting what they already have in other systems already implemented in BI. Yeah, so it's more of a way to speed up the process, like human in the loop, a human still there to be involved, but you can automate a significant amount of the manual work. Yeah, that's first, I would say, benefit to start with this approach. And then the second benefit, which we're working on, is that because we are tracking the underlying metadata, when there are changes such as new calculations being added on the BI front or underlying tables missing, things like that, these operational issues and having these semantic models to be up to date with the current data model is another piece that will make the semantic model scale with usage.

24:53So you mentioned MCP earlier. Can you talk a little bit about what you're doing, what your MCP server does and how people use that? Yeah. So our MCP server today is more of an interface of SelectStar. We have, I think, four or five tools today. One is for searching the metadata. The second one is to get asset details. And then third one is getting lineage and traversing the lineage. So just searching the metadata, Like I've had this kind of like a test before. I wanted to understand what the customer distribution looked like. And I asked my cloud desktop that was connected to our MCP. It will start from getting all the metadata, which had like more than two, 300 different tables.

25:40But from there, using the SelectStar's popularity score and other relevancy metrics, it would narrow it down to like 20. And then from there, it will pick the tables and columns that it would use to create a query and will execute the query to get that result. And from here, I just talked about like search metadata front. Every time that there is a table, then it would use an MCP tool for the get asset details that will get back all the information about that table, including the descriptions, example queries and joins, and when it was updated the last. So those are all something that MCP, sorry, the cloud was looking into to determine and also use those examples to actually put together the query.

26:30And then other tools like lineage, walking through the lineage and then checking the lineage, usually that is used for checking the impact. So if I were to update my DBT model or SQL query, while it's doing that, it can check if there will be any downstream impact from changing the column names or dropping a column. And it can also bring a list of users or owners that may get impacted and needs to be notified from making that change. So those are some of the areas where we've seen a lot of our customers using SelectStar MCP server for with their cloud or cursor, different IDEs that they use for AI work.

27:17Right. So this gives them an interface to kind of speak natural language, but be able to interface with the data. That's right. And we are hearing that this has been a really great addition because they've been using DBT or Snowflake MCP or their own homegrown MCP to just execute queries or having it to grab the schema metadata. but the schema metadata alone does not provide the queries that they want. The accuracy only really came after having SelectStar starting to provide this direction of popularity score, lineage, example queries, all the documentation, so on and so forth. Can you share anything around the accuracy boost that you get from using this approach versus only having the bare bones schemas?

28:06I would say this is not something that we have a scientific measure for, other than the anecdotes and the numerous customer interviews that we've done. And we've been watching customers in terms of how they've been using it. But it's more of like, if you start using it, then you would never go back to pre-SixTime CPUs, basically what we've seen. Why is it that natural language to SQL against real-world databases is difficult? Well, I think that's a really good question. I think there are multiple reasons why. If I think about just a lot of language models directly, we are just at this point only because the foundation models have been trained with just the whole world's data.

28:52All the books and written literature and every single one of them are almost just examples of how language has been used. It's not just because you're training the system with instructions of how to speak a language. It's not because you put a rule of how this should work. It really kind of comes from just having a lot of data or examples of how things have been used. And I think this is why example queries come in as one of the parts that makes the query accuracy much higher. I think the other part is just when data models, as data models get bigger. So if we're just talking about relational database with everything is completely normalized and all the names of columns and tables are very accurate, then it might be easy enough to get accuracy with SQL.

29:50I mean, this is why I think we are starting to get to a really high marks on like Spider or any of the industry benchmarks for it. But it's only the real world data when you're trying to use any of the benchmarks, it actually fails. And that really comes from the real world data is a lot more messy. There are a lot of similar looking tables and columns and how they are being used. There are also like second, third level calculations and metrics that are built on top, you know, that you can easily find in a lot of organizations. I think a lot of those all contribute to complexity that makes it easier for LLMs to hallucinate than actually generating the seemingly easy queries.

30:38But I think it fails because of that. Yeah, I think that makes sense. I mean, I always say that foundation models are really, really smart about kind of general information, but they're really dumb when it comes to your specific, you know, business information because they've never trained on it. So to get value out of them for specific tasks, it's all about like, how can you correctly contextualize the prompt? And if you're doing this kind of, you know, natural language, this equal generation against complex data models that exist in, you know, your warehouse or your lake house or something like that, then without the correct contextualization of how that data, essentially, how do you encapsulate the tribal knowledge that people have within the company?

31:19If you can't feed that into the model, then there's not really a way for the model to probably accurately run a reasonably complex query against it. Yeah, I think that's really well put. Everyone says context is the king. And for data, how do you structure that context that's actually relevant for SQL generation and analysis of data? I think has a particular flavor to it. And I think that is primarily what we've been focused on because we understand that something like popularity or lineage has a very specific implication of how the data should be retrieved or what type of impact it will have on the use of the data.

Read the full transcript

32:06Have you thought about extending any of this approach? So it sounds like, you know, if I'm using Snowflake or something like that, then, you know, I can run SelectStar against my Snowflake. I can go through some process to make sure that what it produces is accurate. And then I can, you know, start to use something like Cloud Desktop and use your MCP server to explore that data in an actual language way. But what about situations where I might want to pull data from other types of systems, not necessarily the warehouse, but I might want to talk to, I don't know, like a SaaS API endpoint, or maybe even a transactional database.

32:42Is there potentially a role for this approach to extend beyond just the understanding of the warehouse data? Yeah, yeah, for sure. I think there are now different ETL systems that we connect to, as well as applications that we're starting to connect to. And I'm seeing that as we add more integrations, so it's not just data warehouse queries, but even in the future, I think we'll be able to start generating dashboards in Power BI and Tableau once they have their own MCCP server, for example. I think that is kind of like the future that we see. With some of the stuff that you're doing around the MTP server, given that you're primarily serving metadata, is it the same?

33:28I guess, do you need to be concerned about what a specific user is accessing? Or is that more going to be a security requirement on where they ultimately are executing that query because that's where the actual data lives? Yeah, that's an interesting question. So right now, it really kind of comes down to the end user role of where the query gets executed. We do have policy based access control support, so that you can, you know, limit the user to query or even just look up metadata within a certain set of schema tables or logical grouping that you may have. But in terms of the actual query execution, we're a little bit decoupled in a way where we're leaving that to the data warehouse user because that's where the query gets executed.

34:18We will generate the query and you can limit the query to only access certain parts. But in terms of security perspective of end user querying, this is something that we kind of offload to the data warehouse side today. And then are you offloading the context optimization problems to the engineer that's building this application? Because I would think that some of this metadata could get pretty big, where it starts to eat up a reasonable amount of the context window. So how does that optimization work? So if the engineer is using the MCP, this is just really kind of like, you know, we control it on our end.

35:00When you say context optimization, it really, I guess, comes down to what context we are exposing. So we have our own embedding that we use for our Ask AI, which is the AI assistant of SelectStar. But that whole context isn't something necessarily like we expose fully to the developer. We do this through the MCP server. So what I'm saying is we don't necessarily put the full embedding of raw metadata for AI agents to use, but rather have the MCP server to provide that information upon request by the agent. Right. So right now there isn't really a challenge regarding, you know, fitting into the context window.

35:45But I think the piece that might actually be interesting to you in this regard is a semantic model generation. So we do have a way to now summarize and build a semantic model for the customer so that the customer can basically use that and feed that into their AI application. And I would say that we haven't gotten to a point where it is so large that it doesn't fit into the context window, but it's fairly early days. So we've been just testing with a number of customers on this, but haven't really run into that issue. Great. In terms of what's next for you guys, where's your focus? Is there anything that you can talk about in terms of the challenges that you're working on now or things that you have coming out relatively soon?

36:31Sure. Yeah, first and foremost, the semantic model is a big part. We have seen that this really helps the text-to-SQL approaches on building an alum to speak the language and surface business questions really well. So we are looking at ways for this to be more general and available for more customers today. So that's one part. The other part is having select stars ask AI to kind of use that model and courting the data directly for the end users so that more users can just ask questions to their data and then about their data and get answers right away. And last but not least, we have different agent workflows coming up that really helps building the business context metadata more automatically.

37:23I'm talking about we already have ways that we're starting to do a lot of auto-documentation of data assets, but a lot of things that we're looking at that are coming up would be tagging the data assets and assigning ownership or propagating different ways of documentation, which we already do. But putting it into the hands of an agent that maintains the governance is really kind of like the direction we're heading today. What are your thoughts on where the value of metadata is going? Historically, we've put a lot of value in the data that a business collects. Historically, when it came to databases and warehouses, there was tight coupling between the compute and the storage.

38:11Then eventually, we separated those things. Now, we have these open table formats of iceberg and delta tables. and we're getting to a place where that data might actually exist just in some cloud storage bucket that's outside of where the actual compute runs. And people want to own the compute work, and there's not as much value attributed to the hosting of the data itself. But now I think one of the things I'm seeing happening in the industry is that a lot of the big data players, they really want to own the catalog and the metadata. So is the new oil of data, especially in the world of AI, really all about the metadata?

38:52I would say it is the map of where the data is, and that's why metadata is being taken a look at now. It tells you what exists and what's really important. And for cloud provider perspectives, I think it's really to expand more capabilities under the same umbrella. And that's why a lot of the cataloging or these metadata features are being introduced by larger players as well. But yeah, nonetheless, I think because if you have the map, you can actually leverage that for operational purposes, like automating impact analysis and let's say PR or letting the downstream users know of what's going to change.

39:40or even using something like popularity to have an AI agent to write the right correct query and pick the right columns and tables. There's just a ton of different things that you can add when you have this context. And I think this is why a big focus is starting to be put on metadata and also high quality of metadata. Yeah, that's great. I wrote an article recently about sort of comparing the semantic web to the world of large language models and use this analogy of how things like the semantic web and ontologies and sparkle and all the things that kind of came from that world was like essentially an architect's view of the world where we're going to predefine these things and sort of architect the structure of the web at the time.

40:28And then with, you know, foundation models, there are much more of like an explorer where they don't have that predefined structure. They're just kind of going out, stumbling around and figuring out based on the patterns of behavior, the way that we write, what are the associations between these things. But ultimately to make them more useful and accurate and prevent things like hallucinations, they need a map. And that map can be things like, you know, ontologies or in the context of what we're talking about, metadata and sort of the eventually. Yep, exactly. And having that up-to-date knowledge graph of the data is really the key.

41:08And in order to make it accurate and also to scale for multiple use cases and as the data changes underneath. Yeah, absolutely. Shinji, I want to thank you so much for your time and for coming back. I really enjoyed this. Thanks so much, Sean. Cheers.

41:30Thank you.

From the publisher

A common challenge in data-rich organizations is that critical context about the data is often hard to capture and even harder to keep up to date. As more people across the organization use data and data models get more complex, simply finding the right dataset can be slow and create bottlenecks. Select Star is a

The post Context-Aware SQL and Metadata with Shinji Kim appeared first on Software Engineering Daily.

More from Software Engineering Daily

All 195 episodes
Context-Aware SQL and Metadata with Shinji KimSoftware Engineering Daily · 42 min
Listen in VO