Foundation Models for Structured Data

23 Jun 2026 · 44 min · 19 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

The episode argues that “foundation models” should be built for structured relational/tabular enterprise data, not just text/images. It introduces “relational deep learning,” treating databases as graphs and using transformer-style attention over rows/columns and across linked tables to reduce manual feature engineering and task-specific model training.

Key claims

traditional predictive modeling (fraud, loan approvals, churn, recommendations) relies on slow, brittle feature engineering and separate models per task; tabular data needs modality-specific neural architectures because it’s quantitative and enterprises store data across many interconnected tables; graph transformers generalize attention to relational databases; a pre-trained, predictive-task-agnostic model can be prompted with a “predictive query” (e.g., churn defined as “zero transactions in next 30 days”) and then fetch subsamples from the database for inference.

Notable examples

fraud detection, loan repayment risk, hospital readmission risk, customer lifetime value/churn, recommendation scoring, and “text-to-SQL” being insufficient for forward-looking prediction.

Guests

Yuri (Sean Falconer’s guest), Jura Leskowitz—Stanford CS professor (AI lab), former chief scientist at Pinterest, investigator at Chan Zuckerberg Biohub, co-founder of Kumo AI; his work focuses on AI over structured tabular relational data and graph-based models used at Meta/Facebook and YouTube.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Understanding Predictive Modeling

1:26 to 2:14

Discussion on the core aspects and importance of predictive modeling.

“This episode is hosted by Sean Falconer.”

The Transition from Industry to Academia

2:14 to 4:12

Exploration of the speaker's shift back to academia and its advantages.

“And in my career, the models we've developed are today used at Facebook, at Meta, at YouTube, and at a number of different places.”

Predictive Modeling: Evolution and Challenges

4:12 to 9:46

In-depth discussion on predictive modeling's past, current methods, and challenges in feature engineering.

“We don't have all that that is in the industry.”

Neural Networks and the Future of Machine Learning

9:46 to 12:41

Exploring the potential neural network revolution in machine learning for structured data.

“And now majority of the work goes in building that scoring function that takes the user profile and the item profile and gives me the prediction.”

Challenges in Applying Neural Networks to Structured Data

12:41 to 14:00

Discussion on the unique challenges of applying neural networks to structured relational data.

“So why is this kind of hard in practice?”

Understanding the Unique Challenges of Tabular Data

14:00 to 17:48

Learn about the distinct nature of tabular data and its implications for machine learning.

“So I think it's a different data modality.”

Generalizing Attention Mechanisms for Tables

17:48 to 19:16

Discover how generalized attention mechanisms can improve predictions across interconnected tables.

“But if you think about SQL, SQL is summarizing what has happened last month, what has happened last week.”

Graph-Based Approaches to Database Modeling

22:20 to 24:28

Examine how traditional database concepts overlap with graph-based machine learning.

“Is this something fundamentally different or is it using a similar concept as essentially the basis for doing this deep learning?”

Training and Querying with Structured Data

24:28 to 28:00

Understand how to train models with structured data and make predictions using database queries.

“and include columns and categorical values and all kinds of geographic information, all kinds works nicely in this framework.”

Understanding Churn Prediction with Structured Data

28:00 to 29:40

Learn how to effectively use models for churn prediction by defining tasks and specifying queries.

“So how do I actually ask questions about churn prediction?”
Show all 19 chapters

Model Operation and Efficiency in Predictive Tasks

29:40 to 31:20

Explore how models fetch data from databases and the efficiency of inference in relational deep learning.

“hoc querying, we see is very useful for, let's say, humans using this type of models or agents using this type of models where you don't know the question ahead of time.”

Challenges in Graph Structures for Databases

31:20 to 33:00

Understand the complexities of converting linear data structures to graph formats for AI workloads.

“And on the other side, you need a GPU for the results to come out.”

Evolution of Graph Neural Networks

33:00 to 34:40

Learn about the development and functioning of graph neural networks and their application in data prediction.

“So I think the graph neural networks were invented maybe 10 years ago, something like that, even a bit less.”

Generalization vs. Purpose-Built Models

34:40 to 36:20

Discuss the effectiveness and quality differences between generalized models and purpose-built models in predictive analytics.

“The attention mechanism is so flexible and so contextualized and context-dependent that it really allows for much better generalization and for very effective pre-training.”

The Power of Data-Driven Learning

36:20 to 38:00

Explore the advantages of allowing neural networks to learn directly from raw data instead of predefined rules.

“that then tells me, is this an elephant or not?”

Learning Through Experience vs. Rules

38:00 to 39:40

Examine how learning through experimentation resembles human learning more closely than fixed rules.

“And I think the reason is the following, right?”

Applying Graph Structures Across Various Domains

39:40 to 42:01

Discover how graph structures can be utilized in diverse fields such as biology and fraud detection.

“And then, of course, right, I think when you talk about rules, I would say those are also in some sense important.”

Exploring Multi-Modality in Foundation Models

42:01 to 43:34

Learn about the advancements in using foundation models for various data types.

“In the end, it's kind of mathematically, it's all the same.”

Challenges and Opportunities in Structured Data

43:35 to 44:46

Discover the potential of foundation models in transforming structured data analysis.

“And also, right, just shows that, you know, the two of us right now are just talking.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Predictive modeling is a core element in modern systems, and powers capabilities such as fraud detection, loan approvals, and recommendation systems. These systems typically operate on structured, relational data stored in enterprise databases, with rows, columns, and interlinked tables. While computer vision and natural language processing have undergone a neural network revolution, the tabular data layer underpinning predictive modeling still largely relies on manual feature engineering and task-specific models. Relational deep learning proposes a new approach. It treats databases as graphs and applies transformer-style attention mechanisms directly over structured relational data.

0:42Researchers are now building foundation models for tabular data that aim to generalize across predictive tasks without painstaking feature engineering. Jura Leskowitz is a professor of computer science at Stanford University, and he previously served as chief scientist at Pinterest and was an investigator at the Chan Zuckerberg Biohub. Most recently, he co-founded the machine learning startup Kumo AI. In this episode, you rejoin Sean Falconer to discuss the limitations of traditional predictive modeling, why structured enterprise data requires its own modality-specific neural architectures, how graph transformers generalize attention to relational databases, and more.

1:26This episode is hosted by Sean Falconer. Check the show notes for more information on Sean's work and where to find him.

1:45Yuri, welcome to the show. Thanks for having me. Great to be here. I'm really excited to get into, I think there's a variety of topics we can dive into today, but maybe before we get there, just kind of grind on the audience a little bit. Like, who are you? What do you do? What's sort of your background? Yeah, I am primarily a professor at the computer science department AI lab at Stanford, been there 15 years. My research focus on AI, especially AI over structured tabular relational type data. I work a lot with graph data. And in my career, the models we've developed are today used at Facebook, at Meta, at YouTube, and at a number of different places.

2:26In my career, I was also a chief scientist at Pinterest for six years, grew Pinterest from 150 employees post IPO and was basically building large scale AI machine learning platforms there. What drew you back to academics after having experience at a place like Pinterest? I think Stanford is the most amazing place. It's where the future happens. And what is also amazing about Stanford is that it has this allow us to kind of flow between industry and academia and really understand what are the biggest problems out there that are worth solving. We go to the industry, we come back with the ideas, the students at Stanford are amazing.

3:03And I would really say kind of the future happens at Stanford. And that's the most exciting part to me. Yeah, I spent some time at Stanford myself as a student and then was drawn out to industry and never found my way back. But I guess now that there's so much going on in the space of artificial intelligence, especially with companies, you have your open AIs, your anthropics of the world, and every large cloud provider doing amazing work in the space. How do you think in terms of where research is going to make significant contributions versus where maybe the private sector of companies are going to make significant contributions?

3:39That's a great question. I would say research academia is different, right? Like we cannot compete on scale. We cannot compete on pushing products to customers and things like that. That's why we have startups. That's why we have industry. And whenever we make some new research breakthrough and we want the world to see that, we spin off a company, we spin off a startup and scale it up there. But then at the same time, there is huge value in academia and research and education because the risk profile for us is very different, right? Like we can truly explore. We can truly fail. We don't have performance reviews.

4:13We don't have all that that is in the industry. So we can always kind of ask about what are the paths not walked yet? What are the interesting new directions that maybe the industry is too conservative to take? Can we show the path there? And there's been examples of this throughout history where basically academia found a new path or researchers found a new path where the industry was just like plowing forward full scale. And right now it's similar in, I would say, in the field of AI. Of course, the frontier labs are making humongous progress, scaling up these models and so on. But there is so much unexplored.

4:45There is so much more to do. And that's what we are focusing in on our research. Yeah, I mean, I think a big part of that would be there's no sort of commercial obligation in academics. You could chase a problem that may or may not ever have some sort of commercial application just because it's an important problem, or at least to the individual to explore and maybe mean something in the long run to how we think about ourselves, our own intelligence, some other type of scientific endeavor. Exactly. And even with that, right, I would say that we are very careful what kind of questions we ask. And we always ask ourselves, if we solve this problem, who's going to care about it?

5:20Who can benefit from the solution? So being connected to the real world is a very important part of the way we think of the research we are doing. I want to get into this a little bit and first talk a little bit about predictive modeling, which I think is something that has a long, rich history in machine learning, artificial intelligence. There's fairly simplistic ways of doing some form of predictive modeling. And then there's very sophisticated approaches as well. Can you give a little bit of background in terms of sort of the history? What are we talking about when we say predictive modeling?

5:53Why does that matter? And what is the history of the discipline? Yeah, that's a very interesting question. And I think as we go through, it will become clear why this is interesting. But predictive modeling has been around forever. And it's all about forecasting risk estimation. It's about filling in some information that doesn't exist yet. And where does this matter? This truly matters for, let's call it quantitative decision-making. Right. If I go ask for a loan, there is a predictive model that estimates the probability that I'm going to pay back that loan. If I'm in a hospital, the hospital wants to estimate what's the risk that if I get discharged that I get readmitted.

6:31Right. If I am, let's say, dealing with customers, I want to estimate what is the lifetime value of the customer. I want to estimate how likely this customer is going to churn. I want to estimate what next product or what next show item to recommend to the customer. If I'm a, let's say, financial institution, I want to estimate what's the likelihood that this particular transaction is fraudulent. When a, let's say, user logs in, I need to estimate how likely is this a stolen identity? Somebody else is logging in and things like that. These are all predictive type problems where based on the historical patterns, based on the data, we want to estimate something that we don't know yet.

7:10we want to forecast something and so on. And this has been around for a very long time and people have been building machine learning, statistics, data science, have been building these predictive models for the last 20, 30 years. And the point is that every percentage improvement in accuracy of these models means humongous business impact, right? Even 1-2 % improvement can have humongous business impact. How would I think about this in terms of with forecasting, say it's like financial forecasting, I'm taking a bunch of history and I figure out what is the function that scribes that history. And then I'm projecting that out to see where could this, I don't know, trend in the stock market go or something like that.

7:51How is it when you think of something like classification? So I want to go and take a bunch of history as my training set, and I'm going to use that to train some sort of classifier to figure out whether an email is spam or not. Is that predictive modeling or is that a different type of classification of how we would think about that AI model? This is all what I would say falls under predictive modeling. Both examples, time series forecasting, any kind of classification, churn modeling, as you gave the example, where the idea is that based on some historic patterns, you are trying to forecast, is this person going to cancel subscription in the next month?

8:28And because we don't have the information what is happening next month, we have to forecast. If we take a specific example, like a recommendation systems, I think Amazon's been pretty famous for using recommendation systems for products for a long time. How do those systems typically work? Yeah. So the reason we started opening this is because this field has remained practically unchanged for the last 20, 30 years, right? It's all based on this idea that you bring in, let's say, a data point, a unit of something that is described with a set of characteristics or a set of features, and then you are making that prediction, right?

9:05So, for example, for a recommender system, the idea would be that you have a description of the user, which would be maybe how long ago did the user register? When was the last time the user logged in? What were the last seven products the user visited? And what categories are these products from? And so on. And you have this kind of profile of the user. And then you would say, okay, I also have to now build a profile of the product. and now I need to learn some function that takes the profile of the user, profile of the product and tells me how likely is, let's say, the user to purchase that product.

9:36And then if I find top 10 products that the user is most likely to purchase, I show those to the user and my sales go up 20, 30, 50%, right? That's how this is generally done. And now majority of the work goes in building that scoring function that takes the user profile and the item profile and gives me the prediction. And if you say, how accurate is this prediction going to be? You have two aspects to it. One is how accurately do I build the user and the product profile? Let's call it this way. And then how powerful that predictive model on top is. And the point is that this is like super painful, super slow, and super manual to do, right?

10:15Like you need to hire a team of data scientists. They need to build these profiles. This is called feature engineering. They come up with some historical summaries of the user activity, put those as into the user profile, they do something similar with the product, and then they create these training data sets to build the models on top. And the models on top can be this kind of decision tree, XGBoost, CADBoost type things that work kind of well in practice and are completely respectable, as well as to more sophisticated neural network approaches and so on. But the bottom line is that you need about two full-time people to build a single model, right?

10:52Like if you say how expensive this is, it's like I need two employees to support one model. So if I now want to have 10 models in production that are making these decisions on the fly, I need that number times two number of people to support that. And if I build like my e-commerce recommendation system, I also need fraud detection. I can't just pick up my recommendation system model and apply it to fraud detection. I got to go train like essentially a new model. I got to do feature engineering just for the fraud detection. probably use maybe even a different type of model to train and test against, and then probably operationalize that model with a different set of people.

11:28Exactly. Exactly. And I would say now, what is the exciting thing? Why are we talking about the past? I think the exciting thing here is that this entire area hasn't seen real progress in the last 20, 30 years. And if you think about what has happened in the broader AI ecosystem is that we went fully neural network. What I mean by that is on the, let's say, In computer vision, we used to do some edge detection, some feature detection, and then build a classifier on top to say what's in the image. We don't do that anymore, right? Today, a neural network just learns directly from the pixels of the image.

12:01The same thing, I would say, happened on the, let's call it, natural language processing area, right? Where we used to do all kinds of parsing and feature extraction from sentences to try to say something about what the text is saying. And today the attention mechanism just attends over the tokens of the text and kind of the AI, the reasoning is born, right? So I think the exciting thing here is that machine learning hasn't gone through this neural network revolution. And that's the exciting new thing here is that there is the neural network revolution for machine learning ready to happen that completely changes how we are building these models, how accurate they are, how much feature engineering it takes, and things like that.

12:41Mm-hmm. So why is this kind of hard in practice? And why couldn't we just take, we've had this revolution around things like large language models that understand text very well, and now they understand images and audio files and even video. Can we not take those and just apply them to this problem? I think the argument is the following. I would say, if you ask what kind of data are we using when we are making this predictive modeling, predictive problems, it's structured relational data, right? So this is data that is stored in tables that are interlinked with primary foreign key relations. This is usually stored in a database in some structured form, right?

13:19And this is the most useful data that enterprises have because this is kind of the ground truth of the enterprise. All the events, all the activity, it's all stored there, right? So now what I'm saying is if you say for images, to process images, we have special neural network to process images. To process text, We have special neural networks. And to process this tabular data, we need special neural networks. It's a different data modality. Image has its own set of networks that are different architectures trained in certain ways. Text has a set of networks trained in specific ways and so on.

13:53And the tabular data also needs its own set of neural network architectures and its own way to train this that can directly kind of attend over this structured tabular data. So I think it's a different data modality. That's why it needs a different approach. Right. I mean, I think with something like text, where there's probably billions, if not trillions of examples of how sentences come together, and I can take a document from one place and a document from another place, and learning one of those documents can probably help me infer something about the other document. I think with the way I think about this problem around tables, rows, columns in a database is, is there that much I can, from a pattern standpoint, like learn from one table versus another table?

14:38Like those patterns seem like they could be fundamentally different in terms of how the person has modeled the data. So it makes a lot of sense that this is sort of a different class of problem, but how do you go about, I guess, like attacking that problem? And does this make sense in terms of the patterns from one table, not necessarily allowing you to infer something about the patterns from another table? That's a great question. I think the way to think of this is as you have this private data, right? You have patterns, properties in it that are unique to you. I think another point to make is also like you cannot just textify a table and give it to a large language model, right?

15:12Large language models are amazing at what I would say qualitative human-like reasoning, but they are not really good with numbers if you want to say it very simply. And if you think about what are we storing in tables, we store quantitative data. So we need to do quantitative reasoning, not qualitative reasoning over huge amounts of data. Another point I think that is important here is to say that no enterprise has data in a single table, right? You have data spread across multiple tables. You usually would have your customer catalog. You would have your product catalog. You would have a set of transaction records.

15:46You would have your website browsing click data. You would have your supplier data. You would have your returns data. And all these tables are interconnected. And the only way to learn over this is to learn over this collection of tables as they are interconnected with each other. And maybe to say more, right, in the past, the way we deal with this, we would say, oh, let's take the user table. Let's take the transaction table. Let's join them. And then somehow summarize the number of purchases, the number of transactions you had in the last time period. But there is kind of an infinite way of creating these summaries.

16:22I can count, I can sum over some time period, over shorter time period, in the mornings, in the evenings. I can add the prices, I can look at product categories, right? So the number of ways you can summarize the data is kind of blows up. And the problem with machine learning is that we kind of predetermine how to summarize the data before we start building the model. So the way to make these things better is to generalize the attention mechanism to attend over the raw events in the database and learn how to summarize them to give you that prediction. And that's the key differentiator. You don't need to be joining tables anymore, but let the attention mechanism.

17:01Very similarly, as in a language model, it attends over the previous words to say, OK, what's truly the meaning of this word? But here we are saying, if we are making a prediction, if we are filling in, let's say, some cell in a table, let's attend over the other rows in the same table, other columns, other tables far out, and figure out how to bring all that information so that we can make the accurate prediction. So it's really about generalizing the notion of attention mechanism to this structured, multi-tabular data. And if we have that, what does that unlock from an application standpoint?

17:33Today, I feel like a lot of people are trying to apply large language models to databases for the purposes of being able to do intelligent natural language to SQL conversion. Does this yield a better version of that? That's a great point, right? So text to SQL is amazingly useful. But if you think about SQL, SQL is summarizing what has happened last month, what has happened last week. So you're kind of aggregating some past and maybe creating a dashboard to understand historical trends. And then maybe you can use those historical trends to do some kind of qualitative decision making about what to do tomorrow.

18:11But if you think about predicting transaction fraud, deciding which customers to send an offer to, and so on, for that you need predictive modeling, right? So text to SQL won't get you anywhere with that. You may use SQL to generate historical patterns and then build a model on top. But as we talked about that, that's super brittle, manual, and takes a lot of time. So the approach we invented in my research group at Stanford, and then we founded a startup around it called Kuma.ai, is this notion of relational deep learning, where we basically take the transformer architecture and generalize it so that it can attend over this structured relational enterprise data.

18:55And the key to this approach is to think of the data as a graph, to think of your enterprise's data as a set of connections between the entities in your database, right? To think of the tables, how they are interconnected as a graph, and then generalize the attention mechanism to be able to attend over this relational, structured information. Agents are getting smarter every day. But even the smartest agents get stuck without the right context and the right tools. That's where Notion comes in. With the recent launch of custom agents, Notion became the collaborative AI workspace where teams and agents work side by side.

19:31And now, their new developer platform is turning that workspace into infrastructure developers can build on. Most agent platforms are single player, making you stand up your own infrastructure just to start. Notion's developer platform flips both. You get primitives to sync any data source in, give your custom agents tools that plain MCP can't deliver, and orchestrate agents like Claude or Codex alongside your team. The CLI authenticates in one line, and workers run on Notion's runtime, so there's nothing to provision. Because the workspace your team lives in is the same thing you build on, permissions and governance come standard.

20:07Write your code. Deploy. Done. Learn more about Notion's developer platform today at notion.com slash SED. That's all lowercase letters. Notion.com slash SED to try Notion's developer platform today. And when you use our link, you're supporting our show. Notion.com slash SED. If you're running Postgres in production, you've probably felt the moment analytical queries start fighting your transactional workload. Most teams end up adding a second database and all the pipeline complexity that comes with it. Tiger Data, creators of TimescaleDB, takes a different approach. We extend Postgres with hybrid, row, and columnar storage, so one table handles both writes and analytical scans.

20:47Native compression cuts storage costs up to 95%. Continuous aggregates keep dashboards live without bash jobs. And it scales to petabytes without you re-architecting. Companies like Cloudflare, Octave Energy, Schneider, Axpo, and Floco run production workloads on Tiger Data today. No stale data. No second system to operate, just Postgres. Managed for you, ready for the workload you're building toward. Try it free at tigerdata.com. You know Fidelity is a financial services leader. But did you know that Inside Fidelity is a community of technologists working together to shape the future of finance and tech?

21:21Fidelity is always investing in tomorrow, from emerging tech to cutting-edge tools that will transform what comes next. Their technologists are encouraged to keep learning so they can expand their skill sets, explore new ground, and stay ahead of this rapidly evolving industry. And right now, Fidelity is hiring technologists to join their team. Fidelity technologists get the best of both worlds, startup energy that's grounded in the stability of a financial institution. That means support, resources, and amazing benefits. Bring your skills to a culture where you're empowered to dream big and build the tech that drives an organization and makes a real impact on people's lives.

22:01Find out more at tech.fidelitycareers.com. That's tech.fidelitycareers.com. Fidelity is an equal opportunity employer. So, I mean, thinking of a database as a graph, I think people in the data modeling world have been using those concepts for a long time. Like, how is this different? Like, if you think in sort of designing schemas and things, people used to build like entity relationship diagrams where essentially you have your tables as a node and then you have relationships defined in terms of edges across foreign keys and so forth. Is this something fundamentally different or is it using a similar concept as essentially the basis for doing this deep learning?

22:42Yeah, it's a great point. Any database is a graph, as you've said. And these concepts have been around for a very long time. But I feel like nobody has put one plus one together in a sense. We've been working on this graph-based machine learning, graph transformers and things like that for a long time. but it was mostly applied to social networks. And people kind of that community didn't realize that actually the database is any database is a graph. And I think the database community was so stuck in feature engineering and running historical SQL queries that they did not think about, oh, how can we take these AI tools and apply them to the database, right?

23:19So what I'm saying in some sense is supernatural, right? We knew for a very long time, the database is a graph, nothing changes, but what in some sense changes is that now the field of graph machine learning is not this obscure, it only applies to social networks type thing, but it applies to any database. And the benefit that comes with it is that is the same neural network revolution that we have in computer vision, that we have in text, natural language understanding, videos, and so on, now to another data modality, which is the structured tabular data modality. With the same set of benefits and the same set of amazing outcomes that we have already seen play out in other data modalities.

24:00And does this only sort of work against like a traditional database? Or could you potentially generalize this to other tabular forms of data like spreadsheets? Yeah, I mean, where the data sits, it kind of doesn't matter that much. It can be in a database, it can be in Salesforce, it can be in Databricks, can be in Snowflake, can be in spreadsheets, can be in JSON. as long as it's structured, semi-structured, and it may include images and it may include text and include columns and categorical values and all kinds of geographic information, all kinds works nicely in this framework. How does training work and how is it different than training with traditional transformer models?

Read the full transcript

24:41Yeah, that's a great point. So far we said two things, right? The first thing we said was any set of enterprise data is structured, semi-structured, is usually stored in these relational tables. These relational tables are a graph. So now what we can do is we can take, in the old days, we would take a graph neural network and apply it to a graph. But in today's age, we take transformers, in particular graph transformers that have the generalized notion of attention and positional encodings that can attend over this structure to give you the prediction. Now that, okay, we have these two steps. Now the third step is how do we train?

25:16And we have, I would say, two options here. One option is to actually build a pre-trained foundation model that is database schema and predictive task agnostic. So what this means is that now you have a single pre-trained model that can connect to any structured data because it just represents and thinks of it as a graph. And then you get basically the specification of the predictive task on the fly and the model is able to make you that prediction. So this means that you don't even need to be building task-specific models on the fly, but you basically have the same type of chat GPT type experience, but now for predictive type of questions, right?

25:54So for predictive modeling type questions, for churn, fraud, readmission prediction, all kinds of advertising use cases, customer 360 use cases, marketing, sales, and so on. And the way this is trained, as you asked, is it's trained basically to teach the model how to do in-context learning, right? So basically how to look at subsets of your database, use those examples to then generalize and be able to make the prediction. And when we say what are the subsets of the database, are these small local subgraphs around the entity or around the entity that you're making the prediction about. So if you're doing in-context learning, is there learning, though, that goes on in order to first force sort of this attention mechanism?

26:38Or, you know, when I think of training traditional large language models against text, sort of the input is all these massive amounts of text that then become tokens. Those are essentially the inputs of the model that help adjust the weights. Is something similar going on? I think I'm not quite following before the context learning. So what you have to do here is you also have to amass humongous amount of pre-training data. This pre-training data now needs to be structured. So you need to amass a number of different tabular data, databases, things like that. And then you also need to define a number of different predictive tasks on top of this data.

27:14What is interesting is that we have just shown, we just published a paper about a method called Plural, where we can basically synthetically generate a lot of data and then train these models on top of that. But you are exactly right. The same way as in, let's say, training large language models, you need large amounts of text. Here, you need large amounts of structured tabular relational data. In terms of, say, like a user interacting with this, they want to be able to do predictions against their own databases or tabular data. What is the input there? Because, again, with going back to the large language model, my input is essentially text that then becomes tokens as the input, which is the same thing that the model started as a training corpus.

27:56here it's a bunch of, you know, database data, you know, graph structures of these databases. So how do I actually ask questions about churn prediction? Yeah, great question. So at the setup time, you need to connect the model to your database and say, here's my database. These are my tables. And then to prompt the model, you need to specify the task. You can specify the task in natural language, or you can specify it in a domain specific structured language that we called predictive query. And to say, I want to predict churn, you need to specify what does that truly mean. So you could say, I want to predict whether the count of transactions or count of purchases of this particular user is going to be zero over the next 30 days.

28:41That's the proper definition of churn. And if you define it this way, then what the model is going to do, it's going to go into the database, take a set of subsamples from the database, and then send them through the neural network to give you the prediction for that specific, let's say, user or customer that you want to say, what's the likelihood they will have zero transactions in the next 30 days, right? So the way this operates is that the model basically goes, fetches data from the database, and then this data from the database is sent through a frozen model to get an output in the end. So the point is, there is no task-specific training needed here.

29:18There is no need to train the model for your specific database. It's very similar to large language models where the model is pre-trained. You give it text, it understands text, you ask it question, you get the answer. So here is similar. The model goes to your data, understands the data, and gives you this forward-looking predictive answer. Of course, just to say, right, now if you write this kind of ad hoc querying, we see is very useful for, let's say, humans using this type of models or agents using this type of models where you don't know the question ahead of time. If you are a large bank and you need to predict 1 million times per second whether a given transaction is fraudulent or not, you would, of course, go and maybe fine-tune a smaller version of the model to make this faster and more cost-effective.

30:04So you can use the big pre-trained model, or you can fine-tune it for a specific task to get more speed. Again, similar to what we see in large language models. Is the mathematics behind inference for these relational deep learning neural networks similar to that with sort of large language models? You mean in terms of the model sizes and things like that? Yeah, the model size and I guess the matrix computation that you're going to be doing behind the scenes. Yeah. The bottom line is what is the cost of inference and how does that sort of compare to the other models that now have become sort of familiar to people at large?

30:39That's great, right? So these models are transformer attention-based models. So you need GPUs to run them effectively. They are smaller than large language models. So the amount of compute that is needed is way smaller. So what this also means then that predictions and output is faster and cheaper. So that's the way I would describe this. What's the difference? The difference is that you need to have a very strong data backend where the data is represented as a graph. So you need a special, let's say, graph engine that is tuned for this type of AI workloads that allows you to make these computations scalably and quickly.

31:20And on the other side, you need a GPU for the results to come out. But since the models are smaller than large language ones, you know, there's, let's say, sub-billion parameter models or something like that, you can run them quite efficiently and quite quickly. What's involved with getting that graph structure for your own database? is? Yeah, so what we have built at our startup Kumo.ai is exactly the infrastructure that allows you to basically take in any database, the engine will internally optimize that representation into this graphical form, so that then these graph transformers can efficiently be run on top.

31:55And that's the hard part, right? Like the linear structures are kind of easier because you can just kind of feed them through. Graphs are hard because they're just these interconnected sets of objects. You cannot chop them. You cannot linearize them. So you really need to be able to kind of do this quick breadth-first searches or subsampling of subgraphs over this database. And that's the hard part. It's basically building the optimized infrastructure that allows you to do this at scale of tens, hundreds of billions of nodes and edges. Is that the equivalent to, I guess, like with text generating like the embedding?

32:30That's very interesting. Yeah, exactly. What the model does internally, it generates the embedding, right? So now the embedding in the text kind of captures the, let's say, semantic meaning of the word, if you think of it this way. Here in this database graph view of the world, now you have the embeddings of your entities that capture the semantic meaning of those entities that is predictive of the downstream activity, whatever is the task that you are trying to solve. You know, the concept of like a graph neural network, how long has that concept been around for? So I think the graph neural networks were invented maybe 10 years ago, something like that, even a bit less.

33:10That's when the field started. And I would say a graph neural network, it's this idea that when, if you have the data organized in a graph, when a node wants to compute something about itself or make a prediction, maybe this is a user is a node, then not only we are using the information about that node, but you also use the information from neighbors and neighbors of neighbors and so on. And the idea is that basically now neighbors are passing information messages to their neighbors and this way kind of to the node in the center that is of interest. That has worked really well for small scale task specific models.

33:43But now the field has moved forward to basically these generalized graph transformer type architectures where we are not passing information across the edges, but it's actually the attention mechanism that attends over the center node, It's neighbors, neighbors of neighbors, and so on. And you can see how this nicely applies to the database world where if we're making a prediction about an entity, maybe a user, maybe a product, we are then attending over the nearby tables and one hop tables, two hop tables, and bring that information all in the same place. So is the value there of using the attention mechanism to sort of apply, distribute these weights over the graph versus using the traditional message passing that's happening in a graph neural network is that essentially it's more dynamic.

34:28So I don't have to have more of a purpose-built model versus a generalized model that I can apply to any essentially database problem in this particular context. That's exactly the right intuition, right? The attention mechanism is so flexible and so contextualized and context-dependent that it really allows for much better generalization and for very effective pre-training. This then means that these types of models can really be applied to a very large, diverse set of use cases and can generalize to them very effectively. How does the performance of these more general models compare to if I was to go and take more of a traditional approach where I take my two data scientists and I go and I build more of a purpose-built model using my domain-specific data, build that model?

35:18Obviously, it's more rigid, but do I get, is there a quality difference? There is a quality difference. And we shouldn't be surprised that there is a quality difference. And the way it works is the same as it worked in computer vision, where we went to basically, and in LLMs as well, where we are basically going to this superhuman level performance. You could say, hey, I'm so good at recognizing any, like, let's say in computer vision, You could say, I want to build a model that detects whether there's an elephant on the image or not. And you could be saying, hey, I studied zoology. I know elephants inside out.

35:53I'll build a perfect elephant detector. And if somebody would go and say, I know how to detect elephants. I'll build these amazing features to detect elephants. Everyone would be laughing at them. And now, you know, when we go to tabular data, I think the outcome is similar, right? Like it's very hard to engineer perfect features, you know, that give you that accurate prediction. The same way as it's very hard to engineer perfect features that detect whether something is an elephant or not. Right. So if you train in a purely data driven way, a neural network that starts attending from raw pixels, in our case, this would be a row, rows and columns and cells of the database and learns how to aggregate all that information into this, you know, representation, this embedding.

36:38that then tells me, is this an elephant or not? Or will this user churn or not? It's kind of the same effect, right? So the point is that with these types of models, we can get to the superhuman level of performance. And what I'm trying to say, we shouldn't be surprised by it because we've seen it before, just as I was now giving this example with computer vision that today nobody is surprised about. And for the machine learning database, structured data world, the same revolution is out there. And it's not surprising. It's amazingly natural. That's what I'm trying to say. Kind of just, you know, take a step back and looking at sort of the larger trend that's happened around the field of AI, especially in like probably the last decade or so where we've gone, you know, full neural networks.

37:20It's a lot more sort of this like bottoms up learning where we're kind of learning the patterns, you know, versus something where if you look at like the early days of like rule-based systems where we're kind of trying to encode all the rules that we could imagine into some sort of tree structure that we didn't navigate. Why do you think that this approach has ultimately been so much more successful than trying to kind of take our knowledge of that system and then encode it into a set of rules that a computer can execute? That's a great question. And exactly what you just said also applies, let's say, to this predictive modeling world, right?

37:56Where we cannot anticipate all the rules, all the hard-coded signals that give us that prediction. And I think the reason is the following, right? The world is so unique, so diverse. There is so many inputs that it's very hard to preconceive all of them and, you know, write them down as rules or signals or whatever we are talking. And, you know, data that kind of captures the world is the king. And let the neural network learn directly from the data, directly from those raw signals, as noisy as they are, how to combine them into the final signal. So I think the reason it works so well is because when we are designing these rules as humans, we don't understand, let's say, the noise.

38:42We don't kind of really understand the richness. And also, right, our visual and our perception neural networks are already kind of learning how to take these raw, noisy inputs that our bodies are sensing into some higher level representations, right? So it's very hard for us to reconstruct, right? It's all like the way we think it's through the neural network. So rather than having one neural network build deterministic rules, we should just train our artificial neural network on the same set of inputs and it will be very good. Yeah. I mean, in a lot of ways, it seems to, I guess, mimic the behavior of humans or maybe even animals to some degree in terms of how they learn.

39:22Like, it's not like, you know, if you have a baby and it's trying to learn how to navigate the world, you know, the parent is giving them like, here's your set of rules that you need to like execute in order to crawl across the floor. Like they're kind of doing that more through, I would say like plain experimentation and like the, it's almost like a reinforcement warning or something like that. Exactly. And then, of course, right, I think when you talk about rules, I would say those are also in some sense important. But the way I think of those is more like explanations, right? When the model, when the neural network makes a decision, makes a recommendation, you can always ask why.

39:55And at that time, it's important for it to give you some general rules, some general reasoning, what led it believe to make this type of prediction, recommendation. Also, what is important, I think, in this machine learning or this predictive modeling side of things is having very strong accuracy estimates. You cannot have hallucinations. It really needs to be kind of data-driven and rooted in the patterns that are in the data. And with these models, you can achieve that as well. Now, this idea of the relational deep learning, you're applying this to databases, helping understand tabular structure, but there's lots and lots of other domains that can be described in terms of graphs.

40:37So can this approach be applied to other sort of domains? Even if I think about object-oriented programming, I can describe the class structure and the relationship between certain objects as a graph. I'm sure there's things in, you take this to the world of biology and biological structures, I'm sure a lot of those things could be described as graphs. Is there a way to apply this same approach to other types of domains of problems? Oh, definitely. I think a lot is very interesting, right? If you think about biological structures, molecules, proteins, and things like that, those are essentially graph structures, right?

41:11And even like models like AlphaFold and so on, they are really, in the end, learning to reason over this graph of amino acids and this graph of proximities as the protein starts to fold. The same happens at the scale of, let's say, small molecules for drug development and so on. So I would say this, the graph view of the world becomes very important. Whenever you basically have a set of entities, objects, this can be atoms, this can be amino acids, can be users, products, whatever those entities are that interact with each other, right? Like, and, you know, now what are you predicting? You can be predicting or estimating the toxicity of the molecule.

41:49You can be outputting the 3D structure, the 3D coordinates of the protein, or you can be doing fraud detection over this interconnected graph of entities to understand whether a transaction is fraudulent or whether the account has been taken over or something like that. In the end, it's kind of mathematically, it's all the same. And that's what makes this so beautiful. With foundation models today, more and more of these are multi-modality where they can handle images, audio files, text, of course, video even. Are we eventually going to move to a world where also essentially tabular data or semi-structured data is part of one of those modalities?

42:33I think so. I think that's where this is going, right? You cannot textify an image. You cannot textify a table. It just makes it kind of super hard to learn over that, right? But in principle, you could. You could just write the image as a sequence of RGB values in ASCII and say, you know, tell me, is this a cat or not, right? Like, so, but the point is we figured out that we need special domain-specific, modality-specific encoders. And I think where this is going is now that, right, basically we have these modality-specific encoders on top of the reasoning-based large language models, right? For images, for videos, and of course now also for this structured tabular data.

43:14Yeah, I mean, even in audio, we saw a huge performance increase when we stopped textifying audio and we actually trained models directly against the audio because there's so much nuance in audio versus what might be just available in the text. So it ends up, I think, leading to, when you textify it, essentially, you run into problems where the model doesn't really understand certain pauses and stuff like that. I think it's a beautiful example. And also, right, just shows that, you know, the two of us right now are just talking. So everything we say should be just captured in words, but it's actually not, right?

43:45Like you see if when you work on the raw signal, on the raw speech signal, the performance, there is more information there, the performance goes up. Yeah, absolutely. Well, Yuri, is there anything else you'd like to share? You know, what is sort of some of the, from your perspective, like the big problems that need to be solved in this space? I think the exciting part is that, you know, I feel like the structured data world has been kind of a bit left behind, right? People know how to run SQL. And then for everything else, they feel like they need to build manual, machine learning models super painfully and super slowly.

44:20And I think what we discussed today, the exciting thing is that we have now foundation models for structured relational data. For example, Kumo RFM is one such example. We have this approach of relational deep learning. We have the transformer architectures that can now learn directly over this structured tabular data. So actually, the structured data world that has humongous business impact is ready for the AI revolution to actually take place there as well. And today we shed some light on it, and that's what I'm very excited about. Yeah, absolutely. Well, thank you so much for being here. This was really, really interesting.

44:55Yeah, amazing. Thank you for having me. Cheers.

From the publisher

Predictive modeling is a core element in modern systems, and powers capabilities such as fraud detection, loan approvals, and recommendation systems. These systems typically operate on structured, relational data stored in enterprise databases, with rows, columns, and interlinked tables. While computer vision and natural language processing have undergone a neural network revolution, the tabular data

The post Foundation Models for Structured Data appeared first on Software Engineering Daily.

More from Software Engineering Daily

All 195 episodes
Foundation Models for Structured DataSoftware Engineering Daily · 44 min
Listen in VO