In short
Relational foundation models for enterprise data. Jure Leskovec argues that instead of flattening relational databases into single-table features, models should learn directly over multi-table relational structure by treating tables as a graph and using graph neural networks/transformers. He also describes Kumo’s “relational foundation model” (RFM2): a frozen, pre-trained model that performs in-context learning over database subgraphs to answer arbitrary prediction tasks without training or feature engineering.
Guest
Jure Leskovec, co-founder and chief scientist at Kumo; professor at Stanford University. Research includes AI for science via AI Virtual Cell (self-supervised foundation models for cells/patients/molecules using single-cell RNA-seq) and computational social science on network-based modeling (e.g., COVID spread).
Key claims
RFM2 can make accurate predictions on “any database and any predictive task” without model training; relational deep learning removes manual feature engineering and improves accuracy via feature discovery over graphs; gains reported as ~5% over supervised SOTA on benchmarks and ~12% after fine-tuning; works well with cold-start/noisy/incomplete data.
Notable examples
fraud detection/anti-money laundering; recommender systems (link prediction for ads/products); DoorDash restaurant recommendations/notifications; Reddit ad click-through rate lift; Coinbase fraud models over the Bitcoin blockchain; Databricks/Snowflake sales conversion modeling.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOIntroduction to Relational Foundation Models
0:00 to 0:45
Learn about the recently released relational foundation model and its capabilities.
“The recent breakthrough that we had and we just released in the second version is our, what we call a relational foundation model.”
Exploring Jure's Research Focus
1:08 to 1:20
Discover Jure's research interests at Stanford and his focus areas.
“We're going to be digging into your work on relational learning.”
AI Virtual Cell and Biomedical Data
1:21 to 3:02
Learn about the AI Virtual Cell project and how it aids in biomedical discoveries.
“Tell us a little bit about your research focus.”
Training Models for Patient Representations
3:03 to 4:48
Understand how models aggregate data from molecules to represent patients.
“are much more robust, driven purely from the data.”
Single-Cell RNA-seq Data Insights
4:49 to 6:28
Explore the significance of single-cell RNA-seq data in understanding cellular behavior.
“Like cell types, cell states, relationships between them.”
Linking Blood Analysis to Patient Health
6:29 to 7:38
Discover how single-cell analysis from blood samples can indicate overall health.
“depending on its type, depending on its state, and things like that.”
Relating Biological and Relational Data
7:39 to 7:50
Jure draws parallels between biological networks and relational data structures.
“Let me tell you a story why this is not so different.”
Understanding Network Interactions in Biology
7:51 to 9:26
Learn how biological systems can be viewed as networks of interactions.
“And, you know, where I started doing this was actually in a third domain, which is computational social science, right?”
Transforming Structured Data with AI
9:27 to 11:04
Explore the transformation of structured data through new AI techniques.
“It's again, cells coming together, talking to each other, organizing in a given way so that my skin has a given structure, it has a given set of layers and so on.”
The Evolution of Machine Learning Approaches
11:05 to 13:18
A deep dive into how machine learning models have evolved over the years.
Show all 33 chapters
Relational Deep Learning Explained
13:19 to 14:00
Discover the principles behind relational deep learning and its implications.
“and when we came up with this idea of relational deep learning, our goal was to fundamentally disrupt this and say, hey, why can't I just learn directly over raw relational data?”
Understanding Multitabular Data and Graph Neural Networks
14:00 to 17:48
Learn how to use graph neural networks for effective predictions on multitabular data.
“bought product id that at this time for this price for example right and that's a three table super simple schema and of course organizations have schemas of 50 60 tables and more depending on their complexity.”
Applications of Graph Neural Networks in Fraud Detection
17:48 to 23:56
Explore the various applications of graph neural networks in detecting fraud and predicting customer behavior.
“Rather, just bring the raw data, have a neural network and get better results that way.”
Benchmarking Multi-Table Prediction Problems
23:56 to 26:06
Discuss the lack of benchmarks for multi-table prediction problems and the efforts to create them.
“problems or their established benchmarks for multi-table prediction problems?”
The Relationship Between Research and Kumo
26:06 to 28:00
Discover how research in relational deep learning informs the commercial platform Kumo.
“Once you aggregate, you lose information.”
Introduction to Relational Foundation Models
28:00 to 29:10
Understanding the breakthrough of relational foundation models and their capabilities.
“But the recent breakthrough that we had and we just released in the second version is what we call a relational foundation model.”
In-Context Learning Explained
29:10 to 30:20
Exploring how in-context learning allows the model to make predictions.
“Because it's easy to say, oh, it's a foundation model, you who who, right?”
Extracting Examples for Predictions
30:20 to 33:10
How the model extracts labeled examples from databases to make predictions.
“And then this is passed through the relational foundation model architecture forward to kind of label the unlabeled graph, right?”
Mechanics of the Prediction Process
33:10 to 35:50
Details on the mechanics of how the model makes predictions without traditional training methods.
“It's just a single forward pass in which the neural network kind of in itself builds the model in a sense that gives us the accurate prediction, right?”
Architectural Innovations in Modeling
35:50 to 38:20
Discussing innovations in model architecture and their implications for relational data.
“I feel like I've gone the full cycle from that's an outlandish claim to, oh, yeah, I can see how that will work.”
Data Requirements and Model Efficiency
38:20 to 41:20
Clarifying data requirements for effective model operation and prediction accuracy.
“at the individual cell levels of a database.”
Generating In-Context Examples
41:20 to 42:00
The importance of generating in-context examples for optimal model performance.
“So in a sense, you'd say, oh, if I'm, I know, predicting fraud for, I don't know, for me, then you could say, oh, let me put some other Stanford professors in my in-context examples.”
Understanding the Model and System Integration
42:00 to 43:35
Learn the importance of both the model and system in generating in-context examples.
“I'm thinking about the line between kind of the model and the system.”
Performance Evaluation and Benchmarking
43:35 to 45:52
Discover how the foundation model improves state-of-the-art performance on benchmarks.
“It's like just build the best model you can and see how high you can get.”
Addressing Cold Start Problems
45:52 to 47:51
Explore how the model handles cold start problems and predictions with limited data.
“Can you talk a little bit about the process for deploying it?”
Deployment Processes in Real-World Use Cases
47:51 to 48:58
Learn about the deployment processes of the model at DoorDash and Reddit.
“super optimized feature engineered pipeline.”
Enhancing Click-Through Rates Through Collaboration
48:58 to 51:15
Understand how collaboration with Reddit improved advertising models and click-through rates.
“I don't think we tried turning those off yet.”
Explainability and Interpretability in Models
51:15 to 54:38
Discover approaches to explainability in models and the importance of interpretability.
“You mentioned we were talking about the recommendations.”
Challenges in Implementing Predictive Models
54:38 to 56:04
Examine the challenges related to integrating predictive models into business processes.
“Another, right, like maybe that's one use case.”
Integrating Models into Downstream Systems
56:04 to 58:10
Learn how to connect AI models with operational applications effectively.
“There is still the problem of how now that we have the model, how are we pushing that to, I wouldn't say to production.”
Predictive Modeling for Customer Behavior
58:10 to 1:00:06
Understand the significance of predictive modeling in customer interactions.
“But every one of them has a bit different definition of what churn is.”
Post-Training and Fine-Tuning AI Models
1:00:06 to 1:01:26
Explore the benefits of post-training and fine-tuning for AI models.
“So that's one reason you would want to, let's say, train, even in a task-agnostic way, over the underlying data to better capture distributions, to better learn priors.”
Challenges in Automated Model Building
1:01:26 to 1:05:10
Discover the challenges faced by AI in automated model generation.
“We are very excited about agents, both basically surfacing these two agents as tools The second thing is right now, right, like the coding agents are out there.”
Transcript
Automatic transcript. May contain errors.0:00The recent breakthrough that we had and we just released in the second version is our, what we call a relational foundation model.
0:10Jure Leskovec:And that's a pre-trained foundation model that can reason over structured relational data. And it's crazy what this model can do. It can make accurate predictions on any database and any predictive task without any model training.
0:45All right, everyone. Welcome to another episode of the TwiML AI podcast. I am your host, Sam Charrington. Today, I'm joined by Jure Leskovitz. Jure is co-founder and chief scientist at Kumo and a professor at Stanford University. Before we get going, be sure to hit that subscribe button wherever you're listening to today's show. Jure, welcome to the podcast.
1:08Jure Leskovec:It's great to finally connect with you. Yeah, great to be here. I'm looking forward to our chat. We're going to be digging into your work on relational learning. as well as some of the other interesting things you're up to at Stanford and around AI for science and more. But let's start there. Tell us a little bit about your research focus. Yeah, great. So I'm professor at Stanford here in the computer science department. You know where the future happens, I like to say. So there's always exciting research going on. Our focus recently has been, I would say, on two areas. First is AI for science.
1:53Jure Leskovec:And in particular, we have a project that we call AI Virtual Cell, where we are basically building next generation foundation models that allow us to represent human cells, patients, as well as individual molecules in cells and allow us to reason across this complex biomedical data. for discovering new cancer therapies, molecule design, reasoning about all different biomedical data modalities, objects, and how they interact with each other to help speed up science. So it's everything from foundation models at the lower level of understanding proteins to then models that aggregate the, let's say, the molecules in the cell to represent a single cell.
2:47Jure Leskovec:And then the next level models that now say, oh, you know, a tissue or a patient is a collection of cells. Cells are a collection of molecules. Let's build models that just aggregate all this knowledge in a very faithful representation, let's say, of a patient. And that helps a lot because now the representations we have are much more robust, driven purely from the data. no biology in some sense is inserted in the model everything is emergent out of the data and it's amazing how much we can we can learn from that so that i would say is one line of one line of work we've been working on and and i i can't help but hit pause and ask like do you train this all end to end are you training an individual model or representation at a time and then aggregating it after you've got these models defined?
3:39Jure Leskovec:Great question. So the way we are doing it right now is actually, the first scientific question was, is this even possible, right? Could you say a cell is a representation of molecules that are inside the cell? So now let's say molecules inside the cell are the proteins. I can use the protein language model to now represent every protein in the cell. And now the cell needs to aggregate information from all these proteins to say, I am a cell, this is my state. Now that you have a representation of the cell, you can build, let's say, a patient-level model that says a patient is a collection of cells in given states that are composed from the proteins that are in there, and are we able to kind of collect this information over these orders of magnitude different scales to get a strong data-driven representation of the, let's say, underlying patient in this example.
4:32Jure Leskovec:And the interesting thing is that this is purely doable and it's trained purely in an unsupervised, self-supervised way, right? So you don't need to insert any human bias, any human knowledge of biology. The biology emerges from the data itself, right? Like cell types, cell states, relationships between them. that kind of human biology, how we describe it, actually emerges directly from the data, right? So the model learns how to best describe the underlying processes and phenomena without us pushing it on it from the top. That's kind of the exciting, interesting kind of emergent capability there.
5:19And tell me if this question makes sense. I think it's related to the way you're describing the training process. But is the data set that you're training on mechanistic in nature or behavioral in nature? In the sense of like, are you observing some behaviors of cells and then training on that data? And there, you know, some kind of faithful representation of mechanisms are emergent? Or does the data have mechanistic properties to it?
5:49Jure Leskovec:The data we are using in this case is called single-cell RNA-seq data. This is data that large international consortia are collecting. But basically what it says is that you can take some sample from some, let's say, some tissue. And then for every cell in that sample, you measure the number of different protein molecules inside that cell. So every cell is now represented by a 20 ,000-dimensional vector that tells me the abundance of that specific protein in that specific cell. And every cell has different, let's say, ratios of these proteins depending on its type, depending on its state, and things like that.
6:37Jure Leskovec:So that's the raw input data. And then, of course, because we know what the protein is, We can actually bring the protein information through ESM or through AlphaFold. And now it gets, and then it gets very interesting. So being protein-based, that brings in both mechanism and behavior. Exactly, exactly, exactly. And then, of course, you can connect this all the way to the phenotype because what we are doing now, you know, is we can take a single drop of blood from a patient. And rather than running kind of a classical blood screen, we can do this single cell RNA-seq analysis. So now we basically can profile every single cell inside a drop of blood.
7:17Jure Leskovec:And why blood is interesting is because it circulates through the entire body. It kind of captures the state, the immune state of the entire body. So we are able to detect diseases, understand patient trajectories and things like that just from this digital twin of a single drop of blood. Super interesting. Also very different from the other thing that you focus on, which is relational data. Yeah. Let me tell you a story why this is not so different. Okay. Okay, so what I'm really, you know, what I'm excited kind of fundamentally or how do I approach things is, you know, to always kind of take them apart and understand how different parts interact and how different parts work together.
8:05Jure Leskovec:And, you know, where I started doing this was actually in a third domain, which is computational social science, right? I was very excited about how do people interact with each other. And when I started my research career, kind of social media just started as a phenomenon. And my view at that point was like, I can use social media as a telescope into human behaviors. So I now can study human behavior through these digital traces that people produce through using cell phones, through using social media and so on. And it's all about networks, graphs of people interacting with each other. And what that, for example, allowed us to do is not only understand phenomenon on social media, but also, for example, model the spread of COVID pandemic super accurately.
8:55Jure Leskovec:And we were able to computationally analyze and predict how the virus will spread as we reopen the economy. If we increase the occupancy levels at different, you know, restaurants, gyms, churches, whatever the locations, to say this is how the virus would spread, this is what you can do. And that underlying was a network. So now, what is biology? Biology is a network, right? It's all about molecules coming together to do something in a cell. And then what's a tissue? It's again, cells coming together, talking to each other, organizing in a given way so that my skin has a given structure, it has a given set of layers and so on.
9:40Jure Leskovec:So that's essentially a network, a graph of interactions as well. And then you mentioned relational data, so data that sits in tables in a database that every enterprise in the world has, and it's kind of the most valuable data. That's also a graph of, or capturing a graph of interactions of different entities inside that organization. When I look into some of the work you're doing around the relational deep learning, it calls to mind other conversations I've had that focus on like deep learning for tabular data, But that tends to be focused on like that single table as opposed to these relationships that arise in, you know, an enterprise data where you've got, you know, different tables that are linked by keys and whatnot.
10:33You can talk a little bit about how these two areas of research and practice relate to one another.
10:40Jure Leskovec:Let me explain, right? So maybe first, if we think about machine learning, it hasn't really changed over the last, I would say, 30 years. No, not really.
11:00Jure Leskovec:We used to have, I don't know, we had decision trees, then we had support vector machines, then people like logistic regression, then people are like, oh, we'll build these neural networks. then we said oh we have gradient boosted trees they are better and and things like that right but fundamentally it has always been you have your data you feature engineer this single table of your features you add a label and now you train some supervised model that from the features predicts predicts that label right and we've been doing that over and over again and maybe this predictive model you know it's a it's a deep model we would call but um it's a neural network but what i would argue is it that ai has not transformed this structured data space in the same way as uh computer vision or natural language understanding have been fundamentally transformed by ai okay and let me let me quantify what do i mean by that right like what was the big breakthrough both in computer vision as well as in natural language understanding it was about let's build neural networks that learn directly on the road data right in the old days you would do in computer vision you would do all kind of feature engineering shift features gabor filters and it'd be like i'll describe this image as well as i can so i can then predict you know is there a is there a car on the image or not right um in in nlp was similar right like we we you know ibm uh won jeopardy with uh with their system but it was all super hand engineered manual and so on but you know kind of it worked right but it took 300 people to build build it and was very you know was was great but was very kind of uh brittle right so again the transformers they just learn over tokens no grammar, no syntax no, it's just, you know, learn over tokens, right?
12:59Jure Leskovec:Again, a neural network directly on the raw data The same thing is actually not happening on structured tabular data Right there, we don't learn on raw data We run all these SQL queries, all this ETL all this feature engineering to then come up with a set of signals from which we, let's say, try to predict something and when we came up with this idea of relational deep learning, our goal was to fundamentally disrupt this and say, hey, why can't I just learn directly over raw relational data? And why do we always have to learn over data in a single table? And the point is that as I take this multitabular data, and just to be very precise, right, what's a good example of this?
13:47Jure Leskovec:It could be like I have a set of customers, I said have a set of products so these are two tables each customer has an id each product has an id and maybe I have a third table that's a set of transactions that says customer id this bought product id that at this time for this price for example right and that's a three table super simple schema and of course organizations have schemas of 50 60 tables and more depending on their complexity. So our question was, how could I just learn directly with a neural network over this multitabular data? And the answer is, you know, kind of surprisingly simple, is to say, just think of the database, think of these tables as a graph of relationships between the entities in the database.
14:39Jure Leskovec:So this would mean in my, you know, I'm a graph person, so I like to think in terms of graphs, right? So graphs are composed of vertices, the nodes. This would be my users, would be my products, would be my transactions, and so on. So this would be now the nodes. And then the connections are just saying this user ID was part of this transaction that was part of that product. And now we have a path from a user to the transaction to the product. And then, you know, another user or another transaction is another path in this very simplistic graph. And now that we have a graph, we can basically apply graph deep learning, like graph neural networks, which is a way to generalize deep learning to graph structured data and just train over that to get an accurate prediction.
15:30Jure Leskovec:And what happens is two things happen. The first thing that happens is you don't have to do manual feature engineering anymore, right? So it's much faster, it requires much less effort to train these models. And the second thing that happens is your models are more accurate. And then you say, why can my models be more accurate? And the answer is very similar to what happens in computer vision, right? If you are saying, I am a human, I know what a car is, So I will build perfect features that detect whether there is a car on the image or not. I know cars. I drive them. I'm such a car expert. I can build the best features for detecting cars.
16:13Jure Leskovec:Nobody in the right mind claims that, right? But, you know, in machine learning, data science prediction, people are still saying, no, I'm the domain expert. I'll engineer the features. your features are just some arbitrary human biased summary statistic of your data that you know you kind of dreamt up with put put it as a feature in the in in your training table retrain the model and then you saw whether that increased the accuracy or not right and a neural network that trains with gradient descent is able to do so much more nuanced almost like feature discovery by basically attending over this graph to extract much more signal so we see this double digit increases in model accuracy because the neural network is able to extract more signal out of the raw data right and i will just you know full transparency right if you are working on a super simple problem that falls on a line then no neural network is ever going to be better than a linear model right so what i'm basically trying to say i cannot guarantee that always you will get better performance because sometimes the data is linear and if you happen to train the linear model to it you already have good performance there's nothing more you can do right but majority of the data is not linear, is much more complex, and that's where the benefit happens.
17:49Jure Leskovec:So that's kind of the key idea behind relational deep learning, is that now we can have neural networks just learn directly on the raw database data, don't need to build this manual feature pipelines and feature stores that are super painful and lead to so many different kind of bugs and inconsistencies and information leakage and time travel makes putting models in production super hard. Rather, just bring the raw data, have a neural network and get better results that way. So that's kind of, let's say, the philosophy and the reasons why we are doing this. Can you give us some examples of the types of things you're trying to predict with these models?
18:36Are you trying to predict things that are primarily about structure or are you trying to predict individual values? How do you think about what the models are capable of?
18:47Jure Leskovec:That's a great question. The way I describe the framework right now, it's very generic in a sense that you can bring any set of tables, any set of connections between them, any set of columns. The underlying mathematical representation kind of remains the same. And of course, the underlying graph changes, but the graph neural network or the graph transformer can be applied to that. So what would you want to predict? Depends on the data. If you have a transaction graph, for example, where we see great results is on fraud. All kinds of fraud detection, anti-money laundering, account-level fraud, transaction-level fraud.
19:27Jure Leskovec:work is beautifully right you just bring this heterogeneous multitabular data together and just learn over it what fraud is and fraud is interesting because it's so non-stationary you know fraudsters are trying to game the system all the time so as a as a machine learning engineer you are always behind your model is always deteriorating and you're like okay how do i design a next feature how do i design the next feature you have a neural network if you just pick the signal directly out of the raw data so fraud is an example fraud you can think of let's say as a classification task then you can think a lot around regression type tasks for example for customer behavior in terms of customer churn next best action things like that and right and then you can also think about in graph terms about link prediction so predicting links between two types of entities.
20:24Jure Leskovec:What's the canonical task there is the recommender system, because it's predicting a link between the customer, the user, and the product. So we've seen great uses of this in recommender systems for ads, product recommendations, and things like that. And historically, when I've talked to folks about deep learning, machine learning for or tabular data, the results were, I don't know the best way to characterize this. Like, I always get the impression that we're not quite there yet. And would you say the same is true for what you're doing? Or is it an issue of like, there was this missing link and that missing link is the graphical structure and now we have it and we're able to do much more?
21:14I'm trying to kind of, you know, ground what you're, you know, working on and saying with, you know, this kind of broader results of applying, you know, these techniques that have shown, you know, to be extremely effective with text and images to tabular data.
21:33Jure Leskovec:That's, I think that's a great point, right? Like when we say tabular data and tabular machine learning, this is the community that works on single table problems, right? The data has already been flattened, pre-formatted, summarized to fit in a single table. We're trying to get better results than we might get with XGBoost or something like that, right? Yeah, right. And what we see there is that these models and architectures and so on, kind of deep learning, right, didn't really displace XGBoost. XGBoost is still kind of the workhorse. maybe on individual examples yeah you can do better but it's still the the the workhorse and the reason i'm kind of less interested in this single table problem is because that's not the right problem to solve i don't know any organization that has older data in a single table right so the hard part and where the information gets lost is when you go from this rich relational structure into the single table and once you are in a single table you know then we are you know then we are kind of talking almost like second order effects did you use this architecture did you use that architecture did you use this tabular model or this tabular foundation model or not right like all the information is there in that single table and all the methods are about equally good at extracting it right i think where the difference happens is if you actually make a step back and say hey single table model or is not the heart is not the heart part is not where is not in a sense uh general or realistic enough where you need to go you need to go to the multi-table setting because that's truly now the raw data you have it's not some summarized featurized data it's the raw data and and there is much more signal there that got dropped when the data got flattened into summarizing to a single table.
23:36Jure Leskovec:So to me, single table problems are, you know, are solved. I think the differences are kind of second order effects. What is unsolved is the multi-table problem. That's where the wins are being hidden. And so how do you think about benchmarking performance for these types of problems or their established benchmarks for multi-table prediction problems? Actually, there is quite a lot of single table data out there because of all the history of machine learning. And I think even when people develop new benchmarks from raw data, they just release that single table because everyone learns on the single table, right?
24:19Jure Leskovec:So what we did actually at Stanford, we were like, okay, so where is a multi-table benchmark? And there is no multi-table benchmark and even if you look at kegel out of thousands of competitions of on kegel you know there are four that are multi-table all the others are features have already been engineered for you there is a single table and you start you know begging and boosting and and creating tricks until until you win right um so we created a benchmark at stanford by collecting and curating open, multi-tabular data sets that we were able to find on the web. We call it RELBench. We have now two versions of RELBench.
25:00Jure Leskovec:It's about 40 different predictive tasks over, I think, about 10, 15 different databases. And then what's also interesting is that SAP, the big German IT company, they released a benchmark, a multi-tabular benchmark of enterprise data called SALT. So those are, I would say, the two big tabular or multi-tabular, so relational benchmarks. SAP from SALT, SALT from SAP, and the Relban line of work that we've been doing and promoting here through Stanford. It also makes me wonder if there's a way to reuse existing benchmarks by like denormalizing, you know, wide single tables or something like that.
25:53Is that something you've looked into?
25:55Jure Leskovec:That's a great point. Like you can try to denormalize, but if you think about it, you can only denormalize one-to-one relations. As soon as you have many to one, you have to aggregate. And that's the key. Once you aggregate, you lose information. That's where you've lost information. Exactly. Exactly. And I know I can dwell on this point a bit, right? Imagine you are doing a churn model, right? So a customer churn model could be, I have a customer and here are historic transactions of the customer. I need to aggregate them. So first I say, I'll count how many purchases you made last month. And then I'll maybe take the median price of those purchases.
26:36Jure Leskovec:And then, you know, some other data scientist says, no, no, let's take the cheapest price of everything you want, right? And then somebody says, no, no, you should take the most expensive one. And then somebody says, no, it's the average. Another person says, oh, but distributions are skewed. We should take the media. Then another person wakes up and says, hey, it's about shopping in the morning. That's what's predictive of church. Let's add another feature, right? And then somebody says, oh, but we have to account for holidays. You're like, just give me the data. You know, like, that's what I mean, right?
27:07Jure Leskovec:And then you're like, oh, holidays. People sleep longer on holidays. Let's now create a new feature that accounts for holidays. Oh, but then there is summer daylight change. Let's account for that. You see how kind of ridiculous this gets? Just attend over the transactions and let the attention figure out what predict should. When I introduced you, I mentioned that you're a co-founder at Kumo in addition to the research. Talk about the relationship between the research and what you're doing at Kumo. What we built at Kuma is a commercial enterprise-grade platform that allows us to do large-scale relational deep learning models.
27:44Jure Leskovec:And we are using this platform to two effects. One is to allow partners, customers to train, tune single-task models over the multitabular relational data. and I can talk about that part. But the recent breakthrough that we had and we just released in the second version is what we call a relational foundation model. And that's a pre-trained foundation model that can reason over structured relational data. And it's crazy what this model can do. So what this model can do, It can make accurate predictions on any database and any predictive task without any model training. And I find that proposition to be almost outlandish.
28:44Like they're just numbers with some unknown relationship. And you're going to say that you're going to train a model on just the relationship between random business numbers. And it's going to work in some unknown use case. Make that make sense to me.
29:00Jure Leskovec:Thank you. Thank you. I think it's great. I think as I say this, people who listen should be like, what is this guy talking? So thank you. So I agree, right? Because it's easy to say, oh, it's a foundation model, you who who, right? Great. But then, okay, what does it really do? So here's maybe how to think about this. So the key here is to do in-context learning, right? The same way as a language model does in-context learning, where I give it a prompt, I give it the information, I give it a task, and then it gives me the answer. So what we do here is the system has several components. So there is the database.
Read the full transcript
29:41Jure Leskovec:And then there needs to be a way for me to instruct the pre-trained foundation model what kind of prediction I want. right i want to say predict me the sum of purchase prices over the next one month for this particular customer and that maybe is like how much i'm predicting how much the customer is going to spend or i'm saying predict me you know transaction dot is fraud equals true for transaction id this much okay so this would be like predict me whether the transaction is fraudulent for this particular transaction id right so i have a way to specify my predictive task and now what the system does the system now goes into the database it extracts a set of labeled in context examples that then get passed through a pre-trained neural network to make a prediction okay so now when i say a set of labeled in context examples this means that you can take the task for the example of fraud i've got historical fraud that's already been labeled and i've got some new transactions coming in that don't have that label attached that i'm trying to to predict for example okay so let's do fraud fraud might be easier yes so the way this would work right if i say predicting the probability of fraud the system would go into your into your database and extract previous transactions for which we know whether they are fraudulent or not for each of those transactions we would then extract kind of the subgraph of entities around it okay so now what the relational foundation model gets on the input is a set of historical subgraphs of previous transactions and their fraud labels, plus the new transaction that is unlabeled.
31:35Jure Leskovec:We don't know it's fraud. And then this is passed through the relational foundation model architecture forward to kind of label the unlabeled graph, right? Like the unlabeled transaction and the graph around it. I'm not sure now that fraud is a good example because it's kind of, I can see how that could work. Like you've collected the graph around these known points and you're asking a model to infer relationships that might lead to this one individual label. And so maybe, I think maybe, and this may be where you're going, like I think a regression type of a problem would strike me as more challenging than a classification problem.
32:26Jure Leskovec:Yeah, I think the key here is, right, what are maybe the key components here is first is that you have a language where you specify the task. We can go generate almost like a mini labeled training data set, these in-context examples. and then you have a pre-trained model that is able to take these in-context examples, these subgraphs that have some certain columns and tables and so on, is able to encode them in a domain-agnostic way. And then the neural network is able to essentially build a predictive model in its brain in a forward path to give you accurate prediction. Right, right, right. Right.
33:09So it's not necessarily about like some universal understanding of numbers or what have you. It's about being able to identify the right relationships between numbers that it hasn't seen before, query the right, you know, examples and create the right universe and then formulate that as the right, I guess, like inference request or something.
33:34Jure Leskovec:exactly so there i would say two aspects to this one is how can you take data and encode it in a domain agnostic way right because we can take any database any set of uh any set of columns the model needs to be able to encode that in a in a universal way and now that it's been encoded then the second step is to perform in context learn so it means that the modeling its brain needs to be able to build a model, right? There's no training. There's no backpropagation. There's no gradients. It's just a single forward pass in which the neural network kind of in itself builds the model in a sense that gives us the accurate prediction, right?
34:19Jure Leskovec:So no training is necessary. No hyperparameter optimization is necessary. No feature engineering is necessary. All you need is a raw database and a way to specify the task. Does the model require some type of memory structure, blackboard or something in order to, you know, do a scratch work to come up with a representation or is this all like thought traces or something like that? No, no, no. This is not an agent. This is a single forward pass of a transformer-like neural network, right? So this is purely inside the neural network. There is no agent. There is no memory. There is no scratch pad.
35:02Jure Leskovec:There is no, let me do this, let me do that. The answer is truly a single forward pass of a neural network. There is no loop, nothing like that. So you get the answer in, I know, 0.2 seconds, half a second, whatever the time be. It's really a single forward pass of a pre-trained, frozen neural network. There is no language model here, right? This is kind of technology that's parallel or complementary to language models, right? You cannot textify a database and then go to ChatGPT and say, hey, what do you think? How likely is this transaction to be fraudulent? You get horrible results, right? So this is, yeah, frozen pre-trained architecture that allows you to do that.
35:51I feel like I've gone the full cycle from that's an outlandish claim to, oh, yeah, I can see how that will work. I don't know. It's still kind of crazy that it works.
35:59Jure Leskovec:No, it's interesting, right? And when we test this on data sets that are locked away and hidden and the model has never been trained on and on tasks that we haven't even thought about, we see a gain over best supervised models out there. If you would go and say, I'll hire a data scientist, they'll spend several weeks building the model, tuning the model, the latest neural networks, whatever, it's still a couple of percentage points worse. And then if you fine-tune the, let's say, the foundation model on more data for the specific task, then you get to this superhuman accuracy performance that, you know, present manual or semi-manual or agentic solutions are just not able to attain.
36:49That is the RFM2, Kumo RFM2, the relational foundation model. You also recently published that iClear relational graph transformer. Is the one based on the other or are they independent lines of research?
37:08Jure Leskovec:What I would say is at Stanford, we are pushing forward in the open new architectural improvements, understandings and as much as we can as academics we release everything open source we talk about everything and then of course what happens inside the company is that some of these innovations that we put out also also kind of diffuse inside I would say that internally the architecture we are using is a bit different it's composed of two different parts. The first part is basically it's the encoding or the attention mechanism over this set of tables. And then the second part is this in-context learning type machine.
38:02Jure Leskovec:There are two papers that are relevant here. One is the relational graph transformer that we mainly use for supervised fine-tuning type tasks. But then another paper we also published at iClear it's called relational transformer. And that one actually allows for in-context learning. So that one does attention all the way at the individual cell levels of a database. And essentially you have three types of attention. You have attention over a given column. So if you are interested in a cell, by attending over other cells in that same column, it kind of gives you a sense of a distribution, right?
38:42Jure Leskovec:Then we attend over the cells in a row. And that kind of then gives you a sense of what's the information in that row. And then we also have a graph-based attention mechanism that allows you to say, oh, this is a user and these are all their transactions. And then each transaction is a row and each row has columns this way. So this means that we can be attending over millions, tens of millions of cells. And the beautiful thing is that our attention mechanism, because of the graph, has much more structure. So the attention mechanism is never quadratic. And this means we can compute much more effectively.
39:21Jure Leskovec:And to do good reasoning, you really need humongous context sizes, right? Even the largest LLMs today, I know, go to a million tokens. For us, a million tokens is small. So I was going to ask, are there data requirements or shapes or use cases that this works well for or conversely doesn't work well for? It sounds like part of that is size, like you need a lot of data in order for this to work. Is that fair? i would actually uh maybe push back on that a bit actually because the model is pre-trained it can do amazing things where you have very little data because it's you know like training models from scratch yeah requires a lot of data but once the model is pre-trained it kind of knows what functions kind of appear in nature so it means that you can give it few examples and it's going to give you very accurate predictions more accurate predictions than something some you know supervised model that you have to trade.
40:33I think I picked that idea up based on you saying that the context that you work with is typically large. Is that saying that when you have a lot of data available, you can use it, but you don't necessarily need it?
40:48Jure Leskovec:Exactly. And then once, if you have a lot of data, you can either increase the context size and by increasing the context size, you get more accurate predictions. Or if you are saying, oh i'm doing fraud you can just fine tune your model for fraud in a sense that you don't even have to do in context learning because you know your data you know the task you just tune the model for that single task and then the model can be smaller much more efficient to run and also more accurate because it doesn't have to you know almost like re-learn the task every single time because you give it the in context examples what we see works best is some mixture of pre-training and in-context examples because the way you're choosing context examples can actually depend on what the target entity is, right?
41:36Jure Leskovec:So in a sense, you'd say, oh, if I'm, I know, predicting fraud for, I don't know, for me, then you could say, oh, let me put some other Stanford professors in my in-context examples. Let me put some other Bay Area folks in here because, you know, that's kind of the, I don't know, the peer group or the most useful examples from which you can learn to make accurate predictions about, I don't know, me being a fraudster. I'm thinking about the line between kind of the model and the system. The system is what is constructing the in-context examples, and the model is just that forward pass. And you need both, right?
42:10Jure Leskovec:I think is important, right? Because somebody has to generate these in-context examples. You won't generate them manually, right? And is that part also learned or is that, you know, kind of a formulaic graph traversal or something else? Somewhere in between, you can do it as a form, kind of just as a graph traversal and a bit of kind of time travel, right, to generate the forward-looking labels. But of course, how you do that and what in-context examples you generate makes all the difference. So there's a lot that goes into that to get top performance. And so speaking of performance, you talked a little bit about some of the challenges with collecting benchmarks, but how do you find performance relative to those benchmarks and also, you know, more importantly in the real world?
43:11Yeah.
43:12Jure Leskovec:So I can say, right, like we have a white paper on KumoRFM2 that people can read with a bunch of different benchmarks. What we see is that the foundation model by itself improves state-of-the-art over all supervised models ever published on this benchmark, right? So the baseline is very high. It's like just build the best model you can and see how high you can get. So the foundation modeling improves that, I think, for about 5 % relative to the accuracy. And then if you further tune the model, meaning if you would fine tune it and do some grain and base updates, then the performance goes to 12 % over the state of the art.
43:59Jure Leskovec:And those are quite sizable gains, especially if you think about putting this in production in recommender systems or fraud detection where, you know, every single digit performance in increasing accuracy can mean millions, tens of millions in business impact. maybe the second thing i would say is where we see these methods also shine is with noisy and incomplete data cold start problems because of the relationships because of the relational structure the model is able to much better kind of hone in and be much more robust to the data missingness data corruption and things like that so we've also done quite a lot of analysis around understanding and like how this performs in real world data, sparse data, small amounts of data, noise, incompleteness, irrelevant columns and things like that.
44:59And when you mention cold start, like that suggests, hey, I want to start identifying fraudulent transactions, but I have no labels. I just have a bunch of data. Can you tell me where I should start looking? Like, Does it work for that kind of problem?
45:15Jure Leskovec:Yeah, maybe I should quantify what cold start means. Usually cold start would mean when a new user shows, a new product shows up, right? So you still need to have some historical labels. I'm not, you still need some historical labels, but usually, you know, prediction is easy once you have a lot of data about a given user or a lot of data about a given product. But when the product is fresh or when the user is fresh, you are data poor. That's what technically is called cold start problems. So I still need historical labels, but to make reliable predictions, I don't need much data. I think I saw somewhere that the system is deployed at places like DoorDash and others.
45:58Can you talk a little bit about the process for deploying it?
46:04Jure Leskovec:Yeah, great question. So the system, the platform, we can deploy it in many different ways. We can, you know, run it as a SaaS, basically as a compute platform. We can deploy it in people's private public clouds, like inside, we call it virtual private cloud. So all the data stays with the customer. There's a bunch, there's, I would say, a bunch of different deployments depending on what organizations like and prefer. and then yeah in terms of let's say deployments or use cases right at uh at doordash it's uh restaurant recommendations and the notification system which user gets what notification at what time of day and things like that and we've seen you know uh revenue impacting hundreds of millions of dollars um another another great client we work with is reddit so the advertising models on Reddit are built on top of or are built with Kuma.
47:05And it was nearly a double digit
47:10Jure Leskovec:increase in ad click-through rates. So basically the revenue, yeah, it's like unbelievable, right? Usually an entire team, you know, like increases maybe 1 % that accuracy year over year, right? Because click-through rate is... And this is your original point about like, you know, domain expertise and like manual features. Like you would imagine that they've been working on this for a long time and they've kind of squeezed a lot of the juice out of that lemon. But, you know, here comes the machine. Exactly, exactly. And it's actually interesting. And I mean, you know, we have a great collaboration and great relationship with the Reddit team and they're amazingly sophisticated.
47:48Jure Leskovec:And of course, they build their own super optimized feature engineered pipeline. And then the way we do it there actually is that we said, okay, let's take your data represent it as a graph and let's create embeddings for users subreddits ads and things like that so now these embeddings actually get appended to their to their own features right and even with that there was there was a huge increase in the click-through rate because this signal that the neural network learn was kind of complementary to what the human feature engineering already had. So actually the model that is in production is combined from the neural network embeddings by graph embeddings by Puma, as well as the manual feature engineering.
48:37Jure Leskovec:So that's been a great collaboration. So it's add the recommendation, click-through rate prediction, if you want to think of it that way. You know, or have you looked at like if their manual features really make a difference? Like, Is that a feel-good thing, like you left them in there because they had them? Or do they provide lift that's been measured? Good question. I don't think we tried turning those off yet. But it's a very interesting question. Sometimes you still want to have those features in there, not maybe for the model accuracy, but because you have so many business rules, you know, like these advertising systems, they are not just pure optimization plays.
49:26Jure Leskovec:There is so many kind of other business rules that need to trigger for the ad to be actually shown to the user. So you kind of need sometimes those signals to be able to trigger business rules anyway. And I wonder in this case and with those hand-engineered features and more broadly with RFM, you know, what kind of explainability story there might be. That's another reason why people like XGBoost is that those trees are fairly interpretable and that's been a challenge with, you know, transformer-based networks. Yeah, that's a great point. Actually, I would say we do explainability really well.
50:11Jure Leskovec:And it's even more models, I would say, are even more explainable than these tree-based models because in tree-based model all you can get is you get a ranked list of features right so you can only explain predictions by the with the features you you engineered what we can do is we can do this and those features might be you know wrong understandings of the data or incomplete yeah exactly so what we do is we do the we can do because the model is fully differentiable we can basically run the model backwards and we see what tables, what columns, what cells the model is attending over. And we get this structure-based explanation.
50:47Jure Leskovec:But then we use a large language model to say, here's where the attention is. This is what the columns are. And this is what the semantics is. And then we generate a text-based explanation that is like super readable, super friendly. Like a text-based explanation of a saliency map of the data or something like that? Think of it maybe that way, right? But the LLM is kind of enriching it with all the human world background knowledge and so on. So it gets very actionable. So you were talking about use cases. You mentioned we were talking about the recommendations. One last one, which is around fraud.
51:24Jure Leskovec:So we've seen great results with fraud. Here we've been partnering with an amazing team at Coinbase. So we have these models running in production at Coinbase on the entire Bitcoin blockchain network. Right. So also we can scale to, you know, the size of the entire edit, to the size of the entire Bitcoin or Coinbase, you know, right. So these methods really scale. But then, you know, with some clients like, let's say Databricks or actually Snowflake, they are using us to run their sales models, you know, predicting what the customer is going to buy next, which customer is going to convert into a paying customer.
52:05Jure Leskovec:and this allows them to optimize their sales team, right? And if you think about sales team data, that data is smaller because the sales teams are, you know, hundreds, maybe, you know, thousand people or so. So you can do well on small data as well. You know, there are aspects of it that sound free lunchy. Like, how am I paying for my lunch? Like, what's the, you know, is it, yeah, maybe I'll leave it there for you to answer. Yeah, I mean, at the end, what is being paid for is compute. What we are really doing is, right, like machine learning, if you think about it, is really a CPU compute based, right?
52:48Jure Leskovec:Majority of the compute that happens, except maybe the final neural network training, happens on the CPU. What we are doing is taking that workload from the CPU to the GPU. so now the amount of gpu compute is is is larger because it's computed over the raw data not over the summaries generated on the cpu um and and that's uh that's i would say what the uh what the what the cost what the cost is uh in the end of course these models and um um don't you know are not in trillions of parameters they are billion parameter type ones they can be quite small So they're actually quite efficient and cheap to run because, you know, the reason we are making predictions is to make decisions, right?
53:39Jure Leskovec:Our commander system is making decisions what to show to every user. So we are like the speed of those decisions is that, you know, tens of thousands, hundreds of thousands, millions of times per second. So performance cost really, really matters. Yeah, I think I was also trying to get at limitations. and if you had someone come to you with a problem that, you know, was in fact like multi-tabular relational, you know, what might be, you know, some reasons why, you know, you ultimately, you know, tell them that it's probably not a good fit? The way I would say is we know if we use our technology, we'll be at least on par or better to what is already there, right?
54:27Jure Leskovec:now where we see bottlenecks usually we see bottlenecks in actually getting the value out like connecting those predictions to some decision making downstream business process so that the value can be reliably measured that's been i think the biggest bottleneck right in a sense that models are are built developed they work great but then engineering teams need to hook them up to actually surface those predictions or make decisions based on those predictions. Another, right, like maybe that's one use case. Another use case is sometimes like where we shine or where the technology shines in these predictive, well-defined, predictive type problems that can be mathematically well formulated and optimized.
55:16Jure Leskovec:If it's more about, hey, we want to understand the patterns, We want to understand what is happening in the past. That is much more, you know, this kind of traditional data analytics-y type things or some pattern detection type thing that our platform and what we discussed is maybe not a good use case. So that's, I would say, is, you know, some examples. You need to know what you are predicting. You need to be able to formulate that. I need to be able to measure accuracy, and then we can optimize. And you're not going to solve the traditional data science problem, like in organizations, if you've got a model, how do you use it?
56:03That's still going to exist.
56:06Jure Leskovec:Exactly. There is still the problem of how now that we have the model, how are we pushing that to, I wouldn't say to production. That's easy. The question is, how do we connect it with the downstream app or the downstream system so that actually somebody is acting on these predictions, right? And where we also see, I would say, a lot of traction recently is in agentic workloads, right? Because agents need to make decisions to take actions. And now you can make decisions based on this, you know, LLM-based common sense. But for anything more, like the best way to make decisions is to estimate or predict their downstream effect.
56:51So now I'm envisioning this model sitting behind a tool interface that an agent can call to, you know, when it needs to make a prediction about the data.
57:01Jure Leskovec:Exactly, exactly. right and even if you think about let's say a customer support agent you call me in i need to estimate what's your lifetime value how likely are you to churn i will respond differently what's the best offer for me to give you how i need to actually ask a counterfactual question if i make you this offer how will that make you happier and and and these are all predictive problems i cannot just hallucinate them or you know ask chat gpt it will do something reasonable right like something something I would say common-sensy, but that's far away from optimal. So these predictions, this reasoning over this structured relational data that captures the patterns, behavior of, let's say, customers inside the organization is crucial to make accurate decisions, right?
57:49Jure Leskovec:And as we are deploying these agents, we cannot be building now separate models for each of these and pre-anticipate the questions, right? the beauty of the foundation model is that you can ask any question and one thing i want to say here it's like just to show how big the problem is right like if you think about a organization like maybe like uh you know think of let's say sap right uh sap has i think i know 70 100 000 customers each of their customers has structured data because it's a it's an organization every one of them um uh changes the schema a bit so everyone has their own data has their own schema and every one of those wants to do a churn prediction, wants to do churn prediction.
58:33Jure Leskovec:But every one of them has a bit different definition of what churn is. So now can you hire 70 ,000 data scientists that are going to build per client churn model with the client's data and the client specification what the churn means? You can't. A foundation model, an agent can just ask, predict me probability of churn under this definition under under this data and you get the answer half a second later that raises a question for me it around uh like post-training the model or fine-tuning the model does it is there any value to um like intermediate it fine tuning. I think what I'm envisioning is like you partner with SAP, SAP has, you know, hundreds of these modules.
59:27There's like a supply chain module and there's a, you know, churn module and some other thing. Like, does it make sense to tune on, you know, the use case, separate from the individual customer's data or does the foundation, the breadth of the foundation model already capture all of the information at that use case level of abstraction and really you're only improving it if you're looking at a specific customer's data?
1:00:01Jure Leskovec:That's a great question. Actually, that's something we are deeply looking into right now. I would say there's several reasons why to post-train. One reason you would want to post-train, even in the in-context learning scenario, is to better learn the distribution of the underlying data, to better understand the distribution of the underlying data, so that prediction, like the data gets better encoded, and prediction later will be more accurate. So that's one reason you would want to, let's say, train, even in a task-agnostic way, over the underlying data to better capture distributions, to better learn priors.
1:00:41Jure Leskovec:another reason why you would want to fine-tune is for cost cost reasons because if i fine-tune for a specific task i don't need to do icl right because now i don't now my context is is much smaller i don't need to bring in the label data i just bring in the entity i want to predict on now the the attention is smaller it's faster it's uh it's uh cheaper cheaper to run if i have large amounts of data, the model can learn a lot from it. So I would say there is a spectrum, there is a continuum of what you can do and the benefits of it is kind of different depending on where on this continuum are you seeing.
1:01:25So what's next?
1:01:27Jure Leskovec:Yeah, what's next? We are very excited about agents, both basically surfacing these two agents as tools The second thing is right now, right, like the coding agents are out there. But what we see is that coding agents require a proper abstraction and a proper infrastructure to be able to be effective. Right. And for example, if you could say, hey, why don't I just, you know, give this modeling task to CloudCode and CloudCode will build the model for me. So, you know, what's the big deal? And when we do that, what we see, we've run this internally, right, is that these models write thousands of lines of code, but there are these like super subtle data science mistakes.
1:02:16Jure Leskovec:mistakes so for example we've done this uh together uh together uh with expedia um and you know when when it was a account level fraud and mistakes for example the agent make was that when it created features for that given account it created it aggregated the transactions till midnight not till the current time right so it said oh today is i don't know uh april 30th so we'll use the data up to midnight of april 30th not actually saying hey it's actually 10 a.m 10 10 a.m on april 30th we can only use data up to here right so that's information leakage was a little mistake in there another mistake it made was that you know we did it at a transaction level instead of the account level and these are like these subtle mistakes that really you need the human in there.
1:03:11Jure Leskovec:But if you give it this more like higher level Kumo-like API, then it's able to do the same work in about 50 lines of code. No mistakes. The task in this case is to do what? Like I thought the task that you were describing was to code up something like what Kumo is trying to do the task is build me a account level fraud fraud detection model over this data and so what you're what you're proposing is like as opposed to trying to the agent trying to code it up from scratch you create some kind of skill or something like that that teaches it how to use kumo to get the same information yeah or what i'm saying is agents you know they they can go they can autonomously maybe make two steps but not hundred steps so now when i ask it for a task i can say hey here's pytorch go build me the model that that's you know takes thousand lines of code to build a model with python i could say here's hgboost build me a model that takes about engineer features and so on that's that takes about 500 lines of code and and you know uh or i can say using the kuma api go build me the model that only takes 50 lines of code and the the the now if you think of this analogy of steps you know 50 lines of code is maybe like two steps 500 uh lines uh lines of code is uh is 20 steps right and a lot you can get quite lost navigating i don't know 20 steps in this right i think the observation is is more general i'm like the observation is that for agents to be effective they need apis that are agent or agentic friendly okay awesome well yuri thank you so much for jumping on and catching us up to what you and Kuma are up to.
1:05:26Super cool stuff.
1:05:27Jure Leskovec:Yeah, thank you so much for the conversation and very insightful questions. Awesome. Thanks so much.
1:05:50Thank you.
From the publisher
In this episode, Jure Leskovec, co-founder and chief scientist at Kumo and professor of computer science at Stanford, joins us to explore two fronts of his work: AI for science and relational deep learning. We begin with AI Virtual Cell, a multiscale effort to learn data-driven representations from proteins to cells to patients using single-cell RNA-seq data, protein language models like ESM, and structure models like AlphaFold—without hand-encoding biology. Jure then dives into relational deep learning, reframing enterprise databases as graphs and training neural networks directly on raw multi-table data. He explains Kumo’s Relational Foundation Model (RFM2), which performs in-context learning over subgraphs to make accurate predictions on new databases and tasks with no training, and how this approach benchmarks against RelBench and other multi-table datasets. We also discuss real-world deployments at companies like Reddit, DoorDash, and Coinbase, explainability via attention over tables and columns, integration with agentic systems, deployment options, and practical limitations.
The complete show notes for this episode can be found at https://twimlai.com/go/768.




