AI revolution finally comes to Relational foundational models for structured data

13 Dec 2025 · 15 min · 10 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Relational foundation models (RFMs) for structured enterprise data—treating databases as graphs so a frozen, pre-trained transformer can make fast, forward-looking predictions without manual feature engineering or task-specific training.

Guest backgrounds

No guests are named in the transcript; it’s a host-led “Deep Dive” discussion with referenced “experts” and “one source.”

Key claims

LLMs fail on structured-data inference because they’re next-token predictors, not cause/effect forecasters. RFMs use relational graph transformers over primary/foreign-key links, enabling ~200ms predictions via a single forward pass. Setup takes days; labels are generated with “time-traveling” windows to avoid temporal leakage. Reported accuracy gains: often 10–20% better than traditional tools.

Notable examples

DoorDash recommendations (+30% vs internal model); Reddit ad-click prediction (equivalent of 4–5 years of improvement in months); dating app recommendations improved +15% by adding image embeddings.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

The Challenge of Structured Data

0:45 to 1:20

Exploration of the limitations of AI in processing structured business data.

“And there's a really fundamental reason why.”

The Inefficiencies of Current Workflows

1:20 to 3:08

Discussion on the lengthy and costly processes of data modeling and feature engineering.

“For anyone listening right now, the business leader or data scientist, what's the immediate just agonizing pain point of sticking with that approach?”

The Rise of Relational Foundation Models

3:08 to 4:35

Introduction to relational foundation models and their innovative approach to data.

“So it was just running really fast in the wrong direction.”

Speed and Efficiency of RFMs

4:35 to 6:26

How RFMs provide rapid predictions and eliminate traditional training time.

“Your users are nodes, your products are nodes, your transactions are nodes.”

The Role of Human Input

6:26 to 8:06

Emphasis on the necessity for human involvement in setting up relational models.

“That is precisely what the evidence shows.”

Automating Label Generation

8:06 to 10:00

Explanation of the automated system for generating clean training sets.

“It's an automated way to build a clean training set.”

Real-World Applications and Success Stories

10:00 to 11:28

Examples of companies successfully using RFMs for improved predictions.

“There's an example from a large dating app.”

Implications for Data Scientists

11:28 to 13:06

Discussion on how RFMs impact the role of data scientists moving forward.

“If they're meant to make autonomous decisions, where does an RFM fit in?”

The Future of RFMs and Enterprise Data

13:06 to 14:00

Exploration of the transformative potential of RFMs across industries.

“To finish, let's go back to the core mystery.”

Exploring Advantages in Market Applications

14:00 to 14:33

Discover how relational foundational models can transform quant trading and sports betting.

“I mean, what kind of massive, maybe unfair advantage could this give you in a market like quant trading?”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Jure Lescovec:Welcome back to the Deep Dive. Our mission today is to cut through the buzzwords, to really get into the breakthrough that's happening in, well, the most overlooked part of the enterprise. Yeah. Your structured business data. I mean, for the last decade, we've seen AI just conquer text, right? It's revolutionized images, it's generating code. But that critical data, the ground truth of your business, your customer records, transactions, all of it stored in these, you know, structured data warehouses, that's been stubbornly stuck. It has. It feels like it's relying on tech from the early 90s. We're here to figure out what finally changes that.

0:35And it hasn't been from a lack of trying. People have tried and they've spent millions trying to force structured data into large language models basically by turning tables into these giant long strings of text.

0:47Jure Lescovec:And that didn't just fail. It failed spectacularly. It did. And there's a really fundamental reason why. LLMs at their core are built for next token prediction. They're asking what's the most probable next word? But business forecasting, that requires extrapolation. It's about inference, about cause and effect. If you're predicting customer churn, you need to know why something is happening, not just what word comes next. That structural understanding is missing. And for high stakes business predictions, that's a fatal flaw. Okay, let's unpack this. So if foundation models couldn't crack it, that means the status quo is still that old manual workflow.

1:24Jure Lescovec:For anyone listening right now, the business leader or data scientist, what's the immediate just agonizing pain point of sticking with that approach? Oh, the pain point is it's monumental cost, glacial speed and incredible uncertainty. Just think about the workflow. Yeah. You hire a team of very expensive data scientists. They spend weeks, sometimes months, doing manual feature engineering, literally handcrafting metrics like count the transactions in the last three weeks or, you know, average spend by this category. Then they have to build a training data set, train a very specific model for that one task, and then try to get it into production.

2:01The source material here is, frankly, shocking. Developing one single model for one task. It often takes a solid six months of effort.

2:10Jure Lescovec:Six months for a single model. That's an eternity in a fast-moving business. It is. And the expense doesn't stop. Once you deploy it, maintaining that one custom model takes, on average, two full-time employees. Just to keep it running. just to keep the features fresh and manage model decay. And here's the kicker, the thing that just destroys ROI. More than half of all custom models developed never even make it to production. They just die in development. Wow. It sounds less like a high-tech team and more like a custom factory that just throws out half its inventory. We did have automation attempts before this, though.

2:43Jure Lescovec:AutoML, for instance. Why didn't that solve the problem? AutoML, it promised a lot, but it was just too unintelligent. One source called it too brute force. It didn't understand the data structure. It was just running these massive for loops, you know, testing every possible hyperparameter, running random SQL queries and giant table joins, just hoping it would stumble onto something that worked. It was all optimization, but with zero intelligence. So it was just running really fast in the wrong direction. Now, before we get to the core breakthrough, there were some early signs of life in a slightly easier area.

3:17Jure Lescovec:Time series foundation models, TSFMs. They were kind of a signpost, weren't they? Absolutely. TSFMs were a necessary step. They took the transformer architecture and applied it to time series data like, say, stock prices or sensor readings, treating it like a sequence of tokens. And that worked because time data is inherently sequential. But that only solves for single ordered lines of data. The real challenge is handling an entire database with all its messy interlinked tables. Which brings us to today. Here's where it gets really interesting. We're moving beyond those linear sequences to the real revolution for enterprise data, the relational foundation model, or RFM.

3:55The RFM is the core innovation. And crucially, when we say foundation model, we mean a true pre-trained frozen model. It has never seen your company's specific customer list or your transaction data. You connect this generalist model to your existing data warehouse, your collection of maybe 5, 10, 20 tables all linked up, and it is immediately ready to make predictions.

4:19Jure Lescovec:That concept, connecting a general model to my super specific proprietary data and expecting instant results, that's where I get a little skeptical. What's the fundamental tech shift that lets it reason over all these arbitrary connections? The shift is architectural. It's a conceptual leap. The key insight is realizing that any database, any set of tables linked by those primary and foreign keys is fundamentally a graph. Exactly. Your users are nodes, your products are nodes, your transactions are nodes. And those keys, they're the edges connecting everything. This isn't your grandfather's data model.

4:51This is the next generation of graph neural networks, what they call relational graph transformers. They're trained on structure, not just sequence.

5:00Jure Lescovec:That graph analogy is crucial. It reminds me of computer vision. Before deep learning, if you wanted to detect, say, defects on a conveyor belt, you hired engineers to manually code features like sharp corners or texture. It was all manual. Yes, exactly. Now the model just learns from raw pixels. That's the perfect analogy. RFMs do for data science what deep learning did for vision. Instead of a human writing complex SQL joins to define a feature, the relational graph transformer just attends directly over the raw tables and events. It learns the features automatically by following the paths in that graph.

5:33The whole discipline of manual feature engineering as well is essentially eliminated.

5:38Jure Lescovec:Okay, let's challenge the claims on speed. You're saying it attends over all this complexity instantly. How fast are we talking? We're talking about predictions in about 200 milliseconds. Seriously? Near real time. Even for a complex query like fraud risk on a new transaction that needs to look up a customer's five-year history across 20 different tables. And that speed holds up because the model is frozen, right? It's not training when I ask it a question. Exactly. It's not training at query time. It gets the task. It generates some in-context examples for itself, and then it just runs a single forward pass.

6:09There's no gradient descent, no iterative training. So months of human labor and machine training are just gone. You get the result almost instantly.

6:19Jure Lescovec:So if I'm the data scientist, the productivity game is just immense. You're saying a zero-shot prediction with no fine-tuning on my company's data is as accurate as a custom model that I, a PhD, spent a full month's building by hand. That is precisely what the evidence shows. And in many cases, it's not just parity. We're talking about superhuman performance. The attention mechanism can find these subtle signals across huge relational paths that human-designed features just miss. The claim is that these RFMs are often 10 to 20 % more accurate than traditional tools. So what does the workflow look like?

6:54Jure Lescovec:If feature engineering is gone, I can't be using standard SQL because SQL is backward-looking. You're right. It uses a domain-specific language, a DSL, that looks a lot like SQL, but it starts with the command predict. So you're not asking, show me all users who bought X last month. you're asking, predict 60-day churn for this specific user. It's a fundamental shift from historical reporting to forward-looking inference. But the human still has a critical role here. It's not just plug and play, right? Absolutely not. The system needs that human input for the initial semantic model. An expert has to define which tables to use, what the key relationships are, and most importantly, what the task means.

7:33You have to tell the RFM what churn actually means at your company. Is it zero purchases, zero logins? That setup takes a couple of days just to make sure the system speaks your specific business language.

7:45Jure Lescovec:And once that's set, you mentioned fine-tuning for high-value tasks. That sounds like a heavy lift, especially generating all the labels from historical data. That's where the platform is really clever. Most business tasks are time-dependent, right, like churn or recommendations. So the platform automates the label generation using what they call a time-traveling mechanism. Okay, what exactly is a time-traveling mechanism? You can't just leave that hanging. Right. It's an automated way to build a clean training set. The system slides a window across your historical data. It uses the past to make a prediction, and then it looks into the future window to see what actually happened.

8:20That's the ground truth label. This completely prevents what's called temporal leakage, so the model never accidentally trains on data from the future. The data scientist just describes the task, and the platform handles all the complexity of generating thousands of clean training examples.

8:36Jure Lescovec:That is a huge reduction in grunt work. Let's get to the real-world evidence. What are the aha moments that prove this isn't just theory? Oh, we have some blockbuster examples from companies whose entire business model relies on this kind of modeling. Take DoorDash. Okay. They use the RFM to tackle their core problem, recommending a new restaurant a user has never ordered from. This is a problem they have optimized for years. The RFM delivered a 30 % improvement over their existing flagship internal model. 30 % over a mature, heavily optimized recommendation engine. How is that even possible? It's because the RFM finds correlations that manual feature engineering just misses.

9:16The manual model might rely on an aggregated feature, like average distance to pass restaurants. The RFM might discover something way more nuanced, like users who ordered a specific vegetarian dish on a Tuesday night and also used this one in-app feature three weeks ago are 30 % more likely to try this new French place. It finds those subtle multi-hop pathways in the graph that no human would ever think to art code.

9:43Jure Lescovec:What about in the ad space? Reddit used this for their advertising models, predicting ad clicks, which is their direct revenue driver. Now, typical improvement in ad-click prediction is maybe 1 % to 2 % a year. Using the RFM, they got the equivalent of four or five years of internal model improvement in just a couple of months. That's a massive acceleration. A massive acceleration. And it's multimodal, too. It's not just numbers. There's an example from a large dating app. They enhanced their recommendation system by simply adding an image column, you know, using computer vision embeddings. Just by adding that image data to the relational graph, their prediction accuracy immediately jumped by 15%.

10:19Jure Lescovec:Okay, before we get too carried away, let's talk about reality. Garbage in, garbage out is a real problem. What happens when the data is just messy? Missing values, inconsistent entries like CA versus Califf. And that's a great point. While no system is totally immune to bad data, the RFM is way more resilient to that kind of noise. Because it's reasoning over the graph's connections, it can sort of borrow information from nearby data points. If one user is missing a postal code, the model can infer a probability based on the postal codes of their nearest neighbors or the transactions they share.

10:55It learns to reason through the mess. And to build trust, transparency is key. Because transformers attend over raw data paths, you can run the system backward to get data-rooted explanations. It will tell you exactly which row, which column, which specific event contributed to its prediction.

11:10Jure Lescovec:So you get a literal audit trail back to the source data. Precisely. And then those structured factual explanations like prediction based on these 17 transactions are fed into an LLM. The LLM then synthesizes that proof into readable, non-hallucinated text for the business user. It's transparent and it's trustworthy. Okay, if we connect this to the bigger picture, AI agents are supposedly the future. If they're meant to make autonomous decisions, where does an RFM fit in? An agent is only as good as the data it acts on. It needs decision power that's rooted in up to the second data and really fast predictions, not just static text.

11:46The RFM basically acts as the high-speed structural data brain for the agent. Think about the insurance industry. You have an autonomous agent trying to proactively retain customers. It needs to estimate the churn risk for thousands of people and then instantly recommend the single best offer, discount, and upsell, whatever, to send them to stop that churn.

12:06Jure Lescovec:And that used to be two huge separate data science projects, building the risk model, then running tests on the offers. Exactly. Months of work. Now, the agent can query the RFM to handle both the risk scoring and the personalized offer in a single 200 millisecond pass. The agent then immediately moves to the high-value step ready and sending the communication bypassing all that manual data science work. This sounds like a technology that should deeply worry data scientists, given how much it automates. How does this change the career path? The consensus from the sources is pretty clear. It's not about displacement.

12:39It's about enablement, massive enablement. It takes the data scientist away from all the drudgery, the feature engineering, the hyperparameter tuning, maintaining pipelines. It basically turns a data scientist from a high-paid feature engineer into a rock-star business strategist. Their new job is to focus on that semantic model, defining the business problem, figuring out what to predict, analyzing the impact. They stop worrying about Python libraries and start driving business value.

13:06Jure Lescovec:That's a powerful way to put it. To finish, let's go back to the core mystery. Why does connecting a general model to a unique corporate database work so well? What's the final intuition here? The experts point to two things. First, we know the transformer architecture is just incredible at generalization. But second, there's this belief that the world of enterprise data is actually less complicated than we think. While DoorDash's data looks different from Reddit's, the underlying relational patterns, the structure of a user, a transaction, they might just collapse into a smaller, more universal set of patterns that this architecture is perfectly suited to learn.

13:40Jure Lescovec:So what does this all mean? The technology is here to transform flagship models in every industry, retail, fraud detection, risk management. It's about unlocking the real value that's been trapped in enterprise data for decades. And given the RFM's ability to analyze this incredibly complex relational data and find signal that exceeds human intuition, it opens up some profound thought experiments. I mean, what kind of massive, maybe unfair advantage could this give you in a market like quant trading? Where analyzing interrelated economic factors is the whole game. Or even sophisticated sports betting, where you're looking at deep relational data on players, teams, climate.

14:21Jure Lescovec:If the model can see the entire graph and find connections that humans and traditional models just can't, the possibilities go way beyond typical enterprise use cases. Of course, we have to stress, these are just thought experiments, not investment advice. That is all the time we have for this deep dive. Thank you for joining us as we explore the relational foundation model. We'll see you next time.

From the publisher

We discuss an interview with Jure Lescovec, co-founder of kumu.ai and a computer science professor at Stanford, regarding the application of foundation models to structured enterprise data. Lescovec explains that traditional **machine learning** methods for this type of data are manual, expensive, and time-consuming, contrasting them with new relational foundation models that leverage a **graph-based approach** to eliminate the need for manual **feature engineering** and **model training**. The technology, which is a next-generation form of **graph neural networks**, is designed to provide rapid, accurate predictions for tasks like churn prediction, forecasting, and recommendation systems by connecting directly to databases and representing them as graphs for **attention mechanism** processing. The discussion emphasizes that the goal is not to displace data scientists but to enhance their productivity by providing a powerful tool capable of achieving **superhuman accuracy** with proper fine-tuning, as demonstrated through successful use cases at companies like DoorDash and Reddit.

More from Best AI papers explained

All 475 episodes
AI revolution finally comes to Relational foundational models for structured dataBest AI papers explained · 15 min
Listen in VO