Impression Share Prediction: An Offline Evaluation Task for Ranking Systems

25 Aug 2026 · 22 min · 16 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

How Meta researchers improve offline evaluation of ranking/recommendation algorithms by predicting “impression share shift” across objective buckets (clicks, video views, conversions) before deployment, using structural causal modeling to handle the counterfactual problem and the fragile first-hour dynamics of live auction systems.

Guest backgrounds

Two podcast hosts (no guest identities given in transcript). They discuss platform engineering and causal inference concepts (offline evaluation, pacing controllers, SCM, transformers).

Key claims

Standard offline accuracy metrics (e.g., NE/RIG/ERC) can look great while still causing harmful redistribution in live auctions; even a ~1% shift away from high-value conversions can materially degrade total system value. A causal graph that omits delivery-capacity-to-model-state backdoor paths enables observational-data identification.

Notable examples

Basketball free-throw analogy; social/shopping feed changing overnight; first 0–1 hour “ghost of prior equilibrium” causing a random-forest snapshot predictor to perform >20% worse on unseen models, while an encoder-conditioned patch TST transformer (two-hour “dash cam” capacity history) improves accuracy by 22.1% and catches up after ~48 hours.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

The Stakes of Ranking Algorithms

0:53 to 1:44

Exploring the challenges of deploying new ranking algorithms.

“And today, our mission for this deep dive is to explore exactly how tech giants are trying to figure out if a new algorithm is going to completely break that ecosystem before they ever even unleash it on you.”

The Flaw in Offline Evaluations

1:44 to 2:52

Examining why traditional offline evaluations of algorithms fail.

“We're looking at some truly groundbreaking data from a team at Meta, and they are tackling a massive challenge called impression share prediction.”

Basketball Analogy for Algorithms

2:52 to 3:51

Using a basketball player analogy to highlight evaluation flaws.

“It's a stand-in metric for what the engineers actually care about, which is downstream utility.”

Understanding Auction Dynamics

3:51 to 5:06

Delving into the competitive auction system in ranking algorithms.

“It doesn't tell you if they refuse to pass the ball or, you know, how the defense reacts to them or if their mere presence on the court exhausts the other players.”

The Fragility of Algorithmic Ecosystems

5:06 to 6:23

Discussing the delicate balance of impression distribution.

“A pacing controller's job is basically to make sure a campaign doesn't blow its entire daily budget in the first five minutes of the morning, right?”

Meta's Research Goals

6:23 to 7:20

Exploring Meta's pursuit of predicting impression share shifts.

“So a one percent shift can essentially tank the overall value the platform is generating.”

Causal Modeling in Prediction

7:20 to 8:13

Understanding how structural causal models aid in predictions.

“Aren't we just gazing into a crystal ball at a reality that hasn't happened yet?”

The Causal Graph's Structure

8:13 to 10:00

Explaining the importance of the causal graph in evaluation.

“Let's map this out for the listener, because the mechanics here are fascinating.”

Eliminating Backdoor Paths

10:00 to 11:13

Highlighting why backdoor paths complicate causal inference.

“What's fascinating here is a specific intentional omission in their causal graph.”

Building the Predictive Framework

11:13 to 12:10

Discussing the monumental scale of Meta's predictive framework.

“But because a model's brain is fixed independently of the minute-by-minute capacity, the researchers proved mathematically that you can identify true causal effects using just regular, everyday observational data.”
Show all 16 chapters

Comparing Predictive Tools

12:10 to 14:01

Contrasting two machine learning methods for predictions.

“What exactly are they handing over to the predictor to make this forecast?”

Analyzing Transformer Capabilities

14:01 to 15:18

Learn how the encoder-conditioned patch TST transformer processes data.

“The second contender, however, was entirely different.”

Performance Comparison of Models

15:18 to 18:22

Discover the contrasting performances of different predictive models.

“Well, the researchers evaluated them in two distinct scenarios.”

Understanding the Zero-to-One Hour Phenomenon

18:22 to 19:30

Explore the dynamics of new models entering systems and their impacts.

“In that same chaotic zero-to-one-hour window, the transformer recovers the gap perfectly, achieving a 22.1 % improvement in accuracy.”

Evaluating Offline Scorecards and Predictions

19:30 to 20:23

Examine how traditional offline evaluations fall short in real-time scenarios.

“That is exactly why your feed might behave strangely before finding its groove.”

The Future of Predictive Models

20:23 to 21:41

Discuss the potential advancements in predictive AI models and their implications.

“They can catch that 1 % shift away from purchases before it costs the platform millions of dollars.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00What if I told you that an algorithm perfectly predicting what you want to see could actually be the exact same algorithm that costs a tech platform millions of dollars in a matter of hours. Yeah. I mean, it sounds completely backward. Right. Because we usually think that, you know, if an AI correctly guesses what we want to click on, it's a success. Exactly. You'd assume it's a huge win for the platform. But the reality of how your attention gets distributed across a live feed is, well, it's far more volatile than that. I mean, have you ever noticed how a social media or a shopping app suddenly changes what it shows you overnight?

0:37Oh, all the time. Yeah, like the feed acts a little erratic, prioritizing completely different types of content before finally settling into a new groove maybe a day or two later. And you're actually witnessing a massive economic ecosystem recalibrating in real time when that happens. It's wild. And today, our mission for this deep dive is to explore exactly how tech giants are trying to figure out if a new algorithm is going to completely break that ecosystem before they ever even unleash it on you. It is arguably one of the most high stakes puzzles in the tech industry right now. No pressure.

1:11Right. Right. I mean, when companies roll out a new ranking algorithm, they're absolutely terrified of those unintended ripple effects. Because they could lose millions. Exactly. So they obviously evaluate new AI models before flipping the switch on live traffic. But the fascinating part is discovering that the way the industry has historically evaluated these algorithms is just fundamentally flawed. Flawed how? Well, especially when it comes to the complex reality of a live breathing auction system. And that brings us to the research we are unpacking today. We're looking at some truly groundbreaking data from a team at Meta, and they are tackling a massive challenge called impression share prediction.

1:53Yeah, and to really grasp why this new research matters, we first have to understand like the blind spot in the old way of doing things. It centers around this concept of offline evaluation. Okay, let's unpack this. Normally when engineers want to know if a new candidate algorithm is any good, they look backward. They do. They take a massive data set of historical logs. Basically a record of what users did yesterday. Yeah, exactly. And they run the new algorithm against that old data to see if it would have correctly predicted the user's behavior. Okay. And they use a whole suite of mathematical scorecards for this, right?

2:26Yeah. Like things you might see abbreviated in the data as NE, RIG, or ERC. Right, which sounds complicated, but for our purposes, we can really just think of them as historical accuracy scores. They are essentially just asking, did this model correctly guess the outcome? Yes. And we intuitively assume that a highly accurate model is, you know, a good model to put into production. But predictive accuracy is just a surrogate, isn't it? It is. It's a stand-in metric for what the engineers actually care about, which is downstream utility. Downstream utility. Yeah, which is basically the total value or the final aggregate outcome generated by the entire system over time.

3:04Okay. And a model that is perfectly accurate in isolation might behave like a bull in a china shop when you drop it into a live environment. That makes a lot of sense. You know, while reading through the findings, the best way I could wrap my head around this blind spot was by, well, comparing it to evaluating a basketball player. Oh, I like that. How so? Imagine you're scouting a new player, and the only metric you use to decide if they start in tonight's game is how many free throws they can make in an empty gym. Yeah, and if they hit 99 out of 100, the traditional offline metrics would say, well, this is our star player, put them in.

3:42Right. But that free throw percentage tells you absolutely nothing about how that player alters the passing dynamics of the team during a live, high-stakes game. Nothing at all. It doesn't tell you if they refuse to pass the ball or, you know, how the defense reacts to them or if their mere presence on the court exhausts the other players. Taking that analogy further, in the world of ranking algorithms, that shared court is a highly competitive multi-agent auction system. Okay, so it's crowded. Very. Every time you refresh your feed, multiple models are operating simultaneously, bidding for the chance to show you something.

4:17Wow, just on a single refresh. Yeah. And the impressions, the actual videos, ads, or posts you end up seeing are organized into what the researchers call objective buckets. Objective buckets, meaning the algorithm's specific underlying goal for that particular piece of content. Exactly. So one objective bucket might be optimized strictly for generating clicks. Okay. Another bucket might be optimized for sustained video views. Like getting someone to watch for more than three seconds. Right. And a third might be a really high value bucket trying to drive actual conversions like getting someone to buy a product or sign up for a newsletter.

4:55So these models are all feeding into a shared auction system. They are. And that system is governed by pacing controllers and a strictly finite amount of delivery capacity. Wait, I want to pause on pacing controllers for a second because that feels crucial. A pacing controller's job is basically to make sure a campaign doesn't blow its entire daily budget in the first five minutes of the morning, right? That's spot on. It meters out the delivery capacity over the course of the day. So it acts like a valve. Yes, exactly like a valve, constantly opening and closing based on how much capacity is left and how fast it's being consumed.

5:28Okay. All these models with their different objective buckets are drawing from the exact same pool of resources through those valves. And this is where that empty gym free throw metric completely falls apart. Completely. Because a new algorithm might be incredibly accurate at predicting clicks in the historical data. But when you put it in the live auction, it might be so aggressively confident about those clicks that it triggers the pacing controllers to throttle other campaigns. Yes. Radically redistributing where the overall impressions go. Wow. And what's deeply unsettling about this for platform engineers is the sheer fragility of that ecosystem.

6:05How fragile are we talking? Well, the research notes that even a microscopic one percentage point shift of impressions away from a high priority objective bucket. Like those high value conversions. Right. Away from conversions and into a lower priority bucket like near clicks. That tiny shift can materially degrade the aggregate outcomes of the entire system. So a one percent shift can essentially tank the overall value the platform is generating. Yes. That feels incredibly fragile. I mean, it means a new model could score an A-plus on all those standard offline accuracy tests. But the moment it goes live, it accidentally diverts 1 % of the traffic away from real purchases and toward just empty clicks.

6:44And the company is suddenly losing massive amounts of value without really knowing why. And standard offline evaluation gives you zero visibility into this dangerous redistribution. Zero. You are completely blind to the impression share shift until you put the model into a live A-B test on real users. And by the time you realize what's happening, the financial or experiential damage is already done. Exactly. Which brings us to the core mission of the meta team's research. They want to build a system that predicts that impression share shift offline before the new model ever touches the live auction.

7:18That's the holy grail, yeah. But I'm stuck on something here. If this new candidate model has literally never interacted with the live system, if it has never bumped up against a pacing controller and its confidence hasn't been shaped by real world friction, how do we mathematically predict its footprint? It sounds impossible. Yeah. Aren't we just gazing into a crystal ball at a reality that hasn't happened yet? You are touching on what statisticians call the counterfactual conundrum. The counterfactual conundrum. Yes, you are trying to predict an alternative reality. Okay, so how did they solve this without just blindly guessing?

7:56To solve this, the researchers utilized a structural causal model, or SCM. A causal model, okay. Right. Instead of just looking at basic correlations, they mathematically mapped out the entire auction ecosystem using a causal graph. We're drawing a map of how everything connects. Exactly, tracking how variables influence each other from one minute to the very next. Let's map this out for the listener, because the mechanics here are fascinating. What are the actual variables driving this causal graph? So there are three main variables moving through time that we really need to track. Okay, what's the first one?

8:27First, you have the model state. You can think of this as the intrinsic brain of the algorithm. Brain. Right. It is the distribution of prediction scores. Essentially how confident the model is about the various items it wants to show. Okay, got it. And the second? Second, you have the impression share. This is the final outcome. Meaning who actually wins. Yes. It is the actual fraction of impressions that end up landing in each of those objective buckets, like clicks or views. Okay, so the model state is the brain bidding in the auction. And the impression share is who actually wins the auction.

9:04It gets shown to the user. What's the third variable? The delivery capacity state. Right. This is that shared pool of resources constrained by those pace and controller valves we talked about earlier. So we have the model state, the impression share, and the delivery capacity. How do they push and pull on each other in this ecosystem? It basically operates as a continuous tangled feedback loop. The loop. Right. The model state influences who wins the auction, which directly dictates the impression share. Meanwhile, the delivery capacity constrains who is even eligible to bid, which also impacts the impression share.

9:37This sounds chaotic. It is. And finally, the impressions that actually get served deplete the shared capacity, which means the delivery capacity state carries forward and changes the conditions for the very next minute. It's an endless cycle. The model drains the capacity, which changes the capacity, which alters the auction for the model on the next refresh. It is, except for one crucial detail. What's fascinating here is a specific intentional omission in their causal graph. An omission. What did they leave out? When you look at how they map the math, there is absolutely no causal link pointing from the system's delivery capacity backward to the model's fundamental state.

10:16Meaning the current weather of the auction, like weather capacity is high or low, doesn't reach back in time and rewrite the underlying architecture of the algorithm. Exactly. A model's intrinsic properties, its neural network architecture, its foundational training, those are fixed before it enters the auction. It's a baked cake. Yes, a baked cake. The system's capacity state does not change the model's fundamental code. And because that connection is missing, it eliminates what statisticians call backdoor paths. Backdoor paths. Why does eliminating a backdoor path matter so much in this context?

10:53Well, in many causal inference problems, everything influences everything else. Right. When you have those tangled backdoor paths, you have to use incredibly complex mathematical acrobatics, things like inverse propensity weighting, which is essentially trying to mathematically unbake a cake when you don't even know the recipe just to figure out what caused what. That sounds like a nightmare. It is. But because a model's brain is fixed independently of the minute-by-minute capacity, the researchers proved mathematically that you can identify true causal effects using just regular, everyday observational data.

11:27Okay, so the theoretical causal math gives them a clean map to follow. But, I mean, having a map is one thing. Building a machine that can actually predict the future 24 hours in advance is another. How did they pull this off practically? The scale of the statistical learning framework they built is just monumental. How monumental. They gathered about five weeks of historical A-B test data. They looked at 150 different candidate models across multiple model families. Wow. And from that, they captured 1.8 million minor level snapshots of auction statistics to train their predictors. 1.8 million snapshots.

12:02So when they want to predict what a new model will do, they have to feed this predictor a specific set of inputs at time zero. Right. What exactly are they handing over to the predictor to make this forecast? At time zero so. The moment they hit evaluate, the predictor receives two distinct sets of features. First, it gets the model confidence feature. Confidence feature. Right. This is a 22-dimensional summary of the new algorithm's brain. 22 dimensions. Wow. Yeah, it includes a histogram of its prediction scores, the mean, the variance, and a vector detailing exactly which objective buckets the model is even allowed to target.

12:37It's essentially the psychological profile and the playbook of the new player stepping onto the court. Exactly. And the second set of features. Yeah, what's the second set? The capacity features. This is the snapshot of the court itself. Okay. It tracks the total capacity, how much has already been consumed, what's remaining, the current pacing multiplier, and the number of active campaigns currently fighting for space. So we have the state of the new model and the state of the current system. Right. And to figure out how to best predict the outcome, the researchers pitted two completely different machine learning tools against each other.

13:11They did. And the contrast between how these two tools approach the problem is striking. Oh, totally. How would you visualize the difference between the two? Well, looking at the data, the first tool was a random forest regressor. Right. I would compare the random forest to taking a single high-resolution photograph of a chaotic city intersection at rush hour. You look at that single still frame at time zero. You see exactly where all the cars are. You note their positions. And from that one photograph, you try to predict what the traffic flow is going to look like for the next 24 hours. That's a great way to put it.

13:45it relies entirely on the instantaneous snapshot. Yeah. It takes the model's confidence profile, overlays it onto the capacity features of that exact minute, and generates a prediction of the impression shares. Right. It has zero concept of momentum. None. The second contender, however, was entirely different. It was an encoder-conditioned patch TST transformer. Quite a mouthful. Yeah. But instead of taking a single photograph, this transformer is like setting up a dash cam and watching a two-hour video of that same intersection before you have to make your prediction. And not just watching the video, but analyzing it in a very specific way.

14:24How so? The transformer ingests two solid hours of minute-level capacity history leading up to time zero. Okay. It takes that history and breaks it down into 15-minute chunks or patches. Hence the patch part of the name. Exactly. It then processes those patches through a two-layer transformer network to extract the momentum of the system. Wait, why are the 15-minute patches so important? Why not just feed it the whole two hours as one big block? Because chunking it into patches allows the AI to calculate velocity and acceleration. Oh, interesting. Right. By comparing one 15-minute patch to the next, the transformer can see if capacity is draining faster than it was an hour ago, or if a pacing controller is starting to clamp down.

15:04So it understands the flow and trajectory of the arena, not just the static frozen state. Yes. Okay, so we have the static photograph versus the two-hour dash cam. How did these two tools actually perform when they were put to the test? Well, the researchers evaluated them in two distinct scenarios. First, on models the predictor had already seen during its training, and then on completely brand new unseen models. Let's start with the seen models. These are algorithms that have already spent some time interacting with the live system, right? Right. And in this scenario, the random forest, the single photograph approach, actually performed exceptionally well.

15:41Really? Yeah. It reduced the L1 prediction error by nearly 49 % compared to a constant baseline. And L1 error is essentially just a measure of how far off the prediction was from reality. Exactly. Lower error means it was highly accurate. So a 49 % reduction is a massive win for a simple snapshot. It is. The researchers wanted to know why it was so successful, so they performed an ablation study. An ablation study. That's essentially taking the prediction model apart piece by piece, removing variables one at a time to see which specific piece of data was doing the heavy lifting. And they found that the candidate model's own score distribution, its internal confidence, accounted for almost 43 percent of that accuracy gain.

16:28Meaning if a model is already integrated into the system, its own internal confidence signals are the best indicator of where traffic will go. Yes. The system capacity just adds a little bit of helpful context. Here's where it gets really interesting, though. Oh, yeah. Because this was only for models that were already integrated. The holy grail of this research is predicting the counterfactual. Right. The unknown. What happens when you test a held out model, a candidate algorithm that has never been seen before, stepping onto the live court for the very first time? This is where we see the true offline to online gap.

17:01When a completely unseen model enters the system, the most chaotic, critical window is the very first zero to one hour. The first hour phenomenon. Yes. And in that critical first hour, the single photograph random forest completely face plants. Ouch. Yeah. It actually performs over 20 % worse than if you had just used a basic static baseline. It performs worse than doing nothing at all. Yes. How does it fail so spectacularly when it was so accurate just a minute ago on the scene models? Because of the snapshot. Okay. When the random forest takes its picture at time zero, the system's capacity state is still deeply polluted by the ghost of the prior model's equilibrium.

17:42The ghost of the prior model. Right. Yeah. The ecosystem hasn't realized there's a new player on the court yet. So the random forest looks at that polluted snapshot and makes wildly inaccurate assumptions about how the new model will behave based on the old model's traffic patterns. It's assuming the cars will drive exactly the way they did yesterday, completely blind to the fact that there's a massive new variable on the road. Precisely. But the transformer, the dash cam approach, handles this friction beautifully. Because of the patches. Because it ingested that two-hour rollout of recent auction dynamics.

18:15It understands the underlying momentum of the system. So it isn't fooled by the static ghost of the old equilibrium. Not at all. In that same chaotic zero-to-one-hour window, the transformer recovers the gap perfectly, achieving a 22.1 % improvement in accuracy. Wow. That is a swing of over 42 percentage points compared to the random forest. It's a huge difference. Just by giving the AI a two-hour memory of the system's velocity, it can accurately forecast how a completely unseen algorithm is going to aggressively reroute millions of impressions. And what's truly illuminating is observing what happens after that first hour.

18:52What happens? As time goes on, the auction system starts to absorb and adapt to the new model. By day two so, 48 hours later, the ecosystem has fully adapted. The ghost of the old equilibrium has been entirely fleshed out. Right. And at that point, the simpler random forest actually catches back up and becomes highly accurate again. It's fascinating to connect this back to our daily experience on the Internet. We started by talking about that jarring feeling when an app changes its algorithm overnight. Yeah, that weird erratic phase. This zero to one hour window, this chaotic adaptation phase where the old ecosystem is dying and the new one hasn't quite settled.

19:30That is exactly why your feed might behave strangely before finding its groove. Exactly. You are literally witnessing the auction dynamics recalibrating around you. You are living inside the counterfactual as it becomes reality. That is just incredible. So if we zoom out, what are the ultimate takeaways here? For years, tech companies have relied on these isolated offline scorecards to evaluate algorithms. Right. But those scorecards are entirely blind to impression shifts. They assume a highly accurate AI is a good AI, completely ignoring how its overconfidence might drain resources from high-value objectives in a live auction.

20:05But by utilizing a structural causal model, we now have a mathematical map. The map without the backdoor paths. Exactly. We know that a model's intrinsic architecture isn't rewritten by the system's capacity, which allows engineers to use observational data to map the future. And by feeding that data into an encoder-conditioned transformer, giving the predictor a memory of the recent dynamics rather than a static snapshot, they can predict exactly how a new algorithm will redirect traffic before it goes live. Which is huge. They can catch that 1 % shift away from purchases before it costs the platform millions of dollars.

20:42And the researchers did point out an incredibly compelling future direction for this work, which I think is worth mentioning. What's that? Right now, this transformer is performing direct prediction over a 24-hour window. But they suggest the ultimate goal is moving toward a fully sequential Bayesian model. A Bayesian model in this context, meaning what exactly? It means a system that propagates uncertainty forward across time, step by step. Instead of just predicting tomorrow, it continually simulates the cascading ripple effects of every tiny change, infinitely forward, constantly updating its own probabilities as it goes.

21:17Okay, that sounds like we are bordering on science fiction. It does, which leads to a thought I want to leave you with. Right. If the industry eventually perfects this, if an AI can flawlessly simulate multi-agent economic ecosystems, minute-by-minute auction dynamics, and human capacity constraints days or weeks in advance, will live A-B testing on actual users eventually become obsolete? Wow. Will the internet of the future, every subtle change in your feed, every new product recommendation be entirely shaped, fought over and perfected inside invisible offline simulations before you ever even open the app?

21:52A completely invisible parallel internet running constant simulations just to make absolutely sure the free throws will translate to the live game. That is a heavy thought to end on. Thank you for joining us on this deep dive. Next time you scroll and feel like the ground just shifted beneath your feet, remember, there is a highly constrained, incredibly fragile auction happening behind every single impression. Keep questioning what you see.

From the publisher

Researchers from Meta Platforms propose a novel offline evaluation task called impression share prediction to better anticipate how new ranking models redistribute traffic across different business objectives. Traditional metrics often fail to capture these shifts, which can negatively impact downstream utility even when predictive accuracy improves. To address this, the authors developed a structural causal model that identifies how model signals and delivery capacity interact to determine impression allocation. Their framework includes a Random Forest regressor for established models and a specialized encoder-conditioned architecture to handle the complex dynamics of newly introduced models. This system significantly reduces prediction error compared to standard baselines, particularly during the critical first hour of a model's deployment. Ultimately, this approach provides practitioners with vital visibility into a candidate model's allocation behavior before proceeding to expensive online A/B testing.

More from Best AI papers explained

All 475 episodes
Impression Share Prediction: An Offline Evaluation Task for Ranking SystemsBest AI papers explained · 22 min
Listen in VO