LLM Evaluation as Tensor Completion: Low-Rank Efficiency and Uncertainty Quantification

12 Apr 2026 · 19 min · 12 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

How to evaluate LLMs from pairwise “chatbot arena” votes using a low-rank tensor model, fixing statistical distortions from sparse/noisy, non-uniform sampling and unequal information across matchups; includes uncertainty quantification via score whitening and inverse probability weighting.

Guests

Jia Chun-Li, David Sinchi Levi, Real Wei Sun (researchers/authors of the discussed MIT/Purdue framework).

Key claims

Win-rate ELO-style leaderboards are statistically misleading because votes are binary outcomes from a Bradley-Terry-Luce latent-strength model with zero-sum normalization; informative comparisons occur near 50-50, while lopsided matches add little Fisher information yet dominate data. Low-rank latent performance tensors enable borrowing strength across tasks; score whitening equalizes Fisher information to avoid a non-commutativity bottleneck; IPW corrects for frontier-model sampling bias.

Notable examples

Grandmaster vs novice chess analogy; arena categories like “Code Technical” (~19k observations) vs “Code General” (~2.6k); improved coverage/narrower confidence intervals vs naive baselines.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

The Challenge of AI Evaluation

0:50 to 2:25

Discussion on the murky landscape of AI evaluation data and the transition from traditional diagnostics.

“So Jia Chun-Li, David Sinchi Levi, and Real Wei Sun, they realized that to fix this broken X-ray, we have to kind of completely change the math behind AI leaderboards.”

Understanding the LMSYS Chatbot Arena

2:25 to 4:12

Explaining the setup and significance of the LMSYS chatbot arena for AI comparison.

“And in this framework, the probability that model A beats model B depends entirely on the difference between their hidden latent scores.”

Statistical Illusions in AI Rankings

4:12 to 6:26

Investigating how treating human votes equally distorts AI capability rankings.

“In the Bradley-Terry-Luce model, the Fisher information tells us how much statistical weight a single observation carries.”

The Importance of Hidden Factors

6:26 to 8:00

Exploring the significance of hidden factors in AI model comparisons and their impact on results.

“For example, maybe one foundational factor maps to logical reasoning ability.”

Introducing the Low-Rank Tensor Model

8:00 to 9:23

How the new low-rank tensor model aims to address AI evaluation challenges.

“If you try to model every single unique quirk and blind spot independently, you run straight back into the sparsity problem.”

Addressing the Sampling Pattern Chaos

9:23 to 11:43

The imbalance in AI evaluation data due to user-driven sampling patterns is dissected.

“In basic math, A times B equals B times A.”

The Technical Roadblock in Ranking

11:43 to 13:44

Understanding the challenges faced in calculating uncertainty around AI rankings.

“Mechanically, how does whitening fix the loop?”

Score Whitening and Inverse Probability Weighting

13:44 to 14:00

How score whitening and inverse probability weighting help stabilize AI ranking.

“But whitening only fixes the math if we assume the data is there to be whitened.”

Understanding Inverse Probability Weighting

14:00 to 16:33

Learn how inverse probability weighting addresses sampling bias in AI evaluation.

“Even if you smooth out the fissure information, the sampling itself is still totally skewed towards a handful of models and tasks.”

The Importance of Confidence Intervals

16:34 to 16:59

Discover how confidence intervals impact the reliability of AI model scores.

“The gap between two scores might just be a statistical illusion caused by low fissure information and skewed sampling.”
Show all 12 chapters

Recap of AI Evaluation Techniques

17:00 to 17:44

Review key AI evaluation concepts including low-rank factors and score whitening.

“and why lopsided matchups give us almost no fissure information.”

Philosophical Implications of AI Evaluation

17:45 to 18:43

Reflect on what AI's statistical evaluation means for understanding human intelligence.

“Are your own complex professional and creative skills secretly just a low-rank tensor waiting to be mapped?”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00You know, usually when we talk about a diagnostic test, there's this expectation of precision. Like if you break your arm, the x-ray shows that jagged white line and the doctor just points and says, yep, there it is. Right. It's binary. Exactly. Broken or not broken. It's clean. But then you step into the world of evaluating artificial intelligence and suddenly that x-ray machine is, well, it's completely broken. Oh, completely. We are looking at a diagnostic landscape that is just incredibly murky because, you know, you want to know which AI is the best, maybe for your own coding projects or like writing tasks or you look at the data.

0:34But the raw data we use to figure that out is an absolute mess. It is the definition of diagnostic muddy waters. I mean, the data powering our understanding of these models is sparse, it's noisy, and it is overwhelmingly non-uniform. Which is exactly why our deep dive today focuses on a new framework from researchers at MIT and Purdue. So Jia Chun-Li, David Sinchi Levi, and Real Wei Sun, they realized that to fix this broken X-ray, we have to kind of completely change the math behind AI leaderboards. We really do. And if you're listening and you've ever wanted to know which AI is currently winning the arms race, you've probably checked out the LMSYS chatbot arena.

1:14Oh, yeah. Everyone uses it. Right. And the setup is pretty straightforward. You type a prompt, two anonymous AI assistants generate responses side by side, and you just vote for the best one. Yep. Assistant A, Assistant B, or a tie. Exactly. It's crowdsourced, pairwise benchmarking. And it has become the gold standard precisely because it evaluates these models with real users and real prompts rather than, you know, just feeding them some static standardized test. It feels very empirical. It does feel empirical, but relying on that raw human preference data to rank these models is actually statistically treacherous.

1:51Treating all those human votes as equal and just calculating a simple win rate, it creates massive statistical illusions. Huge illusions, yeah. So, okay, let's unpack this. Why exactly does treating every vote equally warp our understanding of an AI's capability? Well, the underlying statistical object of these leaderboards isn't actually a flat score, like, say, a grade on a math test. It's actually a collection of latent or, you know, hidden pairwise comparison strengths. Statisticians model this using something called the Bradley Terry Loose model. The BTL model, right. Right. And in this framework, the probability that model A beats model B depends entirely on the difference between their hidden latent scores.

2:33You don't get a direct measurement of an AI's absolute intelligence. You only get the binary outcome of that specific battle. Wait, so we never actually see the real power level of the AI? No, never. We just see who won the arm wrestling match on a given Tuesday. That is exactly it. And because you only ever see the relative difference between two competitors, you can't actually pinpoint their absolute scores mathematically, at least not without imposing a constraint. Constraint, like what? Researchers use a zero-sum normalization along the model dimension. They essentially force the math to assume that within a specific task category, all the model scores sum up to zero.

3:10Oh. Yeah. That anchors the math so you can actually calculate relative differences. But, you know, anchoring the math doesn't solve the absolute chaos of the sampling pattern. Right. Because, well, if you're listening to this and you've ever voted on the chatbot arena, think about what you actually test. You're probably testing the newest, shiniest things. Yeah, exactly. Popular frontier models receive a disproportionate number of battles. Yeah, and the prompts are driven entirely by whatever users find interesting today, not by some structured experimental design. Which creates a severe imbalance.

3:44It makes me think of, like, chess. Okay, how so? If a grandmaster plays a total novice, the outcome tells us almost nothing new. We know the grandmaster is going to win. The information value of that match is practically zero. Yeah. But if two grandmasters play each other, every single move, every decision is rich with new data about their relative skills. That is a perfect analogy. And statisticians quantify that specific concept using local Fisher information. Fisher information. Right. In the Bradley-Terry-Luce model, the Fisher information tells us how much statistical weight a single observation carries.

4:20And that weight is highest when the two models are evenly matched. So when it's a toss-up. Exactly. When the outcome is close to 50-50, the mathematical variance of the logistic curve is at its peak. And that high variance actually provides high Fischer information. Because we are genuinely learning something from the uncertainty. Yes. But when you have lopsided comparisons like a frontier model absolutely crushing an old baseline model, the outcome is near deterministic. Right. Those matchups give you almost no Fischer information. They carry very little statistical weight, yet they clog up the entire data set.

4:56And because the leaderboard is driven by human curiosity, we end up with a mountain of data, but a huge chunk of it is those low-information matchups, or, you know, matchups concentrated on just two or three models. Which leaves the rest of the data incredibly patchy. Very patchy. So, how do these MIT and Purdue researchers propose we fix it? Because we can't just evaluate every single model on every single task independently. For a lot of these matchups, we just don't have enough high-quality data. Well, the core innovation in this paper is moving away from looking at tasks in isolation. They introduced the concept of representing model performance as a low-rank latent score tensor.

5:35Okay, let's break that down for the listener, because tensor is a heavy word. A tensor, in this context, is basically an array of data, but it extends beyond a standard 2D spreadsheet, right? Exactly. A standard 2D matrix might just list models on one axis and broad task categories like math or creative writing on the other. Right. But a tensor expands into multiple dimensions. You can add a third dimension for specific user groups, a fourth for evaluation criteria, or a fifth for languages. Oh, wow. Yeah, it becomes a massive multidimensional cube mapping out every facet of latent abilities across every possible scenario.

6:08And the key phrase they use in the paper is low rank. What does it mean for this giant multidimensional cube of AI abilities to be low rank? So, the low rank assumption posits a structural theory about how intelligence, or at least AI performance, operates under the hood. Okay. It assumes that an LLM's performance across thousands of different, highly specific microscopic scenarios is actually driven by a relatively small number of underlying latent factors. Like what? For example, maybe one foundational factor maps to logical reasoning ability. Another maps to coding syntax proficiency. Another to style sensitivity.

6:47By assuming the massive tensor is low rank, you compress the complexity. You drastically reduce the effective dimensions the math needs to calculate. So the math can say, ah, Model A has a high underlying factor for logical reasoning, so even though we don't have many battles of it writing C++ code, we can infer it will probably be pretty good at it. Exactly. That's the mechanism of borrowing strength across different contexts. And they proved this works with real data, too. Yeah, they filtered the ARENA dataset down to 30 models and 10 task categories, yielding over 81 ,000 pairwise observations.

7:18By applying this low-rank tensor model, it vastly outperformed a naive per-category baseline that just treats every single task as an isolated island. I mean, I have to push back a little here, though. Compressing an AI's vast trillion parameter knowledge down to just a few underlying low-rank factors, it feels like it might oversimplify reality. It's a fair point. Like, doesn't a model sometimes just have a bizarre, highly specific blind spot? Like, it's brilliant at complex physics equations, but uniquely terrible at formatting a simple bulleted list. AI models absolutely have idiosyncratic blind spots, and the researchers acknowledge that reality.

7:56But this is a classic bias-variance trade-off in statistics. How so? If you try to model every single unique quirk and blind spot independently, you run straight back into the sparsity problem. You simply don't have enough user data to measure every model on every hyper-specific quirk with any degree of confidence. Oh, I see. So the low-rank assumption might smooth over a few hyper-specific anomalies, but it is statistically essential to make inference feasible across high dimensions. Otherwise, you just completely drown in the noise of the data. Okay, that makes sense. So we've established this beautiful, multidimensional, low-rank tensor to map AI abilities, and it helps us borrow strength across sparse data.

8:38But this is where the paper gets really intense, because when the researchers try and actually calculate the uncertainty around these rankings, they hit a massive technical roadblock. They did. They needed to find the semi-parametric efficiency bound. And what is that? That is the absolute fundamental limit of how precise our estimates can mathematically be. To do that, they have to use an information operator and project it onto the low-rank tangent space of the tensor. And here's where it gets really interesting, because the math fundamentally breaks. The central technical bottleneck is something called non-commutativity.

9:11Yes. The information operator does not commute with the projection onto the tangent space. Okay, I'm going to jump in here because information operators not commuting with tangent space projections that is heavy. Let's unpack commutativity for a second. Sure. In basic math, A times B equals B times A. Three times four is the same as four times three. That's commutativity. But in higher level linear algebra, the order of operations matters. A times B gives you a totally different result than B times A. Exactly. Standard matrix completion, like the algorithms that predict what movie you want to watch on a streaming service, assumes the data's geometry is what statisticians call isotropic.

9:52Isotropic, meaning uniform. Yes, uniform. In an isotropic space, operators commute smoothly. But our AI data, because of the wildly fluctuating Fisher information we talked about earlier, is anisotropic. Okay, let's visualize why this anisotropic geometry breaks the leaderboard math. Think of it like driving through a city. I like this analogy. If the space is isotropic, it's like driving on a massive salt flat. The speed limit is the exact same in every single direction. But our data is an anisotropic city grid. Right, where the local speed limits change constantly depending on the neighborhood.

10:23Exactly. The speed limit changes based on two things at once. First, the type of car you're driving, which represents the Fisher information, or how evenly matched the AI models are. Yep. And D, the specific neighborhood you are driving through, which represents the tangent space of that specific task. If you calculate the speed limit first, you might assume you have a flat road. But if you map the neighborhood first, you realize you're actually driving up a mountain. And that's a huge problem because the algorithm statisticians use to find the true rank and calculate uncertainty require going backwards from the destination to the start.

11:01And this is the loop. Because the route changes depending on which direction you calculate it, you can't just apply a universal formula. No, you can't. When statisticians try to invert the matrix to find the uncertainty bounds, the fact that the operators don't commute creates a dimension-dependent bottleneck. The math blows up. It gets trapped. It makes figuring out exactly how confident we are in a model's rank seem totally impossible when the data scales up. Right. They are staring at this bumpy, unpredictable, anisotropic mathematical grid, and they can't run their standard debiasing algorithms.

11:34So they had to invent a workaround to force the geometry to behave. Just... They introduced a method called score whitening. Score whitening. Mechanically, how does whitening fix the loop? It tackles the heterogeneous fissure information directly. Remember, some battles are highly informative, carrying a lot of weight, and some are near deterministic and carry almost none. Right, the grandmaster versus novice problem. Exactly. Score whitening equalizes the local fissure information across all comparisons. They do this by normalizing the score. In statistical terms, they divide the derivative of the log likelihood by the variance.

12:11Wait, I need to pause you there. Dividing by the variance, meaning dividing by the Fisher information, isn't that like a statistical volume knob? Sort of, yeah. Because if a grandmaster crushes a novice, the Fisher info is low. The variance is low. If we are turning down the volume on the highly informative matches and dividing by a small number to turn up the volume on the boring lopsided matches, Aren't we artificially making a boring match look incredibly important? Doesn't that skew reality? I know. It's highly counterintuitive, but no, it doesn't skew reality. The key to the math is that the whitened score retains a conditional mean of zero.

12:46Okay. Because the core expected value doesn't shift, it doesn't introduce bias into the truth of the data. Instead, it acts as a semi-parametric preconditioning. Okay. Semi-parametric preconditioning is a mouthful. For those of us without a PhD in statistics, are you saying it basically just irons out the wrinkles in the data? Pretty much. It actively smooths the bumpy city grid back into a flat salt flat. By equalizing the fissure information, the effective operator becomes isotropic, becomes uniform on the tangent space. So it removes the commutativity bottleneck entirely. P times B can behave predictably again.

13:21Yes. Because the effective operator is now isotropic, they can mathematically construct a one-step de-biased estimator. This estimator achieves what's known as asymptotic normality, and it does so at the optimal sample complexity scale. That's amazing. They found a way to make the matrix inversion stable without losing the core truth of who won the matchups. It really is a brilliant piece of mathematical engineering. But whitening only fixes the math if we assume the data is there to be whitened. What do you mean? What about the fact that nobody wants to test the boring models? The chatbot arena is still overwhelmingly flooded with GPT-4 and clod prompts.

14:00Even if you smooth out the fissure information, the sampling itself is still totally skewed towards a handful of models and tasks. Ah, right. Well, the researchers didn't ignore the sampling bias. they extended their score whitening technique with inverse probability weighting, or IPW, specifically to handle the non-uniform sampling. How does IPW actually work in practice, like on the real leaderboard? Let's take the real-world data set. The task category labeled Code Technical has over 19 ,000 observations. It's heavily tested. Sure. But a category like Code General has only about 2 ,600. IPW literally weights the data inversely to its probability of being sampled.

14:42Okay, I think I see where this is going. It multiplies the influence of the rare code general matchups and fractionally reduces the overwhelming volume of the code technical matchups. It mathematically re-weights the entire data set to mimic a perfectly uniform design. So the model gets to treat the data as if users were perfectly democratic in what they tested, further smoothing out those anisotropic bumps. Exactly. And they push the framework even further by extending it to nonlinear estimators. Translating that into plain English, what does a nonlinear estimator give us that a flat score doesn't?

15:14Instead of just calculating a flat score gap, saying model A has an ELO of 1 ,200 and model B has an ELO of 1 ,150, a nonlinear estimator allows you to calculate the precise probability of an event. Oh. Yeah. You can figure out the exact percentage chance that a specific model will beat another model on a highly specific task, like creative writing in French. It translates the abstract multidimensional tensor back into a real-world, highly specific probability, complete with accurate confidence intervals. And when they unleashed this entire statistical toolkit on the actual chatbot arena dataset, the low-rank tensor, the score whitening, the IPW, how did it perform against the way we do things now?

15:57The efficient estimator produced significantly narrower confidence intervals and better coverage compared to standard naive baselines. Meaning we aren't just guessing based on a noisy win rate anymore. We can actually trust the rankings. It provides a principled statistical framework for truly quantifying our uncertainty. When we make the claim Model A is better than Model B, we can back it up mathematically, even with messy, sparse data. So bringing it all back to you, the listener, the next time you check an AI leaderboard to decide which model you're going to rely on for your work, your coding, or your writing, don't just look at that raw surface level score.

16:32Right, because it might be misleading. Exactly. The gap between two scores might just be a statistical illusion caused by low fissure information and skewed sampling. You need to realize that the confidence interval around that score is what actually dictates if the model is reliably better. And thanks to techniques like tensor completion and score whitening, the industry is getting much closer to showing you the true capabilities of these tools. That's a huge step forward. It really is. To quickly recap our journey today, we learned why raw pairwise AI battles are statistically messy and why lopsided matchups give us almost no fissure information.

17:09We saw how researchers map AI abilities into multidimensional tensors, assuming an AI's vast knowledge is driven by a few low-rank latent factors. We explored the headache of anisotropic geometry, where the order of mathematical operations traps you in a loop, and we saw how the brilliant volume knob of score whitening flattens the math out so we can actually get reliable confidence intervals. Beautifully summarized. Thanks. But, you know, I want to leave you with a final thought to mull over building on that low-rank assumption. The math here works specifically because we assume an AI's intelligence can be compressed into just a few latent factors.

17:44If we can mathematically prove that the vast complex capabilities of advanced AI models, models that have ingested the entire Internet, can be accurately mapped this way, well, what does that say about human intelligence? That's an interesting thought. Right. Are your own complex professional and creative skills secretly just a low-rank tensor waiting to be mapped? Are we all just a combination of five or six underlying latent variables interacting in different dimensions? or is there a fundamental anisotropic geometry to the human mind? A complexity that resists the math. Exactly. A complexity that constantly changes its own speed limits, fundamentally resisting being whitened or compressed into a tidy statistical matrix.

18:27We started today talking about how AI evaluation is like looking at a broken x-ray machine. But maybe as we refine these statistical lenses, we aren't just building a better x-ray for artificial intelligence. Maybe we are slowly building a mirror. That is heavy stuff. It is. Thanks for joining us on this deep dive. Keep questioning the data and keep exploring.

From the publisher

This paper introduces a rigorous statistical framework for evaluating Large Language Models (LLMs) by treating the problem as a low-rank tensor completion task. The researchers address the challenges of chatbot leaderboards, such as those on platforms like Chatbot Arena, which rely on noisy and sparse human preference data from pairwise model comparisons. By assuming that model performance across various tasks and contexts is driven by a small number of latent factors, the authors demonstrate how to "borrow strength" across categories to improve accuracy. They develop semiparametric efficiency bounds and a debiased one-step estimator to provide reliable confidence intervals and uncertainty quantification for model rankings. To resolve technical bottlenecks caused by non-uniform sampling, they introduce a score-whitening method that stabilizes inference across heterogeneous matchups. Their findings offer a principled approach to constructing more robust, statistically sound leaderboards for the rapidly evolving field of AI evaluation.

More from Best AI papers explained

All 475 episodes
LLM Evaluation as Tensor Completion: Low-Rank Efficiency and Uncertainty QuantificationBest AI papers explained · 19 min
Listen in VO