Correct Looks Better: Pairwise Comparisons Reveal Accuracy Rankings

13 Jun 2026 · 20 min · 9 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

How pairwise comparisons can rank AI model accuracy even when the “judge” model lacks subject-matter knowledge, using hidden structural signals rather than formatting.

Guests

No named guests; the episode is hosted by two speakers discussing the Max Planck Institute paper “Correct Looks Better.”

Guest backgrounds

Not specified in the transcript.

Key claims

Pairwise “blind taste tests” (ELO/Bradley-Terry rankings) correlate with ground-truth accuracy (Spearman > 0.9 on 4/5 benchmarks). Style/formatting bias is negligible after statistical correction. A weak judge can still rank better models because it penalizes “echo” (repetitive rambling caused by poor stop-token/probability calibration).

Notable examples

Benchmarks MMLU Pro, GPQA Diamond, GSM8K (freeform). In Big-Bench, injecting echo into an otherwise correct answer drops win rate from 65% to 29% (and adding echo to the other answer raises it to 81%). Echo is identified as the “hidden tell of incompetence.”

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

The Complexity of AI Evaluation

0:41 to 2:15

Discover how AI models are evaluated and the limitations of human grading.

“We're looking at how models that are totally clueless can miraculously and accurately judge the intelligence of machines that are vastly smarter than they are.”

Understanding Pairwise Comparisons

2:15 to 4:35

Learn about the pairwise comparison method used to evaluate AI models.

“So the industry has turned to using automated systems to grade other automated systems.”

Concerns in Ranking Systems

4:35 to 5:30

Examine the risks of ranking models based on superficial formatting.

“We might be judging a high-stakes college math competition entirely based on which student has the neatest handwriting.”

Accuracy of Pairwise Comparisons

5:30 to 8:15

Understand how pairwise comparisons can accurately rank AI models.

“if the judge is measuring reality or just style.”

The Judge's Performance Mystery

8:15 to 10:40

Explore the surprising accuracy of a poor-performing judge in pairwise tests.

“Even without an answer key, the judge model is consistently, overwhelmingly preferring the more accurate model.”

Identifying Echo as a Hidden Signal

10:40 to 14:03

Learn about 'echo' as a tell for model accuracy in pairwise comparisons.

“Which forces a logical deduction here, because we know for a mathematical fact that the weak judge does not understand the factual material.”

Echo and Its Impact on AI Accuracy

14:03 to 16:01

Learn how structural issues like echo affect AI performance evaluations.

“This wasn't just a loose correlation they noticed.”

The Evaluation Frontier and Human Progress

16:01 to 17:59

Understand the concept of the evaluation frontier and its implications for AI.

“We are rapidly approaching what researchers call the evaluation frontier.”

Human Pairwise Comparisons in Decision Making

17:59 to 19:42

Explore how humans also evaluate information through pairwise comparisons.

“We started by asking if blind taste tests for artificial intelligence are just rewarding pretty formatting and neat handwriting.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Imagine handing a highly complex graduate level calculus exam to a toddler. Like a literal drooling toddler. Okay, I'm picturing it. It's a disaster. Right. If you ask that toddler to grade a single paper and tell you if it passes, they're going to fail catastrophically. Probably just scribble on it with a crayon. Yeah, of course. But suppose you take two calculus papers, hold them side by side, and ask the toddler to point to the better one. And somehow, time after time, the toddler perfectly points out the A-plus student. Which sounds completely impossible. Exactly. But today, we are taking you on a deep dive into an artificial intelligence paradox that operates exactly like that toddler.

0:41We're looking at how models that are totally clueless can miraculously and accurately judge the intelligence of machines that are vastly smarter than they are. It really is arguably one of the most perplexing phenomenons currently baffling the AI industry. We're looking at a brilliant piece of research from the Max Planck Institute today titled, Correct Looks Better. Such a great title, by the way. It is, yeah. And what this paper does is peel back the layers on how these systems actually evaluate one another. Because, you know, as the models we interact with every day become more sophisticated, they generate these beautifully formatted, incredibly confident answers.

1:16Right. They sound amazing. Exactly. But beneath that seductive illusion of competence, we desperately need to know if they're actually correct or if they're just engineered to be exceptionally articulate. Which is a terrifying thought, honestly, when you're using these tools to summarize dense financial reports or, I don't know, explain quantum entanglement. If you as a human struggle to tell the difference between eloquence and factual accuracy, imagine the headache for the researchers actually building these things. Oh, absolutely. So our mission today is to uncover the hidden mechanics of how these systems are graded.

1:51We're going to explore why the entire tech industry is relying on clueless judges running entirely on, well, vibes. And, you know, the necessity of this comes down to pure scale. Human evaluation simply doesn't work anymore. It's too slow, right? Yeah, right. Exactly. We don't have the millions of hours or the hyper-specialized expertise required to sit down and grade every single response these complex neural networks generate. Right. So the industry has turned to using automated systems to grade other automated systems. Makes sense. But as the findings in this research reveal, the way that grading actually happens beneath the surface defies a lot of our basic assumptions about what it means to recognize intelligence.

2:32Okay. Let's unpack this. Because before we can truly appreciate the crazy discovery these researchers made, we have to understand the baseline of how these neural networks are tested in the first place. Yeah. We need to set the stage. My initial assumption was always that researchers just hand a model, a massive standardized test, run the outputs against a rigid answer key, and score it out of 100. In an ideal world, perhaps. But the current gold standard for evaluating generative models is a process known as pairwise comparisons. Pairwise comparisons. Right. Right. So imagine you have a prompt, perhaps a highly abstract logic puzzle or a creative writing task.

3:08Instead of attempting to build an automated grader that can rigidly assess an open-ended essay against a strict rubric, researchers generate two different answers to the same prompt. Okay. So one from Model A and one from Model B. Precisely. They put those two answers side by side, entirely blind, and they ask a third entity, the judge, a very simple question. Which of these two answers do you prefer? So it is quite literally a blind taste test, like a Pepsi versus Coke challenge. But instead of soda, we're tasting math, logic and reasoning. That is the perfect way to conceptualize it, actually.

3:43They run this blind taste test thousands of times across massive data sets, pitting different models against one another in head-to-head matchups. They're building a whole tournament bracket, basically. Yeah. They take all of those aggregated wins and losses and compile them into a ranking system. Usually they use an ELO rating system or a Bradley Terry model. Oh, I heard of that. Yeah, it's the exact same mathematical framework used to rank chess grandmasters or competitive video game players based on their win-loss records against opponents of varying skill levels. On paper, that makes logical sense.

4:16I mean, if you beat a highly ranked chess grandmaster, your score goes up more than if you beat a novice. Exactly. But this brings up that massive raging debate in the tech community. because if the judge in this scenario doesn't have an answer key, if it is purely picking the response it prefers, aren't we just measuring formatting? Ah, yes, the formatting problem. Right. We might be judging a high-stakes college math competition entirely based on which student has the neatest handwriting. And that concern is exactly what keeps developers up at night. The fear is that these preference rankings are completely divorced from underlying accuracy.

4:53Because they're just picking the pretty one. Exactly. Critics have long argued that a judge might just be rewarding superficial stylistic cues. It might prefer answers that use bold text, or nicely formatted bullet points, or just a particularly polite sycophantic tone. Or worse, it might just default to picking whichever answer is longer. I know I used to do that in middle school. We all did, and the fear is that the preferred answers simply look better without actually possessing any factual substance. And because we usually use these pairwise comparisons on open-ended tasks where there is no clear right or wrong answer, it's been nearly impossible to prove if the judge is measuring reality or just style.

5:33So to figure out if it is all just handwriting and formatting, the researchers behind this study did something incredibly clever. They really did. They decided to run this blind taste test using questions where we do possess the absolute undeniable ground truth answer. They pulled in five of the most brutal, rigorous benchmarks in the industry. We're talking about tests like MMLU Pro, GPQA Diamond, and the GSM8K. We should pause and emphasize just how difficult these benchmarks are. GPQA Diamond, for instance, stands for the Google Proof Q &A dataset. Google Proof. So you literally can't just search for the answer.

6:08Right. These are questions crafted by experts in biology, physics, and chemistry. They are so difficult that PhDs in those fields often fail to answer them correctly, even with full access to the internet. Wow. And GSM8K is a massive data set of grade school math word problems that require incredibly tricky multi-step logical reasoning, where one wrong assumption cascades into a completely wrong answer. So they took these notoriously brutal tests and stripped away the multiple choice formatting to make them freeform questions. They had a variety of models generate answers and then they brought in their judge.

6:43Right, the taste tester. Yeah. In this specific experiment, they used a publicly available model called GTT-AUS-20B to pick the winners in those head-to-head pairwise matchups. Completely blind, no access to the actual answer key. And the alignment they discovered was staggering. Statistically, they achieved a Spearman-rank correlation consistently above 0.9 across these rigorous benchmarks. Okay, let me stop you there to make sure I'm visualizing this correctly. A Spearman rank correlation is essentially a mathematical way of measuring how perfectly two separate ordered lists line up with each other, correct?

7:17Precisely. A perfect match between two lists is a 1.0. A score above 0.9 is phenomenal. That's crazy high. It is. It means that if you rank the model strictly by who gave the most mathematically or factually correct answers against the hidden answer key, and then you rank them by who the blind judge simply preferred in the taste test, you get nearly the exact same leaderboard. The researchers use this other stat called Kendall's distance to visualize it, which helped me wrap my head around it. Kendall's distance basically measures how many adjacent items in a list you'd have to swap to make two lists perfectly identical.

7:53Right, like physically moving names up and down a scoreboard. Yeah. On four out of the five brutal benchmarks they tested, you would only need to swap about 6 % to 8.9 % of the model pairs to perfectly match the true ground truth accuracy rankings. What that Kendall's distance metric proves is that pairwise comparisons are absolutely not just measuring formatting. They are capturing a massive amount of genuine signal about correctness. Even without an answer key, the judge model is consistently, overwhelmingly preferring the more accurate model. Here's where it gets really interesting, though.

8:27I have to push back on this premise. Oh, go ahead. If the judge is just blindly picking A or B, and we know for a fact it does not possess the answer key, how is it miraculously recreating the exact same ranking we get when we do have the answer key? I mean, if the question involves PhD-level quantum physics and the judge doesn't have the physics textbook, it should theoretically just be guessing. That is the pivotal question of the entire paper, because to answer your question, the researchers decided to look at what happens when the judge is actually completely terrible at the test it's grading.

9:00They dove into the data from the SimpleQA benchmark, which is designed to be a very difficult factual accuracy test. First, they used a highly capable strong reasoning model, specifically O3. That strong model gets about 59 % of the questions right on its own. Which makes logical sense. When this strong model acts as the judge and evaluates the pairs, it produces great rankings, it understands the material, so it grades well. Exactly. But then they swapped out the judge. They brought back that publicly available GPT-AUS-20B model. And this model utterly fails the simple QA test. Oh, no. On its own, it only gets a 4.9 % of the questions correct.

9:40It is statistically clueless. Wow, 4.9%. That's rough. Yeah. And they proved it by taking this clueless 4.9 % judge and asking it to act as a direct grader. They gave it a single answer from a model and asked it one question. Is this correct? Yes or no? Which brings us back to handing the calculus paper to the toddler. The single grading approach must have been a complete disaster. Catastrophic. The rankings derived from that direct grading were entirely scrambled. It penalized brilliant models and rewarded terrible ones. The Kendall's distance essentially doubled. It was worse than useless as a direct grader.

10:14But, and this is the wild part for you listening, when they ask that exact same clueless 4.9 % judge to look at two answers side by side and just pick the better one, its rankings snapped right back into alignment with the true accuracy, the pairwise miracle. Yeah. The exact same weak system that cannot grade a single paper to save its life can suddenly rank the entire classroom with incredible accuracy just by looking at the papers two at a time. Which forces a logical deduction here, because we know for a mathematical fact that the weak judge does not understand the factual material. It doesn't know the chemistry or the physics.

10:49Right. It physically can't. Therefore, it absolutely must be looking at something else. It has to be relying on a hidden signal, some kind of subconscious tell that separates the smart models from the bad ones. What's fascinating here is that the researchers systematically eliminated the usual suspects to find that signal. They initially assumed, just as we did, that the judge was biased toward formatting. Right, the bold text and bullet points. Right, maybe the smarter models just use more lists and headers. Or maybe there is self-preference, the judge simply picking models built by the same developers that built it.

11:20But they ran the math. They built mathematical frameworks to statistically correct for formatting bias and self-preference. Meaning they neutralized the handwriting advantage entirely. So did the rankings fall apart. The rankings barely moved. Those stylistic biases had only a negligible effect. They were not the driving force. Wow. Yeah. So to find the real culprit, the researchers had to zoom in on a specific massive subset of the data, the non-discriminative pairs. Okay. If a pair is non-discriminative, that means the correctness of the answer isn't a factor. So these are matchups where both Model A and Model B gave the objectively correct answer or both gave the objectively incorrect answer.

12:01Exactly. And this is not some fringe edge case. These non-discriminative pairs made up about 58 % of all the comparisons in the research. Over half. Right, over half. In 58 % of the matchups, correctness was completely off the table as a differentiator. The judge had to be falling back on some other signal to pick a winner. And the researchers found it. The hidden tell of incompetence. They call it echo. Echo. Let's explain echo clearly because this is the secret sauce. Echo is what happens when a model gives an answer, but then it fails to stop generating text. It just keeps talking. It is a fundamental failure of the system's internal stopping mechanism.

12:42I want to clarify that mechanism. When we talk about a stopping mechanism, we're referring to a stop token, right? That is basically the system's internal brake pedal. That is exactly what it is. To understand why a weak judge mathematically penalizes ECHO, we have to briefly look at how language models actually work. They are next token predictors. They calculate the probability of the next word. A highly intelligent, well-calibrated model calculates the correct answer and then immediately assigns the highest mathematical probability to outputting its internal stop token. It puts on the brakes.

13:13But a poorly calibrated model doesn't have that confidence. It might stumble into the correct answer. Its probability distribution is a mess. Exactly. So instead of hitting the brake pedal, it assigns high probability to repeating itself. Yeah. It might say, the answer is 42. But then it echoes. The answer is 42. The answer is 42. Yes. Or it will start inventing fake question and answer templates to fill the void. The answer is 42. Question, what is 2 plus 2? Answer, 4. It just rambles. Precisely. And this isn't about the model feeling nervous in a human sense. It is a structural mathematical failure of probability calibration.

13:52Repetitive rambling text structurally looks like low probability, poor quality output within a neural network. And the researchers proved causally that judges heavily penalized this structural failure. The causal proof on this was brilliant. This wasn't just a loose correlation they noticed. They went into the Big Bench hard benchmark data. They found matchups with two perfectly clean, correct answers, A and B. Answer A was winning the blind taste test 65 % of the time. It's a solid win rate. Right. Then the researchers played God. They artificially injected echo into answer A. They simply copy pasted its answer a few times at the end of its response to simulate that failure of the stop token.

14:32And answer A's win rate plummeted from 65 % down to 29%. That's a massive drop. By artificially making the text structurally resemble a poorly calibrated output, they caused it to lose. And conversely, when they added echo to answer B instead, answer A's win rate shot up to 81%. The judge heavily penalizes echo, especially when both answers are otherwise equally correct or incorrect. So it's basically a massive red flag for the judge. Exactly. Avoiding echo acts as a proxy for competence. A system that structurally knows when to hit the brake pedal is generally a system that actually possesses the internal logic to know what it is talking about.

15:08So we have essentially solved the mystery of the toddler grading the calculus test. The toddler isn't looking at the complex integration or the derivatives. The toddler is penalizing the structural mess on the page. Exactly. The weak judge cannot verify the facts of the GPQA diamond biology questions, but its neural network is highly attuned to the structural fingerprint of a poorly calibrated model. It punishes the echo. It highlights beautifully why pairwise comparison works, even when the judge is entirely ignorant of the subject matter. Conciseness and the absence of repetitive rambling are incredibly highly correlated with factual accuracy across all these benchmarks.

15:50So what does this all mean? Why should you, listening to this deep dive on your commute or while cooking dinner, care about Bradley Terry scores, non-discriminative pairs, stop tokens, and AI echo? If we connect this to the bigger picture, this entire phenomenon directly impacts the future of human progress. We are rapidly approaching what researchers call the evaluation frontier. The evaluation frontier. Yes. This is the threshold where these artificial systems become vastly smarter than any living human expert in specialized, highly complex domains. I want to ground that in a concrete scenario.

16:22We're talking about like an AI trying to design a novel protein structure to cure a specific type of cancer or an AI running advanced material science simulations to invent a new type of solid state battery. Exactly. When a system generates a completely novel cure for a disease or solves a fusion reactor physics equation, there is no answer key. The ground truth does not exist yet. We haven't run the multi-year clinical trials or built the battery in a lab. Right. So if the smartest human scientists look at the AI's math and say, this is too complex for us to verify directly, how do we know if it's safe to proceed?

16:58How do we evaluate a superintelligence when we humans are the toddlers? Oh, wow. The paradigm shifts entirely. We literally become the weak judges. We do. And this research proves an incredible, almost paradoxical concept. We can reliably rank systems by accuracy without actually knowing the absolute accuracy of the outputs. That's wild to think about. It is. We can use slightly weaker, older systems to safely evaluate the next generation of superintelligent tools simply by having them compare notes. If we put two superintelligent systems head-to-head on that cancer protein problem, a slightly weaker AI judge or even a human judge relying on these automated tools can reliably point to the better answer by reading the structural fingerprints of competence.

17:42We don't need to possess the absolute truth to mathematically figure out who is closer to it. That is incredibly reassuring. It means the concept of evaluation scales. We aren't going to be totally flying blind when these technologies surpass our own cognitive limits. Right. We have a path forward. Let's do a quick recap of the journey we've been on. We started by asking if blind taste tests for artificial intelligence are just rewarding pretty formatting and neat handwriting. But by digging into the ground truth benchmark data, we found out they aren't. Right. They really aren't. Through hidden structural signals like echo, the failure of internal probability calibration that leads to rambling, even incredibly weak, clueless systems can perfectly rank the intelligence of vastly stronger ones just by looking at them two at a time.

18:30Pair-wise, comparisons actually work. This raises an important question, though, because as we marvel at how these artificial neural networks use structural proxies to evaluate intelligence they cannot possibly comprehend, we really have to look in the mirror. Don't we humans do the exact same thing? Oh, we absolutely do. Think about the last time you listened to two financial advisors debating the intricacies of macroeconomic policy, or two politicians arguing over the cascading effects of a complex tax law, or two specialists discussing a highly nuanced medical treatment. You just kind of zone out.

19:04Right. The topic was likely way over your head. You didn't possess the ground truth answer key. So were you actually evaluating their factual accuracy? Or were you just running a subconscious pairwise comparison? Were you just looking for the human equivalent of Echo who seemed defensive, who repeated their talking points, who nervously rambled, and simply choosing the person who structurally sounded like they knew when to stop talking? That is such a piercing thought. We are all just weak judges running our own pairwise comparisons every single day, hoping we're picking up on the right structural fingerprints of confidence.

19:37Something for you to mull over the next time you ask a human expert for advice. Thank you so much for joining us on this deep dive into the hidden mechanics of evaluation. We love unpacking these complex ideas with you. Stay curious, watch out for the echo, and we will catch you on the next one.

From the publisher

This research explores whether pairwise comparisons used to rank generative models actually reflect ground-truth accuracy. By converting multiple benchmarks into free-form formats, the authors found that Elo-style rankings achieve a remarkably high correlation with objective correctness. Surprisingly, this alignment remains strong even when the judge model is weaker than the candidates it evaluates, outperforming direct grading methods. While critics often worry about judge biases or stylistic cues, the study demonstrates that these factors have a minimal impact on the final model hierarchy. Furthermore, the paper identifies "echo"—or repetitive output—as a key reason why judges prefer one answer over another when both are technically correct. Ultimately, the results suggest that relative preferences are a robust and reliable proxy for absolute accuracy in competitive model evaluation.

More from Best AI papers explained

All 475 episodes
Correct Looks Better: Pairwise Comparisons Reveal Accuracy RankingsBest AI papers explained · 20 min
Listen in VO