FUSE: Ensembling Verifiers with Zero Labeled Data

14 May 2026 · 20 min · 8 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

The episode explains FU-SE (Fully Unsupervised Score Ensembling), a Stanford/Google method for ranking AI-generated answers using multiple AI verifiers without any labeled data or answer keys. It addresses “test-time scaling” where models sample many candidate responses and rely on verifiers/reward models to pick the best.

Guest backgrounds

No guest names or bios are provided in the transcript.

Key claims

Naively averaging or majority-voting verifier scores fails because verifiers are query-dependent and often hallucinate in correlated ways. FU-SE uses triplet conditional independence (TCI) plus spectral binarization to make verifier outputs satisfy assumptions, then estimates each verifier’s accuracy, generates pseudolabels (synthetic answer keys), and trains a logistic-regression aggregator to select the final response.

Notable examples

Humanity’s Last Exam (649 questions; random 52.1%, naive ensemble 51.4%, FU-SE 54.3%); IMO shortlist (123 problems with independence; naive ensemble 63.8%, FU-SE matches).

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Understanding the FUSE Method

1:21 to 2:18

Learn about the FUSE method's innovative approach to grading AI responses.

“Our mission for you today is to explore a fascinating new research paper from Stanford and Google researchers.”

The Ineffectiveness of Naive Ensembles

2:18 to 4:43

Examine why traditional methods for averaging AI scores often fail.

“Like we used to just prompt an AI and it would spit out its first guess, just a one-shot deal.”

Limitations of Traditional Crowdsource Assumptions

4:43 to 7:17

Discover the flaws in classical assumptions about judge independence.

“You have this shifting, unstable population of toddlers and experts, and it changes depending on the exact question being asked at any given millisecond.”

Innovations of FUSE: From Correlation to Triads

7:17 to 11:22

Understand how FUSE shifts the analysis from individual judges to groups of three.

“The math gets tricked because the judges are essentially colluding in their wrongness.”

Real-World Application and Testing of FUSE

11:22 to 14:00

Explore how FUSE performs against other models in rigorous testing scenarios.

“Once IFYES knows the secret skill level of every judge, it creates what the paper calls pseudolabels.”

Exploring the Performance of FUSE

14:00 to 18:06

Learn how FUSE effectively filters out flawed AI judges to improve performance.

“If an AI generates responses and you just pick one at random, a metric called pass at one, you score 52.1%.”

A Major Turning Point in AI

18:06 to 18:35

Understand the implications of reduced reliance on human verification in AI.

“We possess the mathematical tools to extract genuine truth from a panel of deeply flawed AI models purely through unsupervised algorithms.”

The Future of Unsupervised AI

18:35 to 20:15

Explore the potential of AI systems to self-grade and innovate without human input.

“We are rapidly moving past the area where AI requires constant human handholding to grade its work.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Imagine taking like the hardest exam of your entire life. You sweat through hours of advanced physics, complex coding, obscure logic puzzles. Oh, yeah. Just a total nightmare scenario. Right. And you finally hand your paper in, feeling this mix of dread and relief. But then you look up and you realize the teacher staring back at you doesn't have the answer key. Which is terrifying. Yeah. In fact, the teacher has no idea how to solve these problems either. And, well, that exact crisis is playing out right now at the absolute bleeding edge of artificial intelligence. Oh, it really is. I mean, we are building AI models that are taking the hardest academic and professional tests ever conceived by humanity.

0:40And the sheer volume of these highly complex answers they generate is it's so massive that paying human experts to sit down and grade this homework is becoming totally impossible. Exactly. It is way too expensive. And the process is just entirely too slow for the pace of development. The industry has basically hit the ultimate bottleneck. Right, because we have engineered machines capable of generating thousands of, you know, highly technical potential solutions in just a matter of seconds. But we simply lack the human brainpower and, frankly, the financial budget to actually check their work.

1:12We are generating intelligence faster than we can verify it. So how do you grade the smartest student in the room when nobody holds the answers? Well, welcome to the Deep Dive. Our mission for you today is to explore a fascinating new research paper from Stanford and Google researchers. It's a really incredible piece of work. It is. We are unpacking a method they call FU-SE, which stands for Fully Unsupervised Score Ensembling. We're going to look at the trap of relying on flawed AI judges and explore how an incredibly clever mathematical trick allows us to find the most accurate judges without ever knowing the right answers ourselves.

1:51And we'll see how this exact method is currently dominating the absolute hardest academic tests on Earth. Okay, let's unpack this. So to understand the necessity of UC, we really need to examine the secret sauce driving the most recent wave of AI advancements. The field has shifted heavily toward a concept called test time scaling. And specifically, they are relying on a technique known as best event sampling. Actually, I want to pause on best event for a second because it really represents a total paradigm shift. It does, yeah. Like we used to just prompt an AI and it would spit out its first guess, just a one-shot deal.

2:28But now, instead of giving you its first answer, the AI generates, say, 100 different responses to your single query. Right. It basically thinks through 100 distinct logical paths. But that immediately creates a new hurdle. You have 100 potential answers and you can only show one to the user. Exactly. So developers introduce what they call a verifier. Which is another AI model. Yes, often called a reward model. And its sole purpose is to score those 100 responses and crown the winner. The theory is actually quite sound. Yeah, you have one AI act as the brainstorming engine and a second AI act as like the strict editor.

3:04Right. But the structural flaw in that pipeline, unfortunately, is that these AI verifiers are notoriously imperfect. Very imperfect. And because human ground truth labels are prohibitively expensive at this kind of scale, developers attempt to bypass the imperfection by bringing in a whole panel of AI judges. So they aggregate the scores from, what, 30 or 40 different verifiers? Yeah, hoping to find some sort of reliable consensus. But that creates an absolute disaster in practice. I mean, it's like tossing a PhD level advanced calculus exam into a room where half the people are actual mathematicians and the other half are very confident toddlers armed with red markers.

3:45That is a very accurate analogy. Right. And if you use what developers call a naive ensemble, which just means mathematically averaging all their scores together, you get a mess. Or if you take a simple majority vote. The sheer volume of toddlers randomly scribbling on the paper will completely drown out the actual mathematicians. You literally just get noise. Because the naive ensemble assumes every voice in the room deserves equal weight. Which is crazy. It is. When you treat the scores from highly unreliable, hallucinating verifiers on par with the brilliant ones, the final selection just degrades significantly.

4:19The noise overpowers those isolated signals of truth. And the real nightmare here is that a judge isn't universally good or bad. Verifier strength is heavily query dependent. Yes, that is a huge factor. Like a verifier that acts like a tenured math professor when judging a calculus problem might instantly turn into one of those toddlers when asked to evaluate a piece of poetry. Or a niche software coding problem. Right. You have this shifting, unstable population of toddlers and experts, and it changes depending on the exact question being asked at any given millisecond. That query dependence means a static list of trusted judges is impossible to maintain.

4:57You know, you cannot simile identify the five best models and just use them forever. Because their reliability fluctuates with the topic. Exactly. Which brings us to a massive wall. We literally cannot deploy these advanced models reliably unless someone finds a way to grade the graders. And do it affordably. Right. If we can't afford human answer keys, and we can't just average the AI judges because their terribleness changes depending on the question, How do we figure out which AI judges to trust? Well, FUSC was engineered to solve that exact dilemma. Okay. It operates as an elegant, purely mathematical method that figures out how to weigh these AI judges dynamically on a question-by-question basis.

5:35And the crazy part is it does this completely unsupervised. The critical factor is that it achieves this with 0 % labeled data, 0 ground truth. It literally never sees an answer key. It is grading the graders in total darkness. Let's dig into the engine of this thing because it honestly sounds like magic to me. It really does at first glance. So the paper mentions that FUSE builds on older statistical literature for crowdsourcing, specifically pointing back to a 2015 paper by Jaff and colleagues. Yeah, the challenge of filtering truth from a noisy crowd has a long statistical history. Older crowdsourcing methods fundamentally relied on a strict mathematical assumption known as joint conditional independence or JCI.

6:17Right. And JCI basically assumes that all the judges make mistakes in total isolation from one another. Like if Judge A gets a question wrong, it has absolutely no bearing on why Judge B got it wrong. Their errors are independent. That's the classical assumption. And I can see how that might work if you survey, you know, a random group of humans on the street. But I was reading about how these large language models are trained and they all read the exact same Internet. Yes, they do. If there is a common misconception on a massive Wikipedia page or a popular Reddit thread, five different AI judges are all going to confidently make the exact same mistake.

6:55They're highly correlated. Exactly. That shared pre-training data creates identical blind spots. Right. And the reinforcement learning pipelines compound the issue, causing multiple models to fail in the exact same manner. So you can't use the old math. No, applying the old 2015 JCI method to AI verifiers causes the entire algorithm to break down. The math gets tricked because the judges are essentially colluding in their wrongness. They're all confidently wrong together. Exactly. So FUSI steps away from that outdated model. Instead of requiring joint conditional independence, FUSI introduces its true innovation by relying on a much weaker, more realistic assumption.

7:36Which is? Triplet conditional independence or TCI. Striplet conditional independence. Yeah. So we're moving from looking at the entire crowd of judges as one big group to analyzing them in triads, like groups of three. Yes. By breaking the panel down, a FUSI examines the second and third order covariance tensors of the verifier scores. Okay, hold on. I need to translate covariance tensors for a second because I'm totally stuck on this. Sure, yeah. We are essentially talking about measuring how sets of two and sets of three verifiers agree or disagree with each other across different generative responses.

8:09Yes, tracking their variance. But how does looking at them in groups of three suddenly bypass their shared bias? It still sounds like mathematical alchemy. If I have three blind judges, analyzing their variance doesn't magically give me the truth. What's fascinating here is how TCI leverages those triads to isolate the signal. Think about the underlying math of agreement. Okay, if Judge A and Judge B always make the same mistake, their covariance is high. Right, they agree a lot. But if you introduce Judge C into the equation and monitor how the three of them interact across a hundred different potential answers, the statistical patterns begin to shift.

8:46Oh, I see. TCI assumes that conditional on the true answer, the errors of any three judges don't perfectly correlate in a way that permanently hides the truth. So they won't all fail the exact same way every single time. Right. The covariance tensor is essentially a 3D matrix tracking those shifting allegiances. By mapping out where these triplets overlap and where they diverge, the math begins to separate the correlated hallucinations from the genuine signal. Okay. I want to try an analogy here. Yeah. Is it somewhat like triangulating a cell phone signal? Oh, I like that. Yes. Because if you only have one cell tower, you just have a massive, unhelpful radius, right?

9:26You know they're out there, but no idea where. Exactly. And two towers give you a narrower overlap, but there's still ambiguity. But the moment you introduce a third tower, the intersecting lines allow you to pinpoint an exact location you can't physically see. The geometry of your analogy holds up incredibly well. Instead of physical distance, we are mapping the variants in their agreements and disagreements. And the third point of reference breaks the ambiguity. Yes. But if U.S. doesn't just run raw scores through this triad filter... Why not? Because the raw probability scores output by AI judges are often terribly calibrated.

10:03Oh, right. Like an AI might claim to be 90 % confident in a completely wrong answer? Exactly. Because of that poor calibration, the raw data frequently violates even the weaker TCI rule. So the data is entirely too messy to triangulate. How does it clean up the data before the math breaks? FUSC adaptively transforms those scores using spectral algorithms. Spectral algorithms. Yes. It analyzes the raw probability distributions and mathematically warps or binarizes them. Binarizes them. So it turns them into ones and zeros. Exactly. It finds specific numerical thresholds to turn those wishy-washy percentage scores into hard yes or no votes.

10:41It forces the judges to take a definitive stance. Yes. By finding the perfect threshold to binarize the data, it actively forces the data structure to satisfy that triplet conditional independence rule. Wow. So it molds the environment so the triangulation can actually work. The spectral algorithm acts as a formatting tool, preparing the data for the final evaluation. Once the scores are transformed and TCI is mathematically satisfied, FUS calculates the estimated accuracies for every single judge in the panel. So it figures out who's good and who's bad. It isolates the sensitivities and specificities.

11:16It identifies exactly who is highly skilled at spotting a correct answer and who is merely guessing. Which leads to the final mechanism. Once IFYES knows the secret skill level of every judge, it creates what the paper calls pseudolabels. Right. And these are essentially temporary, highly educated guesses at what the real answer key should look like based entirely on the math we just discussed. Generating those pseudolabels provides a synthetic ground truth. It builds its own answer key. Yes. And FUS takes those pseudolabels and uses them to train a final aggregator, specifically a parametric model like logistic regression.

11:54This freshly trained aggregator then looks back at the original matrix of a hundred generated responses and selects the optimal final answer to present to the user. That is an incredibly elegant loop. It uses the variance of the triplets to find the underlying truth, creates a temporary answer key from that truth, and uses that key to train a final boss judge that makes the ultimate decision. Exactly. And it does all of this on the fly, entirely unsupervised the moment the user asks the question. Within seconds. But I want to push on the practical reality of this. Mathematical elegance is great in a vacuum, but how does this blind process perform out in the wild against models that actually get to cheat?

12:34That's the real test, and the researchers constructed a rigorous gauntlet to test exactly that. They pitted FUSE against unsupervised baselines, like the naive ensemble and simple majority vote. But the real trial was testing FUSE against semi-supervised models. Right. The heavyweight in that category is a method known as Weaver. And we need to emphasize the massive advantage Weaver brings to the table. Weaver is semi-supervised, meaning it gets to see 5 % of the human-labeled ground truth answers. It's a huge advantage. Yeah, it gets a sneak peek at the actual answer key, allowing it to calibrate its weights and learn the biases of the specific test.

13:11Whereas Fuse gets absolutely 0%, it walks into the room completely blind. That 5 % calibration data is an immense head start for Weaver. And the testing ground for these models spanned a massive variety of data sets, pushing into the absolute frontier of difficulty. Here's where it gets really interesting. Let's talk about the first major highlight from the paper, Humanity's Last Exam, or HLE. Oh, this is a brutal benchmark. Absolutely brutal. We are talking about 649 questions spanning highly technical fields, graduate-level physics, obscure biology, advanced humanities. It is designed specifically to be unsaturated by frontier models.

13:50Models like Gemini 3 Pro and GPT-4 struggle immensely on this benchmark. It represents the current ceiling of machine capability. Yes. And the stakes on humanity's last exam are perfectly illustrated by the baseline scores. If an AI generates responses and you just pick one at random, a metric called pass at one, you score 52.1%. Just slightly better than a coin flip. Right. But if you implement a standard naive ensemble to grade the answers, gathering all the AI judges and averaging their scores, the performance drops to 51.4 percent. Which is literally worse than random guessing. The confident toddlers completely overran the room and the flawed judges dragged the entire consensus down.

14:33The drop in performance vividly demonstrates the danger of unweighted averaging on complex tasks. The hallucinated noise completely swallows the factual signal. However, FUSC, operating without any access to the answer key, successfully pushes through that noise. It does. It mathematically identifies the few verifiers who actually comprehend the physics or biology questions, downweights the confused models, and achieves a score of 54.3%. That's amazing. It beats the random baseline and decisively defeats the naive ensemble. It actively weeded out the bad judges without knowing what the right answers were.

15:10Exactly. But the researchers didn't stop there. They threw a fascinating curveball to see if FUSE would break under different conditions. They tested it on a second elite benchmark, the IMO shortlist. The International Mathematical Olympiad shortlist contains 123 elite math problems. And these aren't just standard problems, right? No. Crucially, human experts actively modified these specific problems to prevent AI models from simply regurgitating memorized answers from their training data. They can't just recite Wikipedia. Exactly. The models are forced to genuinely reason through the mathematics step by step.

15:45And there is a brilliant structural twist hidden inside this specific data set. On the IMO shortlist, the set of AI verifiers provided by the researchers are totally different from the standard crowd. Yes. In this isolated scenario, the verifiers actually are highly independent from one another, and they happen to be roughly equally skilled. It creates a rare anomaly where the underlying assumption of the naive ensemble is actually valid. Wow. Because the verifiers are equally skilled and make truly independent mathematical errors, averaging their scores works perfectly. The noise cancels out.

16:20It does. On this specific test, the naive ensemble acts as a perfect oracle, achieving a massive score of 63.8%. So the naive ensemble is king here. The simple average wins. My immediate thought is that ESE, with its complex covariance tensors, spectral binarizations, and pseudolabels, is going to overcomplicate things. Does a complex method break when a simple average is all you need? If we connect this to the bigger picture, the IMO shortlist results reveal ESE's greatest strength, its dynamic adaptability. Adaptability. Yes. Other sophisticated data-dependent methods, including those semi-supervised models like Weaver that had access to ground truth labels, end up overthinking the problem.

17:03Oh, really? Yeah, they search for complex correlations and hidden weights that simply do not exist in this clean data set, and consequently they fall behind the Oracle score. They outsmart themselves. They try to find a pattern in the noise that isn't there. FUSE avoids that trap entirely. It analyzes the variance through its triplet matrices, recognizes that the verifiers are equally reliable and independent, and perfectly matches the 63.8 % Oracle score. Incredible. It calibrates to the exact environment it finds itself in. So it scales to the messiness of humanity's last exam, but it also scales back to handle the clean data of the Math Olympiad.

17:41And when you zoom out and look at the aggregate results across massive models like LAMA-38B and 70B, across all the tested benchmarks, FUSE wins 27 out of 40 comparisons against models that had human labels. It's a huge victory for unsupervised methods. It is actively beating methods that get a 5 % cheat sheet, proving an incredible level of versatility. The data emphatically validates the premise. We possess the mathematical tools to extract genuine truth from a panel of deeply flawed AI models purely through unsupervised algorithms. Without requiring a single human to verify the output. Not one.

18:19So what does this all mean? If you are tracking the trajectory of artificial intelligence, you need to understand that the single largest bottleneck in the industry right now, the astronomical cost and time required for human verification is actively being solved. It's a major turning point. We are rapidly moving past the area where AI requires constant human handholding to grade its work. U.S. proves that an AI ecosystem can accurately and dynamically self-police its own reasoning at an elite academic scale, completely unsupervised. The industry is undergoing a fundamental shift in how we execute test time scaling.

18:54Developers no longer need to fear the noisy toddlers in the room. We now possess the mathematical frameworks to seamlessly filter them out and elevate the domain experts, regardless of what niche question is being asked. I want to leave you with a final thought that really stretches the implications of what we've unpacked today. Oh, this is the big one. We just established that mathematical frameworks like FUSE allow AI systems to perfectly score, rank, and filter their own unseen outputs without needing human answer keys. Yes. What happens when developers leverage this exact unsupervised mechanism to continuously generate pristine, synthetic training data?

19:32It becomes a feedback loop. Right. If an AI ecosystem can perfectly grade millions of its own complex thoughts, keeping only the absolute brilliance and discarding the hallucinations, could this fully unsupervised self-grading loop be the exact spark required to train subsequent AI models that drastically surpass human intelligence? It's very possible. Are we witnessing the exact transition point where our historical role permanently shifts from being the teachers of AI strictly to the observers? The implications for self-improving systems are profound. The student has essentially reverse-engineered the grading rubric.

20:07Thank you for joining us for this deep dive into FUSE and the future of unsupervised AI. Keep questioning the consensus, and we'll see you next time.

From the publisher

This paper introduces Fully Unsupervised Score Ensembling (FUSE), a novel framework designed to improve the accuracy of large language model (LLM) outputs without requiring human-labeled data. By aggregating scores from multiple imperfect verifiers, FUSE identifies the most reliable responses during the inference process, a technique known as test-time scaling. The method addresses the limitations of traditional ensembling by mathematically adjusting for statistical dependencies between verifiers that typically hinder unsupervised performance. Experimental results demonstrate that FUSE frequently matches or exceeds the performance of semi-supervised models that have access to ground truth labels. This effectiveness is validated across diverse benchmarks, ranging from academic datasets like MMLU to highly difficult math and logic exams. Ultimately, FUSE offers a scalable, cost-effective solution for filtering synthetic data and enhancing model reliability in complex reasoning tasks.

More from Best AI papers explained

All 475 episodes
FUSE: Ensembling Verifiers with Zero Labeled DataBest AI papers explained · 20 min
Listen in VO