Position: Uncertainty Quantification in LLMs is Just Unsupervised Clustering

7 Jul 2026 · 22 min · 12 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

The episode argues that uncertainty quantification (UQ) in large language models is mis-specified: mainstream UQ methods effectively measure internal consistency via unsupervised clustering of sampled generations, not external factual correctness. It claims this causes “confident hallucinations” and a deceptive safety signal for operators.

Guest backgrounds

No guests are named in the transcript; it’s a host-led discussion of an ICML 2026 position paper.

Key claims

UQ is a category error; confidence reflects clustering tightness in latent space, not epistemic uncertainty about reality. UQ is hyperparameter-sensitive (e.g., temperature changes confidence without changing knowledge), runs an internal evaluation cycle, and lacks ground truth anchors.

Notable examples

cockpit/map analogy; medical diagnostics and autonomous systems; “polygraph” analogy; confident hallucinations where wrong answers are stable across samples and get ~99% certainty.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Understanding Uncertainty Quantification in AI

1:25 to 3:20

Discover the role of uncertainty quantification in AI and its critical flaws.

“We are looking at a brand new position paper, and this was actually accepted at the 2026 International Conference on Machine Learning or ICML.”

The Mechanics of UQ and its Pitfalls

3:20 to 5:12

Examine how UQ operates and the significant category errors it introduces.

“Like we think we are measuring the AI's grasp of reality, but we are actually measuring something entirely different.”

Confident Hallucinations and Their Dangers

5:12 to 7:09

Learn about 'confident hallucinations' and their implications for AI reliability.

“External correctness, on the other hand, means that specific pattern actually maps to the objective reality of the world.”

Pathologies of Current UQ Deployments

7:09 to 9:43

Explore the three major pathologies affecting current UQ systems in AI.

“But if you deploy it wrapped in a UQ system that outputs these authoritative sounding confidence scores, human operators naturally lower their guard.”

Proxy Metrics and Their Limitations

9:43 to 12:34

Understand the problems with using proxy metrics for evaluating AI accuracy.

“The uncertainty score is measuring the mechanical tuning of the model, not the actual difficulty or factuality of the question.”

The Need for a Paradigm Shift in UQ

12:34 to 13:20

Discuss the urgent need for a reevaluation of AI safety protocols and UQ methods.

“The proxy is entirely decoupled from external reality.”

Proposed Reforms for UQ in AI

13:20 to 14:01

Learn about proposed frameworks for improving the measurement of AI success.

“Dijin Chen and the research team clearly titled this a position paper because they are advocating for a massive pivot in how the industry approaches this problem.”

Evaluating Success in LLMs

14:01 to 14:58

Learn about the need for new evaluation metrics in LLMs.

“The first fix focuses entirely on changing how we measure success.”

Structural Mechanism Changes for Uncertainty

14:59 to 15:50

Explore the need for fundamental changes in LLM architecture.

“Implementing structural mechanism changes to allow for native uncertainty.”

Integrating Objective Truth in Confidence

15:51 to 17:55

Understand the importance of aligning AI confidence with real-world facts.

“Does the model need a separate loss function during pre-training that specifically penalizes overconfidence on sparse data?”
Show all 12 chapters

Pathologies of Current UQ Practices

17:56 to 20:03

Discover the flaws in current uncertainty quantification practices.

“To guide that back to the listener's perspective for a second, the ultimate goal here is to make AI self-awareness actually reliable.”

The Challenge of Native Uncertainty in AI

20:04 to 21:23

Contemplate the theoretical limits of uncertainty in LLMs.

“What's fascinating here is how this whole breakdown reminds us how vital human critical thinking is.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00So imagine you are sitting in the cockpit of a commercial airliner. Okay. And you're flying through just an absolute nightmare of a storm. Like you can't see anything out the window. Yeah. It is pitch black. Rain is hammering the glass. The turbulence is severe. Sounds terrifying. Right. So you turn to your co-pilot and you ask, are you sure this is the right way to the runway? Yeah. And they look you dead in the eye and say, absolutely, I am 100 % certain. Oh, wow. Okay. So you'd probably feel a lot better. Exactly. You exhale, you trust them. But what if you found out after landing that they were only so confident because they checked their own like horribly outdated map three times in a row rather than actually looking at the radar?

0:44Oh, I mean, that blind confidence is suddenly completely terrifying. Yeah. And today we are exploring that exact scenario, but inside the neural networks of large language models. Yeah, because that feeling of safety in the cockpit, it just evaporates instantly when you realize the co-pilot's confidence has literally no connection to external reality. And, you know, in high stakes, AI deployments think medical diagnostics or autonomous systems, legal analysis, an overly confident model that is disconnected from factual reality is, frankly, infinitely more dangerous than a model that just admits it doesn't know the answer.

1:21So true. Welcome to this deep dive, everyone. We are thrilled you're joining us today because our mission today is to explore a fascinating and honestly pretty alarming blind spot in how we evaluate AI safety. It really is a big one. It is. OK, let's unpack this. We are looking at a brand new position paper, and this was actually accepted at the 2026 International Conference on Machine Learning or ICML. Yeah, very prestigious track. Extremely. It's authored by Tijin Chen, Long Kao Da, Zhao Yu, and Huawei. And the title gets right to the point. It's called Position. Uncertainty quantification in LLMs is just unsupervised clustering.

1:59Yeah, and what this research team is doing here is, well, they're pulling back the curtain on the primary safety nets that the tech industry currently relies on. They're showing us that the mathematical foundations of these nets might just be a total illusion. I mean, they are fundamentally questioning the entire way we measure an AI's self-awareness. So let's actually define that safety net for the listener before we go ahead and tear it down. No idea. The industry term is uncertainty quantification or UQ. So like if we deploy an LLM in a hospital routing system, we desperately need to know when the model is just guessing.

2:35Yes, absolutely. So UQ is supposed to be the gauge that flashes red and says, you know, my confidence score is low here. Route this to a human doctor. Right. It's the fail safe. Exactly. But the actual mechanism for generating that score is exactly what this paper is challenging. Yeah, because in a traditional machine learning pipeline, uncertainty quantification is treated as the gold standard for risk mitigation. Right. The assumption is that the UQ system accurately reflects the model's epistemic uncertainty. Epistemic meaning? Meaning it calculates how much the model actually knows about a given topic based on its training data.

3:09We expect the UQ score to act as this reliable bridge between the model's internal probability calculations and actual physical factual reality. But Chen and the research team, they argue that this entire field of UQ actually suffers from a massive category error. Yes. Like we think we are measuring the AI's grasp of reality, but we are actually measuring something entirely different. A category error is really the perfect framing for this because the researchers demonstrate that mainstream UQ methods for language models, so things like self-reflection mechanisms or semantic entropy calculations, they're essentially just unsupervised clustering algorithms operating in disguise.

3:49Wait, okay, let's break down the mechanics of that because when we talk about unsupervised clustering in this context, we aren't talking about something simple. No, not at all. We are talking about mapping outputs in high dimensional vector space. Right. So the model generates a batch of possible answers to your prompt, right? Right. Then it converts those answers into mathematical representations and then measures the geometric distance between them. That is the literal mechanism, yeah. Yeah. Because when you query an LLM, it doesn't just have one definitive answer waiting in a database somewhere.

4:21Right, it's not a search engine. Exactly. It generates multiple possible pathways by predicting sequences of tokens based on probabilities. So standard UQ methods, they sample several of these generated pathways. Okay. Then using those clustering algorithms you mentioned, it looks at how tightly grouped those pathways are in its latent space. I see. So if the generated answers are highly semantically similar, meaning if they cluster really tightly together mathematically, the UQ system registers low entropy, it looks at that tight cluster and outputs a really high confidence score. Okay, wow. Well, this means the AI isn't checking its answer against an external database of facts at all.

5:01Nope. It is exclusively checking its answers against its other generated answers. Yes, exactly. It's quantifying the internal consistency of its generations rather than their external correctness. And this is the crux of the category error right here, because internal consistency simply means the model's output pathways are converging on the same linguistic pattern repeatedly. Right. External correctness, on the other hand, means that specific pattern actually maps to the objective reality of the world. Which is what we actually care about. Right. We are deploying these models under the desperate assumption that we are measuring external correctness, but the architecture of the UQ systems only allows them to measure internal consistency.

5:43Here's where it gets really interesting. Are you saying the AI is basically living in its own echo chamber? If it repeats a lie enough times, it just assumes it's the truth. Yes. Oh, exactly. And this exposes the most dangerous phenomenon highlighted in the paper, which is the failure to detect what they call confident hallucinations. Confident hallucinations. Yeah. So a standard hallucination is when a model generates an output that is factually incorrect just because it traversed a weird probability path. Right. It just makes something up. Exactly. But a confident hallucination happens when the model generates an incorrect answer, And its internal mathematical pathways are highly stable and tightly clustered around that error.

6:24So the model is dead wrong. Dead wrong. But because it is consistently dead wrong across multiple generated samples, the unsupervised clustering algorithm just slaps a 99 % certainty label on it. Exactly. Which brings us right back to your cockpit analogy. Oh, right. The co-pilot is staring at three identical, deeply flawed maps. And because the maps all agree with one another, which is internal consistency, right? The co-pilot declares absolute certainty about the flight path. Wow. The system is structurally blind to the factual reality outside the window. That creates a deeply deceptive sense of safety for anyone operating the monitor.

7:02Because if you deploy an LLM without any safety net at all, you treat it like an intern, you know? Yeah. You double check every single line of code or text at output. Right, you're paranoid. Exactly. But if you deploy it wrapped in a UQ system that outputs these authoritative sounding confidence scores, human operators naturally lower their guard. They do. The human critical thinking basically gets outsourced to a fundamentally broken gauge. And that deceptive safety is arguably a greater systemic risk than the underlying hallucinations themselves. Yeah, I can see that. And what makes this research so compelling is that it doesn't just, like, point out this philosophical flaw.

7:39It breaks down the specific mechanical failures that happen when you build a safety system this way. Right. The authors actually identify three critical pathologies that plague current UQ deployments. Okay, let's dive into those. Because if the foundation is just unsupervised clustering of generated tokens, the first place it seems likely to break down is in the generation rules themselves. Hold on. We know that engineers constantly tweak hyperparameter dials like temperature or top sampling to change how an LLM outputs text. So if the UQ score relies entirely on how the text groups together, tweaking those dials must throw the whole confidence score into chaos.

8:16Right. That mechanical reality is exactly what the authors identify as the first major breakdown. They call it a hyperparameter sensitivity crisis. The sensitivity crisis. Yeah. Yeah. Let's look at the temperature parameter as an example. So temperature dictates the strictness of the probability distribution for the next token. A low temperature forces the model to pick only the most mathematically likely words, which leads to very rigid, repetitive text. Makes sense. But a high temperature flattens that distribution. It allows the model to pick slightly less probable words, creating more variance.

8:51Meaning if I ask the model a medical question at a low temperature, the generated samples will naturally look nearly identical. Yes. So the clustering algorithm will see a tight group and output a high confidence score. Exactly. But if I ask the exact same model, the exact same question, and simply bump the temperature up slightly, the generated paths will spread out. The clustering algorithm will see a wider spread and suddenly declare the model as highly uncertain. You've got it. The uncertainty score changes dramatically without the actual foundational knowledge of the model changing at all.

9:25That is wild. It is. The UQ system is unbelievably fragile because it is tethered to these generation parameters. You literally cannot deploy a medical diagnostic AI where the safety net's reliability fluctuates based on a tiny decimal tweak in a configuration file. No, you definitely can't. The uncertainty score is measuring the mechanical tuning of the model, not the actual difficulty or factuality of the question. So if the hyperparameter dials completely dictate the uncertainty score, then the model isn't actually evaluating its own knowledge at all. Nope. It's just grading its own math based on a rubric that someone arbitrarily adjusted.

10:05Exactly. Which I think shifts us directly into the second breakdown they talk about. Yeah. So the authors refer to this as the internal evaluation cycle. Because the UQ methods are just unsupervised clustering algorithms looking at generated logits, the system completely conflates mathematical stability with truth. The model is evaluating its own uncertainty by looking at its own outputs. It literally cannot distinguish between a stable truth and a highly stable lie. It operates a lot like a polygraph test, doesn't it? That's a great way to think about it. Because a polygraph doesn't actually measure whether a statement is historically factual.

10:40It just measures the physiological stability of the person speaking. Right. Heart rate, sweating, that kind of thing. Right. So if a pathological liar has rehearsed a fabricated story so deeply that their heart rate and breathing remain perfectly stable, the polygraph registers the statement as true. Yes. The UQ system is just a polygraph measuring the model's mathematical stability. It's completely ignoring the factual reality of the statement. And that conflation of stability and truth is a massive structural trap. I mean, we are using epistemology terms, words like knowledge and uncertainty, to describe what is purely a density estimation in latent space.

11:19But since the model is completely trapped in this closed-loop evaluation cycle, it has no way to verify anything outside of itself. It is navigating entirely by its own internal compass. And that leads directly to the third pathology identified in the paper, which is a fundamental lack of ground truth. Current UQ methods cannot touch objective reality. Because they have no external anchor, the systems are forced to rely entirely on what we call proxy metrics to evaluate uncertainty. And a proxy metric is basically what you use when you can't measure the actual thing you care about. Like if we care about truth, but the model can only measure semantic similarity between generated samples, it just uses similarity as a proxy for truth.

12:05Yes. We are essentially using a compass near a giant electromagnet. Oh, I like that. Yeah. So the compass needle points very confidently in a specific direction, and the internal mechanics of the compass are working perfectly, right? The needle is responding to the magnetic field. Right. It's doing its job. But we use the needle's stability as a proxy metric for our heading. And because the compass is entirely disconnected from true magnetic north, that proxy metric is actively leading us off a cliff. Wow. The proxy is entirely decoupled from external reality. So what does this all mean? If we are evaluating uncertainty using unstable proxies instead of ground truth, aren't we basically building a house on quicksand?

12:49If we connect this to the bigger picture, yeah. As long as we evaluate uncertainty this way, we can never truly trust an LLM in a high-stakes scenario. We just can't. Because we are building critical infrastructure on type of this, we're integrating LLMs into legal research, financial modeling, software engineering, all while relying on proxy metrics that actively break down under pressure. It paints a fairly grim picture of the current state of AI deployment, I won't lie. It really demands an immediate reevaluation of our safety protocols. Absolutely. But fortunately, the paper doesn't just map out the quicksand.

13:22They don't leave us hanging. Right. Dijin Chen and the research team clearly titled this a position paper because they are advocating for a massive pivot in how the industry approaches this problem. Right. So if the current UQ paradigm is fundamentally broken, how did the authors suggest we rebuild the architecture to actually anchor these models to reality? Well, the roadmap they propose requires a complete paradigm shift. I mean, this isn't about just adjusting the temperature dials or tweaking the clustering thresholds a little bit. We're past that. Way past that. They outline three major structural fixes to escape this internal evaluation cycle.

14:00Okay, what's the first one? The first fix focuses entirely on changing how we measure success. We have to adopt completely new and better evaluation metrics. Basically, stop rewarding the polygraph for measuring a stable heart rate. Exactly. The industry needs evaluation frameworks that actively test a model's ability to recognize external factual reality. We have to start using proxy metrics like internal consistency to grade how well a UQ system is working. Because if an evaluation benchmark only tests whether a model generates similar tokens across multiple runs, it is a useless benchmark for real-world safety.

14:37Right. It doesn't tell us if it's true. Exactly. We need metrics centered on information retrieval and actual ground truth validation. But changing the testing metrics is really only half the battle, right? Because if the underlying architecture of the LLM still only knows how to predict the next mathematically likely token, it will still just be a highly sophisticated guessing machine. Which brings us to the second and arguably most mechanically complex step on their roadmap. Go ahead. Implementing structural mechanism changes to allow for native uncertainty. Native uncertainty. Because right now, uncertainty quantification is largely an afterthought.

15:15Oh, completely. We build these massive, multi-billion parameter auto-regressive models designed specifically to generate confident-sounding text. Then, like, after the fact, we retrofit them with clustering algorithms or self-reflection prompts to try and guess when the model is unsure. Yeah. The safety net is basically bolted onto the side. And the authors argue that this retrofitting is the root of the problem. We need structural mechanism changes at the base level of the architecture. Wow. The concept of, I don't know, has to be a mathematically fundamental state within the neural network itself.

15:48But what does that actually look like at the architectural level? Like, are they suggesting we change the transformer architecture itself? Does the model need a separate loss function during pre-training that specifically penalizes overconfidence on sparse data? Yes, they are advocating for exploring those exact kinds of deep architectural shifts. Really? Yeah, it involves looking into epistemic neural networks or Bayesian approaches where the model doesn't just output a single token probability, but actively computes a distribution of its own ignorance based on its training data coverage. Wait, a distribution of its own ignorance?

16:23Yes. So instead of generating five answers and checking if they match, the AI's core engine would inherently possess a built-in doubt vector. Oh, why? If it encounters a query that falls entirely outside its high-density training clusters, the internal attention mechanisms would flag that sparsity natively before generation even begins. That would be a massive leap forward. The model would literally pause its generation pipeline and output a mathematically grounded, I lack, the internal architecture to connect these specific data points rather than hallucinating a confident bridge between them.

16:58Exactly. It transforms the model from a blind token predictor into a system with actual boundaries. Which is huge. It is. And that leads to the third and most vital fix proposed by the researchers, which is anchoring all verification in objective truth. Right. Because if we successfully build a model with native uncertainty, we still have to ensure that its internal doubt actually aligns with the physical world. And the researchers are drawing a hard line here. This is my favorite part of the paper, honestly. It's just so elegant. They demand that a model's confidence must serve as a reliable proxy for reality itself, not just a reflection of its internal math or training distribution.

17:37Right. Anchoring an objective truth means that our verification processes cannot rely on the LLM's latent space. We have to integrate external data sets, fact-checking APIs, and sophisticated retrieval augmented generation, or R pipelines directly into the uncertainty calculation. it's brilliant because it forces the model to actually look outside. Right. To guide that back to the listener's perspective for a second, the ultimate goal here is to make AI self-awareness actually reliable. Yes. The confidence score cannot be finalized until the model is forced to look outside its own echo chamber.

18:12Exactly. The system has to cross-reference its proposed answer with an immutable external knowledge graph. It is entirely about aligning the machine's confidence with the actual factual state of the world so that when the AI says, I don't know, we can trust it, and when the system says, I am 100 % certain, it is because the generated output has been explicitly mapped to a verifiable ground truth, not just because the clustering algorithm found a dense pocket of similar tokens. This really changes the entire conversation around AI safety. Let's recap what we've mapped out for everyone listening today.

18:44Sure. We started by examining uncertainty quantification, which is the primary gauge we use to determine if a large language model is safe to deploy in high-stakes environments. Right. And we discovered through this ICML 2026 position paper by Chen and their team that this gauge is currently suffering from a massive category error. Because we operate under the assumption that UQ measures external correctness, but mechanically it is basically just unsupervised clustering that measures internal mathematical consistency. Exactly. And this architectural flaw blinds the system to confident hallucinations, creating a deceptive sense of safety for human operators.

19:23Yeah. We explored how this manifests in three distinct pathologies. A hypoparameter sensitivity crisis where tiny tweaks completely alter the safety score, an internal evaluation cycle that conflates mathematical stability with the factual truth, and a reliance on unstable proxy metrics due to a total lack of ground truth. We are navigating using a compass that is actively drawn to a magnet rather than true north. Perfectly said. But the authors offer a structural roadmap out. By adopting rigorous new evaluation metrics, engineering structural changes into the neural networks to allow for native uncertainty, and explicitly anchoring all confidence verification in objective, external truth, the industry can begin to build models that actually know what they don't know.

20:08What's fascinating here is how this whole breakdown reminds us how vital human critical thinking is. Oh, absolutely. Especially when we are confronted with absolute certainty from a machine. Every time an LLM outputs a definitive answer, we have to remember the mechanics under the hood. Is this confidence based on external reality or is it just a tight cluster of mathematically similar tokens? It highlights the danger of confusing fluency with accuracy. It really does. And honestly, this raises an important question, one that sits at the very frontier of machine learning. What's that? Well, large language models are, at their absolute core, autoregressive engines structurally designed to predict the next most likely sequence of tokens based on statistical distribution.

20:50Right. That's their fundamental nature. Right. So if that is their nature, is it even theoretically possible to build true epistemological native uncertainty into them? Or will we always, on some level, just be fighting against an architecture that is perfectly mathematically optimized to convince itself and us that it is right? Wow. That is the defining question for the next decade of artificial intelligence. Can an engine built entirely to predict words ever truly comprehend the boundaries of its own truth? It's a tough one. It is. And that is definitely something for you, the listener, to chew on and explore as this technology continues to integrate into our daily lives.

21:30Absolutely. We want to thank you for joining us on this deep dive. Keep pushing back against the illusion of confidence. Keep demanding to see how the mathematical maps are drawn. And most importantly, stay curious. Because the next time a complex system tells you what is absolutely certain, you'll know exactly what mechanical realities to question before you trust it to fly the plane.

From the publisher

This research paper argues that current methods for Uncertainty Quantification (UQ) in large language models are fundamentally flawed because they function as unsupervised clustering rather than measures of factual accuracy. The authors contend that these techniques merely track internal consistency, which fails to identify confident hallucinations where a model is consistently wrong. This reliance on internal stability creates a false sense of security and suffers from issues like hyperparameter sensitivity and a lack of objective ground truth. To fix these problems, the paper proposes a paradigm shift that anchors model confidence in external reality and objective verification. Ultimately, the researchers provide a roadmap for the community to develop more reliable metrics for ensuring AI safety in high-stakes environments.

More from Best AI papers explained

All 475 episodes
Position: Uncertainty Quantification in LLMs is Just Unsupervised ClusteringBest AI papers explained · 22 min
Listen in VO