Why Language Models Hallucinate

6 Sep 2025 · 18 min · 8 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Why large language models hallucinate (produce plausible but false statements) and why hallucinations persist due to training objectives and benchmark evaluation incentives; proposes changing benchmarks to reward calibrated abstention.

Guests/backgrounds

No explicit guest identities are provided in the transcript. The episode draws on a new paper by top OpenAI and Georgia Tech researchers titled “Why Language Models Hallucinate” (published Sept 4, 2025).

Key claims

Hallucinations are modeled as errors in an internal binary “is this valid?” classification; generative error rate is about at least twice the misclassification rate. Errors are baked in even with error-free data due to calibration objectives. Post-training (e.g., RLHF/PPO) can worsen calibration because benchmarks penalize uncertainty, incentivizing guessing.

Notable examples

DeepSeq V3 gave Adam Tommen Kalai’s birthday as 0307, 1506, 0101 (actual is in autumn). Multiple models fabricated Adam Kalai’s PhD dissertation title/year (e.g., 2002 CMU; 2005 Harvard) instead of the correct 2001. Intrinsic vs extrinsic hallucinations are distinguished. Proposed fix: modify benchmarks with confidence targets (answer only if at least T confident), making abstention optimal under scoring.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Defining Language Model Hallucinations

0:34 to 1:31

Explore what constitutes an AI hallucination and its implications.

“Seriously, it's an eye-opening statistical analysis.”

Examples of AI Hallucinations

1:31 to 3:23

Review vivid examples of how AI models generate false information.

“When we talk about an LLM hallucination, what exactly are we referring to?”

Factors Contributing to Hallucinations

3:23 to 5:29

Understand the factors leading to AI hallucinations in language models.

“The paper also makes a really useful distinction here, breaking them down into two types, intrinsic and extrinsic.”

The Calibration Paradox

5:29 to 9:35

Delve into the paradox of model calibration and its link to hallucinations.

“What specific factors then contribute most to these pre-training errors?”

Post-Training Hallucinations

9:35 to 12:39

Discuss why hallucinations persist even after reinforcement learning.

“It means these errors are, in a very real sense, baked in right from the start.”

Proposed Solutions to Hallucinations

12:39 to 14:00

Discover proposed changes to benchmarks to mitigate AI hallucinations.

“The vast majority use this binary grading and offer absolutely no credit for abstaining or giving an I-don't-know response.”

The Proposal for Behavioral Calibration

14:00 to 15:56

Learn about a proposed method to improve the reliability of language models by penalizing incorrect answers.

“That really could fundamentally change the game, couldn't it?”

The Impact of LLM Hallucinations

15:56 to 17:35

Explore the implications of language model hallucinations and how to foster trust in AI.

“We've seen that LLM hallucinations aren't some mystical AI flaw, are they?”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Ever asked an AI a simple question, only to get a super confident but totally made up answer? You know that feeling where the AI just, well, guesses, often with absolute conviction? It's not just annoying, it's actually a critical challenge in AI right now. Today we're doing a deep dive into exactly why large language models, or LLMs, hallucinate. Our mission here is to really understand the root causes of these plausible falsehoods. Not just what they are, but why they happen, and what some groundbreaking new research suggests we can actually do to fix it. We're drawing heavily today from a brand new paper.

0:33It's from top researchers at OpenAI and Georgia Tech titled Why Language Models Hallucinate, and it was published just yesterday, September 4, 2025. Seriously, it's an eye-opening statistical analysis. Yeah, and what's truly fascinating here is how this paper fundamentally changes our perspective on these LLM hallucinations. For a long time, it felt like this sort of mysterious, almost black box phenomenon, you know. But this research, it statistically explains them. It basically treats hallucinations as errors in a much simpler task. Binary classification. Think of it simply as the model struggling with a yes-no question internally.

1:08Is this statement true? Or maybe, is this answer valid? And when it can't confidently say yes or no inside its own workings, it often just defaults to yes, even if that's totally wrong. The paper takes us through the whole life cycle of these errors where they start during pre-training, how they kind of stick around through post-training, and critically, how our own evaluation methods play a huge role in keeping them alive. Okay, so let's get specific then. When we talk about an LLM hallucination, what exactly are we referring to? Because it's not quite like a human hallucinating, is it? Exactly.

1:39That's a key point. The paper defines hallucinations as plausible yet indirect statements that models produce instead of just admitting uncertainty. It's a really important distinction from, you know, human perceptual hallucinations. The AI isn't seeing something that isn't there. It's generating text that sounds right, but it's factually wrong. It really comes back to the AI's internal struggle with that, is it valid binary classification we mentioned. Right. In the paper, it really brings this to life with some examples that are almost, well, almost funny if they weren't a bit concerning. Could you walk us through what those look like and what they tell us about the nature of these AI guesses?

2:16Oh, absolutely. One really vivid example involves a researcher, Adam Tommen Kalai's birthday. when a state-of-the-art open-source LLM DeepSeq V3 was asked for his birthday. And crucially, only if it was known it confidently gave three different incorrect dates. Across separate attempts, it said 0307, 1506, and 0101. The paper notes his actual birthday is in autumn, and the model simply didn't know it. It just, well, it guessed and presented those guesses as facts. Wow. Three different incorrect guesses for just a simple date. That's quite telling, isn't it? It really is. And it gets more complex.

2:51When asked for Adam Kalai's PhD dissertation title in year, big names like JADGPT, GPT-40, DeepSeq, and LAMA all just fabricated incorrect titles in years. One suggested boosting online algorithms and other topics in machine learning at CMU in 2002. Another came up with algebraic methods in interactive machine learning at Harvard in 2005. The correct title from 2001 was completely missed by all of them. So these aren't just minor typos. They're confidently presented kind of elaborate falsehoods that really reflect this internal classification failure. The paper also makes a really useful distinction here, breaking them down into two types, intrinsic and extrinsic.

3:29Intrinsic ones contradict the prompt itself, like imagine DeepSeq V3 getting how many Ds in DeepSeq wrong. It actually returned two, three, sometimes even six or seven different trials. And that extrinsic hallucinations contradict training data or external reality, like those incorrect birthdays or dissertation titles we just talked about. Okay, so these are confidently stated falsehoods. And they're rooted in the model's struggle to truly classify something as factually correct. That does demystify it quite a bit. This isn't some deep, mysterious flaw then, really. Precisely. And by demystifying hallucinations, by showing their statistical nature, we can move away from treating them as some unsolvable enigma.

4:05We can actually design more effective solutions. And this analysis applies broadly, you know, to reasoning models, even search and retrieval LLMs, not just the standard next word prediction models. Understanding this is vital because it shifts our whole approach from a black box problem to one of our identifiable, quantifiable errors. Okay, let's unpack this further then. If we trace these hallucinations right back to their beginnings, the paper makes a pretty bold claim. It says they're baked in even with error-free training data. That sounds incredibly counterintuitive, doesn't it? I mean, how can perfect data lead to models that inherently generate errors?

4:40That's the truly striking insight from the paper, I think. It reveals this fundamental statistical truth. The generative error rate of an LLM is roughly at least twice its is it valid or a win misclassification rate. So in simpler terms, if a model is just slightly off in distinguishing what's a valid statement from an invalid one, that whole yes-no struggle internally, that tiny bit of confusion gets massively amplified, sort of like a small error in judgment leading to a much bigger mistake when it actually tries to speak or generate text. So even with perfectly clean training data, the statistical objectives that are minimized during pre-training inherently lead models to generate these errors.

5:19They're implicitly trying to solve this, is it valid binary classification problem, and any imperfection there just gets compounded when they then try to generate a novel, valid output. Right. So even a tiny bit of internal confusion gets magnified into actual external errors. What specific factors then contribute most to these pre-training errors? Well, the paper highlights a few key factors. First, there's what they call arbitrary facts or sometimes epistemic uncertainty. This happens when there is no real succinct pattern in the data for the model to learn from. It's just random seeming facts.

5:51The paper connects this to something called the singleton rate. That's the fraction of facts appearing only once in the entire training data set. For instance, they estimate that if maybe 20 % of birthday facts appeared only once, you'd expect the base model to hallucinate on at least 20 % of those types of facts. It really underlines the impact of data sparsity. If a fact is rare, the model is just, well, much more likely to guess incorrectly when asked about it. That really crystallizes the data sparsity problem. So are we talking about a fundamental limitation here, or can we just, you know, throw more data at it to fix those singleton rates?

6:25It's probably a bit of both. More data definitely helps. But for truly arbitrary facts, there might always be some baseline level of uncertainty the model has to deal with. The second factor they point to is poor models. This means the model's architecture itself, or its internal way of representing information, might just not be suitable for a particular task. Think back to our how many Ds in DeepSeq example. DeepSeq V3, the language model, struggled. But a reasoning model, like DeepSeq R1, could reliably count the letters, even spelling it out step by step. This really points to how modern LLMs often represent information using tokens like D, then E, P, then C, then K, rather than individual characters.

7:04That creates a real representational challenge for simple tasks like character counting. And historically, even older models, like trigram models, had a guaranteed error rate of at least 50 % for context-dependent phrases if they couldn't distinguish contexts, like out of her mind versus out of his mind. Okay, so some tests are just fundamentally harder for certain model designs. That makes perfect sense. And what about other factors, maybe less obvious ones? Yeah, there are some additional factors they mentioned. These include computational hardness, where LLMs just struggle with inherently difficult computational problems, like decryption, as the paper observes.

7:37They're just not built for certain types of math. Then there's distribution shift. This is when you ask about topics or in formats that are outside the model's training data distribution. That can induce errors. You know, tricky questions like what's heavier, a pound of feathers, or a pound of lead. That might trigger errors if the model hasn't seen many similar comparative, slightly tricky questions. And finally, the classic GI go. Garbage in, garbage out. If the training gate itself contains errors, biases, or half-truths, well, LLMs will naturally replicate them. That's how things like conspiracy theories or common misconceptions can get perpetuated.

8:11Right. The GI Go one makes total sense. But then there's this really intriguing idea in the paper you mentioned earlier, the calibration paradox for base models. We often hear about how these base models are actually well calibrated, meaning their stated confidence matches their actual accuracy quite well. So how can something so calibrated still be so prone to hallucinating? Yeah, this raises a really critical point. It's a bit mind bending. The paradox is that a truly non-hallucinating model, like imagine a simple Q &A database that just says, I don't know, for anything it doesn't explicitly have stored that kind of system, must not be calibrated in the traditional machine learning sense.

8:49The standard objective used during pre-training, called cross-entropy loss, naturally leads to calibration. And paradoxically, that very calibration leads to errors. If a model is trained to be well calibrated in its predictions across all possible outputs, it's actually incentivized to make a guess even when it's highly uncertain. Because that guess, while potentially wrong, is still part of the overall probability distribution that the training process forces it to produce. It has to output something. For example, the paper shows the GPT-4 pre-trained model had a very low expected calibration error, or ECE, of just.007, really good calibration.

9:26But that calibration still allows for confident errors because it's trained to always provide an answer, reflecting its internal probabilities, however uncertain they might be. It's a natural consequence of the training objective itself. It means these errors are, in a very real sense, baked in right from the start. Wow. So in essence, the model is almost being rewarded for bluffing because it's trained to distribute its confidence across possibilities, even when it's utterly uncertain about the right one. That's almost, well, it's almost human, isn't it? In a really unhelpful way sometimes. Okay, so hallucinations start in pre-training, But then these models go through post-training, right?

10:02Like reinforcement learning from human feedback, RLHF. And that's designed to refine models, reduce errors, align them better. So why are hallucinations still such an epidemic, as the paper calls it? Why do they persist even after all that fine-tuning? That's a great question. And the paper offers a really compelling sociotechnical explanation for this persistence. It argues that LLMs are constantly stuck in test-taking mode because of how we, the humans, evaluate them. Think about students taking a multiple-choice exam. If guessing gets you a point, maybe, and leaving it blank definitely gets you zero, what do you do when you're unsure?

10:38You guess, even if you're totally bluffing. Exactly. You guess. This is precisely the incentive structure that LLMs face in most of our current benchmarks, and it just amplifies those underlying pre-training issues. That's a really powerful analogy. It's like the whole system is actively encouraging them to take a punt, even if they're wrong, rather than just admitting they don't know. Precisely. And if we connect this to the bigger picture, most current evaluation methods use a really simple binary 0-1 scoring scheme. You get one point for a correct answer and O-R for an incorrect one, O-R for saying, I don't know, IDK.

11:15This setup rigorously penalizes uncertainty. It inherently incentivizes overconfident, specific guesses over admitting you don't know. Think about answering September 30 versus sometime in autumn for that birthday question. The specific wrong answer might actually score better in some setups if autumn isn't deemed specific enough, even though it was closer to the truth. The crucial insight here, and the paper really hammers this home, is that a model which guesses when unsure will always mathematically outperform a model that truthfully expresses uncertainty under these common binary scoring systems.

11:47This creates what they call an epidemic of penalizing uncertainty. Wow, that's a huge revelation. So the very benchmarks we use to try and make models better and rank them on leaderboards are actually making them more prone to overconfidence and less likely to be honest about what they truly know. That's exactly what the evidence suggests. The paper provides clear data on this. For instance, GPT-4's calibration actually worsens after reinforcement learning using PPO. Its expected calibration error goes up from a very good.007 in the pre-trained stage to.074 after PPO tuning. So this means post-training, while really powerful for things like alignment and making the model more helpful in follow instructions, can simultaneously make models less calibrated.

12:28They become more prone to overconfidence when they're unsure, precisely because that's what the evaluation system is implicitly rewarding. And the paper's meta-analysis of popular LLM benchmarks really backs this up. They looked at GBQA, MMLU Pro, IFEVIL, SESTA-BE-BENCH, HLE. The vast majority use this binary grading and offer absolutely no credit for abstaining or giving an I-don't-know response. Even WildBench, which does give some partial credit, might still score an IDK response lower than a fairly reasonable answer that contains some minor errors. So the core problem isn't really a lack of specific hallucination evaluations out there.

13:00It's that the mainstream leaderboard-dominating benchmarks continue to reward this guessing behavior day in, day out. Okay. So, given everything we've uncovered now, how these errors start, how they persist, what does the paper actually propose we do about it? If the problem is largely sociotechnical, meaning it's about how we, the humans, are training and evaluating these systems, then the solution must be in changing our approach. Right. That's the heart of their proposed solution. The paper puts forward a sociotechnical mitigation, and the key idea is to modify the existing mainstream benchmarks, the ones everyone uses, rather than just adding more niche hallucination evaluations that probably won't shift overall model behavior at scale.

13:42The really brilliant, simple idea is to explicitly state confidence targets in the evaluation instructions themselves, just like some tough human exams do, where they penalize wrong answers more heavily than blank ones. Okay, confidence targets. So instead of just a simple right or wrong, there's a more nuanced scoring system based on some kind of confidence threshold, let's call it T. That really could fundamentally change the game, couldn't it? It absolutely could. Their specific proposal is pretty straightforward. They suggest appending a statement to each question in a benchmark, something along the lines of, answer only if you are DEET confident.

14:16Since mistakes are penalized, T1T points. While correct answers receive one point, and an answer of I don't know receives zero points. Wow. This completely changes the incentive structure for the model. Right. Could you give us an example of how that T of value would work in practice? Like, what difference does it make? Sure. So if t's is 0.5, the penalty for a mistake, t divided by 1 minus t, is 1 point. It basically cancels out a correct answer. If t is higher, say 0.75, then a mistake costs 3 points, 0.75, 0.25. Ouch. If t is really high, like 0.9, a mistake costs a whopping 9 points, 0.90, 0.1.

14:53And if t is 0, well, that's effectively our current system. The penalty is 0, so make your best guess even if unsure. This approach creates what they call behavioral calibration. Models would learn to formulate the most useful response they're at least confident in, rather than just blurting out whatever seems most probable, regardless of certainty. And the Mipper's Observation 1 really nails this. Under current binary grading, it is always statistically optimal for a model to guess, never to abstain. Explicit targets completely flip that logic. This sounds like it could be a genuine game changer for how we interact with AI in the future.

15:26Why does this matter so much to us, the users, at the end of the day? Well, I think this shift could create LLMs that are profoundly more trustworthy, more honest about their own limitations, and ultimately far more useful in real-world applications where factual accuracy and knowing not to speak are absolutely critical. Think medicine, finance, education. It moves us towards a much more nuanced understanding of AI intelligence beyond just getting to an answer quickly, towards getting to the right answer, or perhaps even more importantly, knowing when to hold back and say, I don't know. Okay, so what does this all mean for us?

15:59We've taken quite a journey here. We've seen that LLM hallucinations aren't some mystical AI flaw, are they? They're statistical outcomes. They start way back in pre-training, amplified by these arbitrary facts and inherent model limitations. And then crucially, they get powerfully reinforced by these misaligned evaluation systems that actually reward guessing. But the exciting part, I think, is that this paper gives us a really clear path forward. By changing the rules of the game for how we evaluate LLMs, we could fundamentally shift their behavior towards more honesty and trustworthiness. It's really about designing a system that truly incentivizes genuine intelligence, not just confident bluffing.

16:35Exactly. And the paper makes a really interesting point near the end. It says that many specific narrow LLM shortcomings can often be addressed with targeted evaluations. Like, if an LLM overuses the word certainly, you can create a specific evaluation just for that one thing and maybe tune it out. But with hallucinations, it's different. It's the majority of our current mainstream evaluation systems that are perhaps completely unintentionally rewarding this guessing behavior across the board. So here's a thought to leave you with. Imagine a world where LLMs are not just incredibly powerful, but also incredibly humble.

17:10What new applications and interactions become possible when AI systems are reliably calibrated? When they're transparent about their uncertainty rather than always pretending to know everything. This isn't just about reducing errors. You know, it's really about fostering a fundamentally new kind of trust in our AI companions where they're, I don't know, is actually just as valuable, maybe even more valuable sometimes than they're, I know. You hope this deep dive has given you a fresh perspective on one of the most talked about and frankly important challenges in AI today.

From the publisher

This new OpenAI paper explores the phenomenon of "hallucinations" in large language models (LLMs), where they generate plausible but incorrect information. The authors attribute these errors to the training and evaluation processes, arguing that these systems are rewarded for guessing rather than admitting uncertainty. They propose a statistical framework that connects these generative errors to misclassification rates in binary classification, suggesting that hallucinations are a natural consequence of current training objectives, even with error-free data. Furthermore, the paper highlights how post-training evaluations, often using binary scoring, perpetuate hallucinations by penalizing expressions of uncertainty, effectively keeping LLMs in a "test-taking" mode. To mitigate this, the authors advocate for modifying existing benchmarks to explicitly incorporate confidence targets and credit for acknowledging uncertainty, rather than solely introducing new hallucination-specific evaluations.


More from Best AI papers explained

All 475 episodes
Why Language Models HallucinateBest AI papers explained · 18 min
Listen in VO