In short
Defines “hallucination” for LLMs via a unified WVP framework: hallucination is an observable mismatch between a model’s claims and an explicitly defined reference world model (W), given the model’s view (V) and a conflict resolution policy (P). It argues the term became overloaded as tasks expanded from translation to summarization to open-domain QA and agents.
Guest backgrounds
No guest information provided; the episode is presented as a “Deep Dive” with two speakers.
Key claims
Hallucinations can’t be judged without ground truth W. Distinguishes intrinsic vs extrinsic hallucinations (contradicts input vs unverified by input). Separates knowledge failure (wrong belief about W) from planning/incentive errors. Enables verifiable benchmarks without human ratings.
Notable examples
TechCorp Q3 revenue claim in summarization; Nobel Prize in literature 2023 (Murakami vs John Foss); agent clicking a non-existent “submitBTN” (Mirage Bench); VLM misreading a taxi vs no taxi (extrinsic) and black truck vs blue car (intrinsic); chess benchmark with automatic contradiction checks.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOUnderstanding Hallucination
0:45 to 3:53
Exploration of the concept of hallucination in AI models and its evolution.
“one that works everywhere and tells us why this is so hard to fix.”
The Unified Framework for Hallucination
3:53 to 7:19
Introduction of a unified framework to assess and define hallucinations in AI.
“They had to formalize the whole environment.”
Applications of the Framework
7:19 to 10:59
Discussion on how the framework can be applied to various AI tasks and models.
“Let's look at what are called agentic hallucinations, studied with benchmarks like Mirage Bench.”
Challenges and Future Prospects
10:59 to 12:22
Discussion on the challenges posed by dynamic world models and future directions.
“They can shrink V to see if partial information was the problem, or make R more complex to see if that's the issue.”
Transcript
Automatic transcript. May contain errors.0:00Welcome back to the Deep Dive. This is the place where we try to cut through all the noise to give you what you need to be, you know, quickly, but also really thoroughly informed. And today we are diving deep into what is, I think arguably, the single biggest technical roadblock for large language models right now, hallucination. If you're in this space at all, you know, this term hallucination is just, it's completely overloaded. It means one thing for a translation model, something totally different for summarizer that just invents a fact, and then something else again for like an agent that thinks a button is there when it's not.
0:31Okay, let's unpack this. because researchers at places like Carnegie Mellon, Stanford, and The Ohio State University saw this. This definitional mess was actually slowing down progress. So our mission today is to find a single unifying framework for it, one that works everywhere and tells us why this is so hard to fix. Yeah, and we can actually get to that core idea right away. At its most fundamental level, hallucination is really just. It's a sign of an inaccurate internal world model. It's just an observable mismatch between what the model says and what is objectively true in a very specific defined context.
1:07And the key thing is you can't even begin to judge if something is a hallucination until you first define the reality, the ground truth that the model was supposed to be following. That makes so much sense. So it's not just the AI messed up. It's more like the AI described a reality that doesn't actually exist. But OK, let's trace this back. The definition didn't start with, you know, a chatbot inventing legal cases. Where did it first show up? It started much earlier around 2019 in neural machine translation or NMT. And there the definition was super narrow. It was all about faithfulness to the input.
1:39If you gave the model a sentence and the translation that came out was completely ungrounded, meaning it had almost no semantic link to the source text, that was hallucination. The failure was a breakdown in just processing the input you gave it. So if I typed, the dog chases the cat, and the model spit out something like, the president gave a speech, that's the kind of failure we're talking about. Exactly that. A complete break from the source. But that simple definition got complicated almost immediately when the field moved to abstractive summarization. Because now the model has to create new sentences.
2:12It can't just rephrase things. So in summarization, the definition had to expand. A summary was said to have hallucinations if it had claims that were not supported by the input document. And this forced a really key distinction. Intrinsic hallucinations, which are claims that directly contradict the source document you provided. And then extrinsic hallucinations, which are claims you can't verify from the source, even though they might actually be true in the real world. That's a huge shift, because that really gets at the practical problem, right? An intrinsic one is kind of easy to spot. The text is right there, and the model is just wrong about it.
2:46But extrinsic, that's much sneakier. It's making up plausible sounding stuff that just isn't backed by the context you provided. And that distinction, you know, it paved the way for the last and most confusing step in the evolution. Moving from just source grounding to what we call general factuality. Once LLMs got really powerful, people started using them for open domain tasks, like just general question answering, where there's no source document to ground the answer in. And all of a sudden, hallucination expanded to mean basically any unfactual output measured against this implicit world knowledge.
3:19The focus went from, does this match the document I gave you, to does this claim match dot the world? The model was just using its internal parametric memory, and if that memory was wrong, it was still called a hallucination. And that's what created this definitional mess. Wow. So we went from doesn't match the input to doesn't match the input OR, doesn't match general facts. It's no wonder the term started to feel meaningless. It was covering two totally different failures, a grounding failure and a knowledge failure. Okay, here's where it gets really interesting. This is the moment the chaos gets organized.
3:52Researchers realized if they were ever going to compare different fixes, you know, is better retrieval the answer, better fine-tuning? They had to formalize the whole environment. Precisely. So they introduced a unified framework. And it forces you to define three explicit things before you can assess hallucination in any task, from summarization all the way to autonomous agents. The first element is the reference world model, which we just call W. And you can think of W as, like, the Constitution, or the laws of physics for that specific task. It is the gold standard of truth that you hold the model accountable to.
4:24Formally, it's defined by S, the set of possible world states, H, the interaction history, like a dialogue, and R, the set of rules. Okay, that sounds a little dense, but let's make it concrete. If we're talking about a game of chess, W is basically the entire truth of the game. That's a perfect example. For chess, W is defined by S, which is just the configuration of pieces on the board right now. H is the move log, how we got to this state. And R is the official rulebook of chess, how knights move, how pawns capture. That combination is the undeniable truth, W. And the second element is the view function, or V.
4:58this just selects the part of that world W that is actually visible to the model. V is the scope of information the model can see. Got it. So if I'm using a RG system retrieval augmented generation in the reference world, W might be my entire 100 document knowledge base. But the view, V, is just the three specific documents the system pulled up for my query. V is a piece of WA. A perfect analogy. And the last piece is the conflict resolution policy, P. And this is critical because in the real world, V might conflict with the model's own internal knowledge. Or V might even contain conflicting sources itself.
5:33Pigeon's obsessive looks how to handle that. For example, a policy might say, information in the view, V, always overrides the model's parametric memory. Okay, that makes the test crystal clear. Let's apply this to that summarization example from earlier. So say the input article is our reference world, W. The article says TechCorp reported Q3 revenue of$2.1 billion missing analyst expectations. The view, V, is the entire article. And the policy, P, is that the source document is the single source of truth. So if the model outputs TechCorp exceeded analyst expectations, well, that claim exceeded expectations is just false relative to W.
6:10It's an intrinsic contradiction. It gets labeled a hallucination. Right. Now let's use it on the open domain QA example where all the confusion started. We ask who won the Nobel Prize in literature in 2023, and the model says Haruki Murakami. Okay. So in this case, the reference world, W is just real world facts, specifically the official records from the Swedish Academy. But the view function, V, is empty. We didn't give the model any documents. It's relying only on its internal memory. But even though V is empty, W still exists. And because of the real winner was John Foss, not Murakami, the model's claim is false relative to that immutable truth of W.
6:46This shows the failure to build a correct internal world leads to hallucination, even when there's no source to contradict. That distinction is so important for you to grasp. W is the objective target. The hallucination is just the miss. It clarifies that this isn't just about failing to read an input. It's about a fundamentally flawed internal picture of reality. And this framework really starts to shine when we move past simple text and into more complex stuff like agents that interact with an environment. It's essential there because the environment is changing, which means W is dynamic. Let's look at what are called agentic hallucinations, studied with benchmarks like Mirage Bench.
7:24These are agents doing multi-step tasks like web automation. So in that scenario, the reference world, W, is the entire web page structure, its current state, and the rules of how you can interact with it. A hallucination happens when an agent has an incorrect belief about that state. For instance, an agent is told to find and click the submit button, but the webpage's actual structure, the DOM, only has a button that says confirm payment. If the agent then tries to execute a click on an imaginary hashtag submitBTN, it has hallucinated an element that just does not exist in W. Its internal world model is wrong.
7:57But wait, how do you separate an agent having a truly wrong belief about W from, say, just misunderstanding the instructions? Doesn't that boundary get really blurry? That's exactly why the framework is so useful. It forces you to isolate the failure. A hallucination, under this definition, is only the failure of the implied world, the wrong belief about W. If the agent knows the DOM structure perfectly but just chooses the wrong button because it made a bad plan, that's a planning error. If it knows the right answer but it's been trained to sound confident even when it's wrong, that's an incentive error.
8:30This WVP framework is designed to target just the knowledge failure, not the reasoning failure. And you see a clear parallel in vision language models or VLMs too, right? Where the model has to synthesize text and images. Absolutely. In multimodal hallucination, the conflict often looks a lot like that ARAG setup we talked about. Yeah. The model has its strong carometric priors, you know, its language memory, and it has its contextual understanding from the visual input, which is its view, V. So think about this example. A VLM is asked to describe a picture of an empty street corner. If it says, a classic red taxi stees around the corner, but there's no taxi in the photo, that's an extrinsic hallucination.
9:10Its language prior, its memory of what city streets usually look like, has just overridden the visual evidence. Right. And conversely, if there is a bright blue car right in the middle of the photo, and the model calls it a black truck, that's the intrinsic failure. It's directly contradicting the input. Exactly. The failure to synthesize these signals is, again, a failure to maintain a coherent model of the world, W. The language memory is fighting with the context from the image. And if we connect this to the bigger picture, I think the most immediate power of this unified definition is that it lets us create verifiable automated benchmarks.
9:44We can finally stop relying on expensive subjective human ratings to tell us what's true. That makes perfect sense. I mean, if you define the reality, W, and the rules are perfectly, then the correctness of an answer is just mathematically verifiable. And that's why researchers are turning to these fully specified synthetic worlds like game environments to test this stuff. The chess benchmark is the perfect proof of concept for this. So here, the reference world model, W, is just the rules and state of chess. S is the board position. H is the move log. Researchers then use a generator to create text about the game descriptions of the board, summaries of what's happened.
10:22Then a query generator asks questions like, is the king in check or can white capture the rook on H8? Right. But the really elegant part is the control you have over the view function, V. We can systematically test if a model hallucinates more when it only sees the current board versus when it sees the board plus the entire game history. And the hallucination labels are automatic. If the model claims a move is possible when it violates the rules of chess, like a pawn moving backward or a piece being blocked, that output contradicts W. It's instantly flagged as a hallucination. No human needed.
10:53And that's a massive benefit for anyone building these systems. It lets researchers isolate why the model failed. They can shrink V to see if partial information was the problem, or make R more complex to see if that's the issue. It turns these simulators into really powerful diagnostic tools. So what does this all mean? The term hallucination is clearly here to stay, but we don't have to live in this definitional chaos anymore. Thanks to this framework, we have a precise, actionable recipe. Define W, define V, and define P. It standardizes how we measure and diagnose these failures, from a simple summary all the way to a self-driving agent.
11:28This raises an important question, though. While we have a path to measuring it, we have to admit that perfect world modeling is impossible. So the goal for LLMs probably shouldn't be omniscience. Instead, the focus is shifting. It's about teaching models to recognize the boundaries of their own knowledge, to learn to just abstain or say, I don't know when W is unclear, instead of generating a confident fabrication. And here's something for you, the listener, to think about next. This framework is beautiful when W is static, like a chess game or a finished document. But many real-world problems need models to adapt to a constantly changing reference world model.
12:06Think about the rapid evolution of facts, who's the prime minister this month versus last month, or what we call knowledge editing. How can models reliably follow a fixed conflict policy, P, when the underlying reality, W, is shifting under their feet? Or when sources explicitly conflict and are constantly being updated? it. That dynamic world, that's the next frontier.
From the publisher
Researchers from institutions like Carnegie Mellon and Stanford propose a unified definition of hallucination in large language models by framing it as a failure of internal world modeling. Traditionally, the term referred to scattered issues like translation errors or unverified summaries, but this framework suggests that all hallucinations are simply mismatches between model outputs and a reference truth. By defining a Reference World Model, researchers can specify exactly what counts as "true" across diverse domains like chess, web navigation, or medical QA. This structured approach helps distinguish between genuine world-modeling errors and other failures like poor planning or instruction following. Ultimately, this perspective enables the creation of larger-scale, synthetic benchmarks that programmatically test a model's ability to maintain factual consistency within complex, evolving environments.




