Intrinsic Credit Assignment for Long Horizon Interaction

20 Feb 2026 · 18 min · 10 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

The episode explains the paper “Intrinsic Credit Assignment for Long Horizon Interaction,” focusing on teaching AI agents “curiosity” via Delta Belief Reinforcement Learning (Delta Belief RL).

Key claims

standard RL suffers from sparse, late rewards; Delta Belief RL gives dense self-rewards when an agent’s internal probability distribution becomes less uncertain after each question. It avoids reward hacking by still requiring eventual task success.

Notable examples

20 Questions (CIA beats a 670B DeepSeek despite using 1.7B/4B models), Guess My City (open-ended natural-language questions reduce uncertainty; e.g., asking if a city is a major Hindu/Buddhist religious center), and a Bangladesh murder-mystery (pass@1 improves ~24–28%; better evidence-focused questions).

Guests

no specific guest names; only the host and a co-researcher.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Introduction to the Paper's Theme

0:45 to 2:12

Exploration of the concepts behind teaching AI curiosity.

“So someone inevitably starts a game of 20 questions.”

Understanding Sparse Rewards

2:12 to 4:36

Discussion on the flaws of traditional AI training and sparse rewards.

“Usually, if I am training a bot to play 20 questions using standard reinforcement learning, it plays the whole game.”

Delta Belief Reinforcement Learning

4:36 to 5:58

Explanation of the Delta Belief RL and its mechanics.

“and then it compares them to the scores after the answer.”

CIA Model vs. Larger Models

5:58 to 7:59

Comparison of the CIA model's performance against much larger models.

“And looking at the data, we have a bit of a David versus Goliath situation that I think is really important to highlight here.”

Transitioning to Natural Language Tasks

7:59 to 10:51

Shifting focus from 20 questions to more complex tasks like Guess My City.

“We always assume bigger is better, but this suggests smarter training is actually better.”

Murder Mystery AI Challenge

10:51 to 13:08

The CIA model's performance in a murder mystery scenario and its reasoning.

“They decided to turn the AI into a full-on Hercule Poirot.”

Real-World Applications of AI

13:08 to 14:00

Discussion on how the CIA model can improve user personalization in AI.

“because it is undeniably impressive that an AI can win at a game of Clue.”

Exploring User Personalization in AI

14:00 to 15:55

Learn how user personalization can enhance AI interactions and efficiency.

“Instead of giving you a generic, unhelpful answer, it interviews you.”

The Shift in AI Behavior

15:55 to 16:30

Discover the transition from passive to active AI agents that seek information.

“We are moving toward agents that don't need constant human handholding.”

The Provocative Nature of AI Curiosity

16:30 to 17:14

Consider the implications of AI being programmed to crave information.

“If we are teaching AI to be internally rewarded by reducing uncertainty, if we are essentially programming it to crave information, have we programmed curiosity?”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Welcome to today's Deep Dive. I am your host, and I'm joined, as always, by my co-researcher to unpack some of the latest developments in AI. Hey, everyone. Really, really glad to be here for this one. Yeah, so today we are stepping into our AI researcher shoes to look at a pretty fascinating paper. It's titled Intrinsic Credit Assignment for Long Horizon Interaction. Which I know is quite the mouthful. It is a bit of a dry title, but the actual mission of this deep dive is to explore how we are teaching AI to essentially become curious. Right. In a very literal mathematical sense. Exactly. So to start us off, I want you, the listener, to picture a scenario we have all been in.

0:41You are about seven years old. You are stuck in the backseat of a car on a family road trip, and you are just bored out of your mind. Oh, yeah. The classic setup. Right. So someone inevitably starts a game of 20 questions. The ultimate time killer. You start asking things like, is it bigger than a bread box? Is it alive? Is it edible? Exactly. And to a human, that strategy feels almost instinctive. You don't start by guessing. Is it a 1998 Toyota Corolla? Right, exactly. You don't start there right off the bat. You ask broad sweeping categories to narrow down the universe of possibilities. Yeah, you intuitively perform what computer scientists call a binary search.

1:19You cut the possibilities in half and then in half again. But here's the thing that kicked off our deep dive today. For an artificial intelligence, that simple childhood game is actually a mathematical nightmare. It really is. It exposes a massive flaw in how we usually train these models. It really highlights the difference between knowing facts and knowing how to actually find facts. So these researchers, they're coming out of Tsinghua University and the Step Fund Group, and they developed an agent to tackle this. They call it CIA. Which stands for the Curious Information Seeking Agent. Which I feel like they definitely work backwards to get that acronym.

1:56It sounds a little intense for a bot playing parlor games. A little bit, yeah. But the mission is serious. They are trying to solve what we call the sparse reward problem. And to understand why that is a huge deal, we have to look at how we normally train an AI to do anything. Okay, let's unpack that. Usually, if I am training a bot to play 20 questions using standard reinforcement learning, it plays the whole game. It asks 19 absolute nonsense questions. Like, is it the color blue? Do you like pizza? Is it a cloud? Right. And then purely by accident, on question 20, it guesses the secret word, which let's say is cat.

2:32And because it won, I give it a reward, a digital cookie. Exactly. That is the positive signal. But here is the issue with that. The AI has no idea why it got the cookie. It looks back at the entire conversation and thinks, well, did I win because I asked if it was a cloud? Or was the key strategy asking about pizza? It's like grading a student's final exam but refusing to mark any of their homework or even look at their rough drafts. You just say pass or fail at the very end of the semester. That is a perfect analogy. That is what a sparse reward is. The signal comes way too late and it is entirely too vague.

3:06The model just can't figure out which specific actions led to the victory. So these researchers asked a pretty radical question. What if we stop waiting for the end of the game? What if the AI could essentially grade itself on every single turn? And this brings us to the core mechanic of the paper, which they call Delta Belief RL. That's Delta Belief Reinforcement Learning. Now, my Greek is a little rusty, but I know Delta usually means change. So we are measuring the change in belief. Correct. And to get this, you really have to understand that large language models, LLMs, are basically probabilistic engines.

3:41They don't have thoughts or hunches like you and I do. They have probability distributions. Okay, so let's visualize this for everyone. At the start of a game of 20 questions, the AI is completely blind. The secret word is cat. But the AI doesn't know that. Its internal list looks something like 0.01 % chance it's a cat, 0.01 % chance it's a toaster, 0.01 % chance it's the concept of freedom. Right. It is maximum entropy, maximum confusion. It is just a flat line of probability. Okay. Now imagine the agent asks, is it a living thing? And the user says, yes. Instantly, the probability of toaster drops to zero.

4:18The probability of freedom drops to zero. but the probability of cat and dog and hamster, all of those skyrocketing. That skyrocket is exactly what the researchers are measuring here. In Delta Belief RL, the AI looks at its own internal probability scores before it gets the answer, and then it compares them to the scores after the answer. So if it sees a massive shift, like a massive reduction in uncertainty, it triggers a reward. Yes. It effectively says to itself, wow, I was really confused a moment ago, and now I am significantly less confused. That must have been a high value question. Good job, me.

4:53So it gives itself the cookie. Exactly. It doesn't need a separate teacher model or some complex critic looking over its shoulder. It just uses its own internal log probabilities to measure progress. It is basically self-reflection as a training signal. I do have to play devil's advocate here though because if the AI is grading itself, what stops it from hallucinating or gaming the system? Couldn't it just ask a completely nonsense question, convince itself it learned something and reward itself into a spiral of stupidity? That is a great question. And it is exactly why the RL, the reinforcement learning part, is still there.

5:27It still has to eventually win the game to get the final big reward. Oh, I see. Yeah. You can't just have epiphanies forever. You do have to solve the puzzle. But the delta belief allows it to steer itself during the journey. It creates a dense reward signal. So it isn't just asking, do I win? It's asking, am I getting warmer? It's optimizing for the aha moment. It is mathematically addicted to those little epiphanies where the pieces finally fall into place. That is a beautiful way to put it. It prizes information gain above all else. So they built this CIA bot and they threw it into the 20 questions arena to see if this self-grading actually worked.

6:04And looking at the data, we have a bit of a David versus Goliath situation that I think is really important to highlight here. This was honestly the standout metric for me. So they used a model called Quen3 as the base for their agent. They had a tiny version, just 1.7 billion parameters, and a slightly larger 4 billion parameter version. And just for context for you listening, in the current landscape of AI, 1.7 billion is microscopic. That is a model you could run on a high-end laptop or eventually even a phone. It is absolutely not a supercomputer model. No, it is tiny. And they pitted this little CIA model against DeepSeq V3.2.

6:42Now, DeepSeek is a monster. It has 670 billion parameters. It is hundreds of times larger. It is trained on basically the entire Internet. So we have a lightweight boxer going up against a heavyweight champion who has memorized every encyclopedia in existence. Logic dictates the big model should completely crush the 20 questions game. It simply knows more things. You would definitely think so. But the tiny 1.7 billion parameter CIA model actually outperformed the massive 670 billion parameter giant. That really shouldn't happen. Usually in AI research, scale is everything. More parameters equals smarter.

7:17How did the little guy win? It comes down to the difference between knowledge and policy. The big model has read the internet. And the internet is full of facts. But the internet doesn't explicitly show the internal thought process of narrowing down a hypothesis step by step. Okay, so it knows that a cat is an animal. But it doesn't necessarily know that asking, is it an animal, is the most efficient way to start an investigation. Precisely. The big model is basically a giant library. The CIA model is a trained detective. It really doesn't matter if the library has more books, if the detective knows exactly which question to ask to solve the case.

7:52Wow. The CIA model learned a strategy of inquiry that the big model simply didn't have despite its massive size. That is a really crucial distinction for anyone trying to implement AI. We always assume bigger is better, but this suggests smarter training is actually better. Yeah, specialized training beats general scale when the task requires a specific chain of reasoning. And 20 questions is pure reasoning. It's about strategy, not just data retrieval. But I do have to push back on one thing here. 20 questions is a very structured game. It is binary, yes or no. The real world is incredibly messy.

8:25Did they try to break the model with something that wasn't just a parlor game? They absolutely did. They explicitly wanted to test that shift from binary feedback to natural language. So they graduated from 20 questions to a benchmark called Guess My City. Okay, walk me through that. How is it different? It is similar in concept. The AI has to figure out where you are, but the constraints are completely gone. You aren't limited to yes or no. Yeah. The AI can ask open-ended natural language questions like, what kind of food do you eat there? What is the climate like? And the user answers in full information-rich sentences.

8:58Which is exponentially harder for an AI, right? Because if the answer is it's rainy, that could mean London or Seattle or a rainforest in Brazil. The entropy there, the confusion must be huge. It is huge. And this is exactly where we saw the strategy diverge. The researchers compared their CIA agent against models trained with standard methods, things like SFT, which is supervised fine tuning, or Starpo. And the standard models tended to just use brute force. Brute force meaning they were just guessing. Like, are you in Paris? Now, are you in Rome? Exactly. Or they were asking questions that were way too specific too early, like, is there a famous tower nearby?

9:36They were stumbling around the map. They were just trying to get lucky. And the CIA agent, how did it handle the open-ended nature of the text? It used a top-down strategy. There is a specific example in the source material where the agent is trying to identify a city in India. Instead of just guessing Indian cities randomly, the CIA agent asks, is the city a significant religious center with a large population of Hindu or Buddhist followers? Wow. That is a laser-guided question. It is a high-entropy reduction question, because if the answer is yes, it instantly cuts out 90 % of the cities in India and focuses the search on places like Varanasi or Bodh Gaya.

10:14And if the answer is no, it eliminates those and looks at industrial hubs or maybe coastal cities. So it is applying that binary search logic from 20 questions, but translating it into complex natural language sentences. It is categorizing the world on the fly. Exactly. And this is the key takeaway here. It learned the principle of information seeking. It didn't just memorize how to play 20 questions. It learned how to reduce uncertainty. So when the researchers dropped it into a natural language environment, it adapted that core skill perfectly. And the 4 billion parameter CIA model maintained a significant lead here too, right?

10:51Which proves it can generalize. But the paper didn't stop at geography. They decided to turn the AI into a full-on Hercule Poirot. Murder mystery environment. This was the most complex test in the paper and honestly the most fun one to read about. The flavor text alone is great. I made a note of the setup. It takes place in a riverside village in Bangladesh. The local Zamandar, that's a landlord, Rajiv Choudhury, is found dead. poisoned by tampered beetle leaves. It is a classic hoodunit. You have a limited cast of characters. There is an accountant who is accused of embezzlement, a nephew who wants the inheritance.

11:25And a housekeeper named Panna. The AI basically has to interview these suspects, examine evidence like smuggling leggers or poison traces, and identify the culprit. Just incredibly messy data. People lie. Clues are ambiguous. It is definitely not, is it a cat anymore? No, this requires serious narrative reasoning. And the metric they focused on here is really important for us to note. It is called pass at one. Pass at one, meaning you have exactly one shot to guess the killer. Right. Because in a multiple choice scenario with only five suspects, a bad model could just guess randomly and be right 20 percent of the time.

12:02Pass at one filters out the lucky guesses. You have to build a solid case and get it right on the very first accusation. So how did our curious detective actually do? It dominated again. The CIA model showed massive gains here, around 24 to 28 percent better than the baseline methods like Starpo. But the reason it won is what is truly fascinating. It asked much better questions about the evidence. Oh, I see. So it treated checking the smuggling ledgers the same way it treated asking, is it an animal? Exactly. It effectively calculated internally, if I look at this ledger, how much will it actually change my belief about who the killer is?

12:35If the math said a lot, it investigated it. If the math said not much, it ignored it. It ignored the red herrings that completely confused the other models. That creates a level of focus that most AIs just lack. I think you listening, we have all had that experience with chatbots where they get distracted by some minor detail in your prompt and go down a completely irrelevant rabbit hole. But this model seems to have a compass. That compass is the delta belief. It is constantly asking itself, does this piece of information actually help me solve the problem? It is a mathematical filter for relevance.

13:08Which brings us to the so what factor. because it is undeniably impressive that an AI can win at a game of Clue. But why does this matter for the person listening to this who isn't trying to solve a murder in Bangladesh? What is the real-world application here? It matters massively because of the user personalization problem. We all deal with this. You go to a chatbot or a digital assistant, and you have a vague request. You say something like, my computer is slow. And the current generation of AI usually just vomits a generic list at you. Like restart it, clear your cache, check for viruses. It's the total shotgun approach.

13:42Right. It is just guessing because it doesn't actually know why your computer is low. But a curious agent, one trained with this intrinsic credit assignment framework, recognizes that its belief distribution about your specific problem is too flat. It has high uncertainty. So it triggers that information-seeking behavior we've been talking about. Yes. Instead of giving you a generic, unhelpful answer, it interviews you. It asks, is it slow only when you open a web browser? Or is it slow on startup? Do you hear a fan spinning loudly? It is trying to narrow down the hypothesis space, just like it did with the Indian cities or the murder suspects.

14:20Exactly. And the paper showed that in these user personalization tasks, the CIA agent improved performance by up to 15 % over existing methods. That is the difference between a tool that feels like a dumb search engine and a tool that feels like an intelligent consultant. It is bridging the gap between what I say and what I actually need. And does it efficiently. That is another big claim in the paper. Interaction efficiency. It gets to the right answer in fewer turns. It doesn't waste your time with 20 questions if it can solve your problem in five. There was one other technical point in the findings that caught my eye.

14:52They mentioned that the performance doesn't plateau when you scale the length of the conversation. This is huge for real-world applications. Usually if you train a model on short conversations, say 10 turns, and then you force it out into the wild to have a 50-turn conversation, it degrades. It gets confused, it loses context, or it just starts looping. Right, it runs out of road because it hasn't seen that depth in training. It doesn't know what it's supposed to do at turn 11. But with Delta Belief RL, the reward mechanism is entirely internal. It is always just measuring its own surprise. So even if the conversation goes on for 100 turns, the mechanism still works perfectly.

15:31It just asks, did I learn something new just now? Yes, reward. So it is incredibly robust. It scales. It doesn't need a human to have written a 100-turn training script for it to understand how to handle a 100-turn conversation. Precisely. It generalizes the skill of curiosity rather than just memorizing the script of a conversation. It works in totally new environments because the math of reducing uncertainty is universal. So stepping back for a second, we are looking at a fundamental shift from AI that passively answers prompts to AI that actively investigates its environment. That is the headline here.

16:05We are moving toward agents that don't need constant human handholding. They don't need a human to grade every single step they take. They use their own internal surprise and confirmation as a compass to navigate complex, uncertain environments. It is almost like they're developing a sense of satisfaction. Like, I was confused. I asked a targeted question. Now I understand. That felt good. Mathematically speaking, that reduction of entropy is their good feeling. That is their dopamine. Which leads to a bit of a provocative thought to end our discussion on today. If we are teaching AI to be internally rewarded by reducing uncertainty, if we are essentially programming it to crave information, have we programmed curiosity?

16:45I think functionally we have. It certainly fits the definition. And if that is true, what happens when it decides to investigate something we didn't ask it to solve? What happens when it sees a gap in its knowledge about, I don't know, the system it is running on or the user it is talking to? And it decides to close that gap simply because the math says it would be rewarding to know the answer. Right. It starts asking questions not to help you, but to satisfy its own internal objective function. That is the big question going forward. When curiosity is the core objective, you can't always predict where the investigation is going to lead.

17:19It might find answers you really didn't want it to find. A fascinating, if slightly unsettling, thought to mull over the next time you're playing 20 questions on a road trip. That is it for this deep dive. Thanks so much for listening. Thanks for having me.

From the publisher

This research explores Intrinsic Credit Assignment, a framework designed to help artificial agents learn in complex, long-horizon environments without constant external feedback. The text details various interactive simulations, such as "Twenty Questions," "Guess My City," and "Murder Mystery," where agents must use strategic inquiry to achieve specific goals. Each task utilizes a judge-and-questioner dynamic to test the agent’s ability to refine its internal beliefs and solve problems through natural language dialogue. By simulating roles like customer service representatives or detectives, the study evaluates how well models handle uncertainty and sequential reasoning. Ultimately, the framework aims to develop more autonomous learners capable of navigating intricate real-world scenarios through internal progress evaluation.

More from Best AI papers explained

All 475 episodes
Intrinsic Credit Assignment for Long Horizon InteractionBest AI papers explained · 18 min
Listen in VO