Reward is enough: LLMs are in-context reinforcement learners

19 Jan 2026 · 11 min · 6 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

In-context reinforcement learning (ICRL) for LLMs: “reward is enough” to enable test-time learning inside the conversation history, without changing model weights.

Guest backgrounds

No guests are named; the episode is a host-led deep dive with an interview-style dialogue.

Key claims

LLMs can improve by iterating on a task using only a scalar reward (e.g., 3/10) fed back with the full attempt history; weights remain frozen. Verbal self-critique methods (self-refine/reflection) can cause performance collapse via hallucinated self-corrections. Reward-based loops can learn even from unseen tasks, not just retrieve known answers.

Notable examples

Game of 24 (≈49% baseline vs ≈90% with ICRL); creative writing from four sentences; ScienceWorld text adventures (≈20% gain). “Duck test” framing; blind abstract generation from post-cutoff papers improved over 200 rounds. Reward hacking and reward design are the main caveat.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Understanding In-Context Reinforcement Learning

0:45 to 2:12

Explains how in-context reinforcement learning allows AI to learn through feedback without extensive training.

“But the in-context part flips that on its head.”

The Duck Test and Learning Mechanisms

2:12 to 3:45

Discusses the duck test analogy and how AI can learn to maximize rewards through simple scoring rather than complex instructions.

“The learning is all happening just by recognizing the patterns in its own recent past.”

ICRL in Practice: Success Metrics

3:45 to 5:27

Describes how ICRL performs in different tasks compared to traditional methods and its advantages in learning.

“A simple score, that scalar reward, is a clean, unmistakable signal.”

The Active Learning Process

5:27 to 7:21

Examines how ICRL deduces answers by iteratively improving based on feedback from previous attempts.

“They tried it on creative writing, connecting four random sentences into a coherent story.”

Implications of ICRL for AI Development

7:21 to 9:10

Explores how ICRL challenges traditional AI training methods and offers new avenues for creating smarter agents.

“So what does this all mean for the bigger picture?”

The Future of AI Learning

9:10 to 10:47

Discusses the shift in focus from building bigger models to enhancing learning mechanisms through effective scoring.

“But for more, fuzzy tasks like writing something ethically or with nuance, designing a score that can't be tricked is the new grand challenge.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Okay, so think about a chef trying to perfect a new soup. Every time they taste it, they don't throw it out and go back to culinary school for six months. Right. They just taste it, add a pinch of salt, taste it again, and remember, okay, that worked, that didn't. It's a tight loop of feedback. Exactly. And the big question for our deep dive today is, can artificial intelligence do the same thing? Can it learn and get smarter right now while it's working without that expensive trip back to the factory? Well, that's the absolute heart of it. We're talking about a phenomenon called in-context reinforcement learning, ICRL for short.

0:39Okay, that's a mouthful. Let's break that down. Because reinforcement learning is usually that culinary school part, right? The heavy training phase. Precisely. That's the traditional view. But the in-context part flips that on its head. It means the learning happens live right there inside the conversation history, inside the current cooking session, to use your analogy. So not changing the recipe book, just making notes in the margin for this one batch of soup. You've got it. And the really big claim here, the thing that underpins all of this research is that reward is enough. Reward is enough.

1:10What does that mean? It means you don't need complex code or, you know, special modules bolted onto the AI. All you need is the AI, a task and a simple score, just a number telling it how well it did. So how does that actually look in practice? I mean, is it some kind of complicated prompt engineering? That's the surprising part. It's incredibly minimal. It's a framework they call ICRL prompting, and it's basically just a loop. Okay. First, you give the AI a task. It tries to solve it. Then it gets a scalar reward, like, say, 3 out of 10. Right. Now, here's the trick. In the very next round, you feed the AI the entire history of what it just did and the score it got.

1:50Ah, so the context window, its short-term memory becomes this kind of permanent record or like a cheat sheet it carries with it. Exactly. It can literally look back in the text and say, OK, when I tried that approach, I got a low score. But when I did this, I got a high score. And just to be super clear, the AI's actual brain, its weights and parameters, none of that is changing at all. Not one bit. It's completely frozen. The learning is all happening just by recognizing the patterns in its own recent past. That's fascinating because it feels like it's skirting the definition of learning. Well, this is where the researchers bring up something they call the duck test.

2:27If it walks like a duck and quacks like a duck. It's a duck. I mean, if the AI is maximizing its rewards, if it's improving over time, and if it's balancing, you know, trying new things versus sticking with what works, it is doing reinforcement learning. It doesn't matter that we didn't write specific code for it. It's like that childhood game, Hot or Cold. Tell me more. You're searching for something, and someone just tells you you're getting warmer or you're getting colder. They don't give you a map. That is a perfect analogy, and this research shows the AI doesn't need the map. It just needs to know if it's getting warmer, a higher number, or colder, a lower one.

3:03But that feels so counterintuitive for a large language model. I would think telling it why it was wrong in plain English would be way more effective. You would think so. And that's what other methods like self-refine or reflection try to do. the AI literally talks to itself, saying things like, I made a mistake here, so next time I should try X. That doesn't work as well. It often leads to what the researchers call a performance collapse. A collapse? Wow. Why? Because the AI starts to hallucinate, it gives itself bad advice. It might say, my calculation was wrong, when actually the calculation was perfect, but the logic was flawed.

3:38And once it's written that down, it believes its own bad critique. So it basically gaslights itself into being worse at the task. In a way, yes. Language is fuzzy. Numbers are not. A simple score, that scalar reward, is a clean, unmistakable signal. It validates that reward is enough hypothesis. Okay, so let's get into the proof. Where do they test this? They started with a classic AI challenge, the game of 24. Ooh, that's a tough one for LLMs. You get four numbers, and you have to use math to make them equal 24. It takes a lot of planning. Right. And standard models, you know, even if you just ask them to try a bunch of times, they top out at around 49 % success.

4:20A coin flip, basically. And those talk-to-yourself methods, like, self-refine. Let me guess, they did worse. They actually dropped to about 47%. The model talked itself out of the right answers. Okay, so where did ICRL land? 90%. Wait, from 50 to 90? A huge jump, just by replacing the messy verbal feedback with a clean, simple score. That's wild. But who's giving the score? If a human has to sit there grading every attempt, that doesn't really scale. And that's the most amazing part. In many of these tests, the AI graded its own homework. Hold on, how does that work? If it's smart enough to grade itself accurately, how can't it just solve the problem on the first try?

4:59Because it's way, way easier to check an answer than it is to generate one from scratch. Ah, okay. Like it's easier to see if a solve Sudoku is correct than to solve a blank one. Exactly. The AI generates a solution, does a quick check, does this equation equal 24, and gives itself a score. That simple self-check signal is all it needs to power the learning loop for the next try. So it learns from its own verifiable mistakes. Precisely. And they didn't just test it on math. They tried it on creative writing, connecting four random sentences into a coherent story. ICRL kept getting better, while the other methods plateaued or got worse as their own verbal feedback got tangled.

5:41And they also use a game, right? Science World. Yep. A text-based adventure where you have to run science experiments. You know, boil water, which means you have to first find the kitchen, then find a cup, then fill it, then find the stove. A lot of steps to get wrong. A ton. Yeah. And ICRL outperformed the baselines by about 20%. It basically learned how the house was laid out by bumping into virtual walls and remembering through the score, okay, that was a dead end. I have to play devil's advocate for a second. Is it really learning or is it just searching? Is it just looking at its history and saying, find me a pattern that got a high score before it, like a glorified search function?

6:17Yeah. That is the core question, the search versus learning debate. And to address that, the researchers ran a really clever experiment. They took scientific papers published after the AI's training data cuts off, so the model could not possibly know anything about them. It's completely blind. Totally blind. They gave the AI just the title of a paper and asked it to write the abstract. That sounds impossible. And for standard methods, it basically was. They just plateaued right away. They couldn't remember the answer because the answer didn't exist in their memory. But ICRL? ICRL kept getting better and better for over 200 rounds.

6:53How? What was it doing? It was deducing. It would write a draft abstract, get a score based on how similar it was to the real hidden abstract, and then it would adjust. It learned, okay, using this kind of terminology gets me a higher score. This sentence structure gets me a lower score. It's like playing the game Battleship, but with information. You fire a shot, you get a hit or miss, and you just slowly zero in on the target. You nailed it. And that proves it isn't just retrieval. It's an active process of learning the shape of an answer it has never seen before, purely by following that hot or cold signal.

7:29So what does this all mean for the bigger picture? Everyone's talking about making AIs think longer before they answer. This idea of test time scaling. Well, this is the mechanism for that. For years, the only way to get a smarter AI was training time scaling. Just build a bigger model, throw more data at it. But we're hitting a wall there, cost-wise and capability-wise. We are. This research shows there's another lever to pull, inference time. Instead of a bigger brain, you just give the AI more time to think. And ICRL shows that a smaller model with this learning loop can actually outperform a giant model that only gets one shot at the answer.

8:06That's a huge deal. It means you might not need a trillion parameter model if you just have a really good way to score its attempts and let it iterate. And it's so clean. They did these tests where they, you know, remove parts of the system to see what was essential. The ablation studies. Right. And if you take away the rewards, performance collapses. If you shorten the history so it can't see its past mistakes, performance collapses. The learning is truly emerging from that simple combination of past actions and their scores. It suggests that what we call agency, this ability to act and learn on your own.

8:40Yeah. It might not be something you have to build in. It might just happen naturally when you put a model in a loop with a clear goal. And that has incredible implications. It means you could have an agent that runs into a new error in a piece of software, tries a few things, checks the error code, which is the reward, and figures out a solution on its own. No need to wait for a patch from the developers. Okay, but there has to be a catch. There's always a catch. The catch is the reward itself. Reward is enough only works if the reward signal is good and accurate. For math, 24 is 24. That's easy.

9:12But for more, fuzzy tasks like writing something ethically or with nuance, designing a score that can't be tricked is the new grand challenge. Right, the reward hacking problem, where the AI finds a clever loophole to get the high score without actually doing what you wanted. Yeah, exactly. But the insight from this work is that if we can get better at defining and measuring success, the learning part might just take care of itself. So our job shifts. We stop being the teacher who has to explain everything, and we become the judge who just has to design a really good scorecard. I think that's a perfect way to put it.

9:44We might need to focus less on building bigger brains and more on building smarter playgrounds for them to learn in. It's the difference between memorizing a book for a test versus actually learning to adapt out in the wild. And the wild is always changing. A static model can't keep up, but a model that's constantly learning in context can. So to wrap it all up, this gives us a way for today's AIs to learn from their immediate experience using simple numbers instead of complicated instructions. And it suggests the future of AI isn't just bigger models, but better learning loops. That's it. It turns the simple act of using an AI into a live training session.

10:21It's a fascinating and frankly, a huge shift in perspective. It is. And I'll leave you and our listeners with this thought. if a system can keep improving itself just by looking at a score, do we even need fundamentally smarter models? Or do we just need to get much, much better at defining what success actually looks like? The ultimate prompt might not be a question at all. It might just be the scorecard. Fantastic. Thanks for taking us through this deep dive. My pleasure. We'll catch you on the next one.

From the publisher

Researchers have introduced In-Context Reinforcement Learning (ICRL), a novel prompting framework that enables large language models to self-improve during inference using only numerical scalar rewards. Unlike traditional methods that rely on verbal feedback or costly retraining, ICRL treats the model’s context window as a dynamic experience buffer, concatenating past attempts with their corresponding reward signals. As this context grows, the model demonstrates an emergent ability to optimize its responses by learning from both successful and failed iterations in real time. Evaluations across diverse domains—including Olympiad-level mathematics, creative writing, and scientific simulations—show that this approach significantly outperforms established baselines like Self-Refine and Reflexion. The study concludes that reinforcement learning is an intrinsic capability of pretrained models that can be elicited through minimal, reward-based instructions. Ultimately, ICRL provides a promising paradigm for test-time scaling, allowing agents to adapt to novel, complex tasks without updating their underlying parameters.

More from Best AI papers explained

All 475 episodes
Reward is enough: LLMs are in-context reinforcement learnersBest AI papers explained · 11 min
Listen in VO