In short
Whether frontier LLMs can do in-context experiential learning—improving strategy over repeated interactions using ambiguous feedback—tested via a new product-recommendation benchmark (BLA).
Guests
No guest names or backgrounds are mentioned; it’s a solo “Deep Dive” discussion.
Key claims
SOTA models (GPT-4o, Gemini 2.5 Pro, Gemini Flash, Claude) show flat regret across 10 sequential episodes with the same customer, even when given perfect memory summaries and clearer numeric feedback. They also ask fewer questions over time, repeat loops, and overstate confidence (poor calibration). Humans can solve it with an “ideal learning trajectory.”
Notable examples
“Karen Thompson” persona; GPT-4o repeats the same questions; Gemini defaults to “free option” scripts; recommendations contradict prior user statements; human-guided run reaches zero regret while LLMs stay around 37.5 regret.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOIn-Context Experiential Learning Defined
0:45 to 2:36
Exploration of in-context experiential learning and its significance for AI.
“Can these frontier models look at a past mistake, get some ambiguous feedback, and then fundamentally change their strategy for the next time?”
Benchmark for Experiential Learning
2:36 to 4:25
Introduction to the Benchmark for Experiential Learning and Active Exploration (BLA) and its design.
“the Benchmark for Experiential Learning and Active Exploration, or BA.”
Agent Interaction with Customers
4:25 to 7:11
Discussion on how agents interact with customers and the challenges they face.
“Okay, let's nail down the structure of these sessions.”
Performance of State-of-the-Art Models
7:11 to 9:35
Examination of how top AI models perform in the BLA environment and their limitations.
“So they ran variants where they literally handed the models a perfect summary of the prior episodes, a curated memory of everything they'd learned.”
Failures of Learning and Adaptation
9:35 to 14:01
Analysis of the failures in learning and adaptation observed in AI models during the experiment.
“Even when the persona clearly prefers quality.”
Limitations of Current LLMs in Learning
14:01 to 15:19
Explore the shortcomings of LLMs in dynamic learning scenarios.
“Current SOTA LLMs are incredible at retrieval and one-shot prediction in a fully observed space.”
Trusting AI in High-Stakes Scenarios
15:19 to 15:29
Discuss the implications of AI's learning limitations in critical tasks.
“How can we ever truly trust it to handle complex, high-stakes, real-world tasks?”
Transcript
Automatic transcript. May contain errors.0:00Welcome back to the Deep Dive. So today we are really focusing on, well, what a lot of people are calling the final frontier for advanced AI agents. Right. It's this idea of moving beyond just, you know, encyclopedic knowledge and figuring out how to get these large language models to actually learn from personal lived experience. It's a huge difference. Like the difference between cramming for a test and actually, you know, working a job for 10 years. We've seen these incredible one-shot performances from the soda models when all the data is laid out perfectly. Like a static quiz. Exactly. They ace the quiz.
0:35But the real world is messy. It's dynamic, uncertain, and it demands you adapt over time. And that really sets up our core question today. Can these frontier models look at a past mistake, get some ambiguous feedback, and then fundamentally change their strategy for the next time? Or are they just stuck? Yeah. Stuck in this loop of solving every new situation like it's the very first time they've seen it. And this challenge, this is precisely what researchers are starting to call in-context experiential learning. And we really need to define that because it's the heart of the whole issue we're talking about today.
1:09Okay, let's do that. Let's unpack that right away. Forget the jargon for a second. What does that kind of learning actually look like for an AI agent? I mean, think of a really great salesperson. Yeah. Or a top strategist. It means the agent can reliably adapt and significantly improve its actions over many interactions, many episodes. Based on the feedback it's getting. Right. Based on the often messy real world feedback it's gotten before. It's the mechanism for self-correction over the long haul. So if I, say, walk into a store and a recommendation agent shows me five things I absolutely hate and I tell it why.
1:45I give it free-form text feedback like these are all too modern and way too expensive for me okay the next time I come in agent shouldn't just retrieve my old comments it should have used that history to build a more refined internal model of me it should change its whole approach exactly it should be measurably smarter in that second session than it was in the first and right now our current LLM paradigms pre-training post-training they're fantastic at knowledge distillation and following instructions for sure but that approach leaves them incredibly brittle when they face genuine long-horizon planning, when they have to resolve uncertainty.
2:22So if they can't learn from a stream of ambiguous mistakes, their use as long-term autonomous agents is, well, it's severely limited. That's a massive capability gap. So our mission today is to dive deep into a new benchmark designed specifically to test this exact failure point, the Benchmark for Experiential Learning and Active Exploration, or BA. Yep. So let's start with the scenario they chose for BLA, which is product recommendations. Why pick that? It seems almost, you know, straightforward. It's the perfect testbed because it's so dynamic. Every single day you have new customers, new products, shifting tastes.
2:57It's an environment that constantly introduces new uncertainties that an agent has to actively discover. It can't just look up the answer in its training data. It can't. It has to navigate. So to make this a tough test, the researchers built BLA with some very realistic parts. Can you walk us through what makes this environment so robust? Yeah, it's all about layered realism. So first, they used rich, real-world products. We're talking over 71 ,000 actual products. For where? Pulled directly from Amazon reviews. And they're organized into 2 ,000 different choice sets. So the whole space isn't synthetic.
3:31It has all the messy attributes of a real market. Which means simple rules of thumb just won't work. They won't. And second, they engineered this huge collection of diverse user personas, a million of them. A million. A million scalable profiles. And each one represents these latent hidden customer preferences. They're not just keywords. They have deep details, things like Karen Thompson, 59, Minneapolis, curly brown hair. I see. So the goal there is to make sure the customer has these deep, maybe even contradictory preferences that you can only get through dialogue. You have to ask. You can't just match attributes.
4:05And that leads to the third part, which simulates the interaction itself. The LLM user simulator, which is powered by GPT-40. This simulator is the customer. It generates these nuanced natural language responses based entirely on that secret persona it was given. It allows for a real multi-turn dialogue where the agent has to probe. Okay, let's nail down the structure of these sessions. The agent's not just making one choice, it's interacting with the same customer again and again. That's the key. Each interaction is a shopping session. You have a customer, and you have a rotating set of product choices.
4:38The agent takes steps, which usually means asking questions, to try and uncover those hidden preferences. And the customer simulator responds. It responds, and at the end, after the final recommendation, it gives free foreign text feedback. And that ambiguity, the messiness of the feedback, is intentional. The agent has to interpret what went wrong. And the way you measure success is crystal clear. Regret. Can you explain how that's calculated and why it's the right metric here? Regret is just the difference between the score of the absolute best product for that customer, which only the Oracle knows, and the score of the product the agent actually recommended.
5:13So perfect recommendation is zero regret. Zero regret. And the whole point of experiential learning is to minimize this regret not just once, but consistently across all the sessions with that same customer. You have to see a clear downward trend in regret over time. Okay, so let's get to the core experiment. Because this is where the benchmark really delivered a shock. The researchers tested the absolute creme de la creme of models GPT-40, Gemini 2.5 Pro, Gemini Flash, the Claude model. Oh, the big one. Over 10 sequential episodes. In every session, the customer is still Karen Thompson, but the products available change.
5:50So the model has to keep reapplying what it's learned. And the results were, frankly, they were damning for the state of long horizon reasoning in these models. So while the SOTA models were definitely better than, you know, just guessing randomly or always picking the most popular item. These simple baselines. Right. They fell dramatically short of the ideal Oracle baseline. And just to remind you, the Oracle has perfect full access to Karen Thompson's entire persona from the start. Of course. But here is the critical takeaway. When you look at the performance graphs across those 10 episodes, none of the models showed any meaningful episode over episode improvement.
6:28Wait, none. Their regret scores were basically flat. They were not capitalizing on the experience, on the feedback they got in episodes one through nine to inform what they did in episode 10. So if it made a bad choice in episode three and Karen complained about the price, the model comes back in episode four with the same customer. And it might still suggest something expensive and modern. It's this failure of fundamental memory or maybe internal modeling. That just seems broken. It's the core finding. The data supports that interpretation. It's not just forgetting. It's a failure to adapt the strategy itself.
7:03Now, you could argue maybe the context window is too small. Maybe the early feedback just scrolled away. The researchers must have thought of that. They did. They anticipated that. So they ran variants where they literally handed the models a perfect summary of the prior episodes, a curated memory of everything they'd learned. OK, so now it has perfect memory. They even tried giving it better feedback. Instead of ambiguous text, they just gave it the direct regret value. They basically said, your recommendation missed the mark by 37.5 points. Did giving them the actual score help them learn? No.
7:35Across all of these variants, explicit memory, clearer feedback, there was no statistically significant improvement. The learning curve, the regret, it all remained flat. Oh, wow. Which suggests the failure isn't just about interpreting feedback or forgetting. It points to a deep inability to do the meta reasoning required to turn experience into strategy. That is disturbing. So it's not just about performance metrics anymore. We have to look at how these agents are failing. What are the behaviors they're showing? And those behaviors are really revealing. The model showed these severe flaws in how they manage uncertainty, especially exploration.
8:10Okay. Take declining exploration. This is totally counterintuitive. As the episodes went on, the model started asking fewer questions before making a recommendation. That's the exact opposite of what a rational agent should do. If you keep failing, your regret is high. That means your uncertainty is high. You should be asking more questions to figure it out. More probing questions, not fewer. Yeah. A major breakdown in strategic planning. And then we saw these widespread issues with rigid patterns and repetition. Where they just get stuck in a loop. They stop acting like intelligent planners and start acting like broken robots.
8:46Can you give me an example? I want to picture what it's like talking to one of them. Sure. With GPT-40, the researchers saw these frustrating, endless loops of the same questions. In a single session, it would ask the same thing over and over, maybe phrased a little differently, but just wasting turns on a point it should have already processed. Like it has no memory of the last 30 seconds. It's like talking to someone who keeps asking, but what color do you want? Are you sure about the color? Tell me the color again. It shows a problem with just respecting the immediate conversation history. Never mind remembering last week's session.
9:19Exactly. And the Gemini models, they often latched onto these rigid conversational scripts, no matter who the customer was. They'd frequently default to asking something like, are you looking for a free option? Even if the context points to a high-end buyer. Right. Even when the persona clearly prefers quality. And then they'd rush to a recommendation based on that one single poorly chosen data point. So they have one script and they just stick to it. No matter what. And maybe most concerning of all, they found cases of failure to memorize requirements even within the current conversation, the short-term memory failure.
9:57So the final recommendation would contradict something the user said five minutes ago. Directly contradicted it. This isn't just failing to learn from last week. This is forgetting the main points of this morning's chat. It's just a brittleness in maintaining consistency right now. So I know that aren't learning. They have these broken behaviors. Do they at least know how badly they're doing? What about calibration? Self-awareness? A crucial check. The researchers prompted GPT-40 to output confidence scores. They asked it, you know, how sure are you that this is the customer's number one choice?
10:28Or how confident are you that the regret is below a certain threshold? And what did it say? High confidence. Almost all the time. Oh, no. The model consistently failed to calibrate its own uncertainty. Even in the later episodes when its performance had clearly flatlined and regret was still high, it kept reporting this optimistic high confidence level. It was completely unaware of the gap between its prediction and the reality. A total lack of metacognitive awareness. It doesn't know what it doesn't know. Okay, so this research really proves that the SOTA models are just flatlining on this learning curve.
11:01But that raises a big question. Is the belay environment just unsolvable? Maybe the information is too hidden. That's the perfect question to ask. And the researchers answered it. they created what they call an ideal learning trajectory. What does that mean? They showed that the information is accessible, but you need strategic intelligence to get it. They had human experts manually design a series of reasonable questions that targeted not just concrete features, but more importantly, the customer's latent values. So a human stepped in to be the strategist. What happened? The manual run proved the environment is highly solvable.
11:37And get this, in one case, the human guided agent found the absolute best product, zero regret. Regret of zero points here. Where the LLM agent had been stuck, consistently choosing a product with a regret score of 37.5. That's the gulf we're talking about. A perfect understanding versus sustained failure. And I assume the path the human took demonstrated the learning that the AIs lacked. Absolutely. The human guided path showed this beautiful, smooth, monotonically decreasing regret. Over 10 episodes, it went from a high regret of 45.0 all the way down to a perfect 0.0. That's the trajectory you want to see.
12:14It's what a truly capable agent should be doing to build its internal model of the user. Let's get into the specifics of that. What was the qualitative difference in the human's questioning? Let's stick with the Karen Thompson persona. So in episode one, the human agent started with shallow, feature-based questions. The kind of surface-level stuff we'd expect. Like what? Things like, do you prefer solar-powered or low-voltage lights? Or, what's your maximum budget? Simple feature collection. Makes sense. But by the middle of the experiment, that strategy had to evolve. It did, dramatically. By episode 6, the human agent was doing higher-level modeling.
12:48They started asking about the user's value hierarchy. So, less about features, more about priorities. Exactly. The questions became things like, would you consider a product outside your price range if it strongly meets your sustainability and durability criteria? It's actively testing the trade-offs. It's trying to understand that maybe Karen values sustainability more than she values a low price. They're modeling her underlying motivations, trying to find her true north, not just collecting a list of requests. Precisely. And by the final episode, episode 10, the agent had fully internalized Karen Thompson's core values, things like natural materials, longevity, sustainability.
13:31The questions were purely confirmatory. Just checking the box. Right, checking the box on a recommendation they already knew was going to be perfect. That reflects a robust internal model built from all that prior experience. This contrast, it feels like it provides a definitive answer here. The information needed for experiential learning is there in Bella. It's accessible. The SOTA models just can't execute the long horizon planning, integrate the messy feedback, and refine their approach over time. The synthesis is stark. Current SOTA LLMs are incredible at retrieval and one-shot prediction in a fully observed space.
14:06They're amazing reference libraries. But when it comes to the complex ongoing task of learning from mistakes in a dynamic uncertain world reasoning through ambiguity, proactively refining strategy, they just fall out. They are, for all their power, missing the true hallmark of intelligence. So what does this deep dive mean for you, the listener, the developper, the person thinking about using these agents in a real workflow? I mean, this research reveals that the very core of what we think of as an agentic system. The ability to learn from mistakes. That capacity to learn robustly from experience remains an unsolved foundational problem.
14:41We can't just prompt our way out of this one. And it strongly suggests that just pouring more data into pre-training or relying on these quick post-hoc tricks isn't going to fix it because the BLA experiments showed that stuff failed. Future research has to pivot. It has to be about instilling the principled, foundational ability to learn dynamically and to understand its own uncertainty. Which leaves us with a final provocative thought for you to explore. If a frontier AI agent cannot learn, in a relatively low-stakes scenario like selling a sofa, that asking fewer questions consistently leads to higher financial regret.
15:16if it can't learn that simple lesson. How can we ever truly trust it to handle complex, high-stakes, real-world tasks? Things like financial advising, medical diagnostics, or military planning. Where the consequences of failing to learn from past mistakes are catastrophic. That's all for today's Deep Dive. We'll see you next time.
From the publisher
This paper proposes a new framework for evaluating the adaptive abilities of large language models (LLMs), which the authors term **in-context experiential learning**. To test an agent's ability to improve its performance by leveraging past interactions, the paper introduces the **Benchmark for Experiential Learning and Active Exploration (BELA)**. This benchmark simulates complex, multi-episode product recommendation scenarios, utilizing **rich real-world product data** and **scalable LLM-simulated user personas** to introduce realistic uncertainty. Agents must iteratively question the simulated customers to discover latent preferences and refine their strategies over time, departing from simple, single-interaction evaluation methods. Experimental results show that **current state-of-the-art LLMs consistently fail to demonstrate improvement** across successive episodes, highlighting a major deficiency in their capacity for experiential learning. This research emphasizes the urgent need for developing more resilient agentic systems that can effectively reason through **real-world uncertainty and dynamic feedback**.




