In short
“Age of AI Agency” and the LIME/LiMI research (“Less is More for Intelligent Agency”) arguing that autonomous AI agents can be trained with far fewer examples by using high-quality “trajectories” rather than massive datasets.
Guest backgrounds
No guests are named; it’s a host “Deep Dive” discussion of the LiMI paper.
Key claims
Agency differs from language learning; autonomy emerges from the depth/quality of demonstrations. LiMI trains with 78 total trajectories (not 78k), each capturing reasoning, tool use, and environment feedback/corrections. 78 samples outperform models fine-tuned on 10,000 samples.
Notable examples
Vibe coding (Gamoku web app + minimax/alpha-beta AI) and research workflows (scientific function discovery with order-of-magnitude lower loss on first attempt; “NBA scenarios” requiring web search and multi-constraint synthesis). AgencyBench score: 73.5% vs GLM-4.5 code 47.8%.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOUnderstanding AI Agency
0:45 to 3:18
Exploring the shift from traditional AI to autonomous AI agents.
“And that's what this research we're diving into today challenges.”
The LIME Approach Revealed
3:18 to 4:52
Introduction of LIME and its core finding that less data can lead to more sophisticated AI agency.
“Strategic curation beats data abundance.”
Trajectories and Their Importance
4:52 to 7:05
Discussing the significance of using high-quality trajectories in training AI.
“Learning from mistakes, adapting the plan based on feedback, that's core to real agency.”
Performance Comparison of Models
7:05 to 9:01
Comparative performance of LIME against other large models and highlighting its efficiency.
“It's specifically designed to test these agentic capabilities, the planning, the tool use, the problem-solving across both Vibe coding and research workflow tasks.”
Real-World Applications and Tasks
9:01 to 12:32
Examples of tasks where LIME showcased superior agency in coding and research.
“They found that LiMI also generalized better on other benchmarks, testing things like tool use, general coding, even data science tasks.”
Transcript
Automatic transcript. May contain errors.0:00Welcome to the Deep Dive. Today we're digging into something that feels like a real turning point. We're calling it the Age of AI Agency. Exactly. We're moving past AI systems that just, you know, think the ones great at reasoning or generating text. Right. The large language models we've gotten used to. Yeah. And we're shifting towards systems that actually work. AI that can execute tasks, use tools, solve complex problems all on their own autonomously. And the conditional way to get there, the scaling law approach. Yeah. Well, everyone assumed you'd need just unbelievable amounts of data, right?
0:33Yeah. More data, better AI. That's been the mantra. More data, more parameters, bigger compute. It worked incredibly well for language, for reasoning. But does it hold up for genuine autonomy, for an AI that needs to, say, manage a GitHub project or design a scientific experiment? Well, that's the big question. And that's what this research we're diving into today challenges. It's called LAME, Less is More for Intelligent Agency. Less is more. Okay, I'm intrigued. So the mission today is to unpack how this LIME approach flips the script on needing massive data sets for AI agents. Precisely. The core finding, and it's pretty stunning, is that really sophisticated agency, this ability to do complex things, can emerge from a surprisingly small number of very strategically chosen examples.
1:21Okay, hit me with the numbers then. What are we talking about? We're talking state-of-the-art results using just 78, 78 training samples. 78, not 78 ,000. No, 78 total. It sounds almost unbelievable, I know. Okay, let's define terms first. When we say agency with a capital A, what exactly are we talking about? What does that look like in practice? So agency here means the AI isn't just responding to a prompt. It's acting like an autonomous agent. It can figure out there's a problem, come up with ideas on how to solve it. Like forming a hypothesis? Exactly. Formulating hypotheses and then actually executing a plan, using tools, interacting with its environment, maybe even asking for clarification if it needs it.
2:02It's proactive problem solving. And the research focused on some pretty challenging areas to test this, right? Not just simple tasks. Definitely not simple. They picked two really complex domains. First was something called Vibe Coding. Think real world software development. Messy stuff. Debugging, working with existing code bases, using Git. All of that. Iterative problem solving, collaborating maybe through GitHub pull requests. The second domain was research workflows. Okay, like science. Yeah, the actual process. Searching literature, analyzing data, maybe designing experiments, trying to pull real insights from scattered information, really complex stuff.
2:40So given how complex vibe coding and research workflows are, why wouldn't more data be better? This goes against the grain. What's the core idea behind LandLine? It's what they call the agency efficiency principle. The central idea, the hypothesis, is that learning agency is fundamentally different from learning language. How so? Language models get better with sheer volume, seeing countless examples of text and code. But agency, they argue, emerges from the quality and depth of the examples, not the quantity. It's about seeing really good demonstrations of autonomous behavior. So less about seeing a million lines of code, more about seeing one really complex coding problem solved, start to finish, with all the back and forth.
3:19Exactly. Strategic curation beats data abundance. It's about the richness of the demonstration. Has this idea, quality over quantity, shown up elsewhere? Is there a precedent for this? Oh, absolutely. This isn't totally out of the blue. Think about Lemma that showed you could get great chat alignment with just about a thousand curated examples. Right, I remember that. And Limo did something similar for mathematical reasoning, using only around 817 high-quality samples. So Limo is kind of building on that trend, applying it to this much broader, more complex idea of agency and action. Okay, let's get to the heart of it then.
3:57Those 78 samples. What made them so special? They must be incredibly dense or detailed. They are. They're not just question-answer pairs. They call them trajectories. A trajectory captures the entire workflow, the whole journey of solving one of these complex tasks. The whole journey. What does that include? It's got three key parts captured sequentially. First, there's the model reasoning. This is like the AI's internal thought process, its plan, its analysis, why it's doing what it's doing. Okay, the thinking part. Right. Second, model tool calling. This is the execution, the AI actually using a tool, like running code, calling an API, searching a database.
4:35The doing part. Exactly. And third, crucially, environment observation. This is the feedback. What happened when the tool ran? Did it work? Did it error out? Did the user step in and say, wait, that's not quite right? Ah, so it includes corrections and failures too. Yes, that's critical. Learning from mistakes, adapting the plan based on feedback, that's core to real agency. And these trajectories captured all of that messy reality. You mentioned complexity. I saw a note that the longest trajectory was something like 152 ,000 tokens. Yeah, 152 ,000. That's huge. It shows just how deep these interactions went.
5:11Multiple turns, complex reasoning, tool use failures, successes. It's that density that provides the powerful learning signal. It's learning the process, not just memorizing facts. Precisely. Learning how to navigate complex, long-horizon problems. But hang on a second. If one sample is 152 ,000 tokens long and you only need 78, isn't creating one of those incredibly detailed trajectories super expensive? Maybe even more expensive in terms of human effort than scraping millions of web pages. That's a really sharp point. The cost definitely shifts. It's less about massive compute for training on huge data sets and more about expert human labor to create and curate these high quality trajectories.
5:54So you're trading compute cost for expert curation cost? Kind of. But the argument is that the value you get from that curated data is orders of magnitude higher for teaching agency. You're investing in teaching the model how to think and act autonomously, which seems to be way more effective than just showing it more raw data. The results seem to bear that out. OK, makes sense. How did they actually get these 78 golden trajectories? Where did the problems come from? They used a pretty clever mix to ensure they were realistic. 60 of the problems came directly from real-world tasks that professional developers and researchers actually face, sometimes pulled right from, like, method sections in papers.
6:31So genuine problems. Right. And the other 18 were synthesized, but in a really interesting way, they used GPT-4. Actually, I think the people mentioned GPT-5 level analysis. Yeah, analyzing real GitHub pull requests from very popular repositories, ones with over 10 ,000 stars. So they were generating complex, realistic coding problems based on actual high-stakes development work. Ecological validity, they call it. Okay, this setup is fascinating. Extremely high-quality, curated data capturing complex workflows. Now, the payoff. How did Lemmai actually perform? They used something called AgencyBench.
7:06Yep, AgencyBench. It's specifically designed to test these agentic capabilities, the planning, the tool use, the problem-solving across both Vibe coding and research workflow tasks. And the headline number for Lemmai. Lemmy achieved an average score of 73.5 % across Agency Bench, which is, frankly, remarkable. Okay, 73.5%. How does that compare? This is where we see if the less is more thing really holds up against the big guys. This is where it gets really interesting. The comparison models, big foundation models, they struggled. GLM 4.5, a very capable model, scored 45.1%. Okay, so Lara is significantly ahead already.
7:42Way ahead. And others, like Quen3, Kimi, they were down in the mid-20s. DeepSeq v3.1 was barely above 10%. Wow. So Lemurai isn't just a little better, it's in a different league on these agency tasks. It really suggests that the generalist training of those huge models doesn't automatically translate into effective autonomous agency without specific high-quality agentic training data. Let's zero in on that data efficiency proof. You mentioned comparing Lemion with its 78 samples against a model trained specifically for these tasks, but with lots more data. Right. They compared Lemion to GLM 4.5 code, which was fine-tuned for coding tasks using 10 ,000 samples, a pretty standard approach.
8:2110 ,000 samples versus 78. What was the performance gap? Lemion with its 78 samples scored 73.5%. The GLM 4.5 code model with 10 ,000 samples scored 47.8%. Let me just process that. IMI got a score over 25 percentage points higher. It's a 53.7 % improvement in performance. While using, what is that, over 128 times fewer samples. Exactly. 128 times less data for massively better performance on complex agency tasks. That's the headline finding. It really drives home the agency efficiency principle. Quality over quantity, dramatically so. That is genuinely game-changing. It implies that the type of learning signal is just fundamentally more important than the volume for this kind of capability.
9:06It seems so. And it wasn't just on AgencyBench. They found that LiMI also generalized better on other benchmarks, testing things like tool use, general coding, even data science tasks. So it learned a robust capability, not just how to solve those 78 specific problems. OK, let's make this more concrete for everyone listening. Can we walk through a couple of the actual tasks, tasks that really show off this agency thing in action? Maybe start with that coding example, the Gomoku game. Ah, yes, the Gamoku Expert AI task. This was a beast. It wasn't just write a function. The agent had to build a whole web front end for the game.
9:38Like HTML, CSS, JavaScript. Yep. Then add tricky game logic, like detecting win conditions, handling multiplayer. And the real kicker was implementing an advanced AI opponent using algorithms like Minimax with alpha-beta perning. Okay, any developer listening knows Minimax is non-trivial. It requires careful planning. How did the standard big model do? The base GLM model, the one trained on scale, it struggled. It got stuck trying to render the board correctly. Then it failed when trying to implement different AI difficulty levels. It needed the human user to step in with hints multiple times.
10:14So it wasn't truly autonomous? Not really. LayMai, on the other hand, it managed to complete all the subtasks, including the complex AI logic, without needing interactive hints. It just worked through it much more autonomously. That's a clear difference in planning and execution. Okay, let's switch domains. How about a research task? Task 8, the scientific function discovery. Right. This one was about fitting a mathematical function to some scientific data, but the goal was extreme precision, finding the function that minimized the error, the loss, down to a really tiny target value. So it involves iterative refinement, trying things, seeing the results, adjusting?
10:50Exactly. Hypothesis testing in action. The base GLM model, even with several nudges from the user, eventually got the loss down, but only so far. And Lemai. Lemai achieved a final loss that was an order of magnitude smaller, basically. Ten times better precision, and it did it on its very first attempt. No nudges needed. Wow. Okay, that's just much better intuition. Or at least much more efficient exploration of the solution space. Better internal planning and prediction, definitely. It wasn't just randomly guessing. It seemed to understand the path to a good solution much more directly. One more task nine, the NBA scenarios.
11:27This sounds like a test pulling together lots of different facts. Oh, it did. It was designed to be really tricky, needing the agent to search the web and combine multiple complex conditions. Things like finding a player who met criteria A, B, C, and D. Like all-star status, specific trade details, maybe even weird unrelated facts. Exactly. One condition involved finding players traded for specific assets who also shared a first name with a member of the band from the White Album. Okay, that's intentionally convoluted, the kind of thing that makes humans pull their hair out. How did the models cope?
11:58The base model really struggled here. It gave incorrect answers initially on some subtasks and needed the maximum number of hints allowed to get things right. MMA. On a particularly complex one about Paul George's trade and achievements, Lemon nailed the correct answer on the first try. No hints. For another tricky one about James Harden, it still got it right but required significantly fewer internal reasoning steps. Just cleaner, more efficient synthesis of information. So across coding, scientific discovery, and complex information retrieval, LEMI consistently showed better planning, execution, and autonomy, all from just 78 examples.
12:37That's the story. Okay, let's wrap this up. This deep dive into LEMI really feels like it shifts the ground under our feet regarding AI development. The big takeaway for me is this agency efficiency principle. Yeah. It strongly suggests that the path to truly capable AI agents, the ones that can do things in the world, isn't just about throwing more data and compute at the problem. It's about the quality of the data. Specifically, those rich, detailed demonstrations, the trajectories capturing complex, successful, self-corrected behavior. Exactly. It's the depth of that learning signal. Investing in curating those high-quality examples seems to unlock fundamental improvements in the AI's ability to reason and plan even before it starts using tools.
13:20That strategic curation is the new leverage point. So for you listening, here's the final thought, the provocative question we want to leave you with. If we can unlock this level of sophisticated autonomy in areas like coding and research using just a handful of strategically crafted demonstrations. Demonstrations that capture the essence of expert problem solving in those fields. Then what other complex human skills could we teach AI by focusing on quality over quantity? What happens if we define and curate the essence of expertise in, say, negotiation or complex project management or even artistic creation?
13:54What complex workflows trajectories should we be meticulously collecting next to unlock a whole new level of AI agency in domains we haven't even considered yet? Food for thought. That's all for this deep dive. Thanks for joining us.
From the publisher
This research paper introduces the Agency Efficiency Principle and a methodology called LIMI (Less Is More for Intelligent Agency), arguing that developing autonomous AI systems requires strategically curating small datasets of high-quality agentic demonstrations rather than scaling data volume. The authors define Agency as the capacity for autonomous reasoning, acting, and tool use in complex workflows, specifically focusing on vibe coding (collaborative software development) and research workflows. Experimental results presented using the AgencyBench benchmark show that the LIMI model, fine-tuned on only 78 curated samples, significantly outperforms state-of-the-art baseline models trained on datasets that are orders of magnitude larger, validating the Less-Is-More hypothesis for agentic intelligence. The document also provides extensive details on the AgencyBench tasks, which involve multi-step, complex problems requiring execution in a command-line interface environment.




