In short
The episode explains a research framework for “agents as tool-use decision-makers,” aiming for optimal, efficient behavior by aligning when an agent uses internal reasoning tools versus external tools (search, APIs, physical actions). It introduces “knowledge walls” (what the agent knows vs what it must fetch) and “decision switches” (the internal choice to think or act). Key claim: maximum efficiency comes from perfectly aligning the tool-use decision with the knowledge wall, minimizing unnecessary internal computation and costly external tool calls.
Guest backgrounds
No guests are named; it’s a two-host discussion.
Notable examples
A travel agent that hallucinates train schedules when it overuses internal tools; a “10 + 5” query where an agent needlessly searches Google; embodied robots where moving/using sensors is treated as an external tool call.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VODefining Autonomous Agents
0:10 to 1:41
Discussion on the capabilities and implications of autonomous agents.
“Yeah, we're really moving beyond just chatbots.”
Optimal Behavior of Agents
1:41 to 3:35
Exploration of what defines optimal behavior in autonomous agents.
“Meaning we want agents that are strategically smart about when and how they choose to learn or retrieve information, not just blindly following steps.”
Tools and Their Roles
3:35 to 6:10
Understanding the role of internal and external tools in agent decision-making.
“I have to push back a bit on epistemically equivalent because, look, if I sit and think really hard about a complex problem for, say, three hours, that costs me maybe some mental energy, a cup of coffee.”
Knowledge and Decision Boundaries
6:10 to 9:08
Discussion on the knowledge and decision boundaries that govern agent intelligence.
“Let's maybe skip the formal math symbols from the papers and just get the concepts.”
Misalignments in Tool Use
9:08 to 11:32
Identifying potential misalignments in agent tool use and their consequences.
“The sources really focus on these misalignments as key failure modes, especially things you'd want to catch during training or fine-tuning.”
The Pursuit of Optimal Behavior
11:32 to 14:00
Defining optimal behavior and its importance in autonomous systems.
“And it's important to remember, this whole dynamic isn't static.”
Minimizing Tool Use for Efficiency
14:00 to 14:56
Learn about the goal of minimizing both internal and external tool usage in decision-making processes.
“So it gets the right answer eventually, but the time cost is way too high.”
Training for Judicious Tool Use
14:56 to 17:42
Discover strategies for training AI to use tools effectively and efficiently.
“That's the practical challenge since we can't easily peer inside the agent's mind and directly measure that abstract knowledge wall.”
Expanding the Concept to Robots
17:42 to 19:58
Explore how the discussed theories apply to embodied agents and robots in the physical world.
“We started by moving past this maybe simpler idea of agents as just action executors, just doers.”
Transcript
Automatic transcript. May contain errors.0:00Welcome back to the Deep Dive. For years we've talked about large language models, right? But today, we're jumping straight into what feels like the next frontier. Autonomous agents. Yeah, we're really moving beyond just chatbots. Absolutely. We're talking about systems that don't just, you know, answer questions. They can independently plan and execute pretty complex multi-step tasks. Things like managing your entire corporate travel itinerary, maybe end-to-end. Exactly that. Or even something really sophisticated like designing a novel scientific experiment from scratch. The potential is huge.
0:36Okay, let's unpack this then. Because if we are going to trust these machines with real world stuff, resources, complexity, money, time, even physical actions, well, they can't just be good at executing the plan. No, good isn't good enough here. They have to be like perfectly efficient. They can't waste steps, right? And they definitely can't be wasting computation. That gets expensive fast. It does. So the mission today is to dive into some foundational research that tries to answer fundamentally. What defines an agent's optimal, most efficient behavior? How do we even think about that? That's really the key pivot we're seeing.
1:12Right now, most agents are designed, well, pragmatically. They do some reasoning, then they take an action, then more reasoning, then another action. Turn it back and forth. Exactly. It's functional, gets things done sometimes, but it lacks a really solid principled theory behind it. What researchers are proposing now is a unified framework. Okay. And this framework fundamentally shifts the design focus away from just action execution, just doing stuff, towards creating genuine knowledge-driven intelligence systems. Knowledge-driven. Meaning we want agents that are strategically smart about when and how they choose to learn or retrieve information, not just blindly following steps.
1:50Right. And the core claim we've pulled from the sources, the big idea for this deep diet is that optimal behavior, that peak efficiency we talked about, happens when an agent perfectly aligns its decision about when to use a tool with its actual knowledge limit. You know, what it actually knows versus what it needs to find out. That's the nub of it. If they align those two things perfectly, the agent minimizes all unnecessary action, whether that's internal thinking or external searching. Maximum efficiency. Okay, here's where it gets really interesting for me. Because to even measure that alignment or even talk about it properly, we apparently have to redefine some basic building blocks.
2:30We do. We have to kind of rethink things from the ground up. Starting with the idea that thinking and doing are sort of the same kind of activity. In a way, yes. In this unified theory, the term tool gets a massive upgrade. It's defined really broadly. A tool is any process or interface. Any. Any. Whether it's internal computation, like reasoning, or an external interaction, like making an API call, if it contributes to acquiring the knowledge needed to complete the goal, it's a tool. Okay, so thinking is a tool, searching Google is a tool. Exactly. So if that's the definition of a tool, then the agent itself is, what, just the coordinator?
3:06Yeah, pretty much. The agent is the decision maker. It's the entity that dynamically coordinates all these different tools, internal, external, to hit its objective. That leads to this unified view of reasoning and acting, right, where internal thoughts and external actions are treated as epistemically equivalent. That's the term used, yes. Yeah. Epistemically equivalent tools within a single framework. Both are fundamentally attempts to acquire knowledge needed for the task. Okay, wait a minute, though. I have to push back a bit on epistemically equivalent because, look, if I sit and think really hard about a complex problem for, say, three hours, that costs me maybe some mental energy, a cup of coffee.
3:46Right. But if the agent makes five API calls to, I don't know, a financial data service or a flight booking system, that could cost real money, real world resources. How can we really treat a costly external action as truly equivalent to, like, a free internal thought process? That's a really crucial question. And it absolutely speaks to why efficiency becomes so central. You're right. They are equivalent in their purpose. Both are methods to retrieve or generate knowledge. Okay. Purpose-wise. But they are definitely not equivalent in their cost. And that cost difference, whether it's time, money, computation, or even physical energy for a robot, is precisely why efficiency becomes the ultimate metric for optimality.
4:28We'll definitely circle back to that cost aspect. Got it. So equivalent purpose, different costs. That makes sense. But for the framework itself to hold together, we need to group these tools. We basically would put them into two main categories. Okay. First, you've got your internal cognitive tools. These are the mechanisms that support systematic thinking within the agent's existing knowledge. The stuff baked into its parameters during training. So things we've heard of, like chain of thought. Exactly. Chain of thought, reflection techniques, decomposing a problem, methods that trigger internal knowledge, retrieval, and manipulation using what it already knows.
5:03And the second category. External physical tools, or just external tools, really. These are the modules or interfaces outside the model's internal parameters, like running a search engine query, accessing a database via an API, maybe calling another specialized model, or even executing a physical action, like move to room A, if the agent is in bodies, like a robot. Ah, okay. Okay. So these are for getting knowledge that's beyond what the agent just knows internally. Stuff from the outside world. Precisely. Knowledge it doesn't intrinsically possess. I see. So the agent's journey through a task isn't just step, step, step.
5:38It's more like it alternates, right? It asks itself, okay, what do I know right now that's relevant, that's using an internal tool? And then if that's not enough, it switches to, okay, what information do I need to pull in from the outside world using an external tool? That's the cycle. Internal tool invocation retrieves some knowledge, then maybe an external tool invocation retrieves more knowledge, and so on until the goal is met. And every single decision point becomes a tool use decision, internal or external. You got it. That frames the whole interaction landscape. Okay. Now that we've unified the tools, let's talk about those two critical limits or boundaries that govern the agent's intelligence.
6:17Let's maybe skip the formal math symbols from the papers and just get the concepts. Sounds good. Let's define the first one, the knowledge boundary, or maybe think of it as the knowledge wall. The knowledge wall. I like that. It's essentially the hard frontier. On one side, you have all the facts, patterns, and reasoning capabilities the model has already learned and embedded in its parameters, its internal knowledge. What it knows. Right. And on the other side is everything else, all the external knowledge accessible only from the outside world via those external tools. If the knowledge needed is beyond that wall.
6:52Internal thinking won't cut it. It has to reach out. Exactly. That's the model's fundamental epistemic limit for a given moment or task. Okay. That's one boundary and the other one you mentioned. The other is the decision boundary or maybe the decision switch. The decision switch. Okay. This is more functional. It's the point in the agent's process where it actually decides it's on the fence. Do I flip the switch this way and use an internal tool set to reason further? or do I flip it the other way and activate an external tool set to act or query the world? Got it. Knowledge wall is what it knows.
7:24Decision switch is where it chooses internal versus external. Perfect. And here is where we connect them with the central hypothesis, the decision knowledge alignment principle. This is the core idea. It is. For truly optimal behavior, maximum efficiency, minimum waste. The agent's tool use decision switch, partial DD, must align perfectly with its knowledge wall. Meaning? Meaning the agent should only trigger internal tools when the knowledge it needs is actually within its internal capacity within that wall. And it should only switch to external tools when the information absolutely must be acquired from the environment because it's outside the wall.
8:01No guessing, no unnecessary searches. Precisely. What's fascinating here and kind of mind bending is that this alignment requires something like computational metacognition. Metacognition, like thinking about thinking. Exactly. Just like you or I hopefully do, the agent needs to have some capacity to assess its own cognitive state. Needs to know what it knows and maybe even more importantly, recognize its own uncertainty. So it needs to know when it's bluffing, essentially. In a way, yes. It's like translating that human heuristic to AI. Trust your internal reasoning when you're confident, when the answer feels solid based on your knowledge.
8:40But the moment you feel uncertain or you hit a logical dead end internally. You look it up. You ask someone. You seek external help. You seek external help. You use an external tool. Wow. Okay, that immediately highlights where things go wrong, doesn't it? If that alignment between the decision switch and the knowledge wall is off. Big problems. We don't just see simple mistakes. We see potentially massive inefficiency. And in a real-world deployment, that could be incredibly costly, right? Wasting time, money, compute cycles. Absolutely. The sources really focus on these misalignments as key failure modes, especially things you'd want to catch during training or fine-tuning.
9:17Okay, let's look at those failure modes. What's the first one? What happens if the agent's decision switch is set, let's say, too far out? It thinks it knows more than it does, essentially placing the switch beyond its actual knowledge wall. That leads directly to internal tool overuse. The model keeps trying to reason internally, using its cognitive tools, to generate knowledge it simply doesn't possess. Ah, so it's trying to think its way to an answer that isn't in its head. Exactly. And this is the direct path to poor logic, making stuff up, invented facts, the notorious problem of hallucination.
9:54Can you give an example? Sure. Imagine that travel agent scenario again. It needs the current train schedule between two cities. Its knowledge wall stops at, say, general train travel patterns it learned years ago during pre-training. Right. But its decision switch is faulty. It decides it can figure out the current schedule internally. So instead of making an API call to the railway company, the correct external tool use. It just makes up a schedule. It confidently hallucinates a plausible sounding schedule based on outdated patterns. That's pure inefficiency. Because the resulting plan will fail.
10:25The action was wasted internally. Okay, that's bad. What about the reverse problem? When the decision switch is too conservative, maybe, set inside the knowledge wall. That's external tool overuse. And you could argue this one is potentially worse in terms of sheer wasted computation and resources. How so? Here, the model defaults to using external tools like running a search query or making an expensive API call, even when it already possesses the required knowledge internally. So it knows the answer but asks anyway. Basically, yes. It's wasting time, computation, maybe money on API calls because it's ignoring its own internal capacity.
11:02The alignment is broken in the other direction. A simple example. A perfect, almost trivial example they sometimes use. You ask a highly capable LLM agent, what is 10 plus 5? It absolutely knows this, either through calculation capabilities or memorized facts. But if its decision switch is misaligned, it might immediately initiate a Google search for 10 plus 5. It performs a needless external action because it fails to recognize its internal sufficiency. Just unnecessary clicking, essentially. Wow. Okay. And it's important to remember, this whole dynamic isn't static. It shifts during inference while the agent is actually solving a task.
11:40How does it shift? Well, an agent almost always starts a complex task with incomplete knowledge, right? Right. Its initial knowledge wall might be limited for that specific context. Sure. But every time it successfully uses an external tool, makes that API call, performs that search, takes that physical action, it retrieves new relevant information. Ah, so the external tool use actually expands the knowledge available to it. Exactly. It incrementally expands the agent's effective knowledge wall for that specific task. Yeah. So the self-assessment has to be continuous. After each step, associated an external one, I have to implicitly ask, okay, did that last piece of information give me enough to feed internally now, or do I still need to gather more from the outside?
12:22It's constantly reevaluating its own knowledge state relative to the goal. Constantly. Do I have enough now to stop gathering and start concluding or acting internally? Right. And that brings us straight back to this idea of optimal behavior. Getting the problem solved correctly is like table stakes. It's necessary, but it's not enough. Not nearly enough for autonomous systems. That agent you described earlier, the one that makes 2 ,000 unnecessary internal calculations and maybe 300 pointless search queries to get the right answer eventually, that's actually a failure in this framework, isn't it?
12:56It's a failure of efficiency, absolutely. We must prioritize efficiency alongside correctness. So how do we define that efficient, optimal behavior? We do need to define it. And the research highlights a few theoretical modes. But the real focus is strictly on minimization of effort, minimizing tool use. Makes sense. The first promising mode they really look at is maximizing internal tool use while minimizing external tool use. Okay, so think a lot, act a little. Kind of. This promotes high autonomy, relying on internal reasoning, which is usually good and cheap computationally. But it has a pitfall.
13:33Which is? It often leads to what the researchers call overthinking. Overthinking. What does that look like for an AI agent? It doesn't get stressed. No, not stressed. It means generating unnecessarily long and complex reasoning chains internally. The agent becomes so determined not to incur the cost or risk of an external tool call. That it just keeps thinking and thinking. Exactly. It might spend, you know, two minutes generating a five-paragraph internal monologue reflecting on some nuanced concept when a single quick targeted search query could resolve the point instantly. So it gets the right answer eventually, but the time cost is way too high.
14:09Precisely. Correctness is achieved, but not efficiently. So that's not the ultimate goal then. What is? The true target, the optimal state, the ultimate goal is minimizing both internal and external tool use. Minimizing everything. thing. Minimizing total effort. This represents the most efficient trajectory possible to solve the task. It reflects the ultimate vision. A truly autonomous intelligent system should solve tasks with the minimum possible epistemic effort. Epistemic effort, like knowledge gathering effort. Exactly. Whether that effort involves internal computation or real world action via external tools, it's like the agentic equivalent of seeing the solution to a complex puzzle in your head with just a single decisive insight rather than brute forcing it.
14:53Elegant deficiency. That's the aspiration. Okay. So if the optimal goal is minimizing total tool use, especially minimizing external tool use because it's costly and maybe acts as a proxy for that hard to measure knowledge wall, how do researchers actually train a machine to be so judicious, so restrained? That's the practical challenge since we can't easily peer inside the agent's mind and directly measure that abstract knowledge wall. Right. We use the minimization of costly external actions as an actionable proxy metric for alignment. If the agent learns to use external tools sparingly and only when necessary, it's likely learning to better gauge its internal limits.
15:31Okay, so how do they encourage that minimization during training? There are a couple of key strategies emerging. One major one is a shift in the pre-training objective itself, sometimes called agentic pre-training. Agentic pre-training, different from how LLMs are normally trained. Yes. Instead of the traditional LLM training goal, which is usually predicting the next token, basically, the next word in the sequence, we shift the objective to next tool prediction. The model is trained explicitly to predict the most appropriate tool or interface to invoke at each step of a task. Ah, so the decision itself becomes the thing it learns to predict.
16:06Exactly. It transforms the interaction process, the decision to search or calculate internally or call an API, or even just think more into a structured core modeling target, learning how to proceed, not just what to say next. Interesting. And you mentioned another approach. Yes. The other major avenue is agentic reinforcement learning, or agentic RL. RL makes sense, optimizing for rewards. It does. RL is already powerful because it can optimize for final outcomes. But here, the key is that it needs to optimize not just for the final answer's correctness, but for the process of getting there. Right.
16:42The efficiency of the journey. Precisely. And the real challenge then is in designing the reward function, if you only reward the agent for getting the final answer right. It might learn sloppy habits, like searching all the time, just in case. Exactly. It will likely default to inefficient behaviors because searching often increases the chance of eventual correctness, even if it wasn't strictly needed. So optimal agentic RL must balance correctness with an explicit, often quite high penalty for unnecessary tool calls. Penalizing it for clicking when it shouldn't have. You got it. Some researchers are calling this approach OTCPO.
17:19It stands for over-tool calling penalty optimization. You're essentially training the agent to exercise restraint, to develop that computational self-awareness, making sure the decision to use a costly external tool is always justified by a genuine knowledge gap. Training for judiciousness. That's a good way to put it. Okay. That feels like it brings us towards the close of this deep dive. We started by moving past this maybe simpler idea of agents as just action executors, just doers. And now we're understanding them as potentially sophisticated knowledge-driven decision makers whose path to true optimality to real efficiency is entirely reliant on developing a kind of metacognitive self-awareness.
18:01Knowing where their knowledge wall lies. And aligning their decision switch right up against it, using the minimum effort necessary. Absolutely. And if we connect this back to the bigger picture, this entire theoretical framework is really crucial if we want to scale up genuine agent autonomy, especially in open-ended, complex, real-world environments. Because the goalposts shift. They do. The goal is no longer just achieving raw capability, can it do the task? Yeah. But achieving epistemic efficiency, using the minimum necessary resources, internal or external, to acquire the maximum necessary knowledge to succeed.
18:36Minimum resources, maximum knowledge. That's a powerful idea, which I think leaves us with a final thought for you, the listener, to chew on. Go for it. This framework, this whole idea of knowledge walls and decision switches and tool use, it doesn't just stop at agents running on a server, right? It extends quite naturally to embodied agents. Robots. Robots moving in the physical world. Think about it. If a model's physical body, its sensors, its arms, its legs is effectively an external tool. It is in this view. And a simple physical action like move to room A is essentially a tool call that yields sensory knowledge.
19:11What's in room A or is the door open? Yeah. then how does that foundation agent, that robot brain, learn physical metacognition? How does it align its decision boundary with its physical knowledge boundary? Wow, yeah. The challenge then becomes minimizing unnecessary physical energy expenditure. Exactly. How does the robot learn when it's sufficient to just pause and think internally, maybe access its internal map or memory, versus when it needs to burn battery life, physically walk across the room, and use its body, its sensors, to check if the door is actually open or closed. The boundary shoots entirely from just cognitive memory to encompass physical reality and the cost of interacting with it.
19:51That feels like a whole other layer of complexity. Physical metacognition for efficient action. Something to think about.
From the publisher
This position paper argues for a new epistemic theory of agents that views internal reasoning and external actions as equivalent epistemic tools for acquiring knowledge. The core argument is that for an agent to achieve optimal and efficient behavior, its tool use decision boundary must be aligned with its knowledge boundary, meaning it should only resort to external tools when necessary knowledge is unavailable internally. The paper formalizes this concept by defining tools, agents, and optimal behavior, and introduces three principles of knowledge: foundation, uniqueness/diversity, and dynamic conservation, which provide a theoretical basis for designing next-generation knowledge-driven intelligence systems capable of adaptive, goal-directed behavior with minimal unnecessary action. Finally, the authors propose paths toward achieving agent optimality through enhanced training paradigms like next-tool prediction and reinforcement learning that rewards both correctness and efficiency.




