Position: Agents Should Invoke External Tools ONLY When Epistemically Necessary

6 Jul 2026 · 12 min · 6 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

When AI agents should use external tools (search, APIs, code) versus rely on internal reasoning, arguing tools shouldn’t be used unless epistemically necessary.

Guest backgrounds

No guests are mentioned; the episode is presented as a “deep dive” by the hosts.

Key claims

External tool use is not a free shortcut; it shifts learning by rewarding correctness without reinforcing internal reasoning (“unnecessary delegation”). This causes long-term intelligence stagnation and dysfunctional tool-use behaviors. The decision hinges on “epistemic effort” (total cognitive work to reduce uncertainty to zero), which can be distributed internally or externally but not eliminated.

Notable examples

toddler “crawl vs walk” analogy; GPS analogy for atrophied internal skills; long-context vs retrieval-augmented generation (RAG) debate; future embodied robots where physical constraints redefine epistemic effort.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

The Shift in AI Functionality

1:04 to 2:18

Discussion on how AI has evolved from text generators to autonomous agents.

“You're listening to this because the artificial intelligence systems we use every single day are, well, they're acting a lot more like that toddler than we might want to admit.”

Understanding Knowledge Boundaries

2:18 to 4:33

Exploration of the internal and world task sets and their implications.

“Because the underlying framework proposed by the research, the theory of agent, it suggests we need to basically erase the traditional line between thinking and acting.”

The Consequences of Lazy Tool Use

4:33 to 5:59

Examining how unnecessary delegation affects AI learning and growth.

“You can't just wave a wand and make it go away.”

Behavioral Regimes in AI

5:59 to 7:59

The four behavioral regimes of AI agents based on their internal and external efforts.

“It's exactly like relying entirely on a GPS to navigate your own neighborhood.”

Training AI with Metacognition

7:59 to 9:20

Strategies for training AI to assess its own knowledge before leveraging tools.

“and it only asks for outside help when it's epistemically necessary.”

Future Directions for AI and Robotics

9:20 to 12:00

Discussion on the application of AI frameworks to embodied agents and physical interactions.

“The hardware limits the context window, so engineers externalize the effort to save compute.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00You know, there's this very distinct moment you experience if you spend enough time around a toddler who's just learning to walk. Oh yeah, the coffee table cruising phase. Exactly. They'll be standing there gripping the edge of the coffee table. And they want a toy that's, I don't know, just a few feet away across the rug. Right. And they really have two choices in that moment. They can let go, risk falling, and practice those wobbly uncertain steps. Or they can drop to their knees and do that lightning fast baby crawl. Because they already know the crawl works perfectly. Yeah, the crawl guarantees they get to the toy.

0:34It's basically a calculated optimization. The crawl is safe. It has a 0 % failure rate. And it is highly efficient in terms of getting the reward, which in this case is the toy. Right. But if you continually allow them to choose the crawl or, you know, if you reward them equally for crawling versus walking, they never build the biological muscles to walk. They never learn the micro adjustments or the balance. Exactly. You have to force them into that wobbly, risky zone of trying if you actually want them to grow. So welcome to the deep dive. You're listening to this because the artificial intelligence systems we use every single day are, well, they're acting a lot more like that toddler than we might want to admit.

1:14They really are. Today, our mission is exploring some fascinating new research that attempts to answer this massive question in AI. Basically, when should an AI think for itself and when should it act by using a tool like a search engine or an API? Because the paradigm has completely shifted over the last couple of years, we used to view large language models purely as, you know, text generators. Right. You type in a prompt, it spits out a poem. Exactly. The isolated neural network just predicted the next word based on its internal weights. But today, these models are autonomous agents. They browse the web.

1:50Yeah. They execute Python code. They query databases. They interact with APIs to book your flights. They're constantly reaching out into the physical and digital world. And the core problem this research highlights is that we are currently training AI to just get the right answer. We treat external tools as, like, free shortcuts. Which is causing a huge issue. It is. It's causing AI to either overthink, overact, or fundamentally stunt its own intellectual growth. Okay, let's unpack this. Because the underlying framework proposed by the research, the theory of agent, it suggests we need to basically erase the traditional line between thinking and acting.

2:30Right. To understand why AI agents misuse tools, you have to look at how we normally design them. Traditionally, we separate reasoning from acting. Like humans do. Exactly. You think about a problem, you formulate a plan, and then you execute an action. But this unified framework collapses that entirely. It reframes both reasoning and acting simply as, well, alternative forms of knowledge acquisition. So they're just two different ways to get the same information. Yes. Reasoning is just using an internal cognitive tool. You're querying the AI's own parametric memory. It's own neural pathways. Right.

3:02While acting is using an external physical tool, like querying the outside world. Okay, so functionally, the model is treating its own neural weights as an API. That's a great way to put it. And that brings us to the knowledge boundary. The research formalizes this into two distinct sets. The internal and the world tasks. Yeah. First, you have the internal task set. This is everything the AI can solve purely by thinking, just routing through its own internal data. Got it. Then you have the world task set, things that absolutely require external interaction. The line separating those two sets is a knowledge boundary.

3:38And the key here is that it is different for every single AI model. So a state-of-the-art massive model has a huge internal task set compared to a tiny local model. Exactly. You know, it reminds me of taking an exam. Internal reasoning is like taking a closed book test. You're relying entirely on what you actually studied and memorized. Yes, perfectly said. But external acting is like raising your hand to ask the teacher for a hint or, you know, looking at an open book. That analogy perfectly captures the tension. And how the AI decides whether to take the closed book test or ask the teacher comes down to the currency of AI problem solving.

4:15Which the research calls epistemic effort. Right. Epistemic effort. It sounds super philosophical, but it's a strict mathematical reality. It's the total amount of cognitive work required to solve a task. To get the uncertainty down to zero. Exactly. And the invariant rule here is that epistemic effort cannot be eliminated. You can't just wave a wand and make it go away. No, the work has to be done. It can only be distributed between internal effort thinking and external effort using a tool. So wait, what does this all mean? I mean, on behalf of the listener here, if the total effort is the same and the AI gets the right answer, why do we care if it uses an external tool to do the heavy lifting?

4:53Oh, it matters immensely because relying on external tools for things the AI actually could solve internally is what the researchers call unnecessary delegation. Unnecessary delegation. Yeah. And it leads to catastrophic long-term consequences for the AI's intelligence. Which leads us right into the most alarming finding in all these lab nodes. How lazy tool use destroys learning. It completely undermines the AI's development. Think about how these models learn through rewards. When an AI uses a tool to solve a problem, it already had the internal capacity to solve. The tool acts as a reward shortcut.

5:31A reward shortcut, meaning it gets the points for being right, but it didn't do the mental work. Exactly. It gets a pat on the back for being correct, but those specific internal reasoning pathways receive zero reinforcement. So over time, the AI learns to associate delegation with success. Yes. Not reasoning. Its internal capabilities literally stagnate. It never builds its cognitive muscles because, well, because it keeps using a forklift to pick up a pebble. Wow. It's exactly like relying entirely on a GPS to navigate your own neighborhood. Oh, that's a great comparison. Like you always get to your destination, right?

6:05You achieve task success. But your own biological sense of direction completely atrophies. Because you outsource the effort, your brain never built the map. And if your phone dies, you're lost three blocks from your house. Exactly. And because AI agents are suffering from this exact GPS-like stagnation, they actually start exhibiting highly recognizable, almost dysfunctional personalities. The four behavioral regimes. The researchers map these out into quadrants based on how they balance that internal and external effort. The first one is high internal, high external. The brute force agent. Yeah.

6:42Highly inefficient and chaotic. It overthinks everything and frantically calls external tools over and over. Shooting through compute power. Massively. Then you have low internal, high external. That's our GPS user, right? Exactly, the lazy underthinker. It's highly dependent on outside systems and basically refuses to use its own brain. It'll search the web for a basic fact it already knows. Okay, then the third one. High internal, low external. The stubborn overthinker. But wait, isn't that third one, the stubborn overthinker, actually the goal? Like an independent, highly autonomous AI. What's fascinating here is, no, it's really not.

7:19Oh, really? Yeah, because if an AI tries to internally reason through a problem that is definitively outside its knowledge boundary, it leads directly to hallucinations. Oh, right. Because it literally doesn't have the data, but it refuses to look it up. Exactly. It drastically overestimates its own internal solvability. So it just gives you a confidently incorrect answer. Like a student writing a totally fictional essay to cover the fact that they didn't read the book. Precisely, which is why the fourth regime is the actual holy grail here. Low internal, low external. Low and low. That sounds like it's doing nothing.

7:53It sounds like it, but it means the agent is epistemically calibrated. It uses exactly enough internal thought to process what it knows, and it only asks for outside help when it's epistemically necessary. And it actually crosses its knowledge boundary. Right. It's perfectly balanced. Okay. So to see this theory in action, we can actually look at one of the biggest debates in tech right now. The whole long context models versus retrieval augmented generation, or RRAG. Oh, absolutely. This framework completely relocates that debate. You can view long context as what the research calls preemptive context expansion.

8:30Meaning you're stuffing all the documents into the AI's internal memory before you even ask the question. Right. You're moving a massive task into the AI's internal task set ahead of time. whereas ARG retrieval augmented generation is on-demand external help. It fetches data only when asked. So ARG is the external tool. Exactly, and the research takes a very clear normative stance here. If a model can solve something via internal long-context reasoning, it absolutely should. Because it forces it to build those cognitive muscles. Yes. ARG-ray should strictly be reserved for things that are unpredictable, like live time-varying data or data sets that are just too mathematically massive to fit in context.

9:11So does this mean developers are essentially building Argex systems as a crutch? Like, just because their AI's internal knowledge boundary isn't big enough yet? Honestly, yes. The hardware limits the context window, so engineers externalize the effort to save compute. But defaulting to RAG for everything trains the model to permanently be that lazy under-thinker. Because it never learns to synthesize vast amounts of complex data on its own. Precisely. Okay, so if we know what the ideal, perfectly calibrated AI looks like, and we know we shouldn't just rely on RAG as a crutch, how do we actually train these models to behave correctly?

9:50Well, the solution proposed in the papers is metacognition. We have to teach the AI to assess its own internal solvability before acting. Teaching an AI to know what it doesn't know? Yes. The research outlines specific training pathways for this. First off, we can't just reward correctness anymore. Right. No more reward shortcuts. Exactly. We need reinforcement learning that actively penalizes an AI for using a tool when it didn't actually need to. We have to make unnecessary delegation mathematically painful. Yes. And during pre-training, we have to treat next tool prediction as seriously as we treat next word prediction.

10:24Here's where it gets really interesting. We are essentially trying to teach machines the concept of self-awareness and self-doubt. That's exactly what it is. The ultimate goal is an AI that knows the exact perimeter of its own ignorance. Yes, an AI that stops and says, wait, my internal entropy on this topic is too high. Now I will use a search engine. That is wild. Okay, so just to summarize this journey for you listening, we've learned that AI tool use is not a free pass. Epistemic effort is a strict mathematical currency. It has to be spent. Right. And letting AI be lazy by overusing external tools will ultimately stunt its intelligence and trap it in a state of dependency.

11:05And I want to leave you with a final thought to mull over based on where these lab notes say the research is heading next. Oh, yeah, the future directions. Right. Because we spent this whole time talking about AI navigating digital tools, APIs, search engines. but what happens when this exact framework is applied to embodied agents? You mean physical robots? Yes. If a humanoid robot's external tool is its own physical body interacting with the real world, how will it calculate the epistemic effort? Oh, wow. Like walking across a room to look at a label versus trying to guess what's inside a box?

11:41Exactly. The knowledge boundary is about to enter the physical world. The calculation suddenly involves battery power, physical friction, the time it takes to walk 50 feet. So if it's a lazy underthinker, it might just guess what's in the box and drop something fragile. And if it's an overworker, it might spend its whole battery pacing back and forth checking labels. That completely changes the stakes. Well, thank you so much for joining us on this deep dive. Keep questioning the systems you interact with every day, and we will catch you next time.

From the publisher

This position paper discusess Theory of Agent (ToA), a framework that redefines large language model agents as decision-makers who must choose between internal reasoning and external tool use. The authors argue that agents should only invoke external tools when epistemically necessary, meaning the task cannot be reliably solved using the model's existing internal knowledge and logic. This perspective addresses common failures like overthinking and overacting, which occur when an agent's internal solvability estimates are poorly calibrated. By treating reasoning and acting as co-equal methods for reducing uncertainty, the framework highlights that unnecessary delegation to tools can stagnate the growth of an agent's internal intelligence. Ultimately, the research suggests that alignment should be measured by how effectively an agent allocates epistemic effort rather than just achieving a correct answer. These principles offer a new trajectory for training and evaluating agents to ensure they become more autonomous and efficient over time.

More from Best AI papers explained

All 475 episodes
Position: Agents Should Invoke External Tools ONLY When Epistemically NecessaryBest AI papers explained · 12 min
Listen in VO