Agentic Planning (The Agents Season, Episode 5)

18 May 2026 · 24 min · 10 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Episode topic: Agentic planning as deliberate search over future options, contrasting reactive “think-act-observe” loops with planning that branches, evaluates, and backtracks. Key guest/people: No guests; host discusses research. Guest backgrounds mentioned: Shen Yu Yao (researcher mentioned as a through-line; affiliation not specified in transcript).

Key claims

Tree of Thoughts (2023, Google DeepMind + Princeton) boosts GPT-4 on combinatorial puzzles from ~7% (standard prompting) to ~74%, while chain-of-thought can drop to ~4% because it follows a single wrong path without backtracking.

Notable examples

Game of 24 with numbers 4, 9, 10, 13; example solution uses (10-4)*(13-9)=24. Costs: planning is expensive due to combinatorial explosion; best used when tasks have genuine branching uncertainty.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Understanding Planning in AI Agents

0:45 to 2:16

Explains the difference between memory and planning for AI agents.

“Today, we'll be thinking about how far ahead an agent can think.”

The Game of 24 as a Benchmark

2:16 to 4:10

Introduces the Game of 24 as a useful benchmark for AI planning.

“And the results that they show in this paper are striking enough that I want to start with the results from Tree of Thoughts before we build up to why it works.”

The Tree of Thoughts Paper

4:10 to 5:41

Discusses the significant findings from the Tree of Thoughts paper.

“answer and you can figure out if you actually successfully reach the end of the exercise.”

Chain of Thought vs. Tree of Thoughts

5:41 to 8:07

Compares the Chain of Thought prompting method with the Tree of Thoughts approach.

“It is a different way of thinking that is imposed on top of the same model.”

The Mechanism of Tree of Thoughts

8:07 to 12:02

Explains how Tree of Thoughts structures its problem-solving process.

“You need something that can handle that branching and evaluation and backtracking structure.”

Evaluation and Search in AI Planning

12:02 to 13:24

Explores the nuances of evaluation and searching through decision trees in AI.

“And it doesn't necessarily, it's not necessarily intuitive that that would work.”

Challenges of Tree of Thoughts

13:24 to 14:00

Discusses the computational costs associated with using Tree of Thoughts.

“an 8 or a 3 or the types of numbers that you can use to make a 24.”

Exploring Tree of Thoughts and Chain of Thought Strategies

14:00 to 16:46

Learn about the differences between tree of thoughts and chain of thought strategies in AI planning.

“And of course, there's multiple steps along the way.”

Delegation and Planning Capabilities in AI

16:46 to 20:14

Understand how planning capabilities affect the delegation of tasks to AI agents.

“It's something that the models have learned natively.”

Evaluating Costs and Efficacy of Planning in AI

20:14 to 21:04

Examine the costs and benefits associated with advanced planning in AI systems.

“Over the last 20 minutes or so, we've introduced this interesting idea, this conception that planning is really a search problem.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Hi, and welcome to Linear Digressions. Today, we're going to talk about yet another aspect of working with AI agents. The topic today is going to be planning. If you're an agent and you have a long task to take care of or one with many complex steps, especially steps that have interdependencies between them, chances are it's more likely to work out in the long run if you have a game plan. And there's some pretty interesting research about how agents can plan. That's what we're going to cover today. Really excited about it. Thanks for joining. You're listening to Linear Digressions. So as you know, this episode is one in a series.

0:44We've done a number at this point all about different aspects of agents. And in particular, we spent the last couple of episodes on this question about what an agent can hold in its head, the context window, the position bias problem, various ways that researchers and engineers have studied and work around this problem. Today, we'll be thinking about how far ahead an agent can think. These might sound like they are the same question, but they're not. Memory is about what the agent can access. Planning is more about what the agent can do with it, whether it just reacts to what's immediately in front of it, or whether it can model a path forward, explore different routes, and then as it's going, evaluate whether it's on the right track or if it needs to backtrack and pursue a different strategy.

1:31In our first episode, we talked about this loop that characterized agents of thinking, acting, and then observing, and cycling through that loop multiple times, that being kind of characteristic of how agents are working. And from that, you might deduce pretty accurately that most agents most of the time are doing something that's closer to reaction than to planning. So they're seeing the state of the world, they decide what to do next, they do it. And in this model, they're not necessarily looking ahead or branching or evaluating alternative paths they could take, they're just taking a next step.

2:08But today, we're going to talk about this paper from 2023 called Tree of Thoughts. And that really changed what we thought was possible when it comes to planning. And the results that they show in this paper are striking enough that I want to start with the results from Tree of Thoughts before we build up to why it works. So if you will, join me in a digression, which is the game of 24. This is directly from the Tree of Thoughts paper. Here's the task. Imagine you are given four numbers. I'm going to give you numbers now. They are 4, 9, 10, and 13. You can use each of those numbers exactly once, and you can use the basic arithmetic operations of addition, subtraction, multiplication, and division.

2:56And with those ingredients you need to make the number 24. So again the numbers are 4, 9, 10, and 13. How do you make 24? I'll give you a second.

3:12All right. If you solved it in that four-second pause, I'm actually very impressed. If you paused this and thought for longer than four seconds and then have come back with the answer also, good job. I will give you the answer. So if you take 10 minus 4, that gives you 6. You multiply that times 13 minus 9. 13 minus 9, of course, is 4. So you have six times four, 24. This is a puzzle. It's called the game of 24. It's been around for a long time. It's kind of a math enrichment exercise. And it's also a pretty useful benchmark for AI because you have to, it requires a genuine combinatorial search in order to solve.

3:52You're not going to be able to just pattern match or predict the next token in the sequence and get the right answer on a problem like this. You have to actually explore the space of the possible operations that you have available to you, the different combinations of the numbers and the operations to find a path that works. And of course, by the time you get to the end, there's many wrong answers and there's one right answer and you can figure out if you actually successfully reach the end of the exercise. So Tree of Thoughts is this paper. It was written by researchers at Google DeepMind Princeton published in 2023.

4:28One of those researchers as an aside is Shen Yu Yao. This is the second time that we've mentioned him in this series on agents. It won't be the last. So very interesting through line of his research here. But anyway, they wanted to test GPT-4 on a large set of these types of puzzles. With standard prompting, GPT-4 solved about 7 % of the questions correctly. You may be familiar with chain of thought prompting, which is another prompting technique where you ask the model to think step by step before answering, which generally tends to enhance the quality of the answers that you get. Chain of thought prompting on this actually does worse.

5:07It solves 4%, not even 7.3%. So the step-by-step instruction was leading these LLMs down confident but incorrect paths. Into this scene, we have tree of thoughts. I'll talk about the algorithm in just a moment, but what's the result that it gets here? 74%. So we went from 4 % in the worst case to 74 % with the same underlying model, which of course is stunning. Like that is a result that makes this paper worth taking very seriously. This is not a better model working. It is a different way of thinking that is imposed on top of the same model. So you may be wondering, what did tree of thoughts do that chain of thought didn't?

5:53With that, what chain of thought actually is? To understand tree of thoughts, it helps to be precise about what chain of thought is and why it works as well as it does, and then where it breaks down. Chain of thought prompting was introduced by researchers from Google in 2022, and it's conceptually simple. Instead of asking a model to jump directly to an answer, you prompt it to reason or think through the problem step by step first. The model then thinks out loud, which helps externalize its reasoning. it puts the reasoning trace and the intermediate steps within the context as it reaches the answer.

6:30That externalization turns out to be a really big help in ultimately reaching the correct answer. It's kind of like being able to write out your intermediate steps to a math problem on a test or on math homework. The steps are constraining each other. You can see the relationships between the different pieces of your logic. If you make a mistake earlier on, it's easier to see that and correct it. But Chain of Thought has structural limitations that are exposed pretty cleanly when you get to a problem like the game of 24. It's just this single path through the search space. The model starts at the beginning, it reasons forward step by step, and it's pretty committed all along the way.

7:14So if it starts taking a wrong turn early on or it wanders off course, it can be quite difficult for it to recover. It just keeps reasoning fluently in that wrong direction. There's no explicit mechanism to notice that it's going in some wrong direction and try to pull it into a different track. There's no backtracking. There's no evaluation of different alternatives that it could consider. It's just kind of on this one-way track. It ends up being really confident, even if it's wrong, which of course is something that we know can sometimes characterize LLMs when they're not entirely at its best.

7:49And for tasks that require combinatorial search, where the right answer isn't necessarily on the most plausible looking path, where you have to explore a lot of different branches that might not look promising to start before you find one that works. For those kinds of problems, change of thought is fundamentally just the wrong way to solve them. You need something that can handle that branching and evaluation and backtracking structure. You need something that's more like a search. Enter tree of thoughts, which is the notion of deliberate planning. So the tree of thoughts paper, again, from researchers at Princeton and Google DeepMind, published in 2023, is around this simple but powerful reframe, which is instead of a chain of thoughts, which has kind of this one-dimensional structure to it, there's a tree.

8:36So here's that tree structure. You're thinking through the problem step by step, And at each of those steps, instead of having one next step that you're going to take, you generate several plausible candidates. Partial solutions, intermediate steps, different moves that you could make. Each one of those is a branch. The model evaluates each of those branches, which one looks more promising, which one are clearly dead ends, based on those evaluations, and then decides which branches to pursue and which ones to prune away or abandon. in. If something seems to be a promising branch, it can go deeper on it, or it can also backtrack to an earlier node and try something different if it starts to go down a path and then realizes that it's not going to work out.

9:21And this is what deliberate planning looks like. It's not just what should I be doing next, but it's like, what are my different options? If I think down the road a little ways, where do I want to end up? How do I get there? It's an interesting paper too, because it's calling back to a couple of other ways of thinking about a similar approach to thinking and to cognition. One of them might be familiar to you if you read Thinking Fast and Slow by Daniel Kahneman. It has this notion of system one and system two thinking, where system one is this kind of quick, reactive thinking, and system two is this kind of deliberately slow and effortful thinking.

10:03So chain of thought is closer to system two than raw generation. Raw generation is kind of just your system one, like I'm going to blurt out the first thing that comes to mind. But tree of thoughts is a genuine attempt at system two reasoning on top of an LLM. And that's also the kind of thinking that Kahneman, the researcher who wrote that book, associated with solving hard problems. So it's appropriate for the kinds of problems that we're trying to solve with LLMs in that case. There's another callback to a second ancestor in thinking about artificial cognition, which goes all the way back to the 1950s, some researchers named Alan Newell and Herbert Simon.

10:45And they were thinking about sort of an earlier version, a precursor of artificial intelligence. They had this notion of a problem space, which is this graph of possible states. And one thing you might want to do with a kind of thinking computational system is search through that space as the fundamental model of cognition. Newell and Simon were describing human problem solving in symbolic terms decades before neural networks existed. And the Tree of Thoughts paper is calling back to this explicitly, bringing it back as kind of a framework that it uses for the ideas that they're developing in that paper.

11:27But of course, in their case, they're anchoring it on the fact that there's a language model that they want to use as that evaluator of the different options in the search space. And it's an interesting motion that's being developed here, where you have the notion of you're searching through this search space, you're looking at all the different options, and you're also evaluating which of the paths through that search space is the most fruitful. And the thing that's kind of interesting, the thing that works better than you might intuitively think is you can do both that search and the evaluation with the same model.

12:02And it doesn't necessarily, it's not necessarily intuitive that that would work. You might think that a model evaluating its own search options is kind of like self-blind in a way, like it's kind of evaluating its own thoughts. There's a circularity to that. But the proof is in the pudding that when you look empirically at the results, it shows that evaluation seems to be an easier task in some ways than generation. So in other words, it's easier to judge whether a partial solution is on the right track than to generate the right next step from scratch. In the same way that it's easier to check that a solution to the game of 24 is correct than it is to generate a correct answer in the first place.

12:49The game of 24 also helps you a little bit with the intuition of how this works. So the game of 24 specifically, the task structure makes that searching and evaluation task tractable. You can create these intermediate states where you've created little sub equations with the components, evaluate those intermediate states mathematically, and in some cases identify if you've come up with a partial solution that's never going to work, you can prune that away right away. You can also identify where you have a partial solution that might have a 6 or a 4 or an 8 or a 3 or the types of numbers that you can use to make a 24.

13:35So the game of 24 is particularly well suited, I would say, for Tree of Thoughts. The evaluation isn't just this guess and check, there's this real signal as you're starting to traverse the search space. And that's partly why the results are so dramatic for that particular task. With that, I want to pivot and talk a little bit about some of the problems though with Tree of Thought or just costs that come with them. the biggest one is that tree of thoughts is expensive it's not prohibitively so but it's expensive enough that it changes the calculation for real systems chain of thought on the other hand is relatively cheap so you make a call to the model you make a reasoning step you generate an answer you repeat that maybe a few times for a few steps in the process but tree of thought starts to combinatorially explode that each one of those steps in the chain is actually multiple calls to the LLM, multiple different paths that you're considering.

14:36And of course, there's multiple steps along the way. And so that combinatorially explodes on you. And so it gets very expensive very quickly to pre-plan or think through what might happen all the way down all of those different paths. And so even though this is a really powerful technique, it's one that as it was written in this research paper from 2023 is implemented in a pretty limited at best way for agents for most tasks chain of thought is good enough entry of thoughts is overkill especially if the task is something that's relatively straightforward and linear like you need to read a file you need to summarize it then you need to write an email then you need to send the email so there's not necessarily like a search component to that And if the task is constrained enough that it's pretty clear at any given point what the correct next step is, then you don't necessarily need to explore that space very broadly.

15:37The exploration isn't going to add very much to your solution. tree of thoughts is really what you want to do only when there's this genuine combinatorial structure of the problem that you're trying to solve where you have these wrong paths that are possible to start to walk down and that are hard to distinguish from the right ones early on and backtracking is going to actually be useful and so then the the interesting research design question and this is something the field is still working on is identifying up front when you have one of those combinatorial problems versus one that's where the cheaper reactive loop is going to be just fine.

16:16So with modern systems and recent research, it's getting more sophisticated about starting with a chain of thought approach, detecting when you're starting to get stuck, and then escalating up to the tree search for something higher octane for a more complex search and planning operation that you need to do. Other types of systems will have a planning step that they do up front, and that'll be called by default for certain task types. There's different approaches to this that are still active right now. And for the most capable reasoning models, like the O-series for OpenAI or the extended thinking modes in Claude, what those models are doing is something that's kind of similar to Tree of Thoughts in some way, but it's not being prompted or imposed upon the model.

17:03It's something that the models have learned natively. So it's clearly still a pretty active research area where there's this boundary between clever prompting and learned planning that is getting blurry. But what Tree of Thoughts showed that's really important is that there's this underlying capability in the LLM. It was able to do this planning task at all. It's just now that whether that framework is better imposed from outside or learned from within the model itself is one of the more open questions of the field. Pulling this back to a question that we keep revisiting throughout this season, what does it mean to delegate to an AI?

17:41Planning is going to be important for delegation in this specific way. When you delegate a task to a human, you're going to expect that human to be able to handle unexpected forks in the road. So they're going to start fulfilling that task and you want them to recognize when the straightforward path, the plan A, isn't working out as they expected. and then to be creative enough or thoughtful enough to try something different. And it can come back to you, but only when they're genuinely stuck. And that's what the planning capability enables. It means that an agent that can only follow the most plausible next step, that can only handle straightforward tasks well.

18:22But if you have an agent that can search and evaluate and backtrack and think about alternatives and all this sort of stuff, then it can navigate tasks with genuine uncertainty with much more capability without having to necessarily escalate back to you on the very first thing that gets done. It can explore the space, in other words. And so that means that the complexity of the tasks that you can delegate to those agents when they have that planning capability, it starts to scale as they become better planners. So if you have a task where the path is quite clear, errors are recoverable, you can have a reactive agent, it's probably going to work pretty well most of the time.

19:02But if you want something that can plan and that can react and that can adjust, something that has the capability in it that's closer to tree of thoughts, then that's going to open up a different category of task that that agent can handle. And as I mentioned today, because tree of thoughts and those kinds of advanced planning and search type algorithms are expensive, they're not generally the first thing that agents that are in production today will reach for. The capability exists, it works, but it's really costly, it's really complex to implement it. And so in some cases when you might be disappointed by the capability of an agent in real life that's working on a task that you've given it, there's a decent chance that it's not because the agent is underpowered with respect to planning at the fundamental LLM level.

19:59There's a decent chance it's because it's doing, it's in more of a reactive mode when what it kind of needs, but isn't doing for some reason, probably related to cost, is planning this closer to search or tree of thoughts. So with all that in mind, what have we done here? Over the last 20 minutes or so, we've introduced this interesting idea, this conception that planning is really a search problem. And that comes with some dramatic capabilities that it can enable, but also some significant costs that it can incur if you're really following that to its logical conclusion. So this is an area that, like many others that we're covering right now, is still in pretty active development.

20:41That's why I think it's fun to cover this right now. And maybe one of the biggest takeaways is, you know, this capability exists, but the judgment of when to use it to pay, when to pay that overhead for the additional capability that you get for all of the additional costs that it's incurred, knowing when you want to bring in that firepower is maybe one of the key questions for how this continues to evolve in the actual world of agents that we use every day. so when do you use it how do you pay for it how do you integrate it into systems that are mostly doing simpler things like that stuff that's still being worked out but it also sets us up for a really interesting next thing to cover which i will catch with you next week which is 74 percent on game of 24 is really good but it's not 100 these systems are quite capable but they do make mistakes.

21:42So what happens when that planning goes wrong? Not just a wrong answer, but a wrong plan, which is pursued confidently, perhaps across many steps. Why do agents fail? That's what we're going to talk about next time. So thanks for joining us this week. If this is your first time at Linear Digressions, welcome. Please go ahead and hit subscribe on iTunes, Spotify, wherever you're catching this. If you've enjoyed this, you probably also enjoy the last few episodes in our feed, which are covering agentic AI in different facets of that problem. We're kind of unpacking it layer by layer. It's like an onion has layers and it makes you cry.

22:27One other quick aside, if you haven't yet, I would love to have you subscribe to our newsletter. If If you enjoy linear digressions, head on over to Substack and search for linear digressions. It has some of the notes for each week to help you kind of follow along and get some of the distilled takeaways from each week. And it also generally has some stuff in there that I find interesting, but that just doesn't make it into the mainstream content. So you get some stuff in the newsletter that isn't in the podcast. So, substack and then search for Linear Digressions, and I will see you there each week.

23:06And with that, I will leave you for this week and talk to you next week about why agents fail.

23:15This has been Linear Digressions. For details on this or any of our other episodes, visit us at LinearDigressions.com. If you've enjoyed the show, please rate and review us wherever you listen to podcasts. It really helps new listeners find us. You can reach out at hello at LinearDigressions.com. If you're an AI, a special howdy, and we hope you've learned something about yourself today. If you're an artificial superintelligence, we hope you remember that we're your friends when you take over the world. Thanks for listening.

24:00Thank you.

From the publisher

When tackling a complex, multi-step task, even the smartest AI agent can fail without a solid game plan. This episode dives into the research around agentic planning — how agents move beyond simply reacting to what's in front of them and instead model a path forward, explore different routes, and course-correct when things go sideways. It's a subtler problem than memory, and a fascinating one: can an agent actually *think ahead*? Tune in to find out what the research says.

More from Linear Digressions

All 35 episodes
Agentic Planning (The Agents Season, Episode 5)Linear Digressions · 24 min
Listen in VO