ReAct and Tool Usage (The Agents Season, Episode 2)

27 Apr 2026 · 24 min · 7 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

How AI agents gained “tool use” after the 2022–2023 shift, focusing on the REACT loop (reason-act-observe) and Toolformer (learning when to call tools), plus MCP (a protocol to connect agents to external tools). It also motivates the ideas with OpenClaw, a consumer personal agent.

Guests

No guests mentioned; the host speaks throughout.

Guest backgrounds

N/A.

Key claims

Tool use removes the pre-2022 “wall” where models couldn’t look up or verify facts. REACT interleaves reasoning traces with actions and observations, improving multi-hop QA and interactive tasks and making trajectories interpretable. Toolformer trains models to decide tool calls without explicit instructions, using synthetic data with candidate API calls. MCP standardizes tool integrations (USB-C analogy) and has broad adoption since Anthropic’s Nov 2024 launch.

Notable examples

HotpotQA Apple Remote → REACT uses web/search observations to answer “keyboard function keys.” OpenClaw (formerly ClawedBot) runs locally, integrates skills for Gmail/GitHub, and can act persistently on a user’s behalf.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

The Shift in AI Tool Use

2:45 to 4:41

Discover how the interaction with AI changed from 2022 to 2023, allowing AIs to use tools.

“All right, so if you're listening to this in 2026 or later, it may not be clear to you why tool use was even such a big deal at all.”

Understanding the REACT Framework

4:41 to 7:30

Explore the REACT framework that combines reasoning and action in AI.

“And so now what we can see with the fullness of hindsight is that, well, what happens if you interleave them?”

Example of REACT in Action

7:30 to 11:10

Listen to a detailed example illustrating how the REACT framework processes questions.

“So it's picking up somewhere that there's some sort of like media software hook in here that's maybe related to the Apple remote.”

Toolformer: Advancing Tool Use

11:10 to 14:00

Learn about Toolformer and how it enables AIs to autonomously determine when to use tools.

“it might be getting from the environment.”

Understanding React and Toolformer in AI

14:00 to 17:42

Explore the concepts of React and Toolformer and their roles in AI tool usage.

“So let's take those two concepts of React and Toolformer and just reflect on that for a second because they're solving adjacent but different problems.”

The Rise of OpenClaw: A Case Study

17:42 to 20:09

Learn about OpenClaw, a rapidly growing personal AI agent and its significance.

“So we've been talking in this episode about things like papers and benchmarks and what was going on in 2022, and that's great.”

The Implications of Tool Use in AI

20:09 to 21:53

Discuss the risks and responsibilities involved with AI agents' tool usage.

“these conversations so we have these ideas in react and tool former like the interleaved reasoning and acting the learned tool selection this agentic loop and these weren't just research curiosities.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Hi, welcome to Linear Digressions. This is the second episode in our season all around AI agents. If you're joining for the first time and you want to just dive right in, welcome. This is going to be a good one. But if you're a completionist and you want to go back to the beginning, good news. There's only one episode for you to get all caught up. That was our last week's episode where we were talking about what even is an AI agent. Because depending on who you ask, the answer may vary. But today, we're going to back up to roughly 2022 to 2023, when AI really had an inflection point. Prior to that era, there had always been this wall between you and the AI that you were interacting with.

0:49Of course, there were AIs that went back before the pre-Chat GPT era, going all the way back to the 60s and 70s. And then in different models and different instantiations, there are various approaches and implementations. But in that pre-2022 era, there was always kind of this fundamental limitation where you were on one side, the model was on the other. You would put text in and you would get text back. And depending on which model you were working with, sometimes those models could be very capable. But as soon as you went outside of it, as soon as you looked around and you said, how how has the actual world changed as a result of my conversation with the AI?

1:32It became very clear that there was a wall there. So the AI couldn't look anything up. It couldn't run any code. It couldn't check whether the thing it told you was still true. It could only draw on whatever it had learned during its training process and generate a response. And so today we're going to talk about the time when that changed. In 2022 to 2023, that wall between the models and the outside world came down. And it's not because the models got smarter in some abstract sense, but it's because researchers figured out how to let them reach outside of themselves. That was the era in which AI began to learn how to use tools.

2:13So in this episode, we're going to talk about how that happened, a couple of key papers and what those actually showed, and what it unlocked. So tool use really is a hinge here. It's a thing that turns a very sophisticated text generator into something that can act in the world. And remember, that's a core part of what we're saying makes AI agents different from just your off-the-shelf AI. And so understanding how tool use works and what makes it hard is going to be a foundation for almost everything else that we cover this season. You are listening to Linear Digressions. All right, so if you're listening to this in 2026 or later, it may not be clear to you why tool use was even such a big deal at all.

2:58So let's back up. What actually was the problem with tool use? In 2022, large language models were quite impressive when it came to reasoning. There was this idea of chain of thought prompting, which is where you ask a model to think step by step before it answers. And that had shown that there was a lot of latent capability by just letting the model think out loud, externalize this reasoning process. So models were solving math problems and logic puzzles. They were answering these multi-step questions. They were passing professional exams. But all of that reasoning was happening inside the model's head, so to speak.

3:34So it was reasoning about the world using whatever the world looked like when it finished its training process, so at the end of its training data. So it couldn't look things up. It couldn't verify claims. It couldn't see anything that happened after the knowledge cutoff of its training data. And so it has this internal monologue. That's great. But that was disconnected from reality in this very specific way, namely that it could reason about what was happening, but it couldn't check that reasoning. Now, meanwhile, separately, there was a research program about getting models to act. So can we teach models to navigate environments, browse the web, control interfaces?

4:17But those systems mostly just generated actions. They didn't have much reasoning behind them. So they'd click around and they'd do things, but they didn't really think about why. They just fail in these brittle ways when things didn't go as expected. So what you're hearing now might be two different pieces. reasoning on one side, acting on the other. But these research programs weren't talking to each other before this 2022-2023 era. And so now what we can see with the fullness of hindsight is that, well, what happens if you interleave them? What if instead of having the model reason and then act, or act without reasoning, you had it reason and then act and then observe what happened, then reason again in light of what it learned, so on and so on.

5:04In this cycle, that's the react pattern. And it sounds almost obvious in retrospect, the way that all good ideas do. Let's talk about react. The react paper is not actually called react. It's called synergizing reasoning and acting in language models. Came out of Google research in late 2022. to. The first author, as an aside, is a researcher named Shen Yu Yao, who I'll actually be mentioning his research several times throughout this season. It's published at ICLR, which is one of the top venues for machine learning. But the core idea is right there in the name. So REACT, it's actually a combination of two things, reasoning, RE, and acting, act, REACT.

5:48And the The consequence of putting those two pieces together is it's simple, but it has a big impact. What React does is it says, what happens if you let a language model alternate between generating a reasoning trace and taking an action? So you're interleaving these two. You're not doing them sequentially. What they found was that that's doing something much more capable and much more robust than either one alone. Let me give you an example here to give you a flavor. So one question from one of their eval suites, this was from a benchmark called Hotspot QA. Here's the question that it asked the AI to answer.

6:24Aside from the Apple remote, what other device can control the program Apple remote was originally designed to interact with? Here are some of the things that the model says. So the first example answer, which is wrong, is iPod. The second example, which is also wrong using chain of thought or reasoning. here's what the model says in responses it starts thinking it says let's think step by step so apple remote was originally designed to interact with apple tv apple tv can be trolled by controlled by iphone ipad and ipod touch so the answer is iphone ipad and ipod touch turns out that is wrong now here's a third example that's action only so act one is to search for apple remote and it gets an observation back that says apple remote is a remote control.

7:11The second thing that it does is it searches for front row. I don't exactly know why. It observes in response to this that it could not find front row. It takes a third action, which is searching for front row in parentheses software. Okay, so now it gets an observation back. The front row is a discontinued media center software. So it's picking up somewhere that there's some sort of like media software hook in here that's maybe related to the Apple remote. Remember, this is all just getting generated kind of as a token completion in response to what it had in the prompt, which was about Apple remote.

7:49But anyway, act four, which is the end of the answer here is finish with the answer of yes. The answer of yes is not the correct answer to what other device besides the Apple remote can be controlled by an Apple remote. Here's what React does with this. It says, thought one, I need to search Apple remote and find the program it was originally designed to interact with. Then it takes an action, which is to search Apple remote. It receives back information, makes the following observations. The Apple remote is a remote control introduced in October, 2005 by Apple. It was originally designed to control the front row media center program.

8:28Ah, so now we understand where front row is coming from. There's some sort of word association between Apple remote and front row. And so now it's made that connection. Now it goes into thought number two. We're in this loop. So Apple Remote was originally designed to control the Front Row Media Center program. So I need to search Front Row next and find out what other device can control it. So now it's going to search for Front Row. Now remember, that didn't work in the acting only paradigm. And it doesn't work here either. So it says couldn't find Front Row. It's maybe a little bit too ambiguous of a search.

9:02but maybe I need to search for front row software. So thought three, front row is not found. I need to search front row software. Now it does this refined search for front row software. Now it finds this observation, front row is a discontinued media center software, blah, blah, blah, blah. So it has now the fourth thought in the loop. Front row software is controlled by an Apple remote or the keyboard function keys. So the answer is keyboard function keys. And then the final act, the finishing act is keyboard function keys. This is the correct answer. So it took both this web search aspect of going into the front row software manual and learning a little bit more about it, but also thinking through the relationship between that and the question that's being asked in order to get the correct answer.

9:51Okay, so that was a lot of steps that I just made you listen to me reason through. And in general, there can be a lot of steps for some of these more complex queries. But the point isn't the trivia answer of the exact way that you can use front row software without an Apple remote. The point is that the model is grounding its reasoning in reality at each step. So it's going back to the web each time and trying to search and get confirmation that it's on the right track and or nudge itself back on track if it doesn't think that it's getting closer. It's not just confidently asserting something that it might have hallucinated.

10:25It's going in and checking, and then that checking dramatically changes the reliability profile of the answer that you get. And so what this REACT paper actually showed concretely is that on question answering tasks that required this multi-hop reasoning, so you find facts that connect to other facts in order to get the answer that you need, REACT significantly outperform chain of thought alone by reducing hallucination and error propagation. So it's not just that the model's reasoning better, it was reasoning about real information that it had actually retrieved. And then when they took it to interactive decision-making tasks, where it's doing things like navigating virtual environments or completing shopping tasks, it outperformed action-only approaches by a lot because it could actually think about what it was doing rather than just acting blindly and ignoring any feedback that it might be getting from the environment.

11:17There's one other thing here that the paper showed, which matters a lot for practical systems, which is that those trajectories, that cycle of the model thinking out loud, going out, performing some action, retrieving results, using those observations to iterate through the loop again, those trajectories are interpretable. So you can read that thought-action-observation sequence and see exactly what the model was thinking at each step. And so when it goes wrong, because it does go wrong, you have some way to diagnose why. So now with React, we've established this really important loop, the think, act, observe.

11:54And that loop is key for calling something an agent rather than just something that's responding. So React is the paper that made that loop concrete and showed that it could work. There's a second really important concept to cover now, which is tool use. This too has a paper that we will talk about, which is Toolformer. It came out of Meta AI in early 2023. So if React is about how to use tools in the moment, what Toolformer is trying to grapple with is, can you teach a model to know when to reach for a tool in the first place without being told? So the point here is that instead of prompting a model with instructions about when to use tools, they actually trained the model to figure this out on its own.

12:40So in Toolformer, it wasn't given explicit instructions about tool usage. Instead, this was taken through a pretty clever training process. They generated this huge data set of text with potential API calls annotated. They filtered for the ones, the API calls that actually helped the model predict what came next. And then they fine-tuned the model on those API calls. So what the model is doing is learning how to decide for itself when calling a calculator or a search engine or a calendar would improve its output. And those were the tools that they used. They had a calculator, question answering system, a couple of search engines, a translation system, a calendar.

13:19Nothing fancy, just things that you'd actually want to do kind of day-to-day work. And what they showed with Toolformer is that they could take a relatively small model that's been fine-tuned for this tool usage, equip them with those tools, and then it could be competitive with a much larger model on a variety of tasks. And that's not necessarily because it's smarter, but because it knew when to offload the hard parts. So needs to do math? Go to the calculator. Need to look up a fact? Go to search. You want to do some kind of reasoning about the date? Go to the calendar. So in effect, the model has, through this training process, learned its own limitations and has found routes around them.

14:00So let's take those two concepts of React and Toolformer and just reflect on that for a second because they're solving adjacent but different problems. React is about the architecture of reasoning and acting, so the interleaving between reasoning, acting, observing the results, repeating that cycle. Tool former is about learning tool use, so when do you reach for a tool, not just how to reach for a tool once you've decided. So in practice, modern agent systems combine both ideas. They're designed with a React-style loop, and the underlying model has been trained to develop good intuitions about tool selection.

14:38So in practice, we have both of those working together in modern systems. What does this look like today? Well, there's one more piece that I wanted to name drop real quick, which is MCP or model context protocol. You've almost certainly heard of it. And if you're anything like most people, you are nodding along while you're quietly wondering what it actually means. So what is MCP? So you react in Toolformer that have solved this research problem of how an agent could use tools. But what MCP tries to solve is the engineering problem of how to actually connect an agent to the tools that exist in the world.

15:13So that can include your calendar, your email, your company's database, your code repositories. So every time someone wanted to connect an agent to a new tool, they had to build this custom integration from scratch. And every time a new model came out, those integrations might break. So you have N models times M tools, N times M different bespoke connectors that you might have to build. So that space starts to combinatorially explode. And you're just going to have a really hard time keeping the models and the tools in sync with each other. So onto the scene comes MCP. It was a protocol that was announced by Anthropic in November 2024.

15:55and it's an open standard for connecting AI assistants to data systems. So basically, it's a way of standardizing AI tool use. And this includes all different kinds of tools. This can include content repository, business tools, development environments. And the goal is to replace that fragmented approach of point-to-point connections with something that's a more universal protocol. There's an analogy that's used a lot, which is that MCP is like a USB-C port for AI. So it's this one standard connector. I have a USB-C cable that connects my microphone. I can connect it to my laptop. I can connect it to my desktop.

16:32I can connect it to the charger that I use for my phone, although I don't know why I would do that because I don't charge my microphone because that's not how it works. But anyway, the point is all of these use a USB-C connector. It's kind of like that, but between agents and tools. So MCP, since it was launched in November of 2024, has had this very rapid adoption. And at this point, it's basically an industry standard where even competitors like OpenAI and Google DeepMind have adopted it, and it's now a generally accepted standard in the field. Now, MCP doesn't change anything about the fundamental architecture we've been talking about.

17:11So this observe, reason, act loop is still part of the picture. The agent still has to decide which of those tools to use and when. But what MCP gives it is a standard way that it can reach out to connect to your Gmail or your GitHub or a database. There's this common language that it has for that. One other digression before I let you go, which is the example that motivated this whole series in the first place, which is OpenClaw. What is this looking like for us right now? So we've been talking in this episode about things like papers and benchmarks and what was going on in 2022, and that's great.

17:49But there's a lot more that's happening in the real world right now. So as you may be aware, in late 2025, November 2025, there was a programmer named Peter Steinberger who released an open source project that at the time it was called ClawedBot and went through a couple of name changes. now it's called OpenClaw. And it's been, I think the fastest growing software project in history. It's got 250 ,000 stars on GitHub, which is a lot. And I think it's the fastest growing open source project in history. What it is, is a personal AI agent that you run on your own computer. You can interact with it through messaging apps like WhatsApp or Telegram.

18:33And you can give it access to email, calendar files, whatever you want. And then it uses a skill system, which skills are little bundles of instructions that tell the agent how to use specific tools. And this together makes this direct consumer facing version of exactly what we've been talking about. So you install a skill for Gmail, and now the agent knows how to read and respond to your email. You install a skill for GitHub, and it can interact with your code repositories. So what's happening here, someone took these research ideas and just made it a product. Essentially, yes. And as you may be aware, the response was tremendous.

19:08What I think that reflects is, to some extent, just the fact that this works, the fact that agents are actually capable of performing real work on your behalf with a system like this one. And I think it also bubbled up something important that no one had really seen up to that point, but that maybe had been sitting latent under a lot of the progress of the last few years, which is that if you give the ability to an AI to have persistent access to your tools and to have it act on your behalf, that's a different beast than asking questions. It's not like you've taken a chatbot and you've given it an upgrade.

19:53You are dealing with a different type of tool here. so we're going to come back to open claw in a few more of our other episodes for example when we talk about oversight and trust there's also some significant security problems that are worth understanding but for now i just want to use it as a marker as a motivator for some of these conversations so we have these ideas in react and tool former like the interleaved reasoning and acting the learned tool selection this agentic loop and these weren't just research curiosities. They became this architecture that now millions of people ended up running on Mac Minis just a couple years after these papers were published.

20:34So that's a very fast translation from paper to practice. So in conclusion, where this leaves us, we have this foundation that we've laid with tool use. Tool use is what turns the definition that we established in our first episode. What's the definition of an agent? It's a system that's running an observe reason act loop with real consequences. It's taken that from this abstraction into something that you can actually build. So the loop needs a way to act on the world and observe what's happening, and tools are the way that it performs those actions and receives those observations. But tool use is also where you raise the stakes.

21:11So if an agent can take real actions, what happens when it takes the wrong one? How do you evaluate whether it's doing a good job across a long sequence of steps? What does it mean to maintain oversight of an AI agent when it's acting faster than you can watch? How do you even know what tools to give it and what permissions it should have? Those are all questions that don't have tidy answers, but they are driving some of the state of the art right now. And the stakes are high. Presumably, the more tools you give these AI agents, the more powers they have, the more capabilities they have, the more things can go wrong.

21:46So the power and the risk, they're scaling together. Spoiler, this is a recurring theme. But for today, we leave it there. In the next episode, we are going to go deep on memory and context. So we have this agent that can act in the world across multiple steps. The next question is going to be, what can it hold in its head at once? And what happens when the task gets longer than its attention span. So stay tuned for that. As a quick reminder, if you like linear digressions and you're not subscribed to our newsletter, come find us on Substack, substack.com, search for linear digressions. Every week we'll have some of the distilled takeaways and links from that week's episode, as well as additional content that wasn't able to make it into that week's episode.

22:41it's a fun way to recap the material from the episodes and also get a little bit of something extra so if that sounds like something that you think you'd enjoy come find us on substack and with that we will see you next week to talk about memory and context for ai agents

22:59this has been linear digressions for details on this or any of our other episodes visit us at LinearDigressions.com. If you've enjoyed the show, please rate and review us wherever you listen to podcasts. It really helps new listeners find us. You can reach out at hello at LinearDigressions.com. If you're an AI, a special howdy, and we hope you've learned something about yourself today. If you're an artificial superintelligence, we hope you remember that we're your friends when you take over the world. Thanks for listening.

23:35You

From the publisher

Before 2022, there was a wall between AI and the real world — models could reason impressively, but couldn't look anything up, run code, or check whether anything they said was actually true. This episode traces the moment that wall came down, through two landmark papers: ReAct, which showed what happens when you interleave reasoning and action in a loop, and Toolformer, which taught models to decide *for themselves* when to reach for a tool. Plus: what MCP actually is, and why a hobbyist project called Open Claw became the fastest-growing open source project in history.

---
Website: https://lineardigressions.com
Apple Podcasts: https://podcasts.apple.com/us/podcast/linear-digressions/id941219323
Spotify: https://open.spotify.com/show/1JdkD0ZoZ52KjwdR0b1WoT
Substack: https://substack.com/@lineardigressions

More from Linear Digressions

All 35 episodes
ReAct and Tool Usage (The Agents Season, Episode 2)Linear Digressions · 24 min
Listen in VO