In short
TWIML AI Podcast Episode #726 Summary: Teaching LLMs to Self-Reflect with Reinforcement Learning with Maohao Shen
Overview In this episode of The TWIML AI Podcast, host Sam Charrington speaks with Maohao Shen, a PhD student at MIT, about his recent paper titled “Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search.” The discussion centers around enhancing language model (LLM) reasoning capabilities using reinforcement learning (RL) techniques that allow for self-reflection, self-correction, and exploration of alternative solutions.
Key Concepts
Background of the Research
- Maohao's Research Interests:
- Focus on making AI systems more intelligent and reliable.
- Emphasis on quantifying uncertainty in AI models, especially in high-stakes applications (e.g., healthcare, autonomous driving).
- Aim to move beyond simple chatbot functionality to develop systems capable of complex reasoning.
Satori Framework
- Chain-of-Action-Thought (COAT):
- A structured approach leveraging special tokens (continue, reflect, explore) to guide LLMs through different reasoning actions.
- Enables models to navigate complex reasoning tasks without external supervision.
- Two-Stage Training Process:
- Format Tuning:
- Teaches the model to understand and utilize special action tokens.
- Involves creating demonstration trajectories with the chain-of-action-thought format.
- Reinforcement Learning:
- Optimizes reasoning through trial-and-error self-improvement.
- Enhances the model's ability to self-correct and generalize beyond training data.
Techniques and Innovations
- Restart and Explore:
- A method allowing the model to restart from intermediate reasoning steps, increasing its chances of arriving at a correct answer.
- Reward Design:
- Utilizes the Proximal Policy Optimization (PPO) algorithm, incorporating iterative self-improvement strategies to refine performance.
Key Discussions
Comparison with Other Models
- Satori is positioned alongside ongoing projects like DeepSeq R1, which aims to improve LLM reasoning.
- Comparison with instruct models shows that Satori achieves superior performance with significantly fewer training data, demonstrating its efficiency and effectiveness in reasoning tasks.
Generalization Capabilities
- Despite training primarily on mathematical reasoning, Satori exhibits generalization to other reasoning domains (e.g., common sense, logic, STEM subjects).
- Performance is comparable to larger models, showcasing that Satori is not merely overfitting to the training data but instead learning general reasoning skills.
Future Directions
- Maohao expresses a desire to develop more intelligent AI systems that can handle complex tasks beyond math problems.
- Interest in creating agentic systems that can interact with environments and utilize various tools.
- Exploration of new reinforcement learning algorithms tailored for language model reasoning tasks.
Conclusion The episode offers a compelling look at how reinforcement learning can be employed to enhance the reasoning abilities of language models, resulting in systems that can self-reflect and adapt in ways that mimic human thought processes. Maohao Shen's insights into the Satori project highlight a significant advancement in the field of AI research.
For more details and the full episode, visit the [TWIML AI Podcast](https://twimlai.com/go/726).
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:00Here's the thing. This isn't how humans really naturally solve the hard problems. As humans, we are capable of extended thinking. And more importantly, we can self-reflect. If we realize we have made a mistake, we can backtrack, rethink, and correct the mistake. And sometimes we also propose better solutions. And the motivation is like how we can train the language model to do the similar things.
0:41All right, everyone, welcome to another episode of the TwiML AI podcast. I am your host, Sam Charrington. Today, I'm joined by Maohao Shen. Maohao is a PhD student at MIT. Before we get going, be sure to hit that subscribe button wherever you're listening to today's show. Maohao, welcome to the podcast. Yeah, thanks for having me for this podcast. We are going to be talking about your paper, Satori, Reinforcement Learning with Chain of Action Thought Enhances LLM Reasoning via Autoregressive Search. That is quite a mouthful, but looking forward to digging into it. Maybe we can start by having you share a little bit about your background and research interests.
1:23So I have a broad interest in various AI and machine learning problems. In general, my goal is to make AI system more intelligent and also more reliable. And a large part of my research focuses on quantifying the uncertainty of AI models. Because uncertainty can serve as a signal for users to decide whether they should trust an AI model's predictions. And this is especially important in those high-stake applications such as autonomous driving, health care, where those unreliable or misleading predictions could have serious consequences. And in recent years, I have also been working on multiple language model problems.
2:21And one of the most exciting challenges is moving beyond the chatbot and how to develop the AI system that can think more like humans and tackle those complex problems. And this is my ultimate goal. And Satari is my latest exploration along this direction. And I believe that making AI models with both smarter and more reliable is the key to achieve AGI in the future. What are some of the things you've worked on in the past? So in the past, I worked on several problems relevant to trustworthy machine learning. How to, as I said, how to enhance the uncertainty quantification performance of a machine learning system.
3:18Usually using some auxiliary models building on top of the neural network to predict the uncertainty or using some other algorithms such as Bayesian algorithms to estimate the uncertainty. And this project, Satori, is more like a new direction to be explored. And to talk a little bit about how this particular project fits into some of the current trends in AI research. which we're seeing a lot of development around reasoning models, around test time compute. Recently, DeepSeq R1 explored the use of reinforcement learning to aid in model training. There's so much going on. How does your project fit into that milieu?
4:07So our work is definitely in line with the broader research on improving reasoning in language model. but I would say it takes a different approach from the traditional methods. One of the most well-known techniques is chain-of-south reasoning, where you prompt the model to break down a problem step-by-step. And this kind of structural reasoning is really helpful for handling those complex problems. And previously, people tried improving language model by training the model on a lot of train of salt data. And those data are coming from either human annotation or distillation of the teacher model, for example.
4:57But as you can imagine, that's really expensive in terms of both data collection and training compute. And more recently, instead of just scaling up the training compute, people also started to think about scaling the test time compute. So basically, test time compute, scaling test time compute means they don't require training the model itself. They just push, they try to push the limit of an existing model by letting it generate more response. And a super simple example is majority voting. You ask the model the same question multiple times, get different answers, different response, and pick the one that is most frequent, basically.
5:54And it turns out this is quite effective. It's more accurate than just taking a single answer. And more advanced test and search methods, they usually rely on a reward model to guide the selection process. So instead of just counting answers, you can sample multiple solutions, having a reward model to score the different solutions and pick the optimal solution. And the intuition behind this kind of approach is there's even one correct answer among the samples. Then the reward model can help to find it. So some other approach, even more advanced, they try to score each reasoning step along the way and doing a tree search to find the best reasoning path.
7:06But I mean, what's the downside of this kind of approach? So these methods, they require running multiple models at once. They need to deploy two models at least, which gets expensive. Meaning the model that's scoring as well as the model that's... Right, the reward model and the generator. So two models. This is a two-player system. And they also throw away a lot of useful information you can tell. For example, if you sample 100 times but only keep the best result, you're essentially discarding everything else. This is quite inefficient. And that's why reinforced learning comes in. And this is a, I would say a quite different approach.
7:58So instead of just relying on huge amount of labeled supervised fine tuning data or complex test time search methods, reinforced learning allows the model to learn from its own generated response. And this is also the magic of reinforced learning. Once you define a reward signal, the model can essentially teach itself to get better by using its own knowledge. And this kind of knowledge is obtained from the pre-training stage of Lanxiu model. And we already know this is quite successful for some models like OpenAI's O1 model and DeepSeq R1. And this model already proved that reinforced learning could be used to train language model in an efficient way without requiring too much human supervision.
9:03So, right, so that's pretty much the entire picture about the reasoning model. And so with Satori, we are also inspired by this kind of new direction using reinforced learning. Especially since the OpenAI released a one model, that model showed there was more like there was a new path to be explored. And at the same time, we also realized reinforced learning for reasoning, language model reasoning is quite underexplored in academia. And back to September, we put together a team and started to tackle this tricky problems. And as you know, after a lot of trial and error, over the three months, we finally got Satori working.
10:03Yeah, that's a story about Satori. Right. And you also asked about DeepSeek R1. And I would say that's more like a concurrent work. It's an industry project. And they have done a fantastic job, I would say. And when their report came out, I guess it's late January. And we were actually really surprised to see how many of the high-level ideas overlap with Ours. and at the same time, it was also super exciting because it's like it confirmed that we were on the right track in the working toward more powerful reasoning models. So it's kind of a mixture of two feelings. It's surprising and it's also exciting.
11:02Can you dig into a little bit more the idea of applying RL to evaluating these chain of thought traces? Like how do you set that up from a research perspective? So when I was thinking about this research problem, I always started by thinking about how humans would solve a tricky task step by step. And then I tried to figure out how we can adapt this kind of human thinking process to AI systems. So as I mentioned earlier, a lot of researchers focus on the test and search methods, but I will say that may not be how human really solved the problem in practice. So let me give an analogy. So imagine a student working on a tricky, tricky mass province with a teacher standing nearby.
12:12For example, I'm the student, you are the teacher, and the teacher, the student might try different approaches, and you will pick the most promising solution. or maybe the students write out several reasoning steps and the teacher selects the best one and tell the students to build on that reasoning step and so on and so forth and this kind of method is definitely better than the student just working along with no feedback but here's the thing this isn't how human really naturally solve the hard problems. We don't always need any external guidance. As humans, we are capable of extended thinking, and more importantly, we can self-reflect.
13:09If we realize we have made a mistake, we can backtrack, rethink, and correct the mistake. And sometimes we also propose better solutions. And the motivation is like how we can train the language model to do the similar things, how we can make the language model solving a problem just like the human did. I also mentioned that I did a lot of work about unsending quantification. so actually in the language model community there is also a big challenge it's called hallucination which means the model will generate some response that sounds really competent but actually they are full of factual errors and for example take mass problem as an example so a typical language model might generate a step-by-step solution, output a bunch of equations, derivations.
14:23And even though there are a lot of mistakes, the language model cannot backtrack and catch the errors, even though those errors are actually quite obvious. So, right. So I've been working on a lot of approach to address this. One approach is certainly using unsenting quantification. As I mentioned, we can add some additional modules or algorithms to estimate how confident the model prediction. And if the unsenting is high, we know the answer is unreliable. and if uncertainty is low, then we trust the model predictions. But even though this kind of system is useful, it doesn't resolve the fundamental issue, right?
15:23It just tells us whether we should trust the model or not.
15:32So, but later it's more like OpenAI's AI's O1 model came out in September, and it was a game changer. So they already proved that the reflection and self-correction is not something theoretical. It can actually work in practice. So it's more like I want to know how to really uncover the technical details behind it. And we know the secret sauce of our model is using reinforced learning. Different from the classical training pipeline that teach the model how to reason using a bunch of supervised fine-tuning data, reinforced learning is quite different. It just assigns some reward to the model and led the model to interact with the environment.
16:33And secretly, the model can learn some reasoning patterns and the general abilities for reasoning. Right. So I think that's pretty much the main objective of our paper. We just want to uncover the secret sauce of how to use reinforced learning to really improve the reasoning capabilities of a language model. In the title of the paper, you refer to reasoning via autoregressive search. Expand on the idea of search in this context. What are you getting at there? So we know most language models use transformer architecture, right? most of language model is building on top of transformer. And transformer are autoregressive models.
17:27It sequentially generate a bunch of tokens based on previous tokens. And the search just means a long reasoning process with trials and errors. So combining these two words together, it's like autoregressive search means a single language model performs a human-like reasoning process with self-reflection and self-exploration of new strategies. When it encounters a potential mistake, it can self-identify the arrows and point out the mistake. It can also start a new solution if it realizes the current solution is actually suboptimal. and different from the classical test time search approach that relies on external guidance.
18:24I think the main objective of autoregressive search is teach a single language model to perform a test time search without any external guidance. And is the idea with the test time search that is it different than training the model to change the way it's generating the new token? Like how much evaluation is involved and reflection is involved? And is that all captured in the next token generation process or are there other aspects of the model? Like, you know, we've established that there's not external models, but are there internal modules or other things besides from the token generation pipeline that are coming into play?
19:18Because we know the classical approach, they typically rely on external guidance. There's a reward model to guide this process, like when the model wants to trigger self-reflection or the current solution is suboptimal and the reward model will come in and provide some feedback and select the optimal reasoning stat, for example. But if we want to achieve the similar thing, similar behavior using autoregressive predictions by generating a token in an autoregressive manner, there's a challenge, right? So it's like the traditional chain of thought approach doesn't really achieve this because it just produces a linear sequence of reasoning steps without distinguishing the types of reasoning actions.
20:19So we can treat self-reflection or proposing new strategy as different actions. The COT is more like a robot that can only move straight without the ability to change directions or or stop or rethink the approach. And, but we think about, but let's think about how a human solved the problem. You might go through a few reason steps, pause and start self-reflection. Maybe backtracking, verifying correctness, correcting mistake, or even coming up with new strategies to tackle this problem, right? So it's more like a robot navigating an environment. If it hits a war, it stops and it decides the next movement and explores a different path, different trajectory.
21:18So I think the core idea of our paper is this chain of action salt in this long title. So it explicitly defines different actions the model can take. and the unique technique we introduce is directly leveraging the special token to guide the language model's generation. Because we know autoregressing model can only produce a sequence of tokens, nothing else. But we can still use special tokens as some signal to tell the model how to take the next action. So we introduced a bunch of special tokens and each special token corresponds to one specific actions. And we basically define three types of actions, continue, reflect and explore alternative solutions.
22:23And continue special token encouraged the language model to build upon its current reasoning trajectory by generating the next intermediate step. And reflection tells the model to stop reasoning and verify the correctness of its previous reasoning steps. And exploration signals the model to identify critical errors in reasoning and explore a new solution, basically. So we recycle these three actions. The model can iteratively refine its reasoning and their reasoning process is more like how humans solve the problem. And so then the key challenge of the project is training a model to generate these new special tokens at the right time.
23:19Is that right? I would say the first stage of post-training has a challenge. we need use a format tuning stage as we propose in our paper to get the model familiar with the reasoning format of the chain of action salt reasoning. Or in other words, we should teach the model to get familiar with the special action tokens. When it encounters the special tokens, it needs to make the corresponding response and take the actions. Can you walk us through the training process and how you accomplish these goals? If we want to teach the language model to perform the COAT reasoning, there are actually two challenges.
24:07The first challenge is, as I said, it's called unawareness of the meta-action tokens. So this means that without any training and without any post training or fine tuning, the language model doesn't recognize that encountering these special meta tokens will require reflection or proposing alternative solutions. So we can think about the robot example. So without the training, the robot doesn't know how to properly take actions based on different situations, like how to grasp an object from the table. And the straightforward way to teach the robot, of course, is using some demonstrations provided by humans.
24:59And this technique is exactly called imitation learning. We need to construct some demonstration trajectories with this kind of chain of action salt reasoning format and let the language model emulate the behavior of the demonstration trajectories. And then the problem boils down how to construct this kind of trajectories. Human annotation might be an effective approach, but it's of course really expensive. So in our paper, we instead propose a multi-agent data synthesis framework. So there are two models, basically. It's also a two-player system. There is a generator and there is a critic model.
25:51So given an input problem, the generator will propose solutions and the critic model will step in and provide feedback. And these two models collaborate with each other to solve the tricky problem. And the generator might make mistakes. The critical model aims to identify this mistake and suggest an alternative solution and so on and so forth. So let me make an example. This is more like the scenario that the two students collaborate to solve a problem. or if we are solving a problem by ourselves, it's more like the situation that there are two voices in someone's head and they may or may not agree with each other, right?
26:47That's usually the case. But eventually we hope these two voices will end up one single solution and successfully solve the problem.
27:01you know, through a long debate process, right? The two boys compete with each other, debate with each other, and finally they successfully solve the problem. And we basically want to construct this kind of reasoning data. And then this kind of reasoning data is used to fine-tune a language model, a pre-trained language model. And we call this entire process format tuning. which aims to teach the language model follow the chain of action salt reasoning format. And after this kind of training, the model knows to tackle different problems using different actions triggered by corresponding special tokens.
27:53Is it learning that because that is the way that the critic is communicating to the generator when it needs to take a different action using those special tokens? The critic is more like providing the self-reflection part. It provides the feedback, the suggestions. And this kind of information can be viewed as a self-reflection process. Right. And this is triggered by this reflection special token. And also in practice, we observed that using only 10 ,000 this kind of demonstration trajectories, the model can already learn this chain of action salt format. So we basically have two stages of training.
28:44One is format tuning and the other is reinforced learning. So the main motivation for this reinforced learning stage is the model, although the model learned the chain of action thought format, but it may struggle to generalize its reasoning performance to unseen problems. I think this is expected because the model only ambulates the expert demonstrations. And the quality and the diversity of this demonstration usually, you know, constrain the performance of the model. Right, so in order to mitigate this issue, we need leverage reinforced learning to truly incentivize the model's reasoning capabilities, especially the capability to take different actions with respect to different situations.
29:48So we encourage the model to reason like an agent. It learns from the trials and errors and gradually build a policy that allows the model to get higher reward. But there is a second big challenge, the problem of sparse rewards. So let me explain So language model reasoning Is more like a long horizon Decision making process It needs to output Multiple intermediate steps Takes a lot of actions And finally It gives you a final answer And receives the rewards From the environment But we realize that The reward is only Received at the end And this causes a problem because the language model... It's attribution types of problems.
30:53Right. So if we encounter the problem that the language model cannot successfully solve the problem, there's no any reward. and as long as the model fails, it needs to start from the initial province, the initial state. Meaning so you're not teaching the model to, you're not giving it any partial credit. It's all or nothing. Yeah, so the reward is only at the end because for the mass province, the most effective reward is just checking whether the final answer is correct or incorrect and such reward is at the end, of course. Right. And so to address this challenge, we propose a key technique in our paper, which is called Restart and Exploration.
31:55Restart and Explore, right. So the idea is actually quite simple. Instead of starting from the problem statement or the initial state, we also let the model to start from the intermediate steps, the intermediate states. And it will increase the chance of the model that arrive at the correct answer. And it also encourage the model to trigger self-reflection and self-correcting its mistake. And this is similar to how a robot explore the environment instead of always putting the, placing the robot at the same starting point, we can always place the robot at some intermediate state that the model already, the robot already explored and we let the robot to restart from their intermediate state.
32:54And is a state in this case a sequence of tokens that have been generated? Right. In the language model case, that's the case. It's more like some intermediate reasoning steps. The model already generated in its previous response. Right.
33:16So consider the scenario that this intermediate state is actually a failure state. For example, if we let the robot to move an object from point A to point B, and the robot somehow fails at a certain point of point C at the middle, and we can actually let the robot to restart from an intermediate state, which is this point C, and makes a second attempt of exploration. and in practice we observed that this kind of restart and explore is actually quite effective for the language model to learn how to self-reflect and how to self-correct the mistake right and i think that's that's the that's the main contribution of the reinforced learning part Can you tie together the format tuning and the self-improvement?
34:22Are they totally independent steps or in what ways do they both need to be in place in order for you to achieve the goal? So this is more like you need a pre-training stage for the language model, right? So the pre-training stage aims to inject all the necessary knowledge and information into the language model parameter. and the format tuning as the first stage of all training is more like you teach the model to reason according to a certain reasoning format. And this format is defined using these special tokens and this is the so-called chain of action salt reasoning format. And this training stage is usually small scale because in practice, we only use 10 ,000 of training data.
35:18And the reinforced learning is more like how you can leverage this kind of special reasoning format to really solve the problem, like take different actions and finally solve the problem and get a reward from the reinforced learning. And this stage, this reinforced learning stage is usually large scale. Right, so in practice... How would you characterize the scale of that? In practice, it's more like the format tuning stage, we only use 10 ,000 examples. And reinforced learning, we use more than 300 ,000 examples. But that's actually quite efficient. Why? Because for reinforced learning, we only require the final answer as the target.
Read the full transcript
36:15We don't need the intermediate COT part. So for all the problems, we only need to get the final target, final answer. For example, for the mass problem, then the only thing we need is a final answer. But for the traditional training, the traditional pipeline of training, which usually relies on a large-scale supervised fine-tuning with millions of examples. But that kind of supervised fine-tuning data usually requires human annotation, especially for this intermediate COT part. Is the idea starting from an intermediate state novel in this context, or is that generally not done in RL in general?
37:09It's actually inspired by a classical reinforced learning paper, which is called Go Explore. And that was using this kind of technique to train basically a game agent, an agent that can navigate in the game environment. But this idea hasn't been applied to the language model context. So it's more like we borrowed a lot of insights and ideas from the classical reinforced learning community and apply that to the language model case. And do you have a sense of if you didn't incorporate that idea of starting at an intermediate state, how that would change the number of samples you need or some other factor or the ultimate performance?
38:02So two things. One is if we update this restart and explore technique, we realize the overall benchmark performance will be worse compared to the original version. And second is like we also check the self-correction capability of the model. if we remove this kind of new technique the restart and exploration the model might not be that effective in terms of self-correcting its own mistake so we did conducted some ablation studies we tracked those response where the previous attempt and the second attempt answers are different the answers of these two attempts are different And we evaluated whether the model is actually correcting the wrong answer to the correct answer, or it's actually correcting the correct answer to incorrect answer.
39:07we calculate the ratio between these two scenarios. And we realize that using reinforced learning and this restart and exploration technique, the ratio between these has a big improvement compared to the format tuning model. Basically means only using the first stage of training, the format tuning stage, the model doesn't have any capability to self-correct its mistake. But using large-scale reinforced learning with restart and exploration, it gradually learns the self-correction capability. Can you talk a little bit about the reward design for the reinforcement learning component? I think the main algorithm of reinforced learning is the PPO algorithm.
40:00and we just built on top of the PPO algorithm, the classical PPO algorithm and the novel things involve this restart and exploration. And also we proposed another method called iterative self-improvement. That's also something new in the literature, but that's also inspired from the classical reinforced learning community. And that's corresponding to a classical technique called kickstarting. The main approach is, the main strategy is you train a reinforced learning policy as a teacher policy. And you distill the knowledge from the teacher policy into a student policy. And you start from the student policy and you continue using reinforced learning to improve the student policy.
41:04So why is this effective? Because if you apply reinforced learning to train a single policy, the policy might converge to some local optimal during optimization. And by distilling the teacher policy into student policy, it's more like you change the loss landscape of optimization and it will help the model to jump out of the local optimum. So in practice, it's like we apply this reinforced learning in the first stage and we got a model and we distill this model into a new base model. and starting from the distilled base model, we apply the second round of reinforced learning. So it's more like SFT stage, reinforced learning stage, distillation to a new base model by supervised fine-tuning and starting from the new distilled model, we apply the second round of reinforced learning training.
42:20And in practice, we realized a second round of reinforced learning can further boost the performance. Did you look at a variety of base models or how did you select the base model? We only pick one base model in our experiment. The base model should have the sufficient knowledge, like the knowledge obtained from the pretraining stage. right a weak base model we also right so in our experiment we also conduct some other interesting experiments we use two additional base models which are weaker in terms of the mass reasoning performance but the interesting experiment is we already obtained the strong reinforced learning model checkpoint and we distill the knowledge of this strong model checkpoint into the weak base model.
43:26And we realize this is actually quite effective. The distillation can make base model become a really powerful reasoning model. And this is actually really effective because in practice, we can first train a strong reasoning model using large scale of self-improvement using reinforced learning. And then we can improve the base model performance using this strong teacher model. Right. And this is quite effective. And this is also quite different from the classical way to train the model. It's like we still need to collect a large amount of supervised fine tuning data. and usually that's annotated by human and we train this model using supervised fine tuning.
44:21Can you talk about the various benchmarks you used and the results that you saw? We mainly test our approach in the math domain because we train the model on computation level math problems and this is actually quite common in research communities because math is easy to evaluate. You just check the final answer, whether the final answer is correct. And math problem is also a strong indicator of reasoning abilities.
44:59So we just trained the Satori on math dataset and evaluated the model on several standard benchmark. And the results are quite exciting. So Satori outperforms industry-level instruct models trained on the same base model. And we actually use much fewer training data. So compared to the instruct model, we only use 1 % of the data in the SFT stage. And instead of relying on massive amount of label data, we mainly train the model using large scale reinforcement as I mentioned. Right. So,
45:48and given that we train Satori on mass domains, it's actually not surprising that it performs well on mass benchmarks. But what's really exciting for us is its generalization capability. Even though the model was only trained on math, we find that it actually transfers well to other reasoning domains, like common sense reasoning, logic reasoning, and other STEM subjects like physics and chemistry. And what's even more surprising is that Satori achieved comparable performance to those instruct models that were explicitly trained on a diverse of data sets across multiple different domains. So this means that our model isn't just memorizing those math problem solving skills.
46:43It's actually learning the general reasoning skills. And some of the models you compared it against were much larger in terms of number of parameters and it seemed to do well. We mainly compare with the same scale models like the LAMA 3.1, 8 billion, and these kind of models. And we realize using our training framework, but only training on math domain, we can match the performance with LAMA 3.1 on a bunch of general reasoning domains. Yeah, my interpretation of the table is that for the math benchmarks, the Satori outperformed the comparable models, but the larger models did better. But for out-of-domain, the Satori did better than both the comparable models and the larger models.
47:40Is that right? Some of them. For example, Satori achieved the comparable performance with Lama 3.1, 70 billion, I guess. So actually, we also conducted some ablation study, and the results is showing that this kind of general reasoning ability actually emerged from large-scale reinforced learning. So this is quite exciting because traditionally, language model was more like chatbot. It's trying to memorize patterns from the supervised fine-tuning data. But using reinforced learning, the model behaves more like an agent learning using trials and errors. And it gradually builds more general reasoning skills.
48:32And so how do you see this evolving? Where do you want to take this research direction? So I would say the ultimate goal of the Satori team is to build smart AI systems. Not just language model with strong reasoning skills, but also those agentic framework that can tackle really tricky tasks. And as researchers, we are not just focused on building powerful models or products. We also want to provide the useful insights and explore those fundamental research questions and benefit the research communities. And one of the next big area we are looking at is building a more agentic system. We hope the model to interact with environment to use different tools, for example, search on the website.
49:37And the language model can act like a reasoning engine in this agentic framework. Right. So we just want it to tackle this trickier task, not only limited to mass problem solving. and another research direction is how to develop the new reinforced learning algorithms especially for language models because we know that reinforced learning has been well studied in classical ai fields like robotics and i always believe there are a lot of valuable insights from those classical fields that we can adapt for language models. I think we are still in the early stage of figuring out how to optimize reinforced learning for language model reasoning tasks, but I think there is a huge potential in this space.
50:45Well, Maha, thanks so much for sharing a bit about your project and the research. yeah thanks again for having me today
From the publisher
Today, we're joined by Maohao Shen, PhD student at MIT to discuss his paper, “Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search.” We dig into how Satori leverages reinforcement learning to improve language model reasoning—enabling model self-reflection, self-correction, and exploration of alternative solutions. We explore the Chain-of-Action-Thought (COAT) approach, which uses special tokens—continue, reflect, and explore—to guide the model through distinct reasoning actions, allowing it to navigate complex reasoning tasks without external supervision. We also break down Satori’s two-stage training process: format tuning, which teaches the model to understand and utilize the special action tokens, and reinforcement learning, which optimizes reasoning through trial-and-error self-improvement. We cover key techniques such “restart and explore,” which allows the model to self-correct and generalize beyond its training domain. Finally, Maohao reviews Satori’s performance and how it compares to other models, the reward design, the benchmarks used, and the surprising observations made during the research.
The complete show notes for this episode can be found at https://twimlai.com/go/726.




