In short
Episode Summary: Inside s1: An o1-Style Reasoning Model That Cost Under $50 to Train with Niklas Muennighoff - #721
Podcast Overview
- Podcast Title: The TWIML AI Podcast
- Host: Sam Charrington
- Guest: Niklas Muennighoff, PhD student at Stanford University
- Topic: Discussion of Muennighoff's paper "S1: Simple Test-Time Scaling" and its approach to reasoning models.
Key Concepts Background of the Guest
- Niklas Muennighoff transitioned from business and finance to AI research, motivated by the potential of AI to transform industries.
- He joined Stanford's PhD program after industry experience at Hugging Face.
S1 Model Overview
- Goal: S1 aims to replicate OpenAI's O1 model with a simpler and cost-effective approach.
- Motivations:
- Strong reasoning performance on math and science tasks.
- Improved performance with more compute at test time.
Test-Time Scaling Approaches
- Parallel Scaling: Running multiple computations independently and aggregating results (e.g., majority voting).
- Sequential Scaling: Generating a single reasoning trace, refining answers iteratively.
- Both approaches can be combined for better performance.
Key Techniques
- Budget Forcing: A novel method for allowing the model to think longer on harder problems while optimizing test-time compute.
- Data Curation:
- Initial dataset comprised 59,000 difficult questions, filtered down to 1,000 based on quality, diversity, and difficulty metrics.
- Inspired by the "Less is More" hypothesis from previous research.
Distillation Process
- S1 utilized distillation from models like Google's Gemini and DeepSeek's R1, focusing on extracting reasoning traces and answers.
- There's a trade-off in the quality of the teacher model; better teacher performance leads to better student model outcomes.
Evaluation and Results
- S1 was trained with approximately $50 worth of compute and achieved competitive performance on certain benchmarks.
- Compared to O1, S1 showed better results on some math benchmarks but lagged behind on science questions.
Future Work and Open-Sourcing
- S1 is fully open-sourced, with resources available for replication.
- Future directions include addressing the limitations of scalability and context window issues in reasoning models.
Comparison with Other Models
- S1 vs. R1: S1 focuses on a minimalistic approach aiming for strong reasoning and test-time scaling, while R1 attempts a full replication of O1 with reinforcement learning.
- Hugging Face OpenR1: Aims to replicate R1, likely involving reinforcement learning and more complex methodologies compared to S1's straightforward supervised fine-tuning.
Key Takeaways
- Significance of S1: Offers a low-cost, accessible entry point into advanced reasoning models for researchers and practitioners.
- Future Research Potential: Opportunities exist for further development in model reasoning capabilities and efficiency.
- The Role of Open Science: Emphasizes the importance of open-source research in advancing the field of AI and ML.
Conclusion The episode provides insights into the advancements in AI reasoning models, the motivations behind these developments, and the practical implications of simplifying complex methodologies for broader accessibility in AI research.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:00At a high level, both S1 and R1 are attempts to replicate O1 from OpenAI. And I think the key difference is that for R1, they try to go the full way, replicate the entire pipeline as much as they could. Whereas for S1, we focused on what is the minimal approach, the simplest way to get those two exciting things about R1, which is the strong reasoning performance and the test time scaling. What is the minimal recipe to get that?
0:46All right, everyone, welcome to another episode of the TwiML AI podcast. I'm your host, Sam Charrington, and today I'm joined by Nicholas Munikhoff. Nicholas is a PhD student at Stanford University. Before we get going, be sure to take a moment to hit that subscribe button wherever you're listening to today's show. Nicholas, welcome to the podcast. Thanks for having me. You are welcome to be on. I'm looking forward to digging into the conversation. We are going to be talking about your very recent paper on S1, Simple Test Time Scaling. I'd love to have you share a little bit about your background and kind of how you got started in the field.
1:26You're a first year student there at Stanford, right? I originally did my bachelor's at Peking University and I was focused on business and finance. But then midway through, I realized that artificial intelligence was going to be probably the most important paradigm of our time. So I decided to completely switch my focus and do AI research instead. And then after graduating, I spent some time in industry at Hugging Face and others and eventually landed here in the Stanford PhD program as a first year student. Let's talk a little bit about S1 and how this project came about. you saw what was happening with the O1 models, presumably.
2:05And was that the impetus? Yeah, it was. Interestingly, it was right at the start of the PhD that they announced O1. So it felt like the perfect thing to jump on and do research on this topic. And I think there were two things that we were really excited about with the O1 release. The first one was the really strong reasoning performance that they showed. So gains on math tasks, but also science questions. And the second really exciting one was that those gains seem to get better with more compute at test time or test time scaling. Can you talk about other approaches that folks have taken to dig into test time scaling and try to replicate what OpenAI did with O1?
2:51At a high level, there's two different approaches to test time scaling. One is you scale in parallel, and the other is you scale sequentially. In parallel means that you're running multiple computations independently. So for example, you take one model and run the question through it multiple times, and then you aggregate those answers somehow. The most common way to aggregate them is majority voting, where you just take the majority answer. And similar to classical assembling techniques in machine learning, this will slightly improve the performance of the model because you're not bound to one-off errors or mishaps of the model.
3:28And you can do more fancy things there like have a reward model decide which output to pick and so on. And on the sequential side, you have the model generate a single reasoning trace or reason in a single chain and then iteratively refine what it's doing and improve on its previous current answer. And intuitively, this one should scale better because you're not making potentially the same mistake multiple times in parallel. But the model is reasoning in one continuous chain. So it can build on intermediate steps and try to arrive at a better answer. But the two of them are not exclusive. You can combine them and there is some synergy.
4:10And in the OpenAI 01 blog post, even though the 01 model with its long chains of thoughts that it generates, it is sequential test time scaling. They do show that they can add test time parallel scaling on top of it with majority vote and further improve the performance. So they generate multiple chains in parallel and then aggregate them with majority voting. Is there work being done that looks at, rather than aggregating via voting, aggregating via an additional model processing stage? So maybe, you know, summarization or something like that, or training a model to look at the output of these individual parallel runs and kind of synthesize them.
4:55So there are approaches that leverage reward models in this process. So I think best of N or even some tree-based approaches like rebase, which we also explore in the S1 paper, where there is a second reward model that assigns a reward to the current intermediate or the final output of the model. And then based on that, you just take the one with the highest score or the few ones with the highest score or something like that. And that way, hopefully arrive at a better answer. Yeah, I think I was thinking less of like using the reward model to assign scores and more like using a model to synthesize the text that the other models spits out.
5:36because if the models are, you mentioned how if the models are just generating text in parallel, then you, a disadvantage relative to sequential is that the, you know, other models aren't necessarily learning from the output of the, you know, the other N-1 models. But if you've got a model sitting in front of that, that can potentially learn from all of the outputs, maybe it can generate a higher quality. output oh that seems like a an interesting research idea so maybe we should maybe explore that so it could be i guess it could either be at the very output level or even as you hinted at during the generation of each of these models in parallel maybe there's some synchronization step where all of the models say okay this is where i'm currently at those are the mistakes i've made and then sort of they try to improve based off of that i think that's very exciting um yeah trying to either put another model at the end and improve based on all of these parallel tries.
6:41I haven't seen any work on that yet, but I think it's very exciting. Your work came out right around the time of the DeepSeq R1, you know, excitement. Well, it came out around the time of DeepSeq R1 and the subsequent excitement, I should say can you contextualize s1 with r1 and like how those projects relate to one another i think at a high level both s1 and r1 are attempts to replicate o1 from openai and and i think the key difference is that for r1 they try to go the full way replicate the entire pipeline as much as they could Whereas for S1, we focused on what is the minimal approach, the simplest way to get those two exciting things about O1, which is the strong reasoning performance and the test time scaling.
7:41What is the minimal recipe to get that? And in terms of the results that you saw, I think the headline was that S1 was trained with about$50 worth of compute and matched or exceeded the O1 preview scores on some benchmarks. Can you elaborate on what you saw? So on the cost front, our final training run took 26 minutes on 16 H100s. And if you ran 16 H100s on one online platform, so I looked at Prime Intellect, for example, it costs around, I think,$40 per hour. So that would be 20 for 26 minutes. So very cheap for that final run. And I think that's where... That's the cost for 16? That's a cost for two nodes, 16, yeah, on Prime Intellect.
8:31It's a great platform. And we, of course, also use resources to generate the data and create other parts. But for the final training run, that's all you need. And since we published all of our data sets, anyone should be able to replicate our results with those 26 minutes of 16 H100s. Yeah. I think the performance was better than 01 preview on some math benchmarks, like AME and Math 500. But on GPQA, which is like common PhD science questions, it was worse. And also O1 Preview is an older version. So I think at this point in time, they're already at O3. So of course, need to contextualize the results a little bit.
9:11But nonetheless, I think given the very cheap recipe, it is quite impressive that we could get this strong performance. Awesome, awesome. And we'll talk a little bit later on, I think, about how you might go about evolving your approach to keep up with advances. But I did want to dig into the training recipe in particular. And a couple of the major steps there, one of them was the data set curation. And then this interesting bit that you did with that you called budget forcing. But let's start with the data set. How did you pull that together? Initially, we collected some of the hardest questions we could find online.
9:50So we took some really hard Olympiad questions from covering subjects like math, physics, and others, but also brain teasers from quant interviews and these kinds of difficult problems where you'd think that a lot of reasoning is necessary. And we assembled 59 ,000 of those. And then we went through a process to filter them down to just 1 ,000. And the key criteria in this process were focusing on quality, diversity, and that are very difficult. And the way we quantified these three metrics is that for difficulty, we looked whether previous models could solve them. And if they can't, then we thought these would be very challenging.
10:37So ideally, like a previous poor model cannot solve them, but we still know that they are solvable either by a really strong model being able to solve them or some indication that there's a solution or something like that. And then for diversity, we made sure to sample from very different fields. So not just math, but also astronomy, biology, and various others. And finally, for quality, we had some very simple filters, like ensuring that there is no duplication of questions, that the questions are all properly formatted and so on. And ultimately, this led to our final 1 ,000 example selection.
11:16And that number 1 ,000, did that come from your intuition? You know, let's see what we can do with 1 ,000, or we should be able to get this down to 1 ,000? Like, where did that come from? There is a great prior paper called Lima, Less is More for Alignment, And in that paper, they showed that with just 1 ,000 samples, they were able to train a model that was a pretty good chatbot. So based off of the original Lama models, they used these 1 ,000 samples to fine-tune it. And they framed this as the superficial alignment hypothesis, which is that to align a model, you actually don't need that much data.
11:55You can make it much simpler. And we wanted to revive that, but for reasoning. And that's where the 1 ,000 number came from. And it's also somewhat catchy if we can do 1 ,000. So you generated this data set of 1 ,000 questions or 59 ,000 questions, but then you filtered that down to 1 ,000 questions. You mentioned, and then the next step, I think, was doing distillation. And you did that against Gemini. Can you talk a little bit about distillation and the inspiration for that and, you know, how that process worked? Yeah, of course. Initially, we used the Google Gemini Flash Thinking API, and we just ran it over even the entire 59 ,000 questions.
12:48And then we took from the API the thinking trace and the solution. So at that point in time, they already had a reasoning model, and we directly took the traces and answers from that model. But later, after R1 was released, we switched to the R1 reasoning traces and found that actually they are much better. And we were able to train a slightly better model with that that we just released recently called S1.1. You use the Gemini thinking experimental and I am curious if you've gone back to look at that. And I'm only asking this because in my own experiences with that model, it's changed a bit over time.
13:35And when you talk to Google, it's like it's experimental. We're gonna change it without notice. It's not version, that kind of thing. But when I've used it for coding, I noticed like some pretty dramatic changes in the way it responds to questions over the course of the past, like say five weeks or four weeks? Did you ever go back and look at it again? Yeah. Or did you notice a shift in the course of the project that kind of changed the way you, changed the kind of results you were seeing out of that model? Yeah, the biggest change was, it seems like after our paper got released, they removed the option to get the thinking trace from their API.
14:19and besides that they've also followed up with newer and presumably better thinking models so maybe that also motivated that change and yeah it's changing can you get the thinking trace from those models no i think they removed that option for now okay but that's also partly the reason why r1 is a much better option if you want to generate thinking traces besides the reason that it leads to better performance as we showed with our follow-up. Because you can, because they make the thinking traces available. Exactly, yeah. In talking about the dataset curation, you mentioned difficulty and the criteria was roughly, you know, models have a hard time with these questions, but then the distillation process is let's ask a model these questions and then capture its answer and its reasoning trace.
15:10You know, presumably those are kind of at odds with one another, Like the models are going to generate wrong answers and think about them in wrong ways, but then you're using that information for training. Talk a little bit about that paradox. I think one hypothesis is that you're bottlenecked by the capability of your teacher. So in our case, for Gemini, on one of the benchmarks, Amy, it got 60%. And after distillation, we got 57 with our model. so we decided to stop there at that point in time because it felt like there's not that much room to go further if we keep distilling from Gemini because we're already very close to its performance we're just using its thinking traces and its answers so how much further are we able to go but then with R1 which had much better performance than Gemini we went a lot further and performance improved so I think that is a an important thing to think about is how good is your teacher ideally of course, it's as good as possible.
16:11But not to say that you can't be better. Maybe it is possible to get better than the teacher you're distilling from. But with 1000 samples, that's especially difficult. Yeah, yeah. So the veracity of the answers that the teacher produces is a challenge for this training. Yeah. And it does get many of these questions wrong. So in our case, we actually trained on many wrong questions from the teacher because we didn't find there to be a big difference if the questions are actually correct by the teacher or not. It's much more important to show the model how to do the reasoning and arrive at some answer, even if some of them are wrong.
16:51That's usually fine. Oh, interesting. And did you kind of capture and track a metric, you know, correctness of the training data? And like, what was that percentage? Yeah, we didn't capture that. But I recall a recent paper called Stream of Search, where they also show that with, if I remember correctly, around 50 % of answers being incorrect, they were able to train a better model than one that was only trained on 100 % correct reasoning traces and answers. and part of the reason there was also that i think the incorrect ones they had a lot more reasoning and backtracking because presumably the model knows it's not quite the right time yeah exactly and so it like goes a lot further and that's potentially why it leads to better performance so the intuition is like you're fun what you're trying to do with uh um with the supervision is to teach the model how to think and even if it doesn't get the right answer if you're seeing it kind of go through thinking cycles, that's going to provide some signal to the model.
17:59Yeah. And especially the backtracking in the thinking is really important. So moments where the model goes, wait, is this correct? Let me go through it once more. And these are, of course, more likely with really hard questions, maybe even where the model ends up with the wrong answer. Talk a little bit about the budget forcing aspect of the training process. Yeah. So after training on those 1000 samples, we got a good reasoning model with strong performance on math and science questions similar to 01 and 01 preview. But we didn't yet have the test time scaling. So the second goal, how do we get better performance if we give the model more test time compute?
18:40And we can go into details why this is desirable. But the way we got there is... Well, yeah, let's do that and just start it by elaborating on what you mean. So you have a model that you're training, you're distilling the kind of curated data set and like you're fine tuning this model. When you say you didn't have test time scaling, what does that mean specifically? Yeah, it means that we run the model on those math benchmarks, but it always uses the same compute. So say we ask it, what is 10 plus 10? No matter what we do in the simple setup, it just always does a forward pass through the model and then provides an answer.
19:32there was no clear way at the time for us to change the test time compute so how much time does it spend to think about this answer and and the reason through the question and of course for very hard questions intuitively we'd like the model to think for longer because they are harder it should take more time and for simple questions we don't want it to spend that much test time compute so it should give us an answer very quickly and so can we produce like a plot that It shows that if we provide the model with more of this test time compute, its performance improves continuously. Got it. So relating back to the whole thinking fast and slow idea that kind of kicked off some of this test time scaling work.
20:19Awesome. Awesome. And so how did you implement that? Yeah. So the way we went about it is, again, trying to find a very simple approach. And the one we put forth is called budget forcing. And the idea is we have some test time compute budget. And we usually measure that one in tokens, which is how many text units the model produces. And given this budget, we just force the model to generate that long of a reasoning trace. So either we cut it off. So after it has generated, say, a thousand tokens, we just cut off the thinking. and then the model will be forced to provide its current best guess, current best answer.
21:02And that works well for cutting it short. But the key question for us was how do we extrapolate it? So how do we make the model think longer when it wants to stop and wants to provide its current answer? And for that, we force the model to continue thinking by injecting weight into its current reasoning trace. And when we inject weight into the model's reasoning trace... And to be clear, sorry for interrupting, we're talking about the word or like the token weight, not weights as in parameter weights or anything like that. Yeah, of course. We're talking about the word weight and actually the string, which is then transformed into a token, just right when the model is currently generating its answer to some question we give it.
21:48And so the model might be something like, okay, given what I've reasoned above, the final answer should be three. And then it wants to finish and stop the generation. And what we do is we disallow that and just put wait, and then let the model continue reasoning through it and generate further. And of course, what the model will do is it will be like, wait, is this correct? Let me go through it once more. Because the way we train these models is through next token prediction, where they learn to predict future words. And so if we provided the word wait, then the future word after that is most likely something like, wait, is this correct?
22:28Wait, let me do this once more. And we actually have some of these wait examples, even in the 1000 samples. So some of the reasoning traces, they are naturally something like, wait, let me recheck this. And so the model learns to generate these tokens after wait. And then of course, when it rechecks, it goes through it once more. And if the previous answer was incorrect, then it is likely that it will change it and improve on its answer to get a correct answer. Versus if it is already at a correct answer, you might think that, okay, might this impair the model and lead it to an incorrect answer?
23:06And while this can happen, it's generally less likely. It might do something like, wait, let me double check this. And then it'll be like, it seems my original answer was correct. And the reason we think that it's less likely to change to an incorrect answer is because verification is easier than solving. So when the model verifies its current answer, then it's easier for the model to know whether or not it's already correct. Did you try other words besides wait? Yeah, we tried alternatively. And also no word at all, like just letting the model continue. And wait seemed to provide the best performance, but we definitely haven't exhausted the list of possible words there.
23:47and yeah, I'm excited to see what other people try. And can you elaborate a little bit on how exactly that word is injected in? Is that in kind of the inference generation loop itself or someplace else? Yeah, exactly. It's in the inference loop generation. Specifically, we set the end of thinking token. So the way our generation works is that we put in a question and then after the question, there's a delimiter that denotes the beginning of the reasoning process or thinking process and then there's a end token for this one or end of thinking token after which the model will generate its final answer based on its thinking process and whenever the model tries to generate this end of thinking token delimiter we disallow that so we have it as a stop token in our generation pipeline.
24:41And then instead inject wait right into the model's current generation, and then it'll just continue generating off from there. And of course, after wait, it's very unlikely that the model will just try to put the end of thinking token again, because now a new sentence started, so it'll do some more thinking and reasoning and hopefully get a better answer. And now the insertion of wait doesn't have anything to do with the correctness of the answer, how far the model had gotten or any of that. It's just random. We do it whenever the model wants to finish. So presumably it already has an answer.
25:17And the question is just, is the answer correct or not? And for that, we didn't do any further optimization. But I think you could do something like check with a reward model first or have the model check itself. Is it worth continuing further? Or should I just end the thinking here and provide my final answer? I think there's definitely optimizations there to be done. Okay. So are you always providing some constant number of weights for each generation? Or is the number of weights variable or random? Yeah, we always pick the same number and found in our experiments that up to four can lead to performance gains.
26:01but beyond four times eventually the model will go crazy because it's like getting weighted all the time and then it'll just go off into an infinite loop and bad things happen but up to four times it's still relatively natural and the model will properly use them to improve its performance so yeah we got pretty consistent math gains up to that degree when i'm hearing this i'm hearing like okay for this training run you're always saying wait four times that is training the model to always think, you know, four lengths worth. But I guess what is not true there is that the length of thinking tokens is constant.
26:40The length of thinking tokens is just whatever the model generated based on next token prediction. And so that's going to be variable even if it's always four weights. Yeah, that's a great point. So we can combine it with a maximum, or we could also set a minimum token budget for the thinking trace. But generally we find just adding weight is enough to get us some additional test time compute with better performance and varying that a lot doesn't bring out further gains. And maybe to tie back to the budget forcing and cutting off the generation. So if we put these together, then what we arrive at is like a plot where we have different test time compute budgets.
27:21So either say we cut off at 1 ,000 tokens, we cut off at 2 ,000 tokens, We cut off at 3 ,000 tokens, and then we force wait one time. We force wait two times, three times, four times. And each of those dots has better performance, and we get a very clean test time scaling trend. Meaning each of those is a separate independent test. And so if you're waiting four times, you're not using the cutoff. That's a separate thing? Exactly. We run separate evaluations. So the idea is that ahead of time, you have an idea of how expensive your question is to the model or how much you want the model to think about it.
28:04And based on that, you set your test time compute budget constraints for the model and then let it go with that. You could make it variable as well. So having the model itself decide. But that's kind of already happening in the original thinking trace. So naturally, the model will use different amounts of tokens for different questions. So we already have that. It's just if you additionally want to set a budget constraint to enforce something, you might say the model should answer in one minute or in 10 minutes, then you should use that. Is it minutes or is it tokens? It is tokens. Time-based?
28:39Yeah, it is tokens, but at the end of the day... But you've got a constant tokens per second. Exactly. So depending on how fast your GPUs are or how you're deploying it, you can convert that. I assume in a UI interface, you would turn it into minutes. Or as OpenAI is doing, in their API, they have three parameters for a reasoning effort key. They're low, medium, and high. And so people can choose what they want there. Not exact time control, but also some amount of control. Going back to the dataset, there was a mention in the paper about decontaminating samples. Talk a little bit about that and what that means.
29:18Sure. One big problem in large language model training is that we're training these models in huge amounts of data. And when we evaluate them, there's a chance that some of that evaluation data was in the training data. And then the evaluation isn't really proper anymore because the model has seen these things already. It might just be reciting its training data and not generalizing or really improving performance. And so that's why we employ decontamination. In our case, what we do is we take all of the evaluation samples and check for any matches in our training data. And if there's any match, then we discard the sample from our training data.
Read the full transcript
29:58There's any overlap. And since we only have 1000 samples, that's relatively easy to do. But for pre-training, that's a pretty big problem. You need to decontaminate against trillions of tokens, which is a very expensive process. With regard to budget forcing, you also in the paper compared that to other potential approaches like rejection sampling. Can you talk about kind of the compare contrasts for those other approaches? One very simple idea is to do rejection sampling. And we thought that rather than having to have explicit control before we test the model, we just do it ad hoc after we tested the model.
30:41So we just sample from the model multiple times. And naturally for the same sample, if we, for the same question, if we sample from the model again, so we let it try to solve the question again, then it will solve it slightly differently as long as there's some non-determinism, which usually happens by having temperature sampling or something like that. And then this way we have a length of a different, a chain of a different length. And we do this multiple times and then we just sort them such that we, get chains of different length for all the problem and look at how the scaling behaves. And surprisingly, what we found is that it's actually inverse.
31:21So it scales the exact opposite as we'd like it to. With more test time compute, performance gets worse if we sample and rearrange this way. And my hypothesis for why this is happening is that when we just try to choose the longest generations from the model, then often these are the generations where the model was going off on a wrong track and backtracking multiple times and ultimately ending up with a completely wrong answer. Whereas the ones where it was correct, maybe it just jumped to the right answer right away. And that's why they're pretty short. And so there's this inverse scaling trend for rejection sampling.
31:56And also it's not a very clean method because you can't control it prior to prompting the model, which is what we want, but only after generating multiple times. Now you've released this model and the training data as open source. Is everything that is required to replicate this available in the repo? And have you had any reports of folks successfully doing it? Yeah, we've replicated, we've released everything and there have been people that submitted issues and replicated parts of it. I haven't seen a full replication, but I'm pretty sure it should be possible. Everything is open source and the training only takes 26 minutes.
32:36so if you can afford the GPUs for those 26 minutes which is roughly$20 then you should be able to do that I think one part that we cannot open source is related to the base model because we didn't train that ourselves so we used the Quen model from Alibaba and we don't have the pre-training data for that one but everything we do on top of that model which is our paper is fully open source Yeah. And did you explore other base models? We did not explore other base models, but there have been some people trying smaller versions of QAN. And they also reported some gains. But it seems like QAN is a really great model for these experiments, probably because it had very long pre-training with lots of potentially already reasoning-like questions in there.
33:23And so in this post-training stage, we're able to really extract this behavior with just 1 ,000 samples. In this example, you distilled off of, well, Gemini and then R1 and then used that to train S1 based on Quinn. I'm trying to think of if it makes sense to distill R1 and train a smaller R1 or tune a smaller R1 or something like that. And kind of, you know, based on that, like any thoughts on like crossing model families and that kind of thing and how that might play out? Yeah, for the R1 model, I think they used their DeepSig V3 model as the baseline and then further trained that. And that might be a better choice than our QAN model because it's a lot larger, a lot more powerful.
34:22so we might be able to get better performance if we were to fine-tune that one though it's a lot bigger and a lot more expensive and then there's of course the llama model family which is very popular but based on some very small scale experiments we did and experiments i've heard from others they don't seem to work quite as well for these reasoning tasks and that's why we didn't use them and i haven't seen any other replication attempts of 01 that were based on these models with big success. I remember seeing the tweet, I think, about a paper that looked at, I guess the context was like LLM as judge.
35:05And it seemed to be implying that like asking a model to judge answers from the same type of model like doesn't perform as well because there's some kind of bias there. And I guess that's what I was thinking of. like, are there, you know, biases or interplays, you know, between model families such that, you know, it's either better or worse to train across model families in this way? Yeah, maybe one intuition is that it could be better to cross multiple model families. So you get some of these assembling advantages where you have models with different capabilities and ideally you get the best out of each.
35:44But I haven't seen exact attempts on that one. One more thing on the evaluation side is, I think this is a pretty big issue, especially because we often use ChatGPT or CLOT to do our evaluations. So we ask them, is this correct for questions where maybe there is no ground truth, arithmetic answer or something. And yeah, if we then evaluate the models themselves, maybe there's some bias and that could of course make the evaluations pretty difficult, yeah. And did that come into play in your evals or were your evals relative to benchmarks that had answers. Yeah, our evals all had answers, but because the answers are very difficult to extract sometimes.
36:27So when the model generates a long math reasoning trace and then it uses some formula to express its answer, then it's very hard to just do a basic string match in Python. But rather what we do is we ask OpenAI's models, so maybe GPT-4O, tell us if these two answers are the same. And if yes, then we give it a go. So tell us if the ground truth solution we provided and their answer from a model are equivalent. So we do use them, but probably no bias from that. A couple of things that came up that I wanted to touch on were SFT supervised fine tuning versus reinforcement learning, which is what the R1 folks did, as well as I think there's an effort at hugging face to replicate R1.
37:17I think it's called like Open R1 or something like that. And I was wondering if you can kind of, you know, compare, contrast SFT and RL and also what Hugging Face is doing with regard to replicating R1 relative to what you've done. Yeah, sure. So for the RL versus SFT one, as you mentioned, R1, they did RL as OpenAI also claimed in their blog post and we only did SFT. And while we can get really good performance with SFT, there's some speculation that RL generalizes better and the intuition behind that is that RL incentivizes the model to arrive at a better answer regardless of its approach versus for SFT we force it to follow a certain approach because we train it on the tokens themselves that lead to a certain answer and that could even be an incorrect answer and so potentially with more compute, the RL approach might be better than the SFT approach.
38:20There's an interesting saying, I heard this from Hyeongwon, that might give some more nuance to this, which is, we have this comment saying, give a man a fish, you feed him for a day, teach him how to fish, and you feed him for a lifetime. And teaching him how to fish is a little bit like SFT, so you directly show him how to fish. And then maybe the RL version of that would be teach him the taste of fish and make him hungry. And so the model or the fisher will develop its own methods to fish and just to get at that goal, right? It doesn't matter what exactly is the process. And maybe he'll figure out that the best way to get there is to have a big fish farm rather than fishing himself.
39:04And that maybe captures the why RL might be better at generalizing. But I think there's more research to be done on this. And before you take on that next part, can you dig a little bit deeper into how RL is used in the context of R1? Yeah, so for R1, they also have an SFT stage. So they start with SFT and then follow up with an RL phase. And that leads to their R1 model. They also have another R1-0 model that skips the SFT and directly does RL, but the R1 model performs better. So there's definitely some trade-off there where if you don't have this SFT face, then maybe the model goes on into a very different direction.
39:49One problem, for example, is that the model might not be using the right format you want because you never explicitly force it to use certain tokens. And then the model might just go off and generate text that's really hard to read or hard to format when you try to use it. And how is RL used specifically in the context of the training recipe here? They use an algorithm called GRPO, Group Reference Policy Optimization. And I think it builds on an earlier algorithm from OpenAI called PPO. And that one is used during training of R1. Is it choosing between, you know, generated responses? Is it somehow influencing the generation itself?
40:30Is it like where it is? What is the algorithm optimizing? So with PPO, the way it works is that there's a separate reward model and the model generates something. And then that reward model will tell how good that generation is. And then we optimize based on that. So the model just gets a reward for its full text generation without forcing it to generate some text in particular, but just it can generate whatever it wants, but then it will be graded based on how good that generation is. and that's then backpropagated into the model as a signal. So it wants to maximize that reward. And if the reward model is really good, so it grades text that leads to the right answer and has the right formatting or reasoning trace or whatever we want, then that should lead to a model that will be really, really good.
41:21But of course you can hack the reward model, which is often referred to as reward hacking. Maybe there's some way where the model will generate some really weird things, But nonetheless, that leads to a really high reward because it exploits some weakness of the reward model. And I think that's also why RL is generally a bit harder to get to work and sometimes very unstable. And for replication approaches, it's much simpler to just start with SFT. And it seems like the reward function would have to be like a model inference in order to, you know, judge the generated text. And that sounds very expensive to do at scale.
41:56Yeah, that's a great point. Which is kind of at odds with like R1, you know, being much cheaper than prior efforts. That's a great point. It's also much more expensive because you need to have that second reward model in memory. And depending on whether you use GRPO or PPO or other algorithms, there are other things you need to keep in memory too. So it's a lot more expensive and training is more difficult to get to work. and that's why for I think the simplest approach S1 that shouldn't have any place there but we just go with something that just uses the model itself and simple SFT. And with regards to the hugging face open R1 have you taken a look at that?
42:37Yeah I think that ties back to the discussion on R1 versus S1 which is that for open R1 from what I can tell they're also trying to go all the way to replicate R1 and then potentially O3 and so on. Whereas for S1, we wanted to find out what's the simplest approach. So no RL, pure SFT, whereas for OpenR1, they're going to involve a lot of RL, I presume. But other than that, they're very similar, OpenR1 and S1, in terms of openness and sharing everything. And maybe we'll also go in that direction of OpenR1 with S2 or further. So in a nutshell, it's kind of the difference between replicating the methodology versus replicating the results with a simpler approach.
43:22Yeah, with a simpler approach. And also, I think they're more interested in the reasoning performance. So what's the bottom line performance you get versus for S1, we cared a lot about the scaling. So do we get test time scaling where better performance can be achieved by more test time compute? and I haven't seen OpenR1 focusing on that but we will definitely focus on it I think for S2 and so on. Along those lines like is there a generation time like is there a signal like a tunable parameter that you're giving to the model to tell it that it can spend more time on compute? Yeah I think currently it is only the budget forcing for a final model but we did experiment with other approaches.
44:09One very simple idea we had was, let's just tell the model in the prompt, you've got 10 ,000 tokens, go reason, figure it out. But unfortunately, the model doesn't really care what we tell it and it'll just reason for much longer or do whatever it wants. Even though we tried to SFT that in, so we trained the model on these kinds of prompts and then the reasoning trace we trained it on would always adhere to that limit, but at inference, it didn't care. Maybe it's just an issue of scale. Maybe we need to have a better model in the first place, But that would be a very intuitive way to do it. And we experimented with some other things in the paper, like having steps instead of tokens and so on.
44:49But ultimately, the budget forcing one led to the best scaling. Do you see S1 as the result of, you know, an experiment and interesting for what it teaches you? Or do you think it is a useful model for folks to use for some specific types of problems? Yeah, I think it is useful for a lot of research where knowing what is the exact reasoning data is important. So if you want to look at what parts of the training data led to certain capabilities, then S1, as far as I know, is the best model right now available where the data is open. So for O1, you don't know. And for R1, we don't know. But if you only care about the capabilities, then you should probably use O3 or R1.
45:37But I think similar to the pre-training regime, there is room for one fully open version of this reasoning paradigm. So since R1 and O1, they probably won't release their data or their code fully open source. There's definitely room for one line of work, be it S1, S2, S3, that makes everything fully open and tries to keep pace with the others. So kind of like the Elmo or Ulmo for reasoning. Yeah, exactly. Like, yeah, in pre-training, there's almost sort of trying to do that, whereas Lama and Quen and so on, they don't share their data. Is that the direction that you're trying to go with this? Or, you know, how do you see this project and your research broadly evolving?
46:20Yeah, we're definitely thinking about that. To me, the most... It's a big commitment, I think. It's a big commitment. It's essentially, might be the rest of my PhD, just trying to iterate on that. But I think it is very exciting. It might be a bit non-traditional because usually in PhDs and also many people told me when starting this project is that you want to do something novel. You don't just want to try to replicate other models. And we ended up doing something novel, which is the budget forcing. And I think replication efforts are often much more important and impactful rather than just trying to do something novel for the sake of novelty.
46:57So I'm willing to kind of depart from the norm there. but what I'm really excited about is making the test time scaling that we've shown more clean and scale further so right now there's two key problems with it one is that it does eventually flatten out we briefly touched on this but if you do wait too much then eventually the model will just go into a repetitive loop and it won't get to the correct answer so how can we scale further when we want the model to think for not just minutes, but for hours or days or a few years even for really, really hard questions. And the second part is, which is kind of related to that, is that our model will run out of context window if it thinks for too long.
47:42The technical detail here is that these models have a fixed context window that they can operate on. And if they generate too much, then it doesn't work anymore, at least for most models. Some models have an architectural change that fixes that, but most models still have this limitation. And so for us, we weren't able to scale beyond 32 ,000 tokens with sequential scaling at least, but we were able to use parallel scaling techniques like majority voting to scale further. So we just generate multiple times in parallel. And even though all of them reach 32K max, together, they do a lot more than 32K.
48:17And then we just aggregate them to hopefully arrive at a better answer. So solving these two issues is really exciting to me, such that we can use these models to solve some of our hardest questions like curing cancer and so on. Great, great. Well, Niklas, thanks so much for jumping on and sharing a bit about what you've been up to. Thanks a lot for having me. Thank you.
From the publisher
Today, we're joined by Niklas Muennighoff, a PhD student at Stanford University, to discuss his paper, “S1: Simple Test-Time Scaling.” We explore the motivations behind S1, as well as how it compares to OpenAI's O1 and DeepSeek's R1 models. We dig into the different approaches to test-time scaling, including parallel and sequential scaling, as well as S1’s data curation process, its training recipe, and its use of model distillation from Google Gemini and DeepSeek R1. We explore the novel "budget forcing" technique developed in the paper, allowing it to think longer for harder problems and optimize test-time compute for better performance. Additionally, we cover the evaluation benchmarks used, the comparison between supervised fine-tuning and reinforcement learning, and similar projects like the Hugging Face Open R1 project. Finally, we discuss the open-sourcing of S1 and its future directions.
The complete show notes for this episode can be found at https://twimlai.com/go/721.




