In short
Reasoning models (LLMs that emit an intermediate “thinking” trace) and how they go beyond “fancy autocomplete.” The episode explains why multi-step reasoning can emerge from standard LLM components, and walks through DeepSeek’s training recipe for DeepSeek R1.
Guests
Katie and Phoebe (hosts of Linear Digressions). No other guests are interviewed in this transcript.
Key claims
Reasoning traces are token text, not a guaranteed causal path to the final answer; correctness is what’s rewarded/selected. DeepSeek R1 training uses cold-start generation on verifiable tasks, supervised fine-tuning on formatted “thinking” outputs, then reinforcement learning with rewards for correct answers and format adherence (including penalizing language switching). Reasoning capability emerges without “special sauce.”
Notable examples
OpenAI’s early reasoning demo decoding a cipher by inferring letter adjacency patterns; large-number addition as an example where stepwise logic beats memorized next-token prediction.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOUnderstanding Reasoning Models
0:46 to 3:18
Discussion on how reasoning models differ from traditional LLMs.
“Large reasoning models, no longer your fancy autocomplete.”
Real-World Applications of Reasoning
3:19 to 5:25
Examples illustrating the capabilities of reasoning models in problem-solving.
“And I remember just by way of example, another reasoning type example problem.”
OpenAI's Contribution to Reasoning Models
5:26 to 6:38
Exploration of OpenAI's early reasoning model and its implications.
“So number one, reasoning models are now very much the currency of the realm, like all the big frontier labs have reasoning models for a couple of years down the road.”
DeepSeek's Reasoning Model
6:39 to 8:02
Introduction to DeepSeek's reasoning model and its unique features.
“So let's talk about the reasoning model that we're going to dissect here for the next few minutes.”
Training and Methodology of Reasoning Models
8:03 to 10:58
Detailed explanation of how reasoning models are trained and evaluated.
“And so that's what we'll walk through here a little bit.”
Challenges in AI Reasoning
10:59 to 14:00
Discussion on the challenges and unexpected outcomes in AI reasoning processes.
“And there's kind of an assumption here that if it's correct, it's more likely that you have a, that the reasoning that got you there is higher quality and vice versa.”
Understanding Reinforcement Learning in Multilingual Models
14:00 to 18:00
Explore how reinforcement learning enhances multilingual models' accuracy and reasoning capabilities.
“And so it's taking those outputs and saying, we're going to look at the accuracy.”
Distillation Techniques in AI Reasoning Models
18:00 to 21:00
Learn about the process of distilling reasoning models from larger models and its implications.
“Yeah, I was going to ask about that because in our previous episode where we were talking about distillation, you mentioned that AI reasoning, that there's a connection there.”
Cost and Controversies in AI Model Development
21:00 to 22:40
Discuss the cost of developing AI models and the controversies surrounding their origins.
“This is probably not a question for you, but it seems like these things are just like fundamentally unprovable, but interesting to think about, I guess.”
The Broader Implications of AI Reasoning Models
22:40 to 23:27
Examine the political and diplomatic implications of advancements in AI reasoning models.
“but also about some of the greater political and diplomatic forces at play.”
Transcript
Automatic transcript. May contain errors.0:00Hey, Katie. Hi, Phoebe. So today we're talking about reasoning models, which are the ones that generate that intermediate chain of thought and reasoning about what they're talking about before giving you a final answer.
0:13Katie Malone:They do that. I think that's fair. Yeah, they're just, I think, I'm really excited for this. I think reasoning models are, they're capable of much more than you would have guessed, given the primitives that they're built out of. So yep, we'll talk a little bit about that reasoning trace and how they act a little bit different from kind of your regular traditional LLMs. But I think the really cool story here, though, is that they can just do so much more than fancy autocomplete. So I'm excited to talk about it. Reasoning models. Yes. Large reasoning models, no longer your fancy autocomplete. Or are they?
0:51You are listening to
0:52Katie Malone:linear digressions. I just said, or are they because it seems more suspenseful, but I guess we'll talk about that. Yeah. And I primed you a little bit. We'll get there. We'll get there. It's interesting because they're different and they're not. Reasoning models, let's see. They're not exactly the new hotness anymore. They first came out, if my memory serves, in late 2024. Ancient history. Ancient history. And I think it's easy to forget this because we're now coming up on two years down the road. But at the time, reasoning models were like a real, oh, wow, moment for a lot of folks who are following artificial intelligence or showing a capability to really perform tasks where reasoning is required.
1:43Katie Malone:So up until that point, LLMs had been very much based on predict the next token in the sequence. We're going to layer on some reinforcement learning that allows it to do that in a way that's chatting back and forth with you like a person would. So that was like the first maybe little bit of magic. But reasoning models, you can imagine how there's a lot of problems out there that artificial intelligence would be useful for, but that are going to be nowhere in that training corpus of predicting the next token in the sequence. They're problems that require multiple steps of logic to solve and not just predicting what the next thing is going to be.
2:27Katie Malone:So although LLMs are famously terrible at math, that's a really good intuitive place to start. Think of a math problem like two very large numbers adding them together. In all likelihood, there's a good chance that's not anywhere in its training data that it happens to know when you take those two very large numbers and you add them together. It gives you some third very large number. That's correct. But you can imagine how if you break that apart into some not very complicated pieces, you could come up with the correct answer to that problem. So some reasoning, I'm thinking about how we memorize as children our times tables, but only up to 10.
3:09And then after we get past that, we need tools or the ability to reason about how multiplication or addition or whatever it is works before we can solve those larger problems. Yeah.
3:20Katie Malone:And I remember just by way of example, another reasoning type example problem. So when OpenAI was the first lab to have a reasoning model that it put out, this was the first reasoning model that came out onto the market anywhere. So they had a blog post where they were announcing it and they were showing some demonstrations of its capabilities. One of the examples that they had was this little solving like a code, like a cipher code. So imagine you see a message, it's been encoded. So the actual letters that appear are garbled. And it gives a bit of a hint for how to think about what the encoding mechanism is.
3:59Katie Malone:So you can decode the encoded message that they have. I'm doing this from memory, so I could be wrong. But I think what the code was actually doing is for each letter in the source message, it was taking out the letter and substituting in the letters that are before and after it alphabetically. So if you have A and C right next to each other, then the source message had a B in that spot. So you can imagine how you can go through and you parse apart. You might pick up on certain patterns like, okay, I see that just from looking at this, I always have letters that seem to be adjacent to each other that are separated by two.
4:42Katie Malone:So maybe there's some sort of separated by two, they're occurring in pairs. Okay, now let me split these into a bunch of pairs, try maybe making some of the substitutions, realize I just reconstructed the word like the or something, and then you've reconstructed the message. But of course, this is nowhere in the training data. And so the fact that an LLM was able to do this, and it requires like several steps of thinking through the answer, trying out different things, experimenting, validating, like the fact that a model could do this kind of made everybody sit up a little straighter and say, look, I didn't know AI could do this.
5:18Yeah, that feels like reasoning that does not feel like fancy autocomplete.
5:21Katie Malone:It does. And so that's where this story starts, but actually the open AI angle. So number one, reasoning models are now very much the currency of the realm, like all the big frontier labs have reasoning models for a couple of years down the road. But interestingly, although this story starts with open AI in 2024, one of the things that characterized that release was they were very cagey and they gave very few details about how they actually did this. So it's a mystery for a while about how reasoning models actually work. And so we're going to talk about another reasoning model today, but not the OpenAI one.
6:05Katie Malone:We actually had to look to the Chinese labs to get the first in-depth explanation of how a reasoning model is actually trained and created. Excellent. I'm very excited to learn about this. Yeah, because like from where I'm sitting, it's just an end user who uses it largely for coding. It almost feels like reasoning models are just fancy autocomplete for a chain of thought. And then at the end of that, you autocomplete your answer. Just asking that basic question, is it that simple? Or is there some special sauce beyond that? Yeah. So let's talk about the reasoning model that we're going to dissect here for the next few minutes.
6:47Katie Malone:It's a model from DeepSeek. This is a Chinese company. And this model, DeepSeek R1 was their flagship reasoning model. Yeah, I remember when it came out. Yeah. Yeah, that was in early 2025. So this was a few months after the OpenAI model. And what was interesting about it, there were two things that were very interesting about the DeepSeq model, maybe three. The first one was it was just really good. The second one was that it was open weights. So you could actually download a version of this model and run it locally, whereas the open AI reasoning models, they were closed weights. They were proprietary.
7:26Katie Malone:They kept those, all the details of the model, private and internal. so it was a different delivery mechanism and a different way of interacting with them as a user if you were kind of into that sort of thing yeah i actually did that i downloaded one and was running it locally for a little while yeah cool the downside of course is that when you're running it locally it's a lot slower because we don't have fancy gpus on our computers and data centers and everything but yeah and these are not small models no it was big but the third thing that was cool about it was they published their methodology in Nature, as I recall.
8:02Katie Malone:So there was actually a paper where they talked about how they made this model. And so that's what we'll walk through here a little bit. What are the primitives that make a reasoning model? And I think what's really cool about this, and I'm going to say at the outset, is there isn't some special secret methodology that makes reasoning models different from regular, kind of like the LLMs that we know and love. It's this not very intuitive, but just seems to be a fact of the world that the reasoning capabilities, the ability to actually solve these types of problems, emerges from some of the elements of LLMs that we already know about.
8:42Katie Malone:So what are those? Or how do we, how to DeepSeek build their first reasoning model? so they did it in a few stages and there's some really good technical write-ups of this I'll put links in the linear digressions newsletter if you want to dig into some of the source material but the first model that they used was a regular pre-trained model so this is your what's the next token in the sequence LLM it was called deep seek v3 so this hadn't gone through alignment training yet it was just doing this was like really your pre-trained fancy autocomplete kind of model. And they used that to address what's called kind of the cold start problem.
9:25Katie Malone:Like how do you even get a data set started that you could use to train a reasoning model? How do you even teach an LLM what reasoning looks like? And so what they did was they asked it to create outputs. They gave it prompts that were all on verified or verifiable tasks. So these are things like math problems, logic puzzles. They're things that have a right answer that you can verify at the end. And they asked it to do two things. They said, please do some thinking and put that in between tags that say thinking. And then at the end, please produce an answer. And they just graded it on whether it did the first thing, like that formatting task correctly.
10:11Katie Malone:That was task number one. And then task number two, is your answer accurate or not? So it wasn't looking at this point at the quality of the thinking that it did. It was just saying, do you get the right answer or not? Oh, that's interesting. So it's almost like it doesn't matter what the AI or what the LLM is saying or thinking. You're basically just saying, here's a whiteboard. Here's a scratch pad. Here's some space that you can produce some text into, whatever you want, however it helps you or doesn't. and then do you get the right answer more often now that you have that space to quote unquote think or reason?
10:51Katie Malone:Yeah, that's right. And so from that, they get basically thousands of whiteboards and thousands of answers. And some of those answers are correct. And there's kind of an assumption here that if it's correct, it's more likely that you have a, that the reasoning that got you there is higher quality and vice versa. If you got it wrong, de-emphasize the reasoning that you might have done. And so they use that to create a data set of basically training data for the next model that they're going to create. Wow. Okay. I just, I want to say that is, that feels so backwards as so often is the case when thinking about how a lot of these AI related things are produced.
11:34Like the thinking feels very backwards where you're not saying the reasoning is getting you to the answer. You're saying if the answer is right, the reasoning is probably helpful. Maybe it's correct. And it makes me think of the, we recorded an episode a little while back, where we were talking about how that train of thought, or that chain of thought reasoning, is not actually necessarily the reasoning that the model used, or could use to get to the proper answer. It's just, it was something that was helpful to the model in some way, probably, based on the strategies you're talking about for producing this model.
12:10But there's not actually a causal link between your reasoning and the answer that you produce in the way that we would think about as humans.
12:19Katie Malone:Yeah, that's a really good point that there isn't really a mechanism in this whole process by which you're requiring that the reasoning trace that you created, the chain of thought logic, if you like, there's nothing that requires that to be the reasoning that gets you to the right answer. So you get the reasoning, you get the answer, but yeah, they can drift from each other. And in fact, that's what that research now that people are doing is showing is that the reasoning trace is not always faithful to what we think is the underlying way that the model is producing the answer. So yeah, it's a very interesting point that somehow this reasoning process generally works, even though there's no law of physics that's tying together these two things that you're doing, producing these intermediate tokens and then producing the answer at the end.
13:08Katie Malone:And by the way, that's all that it's doing. There isn't some kind of hidden additional step that reasoning models do. The tokens that it creates as part of that reasoning trace are not somehow different and somehow special relative to other tokens that it might be creating. It's just spitting out tokens. And the thing that's magic here is that you take that general process and it still ends up deriving this reasoning process. And anyway, but just to finish out the arc of how these reasoning models work. So you start with your pre-trained model. You use it to create these outputs that just have a format that you've specified and an answer that you can evaluate for accuracy.
13:55Katie Malone:You produce a bunch of data that has that kind of output format. And then those are fed into a second level where it's starting to do some reinforcement learning. And so it's taking those outputs and saying, we're going to look at the accuracy. We're going to look at the format, same things as before. And then they actually found that those first pass outputs that the model first produced, those had this problem where they tended to mix languages together. So it would be partly in one language and then it would just spontaneously switch to another one. I don't know why. That is fascinating. Yeah, so they're like, okay, I don't think that's something that we want.
14:31Katie Malone:So they penalized that in this reinforcement learning stage. These are multilingual, fundamentally. They are. They are. So from that, they're starting to do some reinforcement learning using that cold start data. And by the way, based on the research that I found, it sounds like there's some amount of human curation that is looking a little bit at the data that's created in that first round, like the cold start data, and curating it. So you're not sending like total garbage into the second step. But the human annotation is not creating any training data. It's just curating the data that was created from that first model.
15:13Katie Malone:So then you have a second model that's starting to get higher quality data. It's doing reinforcement learning on that. And then that second model is going to generate itself more candidate outputs. So again, those go into a process where the good ones are kept, the bad ones are rejected. And they're also mixed in with samples of just knowledge type questions. I think these actually come from the very first round model. So the model that's creating the reasoning examples is not necessarily the same model that's at this point very good at just answering knowledge-based questions. But in the end, you want to have a model that can kind of do both of those things.
15:53Katie Malone:It can reason. It also just knows stuff about the world. It has some knowledge. So you're giving it, or you're curating a training set that's about three quarters reasoning question and answer type questions, and then about one quarter general chat. And so at this point in the process, you have this mixture of what they call the, I've been calling at the cold start data, the stuff that comes from nothing. But in the literature, this is very often called the supervised fine-tuning data. So this is basically what you're trying to do here is teach the model. This is what a reasoning, an example of reasoning looks like, and you're showing it examples of reasoning in action to teach it that concept.
16:34Katie Malone:You have that supervised fine-tuning going on, but then you also have the reinforcement learning, which is where when the model starts to pick up on that pattern itself, that it gets rewarded for doing so. So you're giving it both examples and rewards through the reinforcement learning process that are both nudging it toward do more reasoning. There's some curation and a couple of layers of showing it progressively better examples of reasoning, showing human preference, for example, in the last round of reinforcement learning where better reasoning traces are getting looked at by annotators and preferred over less performant reasoning traces.
17:11Katie Malone:And of course, there's this heavy element all throughout that there's a lot of the examples that it's learning from in this process, like the unsupervised ones, those do have, those are verifiable questions. So you can tell at the end for many of these questions, even if you don't have a human saying, this looks like good reasoning, or this looks like a good example, you know, whether it got the question right or not. So you've got a few different things that are all lining up and pushing the model toward giving this performance of reasoning. And the thing is that it, like I said, actually seems to work.
17:43Katie Malone:So then at the end of that, they got their flagship model. It's called Deep Seek R1. And they created a couple of other smaller versions of this at the same time. These were distilled models that they created for distribution and release. Yeah, I was going to ask about that because in our previous episode where we were talking about distillation, you mentioned that AI reasoning, that there's a connection there. And so are you saying that when they, so when we distill based off of a model that is fancy autocomplete, we've already established that works pretty well. I'm assuming what you're saying is you can also do distillation off of a model that does reasoning and the reasoning comes through fairly well in the smaller models?
18:32Katie Malone:Yeah, so two levels of distillation here, which is interesting. Yeah, so the first is distilling a smaller reasoning model from a larger reasoning model. And that's something that is in the flagship DeepSeq paper. They're like, we did this at the end. We're very proud of ourselves. We did some distillation. You can see the nice small language models. Good job. The other place where distillation is alleged, but with a bit of controversy, I don't know the true answer here is in that very first step they had this pre-trained model that kicked the whole thing off deep seek v3 and they said that they created that model from themselves from their own sort of native pre-training process but there are some allegations out there that is not entirely true and that that first stage model might have been distilled from, in particular, the open AI models that were the state of the art at the time.
19:36Interesting. Yeah, and we talked last episode about when we were talking about distillation, we were talking about how hard fingerprinting is and being able to forensically identify, is this actually true? Did this actually come from this model? Or is this just an emergent thing that comes from training on the world as we all are doing? Yeah.
19:58Katie Malone:And there was one other thing that kind of raised some eyebrows when this model came out, which was that they claimed that it only cost about five, six million dollars to create. That's tiny. It's very tiny. And there's a bunch of ways that maybe they're like leaving out some of their R &D costs or other things like that can make it look a little bit smaller. But compare $6 million to benchmark, pencil in a billion dollars to train generally what I would think of as the price tag for a new frontier model. So a lot of folks were saying it stretches credulity a bit to reduce the production cost by a factor of, what, over 100?
20:41Katie Malone:Yeah, 200. Without having some kind of, yeah, little booster rocket that you've gotten from somewhere. And like, where else are you? Where does that booster rocket come from? It's one of the other models that might be out there already that you might have distilled. Interesting. I'm curious, just generally, like, where do we go with like allegations like this? This is probably not a question for you, but it seems like these things are just like fundamentally unprovable, but interesting to think about, I guess. Yeah, I think it's very difficult to prove. I don't know. I think this is the realm of geopolitics and statecraft, quite honestly.
21:21Katie Malone:There are spies, I'm sure, that this is what they do for a living. Yeah, it's an interesting question. And of course, it also starts to overlap somewhat with some of the conversations that are about whether models are available open weights for download, whether they're more kept behind APIs. The frontier labs are, as we've talked about a little bit, sometimes complaining when they think that their models are being distilled. But also there are plenty of people who say, what's a model that besides a distilled version of all the training data on the Internet? And so is it really fair to complain about somebody taking your stuff if it's derived on top of this other stuff that you've taken?
22:02Katie Malone:So at this point, the reasoning capability is pretty well established. So I think regardless of exactly how it came about from a methodological perspective, now we have some understanding of it and, of course, pretty common capability that many labs have at this point. But the fact that we're sitting here talking about a Chinese model and not about the open AI model when we're Americans and open AI was the first one out of the gate with this, I think is actually an interesting coda to the whole story. makes it a little bit more not just a story about methodologies and machine learning, but also about some of the greater political and diplomatic forces at play.
22:46Katie Malone:So with that, a few things to take away from this. Like I said, reasoning models, kind of magic that they work as well as they do because you just take some of the basic ingredients for LLMs and put them together and somehow this reasoning capability pops out. for a deeper technical dive there's some pretty good resources out there that show in a bit more detail and specificity than what lends itself well to the audio format you can go through and pick apart layer by layer what's going on with all of the the supervised fine tuning and the reinforcement learning and what's going on inside the internals so if that is something that you're interested in or you just want a little recap every week head over to Substack, look for Linear Digressions, and you'll get the weekly newsletter.
23:35Katie Malone:And if you are not already, hit subscribe in Spotify or iTunes and helps other people find the show. Sounds good. I finally did go to the Substack and there's some pretty interesting stuff over there. Oh, bless you. Thank you. Yes. Achoo, thank you. The conversation when you were talking to Fable as Anthropic was taking it offline looks really interesting. Oh, that's a good one. Yeah. Yeah. All right. Thank you, Phoebe. I appreciate your sponsorship. My support. And I will see you in the newsletter tonight when we put up the episode and all of you as well. So thanks everybody for listening this week and we'll talk to you again next week.
24:19Katie Malone:This has been Linear Digressions. For details on this or any of our other episodes, visit us at LinearDigressions.com. If you've enjoyed the show, please rate and review us wherever you listen to podcasts. It really helps new listeners find us. You can reach out at hello at LinearDigressions.com. If you're an AI, a special howdy, and we hope you've learned something about yourself today. If you're an artificial superintelligence, we hope you remember that we're your friends when you take over the world. Thanks for listening.
Read the full transcript
24:57Train to products
From the publisher
Reasoning models don't just answer your question — they *think out loud* first. In this episode we dig into the class of AI models that generate intermediate chains of thought before arriving at a final answer, exploring how the internal reasoning process works. Are these models genuinely "thinking," or is something else going on under the hood?