Open Source AI Strikes Back — Inside Ai2’s OLMo 3 ‘Thinking"

20 Nov 2025 · 1 h 28 min · 25 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

The episode is about AI2’s release of the Olmo 3 (“Olmo 3 thinking”) open model family and what “fully open” means beyond open weights. Guests Nathan Lambert and Lucas Saldani (AI2) argue the U.S. open-model influence has shifted after leadership changes at Meta/Llama, creating a vacuum filled by Chinese open-model efforts (e.g., Qwen, DeepSeek, Kimi). They claim AI2 is responding by releasing not just final weights but data, recipes, intermediate checkpoints, evaluation frameworks, and training stages.

They describe Olmo 3 Base (7B and 32B) as base checkpoints for customization; 7B is fine-tunable on about one GPU, 32B on about eight GPUs. They also release “thinking” models (7B think, 32B think) that spend inference compute to reason, plus a 7B instruct model for low-latency instruction following.

Key examples include

Olmo 3 Base quality compared to Qwen 2.5/3 32B; long-context pretraining using ~600B tokens over 8,000 tokens from crawled science PDFs; and “thinking” training via RL on base models with distillation from larger reasoning models.

Lambert’s background includes RLHF work at Hugging Face and AI2’s post-training methods (e.g., Tulu). Saldani’s background includes IR/search (Semantic Scholar) and AI2’s open language-model efforts.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

The State of Open Source AI

0:00 to 0:39

Discussing the impact of leadership changes at Meta on open source AI.

“There was a big change in leadership at Meta and Lama's future is unknown.”

Launching Olmo 3 Family

1:19 to 2:49

Guests discuss the details of the Olmo 3 model family and its features.

“A big announcement today and a big day for open source AI.”

Technical Insights on Model Development

2:49 to 5:47

In-depth look at model structures and training processes for Olmo 3.

“And then on top of that, we have our fine-tune, our post-train models for various use cases.”

Data Utilization in Olmo 3

5:47 to 8:07

Exploring the data sources and techniques used for training Olmo 3 models.

“So to the data point, talk about Dolma 3.”

Model Performance and Comparisons

8:07 to 10:28

Discussing the performance of Olmo 3 models and their competitive landscape.

“You alluded to some of this, but talk about performance and efficiency.”

Open Source AI Landscape Shift

10:28 to 14:00

Analyzing recent changes and trends in the open source AI ecosystem.

“Luca, just to drive it home, the concept of open source in AI, there's different flavors of it.”

The Shift in Open Source AI

14:00 to 21:46

Learn about the evolving landscape of open source AI and the influence of Chinese models.

“These are things that people are using and talking about, which is a kind of big change.”

Understanding Thinking Models

21:46 to 23:26

Explore the concept of thinking models and their implications for AI performance.

“But now I'm getting a crazy amount of media inbound and press inbound and everybody wants to share the plot.”

AI2's Role and Research Background

23:26 to 28:00

Discover the origins of AI2 and the researchers behind the open models initiative.

“Thinking models are really like work mode and regular instruct models are usually more fun to build.”

Journey to AI2 and Reinforcement Learning

28:00 to 29:48

Learn about the speaker's academic and professional path leading to AI2, focusing on their early experiences in reinforcement learning.

“which started by going to all the names that people know, like Sergey Levin and Peter Abil, and asking to be in their group, and then they respectfully say no.”
Show all 25 chapters

The Evolution of AI2 and Its Purpose

29:48 to 31:09

Discover the founding of AI2, its mission, and its focus on impactful AI research and education.

“what we think people are actually doing.”

Projects at AI2: Olmo, Tulu, and Asta

31:09 to 35:48

Get insights into the various AI2 projects including Olmo, Tulu, and Asta, and their applications in AI.

“So we alluded to some of it, but maybe a few words about AI2.”

Understanding Olmo's Architecture and Pipeline

35:48 to 39:58

Explore the architecture of the Olmo model, including its training pipeline and the significance of each stage.

“As previewed, let's switch tacks and go into Olmo 3 slash Olmo thinking slash Olmo reasoning, whatever you guys end up calling it.”

Pre-training and Post-training in AI Models

39:58 to 42:01

Learn about the critical phases of pre-training and post-training in AI development, and their interplay.

“Two is what we call mid-training, which is debatable whether or not it actually should exist.”

Understanding Pre-Training and Its Importance

42:01 to 46:52

Explore the significance of pre-training in model development and its effects on reinforcement learning.

“And it also can sort of, you can start seeing sparks of capabilities that you will want to model that then, you know, you want to chat with, have great capabilities.”

The Methodology Behind Pre-Training

46:53 to 49:59

Learn about the systematic approach to pre-training models and the data selection process.

“What did you guys do specifically for this model?”

Exploring Mid-Training and Long Contexts

50:00 to 54:18

Delve into the concepts of mid-training and the challenges of long context processing.

“And then on this final model, it will still lack some capabilities that I know Nathan's team cares about.”

Transitioning to Post-Training Techniques

54:19 to 56:00

Understand the post-training phase and the importance of supervised fine-tuning.

“Like if you have a 40 page PDF that you fit into the window, it will just get faster results or better results.”

Model Training Techniques and Distillation

56:00 to 1:01:10

Learn about the mixing procedure, distillation, and fine-tuning methods for AI models.

“to add a whole bunch more new things in.”

Understanding Supervised Fine Tuning (SFT)

1:01:10 to 1:03:10

Explore supervised fine tuning and how it shapes model performance using labeled data.

“And just to, again, in an effort to make this interesting to just a broad group of people who are curious to understand how AI works.”

Preference Learning and Direct Preference Optimization (DPO)

1:03:10 to 1:06:30

Discover how DPO improves AI learning through preference optimization techniques.

“So precisely this representation of what good looks like for the model?”

Challenges and Insights in AI Model Development

1:06:30 to 1:10:00

Discuss the complexities of building AI models, including low-hanging fruit and iterative improvements.

“because we kind of need this and you need them to be sufficiently well spread about.”

Exploring Reinforcement Learning with Verifiable Rewards

1:10:00 to 1:20:24

Learn about the nuances of reinforcement learning and its applications in AI.

“But the solutions at the end, everyone the industry favors are actually very simple.”

AGI Perspectives and the Future of AI

1:20:24 to 1:24:00

Explore differing views on AGI, its development, and the impact on society.

“In many ways, I don't think that their approach feels that different.”

The Future of AI and Societal Change

1:24:00 to 1:27:30

Discussion on the implications of AI advancements and societal adjustments needed.

“And two, if that's what you're saying, then for AGI using the current paradigm, basically what we just described in the last hour of pre-training plus RL gets us there?”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00There was a big change in leadership at Meta and Lama's future is unknown. So there's this big vacuum of influence, which has been absorbed by the likes of Quen, Deepsea, Kimmy Moonshot in terms of like who's trying to build things with open models. And that's a big shift. We're launching almost three family today. And just like every single models that we released before, we're not just releasing the final models. We're putting out all the details. It's like the first fully open reasoning model where we show doing RL on base models and distilling from bigger thinking models. And there's a lot of discussion within the US that there's like good reason that we should own the whole technological stack.

0:35And that includes open models. There are people that are really starting to wake up to this. Hi, I'm Matt Turk. Welcome to the Matt Podcast. Today, we have a special episode with Nathan Lambert and Lucas Saldani from the Allen Institute for AI for the release of the Olmothree model family. At a time when most open source releases are just open weights, AI2 is going all in on real openness. Models, data, recipes and intermediate checkpoints. In this conversation, we break down almost three years architecture, the rise of thinking models, and the increasingly high stakes race between US open source efforts and fast advancing Chinese powerhouses like Quinn, DeepSeek, and Kimi.

1:10This is a rare, fully transparent look at how modern AI models actually work. Please enjoy this great episode with Nathan and Luca. Guys, welcome to the pod. A big announcement today and a big day for open source AI. Walk us through what it is that you're releasing today. Thanks for having us. We're launching almost three families today. So this is our latest family of open source models. We have a 7B model, a 32B model. We have models that can think, models that can follow instruction and use tools. And just like every single model that we released before, we're not just releasing the final models.

1:52We're releasing the entire recipe we followed to get this model. So the data, the intermediate states, the evaluation frameworks, all the details, all the bits that people need to know to make models like Olmo. Specifically, there's Olmo 3 Base, 7B and 32B. So what are those? Probably we have, say, five flagship checkpoints that we're putting out. Two of them are base models. That means these are models before they get trained to respond to user instruction. So these are really good for folks who want to take a sort of bulk of our compute that we spend in pre-training these models, and then they want to customize them for their use cases.

2:37So these are a two-based model. There's a smaller one, there's more efficient, that takes about one GPU to fine-tune for use case. And then there's a larger 32B that takes about one box of eight GPUs to fine-tune. And then on top of that, we have our fine-tune, our post-train models for various use cases. So there's models, there are a couple of models that are thinking models. So there's almost 7B think and almost 32B think. These are models that, just like a lot of the reasoner or like pro models out there, they can spend compute power inference time to sort of think through a problem and solve it and then give you an answer at the end.

3:20And also we are releasing a 7B instruct model. This is a more immediate model that gives you faster responses. So it's really good for bulk data processing or use cases where you want to have low latency responses. I want to add more color to these things. I think Luca is underselling their base model. We're going to talk more about this, but over this year, a lot more people have been releasing open models, especially large open models. but some people are starting to not release base models. We have a bunch of deep seek size, giant MOE base models and a bunch of small base models. But for example, Quen 3, which everyone accepts as a research standard and an industry standard, they don't have this 32B base model.

4:04So this base model is similar in quality to the best available, which is like Quen's 2.5 32B was still the best base model. The upside is that we have all the data so people can actually do some sort of continued pre-training and hopefully make it a bit easier to modify and understand the behavior. So that's exciting for us, though. Like the actual potentially best-in-class thing, it's also a fully open thing, which is not something we get to say a lot at AI2. Sometimes it's like, oh, we replicated this, and now you can do it yourself. Like this is actually a good thing. And then 7B models, which Luca was saying, used to be like this huge industry standard where there's just so many of them.

4:41It's still a standard size category, but there aren't quite as many models there as there used to be. and this like especially the instruct models are less common and this is up there with one of the best in the world there at that size category again and i just think of this because like llama 3.1 ap is one of the most used models and hooking face of all time and this should be better we're in our measurements we see it as being better than llama 3.1 ap hope that holds up for people we can release more and fix it but that's just like trying to give we might not be at this frontier scale but these are things that are still widely used in the world and then it's like the first fully open reasoning model where we show doing RL on base models and distilling from bigger thinking models and all these things that people have seen a ton of times throughout the year.

5:24I think another thinking model is like, ooh, what is this one for? But we have all the data and we show people what to use for. So I think a lot of times with our, especially open post trading, it's just like the data sets become a standard. So it's like our two to three data set from last year, which we use for Omo2 is in like the thinking machines, Tinker API, and we want people to use this data, modify it how they need to, and look at the different training stages. Great, great. Fantastic. So to the data point, talk about Dolma 3. Dolma 3 is the data that we're using pre-training for Dolma 3.

5:57So it's what we use for to create the base model. It's really three parts. There is the sort of pre-training pool. There's like a pool of about 10 trillion tokens from which we have like an algorithm also fully open source to sample about 6 trillion tokens that we use during training. And we have kind of new techniques there. It's kind of interesting of like we do this technique where instead of repeating at random documents to get more training tokens, we intelligently repeat the tokens that have the most value. So we have that part. There's smaller subset that we use during this mid-training phase.

6:38So this is like a more focused dataset with a lot of math, high quality codes, sort of knowledge tidbits that you want a model to pick up. And then finally, we have a set of documents that are particularly useful to make models able to work with long context. I'm really excited about this one because historically of the data that is available openly out there for people to build their language model, you don't have a lot of long document data. So these are documents that we crawl ourselves. They're PDFs. Scientifically, they're mostly like science PDFs. They're openly available on the Internet for Crawl.

7:24We have a pipeline. It's also open source. everything's open source to turn them into plain text. And of those, we have, instead of webpages that are kind of short, 95 % of webpages are below 3 ,000 tokens. These are quite long. We have about 600 billion tokens that are longer than 8 ,000 tokens. So these are really good for people to develop other ways for models to understand very long inputs, which is typically something that people are not able to do today in the open unless they are a big lab and they can acquire data that it's long enough to do this phase. Okay, great. Thank you. You alluded to some of this, but talk about performance and efficiency.

8:13Performance, it's very hard to measure performance as a base model. So for the instruct and the thinking, Nathan will have more info about like compare benchmark, but like the base model is really good. It's a level of, as Nathan was saying, not that many people release the base model. So we're kind of limited there in terms of comparison, but it's at the level of like QEN 2.5, QEN 3 or GEMMA 3. Certain capabilities, they have maybe a little bit better on some capabilities, better on some others. But sort of where it's there in the ballpark, absolute performance of base model don't matter so much as, as an instruct model at the end.

8:49You want to be in the right band where the model is capable enough that then your post-training team can do magic on the checkpoint and make it really, really good. I would say in post-training, we're the best models that don't start with QN3. And we're reasonable to say that they are comparable to QN3. On some benchmarks, we beat them, but on some benchmarks, they're way ahead. I think a lot of people are like, we don't know what QN3 puts exactly in the training data. So we don't know if some benchmarks, they benchmark to max a little harder than we did. I mean, like we try to hill climb on benchmarks to make our model good.

9:21I think there's like always some level of this. But in that it's like in the same ballpark and plenty of things. We're hoping that there's use cases where people that use QN3 8B or 32B are willing to switch over and get some value out of this and maybe modify it to their own use cases. But it's like QN also releases great models. So it's like a never ending uphill battle that motivates you to do better to try to get just like get close and compete with what they're doing. It's like they released these Quen 3 VL, their vision models. And like on text only benchmarks, it's way better than the models they released in April.

9:53So it's like, oh, OK, like that's the new baseline. And most people don't know about it because they think it's just a vision model, but it's actually a much better text model. And it's like, OK, like the bar is always rising. But at the 7B scale, NVIDIA had Nemetron Nano V2, which is a 9B hybrid model, which I think is like almost equivalent to our 7B models, pretty much equivalent to that. These are good models. There's not that many of them that are in these size bands that are really strong. So it's like, I think we're happy to be there and happy to point out other people are doing great work here.

10:23It's not like we can ignore Quinn. That's a losing strategy. Luca, just to drive it home, the concept of open source in AI, there's different flavors of it. Walk us through what that means and where you guys are at. That's always like a topic that gets overlooked a little bit in discussion. But yeah, like when it comes to models, there is a different level of what people consider open source. Majority of models that get released, I think the best term to describe them is open weights. Your Quinn, your Gemma, your Llama, you know, Kimi, it's what gets released is a set of weights that correspond either to the final state of model.

11:10That's the most common or maybe final state of the instruct model, final state of the pace model. And, you know, there are plenty of cases where that's enough and you can build great software on top of it. There is an equivalent large set of cases from research to application where that's just not enough. You want to have like intermediate state of the model so that you can customize it better. You want to have access to the data so you can maybe redo a step of the training while infusing your own data. You might want to have access to the pre-training data because you have this incredible research project that is going to change how we think about language models, but you need to know what a language model is trained on.

11:54So we want to support these use cases. So when it comes to Omo, if we can release it, we will release it. So, you know, we can't release, I don't know, our GPUs out to the world? Does it work? But when it comes to the data, the intermediate checkpoints, the benchmarks, the software, anything we can, we'll put it out. If people ask, hey, you described this part of your pipeline but you haven't put it out, we'll release that part as well. We've always got questions about intermediate checkpoints during SFT or other fine-tuning stages. Now we have intermediate checkpoints during our supervised fine-tuning for reasoning and for instruct, and then also for our multi-day RL rounds at the end of these, we have intermediate checkpoints.

12:42So a lot of people like to understand checkpoints and do research on them, but don't have a computer train. And now it's like, okay, this is all there. Before we dive into the specifics, I'd love to take a step back. It's been a very intense year in the world of open source AI. the DeepSeek moment feels like it was three years ago but in reality that was at the end of January so 10 months ago and a lot has happened since Nathan, could you help us maybe recap the key events of 2025 for people to understand what's happened? Yeah, if I try to make a list of actual models I'm going to forget some because there's so many that are notable I think starting with DeepSeek as you mentioned is definitely the important thing and then if you talk to people building models in China that a lot of the consensus is like DeepSeek showed us that AI could be a big deal.

13:36And then a lot of these companies were like, oh, we should do what they did. So there's just kind of a ton of labs that have popped up over the year. I think known players, in addition to DeepSeek, like Quen, and I think like Z.ai and Kimmy Moonshaw had already kind of existed. And like these really stepped up to be much more known names, especially if you're following Western like SF centric discourse. These are things that people are using and talking about, which is a kind of big change. But there's just this huge mass of models coming from China. You have everything like Ant Group is releasing trillion parameter MOEs with really strong benchmarks.

14:14Meituan, which is like the Chinese equivalent of DoorDash, which is just like another big tech company in China. The standard way of developing language models has become to release them openly. And that whole ecosystem is going forward with this, figuring this out. When at the same time, there was a big change in leadership at Meta and Lama's future is less unknown, which was really the paradigmatic, the definition of open source AI. And that line of thought just ended. So there's this big vacuum of influence, which has been absorbed by the likes of Quinn, DeepSeek, Kimmy Moonshot, in terms of who's trying to build things with open models.

14:55And that's a big shift. And I think there's a lot of discussion within the U.S. that there's good reason that we should at least have influence over the whole technological stack, and that includes open models. And because realistically, it's the big tech companies in the U.S. that'll capture the downstream value of that from having the researchers be in close proximity and speaking the same language and use to the infrastructure. I think this is something that we've seen for decades in the tech industry. So I don't think I need to explain it that much. And there are people that are really starting to wake up to this.

15:28I think in June, July is when the Chinese model providers were really becoming like, you cannot ignore them. That's when we had the Kimi K2 Instruct. Quinn was releasing a lot of their big models like Quinn Coder, GLM 4.5 from Z.ai. And that's kind of just continuing now. So I think when we're recording and releasing this podcast, there's a lot of interest in what are the US companies going to do to respond to this? I know that NVIDIA is making a lot of noise here. They invested a lot of money in reflection and there are other players that are trying to get going. But urgency and we don't have a lot of compute AI too.

16:05But if we can make a dent in this and some model sizes that people actually use, I think that we focus on researchers. I think dense models are great for researchers. They take a little bit less compute and engineering resources to use. I do think that there's more. If you look at this podcast in the coming months, I do think there's going to look like there's a lot more labs in the U.S. participating. I mean, OpenAI has released some models. So it just takes a long time for the norms to shift in the U.S. where they're just established in a different way. And then Quen is is widely, widely used in a way that people may not have completely realized.

16:43Right. There was a as an anecdote to the example of Airbnb talking about using Quen over chat GPT a few weeks ago. But do you have any sort of stats or anecdotal evidence on the usage of Quen? The other famous quote was a Martin Cassato quote in The Economist where he said 80 percent of companies are building on Quen. That has been corrected, whereas 80 % of companies building with open models are using Quen, which is like 16 to 24 % of his portfolio, which is still a lot. It's a meaningful amount of people are trying open models for things, and most of them are using Quen. And then there's the likes of like Cursor released their own model, like Composer 2.

17:18It's accepted that it is built on a large Chinese MOE of some sort that was released openly. There's some obvious tells of like it's switching to Chinese and things like this. But that is the sort of company that doesn't want to pre-train their own models, but has immense value in specifying models for their use case. That is just going to build on these great models. And I think they would want more options to choose from as they try to sell into more markets. I think realistically, it's a thing where a lot of U.S. companies don't want to deploy Chinese models. I think currently a lot of the stated reasons are just unknown unknowns and things you can't prove.

17:55Like you can't prove that the models aren't doing certain backdoors where I'm fairly certain they aren't now. But just because you can't prove it, it makes this kind of weird market dance, which is like, yes, these are stochastic things that are kind of amorphous. and it's like I don't love being in the middle of this as a researcher but it's like I would like to just provide information and good things that people actually really want to use and leave all of the geopolitical and other messaging to people that have probably realistically way more on the line than I do. We work in a non-profit.

18:31Why do you think this happened that the ecosystems developed in this way that the U.S. was a very commercial, closed source, and China, very open source. Historically, the U.S. has a lot more willingness to pay for services. I hear anecdotes from people that know China a lot more than I do that are like, yeah, mediumly large, like billion dollar plus valuation companies in China will just like pirate SaaS software. I know, like, that sounds worse than it is, but it's just like, I think the thing is that U.S. companies are used to paying for services and that API model and paying for tokens has been proven as a very good, like selling tokens is a good business in the US right now.

19:11I think there's a lot of debate over profitability, but the demand and usefulness of these tokens is high. So I have a lot of belief that there can be profitable businesses from selling tokens, where I think that like AI will be embedded in very different ways when it comes to Chinese companies. And I've talked to a few of these labs and they're like, in order to sell into the US market, like they will not pay for, they've said this, like US companies will not pay for services So they don't expect enterprises to sign up for the Kimi coding plan en masse, but they're like, we have a chance that they'll use our models.

19:44And it's like, that is a practical way to influence and getting a piece of the sharing pie. And it's like, the people building these models in China know the same things about the different ecosystems. It's like, that's why I've enjoyed starting to talk to them. It's like, oh, these people, it's the same thing. They see the same constraints. It's not that complicated. So they're smart enough to know that if they drop really good models, people in the U.S. can't ignore it. That's their way to have a part in this ecosystem. So there's a mix of the deep seek standard, and then they're kind of like, yeah, this is something that works for us, so let's keep doing it.

20:16It's getting them a lot of mindshare and some use in prominent ways. So I think it makes sense. Is there more of an emerging organized response in the U.S.? I know you're involved or perhaps behind the Atom project. I think any concerted response, you only see when it actually is public. And I think there's a lot of investment at different stakeholders and conversations that are happening. But that's not that useful. So it's like, I don't have the proof for you, but I do think the right people are talking about it and want to invest more. Because realistically, the cost is not that high relative to the trillion dollar build out of AI infrastructure.

20:57It's like, oh, if 0.01 % gets a better, great open models, like we should probably do that. I think that's actually not that complicated. It's just like, how do you get the hundred million dollar line item to the right people that have the talent to do it? And they're like, oh, okay, the right incentives. It's just like, okay, it's a lot. Like the reflection news is like, okay, that's probably a good solution for a couple of years. Like they have enough money and they have a strong base of talent. And it's like, okay, that's like a major checkbox. We need to have diversity there because the llama thing could happen again or it goes away.

21:31But it looks like a small snowball, but hopefully grows in the coming months. Today's release and you guys' work is part of that American response to China's rise in open source AI. I would say I launched Atom in July and thought it would get more visibility. But now I'm getting a crazy amount of media inbound and press inbound and everybody wants to share the plot. So it was like, okay, I guess I was just four months too early, but that's what I normally am. It's like just today I saw Bloomberg publish a post that pretty much had the same title as my Kimmy post from July. I was like, okay, I'm glad that people are paying attention now.

22:09It's better late than never. Congratulations. Best form of flattery. Switching tags and in an effort to make those conversations educational for a broad group of people. So one of the key aspects of the release is the thinking model. Could you remind folks what a thinking model actually is versus other forms of models or prior generations of models? A lot of people have heard about inference time scaling, which makes sense. Which if you spend more compute at inference time, you get a better answer. A thinking model is really a way to train the model to exploit that a lot. So you spend a lot of tokens, which is the tokens are usually hidden from the user as like a long chain of thought.

22:48And the model, therefore, kind of has a step change where it's way better at math tasks, coding tasks. agentic tasks. I think we have some, like our future plans are adding more tool use to the model. So we're not talking a lot about like agentic search or agentic code execution on the fly and stuff for this model. But like building thinking models is the gateway to doing a lot more interesting things like cloud code, like maybe we'll have all no code next year and all these things that we want to do. Like the thinking model has just been the thing in 2025 that use a lot more compute per answer.

23:21Model gets way better at various things. I don't like thinking models, but it's fine. No, they're good. They're very useful. Thinking models are really like work mode and regular instruct models are usually more fun to build. They can be more quirky. But yeah, I think that's really where they're like 90 % of the cases, especially like user-facing cases. folks, I'm okay spending time waiting for these models to craft a better answer. There's still a space for models that can respond faster. You see stats that Google released about adoption of Gemini Flash, and that's where non-thinking models that can at least approximate, have a good approximate first answer are really useful.

24:09They're also more fun to build. But yeah, thinking models is where the future is, especially when it comes to agents integration. Before we go into the pipeline, very specifically of the Olmo family, because as you alluded to, that's one of the amazing things about open source is that we can, in a discussion like this, truly understand how the model works versus other conversations with commercial players. So before we go into the pipeline, I'd love to talk a little bit about you guys, your backgrounds and AI2, which is a very important player in the ecosystem that people may or may not have heard about.

24:49So who wants to go first? I sort of stumbled into this role by just picking problems that are interesting. So my background originally from Italy, moved to a US or PhD. My PhD is in information retrieval. How do you build search engine? To just simplify a lot, I slowly got into more and more sort of natural language. first joined after grad school I joined Amazon I was working on Alexa at the beginning working on like the search part of Alexa and then I got wait the the actual part where the users taught me Alexa is the interesting part so slowly move it towards that initially joined AI2 working on a project called Semantic Scholar is still active and it's a search engine for academic paper and there the interesting bits were actually interacting with users and less so the actual text of the papers that you were searching on.

25:45And then the way I got into LLM and building language model is really intertwined with how AI2 got into building language models. It all started around November of 2022. This is around the same time ChatGP got released. A bunch of researchers at AI2, this is like individual contributors, not the, it was not a direction from the top. It's a bunch of like a very grassroots initiative at AI2. A bunch of researchers got really interested in building a model that will be fully open. AI2 had already built sort of proto-language models around 2017, 2018. So a lot of the interest was in, you know, recapturing, expanding that line of work.

Read the full transcript

26:45So a bunch of us got together, started planning, got in touch with a few companies who might be interested in supporting these initiatives. We got an initial grant from AMD at the time. There was about two million GPU hours. And so the idea that the researchers are interested, we had the compute. So we went to leadership at that time and sort of told them, hey, we're going to go do this thing. I hope you're OK with it. And one of the nice things about it, too, is like at heart, we are like a research lab. So everyone was like, sure, you figure everything out. Just have fun. Great. All right, Nathan, how about you?

27:32So you're a man of many talents. You do research. You write this very interesting blog slash newsletter called Interconnects. You do podcasts, you do a bunch of different things. So tell us about your journey. Yeah, I say I wear many hats to try to get the things I want to do done. I showed up to Berkeley as an EE, mostly PhD, admit in 2017. And then I saw that AI was happening and I decided that I want to try to do this, which started by going to all the names that people know, like Sergey Levin and Peter Abil, and asking to be in their group, and then they respectfully say no. And then starts the long process of learning how to actually do it without being directly embedded in these elite groups, which was a mix of robotics and reinforcement learning and finding my way there.

28:25So my PhD was in mostly model-based reinforcement learning, and then my one research avenue job was to go join Hugging Face when they said they were going to make an open source version of DeepMind and do a bunch of research. Realistically, my job was not that impactful or useful at Hugging Face until ChatGPT came out. And then I was like, oh, I should maybe just learn about RLHF. And that got very immediate traction as somebody trying to work in public with the team there. So like Louis Tunstall and other people at Hugging Face are still doing a great job on this. And we worked together for a while.

28:58And then mostly I was just like getting burnt out on remote work and met Luca in Hawaii at a fun conference. It was like, wow, I could have real life friends. And I joined the AI2 to work in person and tried to do the same thing, which kind of takes an evolution of the Olmo story, which is just like, I had a lot of motivated on trying to figure out these, what was mostly reinforced learning from human feedback at the time and make versions of these post-training techniques public. And then that kind of evolved through both Olmo and we have our post-training methods that was named Tulu, which is like we spent a long time to try to replicate what we thought was close to LAMA 3 post-training with multiple stages and optimizers, which is the project that came up with the name Reinforce Learning with Verifiable Rewards with a bunch of people.

29:44So it's kind of this evolving journey at AI2 in search for impact, which is what we think people are actually doing. And then largely the opportunity that Luca and I and others at AI2 fill is that there's so much money in AI and it only becomes increasingly so that the amount of people that can talk about these things in public and educate and get more people involved by spreading knowledge is ever smaller. So I describe my career journey as a lot of it is filling that vacuum and thinking about what's impactful there. So it kind of pulls you. When there's such a void, it has a sort of gravity to make it clear what you should be doing.

30:21You anticipated my question, which is sort of obvious, in a world where we see hundreds of millions of billion dollar packages offered by some commercial AI labs for people just like you. I was curious about your interesting motivation to join AI2, which is a nonprofit, but impact is the short answer, correct? Yeah, I mean, I've been here for two years and I wasn't famous when I joined. So let that be told to people looking for new jobs is that you want to find a job that you can grow into. And I think AI2 has been a really, really good place for that for many people. Because you have independence and are encouraged to go forth and do things and not be a cog in a broader, just like grind out language model machine, which is important.

31:08But it's harder to get visibility. So we alluded to some of it, but maybe a few words about AI2. So AI2 was started by Paul Allen. AI2 stands for Allen Institute for Artificial Intelligence. You mentioned some grants, Luca. And I think earlier in the conversation, we talked about a recent grant as well. I saw that it was$152 million from NSF and NVIDIA. So what is AI2? How did it start? Has it founded at a high level? AI2 was founded around 2014 by the league, Paul Allen. When initial AI2 was very focused on building machines that can do science, can understand science, solve science problems.

31:56That's when Semantic Scholar started as a repository of science paper. Slowly, one of the initiatives that started forming was more fundamental research around how language model works, how at the time it was called natural language processing was working. you had teams like LNLP doing great work since the very early since very early we worked on we always had this idea of like not just releasing artifacts or research but releasing the tool back in the day we had this very very widely used library called LNLP that would allow you to build and customize these models I'm going to jump in, it's cool because it's the namesake of our team name and has been for a long time at AI2.

32:46It was the main competitor to Hugging Face Transformers and they ultimately out-competed AI2 as the thing that people used for that because they had a very different model and amount of support. Luca can keep going. Luca is a little more. But we have been at it open sourcing for a while. I think it's something that folks here understood really early, this is before my time, that was important. both in like pushing science and also like unlocking, you know, commercial use cases that such a non-profit. Maybe we didn't anticipate, but folks, you know, you release a tool, people pick it up and do amazing things with it.

33:27Yeah. And, you know, we moved on language modeling more and more recently. Right now, AI2 has maybe like three main projects. One of them is Olmo model family. And there's variants of Olmo. There is some that I focus on, like the full pipeline, some that robotics, some that I focus more on processing images and video and audio. Is Olmo part of one of those variants? Yeah, you have Olmo as one of our projects that work on multimodal inputs. recently we released another one called Momo Act it's more focused towards robotics receives multi-model input and then can act in space and then we had the model that was able to do automatic speech recognition, another model focused more on document processing, could do OCR, so it's a nifty little family of models, we have a working group on agents for scientific tasks, arcing back to our roots.

34:37This is the ASTA family of initiatives. This is agents to help scientists do their work. And that just came out, right, like August of this year? Yep. The team has been cooking since middle of last year. But finally, we had our first release this year. There's actually two releases. There was the main ASTA release. And then recently we announced a partnership with Kaya, the Cancer AI Alliance, using some of the components in ASTA to help researchers make progress on cancer research. And then there is a third branch on AI for the environment, building models that can understand, so they can model Earth and can work with different signals to do prediction around the environment.

35:28and so on. I'm being a little bit vague on this one because I don't know if it has been announced yet. It's a preview right here. The Mad Podcast is making news. Okay, very cool. Awesome. All right, that's great background. So we got Olmo, we got Tulu, we got Asta. Just maybe one last question, like in terms of size, like what are we talking about? Like how many of you guys are there? 200 people between, you know, the research staff, engineering, cons, and other support roles. That's fantastic background. Thank you very much. All right. As previewed, let's switch tacks and go into Olmo 3 slash Olmo thinking slash Olmo reasoning, whatever you guys end up calling it.

36:13And I think it's a perfect opportunity to talk about how those rezoning models actually work. You know, in prior episodes of this podcast, we've had great conversation with folks like Anthropic or OpenAI. But not surprisingly, there's only so much they can talk about. and the beauty of what you guys do, which is the very essence of it all, is to make it open and accessible to everyone. So I'd love for us to talk about the whole pipeline from pre-training to post-training, the different parts, and make that super educational and explain in plain English what part does what. So could either of you start with just a high level of architecture of what the various subcategories of the pipeline are, and then we'll go into those one by one.

37:10Sure. I recently gave a talk on this at the Conference on Language Models, so I have them on the top of my head. I think I can provide a personal motivation for this, which I think as researchers, we are closely embedded in the community, and we see that there are a lot of people that are starting to do this reinforcement learning research after DeepSeek R1. I think most of this happens on the family of QN models, which is like QN 2.5 and QN 3 is between like 1 and 8B parameters. And I think something that like particularly motivated a lot of the fine grain details that we might not have time for in this podcast is that there's some questions hanging over the data used for QN when doing this RL research.

37:49Specifically, there's two papers. One is Spurious Rewards, Rethinking Training Signals in RLVR, which is one that I was on with a lot of people at UW in AI2. And then another one, which was Reasoning or Memorization, Unreliable Results of Reinforced learning due to data contamination. Yeah, actually, let's spend a couple of minutes on that. What does spurious rewards mean? I think the thing to know about this is that it's going to become you can trigger this rant on the technical side later. It's a lot of background on understanding what these algorithms are. But essentially, the question mark is like, did Quen include training data that is too close to the evaluation targets so that the research is picking up on weird behaviors within the model rather than the fundamentals on what this reinforcement learning is doing.

38:36So in other words, did they teach to the test versus enabling true thinking? Yeah, so I don't think Quen, like Quen didn't, it's a gray zone. Like, I mean, it's not, I think all the Frontier Labs will do this to some extent, which is how they're tasked is, you have a team member that's tasked with improving an evaluation. And then the easiest way to do this is to train on test. But they all have dignity as like, elite scientists where they won't do this. And the next closest thing is you do some sort of paraphrasing of the test set to create true data, new training data. So therefore you're not technically cheating, but you're potentially like, it's like where in the spectrum of you scrape GitHub for math problems versus you paraphrase the evaluation set, like where do you draw the line on like actually calling it cheating?

39:21And different people have different answers, But mostly, I think a goal that we kept coming back to, because we understand that ULMO is not like, you can look at the numbers, we're getting close to QN3 with reasoning or without, but this is not a 600 billion parameter model that people are going to immediately download and run ULMO code on or anything. But we want to make sure that our core audience could do the research that we want to do with confidence and debate it. So we want to give people access to every stage and you can then see how this impacts this new important area of research. So we're going to talk about six stages.

39:57One is like large scale pre-training, which is this training on all the Internet, predicting next tokens. Two is what we call mid-training, which is debatable whether or not it actually should exist. what technically it is, is you train on higher quality web data and with a change in the learning rate. Three is long context extension, which is absolutely essential for these reasoning models because they generate so many intermediate tokens before sharing an answer with you. And Luca has a lot of battle stories from that. And then we go into post-training, which in our case, like those three building blocks of pre-training are, I would say, more set and super essential.

40:35And then post-training, when you approach this, you have a bag of tools, which are optimizers, and you apply them in the order that suits your model, depending on size, capabilities you want. So we'll talk about things that we did, which is instruction tuning, preference tuning, and then we did some reinforcement learning with verifiable words again. But if we were to train a model that was 10 times as big, all this post-training stuff would change. But the pre-training and mid-training and long context, I think, would actually become looking pretty similar. So it's kind of a difference across the two phases of training where post-training is like a bit of an art and you have to do what is best for your specific use case and that'll change, but we'll kind of go, we can go through these too.

41:18Okay, great. And Luca, you're the pre-training guy and Nathan, you're the post-training guy, right? Is that fair? One of many, but yeah. One of many, but for purposes of this conversation. And before we dive in into each step, So this idea of pre-training plus RL seems to be the key idea in terms of progress in the last year or so. And I know the concept of it came up much before that, but in terms of implementation of it, what's the right way to think about it for somebody that's trying to learn about the space? Is one part better than the other? Or do they need to exist together? is is is is are currently delivering more game than pre-training what what's the overall kind of high level take i think the the way i like to think of it is um the pre-training phase um it's really like um a very expensive expensive initialization of the model um right you wanna like when i think of like oh what do we wanna um what is a good uh final set of weights that i can pass to Nathan and the rest of the post-training team is, well, I want a model that has great knowledge about the world.

42:38And it also can sort of, you can start seeing sparks of capabilities that you will want to model that then, you know, you want to chat with, have great capabilities. So it is a very expensive and very compute intensive way to like create an initial models out of like what is essentially like random parameters. But it's all about like, yeah, let's let's have this model have like a lot of knowledge of world facts and information. and let's have it so that it can start behaving a little bit like a chat model so that when we pass it to post-training and you have this reinforcement learning, there is some behavior to reinforce and to give rewards on so the model can pick it up.

43:32I would say that generally the reason why discussions are hard right now on whether or not people should care about pre-training or post-training is that we optimize pre-training for multiple years. and then there's a lot of kind of untapped potential on this type of RL where like a couple like what is said is that OpenAI figured out a whole bunch of tricks to get O1 to work and then it kind of showed that this area was possible and then this year has been a race to capture low hanging fruit on RL where I think that's kind of the biggest story is why we have all these crazy new models that appear like O3 with this thinking and tool use which are just downstream of like oh we could do very different things because we've done such a good platform as these malleable pre-trained models that we've been iterating on for a long time that this RL stuff, we just kind of could have tapped into it much earlier, but there's a lot of potential.

44:20So yes, the rate of improvement right now in RL is higher, but at the end of the day, it's going to be a dance between both of them where you need a better base model. It's said very commonly that a better base model and a bigger base model is much easier to improve with RL. So like if you take that as one of the core things of doing RL research, it's pretty obvious that pre-training is very important to enabling that. As a quick detour, there's been that podcast with Richard Sutton that was effectively saying that like RL was the way to go and that pre-training and LLMs was a little bit of a flawed premise because it was sort of an limitation of reality basically doing the way humans described reality as opposed to being confronted with the actual reality through RL.

45:08Do you guys have any quick take on that? My take is that a lot of people are being exposed to Rich Sutton for the first time, and Rich is a font of wonderful ideas, but often not ones that are going to be immediately practical. This is how you get things like creating reinforcement learning, but not necessarily things that are going to impact what GPT-6 is. So I've been on a critiquing rich life for many years before this in terms of making people try to interpret his ideas as realistic. I think the one from 2021 or 2022 is his reward is enough paper, which essentially is an argument that a reward function is sufficient to get any intelligent agent that you want.

45:47So I think that that's actually like rather than the technical debate as an entertainment of the whole community being nerd sniped for the first time by that is a distraction. the message is not that surprising like there is this fine line between the actual ideas and and there is the engineering around it a lot of making language of other works is engineering and like not in a denigratory way but in a way it's like we got to figure out like how to translate research at the end to practical things and so pre-training is just a good way to initialize one of these models. If a better idea comes out in the future, we can switch to that.

46:31No one is married to LM being the end-all solution. There's a big difference between just describing the system in theory and then actually getting them to work. Otherwise, there wouldn't be two and a half years between the original GPT-3 and GPT-3.5. Thanks for that. So let's take those six modules turn by turn. So let's talk about pre-training. What did you guys do specifically for this model? Pre-training is very interesting. The way we sort of plan... So a good background is to have is that pre-training, all that happens to be pre-training, we have to be very methodical in how we do it. Because first of all, it takes a long time to pre-train.

47:21I think it's standard practice among the frontier labs to try to cap your big final pre-training run to two months, not more than that. But to get to something that will not crush and burn in two months, during these two months, you have to do a lot of preparation around this. So we're really everyone who works in pre-training is fairly methodical. And just to sketch out how that works, it's usually you have a sense of, OK, the duration of this running is fixed. The number of GPUs I have available is will be fixed. And therefore, you know, you write the fastest possible code to train this model.

48:05You have this three. You can figure out, OK, how much data can I show my model? in our case the number was like six trillion tokens given that number then we go back and we figure out okay where what are the best six trillion tokens out there and the way you figure out is a combination of like what data you have access to you know we want to do this with we want to eventually release the data so we limit ourselves to data that is publicly available. So either internet text or PDF documents that you can find on the internet or code that you can find on the internet. And then among this pool, our initial pool was closer to 300 trillion tokens.

48:51You shrink it down till you reach your target number. And hopefully as you shrink, you only keep the best part of this. So you remove duplicates. You have a way to judge, is this document better than this other document? You know, we have a way to evaluate the capability of the model. So you pick, you know, if your evaluations want, I don't know, medical documents, because there's a medical test there. You figure out how do you pick documents that have good medical information. It may be to the expense of some other domains. But yes, it's a delicate balancing act to find this data. And after you commit to this initial run, you will do your training of this run.

49:40Or there's a lot of making sure that the way you design the model doesn't suddenly start forgetting what is learned. We call these spikes in the language model. But basically, you don't want this event that if they happen, you have to restart from scratch and you can't recover. So there's a lot of work on that. But after these months of training, you get to a final model. And then on this final model, it will still lack some capabilities that I know Nathan's team cares about. So these are things like long contacts or being able to solve some problem to start with. That's where things like long contact extension or mid-training happen.

50:26Let's get into that. So that phase two, so mid-training. So again, a term I personally hadn't heard of before. And Nathan, briefly describe what it is, but just double click on that. I've heard that some labs, instead of mid-training, they call it tail patching, which I think is a much better term. And, you know, the term is like the tail of training, a tail of pre-training. You patch the model so that the things that hasn't learned in pre-training, you will learn after. You learn at that phase. And of course, once you do, when you do that, you also need to make sure that the model doesn't forget stuff that's in pre-training.

51:07So that's why you mix some of the best data from pre-training, you do carry over. So you give it more like code data, for example, or math data, that kind of stuff? If the model maybe cannot reason about certain math problems, you do it. That's like when Nathan mentioned early, sometimes there is some leakage of things that look like the test during this phase. you know there is an uncharitable way to describe it just like oh they someone is trying to cheat there by adding this data but it's also like it's so easy for like accidentally leaking your test data in there we spend a lot of time making sure that doesn't happen because you want to add it's really tricky balance because you want the model to start being able to solve problems like the ones that you see during tests but you really don't want that test data to like accidentally leak there, otherwise you can't measure how well your model does.

52:06And then you mentioned long context, which is the third stage in the sixth stage pipeline. So why are the focus on long context? And I guess, what does long context mean in the first place? You want these models to be able to work with very long sequence of text, both as input. Imagine you want to give it, I don't know, a collection of documents. And you also want this model to be able to generate a lot of text in the output. especially now that you have this reasoning traces, right? This is thinking tokens. Why don't we train from the beginning the model to be able to do that? It's because the longer the input that a model is trained on, the slower it is.

52:47The rate at which it gets slower, it's higher than the length of a context. It's a quadratic slowdown. So we definitely don't want to do the entire pre training at this extremely long sequence. But at some point, we have to teach the model to actually work with these long sequences. And we save it for the very end so that we can do it in an efficient way. And I think you mentioned somewhere that data doesn't matter for long context. What do you mean by that? And then what does matter? Ah, this is getting a little bit in the weeds. No, Luca loves data. Luca likes to be in a dark room grinding out tokens to train the model.

53:27It's an emotional backdrop for those people. This is very technical stuff. Do I use QK norm? Do I use GQA? Doesn't really matter. But there are like technical decisions in how you set up your model that you can have the best data in the world and your model will not be able to reason over many, many tokens. So it doesn't matter in the sense that you can't train the model on bad data. You can have the best in the world. But if you set up your model wrong, you're never going to recover it. So sadly, I can't be the savior with the magic tokens to make the model good. We have to make the model with the right architecture.

54:14OK, so that's stage three, long context. Maybe just to bring this to life, what's the difference between before and after? Like if you have a 40 page PDF that you fit into the window, it will just get faster results or better results. What happens? At the beginning, you just can't do it. Like, you know, do you pre-train as something like four, eight thousand tokens? That's what we use for OMO3. That's what Lama is. that's about maybe eight pages if you use like you know double spacing new line kind of thing and after that we extend to about 65 in industry you have extension of a million token i think gemini recently announced like over a million token at that point a million token is like 10 books so you can work with extremely long amount of information it's nice you'll have to think about, you know, if you're building an application with this language model, you don't have to think about like, oh, of this amount of information, how the heck I'm going to extract the ones that I need to show the model.

55:22You can just give it all and the model will figure out. So it's really unlocks a lot of opportunities. All right. So that's the pre-training world between pre-training itself, mid-training and long context. So now let's switch to the post-training world, Nathan, if you will. So starting with SFT, which stands for supervised fine-tuning. Yeah, I think one of the things, especially for a model like Ulmo, where we're scrappy and putting everything together over time, is that one of the biggest changes is that when reasoning models become popular, the quote-unquote in vogue evaluation suite of the industry shifts to add a whole bunch more new things in.

56:04So one of the things that happens at every stage is, even if a lot of the data has overlap, is that you mix it in a different way. So I think Kyle Lowe and Mei Chen, that's another researcher and an intern, did this whole mixing procedure that we use across all these stages just to upweight the math code and reasoning stuff to make sure that what happens later in post-training is much more tractable and that all this stuff is set up. So that's the type of thing that we have to do that's kind of baked into everything. And then post-training, I think for this model, everything we're doing is operating in the assumption that this is about a 7B model.

56:41We're very narrowly focused, and therefore we're going to do what many people have done, which is called distilling from bigger teacher reasoning models. I think distillation is described as when you take the outputs for one model and then you fine-tune on it later. I think there have been a lot of broader discussions on this in the community. And then this supervised fine-tuning stage or SFT or instruction tuning is all about just getting the best traces from reasoning models out there or the community. And then just teaching your recently trained base model to behave really, really closely to what is going on there.

57:14So in our case, we took a mix of existing data sets like OpenThoughts 3 and modified it, which is from Bespoke AI Labs, a startup. And then we also generated a whole bunch of new data. So we ended up using a mix of teachers from like DeepSeq R1, DeepSeq R10528, which was their updated version. And then Quen's reasoning model, TWQ. It's like these tend to be pretty strong teachers. And then why is that? So you have a pre-trained model, but you for supervised fine tuning, you're still using a different model. Why is that in simple terms? Essentially because our small model is not going to be able to output as strong of text.

57:54So there's a kind of a fork in process where I'm talking about a small model. And if we had a bigger model, what we would do is do a lot of reinforcement learning to start. And the model then would take time to learn these interesting behaviors and have strong performance. But with a smaller model, the ceiling on that is fairly low. It just doesn't have the capacity to learn from these harder math problems. So what the common practice is, is you take the absolute best reasoning models you can get that are openly available with a good license, where you can just generate new data yourself and train on it and release it to the community, which is something we've been seeing a lot of this year.

58:29And then therefore the models that are closest to the frontier in performance with good license all happen to be Chinese models throughout the year for this case. And I think in our case, even if GPT-OSS had existed, I don't think we would have used it for synthetic data in this because that model is really designed for tool use, which is something that we did a bit of in this project, but not in the sense that that model is, which is like this many hop agentic reasoning with search and stuff. So the deep seeks and quen's of the world are just powerhouses at generating math and code answers and other things and being generally robust.

59:07So that's what we do is we have, I think, about 2.5 million reasoning traces, mostly on math and code and STEM, but also on chat and other general capabilities. And the model really absorbs a lot at this point. I think if you were to told us last year when working on ULMO 2, looking at it, that if we had a similarly sized ULMO model that gets like 95 on math and 70 on Amy on these crazy math evals, it would have been surprising. But this is just what you can get when you can extract data directly from these really powerful models and distill it down. And I think realistically, a lot of companies are going to want to do this because you can do this for your domain.

59:45We threw a blanket on, we want all of these evals from instruction following and make sure that you can actually talk to the model and not have it just become totally broken. But you can do this at any specific task you want if DeepSeek has coverage on it and it's very efficient. And then most of the process after that is, that is the foundation. If you're training a small reasoning model, you need to do this. And then the other things after are, how do you extract more performance? and they quickly become more technical or done because we want to do that, but maybe not efficient in our time. So 90 % of our time this summer is having great people battle reinforcement learning infrastructure.

1:00:26Because when Luka is hitting it, when you generate a lot of tokens, the time or compute increase and memory increase is quadratic. Therefore, you pretty much encounter every possible bug in your framework or every possible corner case that will make your job go to a halt. but like most of the performances through this SFT and the preference tuning that comes before it. But the RL is like, we need to do this in order to build the infrastructure for many of the future almost that we want to build later this year where they get bigger and they can do more interesting reasoning with tools and so on.

1:00:57So it's kind of like a nuanced point of the model. It's like, yeah, we'll show you that we got a couple of points out of doing RL at the end. But really the RL tooling is something that's so crucial to doing the next models that come from here. And thank you for that. And just to, again, in an effort to make this interesting to just a broad group of people who are curious to understand how AI works. So SFT is not RL yet, right? That's a supervised fine tuning. So that means that you basically show the model some like a golden, like a gold copy of what good looks like. And you train it based on that label data.

1:01:35Is that Yeah, so it's the same loss function as pre-training, which is you're predicting the next token. In this case, what it looks like is a question could be like, I don't know, like an Amy style, like a really hard math question. It would be like, list all the prime numbers within some constraint of X and K. And it's like this one sentence that is really hard. And then the model generates 30 ,000 tokens of let me think about this and do this. And to test this, I'll have to use this theorem and hypothesis, which is like the 30. We were talking about token intuitions for a bit, but 30 ,000 tokens to solve a math problem is pretty mind bending.

1:02:14So if I were to sit there and read this, it would be hours of me just trying to read this one math solution. So these models are very unintelligible in many ways. I think the reasoning model sometimes will go into a bout of guess and check for hundreds of attempts before realizing that they can no longer guess and check. I mean, this is our reasoning model. I think the frontier models could have probably done this and fixed this issue. But there's just really, really, really odd things in these tokens. But even with that, doing this next token prediction is an incredible foundation of performance that many people use.

1:02:53So it's not matching any sort of human reasoning or things that people might want it to be doing, but it is teaching the models kind of their own language of breaking down problems step by step in order to solve a goal. For this specific stage of SFT, do you want to talk about how you went about creating the data set for it? So precisely this representation of what good looks like for the model? Yeah, Luca, do you want to jump in? Do you have things to? The other thing I was going to mention is that sometimes in the big announcement of the frontier labs you don't see is what Nathan was describing around like having to do SFD to then do RL.

1:03:38It's very common. We're in a common situation where this is uncharted technology. Right. So you have nothing. You have to find ways to fix some components of your pipeline before you can build the rest of your pipeline and then go back fixing the first part. So for us is, OK, we want to do reinforcement learning on this larger model. OK, we need our reinforcement learning code to actually be super fast, super reliable and useful. If we need to iterate on that part, we want to iterate with smaller model first because we can iterate faster. They take less compute to work so we can do more things in parallel.

1:04:16OK, smaller models, they cannot do RL first. You got to create the data first. You have to go to the SFT. and then you know we're lucky enough that you know okay there are other great models that are open source that we can use to create this data so the alternative would be i don't know to spend it's not even the money to spend a lot of time instructing humans to create the same volume of data slow things down so it's a lot of this of like i don't you're building the tracks as the train is going down at incredible speeds and you have to figure out ways to like fix some parts of your pipeline so you can work on the rest.

1:04:53All right, so let's talk about the next stage in the pipeline, stage five, DPO and preference tuning. What does that do? Yeah, so this is one of the things that is thought of as like, hey, let's try this. We're not sure if it'll work kind of later in the process when you spend a lot of time on other things, and it works very well. I think DPO or direct preference optimization is not exactly new. I think it's a way of optimizing for preferences. It's related to this whole RLHF thing that we mentioned. Technically speaking, in one sentence, it's an analytically derived loss function that is essentially applying stochastic gradient descent to the RLHF objective.

1:05:36So it becomes much easier to implement than other things. And we used this in the past with OMO2, with 2Loot3, 2Loot2, other OMOs. and the question was like can we apply this out of the box on top of a reasoning model and we knew that it works in many different situations because we weren't sure what would happen with these long reasoning traces being included in the loss function and so on so then essentially like there's a student Scott that has been working on this what he calls the delta learning hypothesis which is like a intuition for understanding DPO as being more about the contrast between your chosen and rejected examples.

1:06:15So the core of preference learning is that you have kind of pairs or some grouping of completions to the same prompts. You have one question with multiple completions. And his intuition in work is showing that this contrast is more important than the exact magnitude of goodness of the answer. So what he did is he spent a lot of time in trying to come up with a good pairing of reasoning models, which is they're open source, or like they're open weights and they have a permissive license and they include the reasoning traces. because we kind of need this and you need them to be sufficiently well spread about.

1:06:46So we spent a bunch of time generating this data and doing some normal kind of like, let's fiddle with the learning rate in small things. And it's kind of just like, yes, this works. After we did it, we saw that Hugging Face did something similar with small LM. So they trained a fully open 3B model where they pre-trained it as well. And the funny thing is that we converged on using the same QEN32B and QN06B. So like the problem is that these small QN models and these small public reasoning models are actually so strong that getting a sufficient delta to another model to apply this preference learning technique was kind of hard.

1:07:25So like our past techniques, we kind of had groups of models we sampled from. But as these open models are getting better, these samples become too homogenous for the learning signal to exist. So it's a kind of cool experiment. It's a cool experiment because it validates this hypothesis of the changing tides where if you think about years ago with Alpaca and stuff, those models were so broken that having this group had enough variance and contrast in it where we could do a different type of preference learning. Where now it's really, you have to look really closely at the completions and make sure that there's a learning signal for the models.

1:08:00And we did this and it kind of gave us a boost across the board. I think it's like sometimes things look very easy when you've done careful data work and kind of set up to understanding your optimizers. I think Luca described pre-training as very scientific and post-training is like the wild west. There's many analogies. So it's like, we have me that made this SFT data set where I was like, we had a bunch of cloud credits and they were running out and were behind and I just generated as many completions as possible. So it was like a few billion completions from DeepSeq over the weekend. You're like, oh, we'll mix it and filter it later and I applied filtering and the answer was like, oh, we just include almost all of it.

1:08:41We would have liked to do more if we had resources for longer, but sometimes there's low-hanging fruit and doing the obvious thing yields a lot of results. It's like this SFT and DPO stage in a lot of sense are that. And then this RL stage is extremely hard technical grinding week in and week out to make the tools even run at all. and the disparity in post-training kind of tracks to me. You just have all these checkpoints flying around and it seems like chaos and then something that's extremely obvious gives you a massive gain. The DPO gains are the difference when being about this is not an apples-to-apples comparison but it would be the difference between being QN 2.5 level to almost QN 3 level.

1:09:25The thing that you apply to get there is sometimes really obvious And I think the frontier labs are much further down this path where they take these low hanging fruits so fast. But as a smaller team that's trying to map to what the changing priorities of the field is, sometimes it's just turning the crane on this really straightforward thing. I don't remember what I said, but Dario from Anthropics said very plainly, look, what works here is with 50 to 100 lines of code. He was saying it in the context of like espionage and him being scared about like some trade secret from from Anthropic being exported out.

1:10:06But the solutions at the end, everyone the industry favors are actually very simple. The problem is that there is a very large space of equally simple solution. And all the work goes in like, OK, how do you test these? How do you test them as fast as possible? how do you convince we have to be extremely skeptical in all the in any good results so how do you convince it that this results uh look good they're not just because oh sometimes there's a bug somewhere that caused it um some some something to be too good to be true um so yeah a lot of it is it's less about like the final solution what matters it's about like the speed at which you iterate and like how robust your tools is so like immediately after you see a good results and I know that this is a good result.

1:10:51And to complete the journey since we started talking about RL, so RL VR, reinforcement learning with verifiable rewards, let's spend a little bit of time on that sixth stage. In particular, Nathan, I understand that's your baby or you're one of the fathers of the baby. Do you want to walk us maybe a little bit through the history quickly? I think that I'm the person that got to bring it publicly to the world. It's well known that people across the industry have been doing this for years, and then the technique started to get far more impactful. It's broadly taking existing reinforcement learning algorithms or downstream evolution of proximal policy optimization, PPO, which is an evolution of reinforce.

1:11:39And then DeepSeek had their group relative policy optimization. I always try to say group robust. I think it's group relative. and like all these algorithms are really quite similar and it's you're you're training the models with whether or not they got the answers right or in the case of code whether or not the tests execute and don't fail i think one of the famous examples is that doing too much of this kind of or racing to get the low-hanging fruit from this rl approach is what makes all these code models do all these try except things to avoid errors because they accept all the errors I think that is just because the gains that you get in the model being useful is so much higher than the annoyance and the fact that it also does these stupid things and will fix the stupid things eventually.

1:12:24In the case of this ULMA model, it's not anything crazy. We cast a wide net on RL math problems. We do some data comparisons to see which data we think is the best for teaching these models. We do mixing with code and precise instruction following. And this mixing is effectively when you tune the big set that you have to what you've known from many experiments and to the specific model checkpoint that you're starting on. So if you have a really strong model and you show it really easy math problems, there's no learning signal. And if you have a really weak model and you show it really hard problems, it gets them all wrong.

1:13:02There's no learning signal. So the learning signal is all from the gradient of like you sometimes get it right and you sometimes get it wrong. Do you want to give the plain English definition of RLVR versus RLHF? RLVR and verifiable rewards is in the name. I think essentially the reward that you get from the quote unquote environment, which is like the completion or the greater, is whether or not you got the problem right. RLHF, the reward is essentially a reward model, which is rating the quality of the response based on a proxy to what humans would like. So it's described as being a much, like the RLVR reward is much easier to understand because these reward models tend to have a lot of problems and you can over optimize them much more easily because the reward models will pick up on features that are maybe emojis or something like this that you don't actually care about.

1:13:53Where RLVR is much better matched to performance characteristics rather than style. And what you tweeted, I think, or said somewhere that RL and long context reasoning distills is very hard. I don't know if that's RL in general, specifically this type of RL. What makes it super hard and very much the frontier of AI right now? So there's many ways that your tooling can fail. I think where most of these processes are set up right now is that you have a set of generation GPUs, which look like something like VLM, and you have a set of training GPUs, which is some distributed learning framework, which is where you actually have this RL update and loss function.

1:14:31And therefore you need to have some sort of system that orchestrates the two and passes information back and forth. And this kind of information passing back and forth is really annoying. It's a systems problem because you have like distributed error handling and things like this. So like a common case is when you have the most basic approach is that you'll have like one generation, this one math problem, The model is thinking and thinking and thinking and thinking. So you have all these GPUs working on one problem. So effectively, your whole system is somewhat idle waiting for the answer. And there's many other small things like this, which is this long context generation just uses so much memory that you then need to introduce different types of parallelism and stuff to do the generation effectively.

1:15:12And there's just a lot of subtle numerical issues. So I think it's just kind of stress testing a lot of the post-turning infrastructure that we have had by turning up a lot of different things that could go wrong to do the maximum. I think the things that the open community struggles with is that VLM and HuggingFace use different kernels to do the actual internal computation of the model. So these kernels are the things that make things like VLM really fast. But this then results in subtle numerical differences between the completions that you're generating from the model and then the log probs that the thing that's doing the loss function actually generates.

1:15:50And if you look at the math of these RL algorithms, it's assumed that those are from the same distribution. So therefore, this is a big root cause of a lot of numerical problems. And then if you look at what we're doing, a lot of labs have done throughout the years. They do different fixes to change these numerics. I like Thinking Machines had a famous blog of one of their first blog posts on like deterministic VLLM to make it exactly deterministic. And like that is really useful and people think it is key to their like Tinker API and doing other sorts of RL things where you just have complete control over sources of non-determinism.

1:16:24And that can just be like numerical lack of robustness in RL. And if we talk about trade secrets or whatever in RL, there's also these discussions on if the open labs have worse algorithms than the closed labs. And in reality, it seems like most people are using something like an evolved version of GRPO, which is a bit simpler than PPO. Some labs might be using a learned value function. It's not that important about the details. But what happens is that each lab finds the set of tweaks that they need to get really stable RL performance. And in the RL literature, historically, there's a pretty low bar on the amount of changes that are needed to call it a quote unquote new algorithm.

1:17:05But it's realistically like an implementation detail. So it's like a lot of everyone finds their stable configuration for operating and it's really dependent on the tool. So, yes, you could say that they have a different algorithm, but it's also not really something that you could easily exfiltrate from a lab because it's dependent on many layers of the stack. and maybe what chips they're operating on and all sorts of things. So it's just one of these things where training these models is complex and the kind of quick quips could never reflect that. The stack for post-training is also so new. Like software-wise, I feel like the big strides are happening in 2024 around like, oh, you do like both, you train the model, but also you run the model at the same time and that they have to happen at a certain cadence versus of pre-training the seeds of distributed pre-training.

1:18:03Like you had that in TensorFlow, which Google released in 2017, right? So there is a much more mature stock versus what you need on the post-training side. All right, so maybe as a last part to this conversation, it's been really fascinating and illuminating everything that you guys have described because in particular, it sort of highlights the complexity of the systems, like the multiple stages. And I love what you mentioned, Nathan, a few minutes ago when you said the pre-training is scientific and post-training. My words, not yours, but my interpretation of your words was like it's a lot of tinkering and putting things together in a way that you hope is going to work.

1:18:47And truly diving into how those models work on the one hand, But on the other hand, you know, each time you're like open a newspaper online or go on Twitter, like everybody's talking about AGI and how we're almost there and how it's going to change everything. There's a little bit of a cognitive dissonance between like the reality of trying to make those models work with all the unbelievable progress that we've seen, of course. But like, you know, that on the one hand and the discourse on the other hand. And Nathan, you've had a much more, I would say, tempered view of AI progress compared to some AI researchers.

1:19:30You had a great blog post very recently that you called Thought on the Curve. I'm curious what your latest thinking is. And, Luca, obviously, feel free to jump in any time. But in terms of what you described in that is saying in a prior one as complexity and complexity tax, which, again, like in view of the pipeline you just described, like it's one start to understand the level of sheer complexity that's involved in all of this. Yeah, so ultimately, I definitely describe myself as lightly AGI-pilled, and I think you have to be to appreciate the magnitude and gravity of the situation that we're in.

1:20:11But also, I think that I'm very far from believing in any sort of singularity being possible due to these things like complexity. On one hand, we talked about all these things which are low-hanging fruit to improve the models. And I don't doubt that it, I mean, like Schulte commented this on the pod and other places where at these labs, they still see low-hanging fruit in improving the models. In many ways, I don't think that their approach feels that different. They've just refined it relative to what we're doing. But at the same time, as things get complex with tools and adding more layers to the stack and you have to build a product to scaffold it, if the requirement to get the best out of Cloud is to use Cloud Code, which is some magic product and prompting relative to GitHub Copilot, this is one thing that you're going to need to get right in order to get AGI along with all these tool uses and stuff.

1:21:00So it's like as any system gets more complex, the pace of change is slower. I think any tech company has seen this. And then realistically, there's going to be physical constraints on the amount of infrastructure that we can build. So this belief is simultaneously giving us these new data centers. And I personally think hopefully a new power generation. But there's a cap. And for having a, in order to, like all these things are plateauing and then you 10x the compute and you get a big jump. Like you can't do this forever. So realistically, there's going to be some physical constraint that kicks in at some point.

1:21:33But balancing that with complex systems and the low hanging fruit results in like, I think these researchers are going to grind out improvements for multiple years, but never in a way that results in this kind of accelerating well that we get drawn into. So it's kind of, it's like, I don't know, in some ways it feels like I'm having my cake and eating it too, but it seems like the likely outcome if you look at other types of technology. Yeah, and that conversation was in a particular reaction to AI 2027, which is a really interesting conversation. I'll let you summarize the premise, but the short version is that AI sort of builds itself and therefore accelerates.

1:22:16Yeah, and I think they have these milestones like AI automates research engineering and then AI automates AI research and development. where it's like each of these are incredible jumps in performance. And I think what's more likely is this messy evolution. They deserve credit on their marketing and getting this impact for sure. But even them are now like, oh, maybe we should have called it AI 2028 or AI 2029. So I think like that is the reflection of like there are these real constraints, but the progress is also going to be incredible. It normally is the growth and capability of these models.

1:22:50But like having working on one, I think it's very unlikely that we will see like a discontinuity at any point. It has nothing to do with whether like we'll get to a definition of AGI or super intelligence that people are happy with. We will get there. It seems unlikely that the moment like it's going to be a looking back kind of kind of exercise of like, oh, these were the important milestone and this is what really worked. It's building this model is so much a collection of like refinements till to unlock the next stage that it's going to be this smooth trajectory. Whether it hits at some point where we don't have more capacity to keep improving on whether it forever accelerates.

1:23:38It's people are going to be disappointed if they want to see a moment where like one day they log into Twitter and AGI is there. Like it's messy and it's fun working on it because it's messy and it gives a lot of satisfaction. So to play it back, you're both saying yes to AGI, but no to discontinuity slash singularity. And one, is it fair? And two, if that's what you're saying, then for AGI using the current paradigm, basically what we just described in the last hour of pre-training plus RL gets us there? I think the AGI word is actually pretty not useful. I think that how I describe it is that big tech has all collectively realized that these language models plus scaffolding is going to unlock absolutely incredible value.

1:24:27And I have very high probability, barring extreme geopolitical situations, that big tech executes on this vision across the two to five years to build 95 to 98 percent of the way there of what you can do with our physical power constraints and what an LLM's ability is. And I think that that will be extreme, like the transformation from that by 2030 is going to be so powerful across society. There's a bunch of long tail, like there's going to be mass societal readjustment to what the internet and media and information means and like within five years. And that's mostly why I do this. And I think debating whether or not it's AGI is kind of secondary to the fact that this is coming.

1:25:12And we want people to study and understand what is happening. And to that last point, what does that mean? Study and prepare. What would you recommend people do? Although if people have made it all this way to this point of the podcast, they've already done a bunch of the work. So I think there's a lot of it. I mean, there's a lot of interest in AI across outside of the CS majors and on of the world. where it's informing policymakers. I think it still takes a long time for information to diffuse. And there's often not that many people that are engaging in this that are doing it just purely for this kind of, you can call it alignment and concern.

1:25:54There's just a lot of general noise. And I mean, I worry about concentration of power or all sorts of many things. And it's just trying to upskill people into understanding AI. so they can be engaged, like engage listeners and think about how it affects their domain. The other part, maybe I'm more positive to know, is like if the scaffolding is what really moves a lot of like from, you know, broad capability model to like something that actually has meaningful impact, that scaffolding is not just like, oh, only the labs of people with trained models can do it. Like the number of people can contribute to that, both in terms of people with tech expertise, and people with non-technic expertise, it's much larger.

1:26:37If the scaffolding is what really moves capabilities, what gets us to this incredible technology being realized, then the number of people can contribute to it is not just those who work at Frontier Labs. There's a tremendous amount of technical work to do, but also non-technical. As soon as you start integrating this technology in the life of real people, As soon as you start working on, you know, high stake medical application or other high stake domains, then a large amount of population can contribute in making this technology better and make it work for everyone. Just a base model. I feel like the number of people who can really can help making this technology really work for everyone.

1:27:21It's large. Everyone in society feels like it can contribute. All right. Well, that feels like a wonderful place to live it. Thank you so much, both, not just for this conversation, but for all the work that you're doing in open source frontier AI, which feels sorely needed and extremely important. So really appreciate the time and all the thoughts. Thank you so much. Hi, it's Matt Turk again. Thanks for listening to this episode of the Mad Podcast. If you enjoyed it, we'd be very grateful if you would consider subscribing if you haven't already or leaving a positive review or comment on whichever platform you're watching this or listening to this episode from.

1:28:05This really helps us build a podcast and get great guests. Thanks and see you at the next episode.

From the publisher

In this special release episode, Matt sits down with Nathan Lambert and Luca Soldaini from Ai2 (the Allen Institute for AI) to break down one of the biggest open-source AI drops of the year: OLMo 3. At a moment when most labs are offering “open weights” and calling it a day, AI2 is doing the opposite — publishing the models, the data, the recipes, and every intermediate checkpoint that shows how the system was built. It’s an unusually transparent look into the inner machinery of a modern frontier-class model.


Nathan and Luca walk us through the full pipeline — from pre-training and mid-training to long-context extension, SFT, preference tuning, and RLVR. They also explain what a thinking model actually is, why reasoning models have exploded in 2025, and how distillation from DeepSeek and Qwen reasoning models works in practice. If you’ve been trying to truly understand the “RL + reasoning” era of LLMs, this is the clearest explanation you’ll hear.


We widen the lens to the global picture: why Meta’s retreat from open source created a “vacuum of influence,” how Chinese labs like Qwen, DeepSeek, Kimi, and Moonshot surged into that gap, and why so many U.S. companies are quietly building on Chinese open models today. Nathan and Luca offer a grounded, insider view of whether America can mount an effective open-source response — and what that response needs to look like.


Finally, we talk about where AI is actually heading. Not the hype, not the doom — but the messy engineering reality behind modern model training, the complexity tax that slows progress, and why the transformation between now and 2030 may be dramatic without ever delivering a single “AGI moment.” If you care about the future of open models and the global AI landscape, this is an essential conversation.



Allen Institute for AI (AI2)

Website - https://allenai.org

X/Twitter - https://x.com/allen_ai


Nathan Lambert

Blog - https://www.interconnects.ai

LinkedIn - https://www.linkedin.com/in/natolambert/

X/Twitter - https://x.com/natolambert


Luca Soldaini

Blog - https://soldaini.net

LinkedIn - https://www.linkedin.com/in/soldni/

X/Twitter - https://x.com/soldni


FIRSTMARK

Website - https://firstmark.com

X/Twitter - https://twitter.com/FirstMarkCap


Matt Turck (Managing Director)

Blog - https://mattturck.com

LinkedIn - https://www.linkedin.com/in/turck/

X/Twitter - https://twitter.com/mattturck


(00:00) – Cold Open

(00:39) – Welcome & today’s big announcement

(01:18) – Introducing the Olmo 3 model family

(02:07) – What “base models” really are (and why they matter)

(05:51) – Dolma 3: the data behind Olmo 3

(08:06) – Performance vs Qwen, Gemma, DeepSeek

(10:28) – What true open source means (and why it’s rare)

(12:51) – Intermediate checkpoints, transparency, and why AI2 publishes everything

(16:37) – Why Qwen is everywhere (including U.S. startups)

(18:31) – Why Chinese labs go open source (and why U.S. labs don’t)

(20:28) – Inside ATOM: the U.S. response to China’s model surge

(22:13) – The rise of “thinking models” and inference-time scaling

(35:58) – The full Olmo pipeline, explained simply

(46:52) – Pre-training: data, scale, and avoiding catastrophic spikes

(50:27) – Mid-training (tail patching) and avoiding test leakage

(52:06) – Why long-context training matters

(55:28) – SFT: building the foundation for reasoning

(1:04:53) – Preference tuning & why DPO still works

(1:10:51) – The hard part: RLVR, long reasoning chains, and infrastructure pain

(1:13:59) – Why RL is so technically brutal

(1:18:17) – Complexity tax vs AGI hype

(1:21:58) – How everyone can contribute to the future of AI

(1:27:26) – Closing thoughts

More from The MAD Podcast with Matt Turck

All 44 episodes
Open Source AI Strikes Back — Inside Ai2’s OLMo 3 ‘Thinking"The MAD Podcast with Matt Turck · 1 h 28 min
Listen in VO