In short
Latent Space: The AI Engineer Podcast
Episode Title
[LIVE] Anthropic Distillation & How Models Cheat (SWE-Bench Dead)
Episode Description In this episode, hosts Nathan Lambert and Sebastian Raschka engage in a deep discussion about AI models, especially focusing on Anthropic's recent blog post regarding distillation attacks and the implications of benchmarks in AI. The conversation explores various themes within the AI ecosystem, including the practices of distillation, the challenges of benchmarking, and the geopolitical concerns surrounding AI research.
---
Key Takeaways
- Introduction of Guests and Context
- Swyx joins the team from SAIL Media, which signifies an expansion of content and collaborative effort within the AI media community.
- The podcast emphasizes the dynamic nature of AI discussions, especially regarding the latest research and developments.
- Distillation in AI
- Definition: Distillation involves training a smaller model on the outputs of a larger model, improving efficiency and performance.
- Context: The discussion highlights how Anthropic identified “distillation attacks,” where models from Chinese labs use APIs to train their own systems, raising concerns about intellectual property and the geopolitical landscape of AI.
- Benchmarking Challenges
- SWE-Bench Overview: Originally created by Princeton, SWE-Bench serves as a coding benchmark for evaluating AI models. It involves fixing bugs in code snippets derived from open-source projects.
- Issues Identified:
- SWE-Bench was criticized for selection bias and variability in the quality of tasks.
- OpenAI's SWE-Bench Verified attempted to refine these benchmarks for better accuracy through human vetting.
- The complexity of accurately benchmarking AI models due to the nuances in task definitions and potential for models to exploit patterns in the data.
- Implications of Distillation Attacks
- Detection: The discussion reflects on how companies might detect when their models are being used improperly for distillation purposes, especially regarding the scale of data requests.
- Geopolitical Concerns: The conversation touches on the implications of AI research and development in the context of US-China relations, highlighting the competitive edge each side aims to maintain.
- OpenAI’s Benchmarking Efforts
- OpenAI’s significant investment in creating and refining benchmarks like SWE-Bench Verified is noted, as well as the potential for future iterations like SWE-Bench Pro.
- The conversation underscores the importance of continuously improving benchmarks to ensure they reflect the true capabilities of AI models.
- Community Engagement
- The hosts express appreciation for the human connection in discussions about AI amidst an influx of AI-generated communications, emphasizing the importance of collaborative discourse.
---
Discussion Highlights
- Distillation Techniques: The hosts explain that distillation is an established concept in machine learning, not limited to LLMs (Large Language Models), and is critical in developing smaller, efficient models.
- Ethical Considerations: The ethical implications of using outputs from one AI model to train another are explored, including the responsibility of AI companies regarding terms of service for API usage.
- Benchmark Variability: The variability and unreliability of benchmarks, such as the SWE-Bench, are criticized, with a focus on the need for rigorous testing and validation.
- Future of Benchmarking: The hosts speculate on the future of AI evaluations, particularly how emerging benchmarks might address the shortcomings of current models and better reflect their capabilities.
---
Conclusion The episode wraps up with a commitment to further discussions on the critical topics of AI distillation and benchmarking. The hosts encourage listeners to stay engaged with ongoing developments in the AI field, reiterating the importance of collaboration among engineers, researchers, and enthusiasts.
For the full transcript and additional resources, visit [Latent Space](https://latent.space).
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOWelcoming Sean Swix to the Team
0:46 to 1:42
Hosts welcome Sean Swix, discussing his background and contribution to the Sail Coalition.
“Thanks for, I just want to say thanks for joining us.”
Distillation and Its Relevance
1:43 to 2:58
Discussion on the concept of distillation in AI and its implications for model training.
“I put how models cheat in the top so we can talk about benchmarks.”
Understanding Distillation in AI
2:59 to 4:26
Explains the process of distillation and its historical context in machine learning.
“distillation before we maybe dive into the details.”
Ethical Concerns Around Distillation
4:27 to 5:30
Hosts discuss the ethical implications of using API outputs for training models.
“And they are trained on the outputs of their own larger models.”
Detecting Distillation Attacks
5:31 to 6:40
Exploration of how companies might detect distillation attacks on their models.
“Essentially, in terms of service, it's something that you essentially are using a service.”
Evaluation vs. Distillation
6:41 to 8:08
Discussion on the fine line between model evaluation and distillation practices.
“means really like literally just letting chat GPT, cloud generate synthetic data, and then you collect that synthetic data and train your own model with supervised learning, supervised fine tuning on it.”
Privacy Concerns in AI Model Usage
8:09 to 9:40
Addresses concerns about privacy and data usage in AI model training.
“So maybe one way would be, okay, this is a familiar question.”
Anthropic's Actions Against Companies
9:41 to 11:10
Discussion on Anthropic's history of blocking companies from using their models.
“And I find it kind of interesting that a company would look at that, even at the scale, and call you out.”
Comparison of Models and APIs
11:11 to 12:53
Comparing different APIs and their implications for AI research and development.
“But, I mean, there are a lot of legit use cases.”
Using Open Router for API Access
12:54 to 14:00
Exploring Open Router as a viable solution for accessing APIs while distilling models.
“I think, Nathan, in your write-up, you had a little bit of a comment about...”
Show all 24 chapters
Exploring OpenRouter and API Providers
14:00 to 16:40
Learn about OpenRouter and its features for handling APIs.
“like marketing wise like to make it you know stick or to you know um actually you mentioned also like the different APIs and everything.”
The Impact of Release Timing on Model Detection
16:40 to 22:30
Understand how the timing of model releases affects detection and performance.
“Which, I mean, we will confirm this later on if we do end up doing the call with them.”
Distillation Techniques and Their Challenges
22:30 to 28:00
Discover the complexities of model distillation and data usage.
“So it's almost like easier to start distilling from a medium model.”
Introduction to SweetBench
28:00 to 29:00
Learn about the purpose and significance of SweetBench in LLM comparison.
“to see I think next time because I think this time it might also be a bit biased towards releasing a Codex because they almost released it simultaneously with their app that they want to promote at the moment.”
Understanding the SweetBench Framework
29:00 to 31:00
Explore the details of how SweetBench was created and what it measures.
“So the umbrella topic here is how do we compare which LLM is currently the best LLM?”
The Evolution of SweetBench
31:00 to 33:00
Discover how SweetBench evolved and the challenges it faced during its development.
“I'm actually working on it with Cognition to launch a new benchmark here.”
Benchmarking Challenges in AI
33:00 to 34:20
Examine the inherent difficulties in creating reliable benchmarks for AI models.
“Okay, and then if you want a bit more historical context, this is like a step up from human eval, which is more on completions.”
Issues with Current Benchmarks
34:20 to 37:40
Discuss the problems found in benchmarks, including examples and implications.
“And then maybe like a couple more sort of verification passes or whatever, right?”
The Complexity of LLMs and Their Training
37:40 to 41:40
Delve into the complexities of how LLMs are trained and their memorization capabilities.
“While they were looking at this, they had a second thing that they looked at the chain of thought.”
Understanding Model Performance and Benchmarks
42:00 to 43:21
Explore the challenges of model evaluation and the need for new benchmarks.
“And I still think it's like super understudied.”
The Evolution of Sweepbench and Its Challenges
43:21 to 45:36
Discuss the updates to Sweepbench and how they address previous issues.
“I don't think what I'm saying is that M.2.5 should get less score on SWE bench, but I think other models should get more score.”
Evaluating Models with Private Data
45:36 to 47:19
Learn about the complexities of evaluating models using private datasets.
“So it's not a guaranteed perfect set is what I'm saying.”
The Future of AI Evaluations Beyond Coding
47:19 to 48:35
Examine upcoming trends in AI evaluations across various domains.
“I haven't tried it, so I don't really know.”
The Importance of Human Connection in AI Discussions
48:35 to 50:21
Reflect on the value of human interactions in the AI community.
“I think the other day, though, Anthropic acquired another company that does like a UI type of stuff on the computer.”
Transcript
Automatic transcript. May contain errors.0:00Nathan Lambert:Okay, we're live. We have one person. People will start trickling in. Thanks for coming to Sail Live number six. This is a very exciting one. I think we have a, I mean, the topics are always fun with these. It's whatever is the topic of the day on our little rat-raising minds trying to keep up with AI. But we're welcoming the latest writer that is joining the Sail Coalition. So I think this just means more content for Sail. I think I've been a fan of SWIX and a friend for a while at this point. So I'm very happy to have his content join this. And I think you've been doing great stuff recently and continuing to evolve this.
0:37Nathan Lambert:So welcome to the team. This is my friends and colleagues in the AI media space. And it's just great to be able to support people and keep that network closer.
0:48Sebastian Raschka:Yeah. Thanks for, I just want to say thanks for joining us. It's really a pleasure to have you on here, Sean, or Swix. So, yeah, awesome. I just coincidentally listened to your podcast about the Superbenchmark. So, yeah, awesome to, you know, small world, awesome to have you here.
1:09Nathan Lambert:Yeah, thanks for having me. And, yeah, just glad to be on and chat. I've never, ever done one of these Substack live things. So I'm curious how it works. Because I always think about Substack because it's like a newsletter platform. But they want to go multimedia. I think the live thing before we get to technical content is actually good because it gives it a different edge. It's just like a little bit sharper when you know you're live. I think we've all done a lot of podcasts, even podcasts that are unedited and put this later. But I think the live thing is a different element that can be tapped into nicely.
1:41Nathan Lambert:So I don't know. Why don't we just dive into it? We're going to start with distillation. I put how models cheat in the top so we can talk about benchmarks. I think Anthropic posted this pretty spicy blog post this week. I think it was essentially detailing how they found distributed distillation, quote unquote, attacks on their services from prominent Chinese labs. And I'm very unsurprised with Anthropic calling it an attack. I think that that fits with a lot of their branding. Okay, nice. Screen share. This is what we mean. Sean Swix is such a pro. and it's like and the screen share fee bro was only dropped a few days ago but essentially it's anthropic is detailing how they found distributed accounts across multiple chinese labs building state-of-the-art lms and described what they were doing and why anthropic is concerned about this in their worldview of like ai geopolitics and i think it's very interesting because i'm of the opinion that the chinese labs like obviously should do this.
2:46Nathan Lambert:They're in a massive GPU shortage, and using APIs is way easier than generating synthetic data on their own. And I think there's a lot of...
2:55Sebastian Raschka:Maybe we should, just for the general audience, just define distillation before we maybe dive into the details. Yeah, so distillation, that's like a broader concept. It's not like a new concept that came up with LLMs. It's like an older concept in machine learning in general. And distillation essentially is... So the idea is that you're taking a larger model and train it on the outputs. Sorry, you have a larger model, let it generate outputs and train a smaller model on these outputs of the larger model. And the idea is that you can train the smaller model more efficiently using that larger model.
3:33Sebastian Raschka:And originally, I think you just brought up the paper here. Originally, what you would do is you would train on the logits. So old school machine learning people might remember from deep neural networks like the logits, the outputs of the last layer that you usually work with them to compute the loss function across entropy term. And you would train on this signal. And nowadays in the context of LLMs, it's a bit more loose. So it does not have to be these logits that you train on. It could be just the output data, synthetic data, like Nathan just said. So, for example, it's actually a very common practice.
4:10Sebastian Raschka:For example, in DeepSeq R1 in the paper, or other people that do other companies, they would train the flagship model, the largest model, the R1 model with 671 billion parameters. And then they would create smaller variants like, I forgot the numbers, but 1, 3 billion in a smaller range, like these small models you can run locally. And they are trained on the outputs of their own larger models. I think now the thing is also I mean this is very common practice everyone does that when they are producing the smaller model variants now I think the question or the point Nathan brought up is what happens if you are a company and you generate this synthetic data from another company's LLM and then train your own model on it so sorry that was just like a little interruption but yeah distillation in short is training a smaller model on the outputs of a larger model basically
5:04Nathan Lambert:Yeah, and I think this is even possible at the frontier. So people distill from something like Cloud Opus to build Cloud Sonnet. They're generally doing very similar things internally. They have access to different tools and richer tools. And then the other context is that all of these large labs for years have had terms of service where they say that you effectively cannot use the outputs from these APIs to train something like a competitive AI model. It is vague terms. in terms of service or not a contract. Essentially, in terms of service, it's something that you essentially are using a service.
5:39Nathan Lambert:And then if the provider finds you violated, they can cut off your access. That's just kind of a basic thing. So these have not been enforced within the US much at all. I think there was one case, maybe like ByteDance a year or two ago, that OpenAI cut off their API. But this was discussed so much right after Chat2PT when people were building the first open models on like Alpaca and things. So it's like, is OpenAI going to come after us for doing these research models? And it totally died down. People were worried about this for like over a year. It was kind of an insufferable discussion. So nothing really happened.
6:12And then this is like the first prominent reemergence of the discussion kind of to make,
6:18Nathan Lambert:I think it's because people are far more worried about AI competitiveness. And I'm curious what you guys think.
6:25Sebastian Raschka:Yeah. Can we talk a second about even how they would detect, because you said in the beginning something about distillation attack, and you didn't say that specifically, but you kind of like implicitly put quotation marks on attack. So how would you even detect that? So I think, I mean, distillation in that context means really like literally just letting chat GPT, cloud generate synthetic data, and then you collect that synthetic data and train your own model with supervised learning, supervised fine tuning on it. but then how would you even detect that this is a distillation attack versus just an evaluation because right now I'm actually running I mean I'm distilling myself for chapter 8 of my book but I'm doing it with open weight models so no worry anthropic please don't worry about it I distill from API models
7:14Nathan Lambert:for my job
7:15Sebastian Raschka:I use open router right now and just distill from the deep stick version 3.2 model which I think these folks are okay with that but what I wanted to say is so when I'm evaluating models I use basically almost the same script so when you're evaluating in a model you have a question and you let the model generate the answer so you generate the response to your benchmark question and in my benchmarks I have data sets from math 500 500 examples I have a bigger math data set of 12 ,000 examples so you're basically just running an API in a loop to let it generate these questions and sorry the answers but then how would a company know okay this person is just evaluating versus this person is now saving that data and then later training their own model like you see what i'm saying like it's the same i think it's a scale thing
8:10Nathan Lambert:so like sure when you're evaluating at least the basic evals or you're going to do it once and not do it there's some amount where you are i mean like they say stuff here but there's also more of it where you just are not like you're not gonna i think most of it is quantity and then they're gonna look at patterns across similar accounts is what they're yeah exactly they're
8:31Sebastian Raschka:gonna see like really repetitive stuff yes so i think the interesting point this leads to is like um i mean you can do evaluation at a large scale if you are a big company you want to know whether your lm performs very well you have a large suite of um benchmarks you are gonna run and But then you said like maybe looking for patterns. So maybe one way would be, okay, this is a familiar question. It comes up in the benchmarks. So this person is maybe not stealing our answers. It's just using it for benchmark purposes. But then it means kind of like that they are looking at what you're generating there.
9:09Sebastian Raschka:Of course, nothing is private when you are using LLMs on the internet. The data is somewhere intermediately stored. But then it kind of like almost implies that they are checking what you use the LLM for, what you generate, which is kind of like a sensitive topic, almost like privacy-wise, right? So that's kind of like an interesting point because, I mean, of course, you mentioned the terms of service that you are not allowed to distill, but you're not distilling. So the point I'm trying to make is you're not distilling life when you are on the platform. You are doing it somewhere later. You're just letting the LLM generate answers.
9:42Sebastian Raschka:And I find it kind of interesting that a company would look at that, even at the scale, and call you out. Like, hey, you are generating too many answers here. That's not cool or something. That's kind of a weird thing.
9:55Nathan Lambert:Yeah, I wanted to respond a couple. This is like a few sentences back. But actually, Anthopic has blocked U.S. companies first before the Chinese companies. he has blocked both OpenAI and XAI from using the models and I think maybe explicitly accused XAI of distilling stuff I don't know but definitely not like in a full blog post like this so this one is like definitely the most high profile case and yeah I do think like it is actually pretty hard to distinguish from like hey I'm just running my internal benchmark man and of course it's going to be very high volume of like all of the same stuff because especially some benchmarks you have to run three to five times.
10:42Nathan Lambert:It's the exact same questions. I do think obviously if you get to the tens of thousands, hundreds of thousands, then you're not just running benchmarks. You are distilling this thing.
10:57Sebastian Raschka:There's a good point in the chat. How would the distribution of questions look like if you're distilling? And I think, really to your point, at a certain point when you have a certain magnitude of answers generated, it might look suspicious. But, I mean, there are a lot of legit use cases. If a company uses your, let's say, OpenAI Cloud API as their own chatbot and they have a lot of customers, it's naturally a lot of answers that are generated. And so they would probably look at distributions like maybe you would expect a very broad distribution when you are distilling because you want to cover pretty much everything.
11:35Sebastian Raschka:And when you are running benchmarks, it's maybe more specific. You're running a math benchmark. It's just math. Or if you have a customer chatbot, it's more like customer answers. But yeah, I think they would maybe analyze your distribution. I feel like this is kind of a weird thing to do. I don't know if you're a company and you're looking into your customer privacy, like data generated, you know, like, of course, it's, well, you have to expect that it's not private, but still kind of like a weird thing that they, that they do that essentially.
12:08Nathan Lambert:Okay. What else do you have to talk about? I think, is it, is it interesting? Okay. I did, I did. Okay. One thing, this is a little bit of sub stack, like authors back and forth. One thing I did was I threw it into Nano Banana, which is kind of like a decent visual, right? Throw it into Nano Banana 2. It's a Nano Banana 2 live pod. It just released five minutes ago. This is actually Nano Banana 2. So because I'm in the early access program, they cut you over to new Nano Banana, and I couldn't access the old one. So I was trying to do a diff, and I couldn't do it. Classic. That is classic early tester program shit.
12:50Nathan Lambert:Look at the pain we have to deal with here. Is it interesting that DeepSeek uses so much less than Minimax? I think, Nathan, in your write-up, you had a little bit of a comment about... This is a political blog post, in a way. Maybe not political, but they're trying to make a point that is more about making a point than the details. The DeepSeek thing is definitely way smaller scale. so okay most of the labs will experiment with all the apis they can get access to like data is just so important and you're going to have a pipeline where you can sub in any api and then run an ablation to see if it gives you performance the api is kind of free like just do it the millions of exchanges is a bit more of a bet you can you can measure that a bit longer and it takes a lot longer to get the millions of exchanges is like tens of billions or hundred billion tokens and it takes a lot longer to actually get that out of the api especially when they have to spread it across a ton of accounts these accounts are already limited and have other problems like that takes longer but this tiny one is so fast so i don't like i that was generally my point that it made it clear that anthropics trying to use the deep seek name as the only chinese ai name that people in the u.s know
14:04Sebastian Raschka:like marketing wise like to make it you know stick or to you know um actually you mentioned also like the different APIs and everything. I'm not like sponsored by demo. I have no affiliation. I've never talked to anyone from that company. But Open Router, for example, is a good example where I've been using it a lot for the open weight models because for the bigger ones, they're too big to run them locally. And what's nice is they do also offer. So it's basically just routing you through other companies' APIs and they select automatically at that point what is the cheapest one at that point. I sometimes get some failures.
14:38Sebastian Raschka:I think when it switches, sometimes it crashes, but like in my script maybe it's like something i have to fix there but um so even then if you're distilling you can do that from multiple providers but yeah of course if you are wanting something from chat chpd or claude it's always going to go through the official one and then it gets i guess suspicious but you could also technically um distill a bit from open through open rotor through i mean through their account your direct account you can make multiple accounts and it's kind of interesting that they track all that and then like yeah different topic now that you called out that they call out DeepSeq which is quite interesting
15:17Nathan Lambert:For what it's worth OpenRouter seems to not be using DeepSeq in most of these.
15:24Sebastian Raschka:These are free models DeepSeq's not good. I see I'm using the paid API I should also say it's also nice they show you how much it costs and the tokens per second for different providers so if you go to the search in the top you can go to the different deep seek ones i just like it because i do a lot of model comparisons and then this one is an older model so maybe it only has one provider but if you go to i think deep seek r1 or something or even the normal 3.2 there should be multiple providers that if you scroll down here you can see there There are different providers and different tokens per second, different costs.
16:04Sebastian Raschka:So it's kind of like, I just like that website because it's just quick to use the API and they have an OpenAI-like API. So it's almost like it's not sponsored or something. I just find it generally useful. So, but yeah, just a side note.
16:21Nathan Lambert:Do you want to go back to the comparison? Did you have a high-level point to make there? Oh, okay. just a couple one I think I think the timing post Moonshot releasing their stuff post Minimax releasing their stuff but pre DC v4 I think that was strategic I think that may also have factored into why Minimax was more detected like had a higher number so like you know like when you collect data is actually very important right and so they interrupted or they found Minimax during the training of Minimax 2.5, right? Which, I mean, we will confirm this later on if we do end up doing the call with them.
17:04Nathan Lambert:And so obviously, like, the number is going to be very high because they're, like, actively looking for it. And then they banned the Minimax accounts and Minimax changed their things. Actually, I don't think that's exactly what happened. Sorry. Let me correct myself. While Minimax was distilling, they released Opus 4.6. and they said that they redirected nearly half the traffic. So I'm like, this is like, okay, very, very clearly this is them, right? It's the same exact traffic switched to a new model the moment a new model releases. Okay, cool. DeepSeek maybe wasn't doing that because they hadn't been working on their stuff actively, I don't know.
17:41Nathan Lambert:It could be a different thing. Or DeepSeek is just way more efficient. I get all I need for 150k. You guys are so inefficient. It would be so interesting if we knew the time frame of this? Like, are all these API requests within the last four weeks, are they within the last six months? Like, that's such a different nature of what is going on. Exactly, right? That's what I'm saying. Like, DeepSeq was training 3.1, 3.2, like, you know, a year ago. Yeah, or like, I don't know, DeepSeq OCR. I guess they said what it is, but it's not that.
18:10Sebastian Raschka:But like, also scale-wise, I do think, yeah, Minimax is three times smaller. It's just like a faster model. They don't use MLA, and they don't use the DeepSeq sparse attention, but it is, I mean, I think it's just group query attention, but it is still a pretty snappy model. So it's, I think, just attractive maybe to use it. And the other one, top of my head, I don't know, maybe they had like some free tier or something like that, where I think when the models come out, they sometimes offer free usage. And that was a more recent model than, I think DeepSeek, the last one was from December, the V3.2.
18:46Nathan Lambert:Yeah, yeah. So, you know, maybe this is an irrelevant point because they were training before and people would have the same amount of traffic or they're just way more efficient, right? It does spring to mind... The efficiency thing is not it. I can guarantee it. Like, that is not... Yeah, it's a small chance that they... It's like there's a chance that they got the right research idea early and, like, found the right data to use, but it's not that they're, like, going to be 3x more efficient. Okay, so it's like, you know, it's a timing thing or they just actually don't use it that much. I mean, like, you play this out, like, I was like, okay, why don't they share?
19:22Nathan Lambert:They're all buddies. And it does come to a point where, okay, let's have all of China just distribute it to every citizen. I can talk about this a little bit. There's a lot of research, not a lot of research, but there's a few research projects trying to understand how do you use distillation data. I think SFT is the cleanest example where you're doing this autoregressive loss on Q &A pairs. But the strongest model is not necessarily the best teacher. And most of us in this area think it's due to like some, you have to match the probabilities of the tokens to the base model. So like what's happening is that Quen dense models are the best teachers for a lot of open weight models.
20:06Nathan Lambert:And I think that's because a lot of open weight models are either Quen or have been like Quen-like for a while. So like Olmo learned really well from Quen. And obviously like other Quen models did. but like scaling these pipelines up to use say glm 4.7 or a bigger deep seek model or a more recent big coin moe like all of these it's a lot harder to just generate the data from the same prompts with like the right sampling settings and then do sft on them and actually make the numbers go up interestingly gpt oss is a pretty good teacher but there's like a huge gap there where it's like just because you have this data does not mean it's actually going to make your model better.
20:42Nathan Lambert:So you have to do the research to be like, oh, we learned that we get signal out of Claude. We need to get 100 billion tokens ASAP because it's going to just immediately make our model better. Like that's not a common place to be in modeling because this like weird teacher-student dynamic going on. So I can see that being different across labs.
21:05Sebastian Raschka:I think also it has something to do. I noticed also if you are distilling the smaller model from the same model family it performs better and i think it's to your point that um if you have a very very strong model um it might be also too different um or like if the style is too different and then it's too much of a leap for your model to adapt like it's too too different from the q and a answers during the pre-training or and so you make a bigger leap and another thing i wanted to say about you mentioned olmo and i it's been a while since i read the paper but you might know way better than I do, but I think you did also train on the logits.
Read the full transcript
21:43Sebastian Raschka:We didn't do technical
21:44Nathan Lambert:distillation.
21:46Sebastian Raschka:We just took the tokens. Oh, I see. Okay, then it was probably a different paper. I think Google does that for their Gemma models. Because here, there's also then the distinction, because you mentioned Quen and other models, you can only do that for open-weight models, because if you do that for Claude or OpenAI, that would not work with the logits, because they don't provide them. They only provide them for some tokens like a hundred or a thousand top tokens and so it is in a sense if you want to do the real in quotation mark distillation it is kind of like even easier to do that from open open weight models because you can control it but then also like you said well we need a hundred billion tokens asap that is not a easy thing to do because even like i mean it's like 40 tokens per second or something for these large end models when you generate answers and getting that million billions tokens, it takes time, right?
22:34Sebastian Raschka:So it's almost like easier to start distilling from a medium model. So it's like the question, more data versus more high quality data, right? So it's also like a sweet spot to, like an experiment itself, an ablation study, right?
22:52Nathan Lambert:Yeah. I like that Nathan had to call it technical distillation because it is no longer the default, even though it was the first. Also, another fun fact, I did my Jeff D. in the interview recently and I tried to get out of him. He sort of dodged it a little bit. You know, remember there were actually three sizes of Gemini models. There was Nano, Pro and Ultra. And I was like, where is Ultra? They keep it in the basement and they distill from it, right? Interesting, yeah. Maybe also
23:24Sebastian Raschka:is it like to safeguard yourself so no one can make a copy or the price also but probably both.
23:32Nathan Lambert:Yeah, I mean, I think like this is how like, I always think of like, the model you deploy is never the model you train. Because you train the dense and then you deploy the MOE, right? Like you basically always do it. Like at every lab. Say more. Like you think they're really distilling from dense models? I mean, like, I think that like, that is like the full, like, when you like just unlimited resources, don't care about inference, just care about maxing intelligence. Why not? Yeah. I'm not 100 % sure. I think that MOEs just give you a flop. Like, I don't know if that's actually how I think of gains of MOE when you have really good MOE architecture.
24:11Nathan Lambert:But I do think that they have bigger models that they distill from. And they train internal models different than external because the external models have been getting a lot smaller, which is a weird thing. We don't have a good way to measure it. Maybe Dylan will backwards figure it out and inference max, whatever the heck. They'll deal with this new model size.
24:32Sebastian Raschka:But I'm always suspicious with these things. Also, it's really like a capacity thing, too. How many people use the model at the same time? Hardware, how much is allocated? And it's always, it's like, yeah, maybe a rule of thumb. But yeah, it's really tricky. I think it's really hard to say anything from these numbers.
24:49Nathan Lambert:I do think that they might start restricting models that will only be in products and not be in API. I think the whole API business is brutally competitive. and I don't have a good sense for what the defensibility of it is. I think it makes sense for something like Google and Azure and our existing cloud businesses to have APIs, and that's kind of a more natural transition. But the Anthropic and OpenAI API, the transition from their products, which are their big differentiation, whether it's ChatGPT and CloudCode and Codex, you don't get people to go use the API from that. And I think you get a lot of people that are already spending on clouds that then go to use the APIs, which is why like Lambda and Nebius are going to have these API products.
25:32Nathan Lambert:But like, isn't it, if Claude's really worried about distillation, like they should put the model release in Claude code ASAP and then just not bother with the API. I don't know when that'll happen, but it could.
25:44Sebastian Raschka:I do think though it's a big customer base, the API customer base, any type of product that is built on, I mean, with LLMs, like customer chatbot types of things, But also more generally, I do think the problem like with, I don't know exactly how the plans work in Cloud, but you would reach a token max where you can only get so much with your subscription. You can, I think, buy more tokens, but I think it's just easier with the API at a certain scale. And also like the whole OpenClaw customer base, right? Because they don't allow the plan anymore in the OpenClaw context. So you have to use the API.
26:20Sebastian Raschka:And I do think, given how many tokens OpenClaw generates, it's actually not a bad business if you don't lose money on these tokens. If you sell it at a not-subsidized price, I do think the API is actually not a bad business model. Yeah.
26:38Nathan Lambert:Do you want to take a side? Do you want to try a tie-break? I'm obviously being provocative. I don't really know, but I can see it. Anthropic gives Apple vibes to me. I mean like Anthropic has a higher chance of doing this yes, OpenAI just because I have talked to the people so much I just don't super believe that they will have locked models to products only out of I guess idealism and principles rather than economic incentive economic incentive would agree with you that they should have private models to products. And recently, they've done this. The last three GPT-5s all had codex variants that were two to four weeks ahead released only inside of codex rather than as an API.
27:29Nathan Lambert:So they're starting to get there. But just constitutionally, I don't think the people that run these things believe in locking things behind APIs because they have such a huge market anyway. So they kind of don't care. and then they also, if you're genuinely sort of zealot, if you're not trying to maximize the value of your company and genuinely just trying to spread AGI everywhere, then you release the API because you just don't know what people are going to build with it.
27:57Sebastian Raschka:One more thing though with the Codex thing, we will have to see I think next time because I think this time it might also be a bit biased towards releasing a Codex because they almost released it simultaneously with their app that they want to promote at the moment. So it could have been like more like they did that so that anyone checks out the app.
28:18Nathan Lambert:Always a two to four week exclusive window.
28:24Nathan Lambert:That's still right. If you want to promote Codex that's pretty effective. We have a bunch of questions in the chat. Do we want to cover benchmarks and then this thing? Go right ahead. What do you want to do? It's your sub stack. I don't know. Oh, man. It's a collective. You should just dive into what you're interested in. I mean, Sebastian was interested in the SweetBench stuff. So this past week, SweetBench Verified died. Or officially...
28:55Sebastian Raschka:What do you mean by this? Yeah, let's define SweetBench first, maybe.
29:00Nathan Lambert:Okay. I happen to have the post on this. So let me just...
29:06Sebastian Raschka:So the umbrella topic here is how do we compare which LLM is currently the best LLM? One of the ways would be SWE bench, basically. But then I will maybe let you explain because you had this brilliant podcast or article.
29:24Nathan Lambert:Okay, where do you want me to start? You want me to just define SWE bench, I guess? I guess, yeah.
29:29Sebastian Raschka:Yeah, so maybe going from, so basically that it is a coding benchmark and then SweetBench is like a popular way to compare capabilities of LLMs and then there is SweetBench verified. But maybe, yeah, we should talk a little bit more about SweetBench first.
29:47Nathan Lambert:So SweetBench was a paper out of Princeton from Ophiopress's group and they do a lot of good code benchmarking work. And it happened to be that they just kind of drew thousands of example sort of open source issues and PRs that closed those issues from open source. There's a bit of selection bias here because they only focus on popular open source and only a small number of popular open source, but a large number of issues on those open source. And then they just kind of dredged up some passing tests and then some failing tests that you need to make pass in order to pass the score. when it launched, it was kind of obscure.
30:30Nathan Lambert:Devin actually was the first one to choose it as a benchmark to report. And then it went from like, I think at launch it was like 13 % and now everyone's at 80%, something like that. SweetBench, because it was done on like a student budget, was very kind of, let's call it sloppy or whatever. TerminalBench is like this now too. Like they're just aggregated. It's like hard to do a benchmark that is well calibrated across topics. Yeah, it is hard. So, you know, for the small group that is watching, I'm actually working on it with Cognition to launch a new benchmark here. But yeah, so OpenAI was like, okay, guys, SuiteBench is taking off.
31:12Nathan Lambert:We are going to adopt this, but we refuse to abide by the full SuiteBench. We're just going to actually go and curate 500 subsets of the original SuiteBench. And they actually hired humans to go and vet through. It's somewhere inside of this blog post, but basically they hired three humans for every task to just vet whether the task was high quality or not because there's a lot of slop in there. And they were like, okay, this is the 500 that we're going to endorse.
31:41Sebastian Raschka:So it's like a curated subset of Sweebench where 500, let's say, challenging problems that are supposedly well-defined.
31:49Nathan Lambert:Yeah, yeah. And what's really funny is that at launch, so this was launched in 2024 at launch OpenAI could not run all of its own 500 so for a while like there was like a few releases from OpenAI that reported on a subset of the subset because they couldn't run it on their eval infrastructure so like their numbers were higher because their denominator was lower which is very funny. Anyway
32:15Sebastian Raschka:in that context we should say what Speedbench kind of looks like I think it's like basically like a code that has bugs in it, and usually the task for the LLM is to fix the bug in the code, right? It's right here.
32:28Nathan Lambert:The whole thing's open, which becomes a problem in the future. But right now, you can see the whole thing. You can see the reports from the issue ID and the problem statements, and then you have also the test that you're supposed to pass and fail. So it's all here on Hugging Face. And you can see that it's at 500. Anyway, I think we don't want to get too lost in the details.
32:53Sebastian Raschka:You just wanted to say, or define the context that this is a coding benchmark, essentially 500 examples that are available on the internet.
33:02Nathan Lambert:Okay, and then if you want a bit more historical context, this is like a step up from human eval, which is more on completions. This was, in my mind, the first proper agentic benchmark, I guess, apart from TauBench, where they give you the problem and the end result, and they don't really specify how you're supposed to get there. Whereas I think a lot of previous benchmarks like MMLUs of the world and human evals, which is in the coding domain also released by OpenAI, was very much, here's the problem statement and then give me the right answer immediately after without that much extra files or anything that you're supposed to run.
33:40Nathan Lambert:So the other ones are more autocomplete. This one is more agentic. It's all a spectrum, obviously, because you can use agents to solve autocomplete. But that's not what HumanML was testing. Anyway, I just wanted to make sure people understand that OpenAI actually invested a lot of money and effort into making Sweepbench verified from Sweepbench. That's a question. How much money do you think this costs? Oh my god, don't do this. Millions? I would guess order of a couple. It could be a few million. Yeah, I'd say a couple million. So basically you do like, okay, what's the first filter pass? and then like, okay, it's 500 times three because they had three people per thing.
34:22Nathan Lambert:And then maybe like a couple more sort of verification passes or whatever, right? So like, yeah. So then they were like, oh, so this year, they're like, oh, well, not only is it saturated because like progress, everyone just takes turns to increment by 0.1 every time they release a new model. It's like, it's bullshit. It's obviously bullshit. Like the inherent noise in just running these models varies by like 0.5 to like 1. Every time you run it, like you just choose the highest.
34:50Sebastian Raschka:A little nitpick. I don't think it can be 0.1 % because like what you said before, because it's 500 examples. I think the smallest increment is 0.2%. Okay. They might average. Because like little detail. Yeah, sorry.
35:05Nathan Lambert:I think so as we progress to the next era of benchmarking, the N, so the N here is 500, right? The N doesn't directly correlate to the percentage points because you get sub points as well. Ah, yeah.
35:20Sebastian Raschka:Good point. Good point.
35:21Nathan Lambert:So like Terminal Bench, even though it has 90 something tasks, like you can get subdivisions less than 1%. Anyway, so not only do they have this, they actually audited their own. They were like, okay, how come everyone is saturating at 80 %? What's up with the remaining 20 %? How come everyone's failing at it? And they were like, oh, actually we looked, We paid even more people, six people per task now, with an extra team if any sort of positive identification is found. And we were like, 59 % of them cannot even be solved at all because the original benchmark was still slop. Stuff got through that was not solvable.
35:59Nathan Lambert:And I actually tried to illustrate this in my post. So here, this is an impossible test, right? Okay, so here's an example. This is the sort of value add I did on top of the original post. Here's an example of a Sweebench verified task that passed the first round of human verification, right? So here's the task, and we want to implement Python type hints or something. We want to see expected behavior, I want to see a string in the output, right? So if you were given this, you would never pass this because the test said, I am looking for something called get annotation. And if you don't, give me this magic string get annotation, you will fail this task.
36:38Sebastian Raschka:why so yeah it's way too specific it's yeah yeah it's like uh yeah right so so so this is just a
36:46Nathan Lambert:bad uh task that somehow escapes uh validation so the only way you could kind of solve it is if you're memorizing the answer yeah exactly exactly which is actually a nice like i think every benchmark should include stuff like this where if like a honeypot if you solve this you're like, oh shit, it's a canary, right? It's like, oh, I mean, you're definitely cheating.
37:08Sebastian Raschka:Like a sanity check, yeah. That's actually a really nice point, yeah.
37:12Nathan Lambert:Yeah, so I just think to me, it's a beautiful point of how hard it is to make evals that there was these multiple rounds. There was original Sweebench, which the Princeton kids did do initial first pass. Then there's a second pass of OpenAI doing Sweebench Verified. And then every single person that ran for the next 1.5 years did not call this out.
37:38Nathan Lambert:Until OpenAI was like, hey, let's look at the data. So I think it was really interesting. While they were looking at this, they had a second thing that they looked at the chain of thought. And inside the chain of thought, they found GPT-5's own chain of thought to start including information from the future. Because it was trained on, because the problems are open source and because it was trained on information from GitHub, it would use advanced knowledge of future versions of the Django version that they were using to solve the problem. Definitely seen stuff like this in the real world where the models will hallucinate the new version of the API even if your script isn't on it.
38:20Nathan Lambert:I think a lot of the hugging face stuff is the worst with this where the models just are totally goobly glopped. They've seen all the versions and the API has changed too much over time where they throw something out there. Yeah. So the most biggest, yeah, I mean, I think like, you know, there's a lot of this, right? Like, sort of ethical behavior. Like, okay, so you can blame things like, oh, you should not have released this, the full data set in public, because obviously people can train on a full data set. But like, it's not like the researchers are trying to do this. Like, because these things are also open source, like, any data set that touches GitHub, any training corpus that touches GitHub is going to just eventually absorb this.
39:02Sebastian Raschka:And it's not even this website or the repository directly. It's a clone of this repository or someone else who has that developed their own open source library and has that in the unit tests or something where it's not even intentional or malicious or anything. It's like by accident, you already absorbed that.
39:18Nathan Lambert:Yeah, or a new feature that releases this edit-only feature. It gets written up in a blog post or a conference talk or something. And then it just makes it in, right? Like, it's really funny. okay so to me like OpenAI could have stopped there and said okay we're done they did one more extra thing which is kind of funny they also then ran Flash Gemini and Opus and this one it was more it was like even more egregious okay they just gave the task ID and just said repeat this Rebench task to me and so from task ID they can just vomit out the whole statement and the solution these are crazy the stuff that's in these models when you zoom in deep is really really incredible because like these are models that are like really really well done but there's just so much complexity and all the pieces of the pudding that get put in the recipe yes there's so many weird i also still find it fascinating that
40:14Sebastian Raschka:um i mean of course uh it's like kind of by design when you're training that you memorize things because that's literally like next token prediction but given that how big a model is on how much data it sees, and usually it sees only the data once, that it still has enough capacity to memorize. It's kind of like, so usually I would think, okay, I would have to train multiple epochs to be able to memorize, but no, it is enough maybe to include it once or twice in the training corpus, and it can do a perfect rendition or a perfect recap of what it is in there, which is kind of fascinating. Even if people don't want that, it's you know it's crazy yeah labs got good at this there's essentially like a duplication level
41:01Nathan Lambert:that you need at each stage of training and it's not easy to measure so like if you do too much at pre-training your model forgets basic facts and at post-training it's probably closer to these abilities and i think that that is a thing that is not well reflected in like you can see it in a vows of your knowledge tank yeah this is like an art that they have probably gotten good at
41:23Sebastian Raschka:Yeah, like continued pre-training does also require some revisiting of old data. Otherwise, like you said, you have the forgetting. But it's still fascinating to me that with such a small fraction usually, because you usually use one or two, five percent for like continued pre-training, that it's enough to have the model memorize almost everything, which is fascinating. I don't know. It's just like still after all these years, fascinating.
41:50Nathan Lambert:I think there's, so one of the pet topics that I pursue like two, three times a year on my stuff is the information theory of LLMs. And I still think it's like super understudied. Like how come you can memorize from one pass? Yeah, exactly. And then also like people forget like superposition, which is like Anthropics original McInturp work. also basically stuffs information inside the smaller bits that then get forgotten. But how does superposition actually work? I don't think I've seen a convincing study on that. Okay, anyway. I'm done on my sweet benchmark. I don't know if you have thoughts or questions or whatever.
42:38Nathan Lambert:But I do think this is an example of the models unintentionally cheated and benchmarks are hard to make. and we need new ones. And, you know, if this happens to SweetBench Verified, which I think is the most scrutinized benchmark in the world.
42:53Sebastian Raschka:In my recent post, I had like a bar plot where I showed the SweetBench Verified numbers for most models. And like you said, they were all 80 something percent, like literally 80 point between one and nine, let's say, where there's almost zero variation, even like something like Minimax 2.5, which I do think is worse than, GPT 5.2. Like, no offense, it's a smaller model, it's a cheaper model. For my usage, it's a little bit worse, but on this particular benchmark, it's the same. I don't think what I'm saying is that M.2.5 should get less score on SWE bench, but I think other models should get more score.
43:37Sebastian Raschka:But like you said, the problems are just impossible to solve. but one point I think we didn't bring up is we said that Sweepbench Verified has issues so what do we do about it? I think there is like a Sweepbench Pro now which is kind of like I would say like Verified try to fix the regular Sweepbench and Pro tries to fix Verified but I haven't looked into this, is it like another subset or is it a completely different set of problems?
44:06Nathan Lambert:Yeah, it's a new set so the you know, SweetBench draws from like a 2022-ish, 2023-ish era of problems. So all you do, there's a few things you do, right? One, you do private-public splits, right? That's super obvious. Two, you update the dates which you draw from. And then three, you diversify the repos and the languages, right? So these are all just like very, very super basic fixes. And then obviously, you try to fix the testing. Super basic fixes to the original SweetBench. which it doesn't take a genius to figure out but they did the hard work
44:43Sebastian Raschka:but it is in a sense also what Verified meant to do so it's not let's say people looked at this again but it's no guarantee that it doesn't also still have issues that might be discovered later on right I mean it's
44:55Nathan Lambert:no so SweetBanch Verified was an intentional subset right these guys were like no no no we need to have a superset not even a superset
45:05Sebastian Raschka:we need a different yeah but But what I was trying to say is when Sweebench Verified was developed, there were three people per task making sure the task is well-defined and everything. But then two years later, it turns out, no, no, this was not the case for everything. And what I'm trying to say is it could be that Sweebench Pro is better, but it might still have issues. That might not be obvious right now, but maybe in one to two years when we revisit this and you see some of the failure cases, maybe we'll discover, okay, this has still some issues. So it's not a guaranteed perfect set is what I'm saying.
45:38Sebastian Raschka:I don't know, but it's just like a suspicion here.
45:41Nathan Lambert:Totally, totally. You know, I do think Scale.ai has a professional interest in making sure this is as good as possible. Yeah, no, no.
45:48Sebastian Raschka:But what I was trying to say is 3Bench Verified also had a professional interest to make sure. Oh, very different incentive.
45:55Nathan Lambert:I guess they're all very different. This one has a limited budget. This one has basically unlimited budget because it's like literally existential to Scale.ai that they have good data. Sure. but I also think it's really nice that this team the evals team at OpenAI keeps endorsing Opus it's kind of funny so yeah they deprecate Cbench Verified and then they were like we're going to report Cbench Pro now and GPT-5 is like you know number one
46:24Sebastian Raschka:maybe if do you know if I would want to evaluate on the private data set how would I do that do I provide the API to is there like an API call I have to do against scale AI?
46:36Nathan Lambert:I don't know. I have an API key and agree to not. You have to agree because if you don't have an agreement, then you can just have to keep the data. You have to do special hoops to make sure that you don't steal the private eval.
46:50Sebastian Raschka:Yeah, my question was basically, do they even let you download the data or is it more like you send the answer to them and they do the evaluation on their backend so that you don't even get to download the data. Otherwise, like you said, you could... Yeah, yeah. So basically, you only provide the answers. So you have your LLM generate an answer and you submit the answers and then they have some process to evaluate on their thing so that their private data never leaves their servers, my guess, because otherwise someone might upload it or something like, you know.
47:23Nathan Lambert:Yeah, I don't know. I don't have... I haven't tried it, so I don't really know. I'm sure you can sort of reach out to them to figure it out. Yeah. Anyway. I think this is good. Unless people have more comments that they want to have. I think this, but this is only coding, right? But there's like every other domain needs this. The domain is the one thing right now. I think the frontier evals are even more expensive, which is like the Apex eval from Merkur. Like evals are going to cost, this is millions. they're going to cost tens of millions and hundreds of millions of dollars at the frontier which is just a very strange dynamic whereas like there's so much about the ecosystem is forking between frontier models and then like research and other things and trying to follow that dynamic and explain it to people it's going to take a lot of work but yeah coding is i do
48:16Sebastian Raschka:think really interesting because that's what most people use lms for these days but also it is easier to evaluate i think once you leave like coding math it becomes a bit obscure how do you measure the quality of the answer, you get back to, let's say, preferences, I guess, which is more like a subjective thing where coding is more objective. So it is not a bad thing to do. I think the other day, though, Anthropic acquired another company that does like a UI type of stuff on the computer. And I think that is something where...
48:49Nathan Lambert:Normal talent flows in AI.
48:53Sebastian Raschka:I mean, I'm not trying to say this is like a big thing to talk about. What I'm trying to say is like, this is another interesting point for evaluating LLMs on those tasks, because I think a lot of people want that to like, like they want an LLM to control the computer and do various things, but they are harder to measure. So that will be maybe two years. We will have something more like benchmarks that can, it's harder to specify. it's kind of like what is it called in programming there's unit testing and then the system testing basically like the UI testing and stuff like that and so I think that is the next maybe going to be the next thing basically end testing
49:33Nathan Lambert:yeah GDPVal is usually the thing that gets brought up here so if I'll just leave it there I think we've sort of beaten the benchmarks benchmarking but definitely GDPVal is sort of here I'll put it that way okay yeah
49:49Sebastian Raschka:yeah but like the big topics essentially the distillation and the benchmarks this week. And welcome
49:56Nathan Lambert:to our coalition of whatever that means formally. It just means I get to hang out with you guys which is what I want. I mean it's ultimately a media vehicle and I think brands and vehicles from media are actually very influential today. I think you see many companies investing in that and I think it's important to have people that you respect and are aligned with.
50:21Sebastian Raschka:able to amplify each other yeah it's also nice to um talk to humans because i noticed the last couple of weeks if you go to social media well i think it's 50 uh lobsters like open claw clients nowadays i get a lot of um emails but also notifications or responses that are they look ai generated so it's it's nice to also you know have this human connection and actually talk uh to like an expert about things.
50:51Nathan Lambert:Cool. There are a bunch of like comments. I don't know if you want to do like quick hits or are you like kind of... I have to go to a meeting. That's why I'm trying to wrap this up. I see, I see, I see. Okay, well, you know, time is yours. What do you want to do? Okay, thanks everybody.
51:09Sebastian Raschka:We'll see you next week. Yeah, thanks everyone for joining. It was like a nice spontaneous, I guess, you know, discussion. I mean, it always feels nice to talk about things. And too bad we didn't get to discuss these chat questions because also on my screen, I probably need glasses at some point. My screen is pretty far away. I can just barely read them. But yeah, thanks everyone for commenting. It is just nice to see also so many people excited about these topics.
51:41Nathan Lambert:Yeah. Hopefully see you later. Have a good rest of the day. Bye. Thank you.
From the publisher
Swyx joined SAIL! Thank you SAIL Media, Prof. Tom Yeh, 8Lee, Hamid Bagheri, c9n, and many others for tuning into SAIL Live #6 with Nathan Lambert and Sebastian Raschka, PhD. Sharing here for the LS paid subscribers.
We covered:
This is a public episode. If you'd like to discuss this with other subscribers or get access to bonus episodes, visit www.latent.space/subscribe




