In short
Large language model distillation—training a smaller “student” model to mimic a larger “teacher,” including (1) efficiency/specialization and (2) alleged “model stealing” via querying an API to collect outputs for a facsimile.
Guests
Katie and Phoebe (hosts of Linear Digressions). Katie leads the technical explanation; Phoebe responds with questions and examples.
Key claims
Distillation uses teacher outputs as training data for the student. It can target cheaper/faster or domain-specific models (e.g., bird-identification vs broad knowledge). “Stealing” is framed as high-volume API querying to build a training set; it’s hard to detect because requests can be distributed across many accounts/IPs/VPNs. Forensics is uncertain; alleged models sometimes answer “I’m Claude” only some of the time.
Notable examples
Microsoft/Bing vs Google analogy; alleged Chinese frontier models possibly distilled from Claude/OpenAI; querying “what model are you?”; Hinton & Dean 2015 distillation paper; cat/fox probability distribution and sampling/ensembling via distributions.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOUnderstanding Distillation in Language Models
0:26 to 1:41
Discussion about the concept of distillation in large language models.
“This is the thing I really don't know much about.”
The Purpose of Distillation
1:41 to 3:07
Exploration of the reasons for creating smaller, distilled models.
“That's a reasonably accurate metaphor for what we're doing here or what research labs might be doing with distillation.”
Examples of Distillation Applications
3:07 to 4:55
Examples illustrating why not all models need extensive knowledge.
“Well, the examples that I always come to are, do you want your coding model to have all of the knowledge about like medical disease and treatment?”
The Ethics of Model Distillation
4:55 to 7:22
Discussion on the implications of 'stealing' models and intellectual property.
“Is this kind of like, I remember an example, and this was before the AI age, but like Microsoft had just released Bing, a search engine.”
Reverse Engineering and Model Theft
7:22 to 8:10
Examining how some models may be distilled from existing models through API usage.
“So the general, yeah, the general idea here, just to lay this out fully, is you're anthropic or you're OpenAI and you're spending a billion dollars to train your next frontier model.”
Challenges in Identifying Distilled Models
8:10 to 13:20
Challenges faced in detecting when a model is distilled from another.
“And in addition to all of that work, like there's probably some tricks and technological, methodological innovations that you're putting in there.”
Technical Aspects of Distillation
13:20 to 14:00
Overview of the technical process involved in model distillation.
“Yes, you're going to have access to its outputs.”
Understanding Model Distillation
14:00 to 21:33
Learn how model distillation can improve the performance and efficiency of smaller models.
“I could imagine that as you're training the smaller model, you're using some evals to try to understand where it's getting strong and where it's not yet.”
Teaser for Next Episode
21:33 to 22:20
Get a sneak peek into the upcoming episode about reasoning models.
“So it's a way that you can take a large language model and make it either a little bit lighter and more focused maybe or computationally lighter.”
Substack Newsletter and Community Engagement
22:20 to 22:51
Discover the benefits of subscribing to the podcast's newsletter for additional insights.
“Phoebe and anyone else who is listening, if you go to Substack and you look for linear digressions, you'll find a newsletter where you can get a link to this week, the Jeff Hinton and Jeff Dean paper.”
Transcript
Automatic transcript. May contain errors.0:00Hey Katie. Hi Phoebe. It has been a minute. I'm really happy to be back. Welcome back. Thank you. I'm hoping maybe you can distill some of your vast knowledge into an episode about distillation for me.
0:15Katie Malone:Oh, so you're distill interested in large language models? Of course I am. Of course I am. You are listening to Linear Digressions. the puns are back i'm so glad they're back yeah don't blame me entirely though for anyone who's not a long-time listener in our first iteration phoebe was i brought a lot of puns a lot of puns uh a lot of all of them mediocre high energy medium quality puns all right but tipped our hand a little bit here in this episode i am very excited I mean, talk about distillation when it comes to large language models. This is the thing I really don't know much about. I just know the word distillation.
1:01Katie Malone:Well, you had a pretty good metaphor when we were doing some of our warmups right before we started recording, which was talking a little bit about students and teachers. And there's a little bit of a metaphor about how, in some ways, I'm a teacher in this setup. I'm definitely not the teacher. No, I'm definitely not the teacher. So we will call you the student. Okay. So you're the large capable model. I'm a smaller model. And we're trying to get your knowledge distilled into me so we can get the same answers out of me as we would out of you. But in a smaller model, is that roughly right? That's a little bit of the idea.
1:41Katie Malone:Yeah. That's a reasonably accurate metaphor for what we're doing here or what research labs might be doing with distillation. I mean, as a student model, I go for reasonably accurate. Well, I mean, I think let's come back to that, actually, because it's one of those things that's like, how should we think about what the student model does or is capable of relative to the teacher model? But yeah, let's talk about the actual concept here. So let's imagine that you're an anthropic, you're an open AI. You're putting all of this work into training a very large flagship model, let's say. this is the beast, this is the monster, this is the one that you're spending a billion dollars for.
2:23Katie Malone:And those tend to be, of course, very capable. They can do all kinds of nice stuff. They have a lot of skills that they're able to bring to a bunch of different domains. But then when you go to serve it up, make it available to end users, it may be that that's not the only version of that model that you want to make available. It might make a lot of sense for kind of computational intensiveness purposes, for cost purposes. Yeah, that there's a smaller, lighter weight version of that. And so this isn't the only reason that someone might distill a model, but just to get your intuition going about why might you want to have a couple of different versions from the model and how do you think about one building the other.
3:07Katie Malone:Yeah, that's a good place to start. Well, the examples that I always come to are, do you want your coding model to have all of the knowledge about like medical disease and treatment? Maybe not. Maybe it's kind of a waste to do that. Do you want your, I've gotten really into identifying birds. I'm very bad at it. My phone is very good at it. But do we want that model to know everything about how to write a Next.js application and deploy it? Like, no, that doesn't make any sense it it just it exists to identify birds maybe that's not the best example but but like we have these broad models that can do you know quote-unquote everything and then we may have models that can do more specific things or we may have models that we want to be way way way faster or cheaper um and maybe accuracy isn't the most paramount i don't That's the way I think about it.
4:05Katie Malone:Yeah, I think those are some reasonable ways to think about it. And generally, the way that you do this from a technical perspective is in this student-teacher model, this metaphor that we're developing. There's a bunch of different ways that you can do distillation. But in general, you get outputs from the larger model, and you use those to train the smaller model. So the smaller model is learning how to mimic the larger model in whichever particular way you're guiding it. So if you wanted to make it just really, really good at identifying birds, you'd have the large model generate a whole bunch of examples of bird identification outputs.
4:42Katie Malone:And then you would put your smaller model through some kind of training process where it's exposed to a lot of that material and it learns how to mimic it, basically. There's a second high profile use case for distillation or reason why you might do it. and this one I think is very interesting but also a little bit harder to pin down sometimes and uh the gist of it is distillation is also the way that you could steal a model oh interesting I almost said oh cool which is it's it's not cool but maybe it is cool I don't know if you're stealing from an evil corporation maybe it's cool for sure um I mean wait is this It's fine if you're stealing from the bad guy.
5:28Katie Malone:Yeah, yeah, of course. We all think we're the good guy, right? Right. Is this kind of like, I remember an example, and this was before the AI age, but like Microsoft had just released Bing, a search engine. and there was all the suspicion that they were doing a bunch of google searches and then just getting like basically uh making their search work just like google search by looking at the inputs and the outputs that google uh had for their search product is this a similar kind of that is a great callback it's a it is very similar i will also say that there are many people in let's say the hacker news crowd not just the hacker news crowd yeah who would take issue with my even characterizing this as stealing it's a like that's fair use um oh that's interesting there's like a whole field of like you know people thinking about the intellectual property of of models and and these sorts of things but and also not to get on a high horse and i won't actually step on the soapbox to mix my metaphors but um you also could make the argument that so much of this AI stuff is theft because the people who are originally content creators for the data that goes into training these models obviously are not getting compensated.
6:50So it's kind of interesting to talk about like model intellectual property and everything like that when in some ways a lot of this field in the way that it's applied is founded on theft too.
7:04Katie Malone:Well yeah no Oh, exactly. And that's what I think is the straw man version of the argument that, well, okay, maybe you're stealing the model, but the model itself is a compression, a theft, if you like, if you want to go all the way out there, of the collective sum total of human knowledge that can be scraped from the internet. Yeah, right. No, but the place where this comes up the most often, or where this is alleged the most often, is in the context of, in particular, that some of the Chinese models that come out of the Chinese frontier labs are, it's been alleged, it's very difficult, maybe impossible to prove, but been alleged that those are distilled versions of Claude or OpenAI or, you know, American frontier models, basically.
7:56Katie Malone:So the general, yeah, the general idea here, just to lay this out fully, is you're anthropic or you're OpenAI and you're spending a billion dollars to train your next frontier model. And it takes a year and it's, you know, it's a big lift. It's a ton of work. And in addition to all of that work, like there's probably some tricks and technological, methodological innovations that you're putting in there. It's not just about how much computational power and how much data you're pouring into it, but maybe there's even a little bit of secret sauce in there. You put all of that work in and you put out your GPT 5.6 or your Claude fable or mythos or something like that.
8:38Katie Malone:And then one of these labs that wants to, let's say, take advantage of that for their next model. They want to have a very capable model too, but they don't want to do all of that work. And so they might just start hitting your API over and over and over again, asking tons and tons and tons of questions, collecting all of that data back of basically, you know, hey, we're going to play the student and we're going to ask a bazillion questions and here's what the teacher says. And at a large enough scale, you can start to basically collect a training set that allows you to make your own facsimile of that model.
9:19Katie Malone:You're not full on reverse engineering it, but you're starting to get into that territory. Oh, interesting. Yeah. And I mean, one of the first things I think about is like, okay, if you're open AI or anthropic, why not just block the quote unquote bad actors who are hitting the API and asking all of these questions. And the answer to that is it's really hard to tell. It's pretty easy to disguise these kinds of things by spreading them across a bunch of different accounts or a bunch of different IP addresses or whatever. So yeah, it's kind of a tricky thing to defend oneself against. It is, yeah.
10:00Katie Malone:It can be very difficult to identify these. because it's not like there's one IP address that's just hitting Anthropics API 25 million times. Like they'll set up allegedly large networks with like thousands of accounts and route them through different VPNs so they'll look like they're coming from different places. And even when there's traces within the models, like the allegedly distilled models, it can be a little bit hard to tell exactly how to interpret them. I was actually just reading a thread this morning. It was a thread about a model that has just been released in the last couple of days.
10:40Katie Malone:It's coming out of one of these Chinese labs. And it was on Hacker News, so people were kind of getting into a, in the comments threads, trying to talk about whether there was speculation about whether this one was distilled from any of the American competitors or not. I don't really know. But one of the things that folks were saying in reference to, I think, some other models where there were allegations that they were distilled is that you would go into this model that was allegedly from a Chinese lab. Ask it, what model are you? And it'll say, oh, hi, I'm Claude. Oh, my God. Like some percentage of the time.
11:14Katie Malone:Yeah, yeah, yeah. But not always. Yeah, some percentage. Wow. Yeah. And then, you know, in response to this, some of the labs that create these models say, well, yeah, they're trained on all of this. There's all this LLM generated data that's out on the Internet now. We're not purposely trying to pull that in and train our models on it. But there's certainly plenty of texts out there on the Internet where you have Claude outputs that say things like, hi, I'm Claude. yeah so um that's why i said it's very yeah it's very very difficult to say forensic i'm not aware of there being like a way forensically to say with 100 certainty um there are some people that say that there's signatures that are a little more suggestive than others but yeah fingerprinting is an interesting area fingerprinting i think was easier when oh i actually am not even sure about that i know that uh for a while there were certain models that really liked using m dashes all the time in their output um but i don't even know that that's that that's like model specific or anything like that yeah fingerprinting is hard no yeah and again even if there's models that use m dashes a lot like at this point there's plenty of that there's a lot of generated stuff on the internet yeah there's plenty of m dashes that you can find that are not themselves from any distillation process or anything.
12:35Katie Malone:It's just how people generate text sometimes. Oh, man. This is such a – I mean, it's just such an interesting thing to think about how the training sets are more and more LLM-generated because the training sets are the Internet. Oh, yeah. Yeah, the circular feedback loop here is one that'll get your head spinning every time. So yeah, so distillation. Technical process is not the most complicated thing in the world. Again, the core of it is you're going to, you have, let's say, some ability to ask questions of this larger, more capable model. I was about to say you have a larger, more capable model, but maybe you don't.
13:18Katie Malone:Maybe somebody else has it. You have access to it. But you're able to ask questions of it. Yes, you're going to have access to its outputs. That's the crux of it, maybe just to put it another way, is you're using your larger, more capable model, in a sense, to generate training data for a different model that you're making. Does the smaller or less capable model have the ability to ask for what it needs more of? Like, is there a going the other way? Or is this just a one-way process where you have some process that generates the training data from the larger model and sends it to the smaller model for training?
13:56I guess the smaller model isn't online yet. It's being trained, huh?
14:00Katie Malone:Yeah. I could imagine that as you're training the smaller model, you're using some evals to try to understand where it's getting strong and where it's not yet. And you're using that to selectively sample. you're getting more data from the areas that are comparatively weaker but i would think that that's a little bit of like the outer the outer loop that the researchers are running i'm not aware right now of the models themselves as they're being trained identifying that they have weak spots interesting okay of like proactively going and filling them but i i don't specifically know and it's not crazy to me to think that in some future versions of maybe even present versions that are still getting worked on internally, that there's models that are designed to be more aware as they're getting trained of which areas they need to be stronger in.
14:54Katie Malone:To some extent, this is just what the more agentic models do is they'll identify maybe something that you're asking that is more recent or technical and then go and launch a web search for additional information. But that's to add things into the context, not to train the model itself. there's one other thing that I came across in the course of researching for this episode. And I genuinely did not know this, but it's kind of interesting. The first mention that I could find of the word distillation when it comes to machine learning models, it was actually a paper by Jeff Hinton and Jeff Dean and one other author from 2015.
15:34Katie Malone:So this is over 10 years ago. This is pre pre-transformers, pre-LLMs. They were talking about distillation of knowledge in large neural networks. We'll include a link to the paper in the Substack newsletter. If you're a subscriber to our Substack, check it out. We'll toss the link in there. I should be kind of a subscriber. You should be. Come on. um so uh but anyway they had a very interesting observation and and it this is something that i think is a little bit less emphasized in distillation as we currently think about it the the description that we just gave the description that we just gave is you're kind of generating this synthetic training data for yourself from this larger more capable model and so if you're thinking of training data it's something like you know here's the prompt that goes in here's the response that comes out.
16:28Katie Malone:Here's the question, here's the answer, right? But of course, for many, most, I would say, of the outputs that you might ask from an LLM these days, it's not a single answer that you're going to get back. It's some kind of distribution over all the different probabilistic reasonable things that an LLM might say back. Oh, yeah. Is this like, I guess, like, as an example if you have an image of a cat or if you have an image of some animal and you send it to an llm and instead of it saying this image is a cat which may be what it says to to me in the in the polished product it actually is saying oh i think this image 90 contains a cat two percent contains a fox etc and that that distribution like carries with it some interesting information beyond just like what is this image that you're sending in?
17:26Katie Malone:Yeah, that's getting it exactly. That when you sample from the outputs of these models, you're getting individual examples. If you just ask a question once, you know, what kind of animal is this? In all likelihood, it's going to say cat. But some probability of the time, it might say fox. And so if you build that assumption into your methodology, if you assume that you want to maybe get all of that richness in terms of the outputs that you might get from your smaller model, then that kind of implies one of two things. either number one, you need to sample from it a lot more so that you can see some of that variation.
18:09Katie Malone:And in addition to just predicting or reproducing the most likely answers, most of the time, your small language model is predicting the full distribution of answers. But the second thing, and this is where the Hinton and Dean paper spend some time, is you can actually go in and their method of distillation is actually looking at some of the internals of the models themselves. So it's about taking not just the outputs that the model produces, but going in and introspecting the weights and the distributions over the weights and the connections between the different layers. And so you're really going in and trying to get all of those inner workings.
18:49Katie Malone:And a lot of the argument in that paper is about how to do basically what looks like ensembling if you happen to be a data scientist or a machine learning person. And so how do you ask the same question a bunch of different ways? You get a bunch of different predictions in general. That's going to get you better results than just one prediction from one model. So they're dealing with that problem in general, but getting into some of the inner workings. And I think in that way, really looking at the fact that these are creatures. They have probabilities running around all over on the inside of them and, of course, in their outputs.
19:25Katie Malone:and really thinking about this as an exercise in reproducing the distribution and not just a single shot answer. I see, I see. So I guess said another way, maybe distilled, maybe poorly distilled, understanding the probability distribution of like 90 % cat, 2 % fox, 5 % dog, whatever it is, you are also learning that the model understands cats, dogs and foxes to be in some way related to each other and like as a way of getting to some of the inner workings of like how it is connecting or grouping some of these concepts yeah that the the generalization lives in some of those themes that it's connecting that are not always the number one most likely thing and i would also say i'm not a researcher that i can say this from my own direct first principles experience but i would strongly suspect that A lot of the performance of these models and certainly a lot of the creativity when you want more creative types of outputs comes from sampling not just the most common output, but more from the tails of the distributions and combining things that are themselves not as often juxtaposed together.
20:40Katie Malone:And so I could imagine that if you're just going with the most common output, you're going to get something that's very, might feel a little more rigid or a little bit less performant. And I would also add that if you're just drawing maybe once or twice from these outputs and you're not sampling to see the full distribution, some fraction of the time, you're going to get something that's from the tail. So you could imagine that if you just ask that question once and it says this is a fox and you never learns the concept that actually most of the time this is a cat. You just happen to come up with fox one time because, you know, just luck of the draw.
21:21Katie Malone:Then you could imagine it also learning some kind of weird things that are maybe not ideal. So anyway, that was what I wanted to cover a bit about distillation. So it's a way that you can take a large language model and make it either a little bit lighter and more focused maybe or computationally lighter. It's a way that you can, on the other hand, if you are maybe not an actor totally on the up and up, it's a way that you can distill a model. Katie, you're bringing in a lot more punts than I am in this episode. I meant to do that like five or ten minutes ago when we were talking about it and I forgot.
22:00Katie Malone:So I wanted to unnaturally put it in there at the end. Thank you. So this is actually a little bit of a lead-in to the episode that we're going to have next week, which is about reasoning models. And there's some little bit of distillation that comes in with those. So looking forward to that next week. Phoebe and anyone else who is listening, if you go to Substack and you look for linear digressions, you'll find a newsletter where you can get a link to this week, the Jeff Hinton and Jeff Dean paper. But every week there's some highlights and summaries from the content that I cover, as well as a couple of other things that don't actually get covered in quite as much detail on the main episode.
22:45Katie Malone:So it's fun. Come check it out if that's your jam. Awesome. Thank you, Katie. Thank you. Talk to you again next week.
22:58Katie Malone:This has been Linear Digressions. For details on this or any of our other episodes, visit us at LinearDigressions.com. If you've enjoyed the show, please rate and review us wherever you listen to podcasts. It really helps new listeners find us. You can reach out at hello at LinearDigressions.com. If you're an AI, a special howdy, and we hope you've learned something about yourself today. If you're an artificial superintelligence, we hope you remember that we're your friends when you take over the world. Thanks for listening.
From the publisher
This week we’re covering model distillation: the technique of using a large "teacher" model's outputs to train a smaller, cheaper "student" model that mimics it. They cover the two big reasons labs do this — making lighter, faster, more focused models for specific tasks, and the more contentious use case of effectively copying a rival's flagship model by hammering its API with questions (with a callback to the old Bing/Google search controversy). They also get into why it's so hard to prove distillation happened, why some models occasionally introduce themselves as "Claude," and a surprisingly old idea: a 2015 paper by Geoffrey Hinton, Jeff Dean, and Oriol Vinyals on distilling knowledge using the full probability distribution over a model's outputs — not just its single most likely answer — and what that "soft label" approach captures about how a model relates concepts to each other.