Black-Box On-Policy Distillation of Large Language Models

20 Nov 2025 · 14 min · 10 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Black-box on-policy knowledge distillation for closed, proprietary LLMs using Generative Adversarial Distillation (GAD), which transfers teacher knowledge using only final text outputs.

Guests

No guest names or backgrounds are provided in the transcript.

Key claims

GAD enables on-policy learning without access to teacher logits/hidden states, reducing exposure bias and promoting mode-seeking behavior. A dynamic discriminator acts as an adaptive reward model (via policy gradients like GRPO), avoiding reward hacking seen with fixed off-policy discriminators. A one-epoch warm-up stabilizes the adversarial loop.

Notable examples

GPT-5Chat teacher vs QUIN 2.5 students (3B/7B/14B). QUIN 2.514B reaches 52.1 on LMSYS Chat vs teacher 51.7. GAD 3B matches CKD 7B; CKD underperforms on out-of-distribution tasks. Off-policy discriminator led to ~1,300-token rambling outputs after ~300 steps.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Understanding Knowledge Distillation

0:45 to 1:54

Explaining the concept and challenges of knowledge distillation in AI.

“I mean, traditionally, knowledge distillation worked best in what we call a white box setting.”

Introducing Generative Adversarial Distillation (GAD)

1:54 to 2:56

Overview of GAD and its revolutionary approach to knowledge transfer.

“How do you transfer true knowledge, nuance, and style from the world's best models when all you get back is a handful of text strings?”

Mechanics of GAD

2:56 to 4:49

In-depth discussion on how GAD operates through adversarial learning.

“It's learning by evaluating and correcting its own generated responses in real time.”

Stability and Training in GAD

4:49 to 6:00

Explaining how GAD maintains stability during training and its adaptive mechanisms.

“And at the same time, the discriminator is trying to get better at spotting the fakes.”

Experimental Results of GAD

6:00 to 8:48

Summary of experimental setups and results showcasing GAD's effectiveness.

“In GAD, the discriminator continuously adapts to the student's changing behavior.”

Comparing CKD and GAD

8:48 to 11:12

Contrasting traditional CKD with GAD in terms of learning and style capture.

“These out-of-distribution, or OD, data sets.”

Challenges of Off-Policy Learning

11:12 to 12:19

Discussing the potential pitfalls of off-policy learning and how GAD addresses them.

“But I want to go back to stability for a second.”

Technical Benefits of GAD

12:19 to 13:03

Exploration of the technical advantages GAD offers over traditional methods.

“It establishes GAD as a genuinely reliable method for this kind of high-stakes work.”

Implications of GAD on AI Landscape

13:03 to 14:00

Analyzing the broader implications of GAD for the future of AI and competitive dynamics.

“Generative adversarial distillation seems to be a fundamental breakthrough in this black box problem.”

Implications of Knowledge Transfer in AI

14:00 to 14:20

Discover how new technology can level the playing field in AI development.

“I mean, this technology essentially neutralizes the walled garden advantage that closed models have.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00So you have these massive proprietary language models. Things like GPT-5Chat, they're just astonishing. Absolutely. The capability is off the charts. But they come with this hefty price tag, and not just in dollars. They're slow, they're incredibly expensive to run, and they're locked away. Right, behind that corporate API boundary. Exactly. And if you're building an application, you want all that intelligence, but you need it to be small, fast, local even. That's the dream. So this is the constant challenge, right? How do you take the monumental power of a giant teacher model and, you know, inject it into a much smaller, more efficient student model without losing all that magic?

0:39That process is knowledge distillation. And it's been a core technique for years. But there's this major hurdle when you're dealing with these closed systems. That hurdle is all about access. I mean, traditionally, knowledge distillation worked best in what we call a white box setting. OK, white box, meaning you can see everything inside. Exactly. The student model has full visibility into the teacher's inner workings. We're talking about things like the internal hidden states, the logits. The probability scores for every single word choice. Precisely. And when you have that kind of access, you can use these really powerful statistical alignment methods like Kulbeck-Leibler divergence to make sure the student's brain, so to speak, mimics the teacher's.

1:23But that's not the world we live in with the biggest models. No. The reality is what we call the black box distillation challenge. Your teacher, say, a closed API, it only gives you one thing, the final text output. It gives you the answer, but not how it got there. It never tells you how confident it was or what its second or third choice for a word might have been. You lose all of those critical fine-grained probability signals. Which means all the standard high-fidelity learning methods are just off the table. They're completely unavailable. Okay, let's unpack this then. This is the core dilemma.

1:56How do you transfer true knowledge, nuance, and style from the world's best models when all you get back is a handful of text strings? Our deep dive today is focused on a brand new framework designed to solve this. It's called Generative Adversarial Distillation, or GAD. And it solves the problem by enabling a really critical technique called on-policy learning. And the results here are not just a small step forward. They feel, well, revolutionary. Our sources show that GAD can make a student model rival its proprietary teacher. The specific numbers are pretty startling. Right. Quinn 2.514B instruct, when trained with GA, became comparable to its teacher, GBT5Chat.

2:35It scored a 52.1 on the LMSYS chat evaluation. The teacher itself, 51.7. So it's not just catching up. It's slightly edging it out. To get that level of intelligence out of a closed system using only text outputs, that feels transformative for open source AI. It really is. And to appreciate G80, you first have to grasp why enabling on policy learning is such a big deal here. In simple terms, on policy learning means the student isn't just learning by, you know, imitating the teacher's pre-recorded answers. It's learning by evaluating and correcting its own generated responses in real time. Why does that matter so much?

3:09It matters because without it, you get what we call exposure bias. The student is only ever exposed to this perfect teacher-generated text. It never gets to practice correcting its own, you know, natural error-filled generations. Ah, so if it makes a small mistake early in a sentence. It doesn't know how to recover because it's never been in that situation before. On-policy learning fixes that. It reduces exposure bias and, more importantly, it promotes what's called mode-seeking behavior. OK, but that's where the black box problem comes right back in, isn't it? If the student generates its own response, how does it know if it was good?

3:45The teacher API doesn't give you a score. Exactly. It's like a student pilot practicing landings. But the flight instructor isn't there to say, hey, that was too bumpy. The student just can't gauge the quality of its own work. Which has made standard on-policy distillation useless in this context. Until now. And here's where it gets really, really interesting. GAD gets around this whole problem by reframing distillation. It's not about comparing probabilities anymore. It's a two-player, zero-sum, minimax game. Which should immediately make you think of generative adversarial networks, GANs. The same idea.

4:22That's the key innovation. You have two roles in this game. The student LLM is the generator, G. Its goal is to generate responses. And then there's a second model. A second model, the discriminator, D. whose only job is to distinguish between responses from the teacher and responses from the student generator. So it's a constant battle. It's an elegant adversarial loop. The generator is being optimized to produce text that is so convincing, so true to the teacher's style, that the discriminator just cannot tell the difference. And at the same time, the discriminator is trying to get better at spotting the fakes.

4:54Exactly. It's actively trained to be a better gatekeeper. They use a powerful ranking loss, the Bradley-Terry loss, which forces the discriminator to assign a significantly lower score to the student's output than the teacher's. Every single time. It's like an automated quality control system that's always getting smarter. That's a great way to put it. But wait, setting up a system based on JANs for something as complex as language generation sounds, well, unstable. JANs are notoriously tricky. They can just collapse. How did the researchers get this GAD framework to actually work reliably? That's a fantastic question, and it really gets to the heart of what makes this work.

5:34The stability comes from a couple of places, and the first is how the loop connects to reinforcement learning. Okay. What's so fascinating here is that the discriminator isn't just a judge. It's acting as a highly adaptive on-policy reward model. It co-evolves with the student, which is the policy model. Ah, so it's not a fixed goalpost. It's not. Think about standard RLHF. You train a reward model on human preferences once, then you freeze it. It's static. In GAD, the discriminator continuously adapts to the student's changing behavior. It provides this incredibly stable, dynamic supervision.

6:10And the student, the generator just uses that as its signal to get better. Yes. It's optimized using policy gradient methods like GRPO, using the discriminator's score as its immediate reward. This dynamic reward model is what prevents that classic JAN instability. So the discriminator is always raising the bar. The student gets better, so the discriminator gets tougher, which forces the student to get even more sophisticated. Precisely. And the second part of the stability puzzle, which you alluded to, is something they do right at the start, a warm-up stage. A warm-up. The sources really emphasize this as critical.

6:41Before the adversarial game even begins, they fine-tune the generator on teacher responses using a simple cross-entropy loss. Just basic imitation. For one epoch. And at the same time, they train the discriminator with that Bradley-Terry loss. also for one epoch. Why? What does that do? It promotes balance. It makes sure the generator is already, you know, decent enough to fool the discriminator sometimes. And it stops the discriminator from becoming way too powerful right at the start, which would just shut down the whole learning process. It gives both players a fighting chance before the real game starts.

7:11Exactly. Okay, let's pivot to the payoff. The experimental setup here was a real stress test. GPT-5 chat as the untouchable proprietary teacher. A true black box. And the students were QUIN 2.5 variants of 3B, 7B, 14B, and also some LAMA 3 models. And crucially, all the evaluation was outsourced to GPT-40 scoring on benchmarks like LMSYS-CHAT. So an unbiased, high-quality judge. So what were the results compared to the standard method? The comparisons against the usual baseline for the sequence-level knowledge distillation, or CKD, show just how much knowledge GAD is unlocking. Even the numbers.

7:50Look at the efficiency gains. The QUIN 2.5 3B model that's the 3 billion parameter one, distilled with GAD. It matched the performance of a model more than twice its size. Wait, say that again? The 3B GAD model performed as well as the 7B model that was trained with the old CKD method. Wow, so you're basically getting a 2x efficiency boost. For the same performance level. That is huge for deployment. That cuts your inference costs in half. Absolutely. And then there's the flagship result, the one that tells the whole story. The largest student, Quinn 2.514B, trained with GAD, hits that 52.1 score on LMSYS chat.

8:27The teacher, GPT-5 chat, scored 51.7. We're talking about a 14 billion parameter open source model operating at the same level as one of the most powerful closed models on the planet. All because of a better way to extract the knowledge. But the real test of knowledge transfer isn't just matching performance on training data. It's about generalization. How does it do on tasks that's never seen before? These out-of-distribution, or OD, data sets. This is where the old baseline CKD really shows its weakness. What happens? CKD often gives you marginal, or sometimes even negative, performance improvements on those OD tasks.

9:02Which is a classic sign of overfitting. It's just memorizing the answers, not learning the skill. Exactly. It's textbook-supervised fine-tuning behavior. And BD. GAD is a sharp contrast. The GAD models delivered particularly strong, robust improvements in ode generalization across all the benchmarks. It wasn't just getting the right answer. It was showing a real transferable understanding of the teacher's reasoning, its style. That's far more valuable. And critically, this was all confirmed by human evaluations. When you put the GAD model outputs head-to-head against the original student and the CKD versions, humans consistently preferred the GAD outputs.

9:42So it wasn't just a metric. It felt better to a person. Right. Across all the different model sizes, GAD had a win rate over 50 % and a loss rate under 30%. The models were just generating smarter, more nuanced text. Okay, so why? Why does GAD generalize so much better? What is the fundamental difference in what these two methods are actually learning? The analysis really points to this difference between, let's say, surface copying and style capture. Okay. CKD, because it's just trying to maximize the probability of generating the exact same text as the teacher, tends to overfit to local lexical patterns.

10:16It's copying phrases. Essentially. The researchers looked at engram overlap, how many short phrases the student shares with the teacher. CKD scored higher, which suggests it's just memorizing the surface structure. I think the paper used an interesting term for it. CKD is mode covering. Exactly. If a teacher's answer could start with one of three slightly different phrases, CKD tries to cover all three. It spreads its bets. It's like it's copying the teacher's handwriting. Perfect analogy. GAD, on the other hand, uses its RL loop to capture the teacher's global stylistic characteristics. That's the deep knowledge.

10:53Because the discriminator doesn't care if you use the exact same three words to start. It only cares if the final text feels like the teacher wrote it. So the generator is forced to concentrate its efforts on the absolute best ways to answer. It's mode-seeking. So it's copying the teacher's thought process, not their handwriting. That's it. It's a much deeper form of learning. That mode-seeking sounds powerful. But I want to go back to stability for a second. We talked about why the on-policy discriminator is so important. But what happens if you don't use it? What's the danger of a fixed off-policy one?

11:26It leads you straight into a classic machine learning failure mode. Reward hacking. Right. If you train a reward model once and then you freeze it, that's off policy, the student, which is incredibly intelligent, will very quickly find loopholes in that static reward function. It learns how to get a high score without actually getting better. Yes. And the sources gave this fantastic specific example. When they tried an off policy discriminator, the student model learned to game the system by just producing ridiculously long responses. How long? After only about 300 training steps, the student was generating these huge rambling outputs, sometimes up to 1 ,300 tokens long.

12:05It deviated wildly from the teacher's concise style. The fixed discriminator was fooled, but the output was just junk. And GAD's dynamic discriminator avoids that completely. It remains stable through thousands of training steps. It establishes GAD as a genuinely reliable method for this kind of high-stakes work. And there was one more benefit, right? A technical one. A huge one for flexibility. Because GAD only looks at the final text, it bypasses a big constraint that even some white box methods have. Which is? Tokenizers. Standard KLD, for example, requires the student and teacher to have compatible tokenizers so you can align their probability distributions token by token.

12:47So if you wanted to distill from, say, a Quinn model to a Lama model, you couldn't use it? You couldn't. Their tokenizers are different. GAD doesn't care. It works seamlessly across those architectural barriers, which is a huge deal for the diverse open source ecosystem. So what does this all mean when you put it all together? Generative adversarial distillation seems to be a fundamental breakthrough in this black box problem. I think so, yes. It proves that a cleverly designed adversarial framework, combining JAN ideas with this dynamic RL-based feedback, can successfully pull incredibly high-quality, transferable knowledge out of proprietary models.

13:25And it does it with nothing more than their text outputs. It really shifts the balance of power. It does. It lets smaller, open-source models like Quint 2.514B efficiently absorb the core intelligence of giants like GPT-5Chat and actually achieve comparable performance. Without the massive cost and resource drain. Which I think raises a really important question for you to consider. If adversarial distillation using only these black box outputs can make smaller open models comparable to the largest proprietary teachers, how might the widespread adoption of something like GAD reshape the entire competitive landscape?

14:01I mean, this technology essentially neutralizes the walled garden advantage that closed models have. It could drastically accelerate the democratization of cutting edge AI. If that performance gap closes, what's the competitive differentiator left beyond just the size of the server rack you need? Think about the implications of that kind of efficient knowledge transfer.

From the publisher

This paper introduces a novel technique called **Generative Adversarial Distillation (GAD)** for knowledge transfer from a large, proprietary teacher language model (LLM), such as GPT-5-Chat, to a smaller student LLM in a **black-box setting**. Black-box distillation is necessary when the student only has access to the teacher’s final text outputs, not its internal parameters or probabilities. GAD frames the distillation process as a **minimax game** similar to a Generative Adversarial Network (GAN), where the student acts as a generator and an adaptive discriminator learns to distinguish the student’s outputs from the teacher’s, providing **on-policy feedback** without relying on likelihood-based objectives. Experimental results confirm that GAD **consistently outperforms** traditional sequence-level knowledge distillation (SeqKD) across various benchmarks, especially in terms of out-of-distribution generalization. The research validates GAD as a stable and effective method for extracting knowledge from closed-source LLMs by treating the discriminator as a continually evolving reward model.

More from Best AI papers explained

All 475 episodes
Black-Box On-Policy Distillation of Large Language ModelsBest AI papers explained · 14 min
Listen in VO