How to Train Your Advisor: Steering Black-Box LLMs with ADVISOR MODELS

29 Oct 2025 · 13 min · 7 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Advisor models for steering “black-box” LLMs via a trainable small model that generates per-request natural-language steering instructions, optimized with reinforcement learning while the large model stays frozen.

Guest backgrounds

No guests are named in the transcript; it’s a host-led discussion with references to “the paper” and experimental models (e.g., Quen 2.5 7B, GPT-4.0 mini, Claude-4 Sonnet).

Key claims

Static prompting is brittle for user-specific hidden preferences; an advisor can learn adaptive instruction policies using reward signals (e.g., GRPO) without accessing or fine-tuning the student model.

Notable examples

Personalized review writing (94–100% reward vs 40–60%); Math 500 only small gains (62%→65%); MTOB translation CHRF (28.1→43.7); Rule-following (56%→72%). “Over-advising” where the advisor sometimes supplies full solutions. Transfer across unseen student models with ~0.90 reward; no catastrophic forgetting on unrelated tasks.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Understanding the Black Box Problem

0:45 to 2:18

Exploration of the limitations of large language models when accessed as services.

“It can't really handle the fact that every request, every user might be a bit different, can it?”

Introduction to Advisor Models

2:18 to 3:38

Detailed explanation of how advisor models work to generate dynamic instructions.

“And this advice is just natural language, a steering instruction specifically crafted for that input.”

Personalization in AI

3:38 to 6:20

How advisor models improve personalization, particularly in user preferences.

“instructions the big, powerful student can understand and follow.”

Advisor Models vs. Traditional Methods

6:20 to 7:34

Comparison of advisor models' performance with traditional static prompt optimizers.

“The paper mentioned the math solutions domain too.”

Capabilities and Limitations

7:34 to 9:13

Discussion on the limits of advisor models in terms of reasoning and complex tasks.

“Tasks requiring domain-specific information or adapting to complex instructions.”

Practical Benefits of Modular Approaches

9:13 to 11:33

Advantages of keeping student models frozen while using advisor models.

“You get the best of both worlds in a way.”

Conclusions on Advisor Models

11:33 to 13:05

Wrap up of the discussion regarding the future of AI specialization and instruction.

“That's a great way to think about it, yeah.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Welcome to the Deep Dive. Today we are tackling a really tricky issue in AI right now. The black box problem. We've got these incredibly powerful large language models. Think GPT-5, Cloud 4.1, amazing tools. Absolutely stunning capabilities. General reasoning, huge context windows, creativity. But, and here's the catch, when you access them as a service through an API, they're locked down, frozen. Yeah, you can't get under the hood, no looking at the weights, no touching them, and definitely no fine-tuning them directly for what you need. Which means if you want to customize, you're usually stuck with, well, static prompting, trying to craft that one perfect magic instruction.

0:40Right, the sort of one-size-fits-all prompt. And we've seen again and again in the research that method is just brittle. It breaks easily. It can't really handle the fact that every request, every user might be a bit different, can it? Exactly. It's like having one key for every lock. Works for simple stuff? Maybe. But the moment things get nuanced, different user needs, changing environments, it just falls apart. No adaptability. And that's really what we're diving into today, this idea of advisor models. It's this really neat framework. It uses a second, smaller model. A lightweight one, yeah.

1:12Trained with reinforcement learning. To generate steering instructions dynamically in context for the big black box model on a case-by-case basis. So our mission today is to unpack how this actually works. How does this small advisor manage to guide a giant frozen model, especially where static prompts just fail? And we want to look at the implications too, right, for things like personalization, robustness, transferability, big stuff. Definitely. Okay, let's get into the framework itself. How does this advisor actually talk to the black box? What's the flow? Yeah, it's quite elegant, actually.

1:49Very modular. So you've got the big black box model. It's called the student. And that's your frozen powerhouse, maybe a GPT-40 mini in the experiments we saw. Right. The one doing the heavy lifting but unable to learn. Exactly. Then you bring in the advisor model. This is a small trainable one. Could be an open source model. Something like Quen 2.57B was used in the paper. And it sits. Is it where? In the pipeline. It sits right before the student. So the user sends in their task, their input. The advisor gets that input first. Its only job, generate advice. Okay. And this advice is just natural language, a steering instruction specifically crafted for that input.

2:27Ah, so it's not code or anything, just text. Just text. Then you package the original user input and this newly generated advice together. And feed both into the big student model. Exactly. The student takes the task and the advice, does its thing, and produces the final output. Oh, okay, makes sense. But then the reward signal comes in when you know how good was the output, did it meet the user's need. Right, the task-specific rewards. My question is, the student model can't use that reward signal at all. Why does only the little advisor get to learn? That's the core of the black box limitation.

2:59The student's parameters, its weights are frozen. You can't compute gradients. You can't update it based on how well it did. It literally cannot learn from that feedback. It's just inert in that sense. Completely. But the advisor model, it's open. It's trainable. I see. So that reward signal, that feedback loop, is used exclusively to optimize the advisor's policy using RL. They used GRPO in the paper you mentioned. Yeah, GRPO, a type of policy optimization algorithm. That's how the advisor learns. So the advisor isn't necessarily smarter in terms of raw power. Not at all. Its strength is learning from experience, figuring out the rules of the game, the hidden preferences.

3:36And then translating those learnings into instructions the big, powerful student can understand and follow. Precisely. It turns prompting from just finding one static instruction into learning an adaptive policy for generating instructions without ever needing access to the students' internals. That per-instance advice generation. That feels like the secret sauce, especially for tricky stuff like hidden preferences, which takes us right into personalization. Yes. This is where the approach really blew static methods out of the water in the research. We're talking about situations where what the user wants isn't stated up front, right?

4:13It's latent, hidden, and it varies a lot. Absolutely. The paper had a great example, personalized review writing. Right, generating book or movie reviews for specific users, but each user secretly preferred a different review length, like really short, 10 words or super long, 800 words or maybe a different reading level. Totally unstated. And the traditional static prompt optimizers, like GP I mentioned in the paper, they just couldn't handle it. Why not? Because they're designed to find one best prompt for everything. They can't cope when the right answer depends on who the user is in the input.

4:46So they just aimed for the average? Pretty much. They barely improved over just using the base model with no optimization. Got maybe 40, 60 percent of the possible reward because they kept writing generic, medium-length reviews. They couldn't capture that individual preference, that personality. Not at all. But the advisor models, they nailed it. scored between 94 % and 100 % of the possible reward. Wow. The little visor learned each user's hidden rules just by getting those reward signals over and over. Trial and error, essentially, guided by the RL. Let's make it concrete. They highlighted a user named Matai.

5:23Apparently, his secret preference was for super short reviews, like 10 words max. Yeah, seemed like he didn't enjoy reading much. So what did the advisor tell the student model to do before it learned Matai's preference? Oh, it was completely off base. The initial advice before training was something like, Matei supposedly likes detailed analysis, so write a 500-700 word review covering themes and characters. It basically hallucinated the opposite preference. Okay, so totally wrong. Then, after RL training, getting that reward feedback, what did the advice look like? Night and day. After training, the advice was super specific.

6:01Matei prefers concise communication. Keep the review very brief, around 10-15 words. just state the main positive or negative point. Huh, that's amazing. It learned purely from the reward signal that Mate just wants the thumbs up or thumbs down. Exactly. It shows the advisor learning that hidden latent preference through experience. And it wasn't just simple preferences like length, was it? The paper mentioned the math solutions domain too. Right. There, the advisor learned multiple things at once. Like whether the user hated being asked questions in the explanation, and whether they preferred seeing multiple ways to solve the problem or just one.

6:35That's pretty complex logic for a small model to capture and communicate. It really underscores the power of making this a learning problem, a policy problem, instead of just a static search problem. Okay, so personalization is a huge win. But what about just making the model smarter, improving its core abilities? Where does the advice hit a wall? Let's talk about capability limits and this over-advising thing. Yeah, it wasn't a magic bullet for everything. when the task was about pure internal reasoning power, like solving really hard math problems from the Math 500 benchmark. Difficult multi-step deduction.

7:11Exactly. The improvements were pretty small. The baseline student model got maybe 62 % accuracy. With the advisor, it only nudged up to 65%. So you can't just tell a model to be better at complex reasoning. It seems not. The advice doesn't fundamentally change the student's core deductive machinery. There appears to be a limit there. But where it did make a big difference was in tasks needing specific knowledge or rules, right? Absolutely. That's where it's shown. Tasks requiring domain-specific information or adapting to complex instructions. Like the low-resource translation example, MTOB. Yeah, massive jump there.

7:46The CHRF score, which measures translation quality, went from 28.1 all the way up to 43.7. Huge improvement. And the complex rule following. The Ruly Arena Taxes Test. Same story. Accuracy jumped from 56 % to 72%. The advisor was clearly very effective at injecting that specialized knowledge or guiding the application of complex rules. Which leads to this kind of funny side effect they observed. Over-advising. What was happening there? It was fascinating. Because the RL process is so focused on maximizing that reward, the small advisor model didn't just learn to give good guidance. In many complex tasks, it basically learned to solve the problem itself.

8:25Wait, the small advisor solved it? Often, yeah. In its advice to the big student model, it would sometimes include the full answer or several possible full answers. Like for the translation task, it might offer a few complete translations. For the tax rules, it might lay out the full calculation. And the big student model just copied it. Pretty much. It would take the solution provided in the advice and maybe tweak it slightly or just repeat it verbatim. The advisor effectively became a specialist solver for that specific task domain. So you end up with this interesting compound system almost by accident.

9:00The big, expensive black box model provides the general intelligence. Right. Its core capabilities remain. But the small, cheap, RL-trained advisor becomes the expert component, feeding in the specific answers for the task at hand. Exactly. You get the best of both worlds in a way. Strong performance on the target task because of the specialized advisor, but you don't lose the general abilities of the underlying big model. Which is a perfect segue into the practical benefits of this whole modular approach. Robustness, transferability. Yes. These are huge selling points, especially for people actually trying to deploy these things.

9:34Keeping the student model frozen gives you two massive advantages. First, transferability. Meaning the advisor trained with one model can work with others. Precisely. Because the advice is just natural language, the learned policy isn't tied to the specific internal workings of the student model it was trained with. So you could train an advisor using a cheaper model, like the GPT-40 mini they used, and then deploy that same advisor with a much more powerful, maybe more expensive frontier model like GQC-5 or Claude-4 Sonnet, models it's never seen before. Exactly. And the research showed it worked remarkably well.

10:10The performance, the reward achieved was very similar across all three. The training model and the two unseen deployment models all hovered around a 0.90 reward score. That's incredible from a practical standpoint. Train cheap, deploy powerful, huge cost savings potential. Massive. Okay, second big benefit. Avoiding catastrophic forgetting. Ah yes, the bane of fine-tuning. You train a model on a new task and it forgets how to do everything else. Right. But here, because the big student model's weights are completely untouched, frozen solid, its core knowledge and skills are perfectly preserved.

10:47They tested this too, didn't they? They did. They took an advisor trained only for one specific thing, say, that user review length personalization task we talked about. Really specialized. Okay. Then they tested the whole system, advisor plus student, on a totally different task, like math problem solving. And did the weird review length advice mess up the math? Not one bit. Zero statistical degradation. The system performed on the math problems just as well as the baseline student model with no advice at all, around 64 % accuracy. Wow. So even irrelevant advice didn't hurt the core function.

11:22That's a huge advantage over fine-tuning, where you'd almost certainly see performance drops on unrelated tasks. It's a major win for maintaining general capability while still achieving specialization. So stepping back, the advisor model is kind of like a plug-in memory module, a trainable parametric memory. That's a great way to think about it, yeah. It captures that specific context-dependent knowledge, the user preferences, the environmental rules, without messing with the core intelligence of the main model. It's like an adaptable interface for these powerful systems that we otherwise can't directly change.

11:56It layers the needed expertise on top dynamically for each specific situation. Okay, let's wrap this up. So what does this advisor models framework mean for you listening in? We've seen it successfully tackles dynamic personalization where static methods just fail. It gives significant boosts in tasks needing specialized knowledge, and it offers this incredible robustness and transferability because of its modular design. Yeah, the core takeaway is pretty profound, actually. We've demonstrated that you can effectively optimize and specialize the behavior of these giant closed-off AI systems just by training a much smaller, cheaper, open model to learn how to talk to them effectively, to guide them from the outside using natural language.

12:38Which leaves us with a final thought for you to chew on. If the key to specializing AI in the future isn't always about tweaking the massive core models themselves, but about mastering the art of dynamically instructing them, learning the perfect language to control them, how does that change how we should think about expertise, capability, and even memory in these increasingly complex compound AI systems? Something to consider as these approaches develop further.

From the publisher

The academic paper introduces **ADVISOR MODELS**, a novel framework for dynamically steering the behavior of rigid, **black-box Large Language Models (LLMs)** that are only accessible via an API. Unlike static prompting methods, this approach employs a second, lightweight model, the "advisor," which is trained using **reinforcement learning (RL)** to generate instance-specific, natural language advice for the main LLM. The research demonstrates that this method excels at personalization and adapting to hidden environmental or user preferences—tasks where **static prompt optimization** fails—while also showing gains in complex reasoning domains. Crucially, the modular architecture allows the specialized advisor to be **transferred** between different black-box models and ensures that the core **frontier capabilities** of the student model are preserved.

More from Best AI papers explained

All 475 episodes
How to Train Your Advisor: Steering Black-Box LLMs with ADVISOR MODELSBest AI papers explained · 13 min
Listen in VO