In short
Explains why LLM prompting works using a Bayesian meta-learning view of in-context learning, and compares hard prompt tuning vs prefix tuning (soft prompts), including when prompting fails and how weight tuning (full/LoRA) can succeed.
Guests
No guests mentioned; it’s a solo “Deep Dive” episode.
Key claims
Pre-training makes transformers approximate Bayes-optimal sequential predictors; prompts act as conditioning via a learned inference algorithm. Soft prompts (off-vocabulary vector prefixes) outperform hard prompts because they directly control internal activations. “Prefix Tuning Limitation I”: prompting fails on inherently multimodal tasks because conditioning collapses to a single mode (Dirac delta).
Notable examples
Coin-flip meta-training; single 0.2-bias coin solved Bayes-optimally only by soft PT. Two-coin mixture (0.2/0.8) fails for all prefix tuning (even L25) but is solved by full weight tuning and LoRA. Soft prompts on randomly initialized transformers work surprisingly well; random LSTMs do not.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOUnderstanding Prompting
0:45 to 2:35
Exploring the fundamentals of optimal prompting for LLMs.
“and importantly, where it starts to fall apart.”
Memory-Based Meta-Learning
2:35 to 5:06
Discussing how models learn to predict and adapt through meta-learning.
“The initial conditions, it needs to zero in on the specific task we want it to do now, like setting the coordinates on a map for an explorer who already knows how to navigate.”
Soft vs. Hard Prompting
5:06 to 8:41
Comparison between traditional hard prompting and innovative soft prompting.
“They pre-trained models, both a transformer and an older LSTM, on sequences generated by random coins.”
Applications and Limitations of Soft Prompting
8:41 to 11:40
Investigating scenarios where soft prompting excels and where it fails.
“Prompting is about steering what's already there from pre-training.”
Weight Tuning vs. Prompting
11:40 to 14:01
Examining the trade-offs between weight tuning and prompting methods.
“LSTMs apparently need the training process to build up that kind of sophisticated state update logic.”
Exploring Transferable Task Knowledge in LLMs
14:01 to 14:28
Learn about the potential for prompting to capture abstract, transferable knowledge across different LLM architectures.
“and apply it successfully to different LLMs, maybe even ones with different architectures?”
Transcript
Automatic transcript. May contain errors.0:00Welcome back to the Deep Dive. Today we're really getting into the weeds on something fundamental. What actually makes a prompt for an LLM optimal? You know, for ages, it feels like we've treated prompting more like, I don't know, alchemy. We just throw words at these huge models hoping something clicks. The really amazing thing about today's big models is how fast they learn on the fly, that in-context learning. But we haven't really had a solid theory for why prompting works the way it does. That's exactly what we're tackling today. We're moving past the guesswork. The mission here is to build up a proper conceptual framework for optimal prompting.
0:36And this framework comes from looking at these models through a Bayesian meta-learning lens. We'll be using a really neat, simplified setup, just sequences of coin flips to see where this theory holds up, and importantly, where it starts to fall apart. And the key insight, it's actually pretty elegant when you train a massive model by just making it predict the next token really well, minimizing log loss over tons of data. you essentially force it to become something very close to a Bayes optimal sequential predictor. Whoa. Okay. Hold on. A Bayes optimal predictor. Yeah. Think of it like training at a detective.
1:08You don't teach them about every single possible crime. You teach them the principles of deduction so they could look at the first few clues to prompt and figure out the specific nature of the case they're dealing with right now. Okay. So prompting isn't about teaching the model something brand new. It's about giving the right initial clues, conditioning this predictor that's already been trained to be incredibly smart about inference. Right. Okay, let's unpack that a bit more. You're saying all that massive pre-training doesn't just create a giant dictionary. It creates a machine that's learned the rules for learning new rules super fast.
1:42How does that Bayesian predictor thing actually, like, work inside the network? Yeah, the link is something called memory-based meta-learning. So during pre-training, the network sees a huge mix of different tasks, right? The meta distribution. To do well overall, it can't just memorize. It has to learn a general method or algorithm. This internal method lets it figure out the hidden specifics of the current task just by looking at the input context, the prompt, without actually changing its core programming, its weights. And this learned process is so good that what the network predicts next looks almost exactly like what a formal explicit Bayesian calculation would give you.
2:21The transformer, with its attention mechanisms and all that, is effectively doing Bayesian inference on the fly based on the prompt. Huh. So if the model's already been meta-trained into this, like, super predictor, then prompting is just about giving it the right setup. The initial conditions, it needs to zero in on the specific task we want it to do now, like setting the coordinates on a map for an explorer who already knows how to navigate. Precisely. And finding the optimal prompt becomes a search for the most efficient prefix possible. The shortest string of input that crams the most useful information about your target task into the model's internal state right at the beginning.
2:56Okay. And usually that input is words, right? Normal language tokens, what the paper calls hard tokens. But you mentioned a big breakthrough came from these things called soft prefixes or soft tokens. What are those exactly? How do they compare? Right. So we're looking specifically at prefix tuning. You stick a short sequence of input vectors right at the start of whatever you feed the model. And we compare the classic hard prompt tuning, RPT, searching for the best sequence of actual words or tokens from the model's vocabulary against soft prompt tuning. Soft PT. Now, soft PT uses sequences of just, well, numbers, real valued vectors, high dimensional numerical lists.
3:33And crucially, these get inserted directly into the model's internal embedding space. So they aren't words at all? Nope. They're explicitly off distribution inputs. They're like, ah, secret codes that don't map to any word the model learned during its initial training. Okay, that's where it gets really wild for me. Why use these totally artificial non-language inputs? The whole model is built on language. It feels like trying to talk to it in Martian or something. It does seem counterintuitive, right? Yeah. But mechanistically, these soft prefixes are just way more powerful. Think about hardtoken's actual words.
4:06You're limited by the existing vocabulary, the structure of language. Trying to fine-tune the model's internal state using only words is kind of like, I don't know, trying to perform delicate surgery with a wrench. Okay. You're stuck with the blunt instruments the vocabulary gives you, but soft prompts. They bypass language entirely. They give you direct, high-precision numerical control over the model's internal activations and hidden states. They could push and pull the network's internal workings in ways that no sequence of actual words of the same short length could ever achieve. This lets them inject that PASC information super efficiently, kind of getting around the limitations baked into how the model normally processes language.
4:46So it's less about meaning and more about directly manipulating the machine state. The soft prompt is like a direct neural interface, while the hard prompt is like shouting instructions from across the room. That's a great way to put it. And to test this advantage and the whole Bayesian idea, the researchers used this very clean mathematical setup, coin flips. They pre-trained models, both a transformer and an older LSTM, on sequences generated by random coins. Random coins, meaning? Meaning the bias of the coin used to generate each sequence was chosen uniformly at random. So sometimes it was a fair coin, sometimes heavily biased heads, sometimes tails.
5:24The model had to learn to figure out the bias from the sequence itself. This essentially trains it to be a general predictor for any coin bias. Okay, got it. So that's the pre-training. What about the specific task they wanted to optimize for, the positive case? Right, the success case. That was the single coin task. Here, the coin always had a fixed bias of 0.2, so always leaning towards tails. This cask is definitely one of the possibilities covered by the random coins pre-training. Makes sense. So what happened when they tried to use prefix tuning to make the models nail this specific 0.2 bias coin?
5:57Well, this is where SOFTBT just completely dominated. Using a really short prefix, L6, I think it was, soft prompting, was clearly the best method. And here's the kicker. For both the transformer and the LSTM, SOFTBT was the only prefix tuning method tested that actually reached the theoretical Bayes optimal performance on that single coin task. It achieved perfect prediction for that specific coin bias. Wow. Okay, so these weird non-word vectors were better than any actual phrase they could find. And achieved optimality where the best sequence of actual words, RPT, just couldn't quite reach. It's strong empirical evidence that soft prompts are superior because they can manipulate that internal state more directly, more effectively than language, fitting perfectly with the Bayesian conditioning idea.
6:42OK, so soft PT is the champ. If the target task is simple, like one specific coin, and it's something the model basically saw during pre-training. But if this Bayesian view is right, doesn't it also predict where soft PT should, well, fail? It absolutely does. And this brings us squarely to a fundamental limit of this whole prefix tuning approach. The paper calls it Precicst Tuning Limitation I. A theory predicts that if your target task isn't just one simple thing, but it's inherently multimodal. Multimodal, meaning... Meaning the task itself requires the model to juggle multiple distinct possibilities or modes simultaneously.
7:19Think of it like needing to believe two contradictory things at the same time. Right. So the single coin task was easy, just predict bias 0.2. But what if the target was, say, a 50-50 mix of a 0.2 coin and a 0.8 coin? The model needs to somehow represent both possibilities accurately. Exactly. That's the problem. The Bayesian predictor, when conditioned by the prompt, tends to want to collapse its belief, its posterior distribution onto a single point estimate, what's called a Dirac delta. Basically, it wants to settle on one answer, one PASC explanation. It struggles to keep two distinct modes active in its internal state when prompted.
7:55So it might just average the two or pick one, but it can't properly handle the mixture itself just through prompting. That's the theoretical prediction. And so the next experiment tested exactly that. The target task was the two-coin mixture. Sequences generated half the time by a 0.2 bias coin, half the time by a 0.8 bias coin. The multimodal wall. What happened when they threw soft PT at it? The theory held up. It was a clear failure to reach the optimum. No prefix tuning method. Not soft PT, not hard PT, not even when they tried much longer prefixes like L25 could get anywhere near Bay's optimal performance on this two coin mixture.
8:34Now, soft PT was still the best of the prompting methods. It definitely improved performance over doing nothing, but it hit a hard ceiling well below what perfect performance on that mixture task could look like. That's really telling. It shows the boundary, right? Prompting is about steering what's already there from pre-training. If the behavior you need, like perfectly handling this two-mode mixture, isn't something that conditioned, pre-trained model can represent, then just steering isn't enough. Precisely. You hit the limits of conditioning the existing predictor. But they didn't stop there, did they?
9:04If conditioning fails, what about actually changing the model? Did they compare prompting to actually fine-tuning the weights? They did, and it's the crucial other half of the story. Right. Methods that actually change the model's weights, specifically full weight tuning, full WT, where you retrain everything, and Loro weight tuning, Loro, a more efficient method. Yeah, low-rank adaptation, very popular. Right. Those methods could reach Bayes' optimality on the two-coin mixture task. By actually altering a model's internal update rules, changing his parameters, they could learn to handle that complex, multimodal behavior where just conditioning via prompts failed.
9:42Okay, so that draws a really clear line. Prompting leverages the learned Bayesian inference but is limited by it. Weight tuning can break past those limits by fundamentally changing the model, at least for that specific task. Exactly. Now, before we wrap up, there was this other finding that seemed almost unrelated, but fascinating. The mechanistic surprise. They tried using soft prompts on models that hadn't even been trained. Just random weights. Yes. This was super interesting. It basically takes the whole Bayesian learning story off the table for a moment and just asks, how much power does a soft prompt have purely on the architecture itself before any learning?
10:18And the result was pretty stunning. Soft prompting an untrained transformer one just initialized with random weights could actually achieve performance surprisingly close to optimal on tasks like the two-coin mixture and even the general random coins task. Get out. So a random network guided only by this artificial soft prompt could almost solve these complex prediction tasks. Almost. It suggests that the transformer architecture itself, even with random parameters, somehow contains the latent capability for complex in-context learning algorithms. The soft prompt acts like a key, or maybe like direct instructions to that latent machinery, effectively reprogramming the random network on the fly for the task.
10:56So the architecture itself, the attention, the residual connections, is doing a lot of the heavy lifting even before training optimizes the weights. Training just tunes that inherent capability. That seems to be the implication. The soft prompt finds a way to configure those random weights to do the job, almost like finding the right settings on a very complex, uncalibrated machine. Was this also true for the LSTM? The older architecture? Ah, no. And that's a critical distinction. Soft prompting the untrained LSTM had almost no effect. It barely budged the performance. Okay, wow. This strongly suggests that this amazing ability of soft prompts to work on random networks is specific to the transformer architecture.
11:37The way transformers are built seems inherently amenable to this kind of direct state manipulation via soft inputs, even before learning. LSTMs apparently need the training process to build up that kind of sophisticated state update logic. That architectural difference is huge. So, okay, let's recap this journey. We've seen that prompting works because these massive models learn to be Bayesian predictors during pre-training. Right. Meta-learning makes them inferential engines. And soft prompts beat hard prompts because they can manipulate the model's internal state more directly. These off-distribution inputs are just more efficient injectors of task information.
12:11Yep. Better steering mechanism. But there's a limit. Prompting, even with soft prompts, fails when the target task is too complex. Specifically multimodal, because you're just conditioning the existing predictor, not changing its fundamental capabilities. Exactly. The prefix tuning limitation I. Which leads to the big practical question. If white tuning methods like LoRa can overcome that limitation and reach perfect performance where soft PT fails, like on that two coin mixture, why would anyone still bother with prompting? Why not just Laura tune everything for specific tasks? That's the multi-million dollar question, isn't it?
12:48And there are good reasons. First, weight tuning. Even Laura can sometimes mess up the model's general abilities. You risk catastrophic forgetting of other things it knew. Right. You specialize it too much. Potentially. Second, the in-contact learning mechanism that prompting tats into seems, in practice with giant real-world models, to sometimes generalize more broadly and flexibly than specific weight tuning. Prompting leverages the model's existing broad inferential skills. So prompting is maybe safer for generalization, potentially more flexible across tasks, even if it can't always hit the absolute performance peak for a single complex task?
13:25That seems to be the tradeoff, yes. Okay, so here's a final thought for you to chew on building direct liowness. We know weight tuning, like Loray, creates adaptations that are stuck to that specific model. You can't easily transfer Loray weights trained on one model to another. Right. They're model-specific changes. But soft prompts, they're just sequences of vectors, right? They interact with the model's input state, leveraging that potentially more general meta-learned inference capability. So what if you could tune a really good soft prompt for a task? Maybe an expensive process, but then...
13:57What if that tuned soft prefix could be reused? Could you take that optimal sequence of vectors representing deep task knowledge and apply it successfully to different LLMs, maybe even ones with different architectures? as long as they were trained on broadly similar data. Implying that the prompt captures a kind of abstract, transferable task knowledge. Exactly. It suggests prompting might be tapping into something more fundamental and potentially more generalizable than specific weight changes. It's an open question, but definitely something to think about next time you're crafting a prompt.
From the publisher
The academic paper investigates prompt tuning and in-context learning through a meta-learning and Bayesian lens, positing that optimal prompting can be understood as conditioning Bayesian sequential predictors. The authors detail how meta-trained neural networks, like LSTMs and Transformers, function as Bayes-optimal predictors and explore the theoretical limitations of prompting, particularly for complex, multimodal target task distributions. Empirical experiments on coin-flip sequences confirm these theories, demonstrating that Soft Prompting—using sequences of real-valued vectors—is significantly more effective than hard-token prompts, even showing surprising efficacy in fine-tuning untrained networks. Ultimately, the research provides a fundamental conceptual framework for understanding the mechanisms and constraints of prompt optimization.




