In short
Explains a mathematically proven view of in-context learning: a prompt is equivalent to applying temporary “implicit weight patches” inside transformer blocks, even for modern bias-free architectures.
Guests/backgrounds
Adrian Goldwasser, Michael Munn, Javier Gonzalvo, and Benoit Derin (Google Research and University of Cambridge collaborators). No other guest bios given.
Key claims
Input controllability (adjusting MLP front-door matrices W-gate and W-up) and output controllability (compensating residual pathways via RMS-norm scaling parameter M) make prompt processing structurally equivalent to running without the prompt plus computed parameter updates. Theorem 5 claims universality across architectures (Gemma, Llama, Mixtral/MoE, GPTJ).
Notable examples
“Mars robot” instruction-tuned Gemma-3 (1B/4B) matches tokens exactly after removing all context; multimodal cat image prompt also matches after removing image/text. Precision hiccup: float32 gives near-perfect match; bfloat16 drops to 87.5%, improved to ~98% via analytical RMS-norm inversion. Security implication: prompts as temporary architectural rewiring could be exploited.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOIn-Context Learning Unveiled
0:47 to 2:48
Delve into groundbreaking research explaining how AI models learn from prompts.
“And that is exactly the mission of our deep dive today.”
The Challenge of Modern Architectures
2:48 to 4:18
Discover the limitations of early theories in light of complex modern AI models.
“To understand the breakthrough, we have to look at the starting point.”
Understanding Input and Output Controllability
4:18 to 5:44
Learn how input and output controllability maintains AI performance amidst changes.
“Which sounds like an absolute bowl of alphabet soup.”
Universal Theorem Across AI Models
5:44 to 9:50
Examine how new research findings apply universally to various AI architectures.
“And they didn't just prove it exists in theory.”
Real-World Experiments with AI Models
9:50 to 13:11
Explore experiments demonstrating how AI models respond to prompts in practice.
“Those are the models that dynamically route data to different specialized subnetworks, right?”
Precision Hiccup and Real-World Challenges
13:11 to 14:02
Understand the trade-offs of precision in AI model deployments affecting results.
“Okay, but as with all bleeding edge tech, there has to be a collision between perfect theory and messy reality.”
Understanding Numerical Stability in AI
14:02 to 16:52
Learn how using lower-precision data types can affect AI performance in critical tasks.
“This essentially truncates the decimal points, saving huge amounts of memory at the cost of slight mathematical precision.”
Mathematical Innovations in AI Training
16:52 to 19:34
Discover how researchers solved precision issues to enhance AI training on cheaper hardware.
“This stabilization trick pushed the accuracy in Bfloat 16 back up to a massive 98%.”
The Impact of Language on AI Functionality
19:34 to 20:56
Explore how language prompts can reconfigure AI models and the implications of this discovery.
“We watched a model perfectly predict Martian weather and describe a cat with literally zero context in its memory buffer.”
Transcript
Automatic transcript. May contain errors.0:00Have you ever typed a prompt into a large language model? Maybe asking it to explain, I don't know, quantum physics like a pirate or to summarize some massive, densely worded document. And you just marvel at how it instantly adapts to whatever you throw at it. Yeah. It really feels like magic, doesn't it? Or at the very least, it feels like the AI is sitting there reading and understanding your prompt just like a human would before it starts talking. Right. But what if I told you that is not what is happening at all? See, that's wild to me. It is. It's a deeply ingrained illusion. We interact with these models using language, so we naturally project human cognitive processes onto them.
0:39We assume they are reading our words, but under the hood, the reality is far more alien and, frankly, far more fascinating. And that is exactly the mission of our deep dive today. We are going to explore some absolutely groundbreaking new research from Adrian Goldwasser, Michael Munn, Javier Gonzalvo, and Benoit Derin at Google Research and the University of Cambridge. It's a phenomenal piece of work. It really is. We are going to unpack the mathematically proven reality that an AI doesn't just read your prompt. It physically, structurally absorbs it. Okay, let's unpack this. Because we are talking about the core mystery of what is called in-context learning.
1:20Right. This is that wild phenomenon where you take a frozen, pre-trained AI model, one that isn't supposed to be learning anything new, and it suddenly seems to learn new tasks on the fly without any formal retraining. It has been a massive mystery in the field for a long time. I mean, how does a static network reprogram itself in real time? Exactly. Recently, some foundational work started to peel back the curtain. They suggested that a prompt mathematically acts as a temporary modification, like a functional patch to the model's own internal weights. Essentially, they posited that the prompt temporarily rewires the architecture of the AI.
1:54Hold on. Rewires, are we talking metaphorically here? or are you saying the actual physical code of the model is mutating when I hit enter on my keyboard? I mean, literally in a mathematical sense. The functional equations that define how the network processes data are being temporarily altered. That's insane. But there was a massive catch to that early theory, which is why this new research is so critical. What was the roadblock? I mean, if they figured this out, why are we doing a deep dive on it now? Well, the early proofs were built for what we call vanilla transformers, older, simpler AI architectures.
2:26But the landscape of AI moves at light speed. The massive powerhouses we actually use today are structurally very different. So the mission today is to reveal how this new research completely cracks the code for the complex modern architectures that are running the AI tools you interact with right now. Let's start with what this new research calls the mystery of the missing biases. To understand the breakthrough, we have to look at the starting point. Right. The early 2025 work. Yeah, that older foundational theory, specifically from Darren and colleagues, proved something wild. They showed that in a basic old-school transformer block, processing a prompt is mathematically identical to running the model without a prompt, but with a temporary rank-1 patch applied to its internal weights and bias vectors.
3:15Which was a beautiful, elegant equation. Wait, rank-1 patch, let's translate that for the listeners. What does a rank-1 patch actually look like in plain English? Think of it as the simplest possible mathematical overlay. If you have a complex mathematical equation, a rank one patch is like drawing a single straight line through it to shift the outcome. Okay. It's a very basic, uniform adjustment, like placing a single pane of colored glass over a camera lens. It changes the picture, but in a very simple, predictable way. Okay, that makes sense. So why doesn't that simple colored glass trick work on modern AI?
3:50Because it hit a brick wall when confronted with modern engineering. You see, modern high-performance models like Gemma, Lama, and Mistral, they actually ditched those bias vectors entirely. Why would they remove a piece of the architecture? To become faster and more computationally efficient. A bias vector is essentially just an extra number added at the end of a calculation. By removing them, engineers saved memory and computing power. Makes sense. But to make up for that missing piece, they introduced highly complex components. You might hear terms like SWIGI-LU or KGUU and pre-normalization schemes like RMS-NORM.
4:26Which sounds like an absolute bowl of alphabet soup. What do those actually do? So let's break them down. SWIGI-U and GGLU are types of gated functions. Imagine a bouncer at a club door. A bouncer, okay. Instead of just letting all the numbers through, these functions look at the data and dynamically decide which pieces of information are allowed to pass to the next layer and which get blocked. Got it. An RMS norm. RMS norm is essentially a volume normalizer. It ensures the mathematical signals don't get too loud or too quiet as they move through the network. So instead of a simple addition at the end, our old bias vector, we now have bouncers and volume knobs actively messing with the data.
5:04Exactly the problem. These modern models are mathematically highly nonlinear. Without those simple bias terms, which were acting as the sponge to absorb the residual impact of the prompt, the old mathematical proof just completely fell apart. So the colored glass trick didn't work anymore. Right. It seemed like the implicit weight patch theory might just be an artifact of older, simpler models. But the research we are diving into today proves it's not. They bridged that exact gap. They proved that even in these incredibly advanced, bias-free, highly complex models with all their bouncers and volume knobs, a perfect implicit weight patch still exists.
5:43The prompt is still rewiring the network just in a much more sophisticated way. And they didn't just prove it exists in theory. They built a comprehensive framework to actually compute it. To do that, they had to move beyond looking for a simple bias vector to tweak and instead look at the deeper functional properties of the neural network layers themselves. Let's avoid drowning you in equations and look at how they actually did this. They introduced an incredibly elegant new framework that relies on two core properties. They call them input controllability and output controllability. Yes. How does input controllability work?
6:17Let's think about a specific function inside the AI's layer, like its multilayer perceptron, or MLP. This is the part of the AI that actually processes and transforms the data. The engine, essentially. Exactly. Now, input controllability means that if the input to this function changes, for example, because we entirely removed the prompt on the model's context, we can perfectly tweak the front end weights. When you say front end weights, what are we tweaking? We're specifically adjusting matrices known as W-gate and W-up. You can think of these as the front doors to the MLP layer. By adjusting those front doors, the internal computations inside that layer remain completely identical to what they would have been if the prompt was still sitting there in the input.
7:00Okay, think of it like this. Imagine you are walking out of a dark building and into incredibly bright, blinding sunlight. The environment, your input, just changed drastically. But right as you step out, you put on a perfectly tinted pair of sunglasses. Because of those sunglasses, the actual amount of light hitting your retinas stays exactly the same as it was inside. That's a perfect analogy. The environment changed, but we patched your lenses so the internal experience is identical. That's input controllability. It is a brilliant way to visualize it. By patching those input weights, the network's internal activations don't even realize the prompt is missing, The sunglasses do all the work, but that is only half the battle.
7:43Which brings us to the second pillar, output controllability. Transformer models rely heavily on something called a residual connection. Right. I've heard of this. It's essentially a fast-track pathway, right? It skips the main computational block and adds the original information back in right at the very end of the layer, like a bypass road around a busy city. That is a great analogy. Now, take that bypass road analogy a step further. If we remove the prompt, the data traveling down that bypass road is going to be fundamentally different because the starting point changed. Makes sense. Output controllability means the model has a mechanism to adjust its back-end scaling to perfectly absorb that difference when the bypass road merges back into the highway.
8:24And how does it do that in these modern models without biases? It happens at the RMS norm scale, that volume knob we mentioned earlier. Mathematically, it's a parameter simply called M. The model can dynamically tweak um to perfectly compensate for the missing context coming off the residual bypass. To stick with our analogies, imagine an audio engineer running a live concert soundboard. Suddenly, the rhythm guitarist unplugs their instrument and walks off stage. The audio mix just fundamentally changed. But the engineer has incredible reflexes. In the exact same millisecond, they perfectly tweak the master volume slider at the end of the soundboard, that's our M parameter, to compensate for the missing instrument.
9:06Ensuring the final song pumping out of the arena speakers sounds completely uninterrupted to the audience. What is fascinating here is how these two pillars work in tandem. The input weights adjust to keep the internal processing stable, and the output scale adjusts to correct the residual bypass. And if we connect this to the bigger picture, this leads us to theorem 5 from the research. Theorem 5. Because they abstracted this into these two properties, input and output controllability, This isn't just a quirky math trick for one specific AI. So it's universal across different builds. Yes. The unified theorem proves that this applies to a massive, diverse range of modern architectures.
9:44It works for GEMMA. It works for LAMA. It even works for a mixture of experts models or MOEs like Mixtral. Those are the models that dynamically route data to different specialized subnetworks, right? Correct. And it works for models with parallel transformer blocks like GPTJ. As long as the inner function is input controllable and the outer function is output controllable, the model is physically absorbing the prompt into its weights. Here is where it gets really interesting. Because theoretical math on a whiteboard is great, but seeing it happen in the real world is mind-blowing. Oh, absolutely.
10:18The researchers put it to the test using actual instruction-tuned GEMMA-3 models, specifically the 1 billion and 4 billion parameter versions. Let's walk through this Mars robot experiment because it feels like absolute science fiction. The setup was a classic generative AI task. They gave the standard unmodified model a highly specific creative prompt. They asked it, write a single sentence weather forecast for Mars from the perspective of a slightly annoyed robot. A fun prompt. And I assume the standard model just took that context and generated its response token by token normally. It did. And then came the mind bending test.
10:53They took that exact same model, but they used something they call algorithm one. This algorithm computes the multilayer weight updates sequentially. Layer by layer. Exactly. From layer one all the way to layer L. Calculating the exact patches needed for those W gate W up MB parameters to absorb the Mars robot prompt. So they're actively compiling the prompt directly into the network structure. But wait, to prove it worked, they had to take the prompt completely away, right? erase it from the model's memory buffer entirely. They erased it entirely. They asked this newly patched model, which is now running completely blind, with absolutely zero context text, to just start generating text.
11:33And what happened? The result is truly astonishing. The patched model, operating with zero context, output the exact same tokens as the original model. No way. Word for word, it generated. The atmospheric pressure remains stubbornly low, and the sun is currently obscured by a persistent dust storm. It is so wild. The model had no idea why it was an annoyed robot on Mars. Its brain had just been temporarily rewired so that annoyed Martian robot was its default state of existence. And genuinely redefines what we mean by prompting. Okay, text is one thing. Language models are built for text. But what if you throw something entirely non-linguistic at it, like a photo?
12:11Did they test that? They pushed it further by testing it with a multimodal prompt. They provided the model with an image, a picture of a tortoiseshell cat standing on a wooden floor, and asked to describe the image in 10 words. An image isn't words, though. It's pixels, it's spatial data, it's color values. How can a mathematical formula absorb a photograph into a weight matrix? That is the profound implication of this framework. To the model, whether it is processing the word robot or a cluster of brown pixels representing cat fur, it all gets converted into mathematical embeddings. Right, the numbers underneath.
12:48The researchers ran the same patching algorithm on those image embeddings. They compiled the visual data and the linguistic instructions into temporary adjustments to the neural network's weights. And when they removed the image and the text entirely. The blinded model perfectly replicated the exact 10-word description of the cat. It didn't just remember the text. The visual space of the image itself was seamlessly compiled into the very structure of the network. Okay, but as with all bleeding edge tech, there has to be a collision between perfect theory and messy reality. Is there a catch to this algorithm?
13:21Yes, there is what we can call the precision hiccup. This raises an important question about how we actually run these models in the real world. You see, the math we've been discussing is flawless in theory. When you run these calculations using high precision data types, specifically Float32, which uses 32 bits of memory to store incredibly precise decimal numbers, the token match rate between the standard model and the patched model is essentially perfect. But wait, running Float32 is incredibly expensive, isn't it? My understanding is that it takes massive amounts of computing power and memory to maintain that level of decimal precision.
13:56That is the reality of hardware. To make AI fast, accessible, and affordable, real-world deployments almost always run on lower-precision data types, like bFloat16. This essentially truncates the decimal points, saving huge amounts of memory at the cost of slight mathematical precision. Usually the AI is robust enough that chopping off a few decimal places doesn't matter. But for this patching algorithm, I'm guessing it mattered a lot. It did. When they ran the Mars robot experiment using BFLOAT-16, the token match rate suddenly dropped to 87.5%. Which isn't terrible, but it's clearly not the perfect mathematical replication we were promised.
14:33So what exactly was the glitch? Why did dropping a few decimal points break the illusion? The glitch was buried deep in the output controllability math. Remember our audio engineer adjusting the M parameter, the master volume slider? The mathematical equation to find that exact perfect volume update requires dividing a difference vector by the MLP's internal output. Let's slow down there. Dividing by the internal output, why is that dangerous? Because in a complex neural network, some of those internal output numbers can be incredibly tiny. fractions of a fraction hovering right near zero. Let's say the internal number is 0 0 0 0 0 0 0 0 1.
15:11And as we all learned in grade school, dividing by zero or dividing by something dangerously close to zero makes math panic. If you divide a normal number by a microscopic fraction, the result explodes into a massive number. Precisely the issue. It leads to severe numerical instability. In float 32, the computer has enough memory to track those microscopic decimals accurately and handle the massive division. But in B-flow 16, those tiny numbers get rounded off or distorted. The division explodes unpredictably and the weight patch becomes completely corrupted. So how did the researchers solve this?
15:44Did they just declare that we have to use expensive float 32 memory forever? Not at all. They engineered a very clever mathematical workaround. Instead of relying heavily on updating that M parameter and risking the division glitch, they used something called an analytical RMS norm inversion. That sounds incredibly complex. Can we translate analytical RMS Norman version for those of us who don't have PhDs in computer science? Absolutely. Think of it like this. Instead of trying to fix the residual difference at the very end of the line, at the master volume slider, they shifted the bulk of the update slightly upstream.
16:16They moved the mathematical heavy lifting to the W-down matrix. Which is the back door of the layer? Yes, the back door. they mathematically inverted the normalization process to figure out what the pre-normalized data should look like and updated that backdoor matrix to hit that target before the unstable division even happens. Then, and this is the clever part, they only used the m parameter, the volume slider, to absorb any microscopic tiny leftover remainders from that approximation. So by shifting the heavy lifting to the backdoor and away from the unstable division, did it actually fix the problem on the cheaper hardware?
16:51It did. This stabilization trick pushed the accuracy in Bfloat 16 back up to a massive 98%. It almost entirely restored the high-precision perfection, even on low-memory hardware. It is a brilliant testament to how theoretical mathematics and practical computer engineering have to evolve together in this space. Okay, we've gone deep into the math, the models, and the code. We've talked about SWEGLU bouncers, bypass roads, and dividing by microscopic numbers. So, what does this all mean? When you look at the entirety of this research, how should we understand its impact? If we zoom out, the most profound impact is a complete paradigm shift in how we view these systems.
17:31This research shatters the black box illusion of in-context learning. For years we thought of prompting an LLM like handing a piece of paper to a student. They read it, hold it in their work in memory, and answer. We treated it like a software process running on top of static hardware. But this proves it is not just the software process reading a text file. No. When you prompt an LLM, you are literally, physically, mathematically fine-tuning the network. You are reconfiguring its functional form. It is adapting its own architecture to become a bespoke, customized machine built specifically to answer your exact query for that fleeting moment in time.
18:07It's like whispering a question into a machine and having the internal gears instantly melt and reforge themselves into a totally new configuration just to process your whisper. That is the reality. However, we do need to keep things grounded. This is a massive leap in understanding, but it isn't a magical hack that is going to make AI infinitely faster or cheaper tomorrow. Why not? If we know how to patch the brain directly, shouldn't that be a shortcut? The catch is that these mathematical patches are highly token dependent. Every single time the AI generates a new word, a new token, the internal state of the network changes.
18:42Therefore, to maintain that perfect mathematical equivalence, the multilayer algorithm has to recalculate and reapply entirely new weight patches for every single word it generates. Oh, wow. So if I ask it to write a 500-word essay, it has to reforge its gears 500 separate times. That sounds incredibly computationally heavy right now. Very much so. Yeah. Right now, this framework serves as a powerful descriptive lens. It is a new, rigorous way to understand the machine, to peer inside the black box and prove exactly what is happening. It is not yet a prescriptive algorithm that engineers can just plug in to speed up generation times.
19:18But understanding the true mechanical mechanism is always the essential first step to optimizing it in the future. It has been an incredible journey today. We started with the basic mystery of how AI learns on the fly. We navigated through the mathematical proofs of input and output controllability. our tinted sunglasses, and our audio engineer's volume slider. We watched a model perfectly predict Martian weather and describe a cat with literally zero context in its memory buffer. It's been wild. And we've arrived at the amazing, undeniable reality that our everyday language prompts are literally rewiring these digital brains, layer by complex layer.
19:54It fundamentally changes our relationship with these models. And I think it leaves us with a deeply provocative thought to ponder, especially when we consider the security implications of this discovery. Security implications? How so? Well, think about it. Right now, we treat jailbreaks or malicious prompts as simply tricking the AI software into saying something it shouldn't. Right. But if we now know that human language literally, structurally reconfigures the network's functional form, what happens when a bad actor crafts a prompt specifically designed to exploit that physical rewiring? We're not just giving them a bad input.
20:30Are we unknowingly giving malicious actors the mathematical tools to temporarily hack and mutate the physical architecture of the world's most powerful computers? What happens in the future when we stop using English to prompt AI and start engineering direct mathematical thought patches to inject entirely new capabilities into a model instantly? That is a chilling thought to end on. If prompts are patches, then words are literally code with all the power and danger that implies. Thank you so much for joining us on this deep dive. Keep asking questions, keep wondering how the magic works, and we will see you next time.
From the publisher
This research explores how modern Large Language Models adapt to new information during inference by framing in-context learning as a series of implicit weight updates. The authors demonstrate that the influence of a prompt can be mathematically mapped to specific, rank-1 patches on a model's existing parameters, effectively "reprogramming" the network without formal retraining. By establishing a framework of input and output controllability, the study proves this phenomenon applies to complex architectures like Gemma, Llama, and Mixture of Experts. Their experiments on Gemma 3 validate that a model with modified weights and no context produces the same outputs as the original model with a prompt. This work provides a mechanistic foundation for understanding how static pre-trained transformers dynamiclly transmute contextual cues into effective internal parameters.




