In short
In-context learning in LLMs violates exchangeability: shuffling the order of the same demonstrations can change predictions. The episode argues this is “hallucinating geometry” caused by positional encodings/ROPE, not inability to do statistics.
Key claims
LLMs are Bayesian “in expectation” (averaging over permutations cancels positional noise), but not “in realization” (single prompt order can be biased).
Notable examples
movie-review sentiment flips after swapping example order; retrieval-augmented generation where confidence drops when the key document is pasted third.
Guests
none named in the transcript (hosts discuss sources; no guest bios provided).
Guest backgrounds
not applicable. Key evidence/examples: martingale diagnostics show nonzero “martingale gap”; “Bernoulli microscope” (0/1 sequences) shows error decay with context length and near-perfect statistical error without positional encodings (tiny transformer), but error returns when positional encodings are re-enabled. Period-64 error oscillations align with attention-head dimension and ROPE frequencies; padding shifts the wave. Fix: permutation averaging (k shuffles reduces error variance ~1/sqrt(k)); also “header” prompts (explicitly constrain outputs) reduce extraction failure from ~99% to 0%.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOExploring In-Context Learning Glitches
0:45 to 3:00
Discussion on the glitches in in-context learning and how they affect LLMs.
“And we'll see that it's not just an abstract math problem.”
The Concept of Exchangeability
3:00 to 5:10
Deep dive into the concept of exchangeability in Bayesian statistics.
“It does sound incredibly dense, but it's just a test for statistical consistency.”
Positional Encodings and Their Implications
5:10 to 8:00
Understanding how positional encodings affect LLMs' reasoning.
“To understand language, the model must be obsessed with order.”
Bernoulli Microscope Experiment
8:00 to 11:20
Examination of an experiment that isolates LLMs' statistical reasoning ability.
“The very mechanism that allows the model to speak human language fluently is exactly what prevents it from thinking like a pure statistician.”
Shuffling as a Solution
11:20 to 13:30
Introduction to shuffling prompts to improve LLM accuracy.
“In expectation, which basically means on average, right?”
Real-World Applications of Shuffling
13:30 to 14:01
Practical implications of shuffling for real-world AI applications.
“That's an absolute nightmare for reliability and production.”
Improving Model Accuracy with Shuffling
14:01 to 14:38
Learn how to enhance AI model accuracy by using shuffled prompts.
“So, practically speaking, if you are out there building a medical diagnosis bot or a legal research assistant, you shouldn't just paste your five retrieved documents into the prompt and hope for the best.”
The Impact of Headers on Model Performance
14:39 to 16:28
Discover how adding headers can significantly improve model outputs.
“Now, before we wrap this up, there was one other highly practical tip that came out of this deep dive.”
Understanding LLMs: Geometry and Structure
16:30 to 18:34
Explore how the geometric nature of LLMs affects their reasoning and outputs.
“We started this deep dive talking about a glitch, a failure in the model.”
Adopting a New Mindset for Using LLMs
18:35 to 18:51
Shift your approach to LLMs by focusing on averages rather than single outputs.
“It's the only way to be statistically sound.”
Transcript
Automatic transcript. May contain errors.0:00Welcome back to the Deep Dive. We're doing something a little different today. Yeah, a bit of a pivot from the usual. Right. Because usually, you know, we look at the magic of AI, the crazy things it could do, how it's changing everything. But today we are going to look at, well, the cracks in the magic. The cracks are honestly where it gets the most interesting. Exactly. So we're talking about a very specific, very weird glitch in in-context learning. And it essentially proves that these models are hallucinating. Geometry, which sounds totally absurd. It does. Hallucinating geometry. But we're diving into some fascinating sources today that unpack this exact phenomenon.
0:39And our mission for this deep dive is to figure out why a machine built to process data stumbles over something as simple as the order of a list. Right. And we'll see that it's not just an abstract math problem. It's actually one of the most practical problems in LLM deployment right now. It really challenges the whole idea of whether these models are reasoning at all. Yeah, because the promise of in-context learning, the whole point of it, is that I can give the model a few examples. Let's say three movie reviews with their sentiment labels, positive, negative, whatever. Right, your demonstrations.
1:13And the model just learns the pattern on the fly. I don't have to update its weights. It just gets it. It feels like magic. It infers the task right from the prompt. But here's the problem, the glitch. If I take those exact same three movie reviews, the same data, same labels, and I just shuffle the order, I swap review number one with review number three. The model's prediction for a brand new review often changes. Sometimes significantly. Which is, I mean, it's baffling. It really is. It's terrifying if you're building a business on this. Imagine if you went to a bank for a loan, and the loan officer looked at your W-2, your credit score, your bank statement, and they said, yep, you're approved.
1:53Okay. But then they accidentally knocked the papers off their desk, picked them up in a different order, and said, actually, no, you're denied. You'd think they were completely insane. Right. You'd say they aren't reasoning rationally. And in statistics, there is actually a formal name for this exact rule that the LLM is breaking. It's called exchangeability. Exchangeability. Okay. Let's unpack that for the listeners. Sure. So it's a core concept in Bayesian statistics. It basically says that if you have a bag of independent pieces of evidence, the order in which you pull them out shouldn't change your final conclusion.
2:27Like flipping a coin. Exactly. That's a classic example. If I tell you I flipped a coin three times and got heads, heads, tails. Okay. That is the exact same data set as tails, heads, heads. The order of arrival doesn't change the underlying probability that the coin might be rigged. So if an AI is truly reasoning about the data I give it, it should treat my examples as a set, a bag of facts. Right. But instead, it's treating them as a sequence. And the sources show this happens when you run what they call martingale diagnostics. Martingale diagnostics, yeah. Which sounds like a part I need to get replaced on my transmission.
3:00It does sound incredibly dense, but it's just a test for statistical consistency. A martingale is a mathematical process where your best guess for the future is just whatever your current value is. Okay. In this context, it means that as the model sees more data, its confidence should naturally converge on the truth. It shouldn't be jumping around wildly just because we swapped the order of the inputs it just saw. But the data in these studies shows they're jumping around wildly. They are. And we aren't just talking about old models, even the big state-of-the-art deployments. when researchers measure the martingale gap.
3:36Which is the difference between a perfect statistical engine and the LLM. Exactly. When you measure that gap, it's not zero. The models are highly sensitive to permutation. So this is usually the point where the critics jump in and say, look, see, these things are just stochastic parrots. They don't understand logic. They just predict the next word. But that's where the steep dive gets really interesting, because the evidence suggests the model isn't stupid. It's not that it fundamentally can't do the statistics. It's that it's being sabotaged by its own anatomy. The architecture is the culprit.
4:07Specifically, a mechanism called positional encodings. Okay, let's get into the weeds here. Okay. Because this is the core of it. I know that transformers, the T in GPT, they process data in parallel. Right. They don't read word by word like you or I do. They ingest the whole context window all at once. Correct. And that is exactly what makes them so incredibly fast to train. But there is a downside to swallowing the whole paragraph at once. Which is? Without some extra help, the model has absolutely no idea which word came first. To the raw attention mechanism, a sentence like, the dog bit the man.
4:44Looks exactly the same as the man bit the dog. Exactly. It's just a bag of words. Man, dog, bit. Which is a disaster for speaking English. You really need to know who did the biting. Order is meaning in language. So the engineers originally solved this by adding positional encodings. They basically stamp every single token with a time signature. Like a little sticky note saying, you are word number one. You are word number two, and so on. They inject geometry into the math. I see the conflicts coming. Yeah, it's a fundamental tension. To understand language, the model must be obsessed with order.
5:17But to understand statistics, like our list of independent coin flips or movie reviews. The model must ignore order. And it can't just turn the feature off. It cannot. It is structurally incapable of ignoring position. So when you give it a list of medical records or facts, it is trying to read them like a novel. It's constantly looking for a narrative arc in a list of random data points. It's like an English teacher analyzing a phone book. Like, why is Smith placed before Jones? What is the author trying to say here? That is the perfect analogy. It is over-interpreting the geometry of the prompt.
5:54And that right there is what brace exchangeability. No, what's really cool is the researchers didn't just guess this was happening. They proved it with a controlled setup. They called it the Bernoulli microscope. I love that name. It is a great name. Walk us through it. How do you put a massive LLM under a microscope? Well, natural language is messy, right? Words have double meanings, complex grammar, weird syntax. Sarcasm. Sarcasm, exactly. If you want to see if the model can do pure statistical reasoning, you have to strip all of that away. So they fed the models binary sequences, just zeros and ones.
6:27Literally just 0, 1, 1, 0, and asked the model to predict the next digit. Exactly. Pure pattern matching. No grammar. No vocabulary. Just raw stats. What did they see through this microscope? Two main things. First, they saw what they call decay, meaning as the context length gets longer, as you show the model more zeros and ones, the Martingale gap actually gets smaller. Okay, that makes intuitive sense. If I show you a thousand coin flips, eventually you're going to figure out the bias, even if the order throws you off at first. Right. The sheer volume of data eventually overpowers the positional noise.
7:02But the second finding was the real smoking gun. The ablation study. Yes, an ablation study is where you surgically remove parts of the model to see what happens to its behavior. They trained a tiny transformer from scratch that had no positional encodings whatsoever. So a model with zero concept of time or order? None. To this tiny model, everything is just a jumbled bag. And do you know what happened to the statistical error? I'm guessing it vanished. It collapsed to 10 to the negative 16th power. Which is basically zero. It is effectively perfect. It became a perfect Bayesian reasoning engine.
7:34But I assume a model with no sense of order couldn't write a poem. Oh, couldn't write, hello world. It was completely linguistically illiterate, but statistically perfect. Oh, that's amazing. And then they took that exact same architecture and turned positional encodings back on. And the error came back. Instantly. The error jumped right back up to around 10 to the negative 6th. That is a massive degradation in statistical purity. So that proves it definitively. The very mechanism that allows the model to speak human language fluently is exactly what prevents it from thinking like a pure statistician.
8:09It's a necessary trade-off. In this architecture, you really can't have one without the other. But wait, it gets even weirder. This is the part of the deep dive that really blew my mind. We aren't just talking about random chaotic errors here. When the researchers looked at where the errors were happening, they found a rhythm. A heartbeat. Yeah, a heartbeat in the noise. This is where we see the literal ghost in the machine. If you plot the error rate against the token position in the context window, it's not static noise. It forms a wave. A sine wave. Specifically, what they identified as a period 64 signal.
8:44The error oscillates predictably every 64 tokens. And 64 isn't just a random coincidence here? Not at all. In the models they were testing, 64 is the exact dimension of the attention heads. It also aligns perfectly with the specific frequencies used in ROPE. ROPE. Okay, rotary positional embeddings. That's the standard way modern models like Lama and others handle position now, right? Yes. And without getting too bogged down in the linear algebra of it all, rho p works by rotating the vector representation of a word in a high-dimensional space, and it uses sine and cosine functions to define that rotation.
9:19So when we see a period 64 wave in the model's error rate, we are seeing the literal mathematics of the architecture leaking into the model's prediction. The model isn't actively thinking about the order and getting confused. It's getting physically tripped up because the rotation of the vector at position 64 aligns poorly with the vector at position 128. It's a geometry problem masquerading as a reasoning problem. Exactly. That is just wild. It's like finding out your calculator gives you the wrong answer every 64th time you press a button simply because of the way the physical circuit board is soldered.
9:55That's a remarkably close analogy to the truth. And they actually proved this was geometric, not semantic, by doing a robustness check. Right, the padding experiment. Yeah, they added neutral tokens to the very start of the prompt. Just junk padding. Meaningless filler words. Right. It doesn't change the underlying data or the logic of the prompt at all, but what it does is shift the index of every subsequent token. So token number 10 gets pushed to become token number 15. And sure enough, when they did that, the error wave shifted exactly in lockstep. Wow. So think about this. If you are listening to this, and you're a prompt engineer, and you're tearing your hair out because your prompt just isn't working, It might literally just be that your most important example happened to land on a bad frequency in the model's sine wave.
10:43It is entirely possible. You might literally be standing in a dead zone of the model's attention geometry. That is incredibly frustrating. Oh, absolutely. But also weirdly comforting, because it means the model isn't being irrational or stupid. It's just mechanical. It's deterministic. And the beauty of a deterministic problem is that we can fix it. Which brings us to the solution. How in the world do we fix a problem that is fundamentally baked into the architecture of the model itself? We have to change how we use them. We have to stop treating the model like it's going to be perfect on a single try.
11:16The research proposes we treat it as Bayesian in expectation. In expectation, which basically means on average, right? Exactly. Any single run, what they call a single realization, might be heavily biased by this geometric noise. But if you take the average of all possible orderings, the noise naturally cancels itself out. So the fix is just shuffling. Permutation averaging. It's wonderfully simple in practice. You have your examples, you shuffle them, you run the prompt, you get a probability, then you shuffle them again, run it again. And then you just average the results together. Correct. And the math behind why this works is beautiful.
11:54The standard deviation of the error drops following a specific trend. 1 over the square root of k, where k is the number of shuffles. 1 over the square root of k. That means they're diminishing returns. Right. You don't need to run it 1 ,000 times to get the benefit. The biggest mathematical gains come from the first few shuffles. If you run it 5, 10, maybe 20 times, you knock down a massive amount of that variance. You're essentially smoothing out that period 64 heartbeat. You're forcing the model to ignore the geometry by showing it all the possible geometries. It effectively turns a flawed order-sensitive predictor into a highly reliable Bayesian one.
12:32Now, I know what some listeners are probably thinking right now. They're saying, I don't care about predicting binary digits, zeros and ones. I want to know if this actually matters for my real-world applications. And the sources confirm that it absolutely does. They move from binary to what they call categorical sequences, ABC patterns. And they found the exact same gap and the exact same fix works. But the real kicker for me was the evidence-grounded QA testing. Oh, yeah. The RADG use case, retrieval augmented generation. Which is arguably the most common enterprise use case for AI today. By far.
13:06You give the model a stack of documents, say five news articles or five legal precedents, and you ask a question based on them. And obviously, logically, the order you paste those documents in shouldn't matter at all. It shouldn't, but we know it does. If you put the crucial smoking gun document first, the model might be 90 percent confident in its answer. But if you put that same document third in the prompt, confidence might drop to 60 percent. That's an absolute nightmare for reliability and production. It is. But when the researchers applied permutation averaging, literally just shuffling the order of the evidence documents and averaging the model's output probabilities, the accuracy went up consistently across the board.
13:45They call this the Jensen gain, right? Yeah, he named after Jensen's inequality. Basically, by averaging the predictions over a mixture of random permutations, the model's cross-entropy goes down. Cross-entropy being a measure of how surprised the model is by the truth. Exactly. Lower cross-entropy means it became more confident and more correct. So, practically speaking, if you are out there building a medical diagnosis bot or a legal research assistant, you shouldn't just paste your five retrieved documents into the prompt and hope for the best. you should actually be running that exact same prompt a dozen times with the document shuffled.
14:19If accuracy is your absolute top priority, yes. It increases your inference cost, obviously, because you're making more API calls, but it mathematically guarantees a better, more statistically sound answer. It's the difference between asking one single witness what happened and polling the entire room. That's a great way to put it. Now, before we wrap this up, there was one other highly practical tip that came out of this deep dive. and wasn't about shuffling at all it was about headers oh this is a gem it's such a small detail but the impact they measured was massive tell us about the header versus headerless experiment so when they were testing the binary predictions they tried two different prompt styles to see how the model reacted the first was headerless they just gave the raw list zero one zero one and let the model utter complete the rest which is how honestly most of us cromped every day we just dropped the data in right but the second way was header.
15:12They explicitly wrote a constraint at the very top of the prompt. They said, respond with exactly one of zero, one. They rigidly define the universe of possible answers. Now, usually we do that as developers, just to make it easier to parse the output with our code later. That's what everyone thought. But it turns out it fundamentally changes the model's internal reasoning process. With the headerless prompts, the extraction failure rate was nearly 99%. Wait, 99 % failure? Yes. The model would start rambling or it would lose the count entirely or just start outputting weird, irrelevant text. It got completely lost in the geometry.
15:50But with the header, the failure rate dropped to 0.0%. From 99 % to 0, just by adding a menu of valid options at the top. It strongly suggests that while the model is incredibly sensitive to geometry, it is also desperate for structure. When you constrain the output space explicitly like that, you help the attention mechanism lock onto the actual task. It seems to override or mitigate a lot of that positional confusion. That is massively actionable advice for everyone listening. Always header your prompts. Don't just imply the format you want. Enforce it explicitly. It stabilizes the model. It's not just for your rejects parser.
16:27It actually helps the model's brain focus. So let's zoom out and look at the big picture here. We started this deep dive talking about a glitch, a failure in the model. But the more we talk, the more it feels like we're just finally seeing the reality of how these things actually work. I think that's exactly the right takeaway. We tend to anthropomorphize these models. We want them to be pure reasoning engines that think like we do. But deep down under the hood, they are geometric and they're just processing shapes in high dimensional space. And their shapes have quirks. That period 64 wave isn't a bug in somebody's quote.
17:06It's a natural ripple in the math. The fact that order matters isn't a sign of artificial stupidity. It's literally the price we pay for them to have fluent language. It really reframes the whole conversation we always have about hallucinations. It really does. I mean, think about it. How many times have you been working with an LLM? He gives you a blatantly wrong answer and you just assume, oh, it doesn't know the fact. It's training data is bad. But Maybe it knew the fact perfectly well. Maybe you just happened to paste your bullet points in an order that hit a dead zone in its attention mechanism.
17:35Or you hit a frequency where the positional encodings just completely drowned out the actual data. It makes prompt engineering feel a lot less like an art form and a lot more like radio tuning. You know, we're just turning the dial, trying to find a clear signal through the geometric static. And this research strongly suggests that maybe we should stop trying to find the one perfect prompt. because geometrically the perfect prompt might not even exist due to these fluctuations. Instead, we should be looking for the average prompt. Robustness over perfection. Exactly. Don't trust the single realization.
18:10Trust the expectation. It's a fascinating shift in mindset. We usually think of computers as exact rigid machines. Input A always equals output B. But with LLMs, input A is actually a probability distribution. And we need to start treating them that way in production. If you want the truth from an LLM, don't ask it once. Ask it often. Shuffle the deck every time and take the average. It's a little more work, a little more compute. But the math says it's absolutely worth it. It's the only way to be statistically sound. Well, this has certainly changed how I'm going to use these models tomorrow.
18:42I'm going to be doing a whole lot more shuffling. Just don't get too dizzy. Right. Well, as always, thank you for guiding us through the complexities of the machine. My pleasure. Always fun. And thank you to everyone listening. I want to leave you with a thought. The next time you feel like your AI is gaslighting you or hallucinating facts you know it knows, remember, it might just be the sine waves talking. Try swapping your bullet points around. You might just find the truth hiding in the mix. This has been The Deep Dive. We'll see you next time.
From the publisher
This research explores the discrepancy between transformer in-context learning and Bayesian inference, arguing that models are Bayesian in expectation rather than through every individual realization. While previous studies used martingale diagnostics to question the Bayesian nature of these models, this paper identifies positional encodings as the primary factor that breaks the required exchangeability. By accounting for how architectural design prioritizes sequence order, the authors prove that transformers still achieve near-optimal compression and information-theoretic efficiency when performance is averaged across different orderings. Empirical tests on black-box LLMs and controlled ablations demonstrate that order-induced variance exists but predictably decays as context length increases. Ultimately, the study suggests permutation averaging as a practical and effective method for reducing uncertainty and improving the reliability of model outputs in tasks with exchangeable data.




