A Unifying View of Attention Sinks: Two Algorithms, Two Solutions

16 Jun 2026 · 23 min · 8 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Vision-transformer “attention sinks” show near-vertical stripes in attention maps. The episode argues these stripes aren’t one bug but two coexisting learned algorithms caused by the softmax “forced choice” constraint.

Guests/backgrounds

No guest names or bios are provided in the transcript.

Key claims

(1) “Adaptive NOP” (trash can) sinks dump ~99% attention to a token while driving that token’s value norm near zero, effectively performing a no-op; achieving this costs via spectral spikes or massive activations. (2) “Broadcast” sinks (megaphones) also look like stripes but carry normal payloads and distribute the same global vector to many tokens, producing a rank-one update. Both appear in the same model: early/mid layers use trash-can sinks (often on CLS), deep layers use broadcast sinks (often on random patch tokens), with some heads acting as sinks for ~80% of images.

Notable examples

DinoV2 (large/giant) and open CLIP models; ADE20K segmentation improves mIoU from ~0.166 to ~0.2000 when combining gated attention with register tokens. Gated attention alone leaves deep-layer corruption; register tokens alone get hijacked into trash cans.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Understanding Attention Sinks in AI

0:41 to 4:30

Exploration of the attention sink phenomenon in vision transformers.

“you shared with us about the hidden algorithms inside artificial intelligence, specifically vision transformers.”

The Trash Can Hypothesis

4:30 to 8:42

Discussion of the adaptive NOP algorithm as a workaround for attention mechanisms.

“The technical term of the research is the adaptive NOP, which stands for no operation.”

The Broadcast Hypothesis and Model Dynamics

8:42 to 14:00

Analysis of the broadcast hypothesis and the coexistence of algorithms in AI models.

“Well, here's where it gets really interesting, though.”

Understanding Attention Sinks

14:00 to 15:11

Learn about the dual roles of attention sinks in AI models and their implications.

“Wait, they become dedicated sanitation workers or dedicated broadcasters?”

Exploring Fixes for Attention Issues

15:12 to 18:08

Discover the common fixes used to address attention sink problems in AI models.

“Which perfectly explains the mitigation problem engineers have been facing.”

Emergent Behavior in AI

18:09 to 18:42

Explore how AI repurposes intended design features, leading to unexpected outcomes.

“So the answer is you have to do both simultaneously.”

Impact on Performance and Accuracy

18:43 to 21:00

Understand the implications of separating trash from megaphones for AI performance.

“It translates directly to performance, specifically on highly complex visual tasks.”

Wrapping Up the Insights

21:01 to 22:36

Reflect on the key discoveries regarding AI architecture and its adaptability.

“We started with a visual mystery, a vertical stripe on an attention map that looked like the AI was basically freezing up and staring blankly at a single pixel.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00If you open up the brain of the world's most advanced visual AI, you honestly won't find a clean grid of flawless logic. Which, you know, you'd expect to, right? Right. You'd absolutely expect that. Yeah. I mean, you build a machine, you trace the wire from point A to point B, and it executes exactly what it was engineered to do. Exactly. But instead, inside these models, you'll find this massive, glitchy vertical stripe where the AI essentially just stops looking at the image entirely and stares blankly at a virtual wall. It really is one of the most perplexing anomalies in modern machine learning.

0:36Visually, it just looks like a complete system breakdown. Well, welcome to your custom deep dive. Today, we're unpacking that massive stack of research you shared with us about the hidden algorithms inside artificial intelligence, specifically vision transformers. Yeah, these are the powerful models that are basically learning to, you know, see and interpret the world around us. Right. And our entire mission today is to decode what that weird blank stare actually means, because it turns out it's hiding a massive secret. It really is. And to set the stage for you, we are talking about a phenomenon called an attention sink.

1:08An attention sink? Right. In a vision transformer, the AI breaks an image down into these small puzzle pieces called tokens. So a token might be like a 16 by 16 pixel patch of a dog's ear or maybe a car tire. Okay, got it. And during processing, the AI creates an attention map. This is essentially a graph showing where the model is looking at any given millisecond to gather context. And a normal attention map makes intuitive sense to a human, right? I mean, if the AI is trying to identify a dog, you expect the graph to light up around the dog's eyes, its fur, maybe its tail. Exactly, yeah. But that is not always what happens.

1:45No. No. Instead, across almost all major industry models, researchers keep finding these bizarre near-vertical stripes on the graph. Which is just so weird to look at. It is. It indicates that nearly every single token in the network has suddenly diverted its attention to one single seemingly random piece of data. Wow. Yeah, one token just becomes an absolute black hole for the model's focus. And for a long time, researchers treated these vertical stripes as basically a shared pathology, like a bug. Total. Because the papers you shared note that engineers linked these sinks to all sorts of negative outcomes, things like rank collapse, training instability, and these mysterious massive activations where the math just spikes completely off the charts.

2:29Yes, exactly. But here is the foundational question I really have to ask on behalf of the listener. Because just thinking about this logically, if 99 % of the AI's processing power is staring at one specific token, doesn't that inherently mean that token holds like the most crucial piece of information in the entire image? This raises an important question, and it really gets to the very heart of why this phenomenon was misunderstood for so long. Okay. Because the central challenge is that an attention pattern just by itself is not a mechanistic explanation of what the AI is actually doing. What do you mean?

3:03Well, just because the AI is heavily focused on a piece of data does not necessarily mean it cares about that data. Okay, let's unpack this. How can it intensely focus on something it doesn't even care about? Right, it sounds contradictory. But it all comes down to something called the softmax constraint. The softmax constraint. Yeah. These models use an attention mechanism governed by a mathematical function called softmax, And the defining rule of softmax is that it forces the AI's attention to act as a strict probability distribution. Meaning? Meaning across all the tokens it's looking at, the attention must sum exactly to one or, you know, 100%.

3:43Oh, wow. So an attention head is basically forced to place its probability mass somewhere. Yes. It mathematically cannot output a zero. That is the crux of the issue right there. It is a forced choice scenario. Okay. So if the AI is processing, say, a patch of empty blue sky and a specific attention head realizes I have absolutely no useful context to add here, it cannot simply turn off. Because it has to equal 100%. Exactly. It must point its 100 % at something. Therefore, staring intensely at a single token could mean that token is incredibly important. Or it could mean it's simply the safest place to dump attention when the AI wants to effectively do nothing.

4:20That is wild. And that leads us right into the first of the two hidden algorithms researchers discovered hiding inside this visual glitch. It does. Let's call this the trash can hypothesis. The technical term of the research is the adaptive NOP, which stands for no operation. Right, a no op. Basically, the AI's attention head has no useful information to contribute to the next layer of its thought process. But because of that soft max constraint, it still has to output something. So it rods all its attention to a null token to safely suppress its own update. It is a remarkable emergent workaround.

4:55I mean, the AI essentially invented a trash can because the engineers didn't build one into the architecture. I think of it kind of like being in a massive corporate meeting. Oh, I like this. Yeah. So the boss asks a highly technical question. And you have absolutely nothing useful to add to the conversation. But the boss is scanning the room. And you have to look somewhere, right? You can't just close your eyes. Exactly. So what do you do? you stare blankly and intensely at your notebook. Yes. You're not reading the notebook. The notebook contains no vital information. You're just using it as a prop to avoid making eye contact so you don't get call on.

5:29In this scenario, your notebook is the attention sink. That is a highly accurate way to visualize it. And interestingly, the researchers were able to identify exactly how the AI achieves this notebook staring strategy mathematically. Oh, they could see the math behind the stare. Yeah, they found a diagnostic signature. Because the AI is forced to output a vector, which is, you know, a mathematical package of data, it learns to make the value norm of that sync token nearly zero. And the value norm is essentially the payload, right? Like the actual information content being delivered. Exactly. The attention score is just the delivery mechanism, but the value norm is the package inside.

6:08So by dropping the payload to near zero, the AI can dump 99 % of its attention onto that token, and the resulting update to the entire system is effectively nothing. It's an identity operation. It successfully executes a non-action. But wait, doing nothing inside a forced choice mathematical system actually requires a huge amount of effort, doesn't it? Yeah. Like, it doesn't come for free. Not at all. The cost of doing nothing is incredibly high here. Yeah. To force what we call a hard gate, basically making sure almost all the attention goes to this one specific trash can spot and doesn't leak out to corrupt other tokens that requires extreme math.

6:45Extreme math. Yeah. The AI basically has to distort its own weights. Yeah. And it usually does this in one of two ways. The first is by creating a spectral spike. Okay. Spectral spike. Which means what exactly for the listener? So a spectral spike means the AI develops a massive singular value in its interaction matrix. If you think of the AI's attention mechanism as an antenna trying to pick up signals, a spectral spike is like tuning one specific frequency to be hypersensitive. So the slightest trigger causes it to route everything straight to the trash can. Okay, so if it doesn't build a hypersensitive trigger, its only other option mathematically would be to just scream louder than everything else to drown out the noise, right?

7:25Yeah. Is it doing that? It is, and that is actually the second method, a massive activation. Massive activation. Right. The AI gives the sync token a disproportionately huge mathematical norm. It just inflates the numbers. And to understand why this works, we have to look back at that softmax function we mentioned. The probability thing. Yeah. Softmax calculates probability using exponential functions. So it doesn't just look at the raw numbers. It looks at a base number raised to the power of those activations. Oh, so if one token's value is even just a little bit higher than the rest, The exponential math causes it to absolutely explode into a super attractor.

8:03Precise. It exponentially drowns out the other tokens. Exactly. And this explains a huge lingering mystery in machine learning. Really? Yeah. For years, researchers have been seeing these outlier massive activations in AI models where a random token suddenly spikes with gigantic numerical values. And people thought it was just a glitch. Right. People thought the models were just unstable, but it turns out that's not a bug. That is the energetic cost of the AI enforcing a reliable trash can without distorting the rest of its internal logic. That is wild. It's literally mathematically shouting at the top of its lungs just to enforce a moment of silence.

8:42It highlights how desperately the architecture needs a way to bypass useless information. Right. Well, here's where it gets really interesting, though. We established at the top that there were two different algorithms hiding behind this glitch. We just covered the trash can. Let's pivot to the second one, which the researchers call the broadcast hypothesis. Right, the broadcast hypothesis. Unlike the trash can, this sink acts like the water cooler in an office or, you know, a company megaphone. This specific token aggregates global information from the entire image, bundles it up, and then redistributes it to every other token in the network.

9:19It serves as a vital active communication hub. But wait, I have to challenge this based on the visual data in the research. Okay, go ahead. The paper notes that both of these phenomena look identical on an attention map. They both just look like a giant vertical stripe. They do, yeah. So how on earth can the researchers prove one is a trash can and the other is a megaphone if they look exactly the same to the human eye? Well, you have to look past the superficial attention map. you have to examine the residual stream that's the actual flow of mathematical data between layers. Okay. The diagnostic signature for the broadcast hub is completely different from the no-op trash can.

9:56Remember, the trash can has a near-zero payload, right? Right, the empty package. Exactly. A broadcast sink, on the other hand, has a normal healthy value norm. It is carrying a real dense payload of information. So the megaphone actually has something to say. It does. But the crucial diagnostic clue is how it delivers that payload. How so? Because it is operating as a megaphone, it takes that exact same vector, that one piece of global context, and adds it to multiple tokens simultaneously. Oh, I see. Mathematically, when an AI adds the exact same vector to a whole matrix of tokens, the update matrix becomes what we call low rank, specifically a rank one update.

10:37A rank one update. Okay, for the listener, let's ground that. What does a rank one update actually mean in plain English? Like, think of the AI's internal map of concepts. It's latent space. Sure. If you imagine the AI's latent space as a vast, multidimensional map where all concepts are organized, a rank one update means the AI takes every single token it's currently processing and shifts them all in the exact same direction by the exact same distance on that map. Wow. It's like taking a translucent colored filter and laying it over the entire image. It tints everything with the exact same shade of global context.

11:11That is an excellent analogy. Yes. It increases the mathematical similarity between all the tokens because they've all received the identical memo from the megaphone. Okay, that makes so much sense. So if you are looking at a model and see a vertical stripe where the token has zero payload, you're looking at a trash can. If you see a vertical stripe and it has a normal payload that forces a rank one shift across the board, you are looking at a megaphone. That is such a clever way to pull them apart. So let's take this diagnostic tool out of the theoretical realm and see what happens when researchers tested it in the wild.

11:44Let's do it. The data you shared applies these diagnostics to massive industry standard vision transformers. We're talking about heavy hitters like Dyno V2, both the large and giant versions, as well as open CLOP models. And when they ran the diagnostics on these massive pre-trained models, they uncovered a fascinating structural handoff as the data moves through the network. Walk us through that handoff. What happens to the early layers versus the deep layers? Okay, so in the early layers of the neural network, the AI consistently uses a special pre-assigned token called the CLS token as the sink.

12:21The CLS token. Yes. The CLS, or a class token, is an artificial token engineers add to the start of the sequence specifically to gather global semantic meaning. Like to eventually output the final decision, this image is a golden retriever. Exactly. But as you move deeper into the network's layers, a sharp transition occurs. The CLS token steps down from its sync duties, and random patch tokens, which are literal pieces of the image itself, are forced to take over the job. It's like a relay race. Yes. But why the switch? I mean, if the CLS token was doing a good job, why abandon it? Well. It makes me think of, like, company's master ledger.

12:57Oh, go on. Early on in the fiscal year, the ledger is mostly empty. So employees might scribble notes in the margins, essentially using it as a sink. But once that ledger contains the final critical calculations for the company's survival, the accounts lock it away in a vault. They refuse to write any more garbage on it, which forces the employees to just start scribbling on random desk blotters. Is the AI doing that? That ledger analogy hits the exact mechanism perfectly. The AI is actively protecting the CLS token once its semantic payload becomes too valuable. Wow. In the deep layers, the network realizes, I cannot overwrite this token with a trashcan or a megaphone operation anymore.

13:37It holds the final answer. So it preserves the CLS token and shifts the mechanical burden to a random image patch, sacrificing a piece of the picture to keep the system running. It purposely corrupts a piece of its own vision to save the global context. That's incredible. It is. Furthermore, the analysis reveals that this behavior isn't just a random glitch happening sporadically. Specific attention heads the individual processing units within a layer specialized in this behavior. Wait, they become dedicated sanitation workers or dedicated broadcasters? Precisely. The researchers found that some heads act as sinks for nearly 80 % of the images they process.

14:1580 %? Yeah. They are structurally designated by the AI's own learned weights to play this role permanently. And this brings us to the big reveal in the research, the moment that really recontextualizes everything. When they mapped out the massive Dino 2 giant model layer by layer, mapping out every single sink, they found that both algorithms coexist inside the exact same AI. Yes. It's not that one type of model prefers trash cans and a different model prefers megaphones. No, they operate simultaneously, but in completely different territories of the network. Right. The early and middle layers form a tight, dedicated cluster of trash can.

14:52No-op syncs, desperately suppressing useless updates as the AI figures out the basic shapes and edges of the image. And then right next door, the deeper layers form a dense cluster of megaphone broadcast syncs, sharing high-level global context. They live in the same house. They look identical on a visual graph, but they are executing entirely opposite jobs. Which perfectly explains the mitigation problem engineers have been facing. I mean, they have been tearing their hair out trying to fix attention sinks because they treated them as a single bug. Let's talk about those fixes. Because engineers hate glitches.

15:26They see a vertical stripe. They want it gone. Absolutely. The research outlines two popular independent fixes the industry has tried. The first is called gated attention. This is an architectural change that explicitly gives the AI a mathematical way to output a zero. Right. It adds a literal programmable gate so the AI doesn't have to blow up its math to build a trash can. It could just close the gate and safely pass on no information. It solves the forced choice problem of softmax directly. Right. The second popular fix is the use of registered tokens. This involves artificially injecting four extra completely blank tokens into the AI sequence right at the start.

16:06Okay. Blank tokens? Yeah. The logic here is we know you need a workspace, so here are four blank slates. Use these instead of corrupting the actual image patches. And registered tokens were designed specifically by engineers who were thinking about that broadcast workspace idea. They thought, let's give it four dedicated megaphones. But, and this is hilarious, when the researchers actually observed the AI, it didn't play by the rules. What's fascinating here is how the model reacts to those register tokens. It is perhaps the most revealing finding in the paper. When the team analyzed a Dyno V2 model equipped with four registers, the registers did successfully absorb almost all the sync behavior.

16:46The vertical stripes moved off the image patches and onto the registers, which means the image was protected. But the AI hijacked them. Completely hijacked their intended purpose. It did. The engineers intended them to be communication hubs. But the diagnostics revealed that the AI repurposed the vast majority of those registered tokens to act as no-op trash cans. Wow. It took this highly designed, premium workspace and turned it into a landfill. It is exactly like buying your kid a beautiful, expensive sandbox so they can build intricate sandcastles. castles and instead they just use it as a place to bury the broccoli they refuse to eat at dinner.

17:22That is exactly it. You gave them a creative workspace, but what they desperately needed in that moment was a trash can. A perfect illustration of emergent behavior overriding intended design. And this is the crux of the ultimate solution. Because the AI is actively running two completely different hidden algorithms, neither fix works alone. Let's break down why. If you only use gated attention, what happens? If you only use gated attention, you successfully solve the trash can problem. The AI can just output a zero. But the AI still desperately needs a megaphone for its deep layers. Oh. Since it doesn't have register tokens, it goes right back to corrupting an image patch to broadcast its global context.

18:03And on the flip side, if you only give it registers, it hijacks them to bury the broccoli and you run out of workspaces. Exactly. So the answer is you have to do both simultaneously. You have to combine them. Implementing gated attention gives the AI the direct architectural trash can it wants, completely satisfying the no-op requirement. Which frees up the workspace. Yes. That finally frees up the register tokens so the AI can use them exclusively as the megaphones they were designed to be. It cleanly separates the two conflicting needs. So we've walked through the mystery, the hidden algorithms, and the architectural fix.

18:39But what does this all mean for you, the listener? I mean, if you're building with these tools, integrating them into enterprise software or just relying on AI to analyze complex visual data, why should you care about separating the trash from the megaphones? It translates directly to performance, specifically on highly complex visual tasks. The researchers prove this empirically using Ligipa model. Now, if you test a baseline model on standard global classification like feeding it an image net picture and asking, is this a golden retriever? Combining gating and registers doesn't significantly change the accuracy.

19:12It hovers right around 69%. Which makes sense. Global classification is relatively easy for these massive models. It just needs the overall vibe of the image to guess dog. It doesn't need pixel perfect precision. But the story changes entirely when you evaluate them on complex, dense prediction tasks. Dense prediction. Yeah. These are tasks that require deep, localized spatial understanding, like image segmentation on data sets, such as 8020K or Pascal VOC 2012. So much harder. Much harder. In these scenarios, the AI isn't just naming the image. It has to outline exactly where the dog's fur ends and the fabric of the soca begins.

19:50It literally has to draw borders. Exactly. And on those precise tasks, the performance jumps significantly. On the ADE20K dataset, for instance, the model's performance goes from a.166 score to a.2000 MIU. Wow. That is a highly meaningful leap in spatial accuracy. Wait, MIOU? For anyone not building these models every day, what is that metric actually measuring? Ah, right. MIU stands for mean intersection over union. Mean intersection over union. Yeah. In the simplest terms, it is a grading system for how perfectly the AI colored inside the lines. Oh, I love that. It measures the overlap between where the AI thinks an object is and where the object actually is down to the pixel.

20:29A jump from 0.166 to 0.200 means its borders are getting much sharper and much more accurate. And that jump happens entirely because it's no longer corrupting its own vision to make makeshift tools. If we connect this to the bigger picture by separating the trash from the megaphone, the AI stops overriding the delicate local details of the image. Yeah. It preserves every single local patch representation, making it vastly superior at detailed pixel-level analysis. Okay, let's bring this all together. We started with a visual mystery, a vertical stripe on an attention map that looked like the AI was basically freezing up and staring blankly at a single pixel.

21:09Okay. We unpacked the research you shared to discover how that single visual anomaly is actually hiding two brilliant algorithms. There's the adaptive NOP trash can, where the AI mathematically zeros out its payload to ignore useless data. And then there's the broadcast megaphone, where it uses a rank one update to share vital global context. And we saw how combining gated attention with register tokens finally lets the AI do both jobs efficiently, unlocking a whole new level of spatial awareness. It really represents a profound shift in how we view these systems. We went from seeing a bug to understanding a feature to eventually designing an architecture that naturally supports what the AI is organically trying to do.

21:51It really is a paradigm shift. And it leaves me with one final thought I want you to mull over as we wrap up this deep dive. Oh, lay it on us. We just spent this time exploring how an artificial intelligence creatively hacked its own rigid architecture, hijacking register tokens, blowing up mathematical norms, pulling all these wild tricks just to build a trash can or a megaphone out of sheer computational necessity. Yeah. It really makes you wonder what other glitches or bugs that we are currently trying to patch out of modern neural networks are actually highly sophisticated learned algorithms desperately trying to overcome the rigid boxes we force them to live inside.

22:26It is a compelling question and one that should definitely change how we look at every anomaly moving forward. Until next time, keep looking past the stripes.

From the publisher

This research investigates the nature of attention sinks, which are specific tokens in Transformer models that attract disproportionate attention. The authors reveal that these identical visual patterns actually facilitate two distinct computational algorithms: Adaptive NOP and Broadcast. In the Adaptive NOP mechanism, the model uses a "null" token with near-zero value to suppress updates to the residual stream, essentially performing a "no-op" instruction. Conversely, the Broadcast mechanism uses a sink as a communication hub to aggregate and redistribute global information across the entire sequence. By applying specialized diagnostics to vision transformers (ViTs), the study proves that both mechanisms coexist and often transition from the [CLS] token to specific patch tokens in deeper layers. Finally, the authors demonstrate that combining gated attention with register tokens effectively mitigates these artifacts, leading to significantly improved performance in dense spatial tasks.

More from Best AI papers explained

All 475 episodes
A Unifying View of Attention Sinks: Two Algorithms, Two SolutionsBest AI papers explained · 23 min
Listen in VO