In short
Compares in-context learning (ICL) across MLPs, MLP mixers, and transformers on synthetic regression/classification and relational reasoning tasks, focusing on compute efficiency, context length limits, and out-of-distribution generalization.
Guest backgrounds
No guests are mentioned; it’s a solo “Deep Dive” discussion.
Key claims
ICL is not transformer-exclusive; with enough compute, MLPs/mixers/transformers reach near Bayes-optimal performance. Vanilla MLPs fail on long-context regression beyond ~26 examples (tends toward guessing zero), while mixers avoid this via token-mixing across the context. On relational reasoning (match-to-sample, sphere oddball), MLPs can outperform transformers and generalize better OOD.
Notable examples
Match-to-sample (closest point), sphere oddball (identify odd point under larger-than-trained perturbations), RB-MLP control (dot-product bias helps radial tasks but fails line oddball), same-different with unseen symbols requiring >2^12 symbol diversity.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOChallenging Common Perceptions of ICL
0:45 to 1:20
Discussing misconceptions about in-context learning and MLP capabilities.
“And the results are frankly quite surprising.”
Research Methodology Overview
1:20 to 2:20
Explaining the controlled experiments designed to test various models.
“We're going to unpack how these models compare on ICL tasks using controlled experiments.”
Performance of Various Architectures
2:20 to 3:16
Comparing MLPs, mixers, and transformers in ICL tasks and compute efficiency.
“And how did the different architectures stack up?”
In-Weight Learning vs. True ICL
3:16 to 4:28
Understanding the difference between memorizing and genuine in-context learning.
“The sources pointed out that especially on classification tasks, at the lower end of the compute budget, the vanilla MLPs sometimes actually got a lower error, a better score than the transformers.”
Context Length Limitations in Regression
4:28 to 7:11
Examining the performance drop of MLPs with longer context in regression tasks.
“Okay, so give them enough variety, and even the basic MLP learns the general trick of ICL.”
Classification Task Performance
7:11 to 8:00
Evaluation of model performance in classification tasks with varying context length.
“It struggles with maybe combining numerical information over long sequences, but not categorical information.”
Relational Reasoning Tasks Insights
8:00 to 10:32
How MLPs outperformed transformers on relational reasoning tasks and generalization.
“And historically, people thought MLPs just couldn't handle this kind of relational stuff, right?”
Same-Different Task and Data Diversity
10:32 to 13:08
Analyzing how data diversity impacts MLP generalization in relational tasks.
“It suggests a fundamental limitation in how the transformer was representing or generalizing that geometric relationship.”
Efficiency of MLPs in High Dimensions
13:08 to 14:02
Exploring MLP efficiency in classification and regression tasks with high input dimensions.
“Then the MLP achieved near-perfect generalization to completely unseen symbols.”
Comparative Analysis of MLPs and Transformers
14:02 to 16:10
Learn about the capabilities of MLPs in in-context learning compared to transformers.
“What should we take away from this deep dive?”
Transcript
Automatic transcript. May contain errors.0:00Welcome to the Deep Dive. Today we're really getting into something interesting. We're looking at a stack of research that might just shake up how we think about AI architecture. Absolutely. We're tackling this idea that, you know, transformers are the only models that can really do in-context learning or ICL. Right. ICL, that's that almost magical ability models have now. Learning a totally new task instantly just from the examples you feed it in the prompt. Exactly. Like translating a made-up language, boom, just from seeing a few pairs. And yeah, the common wisdom was always transformers do this thanks to attention.
0:35No weight changes, just pure context. But that's the twist, isn't it? These sources we've been digging into, they paint a very different picture. They actually compare different architectures head to head. And the results are frankly quite surprising. It looks like the multilayer perceptron, the MLP, you know, the simplest, oldest kind of workhorse neural net, can also do ICL. Which is kind of wild. People often dismiss MLPs these days. Totally. This really challenges that whole narrative that you need attention for this kind of quick, flexible learning. It suggests these simpler architectures have more going on than we maybe thought.
1:10Yeah, it suddenly makes these like all MLP alternatives seem much more viable again, right? Precisely. There's a growing interest there and this research definitely fuels that. Okay, so here's the plan for today. We're going to unpack how these models compare on ICL tasks using controlled experiments. Then we'll get into this tricky issue with context length where things get weird. And then the really juicy part, how MLPs actually seem to beat transformers on some classic reasoning tasks. It's pretty compelling stuff. All right, let's start with unpacking ICL itself, but in the specific way these researchers set it up.
1:43They didn't use like messy internet text, did they? No, exactly. They wanted to isolate the core ICL ability, so they designed synthetic tasks, controlled versions of regression and classification. Think predicting the next number in a sequence defined only by examples in the prompt, or classifying points based on a boundary shown just before. So you feed it, say, three input-output examples defining a rule. Right, like three points on a line, or some examples of points inside versus outside a circle. And then you give it a new input, a query, and it has to figure out the output just based on those few examples.
2:17That's ICL in a nutshell for these experiments. No pre-training on that specific rule. No weight updates during the test. Just inference from context. Okay. And how did the different architectures stack up? They looked at compute efficiency, right, using PFLOPs. Yeah, PFLOPs, PETA floating point operations. It's basically a measure of the total computational work done during training. How much effort does it take to get the model smart? And the big finding there was. The big finding was, well, maybe less surprising in hindsight. If you throw enough compute at them, they all get there. MLPs, the MLP mixers, which we'll get to, and transformers, they all achieve pretty comparable near-optimal ICL performance.
2:58Near-optimal meaning like about as good as you can possibly do on that specific task. Exactly. For the regression tasks, they all approached what's called the Bayes Optimal Ridge Mean Squared Error. basically peak performance. So the capacity for ICL isn't exclusive to transformers. That's already a big deal, but what about efficiency, like getting started? Yeah, that was interesting. The sources pointed out that especially on classification tasks, at the lower end of the compute budget, the vanilla MLPs sometimes actually got a lower error, a better score than the transformers. So the simpler network was like cheaper to get off the ground for those tasks.
3:35He was like it. More bang for your buck early on, at least for classification ICL. Okay, so similar peak performance, maybe some early efficiency wins for MLPs. Did they see any common patterns in how these models learned ICO? Yes, they did. All three architectures, MLP, Mixer, Transformers, show this clear transition. They start with what the researchers call in-weight learning, or IWL. IWL sounds like just memorizing stuff. Pretty much. It's learning specific patterns or clusters seen during training and storing that in the weights. But as you increase the diversity of the training data, meaning you show it way more possible patterns, more unique functions or cluster types.
4:15It can't just memorize everything anymore. Right. It's forced to switch strategies. It has to learn the abstract rule from the context examples it sees at test time. That's the jump to true in-context learning, ICL. And all architectures made that jump when the data got diverse enough. Okay, so give them enough variety, and even the basic MLP learns the general trick of ICL. Fascinating. But now, let's get into where the differences really start to show up, the context length issue. Right, this is where things get particularly interesting, and you see a real divergence, especially with the vanilla MLP.
4:49What happened? So specifically in ICL regression tasks, the plain MLP's performance just tanks, as you give it more examples in the context prompt. Tanks how badly? Like really badly. The study found it essentially failed to learn in context once the prompt got longer than about 26 examples. Its performance dropped towards just guessing zero output for everything. Wow. Okay. So more information actually made it dumber. Why would that happen? Why does the simplest model fall apart like that, but only in regression? Well, think about how a basic MLP sees the input. It typically flattens everything into one giant vector.
5:28All the examples and the context just get mushed together. Right. It doesn't really have a built-in way for different parts of that long sequence, different examples to talk to each other or relate to each other easily across the whole input. It loses the structure. And that's where the attention mechanism in a transformer would normally help, right? Letting different parts focus on other relevant parts. Exactly. Attention provides that mechanism. But here's the kicker for the all-MLP story. Ah, the MLP mixer you mentioned earlier. Yes, the MLP mixer. Now, this architecture also avoids that computationally heavy attention mechanism.
6:01Okay, so what does it do differently from a vanilla MLP? It has these special blocks, sometimes called token mixing MLPs. They allow information to flow sort of sideways across the different examples or tokens in the sequence. Ah, so the examples can kind of interact or share information before being processed deeper. Precisely. It allows communication across the context sequence dimension, not just within each example's features. And the result for context length and regression. Strikingly, the mixer doesn't have the vanilla MLP's problem. Its performance stays near that Bay's optimal level, even with really long context prompts and regression.
6:38So the mixer and the transformer both handle long context regression pretty well, but the vanilla MLP stumbles. Exactly. It pinpoints a specific weakness in the very simplest architecture for that task. But hang on, you said this was specific to regression. What about classification? Yeah, that's the other really interesting part. For ICL classification, all of them, vanilla MLP, mixer, transformer, seems mostly fine with longer context. The performance curves were pretty flat as context length increased. Huh. So the vanilla MLP's long context problem is very task-specific. It struggles with maybe combining numerical information over long sequences, but not categorical information.
7:17It seems like it. It suggests it's not just about long inputs in general, but about the specific kind of inference needed for regression over those long inputs. A subtle but important distinction. Okay, this is complex. Let's shift gears now to maybe the most counterintuitive results, the relational reasoning tasks. Ah, yes. This is where MLPs didn't just keep up, they actually seemed to win. We're talking about tasks inspired by classic cognitive psychology experiments. Like match-to-sample, oddball. Things designed to test if a system understands relationships between items. Exactly. Functionally, they are set up as in-context classification tasks.
7:56But the underlying challenge is reasoning. Which item is closest, which one is different, regardless of absolute position. And historically, people thought MLPs just couldn't handle this kind of relational stuff, right? Especially generalizing it. That was definitely a common argument against them. But these studies show something different. What did they find? On these tasks, MLPs performed better than transformers, especially when you considered the compute needed. They learned the relationships more efficiently. And crucially, they generalized better on out-of-distribution tests. Okay, let's take an example.
8:27Match to sample, MTS. What's the task there? You get a query point and several context points. You have to pick the context point that's geometrically closest to the query. It's about relative distance. And the MLP. The vanilla MLP reached a lower loss, meaning better performance, much faster with less compute than the transformer. It just grasped that closeness relationship more easily. And the out-of-distribution part, how did they test generalization? They trained the models, for example, with points arranged within a certain radius, then tested them with points arranged in a different radius, something the model hadn't explicitly seen.
9:01The MLP models, both vanilla and mixer, handled it pretty well. Their performance stayed solid. The transformer, however, kind of fell apart. Its accuracy dropped sharply. Suggesting the transformer had maybe latched onto the specific training setup, not the general rule of closeness. That's what it looks like. It failed to generalize the geometric relationship when the specifics changed. Wow. Okay, what about the oddball task? Similar story, maybe even more dramatic. In sphere oddball, you have a cluster of points, say six identical ones, and one oddball point that's been slightly moved or perturbed.
9:37The task is to identify the oddball. And again, the MLP outperformed. Yes. The vanilla NLP beat the transformer quite convincingly in terms of compute efficiency. And the O test for oddball. This sounds interesting. It was really telling. They trained the models with a certain small perturbation distance for the oddball. Then they tested by massively increasing that distance, moving the oddball much further away than anything seen during training. What happened? The MLPs continued to perform perfectly. They correctly identified the oddball no matter how far away it was. The general rule was learned.
10:11And the transformer. Its performance decayed significantly. When they looked inside, it seemed like the transformer's output score, the logit, for the oddball point kind of hit a ceiling. It couldn't properly represent, this is really far away. So it could tell it was different, but not how different beyond a certain point. A failure in generalizing the magnitude of the relationship. Exactly. It suggests a fundamental limitation in how the transformer was representing or generalizing that geometric relationship. This really points towards different underlying biases in the architectures. Did the researchers explore that idea?
10:45You mentioned something about a specialized MLP. Yes, the Relationally Bottlenecked MLP, or RB MLP. This was a neat control experiment. They basically hard-coded a specific type of relational reasoning into the MLP using dot products, which are great for radial distance calculations like in MTS. So a super specialized MLP designed just for those kinds of tasks. Precisely. And on tasks like MTS and Sphere Oddball, where that dot product bias is perfect, the RB MLP was incredibly efficient, even better than the Vanilla MLP, the perfect tool for the job. But the catch with specialists is... They break when the job changes slightly.
11:21They tested it on Line Oddball, where the points are arranged linearly, and identifying the oddball requires sensitivity to linear structure, not radial distance. And the specialist RBMLP? It completely failed, couldn't solve the task at all. Its strong bias was the wrong bias. But the vanilla MLP? Vanilla MLP actually performed slightly better than the transformer on line oddball. It wasn't perfect, but its generalist nature allowed it to adapt better than the specialist RBMLP or the transformer in that slightly different relational context. So strong bias helps only if perfectly matched, otherwise the generalist approach might win out.
11:59That's fascinating. It really reinforces this idea, sometimes called the bitter lesson in AI. Which it is. Broadly, it suggests that general purpose learning methods that leverage massive computation will ultimately outperform systems that rely heavily on clever, human-designed inductive biases or specialized structures. Scale and data eventually beat intricate design, even if the specialized design seems smarter initially. And the MLP's success here, especially on relational tasks, seems like a case study for that lesson. It also pushes back against a really old criticism of MLPs, doesn't it?
12:33Absolutely. That longstanding claim that MLPs just can't do relational reasoning, especially generalizing to new symbols or items they haven't seen. How did this research tackle that directly? They used the classic same different task. Given pairs of symbols, the model has to say if they're the same or different. The key is generalizing this to symbols it never encountered during training. And could the MLP do it? Yes, but with a crucial condition. Yeah. Data diversity. When they trained the MLP on the same different task, using a massive number of unique symbols specifically, they mentioned needing more than 2 to the power of 12 symbols.
13:08Okay, so a lot of variety. Then the MLP achieved near-perfect generalization to completely unseen symbols. Wow. So it wasn't that MLPs couldn't reason relationally. It was that earlier experiments probably just didn't give them enough diverse data or scale to learn the abstract, same, different concept. That seems to be the implication. The capability was likely there, latent, waiting for sufficient data and feature learning to emerge. It wasn't an inherent architectural flaw for that kind of reasoning. That really reframes decades of AI history. It does. And just one more point from the sources, even forgetting ICL for a moment.
13:43For just standard simple tasks like basic classification or regression, if the input dimension is large, say 64 dimensions, MLPs are just flat out more compute efficient than transformers. The overhead and the attention mechanism really starts to bite on long sequences or high dimensions, even for simple problems. Feed forward is just cheaper. Okay, let's try and wrap this up. What's the big picture here? What should we take away from this deep dive? Well, the core message is that in context learning, this amazing adaptive ability isn't some exclusive magic trick of the transformer architecture.
14:16Right. MLPs, and especially these MLP mixers, they're surprisingly capable contenders. Exactly. They can often match transformers on ICL tasks, sometimes be more compute efficient, and quite strikingly, they seem to generalize better on certain fundamental relational reasoning tasks. But there are paviots, right? This was all done with synthetic controlled data. That's the critical next question. These results provide a powerful theoretical baseline challenging assumptions. But how does this translate to the massive, messy, real-world data sets? Language, vision, the stuff where transformers currently dominate at scale.
14:54We don't know for sure if MLPs would scale the same way on, say, learning to write poetry from context. Precisely. Is the transformer's edge in the real world purely due to its attention mechanism being better suited for that kind of complex data? or is it just that we've invested so much more in scaling that architecture? Or is it something about handling the sheer scale of context in real language? We need more research comparing them in those scaled up, data limited, or more naturalistic settings. Okay, so it leaves us with a really practical, maybe even provocative question. We've seen transformers excel when data is super diverse, quickly learning ICL, but they cost a lot computationally and can stumble on some generalization tests.
15:35MLPs struggle with very long context and regression, but nail certain relational generalizations and are simpler. Yeah, the tradeoffs are becoming clearer. So for you listening, if you had the next giant pile of compute, that next petabyte, where would you place your bet for building the next truly massive AI? Would you keep refining the attention mechanism, trying to overcome its costs and generalization quirks? Or would you take a chance on scaling up these seemingly simpler, sometimes better generalizing all MLP architectures, hoping they can bridge the gap on complex tasks with enough scale?
16:08Definitely something to mull over. We'll be back next time to break down another stack of critical research.
From the publisher
This research paper demonstrates that Multi-Layer Perceptrons (MLPs) can perform In-Context Learning (ICL), an ability often attributed exclusively to Transformer models. The researchers show that MLPs, and related MLP-Mixer models, achieve performance comparable to Transformers on synthetic ICL tasks involving regression and classification. Furthermore, in experiments testing relational reasoning—which is related to ICL classification—MLPs surprisingly outperformed Transformers in terms of both compute efficiency and generalization. These findings suggest that ICL is not solely dependent on attention-based architectures and challenge previous assumptions about the limitations of simple neural networks like MLPs in solving relational tasks. The study encourages further exploration of non-Transformer architectures to better understand the mechanisms of ICL.




