In short
The episode argues that large language models (LLMs) face a “wall” of diminishing returns: scaling up parameters/data yields rapidly worsening reliability and uncertainty, making them unsuitable for rigorous scientific/professional use without fundamental changes.
Guest backgrounds
No guest identities are provided in the transcript; it’s a host-led discussion referencing external researchers/authors.
Key claims
Scaling exponents are tiny (e.g., ~10 billion× compute for 10× error reduction; ~10^20× power for 10× accuracy). Transformer outputs become non-Gaussian (“fat tails”), causing “resilience of uncertainty” and “information catastrophes.” Spurious correlations grow with dataset size, potentially swamping true signal (true/spurious ratio as low as 10^-39 for a 34-bit string). Loss is a “pseudometric,” not a true scientific error measure.
Notable examples
GPT-4.5 rumored 5–10T parameters with 15–30× higher API cost but mainly subjective gains; OpenAI retracted “frontier model” framing. Meta Llama 4 (~2T parameters) underperforms. Random-number/joke anecdotes: models converge on 27; “bad at being random.” AlphaFold is framed as hypothesis-generating with failure rates on unseen structures. Mentions Monte Carlo scaling vs LLM scaling; “AlphaVolve” uses evolutionary algorithms to rigorously test LLM-generated code refinements; warns about “degenerative AI” from synthetic-data feedback loops.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOThe Wall Confronting LLMs
0:45 to 2:26
Discussion on the limitations and challenges that large language models are encountering.
“Okay, here's where it gets really interesting for me.”
Resource Demands of Scaling
2:26 to 4:42
Examination of the astronomical resource and energy costs associated with scaling LLMs.
“And we also see Meta's Lama 4 Behemoth, a$2 trillion parameter model, also underperforming relative to its sheer size.”
Non-Gaussian Behavior and Learning Issues
4:42 to 7:03
Exploration of non-Gaussian outputs and the learning behaviors of LLMs leading to inaccuracies.
“And for power consumption, where the exponent is even smaller, maybe 0.05, the numbers become astronomical.”
Spurious Correlations and Their Impact
7:03 to 9:10
Understanding how spurious correlations in large data sets affect LLM performance.
“The sources share some memorable anecdotes that really drive home this point.”
Limitations of Current AI Approaches
9:10 to 11:28
Discussion on the limitations of current AI methods and the need for new strategies beyond simply scaling up.
“It's a really sobering reevaluation of the current path, isn't it?”
Towards Degenerative AI
11:28 to 14:01
Exploring the concept of degenerative AI and the potential feedback loops of error accumulation.
“The hope there seems to be for more human-like reasoning attitudes, like, you know, pondering before answering, using chain of thought methods to cut error rates and boost reliability.”
Understanding the Scaling Dilemma
14:01 to 15:10
Explore the causal chain from data to information catastrophe in AI.
“And the causal chain leading to potential DAI is laid out quite clearly in the sources.”
Shifting Perspectives on AI Breakthroughs
15:10 to 16:26
Discuss the necessity for a paradigm shift in AI development strategies.
“If data doubles every single year, which is fast, actual information content might only double every 10 years because of these poor scaling laws.”
The Future of AI: Insight vs. Brute Force
16:26 to 17:34
Contemplate the balance between model size and understanding in AI.
“The source material concludes quite strongly that the scientific method itself, it provides precisely the means to sieve off the true correlations from the vastly greater sea of spurious ones.”
Transcript
Automatic transcript. May contain errors.0:00You know, it feels like every other day there's a new headline, right? AI achieving something pretty astounding. We hear about this relentless progress, large language models, LLMs, claiming to exceed human performance, and even things like AlphaFold recently getting Nobel-level recognition in chemistry and physics. So it's really easy to get swept up in the idea that AI is just on this unstoppable exponential march forward. But the material you shared with us today paints a very different and frankly quite sobering picture. It seems to suggest there are fundamental limits these systems are hitting.
0:33Yeah, that's right. What's fascinating here is that while the public narrative focuses heavily on what these models can do, the underlying science from these sources points to a significant fundamental tension. We're going to dive deep into what the authors are calling the wall confronting large language models. Our mission really in this deep dive is to unpack precisely what this wall is, why it's there according to these researchers, and what it truly means for the future of AI, especially, you know, its reliability for rigorous scientific and professional applications. Okay, here's where it gets really interesting for me.
1:06We constantly hear about models with like trillions of parameters. And the initial hype around LLMs suggested this idea of emergence, that they just keep learning better and better as they got bigger, fueling this intense big tech race to just scale up, make them bigger. But this material suggests the cost. I mean, both financially and environmentally for these massive models is just astronomically and apparently for what are relatively limited improvements. So what return are we genuinely getting for this gargantuan demand? That's the core question, isn't it? The authors point out the insatiable desire for electrical energy from these tech giants.
1:42I mean, to the point of potentially needing new nuclear reactors right next to data centers just to power these things, it's kind of mind-boggling. And yet the improvements they're seeing are, well, relatively limited, especially in the areas that truly matter for scientific rigor. Take the rumored GPT 4.5, for example, speculated to have maybe 5 to 10 trillion parameters. Its API cost is apparently an astonishing 15 to 30 times higher than GPT 4.0. Yet despite this massive scale, the qualitative gains seem primarily in subjective areas like creative writing or expressing empathy. But the sources say nothing substantial in verifiable domains like mathematics and science.
2:22In fact, OpenAI even retracted their initial statement calling it not a frontier model. That's telling. Yeah, really telling. And we also see Meta's Lama 4 Behemoth, a$2 trillion parameter model, also underperforming relative to its sheer size. So it's a huge investment for what seems like diminishing returns. So if I'm understanding this correctly, to get even a little bit better accuracy with an LLM, we're talking about resource demands that just, well, they defy common sense compared to what scientists are used to, like with other computational methods. That's a truly profound difference. It really is.
2:54I mean, in traditional scientific computer simulations like grid methods for solving complex differential equations, you understand how errors scale, right? You know, if you double your computational power, your error might reliably halve or quarter. There's a clear, predictable relationship. And all the data, all the parameters, they bear a direct scientific meaning. Exactly. There's a theoretical underpinning. But then you run into this infamous curse of dimensionality, the COD, where accuracy just rapidly deteriorates as dimensions increase, limiting these methods to maybe three dimensions or so.
3:27And that's why Monte Carlo methods are apparently so effective in ultra-high dimensions, because they combat the CODs by sampling hotspots. The sources give this vivid example. A random move in a system of 100 spherical molecules has only a 1 in 10 to the power of 260 chance of hitting the target, tiny. me. Monte Carlo combats this by intelligently focusing resources where they're most needed. It samples the important stuff. You've absolutely hit on the crucial point there. And if we connect this back to the bigger picture, it means that the scope for improvement in LLMs, purely by scaling up size, is described by these authors as absolutely untenable for most scientific applications.
4:06And that's before you even get to the power demands. Untenable. Wow. Let's try to put this into perspective. For Monte Carlo methods, which are considered pretty standard in many scientific fields. To cut an error by a factor of 10, you might need, say, 100 times more compute resources. That's already considered a tough scaling law, equals 0.5. But for LLMs, the scaling is shockingly poor. The exponents are tiny, like 0.1 or less. So to cut an error by just that same factor of 10, these models need, wait for it, 10 billion times more compute. 10 billion. It can't be right. That's what the scaling law suggests, It's 10 to the power of 10.
4:42And for power consumption, where the exponent is even smaller, maybe 0.05, the numbers become astronomical. We're talking 10 to the power of 20 times more power for just one order of magnitude better accuracy. It's just not viable. It almost sounds like a bad joke if it wasn't so serious. Wow. Okay. So it's not just that they're insanely expensive to scale. It's that the very way they learn seems to introduce fundamental issues like increasing uncertainty and this noise of false correlations you mentioned. It really sounds like they're building on incredibly shaky ground, and that wall just keeps getting higher.
5:15Exactly. That's a good way to put it. The core issue seems to lie in the architecture itself. LLM architectures, centered on these transformer models, they transform predictable Gaussian inputs into unpredictable non-Gaussian outputs. Think of Gaussian like a standard bell curve. Most things cluster around the average. Extreme outliers are super rare. Non-Gaussian means those extreme outliers, the fat tails, the distribution, become much more common than you'd expect. Okay, so more weird results, basically. Kind of. It's like finding a giant dinosaur in your backyard every other week instead of once in a million years.
5:49It messes with predictability. Now, this non-Gaussian behavior is actually linked to their learning capability. It's part of how they work. But it also inherently exposes them to potentially catastrophic failures and error pileup. Things can go wrong in big ways. And this leads to something the researchers call resilience of uncertainty, or RO. Essentially, those fat-tailed distributions mean achieving reliable accuracy requires vastly more data compared to more stable Gaussian systems. And this inability to accurately reproduce the extreme ends of distributions, the tails, can lead to what they term information catastrophes.
6:25Information catastrophes, that sounds good. Yeah, it does. And furthermore, the loss function in LLMs, which is often waved around as a measure of success, it isn't really like a conventional scientific error metric. You can't just drive it down to zero and declare victory. In fact, driving the loss too low can actually be detrimental. It can lead to problems like overfitting or mode collapse. So it's described as a pseudometric, a slippery and ill-defined concept when you need real quantitative answers. Right. It's not measuring genuine understanding in the way we might think. it's fascinating how those fundamental properties then kind of bubble up in the LLM's actual behavior.
7:03The sources share some memorable anecdotes that really drive home this point. Apparently, LLMs aren't very good at jokes, like they only have two in their repertoire or something. That's what one source claims, yeah. A bit limited. And when asked for a random number between 1 and 50, models from Anthropic, Google, OpenAI, even xAI's Grok after first giving 42, maybe as a joke, they all tend to give 27. 27. Yes, that specific number keeps cropping up. So they're not just bad at jokes. They're bad at being random. It sort of humanizes their limitations, doesn't it? It makes them less like magic boxes.
7:35And then there's this concept of Potemkin's, this illusion of understanding or holographic accuracy, where the answers look good on the surface until you actually start scratching beneath it. And that's where it gets truly concerning, especially for scientific use. The authors highlight this astonishingly little-known property identified by researchers Kalued and Longo. which is that spurious correlations, false patterns basically, they rapidly increase purely as a function of data set size. It doesn't even matter what the data is, and these spurious correlations overwhelmingly predominate over true correlations in large data sets.
8:10They just swamp the real signal. Wait, so more data means more noise, more false patterns. Exponentially more, yes. This raises a really important question. How much real signal can possibly be extracted when it's drowned in this exponential deluge of noise. For a thin 34-bit string, for instance, the ratio of true to spurious correlations can be as low as 10 to the power of minus 39. 10 to the minus 39. I can't even picture that number. It's practically zero. It's incredibly small. So for every one true insight these models might potentially glean from massive data sets, there are literally trillions upon trillions of false patterns they're also picking up on.
8:52And the authors note this finding is largely ignored by the mainstream AI community, yet it seems absolutely central to the fundamental problems with LLM. So if I'm hearing that right, it fundamentally undermines this whole idea that just throwing more data at an LLM automatically equals more intelligence or better understanding. It's a really sobering reevaluation of the current path, isn't it? And it explains why even the celebrated examples like AlphaFold have these reliability caveats you mentioned. It's like we're being told the car flies, but the small print says, yeah, I might occasionally forget how to land.
9:25That's a good analogy. This lack of generalizability, this inability to consistently perform well on unseen data, is a major concern highlighted in the sources. Machine learning accuracy seems highly dependent on the training data being very homogenous, very similar, and it deteriorates substantially when exposed to real-world, varied information, or even different types of data within the same field. Even within homogenous data, problems like multiscale physics or chaotic dynamics can still trip these models up. And validating on a held-out data set, often the sources argue it's more of a memorization process than a genuine learning one.
10:02The model just gets good at the test, not necessarily the underlying concepts. So while LLMs are often hailed for predicting the unseen and providing plausible responses, the reality is they still make many mistakes. And this is openly acknowledged by the tech companies themselves, usually in the small print. But the number of mistakes is far too many for the standards of accuracy required for most scientific and many other professional domains, like law or education. You can't have mostly right in those fields. Right. Plausible isn't good enough when the stakes are high. Exactly. Even AlphaFold, despite its Nobel-winning prabusin protein structure prediction, it still has substantial failure rates when faced with unseen structures from different scientific sources than the ones it was primarily trained on.
10:46So it provides valuable hypotheses, absolutely. It's a powerful tool. But it doesn't generally replace experimental structure determination. It gives you a really good starting point, a valuable hypothesis, but it's not necessarily the definitive final answer yet. So, OK, this sounds like the big tech industry must be aware of these fundamental limitations, right? They can see the wall, too. And they're trying to move beyond LLMs, maybe by changing the kind of AI, not just making the current ones bigger. That seems to be the direction. We're seeing the emergence of large reasoning models, LRMs, which aim to provide evidence for their answers, and also agentic AI, where you have multiple AI systems working together.
11:28The hope there seems to be for more human-like reasoning attitudes, like, you know, pondering before answering, using chain of thought methods to cut error rates and boost reliability. That's the hope, yes. LRMs try to add justification, but their empirical basis makes quantifying performance maybe even harder. And with egenic AI, the idea is that collaboration might overcome individual model weaknesses. Using techniques like chain of thought or key of tea aims to make the process more transparent and less error-prone. But this raises an important question, which the sources also touch upon. Are we simply building more complex systems on top of fundamentally unstable foundations?
12:04Right. If the base LLM has these inherent scaling and correlation issues, does hooking them together really solve the root problem? The source material suggests that insofar as these new approaches, like agentic AI, still depend heavily on the underlying LLM features we've discussed. It's considered highly unlikely they'll lead to truly sustainable, scalable strategies that overcome the wall. The core issues of scaling, non-Gaussian behavior, and spurious correlations seem likely to persist, regardless of the new packaging. Hmm. So, more complexity, but maybe the same fundamental limits? Possibly.
12:39However, there's also a provocative alternative idea mentioned in the material. Instead of constantly trying to suppress hallucination, maybe we should let LLMs do what generative models are kind of meant to do, which is hallucinate, generate novel, if sometimes weird, stuff. The idea is to channel this generative looseness into exploratory value, and then use other, more rigorous components to handle the evaluation and selection. AlphaVolve is mentioned as an example, where an LLM dreams up code refinements, but then an evolutionary algorithm rigorously tests and guides the process. So you use the creativity, but check it carefully.
13:12It's interesting. Leaning into the weirdness instead of fighting it. So, okay, it's not necessarily an inevitable AI doomsday scenario painted here, but it's definitely a strong warning that our current main approach, just scaling up LLMs, has some serious built-in limits. It really sounds like the authors are arguing the problem isn't AI itself, but how we're building it right now, the principles we're relying on. Exactly. The sources discuss a potential pathway to what they call degenerative AI, or DAI. They define that as the catastrophic accumulation of self-feeding errors and inaccuracies.
13:47And this is seen as particularly likely if LLMs are increasingly trained on synthetic data generated by other AIs. You get this feedback loop of errors. Oh, right, like AI eating its own potentially flawed tail. That sounds like a terrifying feedback loop. It truly does. And the causal chain leading to potential DAI is laid out quite clearly in the sources. Small scaling exponents, SSE, lead to non-Gaussian fluctuations, NGF. These fluctuations cause this resilience of uncertainty, ROU, which can ultimately lead to information catastrophes for service. So the chain is basically SE leads to NGF, which leads to ROU, which can result in IC.
14:24SSE, NGF, ROU, IC. A chain reaction of diminishing returns and increasing uncertainty. Precisely. And it's absolutely critical here to remember that data are often implicitly assumed to be the same as information. But as the sources stress, this is obviously not true. Adding more data can actually decrease useful information. For example, if you add conflicting data or if the data set gets poisoned with fake news or indeed flawed AI generated content, more data isn't always better information. Right. Quality over just sheer quantity. Exactly. And even though the current staling exponents for LLMs are still technically positive, though tiny, it means we haven't hit a full degenerative regime yet.
15:04But we are firmly in a regime of strongly diminishing returns. Unquote. The analogy provided is quite stark. If data doubles every single year, which is fast, actual information content might only double every 10 years because of these poor scaling laws. Even if you factor in something like Huang's law, the idea that GPU power doubles roughly yearly, it still takes that decade to double the useful information gleaned by the LLM. This is because there's this near tie, as they put it, between the exponential growth of resources we throw at it and the exponential decay of true correlations within the data deluge.
15:40Wow. So we're running faster and faster just to stay almost in the same place, information-wise. That's a pretty good summary of the scaling predicament, yes. When I first saw these specific scaling exponents laid out, it completely shifted my perspective on where the real AI breakthroughs need to come from. It's probably not just from bigger models. That's a lot to mull over, definitely. This deep dive has certainly challenged some of the prevailing narratives around LLMs. It's a reminder that sometimes the most advanced solutions might actually require a return to more fundamental principles.
16:10The core message for avoiding DAI and maybe for actually breaking through this wall seems to be, well, don't just rely on brute source. The sources say sacrificing understanding and insight on the altar of brute force and unsustainable scaling boosts the chances that a wall will indeed persist no matter how big the data. That's the crux of it. The source material concludes quite strongly that the scientific method itself, it provides precisely the means to sieve off the true correlations from the vastly greater sea of spurious ones. How? Through the careful construction of world models based on understanding mechanisms, causality, underlying principles, not just pattern matching.
16:46Simply ignoring these principles on the supposition that brute force computation can eventually solve everything is, in their view, doomed to failure. So it's a very clear call for a shift in approach, away from just brute force scaling and empirical trial and error, and towards a greater emphasis on insight and understanding, leveraging the fundamental principles of the scientific method to build more robust AI. So as you, our listener, consider the future of AI and its role in all our lives. Maybe ask yourself, are we just chasing the next biggest model, the highest parameter count? Or are we truly seeking deeper understanding to build genuinely reliable, trustworthy intelligence?
17:22What does it really mean to be well-informed about AI's amazing potential, but also its very real and perhaps fundamental limitations? Thank you for joining us on this exploration today. It's certainly given me a lot to think about.
From the publisher
The article **"The wall confronting large language models"** by P.V. Coveney and S. Succi examines the inherent limitations of Large Language Models (LLMs), arguing that their **scaling laws severely hinder improvements in prediction accuracy**, making it practically impossible to meet scientific standards. The authors suggest that the very mechanism enabling LLMs to learn, specifically their ability to generate non-Gaussian outputs from Gaussian inputs, also contributes to **error accumulation and "information catastrophes."** They contrast LLM scaling with traditional computer simulation methods, highlighting the **drastically diminishing returns** on increased computational resources for LLMs. The paper ultimately posits that ignoring the scientific method in favor of brute-force scaling leads to a **"Degenerative AI" (DAI) pathway**, characterized by low scaling exponents and overwhelming spurious correlations in vast datasets, advocating for greater emphasis on **insight and understanding** to avoid this outcome.




