In short
Dwarkesh Podcast Episode Notes: Will Scaling Work? [Narration]
Overview The episode is a narration of a blog post titled "Will Scaling Work?" by Dwarkesh Patel, discussing the prospects of scaling up large language models (LLMs) and the implications for achieving Artificial General Intelligence (AGI) by 2040.
Key Themes
- Scaling and AGI: The potential for scaling LLMs to achieve AGI and the contrasting views of believers and skeptics.
- Data Bottlenecks: Concerns about the sufficiency of available high-quality data for training advanced AI models.
- Performance Benchmarks: The effectiveness of current benchmarks in assessing true AI intelligence versus memorization capabilities.
Detailed Summary
Introduction
- The episode begins with an introduction to the blog post, which outlines the debate on scaling LLMs.
- The discussion is framed as a dialogue between two characters: the "Believer" and the "Skeptic."
Major Arguments
Will Scaling Work?
- Believer's Perspective:
- Believers argue that scaling LLMs could lead to AGI, suggesting that current models have shown significant performance improvements.
- They believe that with more data, the existing model structures could yield intelligence comparable to humans.
- The ease and efficiency of scaling models is highlighted, with examples of past successes in AI development.
- Skeptic's Concerns:
- Skeptics counter that we may soon run out of high-quality data necessary for effective training.
- They argue that the required computational power (1E35 FLOPs) is far beyond current capabilities and that improvements in data efficiency are insufficient.
- Skeptics express doubts about self-play synthetic data as a viable alternative, citing challenges in evaluation and the massive compute requirements.
Data Challenges
- Skeptics emphasize that current models require vast amounts of data, and without innovative data generation techniques, progress could stagnate.
- The discussion highlights the discrepancy between the data needed for human-level performance and what is currently available.
Benchmarks of Intelligence
- Believer's View: The believer asserts that models consistently improve on benchmarks, which suggests they are becoming more capable.
- Skeptic's View: Skeptics argue that many benchmarks measure memorization rather than true intelligence, pointing out that models perform poorly on complex tasks requiring long-term reasoning.
Performance Scaling
- The believer claims that historical trends show consistent scaling of model performance with increased compute and data.
- Skeptics challenge this, questioning whether improvements in next-token prediction truly correlate with generalized intelligence and problem-solving capabilities.
Conclusions
- Dwarkesh concludes with a personal reflection, expressing a tentative probability estimate of 70% for scaling and algorithmic progress to lead to AGI by 2040, while acknowledging a 30% chance that skeptics are correct about the limitations of current approaches.
Key Takeaways
- Divergent Views: The podcast encapsulates the dichotomy between believers and skeptics regarding the future of AI and the potential for scaling to achieve AGI.
- Importance of Data: A recurring theme is the challenge of data sufficiency and the implications for the advancement of AI technology.
- Evaluating Intelligence: The discussion critiques existing benchmarks and questions their ability to measure true intelligence versus memorization capabilities.
- Future Predictions: The episode ends with a cautious optimism tempered by uncertainty, reflecting ongoing debates in the AI community.
Additional Resources
- Full blog post: [Will Scaling Work?](https://www.dwarkeshpatel.com/p/will-scaling-work)
- Follow on [Twitter for updates](https://twitter.com/dwarkesh_sp)
Conclusion The podcast offers an in-depth exploration of the complexities and nuances surrounding the scaling of AI models. It presents a balanced view of the arguments for and against the feasibility of achieving AGI through current methods, highlighting a critical area of contemporary AI research and discourse.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:00Hey everyone, this is a narration of a blog post I wrote called Will Scaling Work. You can find the full version on my website, dwarcashpattel .com. It was originally published December 26th, 2023. Will Scaling Work. When should we expect AGI? If we can keep scaling LLM's plus plus and get better and more general performance as a result, then there's reason to expect powerful AIs by 2040 or much sooner, which can automate most cognitive labor and speed up further AI progress. However, if scaling doesn't work, then the paths who AGI seems much longer and more intractable for reasons I explained in the post.
0:45In order to think through both the pro and the con arguments about scaling, I wrote the post as a debate between two characters I made up, believer and skeptic. When will we run out of data? Skeptic. We're about to run out of high quality language data next year. Even taking hand -wavy scaling curve seriously implies that we'll need 1E35 flops for an AI that is reliable and smart enough to write a scientific paper. And that's table stakes for the abilities an AI would need to automate further AI research and continue progress when scaling becomes infeasible. Which means we need 5 ohms that is orders of magnitude, more data than we seem to have.
1:32I'm worried that when people hear 5 ohms off how they register it is, oh we have 5x less data than we need, we just need a couple of 2x improvements didn't data efficiency and we're golden. After all, what's a couple of ohms between friends? No, 5 ohms off means we have 100 ,000 times less data than we need. Yes, we will get slightly more data efficient logarithms and multi -modal training will give us more data plus we can recycle tokens on multiple epochs and use curriculum learning. But even if we assume the most generous possible one -off improvements that these techniques are likely to give, they do not grant us the exponential increase in data required to keep up with the exponential increase in compute demanded by these scaling laws.
2:23So then people say we'll get self -play synthetic data working somehow. But self -play has two very difficult challenges. One, evaluation. Self -play worked for off -ago since the model could judge itself based on a concrete win condition. Did I win this game of go? But novel reasoning doesn't have a concrete win condition. And as a result, just as you'd expect, LLMs are incapable so far of correcting their own reasoning. Two, compute. All these math code approaches tend to use various sorts of research where you run LLM on each node repeatedly. Off -ago's compute budget is staggering for the relatively circumscribed task of winning at go.
3:09Now imagine that instead of searching over the space of go moves, you need to search over the space of all possible human thought. All this extra compute need it to get self -play to work and is in addition to the stupendous compute increase already required to scale the parameters themselves. Using the 1E35 flop estimate for human level thought, we need 9 ooms more compute at the top of the biggest models we have today. Yes, you'll get improvements from better hardware and better algorithms, but will you really get a few full equivalent of 9 ooms? Believer. If your main objection to scale working is just a lack of data, your intuitive reaction should not be.
3:55Well, it looks like we could have produced AGI by scaling up a transformer plus plus, but I guess we're going to run out of data first. Your reaction should be holy fuck. If the internet was a lot bigger, scaling up a model whose basic structure I can write down in a few hundred lines of Python code would have produced a human level mind. It's a crazy fact about the world that it's this easy to make big blobs of compute intelligent. The sample over which LLMs are inefficient is mostly just irrelevant e -commerce chunk. We compound this disability by training them on predicting the next token, a loss function which is almost completely unrelated to the actual task we want an intelligent agent to do in the economy.
4:40And despite this minuscule intersection between the abilities we actually want and the terrible loss function and data we train these models with, we can produce a baby AGI that is GPT -4 by throwing just 0 .03 % of Microsoft's yearly revenues at a big scrape of the internet. So, given how easy and simple AGI progress has been so far, we shouldn't be that surprised if synthetic data also just works. After all, the model is just want to learn. GPT -4 has been out for all of eight months. The other AI labs are not only now getting their GPT -4 level models, which means all the researchers are only now getting around to making self -prolet work with current generation models.
5:27And it seems like one of them might have already succeeded. Therefore, the fact that so far we don't have public evidence that synthetic data has worked at scale doesn't mean it can't. After all, RL becomes much more feasible when your base model is capable enough to get the right answer at least some of the time, because now you can reward that one in a hundred times that the model accomplishes the chain of thought required for an extended math proof. Or writes a 500 lines of code needed to complete a full pull request. Soon your one in a hundred success rate becomes 10 in a hundred than 90 in a hundred.
6:03Now you can try the 1000 light and pull request, and not only will the model sometimes succeed, but it will be able to critique itself when it fails. And so on. In fact, this synthetic data bootstrapping seems almost directly analogous to human evolution. R -4iimate ancestors show little evidence of being able to rapidly discern and apply new insights. But, once humans develop language, you have this genetic, cultural co -evolution, which is very similar to the synthetic data self -clay loop for LLMs, where the model gets smarter in order to better make sense of the complex symbolic outputs of similar copies.
6:45Self -play doesn't require models to be perfect at judging their own reasoning. They just have to be better at evaluating reasoning than doing a day -noval, which clearly already seems to be the case. See Constitution AI, for example. Or play around with GPT for a few minutes, and notice that it's better at explaining why what you wrote down is wrong than it is at coming up with the right answer for itself. Almost all the researchers I talk to in the big AI labs are quite confident that they'll get self -play to work. And when I ask why they're so sure, they heed for a moment, as if they're bursting to explain all their ideas.
7:23But then they remember the confidentiality of the thing, and say, I can't tell you the specifics, but there's so much low -hing fruit in terms of what we can try here. Or, as Dario Amadei, the CEO of Anthropic, told me on my podcast, and this is me asking the question, you mentioned that data is likely not to be the constraint. Why do you think that is the case? And Dario responding. There's various possibilities here, and for a number of reasons I shouldn't go into the details, but there's many sources of data in the world, and there's many ways that you can also generate data. My guess is that this will not be a blocker.
8:01Maybe it would be better if it was, but it won't be. Skeptic. Constitutional AI, RLHF, and other RL self -place setups are good at bringing out latent capabilities, or suppressing them when those capabilities are naughty. But no one has demonstrated a method to actually increase the models underlying abilities with RL. If some kind of cell place synthetic data doesn't work, you're absolutely fucked. There's no other way around the data bottleneck. A new architecture is extremely unlikely to provide effects. You would need a jump in sample efficiency much bigger than even LSTMs to transformers.
8:44And LSTMs were invented all the way back in the 90s. So you need a bigger jump than we have gotten out of the past 20 years when all the low -hanging fruit and deep learning has been most accessible. The vibes you're receiving from the people who have an emotional or financial interest in seeing LLM scale can substitute for the complete lack of evidence we have that RL can fix the many ooms shortfall in data. Furthermore, the fact that LLMs seem to need such as to pin this amount of data to get such mediocre reasoning indicates that they simply are not generalizing. If these models can't get anywhere close to human level of performance, with the data human would see in 20 ,000 years, we should entertain the possibility that two billion years worth of data also wouldn't do the trick.
9:35There's no amount of jet fuel that you can add to an airplane to make it reach the moon. Next topic has scaling actually even worked so far. Believe her. What are you talking about? Performance on benchmarks has scaled consistently for eight orders of magnitude. The loss in model performance has been precise down to many decimal places over a million fold increases in compute. In the GP -04 technical report, they say that they were able to predict the performance of the final GP -T4 model for models trained using the same methodology but using at most 10 ,000 times less compute than GP -T4. We should assume that a trend which has worked so consistently for the last eight ooms will be reliable for the next eight.
10:21And the performance which you would achieve from a further eight ooms scale up, or what in performance terms would be equivalent to an eight ooms scale up given the free performance boost we get from algorithmic and hardware progress would likely result in models that are capable enough to speed up AI research. Skeptic. But of course, we don't actually care directly about performance on next token prediction. The models already have human's beat on this loss function. We want to find out whether these scaling curves on next token prediction actually correspond to true progress towards generality.
10:57Believe her. As you scale these models, their performance consistently and reliably improves on a broad range of tasks as measured by benchmarks like MMMU, Big Bench, and Human Eval. Skeptic. But have you actually tried looking at a random sample of MMMU or Big Bench questions? They are almost all just Google search first hit results. They are good test and memorization, not of intelligence. Here are some questions I picked randomly from MMMU and remember these are multiple choice that model just has to choose right answer from a list of four. Question. Which of the following is always true of a spontaneous process?
11:39Answer. The total entropy of the system plus surrounding increases. Question. Who was president of the United States when Bill Clinton was born? Answer. Harry Truman. Now, why is it impressive that a model trained on internet text full of random facts happens to have a lot of random facts memorized? And why does that in any way indicate intelligence or creativity? And even on these contrived and orthogonal benchmarks, performance seems to be plateauing. Google's new Gemini Ultra model is estimated to have almost 5x more compute than GPT -4. But it has performed almost equivalently on an MMMU and Big Bench and other standard benchmarks.
12:18In any case, common benchmarks don't at all measure long horizon task performance. For example, can you do a job over a course of a month? Where LLM's trained a next -door prediction have very few effective data points to learn from? Indeed, as we can see on their performance on Sweet Bench, which measures if LLM's can autonomously complete pull request, they're pretty terrible at integrating complex info over long time horizons. GPT -4 gets a measly 1 .7%, but claw -2 gets a slightly more impressive 4 .8%. So we seem to have two kinds of benchmarks. The ones that measure memorization, recall, interpolation, and these are MMMU Big Bench human eval.
13:02Where these models are already appearing to match or even beat the average human. These tests clearly cannot be a good proxy for intelligence because even a scale maximalist has to admit that models are currently much dumber than humans. And the other type of benchmark we have are the ones that truly measure the ability to autonomously solve problems across long time horizons or difficult abstractions. This is Sweet Benchmark, ARC, where these models aren't even in the running. What are we supposed to conclude about a model, which after being trained on the equivalent of 20 ,000 years of human input, still doesn't understand that if Tom Cruise's mother is merely Fyfer, then merely Fyfer's son is Tom Cruise.
13:47Where his answers are so incredibly contingent in the way and order in which the question is phrased. So it's not even worth asking yet whether scaling will continue to work. We don't even seem to have evidence that scaling has worked so far. Believer. Gemini just seems like a bizarre place to expect a plateau. GPT -4 has clearly already broken through all the pre -registered critiques of connectionism and deep learning by skeptics. The much more plausible explanation for the performance of Gemini relative to GPT -4 is just that Google has not fully caught up to open AI's algorithmic progress.
14:25If there were some fundamental hard ceiling on deep learning and LLMs, shouldn't we have seen it before they started developing common sense, early reasoning and the ability to think across abstractions? What is a primaffatio reason to expect some stubborn limit only between mediocre reasoning and advanced reasoning? Consider how much better GPT -4 is than GPT -3. That's just a hundred X scale up, which sounds like a lot until you consider how much smaller that is than the additional scale up which we could throw at these models. We can afford a further 10 ,000 X scale up on GPT -4, i .e. something that's GPT -6 equivalent before we even touch 1 % of world GDP.
15:07And that's before we count for the pre -training compute efficiency gains, things like mixture of experts, flash attention, new post -training methods, RLI, fine tuning on chain of thought, self -lay, etc. and hardware improvements. Each of these will individually contribute as much to performance as you would have gotten from many ooms of raw scale up, and they have consistently done so in the past. At all these together, you can probably convert 1 % of GDP into a GPT -8 level model. For context on how much society is there willing to spend on new general -purpose technologies? One, British Railway investment at its peak in 1847 was a staggering 7 % of GDP.
15:53Two, even the five years after the Telecommunications Act of 1996 went into effect, Telecommunications companies invested more than $500 billion. That's almost a trillion in today's value, into laying fiber optic cable, adding new switches and building wireless networks. It's possible that GPT -8 aka a model which has the performance of 100 million times scaled up GPT -4 will only slightly be better than GPT -4. But I don't understand why you would expect that to be the case when we already see models figuring out how to think and what the world is like from far smaller scale ups. You know the story from here, millions of GPT -8 copies coating up kernel improvements, finding better hyper parameters, giving themselves boatloads of high quality feedback for fine tuning, so on.
16:45This makes it much cheaper and easier to develop GPT -9, extrapolate this all the way out to the singularity. Next topic, do models understand the world? Believe it, to predict the next token, an LLM has to teach itself all the regularities about the world which lead to one token following another. To predict the next paragraph in a passage from the self -as -gene requires understanding the gene -centered view of evolution. To predict the next passage in a new short story requires understanding the psychology of human characters, and so on. If you trade an LLM on code, it becomes better at reasoning in language.
17:27Now this is just a really studying fact. What this tells us is that the model has squeezed out some deep general understanding of how to think from bringing a shit ton of code. That not only is there some shared logical structure between language and code, but that unsupervised gradient descent can extract this structure and make use of it to be able to better reason. Grading descent tries to find the most efficient compression of its data. The most efficient compression is all the so -the -deepest and most powerful. The most efficient compression of a physics textbook, the one that would likely help you predict how a truncated argument from that book is likely to proceed, it is just a deeply internalized understanding of the underlying scientific explanations.
18:12Skeptic, intelligence involves, among other things, the ability to compress. But the compression itself is not intelligence. Einstein is smart because he can come up with relativity, but Einstein and relativity is not a more intelligent system in this sense that seems meaningful to me. It doesn't make sense to say that Plato was an idiot compared to me plus my knowledge because he didn't have a modern understanding of biology or physics. So, if LLMs are just the compression made by another process, it's too cast at gradient descent, then I don't know why that tells us anything about the LLMs' own ability to make compressions, and therefore why that tells us anything about the LLMs' intelligence.
18:55Believer, an airtight theoretical explanation for why scaling must keep working is not necessary for scaling to keep working. We didn't develop a full understanding of thermodynamics until a century after the C -mangin was invented. The usual pattern in the history of technology is that invention precedes theory and which should expect the same of intelligence. There's not some law of physics which says that Morslaw must continue, and in fact there are always new practical hurdles which imply the end of Morslaw. Get every couple of years, researchers at TMC, Nvidia, Intel, etc. figure out how to solve these problems and give the decades -along trend and extra lease on life.
19:40You can do all this mental gymnastics, but compute and data bottlenecks, the true nature of intelligence, and the vertleness of benchmarks, or you can just look at the fucking line. And the line here is a graphic that shows the transistor count over time, and you know the Morslaw famous exponential growth. Conclusion. All right, and now for the author Egos, here's my personal take. If you were a scale believer over the last few years, the progress we've been seeing would have just made more sense. There is a story you can tell about how GPD4's amazing performance can be explained by some idiom library or lookup table which will never generalize.
20:22But that's a story that none of the skeptics pre -registered. As for the believers, you have people like Ilya, Dario, Gorn, etc., more or less spelling out the slow takeoff you've been seeing due to scaling as early as 12 years ago. It seems pretty clear that some amount of scaling can get us a transformative AI, which is to say, if you achieve the irreducible loss on the scaling curves, you've made an AI that's smart enough to automate most cognitive labor, including the labor required to make smarter AI's. But most things in life are harder than a theory, and many theoretically possible things have just been intractably difficult for some reason or another, fusion power, flying cars, nanotech, etc.
21:08If self -play synthetic data doesn't work, then the models look funt. You're never going to get anywhere near that platonic irreducible loss. Also, the theoretical reason to expect scaling to keep working is murky, and the benchmarks on which scaling seems to lead to better performance have bet debatable generality. So, my tenetive probabilities are 70 % scaling plus algorithmic progress plus hardware advances will get us to AGI by 2040. 30%, the skeptics are right. LLMs in anything even roughly in that on vain is fucked. I'm probably missing some crucial evidence. DEI lives are simply not releasing that much research, since any insights about the signs of AI would leak ideas relevant to building the AGI.
21:57A friend who was a researcher at one of these labs told me that he misses his undergrad habit of winding down with a bunch of papers. Nowadays, nothing worth reading is published. For this reason, I assume that the things I don't know would shorten my timelines. Also, for what it's worth, my day job is a podcaster, but the people who could write a better post are prevented from doing so, either by confidentiality or opportunity cost. So give me a break and let me know what I missed in the comments. Appendix, here are some additional considerations. I don't feel I understand these topics so well enough to fully make sense of what they imply for scaling.
22:38Will models get inside base learning? Believe her, at a larger scale, models would just naturally develop more efficient meta -learning methods. Grocking only happens what you have a large over -parameterized model and beyond the point at which you've trained it to be severely overfit on the data. Grocking seems very similar to how we learn. We have intuitions and mental models of how to categorize new information, and over time, with new observations, those mental models themselves change. Gradient descent over such a large diversity of data will select for the most general and extrapolative circuits.
23:13Hence, we get Grocking, eventually we'll get inside base learning. Skeptic, neural networks have Grocking, but that's orders of magnitude less efficient than how humans actually integrate new explanatory insights. You teach a kid that a sun is at the center of the solar system and that immediately changes how he makes sense of the night sky, but you can't just feed a single copy of Copernicus into a model untrained on any astronomy and have it immediately incorporate that insight into all relevant future outputs. It's bizarre that the model has to hear information so many times in so many different contexts to grok the underlying concepts.
23:50Not only have models never demonstrated insight learning, but I don't see how such learning is even possible given the way we train neural networks with gradient descent. We give them a bunch of very subtle nudges with each example, with the hope that enough such nudges will slowly push them atop the correct tilt. Insight -based learning requires an immediate drag and drop from sea level to the top of Mount Everest. Does primate evolution give evidence of scaling? Believer. I'm sure you could find all sorts of these embarrassing fragilities in chimpanzee cognition, which are far more damning than the reversal curse.
24:27Doesn't mean there are some fundamental limit on primate brains that couldn't be fixed by a 3x scale plus a fine tunic. Indeed, a Susanna Herculano Huzelle has shown, the human brain has endless as many neurons as you'd expect from a scaled up primate brain with a mass of a human brain to have. Rodent and insectivore brains have much worse scaling loss. Relatively bigger brain species in those orders have far fewer neurons than you'd expect just from their brain mass. This suggests that there's some primate neural architecture that's really scalable in comparison to the brains of other species, analogous to how transformers have better scaling laws in LSTMs and RNNs.
25:12Evolution learned, where at least stumbled upon, the bitter lesson when designing primate brains, and the niche in which primates were competing strongly have rewarded marginal increases in intelligence. You have to make sense of all this data coming from your binocular vision, your tool using hands, and all these other smart makines who can talk to you. That's a full post. Thanks for listening. Again, the full blog post and other posts you can find at my website, the warcatchpatel .com. See you next time.
From the publisher
This is a narration of my blog post, Will scaling work?.
You read the full post here: https://www.dwarkeshpatel.com/p/will-scaling-work
Listen on Apple Podcasts, Spotify, or any other podcast platform. Follow me on Twitter for updates on future posts and episodes.
Get full access to Dwarkesh Podcast at www.dwarkesh.com/subscribe




