In short
Dwarkesh Podcast Episode Summary
Podcast Title
Dwarkesh Podcast
Episode Title
Francois Chollet, Mike Knoop - LLMs won’t lead to AGI - $1,000,000 Prize to find true solution
Episode Description
In this episode, host Dwarkesh Patel interviews Francois Chollet, an AI researcher at Google and the creator of Keras, alongside Mike Knoop, co-founder of Zapier. They discuss the newly launched $1 million ARC-AGI Prize, which aims to solve the ARC benchmark—a test designed to measure machine intelligence. Chollet argues why Large Language Models (LLMs) may not lead to Artificial General Intelligence (AGI), and the episode explores various facets of intelligence and the potential pathways toward AGI.
---
Key Discussion Points
- Introduction to the ARC Benchmark
- Definition: The ARC (AI Research Collaborative) benchmark is conceptualized as an IQ test for machine intelligence.
- Unique Design: It focuses on assessing machine intelligence's adaptability rather than its memorization capabilities, requiring basic core knowledge (e.g., physics and basic reasoning).
- Novelty of Challenges: Each puzzle is novel, meaning even if a model could memorize data from the internet, it wouldn’t be able to solve the tasks without reasoning from scratch.
- Challenges for LLMs
- Memorization vs. Intelligence: Chollet argues that LLMs, including future iterations, may excel at memorizing but struggle with genuine reasoning and adaptability.
- Training Limitations: If LLMs can only solve problems they have been trained on, they lack genuine intelligence, which requires the ability to adapt to entirely new situations.
- The Importance of Core Knowledge
- Human Learning: Humans acquire core knowledge early in life, which enables them to tackle novel problems efficiently.
- LLM Limitations: In contrast, LLMs are often stuck in a memorization loop, failing to demonstrate the adaptability seen in human intelligence.
- The Implications of the ARC Prize
- Prize Structure: The $1 million prize includes a $500,000 reward for the first team to achieve an 85% success rate on the ARC benchmark, as well as additional rewards for progressive scores and contributions.
- Encouragement of Innovation: The prize aims to inspire researchers to develop new methodologies capable of handling novel problem-solving without relying solely on memorization.
- Long-Term Vision for AGI
- Combined Approaches: Future AI development may require hybrid models that integrate core memory systems with adaptive reasoning capabilities.
- Resistance to Benchmark Saturation: The goal is to create a benchmark that continuously challenges AI systems, fostering genuine advancements towards AGI.
---
Key Takeaways
- LLMs may not lead to AGI: Key arguments presented by Chollet highlight the limitations of LLMs in achieving true intelligence compared to human cognitive capabilities.
- Novelty in Challenges: The ARC benchmark’s design is intended to emphasize the importance of reasoning and adaptability over raw memorization.
- Potential of Hybrid Systems: Future progress in AI may involve combining different approaches (e.g., deep learning and program synthesis) to enhance adaptability and problem-solving skills.
Additional Resources
- ARC Prize Website: For more information on the prize and participation details, visit [ARC Prize](https://arcprize.org).
- Podcast Platforms: Listeners can access the episode on platforms like [Apple Podcasts](https://podcasts.apple.com/us/podcast/francois-chollet-mike-knoop-llms-wont-lead-to-agi-%241/id1516093381?i=1000658672649) and [Spotify](https://open.spotify.com/episode/7bmeJQOvXGy4LYl6YoiYYP?si=obUSUEwjSA6tkB8EBcb18w).
---
This summary encapsulates the discussion and highlights the significant points made by the guests, providing a comprehensive overview of the issues surrounding AI, LLMs, and the pursuit of AGI.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:00Okay, today I have the pleasure to speak with Francois Cholet, who is a AI researcher at Google and creator of Keras. And he's launching a prize in collaboration with Mike Knuth, the co -founder of Zapier, who will also be talking to in a second, a million dollar prize to solve the ARC benchmark that he created. So first question, what is the ARC benchmark, and why do you even need this prize? Why won't the biggest LLM we have in a year be able to just saturate it? Sure. So ARC is intended as a kind of IQ test for machine intelligence. And what makes it different from most LLM benchmarks out there is that it's designed to be resistant to memorization.
0:39So if you look at the where LLM's work, they're basically this big interpretative memory. And the way you scale up the capabilities is by trying to cram as much knowledge and patterns as possible into them. And by contrast, ARC does not require a lot of knowledge at all. It's designed to only require, with known as core knowledge, which is basic knowledge about things like elementary physics, object -ness, counting, that sort of thing. The sort of knowledge that any four -year -old or five -year -old processes. But what's interesting is that each puzzle in ARC is novel. Is something that you've probably not encountered before, even if you've memorized the entire internet.
1:26And that's what makes it... So in the sweat makes ARC challenging for LLM's. And so far, LLM's have not been doing very well on it. In fact, you've purchased it out working well. Our motorods, discrete program search, program synthesis. So first of all, I'll make a comment that I'm glad that, as a skeptic of LLM, you have put out yourself a benchmark that is it accurate to say that, suppose that the biggest model we have in a year is able to get 80 % on this. Then your view would be, we are on track to AGI with LLM's. How would you think about that? Right. I'm pretty skeptical that we're going to see LLM do each person in a year.
2:07That said, if we do see it, you would also have to look at how this was achieved. If you just train the model and millions or billions of puzzles similar to ARC, so that you're relying on the ability to have some overlap between the tasks that you train on and the tasks that you're going to see at test time, then you're still using memorization, right? And maybe it can work, hopefully ARC is going to be good enough that it's going to be resistant to this sort of attempt and brute forcing. But you never know, maybe it could happen. I'm not saying it's not going to happen. ARC is not a perfect benchmark.
2:44Maybe it has flaws. Flows, maybe it could be hacked in that way. Hmm. So I guess I'm curious about what would GPT -5 have to do that you're very confident that it's on the past AGI? What would make me change my mind about LLM's is basically, if I start seeing a critical mass of cases where you show the model with something it has not seen before, a task that's actually novel from the perspective of its training data, something that's not even training data. And if it can actually adapt on the fly, and this is true for LLM's, but really this would catch my attention with any for AI technique out there.
3:25If I can see the ability to adapt to novelty on the fly, to pick up new skills efficiently, then I would be extremely interested. I would think this is on the past AGI. So the advantage they have is that they do get to see everything. Maybe I'll take issue with how much they are relying on that, but let's suppose that they are relying, obviously they're relying on that more than humans do. To the extent that they do have so much indistribution, to the extent that we have trouble distinguishing, whether an example is indistribution or not, well, if they have everything indistribution, then they can do everything that we can do, maybe it's not indistribution for us.
4:03Why is it so crucial that it has to be out of distribution for them? Why can't we just leverage the fact that they do get to see everything? Right. You're asking basically what's the difference between actual intelligence, which is the ability to adapt to things you've not been prepared for, and pure memorization, like reciting what you've seen before. And it's not just some semantic difference. The big difference is that you can never pre -train on everything that you might see at test time, because the world changes all the time. So it's not just the fact that the space of possible tasks is infinite, and even if you're trained on millions of them, you've only seen zero person of the total space.
4:46It's also the fact that the world is changing every day. This is why we, the human species, is developed in the first place. If there was a true thing as a distribution for the world, for the universe, for our lives, then we would not need intelligence at all. In fact, many creatures, many insects, for instance, do not have intelligence. Instead, what they have is they have in their connectum, in their genes, hard -coded programs, behavioral programs that map some stimuli to a proper response. And they can actually navigate their lives, their environment, in where that's very evolutionary fits.
5:27That way, without needing to learn anything. And while if our environment was static enough, predicateable enough, what would have happened is that evolution would have found the perfect behavioral program, a hard -coded static behavioral program. We'd have written it into our genes, we would have a hard -coded brain connectum, and that's what we would be running on. But no, that's not what happened. Instead, we have general intelligence. So we are born with extremely little knowledge about the world, but we are born with the ability to learn very efficiently. And to adapt in the face of things that we've never seen before.
6:03And that's what makes us unique. And that's what is really, really challenging to recreate in machines. I want to rabbit hole in that a little bit. But before I do that, I'm going to overlay some examples of what an arc -like challenge looks like for the YouTube audience. But maybe for people listening on audio, can you just subscribe? What would an example arc challenge look like? Sure. So one arc puzzle, it looks kind of like an IQ test puzzle. You've got a number of demonstration input output pairs. So one pair is made of two grids. So one grid shows you an input. And the second grid shows you what you should produce as a response to that input.
6:43And you get a couple pairs like this to demonstrate the nature of the task, to demonstrate what you're supposed to do with your inputs. And then you get a new test input. And your job is to produce the corresponding test output. You look at the demonstration pairs. And from that, you figure out what you're supposed to do. And you show that you've understood it on this new test pair. And importantly, in order to the sort of like knowledge basis that you need to approach these changes, is you just need core knowledge. And core knowledge is, it's basically the knowledge of what makes an object basic counting, basic geometry, topology, symmetries, that sort of thing.
7:29So extremely basic knowledge, that's the LMS for sure possesses such knowledge, any child possesses such knowledge. And what's really interesting is that each puzzle is new. So it's not something that you're going to find elsewhere on the internet, for instance. And that means that whether it's as a human or as a machine, every puzzle you have to approach it from scratch. You have to actually reason your way through it. You can just fetch the response from your memory. So the core knowledge, one contention here is we are only now getting multimodal models, who because of the data that are trained on, are trained to do spatial reasoning, whereas obviously not only humans, but for billions of years of revolution, we've had our ancestors have had to learn how to understand abstract physical and spatial properties and recognize the patterns there.
8:26And so one view would be in the next year as we gain models that are multimodal native, that isn't just a sort of second class that is an add -on, but the multimodal capability is a priority, that it will understand these kinds of patterns because that's something which is natively, whereas right now what ArcC is, is some JSON string of 1 ,00, 1 ,00, and it's supposed to recognize the pattern there. And even if you showed a human such as a sequence of these kinds of numbers, it would have a challenge making sense of what kind of question you're asking it. So why I wanted to be the case that as soon as we get multimodal models, which were on the path to unlock right now, they're going to be so much better at archetype spatial reasoning.
9:08That's an empirical question, so I guess we're going to see the answer, within a few months. But my answer to that is, you know, all grades, they're just discrete, two -degree of symbols. They're pretty small, like it's not like, if you flatten an image as a sequence of pixels, for instance, then you get something that's actually very, very difficult to parse. That's not true for Arc because the grades are very small. You only have 10 possible symbols. So there's these two degrees that actually vary to flatten as sequences. And transformers, LLM, they're very good at processing the sequences. In fact, you can show that LLM's do fine with processing arc -lag data by simply fine -tuning LLM on some subsets of the tasks, and then trying to test it on small variations of these tasks.
9:59And you see that, yeah, the LLM can encode just fine solution programs for tasks that are seen before. So it does not really have a problem passing the input or figuring out the program. The reason why LLM's don't do well on Arc is really just the unfamiliarity aspect. The fact that each new task is different from every other task. You cannot, basically, you cannot memorize the solution programs. In advance, you have to synthesize a new solution program on the fly for each new task. And that's free with that at themselves struggling with. So before I do more Devils Advocate, I just want to step back and explain why I'm especially interested in having this conversation.
10:43And obviously the million dollar arc prize, I'm excited to actually play with it myself. And hopefully, the Vesuvius Challenge, which was not Friedman's prize for solving decoding scrolls, the winner of that decoding the scrolls from that were buried in the volcano in the Herculaneon library that was solved by a 22 -year -old who was listening to the podcast Look for Ritor. So hopefully, somebody listening will find this challenge intriguing and find a solution. So I'm, and the reason I, I've had on recently a lot of people who are bullish on LLM's. And I've had discussions with them before interviewing you about how to be explained the fact that LLM's don't seem to be natively performing that well on Arc.
11:24And I found their explanations somewhat contrived. And I'll try out some of the reasons on you. But it is actually an intriguing fact that they actually, these are, some of these problems are relatively straightforward for humans to understand. And they do struggle with them if you just input them natively. All of them are very easy for humans. Like any, any smart human should be able to do 90 % and 95 % on Arc. Smart human. A smart human. But even a five -year -old, so with very little knowledge, they could, they could definitely do over 50%. So let's talk about that because you, I agree that smart humans will do very well on this test.
12:04But the average human will probably do, you know, mediocre. Not true. So we actually tried to use the right humans to describe about 85. That was with Amazon Mechanical Turkwerkers. Right. I'm honestly don't know the demographic profile of Amazon Mechanical Turkwerkers, but I imagine just interacting with the platform that Amazon has set up to do remote work. That's not the median human across the planet, I'm guessing. I mean, the broader point here being that, so we see this spectrum in humans where humans obviously have AGI. But even within humans, you see a spectrum where some people are relatively dumber and they'll do perform work on IQ -like tests.
12:44For example, Ravens' Regressive Matrices sees. If you look at how the average person performs on that and you look at the kind of questions that is sort of mid -ermis, half of people will get it right, half of people will get it wrong. Some of them are like pretty trivial. For us, we might think like this is this is kind of trivial. And so humans have AGI, but from relatively small tweaks, you can go from somebody who misses these kinds of basic IQ test questions to somebody who gets them all right, which suggests that actually if these models are doing natively, we'll talk about some of the previous performances that people tried with these models.
13:13But somebody with a jack col with a 240 million parameter model got 35%. It doesn't that suggest that they're on this spectrum that clearly exists within humans and they're going to be saturated at pretty soon? Yeah, so that's the subject of interesting points here. So there is indeed a branch of LLM approaches suspended by a jack col that are doing quite well, that are in fact state of the art. But you have to look at what's going on there. So there are two things. The first thing is that to get these numbers, you need to pre -train your LLM on millions of generated art tasks. And of course, if you compare that to a five -year -old child looking at art for the first time, the child has never done like you did before.
13:57It's never seen something like black and art tasks before. The only overlap between what they know and what they have to do in the test is core knowledge, is knowing about like counting and objects and symmetries and things like that. And still, they're going to do really well. They're going to do much better than the LLM trained on millions of similar tasks. And the second thing that's something to note about the jack col approach is one thing that's really critical to making the model work at all is test time fine tuning. And that's something that's really missing, by the way, from LLM approaches.
14:33Right now, it's that most of the time when you're using an LLM, it's just doing static inference. The model is frozen and you're just prompting it and then you're getting it in answer. So the model is not actually learning anything on the fly. Its state is not adapting to the task at hand. And what jack col is actually doing is that for every test problem is on the fly is fine tuning a version of the LLM for that task. And that's really what's unlocking performance. If you don't do that, you get like 1%, 2%. So basically, something completely negligible. And if you do test time fine tuning and you add a bunch of tricks on top, then you end up with interesting performance numbers.
15:17So I think what it's doing is trying to dress one of the key limitations of LLM's today, which is the lack of active inference. It's actually adding active inference to LLM's and that's working extremely well actually. So that's fascinating to me. That there's so many interesting rabbit holes there. Should I take them in sequence or deal with them all once? Let me just start. So the point you made about the fact that you need to unlock the adaptive compute slash test time compute. A lot of the scale maximalist, I think this will be interesting rabbit hole to explore with you because a lot of the scaling maximalist have your broader perspective in the sense that they think that in addition to scaling, you need these kinds of things like unlocking adaptive compute or doing some sort of RL to get the system to working.
16:03And their perspective is that this is a relatively straightforward thing that will be added at the top, the representations that a scaled up model has greater access to. No, it's not just the technical detail. It's not a straightforward thing. It is everything. It is the important part. And the scale maximalist argument, it boils down to these people, they refer to scaling loss, which is this empirical relationship that you can draw between how much compute you spend on training a model and the performance you're getting on benchmarks. And the key question here, of course, is, well, how do you measure performance?
16:44What it is that you're actually improving by adding more compute and more data. And well, it's benchmark performance. And the thing is, the way you measure performance is not a technical detail. It's not enough to start because it's going to narrow down the set of questions that you're asking. And so accordingly, it's going to narrow down the set of answers that you're looking for. If you look at the benchmarks we're using for an LMS, they are all memorization -based benchmarks. Like sometimes they are literally just knowledge -based, like a school test. And even if you look at the ones that are explicitly about reasoning, you realize if you look closely that it's in order to solve them, it's enough to memorize a finite set of reasoning patterns.
17:35And then you just reapply them. They're like static programs. LMS are very good at memorizing static programs, small static programs. And they've got this sort of like bank of solution programs. And when you give them a new puzzle, they can just fetch the appropriate program applied. And it's looking like it's reasoning, but really it's not doing any sort of on the flight program synthesis. All it's doing is program fetching. So you can actually solve all these benchmarks with memorization. And so what you're scaling up here, like if you look at the models, they are big parametric curves fitted to the data distribution, which I can't understand.
18:15So they're basically this big interpretive databases, interpretive memories. And of course, if you scale up the size of your database and you cram into it more knowledge, more patterns and so on, you are going to be increasing its performance as measured by memorization benchmark. That's kind of obvious. But as you're doing it, you are not increasing the intelligence of the system one bit. You are increasing the skill of the system. You are increasing its usefulness. It's a scope of applicability, but not its intelligence because skill is not intelligence. And that's the fundamental confusion that people run into is that they're confusing skill and intelligence.
19:00Yeah, there's a lot of fascinating things to talk about here. So skill intelligence, interpolation, I mean, okay, so the thing about they're fitting some manifold into that maps the input data. There's a reductionist way to talk about what happens in the human brain that says that it's just axons firing at each other. But we don't care about the reductionist explanation of what's happening. We care about what the sort of meta at the at the macroscopic level, what happens when these things combine as far as the interpolation goes. So, okay, let's look at one of the benchmarks here. There's one benchmark that does great school math and these are problems that a smart high schooler would be able to solve.
19:44It's called GSM 8K and these models get 95 % on these. Like basically, they always need to memorize it. Let's talk about what that means. So here's one question about from that benchmark. So 30 students are in a class. 150 of them are 12 year olds. One third are 13 year old. One tenth are 11 year olds. How many of them are not 11, 12 or 13 years old? So I agree it's like this is not rocket science, right? You can write down on paper how you go through this problem and a high school kid, at least a smart high school kid, should be able to solve it. Now, when you say memorization, it still has to reason through how to think about fractions and what is the context of the whole problem and then combining the different calculations it's doing.
20:24It depends on how you want to define a reasoning. But there are two definitions you can use. So one is I have available a set of program templates. It's like the structure of the puzzle, which can also generate its solution. And I'm just going to identify the right template, which is in my memory. I'm going to input the new values into the template through the program, get the solution. And you could say this is reasoning. And I say, yeah, sure, okay. But another definition you can use is reasoning is the ability to when you're faced with a puzzle, given that you don't have already a program in memory to solve it, you must synthesize on the fly a new program based on bits of pieces of existing programs that you have.
21:09You have to do on the fly program synthesis. And it's actually dramatically harder than just fetching the right memorize program and replying it. So I think maybe we are overestimating the extent to which humans are so sample efficient. They also don't need training in this way where they have to drill in these kinds of pathways of reasoning through certain kinds of problems. So let's take math, for example. Yeah. It's not like you can just show a baby that axioms of set theory and now they know math, right? So when they're growing up, you have to do years of teaching them pre -algebra. Then you got to do a year of teaching them doing drills and going through the same kind of problem in algebra, then geometry, pre -calculus, calculus.
21:51Yeah, absolutely. So training. Yeah. Isn't that like the same kind of thing where you can't just see one example and now you have the program or whatever. You actually had a drill at these models all started drill with a bunch of returning data. Sure. I mean, in order to do on the fly program synthesis, you actually need building blocks to work from. So knowledge and memory actually tremendously important in the process. I'm not saying it's memory versus reasoning in order to do effective reasoning. You need memory. But it sounds like it's compatible with your story that through seeing a lot of different kinds of examples, these things can learn to reason within the context of those examples.
22:30And we can also see within bigger and bigger models. So that was an example of a high school level math problem. Let's say a model that's like smaller than GPT -3 couldn't do that at all as these models get bigger. They seem to be able to pick a bigger. It's not really a size issue. It's more like a trained data issue in this case. Well, bigger models can pick up these kinds of circuits which smaller models apparently don't do a good job of doing this even if you were to train them on this kind of data. Doesn't that just suggest that if you have bigger and bigger models, they can pick up bigger and bigger pathways or more general ways of reasoning?
23:00Absolutely. But then isn't that intelligence? No, no, it's not. If you scale up your database and you keep adding to it more knowledge, more program templates, then sure it becomes more and more skillful. You can apply it to more and more tasks. But general intelligence is not task -facix skill scaled up to many skills. Because there is an infinite space of possible skills, general intelligence is the ability to approach any problem, any skill and then quickly master it using value or data. Because this is what makes you able to face anything you might have ever encountered. This is what makes this is the definition of a generality.
23:38Like generality is not specificity scaled up. It is the ability to apply your mind to anything or to arbitrary things. And this requires, for now, to this requires the ability to adapt, to learn on the fly efficiently. So my claim is that by doing this free training on bigger and bigger models, you are getting that capacity to then generalize very efficiently. Let me give you an example. So your own company Google, in their paper on Gemini 1 .5, they had this very interesting example where they would give in context, they would give the model the grammar book and the dictionary of a language that has less than 200 living speakers.
24:23So it's not in the free training data. And you just give them the dictionary and it basically is able to speak this language and translate to it, including the complex and organic ways in which language is structured. So a human, if you showed me a dictionary from English to Spanish, I'm not going to be able to pick up the how to structure sentences and how to say things in Spanish. The fact that because of the representations that it has gained through this free training, it is able to now extremely efficiently learn a new language. Doesn't that show that this kind of training actually does increase your ability to learn new tasks?
24:57If you're right, if you were right, LLMs would do really well on archbuzzles because archbuzzles are not complex. Each one of them requires very tool knowledge. Each one of them is very low on complex. You don't need to think they're hard about it. They're actually extremely obvious for humans like even children can do them. But LLMs cannot, even LLMs that have, you know, 100 ,000 times more knowledge than you do, the still cannot. And the only thing that makes arch special is that it was designed with this intent to resist the emoization. This is the only thing and this is the huge blocker for a length performance, right?
25:35And so, you know, I think if you look at LLMs closely, it's pretty obvious that they're not really synthesizing new programs on the fly to solve the tasks that they're faced with. They're very much replying things that they've stored in memory. For instance, one thing that's very striking is LLMs can solve Cesar Cypher, you know, like a Cesar Cypher, like transposing letters to code a message. And well, there's a fairly complex algorithm, right? But it comes up quite a bit on the internet. So, they've basically memorized it. And what's really interesting is that they can do it for a transposition length of like three or five because they're very, very common numbers in examples, provide an internet.
26:22But if you try to do it with an arbitrary number, like nine, it's going to fail. Because it does not encode the generalized form of the algorithm, but only specific cases. It does memorize specific cases of the algorithm, right? And if it could actually synthesize on the fly the solver algorithm, then the value of n would not matter at all. Because it does not increase the product complexity. I think this is true of humans as well. Where what was the study that humans use memorization pattern matching all the time, of course, but humans are not limited to memorization pattern matching. They have this very unique ability to adapt to new situations on the fly.
27:01This is exactly what enables you to navigate every new day in your life. I forget the details, but there was some study that chess grandmasters will perform very well within the context of the moves that accident example, because chess at the highest level is all about memorization chess memorization. Okay, sure. We can leave that aside. What is your explanation for the original question of why can why in context the GPT one, sorry, Gemini 1 .5 was able to learn a language, including the complex grammar structure. Doesn't that show that they can pick up new knowledge? I would assume that it has simply a mind from its extremely extensive and imaginably vast training data.
Read the full transcript
27:40It has mind the required template and then it's just reusing it. We know that there are a very poor ability to synthesize new programs, templates like this on the fly, or even adapt existing ones. They're very much limited to fetching. Suppose there's a programmer at Google, they could go into the office in the morning. At what point are they doing something that 100 % cannot be due to fetching some template that could, even if they, suppose they were an LLM, they could not do if they had fetched some template from their program. At what point do they have to use this so -called externalization capability?
28:11Forget about Google software developers. Every human every day of their lives is full of novel things that have not been prepared for. You cannot navigate your life based on memorization alone. It's possible. I'm sort of denying the premise that they're, you are also great, they're not doing like quote -unquote memorization. It seems like you're saying they're less capable of generalization, but I'm just curious of the kind of generalization they do. If you get into the office and you try to do this kind of generalization, you're going to fail at your job. What is the first point? You're a programmer.
28:44What is the first point when you try to do that generalization? You would lose your job because you can't do the extreme generalization. I don't have any specific examples, but literally, like, take this situation for instance. You've never been here in this room. Maybe you've been in this city a few times, I don't know, but there's a firm on the novelty. You've never been interviewing me. There's a firm on the novelty every hour of every day in your life. It's in fact, by and large, more novelty than any LLM could handle. If you just put LLM in a robot, it could not be doing all the things that you've been doing today.
29:25Take a driving car, for instance, you take a driving car operating in the barrier. Do you think you could just drop it in New York City or drop it in London, where people drive on the left? No, it's going to fail. Not only can you drop, not make it generalize to a change of rules of driving rules, but you can not even make it generalize to a new city. It is a new city. I mean, I agree that self -driving cars aren't age -y. But it's the same type of model that transformers as well. It's also have brains with neurons in them, but they're less intelligent because they're small. We can get into that.
30:07But so I still don't understand a concrete thing of, we also need training. That's why education exists. That's why I just spent the first 18 years of our life doing drills. We have a memory between all not and memory. We are not limited to just the memory. I have been on the firm and said that's necessarily the only thing these models are doing. I'm still not sure what is the task that a remote worker would have to, like, suppose you just have that remote work with an LLM and their programmer. What is the first point that I wish you realized? This is not a human. This is an LLM. What about I just send them a knock puzzle and see how they do?
30:43No, like part of their job, but you have to deal with novelty all the time. If there were more than all the programmers that were replaced and then we're still saying they're only doing memorization, latent programming tasks, but they're still producing a trillion dollars of worth of output in the form of code. Software development is actually a very good example of a job where you're dealing with novelty all the time. Or if you're not, well, I'm not sure what you're doing. So I personally use genetic data of the LLITOL in my software development job. Before an LLM source thing, I was also using Stack Overflow, the LLITOL.
31:21Some people maybe are just copy -pasting stuff from Stack Overflow on our disk to be pasting stuff from an LLM. Personally, I try to focus on problem solving. The syntax is just a technical detail. It was really important is the problem solving. The sense of programming is engineering mental models, like mental representations of the problem you're trying to solve. You can, you know, we have many people can interact with these systems themselves and you can go to chat GPT and say, here's the specification of the kind of program I want. They'll build it for you. As long as there are many examples of this program on LLITOL and Stack Overflow and so on, they will fetch the program for you from their memory.
32:03But you can change arbitrary details. No, it doesn't work. I need it to work on this different kind of server. If that's well true, there would be no software engineers to that. I agree we're not at a full -age AI yet in the sense that these models have, let's say, less than a trillion parameters. A human brain has somewhere on the order of 10 to 30 trillion synapses. I mean, if you were just doing some naive math, you're at least 10x under parameterized. So I agree we're not there yet. But I'm sort of confused on why we're not on the spectrum where, yes, I agree that there's many kinds of general addition they can't do.
32:37But it seems like they're on this kind of smooth spectrum that we see even within humans, where some humans would have a hard time doing an arc type test. We see that, based on their performance on progressive Ravens, matrices type IQ tests. I'm not a fan of IQ tests because for the most parts, you can train all IQ tests and get better at them. So they have very much memorization -based. And this is actually the main pitfall that arc tries not to fall. I'm so lucky. So if all remote jobs are automated in the next five years, let's say, at least that don't require you to be like sort of a service.
33:11It's not like a sales person where you need, you want the human to be talking, but like for example, whatever, in that world, would you say that that's not possible because a lot of what a programmer needs to do definitely requires things that would not be in any pre -training core. But in five years, there would be more self -to -engineers than all today. But I just want to understand. So I'm not sure, I mean, I know how to, I started computer science. If I had become a code monkey out of college, like, what would I be doing? I go to my job. What is the first thing my boss tells me something to do?
33:42When does he realize I'm an LLM? If I was an LLM? Probably on the Jose, you know. Again, if it were true that LLM is good to generalize to novel problems like this and you can actually develop software to solve a problem they've never seen before. You would not need self -to -engineers anymore. In practice, if I look at how people are using LLM in their self -to -engineering job today, they're using it as a stack of a flow replacement. So they're using it as a way to copy paste, code snippets to perform Vy column actions. And this is what they actually need is a database of code snippets. They don't actually need any of the abilities that actually make them self -to -engineers.
34:27When we talk about interpolating between stack overflow databases, if you look at the kinds of math problems or coding problems, maybe to say that they're, maybe let's step back on interpolation and let me ask the question this way. Why can't creativity, why isn't creativity just interpolation in a higher dimension where if a bigger model can learn of more complex manifold, if we're going to use the MLA language, and if you look at read a biography of a scientist, right, it doesn't feel like they're not zero -shotting new scientific theories. They're playing with existing ideas. They're trying to juxtapose them in their head.
35:01They try out some slightly ever in the tree of intellectual descendants. They try out a different evolutionary path. You sort of run the experiment there in terms of publishing the paper, whatever. It seems like a similar kind of thing humans are doing. There's like at a higher level of generalization. And what you see across bigger and bigger models is they can, they seem to be approaching higher and higher level generalization. Where GBT2 couldn't do a great school level math problem that requires more generalization that it has capability for, even that skill, then GBT3 and 4 can. So not quite.
35:32So in GPT4, as a higher degree of skill and higher range of skills, because the same decryptionalization. I don't want to get into the matrix here, but the question of why can't creativity be just interpolation on a higher dimension? I think interpolation can be creative, absolutely. And to your point, I do think that on some level, humans also do a lot of memorization, a lot of reciting, a lot of pattern matching, a lot of interpolation as well. So it's very much a spectrum between pattern matching and true reasoning. It's a spectrum. And humans are never really at one end of the spectrum. They are never doing pure pattern matching or pure reasoning.
36:16They're usually doing some mixture of both. Even if you're doing something that's in vain, reasoning has you like proving a mathematical theorem. As you're doing it, sure you're doing quite a bit of discrete search in your mind, quite a bit of actual reasoning. But you're also very much guided by intuition, guided by pattern matching, guided by the shape of proofs that you've seen before, by your knowledge of mathematics. So it's never really, you know, all of our thoughts, everything we do is a mixture of these sort of like interpollety memorization based thinking, these sort of like type one thinking and type two thinking.
36:56Why are bigger models more sample efficient? Because they have more reusable building blocks that they can lean on to pick up new patterns in their training data. And does that pattern keep continuing as you keep getting bigger and bigger? To the extent that the new patterns, you're giving the model to learn, are good match from what it has learned before. If you present something that's actually novel, that is not in a state of distribution, like an archbazole, fun, sense, eat, we'll fight. Let me make this claim. The program, and this is I think is a very, very useful intuition pump. Why can't it be the case that what's happening in the transformer is the early layers are doing the figuring out how to represent the inputting tokens.
37:38And what the middle layers do is this kind of program search, programs, and this is where they combine the inputs to the, to the, you know, all the circuits in the model where they, they go from the low level representation to a higher level representation during the middle model, they use these programs and they do, they combine these concepts, then what comes out the other end is the reasoning based on that high level intelligence. Possibly, why not? But you know, if these models were actually capable of synthesizing novel programs, however simple they should be able to do arch, because for any arch task, if you write down the solution program in Python, it's not a complex program.
38:20It's extremely simple. And humans can figure that. So why can't I learn to not do it? Okay, I think that's a fair, fair point. And if I turn the question around to you, so suppose that it's the case that in a year, a multimodal model can solve arch, let's say, get 80 percent, whatever the average human will get, then AGI? Quite possibly, yes. I think if you, if you start, honestly, what I would like to see is a NLM type model solving arc at like 80 percent, but after having only been trained on core knowledge related stuff. But human kids, I don't think we're necessarily illustrated none. It's not just that we have an RG, let me erase that.
39:06Only trained on information that is not explicitly trying to anticipate what's going to be in the arc test set. But it isn't the whole point of arc that you can't sort of, it's a new type of intelligence that every single time. Yes, that is the point. So if arc were perfect, flawless benchmark, it would be impossible to anticipate what's in the test set. And you know, arc was released more than four years ago, and so far it's been resistant to memorization. So I think it has to some extent passed a test of time. But I don't think it's perfect. I think if you try to make by hand hundreds of thousands of arc tasks and then you try to multiply them by programmatically generating evaluations and then you end up with maybe hundreds of millions of tasks, just by brute forcing the task space, there will be enough overlap between what you're trying on and what's in the test set that you can actually score very highly.
40:02So you know, with enough scale, you can always cheat. If you can do this for every single thing that's supposedly requires intelligence, then what good is intelligence? Apparently you can just brute force intelligence. If the world, if your life were a static distribution, then sure you could just brute force the space of possible behaviors. You know, the way we think about intelligence, there are several metaphors to use. But one of them is you can think of intelligence as a past finding algorithm in future situation space. Like I know if your family always came development like RTS came development, but you have a map.
40:39And you have a 2D map. And you have partial information about it. There is some further four on your map. There are areas that you haven't explored yet. You know nothing about them. And then there are areas that you've explored, but you only know how they were like in the past. You don't know how they're like today. And now instead of thinking about a 2D map, think about the space of possible future situations that you might encounter and how they're connected to each other. Intelligence is past finding algorithms. So once you set a goal, it will tell you how to get there optimally. But of course, it's constrained by the information you have.
41:21It cannot pass fine in an area that you know nothing about. It cannot also anticipate changes. And the thing is, if you had complete information about the map, then you could solve the past finding problem by simply memorizing every possible path, every mapping from point A to point B. You could solve the problem with pure memory. But the reason you cannot do that in real life is because you don't actually know what's going to happen in the future. Life is ever changing. I feel like you're using words in real memorization, which we would never use for human children. If you're like your kid learns to do algebra and then like now learns to do calculus, you wouldn't say they've memorized calculus.
42:05If they can just solve any arbitrary algebraic problem, you wouldn't say like they've memorized algebra. They say they've learned algebra. Humans are never really doing pure memorization up pure reasoning. But that's only because you're so man -to -leavly labeling when the human does the skill, it's a memorization. When the exact same skill is done by the LLM, as you can measure by these and you can just plug in any sort of map problem. Sometimes humans are doing the exact same as the LLM is doing, which is just, for instance, I know if you learn to add numbers, you're memorizing an algorithm, you're memorizing a program and then you can reapply it, you are not synthesizing on the fly the addition program.
42:37So obviously at some point some human have to figure out how to do addition. But like the way a kid learns it is not that they sort of figure out from the accents of that theory how to do addition. I think what you learn is that it means school is mostly memorization. So my claim is that listen, these models are vastly under parameterized. It's relative to how many flops or how many parameters you have in the human brain. And so yeah, they're not going to be like coming up with new theorems, the smartest humans can. But most humans can't do that either. What most humans do, it sounds like a similar to what you were calling memorization, which is memorizing skills or memorizing techniques that you've learned.
43:17And so it sounds like it's compatible. And let's tell me if this is wrong. Is it compatible in your world if like all the remote workers are gone, but they're doing skills which we can potentially make synthetic data off. So we record everybody's screen and every single remote worker screen, we sort of understand the skills they're performing there. And now we've trained a model that can do all this. All the remote workers are unemployed, regenerating trillions of dollars of economic activity for me, I remote workers. In that world, is there still a memorization regime? So sure, with memorization you can automate almost anything.
43:49As long as it's a static distribution, as long as you don't have to deal with change. Our most jobs part of such a static distribution. Potentially there are lots of things that you can automate. And LLM is an excellent tool for automation. And I think that's, but you have to understand that automation is not the same as intelligence. I'm not saying that LLM's are useless. I've been a huge proponent of deep learning for many years. And for many years I've been saying two things. I've been saying that if you keep scaling up deep learning, it will keep paying off. And at the same time, I've been saying if you keep scaling up deep learning, this will not lead to a GI.
44:24So we can automate more and more things. And yes, this is economically valuable. And yes, potentially there are many jobs. You could automate a well like this. And that would be economically valuable. But you're not still not going to have intelligence. So you can ask, you know, okay, so what does it matter if you can generate all this economic value? Maybe you don't need intelligence after all. Well, you need intelligence. The moment you have to deal with change, with novelty, with uncertainty. As long as you're in a space that can be exactly described in advance, you can just, you can just automate your pure memorization.
44:57In fact, you can always solve any problem. You can always display arbitrary levels of skills on any task without leveraging any intelligence whatsoever. As long as it is possible to describe the problem in its solution very, very precisely. But when they do deal with novelty, then you just call it interpolation, right? And so No, no, no, interpolation is not enough to deal with all kinds of novelty if it were then LLM's would be with BGI. Well, I agree they're not a GI. I'm just trying to figure out how do we figure out we're on the path to GI. And I think sort of correct here is maybe that it seems to me that these things are on a spectrum and we're clearly covering the earliest part of the spectrum with LLM's.
45:43I think so. And oh, okay, interesting. But here's another sort of thing that I think is evidence for this. Grocking, right? So clearly even within deep learning, there's a difference between the memorization regime and the generalization regime where at first, they'll just memorize the data set of, you know, if you're doing modular edition, how to add digits. And then at some point, if you keep training on that, they'll learn the skill. So the fact that there is a distinction suggests that the generalized circuit, the deep learning can learn, there's a regime in enters where it generalizes. If you have an over -parameterized model, which you don't have in comparison to all the tasks we want these models to do right now.
46:20Grocking is very, very old phenomenon. We've been observing it for decades. It's basically an instance of the minimum description length principle where, sure, you can, given a problem, you can just memorize a point -wise input to output mapping, which is completely over -fit. So it does not generalize at all, but it solves the problem on the train data. And from there, you can actually keep improving it, keep making your mapping simpler and simpler and more compressed. And at some point, it will start generalizing. And so that's something called the minimum description length principle. It's decided that the program that will generalize best is the shortest.
47:06Right. And it doesn't mean that you're doing anything other than memorization, but you're doing memorization plus regularization. Right. Okay, generalization. Yeah. And that is absolutely at least to generalization. Right. And then so you do that within one skill, but then the pattern you see here of metal learning is that it's more efficient to store a program that can perform many skills rather than one skill, which is what we might call fluid intelligence. And so as you get bigger and bigger models, you would expect it to go up this hierarchy of generalization where it generalizes to a skill, then it generalizes across multiple skills.
47:38That's correct. That's correct. And at an level, they're not infinitely large. They have only a fixed number of parameters. And so they have to compress their knowledge as much as possible. And in practice, so at a level, they're mostly storing reusable bits of programs like vector programs. And because they have this need for compression, it means that every time they're learning a new program, they're going to try to express it in terms of existing bits and pieces of programs that they've already learned before. And right. Isn't this the generalization? Absolutely. Oh, wait. So this is what, you know, clearly, I'll have some degree of generalization.
48:17Yeah. And this is precisely why is because they have to compress. And why is that intrinsically limited? Why can't you just go it at some point, it has to learn a higher level of generalization, higher level, and then the highest level is the fluid intelligence. It's intrinsically limited because the substrate of your model is a big parametric curve. And all you can do with this is local generalization. If you want to go beyond this, to else broader or an extreme generalization, you have to move to a different type of model. And my paradigm of choice is discrete program search, program synthesis.
48:51So and if you want to understand that, you can sort of like compare it, contrast it with deep planning. So in deep planning, your model is a parametric, a differential world parametric curve. In programs synthesis, your model is discrete graph of operators. So you've got like a set of logical operators like a domain specific language. You're picking instances of it, you're structuring that into a graph. That's a program. And that's actually very similar to like a program you might write in Python or C++ and so on. And in deep planning, you're learning engine because we are doing mesh learning here.
49:29Like we're trying to automatically learn these models. In deep planning, your learning engine is quite in the sense. Right. And quite in the sense is very compute efficient because you have this very strong informative feedback signal about where the solution is. So you know that to make it work, you need a dense sampling of the operating space. You need a dense sampling of the data distribution. And then you're limited to any generalizing within that data distribution. And the reason why you have this limitation is because your model is a curve. And meanwhile, if you look at discrete program search, the learning engine is combinatorial search.
50:13You're just trying a bunch of programs until you find one that actually meets your spec. This process is extremely data efficient. You can learn and generalize a whole program from just one example, two examples, which is why it works so well on arc, by the way. But the big limitations that it's extremely compute inefficient because you're running into combinatorial explosion, of course. And so you can sort of see here how deep planning and discrete program search, they have very complimentary strength and limitations as well. Like every limitation of deep planning has a strength, corresponding strengths in program synthesis and in university.
50:53And I think the past four one is going to be too merged to to basically start doing. So another way you can think about it is so these these parometric curves trained with ground descent, they're great fits for everything that's a system one type thinking like pattern cognition intuition, memorization and so on. And discrete program search is a great fit for type two thinking system two thinking for instance, planning reasoning, a quickly figuring out generalizable model let matches just one or two examples like for an archbuzz or for instance. And I think humans are never doing pure system one or pure system to they're always mixing and matching both.
51:40And right now we have all the tools for system one, we have almost nothing for system two. The way for one is to create a hybrid system. And I think the form is going to take is it's going to be most this system to so the the outer structure is going to be a discrete program search system. But you're going to fix the fundamental limitation of discrete program search, which is combinatorial explosion. You're going to fix it with deep learning. You're going to leverage deep learning to guide to provide intuition in program space to guide the the program search. And I think that's very similar to what you see for instance when when when you're playing chess or when you're trying to prove a theorem is that it's mostly a reasoning thing, but you start out with some intuition about the shape of the solution.
52:31And that's very much something you can get via a deep learning model via deep learning models. They're very much like intuition machines, they're pattern matching machines. So you you start from this shape of the solution and then you're you're going to do actual explicit discrete program search. But you're not going to do it via brute force. You're not going to try things kind of like chronically. And you're actually going to ask another deep learning model for suggestions. Like here's the best likely next step. Here's where in the graph you should be going. And you can also use yet another deep learning model for feedback.
53:11But well, here's what I had so far is it looking good should just backtrack and try something new. So I think this quick program search is going to be the key, but you want to make it dramatically better. All those of magnitude more efficient by leveraging deep learning. And by the way, another thing that you can use deep learning for is of course things like command sense knowledge and knowledge in general. And I think you're going to enter with this sort of system where you have this on the fly synthesis engine that can adapt to new situations. But the way it adapts is that it's going to fetch from a bank of patterns modules that could be themselves curves that could be a differentiable modules and some others that could be algorithmic in nature.
54:00It's going to assemble them via this process that's intuition guided. And it's going to give you for every new situation you might be faced with is going to give you with a generalizable model that was synthesized using very, very tool data. Something like this would sort of arc. That's actually really interesting, uh, uh, prompt because I think an interesting crux here is when I talk to my friends who are extremely optimistic about LLM's and expect AGI within the next couple of years, they also in some sense agree that scaling is not all you need, but that the rest of the progress is undergirded and enabled by scaling and but still you need to add the system to the test time compute, uh, top these models.
54:51And their perspective is that it's relatively straightforward to do that because you have this library or representations that you built up from free training, but it's almost talking like, you know, it's just like skimming through text books. You need some more deliberate way in which it engages with the material it learns. In context learning is extremely sample efficient, but to actually distill that into the weights, you need the model to like talk through the things that sees and then add it back to the weights. As far as the system two goes, they talk about adding some kind of RL setup so that it is encouraged to proceed on the reasoning traces that end up being correct.
55:29And they think this is a relatively straightforward stuff that will be added within the next couple of years. That's an impactful question. Yeah. So I think we see your intuition. I assume is not that. I'm intuition is in fact, this whole like system to architecture is the hard part is the very hard and not obvious part scaling up the interpretive memory is the easy part. All you need is is like it's literally just a big cup. All you need is more data. It's representation of a dataset interpretive representation of the dataset. That's the easy part. The hard part is the architecture of intelligence.
56:02Memory and intelligence are separate components. We have the memory. We don't have the intelligence yet. And I agree with you that well, having the memory is actually very useful. And if you just had the intelligence, but it was not hooked up to an extensive memory, it would not be that useful because it would not have enough material to work from. Yeah. The alternative hypothesis here that a former guest Trenton Brookin advanced is that intelligence is just high -arkly associated memory where higher level patterns when Sherlock Holmes goes into a crime scene and he's extremely sample efficient.
56:36He can just look at a few clues and figure out who was a murderer. And the way he's able to do that is he has learned higher level sort of associations. It's memory in some fundamental sense. But so here's one way to ask a question. In the brain, I suppose the Louis Vu program synthesis, but it is just synapses connected to each other. And so physically it's got to be that you just query the right circuit. It's matter of degree. But if you can learn it, if training in the environment that the human ancestors are trained in means you learn those circuits, training on the same kinds of outputs of humans produce, which to replicate require these kinds of circuits, wouldn't that train the same kind of whatever humans have?
57:19You know, it's matter of degree. If you have a system that has a memory and is only capable of doing local generalization from that, it's not going to be very adaptable. To be really general, you need the memory plus the ability to search to quite some depth to achieve, you know, broader even extreme generalization. You know, like one of my favorites, a psychologist, so Jean -Pierre was the founder of the Elemental Psychology. He had a very good quote about intelligence. He said, intelligence is what you use when you don't know what to do. And it's like, as a human living your life, in most situations you already know what to do because you've been in this situation before, you already have the answer.
58:09And you're only going to need to use intelligence when you're faced with novelty, with something you didn't expect, with something that you weren't prepared for either by your own experience, your own life experience, or by your evolutionary history. This day that you're living right now is different in some important ways from every day you've lived before, but it's also different from any day ever lived by any of your ancestors. And still, you're capable of being functional. How is it possible? I'm not denying that generalization is extremely important and is the basis for intelligence. That's not the correct, the correct, is how much of that is happening in the models.
58:48But let me ask a separate question. We might keep going in the circle here. The difference is intelligence between humans. Maybe the intelligence tests because of reasons you mentioned are not measuring it well, but clearly there's differences in intelligence between different humans. What is your explanation for what's going on there? Because I think that's sort of compatible with my story that there's a spectrum of generality and that these models are climbing up at two human level. And even some humans have it even climbed up to the Einstein level or the Francois level. That's a great question.
59:20There is extensive that intelligence, difference intelligence are mostly genetic in nature, right? Meaning that if you take someone with not the intelligence, there is no amount of training of like training data. You can expose that person to that would make them become Einstein. And this kind of points to the fact that you really need a better architecture. You need a better algorithm. And more training data is not in fact all you need. I think I agree with that. I think maybe the way I might phrase it is that the people who are smarter have an ML language better initializations. It just the newer wiring if you just look at it's more efficient.
1:00:04They have maybe greater density of firing. And so as some part of the story is scaling, there are some correlations between brain size and intelligence. And we also see within the context of quote unquote scaling that people talk about within context of LLMS, architectural improvements where a model like Gemini 1 .5 Flash is performs as well as GPT -4 did when GPT -4 was released a year ago, but is 57 times cheaper on output. So the part of the scaling stories that the architectural improvements are we're in like extremely low -hanging fruit territory when it comes to those. Okay, we're back now with the co -founder of Zapier, Mike Knuff.
1:00:45We had to restart a few times there. And you're funding this prize and you're running this prize with Francois. And so tell me about how this came together. What prompted you guys to launch this prize? Yeah. I guess I've been sort of like AI curious for 13 years. I've been, I co -founded Zapier been running it for the last 13 years. And I think I first got interviews to year work and during COVID. I kind of went down the rabbit hole. I had a lot of free time. And it was right after you published your on -measured intelligent paper, you sort of interviews the concept of AGI. This like efficiency of skill acquisition is like the right definition and the arc puzzles.
1:01:22But I don't think the first Kaggle contest was done yet. I think it was still running. And so I kind of it was interesting. But I just parked the idea. And I had bigger fish to fry. It's Zapier. We're in this middle of this big turnaround of trying to get to our second product. And then it was January 2022 when the chain of thought paper came out that really like awoken me to sort of the progress. I gave a whole presentation to the Zapier on like the GPT -3 paper events. I sort of felt like I had priced in everything that Elms could do. And that paper was really shocking to me in terms of all these.
1:01:53There's latent capabilities that Elms have that I didn't expect that they had. And so I actually gave up my exact team role. It's app. I was running half the company in that point. I went back to be an individual contributor and just do to go do AI research alongside Brian, my co -founder. And all of a sudden that led me to back towards arc. I was looking into it again. And I had sort of expected to see this saturation effect that MMOU has, that GMSK has, 8K has. And when I looked at the scores and the progress that since the last four years, I was really again shocked to see actually we've made very little objective progress towards it.
1:02:30And it felt very, it felt like a really, really important evalon. As I sort of spent the last year asking people, quizzing people about it in sort of my network and community, very people, few people even knew it existed. And that felt like, okay, if it's right that this is a really, really globally, singularly unique, EGI eVAL. And it's different from every other eVAL that exists that are more narrowly measures, AI skill, like more people should know about this thing. I had my own ideas on how to beat the arc as well. So I was working on it some weekends on that and I flew up to meet Francois earlier this year to sort of quiz him, show my ideas.
1:03:08And ultimately, I was like, well, why don't you think more people know about arc? I think you should actually answer that. I think it's a really interesting question. Like why don't you think more people know about arc? Sure. You know, I think benchmarks that gain traction in the research community are benchmarks that are already fairly tractable because the dynamic that you see that some research group is going to make some initial breakthrough. And then this is going to catch the attention of everyone else. And so you're going to get photo papers with people trying to beat the first team and so on.
1:03:38And for arc, this has not really happened because arc is actually very hard for existing AI techniques. Kind of arc requires you to try new ideas. And that's very much the point, by the way. Like the point is not that, yeah, you should just be able to apply existing technology and solve arc. The point is that existing technology has reached a plateau. And if you want to go beyond that, if you want to start being able to tackle problems that you haven't memorized that you haven't seen before, you need to try new ideas. And arc is not just meant to be this sort of like measure of hack loads we are to a GI.
1:04:17It's also meant to be a source of inspiration. I want researchers to look at this puzzle and be like, hey, it's really strange that these puzzles are so simple. And most humans can just do them very quickly. Why is it so hard for existing AI systems? Why is it so hard for all the lamps and so on? And so it's true for our lamps, but arc was actually released before lamps we have really a thing. And the only thing that made it special at the time was that it was designed to be a resistance to memorization. And the fact that it has survived at a lamps and genuine and general so well, it kind of shows that yes, it is actually a resistance to memorization.
1:04:56This is what nerds night me because I went and took a bunch of the puzzles myself. I've like super easy. Are you sure AI can't solve this? That's the reaction in the same one for me as well. And the more you dig in, you're like, okay, yes, there's not just empirical evidence over the last four years that it's unbeaten, but there's theoretical concepts behind why. And I completely agree at this point that new ideas basically are needed to be dark. And there's a lot of current trends in the world that are actually I think working against that happening. Basically, I think we're actually less likely to generate new ideas right now.
1:05:32I think one of the trends is the closing up front to your research. The GP4 paper from opening, I had no technical detail shared. The Gemini paper had no technical detail shared, and like the longer context part of that work. And yet that open innovation, open progress and sharing is what got us to transformers in the first place. That's what got us to elements in the first place. So it's kind of disappointing a little bit actually that like so much front to your work has gone closed. It's really making a bet that like these individual labs are going to have the breakthrough and not the ecosystem is going to have the breakthrough and the sort of the internet open source has shown that that's like the most powerfully innovation ecosystem that's ever existed probably in the entire world.
1:06:08I think that's actually really sad front to your research is no longer being published. If you look back four years ago, well, everything was just open and shared like all the set of the art results were published. And it's not like the case. And it's very much open AI single -handedly changed the game. And I think opening up basically set back progress towards HDI by quite a few years, probably like five to ten years for two reasons. And one is that well, the cause this complete closing down of research, frontier research publishing. But also the trigger this initial burst of hype around LLMs.
1:06:52And now LLMs have sucked the oxygen out of the room like everything everyone is just doing LLMs. And I see LLMs as more than off ramp on the path to a GR actually. And all these new resources, they're actually going to LLMs instead of everything as they could be going to. And if you look further into the past to 2015, 2016, there were like a thousand times fewer people doing AI back then. And yet I feel like the rate of progress was higher because people were exploring more directions. The world felt more open -ended like you could just go and try like have a cool idea of a launch and try it and get some interesting results.
1:07:38So there was this energy. And now everyone is very much doing some variation of the same thing. And the big labs also tried the handle arc, but because they got bad results, they didn't publish anything. Like, you know, people only publish positive results. I wonder how much effort people have put into trying to prompt or scaffold, do some sort of maybe dev and type approach into getting the frontier models and frontier models of today, not just a year ago because a lot of post training has gone into making them better. So cloud through OPS or GPD40 into getting good solutions on arc. I hope that one of the things this episode does is get people to try out this open competition where they have to put in an open source model to compete, but also to like figure out if they're maybe the lake capability is latent in cloud opus and just see if you can show that.
1:08:36I think that would be super interesting. So let's talk about the prize. How much do you win if you solve it? You know, get whatever percent on arc. How much do you get if you get the best division, but don't crack it? So we got a million dollar, actually a little over a million dollars of the price pool running the contest on an annual basis. We're going to, we're starting it today through the middle of November. And the goal is to get 85%. That's the lower bound and human average that you guys talked about earlier. And there's a $500 ,000 prize for the first team that can get to the 85 % benchmark.
1:09:07We're also going to run, we don't expect that to happen this year, actually. One of the early statisticians that's happier, gave me this line that has always stuck with me that the longer it takes, the longer it takes. So my prior is that like arc is going to take years to solve. And so we're going to keep, we're also going to break down and do a progress price this year. So we're, there's a $100 ,000 progress price, which we will pay out to the top scores. So $50 ,000 is going to go to the top objective scores this year on the Kaggle leaderboard, which is we're hosting it on Kaggle. And then we're going to have a $50 ,000 pot set for a paper award for the best paper that explains conceptually the scores that they were able to achieve.
1:09:46And one of the, I think, interesting things we're also going to be doing is we're going to be requiring that in order to win the prize money that you put the solution or your paper out into public domain. The reason for this is, you know, tend to typically with contests, you see a lot of like closed up sharing people are kind of private secret. They want to hold their alphabet of themselves during the contest period. And because we expect it's going to be multiple years, we want to enter a game here. So the plan is, you know, at the end of November, you will award the $100 ,000 prize money to the top progress prize and then use the down time between December, January, February to share out all the knowledge from the top scores and the approaches folks were taking it in order to rebase line the community up to whatever the state of the art is and then run the contest again next year and keep doing that on a yearly basis until we get 85%.
1:10:32I'll give some people some context on why I think this prize is very interesting. I was having conversations with my friends who are very much believers and models that exist today. And first of all, it was intriguing to me that they didn't know about ARC. These are experienced ML researchers. And so you show them the, this has happened a couple of nights ago. We went to dinner and I showed them an example problem. And they said, of course, an LLM would be able to solve something like this. And then we take a screenshot of it. We just put it into our chat GPT app and it doesn't get the pattern.
1:11:02And so I think it's very interesting. Like it is a notable fact. I was sort of playing that was advocate against you on these kinds of questions. But this is a very intriguing fact in them. I think this is as price is extremely interesting because we're going to learn, we're going to learn something fascinating one way or another. So with regards to the 85%, separate from this prize, I'd be very curious if somebody could replicate that result because obviously in psychology and other kinds of fields, which this result seems to be analogous to when you run tests on some small sample of people, often they're hard to replicate.
1:11:36So I'd be very curious if you try to replicate this. What does it average human perform on ARC? Ask for the difficulty on how long it will take to crack this benchmark. It's very interesting because the other benchmarks that are fully saturated like MMUL math. Actually, the people who made them, Dan Hendrix and Colin Burns who did MMUL math, I think they were grad students or college students when they made it. And the goal when they made it just a couple of years ago was that this will be a test of AGI. And of course, it got totally saturated. I know you'll argue that these are test and memorization.
1:12:08But I think the pattern we've seen, in fact, Epoch AI has a very interesting graph that I'll overlay for the YouTube version here where you see this almost exponential where it gets 5%, 10%, 30%, 40%, as you increase the compute across models and then it just shoots up. And in the GBT -4 technical report, they had this interesting graph of the human Epoch problem set, which was 22 coding problems. And they had to graph it on the mean log pass curve, basically because it early on in training, or even smaller models, can have the right idea of how to solve this problem. But it takes a lot of reliability to make sure they stay on track to solve the whole problem.
1:12:52And so you really want to upway the signal where they get a right at least some of the time. We want it 100 times, we want it a thousand. And then so they go from like 1 ,100, 110, and then they just like totally saturated. I guess the question I have, this is all leading up to, is why won't the same thing happen with ARC where people had to try really hard bigger models. And now they figured out these techniques that Jack Colis figured out with only a 240 million parameter language model that can get 35%. Shouldn't we see the same pattern we saw across all these other benchmarks, where you're just like sort of eke out and then once you get the general idea, then you just go all the way to 100.
1:13:27That's an empirical question. So we've seen practice with happens. But what Jack Colis is doing is actually very unique. It's not just portraying an alarm and then prompting it. It's actually trying to do active inference. You do it. It's time -fantuning. And this is actually trying to lift one of the key limitations of the alarms, which is that at inference time they cannot learn anything. They cannot adapt on the fly to where they are and is actually trying to learn. So what he's doing is effectively a form of program synthesis. Because the alarm contains a lot of useful building blocks, like programming building blocks, and by fancy units on the task at test time, you are trying to assemble these building blocks into the right pattern that matches the task.
1:14:15This is exactly what programs this is about. The way we contrast this approach with discrete program search is that in discrete program search, so you're trying to assemble a program from a set of primitives. You have very few primitives. So people working on discrete program search on Arc, for instance, they tend to work with DSLs that have like 100 to 200 primitive programs. So very small DSL, but then they're trying to combine these primitives into very complex programs. So there's very deep depths of search. And on the other hand, if you look at what Jack will is doing with the alarms, is that he has got this sort of like vector program database, DSL, of millions of building blocks in the alarm that are mined by pre -training the alarm, not just on a ton of programming problems, but also on millions of generated Arc like tasks.
1:15:14So you have an extraordinarily large DSL. And then the functioning is very, very shallow recombination of these primitives. So discrete program search very deep recombination, very small set of primitive programs. And the alarm approach is the same, but on the complete opposite end of the spectrum, where you scale up the memorization by a massive factor and you're doing very, very shallow search. But they are the same thing, just different ends of the spectrum. And I think where you're going to get the most value for your compute cycles is going to be somewhere in between. You want to leverage memorization to build up a richer, more useful bank, alternative programs.
1:16:02And you don't want them to be hard -coded, like what we saw for the typical ArcDS. So you want them to be learned from examples. But then you also want to do some degree of deep search. As long as you're only doing very shallow search, you are limited to local generalization. If you want to generalize further more broadly, this depth of search is going to be critical. I might argue that the reason that he had to rely so heavily on the synthetic data was because he used a 240 million parameter model because the Kaggle competition at the time required him to use a P100 GPU, which has a tenth or something of the flops of an H100.
1:16:44And so obviously he can't use. If you believe that scaling will solve these kind of reasoning, then there you can just rely on the generalization, whereas if you're using a much smaller, for context for the listeners, by the way, the Frontier model study are literally a thousand extra than that. And so for your competition, from what I remember, the submission you'll have to submit can't make any API calls, can't go online and has to run on Nvidia Tesla T4 P100. P100. P100. Oh, is it P100? Yeah. Okay. So again, it's like significantly less possible. There's 12 hour runtime limit, basically. There's a forcing function of efficiency in the but it has this thing.
1:17:25You only have 100 test tasks. So do you have a computer available for each task is actually quite a bit, especially if you contrast that with the simplicity of each task? So it would be seven minutes per task, basically, which for people have tried to do these estimates of how many flops does a human brain have. And you can take them with a great assault, but as a sort of anchor, it's basically the amount of flops that H100 has. And I guess maybe you would argue with that well, a human brain can solve this question in faster than 7 .42 minutes. So even with a 10th of the compute, you should be able to do it in seven minutes.
1:17:58Obviously, we have less memory than, you know, like petabytes of fast access memory in the brain. And with these, you know, 29 or whatever gigabytes in this H100. Anyway, I guess the runner question I'm asking is I wish there's a way to also test this prize with some sort of scaffolding on the biggest models as a way to test whether scaling is the path to get to solving arc. Absolutely. So in the context of the competition, we want to see how much progress we can do with limited resources. But you're entirely right that it's a super interesting open question. What could the biggest model add there actually do on arc?
1:18:38So we want to actually also make available a private sort of like one -off track where you can send me to us at EM. And so you can put on it any model you want. Like you can take one of the largest open source models out there, find you need do whatever you want. And just give us an image. And then we run it on the H100 for like 24 hours or something and you see what you get. I think it's worth pointing out that there's two different test sets. There is a public test set that's in the public GitHub repository that anyone can use to train, you know, put it in an open API call, whatever you'd like to do.
1:19:14And then there's the private test set which is the 100 that is actually measuring the state of the art. So I think it is pretty open and interesting to have folks attempt to at least use the public test set and go try it. Now there is an asterisk on any score that reported on against the public test set because it is public, it could have leaked into into the training data. And this is actually what people are already doing. Like you can already try to prompt one of the best models like the latest Jaminar, the latest GPT -4. We stars from the public evolutions set. And you know, again, the primary set, these tasks are available as JSON files on GitHub.
1:19:48These models are also trained on GitHub. So they're actually trained on these tasks. And yeah, that can create uncertainty about if they can actually source some of the tasks is that because they memorize the answer or not. You know, maybe you would be better off trying to create your own private art like a vain novel test set. Don't make the task difficult. Don't make them complex. Make them very obvious for humans. But make sure to make them original as much as possible. Make them unique, different. And see how much your GPT -4 and so on are GPT -5 dozen them. Well, they're having tests on whether these models are being over trained on these benchmarks.
1:20:29Scale recently did this where on the GSM was really interesting. They basically replicated the benchmark with different questions. And so some of the models actually were extremely overfit on the benchmark like Mistral and so forth. But the Frontier models, Cloud and GPT actually did as well on their novel benchmark and then they did on the specific questions that were in the existing public benchmark. So I would be relatively optimistic about them just sort of training on the JSON. I was joking with Mike that you should allow API access but sort of keep a and even more private validation set of these arc questions.
1:21:09And so allow API access. People can sort of play with GPT -4 scaffolding to enter into this contest. And if it turns out maybe later on you run the validation set on the API and if it performs worse than the test set that you allow the API access to originally. That means that OpenAI is training on your API calls and you like go public with this and show them like oh my god they're you know they're like leaked your data. We do want to make we want to evolve the Arc dataset. Like that is that is a goal that we want to do. I think France why you mentioned you know it's not perfect. Yeah no arc is not perfect perfect benchmark.
1:21:40I mean I made it like four years ago over four years ago almost five now. This was in a time before all elams. And I think we learned a lot actually since about with potential flaws. There might be I think there is some more dentancy in the set of tasks which is of course against the goals of the benchmark. Every task is supposed to be unique in practice. That's not quite true. I think there's also every task is supposed to be very enough old but in practice they might not be. They might be structurally similar to something that you might find online somewhere. So we want to keep iterating and release an arc two version later this year.
1:22:18And I think when we do that we're going to want to make the old private test set available. So maybe we won't be releasing it publicly but what we could do is just create a test server where you can query get a task. You submit a solution and of course you can use whatever frontier model you want there. So that way because you actually have to query CPI you're making sure that no one is going to buy accident train on this data. It's unlike like the current public audit which is literally on GitHub. So there's no question about whether the model is actually trained on it. Yes they are because of the train on GitHub.
1:22:54So by sort of like gating access to querying this API with a variety of issues. And then we would see for people who actually want to try whatever technique they have in mind using whatever resources they want that would be a way for them to get an answer. I wonder what might happen. I'm not sure. One answer is that they come up with a whole new algorithm for AI with some things some explicit programs that assist that now we're on a new track. And another is they did something hacky with the existing models in a way that actually is valid which reveals that maybe intelligence is more of getting getting things to the right part of the distribution but then it can reason.
1:23:35And in that world I guess that will be interesting and maybe that'll indicate that you know you had to do something hacky with current models as they get better you won't have to do something hacky. I'm also very going to be very curious to see how these multimodal models if they will perform natively much better at arc -like tests. If arc survives three months from here we'll pull up the price. I think we're about to make a really important moment of like contact with reality by blowing up the price putting a much big price pool against it. We're going to learn really quickly if there's like low -hanging fruit of ideas.
1:24:04Again I think new ideas are needed. I think anyone listening this might have the idea in their head and I'd encourage everyone to like give it a try. And I think as time goes on that adds strength to the argument that we've sort of solved all that in progress and the new ideas are necessary to be dark. That's the point of having a money price is that you attract more people. You get them to try to solve it. And if there's an easy way to hack the benchmark that reveals that the benchmark is throughout then you're going to know about it. In fact that was the point of the original kernel competition back in 2020 for arc.
1:24:36I was running this competition because I had released this data set and I wanted to know if it was hackable, if you could cheat. So there was a small money price at the time that was like 20k and this was right around the same time as GPT -3 was released. So people of course tried GPT -3 on the public data. It's called zero. But I think with the first context, the first contest taught us is that there is no obvious shortcut. Right. And well now there's more money. There's going to be more people looking into it. Well, we're going to find out. We're going to see if the benchmark is going to survive.
1:25:17And if we end up with a solution that is not like trying to brute force the space of possible arc tasks that's just trained on core knowledge, I don't think it's necessarily going to be in and by itself, HGi, but it's probably going to be a huge milestone on the way to HGi. Because what it represents is the ability to synthesize, task, a problem solving program from just two or three examples. And that alone is a new way to program. It's an entirely new pattern for software development where you can start programming potentially quite complex programs that will generalize very well. And instead of programming them by coming up with the shape of the program in your mind and then typing it up, you're actually just showing the computer what add what you want.
1:26:14And you let the computer figure that. I think that alone is extremely powerful. I want to riff a little bit on what kinds of solutions might be possible here. And which you would consider sort of defeating the purpose of arc. And which are sort of valid. Here's one I'll mention, which is my friends Ryan and Buck stayed up last night because I told them about this. And they were like, oh, of course, of course, they keep this. Of course, all of this, and then so they were trying to prompt, I think, Claude Opus on this. And they say they got 25 % on the public arc test. And what they've done, did was have other examples of some of the arc tests and in context, explain the reasoning of why you went from one output to another output.
1:26:57And then now you have the current problem. And I think also maybe expressing the JSON in a way that is more amenable to the tokenizer. And another thing was using the code interpreter. So I'm curious actually, what if you think the code interpreter, which keeps getting better as these models gets smarter, is just the programs synthesis right there. Because what they were able to do was the actual output of the cells that the JSON output, they got through the code interpreter, like write the Python program that gets right out here. Do you think that the programs synthesis kind of researchers are talking about will look like just using the code interpreter in large language models?
1:27:36I think whatever solution we see that we score well is going to probably need to leverage some aspects from deep learning models and LLMs in particular. We've shown already that LLMs can do quite well. That's basically the jack call approach. We've also shown that pure discrete program search from a small DSL does very, very well before jack call this was a state of the art. In fact, it's still extremely close to the state of the art. And there's no deep learning involved at all in these models. So we have two approaches that have basically no overlap that are doing quite well. And they're very much at two opposite ends of one spectrum.
1:28:11Where on one end you have these extremely large banks of millions of vector programs, but very, very shadow recombination, like simplicity group combination. And on the other end you have very simplistic DSLs, very simple, like 100 or 200 primitives, but very deep, very sophisticated program search. The solution is going to be somewhere in between. So the people are going to be winning the art competition and we are going to be making the most progress towards near -term engineer arguments who is that manage to merge the deep learning paradigm and the discrete for unsurpassed time into one elegant way.
1:28:49And you know, you're asked like what would be legitimate and what would be cheating for instance. So I think you want to add the code code interpreter to the system. I think that's great. So the legitimate. The part that would be cheating is try to anticipate what might be in the test, like brute force, the space of possible tasks and then train a memorization system on it. And then rely on the fact that you're generating so many tasks like millions and millions and millions that inevitably there's going to be some overlap between what you're generating and what's in the test set. I think that's defeating the purpose of benchmark because then you can just solve it with that and you need to adapt just by fetching a memorize solution.
1:29:31So hopefully ARC will resist to that, but you know, nothing, no benchmarks necessarily perfect. So maybe there's a way to hack it. And I guess we are going to get an answer soon. I think some amount of fine tuning is valid because these models don't natively think in terms of, especially the language models alone, which the open source models that they would have to use to be competitive here, compete here. They're, you know, they're like natively language. So they like need to be able to think in the in this kind of yes, the arc type way. You want to input core knowledge like arc like core knowledge into the model, but surely you don't need tens of millions of tasks to do this.
1:30:06Like, core analysis is really basic. If you look at some arc type questions, I actually do think they rely a little bit on things I have seen throughout my life. And for the same, like, for example, like something bounces off a wall and comes back and you see that pattern. It's like I played arcade games and I've seen like pong or something. And I think, for example, when you see the flin effect and people's intelligence has measured on variants progressive matrices increasing on these kinds of questions, it's probably a similar story where since now since childhood, we actually see these sorts of patterns in TV and whatever, spatial patterns.
1:30:41And so I don't think this is sort of core knowledge. I think actually this is also part of the quote unquote fine tuning that humans have as they grow up of seeing different kinds of spatial patterns and trying to pattern match to them. I would definitely file that in the core knowledge. Like, core knowledge includes basic physics, for instance, bouncing or trajectories that would be included. But yeah, I think I think you're entirely right. The reason why as a human, you're able to quickly figure out the solution is because you have this set of building blocks, this set of patterns in your mind that you can recombine.
1:31:12Is core knowledge required to attain intelligence and any algorithm you have, does the core knowledge have to be in some sense hard coded or can even the core knowledge be learned through intelligence? Core knowledge can be learned. And I think in the case of humans, some amount of core knowledge is something that you're born with. Like we're actually born with a small amount of knowledge about the well -being of an alien. We are not blanks, but most core knowledge is acquired through experience. But the thing with core knowledge that it's not going to be acquired, like for instance, in school, it's actually acquired very, very early in the first class three to four years of your life.
1:31:48And by age four, you have all the core knowledge you're going to need as an adult. Okay, interesting. So I mean, on the price itself, I'm super excited to see both the open source versions of maybe with a llama, a 7 -B or something, what people can score in the competition itself. Then if to sort of test specifically the scaling hypothesis, I'm very curious to see if you can prompt on the public version of ARC, which I guess won't be compatible. You won't be able to submit to this competition itself. But I'd be very curious to see how if people can sort of crack that and get artworking there. And if that were to update your reviews on Asia.
1:32:22It's only be motivating. We're going to keep running the contest until somebody puts a reproducible open source version into public domain. So even if somebody privately beats the ARC, we're going to still keep the price money until someone can reproduce it and put the public reproducible version out there. Yeah, exactly. Like the goal is to accelerate progress towards the GI. And a key part of that is that any sort of meaningful bits of progress needs to be shared, needs to be public. So everyone can know about it and can try to iterate on it. If there's no shaming, there's no progress. What I'm especially curious about is sort of disaggregating the bets of like, can we make an open version of this versus is this a thing that's just possible with scaling?
1:33:01And we can, I guess, test both of them based on the public and the private version. We're making contact with reality as well with this. Right? We're going to learn a lot. I think about what the actual limits of the compute were. If someone showed up and said, hey, here's a closed source model that like, I'm getting 50 plus percent on, I think that would probably update us on like, okay, perhaps we should increase the amount of compute that we give on the private test set in order to balance some of the decisions initially or someone arbitrary in order to learn about, okay, what do people want?
1:33:25What does progress look like? And I think both of us are sort of committed to evolving it over time in order to be the best, the closest to perfect because we can get it. Awesome. And where can people go to learn more about the prize and maybe give their hand at it? ParkPrize .org. Which to go to those lives today. So, what? No. But $1 million is on this line, people. Good luck. Thank you, guys. We're coming on the podcast. It's super fine to go through all the cruxers on intelligence and get a different perspective. And also to announce the prize here. So this is awesome. Thank you for opening the news.
1:33:51Thank you. Thank you. Fineness.
From the publisher
Here is my conversation with Francois Chollet and Mike Knoop on the $1 million ARC-AGI Prize they're launching today.
I did a bunch of socratic grilling throughout, but Francois’s arguments about why LLMs won’t lead to AGI are very interesting and worth thinking through.
It was really fun discussing/debating the cruxes. Enjoy!
Watch on YouTube. Listen on Apple Podcasts, Spotify, or any other podcast platform. Read the full transcript here.
Timestamps
(00:00:00) – The ARC benchmark
(00:11:10) – Why LLMs struggle with ARC
(00:19:00) – Skill vs intelligence
(00:27:55) - Do we need “AGI” to automate most jobs?
(00:48:28) – Future of AI progress: deep learning + program synthesis
(01:00:40) – How Mike Knoop got nerd-sniped by ARC
(01:08:37) – Million $ ARC Prize
(01:10:33) – Resisting benchmark saturation
(01:18:08) – ARC scores on frontier vs open source models
(01:26:19) – Possible solutions to ARC Prize
Get full access to Dwarkesh Podcast at www.dwarkesh.com/subscribe




