Zapier’s Mike Knoop launches ARC Prize to Jumpstart New Ideas for AGI

2 Jul 2024 · 55 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Podcast Summary: Zapier’s Mike Knoop launches ARC Prize to Jumpstart New Ideas for AGI

Podcast Overview Podcast Title: Training Data Hosts: Sonya Huang and Pat Grady, Sequoia Capital Episode Title: Zapier’s Mike Knoop launches ARC Prize to Jumpstart New Ideas for AGI Episode Description: Discussion on the lack of progress towards AGI despite advances in AI technologies, focusing on the Abstraction and Reasoning Corpus (ARC) and the recently launched ARC Prize competition.

---

Key Themes and Discussions

Introduction to ARC and AGI

  • ARC Benchmark: Introduced by François Chollet in 2019, focusing on efficiency in skill acquisition rather than language or scale.
  • AGI Definition: Knoop emphasizes the need for a clearer definition of AGI, which he aligns with the ability to efficiently learn new skills from minimal examples.

Launching the ARC Prize

  • Motivation: After observing the stagnation in progress towards beating the ARC benchmark, Knoop co-launched the ARC Prize with Chollet, offering over $1 million to incentivize breakthroughs.
  • Competition Goals: Aimed at attracting new ideas and researchers to address the limitations of existing AI models.

Progress and Expectations

  • Current Standing: As of the episode, the state-of-the-art performance on the ARC benchmark is around 39%.
  • Future Predictions: Knoop expresses cautious optimism, suggesting that achieving 50% during the competition is plausible, yet 85% may take much longer.

---

Insights from Mike Knoop Zapier's AI Integration

  • Zapier Overview: Knoop highlights how Zapier integrates AI into its workflow automation platform, enabling non-technical users to leverage powerful AI tools.
  • Impact of AI: Over 50% of Zapier employees utilize AI in daily tasks, significantly enhancing productivity.

The Role of New Ideas in AGI

  • Knoop insists that breakthroughs in AGI will require innovative approaches rather than just scaling existing models.
  • Historical reliance on large language models (LLMs) may be limiting the exploration of new pathways to AGI.

Differentiating ARC from Other Benchmarks

  • Knoop argues that ARC's unique design, which emphasizes novelty and challenges memorization techniques, sets it apart from other benchmarks that are susceptible to being easily beaten as models scale.

---

Audience Engagement and Research Community

  • Diverse Participation: The ARC Prize has attracted participants from various backgrounds, with many competitors coming from outside traditional AI labs, emphasizing the potential for innovative solutions from unconventional thinkers.
  • Public Awareness: Knoop notes that increasing awareness around the ARC benchmark is crucial for encouraging new research ideas and driving progress towards meaningful advancements in AGI.

---

Challenges and Future Directions

  • Current Research Landscape: Knoop discusses how the focus on LLMs and large-scale models has overshadowed the need for fresh perspectives in AGI research.
  • Open Source Commitment: Winning solutions must be shared publicly to foster collaborative growth and accelerate AGI advancements.

---

Conclusion Mike Knoop's insights during this podcast episode underscore the importance of redefining benchmarks and encouraging innovative approaches in the pursuit of AGI. The ARC Prize serves as a pivotal initiative aimed at reigniting interest and creativity in solving complex AI challenges, setting the stage for future breakthroughs that could redefine our understanding of intelligence.

Key Takeaways

  • The ARC Prize aims to address stagnation in AGI progress with a focus on skill acquisition efficiency.
  • Zapier's integration of AI illustrates practical applications of AI in enhancing productivity.
  • A diverse range of participants in the ARC competition may lead to unexpected breakthroughs.
  • Public awareness and open-source sharing of solutions are crucial for advancing AI research.

---

This summary captures the essence of the podcast discussion, highlighting significant concepts and the forward-looking vision for AGI development as articulated by Mike Knoop.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Right now, I think that what I see happening is there's sort of this mythical story of a very bad outcome once we get to like superintelligence rate. It's a very theoretical driven story. It's not grounded in sort of empirical evidence. It's basically based on sort of reasoning our way to like this outcome. And I think the only way that we can really, really effectively and truly set good policies by you have to look at what the systems can and can't do and then regulate or decide make decisions at that point about what it can or can't do. I think anything else is sort of like, you know, you're cutting off potential really, really good futures way too early.

0:45Hi, and welcome to Training Data. We have with us today Mike Nupe, co -founder of Zapier. Mike has recently stepped up to co -found and sponsor the ARC Prize, which is one of the most unique benchmarks in AI that measures a machine's ability to truly learn things intelligently versus just to parrot patterns in the training data. We're excited to ask Mike for an update on how things are going with the ARC Prize two weeks in and to hear his views on why we need radically different approaches and benchmarks to achieve true general intelligence. My thanks for being here today. So we're excited to talk about the Arc HGI initiative.

1:24Before we get into that, I'd love to spend a few minutes on your background at Zapier, because I think Zapier is emerged as probably one of the best examples of what an existing application company can do with the Power of AI. And the way that you guys were sort of early to that, the way that it's now kind of interwoven in the product has been really interesting to watch. So maybe can you just say a few words on what is Appier and what has your approach to AI been as Appier? Yeah, Zapier is a workflow automation platform. We support 6 ,000 different integrations things from Salesforce to Gmail to you know, any basically SaaS software that you can imagine we connect with and I think the anything about Zapier is that's intended to be either easy to use for non -technical users.

2:07So the majority of Zapier customers users would not self -identify with being a programmer or being technical, being an engineer or something like that. Even though I also think of myself as an engineer, and I find lots of interesting use cases for Zapier with the large majority of our users, I don't. And I think that's what is quite special about Zapier, and why people tend to fall in love with it is the feeling of power and leverage that you get as a non -technical user being able to have software do work for you basically. And I think that's in an interesting way, the exact same promise of AI, right?

2:41Like that's what people want of AI. They want software that's just going to do more work for them. And so in many ways, I think the mission of Zapier and sort of the mission and purpose of AI intersect. And I've been sort of like, I guess, called me AI curious since all the way back to college. And I gave like a whole all hands at Zapier camera what year but like when the GPT -3 paper came out and show that to the company. And so I've been sort of tracking and following the long progress. But it really wasn't until I think January of 2022 when I think this was the chain of thought paper when that one came out that I saw that and really surprised me because I thought I'd priced in everything that sort of AI language models could sort of do up to that point.

3:19And this idea of like, you know, let's think stuff I step, this chain of thought technique breaking down, these language models is just tools for reasoning as of, you know, just a one shot sort of completion or a chat engine felt very special and something that I think most people didn't expect that they could do even though the technology down there for over a year at that point. And so that actually that kind of moment caused me to give up my exacting role I was running all of our product, then jeering at the time, and I went back to basically be an individual contributor of the company as an AI researcher alongside my co -founder Brian.

3:50And I have to talk more about sort of that journey, but yeah, I think that's what caused Zapier to be relatively early in terms of AI. What are some of the things that you've put into the product that you're most proud of at Zapier in terms of AI features? You know, at this point, I think Zapier was pretty. I'd say there's probably two main places where we've gotten a lot of value from AI. The first is that over half the company now individual -level uses AI on a daily basis. And I know this because we're actually measuring Zapier's own platform usage of our own company. So this is like we have over half the company is like building, zaps, building automations that use an AI step like a chat chat GPT step in the middle, either to do so content generation or data extraction over unstructured text.

4:34Also to relinching use cases that we can sort of talk about. in fact, one of the top internal use cases is probably getting us about, I think, like a hundred X labor enhancement rate, which is been phenomenal. Well, what is that? Yeah, you want to talk about it? Yeah, yeah. I mean, hundred X improvement. Yeah, I've got to talk about that. I think it's our personal, like, high watermark for what we've been able to achieve using AI internally for, like, an operational perspective. And so Zapier has these things on our website called Zap Templates. their effectively recipes that help users figure out what can zap your due and help them get started.

5:09And these templates are in order to make them their own, they've all been historically handmade because they require a bit of right brain and a bit of left brain. They're very creative. You have to inspire the end user, the customer, the, you know, would be user of what could zap your due for you, what's the outcome, what's the ROI you might get. And then there's also a very technical way they have to be crafted in a built as well. You have the map JSON fields from One integration app to another integration app to make sure it actually works and together That's what actually helps users get started and activates the minute the product and we had a whole of maybe a Millen of these things that we knew we wanted to build that we hadn't built yet because they just they take so much effort the Sort of rate of production for our contractors was about 10 a day up until last summer and We had a person the member of our I think it was our partner marketing team actually a background of other in freelance writing who built a zap, a system of actually several zaps using open a eyes and middle steps here and built an end in system that whenever a new integration got launched on Zapier, it would automatically try to like figure out what are the most interesting zap templates it could be built, write the inspirational use case behind it because there's probably a ton of millions of these things already today with lots of training data and then also do the exacting field mapping as well and we moved the human in that workflow from the do loop to the review loops.

6:28So instead of having those contractors now, actually generating them, thinking really hard about what the use cases should be building that the sort of field mappings, now they're basically reviewing output from the system in a spreadsheet and saying, yes, this one is good, this one's bad, this one's good, this one's bad. And the funny thing is because of the cost of production of generating these things is solo, we don't even try to fix the bad ones. We just throw them away and say, well, just generate another stack and throw them on top. And so that rate of production is about a thousand a day now.

6:54So we've gone from 10 to 8000 a day per contractor. And we've been shipping away steadily at that whole of a million and then keeping up with the launch of integration zones after. I think one of the main things that that showed me was the space you want to look for in businesses if you're thinking about how to deploy I. It's like really up top of final or bottom of the funnel tend to be like you want to get something that's really close to an important conversion for your business. And then, you know, I think if you can identify any manual work that your organization does that has I volume Where human is doing the work.

7:29I think those are sort of opportunities to like introspect and say okay Is there an opportunity to get that human out of a do loop and and sort of craft a system that can do the work But then put the human in the review loop which is still quite needed at the sort of maturity level of the technology today But still phenomenal from our eye perspective Are there any metrics you can share on the impact Zapier AI has had on the overall Zapier business? Yes. The biggest one today is we're just about to hit 10 million AI tasks per month. No, I think we're on run rate for about 10 million AI tasks per month.

8:01And I think what's, I think I would love to be shown wrong if corrected if you know examples. But I think at this point, Zapier might be the biggest automated AI platform in the world. In the sense, you know, there's a lot of researchers, entrepreneurs, builders who are trying these like agentic AI systems where the AI is sort of working without sort of human, you know, in the loop. And yeah, I think at this point with 10 million ad tasks a month, we, exactly, maybe the like biggest example of that in the world right now. Really cool. Can we talk about our Arcage AI? Let's do it. Maybe start with a recap of what is Arcage AI?

8:39Why did you, in front of us, set out to establish this price? Yeah. This was, So follow on to my AI curiosity. So the reason I gave up my exacting role back in 2022, what's I kind of wanted to know for myself? Are we in path rage or not? I felt very important to know for Zapers mission, but all of just as a human, I was very curious and one to know, is this gonna happen? There's definitely some interesting scaling that's happening. Is that sufficient to get to, I think, what naively I had in my head is this super intelligence API. And surprisingly, what I learned is the answer is no. I revisited, actually got to first know a Francois Chalet, who is my co -founder on ArcPrice.

9:21I first heard about him and got exposed to his research factor in COVID actually. It was during 2020, he did a, I think, another podcast where he was explaining his paper, 2019 paper, Unmeasured Intelligence, where he tried to formalize a definition of what a GI actually is. Yeah, I thought it was interesting at the time, but I kind of parked it with lots of other stuff going on. It's happier and doing our own AI product building. But as app you got more into sort of building with AI, I built my intuition of what language models could do, what the apparent limits were, I started getting more into AI eVals and trying to like understand where it's sort of the limits were, what could we expect from a product as on a product building perspective, like where are our products going to like tap out, where should we invest our engineering and research effort versus just wait for the technology to keep scaling mature.

10:06And you know, the thing I found was the most AI eVals were saturating up to human level performance and it was accelerating. And when I went back to look at the ARC, Yval, back from 2019, expected to see a similar trend. And instead, what I found was basically the opposite, not only had it not reached to him in performance yet, but it was actually decelerating over time. And this was like really supremely sort of surprising to me. And maybe it's worth defining, you know, we use these terms of AGI, but like what's the actual correct definition for it, right? I think there's a kind of popular definition in the world today.

10:39Actually, there are probably two schools of thought. I think one school of thought that I see is AGI is undefinedable and we shouldn't even try. This is a quite popular perspective. And I think the other school of thought is, that AGI is the system that can do the majority of economically useful work humans do. This was popularized by OpenAI and the Microsoft deal. This is actually like in their deal together. Like once this is achieved, like OpenAI retains all the future IP, it's very interesting. I think the node coastline actually might get credit for pointing that definition. But nonetheless, I think because of open -air success, that definition has sort of become accepted by a lot of people and is a target a goal we should shoot for.

11:17The challenge is I think it's a fine goal, by the way, and I think current model architecture may be within spitting distance of it. I think it says way more probably about what the majority of humans do for work if it's a true goal than what AGI actually is, though. And Francois defines AGI as the efficiency of acquiring new skill. That's it. And here's like a quick thought experiment I think you can use to chart if like kind of a rock this is we've had AI systems now for you know many years, five plus years that can beat humans at games like go, chess, poker, diplomacy even. And the factor means that you cannot take any one of the systems that was built to beat one of those games and simply retrain it with new data, new experience to beat humans at another game.

12:03Instead, what researchers and builders and engineers have to do is they have to go back to zero. They have to tear it all down, rethink of new algorithms, new architectures, new ideas, of course, new training data as well, often new amounts of scale in order to beat that next game. And yet, this is in complete contrast to how you two both learn, right? I could like sit you both down here, teach you a new card game and probably about an hour. I could probably show you a new board game and get you up to proficiency within a couple hours. And that fact, is what makes, I think it's highly represented what makes you generally intelligent.

12:37It's your ability to very quickly and efficiently sort of gain skill in order to accomplish some up, open -ended, or novel tasks that you've like never encountered before in your life. And that's what's special about Arc. So Arcage, GIs, and Eval, that tries to take that definition and actually measure it. And it was designed specifically to resist the ability to memorize the benchmark, which is very different from most other AI eVals that are out there. Every task is completely novel, and there's a private test at the known scene outside of a handful of people that have taken it to verify that all of the puzzles are solvable.

13:09And that degree of novelty and that degree of not having ever been seen before is what makes ARCA really, really strong benchmark for trying to distinguish between this more narrow AI that can be beaten largely through memorization techniques and EGI, which is a system that can very, very rapidly and efficiently acquire the skill of test time. What is the definition of efficiency? I imagine there's a compute component, a data component. What's the definition of efficiently and efficiently acquired new skill? Yeah, Francoisum, I'll probably do a bad job trying to like summarize his research. If you want to read more, by the way, his on -mode of intelligence paper is like the source of truth for all of this stuff.

13:44I think it's really, really good. And I think one other, before I get to the answer, I think one other important thing to sort of see is that ARC has been unbeaten since 2019. And I think it's endurance to date is probably the strongest set of empirical evidence that the underlying concepts of the definition are correct, which is why I think it's worth paying attention to and why it's such a specialty, though, in a special set of research. So I think for ASIO, I would describe efficiency as the ability for a system to translate from core knowledge priors to being able to attack the search space or task space around it.

14:21A very weak generalizable system is only going to be able to take on a near -term adjacent and tasks to the core data, the core knowledge that that system was trained on, whereas a highly generalizable system is going to be able to have a much larger field of tasks and novelty that it's able to attack and be able to effectively do with a small set of training data. And that's what we hope to see with the eventual solution for Arcade GIs, as well, is that if someone's able to beat it, the goal is to get 85 % on the eVAL. Today's state of the art is, I think, 39 % as we record this. And I think what's special is if someone One thing that can actually be at ARC at the 85%.

14:59That would mean that you've created a computer program that can be trained on this very small set of core knowledge priors, things like goal -directedness, objectness, symmetry rotation. These are sort of things that emerge very early in childhood development. And be able to use those core knowledge priors and recombine them and synthesize them into new programs in order to solve tasks with exacting accuracy that that system has never we've never seen before, never been exposed to it, it's trained in. And that would be a really, really important thing. Particularly if application layer were like the number one problem today is like hallucination, accuracy and consistency.

15:33And that results in this loads of trust, which limits deployment of like really eye right now. You have some peculiar rules for a company for the price. I think there's a limit on how much computer you can use. You can't use the internet. I don't know if you can use GPC4 and closed models. Why put those limits in place? Yeah, so the two big ones are you're right. You so on the the competition so it's on Kaggle and Kaggle and force is no internet and you have limited compute so specifically you get one P100 for 12 hours and No internet means you can't use front -term close models. They're you know available three APIs like Claude, Sonnet or Gemini or Tribute 40 or 40 Maybe a little take -men order.

16:15I think the compute one is maybe more interesting the reason for the compute limit is to target efficiency, first and foremost. Because if there wasn't any compute at all, then you could simply define a GI as saying it's just a system that can acquire skill with no degree of efficiency attached to it. And if that was true, that would mean that this system could brute force basically every possible program, think through every possible future outcome here, generate every possible single archetype of puzzle, and use that in order to sort of win the challenge. And we know that's not actually what happens in human general intelligence.

16:46The thought you can read more in sort of France has paper about why, but the way that I think about it is, you can think about it, you can introspect even yourself while you're taking the arc puzzles and see that when you're trying to solve one of these, that you're not brute forcing every possible transformation from the pattern, trying to recognize the pattern and apply it to the test. Instead, you're using your intuition, you're using your prior experience to try and identify maybe three, four, five possible possibilities of like what the pattern is and then you check them right in your head.

17:14And I think this shows that humans are the sort of efficiency humans have is not brute forcing every possible solution and checking it's actually there's a degree of efficiency. So the computer with it, it sort of forces researchers to reckon with that definition. Now I think it is worth important acknowledging like we don't know exactly how much computer is necessary to beat our kit and we're going to like keep upping the computer bar over time is what I expect. For example, we already upped it, we over to x to it from prior versions of the competition. So I think in prior years you got somewhere between like two and five hours to run on the GPU and we we bump that up to 12 Interestingly all of the state the art techniques are actually maxing out that 12 hour runtime as well So I do expect we'll continue to increase it over time But it but I think it is important tool in order to like force the generality out of the sort of full solution that we're looking for And then new internet is a little more of a practical reason, you know We're trying to reduce cheating reduce contamination reduce overfitting not be able to leak the private test set and And largely just increased confidence that when we reach the 85 % grand prize mark that someone has actually be narc and be able to sort of say that with some sense of sort of authority and confidence that's that's a true statement One of Francois and my goals for our prize is to establish a public benchmark of progress towards or maybe towards a GI or maybe the lack of progress towards a GI and and have it be sort of a trusted public tool that policy makers at students, entrepreneurs, venture capitalists, employees, everyone can look at to get a sense of how close or far are we away from this sort of important technology existing, and then using that insight in order to help try to drive more AI researchers to work again on exploring new ideas, which is something that's unfortunately kind of fallen out of favor in the last several years as LLM -SIP tech -off.

19:03What have you seen or maybe what do you expect to be true about the efforts that are successful or more successful toward ArcAGI that makes them different from what we're seeing out of the frontier models and the big research labs. So it gets into the details of how does an LM work because that's kind of the best frontier AI research slides have been taking less. But here is we're going to scale up language models and that's going to more scale more data is going to get us to AI. And even though that's the dominant story, I actually think it's what most of the labs actually believe internally.

19:34most of them are working on new ideas. So I think there's like an interesting story there, but it is definitely in their interest to sort of promote a very strong narrative of like scales all you need to don't compete with. You know, we're just gonna steamroll you out of the sea here. Yeah, I think there's true competitive dynamics sort of immersion in the market that have, that are, unfortunately, I think shaping a lot of attention, investment, effort away from exploring new ideas. And if it is true that new ideas are needed, which I believe it is, and I think ArcPy and ArcPry show that at least some new idea is needed, then due to the competitive dynamics and emergence of the marketing of the last couple of years, we're headed in the wrong direction.

20:16There's all the frontier research has basically gone close source. The GPD -4 paper had no technical details shared in it. The GPD, the Gemini paper had no technical details shared on its longer context innovation, things like that. Yet this is in direct contrast to the history of how we even got here today, right? The sort of innovation set that led the sort of chain of research that led from, you know, Ilya's sequence to sequence paper, Google, out to Jacobs University, back to Google, then to Alagradford, and back to Ilya at OpenAI. Like, there's like a six or seven year chain of research that only happened because of open sharing, open progress, and open science.

20:50And I think that's a bit unfortunate that we don't, we don't really have that right now. Again, somewhat just due to the market dynamics and commercial successful success of language models kind of forcing a lot of that closed -frontary research up. So, yeah, one of the goals that I've probably just helped counterbalance a lot of those things. You were asking about what difference between what you said resonates because it seems like a lot of the foundation model companies are going down very similar somewhat clearly to find paths. And I'm sure that internally there's all sorts of work being done to find the next breakthrough through an architecture.

21:26But in terms of what's working today, they're all fairly similar paths. And I imagine though, or based on lumps? Yes, and I imagine though, what works for the sake of Arcage .ai is gonna be a little bit of a different shape. And I'm wondering if you're starting to see what shape that may take. Got it. And have a sense for what may be different about this more general architecture than what we're seeing out of the foundation models. Great, so I think, you know, a useful shortcut on how to think about language models is that they are effectively doing very high -dimensional memorization, right? They're able to train and memorize tons and tons of examples and apply them in slightly adjacent contexts.

22:04I don't want to under, I don't want to miss language models too much because I think that they are very special, so they're very magical and something very, that has lots of economic utility, Zapier is an existence proof of that fact alone. So, yeah, I don't want to like throw it under the bus too much. I think there's some really good things that it has unlocked as a technology goes. But there are limits to it. And the sort of limits are not being able to effectively leverage its training data to compose it or combine it at test time to go attack and accomplish novel tasks that had never seen before in its training data.

22:40And that's what ArcShort of shows, right, is that this is like a skill that these thingage models don't sort of possess. I think it's kind of maybe, um, used to the look at the history of the HIFS course so far and maybe where we expected to go. So from when the e -vow was first introduced in 2019, 2020, the first, there was a small cargo accomplished to the brand to kind of get a baseline when it was 20%. And from 20 % to 30%, the techniques that worked were effectively, researchers crafted a handcrafted domain -specific language by looking at the puzzles that were in part of the public test set.

23:11There's two test sets. There's a public set and then the private one that's the say the R's measure on, they looked at the public test and they tried to infer and write down programs and Python code or C -sharp or whatever. What are the like individual transformations that like you do in your head to go from the you know one puzzle to the next. And so that they called this a DSL and then they wrote a brute force search to try and search through all possible like permutations and combinations of those sub -programs in order to like find the general pattern and then apply it at real time. And that got to about 30%.

23:41What's gotten from 30 % up to close to 40 % now is a slightly different technique. This is a jack hole in his approach is effectively using a code -based open source language model and doing test time fine -tuning. So he has some pre -training down on the code -gen model and then at test time, aching the puzzles they get, the novel puzzles that's never been seen before and permutating variations of it and then training this like code -gen based model in order to write that program and find a program that fits the pattern and then apply it at test time, and that's gotten to 40%. I suspect that we probably have, I bet we have the ideas in the air already to get to like the 50 % mark, maybe even a little beyond the 50 % mark without a lot of new innovation.

24:26I bet just unsombling or combining these sort of existing ideas, that's that have already worked toward dark probably gets you about halfway. I think to get to the 85 % market or beyond, I think the solution, the ultimate solution probably looks more like the shape of, at least to solve arc, something that looks like a deep learning guided DSL generator where you have some sort of, instead of hand coding and hand crafting the DSL like ahead of time by trying to infer from the public test set what those like subprograms would be. You need some way to generate that DSL dynamically, right, by looking at the puzzles in real time, and being able to learn from past puzzles and apply that towards future puzzles.

25:06This is also another important thing humans do when they're going through the art set. Sometimes the first or second puzzle are actually a little trickier because you're orienting yourself around, what am I doing, what task am I doing, what does the possible solution space look like, and then as you get further into the task set, they tend to get a little easier because they start, you know, some of the rec, the sort of space that possible transformations is just fine at. So you start kind of recognizing patterns there. And then combining that with some sort of deep learning based programs at this ascension, something that can not brute force all possible programs of how to combine those DSLs together, but something that has some sort of deep learning approach to shape which program traces do you try to generate or test and try against the pattern, and then it kind of goes back to this human introspection of how we take the puzzles, which is we're not brute forcing all possible programs in our head and said we're trying to identify just a handful of likely candidates and then testing those deterministic manner ahead of applying the one that works.

25:59It's really interesting that code generation and programs synthesis kind of underlies all of the methods you just talked about. There's something very special there like programs synthesis is very general allows either actually get closer to that definition of generalize intelligence that you mentioned at the beginning. It's very exacting and I think this is one of the reasons why this solution to arcade GIs going to be useful very quickly. So, you were talking this before. There's a history of toy AI benchmarks over the last 10 -15 years that kind of looked like ARC. There were games, there were puzzles, and really never amounted to much in terms of being beaten.

Read the full transcript

26:39They all got handily beaten as sort of scale emerged, and they really didn't add to or understanding of how to build useful AI systems. And so what's one of the common questions I've been feeling last couple weeks is like, What's different about ARC? Is that likely to just happen here again? And I think the reason why we're likely to see something much more useful, assuming we get a really good solution to ARC from the first grand prize win, is that the number one problem at the application layer, and we see this with Zapier too, with our new AI bots that we launched a couple months ago, been surprising to me in how that has gotten adopted actually by our users.

27:19There's like how you kind of describe it. There's like concentric rings of use cases that you can use AI automation for. And what we're seeing is people are sort of restricting the use cases for the AI bots, where they're sort of fully automated, totally hands off to the use cases, where there's sort of a very low need for user trust. Or where the sort of, let me say that different way, is if it goes wrong, it's not cast -strapped looking bad. So they deployed for use cases like personal productivity activity or team -based like workflow automation, things where, you know, if it's wrong, or it's right only nine out of 10 times, or it takes me, you know, maybe a couple days to like really work with the system to like do the prompt engineering to steer towards getting maybe 95, 99 % reliability.

28:01That's acceptable because the risk of being wrong is just, you know, quite low. In order to get much higher up and expand the number of concentric rings to, you know, moderate risk to high -risk deployment scenarios where we want these systems working autonomously, we're going to need that the main thing that is missing is user conference and the exacting nature of what it can do and what it can't do. This is what Zapier Core Classic gives us, right, is like a deterministic engine to execute automation. So you know, once you build it and set up, it's going to do the exact same thing every single time.

28:29But that's also what makes it fragile and hard to use. And on contrast, you know, these AI Core -based LM systems that are total autonomous of the opposite set of trade -offs, right? They're much easier to use. just steer them, guide them, and fix them entirely to natural language. But because the accuracy is still in exact confidence is low. And I think that's what art gets us. So solution to art at 85 or 100 % means that you've written a computer program that can generalize from like very simple core knowledge priors to solve with exacting accuracy at high percent reliability. These like site unseen puzzles.

29:00And I think that that tool, as a, there will be a new tool in like the program is still in terms of building products and building systems that can achieve that same thing. We're two weeks in, I think, to when you launched ArcAGI Prize. What have you learned so far? What types of people are working and competing on this? Is it like the pedigrid researchers or the big labs? Is it scrappy hustler types? Like, who's competing? How many teams are seven -thing solutions? Yeah, let's see. So the response by the way after launch was phenomenal, was much bigger than we expected. I think we had, we were trending on Twitter twice, during the launch week, the number one Kaggle competition in the world, over a million social views, I think, over all the launch channels.

29:42So just a very phenomenal, I'm really thankful for everyone who helped sort of promote and helped share Arc. Hopefully we actually like can get a solution here and some short time. I think that like the most interesting thing about the folks that are working towards Arc is probably a historic glance here and then what I've seen over the last two weeks. So the historic Lancers is most of the people that have worked on ARC or outsiders to the field. They, like, this is not actually the first year that there's ever been a contest about. There was a past competition called Archithal and I was not sure what's smaller.

30:12I was hosted out of this lab 42 AI lab in Switzerland. And so last year there was actually 300 teams that worked on trying to be dark and again no one had sort of beat it. And almost to my knowledge, all of those teams were effectively individuals or effective, you know, outsiders in some way. They're not, you know, people with big AI labs, their focus with backgrounds and engineering, mechanical engineering or video game programming or physics. Folks, they just kind of got curious and interested in the problem at hand. And I actually think that's more likely than not where the breakthrough for ARK is going to come from.

30:48I think it's going to come from an outsider. Somebody who, like, somebody just thinks a little bit differently or has a different set of life experience, so they're able to cross -pollinate a couple of really important ideas across fields. That's one of the reasons why I put as much money as we did in ARC prize. I felt like the progress was idea rate limited actually. And one of the best ways to sort of increase the amount of ideas is to try and blow up awareness, which is what the launch kind of did. Over the last two weeks, I've kind of seen like two probably like camps of people, at least on Twitter, Emerge.

31:19I think there's one camp of people who are sort of the, you know, they're in it for the mission. They grew with the underlying concept. they think that we do need some new ideas or excited to try and figure out what those are. And then there's a second group of people that are sort of like, I'm gonna prove you wrong. LMs are definitely enough. Scale is definitely what we did. And I'm gonna do my best to go feed this benchmark just using existing off -the -shelf technology and sort of prove you wrong. So I'm actually quite happy for both those camps to exist. One of those approaches is currently up in Litherboards.

31:48All right. So yeah, we can break some sort of news here. So this week, this Thursday, we're launching. I guess when this comes out, I'll have launched just a couple days in the past, a brand new public task leaderboard. We talked about how ARC doesn't allow internet access and there's compute limits. I know personally how unsatisfying that is to not be able to use frontier models though. I also want to know how good can GPD -4O, how good it can collage on it, like do against this benchmark. And also because no compute owner, it's also a bit of a barrier to entry. You have to use up -and -source models, you have to do, just quite a bit of engineering or get to do before you can start just testing and experimenting.

32:26So we're gonna be, we're launching a new public task leader board. It's gonna be a secondary leader board. We're gonna be committing about $150 ,000 for a reproducible fund towards this secondary leader board. And we'll be officially part of the competition this year. You know, we wanna, we wanna like maintain that aspect of like assurance on, you know, cheating and contamination overfitting with the private test set. And that's also the test set that has sort of the most empirical evidence against over last four years. But the secondary leaderboard is going to allow folks to basically submit scores towards it against the 400 public tasks set.

32:59And we'll verify and reproduce the scores locally to sort of ensure good like fitting with the approach. And we'll publish that. And you're right. I think the top score or one of the top scores on that is this guy Ryan Greenblatt. He came out a couple days after the competition launched with a pretty interesting novel approach actually towards beating it and he's using GPT -40. And but not just GPT -40. I think the interesting thing is he created a like an outer loop around 40 where he is using 40 to generate Programs or sample from GPD 40 these like program these reasoning traces to Beat the tasks or identify the patterns then testing these patterns against Attempts against the demonstrations that then finding the one that works on applying it and that's that approach seems to be getting in the like like low 40s, maybe 40 % or 41, somewhere in that range.

33:51And I think it's pretty interesting because I think it's, you know, someone might look at that and I think in it, sort of at first blush say, well, isn't that evidence that skills all you need? And I do think there is something interesting there, right? It's like it's showing that, hey, the more training data these things have, the more sort of, you know, programs that they can spit out that might be kind of right. But it also shows that I think that new ideas are needed still. Like this outer loop is novel. Well, that might actually be Frontier LLM reasoning that Ryan published. And we're going to make all the approach whenever we put similar dark prize.

34:26We're going to open source all the code for all the reproducible solutions so folks can take these and apply them and try to reproduce them. And so we're using private closed model, open source models for the closed private data set. But yeah, I do think it's pretty interesting. How much innovation do you get when you... How much innovation we've got over the last few weeks? It's just a result of putting even just the awareness against the public. Are the folks at the big research labs like why are they not working towards this benchmark? Because it almost like when you explain the benchmark it seems so clear that obviously this is the thing you want to solve.

35:01You don't want to solve the memorizing of the textbook use case. Why do you think the folks at the big research labs aren't trying to solve this benchmark? Or are they? So I am aware of a handful of big ILAbs that have tried in the past several years ago. So this was perhaps at smaller scale with weaker models and things like that. One of the things I would hope is that actually more do you in the future actually love to see if we could make ARC AGI an actual measure on some of the model cards that get reported against future models. I think that would be a really cool thing. We're willing to do it.

35:34So if anyone is listening to this and wants to reach out and make that happen, more than happy to work with them and find some way to do that. If I had to guess, well, let me say, let me not guess unless, say, but what I have more sort of confidence in. I've been serving, once I got exposed to Francois to work again and was sort of deep, thinking, thinking deeply about our cage. I started serving a lot of my like friends and researchers and SF in the Bay area about had they heard of Francois and they heard of Arc. And Francois was a pretty good name recognition because he's really big on Twitter, been big on Twitter for many years.

36:11Probably nine out of 10 people I talked to like knew who Francois Shullay was. Maybe one in 10, two in 10 had heard about the Arc age at Yval. And probably half those were confused because there's like five other A .I. Yval's called Arc. And I had to like sort of do some, so I had to like, you know, disemphiguate with them. So it had really low awareness. This was one of the first things I asked for and so about when I met him for the first time in person this year I asked him why do you think that is why do you think you have such high -werness but arc has such low awareness and His answer effectively was That it's hard, you know the way that benchmarks Get like gain popularity and notoriety is we make progress towards them right researchers are working against it somebody has an idea, they have a breakthrough, they publish that in a paper, that paper gets picked up and cited by others, that generates awareness and attention.

37:04Other researchers say, ooh, interesting. Okay, something might be possible now on this really hard benchmark. And so you get the snowball effect of attention. And because ARC has endured with very low rates of progress, in fact, decelerating progress over the last four years, I think anyone kind of, in the lab looking at that, I would just say, well, maybe the time's not right for you. Maybe we don't have the idea set in the world. Maybe we don't have the scale we need yet in order to sort of beat this thing. And it looks like a toy and it doesn't like, you know, I don't fully understand why, I don't get the necessary importance or how it's qualitatively different.

37:37I haven't just spent that much time. I got a million other benchmarks I could use. And, you know, I think that's somewhat of the dynamic that has existed in the past. And it is one of the, again, reasons why we launched our price, right? I think there are some, there's lots of like market tools you can use to shape markets and shape innovation. And I think prizes do have a narrow spot where prizes can be outrageously effective. And it's where the idea is small. And it's like idea rate limited. One person or a small team can make that breakthrough. And it's very quickly and easily inspecable, reproducible, and built on top of it rapidly.

38:15And all those boxes got checked. And yeah, one of the reasons why I decided to go to our surprise. And you mentioned curiosity around AI or AGI dating back to college that was sparked a few years ago in the context of Zapier and has kind of been nurtured ever since. Beyond the curiosity, I'm curious why this is important, meaning if you could paint the picture of what life looks like for the world post -age AI, where we've defined it as the ability to efficiently acquire new skills. What do you think that version of the future looks like? Like, why is this an important thing to solve? The thing that I feel like I have a unique insight into at this point, having spent a lot of time thinking about Ark in this AGI definition is I suspect the advent of AGI is going to look very differently than most people expect.

39:10Especially of the group who are in the camp that AGI is undefinedable because it's so mythical and scary or big or awesome that we can't even hope to ever define. It's just going to be this magical special thing. It turns out, or something that I believe quite deeply, is the definitions are really important. Because definitions allow us to create benchmarks. And benchmarks allow us as a society to measure progress and set goals towards things that we care about and what to happen. And this idea of efficiently acquiring skill, one of the, you have talked about a handful of times today, but one of the direct near term things that you get from that is you get systems that can do exacting accuracy generalization from a small set of core priors and apply it towards novel solutions.

39:58That is again the number one problem that rate limits AI adoption for more real world use cases today. And so that's the, that's what you're going to see, you're going to see basically like the application layer of AI get like amazingly good at accuracy consistency, low hallucination rates, which is going to allow us to use it in a much more unfettered way and a much more trusted way because of because of like the underlying way that in which it's built. So I think that's like, and the reason I think that's important,

40:32I think that's the reason So what that set of capabilities is going to build on top of into the future, right? There's lots of unknowns, I think, about AI, AI, AI evolves beyond the actual inception moment of a system that can efficiently acquire skill. But I think it's going to be a much more gradual and incremental rollout, whether it's a lot of contact with reality as we build and engineer these systems, which is going to give us, as like a society, a lot of time to update based on what those capabilities, what it can do what it can do and make decisions at that point about how do we want to like deploy this technology where might we as a society say we don't want to deploy it for this set of use cases.

41:15You know, I think that's one of the reasons why I've been so sort of such a opponent I think of of open source H .I. progress with our prize is like right now I think the what I see happening is there's sort of this mythical story of a very bad outcome once we get to like super intelligence right. It's a very theoretical driven story. It's not grounded in sort of empirical evidence. It's basically based on sort of reasoning our way to like this outcome. And I think the only way that we can really, really effectively and truly set good policies by you have to look at what the systems can and can't do and then regulate or decide make decisions at that point about what what a can or can't do.

41:49I think anything else is sort of like, you know, you're cutting off potential really, really good futures way too early. And that's sort of what's happening I think with a lot of this early air regulation where, you know, I'll try to, you know, be the paint the good side of this picture. It's like, hey, I care, you know, maybe the risk of this bad outcome is so high in the future that we should like pause here. I think the risk of that is you've like trimmed off every possible good path of the future way too early. And the reason it's way too early is because we still need new ideas. We need new ideas for researchers.

42:19We need new ideas for students. We need new ideas from young people. We need new ideas from labs. Otherwise, we're actually, there's like a chance that we'd like never actually reach the like high degree of useful AI that we actually want. And so that's kind of my nuanced take, I think, on probably what the advent of AI looks like. I think it's much less likely to be a moment in time and much more likely to be a stair step of technology that we build on on top of past technology. And that creates a lot of moments to update beliefs based on what it can and can do. Do you have any predictions on when we'll cross 85 % on our price?

42:56You know, before the competition started, there was a, the first data scientist we hired at SAP, giving this idea a long time ago to stuck with Mickey. He said, the longer it goes, the longer it goes. And so it's this idea that like the longer something takes the more you should update your prior about, it's going to take longer. So coming into this year, my expectation was like, hey, at least three or four years, probably before we get to the grand prize mark based on sort of like past, basically in the past track record. Having seen what we've seen over the last two or three weeks, though, I think it is quite likely we get to 50 % during this competition period.

43:32I would be very surprised, you know, I'd be surprised in a good way if we actually get to the 85 % grand prize in this competition period. But I think it is not unlikely that we crest the 50 % mark before the end of our middle of November, which is when the contest period ends for 2024. before. And is there a good why now? Because people have been trying at this for five years now. And you're galvanizing interest around it. And now a lot more resources down the world are interested in AI and solving hard problems. But is there a good why now in terms of enabling techniques, technologies, et cetera, that's different now than five years ago when friends who have first kind of defined the benchmark?

44:11If it is true that deep learning is an important part of the solution, right? a deep learning guided program to this ascension, or a DSL that is generated on the fly through deep learning technique. If that's true, the world has a lot of experience now on building and engineering and scaling such systems over the last three or four years, and there's a lot more compute online, which brings the cost down into a territory where some of these things may have just been out of practical cost before. For example, actually, Ryan Greenblatt's solution right now how is maxing out our cost limits we're going to have against the public leaderboard costs $10 ,000 to generate the 8 ,000 reasoning sample traces from GPD40 that he then deterministically checks.

44:58And so that would have been a technique that would not have been possible three or four years ago in any way it's a guard. So I think if it is true that there's a minimum amount of scale, the necessary to be dark, I think, hey, we've gotten more of that in the last three or four years than we had when the first come to SRAM. And then I think the other why now is just, can't just largely due to awareness. If it truly is, actually let me answer in the opposite. Like I think the risk that it is, the reason we launched our prize is that it is actually not why now, right? It's like actually it's not going to happen is the problem.

45:32It's not a why now. It's not like oh the ideas are in the world we just need to like get people to work on it. The risk is that it's actually not why now. And not why now is I think a much more interesting story right around this like LM driven, you know, focused attention on sort of LM solutions only, the closed research due to the competitive dynamics, all of these things have like shifted and shaped attention away from the new ideas and towards scale, towards LM's, towards like, you know, yeah, application layer AI. And I think that we think we need some shaping, reshaping back towards the new idea set.

46:06So hopefully, The Y now is because ARC has now lots of potential to be seen. Do you think LLMs will be part of the solution? I'm curious what you think of it. It seems like in the big research route labs right now, all of the frontier research is around. Let's merge kind of LLMs with the insights that you get from the inference time computer and the Q -star, AlphaGo stuff. I'm curious what you think of that kind of direction of research. I, there's some like pretty interesting research I've come across with like transformers that the transformers are capable as an architecture of representing very deep deductive rates in change with 100 % accuracy.

46:45I think this is interesting. And the challenge with them is actually we just don't have the learning algorithm. Backpropagation is an ineffective learning algorithm in order to teach a transform architecture, a set of network weights that can do deductive reasoning with 100 % accuracy. And so I think it's possible that the systems that we kind of, or at least the core concepts that underlie language models have sufficient capability in order to do this type of reasoning. And we have not yet discovered the like algorithm that can train the model in the right way. We haven't quite discovered the right outer loop around the transformer that is going to do the program.

47:17Since this is engine or the DSL generator, I feel more confident in saying that like deep learning is almost certainly going to be a part of the grand prize in some way. I bet it's that it won't, I'm pretty confident it wouldn't be just like a pure deterministic program is going to be how it gets solved. I think transformers are effectively the technology that is the most, has the highest degree of awareness. In research like literature, there's a lot of hardware now that's going towards accelerating transformers, models I think actually just a roller that was like an ASIC that got announced to recently that's trying to accelerate like the transform architecture.

47:52So do they agree that actually like some degree of like scalar computers necessary to be dark? I think those are like things that I would say I'm bullish on sort of the transform architecture. sure though, I would point out that the search base of alternative architectures is quite rich. Right? We've had maybe like nine or ten now mainline architectures for transformers to LSTM, CNNs, RNNs, XLSTM, state space models. This would suggest that the search base of those architectures is quite rich actually and they all have like slightly different varying properties. So like I think it's certainly possible that someone comes up with some innovation there.

48:27I'm less confident or like bullish that LOMs in their exact form are going to be part of 85 % solution though. I would think it probably like a subcomponent of the architecture instead of the entire application system itself. When somebody does ultimately hit the 85 % threshold, what do you hope they do with the solution? Like what would you like to see out of that person? other than submit it to the leaderboard. So this is one thing we didn't talk about a ton, but one of our prizes goals is to accelerate open progress towards AGI. So we are going to be requiring that in order to claim the prize money, you do have to publicly share and reproduce publicly share reproducible code and put it into the public domain.

49:11And this goes for both the public leaderboard and the official competition leaderboard as well. You know, this is with the spirit of trying to re -accelerate open progress again so that we have research in small ways out in public that other researchers can build on top of and hopefully steer supper would actually, actually, billing a GI and not getting stuck in sort of the plateau that we're at that we're in today. But I think that's probably my first pick up. I've actually seen a handful of people online that have said, hey, if you've got a solution in our KGI, I'll give you this like, you know, a million dollar offer.

49:43or like, I'm pretty much like a real person. Yeah, exactly. Which, you know, at one hand, I'm like, yeah, okay. That's kind of interesting. But, you know, in the other end, I'm kind of like, you know, I think that's like great awareness. And it shows actually like the importance of solving this. I think more people are starting to become aware of the lack of progress, lack of frontier progress. I think Arca is kind of becoming a lighting rod for folks that want a actual measure towards this, I think growing sentiment in the field today. Should we close out with some rapid fire questions? Yeah, let's do it.

50:17Okay. Who do you admire most in AI? I mean, Fransfalsha lays a bit of a cop, I'm answering. I wouldn't have go find a dark prize if I didn't admire and respect his work over the last four years. I mean, I think the two people that I have learned the most from, like directly, have inspired my own belief and work, Rich Sutton and Shalei. Both of them published papers in 2019, Rich Sutton published The Better Lasting, which I think is fairly well known on the industry at this point. I think his idea set is quite right there with maybe the one aspect that the one aspect that has not yet had scale applied to it is architecture itself.

50:53We certainly have unbiased search and learning applied on the inference side and the training side, but every architecture still is a very human handcrafted story and journey to it, which is an important insight, I think, about that. And then, I'm measuring intelligence from Shilay in 2019. And I think both of these papers are in like, or I guess maybe, sentence was more of a blog post. But both of these pieces of writing, I think are very important because history has proven the right as time has gone on. Language models, transformer scale has sort of shown some ideas to be even more true than they were in 2018.

51:24And I think the endurance of ARC has shown Fratz -Wastephanesche rage eye to be more and more true as time has gone on. What is your most contrarian point of view in AI? I feel like everything we talk about today, we'll be able to go to group with the Hanzo. All right, we'll count it. Like, Gale is not all you need. Well, new ideas are needed. What's your favorite AI app, other than Zapier? Let me look and see, what do I have, I have a handful. I think I'm not gonna, like, surprise you with anything. So I've got chat to BT, perplexity and cloud, and I'm a paying user of all three of those services.

52:06is, you know, one interesting thing actually I'll add, I have gotten way more value out of language model based tooling over the last, like, call it six months. Then I ever did in the first aspect when I was starting to start work on it, Zapier. And it's because one of the things are really, really good at, the thing that they're perhaps like best at is summarizing tons of unstructured text and helping be like an educational tutorial for you. So it's significantly ramped up my learning rate on actually building with AI, built these fundamental different architectures. Starting to do model training, something's after you're less than done, but I started to do myself over the last six months to get a sense of that type of work.

52:51Language model and AI tooling has definitely accelerated my learning process. All right, last question. Let's do something optimistic, something that we can all dream about. What change in the world are you most excited to see over the next five or 10 years as a result of AI? I think the I've always wanted to like live in the future. I think that's maybe a something that has always driven me towards like working on like frontier tech. I've always you know, bought the latest gadget always tried the latest app. I think it's led me to work on Zapier and AI and And it's one of the reasons I'm working a GYR now is because I think it's like the biggest thing that you can potentially have an influence on trying to pull forward that future.

53:35You know, I personally get, I think one of the things that feels very limited for me right now is that with the nearer form of AI that we have, if we never get a GI, What that will mean is that we will always be rate limited on developing things by the human that's in the loop. And that means we will never have AI systems that can invent and discover and sort of innovate alongside humans and really pull forward and push forward the frontier in I think a lot of really interesting ways. I understand more about the universe and then discover new pharmaceutical things. You know, in discovery physics, discover how to build AI.

54:26You know, I think we're always going to be rate limited by the human today. And I think if you just sort of care about living in the future and you want to pull forward the good aspects of the future, some form of age eyes is I started to do that. Awesome. Thank you, Mike. Thank you both for having me.

55:08Swings and

From the publisher

As impressive as LLMs are, the growing consensus is that language, scale and compute won’t get us to AGI. Although many AI benchmarks have quickly achieved human-level performance, there is one eval that has barely budged since it was created in 2019.

Google researcher François Chollet wrote a paper that year defining intelligence as skill-acquisition efficiency—the ability to learn new skills as humans do, from a small number of examples. To make it testable he proposed a new benchmark, the Abstraction and Reasoning Corpus (ARC), designed to be easy for humans, but hard for AI. Notably, it doesn’t rely on language.

Zapier co-founder Mike Knoop read Chollet’s paper as the LLM wave was rising. He worked quickly to integrate generative AI into Zapier’s product, but kept coming back to the lack of progress on the ARC benchmark. In June, Knoop and Chollet launched the ARC Prize, a public competition offering more than $1M to beat and open-source a solution to the ARC-AGI eval.

In this episode Mike talks about the new ideas required to solve ARC, shares updates from the first two weeks of the competition, and shares why he’s excited for AGI systems that can innovate alongside humans.

Hosted by: Sonya Huang and Pat Grady, Sequoia Capital 

Mentioned:

Chain-of-Thought Prompting Elicits Reasoning in Large Language Models: The 2019 paper that first caught Mike’s attention about the capabilities of LLMs

On the Measure of Intelligence: 2019 paper by Google researcher François Chollet that introduced the ARC benchmark, which remains unbeaten

ARC Prize 2024: The $1M+ competition Mike and François have launched to drive interest in solving the ARC-AGI eval

Sequence to Sequence Learning with Neural Networks: Ilya Sutskever paper from 2014 that influenced the direction of machine translation with deep neural networks.

Etched: Luke Miles on LessWrong wrote about the first ASIC chip that accelerates transformers on silicon

Kaggle: The leading data science competition platform and online community, acquired by Google in 2017

Lab42: Swiss AU lab that hosted ARCathon precursor to ARC Prize

Jack Cole: Researcher on team that was #1 on the leaderboard for ARCathon

Ryan Greenblatt: Researcher with current high score (50%) on ARC public leaderboard

(00:00) Introduction
(01:51) AI at Zapier
(08:31) What is ARC AGI?
(13:25) What does it mean to efficiently acquire a new skill?
(19:03) What approaches will succeed?
(21:11) A little bit of a different shape
(25:59) The role of code generation and program synthesis
(29:11) What types of people are working on this?
(31:45) Trying to prove you wrong
(34:50) Where are the big labs?
(38:21) The world post-AGI
(42:51) When will we cross 85% on ARC AGI?
(46:12) Will LLMs be part of the solution?
(50:13) Lightning round

More from Training Data

All 110 episodes
Zapier’s Mike Knoop launches ARC Prize to Jumpstart New Ideas for AGITraining Data · 55 min
Listen in VO