Founder Eric Steinberger on Magic’s Counterintuitive Approach to Pursuing AGI

10 Sep 2024 · 51 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Podcast Episode Notes: Founder Eric Steinberger on Magic’s Counterintuitive Approach to Pursuing AGI

Podcast Overview Title: Training Data Description: Conversations with AI builders and researchers to explore the evolving technologies of AI and their implications for technology, business, and society. Episode Title: Founder Eric Steinberger on Magic’s Counterintuitive Approach to Pursuing AGI Episode Description: Discusses Eric Steinberger's journey from a high school student to founder of Magic.dev, focusing on the pursuit of Artificial General Intelligence (AGI) and innovative approaches to software engineering.

Key Participants

  • Sonya Huang - Host, Sequoia Capital
  • Eric Steinberger - Founder and CEO of Magic.dev
  • David Silver - DeepMind researcher
  • Noam Brown - Reinforcement learning researcher at DeepMind and collaborator with Eric
  • Johannes Heinrich - PhD student and mentor to Eric

Episode Structure 00:00 - Introduction

  • Introduction of the host and guest.

01:39 - Eric's Early Life

  • Eric, a wunderkind from Vienna, develops an early obsession with AI through his interest in math.

04:56 - Collaboration with Noam Brown

  • Describes his initial collaboration with Noam Brown at DeepMind and the learning experience.

08:00 - Balancing Multiple Projects

  • Eric discusses the challenges of juggling multiple responsibilities and the importance of focus.

10:37 - AGI To-Do List

  • Eric shares his thoughts on the current state of AGI, emphasizing the need for focused research.

20:35 - Advice for Young Researchers

  • Encourages young researchers to pursue their goals fiercely and be proactive in seeking mentorship.

29:59 - What is Magic?

  • An introduction to Magic.dev and its mission to automate software engineering.

36:12 - Competing Against Established Companies

  • Discusses the challenges of competing against major players in the AI space.

40:10 - AI as a Colleague

  • Eric emphasizes the importance of AI feeling like a colleague, capable of self-management.

44:30 - Lightning Round

  • Quickfire questions and answers.

47:50 - Bonus Round: 200M Token Context Announcement

  • Announcement regarding Magic’s advancements in context processing capabilities.

Key Concepts and Discussions

  • Early Passion for AI: Eric's fascination began at age 14 when he recognized AI as a meaningful pursuit.
  • Collaboration is Key: Importance of mentorship and collaboration in research success.
  • AGI as a Focus: Eric’s conviction that AGI is not far off and the belief that it can be achieved through clever engineering.
  • Importance of Reading: Eric emphasizes the need for voracious reading and understanding of existing research to synthesize new ideas.
  • Automating Software Engineering: Magic’s mission to create AI capable of performing software engineering tasks, which Eric views as a critical step toward AGI.
  • Competitive Landscape: Eric's perspective on competing with larger companies like OpenAI, emphasizing the necessity of owning proprietary models.
  • Ideal Team Size: Eric discusses optimal team sizes for efficiency in research and development, highlighting the need for focus.
  • AI as a Collaborative Colleague: The envisioned future where AI systems operate effectively alongside human engineers, managing tasks autonomously.

Key Takeaways

  • Pursue Your Passion: Following one’s passion relentlessly can lead to significant breakthroughs in research.
  • Synthesis of Ideas: Successful AI research often requires synthesizing ideas from various sources and disciplines.
  • AGI Development is Near: Eric believes that advancements in AI will escalate rapidly, leading to substantial societal impacts within the next decade.
  • Open Source Contribution: Eric’s commitment to open-sourcing Magic’s models and evaluation tools reflects a trend towards collaboration and transparency in the AI community.
  • Focus on Quality and Efficiency: Achieving a high-quality AI assistant requires extensive resources and a clear understanding of the problems to be solved.

Conclusion The conversation with Eric Steinberger reveals not only his impressive background and deep understanding of AI but also his ambitious vision for the future of AGI and the role of Magic.dev in that journey. His insights into research, collaboration, and the importance of quality in AI development provide valuable guidance for aspiring researchers and entrepreneurs in the field.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00The thing that remains to be solved is general domain, long horizon, reliability, and I think you need in friend's time compute test and compute for that. When you try to prove a new theorem in math or when you're writing a large software program or when you're writing an essay of reasonable complexity, you usually wouldn't write a token by token. You'd want to think quite hard about some of those tokens. and finding ways to spend not 1x or 2x or 10x but a million x the resources on that token in a productive way I think it's really important that is probably the last big problem

1:00Hi and welcome to Training Data. I'm delighted to share today's episode with Eric Steinberger, founder and CEO of Magic. Eric has an epic backstory as a researcher, having caught the attention of Noam Brown and becoming one of his research collaborators while still a student in high school. Eric is known for his exquisite research taste as well as his big ambition to build an AI software engineer. We're excited to ask Eric about what it takes to build a full stack company in AI, his ambitions for magic, and what separates a good AI researcher from a legendary one. Eric, welcome to the show. Thank you so much for joining us.

1:37Thank you for having me, Sonny. Okay, so let's start with who's Eric? You're a Vienna -born Wundergind, whose early passion for math turned into a... I think what you described as a full -fledged obsession with AI by age 14. Take us back to age 14, Eric. Like, what were you up to? How did you become so obsessed with AI. Thank you, Zanya. I think I just have my midlife crisis when I was 14. And I was just looking for something meaningful to do. Spent about a year looking at physics, math, bio, medicine, just anything really that seemed valuable to the world. And at some point bumped into just simply the idea of AI.

2:15It hadn't sort of occurred to me until then. And if you could just build a system, a computer system, and it could do all this other stuff for me. Like, great, I don't have to decide. So it felt like my decision paralysis was sort of resolved. It was this weird moment where I could just see the next 30 years of my life unfold in front of me. And I was like, okay, this is clearly what's going to happen. Like I have to do this. And it was quite nice. I like predictability. So it was great to like, no, what the world will look like. And you started loving math. Like, why AI then? I think I'm naturally attracted to math.

2:51It's just what my brain sort of gravitates to. AI just seems useful. The thing that's most important to me is just what is useful for humanity in the world and math is nice but not useful at some point like you know 17 dimensional spheres are probably not going to be the best career choice if you want to be useful so It seemed like something that I could get good at but also just the most important thing ever. And so it was a very clear choice. AI just just like, it was clear 10 years ago. It just, it wasn't close. And now it's close and clear. Can you tell the story of how you got to fair? I think it is such an epic story.

3:32Sure. I mean, so when I started at 14, I didn't really know how to program. I didn't get into programming out of curiosity about computers. as I just wanted to solve AI, basically. So after a couple years of just warming up on my own, I reached out to one of David Silver's PhD students, who was the AlphaGo DeepMind co -founder. And this PhD student, I guess at that point, he was a graduate and sort of worked at DeepMind. I asked him if you could spend a year with me just every two weeks bashing my work, trying to do, so super speed up, mini PhD -like experience where I could just learn how to do research.

4:17I sent him this giant email, you could print it out, I don't know how many pages it would be, but it'd be a lot of pages. I was basically just saying, I want to build this algorithm you made in your PhD. I'm going to beat this algorithm you made in your PhD. Here's a list of 10 ideas. I don't know if they're going to work and I think I need your help to figure that out. And then over a year we eventually got there and he was like as they say Harness, Yannis was kind enough to just patch me every two weeks roughly. And yeah, it was brutal dude. Like, I was like, he always do the standard, you know?

4:52Don't be nice just because I've been high school. Yeah, I was in high school. And then when we were done, I just graduated in high school when I finished. So this is when I finished the project that I was trying to get to with this. And then Noam Brown, who is obviously one of the best RL researchers in the world, reached out. Because he worked on something similar, turns out. And we sort of had some ideas that were very similar. And some ideas that were a little different. And so we just published this and Eri Stoud. And then I got to work with Noam Brown for two years, which was great. And then that continued.

5:25And so I got back to the other year. And you were a high schooler. He was Noam Brown. Well, I mean he published a paper called Deep Counterfactual Regroup Minimization and I published a paper called Single Deep Counterfactual Regroup Minimization and Mine Beat Hissed by a little bit and... So you want to know him? Brown is a high schooler. I think I just graduated and I also took him like three months to write this paper and it took me a couple years. But yeah, I mean slightly. I'm sure he would have gone up with this the next day. but the gap between the two things. But it was just, yeah, it was a obsession.

6:00It's the right word. I do things like 100%. And yeah, so that was a little fun. Keep working in our all with No Brown for a while. And yeah, so that's how I got to fair. No Brown worked at fair at the time, and he reached out. I was actually at university then. And basically just work part time as a researcher at fair file studying. Anyway, that was it. That's awesome. It was a lot of fun. No, it was great. Like the brainstorm ping pong sessions with no brown dude. There's like nothing like this. Or you would just think there's like this problem. And it sort of like, you know, maybe you would like start a six -month research for like no.

6:42No, I mean, I would get on a call and it was like, we just discuss it and it's done. I love that. I love that. What makes him so great as a researcher? I think it's a number of things. As a researcher, generally, from more from a meta -level, he is fantastic at picking the right problems and then spending a long time just grinding to make it better and better and better and better. So he's very good at the whole like compounding thing in research. Also making bets that aren't obviously the right bets when he makes them, because he makes them earlier, I suppose, and making them say differently.

7:16So you're generally very good at picking problems and then attacking them consistently. It's also just very smart, I guess that helps. It works really hard. He used to do 100 -hour week storing his PhD. I don't know if he still does them, but he used to work really, really hard during his PhD. I imagine he still is. Okay, so no more range for you to become a Researcher at Fair while you were still a university student. Yeah, I was from. I was in my first semester, I think, or something. So you were juggling that, you were juggling being a collaborator to know him at the fair, and then you became obsessed with yet another problem, climate change, and I actually started an NGO that is incredibly popular, climate science.

7:56So you just didn't have enough on your plate. That was actually too crazy, that's when I dropped out, I was like this is crazy, this is too much. I can do two things, I cannot do three, it was like my conclusion after I did three months of that, that was terrible. That was fucking awful. Doing all three things. Because you just can't do well at three things. I mean, Elon can, but maybe I'll learn it in 10 years. But I couldn't at the time. So I dropped out at the time. But yeah, I started NGO. I generally think charity stuff is awesome and hugely underappreciated. It's super high status to start a startup.

8:36But I think it should be equally cool to start a charity. You're like helping the world in other ways. And so, yeah, I mean, we started it as a, it's a nonprofit, but we started it like a startup. People were working insanely hard. We had clear objectives. It was, you know, a software product effectively. It was much more similar to a startup with the exception that there was no money and no money out. I can flip. I mean, just very weird, but the, yeah, it's mostly volunteer driven, or I guess is, I just no longer run it. Yeah. Yeah, that was an interesting experience. You'd think I would learn transfer a lot from running a Quote Enquad Company to running a Quote Enquad Company now, but they are so different that I could like, there was no transfer at all between climate science and magic.

9:21Like a thousand volunteer, 20 hardcore engineers. No money at all. I can't tell you how much we raise because it's not enough yet, but, you know, giant. It's like completely different in every imaginable way. So, but yeah, it's not fun. So Eric, climate science became an incredibly successful nonprofit. I wasn't just sending nonprofits. What made you decide to kind of hand over the reins and hand over the torch on that and then go start a company in AI? I just thought AGI was further away when we started it at all. I wouldn't ever have started anything else if I thought AGI was so close. And once I realized it is, there was just like no other.

10:01I mean, my initial thing was always AI. That's what I did as a kid. I care about various issues in the world, but it's none of them are my unique calling in any way. I just, you know, I hopefully have been in a position to donate a bunch of money on whatever, but the thing I care about fundamentally is AGI, and it was like, oh damn it, this is not 20 years away. So I have been running around with this AGI to -do list, which is somewhat of a meme, uh, uh, uh, internally. Because it's sort of like, just like going through it. Then we're trying like to fix all these problems. You have seen it. Yeah.

10:37So we showed it to you. Um, and, uh, I've been running around with a version of this. There's actually like a 2017 or so. I was still in high school. I got it. It's I don't know why but some conference invited me to present my AGI to the list. It was wrong at the time. I was also, I was also sure it was wrong. But, um, at some point, so that was one thing I just couldn't at all figure out. And I don't like Blue Sky Research in the sense of like just staring at a wall and trying to like figure out what the right question is. I really like to have the question and then look for the right answer when starting an intense project because you need to know which direction you run into really plan for it.

11:13And many things seem clear, but it seemed completely unclear how to make these models reason in the general domain. And that became more clear with language models, especially code models. And so, yeah, when I saw just some of these early, some of the early results in this basic, I was like, okay, I know all this stuff from the RL world, I have a bunch of other thoughts. This seems great, like we should just take LMS and make them do the RL stuff. It's a very simple kind of proposal, but I think that's sort of where, I mean, it makes a lot of sense. RL has been going this for 10 years. it works in restricted domains.

11:56If you can make something work in 20 restricted domains, and you have something else that works in a general domain, if you can combine them, maybe you'll get both the X and the Y axis, and then you have your beautiful top right corner of the matrix. It seems like it seems persuable. When something becomes, when something as important as AGI becomes an actually executable to do, it's obviously there are still things to figure out, details of the algorithms, how do you make it efficient, etc, etc. we do everything at all. Many, many things to figure out, but the direction was clear. Yeah. So it seemed like the right moment.

12:32Okay, we're going to circle back to your AGI to do list later because I'm curious about it. Sure. I want to brag about you for minutes because it might be wrong still. I don't know, until we have AGI, it is a hypothetical AGI to do list, but we're frying. I think the research field is tracking pretty closely to your to do list. I want to brag about you for a minute. I think you've been incredibly humble about your background, but you know, as a high school student, you did catch Noam Brown's eye and as, you know, as one of his colleagues at Fair, you became one of his top collaborators, not even just one of many, because there's such talented people that work there, but you're one of his top collaborators.

13:10And, you know, when I speak to folks that know you, they just say extraordinary things about your capabilities as a researcher, your creativity, your work ethic, as far as I can tell you work non -stop. Thank you, texted me at 2am in preparation for this podcast. So I think it's safe to say. No, thank you silent mode. Anyways, I think it's safe to say that you are one of the brightest minds of the current research generation already and will certainly be one of the legends that people talk about for the next decade. And so with that in mind, I'd love to ask you some questions of advice for aspiring researchers.

13:43And so maybe first off, you did it all from a very untraditional background. How did you do it and like do you think that's like what advice would you have to others in your shoes? I can only really speak for it a sort of profile of goals and person I am. I think I was lucky in the sense that I knew very very early with Martinez as we said exactly what I wanted to do with my life. I had no doubt at all and then not certainly can be paralyzing to a lot of people. I also had a very clear sense that I did not at all have a plan B. Like there was no water, a path in life that I would have been even like remotely above the neutral line on.

14:22It had to be build AGI, everything else is completely irrelevant. So, you know, I understand, so for many people like, you know, a well -paying job at Google is a great achievement. I mean, I would, if it's on AGI, it's fine. But you get what I mean, like I just knew that there was nothing else I could do and like be fulfilled in life I look back when I'm 90 and be happy So in a way like burning the boats very very early gives you the opportunity to Just be be like do things that you do otherwise do 10 years later, which again even like I sucked at the beginning It took me two months to understand the first paper.

14:59I tried to understand and I was terrible at programming for a long time. But when you're a teenager, you're like a decent researcher, you don't have to be great. That gets you things like a great mentor who then bachelors you for a year, which was very, very helpful. And then you get better, and you're still young, so you serve your brain shapes more easily maybe. I don't know. So I feel like I benefited a lot from being early. But within that, I'd say just like go for the end goal immediately, doing anything sort of like, oh I'm gonna do a PhD because I need a PhD to get a jub -ass -hole bullshit.

15:36Like you don't, it's just completely bullshit. The other thing is like writing five -page emails to people actually works. Writing like, I get a lot of these like two paragraph meh things now and grateful like get emails but don't, I understand now why people think this stuff doesn't work. It certainly does when you're like, here is how I'm gonna beat your algorithm. Please help me. Five pages, at least in my experience, every single time anyone I I want to help from in this way was very helpful. So I suppose be proactive in seeking like the best people in the world to in a time -efficient manner just distilled their brain into yours and show them that you can make use of that.

16:18If you if you if you tell someone who's very good effectively, hey, I'm gonna make good use of this. If you want to coach someone, I would love to be that person. They'll usually do it. They won't do it for 10 people, but if they do it for one or two, that's enough. You just have to win that seat, I guess. So that's been really helpful in my experience. Also, just not shying away from learning new things. Like, again, I didn't get into programming because I'm curious about computers. I'm not very curious about computers. I just like AI and that computers are the thing that are necessary to do.

16:47So it's fine. I enjoy programming now. It's great. But I wouldn't have gotten into it, I think, if it wasn't But still, you get into it. So don't be shy. We need to be a lot of people who don't know how you'd implement an LLM. And it's kind of crazy to me if you're a researcher, and you couldn't implement charting or whatever. Like it's just insane. So really understanding the whole stack, going down to, but not bottom up, really top down. Here's the thing I care about. This is the problem I want to solve. What do I need? What do I need? And then all the way down. And I know like much more competent people at kernel programming and hardware design or whatever than I could ever dream up to be.

17:27But I understand enough of it to do better work at the top of the stack than I could if I didn't.

17:36So I think fundamentally you need to understand the domain you work in. It's also really good to just read everything. Like I used to read, I don't know, I don't have a precise number, but just every paper, I could, every paper with C basically. And eventually you get so fast at it that you can, like, that's feasible. And you, like, build a database in your head of, like, oh, this is similar to this thing. This is sort of like, this is sort of my eye -opening moment. Or Bill Gates has this interview, like, oh, yeah, you've learned enough things. They're all, like, similar to each other. So it's not linear.

18:03It gets easier. And at that point, I was like, I should read every paper. And so thanks for the advice, Bill. Obviously, it was, like, through a video. I never met him. But the, so I just started reading every paper. And that's really, really helpful because a lot of the best ideas that we had that work really well now at magic. like we're enabled by random things that are like, oh, it would never work without this random thing that I would have to have come up with in tandem. But because I have this database in my head, I can go like, oh yeah, like this. And then so often, like one good idea is enabled by three other ideas that other stuff come up with.

18:35And so it's always just like this composition of stuff. So having a large database is really helpful. Yeah, and then just never stop. Like never, never stop. It takes like ages to do good stuff, to do good work. at any point. I was actually one moment with Johannes, the Deep Rind Research Scientist who mentored me for a year in high school, where we had a version of the algorithm that wasn't very good. It was all right and we're thinking that guys should be published this, like we were both not really happy about it, and he was like close to giving up on me. I was like, well you know, like maybe this is just not gonna work, like I wouldn't want to publish this, and so that was like, dude, fuck you, I'm just gonna get this done, and then we got it done like a one or two later.

19:18And so I think that I remember going on a walk after this and just being like, can I do this? Yeah, I don't know if I can do this. But there is no other option. So I just better get it done. And then I went back home and I started programming again. It was still like sad that day, but the next day was fine again, and they just keep going. So I think you have to, I think that was a pretty formative experience because I actually wasn't sure if I could do it. And then we just did it like super soon after. So I really haven't felt that in the same level of doubt and pressure since then, which has enabled, it's actually beneficial.

19:54You have to be realistic, but if you stop, yeah, I mean, so anyway. So I think those would be the main things. Also be really fucking honest about what you suck at to yourself, because otherwise you're never going to get good at it. like you need to search for the bad things. And instead of like trying, actually, I think like as a researcher, betting on your strengths is good only to the extent that you don't have necessary conditions that are completely missing. Like you can't bet on your strengths if they're not enabled. Is again, like back to the engineering thing, for example. So yeah, I don't know, I'm rambling, but stuff like that.

20:34That's great. That is such a fascinating glimpse since the inner mind of what it takes to be a great researcher, and behind all the glamour of training large models. And so thank you for providing that peak. I'm really glad that you mentioned reading every paper voraciously and having this database in your head. Because one thing I've heard from your collaborators is that your superpower is understanding and absorbing new research. And so I'm curious, do you agree? Do you think that is your superpower as a researcher? What traits do you think have made you such an exceptional researcher. So I think initially in DRL work I did, it was synthesis where I would read every paper and I would go like this thing plus this thing plus that thing with this modification.

21:14I think that's what they would mean. That yes, it was definitely very helpful. I think is a good way to do research. Generally there is enough work for synthesis to be a successful strategy. I guess to an extent it's still that I tried very hard after it. This is actually, in reading, it's like you bring this up. I realized this and I tried very hard to get better at leaps like coming up with totally alien crap that just there's no reference for it at all. And because ultimately like, so if you take like the transformer, for example, like attention existed, the idea of stacking a bunch of LSTM blocks existed and you just have to remove the idea of recurrence, really like that and like a bunch of couple other things that were necessary, right?

21:55Residual streams like the residual update and transformer existence from existed from ResNet. So it's like, it's synthesis, but there is an amount of leap in there to make it all work. It's a little more complex than just taking components and putting them together. You need to come up with new things too. Like, you know, the normalization and the head, square square, the ad rejection, correct? But anyway, everyone now knows this. But roughly, like you should do some organization. So there are some new ideas in there that really help make it work. But it's still a larger amount of synthesis. So I suppose like most good ideas are synthesis, but they're always some like in the best ideas are some leaps and I'm trying to get better at those But still it's mostly I guess like take five things and throw away the stuff that doesn't work in them Make the things work and Don't forget about yeah, I think I think some Some stuff needs leaps but but yeah, I guess like no, that's a recipe like take all of them make them super efficient block context giant, throw our L on it, make it all work together.

22:58It's still mostly synthesis, I guess you're right. Who do you admire most in the research world and like what do you think those folks superpowers are? Shizur, no Shizur. He is. He is. What is his superpower? I guess to an extent synthesis. He is just the best at synthesis. He is also great at everything in the stack. He has no weakness really. He can implement the whole thing on his own if he had to run it. He sees the future in a way that's very unconstrained and I think I run sort of crediting a number of the labs for scaling laws. was this guy made a presentation where he was zipping through essays or a completion of whatever written by like models of various scale.

23:52I was like this is 100 million parameter model. This is a 300 million parameter model. This is a billion parameter model. This is a 5 billion parameter model. This is on YouTube summer. It's hilarious. I knew he could like, well, what if we make this bigger? It's sort of presenting it this hilarious way. And I was also super scientific about it. I think no, this generally just, if I had to put it, it's very intuitive. I think there are a lot of labs and researchers are sort of, I think this is not a bad thing. It's very good. A very eval is driven very mechanical, right? It's sort of very empirical in a way.

24:26No, I'm sort of just nose. He's like, God, this would work. And then it works. So I think that's a superpower. It's just extremely great synthesis. Yes, the larger data base, because he's been around for so long. You just you literally knows everything. I mean he bent the half of the stuff that everyone's doing now There's no one to convince I'd say There are a number of other people I guess Just you you shouldn't feel like out of all the people who are sort of the OGs of deep learning I think you could have hinted there's but deserves by part of most credit just as he like when through all the bashing when And when it was like, oh, it is will never work.

25:08It's like trading tiny, tiny, tiny, tiny, tiny things. They're like, it's will never work. And he's not gonna stuck with it. I think that's that level of grit and belief and something that is now obviously working deserves a huge amount of credit whether Capsuleck's work or not, whatever. Yeah. He, like, you know, it's incredible to come to something like the conclusions that the world is at now. And if you look at some of the older papers, a lot of the ideas that are important now were in there already. So that's important. And I think he just deserves a ton of credit. And Noam Brown had the Army of Noam's.

Read the full transcript

25:47Noam Brown. I should name my kid Noam as well. It's a very good strategy. It's a great strategy actually. I think 100 % of Noam's that are somewhat popular and well known in the research community are great. Yeah, no, he's also amazing. I mean a number of labs were working on what he was working on during his PhD And he basically soloed the thing and was like way better and way faster than labs that put 10 people Including some really famous names on it and if you just book it the paper trail and track record like And I was like like here's like the rest of the field and then no I'm 100x efficiency And then here's the rest of the field that no I'm like does it again?

26:27And it's like it's just consistent. I think the consistency with which he has just bashed out these 100x multipliers in RL data efficiency and computer efficiency is crazy. Yeah. So, no, the no -oom army is pretty good. I want to go back to this concept of, you know, leaps are still needed in research and that you still have this AGI to do list. Yes. What do you think are the most interesting unsolved problems in AI right now? Well, so a lot of it is solved now, I think. And the thing that remains to be solved is general domain, long horizon, reliability, and I think you need in -friends time compute, test time compute for that.

27:09So you'd want, when you try to prove a new theorem in math or when you're writing a large software program or when you're writing an essay of reasonable complexity, you usually wouldn't write a token by token. You want to think quite hard about some of those tokens and finding ways to spend not 1x or 2x or 10x but a million x the resources on that token in a productive way I think it's really important. That is probably the last big problem. That's the last one. I hope so. I think it's reasonable to think that is the last big I mean look over the last few years all of this other stuff got sold like oh can we do multi -model things can we do long Context can we do it all is gone a reasonably smart models, you know they're quite efficient now in terms of cost that I mean It has to be a reality deny our job.

28:07Not see what's coming I mean this is just a this is like a realization through a lot of people in the space But like RL has been doing this some, or like ages. So it's just like so clear that you need to do that. Or maybe you don't need to. Maybe you can get away without doing it, which would be insane. But if you don't need to, it will still help you a lot. Like it's just like, do I want to spend a billion dollars on my pre -training run? And then like a little bit more money on inference. Or do I need to spend 10 billion dollars on my pre -training on an I'd rather, you know, like 10 billion would be great.

28:38But I'm going to be, I'm going to be, I'm going to prefer spending one. And it's bringing like the LLM and RL world together? Is that like a research problem? Like there's still like fundamental like unsolved science problems? Or is that like a you know we have the recipe we just need to do it and have the compute and the data? I think there is no public successful recipe right now. There are good ideas like okay anybody can take best of them, make end large enough it's sort of you know it's not terrible. Yeah. So there are ideas. I don't know that the final idea exists. I think there's just a lot of room up from what is currently known, but there are ideas.

29:17See, I think it's very unlikely that even if you stop progress in research, we would not at some point hit something that everyone would agree as HGI is. It's just that I think we can do better and maybe it couldn't solve Riemann, right? Maybe it couldn't do all these like super -wide things, but it'd be pretty good. And now, like, I'm just curious, like, okay, like, what's the actual, like, like if we did all the things, how good will it get? So I think there is research left to be done. And there are a lot of ideas floating in the world now, everyone's sort of working on this. But I don't know that the current set of ideas is even final, like it will keep moving, I think.

29:58Let's transition to talking about magic. Maybe just what is magic? You've been very mysterious today, so maybe just share a little bit about what you're building. Yeah, I mean, we're trying to automate software engineering. It took us a while to figure out how to train supergiant models. It's a pretty interesting engineering challenge. I mean, fundamentally, we're trying to automate software engineering from the product side. And a subset of that is a model that can build HDI. Because if it's a great software engineer, then it should be able to do everyone's job at magic. If we can do everyone else's job, that would be a subset.

30:34that. So the idea is that you could use this to recursively improve alignment as well as models themselves in a way that isn't bottlenecked by human resources. And you know, there aren't that many non -shazirs in the world. If I had a non -shazier in my computer, I could spin up a million of them and maybe alignment would just be solved. I'm just like simplifying a ton and very idealistic in the statement. I'm happy to turn this whole thing into a scalable oversight podcast if you'd like, but the core idea is like, okay, like if I could just clone what we are doing into a computer and then press yes on the money button to run a cluster, to do the work we would be doing next week, that would be phenomenal.

31:23So I think we sort of were pursuing these two things in tandem where we want to ship something that's a good AI software engineer for people to use. I think one of the going to be one of the first domains to see high levels of automation. I don't like talking around. I don't think the whole assistant pitch is going to last very long once these models are good enough to automate. There's just no way the economy is not going to do that. I think everyone knows this and they're just like, they just don't like talking about it. It's totally fine. We use the olby farmers. We're not farmers. We're fine.

31:51Everyone prefers this. We'll figure out how we out in the economy if it produces the same or were stuff with less inputs. Like we should be able to figure that out. That's not a hard problem in like from the economic principles. You just have to figure out distribution. Anyway, but that's what we're trying to do. We're trying to automate software engineering and as a part of that, automate ourselves in doing the work we want to do. And so the reason they go out for software engineering them is that is the kind of lever that allows you to augment everything else. It's like the MVP of AGI, right?

32:20Like the minimum viable AGI. Yeah. Because that creates everything else. Yeah. Yeah, we wouldn't train something like Sora. Sora is great. You know, fantastic, generic videos. Awesome. It's just not interesting from an AGI perspective, if you believe that models can code themselves soon. Totally. And so out of all the companies that are trying to build an AI software engineer, you were probably the only one that is really taking a very clean and graded approach and training your own models. And that is either insanely brave or insanely crazy and probably a combination of both. I'm curious, I know you love training models, and so I know that's part of it.

33:01But why do you think you need to own the model to get this right? And how do you motivate yourself in the kind of the David versus Goliath of knowing that OpenAI exists and has great people and cares about coding and is great at building models, obviously? How do you think about that entire dynamic? I think you need, well, to build the best model, you need to build the model. And we want to solve these fundamental problems. You can't rely on any, like if the API guys solved it, then what the hell are you, you know, we might as well start the company three years later. And it goes to the point where we started, right?

33:33We started working on this stuff two years ago. So we have, you know, it took us some time to learn how to train these large models. Like it was really, I think it took opening I two years to get from GPT -3 to GPT -4 as well. And I thought we could be like much faster and this is gonna be great. It's a pain. So it's definitely an engineering challenge, but it's necessary. Like it's not like we're doing it just because it's fun or because I like training models. It's a massive financial investment that people trust us with. That it's not like it's one of those like one -to -one ROI investment.

34:11It's like if it works, if it works, it's fantastic. And if it doesn't work, the GPU's ran, and the money is gone. So, like, you're getting a lot of people's trust doing that. It's certainly not something you should do just because it's fun and you enjoy it. Fundamentally, I think the value will accrue at both at the AGI and at the hardware level and never at the application level. There's no incentive at all to offer an API. If the API creates a $100 billion company, you will just build that company internally. And if OpenAI doesn't someone else will, it's just incredibly unimaginable. to me that that would be how you would build these companies in the first place.

34:50So from a business perspective, I don't think that's necessarily the right way. Maybe there's some partnership potentials. You could like, oh, we'll get like special access or whatever. And then we're like, I've some, I don't know. Like cloud computing, right? Like there's, there's been many $10 billion. I mean, it's much, much, much harder to build Netflix and Airbnb and Uber than it is to build a chat interface. Like fundamentally, magic is an application you press download on it. We have a couple guys working on and it's just there. Like it's not, you know, you can build this with like YC pre -seed money.

35:24The modes in, I guess I can just make the API twice is expensive for the next model and then launch my own product and then under cut every, it's really fucked to not own the model. And in this domain. And then any domain that's going to generate a ton of revenue for a single company. In the case where it's distributed, maybe it's fine, but I don't think this will be. So it's necessary both for the market, which is good for us because the market is incentivized to fund folks like us, which it isn't in other domains, like a fund writing like an email assistant. You're not gonna get that funded anymore.

35:57So that's helpful, but fundamentally, the reason we train our own models is because it's necessary for our mission. And I just wouldn't be interested in building like a nice little SaaS wrapper. It's just not like, that's not gonna happen anyway. And I think though about competing against the 800 pound grillas, like you've raised all of money, but some people have raised boat loads. Yeah, there is a lot more money to do. Oh, and some people have a hundred million plus in revenue a year that they could spend as well. It goes beyond even their wants to could raise. Yeah, absolutely. And so how do you motivate yourself to compete in that, in that, you know, reality?

36:35The question is how much does it cost to build AGI and not how much money can you raise? because if you can build AGI for however much you can raise and you're having more might help you but it won't get you there substantially sooner. Like if you have all the right ideas and you get you can build it with a certain amount of hardware like by definition. Okay if someone had like a hundred times more hardware would it be like computing that much faster or whatever but it doesn't seem like a material advantage if you're estimate for how much compute you need to build AGI is not as far as the revenue these companies get generated.

37:06the funding they raise is in fact much lower. So, and I think that is the case. So, it's not by any means accessible. It's very that hard to get that much money, but it's not a hundred billion. It's, if I'm wrong, I'm wrong, and it'll be a hundred billion and we will not have a hundred billion, and that's it. But if we can get to that point where we have AGI and a couple others have AGI, and then the sort of, the benefit of additional computers there at you show in ROI, it's like a reasonably even playing field in terms of additional revenue. You're just gonna bring AGI to the market. You're gonna raise more on it.

37:42So the starting conditions of half this hardware is, and you need sufficient hardware, but you don't need more than sufficient. And so that's a bet that's not a, you don't know, but I think it's a bet with the high enough probability of being right that it is reasonable to compete in the space. And I think it is actually, it is reasonable to think that like the ROI of having quote unquote sufficient funding might be better than the ROI of having like infinite funding early on. And is there like an ideal for investors? That is not for me, for investors. Is there an ideal like team size for researchers?

38:22Is there a certain point at which you reach kind of like diminishing marginal returns of adding on the research? Yeah. So one of my biggest weaknesses, especially early on at Magic, was just scaling the team effectively. like we were very single threaded on a very small number of people doing basically all the work and I think we're getting better at that now. It's also you just need a certain level of maturity of your code base and of your research ideas and everything to properly segment them. So early on I would have said five for that time. Now I would say closer to 20 and I'm not including like folks working on other stuff.

38:59I'm including folks working on like the models and everything. I'd say closer to 20. I could imagine that in a few months I'll say at a slightly larger number, especially when you get into large -scale deployment, you really want to have a very, very good processes around just having high reliability, availability of services that are attached from each other, etc., etc. So then you can segment even more, which obviously stuff we're working on now.

39:27But it sort of grows our time. I don't see it ever exceeding like the tens of people. And right now it's in the low tens, very low tens. But I don't know maybe, it's actually it's a skill to be able to utilize. If you're able to utilize 200 people, you're just a better CEO than I am. The, no seriously, if you can, it's a good skill. And I think part of why I say a smaller number for us is that there is a ton of stuff we just don't do Like where if we built a video model that would just be a separate team They built a video model and like you know that's that's more scaling So so to an extent we're more focused and that's why we're smaller but Also if we could double the team and be twice as fast.

40:07That's I mean I would do it any day back in Or was it late 2022 when I first met you at the time like it was marketing assistance and email assistance since we're all the rage. And you were the first pitch that I heard that was AI that feels like a colleague. And I just remember that really sticking in my brain. So in some sense, you've been thinking about, it's kind of like agents to use a buzzword longer than anyone else. Maybe sharing your vision for that and like what you think it takes to build a great agent. Fundamentally, there are two tiers here. I guess three. What is useless? The next is a system that you have to micromanage.

40:44and then the next is the thing that manages you basically, where it's sort of more like a colleague. You could, I think the layer where it's exactly even, doesn't really exist, because it's sort of this little thin point. Once the model is more competent than you are, you are there to give it guidance on what you want to be accomplished and answer clarification questions, but you'll never have to tell it, like, here's a bug. I'm not saying that this is the one of everything. I'm not saying this is the one of our product, but fundamentally that has to be the goal. The way I feel when I talk to my best engineer, that's how I want to feel when I talk to magic, where we have a discussion.

41:30He's almost always right, and then he just writes the code, and then he someone else reveals it, and then it works. like that experience where my job is exclusively saying like here's kind of what I want and then they help clarify even right like I just want to hear specifically that that you it should feel like that and everything else doesn't matter to the user like what tools Dacia uses how it works does it run locally in the cloud does it need a VM does it have a browser I don't care doesn't fucking matter our problem not your problem you care about your problems getting sold so So fundamentally that's what I think matters to customers and to everything else is dependent on the exact product shape, exact domain, except everything.

42:14And like I'm stubborn as fuck, I just don't want to launch anything that isn't that. We will probably have to, but I just really want to get that thing. Like I want to talk to my computer, go and have lunch and come back and it built AGI. Like that's the end goal, right? And they'll be checkpoints, but yeah. I don't think everything else matters. How you accomplished that is up to an individual company. Yeah, how far away do you think we are from that? I guess maybe break it down into a little bit more. We met in 2022. You learned how to extrapolate Eric's timelines. So maybe one and a half or double everything I say, but I think very soon, like, very small number of years.

43:03I don't want to give a number now, but very small number. Less than 10. Oh, definitely less than 10. I mean, wait less. Wow. OK. Because I'm seeing some of the, like, the sweet agent stuff that just came out. They're like 14 % on sweet bench, which feels like. I mean, 14%. I just don't care about 14%. Like, I mean, we, like, I don't know, if 80 or 90 is good enough. Yep. Like, I think you need 99. Even 96, I don't trust my computer. Like, I don't want to review the code. If I have to, the tier of product where I have to review the code is fundamentally different from the tier of product where I don't have to review and understand the code.

43:40And like, you're not talking about 95 when you don't want to review. You're talking about 99 points something. You're talking about whatever my developers accomplish. Plus some, same as with self -driving cars. So the difference with self -driving cars is like, you die if the crash is. And here you just have to review code. So it's lo and shareable before, but fundamentally you need way, way, way more. And like usually the last few, right? Like the the nines are hard to get. So yeah, but no, I think in, I don't know, I don't. People have, I mean models have surpassed all these benchmarks. I mean, just recently, they're a math benchmark, right?

44:13Like way faster than even like for prediction markets assumed. And like, I don't see that stopping. There's just too much. Like if everyone was stuck and I got realist to some perception in the public, like that, all our GPT -4 is like, oh, I'm the next stop getting much better. Okay, no. Okay, we're gonna close that with a lightning round. One more of the answers. One, what's your favorite AI app, not magic? Probably only invisible ones still, like my spam filter and all this stuff. I love that. The things that keep life working, I think are still at the moment more useful than the sort of AGI -like apps.

44:47Because if you took them away, like life would just be awful. Like recommendation algorithms. for whatever. I think that's really useful. Other than that, yeah, I think whichever you're saying other than the, let's say, other than the programming world, other than magic. I don't actually know. But I'd say whichever model is currently like best, what kind of, it's a very boring answer, but I actually picked this path address, etc. I'd recommendation services first. What paper has been most influential to you? I don't think this paper is relevant at all in the world anymore, but it was the first paper I ever tried to deeply understand.

45:28Like, there's been months on it and re -implement it and everything. And so it was most influential to me as a person, not so much to my current work. And the paper's called DeepStack. It's one of those neural networks plus in perfect information game -solving papers. It's reasonably complex for time at, yeah, so it's for folks who are interested. It's like no one years so to now, but the editors just did rather than type of algorithm. But at the current time, at the right now, in fact, then it was useful. That was very interesting for me because it was just my first touch point with research.

46:05Really, I had no idea how to do research at all, and then that sort of was just, I'm going to dig into this, the way people like Hyperlink spam would be pdf, where you'd like, Revit hole, I did that with this paper. I love that. Okay, that's going to be my weekend reading. Last question. What are you most excited about in AI in the next one, five, and ten years? Just what it's going to, how society is going to integrate with it? I think that's, we're getting to the point now where it's really going to impact, we're in the next one to five years, it's really going to impact how society does stuff and beyond just another tab in your browser that speeds you up by some percentage on some passes.

46:46I think it'll get much more significant in that time frame. And ultimately, you should like the only, I am not one of the intrinsic curiosity type of people. I know most researchers are. I really am not. I just care about the outcome. And that is the outcome. So I'm most excited for the outcome. Eric, thank you for joining us again. Last time I recorded the podcast, we weren't actually able to talk about the thing that got us so excited about magic, which was you had shared with us your long -context eVal, and our own AI researchers had gotten really excited by what you had accomplished on that.

47:21That was what led to us investing in magic in the first place. You just made some exciting new announcements around the eVal. I was hoping you could share it with our audience. Yeah, for sure. Thank you so much. We've been running around with this Hashes eVal for a while. basically just being frustrated by the need for the new site eVal, and you know, everyone keeps complaining about it, and now that we've decided to announce our, where we're currently at in terms of our context work, instead of just like, you know, blah, blah, talking about like all we have so many tokens of context, it felt reasonable to share the eVal as well.

47:59I mean, we've used it in our fundraising obviously, and thanks for backing us, and generally just used it to guide our architecture development and our research. So yeah, it felt right to open source it and that others compare the architectures and their results with ours. And then, you know, we'd be, yeah, so it's exciting to share. And thank you for having back on the talk about it. Yep. Thank you. Can you say word on what's broken about Needle and Haystack and what Urival does differently? Yeah, for sure. You know, with Needle and Haystack, basically, what you're testing is like, find this weird thing, the needle in this giant pool of not weird stuff, the haystack.

48:40And so really all you need to be able to do to do this is to sort of take like a little backpack and walk from the start of the context to the end of the context and like find the weird thing, put it in your backpack and return it. If you have to sort of implicit prior that there is this thing is weird, so you're more likely to remember it. Which sort of means that you actually don't need to remember the whole context window. you don't need to know a little bit. So that allows some levels to look like they're doing long context really well, when really it's not working as well. So we decided to just go the complete opposite, super hard curmud, and just replace everything with random noise.

49:22There is no semantic information at all, because it's just randomly generated letters, basically, just hashes. And if you did something like needle in the he stack in a pool of passes, you really have to know the whole thing. But then what we do is we also do a hop. So it's not just you find this one thing, but you find this one thing, and then you find another thing. As always, you can keep that going. But those two dimensions, I think, really are important quantitative components of context. There are other things you can measure much better in more domain -specific evils. Of course, we care a lot about code, and so we look a lot at that, too, internally.

50:01But I think from a general purpose context evaluation perspective and the reason we chose to open source this evil and only this evil is just that I think this quantifies exactly what you want to measure when you think about long context and everything else is sort of the main specific Yeah, you want to be you don't you want to be forced to remember the whole context window when you're talking about context window Yeah, is it really that big like yeah, totally I remember our own researchers were just blown away by the purity of the of the Eval and how well done it was. And so thank you for what you're doing.

50:36And thank you for open sourcing it, especially the age where long context is becoming more and more important. Congratulations. We love arriving in your web. Yep, cheers. Of course. Thanks, Eric.

50:52Thanks a lot, and I think we will see you back.

From the publisher

There’s a new archetype in Silicon Valley, the AI researcher turned founder. Instead of tinkering in a garage they write papers that earn them the right to collaborate with cutting-edge labs until they break out and start their own.

This is the story of wunderkind Eric Steinberger, the founder and CEO of Magic.dev. Eric came to programming through his obsession with AI and caught the attention of DeepMind researchers as a high school student. In 2022 he realized that AGI was closer than he had previously thought and started Magic to automate the software engineering necessary to get there. Among his counterintuitive ideas are the need to train proprietary large models, that value will not accrue in the application layer and that the best agents will manage themselves. Eric also talks about Magic’s recent 100M token context window model and the HashHop eval they’re open sourcing.

Hosted by: Sonya Huang, Sequoia Capital

Mentioned in this episode:

David Silver: DeepMind researcher that led the AlphaGo team

Johannes Heinrich: a PhD student of Silver’s and DeepMind researcher who mentored Eric as a highschooler

Reinforcement Learning from Self-Play in Imperfect-Information Games: Johannes’s dissertation that inspired Eric 

Noam Brown: DeepMind, Meta and now OpenAI reinforcement learning researcher who eventually collaborated with Eric and brought him to FAIR

ClimateScience: NGO that Eric co-founded in 2019 while a university student 

Noam Shazeer: One of the original Transformers researchers at Google and founder of Charater.ai 

DeepStack: Expert-Level Artificial Intelligence in Heads-Up No-Limit Poker: the first AI paper Eric ever tried to deeply understand

LTM-2-mini: Magic’s first 100M token context model, build using the HashHop eval (now available open source)

00:00 - Introduction
01:39 - Vienna-born wunderkind
04:56 - Working with Noam Brown
8:00 - “I can do two things. I cannot do three.”
10:37 - AGI to-do list
13:27 - Advice for young researchers
20:35 - Reading every paper voraciously
23:06 - The army of Noams
26:46 - The leaps still needed in research
29:59 - What is Magic?
36:12 - Competing against the 800-pound gorillas
38:21 - Ideal team size for researchers
40:10 - AI that feels like a colleague
44:30 - Lightning round
47:50 - Bonus round: 200M token context announcement

More from Training Data

All 110 episodes
Founder Eric Steinberger on Magic’s Counterintuitive Approach to Pursuing AGI Training Data · 51 min
Listen in VO